Overview of VieNeu-TTS
mainVieNeu-TTS is a high-fidelity, on-device Vietnamese Text-to-Speech (TTS) model. The latest version, VieNeu-TTS-v2, features a bilingual (English-Vietnamese) architecture trained on over 10,000 hours of data.
Key capabilities include:
- Bilingual Code-switching: Smooth transitions between English and Vietnamese within a single sentence.
- Zero-shot Voice Cloning: Clone any voice using only 3-5 seconds of audio samples.
- Podcast & Dialogue Modes: Supports multi-speaker interactions with automatic character recognition.
- High Performance: Optimized for GPU (via LMDeploy) and CPU (via GGUF/ONNX).
- Offline Capability: Generates high-quality 24 kHz audio completely offline.