Overview of Qwen3-TTS capabilities
mainQwen3-TTS is a series of speech generation models supporting 10 languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian). It features:
- Powerful Speech Representation: Uses
Qwen3-TTS-Tokenizer-12Hzfor efficient acoustic compression and high-fidelity reconstruction. - Universal End-to-End Architecture: A discrete multi-codebook LM architecture that avoids information bottlenecks found in traditional LM+DiT schemes.
- Low-Latency Streaming: Supports both streaming and non-streaming generation via a Dual-Track hybrid architecture, with end-to-end latency as low as 97ms.
- Intelligent Voice Control: Allows control over timbre, emotion, and prosody using natural language instructions.