Overview of Fish Audio S2 Pro
mainFish Audio S2 Pro is a multimodal TTS model trained on over 10 million hours of audio data, supporting more than 80 languages.
Key Features
- Dual-Autoregressive (Dual-AR) Architecture: An innovative architecture for high-quality speech generation.
- Sub-word Level Control: Supports fine-grained control over tone and emotion using natural language tags (e.g.,
[whisper],[excited],[angry]). - Long Context Support: Native support for multi-speaker generation and multi-turn dialogues with very long contexts.
- RL Alignment: Uses Reinforcement Learning to improve speech naturalness and emotional depth.
- Streaming Performance: Optimized for extreme streaming performance via SGLang.
Model Variants
| Model | Size | Availability | Description |
|---|---|---|---|
| S2-Pro | 4B parameters | HuggingFace | Leading full-featured model with highest quality and stability |