Overview of SenseVoice Nano models
masterThe SenseVoice Nano models in this directory are converted versions of the Fun-ASR-Nano-2512 models from HuggingFace. These models are optimized for high-performance speech recognition in challenging environments and diverse linguistic scenarios.
Key Capabilities:
- Far-field & High-noise Recognition: Optimized for far-distance pickup and noisy environments (e.g., conference rooms, vehicles, industrial sites) with up to 93% accuracy.
- Chinese Dialects & Accents: Supports 7 major dialects (Wu, Cantonese, Min, Hakka, Gan, Xiang, Jin) and covers 26 regional accents (including Henan, Shaanxi, Hubei, Sichuan, Chongqing, Yunnan, Guizhou, Guangdong, Guangxi, etc.).
- Multi-language Support: Supports 31 languages, with specific optimization for East and Southeast Asian languages. It supports free language switching and mixed-language recognition.
- Music Background Recognition: Specifically enhanced to recognize lyrics accurately even when music is playing in the background.