Overview of GPT-SoVITS Features
mainGPT-SoVITS is a powerful few-shot voice conversion and speech synthesis WebUI. Key capabilities include:
- Zero-shot TTS: Instant text-to-speech conversion using only a 5-second audio sample.
- Few-shot TTS: Fine-tune models with as little as 1 minute of training data to improve similarity and realism.
- Cross-lingual Support: Inference in languages different from the training set, including English, Japanese, Korean, Cantonese, and Chinese.
- Integrated WebUI Tools: Includes voice/accompaniment separation, automatic training set segmentation, multi-language Automatic Speech Recognition (ASR) using Fun-ASR-Nano, SenseVoice, and FunASR, and text annotation tools to assist in dataset creation.