What is VibeVoice-ASR and its key features?
mainVibeVoice-ASR is a unified speech-to-text model designed for long-form audio processing (up to 60 minutes in a single pass). It provides structured transcriptions including:
- Who (Speaker): Diarization to identify different speakers.
- When (Timestamps): Precise timing for utterances.
- What (Content): The transcribed text.
Key Capabilities:
- Single-Pass 60-minute Processing: Handles continuous audio within a 64K token length, maintaining global context and speaker tracking.
- Customized Hotwords: Allows users to provide specific names or technical terms to improve recognition accuracy.
- Multilingual & Code-Switching: Supports over 50 languages natively without requiring explicit language settings and handles switching between languages within a single utterance.