Overview of the SALMONN model family
mainSALMONN is a suite of advanced multi-modal large language models (LLMs) designed for various sensory modalities including vision, speech, text, and action. Because the repository acts as a central hub, specific implementations and models are located in dedicated branches.
Key models in the family include:
- ELLSA: An end-to-end model unifying vision, speech, text, and action in a streaming full-duplex framework.
- video-SALMONN 2: An audio-visual LLM for high-quality audio-visual video captioning and general video QA.
- video-SALMONN-o1: A reasoning-enhanced audio-visual LLM.
- SALMONN for speech quality assessment: Specialized for automatic speech quality evaluation.
- video-SALMONN: Speech-enhanced audio-visual LLM.
- SALMONN (Original): Focused on generic hearing abilities for LLMs.