MiniCPM-V and MiniCPM-o Multimodal LLM Series

repository·main·Indexed 12 days ago

https://github.com/openbmb/minicpm-v

Efficient multimodal LLM series designed for high-performance vision, audio, and text understanding on mobile and edge devices. Includes models such as MiniCPM-V, MiniCPM-V-2, MiniCPM-Llama3-V-2_5, MiniCPM-V-2_6, and MiniCPM-o-2_6. Documentation covers inference and evaluation using the opencompass framework and vqaeval for datasets like TextVQA and DocVQA.

Tokens
58.3K
Snippets
129
Records
220
Agent score
96%

What's inside MiniCPM-V

  1. Overview of MiniCPM-V 2.6

    main

    MiniCPM-V 2.6 is an 8B parameter multimodal model built on SigLip-400M and Qwen2-7B. It is designed for high-performance visual understanding, including single-image, multi-image, and video comprehension.

    Key capabilities include:

    • High-Performance Single Image Understanding: Competes with proprietary models like GPT-4o mini and Claude 3.5 Sonnet.
    • Multi-Image & Context Learning: Supports multi-image dialogue and reasoning.
    • Video Understanding: Capable of processing video inputs for detailed temporal and spatial descriptions.
    • Advanced OCR: Handles arbitrary aspect ratios and up to 1.8 million pixels (e.g., 1344x1344).
    • Multilingual Support: Supports English, Chinese, German, French, Italian, Korean, and more.
    • Efficiency: Features high visual token density, requiring only 640 tokens for 1.8 million pixels, optimizing inference speed and memory usage.
  2. Overview of MiniCPM-V 4.0

    main

    MiniCPM-V 4.0 is an efficient multimodal model in the MiniCPM-V series. It is built upon SigLIP2-400M and MiniCPM4-3B, with a total of 4.1B parameters. It maintains strong capabilities in single-image, multi-image, and video understanding while significantly improving inference efficiency.

    Key features include:

    • Leading Visual Capabilities: High performance on OpenCompass (average 69.0), outperforming models like MiniCPM-V 2.6 and Qwen2.5-VL-3B-Instruct, and rivaling closed-source models like GPT-4.1-mini-20250414.
    • Exceptional Efficiency: Optimized for edge devices. It can run smoothly on an iPhone 16 Pro Max with a first token latency as low as 2 seconds and a decoding speed of 17.9 tokens/s.
    • Ease of Use: Supports various inference frameworks and deployment methods.
  3. Overview of MiniCPM-V 4.6

    main

    MiniCPM-V 4.6 is a high-efficiency multimodal model designed for edge deployment. It is built upon SigLIP2-400M and the Qwen3.5-0.8B LLM.

    Key features include:

    • High Efficiency: Uses a ViT internal visual token early compression mechanism to reduce computation by over 50% during the visual encoding stage. It supports 4x and 16x hybrid visual token compression rates.
    • Multimodal Capabilities: Strong performance in single-image, multi-image, and video understanding, comparable to Qwen3.5 2B levels on benchmarks like OpenCompass and OCRBench.
    • Edge Support: Optimized for deployment on iOS, Android, and HarmonyOS.
    • Framework Compatibility: Supports inference via SGLang, vLLM, llama.cpp, and Ollama, and fine-tuning via SWIFT and LLaMA-Factory.
    • Quantization: Available in GGUF, BNB, AWQ, and GPTQ formats.
  4. Overview of MiniCPM-V series

    main

    MiniCPM-V is a series of end-side multimodal Large Language Models (MLLMs) optimized for vision-language understanding. These models accept image, video, and text as inputs to produce high-quality text outputs. The series focuses on achieving strong performance while maintaining efficient deployment capabilities for end-side devices.

    Key models include:

    • MiniCPM-V 2.6: The latest 8B parameter model. It features superior token density, enabling real-time video understanding on devices like iPads. It is designed to compete with or surpass models like GPT-4V, GPT-4o mini, Gemini 1.5 Pro, and Claude 3.5 Sonnet in single image, multi-image, and video understanding tasks.
    • MiniCPM-Llama3-V 2.5: An 8B parameter model built on SigLip-400M and Llama3-8B-Instruct, offering significant improvements over version 2.0.
  5. Overview of MiniCPM-Llama3-V 2.5

    main
    MiniCPM-Llama3-V 2.5 is an 8B parameter multimodal model built on SigLip-400M and Llama3-8B-Instruct. It is designed for high-performance multimodal interaction, featuring strong OCR capabilities (handling images up to 1.8 million pixels), improved instruction-following, and reduced hallucination rates via the RLAIF-V method. It supports over 30 languages including German, French, Spanish, Italian, and Korean.
  6. Overview of MiniCPM-V 2.0

    main

    MiniCPM-V 2.0 is an efficient Large Multimodal Model (LMM) designed for high performance and deployment flexibility. It is built using SigLip-400M and MiniCPM-2.4B, connected via a perceiver resampler.

    Key Features

    • High Performance: Achieves state-of-the-art results among models under 7B parameters on benchmarks like OCRBench, TextVQA, and MME. It outperforms larger models like Qwen-VL-Chat (9.6B) and CogVLM-Chat (17.4B) on OpenCompass.
    • Strong OCR & Scene Understanding: Comparable to Gemini Pro in scene-text understanding and achieves state-of-the-art performance on OCRBench.
    • Trustworthy Behavior: Uses multimodal RLHF (via the RLHF-V technique) to align behavior and reduce hallucinations, matching GPT-4V's performance on Object HalBench.
    • High-Resolution Support: Supports images up to 1.8 million pixels (e.g., 1344x1344) at any aspect ratio using a technique from LLaVA-UHD.
    • Efficiency: Uses a perceiver resampler to compress visual tokens, making it suitable for GPUs, personal computers, and mobile devices.
    • Bilingual Support: Strong multimodal capabilities in both English and Chinese.
  7. Overview of MiniCPM-o 4.5

    main

    MiniCPM-o 4.5 is a high-performance, end-to-end multimodal model with 9B parameters. It is built upon SigLip2, Whisper-medium, CosyVoice2, and Qwen3-8B.

    Key capabilities include:

    • Advanced Vision: High-resolution image processing (up to 1.8M pixels) and high-frame-rate video support (up to 10fps). It excels in OCR and document parsing (OmniDocBench).
    • Natural Speech: Supports configurable Chinese-English bilingual real-time voice dialogue, including voice cloning and role-playing via reference audio.
    • Full-Duplex Interaction: Capable of simultaneous real-time video/audio input processing and synchronized text/speech output generation (the model can "see, hear, and speak" at once).
    • Active Interaction: The model can monitor streams and proactively initiate reminders or comments at a 1Hz decision frequency.
  8. Overview of MiniCPM-V 4.5

    main

    MiniCPM-V 4.5 is a high-performance multimodal large language model (MLLM) with 8B parameters, built on Qwen3-8B and SigLIP2-400M. It is designed to compete with or surpass much larger proprietary models like GPT-4o-latest and Gemini-2.0 Pro in vision-language tasks.

    Key Capabilities:

    • High-FPS & Long Video Understanding: Uses a unified 3D-Resampler to achieve up to 96x compression for video tokens, enabling high-FPS (up to 10FPS) and long video perception without increasing LLM inference costs.
    • Controllable Thinking Modes: Supports switchable Fast Thinking (for efficiency) and Deep Thinking (for complex reasoning) modes.
    • Advanced OCR & Document Parsing: Capable of processing high-resolution images (up to 1.8 million pixels) with any aspect ratio using 4x fewer visual tokens than typical MLLMs. It excels at OCR and PDF document parsing.
    • Multilingual Support: Supports more than 30 languages and features trustworthy behavior (reduced hallucinations).
  9. Overview of MiniCPM-o 2.6

    main

    MiniCPM-o 2.6 is an 8B parameter end-to-end multimodal model built on SigLip-400M, Whisper-medium-300M, ChatTTS-200M, and Qwen2.5-7B. It is designed for real-time voice dialogue and multimodal streaming interaction.

    Key capabilities include:

    • Leading Visual Intelligence: High performance in single-image understanding, multi-image, and video understanding, rivaling proprietary models like GPT-4o and Claude 3.5 Sonnet.
    • Advanced Speech Capabilities: Supports configurable bilingual (Chinese/English) real-time voice dialogue, including emotion/speed/style control, voice cloning, and role-playing.
    • Multimodal Streaming Interaction: Capable of accepting continuous video and audio streams for real-time interaction.
    • Strong OCR & Multilingual Support: Optimized for high-resolution images (up to 1.8 million pixels) and supports over 30 languages.
    • High Efficiency: Features high visual token density (processing 1.8M pixels with only 640 tokens), reducing latency and memory usage for edge devices like iPads.
  10. Overview of OmniLMM-12B

    main

    OmniLMM-12B is an early-release Large Multimodal Model (LMM) built on EVA02-5B and Zephyr-7B-β, utilizing a perceiver resampler layer. It is designed for high performance and trustworthy multimodal interaction.

    Key features include:

    • Strong Performance: Surpasses established LMMs on benchmarks like MME, MMBench, and SEED-Bench.
    • Trustworthy Behavior: Aligned via multimodal RLHF (using the RLHF-V technique) to reduce hallucinations. It ranks #1 on MMHal-Bench and outperforms GPT-4V on Object HalBench.
    • Real-time Multimodal Interaction: Can be integrated with text-only models (like GPT-3.5) to create assistants that process video streams from cameras and speech streams from microphones to emit speech output.

    Note: For better performance and efficiency, it is recommended to use the more recently released models found in the main README.

  11. Overview of MiniCPM-V and MiniCPM-o

    main

    MiniCPM-V and MiniCPM-o are series of multimodal large models designed for high performance and efficient deployment on edge devices (mobile platforms like iOS, Android, and HarmonyOS).

    • MiniCPM-V: Focuses on efficient visual-language understanding for image, video, and text inputs.
    • MiniCPM-o: Extends capabilities to real-time, end-to-end full-modality interaction, supporting streaming video/audio input and text/speech output.

    Key models include:

    • MiniCPM-V 4.6: A 1.3B parameter model optimized for efficiency. It uses ViT early compression technology (based on LLaVA-UHD v4) to reduce visual encoding overhead by 50% and supports 4x/16x mixed visual token compression rates to balance performance and efficiency. It is suitable for deployment on mainstream mobile platforms.
    • MiniCPM-o 4.5: A 9B parameter model designed for full-duplex multimodal real-time streaming interaction. It supports simultaneous streaming of output (speech/text) and input (video/audio), enabling real-time "see, hear, and speak" capabilities and proactive interactions like "active reminders."