Overview of MonkeyOCR
mainMonkeyOCR is a document parsing system based on the Structure-Recognition-Relation (SRR) triplet paradigm. It is designed to simplify multi-tool pipelines while maintaining efficiency by avoiding the use of large multimodal models for full-page processing.
Key model variants include:
- MonkeyOCR-pro-1.2B: A leaner, faster version that outperforms the 3B version in accuracy, speed, and efficiency (specifically on Chinese documents).
- MonkeyOCR-pro-3B: A larger model that achieves high performance on English and Chinese documents, outperforming several closed-source and extra-large open-source VLMs on benchmarks like OmniDocBench.