Overview of YomiToku capabilities
mainYomiToku is a Document AI engine specialized in Japanese document image analysis. It provides full OCR (optical character recognition) and layout analysis to recognize, extract, and convert text and diagrams from images.
Key Features:
- Specialized AI Models: Uses four independent models for text detection, text recognition, layout analysis, and table structure recognition, all optimized for Japanese documents.
- Japanese Language Support: Supports over 7,000 Japanese characters, including vertical text and unique layout structures. It also supports English documents.
- Semantic Extraction: Leverages layout analysis, table structure parsing, and reading order estimation to preserve the semantic structure of the document.
- Versatile Output: Supports conversion to HTML, Markdown, JSON, CSV, and text-searchable PDFs. It can also extract diagrams and images.
- Hardware Efficiency: Optimized for GPU environments (requires < 8GB VRAM). It also features an efficient mode for fast inference on CPUs.