Core features of ExtractThinker
mainExtractThinker provides several native capabilities for Document Intelligence Processing (IDP):
- Extraction with Pydantic: Use Pydantic models for structured data extraction, validation, and prompt engineering.
- Classification & Split: Intelligent document classification and splitting with support for consensus strategies, eager/lazy splitting, and confidence thresholds.
- PII Detection: Automatic detection and handling of sensitive personal information with a privacy-first approach.
- LLM and OCR Agnostic: Ability to switch between different LLM providers and OCR engines based on cost and performance requirements.