ModelRefs / Document Intelligence — Canonical Workflow
Document Intelligence — Canonical Workflow
Canonical Document Intelligence workflow: OCR, multimodal parsing, structured extraction and validation.
Overview
Document intelligence turns PDFs, scans, and forms into structured data by chaining OCR, layout parsing, multimodal understanding, and schema-validated extraction, a pattern often paired with enterprise search and data-analysis workflows downstream, and increasingly relied on to remove manual data entry from back-office processes.
Use this page to check which multimodal models handle your document types and languages, which managed-container or self-hosted extraction architecture fits your volume and latency needs, and what evidence exists for accuracy on layouts, handwriting, and scan quality similar to yours, including scanned or photographed documents rather than clean digital PDFs, and how confidence scoring is surfaced downstream.
Extraction accuracy depends heavily on document quality, layout complexity, language, and schema design — this is provisional decision support, not a guarantee of accuracy on your specific document set. Validate against a representative sample, including edge-case scans and multi-page forms, and keep a human-review checkpoint for low-confidence extractions before they reach downstream systems.
Implementation profile
| Category | multimodal-models |
|---|---|
| Implementation maturity | production |
| Evidence status | partial |
| Primary use cases | ocr, extraction |
| Deployment options | managed-api, self-hosted |
| Architectures | managed-container, serverless-api |
Candidate models with published references
- BGE-M3
- GPT-5
- GPT-5 Mini
- Claude Opus 4
- Llama 4 Scout
- DeepSeek R1
- Mistral Large 2
- Command R+
- o3
- o4 Mini
- Text Embedding 3 Large
- Claude Sonnet 4
Coverage means the model is a candidate worth evaluating for this workflow, not a ranking or a recommendation. Models whose reference pages are still in review are omitted.
Benchmarks relevant to this workflow
miracl, mkqa, mldr, swe-bench, aider-polyglot, gpqa, aime-2025, tau-bench, browsecomp-long-context, longfact-concepts, terminal-bench, mmmu, mmlu-pro, livecodebench.
Relevance is a coverage signal from the canonical registry. Each benchmark only describes its own protocol and date, so confirm the harness matches your workload before treating a score as evidence.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Document Intelligence — Canonical Workflow.