ModelRefs / DocVQA Leaderboard — AI Model Scores
DocVQA Leaderboard — AI Model Scores
Document visual question answering on scanned business docs. Current leaders, methodology, and citation sources for DocVQA.
Overview
Document visual question answering on scanned business docs.
How it is measured: ANLS over test set.
What this benchmark measures
- document-image question answering
- OCR-dependent document understanding
- layout- and structure-sensitive answer extraction
Relevant to:
- document question-answering screening
- document-intelligence evaluation portfolios
- visual extraction workflow evaluation
- invoice, vendor-document, contract, and audit-evidence failure-case screening
- completion-readiness test design for invoice, AP, EHR, legal, and evidence-synthesis workflows
Failure modes it exercises:
- incorrect answers from document images
- failure to use document layout and structure
- OCR and answer-normalization errors
Method and its limits
Answer natural-language questions about document images using the pinned DocVQA task, split, and evaluation metric; report whether OCR, layout annotations, retrieval, or other auxiliary inputs are available.
- The original benchmark focuses on single-page document images and question answering rather than full document parsing or schema extraction.
- ANLS and accuracy are sensitive to answer normalization and do not expose every factual, layout, or provenance error.
- OCR system, image resolution, auxiliary text, task edition, and document-domain mix can materially change results.
Dataset
- Dataset
- DocVQA
- Type
- scanned document images with natural-language questions and short answers
- Freshness
- aging
The primary paper reports 50,000 questions over more than 12,000 document images. Later DocVQA challenges and multi-page variants are separate scopes and must not be mixed with the original task.
How to read this score
Read the score with task edition, split, metric, OCR condition, and document categories. Use document- and question-type errors rather than one aggregate score for workflow decisions.
Similarity to real tasks: Medium — The dataset uses scanned document images and human questions, but the original single-page task does not reproduce long document collections, custom schemas, handwriting, access controls, or downstream automation.
Data contamination risk: Unknown — The public dataset and challenge materials may be present in training corpora; no model-specific overlap audit is registered.
Benchmark gaming risk: Unknown — OCR choice, document-specific tuning, auxiliary text, answer normalization, and task-edition selection can affect results.
What you still need to test yourself
- Test representative document families, layouts, scan qualities, languages, tables, handwriting, multi-page context, and required output schemas.
- Measure field-level accuracy, source grounding, missing-field handling, confidence, human handoff, privacy controls, latency, and cost.
- Test actual invoice fields, vendor evidence, amendments, retention metadata, source reliability, validation rules, and reviewer corrections.
- Define deployment-specific acceptance, monitoring, escalation, rollback, and qualified-review gates before treating document performance as workflow evidence.
This benchmark supports decisions about:
- Screen multimodal systems for bounded question answering over document images.
- Identify layout, OCR, document-type, and answer-normalization failures that require workload-specific extraction tests.
- Inform invoice, vendor-intake, contract-review, and evidence-collection test cohorts without treating QA accuracy as workflow validation.
- Seed document cohorts and failure taxonomies for workflow acceptance testing without satisfying completion-readiness by itself.
Limitations
- DocVQA does not establish reliable schema extraction, multi-page reasoning, handwriting support, or downstream business-process safety.
- A question-answering score does not guarantee complete field recall, correct provenance, calibrated confidence, or privacy compliance.
- It does not establish invoice validity, vendor safety, contract meaning, audit sufficiency, control effectiveness, or permission to act.
- It cannot by itself make an invoice, AP, EHR, legal, or clinical workflow a completion candidate.
Sources reviewed 2026-07-02. Dataset freshness remains unknown; challenge news and site updates do not establish an original-corpus freshness cutoff.
Sources
- DocVQA DocVQA · accessed 2026-07-02
- DocVQA: A Dataset for VQA on Document Images DocVQA authors / WACV 2021 · accessed 2026-07-02
How this benchmark is scored
| Category | vision |
|---|---|
| Maximum score | 100 ANLS |
| Direction | Higher is better |
Primary source: https://www.docvqa.org/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to DocVQA Leaderboard — AI Model Scores.