ModelRefs / DocVQA — AI Glossary

DocVQA — AI Glossary

A visual question answering benchmark testing information extraction and reasoning over scanned documents and forms.

Overview

DocVQA (Mathew et al. 2021) contains 12,767 Q&A pairs over 5,188 document images: invoices, letters, forms, and reports. Tests text localization, reading, and understanding of structured and semi-structured documents. Frontier VLMs score 90–95% ANLS; challenging for models that must localize text in complex layouts.

Reference details

Topicevaluation
Last reviewed2026-06-24

Primary source

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to DocVQA — AI Glossary.

Frequently asked questions

What is DocVQA?

A visual question answering benchmark testing information extraction and reasoning over scanned documents and forms.

What concepts are related to DocVQA?

Closely related concepts include vqa, ocr, chartqa.