ModelRefs / MathVista Leaderboard — AI Model Scores
MathVista Leaderboard — AI Model Scores
Math reasoning over visual contexts (charts, figures, geometry). Current leaders, methodology, and citation sources for MathVista.
Overview
Math reasoning over visual contexts (charts, figures, geometry).
How it is measured: TestMini split; mean accuracy.
What this benchmark measures
- visual perception
- mathematical reasoning over visual contexts
- structured answer generation
Relevant to:
- visual-math evaluation portfolios
- chart and figure reasoning screening for finance-document workflows
Failure modes it exercises:
- incorrect visual interpretation
- incorrect mathematical reasoning
- answer-extraction failure
Method and its limits
Accuracy on visual mathematical questions drawn from 28 existing multimodal datasets and three newly created datasets, with result interpretation pinned to the dataset split, prompt, visual inputs, and answer-extraction protocol.
- The aggregate combines heterogeneous task types, skills, and source datasets, so overall accuracy can hide category-specific weaknesses.
- OCR, captioning, chain-of-thought, program tools, prompt format, and answer extraction can materially affect results.
Dataset
- Dataset
- MathVista
- Type
- visual mathematical reasoning questions across charts, plots, tables, diagrams, document images, puzzles, and scientific figures
- Freshness
- unknown
The primary sources describe 6,141 examples from 28 existing datasets and three new datasets; this does not establish coverage of current financial documents or accounting tasks.
How to read this score
Use category-level and problem-level results under a pinned protocol; do not infer financial-document, accounting, or production reliability from aggregate accuracy.
Similarity to real tasks: Medium — The benchmark includes charts, plots, tables, diagrams, document images, and scientific figures, but does not reproduce an organization's financial records, controls, or review workflow.
Data contamination risk: Unknown — The benchmark consolidates public datasets and ModelRefs has no model-run-specific training-overlap assessment.
Benchmark gaming risk: Unknown — Prompting, auxiliary OCR or captions, tools, answer extraction, and benchmark-targeted tuning can affect results.
What you still need to test yourself
- Test the organization's actual invoices, reports, charts, tables, layouts, currencies, tax formats, and review rules.
- Measure field and figure accuracy, source traceability, abstention, reviewer correction, latency, cost, and control failures separately.
This benchmark supports decisions about:
- Screen multimodal systems for visual mathematical reasoning under matched prompts, inputs, tools, and answer extraction.
- Identify visual, OCR, chart, table, and reasoning failure categories that should appear in internal finance-document tests.
Limitations
- MathVista is not an accounting, compliance, audit, fraud, forecasting, or financial-control benchmark.
- Visual mathematical accuracy does not prove reliable extraction, source traceability, policy adherence, or human-review outcomes on target documents.
Sources reviewed 2026-07-02. Dataset freshness remains unknown; pin the dataset split, repository revision, prompt, auxiliary inputs, tools, and answer-extraction protocol.
Sources
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts MathVista authors / ICLR 2024 · accessed 2026-07-02
- MathVista repository MathVista authors · accessed 2026-07-02
How this benchmark is scored
| Category | multimodal |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
Primary source: https://mathvista.github.io/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MathVista Leaderboard — AI Model Scores.