ModelRefs / ChartQA Leaderboard — AI Model Scores

ChartQA Leaderboard — AI Model Scores

Question answering over real-world charts. Current leaders, methodology, and citation sources for ChartQA.

Overview

Question answering over real-world charts.

How it is measured: Relaxed accuracy across human + augmented splits.

What this benchmark measures

  • chart visual understanding
  • chart question answering
  • logical and arithmetic reasoning over chart data

Relevant to:

  • chart question-answering screening
  • document and dashboard analysis evaluation
  • visual data extraction portfolios
  • financial reporting and variance-chart failure-case screening

Failure modes it exercises:

  • misreading chart labels or values
  • incorrect arithmetic or logical operations over chart data
  • failure across human-authored and generated questions

Method and its limits

Evaluate question answering over chart images using human-authored and machine-generated question sets, with visual, tabular, logical, and arithmetic reasoning under the pinned ChartQA split and metric.

  • Human-authored and generated question subsets have different construction and difficulty characteristics and should be reported separately when available.
  • Relaxed accuracy, answer normalization, OCR, chart rendering, and access to underlying tables can materially affect results.
  • Repository annotations include acknowledged noise, and the chart-type distribution does not cover every real dashboard or visualization.

Dataset

Dataset
ChartQA
Type
chart images with human-authored and generated visual question-answer pairs
Freshness
aging

The paper reports 9.6K human-written and 23.1K generated questions. The official repository separates human and augmented splits and documents chart annotations and known noise.

How to read this score

Inspect human and augmented splits, chart types, answer categories, OCR conditions, and error classes. Aggregate relaxed accuracy is not proof of reliable chart interpretation.

Similarity to real tasks: Medium — The benchmark uses chart images and questions relevant to visual analytics, but a controlled single-chart answer does not reproduce messy dashboards, provenance checks, multi-chart synthesis, or consequential business decisions.

Data contamination risk: Unknown — The dataset and charts are public; no model-specific training-overlap audit is registered.

Benchmark gaming risk: Unknown — Benchmark-specific OCR, table extraction, answer normalization, and tuning can affect performance; no run-level gaming assessment is registered.

What you still need to test yourself

  • Test the target dashboard and report formats, OCR quality, accessibility variants, chart types, units, annotations, and source provenance.
  • Measure extraction accuracy, calculation accuracy, abstention, citation to chart elements, latency, and human-review outcomes separately.
  • Use reconciled internal reports and charts to test materiality, period/entity selection, source-of-truth links, and reviewer correction.

This benchmark supports decisions about:

  • Screen multimodal candidates for chart reading and bounded chart reasoning under a matched ChartQA protocol.
  • Identify chart types, visual features, and arithmetic operations that require deeper workload testing.
  • Expose chart-label, unit, legend, OCR, and arithmetic risks that finance-report and variance-analysis test sets should cover.

Limitations

  • ChartQA does not establish reliable decision-making from enterprise dashboards or high-stakes charts.
  • A high aggregate score can hide failures on uncommon chart types, OCR, units, legends, visual clutter, or multi-step calculations.
  • It does not establish financial-statement accuracy, variance causality, materiality, control effectiveness, or audit evidence.

Sources reviewed 2026-07-02. Dataset freshness remains unknown; repository update dates are not treated as chart-content freshness dates.

Sources

How this benchmark is scored

Categoryvision
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://github.com/vis-nlp/ChartQA

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to ChartQA Leaderboard — AI Model Scores.