ModelRefs / ARC-AGI Leaderboard — AI Model Scores
ARC-AGI Leaderboard — AI Model Scores
Abstraction and Reasoning Corpus — visual grid puzzles testing fluid intelligence. Current leaders, methodology, and citation sources for ARC-AGI.
Overview
Abstraction and Reasoning Corpus — visual grid puzzles testing fluid intelligence.
How it is measured: Pass@2 on private evaluation set; very low ceiling for current LLMs.
What this benchmark measures
- few-shot abstraction
- novel task induction
- grid transformation reasoning
Relevant to:
- fluid-reasoning research
- abstraction evaluation portfolios
Failure modes it exercises:
- failure to infer latent rules from sparse demonstrations
Method and its limits
Infer a transformation rule from a small set of colored-grid input/output examples and produce exact output grids for held-out tasks.
- Exact grid tasks cover a narrow abstraction format and do not directly test language, tools, domain knowledge, or deployment behavior.
- Scores depend on task split, attempt budget, program synthesis or test-time adaptation, and submission protocol.
Dataset
- Dataset
- Abstraction and Reasoning Corpus (ARC-AGI-1)
- Type
- few-shot colored-grid transformation tasks
- Freshness
- aging
The original ARC-AGI-1 corpus uses training, evaluation, and hidden test tasks designed around core knowledge priors; later ARC Prize competition protocols and ARC-AGI-2/3 are separate scopes.
How to read this score
Treat a score as system-plus-protocol performance on the specified split, not a universal intelligence measure.
Similarity to real tasks: Low — ARC tasks isolate abstract grid transformations and intentionally minimize language and specialized world knowledge.
Data contamination risk: Medium — Public training and evaluation tasks create exposure risk, while hidden test tasks reduce direct test-set exposure. Neither condition proves absence of model-specific training overlap.
Benchmark gaming risk: High — Competition scaffolding, test-time compute, task-specific solvers, attempt budgets, and per-task adaptation can materially affect results.
What you still need to test yourself
- Test language, domain knowledge, tool use, uncertainty, and real implementation tasks separately.
- Record scaffolding, compute, attempts, task split, and any adaptation or external data.
This benchmark supports decisions about:
- Compare abstraction systems under the same ARC-AGI split, attempt budget, and scoring protocol.
- Study task-level failure patterns for few-shot rule induction.
Limitations
- ARC-AGI does not test broad language, factual knowledge, safety, cost, or production reliability.
- System-level results are not directly attributable to the base model when scaffolding or test-time adaptation differs.
- No model-specific contamination or training-overlap audit is inferred from hidden tests; canonical scores must encode scaffolding, compute, and attempt-budget differences.
Sources re-reviewed 2026-07-11. ARC-AGI-1 is treated as an aging 2019 benchmark; each result must pin ARC version, split, attempt budget, and competition or evaluation protocol.
Sources
- On the Measure of Intelligence François Chollet · accessed 2026-07-11
- ARC-AGI repository ARC Prize Foundation / François Chollet · accessed 2026-07-11
- What is ARC-AGI? ARC Prize Foundation · accessed 2026-07-11
How this benchmark is scored
| Category | reasoning |
|---|---|
| Maximum score | 100 % solved |
| Direction | Higher is better |
Primary source: https://arcprize.org/
Published results
| Model | Score | Run date | Source |
|---|---|---|---|
| GPT-5 | 32.4 | 2026-05-01 | Aggregated public reports |
| Claude Opus 4 | 24.5 | 2026-05-01 | Aggregated public reports |
| DeepSeek R1 | 21 | 2026-05-01 | Aggregated public reports |
Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to ARC-AGI Leaderboard — AI Model Scores.