ModelRefs / ARC-AGI Leaderboard — AI Model Scores

ARC-AGI Leaderboard — AI Model Scores

Abstraction and Reasoning Corpus — visual grid puzzles testing fluid intelligence. Current leaders, methodology, and citation sources for ARC-AGI.

Overview

Abstraction and Reasoning Corpus — visual grid puzzles testing fluid intelligence.

How it is measured: Pass@2 on private evaluation set; very low ceiling for current LLMs.

What this benchmark measures

  • few-shot abstraction
  • novel task induction
  • grid transformation reasoning

Relevant to:

  • fluid-reasoning research
  • abstraction evaluation portfolios

Failure modes it exercises:

  • failure to infer latent rules from sparse demonstrations

Method and its limits

Infer a transformation rule from a small set of colored-grid input/output examples and produce exact output grids for held-out tasks.

  • Exact grid tasks cover a narrow abstraction format and do not directly test language, tools, domain knowledge, or deployment behavior.
  • Scores depend on task split, attempt budget, program synthesis or test-time adaptation, and submission protocol.

Dataset

Dataset
Abstraction and Reasoning Corpus (ARC-AGI-1)
Type
few-shot colored-grid transformation tasks
Freshness
aging

The original ARC-AGI-1 corpus uses training, evaluation, and hidden test tasks designed around core knowledge priors; later ARC Prize competition protocols and ARC-AGI-2/3 are separate scopes.

How to read this score

Treat a score as system-plus-protocol performance on the specified split, not a universal intelligence measure.

Similarity to real tasks: Low — ARC tasks isolate abstract grid transformations and intentionally minimize language and specialized world knowledge.

Data contamination risk: Medium — Public training and evaluation tasks create exposure risk, while hidden test tasks reduce direct test-set exposure. Neither condition proves absence of model-specific training overlap.

Benchmark gaming risk: High — Competition scaffolding, test-time compute, task-specific solvers, attempt budgets, and per-task adaptation can materially affect results.

What you still need to test yourself

  • Test language, domain knowledge, tool use, uncertainty, and real implementation tasks separately.
  • Record scaffolding, compute, attempts, task split, and any adaptation or external data.

This benchmark supports decisions about:

  • Compare abstraction systems under the same ARC-AGI split, attempt budget, and scoring protocol.
  • Study task-level failure patterns for few-shot rule induction.

Limitations

  • ARC-AGI does not test broad language, factual knowledge, safety, cost, or production reliability.
  • System-level results are not directly attributable to the base model when scaffolding or test-time adaptation differs.
  • No model-specific contamination or training-overlap audit is inferred from hidden tests; canonical scores must encode scaffolding, compute, and attempt-budget differences.

Sources re-reviewed 2026-07-11. ARC-AGI-1 is treated as an aging 2019 benchmark; each result must pin ARC version, split, attempt budget, and competition or evaluation protocol.

Sources

How this benchmark is scored

Categoryreasoning
Maximum score100 % solved
DirectionHigher is better

Primary source: https://arcprize.org/

Published results

ModelScoreRun dateSource
GPT-532.42026-05-01Aggregated public reports
Claude Opus 424.52026-05-01Aggregated public reports
DeepSeek R1212026-05-01Aggregated public reports

Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to ARC-AGI Leaderboard — AI Model Scores.