ModelRefs / ARC-AGI Methodology — Methodology

ARC-AGI Methodology — Methodology

ARC-AGI tests abstract pattern-recognition on grid puzzles designed to resist memorization.

Overview

What it measures: Fluid intelligence — solving novel visual reasoning puzzles from a few examples.

How it works

  • Each task: 3–5 input/output grid pairs demonstrating a rule, then a test input.
  • Model must produce the correct output grid.
  • Tasks are private and rotated to prevent training contamination.

Strengths

  • Strong contamination resistance
  • Tests generalization, not recall

Limitations

  • Small eval set
  • Grid format is unusual for LLMs

Best use cases

  • AGI-progress measurement
  • Reasoning-model stress test

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to ARC-AGI Methodology — Methodology.

Frequently asked questions

What does ARC-AGI measure?

Fluid intelligence — solving novel visual reasoning puzzles from a few examples.

What are its main limitations?

Small eval set Grid format is unusual for LLMs

When should I use this benchmark?

AGI-progress measurement Reasoning-model stress test