ModelRefs / AIME 2024 Leaderboard — AI Model Scores

AIME 2024 Leaderboard — AI Model Scores

American Invitational Math Examination — 15-problem olympiad set. Current leaders, methodology, and citation sources for AIME 2024.

Overview

American Invitational Math Examination — 15-problem olympiad set.

How it is measured: Pass@1 over the 2024 paper; integer answers in [0,999].

What this benchmark measures

  • competition mathematics
  • multi-step symbolic reasoning

Relevant to:

  • hard mathematical reasoning screening
  • reasoning-model evaluation portfolios

Failure modes it exercises:

  • incorrect exact-answer reasoning on bounded math problems

Method and its limits

Exact-match scoring on explicitly identified 2024 AIME I and/or II integer-answer problems; every report must state the form, prompt, sampling, and aggregation protocol.

  • The two forms contain only 30 problems in total, so each item materially changes percentage accuracy.
  • Reports differ on whether they use one or both forms, pass@1, majority vote, or repeated samples.

Dataset

Dataset
2024 American Invitational Mathematics Examination
Type
high-school invitational competition mathematics problems
Freshness
unknown

AIME I and II each contain 15 problems with integer answers from 000 through 999; ModelRefs does not infer a benchmark freshness cutoff.

How to read this score

Report solved items and uncertainty, not only a percentage; compare scores only when form and sampling protocols match.

Similarity to real tasks: Low — AIME problems are rigorous but small, contest-style, exact-answer tasks rather than production mathematics workflows.

Data contamination risk: Unknown — The problems and solutions are public, and no model-specific training-overlap assessment is registered.

Benchmark gaming risk: Unknown — Small public sets are vulnerable to prompt, sampling, and benchmark-specific optimization effects.

What you still need to test yourself

  • Test representative domain mathematics, proofs, tool use, numerical verification, and error recovery.
  • Use hidden or newly authored problems to assess contamination and prompt sensitivity.

This benchmark supports decisions about:

  • Screen exact-answer mathematical reasoning under a pinned 2024 form and protocol.
  • Inspect problem-level errors before selecting a model for further math evaluation.

Limitations

  • A small contest set does not establish broad mathematical competence or production reliability.
  • A percentage without the exact form and sampling protocol is not comparable.

Dataset freshness is not fully confirmed from available source coverage. Contest and publication dates are not treated as model-training cutoffs.

Sources

How this benchmark is scored

Categoryreasoning
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://artofproblemsolving.com/

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to AIME 2024 Leaderboard — AI Model Scores.