ModelRefs / AIME 2025 Leaderboard — AI Model Scores
AIME 2025 Leaderboard — AI Model Scores
The two 2025 AIME competition-mathematics forms used for exact-answer reasoning evaluation. Current leaders, methodology, and citation sources for AIME 2025.
Overview
The two 2025 AIME competition-mathematics forms used for exact-answer reasoning evaluation.
How it is measured: Exact-answer accuracy on AIME I, AIME II, or both; form coverage, tools, sampling, answer extraction, and aggregation must be reported per run.
What this benchmark measures
- competition mathematics
- multi-step exact-answer reasoning
Relevant to:
- recent mathematical reasoning screening
- reasoning-model evaluation portfolios
Failure modes it exercises:
- incorrect exact-answer reasoning on a recent bounded problem set
Method and its limits
Exact-match evaluation on explicitly identified 2025 AIME I and/or II problems; reports must state form, prompt, sampling, answer extraction, and aggregation.
- The combined set is only 30 problems and early reports may cover only AIME I.
- Release recency reduces some exposure pathways but does not prove absence from training, browsing, tools, or benchmark tuning.
- Provider results are not automatically comparable: disclosures vary across form coverage, tool access, reasoning effort, sampling, answer extraction, and parallel-selection or majority-vote aggregation.
Dataset
- Dataset
- 2025 American Invitational Mathematics Examination
- Type
- high-school invitational competition mathematics problems
- Freshness
- aging
Coverage is established for the complete 2025 competition artifact: AIME I was administered February 6 and AIME II February 12, 2025, with 15 integer-answer problems per form (30 distinct problems when both forms are combined). Reports must still identify whether they evaluated AIME I, AIME II, or both; February 12 records completion of the two-form administration, not a model-training cutoff, contamination clearance, or model-score evaluation date.
How to read this score
State the form and denominator, show problem-level outcomes, and avoid comparing one-form results with combined-form results.
Similarity to real tasks: Low — Established from the official competition format: AIME is a proctored, three-hour invitational examination with 15 difficult pre-calculus problems per form and an exact integer answer from 000 through 999 for each problem. The 2025 I and II forms therefore resemble a bounded human contest-mathematics task and provide a useful check of multi-step symbolic reasoning under exact-answer scoring. They do not reproduce most professional mathematics work: no proof or derivation is graded, and the task does not require requirements discovery, current evidence, domain software, numerical validation, peer review, uncertainty reporting, or a reusable technical deliverable. Real-task similarity is supported but low rather than unavailable.
Data contamination risk: High — The complete 30-problem set and answer material became publicly inspectable after the February 2025 administrations, and the small fixed set is now repeatedly used in model reports. This creates high exposure and memorization risk for post-release models and tool-enabled evaluations. The classification describes benchmark-level exposure; no model-specific training-overlap, browsing, tool-access, or post-release contamination audit is inferred.
Benchmark gaming risk: High — A 30-item exact-answer set is highly sensitive to form selection, prompt and answer parsing, reasoning budget, repeated sampling, and aggregation. Anthropic's Claude 4 disclosure, for example, separates standard nucleus-sampled results from higher parallel-test-time-compute results selected by an internal scoring model. Every comparison must therefore pin the evaluated form, tools, sampling, extraction, and aggregation protocol.
What you still need to test yourself
- Test hidden, domain-relevant, tool-assisted, proof, and numerical-verification tasks.
- Assess repeated-run variance, prompt sensitivity, leakage pathways, latency, and cost.
This benchmark supports decisions about:
- Screen recent contest-math performance under a pinned form and reproducible protocol.
- Identify failure patterns for deeper internal math evaluation.
Limitations
- Recency is not proof of uncontaminated evaluation.
- Small-set accuracy does not establish broad or production mathematical reliability.
Dataset publication coverage is confirmed through the February 12, 2025 AIME II administration and is now aging. These official competition dates establish when the two 2025 forms existed, but do not establish model-training isolation, post-release exposure, or the freshness of any reported model score.
Sources
- 2024-25 AIME Thresholds Are Available Mathematical Association of America · accessed 2026-08-11
- MAA Invitational Competitions — American Invitational Mathematics Examination Mathematical Association of America · accessed 2026-08-11
- AIME Problems and Solutions Art of Problem Solving · accessed 2026-06-29
- Introducing Claude 4 Anthropic · accessed 2026-08-11
How this benchmark is scored
| Category | reasoning |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
Primary source: https://maa.org/maa-invitational-competitions/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to AIME 2025 Leaderboard — AI Model Scores.