ModelRefs / AGIEval Leaderboard — AI Model Scores
AGIEval Leaderboard — AI Model Scores
Human-centric standardized exams (SAT, GRE, GMAT, LSAT, civil service). Current leaders, methodology, and citation sources for AGIEval.
Overview
Human-centric standardized exams (SAT, GRE, GMAT, LSAT, civil service).
How it is measured: Zero/few-shot accuracy across English & Chinese exams.
What this benchmark measures
- academic and professional exam reasoning
- bilingual knowledge application
Relevant to:
- broad reasoning evaluation portfolios
- English and Chinese exam-task screening
Failure modes it exercises:
- knowledge and reasoning errors across standardized exam formats
Method and its limits
Zero- and few-shot evaluation over public admission, qualification, law, civil-service, and math-competition exams in English and Chinese.
- Aggregate accuracy can hide large differences by language, exam, prompt, and answer format.
- Public exam content creates model-specific training-overlap and contamination uncertainty.
Dataset
- Dataset
- AGIEval
- Type
- human standardized admission and qualification exam questions
- Freshness
- unknown
The paper describes 20 public exams spanning English and Chinese tasks; it is not a current-world or workplace-performance dataset.
How to read this score
Use matched prompts and exam-level breakdowns; do not treat one aggregate score as general intelligence or professional qualification.
Similarity to real tasks: Low — Tasks originate in human exams, but static questions and exact grading do not reproduce most open-ended production work.
Data contamination risk: Unknown — The source material is public and no model-specific training-overlap audit is registered.
Benchmark gaming risk: Unknown — No benchmark-specific optimization or prompt-tuning audit is registered.
What you still need to test yourself
- Test open-ended, current, domain-specific tasks with the target language and output format.
- Evaluate calibration, safety, cost, latency, and model-specific training overlap separately.
This benchmark supports decisions about:
- Compare broad exam-style reasoning under the same AGIEval task set and prompt protocol.
- Identify languages or exam domains requiring targeted evaluation.
Limitations
- Exam accuracy does not prove professional competence, tool use, safety, or production reliability.
- Cross-report scores may use different subsets, prompts, languages, and answer extraction.
Dataset freshness is not fully confirmed from available source coverage. The paper date is not treated as a knowledge cutoff.
Sources
- AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models AGIEval authors / Findings of NAACL 2024 · accessed 2026-06-29
- AGIEval code, data, and model outputs AGIEval authors · accessed 2026-06-29
How this benchmark is scored
| Category | reasoning |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
Primary source: https://arxiv.org/abs/2304.06364
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to AGIEval Leaderboard — AI Model Scores.