ModelRefs / AGIEval Leaderboard — AI Model Scores

AGIEval Leaderboard — AI Model Scores

Human-centric standardized exams (SAT, GRE, GMAT, LSAT, civil service). Current leaders, methodology, and citation sources for AGIEval.

Overview

Human-centric standardized exams (SAT, GRE, GMAT, LSAT, civil service).

How it is measured: Zero/few-shot accuracy across English & Chinese exams.

What this benchmark measures

  • academic and professional exam reasoning
  • bilingual knowledge application

Relevant to:

  • broad reasoning evaluation portfolios
  • English and Chinese exam-task screening

Failure modes it exercises:

  • knowledge and reasoning errors across standardized exam formats

Method and its limits

Zero- and few-shot evaluation over public admission, qualification, law, civil-service, and math-competition exams in English and Chinese.

  • Aggregate accuracy can hide large differences by language, exam, prompt, and answer format.
  • Public exam content creates model-specific training-overlap and contamination uncertainty.

Dataset

Dataset
AGIEval
Type
human standardized admission and qualification exam questions
Freshness
unknown

The paper describes 20 public exams spanning English and Chinese tasks; it is not a current-world or workplace-performance dataset.

How to read this score

Use matched prompts and exam-level breakdowns; do not treat one aggregate score as general intelligence or professional qualification.

Similarity to real tasks: Low — Tasks originate in human exams, but static questions and exact grading do not reproduce most open-ended production work.

Data contamination risk: Unknown — The source material is public and no model-specific training-overlap audit is registered.

Benchmark gaming risk: Unknown — No benchmark-specific optimization or prompt-tuning audit is registered.

What you still need to test yourself

  • Test open-ended, current, domain-specific tasks with the target language and output format.
  • Evaluate calibration, safety, cost, latency, and model-specific training overlap separately.

This benchmark supports decisions about:

  • Compare broad exam-style reasoning under the same AGIEval task set and prompt protocol.
  • Identify languages or exam domains requiring targeted evaluation.

Limitations

  • Exam accuracy does not prove professional competence, tool use, safety, or production reliability.
  • Cross-report scores may use different subsets, prompts, languages, and answer extraction.

Dataset freshness is not fully confirmed from available source coverage. The paper date is not treated as a knowledge cutoff.

Sources

How this benchmark is scored

Categoryreasoning
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://arxiv.org/abs/2304.06364

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to AGIEval Leaderboard — AI Model Scores.