ModelRefs / MMLU Methodology — Methodology

MMLU Methodology — Methodology

MMLU evaluates broad multi-domain academic knowledge via 4-way multiple choice across 57 subjects.

Overview

What it measures: Breadth of factual and conceptual knowledge from elementary through professional levels.

How it works

  • 57 subject categories ranging from elementary math to professional law.
  • Each item is a 4-way multiple choice question.
  • Models are scored on exact-match accuracy across ~14K questions.
  • Typically evaluated 0-shot or 5-shot.

Strengths

  • Wide subject coverage
  • Cheap to evaluate
  • Well-established baseline

Limitations

  • Saturated by frontier models (>90%)
  • Multiple-choice masks reasoning weaknesses
  • Contamination risk in pretraining

Best use cases

  • General-knowledge baselining
  • Cross-model comparison at parity

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MMLU Methodology — Methodology.

Frequently asked questions

What does MMLU measure?

Breadth of factual and conceptual knowledge from elementary through professional levels.

What are its main limitations?

Saturated by frontier models (>90%) Multiple-choice masks reasoning weaknesses Contamination risk in pretraining

When should I use this benchmark?

General-knowledge baselining Cross-model comparison at parity