ModelRefs / MMMU Leaderboard — AI Model Scores
MMMU Leaderboard — AI Model Scores
Massive Multi-discipline Multimodal Understanding across 30 subjects. Current leaders, methodology, and citation sources for MMMU.
Overview
Massive Multi-discipline Multimodal Understanding across 30 subjects.
How it is measured: Multiple-choice + open accuracy on college-level visual questions.
What this benchmark measures
- multimodal perception and reasoning
- college-level domain knowledge
- reasoning across heterogeneous visual inputs
Relevant to:
- multimodal-model screening
- expert-domain visual reasoning portfolios
- document, chart, diagram, and scientific-image evaluation planning
- limited failure-case planning for clinical documents, legal records, IP filings, and investor reports
Failure modes it exercises:
- incorrect integration of text and visual evidence
- domain-knowledge errors across covered subjects
- reasoning failures on heterogeneous image types
Method and its limits
Evaluate multimodal models on the pinned MMMU split across six disciplines, 30 subjects, and heterogeneous image types using the official evaluation code and a declared zero- or few-shot protocol.
- Aggregate accuracy can hide large differences by discipline, subject, question type, and image type.
- Prompting, image preprocessing, answer extraction, split, and access to interleaved images materially affect comparability.
- MMMU and MMMU-Pro are distinct benchmarks and must not be mixed.
Dataset
- Dataset
- MMMU
- Type
- college-level multimodal questions from exams, quizzes, and textbooks
- Freshness
- aging
The official repository reports 11.5K questions across six disciplines, 30 subjects, 183 subfields, and 32 image types. That breadth does not establish coverage of a deployment's domain or current-world information.
How to read this score
Use discipline-, subject-, question-, and image-type breakdowns under matched protocols. An aggregate MMMU score is not a universal multimodal rank.
Similarity to real tasks: Low — The questions draw from college exams, quizzes, and textbooks and cover diverse image types, but remain controlled academic tasks rather than end-to-end professional workflows.
Data contamination risk: Medium — Questions originate from educational materials and the dataset/evaluation code are public. No model-specific training-overlap audit is registered, so contamination cannot be ruled out.
Benchmark gaming risk: Medium — Benchmark-specific prompting, image preprocessing, answer extraction, split selection, and MMMU-vs-MMMU-Pro confusion can affect results.
What you still need to test yourself
- Test the deployment's actual modalities, image quality, domain terminology, documents, charts, diagrams, safety risks, and human-review path.
- Record split, prompt, shots, image preprocessing, answer extraction, model snapshot, latency, cost, and failure breakdowns.
- Use authorized domain documents and qualified clinical, legal, IP, finance, or investor-relations reviewers to measure source traceability, omission, privacy, and consequential interpretation errors.
This benchmark supports decisions about:
- Screen multimodal candidates under a broad, pinned academic visual-reasoning protocol.
- Identify disciplines and image types that need dedicated workflow evaluation.
Limitations
- MMMU does not establish safe or reliable performance in clinical, financial, legal, engineering, or other professional workflows.
- Broad academic coverage does not substitute for deployment-specific visual quality, grounding, privacy, latency, cost, and oversight evaluation.
- MMMU does not prove clinical interpretation, legal or IP conclusions, medical coding, payer-policy answers, or investor disclosure validity.
- Every score still needs exact split, prompt, shot count, preprocessing, answer-extraction, model-version, and source-scoped run evidence.
Sources re-reviewed 2026-07-11. MMMU is treated as an aging 2023 benchmark; the paper date is a release reference, not a current-world or educational-material freshness cutoff.
Sources
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI MMMU authors / CVPR 2024 · accessed 2026-07-11
- MMMU official repository MMMU-Benchmark · accessed 2026-07-11
How this benchmark is scored
| Category | multimodal |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
Primary source: https://mmmu-benchmark.github.io/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MMMU Leaderboard — AI Model Scores.