ModelRefs / GPQA Diamond — AI Glossary
GPQA Diamond — AI Glossary
A graduate-level multiple-choice benchmark across physics, chemistry, and biology that expert human PhDs answer correctly only ~65% of the time.
Overview
GPQA Diamond (Google DeepMind, 2023) is the hardest open reasoning benchmark. Questions are deliberately 'Google-proof' — web search does not help. It measures genuine expert-level scientific reasoning rather than recall.
Reference details
| Topic | evaluation |
|---|---|
| Also known as | GPQA, Graduate-Level Google-Proof Q&A |
| Last reviewed | 2026-06-24 |
Related terms
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to GPQA Diamond — AI Glossary.
Frequently asked questions
What is GPQA Diamond?
A graduate-level multiple-choice benchmark across physics, chemistry, and biology that expert human PhDs answer correctly only ~65% of the time.
Is GPQA Diamond the same as GPQA?
Yes — GPQA, Graduate-Level Google-Proof Q&A are common aliases for GPQA Diamond.
What concepts are related to GPQA Diamond?
Closely related concepts include evaluation benchmark, mmlu.