ModelRefs / GPQA Methodology — Methodology
GPQA Methodology — Methodology
GPQA tests graduate-level scientific reasoning in biology, chemistry, and physics with expert-validated questions.
Overview
What it measures: Reasoning on novel, Google-resistant graduate-level science problems.
How it works
- ~450 questions written by domain PhDs.
- Validated to be unanswerable by non-experts even with web access.
- Multiple choice with 4 options.
- Diamond subset (~200) is hardest and most cited.
Strengths
- Google-resistant
- Expert-validated
- Discriminates frontier reasoning
Limitations
- Small set increases variance
- Multiple choice still gameable
Best use cases
- Frontier reasoning evaluation
- Reasoning-model differentiation
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to GPQA Methodology — Methodology.
Frequently asked questions
What does GPQA measure?
Reasoning on novel, Google-resistant graduate-level science problems.
What are its main limitations?
Small set increases variance Multiple choice still gameable
When should I use this benchmark?
Frontier reasoning evaluation Reasoning-model differentiation