ModelRefs / GPQA Methodology — Methodology

GPQA Methodology — Methodology

GPQA tests graduate-level scientific reasoning in biology, chemistry, and physics with expert-validated questions.

Overview

What it measures: Reasoning on novel, Google-resistant graduate-level science problems.

How it works

  • ~450 questions written by domain PhDs.
  • Validated to be unanswerable by non-experts even with web access.
  • Multiple choice with 4 options.
  • Diamond subset (~200) is hardest and most cited.

Strengths

  • Google-resistant
  • Expert-validated
  • Discriminates frontier reasoning

Limitations

  • Small set increases variance
  • Multiple choice still gameable

Best use cases

  • Frontier reasoning evaluation
  • Reasoning-model differentiation

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to GPQA Methodology — Methodology.

Frequently asked questions

What does GPQA measure?

Reasoning on novel, Google-resistant graduate-level science problems.

What are its main limitations?

Small set increases variance Multiple choice still gameable

When should I use this benchmark?

Frontier reasoning evaluation Reasoning-model differentiation