ModelRefs / GPQA Diamond Leaderboard — AI Model Scores

GPQA Diamond Leaderboard — AI Model Scores

Graduate-level Google-Proof Q&A in physics, chemistry, and biology. Current leaders, methodology, and citation sources for GPQA Diamond.

Overview

Graduate-level Google-Proof Q&A in physics, chemistry, and biology.

How it is measured: Closed-book, single-answer; Diamond subset is the highest-difficulty tier.

What this benchmark measures

  • graduate-level science reasoning
  • closed-book question answering

Relevant to:

  • hard reasoning evaluation portfolios
  • scientific question-answering screening

Failure modes it exercises:

  • expert-domain reasoning errors in covered science subjects

Method and its limits

Four-option graduate-level science questions, with the Diamond subset representing the questions on which expert validators agreed and skilled non-experts performed worst.

  • Coverage is concentrated in physics, chemistry, and biology.
  • Accuracy does not establish citation quality, tool use, or reliable scientific practice.

Dataset

Dataset
Graduate-Level Google-Proof Q&A
Type
expert-authored closed-book science questions
Freshness
aging

The registered description covers 448 expert-written biology, physics, and chemistry questions in the main set; broader domain coverage is not implied. The date records paper release, not a scientific knowledge cutoff.

How to read this score

Accuracy is directly interpretable within the selected subset, but transfer to production science tasks is uncertain.

Similarity to real tasks: Low — Established from the primary methodology: GPQA contains 448 four-option questions written by graduate-level domain experts in biology, physics, and chemistry, with expert and skilled non-expert validation used to establish difficulty. This resembles a narrow expert knowledge-and-reasoning check, not end-to-end scientific work: the model selects one supplied answer and is not asked to formulate a research question, search or assess literature, design or run experiments, analyze original data, cite evidence, collaborate with experts, or produce a reviewable scientific deliverable. Real-task similarity is therefore supported but low rather than unavailable.

Data contamination risk: Medium — The repository includes a canary string and password-protected dataset archive, which are leakage mitigations, but the dataset is publicly obtainable and no model-run-specific overlap audit is registered.

Benchmark gaming risk: Medium — GPQA Diamond is a common frontier reasoning signal; subset choice, prompt style, answer-order randomization, and provider-reported protocols can affect comparability.

What you still need to test yourself

  • Test retrieval, citation quality, uncertainty, and expert-review behavior on the target scientific workflow.
  • Evaluate domain governance and harmful-error consequences separately.

This benchmark supports decisions about:

  • Screen closed-book reasoning on difficult biology, physics, and chemistry questions.
  • Add a hard expert-domain signal to a broader reasoning evaluation portfolio.

Limitations

  • Not evidence of safe or reliable scientific advice.
  • Does not test retrieval, citations, experimentation, or domain governance.
  • The canary is a leakage mitigation, not proof that a particular model run is uncontaminated; every score still needs model-, subset-, protocol-, and source-scoped run evidence.

Sources re-reviewed 2026-07-11. GPQA is treated as an aging 2023 benchmark; the paper date is known but is not a dataset knowledge cutoff or score-freshness date.

Sources

How this benchmark is scored

Categoryreasoning
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://arxiv.org/abs/2311.12022

Published results

ModelScoreRun dateSource
GPT-585.42026-05-01Aggregated public reports
Claude Opus 480.22026-05-01Aggregated public reports
DeepSeek R179.52026-05-01Aggregated public reports
GPT-5 Mini712026-05-01Aggregated public reports
Llama 4 Scout602026-05-01Aggregated public reports
Mistral Large 2562026-05-01Aggregated public reports
Command R+49.42026-05-01Aggregated public reports

Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to GPQA Diamond Leaderboard — AI Model Scores.