History Problem Core Idea Design Results Impact Quiz Takeaways
Interactive Paper Explainer

Questions Google Can't Answer
GPQA

448 graduate-level science questions, written and validated by PhDs — and deliberately hardened so that skilled non-experts with 30+ minutes of open web access still score barely above chance.

Start Learning Read the Paper ↗
448
Expert questions
65% → 74%
Expert accuracy
34%
Non-expert + web
2023
Rein et al.
History

Benchmarks That Leak

The contamination problem, and the escalation answer: make questions that search cannot solve.

2020-23
Contamination anxieties
MMLU (entry #66) and friends permeate training corpora; leaderboard gains become ambiguous — knowledge or memorization?
2023
Fresh-question strategies
Live benchmarks, held-out rewrites, private test sets — each patch criticized (judges biased, crowdsourcing noisy, private unverifiable).
Nov 2023
🚀 GPQA
Rein et al.: escalate difficulty past searchability — questions where domain experts beat search-equipped non-experts by design. 448 validated questions; three-way expert agreement required.
2023+
The frontier gate
GPQA-diamond becomes the standard hard-science filter in model cards (GPT-4o, Claude, Gemini all report it); expert-level science joins the frontier checklist.
2024-25
The ceiling question
Frontier models pass GPQA-diamond — escalation continues (post-graduate, research-frontier sets); GPQA remains the template for hard-verifiable evals.
Google-Proofing as a Design Process

GPQA's questions are not merely hard — they are hard to look up. Construction: domain experts (PhD-level) write questions requiring genuine reasoning with specialist knowledge; then validation: other experts must answer them correctly AND flag flaws, while skilled non-experts spend 30+ minutes with unrestricted web access trying to solve them by search — if the non-experts succeed, the question is revised or dropped. The surviving 448 questions define a clean separation: expert reasoning, not retrievable trivia — hence "Google-proof."

Chapter 01

Leaderboards Corroded from Inside

The evaluation crisis GPQA was engineered to escape.

🩸
The Contamination Spiral
  • Public benchmarks enter pre-training data — scores inflate without capability
  • Held-out rewrites get attacked as non-representative; judges introduce their own biases
  • Multiple-choice science questions on the public web can be searched — 'closed-book' is unenforceable
  • The field needs hard, verifiable questions that resist both memorization AND lookup
🛡
The GPQA Answer
  • Graduate-level biology, physics, chemistry — written by experts with domain PhDs
  • Expert validation: other PhDs answer; errors found in retrospect are excluded (65% → 74% adjusted)
  • Google-proofing: web-equipped non-experts (30+ min, unrestricted) score only 34% — near chance
  • 448 surviving questions; a 'diamond' subset of most-agreed items as the flagship slice
Analogy — The Locked-Room Exam

A contaminated benchmark is an open-book exam where students photographed the book. GPQA builds a locked room: the questions are written in a language (graduate physics) that a visitor with the library card (Google) still cannot read — the examiner personally verifies the room is sealed by watching search-armed non-experts fail it, again and again.

Chapter 02

The Three-Way Validation

The quality-control loop that defines the dataset.

1️⃣ Expert authors
PhD-level scientists in biology, physics, chemistry write questions requiring real domain reasoning — not textbook recall.
2️⃣ Expert validators
Different experts answer each question under time constraints; items they miss or flag get examined — mistakes identified in retrospect excluded (65% raw → 74% adjusted expert ceiling).
3️⃣ Search-proofing
Skilled non-experts with 30+ minutes and unrestricted web access attempt each question — solvable-by-search items are revised or removed. Non-experts end at 34%.
4️⃣ The diamond cut
The most-validated, most-robust subset — GPQA-diamond — becomes the reported standard: fewer questions, cleaner signal.
The Measured Landscape (2023, from the paper)
Interactive Demo — Building One Google-Proof Question

Follow a question from expert draft through search-hardening to the surviving validated form.

Chapter 03

Why Hard Verifiable Benchmarks Matter

The design lesson that outlived the specific dataset.

The Verification Asymmetry

Hard open-ended generation is expensive to grade; multiple choice is cheap but lookup-prone. GPQA's move: keep the cheap answer format but escalate the question's construction cost — expert hours spent upfront buying automatic, permanent verifiability. That trade (expensive authoring, cheap automated grading, resistance to search) became the template for post-2023 hard evals: ARC-AGI-style puzzles, frontier-math sets, and every "expert-verified" benchmark since. Contamination resistance is bought with human expertise at dataset-construction time, not with clever grading after the fact.

Interactive Demo — What Makes a Question Google-Proof

Tab through question archetypes — which survive the search attack, and why.

Chapter 05

The Search-Proof Gap

Three numbers that defined a clean measurement frontier.

EXPERTS
65% / 74%
raw / excluding identified mistakes
NON-EXPERTS + WEB
34%
30+ minutes, unrestricted access
QUESTIONS
448
biology · physics · chemistry
MODEL GAP (2023)
large
strongest GPT-4 baselines below expert level
Interactive Demo — The Separation Chart

Press run for GPQA's defining separation — expertise versus search versus models (2023).

PopulationAccuracyConditions
Domain experts (PhD)65% (74% adj.)closed book, limited time
Skilled non-experts34%30+ min, unrestricted web
Chance25%4-way multiple choice
GPT-4 + retrieval (2023)well below expertthe measured gap

The separation table: expertise beats search by construction — that margin is what models are graded against.

Legacy

Legacy — The Frontier Filter

GPQA-diamond became the hard-science line on every model card.

🚪 The frontier gatekeeper
GPQA-diamond scores became the standard 'genuinely hard science' number in frontier releases — the benchmark that separated memorization-laced claims from reasoning.
🛡 Contamination economics
The design lesson institutionalized: buy verifiability with expert authoring hours upfront — the template for every expert-built hard eval since.
🔬 Expert-level as a target
By pricing the human ceiling honestly (65/74%), GPQA gave models a real number to chase — and the field a clean definition of 'expert-level science'.
⚠️ What it did NOT solve
448 questions is small (slicing to diamond makes it smaller — statistical power limited); expert-written English-only; and as models pass diamond, escalation moves on — GPQA measures a frontier, it doesn't freeze one.
🛤 Read next
The eval ladder: LiveBench · BrowseComp · MMLU
Test Yourself

Quick Quiz

Check your understanding of the key concepts from GPQA.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ 448 expert-written, expert-validated graduate science questions — biology, physics, chemistry.
✅ Google-proofed behaviorally: search-armed non-experts (30+ min) score 34% ≈ chance; questions they solve are cut.
✅ Honest expert ceilings: 65% raw, 74% excluding identified mistakes.
✅ The 2023 model gap (strong GPT-4 baselines below expert) defined the measurement frontier.
✅ GPQA-diamond became the hard-science line on every frontier model card.
✅ Read it as the template for hard verifiable evals: expert hours buy contamination resistance.