448 graduate-level science questions, written and validated by PhDs — and deliberately hardened so that skilled non-experts with 30+ minutes of open web access still score barely above chance.
The contamination problem, and the escalation answer: make questions that search cannot solve.
GPQA's questions are not merely hard — they are hard to look up. Construction: domain experts (PhD-level) write questions requiring genuine reasoning with specialist knowledge; then validation: other experts must answer them correctly AND flag flaws, while skilled non-experts spend 30+ minutes with unrestricted web access trying to solve them by search — if the non-experts succeed, the question is revised or dropped. The surviving 448 questions define a clean separation: expert reasoning, not retrievable trivia — hence "Google-proof."
The evaluation crisis GPQA was engineered to escape.
A contaminated benchmark is an open-book exam where students photographed the book. GPQA builds a locked room: the questions are written in a language (graduate physics) that a visitor with the library card (Google) still cannot read — the examiner personally verifies the room is sealed by watching search-armed non-experts fail it, again and again.
The quality-control loop that defines the dataset.
The design lesson that outlived the specific dataset.
Hard open-ended generation is expensive to grade; multiple choice is cheap but lookup-prone. GPQA's move: keep the cheap answer format but escalate the question's construction cost — expert hours spent upfront buying automatic, permanent verifiability. That trade (expensive authoring, cheap automated grading, resistance to search) became the template for post-2023 hard evals: ARC-AGI-style puzzles, frontier-math sets, and every "expert-verified" benchmark since. Contamination resistance is bought with human expertise at dataset-construction time, not with clever grading after the fact.
Three numbers that defined a clean measurement frontier.
| Population | Accuracy | Conditions |
|---|---|---|
| Domain experts (PhD) | 65% (74% adj.) | closed book, limited time |
| Skilled non-experts | 34% | 30+ min, unrestricted web |
| Chance | 25% | 4-way multiple choice |
| GPT-4 + retrieval (2023) | well below expert | the measured gap |
The separation table: expertise beats search by construction — that margin is what models are graded against.
GPQA-diamond became the hard-science line on every model card.
Check your understanding of the key concepts from GPQA.
Everything you need to remember about this paper.