History Problem Core Idea Harness Results Impact Quiz Takeaways
Interactive Paper Explainer

Find the Need Across the Web
BrowseComp

1,266 questions whose answers exist but hide: multi-hop, entangled, spread across the web. Short verifiable answers make grading instant — and GPT-4o with browsing scores ~2%.

Start Learning Read the Paper ↗
1,266
Questions
~2%
GPT-4o + browsing
Deep Research
The unlock
2025
Wei et al. (OpenAI)
History

Search Became Answering — Then Stopped

The capability gap between retrieval and persistence.

2023-24
RAG answers everything shallow
Retrieval + generation answers questions answerable in one lookup — the long tail of hard questions stays unanswered.
2024
Deep-research products appear
Search-grounded assistants browse, read, and synthesize — but nobody can measure HOW WELL they actually search.
Apr 2025
🚀 BrowseComp
OpenAI: 1,266 questions demanding persistent navigation — answers short and verifiable, difficulty brutal: GPT-4o with browsing ~2%.
2025+
The harness lessons
The paper's agent designs (reasoning-browsing interleave, page distillation) become the deep-research playbook; BrowseComp grades it.
Persistence as the Measured Capability

The questions are engineered so the answer cannot be fetched — only found: chains of lookups where each hop's query depends on what the previous page revealed ("the third album of the band that the producer of X founded after leaving Y"). The benchmark's own framing: it measures "the important core capability of exercising persistence and creativity in finding information" — analogous to how programming competitions under-measure but usefully benchmark coding. Short, checkable answers keep grading objective while the search process stays free-form.

Chapter 01

One Query, One Answer — Sometimes

What standard search (and RAG) structurally cannot do.

🔍
The Single-Hop Ceiling
  • Standard search + reading answers what one page states — questions needing multi-hop chains fail
  • RAG pipelines retrieve once and generate — no persistence, no re-querying after partial results
  • Agents give up after failed searches instead of reformulating creatively
  • No benchmark measured search PROCESS — only single-shot retrieval quality
🧭
The BrowseComp Answer
  • 1,266 multi-hop questions with entangled, hard-to-find answers
  • Answers deliberately short and verifiable — grading simple and objective
  • Difficulty tuned so retrieval-augmented GPT-4o lands at ~2%
  • The paper ships agent-harness designs that interleave reasoning and browsing — the deep-research recipe
Analogy — The Archivist vs the Search Bar

The search bar is a receptionist: one query, one directory lookup, one answer. A BrowseComp question needs an archivist: follow the citation in box 12 to the letter in folder C, notice the name change in 1962, cross-check the second archive — six hops, three dead ends, one reformulation, and then the answer. The receptionist hangs up at hop one; the archivist persists.

Chapter 02

The Question Design

Multi-hop, entangled, verifiable — the construction triangle.

What a question looks like
  • Each question chains facts: identifying an entity requires finding another first
  • The answer appears on NO single page — it is assembled across hops
  • Answers are short (a name, a date, a number) — trivially checkable against references
  • Difficulty calibrated: even with browsing tools, GPT-4o answers ~2% correctly
The capability triangle
  • Persistence: keep searching through failed queries — the anti-give-up
  • Creativity: reformulate queries from partial evidence — the anti-one-try
  • Navigation: move through pages following leads — the anti-single-fetch
Interactive Demo — One Question, Six Hops

Follow a multi-hop question through a persistent search — watch the reformulations that single-shot search never makes.

Chapter 03

The Harness Lessons

The paper's second contribution: how to build a browsing agent that doesn't fail at 2%.

Design Patterns from the Paper
Interactive Demo — Why RAG Fails Here

Tab through information-access paradigms against a BrowseComp question — the persistence axis separates them.

Chapter 05

2% — and the Harness Delta

The number that exposed browsing as an architecture problem.

QUESTIONS
1,266
multi-hop, entangled, hard-to-find
GPT-4O + BROWSING
~2%
tools present, persistence absent
GRADING
short answers
instant, objective verification
HARNESS
the lever
interleaved reasoning-browsing designs
Interactive Demo — The Honest Scope

What BrowseComp measures and what it deliberately doesn't. Press reveal.

Agent configurationBrowseComp accuracy
GPT-4o, no browsingvery low
GPT-4o + browsing tools~2%
Deep Research-style agent (paper's design)far higher — the harness delta
Simple web browsing agentslow single digits

The paper's own comparison shape: the same underlying model swings enormously with the agent harness — SWE-agent's thesis (entry #79), re-proven on the web.

Legacy

Legacy — Grading Deep Research

The benchmark that gave the deep-research product category its ruler.

🧭 The deep-research ruler
Search-grounded assistants and deep-research systems report BrowseComp scores as the headline capability number — the category's GPQA.
🏗 The harness canon
The paper's interleaved reasoning-browsing + page-distillation designs became standard deep-research architecture — published as benchmark methodology, adopted as product engineering.
📉 The 2% reset
Exposing naive browsing at ~2% killed 'we gave the model a search tool' as a capability claim — persistence became the metric that matters.
⚠️ What it did NOT solve
Short-answer format sidesteps synthesis quality; question distribution is synthetic-entangled rather than user-realistic; web drift complicates reproducibility; and 'creative persistence' can shade into reward-hacking search loops — the measurement arms race continues.
🛤 Read next
The search lineage: RAG · GAIA · Lost in the Middle
Test Yourself

Quick Quiz

Check your understanding of the key concepts from BrowseComp.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ BrowseComp: 1,266 multi-hop questions with short, objectively verifiable answers.
✅ GPT-4o with browsing: ~2% — persistence, not tools, is the missing capability.
✅ The unlock is harness engineering: reasoning-browsing interleave + page distillation.
✅ Answers assembled across hops — no single page contains them.
✅ Scope is honest: measuring information finding, not answer synthesis or ambiguity handling.
✅ Read it as the ruler (and the recipe book) of the deep-research product category.