1,266 questions whose answers exist but hide: multi-hop, entangled, spread across the web. Short verifiable answers make grading instant — and GPT-4o with browsing scores ~2%.
The capability gap between retrieval and persistence.
The questions are engineered so the answer cannot be fetched — only found: chains of lookups where each hop's query depends on what the previous page revealed ("the third album of the band that the producer of X founded after leaving Y"). The benchmark's own framing: it measures "the important core capability of exercising persistence and creativity in finding information" — analogous to how programming competitions under-measure but usefully benchmark coding. Short, checkable answers keep grading objective while the search process stays free-form.
What standard search (and RAG) structurally cannot do.
The search bar is a receptionist: one query, one directory lookup, one answer. A BrowseComp question needs an archivist: follow the citation in box 12 to the letter in folder C, notice the name change in 1962, cross-check the second archive — six hops, three dead ends, one reformulation, and then the answer. The receptionist hangs up at hop one; the archivist persists.
Multi-hop, entangled, verifiable — the construction triangle.
The paper's second contribution: how to build a browsing agent that doesn't fail at 2%.
The number that exposed browsing as an architecture problem.
| Agent configuration | BrowseComp accuracy |
|---|---|
| GPT-4o, no browsing | very low |
| GPT-4o + browsing tools | ~2% |
| Deep Research-style agent (paper's design) | far higher — the harness delta |
| Simple web browsing agents | low single digits |
The paper's own comparison shape: the same underlying model swings enormously with the agent harness — SWE-agent's thesis (entry #79), re-proven on the web.
The benchmark that gave the deep-research product category its ruler.
Check your understanding of the key concepts from BrowseComp.
Everything you need to remember about this paper.