A visual, step-by-step guide to the benchmark that fixed evaluation's blind spot: in the real world the user also acts on the environment — and the agent must guide an unreliable partner through a shared, changing world.
Every prior conversational-agent benchmark made the user a talking head. τ² gives the user hands.
Consider real technical support: "have you tried restarting the router?" The user physically restarts the router. In every prior benchmark, that step simply didn't exist — the agent either could do everything or nothing changed. Dual-control inserts the user as a co-actor with their own tools, and scores the agent on steering the whole two-person system to the goal.
τ-bench exposed agent unreliability; τ² exposes a subtler failure — agents that can't collaborate.
A single-control agent is an autopilot: it flies, you watch. A dual-control agent is a pilot flying with a passenger who has their own controls — it must run the checklist, watch what the passenger actually did (not what they said they did), and recover when they flip the wrong switch. CRM — crew resource management — is the skill being measured.
The paper's own summary: a domain, a generator, a simulator, and a metric — each fixing a predecessor's flaw.
The flagship domain: account, devices, plans, diagnostics — split between company systems and the customer's own hardware.
In τ-bench, simulator mistakes could sink a capable agent — attribution was muddy (retail simulator ≈ 40% error, 12% critical). τ²'s rebuilt simulator reaches ~16% error with ~6% critical in telecom. When the agent fails now, it's (usually) the agent — the measurement got cleaner, and the paper reports telecom-domain user-simulator stats to prove it.
The compositional generator assembles tasks from atomic parts — programmatic diversity with a difficulty dial.
Atomic components combine into chains: more steps, more user-side dependencies, more partial observability — the same domain scales from two-turn tasks to long interleaved sessions. Research can now ask where agents break along an axis, not just that they break.
The paper's sharpest empirical pattern: performance drops when agents move from no-user to dual-control settings.
τ² reframes what conversational agents are: not tools users operate, but teammates users coordinate with.
The hardest problem in dual-control evaluation: when the two-person system fails, who failed?
τ²-bench's lasting idea is an honest accounting unit: the two-actor system. As agents deploy into worlds where humans keep some of the controls — medicine, field ops, home automation — "the agent's score" will increasingly mean "the joint system's score, with attribution." This benchmark is the first to build that machinery properly, and its most quotable stat (40% → 16% simulator error) is really a lesson: before you measure minds interacting, fix your measuring partner.
Check your understanding of the key concepts from the τ²-bench paper.
Everything you need to remember about this paper.