The paper that taught language models to interleave thinking with doing — reasoning traces guide actions, actions gather evidence, and observations ground reasoning. It became the template behind every modern AI agent.
One research thread taught models to act; the other taught them to reason. ReAct is where the two finally met.
Reasoning and acting are two halves of one skill. The reasoning trace tells the agent why it is about to act; the action brings back facts that reasoning alone could never invent. Interleave them, and each half fixes the other's failure mode.
In 2022, an agent could act or it could reason — almost never both at once. Each camp had a signature failure mode.
ReAct prompts a large language model to solve tasks by interleaving thinking and doing. Every step of an episode is plain text — either the model writes it, or a tool writes it back.
The question, a handful of demonstration traces, and the growing episode itself — one flat sequence of text. Nothing is hidden: the "agent" is a language model reading and writing a transcript of its own actions. No fine-tuning is involved; the loop is pure prompting.
ReAct keeps the interface tiny. For question answering, three verbs cover everything; in interactive environments, the verbs come from the world itself.
Every action must be plain text the model can write. A small verb set keeps the prompt short, the format stable, and the trace readable — and adding a new environment means writing new demonstrations, not new code. That modularity is exactly what agent frameworks later inherited.
The honest story: acting grounds the model but can cap recall; reasoning recalls more but hallucinates. The best agent knows when to do which.
| Benchmark | Baselines | CoT (reason-only) | ReAct | ReAct + CoT-SC |
|---|---|---|---|---|
| HotpotQA — multi-hop QA (EM) | standard prompting | slightly higher EM, more confident errors | comparable EM, far fewer hallucinations | ~34.2 EM — best of both |
| FEVER — fact verification | standard prompting | verdicts reasoned from (stale) memory | claims checked against retrieved passages | best of both — accurate & faithful |
| ALFWorld — household tasks | ~45% success (imitation / RL) | — | ~71% success — tens of points up | — |
| WebShop — shopping instructions | IL / RL agents | — | strong task-completion gains | — |
HotpotQA & FEVER query a Wikipedia API; ALFWorld & WebShop are interactive environments, where the paper's comparison points were imitation / RL agents (hence "—" in the CoT column). ReAct alone can slightly trail CoT on exact match when the answer lives inside the model — but it hallucinates far less, and the combination wins overall.
Illustrative trace in the style of the paper's qualitative case studies — the pattern the authors report: CoT confidently wrong, ReAct grounded right.
Four mechanics make the loop more than the sum of its parts — plus one property nobody expected: legibility.
No stack trace, no reward curve — the debug log is the same text the model reasons over. That single design choice is a big reason agent traces became the industry standard for observability.
ReAct didn't just win benchmarks. It named the pattern the industry would build on.
Chain-of-Thought showed that reasoning in text is powerful but ungrounded → ReAct closed the loop with actions → Toolformer taught models to call tools themselves. Three papers, one track: the making of the modern agent.
Check your understanding of the key concepts from the ReAct paper.
Everything you need to remember about this paper.