History Problem Core Idea Action Space Results Why It Works Impact Quiz
Interactive Paper Explainer

Reasoning + Acting
ReAct

The paper that taught language models to interleave thinking with doing — reasoning traces guide actions, actions gather evidence, and observations ground reasoning. It became the template behind every modern AI agent.

Start Learning Read the Paper ↗
3
Loop Steps (Thought · Action · Observation)
4
Benchmarks (HotpotQA · FEVER · ALFWorld · WebShop)
540B
Backbone (PaLM-540B)
2022
Year Published
History

Two Threads, One Loop

One research thread taught models to act; the other taught them to reason. ReAct is where the two finally met.

2017
RL agents in text worlds
Reinforcement-learning agents act in text-based environments, learning from reward signals alone — acting without reasoning, at huge sample cost.
2020–21
Language models become policies
LMs start choosing actions in interactive environments — by 2022, SayCan grounds LM plans on real robots. The reasoning behind each action stays implicit.
2022 · Jan
Chain-of-Thought prompting
Pure reasoning in text — step-by-step traces unlock multi-step problems, but the model never touches the world, so confident hallucination persists. CoT Guide ↗
2022 · Oct
🚀 ReAct (Yao et al.)
Reasoning and acting in one loop: thoughts steer actions, actions fetch evidence, observations correct the plan — all taught with a handful of example traces.
2023
Toolformer & function calling
Models teach themselves to call tools (Toolformer) and APIs ship native function calling — tool traces go mainstream. Toolformer Guide ↗
2024 →
The agent era
AutoGPT-style autonomous loops and agent frameworks (LangChain, LlamaIndex, …) standardize on ReAct-style traces — "agentic AI" enters the vocabulary.
Key Insight

Reasoning and acting are two halves of one skill. The reasoning trace tells the agent why it is about to act; the action brings back facts that reasoning alone could never invent. Interleave them, and each half fixes the other's failure mode.

THE LOOP IN ONE GLANCE
Thought 1  I should search for the entity first.
Action 1  Search[Apex Mountains]
Observation 1  (intro paragraphs of the article…)
Thought 2  The range isn't mentioned — look it up…
Thought (amber) plans · Action (blue) executes · Observation (gray) reports back.
Chapter 01

The Problem — Half a Brain

In 2022, an agent could act or it could reason — almost never both at once. Each camp had a signature failure mode.

🤖
Acting without reasoning
  • Imitation / RL agents map states straight to actions
  • No explicit reasoning to explain why a move was chosen
  • One wrong action compounds — nothing notices the drift
  • Opaque decision chains — hard to debug, hard to trust
  • Limited to behaviors seen (or rewarded) in training
🧠
Reasoning without acting
  • Chain-of-thought never queries the outside world
  • Confidently hallucinated facts read exactly like true ones
  • Knowledge frozen at pre-training time — stale, unchecked
  • Early errors propagate silently to the final answer
  • No way to gather missing evidence mid-chain
🔁
ReAct — both halves, one loop
  • Thoughts plan, track progress, and flag uncertainty
  • Actions retrieve ground truth (Search, Lookup) or change the world
  • Observations correct the plan mid-episode
  • Every step is plain text — readable, debuggable, promptable
  • Few-shot: a handful of example traces teaches the format — no training
Analogy — Compass vs. Eyes
🤖 Acting-only agent
A courier sprinting a route on pure muscle memory — fast, but it never reads a street sign, never notices the wrong turn, and cannot tell you why it turned left.
🧠 Reasoning-only agent
A navigator plotting the whole trip from a beautiful, decades-old map — the plan is impeccable, but the bridge it crosses was demolished years ago.
🔁 ReAct agent
Plan a few streets, look around, re-plan — compass in one hand, open eyes in the other. The map and the street signs finally agree.
Chapter 02

The Core Idea — One Loop

ReAct prompts a large language model to solve tasks by interleaving thinking and doing. Every step of an episode is plain text — either the model writes it, or a tool writes it back.

Thought 1 → Action 1 → Observation 1 → Thought 2 → … → Finish[answer]
💭 Thought
Private reasoning
The model analyzes the current state — what is known, what is missing, what to do next — and writes the reasoning down.
⚡ Action
Tool call
A formatted verb: Search[query], Lookup[term], Finish[answer] — or an environment command like "go to shelf 1".
👁 Observation
Evidence returns
Whatever the tool or environment answers is appended to the context — fresh, external, grounded.
🏁 Finish
Exit
When the thoughts say the goal is met, the agent emits Finish[answer] and the episode ends.
Interactive Demo — Agent Trace Player

A HotpotQA-style question, solved the ReAct way. Press Next step to advance the precomputed trace — watch thoughts plan, actions execute, and observations report back.

Q: What is the elevation range for the Apex Mountains?
●●● react-agent — wikipedia api · PaLM-540B step 0 / 8
Thought = amber · Action = blue · Observation = gray
What the model actually sees

The question, a handful of demonstration traces, and the growing episode itself — one flat sequence of text. Nothing is hidden: the "agent" is a language model reading and writing a transcript of its own actions. No fine-tuning is involved; the loop is pure prompting.

Chapter 03

The Action Space

ReAct keeps the interface tiny. For question answering, three verbs cover everything; in interactive environments, the verbs come from the world itself.

Interactive Demo — The Three QA Verbs

Click a card to inspect the exact text the model writes, and what it gets back.

📚 Wikipedia API
The QA setting for HotpotQA & FEVER. Verbs: Search[entity], Lookup[term], Finish[answer] over a preprocessed Wikipedia dump.
🏠 ALFWorld
Text-based household tasks: go to shelf 1, take candle 1, use candle 1 to examine the letters… The loop is identical — only the verbs change.
🛒 WebShop
A simulated e-commerce store: search [running shoes], click [item], buy now. Scored on how well the purchase matches the instruction.
Why keep it tiny?

Every action must be plain text the model can write. A small verb set keeps the prompt short, the format stable, and the trace readable — and adding a new environment means writing new demonstrations, not new code. That modularity is exactly what agent frameworks later inherited.

Chapter 04

Results — When to Reason, When to Act

The honest story: acting grounds the model but can cap recall; reasoning recalls more but hallucinates. The best agent knows when to do which.

HOTPOTQA (EM)
~34.2
ReAct + CoT-SC, PaLM-540B — best of both
self-consistency decides when to re-act
ALFWORLD (SUCCESS)
~71%
ReAct vs ~45% for imitation / RL baselines
tens of points of absolute gain
HALLUCINATIONS
Far fewer
than CoT on HotpotQA & FEVER — every claim backed by an observation
qualitative finding — illustrative bar
WEBSHOP
Strong
task-completion gains over IL / RL baselines
the same loop, shopping edition
Four Benchmarks, One Loop (PaLM-540B — "~" marks approximate / qualitative values)
BenchmarkBaselinesCoT (reason-only)ReActReAct + CoT-SC
HotpotQA — multi-hop QA (EM)standard promptingslightly higher EM, more confident errorscomparable EM, far fewer hallucinations~34.2 EM — best of both
FEVER — fact verificationstandard promptingverdicts reasoned from (stale) memoryclaims checked against retrieved passagesbest of both — accurate & faithful
ALFWorld — household tasks~45% success (imitation / RL)—~71% success — tens of points up—
WebShop — shopping instructionsIL / RL agents—strong task-completion gains—

HotpotQA & FEVER query a Wikipedia API; ALFWorld & WebShop are interactive environments, where the paper's comparison points were imitation / RL agents (hence "—" in the CoT column). ReAct alone can slightly trail CoT on exact match when the answer lives inside the model — but it hallucinates far less, and the combination wins overall.

Interactive Demo — The Hallucination Showdown

Same claim, two strategies. Chain-of-Thought never leaves the model's head; ReAct checks the world. Press Run both and read the verdict.

CLAIM: "The Apex Mountains were first mapped in 1853." — SUPPORTED or REFUTED?
🧠 CoT — reasoning only
🔁 ReAct — reasoning + acting

Illustrative trace in the style of the paper's qualitative case studies — the pattern the authors report: CoT confidently wrong, ReAct grounded right.

Chapter 05

Why ReAct Works

Four mechanics make the loop more than the sum of its parts — plus one property nobody expected: legibility.

🧠 Reasoning decomposes & tracks
Thoughts break a goal into sub-goals and keep the agent oriented across long episodes — nothing gets silently lost.
🔎 Acting injects fresh evidence
Observations carry information the model never memorized — or that changed since pre-training — straight into the context.
📜 Traces are interpretable
You can read the agent's diary. When it fails, the failure localizes to a step you can inspect, discuss, and fix.
🪄 Few-shot is enough
A handful of demonstration traces teaches the format — no fine-tuning, no reward engineering, no environment-specific code.
🔁 Errors are recoverable
A wrong turn becomes an observation the next thought can react to — re-search, re-plan, continue. Failure is a step, not a wall.
The Agent's Diary — failures localize
Thought 2  The article will mention the range.
Action 2  Lookup[range]
Observation 2  (no matching sentences found) failure — visible & local
Thought 3  Nothing here. Let me search a related entity instead. recovery

No stack trace, no reward curve — the debug log is the same text the model reasons over. That single design choice is a big reason agent traces became the industry standard for observability.

Legacy

Impact — The Agent Template

ReAct didn't just win benchmarks. It named the pattern the industry would build on.

🦜 LangChain & LlamaIndex
The "ReAct agent" became a stock component of the two biggest LLM application frameworks — often the default meaning of the word "agent".
🛠️ Function calling
Modern tool-use APIs — models calling functions and reading results back — descend directly from Thought → Action → Observation traces.
🤖 AutoGPT era
The 2023 autonomous-agent boom standardized on ReAct-style loops — ambitions scaled faster than reliability did.
📊 Agent benchmarks
Eval suites like AgentBench and its successors grade models inside ReAct-style loops — the paper's format became the test harness.
💬 "Agentic AI"
The vocabulary of the agentic-AI era — traces, tools, loops, observations — descends from this paper's framing.
🔗 Next: Toolformer
Where ReAct is prompted, Toolformer is self-taught: models learn when to call which tool. Toolformer Guide ↗
Follow the Thread

Chain-of-Thought showed that reasoning in text is powerful but ungrounded → ReAct closed the loop with actions → Toolformer taught models to call tools themselves. Three papers, one track: the making of the modern agent.

Test Yourself

Quick Quiz

Check your understanding of the key concepts from the ReAct paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ One loop: Thought → Action → Observation, repeated until the agent emits Finish[answer].
✅ Every step is plain text — the model learns the format from a handful of demonstration traces (few-shot prompting, zero training).
✅ QA action space: Search[entity], Lookup[term], Finish[answer] over a Wikipedia API.
✅ The same loop drives ALFWorld (household tasks) and WebShop (shopping) — only the verbs change.
✅ Grounded observations ⇒ far fewer hallucinations than CoT; combined with CoT self-consistency, ~34.2 EM on HotpotQA (PaLM-540B).
✅ ~71% vs ~45% success on ALFWorld over imitation / RL baselines — and the trace format became the template for modern agents and function calling.