History Problem Core Idea Why Results Impact Quiz Takeaways
Interactive Paper Explainer

Fails, Writes It Down,
Tries Again
Reflexion

Instead of updating weights from trial and error, the agent converts failures into natural-language lessons, stores them in episodic memory, and reads them before the next attempt — 91% pass@1 on HumanEval, past GPT-4's 80%.

Start Learning Read the Paper ↗
91%
HumanEval pass@1
80% → passed
GPT-4 baseline
0
Weight updates
2023
Shinn et al.
History

Trial and Error, Without Gradients

Agents learn from retries — the question was always where the learning lives.

2022
Agents need experience
LLM agents (ReAct, entry #53) act in environments but repeat their mistakes: the prompt resets, the lesson evaporates.
2022
RL's cost problem
Classical trial-and-error learning = fine-tuning: thousands of samples, expensive gradients, catastrophic-forgetting risk (the paper's framing).
Mar 2023
🚀 Reflexion
Shinn et al.: feedback converted to TEXT — the agent reflects on failure, stores the lesson in a memory buffer, and conditions the next trial on it. No training, just an outer loop.
2023+
Memory-based learning spreads
Voyager's skill libraries, Memp prompts, agent long-term memory (the X category) — the pattern: experience as retrievable text.
2024+
Verifier feedback loops
Self-correction and test-time iteration (entry #58) inherit the structure: signal → summary → retry, with the model's weights untouched.
Reinforcement in Words

The paper's equivalence claim: verbal feedback can play the role of the reward signal, and episodic memory can play the role of the parameter update. Concretely — after a failed episode, a language evaluator (self-assessment, unit tests, heuristics, or the environment) produces feedback; the agent verbally reflects on what went wrong into a concrete lesson ("the off-by-one came from indexing after the append"); the lesson enters an episodic memory buffer; the next trial's prompt includes it. Learning accumulates across episodes in text — transferable, inspectable, and free of gradients.

Chapter 01

Agents That Forget Their Mistakes

The retry problem: experience that never persists.

🔁
The Amnesiac Loop
  • LLM agents retry a task; each attempt starts from the same prompt — identical mistakes recur
  • Fine-tuning on failures is expensive: many samples, gradient surgery, forgetting risks
  • Sparse scalar rewards say THAT it failed, not WHY — unusable as direct conditioning
  • Environment feedback (stack traces, test output) is noisy, long, and unstructured
🪞
The Reflexion Answer
  • Convert feedback to language: reflect on the failure trace → a short, specific lesson
  • Store lessons in an episodic memory buffer (size-limited, task-scoped)
  • Condition the next attempt on the accumulated lessons — the prompt grows wiser, not the weights
  • Works with multiple feedback sources: scalar values, free-form text, internally simulated self-critique
Analogy — The Pilot's Logbook

A gradient-based agent is a pilot whose brain rewires after each rough landing — effective, expensive, slightly dangerous. A Reflexion agent keeps a logbook: after each rough flight, they write 'wind shear on approach at runway 27 — come in 5 knots hotter' and read the logbook before every takeoff. The pilot's brain is untouched; the practice compounds. And you can audit every lesson — try that with a Hessian.

Chapter 02

The Outer Loop

Four stations around the episode: act, evaluate, reflect, remember.

1️⃣ Act
The agent attempts the task (coding, decision-making, reasoning) inside an environment — action, execution, result.
2️⃣ Evaluate
A feedback source grades the attempt: unit tests (coding), environment heuristics, scalar scores, or the model's own self-critique.
3️⃣ Reflect
Given the trace + feedback, the agent writes a CONCRETE lesson: what failed, why, and what to change — short, specific, natural language.
4️⃣ Remember & retry
The lesson joins the episodic memory buffer; the next attempt's prompt carries the accumulated lessons. Repeat until success or budget.
The headline: HumanEval 91%
  • Reflexion + GPT-4 reaches 91% pass@1 on HumanEval coding
  • Surpassing the then-SOTA GPT-4 baseline at 80% — without any weight updates
  • Feedback source there: Python execution + unit tests — real, verifiable signals
  • Ablations show the lesson QUALITY matters: generic reflections underperform concrete ones
Task spread + limits
  • Sequential decision-making (AlfWorld), decision-making (HotpotQA-style), language reasoning — consistent gains over a baseline agent
  • Memory buffer is size-limited — lessons get evicted; long-horizon accumulation needs management
  • Reflection can rationalize instead of diagnose — the same model grades its own homework (when self-feedback)
Interactive Demo — One Bug, Two Episodes

Follow a failing coding episode through reflection into a memory buffer, then watch the retry inherit the lesson.

Chapter 03

Why Verbal Works

The information-theoretic argument hiding in the design.

From Signal to Summary

A failing unit test carries a stack trace — hundreds of tokens of noise around one causal insight. Fine-tuning on the raw signal would need many samples to distill the pattern; a language model can compress it in one pass into the operative sentence. The reflection is a lossy-but-causal compression of experience: exactly the part worth carrying forward. That compression choice is also the risk — a model that compresses the wrong cause into the lesson will repeat the right mistake — which is why externally-graded feedback (tests, environments) outperforms pure self-reflection in the paper's ablations.

Interactive Demo — Feedback Sources Compared

Tab through the feedback types Reflexion can consume — and their reliability gradient.

Chapter 05

91% Without Training

The number that made the field take memory-based learning seriously.

HUMANEVAL pass@1
91%
surpassing GPT-4's 80% SOTA
WEIGHT UPDATES
zero
learning lives in the prompt's memory
TASK FAMILIES
3+
decision-making, coding, language reasoning
FEEDBACK TYPES
scalar + text
external or internally simulated
Interactive Demo — Reflexion vs Fine-Tuning

Both are 'learning from experience.' Press reveal for the honest comparison the paper set up.

Legacy

Legacy — Memory Is Learning

Reflexion is the founding document of the agent-memory category.

🗂 Category X's origin story
The paper made 'episodic memory of lessons' a named, studied mechanism — the direct ancestor of MemGPT-style management and the memory systems of entries #97-104.
🧪 Self-correction as standard
Retry-with-insight loops became the default agent architecture pattern (Voyager, SWE-agent recovery, test-time iteration — entry #58's foundation).
💬 The verbal-RL thesis
'Feedback as language' reframed reward engineering: what to tell the model, not just what number to give it — now visible in every verifier-feedback design.
🏆 The 91% datapoint
Beating GPT-4's HumanEval SOTA without training made the engineering world notice: prompt-side accumulation is a real learning channel.
⚠️ What it did NOT solve
Self-reflection rationalizes (the model grades its own homework); memory buffers are short-lived and task-scoped; lessons don't compose into transferable skills; and long horizons need the memory management the X-category papers later supply.
🛤 Read next
The memory lineage: MemGPT · A-MEM · ReAct
Test Yourself

Quick Quiz

Check your understanding of the key concepts from Reflexion.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Reflexion converts failure feedback into verbal lessons stored in episodic memory — no weight updates.
✅ The loop: act → evaluate (tests/environment/self) → reflect → remember → retry, wiser each pass.
✅ 91% pass@1 on HumanEval — past GPT-4's 80% — the result that legitimized memory-based learning.
✅ Works with scalar, free-form, external, or internally simulated feedback; external wins on reliability.
✅ Reflection = lossy-but-causal compression of experience into portable text.
✅ Read it as the origin point of the entire agent-memory category (Category X).