History Problem Architecture Reflection Emergence Evaluation Impact Deep Dive Quiz
Interactive Paper Explainer

A Town Where Memories Live
Generative Agents

A visual, step-by-step guide to the Stanford paper that gave 25 LLM-powered residents of a Sims-like town a memory stream, reflections, and daily plans — and watched a Valentine's Day party emerge from a single seed of intent.

Start Learning Read the Paper ↗
25
Agents in Smallville
4
Architecture Layers
1
Seed → Emergent Party
2023
Year Published
History

From Scripted Chatbots to a Living Town

Generative Agents landed in April 2023, right when LLMs had knowledge but no continuity. Here is the road that led to Smallville.

2017
Scripted chatbots
Rule-based assistants follow conversation trees and intent matching. Every session starts from zero — nothing is remembered.
2020
Neural dialogue, no memory
Open-domain chatbots hold a short context window, but personality and history evaporate the moment a session ends.
2022
LLMs: knowledge without continuity
ChatGPT-era models know a lot, yet are stateless — the same question an hour later gets an answer from a stranger.
2023 · Apr
🚀 Generative Agents
Park et al. give 25 agents a full memory architecture and let them loose in a Sims-like sandbox, Smallville, for two game days.
2023 · Summer
The "AI town" wave
Open-source recreations and agent-town demos spread within months — the paper's architecture becomes the reference recipe.
2024 → 25
Memory becomes a research line
Systems like MemGPT and A-MEM, and surveys of LLM-agent memory, build directly on the memory–reflect–retrieve pattern.
Key Insight

Believability over time is a memory problem, not a scale problem. A stateless LLM, asked "what should I do now?", will do something plausible each moment — but eat lunch at 12:00, again at 12:30, and again at 1:00. The fix is not a bigger model; it is a record: store experiences, retrieve what matters, reflect, plan.

12:00 pm — Klaus has lunch at Hobbs Cafe
12:30 pm — Klaus has lunch (again)
01:00 pm — Klaus has lunch (again!)
with planning → lunch at 12:00, research at the library at 1:00, walk in the park at 3:00
The paper's own example of moment-by-moment believability failing without planning.
Chapter 01

The Problem — Stateless Minds

Before this paper, an LLM agent could roleplay a person for one conversation — but could not stay a person across days, relationships, and events.

🫥
A Stateless LLM
  • No memory — every session starts from zero
  • No consistent persona — no record of past choices to stay in character
  • No social life — cannot form relationships or share news with others
  • Plausible per moment, incoherent over time (the triple lunch)
  • Cannot generalize from raw experience into higher-level conclusions
🏘️
The Generative Agent Stack
  • Memory stream — every experience recorded in natural language
  • Retrieval — relevance + recency + importance selects what matters
  • Reflection — periodic higher-level inferences, stored like memories
  • Planning — day schedules, revised when the world intervenes
  • 25 agents running this loop together in one simulated town
The Analogy

A raw LLM is like a person with anterograde amnesia — brilliant in the moment, unable to form new memories. The architecture hands them a diary and a routine: write down every experience (memory stream), look up what matters (retrieval), reread and draw conclusions (reflection), and plan tomorrow from what the diary says (planning). With the diary, an agent you met yesterday actually remembers you today.

Chapter 02

The Architecture — A Stack of Memory, Reflection, and Plans

The paper's core contribution is an architecture that turns a stateless LLM into an agent that remembers, synthesizes, and plans. Everything runs on gpt-3.5-turbo (ChatGPT) inside a Phaser game engine sandbox.

📝
1 · Memory Stream
Every observation is stored as a natural-language record with a timestamp and an LLM-scored importance (1–10). The stream only grows.
🔍
2 · Retrieval
Each memory is scored on relevance + recency + importance; the top-ranked memories that fit the context window enter the prompt.
💭
3 · Reflection
When recent importance crosses a threshold, the agent synthesizes higher-level inferences — which re-enter the stream as memories.
🗓️
4 · Planning + Reaction
Days are planned top-down (5–8 chunks → hours → 5–15 min), stored like memories, and regenerated when something happens.
perceive → store in stream → retrieve top memories → act / converse → (importance sums high?) → reflect → (re)plan
score = αrecency·recency + αimportance·importance + αrelevance·relevance
The retrieval score that decides which memories enter the agent's context. In the paper, all α = 1.
recency
Exponential decay
0.995 raised to the power of the game-hours since the memory was last retrieved — the morning stays in the attentional sphere.
importance
LLM poignancy score
An integer from 1 (mundane, e.g. brushing teeth) to 10 (poignant, e.g. a breakup) assigned when the memory is created.
relevance
Cosine similarity
Similarity between the embedding of the memory and the embedding of the current query situation.
[0, 1]
Min-max normalization
Each of the three scores is min-max normalized across the stream, then summed with equal weights.
Interactive Demo — Memory Stream Retrieval (Centerpiece)

You are Isabella's retrieval function, on the morning of Feb 14 — one hour before she finalizes plans for the Valentine's Day party (5–7 pm at Hobbs Cafe). Adjust the weights and watch the stream re-rank. The top-3 memories are what Isabella "thinks about".

PRESETS:
α recency: 1.0
α importance: 1.0
α relevance: 1.0

Timestamps, recency decay (0.995/game-hour), and importance follow the paper's implementation. Relevance values are illustrative stand-ins for embedding cosine similarity; memories marked (✓) are drawn from the paper's own examples.

Chapter 03

Reflection — When Memories Become Insight

Raw observations pile up but do not generalize. Reflection synthesizes them into higher-level thoughts that re-enter the memory stream — recursively building a "self".

How a Reflection Forms
  1. Trigger: when the importance scores of recently perceived events sum past a threshold (150 in the paper), the agent reflects — roughly 2–3 times a day.
  2. Question: the LLM reads the 100 most recent records and proposes salient questions, e.g. "What topic is Klaus Mueller passionate about?"
  3. Synthesis: the questions retrieve relevant memories (including earlier reflections), and the LLM extracts insights that cite their evidence: "Klaus Mueller is dedicated to his research on gentrification (because of 1, 2, 8, 15)".
  4. Storage: the insight is stored in the memory stream like any observation — with pointers to the memories it cited — so later retrievals can use it.
Why It Matters — the Gift Question

Asked what to buy Wolfgang for his birthday, Maria without reflections says she does not know what he likes — despite many interactions stored as raw observations.

Maria (no reflection): "I don't know what Wolfgang likes…"

With reflections in her stream, Maria synthesizes what her memories imply:

Maria (with reflection): "Since he's interested in mathematical music composition, I could get him some books… or maybe special software."

Both quotes are from the paper's evaluation (§6.5.3).

Interactive Demo — Reflection Cascade

Run Klaus's reflection: watch seed memories become a question, then an insight, then a new memory — and finally see it change a future decision.

Seed memories, the generated question, and the cited insight are the paper's real examples; the final decision flip (Wolfgang → Maria) is the paper's §4.2 outcome. Verdict lines are paraphrased.

Chapter 04

Emergence — A Party Nobody Planned

The paper's most-loved result: give one agent one seed of intent, and a town's worth of invitations, decorations, crushes, and RSVPs follow — with no user intervention.

The Valentine's Day Party — a 4-beat story (all events from the paper)
THE SEED · FEB 13
One user-specified notion
Isabella Rodriguez, who works at Hobbs Cafe, is initialized with a single intent: throw a Valentine's Day party from 5 to 7 pm on February 14. Nothing else is scripted.
FEB 13 · DAYTIME
Invitations + decoration
Isabella invites friends and customers whenever she sees them, spends the afternoon decorating the cafe, gathers materials, and enlists help. Maria, a close friend, agrees to help decorate.
FEB 13 · NIGHT
A crush acts
Maria's character description says she has a crush on Klaus. That night, she invites Klaus to join her at the party — and he gladly accepts.
FEB 14 · 5 PM
The party happens
Five agents — including Klaus and Maria — show up at Hobbs Cafe at 5 pm and enjoy the festivities. Meanwhile 12 agents in total had heard about the party through conversation chains.
INFO DIFFUSION · PARTY
52%
agents who knew (1 → 13 of 25), none hallucinated it
INFO DIFFUSION · ELECTION
32%
agents who knew about Sam's mayoral run (1 → 8 of 25)
RELATIONSHIPS
0.74
network density after two days, up from 0.167
COORDINATION
5/12
invited agents who showed up at 5 pm on Feb 14
Interactive Demo — Party Diffusion

Press "Spread the word" to advance time and watch the invitation travel through Smallville by pure conversation — the graph below shows 6 of the 25 agents.

t = 0
AGENTS WHO KNOW:
1 / 25

Node layout and intermediate conversation chains are illustrative; the end counts — 12 agents heard (13 knew, 52%), none hallucinated, 5 of 12 invited attended — are the paper's verified numbers (§7.1).

Chapter 05

Evaluation — Do Humans Believe Them?

The team "interviewed" agents with 25 questions across five categories — self-knowledge, memory, plans, reactions, and reflections — and had 100 human evaluators rank the answers for believability.

FULL ARCHITECTURE
μ 29.89
TrueSkill believability — ranked most believable of all conditions
NO REFLECTION (OBS + PLANS)
μ 26.88
losing reflections drops the ranking measurably
NO REFLECTION, NO PLANNING
μ 25.64
observations only — plausible but shallow answers
HUMAN CROWDWORKERS
μ 22.95
roleplaying the agents after watching their lives — a baseline the full architecture beats
NO MEMORY AT ALL (PRIOR WORK)
μ 21.21
the stateless-LLM-agent condition of earlier papers
Statistical Punchline
  • Every component matters: each removal of reflection / planning / observation significantly lowered believability (Kruskal-Wallis H(4)=150.29, p<0.001; all pairwise differences p<0.001 except crowdworkers vs. no-memory).
  • Huge effect size: full architecture vs. the no-memory condition ≈ 8 standard deviations (Cohen's d = 8.16).
  • Beats the human baseline: 100 evaluators ranked the full architecture's interview answers above crowdworker-authored roleplay.
  • Consistent selves: agents like Abigail Chen introduce themselves consistently — memory keeps answers aligned with identity.
What It Did NOT Solve
  • Cost: simulating 25 agents for two days cost thousands of dollars in token credits and took multiple days of compute.
  • Retrieval misses: Tom remembered what to discuss at the party but not that the party existed.
  • Embellishments: 1.3% of 453 awareness answers were hallucinated — Isabella "confirmed" an announcement Sam never made; Yuriko described her neighbor Adam Smith as the author of Wealth of Nations.
  • Instruction-tuned politeness: agents came out overly formal and cooperative — Isabella rarely said no to others' party ideas, and her interests drifted toward theirs.
Legacy

Impact — The Town That Launched a Genre

Generative Agents turned "agent memory" from a hack into a named architecture — and became the template for a whole wave of LLM-agent systems.

🏘️ The AI-town wave
Open-source recreations and "AI town" demos spread within months of the paper, each reimplementing the memory–reflect–plan loop over a map of agents.
🧠 Agent memory as a research line
The memory stream became the reference design for LLM-agent memory systems — with follow-ups making it hierarchical, compressible, and cheaper.
📄 MemGPT & A-MEM lineage
MemGPT (OS-style paging for LLM memory) and A-MEM (agentic memory that organizes itself) both build on the stream-plus-synthesis idea this paper established.
🎮 Believable NPCs
Games and interactive fiction picked up the recipe: NPCs who remember the player, form opinions, and coordinate — instead of scripted dialogue trees.
🔬 Social simulation
Sociologists and economists used generative agents to prototype social systems and test theories — simulated focus groups and synthetic populations.
💡 The core lesson
Environment + memory = agent. A plain LLM plus a well-designed record-keeping loop already produces believable, consistent, social behavior — no new model weights required.
Deep Dive

The Retrieval Formula Behind a Believable Town

Strip away the Sims aesthetic and the 25 agents run on one equation: score = relevance + recency + importance. What made the Smallville experiment science was the ablation — remove any term and human raters found the agents significantly less believable. The formula, not the LLM alone, is the architecture.

🧮
Three Signals, One Score
  • Relevance — embedding similarity to the current situation: what is useful now
  • Recency — exponential decay, 0.995 per game-hour: what is fresh
  • Importance — the LLM scores each memory 1–10 at write time: what mattered
  • All weights α = 1 — even, un-tuned, and each one load-bearing in the ablation
🧪
What the Ablations Proved
  • Full architecture: believability μ = 29.89 — every removal scored significantly lower
  • Drop reflection and agents lose long-range reasoning ("why" chains)
  • Drop planning and behavior degrades into reactive drifting
  • Emergence demo: from one seed mention, 13/25 agents learned of the party — 5 of 12 invitees attended, unprompted by any user
Interactive Demo — Score the Memory Stream

A real-shaped retrieval moment: Klaus (a pharmacy student) is asked to plan a study group. Watch every memory get scored on the three terms, watch the composite rank them, and watch the top three flow into context. Reflection triggers at importance-sum ≥ 150 — keep an eye on the running total.

SITUATION · "Klaus is organizing a study group for the chemistry midterm."
importance running total: 0 / 150 to reflection
VERDICT
Believability is an information-retrieval property
The uncomfortable reading of the ablations: the LLM was constant across conditions — only what it was shown changed, and humans could tell. Memory selection is where agent personality actually lives. The lineage runs forward through this whole track: MemGPT replaces the scoring function with an OS-style manager, A-MEM lets notes evolve, and LongMemEval finally measures whether any of it works at scale.
⏱️ The 0.995 decay
A memory from 10 game-hours ago retains ~95% of its recency score; from 100 hours, ~61%. Gentle decay — enough to prefer freshness without amnesia.
💭 Reflections cite evidence
A reflection ("Klaus values academic community") links to the memories it came from — provenance-chained synthesis, the idea A-MEM later generalized into note evolution.
🎲 Emergence, carefully
13/25 agents hearing about a party looks like magic, but each hop is one retrieval + one conversation. The magic is that nobody scripted the hops.
🪞 Humans as ground truth
Believability was rated by humans on the agents' interview answers — the evaluation pattern later benchmarks made automatic with LLM judges.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the Generative Agents paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ 25 agents in a Sims-like town (Smallville), each backed by gpt-3.5-turbo with identity, relationships, and memory.
✅ The 4-layer architecture: memory stream → retrieval → reflection → planning + reaction.
✅ Retrieval score = relevance + recency + importance (all α = 1; recency decays at 0.995 per game-hour; importance is LLM-scored 1–10).
✅ Reflections trigger at an importance-sum threshold (150), cite their evidence, and re-enter the stream as memories.
✅ Emergence: from 1 seed, 13/25 agents (52%) learned of the party and 5 of 12 invited attended — no user intervention.
✅ Humans ranked the full architecture most believable (μ = 29.89) — every ablation significantly reduced believability.