History Problem Taxonomy Types Operations Challenges Impact Deep Dive Quiz
Interactive Paper Explainer

How Agents Remember
The Memory Survey

A visual, step-by-step guide to "A Survey on the Memory Mechanism of Large Language Model based Agents" — the 2024 field map by Zhang et al. that organized where agent memory lives, what it stores, and which operations keep it useful.

Start Learning Read the Paper ↗
2
Memory Forms
3
Memory Operations
28
Systems Compared
174
References Cited
History

From Blank Slate to Memory Streams

Agent memory was invented over and over, paper by paper — until this survey drew the map. Here is the road that made the map necessary.

2020
Two kinds of knowing
GPT-3 shows that pre-training bakes world knowledge into weights — parametric memory. RAG (Lewis et al.) pairs those weights with retrieved documents — non-parametric memory.
2022
The blank-slate problem
ChatGPT goes mainstream — and every new session starts from zero. Anything outside the context window is simply gone. The memory gap becomes impossible to ignore.
2023 · Apr
Generative Agents
Park et al.'s simulated town gives each agent a memory stream: retrieval scored by relevance × importance × recency, plus reflection that turns events into higher-level thoughts.
2023 · Oct
MemGPT
Packer et al. treat the LLM like an operating system: a small "main context" plus external storage, with the agent itself paging memories in and out.
2024 · Apr
🚀 This survey
Zhang et al. (Renmin University of China + Huawei Noah's Ark Lab) map 28 memory systems across sources, forms, and operations — and name the open problems. Later accepted to ACM TOIS (2025).
2024–25
The gap-fillers
LongMemEval (ICLR 2025) benchmarks long-term memory with 500 questions; A-MEM (NeurIPS 2025) makes the memory system itself agentic.
Key Insight

An LLM's weights already remember — pre-training is a kind of memory. What agents add is experience: failures, preferences, and skills accumulated across interactions. The survey opens with Elie Wiesel: "Without memory, there is no culture."

THE SURVEY'S TWO DEFINITIONS
narrow = history of the current trial
broad = + past trials + external knowledge
A "trial" is one full attempt at a task; a "step" is one action–observation turn.
39
pages long
4
comparison tables
28
memory systems mapped
174
references cited
Chapter 01

The Problem with Ad Hoc Memory

By 2024, dozens of papers had built agent memory — each from scratch, each in its own words. Promising designs were scattered and incomparable.

🧩
A Field of One-Off Designs
  • Every agent paper invents its own memory mechanism from scratch
  • Designs scattered across papers, described in incompatible vocabularies
  • No shared answer to "where does memory live?" — prompt, file, or weights?
  • No fair comparison: is MemGPT's memory "better" than MemoryBank's? At what?
  • Hard-won design patterns stay buried inside individual systems
🗺️
What One Survey Fixes
  • A single taxonomy: memory sources, memory forms, memory operations
  • 28 systems placed side by side in 4 comparison tables
  • Shared vocabulary: write · manage · read; textual vs. parametric
  • Common design patterns abstracted for future memory builders
  • Open challenges named — including the missing memory benchmark
Analogy — Zookeepers Without a Species Catalog

Imagine a zoo where every keeper invents their own name for every animal — "the striped horse", "the river wolf", "the long-nosed grey giant". Each description is accurate, yet nobody can tell whether two keepers are describing the same species, or which animals thrive in which enclosure. Research on agent memory was that zoo. This survey is the field guide: one classification — sources (what it feeds on), forms (its body plan), and operations (how it behaves) — that finally lets the keepers compare notes.

Chapter 02

Where Memory Lives

The survey's first big cut: memory can live outside the weights as text, or inside the weights as parameters. Everything else branches from there.

Agent Memory  Mt = f ( ξt , Ξ , D )
ξt
Inside-trial
Everything the agent has seen and done in the current attempt — the "narrow" memory.
Ξ
Cross-trial
Experiences from earlier attempts and earlier tasks — including the failures.
D
External knowledge
Static sources outside the interaction loop — files, databases, wikis.
f
Operations
Writing, management, and reading — the verbs that turn sources into usable memory.
📝 Textual Form — memory as words

Information is kept explicitly, in natural language or structured records (tuples, databases). It is what most systems actually use: interpretable, easy to edit, fast to write — but every read costs prompt tokens.

complete history recent window retrieved entries external knowledge
⚙️ Parametric Form — memory as weights

Information is encoded into the model parameters — via fine-tuning (bake knowledge in with SFT/LoRA) or knowledge editing (surgically change specific facts). Reading is free at inference; writing is slow, and the survey calls this direction under-researched.

fine-tuning knowledge editing hybrid: mix both
Interactive Demo — The Memory Taxonomy Tree

Click any leaf to see what it stores, which real agents use it, and the trade-offs the survey documents (Table 2).

What the 28 Surveyed Systems Actually Use (Table 2)
RETRIEVED INTERACTIONS
17/28
61% of systems — the mainstream
COMPLETE HISTORY
11/28
39% keep the whole trajectory
EXTERNAL KNOWLEDGE
7/28
25% consult outside sources
RECENT WINDOW
5/28
18% cache only the newest
FINE-TUNING
4/28
14% write into the weights
KNOWLEDGE EDITING
1/28
4% — only MAC in 2024

Counted from the survey's Table 2 (28 systems; a system can appear in several columns). Retrieval wins; editing barely exists yet.

Chapter 03

Four Kinds of Remembering

The survey motivates memory with cognitive psychology (§4.1). The field often borrows that lens — working, episodic, semantic, procedural — to describe what an agent remembers. Learn all four; they are everywhere in agent papers.

⚡ Working Memory

The active scratchpad — whatever is in the prompt right now. Small, fast, and constantly overwritten.

Agent example: MemGPT's "working context" — recent history held in main context.
survey term: inside-trial
📅 Episodic Memory

Specific past events with a "when" attached — what happened, to whom, in which session.

Agent example: Generative Agents' memory stream logs each event as it happens.
survey term: cross-trial
📚 Semantic Memory

General facts, stripped of time and place — "Paris is the capital of France".

Agent example: Huatuo fine-tunes medical knowledge into weights; GITM consults a knowledge base.
survey term: external / parametric
🛠️ Procedural Memory

Skills and policies — how to act, not what happened. Learned from experience, applied automatically.

Agent example: Voyager's skill library of verified code; Reflexion's distilled lessons.
survey term: cross-trial, distilled
⚠️ Label Carefully — What the Survey Actually Adopts

This four-way vocabulary is the cognitive-science lens, not the survey's own scheme. The survey cites cognitive psychology as motivation — "following human's working mechanisms to design the agents is a natural and essential choice" (§4.1) — and covered systems borrow its words (MemGPT's "working context", RecAgent's "short-term memory"). But its organizing taxonomy is sources → forms → operations. Use the cognitive lens to build intuition; use the survey's axes to compare systems.

Interactive Demo — Which Memory Type?

Five real-ish agent moments. Pick the cognitive memory type each one exercises — the sorter keeps score and explains every verdict.

Chapter 04

The Memory Lifecycle

The survey splits memory into three operations — writing, management (merging + reflection + forgetting), and reading. Real systems are pipelines that compose them.

✍️ Memory Writing
store("429 on POST /orders")
Decide what to keep: raw observations or summaries. TiM extracts entity relations; MemoChat stores topic keys; MemGPT lets the agent edit its own memory.
🔗 Merging
merge(S1, M1) → R1
Fold redundant entries together. GITM summarizes key actions across many plans into common reference plans; Voyager refines its skill library from environment feedback.
💭 Reflection
reflect(events) → insight
Generate higher-level thoughts from accumulated events. Generative Agents triggers reflection when enough events pile up; MemoryBank distills daily insights and personality takeaways.
🌫️ Forgetting
decay(P2, t) → forgotten
Drop the stale and irrelevant. MemoryBank updates memories with an Ebbinghaus-inspired forgetting curve; knowledge editing can also remove "bad memory" from weights.
🔍 Memory Reading
read(query) → top-K
Fetch what matters for the current state. ChatDB reads by generating SQL ("Chain-of-Memory"); ExpeL pulls the top-K most similar successful trajectories from a FAISS store.
Interactive Demo — The Memory Lifecycle (stepwise)

An agent hits a 429 error while placing Alice's order. Press Play and watch one memory chip travel the full pipeline: written, reflected into an insight, retrieved on demand — while a stale memory decays away.

How Six Landmark Systems Compose Memory (survey Tables 1–3)
SystemSourcesFormSignature Operation
Generative Agents (2023)inside-trialtextual — retrievedreflection: events → higher-level thoughts; retrieval = relevance × importance × recency
MemGPT (2023)inside-trialtextual — recent + retrievedOS-style virtual context management: the agent pages its own memory in and out
MemoryBank (2023)inside-trialtextual — retrievedEbbinghaus-inspired forgetting + daily summary insights (all five operations ✓, Table 3)
ChatDB (2023)inside + externaltextual — retrieved (symbolic)reading via agent-generated SQL — a "Chain-of-Memory"
Voyager (2023)inside-trial + externaltextual — retrievedenvironment-feedback-driven refinement of a growing skill library
ExpeL (2023)inside + cross-trial + externaltextual — complete + retrieved + externaldistills cross-trial insights from the top-K most similar successful trajectories

Every fact above is read off the survey's Tables 1–3 and its representative-studies text. MemoryBank and RecAgent are the only two systems with a ✓ in all five operation columns.

Chapter 05

What's Still Unsolved

The survey ends where the real work begins. Six named open problems — plus one honest admission about what a survey cannot do.

🤔 What to write
Which observations deserve storage? Raw streams are long and noisy; over-summarizing throws away detail. The survey calls information extraction "vital" (§5.3.1).
🕰️ When to forget
Forgetting protects relevance but risks losing what matters. When to decay, merge, or drop a memory — and how to measure the damage — is open (§5.3.2).
📏 No memory benchmark
"To our knowledge, there are no open-sourced benchmarks tailored for the memory modules in LLM-based agents" (§6.3). Direct evaluation stays mostly subjective — coherence and rationality judged by humans.
⚙️ Parametric friction
Writing to weights is slow and data-hungry; editing risks side effects; both are hard to interpret — a problem for high-trust domains like medicine (§8.1).
👥 Multi-agent memory
Synchronized knowledge bases, memory-driven communication, and information asymmetry in competitive settings are barely explored (§8.2).
♾️ Lifelong scale
Temporality, memory overlap, and storing/retrieving across an agent's whole "lifetime" strain every current design (§8.3).
What the Survey Did NOT Solve
Legacy

Impact — Memory Becomes a Field

A survey's product is shared structure. This one gave agent memory a vocabulary, a to-do list, and a measuring stick.

🗺️ A shared vocabulary
Sources · forms · operations became the standard way to describe — and compare — agent memory designs.
📏 Benchmark blueprint
The flagged "no open-sourced benchmark" gap was filled by LongMemEval (ICLR 2025): 500 questions across 5 long-term memory abilities.
🧠 The agentic-memory lineage
A-MEM (NeurIPS 2025) attacks the fixed-structure, fixed-operations limits that this survey catalogued — memory that organizes itself.
📦 Memory as a product
Persistent "memory" features in consumer AI assistants descend directly from the designs this survey organized.
🔄 A living map
The authors maintain a companion repo (nuster1128/LLM_Agent_Memory_Survey) tracking new papers; the survey was accepted to ACM TOIS in 2025.
🎓 The design-space lesson
Where memory lives, what it stores, and which operations run on it are independent axes — choose each one deliberately, not by accident.
Deep Dive

Write, Manage, Read: Memory as a Lifecycle

The survey's most durable idea is that every memory system, however branded, reduces to three operations on two storage forms — and that systems differ mainly in which operations they invest in. Once you see the lifecycle, "memory" stops being a feature and becomes a design surface with named coordinates.

🗺️
The Three Design Axes
  • Sources — inside-trial (this conversation), cross-trial (past sessions), external (retrieved corpora)
  • Forms — textual (explicit stores) vs parametric (inside the weights)
  • Operations — write, manage, read: the lifecycle every architecture walks
  • 17 of 28 surveyed systems store retrieved interactions — retrieval is the mainstream; only MAC edits knowledge in weights
⚖️
The Two Forms, Honestly Priced
  • Parametric reads are free at inference and scale with the model — but writing is slow, risky, and opaque
  • Editing weights courts catastrophic forgetting: today's write can erase yesterday's knowledge
  • Textual memory pays a retrieval cost per read, but writes are instant, auditable, and erasable
  • Management — merging, reflection, forgetting — is where textual systems win and most systems under-invest
Interactive Demo — Walk One Memory Through the Pipeline

A single memory — "user prefers vegetarian meals" — travels the lifecycle. Step through each operation and watch what the surveyed systems actually do there: consolidation on write, merging and reflection in management, scored selection on read. The card updates as the memory transforms.

VERDICT
The survey's loudest gap became the field's next paper
The authors' pointed observation — no benchmark existed for agent memory — was answered within months by LongMemEval. That is what a good survey does: it doesn't just sort the literature, it publishes the to-do list. For the systems mapped here, read onward: MemGPT (management as an OS), A-MEM (notes that evolve), Generative Agents (the origin of scored reading).
➕ Management = merging + reflection + forgetting
Consolidation operations inspired by human sleep: merge duplicates, reflect on patterns, decay the stale. Skipping management is how memory stores become landfills.
🔍 Reading, mostly nearest-neighbor
The dominant read is retrieval by similarity — sometimes time-aware, rarely structure-aware. A-MEM and LoCoMo-Plus both attack exactly this conservatism.
🧠 Parametric memory's lonely champion
Only MAC among 28 systems edits knowledge in weights. The rest voted with their architecture: explicit text you can audit beats implicit weights you cannot.
📚 A living companion
39 pages, 4 comparison tables, and a maintained repo — the survey is organized as a reference work, which is why the taxonomy still frames papers written two years later.
Test Yourself

Quick Quiz

Check your understanding of the survey's taxonomy, operations, and open challenges.

Reference

Key Takeaways

Everything you need to remember about this survey.

✅ Agent memory has three design axes: sources (inside-trial / cross-trial / external), forms (textual / parametric), operations (write / manage / read).
✅ Management = merging + reflection + forgetting — consolidation operations inspired by how human brains abstract.
✅ Retrieval is the mainstream: 17 of 28 surveyed systems store retrieved interactions; only MAC edits knowledge in weights.
✅ Parametric memory reads for free and scales, but writing is slow, risky (catastrophic forgetting), and opaque.
✅ The survey's loudest gap — no memory benchmark — was later filled by LongMemEval (ICLR 2025): 500 questions, 5 abilities.
✅ Zhang et al., arXiv 2404.13501 (Apr 2024; ACM TOIS 2025) — 39 pages, 4 tables, and a living companion repo.