A visual, step-by-step guide to "A Survey on the Memory Mechanism of Large Language Model based Agents" — the 2024 field map by Zhang et al. that organized where agent memory lives, what it stores, and which operations keep it useful.
Agent memory was invented over and over, paper by paper — until this survey drew the map. Here is the road that made the map necessary.
An LLM's weights already remember — pre-training is a kind of memory. What agents add is experience: failures, preferences, and skills accumulated across interactions. The survey opens with Elie Wiesel: "Without memory, there is no culture."
By 2024, dozens of papers had built agent memory — each from scratch, each in its own words. Promising designs were scattered and incomparable.
Imagine a zoo where every keeper invents their own name for every animal — "the striped horse", "the river wolf", "the long-nosed grey giant". Each description is accurate, yet nobody can tell whether two keepers are describing the same species, or which animals thrive in which enclosure. Research on agent memory was that zoo. This survey is the field guide: one classification — sources (what it feeds on), forms (its body plan), and operations (how it behaves) — that finally lets the keepers compare notes.
The survey's first big cut: memory can live outside the weights as text, or inside the weights as parameters. Everything else branches from there.
Information is kept explicitly, in natural language or structured records (tuples, databases). It is what most systems actually use: interpretable, easy to edit, fast to write — but every read costs prompt tokens.
Information is encoded into the model parameters — via fine-tuning (bake knowledge in with SFT/LoRA) or knowledge editing (surgically change specific facts). Reading is free at inference; writing is slow, and the survey calls this direction under-researched.
Counted from the survey's Table 2 (28 systems; a system can appear in several columns). Retrieval wins; editing barely exists yet.
The survey motivates memory with cognitive psychology (§4.1). The field often borrows that lens — working, episodic, semantic, procedural — to describe what an agent remembers. Learn all four; they are everywhere in agent papers.
The active scratchpad — whatever is in the prompt right now. Small, fast, and constantly overwritten.
Specific past events with a "when" attached — what happened, to whom, in which session.
General facts, stripped of time and place — "Paris is the capital of France".
Skills and policies — how to act, not what happened. Learned from experience, applied automatically.
This four-way vocabulary is the cognitive-science lens, not the survey's own scheme. The survey cites cognitive psychology as motivation — "following human's working mechanisms to design the agents is a natural and essential choice" (§4.1) — and covered systems borrow its words (MemGPT's "working context", RecAgent's "short-term memory"). But its organizing taxonomy is sources → forms → operations. Use the cognitive lens to build intuition; use the survey's axes to compare systems.
The survey splits memory into three operations — writing, management (merging + reflection + forgetting), and reading. Real systems are pipelines that compose them.
| System | Sources | Form | Signature Operation |
|---|---|---|---|
| Generative Agents (2023) | inside-trial | textual — retrieved | reflection: events → higher-level thoughts; retrieval = relevance × importance × recency |
| MemGPT (2023) | inside-trial | textual — recent + retrieved | OS-style virtual context management: the agent pages its own memory in and out |
| MemoryBank (2023) | inside-trial | textual — retrieved | Ebbinghaus-inspired forgetting + daily summary insights (all five operations ✓, Table 3) |
| ChatDB (2023) | inside + external | textual — retrieved (symbolic) | reading via agent-generated SQL — a "Chain-of-Memory" |
| Voyager (2023) | inside-trial + external | textual — retrieved | environment-feedback-driven refinement of a growing skill library |
| ExpeL (2023) | inside + cross-trial + external | textual — complete + retrieved + external | distills cross-trial insights from the top-K most similar successful trajectories |
Every fact above is read off the survey's Tables 1–3 and its representative-studies text. MemoryBank and RecAgent are the only two systems with a ✓ in all five operation columns.
The survey ends where the real work begins. Six named open problems — plus one honest admission about what a survey cannot do.
A survey's product is shared structure. This one gave agent memory a vocabulary, a to-do list, and a measuring stick.
The survey's most durable idea is that every memory system, however branded, reduces to three operations on two storage forms — and that systems differ mainly in which operations they invest in. Once you see the lifecycle, "memory" stops being a feature and becomes a design surface with named coordinates.
Check your understanding of the survey's taxonomy, operations, and open challenges.
Everything you need to remember about this survey.