A visual, step-by-step guide to the paper that gave LLM agents a Zettelkasten — memory that writes structured notes, links them into a growing knowledge network, and rewrites older notes when new experience arrives (Xu et al., Rutgers — NeurIPS 2025).
Agent memory has been through three eras — store everything, manage pages, and now self-organize. The arc starts in a paper slip-box from the 1950s.
A slip-box is powerful because of organization, not storage. Every new note is written in the context of the notes it relates to, and older notes are re-interpreted when new ones arrive. A-MEM turns exactly this loop into an algorithm for LLM agents.
LLM agents can call tools and plan, but their memory systems mostly just store and retrieve. The structure is fixed at design time — and nothing inside the store ever reorganizes.
A-MEM's design follows the Zettelkasten method. Three slip-box principles become three mechanisms — starting with note construction.
One LLM call with a construction prompt turns a raw sentence into a structured, retrievable, linkable note. In the experiments the text encoder is all-minilm-l6-v2, and top-k retrieval uses k=10 by default.
When a note joins the network, A-MEM shortlists neighbors with embedding similarity — then lets the LLM decide which connections actually mean something.
Embedding retrieval is a cheap first filter that scales to large collections; the LLM-driven step is what catches subtle patterns, causal relationships, and conceptual connections that raw similarity misses.
Related notes interconnect through their similar contextual descriptions — like the entry points of a Zettelkasten. One note can live in multiple boxes at once. When any note in a box is retrieved, the rest of the box is automatically accessible too.
A query is embedded with the same encoder, the top-k most similar notes are retrieved — and then linked notes in the same box are automatically accessed. One seed hit plus hops along links: multi-hop reach without re-reading the whole history.
The key differentiator: adding a note can update the notes it links to. Their contextual descriptions, keywords, and tags get rewritten — so the whole network stays coherent as experience accumulates.
The LLM sees the new note, the surrounding neighborhood, and the old note — then decides whether the old note's context, keywords, and tags should change to absorb the new evidence.
A note written in January ("user is lactose-intolerant") can quietly become incomplete by June. Evolution folds new evidence into old notes, letting the system discover higher-order patterns across memories — the paper's words: "mimicking human learning processes". Memory becomes a living network, not a static store.
On LoCoMo's long conversations (avg. ~9K tokens, up to 35 sessions, 7,512 QA pairs across 5 question types) and DialSim's multi-party TV-show dialogues, A-MEM beats ReadAgent, MemoryBank, MemGPT, and full-context baselines across six foundation models.
| Backbone | Full-context (LoCoMo) | Best memory baseline | A-MEM | Gain |
|---|---|---|---|---|
| GPT-4o-mini | 25.02 | 26.65 (MemGPT) | 27.02 | +0.4 · multi-hop 45.85 vs 25.52 |
| GPT-4o | 28.00 | 30.36 (MemGPT) | 32.86 | +2.5 · multi-hop 39.41 vs 17.29 |
| Qwen2.5-1.5B | 9.05 | 11.14 (MemoryBank) | 18.23 | +7.1 · 1.6× |
| Qwen2.5-3B | 4.61 | 5.07 (MemGPT) | 12.57 | +8.0 · 2.7× |
| Llama 3.2-1B | 11.25 | 13.18 (MemoryBank) | 19.06 | +5.9 · 1.4× |
| Llama 3.2-3B | 6.88 | 6.19 (MemoryBank) | 17.44 | +10.6 · 2.5× |
A-MEM ranks first on every category for all four open 1–3B models (method ranking 1.0) — and first on average for both GPT models, with the margin concentrated in multi-hop questions.
Link generation is the foundation; memory evolution adds the refinements that push multi-hop reasoning from 31.24 to 45.85 — the two modules are complementary by design.
The four open 1–3B models see 1.4–2.7× average gains, while on GPT-4o-mini/-4o the full-context and MemGPT baselines stay competitive on single-hop and adversarial questions (strong parametric knowledge). Interpretation: a self-organized note network substitutes for parametric knowledge — the weaker the backbone, the more structure pays off.
A-MEM reframed agent memory from a storage problem to an organization problem — and connected 70 years of knowledge-management practice to LLM architecture.
A-MEM imports a 70-year-old idea from personal knowledge management — the Zettelkasten: atomic, linked notes that grow smarter as you add more. The twist that matters for agents: when a new note arrives, older notes' descriptions are rewritten in light of it. Memory stops being a log and becomes an interpretation.
Check your understanding of the key concepts from the A-MEM paper.
Everything you need to remember about this paper.