History Problem Zettelkasten Linking Evolution Results Impact Deep Dive Quiz
Interactive Paper Explainer

Notes That Think Back
A-MEM

A visual, step-by-step guide to the paper that gave LLM agents a Zettelkasten — memory that writes structured notes, links them into a growing knowledge network, and rewrites older notes when new experience arrives (Xu et al., Rutgers — NeurIPS 2025).

Start Learning Read the Paper ↗
3
Agentic Memory Mechanisms
6
Foundation Models Tested
93%
Fewer Tokens (up to)
2025
NeurIPS Publication
History

From Paper Slips to Living Memory

Agent memory has been through three eras — store everything, manage pages, and now self-organize. The arc starts in a paper slip-box from the 1950s.

1950s → 1990s
Zettelkasten (Niklas Luhmann)
The sociologist files ~90,000 numbered index cards — one atomic idea each, cross-referenced by hand. He calls the slip-box his "conversation partner"; it helps produce 70+ books.
2020
RAG (Lewis et al.)
Chunk documents, embed them, retrieve the top-k by similarity at query time. Great for static knowledge — but the store never reorganizes itself.
2023 · Apr
Generative Agents (Park et al.)
Agent memory becomes a stream: every observation is appended and scored by recency, importance, and relevance, with "reflection" synthesizing higher-level thoughts.
2023 · Oct
MemGPT (Packer et al.)
The LLM context becomes an operating system: the model itself pages memory between a small main context and external storage.
2024
Memory becomes a research field
Mem0 adds graph layers to agent memory, and benchmarks like LoCoMo and LongMemEval start measuring long-term conversational memory.
2025 · Feb
🚀 A-MEM (Xu et al., Rutgers)
Memory writes itself: notes with LLM-extracted attributes, self-generated links, and older notes that get rewritten as new ones arrive. Published at NeurIPS 2025.
Key Insight

A slip-box is powerful because of organization, not storage. Every new note is written in the context of the notes it relates to, and older notes are re-interpreted when new ones arrive. A-MEM turns exactly this loop into an algorithm for LLM agents.

THE A-MEM LOOP FOR EVERY NEW MEMORY
new memory → ① CONSTRUCT note → ② LINK to neighbors → ③ EVOLVE old notes → repeat
No predefined schema, no fixed workflow — the structure emerges from the content.
Chapter 01

The Problem with Static Memory

LLM agents can call tools and plan, but their memory systems mostly just store and retrieve. The structure is fixed at design time — and nothing inside the store ever reorganizes.

🗄️
Static Memory Stores
  • Developers must predefine storage structure, write points, and retrieval timing
  • Memories are written once and never re-organized — no consolidation
  • Relationships exist only as embedding similarity, computed at query time
  • Graph-database variants still rely on predefined schemas and relation types
  • Fixed operations generalize poorly across diverse tasks and long horizons
🕸️
A-MEM's Agentic Memory
  • The LLM itself extracts each note's keywords, tags, and contextual description
  • Links are decided at write time and stored — a persistent, growing network
  • New memories can rewrite older notes' descriptions and attributes
  • No predefined schema: structure emerges from the content itself
  • One adaptive mechanism across tasks, without workflow changes
Analogy — The Library That Never Re-Catalogs
SYSTEM A — THE STATIC LIBRARY

Every book is dropped on a shelf in arrival order; the catalog is written once and never touched again. To answer a question you either read the whole library (a full context window) or trust a keyword match that cannot see across topics. Nothing inside ever changes.

SYSTEM B — THE RESEARCHER'S SLIP-BOX

Each new slip is filed next to related slips, arrows are drawn between them, and an old note's summary is rewritten when a new slip changes what it means. To answer a question: start anywhere and follow the thread.

A-MEM is System B. Plain vector stores — and even graph schemas frozen at design time — are System A. The paper's critique: memory systems need "sophisticated memory organization", not just storage and retrieval.

Chapter 02

The Zettelkasten Blueprint — Atomic Notes

A-MEM's design follows the Zettelkasten method. Three slip-box principles become three mechanisms — starting with note construction.

1️⃣ Atomic
One idea per slip, self-contained, so any note can be retrieved on its own. → Note construction: the LLM distills each observation into a structured note.
2️⃣ Linked
Slips reference each other; meaning lives in the web, not the slips. → Link generation: the LLM decides which existing notes relate to the new one.
3️⃣ Evolving
Old slips get rewritten as understanding grows. → Memory evolution: related old notes are updated when a new note arrives.
mi = { ci, ti, Ki, Gi, Xi, ei, Li }
c
Content
The raw interaction text, kept verbatim.
t
Timestamp
When the interaction happened.
K
Keywords
LLM-extracted key concepts of the memory.
G
Tags
LLM-generated labels for categorization.
X
Contextual description
The LLM's own summary of what the memory means — the piece that later evolves.
e
Embedding
Dense vector of concat(c, K, G, X) for fast similarity search.
L
Links
The set of other notes this note connects to.
What a Constructed Note Looks Like
OBSERVATION  "User asked for a quiet restaurant for dinner"
──────────────────────────────────────────────────
c  User asked for a quiet restaurant for dinner
t  2025-06-06 18:42
K  [quiet, restaurant, dinner]
G  [dining, preference]
X  User prefers calm, low-noise dining venues; avoids loud or crowded restaurants.
e  fenc(concat(c, K, G, X))  → dense vector
L  [ ]  ← filled in by link generation…

One LLM call with a construction prompt turns a raw sentence into a structured, retrievable, linkable note. In the experiments the text encoder is all-minilm-l6-v2, and top-k retrieval uses k=10 by default.

Interactive Demo — Build the Slip-Box (Centerpiece)

Feed four observations into an empty memory, one at a time. Watch each note get constructed (attribute chips), linked (green connectors), and — on the last note — watch an older note get rewritten.

① construction → ② linking → ③ evolution
The slip-box is empty. Add observations one by one.
Chapter 03

Link Generation — the Web Decides

When a note joins the network, A-MEM shortlists neighbors with embedding similarity — then lets the LLM decide which connections actually mean something.

How a New Note Finds Its Neighbors
new note mn
① similarity  sn,j = en·ej / (|en||ej|)  — cosine over all notes
② shortlist  top-k nearest notes (k = 10 in the experiments)
③ LLM judgment  prompt: do these notes share attributes or context?
④ write links  Ln updated — the note joins one or more "boxes"

Embedding retrieval is a cheap first filter that scales to large collections; the LLM-driven step is what catches subtle patterns, causal relationships, and conceptual connections that raw similarity misses.

"Boxes" — Zettelkasten Entry Points

Related notes interconnect through their similar contextual descriptions — like the entry points of a Zettelkasten. One note can live in multiple boxes at once. When any note in a box is retrieved, the rest of the box is automatically accessible too.

Retrieval — Query, Then Follow the Thread

A query is embedded with the same encoder, the top-k most similar notes are retrieved — and then linked notes in the same box are automatically accessed. One seed hit plus hops along links: multi-hop reach without re-reading the whole history.

Interactive Demo — Static RAG vs A-MEM Retrieval Race

Same memory store, same query. The left column retrieves the top-3 chunks by cosine similarity only; the right column retrieves one note and follows its links. Run it and see what each system puts in the prompt.

QUERY: "Plan a dinner for the user"
STATIC RAG — top-3 by similarity
A-MEM — retrieve + follow links
Chapter 04

Memory Evolution — Notes That Rewrite Themselves

The key differentiator: adding a note can update the notes it links to. Their contextual descriptions, keywords, and tags get rewritten — so the whole network stays coherent as experience accumulates.

The Evolution Step
for each neighbor mj of the new note mn:
  mj* ← LLM( mn ∥ other neighbors ∥ mj )
  evolved note mj* replaces the original mj in memory

The LLM sees the new note, the surrounding neighborhood, and the old note — then decides whether the old note's context, keywords, and tags should change to absorb the new evidence.

Why This Matters

A note written in January ("user is lactose-intolerant") can quietly become incomplete by June. Evolution folds new evidence into old notes, letting the system discover higher-order patterns across memories — the paper's words: "mimicking human learning processes". Memory becomes a living network, not a static store.

Evidence It Works
  • Ablation on GPT-4o-mini: removing evolution drops multi-hop F1 from 45.85 → 31.24; removing links too drops it to 24.55
  • t-SNE plots show A-MEM notes forming coherent clusters vs a dispersed baseline without links/evolution
  • As more memories arrive, the network develops "increasingly sophisticated knowledge structures"
Interactive Demo — Note Evolution Viewer

Pick a note and read its version history. Each time a related observation arrives, the LLM rewrites the note's contextual description — highlighted phrases show what changed.

Chapter 05

Results — Self-Organization Pays Off

On LoCoMo's long conversations (avg. ~9K tokens, up to 35 sessions, 7,512 QA pairs across 5 question types) and DialSim's multi-party TV-show dialogues, A-MEM beats ReadAgent, MemoryBank, MemGPT, and full-context baselines across six foundation models.

GPT-4o-mini · MULTI-HOP F1
45.85
vs 25.52 (MemGPT) and 18.41 (full context)
≈2× the best baseline where reasoning chains matter
QWEN2.5-3B · AVG F1
12.57
vs 4.61 for the full-context baseline
a 2.7× gain — the biggest relative jump of the six
DIALSIM · F1
3.45
+35% over LoCoMo (2.55), +192% over MemGPT (1.18)
best on every metric: F1, BLEU-1, ROUGE-L/2, METEOR, SBERT
TOKENS PER ANSWER
2,520
vs 16,910 for full context (GPT-4o-mini)
85–93% token reduction; 1,216 with GPT-4o
Average F1 on LoCoMo, by Backbone (Table 1)
BackboneFull-context (LoCoMo)Best memory baselineA-MEMGain
GPT-4o-mini25.0226.65 (MemGPT)27.02+0.4 · multi-hop 45.85 vs 25.52
GPT-4o28.0030.36 (MemGPT)32.86+2.5 · multi-hop 39.41 vs 17.29
Qwen2.5-1.5B9.0511.14 (MemoryBank)18.23+7.1 · 1.6×
Qwen2.5-3B4.615.07 (MemGPT)12.57+8.0 · 2.7×
Llama 3.2-1B11.2513.18 (MemoryBank)19.06+5.9 · 1.4×
Llama 3.2-3B6.886.19 (MemoryBank)17.44+10.6 · 2.5×

A-MEM ranks first on every category for all four open 1–3B models (method ranking 1.0) — and first on average for both GPT models, with the margin concentrated in multi-hop questions.

Ablation — Where the Gains Come From (GPT-4o-mini, multi-hop F1)
no links, no evolution
24.55
links only (w/o evol.)
31.24
full A-MEM
45.85

Link generation is the foundation; memory evolution adds the refinements that push multi-hop reasoning from 31.24 to 45.85 — the two modules are complementary by design.

Backbone Analysis — Smaller Models Gain More

The four open 1–3B models see 1.4–2.7× average gains, while on GPT-4o-mini/-4o the full-context and MemGPT baselines stay competitive on single-hop and adversarial questions (strong parametric knowledge). Interpretation: a self-organized note network substitutes for parametric knowledge — the weaker the backbone, the more structure pays off.

Cost & Scale
  • ~1,200 tokens per memory operation; under $0.0003 per operation
  • 5.4s processing with GPT-4o-mini; 1.1s with local Llama 3.2-1B
  • Storage grows O(N); retrieval 0.31μs at 1K notes → 3.70μs at 1M notes
What A-MEM Did NOT Solve
Legacy

Impact — Memory That Writes Itself

A-MEM reframed agent memory from a storage problem to an organization problem — and connected 70 years of knowledge-management practice to LLM architecture.

🧠 Self-organizing memory
Memory becomes an active process: notes generate their own context, links, and updates — no predefined schema, no fixed operations in the workflow.
🗂️ The Zettelkasten bridge
A 20th-century note-taking method became a 2025 architecture — evidence that human knowledge-management practice is a rich design space for agents.
🕸️ Note-graph retrieval
Retrieval means following the thread: linked "boxes" give multi-hop access to related memories without re-reading entire histories.
📈 2025's memory ecosystem
A-MEM lands amid Mem0, MemGPT, MemoryBank, and a wave of surveys and benchmarks — the year agent memory became its own research field.
🤏 Small-model empowerment
1.4–2.7× gains on 1–3B backbones: structured external memory substitutes for scale, making capable agents cheaper to run.
🏭 Production path
A production-ready A-MEM ships inside AIOS (the LLM-agent operating system), and the benchmark evaluation code is open-source on GitHub.
Deep Dive

Notes That Rewrite Themselves

A-MEM imports a 70-year-old idea from personal knowledge management — the Zettelkasten: atomic, linked notes that grow smarter as you add more. The twist that matters for agents: when a new note arrives, older notes' descriptions are rewritten in light of it. Memory stops being a log and becomes an interpretation.

🗂️
What a Note Knows About Itself
  • Every note = {content, timestamp, keywords, tags, contextual description, embedding, links}
  • Attributes extracted by the LLM itself at write time — no hand schema
  • Links are LLM decisions over embedding-shortlisted neighbors; "boxes" group related notes
  • The rewrite loop: new evidence arrives → older descriptions update → retrieval quality compounds
🧾
What It Costs, and What It Risks
  • Every write triggers attribute extraction + link decisions + possible rewrites — <$0.0003 per op, but ops multiply
  • Evolution is a bet: rewriting can also drift notes away from what was actually said
  • Retrieval still rides embeddings — the semantic-shortcut limits LoCoMo-Plus exposes remain upstream
  • Links accumulate: neighborhood context grows with the store — curation is unsolved at million-note scale
Interactive Demo — Watch a Memory Evolve

Three notes arrive over three sessions. Watch the first note's contextual description rewrite itself as evidence accumulates, watch links form and thicken, and then run the multi-hop query that needs all three notes at once — the retrieval pattern where A-MEM beat ReadAgent, MemoryBank, MemGPT and full-context across six backbones.

VERDICT
Evolution is the differentiator; efficiency is the by-product
Multi-hop F1 of 45.85 with GPT-4o-mini — and 2.7× average improvement on a 3B backbone — while spending up to 93% fewer tokens per answer: because the notes already did the synthesis, the model reads a curated neighborhood instead of a transcript. O(N) storage keeps the bet affordable. The Zettelkasten lineage is explicit design honesty: the paper cites the method's human origin. Where the lineage continues: Generative Agents (proto-evolution via reflection), MemGPT (capacity management), and the benchmarks that grade it: LongMemEval and LoCoMo-Plus.
🔗 Links as retrieval pre-computation
A multi-hop query needs notes that do not resemble each other. Links are the bridge infrastructure embeddings cannot build — relation decided at write time, exploited at read time.
✍️ The description is the index
Raw content stays immutable; the contextual description rewrites. That separation — evidence preserved, interpretation updated — is what makes evolution auditable.
📦 Boxes
Zettelkasten "boxes" surface each other at retrieval: related neighborhoods arrive together. Structured browsing, not just nearest-neighbor.
💰 $0.0003 per operation
Memory management priced at production scale: a million-note lifetime of operations costs less than one frontier-model page of context. The economics complete the argument.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the A-MEM paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ A-MEM (Xu et al., Rutgers — NeurIPS 2025) gives LLM agents Zettelkasten-style memory: atomic, linked, evolving notes.
✅ Every note = {content, timestamp, keywords, tags, contextual description, embedding, links} — attributes extracted by the LLM itself.
✅ Links are LLM decisions over embedding-shortlisted neighbors at write time; "boxes" group related notes and surface each other at retrieval.
✅ Memory evolution rewrites older notes' descriptions as new evidence arrives — the differentiator vs static stores.
✅ Beats ReadAgent, MemoryBank, MemGPT and full-context on LoCoMo across 6 backbones (multi-hop F1 45.85 on GPT-4o-mini; 2.7× average on Qwen2.5-3B) and on DialSim.
✅ Up to 93% fewer tokens per answer, under $0.0003 per memory operation, O(N) storage — practical at million-note scale.