History Problem Architecture Functions Modes Results Impact Deep Dive Quiz
Interactive Paper Explainer

Paging Memory Into Context
MemGPT

A visual, step-by-step guide to the paper that treated the LLM context window as RAM — and taught the model to manage its own memory like an operating system, paging facts between a small window and unlimited "disk" via self-generated function calls.

Start Learning Read the Paper ↗
2
Memory Tiers (RAM + Disk)
4+
Self-Editing Memory Functions
93.4%
Deep Memory Retrieval Acc.
2023
Year Published (October)
History

From Fixed Windows to Virtual Memory

LLMs have always had a "RAM problem". MemGPT's answer didn't come from a new architecture — it came from operating-system design.

2017
The Transformer (Vaswani et al.)
Attention-only architecture — with fixed-length context from day one. Everything outside the window simply doesn't exist for the model.
2019–2020
Long-context attempts
Transformer-XL, Longformer and other sparse / efficient-attention variants attack the problem — but self-attention cost grows quadratically with window size.
2022
Lossy chat memory
Chat products keep rolling summaries of older turns. Details quietly vanish — a summary of a summary loses the thread, and there is no way to look the original back up.
2023 · Apr
Generative Agents (Park et al.)
A memory stream + reflection gives simulated agents long-term memory — an influential first structured memory for LLM agents, cited by the MemGPT paper.
2023 · Oct
🚀 MemGPT (Packer et al., UC Berkeley)
Virtual context management: borrow the OS's virtual-memory playbook so a fixed-context LLM can page its own memory between "RAM" and "disk" with function calls.
2024 →
Letta & the memory ecosystem
The MemGPT project grows into Letta — and agent memory (mem0, A-MEM and successors) becomes a whole product and research category.
Key Insight

Operating systems solved this exact problem decades ago: give every program the illusion of unlimited memory by transparently paging data between fast, small RAM and slow, huge disk. MemGPT applies the same trick to the LLM context window — with a twist: the model itself is the memory manager.

THE OS → MEMGPT DICTIONARY
main memory (RAM) → main context
disk → external context
virtual memory → virtual context management
page eviction → queue flush + summary
memory pressure → ~70% context warning
interrupts → events that trigger inference
The paper's own name for the technique: "virtual context management".
Chapter 01

Context Windows Are Tiny RAM

The paper opens with a brutal observation: an LLM's "memory" is a small read buffer — while agents need something more like a filing system.

💾
The World Doesn't Fit in RAM
  • Widely-used open models of the era: ~2k–4k tokens — roughly 20–60 messages
  • Even 128k windows drown: a single SEC Form 10-K can pass the million-token mark
  • Extending attention is quadratic in compute and memory cost
  • Rolling summaries are lossy — early details silently disappear
  • Plain retrieval adds text, but can't evict, reorganize or curate context
🗄️
MemGPT's Solution: Virtual Memory
  • Two-tier hierarchy: main context ("RAM") + external context ("disk")
  • The LLM itself pages data in and out — via function calls
  • Memory-pressure warnings and queue flushes, just like an OS
  • One fixed-context model, the illusion of unbounded context
  • The same design works for endless chats and million-token documents
How Small Is "RAM"? — Table 1 of the Paper (data collected 1/2024)
Model / APIOpen?Context Tokens≈ Messages*
Llama (1)✓2k20
Llama 2✓4k60
GPT-3.5 Turbo (release)✗4k60
Mistral 7B✓8k140
GPT-4 (release)✗8k140
GPT-3.5 Turbo✗16k300
GPT-4✗32k~600
Claude 2✗100k~2,000
GPT-4 Turbo✗128k~2,600
Yi-34B-200k✓200k~4,000

*Assuming a ~1k-token preprompt and ~50 tokens per message. Even the big numbers run out: ~600 messages fills a 32k window — and a "personal assistant" is supposed to last for months or years.

The Analogy — Your Desk vs the Filing Cabinet

Think of the context window as your desk: you can only work with what's on it right now. The rest of your knowledge lives in a filing cabinet (external storage). A person with a small desk and a good filing system beats one with a huge, messy desk — as long as someone keeps filing and retrieving intelligently. MemGPT's bet: that "someone" can be the LLM itself, swapping pages in and out exactly like an OS kernel does between RAM and disk.

Chapter 02

The Architecture — Two Tiers of Memory

MemGPT wraps a fixed-context "LLM processor" in a hierarchical memory system it manages itself. Everything the model can see right now lives in main context; everything it can fetch lives outside.

main context = system instructions + working context + FIFO queue
external context = recall storage + archival storage
📖
System Instructions
Read-only (like kernel code): what MemGPT is, what each tier is for, and how to call the functions.
📝
Working Context
A fixed-size read/write block — the model's own notes about you and the task, editable only via function calls.
📬
FIFO Queue
Rolling message history. Its first slot holds a recursive summary of evicted messages.
🗂️
Recall Storage
The full message database — every message ever seen, searchable and re-insertable into the queue.
🗃️
Archival Storage
A vector database for arbitrary-length text — facts, notes, whole documents. Unbounded, searchable.
Main Context — "In RAM" (inside every LLM call)
system instructions
read-only · loaded on every call
working context
read/write · the model's self-edited notes
FIFO queue
rolling message history · evicts under pressure

Everything here counts against the fixed context-window budget.

External Context — "On Disk" (persists across calls)
recall storage
the MemGPT message database — the full chat log
archival storage
vector DB (PostgreSQL + pgvector in the experiments) for arbitrary text

Indefinitely large — reached only through function calls.

↑↓ function calls page data between the tiers — and the LLM decides what, when and why
⚠️ ~70%: Memory Pressure
When prompt tokens pass the "warning token count", the queue manager inserts a system message telling the LLM to save important queue data before eviction hits — the paper's "memory pressure" warning.
🧹 100%: Queue Flush
At the "flush token count" the queue manager evicts the oldest messages (e.g. ~50% of the window) to recall storage and writes a new recursive summary at the queue front.
🔌 Interrupts & Events
User messages, system alerts (context warnings, "document uploaded") and timed events all trigger LLM inference — MemGPT can even run unprompted on a schedule.
🔗 Function Chaining
A flag like request_heartbeat=true immediately re-invokes the processor after a function returns, so the model can chain several calls before yielding back to the user.
Interactive Demo — Context Window Simulator (the OS Paging Loop, Live)

This is MemGPT's core mechanism, made visible. The main context holds 8 message slots (the "RAM"). Send messages until it fills: older turns get paged out to external storage (amber flash). Then ask a question about something old — MemGPT searches storage and pages it back in (green flash).

MAIN CONTEXT — "RAM" (fixed 8-slot window)
⚠ warning ~70% · flush at 100%
↓ function calls page data in & out ↓
EXTERNAL CONTEXT — "DISK" (recall + archival storage)
Chapter 03

Self-Editing Memory Functions

MemGPT's LLM doesn't just chat — it operates on its own memory. The paper calls this "self-directed" editing: every store, search and eviction is a function call the model chose to make.

Write to Disk — archival_memory_insert
archival_memory_insert(
  "User is vegetarian; Rome trip in April"
)

Stores text indefinitely in archival storage. The model typically fires this right after a memory-pressure warning — deciding, on its own, what's worth keeping.

Search the Disk — archival_memory_search
archival_memory_search(
  "Rome budget hotel", page=1
)

Vector search over archival storage. Results are paginated, so a query can never overflow the context window — and the model can keep turning pages.

Search Your Own History — conversation_search
conversation_search(
  "flight arrival time"
)

Searches recall storage — the full message database. Matching messages are re-inserted at the back of the FIFO queue: a page-in of your own past.

Rewrite Your Notes — Working-Context Edits
[self-edit] "prefers window seat"
→ "prefers aisle seat"

Main context isn't static: the model updates its working-context notes as facts change — merging, correcting and reorganizing what it "knows" about you.

Who Runs the Memory? (The Twist)

In a classic RAG pipeline, a separate retriever decides what the model sees. In MemGPT, the LLM processor emits the function calls itself; a function executor validates and runs them, then feeds the results — including runtime errors — back into context. The paper's phrase: memory edits and retrieval are "entirely self-directed". If a call fails (e.g. inserting into an already-full context), the model sees the error and adapts.

Interactive Demo — Memory Function Console

Click a scenario to see the function call the LLM generates, the result it gets back, and which memory tier lights up (green). This is self-editing memory: the model, not a human, runs the memory operations.

MAIN CONTEXT (RAM)
📖 system instructions
📝 working context
📬 FIFO queue
EXTERNAL CONTEXT (DISK)
🗂️ recall storage
🗃️ archival storage
Chapter 04

Two Modes — Chat & Document Analysis

The paper evaluates the same OS-inspired system in the two domains where fixed windows hurt most: assistants that must remember you for months, and documents far bigger than any window.

Mode A — MEMGPT-CHAT

A persistent dialogue assistant with long-term memory across sessions: it remembers facts, preferences and events, reflects on them, and evolves its persona over time. Benchmarked on the Multi-Session Chat (MSC) dataset — five sessions per persona — plus two new tasks the authors introduced: deep memory retrieval and conversation openers.

SESSION 1 · "I'm vegetarian, and I'm planning a Rome trip in April"
→ archival_memory_insert("vegetarian; Rome in April")
SESSION 6 · days later · "Where did we decide to stay?"
→ conversation_search("Rome")
✓ "Trastevere — the ~$140/night B&B we found."

The chat personalization loop: save facts when the window fills, page them back when the user needs them.

Mode B — MEMGPT-DOC

A document-analysis agent that works beyond the context window: the whole corpus (in the experiments, a late-2018 dump of Wikipedia, embedded into a vector database) is loaded into archival storage, and the model pages in relevant chunks itself. Its real preprompt begins: "You are MemGPT DOC-QA bot… remember to keep searching if you can't find the answer."

FIXED-WINDOW READER
Top-K documents must fit the window. If the retriever's top-K misses the gold passage, it is simply never seen. Truncating docs to squeeze in more K degrades things further.
MEMGPT READER
The model keeps issuing archival_memory_search and turning pages — it can walk the retriever's full ranking, so the gold passage can surface even outside the top dozen.
Interactive Demo — Fixed Window vs MemGPT (A/B Recall Race)

Same model, same 4-slot window, same 9-turn conversation — only the right column manages its memory. Watch what happens to the fact mentioned in turn 1 when the final question arrives.

① FIXED WINDOW — no memory manager
DROPPED — GONE FOREVER:
② MEMGPT — pages its own memory
ARCHIVED — STILL SEARCHABLE:
Honest Notes on Doc Mode
Chapter 05

Results — Paging Wins

All numbers below are from the paper's tables. The pattern is consistent: MemGPT turns a weaker fixed-context baseline into a far stronger memory system.

GPT-3.5 TURBO (16k) → MEMGPT
66.9%
deep memory retrieval acc. — vs 38.7% fixed-context
ROUGE-L (R): 0.394 → 0.629
GPT-4 (8k) → MEMGPT
92.5%
deep memory retrieval acc. — vs 32.1% fixed-context
ROUGE-L (R): 0.296 → 0.814
GPT-4 TURBO (128k) → MEMGPT
93.4%
deep memory retrieval acc. — vs 35.3% fixed-context
ROUGE-L (R): 0.359 → 0.827
NESTED KV RETRIEVAL
4 / 4
nesting levels solved by MemGPT (GPT-4) — fixed GPT-4 hits 0% by level 3
140 UUID key-value pairs ≈ 8k tokens · multi-hop lookups via repeated function queries
DOC QA (FIG. 5)
Flat
accuracy vs. context length — unaffected as documents grow
Truncation-based baselines degrade as compression grows; MemGPT ≈ same with GPT-4 & GPT-4 Turbo
Nested KV Retrieval — Multi-Hop Lookups
SystemBehavior
GPT-3.5 (fixed)0% accuracy at just 1 nesting level
GPT-4 / GPT-4 Turbo (fixed)better, but 0% by 3 nesting levels
MemGPT (GPT-4 Turbo / 3.5)beats their baselines, but drops off at 2 levels
MemGPT (GPT-4)unaffected — solves all 4 levels via repeated function queries

"Nested KV": values may themselves be keys, so the agent must chain lookups — the paper's test of collating information across sources.

Conversation Openers — Engagement (Table 3)
MethodSIM-1SIM-3SIM-H
Human (gold opener)0.8000.8001.000
GPT-3.5 Turbo0.8300.8120.817
GPT-40.8680.8430.773
GPT-4 Turbo0.8570.8280.767

Baseline rows shown. With MemGPT layered on the same models, the paper reports openers that perform similarly to and occasionally exceeding the hand-written human openers — storing facts in working context was key.

Document QA — What Happened (Fig. 5)
  • Task: NaturalQuestions-Open questions over a late-2018 Wikipedia dump, retriever-reader setup, 50 sampled questions, LLM-judge scoring.
  • Fixed baselines are capped by the retriever: if the gold article isn't in the top-K, it's never seen.
  • Truncation lets fixed models take more documents, but accuracy falls as each document shrinks.
  • MemGPT iteratively pages through the retriever's ranking — performance is unaffected by context length.
  • The full pipeline (embeddings + dataset) was publicly released — embeddings for 20M Wikipedia articles.
What MemGPT Did NOT Solve
  • It stops paging early: the model often quits before exhausting the retriever database, leaving recall on the table.
  • Retriever-bound: embedding-search quality still caps doc-QA accuracy — MemGPT only mitigates it.
  • Needs strong function calling: GPT-3.5-based MemGPT degrades on doc QA; even GPT-4-Turbo-based MemGPT underperformed GPT-4-based MemGPT on nested KV.
  • Self-editing is judgment: whatever the LLM decides not to save is gone — memory quality is only as good as the model's own decisions.
Legacy

Impact — Memory Becomes a Product

MemGPT turned "context management" into "memory management" — and a research idea into an ecosystem.

🏢 Letta
The MemGPT project grew into Letta, founded by the paper's authors — the OS-inspired system became a product line for memory-embedded agents.
🧩 mem0 & memory layers
Memory-for-agents became a product category: mem0 and similar memory layers let any app add persistent, tiered agent memory.
🖥️ The OS paradigm
Paging, interrupts, memory pressure and hierarchical tiers became standard vocabulary for agent design — A-MEM and later memory architectures build on the same lens.
✏️ Self-editing memory
Agents that write and rewrite their own memory via function calls — rather than passively receiving retrieved text — trace their lineage to this paper.
📄 Long-doc analysis
Paging relevant chunks of huge documents in and out of context became the standard pattern for deep Q&A beyond the window.
💡 The lesson
"Context management is memory management": treat the window as RAM, add a disk, and let the model run the swap — don't just make windows bigger.
Deep Dive

Context Is RAM: The Economics of Paging

MemGPT's OS metaphor is doing real engineering work. Main context = RAM (system instructions + working context + a FIFO queue); recall and archival storage = disk. The LLM is the kernel — and every memory operation it triggers has a token price. The paper's results are an accounting argument: self-paging buys capability that no fixed window can afford.

💱
What the Ledger Buys
  • Deep memory retrieval: 92.5% vs 32.1% for fixed-context GPT-4 on the same test
  • Nested KV lookups solved at all 4 depth levels — fixed GPT-4 hits 0% by level 3
  • ~70% of runs hit memory-pressure warnings — the OS noticing, not the user suffering
  • 100% of queue flushes resolved via recursive summarization — the system self-heals
🧾
What the Ledger Costs
  • Every self-edit is an LLM call — memory management spends the same currency as the task
  • Recursive summaries are lossy: paging trades precision for capacity, like every OS
  • Interrupts and heartbeat function-chaining add latency to ordinary turns
  • Working-context size is a tuning knob with no free setting: too small thrashes, too large wastes
Interactive Demo — The Paging Ledger, Live

A MemGPT-style session in miniature. Press send and watch the main-context budget: messages land in the FIFO queue, pressure builds, and the kernel must choose — page out via self-edit, or overflow. The meters on the right track the cost accounting that is the paper's hidden subject.

MAIN CONTEXT · system 400 tk + working 200 tk + queue
0 / 1900 queue tokenssafe
VERDICT
Virtual memory, not bigger memory
The 92.5% vs 32.1% gap is the whole argument: same model, same window — one of them thinks it has a disk. And the lineage is the point of this track: the retrieval-scored stream of Generative Agents became a managed hierarchy here, became evolving notes in A-MEM, and got measured honestly in LongMemEval. The project itself became Letta — memory as a product category.
🧠 The kernel is the model
Memory edits are "entirely self-directed" function calls: the LLM decides what to page, summarize, or archive. No external controller — the intelligence manages its own memory.
📦 FIFO discipline
The queue evicts oldest-first, but the eviction is a recursive summary into recall storage — nothing is dropped, everything is compressed. 100% of flushes resolved this way in testing.
💡 Heartbeats & interrupts
Function chaining via "heartbeat" events and interrupt-driven notifications are lifted from OS design — agent frameworks today (AutoGen lineage, Magentic-One) still ship this vocabulary.
📉 Thrash risk
A too-small working context causes page storms — self-edits triggering self-edits. The paper's ~70% pressure-warning rate is a reminder this is a real operating point, not a free lunch.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the MemGPT paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ The metaphor: context window = RAM, external context = disk, and MemGPT = the OS that pages between them ("virtual context management").
✅ Main context = system instructions + working context + FIFO queue; external context = recall storage + archival storage.
✅ The LLM itself is the memory manager — memory edits and retrieval are "entirely self-directed" function calls.
✅ OS mechanics throughout: ~70% memory-pressure warnings, 100% queue flushes with recursive summaries, interrupts, and function chaining (heartbeats).
✅ Results: deep memory retrieval 92.5% vs 32.1% (GPT-4); nested KV solved at all 4 levels where fixed GPT-4 hits 0% by 3.
✅ Legacy: the project became Letta, seeded the agent-memory ecosystem (mem0, A-MEM, …), and made "memory layer" a product category.