A visual, step-by-step guide to the paper that treated the LLM context window as RAM — and taught the model to manage its own memory like an operating system, paging facts between a small window and unlimited "disk" via self-generated function calls.
LLMs have always had a "RAM problem". MemGPT's answer didn't come from a new architecture — it came from operating-system design.
Operating systems solved this exact problem decades ago: give every program the illusion of unlimited memory by transparently paging data between fast, small RAM and slow, huge disk. MemGPT applies the same trick to the LLM context window — with a twist: the model itself is the memory manager.
The paper opens with a brutal observation: an LLM's "memory" is a small read buffer — while agents need something more like a filing system.
| Model / API | Open? | Context Tokens | ≈ Messages* |
|---|---|---|---|
| Llama (1) | ✓ | 2k | 20 |
| Llama 2 | ✓ | 4k | 60 |
| GPT-3.5 Turbo (release) | ✗ | 4k | 60 |
| Mistral 7B | ✓ | 8k | 140 |
| GPT-4 (release) | ✗ | 8k | 140 |
| GPT-3.5 Turbo | ✗ | 16k | 300 |
| GPT-4 | ✗ | 32k | ~600 |
| Claude 2 | ✗ | 100k | ~2,000 |
| GPT-4 Turbo | ✗ | 128k | ~2,600 |
| Yi-34B-200k | ✓ | 200k | ~4,000 |
*Assuming a ~1k-token preprompt and ~50 tokens per message. Even the big numbers run out: ~600 messages fills a 32k window — and a "personal assistant" is supposed to last for months or years.
Think of the context window as your desk: you can only work with what's on it right now. The rest of your knowledge lives in a filing cabinet (external storage). A person with a small desk and a good filing system beats one with a huge, messy desk — as long as someone keeps filing and retrieving intelligently. MemGPT's bet: that "someone" can be the LLM itself, swapping pages in and out exactly like an OS kernel does between RAM and disk.
MemGPT wraps a fixed-context "LLM processor" in a hierarchical memory system it manages itself. Everything the model can see right now lives in main context; everything it can fetch lives outside.
Everything here counts against the fixed context-window budget.
Indefinitely large — reached only through function calls.
MemGPT's LLM doesn't just chat — it operates on its own memory. The paper calls this "self-directed" editing: every store, search and eviction is a function call the model chose to make.
Stores text indefinitely in archival storage. The model typically fires this right after a memory-pressure warning — deciding, on its own, what's worth keeping.
Vector search over archival storage. Results are paginated, so a query can never overflow the context window — and the model can keep turning pages.
Searches recall storage — the full message database. Matching messages are re-inserted at the back of the FIFO queue: a page-in of your own past.
Main context isn't static: the model updates its working-context notes as facts change — merging, correcting and reorganizing what it "knows" about you.
In a classic RAG pipeline, a separate retriever decides what the model sees. In MemGPT, the LLM processor emits the function calls itself; a function executor validates and runs them, then feeds the results — including runtime errors — back into context. The paper's phrase: memory edits and retrieval are "entirely self-directed". If a call fails (e.g. inserting into an already-full context), the model sees the error and adapts.
The paper evaluates the same OS-inspired system in the two domains where fixed windows hurt most: assistants that must remember you for months, and documents far bigger than any window.
A persistent dialogue assistant with long-term memory across sessions: it remembers facts, preferences and events, reflects on them, and evolves its persona over time. Benchmarked on the Multi-Session Chat (MSC) dataset — five sessions per persona — plus two new tasks the authors introduced: deep memory retrieval and conversation openers.
The chat personalization loop: save facts when the window fills, page them back when the user needs them.
A document-analysis agent that works beyond the context window: the whole corpus (in the experiments, a late-2018 dump of Wikipedia, embedded into a vector database) is loaded into archival storage, and the model pages in relevant chunks itself. Its real preprompt begins: "You are MemGPT DOC-QA bot… remember to keep searching if you can't find the answer."
All numbers below are from the paper's tables. The pattern is consistent: MemGPT turns a weaker fixed-context baseline into a far stronger memory system.
| System | Behavior |
|---|---|
| GPT-3.5 (fixed) | 0% accuracy at just 1 nesting level |
| GPT-4 / GPT-4 Turbo (fixed) | better, but 0% by 3 nesting levels |
| MemGPT (GPT-4 Turbo / 3.5) | beats their baselines, but drops off at 2 levels |
| MemGPT (GPT-4) | unaffected — solves all 4 levels via repeated function queries |
"Nested KV": values may themselves be keys, so the agent must chain lookups — the paper's test of collating information across sources.
| Method | SIM-1 | SIM-3 | SIM-H |
|---|---|---|---|
| Human (gold opener) | 0.800 | 0.800 | 1.000 |
| GPT-3.5 Turbo | 0.830 | 0.812 | 0.817 |
| GPT-4 | 0.868 | 0.843 | 0.773 |
| GPT-4 Turbo | 0.857 | 0.828 | 0.767 |
Baseline rows shown. With MemGPT layered on the same models, the paper reports openers that perform similarly to and occasionally exceeding the hand-written human openers — storing facts in working context was key.
MemGPT turned "context management" into "memory management" — and a research idea into an ecosystem.
MemGPT's OS metaphor is doing real engineering work. Main context = RAM (system instructions + working context + a FIFO queue); recall and archival storage = disk. The LLM is the kernel — and every memory operation it triggers has a token price. The paper's results are an accounting argument: self-paging buys capability that no fixed window can afford.
Check your understanding of the key concepts from the MemGPT paper.
Everything you need to remember about this paper.