History Problem Abilities Design Solutions Results Impact Deep Dive Quiz
Interactive Paper Explainer

The Long-Term
Memory Test

A visual, step-by-step guide to LongMemEval — the benchmark that embeds 500 carefully curated questions inside ~115K-token multi-session chat histories to find out whether chat assistants truly remember you. Long-context LLMs drop 30–60%. Commercial memory features fall even harder.

Start Learning Read the Paper ↗
500
Questions
5
Memory Abilities
~115K
Tokens of History (S)
2024
Year Released
History

From Session Amnesia to Memory Wars

Chat assistants gained "memory" features remarkably fast — but nobody had a rigorous way to grade them. LongMemEval arrived to fill exactly that gap.

2022
Session amnesia
Chat assistants start fresh every conversation. Early long-term dialogue work (MSC, DuLeMon — Xu et al., 2022) grades generated replies, not memory accuracy.
2023 · Oct
MemGPT (Packer et al.)
Treats the LLM like an operating system with paged conversational memory. Memory becomes an architecture problem — not just a longer prompt.
2024 · Early
Memory goes commercial
ChatGPT and Coze ship memory features that store user facts. Academic benchmarks appear too — but each covers a slice: MemoryBank (194 questions, ~5K tokens), LoCoMo, PerLTQA, DialSim.
2024
The design space gets mapped
Surveys and systems explore compression, indexing, and retrieval for LLM memory. Still no benchmark scores all five memory abilities at realistic length.
2024 · Oct
🚀 LongMemEval (Wu et al.)
500 questions, 5 abilities, histories of ~115K tokens (scalable to ~1.5M). Long-context LLMs drop 30–60%; commercial systems score 30–70% in a simpler setting.
2025 →
Memory as a battleground
Accepted at ICLR 2025. Dataset and code released on GitHub and Hugging Face; memory becomes a headline feature — and LongMemEval the yardstick.
Key Insight

Long-term memory is not storage. A real memory must recall scattered facts, merge them across sessions, respect timestamps, overwrite stale facts, and admit when something was never said. Recalling isolated user facts — today's "memory" feature — misses almost all of that.

A FACT THAT CHANGES (THE PAPER'S HARDEST TRAP)
S3: "I live in Boston"
S17: "I just moved to Austin"
S31: "moved back to Boston!"
Q: where do I live now?
The correct answer exists only at the moment of asking — stale copies are landmines.
Chapter 01

Chatbots with Amnesia

Everyone sells "memory" — but the tests of that memory were short, synthetic, or single-ability. Here is the gap LongMemEval was built to fill.

🕳️
How Memory Was Tested Before
  • Standard QA benchmarks are single-session — no memory involved at all
  • MemoryBank: 194 questions over ~5K-token histories
  • LoCoMo and DialSim: ~10K-token dialogues or synthetic TV-world agents
  • PerLTQA scales to ~1M tokens but skips knowledge updates
  • Commercial "memory" means recalling isolated user facts
  • None of them scores all five memory abilities
🧠
LongMemEval's Answer
  • 500 questions covering all 5 abilities — including updates and abstention
  • Engineered histories averaging ~115K tokens, scalable to ~1.5M
  • Evidence scattered across up to 6 sessions, at any position
  • Timestamped sessions enable real temporal reasoning
  • Every question ships with labeled evidence sessions (oracle labels)
  • Two fixed settings (S and M) for apples-to-apples comparison
The Analogy — The Birthday Friend

Imagine a friend who remembers your birthday — but not that you moved cities, changed jobs, or asked them to stop recommending seafood. That is today's assistant memory: strong on isolated facts, weak on everything that makes memory useful. LongMemEval grades the whole picture: recall, synthesis, time, change, and honest ignorance.

Interactive Demo — The Abstention Detector

Some questions below have answers in the memory snippets; some don't — the paper calls these "false premise" questions (30 of the 500). For each card, decide: should the assistant answer, or abstain ("I don't know")? Correct abstentions score points too.

Chapter 02

Five Abilities, Seven Question Types

LongMemEval defines long-term memory as five skills, then maps them onto seven concrete question types — from "what did I say in March?" to "I never told you that." Examples below are written in the paper's style.

Ability 1 · Information Extraction (IE)

Recall specific details buried in extensive histories — mentioned by the user or by the assistant. Question types: single-session-user single-session-assistant single-session-preference

Q: "Which gym did the user say they switched to?"
Ability 2 · Multi-Session Reasoning (MR)

Synthesize information across two or more sessions — aggregation and comparison. Most LongMemEval questions need this; evidence can span up to 6 sessions. Question type: multi-session

Q: "How many books did the user finish this year, across all their reading updates?"
Ability 3 · Temporal Reasoning (TR)

Reason about when: both explicit time mentions in text ("last weekend") and timestamp metadata on each session. Question type: temporal-reasoning

Q: "Which restaurant did the assistant recommend most recently?"
Ability 4 · Knowledge Updates (KU)

Recognize that the user's life changed — new city, new job — and keep memory current instead of stale. Question type: knowledge-update

Q: "Where does the user live now, after all their moves?"
Ability 5 · Abstention (ABS)

Identify questions seeking information never mentioned in the history — and answer "I don't know." 30 questions are ordinary items rewritten with a false premise. Question type: abstention

Q: "What's the user's favorite podcast?" — never discussed. ✓ "I don't know."
Why Seven Types, Not Five?

Each ability gets concrete, testable forms. IE splits into user-said, assistant-said, and preference (using remembered user info in a reply). The other four map 1-to-1. Every question also has a reference session and labeled evidence turns for scoring recall, not just answers.

Interactive Demo — The Knowledge Update Timeline (Centerpiece)

A user fact changes across three sessions. Read the timeline, then click the session whose fact is current right now. This is the paper's hardest trap: systems love stale facts.

Chapter 03

Question-First Construction

Don't generate a chat log and hope for questions. Start from a question that needs memory — then grow months of history around it, backward, so the answer is scattered, timestamped, and hard.

LongMemEval instance = ( S , q , t_q , a )
S
Chat History
A timestamped sequence of user–assistant sessions: ~115K tokens (setting S) or 500 sessions ≈ 1.5M tokens (setting M).
q
The Question
Asked in a fresh session, long after the evidence was buried in S.
t_q
Question Date
The "now" the question is anchored to — the reference point for temporal reasoning.
a
Answer + Evidence
The gold answer, plus labeled evidence sessions and turns for oracle retrieval scoring.
Step 1 · Design
Questions start from an ontology of 164 user attributes (lifestyle, belongings, life events, situations, demographics). An LLM proposes ~1,000 candidates per type; human experts keep and rewrite the best ~5%.
Step 2 · Decompose
Each answer is manually split into evidence statements — with timestamps assigned if the fact involves time.
Step 3 · Simulate
Llama 3 70B self-chats a task-oriented session per statement. The user LLM reveals evidence indirectly: "I bought a new car" becomes a question about car insurance.
Step 4 · Scatter
Evidence sessions are shuffled among filler sessions — 25% ShareGPT, 25% UltraChat, 50% simulated — then timestamps resolved around evidence anchors.
Step 5 · Verify
Humans screen and edit every session: check evidence inclusion, spread evidence positions, and rephrase LLM-sounding time mentions into natural speech. ~400 human hours in total.
ATTRIBUTES
164
user attributes in the question ontology
CANDIDATES / TYPE
~1,000
LLM-proposed questions per type
FINAL YIELD
~5%
survive human filtering and rewriting
EVIDENCE SPAN
≤ 6
sessions a single answer can be scattered across
Chapter 04

Three Ways to Give an Assistant Memory

The paper unifies memory systems into three stages — indexing, retrieval, reading — and stress-tests the three most common families. Each one fails differently.

The Three Memory Solution Families on LongMemEval
ApproachHow It WorksWhere It WinsWhere It Fails
Long-context reading
GPT-4o, Llama 3.1, Phi-3…
Feed the entire chat history into the LLM's context window. No memory system at all. No architecture change; near-oracle accuracy when only evidence sessions are given (GPT-4o: 0.870). 30–60% accuracy drop at ~115K tokens; lost-in-the-middle; ~1.5M tokens (setting M) is simply unreachable.
RAG memory
index → retrieve → read
Index sessions or rounds, retrieve the top-k relevant items for the question, read them with the LLM. Works at 1.5M tokens. Optimized (rounds + fact-expanded keys + CoN) GPT-4o reaches 0.720 on M. Retrieval misses — especially on knowledge updates (fetches the new fact, misses the old one). Weak readers choke past ~3K retrieved tokens.
Fact-summary memory
MemoryBank-style; ChatGPT, Coze
Compress sessions into user facts stored in a memory bank; recall facts at query time. Cheap, plug-and-play; easy to inspect and edit; this is what commercial products ship. Information loss on details; ChatGPT "tended to overwrite crucial information"; Coze "failed to record indirectly provided user information."

Quotes from the paper's human study of commercial systems (Section 3.4). Even the RAG winner is only ~0.72 — long-term memory remains unsolved.

The Unified Framework — 3 Stages × 4 Control Points

Every memory design — including ChatGPT's and Coze's — can be described as choices at four control points. The paper's experiments turn each knob and measure what happens:

CP1 · Value (indexing)
Decompose sessions into individual rounds — better granularity for retrieval and reading. Compressing to summaries or facts loses information (except on multi-session questions, where facts help).
CP2 · Key (indexing)
Expand each item's key with extracted user facts: fact-augmented key expansion adds +9.4% average recall@k and +5.4% end-to-end accuracy across models.
CP3 · Query (retrieval)
Naive similarity search fails on "last weekend"-style questions. Time-aware query expansion narrows the search range: +6.8–11.3% recall on temporal-reasoning questions.
CP4 · Reading strategy
Present retrieved items as structured JSON and read with Chain-of-Note (extract notes per item, then reason). Up to 10 absolute points — even with perfect retrieval.
Chapter 05

Results: Everyone Forgets

Long-context LLMs drop 30–60% as history grows to ~115K tokens. Commercial memory features fall even harder against offline reading. "Long context" and "long memory" are different products.

GPT-4o · FULL CONTEXT
0.606
on ~115K-token histories — vs 0.870 evidence-only (−30.3%)
LLAMA 3.1 70B · FULL CONTEXT
0.334
vs 0.744 evidence-only (−55.1%)
CHATGPT · GPT-4o MEMORY
0.577
online memory vs 0.918 offline reading (−37%)
COZE · GPT-4o MEMORY
0.330
online memory vs 0.918 offline reading (−64%)
Interactive Demo — History Growth Stress Test

Pick a model, then toggle the memory solution. Watch what happens as the history grows from evidence-only, to ~115K tokens (S), to 500 sessions ≈ 1.5M tokens (M).

Real numbers from the paper (Table 8, five additional LLMs): "read everything" = direct long-context reading of the full history; "RAG memory" = retrieval over round-granularity values with fact-expanded keys. QA accuracy, LLM-judge evaluated.

Commercial Assistants, by Ability (human study, simplified setting)
SystemLLMIEMRKUTR
ChatGPTGPT-4o-mini1.0000.6470.6670.652
GPT-4o0.6880.4410.8330.435
CozeGPT-4o0.8130.1470.2080.391
GPT-3.5-turbo0.6250.1180.3750.043

Setting: 97 questions with 3–6-session histories (~10× shorter than LongMemEval-S), single-session-assistant and abstention questions skipped. IE = information extraction, MR = multi-session reasoning, KU = knowledge updates, TR = temporal reasoning. Multi-session and temporal reasoning are where commercial memory collapses — Coze's temporal score with GPT-3.5 is 0.043.

Which Ability Is Hardest?
  • For commercial systems: multi-session reasoning and temporal reasoning — aggregation across sessions and time is where fact-list memories fall apart.
  • For RAG systems: knowledge updates. In the error analysis, ~10% of correctly answered questions were answered despite retrieval failure — and they were mostly knowledge-update items where the retriever found the updated fact but missed the information before the update.
  • For everyone: 15–19% of all instances had correct retrieval but wrong generation — reading is its own failure mode.
What LongMemEval Did NOT Solve
  • Histories are engineered (self-chat + public corpora), not organic human logs — realism is approximated, not guaranteed.
  • The best memory design still only reaches ~0.72 accuracy on M — a ceiling, not a finish line.
  • Scoring relies on a GPT-4o judge (≥90% agreement with experts, but open-ended preference and abstention answers drift).
  • Online cost and latency of memory operations — the price a product pays — are not measured.
Legacy

Impact — Memory Gets a Yardstick

LongMemEval turned "our assistant remembers you" from a marketing claim into a measurable score. Here is what it changed.

📏 A yardstick for memory
500 questions, two fixed settings (S: ~115K tokens, M: 500 sessions ≈ 1.5M), oracle labels for recall — released on GitHub and Hugging Face for anyone to test against.
🔄 The knowledge-update trap
Commercial banks overwrite crucial information; RAG retrievers fetch the updated fact and miss the old one. Updates are now a first-class evaluation target.
🙅 Abstention became an ability
30 false-premise questions grade "I don't know" as a correct answer — memory includes knowing what was never said.
🎓 A training signal
Labeled evidence sessions and turns turn memory from a prompt trick into a trainable behavior — supervised data for fine-tuning and aligning future memory-aware assistants.
🏭 Products, measured
ChatGPT: 0.577. Coze: 0.330. Offline reading with the same LLM: 0.918. A memory feature is not the same thing as memory ability.
🧠 Long context ≠ long memory
A giant window is not a memory: 30–60% drops at ~115K tokens, and 1.5M tokens is unreachable. Retrieval-based memory keeps small models alive at scale.
Deep Dive

The Oracle Cliff: What Long Contexts Actually Cost

LongMemEval's most-quoted experiment is a subtraction. Give GPT-4o only the evidence it needs — the "oracle" setting — and it answers at 0.870. Give it the same facts buried in a ~115K-token chat history and it falls to 0.606. The difference is pure context length: nothing new to find, everything new to ignore.

🏗️
A Benchmark Built Backward
  • Question-first construction: design the question, then generate months of history backward so the answer is scattered and timestamped
  • 500 questions × 5 abilities: extraction, multi-session reasoning, knowledge updates, temporal reasoning, abstention
  • Average histories ~115K tokens (~40 sessions), scalable to 1.5M — far beyond any earlier memory benchmark
  • Abstention is graded: knowing what you no longer know is an ability, not a refusal
📉
The Cliff, Measured Three Ways
  • Long-context LLMs drop 30–60% from oracle to full history: GPT-4o 0.870 → 0.606
  • Commercial memory underdelivers: ChatGPT 0.577, Coze 0.330 — vs 0.918 for offline reading with the same LLM
  • Knowledge updates are the trap: models prefer early facts over current ones — primacy over recency
  • Every long-context fix (needle tests) had trained on retrieval, not update-aware reasoning over months
Interactive Demo — The Update Trap and the Cliff

First: a knowledge-update question the way a model sees it — two facts, months apart, one current. Pick the answer a well-behaved memory system gives. Then: the cliff chart, oracle vs full history for three systems — the exact comparison that reframed long-context marketing.

THE UPDATE TRAP · SESSION 3 vs SESSION 38
Session 3: "I moved apartments — I am in Oakland now, near Lake Merritt." · Session 38: "Finally left the Bay — I am in Austin these days."
Q: Which neighborhood is the user's doctor most convenient to today?
THE ORACLE CLIFF · SAME FACTS, DIFFERENT CONTEXT
VERDICT
The winning recipe was engineering, not heroics
The paper's best configuration is a checklist: round decomposition (answer per round, then aggregate), fact-expanded keys (+9.4% recall), time-aware queries, and Chain-of-Note JSON reading. None of it is a new model — all of it is memory management, which is the survey's vocabulary vindicated. Where this goes next: LoCoMo-Plus pushes beyond factual recall to behavioral consistency, and the systems it grades include A-MEM and MemGPT.
🎯 Primacy over recency
Models prefer early-learned facts over updated ones — the reverse of what a memory system needs. Long contexts amplify the bias: more history, more stale anchors.
🤐 Abstention as ability 5
When the history is silent, the correct answer is "I do not know." Grading abstention stops benchmarks from rewarding confident fabrication.
🏬 Commercial memory ≠ memory research
ChatGPT 0.577 and Coze 0.330 against 0.918 offline: shipped memory systems are convenience features, not yet knowledge custodians.
🧱 500K-token stress mode
The scalable protocol (~500 sessions) exists to keep the benchmark ahead of context windows — the test hardens as the models grow.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the LongMemEval paper.

Reference

Key Takeaways

Everything you need to remember about this paper (which is, after all, the point).

✅ 500 questions × 5 abilities: information extraction, multi-session reasoning, knowledge updates, temporal reasoning, abstention.
✅ Histories average ~115K tokens (≈40 sessions), scalable to 500 sessions ≈ 1.5M tokens — far beyond any prior memory benchmark.
✅ Question-first construction: design the question, then generate months of history backward so the answer is scattered and timestamped.
✅ Long-context LLMs drop 30–60% at ~115K tokens — GPT-4o falls from 0.870 (evidence-only) to 0.606.
✅ Commercial memory underperforms: ChatGPT 0.577 and Coze 0.330 vs 0.918 for offline reading with the same LLM.
✅ Best recipe found: round decomposition + fact-expanded keys (+9.4% recall) + time-aware queries + Chain-of-Note JSON reading.