A visual, step-by-step guide to LongMemEval — the benchmark that embeds 500 carefully curated questions inside ~115K-token multi-session chat histories to find out whether chat assistants truly remember you. Long-context LLMs drop 30–60%. Commercial memory features fall even harder.
Chat assistants gained "memory" features remarkably fast — but nobody had a rigorous way to grade them. LongMemEval arrived to fill exactly that gap.
Long-term memory is not storage. A real memory must recall scattered facts, merge them across sessions, respect timestamps, overwrite stale facts, and admit when something was never said. Recalling isolated user facts — today's "memory" feature — misses almost all of that.
Everyone sells "memory" — but the tests of that memory were short, synthetic, or single-ability. Here is the gap LongMemEval was built to fill.
Imagine a friend who remembers your birthday — but not that you moved cities, changed jobs, or asked them to stop recommending seafood. That is today's assistant memory: strong on isolated facts, weak on everything that makes memory useful. LongMemEval grades the whole picture: recall, synthesis, time, change, and honest ignorance.
LongMemEval defines long-term memory as five skills, then maps them onto seven concrete question types — from "what did I say in March?" to "I never told you that." Examples below are written in the paper's style.
Recall specific details buried in extensive histories — mentioned by the user or by the assistant. Question types: single-session-user single-session-assistant single-session-preference
Synthesize information across two or more sessions — aggregation and comparison. Most LongMemEval questions need this; evidence can span up to 6 sessions. Question type: multi-session
Reason about when: both explicit time mentions in text ("last weekend") and timestamp metadata on each session. Question type: temporal-reasoning
Recognize that the user's life changed — new city, new job — and keep memory current instead of stale. Question type: knowledge-update
Identify questions seeking information never mentioned in the history — and answer "I don't know." 30 questions are ordinary items rewritten with a false premise. Question type: abstention
Each ability gets concrete, testable forms. IE splits into user-said, assistant-said, and preference (using remembered user info in a reply). The other four map 1-to-1. Every question also has a reference session and labeled evidence turns for scoring recall, not just answers.
Don't generate a chat log and hope for questions. Start from a question that needs memory — then grow months of history around it, backward, so the answer is scattered, timestamped, and hard.
The paper unifies memory systems into three stages — indexing, retrieval, reading — and stress-tests the three most common families. Each one fails differently.
| Approach | How It Works | Where It Wins | Where It Fails |
|---|---|---|---|
| Long-context reading GPT-4o, Llama 3.1, Phi-3… |
Feed the entire chat history into the LLM's context window. No memory system at all. | No architecture change; near-oracle accuracy when only evidence sessions are given (GPT-4o: 0.870). | 30–60% accuracy drop at ~115K tokens; lost-in-the-middle; ~1.5M tokens (setting M) is simply unreachable. |
| RAG memory index → retrieve → read |
Index sessions or rounds, retrieve the top-k relevant items for the question, read them with the LLM. | Works at 1.5M tokens. Optimized (rounds + fact-expanded keys + CoN) GPT-4o reaches 0.720 on M. | Retrieval misses — especially on knowledge updates (fetches the new fact, misses the old one). Weak readers choke past ~3K retrieved tokens. |
| Fact-summary memory MemoryBank-style; ChatGPT, Coze |
Compress sessions into user facts stored in a memory bank; recall facts at query time. | Cheap, plug-and-play; easy to inspect and edit; this is what commercial products ship. | Information loss on details; ChatGPT "tended to overwrite crucial information"; Coze "failed to record indirectly provided user information." |
Quotes from the paper's human study of commercial systems (Section 3.4). Even the RAG winner is only ~0.72 — long-term memory remains unsolved.
Every memory design — including ChatGPT's and Coze's — can be described as choices at four control points. The paper's experiments turn each knob and measure what happens:
Long-context LLMs drop 30–60% as history grows to ~115K tokens. Commercial memory features fall even harder against offline reading. "Long context" and "long memory" are different products.
| System | LLM | IE | MR | KU | TR |
|---|---|---|---|---|---|
| ChatGPT | GPT-4o-mini | 1.000 | 0.647 | 0.667 | 0.652 |
| GPT-4o | 0.688 | 0.441 | 0.833 | 0.435 | |
| Coze | GPT-4o | 0.813 | 0.147 | 0.208 | 0.391 |
| GPT-3.5-turbo | 0.625 | 0.118 | 0.375 | 0.043 |
Setting: 97 questions with 3–6-session histories (~10× shorter than LongMemEval-S), single-session-assistant and abstention questions skipped. IE = information extraction, MR = multi-session reasoning, KU = knowledge updates, TR = temporal reasoning. Multi-session and temporal reasoning are where commercial memory collapses — Coze's temporal score with GPT-3.5 is 0.043.
LongMemEval turned "our assistant remembers you" from a marketing claim into a measurable score. Here is what it changed.
LongMemEval's most-quoted experiment is a subtraction. Give GPT-4o only the evidence it needs — the "oracle" setting — and it answers at 0.870. Give it the same facts buried in a ~115K-token chat history and it falls to 0.606. The difference is pure context length: nothing new to find, everything new to ignore.
Check your understanding of the key concepts from the LongMemEval paper.
Everything you need to remember about this paper (which is, after all, the point).