History Problem Core Idea Breakdown Results Impact Quiz Takeaways
Interactive Paper Explainer

Memory Across Months of Chat
LoCoMo

Prior benchmarks stopped at five sessions. LoCoMo's machine-human pipeline builds conversations of 300 turns across up to 35 sessions — grounded in personas and temporal event graphs — and asks models to actually remember.

Start Learning Read the Paper ↗
300
Turns per conversation
35
Sessions max
3
Task families
2024
Maharjan et al.
History

Conversations That Forget

The five-session ceiling on conversational memory research.

2017-22
Short-dialogue benchmarks
PersonaChat, MultiWOZ-class: a handful of turns, one or two sessions — memory within a snippet.
2022-23
Long-context models arrive
LLMs handle 100k+ tokens — but existing dialogue evals never test memory across dozens of separated sessions.
Feb 2024
🚀 LoCoMo
Maharjan et al.: a machine-human pipeline — LLM agents grounded in personas and temporal event graphs converse over up to 35 sessions; humans verify long-range consistency.
2024+
The memory benchmark wave
LongMemEval (entry #101), LoCoMo-Plus (entry #104), and production memory systems adopt its question-type probe structure.
Grounded Fiction, Verified

The pipeline: two LLM-driven agents converse, each grounded in a persona and a temporal event graph — the graph ensures events, dates, and relationships stay consistent across months of sessions; agents even share and react to images (multimodal memory). Human annotators verify and edit for long-range consistency and graph grounding. The result: conversations averaging 300 turns over up to 35 sessions (~9K tokens) — with structured probes over them: question answering (five probe types), event summarization, and multimodal dialogue generation.

Chapter 01

Chatbots with Amnesia

What five-session benchmarks couldn't see.

🧠
The Short-Horizon Habit
  • Dialogue benchmarks cap at ~5 sessions — memory beyond a snippet is untested
  • Long-context models: can they USE context that is months of scattered chat sessions?
  • RAG systems: is retrieval over conversation history actually effective at this span?
  • No controlled ground truth existed for 'what happened when' across a long relationship
🕰
The LoCoMo Answer
  • Persona + temporal event graph grounding: consistent, structured long-range histories
  • Conversations of 300 turns across up to 35 sessions (~9K tokens), human-verified
  • Probes: QA (five question types incl. adversarial and multimodal), event summarization, dialogue generation
  • Findings: long-context LLMs and RAG both struggle in very long-term settings — the gap is measurable now
Analogy — The Yearbook Test

A five-session benchmark is quizzing someone on this week's group chat. LoCoMo is handing them a yearbook of a friendship — 35 months of moments, people who came and went, promises made in March and broken in October — and asking pointed questions about all of it. Humans with the transcript do fine; models discover how much of a relationship they can actually hold.

Chapter 02

The Pipeline

From event graph to verified conversation — the construction machinery.

Construction
  • Define personas + temporal event graphs — the structured skeleton of each relationship
  • LLM agents converse grounded in the graph — events, dates, and facts stay consistent
  • Agents share and react to images — multimodal long-term memory included
  • Human verification/editing for long-range consistency and grounding
The probes
  • Question answering — five types: single-hop, multi-hop, temporal, open-domain, and adversarial (unanswerable)
  • Event summarization — condense long histories faithfully
  • Multimodal dialogue generation — continue with image context intact
The Findings

Across very long-term dialogues, long-context LLMs and RAG-based memory both struggle: single-hop recall is decent; multi-hop and temporal reasoning degrade badly; adversarial questions (no answer exists) tempt confabulation. The paper's position: capability gaps at conversation scale are structured and measurable — the field had been optimizing memory for the wrong horizon. Its question-type design (especially multi-hop + adversarial) became the template for LongMemEval (entry #101).

Interactive Demo — Five Kinds of Remembering

Tab through the question types over one 35-session friendship — each probes a different memory faculty.

Chapter 03

Where Memory Breaks

The failure decomposition — which probe types hurt most.

Difficulty Ladder
Interactive Demo — Building One Conversation

Follow the machine-human pipeline — from event graph to verified 35-session history.

Chapter 05

The Long Horizon, Measured

What 35 sessions of ground-truthed conversation revealed.

CONVERSATIONS
300 turns
up to 35 sessions, ~9K tokens, human-verified
PROBES
5 QA types
plus summarization + multimodal generation
LONG-CONTEXT LLMS
struggle
multi-hop and temporal degrade
RAG MEMORY
struggles
retrieval alone insufficient at conversation scale
Interactive Demo — Why Not Just Longer Contexts?

128k tokens could hold the whole conversation. Press reveal for why that isn't the answer.

Probe typeWhat it testsTypical failure
Single-hop QAdirect recallsearch misses the session
Multi-hop QAchaining scattered factsretrieval finds parts, reasoning fails to compose
Temporal QAevent ordering across monthsconfuses the session timeline
Adversarial QAabstention when no answer existsconfabulates a plausible one
Event summarizationfaithful condensationdrops or merges distinct events

The probe taxonomy LoCoMo introduced — adopted nearly wholesale by the memory-benchmark generation that followed.

Legacy

Legacy — The Exam for Memory Systems

LoCoMo set the probe structure the memory category builds against.

📝 The probe canon
Five QA types + summarization became the standard memory-eval battery — LongMemEval (entry #101) and LoCoMo-Plus (entry #104) extend it.
🧵 The pipeline pattern
Graph-grounded conversation generation with human verification became the accepted way to build long-horizon ground truth — synthetic but checkable.
📉 The honest gap
Showing long-context AND RAG both struggling moved memory research from retrieval tricks to system design (Mem0, A-MEM — entries #102-103).
⚠️ What it did NOT solve
Conversations are synthetic personas — real-user messiness differs; ~9K-token histories are modest by production standards; and English-only coverage limits cross-linguistic claims.
🛤 Read next
The memory category: LongMemEval · Mem0 · LoCoMo-Plus
Test Yourself

Quick Quiz

Check your understanding of the key concepts from LoCoMo.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ LoCoMo: 300-turn conversations across up to 35 sessions — human-verified, graph-grounded.
✅ Pipeline: persona + temporal event graph agents converse; humans audit long-range consistency.
✅ Probes: 5 QA types (incl. adversarial), event summarization, multimodal dialogue generation.
✅ Both long-context LLMs and RAG struggle — multi-hop, temporal, and abstention degrade.
✅ The probe taxonomy became the memory-category standard (LongMemEval, LoCoMo-Plus).
✅ Read it as the exam that turned agent memory into a systems discipline.