Prior benchmarks stopped at five sessions. LoCoMo's machine-human pipeline builds conversations of 300 turns across up to 35 sessions — grounded in personas and temporal event graphs — and asks models to actually remember.
The five-session ceiling on conversational memory research.
The pipeline: two LLM-driven agents converse, each grounded in a persona and a temporal event graph — the graph ensures events, dates, and relationships stay consistent across months of sessions; agents even share and react to images (multimodal memory). Human annotators verify and edit for long-range consistency and graph grounding. The result: conversations averaging 300 turns over up to 35 sessions (~9K tokens) — with structured probes over them: question answering (five probe types), event summarization, and multimodal dialogue generation.
What five-session benchmarks couldn't see.
A five-session benchmark is quizzing someone on this week's group chat. LoCoMo is handing them a yearbook of a friendship — 35 months of moments, people who came and went, promises made in March and broken in October — and asking pointed questions about all of it. Humans with the transcript do fine; models discover how much of a relationship they can actually hold.
From event graph to verified conversation — the construction machinery.
Across very long-term dialogues, long-context LLMs and RAG-based memory both struggle: single-hop recall is decent; multi-hop and temporal reasoning degrade badly; adversarial questions (no answer exists) tempt confabulation. The paper's position: capability gaps at conversation scale are structured and measurable — the field had been optimizing memory for the wrong horizon. Its question-type design (especially multi-hop + adversarial) became the template for LongMemEval (entry #101).
The failure decomposition — which probe types hurt most.
What 35 sessions of ground-truthed conversation revealed.
| Probe type | What it tests | Typical failure |
|---|---|---|
| Single-hop QA | direct recall | search misses the session |
| Multi-hop QA | chaining scattered facts | retrieval finds parts, reasoning fails to compose |
| Temporal QA | event ordering across months | confuses the session timeline |
| Adversarial QA | abstention when no answer exists | confabulates a plausible one |
| Event summarization | faithful condensation | drops or merges distinct events |
The probe taxonomy LoCoMo introduced — adopted nearly wholesale by the memory-benchmark generation that followed.
LoCoMo set the probe structure the memory category builds against.
Check your understanding of the key concepts from LoCoMo.
Everything you need to remember about this paper.