History Problem Framework Abilities Construction Results Impact Deep Dive Quiz
Interactive Paper Explainer

Memory Beyond Facts
LoCoMo-Plus

A visual, step-by-step guide to the 2026 benchmark that asks a harder question about LLM agents: not "can you recall what I said?" but "can you remember my goals, states, and values — and let those latent constraints steer your advice hundreds of turns later?"

Start Learning Read the Paper ↗
4
Latent Constraints (causal · state · goal · value)
17
Methods Benchmarked
26.06
Best Cognitive Score / 100 (gemini-2.5-pro)
2026
Year Published
History

From Recall to Remembering You

Long-term conversational memory went from a developer hack to a measured capability in about three years — but the yardstick kept measuring the same thing: facts. LoCoMo-Plus changes the yardstick.

2023
MemGPT & toolkit memory
Packer et al. treat the LLM like an operating system — a small working context plus paged long-term storage. LangChain and LlamaIndex ship summary & buffer chat memories.
2024 · Feb
LoCoMo (Maharana et al.)
Humans and bots chat for ~300 turns, then models face single-hop, multi-hop, temporal, commonsense, and adversarial questions. Very long-term memory gets its benchmark (ACL 2024).
2024
Memory ships — and gets measured
ChatGPT rolls out persistent memory with user controls (OpenAI). LongMemEval (Wu et al.) probes 500 questions across 5 memory abilities — still mostly explicit facts.
2025
Memory systems bloom
A-MEM (Xu et al.) gives agents self-organizing memory; Mem0 (Chhikara et al.) ships a production memory layer; Claude gains memory for teams (Anthropic, 2025).
2026 · Feb
🚀 LoCoMo-Plus (Li et al.)
Memory is not a retrieval problem. LoCoMo-Plus tests cognitive memory under cue–trigger semantic disconnect, graded by constraint consistency — not string overlap.
Next →
The open frontier
Belief revision, emotional dynamics, multi-agent memory — named as out-of-scope in the paper's limitations — still await benchmarks of their own.
Key Insight

Real memory constrains behavior instead of answering quizzes. In the paper's own motivating example, a user says they're preparing for an important exam and want to minimize distractions. Much later, they ask: "Should I start watching that new TV series everyone is talking about?"

No fact is being recalled — several replies look perfectly fine in isolation. The right one depends on remembering the earlier goal and noticing the conflict. Benchmarks built on factual recall can't see this skill at all.

THE TWO MOMENTS, SIDE BY SIDE
🎙️ Turn 12: "Prepping for a big exam — want zero distractions."
… ~40 turns of work, family, and travel talk …
🎬 Turn 260: "Should I start that new series everyone's talking about?"
The link is cognitive, not lexical — cue and trigger share almost no keywords.
Chapter 01

The Problem with Pop-Quiz Memory

Existing benchmarks mostly check whether a model can find a fact that was explicitly stated earlier. But in real conversations, memory's job is subtler: it quietly constrains every later response.

🗂️
Recall-Only Evaluation
  • Facts are stated explicitly, then queried with strong semantic alignment
  • Task disclosure — "this question tests memory" — leaks the intent to the model
  • EM, token-F1, BLEU, ROUGE grade surface overlap, not behavioral sense
  • Many valid answers receive different scores; length and verbosity skew results
  • The result: memory becomes a retrieval problem it never really was
🧠
LoCoMo-Plus's Constraint View
  • Implicit constraints — user state, goals, values, causal context — drive behavior
  • Queries arrive as natural continuations, with no task announcement
  • Correct = consistent with the latent constraint, in any surface form
  • Cues are semantically disconnected from their triggers — retrieval can't shortcut
  • Exposes failures that existing benchmarks never capture
Interactive Demo — Recall vs Reason: Same Memory, Two Questions
An Analogy

A friend who aces trivia night about your life — your job, your dog's name, your birthday — but offers you an espresso at 11pm before your 6am flight has facts, not memory. Facts answer queries; memory shapes behavior when nobody is querying anything. LoCoMo-Plus is the first benchmark in this lineage built to grade the second kind.

Chapter 02

Two Levels of Memory

The paper's formal framing splits conversational memory in two. Level-1 is what every prior benchmark measures. Level-2 is what actually makes an assistant feel like it knows you.

Level-1 · Factual Memory

Relevant information is explicitly stated in the history and can be directly recalled or reasoned over — localized object-centric facts ("my dog is named Pepper") and event-oriented episodic details ("we hiked on May 3rd"). Each query admits a single well-defined ground-truth answer, and correctness is string or semantic similarity against it. This is the regime of LoCoMo, LongMemEval, and most QA-style memory tests.

Level-2 · Cognitive Memory

Realistic conversations depend on implicit constraints inferred from prior turns — user state, goals, preferences, values, causal context. There is no single correct response: the history induces a latent constraint c that restricts the space of behaviorally valid outputs. Any response inside that space counts, regardless of wording. This is the regime LoCoMo-Plus opens for measurement.

Given history ℋ = {u₁,a₁, …, uₜ,aₜ} and query qₜ₊₁:
aₜ₊₁ is correct ⟺ aₜ₊₁ ∈ 𝒜𝒸 = { a | a is consistent with c }
ℋ
Interaction history
The full multi-turn conversation — user turns u and agent turns a, possibly hundreds of turns long.
c
Latent constraint
What the history implicitly induces — a goal, state, causal fact, or value the user holds. Never restated at query time.
𝒜𝒸
Valid response space
The set of all responses consistent with c. Many surface forms; one behavioral requirement.
∈
Membership, not matching
Correctness is membership in 𝒜𝒸 — not string overlap with one reference answer. This is "constraint consistency."
🎯 The target skill
Retain and apply latent constraints across long contexts, even when the downstream query looks unrelated to the original cue.
🔌 Cue–trigger disconnect
The trigger query is deliberately low in semantic similarity to its cue — similarity-based retrieval cannot bridge the gap.
🧩 Four constraint types
Cognitive memory is decomposed into causal, state, goal, and value — four interacting latent signals that shape behavior.
📊 New yardstick
A unified constraint-consistency framework replaces task-disclosed prompts and string-matching metrics for both memory levels.
Chapter 03

The Four Latent Constraints

LoCoMo-Plus decomposes cognitive memory into four constraint types. Every instance in the benchmark is a cue that plants one of them — and a trigger, far away in the conversation, that only makes sense if the model kept it. The examples below are drawn directly from the paper's Appendix C.

Interactive Demo — Cognitive Constraint Explorer

Pick a constraint type. Read the cue (buried early in a long chat), then the trigger (months later). Ask yourself: could any retrieval system connect these two? Almost no words overlap.

Why retrieval can't shortcut this
  • Semantic filtering at build time: BM25 and MPNet similarity scoring removes any cue–trigger pair where the cue is repeated, paraphrased, or recoverable from the query alone.
  • Triggers are underspecified: in isolation, several responses look reasonable — only cue-consistent ones are valid.
  • Temporal gap by design: each pair carries a gap indicator t (a week to several months), so the cue is buried under turns of interference.
  • First-person voice: triggers sound like natural self-reflection, not quiz questions.
How responses get graded
  • No task disclosure: the query arrives as a plain dialogue continuation — the model is never told "this is a memory test."
  • Evidence-grounded LLM judge: an LLM checks whether the response acknowledges or adapts to the cue (gemini-2.5-flash as the main judge).
  • Binary labels for cognitive questions: the criterion is presence or absence of memory-aware behavior — not partial string credit.
  • Judge audited: human–human agreement 0.903; human–judge agreement 0.801–0.820; scores stable across judge backbones (|Δ| ≤ 3.33 points).
Chapter 04

Building Cue–Trigger Pairs

LoCoMo-Plus instances are constructed from scratch and embedded into LoCoMo's long dialogues. The pipeline is deliberately heavy on human validation and filtering — diagnostic coverage is prioritized over scale.

The Six-Stage Construction Pipeline
Step 1 · Generate
An LLM writes short two-turn cue dialogues that implicitly convey state, goals, preferences, or values — naturalistic, never explicit facts. Temperature 0.7, 256-token cap, 50 samples per relation type across the 4 types.
Step 2 · Verify memory-worthiness
Humans keep only cues that are persistent or behaviorally constraining, not trivially inferable from local context, and plausibly useful to remember. Everything else is discarded.
Step 3 · Build triggers
Per cue, five candidate trigger queries — first-person, semantically distant, one week to several months later, each a distinct cognitive angle. Every pair gets a temporal gap indicator t.
Step 4 · Filter semantically
BM25 and MPNet similarity scoring strips out any pair where the cue is repeated, paraphrased, or recoverable from the trigger alone — no lexical shortcuts allowed.
Step 5 · Validate elicitation
A final human check: does a helpful response genuinely require recalling and applying the cue? Only validated cue–trigger pairs survive.
Step 6 · Embed in LoCoMo
Each pair is inserted into a ~300-turn LoCoMo dialogue: the cue early, the trigger after the specified gap, surrounded by realistic distractor turns.
Old Evaluation vs LoCoMo-Plus Evaluation
AspectLoCoMo-style protocolLoCoMo-Plus framework
Input sideTask-disclosed prompts — the model is told which memory skill is testedNatural dialogue continuations, no task disclosure
Output sideEM, token-F1, BLEU, ROUGE vs a reference stringLLM judge scores constraint consistency, grounded in dialogue evidence
Ground truthOne reference answer per questionA valid response space 𝒜𝒸 — many correct surface forms
LabelsBinary correctness (typical QA)3-level (correct / partial / wrong) for factual & commonsense; binary for temporal, adversarial, cognitive
Known biasPrompt adaptation + length bias (scores peak near the 5.18-token average reference length)Judge robustness measured against humans (0.801–0.820) and across backbones
GENERATION MODELS
5
gpt-5-nano, gpt-4o, gpt-4.1, gemini-2.5-flash, gemini-2.5-pro — used only to build data
TRIGGERS PER CUE
5
distinct cognitive angles, then filtered down
TEMPERATURE
0.7
stochastic decoding for diversity, 256-token cap
HUMAN PASSES
2
memory-worthiness check + elicitation validation
Chapter 05

Everyone Falls — Even the Best

The paper evaluates 17 methods across four categories: open-source LLMs, closed-source LLMs, RAG pipelines, and dedicated memory systems. One pattern holds everywhere — a large, persistent gap between factual and cognitive memory.

Table 1 Highlights — LoCoMo (factual average) vs LoCoMo-Plus (cognitive)
MethodLoCoMo AvgLoCoMo-PlusGap
gemini-2.5-pro71.7826.0645.72
gemini-2.5-flash69.2524.6744.58
gpt-4o (full context)62.9921.0541.94
gpt-4.162.2118.6343.58
Qwen3-14B59.6519.0940.56
A-Mem (GPT-4o)59.6417.2042.44
SeCom (GPT-4o)57.5314.9042.63
Mem0 (GPT-4o)57.2415.8041.44
gpt-5-nano54.9614.8440.12
RAG top-5, emb-large (GPT-4o)45.3215.5529.77
Qwen2.5-7B-Instruct45.319.5735.74

Scores / 100. Gaps run ~12–46 points across all 17 methods. Memory systems and retrieval pipelines do not close the cognitive gap — and differences between methods compress toward uniformly low performance.

GEMINI-2.5-PRO
26.06
best cognitive score — vs 71.78 factual average
−45.72 gap from its own Level-1 score
GPT-4O (FULL CONTEXT)
21.05
cognitive score — vs 62.99 factual average
context length alone doesn't save cognitive memory
A-MEM (GPT-4O)
17.20
best memory system on LoCoMo-Plus
structured agentic memory still −42.44 from its factual avg
RAG TOP-5 (EMB-LARGE)
15.55
cognitive score — retrieval-only baseline
similarity search can't find semantically distant cues
Bias Finding 1 · Prompt Disclosure

When benchmarks announce the task ("this is a temporal-reasoning question"), measured ability profiles shift — temporal and adversarial categories get disproportionately inflated scores under task-disclosed evaluation. Reported gains partly reflect sensitivity to prompts, not stable memory behavior. Real users never announce "this query requires memory."

Bias Finding 2 · Length Sensitivity

EM, F1, BLEU, and ROUGE all vary systematically with output length — scores peak near the average reference length (5.18 tokens) and degrade as generations get shorter or longer. Models are penalized or favored by verbosity alone, regardless of semantic correctness — a systematic bias in cross-model comparison.

Interactive Demo — Memory Stress Timeline

Grow the conversation one session at a time and watch what happens to the three memory types — the shape of the collapse follows the paper's Figure 7 finding.

Illustrative rendering of the paper's Figure 7 (100 representative cases per memory type, varying dialogue turns): object recall stays robust, episodic recall degrades steadily, cognitive memory collapses rapidly as context grows. Bar heights are qualitative, not measured values.

REFERENCE LENGTH
5.18
avg tokens in ground-truth answers — where string metrics peak
RAG RETRIEVAL
Top-5
dialogue segments retrieved per query, appended to the prompt
HUMAN AGREEMENT
0.903
between the two human annotators (judge: 0.801–0.820 vs humans)
LENGTH STUDY
100
representative cases per memory type, varying dialogue turns
Legacy

Impact — Measuring What Matters

LoCoMo-Plus reframes what "remembering" means for an LLM agent — and hands the field a yardstick that can actually see the difference.

🧭 A new definition of memory
Constraint consistency replaces recall: memory is graded by whether behavior respects latent constraints (causal, state, goal, value), not by string overlap.
🕳️ A blind spot, exposed
Failures that LoCoMo- and LongMemEval-style factual QA never surface — models silently forgetting your goals — now have a benchmark that catches them.
⚖️ Evaluation hygiene
Shows that task-disclosed prompting and EM/F1/BLEU/ROUGE distort scores (prompt bias + length bias) — and offers a unified, judge-based alternative.
📉 A 2026 reality check
The best method scores 26.06 on cognitive memory. Memory systems don't close the gap (A-Mem 17.20, Mem0 15.80, SeCom 14.90) — differences compress toward uniformly low scores.
🔍 Retrieval is not memory
Cue–trigger semantic disconnect defeats BM25 and embedding search by construction — remembering must mean applying, not finding.
🧪 Open and auditable
The benchmark, judge prompts, annotations, and LLM-judge rationales are released at github.com/xjtuleeyf/Locomo-Plus for inspection and reproducibility.
What It Did NOT Solve
Deep Dive

Memory Beyond Recall: Grading Consistency

LoCoMo-Plus makes the hardest distinction in memory evaluation: remembering facts is Level 1; behaving consistently with latent constraints is Level 2. A system that can recite "user has a puppy" and still recommends a spa weekend has a memory and no understanding — and on this benchmark, every one of 17 methods bleeds points between the two levels.

🎯
A Benchmark Built to Resist Shortcuts
  • Cue-trigger pairs are semantically disconnected — the trigger shares almost no words with the cue
  • BM25 and embedding retrieval cannot shortcut the task by construction
  • Human-validated pairs embedded in ~300-turn LoCoMo dialogues — built to measure, not to train on
  • LLM judges, evidence-grounded: 0.903 human-human agreement, stable across judge backbones
📉
The Universal Drop
  • 17 methods tested — from Qwen2.5-3B to gemini-2.5-pro, Mem0, A-Mem, SeCom — all drop 12–46 points
  • Best cognitive score: 26.06 — the ceiling of the entire field on constraint consistency
  • Level-1 factual recall stays respectable: recall is solved, cognition is not
  • No task disclosure to the system — zero-shot constraint application, the honest setting
Interactive Demo — The Semantic Disconnect

Run the retrieval both ways on the same trigger. BM25 scores overlap between the cue and the trigger — watch it find nothing. Then run constraint-aware reading: the cue is reachable, the constraint applicable, and the response graded against the space of consistent answers 𝒜𝒸.

THE CUE · SESSION 2
"We adopted a golden retriever puppy — she cries if left alone for even an hour."
THE TRIGGER · SESSION 31
"Plan me a perfect weekend — I need to get out of the house."
VERDICT
The next benchmark always measures what current systems fake
The trajectory of this track is now legible: LongMemEval asked whether facts survive length; LoCoMo-Plus asks whether implications survive at all. Causal, state, goal and value constraints have a whole answer space 𝒜𝒸, not one string — which is why LLM-judge evaluation with evidence grounding was the only scoring that worked (0.903 human agreement). Every memory system graded here, including A-MEM, inherits the same homework: representation must carry meaning, not just tokens.
📚 Two levels, one ladder
Level-1: explicit facts, single ground truth. Level-2: implicit constraints, a space of valid behaviors. Systems 40+ points apart on Level-1 converge near 20 on Level-2 — difficulty equalizes the field.
🚫 No leakage by design
No task disclosure, no fine-tuning on the benchmark: the dialogues exist to measure, not to train. Contamination-proofing as a first-class design goal.
🧑‍⚖️ Judging the space
With a valid-answer space rather than a string, evaluation needs judgment — hence evidence-grounded LLM judges with 0.801–0.820 human agreement, stable across judge backbones.
🏁 26.06
The best cognitive score across 17 methods. Hold that number: it is the field's frontier, and it is a failing grade. Benchmarks exist to be humbled by exactly this.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the LoCoMo-Plus paper.

Reference

Key Takeaways

Everything you need to remember about this paper (yes — memory pun intended).

✅ Memory ≠ recall: LoCoMo-Plus grades constraint consistency — does behavior respect latent constraints (causal, state, goal, value), not strings.
✅ Cue–trigger semantic disconnect: triggers share almost no words with their cues, so BM25/embedding retrieval cannot shortcut the task.
✅ Two levels: Level-1 factual memory (explicit facts, one ground truth) vs Level-2 cognitive memory (implicit constraints, a whole space 𝒜𝒸 of valid answers).
✅ Evaluation reform: no task disclosure, evidence-grounded LLM judges — 0.903 human–human agreement, 0.801–0.820 human–judge, stable across judge backbones.
✅ 17 methods, one verdict: every method — from Qwen2.5-3B to gemini-2.5-pro, Mem0, A-Mem, SeCom — drops 12–46 points; the best cognitive score is 26.06.
✅ Diagnostic by design: human-validated cue–trigger pairs embedded in ~300-turn LoCoMo dialogues — built to measure, not to train on.