A visual, step-by-step guide to the 2026 benchmark that asks a harder question about LLM agents: not "can you recall what I said?" but "can you remember my goals, states, and values — and let those latent constraints steer your advice hundreds of turns later?"
Long-term conversational memory went from a developer hack to a measured capability in about three years — but the yardstick kept measuring the same thing: facts. LoCoMo-Plus changes the yardstick.
Real memory constrains behavior instead of answering quizzes. In the paper's own motivating example, a user says they're preparing for an important exam and want to minimize distractions. Much later, they ask: "Should I start watching that new TV series everyone is talking about?"
No fact is being recalled — several replies look perfectly fine in isolation. The right one depends on remembering the earlier goal and noticing the conflict. Benchmarks built on factual recall can't see this skill at all.
Existing benchmarks mostly check whether a model can find a fact that was explicitly stated earlier. But in real conversations, memory's job is subtler: it quietly constrains every later response.
A friend who aces trivia night about your life — your job, your dog's name, your birthday — but offers you an espresso at 11pm before your 6am flight has facts, not memory. Facts answer queries; memory shapes behavior when nobody is querying anything. LoCoMo-Plus is the first benchmark in this lineage built to grade the second kind.
The paper's formal framing splits conversational memory in two. Level-1 is what every prior benchmark measures. Level-2 is what actually makes an assistant feel like it knows you.
Relevant information is explicitly stated in the history and can be directly recalled or reasoned over — localized object-centric facts ("my dog is named Pepper") and event-oriented episodic details ("we hiked on May 3rd"). Each query admits a single well-defined ground-truth answer, and correctness is string or semantic similarity against it. This is the regime of LoCoMo, LongMemEval, and most QA-style memory tests.
Realistic conversations depend on implicit constraints inferred from prior turns — user state, goals, preferences, values, causal context. There is no single correct response: the history induces a latent constraint c that restricts the space of behaviorally valid outputs. Any response inside that space counts, regardless of wording. This is the regime LoCoMo-Plus opens for measurement.
LoCoMo-Plus decomposes cognitive memory into four constraint types. Every instance in the benchmark is a cue that plants one of them — and a trigger, far away in the conversation, that only makes sense if the model kept it. The examples below are drawn directly from the paper's Appendix C.
LoCoMo-Plus instances are constructed from scratch and embedded into LoCoMo's long dialogues. The pipeline is deliberately heavy on human validation and filtering — diagnostic coverage is prioritized over scale.
| Aspect | LoCoMo-style protocol | LoCoMo-Plus framework |
|---|---|---|
| Input side | Task-disclosed prompts — the model is told which memory skill is tested | Natural dialogue continuations, no task disclosure |
| Output side | EM, token-F1, BLEU, ROUGE vs a reference string | LLM judge scores constraint consistency, grounded in dialogue evidence |
| Ground truth | One reference answer per question | A valid response space 𝒜𝒸 — many correct surface forms |
| Labels | Binary correctness (typical QA) | 3-level (correct / partial / wrong) for factual & commonsense; binary for temporal, adversarial, cognitive |
| Known bias | Prompt adaptation + length bias (scores peak near the 5.18-token average reference length) | Judge robustness measured against humans (0.801–0.820) and across backbones |
The paper evaluates 17 methods across four categories: open-source LLMs, closed-source LLMs, RAG pipelines, and dedicated memory systems. One pattern holds everywhere — a large, persistent gap between factual and cognitive memory.
| Method | LoCoMo Avg | LoCoMo-Plus | Gap |
|---|---|---|---|
| gemini-2.5-pro | 71.78 | 26.06 | 45.72 |
| gemini-2.5-flash | 69.25 | 24.67 | 44.58 |
| gpt-4o (full context) | 62.99 | 21.05 | 41.94 |
| gpt-4.1 | 62.21 | 18.63 | 43.58 |
| Qwen3-14B | 59.65 | 19.09 | 40.56 |
| A-Mem (GPT-4o) | 59.64 | 17.20 | 42.44 |
| SeCom (GPT-4o) | 57.53 | 14.90 | 42.63 |
| Mem0 (GPT-4o) | 57.24 | 15.80 | 41.44 |
| gpt-5-nano | 54.96 | 14.84 | 40.12 |
| RAG top-5, emb-large (GPT-4o) | 45.32 | 15.55 | 29.77 |
| Qwen2.5-7B-Instruct | 45.31 | 9.57 | 35.74 |
Scores / 100. Gaps run ~12–46 points across all 17 methods. Memory systems and retrieval pipelines do not close the cognitive gap — and differences between methods compress toward uniformly low performance.
When benchmarks announce the task ("this is a temporal-reasoning question"), measured ability profiles shift — temporal and adversarial categories get disproportionately inflated scores under task-disclosed evaluation. Reported gains partly reflect sensitivity to prompts, not stable memory behavior. Real users never announce "this query requires memory."
EM, F1, BLEU, and ROUGE all vary systematically with output length — scores peak near the average reference length (5.18 tokens) and degrade as generations get shorter or longer. Models are penalized or favored by verbosity alone, regardless of semantic correctness — a systematic bias in cross-model comparison.
LoCoMo-Plus reframes what "remembering" means for an LLM agent — and hands the field a yardstick that can actually see the difference.
LoCoMo-Plus makes the hardest distinction in memory evaluation: remembering facts is Level 1; behaving consistently with latent constraints is Level 2. A system that can recite "user has a puppy" and still recommends a spa weekend has a memory and no understanding — and on this benchmark, every one of 17 methods bleeds points between the two levels.
Check your understanding of the key concepts from the LoCoMo-Plus paper.
Everything you need to remember about this paper (yes — memory pun intended).