Research
Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA
Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA Overview Research area: Long-term memory retrieval for LLM agents that converse with a user across many sessi

- arXiv
- 2610.09348
- Published
- 2026-10-07
- Authors
- Yufeng Li, Shuxin Li, Zhenhua Xu, Junxian Li, Peng Zeng, Sheng Yao, Changting Lin, Gaolei Li, Ran Bi, Meng Han
AI summary
Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QAOverview
- Research area: Long-term memory retrieval for LLM agents that converse with a user across many sessions, evaluated on multi-session question answering benchmarks (LoCoMo and LongMemEval-S).
- Technical level: Advanced. The paper combines an evidence-law framing of "sufficiency," a Formal Concept Analysis (FCA) prefilter, set-level LLM-judged metrics, and multi-view budgeted retrieval.
- Scope: The paper argues that relevance-ranked memory retrieval cannot guarantee an answerable record, formalizes memory retrieval as construction of a sufficient memory set, and proposes Budgeted Flat Reconstruction (BFR) to build such sets over a fixed flat memory store under a retrieval-call budget.
What This Paper Is About
LLM agents accumulate interaction histories too long to keep in a context window, so they store the past externally and answer each question from a handful of retrieved records. Existing memory systems (Mem0, A-Mem, CoM, MRAgent and others) pick those records by lexical or embedding relevance, but a top-ranked set can be individually relevant to the topic while jointly missing the one complementary fact the answer needs — for example, retrieving the February 5 order date for a remote shutter release but not the February 10 arrival date, so the five-day interval cannot be computed. Drawing on the distinction in legal evidence scholarship between relevance (a property of an item) and sufficiency (a property of the assembled record), the paper's goal is to make set-level sufficiency measurable and to build retrieval methods that construct a set from which the answer can actually be determined.
Key Contributions
- Reframing and measuring set-level sufficiency. The paper formulates memory retrieval as construction of a sufficient memory set and makes sufficiency measurable with three instruments: LLM-Suff@Set (a blinded, reference-conditioned LLM judgment of whether the retrieved set supports the answer), Gold Hit (whether the set reaches at least one annotated supporting record), and Turn Hit (whether it reaches at least one annotated answer-bearing turn).
- The BFR two-stage method. Budgeted Flat Reconstruction builds sufficient sets over a fixed flat store: Stage I applies Formal Concept Analysis for Memory Selection (FCA-MS) to decompose the question into requirements and select a compact complementary subset, grounded to source records as an initial set; Stage II repeatedly acquires unseen records through text access (BFR-Text) or text, entity, and session views (BFR-MV) until a fixed budget is exhausted.
- Controlled same-store experiments on two benchmarks. Experiments on LoCoMo and LongMemEval-S hold the memory store, answerer, and judge fixed within each benchmark and compare BFR against same-store adaptations of Mem0, A-Mem, CoM, and MRAgent, showing improvements in both answer quality and evidence coverage.
- Attribution of where the gains come from. Ablations and cost curves attribute most of the improvement to broader evidence acquisition beyond the initial candidate pool rather than to adaptive state control or reselection within the pool, and separate evidence composition from evidence acquisition.
Main Findings
- BFR leads same-store baselines on LoCoMo (n = 200). Overall Stem F1 / LLM-Judge: BFR 68.3 / 70.5, versus MRAgent-Flat* 58.0 / 64.5, Mem0 55.3 / 59.0, CoM* 50.0 / 54.0, and A-Mem 48.7 / 53.0. BFR leads on Multi-Hop (44.6 / 48.5), Temporal (78.1 / 80.0), and Single-Hop (75.7 / 77.1). The small Open Domain slice is reported without a separate type-level conclusion.
- BFR leads on LongMemEval-S (n = 500). Overall Judge accuracy rises from 72.4% (Mem0) to 82.2% (BFR), and Turn Hit reaches 91.4%. Other systems: A-Mem 68.6% Judge / 82.2% Turn Hit, CoM* 68.6% / 80.6%, MRAgent-Flat* 71.4% / 78.8%, Mem0 72.4% / 64.8%. BFR's Temporal question accuracy is 84.2%, versus 65.4% for Mem0, 63.2% for MRAgent-Flat*, and 57.1% for A-Mem and CoM*.
- Set-level sufficiency improves, not just coverage. LLM-Suff@Set on LongMemEval-S: BFR-MV 65.0%, BFR-Text 63.0%, MRAgent-Flat* 56.0%, CoM* 49.0%, Mem0 48.0%, A-Mem 38.0%. On LoCoMo: BFR-MV 63.5%, BFR-Text 62.0%, MRAgent-Flat* 54.5%, Mem0 45.5%, CoM* 44.0%, A-Mem 36.5%.
- Annotated-evidence proxies do not equal sufficiency. On LongMemEval-S, A-Mem has higher Turn Hit (82.2%) than Mem0 (64.8%) but lower LLM-Suff@Set (38.0% vs. 48.0%). On LoCoMo, BFR-Text has higher overall CER while BFR-MV has higher LLM-Suff@Set. The illustrative delivery-time table shows each singleton record has Gold Hit 1 but LLM-Suff@Set 0, and only the pair reaches LLM-Suff@Set 1.
- Stage I alone is not enough. FCA-MS alone raises LoCoMo Gold Hit from 50.0% to 51.0% and LongMemEval-S Turn Hit from 82.2% to 85.0%, but trails both completion schedules (LoCoMo Gold Hit: BFR-Text 72.5%, BFR-MV 73.5%; LongMemEval-S Turn Hit: BFR-Text 88.6%, BFR-MV 91.4%; FCA-MS alone 51.0% and 85.0%).
- Both stages contribute to answer quality. Ablation on LoCoMo Stem F1 / Judge and LongMemEval-S Judge / Turn Hit: FCA-MS alone 48.6 / 53.5 and 70.4 / 85.0; with BFR-Text 67.8 / 70.0 and 81.6 / 88.6; with BFR-MV 68.3 / 70.5 and 82.2 / 91.4.
- Completion closes type-specific evidence gaps on LoCoMo. Complete Evidence Recall / LLM-Suff@Set overall: FCA-MS 40.5 / 42.5, BFR-Text 63.0 / 62.0, BFR-MV 62.0 / 63.5. On Single-Hop (n = 109), CER goes from 41.3 to 71.6 (Text) and 70.6 (MV); on Temporal (n = 50), from 66.0 to 82.0 for both.
- Most of the gain arrives early in the budget. On LoCoMo (n = 200), most of the answer-quality gain over FCA-MS appears after the first completion call. BFR-Text C4 and C6 coincide at four realized calls. BFR-MV reaches its highest scores at C9 but requires nearly nine calls, and its curves are not monotonic between budgets. The C6 configuration is the fixed default operating point, not a sweep-selected maximum.
- A structured-memory boundary comparison is reported, not as a fair baseline. On the LoCoMo conv-30 overlap (n = 81) under a unified extractive answerer and judge: Full-MRAgent 88.9% Gold Hit / 95.1% Judge, CoM 80.2% / 85.2%, BFR-Text 77.8% / 81.5%, MRAgent-Flat 67.9% / 72.8%. Mean evidence-set sizes: 2.4 (Full-MRAgent), 14.2 (MRAgent-Flat), 34.0 (BFR-Text). Full-MRAgent reaches gold on 12 questions BFR-Text misses, with the reverse on 3; against MRAgent-Flat, the full pipeline is correct alone on 20 questions versus 2.
- The motivating failure case is traced end to end. On LongMemEval-S question b3c15d39, lexical top-k and FCA-MS both reach the order turn but miss the arrival turn (50% case evidence recall, wrong answer), while BFR-MV reaches both turns (100% recall, correct "5 days").
- Gold-free sufficiency estimates are not reliable stopping signals. The authors report that human checks show gold-free sufficiency estimates remain too unreliable to serve as stopping criteria, which is why budget exhaustion is used instead.
Methodology in Plain English
The authors keep one fixed flat memory store per benchmark and change only how evidence is assembled from it, so that answerer and judge stay constant across all compared systems.
Stage I — composing complementary evidence (FCA-MS). The agent memory system is asked for a candidate pool larger than its usual top-k cutoff. An LLM decomposes the question into answer-oriented requirements — distinct facts, bridges, list items, and temporal constraints. A Formal Concept Analysis prefilter narrows that pool using the requirements, and then an LLM selects the smallest memory subset that jointly covers the requirements. Only those selected notes are grounded back to source turns to form the initial evidence set. Importantly, the retriever's original top-k is treated as a reference cutoff, not as a retained base set.
Stage II — completing missing support under a budget. The initial set becomes a cumulative evidence state. At each round the method issues retrieval actions that exclude everything already in the set, so acquisition is duplicate-free. BFR-Text simply re-queries the text index with the original question and takes the next unseen hits (k_keep = 8), which walks past the first page without query rewriting. BFR-MV instead cycles through text, entity, and session views, where the entity action uses the first person named in the question and the session action uses the latest retained record as an anchor. Retention ranks returned records by relevance and by novelty against the current set and keeps at most k_keep. The main budget allows three rounds, six calls, and 30 new records; stopping happens when the budget is exhausted, because the offline sufficiency measures require a reference answer or gold annotations and are therefore unavailable at inference time.
Measurement. Because no oracle declares a set sufficient at inference time, the paper measures sufficiency offline with a blinded, reference-conditioned LLM adjudication (LLM-Suff@Set), plus Gold Hit and Turn Hit as coverage proxies, and, on LoCoMo, Evidence Recall and Complete Evidence Recall over the released supporting-record annotations. Answer quality uses Porter stem F1 and LLM-Judge on LoCoMo and Judge on LongMemEval-S, with gpt-5.5 at temperature zero using prompt mem0_accuracy_v1. Evaluation uses a fixed stratified 200-question LoCoMo development subset and all 500 LongMemEval-S questions.
Why This Matters
Impact on research. The paper shifts the objective of long-term memory retrieval from ranking individually relevant records to constructing a set whose members jointly satisfy a question's information requirements. It supplies a vocabulary (composition versus completion), measurable set-level instruments (LLM-Suff@Set, Gold Hit, Turn Hit, ER, CER), and evidence that coverage proxies can move in the opposite direction from set-level sufficiency — a caution for anyone who reports only retrieval hits as a proxy for answerability.
Real-world applications:
- Personal assistants and companion agents that must answer questions spanning months of conversation, where a needed fact sits in a different session from the topically similar ones.
- Customer-support and CRM copilots reconstructing the history of a case across tickets, where an early order record and a later resolution record must be combined.
- Enterprise knowledge assistants over long document or message archives, where a question depends on a fact list plus a temporal constraint.
- Multi-session agent workflows that need an auditable evidentiary record behind an answer rather than a list of similar snippets.
Industry relevance. The method is implemented at inference time over an existing flat store and does not require rebuilding memory representation, so the paper frames it as a practical complement to systems like Mem0, A-Mem, CoM, and MRAgent. The reported budget is explicit (three rounds, six calls, 30 new records), and the cost curves show most of the answer-quality gain appearing after the first completion call — relevant for teams trading retrieval latency and call cost against accuracy. The boundary comparison also indicates that structured stores can outperform flat acquisition on the 81-question overlap slice, so store design remains a separate lever from retrieval policy.
Future Directions
- Develop reliable inference-time stopping. The paper states that human checks show gold-free sufficiency estimates are too unreliable to serve as stopping criteria, so budget exhaustion is used instead; a trustworthy at-inference sufficiency test remains open.
- Extend the comparison beyond flat stores. The conv-30 overlap slice shows Full-MRAgent at 88.9% Gold Hit / 95.1% Judge with mean evidence-set size 2.4 versus BFR-Text at 77.8% / 81.5% with size 34.0, but store, tools, and derived facts move together in that comparison. The paper notes a causal graph-edge claim would require further work (the truncated text ends mid-sentence here).
- Explain and stabilize the multi-view budget behavior. BFR-MV is not monotonic across budgets and only reaches its highest LoCoMo scores at C9 with nearly nine calls, while BFR-Text C4 and C6 coincide at four realized calls; understanding this schedule behavior is a natural follow-up.
- Test whether requirement decomposition and multi-view access transfer to other benchmarks and store types, given that the reported evidence comes from LoCoMo (200-question subset) and LongMemEval-S (500 questions), with LoCoMo's Open Domain slice containing only eight questions and interpreted descriptively.
Target Audience
Researchers and engineers working on LLM agent memory, retrieval-augmented generation, and long-conversation QA will get the most from this paper, particularly those who design or evaluate retrieval pipelines and need metrics beyond top-k relevance. It is also useful for evaluation-focused readers interested in set-level sufficiency judgments and evidence-coverage proxies, and for practitioners deploying multi-session assistants who must decide retrieval budgets. The paper assumes familiarity with retrieval metrics, embedding or lexical ranking, and LLM-as-judge evaluation; the Formal Concept Analysis component and the evidence-law mapping in the appendix are the most technical parts.
Authors’ abstract
LLM agents that interact with a user across many sessions accumulate histories that exceed their context window, so they store past interactions in an external memory and answer each question from a small set of retrieved records. Existing memory systems rank records by lexical or embedding relevance, yet the top-ranked memories can each be relevant while jointly omitting a complementary fact that the answer requires, especially for multi-session and temporal questions. Drawing on the distinction between relevance and sufficiency in legal evidence scholarship, we recast memory retrieval as constructing a sufficient memory set. To operationalize this view, we introduce a blinded LLM judgment over the retrieved set, together with Gold Hit and Turn Hit as evidence-coverage proxies. We then propose Budgeted Flat Reconstruction (BFR), which builds sufficient sets over a fixed flat memory store in two stages. Specifically, we first apply Formal Concept Analysis for Memory Selection (FCA-MS) to decompose the question into information requirements and select a compact candidate subset that jointly covers them. Then, we repeatedly acquire unseen records through deeper text search or complementary entity and session views, stopping when the budget is exhausted. Experiments on LoCoMo and LongMemEval-S show that BFR outperforms same-store adaptations of recent agent-memory systems in both answer quality and evidence coverage. Specifically, on LongMemEval-S it raises judged accuracy from 72.4% to 82.2% and Turn Hit to 91.4%.