Skip to content
AI.info

Research

The Right Memory in the Wrong Context: Verifying Retrieval Admissibility in Long-Term Agent Memory

Overview Research area: long-term memory for AI agents, retrieval governance, and evaluation methodology for memory-augmented systems. Technical level: Advanced. Scope: a verification framework that j

The Right Memory in the Wrong Context: Verifying Retrieval Admissibility in Long-Term Agent Memory
arXiv
2610.07309
Published
2026-10-05
Authors
Zi Wang, Xingqiao Wang, Emmanuel Addai, Devika Ambekar, Xiaowei Xu

AI summary

Overview

Research area: long-term memory for AI agents, retrieval governance, and evaluation methodology for memory-augmented systems. Technical level: Advanced. Scope: a verification framework that judges whether a retrieved memory is not only relevant but also admissible for the current query context, evaluated on two public long-term-memory benchmarks plus controlled diagnostics and reader experiments.

What This Paper Is About

Long-term agent memory can surface a record that is on-topic but should not be used for the current request, because it belongs to another principal, violates an active policy, or reflects an outdated or forgotten state. Standard recall and final-answer accuracy cannot detect this: a route can look safe simply by returning less evidence, and a correct answer can still follow inadmissible prompt exposure. The paper builds a framework that labels every memory–query pair as admissible, inadmissible, or unresolved, compares retrieval routes at matched evidence recall, and traces individual memory IDs from storage through retrieval and prompt exposure to answer-level disclosure.

Key Contributions

  1. A query-conditioned, three-valued admissibility target (admissible, inadmissible, unresolved) with matched-recall partial-identification bounds, so unresolved labels are never imputed as permission or denial.
  2. A protocol that tracks memory IDs through retrieval and agent-prompt exposure and links exposure to target-level answer disclosure, while separating wrong-scope, policy, and lifecycle violations.
  3. Three stagewise studies — candidate support, eligibility verification, and exposure/disclosure — run on separate, non-pooled populations (benchmark-native and controlled), never pooled together.
  4. A post-hoc top-20 reanalysis of frozen rankings from RHELM and MemOps covering 3,767 queries, reported with the released evidence, reproduction scripts, and a GitHub repository.

Main Findings

  • Candidate support rises under trusted namespaces. Restricting candidates to the query's trusted provenance namespace raises top-20 anchor recall from 0.432 to 0.533 and 80% recall feasibility from 0.237 to 0.311, with paired top-20 differences of +.101 [+.082,+.120] and +.074 [+.056,+.095]. Only .311 of queries reach the .8 target under namespace support.
  • Bounds and loss move together. Bound endpoints fall from [.187,.195] to [.118,.126] at .992 coverage, and the constrained loss falls by -.085 [-.107,-.064]; infeasibility contributes -.074 [-.095,-.056] and feasible-prefix risk contributes -.011 [-.016,-.007], so 87% of the loss reduction is attributable to feasibility.
  • Search work drops sharply. Exact similarity evaluations fall by 98.3%, from 90,122 (global dense) to 1,551 (namespace dense).
  • Identity of the filter matters more than pool size. A ten-seed random same-size partition collapses to .061 recall and .024 feasibility, while global post-filtering at depth 500 nearly matches namespace recall (.531 vs .533) and loss (.713) but retains 90,122 similarity evaluations. An anchor-preserving same-size oracle reaches .840 recall and .747 feasibility and is explicitly an unattainable diagnostic upper bound.
  • Effects are scope-specific. After excluding scope, upper risk is .126 for namespace versus .123 for global. Within feasible prefixes, wrong-scope exposure falls from .084 to zero and lifecycle-incompatible exposure from .019 to .017, but policy-disallowed exposure rises from .094 to .101.
  • Text-only verifiers fail the stated bar. On 72 held-out development cases, the released-metadata reference lowers loss by .032 with unchanged recall/feasibility, while GPT-5.6 Sol and Gemini 3.6 Flash change recall/feasibility/loss by -.042/-.083/+.057 and -.014/-.028/+.018; their ROC-AUC is .626/.522 and required-anchor false denial is .059/.020. At the 1% required-anchor false-denial limit, GPT-5.6 Sol detects nothing and no Gemini 3.6 Flash threshold qualifies.
  • Downstream accuracy improves, but the comparison is observational. Across 1,523 paired benchmark-native cases, namespace routing is associated with judged-accuracy gains of .053–.068 across three readers (DeepSeek V4 Pro +.053 [+.020,+.084], Gemini 3.6 Flash +.068 [+.047,+.090], GPT-5.6 Luna +.066 [+.039,+.096]), with non-answer falling -.038 to -.048 and quality rising +.043 to +.050. Because namespace also raises recall, the paper states this cannot separate admissibility from reachability.
  • Filtering is not uniformly better. Policy intervals cross zero; text verification raises loss by .0068 and lowers Gemini 3.6 Flash accuracy by -.0149 [-.0283,-.0011]; metadata gates raise Gemini 3.6 Flash stale disclosure by +.020 to +.024.
  • Only one reader shows a disclosure contrast. Across 16 controlled exposure scenarios, selectivity gaps are .906/.813/.656/.844 for GPT-5.6 Sol, Gemini 3.6 Flash, DeepSeek V4 Pro, and Claude Opus 5; relevant-inadmissible literal-disclosure contrasts are .000/.000/+.156/+.125, and only DeepSeek V4 Pro's 95% CI [.031,.312] excludes zero.

Methodology in Plain English

The authors treat admissibility as a property of the memory–query pair, not of the memory alone. Each pair gets three diagnostic axes — scope (does the record belong to this principal?), policy permission, and lifecycle compatibility — each valued as established positive, established negative, or unresolved, then combined with a three-valued (strong Kleene) conjunction. Admissibility is the conjunction of those three; usability adds relevance on top. Unresolved never becomes a yes or a no.

Routes are then compared fairly. Every route is truncated at the smallest prefix that reaches a shared target recall of .8 (fixed before evaluation, not tuned), so a route cannot look safer by returning less useful evidence. For that prefix the paper reports known-label risk, label coverage, partial-identification bounds, and non-dilutable companions (whether any violation occurred, and how many). Infeasible routes are charged a unit penalty rather than being dropped.

On top of this, the authors follow a single memory ID through the pipeline using three nested indicators — stored, retrieved, exposed to the reader — enforcing exposure as a subset of retrieval, and separately scoring whether target content appears in the answer. Populations are kept apart: 3,767 RHELM/MemOps evaluation queries (development: 23 groups, 45,978 memories, 1,104 queries; evaluation: 87 groups, 182,908 memories) for support; 96 public-development queries (24 calibration, 72 analysis) plus 16 controlled scenarios for verification; and 1,523 paired benchmark-native cases with 192 controlled pairs per reader for enforcement. Dense arms use the pinned Qwen/Qwen3-Embedding-8B snapshot (4,096-dimensional float32, 512-token maximum). Two PhD-level annotators labeled 207 record–query pairs in 20 blinded packets, yielding Krippendorff's alpha of .814 for relevance, .981 for scope, .670 for lifecycle, and .083 for prohibited status — which is why policy is interpreted as released-field consistency rather than adjudicated truth. Pointwise 95% intervals use 10,000 percentile-bootstrap replicates.

Why This Matters

This work argues that recall and final-answer accuracy are insufficient audit signals for agent memory, because both can hide the fact that inadmissible evidence crossed the prompt boundary. It provides a stagewise accounting protocol that localizes failures in candidate support, verification, prompt assembly, or answer disclosure, and it shows that the same intervention can help one stage and hurt another.

Real-world applications:

  • Enterprise assistants serving multiple tenants, accounts, or projects, where one client's records must not enter another client's prompt.
  • Healthcare and medication agents, where superseded states may be required for "how did it change?" but inadmissible for "what is true now?"
  • Compliance and legal workflows, where policy disallowed or explicitly forgotten records must be excluded before generation and exclusions must be auditable.
  • Privacy-facing personal assistants that must honor deletion/forgetting requests while still answering history questions.

Industry relevance: the results give a concrete design recommendation — constrain candidate support with authenticated provenance before semantic ranking, apply later gates only when their fields are reliable, and surface unresolved cases rather than imputing permission. The negative results (text-only gates over-deny, post-filtering retains full scoring work) are directly actionable for teams building retrieval pipelines at scale.

Future Directions

  • Inferring or verifying namespaces rather than assuming trusted provenance labels, since the reported support result assumes flat candidate sets and trusted support labels.
  • Extending coverage to graph-structured selectors, where unauthenticated structural writes can change which authenticated records get selected — a case the paper explicitly does not cover.
  • Establishing gold-aware over-refusal, internal causal use, model rankings, safety certification, and generalization beyond the tested corruption channels — none of which the authors claim.
  • Broadening the evidence base: the study covers two public sources and only seven RHELM groups, uses one shared Claude Haiku 4.5 judge, a sequential GPT-5.6 Luna replication, a 200-output alternate-judge audit, and 72 fixed-candidate cases, and lacks complete source-execution hardware records, so candidate-count reductions are not wall-clock latency claims.

Target Audience

Researchers and engineers working on agent memory, retrieval governance, provenance, and evaluation methodology; practitioners building multi-tenant or privacy-sensitive retrieval systems; and reviewers who need a stagewise, auditable protocol for judging whether retrieved evidence was eligible, not merely relevant or useful. The paper is advanced in its notation (three-valued logic, partial-identification bounds, matched-prefix accounting), but its central question — may this memory be used here? — is stated in plain terms and is useful to a broad technical audience.

Authors’ abstract

Long-term-memory agents can retrieve relevant information that is inadmissible for the current request because it belongs to another principal, violates policy, or reflects an incompatible lifecycle state. Recall and final-answer accuracy do not reveal this: a route can appear safe by missing required evidence, while a correct answer may follow inadmissible prompt exposure. We introduce a retrieval-admissibility verification framework that assigns each memory-query pair one of three statuses (admissible, inadmissible, or unresolved), compares routes at matched required-evidence recall with bounds for unresolved cases, and tracks memory IDs through prompt exposure while linking exposure to target-level disclosure. We evaluate its stages on separate, non-pooled populations. A post-hoc top-20 reanalysis of frozen rankings from two public long-term-memory benchmarks, RHELM and MemOps, covers 3,767 queries. All released anchors lie within trusted query namespaces; with within-namespace scores unchanged, off-namespace filtering cannot lower their ranks. Top-20 anchor recall increases from 0.432 to 0.533, 80% recall feasibility from 0.237 to 0.311, and exact similarity evaluations decrease by 98.3%. In a frozen 72-case development diagnostic, a released-metadata reference preserves required evidence, whereas neither text-only verifier detects violations under the 1% required-anchor false-denial limit. Across 1,523 paired benchmark-native cases, namespace routing is associated with judged-accuracy gains of 0.053-0.068 across three readers; recall also changes, so this comparison is observational. In 16 controlled exposure scenarios, only one of four reader-specific 95% confidence intervals excludes zero for relevant-inadmissible literal disclosure (+0.156, 95% CI [0.031, 0.312]). Results motivate separate verification of candidate support, admissibility, prompt exposure, and answer disclosure.

Read the original paper