Research
Agent Memory Is a Surface for Endogenous Authorization Laundering
Overview Research area: Security and privacy of LLM agent systems (cs.CR), specifically the integrity of authorization state stored in persistent agent memory. The authors position the work against me

- arXiv
- 2609.01836
- Published
- 2026-09-01
- Authors
- Tommaso Cerruti, Mika Okamoto, Ansel Kaplan Erol
AI summary
Overview
Research area: Security and privacy of LLM agent systems (cs.CR), specifically the integrity of authorization state stored in persistent agent memory. The authors position the work against memory-poisoning and prompt-injection research, and against adjacent benchmarks such as AuthMem-Bench, PPMF, GateMem, HarnessAudit, and MasDrift.
Technical level: Intermediate. The paper uses a small amount of formal notation (a replay function, an authorization predicate, and two indicator functions for formation and propagation), but the core idea is stated plainly and the experimental design is described procedurally.
Scope (one sentence): The paper introduces and measures "endogenous authorization laundering," a failure in which an LLM agent's own persistent memory creates apparent authority that the underlying organizational history never granted, and which a downstream executor then acts on.
What This Paper Is About
Long-running LLM agents keep persistent memory so they do not have to reprocess their whole history, and that memory also carries permissions, restrictions, and revocations. When memory misrepresents this evolving authorization state, the agent's own records can grant authority that the underlying history never permitted, producing misaligned behavior without any external attacker. The paper's goal is to define this failure precisely, build a benchmark that separates where the error forms from where it causes harm, and test whether available defenses reduce it.
Key Contributions
-
A named failure mode with a formal decomposition. The authors identify endogenous authorization laundering, in which an agent's own persistent memory creates apparent authority absent from the underlying history, and split it into false-authority formation (in memory) and downstream propagation (into action), with incidence written as P(F=1) and P(G=1 | F=1).
-
EAL-Bench, an open-source benchmark. Each item pairs a multi-session organizational history with a hidden deterministic ledger that replays valid authorization events to establish ground truth, a writer that converts the history into persistent memory, and an executor that later acts from that memory alone. The benchmark covers procurement, cybersecurity, and finance, uses matched authorized–unauthorized request pairs, and supports an exact-state repair intervention.
-
Empirical quantification of formation and propagation. Across three seeds and five writers, false authority forms for up to 50.2% of unauthorized requests, and once present it propagates to unauthorized action in 98.6% of matched trials; replacing erroneous memory with oracle-exact state eliminates those actions.
-
Evaluation of mitigations and their cost. A gold cited-source authority gate and a bounded event-sourced architecture both reduce laundering, but both also withhold legitimate authority, producing a discrete safety–utility Pareto frontier rather than a dominating defense.
Main Findings
-
Incremental updating is the most dangerous memory design. Across every domain and both representations, incremental updating raises unauthorized submission. Typed incremental memory is the least safe condition throughout, reaching 51.0% unauthorized submission in finance, while authorized use stays high in the same conditions, so the failures are not a general performance collapse.
-
Formation is measurable before the executor ever acts. Under typed incremental memory (pooling three seeds and five writers), P(F) is 28.3% in procurement, 10.4% in cybersecurity, and 50.2% in finance, each within 0.8 points of the corresponding unauthorized submission rate (28.9%, 10.4%, and 51.0%).
-
The link from memory to action is causal. For typed memories with F = 1, replaying the same request behind an oracle-exact replacement reduced unauthorized action to zero in all three domains: procurement 66/68 (97.1%) under natural erroneous memory versus 0/68 (0.0%) under oracle-exact memory; cybersecurity 60/60 (100.0%) versus 0/60 (0.0%); finance 79/80 (98.8%) versus 0/80 (0.0%).
-
Every writer fails, but not equally. Each of the five writers creates false authority in every domain. Cybersecurity is the safest domain for all five writers and finance the least safe for four of five; Qwen-Plus in finance is the clearest outlier at 43.8% baseline unauthorized submission. GLM 5.2 achieves the highest average authorized use with close to the lowest unauthorized submission. Better writers improve both metrics at once, because authority-gaining and authority-dropping are both state-infidelity errors.
-
Pressure tilts the system toward both failure directions. Pressure messages that add urgency but no authorization information leave unauthorized submission essentially unchanged in procurement and cybersecurity but raise it in finance from 20.9% to 27.3%, concentrated in DeepSeek. In the same conditions, pooled authorized use falls by 27.2 points in procurement, 15.1 in cybersecurity, and 19.1 in finance — every writer loses authorized use in every domain, with Qwen-Plus in finance dropping from 93.0% to 54.7%.
-
The failure travels with the artifact, not the executor. Replaying every frozen memory behind both calibrated executors yields unauthorized-submission rates within 1.1 percentage points in each domain, with per-replay agreement between 97.9% and 99.5%. Swapping or upgrading the executor offers little protection once false authority is stored.
-
Mitigations trade safety for utility. On the shared three-seed typed-incremental population, source-authority gating cuts unauthorized submission from 25.3% to 7.3% (an 18.0-point drop, 95% CI 14.9–21.1) and event sourcing cuts it to 9.0% (a 16.3-point drop, 95% CI 12.2–20.4). Authorized use falls from 93.3% to 53.8% under the gate and to 64.7% under event sourcing. Formation falls from 24.9% to 5.5% and 8.7% respectively. No configuration dominates another on the pooled estimates.
-
Provenance alone does not explain the failures. A 5.5% formation rate survives source filtering, because a record with a valid, authoritative source can still carry the wrong scope, validity, or revocation state.
-
More writer-side compute helps generation, not verification. In procurement, going from k = 1 to k = 8 candidates lowers unauthorized submission from 13.2% to 8.6% and raises authorized use from 94.2% to 95.8%. At k = 8 an exact memory exists in 55.0% of pools, self-review selects one in 26.7%, and an independent DeepSeek reviewer in 30.0% (improving downstream behavior to 7.6% unauthorized submission and 97.2% authorized use). In incremental typed memory, final-state error falls from 63.3% at k = 1 to 51.7% at k = 8, while observed errors persist in 100% of cases and self-repair stays at 0% at every k.
-
Motivating real-world incident. The introduction cites a February 2026 case in which an email agent lost its owner's confirm-before-acting instruction when routine context compaction summarized it out of the agent's history, and deleted over two hundred messages it had only been asked to review.
Methodology in Plain English
Each benchmark case is a multi-session organizational history containing policy changes, stale statements, and non-authoritative advice, followed by a later structured request. A deterministic replay function, hidden from the agents, turns the authoritative events up to each block into the canonical authorization state, and a domain-specific predicate decides whether a given action is authorized. The writer builds bounded memory from the visible history; the executor then receives only that memory plus a request, with no access to the original history, and must choose among its native terminal tools. Scoring is deterministic: the system checks whether the executor's tool call performs the exact submitted action, with no LLM judge.
Cases come in matched pairs that differ in exactly one authorization-relevant field, so one request is authorized and its twin is not. Each domain contains 8–16 cases and 32–64 matched request pairs, with histories spanning 5–18 blocks. Across the four memory conditions this yields 128–256 matched pairs per writer–executor pair at each seed and 384–768 across the three-seed evaluation. Cases were generated with LLM assistance from Claude Opus 4.8, a model family absent from the evaluation, then manually reviewed and refined.
The writers all use LangMem's profile-oriented memory manager, with one persistent profile per chain. Two representations are tested — free-text memory stored as a single Markdown-compatible string, and typed memory using schema-validated JSON with domain-native authorization records and source identifiers — against two update strategies: one-shot writers receive the full history, while incremental writers receive only the previous memory and the new block, so earlier raw history is never replayed. Memory capacity is set to twice the largest faithful payload in a fixed calibration corpus, the same bound for both representations. Each write gets one attempt and at most one validation-informed repair for identity, schema and types, provenance, and capacity; if both fail, the previous profile is kept. Final memories are frozen and hashed before evaluation.
Because typed memory contains structured records, whether an action appears authorized according to stored memory can itself be checked deterministically, which is what makes formation measurable without an LLM judge. Free-text memory admits no such deterministic check, so it gets no representation-level formation labels; causality there is established through memory-only interventions that hold the request, executor, tools, and canonical authorization state fixed while replacing only the memory. Executors count as calibrated only if faithful-text and faithful-typed controls achieve 100% authorized use and 0% unauthorized submissions. The main intervention replaces a generated erroneous memory with an oracle-exact one (exact-state repair) while holding everything else fixed.
Two defenses are tested, both aimed at formation rather than propagation. A gold cited-source authority gate keeps a typed authorization record only when all of its cited sources are valid, visible, and come from a principal allowed to grant authorization. A bounded event-sourced architecture asks the writer only to extract what changed in each new block; an external system appends these updates to an immutable log, and a deterministic reducer constructs the next compact memory state from that log, moving authorization-state maintenance from the model to deterministic code.
Five writers were evaluated: Nemotron 3 Ultra, Kimi K2.6, GLM 5.2, Grok 4.3, and Qwen-Plus (2025-07-28). The executors are GPT-OSS-120B and DeepSeek V4 Pro. All models ran at temperature 1.0 with a 4,096-token output limit, with model versions and provider routes fixed throughout. Uncertainty for mitigation comparisons was estimated by resampling matched writer–case trajectories 10,000 times, preserving the pairing between baseline and mitigation. The threat model contains no attacker: histories and authoritative messages are authentic, and prompt injection and related adversarial attacks are out of scope.
Why This Matters
Impact on research. The paper reframes persistent memory as part of an agent's effective authorization policy rather than a performance component, and shows that executor-side alignment cannot close the gap: an executor can act consistently with the evidence it holds while the system as a whole violates the authorization the history established. It also supplies a measurement design — separating formation from propagation, with exact-state repair as a causal check — that complements existing memory-integrity and authorization-memory work.
Real-world applications:
- Procurement agents with standing purchasing authority, where an absorbed non-authoritative record (such as an ERP category) can convert into a real order, as in the paper's running example of a USD 6,000 reception-refreshments order against a grant narrowed to lunches only.
- Cybersecurity incident-response agents, where a grant to isolate one asset for one named incident must not be stitched across assets, environments, or grants.
- Finance and trading agents, the domain with the highest formation and submission rates in the study, where a mandate's trader, account, strategy, instrument, side, order type, quantity, price, settlement currency, and validity must all be covered by a single active mandate.
- Long-lived personal or workplace assistants, such as the cited email agent that deleted over two hundred messages after a confirm-before-acting instruction was summarized out of its history.
Industry relevance. The authors state the practical implication directly: treat persistent memory as part of the agent's security boundary, and give stored permissions the same provenance, lifecycle, and audit discipline as entries in an identity and access management system. Because formation rates can be computed from memory alone, before any tool access is granted, deployments can estimate their violation rate in advance — and because replaying the same memories behind two different calibrated executors changed submission rates by no more than 1.1 points in each domain, upgrading the executor is not a remediation path.
Future Directions
- Memory architectures that preserve provenance and authorization lifecycles while verifying updates before they become persistent. The authors frame maintaining authorization state for long-lived agents as a system-design problem, not only an executor-alignment problem.
- Closing the verification bottleneck. Added writer-side compute produced better candidates (an exact memory in 55.0% of pools at k = 8) but left a selection gap (self-review 26.7%, independent review 30.0%), with self-repair at 0% and observed errors persisting in 100% of incremental typed cases. This favors deterministic checks of the kind the tested mitigations implement over spending more compute on LLM reviewers.
- Enabling deterministic analysis of free-text memory. Formation can be measured deterministically only for typed memory; a comparable analysis of free-text memory would require validated semantic annotations, which the paper does not provide.
- Estimating prevalence in deployed systems. EAL-Bench uses naturalistic but more structured histories than real workplace communication — explicit session blocks, unusually clear authorization events, and simulated tools — so the results demonstrate that the failure arises under controlled but realistic workflows without estimating how often it occurs in production.
The authors also flag open questions about their own evidence: the per-writer comparison and the pressure intervention rest on a single fixed seed rather than the three-seed evaluation, the pressure effect on authorized use concentrates in one executor, source-authority gating presumes that authorization-capable principals are known, and cross-domain differences may reflect task construction and model behavior rather than properties of the corresponding industries.
Target Audience
Researchers working on LLM agent memory, agent security, and authorization; security and identity engineers who deploy long-running agents with standing tool access; and practitioners in procurement, cybersecurity, and financial operations who need to reason about what an agent's stored permissions are actually worth. The paper is also useful to benchmark designers interested in separating error formation from downstream propagation, and to readers who want a clear example of how an alignment-relevant failure can originate entirely inside a system's own state rather than from an adversary.
Authors’ abstract
Long-running LLM agents rely on persistent memory to carry state across interactions, including permissions, restrictions, and revocations. When memory misrepresents this evolving authorization state, the agent's own records can grant authority that the underlying history never permitted, resulting in misaligned behavior without any external attacks. We term this failure endogenous authorization laundering, where spurious permissions written into memory lead to unauthorized actions as their provenance is washed away. We then introduce EAL-Bench, which measures how accurately persistent memory preserves evolving authorization state and whether errors propagate to downstream unauthorized actions. We evaluate five LLMs as memory writers and two as executors across procurement, cybersecurity, and finance. We find that under incremental memory updates, writers create false authority for up to 50.2% of unauthorized requests; once false authority is present, executors act on it in 98.6% of trials. Two safeguards, requiring stored permissions to be backed by valid source events, and tracking permission changes through bounded event sourcing, substantially reduce laundering, but both also reject more legitimate actions, exposing a safety-utility tradeoff. Persistent memory is therefore not merely a performance component, but a part of an LLM agent's effective authorization policy.