Research
CoEM: Empowering Long-Context Reasoning with Commit-on-Evidence Memory
Overview Research area: Long-context reasoning with large language models, specifically recurrent (chunk-by-chunk) memory agents that carry a bounded textual memory across reading steps. Technical lev

- arXiv
- 2609.36935
- Published
- 2026-09-29
- Authors
- Jingguang Li, Yebo Wu, Zuyi Guo, Kailang Ma, Xianjie Dai, Han Zheng, Benwang Chen, Li Li, Can Rong, Heye Huang
AI summary
Overview
Research area: Long-context reasoning with large language models, specifically recurrent (chunk-by-chunk) memory agents that carry a bounded textual memory across reading steps.
Technical level: Advanced. The paper assumes familiarity with reinforcement learning (clipped policy-gradient objectives, group-relative advantages, KL regularization), partially observable Markov decision processes, natural language inference verifiers, and multi-hop question answering evaluation.
One-sentence scope: The paper proposes Commit-on-Evidence Memory (CoEM), a recurrent memory agent that defers irreversible compression by keeping unresolved source excerpts verbatim in a bounded "pending set" and uses a learned Promote/Keep/Drop policy plus a frozen verifier to decide when to commit evidence into memory.
What This Paper Is About
Existing recurrent memory agents read a long document stream chunk by chunk and immediately summarize what matters into a bounded textual memory. The problem is that a detail which looks irrelevant when first encountered may become essential several chunks later, and once it is compressed away or discarded it cannot be recovered. CoEM's goal is to let the model postpone that irreversible decision, holding short verbatim excerpts aside until later context clarifies whether they matter, while still staying inside a fixed memory budget.
Key Contributions
-
Identification of a failure mode. The authors show that premature information compression in recurrent memory agents irreversibly discards evidence that appears unimportant when first observed but later becomes necessary. Across three multi-hop QA benchmarks, they report that 32.03%–44.53% of randomly sampled questions exhibit this "delayed evidence relevance."
-
The CoEM architecture. A fixed context-memory budget is split into a committed memory of compact, source-supported facts and a pending set of short verbatim source excerpts. At each new chunk, a trained policy chooses Promote (convert pending evidence into a committed fact), Keep (retain it unchanged for later), or Drop (release its capacity) for each pending item.
-
Source-verified commitment. A frozen DeBERTaV3-small NLI cross-encoder checks that a proposed fact is entailed by its cited excerpts before it enters committed memory. Memory updates are treated as atomic transactions: if any new entry fails verification or the resulting memory exceeds its budget, the entire update is rejected.
-
Step-level evidence rewards. The policy is trained with reinforcement learning that combines a final answer reward with a dense, step-level evidence reward penalizing verifier-rejected facts and invalid operations, using a return-to-go so that later verification feedback propagates back to earlier retention and commitment decisions.
Main Findings
-
Consistent gains at every length and both backbones. With Qwen3.5-9B, CoEM's margin over the strongest memory baseline grows from 1.3–2.1 F1 points at 50 documents to 10.4–11.4 F1 points at 6,400 documents. The same trend holds with Qwen3.5-4B.
-
Slower degradation as context grows. Over the range of 50 to 6,400 documents, CoEM's answer F1 declines by only 1.0–2.1 points, compared with a 9.6–11.3-point decline for ReMemR1.
-
Large advantage over general-purpose LLMs on distractor-heavy inputs. At 1,600 documents, CoEM with Qwen3.5-9B outperforms the general-purpose LLMs by 17.5 F1 points on HotpotQA and 21.0 F1 points on both 2WikiMultiHopQA and MuSiQue. Inputs above 1,600 documents exceed those models' configured context window.
-
Generalization to unseen datasets. Without additional training, CoEM performs best at every input length on 2WikiMultiHopQA and MuSiQue, and ReMemR1's native archive-enabled configuration still trails CoEM by 2.3–4.3 average F1 points.
-
Handling deliberately delayed evidence. When identical passages are reordered so the source-to-bridge gap reaches 128 chunks, CoEM reaches 74.5 F1 with Qwen3.5-4B, 25.3 points above the strongest baseline. Because the available evidence is unchanged, the widening gap supports the value of deferring decisions.
-
Pending-set occupancy stays within budget. In a 32-step 2WikiMultiHopQA episode at 1,600 documents with a 256-token pending budget, occupancy stays within budget with three capacity events and a terminal decrease to zero, as promotions and drops release space.
-
Capacity pressure under stress. Forced drops per admitted pending record rise from 0.5% under ordinary ordering to 1.8% with a 128-chunk gap and 7.6% under source-dense stress.
-
The 256/768 split is a sweet spot. Retraining Qwen3.5-4B with different pending-set allocations at a fixed 1,024-token total, answer F1 on HotpotQA at 800 documents peaks at 85.5 with a 256-token pending budget.
-
A lightweight verifier is nearly as faithful and far faster. On 2,000 blinded promotion attempts, the calibrated DeBERTaV3-small verifier achieves 97.2% faithfulness at 4.3 ms per source–fact pair, versus Qwen3.5-9B's 97.5% at 295.0 ms per pair — roughly 69 times lower per-pair latency.
-
Competitive inference cost. At 6,400 documents, CoEM takes 548 s per question, 24–42% less than MemAgent and both ReMemR1 variants. It is 15.6% slower than GRU-Mem but gains 8.8 F1 points.
-
Ablations confirm each component. With Qwen3.5-4B at 800 documents, removing RL costs 18.0/21.0 F1 points (HotpotQA/2Wiki); removing pending retention and verification costs 12.1/14.3; adding verification to that eager variant recovers 3.8/4.2 but remains 8.3/10.1 below CoEM; removing the verifier while keeping pending retention costs 4.0/4.8 F1 and 10.5 percentage points of faithfulness; replacing step-level evidence rewards with outcome-only RL costs 5.1/6.2; and replacing the learned policy with a rigid two-chunk delay costs 3.2/3.9. Full CoEM scores 85.5/76.1 F1 with 97.2/96.7 faithfulness.
Methodology in Plain English
The authors build on the recurrent-memory setup from prior work: a model reads a stream of 5,000-token chunks one at a time and carries a bounded memory forward, then answers after the last chunk. They reframe this as a partially observable problem — future chunks are unseen and past chunks cannot be revisited.
Their key change is to split the fixed 1,024-token memory into two parts. The committed memory (768 tokens) holds compact facts with references to their source. The pending set (256 tokens) holds short, verbatim excerpts whose relevance is not yet clear. When a new chunk arrives, the policy looks at the question, the new chunk, and the previous state, and for each pending excerpt decides to Promote, Keep, or Drop. Keeping has no age limit; if the pending set is full, the system resolves the oldest items until the new evidence fits. At the final chunk, Keep is disabled, so the pending set is emptied and the answer is generated from committed memory alone.
Before any fact is committed, a frozen NLI verifier (DeBERTaV3-small, acceptance threshold 0.90) checks whether the cited excerpts entail the proposed fact, and a validity check confirms the cited spans are still available and match the source text exactly. Removing old facts to make room is handled as a single atomic transaction: if anything fails verification or the budget check, nothing changes.
To train the policy, the authors sample eight complete rollouts per question with a frozen old policy, then compute two advantages: one from the final answer reward, and one from a step-level evidence reward that counts verifier rejections and invalid operations. The evidence reward is accumulated as a return-to-go, so a rejection at a later step lowers the advantage credited to earlier admission and Keep decisions. Both advantages are centered within the rollout group and blended by a weight α, then applied through a clipped policy-gradient objective with a KL penalty against the frozen reference policy. Training used 32,768 HotpotQA samples each paired with 200 documents on eight NVIDIA H800 GPUs; evaluation sampled 128 held-out questions per dataset at eight context lengths from 50 to 6,400 documents, with a 9,216-token serving window.
Why This Matters
Impact on research. The paper reframes long-context memory as a question of when to compress, not only what to remember. It gives a concrete mechanism (a bounded pending set), a verification gate, and a dense reward signal for training deferred decisions, and it quantifies a failure mode — delayed evidence relevance, reported at 32.03%–44.53% of sampled questions — that prior recurrent memory agents cannot address by design.
Real-world applications:
- Cross-document multi-hop question answering, such as legal, regulatory, or investigative research where a fact in one document only becomes meaningful after reading another.
- Repository-level debugging and code understanding, where a seemingly unrelated source file becomes the key to a bug once a later file is read. The authors explicitly name code as a natural extension.
- Scientific evidence synthesis, where findings distributed across distant papers must be combined and where provenance of each committed claim matters.
- Long-horizon agent pipelines that must operate under a fixed context budget rather than an ever-growing transcript.
Industry relevance. The method runs inside a fixed 1,024-token memory budget and a 9,216-token serving window, which directly bounds serving cost — the practical bottleneck for deploying long-context agents. The 4.3 ms per-pair cost of the verifier makes source-grounded memory affordable at scale, and the reported 548 s per question at 6,400 documents is competitive with, and in several comparisons cheaper than, existing recurrent-memory baselines. Faithfulness rates near 97% also matter for any deployment where an agent's stored memory must be traceable to a source.
Future Directions
-
Extending beyond question answering. The authors identify code and scientific documents as the natural next domain, where evidence structure and provenance requirements differ from multi-hop QA.
-
Relaxing the fixed pending budget. Forced drops rise from 0.5% to 7.6% under source-dense stress at a 256-token pending budget, so how to allocate or adapt that budget when unresolved evidence accumulates remains open.
-
Replacing the frozen NLI verifier. The current verifier trades false acceptances for false rejections after calibration; a stronger or learned verifier could reduce both error types without the 295.0 ms per-pair cost of a large model.
-
Deciding when to commit, beyond a learned policy. The ablation shows a rigid two-chunk delay loses 3.2/3.9 F1 points versus the learned policy, suggesting there is room to study alternative commitment-timing mechanisms and the theoretical conditions under which deferred commitment helps.
Target Audience
Researchers and engineers working on long-context LLM inference, agent memory architectures, and retrieval-augmented or recurrent memory systems. It is most useful to those already familiar with reinforcement learning for LLM agents and with multi-hop QA benchmarks such as HotpotQA, 2WikiMultiHopQA, and MuSiQue. Practitioners building production agents under fixed context budgets will find the efficiency and faithfulness analyses directly actionable, while readers new to the area will need background in memory-augmented LLMs and policy-gradient methods to follow Sections 3.3 and 3.4.
Authors’ abstract
Long-context reasoning is essential for complex and long-horizon tasks, yet the performance of large language models (LLMs) degrades as context length increases. Recent approaches address this by processing input chunk by chunk while maintaining a bounded textual memory in model context. However, premature information compression can discard critical details essential for subsequent reasoning. In this paper, we introduce Commit-on-Evidence Memory (CoEM), which learns when to convert source evidence into compact memory facts. Specifically, under a fixed context-memory budget, CoEM preserves potentially useful source excerpts verbatim in a pending set, allowing subsequent context to clarify their relevance before irreversible compression. As new context arrives, a learned policy revisits each pending excerpt and decides whether to promote it to the committed memory, retain it for further consideration, or discard it. A frozen verifier ensures proposed facts are accepted only if supported by retained excerpts and current context. To further guide effective memory management, we train this policy using reinforcement learning by combining fine-grained, step-level evidence rewards with final answer rewards. Extensive experiments demonstrate that CoEM consistently improves long-context reasoning. When evaluated on 6,400 documents long-context input, CoEM outperforms the strongest memory baseline by 10.4-11.4 F1 points on Qwen3.5-9B. Code repository: https://github.com/benmagnifico/CoEM.