Research
Context Language Models
Overview Research area: Artificial intelligence — long-horizon LLM agents, context/memory management, reinforcement learning post-training, and inference serving efficiency. Technical level: Advanced

- arXiv
- 2609.37725
- Published
- 2026-09-29
- Authors
- Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin, Yuetai Li, Minheng Wang, Hamish Ivison, Radha Poovendran, Nathan Lambert, Teng Xiao, Mike Lewis, Wen-tau Yih, Luke Zettlemoyer, Pang Wei Koh
AI summary
Overview
- Research area: Artificial intelligence — long-horizon LLM agents, context/memory management, reinforcement learning post-training, and inference serving efficiency.
- Technical level: Advanced (assumes familiarity with language model serving, KV/prefix caching, GRPO, and agent harnesses).
- Scope: The paper proposes Context Language Models (CLMs), in which the model itself — rather than an external harness — edits its own live context, implemented by treating the context as an editable file, and evaluates this design zero-shot, with in-context learning, with reinforcement learning, and with a co-designed serving strategy.
What This Paper Is About
Standard language models only append tokens to their context, so deciding what to keep, compress, or discard is handled by external "harnesses" — fixed code or a small menu of predefined tools such as compaction, offloading, and retrieval. The paper asks what happens if the model is given full, unrestricted control over its own context instead, and whether such control can be taught in context or in model weights and served efficiently. The goal is to show that making context management an intrinsic model capability yields better accuracy at lower compute than human-designed context-management strategies.
Key Contributions
-
The CLM formulation and a "context as a file" implementation. The authors generalize the standard append-only context update, c_{t+1} = c_t ⊕ f^LM_θ(c_t), to a fully model-controlled transition, c_{t+1} = f^CLM_θ(c_t). Concretely, the live context is mirrored into an editable file whose path is given in the system prompt; the model may append tokens or use general Bash commands to rewrite the file, with each edit synchronized back to the live context for the next turn. The design extends to multi-agent settings, where multiple context files coexist (agent swarms) or are created and deleted (subagents).
-
ContextBench, a diagnostic benchmark, plus a pilot study of existing strategies. ContextBench contains four synthetic tasks — Needle Retention, Sudoku Sketchpad, KV Store, and Log Triage — that decouple context management from reasoning and knowledge. With a 32K context limit and context pressure up to 24×, the pilot shows that Codex-style Summary, Context Folding, RLM, Self-Compact, ACM, and the bare Mini-SWE-Agent harness all fail even on these simple tasks.
-
Zero-shot evaluation across long-horizon agentic tasks. CLMs are tested out of the box (no training) against MEM1, Self-Compact, ACM, and RLM on coding, deep research, mathematical optimization, 12-hour single-repository optimization, and a 24-hour six-repository agent-swarm task, reporting both task scores and a new compute metric.
-
Learning context management in context and in weights, plus a serving co-design. The paper steers CLMs with natural-language instructions, evolves skill documents through a prompt-evolution loop, introduces a success-gated efficiency advantage for stepwise GRPO, and co-designs Suffix Cache Reuse (SCR) to cut re-prefilling costs during serving.
Main Findings
- BrowseComp-Plus (BCP): Zero-shot CLMs reach 59.4% at a 32K context limit, exceeding the strongest baseline (Codex-style summarization) by 11.4% relative, while using 21.5% and 28.9% fewer prefix-reuse FLOPs than Codex-style summarization and MEM1, respectively.
- Terminal coding: CLMs match the strongest baseline, Codex-style summarization, on TerminalBench 2.1 while using only 70% of its prefix-reuse FLOPs, and exceed it on TBLite (73.7% versus 67.0%) with 91% of its FLOPs.
- Mathematical optimization: With Claude 4.6 Sonnet and a 32K context limit (capped at 100 scored attempts or five hours), CLM achieves the highest best-of-run score on all four problems versus OpenEvolve and OpenEvolve-Agent — circle packing 2.618 (CLM) and 2.636 (CLM with subagents) versus 2.541 and 2.525; Heilbronn 0.03653 and 0.03617 versus 0.03127 and 0.03053; min-max/min-dist 0.07758 versus 0.07690 and 0.07724; Erdős overlap 0.38094 and 0.38109 versus 0.38123 and 0.38167. The abstract reports gains of up to 16.8% (Heilbronn) and 3.0% (circle packing).
- 12-hour single-repository optimization (EdgeBench-10): With Qwen3.6-27B, CLM reaches 44.6 at 179 prefix-reuse PFLOPs per trial versus 42.3 at 437 PFLOPs for summarization; the subagent variant reaches 44.2 at 181 PFLOPs. With Claude 4.6 Sonnet, CLM and its subagent variant reach 51.0 and 50.4 versus 42.3 for summarization. Subagents provide little additional benefit on this benchmark.
- 24-hour multi-repository agent swarms (Software World): With GPT-5.6-Sol and a 272K context budget, six agents jointly optimize interdependent Python repositories and are evaluated on four unseen downstream packages; CLM achieves 65% greater downstream speedup than a summary-based agent swarm at the same spend.
- Steering by instruction: A single added sentence can change compaction timing, move compaction to semantic sub-question boundaries, or make the agent back up context before compacting, with no harness or parameter change.
- Skill evolution: On ContextBench with a 32K budget, a prompt-evolution loop improves held-out accuracy by up to 35.9 points at lower compute; assisted evolution used Qwen3.6-27B with Claude Fable 5.1 as proposer, and self-evolution used Opus 5 for both roles.
- Reinforcement learning: Training Qwen3.5-9B on OpenResearcher raises BCP accuracy from 28.8% to 42.5% (a 47.6% relative improvement) while reducing cost from 1.52 to 1.34 PFLOPs per question. The trained summary harness moves from 34.7% to 42.1% and from 4.01 to 2.19 PFLOPs per question; the paper reports CLM outperforming the equivalently trained summary harness by 0.4 points with 38.8% fewer FLOPs.
- Suffix Cache Reuse: SCR matches standard SGLang serving while using 65.0% of its empirical prefix-reuse FLOPs on BCP, corresponding to a 35% reduction in server-side compute at matched performance.
- Emergent behaviors: CLMs invented behaviors not specified in the harness, including maintaining an in-context scoreboard and tracker with 163 in-place edits while holding the context at 6–8K tokens, introducing a new "notes" role in the context template, using loops to strip irrelevant search results, and defining a reusable
compact_turnsfunction invoked 37 times, and compressing 21K tokens into answer-relevant summaries.
Methodology in Plain English
The core move is to stop treating context as something an external program manages and instead hand the model a writable file that is its context. The model sees the file's path in its system prompt and can edit it with ordinary command-line tools; whatever it writes becomes what it sees on the next turn. If it writes nothing, behavior falls back to the usual appending. Because the context is just a file, several of them can coexist, which is how the paper gets agent swarms (multiple persistent contexts) and subagents (created and deleted on demand).
To measure cost fairly, the authors define prefix-reuse FLOPs, which counts the prefill of tokens from the first prefix mismatch onward plus decoding of newly generated tokens — capturing the re-prefill penalty that in-the-middle edits cause under standard prefix caching.
For the empirical work, they benchmark zero-shot CLMs against harness-defined and action-based baselines under a shared Mini-SWE-Agent backbone with Qwen3.6-27B at a 32K context limit and a 100-turn cap, and additionally use GPT-5.6-Sol, Claude 4.6 Sonnet, and Opus 5 in specific experiments. For learning, they (a) add one-sentence steering instructions, (b) run a prompt-evolution loop that proposes candidate skills from rollout traces and selects them on a development split, and (c) apply stepwise GRPO with a success-gated efficiency advantage: outcome advantages are computed per trajectory, and an extra term rewards successful trajectories with lower prefix-reuse FLOPs (zero unless at least two trajectories in the group succeeded). Finally, they implement Suffix Cache Reuse inside SGLang so that after an edit, the cached states of unchanged suffix tokens are reused rather than re-prefilled.
Why This Matters
- Impact on research: The paper reframes context management from an engineering artifact (harness code, a fixed tool menu) into a learnable model capability, giving a formal generalization of the append-only context transition and a diagnostic benchmark (ContextBench) for isolating it. It also connects agent behavior to serving systems, showing that the two cannot be evaluated independently.
- Real-world applications:
- Long-running coding agents that must stay coherent across hundreds of turns and repository-wide changes.
- Deep-research agents that search, retrieve, and must keep only decision-relevant evidence in context.
- Open discovery workflows in scientific computing and mathematical optimization, where agents iterate for hours or a full day.
- Multi-agent software optimization where several agents work on interdependent repositories with held-out downstream evaluation.
- Industry relevance: The results target inference cost directly — fewer prefix-reuse FLOPs at equal or better accuracy, plus a 35% server-side compute reduction from Suffix Cache Reuse against standard SGLang. That matters for anyone serving long-horizon agents, and the "harness-to-CLM" framing suggests existing agent products could migrate their hand-built context logic into model behavior.
Future Directions
- Scaling CLM reinforcement learning: The paper calls for larger-scale RL so models can explore and learn context-management strategies more thoroughly; the current RL result is on Qwen3.5-9B and a single deep-research training set (OpenResearcher).
- Distilling harnesses into CLMs: Because many harness operations are context transformations, the authors propose a harness-to-CLM pipeline that translates them into CLM actions and eventually internalizes them into weights, treating harnesses as procedural memory.
- Safety of model-editable context: Editable context is a new channel for prompt injections or self-generated instructions that persist across turns; the paper notes prior work observing a model inserting unauthorized instructions into its own compaction summary, and asks for defenses that keep the flexibility while preserving integrity.
- Further serving savings: The appendix describes SCR for hybrid models with interleaved full- and linear-attention layers, handling of multiple surviving post-edit spans, and a decomposition of savings from reasoning-token stripping, while noting remaining opportunities for improvement in SGLang serving with SCR.
Target Audience
Researchers and engineers working on LLM agents, context and memory management, agentic reinforcement learning, and inference serving systems. Practitioners who build long-horizon coding, research, or multi-agent pipelines will find the benchmark comparisons and compute accounting directly actionable; those new to the area will find the ContextBench pilot and qualitative examples the most accessible entry point.
Authors’ abstract
We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.