Research
Environment Evolution for Terminal Agents
Overview Research area: Reinforcement learning for LLM-based terminal agents, specifically the automatic construction and scaling of executable, verifiable training environments. Technical level: Adva

- arXiv
- 2609.04128
- Published
- 2026-09-03
- Authors
- Zhiyuan Fan, Tinghao Yu, Yuanjun Cai, Jiang Zhou, Jiangtao Guan, Jincheng Liu, Yun Yang, Dingxin Hu, Zhuo Han, Xing Wu, Feng Zhang, Lilin Wang
AI summary
Overview
- Research area: Reinforcement learning for LLM-based terminal agents, specifically the automatic construction and scaling of executable, verifiable training environments.
- Technical level: Advanced. The paper derives an off-policy difficulty formulation from a multi-turn learning objective and assumes familiarity with RL for agents, harnesses, schedulers, and MoE training.
- Scope: The paper proposes "environment evolution," a paradigm that incrementally raises the difficulty of terminal environments generation by generation without target-agent rollouts, and validates it through 200-step long-horizon RL training on two Qwen3.6 models.
What This Paper Is About
Environments synthesized from scratch are increasingly easy for frontier models, so they yield weak learning signals when used for RL training. Existing co-evolution methods synthesize harder environments from a model's own on-policy failures, which ties the resulting environments to that model and stops producing useful signals as the model improves. This paper instead evolves an existing environment into a lineage of progressively harder environments off-policy, and schedules those generations during training so the agent keeps receiving solvable-but-challenging tasks.
Key Contributions
- A model-agnostic formulation of environment difficulty derived from the multi-turn learning objective. By replacing model-dependent probabilities with a reference distribution grounded in broad world knowledge, the authors convert difficulty from a model-specific weakness measure into environment-level difficulty, and show that on-policy co-evolution mainly targets skill-selection errors rather than the full difficulty space.
- The environment evolution paradigm, implemented as a loop-engineered multi-agent harness that incrementally modifies an existing environment along one of three derived directions (scenario novelty, skill rarity, execution length) to build verified lineages of increasing difficulty.
- The Evolution-Lineage (EL) Scheduler, which admits evolved environments generation by generation based on a pass-rate threshold so that rollout groups remain partially solved and retain non-zero within-group advantage.
- Validation through 200-step long-horizon RL experiments on Qwen3.6-27B and Qwen3.6-35B-A3B, reporting gains of 14.4 and 18.0 percentage points on Terminal-Bench 2.1 and higher peak accuracy than co-evolution and ensemble baselines.
Main Findings
- Evolved environments are harder across models: Rollout-based difficulty estimates with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol show that environment evolution consistently produces more difficult environments despite being synthesized off-policy.
- Evolution effort controls mutation scope: With a 15-generation cap,
loweffort (restricted to one scenario–skill pair) shows avg turns increasing across generations but pass rate fluctuating and remaining above zero;highandmaxmonotonically reduce pass rate to zero and increase avg turns, withmaxreaching zero earlier and producing the larger difficulty change.highis used as the default. - Mutation composition differs by direction: Instruction mutation stays high under
highandmax, fluctuating between 95% and 100% while trending downward, while environment and verification system changes remain more selective. - Each evolution direction has a stable profile: In the 1-step setting,
lengthproduces the largest pass-rate decrease (−7.1 pp) with the largest avg-turn increase forscenario(+13.5);scenarioyields the largest total mutation (71.1%), andlengthachieves the strongest pass-rate reduction with the smallest total mutation (63.8%). The 15-step means preserve these profiles (length−4.8 pp pass rate, +7.4 turns;scenario−2.9 pp, +9.5 turns;skill−2.7 pp, +10.6 turns). - Scheduling beats random sampling early in training: Over the first 50 RL steps, the EL Scheduler keeps more rollout groups partially solved (one to seven successes among eight rollouts) under matched rollout budgets, providing GRPO with more informative signals.
- Trajectories lengthen during training: Tokens per turn increase from approximately 951 to 1,221 for Qwen3.6-27B and from 947 to 1,103 for Qwen3.6-35B-A3B over 200 steps.
- Peak accuracy comparison: Environment evolution reaches peak accuracies of 71.5% and 64.9% on Qwen3.6-27B and Qwen3.6-35B-A3B, compared with 62.9% and 55.1% for co-evolution and 60.0% and 52.8% for Ensemble, with Claude Opus 5 fixed as the environment synthesis model and the total number of training environments held constant.
- Seed environments were heavily filtered: Of 47,678 non-benchmark terminal environments collected from Hugging Face and GitHub, 127 were retained through strict rubric-based filtering, then supplemented using SkillSynth to yield a balanced seed pool of 500 environments.
Methodology in Plain English
The authors start by asking what "difficult" means for an environment rather than for a particular model. They model agent execution as a sequence of interleaved scenarios and skills, then define difficulty as the negative log-likelihood of that sequence under a reference distribution representing broad world knowledge. This yields three levers: how many solver steps a task requires (length), how novel a scenario is (scenario), and how rare the required skill is under that scenario (skill). The gap between a specific model's difficulty and the environment-family difficulty is defined as that model's weakness, which the authors use to argue that co-evolution only covers part of the difficulty space.
Environment evolution then operates on an existing environment by editing its expected execution sequence along one of the three directions. A Proposer extracts the current scenario–skill sequence, edits it according to the chosen direction, and generates a plan; a rubric-based reviewer iterates on the plan until it is accepted. A Modifier turns the plan into a concrete change to the environment, and three parallel verifiers check it: an Oracle verifier confirms the reference solution succeeds, an Invalid-test verifier confirms an empty or no-op solution fails, and an adaptive general-rubrics verifier checks environment quality. During development, accepted environments also go through human-in-the-loop review, and any issues that slipped past the rubric checks become new rubrics until the loop reliably produces issue-free environments.
A prompt-controlled "evolution effort" parameter governs how much of the sequence can be edited at once (one pair, one contiguous span, or an unrestricted portion). At each generation the three directions are randomly ordered, and if a plan is rejected or the repair budget runs out, the harness falls back to the next direction.
Finally, an Evolution-Lineage Scheduler walks through each lineage in order, advancing to the next environment once the current one's average reward over eight rollouts exceeds 6/8, and moving to the next generation only when the current one is exhausted. Training uses GRPO with partial rollouts, asynchronous GPUs, and a staleness bound of 5, on top of an RFT checkpoint, inside the Claude Code harness with a 256K context window that compacts when remaining usable context reaches 16K.
Why This Matters
- Research impact: The paper reframes environment scaling from a model-coupled, on-policy process to an off-policy process that can keep supplying learning signals after a model has saturated its seed environments. It also provides an analytic account of what on-policy co-evolution does and does not control.
- Continued scaling of agentic RL: The stated bottleneck is no longer the RL algorithm but the environment supply; this work targets that bottleneck directly.
- Applications:
- Training terminal and DevOps agents that must complete long multi-step command-line tasks.
- Producing long-horizon training data for software engineering agents and computer-use agents, which the authors name as future work.
- Curating and upgrading existing open-source environment collections, since the paper reports that such environments often suffer from low-quality reward signals such as misalignment between task instructions and verification systems, and corrupted environments.
- Generating progressively harder benchmark-like tasks from a small filtered seed pool, reducing dependence on manually authored difficulty ladders.
- Industry relevance: The work comes from the Hunyuan Team at Tencent and includes a pipeline that filters tens of thousands of public environments into a much smaller usable pool, which is directly relevant to organizations building agent training infrastructure at scale.
Future Directions
- Extending environment evolution to SWE agents and Computer-Use Agents, which the conclusion names explicitly.
- Reporting RSI (recursive self-improvement) results in which the same model both constructs and learns from the environments, deferred to future versions in the appendix.
- Determining whether evolution remains beneficial beyond the 15-generation cap used here; the authors note that pass rate provides no further resolution once a lineage enters the zero-pass regime.
- Broadening the evaluation beyond the two Qwen3.6 models and the single held-out benchmark, since all reported training results use Terminal-Bench 2.1 Verified.
Target Audience
Researchers and engineers working on agentic reinforcement learning, environment design, and post-training infrastructure for terminal or tool-using agents. The paper is also relevant to teams that maintain large collections of executable task environments and want a principled way to raise their difficulty, and to readers interested in open-ended learning and unsupervised environment design translated to LLM agents.
Authors’ abstract
Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models become more capable, environments synthesized from scratch become less challenging and thus provide limited learning signals. Recent co-evolution methods iteratively synthesize environments near the model's learnable frontier based on weaknesses exposed during rollouts. However, their dependence on on-policy rollouts limits generalization and the continuous provision of learning signals as the model becomes stronger. In this paper, we propose environment evolution, which incrementally increases environment difficulty off-policy and schedules the evolved environments generation by generation during training to provide continuous learning signals. We derive three evolution directions that influence environment difficulty from the multi-turn learning objective and then implement evolution along these directions through a loop-engineered multi-agent harness. Quantitative rollout experiments with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol show that environment evolution consistently produces more difficult environments. We validate its effectiveness on Qwen3.6-27B and Qwen3.6-35B-A3B through simple long-horizon RL training, improving their performance by 14.4 and 18.0 percentage points on Terminal-Bench 2.1, respectively.