Research
QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
Overview Research area: Machine learning systems and reinforcement learning (RL) infrastructure for large language model (LLM) agents, specifically the training of agents on "xLong-horizon" software e

- arXiv
- 2609.33848
- Published
- 2026-09-27
- Authors
- Weiqi Wang, Yuxin Zhou, Mouxiang Chen, Siyuan Zhang, Yi Zhang, Yuyan Luo, Zhiyu Yin, Chencan Wu, Jiemin Jiang, Wentao Yao, Chujie Zheng, JianWei Zhang
AI summary
Overview
Research area: Machine learning systems and reinforcement learning (RL) infrastructure for large language model (LLM) agents, specifically the training of agents on "xLong-horizon" software engineering tasks.
Technical level: Advanced. The paper assumes familiarity with online RL for LLMs, group-relative advantage estimation, GPU parallelism layouts, and distributed training/serving systems.
Scope: The paper presents QwenGyre, an end-to-end framework for online RL over black-box agent harnesses at xLong horizons, combining an elastic GPU scheduler with a trajectory processor, and evaluates it on three agentic benchmarks using Qwen 3.6 122B and Qwen 3.8 2.4T.
What This Paper Is About
LLM agents are increasingly asked to perform tasks that run for hours, involve hundreds of model–environment interactions, and process around 1M tokens per rollout. Applying online RL to these executions creates two problems: rollout durations vary so much that GPUs sit idle waiting for a few slow stragglers, and the way agent harnesses compact history, delegate to sub-agents, and retry failed paths turns a single execution into a branching graph whose expansion into training samples produces massive redundancy. QwenGyre aims to solve both — by moving GPUs between rollout and training as demand shifts without killing live executions, and by converting branching executions into a bounded, deduplicated set of training trajectories.
Key Contributions
-
Identification of the core obstacles. The paper frames the scheduling and trajectory-processing challenges of online RL for xLong agent executions, showing why standard Colocate and Async GPU placement policies underutilize hardware at these horizons and why black-box harnesses break the append-only trajectory assumption.
-
An elastic scheduler. QwenGyre reallocates GPUs between rollout and training as demand changes, lets newly available GPUs join an ongoing training batch via centralized dynamic data parallelism (satellites pull the core cell's weights over RDMA and start training mid-batch), and balances training work across cells using streamed micro-steps so participating cells finish a batch at similar times.
-
A trajectory processor. It preserves each model output's original conditioning context using token-in, token-out (TITO) recording, organizes calls into prefix-sharing trajectory trees, evaluates partial work after timeouts, selects at most
J_maxtrajectories per execution by role priority, counts shared outputs once, and averages token losses within each execution. -
Empirical demonstration on a flagship model. QwenGyre improved NL2RepoBench pass rate for Qwen 3.8 2.4T from 52.5% to 58.5% in 48 training steps, with end-to-end speedups of up to 1.85x over Colocate and 1.78x over Async across evaluated workloads.
Main Findings
-
Flagship-model training gain: On NL2RepoBench, QwenGyre raised Qwen 3.8 2.4T's score by 6.0% absolute (52.5% → 58.5%) over 48 training steps, using 700K tokens per rollout.
-
End-to-end speedups: The abstract reports up to 1.85x over Colocate and 1.78x over Async. The introduction reports a range of 1.38x to 1.78x over Async and 1.21x to 1.85x over Colocate across reported configurations, while matching baseline training scores at equal GPU budgets and the same staleness constraints.
-
xLong execution profile is genuinely extreme: With Qwen 3.6 122B on NL2RepoBench at
E[d] = 1.5, mean durations were 1.93 hours per rollout execution and 2.96 hours per query, and 9.51% of queries took at least four hours. For Qwen 3.8 2.4T, the duration histogram's final bin collects runs at or above 4.45 hours. -
Straggler and termination clustering: The paper reports harness and overall timeout rates beginning at 6.25% (the sentence is truncated in the available content, so the full breakdown is not reported here), and argues that clustered terminations sharply reduce inference demand while a few long tails keep GPUs idle.
-
Non-linear trajectories are the norm: A single execution contains main-agent and sub-agent trajectories, with compaction producing summary branches; the paper's Figure 5 characterization of leaves per rollout tree and tool calls per rollout execution supports the claim that simple linear histories do not survive at these horizons.
-
Role-priority selection matters: Candidate trajectories are ranked main > main-summary > sub-agent > sub-agent-summary, on the reasoning that main-agent turns are what task-level rewards most directly credit, while summary trajectories manage context rather than act on the task.
-
Partial credit is usable: Executions that time out can still receive a valid score when their preserved work is assessable. Unresolved assessments are excluded from training and are treated as distinct from a valid zero reward; rare complete failures with no assessable work receive a dummy trajectory with zero reward.
-
Elastic roles do shift in practice: The paper's execution profile for QwenGyre on NL2RepoBench with Qwen 3.6 122B at
E[d] = 1.5shows cell roles changing over the first three bursts (steps 1–9), with Cell 7 remaining in rollout throughout the observed interval. -
Not reported in the available content: Numerical training results for DeepSWE and TerminalBench, the complete timeout breakdown, and the full results tables are truncated; only their setup (24 nodes, six cells, inference concurrency limit 192 per cell) is described.
Methodology in Plain English
The researchers built a training system that sits outside the agent harness rather than inside it. The harness — Claude Code 2.1.220 in these experiments — is treated as a black box that talks to the model through a proxy; the framework records every model call and receives the task-level evaluation result without modifying the harness's control flow.
To keep GPUs busy, they split hardware into "elastic cells" that share one training parallel layout. At any moment a cell can either serve rollouts or run a full training replica. A single "core" cell holds the authoritative weights and is the only one that owns optimizer state; "satellite" cells pull the core's weight snapshot over RDMA and can join a batch that is already in progress. The scheduler tracks three counters — executions dispatched, launched, and finished — and uses the difference between dispatched and finished (the "waterlevel") to decide when another cell can safely leave rollout. A cell only switches if the remaining rollout capacity still covers outstanding executions, and only after a supply check confirms enough ready groups exist. When role changes happen, the proxy reroutes in-flight model requests to other engines and can migrate KV-cache blocks so prefixes do not need recomputation.
To tame trajectory explosion, the proxy records exact input and output token IDs plus behavior log-probabilities for every call, and arranges them into a tree whose nodes share common prefixes. Diverging contexts become branches; only new content is encoded. At admission time, the system masks any token the policy did not generate, then repeatedly draws the highest-priority leaf class that still has unmasked trainable tokens, with probability proportional to its unmasked count. Each draw fixes that trajectory's targets and masks its whole root-to-leaf path, so a shared prefix is never trained twice. Losses are averaged over trainable tokens within an execution and then over executions in the batch, so adding trajectories does not inflate an execution's weight.
Training then proceeds in streamed micro-steps: a packing buffer pulls the oldest ready group, materializes its selected trajectories, and cells atomically claim micro-steps as compute frees up, using a continuous one-forward-one-backward pipeline.
Why This Matters
Impact on research: The paper argues that xLong-horizon RL is not just a scaled-up version of existing agentic RL — the scheduling assumptions of Colocate and Async designs, and the append-only trajectory assumption of simple agent loops, both break down. Framing the harness as a black box and the recording layer as the training interface offers a template for training against closed, rapidly evolving deployment harnesses rather than reimplementing them inside the RL framework.
Real-world applications:
- Automated repository construction from natural-language specifications (the NL2RepoBench setting).
- Software engineering inside existing repositories, such as issue resolution and repository-level patching (the DeepSWE setting).
- Multi-turn terminal and shell-based task automation (the TerminalBench setting).
- Codebase-wide refactoring and system reimplementation, which the paper cites as motivating xLong workloads.
Industry relevance: The paper comes from Alibaba Token Hub at Alibaba Group, with co-authors from the University of Science and Technology of China and Tsinghua University. The results are reported on models at production scale (Qwen 3.6 122B and Qwen 3.8 2.4T) over large node budgets (up to four cells of 48 nodes). Speedups of this size directly translate into training throughput and cost for organizations running RL on long-horizon agents, and the elastic design addresses a failure mode — GPUs idle while a handful of executions finish — that any team running such workloads would recognize.
Future Directions
-
Generalizing beyond the Claude Code harness: The role-priority ordering is specified concretely for a Claude Code harness (main, main-summary, sub-agent, sub-agent-summary). How the ordering should be defined for other harnesses with different delegation or compaction patterns is left open.
-
Tuning the elastic control parameters: The scheduler's behavior depends on
rho_min(the fraction of the batch that must be continuously ready before a transition), the number of cellsK, and cell granularity. The paper includes a cell-granularity ablation comparing eight 4-node cells against four 8-node cells under the same 32-node budget, but the general trade-off between cell size and reallocation responsiveness is not fully resolved. -
Handling unresolved assessments at scale: Executions whose evaluation budget expires without a valid reward are excluded from training, and complete failures receive a zero-reward dummy trajectory. How often this happens and its effect on group-relative advantage estimation for the
Nrollouts in a group is an obvious next question. -
Extending evidence to the remaining benchmarks: The available content reports end-to-end speedup and score figures for NL2RepoBench; DeepSWE and TerminalBench are described in the setup but their numerical outcomes are truncated, leaving their results to be confirmed in the full paper.
Target Audience
This paper is most useful to machine learning systems engineers and RL infrastructure researchers who build or operate training pipelines for LLM agents; to applied scientists working on agentic RL who need to understand why long-horizon rollouts break conventional scheduling and trajectory assumptions; and to engineering leads deciding how to allocate GPU resources for agent training at production scale. Readers without a background in distributed training, RL advantage estimation, or agent harness design will find the systems sections demanding, though the problem framing and the headline results are accessible without it.
Authors’ abstract
Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental challenges: (1) severe execution variance and prolonged rollout delays cause massive GPU idling; and (2) complex non-linear branching generates massive trajectory redundancy, crippling training efficiency. To address these, we presents QwenGyre, an end-to-end framework for xlong-horizon online RL. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs. Scaled to our flagship model, Qwen~3.8 2.4T, with 700K tokens per rollout, QwenGyre yields a 6.0% absolute gain on NL2RepoBench (52.5% $\to$ 58.5%) in 48 steps. Across our evaluations on diverse domains of training datasets, QwenGyre delivers up to $1.85\times$ and $1.78\times$ speedups over Colocate and Async, respectively.