Research
ROSS: Relearning from Self-Generated Rollouts through Selective Supervision
ROSS: Relearning from Self-Generated Rollouts through Selective Supervision Overview Research area: Large language model post-training, specifically reinforcement learning (RL), on-policy distillation

- arXiv
- 2609.35954
- Published
- 2026-09-28
- Authors
- Zhiwei Zhang, Huayu Deng, Fei Zhao, Jiayan Fu, Bin Liang, Kam-Fai Wong, Mu Chuan
AI summary
ROSS: Relearning from Self-Generated Rollouts through Selective SupervisionOverview
- Research area: Large language model post-training, specifically reinforcement learning (RL), on-policy distillation (OPD), and agentic RL, and how to reuse the rollout data these procedures produce.
- Technical level: Advanced. The paper assumes familiarity with RL post-training, supervised fine-tuning (SFT), loss masking, and the distinction between on-policy and off-policy data.
- Scope: A single paper proposing ROSS, a method that selects which saved historical rollouts to reuse and which token spans inside them to supervise, validated on Qwen3.6-35B-A3B across mathematics, code generation, instruction following, and software engineering.
What This Paper Is About
Post-training methods such as RL and on-policy distillation continuously generate rollouts from the model itself, but once the policy advances, those saved rollouts are normally treated as stale and thrown away. The authors argue that these historical rollouts can still contain behaviors a later checkpoint is compatible with but no longer expresses reliably, while also containing mistakes, abandoned attempts, and redundant actions that should not be imitated. ROSS addresses both levels of selection: it keeps a historical trajectory's full text as context but applies training loss only to the specific model-generated segments judged worth imitating.
Key Contributions
- Historical self-generated rollouts as a reusable training resource. The paper establishes that rollouts from earlier checkpoints can retain useful behaviors under an evolving policy and can be relearned from later, in a separate offline SFT stage.
- The ROSS method. ROSS applies two-level selection: a trajectory-level outcome verifier decides which historical rollouts to keep, and an LLM-based annotation procedure proposes and audits which model-generated token spans inside those rollouts receive imitation loss, while the complete original prefix is preserved as context.
- Validation across three post-training settings. ROSS is tested as an additional offline SFT stage after domain-specific RL, multi-teacher on-policy distillation (MOPD), and agentic RL, with results reported on mathematics, code generation, instruction following, and long-horizon agentic tasks.
- Analysis of what selective supervision removes. The paper classifies the masked content across domains and measures the target exclusion rate (TER) over the course of RL training, showing that excluded supervision persists as the policy evolves.
Main Findings
- Compatibility of historical rollouts. On 100 mathematical problems, with one response per problem from historical rollouts, current-policy rollouts, and the external teacher GLM-5.2, historical rollouts had substantially lower token-level negative log-likelihood (NLL) under the current policy than external-teacher responses, and remained close to current-policy rollouts.
- Complementarity (under-consolidated success). Across the 100 most difficult training problems with 32 current-policy rollouts each, 98 problems had at least one successful historical rollout. The current policy solved 97 of these, with average pass@32 of 99.0% but average pass@1 of only 37.7%, and 74 of 98 problems had a per-problem success rate of at most 50%.
- Segment-level complementarity. The paper gives an example where a historical rollout reaches the correct count 1007 by recovering from an intermediate 2^10 = 1024 counting error, while a later rollout retains the uncorrected count of 1024 — showing that useful and undesirable reasoning can coexist in one trajectory.
- Domain-specific RL gains. After Math RL, the Math average improved from 75.19 to 76.98; after Code RL, the Code average improved from 42.09 to 44.21. ROSS was the best relearning method in both settings.
- MOPD gains. The six-benchmark MOPD average improved from 58.40% to 62.20%, outperforming Continued MOPD (59.16) and Positive-Rollout SFT (57.70).
- Agentic gains. On SWE-bench Verified, ROSS improved the resolved-issue rate from 64.20 to 68.40, a gain of 4.20 points over the Upstream RL checkpoint, versus 65.20 (+1.00) for Positive-Rollout SFT.
- Cross-domain preservation. ROSS preserved domain-specific gains while improving several off-domain capabilities: after Math RL it raised Avg. Code from 41.21 to 44.24 and IFBench from 33.30 to 34.60; after Code RL it improved Avg. Math from 74.19 to 75.03. It also improved the overall BFCL multi-turn agentic tool-use score from 44.12 to 45.62 and from 46.25 to 49.38.
- Trajectory filtering already helps. ROSS w/o mask (same retained examples, all eligible tokens supervised) improved over Upstream by 0.19 points on LCB Gen and 1.72 points on OJBench, and raised the MOPD average from 57.70 to 58.49, reversing degradation seen with Positive-Rollout SFT on Code.
- Masking accounts for the remaining gains. With the replay set fixed, ROSS improved over ROSS w/o mask by 3.71 points on MOPD, 1.91 on LCB Gen, 0.43 on OJBench, and 0.66 on Math.
- Relearning transfers across initializations. Using identical trajectories, masks, and SFT configurations, ROSS initialized from Base performed comparably to ROSS initialized from the Upstream checkpoint across all evaluated settings, despite inheriting none of the original RL/MOPD parameter updates.
- What masking removes differs by domain. From 1,185 sampled trajectories classified with GLM-5.2, Math masked excerpts were dominated by mistakes followed by a visible correction (72.7% early, 68.3% late), while Code was dominated by redundant exploration (67.9% early, rising to 76.6% late).
- Excluded supervision persists. Over the first 100 rollout steps, TER among retained verifier-positive responses fell only modestly from 11.60% to 9.27% in Math, and rose from 22.65% to 28.07% in Code, even as policy entropy decreased.
- Undesirable modes can be suppressed. In a separate 4B Math RL run where the Upstream checkpoint developed a repetitive-answer mode, ROSS suppressed this behavior while improving Avg. Math from 52.05% to 63.65%.
Methodology in Plain English
The starting point is an already-finished training run — math RL, code RL, MOPD, or agentic RL — together with the rollouts that run saved along the way. ROSS treats the final checkpoint as its starting model and adds an offline SFT stage on that saved history.
Two selections happen in sequence. First, a trajectory-level selector keeps only rollouts whose outcome a verifier marks as successful, where the verifier can be an exact-answer check, an executable checker, a unit-test suite, or an environment success signal. Second, for each retained trajectory, an LLM reviewer (GLM-5.2 in high-thinking mode) examines the task, the full trajectory, and the verifier evidence, and returns an audit status plus ordered intervals of model output that are locally correct, self-contained, and behaviorally useful. Deterministic checks then confirm source consistency and token alignment, and the intervals are compiled into a token-level mask.
The key design detail is that masking a token removes it from the loss but not from the sequence. Every selected token is still predicted from the complete original prefix — including earlier mistakes, abandoned attempts, and environment feedback — so the model sees the exact state in which the continuation originally occurred. Training is masked teacher forcing on those selected tokens only. Because the method needs only saved checkpoints and saved rollouts, no new policy rollouts are generated for the relearning stage. The annotation model supplies selection decisions, not replacement target responses; the saved response text is preserved.
Comparisons isolate each component. Positive-Rollout SFT supervises all eligible tokens in all verifier-positive trajectories. ROSS w/o mask uses the same post-annotation subset as ROSS but supervises all eligible policy-generated tokens, so the difference between ROSS and ROSS w/o mask is purely the token-level mask.
Setup specifics: the model is Qwen3.6-35B-A3B, a mixture-of-experts model with approximately 3B active parameters. Single-turn experiments use a no-thinking configuration and report results after three SFT epochs; the agentic experiments retain thinking traces with an eight-epoch configuration and a longer context. The upstream single-turn pools are 3,000 Math prompts from DAPO-Math-17K, 12,000 Code prompts from CodeI/O, and 10,000 instruction-following prompts from the data released with Nemotron-Cascade 2. Retained examples were 27,848 (Math RL), 144,702 (Code RL), and 31,228 (MOPD), versus 35,700, 152,454, and 42,558 verifier-positive examples for Positive-Rollout SFT. For the agentic setting, training uses the OpenSWE set from daVinci-Env after removing tasks with exposed .git histories or other reward-hacking artifacts, yielding 4,048 training tasks, with a Codex CLI agent acting through Harbor on containerized repositories and SWE-bench Verified reserved for evaluation.
Why This Matters
- Impact on research: The paper reframes a byproduct of post-training — the growing record of self-generated experience — as a reusable asset rather than waste, and argues that training data selection should operate at the sub-trajectory level rather than treating a rollout as an indivisible unit. It also shows these gains do not require additional policy rollouts, since ROSS is an offline SFT stage on already-saved data.
- Real-world applications:
- Code-generation and software-engineering agents, where the paper reports gains on LCB Gen, OJBench, and SWE-bench Verified.
- Multi-turn tool-use and agentic assistants, where ROSS improved the BFCL overall score.
- Mathematical reasoning systems, evaluated on AIME 2025, AIME 2026, and HMMT-November 2025.
- Instruction-following assistants, measured by IFBench.
- Industry relevance: Practitioners who already run RL or distillation pipelines accumulate saved rollouts as a matter of course. ROSS suggests a way to extract further performance from that existing data without paying for new rollout generation, and it can be applied on top of different upstream procedures (domain-specific RL, MOPD, and agentic RL) with a fixed recipe.
Future Directions
- Why masking matters at scale. Since the target exclusion rate in Code rose from 22.65% to 28.07% rather than shrinking as RL progressed, it remains an open question whether masking needs grow, shrink, or shift in character over much longer training runs.
- Transfer across initializations. Base-initialized ROSS matched Upstream-initialized ROSS on the evaluated settings; how far this transferability extends, for example to entirely different base models or training recipes, is not established by the paper.
- Annotation cost and reliability. The within-trajectory selection depends on an LLM reviewer (GLM-5.2) plus deterministic validation; the paper does not report how sensitive the results are to the choice of reviewer or to reviewer errors that pass validation.
- Beyond the tested domains. Evaluation covers mathematics, code generation, instruction following, software engineering, and multi-turn tool use, so whether the same selective-supervision recipe generalizes to other long-horizon or multimodal agentic settings remains untested here.
Target Audience
Researchers and engineers working on LLM post-training — particularly those running RL, on-policy distillation, or agentic RL pipelines and looking for ways to reuse the rollouts those pipelines already produce. It is also relevant to readers interested in data selection, process-level supervision, and loss masking, and to practitioners building code or software-engineering agents who want an offline SFT stage that adds capability without new rollout generation.
Authors’ abstract
Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating finer-grained selective supervision. We introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), which preserves the full historical trajectory as context while applying loss only to selected model-generated continuations. Across domain-specific reinforcement learning, multi-teacher on-policy distillation, and agentic reinforcement learning, ROSS consistently improves upstream checkpoints and outperforms baselines across mathematics, code generation, instruction following, and software engineering. On Qwen3.6-35B-A3B, ROSS improves the six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%. These results show that self-rollout training leaves behind reusable behavioral experience that can yield further gains through offline supervised fine-tuning (SFT), without additional policy rollouts.