Research
Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR
Overview Research area: Reinforcement Learning with Verifiable Rewards (RLVR) for large language model reasoning, specifically cross-model trajectory sharing and off-policy correction. Technical level

- arXiv
- 2609.37868
- Published
- 2026-09-29
- Authors
- Doohyuk Jang, Yoonsik Park, Gyouk Chu, Sihwan Park, Eunho Yang
AI summary
Overview
Research area: Reinforcement Learning with Verifiable Rewards (RLVR) for large language model reasoning, specifically cross-model trajectory sharing and off-policy correction.
Technical level: Intermediate — readers should be comfortable with policy-gradient RL, GRPO-style group-relative advantages, importance ratios, and PPO-style clipping, though the paper explains each of these.
Scope: The paper introduces GRAFT, a framework in which two heterogeneous LLMs exchange rollout groups during RLVR training so that each can learn from prompts its peer solves but it fails entirely, and reports results across three model pairs and five mathematical reasoning benchmarks.
What This Paper Is About
RLVR methods such as GRPO learn only from the model's own sampled responses, so when every response in a group fails the group-relative advantage is zero and the prompt yields no gradient signal. The paper observes that different models frequently succeed on complementary prompts — successes one model never samples may already exist in another model's rollouts — and builds a framework that transfers those peer trajectories while explicitly controlling the mismatch between models.
Key Contributions
- Complementary prompt selection. GRAFT defines the candidate set as prompts where the receiver's whole rollout group fails (k_B(q)=0) and the peer group contains at least one success but is not entirely correct (1 ≤ k_A(q) < n), ensuring both a verified success and nonzero reward variance in the transferred group.
- Balanced bidirectional exchange. Because candidate sets can differ in size between directions, GRAFT uses m = min(|C_A→B|, |C_B→A|) as a common selection target, ranking candidates by the source model's success count and retaining ties at the boundary.
- Off-policy-aware peer updates. Transferred groups keep their source-computed advantages; each peer response is weighted by a bounded sequence-level compatibility score (a gate at threshold δ plus a cap at 1), and updates use token-level importance ratios computed on the receiver's own tokenization with asymmetric clipping. Peer-containing minibatches are processed after the receiver's on-policy minibatches.
- Empirical validation across pairs, budgets, and storage modes. GRAFT is tested against GRPO at n=8, 16, and 32, against HACPO and SGT, and in a variant that reuses stored peer trajectories instead of co-training.
Main Findings
- Complementarity exists in practice: In independent GRPO runs, SmolLM3-3B-Base solves 47.9% of the prompts on which Qwen3-1.7B-Base fails across all eight rollouts, and Qwen3-1.7B-Base solves 18.7% of SmolLM3-3B-Base's all-fail prompts.
- Consistent gains over GRPO: Across three heterogeneous model pairs and five mathematical benchmarks (MATH500, AIME2024, AIME2025, AMC23, Minerva), GRAFT improves both models in every pair over GRPO at n=8, gaining 2.1 points on average and up to 4.5 points in model-level average performance. Per-block average-score gains range from +0.66 to +4.46.
- Competitive with much larger rollout budgets: On Pair 1, both models outperform GRPO with 4× the rollouts (n=32); on Pair 2 both are comparable to it. In two of the three pairs, both models match or exceed GRPO trained with n=32.
- Stronger than cross-model baselines: GRAFT outperforms HACPO and SGT by 4.0 and 1.5 points on average, respectively. HACPO falls below budget-matched GRPO in five of six model blocks, by as much as −4.17; SGT is inconsistent, ranging from −0.53 to +2.28.
- Better compute trade-off: GRAFT reaches a pair-mean score of 35.25 at 40.9 GPU-hours, exceeding GRPO (n=32) by 1.18 points at 0.45× its compute and GRPO (n=16) by 2.48 points at comparable compute.
- Gains survive without co-training: Reusing stored peer trajectories improves over GRPO (n=8) in all six model blocks by 1.78 points on average, versus 2.11 for online exchange — 84% of the online gain. Stored trajectories cut compute by 27–76%; on Pair 1, obtaining both selected receiver checkpoints requires 23.2 GPU-hours excluding the prior GRPO runs that collected the peer logs.
- Ablations confirm each component: Removing compatibility weighting costs SmolLM3 −8.36 points, the largest degradation in the component ablation; removing the floor costs Qwen3 −3.56. Removing balancing (−2.00/−0.48), replacing the token-level ratio with a sequence-level ratio (−1.85/−0.45), random prompt selection (−2.40/−1.58), random admission (−1.86/−1.09), and peer-first ordering (−2.45/−1.71) all degrade performance. Pooling peer and self groups (−3.99/−2.66) and transferring successes only (−7.84/−3.93) also underperform full GRAFT.
- Selection and update rule must be paired: With GRAFT's prompt selection fixed, HACPO, SGT, and LUFFY-style updates all trail full GRAFT; using GRAFT's update with SGT's unbalanced selection leaves a gap of 2.00/0.48 points.
- Gain magnitude tracks complementarity: Improvements are largest on Pair 1 and smallest on Pair 3.
Methodology in Plain English
Two models are trained at the same time on the same prompts and the same verifiable binary reward, but with separate parameters and gradients; they exchange only sampled response strings, generation log-probabilities, and rewards. For each prompt, each model samples n=8 responses. When one model's group is entirely wrong and the other's group contains a mix of right and wrong answers, GRAFT hands the whole peer group to the failing model, keeping the peer's own advantages so that the right/wrong contrast is preserved rather than pooling rewards across models.
Because the two models may use different tokenizers, the peer's text is re-tokenized with the receiver's tokenizer. Before a peer response is admitted, a sequence-level compatibility score compares average token log-likelihoods under the receiver and under the generating peer; responses scoring at or below a threshold δ=0.8 are dropped, and admitted responses are weighted by min(s, 1). The update then uses token-level importance ratios whose denominator is always the receiver's own behaviour policy, so the ratio tracks only the receiver's change while the sequence-level weight carries the cross-model mismatch. Peer-containing minibatches are ordered last, so that clipping is already active by the time they are processed. Training uses verl with Ray, FSDP, and vLLM rollout on the 7,500 problems in the MATH training split, with learning rate, training steps, and data held fixed across methods. Evaluation reports pass@1 on the five benchmarks, each checkpoint evaluated over five runs with 8 samples per prompt, reporting mean and standard deviation.
Why This Matters
Impact on research: The work reframes the all-fail group, a known failure mode of GRPO, as an opportunity for cross-model exchange rather than something to be discarded, resampled at extra cost, or reshaped without new information. It also separates within-receiver policy change from cross-model mismatch analytically and shows that treating them with different mechanisms (token-level clipping and sequence-level gating) outperforms applying either prior cross-model recipe to the same selected prompts.
Real-world applications:
- Training several small open-weight models on limited rollout budgets and letting them improve each other instead of paying for more rollouts each.
- Reusing trajectory logs already collected from earlier training runs as a supervision source for a new model, without keeping the donor model resident in memory.
- Cost-constrained post-training of reasoning models where GPU-hours are the binding constraint.
- Building mixed-model training pipelines where models of different sizes, tokenizers, and pretraining corpora are co-trained rather than distilled into a single designated student.
Industry relevance: The stored-trajectory setting is directly relevant to practitioners who already have logs from completed RLVR runs and want additional returns from them; the measured 27–76% compute reduction and 84% retention of the online gain quantify that value. The reliance only on responses, log-probabilities, and rewards — no shared tokenizer, no architecture change, no stronger teacher — makes the method applicable to heterogeneous model fleets.
Future Directions
- Scaling beyond pairs: The paper studies only two-model pairs; exchange among more than two peers is left to future work.
- Beyond verifiable rewards: The method depends on binary verifiable rewards, and the authors leave domains without such rewards to future work.
- Better cross-tokenizer mismatch measures: The compatibility score is explicitly described as an empirical proxy from average token log-likelihoods, not an exact cross-tokenizer importance ratio.
- Dependence on complementarity and model scale: Gains are smallest on Pair 3, and the study covers only base models of at most 3B parameters, so how the approach behaves when models are more similar or larger remains open.
Target Audience
Researchers and engineers working on RLVR and post-training of reasoning LLMs, particularly those interested in off-policy correction, multi-model or collaborative training, and rollout-budget efficiency. It is also useful to practitioners who run GRPO-style pipelines at scale and encounter all-fail groups, and who want to extract more value from existing rollout logs. Readers without a background in policy-gradient methods will need to consult the preliminaries section, but the paper's design questions — which trajectories to transfer and how to control their influence — are stated in accessible terms.
Authors’ abstract
Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, successful trajectories missing from one model's rollouts may already have been discovered by another. Indeed, we observe that heterogeneous models often succeed on complementary prompts, creating opportunities for mutual learning without a designated stronger teacher. To exploit this complementarity, we propose GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories), an off-policy-aware framework that replaces all-fail groups with informative peer groups. GRAFT transfers both successful and unsuccessful peer responses with peer-computed advantages, while controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping. Across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT consistently improves both models over GRPO with the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points in model-level average performance. Stored peer trajectories preserve most of the gains, improving over GRPO by 1.8 points on average without simultaneous co-training.