Research
Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding
Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding Overview Research area: Reinforcement learning for large language model reasoning, specifically reinforceme
- arXiv
- 2610.07342
- Published
- 2026-10-05
- Authors
- Hoang Phan, Minh Pham, Chau Pham, Chinmay Hegde, Trung Le, Qi Lei
AI summary
Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale ScaffoldingOverview
- Research area: Reinforcement learning for large language model reasoning, specifically reinforcement learning with verifiable rewards (RLVR), plus vision-language model post-training.
- Technical level: Advanced. The paper assumes familiarity with policy-gradient methods (PPO, GRPO, REINFORCE++), KL regularization, advantage normalization, and supervised fine-tuning objectives.
- Scope: The paper introduces Rationale-Guided Policy Optimization (RGPO), a training framework that uses ground-truth rationales as temporary, adaptively revealed scaffolds inside an otherwise standard RLVR loop, and evaluates it on synthetic problems, a 3B language model, and 2B/7B vision-language models.
What This Paper Is About
On-policy RL for reasoning stalls when a model is trained on problems it cannot yet solve: every sampled rollout is wrong, rewards become uniform, and the gradient carries almost no information. RGPO addresses this by reusing available reference rationales (solution hints) not as imitation targets, but as temporary help that lets the model produce a better answer, after which only the higher-reward, model-generated answer is trained on using the original unguided prompt. The amount of help is adjusted automatically per example based on whether the model currently succeeds.
Key Contributions
- Rationale-guided policy optimization (RGPO): A framework that augments the standard RLVR loop with a second, hint-assisted refinement pass. Reference rationales serve as temporary scaffolds rather than fixed imitation targets, and the paper emphasizes that this removes the requirement that off-policy data match the format of the RL task.
- Adaptive hint mechanism: The scheduler increases or decreases the amount of revealed solution guidance according to the model's current ability on each individual example, concentrating help on hard problems and annealing it away on problems the model can already solve.
- Reward gate over guided responses: A guided response is transferred into training only if its verifier reward exceeds the best unguided rollout for the same prompt by a margin, so the hint is treated as a proposal distribution rather than assumed-correct supervision.
- Cross-modal empirical validation and theory: Experiments on language models and vision-language models show training-sample-efficiency gains over standard RLVR and over hybrid on-policy/off-policy baselines, supported by a controlled simulation and a probabilistic analysis of when guidance helps under sparse rewards.
Main Findings
- Language model results: With Qwen2.5-3B trained on the DAPO-Math-27K split for 3 epochs, RGPO achieves the best result on every in-distribution benchmark, raising the in-distribution average from 19.7 with GRPO to 26.8 and outperforming the strongest prior baseline, HINT, by 5.0 points (26.8 vs. 21.8). Gains are largest on MATH-500 and Minerva Math, where RGPO improves over HINT by 6.6 and 8.7 points respectively.
- Out-of-distribution transfer: RGPO obtains the highest scores on ARC, GPQA, and MMLU, with an out-of-distribution average of 32.2 compared with 29.9 for the best baseline (a 2.3-point improvement). RGPO's individual out-of-distribution scores are ARC 51.2, GPQA-Diamond 14.7, and MMLU-Pro 30.8.
- Vision-language results: On MMStar-Math, RGPO reaches 70.1, outperforming the base Qwen2-VL-7B-Instruct model by 23.7 points. On MathVista-Math, RGPO achieves 73.65 versus the base model's 41.11, with the best subset scores in GEO (73.47), ALG (73.62), and GPS (73.89); the gain on the Textbook Question Answering subset is described as relatively modest (70.35).
- Improved data efficiency: On Qwen2-VL-2B with GRPO, RGPO reaches the baseline's performance in 2.95× fewer training steps and achieves higher final accuracy. The improvement is described as even larger against REINFORCE++, which the paper says struggles to learn effectively on that model.
- Controlled simulation: Unguided GRPO reaches 99.1% test accuracy on easy synthetic problems but remains at 0.0% on hard problems after three epochs, while RGPO reaches 99.5% hard test accuracy while preserving easy-problem performance.
- More guidance is not always better: Directly training on the imperfect reference plateaus at 12.8% hard test accuracy because its systematic errors are reinforced; removing RGPO's reward gate produces essentially the same behavior. Training with the full rationale available lets the model exploit the exposed final answer, but performance falls to the 1% chance level when the rationale is removed at evaluation.
- Few verified trajectories suffice: In the main simulation setting, only 7.4 guided responses are accepted per run on average, yet these are sufficient to bootstrap learning because the same hard decisions recur across problems.
- Always-on hints create shortcuts: Compared with a full-hint setting, the full-hint model rapidly inflates training reward while producing extremely short outputs, and generalizes poorly; its policy-gradient loss and gradient-norm curves show large early spikes and sharp fluctuations, whereas RGPO improves more gradually and achieves stronger no-hint validation performance.
- Format robustness: With the concise DeepSeek-R1-style prompt, RGPO reaches a higher reward plateau than GRPO while maintaining a comparable entropy trajectory. Under the structured Mulberry-style template, a hybrid method the paper calls AHPO shows a delayed reward increase followed by a sharp entropy collapse. The paper notes a 2B model "struggles to optimize under the Mulberry-style template."
- Theory: At an equal response budget, RGPO's probability of finding at least one successful trajectory exceeds the plain baseline's exactly when the guided success probability exceeds the unguided one. The hint level has negative drift once the success probability exceeds p* = 1 − 2^(−1/K), and the expected time to reach a target success level is O(log 1/p₀) for RGPO versus Ω(1/p₀) for the plain baseline.
Methodology in Plain English
RGPO extends a normal RLVR training run over N+1 epochs. The first epoch is fully unguided: the model samples responses and is trained with the usual RL objective (the paper uses GRPO and is also compatible with PPO or REINFORCE++). The reference rationale for each example is split into N ordered segments of roughly 1/N of the solution each, using a deterministic rule such as character count.
After the first epoch, each example is assigned a hint level based on whether any rollout was correct. If none were correct, the next epoch reveals one more segment of the reference solution; if at least one was correct, one segment is removed. This creates a per-example curriculum that drifts toward the smallest amount of help needed to keep the problem learnable — hard examples gradually receive more of the solution, easy examples are annealed back toward no help.
In each subsequent epoch, the model runs two passes: the ordinary unguided rollout from the original prompt, and a guided revision rollout from a reprompt that includes the original problem plus the revealed hint prefix and verifier feedback. The guided prompt is used only to generate candidate answers. A guided answer is accepted into an auxiliary buffer only if its verifier reward beats the best unguided rollout for that example by a margin δ (with δ = 0 in the simplest exact-reward setting). The auxiliary supervised loss maximizes the likelihood of accepted guided answers conditioned on the original prompt only — not on the hint-augmented prompt — so that the benefit is transferred back into the deployment condition and the model is not taught to depend on hints at inference. The total objective is the RL loss plus a weighted auxiliary loss, where the weight λ_t controls the auxiliary term.
Why This Matters
Impact on research: RGPO targets the central bottleneck of on-policy RLVR — reward sparsity caused by capacity–difficulty mismatch — without requiring off-policy data to be reformatted to match the RL task, and without needing rejection sampling from a stronger model. It offers a general recipe that can be layered on top of existing algorithms, and it links its design to cognitive-science findings about guidance timing (productive struggle, the "assistance dilemma," fading worked examples).
Real-world applications:
- Mathematics education and automated tutoring, where hints must be revealed selectively rather than given away.
- Multimodal problem solving such as geometry, chart, and diagram reasoning, demonstrated on MMStar-Math and MathVista-Math subsets.
- Graduate-level scientific question answering and domain QA, reflected in the GPQA-Diamond and MMLU-Pro evaluations.
- Post-training of smaller or lower-capacity models, where the paper reports the largest data-efficiency gains and where unguided RLVR most often stagnates.
Industry relevance: The framework is designed for practical post-training pipelines: it plugs into existing RLVR stacks, was implemented in the verl framework, and the authors report that RGPO's cross-modal capabilities allow training to exploit information already available as reference rationales without matching the RL task format. The reported gains are about training-sample efficiency, which matters directly for compute cost. Code is released at https://github.com/VietHoang1512/rgpo.
Future Directions
- Better hint segmentation: The experiments use raw character counts to split rationales into 1/N segments; the paper explicitly notes that partitioning could instead follow token counts, sentence boundaries, or reasoning steps, leaving open how much step-aware segmentation would help.
- Sensitivity to rationale quality: The simulation deliberately makes the reference systematically incorrect for three of eight hard decision types, and naive training on it plateaus at 12.8% hard accuracy. How RGPO behaves with noisier, partially relevant, or self-generated rationales is an open question.
- Improving the transfer of guided responses: Only 7.4 guided responses were accepted on average in the main simulation setting; understanding how to raise the yield of useful guided trajectories, and how the margin δ should be tuned, are natural follow-ups.
- Scaling and broader task coverage: The empirical study covers a 3B language model and 2B/7B vision-language models with 3 epochs each. Whether the advantages persist at larger scales, over longer training, and in domains such as long-horizon agentic tasks or code generation is not reported.
Target Audience
Researchers and engineers working on reinforcement learning post-training of large language and vision-language models, especially those dealing with reward sparsity, hard-example stagnation, or hybrid on-policy/off-policy training. It is also relevant to practitioners who have access to reference solutions or expert traces and want to use them without forcing format-matching imitation, and to readers interested in how cognitive-science principles about guidance and scaffolding are being operationalized in RL training loops.
Authors’ abstract
On-policy reinforcement learning has become a central paradigm for improving the reasoning abilities of large language models. However, its effectiveness is often limited by reward sparsity: when a model fails to discover correct trajectories for difficult problems, the optimization process receives little useful signal and may stagnate. Existing approaches mitigate this issue by incorporating off-policy demonstrations, expert traces, or model-generated solutions, but they typically require the auxiliary data to match the format of the reinforcement-learning task, often relying on rejection sampling from stronger models to obtain suitable training trajectories. We introduce Rationale-Guided Policy Optimization (RGPO), a framework that adaptively leverages ground-truth rationale information according to the model's current capability while preserving its freedom to explore. Rather than treating reference solutions as fixed imitation targets, RGPO uses them as temporary scaffolds: rationales help the model generate improved responses, after which only higher-reward, model-generated solutions are transferred back to the original unguided setting. This design allows training to exploit available ground-truth information without requiring off-policy data to follow the same format as the RL task. Across both language-only and vision-language reasoning settings, RGPO consistently improves performance over RLVR baselines, and ablation studies show that adaptive rationale guidance is a key contributor to these gains. These results suggest that RGPO offers a practical and general approach for reducing reward sparsity, stabilizing reinforcement learning, and improving reasoning performance in both text-only and multimodal models.