Research
Semifactual Credit-Augmented Policy Optimization
Overview Research area: Reinforcement learning with verifiable rewards (RLVR) for large language model reasoning, specifically token-level credit assignment and robustness to task-irrelevant prompt fe

- arXiv
- 2609.40360
- Published
- 2026-09-30
- Authors
- Junshu Pan, Zhizhang Fu, Shulin Huang, Yiran Ding, Zifan Cheng, Wenqi Shao, Qiaosheng Zhang, Yue Zhang
AI summary
Overview
Research area: Reinforcement learning with verifiable rewards (RLVR) for large language model reasoning, specifically token-level credit assignment and robustness to task-irrelevant prompt features.
Technical level: Advanced. The paper assumes familiarity with policy-gradient methods, group-relative advantages, importance-ratio clipping, and autoregressive decoding.
Scope: The paper diagnoses token-level sensitivity of LLM reasoning to answer-preserving ("semifactual") prompt perturbations, then converts that measured instability into a training-time credit signal for a GRPO variant called Semifactual Credit-Augmented Policy Optimization (SCAPO), evaluated on two Qwen3 base models across mathematics and out-of-distribution benchmarks.
What This Paper Is About
RLVR methods such as GRPO reward a response based only on whether its final answer is verifiable, applying the same outcome-derived advantage to every token in that response. The authors argue this is too coarse: if some tokens are spuriously sensitive to irrelevant prompt wording, uniform positive credit may reinforce that spurious dependence alongside genuine reasoning.
The paper's goal is to measure this token-level sensitivity using semifactual prompt interventions that change incidental prompt features while preserving the problem and its answer, and then to use that measurement to refine how credit is assigned during RL training.
Key Contributions
-
A token-level diagnostic of spurious dependence. Using semifactual prompt interventions on a frozen model, the authors quantify probability drift per response token, show the sensitivity is highly heterogeneous across tokens and lexical categories, and demonstrate that suppressing high-drift token candidates during decoding raises accuracy without any weight update.
-
SCAPO, a token-level credit-augmented GRPO variant. SCAPO teacher-forces fixed responses under the original prompt and its perturbations, converts probability drift into a group-relative stability score, keeps only the negative part of that score, and adds it as a stop-gradient correction to the GRPO advantage. It requires no process supervision and no external reward model.
-
Empirical validation at two model scales. On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improves AIME 2024–2026 accuracy over GRPO by 5.63 and 4.17 percentage points respectively, and achieves the best results on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks among the compared methods.
-
Ablations isolating what matters in the credit signal. Controlled variants show that the source of the signal (semifactual vs. counterfactual vs. shuffled) and its sign (negative-only vs. all-sign vs. positive-only) both affect results, supporting the specific design choices.
Main Findings
-
Uniform advantages understate token-level sensitivity: Under GRPO, the same advantage is applied to every valid token in a response, so it cannot distinguish tokens that depend on task-irrelevant prompt features from tokens that carry the reasoning.
-
Semifactual drift is heterogeneous: On 870,586 response token drifts collected from 1,000 DAPO-Math-17K questions with frozen Qwen3-4B-Base, 12.82% of positions show no recorded change, while nonzero drifts span several orders of magnitude.
-
Sensitivity differs sharply by token category: Reflection markers show mean drift 2.74× the overall mean (0.01640) and discourse connectives 2.32×, whereas mathematical symbols show 0.47× and numbers 0.37×.
-
Suppressing unstable tokens helps decoding without weight updates: A
d-filtered decoding strategy that ranks non-EOS candidates by drift and masks the longest prefix whose cumulative probability mass does not exceed 0.8 raises frozen Qwen3-4B-Base accuracy from 15.8% to 30.0% (+14.2 points) on the 1,000-question diagnostic panel. -
SCAPO improves in-distribution mathematics: AIME 2024–2026 accuracy rises from 22.08 to 27.71 at 4B and from 7.71 to 11.88 at 1.7B. HMMT 2025–2026 improves by 4.46 and 2.98 points at the two scales. SCAPO achieves the highest accuracy on six of eight mathematics metrics at 4B and all eight at 1.7B, with the highest average accuracy at both scales (39.65 at 4B versus GRPO's 34.93; 23.08 at 1.7B versus GRPO's 17.86).
-
Gains extend out of distribution: SCAPO has the highest accuracy on all three out-of-distribution benchmarks at both scales. GPQA-Diamond improves from 36.26% to 42.45% at 4B and from 27.42% to 31.25% at 1.7B, gains of 6.19 and 3.83 points. Average out-of-distribution accuracy rises from 26.48 to 30.72 at 4B and from 13.70 to 17.22 at 1.7B.
-
Faster convergence: SCAPO reaches GRPO's final AIME accuracy in fewer than half as many policy optimization steps at both model scales.
-
Ablations support the design: On Qwen3-1.7B-Base, SCAPO averages 17.65 across five competition-level benchmarks, versus 13.65 for a random-shuffle control, 11.68 for a counterfactual (answer-changing) control, 15.65 for all-sign corrections, and 14.05 for positive-only corrections. On AIME 24–26, SCAPO leads the shuffle and counterfactual controls by 3.55 and 5.56 points.
-
Probing is computationally cheap: Semifactual probing and credit construction account for 1.8% of measured training runtime in the 4B experiment, because probes reuse sampled responses via teacher forcing and are confined to the initial credit augmentation phase.
Methodology in Plain English
The authors start from a simple question: if you change the wording of a math problem without changing the math, how much do the model's token probabilities move? They build four kinds of answer-preserving perturbations for each training problem — a paraphrase, a single-character typo in an ordinary English word, a light scenario wrapper, and an appended irrelevant snippet — generated with GPT-5.5 under strict instructions to preserve all quantities, expressions, constraints, and the requested answer.
To measure drift, they sample a response once under the original prompt, then force the model to re-score that exact same response under each perturbed prompt. For every token position, the change in that token's probability is converted into a bounded-symmetric distance, which normalizes by the local mean probability so that low-probability tokens are not unfairly penalized, and stays within [0, 2] for numerical robustness. Averaging over the four perturbations gives one drift value per token.
The training method reuses this machinery. For each prompt group, SCAPO computes drift for all tokens across the group's responses, standardizes the negated drift separately per perturbation type, averages the standardized scores, and standardizes again. The result is a within-group relative stability score. Because stability is not the same as correctness, only the negative part of that score is kept, so tokens are never rewarded merely for being stable — they can only lose advantage for being unstable. This correction is detached from the gradient, scaled by λ, and added to the GRPO advantage, which yields a per-token advantage that is always less than or equal to the original group advantage.
The intervention is applied only during early training, to shape which trajectories get reinforced, after which standard GRPO takes over. In the experiments, λ₀ = 0.01 with augmentation for the first 120 steps at 4B and 200 steps at 1.7B. Training uses DAPO-Math-17K, a rollout batch size of 128, an update batch size of 64, 8 rollouts per prompt, a maximum response length of 16,384 tokens, and a constant learning rate of 10⁻⁶. Evaluation samples at temperature 0.7 and top-p 0.9, scores with Math-Verify, and compares against GRPO, GSPO, SAPO, CF-GRPO, and FIPO under matched settings.
Why This Matters
Impact on research. The paper reframes robustness in RLVR from a post-hoc evaluation concern into a training-time signal. It shows that a property usually used to diagnose brittle reasoning — sensitivity to answer-preserving prompt changes — can also shape learning, and it does so without process annotations, Monte Carlo rollouts, or a separate reward model. It also provides a token-level counterpoint to the response-level advantage design that dominates critic-free RLVR methods, and reports that the resulting gains transfer to GPQA-Diamond, a graduate-level science benchmark outside the mathematical training domain.
Potential real-world applications (not all evaluated in this paper):
- More reliable mathematical and quantitative reasoning assistants where the same problem may be phrased many different ways by different users.
- Robustness engineering for deployed reasoning models that must handle users' paraphrases, typos, and irrelevant surrounding context.
- Automated grading or tutoring systems where spurious sensitivity to prompt framing would translate into inconsistent feedback.
- Training pipelines for scientific and technical question answering, using the GPQA-Diamond transfer result as early evidence.
Industry relevance. The augmentation adds roughly 1.8% to measured training runtime in the reported 4B experiment and requires no additional human annotation or reward model, which makes it practical to bolt onto existing GRPO-style RLVR pipelines. The reliance on a small set of precomputed perturbations per training prompt also fits existing data-preparation stages.
Future Directions
-
Scale and architecture. SCAPO was only tested on dense Qwen3 models of 1.7B and 4B parameters trained on DAPO-Math-17K. Whether the effect holds at larger scales, with mixture-of-experts and hybrid architectures, and with broader training data is explicitly left open.
-
Domains beyond mathematics. The paper notes GPQA-Diamond as evidence of transfer, but the training domain remains competition mathematics. Extending the credit signal to code, scientific reasoning, or open-ended tasks is untested.
-
Perturbation design and cost. The perturbations are generated by GPT-5.5 and fixed at K = 4 per prompt. How sensitive SCAPO is to the number, type, and quality of perturbations — and whether semantically targeted perturbations would help more — is not reported.
-
Scheduling and strength of the credit signal. The augmentation is applied for a fixed early window with a constant
λ₀. Whether adaptive schedules, drift-based curriculum, or different stopping criteria would improve results is not investigated; the paper only provides an analysis under a stochastic-gradient model in its appendix.
Target Audience
Researchers and engineers working on RLVR, post-training, and reasoning-oriented LLM fine-tuning will get the most from this paper, particularly those implementing or extending GRPO-style pipelines. It is also relevant to readers interested in causal and invariance-based analyses of model behavior, and to practitioners who care about robustness to prompt perturbation without reweighting. The diagnostic section (Section 2) is readable with moderate background; the method and experimental sections require comfort with policy-gradient objectives and advantage estimation.
Authors’ abstract
Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights. These findings highlight a limitation of Group Relative Policy Optimization (GRPO), which assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful reasoning. Motivated by this observation, we introduce Semifactual Credit-Augmented Policy Optimization (SCAPO), a causally inspired variant of GRPO that incorporates semifactual stability into token-level credit assignment. SCAPO measures token probability drift for fixed responses under semifactual interventions and uses normalized stability scores to reduce advantages for relatively unstable tokens during early training, while granting no additional credit for stability alone. On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improves AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points, respectively. At both model scales, SCAPO achieves the best results on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks among the compared methods. These results suggest that semifactual stability provides an effective training signal for improving reasoning and generalization through finer-grained credit assignment in RLVR. The code is available at https://github.com/DtYXs/SCAPO.