Research
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
Overview Research area: Reinforcement learning with verifiable rewards (RLVR) for large language models, specifically policy-gradient optimization of the Pass@K objective. Technical level: Advanced. T

- arXiv
- 2510.23049
- Published
- 2025-10-27
- Authors
- Christos Thrampoulidis, Sadegh Mahdavi, Wenlong Deng
AI summary
Overview
Research area: Reinforcement learning with verifiable rewards (RLVR) for large language models, specifically policy-gradient optimization of the Pass@K objective.
Technical level: Advanced. The paper is a theory and derivation paper built on gradient computations, surrogate objectives, and binary-reward statistics.
Scope: A single-paper unification showing that advantage-shaping methods (which directly modify GRPO advantage scores) and direct REINFORCE-style Pass@K optimization are equivalent to optimizing implicit surrogate rewards, plus a recipe for designing new algorithms from a chosen surrogate reward.
What This Paper Is About
Large language models on math and coding tasks are evaluated with Pass@K — whether at least one of K generated solutions is correct — but most policy-gradient methods such as REINFORCE, RLOO, and GRPO optimize a single-attempt 0/1 reward, creating a mismatch between training and testing. Two research lines attack this: direct Pass@K policy gradients with reweighted advantages, and "advantage shaping" that directly edits GRPO's advantage scores. This paper shows the two lines are two sides of the same coin, by revealing that each corresponds to maximizing a different surrogate reward function of the per-example reward.
Key Contributions
-
Direct Pass@K optimization (C-1). Building on prior work, the authors show the exact Pass@K policy gradient equals the 0/1 gradient reweighted by per-example Fail@(K−1) probabilities. This yields Monte Carlo unbiased estimators REINFORCE_K and RLOO_K, plus a biased variant GRPO_K, all expressible as asymmetric reweightings of their 0/1-reward counterparts.
-
Advantage shaping = surrogate reward maximization (C-2). Reverse-engineering the advantage-shaping algorithm of prior work shows it asymptotically (in the number of generated responses) maximizes the surrogate reward (2/K)·arcsin(sqrt(ρ_{K,θ})). Conversely, the authors give a forward-engineering recipe: differentiate any differentiable surrogate F:[0,1]→ℝ, replace F′(ρ_θ) with F′(ρ̂), and replace the population reward gradient with its RLOO proxy, yielding shaped advantages A_i^F = F′(ρ̂)·sqrt(ρ̂(1−ρ̂))·A_i^GRPO.
-
Reward-level regularization (C-3). The common heuristic of downweighting "easy" examples and upweighting "hard" ones is interpreted as regularized surrogate reward maximization. Multiplying GRPO gradients by 1−ρ̂ is shown to be effectively equivalent to optimizing the regularized surrogate arcsin(sqrt(ρ_θ)) + sqrt(ρ_θ(1−ρ_θ)).
-
A bidirectional framework and unifying table. Table 1 maps advantage scores ↔ weighted empirical gradients ↔ population surrogate reward for RLOO, GRPO, Skew-R (this work, effectively equivalent to Kimi 1.5 Prioritized Sampling), RLOO_K, GRPO_K, and the advantage-shaping variant.
Main Findings
-
Pass@K gradient is a reweighting of the 0/1 gradient. The population per-example Pass@K gradient equals K·(1−ρ_{K−1,θ}(x,a))·G_{0/1}(θ;(x,a)), where the underbraced term is the Fail@(K−1) probability.
-
All the algorithms share one shape. Every method in the study can be written as w₊∇̂₊ − w₋∇̂₋, where ∇̂₊ and ∇̂₋ are the average empirical log-probability gradients over correct and wrong responses, and w± are effective gradient weights.
-
GRPO is itself a surrogate-reward method. When K=1, GRPO is (up to clipping) RLOO applied to the surrogate reward F(ρ_θ) = 2·arcsin(sqrt(ρ_θ)), a variance-stabilizing transformation of the 0/1 reward.
-
The advantage-shaping algorithm implicitly optimizes an arcsin transform. The GRPO variant of prior work rescales the vanilla GRPO empirical gradient and, for sufficiently large sample size N ≫ K, corresponds to maximizing the per-example surrogate reward (2/K)·arcsin(sqrt(ρ_{K,θ})), which is strictly monotone with the Pass@K reward and shares its optimum.
-
Symmetric versus asymmetric reweighting distinguishes the two views. Advantage shaping uses the same multiplier ω̃_K for correct and wrong responses, whereas direct optimization uses asymmetric weights f̂_{K−1}⁺ ≠ f̂_{K−1}⁻. Direct optimization is described as more aggressive at dampening contributions when the empirical reward is high and amplifying more when it is low.
-
Variance stabilization is not coincidental. The arcsin transform satisfies Var(sqrt(M)·arcsin(sqrt(X/M))) ≈ 1/4 independent of p for X ~ Bin(M, p). The authors note that the advantage-shaping method's authors arrived at their approach via batched sampling, where N responses are partitioned into N/K groups of size K and the number of successful groups is binomial with M = N/K trials.
-
Hard-example upweighting is reward-level, not policy- or parameter-level, regularization. It is framed as a data-fitting term F plus an additive regularizer Ω that encourages keeping probability mass on wrong responses, so alternative solution paths may generalize beyond the training set.
-
A concrete instantiation of the recipe. The authors report instantiating the forward-engineering recipe with entropy regularization to yield a new tunable advantage-shaping algorithm (details are in the truncated portion of the paper).
-
No empirical benchmarks are reported in the available content. The paper content provided contains derivations, tables, and figures but no dataset sizes, benchmark scores, or model evaluations; the truncated text ends mid-derivation in Section 5.2.
Methodology in Plain English
The authors start from a finite training set of problem–answer pairs, sample N responses per problem from the model, and score each response with a 0/1 correctness reward. They track two summary statistics per problem: the empirical 0/1 reward and the empirical Pass@K reward, along with the number of correct and wrong responses.
From there they take three steps. First, they differentiate the Pass@K objective directly, using the identity that the expected Pass@K reward equals 1 − (1 − ρ_θ)^K, which produces a gradient equal to the 0/1 gradient multiplied by K and the Fail@(K−1) probability. They build unbiased estimators of that product using leave-one-out versions of the Fail@(K−1) term, and simplify the result into two scalers, f̂_{K−1}⁺ for correct responses and f̂_{K−1}⁻ for wrong responses.
Second, they take an existing advantage-shaping method, rewrite its advantage scores in the paper's own notation, and algebraically simplify the resulting gradient until it reads as a single multiplier times the vanilla GRPO gradient. Because a gradient that looks like a positive function of ρ̂ times the vanilla GRPO gradient matches the chain rule for a transformed objective, they can read off the surrogate reward the method must be optimizing — here, an arcsin transform of the Pass@K reward. They then run the logic in reverse: start with the surrogate, differentiate it, and substitute empirical quantities to recover advantage scores, which reproduces the existing method and gives a general recipe.
Third, they apply that same recipe to regularized objectives — a data-fitting surrogate plus a term that keeps probability mass on wrong answers — and show that familiar hard-example upweighting heuristics fall out as special cases.
Why This Matters
Impact on research. The paper supplies a common language for a scattered literature. Instead of treating REINFORCE-style Pass@K gradients and GRPO advantage shaping as rival techniques, it treats them as choices of surrogate reward. This converts algorithm design from ad hoc "shaping" of advantage numbers into choosing an objective, and it gives a way to reverse-engineer what any reweighting scheme is actually optimizing. It also claims applicability beyond the original Pass@K motivation, positioning surrogate rewards as a general design mechanism for RLVR policy-gradient optimization.
Real-world applications:
- Training coding assistants where a user retries a prompt several times, so the metric that matters is whether at least one attempt passes the tests.
- Training math reasoning models graded by a final numeric answer, where the paper's example setup has a problem x and reference answer a.
- Filtering and weighting training problems in large reasoning-model pipelines, so that already-solved "easy" problems are downweighted and "hard" ones upweighted.
- Reusing existing GRPO codebases with a small change: the recipe says a practical implementation can multiply the GRPO advantages (before clipping) by the prefactor F′(ρ̂)·sqrt(ρ̂(1−ρ̂)).
Industry relevance. GRPO is a widely used RLVR algorithm, and the paper's claim that its reweighting corresponds to an arcsin-transformed 0/1 reward means teams can reinterpret and tune existing training runs without rewriting their training loop. The reward-level regularization view also gives a principled explanation for the popular practice of prioritizing hard examples, such as the Kimi 1.5 Prioritized Sampling behavior that the paper's Skew-R is said to be effectively equivalent to.
Future Directions
- Formalizing the variance-stabilization connection. The authors explicitly call it an open direction to formalize the relationship between variance-stabilizing transformations and optimization stability when empirical gradients are applied to transformed rewards.
- Testing the reward-level regularization view empirically. The paper instantiates its recipe with entropy regularization to produce a new tunable advantage-shaping algorithm; whether this improves over the heuristic it explains is left to evaluation.
- Extending the bidirectional mapping beyond binary rewards. The authors state that they generalize the reverse/forward correspondence to a broad class of binary-reward RLVR gradients; whether the framework carries over to other reward structures is not established in the available content.
- Bridging to optimization theory. Section 8 connects the surrogate-reward formulation of RLVR to conditional stochastic optimization in the optimization literature, suggesting further theory work at that interface.
- Practical limitations. Section 7 collects practical considerations and limitations of the surrogate-reward viewpoint, which the truncated content does not detail.
Target Audience
Researchers and graduate students working on reinforcement learning for large language models, particularly those building or analyzing GRPO-style algorithms and Pass@K objectives. It is also relevant to engineers who maintain RLVR training pipelines and want a principled basis for advantage reweighting and hard-example prioritization. Readers need comfort with policy-gradient derivations, advantage functions, and basic probability (binomial sampling, variance-stabilizing transforms); it is not an introductory paper.
Authors’ abstract
This note reconciles two seemingly distinct approaches to policy gradient optimization for the Pass@K objective in reinforcement learning with verifiable rewards: (1) direct REINFORCE-style methods, and (2) advantage-shaping techniques that directly modify GRPO. We show that these are two sides of the same coin. By reverse-engineering existing advantage-shaping algorithms, we reveal that they implicitly optimize surrogate rewards. We specifically interpret practical "hard-example up-weighting" modifications to GRPO as reward-level regularization. Conversely, starting from surrogate reward objectives, we provide a simple recipe for deriving both existing and new advantage-shaping methods. This perspective provides a lens for RLVR policy gradient optimization beyond our original motivation of Pass@K.