Skip to content
AI.info

Research

CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning

Overview Research area: Reinforcement learning for large language models, specifically multi-reward advantage estimation in Group Relative Policy Optimization (GRPO). Technical level: Advanced. The pa

CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning
arXiv
2609.36820
Published
2026-09-29
Authors
Wenbin Hu, Huihao Jing, Haochen Shi, Yuxuan Liu, Haoran Li, Yangqiu Song

AI summary

Overview

Research area: Reinforcement learning for large language models, specifically multi-reward advantage estimation in Group Relative Policy Optimization (GRPO).

Technical level: Advanced. The paper builds its argument on a covariance/variance-of-a-sum identity, Pearson correlation matrices, and gradient-direction arguments; readers need working familiarity with policy-gradient methods and basic statistics.

Scope: The paper proposes CorrGRPO, a modification to GRPO's advantage normalization that replaces covariance terms with Pearson correlations, and evaluates it on code generation, tool calling, and agent utility-versus-security across models from 0.5B to 8B parameters.

What This Paper Is About

GRPO trains reasoning models by centering rewards and dividing by their within-group standard deviation. When several reward components are summed, that standard deviation mathematically equals the square root of the sum of all pairwise reward covariances, so the size of each update implicitly depends on how the rewards co-vary. The problem is that covariance mixes dependence with scale: a reward component that simply varies over a larger numerical range can dominate the normalization even when it is weakly correlated with the others, muting the influence of smaller-scale rewards. CorrGRPO fixes this by normalizing the pairwise terms into Pearson correlation coefficients while leaving the centered total reward in the numerator unchanged.

Key Contributions

  1. A covariance perspective on multi-reward GRPO. The paper shows that GRPO's denominator is the sum of all entries of the reward covariance matrix, so advantage magnitudes are implicitly scaled by reward dependence, and demonstrates that large-scale reward components can dominate this denominator.

  2. Correlation-normalized advantage estimation (CorrGRPO). CorrGRPO replaces the covariance terms in the denominator with sample Pearson correlation coefficients, giving every nonzero-variance reward an equal diagonal contribution of one, while preserving the centered total reward and the relative reward weights in the numerator.

  3. Experiments across three multi-reward domains. Code generation, tool calling, and agent utility versus security are evaluated with models from 0.5B to 8B parameters, showing gains over GRPO and relevant variants and an outward expansion of the empirical Pareto frontier between reward objectives.

  4. Compatibility and gradient analysis. Appendix A shows CorrGRPO scales each group's gradient by a positive coefficient c_q while preserving its direction, and that the method can be combined with DAPO, CISPO, and GDPO.

Main Findings

  • GRPO's denominator is a covariance sum. For r reward components, the advantage decomposes as the centered total reward divided by the square root of the sum of all pairwise covariances, so stronger positive correlations shrink advantages and weaker or negative correlations enlarge them.

  • Scale can dominate dependence. In the paper's four-trajectory, three-reward example, rewards r₁ and r₂ are strongly correlated (ρ̂₁₂ ≈ 0.9756) while r₃ is weakly correlated with both (ρ̂₁₃ = ρ̂₂₃ ≈ 0.1098). Yet r₃'s variance of 14.1696 alone accounts for approximately 91.1% of the covariance sum S = 15.5520, and its covariances with the other rewards (both 0.1728) exceed the covariance of 0.1707 between the two strongly correlated rewards.

  • CorrGRPO rebalances that example. Under correlation normalization, the strong 0.9756 relationship contributes approximately 8.89 times as much to the denominator sum as each 0.1098 weak relationship, so the small-scale correlated rewards are no longer downweighted.

  • Coding results. CorrGRPO achieves the highest LeetCodeDataset Pass@1 and highest average benchmark Pass@1 at every evaluated model scale (0.5B, 1.5B, 3B, 7B Qwen2.5-Coder-Instruct), outperforming GRPO and GDPO. Average Pass@1 rises over GRPO by 2.09, 0.80, 2.27, and 4.21 percentage points at 0.5B, 1.5B, 3B, and 7B respectively. Its 7B model reaches the highest Pass@1 at 24.12%, its 0.5B model the highest Efficiency at 67.86%, and its 3B model occupies an intermediate frontier point with 14.47% Pass@1 and 60.00% Efficiency. CorrGRPO accounts for every nondominated configuration in the correctness–efficiency tradeoff across evaluated methods and scales.

  • Tool-calling results. CorrGRPO achieves the highest RLLA-4K and API-Bank average all-exact scores for every backbone. RLLA-4K all-exact rises from 63.38% to 67.61% (Qwen2.5-7B), 57.75% to 59.15% (Qwen3-4B-Thinking-2507), and 61.97% to 63.38% (Qwen3-8B). API-Bank averages improve by 1.14, 3.94, and 2.26 percentage points; the largest single gain is Qwen3-4B-Thinking on API-Bank v3, from 52.24% to 64.08%. Overall averages rise by 2.68, 2.67, and 1.84 percentage points.

  • Agent utility versus security. ASB Joint Accuracy improves over GRPO from 14.88% to 18.13% (Qwen2.5-3B), 31.38% to 47.88% (Qwen2.5-7B), and 57.38% to 58.63% (Qwen3-8B). On InjecAgent, overall attack success rate falls from 8.80% to 5.52%, 13.53% to 6.17%, and 13.94% to 7.98%. On AgentDojo, Joint Accuracy improves for Qwen2.5-3B and Qwen2.5-7B, with the largest gain at 3B, from 60.13% to 89.62%.

  • Gradients keep their direction. The ratio between CorrGRPO and GRPO advantages is a positive group-level coefficient c_q computed from the same group's reward statistics, so CorrGRPO preserves the direction of each group's nonzero reward-driven gradient and leaves clipping thresholds unchanged, while adapting gradient magnitude to reward dependence.

  • Slower entropy decay. Training curves indicate CorrGRPO preserves higher policy entropy over training, most visibly on AgentDojo and LeetCodeDataset, though the paper notes this pattern varies across settings.

Methodology in Plain English

The authors start from an algebraic fact rather than a new algorithm. GRPO divides a centered total reward by the group standard deviation of that total. Because the variance of a sum equals the sum of all pairwise covariances, the denominator silently measures how much the reward components move together. Each covariance equals the product of the two components' standard deviations times their Pearson correlation, which is where scale sneaks in: a reward that varies over a large range inflates both its own variance term and its covariance with everything else.

CorrGRPO keeps the numerator exactly as GRPO has it — the centered total reward, so the reward weights chosen by the designer are untouched — and swaps each covariance in the denominator for the corresponding Pearson correlation coefficient. Every reward with nonzero within-group variance then contributes one to the diagonal, and off-diagonal entries carry only the strength and sign of the relationship. Zero-variance components have their rows and columns set to zero, including diagonal entries, which keeps the matrix positive semidefinite and the denominator strictly positive when ε > 0.

The evaluation covers three domains chosen because they naturally involve multiple rewards that either reinforce each other or trade off. Coding uses seven weighted binary and continuous rewards (format, syntax, compilation, runtime, all-tests-passed, AST similarity, and efficiency), trained on the 2,641-problem LeetCodeDataset training split and tested on its 228-problem split plus HumanEval, MBPP, and LiveCodeBench v6. Tool calling uses four reward components for function names, parameter names, parameter values, and format, trained on 3,920 examples from RLLA-4K with 80 test examples and evaluated on API-Bank v1–v3. Agent security uses equal-weight utility and security rewards, trained on a deterministic AgentDojo split of 1,584 training cases and 411 test cases (21 clean, 390 attacked), and evaluated on ASB (400 paired clean and attacked cases per model) and InjecAgent.

Why This Matters

Impact on research. The paper reframes a widely used normalization step as an implicit covariance-weighting mechanism, which gives multi-reward RL practitioners a concrete diagnostic: if one reward has a much larger within-group spread than the others, GRPO's update scaling may be driven by that reward's scale rather than by how the rewards actually relate. It also connects advantage estimation to known multi-reward methods such as MO-GRPO, GDPO, and RDPO by positioning correlation normalization as an alternative to variance-based reweighting and Mahalanobis whitening.

Real-world applications:

  • Code assistants trained to be both correct and fast, where the correctness–efficiency tradeoff is exactly the Pareto frontier the paper measures.
  • Tool-using and function-calling agents, where function selection, parameter naming, and parameter values are separate signals with different numerical ranges.
  • Agent deployments exposed to indirect prompt injection, where resisting attacks and completing legitimate tasks pull in opposite directions.
  • General multi-objective alignment pipelines, where objectives such as helpfulness and safety must be balanced rather than collapsed into one hand-tuned scalar.

Industry relevance. The change is confined to the advantage denominator, requires no extra value model, and is reported as compatible with existing objectives such as DAPO, CISPO, and GDPO, so it can be adopted inside existing GRPO training stacks. Because it equalizes the influence of differently scaled reward components, it reduces the amount of reward-scale tuning needed when teams add new reward signals.

Future Directions

  • Continuing the entropy investigation. The paper observes slower entropy decay under CorrGRPO and cites prior work linking entropy collapse to diminished exploration and performance saturation, but does not establish a causal account of why correlation normalization sustains entropy.

  • Broadening method integration. CorrGRPO is described as compatible with DAPO, CISPO, and GDPO, but the CISPO integration is the one evaluated, and only on Qwen2.5-Coder-7B-Instruct in Appendix F.1; other combinations remain to be tested at scale.

  • Extending the theoretical account. The gradient-direction and quadratic-norm arguments are developed in the appendices under stated assumptions; how the per-group coefficient c_q redistributes relative contributions across groups in the batch gradient is noted as a possibility but not resolved.

  • Reward specification and real-world safety. The ethics statement notes that correlation-based normalization does not correct biased or misspecified rewards, and that benchmark gains against prompt injection do not establish safety against unseen attacks or in real deployments, leaving application-specific evaluation as an open requirement.

Target Audience

Reinforcement learning researchers and engineers who train language models with multiple reward signals, particularly those working on GRPO-style policy optimization, multi-objective alignment, or reward design. It is also relevant to practitioners building coding assistants, tool-calling agents, and agent security defenses, and to readers interested in how normalization choices in advantage estimation shape training dynamics. Readers without a background in policy-gradient methods or basic statistics will need to consult the cited GRPO work first.

Authors’ abstract

Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation. The corresponding variance equals the sum of all pairwise reward covariances. For a fixed centered reward, larger aggregate covariance produces smaller advantages, and vice versa, allowing update magnitudes to adapt to reward dependence. However, correlated rewards with large scales can dominate this normalization and suppress signals from smaller-scale rewards. We propose Correlation-Normalized GRPO (CorrGRPO), which normalizes pairwise covariances into Pearson correlation coefficients. CorrGRPO keeps the centered total reward unchanged while balancing the influence of differently scaled rewards on the correlation-based normalization. This allows advantage magnitudes to adapt to reward correlations without the normalization being dominated by large-scale reward components. We compare CorrGRPO with GRPO and other variants on code generation, tool calling, and agent security, using models ranging from 0.5B to 8B parameters. These tasks all involve multiple rewards that can improve together or present tradeoffs. Results show improvements across three domains, including code generation, tool calling, and agent security. Our code is available at https://github.com/HKUST-KnowComp/CorrGRPO.

Read the original paper