Skip to content
AI.info

Research

Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL

Overview Research area: Multi-reward reinforcement learning for large language model post-training, specifically the aggregation of multiple reward signals inside group-relative policy optimization. T

Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
arXiv
2610.00574
Published
2026-09-30
Authors
Tong Zheng, Skylar Zhai, Zhan Cheng, TianMing Sha, Youling Huang, Shuo Zhou, Shaotong Qi, Jingcheng Liang, Xuwei Ding, Pengcheng Xu

AI summary

Overview

Research area: Multi-reward reinforcement learning for large language model post-training, specifically the aggregation of multiple reward signals inside group-relative policy optimization.

Technical level: Advanced. The paper combines a theoretical derivation of a density–energy relation for group-normalized advantages with empirical training studies on tool calling and mathematical reasoning.

Scope: The paper identifies why some reward dimensions receive weaker learning signals than others after reward-wise normalization, proposes a density-based correction called DARA, and tests whether that correction accelerates learning without hurting final task performance.

What This Paper Is About

When large language models are trained with several reward signals at once (for example, correctness, output format, and response length), existing methods such as GDPO normalize each reward separately to preserve its relative information. Even so, the objectives learn at uneven rates: some reward dimensions dominate while others are starved of useful learning signal. This paper asks what optimization mechanism produces that imbalance, and whether reward weights can be derived from that mechanism instead of tuned by hand.

Key Contributions

  1. A quantitative account of reward contribution. The authors define advantage energy as the sum of a reward's squared advantages over a batch, and define active-group density as the fraction of rollout groups in which a reward produces nonzero relative advantages. Under idealized GDPO normalization they prove that advantage energy is proportional to active-group density, exposing a residual batch-level signal imbalance that survives reward-wise normalization.

  2. A density-based aggregation method. From that relation they derive an inverse-square-root correction, sqrt(π_ref / π_k), that matches the energy of less frequently active rewards to the highest-density reward in the batch. This yields DARA (Density-Aware Reward Aggregation), with two variants: DARA-Sym, which scales positive and negative advantages alike, and DARA-Asym, the main method, which applies the extra amplification only to positive advantages. A cap w_max limits amplification.

  3. An empirical evaluation of learning speed and stability. Experiments on tool calling and mathematical reasoning show DARA reaches high format compliance in up to 26% fewer training steps and near-saturated length compliance in up to 65% fewer steps than GDPO, while remaining competitive in final performance.

  4. A released implementation. The code is available at github.com/zhaihaotian/DARA.

Main Findings

  • Advantage energy scales linearly with active-group density. Under idealized GDPO normalization (with ε = τ = 0, zero advantages for constant-reward groups, and sample standard deviation with divisor G − 1), every active group contributes exactly G − 1 units of energy, so E_k = B·π_k·(G − 1). GDPO equalizes the contribution of each active group, but the batch total still scales with how often a reward is active.

  • A binary reward's density is driven by success probability, not just difficulty. A rollout group is active with probability 1 − p^G − (1 − p)^G, which grows with group size G for any success probability p. Very hard rewards produce mostly all-failure groups and near-saturated rewards produce mostly all-success groups; rewards in between are active far more often.

  • Convergence tracks active-group density. Median format reward reaches 0.8 at steps 14–15 for both DARA variants, versus 19 for GDPO and 34 for GRPO. In the group-size study (G ∈ {4, 8, 16, 32} at 2,048 responses per step), larger G raised format active-group density for all methods and made the format reward reach 0.8 earlier. GDPO learned the format objective slowly and inconsistently at small G but reliably at G = 16 and 32; DARA reached 0.8 earlier than GRPO and GDPO at every group size.

  • Faster convergence transfers to downstream tool calling. At step 60 on Qwen2.5-1.5B-Instruct, DARA-Sym and DARA-Asym achieved the highest Average Accuracy (50.94% and 50.69%) and Average Format (96.06% and 96.02%) among compared methods, while baselines reached at most 50.50% Average Accuracy and 90.11% Average Format. At step 100 on the 1.5B model, DARA-Sym reached 51.17% Average Accuracy and 97.55% Average Format and DARA-Asym 50.59% and 97.90%, versus 48.46%/81.01% for GRPO and 50.37%/97.11% for GDPO. On the 3B model, DARA-Sym achieved the highest Average Format (97.22%) and near-best Average Accuracy (54.15%), while DARA-Asym performed on par with GDPO (53.63% accuracy, 96.50% format).

  • DARA resists interference from an added third reward. Adding a length reward (one when the <think> block contains at most 16 words) lowered GDPO's Average Format from 97.11% to 94.39% and Multi-Turn format from 91.40% to 83.17%. DARA changed far less: 97.90% to 96.93% for DARA-Asym and 97.55% to 97.43% for DARA-Sym.

  • The benefit grows as a reward approaches saturation. At moderate length-compliance thresholds methods differ little — on Qwen3-4B-Instruct, DARA-Asym, DARA-Sym, and GDPO all reach the 80% threshold at step 20. At 99% compliance they reach it at steps 66, 59, and 111 respectively; on DeepSeek-R1-7B at steps 56, 50, and 142, corresponding to 41–65% fewer steps than GDPO.

  • Held-out performance follows the same pattern. On DeepSeek-R1-7B, both DARA variants first reach 95% held-out length compliance at step 50 versus step 70 for GDPO. All methods show an initial accuracy drop when the length constraint is enforced, followed by partial recovery.

  • Accuracy–compliance trade-off. At step 50, DARA-Sym achieved the lowest average Exceed rate on all three models (5.18%, 1.69%, and 1.09% on DeepSeek-R1-1.5B, Qwen3-4B-Instruct, and DeepSeek-R1-7B). DARA-Asym retained higher accuracy (45.83%, 60.30%, 58.31%) with Exceed rates of 8.16%, 4.38%, and 3.73%. On both DeepSeek-R1 models, DARA-Asym achieved the highest Joint score among compared methods (45.14% and 57.83%).

  • The gain is not simply a larger fixed weight. Giving the sparse length reward a fixed weight of five (versus one for correctness) on DeepSeek-R1-1.5B improved length compliance but cost accuracy: DARA-Asym's Exceed fell from 4.21% to 3.14% while Acc dropped from 48.03% to 46.84%; DARA-Sym's Exceed moved only from 0.89% to 0.87% while Acc fell from 46.71% to 45.97%.

  • The gap narrows after full convergence. Extending training to step 100 on Qwen3-4B-Instruct, DARA-Asym reached 62.17% Acc versus 62.99% for GRPO while reducing Exceed from 0.70% to 0.25%. On Qwen3-4B-Thinking it reached 64.23% Acc versus 64.82% for GRPO, reducing Exceed from 1.86% to 0.94%.

  • Energy measurements on logged batches. At step 40 of mathematical reasoning training, the length reward was active in fewer groups than correctness, and under GDPO its energy was 31–46% of correctness; DARA matched the two. Under GDPO the length-to-correctness energy ratio declined toward saturation and stayed below 0.04 from the steps at which the required weight exceeded the cap w_max = 5 (steps 59, 62, and 48 for the three models), whereas the DARA median in that phase was 0.39–0.49.

Methodology in Plain English

The authors start by looking at how much "signal" each reward actually contributes to a policy update. They reason that a reward can only teach the model something in a rollout group where the sampled responses differ on that reward — if every response gets the same score, the group-normalized advantage is zero and the reward says nothing. They call the fraction of such informative groups the reward's active-group density, and they show mathematically that, under GDPO-style normalization, the total squared advantage a reward contributes to a batch is proportional to that density.

That relation gives a natural fix. If a rarely useful reward should contribute as much signal as the most frequently useful reward, its advantages should be scaled by the square root of the density ratio — a correction that grows as a reward gets sparser. The authors cap this scaling at w_max = 5 to avoid extreme amplification when a reward is active in very few groups, and give a weight of 1 to rewards with zero measured density.

They then build DARA, which computes these weights fresh from each rollout batch. Because the weights are recomputed every batch, they follow the changing activity of each reward throughout training. Two variants are tested: DARA-Sym applies the weight to positive and negative advantages alike, while DARA-Asym applies the extra amplification only to positive advantages, keeping negative coefficients at their original GDPO scale so conflicting reward dimensions are not amplified in the wrong direction. DARA-Asym is the main method. Critically, DARA only changes the reward aggregation step; the underlying policy optimization objective is unchanged.

Evaluation covers two domains. For tool calling, models are trained on the ToolRL dataset (3,920 training examples, 80 held-out validation examples) with a binary format reward and a correctness reward in [−3, 3], then tested on BFCL-v4 across Live, Non-Live, and Multi-Turn tasks. For mathematical reasoning, models are trained with DAPO-style training on the DeepScaleR-Preview dataset using the DeepSeek-R1 prompt format, with a maximum response length of 8,000 tokens and two binary rewards (correctness and staying at or under 4,000 tokens), then evaluated on MATH-500, AIME 2024, AMC 2022/2023, Minerva, and OlympiadBench. Checkpoints are tested every 10 training steps during the first 100 steps, and a "stable crossing" requires the trailing 10-step mean to stay above a target for the following 20 optimization steps.

Why This Matters

Impact on research. The paper reframes uneven multi-reward learning as a measurable batch-level quantity rather than a vague tendency of models to "prefer easier objectives." It shows that reward-wise normalization, which already fixed reward collapse, still leaves a systematic imbalance, and it provides a principled derivation for reward weights instead of repeated empirical tuning. The analysis is orthogonal to policy-update mechanisms, so it can in principle be combined with other GRPO variants.

Real-world applications:

  • Tool-calling agents that must simultaneously produce well-formed structured output, select the right functions, and fill in correct arguments reliably.
  • Reasoning assistants that must give correct answers under a latency or cost budget enforced through a response-length constraint.
  • Safety or style constraints layered on top of a task objective, where the constraint reward is sparse because most outputs already satisfy it.
  • Any multi-objective post-training pipeline where practitioners currently hand-tune reward coefficients and want a batch-adaptive alternative.

Industry relevance. Multi-reward post-training is standard practice in deployed LLM systems, and reward-weight tuning is a recurring cost. DARA addresses a practical failure mode — a reward that becomes almost inactive as the model saturates it, exactly when further tightening is wanted — and does so without changing the optimization objective, making it relatively easy to drop into an existing GDPO-style training stack.

Future Directions

  • Interaction with the cap. The authors note that the cap w_max can prevent exact energy matching when the required correction exceeds it; how to set or schedule the cap, and what is lost when it binds, is left open.
  • Choosing between symmetric and asymmetric calibration. DARA-Sym achieves the lowest Exceed rates while DARA-Asym retains higher accuracy, and the paper reports these as different trade-offs rather than resolving which to prefer. Understanding when each is appropriate remains an open question.
  • Why accuracy leads differ across model scales. Pass@1 accuracy is described as more method-dependent, with different aggregation strategies leading on different model scales (for example, GRPO gives the best step-50 result on Qwen3-4B-Instruct). The source of this scale dependence is not explained.
  • Generalization beyond two domains. The evaluation covers tool calling and mathematical reasoning with two and three rewards; whether the density–energy relation and the calibration rule extend to other reward structures, larger numbers of rewards, or non-binary reward shapes is not established.

Target Audience

Researchers and practitioners working on reinforcement learning for LLM post-training, especially those using GRPO-style group-relative methods with multiple reward signals. It will be most useful to readers comfortable with advantage estimation and reward normalization who want a mechanistic explanation of uneven objective learning, and to engineers seeking a principled alternative to hand-tuned reward weights. Readers looking only for a drop-in implementation can focus on Section 3.3 and the released code; readers interested in the theory will find the derivations in Appendix A.

Authors’ abstract

Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative advantages. This reveals a residual batch-level signal imbalance and provides a basis for calibrating reward contributions. Based on this relation, we propose Density-Aware Reward Aggregation (DARA). We derive an inverse-square-root density correction that gives greater weight to signals from less frequently active rewards. DARA computes its weights from each rollout batch, adapting to changes in reward activity throughout training without modifying the underlying policy optimization objective. Experiments on tool calling and mathematical reasoning show that DARA learns the targeted behaviors faster than GDPO, reaching high format compliance in up to 26% fewer training steps on tool calling and near-saturated length compliance in up to 65% fewer steps on mathematical reasoning, while remaining competitive in final performance. Our code is available at https://github.com/zhaihaotian/DARA.

Read the original paper