Research
Balancing Rewards in Text Summarization: Multi-Objective Reinforcement Learning via HyperVolume Optimization
Overview Research area: Natural Language Processing — text summarization, large language model post-training, and multi-objective reinforcement learning (MORL). Technical level: Advanced. The paper as

- arXiv
- 2510.19325
- Published
- 2025-10-22
- Authors
- Junjie Song, Yiwen Liu, Dapeng Li, Yin Sun, Shukun Fu, Siqi Chen, Yuji Cao
AI summary
Overview
Research area: Natural Language Processing — text summarization, large language model post-training, and multi-objective reinforcement learning (MORL).
Technical level: Advanced. The paper assumes familiarity with reinforcement learning for LLMs, group relative policy optimization (GRPO), Pareto optimality, and hypervolume indicators in multi-objective optimization.
Scope in one sentence: The paper proposes hypervolume optimization (HVO), a GRPO-based multi-objective RL strategy that replaces the usual weighted-sum reward with a hypervolume-based reward so that a 7B LLM produces summaries that are simultaneously strong across coherence, consistency, fluency, and relevance.
What This Paper Is About
Summary quality is judged on several dimensions at once — coherence, consistency, fluency, and relevance — and pushing one dimension up often drags another down. Most reinforcement learning approaches to summarization optimize a single reward signal, or combine multiple rewards by hand-tuned linear weighting, which the authors argue handles dependencies between objectives poorly. The goal is a training method that steers an LLM toward the Pareto front of these dimensions without supervised fine-tuning or a cold start, producing summaries that are balanced rather than lopsided.
Key Contributions
- HVO, a multi-objective RL strategy built on GRPO for text summarization, which integrates multi-dimensional UniEval rewards into a hypervolume computation and optimizes toward hypervolume maximization, progressively approximating the Pareto optimal frontier — without requiring supervised fine-tuning or cold-start initialization.
- Empirical gains over GRPO on representative benchmarks: higher hypervolume (HV) and overall UniEval scores on both CNN/DailyMail and BillSum, with the 7B HVO model performing comparably to GPT-4 on the two benchmarks.
- A new length constraint mechanism to counter the training instability and summary length collapse that the authors observe in vanilla GRPO when UniEval is used as the sole reward, improving training stability and conciseness.
- Public code release at https://github.com/ai4business-LiAuto/HVO.git.
Main Findings
- HVO beats GRPO on hypervolume and overall score on both datasets. On CNN/DailyMail, Qwen2.5 7B with HVO reaches an HV score of 3.258 and an overall of 0.943, versus GRPO's 1.938 and 0.922. On BillSum, HVO reaches HV 1.961 and overall 0.928, versus GRPO's 1.358 and 0.912. HV scores are reported in units of 10⁻³ per Equation 3.
- Balanced dimensions, not just a higher average. On CNN/DailyMail, HVO scores 0.961 coherence, 0.926 consistency, 0.951 fluency, and 0.934 relevance, with an overall standard deviation of 0.016 compared with GRPO's 0.023.
- A 7B model with HVO is competitive with GPT-4. GPT-4-turbo zero-shot scores 0.921 overall on CNN/DailyMail and 0.904 on BillSum, while Qwen2.5 7B with HVO scores 0.943 and 0.928 respectively. The paper attributes GPT-4's edge to specific dimensions — coherence (0.967) and fluency (0.945) on CNN/DailyMail, and coherence (0.973) and relevance (0.971) on BillSum — but reports that HVO is stronger in overall performance and dimension balance. In the reported table, HVO's CNN/DailyMail fluency of 0.951 is above GPT-4's 0.945.
- HVO produces shorter summaries. The scatter plot of overall score against completion length shows HVO achieving the highest overall score while maintaining a shorter completion length than the alternatives; specific length values are not reported in the provided content.
- Training dynamics differ from GRPO. On the 500-example validation set tracked during CNN/DailyMail training, GRPO prioritizes fluency and relevance early and under-weights consistency, constraining consistency optimization; HVO optimizes all objectives more evenly. HVO's standard deviation is large and persists longer during training, which the authors argue makes the advantage signal more significant and opens a larger strategy space.
- Fine-tuned PEGASUS lags on the multi-dimensional metrics. PEGASUS with SFT scores 0.830 overall (std 0.015) on BillSum and 0.843 (std 0.121) on CNN/DailyMail, the lowest overall among the compared methods on the latter.
- Zero-shot scaling is not monotonic. On CNN/DailyMail, Qwen2.5 1.5B scores 0.872 overall, 7B 0.879, 14B 0.881, and 32B 0.897; on BillSum, 1.5B scores 0.896, 7B 0.892, 14B 0.904, and 32B 0.905.
Methodology in Plain English
The starting point is GRPO, a reinforcement learning method that, for each prompt, samples a group of candidate outputs (here, 8 generations), scores them, and normalizes rewards within the group to build an advantage signal that updates the policy. HVO keeps this structure but changes how the reward is computed.
Instead of combining dimension scores with fixed weights (the reward would simply be the weighted sum of coherence, consistency, fluency, and relevance), HVO treats each summary as a point in a multi-dimensional objective space and computes a hypervolume-style score for it. The intuition: when two summaries have similar weighted-sum scores, the one whose dimensions are more balanced gets the higher hypervolume value. Each dimension's score is first shifted by the minimum value in the group and bounded between small constants δ (to avoid zeros) and ε (an upper bound), then multiplied across dimensions with an exponent of −wᵏ. Because hypervolume-based evaluation is Pareto-consistent, optimizing it should push the policy steadily toward the Pareto front rather than trading one dimension off against another.
The authors also observed that using UniEval alone as the reward caused training instability and "length collapse," where generated summaries shrink pathologically. They added a conciseness term that scores how close a summary's compression ratio (input length divided by output length) is to the average compression ratio of human-written reference summaries in the training set. This term decreases slowly near a perfect match to preserve exploration, then falls off steeply for large deviations, controlled by a steepness parameter λ and an offset ρ.
Rewards come from UniEval, a multi-dimensional evaluation tool that the paper describes as strongly correlated with human judgment. Training uses a "R1-Zero-like" paradigm: GRPO applied directly to base LLMs with no supervised fine-tuning step. Hyperparameters: train_batch_size = 64, num_generations = 8, max_grad_norm = 0.4, learning_rate = 5e-7, all rewards weighted equally at 1.0; for HVO, ε = 0.99, δ = 0.1, ρ = 16, λ = 2, and wᵏ defaulted to −1 to align with GRPO. Evaluation uses UniEval per dimension plus their average ("overall"), and the HV score from Equation 3.
Why This Matters
Impact on research. The paper shows a route to multi-objective RL for LLMs that avoids the pairwise gradient projections used by prior work such as MDO, which the authors say is too computationally expensive to integrate into LLMs. It also provides a concrete alternative to manually weighted reward sums, which the paper argues handle inter-objective dependencies poorly. The reported comparison against GRPO and GPT-4 on two established benchmarks gives a reference point for future work on balanced summarization, and the code is public.
Real-world applications:
- News summarization, where a system must stay faithful to the source (consistency) while remaining readable and on-topic — the CNN/DailyMail setting.
- Legislative and legal document summarization, where summaries of bills must preserve detail and coherence — the BillSum setting.
- Enterprise document workflows that require concise but balanced digests of long reports, where excessive length is itself a cost.
- Any summarization deployment where a smaller open-weight model (7B) is preferred over an API-based frontier model for cost, latency, or data-privacy reasons, since HVO's 7B model is reported as comparable to GPT-4 in overall score.
Industry relevance. The authors are affiliated with AI-4-Business of Li Auto Inc. in Beijing, and the paper's emphasis on avoiding supervised fine-tuning, keeping generation short, and matching a much larger closed model with a 7B open-weight model reflects practical deployment constraints rather than purely academic ones.
Future Directions
- Broaden the objective set and reward models. The current work uses four UniEval dimensions; whether HVO scales to additional objectives (faithfulness, factuality, style) or to other reward models is not reported.
- Validate beyond automatic metrics. All reported results use UniEval and hypervolume; human evaluation is not reported, so the correlation between HVO's hypervolume gains and human preference remains an open question.
- Explain and exploit the training-dynamics effect. The paper observes that HVO's larger and more persistent standard deviation widens the advantage signal and the explored strategy space, but does not isolate which components (hypervolume reward, conciseness term, or their interaction) drive the stability improvement.
- Push toward the Pareto front on larger models and more datasets. Only the 7B model is compared between GRPO and HVO, and only two datasets are used; whether the approach transfers to larger or non-Qwen models, and to domains such as dialogue or scientific summarization, is untested.
Target Audience
Readers who will benefit most are NLP and machine learning researchers working on reinforcement learning for LLM post-training, and practitioners building summarization systems who need multi-dimensional quality rather than a single scalar reward. It is also relevant to engineers choosing between fine-tuning a smaller open model and calling a frontier API, and to multi-objective optimization researchers interested in hypervolume indicators applied to language generation. The paper is written for an audience already comfortable with GRPO, Pareto optimality, and policy-gradient objectives.
Authors’ abstract
Text summarization is a crucial task that requires the simultaneous optimization of multiple objectives, including consistency, coherence, relevance, and fluency, which presents considerable challenges. Although large language models (LLMs) have demonstrated remarkable performance, enhanced by reinforcement learning (RL), few studies have focused on optimizing the multi-objective problem of summarization through RL based on LLMs. In this paper, we introduce hypervolume optimization (HVO), a novel optimization strategy that dynamically adjusts the scores between groups during the reward process in RL by using the hypervolume method. This method guides the model's optimization to progressively approximate the pareto front, thereby generating balanced summaries across multiple objectives. Experimental results on several representative summarization datasets demonstrate that our method outperforms group relative policy optimization (GRPO) in overall scores and shows more balanced performance across different dimensions. Moreover, a 7B foundation model enhanced by HVO performs comparably to GPT-4 in the summarization task, while maintaining a shorter generation length. Our code is publicly available at https://github.com/ai4business-LiAuto/HVO.git