Skip to content
AI.info

Research

Diffusion Reward Models

Diffusion Reward Models Overview Research area: Reward modeling and alignment of large language models (LLM post-training, RLHF), combining diffusion generative modeling with preference learning. Tech

Diffusion Reward Models
arXiv
2609.33803
Published
2026-09-27
Authors
Xiangyang Wang, Bingxiang He, Zeyuan Liu, Jiaze WangZiqing Qiao, Yuxin Zuo, Huan-ang Gao, Cheng Qian, Wenbin Zhang, Ran Li, Youbang Sun, Ning Ding, Yuanchun Shi, Zhiyuan Liu, Chaojun Xiao, Chun Yu

AI summary

Diffusion Reward Models

Overview

Research area: Reward modeling and alignment of large language models (LLM post-training, RLHF), combining diffusion generative modeling with preference learning.

Technical level: Advanced. The paper assumes familiarity with RLHF, Bradley–Terry preference losses, classifier-free guidance, and denoising diffusion / DDIM sampling.

Scope: The paper introduces DRM (Diffusion Reward Model), a reward head that replaces the conventional scalar value head with a lightweight Diffusion Transformer that models the full conditional reward density p(r | x, y), and evaluates it across five reward-model benchmarks plus downstream RLHF.

Paper metadata: arXiv:2609.33803v2 [cs.LG], published 2026-09-27, license CC BY 4.0. Primary affiliations are Tsinghua University, with The Chinese University of Hong Kong and University of Illinois Urbana-Champaign. Code and model links are https://github.com/thunlp/DRM and https://huggingface.co/Teburile/DRM.

What This Paper Is About

Most reward models for LLM alignment compress every prompt–response pair into a single scalar score, or into a distribution from a fixed parametric family such as a Gaussian or a fixed quantile grid. The authors argue this conflicts with human preference, which is inherently multimodal (in the statistical sense of having multiple modes or local peaks) because annotators disagree systematically over values, rubric interpretation, and helpfulness–harmlessness trade-offs. DRM instead recasts reward modeling as conditional density estimation over p(r | x, y) using a diffusion head that imposes no parametric assumption on the output distribution.

Key Contributions

  1. Diagnosis of a shared limitation. The authors identify that scalar reward models, multi-attribute reward models, and parametric-distributional reward models all commit to a fixed output-distribution family, and propose recasting reward modeling as density estimation over p(r | x, y).

  2. The DRM architecture. A lightweight Diffusion Transformer (DiT) reward head is placed on top of a frozen LLM encoder, denoising Gaussian noise into a K-dimensional reward vector with no parametric form imposed on the output. A single architecture handles both multi-attribute regression (K > 1) and pairwise preference data (K = 1) via a distributional Bradley–Terry objective.

  3. A new test-time scaling axis. At inference, N samples form an empirical reward distribution that can be aggregated into a scalar, a variance, or quantiles. This yields a reward axis for test-time scaling (more diffusion samples for a fixed pair) that is unique to DRM, alongside the conventional response axis (Best-of-N over candidate responses).

  4. Empirical validation across five benchmarks plus RLHF. DRM outperforms matched baselines, remains competitive with substantially larger models, recovers multimodal reward structure associated with human disagreement, and improves downstream policy performance when used as the training-time reward.

Main Findings

  • DRM beats matched baselines under identical data and backbone. Trained on the ArmoRM multi-attribute corpus with the FsfairX-LLaMA3-RM-v0.1 encoder, DRM-Multi-8B reaches an average of 66.2 across five benchmarks, exceeding the scalar multi-attribute head ArmoRM (62.3) and the parametric-quantile head QRM (64.1). The largest gain is on RMB Pairwise, where DRM-Multi-8B reaches 78.0.

  • Per-benchmark scores for the two DRM variants. DRM-Multi-8B: RewardBench v2 65.6, PPE Pref 62.5, PPE Corr 63.8, RMB Pairwise 78.0, RM-Bench 68.8, JudgeBench 58.6, average 66.2. DRM-Pref-8B: 65.7, 63.0, 62.5, 78.2, 68.1, 57.1, average 65.8.

  • Competitive with much larger models. DRM stays competitive with Eurus-RM-7B (61.8 average), Skywork-Reward-Llama-3.1-8B-v0.2 (64.8), Llama-3.1-Nemotron-70B-Reward (67.8), DeepSeek-GRM-27B (65.6), and GPT-4o (67.7), and is described as approaching GPT-4o at far lower inference cost. Claude-3.5-Sonnet is listed at 68.5 and Skywork-Reward-V2-Qwen3-8B at 76.9.

  • Matches parametric distributional models without their assumptions. DRM-Multi-8B's 66.2 average is comparable to URM-LLaMa-3.1-8B (66.1) and QRM-Llama3.1-8B-v2 (64.1), while making no distributional assumption about output shape. LDL-Reward-Gemma-2-27B-v0.1 is listed at 67.0, and QRM-Gemma-2-27B at 60.1.

  • The framework transfers across supervision types. DRM-Multi-8B (66.2) and DRM-Pref-8B (65.8) perform similarly despite one being trained on multi-attribute regression and the other on pairwise preference data.

  • Balanced helpfulness and harmlessness. In RMB Best-of-N evaluation, DRM's average difference between Helpfulness and Harmlessness is less than 5 points, in contrast to reward models that score high on one dimension but noticeably lower on the other.

  • Correctness judgment. DRM-Multi-8B reaches 58.6 on JudgeBench and 63.8 on PPE Correctness, with stable performance on subtasks such as math, MMLU, and MBPP. The paper states this outperforms GPT-4o and Claude-3.5-Sonnet; Table 1 lists GPT-4o at 59.8 and Claude-3.5-Sonnet at 64.8 on JudgeBench.

  • Human annotations are genuinely multimodal. On HelpSteer2-Disagreements, helpfulness ratings have a range of at least 2 in 43.49% of examples, separated clusters in 28.20%, and low-high polarization in 24.34%; correctness shows 41.99%, 27.79%, and 24.24%. On MultiPref, 45.64% of overall-preference examples contain opposite-side judgments and 37.23% show separated clusters (helpfulness: 42.77% and 34.37%); 87.11% of harmlessness examples receive fully consistent annotations.

  • DRM aligns with empirical human reward distributions. Measured against repeated human ratings by Wasserstein distance, JS divergence, and L1 distance, DRM achieves the best Wasserstein distance on helpfulness (0.804 vs. 1.030 for an empirical prior, 1.032 for a global Gaussian, and 1.032 for a pointwise baseline) with JS 0.215 and L1 0.907. On correctness, DRM achieves the best Wasserstein distance at 0.846.

  • Multimodality tracks human disagreement. As the human rating range grows, the share of DRM outputs classified as multimodal rises from 37.6% (range at most 1) to 55.5% (range at least 2) to 63.2% (low-high polarized) for helpfulness, and from 37.6% to 54.7% to 62.0% for correctness.

  • Uncertainty-aware rejection improves retained accuracy. Ranking examples by DRM's distributional uncertainty and rejecting the most uncertain raises PPE Correctness accuracy by 2.81 percentage points on average from 100% to 70% coverage (GPQA +3.59, MATH +6.20, MMLU-Pro +1.91, IFEval +1.36, MBPP+ +0.99), and improves RMB by 4.56 to 7.31 points across Helpfulness/Harmlessness Best-of-N and pairwise settings (Helpfulness BoN +6.48, Harmlessness BoN +6.99, Helpfulness Pairwise +7.31, Harmlessness Pairwise +4.56).

  • DRM improves downstream RLHF. Using allenai/Llama-3.1-Tulu-3-8B-SFT as the shared actor and UltraFeedback prompts, DRM-Multi RLHF reaches 2.0 on Arena-Hard v2 and 74.8 on MT-Bench, versus 1.3 and 73.6 for FsfairX RLHF, 1.0 and 71.5 for ArmoRM RLHF, and 1.0 and 59.4 for the SFT baseline.

  • Lower-confidence-bound aggregation. The paper describes an LCB score, LCB_λ = μ − λσ, with λ = 0.4, to penalize candidates with high mean reward but unstable reward distributions at full coverage. The text is truncated at this point, so the resulting LCB numbers are not reported in the available content.

Methodology in Plain English

The system has two parts. First, a frozen LLM encoder reads the prompt and response as a single sequence and outputs the hidden state of the last token, which serves as a fixed semantic summary of the pair. These hidden states are precomputed offline, so the expensive language model is never retrained.

Second, on top of that frozen representation sits a small Diffusion Transformer reward head. Instead of predicting a single score, it learns to reverse a noising process: it starts from random Gaussian noise and iteratively denoises it into a reward vector, conditioned on the encoder's hidden state through adaptive layer normalization. Training uses a masked denoising loss on multi-attribute data, where the mask excludes attributes that were not annotated for a given example, so a single model can absorb datasets with heterogeneous label schemas across a unified 19-dimensional reward space.

For pairwise preference data, absolute labels do not exist, so the authors build symmetric pseudo-reward targets centered at zero with a fixed margin Δ = 1 between chosen and rejected responses, and add a Bradley–Terry ranking loss applied to the denoised reward estimates. The two losses are combined, with λ_BT = 0.5.

At inference, the head draws N = 32 samples using DDIM with 10 steps and classifier-free guidance at scale ω = 7, forming an empirical reward distribution. The mean gives a scalar usable by standard RLHF and Best-of-N protocols; the variance, quantiles, and other distributional statistics support uncertainty-aware rejection and risk-sensitive ranking.

Why This Matters

Impact on research. The paper reframes reward modeling as density estimation rather than point prediction, and shows that a small, non-parametric diffusion head can match or beat parametric distributional heads and larger discriminative and generative reward models under matched data and backbone. It also contributes evidence that annotator disagreement is genuinely clustered and polarized rather than noise around a mean, which is a useful data-side justification for distributional reward modeling.

Real-world applications:

  • RLHF and preference optimization pipelines, where DRM can serve directly as the training-time reward signal, as demonstrated by improved Arena-Hard v2 and MT-Bench scores.
  • Best-of-N selection and inference-time scaling, where the reward axis (more diffusion samples for a fixed pair) tightens scoring precision without changing the candidate set.
  • Selective prediction and human review triage, where the reward distribution flags low-confidence decisions for abstention or escalation to human annotators.
  • Risk-sensitive ranking at full coverage, where lower-confidence-bound aggregation penalizes candidates with high but unstable average reward.

Industry relevance. Because the encoder is frozen and precomputed and the head is small (hidden size 384, 3 blocks, 6 heads), DRM separates expensive language representation learning from cheap reward distribution modeling. That makes distributional reward modeling practical to train, and it gives deployed systems a way to express and act on uncertainty rather than forcing every judgment through a single number.

Future Directions

  • Scaling the diffusion head and the training set. DRM-Multi-8B and DRM-Pref-8B differ by only 0.4 points on average (66.2 vs. 65.8); the authors defer analysis of this gap to a size-matched experiment, leaving open how the head behaves at larger training scale.
  • Exploring the reward axis more systematically. The paper demonstrates that increasing N improves scoring precision, but the full trade-off between diffusion sampling cost, guidance scale ω, and downstream accuracy is not characterized in the available content.
  • Quantifying the benefits of LCB and distribution-aware ranking. The truncated text introduces LCB with λ = 0.4 but the reported results are not available, so the magnitude of full-coverage ranking gains remains an open question.
  • Broader human-disagreement validation. The distributional validation uses HelpSteer2-Disagreements and MultiPref; whether the same multimodal structure is recovered on other annotation regimes, domains, and languages is not addressed.

Target Audience

This paper is most useful to researchers and engineers working on LLM alignment, RLHF reward modeling, and preference learning, particularly those interested in uncertainty-aware or distributional reward signals. It also suits practitioners building Best-of-N, selective prediction, or risk-sensitive ranking systems, and diffusion-model researchers looking for an application beyond text and image generation. Readers without background in diffusion models, Bradley–Terry losses, and RLHF will need to consult the cited background material first, given the paper's mathematical treatment.

Authors’ abstract

Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonably judged in many ways, and no single family covers all of them. To better fit this structure, we introduce DRM, a Diffusion Reward Model that recasts reward modeling as conditional density estimation over $p(\mathbf{r}\mid x,y)$. Conditioned on a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector, placing no parametric assumption on the output distribution and naturally representing its multimodal structure. A single architecture handles both multi-attribute regression and pairwise preference data, and at inference $N$ samples form an empirical reward distribution that can be aggregated into a scalar, a variance, or quantiles. Across five benchmarks, DRM matches or surpasses baselines under matched data and backbone, stays competitive with much larger discriminative, distributional, and generative RMs despite its modest training scale, and recovers multimodal reward structure where conventional heads collapse to a point. Uncertainty-aware rejection and lower-confidence-bound (LCB) aggregation further demonstrate that DRM can exploit distributional information beyond a scalar reward to improve reward-model decisions. Downstream RLHF experiments additionally show that using DRM as the training-time reward leads to improved policy performance, directly validating the practical benefit of diffusion-based reward modeling for RLHF training.

Read the original paper