Skip to content
AI.info

Research

Entropy Aware Reward Guidance for Diffusion Language Model Alignment

Overview Research area: Alignment of discrete diffusion large language models (dLLMs) using reward-model gradients — a crossover of diffusion modeling, test-time adaptation, and post-training / reinfo

Entropy Aware Reward Guidance for Diffusion Language Model Alignment
arXiv
2602.05000
Published
2026-02-04
Authors
Atula Tejaswi, Litu Rout, Constantine Caramanis, Sanjay Shakkottai, Sujay Sanghavi

AI summary

Overview

Research area: Alignment of discrete diffusion large language models (dLLMs) using reward-model gradients — a crossover of diffusion modeling, test-time adaptation, and post-training / reinforcement learning.

Technical level: Advanced. The paper assumes familiarity with masked diffusion, posterior sampling, straight-through estimators, and reward-model fine-tuning.

Scope: One sentence — the paper introduces an entropy-aware way to feed reward models so that gradients remain both optimizable and interpretable, evaluated on 7B-plus dLLMs for test-time steering and reward-guided post-training.

Note: the provided paper content is truncated inside Section 4.3 ("motivating furthe…"), so results reported after that point are not available here.

What This Paper Is About

Discrete diffusion language models generate text by iteratively unmasking tokens, but their outputs are discrete tokens — so you cannot backpropagate reward-model gradients through them the way you can in continuous image diffusion. Existing fixes either feed the reward model soft, averaged token embeddings (which the reward model never saw during training, making its scores unreliable) or feed it real sampled tokens while faking the gradient with a straight-through estimator (which creates a mismatch between where the reward is measured and where the gradient is applied). This paper asks how to get a strongly optimizable reward gradient without ever handing the reward model an input it cannot interpret, and answers it by interpolating between the two representations token-by-token according to the diffusion model's own predictive entropy.

Key Contributions

  1. EntRGi (Entropy-aware Reward Guidance): a mechanism that dynamically interpolates between continuous token relaxations and sampled hard tokens on a token-by-token basis, using the diffusion model's per-token predictive entropy as the interpolation weight. The weighting is w^l = H(q^l) / log K, which recovers the Expectation method at w^l = 0 and the prior APS method at w^l = 1.

  2. A theoretical framing of the trade-off: the paper defines a vocabulary error D^l = min_k ||e_soft^l − E_k^R||, which penalizes inputs far from real token embeddings, and an approximation error E^l = ||e_hard^l − e_soft^l||, which penalizes gradient/evaluation mismatch. It shows EntRGi strictly reduces the approximation error below APS at every entropy level while keeping the vocabulary error bounded by (1 − w^l) · ||e_hard^l − e_soft^l||, which is small at low entropy and goes to zero as w^l → 1.

  3. RGRL (Reward Guided Reinforcement Learning): a post-training recipe that generates reward-guided completions with Algorithm 1 and then fine-tunes on them via self-distillation, as opposed to standard RL methods that do not guide generation. It works with either APS guidance (RGRL-APS) or EntRGi guidance (RGRL-EntRGi), with the latter providing higher gains.

  4. Empirical validation at scale: results on up to two 7B-plus parameter models (Dream-v0-Instruct-7B and LLaDA-8B-Instruct), 5 multi-skill datasets, and 4 reward models, covering test-time adaptation and post-training, including a mismatched-tokenizer setting.

Main Findings

  • Gradient-based guidance beats Best-of-N everywhere. Across all benchmarks, BoN, Expectation, APS, and EntRGi are compared, and all gradient-based methods consistently outperform BoN (which relies on zeroth-order sampling over randomly generated trajectories).

  • EntRGi beats APS in test-time adaptation. EntRGi achieves a relative improvement of approximately 33% over APS in reward-model-judged output quality, and improves the LMUnit score on RewardBench-2 from 4.19 (APS) to 4.22, and on RM-Bench from 4.01 to 4.06, while also achieving higher Top@1 reward across all tasks.

  • The advantage widens at higher temperature. At τ = 0.7, EntRGi reaches Reward-Bench-2 Top@1 3.91 / Avg@4 2.20 / LMUnit 4.25, JudgeBench Top@1 2.44 / Avg@4 0.02 / LMUnit 3.98, and RM-Bench Top@1 5.70 / Avg@4 3.41 / LMUnit 4.04 — while APS noticeably degrades (RM-Bench Top@1 5.11, JudgeBench Avg@4 −0.63).

  • Straight-through is critical at high-entropy positions. Setting w = 0 at high-entropy positions reduces EntRGi to the Expectation baseline; EntRGi consistently outperforms Expectation, showing that STE-style hardness matters in that regime. At t = T, per-token entropy is typically high at most positions due to limited context.

  • EntRGi works under tokenizer mismatch. On LLaDA-8B-Instruct, whose tokenizer overlap is 45% with Qwen3-0.6B and 55% with Llama-3.2-1B (the Table 2 caption describes this as 45% mismatch with Llama and 55% with Qwen), gradient methods still beat BoN and EntRGi performs best. With Skywork-Reward-V2-Qwen-3-0.6B as reward, EntRGi gets Reward-Bench-2 Top@1 2.80 versus APS 2.35 and BoN 1.77; with Skywork-Reward-V2-Llama-3.2-1B, EntRGi gets 6.40 versus APS 6.32 and BoN 5.34.

  • RGRL improves sample efficiency dramatically. Controlling for training steps, RGRL shows a relative improvement of up to 70% (+0.90 absolute improvement) over diffu-GRPO on WildChat-IF with Dream.

  • Compute cost is a caveat. Faster convergence in wall-clock time was observed in only 1 out of 4 settings, which the authors attribute to the overhead of differentiating through the reward model.

  • Gains scale inversely with initial reward. Improvements are largest on WildChat-IF, smaller on lmsys-chat-1M, and smallest on Magpie-Ultra, where initial rewards are already nearly positive. Under tokenizer mismatch, absolute gains reach up to +0.52 on WildChat-IF.

  • EntRGi reduces early-step approximation error. Measuring the L2 discrepancy between the reward model input and the soft embedding across timesteps (averaged over sequence length L = 128 and 32 prompts, with maximum possible entropy log K ≈ 11), APS shows large error at moderate-to-high entropy (entropy ≈ 4–6) in early decoding, while EntRGi trades vocabulary error against reward reliability to reduce it; both converge to zero as denoising progresses.

  • Larger reward models help everyone. From 0.6B to 4B parameters (evaluated with M = 3 and τ = 0.7), APS improves from an average LMUnit score of 4.00 at 0.6B to 4.08 at 4B, while EntRGi improves from 4.04 to 4.12, with EntRGi ahead at each size.

  • Too many gradient steps causes over-optimization. Increasing M from 1 to roughly 3–4 improves reward and LMUnit on JudgeBench and RM-Bench before degrading; LMUnit collapses beyond M = 4. On Reward-Bench-2, reward scores roughly improve up to M = 5. The experiments use M = 3 for all datasets.

Methodology in Plain English

The starting point is a masked diffusion language model that begins with a fully masked sequence and progressively unmasks tokens, committing each unmasked token permanently. At each step, masked positions hold a probability distribution over the vocabulary.

To steer generation toward high reward, you want to nudge those distributions using gradients from a reward model. The catch is that the reward model is a language model trained on real tokens only.

The authors' move is to construct, for each masked position, an input vector that is a weighted blend of two things: the probability-weighted average of token embeddings (smooth, differentiable, but unfamiliar to the reward model) and the embedding of an actually sampled token (familiar, but non-differentiable). The blend weight is the model's normalized predictive entropy at that position, H(q^l)/log K. The stop-gradient operator is applied to the difference term, so the reward is always evaluated at a point that is mostly composed of real token embeddings, while the gradient still flows back to the logits.

Concretely: when the model is confident (low entropy), the input is close to the smooth average, so gradients are accurate. When the model is uncertain (high entropy), the input snaps toward a sampled real token, so the reward model is not asked to score something out of distribution. The logits at masked positions are then updated by gradient ascent on the reward, for M inner steps per denoising step, and tokens are committed from the updated distribution according to the model's usual low-entropy selection logic.

For post-training, the same guided generation is used to self-distill: draw a prompt, produce N reward-guided completions, and update the model parameters to increase the likelihood of those completions — dense feedback rather than the scalar-reward signal used by policy-gradient RL.

For the mismatched-tokenizer case, the paper defines the vocabulary intersection and sets reward embeddings for absent tokens to zero, so only matched tokens contribute to the gradient.

Why This Matters

Impact on research. The work reframes a known engineering annoyance — discrete outputs break differentiability — as a quantifiable trade-off between two named error terms, and then shows a single scalar (predictive entropy) can steer between them. It provides a principled bridge from continuous-diffusion reward guidance into the discrete diffusion language model setting, and it positions dense reward gradients as a viable alternative to supervised fine-tuning and scalar-reward RL for dLLMs.

Real-world applications:

  • Steerable assistants built on diffusion language models, where a single open-weight dLLM can be adapted at inference time toward safety, factuality, or instruction-following preferences without retraining.
  • Post-training pipelines that need reward models with different tokenizers than the base model, a common situation when reusing off-the-shelf reward models such as the Skywork family.
  • Domains with expensive or non-differentiable-feeling objectives translated into reward models, such as biological sequence diffusion (the related DRAKES work is cited for fine-tuning biological sequence diffusion models).
  • Controllable generation where the best-of-many sampling budget is prohibitive, since guided search outperformed Best-of-N across all reported benchmarks.

Industry relevance. The recipe is drop-in: it needs only a diffusion LLM and a reward model, both increasingly available as open weights. The practical caveats are explicit — extra compute from reward-model backpropagation, dataset-dependent optimal M, and over-optimization risk past M = 4 — which matters for teams deciding between this and a scalar-reward RL pipeline.

Future Directions

  1. Tuning M adaptively rather than per dataset. Optimal gradient-step counts vary across datasets (the experiments fix M = 3, while Reward-Bench-2 improves up to M = 5 and LMUnit collapses beyond M = 4), and the truncated text hints at further discussion of this.

  2. Alternative weighting schemes. The authors state that alternative weighting mechanisms are ablated in Appendix C.3, leaving open whether other uncertainty measures or calibration signals outperform H(q^l)/log K.

  3. Better handling of mismatched tokenizers. Non-overlapping tokens currently receive no gradients (their reward embeddings are set to zero), which likely explains why absolute post-training gains are smaller under mismatch (+0.52 versus +0.90).

  4. Closing the wall-clock gap. RGRL converged faster in wall-clock time in only 1 of 4 settings, so reducing the cost of differentiating through the reward model is an open engineering question.

  5. Generalization beyond the evaluated suite. The work covers 5 multi-skill datasets; the authors' own observation that gains shrink as initial reward rises suggests headroom is concentrated on harder, lower-initial-reward prompts.

Target Audience

Researchers and engineers working on diffusion language models, inference-time alignment, and reward-model-based post-training. A reader needs working knowledge of masked diffusion, gradient-based guidance, and straight-through estimators to follow the error analysis. Practitioners interested only in results can read the tables and figures, but the reasoning about why EntRGi works requires the entropy argument. Those comparing alignment pipelines for open-weight dLLMs — especially across mismatched tokenizers — will find the most direct value.

Authors’ abstract

Reward guidance, also known as posterior sampling, is a popular method for test-time adaptation and post-training in continuous diffusion models. In this paper, we study reward guidance for discrete diffusion language models; now, one cannot differentiate through the natural outputs of the model because they are discrete tokens. We introduce a novel mechanism called EntRGi (Entropy aware Reward Guidance) to address this issue. EntRGi dynamically interpolates between continuous token relaxations and sampled hard tokens, on a token-by-token basis, using the diffusion model's predictive entropy. We demonstrate that EntRGi maintains both reward model reliability and optimization accuracy, while existing approaches sacrifice one for the other. We empirically validate our approach on 7B-parameter diffusion language models across two settings: (1) test-time adaptation, and (2) RGRL (Reward Guided Reinforcement Learning), our recipe for post-training on reward-guided data, showing consistent improvements over state-of-the-art methods. Our code is available at https://atutej.github.io/entrgi-rgrl

Read the original paper