Research
Best-of-$N$ Guidance for Test-time Diffusion Alignment
Overview Research area: Test-time alignment of diffusion models using human-preference reward models, specifically Best-of-N sampling, Sequential Monte Carlo, and guidance-based methods. Technical lev

- arXiv
- 2610.05108
- Published
- 2026-10-04
- Authors
- Richard Lee Kim, Yeongmin Kim, Gyuwon Sim, Taekyu Kim, Minsang Park, Il-chul Moon
AI summary
Overview
Research area: Test-time alignment of diffusion models using human-preference reward models, specifically Best-of-N sampling, Sequential Monte Carlo, and guidance-based methods.
Technical level: Advanced. The method itself is conceptually simple, but the framing relies on diffusion SDEs, reward-tilted distributions, and self-normalized importance sampling.
Scope: The paper proposes Best-of-N Guidance (BoNG), a rollout-free, sample-based guidance method that moves Best-of-N selection inside the reverse diffusion process, and evaluates it on GenEval prompts with ImageReward, CLIP Score, and GenEval metrics across SD v1.5 (0.9B), SDXL (2.6B), and FLUX.1-dev (12B) backbones.
What This Paper Is About
Diffusion models generate images well but often fail to match human preferences as scored by a reward model. Best-of-N (BoN) sampling is a strong, simple fix: draw N independent samples and return the highest-reward one. Its weakness is that the reward is used only at the very end, so generation itself is never improved, and only one output can be served. The paper's goal is to inject the Best-of-N principle into the denoising trajectory itself, so that both the single best output and the average quality of all outputs improve without extra rollouts.
Key Contributions
- BoNG algorithm: A method that performs online Best-of-N selection over denoising particles at intermediate timesteps and uses the selected highest-reward particle as a guidance signal to steer the rest of the population, with no additional feed-forward passes of the score network.
- Theoretical connection: Lemmas 4.1 and 4.2 reduce reward guidance to estimating the reward-tilted posterior mean via a self-normalized importance sampling (SNIS) estimator, and Theorem 4.3 shows that BoNG's hard BoN selection realizes the large-λ limit of that estimator, matching the BoN guidance term.
- Unified view of prior work: Lemma 4.2 shows LiDAR corresponds to proposal distribution q(x₀) = p_θ(x₀|c) and Tilt to q(x₀) = p_θ(x₀|x_t, c), placing prior sample-based guidance methods in one framework.
- Efficiency and multi-output capability: A rollout-free method that is reported as achieving a 1.3× ImageReward score of the latest sample-based guidance method with a 1.6× speedup, released at https://github.com/aailab-kaist/BoNG.
Main Findings
- Single-output dominance: Over 36 empirical comparisons, BoNG achieves the best performance in 29 cases, ranking first in 80.56% of comparisons against SMC and Vanilla BoN sampling.
- Small budgets suffice: BoNG with N = 4 outperforms Vanilla BoN with N = 8 in GenEval, and gains appear already at N = 2 across SD v1.5 (DDIM 50 steps and DDPM 100 steps) and SDXL (DDIM 50 steps and DDPM 100 steps).
- Beat SMC despite similar cost: On SD v1.5 with DDPM 100 steps at N = 4, BoNG reaches IR 0.894 versus SMC 0.826 and Vanilla 0.742. The explanation given is that SMC collapses particles through repeated resampling, while BoNG preserves all particles and steers them via guidance.
- Better multi-output quality: On SD v1.5 (Table 2, four generated images), BoNG achieves IR 0.506, CLIP 0.280, GenEval 0.511 at 8.06 s and 8.90 G memory, compared with UG (0.326 / 0.262 / 0.355 at 58.36 s, 28.16 G), DATE (0.364 / 0.274 / 0.438 at 32.89 s, 24.71 G), and LiDAR (0.384 / 0.278 / 0.478 at 13.41 s, 8.90 G).
- Budget-matched win over LiDAR: Under matched end-to-end wall-clock time, BoNG (N = 4) reaches IR 0.894, CLIP 0.286, GenEval 0.555 in 8.06 s, versus LiDAR (N = 4, n = 9) at 0.866 / 0.286 / 0.549 in 8.24 s.
- Transfers to modern backbones: On FLUX.1-dev (12B) with the flow-matching Euler 28-step sampler, BoNG beats Vanilla at every N: N = 4 gives IR 1.370, CLIP 0.286, GenEval 0.700 versus Vanilla's 1.353 / 0.287 / 0.696.
- Works with different rewards: With HPSv2 as the guidance reward, BoNG improves mean scores (HPSv2 0.276 versus Vanilla 0.263; IR 0.316 versus −0.018; GenEval 0.487 versus 0.427). With IR–CLIP reward mixing (Z-score normalized), increasing the CLIP weight raises CLIP scores and increasing the IR weight raises IR scores, producing a trade-off frontier.
- Composable with SMC: Applying BoNG on top of SMC improves both the mean and the BoN reward at every particle count.
- Better reward–diversity frontier: BoNG is reported to achieve a consistently better reward–diversity Pareto frontier than SMC, because guidance modifies trajectories gradually rather than duplicating and deleting particles.
Methodology in Plain English
BoNG keeps a population of N particles being denoised together. At selected timesteps, it asks each particle for its one-step estimate of the clean image (the Tweedie estimate), scores all N estimates with the reward model, and picks the highest-reward one. That winner becomes a moving target: the winner itself keeps denoising normally, while every other particle is nudged toward it using an extra term in the score function, proportional to the difference between the winner's clean estimate and its own. The result is an asymmetric interaction in which one particle leads and the others follow, without copying particles and without generating extra lookahead samples. Two knobs control behavior: which timesteps activate guidance (the guidance window, analogous to the reward-strength λ) and a guidance scale multiplying the term. Evaluation uses GenEval prompts with ImageReward, CLIP Score, and GenEval score, comparing against Vanilla BoN and SMC for single-output and against UG, DATE, and LiDAR for multi-output, with cost measured on a single A100 GPU for generating four images.
Why This Matters
Impact on research: The paper gives a theoretical bridge between two previously separate families — Best-of-N selection and reward-tilted-distribution sampling — showing that hard selection is the large-λ limit of an importance-sampling estimator of the tilted posterior mean. That reframing could generalize beyond diffusion to any iterative generator with a scoring model.
Real-world applications:
- Text-to-image product and marketing asset generation, where every generated candidate needs to be usable, not just the best one.
- Creative design tools where a user wants several prompt-aligned options with diverse layouts rather than near-duplicates.
- Content pipelines using larger backbones such as FLUX.1-dev (12B) where the reported gains persist.
- Any deployment using a different preference reward (for example HPSv2) as the quality signal.
Industry relevance: BoNG's cost profile is its main practical selling point. It is reported at 8.06 s and 8.90 G for four images on SD v1.5, essentially matching vanilla sampling (7.07 s, 8.90 G) while outperforming methods that cost 13.41 s to 58.36 s and up to 28.16 G. Sample-based guidance methods normally carry rollouts that make them 2–3× slower than vanilla diffusion sampling; removing that overhead matters for serving latency and memory budgets.
Future Directions
- Hyperparameter selection: The guidance timestep window and the guidance scale are the two control knobs, but the paper does not report a rule for choosing them; an automatic or adaptive schedule is the obvious next step.
- Extending the finite-λ soft variant: Appendix E derives a soft-SNIS relaxation that interpolates between the proposal-set average and the BoN target, with a λ sweep verifying predicted reward–diversity scaling; making this practical at scale remains open.
- Scaling beyond images: The paper demonstrates flow-matching backbones via FLUX.1-dev, but not other modalities; whether the particle-interaction idea transfers to video or other iterative generators is untested here.
- Combining with other alignment families: BoNG on top of SMC already improves mean and BoN rewards at every particle count, suggesting further composition with fine-tuning or parameter-scaling approaches is unexplored.
Target Audience
Researchers and engineers working on diffusion model alignment, inference-time scaling, or reward-guided generation. It is most useful to readers already comfortable with diffusion sampling, SMC, and importance sampling, and to practitioners deciding whether test-time alignment can be deployed without the 2–3× rollout penalty of prior sample-based guidance methods.
Authors’ abstract
Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model. A simple yet effective algorithm for test-time alignment is Best-of-$N$ (BoN) sampling, which draws $N$ i.i.d. samples from a pre-trained diffusion model and outputs the single highest-reward sample. Despite its empirical success, BoN makes limited use of reward information, as it is incorporated only at the final selection stage without influencing the reverse diffusion trajectory during sampling. Consequently, BoN sampling does not improve the average alignment of generated samples and is primarily suited to single-output settings. We propose Best-of-$N$ Guidance (BoNG), a novel method that integrates the principle of BoN sampling directly into the reverse diffusion process. BoNG performs online BoN selection over denoising particles and adjusts the reverse diffusion process to steer the particle population toward higher-reward regions during generation. Specifically, by introducing an asymmetric guidance interaction among denoising particles, BoNG uses the current BoN particle as a guidance signal to the rest of the particle population. This particle-level interaction reshapes the sampling process toward higher-reward regions, enabling BoNG to improve not only the final best sample beyond Vanilla BoN sampling, but also the average quality of generated samples. Over 36 empirical comparisons, BoNG achieves the best performance in 29 cases, ranking first in 80.56% of the comparisons against SMC and Vanilla BoN sampling. BoNG also supports multi-output capability, achieving 1.3$\times$ ImageReward score of the latest sample-based guidance method with a 1.6$\times$ speedup. We release the code at https://github.com/aailab-kaist/BoNG.