Skip to content
AI.info

Research

When Are Two Scores Better Than One? Investigating Ensembles of Diffusion Models

Overview Research area: Generative modeling — specifically ensemble methods applied to score-based diffusion models. Technical level: Advanced. The paper combines empirical benchmarking with stochasti

When Are Two Scores Better Than One? Investigating Ensembles of Diffusion Models
arXiv
2601.11444
Published
2026-01-16
Authors
Raphaël Razafindralambo, Rémy Sun, Frédéric Precioso, Damien Garreau, Pierre-Alexandre Mattei

AI summary

Overview

  • Research area: Generative modeling — specifically ensemble methods applied to score-based diffusion models.
  • Technical level: Advanced. The paper combines empirical benchmarking with stochastic-differential-equation theory (reverse-time SDEs, denoising score matching) and formal propositions about density composition.
  • Scope: A combined theoretical and empirical investigation of whether aggregating the score predictions of multiple diffusion models improves generative quality on image and tabular data.

What This Paper Is About

Diffusion models are usually improved by making a single network bigger or training it longer, which is expensive. Ensembling — training several models and combining their outputs — is a standard, cheap way to improve supervised models, but nobody had systematically checked whether it works for unconditional score-based diffusion models.

The paper asks a simple question: if you take several independently trained diffusion models and combine their score estimates at every denoising step, do you actually get better samples? The answer the authors report is largely no, at least for perceptual image-quality metrics.

Key Contributions

  1. A framework for ensembling diffusion models. The authors describe how to build K diverse score predictors (via Deep Ensembles and Monte Carlo Dropout for neural networks, and via Random Forests for tabular data) and how to aggregate them, introducing and comparing several aggregation rules: arithmetic mean, geometric mean, median, the sum of scores, Mixture of Experts, Alternating Sampling, and Dominant Feature.

  2. A theoretical analysis of score averaging. They prove a monotonicity result for the DDSM (Denoising Diffusion Score Matching) loss — adding a model to an equally-weighted ensemble does not increase the expected loss — and then show that a natural interpretation of that averaging (sampling from a Product-of-Experts) is incorrect, because diffusing and composing densities do not commute.

  3. A broad empirical evaluation. Across CIFAR-10 and FFHQ and a range of aggregation rules, ensembling does not significantly improve FID or KID, despite consistently reducing the score-matching loss. Tabular data is examined through random forests; the paper states that one aggregation strategy outperforms the others there (the visible portion of the text is truncated before the tabular numbers).

  4. Identification of conditions and explanations. The authors highlight two settings where ensembling does help — Mixture of Experts when members are sufficiently complementary, and one decision-tree aggregation method that alleviates an underestimation bias — and investigate model diversity and the mismatch between the score-matching loss and perceptual quality as explanations for the general failure.

Main Findings

  • Better training loss, no better images: Ensembling consistently reduces the DDSM objective, but this reduction does not translate into consistent gains in perceptual metrics such as FID and KID. The authors state that perceptual improvements are "at best marginal."

  • Ensembles rarely beat the best single model: On CIFAR-10, the best individual model reached FID 4.65 and KID 0.0006, while the arithmetic mean ensemble of K = 5 reached FID 4.98 and KID 0.0007, and the geometric mean reached FID 4.97 and KID 0.0008. The mean of individual models was FID 4.79 (SD 0.08) with KID 0.0008 (SD 0.0001), and its DDSM loss was 0.0284 (SD 0.0005).

  • Small loss reductions do appear: On CIFAR-10 the arithmetic-mean ensemble reduced the DDSM loss to 0.0280 (SD 0.0005) versus 0.0284 (SD 0.0005) for the average individual model; the geometric mean also reached 0.0280 (SD 0.0005) and the median 0.0281 (SD 0.0006).

  • Aggregation rules mostly look the same: Arithmetic mean, geometric mean, and median produced very similar scores on CIFAR-10 (FID 4.98, 4.97, 4.98 respectively) and on FFHQ (FID 23.44, 24.32, 23.13). The paper states that the four aggregation schemes combining models at each timestep show no significant differences.

  • Dominant Feature is clearly worse: Selecting the largest-magnitude score coordinate across experts degraded quality sharply: FID 7.88 and KID 0.0041 on CIFAR-10, and FID 46.79, KID 0.040 on FFHQ, with a DDSM loss of 0.0287 (SD 0.0004) on CIFAR-10.

  • Random-selection methods were competitive but not decisive: On FFHQ with K = 4, Mixture of Experts obtained FID 20.36 and KID 0.011, and Alternating Sampling FID 23.07 and KID 0.012, compared to FID 21.7 and KID 0.012 for the best individual model and 24.20 (SD 5.70) with KID 0.015 (SD 0.003) for the average individual model.

  • Ensembling can even hurt: The authors report that on CIFAR-10 (shown in an appendix figure) ensembling can perform worse than the weakest individual model.

  • Theory confirms the loss improves but explains the gap: Proposition 4.1 shows that, when the individual score outputs are identically distributed conditional on the noisy input and timestep, the expected DDSM loss of a (K+1)-model ensemble is less than or equal to that of a K-model ensemble. Proposition 4.2 shows that for centered Gaussian initial densities with diagonal covariances, noising and taking a Product-of-Experts do not commute unless all the covariances are identical — i.e., the averaging step only reduces to a PoE when the ensemble collapses to essentially one model.

  • Guidance is implicated too: Because of the non-commutativity result, the paper argues that classifier-free guidance (Equation 6) does not exactly produce its target distribution, and should be viewed as a heuristic approximation.

Methodology in Plain English

The authors train multiple diffusion models separately, each with a different random initialization, on the same datasets. At sampling time, instead of asking a single network for the score at each denoising step, they ask all K networks and combine their answers with a chosen rule — averaging them, taking a geometric mean, taking a median, picking one expert at random per sample, picking a different expert per noise level, or taking the largest-magnitude coordinate across experts. They then measure FID and KID (perceptual image-quality metrics) and the training objective itself.

For the neural-network experiments they use Deep Ensembles (independently trained U-Nets) as the main setup, plus Monte Carlo Dropout, which keeps dropout active at test time on a single trained model to get multiple stochastic forward passes without retraining. For comparison with a non-neural approach, they use Random Forests in a Forest-VP-style diffusion setup for tabular data, varying the number of trees and the aggregation scheme.

Evaluation is done on CIFAR-10 at 32×32 and FFHQ at 256×256, with FID and KID computed on 10k samples rather than the standard 50k. The authors note this inflates FID (they call it FID-10k but write FID for short). For FID they report confidence intervals computed over all subsets of the model pool for a given ensemble size K, and for the DDSM loss they report intervals from 10 runs. Section 5.3 estimates FID intervals by bootstrap; KID follows the same procedure. On the theory side they prove a monotonicity statement for the DDSM loss via Jensen-style arguments and build an explicit Gaussian counterexample within the Variance Preserving SDE framework to show that diffusion and Product-of-Experts composition do not commute.

Why This Matters

Impact on research. The paper undercuts a common intuition that ensembling should be a free lunch for diffusion models, and separates two things that are often conflated: lower score-matching loss and better perceptual sample quality. It also supplies a caveat that extends beyond ensembling to guidance and other model-composition techniques, since those methods rely on the same kind of score summation.

Real-world applications. The paper's introduction situates diffusion models in several applied domains:

  • Image generation and text-to-image synthesis.
  • Video modeling.
  • Graph and molecular design.
  • Tabular data generation and imputation.

Industry relevance. Training multiple diffusion models is expensive in compute, memory, and inference time. A negative result here matters commercially: practitioners who were considering ensembling as a cheap alternative to scaling a single model now have evidence that it mainly buys a better training objective, not better images. The appendix attempts reported in the paper — using score averaging to compensate for discretization error with fewer DDIM steps, and averaging only during early sampling steps — also failed to beat individual models, which matters for anyone hoping to cut inference cost. On the positive side, the paper reports small benefits for Mixture of Experts when the models are complementary and for one decision-tree aggregation rule used on tabular data. The authors release their code at a public GitHub repository.

Future Directions

  • Closing the gap between loss and perception. The paper explicitly investigates the mismatch between the score-matching loss and perceptual quality metrics but leaves open what a better ensembling objective would look like; the authors state that progress "will require going beyond this aggregation strategy."

  • Diversity design. Section 5.2.3 examines model diversity as a factor, motivated by prior work showing that diffusion models with different initializations or architectures converge to similar functions. How to construct genuinely complementary diffusion models remains unresolved.

  • Correcting composition-based methods. Given that score averaging does not sample from a Product-of-Experts and that classifier-free guidance does not exactly target its stated distribution, the paper points to correction mechanisms such as MCMC-based correctors and the Feynman–Kac equation as directions for making composition exact.

  • Extending the tabular analysis. The paper states that on tabular data one aggregation strategy outperforms the others and that it alleviates an underestimation bias in decision-tree aggregation; generalizing and explaining this behavior beyond the settings tested is a natural next step.

Target Audience

This paper is most useful to machine-learning researchers and graduate students working on diffusion models, generative modeling, or ensemble methods, and to practitioners who are considering ensembling as a cheaper alternative to scaling a single model. It will also interest anyone working on guidance and model composition, since the non-commutativity result applies there. The theoretical sections assume comfort with stochastic differential equations and score matching; the experimental sections are readable by a broader audience with basic knowledge of FID and KID.

Authors’ abstract

Diffusion models now generate high-quality, diverse samples, with an increasing focus on more powerful models. Although ensembling is a well-known way to improve supervised models, its application to unconditional score-based diffusion models remains largely unexplored. In this work we investigate whether it provides tangible benefits for generative modelling. We find that while ensembling the scores generally improves the score-matching loss and model likelihood, it fails to consistently enhance perceptual quality metrics such as FID on image datasets. We confirm this observation across a breadth of aggregation rules using Deep Ensembles, Monte Carlo Dropout, on CIFAR-10 and FFHQ. We attempt to explain this discrepancy by investigating possible explanations, such as the link between score estimation and image quality. We also look into tabular data through random forests, and find that one aggregation strategy outperforms the others. Finally, we provide theoretical insights into the summing of score models, which shed light not only on ensembling but also on several model composition techniques (e.g. guidance).

Read the original paper