Research
A Diffusive Classification Loss for Learning Energy-based Generative Models
Overview Research area: Generative modeling / statistical machine learning, specifically energy-based models (EBMs), diffusion models, and stochastic interpolants. Technical level: Advanced. The paper

- arXiv
- 2601.21025
- Published
- 2026-01-28
- Authors
- RuiKang OuYang, Louis Grenioux, José Miguel Hernández-Lobato
AI summary
Overview
Research area: Generative modeling / statistical machine learning, specifically energy-based models (EBMs), diffusion models, and stochastic interpolants.
Technical level: Advanced. The paper relies on stochastic differential equations, score matching, Fokker–Planck relations, and asymptotic statistical theory.
Scope: The paper introduces the Diffusive Classification (DiffCLF) loss, a supervised-classification-across-noise-levels objective that trains time-dependent energy-based generative models without the mode-blindness of pure score matching and without the cost of maximum likelihood.
What This Paper Is About
Score-based generative models (diffusion models and stochastic interpolants) are usually trained by fitting the score, the gradient of the log-density. That works well for sampling, but it never recovers the actual energy landscape, and score-only losses are "mode blind": two distributions with the same modes but different mixture weights get nearly identical scores. The paper asks whether you can learn the energies themselves directly, cheaply, and without mode blindness, so that the resulting models can be used for composition, Boltzmann Generator sampling, and free-energy estimation.
Key Contributions
- The DiffCLF objective. A multi-class classification loss over noise levels N that treats the parameterized marginal densities as class-conditional distributions and the noise levels as classes, trainable alongside standard denoising score matching (DSM).
- Theoretical guarantees. The authors prove that DiffCLF recovers the ground-truth marginals at optimality, that the joint objective L_DSM + L_clf has a unique minimizer at the true marginals, and that the Monte-Carlo minimizer is consistent and asymptotically normal. They also show the asymptotic covariance shrinks as the number of classification levels N grows (Corollary D.9).
- Connections to prior work. They show the binary DiffCLF loss converges to Time Score Matching in the continuous-time limit (Proposition 4.1), relate DiffCLF to Noise Contrastive Estimation, and argue that Fokker–Planck and Bayes-rule self-consistency regularizers remain mode blind.
- Empirical validation across diffusion models and stochastic interpolants, including synthetic Gaussian mixtures, molecular benchmarks, model composition, and Boltzmann Generator sampling.
Main Findings
- Mode blindness is real and measurable. In Figure 1, Gaussian mixtures with weights 2/3–1/3 versus left-mode weights ranging in [0.2, 0.8] produce nearly identical scores at t₁ = 0.1, but the 3-class classification posterior probabilities (t₂ = 0.5, t₃ = 0.7) vary with the mixture weights. DiffCLF therefore sees information the score discards.
- DiffCLF matches DSM on sample quality while being far better calibrated. On the synthetic 40-mode Gaussian mixture (MOG-40) benchmark with a variance-preserving diffusion model, DiffCLF matches DSM in Fisher divergence and Maximum Mean Discrepancy (reported ×100) while keeping the classification loss near ~4.0–4.4 across dimensions 8, 16, 32, 64, and 128. DSM's classification loss degrades from 9.19 ± 0.33 at d = 8 to 383.53 ± 35.99 at d = 128.
- CtSM+DSM does not close the gap. The Conditional Time Score Matching baseline reaches classification losses of 6.80 ± 0.86 (d = 8) up to 20.86 ± 4.93 (d = 128), worse than DiffCLF throughout, and its MMD is much larger (e.g. 19.41 ± 0.77 versus 0.69 ± 0.59 at d = 8).
- DiffCLF learns log-densities that track the truth over time. In Figure 3 (a stochastic interpolant between a bi-modal and a 40-mode Gaussian mixture), only DiffCLF shows consistently high R² agreement with true log-densities across t ∈ (0, 1); DSM and DSM + CtSM do not.
- Downstream sampling improves. Figure 4 shows Langevin dynamics samples using the t = 0 energy on three benchmarks: Müller-Brown (MB), Alanine Dipeptide (ALDP), and Chignolin. For MB the paper shows a sample histogram; for ALDP, a torsion-angle histogram (φ, ψ); for Chignolin, a histogram of the first two TIC axes.
- Composition works. Figure 5 shows OR and AND model composition using 512-step Sequential Monte Carlo, comparing DiffCLF (orange) against DSM (purple) against ground truth (green).
- The cost is small. DSM needs two network evaluations per sampled time; the N-class DiffCLF needs N; the combined objective needs N + 1, i.e. N − 1 more than DSM, and only one extra evaluation in the binary case.
Methodology in Plain English
The authors work in a general noising framework where an observation Y_t equals a latent signal X_t plus Gaussian noise scaled by a schedule γ(t). Instead of only fitting the gradient of the log-density (the score), they parameterize an energy U_t^θ(y) plus a learnable bias F_t^θ that acts as a flexible log-normalizing constant. They then assign a class label to each noise level and ask a softmax classifier to tell which level a sample came from. The class-conditional densities are the model's own energies. The cross-entropy of this classifier is the DiffCLF loss, and because the softmax compares energies across levels, the loss depends on the values of the energies, not just their gradients — which is exactly what defeats mode blindness.
Since DiffCLF alone has a non-unique minimizer (any positive reweighting c(y) of the true density also minimizes it), the authors pair it with standard denoising score matching, and prove this combination pins down the true marginals uniquely. In practice F_t^θ is just a bias added to the last neural-network layer, so implementation is trivial. They then compare against DSM alone and against a Conditional Time Score Matching baseline under equal compute.
Why This Matters
Research impact. The paper gives a cheap, provably consistent route to learning full energy landscapes inside score-based generative models, and clarifies theoretically why Fokker–Planck and Bayes-rule regularizers do not fix mode blindness. It connects EBM training, noise contrastive estimation, and time-score matching into a single picture.
Real-world applications (as described or implied by the paper):
- Boltzmann Generators for molecular simulation, where the sequence of learned intermediate densities is used with annealed importance sampling, sequential Monte Carlo, or resampling to sample a target Boltzmann distribution.
- Compositional generation, combining models trained on different targets into mixtures or products via OR / AND composition without retraining.
- Free-energy difference estimation (thermodynamic integration, MBAR), where learned intermediate potentials replace hand-designed annealing paths.
- Molecular conformational sampling, as illustrated by the Müller-Brown, Alanine Dipeptide, and Chignolin Langevin-dynamics experiments.
Industry relevance. Accurate energies matter wherever a generative model must be reweighted, corrected, or combined with physics-based targets — drug discovery and molecular dynamics, materials science, and any pipeline that needs calibrated probabilistic outputs rather than just plausible samples. The low overhead (only N − 1 extra network evaluations versus DSM) makes adoption practical.
Future Directions
- Discrete and non-Euclidean data. The authors note in Appendix F that DiffCLF only requires the parameterized marginals, so it should extend to continuous-time Markov chains for discrete diffusion; this is sketched rather than fully benchmarked.
- Beyond cross-entropy. Appendix E generalizes the classification objective using Bergman divergences for learning density ratios, leaving a family of alternative losses largely unexplored.
- Scaling the number of classification levels N. The theory predicts the asymptotic covariance shrinks as N grows, but the practical trade-off between N, compute, and accuracy is only partially explored in the reported experiments.
- Integrating known marginals. The paper notes that whenever some marginals are known exactly (p_T in diffusion models, p_0 and p_1 in stochastic interpolants), they can be plugged into the framework; quantifying the benefit is left open.
Target Audience
Researchers and graduate students in generative modeling, statistical machine learning, and computational physics/chemistry who already understand score-based models and want energies they can actually use. Practitioners building Boltzmann Generators, compositional samplers, or free-energy estimators will benefit most, but the paper assumes comfort with SDEs, score matching, and asymptotic estimation theory.
Authors’ abstract
Score-based generative models have recently achieved remarkable success. While they are usually parameterized by the score, an alternative way is to use a series of time-dependent energy-based models (EBMs), where the score is obtained from the negative input-gradient of the energy. Crucially, EBMs can be leveraged not only for generation, but also for tasks such as compositional sampling or building Boltzmann Generators via Monte Carlo methods. However, training EBMs remains challenging. Direct maximum likelihood is computationally prohibitive due to the need for nested sampling, while score matching, though efficient, suffers from mode blindness. To address these issues, we introduce the Diffusive Classification (DiffCLF) objective, a simple method that avoids blindness while remaining computationally efficient. DiffCLF reframes EBM learning as a supervised classification problem across noise levels, and can be seamlessly combined with standard score-based objectives. We validate the effectiveness of DiffCLF by comparing the estimated energies against ground truth in analytical Gaussian mixture cases, and by applying the trained models to tasks such as model composition and Boltzmann Generator sampling. Our results show that DiffCLF enables EBMs with higher fidelity and broader applicability than existing approaches.