Research
Optimizing Input of Denoising Score Matching is Biased Towards Higher Score Norm
Overview Research area: Generative modeling — specifically the theory of denoising score matching (DSM) in diffusion models, and the practice of using the DSM loss to optimize things other than the sc

- arXiv
- 2511.11727
- Published
- 2025-11-13
- Authors
- Tongda Xu
AI summary
Overview
Research area: Generative modeling — specifically the theory of denoising score matching (DSM) in diffusion models, and the practice of using the DSM loss to optimize things other than the score network's own parameters.
Technical level: Advanced. The paper is a short theoretical workshop paper built around score-function algebra and the DSM/ESM equivalence proof.
Scope: A single-author workshop paper (arXiv:2511.11727v1, cs.LG, 13 Nov 2025, presented at the workshop "Frontiers in Probabilistic Inference: Sampling Meets Learning") that proves a bias term arises when DSM is used to optimize a diffusion model's conditional input or its data distribution, and catalogs the application areas affected.
What This Paper Is About
Denoising score matching is theoretically interchangeable with exact score matching, but only when the thing being optimized is the score network's parameters. Many recent systems instead use the DSM loss to optimize other quantities — the conditioning input fed to a diffusion model, or the data distribution itself. This paper shows that in those cases the equivalence no longer holds, and that the leftover term pushes the optimization toward distributions with a larger score norm.
Key Contributions
-
Identifies a broken equivalence. The paper shows that when the conditional input
cof a diffusion model is optimized alongside the network parametersθ, the classic DSM–ESM equivalence from Vincent (reference 8) fails, because terms that were previously constant now depend onc. -
Derives the explicit bias term (Theorem 2). Minimizing the DSM objective over both
θandcis shown to be equivalent to minimizing the ESM objective minus a termC₂ = E_{q(x_t|x)p(x|c)}[ ½ ‖∇_{x_t} log q(x_t|c)‖²]. Because that subtracted term is being minimized, its negation is being maximized — the optimization is biased toward higher score norm. -
Extends the result to the data distribution (Corollary 1). The same argument applies when a pre-trained diffusion model is used to optimize an input distribution
p(x), where the bias term becomesE_{q(x_t|x)p(x)}[ ½ ‖∇_{x_t} log q(x_t)‖²]. -
Maps the affected literature. The paper lists four domains with recent works it says are affected: auto-regressive generation (MAR, MetaQuery, and references 11–18), image compression (CDC, PerCo, FlowMo together with VQ-VAE), text-to-3D generation (DreamFusion, and references 22–23), and diffusion-based inverse problem solving (RED-Diff, and references 25–26). The author states this list is by no means exhaustive.
Main Findings
-
The DSM/ESM equivalence is conditional on what you optimize. Theorem 1 (attributed to reference 8) states that DSM and ESM are equivalent when optimizing
θ. The paper's Theorem 2 shows this breaks whenθand the conditioncare optimized jointly: the two differ by theC₂term. -
The
C₃term drops out because of a Markov chain. In the decomposition of the DSM loss into ESM plus residual terms, the termC₃ = E_{q(x_t|x)p(x|c)}[ ½ ‖∇_{x_t} log q(x_t|x,c)‖²]reduces to a form that does not depend onθorc, becausex_t → x → cforms a Markov chain. It therefore has no gradient with respect toc. -
The surviving bias favors larger score norm. Since minimizing the DSM loss over
camounts to minimizing ESM while maximizingC₂ = E_{q(x_t|x)p(x|c)}[ ½ ‖∇_{x_t} log q(x_t|c)‖²], the paper's central claim is that the optimization is pulled toward configurations with a higher score norm. -
The same bias appears when optimizing a data distribution. Corollary 1 states that when a pre-trained diffusion model is used to optimize
p(x), the DSM loss equals the ESM loss minusE_{q(x_t|x)p(x)}[ ½ ‖∇_{x_t} log q(x_t)‖²]— structurally the same bias as in the conditional case. -
No empirical evaluation is reported. The paper is theoretical, with the main result proved in Appendix A. No datasets, benchmarks, model architectures, training runs, or quantitative measurements are reported in the content provided.
Methodology in Plain English
The argument proceeds analytically rather than experimentally.
First, the paper writes out the ESM loss (matching the network's output to the true score of the noisy distribution) and notes it is intractable, which is why DSM (matching the network to the score of the simple Gaussian noise kernel) is used instead. It recalls the known result that the two are equivalent when only the network parameters are being learned.
Then it expands the DSM objective algebraically and separates it into pieces. Two of those pieces do not involve the network parameters at all. Under the classical setting they are constants that vanish from the optimization. But the paper points out that one of them — the expected squared norm of the marginal score ∇ log q(x_t|c) — does depend on the conditioning variable c. So the moment you start optimizing c, that previously harmless constant becomes an active objective term with the wrong sign.
The proof in Appendix A makes the argument rigorous by applying the log-derivative trick twice to convert a term involving the marginal score into a term involving the tractable conditional score, then substituting back and comparing the two losses side by side. The final step uses the Markov structure x_t → x → c to show the remaining residual term is gradient-free with respect to c. A one-line parallel argument yields the data-distribution corollary.
The last section is a literature-mapping exercise: identifying the pattern (DSM used to optimize c or p(x) with a pre-trained or jointly trained model) and listing the works that match it.
Why This Matters
Impact on research: A large family of recent methods treats the diffusion/DSM loss as a generic differentiable objective for learning representations, conditions, or even the data itself. This paper says that when used that way, the objective is not the score-matching objective it appears to be — it carries an extra, one-sided pressure toward high score norm. That is a theoretical caveat that applies to an entire methodological pattern rather than to one system.
Real-world applications named in the paper as affected:
- Auto-regressive image generation without tokenizers: MAR uses DSM to optimize
c, the output of the auto-regressive model; MetaQuery uses diffusion loss to optimize the connector output in MLLM models. - Image compression: CDC jointly minimizes the bitrate of
calongside the diffusion loss; PerCo optimizescwith diffusion loss and uses the resulting gradient to refine a VQ-VAE; FlowMo uses that gradient to jointly train a VQ-VAE. - Text-to-3D generation: DreamFusion uses a pre-trained 2D diffusion model to optimize an input distribution
p(x), then uses the gradient to optimize a 3D representation that renders it. - Inverse problems: RED-Diff adopts the DreamFusion loss to optimize
p(x), solving tasks such as super-resolution and de-blurring.
Industry relevance: These application areas span generative content creation, learned compression codecs, 3D asset generation, and image restoration — all commercially active. If the bias is real, practitioners in these areas may be optimizing something subtly different from what their loss function suggests, and might benefit from correcting or re-weighting the objective.
Future Directions
-
Quantify the practical severity. The paper establishes the bias theoretically but reports no experiments. Whether the bias materially degrades outputs — or is negligible in practice — is left open.
-
Design a debiased objective. Since the offending
C₂term is written explicitly, a natural follow-up is a corrected loss that subtracts or neutralizes it, and a study of whether that improves the affected methods. -
Connect score norm to observable failure modes. The paper's claim is that optimization pushes toward higher score norm. What that means visually or perceptually — over-sharpening, mode bias, artifacts, or something else — is not addressed and would be a natural empirical question.
-
Widen the audit. The affected-works list is stated to be non-exhaustive and covers four domains. A systematic survey of which diffusion-loss-based methods are and are not subject to the bias, perhaps with a common diagnostic, would clarify the scope.
Target Audience
Researchers and engineers working on diffusion models, score-based generative modeling, or any pipeline that repurposes a diffusion/denoising loss as a training signal for something other than the score network — for example, learned tokenizers, auto-regressive generators, neural compression codecs, text-to-3D systems, and plug-and-play inverse-problem solvers. It is most useful to readers comfortable with score functions, score matching, and the DSM/ESM equivalence; readers without that background will need the surrounding literature (Vincent, reference 8, in particular) to follow the derivation.
Authors’ abstract
Many recent works utilize denoising score matching to optimize the conditional input of diffusion models. In this workshop paper, we demonstrate that such optimization breaks the equivalence between denoising score matching and exact score matching. Furthermore, we show that this bias leads to higher score norm. Additionally, we observe a similar bias when optimizing the data distribution using a pre-trained diffusion model. Finally, we discuss the wide range of works across different domains that are affected by this bias, including MAR for auto-regressive generation, PerCo for image compression, and DreamFusion for text to 3D generation.