Skip to content
AI.info

Research

Infinite-dimensional generative diffusions via Doob's h-transform

Overview Research area: Generative modeling with diffusion models, specifically extending diffusion-based generative models to infinite-dimensional function spaces (stat.ML / machine learning theory).

Infinite-dimensional generative diffusions via Doob's h-transform
arXiv
2602.06621
Published
2026-02-06
Authors
Thorben Pieper-Sethmacher, Daniel Paulin

AI summary

Overview

Research area: Generative modeling with diffusion models, specifically extending diffusion-based generative models to infinite-dimensional function spaces (stat.ML / machine learning theory).

Technical level: Advanced. The paper relies on infinite-dimensional stochastic analysis (mild solutions of SDEs on Hilbert space, cylindrical Wiener processes, Cameron-Martin spaces, Doob's h-transform, Girsanov theorem, Wasserstein bounds).

Scope: The paper derives a rigorous framework for building generative diffusion models on an infinite-dimensional Hilbert space by steering a reference diffusion toward the target measure via Doob's h-transform, rather than by time-reversing a noising process, and validates it on a synthetic Gaussian mixture and on MNIST-SDF data.

What This Paper Is About

Standard diffusion models generate data by noising samples until they become approximately Gaussian, then reversing that process. This fails to behave well when data is really an infinite-dimensional object (functions, images at arbitrary resolution, physical fields) because the noising process may never converge, and reversing infinite-dimensional SDEs is hard to justify. The authors instead force a reference diffusion process to hit the target distribution at a chosen time using a change of measure called Doob's h-transform, which works for any time horizon and is shown to be well defined under explicit, checkable conditions.

Key Contributions

  1. A rigorous infinite-dimensional generative diffusion framework. The paper constructs generative diffusions on an infinite-dimensional Hilbert space via Doob's h-transform, deriving the forced process as a mild solution of an SDE with an additive steering term, under verifiable assumptions (including absolute continuity of the target with respect to a Gaussian reference measure and existence of transition densities).

  2. Removal of the time-reversal requirement. Because the construction forces the reference process to the target at any final time T > 0, the method avoids pathologies arising when a noising process does not converge to the prescribed reference measure, and it is well defined for arbitrary T. A Langevin initialisation step from the modified initial measure corrects for prior/data mismatch when T is small.

  3. A trainable score-matching objective with a correspondence theorem. Proposition 4.1 shows that minimising a specific loss is equivalent to minimising the Kullback-Leibler divergence between the path measure of the forced process and that of the parametrised approximation. Corollary 4.4 shows the minimiser recovers the forced process exactly when the parametrised class is rich enough.

  4. A new noising process (VP-SPDE) and error bounds. The Variance Preserving SPDE mimics the temporal noise scheduling of the VP-SDE while accounting for spatial structure via the covariance operator C, with parameters γ ∈ (0,1] and β(t). Proposition 6.1 gives a Wasserstein-2 bound decomposing error into an initialisation term, a training-loss term, and a numerical discretisation term.

Main Findings

  • The forced process exists for the VP-SPDE. Proposition 5.1 states the VP-SPDE has a unique mild solution satisfying the transition-density assumption, and Corollary 5.2 states the resulting h-function exists with a well-defined Fréchet derivative, so the h-transformed process is well defined.

  • Variance is preserved in the forced process. Remark 5.4 shows that if the target μ has covariance operator C, the marginals of the forced process have constant covariance operator C — this is the sense in which the VP-SPDE is "variance preserving."

  • The score-matching loss is equivalent to a KL minimisation. Proposition 4.1 establishes the if-and-only-if relation between minimising the loss and minimising D_KL of path measures; Corollary 4.4 shows exact recovery of the forced path measure when the parametrised score class contains the truth.

  • The error bound omits a term that grows with T. Remark 6.3 compares Proposition 6.1 with Theorem 14 of Pidstrigach et al. (2024): the new bound replaces a W₂(μ_data, N(0,C))·exp(−T/2 + L²T/4) term with ε_Init·exp(LT), which can be reduced by running more Langevin steps.

  • Discretise-last robustness on a Gaussian mixture. On μ = α N(u,C) + (1−α) N(−u,C) on L²([0,1]) with α = 0.1 and a Matérn covariance, the finite-dimensional noising-denoising model (ND) diverges as the approximation dimension D grows while the function-space models do not. At D = 500 with N = 5000 samples, averaged sliced Wasserstein distances to the target were: at T = 1, ND 0.274, HND 0.117, FCN (ours) 0.258, FNS (ours) 0.055; at T = 0.2, ND 0.692, HND 0.51, FCN (ours) 0.123, FNS (ours) 0.053.

  • Robustness to short noising time. The HND model collapsed at the limited noising time T = 0.2, while the forced models FCN and FNS stayed robust. The forced noise-scheduled model (FNS) consistently outperformed the forced constant-noise model (FCN) in this synthetic example.

  • Competitive but not best FID on MNIST-SDF. Trained at 64×64 resolution (upsampled from 32×32) with γ = 0.05 and a cosine noising schedule, the VP-SPDE's FID is described as competitive with Denoising-Diffusion-Operators (DDO) (Lim et al., 2025) but not fully matching its reported performance. The authors attribute this to limited hyperparameter tuning due to limited computational resources. The specific FID values from Table 2 are not included in the available content.

Methodology in Plain English

Instead of the usual "add noise, then reverse it" recipe, the authors start with a simple reference diffusion process driven by a Wiener process and ask: how do we tilt this process so that, at a chosen time T, its distribution equals the data distribution? The answer is Doob's h-transform, an exponential change of measure. This change of measure both adds a steering drift to the SDE and modifies the starting distribution, so samples must be drawn from a corrected initial measure (via an infinite-dimensional Langevin sampler) before simulating forward.

Because the function h — and hence the steering term s(t,x) = D_x log h(t,x) — is unknown, the authors fit it with a neural network by minimising a loss that closely resembles denoising score matching. When the reference diffusion is linear and the initial distribution is Gaussian or a Dirac measure, the required diffusion bridges are tractable Ornstein-Uhlenbeck bridges with closed-form Gaussian scores; for nonlinear drifts the authors point to guided proposals and importance sampling. A new reference process, the VP-SPDE, is designed so the noise is preconditioned by the covariance operator of the target, keeping the diffusion well defined in infinite dimensions; discretisation happens only at implementation time (in the spectral eigenbasis of C). Finally, they bound the Wasserstein-2 distance between the data measure and the sampled measure, separating initialisation error, training loss, and numerical discretisation error.

Why This Matters

Impact on research. The paper provides a mathematically rigorous alternative to time-reversal-based infinite-dimensional diffusion models, giving explicit verifiable conditions under which the construction is valid, plus quantitative bounds to the target measure. It also connects generative diffusion modeling to the long-standing "discretise last" paradigm from Bayesian inverse problems, and it removes the need for long-time convergence of a noising process.

Real-world applications:

  • Image generation at arbitrary resolution, where pixels are a discretisation of an underlying continuous image (the paper's theory cites image and video data as the smooth-data regime).
  • Geometric shape modeling, where shapes are represented as functions such as signed distance fields (demonstrated on MNIST-SDF).
  • Spatiotemporal physical process modeling, such as weather or climate fields.
  • Function-space measures in Bayesian inverse problems and data assimilation, where targets are distributions over functions.

Industry relevance. Models that behave consistently as resolution or state dimension increases are valuable wherever data is stored or generated at multiple resolutions, since they avoid the retraining and degradation that plague finite-dimensional models when discretisation is refined. The demonstrated MNIST-SDF results at 64×64 and comparison against DDO, GANO, and MultiDiff place the method directly alongside existing function-space generative models.

Future Directions

  • How to choose γ. Remark 5.3 notes that larger γ means rougher noise, that γ = 1 (white noise) causes high-frequency data information to be lost almost immediately, and that γ ≈ 0 performed better in experiments. The paper explicitly calls for further research on how γ should be chosen based on the structure of μ.

  • Nonlinear reference diffusions. Corollary 5.5 permits nonlinear F, but the authors restrict applications to the linear VP-SPDE because estimating the loss is more complex. Extending the empirical work to nonlinear drifts via guided proposals and importance sampling is left open.

  • Better hyperparameter tuning and larger-scale evaluation. The authors attribute their FID gap relative to DDO to limited tuning under limited computational resources, suggesting headroom on real data.

  • The smooth-data regime. Section 6.1 shows the method remains valid when μ is supported on the Cameron-Martin space after an arbitrarily small noising step, producing a regularised target μ_ε; how best to use this in practice is a natural follow-up.

Target Audience

Researchers and graduate students working on diffusion models, generative modeling in function spaces, and the theory of infinite-dimensional SDEs, as well as practitioners in scientific machine learning who need generative models that stay stable as resolution increases. The paper assumes familiarity with stochastic analysis on Hilbert spaces, so it is best suited to readers with an advanced mathematical background; those primarily interested in applied results can read the applications section and the sliced Wasserstein comparisons.

Authors’ abstract

This paper introduces a rigorous framework for defining generative diffusion models in infinite dimensions via Doob's h-transform. Rather than relying on time reversal of a noising process, a reference diffusion is forced towards the target distribution by an exponential change of measure. Compared to existing methodology, this approach readily generalises to the infinite-dimensional setting, hence offering greater flexibility in the diffusion model. The construction is derived rigorously under verifiable conditions, and bounds with respect to the target measure are established. We show that the forced process under the changed measure can be approximated by minimising a score-matching objective and validate our method on both synthetic and real data.

Read the original paper