Research
A Connection Between Score Matching and Local Intrinsic Dimension
Overview Research area: Machine learning — generative modeling (diffusion and flow matching), with connections to signal processing, information theory, and geometric data analysis. Technical level: I

- arXiv
- 2510.12975
- Published
- 2025-10-14
- Authors
- Eric Yeats, Aaron Jacobson, Darryl Hannan, Yiran Jia, Timothy Doster, Henry Kvinge, Scott Mahan
AI summary
Overview
- Research area: Machine learning — generative modeling (diffusion and flow matching), with connections to signal processing, information theory, and geometric data analysis.
- Technical level: Intermediate. The experimental results are accessible, but the paper's core claims rest on measure-theoretic proofs (in the appendix) about how score-matching losses behave near a data manifold.
- Scope: The paper proves that the denoising score matching loss is lower bounded by the local intrinsic dimension (LID) of the data manifold, connects that result to the implicit score matching loss and to existing estimators such as FLIPD and the normal bundle method, and tests the resulting estimator on synthetic manifolds and on Stable Diffusion 3.5 and Stable Diffusion 2.
What This Paper Is About
High-dimensional data produced by natural processes usually lies on a much lower-dimensional structure, and the number of dimensions needed to locally describe that structure is the LID. Recent research showed that diffusion models can recover LID, but the leading methods need many forward passes through the model or require gradient computation, which makes them expensive in compute- and memory-constrained settings. The authors ask whether the denoising score matching loss — a quantity a model already computes during training — can itself serve as a fast, scalable LID estimator.
Key Contributions
-
A lower bound from the denoising loss. The authors prove (Theorem 3.1) that, for a sufficiently small noise level, the expected denoising score matching loss is greater than or equal to the intrinsic dimension d of the manifold, so the loss can be read as an estimate of LID. They extend this to stratified manifolds, where the loss is lower bounded by the expected dimension across submanifolds.
-
A bridge to implicit score matching and FLIPD. The authors prove (Theorem 3.3) that the implicit score matching loss is lower bounded by the negative normal dimension, −(n − d), and then show that expected FLIPD is also lower bounded by the LID because of its close relationship to that loss.
-
A bridge to the normal bundle estimator. They define an "error bundle" method analogous to the normal bundle (NB) method but built from error vectors, and show that the trace of its scaled Gram matrix equals the denoising score matching loss at a point.
-
Experiments on manifolds and real generative models. They benchmark the denoising loss against FLIPD and three non-parametric methods on known-LID manifolds, and test it with Stable Diffusion 3.5 medium and Stable Diffusion 2, reporting accuracy, peak GPU memory, and behavior under quantization.
Main Findings
-
The denoising loss beat FLIPD in average accuracy across architectures and noise levels. On the manifold benchmark, the DiT-based denoising loss reached average MAE of 7.00 (σ = 0.01), 4.48 (σ = 0.02), and 3.22 (σ = 0.05), while DiT-based FLIPD reached 10.16, 7.00, and 4.32. With an MLP, the denoising loss averaged 9.01, 8.22, and 7.48 versus 10.80, 11.22, and 8.40 for FLIPD.
-
Parametric methods with the DiT architecture outperformed non-parametric ones. Of the non-parametric estimators, ESS performed best, with average MAE of 7.84 and 7.12 for k = 50 and k = 100. MLE averaged 20.18 and 22.01, and TwoNN averaged 21.71 and 20.08.
-
The gap was largest on highly curved manifolds. On the 32-dimensional "Nonlinear" manifold, the DiT denoising loss recorded MAE of 12.47, 7.54, and 2.15 across the three noise levels, while FLIPD recorded 26.54, 20.44, and 12.87. The authors hypothesize manifold curvature affects the divergence term in FLIPD. On the 16-HyperSphere (d = 16, n = 64), the pattern reversed at some noise levels: FLIPD reached 0.70 at σ = 0.05 against 2.58 for the denoising loss.
-
Memory scaled far better for the denoising loss than for FLIPD. In the hypersphere scaling experiment, where ambient dimension is twice the true LID, peak GPU memory for FLIPD grew rapidly because of its reliance on gradient computation, while the denoising loss grew slowly. ESS failed to scale, reaching MAE of approximately 25 when n = 256 and d = 128.
-
The denoising loss is more accurate and cheaper when the normal bundle / error bundle method runs out of samples. Figure 1(b) illustrates a 128-dimensional manifold in a 256-dimensional ambient space, where the denoising loss remains accurate at small sample counts such as m = 8, whereas the error bundle (and by extension the normal bundle) method needs at least as many samples as the LID (respectively, the normal dimension).
-
On Stable Diffusion 3.5 medium, the two estimators were highly correlated, with FLIPD systematically higher. Using 500 256×256 images of "a photo of a cat" generated with 28 sampling steps and a guidance level of 3.5, the authors found FLIPD estimates higher on average than denoising loss estimates at each noise scale, giving a line of best fit with slope below 1 in the scatter plot. At low noise scales the data occupied most of the 16384 dimensions of the latent space — still under 8.4% of the ambient dimension of the images — and at high noise scales the scaled data appeared as a 0-dimensional point.
-
Under quantization, the denoising loss degraded less than FLIPD. Averaging across all 500 images and comparing float16 and bfloat16 against float32, the change in LID estimates was smaller for the denoising loss than for FLIPD, which the authors attribute to accumulated error in the gradient computation. Float16 produced lower MAE from float32 than bfloat16, suggesting LID estimation with SD-3.5 benefits more from higher precision than higher dynamic range. Peak GPU memory for the denoising loss was roughly 60% of FLIPD's on batches of 10 images.
-
On Stable Diffusion 2, FLIPD produced invalid negative LID estimates at higher noise levels. SD2 is a latent diffusion model with a U-Net architecture, and the authors attribute the negative values to the tendency of such networks to parameterize functions with high Lipschitz constants. They hypothesize this did not occur with SD3.5 because of the flow matching noise parameterization ε_θ(x, σ) := (1 − σ)v_θ(x, σ) + x.
-
The constants in the loss equivalences have a geometric interpretation. The authors argue that C_DSM is the negative average LID of the data and C_ISM is the average normal dimension, since the minimum of the explicit score matching loss is 0 and training amortizes the pointwise loss over a dataset.
Methodology in Plain English
The authors start from a standard setup in which a score model is trained at a single, sufficiently small noise level, so that the density on the manifold looks locally constant and the manifold's curvature is negligible over the region where the Gaussian perturbations have meaningful probability mass. They decompose the denoising error into components along the manifold's tangent space and its normal space at each point. Noise along the normal directions is essentially unlearnable and contributes roughly 1 per dimension, while noise along the tangent directions is predictable and contributes roughly 0, so summing the expected squared error recovers the number of tangent dimensions — the LID. The same decomposition, applied to the implicit loss, yields a bound involving the normal dimension.
To connect this to prior work, they rewrite FLIPD in terms of the implicit score matching loss plus a score-norm term plus the ambient dimension n, which lets them transfer the implicit bound to FLIPD. They also define an error bundle method using error vectors rather than noise predictions, and show its Gram matrix trace is exactly the denoising loss.
For experiments, they pull manifolds of known LID from the scikit-dimension package and measure mean absolute error against the true LID over 2000 data points. Each manifold is represented by 2000 uniformly sampled points, and models are trained for 50000 batches of size 100 with a cosine annealed learning rate schedule. They train both a diffusion transformer (patch size 4, hidden dimension 128, 16 attention heads, 8 layers) and an MLP with skip connections using a flow matching objective, convert flow predictions to noise predictions, and evaluate at σ = 0.01, 0.02, and 0.05, scaling data by (1 − σ) to match the training schedule. The denoising loss uses 8 Gaussian noise samples and FLIPD uses 8 Rademacher samples for the divergence estimate. Everything runs in PyTorch on a single NVIDIA H100 80GB GPU. For the Clifford torus, dimensions are randomly permuted so the patch-based DiT must use attention to learn the structure, and all manifold data is centered and scaled by σ_A⁻¹, where σ_A is the maximum per-feature standard deviation.
For the image experiments, they apply the same two estimators to SD-3.5 medium from the diffusers library and to SD2, estimating LID in latent space with the models' own noise parameterizations.
Why This Matters
The work reframes a loss that generative models already compute during training as a practical measurement of data geometry, which removes the need for many forward passes or any gradient computation at estimation time. For research, it ties three previously separate lines together — denoising score matching, implicit score matching, and the leading parametric estimators FLIPD and the normal bundle method — and gives the constants in the score matching equivalences a concrete geometric meaning. It also offers a hypothesis about connecting likelihood attribution in diffusion and flow models to a learned normal dimension along the ODE solution.
Real-world applications the paper's framing touches on:
- Anomaly detection, where the paper notes LID has been leveraged in engineering practice.
- Clustering and segmentation, also cited as existing uses of LID.
- Compression, since LID determines bounds on how locally compressible a distribution is.
- Statistical efficiency in deep learning, where lower-dimensional structure is argued to make learning and generalization easier.
Industry relevance centers on cost. Diffusion and flow models are increasingly deployed at scale, and an LID estimator that avoids gradient computation and holds up under half precision occupies a fraction of the memory — roughly 60% of FLIPD's peak in these experiments — while showing less degradation when a model is quantized for deployment.
Future Directions
-
Removing the experimental constraints the authors list. Their limitations section notes experiments used only one H100 80GB GPU without distributed computation for larger batches, and that they did not quantize below half precision. Both leave room for testing whether the memory advantage grows with batch size or lower precision.
-
Applying "knee search" to LID curves. The authors report average statistics over a few hyperparameters rather than performing the knee search on LID curves used in prior work such as FLIPD.
-
Testing the curvature hypothesis. The authors hypothesize that FLIPD's larger error on highly curved manifolds comes from the effect of curvature on its divergence term, and that SD2's negative estimates come from high Lipschitz constants in the U-Net. Neither hypothesis is verified in the paper.
-
Connecting likelihood to normal dimension at each ODE step. The authors hypothesize that higher likelihood attribution may be associated with a higher learned normal dimension at each point along the ODE solution used for likelihood computation, an idea they raise but do not test.
Target Audience
Researchers and practitioners working on diffusion and flow-based generative models, intrinsic dimension estimation, or manifold learning; readers interested in the theory connecting score matching objectives to data geometry; and engineers who need to estimate data dimensionality under tight memory or compute budgets, including those deploying quantized models. A background in probability, score matching, and basic differential geometry will help with the theoretical sections.
Authors’ abstract
The local intrinsic dimension (LID) of data is a fundamental quantity in signal processing and learning theory, but quantifying the LID of high-dimensional, complex data has been a historically challenging task. Recent works have discovered that diffusion models capture the LID of data through the spectra of their score estimates and through the rate of change of their density estimates under various noise perturbations. While these methods can accurately quantify LID, they require either many forward passes of the diffusion model or use of gradient computation, limiting their applicability in compute- and memory-constrained scenarios. We show that the LID is a lower bound on the denoising score matching loss, motivating use of the denoising score matching loss as a LID estimator. Moreover, we show that the equivalent implicit score matching loss also approximates LID via the normal dimension and is closely related to a recent LID estimator, FLIPD. Our experiments on a manifold benchmark and with Stable Diffusion 3.5 indicate that the denoising score matching loss is a highly competitive and scalable LID estimator, achieving superior accuracy and memory footprint under increasing problem size and quantization level.