Research
Latent Target Score Matching, with an application to Simulation-Based Inference
Latent Target Score Matching, with an application to Simulation-Based Inference Paper: arXiv:2602.07189v1 [cs.LG], published 2026-02-06, license CC BY 4.0 Authors: Joohwan Ko (University of Massachuse

- arXiv
- 2602.07189
- Published
- 2026-02-06
- Authors
- Joohwan Ko, Tomas Geffner
AI summary
Latent Target Score Matching, with an application to Simulation-Based InferencePaper: arXiv:2602.07189v1 [cs.LG], published 2026-02-06, license CC BY 4.0 Authors: Joohwan Ko (University of Massachusetts Amherst), Tomas Geffner (NVIDIA)
Overview
Research area: Generative modeling and simulation-based inference — specifically score-based diffusion model training objectives for models with latent variables.
Technical level: Advanced. The paper relies on stochastic differential equations, score matching identities, and Bayesian inference terminology, though the central idea can be followed with intermediate knowledge of diffusion models.
Scope in one sentence: The paper introduces a low-variance diffusion training target derived from joint scores over latent variables, and evaluates it on three simulation-based inference tasks.
What This Paper Is About
Diffusion models are typically trained with denoising score matching (DSM), which becomes high-variance as the diffusion noise level approaches zero, degrading accuracy. Target Score Matching (TSM) fixes this but requires access to the clean data score, which is unavailable when the model contains latent variables. The paper proposes Latent Target Score Matching (LTSM), which instead uses the joint score over the variables of interest and the latents to build a low-variance unbiased estimator of the marginal score, and combines it with DSM through a time-dependent mixture.
Key Contributions
-
A latent target score identity (Proposition 3.1). The paper shows that under the variance-preserving SDE (VP-SDE), which diffuses the variable of interest θ but keeps the latent z fixed, the marginal score equals (1/α(t)) times the conditional expectation of the joint score ∇_θ0 log p(θ0, z), where the expectation is taken over θ0 and z given θ_t.
-
The LTSM training objective (eq. 3.2). Using that identity, the paper defines a regression target y_LTSM(θ0, z, t) = (1/α(t)) ∇_θ0 log p(θ0, z), yielding a low-variance, unbiased training objective that requires only samples from the joint p(θ, z) and the ability to evaluate the joint score — not the clean marginal score needed by TSM.
-
A mixture objective with an analytic optimal weight (Proposition 3.2, eq. 3.3–3.5). Because LTSM's variance grows at larger noise levels while DSM's is stable there, the paper mixes the two targets with a time-dependent weight w_t and derives the variance-minimizing w_t* in closed form. In practice w_t = σ(MLP(t)) is learned jointly with the score network, avoiding expectation estimates.
-
An empirical study on simulation-based inference. The paper evaluates DSM, LTSM, and the mixture on three SBI simulators — a Gaussian model with exact scores, a Mixture of Categoricals, and a Generalized Galton Board — measuring regression-target variance, score error, and maximum mean discrepancy (MMD) of posterior samples.
Main Findings
-
Complementary variance profiles. Figure 1 (left) shows DSM's regression-target variance grows large as t → 0, while LTSM stays well-conditioned at low noise and grows only for larger t. The mixture maintains low variance across all noise levels.
-
Best score accuracy from the mixture. On the Gaussian task, where the true score can be computed analytically, the ℓ1 error between the learned score s_ψ(θ_t, t, x) and ∇ log p_t(θ_t | x) is lowest for the mixture (Figure 1, right), which uses LTSM's accuracy at small t and DSM's robustness at larger t.
-
Better posterior samples. MMD results in Figure 2 show the MIX method consistently producing better posterior samples than DSM across all three tasks and for several different potential observations x*, with the left column averaging over five observations x.
-
Largest gains at small simulator budgets. Improvements are most significant for smaller simulator budgets, indicating better sample efficiency — the paper frames this as the practically important regime because generating large datasets by calling simulators is expensive.
-
Learned weights track the analytic optimum. Appendix A.5 reports that the learned w_t = σ(MLP(t)) is low near t ≈ 0 and rises toward 1 as t → 1, closely following the Monte Carlo-estimated variance-optimal w_t* (Figures 4 and 5).
-
No numeric MMD or dataset-size values are reported in the provided text. The paper states the metrics and qualitative ordering but the truncated content does not include the specific MMD numbers, simulator-budget sizes, number of classes K, number of rows R, or network/training hyperparameters.
Methodology in Plain English
The authors consider settings where a system has variables of interest θ plus auxiliary latent variables z, and where the joint density p(θ, z) and its gradient with respect to θ are available, even though the marginal p(θ) is not. This is the "gray-box" simulator setting common in scientific inference, where the full joint p(θ, z, x) = p(θ) p(z | θ) p(x | θ, z) is tractable but the marginal likelihood p(x | θ) is not.
Their approach diffuses only θ under the VP-SDE and leaves z and x fixed. They train a conditional score network s_ψ(θ_t, t, x) to approximate ∇_θt log p_t(θ_t | x). Rather than regressing against the standard DSM target ∇_θt log p_t(θ_t | θ0) — which is noisy at low noise — they regress against the joint-score-derived target, obtained by forming the joint log-density and backpropagating to θ0 via automatic differentiation. Because this target is unbiased and well-behaved at small t but noisier at large t, they blend it with the DSM target using a weight that varies with diffusion time; the blend is still an unbiased estimator of the true score for any weight. The variance-optimal weight has a closed form in terms of second moments of the two targets, but in experiments they simply learn it as a sigmoid of a small MLP of t, trained end-to-end with no separate variance objective.
Evaluation proceeds in three stages: comparing conditional variance of the regression targets as a function of diffusion time; comparing learned scores against the exact analytic score on the Gaussian task; and comparing posterior sample quality via MMD with a Gaussian/RBF kernel, where the kernel bandwidth is chosen once per simulator and observation by the median heuristic on reference posterior samples and held fixed across methods and budgets. Reference posteriors for the non-Gaussian tasks come from rejection sampling over simulator draws. All methods share the same architecture, noise schedule, and training budgets.
Why This Matters
Impact on research. Training diffusion models in low-noise regimes is a known weak point, and this paper extends the low-variance toolkit from settings with known clean scores to the much broader class of latent-variable models. It also connects score matching directly to gray-box simulation-based inference, where joint information is available but marginal likelihoods are not.
Real-world applications (drawn from the domains the paper lists):
- Simulation-based inference in science, where expensive simulators produce observations and the goal is a posterior over parameters.
- Coarse-graining in structural biology, where some variables are modeled explicitly and others treated as latent.
- Backbone protein design, cited as a setting with latent or auxiliary components alongside the modeled variables.
- Inference on a few parameters of interest from mechanistic models, with all remaining structure treated as nuisance variables.
Industry relevance. Sample efficiency is framed as the key practical concern: simulators are expensive to call, so a training objective that reaches good posterior quality from fewer simulator calls has direct value. One author is affiliated with NVIDIA, suggesting relevance to accelerated computing and generative-model tooling.
Future Directions
The paper does not explicitly list future work, so the following are open questions the work raises:
- Scaling beyond the evaluated simulators. The experiments cover three simulators (Gaussian, Mixture of Categoricals, Generalized Galton Board) with diffusions over θ only; whether the same identities and variance profile hold for high-dimensional θ, long simulator rollouts, or diffusing additional variables.
- How the diffusion process should treat latents. LTSM is derived for a VP-SDE that diffuses θ and keeps z fixed. The consequences of instead diffusing z, or of using a different forward process or noise schedule, are not addressed in the provided text.
- Weight-learning alternatives. The paper derives the variance-optimal w_t* but learns w_t jointly with the score network in all main experiments; the comparative value of directly estimating w_t* versus learning it, and its sensitivity to the MLP parameterization, is left open.
- Extension beyond the SBI setting. The identity applies to any latent-variable model with an accessible joint score, but the paper's empirical evidence is confined to SBI tasks, so generality to other latent-variable generative modeling problems remains untested here.
Target Audience
This paper is most useful to machine learning researchers working on diffusion models and score matching, and to practitioners of simulation-based inference who have access to gray-box simulators exposing joint densities — particularly in the physical and life sciences. Readers need comfort with SDEs, score identities, and Bayesian posterior inference; those seeking a high-level conceptual introduction to diffusion models will find the formalism dense.
Authors’ abstract
Denoising score matching (DSM) for training diffusion models may suffer from high variance at low noise levels. Target Score Matching (TSM) mitigates this when clean data scores are available, providing a low-variance objective. In many applications clean scores are inaccessible due to the presence of latent variables, leaving only joint signals exposed. We propose Latent Target Score Matching (LTSM), an extension of TSM to leverage joint scores for low-variance supervision of the marginal score. While LTSM is effective at low noise levels, a mixture with DSM ensures robustness across noise scales. Across simulation-based inference tasks, LTSM consistently improves variance, score accuracy, and sample quality.