Research
One Inverse Step is a Convex Program: Bayes-Limit Calibration of Diffusion Inversion
Overview Research area: Statistical machine learning, specifically the theory of diffusion models — the geometry of DDIM inversion, score-function Jacobians, and the calibration of geometry probes run
- arXiv
- 2608.23094
- Published
- 2026-08-24
- Authors
- Gordei Verbii
AI summary
Overview
- Research area: Statistical machine learning, specifically the theory of diffusion models — the geometry of DDIM inversion, score-function Jacobians, and the calibration of geometry probes run on pretrained diffusion models.
- Technical level: Advanced. The paper is written in the language of convex analysis, Riemannian/differential geometry (reach, focal radius, shape operator, Fermi coordinates) and stochastic processes, and assumes familiarity with DDPM/DDIM, Tweedie's formula and Picard iteration.
- Scope in one sentence: The paper shows that a single implicit DDIM inversion step is the stationarity condition of an explicitly constructed strongly convex potential whose modulus is fixed by the noise schedule alone, derives closed-form Bayes-limit values for every readout of that step, and then uses those values as a calibration curve plus a set of hypothesis-free falsification tests on both exact and trained denoisers.
What This Paper Is About
A widely used way to ask whether a pretrained diffusion model has learned the local geometry of its data is to take one implicit DDIM inversion step, differentiate the network once, and read geometry off the Jacobian. This paper asks what that probe is actually measuring, and answers that in the Bayes limit every quantity it reports has a closed-form value determined by the manifold dimension, ambient dimension, noise level and mean curvature. Turning the probe from an estimator into a calibrated instrument then makes its failures attributable: the paper separates the solution of the step, the solver's orbit, and the convergence domain, showing the geometry lives only in the third.
Key Contributions
-
One inverse step is a convex program (Theorem 1). The inversion residual is a gradient field, F(x) = x − G(x) = ∇Ψ_t(x), with ∇²Ψ_t = e^{−h_t} I + ρ_g* C_{t+1}/σ_{t+1}² ⪰ e^{−h_t} I. The solution is therefore unique for every data law and schedule, Picard iteration is unit-step gradient descent on Ψ_t, and the condition number is e^{h_t} — independent of the data, of d, of D and of curvature.
-
A contraction theorem with a schedule-only constant (Theorem 4). Under Assumptions 1 and 2, ρ_g ≤ ρ_g* = |B_t|/σ_{t+1} = 1 − e^{−h_t} < 1 for every schedule with ᾱ_T > 0, with the identity ρ_g = ρ_g* 𝖲_{t+1}. The constant itself is unconditional and stays below 0.326 on the standard DDPM schedule.
-
The convergence domain is a curvature (Theorem 5). Both Picard stability boundaries are fixed fractions of the manifold's focal radius: the oscillation shell at w_osc = 1/2 (schedule-free) and the divergence shell at w_div = 1/(1 + ρ_g*), with a measured finite-noise correction ½ σ̃_{t+1}² ‖II‖² that is free of the data law, density, dimension and codimension.
-
A Bayes-limit calibration table and three hypothesis-free falsifications (§5.1–5.2). Every readout of the one-step probe is given an exact value, verified on three integrable manifolds to machine precision in codimension 1 and 2 (Table 1), including two corrections to values assumed in the literature: the Bayes residual at the noiseless point is (σ_t/2)‖H‖/√ᾱ_t rather than 0, and the tangential scaling exponent is −1 rather than 0. Three tests that follow from C_t ⪰ 0 alone — σ_t λ_max(sym J) ≤ 1, skew Q = 0, and σ_t tr J ≤ D — are run with an oracle control that separates a broken estimator from a broken model.
Main Findings
-
Uniqueness is unconditional, but multiplicity is a certificate of model error. At the Bayes limit the fixed point of one inversion step is unique for every data law and every schedule with ᾱ_T > 0. A trained score admits a second solution only if σ_{t+1} λ_max(sym J(ξ, t+1)) ≥ 1/ρ_g* = 1/(1 − e^{−h_t}) for some ξ on the segment. Roots born at such a fold carry multiplier ρ_g* σ_{t+1} λ_max > 1, so they are repelling and invisible to Picard. On a two-atom Bayes law with amplitude-scaled score ε_φ = γ ε*, the root count moves from 1 to 3 exactly where predicted, bracketed at ρ_g* γ ∈ (0.996, 1.050] — a threshold verified in an analytic model, and no trained score in the study was tested against it.
-
The solver can fail even where the solution is unique. Because Picard's step size is pinned at 1 by the algebra of the map rather than chosen, Picard is unstable wherever λ_max(∇²Ψ_t) > 2; damping below 2/λ_max cures it, with optimum η* = 2/(μ + Λ_t). Oscillation therefore certifies nothing about the model. Damping with η < 2/Λ_t converges at every ρ_g, at rate tanh(h_t/2) when optimal inside the tube.
-
Contraction is a schedule constant, not a data property. Evaluating ρ_g* = 1 − e^{−h_t} at every step gives max_t ρ_g* = 0.326, attained at the first step, on the standard DDPM schedule (T = 1000, β linear on [10⁻⁴, 0.02]). The coarse T = 100, β_end = 0.05 schedule reads 0.623, not the 1.65 of |B_t|/σ_t = A_t c, which belongs to the explicit map and has no fixed point to contract to. The other five families satisfy the identity to 9.4 × 10⁻¹⁶ and stay strictly below 1 at strides 1 to 100, with ρ_g* spanning 0.024 to 0.971.
-
Geometry lives in the convergence domain, and exact scores reproduce both shells. On the scale-free depth w = r κ_max, the oscillation shell sits at w = 1/2 independent of the schedule, and w_div = 1/(1 + ρ_g*) — on DDPM-linear at strides 1/10/20/50/100, w_div = 0.754/0.560/0.534/0.515/0.507 while w_osc = 1/2 at every one. Exact scores reproduce both boundaries to within 0.54% on three classes. The displacement must be inward: an outward ray reads λ_max(C_t)/σ_t² = 1/(1 + w), crosses no threshold and returns a confident null — on the exact score it resolves both branches at all 18 depths and reports no shell, while its inward twin reads a multiplier 8.66× larger.
-
No trained score shows a shell — a derived limitation, not a null result. The Fermi window conflicts with the model's own training support by 3.6–5.6 times, and the trained Hessian-Lipschitz constant is 2–12% of the curvature the law reads, 0 on a ReLU net. The Bayes reference is valid only inside a two-sided window (g_t ≤ 0.1 and σ̃_t ≤ ½ reach(M)), which voids 34 of 55 exact-score settings.
-
The unconditional ceiling is violated by every trained image model. From Cov(x_0 | x_t) ⪰ 0 alone, σ_t λ_max(sym J) ≤ 1 holds for the exact score to 3 × 10⁻⁷ — the largest of 55 exactly computed cloud spectra being 1.0000003 — but is crossed in 10 of 110 cloud settings up to 1.1323, in 36 of 36 CIFAR-10 settings by 1.2586 to 2.4119, and in 9 of 9 CelebA-HQ-256 settings by 1.6146 to 4.6637. Since a shifted power iteration returns ‖Av‖ on a positive-semidefinite operator, every image value is a certified lower bound on its own excess. The arm of the test that computes exactly 110 cloud settings crosses in 10, and of 45 image settings, 45 are crossed.
-
𝖲_t is not readable as a stiffness on images. The skew index is exactly 0 at Bayes (J* is a posterior covariance) but puts up to 43% of s_1(J) into the antisymmetric part on CIFAR-10, against 2 × 10⁻⁶ on the exact score. The signed pair (σ_t λ_max(sym J), −σ_t λ_min(sym J)) is a necessity, not an improvement: on the von Mises circle (ρ_0 ∝ e^{23 cos θ}, σ_t = 0.2 < reach(M) = 1) the Bayes value at the antipode is s_1(J*) = 48.6 against 1/σ_t = 5, and 𝖲_t = 9.7 while the ceiling still holds.
-
Where 𝖲_t is readable, it explains the measured contraction. At the smallest guard-clean cloud setting (σ_t = 0.0426) the exact score reads 𝖲_t/𝖲_t* within 8 × 10⁻⁸ of 1 on the two closed, uniformly sampled classes, while the epoch-300 network reads 0.050–0.055 and the epoch-50 one 0.024–0.034 — an 18 to 42× deficit against a reference computed in the same harness. At the CIFAR probe ρ_g* = 0.1180 and the measured Picard rate is 0.1409–0.1498, one-sided high exactly as predicted; at stride 20 Picard measurably diverges where the Bayes map contracts at 0.7134.
-
The calibration table resolves, not just reports. All ratios in Table 1 read 1.0000, so the evidence that the table resolves anything is its last row: over a factor 4 in κ and a decade in σ_t, the two ratios leave 1 at the predicted orders 2 and 4 (2.004–2.013 and 4.006–4.017). The identities μ(∇²Ψ_t)e^{h_t} and ρ_g*/(1 − e^{−h_t}) hold to 1 ± 2·10⁻¹⁶ and 1 ± 9.4·10⁻¹⁶; 2w_osc and w_div(1 + ρ_g*) hold to 1 − 3.1·10⁻¹⁰ and 1 + 1.7·10⁻⁵. The Kantorovich constant a/e^{h_t} = 0.9947 on ten exact-score settings.
-
Flattened dimension readouts measure the wrong target. A point-cloud sample of n points in ℝ³ flattened to D = 3n has support dimension nd ∈ {128, 256}, not d ∈ {1, 2}, so comparisons against d are mis-specified. At the prescribed K = 4D = 1536, no rank readout on a trained model survives: the largest consecutive gap ratio is 1.02–1.58 over all ten trained settings against a pure-noise null of 1.022 at the same (K, D), and every one abstains — including the sphere, whose codimension-128 gap ratio of 5.06 at K = 1.08D falls to 1.16 at index 373.
-
The oracle control makes every negative attributable. Replacing ε_φ with the exact score under the identical harness, probe cloud, flattening and budget at σ_t = 0.0426 gives d̂(x°) = 255.99999, 255.99999, 256.99, 246.32, 128.16 on the sphere, two spheres, torus, swiss roll and trefoil against flattened targets 256, 256, 256, 256, 128; the rank statistic recovers the true codimension on every closed class and the exponent split on four of five. The epoch-300 model reads 382–383 on all five classes and the epoch-50 one 391–410, i.e. the ambient D. Both oracle misses are the swiss roll's free boundary.
-
A negative result about a positive result. The descent-based estimator of the Łojasiewicz exponent returns θ = 1/2 with R² = 1 (measured R² = 1.000000) whenever one scale governs the fit window, in particular under the divergent step the published protocol implies. It appears on all three manifolds regardless of d or topology, which already shows the exponent cannot discriminate geometry; the paper gives the field-based estimator that does carry information and measures it on CelebA-HQ-256.
-
Corrections and scope clauses. The sampled-point readout d̂ has expectation D − ½ δ_t for every data law and every model, so at Bayes it is exactly D — the FLIPD bound 𝔼[FLIPD] ≥ d is not merely loose, its exact value is D. Only at the noiseless point x° does it recover the manifold, and there with a signed curvature correction. Assumption 2 (λ_max(C_t(ξ)) ≤ 2σ_t²) is checkable: inside the tube of a uniform M the exact Fermi value is λ_max(C_t) = σ_t²/(1 − r κ_max), so it holds iff r ≤ 1/(2κ_max), half the focal radius, with margin 2 for every log-concave law, and is conservative by 2/(1 + ρ_g*) = 1.51 at the first step of the standard DDPM schedule. The deficit in 𝖲_t vanishes automatically in high codimension — at D = 3072 for every d ≤ 76 — so it is measurable only in low codimension.
-
Solver-independence is measured. Over 70 settings Picard, Anderson and Newton–Krylov agree to 1.7 × 10⁻⁴ of the local noise scale and every readout is solver
Authors’ abstract
One implicit DDIM inversion step is the cheapest probe of whether a pretrained diffusion model encodes local manifold geometry. It is the stationarity condition of an explicit potential, $x-G(x)=\nablaΨ_t(x)$, strongly convex at the Bayes limit with modulus exactly $e^{-h_t}$ for the step's log-SNR gap $h_t$ $-$ for every data law, schedule and point, with no manifold, reach or unimodality hypothesis. Three consequences must be kept apart. (i) The solution is unique at the Bayes limit; a second one requires the trained score to violate the posterior-covariance bound by $1/(1-e^{-h_t})$, a hypothesis-free certificate of model error; the same bound makes contraction a schedule constant, $ρ_g^{\star}=1-e^{-h_t}<0.326$ throughout the standard DDPM schedule. (ii) The solver can still fail: Picard iteration is unit-step gradient descent on $Ψ_t$, unstable wherever $λ_{\max}(\nabla^2Ψ_t)>2$, so oscillation certifies nothing; damping below $2/λ_{\max}$ cures it. (iii) The geometry lives in the convergence domain: on the scale-free depth $w=rκ_{\max}$ the oscillation shell sits at $w=\tfrac12$, schedule-free, and the divergence shell at $w=1/(1+ρ_g^{\star})$, with a measured finite-noise correction in $\|\mathrm{II}\|^2$. Exact scores reproduce both to within $0.54\%$ on three classes; no trained score we probe shows a shell $-$ a derived limitation, not a null result: the Fermi window conflicts with the model's own training support by $3.6$-$5.6\times$, and the trained Hessian-Lipschitz constant is $2$-$12\%$ of the curvature the law reads, $0$ on a ReLU net. Finally the unconditional ceiling $σ_tλ_{\max}(\mathrm{sym}\,J)\le1$, from $\mathrm{Cov}(x_0\mid x_t)\succeq0$ alone, holds for the exact score to $3\times10^{-7}$ but is violated in all DDPM CIFAR-10/CelebA-HQ-256 settings, by $1.26$-$4.66\times$.