Skip to content
AI.info

Research

Deceptron: Learned Local Inverses for Fast and Stable Physics Inversion

Deceptron: Learned Local Inverses for Fast and Stable Physics Inversion Overview Research area: Machine learning for physical inverse problems (PDE inversion, system identification, imaging), intersec

Deceptron: Learned Local Inverses for Fast and Stable Physics Inversion
arXiv
2511.21076
Published
2025-11-26
Authors
Aaditya L. Kachhadiya

AI summary

Deceptron: Learned Local Inverses for Fast and Stable Physics Inversion

Overview

Research area: Machine learning for physical inverse problems (PDE inversion, system identification, imaging), intersecting learned preconditioning, automatic differentiation, and classical numerical optimization.

Technical level: Advanced. The paper assumes familiarity with Jacobians, JVP/VJP products, Gauss–Newton and Levenberg–Marquardt methods, Armijo backtracking line search, and projection onto constraint sets.

Scope: The paper introduces a lightweight bidirectional module (the Deceptron) that learns a local inverse of a differentiable forward surrogate, uses it to precondition inverse updates (D-IPG), and previews a 2D unrolled variant (DeceptronNet v0), evaluated on Heat-1D initial-condition recovery, a damped oscillator inverse problem, a 2D PSF task, and Kodak24.

What This Paper Is About

Many inverse problems in the physical sciences require recovering unknown inputs from indirect measurements, and the standard approach — minimizing a data misfit in input space — is often ill-conditioned, so gradients are poorly scaled and progress depends heavily on step size. The paper's goal is to learn a local inverse of a differentiable forward surrogate and use that learned inverse to pull output-space descent steps back into input space, making the iteration faster and more stable than projected gradient while remaining competitive with Gauss–Newton. A secondary goal is to preview a compact unrolled 2D corrector, DeceptronNet v0, that performs few-step image corrections.

Key Contributions

  1. The Deceptron module and Jacobian Composition Penalty (JCP). A bidirectional parameterization with a forward map f_W(x) = σ(Wx + b) and a reverse map g_V(y) = σ̃(Vy + c), trained with a composite loss that includes a supervised fit, forward–reverse consistency, a cycle term, a spectral penalty ‖WᵀW − I‖_F², a soft bias tie ‖b + c‖₂², an optional ‖VW − I‖_F² term, and a JCP term that encourages J_g(f(x)) J_f(x) ≈ I via JVP/VJP probes. V and Wᵀ are explicitly not tied.

  2. D-IPG (Deceptron Inverse-Preconditioned Gradient). A solver that takes a descent step in output space (y_{t+1}^prop = y_t − α r_t), pulls it back through g_V, and applies projection under the same Armijo backtracking and stopping rules as the baselines, changing only the update direction.

  3. A runtime diagnostic, RJCP. An unbiased Hutchinson-style estimator of ‖J_g(f(x))J_f(x) − I‖_F² computed with JVP/VJP products, used to monitor how well the learned reverse map acts as a local left inverse at training and inference time.

  4. DeceptronNet v0. A single-scale, 2D unrolled corrector with a UNetSmall (3→32→1) architecture, fixed depth N = 6, learnable gain α_t = σ(γ_t) ∈ (0,1), initialization x₀ = ↑y, and projection onto [0,1].

Main Findings

  • Iteration reduction versus projected gradient: On Heat-1D initial-condition recovery, D-IPG reaches the fixed normalized tolerance (ε = 0.30) with approximately 20× fewer iterations than projected gradient, and approximately 2–3× fewer on the Damped Oscillator problem.

  • Comparable to Gauss–Newton in iterations, cheaper per step: In Table 1 (Heat-1D hard, mean ± std iterations), x-GD used 58.2 ± 28.9, D-IPG 2.8 ± 1.0, and GN/LM 2.8 ± 0.9. Final unnormalized RMSE was 0.045 (x-GD), 0.010 (D-IPG), and 0.009 (GN/LM). Acceptance rates were 1.00, 0.58, and 0.97 respectively.

  • Oscillator results: Table 1 reports x-GD 58.2 ± 52.1 iterations, D-IPG 24.6 ± 27.2, and GN/LM 17.3 ± 15.7, with final RMSE 0.356, 0.368, and 0.353 and acceptance rates 1.00, 0.64, and 0.69.

  • Per-step cost advantage: On Heat-1D (Table 2), median iterations [IQR] were 49.0 [38.2, 80.0] for x-GD at 0.43 ms/iter and 0.026 s mean time-to-ε; 3.0 [2.0, 3.0] for D-IPG at 0.51 ms/iter and 0.001 s; and 3.0 [2.0, 3.0] for GN/LM at 3.82 ms/iter and 0.011 s. D-IPG therefore matches GN/LM in iteration count but has much lighter iterations.

  • Oscillator timing and success: On the Oscillator (Table 3), median iterations were 65.0 [1.0, 104.5] for x-GD (success 0.50, 0.45 ms/iter, 0.004 s), 28.0 [1.0, 34.0] for D-IPG (success 0.45, 1.28 ms/iter, 0.001 s), and 16.5 [1.0, 33.2] for GN/LM (success 0.50, 4.22 ms/iter, 0.007 s).

  • Difficulty sweep: In the Heat-1D difficulty sweep, both x-GD and D-IPG require the most iterations at the medium setting, but D-IPG consistently converges in far fewer steps and shows its largest relative gain at the hard setting, roughly an order of magnitude fewer iterations than x-GD.

  • JCP effect: Enabling JCP reduces RJCP by several orders of magnitude. With JCP active on Heat-1D, the method reaches tolerance in about 2.6 iterations with final RMSE 0.007; disabling JCP raises the composition residual to 457.7 and requires roughly 3.8 iterations with higher error.

  • Ablations (Table 5): Removing JCP gave D-IPG 3.25 ± 1.18 iterations, RMSE 0.0159, acceptance 0.606 (vs. x-GD 56.3 ± 37.0, 0.0436, 1.000). Tying V = Wᵀ degraded conditioning: D-IPG 16.2 ± 8.55 iterations, RMSE 0.0894, acceptance 0.061 (vs. x-GD 106.9 ± 41.3, 0.0444, 1.000). Removing reconstruction and cycle terms did not harm convergence (D-IPG 2.60 ± 0.92, RMSE 0.0086), indicating the preconditioning effect is driven by the local inverse property rather than the auxiliary reconstruction losses.

  • Theoretical link to Gauss–Newton: The paper derives that ‖Δx_dipg − Δx_GN‖ ≤ α · (‖J_g J − I‖₂ / σ_min(J)) · ‖r‖, so as ‖J_g J − I‖₂ → 0, the D-IPG direction converges to the Gauss–Newton direction up to the scalar step size α, for residual components in range(J).

  • 2D PSF results (Table 4): LM (true model) required a mean 69.25 iterations with mean image RMSE 0.0883; x-GD (true model) 80.00 iterations with 0.1271; DNet v0 (unrolled N = 6) reached the target in 6.00 iterations with 0.0640 RMSE.

  • Kodak24 results (Table 7): Under Normal conditions (σ = 3.0), DNet achieved RMSE 0.0258 in 6 iterations and 0.0067 s, versus LM 0.0209 in 80 iterations and 0.0058 s, X-GD 0.0265 in 80 iterations and 0.0054 s, and L-BFGS 0.0276 in 80 iterations and 0.173 s. Under Hard conditions (σ = 4.0), DNet reached 0.0575 in 6 iterations and 0.0046 s. Under σ-Mismatch (train σ = 4.0, eval σ = 3.0), DNet reached 0.0525 in 6 iterations and 0.0045 s.

  • Diagnostics track convergence: During training, RJCP decreases steadily alongside validation error, and lower RJCP empirically correlates with fewer iterations to tolerance. The paper notes RJCP is a local diagnostic only and does not imply global invertibility.

  • Fairness protocol: All solvers used the same projector Π_C, Armijo rule with c = 10⁻⁴ (up to eight halvings), relaxation ρ = 0.4, a shared initial step size of 1.0, a maximum of 200 iterations, shared backtracking f evaluations, deterministic fixed seeds, and no proximal or smoothing heuristics. For DeceptronNet, all methods shared initialization, clamping, residual-based stopping at 0.3 r₀, an iteration budget of 80, and Armijo backtracking.

Methodology in Plain English

The forward process maps an unknown input to measurements. The authors train a small neural module that contains both a forward map and a reverse map, and they explicitly encourage the reverse map to undo the forward map locally. That is done by requiring that, when you apply the forward map's local behavior and then the reverse map's local behavior to the same direction, you end up where you started; the JCP term enforces this using random probe vectors and only a few JVP/VJP products, so no explicit Jacobians are ever formed.

At solve time, the algorithm does not take a gradient step directly on the unknown input. Instead it computes the residual between the predicted and observed measurements, takes a step in that measurement space, and pushes the result back through the learned reverse map. The resulting candidate is mixed with the previous iterate by a relaxation factor, projected onto the constraint set, and accepted or rejected by the same Armijo backtracking rule the baselines use. Because only the direction changes, comparisons remain fair. The paper argues that when the learned reverse map acts like the pseudoinverse of the forward Jacobian, the resulting direction approximates Gauss–Newton, but without solving linear systems at each iteration.

For the 2D extension, DeceptronNet unrolls a fixed number of correction steps through a compact U-Net that takes upsampled measurements, upsampled residuals, and the current estimate as inputs, predicts an image-space correction, and scales it by a learned gain before projecting back to [0,1].

Why This Matters

Impact on research: The paper offers a middle path between slow first-order projected gradient and expensive second-order methods. It provides an interpretable diagnostic (RJCP) that lets practitioners check whether a learned preconditioner is actually functioning as a local inverse, and it frames the learned inverse as a corrective accelerator rather than a full solver replacement. The derived bound connecting D-IPG to Gauss–Newton gives a formal reason why reducing composition error should translate into faster convergence.

Real-world applications:

  • Recovering initial conditions or parameters in physical simulations, such as heat diffusion or damped oscillators.
  • System identification, where unknown parameters of a system must be inferred from indirect measurements.
  • Imaging and deblurring, where the forward model combines blur, mild nonlinearity, downsampling, and Poisson-like noise (as in the 2D PSF task) or real photographic inverse tasks (Kodak24).
  • Scientific pipelines that require large parameter sweeps, where faster, better-conditioned solvers reduce compute cost.

Industry relevance: The per-iteration cost profile is significant for practitioners: D-IPG used 0.51 ms/iter on Heat-1D versus 3.82 ms/iter for GN/LM, while DeceptronNet reached tolerance in a fixed 6 steps, making runtime more predictable. Fixed-depth learned correctors are attractive for deployment settings where bounded, reproducible compute matters. The paper also warns that misuse of learned surrogates outside their validity domain can yield overconfident reconstructions, and recommends explicitly reporting surrogate ranges and RJCP metrics.

Future Directions

  • Extending DeceptronNet from the single-scale 2D prototype to multi-scale operators and higher-dimensional tasks, where richer architectures and multi-scale design could overcome the expressivity limits noted in Appendix D.
  • Testing on more realistic noise models and broader physical models beyond the blur, mild nonlinearity, downsample, and Poisson-like noise used in the 2D experiments.
  • Addressing the stated limitations: reliance on a reasonably accurate surrogate, locality of linearization, sensitivity to the scheduling of step-size gains α_t, and sensitivity to the timing and weighting of the JCP term, which may need manual tuning.
  • Handling highly non-identifiable regimes, where multiple inputs map to similar measurements — the paper states the method cannot resolve global ambiguity and only improves local conditioning, and suggests projected RJCP for under-determined regimes as an open direction.

Target Audience

Researchers and practitioners working on computational inverse problems, PDE-constrained optimization, and machine learning for the physical sciences, particularly those already familiar with Gauss–Newton, Levenberg–Marquardt, and projected gradient methods. It is also relevant to numerical optimization researchers interested in learned preconditioners, to imaging and system-identification engineers seeking faster reconstruction pipelines, and to readers interested in diagnostics that verify whether a learned module behaves as intended rather than as a generic regularizer. Beginners would need background in Jacobians, automatic differentiation, and line-search methods before the method section is accessible.

Authors’ abstract

Inverse problems in the physical sciences are often ill-conditioned in input space, making progress step-size sensitive. We propose the Deceptron, a lightweight bidirectional module that learns a local inverse of a differentiable forward surrogate. Training combines a supervised fit, forward-reverse consistency, a lightweight spectral penalty, a soft bias tie, and a Jacobian Composition Penalty (JCP) that encourages $J_g(f(x))\,J_f(x)\!\approx\!I$ via JVP/VJP probes. At solve time, D-IPG (Deceptron Inverse-Preconditioned Gradient) takes a descent step in output space, pulls it back through $g$, and projects under the same backtracking and stopping rules as baselines. On Heat-1D initial-condition recovery and a Damped Oscillator inverse problem, D-IPG reaches a fixed normalized tolerance with $\sim$20$\times$ fewer iterations on Heat and $\sim$2-3$\times$ fewer on Oscillator than projected gradient, competitive in iterations and cost with Gauss-Newton. Diagnostics show JCP reduces a measured composition error and tracks iteration gains. We also preview a single-scale 2D instantiation, DeceptronNet (v0), that learns few-step corrections under a strict fairness protocol and exhibits notably fast convergence.

Read the original paper