Research
Automated Discovery of Conservation Laws via Hybrid Neural ODE-Transformers
Overview Research area: Machine learning for scientific discovery — specifically symbolic regression and invariant/conservation-law discovery from observational data. Technical level: Intermediate. Th

- arXiv
- 2511.00102
- Published
- 2025-10-30
- Authors
- Vivan Doshi
AI summary
Overview
Research area: Machine learning for scientific discovery — specifically symbolic regression and invariant/conservation-law discovery from observational data.
Technical level: Intermediate. The paper assumes familiarity with Neural ODEs, Transformers, and reinforcement learning fine-tuning, though the three-stage pipeline is described clearly enough for a reader with general ML background.
Scope: This workshop paper (MATH-AI: The 5th Workshop on Mathematical Reasoning and AI; arXiv:2511.00102v1 [cs.LG], 30 Oct 2025, licensed CC BY 4.0) proposes and empirically tests a three-stage hybrid pipeline for recovering conserved quantities from noisy trajectory data.
What This Paper Is About
Conservation laws — quantities like energy or angular momentum that stay constant as a system evolves — are central to physics, but finding them automatically from messy, noisy measurements is hard. Existing approaches either learn accurate but opaque dynamics models or search for symbolic expressions that tend to break down under noise. The author's goal is to combine both strengths: learn a clean continuous model of the dynamics first, then search for symbolic invariants of that learned model, then certify the results numerically.
Key Contributions
- A decoupled three-stage architecture that first learns a system's vector field with a Neural ODE, then uses a Transformer to generate symbolic candidate invariants conditioned on that learned model, rather than searching directly on raw trajectories.
- A symbolic-numeric verification module that uses exact symbolic differentiation (via a library such as SymPy) plus dense numerical evaluation to filter candidates, serving as a strong certificate that a candidate is a true invariant of the learned dynamics rather than an artifact of noise.
- An empirical demonstration across three canonical physical systems showing substantially higher discovery rates than two baselines (PySR and an end-to-end Transformer), with reported 95% Wilson confidence intervals over 20 runs.
- Ablations and failure-mode analysis showing that removing the Neural ODE module causes a sharp performance drop, and that failure correlates with poor Neural ODE fidelity (validation MSE above 10⁻³).
Main Findings
- Harmonic oscillator (energy): The hybrid method succeeded in 95% of runs [75, 100], versus 75% [51, 91] for PySR and 60% [36, 81] for the end-to-end Transformer.
- Pendulum (energy): Hybrid 90% [68, 99], PySR 60% [36, 81], end-to-end 55% [32, 77].
- Kepler problem (energy): Hybrid 70% [46, 88], PySR 15% [3, 40], end-to-end 5% [0, 25].
- Kepler problem (angular momentum): Hybrid 80% [56, 94], PySR 20% [6, 44], end-to-end 10% [1, 32].
- Noise robustness: On the harmonic oscillator, the method maintained a discovery rate above 70% at 10% noise, while baseline performance collapsed below 20%.
- Ablations: Removing the Neural ODE module caused a sharp performance drop. Omitting the Transformer's pre-training stage also degraded performance because the model struggled to produce syntactically valid expressions.
- Failure mode: When the Neural ODE underfits (validation MSE above 10⁻³), the Transformer generates spurious invariants — for example, on a poorly learned pendulum it produced Ĉ = 0.8(p² + q²) + 0.3 sin(q), which passed verification for the learned model but deviated by 15% on true trajectories.
- Baseline failure mode: Baselines often produced overly complex expressions that fit trajectory noise rather than the underlying dynamics.
- Preliminary chaotic-system result: On the Lorenz system (σ = 10, ρ = 28, β = 8/3), the method discovered the dissipation relation V̇ = −σx² − y² − βz² in 45% of runs versus 10% for baselines, but struggled with the two quadratic invariants.
Methodology in Plain English
The pipeline has three sequential parts.
First, a Neural ODE learns the system's dynamics from observed trajectories. A neural network represents the unknown vector field, and its parameters are optimized by comparing trajectories produced by an adaptive-step numerical ODE solver (e.g., Dopri5) against the observed data, using the adjoint sensitivity method for memory-efficient backpropagation. The result is a continuous model robust to irregular sampling. The implementation used a 4-layer MLP with 128 hidden units and Swish activations, trained for 200 epochs with Adam at a learning rate of 10⁻³ and batch size 64, until validation MSE fell below 10⁻⁵.
Second, a Transformer proposes symbolic candidate invariants. A quantity C(z) is conserved when ∇zC(z) · f_θ(z) = 0, and the Transformer is trained to produce expressions satisfying this. It is first pre-trained on a large corpus of mathematical expressions to learn syntactic priors over a grammar of variables (x, y, v_x, v_y) and operators {+, −, *, /, sin, cos, pow}. It is then fine-tuned with Proximal Policy Optimization (PPO), receiving state-derivative pairs (z_j, f_θ(z_j)) and generating a candidate Ĉ(z). The reward is R(Ĉ) = exp(−λ₁ · err) + λ₂ · ||∇Ĉ||₂, where "err" is mean squared invariance error over a batch of points and the second term is a non-degeneracy penalty discouraging trivial solutions. The candidate generator was a 6-layer Transformer fine-tuned for 50 epochs.
Third, a verifier certifies candidates. It computes the exact gradient ∇zĈ(z) symbolically, then numerically evaluates |∇zĈ(z) · f_θ(z)| over a dense uniform grid of 10,000 points sampled from the convex hull of the training data. If the maximum value falls below a strict threshold (e.g., 10⁻⁶), the candidate is certified as an invariant of the learned model.
Experiments covered the harmonic oscillator, pendulum, and 2D Kepler two-body problem, with trajectories generated at 2% Gaussian noise. A discovery counted as successful if the expression was functionally equivalent to ground truth, non-trivial, and met an RMSE threshold. Baselines were PySR and an end-to-end Transformer. Compute was a single NVIDIA RTX 3090 GPU (24GB): roughly 1–2 hours to train the Neural ODE per system, 2–3 hours for Transformer fine-tuning, and 5–15 minutes per candidate for verification parallelized across 100 grid points, totaling 3–6 hours per experimental run.
Why This Matters
Impact on research: The paper argues for modular "learn-then-search" pipelines over end-to-end symbolic regression for invariant discovery. It also provides a candid account of the trade-off that dominates the field: end-to-end models risk overfitting noise, while physics-informed models (such as Hamiltonian Neural Networks, which require pre-specifying a Hamiltonian structure) generalize well but sacrifice discovery potential. The verification module introduces a numerical certificate step that the author positions as a guard against noise-driven false discoveries — though the paper is explicit that the certificate applies to the learned model, not the true system.
Real-world applications (as framed or suggested by the paper):
- Automated scientific discovery from noisy experimental or observational data, where underlying physical laws are obscured.
- Systems biology, which the paper names as a domain where the method's denoising properties may be especially valuable.
- Econometrics, also named in the paper's future-work discussion as a target domain.
- Physical simulation and modeling of mechanical systems, illustrated by the harmonic oscillator, pendulum, and Kepler benchmarks.
Industry relevance: The paper claims no theoretical results, no dataset or asset release, and no detailed societal-impact discussion (the checklist answers these as N/A or No). Code and data are not publicly hosted; the author states they will be made available to researchers upon reasonable request. Industry relevance is therefore indirect — it lies in the general promise of extracting interpretable governing equations from messy sensor or simulation data rather than in a deployed tool.
Future Directions
- More robust ODE learning: The author plans to explore architectures incorporating equivariant layers or symplectic integrators to better handle structured systems, since success depends on learning an accurate ODE model.
- Formal verification: Applying tools such as alpha-beta CROWN to provide provable certificates for discovered invariants with respect to the learned dynamics f_θ(z), replacing the numerical grid check.
- Closing the gap between learned and true invariants: Quantifying the difference between invariants of the learned proxy model and invariants of the real system is described as a key open challenge.
- Scaling and new domains: Extending to higher-dimensional systems and to real-world data from areas like systems biology or econometrics, plus developing specialized techniques for chaotic dynamics, where the paper reports only partial success (e.g., failure on the Lorenz system's two quadratic invariants).
Additional stated limitations: the approach struggles with stiff or chaotic systems, and all experiments reported are on well-behaved, low-dimensional systems.
Target Audience
Researchers working at the intersection of machine learning and the physical sciences — particularly those interested in symbolic regression, Neural ODEs, equation discovery, and AI-assisted scientific discovery. It is also useful for practitioners weighing modular, verifier-backed pipelines against monolithic end-to-end models, and for reviewers or students seeking a concise statement of the noise-robustness trade-offs in this subfield. Readers seeking proven theoretical guarantees or released code and benchmarks will not find them here, as the paper is an empirical workshop contribution that reports neither.
Authors’ abstract
The discovery of conservation laws is a cornerstone of scientific progress. However, identifying these invariants from observational data remains a significant challenge. We propose a hybrid framework to automate the discovery of conserved quantities from noisy trajectory data. Our approach integrates three components: (1) a Neural Ordinary Differential Equation (Neural ODE) that learns a continuous model of the system's dynamics, (2) a Transformer that generates symbolic candidate invariants conditioned on the learned vector field, and (3) a symbolic-numeric verifier that provides a strong numerical certificate for the validity of these candidates. We test our framework on canonical physical systems and show that it significantly outperforms baselines that operate directly on trajectory data. This work demonstrates the robustness of a decoupled learn-then-search approach for discovering mathematical principles from imperfect data.