Research
CODE: A global approach to ODE dynamics learning
CODE: A global approach to ODE dynamics learning Overview Research area: Scientific machine learning / data-driven discovery of dynamical systems (learning ordinary differential equations from data).

- arXiv
- 2511.15619
- Published
- 2025-11-19
- Authors
- Nils Wildt, Daniel M. Tartakovsky, Sergey Oladyshkin, Wolfgang Nowak
AI summary
CODE: A global approach to ODE dynamics learningOverview
Research area: Scientific machine learning / data-driven discovery of dynamical systems (learning ordinary differential equations from data).
Technical level: Advanced. The paper assumes familiarity with ODEs, function-space approximation, polynomial chaos expansions, kernel methods, and NeuralODEs.
Scope in one sentence: The paper introduces ChaosODE (CODE), a global orthonormal polynomial (arbitrary Polynomial Chaos) representation of an ODE's right-hand side, and benchmarks it against NeuralODE and KernelODE on the Lotka–Volterra system under scarce data, noise, and unseen initial conditions.
What This Paper Is About
When scientists want to learn the governing dynamics of a physical system directly from measurements, they usually only have state observations at a few discrete, sparsely sampled time points, with noise, and no derivative information. The paper asks how different choices of the function space used to represent the ODE's right-hand side affect the ability to learn, and especially to extrapolate, in this difficult sparse-and-noisy regime. It proposes ChaosODE, which represents the dynamics with a data-driven, globally orthonormal polynomial basis, and compares it to the flexible-but-local NeuralODE and KernelODE alternatives.
Key Contributions
-
A new RHS representation: ChaosODE (CODE). The paper uses an arbitrary Polynomial Chaos Expansion (aPCE) as a global, orthonormal polynomial representation of the ODE right-hand side. Unlike a full polynomial basis, the basis is orthonormal with respect to the data distribution (the distribution of observed states), and the resulting model has relatively few coefficients that define the dynamics over the entire allowable domain.
-
A systematic comparison of three RHS ansatz spaces. NeuralODE (neural network), KernelODE (kernel/Gaussian-process-mean regression), and ChaosODE (polynomial chaos) are compared within the same Universal Differential Equation framing on the Lotka–Volterra benchmark.
-
An evaluation protocol that goes beyond interpolation. Three comparison setups are defined:
ex-it(in training time, same initial condition),ex-oot(out of training time, same initial condition), andex-ood(out of distribution, new initial condition), so that extrapolation to unseen initial states is explicitly tested rather than left unexplored. -
Practical optimization guidelines for robust dynamics learning. The paper documents a full pipeline (multiple- vs. single-shooting, discretize-then-optimize, initial RHS parameter estimation, PSO, CMA-ES, Quasi-Newton refinement) and states that the guidelines are illustrated in accompanying code.
Main Findings
-
CODE generalizes to unseen initial conditions. In scenario S1 (perfect-information baseline with 144 training observations, dense sampling, no noise, starting from globally pretrained coefficients), CODE reached an in-training MSE of 1.23 × 10⁻⁶, an out-of-training-time MSE of 2.83 × 10⁻⁶, and an out-of-distribution MSE of 0.0119 on new initial conditions.
-
The polynomial basis matches polynomial dynamics. Because the Lotka–Volterra right-hand side consists of the polynomial terms αx, βxy, γxy, and δy, the polynomial chaos basis can represent these dynamics exactly. The reported ex-ood error is attributed to accumulating floating-point errors from numerical integration over the extended horizon and to optimization possibly converging to a near-optimal local minimum.
-
Local data can degrade global approximations. S1 demonstrates that fine-tuning on a single trajectory improves accuracy along that trajectory while deteriorating performance in regions of state space not covered by the training data, i.e., local data can "carry away" an initially accurate global approximation.
-
High flexibility hurts extrapolation in scarce, noisy data. The abstract reports that NeuralODE and KernelODE, with their high flexibility, show degraded extrapolation capability under scarce data and measurement noise, and that CODE shows advantages over them.
-
Local methods behave locally. The paper argues NeuralODE interpolation within the training domain is effectively arbitrary in its extrapolation because the effective ansatz function is not understood, and that KernelODE extrapolation depends entirely on the chosen kernel function and its length scale(s), with accuracy highest near collocation points.
-
Transformed solution to polynomial dynamics. The paper notes the ex-ood MSE should theoretically be zero for CODE on Lotka–Volterra given its exact representation, and the nonzero value reflects numerical and optimization effects described above.
-
Results for individual methods in S1, and the quantitative outcomes of scenarios S2 (data effect), S3 (noise effect), and S4 (orthonormal vs. standard basis) are not reported in the provided content. The available text describes what these scenarios are designed to investigate but does not include their numerical results, and no specific noise levels or minimum data requirements are reported in the available content.
Methodology in Plain English
-
Frame the problem. Observations are noisy state measurements at discrete times: (tᵢ, x̂ᵢ) for i = 1…N, with x̂ᵢ = xᵢ + εᵢ. Since time intervals can be arbitrary, derivatives cannot simply be approximated by finite differences, so the state is assumed to follow an autonomous ODE, dx/dt = f(x(t)), with an initial condition.
-
Choose a space for f. Three options are compared: a neural network (NeuralODE), a kernel expansion with trainable values at pilot points (KernelODE), and a truncated orthonormal polynomial expansion (ChaosODE). For CODE, the polynomials are constructed from empirical moments of the state distribution using arbitrary polynomial chaos, normalized for numerical stability, and combined into multivariate tensor-product bases up to total degree N_max, giving M = (n + N_max choose n) basis functions.
-
Optimize through the ODE solution. Since only state measurements are available, training requires integrating the estimated dynamics. The paper adopts a discretize-then-optimize approach with a fourth-order Runge–Kutta (RK4) scheme and automatic differentiation, rather than optimize-then-discretize (adjoint) methods.
-
Cope with instability. Single-shooting over the full time interval is prone to divergence — especially for polynomials, which grow unboundedly outside their stable region. The paper therefore uses multiple shooting to obtain robust, segmented initial approximations, followed by a single-shooting refinement to smooth the solution.
-
Initialize well. A kernel regression surrogate is fitted to the time series; because kernel regression is closed under linear operations, time derivatives can be computed at arbitrary points, turning the problem into a simple regression to give all three RHS models sensible starting coefficients.
-
Run a staged optimization. Particle Swarm Optimization in a single-shooting setting to reach a stable, non-diverging solution; then CMA-ES in a multiple-shooting setting for gradient-free exploration; then a Quasi-Newton method in a multiple-shooting setting for exploitation; finally a single-shooting Quasi-Newton refinement to smooth the segmented solution. This was implemented with the Julia libraries Manopt.jl and Optim.jl.
-
Evaluate. Performance is measured with mean squared error between simulation results and observed data, across the three setups
ex-it,ex-oot, andex-ood, and across four scenarios S1–S4.
Benchmark details as reported: Lotka–Volterra with α = 1.5, β = 1, γ = 1, δ = 3, initial condition x₀ = (1.0, 1.0)ᵀ, following the conditions of Chen et al. (2019). Figures 1 and 2 illustrate the true solution for N = 35 training data points. Evaluation settings are: ex-it with t ∈ (0.0, 7.0) and u₀ = (1.0, 1.0)ᵀ; ex-oot with t ∈ (0.0, 14.0) and u₀ = (1.0, 1.0)ᵀ; ex-ood with t ∈ (0.0, 14.0) and u₀ᵒᵒᵈ = (0.5, 0.5)ᵀ.
Why This Matters
The paper argues that the choice of representation space for the right-hand side is a fundamental modeling decision that determines the space of achievable solutions, not a mere implementation detail. If problem-tailored RHS spaces improve learning, then the common default of a highly flexible neural network may be the wrong choice in the sparse, noisy, low-data settings that dominate many practical applications — the paper cites contaminant transport of TCE and PFAS as examples with far fewer observations than the hundreds of data points typical neural and kernel frameworks require.
Real-world applications (as motivated by the paper's framing):
- Hydrology and water resources — learning dynamics of environmental systems, following cited work on learning ODEs in hydrology.
- Geosciences — data-driven modeling of Earth system dynamics, as cited in the paper's overview of the field.
- Contaminant transport — the paper specifically names TCE and PFAS transport as settings where observations are few.
- Physics, chemistry, and biology — the paper notes polynomial dynamics are common across these fields, which is where a polynomial-chaos RHS is naturally suited.
Industry relevance: Any organization fitting dynamical models to sparse, noisy sensor data — environmental regulation and monitoring, chemical and process engineering, and ecological or resource management — benefits from knowing when a global, interpretable, low-coefficient model outperforms a flexible black-box one. The accompanying optimization guidelines and code are aimed at practitioners who need such models to train reliably rather than to diverge.
Future Directions
-
Whether CODE's advantage holds beyond polynomial dynamics. Its strong reported performance is tied to Lotka–Volterra having a polynomial RHS, which the polynomial basis can represent exactly. The paper's own framing leaves open how the approach behaves for non-polynomial dynamics where exact representation is impossible.
-
Scaling to higher dimensions. The paper explicitly notes that the number of basis functions grows combinatorially with dimension and polynomial degree, which "limits the method to moderate-dimensional problems." Extending CODE to higher-dimensional state spaces is an unresolved question.
-
How much data is actually needed. Scenario S2 is designed to identify minimum data requirements for reliable model performance, but the outcome is not reported in the available content — determining those thresholds remains an open question the design raises.
-
Robustness under noise. Scenario S3 is designed to analyze the impact of data noise on training robustness; the specific noise levels tested and the resulting degradation curves are not reported in the available content, and how CODE's polynomial basis behaves when the empirical moment estimates used to build the basis are themselves corrupted by noise is a natural follow-up.
-
Systematic tests on unseen initial conditions. The paper observes that tests on out-of-distribution, previously unseen initial states are rarely conducted in the NeuralODE literature, and its
ex-oodsetup is proposed as a remedy; broader adoption and standardization of this evaluation is an implicit open agenda.
Target Audience
Readers best served by this paper are:
- Scientific machine learning and physics-informed ML researchers comparing ansatz spaces for data-driven ODE learning, particularly those working on Universal Differential Equations.
- Uncertainty quantification researchers familiar with polynomial chaos expansions who want to see aPC repurposed from uncertainty propagation to dynamics learning.
- Computational scientists and engineers in hydrology, geosciences, and environmental modeling who fit dynamical models to sparse, noisy measurements and need robust training recipes.
- Practitioners of NeuralODE and kernel-based ODE learning who want to understand the extrapolation limitations of high-flexibility RHS representations and the optimization pitfalls (solver divergence, sensitivity) that come with optimizing through an ODE solve.
A background in ODEs, numerical integration, and either neural or kernel regression is assumed; the paper is not written as an introduction for newcomers.
Authors’ abstract
Ordinary differential equations (ODEs) are a conventional way to describe the observed dynamics of physical systems. Scientists typically hypothesize about dynamical behavior, propose a mathematical model, and compare its predictions to data. However, modern computing and algorithmic advances now enable purely data-driven learning of governing dynamics directly from observations. In data-driven settings, one learns the ODE's right-hand side (RHS). Dense measurements are often assumed, yet high temporal resolution is typically both cumbersome and expensive. Consequently, one usually has only sparsely sampled data. In this work we introduce ChaosODE (CODE), a Polynomial Chaos ODE Expansion in which we use an arbitrary Polynomial Chaos Expansion (aPCE) for the ODE's right-hand side, resulting in a global orthonormal polynomial representation of dynamics. We evaluate the performance of CODE in several experiments on the Lotka-Volterra system, across varying noise levels, initial conditions, and predictions far into the future, even on previously unseen initial conditions. CODE exhibits remarkable extrapolation capabilities even when evaluated under novel initial conditions and shows advantages compared to well-examined methods using neural networks (NeuralODE) or kernel approximators (KernelODE) as the RHS representer. We observe that the high flexibility of NeuralODE and KernelODE degrades extrapolation capabilities under scarce data and measurement noise. Finally, we provide practical guidelines for robust optimization of dynamics-learning problems and illustrate them in the accompanying code.