Research
Predicting symbolic ODEs from multiple trajectories
Overview Research area: Machine learning for scientific discovery — specifically transformer-based symbolic regression for recovering ordinary differential equations (ODEs) from data. Technical level:

- arXiv
- 2510.23295
- Published
- 2025-10-27
- Authors
- Yakup Emre Şahin, Niki Kilbertus, Sören Becker
AI summary
Overview
Research area: Machine learning for scientific discovery — specifically transformer-based symbolic regression for recovering ordinary differential equations (ODEs) from data.
Technical level: Advanced (assumes familiarity with transformers, sequence-to-sequence modeling, multiple instance learning, and symbolic regression).
Scope: The paper introduces MIO (Multiple Instance-ODEFormer), a transformer-based model that infers symbolic ODEs from multiple observed trajectories of the same dynamical system, and compares several strategies for aggregating information across those trajectories.
What This Paper Is About
Recovering the governing differential equations of a dynamical system from observed data is a central goal of scientific modeling, because equations are interpretable and generalize across system states. Existing transformer-based approaches such as ODEFormer can only process one trajectory at a time, even though multiple trajectories of the same system are often available. MIO addresses this limitation by combining multiple instance learning with transformer-based symbolic regression so that several trajectories jointly inform a single symbolic prediction.
Key Contributions
- MIO architecture: A transformer model that extends ODEFormer with an aggregator block placed between the encoder and decoder, allowing variable numbers of trajectory instances per system to be fused into a single system-level latent representation before decoding a symbolic ODE.
- Systematic comparison of aggregation strategies: The paper evaluates mean pooling, attentive pooling, time-agnostic attention pooling, and time-aware attention pooling, showing that the parameter-free mean pooling performs on par with the most sophisticated attention-based alternatives and clearly outperforms attentive pooling.
- Instance-count-aware training and evaluation: The authors train separate models per number of instances (1, 2, 3, 4) on 2D and 3D systems, demonstrating that training with multiple instances is what unlocks the benefit of extra instances at test time — models trained on a single instance actually degrade when given more.
- Inference-time rescaling for multiple instances: A shared scaling factor computed from the root mean square (RMS) of the initial values across all trajectories, which keeps all instances on a consistent scale so that a single symbolic prediction can be reliably mapped back to the original scale.
Main Findings
- Mean pooling is competitive: In Experiment 1, mean pooling, time-aware attention pooling, and time-agnostic attention pooling perform on par across all tested numbers of dimensions and instances; attentive pooling is clearly inferior. Simple mean aggregation is only marginally worse than the best attention-based pooling.
- Instances help generalization but hurt reconstruction: Increasing the number of instances improves performance on the generalization task (fitting one unseen trajectory) yet hurts reconstruction (fitting all observed input instances), because reconstruction requires the predicted ODE to fit more trajectories simultaneously.
- Performance declines with system dimensionality: In Experiment 1, performance degrades as dimensionality grows; the authors hypothesize that the number of training samples needed scales with dimension, while this count is roughly equal across dimensions in their dataset.
- Multi-instance training is essential: In Experiment 2, models trained on 2N–4N instances improve substantially as more instances are provided at test time, and performances between those models differ only marginally. In contrast, models trained on a single instance (1N models) get worse with additional instances.
- Robustness across noise: With multiplicative Gaussian noise sampled independently per time step from N(1, σ²) with σ = 0.05, MIO's absolute performance remains robust across noise levels while PySR, SINDy, and FFX degrade as noise grows.
- Outperforms all baselines: PySR, SINDy, FFX, and ODEFormer are all clearly outperformed by MIO on the generalization task. PySR, SINDy, and FFX reflect the same qualitative trends as MIO but suffer more from noise.
- Larger gain than top-n selection: ODEFormer (max) — a top-n evaluation that favors ODEFormer — shows a linear performance increase with instances; MIO's gain from one to two instances is larger than ODEFormer (max)'s, indicating MIO uses additional information beyond what a top-2 evaluation allows.
- Diminishing returns beyond two instances: The largest performance gain comes when moving from one to two instances, with smaller improvements thereafter; the authors hypothesize that the tested systems are too low-dimensional in their behavioral complexity to require more than two trajectories.
Methodology in Plain English
The researchers first generate a large synthetic corpus of differential equations. A mathematical expression is sampled, interpreted as the right-hand side of an ODE, and solved numerically (on the time interval [1, 10] with 100 equidistant support points using scipy.integrate.odeint, with rtol = 10⁻³ and atol = 10⁻⁶). Each system is given four initial values sampled from a standard normal distribution, and solutions where any component exceeds an absolute value of 10² are discarded. This produces paired data: numerical trajectories and their corresponding symbolic ODE expressions.
The model keeps ODEFormer's trajectory embedder, encoder, and decoder unchanged. Numbers are tokenized by rounding to four significant digits and splitting each float into sign, mantissa, and exponent — three tokens per value, yielding a vocabulary of 10,203 tokens. Systems smaller than the maximum dimension (D_max = 4) are zero-padded.
The novel piece is an aggregator block between encoder and decoder. The encoder processes each trajectory instance separately (like mini-batch elements) to produce instance-specific latents. The aggregator then fuses these latents into one system representation, and the decoder produces a single symbolic prediction per system. The authors keep track of how many instances belong to each system so only instances of the same underlying ODE are combined.
Four aggregation strategies are compared. Mean pooling simply averages the instance embeddings. Attentive pooling first aggregates over time using a 4-layer transformer encoder whose class-token embedding condenses each instance, then performs a softmax-weighted average over instances with learnable parameters. Time-agnostic attention pooling uses a learnable query in a cross-attention layer over all instances at once, avoiding the softmax competition between instances but ignoring temporal structure. Time-aware attention pooling concatenates instance embeddings, time-condensed embeddings, and a class token into a (2n+1)-wide tensor, processes it with a 4-layer transformer encoder, and uses the class-token output as the system representation.
Training uses Adam. Experiment 1 uses a cosine annealing schedule with 1,000 warm-up steps, cycles every 30,000 steps, period multiplier 1.1, shrinkage 0.75, learning rate oscillating between 2×10⁻⁴ and 1×10⁻⁹, and batch size 55 (about 135 unique ODE systems and roughly 180,000 numerical tokens per batch). Experiment 2 uses batch size 100 with the Noam scheduler and 2,000 warm-up steps, with a maximum learning rate of 4×10⁻⁴ for 2D models and 1×10⁻⁴ for 3D models, trained for ~1.95M steps. Inference uses beam sampling with temperature 0.1 and beam size 20.
Performance is measured as accuracy: the fraction of predictions for which the R² score between noiseless ground-truth and predicted trajectories exceeds 0.9. For reconstruction, a separate R² is computed per observed instance and the minimum must exceed the threshold. Two tasks are used: reconstruction (does the predicted equation reproduce the observed input trajectories?) and generalization (does it reproduce trajectories from a new initial value not used for inference?).
Why This Matters
The paper targets a practical bottleneck in data-driven science: single trajectories are often insufficient to identify governing equations because of noise, sparsity, or structural ambiguity — a problem documented even for linear systems and especially for nonlinear ones. By exploiting the repeated measurements that are commonly available in real experimental settings, MIO extracts more reliable symbolic models without any additional fitting time at inference.
Potential real-world applications (grounded in the paper's framing of repeated measurements, though the paper itself does not test any specific domain):
- Repeated-measures experimental designs, the scenario the authors cite explicitly, where the same system is observed under multiple initial conditions and all observations should inform one equation.
- Biological and biomedical systems, where population- or subject-level recordings of the same process (e.g., gene regulatory dynamics, pharmacokinetics) are collected across individuals and the underlying dynamics are shared.
- Physical sciences and engineering, where laboratory systems — chemical kinetics, mechanical oscillators, electrical circuits — are characterized by many repeated experimental runs, often noisy and irregularly sampled.
- Climate and environmental monitoring, where the same dynamical process is observed from multiple sensors, stations, or ensemble members and a single interpretable equation is desired.
Industry relevance: Many industrial settings generate repeated runs of the same process — condition monitoring of machinery, process control in chemical plants, or battery and thermal system characterization — where a compact symbolic model is more useful than a black-box predictor because it supports extrapolation, control design, and explanation. The paper's finding that a cheap, parameter-free mean-pooling aggregator suffices is directly relevant to deploying such models at scale, since it avoids the memory and compute cost of more elaborate attention-based fusion. Limiting factors noted by the authors are that evaluation is confined to test sets drawn from the same distribution as the training data, and scalability to higher dimensions remains untouched.
Future Directions
- Unifying the current model zoo: The authors train separate models per system dimension and per number of instances (8 models in Experiment 2). They explicitly name as future work the goal of unifying this into a single, generalizable model that handles varying dimensionality and instance counts.
- Explaining the diminishing returns: The paper asks why performance gains diminish after observing more than two instances, hypothesizing that the tested systems are too low-dimensional in behavioral complexity to need more than two trajectories. Testing this on higher-dimensional or qualitatively richer systems is an open question.
- Scalability to higher dimensions: The authors call this a "highly relevant open challenge that this work did not yet touch upon." Their results already show performance degrading with system dimensionality, and 4D was considered out of reach in Experiment 2.
- Data-distribution-agnostic evaluation: The current evaluation is limited to test sets drawn from the same distribution as training. The authors frame their contribution as understanding aggregation strategies rather than building a distribution-agnostic model or one for a particular application domain, leaving out-of-distribution generalization open.
Target Audience
Researchers and graduate students working at the intersection of machine learning and the physical sciences, particularly those interested in symbolic regression, equation discovery, and physics-informed or structure-aware deep learning. It is also relevant to practitioners who have repeated trajectory measurements of a dynamical system and want to recover interpretable governing equations, and to readers following the ODEFormer line of work who want to understand how multiple-instance learning can be layered onto sequence-to-sequence symbolic regression.
Authors’ abstract
We introduce MIO, a transformer-based model for inferring symbolic ordinary differential equations (ODEs) from multiple observed trajectories of a dynamical system. By combining multiple instance learning with transformer-based symbolic regression, the model effectively leverages repeated observations of the same system to learn more generalizable representations of the underlying dynamics. We investigate different instance aggregation strategies and show that even simple mean aggregation can substantially boost performance. MIO is evaluated on systems ranging from one to four dimensions and under varying noise levels, consistently outperforming existing baselines.