Research
RealPDEBench: A Benchmark for Complex Physical Systems with Real-World Data
Overview Research area: Scientific machine learning / AI for physical systems — specifically, benchmarking surrogate models for complex spatiotemporal physical systems with real-world measured data. T

- arXiv
- 2601.01829
- Published
- 2026-01-05
- Authors
- Peiyan Hu, Haodong Feng, Hongyuan Liu, Tongtong Yan, Wenhao Deng, Tianrun Gao, Rong Zheng, Haoren Zheng, Chenglei Yu, Chuanrui Wang, Kaiwen Li, Zhi-Ming Ma, Dezhi Zhou, Xingcai Lu, Dixia Fan, Tailin Wu
AI summary
Overview
- Research area: Scientific machine learning / AI for physical systems — specifically, benchmarking surrogate models for complex spatiotemporal physical systems with real-world measured data.
- Technical level: Intermediate. The high-level motivation is accessible to any ML reader, but the paper assumes familiarity with neural operators, PDE surrogates, and fluid/combustion physics terminology.
- Scope: In one sentence: the paper introduces RealPDEBench, the first scientific ML benchmark that pairs real-world measurements with matched numerical simulations across five complex physical systems, and uses it to quantify the gap between simulated and real data and the value of simulated pretraining.
Note on completeness: the provided paper text is truncated (it cuts off mid-sentence in the low-/mid-/high-frequency fRMSE analysis) and the appendices it repeatedly refers to (notably Appendix B, B.3, B.4, D, E, A.1, A.2) are not included, so several details are reported here only at the level the main text gives.
What This Paper Is About
Scientific ML models for predicting how physical systems evolve are almost always trained and tested on data produced by numerical solvers, not on measurements from the real world, because real-world data is expensive to collect. The authors argue this leaves open the key question of how these models actually perform on reality, and they build a benchmark of paired real and simulated data covering fluid dynamics and combustion to answer it.
Key Contributions
- Data. Five real-world measured datasets with paired simulated datasets, covering more than 700 trajectories that each exceed 2000 frames (736 trajectories in total), across five scenarios: Cylinder, Controlled Cylinder, Fluid-Structure Interaction (FSI), Foil, and Combustion. Each scenario includes multiple system parameters, and for every parameter setting both real-world and simulated data of equal temporal duration are provided. The Combustion dataset consists of experimental and numerical cross-sectional data of 3D swirl-stabilized NH3/CH4/air flames.
- Tasks. Three prediction settings, all of them evaluated on real-world data: (i) training on simulated data, (ii) training on real-world data, and (iii) pretraining on simulated data followed by finetuning on real-world data. For each dataset there are N real-world samples and N simulated samples, with n used for training, (N−n)/2 for validation and (N−n)/2 for testing, split at the parameter level, with the validation and test sets fixed to real-world samples across all three tasks.
- Metrics. Nine evaluation metrics split into data-oriented ones (RMSE, MAE, Relative L2 Error, coefficient of determination R², and an Update Ratio measuring how many finetuning versus from-scratch updates are needed to reach the best real-world-training RMSE) and physics-oriented ones (Fourier Space Error at low/mid/high frequency bands, Frequency Error, Kinetic Energy Error, and Mean Velocity Profile Error).
- Baselines and framework. Ten baselines: nine data-driven scientific ML models and one traditional method. The ML models include U-Net, CNO, DeepONet, FNO, WDNO, MWT, GK-Transformer, Transolver, and DPOT (a pretrained PDE foundation model, described in the text as a 30M small model and a 509M large model, both finetuned on these datasets); the traditional method is DMD. All datasets and baselines are integrated into a unified, modular PyTorch codebase with common
RealDatasetandModelmodules.
Main Findings
- There is a substantial gap between simulated and real-world data. The gaps appear in three ways: different modalities (simulated data usually contain more modalities than real-world data because of measurement limitations), different error behavior (models trained on simulated data generalize poorly to real data even when physical parameters match), and different error sources (measurement noise versus numerical errors caused by simplified physics, ideal conditions, or modeling such as Large Eddy Simulation and discretization with second-order convergence).
- Real-world training beats simulated training by a large margin. When tested on the same real-world test set, real-world training improves Relative L2 Error over simulated training by 9.39% to 78.91%. On the Cylinder dataset the ML-average RMSE is 0.0752 for simulated training, 0.0692 for real-world training, and 0.0613 for real-world finetuning.
- Simulated training misses periodicity. On the Controlled Cylinder dataset, simulated training yields much higher Frequency Errors than real-world training, indicating simulated data cannot perfectly capture periodicity in real-world systems.
- Simulated pretraining nevertheless helps. The real-world finetuning column of the main results table shows lower errors than real-world training in most cases, meaning pretraining on abundant simulated data and then finetuning on scarce real data beats training on the same amount of real data alone. On the Combustion dataset, the validation RMSE curve for finetuning drops much faster than for real-world training.
- Finetuning converges in fewer updates. The Update Ratio (finetuning updates divided by from-scratch training updates needed to reach the best real-world-training RMSE) is below 1 for most datasets and baselines. Reported ML averages per dataset are 0.5666 (Cylinder), 0.6504 (Controlled Cylinder), 0.4964 (FSI), 0.5566 (Foil), and 0.7559 (Combustion).
- Large pretrained foundation models perform best overall. In the RMSE-versus-Frequency-Error trade-off plot, DPOT-L-FT (the large pretrained model) is closest to the origin, reflecting the benefit of large-scale PDE pretraining and more parameters.
- Convolution-based models do well on pixel-level error. U-Net and CNO achieve lower RMSE, which the authors attribute to the tasks resembling image processing. MWT shows advantages in learning periodicity thanks to its multiwavelet transform. Because most models are trained with data-oriented losses such as MSE, they tend to excel on local features and are weaker on physics-oriented metrics that involve global features.
- Long-horizon behavior differs by model. Under 1, 2, 3, 5, and 10 rounds of autoregressive prediction on the Cylinder dataset, CNO performs well in single-round prediction but its error grows faster than other methods as rounds increase, suggesting error accumulation. Under MVPE at 10 rounds, DMD shows limitations while the large DPOT model performs best.
- The traditional method is included as a reference. DMD, which has no training process, is reported only in the real-world finetuning column. On Cylinder its RMSE is 0.0862 (Rel L2 0.3590, fRMSE 0.0114), and it is outperformed by most learning baselines.
Methodology in Plain English
The authors start from the observation that measuring real physical systems and simulating them produce different kinds of data, and that both are needed to judge scientific ML honestly. They therefore collect real measurements and, for the exact same parameter settings, run numerical simulations of equal temporal duration.
Real-world measurement. Fluid datasets are measured in circulating water tunnels using Particle Image Velocimetry. Fluorescent hollow-glass microspheres of 10 micrometers are seeded into the water and illuminated by a continuous laser forming a 2 mm thick laser layer on the shooting surface, so fluid velocity is inferred from particle velocities recorded by high-speed cameras and processed with PIVLab. The structure under study is mounted on a rail driven by a stepper motor, which allows controlled motion. Combustion data are measured with flame chemiluminescence imaging: a swirl combustor is fed air, NH3 and CH4 through mass flow controllers that set the injection ratio, the mixture is ignited, and light intensity is recorded with an OH* CL camera.
Simulated data. Simulated data are generated with Computational Fluid Dynamics using the Finite Volume Method and the Immersed Boundary Method, with Lilypad as the solver for 2D experiments and Waterlily, an efficient GPU-based 3D solver, for 3D simulations. For Combustion, the authors run a three-dimensional implicit unsteady Large Eddy Simulation of thermoacoustic instabilities in a swirl-stabilized flame, using the Eddy Dissipation Concept to model turbulence-chemistry interaction.
Storage and splits. All data are stored in HDF5, with each file holding one trajectory sampled at equal time intervals on a uniform spatial grid as a NumPy array of shape (T, X, Y) with C modalities, plus the corresponding system parameters such as Reynolds number, oscillation frequency and equivalence ratio.
Training and evaluation. Every model is evaluated on real-world data. To make simulated training more realistic, the authors add noise and randomly mask modalities that are not measured in the real-world data. For each task they compute the nine metrics described above, and they additionally run autoregressive evaluations with 1, 2, 3, 5 and 10 rounds, where each predicted block of T steps is fed back as input so that error is measured over N·T steps.
Why This Matters
Impact on research. Benchmarks in this area (PDEBench, the Well, fluid-focused datasets from Luo et al. and Tali et al., and the concurrent REALM framework for multiphysics reactive flows) mostly rest on simulated data. RealPDEBench adds paired real measurements, which makes it possible to study sim-to-real transfer, learning from noisy data, and how measurement limitations affect model quality in a controlled, reproducible way. The authors open-source the datasets, benchmark and instructions at https://realpdebench.github.io/.
Real-world applications (from the scenarios the paper covers):
- Unsteady wake flows behind cylinders, including the classical Kármán vortex street, relevant to any bluff-body flow problem.
- Flow control, since the Controlled Cylinder dataset uses active external forcing with periodic sinusoidal control at different frequencies.
- Fluid-structure interaction such as bridges under wind loading and offshore platforms in ocean currents, including lock-in phenomena and galloping instabilities.
- Airfoil/hydrofoil design and marine engineering optimization, using the Foil dataset's angles of attack and Reynolds numbers.
- Combustion and thermoacoustic instability control in aircraft and space engines, using the swirl-stabilized NH3/CH4/air flame data where concentration, temperature and other fields must be predicted.
Industry relevance. Any sector that currently relies on either expensive experiments or unvalidated simulations — aerospace and propulsion, energy and combustion, marine and offshore engineering, and industrial flow design — can use this benchmark to judge whether an ML surrogate is trustworthy on measured data before deployment. The finding that simulated pretraining speeds up convergence and improves accuracy is directly actionable for teams with limited real measurement budget.
Future Directions
- How to best combine real and simulated data. The authors state their work is a foundation for further exploring how to combine the advantages of both data sources and achieve improved models; the three task settings are a starting point rather than a solved recipe.
- Better sim-to-real transfer. Closing the gap revealed by the 9.39% to 78.91% Relative L2 differences, and handling noise and unmeasured modalities more principledly than the noise-injection and random-masking strategies used here.
- Extending coverage. The modular
RealDatasetandModeldesign is explicitly intended to make it straightforward to add new datasets, scenarios and baselines, including models beyond the ten tested. - Understanding measurement limitations. The paper names "understanding how limitations of measurement techniques influence model performance" as a task this benchmark makes possible, but does not resolve it — the relative influence of camera noise, non-uniform incoming flow and other measurement artifacts on model ranking remains open.
- Scaling of PDE foundation models. DPOT-L performs best, but whether further scaling or different pretraining objectives consistently improve physics-oriented metrics such as Frequency Error and mean velocity profile accuracy is not established.
Target Audience
Researchers and practitioners in scientific machine learning who build or evaluate surrogate models for PDE-governed systems; developers of neural operators and PDE foundation models who need an out-of-distribution, real-data test bed; fluid dynamicists and combustion researchers interested in data-driven prediction of wake flows, flow control, fluid-structure interaction, foils, and swirl-stabilized flames; benchmark designers who want a template for pairing real measurements with matched simulations; and engineering teams in aerospace, energy, marine and propulsion sectors assessing whether ML surrogates are ready for deployment.
Authors’ abstract
Predicting the evolution of complex physical systems remains a central problem in science and engineering. Despite rapid progress in scientific Machine Learning (ML) models, a critical bottleneck is the lack of expensive real-world data, resulting in most current models being trained and validated on simulated data. Beyond limiting the development and evaluation of scientific ML, this gap also hinders research into essential tasks such as sim-to-real transfer. We introduce RealPDEBench, the first benchmark for scientific ML that integrates real-world measurements with paired numerical simulations. RealPDEBench consists of five datasets, three tasks, eight metrics, and ten baselines. We first present five real-world measured datasets with paired simulated datasets across different complex physical systems. We further define three tasks, which allow comparisons between real-world and simulated data, and facilitate the development of methods to bridge the two. Moreover, we design eight evaluation metrics, spanning data-oriented and physics-oriented metrics, and finally benchmark ten representative baselines, including state-of-the-art models, pretrained PDE foundation models, and a traditional method. Experiments reveal significant discrepancies between simulated and real-world data, while showing that pretraining with simulated data consistently improves both accuracy and convergence. In this work, we hope to provide insights from real-world data, advancing scientific ML toward bridging the sim-to-real gap and real-world deployment. Our benchmark, datasets, and instructions are available at https://realpdebench.github.io/.