Skip to content
AI.info

Research

Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers

Overview Research area: Computer vision — autoregressive video prediction, with a focus on transformers and physics-based simulation data (partial differential equations, or PDEs). Technical level: Ad

Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
arXiv
2510.20807
Published
2025-10-23
Authors
Dean L Slack, G Thomas Hudson, Thomas Winterbottom, Noura Al Moubayed

AI summary

Overview

Research area: Computer vision — autoregressive video prediction, with a focus on transformers and physics-based simulation data (partial differential equations, or PDEs).

Technical level: Advanced. The paper assumes familiarity with transformer architectures, self-attention, ViT-style patch embedding, and generative video models.

One-sentence scope: The paper introduces PSViT, a pure pixel-space transformer that predicts future video frames of physical simulations, compares spatiotemporal attention layouts, and probes its internal representations for encoded physics.

What This Paper Is About

Most video generation models compress frames into a learned latent space before predicting the future, and they are usually judged on how good the output looks rather than whether the physics is right. The authors argue this hides a real weakness: these models lose physical accuracy quickly over time. They build a deliberately simple, end-to-end transformer that predicts video directly in continuous pixel space, and they measure how long its predictions stay physically faithful by tracking object positions over time.

Key Contributions

  1. A simple end-to-end pixel-space transformer. PSViT is a pure-transformer, U-Net-style architecture for autoregressive video prediction that omits complex architectural priors and specialised training objectives, and instead tests a range of patch-wise spatiotemporal self-attention strategies.

  2. Evidence that pixel-space modelling extends physical accuracy. Using an object-tracking divergence metric, the authors report that their model extends the time horizon of temporally accurate predictions on PDE-driven sequences compared to existing approaches, while also being competitive on the Moving MNIST and BAIR benchmarks — which they present as evidence of limits in some compressed latent-space models.

  3. Interpretability and probing of learned physics. The authors identify network regions that encode measurable physical dynamics and probe internal model representations to estimate out-of-distribution simulation parameters, arguing this shows a learned encoding of underlying physics not directly extractable from pixel space.

Main Findings

  • Up to 50% longer physically accurate predictions. The authors report that their approach significantly extends the time horizon for physically accurate predictions by up to 50% compared with existing latent-space approaches, while maintaining comparable performance on common video quality metrics.

  • Separating space and time beats joint attention. The paper reports that combining spatial and temporal information in a single joint self-attention operation (Joint-ST) performed considerably worse than isolating spatial and temporal operations into separately parameterised layers. Among the tested strategies, global-space + local-time (GS+LT, average divergence 3.08) and global-space + self-time (GS+T, average divergence 3.06) significantly outperformed local-space + local-time (LS+LT, average divergence 3.36) across all datasets. GS+T was selected as the standard configuration because it performs comparably to GS+LT while scaling linearly with sequence length.

  • Learnable positional encodings won. LPE achieved the best average divergence (3.05), compared with RoPE (3.08) and APE (3.21), so LPE is used in all PSViT configurations thereafter.

  • Patch size 8 was the sweet spot. Preliminary testing found P = 8 gave a better performance-efficiency trade-off than P = 16 (used in the original ViT work); larger patches were parameter-inefficient and smaller patches (e.g. P = 4) degraded performance and consumed prohibitively high memory.

  • Bigger models helped most on the 3D data. PSViT-medium (84M parameters) achieved an average divergence of 2.90 versus 3.05 for PSViT-small (49M). The authors note that increasing model size had a bigger impact on the 3D datasets than the 2D physics simulations, suggesting the extra parameters handle increased visual complexity rather than solely PDE dynamics.

  • Competitive divergence against larger baselines. Average divergence scores at t = 50: PSViT-medium 2.90, CV-VAE (160M) 3.13, SimVP (33M) 3.16, Diffusion Transformer (1B) 3.19, MAGVIT (300M) 3.24.

  • Better output-quality metrics. Average SSIM: PSViT-medium 0.9943, SimVP 0.9908, MAGVIT 0.9901. Average PSNR: PSViT-medium 54.56, MAGVIT 53.44, SimVP 53.23.

  • Mixed results on standard benchmarks. On BAIR (FVD, lower is better), PSViT-medium scored 64.1, behind MAGVIT (62.4), Diffusion Transformer (61.0) and CV-VAE (63.6) but ahead of SimVP (67.1). On Moving MNIST (SSIM, higher is better), PSViT-medium scored 0.963, ahead of SimVP (0.948), CV-VAE (0.945), MAGVIT (0.938) and the Diffusion Transformer (0.950).

  • Divergence tracks physical complexity. All tasks show low object divergence up to t = 20. The authors report that divergence rate for Roller, Moon and Pendulum correlates with the number of PDE variables — Pendulum, with five PDE variables, performs worst, versus Roller with two. The 3D Balls dataset showed similar performance across all models, which the authors attribute to its simple dynamics (objects share a constant velocity across sequences; only size and initial trajectory vary).

  • More context is not always better. When varying the number of input context frames, the maximum context size did not achieve the best performance for three of the datasets studied, indicating a limitation in handling increased context length. Context length affected when divergence begins but not necessarily its rate afterwards.

  • Quality metrics can miss physical failure. The authors observe a large relative drop in performance on the 3D and CLEVRER datasets when measured by object divergence that is not captured as markedly by L1, SSIM and PSNR — which they say makes the divergence metric an important factor in assessing accuracy of the underlying physics.

  • Register tokens explored. The model can add learnable register tokens appended to patch sequences at each time-step, intended as "memory" slots aggregating non-local, sequence-level signals such as global boundary conditions or slowly varying PDE modes; these are discarded before output frames are generated.

Methodology in Plain English

The researchers treat video frames as sequences of image patches, just as a language model treats text as sequences of tokens. Each frame is split into non-overlapping patches (size 8×8), flattened, and linearly embedded into a higher-dimensional space, with separate learnable positional encodings for time and space added on top.

The transformer backbone alternates between two kinds of attention. Spatial attention lets patches within a single frame look at other patches in that same frame. Temporal attention lets a patch look at the same spatial position in earlier frames, and it is causally masked so that no information from future frames can leak in. The two are separated by a feed-forward network. To keep computation manageable, the model is arranged like a U-Net: local space-time transformer blocks operate at full patch resolution, patch merge operations progressively reduce the effective resolution, and a global space-time block runs on the reduced representations, with skip connections preserving features through the reverse decoding path.

Training is fully self-supervised: the model sees the whole video minus the final frame and is trained to predict each frame from the ground-truth frames that precede it, so predictions never condition on the model's own outputs during training. Every model was trained with the Adam optimiser, a weight decay factor of 1e-4, batch size 32, the SSIM loss function, pixel values normalised to [0, 1], and a maximum of 2000 epochs on an NVIDIA A100 GPU, with no data augmentation.

Evaluation centres on a physics-aware metric rather than just image quality. The authors track the centroid of each object in predicted versus ground-truth frames and compute a rolling average 2D Euclidean pixel distance (the "Divergence" score), normalised to 128×128 resolution, reported at t = 50. SSIM and PSNR are reported as averages over the first 5 output frames.

Datasets used include four PDE simulation datasets (Moon, Pendulum, Roller, 3D Balls) at 128×128 with an 80/10/10 train/validation/test split and no overlapping initial conditions or parameters across splits, plus the CLEVRER colliding-objects dataset and the DPI-Net Fluid simulation (both centre-cropped and downsampled to 128×128), and the Moving MNIST and BAIR benchmarks. Comparisons are against MAGVIT (300M), Latent Diffusion Transformer Open-Sora (1B), CV-VAE (160M) and SimVP (33M).

Why This Matters

Impact on research. The paper challenges the widely used encoder-predictor-decoder assumption that compressing video into a latent space is necessary for good generation. It argues that latent models can produce high-quality video for stochastic datasets like BAIR but fall short on physical coherence over time, and it shows that a simple, parameter-efficient pixel-space model with 49M–84M parameters can beat 160M–1B parameter baselines on a physics-aware metric. It also contributes an interpretability angle: probing experiments suggest the network encodes PDE parameters that generalise to out-of-distribution parameters, something not directly readable from pixels.

Real-world applications (as identified by the authors):

  • Simulating fluid dynamics
  • Weather forecasting
  • Robot motion planning
  • Future scenario prediction for autonomous driving
  • Traffic prediction

Industry relevance. Any domain where a forecast must remain physically plausible over many time-steps — simulation, engineering, robotics, mobility and climate-adjacent industries — could benefit from a model that stays accurate longer without requiring pretrained encoders, complex training strategies or latent feature-learning components. The paper's claim of parameter efficiency relative to large latent baselines is directly relevant where inference and training cost matter.

Future Directions

  • Handling longer context. The finding that the maximum context size did not achieve the best performance for three of the datasets indicates a limitation in handling increased context length. Extending effective context handling is an open problem.
  • Scaling to more complex physics. Divergence rate correlated with the number of PDE variables, and larger models helped most on the visually complex 3D data. Whether these trends hold for more variables or harder simulations is untested here.
  • Better probing of encoded physics. The interpretability experiments generalise to out-of-distribution simulation parameters; the paper frames the model as a platform for further attention-based spatiotemporal modelling, leaving richer use of these internal representations open.
  • Training and efficiency trade-offs. The authors train without data augmentation and report model training and inference costs only in an appendix; broader exploration of training strategies, loss functions and resolutions beyond 128×128 is left open.

Target Audience

Researchers and practitioners in video generation, world models and physics-informed machine learning who are interested in whether latent compression is necessary for video prediction. It is also relevant to those working on model interpretability and probing of learned physical representations, and to engineers who need long-horizon, physically coherent forecasts at modest parameter counts. Readers without a background in transformer architectures will find the model details dense, though the central comparison — pixel space versus latent space, measured by object tracking rather than image quality — is accessible.

Authors’ abstract

Inspired by the performance and scalability of autoregressive large language models (LLMs), transformer-based models have seen recent success in the visual domain. This study investigates a transformer adaptation for video prediction with a simple end-to-end approach, comparing various spatiotemporal self-attention layouts. Focusing on causal modeling of physical simulations over time; a common shortcoming of existing video-generative approaches, we attempt to isolate spatiotemporal reasoning via physical object tracking metrics and unsupervised training on physical simulation datasets. We introduce a simple yet effective pure transformer model for autoregressive video prediction, utilizing continuous pixel-space representations for video prediction. Without the need for complex training strategies or latent feature-learning components, our approach significantly extends the time horizon for physically accurate predictions by up to 50% when compared with existing latent-space approaches, while maintaining comparable performance on common video quality metrics. In addition, we conduct interpretability experiments to identify network regions that encode information useful to perform accurate estimations of PDE simulation parameters via probing models, and find that this generalizes to the estimation of out-of-distribution simulation parameters. This work serves as a platform for further attention-based spatiotemporal modeling of videos via a simple, parameter efficient, and interpretable approach.

Read the original paper