Research
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
Overview Research area: Machine learning for spatiotemporal modeling, specifically geometric positional encodings for Transformers applied to molecular dynamics and video prediction. Technical level:

- arXiv
- 2609.33804
- Published
- 2026-09-27
- Authors
- Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu
AI summary
Overview
Research area: Machine learning for spatiotemporal modeling, specifically geometric positional encodings for Transformers applied to molecular dynamics and video prediction.
Technical level: Intermediate to Advanced. The core idea rests on Minkowski geometry, Lorentz transformations, and the indefinite metric signature diag(−1, +1, …, +1), so some familiarity with special relativity notation or Lie-group reasoning helps. The empirical sections are accessible without that background.
Scope: The paper proposes a relative positional encoding that couples time and space by parameterizing Lorentz transformations with joint spatiotemporal coordinates, and evaluates it on one microscopic task (protein–ligand molecular dynamics on the MISATO dataset) and one macroscopic task (KTH video prediction) within a single task-agnostic Transformer framework.
What This Paper Is About
Models for spatiotemporal data tend to sit at two extremes: physics-motivated approaches that bake in dynamical laws (recurrent, convolutional, or differential-equation propagators) and thereby gain strong priors at the cost of flexibility, or learning-motivated Transformer architectures that are flexible but leave the coupling between time and space implicit in learned attention. The authors ask whether it is possible to keep the flexibility of data-driven learning while adding an explicit geometric prior for how time and space relate. Their answer is to make positional encoding the site of that prior: treat each token as an event in a joint time–space coordinate system and use those coordinates to generate Lorentz transformations on query and key features, so that attention scores depend only on the relative spacetime displacement between tokens.
Key Contributions
-
Minkowski Positional Encoding (MinkowskiPE): A relative positional encoding in which each token's spatiotemporal coordinate x_i = (t_i, r_i) parameterizes a Lorentz transformation applied to its query and key features. The indefinite metric η = diag(−1, +1, …, +1) distinguishes the temporal direction from the spatial directions, and the metric factor η inserted into the key transformation makes the ordinary Euclidean dot product realize a Minkowski bilinear interaction.
-
Translation invariance by construction: Because the transformation is homomorphic (Λ(x₁ + x₂) = Λ(x₁)Λ(x₂), Λ(0) = I), the positional contribution to the query–key score reduces to q_i^T η Λ(w^T Δx_ij) k_j, depending on position only through the relative displacement Δx_ij = x_j − x_i. A group-theoretic analysis in Appendix A shows that for the (1+1)-dimensional Lorentzian feature blocks used here, any smooth homomorphism from the additive coordinate group into SO⁺(1,1) necessarily reduces to a learnable scalar projection followed by a one-parameter Lorentz boost.
-
A practical multi-head parameterization: For head dimension d_h, query and key features are partitioned into P = d_h/2 two-dimensional blocks, with a learnable projection matrix W ∈ ℝ^(D×P) shared across all Transformer layers and attention heads. The block-wise boosts are written in an element-wise form using cosh and sinh, and the construction retains the standard dot-product attention interface, so it remains compatible with efficient attention implementations such as FlashAttention.
-
Cross-scale empirical validation: The same positional formulation is applied unchanged to molecular dynamics (three-dimensional atomic motion) and video prediction (two-dimensional visual scenes), covering both microscopic and macroscopic spatiotemporal prediction.
Main Findings
-
Single-trajectory molecular dynamics: Across ten protein–ligand systems and three metrics on the MISATO dataset (10 observed frames predicting 20 future frames), MinkowskiPE achieves the best result in 17 of 30 evaluations and ranks among the top two in 26 of 30. It achieves the best performance across all three metrics on 4K6W, 1XP6, 4G3E, and 3B9S. Results are averaged over three random seeds.
-
Multi-trajectory molecular dynamics: On MISATO-100, MISATO-1000, and MISATO-Full (80 observed frames predicting 20 future frames), MinkowskiPE achieves the best performance on all nine evaluations, covering both coordinate accuracy and structural metrics, and the advantage persists as the number and diversity of training trajectories increase.
-
Video prediction: On KTH under the OpenSTL protocol (10 grayscale frames at 128×128 predicting 20 frames), MinkowskiPE achieves the best performance on MSE, MAE, PSNR, and LPIPS, while remaining competitive on SSIM. Its best run reports MSE 35.66, MAE 347.7, SSIM 0.9078, PSNR 28.18, and LPIPS 0.18490 with 2.5M parameters. The three-seed mean reports MSE 35.78, MAE 347.9, SSIM 0.9079, PSNR 28.15, and LPIPS 0.18342.
-
Parameter efficiency: Compared with the strongest baseline result for each metric, MinkowskiPE reduces MSE, MAE, and LPIPS by 9.9%, 5.7%, and 1.7% respectively while using only 2.5M parameters. The paper states this is roughly one-tenth as many parameters as the lowest-MSE baseline.
-
Controlled ablation: Removing only the positional encoding while keeping architecture, optimization, and training configuration unchanged degrades every evaluated metric. On MISATO-1000, MinkowskiPE reduces MAE by 2.7% and Matching by 3.9%, and improves Stability by 1.2%. On KTH it reduces MSE by 5.7% and MAE by 3.6%, and improves LPIPS by 4.5%, with consistent gains in SSIM and PSNR. All ablation results are averaged over three random seeds.
-
Interpretation of the geometry: The authors emphasize that MinkowskiPE does not assume the underlying data obey relativistic dynamics. Minkowski geometry is used as a representation-level inductive bias, not as a symmetry imposed on the data.
-
Galilean obstruction: Appendix C analyzes a construction based on finite-dimensional unitary token-wise transformations and shows that exact Galilean invariance imposes a strong structural constraint: the resulting pairwise interaction cannot retain the Galilean-invariant spatial component of the relative group element. The authors state this does not rule out Galilean positional encodings generally.
Methodology in Plain English
The approach starts from a simple observation: both an atom in a molecular simulation and an image patch in a video can be described as an "event" with a time coordinate and a spatial coordinate. In molecular dynamics the coordinate is (t, x, y, z); in video it is (t, u, v) on a latent grid.
Rather than injecting time and space through separate branches or hand-designed attention patterns, the authors use those coordinates to modify the query and key vectors inside attention. Each token's spacetime coordinate is compressed by a learned linear projection into a scalar, and that scalar controls a "boost" — a hyperbolic rotation built from cosh and sinh — applied to two-dimensional slices of the query and key features. The key side additionally gets multiplied by a Minkowski metric factor with a −1 in the temporal slot and +1 in the spatial slots. That sign difference is what makes time and space enter with opposite signs.
The algebraic payoff is that the two absolute transformations cancel down to a transformation of the relative displacement between the two tokens. Attention then depends only on how far apart two events are in spacetime, not on where they sit in absolute coordinates — a translation-invariance property obtained from the group structure rather than from adding a relative-bias table.
For implementation, the head's feature dimension is split into blocks, each block receiving its own learned direction in spacetime. A single projection matrix is shared across layers and heads. The element-wise cosh/sinh form means the modification stays inside the standard dot-product attention computation.
Architecturally, the molecular model treats each ligand atom at each observed time step as a token, applies frame-wise cross-attention from ligand atoms to static protein residues followed by global self-attention over ligand event tokens, and decodes future displacements with an MLP from the last observed frame. The video model encodes each frame into a 16×16 latent feature grid, treats each latent feature as an event token, concatenates observed tokens with future tokens initialized from the last observed latent map and augmented with learnable future-step embeddings, processes everything with an eight-layer global Transformer, and decodes back to frames with a convolutional decoder.
Baselines for molecular dynamics are VerletMD, GNN-MD, DenoisingLD, NeuralMD-ODE, NeuralMD-SDE, and BioMD. Video baselines reported from OpenSTL are ConvLSTM, MIM, PredRNN variants, SimVP variants, and TAU.
Why This Matters
Impact on research. The paper reframes positional encoding as a design axis for spatiotemporal Transformers, rather than treating it as a preprocessing detail. It demonstrates that a geometric prior can be inserted at the representation level without prescribing a dynamical law — sidestepping the usual trade-off between rigid physics-based models and purely implicit attention-based ones. It also provides a group-theoretic characterization showing exactly what constraints a homomorphic Lorentzian positional encoding must satisfy, and an analysis of why an analogous Galilean construction faces an obstruction.
Real-world applications:
- Structure-based drug discovery: Protein–ligand binding dynamics prediction, the exact task evaluated here on MISATO, is directly relevant to understanding how candidate molecules move and settle into binding pockets.
- Video forecasting for monitoring and safety: The tested KTH task involves human-action video, a proxy for surveillance, sports analysis, and traffic-scene forecasting.
- Embodied and robotic dynamics: The authors name embodied dynamics as a natural extension, where an agent's observations carry both time stamps and spatial positions.
- Weather and climate forecasting: Also named by the authors as a target domain, where spatiotemporal fields evolve across scales and geospatial coordinates.
- Generative spatiotemporal models: The event-based token formulation could serve as a positional backbone for generative video or trajectory models.
Industry relevance. The method requires only a lightweight modification to query–key transformations and preserves the standard attention interface, so it can be dropped into existing efficient attention pipelines. The reported parameter efficiency — 2.5M parameters on KTH versus the baselines' 12.2M to 39.8M — makes it attractive where inference cost or deployment constraints matter. The claimed statistical consistency gains on molecular benchmarks (Stability scores of 85.45 on MISATO-100, 86.01 on MISATO-1000, and 86.29 on MISATO-Full) are the kind of property that matters in simulation-based discovery workflows, where unstable predictions are costly.
Future Directions
-
Data-adaptive geometry: The authors note as a limitation that MinkowskiPE uses a fixed geometric bias rather than learning the geometry from data, and flag data-adaptive geometric structures as a direction for future work. Different applications may benefit from different inductive biases.
-
More expressive multi-parameter Lorentz constructions: The current construction uses learned scalar projections followed by one-parameter Lorentz boosts. Richer multi-parameter constructions could capture more interactions among temporal and spatial directions, but they require mutually commuting Lorentz generators and are therefore substantially more constrained — a tension the paper leaves open.
-
Alternative symmetry groups: The paper analyzes a Galilean construction in Appendix C and finds that exact Galilean invariance prevents the pairwise interaction from retaining the Galilean-invariant spatial component of the relative group element. Whether other classes of Galilean positional encodings can avoid this obstruction is an open question, and Galilean symmetry is noted as a natural candidate for non-relativistic physical settings.
-
Broader domains: Extending the framework to embodied dynamics, weather forecasting, and generative models is explicitly suggested, along with exploring how the same formulation behaves under different spatiotemporal tokenizations.
Target Audience
Researchers and practitioners working on spatiotemporal modeling, video prediction, and machine learning for molecular simulation, particularly those interested in geometric or symmetry-aware inductive biases for Transformers. It will also interest readers working on positional encoding design more broadly, and anyone building efficient attention architectures who wants a physically motivated positional mechanism that preserves the standard attention interface. The group-theoretic appendices are aimed at readers with a background in Lie groups or Lorentz geometry, while the experimental sections are readable by anyone familiar with standard sequence-prediction benchmarks.
Authors’ abstract
Modeling spatiotemporal coupling is a key challenge in building physical intelligence across scales, from microscopic to macroscopic. Existing models capture such structure broadly through physics-motivated dynamical formulations or learning-motivated architectures. The former provide stronger priors but may constrain flexibility, whereas the latter are more flexible but leave the spatiotemporal coupling largely implicit. We therefore seek an approach that combines flexible learning with an explicit geometric bias for jointly modeling time and space. To this end, we propose Minkowski Positional Encoding (MinkowskiPE), which uses joint temporal and spatial coordinates to parameterize Lorentz transformations applied to query and key features. With MinkowskiPE, the query-key attention score depends on position only through the relative spacetime displacement between the two tokens and is therefore invariant to global translation of the coordinates. This paradigm retains the standard dot-product attention interface and remains compatible with efficient attention implementations. We evaluate MinkowskiPE on microscopic molecular dynamics and macroscopic video prediction tasks, achieving the best results on all nine multi-trajectory molecular evaluations and reducing KTH video-prediction MSE by 9.9% relative to the best baseline while using roughly one-tenth as many parameters.