Skip to content
AI.info

Research

Can Image-To-Video Models Simulate Pedestrian Dynamics?

Recent high-performing image-to-video (I2V) models based on variants of the diffusion transformer (DiT) have displayed remarkable inherent world-modeling capabilities by virtue of training on large sc

Can Image-To-Video Models Simulate Pedestrian Dynamics?
arXiv
2510.17731
Published
2025-10-20
Authors
Aaron Appelle, Jerome P. Lynch

Authors’ abstract

Recent high-performing image-to-video (I2V) models based on variants of the diffusion transformer (DiT) have displayed remarkable inherent world-modeling capabilities by virtue of training on large scale video datasets. We investigate whether these models can generate realistic pedestrian movement patterns in crowded public scenes. Our framework conditions I2V models on keyframes extracted from pedestrian trajectory benchmarks, then evaluates their trajectory prediction performance using quantitative measures of pedestrian dynamics.

Read the original paper