Skip to content
AI.info

Research

PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation

PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation Overview Research area: Computer vision — video generation evaluation, pedestrian/crowd simulation, and world-model benchmarkin

PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation
arXiv
2510.20182
Published
2025-10-23
Authors
Aaron Appelle, Jerome P. Lynch

AI summary

PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation

Overview

  • Research area: Computer vision — video generation evaluation, pedestrian/crowd simulation, and world-model benchmarking.
  • Technical level: Advanced (assumes familiarity with diffusion-based video models, multi-object tracking, structure-from-motion, and metric depth estimation).
  • Scope: The paper proposes PEDRA (PEdestrian Dynamics Realism Assessment), an evaluation protocol that tests whether text-to-video (T2V) and image-to-video (I2V) models behave as implicit simulators of multi-agent pedestrian dynamics, using twelve metrics across trajectory kinematics, social interaction, and video fidelity.

What This Paper Is About

Pedestrian simulation today depends on hand-crafted, expert-tuned models that are hard to scale and generalize, while video generation models now produce visually convincing footage and are being explored as general-purpose world simulators. Existing video benchmarks mostly judge single-subject realism, so nobody has systematically tested whether generated videos contain physically and socially plausible behavior for multiple interacting people. PEDRA fills that gap by extracting metric bird's-eye-view (BEV) pedestrian trajectories from generated videos and comparing their dynamics against real pedestrian datasets.

Key Contributions

  1. A dual-condition evaluation protocol. PEDRA benchmarks both I2V (conditioned on start frames from ETH and UCY, enabling direct comparison with ground-truth video through known homographies) and T2V (using a structured prompt suite spanning crowd density and interaction type), with a code release at https://github.com/aaronappelle/PEDRA.
  2. A trajectory extraction pipeline for synthetic scenes without known camera parameters. The method combines Visual Geometry Grounded Transformer (VGGT) for camera intrinsics/extrinsics and unscaled depth, Depth Pro for metric depth, RANSAC-based per-frame scale alignment, and an anthropometric height prior to recover meter-scale 2D BEV trajectories from pixel-space tracks.
  3. A twelve-metric realism suite. Metrics are grouped into trajectory kinematics (velocity, acceleration, distance), social interaction (collision, stationary, population, flow, nearest-neighbor distance), and video fidelity (disappearance, MOT confidence, 3D geometric confidence), with EMD-based distributional comparison to ground truth for I2V and reference ranges from ten public pedestrian benchmarks for T2V.
  4. A structured T2V prompt suite and benchmark of five state-of-the-art models. Prompts cover nine density/interaction combinations, with 20 LLM-generated (Gemini 2.5 Pro) scene descriptions per category and 5 repetitions each.

Main Findings

  • No single model dominates. In I2V, HunyuanVideo scores best on velocity (0.419), acceleration (0.639), and distance (0.288) EMD; LTX-Video produces the most trackable pedestrians (MOT confidence 0.503, disappearance EMD 0.130) and the best nearest-neighbor distance EMD (0.051); Open-Sora 2.0 best matches ground-truth flow (0.180); Wan2.1 best matches the low ground-truth collision rate (0.029).
  • Models respond to semantic prompts in T2V. Crowded prompts produce far higher population than sparse prompts (74.75 vs. 4.79 agents per frame on average), and directional prompts yield higher mean velocity (0.66 m/s) than multidirectional (0.55 m/s) or converging (0.47 m/s) prompts.
  • Collision rates rise with density and exceed real-world levels. Crowded scenes average 7.57% of pedestrian detections in collision versus 1.71% for sparse; the real-world reference is 1.19%. Open-Sora 2.0 has the lowest T2V collision rate among tested models but still exceeds real-world levels by more than twofold. Collisions are defined as agents within 0.1 meters.
  • Speeds are underestimated, accelerations are off. The real-world reference mean velocity is 0.91 m/s; model means range from 0.38 m/s (Open-Sora 2.0) to 0.80 m/s (LTX-Video), with LTX-Video closest to reference but showing unrealistically high accelerations (1.19 m/s²).
  • Population is overestimated by all models. Reference mean population is 13.77; model means run from 22.79 (Open-Sora 2.0) to 56.83 (Wan2.1).
  • HunyuanVideo's flow is physically implausible. Its mean flow rate exceeds the reference by more than three times.
  • Persistent failure modes: merging and disappearance. Pedestrians merge or vanish mid-trajectory, walk through each other as if other agents were absent, and become untrackable fluid-like pixel masses — most often in the "crowded" and "multidirectional" T2V prompts and for small background agents.
  • Scene adherence fails in specific cases. Models sometimes ignore negative prompts meant to keep the camera static, cause time-lapse streaking that blurs agents, animate parked cars in pedestrian zones, and — in HunyuanVideo's case — reduce agent counts over the 5-second clip as pedestrians vanish.
  • Some zero-shot scene understanding appears. Models occasionally show unexpected competence, such as an arriving train opening its doors to entering pedestrians, a converging funnel when a "crowd of shoppers" exits a market, and plausible walking paths on a pier that avoid water and furniture.
  • Architecture and training data shape results. All five models use the DiT backbone but differ in VAE compression: LTX-Video and Open-Sora 2.0 use high spatial compression, which may blur individual agents in dense crowds, while Wan2.1, CogVideoX, and HunyuanVideo use moderate compression that better preserves detail at higher cost — though CogVideoX still showed high distortion and inconsistency. WAN explicitly removes "crowded street scenes" and HunyuanVideo filters out videos with more than five people during fine-tuning for certain downstream tasks, which may explain weaker dense-scene quality.

Methodology in Plain English

The authors generate videos in two ways. For image-to-video, they take 530 unique non-overlapping start frames sampled at 5-second intervals from five ETH/UCY scenes (ETH, HOTEL, UNIV, ZARA1, ZARA2) and generate video from each frame plus a constant text prompt of "A stationary overhead view of pedestrian movement." For text-to-video, they define nine prompt categories along two axes — crowd density (sparse, moderate, crowded) and interaction type (directional, multidirectional, converging/diverging) — and ask Gemini 2.5 Pro to write 20 distinct scene descriptions per category, always requesting a stationary camera. They sample 5 repetitions per prompt, yielding 900 videos (1.25 hours) per model.

To turn pixels into trajectories, a multi-object tracker (FairMOT) finds and follows pedestrians, and the bottom-midpoint of each bounding box is treated as the ground contact point. For I2V, the known ETH/UCY homography matrices map those pixel tracks directly into BEV world coordinates. For T2V, no camera information exists, so the authors estimate camera intrinsics and extrinsics and an unscaled depth map with VGGT, estimate metric-scale depth with Depth Pro on keyframes, and then compute a per-frame scale factor by RANSAC-based robust alignment of the two depth maps, minimizing a Huber loss. They validate the scale with an anthropometric prior: each person's world height is computed as h_pixels · Z_cam / f_y, and if the mean height falls outside (1.4, 2.0) meters, they correct the scale factors to make the mean 1.7 m. Samples are discarded when the two depth estimates are inconsistent beyond a simple scale difference.

Realism is then scored with twelve metrics. For T2V, values are reported in real-world units and compared against reference ranges computed from ten public pedestrian benchmarks (ETH, UCY, PETS-2009, SDD, Grand Central, HERMES, KITTI, Edinburgh, Town Center, WildTrack, all accessed via OpenTraj). For I2V, distributions are compared to ground truth using Earth Mover's Distance normalized by ground-truth mean and standard deviation. To reduce label bias, ground-truth videos are reprocessed with the same tracking pipeline rather than using original manual annotations. Static camera viewpoints are filtered using pyramidal Lucas-Kanade optical flow.

Five state-of-the-art models with both I2V and T2V variants were tested — Wan2.1, CogVideoX1.5, HunyuanVideo, LTX-Video, and Open-Sora 2.0 — standardized to roughly 5-second clips (the maximum for Open-Sora 2.0 and HunyuanVideo), at a resolution as close as possible to the start image for I2V and 720p for T2V. All ran with suggested default hyperparameters on four NVIDIA H200 GPUs, with generation times varying between 2 and 8 minutes per video.

Why This Matters

  • Research impact: PEDRA shifts video-generation evaluation away from single-score visual quality and vision-language-model judging toward grounded, physically and socially specific measurement against real pedestrian data, establishing a baseline for multi-agent world modeling.
  • Real-world applications:
    • Autonomous driving perception and planning systems that need plausible pedestrian behavior for testing.
    • Emergency planning and evacuation analysis, where crowd dynamics must be physically credible.
    • Urban design and public-space layout studies that rely on simulated pedestrian flow.
    • Human-robot interaction and computer graphics, where believable crowds are needed for safe deployment or animation.
  • Industry relevance: Identifying that common data-filtering choices (removing crowded street scenes, excluding videos with more than five people) trade multi-agent fidelity for single-subject clarity gives model developers concrete, actionable signals about where current training pipelines limit crowd realism.

Future Directions

  1. Fix agent-level integrity. Merging, mid-trajectory disappearance, and agents passing through one another remain the dominant failure modes, especially for small background pedestrians — suggesting a link between representation scale and dynamic consistency that needs targeted solutions.
  2. Extend beyond 5-second clips. The current 5-second horizon precludes analysis of long-range navigation and how simulation fidelity degrades over time, both of which traditional crowd simulators capture.
  3. Reduce extraction pipeline noise. The multi-stage trajectory extraction can introduce label noise, particularly in T2V metric scale estimation, which future work could address with better depth-scale recovery or stronger anthropometric validation.
  4. Broaden model and scene coverage. The study is limited to a representative set of current models; expanding to more architectures and wider scene variety would clarify whether observed trade-offs are systematic.

Target Audience

Researchers and engineers working on video generation, world models, crowd simulation, and autonomous systems who need a rigorous way to test whether generated multi-agent behavior is physically and socially plausible. It is also valuable to practitioners in urban planning, robotics, and graphics who might consider replacing hand-tuned crowd simulators with video generation models, and to model developers who want to know which training and compression choices limit dense-scene realism.

Authors’ abstract

Pedestrian simulation traditionally relies on expert-tuned, hand-crafted models that limit scalability and generalization. Meanwhile, large-scale video generation models have achieved high visual realism across diverse settings, motivating exploration of their potential as general-purpose world simulators. Existing benchmarks primarily assess single-subject realism rather than scenes with multiple interacting people, leaving the plausibility of multi-agent dynamics in generated videos untested. We propose a rigorous evaluation protocol to benchmark text-to-video (T2V) and image-to-video (I2V) models as implicit simulators of pedestrian dynamics. For I2V, we leverage start frames from established datasets to enable direct comparison with ground truth videos, while for T2V we design a prompt suite covering varied crowd densities and interaction types. A key component is a method to reconstruct 2D bird's-eye view trajectories from pixel-space without known camera parameters. Our analysis shows that leading models exhibit effective priors for plausible multi-agent behavior, though issues such as merging and disappearing pedestrians reveal limits to their physical consistency.

Read the original paper