Skip to content
AI.info

Research

Stereo4DWalker: Learning 4D-aware Embodied Urban Navigation from Internet Stereo Videos

Overview Research area: Computer vision and embodied robotics — specifically vision-based urban navigation, stereo/4D scene understanding, and learning from Internet video. Technical level: Advanced.

Stereo4DWalker: Learning 4D-aware Embodied Urban Navigation from Internet Stereo Videos
arXiv
2512.10956
Published
2025-12-11
Authors
Wentao Zhou, Xuweiyi Chen, Vignesh Rajagopal, Jeffrey Chen, Rohan Chandra, Zezhou Cheng

AI summary

Overview

Research area: Computer vision and embodied robotics — specifically vision-based urban navigation, stereo/4D scene understanding, and learning from Internet video.

Technical level: Advanced. The paper assumes familiarity with transformers, depth estimation, dense point tracking, visual odometry, and navigation benchmarks.

Scope: This paper proposes Stereo4DWalker, a stereo-input navigation transformer that injects explicit 4D (geometry plus motion) structure from pretrained vision foundation models, and accompanies it with StereoUrban, a 600-clip, 60-hour Internet stereo navigation dataset with automatically generated action labels.

What This Paper Is About

Most urban navigation models take a single monocular camera image and directly predict a robot action, hoping that understanding of 3D geometry and scene motion will emerge implicitly from large amounts of pixel-to-action supervision. That supervision is expensive and hard to collect, and monocular input introduces depth-scale ambiguity that amplifies action-label noise. The goal of this work is to instead feed stereo imagery into a navigation transformer and explicitly construct 4D representations of geometry and motion using pretrained vision foundation models, so that accurate navigation requires far less labeled data.

Key Contributions

  1. StereoUrban dataset. A large-scale Internet stereo navigation dataset of pedestrians across global metropolitan cities: 600 high-resolution stereo first-person walking clips totaling 60 hours of footage, spanning locations such as San Francisco, Madrid, and Tokyo. The authors state it is the first large-scale, real-world stereo navigation benchmark for urban scenes collected from Internet videos, and describe an automated filtering and scalable action annotation pipeline (Qwen2-VL for filtering, MAC-VO stereo visual odometry for pseudo-labels).

  2. Stereo4DWalker model. A 4D-aware embodied urban navigation framework that takes rectified stereo (or monocular) frames and integrates explicit 4D scene structures from vision foundation models into a navigation transformer via a 4D-conditioned attention mechanism.

  3. Data-efficient transfer. A demonstration that the model surpasses state-of-the-art navigation performance using only 1.5% of the training data, along with aligned improvements on established benchmarks and real-world robot deployments.

Main Findings

  • Surpasses state of the art with 1.5% of the training data. The paper reports that Stereo4DWalker surpasses prior state-of-the-art performance while using only 1.5% of the training data, and that the model empowered by mid-level vision capabilities can surpass a CityWalker model trained for over 2,000 hours using only 30 hours of data with the same number of epochs.
  • Fine-tuned monocular benchmark gains. On the monocular benchmark (CityWalker benchmark), the fine-tuned model improves MAOE by an average of 4–13% and arrival rates by 1–19%, reaching the best scores in every scenario except "turn." Reported "All" values: L2 1.00 m, MAOE 11.0 degrees, Arrival 88.9%, versus CityWalker at 1.07 m, 11.5 degrees, and 87.8%.
  • Monocular variant already beats monocular baselines on the stereo benchmark. Without stereo, the authors' monocular model reduces average L2 error by 17–73%, MAOE by 11–48%, and increases arrival rates by 3–24% relative to the monocular baselines.
  • Stereo adds further gains. Stereo training reaches the best reported numbers, reducing overall L2 error by 18–73%, MAOE by 22–54%, and arrival rates by 3–25% compared to monocular setups. The stereo variant's "All" scores are L2 0.62 m, MAOE 2.9 degrees, Arrival 94.6%, versus CityWalker's 0.76 m, 3.7 degrees, and 91.2%.
  • Turns remain the hard case. The "turn" scenario is an exception: the authors hypothesize the relative weakness on turns arises from data imbalance toward straight segments and amplification of small orientation errors during sharp heading changes, and they note that stereo arrival in "turn" lags behind monocular baselines despite comparable orientation error.
  • Real-world deployment gains. With 14 trials per motion pattern, the fine-tuned model achieves success rates of 73.8% overall, 85.7% forward, 71.4% left turn, and 64.3% right turn, compared with CityWalker at 50.0% overall, 57.1% forward, 42.9% left turn, and 50.0% right turn. Fine-tuned Stereo4DWalker improves performance by an average of 23.8% over the Forward, Left turn, and Right turn scenarios.
  • Ablation confirms each 4D component helps. On the CityWalker teleoperation benchmark, average MAOE drops from 17.55 with no patch tokens, depth, or tracking, to 16.90 (-0.65) with patch tokens, 16.23 (-1.32) with patch tokens plus depth, and 15.77 (-1.78) with patch tokens plus depth plus tracking. The paper attributes a 3.7% improvement to using all patch tokens instead of a single [CLS] token, a further 4.0% reduction to adding depth, and an additional 2.8% reduction to adding tracking.
  • Modest inference overhead. Stereo4DWalker requires 2.89 GB of VRAM and 0.2 s per sample on an A100 GPU, versus CityWalker at 1.68 GB and 0.06 s per sample. The authors argue this is manageable because the model predicts five future waypoints spanning a five-second horizon, allowing a one-second inference interval.

Methodology in Plain English

The researchers treat urban navigation as position-goal waypoint prediction: given a short window of recent visual observations and positions plus a sub-goal waypoint, the model predicts the next N waypoints, and once the sub-goal is reached, the next waypoint from a global planner such as A* becomes the new target.

For the data, they harvested stereoscopic city-walking videos from the Internet. Because much online walking footage includes pauses, conversations with bystanders, or shopping and sightseeing, they used a vision-language model, Qwen2-VL, to review frames and captions and keep only clips showing egocentric, target-oriented forward locomotion. Trajectory labels were then generated automatically with stereo visual odometry, specifically MAC-VO, validated on teleoperated navigation sequences with LiDAR-based SLAM ground truth, so no manual annotation or language-based prompting is needed.

For the model, each frame is tokenized with DINOv2, but unlike prior navigation models that compress a frame to a single [CLS] token, Stereo4DWalker keeps all patch tokens to preserve fine-grained spatial structure. Depth comes from DepthAnythingV2 for monocular input and MonSter++ for stereo alignment; the depth map is encoded by a simple CNN into depth embeddings that are concatenated with the DINOv2 patch tokens. Temporal motion comes from CoTracker3 dense pixel tracks, processed by a track transformer inspired by TrackTention through three operations: cross-attention that pools image evidence into track features, self-attention along time to smooth each track's evolution, and a second cross-attention that writes updated track information back into the image tokens. The unified token sequence then goes through global self-attention (with the target token and trajectory tokens included) and target-token attention, and finally an arrival head predicts the probability of reaching the sub-goal while an action head predicts the next waypoints. Training minimizes a composite loss combining a waypoint loss, an arrival loss weighted at 1.0, and a direction loss weighted at 10.0.

Evaluation uses three metrics: maximum average orientation error (MAOE, the worst per-step heading error over the horizon, averaged across samples), arrival accuracy (the fraction of samples coming within r = 5 m of the target within K = 5 steps), and Euclidean L2 distance. Baselines are GNM, ViNT, NoMaD, and CityWalker, all of which were fed left images on the stereo benchmark, plus a monocular variant of the authors' own model for fairness. The real-world setup uses a Clearpath Jackal J100 robot communicating with a remote GPU server via FastAPI, with a ROS 2 low-level controller converting predicted short-horizon trajectories into velocity commands; a trial counts as successful when the robot arrives within 1 m of the target and stays there, and collisions count as failures.

Training data for the robot experiments was collected via teleoperation using a Clearpath Jackal equipped with an Ouster LiDAR and a ZED 2i stereo camera, with LiDAR-based SLAM providing ground-truth pose and wheel odometry providing continuous relative motion estimates.

Why This Matters

Impact on research: The paper argues that explicitly modeling geometry and motion through stereo, depth, and dense tracking gives navigation transformers strong inductive biases, so they need far less pixel-to-action supervision. It positions this as a counterpoint to the dominant end-to-end monocular pixel-to-action paradigm and as evidence that classical mid-level vision representations remain relevant for modern robot navigation foundation models.

Real-world applications:

  • Last-mile delivery robots operating in dense, unstructured pedestrian environments, which the introduction names as the motivating application.
  • Sidewalk-following and socially compliant navigation, including staying on sidewalks, obeying traffic signals, and maintaining appropriate interpersonal distance.
  • Urban robot platforms that need to navigate around crowds, crossings, and detours while carrying only passive, low-cost stereo sensing.
  • Teleoperation-to-autonomy pipelines, where expert demonstrations collected with a LiDAR-plus-stereo robot can be used to fine-tune a deployment policy.

Industry relevance: Stereo cameras are described as a practical source of metric geometry with passive sensing, low cost, and simple deployment, which matters for robots that cannot carry LiDAR-grade sensors. The paper also reports a concrete deployment story — a Clearpath Jackal J100, remote GPU inference over FastAPI, and ROS 2 control — plus an inference footprint of 2.89 GB VRAM and 0.2 s per sample, which is the kind of detail that determines whether a model can run on a real robot's compute budget.

Future Directions

  • Extending 4D-aware navigation beyond ground urban navigation to broader robotic tasks, which the authors explicitly say are not yet fully explored, including mobile manipulators and aerial robots.
  • Making the data curation pipeline more scalable so larger and more diverse datasets can be produced with lower processing overhead.
  • Addressing the persistent weakness on turning scenarios, which the authors attribute to data imbalance toward straight segments and to the amplification of small orientation errors during sharp heading changes, suggesting either more turning-rich data or turn-aware objectives.
  • Reducing the gap between stereo and monocular inference cost, since the stereo model's 2.89 GB and 0.2 s per sample is roughly double CityWalker's 1.68 GB and 0.06 s per sample.

Target Audience

Researchers and engineers working on embodied AI, robot navigation foundation models, and vision-based robotics who already understand transformers and depth estimation. It is also relevant to practitioners building autonomous delivery or urban service robots who care about data-efficient training and stereo-based perception, and to dataset builders interested in automatically mining and labeling navigation supervision from Internet video.

Authors’ abstract

Despite rapid progress, embodied navigation in dynamic and unstructured urban environments remains brittle. Most existing approaches directly map monocular visual inputs to actions through end-to-end pixel-to-action training, assuming that accurate spatiotemporal (4D) scene understanding will emerge implicitly. While appealing, this paradigm requires large amounts of pixel-to-action supervision that are difficult to obtain. This challenge is amplified in dynamic, unstructured settings, where robust navigation requires precise 4D scene modeling. To address these limitations, we present Stereo4DWalker, a 4D-aware embodied navigation model that leverages stereo inputs and explicitly builds structured 4D representations of geometry and motion. These 4D structures are integrated into the navigation transformer through simple yet effective 4D-conditioned attention layers. To support scalable training, we curate a large-scale stereo navigation dataset with automatically annotated actions from Internet stereo videos. Our experiments show that Stereo4DWalker surpasses state-of-the-art performance using only 1.5% of the training data, highlighting the effectiveness of explicit 4D visual modeling for data-efficient and robust urban navigation.

Read the original paper