Skip to content
AI.info

Research

SPW-Nav: A Streaming Panoramic World Model for Language-Guided Navigation

Overview Research area: Computer vision / generative world models — specifically language-guided panoramic (360°) video generation for navigation. Technical level: Advanced. The paper assumes familiar

SPW-Nav: A Streaming Panoramic World Model for Language-Guided Navigation
arXiv
2610.08941
Published
2026-10-06
Authors
Yunheng Liu, Ziqi Cai, Siqi Yang, Yimu Wang, Minggui Teng, Jiaming Tan, Shuchen Weng, Erwin Wu, Kaipeng Zhang, Boxin Shi

AI summary

Overview

  • Research area: Computer vision / generative world models — specifically language-guided panoramic (360°) video generation for navigation.
  • Technical level: Advanced. The paper assumes familiarity with diffusion transformers, SE(3) camera poses, equirectangular projection, few-step distillation, and autoregressive streaming generation.
  • Scope: SPW-Nav is a streaming panoramic world model that reads a movement instruction in the panorama it previously generated, converts it into a camera trajectory, and streams one minute of 2K 360° video in real time from a single panorama.

What This Paper Is About

Existing panoramic video generators follow trajectories that are fixed in advance, and interactive world models are driven by low-level controls such as keyboard inputs in flat perspective views. Neither can understand an instruction that refers to scene content, such as "walk to the door behind the camera." The paper's goal is to close this semantic gap: to build a model that interprets each language instruction inside the panorama it has just generated and continues a 360° video stream accordingly, while switching instructions on the fly.

Key Contributions

  1. SPW-Nav, the first streaming panoramic world model for language-guided navigation. It generates one-minute 2K 360° streams (1024 × 2048) at 26.6 FPS from a single panorama. The authors also build SPW-NavSet, 233 hours of panoramic video with camera trajectories and verified instructions.
  2. Spherical rotation decoupling (SRD) and pose-aligned conditioning (PAC). SRD applies the commanded rotation exactly by resampling the sphere so the video backbone only models translation; PAC expresses translation relative to the preceding chunk's first frame, keeping pose inputs bounded however far the navigation travels.
  3. Noise-tail trajectory learning (NTL) and opposite-trajectory guidance (OTG). NTL concentrates training and distillation on high noise levels where pose is the dominant motion cue; OTG contrasts the commanded translation with its reverse at the coarsest scale to strengthen translation following.
  4. A streaming generator. The video backbone is made autoregressive with a multi-term memory (recent latents, two earlier ones, sixteen older ones, plus a fixed first-frame anchor) and distilled to two steps per pyramid scale.

Main Findings

  • Language-guided navigation (SPW-NavBench, Table 1): SPW-Nav reaches a TDE of 56.66° against at least 62.07° for baselines, an FVD of 351.6 against at least 628.8, and an RE of 3.02° that is less than one quarter of any baseline's. It leads on every measure except ATE (6.28), where PanoWorld's keyword-matched presets supply a predefined trajectory shape while SPW-Nav infers the trajectory from the instruction and scene.
  • Navigation fine-tuning matters: replacing the fine-tuned backbone with its base model raises TDE to 100.74° and lowers MA to 0.028.
  • Camera-controlled generation (Table 2): Given ground-truth trajectories, SPW-Nav rotates 13.59° for a commanded 14.01°, while the panoramic baselines rotate by less than 2° and NWM overshoots. Its RE (1.84°) is less than one seventh of PanoWorld-A.'s. It achieves the best video quality on all three measures (PSNR 16.48, LPIPS 0.382, FVD 308.0). PanoWorld-A. reaches a lower ATE (3.93 vs 5.00) and higher MA (0.726 vs 0.631), which the authors attribute to its specialization in fixed-heading 960 × 480 clips versus SPW-Nav's 2K arbitrary-turn streaming.
  • Navigation accuracy: On SPW-NavBench's 40 explicit and 60 goal instructions, the backbone interprets all explicit instructions correctly and localizes 62% of visible targets. On the held-out test scene, fine-tuning raises accuracy from 0.663 to 0.925 for explicit instructions and from 0.154 to 0.692 for goal instructions. Without the panorama, goal localization drops to zero. On motion instructions it matches the commanded direction for 98% of instructions with a mean distance error of 0.03 m.
  • Long-horizon consistency (Table 3): Across 40 closed routes of twelve chunks, SPW-Nav closes the loop with a rotation error of 1.62° versus 8.91° for HunyuanWorld-Voyager (next best) and up to 39.52° for PanoWorld-A., and reaches the lowest translation closure error at 30.7%. Matrix-Game 3.0 has a higher memory gain (0.249 native perspective, 0.204 as a panorama, vs 0.173) but its closure rotation error exceeds 22° in both settings.
  • Instruction switching: In 20 scenes with four consecutive instructions spanning twelve chunks (about 25 seconds), switches cause no visible jump — inter-frame LPIPS across instruction boundaries is no higher than within segments. FVD is higher after the first segment (224.9 in segment 1, rising to 435.5 in segment 4).
  • Ablations (Table 4): Adding OTG cuts TDE from 43.53° to 24.30°. With OTG disabled, removing NTL raises RE from 0.99° to 26.58°, drops PSNR by more than 3 dB, and raises FVD from 395.6 to 512.2 — so NTL is needed for both motion accuracy and OTG's benefit. Without SRD (also trained without NTL), RE rises from 1.84° to 31.31° and FVD from 308.0 to 596.5, although that comparison does not isolate SRD from NTL.
  • Versus perspective generators (Table 73): On a 60° crop, SPW-Nav has the lowest rotation errors (RE 2.23°, RPE_R 0.321) and second-best video quality (PSNR 17.12, LPIPS 0.352, FVD 278.8), behind CameraCtrl's PSNR of 17.31. HunyuanWorld-Voyager follows translation more accurately through depth-based reprojection.
  • Speed: With eight H200 GPUs and sparse attention, SPW-Nav generates 26.6 FPS, faster than its 16 FPS playback rate, so new instructions can be issued while the stream plays.

Methodology in Plain English

Step 1 — Turn language into motion. A vision-language backbone (Qwen3-VL-8B with rank-16 LoRA on its language model's attention layers and a frozen vision encoder) looks at the last panorama the system generated plus the new instruction. Every instruction becomes one of two things: a displacement target (up to three waypoints given in meters and degrees relative to the current camera, which covers multi-step movement, stopping, and corrections) or a goal target (a bounding box around the named object or place, plus an approach distance). For a goal, the camera turns to the bearing of the box's horizontal center and then moves forward by the approach distance. Either target is converted into an SE(3) camera trajectory interpolated over the segment.

Step 2 — Handle rotation exactly, not by synthesis. A panoramic turn only rearranges what is already visible; it reveals no new content. So the model rerenders the panorama on the sphere with bilinear sampling that wraps the longitude seam, applying the rotation outside the network. Training frames are resampled to the first frame's orientation, which reduces the trajectory to pure translation. At inference, the rotation is restored after decoding. This means the generator never has to learn to synthesize a turn.

Step 3 — Keep pose inputs small. World-frame poses grow with distance traveled and differ by location. Pose-aligned conditioning instead states each chunk's poses relative to the first frame of the preceding chunk, which sits in memory. That frame's pose is therefore always visible and nearby, so the inputs stay in the training range and a relative translation of (1, 0, 0) always means "move right in the current view." Each of the 9 latent frames per chunk receives a flattened 3 × 4 pose vector in R¹², injected through a zero-initialized linear layer in every transformer block so the pretrained backbone is unchanged at initialization.

Step 4 — Train for the noise levels that matter. Noise-tail trajectory learning draws a fraction λ = 0.8 of noise levels from the tail [0.9, 0.98], because at high noise the sample shows almost no layout and the pose becomes the main signal for where content should move; at low noise the layout is already visible and the pose adds little. This applies both to the autoregressive teacher and to the few-step student. Opposite-trajectory guidance (κ = 2) contrasts the model's velocity under the commanded pose with its velocity under the negated pose at the coarsest scale, cancelling shared components such as scene appearance.

Step 5 — Stream. The backbone is Wan2.2-TI2V-5B fine-tuned with rank-256 LoRA, producing 1024 × 2048 video at 16 FPS playback. It is trained without language on rotation-free videos in autoregressive, pyramid, and trajectory-conditioned stages, then distilled into a few-step student via ODE regression and distribution matching distillation with a critic. Chunks are 33 frames, denoised from coarse to fine over three pyramid scales at two steps each (four for the first chunk) — six network evaluations per chunk after the first, plus two for guidance. Because generation runs in the rotation-free orientation and each remembered panorama covers all directions, rotation never disturbs the memory.

Data. SPW-NavSet combines 8,527 indoor videos rendered with Habitat from 29 indoor scene configurations (three from HM3D, ten from Replica, sixteen from ReplicaCAD) with 9,609 real outdoor MUGEN clips from 2,856 source videos, for 18,136 clips and about 233 hours at 30 FPS. Clips whose camera stays static or only rotates in place are removed and the rest are resampled to be rotation-free. Instruction annotations combine explicit instructions derived from trajectories with goal instructions verified by two vision-language teachers, spanning 68 unique goals from four indoor scenes. SPW-NavBench holds 100 real-world outdoor clips (from 94 MUGEN source videos, each with an 81-frame window), with 40 explicit and 60 goal instructions, sharing neither clips nor source videos with any training data.

Why This Matters

Impact on research. The paper argues that rotation in a panoramic world should not be synthesized at all — it should be resampled exactly — and that this decoupling frees the generative model to focus on translation. It also shows that the noise levels a model trains on are a lever for motion fidelity, not just image fidelity. The release of SPW-NavSet (333 hours of trajectory-annotated panorama) and SPW-NavBench gives the community a benchmark for language-driven panoramic navigation, where previously no panoramic generator understood semantic movement instructions.

Real-world applications:

  • Interactive 3D scene exploration and virtual reality, where a user or agent says where to go and the surrounding world streams around them.
  • Embodied agent training, supplying synthetic but instruction-following egocentric panoramas for robots and agents to learn from.
  • Content creation and virtual production, generating one-minute 2K 360° footage from a single input panorama without a captured trajectory.
  • Scene previewing and digital twins, letting a user walk a reconstructed environment using natural language rather than a prerecorded camera path.

Industry relevance. The system generates at 26.6 FPS on eight H200 GPUs, which is faster than its own 16 FPS playback rate — the threshold for genuinely interactive use, where a new instruction can be given mid-stream. Because the video backbone consumes camera trajectories rather than language, the interface is flexible: motion targets could come from clicks on the panorama or external planners without retraining. The explicit pose representation also means generation and navigation can be learned from complementary, separately curated data.

Future Directions

  • Distant revisit consistency. The authors note that consistency on distant revisits and the trajectory shape of translation remain open, and suggest a pose-indexed 3D memory to make revisits more reliable — though better memory mechanisms operating inside the video model itself might achieve the same.
  • Improving translation trajectory shape. More diverse training trajectories could improve how well the generated path matches the commanded path, where PanoWorld-A. currently achieves lower ATE and higher MA in the camera-controlled setting.
  • Alternative motion interfaces. Since the video backbone takes camera trajectories rather than language, motion targets could be supplied by clicks on the panorama or by external planners.
  • Narrowing the specialization–generality trade-off. The paper frames the ATE/MA gap versus PanoWorld-A. as a trade-off between a specialist in fixed-heading 960 × 480 clips and SPW-Nav's general 2K arbitrary-turn streaming; closing that gap is an implicit next step.

Target Audience

Readers who will benefit most are researchers and engineers working on generative world models, panoramic and 360° video synthesis, and vision-language navigation — especially those interested in streaming or interactive video generation and in how camera pose should be represented inside a diffusion backbone. The paper is also relevant to practitioners building VR exploration tools or embodied-agent training pipelines, and to benchmarking researchers who need a trajectory-annotated panorama dataset such as SPW-NavSet and SPW-NavBench. It is an advanced paper: it assumes fluency with diffusion transformers, SE(3) pose algebra, equirectangular projection, and few-step distillation.

Authors’ abstract

Language-guided panoramic video generation benefits various downstream applications, such as interactive 3D scene exploration, virtual reality experiences, and embodied agent training. Existing panoramic generators follow predefined trajectories, and interactive world models act through low-level actions in perspective views. We propose SPW-Nav, a streaming panoramic world model that understands movement instructions and streams one minute of 2K 360-degree video in real time from a single panorama. SPW-Nav interprets each instruction in the previously generated panorama as camera motion. Spherical rotation decoupling applies rotation exactly on the sphere, pose-aligned conditioning keeps translation inputs bounded over long streams, and a multi-term memory with a few-step generator continues the scene as instructions change. We also build SPW-NavSet, panoramic videos with camera trajectories and verified instructions. Driven by language, SPW-Nav outperforms prior panoramic generators in camera-following accuracy and video quality, and supports on-the-fly instruction switching.

Read the original paper