Skip to content
AI.info

Research

PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation

Overview Research area: Vision-and-language navigation (VLN), specifically navigation in continuous environments (VLN-CE) using vision-language models (VLMs), with panoramic (equirectangular, 360° × 1

PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation
arXiv
2609.34759
Published
2026-09-28
Authors
Zhen Wang, Changpeng Wang, Zhe Liu, Zhangyang Qi, Yuxiang Lu, Zimo Zeng, Donglian Qi, Xi Chen

AI summary

Overview

Research area: Vision-and-language navigation (VLN), specifically navigation in continuous environments (VLN-CE) using vision-language models (VLMs), with panoramic (equirectangular, 360° × 180°) RGB observations.

Technical level: Intermediate. The paper assumes familiarity with VLN benchmarks (R2R-CE, RxR-CE), VLM policy training, action-space discretization, and standard navigation metrics (NE, OS, SR, SPL, nDTW), but each component is explained at a level an informed graduate student can follow.

Scope: The paper diagnoses why simply swapping perspective cameras for panoramas fails to improve VLN, then proposes PanoVLN — a 4B-parameter, RGB-only system that adapts action prediction, training supervision, and visual representation to panoramic input, and validates it in simulation and on a quadruped robot.

What This Paper Is About

The authors ask whether a 360° panorama gives an instruction-following navigation agent an advantage over the perspective images used by most prior VLN work, on the intuition that seeing more of the scene should support better-informed decisions. They find that a direct input substitution — keeping the same model, training data, and execution procedure but feeding panoramas — yields only limited gains and can even reduce navigation performance. PanoVLN is their response: a set of targeted changes to how the model predicts actions, how training routes and instructions are constructed, and how panoramic visual features are represented.

Key Contributions

  1. Longer action-sequence supervision plus confidence-guided execution (CGE). The policy is trained to predict the next H = 18 expert actions, and at inference CGE uses per-action uncertainty to decide how many predicted actions to execute before re-observing, with an uncertainty budget B = 1.2 and a minimum execution of E_min = 4 actions.

  2. A decision-centric training dataset of 98K trajectories across 800 HM3D scenes. Routes are built to pass through frequent branching points (where at least two visible traversable paths lead to different areas, excluding the incoming path), paired with instructions generated and verified for motion consistency, choice grounding, and stop grounding, plus denser sampling at sustained-turn onsets and near termination.

  3. Geometry-aware panoramic visual representation. The model fuses the VLM's semantic tokens with geometric features from a frozen PanoVGGT encoder, aligned by resampling and grouping in ERP coordinates over the same regions as the current visual tokens, using residual fusion V̄_t = V_t + α·f_ψ(G_t) with α = 0.2 — adding spatial cues without adding visual tokens or depth input.

  4. State-of-the-art results with a 4B backbone and RGB-only input, plus real-world quadruped deployment. PanoVLN reports 77.3% SR on R2R-CE Val-Unseen and 78.0% on RxR-CE Val-Unseen, exceeding the previous state of the art by 11.9 and 8.7 percentage points, and transfers to a Unitree Go2 without scene-specific fine-tuning.

Main Findings

  • Panorama input alone is not enough: Replacing perspective images with ERPs under the same training and inference setup "yields only limited gains and can even reduce navigation performance," which motivates all three adaptations.

  • Panoramic policies prefer much longer prediction horizons than perspective policies: In the horizon study (fixed VLM, visual-token budget, and random-start sampling; 4,000 updates; geometry-free policies), ERP policies trail perspective policies at short horizons but overtake them as targets lengthen. Perspective performance peaks at H = 4, while ERP favors H = 18. ERP improves from H = 8 to H = 18 with execution fixed at six actions.

  • Longer supervision is not the same as longer execution: With a geometry-free H = 18 checkpoint, CGE achieves NE 4.10, OS 71.0, SR 66.6, SPL 61.5 on R2R-CE and NE 4.06, SR 66.5, SPL 56.4, nDTW 70.8 on RxR-CE, outperforming fixed prefixes of 1, 6, 12, and 18 actions and random 1–18 action prefixes. For comparison, fixed 18-action execution drops to SR 57.2 (R2R-CE) and 54.9 (RxR-CE).

  • Turn- and termination-aware sampling improves termination reliability: Against random starts, the authors' sampling raises SR from 56.4 to 62.8 and SPL from 49.5 to 57.2 on R2R-CE, with NE falling from 5.37 to 4.40 and OS roughly unchanged (70.7 vs 69.9) — which the authors read as more reliable stopping in the goal region.

  • PanoVGGT is the best geometry encoder tested: Against no encoder, UniK3D, DA², and DAP under the same fusion design, token budget, and CGE, PanoVGGT gives SR 68.6 on R2R-CE and 67.1 on RxR-CE, the highest of the four encoders. Alternative encoders show mixed effects.

  • Contributions of the data sources are cumulative: Adding DAgger to R2R-CE and RxR-CE raises SR by 5.3 and 7.0 points and SPL by 4.2 and 6.7 points; adding the constructed decision-centric dataset then improves SR by a further 3.4 and 3.9 points and SPL by 2.7 and 2.0 points, reaching the full-model numbers (SR 77.3 / 78.0; SPL 70.6 / 65.9; RxR-CE nDTW 73.3, up from 71.4).

  • Instruction quality matters independently of routes: ScaleVLN-rewrite, which keeps the original routes but regenerates instructions with the authors' instruction pipeline, improves over ScaleVLN at matched trajectory counts; the full pipeline yields further gains across training scales.

  • Results are not tied to a single backbone: With all four training sources, PanoVGGT, CGE, visual history, and H = 18 fixed, Qwen3-VL-4B already reaches SR 76.2 (R2R-CE) and 75.7 (RxR-CE); swapping in Qwen3.5-4B improves SR by 1.1 and 2.3 points. Qwen2.5-VL-7B reaches SR 75.9 and 76.3.

  • Real-world efficiency improves on a quadruped: On trial averages, PanoVLN records 86.4 s navigation time, 25.7 cm/s speed, 13.6% waiting, 5.4 pauses, 7.4 policy calls, and 1.08 s latency — versus NaVid (135.2 s, 12.9 cm/s, 23.2%, 15.7, 29.4, 0.90), NaVILA (194.0 s, 8.2 cm/s, 24.7%, 29.6, 41.3, 1.05), StreamVLN (117.8 s, 13.3 cm/s, 27.1%, 8.3, 34.3, 0.58), and JanusVLN (412.9 s, 5.3 cm/s, 39.2%, 95.3, 101.7, 1.32). The paper notes JanusVLN requests a prediction for each action, while StreamVLN replans every four actions.

  • Substantial direction changes are common in reference paths: Of 54,457 turns of at least 30° in R2R-CE and RxR-CE, the 45°–180° categories account for 53.3%, meaning many outgoing directions fall outside a forward 90° view — context that ERP observations retain before rotation.

Methodology in Plain English

The team starts from a conventional VLM navigation policy that takes an instruction, the current camera image, and a sampled history of images, and autoregressively emits discrete actions (forward 25 cm, left/right 15°, stop). They first build a "panoramic baseline" by swapping in equirectangular panoramas with the agent's heading centered at the image middle and the left and right edges joined behind the agent — and observe that this alone does not help.

Three changes follow. First, rather than training the model to predict one or a few next actions, they train it to emit 18 actions at a time, which lets a single panorama supervise route choices and subsequent movement. Because later predictions in a sequence are less trustworthy, they compute each predicted action's uncertainty as the negative log probability of the chosen action, accumulate uncertainty across the prefix, and execute the longest prefix whose cumulative uncertainty stays within a budget (B = 1.2), always executing at least four actions before re-observing.

Second, because branching points are sparse in existing VLN data, they build new training data: they partition walkable space in 800 HM3D scenes using the navigation mesh, define branching points as having at least two visible traversable paths to different areas excluding the incoming path, sample endpoints in different areas, retain routes through branching points, filter out candidates with mesh holes or incomplete geometry, remove near-duplicates, and convert surviving routes into primitive action sequences with replay verification. Instructions are generated by Qwen3.8-27B from first-person video with the expert path marked, plus eight-view compass images at branching and arrival segments, then verified for motion consistency, choice grounding, and stop grounding. Training states are sampled on a stride-six grid (for H = 18, reducing overlap between adjacent targets from 17 to 12 actions), with extra states at sustained-turn onsets and near termination.

Third, since VLM visual features mostly capture appearance rather than layout, they run a frozen PanoVGGT encoder on the same RGB panorama, align its features to the same ERP regions covered by the VLM's current tokens, project them through an MLP, and add them residually to the visual tokens — keeping token count and order unchanged, and training the VLM and projector while freezing the geometry encoder.

Evaluation uses Habitat in Matterport3D scenes on R2R-CE and RxR-CE Val-Unseen, reporting NE (meters) and OS, SR, SPL, and nDTW on a 0–100 scale, with OS defined as coming within 3 m of the goal and SR additionally requiring stopping there. The default configuration is Qwen3.5-4B with PanoVGGT and H = 18 with CGE, trained for one epoch with Fused AdamW, effective batch size 128 (8 GPUs × 4 examples × 4 accumulation steps), vision-encoder learning rate 2×10⁻⁶ and language/merger/projector learning rate 2×10⁻⁵, cosine schedule with 3% warmup, weight decay 0.01, gradient clipping at max norm 10, bfloat16, seed 42, and gradient checkpointing enabled. Current ERPs are resized to 960×480 and history frames to 448×224 before cropping 20° from each pole, with geometry encoder input at 1036×518 over the full panorama, and up to 10 earlier frames drawn from the latest 100 observations.

Real-world tests deploy on a Unitree Go2 with an Insta360 X5 mounted 1.5 m above ground and a remote RTX 3090, using about 12 GB of GPU memory and averaging 272 ms of network overhead per call, with synchronous execution (the robot waits for a response, executes, then requests the next prediction). They compare NaVid, NaVILA, StreamVLN, JanusVLN, and PanoVLN on 20 shared instruction–route pairs per setting, without scene-specific fine-tuning, in Hallway (successive turns), Office (clutter and room transitions), and Campus (gardens, sports fields, plazas).

Why This Matters

The paper's central claim is not that panoramas are better cameras, but that broader visibility should change how a navigation policy is supervised and executed — a design principle that applies beyond this specific system. It also reports a 4B-parameter, RGB-only model that surpasses larger or multi-modal baselines on R2R-CE and RxR-CE Val-Unseen, and demonstrates transfer from indoor simulation training to indoor and outdoor quadruped deployment without scene-specific fine-tuning. Notably, the authors disclose in Appendix A that GPT-5.6

Authors’ abstract

Recent vision-language models (VLMs) have advanced vision-and-language navigation (VLN), enabling models to predict navigation actions from visual observations and language instructions. In this work, we explore VLN with panoramic observations and introduce PanoVLN. The motivation is straightforward: more complete visual context should enable better-informed navigation decisions. For example, a panorama can reveal a passage outside a perspective camera's field of view, allowing the model to identify the intended route without additional exploration. However, we find that simply replacing perspective images with panoramas yields only limited gains. Our diagnosis suggests that fully exploiting wider visibility requires modifications to action prediction, training supervision, and visual representation. First, wider visibility supports longer-horizon action planning. We make the model predict longer action sequences, enabling larger turns and subsequent movement from a single panorama. Specifically, we introduce a confidence-guided execution (CGE) strategy that dynamically determines how many predicted actions to execute before replanning. Second, wider visibility also brings more complex route choices. We therefore construct training routes with frequent branching points and clear instructions to provide targeted supervision for route selection. Third, panoramic navigation requires understanding spatial relationships across viewing directions, beyond recognizing individual landmarks. We combine semantic and geometric features from RGB panoramas to capture both scene content and spatial layout without adding visual tokens. With a 4B backbone and RGB-only input, PanoVLN surpasses the previous SOTA by 11.9% and 8.7% in success rate on R2R-CE and RxR-CE Val-Unseen. Real-world experiments on a quadruped further demonstrate faster navigation with fewer pauses than prior VLN methods.

Read the original paper