Research
DroneWAM: Efficient World Action Model for Drone Visual Navigation
DroneWAM: Efficient World Action Model for Drone Visual Navigation Overview Research area: Computer vision and robot learning, specifically world-action models for autonomous drone (UAV) visual naviga

- arXiv
- 2609.33148
- Published
- 2026-09-27
- Authors
- Liang Yao, Fan Liu, Hongbo Lu, Wei Xu, Jianyu Jiang, Yijun Shen, Chuanyi Zhang, Pai Peng
AI summary
DroneWAM: Efficient World Action Model for Drone Visual NavigationOverview
Research area: Computer vision and robot learning, specifically world-action models for autonomous drone (UAV) visual navigation, with a secondary contribution in simulated dataset construction.
Technical level: Advanced. The paper assumes familiarity with latent world models, JEPA-style predictive representation learning, transformer-based planners, and preference optimization (DPO).
Scope: The paper proposes a latent world-action model that predicts future states in representation space rather than pixels, compresses visual tokens with a pretrained Resampler, and adaptively decides how many imagined rollout steps to perform per scene, evaluated on a newly built 6-DoF simulated drone navigation dataset.
What This Paper Is About
Drone visual navigation requires an agent to anticipate how candidate control commands will change future observations, and to act on those predicted consequences, all within tight onboard compute limits. Existing aerial world models and world-action models predict the future explicitly and run a fixed number of imagined rollout steps for every scene, so predictive reasoning becomes expensive and is allocated uniformly regardless of how ambiguous the current scene actually is. DroneWAM's goal is to make this predictive process both accurate and economical by compressing the representation being propagated and by spending more imagination only where a scene demands it.
Key Contributions
-
DroneWAM, an efficient world-action model for drone visual navigation that combines JEPA-based latent prediction, a pretrained Resampler that compresses dense encoder features into 128 latent tokens, and a preference-trained Gate for adaptive rollout depth. The paper positions this as a practical modeling framework for predictive navigation under constrained onboard computation.
-
DroneNav-6D, a simulated 6-DoF aerial navigation dataset built with an automatic data generation pipeline, containing synchronized RGB observations, flight commands, 6-DoF trajectories, and randomized wind disturbances across five simulated worlds. It is offered as a scalable resource for training and evaluating aerial world-action models beyond conventional 4-DoF motion (translation and yaw).
-
An empirical demonstration that DroneWAM achieves both higher navigation accuracy and lower inference cost than existing world-model and world-action baselines, in both open-loop trajectory prediction and closed-loop receding-horizon navigation.
-
Evidence that adaptive rollout generalizes beyond aerial navigation, shown by applying the same rollout strategy to the LIBERO robotic manipulation benchmark, where inference latency drops while action prediction accuracy slightly improves.
Main Findings
-
Best open-loop trajectory accuracy at the lowest latency. On DroneNav-6D, DroneWAM records ATE 1.1062, Rel.ATE 0.0578, RPE 0.3470, Rel.RPE 0.0180, and 483.29 ms inference. Compared with FastWAM (ATE 1.8366, Rel.ATE 0.1919, 712.33 ms), this is a 39.8% reduction in ATE and a 32.2% reduction in inference latency. Other baselines reported are DINO-WM (ATE 9.0923, 29,622.24 ms), V-JEPA 2-AC (ATE 27.4317, 64,617.72 ms), LeWM (ATE 4.2259, 654.37 ms), and NWM (ATE 3.7580, 403,827.78 ms).
-
Closed-loop gains without task completion. In closed-loop evaluation, DroneWAM reaches task progress 0.774 versus 0.510 for FastWAM and 0.500 for NWM, with the lowest trajectory errors and MSE (91.96) and latency 483.29 ms. The paper states explicitly that none of the evaluated methods completes the task, so the system is not yet ready for real-world deployment. Closed-loop ATE is 83.4461 (Ours), 88.6305 (NWM), 119.8063 (FastWAM).
-
Adaptive rollout cuts prediction depth and improves accuracy simultaneously. Average rollout depth falls from 8 to 4.58 while trajectory accuracy improves. Adaptive rollout lowers latency by 22.4% over fixed-depth rollout. In the strategy comparison, DroneWAM reaches Rel.ATE 0.0578 at an average depth of 4.578, while random stopping at a nearly identical average depth of 4.581 yields Rel.ATE 0.0722, and fixed depth 8 yields 0.0688. A "Latent Margin" baseline reaches 0.0887 at an average depth of 5.625.
-
Token compression is nearly lossless between 512 and 128 tokens. Compressing from 512 to 128 Resampler tokens reduces latency from 591.28 ms to 483.29 ms (an 18.3% reduction) with Rel.ATE changing from 0.0574 to 0.0578 and Rel.RPE from 0.0179 to 0.0180 — absolute increases of 0.0004 and 0.0001. Further compression to 16 tokens lowers latency to 405.16 ms but degrades accuracy to Rel.ATE 0.0659 and Rel.RPE 0.0198.
-
Component ablation confirms complementary savings. Adding the Planner takes Rel.ATE from 0.6621 to 0.0584 at 656.12 ms; adding the Resampler moves latency to 582.65 ms with Rel.ATE 0.0603; adding Adaptive Rollout reaches 483.29 ms with Rel.ATE 0.0578 and Rel.RPE 0.0180. The Resampler and adaptive rollout reduce computation from complementary spatial and temporal directions.
-
More prediction is not always better. Evaluating every sample at all rollout depths from K = 0 to K = 8 and grouping by the depth with lowest trajectory error shows that optimal depths vary substantially across samples rather than concentrating at the maximum depth.
-
Adaptive rollout transfers to LIBERO manipulation. Latency falls from 890.8 ms at fixed K = 8 to 411.9 ms, a 53.8% reduction, while normalized action error improves from 0.2685 to 0.2652 and gripper accuracy from 89.83% to 90.35%. The improvement over fixed K = 8 holds across all four LIBERO suites.
-
The preference-trained Gate does not visibly overfit to training environments. Under environment-level cross-validation, seen and unseen results stay closely aligned for Rel.ATE and Rel.RPE across all five environments, with a maximum absolute gap of 0.0034 and 0.0010 respectively, and unseen errors lower in seven of ten environment–metric comparisons.
-
Qualitative trajectories are more faithful. In four representative scenes, DroneWAM stays closer to the ground-truth trajectory than FastWAM, which shows larger directional and accumulated drift as prediction lengthens.
-
Dataset scale. DroneNav-6D contains 3,837 trajectory clips, 1,592,170 RGB observations, 63.03 hours of recorded flight, and 2,004.11 km of traveled distance, at a median recording rate of approximately 7.00 Hz, split into 3,453 training and 384 validation clips across five worlds (CITYBIM, NYC, RuralAustralia, SnappyRoads, Snowy Mountain).
Methodology in Plain English
Prediction in latent space instead of pixels. DroneWAM follows the JEPA idea: rather than generating future images, it predicts future representations. A frozen ViT-L encoder turns observations into features, and a lightweight Resampler compresses those dense features into 128 latent tokens. The Resampler is pretrained together with a small Decoder using feature reconstruction and temporal consistency objectives; after pretraining the Decoder is discarded and the encoder and Resampler are frozen. Because the same compact latent is carried through every imagined step, this compression pays off repeatedly.
Predictor and Planner. Given the recorded context — observed latents, motion states, and executed action history — plus a latent of the goal image, the Predictor autoregressively predicts up to K_max future latent states. Future ground-truth states and actions are masked during training, so the Predictor only uses what would be available online. At each imagination depth, the Planner produces a complete H-step 6-DoF action plan (H = 8 in the implementation), which is integrated from the current pose into a predicted trajectory. World prediction is supervised by a temporally discounted latent loss, and the Planner by relative absolute trajectory error (rATE) and relative pose error (rRPE) summed over all imagination depths.
Learning when to stop. After training the world-action model, the researchers freeze it and score every rollout depth on every sample with a cost that combines trajectory error and rollout computation: rATE/ε_A + rRPE/ε_R + λ_c·(k/K_max). The lowest-cost depth becomes the preferred sample and the rest become rejected samples. A five-member marginal-value Gate, taking only inference-time information (observed latent, prediction and action-plan history, changes between adjacent predictions, recent motion, velocity, and depth), is then optimized with a DPO-style reference-adjusted reward. A disjoint calibration subset sets a stopping threshold η_k per depth. At inference, the rollout stops once the Gate's high-reward score crosses the threshold, and the highest-scoring plan seen before termination is the one executed — only its first action is run before the loop repeats in a receding-horizon manner.
Dataset construction. DroneNav-6D is generated with Unreal Engine 4.27 and AirSim. A multirotor with a forward-facing RGB camera ascends to a cruising altitude, waypoints sampled within scene-specific flight boundaries are connected into smooth 3D trajectories, and the flight is executed through velocity and heading commands while RGB observations, position, orientation, and issued commands are recorded. Random wind disturbances are injected, producing lateral drift, attitude changes, and controller corrections, which are captured in the paired command-and-pose records. Episodes are filtered on collisions, prolonged immobility, trajectory length, frame intervals, and action validity.
Training setup. Base world-action model training runs 30 epochs of AdamW in bfloat16 (peak learning rate 7.5 × 10⁻⁵, effective batch size 128), followed by five epochs of variable-depth Predictor–Planner adaptation, then 30 epochs of Gate training on eight A100-80GB GPUs. The paper reports that code and data will be released.
Why This Matters
Impact on research. The paper reframes efficiency as a property of the predictive process itself, not just of the backbone network. Two of its ideas are portable: compressing the latent that gets propagated repeatedly, and learning a policy over how much imagination to spend. The finding that per-sample optimal rollout depth varies widely, and that a preference-trained Gate beats random early stopping at identical average depth, is a direct challenge to the common assumption that a longer fixed horizon is strictly better. The DroneNav-6D dataset also addresses a specific gap: most established aerial benchmarks use 4-DoF motion over translation and yaw, which leaves roll and pitch — both of which directly change camera viewpoint and visual dynamics — poorly covered.
Real-world applications.
- Autonomous drone delivery and inspection in urban or industrial settings, where a drone must re-plan continuously from a forward-facing camera with limited onboard compute.
- Search and rescue in unstructured terrain such as snowy mountains or rural areas, where ambiguous scenes could benefit from deeper foresight while simple stretches of flight do not.
- Agricultural monitoring and survey flights, where long, roughly straight traverses dominate and shallow prediction is likely sufficient.
- Sim-to-real training pipelines: DroneNav-6D's automatic generation process, including wind disturbances, offers a way to pre-train predictive navigation models before physical testing.
Industry relevance. Inference latency and onboard compute are the practical bottlenecks for deploying predictive models on drones, and this paper treats milliseconds and GPU wall-clock time as first-class metrics alongside trajectory error. The latency reductions reported — 32.2% versus FastWAM in open-loop, 18.3% from token compression, 22.4% from adaptive rollout — are the kind of numbers that determine whether a model fits a flight controller budget. The demonstrated transfer to LIBERO suggests the adaptive-computation idea is not specific to aerial navigation and could be relevant to robotic manipulation more broadly.
Future Directions
-
Closing the gap to actual task success. The paper reports that no evaluated method completes the closed-loop task, and lists improving success rates and robustness as prerequisites for practical use. What changes to representation, training coverage, or control integration would move progress from 0.774 to full completion remains open.
-
Removing the fixed token budget and bounded horizon. The limitations section notes that DroneWAM's fixed token budget and bounded imagination horizon may need adjustment across scenes and computational budgets, and that performance depends on what the pretrained encoder and resampler retain. Making both adaptive is a natural extension of the adaptive-rollout idea.
-
Validation on physical platforms. All reported results are in simulation. The paper states that practical use requires validation on physical platforms, and no real-world flight results are reported here.
-
Broadening the Gate's supervision and evaluation. The Gate is trained on preference pairs and shown not to overfit across five simulated environments under environment-level cross-validation. Whether the same preference formulation holds across more diverse environments, task families, and compute budgets — and how sensitive it is to the cost weight λ_c and the calibration split — is not established.
Target Audience
Researchers and engineers working on world models, world-action models, and model-based reinforcement learning; drone and UAV autonomy practitioners concerned with onboard inference cost; and robot learning researchers interested in adaptive computation allocation. Readers building vision-language navigation systems or simulation-based aerial datasets will also find the DroneNav-6D construction pipeline relevant. A working understanding of latent predictive models and transformer-based planning is needed to follow the method section comfortably.
Authors’ abstract
World-action models give visual navigation agents a way to anticipate how candidate actions will change future observations and to act from the predicted consequences. For drones, this capability must operate under tight accuracy and efficiency constraints. We present DroneWAM, an efficient world-action model for drone visual navigation. DroneWAM adopts a JEPA-based architecture to model future states directly in representation space, avoiding the cost of explicit future image generation. A pretrained Resampler further compresses dense encoder features into fewer latent tokens, reducing the computation repeated at each imagined step. We also introduce adaptive rollout, where a preference-trained Gate adaptively allocates prediction depth according to the current scene. To support learning under richer aerial motion, we construct DroneNav-6D, a simulated visual navigation dataset with synchronized RGB observations, 6-DoF flight trajectories, control commands, and randomized wind disturbances. On DroneNav-6D, DroneWAM achieves the best trajectory accuracy among the compared methods. Adaptive rollout further reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy, demonstrating that predictive computation can be allocated more effectively across scenes. \href{https://github.com/1e12Leon/DroneWAM}{Codes and data} will be released.