Research
AirNav: A Large-Scale UAV Vision-and-Language Navigation Dataset with Natural and Diverse Instructions
Overview Research area: UAV (drone) vision-and-language navigation (VLN), benchmark dataset construction, and multimodal large language model (MLLM) training for embodied navigation. Technical level:

- arXiv
- 2601.03707
- Published
- 2026-01-07
- Authors
- Hengxing Cai, Yijie Rao, Ligang Huang, Zanyang Zhong, Jinhan Dong, Jingjun Tan, Changhao Nai, Jue Hou, Wenhao Lu, Renxin Zhong
AI summary
Overview
Research area: UAV (drone) vision-and-language navigation (VLN), benchmark dataset construction, and multimodal large language model (MLLM) training for embodied navigation.
Technical level: Intermediate. The paper is readable without deep reinforcement-learning background, but the training and reward-design sections assume familiarity with supervised fine-tuning, GRPO-style policy optimization, and standard VLN evaluation metrics.
Scope (1 sentence): The paper introduces AirNav, a 137K-sample UAV VLN benchmark built from real urban aerial data with human–LLM-generated instructions across 10 user personas, plus AirVLN-R1, a two-stage fine-tuned navigation model that reaches 51.82% success rate on the test-unseen split.
What This Paper Is About
UAV navigation guided by natural language needs data that is simultaneously realistic, naturally worded, and large enough to train and test models on. Existing UAV VLN benchmarks typically deliver only two of those three: some use real aerial imagery but give only target descriptions with no intermediate steps, others include step-by-step instructions but are template-generated or simulation-only, and most are small or linguistically narrow.
AirNav aims to fix that by building a benchmark from real urban aerial data with 137K navigation samples, instructions written by a human–LLM collaborative pipeline with 10 distinct user personas, and a unified evaluation of baselines running from traditional models to MLLMs. The authors also train AirVLN-R1 on this data and test it on a physical UAV.
Key Contributions
- AirNav benchmark: A large-scale UAV VLN benchmark built on real urban aerial data with 137K navigation samples and natural, diverse instructions produced by a human–LLM collaborative pipeline using 10 user personas.
- Unified evaluation: A systematic comparison of representative approaches, from traditional models to multimodal large language models, under unified metrics with open-source implementations released.
- AirVLN-R1 model: A navigation model combining supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT) that achieves state-of-the-art results on AirNav, including a 51.82% success rate on the test-unseen split.
- Real-world deployment: Deployment of AirVLN-R1 on a physical UAV platform with indoor and outdoor navigation experiments, offering initial evidence for practical deployment and sim-to-real transfer.
Main Findings
-
AirNav is larger and more linguistically diverse than prior UAV VLN benchmarks. AirNav has 137K samples, a 4 DoF action space, sub-goals, and a vocabulary size of 20.3K. In comparison, OpenFly has 100k samples and a 15.6K vocabulary; CityNav has 32,637 samples; OpenUAV has 12,149; AerialVLN has 8,446; AVDN has 3,064; and LANI has 6,000 with a 2.3K vocabulary. Among these, only AVDN, CityNav, and AirNav are built from real data (OpenFly uses virtual plus real data).
-
Instructions are rated more natural than prior datasets. An LLM-based automated assessment scored 2,000 instructions per dataset on a 5-point scale using GPT-4o. AirNav achieves the highest naturalness score (3.75). Human validation of this evaluation produced a Krippendorff's alpha of 0.70 (inter-annotator agreement) and a Spearman's rho of 0.74 (human–LLM correlation).
-
Personas produce measurable stylistic differences. Different personas show clearly separated medians and distribution ranges in instruction length. Retired Adults tend to produce longer, more explanatory instructions, while University Students and Advanced Navigation Users generally prefer concise expressions.
-
Traditional methods fail on AirNav. Seq2Seq and CMA achieve success rates below 2% and 6% respectively. Among zero-shot MLLMs, LLaMA-3.2-11B-Vision also yields low SR, GPT-4o performs better but remains limited on unseen scenarios, and Qwen3-VL-235B-A22B reaches the highest SR among all zero-shot baselines.
-
Fine-tuned navigation models beat general-purpose ones. π0, fine-tuned on AirNav, achieves only 6.29% SR on val-seen, suggesting VLA models built for robot manipulation do not transfer well to UAV VLN. Uni-NaVid, fine-tuned on AirNav, reaches 20.56% SR on val-seen and 15.89% on test-unseen.
-
AirVLN-R1 leads across all metrics. AirVLN-R1 achieves 51.76% SR on val-seen, 51.32% on val-unseen, and 51.82% on test-unseen, with NE of 39.7, 40.8, and 39.8 respectively, and SPL of 50.61, 50.11, and 50.66. Built on Qwen2.5-VL-7B, it outperforms larger models such as Qwen2.5-VL-32B and Qwen3-VL-235B-A22B.
-
Both training stages are necessary. SFT-only reaches 44.27% SR on val-seen but leaves room for improvement on val-unseen and test-unseen. RFT-only struggles due to sparse reward signals without SFT initialization. SFT+RFT gives the best and most stable results across all splits.
-
History sampling strategy matters. Progressive Interval Sampling achieves the best performance (51.82% SR on test-unseen), outperforming Uniform-K (49.83%) by balancing short-term perception with long-range context.
-
Reward components contribute differently. Removing the Subgoal State Alignment reward causes the largest drop (51.82% to 46.00%). Removing the Stop Consistency reward drops performance to 47.65%. The Format reward contributes to training stability with a smaller but consistent effect.
-
Reward weighting is robust. A 3×3 grid search over reward weights λ1 and λ2 produced a maximum SR fluctuation within 1.5 percentage points.
-
Persona difficulty varies. Teacher/Educator and Advanced Navigation User instructions yield the top two SR values, while Retired Adult instructions (verbose, repeated context) and Primary School Student instructions (concise but lacking spatial detail) achieve comparatively lower SR.
-
Real-world results track simulation. Across 200 real navigation tasks (100 indoor, 100 outdoor), Seq2Seq and CMA succeeded on only 1 and 2 tasks respectively. AirVLN-R1 reached SR = 53/200 with the lowest NE (70.6).
Methodology in Plain English
The benchmark is built in four steps on top of existing data. The authors start from the SensatUrban and CityRefer datasets and use the CityFlight environment, which aligns SensatUrban data with OpenStreetMap. First, a start point is randomly sampled and a navigation target is chosen; an MLLM writes a target description that is filtered for ambiguity. Second, the MLLM identifies landmarks between start and target, with a maximum distance constraint between consecutive landmarks and a semantic refinement pass that verifies and rewrites landmark descriptions. Third, a look-ahead strategy generates executable action sequences for each path segment, which are concatenated into a full trajectory. Fourth, GPT-4o receives the trajectory, map, and node positions/descriptions and writes the navigation instruction, guided by persona settings and few-shot examples of human-written navigation instructions.
To keep quality up at scale, the authors run an iterative inspect-revise loop: sample outputs, categorize errors (ambiguous landmarks, visually indistinguishable targets, excessive gaps between landmarks), update prompts and filters, regenerate, and repeat until the sampled error rate stabilizes.
For the model, each step feeds the model the instruction, UAV coordinates and heading, prior action history, the current first-person image, and a set of historical images chosen by Progressive Interval Sampling so recent steps are densely sampled and distant ones sparsely. The model outputs a variable-length sequence of up to 8 discrete actions from {move forward, turn left, turn right, stop}. Training runs in two stages: supervised fine-tuning on the training set, then reinforcement fine-tuning with GRPO against a combined reward of distance-to-subgoal progress, heading-angle alignment, stop-decision consistency, and output formatting. AirVLN-R1 is built on Qwen2.5-VL-7B and trained on an 8×A100 GPU server.
Why This Matters
Impact on research. The paper addresses a specific bottleneck: the absence of a benchmark that simultaneously offers real aerial scenes, process-level natural instructions, and scale. It also releases open-source implementations and a unified evaluation protocol, which makes head-to-head comparison across traditional, zero-shot MLLM, and fine-tuned navigation models easier. The finding that a 7B model with task-specific SFT+RFT beats much larger zero-shot MLLMs is a meaningful data point for the community.
Real-world applications (from the paper):
- Emergency rescue
- Urban patrol
- Infrastructure inspection
- Search-and-rescue (listed in the ethics section as a beneficial use)
Industry relevance. The real-world experiments on a physical UAV platform, including a comparison of inference latency and GPU memory against baselines, speak directly to deployability rather than benchmark scores alone. The paper's own framing of the discrete-action gap and the sim-to-real gap, plus the resource-cost discussion, gives practitioners a clearer picture of where such models can and cannot be plugged into real flight stacks.
Future Directions
-
Broader geographic and environmental coverage. AirNav is built on SensatUrban and CityRefer, so city styles, infrastructure layouts, and annotation schemes are constrained by those sources. Generalization across different cities, countries, and seasonal conditions remains unvalidated, and aerial imagery may not capture complex altitude variations, occlusion patterns, and extreme lighting.
-
Closing the discrete-action gap. The discrete action set with sequences of up to eight steps supports stable training and fair comparison but cannot represent fine-grained motions or dynamic constraints of continuous UAV control. Long-horizon navigation and precise target approach may suffer from trajectory approximation errors.
-
Reducing the sim-to-real gap. Real-world evaluation covers 200 navigation tasks with limited environmental diversity, and perception noise, viewpoint variation, and control uncertainty persist. Benchmark gains do not automatically translate into stable real-world gains.
-
Reproducibility of closed-source baselines. The evaluation includes GPT-4o, which is subject to API access restrictions and version changes, so exact reproduction of reported results may be limited.
Target Audience
Researchers and engineers working on UAV autonomy, embodied navigation, vision-and-language navigation, and multimodal foundation models. It is most useful to those who need a benchmark for training or comparing navigation agents, to teams evaluating MLLMs on embodied spatial reasoning, and to practitioners assessing whether VLN models transfer to physical drone platforms. Readers focused on dataset construction methodology, instruction generation with LLMs, or reinforcement fine-tuning for sequential decision-making will also find the pipeline and reward design sections relevant.
Authors’ abstract
Existing UAV vision-and-language navigation (VLN) benchmarks rarely provide realistic aerial scenes, natural process-level instructions, and sufficient scale simultaneously, making it difficult to systematically train and evaluate UAV VLN agents under realistic settings. To address this, we propose \textbf{AirNav}, a large-scale benchmark built on real urban aerial data, comprising 137K navigation samples with natural and diverse instructions generated via a human--LLM collaborative pipeline with 10 user personas. We conduct a systematic evaluation of representative approaches on AirNav, ranging from traditional models to multimodal large language models (MLLMs), under unified metrics with open-source implementations. We further propose \textbf{AirVLN-R1}, trained via supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT), achieving state-of-the-art performance with a 51.82\% success rate on the test-unseen split. Real-world experiments on a physical UAV platform provide preliminary evidence of sim-to-real transferability, and our dataset and code are publicly available.