Skip to content
AI.info

Research

FlyCo: Foundation Model-Empowered Drones for Autonomous 3D Structure Scanning in Open-World Environments

Overview Research area: Aerial robotics — autonomous drone-based 3D scanning, and the integration of foundation models (FMs) into robot perception and planning. Technical level: Advanced. The paper as

arXiv
2601.07558
Published
2026-01-12
Authors
Chen Feng, Guiyong Zheng, Tengkai Zhuang, Yongqian Wu, Fangzhan He, Haojia Li, Juepeng Zheng, Shaojie Shen, Boyu Zhou

AI summary

Overview

Research area: Aerial robotics — autonomous drone-based 3D scanning, and the integration of foundation models (FMs) into robot perception and planning.

Technical level: Advanced. The paper assumes familiarity with SLAM, coverage planning, point-cloud completion, and vision-language models, although the core architecture can be understood at a conceptual level.

One-sentence scope: FlyCo is a foundation-model-empowered drone system that turns low-effort human prompts (text descriptions and sparse visual annotations) into autonomous, efficient, and collision-free 3D scans of user-specified structures in unknown open-world environments.

What This Paper Is About

Existing drone scanning systems either rely on pre-defined flight patterns, two-stage pipelines with offline human processing, external priors such as satellite imagery or CAD models, or exploration algorithms that require user-specified 3D bounding boxes. All of these depend on high-effort human setup or idealized assumptions. FlyCo's goal is to replace those priors with foundation-model world knowledge, so a drone can identify an intended structure amid clutter, predict its unseen geometry from partial views, and plan an efficient, safe coverage path — all from an abstract prompt.

Key Contributions

  1. A principled FM-empowered system architecture. FlyCo integrates generalizable foundation models with flight skills through a closed-loop perception–prediction–planning framework, enabling prompt-driven scanning of complex structures in diverse, unknown environments without human intervention or prior environmental knowledge.

  2. A prompt-grounded perception module. It combines visual segmentation foundation models (SAM2) with vision-language knowledge (BEiT3) via an early-fusion scheme, and adds FM-tailored streaming-efficient adaptation plus cross-modal refinement to boost efficiency and robustness across viewpoints and time. The streaming adaptation uses a sliding-window memory bank retaining only the most recent W_m segmentation results, together with lightweight quantization and graph-level compilation for onboard inference. Cross-modal refinement uses a linear Kalman filter over 2D mask motion, a combined keyframe score (box IoU temporal consistency plus pixel-level IoU geometric consistency), and a cross-modal re-initialization triggered when box IoU falls below a threshold κ_reinit.

  3. A multi-modal surface prediction pipeline. A minimal-bias, post-fusion network fuses partial geometric cues (which anchor outputs to physical measurements) with visual and textual cues distilled from FMs (which drive generalization), producing metrically grounded shape completion. Four supporting designs are highlighted: FM-based automatic data generation for annotation-free, large-scale multi-modal learning; adaptive densification to handle structures at different scales; and lightweight partial-surface regularization during training to keep predictions faithful to observed geometry and improve temporal stability.

  4. A prediction-aware hierarchical planner. Global planning injects consistency awareness to stay coherent with historical paths, and accelerates NP-hard coverage planning using inlier-driven intersection parity checking, gravitational-like viewpoint pruning, and a two-level decomposition strategy. A local layer applies viewpoint-constrained trajectory optimization for high-frequency replanning, preserving information completeness while responding to emerging obstacles.

A fifth stated contribution is extensive validation via real-world in-the-wild flights, large-scale simulation benchmarks against state-of-the-art (SOTA) methods, and ablation studies.

Main Findings

  • Benchmarked advantage over SOTA in complex 3D simulated scenarios: FlyCo consistently outperformed competitors with lighter human input, 1.25–3× reduced flight time, 4.3–56.2% higher target coverage, and 11.6–50.1% higher success rate. (The paper does not name the specific simulation benchmarks or report dataset sizes in the provided content.)

  • Real-world in-the-wild validation: Field trials at unmapped wild sites covered four demonstrated scenarios — (A) an arch bridge linking buildings, (B) a large concert hall, (C) a castle gate, and (D) a red-brick building. FlyCo reliably scanned user-specified structures from simple prompts.

  • Cross-modal refinement is necessary: The paper reports that naively projecting 2D segmentation masks onto 3D LiDAR points fails under strong viewpoint changes and occlusions during long-horizon flights, motivating the geometric clustering and bidirectional refinement mechanism.

  • Ablations confirm each component's contribution: Comprehensive ablations are reported to validate the design choices of the individual modules.

  • Latency mismatch is resolved architecturally: Asynchronous operation of the three modules removes mutual blocking, and the planner continually harvests low-frequency FM inference outputs, reconciling slow FM reasoning with real-time flight needs.

  • Termination is autonomous: The system detects when the target surface is sufficiently covered and ends the mission, delivering scanned data to the user.

Methodology in Plain English

The design is modeled on how an expert human pilot handles a scanning request such as "scan the castle in the valley": identify the intended structure, imagine its full layout from partial views, and coordinate sensing with motion. FlyCo splits that behavior into three cooperative modules rather than one end-to-end learned policy — a deliberate choice to avoid the massive expert-flight datasets that end-to-end FM training would require.

  1. Perception — "what to scan." Text prompts and visual annotations from the user are fused with streaming RGB and LiDAR data. A density-based clustering algorithm (Ester et al. 1996) partitions the initial LiDAR cloud into groups; clusters that project inside the initial 2D mask are labeled as target. Clusters are updated incrementally, and the 3D target cluster both selects the most temporally stable and geometrically representative keyframes for SAM2's memory attention and detects segmentation drift, repairing it by projecting farthest-point-sampled target points back into the image as virtual prompts.

  2. Prediction — "what is not yet seen." The segmented visual and geometric observations plus the original text prompt feed a post-fusion network that completes the target's full 3D surface, remaining faithful to the parts already measured. Training data is synthesized automatically by an FM-based generation tool, so no manual annotation is needed.

  3. Planning — "how to cover it efficiently and safely." The predicted geometry converts the open-ended scanning task into a structured coverage problem with global context. The planner generates coverage viewpoints on the prediction, prunes and decomposes the problem for real-time solvability, and hands global guidance to a local optimizer that produces a five-degree-of-freedom trajectory (drone position plus camera gimbal pitch and yaw) while avoiding obstacles.

The loop repeats as new sensor data arrives, and the paper formalizes the underlying tension as a multi-objective problem over flight time T(x), incompleteness 1 − C(x) with C(x) ∈ [0,1], and a qualitative human-effort measure E(h). The authors state this formulation is illustrative and not explicitly optimized — it exists to state the design principles.

Why This Matters

Impact on research: The paper argues that naively mirroring FM-based ground-robot systems fails for aerial scanning, because end-to-end learning is data-hungry and FM inference latency conflicts with real-time flight. FlyCo offers a modular alternative: FMs handle perception and prediction, while optimization handles long-horizon planning and real-time response. The modular interfaces are explicitly presented as a flexible, extensible blueprint that can absorb future foundation models and robotics advances.

Real-world applications (the three use cases cited in the paper plus the demonstrated scenarios):

  • Infrastructure inspection (the paper cites BeyondVision 2023)
  • Disaster response (Esri 2023)
  • Heritage documentation (Mapalytics 2024)
  • Inspection of specific structures such as arch bridges, concert halls, castle gates, and buildings, as demonstrated in the field trials

Industry relevance: The paper positions FlyCo against commercial mapping and inspection tools from DJI, Skydio, Litchi, and PIX4D, which generate parameterized lawnmower, orbit, or grid patterns after a user draws a polygon or box. It also contrasts with research lines such as two-stage coarse-scan-then-offline-planning pipelines and exploration methods like Star-Searcher (Luo et al. 2024) that require 3D bounding boxes. Lower human effort, fewer re-flights, and fewer post-flight human evaluations are the practical selling points.

Future Directions

  • Absorbing future foundation models and robotics advances. The authors explicitly frame FlyCo's modular perception–prediction–planning loop and interfaces as a blueprint that can be upgraded by direct module replacement.
  • Releasing code and development environments. The paper states that code and development environments will be released, which would enable reproduction and extension.
  • Closing the generalization gap noted for prior prediction work. The paper observes that earlier prediction-based scanning methods were evaluated mainly in small-scale simulations or simplified environments with no target-irrelevant obstacles and restricted structure diversity, leaving zero-shot generalization to real-world scenes unclear — FlyCo tackles this, but the question of sustained robustness in the wild remains open.
  • Extending beyond camera-limited coverage. The formulation notes that because the LiDAR's effective range typically exceeds the camera's, scanning completeness is effectively governed by camera coverage, which frames sensing coverage as the practical bottleneck to push further.

Target Audience

Robotics researchers working on aerial autonomy, integrated planning and learning, and autonomous agents; practitioners building inspection, mapping, or surveying drones; and researchers studying how to deploy foundation models on physical robots without end-to-end expert demonstration data. The paper's keywords are Aerial Systems: Perception and Autonomy, Aerial Systems: Applications, Integrated Planning and Learning, and Autonomous Agents.

Authors’ abstract

Autonomous 3D scanning of open-world target structures via drones remains challenging despite broad applications. Existing paradigms rely on restrictive assumptions or effortful human priors, limiting practicality, efficiency, and adaptability. Recent foundation models (FMs) offer great potential to bridge this gap. This paper investigates a critical research problem: What system architecture can effectively integrate FM knowledge for this task? We answer it with FlyCo, a principled FM-empowered perception-prediction-planning loop enabling fully autonomous, prompt-driven 3D target scanning in diverse unknown open-world environments. FlyCo directly translates low-effort human prompts (text, visual annotations) into precise adaptive scanning flights via three coordinated stages: (1) perception fuses streaming sensor data with vision-language FMs for robust target grounding and tracking; (2) prediction distills FM knowledge and combines multi-modal cues to infer the partially observed target's complete geometry; (3) planning leverages predictive foresight to generate efficient and safe paths with comprehensive target coverage. Building on this, we further design key components to boost open-world target grounding efficiency and robustness, enhance prediction quality in terms of shape accuracy, zero-shot generalization, and temporal stability, and balance long-horizon flight efficiency with real-time computability and online collision avoidance. Extensive challenging real-world and simulation experiments show FlyCo delivers precise scene understanding, high efficiency, and real-time safety, outperforming existing paradigms with lower human effort and verifying the proposed architecture's practicality. Comprehensive ablations validate each component's contribution. FlyCo also serves as a flexible, extensible blueprint, readily leveraging future FM and robotics advances. Code will be released.

Read the original paper