Research
A Multi-Drone Multi-View Dataset and Deep Learning Framework for Pedestrian Detection and Tracking
Overview Research area: Computer vision for multi-drone (aerial) pedestrian detection and tracking; multi-view perception, bird's-eye-view (BEV) feature fusion, and surveillance dataset construction.

- arXiv
- 2511.08615
- Published
- 2025-11-06
- Authors
- Kosta Dakic, Kanchana Thilakarathna, Rodrigo N. Calheiros, Teng Joon Lim
AI summary
Overview
Research area: Computer vision for multi-drone (aerial) pedestrian detection and tracking; multi-view perception, bird's-eye-view (BEV) feature fusion, and surveillance dataset construction.
Technical level: Advanced. The paper assumes familiarity with camera projection geometry, homography estimation, RANSAC, probabilistic occupancy maps, convolutional encoders, and multi-object tracking metrics (MOTA, MOTP, IDF1).
Scope in one sentence: The paper introduces MATRIX, a synthetic eight-drone multi-view pedestrian dataset with continuously moving cameras, and a deep learning pipeline that handles dynamic camera calibration, view registration, and BEV feature fusion for detection and tracking in occluded urban scenes.
What This Paper Is About
Existing multi-view pedestrian detection and tracking systems mostly assume cameras that stay in a fixed place, which makes it easy to line up views and predict where people will go. Drones break that assumption because the cameras are constantly moving, views shift, and people get hidden behind buildings. The paper builds a dataset and a processing pipeline designed specifically for that moving-camera setting, then tests how well both static-camera methods and the new method hold up when the scene gets crowded and obstructed.
Key Contributions
- The MATRIX dataset (Multi-Aerial TRacking In compleX environments): synchronized footage from eight drones with continuously changing positions in an urban environment, with comprehensive detection and tracking annotations, released at https://github.com/KostaDakic/MATRIX/tree/main.
- A dynamic camera calibration system that continuously updates extrinsic parameters as the drones move through the surveillance area.
- An efficient multi-view feature fusion pipeline that keeps detection and tracking performance high while adapting to rapid scene changes and viewpoint variation.
- A comprehensive evaluation covering both simple and complex environments with major occlusions, transfer-learning generalization to modified scenarios, and systematic camera-dropout tests.
Main Findings
- Static-camera methods fail in the complex scene: Methods designed for fixed camera setups maintain over 90% detection and tracking precision and accuracy metrics in a simplified MATRIX environment (no obstruction, 10 pedestrians, much smaller observational area), but their performance degrades significantly in the complex environment.
- The proposed pipeline stays robust: It maintains approximately 90% detection and tracking accuracy and successfully tracks approximately 80% of trajectories under the challenging conditions.
- Transfer learning helps substantially: In the transfer-learning experiments, the pretrained model achieved much higher detection and tracking accuracy than training the model from scratch.
- Camera failures degrade performance gracefully: Systematic camera-dropout experiments showed a graceful degradation pattern rather than a collapse, which the authors present as evidence of practical robustness for real deployments.
- Dataset gap identified: Existing drone datasets are mostly single-drone or limited to 2–4 cameras (MDOT: 2–3 drones; MDMT: 2 drones; MUMO: 4 cameras), while established multi-view datasets such as WildTrack (7 cameras) and MultiViewX (6 cameras) use static cameras. MATRIX combines eight-camera coverage with drone mobility.
Methodology in Plain English
The data is generated in simulation rather than filmed in the real world. The authors build two urban scenes in Unreal Engine 5 and use AirSim to fly eight drones carrying 1920×1080 cameras with a 70-degree field of view. Drone flight is confined to ±3 m in x and y and 7–8 m in z. A PID controller keeps the cameras aimed at the scene center (gains Kp = 0.0002, Ki = 0.00002, Kd = 0.0001), and velocity-based position control limits movement to v_max = 1.5 m/s to keep flight smooth. Data is captured every 0.5 seconds.
Because the drones move, the camera pose must be re-estimated constantly. Checkerboard patterns placed around the scene provide known 3D points that are matched to their 2D image locations, so intrinsic and extrinsic parameters can be recovered on the fly. Ground-truth annotations are generated automatically: probabilistic occupancy maps estimate where pedestrians are on the ground plane, and ray-casting in Unreal Engine checks line of sight between each drone and each pedestrian — if geometry blocks the ray, the pedestrian is flagged as not visible. Annotations include 2D boxes, 3D world coordinates, visibility flags, consistent person IDs, and per-frame calibration parameters.
The tracking framework takes synchronized frames from all drones and runs them through two parallel streams — one holding a stable reference view, one processing the incoming frame. Each view is projected onto the ground plane as a BEV representation using a time-varying homography (valid because pedestrians are tracked at foot level, z = 0). The current view is then aligned to the reference using GFTT-AffNet-HardNet keypoints, symmetric nearest-neighbor matching with Lowe's ratio test (threshold typically 0.8), and RANSAC homography estimation; a confidence score is computed as the ratio of inliers to total matches. The aligned BEV features from all drones are stacked and compressed, then fused with historical frames for temporal consistency, before specialized heads decode detections and tracking identities. Because RANSAC is stochastic, the authors report their results with standard deviations over multiple runs using different random seeds.
Why This Matters
Impact on research: The paper targets a specific blind spot — nearly all multi-view benchmarks use fixed cameras, so algorithms that look strong on those benchmarks have never been stress-tested against moving viewpoints, dense crowds, and architectural occlusion. MATRIX gives the field a benchmark where static-camera assumptions visibly break down, and the reported graceful degradation under camera dropout gives a baseline for fault-tolerant multi-robot perception.
Real-world applications:
- Urban surveillance and public-space security, where drones must monitor crowds around buildings and other obstructions.
- Event crowd management and safety monitoring, where pedestrian density is high and occlusion is constant.
- Search and rescue or disaster response, where camera positions shift and individual cameras may fail.
- Traffic and pedestrian-flow analysis in cities, where fixed infrastructure may be sparse or unavailable.
Industry relevance: Drone-based monitoring, smart-city sensing, and edge-computing platforms all need detection and tracking that survives moving sensors and partial hardware failure. The paper explicitly frames computational and real-time processing requirements of aerial platforms as a constraint, and notes the importance of efficiency and parallelization for mobile edge computing.
Future Directions
- Closing the simulation-to-reality gap: MATRIX is generated in Unreal Engine 5 with AirSim under ideal weather and lighting. Whether these results transfer to real drone footage with sensor noise, wind, and variable illumination is an open question.
- Handling more severe camera loss: Camera dropout testing showed graceful degradation; how the pipeline behaves when a large fraction of the eight drones fails simultaneously, or when failures are correlated, is not established in the available content.
- Scaling to higher density and larger areas: The complex variant covers 40 pedestrians in a 30×30 m area. Behaviour under denser crowds or wider operational areas is untested in the content provided.
- Improving the stochastic parts of the pipeline: RANSAC introduces run-to-run variability in detection and tracking, which the authors quantify with standard deviations across seeds. Reducing or removing that variance is a natural engineering target.
Target Audience
Researchers and graduate students in computer vision working on multi-view detection, multi-object tracking, or multi-camera surveillance; robotics and aerial-systems engineers building drone-based monitoring; and practitioners in smart-city, security, or edge-computing roles who need to reason about how tracking pipelines behave when cameras move or fail. Readers without a background in camera geometry, homography estimation, and tracking metrics will need to consult the cited background work (MVDet, MVDeTr, EarlyBird, POM-based methods) first.
Authors’ abstract
Multi-drone surveillance systems offer enhanced coverage and robustness for pedestrian tracking, yet existing approaches struggle with dynamic camera positions and complex occlusions. This paper introduces MATRIX (Multi-Aerial TRacking In compleX environments), a comprehensive dataset featuring synchronized footage from eight drones with continuously changing positions, and a novel deep learning framework for multi-view detection and tracking. Unlike existing datasets that rely on static cameras or limited drone coverage, MATRIX provides a challenging scenario with 40 pedestrians and a significant architectural obstruction in an urban environment. Our framework addresses the unique challenges of dynamic drone-based surveillance through real-time camera calibration, feature-based image registration, and multi-view feature fusion in bird's-eye-view (BEV) representation. Experimental results demonstrate that while static camera methods maintain over 90\% detection and tracking precision and accuracy metrics in a simplified MATRIX environment without an obstruction, 10 pedestrians and a much smaller observational area, their performance significantly degrades in the complex environment. Our proposed approach maintains robust performance with $\sim$90\% detection and tracking accuracy, as well as successfully tracks $\sim$80\% of trajectories under challenging conditions. Transfer learning experiments reveal strong generalization capabilities, with the pretrained model achieving much higher detection and tracking accuracy performance compared to training the model from scratch. Additionally, systematic camera dropout experiments reveal graceful performance degradation, demonstrating practical robustness for real-world deployments where camera failures may occur. The MATRIX dataset and framework provide essential benchmarks for advancing dynamic multi-view surveillance systems.