Research
CageDroneRF: A Large-Scale RF Benchmark and Toolkit for Drone Perception
Overview Research area: Radio-Frequency (RF) sensing for Unmanned Aerial Vehicle (UAV, or "drone") detection and identification, positioned within computer vision through the use of spectrogram images

- arXiv
- 2601.03302
- Published
- 2026-01-06
- Authors
- Mohammad Rostami, Atik Faysal, Hongtao Xia, Hadi Kasasbeh, Ziang Gao, Huaxia Wang
AI summary
Overview
Research area: Radio-Frequency (RF) sensing for Unmanned Aerial Vehicle (UAV, or "drone") detection and identification, positioned within computer vision through the use of spectrogram images for classification and object detection.
Technical level: Intermediate. The paper is readable without deep RF expertise, but it assumes familiarity with neural network training, object detection annotation formats, and basic signal-processing concepts such as I/Q sampling and the Short-Time Fourier Transform.
Scope: The paper introduces CageDroneRF (CDRF), a large-scale benchmark dataset and open-source toolkit for RF-based drone detection, classification, open-set recognition, and object detection, built from real-world Faraday-cage and outdoor captures with a raw-signal augmentation pipeline that keeps detection labels consistent.
What This Paper Is About
Machine learning for RF-based drone detection has lagged behind other sensing domains because public datasets are narrow in class diversity, often distributed only as preprocessed or magnitude-only data, and rarely include standardized tooling. As a result, models trained on them can score near-perfectly on easy, clean benchmarks yet degrade badly in realistic conditions with Wi-Fi/Bluetooth interference, spectrum crowding, and low Signal-to-Noise Ratio (SNR). The paper's goal is to close this gap by releasing a combined dataset-and-toolkit resource: approximately 500+ GB of raw complex I/Q recordings spanning 39 classes and 23 drone models across two collection environments, paired with an augmentation and evaluation stack that automatically regenerates spectrograms and recomputes bounding-box annotations whenever a signal-level transform or spectrogram parameter is changed.
Key Contributions
-
The CDRF benchmark dataset. Real-world RF captures from a customized Faraday-cage facility and from outdoor recordings at the Rowan University campus, with standardized spectrogram generation and rich per-sample metadata. The dataset totals approximately 500+ GB of raw data covering 39 unique classes and 23 distinct commercial and hobbyist UAV models, their remote controllers, and operational variants (e.g., armed vs. unarmed states), plus a substantial number of no-drone recordings and captures of multiple drones operating simultaneously.
-
A raw-signal (I/Q-level) augmentation pipeline. Controlled SNR injection via complex additive white Gaussian noise (AWGN), interferer mixing at a controllable ratio, and frequency shifting via multiplication with a complex exponential, all applied to complex baseband I/Q before time-frequency conversion. For detection, frequency shifts trigger exact recomputation of YOLO-format bounding boxes with correct wrap-around behavior on the spectrogram's frequency axis.
-
SNR-structured dataset variants and negative samples. Programmatic creation of SNR-stratified splits (including noise-only spectrograms and per-SNR subdirectories) for robustness sweeps, plus a curated corpus of real-world "no-drone" segments designed to support well-calibrated binary detectors and hard-negative mining.
-
Interoperable, dataset-agnostic tooling and baselines. Open-source utilities for dataset creation, metadata generation, cleaning of third-party releases (e.g., Roboflow-style layouts), and spectrogram rendering; PyTorch baselines (binary and multi-class) built on ResNet-18 with dataloaders for both spectrogram images and array/pickle formats; an implementation of open-set recognition for previously unseen classes; and a lightweight patch exposing per-detection class probabilities from YOLO for calibration and open-set analysis. The code, dataset, and trained models are available at https://github.com/DroneGoHome/U-RAPTOR-PUB.
Main Findings
-
Existing datasets share a "realism gap." The paper surveys DroneRF (three drone models in a controlled environment), DroneDetect/DroneDetect V2 (seven drones with explicit Wi-Fi/Bluetooth co-channel interference, raw complex I/Q via a Nuand BladeRF SDR with GNURadio), Cardinal RF (restricted public availability, limiting reproducibility), VTI_DroneSET_FFT (three DJI models, distributed as preprocessed .mat files), the Noisy Drone RF Signal Dataset (synthetically controlled SNR ranges such as −20 to 30 dB, distributed as preprocessed tensors), and RFUAV (the largest existing benchmark at ∼1.3 TB of raw complex I/Q from 37 UAV types, but predominantly high-SNR, with low-SNR and interference conditions synthesized post-hoc rather than captured in situ).
-
Raw-to-annotation traceability is the most consequential differentiator. In prior benchmarks that ship raw I/Q alongside pre-generated spectrograms and labels, there is no automated, reproducible link between signal and annotation. Changing noise, frequency shift, or interferer mixing invalidates pre-existing boxes, and adjusting sampling rate, FFT size, window length, or hop size changes time and frequency resolution so that every pixel coordinate and therefore every bounding box changes. CDRF instead maintains a complete parameterized chain from raw I/Q through augmentation to detection annotation, automatically producing new spectrogram images and recomputed bounding boxes — including correct wrap-around handling when a shifted box straddles the top/bottom edge, which is split into two valid boxes whose heights sum to the original.
-
CDRF deliberately uses a lower sampling rate than RFUAV. CDRF adopts a 20 MHz sampling rate chosen for edge-device compatibility, whereas RFUAV's 100 MHz (100 MS/s) complex sampling rate imposes computational and memory demands the paper describes as impractical for real-time, resource-constrained deployments where capture, spectrogram generation, and inference must all fit within tight budgets.
-
Dual-environment collection is novel among RF drone datasets. CDRF combines Faraday-cage isolation for clean reference captures with open-campus outdoor recordings under natural interference. Representative outdoor spectrograms from the Autel EXOII, DJI Inspire1, DJI Mavic3, and DJI Phantom3 Advanced show visibly brighter and more cluttered backgrounds than indoor captures.
-
Signals are augmented at baseband to preserve RF physics. SNR conditioning adds complex AWGN with noise power set to σ² = P_s / SNR_lin for a target SNR_dB; frequency shifting multiplies by exp(j2πΔf n / F_s), inducing a vertical translation of Δy = Δf / F_s in normalized image coordinates; interferer mixing combines two normalized signals as norm(x₁ + α x₂) with α ∈ [0, 1] and optional independent frequency shifts on each source.
-
A specific annotation policy governs detection labels. CDRF uses YOLO-format annotations on spectrograms with a whole-signal policy: contiguous RF emission bursts from a device are annotated as a single object. Boxes begin and end on an ON state, and OFF states at the beginning and end of a spectrogram are excluded. If bursts occupy less than 10% of the spectrogram's time extent, or if only a single ON state exists in the entire spectrogram, no bounding box is assigned and the spectrogram is labeled as background. Remote controller signals receive two annotation types: a single box covering all bursts, and per-channel boxes for bursts transmitted on the same frequency channel.
-
Spectrogram generation parameters are fully exposed. The default pipeline uses scipy.signal.stft with a Hann window, FFT_SIZE = 1024, overlap = 128 samples, two-sided spectra with DC centered via fftshift, SAMPLING_RATE = 20 MHz, and NUM_FFT_SPEC = 1500 time bins per sample. Multiple colormaps (viridis, plasma, inferno, magma, cividis, gray, hot) are offered to test color sensitivity, with normalization to [0, 1] before colorization. The paper states that spectrogram-based models have been found more robust than raw 1D I/Q pipelines under low SNR and co-channel interference, particularly in ISM bands congested by Wi-Fi and Bluetooth.
-
Benchmark results are not reported in the provided content. The paper states that Section V presents the experimental setup and benchmark results for drone detection, single-label, open-set, and hierarchical classification, and that Section VII concludes with future research directions. The provided text truncates at Section IV-F ("Reproducibility and Metadata"), so no accuracy, precision, recall, or other quantitative model results are available here.
Methodology in Plain English
The researchers first built a hardware capture platform. An SDR-based receive chain uses a Universal Software Radio Peripheral (USRP) B200-mini driven by GNU Radio v3.10, connected to a dual-band omnidirectional antenna on a static mount, with a Lenovo IdeaPad Flex 5 (8-core Intel Core i7-1165G7 CPU, 16 GB memory, Ubuntu 24.04.2 LTS) as the data sink. To isolate drone signals, they constructed a customized Faraday cage and placed both the SDR card and the receiving antenna inside it. During indoor capture, the drone is powered on outside the cage at a very close distance to the antenna while its remote controller is kept in an adjacent room — so drone signals arrive strong while interfering sources are attenuated and remote-controller signals are largely excluded. The SDR gain is set to 50 dB indoors and 76 dB outdoors, with a 20 MHz sampling rate in both environments, and sweep scanning is used across bands such as 900 MHz, 2.4 GHz, and 5.8 GHz. The SDR center frequency is tuned to the UAV channel so that signals with 20 MHz bandwidth or less are fully captured and the resulting spectrogram contains a centralized signal.
Recordings are stored in .dat files whose names or directory paths encode the manufacturer, model, center frequency, bandwidth, and operational mode. Outdoor directory names encode an extended set of fields: device, status, environment, SDR gain, splitter flag, recording duration, distance, altitude, center frequency, drone center frequency, bandwidth, SNR, sampling rate, and recording directory. A GNU Radio signal playback module lets the team reload files to verify that recorded signals match the predefined parameters and to estimate SNR and channel conditions.
For machine learning, raw complex I/Q streams are converted into time-frequency representations with the STFT, squared to a power spectrogram, converted to dB, and mapped to RGB with a perceptual colormap, then exported as PNG images alongside per-sample metadata. Before any of that, augmentation is applied on the baseband: noise is added for a target SNR, signals are mixed with interferers, and frequency shifts are injected. Because the transformations happen before the spectrogram is drawn, the resulting images reflect genuine RF phenomena rather than image-level edits — and because the software knows exactly what transform was applied, it can analytically recompute the YOLO bounding boxes, including splitting boxes that wrap around the frequency axis and filtering out boxes below a configurable minimum height. The released toolkit also normalizes third-party dataset layouts (such as Roboflow-style hierarchies with images/ and labels/ folders and train/val/test splits) so the same augmentation and training code can run on other public benchmarks.
Why This Matters
Impact on research. The paper argues the field is shifting from model-centric to data- and system-centric evaluation. CDRF's explicit tiering of SNR and interference conditions, release of raw I/Q for reproducible reprocessing, and label-consistent transforms give researchers a way to stress-test models rather than report inflated scores from clean conditions. The dataset-agnostic tooling is intended to make results comparable across benchmarks, and the open-set recognition implementation addresses the practical case where a detector meets a drone model it has never seen.
Real-world applications:
- Airport perimeter and airspace monitoring, motivated by the high-profile disruptions near airports cited in the paper.
- Protection of critical infrastructure and sensitive facilities, where unauthorized flights have been repeatedly reported.
- Day/night and non-line-of-sight surveillance, since most consumer and professional drones maintain continuous RF links for command-and-control and video transmission.
- Low-cost, edge-deployable counter-UAV sensing, enabled by the deliberate 20 MHz sampling choice and lightweight YOLO-based detection on spectrograms.
Industry relevance. The work is a direct collaboration between Rowan University's Department of Electrical and Computer Engineering and AeroDefense, and the dataset, code, and trained models are publicly released, with a separate data-request route. That makes it relevant to counter-UAV vendors, SDR and edge-hardware developers, spectrum regulators dealing with congested ISM bands, and integrators building multi-sensor fusion systems, where RF is one complementary modality alongside radar, vision, acoustics, LiDAR, and thermal sensing.
Future Directions
- Real-world generalization under stress. The paper frames the central open problem as moving from narrow, clean benchmarks to realistic evaluation under varied SNRs, interference, and frequency offsets. Whether CDRF-trained models actually transfer better to operational deployments is the question the benchmark is built to answer.
- Detection of autonomous drones. The paper notes a key limitation of RF sensing: drones that operate autonomously and lack continuous RF communication are not covered by this modality, which points toward multimodal fusion with radar, vision, acoustics, or other sensors.
- Cross-benchmark and parameter-transfer studies. Because all spectrogram parameters (sampling rate, FFT size, segment length, hop size, colormap, normalization) are exposed, an open line of work is ablating them and re-rendering datasets under different hardware targets — something the paper argues was practically impossible with prior fixed, pre-rendered releases.
- Calibration and open-set analysis. The YOLO patch exposing per-detection class probabilities and the open-set recognition implementation are positioned as enablers for calibration studies and handling previously unseen classes during inference; the paper states that Section VII discusses further future research directions, but those specifics are not included in the provided content.
Target Audience
This paper is most useful to researchers and engineers working on RF-based drone detection and counter-UAV systems, particularly those who need raw I/Q data and want to control the preprocessing pipeline themselves. It will also benefit machine learning practitioners applying classification and object detection to time-frequency representations, benchmark builders interested in label-consistent augmentation design, and industry teams evaluating edge-deployable RF sensing. Readers looking primarily for quantitative detection or classification results should note that the benchmark results section is not included in the content analyzed here.
Authors’ abstract
We present CageDroneRF (CDRF), a large-scale benchmark for Radio-Frequency (RF) drone detection and identification built from real-world captures and systematically generated synthetic variants. CDRF addresses the scarcity and limited diversity of existing RF datasets by coupling extensive raw recordings with a principled augmentation pipeline that (i)~precisely controls Signal-to-Noise Ratio (SNR), (ii)~injects interfering emitters, and (iii)~applies frequency shifts with label-consistent bounding-box recomputation for detection. The dataset spans a wide range of contemporary drone models, many of which are unavailable in current public datasets, and diverse acquisition conditions, derived from data collected at the Rowan University campus and within a controlled RF-cage facility. CDRF is released with interoperable open-source tools for data generation, preprocessing, augmentation, and evaluation that also operate on existing public benchmarks. It enables standardized benchmarking for classification, open-set recognition, and object detection, supporting rigorous comparisons and reproducible pipelines. By releasing this comprehensive benchmark and tooling, we aim to accelerate progress toward robust, generalizable RF perception models.