Skip to content
AI.info

Research

FOD-S2R: A FOD Dataset for Sim2Real Transfer Learning based Object Detection

Overview Research area: Computer vision for aviation safety — object detection datasets, synthetic data generation, and Sim2Real transfer learning. Technical level: Intermediate. Scope: The paper intr

FOD-S2R: A FOD Dataset for Sim2Real Transfer Learning based Object Detection
arXiv
2512.01315
Published
2025-12-01
Authors
Ashish Vashist, Qiranul Saadiyean, Suresh Sundaram, Chandra Sekhar Seelamantula

AI summary

Overview

  • Research area: Computer vision for aviation safety — object detection datasets, synthetic data generation, and Sim2Real transfer learning.
  • Technical level: Intermediate.
  • Scope: The paper introduces FOD-S2R, a paired real-and-synthetic dataset of Foreign Object Debris (FOD) inside a simulated aircraft fuel tank, and benchmarks anchor-based and anchor-free detectors plus Sim2Real adaptation strategies on it.

What This Paper Is About

Foreign Object Debris — loose nuts, bolts, tools, metal fragments — inside an aircraft fuel tank can cause fuel contamination, system malfunctions, and higher maintenance costs, but detecting it is hard because fuel tanks are enclosed, poorly lit, cluttered, and full of structures that resemble debris. Existing FOD datasets were collected outdoors, on runways and aprons, so they do not capture the visual constraints of a closed tank interior, and collecting real annotated data inside aviation structures is expensive and restricted by safety rules. The paper's goal is to build a dataset that pairs a limited real-world subset with a large synthetic subset and to test whether synthetic images can improve real-world FOD detection.

Key Contributions

  1. A hybrid FOD dataset for enclosed environments. FOD-S2R contains 6,250 high-resolution images: 3,114 real-world images captured in a custom-built fuel tank replica and 3,137 synthetic images generated in Unreal Engine, spanning 14 classes. The paper states it is the first dataset to systematically evaluate synthetic data for FOD detection in confined, closed structures.
  2. A hybrid annotation pipeline. Initial manual annotations were used to train a YOLOv11 model, which then generated annotations for the rest of the dataset, with every box iteratively reviewed by hand for alignment, non-overlap, and tight fitting to object borders. The synthetic subset uses Unreal Engine Blueprint scripting to produce bounding boxes automatically.
  3. A benchmark of state-of-the-art detectors. Anchor-based and anchor-free models — YOLOv12, YOLOv11, YOLOv8, YOLOv5, RT-DETR, DDQ, Faster R-CNN, RetinaNet, and RF-DETR — were evaluated on both the real and synthetic subsets, including scale-aware metrics for small, medium, and large objects.
  4. A Sim2Real transfer study. Two adaptation orders were tested: synthetic pretraining followed by real fine-tuning, and the reverse, using RF-DETR and YOLOv12.

Main Findings

  • A measurable domain gap exists. For YOLOv11, mAP50 is 0.845 on synthetic data and rises to 0.929 on real images, while mAP50:95 increases from 0.568 to 0.715. YOLOv5 and YOLOv8 show the same pattern, indicating real samples carry more detectable structural cues.
  • YOLOv12 leads on the synthetic subset. It reports mAP50 = 0.94 and mAP50:95 = 0.718 on synthetic data, and 0.903 and 0.682 on real data.
  • RF-DETR is strongest overall on real data. It reaches mAP50 = 0.930 and mAP50:95 = 0.702, and the highest reported mAP75 = 0.782 on real images. Its small-object score mAPS = 0.383 shows degraded sensitivity to small debris such as washers or cable ties.
  • Small-object detection is the recurring bottleneck. RT-DETR scores mAPS = 0.640 on synthetic data but drops to mAPS = 0.601 on real data. DDQ performs worst overall, with real mAP50 = 0.471 and mAP50:95 = 0.360.
  • Synthetic-first training gives the best real-world result. Pretraining on synthetic data and fine-tuning on real images reached mAP50 = 0.931 and mAP50:95 = 0.740 on the real test set, and raised small-object performance from mAPS = 0.383 in the real-only baseline to mAPS = 0.676.
  • Reversing the order hurts. Real-first, synthetic-second adaptation produced mAP50 = 0.913 but mAP50:95 = 0.712, which the authors attribute to the synthetic domain lacking the high-frequency surface irregularities of real fuel tanks. Note that Table II's row labels appear swapped relative to this narrative — the table assigns the 0.931 / 0.740 numbers to the "train real, fine-tune synthetic" row, while the paper's text assigns them to the synthetic-first strategy.
  • YOLOv12 shows a much smaller ordering effect. Synthetic-then-real fine-tuning gives mAP50 = 0.902 and mAP50:95 = 0.69, versus 0.898 and 0.679 for real-then-synthetic, a nearly negligible difference.
  • Cross-domain drops of 10–20%. Several models lose 10–20% on mAP50:95 when transitioning between domains, confirming generalization without adaptation is nontrivial. The conclusion states that integrating synthetic data improves detection performance by 15%.

Methodology in Plain English

The authors built a digital twin of a fuel tank. A detailed CAD model was created in Blender using real tank dimensions, including joints, rivets, and internal structural features, then imported into Unreal Engine 5. Textures representing grease stains, surface irregularities, scratches, and damaged rivets were applied, and FOD objects were placed at arbitrary positions. Camera pose, field of view, and lighting were controlled — one illustrated setting uses a field of view of 70 with light intensity of 9 cd, compared against a field of view of 25 with 3 cd — along with variation in FOD color, fuel tank color (yellow, light yellow, gray), and object size. Bounding boxes were produced automatically through Blueprint scripts.

For the real subset, no actual aircraft fuel tank was available, so a controlled replica setup was constructed with structural reinforcements that create realistic occlusions. Images were captured with a Canon DSLR at 25-megapixel resolution from multiple distances, zoom levels, top-down and horizontal perspectives, plus frames extracted from video. FOD items of metal, plastic, rubber, cloth, and electronic materials were placed in partially hidden locations under low-light, direct reflection, and diffuse lighting.

All models used a ResNet-50 backbone pretrained on ImageNet, implemented in the Detectron2 framework on two NVIDIA RTX 2080 GPUs with 12 GB each. Data was split 70% train, 10% validation, 20% test, tiled into 640 × 640 pixel tiles that each contain at least one FOD instance. Augmentation included saturation perturbation between −20% and +20%. Experiments ran for 100 epochs with a fixed learning rate of 0.001 and batch size of 1. Three training configurations were compared: real-only, synthetic-only, and synthetic-pretrained then fine-tuned on real data.

Why This Matters

This is the first FOD dataset aimed at enclosed, internally complex aviation spaces rather than open runways, and it provides both the data and a structured benchmark for the Sim2Real problem in that setting. The headline practical claim is that synthetic data can serve as a scalable initialization resource, reducing the amount of expensive, safety-restricted real annotation needed for safety-critical inspection.

Real-world applications:

  • Aircraft fuel tank inspection, where automated detection could catch debris before it contaminates fuel or damages systems.
  • Robotic or borescope-based internal inspection of tanks, bays, and other confined structures where a camera is the only access.
  • Runway and apron FOD detection, since the benchmarked detectors and Sim2Real workflow apply to the older outdoor problem too.
  • General industrial enclosed-space inspection, such as machinery housings and maintenance bays, where backgrounds resemble defects.

Industry relevance: the work targets aviation maintenance organizations that currently rely on manual inspection, and it demonstrates a pipeline — synthetic pretraining plus limited real fine-tuning — that lowers data collection cost while improving generalization to unseen viewpoints and object configurations.

Future Directions

  • Verify and extend the Sim2Real ordering result. Only RF-DETR and YOLOv12 were used for adaptation experiments, and the Table II label ordering conflicts with the narrative, so the synthetic-first advantage needs confirmation across more architectures.
  • Validate on actual aircraft fuel tanks. The real subset came from a replica because no real fuel tank was available; testing on operational tanks would reveal how much of the remaining domain gap is due to the replica itself.
  • Improve small-object detection. The persistent weakness is mAPS, with RF-DETR as low as 0.383 on real data; better strategies for tiny metallic debris such as washers and fragments are unresolved.
  • Explore richer adaptation techniques. The paper compared only two fine-tuning orders despite discussing domain randomization, leaving other Sim2Real methods, greater lighting and occlusion diversity, and depth or video-based detection as open directions.

Target Audience

Computer vision researchers working on object detection, dataset construction, and Sim2Real transfer; aviation safety and maintenance engineers interested in automated FOD detection; and practitioners building inspection systems for confined industrial environments where real annotated data is scarce.

Authors’ abstract

Foreign Object Debris (FOD) within aircraft fuel tanks presents critical safety hazards including fuel contamination, system malfunctions, and increased maintenance costs. Despite the severity of these risks, there is a notable lack of dedicated datasets for the complex, enclosed environments found inside fuel tanks. To bridge this gap, we present a novel dataset, FOD-S2R, composed of real and synthetic images of the FOD within a simulated aircraft fuel tank. Unlike existing datasets that focus on external or open-air environments, our dataset is the first to systematically evaluate the effectiveness of synthetic data in enhancing the real-world FOD detection performance in confined, closed structures. The real-world subset consists of 3,114 high-resolution HD images captured in a controlled fuel tank replica, while the synthetic subset includes 3,137 images generated using Unreal Engine. The dataset is composed of various Field of views (FOV), object distances, lighting conditions, color, and object size. Prior research has demonstrated that synthetic data can reduce reliance on extensive real-world annotations and improve the generalizability of vision models. Thus, we benchmark several state-of-the-art object detection models and demonstrate that introducing synthetic data improves the detection accuracy and generalization to real-world conditions. These experiments demonstrate the effectiveness of synthetic data in enhancing the model performance and narrowing the Sim2Real gap, providing a valuable foundation for developing automated FOD detection systems for aviation maintenance.

Read the original paper