Research
Real-Time Wildfire Localization on the NASA Autonomous Modular Sensor using Deep Learning
Overview Research area: Computer vision for remote sensing — specifically deep learning-based wildfire detection and pixel-level segmentation on multi-spectral, high-altitude aerial imagery. Technical

- arXiv
- 2601.14475
- Published
- 2026-01-20
- Authors
- Yajvan Ravan, Aref Malek, Chester Dolph, Nikhil Behari
AI summary
Overview
Research area: Computer vision for remote sensing — specifically deep learning-based wildfire detection and pixel-level segmentation on multi-spectral, high-altitude aerial imagery.
Technical level: Intermediate. The paper is readable for someone with basic familiarity with convolutional neural networks and remote sensing, though the spectral band discussion (SWIR, IR, thermal, reflectance calibration) assumes some domain background.
Scope in one sentence: The paper introduces a human-annotated, 12-channel dataset from the NASA Autonomous Modular Sensor (AMS) and trains a two-tier classification-plus-segmentation model that localizes active wildfire in high-altitude aerial imagery in simulated real time.
What This Paper Is About
Current wildfire mapping depends heavily on human firespotters who manually label fire perimeters from infrared aerial imagery, with results sometimes reaching response teams up to 12 hours after the flight. Existing automated approaches rely on satellite imagery, low-altitude UAV footage, RGB-only data, or offline processing — none of which match the operational conditions of modern high-altitude wildfire missions. This paper builds a deep learning pipeline that classifies whether fire is present and, if so, segments the exact fire perimeter, using the same medium-to-high altitude (3-50 km) multi-spectral imagery that US wildfire agencies already collect.
Key Contributions
-
A manually annotated AMS dataset. The authors created and published a human-labeled wildfire dataset from the NASA Autonomous Modular Sensor, combining 12 spectral channels including infrared (IR), short-wave IR (SWIR), and thermal bands. The abstract states imagery came from 20 wildfire missions and produced over 4000 images; Section 3.1 states 18 observation missions spanning 2006 to 2019 were used, yielding 4259 training patches (some with overlap) and 85 testing patches (no overlap). The dataset deliberately includes difficult cases: cloud and smoke occlusion, easily confused false positives, and nighttime imagery.
-
A two-tier real-time localization model. A classification network first decides whether a 256×256-pixel patch contains fire; only if it does does the segmentation network run. The paper reports the combined system runs at over 100 fps on NVIDIA T1200 laptop GPUs, and that each AMS flight took roughly 5-10 minutes and generated 200-300 patches of size 256×256 pixels.
-
Evidence that multi-spectral and human-annotated data both matter. The authors benchmark against Pereira et al.'s Landsat-trained segmenter and against classical color-rule algorithms, and run spectral ablations showing that bands 9, 10, 11, and 12 carry most of the useful fire information.
-
A description of a simulated real-time inference pipeline. Test images are split into a grid of non-overlapping patches fed in one at a time from left-to-right and top-to-bottom, mimicking the line-by-line scanning behavior of wildfire observation aircraft.
Main Findings
-
Headline performance: In Table 3, the classifier reached 96.8% accuracy, 82.1% precision, 77.8% recall, and 1.19 ms inference time. The segmenter reached 96.0% accuracy, 58.6% precision, 84.0% recall, 74.0% IoU, and 7.75 ms inference time. (The abstract rounds these to 96% classification accuracy, 74% IoU, and 84% recall.) Results were averaged over 10 training seeds.
-
It beats the satellite-trained baseline: Pereira et al.'s segmenter scored 95.5% accuracy, 53.2% precision, 74.8% recall, and 47.5% IoU on the AMS test set. The authors attribute their improvement to human-annotated data and IR spectral data, and note it also implies a domain gap between satellite imagery and airborne high-altitude imagery.
-
The trimmed model also wins: When retrained on only bands 10, 9, and 2 — directly comparable to the Landsat network's band selection — the authors report an IoU of 75% versus 47.5%, suggesting that the human annotations alone provide a large part of the gain.
-
Thermal and IR bands matter most: In the classification ablation, bands 1-8 (non-thermal) achieved 89.8% accuracy, 30.0% precision, and 4.4% recall, while bands 5, 3, 2 achieved 89.5% accuracy, 12.2% precision, and 7.8% recall. Bands 9 and 10 alone performed comparably to the full model, while bands 11 and 12 showed close but notably worse recall for classification and worse precision for segmentation.
-
The single best ablation combination: Bands 10, 9, 2 gave 97.4% accuracy, 83.3% precision, and 84.4% recall for classification, and 96.6% accuracy, 64.0% precision, 76.8% recall, and 75.0% IoU for segmentation.
-
Highest precision from thermal alone: Band 12 alone produced 97.5% precision for classification (with 63.3% recall), and band 10 alone produced 76.4% IoU for segmentation.
-
Detection under obscuration: The paper states the model detects active wildfire at nighttime and behind clouds, and can distinguish false positives — cases where RGB-only approaches would fail.
-
Fire is a tiny fraction of each image: Active wildfire averages roughly 2% of an image, and only approximately 18% of patches contain more than 0.5% active fire pixels. The paper notes 82% of the dataset contains no fire, which is what makes the two-tier skip mechanism useful.
-
Speed rationale: The segmentation network contains 500 times the number of parameters in the classification network and has about 7× the inference time, so skipping segmentation on fire-free patches saves substantial computation.
Methodology in Plain English
The authors took raw flight data from the NASA AMS sensor and normalized it. For non-thermal bands (1-11), calibrated radiance was divided by band-averaged solar irradiance and clipped between 0 and 1. For thermal band 12, radiance was converted to brightness temperature and clipped between 250 and 500 K. All data was resampled to a consistent 10 m per pixel. Only bands 1-12 were used, since bands 13-16 are low-gain versions of bands 1-12 with less temperature resolution.
Because aerial missions are expensive and datasets are small, the authors exploited the large area each flight covers: 250-350 non-overlapping 256×256 patches can be cut from each image, and patches from the same image can look very different in terrain, smoke, and development. This patching also matches how an aircraft scans laterally, producing the image in series.
Humans then inspected every available spectral band to hand-mark active fire pixels. Two networks were trained. The classifier is a three-layer encoder (convolution plus ReLU, batch normalization, max pooling), ending in global max pooling and a dense sigmoid layer, trained for 500 epochs with binary cross-entropy and Adam at an initial learning rate of 7.5×10⁻⁴, with positive examples resampled to offset class imbalance. The segmenter uses two convolution/downsampling layers followed by two deconvolution layers with skip connections, trained only on positive examples with weighted binary cross-entropy (weight 50) and Adam at 3×10⁻⁴.
At inference the two networks are chained: the classifier gates the segmenter. For the real-time simulation, test images are chopped into a grid of 256×256 patches and fed through one at a time at the rate they were acquired according to flight logs, with blank masks emitted where the classifier says no fire.
Why This Matters
Impact on research: This is one of the few public, human-annotated, pixel-wise wildfire segmentation datasets at medium-to-high altitude with a wide spectral range, filling a gap between low-altitude UAV datasets and satellite datasets. The ablations give concrete guidance about which spectral bands carry fire information, and the benchmark against a satellite-trained model quantifies the domain gap between spaceborne and airborne remote sensing for this task.
Real-world applications:
- Wildfire incident response: automatically generating fire perimeter products that today require trained human firespotters, potentially cutting the lag between overflight and response.
- Nighttime and obscured-condition mapping: since the model uses IR and thermal bands, it can operate when visible-spectrum imagery is unusable due to darkness, smoke, or cloud cover.
- Sensor design prioritization: the finding that bands 9-12 dominate suggests future instruments should prioritize SWIR, IR, and thermal imaging.
- Human-in-the-loop triage: the paper notes that an initial deployment could run only the classifier so experts can ignore irrelevant portions of an image while still hand-labeling patches that contain fire.
Industry relevance: The work targets the operational modality actually used by the US Forest Service's NIROPS program, California's FIRIS, and Colorado's DFPC Multi-Mission Aircraft — line-scanner systems producing geo-rectified TIFFs and shapefiles, with scan volumes around 100k-400k acres per hour across multiple large fires per sortie. A model that runs at over 100 fps on laptop-class GPUs is plausible to deploy on existing aircraft processing hardware rather than requiring new ground infrastructure.
Future Directions
- Minimizing false negatives: the authors state that for field deployment, high recall matters more than precision because missed fire is more consequential than a false alarm.
- Pipeline integration: they flag that data preprocessing needed to feed the model may not be negligible for real-time performance once embedded in an operational system.
- Human-in-the-loop automation: the paper argues humans remain necessary until models like this become robust and reliable, so cooperative workflows need study.
- Applying newer computer vision architectures: the authors note that many recent techniques have yet to be applied to this problem domain.
Target Audience
Researchers and engineers working on remote sensing, aerial and satellite computer vision, and disaster response automation — particularly those interested in multi-spectral segmentation or in bridging machine learning research and operational wildfire management. It is also relevant to wildfire agency technologists evaluating whether automated perimeter detection could augment or partially replace manual firespotter workflows, and to dataset builders looking for an example of patching-based augmentation for scarce aerial imagery.
Authors’ abstract
High-altitude, multi-spectral, aerial imagery is scarce and expensive to acquire, yet it is necessary for algorithmic advances and application of machine learning models to high-impact problems such as wildfire detection. We introduce a human-annotated dataset from the NASA Autonomous Modular Sensor (AMS) using 12-channel, medium to high altitude (3 - 50 km) aerial wildfire images similar to those used in current US wildfire missions. Our dataset combines spectral data from 12 different channels, including infrared (IR), short-wave IR (SWIR), and thermal. We take imagery from 20 wildfire missions and randomly sample small patches to generate over 4000 images with high variability, including occlusions by smoke/clouds, easily-confused false positives, and nighttime imagery. We demonstrate results from a deep-learning model to automate the human-intensive process of fire perimeter determination. We train two deep neural networks, one for image classification and the other for pixel-level segmentation. The networks are combined into a unique real-time segmentation model to efficiently localize active wildfire on an incoming image feed. Our model achieves 96% classification accuracy, 74% Intersection-over-Union(IoU), and 84% recall surpassing past methods, including models trained on satellite data and classical color-rule algorithms. By leveraging a multi-spectral dataset, our model is able to detect active wildfire at nighttime and behind clouds, while distinguishing between false positives. We find that data from the SWIR, IR, and thermal bands is the most important to distinguish fire perimeters. Our code and dataset can be found here: https://github.com/nasa/Autonomous-Modular-Sensor-Wildfire-Segmentation/tree/main and https://drive.google.com/drive/folders/1-u4vs9rqwkwgdeeeoUhftCxrfe_4QPTn?=usp=drive_link