Skip to content
AI.info

Research

PyroAdapt: Adapting Wildfire Prediction under Spatial Heterogeneity and Temporal Shift

Overview Research area: Machine learning for environmental hazard prediction — specifically wildfire occurrence forecasting under distribution shift, combining retrieval-augmented adaptation, pairwise

PyroAdapt: Adapting Wildfire Prediction under Spatial Heterogeneity and Temporal Shift
arXiv
2605.12435
Published
2026-05-12
Authors
Enyi Jiang, Wu Sun

AI summary

Overview

Research area: Machine learning for environmental hazard prediction — specifically wildfire occurrence forecasting under distribution shift, combining retrieval-augmented adaptation, pairwise ranking losses, and preference-optimization theory.

Technical level: Intermediate. The empirical setup and motivations are accessible, but the core machinery (pairwise ranking objectives, residual pairwise DPO, score-gap gradient analysis, transductive retrieval) assumes familiarity with loss functions and ranking metrics. The paper includes a fairly mathematical theory section (Propositions 1, 2, 4, 5).

Scope in one sentence: The paper proposes PyroAdapt, a pretrain–retrieve–rank framework that adapts a historically pretrained wildfire risk model to a new year using unlabeled target covariates and retrieved historical labels, and evaluates it on a nine-cell Yosemite grid (temporal shift) and the full 666-cell California domain (spatial prioritization under a daily budget).

What This Paper Is About

Wildfire occurrence prediction is hard because fire-positive location–days are extremely rare compared with non-fire ones, because the relationship between weather/fuel conditions and fire changes from year to year, and because the same conditions produce different risk in different landscapes. Models trained on historical fire records can therefore score poorly in the year they are actually deployed. The authors build PyroAdapt to close that gap: it keeps target-year labels hidden, retrieves historically labeled examples whose conditions resemble the target batch, and fine-tunes the pretrained model with a pairwise ranking loss so that fire-prone locations get ranked above non-fire locations.

Key Contributions

  1. A unified adaptation framework. PyroAdapt combines historical risk pretraining, target-conditioned KNN retrieval of labeled historical neighbors, and ranking-based fine-tuning, all without using target-period labels. The retrieved set is a deduplicated union of each target input's k nearest historical neighbors, so target density does not implicitly reweight the loss.

  2. A budgeted-ranking certificate and a selective target. The authors analyze direct ranking, residual pairwise DPO (RDPO), and selective ranking through one score-gap formulation that sets the loss target τ to 0, the frozen pretrained gap, or the positive part of that gap. They show that same-day pairwise losses with nonnegative targets bound the number of avoidable misses under a fixed daily budget (Proposition 5), and give selective ranking an additional positive-reference-margin certificate.

  3. Gradient-allocation theory for the three targets. Proposition 1 shows how decreasing α increases the initial gap-gradient magnitude on pairs the reference model ordered wrongly, while increasing α increases it on pairs it ordered correctly; at α = 1 the magnitude equals 1/2 for every pair. Proposition 2 bounds how sensitive the per-pair gradient weight is to the choice of target, and Proposition 4 links the binary-label RDPO preference-risk minimizer to cost-sensitive logit adjustment.

  4. Daily prioritization over a spatially heterogeneous grid. For California the authors condition risk on terrain, learned four-dimensional ecoregion embeddings, and prior-year fire rates, and restrict pairs to same-day cells so the ranking term cannot gain by raising all scores on a high-risk date.

Main Findings

  • Ranking beats continued focal fine-tuning in California. On the full 666-cell, 0.25° grid in 2022, daily average precision rises from 21.62% for continued focal fine-tuning to 24.35% (direct ranking), 24.47% (RDPO), and 24.57% (selective ranking). Daily Top5% recall rises from 18.70% to 22.25%, 22.11%, and 22.79% respectively. All three exceed the pretrained spatial focal model (21.10% daily AP, 17.98% daily R@5%) and an XGBoost baseline (14.24% daily AP, 11.64% daily R@5%).

  • More captures at the same alarm budget. With a fixed daily budget of 34 cells (5% of the 666-cell domain, totaling 12,410 alarms across the 2022 evaluation), selective ranking captures a mean of 3,210.3 neighborhood-smoothed positive cell–days versus 2,866.0 for continued focal — 344 additional captures, roughly 12%. Direct and RDPO add 311 and 301 respectively. The authors note these are smoothed positive cell–days, not independent incidents.

  • Larger gains on the most severe fires. Among 2,193 raw-DM-positive test examples, evaluated at each policy's historical F1-selected threshold, selective ranking raises recall by 39.70, 28.18, and 20.50 percentage points within the highest-dry-matter 5%, 10%, and 20% strata; direct ranking gains 38.48/27.27/20.27 and RDPO 38.18/26.97/19.89. Alarm rates are unmatched in this diagnostic, so the authors state it does not establish severity forecasting or equal-budget severe-event discrimination.

  • Ranking gains persist under temporal shift in Yosemite. On the fixed 3×3 grid in 2021, direct ranking raises AUROC from 70.32% (pretrained) to 71.23% and AUPRC from 21.23% to 22.82%; selective ranking reaches 71.51% AUROC and 22.94% AUPRC; the pointwise binary-label RDPO baseline reaches 71.00% and 22.59%.

  • Rolling evaluation supports explicit pairwise fine-tuning. Across 2019–2021, direct ranking attains mean AUROC of 74.19% and AUPRC of 24.30%, versus 73.12%/22.85% for the pretrained model and 73.08%/22.08% for continued focal. Relative to binary-label RDPO, direct ranking gains 2.06 AUROC, 2.25 AUPRC, 3.69 R@20, and 5.97 R@30 percentage points (differences computed before rounding). The authors note the evaluated binary-label RDPO configuration with no focal supervision and a fixed schedule shows unstable performance across test years, and that this is specific to that configuration rather than a general limitation of RDPO.

  • Local adaptation beats training from scratch. A joint focal + direct-ranking model trained from random initialization on all historical rows attains rolling AUROC of 71.73% and AUPRC of 21.43%, compared with 74.19% and 24.30% for local direct adaptation — a paired local-minus-joint gain of 2.46 AUROC points and 2.86 AUPRC points.

  • Retrieval itself contributes, modestly. A three-seed continued-focal ablation at exactly 6,800 updates compares spatial KNN retrieval against a size-matched random subset and the full historical sample. KNN improves mean daily AP by 0.66 points over random and 0.36 over full-history adaptation, with positive paired differences in every seed. The authors describe Top5% recall gains from this ablation as less conclusive.

  • Data scale used. California pretraining uses 2002–2021 data and targets 2022; a fixed 20% historical sample contains 972,892 cell–days; k=5 retrieval queries all 243,090 unlabeled target covariates, producing 363,643 unique historical examples for fine-tuning; evaluation covers 14,760 positives across 351 fire dates. Yosemite's main split has 62,460 training grid–days through 2020 (7,613 positives) and 3,285 target grid–days in 2021 (363 positives).

Methodology in Plain English

The pipeline has three stages.

Pretrain. A TabNet model is trained with focal loss on labeled historical years to produce a risk scorer and a checkpoint. This checkpoint serves two roles: it initializes every adaptation variant, and — frozen in evaluation mode — its scores define the loss targets used by RDPO and selective ranking. Reference scores are detached so gradients only update the adapted policy.

Retrieve. Continuous covariates are standardized with historical training statistics. Each unlabeled target input finds its k=5 nearest labeled historical inputs in that covariate space (k is fixed at 5 throughout). The union of all retrieved rows, deduplicated, becomes the local adaptation set. This is a sample-selection step, not label propagation: target inputs never inherit a neighbor's label. The authors acknowledge the method assumes useful overlap between historical and target covariates, and that KNN returns neighbors even when they are far away, so it cannot reconstruct a target regime simply absent from the archive.

Rank. Retrieved historical labels are converted into positive–negative pairs, and fine-tuning minimizes a softplus-over-score-gap loss with no focal supervision. The three objectives differ only in the target gap: direct ranking uses zero, RDPO uses the frozen pretrained gap, and selective ranking uses the positive part of that gap (equivalently, RDPO on correctly ordered pairs and direct ranking on misordered ones). For California, pairs are drawn from the same day only, so an additive date offset cancels and the loss must resolve which cells within a date are riskier, matching the daily top-B decision. Location-dependent inputs include previous-day weather, coordinates and terrain, prior-year fire-rate history from strictly earlier years, seasonality, and a learned four-dimensional ecoregion embedding; retrieval distances use standardized continuous covariates plus a fixed ecoregion one-hot vector. Hyperparameters in California are β_rank = λ_rank = 0.1 and λ_s = 0.05. After adaptation, the retrieved set and frozen reference are discarded; inference is a single forward pass.

Why This Matters

Impact on research. The paper frames hazard prediction as an adaptation problem rather than a static classification problem, and gives a single score-gap lens under which direct ranking, RDPO, and selective ranking are special cases. That unification — plus the gradient-allocation propositions and the budgeted-miss certificate — offers a reusable template for other rare-event, spatially heterogeneous, temporally shifting prediction tasks where only top-k decisions matter operationally.

Real-world applications:

  • Allocating a fixed number of daily fire-monitoring or patrol slots across a large jurisdiction, where the operational question is which locations to watch, not a probability for every pixel.
  • Triaging field inspection or fuel-treatment planning in the specific regions and ecoregions that rank highest on a given day.
  • Flagging days and places likely to host extreme fire events, using the retrospective high-dry-matter recall diagnostic as an early sensitivity check.
  • Adapting an existing operational risk model to a new fire season without waiting for that season's labels, since the method is batch-transductive but label-free for the target period.

Industry relevance. Relevant to electric utilities, insurers, forestry and land-management agencies, satellite-based fire-monitoring vendors, and emergency-management operations centers — any organization that must convert a finite daily monitoring or response capacity into a ranked list of locations. The explicit budget framing (34 cells per day, 12,410 alarms) is the kind of constraint those groups actually plan against.

Future Directions

  • Handling target regimes absent from the archive. The authors state retrieval cannot reconstruct a target regime not represented historically. A natural next step is a method that detects and flags such cases rather than always returning k neighbors.
  • Regional fairness auditing. The ethics statement notes that location prioritization can disadvantage regions with sparse historical support, since history features shrink such cells toward a statewide prior, and recommends auditing performance by region rather than in aggregate. That audit is not reported here.
  • Equal-budget severe-event discrimination. Because the high-dry-matter diagnostic uses unmatched alarm rates and each policy's own historical F1-selected threshold, the paper explicitly does not establish severity forecasting. Matching alarm rates would test whether the gains survive a fair severity comparison.
  • Extension beyond the evaluated configurations. The authors report unstable year-to-year performance for the specific binary-label RDPO configuration tested, flag the Top5% recall gains in the retrieval ablation as inconclusive, and mention broader-area extensions in Appendix E — all open threads for follow-up work.

Target Audience

Machine learning researchers working on distribution shift, retrieval-augmented adaptation, and learning-to-rank; environmental and climate data scientists building operational fire-risk systems; and practitioners in utilities, insurance, land management, and emergency response who need to prioritize a limited daily monitoring budget across a large, ecologically diverse domain. Readers without a ranking-loss background will still follow the problem framing and empirical results, but will need to work through the theory section for the gradient-allocation arguments.

Authors’ abstract

Prediction of wildfire occurrence is a rare-event problem compounded by spatial heterogeneity and temporal distribution shift, as fire occurrences are vastly outnumbered by non-occurrences, and predictor--fire relationship varies across space and time. Models trained on historical fire data may perform poorly under new conditions and require adaptation to the target distribution before operational use. We propose PyroAdapt, a pretrain--retrieve--rank framework that adapts a pretrained model to target conditions by retrieving historical locations with similar conditions and fine-tuning on the retrievals through risk ranking. For spatial adaptation, we condition risk on terrain, ecoregion embeddings, and fire rates, accounting for spatial context in the retrieval, and learn risk ordering from same-day fire--nonfire cell pairs. We compare direct ranking, residual pairwise DPO (RDPO), and selective ranking through a unified score-gap formulation that characterizes their gradient allocation. Over California (discretized into 666 0.25x0.25 grid cells), these objectives raise daily average precision from 21.62% for continued focal fine-tuning to 24.35--24.57%, and Top5% recall from 18.70% to 22.11--22.79%. Under a fixed daily detection budget of 34 cells (5% area), selective ranking captures 344 additional positive cell--days. For fires in the top 5%/10%/20% of dry matter consumption, selective ranking raises recall by 39.70/28.18/20.50 percentage points, respectively. Furthermore, rolling evaluations over Yosemite show that the gains from ranking persist under temporal distribution shift. Together, these results show that PyroAdapt prioritizes the most fire-prone locations under a daily budget constraint and detects more extreme fire events.

Read the original paper