Skip to content
AI.info

Research

Power Ensemble Aggregation for Improved Extreme Event AI Prediction

Power Ensemble Aggregation for Improved Extreme Event AI Prediction Overview Research area: Machine learning for weather and climate prediction, specifically the classification of extreme heat events

Power Ensemble Aggregation for Improved Extreme Event AI Prediction
arXiv
2511.11170
Published
2025-11-14
Authors
Julien Collard, Pierre Gentine, Tian Zheng

AI summary

Power Ensemble Aggregation for Improved Extreme Event AI Prediction

Overview

Research area: Machine learning for weather and climate prediction, specifically the classification of extreme heat events and the statistical aggregation of ensemble forecasts.

Technical level: Intermediate. The paper assumes familiarity with binary classification metrics (AUC), ensemble forecasting, and convolutional neural networks, but its central idea — the power mean — is mathematically simple.

Scope: The paper tests whether a non-linear power-mean aggregation of ensemble members from a generative weather model improves the detection of extreme heat events compared with averaging those members.

What This Paper Is About

Forecasting extreme events such as heat waves with machine learning is difficult because extremes are rare and behave differently from the bulk of the data, so models trained on average conditions tend to under-predict them. This paper reframes the problem as a binary classification task — will surface air temperature exceed its q-th local quantile at a given place and future day? — and asks whether the way a generative model's ensemble members are combined into a single score changes how well extremes are detected. The authors propose replacing the standard ensemble mean with a power mean, controlled by an exponent p, that can deliberately place more weight on members predicting an extreme.

Key Contributions

  1. A power-mean aggregation rule for extreme-event classification. The authors define the ensemble score as the p-th power mean of the individual member scores, s_i = Φ(x_i), rather than the arithmetic mean, with p ≥ 1 acting as a tunable hyperparameter that interpolates between averaging and taking the maximum.

  2. A generative weather model built for the task. They adapt a U-Net-style convolutional model on a cube-sphere grid, made generative by injecting Perlin noise into the inputs, and train it with a CRPS loss to produce ensembles of local temperature anomalies (n = 50 members).

  3. An empirically characterized scaling law for the optimal exponent. Across quantile thresholds, the optimal power exponent p_opt is found to increase almost perfectly exponentially with the quantile q, giving a simple way to estimate the exponent for any chosen threshold.

  4. A benchmark comparison against GraphCast and baselines. The power-mean method is evaluated against mean prediction, the deterministic GraphCast model, and a persistence baseline across quantiles and forecast lead times.

Main Findings

  • The power mean beats the ensemble mean for every quantile tested. On the test dataset (2015–2018), using 7-day-ahead predictions, the power-mean aggregation outperformed the mean prediction method for all quantile thresholds examined.

  • Effectiveness grows with the severity of the extreme. The relative improvement (RI), defined as 100 × (AUC_popt − AUC_mean pred) / AUC_mean pred, increases with the quantile threshold q, meaning the method helps most for the most extreme events.

  • The optimal exponent scales exponentially with the quantile. When p_opt was computed for different thresholds on the validation set (2010–2015), it increased almost perfectly exponentially with q, which the authors present as a simple estimate usable for any quantile.

  • A concrete example at q = 0.9. For the AUC-vs-p curve at q = 0.9 with a 7-day lead time, the optimal power exponent is p_opt ≃ 18.3 and the relative improvement is RI_opt ≃ 2.67%.

  • The AUC at p = 1 is not identical to the ensemble-mean AUC. The authors attribute this to the non-linearity of the normal quantile function Φ.

  • Advantage grows with forecast lead time. With exponents optimized at 7 days but applied to other lead times, the power-mean method outperformed the mean method for all quantiles and lead times, and the relative improvement appeared to increase with lead time as well as with quantile.

  • The simple generative model can overtake GraphCast on hard cases. For high quantiles and long lead times, the power-mean method applied to the authors' simpler model outperformed GraphCast, even though GraphCast is described as having around 36.7 million parameters versus approximately 1.2 million for the authors' model.

  • Raw forecast skill remains lower than GraphCast. On the global anomaly RMSE metric, the authors' generative model (using its mean prediction) was always worse than GraphCast, but better than persistence and remaining below the climatology reference out to 12 days.

Methodology in Plain English

The authors started from ERA5 reanalysis data (via the WeatherBench2 cloud datasets), downloaded at 1.5° spatial resolution, 6-hourly, covering 1990–2020, then resampled to daily resolution to remove the diurnal cycle and focus on longer-timescale extremes. They regridded the data onto a 6×48×48 gnomonic equiangular cube-sphere grid to avoid pole singularities, using the TempestRemap software and, where convenient, k-nearest-neighbors interpolation with k = 4.

They defined an extreme heat event locally: a temperature counts as q-extreme if its local anomaly x satisfies Φ(x) ≥ q, where Φ is the standard normal CDF and the anomaly is the temperature minus the local climatological mean, divided by the local climatological standard deviation. Using local rather than global definitions lets the method detect extremes anywhere and at any time of year, rather than only in permanently hot regions.

The prediction task is binary classification: given the atmosphere on day d and the preceding few days, predict whether surface temperature at each location will exceed its q-th local quantile on day d + Δd. Performance is measured with the area under the ROC curve (AUC).

Their model is a U-Net-style CNN adapted to the cube-sphere grid, taking atmospheric and ground variables from days d, d−1, d−2 and d−3 as input and outputting surface air temperature local anomalies for days d+1 through d+12. To make it generative, they inject Perlin noise — a smooth, spatially coherent gradient noise — into the network's input, following prior work, generating 3D Perlin noise on a unit cube and slicing the 2D surface. Their specific twist is randomizing the gradient vector amplitudes with a log-normal distribution so the noise can exceed the default −1 to 1 bounds and better capture extremes, and using a learned convolutional "noise modulator" to shape the noise before mixing it into the input. The model was trained on 1990–2010 data with a CRPS loss, using the Adam optimizer at a learning rate of 10⁻³.

The aggregation step is where the paper's core idea lives. Each ensemble member's predicted local anomaly is converted into a score via the normal CDF, and the members' scores are combined with a power mean controlled by p ≥ 1. When p = 1 this is the arithmetic mean; larger p shifts weight toward the highest-scoring (most extreme) members, approaching a maximum as p grows. The authors optimize p on a validation period (2010–2015) at a fixed 7-day lead time, then evaluate on 2015–2018.

Why This Matters

Impact on research. The paper shows that aggregation choice is a first-class modelling decision for extreme-event prediction, not a detail. Improving the detection of extremes may not require a bigger or more sophisticated forecasting model — a domain-specific, tunable aggregation can extract more from an existing generative model, and the exponential scaling of p_opt with q gives other researchers a simple starting heuristic. The finding that a relatively small model can outperform GraphCast on high-quantile, long-lead cases suggests that ensembling and aggregation deserve more attention than raw deterministic skill alone.

Real-world applications

  • Heat-wave early warning. More accurate detection of days expected to exceed local temperature quantiles could feed preparation and emergency planning systems.
  • Public health preparedness. The impact statement specifically links heat-wave forecasting to reduced mortality and health burdens.
  • Infrastructure and energy planning. Anticipating extreme heat supports decisions about power demand and grid stress, though the paper does not test these directly.
  • Transfer to other hazards. The authors note the method could be generalized to other weather variables to predict droughts, heavy rains and storms, since the aggregation is model-agnostic.

Industry relevance. Any organization running generative or ensemble weather models — national meteorological services, energy traders, insurers, agricultural planners, logistics operators — could apply a power mean at inference time with essentially no additional training cost, which makes it an unusually cheap potential improvement.

Future Directions

  • Apply the method to stronger base models. The authors explicitly list applying power aggregation to "more complex and effective baseline models" as future work, since their own model was deliberately light.
  • Extend beyond surface air temperature. Generalizing to other weather variables would enable prediction of droughts, heavy rainfall and storms, and would require moving from univariate to multivariate extreme definitions (for example, copula-based climatologies).
  • Use dynamic rather than static climatology. The anomalies used here are based on a static climatology that ignores climate change — which the authors flag as a major limitation, since climate change is expected to significantly alter extreme events.
  • Richer extreme-event definitions and metrics. The authors note their definition of extremes is simple and does not capture the full complexity of real events, and that AUC is not necessarily the best metric — socio-economic considerations could be integrated into evaluation.

Target Audience

This paper is most useful for machine learning researchers working on climate and weather applications, particularly those building or evaluating generative ensemble forecast models. It is also relevant to meteorologists and climate scientists interested in statistical post-processing of ensembles, and to practitioners — such as national weather services or risk analysts — looking for low-cost inference-time improvements to extreme-event detection. Readers without any background in classification metrics or neural networks will need to consult the referenced literature, but the mathematical core of the method is accessible to anyone comfortable with basic statistics.

Authors’ abstract

This paper addresses the critical challenge of improving predictions of climate extreme events, specifically heat waves, using machine learning methods. Our work is framed as a classification problem in which we try to predict whether surface air temperature will exceed its q-th local quantile within a specified timeframe. Our key finding is that aggregating ensemble predictions using a power mean significantly enhances the classifier's performance. By making a machine-learning based weather forecasting model generative and applying this non-linear aggregation method, we achieve better accuracy in predicting extreme heat events than with the typical mean prediction from the same model. Our power aggregation method shows promise and adaptability, as its optimal performance varies with the quantile threshold chosen, demonstrating increased effectiveness for higher extremes prediction.

Read the original paper