Skip to content
AI.info

Research

ForeSWE: Forecasting Snow-Water Equivalent with an Uncertainty-Aware Attention Model

Overview Research area: Machine learning for hydrology and climate — spatio-temporal probabilistic forecasting of Snow-Water Equivalent (SWE), at the intersection of attention-based deep learning and

ForeSWE: Forecasting Snow-Water Equivalent with an Uncertainty-Aware Attention Model
arXiv
2511.08856
Published
2025-11-12
Authors
Krishu K Thapa, Supriya Savalkar, Bhupinderjeet Singh, Trong Nghia Hoang, Kirti Rajagopalan, Ananth Kalyanaraman

AI summary

Overview

Research area: Machine learning for hydrology and climate — spatio-temporal probabilistic forecasting of Snow-Water Equivalent (SWE), at the intersection of attention-based deep learning and Gaussian process regression.

Technical level: Advanced. The paper assumes familiarity with self-attention, sequence-to-sequence modeling, Gaussian processes, and hydrological evaluation metrics such as Nash-Sutcliffe Efficiency.

Scope: The paper introduces ForeSWE, an uncertainty-aware attention model that forecasts SWE at 512 Snow Telemetry (SNOTEL) stations in the Western US for both daily (10-day) and weekly (4-week) horizons, and evaluates it against seven other models on accuracy and on the quality of its prediction intervals.

What This Paper Is About

Snow-Water Equivalent measures how much water is stored in a snowpack, and in snow-dominant watersheds of the Western U.S. it drives 50–80% of annual streamflow. Forecasting SWE is hard because it varies across space and time under the influence of topography, weather, and the phase of the snow season, and because existing methods neither exploit spatial and temporal correlations well nor say how confident they are in a given forecast. ForeSWE addresses both gaps by combining a spatio-temporal attention model with a Gaussian process prediction head that produces a calibrated prediction interval alongside each forecast.

Key Contributions

  1. An attention-based deep learning model parameterized with a new spatio-temporal attention module designed to capture correlation in space and time as well as interaction between attributes in the SWE forecasting context.
  2. A probabilistic augmentation using a Gaussian process as the prediction head, enabling spatio-temporal correlation learning and uncertainty quantification of SWE forecasts.
  3. A standalone sparse raw Gaussian process implementation (Raw-GP) that serves as a GP-only baseline for the SWE forecasting problem.
  4. A thorough experimental evaluation against various spatial and/or temporal machine learning approaches, comparing both accuracy and the quality of uncertainty estimates across daily and weekly horizons.

Main Findings

  • Daily forecasting accuracy: ForeSWE and Raw-GP achieve the best NSE values with comparable performance, reaching NSE above 0.75 in 99.2% and 99.6% of locations respectively.

  • Weekly forecasting accuracy: ForeSWE performs best, achieving NSE above 0.75 at over 427 locations, while Raw-GP achieves comparable NSE at only 340 locations. The authors attribute Raw-GP's decline to operating in raw feature space, where it over-smooths short-scale variation and over-emphasizes temporal over spatial correlation.

  • Temporal vs. spatial effects by horizon: For daily forecasting, temporal models perform better, indicating the prominence of temporal effects for near-term forecasting. For weekly forecasting, Sp-Att outperforms most temporal models such as LSTM and Transformer, underscoring the significance of spatial correlation in relatively long-term forecasting. Tp-Att's relative performance is comparable to or better than Raw-GP for the weekly horizon.

  • Uncertainty quality (daily, Table 1): ForeSWE has better uncertainty estimates than Raw-GP across all metrics for all test years 2015–2019. ForeSWE NLL values were 6.94 (2015), 7.48 (2016), 7.48 (2017), 6.66 (2018), 6.49 (2019), versus Raw-GP values of 11.97, 16.85, 22.70, 17.46, and 18.42. ForeSWE ECE values were 0.14, 0.15, 0.18, 0.12, 0.15 versus Raw-GP 0.362, 0.38, 0.41, 0.37, 0.40. Coverage was 80.88, 79.57, 76.12, 82.53, 79.84 for ForeSWE versus 58.71, 56.62, 53.89, 57.15, 54.54 for Raw-GP.

  • Coverage falls at longer horizons: Both methods have lower coverage in the weekly setting than in the daily setting, which the authors describe as expected since forecasting longer horizons is generally associated with higher uncertainty.

  • Seasonal error pattern: Forecasting accuracy is high in the active snow accumulation phase (February, March) across all location groups. March marks the onset of melting in low-SWE locations (group 1), which accounts for a slight increase in error. April is when most locations reach peak SWE with an accelerated melt phase, and forecast errors are higher for groups 1 and 2, while groups 3 and 4 have a median relative bias close to 0%. In May, groups 3 and 4 show increasing relative bias with horizon as most snow has melted.

  • Model design ablation set: Eight models were compared in the design ablation study — LSTM, Raw-GP, Sp-Att, Tp-Att, Transformer, TFT, NVA-Base, and ForeSWE. NVA-Base is a trimmed-down version of the proposed model without the variable aggregation and GP components.

Methodology in Plain English

The problem is framed as sequence-to-sequence prediction: for each location, the model ingests the last k days of historical observations across f attributes and outputs SWE values for the next h days. For every location on a given day, the model builds two kinds of embeddings. A location embedding combines daily spatial features and key static spatial features (such as latitude/longitude, southness, and elevation) with prompt vectors that describe context such as weather patterns, vegetation, and elevation range. Separately, each attribute's historical vector is embedded and then aggregated, using self-attention where the query comes from the location's spatial attributes plus prompts and the keys and values come from its historical daily observations. This lets the model capture how attributes interact over time for that location.

Because locations also influence each other, the per-location representations from C different temporal windows are concatenated and passed through a "modified spatial attention" step. This modification adds Haversine distance and angularity between location pairs into the attention weights, with learnable parameters governing how much each contributes — an explicit bias toward spatial proximity.

Once this spatio-temporal attention model is trained on actual SWE, its dense representations (dimension 1024) are reduced to a much smaller dimension (8) and the original prediction head is replaced by a Gaussian process. The GP is formulated as a τ-component linear co-regionalized prior with RBF kernel components, and it retains a separate temporal correlation term so that dependencies can evolve differently over short and long horizons rather than being forced into a single averaged time scale. The GP's predictive distribution yields both a mean forecast and a variance, from which an α-prediction interval is computed; the experiments use α = 0.95. Because exact GP training and inference scale cubically in the number of spatio-temporal training points, the authors adopt existing sparse approximations to bring the cost back to linear in training dataset size.

Data came from 822 SNOTEL stations across 28 water years (1991–2019); stations with more than 10% missing snow observations in any given year were filtered out, leaving 512 stations. About 180 days of daily data starting December 1 were used per year, giving 2,580,480 (= 512 × 28 × 180) location-year-day combinations. Training and testing used 25 years from 1994 onward, split into 20 training years and 5 test years, with 1991–1993 held as a buffer for historical inputs; the test water years are 2015 through 2019, chosen consecutively so those data points never appear in training, and spanning the driest (2015) to the wettest (2017) average SWE. Locations were binned into four groups by averaged peak SWE. Implementation used PyTorch v2.0.1 for the LSTM and attention models and GPyTorch v1.12 for the Gaussian process; code and data are available at https://github.com/Krishuthapa/SWE-Forecasting.

Why This Matters

SWE forecast information is critical but, as the authors state, is not yet available in an operational context, even though current-state SWE information is already widely used by local and federal water agencies including reservoir operators and irrigation districts. By producing a prediction interval with every forecast, ForeSWE targets a gap that matters directly for decision-making under uncertainty — the authors note that accounting for uncertainty is essential for optimizing planning decisions about resource deployment. The work also serves as a platform for deployment and feedback by the water management community.

Real-world applications:

  • Multi-day flood risk management: Multi-day forecasts help anticipate rapid snowmelt and potential flooding, enabling real-time interventions such as reservoir drawdowns.
  • Subseasonal water allocation: Multi-week forecasts of SWE and peak SWE inform subseasonal planning and allocation decisions across agricultural, ecological, and hydropower sectors.
  • Streamflow forecasting inputs: SWE serves as a key input to streamflow forecasts and improves sub-seasonal climate outlooks by capturing land–atmosphere feedbacks.
  • Operational planning by water agencies: Reservoir operators, irrigation districts, and the NRCS dashboard ecosystem could integrate probabilistic SWE forecasts into planning workflows.

Industry relevance: The paper explicitly names hydropower generation, agricultural planning, and recreational activities such as skiing season operations as beneficiaries of weekly forecasts, and flood risk management as the beneficiary of daily forecasts. The authors frame improvements in SWE forecasting as supporting more equitable and sustainable water use, reducing societal impacts of extreme hydrologic events, and strengthening the resilience of communities and ecosystems dependent on predictable water availability.

Future Directions

  • Operational deployment and community feedback: The authors state their findings have provided a platform for deployment and feedback by the water management community, implying the next step is real-world operational use and iteration based on that feedback.
  • Extending spatial and temporal scales: The paper evaluates a 10-day daily horizon and a 4-week weekly horizon; whether ForeSWE's advantage over Raw-GP holds at other scales is not addressed.
  • Combining models rather than selecting one: The results show ForeSWE's relative bias against the temporal LSTM model for Group 4 in May, and a relative bias plot with an ensembling of ForeSWE and LSTM, suggesting model combination as a direction worth pursuing.
  • Improving coverage at long horizons: Both ForeSWE and Raw-GP lose coverage in the weekly setting relative to daily, leaving open the question of how to tighten prediction intervals for longer horizons.

Target Audience

Hydrologists and water resource managers interested in probabilistic snowpack forecasting; machine learning researchers working on spatio-temporal and uncertainty-aware models; climate and earth-system scientists applying deep learning to physical processes; and anyone building decision-support tools where forecast confidence intervals matter as much as the forecast value itself. The paper is written for readers comfortable with attention mechanisms and Gaussian processes, though the introductory problem framing is accessible to a broader water-management audience.

Authors’ abstract

Various complex water management decisions are made in snow-dominant watersheds with the knowledge of Snow-Water Equivalent (SWE) -- a key measure widely used to estimate the water content of a snowpack. However, forecasting SWE is challenging because SWE is influenced by various factors including topography and an array of environmental conditions, and has therefore been observed to be spatio-temporally variable. Classical approaches to SWE forecasting have not adequately utilized these spatial/temporal correlations, nor do they provide uncertainty estimates -- which can be of significant value to the decision maker. In this paper, we present ForeSWE, a new probabilistic spatio-temporal forecasting model that integrates deep learning and classical probabilistic techniques. The resulting model features a combination of an attention mechanism to integrate spatiotemporal features and interactions, alongside a Gaussian process module that provides principled quantification of prediction uncertainty. We evaluate the model on data from 512 Snow Telemetry (SNOTEL) stations in the Western US. The results show significant improvements in both forecasting accuracy and prediction interval compared to existing approaches. The results also serve to highlight the efficacy in uncertainty estimates between different approaches. Collectively, these findings have provided a platform for deployment and feedback by the water management community.

Read the original paper