Skip to content
AI.info

Research

Rethinking deep learning: linear regression remains a key benchmark in predicting terrestrial water storage

Overview Research area: Machine learning applied to hydrology and land surface science, specifically the prediction of terrestrial water storage (TWS). Technical level: Intermediate. The abstract assu

Rethinking deep learning: linear regression remains a key benchmark in predicting terrestrial water storage
arXiv
2510.10799
Published
2025-10-12
Authors
Wanshu Nie, Sujay V. Kumar, Junyu Chen, Long Zhao, Olya Skulovich, Jinwoong Yoo, Justin Pflug, Shahryar Khalique Ahmad, Goutam Konapala

AI summary

Overview

Research area: Machine learning applied to hydrology and land surface science, specifically the prediction of terrestrial water storage (TWS).

Technical level: Intermediate. The abstract assumes familiarity with LSTM networks, Transformers, land surface models, and remote sensing data assimilation, but the central argument is conceptual rather than mathematically dense.

Scope: A benchmark comparison arguing that simple linear regression can outperform deep learning models for predicting terrestrial water storage, and that traditional statistical models should be standard baselines in deep learning development.

What This Paper Is About

Deep learning models such as LSTMs and Transformers have become popular in hydrology and often beat traditional physical models. This paper asks whether that advantage holds for predicting land surface states like terrestrial water storage, which are shaped by both natural variability and human-driven changes. The authors test that assumption by comparing linear regression against two deep learning architectures on a globally representative dataset.

Key Contributions

  1. A direct benchmark comparison of linear regression against LSTM and Temporal Fusion Transformer models for terrestrial water storage prediction, using the globally representative HydroGlobe dataset.
  2. Evidence that linear regression outperforms the more complex deep learning models on this task, challenging the assumption that architectural complexity translates into better predictive skill for TWS.
  3. A methodological argument that traditional statistical models should be included as baselines whenever deep learning models are developed and evaluated in this domain.
  4. A call for globally representative benchmark datasets that capture both natural variability and human interventions, rather than natural dynamics alone.

Main Findings

  • Linear regression as a robust benchmark: The paper reports that linear regression outperforms both LSTM and Temporal Fusion Transformer models for TWS prediction.
  • Complexity is not automatically superior: Despite their strong track record in other hydrological tasks, the deep learning models did not show an advantage for this particular land surface state. The abstract does not report the magnitude of the gap or any performance metrics.
  • TWS is a difficult target for deep learning: The authors attribute the unclear benefit of deep learning to the fact that terrestrial water storage is governed by many interacting factors, including natural variability and human-driven modifications.
  • Dataset design matters: Two versions of HydroGlobe were used: a baseline derived only from a land surface model simulation, and an advanced version that incorporates multi-source remote sensing data assimilation. The abstract does not state how the two versions differed in results.
  • Benchmarking gap: The paper argues that the field lacks globally representative benchmark datasets that reflect the combined influence of natural processes and human interventions, and that this gap limits fair evaluation of deep learning models.

Methodology in Plain English

The researchers used HydroGlobe, an open-access dataset designed to be globally representative. It comes in two forms: a baseline built purely from a land surface model simulation, and an advanced version that blends in observations from multiple remote sensing sources through data assimilation.

On this dataset, they set up a prediction problem for terrestrial water storage and compared three approaches: a simple linear regression, an LSTM (a type of recurrent neural network suited to sequences), and a Temporal Fusion Transformer (a more elaborate deep learning architecture designed for multi-horizon time series forecasting). Performance was judged on how well each model predicted TWS. The abstract does not describe the input variables, the prediction horizon, the spatial resolution, or the specific evaluation metrics used.

Why This Matters

Impact on research: The paper pushes back on the default assumption that deep learning is the right tool for every geoscientific prediction problem. It argues for stronger benchmarking discipline: before claiming that a complex model advances the field, researchers should demonstrate that it beats simple statistical alternatives. It also highlights a data problem, namely that benchmark datasets often fail to represent the human-driven side of water storage change, which can make model comparisons misleading.

Real-world applications:

  • Water resource planning: Reliable TWS estimates support decisions about water availability, allocation, and long-term supply in regions where ground observations are sparse.
  • Drought and flood monitoring: TWS is a core indicator of accumulated water deficit or surplus, relevant to early warning systems.
  • Groundwater depletion tracking: Because TWS integrates surface, soil, and groundwater storage, better prediction supports monitoring of stressed aquifers.
  • Satellite product evaluation: The HydroGlobe framing of simulation versus data assimilation is relevant to teams building and validating remote sensing based water storage products.

Industry relevance: Water utilities, agricultural planning services, reinsurance and risk modeling firms, and geospatial analytics companies all depend on water storage estimates. If a simple model performs as well as or better than a deep learning model, that lowers computational cost, reduces the need for large training datasets, and makes operational forecasting systems easier to maintain.

Future Directions

  • Determine whether linear regression's advantage holds for other land surface states beyond terrestrial water storage, or whether it is specific to this variable.
  • Build benchmark datasets that more explicitly encode human interventions, such as irrigation, reservoir operation, and groundwater pumping, alongside natural variability.
  • Clarify the conditions under which deep learning architectures do provide an advantage in hydrological prediction, given that they succeed in other tasks but not this one.
  • Establish community conventions for which statistical baselines should be mandatory when evaluating new deep learning models in the geosciences.

Target Audience

Hydrologists and land surface scientists working on water storage estimation; machine learning researchers who apply deep learning to Earth system problems; developers of remote sensing and data assimilation products; and reviewers or practitioners who need guidance on choosing between simple statistical models and complex neural architectures for environmental prediction. The paper is also relevant to research groups building benchmark datasets, since its central recommendation concerns how those datasets should be designed and how models should be compared against them.

Authors’ abstract

Recent advances in machine learning such as Long Short-Term Memory (LSTM) models and Transformers have been widely adopted in hydrological applications, demonstrating impressive performance amongst deep learning models and outperforming physical models in various tasks. However, their superiority in predicting land surface states such as terrestrial water storage (TWS) that are dominated by many factors such as natural variability and human driven modifications remains unclear. Here, using the open-access, globally representative HydroGlobe dataset - comprising a baseline version derived solely from a land surface model simulation and an advanced version incorporating multi-source remote sensing data assimilation - we show that linear regression is a robust benchmark, outperforming the more complex LSTM and Temporal Fusion Transformer for TWS prediction. Our findings highlight the importance of including traditional statistical models as benchmarks when developing and evaluating deep learning models. Additionally, we emphasize the critical need to establish globally representative benchmark datasets that capture the combined impact of natural variability and human interventions.

Read the original paper