Research
Assessing Electricity Demand Forecasting with Exogenous Data in Time Series Foundation Models
Overview Research area: Time-series forecasting / machine learning for energy systems, specifically electricity demand forecasting using time-series foundation models with exogenous (external) variabl

- arXiv
- 2602.05390
- Published
- 2026-02-05
- Authors
- Wei Soon Cheong, Lian Lian Jiang, Jamie Ng Suat Ling
AI summary
Overview
Research area: Time-series forecasting / machine learning for energy systems, specifically electricity demand forecasting using time-series foundation models with exogenous (external) variables.
Technical level: Intermediate. The paper assumes familiarity with forecasting metrics like MAPE, transformer architectures, and the distinction between zero-shot and fine-tuned evaluation, but its central argument is accessible to anyone working in energy analytics.
Scope: A systematic empirical comparison of five time-series foundation models (MOIRAI, MOMENT, TinyTimeMixers, ChronosX, Chronos-2) against a RevIN-LSTM baseline across Singaporean and Australian electricity markets at hourly and daily granularities, under three feature configurations (all features, selected features, target-only).
What This Paper Is About
Time-series foundation models are pre-trained on massive, heterogeneous collections of time series and are marketed as general-purpose forecasters, but it is unclear whether they can actually exploit the weather and calendar variables that electricity demand forecasting depends on. This paper tests that question directly by running the same forecasting task across two very different electricity markets and measuring whether adding external features helps or hurts each model. The goal is to determine whether foundation models are a genuine advance for energy forecasting or whether simpler models remain sufficient.
Key Contributions
-
A controlled multi-model, multi-region benchmark. The paper evaluates MOIRAI, MOMENT, TinyTimeMixers (TTM), ChronosX, and Chronos-2 alongside a RevIN-LSTM baseline on four datasets (Singapore hourly and daily, Australia hourly and daily), using identical sliding-window sampling with a step size of 1 and a 512-step context length.
-
Isolating the effect of exogenous features. Each model is run under three configurations — all features (30 variables), selected features (7-10 variables chosen via Spearman correlation and confirmed with Granger causality tests), and target-only — so the marginal value of external data can be measured in percentage improvement terms.
-
Connecting architecture to feature utilization. The paper attributes differences in exogenous-feature effectiveness to specific design choices: TTM's channel-mixer block enabled during fine-tuning, Chronos-2's Group Attention across channels, MOIRAI's Any-variate Attention, MOMENT's reliance on linear probing, and ChronosX's adapter modules.
-
A geographic-context argument. It shows that the benefit of foundation models is conditional on climate predictability, with stable Singapore favoring the simple LSTM and variable-climate Australia favoring the foundation models.
Main Findings
-
Chronos-2 is the strongest foundation model, zero-shot. Chronos-2 consistently achieved the lowest MAPE among foundation models across both regions and most horizons, in both hourly and daily settings, without any fine-tuning (the authors note it was not fine-tuned due to time constraints and fine-tuning being an experimental feature).
-
The RevIN-LSTM baseline often wins in Singapore. Without exogenous features, RevIN-LSTM frequently outperformed all foundation models in Singapore for short-term horizons up to 48 hours, and remained competitive at longer horizons. In the daily setting it was weaker, staying competitive only at shorter horizons and tapering off after horizon 30.
-
Absolute errors are far lower in Singapore. Hourly MAPE ranged from 0.44% to 3.14% in Singapore versus 1.68% to 9.22% in Australia's ACT region. Daily MAPE ranged from 1.35% to 5.13% in Singapore versus 3.07% to 13.11% in Australia.
-
Exogenous features are underused in Singapore, better used in Australia. In the hourly Singapore setting, most models preferred univariate prediction. In Australia, with the exception of MOMENT and MOIRAI-ZS, all models improved consistently with exogenous information, with the largest gains from RevIN-LSTM.
-
TTM shows the most volatile response to features. In hourly Singapore, TTM improved by up to 21.1% but also deteriorated by up to 20.0% depending on the horizon.
-
Daily granularity favors exogenous features more than hourly. In daily Singapore, TTM gained between 1.62% and 14.15% and Chronos-2 between 1.41% and 26.41% from feature inclusion. MOIRAI-ZS recorded increases above 60% at longer horizons in Singapore, MOMENT up to 10%, and RevIN-LSTM a single 36.02% increase with all features at horizon 365.
-
ChronosX behaves unpredictably and fails at the shortest horizon. It showed drastic performance reductions when using exogenous features at horizon 1 across all experiment configurations, with some hourly figures reported as "< -100" percent. The paper notes ChronosX with no features (zero-shot Chronos) only supports forecasting up to a horizon of 64.
-
Feature-set preference is not stable. Chronos-2 preferred all features in Australia-hourly and Singapore-daily, but preferred selected features in Australia-daily. TTM and Chronos-2 slightly preferred the complete feature set in Australia-hourly, suggesting those architectures handle higher-dimensional inputs when predictive signal is present.
-
MOIRAI's theoretical advantage did not translate into results. Despite its Any-variate Attention layer, fine-tuned MOIRAI remained the poorest performer overall, and its zero-shot variant deteriorated across most settings when exogenous features were added.
-
MOMENT's inconsistency was expected. The authors attribute it to MOMENT not being designed for cross-channel modeling, relying instead on linear probing of output embeddings.
Methodology in Plain English
The researchers collected electricity demand data from Singapore's Energy Market Authority and Australia's AEMO, plus weather data from World Weather Online and air quality data from government portals. They built four datasets: Singapore and Australia, each at hourly and daily resolution. Missing values were filled by linear interpolation for continuous variables and forward filling for categorical ones, and the data was split 60/20/20 into train, validation, and test sets. Training set sizes were 36,821 (Singapore hourly), 1,535 (Singapore daily), 47,333 (Australia hourly), and 1,972 (Australia daily); validation/test sizes were 12,273, 511, 15,777, and 657 respectively.
Each input window contains 512 time steps (the context length many foundation models use), with a step size of 1. Forecast horizons were [1, 12, 24, 48, 72, 168, 336] for hourly data (up to 14 days) and [1, 7, 14, 30, 60, 180, 365] for daily data (up to one year). Accuracy was measured with Mean Absolute Percentage Error (MAPE).
To test the value of external data, each model was run three ways: with all features (30 variables), with a curated subset (7-10 variables picked by Spearman correlation with the target and validated with Granger causality tests at lags below the 512 context length), and with the target only. All models were trained with MSE loss, the Adam optimizer at a learning rate of 0.001, batch size 16, up to 100 epochs with early stopping patience of 10, on a single A40 GPU, using default parameters from each model's GitHub repository where possible. TTM used different published revisions depending on horizon, and Chronos-2 was evaluated zero-shot only.
Why This Matters
The paper pushes back on the assumption that larger pre-trained models are universally better. It shows that the value of a foundation model depends jointly on architecture, forecast horizon, and the geographic predictability of the market — a conclusion that matters for how utilities and grid operators choose their forecasting stack.
Impact on research: It identifies a specific gap — prior foundation-model benchmarks in electricity forecasting used inherently univariate models — and frames the case for energy-domain foundation models pretrained on diverse electricity and energy contexts with explicit architectural attention to weather-demand relationships. It also flags an architectural tension: the models that handle exogenous variables in this study are primarily encoder-only, while the field appears to be converging on decoder-only designs such as TimesFM-2 and MOIRAI-2.0.
Real-world applications:
- Short-term operational forecasting for grids in stable climates, where a small RevIN-LSTM may match or beat a large foundation model at far lower cost.
- Demand response program scheduling, where accurate short-horizon load prediction determines when to call on flexibility.
- Infrastructure planning and maintenance scheduling in regions with highly variable weather, where exogenous features demonstrably add predictive value.
- Feature-selection pipelines for forecasting systems, since the paper shows that curtailing a 30-variable input to 7-10 variables can preserve or improve accuracy for some models.
Industry relevance: The authors argue that real-time grid operations may find foundation-model overhead prohibitive, because attention-based approaches scale quadratically with the number of input channels and demand substantially more memory and compute than the RevIN-LSTM baseline during both fine-tuning and inference. TTM's lightweight MLP-based architecture is noted as a more favorable trade-off. The suggested mitigations are rigorous preliminary feature selection, sparse attention mechanisms, and state-space models.
Future Directions
-
Build energy-domain-specific foundation models. The authors call for pretraining on diverse electricity and energy forecasting contexts with explicit architectural attention to weather-demand relationships, rather than relying on general-purpose models trained on heterogeneous multisubject data.
-
Determine which mechanisms best capture exogenous dependencies. The paper asks explicitly which architectural approaches most reliably model covariate relationships, given that adaptive designs such as TTM's channel-mixing and Chronos-2's grouped attention worked while others did not.
-
Investigate decoder-only models with exogenous inputs. Since encoder-only architectures dominate among the models that succeeded here, the authors identify decoder-only autoregressive forecasting with covariates as a promising but potentially complex direction, given the difficulty of high-dimensional multivariate forecasting in an autoregressive manner.
-
Test pretraining value in low-data regimes. The authors acknowledge that their experiments supply substantial training data, which may reduce the relative advantage of pretraining, and suggest this limits the conclusions that can be drawn about foundation models in data-scarce settings.
Target Audience
Energy analysts, grid operators, and forecasting engineers deciding between foundation models and conventional deep learning for load prediction; machine learning researchers studying multivariate and covariate-aware time-series architectures; and benchmark designers who need to understand why foundation-model evaluations that ignore geographic and climatic context may produce misleading conclusions.
Authors’ abstract
Time-series foundation models have emerged as a new paradigm for forecasting, yet their ability to effectively leverage exogenous features -- critical for electricity demand forecasting -- remains unclear. This paper empirically evaluates foundation models capable of modeling cross-channel correlations against a baseline LSTM with reversible instance normalization across Singaporean and Australian electricity markets at hourly and daily granularities. We systematically assess MOIRAI, MOMENT, TinyTimeMixers, ChronosX, and Chronos-2 under three feature configurations: all features, selected features, and target-only. Our findings reveal highly variable effectiveness: while Chronos-2 achieves the best performance among foundation models (in zero-shot settings), the simple baseline frequently outperforms all foundation models in Singapore's stable climate, particularly for short-term horizons. Model architecture proves critical, with synergistic architectural implementations (TTM's channel-mixing, Chronos-2's grouped attention) consistently leveraging exogenous features, while other approaches show inconsistent benefits. Geographic context emerges as equally important, with foundation models demonstrating advantages primarily in variable climates. These results challenge assumptions about universal foundation model superiority and highlight the need for domain-specific models, specifically in the energy domain.