Research
Day-Ahead Electricity Price Forecasting for Volatile Markets Using Foundation Models with Regularization Strategy
Overview Research area: Day-ahead electricity price forecasting (EPF) using time series foundation models (TSFMs), applied to half-hourly wholesale market data from Singapore. Technical level: Interme

- arXiv
- 2602.05430
- Published
- 2026-02-05
- Authors
- Kritchanat Ponyuenyong, Pengyu Tu, Jia Wei Tan, Wei Soon Cheong, Jamie Ng Suat Ling, Lianlian Jiang
AI summary
Overview
Research area: Day-ahead electricity price forecasting (EPF) using time series foundation models (TSFMs), applied to half-hourly wholesale market data from Singapore.
Technical level: Intermediate. The paper assumes familiarity with forecasting metrics (MAE, RMSE, MAPE), deep learning baselines, and the idea of pre-trained foundation models, but its central mechanisms — spike removal, log transformation, lag features — are explained in accessible terms.
Scope: The paper evaluates 37 model configurations (eight TSFM families plus statistical and deep learning baselines, each under several training strategies) for 24-hour-ahead price prediction, and proposes a three-stage spike detection and regularization pipeline called STL-KF.
What This Paper Is About
Electricity prices in markets like Singapore are volatile, seasonal, and punctuated by extreme spikes, which makes accurate day-ahead forecasting hard for both classical statistical models and deep learning models. The authors ask whether time series foundation models — normally used for general forecasting tasks such as traffic or weather — can beat those traditional approaches on this specific, high-frequency, spike-prone problem. They also propose a preprocessing strategy that cleans extreme price spikes before training so models learn the underlying trend rather than overfitting to rare anomalies.
Key Contributions
-
A hybrid spike detection and regularization strategy (STL-KF). The method combines Season-Trend decomposition using LOESS (STL) at weekly and monthly levels, a Kalman Filter enhanced with Huber loss functions for outlier robustness, and an adaptive thresholding mechanism whose acceptance bounds widen or narrow based on the Kalman Filter's estimated uncertainty. The scaling factor is set to λ = 3.
-
A broad, standardized benchmark of TSFMs against classical baselines for EPF. Eight foundation model families (MOIRAI, MOMENT, TTMs, TimesFM, Time-MoE, Timer-XL, Lag-Llama, Chronos) are compared against ARIMA, LSTM, CNN-LSTM, PatchTST, and Amplifier, yielding 37 models in total across six evaluation strategies.
-
A systematic study of evaluation and preprocessing choices. The paper quantifies the effect of zero-shot versus fine-tuned settings, univariate versus multivariate fine-tuning, log transformation of prices, and the addition of lag and exogenous features.
-
A practical comparison of accuracy against computational cost. The discussion contrasts the quadratic complexity of Transformer attention with the MLP-Mixer design of TTMs as a latency-friendly alternative for real-time use.
Main Findings
-
Foundation models win overall: Across all forecasting settings, TSFMs achieved lower average errors (29.76 MAE and 13.99% MAPE) than statistical models (34.29 MAE, 18.35% MAPE) and deep learning models (32.13 MAE, 16.47% MAPE). The abstract reports up to 37.4% improvement in MAPE.
-
Multivariate fine-tuned TTMs is the best single model: TTMs (mft) recorded the lowest MAE (27.86) and MAPE (11.94%) on the test set, with RMSE of 192.19. The three top-performing configurations were all TSFMs under univariate or multivariate fine-tuning.
-
Univariate fine-tuning is often not worth it: TTMs improved only from 12.62% MAPE (zero-shot) to 12.41% (univariate fine-tuned), a marginal 0.21% gain for extra training. Lag-Llama worsened from 12.63% to 13.54%, Timer-XL from 13.31% to 16.27%, and TimesFM from 13.32% to 15.03%.
-
Spike regularization substantially improves forecasts: Using STL-KF before modeling produced up to a 61.67% improvement in MAPE (based on LSTM) and an average of 28.69% across all models, with metrics computed against the original raw price values.
-
Log transformation and features help: Log transformation reduced MAPE by an average of 9.64%, and adding lag and all other features (including weather) provided an additional 3.80% reduction.
-
MOMENT zero-shot was an outlier: MOMENT (zs) produced 61.16 MAE, 38.82% MAPE, and 236.47 RMSE, and was excluded from the TSFM average as an anomaly.
-
Recurrent models struggled: LSTM performed poorly relative to transformer-based models, attributed to difficulty capturing rapid fluctuations in high-frequency volatile prices.
-
Larger models showed lower spike error: Time-MoE and Timer-XL achieved RMSE values of 190.95 and 189.86 respectively, described as indicating lower spike errors, though their MAPE was higher than the best TTMs configuration.
-
Efficiency trade-off: Self-attention layers have quadratic complexity O(L²), which the authors say causes substantial latency especially for autoregressive decoder-only models, whereas TTMs use a lightweight MLP-Mixer backbone with adaptive patching at 0.8M parameters.
Methodology in Plain English
The authors collected half-hourly electricity price and demand data for Singapore from 1 January 2021 to 31 December 2024 from the Energy Market Company and Energy Market Authority, plus temperature, humidity, and heat index from the OpenWeather API, and public holiday information. The dataset contains 69,600 samples before splitting, divided into 70% training, 20% validation, and 10% testing.
Before modeling, they clean the data in three steps. First, STL removes weekly and monthly seasonal patterns, isolating non-seasonal deviations. Second, a Kalman Filter with a Huber loss estimates the underlying price state while limiting the influence of outliers. Third, an adaptive threshold checks whether each observed price falls outside bounds computed from the predicted state plus the seasonal component, plus or minus λ times the square root of the summed system error covariance and measurement noise covariance, with λ = 3. Prices outside the bounds are flagged as spikes and replaced with the filtered state estimate. The approach is designed to widen the acceptance range automatically when the market is unstable, reducing false positives.
They then build 14 features: lag features at t-1, t-2, t-4, t-24, t-48, t-96, t-192, and t-336; a natural log transformation of prices; and six exogenous variables chosen by correlation analysis — demand, hour of the day, temperature, humidity, heat index, and weekend indicators. Everything is min-max normalized.
For forecasting, the time series is cut into overlapping windows with a stride of one. Each window looks back 512 time steps (10 days and 16 hours) and forecasts 48 steps ahead (24 hours). Models are run under six strategies: zero-shot, log-price zero-shot, univariate fine-tuned, train-from-scratch, log-price train-from-scratch, and multivariate fine-tuned. Performance is measured with MAE, RMSE, and MAPE, with MAPE prioritized because relative error is more informative across varying price levels.
Why This Matters
Impact on research. This is among the first studies to test time series foundation models specifically on day-ahead electricity price forecasting in a volatile, high-granularity market. It shows that pre-trained temporal representations transfer well to this domain and that a relatively small foundation model (TTMs at 0.8M parameters) can beat much larger ones (MOMENT at 346M, TimesFM at 500M, Time-MoE at 453M) as well as classical baselines. It also provides a caution: naive univariate fine-tuning can degrade foundation model performance by stripping away the multivariate structure they were pretrained to use.
Real-world applications:
- Grid operators scheduling generation and managing reserve margins with more reliable day-ahead price signals.
- Energy traders setting bids in half-hourly real-time auctions, where relative accuracy at different price levels matters more than absolute error.
- Policymakers and regulators assessing market volatility and the effect of fuel price shocks or geopolitical events on imported natural gas-dependent markets.
- Forecasting practitioners in other volatile markets who can adopt the STL-KF preprocessing step independently of which model they use, given the 28.69% average MAPE improvement reported.
Industry relevance. The paper explicitly frames its contribution as practical guidance for identifying suitable models for a given use case. The computational efficiency discussion supports deployment decisions: Transformer-based models offer strong zero-shot generalization but incur attention overhead, while TTMs are positioned as a balanced choice for real-time applications because they achieve a global receptive field without attention heads.
Future Directions
-
Extending the methodology to other electricity markets. The authors state this as planned future work, which would test whether the STL-KF plus TSFM combination generalizes beyond Singapore's imported-gas-dependent market.
-
Probabilistic forecasting and uncertainty analysis. The paper proposes covering the likelihood of extreme spike events and their quantile ranges, moving beyond point forecasts.
-
Understanding why fine-tuning sometimes hurts. The observed degradation for Timer-XL, TimesFM, and Lag-Llama under univariate fine-tuning raises the open question of how to fine-tune foundation models without discarding the multivariate structure they were pretrained on.
-
Quantifying the efficiency trade-off. The computational efficiency section is qualitative — it discusses O(L²) attention complexity and low-latency matrix operations but does not report measured latency or throughput figures. Whether TTMs' accuracy advantage holds under hard real-time constraints in production remains untested here.
Target Audience
This paper suits energy market analysts and forecasting engineers who want a practical model-selection guide for volatile price data; machine learning researchers interested in whether foundation models transfer to domain-specific, non-stationary time series; and grid operators or trading desks evaluating whether to adopt spike-regularization preprocessing and foundation models over ARIMA, LSTM, or CNN-LSTM pipelines. Readers without a forecasting background can still follow the preprocessing logic and the headline comparisons, while the fine-tuning and feature ablations are aimed at more technical readers.
Authors’ abstract
Electricity price forecasting (EPF) is essential for energy markets stakeholders (e.g. grid operators, energy traders, policymakers) but remains challenging due to the inherent volatility and nonlinearity of price signals. Traditional statistical and deep learning (DL) models often struggle to capture complex temporal dependencies and integrate heterogeneous data effectively. While time series foundation models (TSFMs) have shown strong performance in general time series forecasting tasks, such as traffic forecasting and weather forecasting. However, their effectiveness in day-ahead EPF, particularly in volatile markets, remains underexplored. This paper presents a spike regularization strategy and evaluates a wide range of TSFMs, including Tiny Time Mixers (TTMs), MOIRAI, MOMENT, and TimesFM, against traditional statistical and DL models such as Autoregressive Integrated Moving Average (ARIMA), Long-short Term Memory (LSTM), and Convolutional Neural Network - LSTM (CNN-LSTM) using half-hourly wholesale market data with volatile trends in Singapore. Exogenous factors (e.g. weather and calendar variables) are also incorporated into models where applicable. Results demonstrate that TSFMs consistently outperform traditional approaches, achieving up to 37.4% improvement in MAPE across various evaluation settings. The findings offer practical guidance for improving forecast accuracy and decision-making in volatile electricity markets.