Skip to content
AI.info

Research

Shapelets-Enriched Selective Forecasting using Time Series Foundation Models

Overview Research area: Time series forecasting with pre-trained time series foundation models (TSFMs), combined with selective prediction (also called reject-option learning) and shapelet-based time

Shapelets-Enriched Selective Forecasting using Time Series Foundation Models
arXiv
2601.11821
Published
2026-01-16
Authors
Shivani Tomar, Seshu Tirupathi, Elizabeth Daly, Ivana Dusparic

AI summary

Overview

Research area: Time series forecasting with pre-trained time series foundation models (TSFMs), combined with selective prediction (also called reject-option learning) and shapelet-based time series pattern mining.

Technical level: Intermediate. The paper assumes familiarity with forecasting metrics such as MSE, fine-tuning, and dictionary learning, but the core idea (drop predictions you do not trust) is easy to grasp.

Scope: The paper proposes a shapelet-guided selective forecasting framework that lets a user discard the TSFM predictions most likely to be wrong, and evaluates it with the Tiny Time Mixer (TTM) model on six benchmark datasets.

Paper details: arXiv:2601.11821v1 [cs.LG], 16 Jan 2026, by Shivani Tomar, Seshu Tirupathi, Elizabeth Daly and Ivana Dusparic (Trinity College Dublin and IBM Research, Dublin).

What This Paper Is About

Time series foundation models such as Chronos, Tiny Time Mixers (TTM), Moirai, TimesFM, LLMTime and GPT4TS report strong average zero-shot forecasting accuracy, but they often produce unreliable predictions on specific segments of data characterized by significant or abrupt changes. The paper's goal is to build a selection mechanism that identifies those unreliable segments using time series shapelets, so that a user can discard those predictions and obtain a lower overall error, along with a realistic picture of what the model can and cannot forecast.

Key Contributions

  1. A selective forecasting framework for TSFMs guided by time series shapelets, which identifies occurrences of specific patterns in the data that coincide with potentially high prediction errors on target-domain datasets.
  2. A sample selection module that uses distance-based similarity between learned shapelets and test samples to decide which predictions to discard at inference time.
  3. Extensive experiments using TTM on six benchmark datasets, showing reduced mean squared error (MSE) averaging 22.17% for zero-shot and 22.62% for full-shot fine-tuned models.
  4. Comparison against random selection baselines at the same 20% drop percentage, with gains of up to 21.41% (zero-shot) and 21.43% (full-shot fine-tuned) on one of the datasets, plus ablations over the error threshold and the drop percentage.

Main Findings

  • Shapelet-guided selection reduces MSE on most datasets: Across the six benchmark datasets, selective forecasting with shapelets achieved lower MSE than the unmodified zero-shot and full-shot fine-tuned baselines. Table 2 reports, for example, Exchange Rate falling from 0.0790 (zero-shot) and 0.0776 (full-shot fine-tuned) to 0.0492 and 0.0484 respectively, and Traffic falling from 0.1766 and 0.1173 to 0.1407 and 0.0933.
  • Average error reduction figures: The authors report an average reduction in overall error of 22.17% for the zero-shot model and 22.62% for the full-shot fine-tuned model.
  • Comparison against random selection: Random selection at 20% drop percentage produced MSEs such as 0.0418 (zero-shot) and 0.0408 (full-shot fine-tuned) on ETTh1, 0.0615 and 0.0518 on ETTm2, and 0.0626 and 0.0616 on Exchange Rate. The paper states the random-selection baseline performs slightly better on two datasets but leads to significantly higher error on the remaining datasets, and that the proposed approach outperforms random selection by up to 21.41% (zero-shot) and 21.43% (full-shot fine-tuned) on one of the datasets.
  • Four out of six datasets improved: The paper reports that for four of the six datasets, selective forecasting yielded lowered error while discarding unreliable predictions.
  • Discarded versus retained samples differ measurably by distance: On Exchange Rate, a discarded test sample matched its closest shapelet at a z-normalized Euclidean distance of 4.46, while a retained sample's closest shapelet distance was 7.44. Similar patterns are reported for ETTh1 (5.47 vs 7.86), ETTh2 (2.78 vs 6.16), Traffic (1.43 vs 2.18), ETTm1 (2.23 vs 7.67) and ETTm2 (1.27 vs 7.38).
  • Error threshold ablation: Varying the multiplier δ in the threshold (mean error + δ × std) on Exchange Rate and ETTh2 gave MSEs of 0.0526 and 0.1013 at δ = 1.0; 0.0516 and 0.1040 at δ = 1.5; 0.0484 and 0.0992 at δ = 2.0; 0.0516 and 0.1028 at δ = 2.5; and 0.0507 and 0.1017 at δ = 3.0. Higher thresholds select fewer validation samples, risking missed patterns.
  • Drop percentage ablation: Increasing the drop percentage over {10%, 20%, 30%, 40%, 50%} with the full-shot fine-tuned model decreased MSE monotonically for every dataset. For example, ETTh1 goes 0.0482 → 0.0442 → 0.0400 → 0.0344 → 0.0312, and Traffic goes 0.1031 → 0.0933 → 0.0787 → 0.0669 → 0.0567.
  • Limitation observed under distribution shift: For ETTh1, high-error test samples at the start and end of January 2018 were not discarded, which the authors attribute to distribution shift between the validation and test sets, meaning validation-learned shapelets cannot represent later pattern changes.
  • Illustrative threshold values: On the Exchange Rate dataset the mean validation error was 0.4506 and δ was set to 2, giving an error threshold of 1.6764.

Methodology in Plain English

The framework has four stages:

  1. Measure baseline unreliability. A pre-trained TTM is run zero-shot on the validation split of the target dataset. The mean MSE is computed, and an error threshold is set as the mean error plus δ times the standard deviation (δ = 2 in the main experiments). Validation windows whose error exceeds this threshold are collected into a smaller set.
  2. Learn shapelets from those high-error windows only. Using shift-invariant dictionary learning (following Zheng et al. 2016), the method learns a dictionary of basis elements — shapelets — plus sparse coding coefficients. The number of shapelets (K), their length (q) and the sparsity regularization parameter λ are chosen by grid search over 3-fold cross-validation on the validation split. The dictionary is then pruned to the TopK shapelets ranked by the sparse coding matrix, to remove redundant or self-similar shapelets.
  3. Score test samples by similarity. Each test sample is compared with each shapelet using z-normalized Euclidean distance, producing a distance matrix. The authors note that they use shapelets not to identify classes, as in classification, but purely to uncover representative patterns.
  4. Selectively discard. The distance matrix is sorted ascending, and the samples closest to the shapelets are dropped, up to a user-defined drop percentage (dp), with their prediction error marked as 0. The remaining test samples are forecast using either the zero-shot or the full-shot fine-tuned model. The drop percentage acts as a coverage control.

Experimental setup details: TTM-Base is used with context length 512 and forecast length 96. Full-shot fine-tuning updates the model head while keeping backbone parameters frozen, using hyperparameters (head dropout, batch size, learning rate) taken from the validation performance reported in the original TTM paper. All experiments are run with three random seeds and averaged, with dp = 20% and δ = 2. Datasets are split into train/validation/test in the same ratio as the original TTM paper. The six datasets are ETTh1 (1 hour, length 17420), ETTh2 (1 hour, 17420), ETTm1 (15 min, 69680), ETTm2 (15 min, 69680), Exchange Rate (1 day, 7588), and Traffic (1 hour, 17544). Forecasting is univariate (one channel).

Why This Matters

The paper addresses a practical gap: foundation models look good on average but can silently fail on exactly the turbulent segments practitioners care about. Rather than claiming better raw accuracy everywhere, it offers users an explicit control knob — how much of the prediction set to abstain from — plus interpretable visual evidence of which patterns trigger abstention. This matters for research because it connects the well-established selective prediction literature with the newer TSFM literature, where jointly training a reject function (as in SelectiveNet) is not feasible for already pre-trained models.

Real-world applications:

  • Energy and electricity forecasting — the ETT datasets are electricity transformer temperature measurements, relevant to grid planning and load management.
  • Traffic management — the Traffic dataset records hourly road occupancy rates on San Francisco freeways using 862 sensors, relevant to congestion and routing systems.
  • Financial forecasting — the Exchange Rate dataset covers daily rates for 8 currencies against USD from 1990 to 2016, where unreliable forecasts on abrupt moves are costly.
  • Weather and other abrupt-change domains — the abstract names traffic, energy and weather as target domains where unique trends make TSFM predictions unreliable.

Industry relevance: Automated abstention with a user-set drop percentage is a natural fit for deployed forecasting pipelines that need a coverage-versus-certainty trade-off and a lightweight model (TTM) that can be fine-tuned in resource-constrained environments.

Future Directions

  • Adaptive shapelet learning. The authors explicitly state that future work will explore expanding the approach to encompass adaptive shapelet learning, to handle distribution shift between validation and test sets — the failure mode observed on ETTh1.
  • Handling distribution shift more broadly. Because shapelets are learned once from the validation split, patterns that emerge later in the test period are not represented; developing methods that update or refresh shapelets over time is an open question.
  • Extending beyond TTM. The framework is demonstrated only with TTM as the prediction model; whether it transfers to other TSFMs such as Chronos, Moirai, TimesFM, LLMTime or GPT4TS is not reported.
  • Choosing drop percentage and threshold in a principled way. The ablations show monotonic improvement as more predictions are dropped, which raises the question of how to pick an operating point given a coverage requirement or a downstream cost model.

Target Audience

Researchers and practitioners working on time series forecasting with foundation models, selective prediction and abstention mechanisms, and interpretable time series pattern mining will benefit most. It is also useful for applied machine learning engineers in energy, traffic and finance who deploy forecasting models and need a way to flag or suppress unreliable predictions. Readers should be comfortable with MSE, fine-tuning concepts, and basic dictionary learning notation.

Authors’ abstract

Time series foundation models have recently gained a lot of attention due to their ability to model complex time series data encompassing different domains including traffic, energy, and weather. Although they exhibit strong average zero-shot performance on forecasting tasks, their predictions on certain critical regions of the data are not always reliable, limiting their usability in real-world applications, especially when data exhibits unique trends. In this paper, we propose a selective forecasting framework to identify these critical segments of time series using shapelets. We learn shapelets using shift-invariant dictionary learning on the validation split of the target domain dataset. Utilizing distance-based similarity to these shapelets, we facilitate the user to selectively discard unreliable predictions and be informed of the model's realistic capabilities. Empirical results on diverse benchmark time series datasets demonstrate that our approach leveraging both zero-shot and full-shot fine-tuned models reduces the overall error by an average of 22.17% for zero-shot and 22.62% for full-shot fine-tuned model. Furthermore, our approach using zero-shot and full-shot fine-tuned models, also outperforms its random selection counterparts by up to 21.41% and 21.43% on one of the datasets.

Read the original paper