Skip to content
AI.info

Research

Pre-trained Forecasting Models: Strong Zero-Shot Feature Extractors for Time Series Classification

Pre-trained Forecasting Models: Strong Zero-Shot Feature Extractors for Time Series Classification Overview Research area: Time series machine learning, specifically transfer learning and foundation m

Pre-trained Forecasting Models: Strong Zero-Shot Feature Extractors for Time Series Classification
arXiv
2510.26777
Published
2025-10-30
Authors
Andreas Auer, Daniel Klotz, Sebastinan Böck, Sepp Hochreiter

AI summary

Pre-trained Forecasting Models: Strong Zero-Shot Feature Extractors for Time Series Classification

Overview

Research area: Time series machine learning, specifically transfer learning and foundation models — using pre-trained forecasting models as frozen feature extractors for time series classification (TSC).

Technical level: Intermediate. The paper is readable without deep transformer knowledge, but familiarity with embedding extraction, pooling strategies, and standard benchmarks (UCR/UEA) helps.

Scope: A large empirical study evaluating whether frozen, pre-trained forecasting models produce representations that transfer, zero-shot, to classification tasks across 127 univariate and 30 multivariate benchmark datasets.

What This Paper Is About

Recent time series foundation models are mostly pre-trained for forecasting, while classification models are usually pre-trained specifically for classification. This raises the question of whether forecasting representations are general enough to transfer to other tasks. The authors test this by freezing a range of pre-trained forecasting models, extracting embeddings from them without any fine-tuning, and training simple classifiers on top to see how well those representations support time series classification.

Key Contributions

  1. Demonstrating that forecasting representations transfer. The authors show that embeddings from pre-trained forecasting models reach classification accuracy on par with, and in some cases above, models pre-trained explicitly for classification — even though the forecasting models never saw the benchmark training data.

  2. A systematic analysis of representation extraction design choices. They ablate how to aggregate hidden states across layers and along the sequence dimension, and how to combine per-variate embeddings for multivariate data, giving practical guidance for using forecasting models as feature extractors.

  3. Two model-agnostic embedding augmentations. They propose adding absolute sample statistics (mean, standard deviation, minimum, maximum over patches, with k=8 in main experiments) and embeddings of the first-order differenced series to the model's own representation.

  4. Evidence that forecasting and classification performance are correlated. A trend analysis against CRPS on the GiftEval forecasting benchmark suggests better forecasters tend to make better feature extractors, with acknowledged noise in the trend.

Main Findings

  • Forecasting models are competitive or better as zero-shot extractors. With a Random Forest classifier, TiRex reaches 0.80 overall accuracy without augmentations and 0.80 with both augmentations (univariate 0.80 → 0.81, multivariate 0.74 → 0.74). Mantis, a model pre-trained specifically for classification, reaches 0.78 overall (0.79 univariate, 0.74 multivariate). Chronos Bolt (Base) goes from 0.76 to 0.78 overall with augmentations. DTW baselines score 0.73 overall (1-NN) and 0.71 (3-NN).

  • The zero-shot caveat matters. The authors note the forecasting models had no access to benchmark training data during pre-training, while Mantis, NuTime, and Moment did — making the forecasting results more striking. NuTime scores 0.67 overall, Moment (Large) 0.62, Moment (Base) 0.64, Moment (Small) 0.62.

  • Forecasting and classification performance correlate. Figure 2 shows higher classification accuracy tracking lower CRPS on GiftEval, though with considerable noise and notable under-performance from Chronos (possibly due to missing patch processing) and Toto.

  • No single architecture paradigm wins. Encoder, decoder, and encoder-decoder models appear among both the top and bottom performers. TiRex, the only non-Transformer model evaluated, performs best; the authors suggest its state-tracking capability may explain a better general representation.

  • Aggregation choices matter. Mean pooling over the sequence dimension plus concatenation over layers was the top-performing strategy across almost all models, and no other combination performed significantly better. For multivariate data, concatenating per-variate embeddings consistently beat mean or max pooling across all tested models.

  • Larger models usually help. In almost all cases, larger model sizes performed better, mirroring forecasting trends — with the exception of Moment, where the base model outperformed the large version (0.64 vs 0.62 overall).

  • Augmentations help, but significance varies. Both statistics and differencing augmentations improve most models individually, and the combination most often gives the best result, though the statistical significance of the gains varies.

  • Classifier choice changes the ranking at the low end. With a linear model the ranking largely holds (TiRex 0.77 overall, Moirai Large 0.77, Mantis 0.76), but with the simplest classifier (1-NN) forecasting models no longer outperform Mantis (Mantis 0.76, TiRex 0.74, Moirai Large 0.75). The authors hypothesize that weaker classifiers cannot transform the feature space as effectively, and that classification-specific embeddings are already better aligned with the task.

Methodology in Plain English

The researchers treat a pre-trained forecasting model as a fixed encoder. A time series goes in; hidden states from inside the network come out; those hidden states are pooled into a single fixed-size vector; and a simple classifier — Random Forest, a single linear layer, or 1-NN — is trained on that vector for each dataset. The forecasting model's weights are never updated.

Because forecasting models have no canonical way to produce one embedding per series, the authors make explicit choices and then test them. They pool across the sequence dimension (mostly by taking the mean), and combine representations from different layers (mostly by concatenating them). For multivariate data, they run the often-univariate models on one variate at a time and concatenate the results. They then add two hand-crafted signals to the embedding: basic statistics of the raw series (mean, standard deviation, min, max over 8 patches), which instance normalization would otherwise erase, and a second embedding computed on the differenced series to expose step-to-step changes.

Evaluation uses the UCR archive (127 univariate datasets) and the UEA archive (30 multivariate datasets) with predefined train/test splits. They excluded 5 datasets with sample lengths over 2048 (MotorImagery, HandOutlines, StandWalkJump, EigenWorms, Rock) and 2 more for processing problems (InsectWingbeat, PLAID). Statistical comparisons use critical difference plots with pairwise Wilcoxon signed-rank tests and Holm correction at α = 0.1. Where models failed computationally, the DTW baseline score was substituted to avoid skewing the aggregates.

Why This Matters

Impact on research. The results challenge the assumption that pre-training objectives must match the downstream task. If a forecasting objective yields representations that rival classification-specific pre-training, then a single model may serve as a genuinely general-purpose time series foundation model, echoing how language and vision foundation models transfer across tasks.

Real-world applications (drawing on the time series types the authors say the benchmarks cover — sensor, audio, motion, and health data):

  • Health monitoring: classifying ECG, heartbeat, or motion-capture recordings using a model that was never trained on labeled medical data.
  • Sensor and industrial systems: labeling activity or device states from wearable and IoT sensor streams.
  • Audio and speech classification: applying the same frozen encoder to spoken-digit or audio-derived series.
  • Low-label settings: organizations with little labeled data can reuse one pre-trained encoder plus a cheap classifier rather than training task-specific models.

Industry relevance. The approach is practical: no fine-tuning, no per-dataset training of a large model, just embeddings plus a Random Forest or linear layer. That translates to lower compute and simpler deployment pipelines. For vendors already offering forecasting foundation models, this work expands the addressable use cases of their existing weights.

Future Directions

  • Fine-tuning. The study deliberately stays zero-shot; the authors note fine-tuning could improve results but was excluded because optimal strategies would likely be model-specific.
  • Direct comparison to supervised and task-specific classifiers. The paper omits this comparison, arguing that prior work already shows the pre-trained classification models they evaluate are competitive with those baselines.
  • Other downstream tasks. Probing generalizability on tasks such as anomaly detection is explicitly proposed.
  • Understanding why forecasting transfers. The correlation in Figure 2 is noisy, with unexplained under-performance from Chronos and Toto, and the claim that TiRex's state-tracking ability produces better general representations remains a hypothesis rather than a tested mechanism.

Target Audience

Researchers and practitioners working on time series foundation models, transfer learning, and representation learning, as well as applied machine learning engineers who need to classify time series but lack large labeled datasets. It is also relevant to anyone evaluating whether a forecasting pre-training objective is worth the investment as a general-purpose backbone.

Authors’ abstract

Recent research on time series foundation models has primarily focused on forecasting, leaving it unclear how generalizable their learned representations are. In this study, we examine whether frozen pre-trained forecasting models can provide effective representations for classification. To this end, we compare different representation extraction strategies and introduce two model-agnostic embedding augmentations. Our experiments show that the best forecasting models achieve classification accuracy that matches or even surpasses that of state-of-the-art models pre-trained specifically for classification. Moreover, we observe a positive correlation between forecasting and classification performance. These findings challenge the assumption that task-specific pre-training is necessary, and suggest that learning to forecast may provide a powerful route toward constructing general-purpose time series foundation models.

Read the original paper