Research
Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting
Overview Research area: Machine learning evaluation methodology, specifically benchmarking of time-series forecasting models (including time-series foundation models). Technical level: Intermediate. T

- arXiv
- 2603.08707
- Published
- 2026-03-09
- Authors
- Azul Garza, Renée Rosillo, Rodrigo Mendoza-Smith, David Salinas, Andrew Robert Williams, Arjun Ashok, Mononito Goswami, José Martín Juárez
AI summary
Overview
Research area: Machine learning evaluation methodology, specifically benchmarking of time-series forecasting models (including time-series foundation models).
Technical level: Intermediate. The paper assumes familiarity with forecasting metrics such as MASE and CRPS and with foundation-style forecasting models, but its central argument is conceptual rather than mathematically heavy.
Scope: The paper proposes and describes Impermanent, a live benchmark instantiated on GitHub open-source activity that scores forecasts sequentially over time to test whether forecasting models generalize as data distributions shift, rather than on a fixed held-out test set.
What This Paper Is About
Most time-series forecasting benchmarks use static train-test splits, which can be contaminated when foundation models are pretrained on data that later appears in test sets, or when test scores are used for model selection. This makes it hard to tell whether a model truly generalizes or simply performs well on a frozen slice of data. The paper builds a benchmark where forecasts must be issued before the outcomes exist and are scored only after new data arrives, so the target itself keeps changing over time.
Key Contributions
-
A live evaluation protocol for temporal generalization. The authors introduce what they describe as the first live benchmark designed specifically to evaluate temporal generalization in time-series forecasting, scoring forecasts sequentially over time on continuously updated data streams.
-
A deployment-faithful, leak-proof loop. At each cutoff time, models must produce forecasts before the corresponding ground truth exists. Cutoffs are spaced exactly one horizon apart, and the most recent cutoff is always excluded because observations may still be incomplete.
-
A concrete instantiation on GitHub activity. The benchmark uses GH Archive event streams covering four event types (issues opened, pull requests opened, push events, and new stargazers) across the top 400 repositories by star count, with stratified repositories by activity level and a repository-event-type pair defining a univariate series in the current release.
-
Open, automated infrastructure and leaderboards. All data pipelines, evaluation code, and leaderboard infrastructure are described as open and fully automated, with code at
https://github.com/TimeCopilot/impermanentand a live dashboard athttps://impermanent.timecopilot.dev.
Main Findings
-
Static benchmarks measure the wrong thing. Widely used benchmarks such as GIFT-Eval, FEV, and the Monash Forecasting Repository evaluate on held-out splits drawn from the same distribution as training data, which tests cross-sectional generalization but not whether performance persists across time.
-
The dataset is genuinely non-stationary. The GitHub streams mix smooth, trend-like behavior with spiky, volatile behavior. Exploratory views on a random sample of 25 repositories show pronounced intermittency (long zero-count stretches), burstiness (isolated spikes), level shifts, and heterogeneous scales across repositories and event types.
-
Spectral descriptors separate the event types. Using spectral centroid (typical frequency) and normalized spectral entropy (spectral spread), the paper reports that different event types populate different regions of that space: pull-request and push series concentrate in mid-to-high centroid regions with moderate-to-high entropy, while stargazer series span from low-centroid/low-entropy to high-centroid/high-entropy regimes.
-
Foundation models lead an early snapshot. In the snapshot up to February 12th, 2026, pretrained foundation models occupy the top four positions, and TimesFM leads on three of the four leaderboard columns, with a median scaled MASE of 0.609 and median scaled CRPS of 1.055.
-
Point accuracy and probabilistic calibration diverge. SeasonalNaive achieves a competitive MASE rank of 5.385 while showing poor probabilistic calibration with a CRPS rank of 9.495. AutoETS and AutoARIMA attain CRPS ranks (5.864 and 5.840) comparable to DynOptTheta (6.088) despite weaker point accuracy.
-
Full leaderboard values. Median scaled MASE and CRPS, followed by mean MASE rank and mean CRPS rank, are reported as: HistoricAverage 4.740 / 3.669, rank 9.943 / 8.401; SeasonalNaive 1.272 / 2.950, rank 5.385 / 9.495; Prophet 4.264 / 6.713, rank 9.791 / 8.638; AutoCES 2.272 / 2.385, rank 7.293 / 6.433; AutoARIMA 3.157 / 2.258, rank 7.842 / 5.840; AutoETS 2.802 / 2.232, rank 7.119 / 5.864; DynOptTheta 1.522 / 2.494, rank 5.838 / 6.088; Chronos 0.789 / 2.341, rank 3.340 / 4.348; Moirai 0.786 / 2.153, rank 3.028 / 4.173; TiRex 0.757 / 2.270, rank 2.938 / 4.223; TimesFM 0.609 / 1.055, rank 2.979 / 2.041.
-
The snapshot is not the conclusion. Because models are scored sequentially over an evolving stream, the paper states these rankings will shift as new cutoffs accumulate, making it possible to track whether early advantages persist rather than treating one leaderboard snapshot as definitive.
Methodology in Plain English
The benchmark draws from GH Archive, a public stream of GitHub event data. The authors filter for four event types, aggregate counts per repository with DuckDB, and roll the counts up into daily, weekly, and monthly granularities, subject to completeness thresholds of 90%, 95%, and 99% of constituent hours respectively. Each repository-event-type pair is a univariate series.
Evaluation follows a rolling-origin, prequential pattern. At each cutoff date, every model receives a context window of historical observations and must produce point and quantile forecasts for the next horizon before any ground truth is available. The protocol parameters are: for hourly data, horizon 24, max context 1,024, cutoff step 24 hours, first cutoff 2026-02-08; for daily data, horizon 7, max context 512, cutoff step 7 days, first cutoff 2026-01-04; for weekly data, horizon 1, max context 114, cutoff step 1 week, first cutoff 2026-01-04; for monthly data, horizon 1, max context 24, cutoff step 1 month, first cutoff 2025-10-01. (The main text describes the frequencies with horizons hourly h=24, daily h=7, weekly h=4, and monthly h=1, which differs from the weekly horizon of 1 listed in Table 1.)
Accuracy is measured with MASE, and the predictive distribution is measured with a scaled CRPS approximated from nine quantile levels, Q = {0.1, 0.2, ..., 0.9}. Both metrics are computed per series and aggregated per subdataset using the median. To make scores comparable across subdatasets, each metric is divided by the zero model's score, with a floor at the 10th percentile of strictly positive zero-model scores to prevent unstable ratios when the zero-model score is very small.
Eleven models are evaluated in three groups. Baselines: ZeroModel, HistoricAverage, SeasonalNaive. Statistical models: AutoARIMA, AutoETS, AutoCES, Dynamic Optimized Theta, and Prophet, all run on CPU. Foundation models: Chronos-2, Moirai 2.0-R-Small, TimesFM 2.5, and TiRex, each run on an A10G GPU with batch size 64. Only open-source time-series foundation models with released weights and reproducible inference code are included, and all models run through TimeCopilot.
The pipeline runs as serverless jobs on Modal with artifacts stored on Amazon S3. Forecasting dispatches one job per (model, cutoff) pair, with CPU models on 32-core instances and foundation models on NVIDIA A10G GPUs, up to 125 containers in parallel. Evaluation scores stored forecasts once ground truth arrives and rebuilds the leaderboard by reading all per-cutoff metric files, scaling by the zero-model baseline, and writing a single result Parquet. The full cycle is triggered weekly, and every stage is idempotent so re-runs skip completed work and new models can be added without reprocessing history.
Why This Matters
Impact on research. Static benchmarks cannot distinguish genuine transfer from memorization, data leakage, or test-set contamination—concerns the paper notes grow more relevant as foundation models scale and training-data curation becomes less transparent. A live benchmark changes the object of measurement from one-off accuracy on a frozen test set to sustained accuracy, robustness to distributional shift and shocks, and the stability of model rankings under ongoing change.
Real-world applications.
- Capacity planning and infrastructure scaling for services whose usage patterns are shaped by releases, external events, and behavioral shifts.
- Operational monitoring and incident-response forecasting, where data streams are intermittently active and prone to bursts and level shifts.
- Demand and engagement forecasting for platforms whose user activity is driven by community dynamics rather than stable seasonal cycles.
- Any deployment setting where a model must keep being re-scored as new data arrives, rather than being validated once before launch.
Industry relevance. The protocol reflects how forecasting is actually deployed: forecasts are issued before outcomes exist, and their quality is judged as reality unfolds. The benchmark also includes non-neural baselines and statistical models alongside foundation models, which keeps comparisons grounded for practitioners deciding whether a pretrained model is worth its inference cost.
Future Directions
- Expand to additional live data streams beyond GitHub activity, since the framework is explicitly designed to support broader future development.
- Enrich forecasting tasks with auxiliary contextual information, moving beyond the current setting where each repository-event-type pair is a univariate series.
- Capture cross-repository co-movement and incorporate additional covariates, which the paper identifies as important next steps for the current release.
- Use longer evaluation horizons to better understand performance stability and ranking dynamics over time, and to test whether benchmark performance in static settings translates into reliable performance after deployment.
Target Audience
Researchers and practitioners who build or select time-series forecasting models, especially those working with time-series foundation models and needing evidence that claimed generalization holds up over time. It is also relevant to benchmark designers and evaluation researchers concerned with contamination, leakage, and reproducibility, and to machine-learning engineers who deploy forecasting systems and must monitor their performance after launch. Readers looking for detailed per-subdataset or per-frequency breakdowns of the leaderboard will not find them in this paper; the reported results are aggregated across all subdatasets, frequencies, and cutoff dates.
Authors’ abstract
Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style models. While these models often claim broad generalization, existing evaluation protocols provide limited evidence. Indeed, most current benchmarks use static train-test splits that can easily lead to contamination as foundation models can inadvertently train on test data or perform model selection using test scores, which can inflate performance. We introduce Impermanent, a live benchmark that evaluates forecasting models under open-world temporal change by scoring forecasts sequentially over time on continuously updated data streams, enabling the study of temporal robustness, distributional shift, and performance stability rather than one-off accuracy on a frozen test set. Impermanent is instantiated on GitHub open-source activity, providing a naturally live and highly non-stationary dataset shaped by releases, shifting contributor behavior, platform/tooling changes, and external events. We focus on the top 400 repositories by star count and construct time series from issues opened, pull requests opened, push events, and new stargazers, evaluated over a rolling window with daily updates, alongside standardized protocols and leaderboards for reproducible, ongoing comparison. By shifting evaluation from static accuracy to sustained performance, Impermanent takes a concrete step toward assessing when and whether foundation-level generalization in time-series forecasting can be meaningfully claimed. Code and a live dashboard are available at https://github.com/TimeCopilot/impermanent and https://impermanent.timecopilot.dev.