Skip to content
AI.info

Research

What Does an Observability Foundation Model Know?

Overview Research area: Interpretability of time-series foundation models (specifically an observability forecasting foundation model), combining linear probing, control baselines, and causal-style in

What Does an Observability Foundation Model Know?
arXiv
2610.05577
Published
2026-10-04
Authors
Dhyey Dharmendrakumar Mavani, Rian Atri, Tairan Ji

AI summary

Overview

Research area: Interpretability of time-series foundation models (specifically an observability forecasting foundation model), combining linear probing, control baselines, and causal-style interventions.

Technical level: Intermediate. The paper assumes familiarity with linear probes, frozen representations, macro-F1, and R-squared, but the argument is stated in plain terms.

Scope: The paper audits what is linearly recoverable from the frozen hidden states of Toto, an observability forecasting foundation model, on the Benchmark of Observability Metrics (BOOM), and separately tests whether that recoverable information is actually used in forecasting.

What This Paper Is About

A linear probe can show that a label (such as sampling cadence or metric type) can be read out of a model's hidden states, but that alone does not prove the model learned something the raw input does not already reveal, nor that the model uses that information when it forecasts. The authors build an audit that asks three questions per label: is it more recoverable from Toto's residuals than from the input window, does that depend on Toto's trained weights and their order, and does moving the readout move the forecast. The goal is a label-by-label profile of "what the model knows," defined operationally as held-out linear recoverability, rather than a single verdict.

Key Contributions

  1. An audit design for forecasting representations. Every held-out probe score is paired on the same windows with input baselines (a six-statistic linear probe and four raw-window models trained from scratch) and backbone baselines (Toto with randomly initialized weights and Toto with its 12 trained transformer blocks run in a non-identity order), plus a separate donor exchange that tests forecast response.

  2. A label-by-label result on BOOM across five series-disjoint resplits. Cadence and metric type are more linearly recoverable from Toto's residuals than from every tested raw-window model in all five resplits; domain is nearly tied; cardinality is recovered far better from the raw window.

  3. Bounds on how far probe readouts can be interpreted. In both Toto and MOMENT-base, exchanging residuals moves the targeted readout without consistently moving the forecast endpoint, and a BOOM-trained coordination probe has negative zero-shot R-squared on the tested external benchmarks.

  4. A reproducibility artifact. Code and audited results are publicly available at https://github.com/DhyeyMavani2003/toto-interp/tree/neurips-2026-v1.

Main Findings

  • Cadence is the largest gap. Toto reaches macro-F1 0.766 ± 0.024 versus 0.633 ± 0.047 for the strongest raw-window model (FNO), a paired gap of +0.133 ± 0.063, positive in 5/5 resplits. The six-statistic probe reaches 0.409 ± 0.047.

  • Metric type is a smaller but consistent gap. Toto 0.545 ± 0.037 versus 0.498 ± 0.007 for the strongest raw-window model (GBDT), a paired gap of +0.047 ± 0.036, positive in 5/5 resplits. The six-statistic probe reaches 0.347 ± 0.030.

  • Domain is nearly tied. Toto 0.484 ± 0.011 versus GBDT 0.482 ± 0.040, a paired gap of +0.002 ± 0.040, positive in only 2/5 resplits. Toto's accuracy is lower than the GBDT's (0.534 versus 0.573, 0/5 resplits).

  • Cardinality reverses. Toto 0.471 ± 0.029 versus FNO 0.961 ± 0.016, a paired gap of -0.490 ± 0.044, positive in 0/5 resplits. The CNN, FNO, and patch Transformer all reach macro-F1 near 0.95, and even the two baselines that see neither variate channels nor a coverage channel still beat Toto: the six-statistic probe at 0.635 ± 0.027 and the GBDT at 0.663.

  • Backbone baselines are clearly below the pretrained model on the two strong labels. Random initialization gives cadence 0.547 ± 0.040 and block permutation 0.567 ± 0.020, versus Toto's 0.766; the shuffled-label floor is 0.486 ± 0.011. For metric type, random init is 0.398 ± 0.035, block perm 0.420 ± 0.016, shuffled labels 0.321 ± 0.032.

  • Common-support probes preserve two gaps. Within a fixed stratum of the other labels, cadence (Infrastructure/gauge) gives Toto 0.697 ± 0.060 versus random init 0.586 ± 0.121 and the six-statistic probe 0.485 ± 0.137, winning 5/5 against both. Metric type (Application Usage/Short) gives 0.502 ± 0.053 versus 0.395 ± 0.048 and 0.371 ± 0.034, winning 5/5 against both. Domain (Short/gauge) gives 0.406 ± 0.073 versus 0.350 ± 0.091 and 0.376 ± 0.087, winning 4/5 and 3/5.

  • Labels are correlated, unevenly. Domain and metric type have series-level Cramér's V = 0.402 ± 0.011, while cadence is nearly independent of both (0.055 ± 0.012 with domain and 0.050 ± 0.018 with metric type).

  • MOMENT-base shows related readouts. On the same resplits, MOMENT's cadence (0.632 ± 0.113), metric type (0.499 ± 0.022), and domain (0.437 ± 0.014) all beat its random initialization (0.556, 0.452, 0.409) and the six-statistic probe (0.401, 0.351, 0.337) in 5/5 resplits, with gaps of +0.231 ± 0.044, +0.148 ± 0.030, and +0.100 ± 0.023. Its cardinality readout (0.259 ± 0.010) matches random initialization (0.257 ± 0.018) and trails the six-statistic probe (0.668 ± 0.013) with a gap of -0.409 ± 0.012 (0/5).

  • Moving the readout does not consistently move the forecast. In Toto's donor exchange over 40 triples, the probe check is 0.655 ± 0.136 and is the same at every blend; the fraction of targets whose forecast is burstier under the high-burst donor is below one half on average at every blend (0.480 ± 0.086 at β = 0.25, 0.430 ± 0.046 at β = 0.50, 0.390 ± 0.127 at β = 1.0) and above one half in at most one resplit per blend. Median WAPE ratios favor the high-burst donor at the point estimate (0.969 ± 0.062, 0.952 ± 0.108, 0.903 ± 0.155) but their half-widths span one. MOMENT's matched interchange shows the same dissociation.

  • The tested dynamic label is not better recovered from Toto's residuals than from simple statistics. Toto's future-burstiness probe reaches R-squared 0.106 ± 0.100 versus 0.111 ± 0.066 for the six-statistic probe (above it in 3/5 resplits).

  • Zero-shot coordination-probe transfer fails. The BOOM-trained coordination probe, applied without refitting, has R-squared -0.938 ± 0.654 across 11 FEV configurations and -17.668 ± 10.896 across four LSTF datasets, worse than predicting the mean in both cases.

  • Held-out-combination stress tests are reported as coverage ranges, not confidence intervals. For example, cadence with domain × metric type held out gives Toto 0.865 [0.627, 1.000] versus random init 0.541 [0.250, 1.000]; cadence by domain gives 0.654 [0.513, 0.766] versus 0.511 [0.420, 0.575].

Methodology in Plain English

The authors take a public observability foundation model, Toto (the Datadog/Toto-Open-Base-1.0 checkpoint), and freeze it. They split a 1024-observation context into 16 non-overlapping 64-observation patches and read activations at the patch embedding and blocks 0–11. On those frozen activations they train simple probes: logistic regression for categorical labels, ridge regression for continuous ones, with the view chosen on validation series.

For every label, they compare the probe score against two families of controls. The input baselines ask whether the label is already visible in the raw window: a linear probe on six summary statistics of the last patch, and four models trained from scratch on the same windows (a CNN, a Fourier neural operator, a patch Transformer, and a gradient-boosted tree on 15 engineered features). The backbone baselines ask whether the result depends on Toto's trained weights: Toto with randomly initialized weights, and Toto with its 12 pretrained blocks run in a shuffled order.

To test whether the model actually uses what the probe reads, they run a donor exchange. They fit a future-burstiness probe at a fixed layer, pair low-burstiness target windows with high-burst donors and randomized donors from other series, and blend donor residuals into the target at weights of 0.25, 0.5, and 1.0. They check that the exchange moves the readout as intended, then check whether the forecast follows. Separately, they take a probe trained on BOOM and apply it unchanged to external data to test zero-shot transfer.

The resplit is the unit of analysis: five seeded, series-disjoint 70/15/15 resplits (seeds 42–46) of the 2,807 BOOM series, with training capped at the first 500 shuffled training series, giving 500/421/422 train, validation, and test series per resplit.

Why This Matters

Impact on research. The paper gives a template for separating decodability from use in time-series foundation models, where probing work is much thinner than in language models. It shows that a probe score is only interpretable relative to the strongest baseline it beats, and that reporting representation and forecast outcomes side by side keeps the boundary between "recoverable" and "used" visible. It also demonstrates cross-model recurrence: Toto and MOMENT-base show related cadence, metric-type, and domain readouts, though the authors note this does not rank architectures and MOMENT's baselines are narrower.

Real-world applications.

  • Observability and monitoring platform builders deciding whether a foundation model's internal signals can be trusted for tasks like routing, labeling, or alerting on telemetry series.
  • Practitioners deploying TSFMs zero-shot who need to know which series attributes a model can distinguish before relying on it.
  • Benchmark and evaluation designers who need a paired-baseline design for probing claims rather than standalone probe scores.
  • Model-control researchers considering steering or editing a forecasting model's internals, for whom the negative forecast-response result is a caution against treating a probe readout as a control knob.

Industry relevance. The setting is industrial telemetry from an observability platform, so the labels audited (sampling cadence, metric type, domain, series cardinality) are exactly the metadata that monitoring systems use to organize series. The finding that cardinality is far better recovered from the raw window than from the model's residuals, and that cadence and metric type are better recovered from the residuals, tells platform builders where a learned representation adds information over simple feature extraction and where it does not.

Future Directions

  • Repeat the paired design across independent telemetry corpora and TSFMs, with comparable raw-window models and label definitions, since five resplits of one corpus measure split variability rather than uncertainty over independent observability data.

  • Resolve the pretraining overlap question. Toto was pretrained on an internal observability corpus and the overlap with BOOM is unknown, so the source of the observed signal remains unattributed.

  • Strengthen the intervention. A next Toto test would match donors on context and taxonomy as MOMENT's interchange does, preserve the target context while changing the intended factor, include sham exchanges, and report several forecast endpoints.

  • Broaden the label and protocol coverage. The cadence label covers only the Short and Medium buckets, the common-support probes omit the raw-window models, and controlled tests that vary context and patch scales would help explain the cadence result. Whether external series encode coordination is a separate question that only a probe refit on those series could test.

Target Audience

Researchers working on interpretability and probing of time-series foundation models; machine-learning engineers building or evaluating observability and forecasting systems; benchmark and evaluation designers who need paired-baseline methodology; and graduate students looking for a compact example of how to separate representation-level recoverability from evidence of model use. Readers without a background in linear probing or forecasting evaluation will find the core argument accessible at an intermediate level, though the tables and protocol appendices assume familiarity with macro-F1, R-squared, and cross-validation conventions.

Authors’ abstract

A linear probe can show that a label is recoverable from a model's hidden states, but not whether that goes beyond what the input already reveals, or whether the model uses it. We audit Toto, an observability forecasting foundation model, on the Benchmark of Observability Metrics (BOOM) across five series-disjoint resplits, comparing linear probes on its frozen residual stream with models that read the raw input window and with Toto's architecture stripped of its trained configuration. Short-vs-medium cadence and metric type are more linearly recoverable from Toto's residuals than from the strongest raw-window model in every resplit (macro-F1 0.766 vs. 0.633 and 0.545 vs. 0.498). Domain is nearly tied, and series cardinality is recovered far better from the raw window. MOMENT-base shows related cadence, metric-type, and domain readouts. Recoverability is not use: exchanging Toto's residuals with those of high-burst donors moves a future-burstiness readout as intended but does not make forecasts consistently burstier than a randomized donor. A BOOM-trained coordination probe has negative zero-shot R^2 on the tested external benchmarks. We report each label against its strongest baseline.

Read the original paper