cs.LGOct 4, 2026

What Does an Observability Foundation Model Know?

Authors: Dhyey Dharmendrakumar Mavani, Rian Atri, Tairan Ji

Organizations: Independent · Keiji AI

Abstract

A linear probe can show that a label is recoverable from a model's hidden states, but not whether that goes beyond what the input already reveals, or whether the model uses it. We audit Toto, an observability forecasting foundation model, on the Benchmark of Observability Metrics (BOOM) across five series-disjoint resplits, comparing linear probes on its frozen residual stream with models that read the raw input window and with Toto's architecture stripped of its trained configuration. Short-vs-medium cadence and metric type are more linearly recoverable from Toto's residuals than from the strongest raw-window model in every resplit (macro-F1 0.766 vs. 0.633 and 0.545 vs. 0.498). Domain is nearly tied, and series cardinality is recovered far better from the raw window. MOMENT-base shows related cadence, metric-type, and domain readouts. Recoverability is not use: exchanging Toto's residuals with those of high-burst donors moves a future-burstiness readout as intended but does not make forecasts consistently burstier than a randomized donor. A BOOM-trained coordination probe has negative zero-shot R^2 on the tested external benchmarks. We report each label against its strongest baseline.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Dynamical Parameters: An Interpretability Framework for Time-Series Foundation Models

    Sep 28, 2026Kang Yang, Gaofeng Dong, Liying Han +1Time Series Foundation ModelsHidden States

  2. Toto 2.0: Time Series Forecasting Enters the Scaling Era

    May 19, 2026Emaad Khwaja, Chris Lettieri, Gerald Woo +10Time Series Foundation ModelsTime Series Forecasting