A linear probe can show that a label is recoverable from a model's hidden states, but not whether that goes beyond what the input already reveals, or whether the model uses it. We audit Toto, an observability forecasting foundation model, on the Benchmark of Observability Metrics (BOOM) across five series-disjoint resplits, comparing linear probes on its frozen residual stream with models that read the raw input window and with Toto's architecture stripped of its trained configuration. Short-vs-medium cadence and metric type are more linearly recoverable from Toto's residuals than from the strongest raw-window model in every resplit (macro-F1 0.766 vs. 0.633 and 0.545 vs. 0.498). Domain is nearly tied, and series cardinality is recovered far better from the raw window. MOMENT-base shows related cadence, metric-type, and domain readouts. Recoverability is not use: exchanging Toto's residuals with those of high-burst donors moves a future-burstiness readout as intended but does not make forecasts consistently burstier than a randomized donor. A BOOM-trained coordination probe has negative zero-shot R^2 on the tested external benchmarks. We report each label against its strongest baseline.
Figures & tables
Figure 1 : Held-out test macro- F1 of Toto, its baselines, and the shuffled-label floor for each taxonomy label (means ± half-widths over five resplits; § 4 ). The dashed line marks Toto’s mean; the orange square is the strongest raw-window model for that label.
Figure 2 : Method. (a) Linear probes on frozen Toto residual views (pretrained backbone, or the randomly initialized and block-permuted backbone baselines) versus the input baselines (raw-window models and a six-statistic probe), compared as paired gaps on held-out series. (b) Layer-11 donor exchange, read by the probe and the forecast.
A. Input baselines
Label
Toto
Six-stat. probe
Strongest raw-window model
Paired gap
Wins
Cadence
0.766±0.024
0.409±0.047
FNO 0.633±0.047
+0.133±0.063
5/5
Metric type
0.545±0.037
0.347±0.030
GBDT 0.498±0.007
+0.047±0.036
5/5
Domain
0.484±0.011
0.371±0.037
GBDT 0.482±0.040
+0.002±0.040
2/5
Cardinality
0.471±0.029
0.635±0.027
FNO 0.961±0.016
−0.490±0.044
0/5
B. Backbone baselines and shuffled-label floor
Table 1: Held-out test macro- F1 on BOOM (means ± half-widths over five resplits). A: Toto against the input baselines. The strongest raw-window model is chosen by mean validation macro- F1 ; the paired gap is Toto minus that model, and wins count the resplits in which it is positive. B: backbone baselines and the shuffled-label floor. All raw-window models are in App. D .
Figure 3 : Held-out test macro- F1 in each resplit for Toto and the strongest raw-window model (thin lines: solid where Toto is higher, dashed where the raw-window model is higher) and their five-resplit means (bold). Each panel has its own vertical scale.
Target
Fixed stratum
Toto
Random init.
Six-stat.
Wins
Cadence
Infrastructure/gauge
0.697±0.060
0.586±0.121
0.485±0.137
5/5; 5/5
Metric type
Application Usage/Short
0.502±0.053
0.395±0.048
0.371±0.034
5/5; 5/5
Domain
Short/gauge
0.406±0.073
0.350±0.091
0.376±0.087
4/5; 3/5
Table 2: Common-support probes: held-out test macro- F1 within one fixed stratum of the other two labels (means ± half-widths over five resplits; classes balanced by series). Test support is 11–19 (cadence), 17–26 (metric type), and 11–18 (domain) series per class. Wins are against random initialization and against the six-statistic probe.
Label
MOMENT
Random init.
Six-stat.
Gap to six-stat. (wins)
Cadence
0.632±0.113
0.556±0.105
0.401±0.099
+0.231±0.044 (5/5)
Metric type
0.499±0.022
0.452±0.029
0.351±0.042
+0.148±0.030 (5/5)
Domain
0.437±0.014
0.409±0.020
0.337±0.034
+0.100±0.023 (5/5)
Cardinality
0.259±0.010
0.257±0.018
0.668±0.013
−0.409±0.012 (0/5)
Table 3: MOMENT-base held-out test macro- F1 on the same five resplits (means ± half-widths; views chosen on validation series). The last column is the paired gap to the six-statistic probe and the number of resplits in which it is positive.
Figure 4 : Probe check versus forecast endpoint across the residual blend β . The shaded band is the probe check (mean ± half-width), which is the same at every blend; points are resplits, and the line joins the forecast-endpoint means ± half-widths. The dashed line is parity.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Parameters
Input
Training
Checkpoint selection
CNN
39.5k
masked ctx + coverage
AdamW, 20 ep, batch 16
best val epoch
FNO
105.9k
masked ctx + coverage
AdamW, 20 ep, batch 16
best val epoch
Patch Transformer
246.7k
masked ctx + coverage
AdamW, 20 ep, batch 16
best val epoch
GBDT
–
15 eng. features
HistGradBoost, 300 iter
none (fixed budget)
Appendix
Table 4: Raw-window baseline models. Neural models see the masked full context plus a coverage channel; GBDT sees 15 engineered features.
Target
Held-out support
Toto macro- F1 range
Random init. range
Cadence
Domain × metric type
0.865[0.627,1.000]
0.541[0.250,1.000]
Cadence
Domain
0.654[0.513,0.766]
0.511[0.420,0.575]
Metric type
Domain × cadence
0.258[0.130,0.472]
0.179[0.135,0.224]
Metric type
Domain
0.281[0.188,0.505]
0.220[0.161,0.266]
Appendix
Table 5: Rotated within-BOOM held-out-combination tests. Each bracket spans scores from distinct held-out values across the five resplits; it is not an uncertainty interval.
Blend β
Forecast burstier
Resplits >0.5
Median WAPE ratio
0.25
0.480±0.086
1/5
0.969±0.062
0.50
0.430±0.046
0/5
0.952±0.108
1.00
0.390±0.127
1/5
0.903±0.155
Appendix
Table 6: Toto donor exchange: fraction of the 40 targets per resplit whose forecast is burstier under the high-burst than under the randomized donor, and the median WAPE ratio (high-burst/randomized). Means ± half-widths across resplits; the same triples are used at every blend. Donors come from other series and are not matched on covariates or taxonomy. Probe check: 0.655±0.136 at every blend.
Benchmark set
Entries
Coordination R2
FEV
11
−0.938±
0.654
LSTF
4
−17.668±
10.896
Appendix
Table 7: Zero-shot coordination-probe transfer (BOOM-trained probe, no refitting). R2 is macro-averaged over the datasets of each set within a resplit, relative to a constant-mean reference; means ± half-widths across resplits.
Dynamic label
MOMENT R2
Random init.
Six-stat.
MOMENT above six-stat.
Coordination
0.092±
0.047
0.083±
0.041
0.071±
0.041
3/5
Current burstiness
0.024±
0.025
0.007±
0.018
0.057±
0.620
2/5
Future burstiness
0.015±
0.027
0.002±
0.021
−0.005±
0.082
2/5
Shift risk
−0.006±
0.067
−0.018±
0.062
0.021±
0.076
1/5
Appendix
Table 8: MOMENT dynamic readouts on held-out BOOM series. Views are selected on validation R2 and evaluated once on test series. The near-constant sparsity targets are omitted; their six-statistic R2 values are unstable. Means ± half-widths (§ 4 ).
Blend β
Lower future-MAE fraction
Resplits >0.5
0.25
0.410±0.084
0/5
0.50
0.365±0.109
0/5
1.00
0.405±0.104
1/5
Appendix
Table 9: MOMENT matched interchange. Each resplit evaluates 40 held-out target/real/matched-null triples with the same triples reused at each blend. Probe check (blend-invariant): 0.590±0.071 , above 0.5 in 5/5 resplits. The lower future-MAE fraction is the forecast endpoint. Means ± half-widths (§ 4 ).
Target
Toto
Random init.
Block perm.
Cadence
0.766±0.024
0.547±0.040
0.567±0.020
Metric type
0.545±0.037
0.398±0.035
0.420±0.016
Domain
0.484±0.011
0.362±0.017
0.394±0.011
Cardinality
0.471±0.029
0.250±0.025
0.428±0.127
Appendix
Table 10: Held-out test macro- F1 of Toto and all raw-window and backbone baselines over five resplits. Neural raw-window checkpoints are selected on validation data. Means ± half-widths (§ 4 ).