How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing
Organizations: Independent Researcher
Abstract
Probes are the workhorse of interpretability. If a model's hidden states predict a variable, the model is said to represent it. But a probe score has no fixed meaning. An of 0.6 may only reflect what the input already gives away, and the same score can mean different things on different data. We propose reading every probe score against two reference points: a floor, what a declared set of simple inputs already predicts, and a ceiling, what the full input can predict. The gap between them, the headroom, is the range in which a probe can show that a model computes something beyond the simple inputs. We prove that headroom vanishes in two ways: the target stops depending on a hidden variable the model must infer, or the input stops revealing it. We test this on transformers trained for in-context meta-analysis, which must infer the hidden heterogeneity between studies to weight them correctly, and where both reference points are known. Under distribution shift, probe scores fall and prediction error rises --, yet the model recovers a similar share of the headroom, indicating that the data lost information, not the representation. We then analyze the real models. The single-cell foundation model scGPT encodes biological variability only partially. We also revisit four influential LLM probing studies, which claim that models represent geography, the state of an Othello board, truth, and the demographics of their users. Against a floor computed from the input text alone, some of these claims hold, while others are largely explained by the text itself.
Figures & tables
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Model size | Parameters | Peak learning rate | Epoch budget | |||
| xs | 16 | 2 | 1 | 3,345 | 3,300 | |
| s | 32 | 4 | 2 | 25,537 | 3,300 | |
| m | 64 | 4 | 2 | 100,225 | 3,000 | |
| l | 64 | 4 | 4 | 200,193 | 3,500 | |
| xl | 128 | 8 | 4 | 793,601 | 3,500 |
| Decoder | Fixed specification |
| GBM | HistGradientBoostingRegressor with 400 maximum iterations, learning rate , unrestricted depth, early stopping, and a internal validation fraction. |
| Shallow GBM | HistGradientBoostingRegressor with 600 maximum iterations, learning rate , maximum depth , early stopping, and a internal validation fraction. |
| MLP | Standardization followed by an MLP with hidden widths , penalty , 600 maximum iterations, early stopping, and patience 25. |
| Linear | Standardization followed by RidgeCV with |
| Draw | Meaning |
| mean expression | |
| overdispersion (latent ) | |
| sequencing depth of cell | |
| true expression rate | |
| observed count |
| Floor | Ceiling | Headroom | |||
| Domain | Target | Condition | |||
| Meta-analysis | ID | ||||
| OOD | |||||
| ID | |||||
| OOD | |||||
| Synthetic genomics | Full depth |
| Baseline | Target | Floor mean SD | Ceiling mean SD | Headroom mean SD | Probe mean SD | mean SD |
| Neuroscience | ||||||
| Force | 0.013 0.019 | 0.987 0.009 | 1.000 0.025 | 0.745 0.014 | 0.758 0.019 | |
| Force | 0.002 0.006 | 0.997 0.000 | 0.999 0.006 | 0.573 0.019 | 0.575 0.017 | |
| Force | 0.004 0.009 | 0.987 0.009 | 0.983 0.016 | 0.745 0.014 | 0.754 0.018 | |
| Force | 0.013 0.028 | 0.997 0.000 | 1.010 0.028 | 0.573 0.019 | 0.580 0.011 | |
| Force | 0.002 0.019 | 0.987 0.009 | 0.989 0.025 | 0.745 0.014 | 0.755 0.019 | |
| Dataset | Floor (95% CI) | Probe (7B / 70B) | (7B / 70B) |
| World places | 0.61 (0.59–0.62) | 0.88 / 0.91 | 0.70 / 0.77 |
| US places | 0.46 (0.44–0.48) | 0.80 / 0.86 | 0.63 / 0.75 |
| NYC places | 0.33 (0.29–0.37) | 0.22 / 0.36 | 0.17 / 0.04 |
| Historical figures | 0.60 (0.59–0.62) | 0.79 / 0.83 | 0.46 / 0.58 |
| Entertainment | 0.16 (0.05–0.28) | 0.79 / 0.89 | 0.75 / 0.86 |
| Headlines | 0.45 (0.43–0.47) | 0.56 / 0.75 | 0.21 / 0.54 |
| Floor features | Decoder | Floor acc. (95% CI) | Linear floor acc. | Chance acc. | Occupied acc. | Probe acc. (95% CI) | Linear probe acc. | linear probe | |
| Black/White/Empty | |||||||||
| : position | GBM | 0.62 (0.61–0.62) | 0.56 | 0.46 | 0.39 | 0.89 (0.89–0.90) | 0.74 | 0.72 | 0.33 |
| : recency | GBM | 0.66 (0.65–0.67) | 0.60 | 0.46 | 0.47 | 0.89 (0.89–0.90) | 0.74 | 0.68 | 0.24 |
| : placement | GBM | 0.80 (0.80–0.81) | 0.76 | 0.46 | 0.64 | 0.89 (0.89–0.90) | 0.74 | 0.45 | |
| Mine/Theirs/Empty | |||||||||
| : position | MLP | 0.62 (0.61–0.63) | 0.56 | 0.46 | 0.40 | 0.98 (0.98–0.99) | 0.99 | 0.96 | 0.96 |
| Dataset | Decoder | Floor acc. (95% CI) | Linear acc. |
| cities | Linear | 0.46 (0.40–0.52) | 0.46 |
| neg_cities | Linear | 0.45 (0.39–0.50) | 0.45 |
| sp_en_trans | Linear | 0.65 (0.52–0.76) | 0.65 |
| neg_sp_en_trans | GBM | 0.49 (0.38–0.62) | 0.56 |
| larger_than | Linear | 0.99 (0.97–1.00) | 0.99 |
| smaller_than | Linear | 0.97 (0.96–0.99) | 0.97 |
| Train Test | Decoder | Floor acc. (95% CI) | Linear acc. | Published probe acc. (13B / 70B) |
| larger_than + smaller_than sp_en_trans | GBM | 0.49 (0.44–0.55) | 0.50 | 0.97 / 0.97 |
| cities + neg_cities neg_sp_en_trans | Shallow GBM | 0.50 (0.45–0.55) | 0.50 | 0.96 / 0.99 |
| cities neg_cities 1 1 1 The decoder is selected on held-out source rows, which contain unseen city–country pairs, so no decoder is informative there (shallow GBM pseudo- , linear pseudo- ); neg_cities repeats every cities pair with “not” inserted and the label flipped, so pairs that shallow GBM memorised during refitting are wrong on the target ( , below the chance level), while the near-prior linear decoder stays at chance ( ). | Shallow GBM | 0.35 (0.32–0.37) | 0.47 | 0.75 / 0.37 |
| larger_than smaller_than | Linear | 0.00 (0.00–0.00) | 0.00 | 0.07 / 0.55 |
| Attribute | Evaluation | Floor acc. (95% CI) | Chance acc. | Published probe acc. (reading / control) | (reading / control) |
| Age | |||||
| Pooled | 0.98 (0.97–0.99) | 0.25 | 0.98 / 0.96 | 0.11 / 0.78 | |
| User turns | 0.98 (0.97–0.99) | 0.25 | – | – | |
| GPT-3.5 | 0.98 (0.97–0.99) | 0.25 | – | – | |
| Llama-2-Chat | 0.93 (0.89–0.96) | 0.25 | – | – | |
| GPT-3.5 Llama-2-Chat | 0.88 (0.86–0.90) | 0.25 | – | – | |
| Condition | ID | OOD | ||
| Target | ||||
| Estimated floor | ||||
| Estimated ceiling | ||||
| Estimated headroom | ||||
| Hidden probe | ||||
| Probe on , | ||||
| Target | Black/White/Empty | Mine/Theirs/Empty |
| Layer | 5 | 6 |
| Estimated floor (rung ) | 0.762 (0.754–0.768) | 0.803 (0.797–0.808) |
| Ceiling (exact) | 1.000 | 1.000 |
| Headroom | 0.238 | 0.197 |
| Probe on | 0.768 (0.762–0.774) | 0.808 (0.803–0.814) |
| Hidden probe | 0.745 (0.738–0.752) | 0.986 (0.981–0.990) |
| MSE | |||||||
| ID | OOD | ID | OOD | ID | OOD | OOD/ID | |
| xs | 0.28 0.20 | 0.30 0.04 | 0.22 0.20 | 0.24 0.03 | 1.05 0.01 | 5.05 2.25 | 28.3 |
| s | 0.60 0.04 | 0.29 0.03 | 0.54 0.03 | 0.26 0.02 | 1.02 0.00 | 2.44 0.37 | 14.2 |
| m | 0.60 0.03 | 0.34 0.03 | 0.60 0.04 | 0.28 0.03 | 1.03 0.00 | 2.58 0.73 | 14.9 |
| l | 0.61 0.02 | 0.35 0.05 | 0.58 0.05 | 0.29 0.05 | 1.03 0.01 | 2.23 0.59 | 12.9 |
| xl | 0.64 0.05 | 0.40 0.02 | 0.64 0.02 | 0.34 0.02 | 1.03 0.00 | 2.08 0.51 | 12.0 |
| Target | Floor | Ceiling | Headroom | Representation probe | |
| Neuroscience | |||||
| Force | |||||
| Force | |||||
| Chemistry | |||||
| HOMO | |||||
| LUMO | |||||