Erased, Rerouted, or Rescaled? Post-Training and the Causal Quotient of a Language Model's Belief State
Organizations: The University of Tokyo · The Hong Kong University of Science and Technology · RWTH Aachen University · Harbin Institute of Technology, Shenzhen
Abstract
What happens to information a pretrained model already encodes when post-training no longer rewards using it? The common language of representation compression conflates three fates: information may be erased, rerouted away from the decision while still represented, or rescaled to occupy less variance while still represented and used. We make these fates identifiable in models whose pretraining recovers Bayesian belief states. A reward that reads only a coarse function of the hidden state defines an exact reward-null kernel. The kernel lets us separately measure whether the information remains recoverable, whether decisions causally depend on it, and how much activation variance it occupies. Theory says what is protected: KL-anchored reinforcement learning preserves the reference policy's log-odds among equally rewarded outputs, supervised and unanchored objectives carry no such constraint, and spectral compression implies neither erasure nor loss of use. In controlled worlds, post-training mostly reroutes or rescales reward-null information and leaves it decodable. Without an anchor decisions can stop using it although the representation survives, and with one they keep using it. Erasure appears only under prolonged weight decay, for distinctions that neither reward nor next-token prediction can see. Open language models show the same dissociation: in-context belief geometry stays decodable under late-layer spectral compression, and within-class behavior depends on the anchor. Post-training thus selects a causal quotient of the pretrained belief state: the reward defines decision-equivalence, the anchor and the state update protect part of what it ignores, and optimization decides whether the rest is erased, rerouted, or rescaled.
Figures & tables
| Question about | Operational meaning | |
|---|---|---|
| Is it still represented? | linear recoverability of the belief’s coordinates along | |
| Does the decision still use it? | causal sensitivity of the policy to | |
| How much space does it occupy? | share of the activation variance carried by |
| Variant | What it isolates |
|---|---|
| Worlds | |
| separable | independent task and nuisance chains: the filter needs none of the kernel, and half of it is next-token-invisible |
| visible control | the same with the whole kernel next-token-visible |
| entangled | filtering the task variable requires the nuisance belief |
| Mess3 | the process of Shai et al. (2024) |
| Action designs | |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.