A useful visual representation is a statistic of the observed past that retains the latent factors shared with the future and discards patch-private noise. We formalize this requirement as a conditional predictive sufficient statistic (CPSS). Under a shared-factor model of image patches, the mutual information between the past and the next patch equals the information the past carries about the shared factor, up to a remainder that the next patch itself fails to reveal. Predicting the next patch embedding with a cosine loss is maximum likelihood for a von Mises-Fisher model of that embedding's direction, and is therefore a tractable surrogate for the predictive information. The same population loss is also minimized by a constant embedding, so stop-gradient does not by itself select the sufficient statistic; it only blocks the symmetric gradient that implements the constant solution in one step. The regression target is a shallow embedding, which forces the network output back into that shallow range and leaves the sufficient statistic in intermediate blocks. Small causal Transformers on MNIST and CIFAR-10 are used as diagnostics, not as a leaderboard. On MNIST the future shift and the stop-gradient move probe accuracy by tens of points, and the CPSS readout peaks before the output. On CIFAR-10, with the same short budget and no augmentation, every objective lands near a linear classifier on pixels. What still matches the derivation is the geometry: the CPSS output is a worse readout than its best intermediate block, next-pixel regression does not pay that penalty, and removing the stop-gradient collapses the effective rank of the embedding even when the pretext loss looks perfect.
Figures & tables
Run
Removed piece
Predicted probe
CPSS
none
mid-layer peak
nostop
stop-gradient
collapse, low accuracy
noshift
future shift
loss solved, probe weak
pixel
embedding target
weaker class statistic
Table 1: Each control removes one hypothesis of the derivation. The predicted probe change is relative to CPSS.
Run
Loss
Emb.
Best
Out
Rank
MNIST
CPSS
−0.926
15.0
87.3 (2)
77.2
17.1
nostop
−1.000
14.1
60.0 (6)
60.0
4.0
noshift
−1.000
14.7
50.1 (4)
45.6
14.6
pixel
0.027
15.1
89.6 (2)
87.1
17.1
CIFAR-10
Table 2: Last-token linear probe (%). Emb. is block 0, Out is the final block, Best is the highest probe and its block index. Rank is the effective rank of the patch embedding. A linear classifier on raw pixels scores 91.6% on MNIST and 32.1% on CIFAR-10.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Data
Run
0
1
2
3
4
5
6
MNIST
CPSS
15.0
84.1
87.3
86.2
84.6
80.7
77.2
MNIST
nostop
14.1
34.3
47.4
49.0
49.0
49.2
60.0
MNIST
noshift
14.7
15.7
43.2
46.4
50.1
49.1
45.6
MNIST
pixel
15.1
83.8
89.6
87.8
87.0
86.6
87.1
CIFAR-10
CPSS
20.0
28.3
31.2
32.2
31.7
29.4
28.2
CIFAR-10
nostop
18.8
21.9
26.2
27.1
28.3
27.1
32.7
Appendix
Table 3: Last-token linear probe accuracy (%) at every block.
Data
Run
0
1
2
3
4
5
6
MNIST
CPSS
17.1
90.6
102.1
96.7
74.8
56.9
48.0
MNIST
nostop
4.0
11.3
15.5
17.2
17.9
18.3
52.3
MNIST
noshift
14.6
16.7
20.2
20.6
20.4
20.0
19.5
MNIST
pixel
17.1
93.8
103.5
84.2
75.2
71.5
80.5
CIFAR-10
CPSS
21.0
57.0
59.9
57.4
45.7
36.1
29.6
CIFAR-10
nostop
2.1
2.8
3.2
3.4
3.4
3.5
37.9
Appendix
Table 4: Effective rank of the last-token state at every block.