Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation
Organizations: King Abdullah University of Science and Technology (KAUST)
Abstract
Test-time training (TTT) lets a model store information in its weights during inference. When the model learns from its own output, however, each update also changes the model that generates the next training example. Across 128K-token streams, retaining generated-text updates worsens prediction on independent human-written text with three TTT-E2E model configurations (labeled 125M, 760M, and 3B). The same failure occurs when Adam updates Qwen3-4B's existing weights. The same update mechanisms can improve on real text, so writing itself is not the failure. Three matched comparisons trace the causal pathway. Fixed Generation removes over 98% of the damage at 125M and 760M by using a frozen model to generate training chunks. Recorded Replay separates the loss caused by reading degraded text from the additional loss stored by updating on it. A paired one-update comparison then shows the local conflict: an update predicts its source better but new real text worse. This cost grows after Closed Loop adaptation, with a few trajectories accounting for most large failures. Finally, Settlement evaluates the candidate state on independent real text before commitment. It leaves mean endpoint gaps of 0.07 and -0.02 nats at 125M and 760M while retaining real-text adaptation. These results motivate checking prediction on independent evidence before retaining an update.
Figures & tables
| Model | ( worse) | 95% Interval |
|---|---|---|
| 125M TTT-E2E | ||
| 760M TTT-E2E | ||
| 3B TTT-E2E |
| Adam LR | Persistent self-writing 95% CI; worse | Real-text state range; better |
|---|---|---|
| Model | Fixed Generation | Closed Loop harm removed |
|---|---|---|
| 125M | ||
| 760M |
| Scale | Receiver operation on the same recorded text | Real-text NLL above control |
|---|---|---|
| 125M | Read only | |
| 125M | Read and write | |
| 125M | Additional cost of writing | |
| 3B | Additional cost of writing |
| Preceding History | Training-Chunk NLL ( better fit) | Next Real-Text NLL ( worse) | Future Repeated-4 ( more repetition) |
|---|---|---|---|
| Writes Off History | |||
| Closed Loop | |||
| Fixed Generation |
| History | 125M | 760M |
|---|---|---|
| Closed Loop | ||
| Writes Off |
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Question | Shared configuration | Controlled comparison |
|---|---|---|
| Long horizon | 125M, 760M, 3B; canonical six books; five seeds; 128K; width 8 | Closed Loop versus Writes Off |
| Feedback path | 125M and 760M; canonical six books; five seeds; 128K; Fixed Generation added | Closed Loop versus Fixed Generation |
| Read vs. write | Fixed recorded tokens; 125M: eight receiver books, five seeds; 3B: six source and six disjoint receiver books | Read Only versus Read + Write |
| Single update | Eight books and four stream positions; identical state and source text in each branch | Keep versus discard one update |
| Decoder/exposure | Decoder: canonical setup; exposure: one eight-book long-stream set, five seeds | Decoder endpoints versus a common Writes Off reference; evenly spaced versus bursty real text |
| Validation | Scale-specific long-book sets; each scale has its own paired Writes Off trajectory | Closed Loop versus Settlement |
| Policy | Final clean NLL | Final relative drift | Tail Distinct-2 | Tail Repeated-4 |
|---|---|---|---|---|
| Closed Loop | ||||
| Writes Off | ||||
| Fixed Generation | ||||
| Real-Text Learning | — | — |
| Model | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|---|---|---|
| 760M, DCLM | 3.69 | 1.44 | 1.15 | 1.23 | 1.18 | 1.19 | 1.14 | 1.18 |
| 3B, DCLM | 10.35 | 4.12 | 1.30 | 1.33 | 1.29 | 1.32 | 1.30 | 1.28 |
| 3B, Books finetune | 8.83 | 2.67 | 1.44 | 1.69 | 1.44 | 1.41 | 1.37 | 1.46 |
| Model | Condition | Initial NLL | Final NLL | Final excess over Writes Off [95% CI] |
|---|---|---|---|---|
| 125M | Writes Off | 3.3469 | 3.3845 | 0 |
| Fixed Generation | 3.3469 | 3.4361 | .0516 [.0248,.0772] | |
| Closed Loop | 3.3469 | 6.3942 | 3.0097 [1.8575,4.1972] | |
| 760M | Writes Off | 2.7913 | 2.9383 | 0 |
| Fixed Generation | 2.7913 | 3.0116 | .0733 [.0538,.0939] | |
| Closed Loop | 2.7913 | 8.9405 | 6.0021 [4.2695,7.7347] |
| Control | Endpoint gap to canonical Writes Off | Late slope (nats/evaluation) |
|---|---|---|
| top- | ||
| top- | ||
| typical- | ||
| full support | ||
| , top- | ||
| repetition penalty |
| Real fraction | Schedule | Real slots | Max self burst | [95% CI] | [95% CI] |
|---|---|---|---|---|---|
| 0% | Even | 0 | 105 | ||
| 5% | Even | 5 | 17 | ||
| 10% | Even | 11 | 8 | ||
| 20% | Even | 21 | 4 | ||
| 31% | Even | 33 | 3 | ||
| 31% | Bursty | 33 | 12 |
| Source book | Mean | Leave-one-out |
|---|---|---|
| 12 | .1862 | .1940 |
| 13 | .1953 | .1922 |
| 14 | .0617 | .2189 |
| 15 | .1331 | .2046 |
| 16 | .1602 | .1992 |
| 17 | .4196 | .1473 |
| History | source NLL | real NLL | future repeated-4 |
|---|---|---|---|
| Closed Loop | |||
| Writes Off | |||
| Fixed Generation |
| History | Position 1 | 33 | 65 | 97 |
|---|---|---|---|---|
| Closed Loop | .0062 | .0087 | .0250 | .1126 |
| Writes Off | .0063 | .0044 | .0038 | .0037 |
| Fixed Generation | .0042 | .0008 | .0011 | .0012 |
| Model / history | Mean | |||
|---|---|---|---|---|
| 125M / Closed Loop | ||||
| 125M / Writes Off | ||||
| 760M / Closed Loop | ||||
| 760M / Writes Off |
| Condition | NLL change | Relative drift | Distinct-2 | Repeated-4 |
|---|---|---|---|---|
| Writes Off | -.060 | .00000 | .140 | .831 |
| Adam | -.034 | .00627 | .187 | .775 |
| Adam | 1.171 | .02399 | .008 | .991 |
| Adam LR | [range of two state means] | State 0 | State 1 |
|---|---|---|---|
| Source stage | Median | Mean | Standard deviation | Cases above .5 nats |
|---|---|---|---|---|
| 1 | .0422 | .0439 | .025 | 0/72 |
| 17 | .0040 | .0898 | .385 | 3/72 |
| 33 | .0039 | .1399 | .620 | 3/72 |
| 49 | .0042 | .3135 | 1.009 | 6/72 |
| 65 | .0047 | .4855 | 1.274 | 9/72 |
| 81 | .0105 | .4858 | 1.267 | 9/72 |
| Policy | Final NLL | [95% interval] | Distinct-2 |
|---|---|---|---|
| Writes Off | 3.5767 | 0 | .8425 |
| Closed Loop | 6.4620 | 2.8853 [2.0072,3.8310] | .0082 |
| Repetition weighting | 5.1404 | 1.5637 [1.1858,1.9825] | .0469 |
| Random dose | 5.6325 | 2.0558 [1.6339,2.4946] | .0086 |
| Uniform dose | 6.0628 | 2.4861 [1.9510,3.0939] | .0155 |
| Read-only real context | 4.9746 | 1.3979 [1.0988,1.7140] | .0730 |
| Excess NLL change [95% CI] | Real-text benefit [95% CI] | |
|---|---|---|
| Read-only content | Final NLL | Difference from no-op [95% CI] |
|---|---|---|
| No-op | 6.4042 | 0 |
| Self-generated | 6.7214 | .3172 [-.4775,1.1963] |
| Random tokens | 7.3969 | .9927 [-.2699,2.0992] |
| Shuffled real | 5.5680 | -.8362 [-1.7643,.0324] |
| Fixed Generation text | 5.7912 | -.6130 [-1.5127,.2200] |
| Real text | 4.8349 | -1.5693 [-2.3011,-.8433] |
| Model / books | Closed Loop minus Writes Off | Settlement minus Writes Off | Generated admitted |
|---|---|---|---|
| 125M / 12 | 3.183 [2.106,4.253] | .067 [-.013,.169] | 22/936 |
| 760M / 4 | 1.263 [.799,1.618] | -.021 [-.058,.014] | 18/312 |
| Without validation | Excess [95% interval] | With Settlement | Excess [95% interval] |
|---|---|---|---|
| Closed Loop | 3.06 [.67,7.02] | Settlement | .04 [.00,.09] |
| Quarter Dose | .90 [.32,1.97] | + Quarter Dose | .02 [-.01,.04] |
| Read-Only Real Context | 1.04 [.19,2.59] | + Real Context | -.02 [-.03,.00] |
| Both controls | .28 [.13,.53] | + Both controls | .00 [-.01,.02] |
| Policy | Real admitted | Benefit [95% interval] |
|---|---|---|
| Reject All | ||
| Write All | ||
| Settlement |
| Policy | Candidate judged against | State committed | ( ) | ( ) |
|---|---|---|---|---|
| Ordinary writing | — | |||
| Joint commit | (each ) | |||
| Sequential | ||||
| Settlement | on next |
| Observed error | Generated kept | Real discarded | Source Masking NLL | Settlement NLL |
|---|---|---|---|---|
| .000 | 0 | 0 | 3.8072 | 3.8196 |
| .053 | 4 | 2 | 3.8116 | 3.8230 |
| .124 | 9 | 5 | 3.8129 | 3.8248 |
| .212 | 15 | 9 | 3.8377 | 3.8286 |
| .389 | 26 | 18 | 3.8584 | 3.8215 |
| .619 | 41 | 29 | 3.9242 | 3.8270 |
| Policy | Reward | Exact success | Retained blocks |
|---|---|---|---|
| Writes Off | |||
| Closed Loop | |||
| Fixed Generation | |||
| Settlement |
| Success rate (%) | Accepted | ||||
| Scale | Policy | Overall (274) | Seen (140) | Unseen (134) | blocks |
| – | Writes Off | – | |||
| 8 | Closed Loop | All updates | |||
| Settlement | |||||
| 16 | Closed Loop | All updates | |||
| Settlement | |||||