Pretraining Latent Information Feedback Transformers with Teacher Supervision
Organizations: Blavatnik School of Computer Science and AI, Tel Aviv University · The Hebrew University of Jerusalem
Abstract
Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discard alternative continuations. In this work, we remove this bottleneck during pretraining, introducing the LIFT (Latent Information Feedback Transformer) architecture and training method which enable LMs to propagate state across generation. We achieve this by turning recurrent-state learning into a teacher-forced prediction problem: each input token is paired with an information-dense state, derived from the next-token distribution of an off-the-shelf pretrained LM. The model, extended with a small number of additional parameters, is then trained to predict both the next token and the next state. As the input states are precomputed, pretraining remains fully parallel across positions. At inference, the model's own predicted states are fed back, with a minor computational overhead that decreases with model size. Experiments with pretrained models ranging from 135M to 1B parameters show that LIFT consistently outperforms standard Transformers and baselines on language modeling, downstream reasoning tasks, and procedural tasks under token-matched budget, while being on par with or ahead of compute-matched Transformers. Moreover, a controlled study on a state-tracking task shows that a tiny LIFT outperforms same-size Transformers trained on 8x more data, even when trained with the states of a Transformer that fails the task. Overall, we show that LMs can learn to exploit deep-to-shallow feedback during pretraining via scalable teacher supervision.
Figures & tables
| Model | Tokens | PPL | Arith. BPB | MC acc. % | Gen. F1/EM % | RC BPB | |
|---|---|---|---|---|---|---|---|
| 135M | LIFT | 13.4B | 30.65 | 2.29 | 39.4 | 15.5 | 1.078 |
| LIFT w/o states | 13.4B | 32.44 | 2.42 | 38.7 | 12.0 | 1.105 | |
| Transformer, token-match | 13.4B | 32.43 | 2.46 | 38.1 | 13.6 | 1.110 | |
| Transformer, compute-match | 19.3B | 31.04 | 2.50 | 39.5 | 14.1 | 1.080 | |
| Distillation | 13.4B | 31.77 | 2.36 | 39.1 | 12.6 | 1.089 | |
| 350M | LIFT | 34.9B | 23.82 | 2.12 | 45.0 | 23.2 | 0.921 |
| Token-matched 10B tokens | Compute-matched FLOPs | |||||||||||
| Model | FLOPs | PPL | Arith. | MC % | Gen. % | RC | Tokens | PPL | Arith. | MC % | Gen. % | RC |
| LIFT | 1.44e19 | 19.42 | 2.18 | 36.7 | 7.8 | 1.705 | 16.2B | 18.04 | 2.18 | 38.0 | 9.5 | 1.629 |
| Transformer | 1.02e19 | 20.30 | 2.35 | 36.2 | 6.7 | 1.761 | 23.0B | 18.02 | 2.30 | 38.1 | 8.5 | 1.622 |
| T 2 MLR (our run) | 2.32e19 | 19.92 | 2.27 | 36.5 | 7.8 | 1.737 | 10.0B | 19.92 | 2.27 | 36.5 | 7.8 | 1.737 |
| T 2 MLR (HF) | 2.35e19 | 21.55 | 2.38 | 35.3 | 6.5 | 1.816 | 10.1B | 21.55 | 2.38 | 35.3 | 6.5 | 1.816 |
| Multi-pass Transformer | 1.31e19 | 20.40 | 2.25 | 36.3 | 6.5 | 1.773 | 17.8B | 18.58 | 2.25 | 37.6 | 7.7 | 1.675 |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| 135M | 350M | 1B | |
| Non-embedding parameters | 134.3M | 348.9M | 1,073.7M |
| Total parameters | 237.0M | 490.2M | 1,484.9M |
| Hidden dimension | 1024 | 1408 | 2048 |
| Layers | 8 | 11 | 16 |
| Attention heads | 8 | 11 | 16 |
| Embeddings | tied | tied | untied |
| Suite | Tests | Items | Shots | Scoring | Group |
| Core 8: ARC-Easy, ARC-Challenge, CommonsenseQA, HellaSwag, OpenBookQA, PIQA, SocialIQA, WinoGrande | Science and commonsense questions | 500–1,267 | 5 | cloze acc. | MC |
| MMLU | 57 academic subjects | 14,042 | 5 | cloze acc. | MC |
| Basic skills (OLMES) | Arithmetic, string operations, pattern continuation, coding, logical reasoning, common knowledge | 5,967 | 5 | cloze acc. | MC |
| Generative QA: CoQA, SQuAD, Jeopardy, Natural Questions, DROP | Conversational, extractive and closed-book QA | 7,983 / 1,000 / 2,116 / 1,000 / 1,000 | 5 (CoQA 0) | F1 | Gen. |
| TriviaQA | Closed-book trivia | 7,993 | 5 | F1 | Gen. |
| BBH | 27 reasoning tasks with chain of thought | 6,511 | 3 | exact match | Gen. |
| Source | LIFT | w/o states | Tr. token | Tr. compute | Distillation | Seqs |
| 135M | ||||||
| C4 | 30.65 | 32.44 | 32.43 | 31.04 | 31.77 | 971 |
| Common Crawl | 31.55 | 33.44 | 33.55 | 32.01 | 32.63 | 50 |
| Books | 33.58 | 35.93 | 35.90 | 33.85 | 34.86 | 50 |
| Wikipedia | 19.64 | 20.70 | 20.78 | 19.87 | 20.39 | 50 |
| peS2o | 17.85 | 18.93 | 19.00 | 18.12 | 18.56 | 51 |
| PPL | 135M | 350M | 1B |
|---|---|---|---|
| vs. Transformer, token-match | -1.77 [-1.85, -1.70] | -1.33 [-1.39, -1.28] | -0.99 [-1.05, -0.93] |
| vs. Transformer, compute-match | -0.38 [-0.44, -0.33] | -0.50 [-0.55, -0.46] | -0.36 [-0.40, -0.32] |
| vs. Distillation | -1.11 [-1.17, -1.06] | -0.70 [-0.74, -0.66] | -0.59 [-0.63, -0.55] |
| Suite | LIFT 135M@29.6B | w/o states 135M@29.6B | Tr. token 13.4B | Tr. compute 19.3B | Distillation 13.4B | 2 tokens 29.6B |
|---|---|---|---|---|---|---|
| Core 8 (cloze) | 44.0 | 43.3 | 42.5 | 43.9 | 43.9 | 44.7 |
| MMLU (cloze) | 28.1 | 27.9 | 27.3 | 28.0 | 27.6 | 28.2 |
| Basic skills (cloze) | 46.0 | 44.8 | 44.6 | 46.4 | 45.8 | 47.2 |
| Generative QA | 13.4 | 11.3 | 11.4 | 13.8 | 10.7 | 15.5 |
| TriviaQA | 13.6 | 13.3 | 13.8 | 14.8 | 14.3 | 14.8 |
| BBH (CoT) | 19.5 | 11.3 | 15.6 | 13.8 | 12.9 | 14.7 |
| Suite | LIFT 350M@77.8B | w/o states 350M@77.8B | Tr. token 34.9B | Tr. compute 49.9B | Distillation 34.9B | 2 tokens 77.8B |
|---|---|---|---|---|---|---|
| Core 8 (cloze) | 51.2 | 49.8 | 48.9 | 50.8 | 50.4 | 51.7 |
| MMLU (cloze) | 30.2 | 30.1 | 30.3 | 31.0 | 30.6 | 31.1 |
| Basic skills (cloze) | 53.8 | 52.7 | 50.3 | 50.4 | 52.4 | 52.6 |
| Generative QA | 27.3 | 24.0 | 21.5 | 23.1 | 24.8 | 26.6 |
| TriviaQA | 21.9 | 20.6 | 20.1 | 21.7 | 21.4 | 23.6 |
| BBH (CoT) | 20.3 | 18.4 | 21.3 | 21.3 | 20.1 | 21.9 |
| Suite | LIFT 1B@222B | LIFT 1B@4001B | w/o states 1B@222B | w/o states 1B@4001B | Tr. token 107.4B | Tr. compute 152.4B | Distillation 107.4B | 2 tokens 222B |
|---|---|---|---|---|---|---|---|---|
| Core 8 (cloze) | 59.4 | 58.3 | 58.3 | 56.7 | 58.0 | 59.0 | 58.5 | 60.2 |
| MMLU (cloze) | 35.4 | 35.1 | 34.7 | 34.3 | 34.4 | 34.9 | 35.0 | 35.6 |
| Basic skills (cloze) | 62.6 | 62.2 | 60.6 | 61.5 | 59.3 | 60.9 | 62.5 | 63.0 |
| Generative QA | 42.2 | 42.1 | 39.3 | 39.5 | 38.6 | 40.5 | 41.4 | 43.5 |
| TriviaQA | 37.3 | 37.6 | 36.1 | 35.7 | 34.9 | 37.9 | 36.7 | 41.2 |
| BBH (CoT) | 25.6 | 27.8 | 23.2 | 26.2 | 25.2 | 26.0 | 26.4 | 27.4 |
| Model | PPL | MC % | Gen. % | RC (BPB ) | Basic skills | Pattern % | Pattern BPB |
|---|---|---|---|---|---|---|---|
| LIFT | 30.57 0.08 | 39.5 0.1 | 15.0 0.6 | 1.076 0.006 | 1.326 0.018 | 53.2 0.7 | 1.339 0.045 |
| LIFT w/o states | 32.39 0.07 | 38.8 0.2 | 11.9 0.5 | 1.105 0.005 | 1.364 0.007 | 53.2 0.7 | 1.380 0.046 |
| Transformer, token-match | 32.44 0.02 | 38.4 0.5 | 13.0 0.6 | 1.112 0.004 | 1.390 0.011 | 50.7 2.0 | 1.427 0.051 |
| Transformer, compute-match | 31.04 0.00 | 39.2 0.3 | 13.6 0.6 | 1.085 0.005 | 1.348 0.036 | 52.2 0.5 | 1.393 0.003 |
| Distillation | 31.81 0.04 | 39.2 0.4 | 13.0 0.4 | 1.096 0.006 | 1.361 0.019 | 50.5 1.4 | 1.373 0.022 |
| Number of operands | |||||||||||||
| Model | Tokens | Teacher (size@tokens) | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | Avg. | |
| 135M | LIFT | 13.4B | 135M@29.6B | 2.03 | 2.13 | 2.12 | 2.27 | 2.27 | 2.35 | 2.36 | 2.42 | 2.42 | 2.29 |
| LIFT | 13.4B | 350M@38.4B | 2.23 | 2.21 | 2.17 | 2.26 | 2.25 | 2.32 | 2.35 | 2.38 | 2.41 | 2.29 | |
| LIFT w/o states | 13.4B | 135M@29.6B | 2.24 | 2.32 | 2.26 | 2.37 | 2.37 | 2.47 | 2.49 | 2.55 | 2.56 | 2.42 | |
| LIFT w/o states | 13.4B | 350M@38.4B | 2.33 | 2.28 | 2.24 | 2.44 | 2.45 | 2.58 | 2.59 | 2.63 | 2.64 | 2.48 | |
| Transformer, token-match | 13.4B | – | 2.02 | 2.28 | 2.29 | 2.42 | 2.45 | 2.55 | 2.57 | 2.58 | 2.59 | 2.46 | |
| Number of operands | |||||||||||||
| Model | Tokens | Teacher (size@tokens) | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | Avg. | |
| 135M | LIFT | 13.4B | 135M@29.6B | 7.0 | 5.0 | 3.4 | 2.4 | 3.2 | 2.5 | 2.2 | 3.0 | 1.9 | 3.0 |
| LIFT | 13.4B | 350M@38.4B | 5.5 | 4.4 | 4.4 | 3.3 | 4.0 | 2.5 | 2.2 | 2.5 | 3.6 | 3.4 | |
| LIFT w/o states | 13.4B | 135M@29.6B | 6.0 | 4.5 | 3.8 | 3.1 | 3.1 | 2.7 | 1.7 | 3.1 | 2.8 | 3.2 | |
| LIFT w/o states | 13.4B | 350M@38.4B | 5.0 | 4.7 | 3.5 | 3.3 | 3.9 | 2.0 | 2.3 | 2.8 | 2.3 | 3.1 | |
| Transformer, token-match | 13.4B | – | 7.0 | 4.5 | 4.1 | 3.0 | 2.6 | 3.0 | 2.4 | 2.0 | 2.9 | 3.2 | |
| 135M | 350M | 1B | ||||
|---|---|---|---|---|---|---|
| Model | Acc. % | BPB | Acc. % | BPB | Acc. % | BPB |
| LIFT | 53.2 | 1.356 | 64.4 | 1.134 | 72.1 | 0.851 |
| LIFT w/o states | 52.6 | 1.383 | 63.7 | 1.121 | 70.6 | 0.965 |
| Transformer, token-match | 50.7 | 1.443 | 58.4 | 1.216 | 66.1 | 0.948 |
| Transformer, compute-match | 52.6 | 1.396 | 58.1 | 1.144 | 67.6 | 0.955 |
| Distillation | 48.9 | 1.392 | 60.7 | 1.153 | 69.7 | 0.890 |
| Model | PPL | Arith. | MC % | Gen. % | RC |
|---|---|---|---|---|---|
| Token-matched (10B tokens) | |||||
| LIFT | 19.42 0.02 | 2.182 0.007 | 36.7 0.27 | 7.8 0.61 | 1.705 0.023 |
| Transformer | 20.30 0.03 | 2.345 0.156 | 36.2 0.19 | 6.7 0.21 | 1.761 0.002 |
| T 2 MLR (our run) | 19.92 0.06 | 2.267 0.101 | 36.5 0.04 | 7.8 0.94 | 1.737 0.010 |
| Multi-pass Transformer | 20.40 0.04 | 2.249 0.064 | 36.3 0.36 | 6.5 0.25 | 1.773 0.017 |
| Compute-matched ( FLOPs) | |||||
| Token-matched 10B tokens | Compute-matched FLOPs | |||||||||
| Suite | LIFT | Trans- former | T 2 MLR (ours) | T 2 MLR (HF) | Multi- pass | LIFT | Trans- former | T 2 MLR (ours) | T 2 MLR (HF) | Multi- pass |
| Core 8 (cloze) | 42.0 | 40.7 | 42.5 | 40.7 | 41.9 | 43.8 | 44.1 | 42.5 | 40.7 | 43.4 |
| MMLU (cloze) | 27.4 | 27.5 | 27.2 | 26.5 | 26.9 | 27.9 | 28.3 | 27.2 | 26.5 | 27.9 |
| Basic skills (cloze) | 40.4 | 39.8 | 39.6 | 38.7 | 40.1 | 42.1 | 42.9 | 39.6 | 38.7 | 41.4 |
| Generative QA | 5.2 | 4.5 | 5.7 | 3.8 | 4.6 | 8.0 | 7.8 | 5.7 | 3.8 | 6.6 |
| TriviaQA | 4.1 | 4.2 | 4.0 | 4.2 | 3.6 | 3.9 | 3.7 | 4.0 | 4.2 | 4.2 |
| Model | Size | Tokens | Teacher (size@tokens) | Teacher PPL | PPL | Arith. | MC % |
| LIFT | 135M | 13.4B | 135M@29.6B | 29.80 | 30.65 | 2.29 | 39.4 |
| LIFT | 135M | 13.4B | 350M@38.4B | 24.91 | 30.40 | 2.29 | 39.5 |
| LIFT w/o states | 135M | 13.4B | 135M@29.6B | 29.80 | 32.44 | 2.42 | 38.7 |
| LIFT w/o states | 135M | 13.4B | 350M@38.4B | 24.91 | 32.40 | 2.48 | 39.0 |
| Transformer, token-match | 135M | 13.4B | – | – | 32.43 | 2.46 | 38.1 |
| LIFT | 350M | 34.9B | 350M@77.8B | 23.49 | 23.82 | 2.12 | 45.0 |
| Component | FLOPs / token | B (MFLOPs) | % of a step | Scaling |
| Fusion block, Eq. ( 5 ), forward + backward | , falls with depth | |||
| Soft-token sum , forward + backward | ||||
| Forward KL on the top- support + tail, forward + backward | ||||
| Prefix state dropout, learned bias | – | |||
| Trunk total, per training token | ||||
| Adaptation phase: extra no-gradient forwards on average | of those steps | fixed by the schedule |
| Model | Micro-batch / GPU | Peak memory |
|---|---|---|
| LIFT | 8 | 35.4 GB |
| T 2 MLR(13,18) | 8 | 44.7 GB |
| Multi-pass Transformer | 8 | 62.9 GB |
| LIFT | 4 | 18.5 GB |
| T 2 MLR(13,18) | 4 | 23.0 GB |
| Multi-pass Transformer | 4 | 32.1 GB |
| Method | Fed-back signal | Enters at | Token | Training-time states | Stage |
| Feedback Transformer ( Fan et al., 2021 ) | All layers’ states of past positions | Attention of every layer | kept | own, sequential | pretraining |
| RMT ( Bulatov et al., 2022 ) | Output memory tokens of the previous segment | Input of the next segment | kept | own, sequential over segments | pretraining |
| LCKV ( Wu & Tu, 2024 ) | Top-layer keys and values | Attention of every layer | kept | own, iterative parallel passes | pretraining |
| Turbo Connection ( Tang & Lu, 2026 ) | Higher-layer states of the previous token | Lower layers | kept | own, sequential in groups | fine-tuning |
| PonderLM ( Zeng et al., 2026b ) | Top- soft token of its own prediction | Input of the same position | kept | own, extra passes per position | pretraining |
| PonderLM-2 ( Zeng et al., 2026a ) | Last hidden state | An added input position | kept | own, Jacobi passes | pretraining |