Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discard alternative continuations. In this work, we remove this bottleneck during pretraining, introducing the LIFT (Latent Information Feedback Transformer) architecture and training method which enable LMs to propagate state across generation. We achieve this by turning recurrent-state learning into a teacher-forced prediction problem: each input token is paired with an information-dense state, derived from the next-token distribution of an off-the-shelf pretrained LM. The model, extended with a small number of additional parameters, is then trained to predict both the next token and the next state. As the input states are precomputed, pretraining remains fully parallel across positions. At inference, the model's own predicted states are fed back, with a minor computational overhead that decreases with model size. Experiments with pretrained models ranging from 135M to 1B parameters show that LIFT consistently outperforms standard Transformers and baselines on language modeling, downstream reasoning tasks, and procedural tasks under token-matched budget, while being on par with or ahead of compute-matched Transformers. Moreover, a controlled study on a state-tracking task shows that a tiny LIFT outperforms same-size Transformers trained on 8x more data, even when trained with the states of a Transformer that fails the task. Overall, we show that LMs can learn to exploit deep-to-shallow feedback during pretraining via scalable teacher supervision.
Figures & tables
Figure 1: We remove the information flow bottleneck in Transformers (left) by introducing the LIFT architecture and pretraining approach (middle), which allows state propagation at inference (right).
Figure 2: S5 accuracy at increasing task lengths N (linear for N≤16 , logarithmic for N>16 ), of Transformers (solid lines) versus LIFT models (dashed lines). Standard error is over seeds.
Model
Tokens
PPL ↓
Arith. BPB ↓
MC acc. %
Gen. F1/EM %
RC BPB ↓
135M
LIFT
13.4B
30.65
2.29
39.4
15.5
1.078
LIFT w/o states
13.4B
32.44
2.42
38.7
12.0
1.105
Transformer, token-match
13.4B
32.43
2.46
38.1
13.6
1.110
Transformer, compute-match
19.3B
31.04
2.50
39.5
14.1
1.080
Distillation
13.4B
31.77
2.36
39.1
12.6
1.089
350M
LIFT
34.9B
23.82
2.12
45.0
23.2
0.921
Table 1: Performance of LIFT LMs against non-feedback Transformer baselines at three scales.
Figure 3: LIFT performance on arithmetic and language modeling. (a) LIFT shows a consistent advantage on arithmetic that persists as task complexity increases, outperforming a baseline trained on 2× more tokens. (b) Throughout training, LIFT is more token-efficient than a vanilla Transformer, and the advantage grows with training: each horizontal segment marks the additional tokens the Transformer needs to reach the same loss.
Token-matched 10B tokens
Compute-matched 2.3×1019 FLOPs
Model
FLOPs
PPL ↓
Arith. ↓
MC %
Gen. %
RC ↓
Tokens
PPL ↓
Arith. ↓
MC %
Gen. %
RC ↓
LIFT
1.44e19
19.42
2.18
36.7
7.8
1.705
16.2B
18.04
2.18
38.0
9.5
1.629
Transformer
1.02e19
20.30
2.35
36.2
6.7
1.761
23.0B
18.02
2.30
38.1
8.5
1.622
T 2 MLR (our run)
2.32e19
19.92
2.27
36.5
7.8
1.737
10.0B
19.92
2.27
36.5
7.8
1.737
T 2 MLR (HF)
2.35e19
21.55
2.38
35.3
6.5
1.816
10.1B
21.55
2.38
35.3
6.5
1.816
Multi-pass Transformer
1.31e19
20.40
2.25
36.3
6.5
1.773
17.8B
18.58
2.25
37.6
7.7
1.675
Table 2: Comparison with Jacobi-iteration feedback Transformers at matched tokens and matched compute. We report results for T 2 MLR using the official checkpoint ( “HF” ) and our reproduction.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
135M
350M
1B
Non-embedding parameters
134.3M
348.9M
1,073.7M
Total parameters
237.0M
490.2M
1,484.9M
Hidden dimension d
1024
1408
2048
Layers
8
11
16
Attention heads
8
11
16
Embeddings
tied
tied
untied
Appendix
Table 3: Model and training configurations. Parameter counts exclude the RMSNorm vectors; the 1B counts both embedding matrices, the smaller models one tied matrix.
Arithmetic, string operations, pattern continuation, coding, logical reasoning, common knowledge
5,967
5
cloze acc.
MC
Generative QA: CoQA, SQuAD, Jeopardy, Natural Questions, DROP
Conversational, extractive and closed-book QA
7,983 / 1,000 / 2,116 / 1,000 / 1,000
5 (CoQA 0)
F1
Gen.
TriviaQA
Closed-book trivia
7,993
5
F1
Gen.
BBH
27 reasoning tasks with chain of thought
6,511
3
exact match
Gen.
Appendix
Table 4: The downstream suite. Items: scored instances (per task for the core tasks); Shots: in-context examples; Group: the Table 1 column, or the reason a suite is reported only in the appendix. Task sources are cited below the table.
Figure 4: The same two-layer Transformer on S5 at four training budgets, darker lines for longer training. Bands show standard error over three seeds. Training beyond 100K steps ( 20× ) brings no further gain.
Source
LIFT
w/o states
Tr. token
Tr. compute
Distillation
Seqs
135M
C4
30.65
32.44
32.43
31.04
31.77
971
Common Crawl
31.55
33.44
33.55
32.01
32.63
50
Books
33.58
35.93
35.90
33.85
34.86
50
Wikipedia
19.64
20.70
20.78
19.87
20.39
50
peS2o
17.85
18.93
19.00
18.12
18.56
51
Appendix
Table 5: Perplexity by source on the OLMo 2 perplexity evaluation set: the full C4 validation slice (971 sequences of 1,024 tokens) and 10% of each other source (at least 50 sequences). Columns as in Table 1 ; lower is better, bold is the best in a row.
Δ PPL
135M
350M
1B
vs. Transformer, token-match
-1.77 [-1.85, -1.70]
-1.33 [-1.39, -1.28]
-0.99 [-1.05, -0.93]
vs. Transformer, compute-match
-0.38 [-0.44, -0.33]
-0.50 [-0.55, -0.46]
-0.36 [-0.40, -0.32]
vs. Distillation
-1.11 [-1.17, -1.06]
-0.70 [-0.74, -0.66]
-0.59 [-0.63, -0.55]
Appendix
Table 6: Evaluation-set uncertainty of the perplexity gaps of Table 1 , for the paper’s training seed: the gap of LIFT to each Transformer baseline with a paired bootstrap 95% interval over the evaluation sequences (both models scored on the same text; 5,000 resamples). This measures how much a gap depends on the held-out text, at a fixed training run; it is not seed variance, which Table 10 reports for 135M. Negative is in favor of LIFT .
Suite
LIFT 135M@29.6B
w/o states 135M@29.6B
Tr. token 13.4B
Tr. compute 19.3B
Distillation 13.4B
2 × tokens 29.6B
Core 8 (cloze)
44.0
43.3
42.5
43.9
43.9
44.7
MMLU (cloze)
28.1
27.9
27.3
28.0
27.6
28.2
Basic skills (cloze)
46.0
44.8
44.6
46.4
45.8
47.2
Generative QA
13.4
11.3
11.4
13.8
10.7
15.5
TriviaQA
13.6
13.3
13.8
14.8
14.3
14.8
BBH (CoT)
19.5
11.3
15.6
13.8
12.9
14.7
Appendix
Table 7: OLMES results at 135M for the rows of Tables 1 and 16 : the suites of the three Table 1 groups, then the suites left out of it (at chance or floor for every model, and BoolQ). Accuracy % for the multiple-choice, generative and excluded rows; bits per byte (lower is better) for the BPB rows; suites as in Table 4 . Gray: the 2 × -token teacher, outside the bold rule. Bold: best in the row.
Suite
LIFT 350M@77.8B
w/o states 350M@77.8B
Tr. token 34.9B
Tr. compute 49.9B
Distillation 34.9B
2 × tokens 77.8B
Core 8 (cloze)
51.2
49.8
48.9
50.8
50.4
51.7
MMLU (cloze)
30.2
30.1
30.3
31.0
30.6
31.1
Basic skills (cloze)
53.8
52.7
50.3
50.4
52.4
52.6
Generative QA
27.3
24.0
21.5
23.1
24.8
26.6
TriviaQA
21.9
20.6
20.1
21.7
21.4
23.6
BBH (CoT)
20.3
18.4
21.3
21.3
20.1
21.9
Appendix
Table 8: OLMES results at 350M for the rows of Tables 1 and 16 : the suites of the three Table 1 groups, then the suites left out of it (at chance or floor for every model, and BoolQ). Accuracy % for the multiple-choice, generative and excluded rows; bits per byte (lower is better) for the BPB rows; suites as in Table 4 . Gray: the 2 × -token teacher, outside the bold rule. Bold: best in the row.
Suite
LIFT 1B@222B
LIFT 1B@4001B
w/o states 1B@222B
w/o states 1B@4001B
Tr. token 107.4B
Tr. compute 152.4B
Distillation 107.4B
2 × tokens 222B
Core 8 (cloze)
59.4
58.3
58.3
56.7
58.0
59.0
58.5
60.2
MMLU (cloze)
35.4
35.1
34.7
34.3
34.4
34.9
35.0
35.6
Basic skills (cloze)
62.6
62.2
60.6
61.5
59.3
60.9
62.5
63.0
Generative QA
42.2
42.1
39.3
39.5
38.6
40.5
41.4
43.5
TriviaQA
37.3
37.6
36.1
35.7
34.9
37.9
36.7
41.2
BBH (CoT)
25.6
27.8
23.2
26.2
25.2
26.0
26.4
27.4
Appendix
Table 9: OLMES results at 1B for the rows of Tables 1 and 16 : the suites of the three Table 1 groups, then the suites left out of it (at chance or floor for every model, and BoolQ). Accuracy % for the multiple-choice, generative and excluded rows; bits per byte (lower is better) for the BPB rows; suites as in Table 4 . Gray: the 2 × -token teacher, outside the bold rule. Bold: best in the row.
Model
PPL ↓
MC %
Gen. %
RC (BPB ↓ )
Basic skills ↓
Pattern %
Pattern BPB ↓
LIFT
30.57 ± 0.08
39.5 ± 0.1
15.0 ± 0.6
1.076 ± 0.006
1.326 ± 0.018
53.2 ± 0.7
1.339 ± 0.045
LIFT w/o states
32.39 ± 0.07
38.8 ± 0.2
11.9 ± 0.5
1.105 ± 0.005
1.364 ± 0.007
53.2 ± 0.7
1.380 ± 0.046
Transformer, token-match
32.44 ± 0.02
38.4 ± 0.5
13.0 ± 0.6
1.112 ± 0.004
1.390 ± 0.011
50.7 ± 2.0
1.427 ± 0.051
Transformer, compute-match
31.04 ± 0.00
39.2 ± 0.3
13.6 ± 0.6
1.085 ± 0.005
1.348 ± 0.036
52.2 ± 0.5
1.393 ± 0.003
Distillation
31.81 ± 0.04
39.2 ± 0.4
13.0 ± 0.4
1.096 ± 0.006
1.361 ± 0.019
50.5 ± 1.4
1.373 ± 0.022
Appendix
Table 10: The 135M rows of Table 1 over three training seeds (6198 / 6199 / 6200): mean ± sd of each column. We regard a difference between two models as beyond seed noise when their mean ± sd ranges do not overlap. Pattern: the basic-skills pattern-continuation task (534 questions).
Number of operands n
Model
Tokens
Teacher (size@tokens)
2
3
4
5
6
7
8
9
10
Avg.
135M
LIFT
13.4B
135M@29.6B
2.03
2.13
2.12
2.27
2.27
2.35
2.36
2.42
2.42
2.29
LIFT
13.4B
350M@38.4B
2.23
2.21
2.17
2.26
2.25
2.32
2.35
2.38
2.41
2.29
LIFT w/o states
13.4B
135M@29.6B
2.24
2.32
2.26
2.37
2.37
2.47
2.49
2.55
2.56
2.42
LIFT w/o states
13.4B
350M@38.4B
2.33
2.28
2.24
2.44
2.45
2.58
2.59
2.63
2.64
2.48
Transformer, token-match
13.4B
–
2.02
2.28
2.29
2.42
2.45
2.55
2.57
2.58
2.59
2.46
Appendix
Table 11: Arithmetic by number of operands: bits per byte of the answer (lower is better), for the rows of Tables 1 and 16 plus the vanilla Transformer trained on 2 × the tokens (Figure 3(a) ); 1,000 expressions per operand count (200 at two operands), three-shot. Avg. is over all expressions (total bits over total answer bytes), the number reported in Table 1 . Bold: best in the column within a size.
Number of operands n
Model
Tokens
Teacher (size@tokens)
2
3
4
5
6
7
8
9
10
Avg.
135M
LIFT
13.4B
135M@29.6B
7.0
5.0
3.4
2.4
3.2
2.5
2.2
3.0
1.9
3.0
LIFT
13.4B
350M@38.4B
5.5
4.4
4.4
3.3
4.0
2.5
2.2
2.5
3.6
3.4
LIFT w/o states
13.4B
135M@29.6B
6.0
4.5
3.8
3.1
3.1
2.7
1.7
3.1
2.8
3.2
LIFT w/o states
13.4B
350M@38.4B
5.0
4.7
3.5
3.3
3.9
2.0
2.3
2.8
2.3
3.1
Transformer, token-match
13.4B
–
7.0
4.5
4.1
3.0
2.6
3.0
2.4
2.0
2.9
3.2
Appendix
Table 12: Arithmetic by number of operands: exact-match accuracy % of the teacher-forced argmax, for the expressions and rows of Table 11 . Avg. is over all expressions. Bold: best in the column within a size.
135M
350M
1B
Model
Acc. %
BPB ↓
Acc. %
BPB ↓
Acc. %
BPB ↓
LIFT
53.2
1.356
64.4
1.134
72.1
0.851
LIFT w/o states
52.6
1.383
63.7
1.121
70.6
0.965
Transformer, token-match
50.7
1.443
58.4
1.216
66.1
0.948
Transformer, compute-match
52.6
1.396
58.1
1.144
67.6
0.955
Distillation
48.9
1.392
60.7
1.153
69.7
0.890
Appendix
Table 13: Pattern continuation, the basic-skills task of OLMES (534 five-shot questions, e.g., 2 4 6 8 → 10 ): cloze accuracy % and bits per byte of the answer (BPB, lower is better), for the rows of Table 1 at the paper’s training seed, and the Transformer trained on 2 × the tokens (gray, outside the bold rule). Three-seed statistics at 135M are in Table 10 . Bold: best in column.
Model
PPL ↓
Arith. ↓
MC %
Gen. %
RC ↓
Token-matched (10B tokens)
LIFT
19.42 ± 0.02
2.182 ± 0.007
36.7 ± 0.27
7.8 ± 0.61
1.705 ± 0.023
Transformer
20.30 ± 0.03
2.345 ± 0.156
36.2 ± 0.19
6.7 ± 0.21
1.761 ± 0.002
T 2 MLR (our run)
19.92 ± 0.06
2.267 ± 0.101
36.5 ± 0.04
7.8 ± 0.94
1.737 ± 0.010
Multi-pass Transformer
20.40 ± 0.04
2.249 ± 0.064
36.3 ± 0.36
6.5 ± 0.25
1.773 ± 0.017
Compute-matched ( 2.3×1019 FLOPs)
Appendix
Table 14: The models of Table 2 over three training seeds (42 / 43 / 44): mean ± sd of each column. T 2 MLR (HF) is not listed: it is the authors’ single released checkpoint, so it has no seed statistics. We regard a difference between two models as beyond seed noise when their mean ± sd ranges do not overlap. T 2 MLR (our run) sets the compute budget, so its row is the same under both budgets. Bold: the best mean within each budget. Per-suite results of seed 42 are in Table 15 .
Token-matched 10B tokens
Compute-matched 2.3×1019 FLOPs
Suite
LIFT
Trans- former
T 2 MLR (ours)
T 2 MLR (HF)
Multi- pass
LIFT
Trans- former
T 2 MLR (ours)
T 2 MLR (HF)
Multi- pass
Core 8 (cloze)
42.0
40.7
42.5
40.7
41.9
43.8
44.1
42.5
40.7
43.4
MMLU (cloze)
27.4
27.5
27.2
26.5
26.9
27.9
28.3
27.2
26.5
27.9
Basic skills (cloze)
40.4
39.8
39.6
38.7
40.1
42.1
42.9
39.6
38.7
41.4
Generative QA
5.2
4.5
5.7
3.8
4.6
8.0
7.8
5.7
3.8
6.6
TriviaQA
4.1
4.2
4.0
4.2
3.6
3.9
3.7
4.0
4.2
4.2
Appendix
Table 15: OLMES results for the feedback-Transformer comparison (Table 2 ), every suite of the scan; the T 2 MLR columns coincide across the two budgets. Rows and units as in Table 7 . Bold: best within a budget.
Figure 5: Teacher ablation at 135M: spending LIFT ’s extra forward pass. Each bar is a model’s perplexity improvement on the full C4 validation subset over the compute-matched Transformer, which spends that compute on more tokens (the zero line; right is better): the mean over three training seeds of the per-seed difference, with whiskers of ±1 sd.
Figure 6: Channel width at 135M: improvement in perplexity over the compute-matched Transformer (the zero line; up is better) for k from 1 to 2048 , with the model’s own states fed back ( LIFT ) and for the same model run without states ( LIFT w/o states). Dashed: the token-matched Transformer. Every step in k is beyond the paired bootstrap 95% interval over evaluation sequences (Table 6 ’s protocol).
Model
Size
Tokens
Teacher (size@tokens)
Teacher PPL
PPL ↓
Arith. ↓
MC %
LIFT
135M
13.4B
135M@29.6B
29.80
30.65
2.29
39.4
LIFT
135M
13.4B
350M@38.4B
24.91
30.40
2.29
39.5
LIFT w/o states
135M
13.4B
135M@29.6B
29.80
32.44
2.42
38.7
LIFT w/o states
135M
13.4B
350M@38.4B
24.91
32.40
2.48
39.0
Transformer, token-match
135M
13.4B
–
–
32.43
2.46
38.1
LIFT
350M
34.9B
350M@77.8B
23.49
23.82
2.12
45.0
Appendix
Table 16: Teacher choice: LIFT and LIFT w/o states per teacher (size@training tokens) against the token-matched vanilla Transformer. Tokens : the training budget of every model in the row block, 5× the Chinchilla budget of its size. Teacher PPL : the teacher’s own perplexity on the same evaluation set. The results hold on every metric, and barely change, with a larger teacher or one trained much longer. Other columns as in Table 1 . Bold: best in column within a size.
Figure 7: Relative gain of LIFT with parallel prefill (blue points) by the number of refinement passes over the prompt, against LIFT ( 100% ) and LIFT w/o states ( 0% ). The token- and compute-matched Transformers are placed on the same scale; below 0% , a model is behind LIFT w/o states.
Component
FLOPs / token
1 B (MFLOPs)
% of a 6N step
Scaling
Fusion block, Eq. ( 5 ), forward + backward
66d2
277
4.3%
11d2/N=11/(16L) , falls with depth
Soft-token sum E⊤si , forward + backward
4kd
8
0.13%
k/(24Ld)
Forward KL on the top- k support + tail, forward + backward
≈5∣V∣+20k
0.5
<0.01%
≈5∣V∣/(96Ld2)
Prefix state dropout, learned bias
0
0
0
–
Trunk total, per training token
286
4.4%
Adaptation phase: 1.5 extra no-gradient forwards on average
1.5(2N+22d2+2kd)
3,366
+52% of those steps
fixed by the schedule
Appendix
Table 18: Compute added per token by LIFT at 1 B, excluding the teacher forward. The trunk is the first 90% of training; the adaptation phase is the last 10% (§ 3.2 ).
Model
Micro-batch / GPU
Peak memory
LIFT
8
35.4 GB
T 2 MLR(13,18)
8
44.7 GB
Multi-pass Transformer
8
62.9 GB
LIFT
4
18.5 GB
T 2 MLR(13,18)
4
23.0 GB
Multi-pass Transformer
4
32.1 GB
Appendix
Table 19: Peak memory per GPU of LIFT , T 2 MLR and the Multi-pass Transformer at 135M (SmolLM2 backbone, 2,048-token sequences, B200). Memory is the peak allocated by PyTorch. The Multi-pass Transformer’s peak is measured on a three-pass batch, which determines the memory needed to run its full schedule.
Method
Fed-back signal
Enters at
Token
Training-time states
Stage
Feedback Transformer ( Fan et al., 2021 )
All layers’ states of past positions
Attention of every layer
kept
own, sequential
pretraining
RMT ( Bulatov et al., 2022 )
Output memory tokens of the previous segment
Input of the next segment
kept
own, sequential over segments
pretraining
LCKV ( Wu & Tu, 2024 )
Top-layer keys and values
Attention of every layer
kept
own, iterative parallel passes
pretraining
Turbo Connection ( Tang & Lu, 2026 )
Higher-layer states of the previous token
Lower layers
kept
own, sequential in groups
fine-tuning
PonderLM ( Zeng et al., 2026b )
Top- k soft token of its own prediction
Input of the same position
kept
own, extra passes per position
pretraining
PonderLM-2 ( Zeng et al., 2026a )
Last hidden state
An added input position
kept
own, Jacobi passes
pretraining
Appendix
Table 20: Methods that feed information back across generation steps. Training-time states : how the fed-back states are obtained during training; own means generated by the model being trained. † arXiv preprint.
We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden state from the previous token as recurrent memory for the next token. Because this state is already computed during ordinary decoding, LRT introduces a cross-token, cross-layer latent pathway while preserving the standard attention mechanism, KV-cache interface, and one model forward per generated token. To pretrain this recurrence without sequentially unrolling the full sequence, we introduce interleaved parallel training: one full-sequence initialization forward constructs a shared buffer, followed by sequential refinement of disjoint position subsets with parallel computation within each subset. This provides every token with recurrent-memory-aware supervision at approximately 2x ideal token compute. Across 1.3B- and 2.1B-parameter nanochat-style backbones and a wide range of training budgets, LRT improves both BPB and CORE under matched effective compute. Additionally, LRT outperforms two-forward PonderLM-2 and matches a three-loop Transformer in BPB, while retaining one-forward-per-token decoding with 9% latency overhead over the standard Transformer.
Zeyi Huang, Xuehai He, LiLiang Ren +8
Microsoft · University of Washington · University of Wisconsin-Madison
Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth. Dense attention gives each token broad horizontal access to the past, but the vertical feedback channel between decoding steps remains narrow: only the sampled token returns to the bottom of the stack, while the top-layer hidden state is discarded. We introduce the full-bandwidth transformer, which widens this channel with latent feedback: at each decoding step, the previous top-layer hidden state is fused with the sampled token embedding through a gated linear unit and fed back as the next input. Latent feedback lets non-verbalized computation re-enter the stack with a renewed depth budget, while preserving the standard transformer architecture, KV cache, and language-modeling objective. To train full-bandwidth transformers without losing parallel teacher forcing, we use a scheduled multi-pass objective that introduces latent feedback late in pretraining and mixes a small fraction of deeper feedback passes for stability. We train 1B-parameter full-bandwidth transformers on up to 400B tokens and find that latent feedback improves validation loss, 5-shot language-model evaluation, math and coding generation, and instruction-tuned performance. With negligible per-token decoding overhead, full-bandwidth transformers match or approach standard transformers trained with roughly 1.5x more tokens, and manage to produce shorter reasoning when no off-policy templates are provided.
Xi Wang, Ziyang Cai, Zheng Zhan +5
Johns Hopkins University · Princeton University · Microsoft
In contrast to RNNs, which compress their history into a single hidden state, Transformers can attend to all past tokens directly. However, standard Transformers rely solely on the hidden state from the previous layer to represent the entire context. We show that this design creates pressure toward representation collapse and can degrade performance. To address this issue, we introduce Layer-Integrated Memory (LIMe), a lightweight extension that leverages existing key-value buffers and learns per-head, per-layer routing weights to integrate representations from previous layers. Across language modeling, synthetic reasoning, and deep architectures, LIMe improves perplexity per FLOP in the studied regimes and yields strong gains on synthetic tasks while preserving higher value-vector entropy and token separability. Finally, learned routing weights reveal systematic reuse of local and long-distance features, showing how LIMe enriches attention-time memory without increasing hidden-state size. Code is available at https://github.com/corl-team/lime.