Telescopic Language Models
Organizations: University of Cambridge · University of British Columbia · Google
Abstract
One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a nested-capacity Transformer supervised by stochastic prefix supervision with a full anchor. At every step, one randomly truncated prefix of the capacity axis is trained against the full next-token target, alongside one full-capacity pass, so the trained artifact is a valid language model at every depth. Two forward-backward passes per step, no architectural change, nothing extra at inference. Fixed-exit suites such as Matryoshka Language Model Suites (MLMS) occupy one point in this design space, and the point has a cost: supervising only a few fixed exits leaves the nested model at chance level everywhere else (perplexity 10^2-10^5 in our baselines). On a 200M proxy suite (20B FineWeb-Edu tokens, identical data stream for all methods), a single TLM run is a valid language model at every one of its twenty layer prefixes, in perplexity and on perplexity-sensitive downstream tasks, reducing the area under the quality-budget curve by 43-44% relative to the fixed-exit suites while matching them at full capacity, at ~12% lower GPU cost per run. The prefix sampling density is a dial: concentrating it on a few depths recovers fixed-exit quality there at the price of the continuum, so the operating points become a training-time choice rather than an architectural one. These results indicate that the training objective, not the nesting itself, is what makes a model elastic.
Figures & tables
| Method | 200M PPL | AULB | ||
|---|---|---|---|---|
| MLMS (best suite) | 14.98 | 5.90 | 21.8 | 27.8 |
| TLM ( uniform) | 14.99 | 3.28 | 38.9 | 47.3 |
| TLM ( log-unif.) | 14.99 | 3.45 | 34.9 | 48.5 |
| TLM ( exit-grid) | 14.65 | 6.21 | – | 25.7 |
| vanilla twin (ref.) | 13.98 | – | – | – |
| Method | GPU-h | operating points | GPU-h / point |
| vanilla suite (3 models) | 216 | 3 | 72 |
| MLMS suite, no distill (1 run) | 149 | 3 | 50 |
| MLMS suite, +distill (1 run) | 182–185 | 3 | 61–62 |
| TLM, uniform (1 run) | 131 | 20 | 6.5 |