Universality and Generalization of Causal Transformers Across Context Lengths
Organizations: Doshisha University, RIKEN AIP · Rice University · CNRS, ENS, PSL Université
Abstract
Long contexts are central to modern transformer systems, but most expressivity results choose a different network for each fixed sequence length. We study whether one masked transformer can approximate causal token-to-token maps uniformly over sequences of arbitrary length sampling a fixed normalized horizon. To relate sampling resolutions, we model tokens by -Hölder sequences or, more generally, a common modulus of continuity. Our notion of continuity across resolutions characterizes the causal families admitting uniform approximation on these compact input classes by a single transformer with length-independent parameters. The result extends to the infinite-length mean-field limit, where tokens form continuous curves and masked attention becomes a causal time integral. For bounded regression with target maps satisfying a -smooth stability condition defined using regular test functions, quantitative approximation yields a generalization bound: exact empirical risk minimization over suitably sized bounded-weight transformers gives root mean-square prediction error from iid labeled sequences. The bound holds at fixed confidence on the same sampling distribution, with the token dimension and no maximum-length factor. Finally, experiments on physical time series support the Hölder-regular token model at observed scales, with dataset-dependent fitted exponents, whereas text input embeddings provide a contrasting case. Native and dense sampling, shuffled controls, and refinement checks delimit this empirical regularity regime.
Figures & tables
| Domain | Input/dense | Native | Shuffled | |
|---|---|---|---|---|
| Jena weather | 8 | |||
| Beijing PM 2.5 | 6 | – | ||
| Appliances | 4 | |||
| Road traffic | 3 | – |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| data type | representation | embedding map | |
|---|---|---|---|
| Scalar series | Chronos-Bolt-tiny | 256 | Standardized 16-value patch plus 16 observation-mask entries, pretrained residual ReLU MLP; content only, evaluated at stride 16 and diagnostic stride one. |
| NLP | BigBird-RoBERTa | 768 | Learned word and token-type embeddings, linearly interpolated learned absolute position, pretrained input LayerNorm. |
| domain | dense stride 1 | native stride 16 | lags |
|---|---|---|---|
| Jena weather | 7 | ||
| Beijing PM2.5 | – | – | 2 |
| Building appliances | 4 | ||
| Road traffic | – | – | 2 |
| Transformer oil temp. | 4 | ||
| Transformer load | 4 |
| domain | ||||||
|---|---|---|---|---|---|---|
| Jena weather | 32,769 | 1 | 8 | .89 | .421 | |
| Beijing PM2.5 | 1,025 | 1 | 6 | .73 | .365 | |
| Building appliances | 4,097 | 1 | 4 | .61 | .101 | |
| Road traffic | 1,025 | 1 | 3 | .60 | .178 | |
| Transformer oil temp. | 4,097 | 1 | 4 | .85 | .291 | |
| Transformer load | 4,097 | 1 | 4 | .77 | .176 |
| Source | Shuffled | ||
|---|---|---|---|
| WikiText-103 | – | ||
| AG News | – | ||
| IMDb | – | ||
| L a T e X control | – | ||
| Shakespeare drama | |||
| KJV Bible |
| BigBird (12 bidirectional blocks) | |||
|---|---|---|---|
| Source | Input | Block 8 | Final |
| WikiText-103 | |||
| AG News | |||
| IMDb | |||
| L a T e X control | |||
| DistilGPT2 (6 causal blocks) | |||