Long contexts are central to modern transformer systems, but most expressivity results choose a different network for each fixed sequence length. We study whether one masked transformer can approximate causal token-to-token maps uniformly over sequences of arbitrary length sampling a fixed normalized horizon. To relate sampling resolutions, we model tokens by α-Hölder sequences or, more generally, a common modulus of continuity. Our notion of continuity across resolutions characterizes the causal families admitting uniform approximation on these compact input classes by a single transformer with length-independent parameters. The result extends to the infinite-length mean-field limit, where tokens form continuous curves and masked attention becomes a causal time integral. For bounded regression with target maps satisfying a β-smooth stability condition defined using regular test functions, quantitative approximation yields a generalization bound: exact empirical risk minimization over suitably sized bounded-weight transformers gives root mean-square prediction error O((loglogN/logN)β/(d+2)) from N iid labeled sequences. The bound holds at fixed confidence on the same sampling distribution, with d the token dimension and no maximum-length factor. Finally, experiments on physical time series support the Hölder-regular token model at observed scales, with dataset-dependent fitted exponents, whereas text input embeddings provide a contrasting case. Native and dense sampling, shuffled controls, and refinement checks delimit this empirical regularity regime.
Figures & tables
Figure 1: Representation-dependent finite-scale regularity. Worst-increment curves against normalized lag h (horizontal axis): physical time in (a)–(c), token position in (d). Curves are normalized per window by the finest-lag increment before averaging; shading is one sample SD. This preserves each window’s slope and fixes the first point at one with zero SD. Raw-patch and shuffled controls probe patch geometry and temporal organization; L a T e X and Python are technical text controls. Dashed lines mark the fit cutoff; all panels have lower ordinate 0.85 .
Domain
m
Input/dense
Native
Shuffled
Jena weather
8
.335±.037
.167±.040
.005±.003
Beijing PM 2.5
6
.257±.085
–
.023±.016
Appliances
4
.087±.010
.020±.017
.015±.012
Road traffic
3
.303±.009
–
.035±.011
Table 1: Pretrained content comparison. Fine-band α∞ , mean ± sample SD across m windows. Physical rows use Bolt content; WikiText-103 uses BigBird input. Shuffling precedes patching. Native dashes denote insufficient scales; the text shuffled control is unmeasured.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
data type
representation
d
embedding map
Scalar series
Chronos-Bolt-tiny
256
Standardized 16-value patch plus 16 observation-mask entries, pretrained residual ReLU MLP; content only, evaluated at stride 16 and diagnostic stride one.
NLP
BigBird-RoBERTa
768
Learned word and token-type embeddings, linearly interpolated learned absolute position, pretrained input LayerNorm.
Appendix
Table 2: Pretrained front ends. Every row uses Equations ( 23 )– ( 24 ); d is the increment dimension.
domain
dense stride 1
native stride 16
lags
Jena weather
.157±.039
.167±.040
7
Beijing PM2.5
–
–
2
Building appliances
.020±.022
.020±.017
4
Road traffic
–
–
2
Transformer oil temp.
.133±.044
.175±.041
4
Transformer load
.025±.027
.043±.055
4
Appendix
Table 3: Matched physical fine band for the same Bolt map. Mean α∞± sample SD; the last column counts dyadic lags. Native cadence uses a subset of the dense starting positions.
domain
N
d
m
α∞
R2
α.99
Jena weather
32,769
1
8
.253±.040
.89
.421
Beijing PM2.5
1,025
1
6
.139±.061
.73
.365
Building appliances
4,097
1
4
.041±.020
.61
.101
Road traffic
1,025
1
3
.172±.019
.60
.178
Transformer oil temp.
4,097
1
4
.134±.027
.85
.291
Transformer load
4,097
1
4
.129±.025
.77
.176
Appendix
Table 4: Raw physical-signal worst-increment regressions. N is waveform length, d the number of measured coordinates, and m the number of windows. Slopes are mean ± sample SD on h≤1/32 ; R2 is the mean fit quality, and α.99 the mean 99th-percentile sensitivity slope.
Figure 2: Calibration of finite-band slopes. Twelve realizations per scenario, with mean ± sample SD. The ordering of fractional-Brownian roughness is visible, but the positive slope of the jump control after patch encoding prevents interpreting the regression as an asymptotic certificate.
Figure 3: Exact candidate Hölder prefactors on nested restrictions. Each thin curve is a domain’s mean ratio An(α)/A65(α) , for raw endpoints or dense Bolt content. Domains are not pooled as independent samples. The finite cadence and central subinterval limit the scope of this stability check.
Source
α∞
α2
Shuffled α∞
WikiText-103
0.0122±0.0012
0.0188±0.0006
–
AG News
0.0126±0.0018
0.0191±0.0003
–
IMDb
0.0130±0.0008
0.0180±0.0001
–
L a T e X control
0.0119±0.0012
0.0148±0.0004
–
Shakespeare drama
0.0106±0.0018
0.0173±0.0002
0.0109±0.0017
KJV Bible
0.0115±0.0005
0.0178±0.0001
0.0117±0.0024
Appendix
Table 5: The text screen does not reveal a more regular input regime. Mean ± sample SD on the fixed fine band, using identical full-input BigBird embeddings. Shuffled controls are measured for all three new sources; dashes denote unmeasured controls at this long context. The original L a T e X stream is a repeated technical control, as described above.
Figure 4: BigBird’s multiscale geometry varies with depth. All depths are overlaid on the same log–log axes, using the same eight windows per source. Curves are ratios of mean worst increments as in ( 242 ); purple and yellow mark the input and final block. The dotted line is the fixed fine-band cutoff h=1/32 . Intermediate blocks have steeper finite-band curves, but the gain does not persist to the final block.
BigBird (12 bidirectional blocks)
Source
Input
Block 8
Final
WikiText-103
.0130±.0016
.0308±.0027
.0023±.0060
AG News
.0141±.0015
.0306±.0042
.0039±.0082
IMDb
.0143±.0014
.0279±.0037
.0106±.0082
L a T e X control
.0143±.0019
.0320±.0043
.0140±.0086
DistilGPT2 (6 causal blocks)
Appendix
Table 6: Effective maximum-increment slopes across depth. Mean ± sample SD across eight matched natural-order windows per source, on h≤1/32 . BigBird’s final state is block 12; DistilGPT2’s is the output after the separate final LayerNorm. SD is descriptive, not a confidence interval: windows may share documents or repeated content.
Figure 5: Depth, token order, and the choice of increment statistic. Mean fine-band slopes at every block for BigBird and causal DistilGPT2. Solid lines use natural order, dashed lines use shuffled inputs; shading is one sample SD for natural-order windows. Top: maximum increments. Bottom: RMS increments, a weaker statistic. DistilGPT2’s +LN point applies the final normalization to block six. Architectures, tokenizers, and context lengths differ, so cross-model differences do not isolate an effect of masking.
We develop a quantitative statistical theory of transformers in the large-context regime by adopting the abstraction of contextual flow maps (CFMs): dynamical systems that evolve a distinguished token in the presence of a contextual measure across a stack of attention blocks. Within this framework, the finite-context model approximates an idealized infinite-context system in which the contextual measure is replaced by its underlying population, so that the context length n becomes a statistical resource. Exploiting the McKean--Vlasov structure of the dynamics and the classical machinery of propagation of chaos, we establish a forward bound controlling the deviation between the finite- and infinite-context CFMs uniformly along depth, and a backward bound controlling the deviation between the corresponding training trajectories uniformly across iterations of online gradient descent. Both bounds achieve the optimal Wasserstein rate n−1/d for general CFMs and parametric rate n−1/2 for a restricted class of CFMs that includes transformers as a special case. The analysis rests on a new Eulerian adjoint formulation of the loss gradient and stability estimates for the resulting forward--adjoint system, both of which may be of independent interest.
We explore the expressive power of Transformers by establishing precise approximation error upper and lower bounds for Hölder class. Specifically, a new approximation upper bound is derived for the standard Transformer architecture equipped with Softmax operators, ReLU activation functions, and residual connections. We prove that a Transformer network composed of at most O(ε−d0/α) blocks can approximate any bounded Hölder function with d0-dimensional input and smoothness α∈(0,1] under any accuracy ε>0. In the case of approximation lower bounds, leveraging the VC-dimension upper bound, we are the first to rigorously prove that Transformers demand for at least Ω(ε−d0/(4α)) blocks to achieve the ε approximation accuracy. As a final step, we extend the derived results for standard Transformers to a general regression task and establish the corresponding excess risk rates demonstrating Transformers' empirical effectiveness in real-world settings.
Xin He, Yuling Jiao, Xiliang Lu +1
School of Mathematics and Statistics, Wuhan University, Wuhan, China · School of Artificial Intelligence, Wuhan University, Wuhan, China · Hubei Key Laboratory of Computational Science, Wuhan University, Wuhan, China
Recent advancements in transformer length generalization theory enable us to reliably predict when a transformer can learn to solve a task. In particular, the C-RASP hypothesis (a formalized version of the so-called RASP-l conjecture) posits that transformers length-generalize on a task if and only if a solution is expressible in the C-RASP language. While this hypothesis has strong empirical validation, theoretical problems arise from the fact that no computable length generalization bounds exist for C-RASP, alongside the discovery of seemingly contradictory experiments. To address these problems, we refine the C-RASP hypothesis utilizing the recently-proposed fragments C-RASP+ and C-RASP1. These fragments have computable length generalization bounds, though in the worst case requiring an extremely large (double exponential) sample size. It is an open question whether these sample size bounds are tight. In this paper, we resolve this open question by providing an exponentially tighter bound. In doing so, we show a polynomial length generalization bound for transformers if we adopt compressed strings, via a novel connection to power words. As an application, we show how this yields a fine-grained analysis of the C-RASP conjecture that resolves contradicting experimental evidence against it.
Georg Zetzsche, Hongjian Jiang, Andy Yang +4
Max Planck Institute for Software Systems (MPI-SWS) · RPTU Kaiserslautern-Landau · University of Notre Dame