The Dobrushin coefficient of each attention-rollout factor satisfies κ(21(I+A))≤21(1+κ(A)), and multiplying these inequalities over layers bounds the coefficient of the whole rollout. We characterise exactly when the layerwise bound is tight: equality holds if and only if some token pair attaining κ(A) is mutually self-dominant - each of the two attends to itself at least as strongly as the other attends to it. The condition is far from automatic: uniformly random stochastic matrices satisfy it only 24-30% of the time. When tested on the head-averaged attention of each individual input and restricted to content tokens - image patches, words or tabular features, excluding cls, register and separator tokens - the condition holds for essentially every input at every layer of DINOv2 (three model sizes), RoBERTa and DistilBERT. In the supervised models DeiT-B and ViT-B/16 it holds for 91% and 64% of input-layer pairs respectively, with all failures occurring late in depth. In FT-Transformer trained on two standard tabular benchmarks it holds for only 11-44% of input-layer pairs. The special tokens account for almost all failures in DINOv2 and the language models: when they are included, the condition holds for only 82-97% of input-layer pairs.
Figures & tables
per input
averaged
model
training
content tokens
all tokens
heads
exact layers
DINOv2-S
self-supervised
6144/6144 ( 100% )
91.2%
97.2%
12/12
DINOv2-B
self-supervised
6137/6144 ( 99.9% )
83.8%
95.9%
12/12
DINOv2-L
self-supervised
12285/12288 ( 99.98% )
81.9%
91.4%
24/24
DINOv2-B, no registers
self-supervised
6141/6144 ( 99.95% )
97.4%
94.0%
12/12
DeiT-B
supervised
5582/6144 ( 90.9% )
90.7%
82.3%
11/12
Table 1: How often condition ( 8 ) holds: per input, the object of the depth bound of Biccari et al. (2026) , and for attention averaged over inputs. Content tokens : share of input–layer pairs for which the input’s head-averaged attention matrix, restricted to content tokens (M4), satisfies the condition, across 512 images or sequences per pretrained model. All tokens : the same share without restricting to content tokens. Heads : share of head–input–layer triples — one head, at one layer, on one input — for which that head’s attention matrix, restricted to content tokens, satisfies the condition, across the same 512 inputs; not computed for FT-T. Averaged : number of layers for which ⟨Aˉ(ℓ)⟩ , restricted to content tokens, satisfies the condition (Section 4.2 ). The content tokens are the 256 patches of DINOv2, the 196 patches of DeiT-B and ViT-B/16, the 126 words of the language models, and the 8 or 28 features of FT-Transformer (FT-T, Section 4.3 ), which is evaluated over three seeds with 4000 test rows each.
Figure 1: Percentage of the 512 images or sequences whose head-averaged attention matrix at that layer satisfies condition ( 8 ), layer by layer, against relative depth ℓ/L , for the pretrained models of Table 1 ( 512 images or sequences each), on content tokens (top) and on all tokens (bottom). In the top-left panel the four DINOv2 curves lie on top of one another at 100% . Appendix C gives the exact number of failing inputs at every layer where there are any.
Figure 2: Doeblin bound 1−Nε against the directly computed coefficient, per layer, for ⟨Aˉ(ℓ)⟩ on the patch subchain of DINOv2-S, -B and -L. The shaded gap is what the entry-wise route gives up.
Figure 3: Leave-one-out score of the top carrier at each layer, by token role (log scale), for ⟨Aˉ(ℓ)⟩ of DINOv2, all tokens included. Register and cls carriers mostly sit one to two orders of magnitude above patch carriers; the exception is layer 23 of DINOv2-L.
model
L
layers with reg/ cls carrier
maxℓΔ
median Δ (patch carrier)
modal carrier
DINOv2-S
12
6
0.163
8.3×10−4
reg#1 ( 4/6 )
DINOv2-B
12
7
0.135
1.7×10−3
reg#4 ( 5/7 )
DINOv2-L
24
13
0.156
1.2×10−3
reg#1 ( 8/13 )
Table 2: Leave-one-out token scores for ⟨Aˉ(ℓ)⟩ of DINOv2, all tokens included. For each token i , Δi is the drop in κ when i is excluded from the maximum, and the layer’s carrier is the token with the largest Δi . Columns: the number of layers whose carrier is a register or cls ; the largest Δ over all layers; the median Δ at layers whose carrier is a patch; and the modal carrier, the token that is the carrier most often, with the number of register-or- cls layers it carries.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
configuration
N
metric
FT-T
authors
boosting
linear
features
all
averaged
California, 3 blocks
9
RMSE
0.473 – 0.475
0.463 – 0.465
0.463
0.731
18.7%
16.1%
2/6
California, 6 blocks
9
RMSE
0.468 – 0.475
—
0.463
0.731
11.3%
9.8%
3/15
Higgs, 3 blocks
29
accuracy
0.725 – 0.728
0.721 – 0.724
0.716
0.630
44.4%
32.3%
3/6
Appendix
Table 3: FT-Transformer (FT-T), three seeds per configuration. Performance entries give the lowest and highest test value over the seeds: our runs (FT-T), the authors’ own runs with the same seeds (available for their default of three blocks), gradient boosting and a linear model. N is the number of tokens, including cls . Features and all : share of row–layer pairs for which the row’s head-averaged attention matrix satisfies condition ( 8 ), restricted to the feature tokens and on all tokens, over the 4000 test rows and three seeds. Averaged : number of tested layers, over the three seeds, for which ⟨Aˉ(ℓ)⟩ , restricted to the feature tokens, satisfies it.
model
layers
failing inputs out of 512 , by layer
content tokens
DINOv2-S
12
none
DINOv2-B
12
layer 1 : 7
DINOv2-L
24
layer 1 : 3
DINOv2-B, no registers
12
layer 1 : 2 ; layer 11 : 1
DeiT-B
12
layer 11 : 50 ; layer 12 : 512
Appendix
Table 4: Layers at which some input fails condition ( 8 ), with the number of failing inputs out of 512 , on content tokens and on all tokens. All other layers are exact for every input.
Figure 4: Per-input test on content tokens at several input sizes; N is the number of content tokens. Left: percentage of input–layer pairs whose head-averaged attention matrix satisfies condition ( 8 ). Right: median gap among the failing pairs, as a percentage of 21κ(Aˉ) , for the two models with enough failures for a median; RoBERTa is shown from N=126 .