cs.LGOct 4, 2026

An equality condition for the Dobrushin bound on attention rollout and how often it holds in trained transformers

Authors: Przemysław Rola

Organizations: Department of Mathematics, Kraków University of Economics

Abstract

The Dobrushin coefficient of each attention-rollout factor satisfies κ(12(I+A))≤12(1+κ(A))κ(\frac12(I+A))\le\frac12(1+κ(A)), and multiplying these inequalities over layers bounds the coefficient of the whole rollout. We characterise exactly when the layerwise bound is tight: equality holds if and only if some token pair attaining κ(A)κ(A) is mutually self-dominant - each of the two attends to itself at least as strongly as the other attends to it. The condition is far from automatic: uniformly random stochastic matrices satisfy it only 24-30% of the time. When tested on the head-averaged attention of each individual input and restricted to content tokens - image patches, words or tabular features, excluding cls, register and separator tokens - the condition holds for essentially every input at every layer of DINOv2 (three model sizes), RoBERTa and DistilBERT. In the supervised models DeiT-B and ViT-B/16 it holds for 91% and 64% of input-layer pairs respectively, with all failures occurring late in depth. In FT-Transformer trained on two standard tabular benchmarks it holds for only 11-44% of input-layer pairs. The special tokens account for almost all failures in DINOv2 and the language models: when they are included, the condition holds for only 82-97% of input-layer pairs.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Convergent Stochastic Training of Attention and Understanding LoRA

    May 8, 2026Zhengkai Sun, Dibyakanti Kumar, Alejandro F Frangi +2Attention Layers

  2. A Controlled Study of Attention-Only Transformers

    Jul 20, 2026Henry Ndubuaku, Karen Mosoyan, Jakub Mroz +5Transformer AttentionTransformer Encoder