cs.LGSep 30, 2026

Attention Kernels for Learning Maps Between Heavy-Tailed Measures

Authors: Kailen Hargenrader, Edoardo Calvello, Bohan Chen

Organizations: Computing and Mathematical Sciences California Institute of Technology · Lawrence Berkeley National Laboratory University of California, Berkeley ICSI

Abstract

Operator learning on probability measures can be accomplished with transformers. For measures with polynomial tails, the exponential weighting in softmax can make the corresponding measure-level attention integrals diverge. This motivates replacing the exponential with slower-growing functions. We construct two benchmarks for operator learning on measures with closed-form targets. We use these benchmarks to study attention kernel growth and data transformation in post-norm transformers. Without data transformation, the softmax models exhibit ensemble collapse on both heavy-tailed benchmarks, while the three slower-growing kernels avoid collapse. Symlog preprocessing allows softmax to avoid collapse on the matrix inverse task but not on the sheared swap task. On the Gaussian control, all four kernels perform similarly. We also examine how sample size affects the sensitivity of empirical energy and Wasserstein distances to tail differences. These results support slower-growing attention kernels as an effective design choice for post-norm transformers learning from heavy-tailed ensembles.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Functional Attention: From Pairwise Affinities to Functional Correspondences

    May 29, 2026Jiefang Xiao, Maolin Gao, Simon Weber +2Transformer AttentionDynamic Attention

  2. Pretraining Transformers with Quantized Softmax in Attention

    Sep 27, 2026Shangzhen Zhu, Muyan Hu, Tomasz KozlowskiTransformer AttentionGumbel-Softmax Relaxation

  3. Why Post-Norm Transformers Collapse: Attention Amplification and Gradient Repair Failure

    Aug 10, 2026Xingjian Wang, Qingyu Han, Xiaodong Luo +1Batch NormalizationTransformer Encoder