cs.LGOct 7, 2026

The Identifiability and Observability of Deep Normalized Attention

Authors: Pranav Venkata Konda

Organizations: Columbia University

Abstract

We study which parameters of deep, unmasked, single-head attention are determined by its input--output function. For known positive nonconstant real-analytic normalizers, the function generically determines the effective scores and combined value map up to the signs induced by even normalizers. This proves the real-analytic case of a conjecture of Henry--Marchetti--Kohn, including softmax. We then classify exceptional fibers under explicit normalizer conditions, identifying when collapse makes later scores unobservable, and establish sharp Taylor orders for local identification. Near simultaneous query/key collapse, we compute the complete native Jacobian decay spectrum on separating finite input banks. For common first nonconstant normalizer degree kk, layer ii has contact order 2k3i−1−12k3^{i-1}-1, with exact multiplicities and kernel dimension. High-precision and automatic differentiation calculations illustrate the resulting loss of numerical sensitivity.

Figures & tables

Explore similar work

CardsList
  1. Rank, Head-Channel Non-Identifiability, and Symmetry Breaking: A Precise Analysis of Representational Collapse in Transformers

    Apr 26, 2026Giansalvo CirrincioneMulti-Head AttentionTransformer Architectures

  2. Geometric and Spectral Alignment for Deep Neural Network I

    May 4, 2026Ziran Liu, Wei Wang, Jinhao Wang +5Spectral NormJacobian

  3. The Routing and Filtering Structure of Attention

    May 12, 2026Shafayeth Jamil, Rehan KapadiaKimi Delta AttentionDecomposition