cs.LGSep 21, 2026

Prescriptive SVD-Inspired Attention via Spectral Energy Retention

Authors: Vasileios Arampatzakis, Vasileios Sevetlidis, George Pavlidis

Organizations: Athena Research Center, Greece

Abstract

Self-attention is central to modern Transformer architectures, but its dense dot-product formulation makes it difficult to identify which internal directions are structurally important and which can be modified without disrupting the model. SVD-Inspired Attention (SVDA) addresses part of this problem by introducing a learned diagonal spectrum into the query-key score interaction, making latent attention directions explicitly inspectable through indicators such as spectral entropy, effective rank, sparsity, alignment, selectivity, and perturbation response. This paper examines the transition from diagnostic interpretation to operational intervention. A diagnosis--intervention--verification framework is proposed, and one intervention is evaluated: spectral energy retention in the attention-score pathway. Across FashionMNIST, CIFAR-10, CIFAR-100, and Food-101, the ρ=0.90ρ=0.90 prescription removes 24.5--53.7% of score directions, reduces parameters by 2.6--4.3%, and reduces estimated MACs by 2.8--5.4%. The paired mean accuracy change of the dimension-reduced model ranges from −0.03-0.03 to +0.05+0.05 percentage points over three seeds. These results support SVDA as an intrinsically interpretable attention mechanism whose learned spectrum exposes an operational coordinate system for deterministic and verifiable modification of attention-score formation.

Explore similar work

May 12, 2026cs.LG

The Routing and Filtering Structure of Attention

The attention interaction matrix QK⊤QK^{\top} contains two entangled computations: a skew-symmetric component that redistributes information between positions (routing) and a symmetric component that scales mutual relevance (filtering). We decompose 1776 heads across five pretrained transformers and find routing operating at low rank, well below the routing capacity allocated by the weight kernel. We introduce SS-DD attention as a diagnostic parameterization that disentangles routing from filtering by construction with guaranteed stability (Re(λ)≤0\mathrm{Re}(λ) \le 0) and trains stably without layer normalization. When disentangled and unnormalized, routing self-organizes into a spectral cascade, effective rank 22 at the first layer, expanding with depth across six scales from 7M to 355M parameters. The cascade predicts where attention can be simplified: linearizing the first seven layers of 125M SS-DD attention costs <5%{<}5\% perplexity, whereas standard attention collapses under the same intervention. The linearizable region widens with depth. Replacing the first four layers with ELU+1 linear attention reaches within 1.4%1.4\% of baseline at full head dimension. Cascade-allocated architectures trade attention parameters for perplexity (47%−65%47\%-65\% fewer attention parameters at +3.9%+3.9\% to +8.4%+8.4\% PPL). The routing-filtering decomposition makes the spectral budget legible; the cascade makes it actionable.
Shafayeth Jamil, Rehan Kapadia
Jul 31, 2026cs.CV

Interpretability-Guided Soft Pruning of Attention Heads in Vision Transformers

Vision foundation models, such as DINOv2, learn highly expressive representations but rely on massive, opaque architectures that demand substantial computational power and memory. To provide an interpretable-guided and efficient solution to this issue, we first propose a spectral analysis and new visualization technique for individual attention heads based on the Laplacian eigenvectors of their attention maps. Building upon recent observations regarding the block structure of Vision Transformers, we perform semantic clustering of attention heads and identify functional redundancies. Leveraging these insights, we introduce SAPER (Soft Attention PrunER), an end-to-end differentiable pruning framework based on the LapSum Soft Top-K approach. Extensive experiments on ImageNet-1K demonstrate that SAPER achieves a highly favorable accuracy-efficiency trade-off, outperforming the competitive RAPTOR baseline in FLOPs reduction while preserving strong classification performance.
Kamil Książek, Piotr Suszyński, Michał Jan Włodarczyk +2
May 6, 2026cs.LG

Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics

When a language model processes a hallucinated response, its attention routing tends to fail in one of two shapes: over-concentrating on a narrow set of positions, or spreading so diffusely that relevance is diluted, and the shape of the failure carries diagnostic signal. We study these shapes as a diagnostic characterization, computed from attention matrices under \emph{forced scoring} of benchmark-labeled responses rather than during live generation. A widely used family of spectral methods analyzes the symmetric component of the degree-normalized attention operator, which governs transport \emph{capacity}; we prove that every transpose-invariant spectral diagnostic of this operator is structurally \emph{orientation-blind} (it cannot distinguish an operator from its transpose, and therefore cannot detect information-flow direction), with a converse to the blindness theorem bounding any Lipschitz diagnostic's transpose sensitivity by the asymmetry coefficient GG. Pairing this with a closed-form bipartite-Cheeger landscape for canonical causal architectures, we show that uniform causal attention satisfies an nn-independent floor φ≥1/5φ\ge 1/5, while window attention pierces the floor as O(w/n)O(w/n); failure modes are shape-different, not just value-different. This floor is an idealized-architecture benchmark, not an empirical attractor: the fraction of real attention heads that pierce it is itself an architectural signature. The resulting two-axis diagnostic (φφ for capacity, GG for direction) yields a falsifiable polarity prediction: bottleneck- and diffuse-dominated benchmarks should exhibit opposite polarity. Under length-controlled evaluation, transport features retain interpretable signal (0.62-0.84 LC-AUROC) across the tested decoder-only, encoder-only, and encoder-decoder models, with polarity reversing as predicted between HaluEval and MedHallu.
Dominik Dahlem, Diego Maniloff, Mac Misiura