cs.LGAug 22, 2026

TANGO: Treating Tokens as Operators

Authors: Joshua Nunley

Organizations: Department of Informatics Luddy School of Informatics, Computing, and Engineering Cognitive Science Program Indiana University Bloomington

Abstract

Transformers separate cross-token mixing in self-attention from token-wise transformation in feed-forward networks. We ask whether combining these operations can lower predictive loss under fixed data and parameter budgets. To do so, we introduce the Token-Aggregated Nonlinear Gating Operator (TANGO) model. TANGO computes a nonlinear feature-wise gate at each source token. Attention averages these gates for each destination. The average modulates a linear projection of the destination and forms the diagonal core of a source-conditioned linear operator. We test this proposal by comparing full-prefix and windowed TANGO with looped and untied Transformers, the Gated Attention Unit (GAU), and Fast Linear Attention with a Single Head (FLASH) on web text, Lean formal mathematics, DeepMind Mathematics, and code. The comparison uses two parameter scales, two depths, and three seeds. Checkpoints are selected on development data and evaluated on held-out test data. At matched parameters and training data, full-prefix TANGO has the lowest mean test negative log-likelihood in all 16 settings. In eight additional combinations of size and dataset, its development loss never increases as depth rises from 4 to 8 to 16, whereas the looped Transformer's loss increases in four. Full-prefix TANGO is computationally expensive because it averages wide gates over every visible source. To reduce this cost, we evaluate a variant with three narrower gated-projection sets assigned to the first, middle, and last applications. Across four FineWeb-Edu settings, this variant achieves 3.26 to 3.45 times the throughput of TANGO and 75% to 96% that of the looped Transformer. Its mean development negative log-likelihood is lower than TANGO's in three settings and 0.023 higher in the fourth, while remaining lower than both Transformer baselines in all four.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Apr 18, 2026cs.LG

Adaptive Computation Depth via Learned Token Routing in Transformers

Standard transformer architectures apply the same number of layers to every token regardless of contextual difficulty. We present Token-Selective Attention (TSA), a learned per-token gate on residual updates between consecutive transformer blocks. Each gate is a lightweight two-layer multi-layer perceptron (MLP) that produces a continuous halting probability, making the mechanism end-to-end differentiable with 1.7% parameter overhead and no changes to the base architecture. Notably, TSA learns difficulty-proportional routing without any explicit depth pressure: even at λ=0λ=0 (no depth regularisation), the task-loss gradient alone drives the router to skip 20% of token-layer operations. On character-level language modeling, TSA saved 14-23% of token-layer operations (TLOps) across Tiny-Shakespeare and enwik8 at <0.5% quality loss. At matched efficiency, TSA achieved 0.7% lower validation loss than early exit, and the learned routing transfers directly to inference-time sparse execution for real wall-clock speedup.
May 4, 2026cs.LG

Projection-Free Transformers via Gaussian Kernel Attention

Self-attention in Transformers is typically implemented as softmax(QK⊤/d)V\mathrm{softmax}(QK^\top/\sqrt{d})V, where Q=XWQQ=XW_Q, K=XWKK=XW_K, and V=XWVV=XW_V are learned linear projections of the input XX. We ask whether these learned projections are necessary, or whether they can be replaced by a simpler similarity-based diffusion operator. We introduce \textbf{Gaussian Kernel Attention} (GKA), a drop-in replacement for dot-product attention that computes token affinities directly using a Gaussian radial basis function (RBF) kernel applied to per-head token features. Each head learns only a bandwidth parameter σhσ_h, while a single output projection WOW_O preserves compatibility with the standard Transformer interface. GKA can be interpreted as normalized kernel regression over tokens, linking modern Transformer architectures to classical non-local filtering and kernel smoothing methods. We evaluate GKA in both vision and language modeling settings. For autoregressive language modeling within the \texttt{nanochat} framework, we implement causal masking and sliding-window constraints by masking and renormalizing the Gaussian kernel. At depth 20, a GKA model with 0.42×0.42\times the parameters and 0.49×0.49\times the total training FLOPs of a standard attention baseline trains stably, exhibits a near-zero train-validation gap, and demonstrates competitive behavior on standard benchmarks, albeit with higher bits-per-byte (BPB) at this compute scale. Overall, GKA provides a minimal, interpretable attention mechanism with an explicit locality scale, offering a dimension in the accuracy-efficiency trade-off for Transformer design.
May 21, 2026cs.LG

Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention

Standard transformer attention computes pairwise similarity between queries and keys, treating all tokens as equally salient regardless of their intrinsic informational content. In turbulent fluid dynamics, coherent structures -- the energetically dominant, spatially organized patterns that persist amid background chaos -- carry a disproportionate fraction of total energy and govern all transport. We propose that tokens play an analogous role in transformer attention: informationally dense positions (morphological boundaries, syntactic heads, discourse markers) concentrate spectral energy and should attract proportionally more attention than background tokens (function words, repeated patterns, low-information filler). We propose Energy-Gated Attention (EGA): a simple modification that gates value aggregation by the spectral energy of key token embeddings, computed by a single learned linear projection that discovers the dominant spectral mode of the embedding field. On TinyShakespeare, EGA achieves +0.103 validation loss improvement with only 12,480 additional parameters (<0.26% overhead) and no measurable computational cost. The result is consistent on Penn Treebank (+0.101), demonstrating dataset independence. A systematic ablation across three wavelet families (fixed Morlet, Daubechies db2/db4, and a parametric Morlet) establishes that fixed structured bases are suboptimal -- the optimal energy direction is data-adaptive and non-sinusoidal -- while identifying learned wavelet packets as a promising open direction. The learned energy threshold converges to tau ~= 0.35 independently of initialization, corresponding to the fraction (~36%) of tokens carrying above-average spectral energy in English text, a stable linguistic property consistent with the fraction of content words in running English text.