cs.LGFeb 6, 2026

Attention-Mass Condensation for Sparse Decoding

Authors: Jorge L. Ruiz Williams

Organizations: NaNZeta LLC

Abstract

Attention-mass concentration creates an opportunity for sparse decoding, but retained mass alone does not guarantee a stable greedy decision: retrieval error, omitted value directions, and recursive decoding all matter. We formalize this distinction with an exact omitted-mass identity and a sufficient downstream margin condition, then characterize a query-dependent mean-pooled block selector. On Qwen2-0.5B, a paired fresh-selection sweep covers supports of 97--769 positions, contexts of 2K--16K, and five prefixes per context. The primary exact-match result is that none of 60 runs remains identical to dense decoding through 128 tokens. Distributional quality is distinct: for supports of at least 193, seven of nine context-support conditions have median teacher-forced continuation perplexity changes within 5% of dense, but prompt-level ranges include severe 16K outliers. All seven runs with teacher-forced match below 70% have perplexity increases above 100%; these observations come from two prefixes and suggest a warning regime, not a general threshold. The measured perplexity is teacher-forced on the dense model's own continuation, not the sparse model's free-running output. Separate retrieval and attention-mass probes illustrate why captured mass alone is not a retrieval or decision guarantee. Isolated operator timings do not establish matched-quality acceleration or end-to-end serving speed.

Explore similar work

CardsList
  1. EntmaxKV: Support-Aware Decoding for Entmax Attention

    May 20, 2026Gonçalo Duarte, Miguel Couceiro, Marcos V. TrevisoDynamic Sparse AttentionLarge Language Model Decoding