cs.CLAug 19, 2026

WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing

Authors: Wenbo Zhang, Xiang Ren

Organizations: University of Southern California

Abstract

When generating text, a Transformer produces representations of past tokens at every layer, but each layer can normally use only representations from the same depth. This restriction prevents the model from fully reusing information it has already computed. We introduce WhiteMatter, which allows every layer to draw on past-token representations from any depth. A learned mixer selects the most useful depths for the current context and combines their representations into shared key-value (KV) cache channels. Sharing these channels across layers can reduce the cache size. Given the same number of training tokens, WhiteMatter with a full-size cache performs comparably to a standard Transformer with 50% more layers. With half the KV cache, WhiteMatter outperforms matched standard Transformers at two model scales, up to 1.3B parameters. Cross-layer connections, however, introduce dependencies that slow training and prompt processing. We address this problem with cyclic iteration, which updates interleaved groups of tokens in turn while processing the tokens within each group in parallel. On a reference model trained with exact autoregressive execution, cyclic iteration converges 12.5x faster than standard Jacobi iteration.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Depth-Attention: Cross-Layer Value Mixing for Language Models

    Jun 3, 2026Boyi Zeng, Yiqin Hao, Zitong Wang +7Layer-WiseCross-Layer Interactions

  2. You Do Not Fully Utilize Transformer's Representation Capacity

    Feb 13, 2025Gleb Gerasimov, Yaroslav Aksenov, Nikita Balagansky +2Transformer ArchitecturesLanguage Modeling

  3. The Recurrent Transformer: Greater Effective Depth and Efficient Decoding

    Apr 23, 2026Costin-Andrei Oncescu, Depen Morwani, Samy Jelassi +3Transformer ArchitecturesKv-Cache Management