cs.LGSep 28, 2026

Shallow Queries, Mature Values: Depth-Asynchronous Self-Speculation for Looped Transformers

Authors: Guanghao Li, Zihan Su, Hao Yu, Jinyang Jiang, Tao Ren, Zehao Li, Feng Lu, Ming Tang, +1 more

Organizations: Tsinghua University · Southern University of Science and Technology · The University of Hong Kong · Peking University · Shenzhen University of Advanced Technology

Abstract

Looped Transformers reuse a shared block across recurrent depths, making autoregressive decoding expensive because every generated token requires many sequential recurrent passes. Self-speculative decoders reduce this cost by drafting at an early depth and verifying at full depth, but typically bind draft computation to prefix representations from the same recurrent depth. We find that queries and keys approach their final-depth representations earlier than values, and controlled prefix-channel interventions show that mature values substantially improve shallow draft predictions. Motivated by this asymmetry, we introduce Depth-Asynchronous Self-Speculation (DAS), which decouples the depth of draft computation from the depth of verified-prefix representations it reads. Its Mature-V primitive lets shallow queries retrieve full-depth prefix values without additional recurrent computation. We further develop DAS-Wave, which combines depth-asynchronous prefix reads with carried parallel refinement, progressive block growth, and an independent full-depth verifier. Across four recurrent-model checkpoints and mathematics and code workloads, DAS-Wave achieves 4.00--6.96×\times mean throughput speedup over paired full-depth autoregressive decoding in the same inference stack. These results identify prefix-information depth as an effective design axis for recurrent self-speculation.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers

    Sep 15, 2026SangLyul Cho, Langqing Cui, Sehoon Kim +2Speculative DecodingTransformer Architectures

  2. ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference

    Sep 16, 2026Amir Ziashahabi, Hossein Entezari Zarch, Lei Gao +2Speculative DecodingLLM Inference Optimization

  3. Approximate Speculative Decoding

    Aug 4, 2026Yuannuo Feng, Zegang Peng, Yuxin Xie +5Speculative DecodingTruncation