cs.LGSep 20, 2026

LumoTree: Path-Parallel Speculative Verification for Hybrid Language Models

Authors: Zhiyuan Ma

Abstract

Tree speculative decoding for hybrid language models must preserve one coherent continuation across recurrent, convolution, and attention state. We present LumoTree, a verifier that executes recurrent paths in parallel, reuses state tiles within each path, and coordinates native recurrent replay, convolution-history gathering, and attention-cache remapping through a shared logical tree. Fused candidate selection, GPU-resident acceptance, and grouped split-K attention support the verification cycle. Component experiments show exact candidate-selection parity and recurrent agreement within paired error bounds. An exploratory Qwen3.8-27B NVFP4 deployment on a single NVIDIA DGX Spark (GB10) records 25.63 pooled tokens/s on ten SWE-bench Verified Astropy tasks. The results characterize component-level numerical agreement and coding-agent deployment; complete-model continuation and controlled application speedups remain open.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Bole: Efficient Tree Speculation for Hybrid-Attention Language Models

    Aug 3, 2026Li Wang, Yi Su, Xiabao Wu +9Speculative DecodingSelf-Speculative

  2. Provably Shorter Scratchpads in Hybrid DeltaNet-Attention Decoders

    May 15, 2026Tomasz SteiferKimi Delta AttentionGated Deltanet

  3. Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding

    May 31, 2026Xin Su, Dawid Majchrowski, Fangyuan Yu +5Speculative DecodingDecoding