cs.LGSep 20, 2026

LumoTree: Path-Parallel Speculative Verification for Hybrid Language Models

Authors: Zhiyuan Ma

Abstract

Tree speculative decoding for hybrid language models must preserve one coherent continuation across recurrent, convolution, and attention state. We present LumoTree, a verifier that executes recurrent paths in parallel, reuses state tiles within each path, and coordinates native recurrent replay, convolution-history gathering, and attention-cache remapping through a shared logical tree. Fused candidate selection, GPU-resident acceptance, and grouped split-K attention support the verification cycle. Component experiments show exact candidate-selection parity and recurrent agreement within paired error bounds. An exploratory Qwen3.8-27B NVFP4 deployment on a single NVIDIA DGX Spark (GB10) records 25.63 pooled tokens/s on ten SWE-bench Verified Astropy tasks. The results characterize component-level numerical agreement and coding-agent deployment; complete-model continuation and controlled application speedups remain open.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Aug 3, 2026cs.DC

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models

Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autoregressive decoding remains memory-bound. Tree speculative decoding offers an attractive acceleration path, but existing tree-speculation systems are designed around the key--value caches of full-attention models. On hybrid models, they traverse recurrent layers branch by branch and materialize a full state for every proposal node, causing verification latency and transient memory to scale poorly with tree and batch sizes. We present Bole, a kernel--runtime co-design that enables efficient tree speculation for hybrid-attention LLMs. Bole transforms the linear-attention recurrence into a tree-structured closed form and realizes it with a resource-efficient GPU kernel, verifying all proposal nodes in parallel and accelerating linear-attention tree verification by 3.4--7.7×\times. It losslessly encodes speculative state updates as token-level factors and reconstructs only the state selected after sampling, reducing transient state memory by 82--99×\times and freeing GPU capacity for KV caches. Its integration into SGLang, a widely deployed production LLM serving engine, couples efficient state management with a batch-wide verification budget calibrated to the complete hybrid forward. Across four models, two GPU platforms, and diverse datasets, Bole delivers up to 4.72×4.72\times the offline decode throughput of autoregressive decoding and up to 2.03×2.03\times that of the strongest tree-speculative baseline. Under online agent workloads, it reduces TTFT and TPOT by up to 67.667.6% and 49.949.9%, respectively, over the strongest tree-speculative baseline.
May 15, 2026cs.LG

Provably Shorter Scratchpads in Hybrid DeltaNet-Attention Decoders

We investigate the expressive power of hybrid recurrent-attention decoders, a class of architectures used in recent open-source language models such as Qwen3-Next and its successors. These models combine Gated Attention heads with recurrent Gated DeltaNet heads. Is there a formal advantage, in terms of model expressivity or efficiency, to such a hybrid architecture? We show that there is. We define parity-conditioned retrieval task and show that under constant-precision assumption, a Qwen-style hybrid of Gated DeltaNet and Gated Attention solves this task with a constant scratchpad, or equivalently O(1)O(1) chain-of-thought steps. In contrast, no similar solution exists for pure Gated DeltaNet models, while pure Gated Attention requires at least a polynomial scratchpad.
May 31, 2026cs.CL

Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding

Large Language Model (LLM) generation remains expensive because autoregressive decoding calls the model once for each new token. Speculative decoding reduces this cost by drafting multiple tokens and verifying them with the target model in one step, but its speedup depends on how many drafted tokens are accepted. Parameter-free draft sources can propose long continuations at low cost in structured and agentic workloads, yet a cache match that looks promising at one generation step may have low payoff at the next. We propose Hybrid Verified Decoding, which predicts the accepted length of a cache draft before verification and uses this payoff estimate to choose between cache verification and a model-based drafter. Across three LLMs and sixteen datasets, Hybrid Verified Decoding is especially effective on agentic workflows, where it outperforms EAGLE3 in every setting with a 2.73x average speedup. Our analysis shows how prompt structure creates cache opportunities, how high-payoff cache drafts concentrate in a small part of the draft space, and how payoff-guided selection reduces sequential decoding work, pointing to runtime draft selection as a promising direction for speculative decoding.