cs.LGFeb 5, 2026

TIDE: Temporal Incremental Draft Engine for Self-Improving LLM Inference

Authors: Jiyoung Park, Hankyu Jang, Changseok Song, Wookeun Jung

Organizations: Moreh, Inc.

Abstract

Speculative decoding can substantially accelerate LLM inference, but realizing its benefits in practice is challenging due to evolving workloads. We present TIDE (Temporal Incremental Draft Engine), a serving-engine-native framework that integrates online draft adaptation directly into high-performance LLM inference systems. TIDE reuses target model's intermediate hidden states generated during inference as training signals for draft adaptation, thereby avoiding additional target model computation and serving-time overhead. It employs adaptive runtime control to activate speculation and draft model training only when beneficial. TIDE exploits heterogeneous clusters by mapping inference and training to appropriate GPU classes. Across diverse real-world workloads, TIDE achieves up to 1.66×\times throughput over no-speculation baselines while recovering performance on misaligned workloads where static draft models degrade throughput. TIDE also reduces training time by up to 3.02×\times and storage requirements by 24×\times compared to existing draft training approaches, and improves system throughput by up to 1.22×\times on heterogeneous GPU clusters.

Figures & tables

Explore similar work

CardsList
  1. MineDraft: A Framework for Batch Parallel Speculative Decoding

    Feb 24, 2026Zhenwei Tang, Arun Verma, Zijian Zhou +4Speculative DecodingLLM Inference Optimization

  2. Test-Time Speculation

    May 10, 2026Avinash Kumar, Sujay Sanghavi, Poulami DasSpeculative DecodingSelf-Speculative

  3. Draft-OPD: On-Policy Distillation for Speculative Draft Models

    May 28, 2026Haodi Lei, Yafu Li, Haoran Zhang +8Speculative DecodingDrafts