cs.CLSep 30, 2026

Recovering Off-Policy Supervision for Speculative Decoding

Authors: Jungseob Lee, Chanjun Park, Sugyeong Eo, Hyeonseok Moon

Organizations: Korea University · Soongsil University · Yonsei University Mirae Campus · Sookmyung Women’s University

Abstract

Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off-policy token invalidates supervision for all subsequent slots in a block. Existing approaches discard these divergent slots, resulting in severe supervision loss. To resolve this problem while preserving the training corpus, we propose a rollout-based training framework that recovers full supervision through two complementary components. The first component, Anchor-Label Relabelling (ALR), replaces corpus labels with distributions from greedy target rollouts, restoring valid supervision across all predicted slots. The second component, In-Rollout Anchors (IRA), places draft blocks directly inside these rollouts to expose the drafter to target-generated context, reusing precomputed rollout features at no additional target cost. Across fixed vision-language and text corpora, our framework increases greedy accepted length by up to 36.5% over DFlash and consistently outperforms erasing baselines. Notably, a single epoch of our method surpasses the best erase schedules. After three epochs, it matches the acceptance length of training on target-regenerated responses. These results show that our framework provides an effective and compute-efficient approach for training speculative drafters on fixed corpora without modifying the original text. Code is available at https://github.com/js-lee-AI/ALR-IRA.

Explore similar work

CardsList
  1. Draft-OPD: On-Policy Distillation for Speculative Draft Models

    May 28, 2026Haodi Lei, Yafu Li, Haoran Zhang +8Speculative DecodingDrafts

  2. SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Drafting

    May 8, 2026Weijie Shi, Qiang Xu, Fan Deng +9Speculative DecodingDrafter

  3. Spec-AUF: Accept-Until-Fail Training under Train-Inference Misalignment for Masked Block Drafters

    Jul 2, 2026Tianjian Yang, Meng LiDrafterSpeculative Decoding