HAWK: Rethinking Multimodal Drafting for Speculative Decoding
Organizations: University of California, Los Angeles · Samsung Research America
Abstract
Speculative decoding has achieved substantial lossless speedups for LLMs, but remains less effective for large vision-language models (LVLMs), where lightweight drafters struggle to use rich multimodal information. A second limitation is that standard distillation supervises the drafter only along the original training trajectory, without modeling how target predictions shift after the drafter's own proposals. As drafting moves away from this trajectory, the drafter can increasingly disagree with the target, reducing acceptance in later steps. We propose HAWK to address both limitations. HAWK uses representation similarity to select informative target layers and learns how to combine their hidden states. For visual information, it directly provides the drafter with compressed visual hidden states from the target model instead of raw visual tokens, making the visual information easier for a shallow drafter to use. HAWK also trains the drafter to capture how target predictions change after its own proposals, improving its agreement with the target during multi-step drafting. On SmolVLM-256M across ten multimodal benchmarks, HAWK raises average acceptance length from 3.32 to 4.08 and speedup from 2.19x to 2.60x over EAGLE-3 under greedy decoding, and from 2.89 to 3.41 and 1.92x to 2.19x under sampling.
Figures & tables
| Setting | SEED-Bench | MM-Vet | MathVista | SQA | GQA | DocVQA | Avg. |
| Acceptance length | |||||||
| No pooling | 4.03 | 3.15 | 2.85 | 2.36 | 3.80 | 2.08 | 3.05 |
| pooling | 4.03 | 3.25 | 2.92 | 2.45 | 3.86 | 2.09 | 3.10 |
| pooling | 4.00 | 3.20 | 2.75 | 2.43 | 3.82 | 2.03 | 3.04 |
| Speedup ratio | |||||||
| No pooling | 2.36 | 2.04 | 2.18 | 1.74 | 2.28 | 1.79 | 2.07 |
| Setting | SEED-Bench | MM-Vet | MathVista | SQA | GQA | DocVQA | Avg. |
| Acceptance length | |||||||
| No auxiliary loss | 4.03 | 3.15 | 2.85 | 2.36 | 3.80 | 2.08 | 3.05 |
| Direct distillation | 4.00 | 3.21 | 2.84 | 2.56 | 3.85 | 2.14 | 3.10 |
| Shift, | 3.94 | 3.17 | 2.98 | 2.60 | 3.79 | 2.16 | 3.11 |
| HAWK | 4.01 | 3.29 | 2.97 | 2.68 | 3.88 | 2.16 | 3.17 |
| Speedup ratio | |||||||
| Setting | SEED-Bench | MM-Vet | MathVista | SQA | GQA | DocVQA | Avg. |
| Acceptance length | |||||||
| EAGLE-3 layers (reference) | 4.03 | 3.15 | 2.85 | 2.36 | 3.80 | 2.08 | 3.05 |
| Uniform layers, learned mix | 4.16 | 3.25 | 2.86 | 2.65 | 3.92 | 2.19 | 3.17 |
| CKA layers, uniform mix | 4.07 | 3.29 | 3.09 | 2.66 | 3.93 | 2.21 | 3.21 |
| CKA layers, learned mix | 4.18 | 3.34 | 3.14 | 2.67 | 3.90 | 2.22 | 3.24 |
| Speedup ratio | |||||||
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Drafter | Early / Middle / Late | Used in |
| EAGLE-3 | / / | Table 1 ; Tables 2 – 3 |
| Hawk | / / | Table 1 |
| CKA coverage, | / / | Table 4 |
| Uniform grid | / / | Table 4 |
| EAGLE-3 (20 ep) | 0.789 | 0.603 | 0.363 | 0.248 | 0.161 |
| Hawk (20 ep) | 0.840 | 0.702 | 0.525 | 0.365 | 0.233 |
| No auxiliary loss (4 ep) | 0.767 | 0.585 | 0.344 | 0.214 | 0.136 |
| Direct distillation (4 ep) | 0.770 | 0.595 | 0.354 | 0.235 | 0.148 |
| Prediction shift (4 ep) | 0.789 | 0.612 | 0.373 | 0.248 | 0.168 |