Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distribution of the backbone is preserved. Every open MTP-family release (MiMo-7B, DeepSeek-V3, Qwen3) trains its heads jointly with the backbone over the full pretraining run of tens of trillions of tokens, thus setting the drafter quality at pretraining time. We ask whether a lightweight post-training pass on target-generated chain-of-thought is enough to reach the same expected throughput speedup on a frozen reasoning model, and study how a serving-time system built on such a checkpoint can be optimized. We present three findings. 1) On a frozen Qwen3-8B with K=3 chained MTP heads, we show that a post-training recipe with plain cross-entropy on ≈2.5B tokens reaches or exceeds the expected speedup of jointly trained MiMo-7B on math, coding and knowledge benchmarks. Our post-training recipe utilizes 103-104× less MTP-training tokens as compared with joint pre-training of MiMO-7B MTP baseline. 2) We propose a chain-aware relaxation of draft token verification rule that allows a bounded drift from backbone language model token distribution. We show that this relaxation lifts expected speedups by +12 to +16% per benchmark while preserving task accuracy. 3) We propose an adaptive controller that dynamically chooses the number of MTP heads to be engaged at inference time and demonstrate recovery of upto 11--14% loss in speedup using fixed maximum MTP draft length.
Figures & tables
MAL
H1/H2/H3 (%)
Benchmark
Ours
MiMo-7B
Ratio
Ours
MiMo-7B
Math reasoning
GSM8K ( Cobbe et al., 2021 )
3.05
2.80
1.09×
86.5/67.4/51.1
87.1/59.1/33.6
MATH500 ( Lightman et al., 2023 )
3.00
2.87
1.05×
86.1/65.6/48.3
89.1/62.4/35.3
AIME24 ( MAA, 2024 )
2.81
2.85
0.99×
82.9/58.5/39.4
88.9/61.7/34.3
AIME25 ( MAA, 2025 )
2.81
2.82
1.00×
83.1/58.6/39.3
88.3/60.6/33.1
Table 1: Mean accepted length and per-head acceptance rates ( H1/H2/H3 , %) at matched per-head capacity. Our MAL generally matches or exceeds MiMo on all benchmarks. Task accuracy is preserved by exact verification.
Rule (deterministic Q=1 )
GSM8K
MATH500
AIME24
AIME25
Avg
Standard (exact)
3.05
3.00
2.81
2.81
2.91
Argmax bypass
3.13
3.09
2.86
2.85
2.98
Typical acceptance ( Cai et al., 2024 ; Meister et al., 2022 )
3.40
3.34
3.14
3.19
3.27
BiLD KL threshold ( Kim et al., 2023 )
3.42
3.34
3.11
unsafe
—
DistillSpec β ( Zhou et al., 2024 )
3.24
3.19
3.00
3.01
3.11
Cactus λ ( Hao and Mou, 2026 )
3.33
3.24
3.15
2.99
3.18
Table 2: Chain-aware approximate verification shows greatest improvement in expected throughput speedup (MAL). For all cells with valid entry, the task accuracy was preserved despite lenient verification.
Observed throughput ratio to fixed K=3 (avg over 10 workloads)
bs 4
bs 16
bs 64
bs 128
Fixed K=6
1.01
1.01
0.86
0.89
Adaptive- K controller
0.94
0.94
0.98
0.99
Table 3: Adaptive controller selects the right MTP depth dynamically for a workload. Table shows averaged speedup ratio w.r.t fixed MTP depth of K=3 across ten workloads.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Variant
Change
GSM8K
MATH500
AIME24
AIME25
Effect of self-distillation
Base recipe
Self-distilled CoT + frozen + CE
3.05
3.00
2.78
2.81
Raw AM-Thinking
Same recipe, external text
2.64
2.67
2.59
2.57
Extensions trained on Row 2’s raw corpus
+ GRU gate
Cross-head routing
2.72
2.50
2.48
2.46
+ Mamba gate
Cross-head routing
3.07
2.97
2.70
2.71
Appendix
Table 4: Ablations. MAL on math reasoning tasks shows most improvement with cross-entropy loss and using self-distilled tokens from Qwen3-8B model. We set K=3 chained heads, exact verification, batch size 4 . Row 1 vs. Row 2 isolates self-distillation; the lower block is extensions trained on the same raw AM-Thinking corpus as Row 2 and should be compared against Row 2, not against the self-distilled base. Mamba gate chaining produces the most improvement in ablation settings over Row 2.
τ
GSM8K ( n=200 )
MATH500 ( n=100 )
AIME24 ( n=30 )
AIME25 ( n=30 )
MAL
acc%
MAL
acc%
MAL
acc%
MAL
acc%
10−15 (ceiling)
3.85
60.5
3.93
32.0
3.86
0.0
3.89
0.0
10−6
3.44
90.5
3.07
67.0
2.94
23.3
2.98
16.7
10−3
3.43
97.0
3.38
81.0
3.25
80.0
3.20
60.0
10−2
3.43
96.0
3.36
79.0
3.14
66.7
3.18
63.3
0.10
3.42
96.5
3.37
80.0
3.10
63.3
3.17
63.3
Appendix
Table 5: Progressive-joint τ sweep on the four math benchmarks (K = 3 checkpoint, batch 4 , thinking mode). τ=10−15 is the accept-everything ceiling; higher τ approaches strict verification. Task accuracy stays at the exact baseline for τ≥10−3 and collapses below; MAL is non-monotone in τ (see text).
τ
MAL
R-1
R-2
R-L
tok/s
10−15 (ceiling)
3.61
9.57
3.08
6.88
245
10−6
2.90
12.21
4.08
8.51
208
10−3
2.60
13.75
4.58
9.52
207
10−2
2.59
13.81
4.61
9.55
208
0.10
2.54
13.81
4.56
9.56
204
0.30
2.48
13.87
4.58
9.69
197
Appendix
Table 6: τ sweep on CNN/DailyMail ( n=500 , K = 3, thinking) with ROUGE-1/2/L (%) against reference highlights. ROUGE tracks quality smoothly, unlike math accuracy.
Figure 1: ROUGE vs. MAL on CNN/DailyMail as τ varies over the eight decades in Table 6 . Curves are nearly flat across the safe corridor and drop toward the ceiling. Contrast with the step-function behaviour of math accuracy in Table 5 : n -gram overlap tolerates local drift; exact final-answer grading does not.
Arm ratio to fixed K=3
bs 4
bs 16
bs 64
bs 128
Fixed K=6
1.01
1.01
0.86
0.89
Adaptive bandit (online)
0.93
0.94
0.98
0.99
Appendix
Table 7: Bandit vs. fixed- K throughput averaged across ten workloads (paired per-sample-median tok/s). Fixed K=6 crashes at high batch; the bandit reaches ≥97% of fixed K=3 there by falling back correctly.
Workload
bs 4
bs 16
bs 64
bs 128
GSM8K
+2 % | 90 %
+3 % | 91 %
−23 % | 96 %
−21 % | 102 %
MATH500
+3 % | 89 %
+4 % | 89 %
−22 % | 96 %
−23 % | 98 %
AIME 2024
−5 % | 90 %
−3 % | 91 %
−1 % | 96 %
−2 % | 96 %
AIME 2025
+1 % | 92 %
+1 % | 93 %
+1 % | 97 %
−0 % | 98 %
HumanEval
+9 % | 86 %
+5 % | 92 %
−4 % | 104 %
+35 % | 74 %
MBPP
−0 % | 96 %
+0 % | 96 %
−8 % | 102 %
−4 % | 94 %
Appendix
Table 8: Per-cell view. Cell format ‘K=6 lift over K=3 | bandit % of best-fixed’. K=6 helps at low batch on math/code and hurts at high batch by 20 – 29% on math/MCQ; the bandit stays close to whichever fixed is best. Mixes: Math (mixed difficulty) draws 120/96/20/20 from GSM8K/MATH500/AIME 2024/2025; Cross-domain draws 96/80/80 from GSM8K/MBPP/MMLU-Redux.
Workload
MAL
H1/H2/H3/H4/H5/H6 (%)
GSM8K
3.91
91/78/65/31/18/10
MATH500
4.05
92/80/68/35/21/13
AIME 2024
3.60
89/72/56/24/13/7
AIME 2025
3.68
89/73/57/27/14/8
Math (mixed difficulty)
3.94
91/78/64/32/18/10
HumanEval
3.90
90/74/58/36/20/12
Appendix
Table 9: K = 6 checkpoint per-head acceptance ( H1 – H6 , %) and MAL at batch 4 across ten workloads (arm fixed_k6 of the bandit sweep).
Multi-token prediction has been shown to increase data density during training, improve downstream text-generation quality, and serves as the defacto approach for self-speculative decoding. Existing foundation and open source models that use MTP heads commit to a static tree-based attention topology throughout the entire generation sequence, meaning the speculation depth, and thus the compute required during verification, stays constant regardless of the context. This is fundamentally misaligned with the entropy patterns of natural language where low-entropy regions often support reliable multi-step drafting, while high-entropy regions require more conservative speculation. To address this, we propose Entropy-guided Multi-Token Prediction (EntMTP), a training-free scheduler that toggles between tree-based attention topologies from a set of task-specific pareto-optimal trees conditioned on a running estimate of local generation entropy. By matching speculation depth to context predictability, EntMTP maximizes expected accepted-token throughput across the full distribution of generated text without sacrificing generation quality. When evaluated across Humaneval, ShareGPT, GSM8k, and Litbench benchmarks, EntMTP consistently achieves a 1.15x speedup against Hydra and peak speedup of 1.36x against Medusa baselines respectively.
Large language model inference is bottlenecked by autoregressive decoding, where each token requires a full forward pass. Multi-token prediction (MTP) offers a promising acceleration path, but existing approaches suffer from a fundamental architectural flaw: the MTP head for the first token competes with the backbone's own language model (LM) head, leading to severe quality degradation when predictions are accepted. We identify this head-backbone competition as the root cause of repetitive and incoherent outputs in prior MTP-based acceleration methods. To address this, we propose Backbone-as-Architect, a design principle where the backbone LM head always generates the first token, and MTP heads are responsible only for subsequent tokens. Building on this principle, we introduce CLP (Collocation-Length Predictor), a lightweight span-level decision layer that predicts how many additional tokens can be safely accepted at each decoding step. CLP uses only a single linear layer (4.6K--7.7K parameters), replacing the over-engineered 1M-parameter gate networks used in prior work. Experiments on Qwen2.5 models (0.5B, 1.5B, 7B) show that CLP achieves 1.20x--1.29x speedup on 1.5B and 1.14x--1.20x on 7B, with zero quality degradation (repetition ratio < 0.02), while gate-based approaches fail to accelerate (1.07x) or produce severely degraded outputs (repetition ratio > 0.5%). We further demonstrate that shorter prediction horizons (k=2) recover 24% higher MTP head accuracy on large models, establishing a scaling-aware design principle. We identify MTP head prediction accuracy as the binding constraint on acceleration and establish a clear roadmap for future improvements.
Multi-Token Prediction (MTP) has emerged as an effective paradigm that augments a shared Large Language Model backbone with auxiliary heads, training the model to predict several future tokens in parallel to enrich its supervision signal and accelerate inference. However, existing training frameworks adopt a rigid, fixed-length prediction horizon, disregarding the highly non-uniform information density of natural language and code. Forcing the auxiliary heads to predict across high-entropy semantic boundaries injects noisy, conflicting training signals; because these heads share the backbone's latent representations, the resulting gradients backpropagate and interfere with the model's core capabilities. We propose AdaMTP, an adaptive training paradigm that dynamically aligns the prediction horizon with the intrinsic predictability of the sequence. At its core, an entropy-based segmentation algorithm leverages the base model to detect sudden surges in uncertainty as semantic boundaries, partitioning sequences into variable-length groups. Each token is assigned an adaptive prediction depth, and a dynamically masked MTP objective suppresses the loss for predictions that cross these boundaries, attenuating the noisy gradients that degrade the backbone. Across mathematical reasoning, code generation, and general benchmarks on three backbones (Llama-3.1-8B, Qwen-2.5-7B, Gemma-3-12B), AdaMTP consistently outperforms standard MTP in both task performance and inference speedup.
Ziqiang Cui, Han Shi, Bowei He +8
City University of Hong Kong · Huawei Technologies · Mohamed bin Zayed University of Artificial Intelligence +2