Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distribution of the backbone is preserved. Every open MTP-family release (MiMo-7B, DeepSeek-V3, Qwen3) trains its heads jointly with the backbone over the full pretraining run of tens of trillions of tokens, thus setting the drafter quality at pretraining time. We ask whether a lightweight post-training pass on target-generated chain-of-thought is enough to reach the same expected throughput speedup on a frozen reasoning model, and study how a serving-time system built on such a checkpoint can be optimized. We present three findings. 1) On a frozen Qwen3-8B with K=3 chained MTP heads, we show that a post-training recipe with plain cross-entropy on ≈2.5B tokens reaches or exceeds the expected speedup of jointly trained MiMo-7B on math, coding and knowledge benchmarks. Our post-training recipe utilizes 103-104× less MTP-training tokens as compared with joint pre-training of MiMO-7B MTP baseline. 2) We propose a chain-aware relaxation of draft token verification rule that allows a bounded drift from backbone language model token distribution. We show that this relaxation lifts expected speedups by +12 to +16% per benchmark while preserving task accuracy. 3) We propose an adaptive controller that dynamically chooses the number of MTP heads to be engaged at inference time and demonstrate recovery of upto 11--14% loss in speedup using fixed maximum MTP draft length.
Figures & tables
MAL
H1/H2/H3 (%)
Benchmark
Ours
MiMo-7B
Ratio
Ours
MiMo-7B
Math reasoning
GSM8K ( Cobbe et al., 2021 )
3.05
2.80
1.09×
86.5/67.4/51.1
87.1/59.1/33.6
MATH500 ( Lightman et al., 2023 )
3.00
2.87
1.05×
86.1/65.6/48.3
89.1/62.4/35.3
AIME24 ( MAA, 2024 )
2.81
2.85
0.99×
82.9/58.5/39.4
88.9/61.7/34.3
AIME25 ( MAA, 2025 )
2.81
2.82
1.00×
83.1/58.6/39.3
88.3/60.6/33.1
Table 1: Mean accepted length and per-head acceptance rates ( H1/H2/H3 , %) at matched per-head capacity. Our MAL generally matches or exceeds MiMo on all benchmarks. Task accuracy is preserved by exact verification.
Rule (deterministic Q=1 )
GSM8K
MATH500
AIME24
AIME25
Avg
Standard (exact)
3.05
3.00
2.81
2.81
2.91
Argmax bypass
3.13
3.09
2.86
2.85
2.98
Typical acceptance ( Cai et al., 2024 ; Meister et al., 2022 )
3.40
3.34
3.14
3.19
3.27
BiLD KL threshold ( Kim et al., 2023 )
3.42
3.34
3.11
unsafe
—
DistillSpec β ( Zhou et al., 2024 )
3.24
3.19
3.00
3.01
3.11
Cactus λ ( Hao and Mou, 2026 )
3.33
3.24
3.15
2.99
3.18
Table 2: Chain-aware approximate verification shows greatest improvement in expected throughput speedup (MAL). For all cells with valid entry, the task accuracy was preserved despite lenient verification.
Observed throughput ratio to fixed K=3 (avg over 10 workloads)
bs 4
bs 16
bs 64
bs 128
Fixed K=6
1.01
1.01
0.86
0.89
Adaptive- K controller
0.94
0.94
0.98
0.99
Table 3: Adaptive controller selects the right MTP depth dynamically for a workload. Table shows averaged speedup ratio w.r.t fixed MTP depth of K=3 across ten workloads.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Variant
Change
GSM8K
MATH500
AIME24
AIME25
Effect of self-distillation
Base recipe
Self-distilled CoT + frozen + CE
3.05
3.00
2.78
2.81
Raw AM-Thinking
Same recipe, external text
2.64
2.67
2.59
2.57
Extensions trained on Row 2’s raw corpus
+ GRU gate
Cross-head routing
2.72
2.50
2.48
2.46
+ Mamba gate
Cross-head routing
3.07
2.97
2.70
2.71
Appendix
Table 4: Ablations. MAL on math reasoning tasks shows most improvement with cross-entropy loss and using self-distilled tokens from Qwen3-8B model. We set K=3 chained heads, exact verification, batch size 4 . Row 1 vs. Row 2 isolates self-distillation; the lower block is extensions trained on the same raw AM-Thinking corpus as Row 2 and should be compared against Row 2, not against the self-distilled base. Mamba gate chaining produces the most improvement in ablation settings over Row 2.
τ
GSM8K ( n=200 )
MATH500 ( n=100 )
AIME24 ( n=30 )
AIME25 ( n=30 )
MAL
acc%
MAL
acc%
MAL
acc%
MAL
acc%
10−15 (ceiling)
3.85
60.5
3.93
32.0
3.86
0.0
3.89
0.0
10−6
3.44
90.5
3.07
67.0
2.94
23.3
2.98
16.7
10−3
3.43
97.0
3.38
81.0
3.25
80.0
3.20
60.0
10−2
3.43
96.0
3.36
79.0
3.14
66.7
3.18
63.3
0.10
3.42
96.5
3.37
80.0
3.10
63.3
3.17
63.3
Appendix
Table 5: Progressive-joint τ sweep on the four math benchmarks (K = 3 checkpoint, batch 4 , thinking mode). τ=10−15 is the accept-everything ceiling; higher τ approaches strict verification. Task accuracy stays at the exact baseline for τ≥10−3 and collapses below; MAL is non-monotone in τ (see text).
τ
MAL
R-1
R-2
R-L
tok/s
10−15 (ceiling)
3.61
9.57
3.08
6.88
245
10−6
2.90
12.21
4.08
8.51
208
10−3
2.60
13.75
4.58
9.52
207
10−2
2.59
13.81
4.61
9.55
208
0.10
2.54
13.81
4.56
9.56
204
0.30
2.48
13.87
4.58
9.69
197
Appendix
Table 6: τ sweep on CNN/DailyMail ( n=500 , K = 3, thinking) with ROUGE-1/2/L (%) against reference highlights. ROUGE tracks quality smoothly, unlike math accuracy.
Figure 1: ROUGE vs. MAL on CNN/DailyMail as τ varies over the eight decades in Table 6 . Curves are nearly flat across the safe corridor and drop toward the ceiling. Contrast with the step-function behaviour of math accuracy in Table 5 : n -gram overlap tolerates local drift; exact final-answer grading does not.
Arm ratio to fixed K=3
bs 4
bs 16
bs 64
bs 128
Fixed K=6
1.01
1.01
0.86
0.89
Adaptive bandit (online)
0.93
0.94
0.98
0.99
Appendix
Table 7: Bandit vs. fixed- K throughput averaged across ten workloads (paired per-sample-median tok/s). Fixed K=6 crashes at high batch; the bandit reaches ≥97% of fixed K=3 there by falling back correctly.
Workload
bs 4
bs 16
bs 64
bs 128
GSM8K
+2 % | 90 %
+3 % | 91 %
−23 % | 96 %
−21 % | 102 %
MATH500
+3 % | 89 %
+4 % | 89 %
−22 % | 96 %
−23 % | 98 %
AIME 2024
−5 % | 90 %
−3 % | 91 %
−1 % | 96 %
−2 % | 96 %
AIME 2025
+1 % | 92 %
+1 % | 93 %
+1 % | 97 %
−0 % | 98 %
HumanEval
+9 % | 86 %
+5 % | 92 %
−4 % | 104 %
+35 % | 74 %
MBPP
−0 % | 96 %
+0 % | 96 %
−8 % | 102 %
−4 % | 94 %
Appendix
Table 8: Per-cell view. Cell format ‘K=6 lift over K=3 | bandit % of best-fixed’. K=6 helps at low batch on math/code and hurts at high batch by 20 – 29% on math/MCQ; the bandit stays close to whichever fixed is best. Mixes: Math (mixed difficulty) draws 120/96/20/20 from GSM8K/MATH500/AIME 2024/2025; Cross-domain draws 96/80/80 from GSM8K/MBPP/MMLU-Redux.
Workload
MAL
H1/H2/H3/H4/H5/H6 (%)
GSM8K
3.91
91/78/65/31/18/10
MATH500
4.05
92/80/68/35/21/13
AIME 2024
3.60
89/72/56/24/13/7
AIME 2025
3.68
89/73/57/27/14/8
Math (mixed difficulty)
3.94
91/78/64/32/18/10
HumanEval
3.90
90/74/58/36/20/12
Appendix
Table 9: K = 6 checkpoint per-head acceptance ( H1 – H6 , %) and MAL at batch 4 across ten workloads (arm fixed_k6 of the bandit sweep).