On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-constrained teacher-guided rollout with maximal coupling and reuses realized accept/correction events to route token-level supervision. Accepted positions retain sampled-token reverse-KL supervision, while correction positions receive direct supervision on the teacher's highest-probability token. Under maximal coupling, the correction probability is exactly TV(p_t, q_t), so the same trust-region radius controls rollout deviation and upper-bounds intervention and specialized-supervision frequency. We further implement an engine-resident speculative verifier that preserves the exact-q trajectory distribution and coupling semantics while improving matched-workload rollout throughput by 4.22x. Across seven mathematical reasoning benchmarks, SAKI improves the matched teacher-guided baseline in Mean@8 and Pass@8 for both 1.7B and 0.6B students. Placement controls and fixed-prefix analysis further support correction-triggered routing as a conflict-adaptive supervision signal.
Figures & tables
Figure 1: Overview of SAKI and its exact- q execution. The TRB bridge defines qt ; maximal coupling either retains a student proposal with RKL supervision or draws a residual rollout correction and routes teacher-Top-1 supervision. The residual token controls the next prefix, whereas the teacher mode controls the update. Bottom: an engine-resident K -token verifier commits proposals only through the first rejection, commits an exact residual correction there, and discards the invalid suffix.
Method
MATH-500
HMMT
AIME26
AIME25
AMC23
Minerva
Olympiad
Average
Teacher (4B)
84.1
15.2
19.2
19.6
60.9
38.0
51.9
41.3
Student (1.7B)
10.2
0.0
0.4
0.4
3.8
3.1
3.2
3.0
OPD
70.7
2.3
6.3
8.3
42.8
27.2
33.7
27.3
ExOPD
69.5
2.7
5.4
7.9
44.4
27.5
32.1
27.1
SKD
64.3
3.4
5.4
4.6
40.9
23.9
29.1
24.5
TRB
70.7
4.9
6.3
8.8
43.1
26.5
34.7
27.9
Table 1: Mean@8 accuracy (%). Average is the mean over the seven benchmarks; bold and underline denote the best and second-best trained method within each student size, respectively.
Method
MATH-500
HMMT
AIME26
AIME25
AMC23
Minerva
Olympiad
Average
Teacher (4B)
94.4
24.2
30.0
26.7
97.5
57.4
70.2
57.2
Student (1.7B)
49.6
0.0
3.3
3.3
20.0
19.9
18.4
16.4
OPD
90.0
9.1
20.0
20.0
72.5
46.3
57.2
45.0
ExOPD
87.6
6.1
20.0
20.0
75.0
45.6
56.7
44.4
SKD
86.6
6.1
16.7
13.3
67.5
44.5
53.6
41.2
TRB
90.4
9.1
16.7
23.3
70.0
44.9
57.9
44.6
Table 2: Pass@8 accuracy (%). Average is the mean over the seven benchmarks; bold and underline denote the best and second-best trained method within each student size, respectively.
Backend
Tok/s
Relative
Student-only reference
7,760
–
External-loop exact- q K8
776
1.00 ×
Engine-resident exact- q K8
3,276
4.22 ×
Table 3: Exact- q rollout throughput under the matched workload above.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Initialization
Qwen3-4B-Base
RL algorithm
GRPO
Training data
Processed DAPO-Math-17K
Training epochs
1
Prompt batch size
64
Reported micro-batch size
64
Appendix
Table 4: Training configuration used to construct the Qwen3-4B-Base-GRPO teacher , following Li et al. (2026) .
Setting
Value
Students
Qwen3-0.6B / 1.7B Base
Teacher
Qwen3-4B-Base-GRPO
Training
200 steps; AdamW; LR 1×10−6
Weight decay / grad clip
0.01 / 1.0
Rollout batch
64 prompts ×8 responses =512
Sampling
T=1.0 , top- p=1.0 , no top- k
Appendix
Table 5: Shared training configuration for the 0.6B and 1.7B students.
Method
Mean@8
Pass@8
TRB
27.90
44.60
Random-TM
28.30
45.50
TV-Weighted-TM
28.24
46.77
Ours
29.00
47.50
Appendix
Table 6: Matched supervision-placement controls for the 1.7B student on the same seven-benchmark suite. Random-TM controls for the supervision budget, while TV-Weighted-TM additionally favors positions with larger local student–guided disagreement.
Step
ΔC1
ΔC16
ΔDKL(T∥S)
ΔH(S)
50
+4.63 pp
+4.54 pp
−0.133
−0.392
100
+4.30 pp
+2.69 pp
−0.154
−0.249
200
+4.26 pp
+2.94 pp
−0.196
−0.269
Appendix
Table 7: Prompt-equal paired differences between correction-triggered teacher-mode supervision and the matched TRB control.
Correction supervision
Mean@8
Pass@8
TRB
27.90
44.60
Teacher-sampled
28.38
45.94
Teacher-mode ( SAKI )
29.00
47.50
Appendix
Table 8: Correction-supervision ablation for the 1.7B student. The coupling-aware variants use the same maximal-coupling routing rule and trust-region schedule, differing only in the supervision target used at correction positions.
Implementation
K
Tok/s
Execution
Token-wise prototype
1
∼ 5
Per-token external RPC
Early external block path
8
165–299
External block verification
Final external loop
8
776
Mature Python/RPC pipeline
Engine-resident backend
8
3,276
GPU-resident bridge/coupling
Appendix
Table 9: Engineering progression of the exact- q rollout implementation . The first two rows reflect early exploratory setups, while the final two rows benchmark mature implementations under the matched workload.
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu
Department of Computer Sciences, University of Wisconsin–Madison · Machine Learning Research, Morgan Stanley · Department of Electrical and Computer Engineering, Rutgers University–New Brunswick