Judge Decoding accelerates LLM inference by relaxing the strict verification of Speculative Decoding, yet it typically relies on expensive and noisy supervision. In this work, we revisit this paradigm from first principles, revealing that the ``criticality'' scores learned via costly supervision are intrinsically encoded in the draft-target distributional divergence. We theoretically prove a structural correspondence between learned linear judges and Kullback-Leibler (KL) divergence, demonstrating they rely on the same underlying logit primitives. Guided by this, we propose a simple, training-free verification mechanism based on KL divergence. Extensive experiments across reasoning and coding benchmarks show that our method matches or outperforms complex trained judges (e.g., AutoJudge), offering superior robustness to domain shifts and eliminating the supervision bottleneck entirely. Code is available at https://github.com/sunshy-1/JuDi
Figures & tables
Figure 1: Relationship between AutoJudge score and token-level KL divergence for token criticality in illustrative cases. Higher AutoJudge scores indicate more critical tokens, which coincide with larger KL divergence and stronger target–draft disagreement (Case 1: inconsistent key numbers). Conversely, low scores correspond to small KL divergence and minor, non-semantic deviations (Case 2: capitalization only).
Figure 2: Judge decoding as learning distributional discrepancy between draft and target models under three paradigms. (a) Manual-annotation pipeline for training critical-token judges; (b) Divergence-point mining without manual labels; (c) We show that a linear judge’s score is empirically correlated with and theoretically connected to distributional divergence (e.g., KL divergence).
Figure 3: Illustration of overextended labeling boundaries in manual annotation, adapted from Bachmann et al. (2025) . Key: Non-critical Tokens , Critical Tokens , Logic-pivoting Tokens . The span-based labeling ( red ) obscures the true logic-pivoting tokens ( red underline ), creating noisy supervision signals.
Figure 4: Efficiency analysis of AutoJudge dataset mining. The substantial growth in GPU hours and memory footprint for larger models highlights a significant scalability bottleneck.
Figure 5: Consistency analysis of critical tokens. The low percentage of consistently critical tokens (Level 4) indicates that heuristic mining is heavily influenced by generation randomness.
Figure 6: Correlation between supervised criticality scores and intrinsic model statistics on GSM8K. Samples are stratified into 10 bins based on AutoJudge scores (x-axis). Top: The mean KL divergence (red line) shows a clear upward trend as criticality increases, whereas model entropy (dashed lines) shows no significant correlation. Bottom: The sample count distribution across probability bins (i.e., AutoJudge scores).
Figure 7: Accuracy and MAT on GSM8K for Llama-3.2-1B-Instruct/Llama-3.1-8B-Instruct (left) and Llama-3.1-8B-Instruct/Llama-3.1-70B-Instruct (right).
Figure 8: Accuracy and MAT on LiveCodeBench for Llama-3.2-1B-Instruct/Llama-3.1-8B-Instruct (left) and Llama-3.1-8B-Instruct/Llama-3.1-70B-Instruct (right).
Figure 9: Accuracy and MAT on MATH-500-Hard for Llama-3.2-1B-Instruct/Llama-3.1-8B-Instruct (left) and Llama-3.1-8B-Instruct/Llama-3.1-70B-Instruct (right).
Figure 10: Accuracy and MAT on MMLU-Pro for Llama-3.2-1B-Instruct/Llama-3.1-8B-Instruct (left) and Llama-3.1-8B-Instruct/Llama-3.1-70B-Instruct (right).
Method
Metric
Llama-3.2-1B-Instruct & Llama-3.1-8B-Instruct
AutoJudge
Threshold
0.01
0.07
0.09
0.14
0.22
Acc
82.50%
80.29%
79.39%
78.40%
75.59%
Speed
39.11
45.06
47.12
48.87
53.70
Speedup
1.02 ×
1.17 ×
1.23 ×
1.27 ×
1.40 ×
KL
Threshold
0.10
0.30
0.40
0.50
0.90
Acc
82.77%
80.96%
81.07%
79.46%
77.13%
Table 1: Inference deployment results with vLLM on GSM8K for Llama - 3.2 - 1B - Instruct/Llama - 3.1 - 8B - Instruct (left) and Llama - 3.1 - 8B - Instruct/Llama - 3.1 - 70B - Instruct (right). Speed is measured in tokens/s, and Speedup is reported relative to standard speculative decoding (i.e., Vanilla SP) .
Figure 11: Accuracy and MAT on GSM8K for Qwen3-0.6B/Qwen3-8B.
Threshold
Full-vocab KL
Top-20
Top-200
Acc
MAT
Acc
MAT
ACC
MAT
Vanilla SP
87.75%
10.61
87.75%
10.61
87.75%
10.61
0.3
87.50%
13.33
86.00% ( ↓ 1.50%)
13.45
86.38% ( ↓ 1.12%)
13.35
0.6
85.50%
17.35
85.62%
17.40
86.05%
16.72
0.7
83.88%
18.94
84.00%
18.72
84.30%
18.90
0.8
83.12%
19.65
83.62%
19.93
83.49%
19.91
Table 2: Output quality and MAT under full-vocab and Top-K KL thresholding. Evaluated on GSM8K with Llama-3.2-1B-Instruct / Llama-3.1-8B-Instruct.
Figure 12: Accuracy–speedup trade-off of Token-KL vs. Step-KL on GSM8K (Llama-3.2-1B-Instruct / Llama-3.1-8B-Instruct).
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 13: Accuracy and MAT on GSM8K for Llama-3.2-1B-Instruct and Llama-3.1-8B-Instruct (with Qwen3-Max as the annotator).
Figure 14: Prompt for critical-token annotation, using Qwen3-Max as the LLM annotator.
Figure 15: Example of Qwen3-Max annotation on the GSM8K (Question ID: 5), with [ERROR _ START] and [ERROR _ END] marking critical tokens in the draft model’s erroneous solution.
Figure 16: Example of Qwen3-Max annotation on the GSM8K (Question ID: 68), with [ERROR _ START] and [ERROR _ END] marking critical tokens in the draft model’s erroneous solution.
Step
Full-vocab KL
Top-K KL
1: Top-K selection
Not needed
O(V) , branch-intensive, irregular access
2: Gather by index
Not needed
O(K) , indirect indexing, breaks coalescing
3: Log-softmax + KL
O(V) , streaming, fusible
O(K) , streaming, fusible
Appendix
Table 3: Operational comparison between full-vocab KL and Top-K KL. Full-vocab KL avoids the irregular memory access introduced by Steps 1–2 and consists entirely of streaming, fused-kernel-friendly operations.
Figure 17: Wall-clock latency (ms) of KL computation under different Top-K configurations. Full-vocab KL is 1.8 × faster than the fastest Top-K variant due to its fully streaming memory access pattern.
Vanilla SP
0.3
0.6
0.7
0.8
w/ fallback
Acc (%)
87.75
87.50
85.50
83.88
83.12
MAT
10.61
13.33
17.35
18.94
19.65
Trigger (%)
N/A
0.0
1.1
2.7
4.9
Trigger/Query
N/A
0
0.11
0.22
0.37
w/o
Acc (%)
87.75
87.60
86.62
85.25
84.38
MAT
10.61
13.45
17.21
18.75
19.80
Appendix
Table 4: Trigger frequency and performance of the Top-1 confidence fallback across KL thresholds. “Trigger (%)” denotes the percentage of token-level decisions where the fallback is activated; “Trigger / Query” denotes the average number of activations per query.
Setting
Metric
Vanilla SP
0.003
0.008
0.02
0.03
w/o fallback
Acc
87.75%
87.12%
85.12%
82.25%
80.38%
MAT
10.61
14.06
16.18
17.26
18.51
w/ fallback
Acc
87.75%
86.60%
84.12%
82.25%
80.75%
MAT
10.61
14.05
15.99
17.87
18.63
Appendix
Table 5: AutoJudge performance with and without Top-1 fallback under different KL thresholds.
δ (Acc Drop)
Acc (%)
Speed (tokens/s)
Speedup
Vanilla SP
82.56
29.45
1.00 ×
≤ 0.1%
82.77
39.58
1.03 ×
≤ 0.5%
82.00
41.91
1.09 ×
≤ 1.5%
81.07
48.35
1.26 ×
≤ 2.5%
79.46
50.96
1.33 ×
≤ 5%
77.13
61.45
1.60 ×
Appendix
Table 6: Quality-constrained speedups on vLLM with Llama-3.2-1B-Instruct/Llama-3.1-8B-Instruct.
δ (Acc Drop)
Acc (%)
Speed (tokens/s)
Speedup
Vanilla SP
93.00
24.19
1.00 ×
≤ 0.5%
92.69
27.21
1.12 ×
≤ 1.0%
92.02
31.77
1.31 ×
≤ 1.5%
91.41
32.78
1.36 ×
≤ 2.5%
90.36
34.29
1.42 ×
≤ 3.0%
90.27
35.10
1.45 ×
Appendix
Table 7: Quality-constrained speedups on vLLM with Llama-3.1-8B-Instruct/Llama-3.1-70B-Instruct.