Judge Decoding accelerates LLM inference by relaxing the strict verification of Speculative Decoding, yet it typically relies on expensive and noisy supervision. In this work, we revisit this paradigm from first principles, revealing that the ``criticality'' scores learned via costly supervision are intrinsically encoded in the draft-target distributional divergence. We theoretically prove a structural correspondence between learned linear judges and Kullback-Leibler (KL) divergence, demonstrating they rely on the same underlying logit primitives. Guided by this, we propose a simple, training-free verification mechanism based on KL divergence. Extensive experiments across reasoning and coding benchmarks show that our method matches or outperforms complex trained judges (e.g., AutoJudge), offering superior robustness to domain shifts and eliminating the supervision bottleneck entirely. Code is available at https://github.com/sunshy-1/JuDi
Figures & tables
Figure 1: Relationship between AutoJudge score and token-level KL divergence for token criticality in illustrative cases. Higher AutoJudge scores indicate more critical tokens, which coincide with larger KL divergence and stronger target–draft disagreement (Case 1: inconsistent key numbers). Conversely, low scores correspond to small KL divergence and minor, non-semantic deviations (Case 2: capitalization only).
Figure 2: Judge decoding as learning distributional discrepancy between draft and target models under three paradigms. (a) Manual-annotation pipeline for training critical-token judges; (b) Divergence-point mining without manual labels; (c) We show that a linear judge’s score is empirically correlated with and theoretically connected to distributional divergence (e.g., KL divergence).
Figure 3: Illustration of overextended labeling boundaries in manual annotation, adapted from Bachmann et al. (2025) . Key: Non-critical Tokens , Critical Tokens , Logic-pivoting Tokens . The span-based labeling ( red ) obscures the true logic-pivoting tokens ( red underline ), creating noisy supervision signals.
Figure 4: Efficiency analysis of AutoJudge dataset mining. The substantial growth in GPU hours and memory footprint for larger models highlights a significant scalability bottleneck.
Figure 5: Consistency analysis of critical tokens. The low percentage of consistently critical tokens (Level 4) indicates that heuristic mining is heavily influenced by generation randomness.
Figure 6: Correlation between supervised criticality scores and intrinsic model statistics on GSM8K. Samples are stratified into 10 bins based on AutoJudge scores (x-axis). Top: The mean KL divergence (red line) shows a clear upward trend as criticality increases, whereas model entropy (dashed lines) shows no significant correlation. Bottom: The sample count distribution across probability bins (i.e., AutoJudge scores).
Figure 7: Accuracy and MAT on GSM8K for Llama-3.2-1B-Instruct/Llama-3.1-8B-Instruct (left) and Llama-3.1-8B-Instruct/Llama-3.1-70B-Instruct (right).
Figure 8: Accuracy and MAT on LiveCodeBench for Llama-3.2-1B-Instruct/Llama-3.1-8B-Instruct (left) and Llama-3.1-8B-Instruct/Llama-3.1-70B-Instruct (right).
Figure 9: Accuracy and MAT on MATH-500-Hard for Llama-3.2-1B-Instruct/Llama-3.1-8B-Instruct (left) and Llama-3.1-8B-Instruct/Llama-3.1-70B-Instruct (right).
Figure 10: Accuracy and MAT on MMLU-Pro for Llama-3.2-1B-Instruct/Llama-3.1-8B-Instruct (left) and Llama-3.1-8B-Instruct/Llama-3.1-70B-Instruct (right).
Method
Metric
Llama-3.2-1B-Instruct & Llama-3.1-8B-Instruct
AutoJudge
Threshold
0.01
0.07
0.09
0.14
0.22
Acc
82.50%
80.29%
79.39%
78.40%
75.59%
Speed
39.11
45.06
47.12
48.87
53.70
Speedup
1.02 ×
1.17 ×
1.23 ×
1.27 ×
1.40 ×
KL
Threshold
0.10
0.30
0.40
0.50
0.90
Acc
82.77%
80.96%
81.07%
79.46%
77.13%
Table 1: Inference deployment results with vLLM on GSM8K for Llama - 3.2 - 1B - Instruct/Llama - 3.1 - 8B - Instruct (left) and Llama - 3.1 - 8B - Instruct/Llama - 3.1 - 70B - Instruct (right). Speed is measured in tokens/s, and Speedup is reported relative to standard speculative decoding (i.e., Vanilla SP) .
Figure 11: Accuracy and MAT on GSM8K for Qwen3-0.6B/Qwen3-8B.
Threshold
Full-vocab KL
Top-20
Top-200
Acc
MAT
Acc
MAT
ACC
MAT
Vanilla SP
87.75%
10.61
87.75%
10.61
87.75%
10.61
0.3
87.50%
13.33
86.00% ( ↓ 1.50%)
13.45
86.38% ( ↓ 1.12%)
13.35
0.6
85.50%
17.35
85.62%
17.40
86.05%
16.72
0.7
83.88%
18.94
84.00%
18.72
84.30%
18.90
0.8
83.12%
19.65
83.62%
19.93
83.49%
19.91
Table 2: Output quality and MAT under full-vocab and Top-K KL thresholding. Evaluated on GSM8K with Llama-3.2-1B-Instruct / Llama-3.1-8B-Instruct.
Figure 12: Accuracy–speedup trade-off of Token-KL vs. Step-KL on GSM8K (Llama-3.2-1B-Instruct / Llama-3.1-8B-Instruct).
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 13: Accuracy and MAT on GSM8K for Llama-3.2-1B-Instruct and Llama-3.1-8B-Instruct (with Qwen3-Max as the annotator).
Figure 14: Prompt for critical-token annotation, using Qwen3-Max as the LLM annotator.
Figure 15: Example of Qwen3-Max annotation on the GSM8K (Question ID: 5), with [ERROR _ START] and [ERROR _ END] marking critical tokens in the draft model’s erroneous solution.
Figure 16: Example of Qwen3-Max annotation on the GSM8K (Question ID: 68), with [ERROR _ START] and [ERROR _ END] marking critical tokens in the draft model’s erroneous solution.
Step
Full-vocab KL
Top-K KL
1: Top-K selection
Not needed
O(V) , branch-intensive, irregular access
2: Gather by index
Not needed
O(K) , indirect indexing, breaks coalescing
3: Log-softmax + KL
O(V) , streaming, fusible
O(K) , streaming, fusible
Appendix
Table 3: Operational comparison between full-vocab KL and Top-K KL. Full-vocab KL avoids the irregular memory access introduced by Steps 1–2 and consists entirely of streaming, fused-kernel-friendly operations.
Figure 17: Wall-clock latency (ms) of KL computation under different Top-K configurations. Full-vocab KL is 1.8 × faster than the fastest Top-K variant due to its fully streaming memory access pattern.
Vanilla SP
0.3
0.6
0.7
0.8
w/ fallback
Acc (%)
87.75
87.50
85.50
83.88
83.12
MAT
10.61
13.33
17.35
18.94
19.65
Trigger (%)
N/A
0.0
1.1
2.7
4.9
Trigger/Query
N/A
0
0.11
0.22
0.37
w/o
Acc (%)
87.75
87.60
86.62
85.25
84.38
MAT
10.61
13.45
17.21
18.75
19.80
Appendix
Table 4: Trigger frequency and performance of the Top-1 confidence fallback across KL thresholds. “Trigger (%)” denotes the percentage of token-level decisions where the fallback is activated; “Trigger / Query” denotes the average number of activations per query.
Setting
Metric
Vanilla SP
0.003
0.008
0.02
0.03
w/o fallback
Acc
87.75%
87.12%
85.12%
82.25%
80.38%
MAT
10.61
14.06
16.18
17.26
18.51
w/ fallback
Acc
87.75%
86.60%
84.12%
82.25%
80.75%
MAT
10.61
14.05
15.99
17.87
18.63
Appendix
Table 5: AutoJudge performance with and without Top-1 fallback under different KL thresholds.
δ (Acc Drop)
Acc (%)
Speed (tokens/s)
Speedup
Vanilla SP
82.56
29.45
1.00 ×
≤ 0.1%
82.77
39.58
1.03 ×
≤ 0.5%
82.00
41.91
1.09 ×
≤ 1.5%
81.07
48.35
1.26 ×
≤ 2.5%
79.46
50.96
1.33 ×
≤ 5%
77.13
61.45
1.60 ×
Appendix
Table 6: Quality-constrained speedups on vLLM with Llama-3.2-1B-Instruct/Llama-3.1-8B-Instruct.
δ (Acc Drop)
Acc (%)
Speed (tokens/s)
Speedup
Vanilla SP
93.00
24.19
1.00 ×
≤ 0.5%
92.69
27.21
1.12 ×
≤ 1.0%
92.02
31.77
1.31 ×
≤ 1.5%
91.41
32.78
1.36 ×
≤ 2.5%
90.36
34.29
1.42 ×
≤ 3.0%
90.27
35.10
1.45 ×
Appendix
Table 7: Quality-constrained speedups on vLLM with Llama-3.1-8B-Instruct/Llama-3.1-70B-Instruct.
LLM-as-a-judge is now the default measurement instrument for open-ended generation, but on the public JudgeBench benchmark even strong instruction-tuned judges barely scrape past random on objective-correctness pairwise items. We introduce RTLC, a three-stage prompting recipe -- Research, Teach-to-Learn, Critique -- that promotes a single black-box LLM into an ensemble-of-thought judge with no fine-tuning, retrieval, or external tools. Stage 1 wraps the input in a fixed pedagogical scaffold porting the Feynman Learning Technique (study → teach → find gaps → simplify) into LLM prompting. Stage 2 draws N=10 independent candidate verdicts at temperature 0.4. Stage 3 acts as its own critic, cross-comparing the candidate set against the original question to emit one critiqued verdict at temperature 0. On JudgeBench-GPT (350 hard pairwise items), Claude 3.7 Sonnet's pairwise accuracy climbs from 64.6% (single-shot vanilla prompt) to 78.6% (RTLC critique-of-10) -- an absolute 14.0-percentage-point gain. RTLC also beats N=10 self-consistency majority voting (77.7%) and a zero-shot first candidate (74.0%). A clean three-step ablation attributes +9.4 pp to the Teach-to-Learn scaffold, +3.7 pp to N=10 marginalisation, and +0.9 pp to explicit critique. We discuss the cost-accuracy frontier (RTLC sits above self-consistency at every working point), the error-budget breakdown across the four JudgeBench categories (knowledge, reasoning, math, coding), and how RTLC composes orthogonally with post-hoc judge-score calibration, with the two interventions compounding multiplicatively in practice.
Reasoning-capable large language models (LLMs) have recently been adopted as automated judges, but their benefits and costs in LLM-as-a-Judge settings remain unclear. Through controlled comparisons between reasoning and non-reasoning judges, we show that explicit reasoning substantially improves judgment accuracy on tasks requiring structured verification (e.g., math and coding), while offering limited or even negative gains on simpler evaluations and incurring significantly higher computational cost. These findings motivate that reasoning should be used selectively rather than universally, with awareness of possible distribution shift. We propose a Robust Adaptive Cost-Efficient Routing (RACER), which dynamically selects between reasoning and non-reasoning judges under a fixed budget by formulating routing as a constrained distributionally robust optimization problem. RACER explicitly accounts for distribution shift via a KL-divergence uncertainty set, admits an efficient primal--dual algorithm, and enjoys theoretical guarantees including uniqueness of the optimal policy and linear convergence. Extensive experiments show that RACER achieves superior accuracy--cost trade-offs under distribution shift.
Wenbo Zhang, Lijinghua Zhang, Liner Xiang +1
Department of Statistics, University of California, Irvine, USA.
LLM-as-Judge systems are widely deployed for automated evaluation, yet practitioners lack reliable methods to know when a judge's verdict should be trusted. Token log-probabilities, the standard post-hoc confidence signal, are unavailable for many commercial LLMs and, even when accessible, saturate above 0.999 with structured JSON output. We introduce VERDI (VERification-Decomposed Inference), a method that extracts confidence from the reasoning trace a structured judge already produces, with no additional inference calls. VERDI decomposes each verification-style evaluation into sub-checks and derives three structural signals: Step-Verdict Alignment, Claim-Level Margin, and Evidence Grounding Score. We combine them with Platt-scaled logistic regression. On three public benchmarks, VERDI achieves AUROC 0.72-0.91 on GPT-4.1-mini and 0.66-0.80 on GPT-5.4-mini. On Qwen3.5-4B/9B/27B, where answer-token logprobs are anti-calibrated (higher confidence on errors, AUROC 0.32-0.49), VERDI achieves 0.56-0.70. We additionally validate on a production system with eight rubrics (AUROC 0.73-0.88 on factual rubrics), demonstrate cross-model transfer (AUROC 0.66-0.69), and show that a 33M-parameter NLI (Natural Language Inference) model provides a scalable alternative to regex extraction.