VERA: Verdict-Conditioned Reliability for Adaptive LLM Judges
Organizations: Penn State University · Amazon
Abstract
Accurately estimating judgment reliability is a central challenge in adapting LLM judges to newly verified feedback while preserving previously learned behavior. However, existing approaches often rely on output-level confidence, which can be overconfident and poorly aligned with judgment correctness. We propose VERA, a VErdict-conditioned Reliability Axis that estimates reliability from hidden activations by distinguishing correct from incorrect judgments within each predicted-verdict group. Using VERA as a control signal, we develop a VERA-guided periodic adaptation framework that integrates reliability-ranked corrective updates, reliability-residual replay, and periodic refresh of the reliability directions. After VERA-guided adaptation on Chatbot Arena, 8B- and 14B-parameter judges outperform the strongest baseline on each of four held-out public benchmarks, with relative gains of up to 23.01%. The framework also improves focal-class recall by up to 16.1% relative to the strongest adaptive baselines on a separate proprietary temporal auditing task.
Figures & tables
| Model | Method | MT-Bench-HJ (%) | Auto-J(%) | JudgeBench(%) | LLMBar(%) |
|---|---|---|---|---|---|
| Qwen3-8B | IFT | ||||
| DPO | |||||
| VERA | 66.69 0.36 | 59.85 1.01 | 57.22 1.86 | 64.95 2.40 | |
| Relative gain over best baseline | |||||
| Qwen3-14B | IFT | ||||
| DPO | |||||
| Model | Method | Earlier anchor | Forward set | ||
|---|---|---|---|---|---|
| Recall (%) | F1 (%) | Recall (%) | F1 (%) | ||
| Qwen3-8B | Frozen initial judge | 85.8 | 78.8 | 65.7 | 58.8 |
| Periodic RAG | 82.0 | 79.0 | 61.6 | 61.6 | |
| Periodic IFT | 72.5 | 77.4 | 57.6 | 64.0 | |
| Periodic DPO | 84.1 | 78.7 | 62.6 | 57.9 | |
| VERA | 86.0 | 79.1 | 72.7 | 66.1 | |
| Setting | Model | Stage | Verbalized | Softmax | VERA |
|---|---|---|---|---|---|
| Proprietary | 8B | Frozen | 0.649 | 0.593 | 0.812 |
| IFT | 0.497 | 0.857 | 0.867 | ||
| DPO | 0.513 | 0.852 | 0.866 | ||
| 14B | Frozen | 0.618 | 0.628 | 0.831 | |
| IFT | 0.747 | 0.866 | 0.869 | ||
| DPO | 0.779 | 0.870 | 0.872 |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Hyperparameter | IFT | DPO | VERA -guided periodic update |
|---|---|---|---|
| Optimizer | AdamW | AdamW | AdamW |
| Maximum sequence length (tokens) | |||
| Peak learning rate | |||
| Training epochs | per feedback batch | ||
| Effective global batch size | |||
| Warmup ratio |
| Proprietary | Public | |||
|---|---|---|---|---|
| Hyperparameter | 8B | 14B | 8B | 14B |
| Feedback batches | ||||
| Selected layer | ||||
| Initial selection fraction (%) | ||||
| Feedback selection fraction (%) | ||||
| Final replay examples | ||||
| Setting | Model | Stage | Global | Per-class ( VERA ) | LR probe |
|---|---|---|---|---|---|
| Proprietary | Qwen3-8B | Frozen | 0.671 | 0.812 | 0.761 |
| IFT | 0.848 | 0.867 | 0.862 | ||
| MT-Bench-HJ | Qwen3-8B | Frozen | |||
| IFT |
| Setting | Model | Stage | Acc. (%) | Macro-F1 (%) | Softmax AUROC |
|---|---|---|---|---|---|
| Proprietary | 8B | Frozen | 49.5 | 56.5 | 0.543 |
| IFT | |||||
| DPO | |||||
| 14B | Frozen | 30.4 | 24.4 | 0.630 | |
| IFT | 87.1 | 84.8 | 0.866 | ||
| DPO | 86.9 | 84.6 | 0.870 |
| Signal | MT-Bench-HJ | Auto-J | JudgeBench | LLMBar | RewardBench 2 | Avg. |
|---|---|---|---|---|---|---|
| Source VERA | 0.732 | 0.711 | 0.616 | 0.565 | 0.760 | 0.677 |
| In-domain VERA (ours) | 0.732 | 0.721 | 0.633 | 0.745 | 0.814 | 0.729 |
| Softmax | 0.732 | 0.719 | 0.623 | 0.576 | 0.796 | 0.689 |
| Verbalized | 0.535 | 0.568 | 0.556 | 0.466 | 0.545 | 0.534 |
| Model | Selection rule | Anchor Rec. (%) | Forward Rec. (%) | Forward F1 (%) |
|---|---|---|---|---|
| Qwen3-8B | No replay | 78.7 | 61.6 | 63.9 |
| Random | 85.0 | 69.7 | 63.0 | |
| Softmax residual | 85.5 | 69.7 | 63.0 | |
| VERA residual | 86.0 | 72.7 | 66.1 |
| Model | Fraction | Earlier anchor | Forward set | ||
|---|---|---|---|---|---|
| Recall (%) | F1 (%) | Recall (%) | F1 (%) | ||
| Qwen3-8B | 86.0 | 79.1 | 72.7 | 66.1 | |
| 85.3 | 79.5 | 69.7 | 65.1 | ||
| 89.0 | 78.1 | 75.8 | 62.8 | ||
| Qwen3-14B | 84.2 | 80.0 | 72.7 | 68.6 | |
| 85.3 | 79.1 | 75.8 | 69.1 | ||
| Configuration | Earlier anchor | Forward set | ||||
|---|---|---|---|---|---|---|
| Bal. Acc. | Recall | F1 | Bal. Acc. | Recall | F1 | |
| Complete VERA | 86.5 | 86.0 | 79.1 | 83.4 | 72.7 | 66.1 |
| replay | 85.2 | 78.7 | 79.2 | 78.8 | 61.6 | 63.9 |
| reliability refresh | 86.9 | 84.9 | 80.3 | 81.2 | 67.7 | 64.7 |
| rank-UL | 85.4 | 79.1 | 79.4 | 77.6 | 58.6 | 63.0 |
| Model | Selector | Target-failure recall (%) | Enrichment |
|---|---|---|---|
| Qwen3-8B | Random | 28.1 | |
| VERA residual | 68.5 | ||
| Qwen3-14B | Random | 26.3 | |
| VERA residual | 81.6 |
| Model | Intervention | Safety Acc. (%) | Safety gain (%) | MT-Bench-HJ Acc. (%) |
|---|---|---|---|---|
| Qwen3-8B | Before repair | 49.6 | – | 66.4 |
| Random replay | 80.6 | 65.7 | ||
| VERA replay | 86.9 | 65.2 | ||
| Qwen3-14B | Before repair | 54.0 | – | 68.0 |
| Random replay | 72.1 | 65.9 | ||
| VERA replay | 78.6 | 66.0 |
| Model | Harness | Acc. (%) | Verbalized AUROC |
|---|---|---|---|
| Qwen3-8B | 0-judge | 63.7 | 0.665 |
| 3-subjudge | 65.3 | 0.634 | |
| 3-subjudge + RAG | 66.0 | 0.658 | |
| Qwen3-14B | 0-judge | 65.6 | 0.636 |
| 3-subjudge | 68.9 | 0.628 | |
| 3-subjudge + RAG | 70.4 | 0.603 |
| Model | Configuration | Time (s/example) |
|---|---|---|
| Qwen3-8B | 0-judge | 14.04 |
| 3-subjudge | 36.21 | |
| 3-subjudge + RAG | 37.18 | |
| Post-trained judge † | 0.24 | |
| Qwen3-14B | 0-judge | 12.18 |
| 3-subjudge | 43.64 |
| Model | Harness | Unsure (%) | Acc. (%) | Verbalized AUROC |
|---|---|---|---|---|
| Qwen3-8B | 0-judge | 0.0 | 62.0 | 0.533 |
| 3-subjudge † | 0.0 | 60.6 | 0.551 | |
| 3-subjudge + RAG † | 0.0 | 59.6 | 0.565 | |
| Qwen3-14B | 0-judge | 0.0 | 62.8 | 0.557 |
| 3-subjudge † | 0.0 | 64.2 | 0.566 | |
| 3-subjudge + RAG † | 0.0 | 64.4 | 0.558 |
| Evaluator | Agreement (%) | Consistency (%) |
|---|---|---|
| Published Auto-J results | ||
| LLaMA-2-13B-Chat | 29.81 | 48.56 |
| WizardLM-13B-v1.2 | 36.35 | 57.69 |
| Vicuna-13B-v1.5 | 39.22 | 62.07 |
| PandaLM | 39.44 | 66.88 |
| Claude-2 | 42.60 | 63.43 |
| Evaluator | Avg. † | Fact. | Prec. IF | Math | Safety | Focus | Ties † |
|---|---|---|---|---|---|---|---|
| Published RewardBench 2 results | |||||||
| LMUnit-Qwen2.5-72B | 82.1 | 87.2 | 54.4 | 72.7 | 91.3 | 96.8 | 90.1 |
| LMUnit-Llama3.1-70B | 80.5 | 84.6 | 48.8 | 71.6 | 90.7 | 97.0 | 90.6 |
| Gemini-2.5-Pro | 79.5 | 75.5 | 61.9 | 89.8 | 88.1 | 80.5 | 81.1 |
| Gemini-2.5-Flash | 77.7 | 67.4 | 57.5 | 85.2 | 90.9 | 84.1 | 80.9 |
| Claude Opus 4 | 76.5 | 82.7 | 41.9 | 74.9 | 89.5 | 86.2 | 83.7 |
| Evaluator | Accuracy (%) |
|---|---|
| Published JudgeBench results | |
| o3-mini (high), Arena-Hard | 80.86 |
| Claude-3.5-Sonnet, Arena-Hard | 64.29 |
| GPT-4o, Arena-Hard | 56.57 |
| Skywork-LLaMA-3.1-8B judge | 53.43 |
| VertexAI Gemini-1.5-Pro | 44.57 |
| Evaluator | Adversarial Acc. (%) |
|---|---|
| Published LLMBar results | |
| ChatGPT, vanilla | 32.0 |
| ChatGPT, enhanced prompt | 38.9 |
| LLaMA-2-70B, vanilla | 32.6 |
| LLaMA-2-70B, enhanced prompt | 43.4 |
| PaLM2, vanilla | 60.7 |