From Judgment Quality to Downstream Utility: Rethinking LLM-as-a-Judge for Open-Ended Tasks
Organizations: School of Information Science and Technology, ShanghaiTech University · State Key Laboratory of General Artificial Intelligence, BIGAI
Abstract
LLM-as-a-Judge is increasingly used to evaluate policy responses on open-ended tasks that lack ground-truth answers. Existing work often directly converts the resulting judgments into reward signals for policy training, paying limited attention to intrinsic judgment quality and largely restricting the use of Judges to training-time supervision. We systematically investigate judgment quality and downstream utility by examining both how judgments are elicited and how they are used. For judgment elicitation, we vary the Judge protocol along three dimensions: verdict granularity, critique usage, and evaluation batching. For judgment usage, beyond policy training, we extend Judge to test-time inference through Best-of-N selection, Judge-guided revision, and beam search. We find that, (i) Surprisingly, judgment quality and downstream utility do not always align. (ii) Judge protocol design substantially affects both intrinsic judgment quality and downstream utility. (iii) Judge guidance effectively converts test-time compute into performance gains, with benefits varying across inference strategies. Our results call for a multifaceted evaluation of LLM Judges on open-ended tasks, encompassing intrinsic judgment quality, and downstream utility.
Figures & tables
| Policy | Training-Time Judge | Verdict Form | Health Bench | RaR- Medicine | Research QA | RaR- Science |
| Qwen2.5-1.5B-Instruct | Qwen2.5-3B-Instruct | T/F | 0.140 0.006 | 0.266 0.004 | ||
| Rating | 0.400 0.071 | 0.333 0.019 | ||||
| gpt-oss-20b | T/F | |||||
| Rating | 0.164 0.003 | 0.285 0.006 | 0.402 0.028 | 0.349 0.005 | ||
| Qwen3-1.7B | Qwen2.5-3B-Instruct | T/F | 0.260 0.005 | 0.558 0.006 | 0.516 0.005 | |
| Rating | 0.315 0.004 |
| Policy | Training-Time Judge | Judge Output | Health Bench | RaR- Medicine | Research QA | RaR- Science |
| Qwen2.5-1.5B-Instruct | Qwen2.5-3B-Instruct | Critique+Verdict | 0.266 0.004 | |||
| Only Verdict | 0.149 0.007 | 0.398 0.019 | 0.340 0.004 | |||
| Verdict+Critique | 0.153 0.005 | 0.275 0.009 | 0.414 0.018 | 0.346 0.002 | ||
| gpt-oss-20b | Critique+Verdict | 0.158 0.014 | 0.383 0.005 | 0.334 0.006 | ||
| Only Verdict | 0.285 0.002 | 0.367 0.006 | 0.332 0.004 | |||
| Verdict+Critique | 0.160 0.010 | 0.282 0.008 | 0.334 0.002 |
| Policy | Training-Time Judge | Judge Output | Health Bench | RaR- Medicine | Research QA | RaR- Science |
| Qwen2.5-1.5B-Instruct | Qwen2.5-3B-Instruct | Independent | 0.266 0.004 | 0.371 0.021 | ||
| Batched | 0.143 0.005 | 0.306 0.005 | ||||
| gpt-oss-20b | Independent | 0.158 0.014 | 0.280 0.005 | 0.383 0.005 | ||
| Batched | 0.335 0.007 | |||||
| Qwen3-1.7B | Qwen2.5-3B-Instruct | Independent | 0.260 0.005 | 0.558 0.006 | 0.516 0.005 | |
| Batched | 0.319 0.003 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Policy | Training-Time Judge | Output | Medicine | Science | ||
| Response-Level Reward Ties | Uniform-Reward Groups | Response-Level Reward Ties | Uniform-Reward Groups | |||
| Qwen2.5-1.5B- Instruct | Qwen2.5-3B-Instruct | Critique+Verdict | 36.3% | 7.3% | 32.5% | 4.7% |
| Verdict Only | 40.4% | 9.5% | 35.8% | 7.7% | ||
| Verdict+Critique | 43.7% | 12.6% | 33.9% | 5.6% | ||
| GPT-OSS-20B | Critique+Verdict | 30.8% | 4.4% | 32.8% | 6.4% | |
| Verdict Only | 30.4% | 4.4% | 30.0% | 5.0% | ||
| Policy | Training-Time Judge | Protocol | Medicine | Science | ||
| Response-Level Reward Ties | Uniform-Reward Groups | Response-Level Reward Ties | Uniform-Reward Groups | |||
| Qwen2.5-1.5B -Instruct | Qwen2.5-3B-Instruct | Independent | 36.3% | 7.3% | 32.5% | 4.7% |
| Batched | 39.4% | 8.3% | 57.0% | 25.3% | ||
| GPT-OSS-20B | Independent | 30.8% | 4.4% | 32.8% | 6.4% | |
| Batched | 31.8% | 5.2% | 32.7% | 6.6% | ||
| Qwen3-1.7B | Qwen2.5-3B-Instruct | Independent | 38.0% | 8.4% | 26.3% | 2.6% |
| Policy | Training-Time Judge | Verdict Form | Health Bench | RaR- Medicine | Research QA | RaR- Science |
| Qwen2.5-7B-Instruct | Qwen2.5-7B-Instruct | T/F | 0.533 | |||
| Rating | 0.318 | 0.599 | 0.585 |
| Policy | Training-Time Judge | Judge Output | Health Bench | RaR- Medicine | Research QA | RaR- Science |
| Qwen2.5-7B-Instruct | Qwen2.5-7B-Instruct | Critique+Verdict | 0.314 | 0.533 | 0.577 | |
| Only Verdict | 0.518 | 0.586 | ||||
| Verdict+Critique | 0.318 | 0.584 | 0.578 |
| Policy | Training-Time Judge | Judge Output | Health Bench | RaR- Medicine | Research QA | RaR- Science |
| Qwen2.5-7B-Instruct | Qwen2.5-7B-Instruct | Single-criterion | 0.533 | 0.577 | ||
| Group-level | 0.320 | 0.582 |
| Code | Generic guidance |
| VERIFY_FACTS | Check factual claims and correct unsupported or contradictory statements. |
| ANSWER_DIRECTLY | State the requested answer directly and keep every part relevant to the question. |
| COMPLETE_REASONING | Fill in missing logical, mathematical, or causal steps needed to support the conclusion. |
| COVER_ALL_PARTS | Address every part of the visible user request. |
| USE_SUPPORT | Support important claims with an appropriate explanation, derivation, or concrete evidence. |
| CHECK_CONSISTENCY | Make the reasoning and final conclusion internally consistent. |