cs.CLMar 7, 2025

No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Authors: Michael Krumdick, Charles Lovering, Varshini Reddy, Seth Ebner, Chris Tanner

Organizations: Kensho Technologies · MIT

Abstract

Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-Judge framework, which uses prompted LLMs to evaluate response quality, is appealing due to its scalability, low cost, and strong correlations with human stylistic preferences. However, it remains unclear how accurately these methods can assess response quality in domains where correctness matters more than style. To address this gap, we introduce the Business and Finance Fundamentals Benchmark (BFF-Bench), a dataset of 160 challenging questions and long-form responses authored by financial professionals. These experts subsequently evaluated the correctness of 1,200 responses generated by a diverse set of LLMs on both BFF-Bench and a challenging subset of MT-Bench. With this expert-annotated dataset of judgments (VERDICTS), we analyze the agreement between a suite of automated grading methods and human experts. While we observe that LLM Judges are more reliable than other grading methods, our findings reveal a clear pattern in LLM Judge performance: when not provided with a correct reference, judges show high agreement with human experts only on questions the judges were able to correctly answer themselves. We demonstrate that providing the judges with expert-written references largely mitigates this issue, highlighting the limits of using LLM-as-a-Judge without any form of human verification.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LLM Judges Can Be Too Generous When There Is No Reference Answer

    Jul 14, 2026Chalamalasetti Kranti, Sowmya VajjalaLarge Language Model JudgesHuman Judgment

  2. Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

    Jun 1, 2026Junjie Chen, Yuxi Dong, Haitao Li +5Llm-As-A-Judge

  3. Flaws in the LLM Automation Narrative

    Jun 9, 2026George Perrett, Javae Elliott, Jennifer Hill +1Large Language Model BenchmarksNarratives