cs.CLSep 21, 2026

LLJ Cards: Best practices for the Use of LLMs as Judges

Authors: Khaoula Chehbouni, Melina Medjdoub, Florian Carichon, Golnoosh Farnadi, Jackie Chi Kit Cheung

Organizations: McGill University · Mila - Quebec AI Institute · Concordia University

Abstract

In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these systems have been widely adopted by researchers and practitioners across a broad range of measurement tasks, driven by their strong performance, scalability, and cost-effectiveness relative to human judgment. However, a growing body of work has shown that the use of LLJs raise concerns about their validity and reliability as evaluators. Existing efforts to address these challenges have largely focused on developing bias-mitigation techniques and refining prompting strategies. While these approaches represent an important step forward, they primarily offer technical fixes and leave a more fundamental challenge unaddressed: the lack of standardized, transparent, and reproducible evaluation practices. In this paper, we introduce LLJ Cards, a framework that synthesizes best practices from measurement theory, natural language generation, and machine learning literature into practical guidelines for LLJ-based evaluations. While LLJs offer a promising path toward scalable evaluation, their effective use requires grounding in rigorous evaluation principles to ensure validity, reliability, and reproducibility. LLJ Cards addresses this need by providing a structured framework for applying these principles in the design and reporting of automated evaluations.

Figures & tables

Explore similar work

CardsList
  1. No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

    Mar 7, 2025Michael Krumdick, Charles Lovering, Varshini Reddy +2Llm-As-A-JudgeGrading

  2. Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?

    May 8, 2026Jane Paik KimLarge Language Model JudgesRater

  3. How Much Do LLM-as-a-Judge Design Choices Matter? A Systematic Comparison of Prompt Designs, Rating Scales, and Models

    Oct 4, 2026Laurène Vaugrante, Thilo Hagendorff