LLM-as-a-Judge

Latest papers 362

All topics
CardsList
  1. RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

    May 20, 2026Zhenwei Tang, Zhaoyan Liu, Rasa Hosseinzadeh +3LLM EvaluationPairwise Comparison

  2. Exploring the Effectiveness of Using LLMs for Automated Assessment of Student Self Explanations in Programming Education

    May 20, 2026Arun-Balajiee Lekshmi-Narayanan, Mohammad Hassany, Peter BrusilovskyLLM-as-a-JudgeProgramming Education

  3. Pseudo-Formalization for Automatic Proof Verification

    May 19, 2026Slim Barkallah, Luke Bailey, Kaiyue Wen +2Mathematical Reasoning BenchmarksLLM-as-a-Judge

  4. LP-Eval: Rubric and Dataset for Measuring the Quality of Legal Proposition Generation

    May 19, 2026Shanshan Xu, Johan Lindholm, Amogh Raina +2Legal NLPLLM Evaluation

  5. Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

    May 18, 2026Leyao Wang, Yanan He, Peng Chen +5AI Agent ReliabilityLLM-as-a-Judge

  6. GRASP: Deterministic argument ranking in interaction graphs

    May 18, 2026Diganta Misra, Antonio Orvieto, Rediet Abebe +1LLM-as-a-JudgeLearning to Rank

  7. Estimating Item Difficulty with Large Language Models as Experts

    May 18, 2026Diana Kolesnikova, Kirill Fedyanin, Abe D. Hofman +2LLM EvaluationLLM-as-a-Judge

  8. QQJ: Quantifying Qualitative Judgment for Scalable and Human-Aligned Evaluation of Generative AI

    May 17, 2026Marjan Veysi, Pirooz Shamsinejadbabaki, Mohammad Zare +1LLM-as-a-JudgeHuman-in-the-Loop Evaluation

  9. Judge Circuits Explain Format-Induced Inconsistency in LLM-as-a-Judge

    May 15, 2026Nils Feldhus, Tanja Baeumel, Elena Golimblevskaia +10LLM-as-a-JudgeLLM Interpretability

  10. Margin-Adaptive Confidence Ranking for Reliable LLM Judgement

    May 14, 2026Gaojie Jin, Yong Tao, Lijia Yu +1LLM EvaluationLLM-as-a-Judge

  11. Small, Private Language Models as Teammates for Educational Assessment Design

    May 14, 2026Chris Davis Jaldi, Anmol Saini, Shan Zhang +3LLM EvaluationLLM-as-a-Judge

  12. Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations

    May 13, 2026Kyo Gerrits, Rik van Noord, Ana Guerberof ArenasCreativity AssessmentLLM-as-a-Judge

  13. Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges

    May 13, 2026Riya Tapwal, Abhishek Kumar, Carsten MapleLLM-as-a-JudgeFaithfulness of Language Model Explanations

  14. Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge

    May 11, 2026Wenbo Zhang, Lijinghua Zhang, Liner Xiang +1LLM-as-a-JudgeCost-Aware Inference

  15. LLM-Based Code Documentation Generation and Multi-Judge Evaluation

    May 11, 2026Ikbel Ghrab, Mohamed Dhieb, Ismail Khenissi +1LLM EvaluationLLM-as-a-Judge

  16. Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges

    May 10, 2026Yanran LiLLM EvaluationLLM-as-a-Judge

  17. Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?

    May 8, 2026Jane Paik KimLLM EvaluationLLM-as-a-Judge

  18. Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts

    May 8, 2026Saloni Garg, Amit SagtaniLLM EvaluationLLM-as-a-Judge

  19. SAGE: Hierarchical LLM-Based Literary Evaluation through Ontology-Grounded Interpretive Dimensions

    May 8, 2026Tianyu Wang, Nianjun ZhouComputational Literary StudiesLLM Evaluation

  20. When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels

    May 7, 2026Sushant Gautam, Finn Schwall, Annika Willoch Olstad +6LLM Safety BenchmarksLLM Evaluation

  21. Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement

    May 7, 2026Jessica Huynh, Alfredo Gomez, Athiya Deviyani +3Inter-Rater ReliabilityLLM-as-a-Judge

  22. Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges

    May 7, 2026Shihao Weng, Yang Feng, Xiaofei XieLLM-as-a-JudgeAI Agent Safety

  23. LaTA: A Drop-in, FERPA-Compliant Local-LLM Autograder for Upper-Division STEM Coursework

    May 6, 2026Jesse A. RodríguezLLM-as-a-JudgeGenerative AI in Education

  24. Counterargument for Critical Thinking as Judged by AI and Humans

    May 6, 2026Tosin Adewumi, Marcus Liwicki, Foteini Simistira Liwicki +3LLM-as-a-JudgeEducational Assessment