LLM-as-a-Judge

Latest papers 362

All topics
CardsList
  1. CRJudgeBench: Can AI Detect Plausible but Invalid Code Reviews?

    Sep 29, 2026Yue Pan, Jiawei Li, Ziyuan Zhang +2LLM EvaluationLLM-as-a-Judge

  2. From Judgment Quality to Downstream Utility: Rethinking LLM-as-a-Judge for Open-Ended Tasks

    Sep 29, 2026Zheng Zhang, Lufei Li, Xinyue Tan +4LLM EvaluationLLM-as-a-Judge

  3. Louder, Longer, Livelier: Acoustic Shortcuts and Underspecified Rationales in Speech LLM Judges

    Sep 29, 2026Mingyue Huo, Shivam Mehta, Bhavin Jawade +2LLM-as-a-JudgeSpeech Quality Assessment

  4. JudgeProfile: Understanding and Steering Subjectivity in LLM Judges

    Sep 29, 2026Qi Cao, Kangning Liu, Xuan Kan +10Human Preference EvaluationLLM Evaluation

  5. Certified Selective Automation of LLM Agent Evaluation

    Sep 28, 2026Chengguang Gan, Yunhao Liang, Qinghao Zhang +1LLM-as-a-JudgeSelective Prediction

  6. Curating Merchant-Matching Training Data with Two Confidence-Gated Local LLM Judges

    Sep 27, 2026Donghao Huang, Jinling Pei, Zhaoxia WangLLM-as-a-JudgeSelective Prediction

  7. Learning Strategies to Break Judges

    Sep 27, 2026Guruprerana Shabadi, Aaditya Naik, Rajeev Alur +1LLM-as-a-JudgeAI Agent Auditing

  8. LLM Judge Validation Under Sparse Overlap: From Inference to Design

    Sep 25, 2026Junxuan Li, Arko Mukherjee, Soumyabrata PalInter-Rater ReliabilityLLM Evaluation

  9. Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams

    Sep 24, 2026Ali Habibullah, Yazan Alshoibi, Mohammad Alshiekh +2Prompt SensitivityLLM-as-a-Judge

  10. Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality

    Sep 23, 2026Jiaju Huang, Hao Yang, Xinyu Ma +6Radiology Report GenerationLLM-as-a-Judge

  11. JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

    Sep 22, 2026Yubo Li, Yidi Miao, Ramayya Krishnan +1LLM-as-a-JudgeLanguage Model Generation Evaluation

  12. GroundedGEO: Auditing the Evidence Gap in Generative Search Rankings

    Sep 21, 2026Yihan Xia, Huiling Fan, Kangrong Zhong +1LLM-as-a-JudgeLLM Reranking

  13. LLJ Cards: Best practices for the Use of LLMs as Judges

    Sep 21, 2026Khaoula Chehbouni, Melina Medjdoub, Florian Carichon +2LLM EvaluationLLM-as-a-Judge

  14. ChartJudgeBench: Evaluating LMM Judges for Chart-to-Code Generation

    Sep 21, 2026Lijian Wu, Henry Hengyuan Zhao, Zijian Zhang +3VLM EvaluationLLM Evaluation

  15. How Many Humans Are 32 LLM Judges Worth?

    Sep 18, 2026Chao Li, Yingying Yu, Yunfeng LiAnnotator DisagreementLLM-as-a-Judge

  16. VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

    Sep 17, 2026Bhavana Akkiraju, Ravi Sastry Kolluru, Sri Charan D +3Audio QALLM-as-a-Judge

  17. I code or AI code: A comparative evaluation of AI-rated scores in classroom observations

    Sep 16, 2026Y. Fong, J. Xiang, T. Y. D. Chan +2LLM-as-a-JudgeEducational Assessment

  18. Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers

    Sep 16, 2026Zihan Chen, Di Zhu, Lei Zheng +1LLM-as-a-JudgeLLM Auditing

  19. A Calibrated Instrument for Measuring How Inference Optimizations Affect Output Quality

    Sep 16, 2026Jerry KaplanLLM QuantizationLLM Evaluation

  20. Skill-based Agentic Evaluation for Real-time Data Science Tasks

    Sep 15, 2026Aniruddha Tamhane, Raghavendra Addanki, Ayushi Aggarwal +4LLM-as-a-JudgeLLM Agent Evaluation