LLM-as-a-Judge

Latest papers 362

All topics
CardsList
  1. Sharding Prevents LLM Oversight Failures and Adversarial Exploitation

    Aug 5, 2026Victor Akinwande, J. Zico Kolter, Aran NayebiScalable OversightLLM Evaluation

  2. RADAR: Rubric-Aware Dependency and Redundancy Analysis for LLM-as-Judge Evaluation

    Aug 3, 2026Divyansh Singh, Reza Davari, Afra MashhadiLLM-as-a-JudgeLLM Auditing

  3. Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias

    Aug 3, 2026Baicheng Lin, Lingxi Jin, Kyung-Seok MinLLM-as-a-JudgeEducational Technology

  4. Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

    Aug 3, 2026Fengxian Ji, Yuke Li, Jingpu Yang +8LLM EvaluationLLM-as-a-Judge

  5. Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets

    Aug 2, 2026Wenhui Chen, Jianlin Chen, Ziyao Lin +2LLM EvaluationLLM-as-a-Judge

  6. A Triple-Robustness Analysis of Retrieval-Augmented Generation for Multi-Hop Requirements Traceability

    Aug 1, 2026Meftun Akarsu, Burak Özdemir, Doğancan Büyükçolak +1LLM-as-a-JudgeGraph Retrieval-Augmented Generation

  7. Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation

    Jul 30, 2026Zheng Wu, Yibo Luo, Pu Zhang +2Human Preference EvaluationLLM-as-a-Judge

  8. Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models

    Jul 30, 2026Shuyi Fan, Boyuan Deng, Mengyu Xu +4LLM EvaluationLLM-as-a-Judge

  9. SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents

    Jul 25, 2026Yang Wan, Zhenhao Zhang, Jierui Wang +1Reward ModelingLLM-as-a-Judge

  10. Reference-Free Evaluation of Reasoning in Open-Ended Question Answering

    Jul 22, 2026Guneet Singh Kohli, Yuxiang Zhou, Michael Sejr Schlichtkrull +2LLM-as-a-JudgeReasoning Evaluation

  11. EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

    Jul 20, 2026Jia-Kai Dong, Yi-Cheng Lin, Hung-yi LeeLLM-as-a-JudgeEducational Assessment

  12. Does Multi-Agent Debate Improve AI Feedback on Research Papers?

    Jul 16, 2026Tomas Havranek, Zuzana IrsovaLLM-as-a-JudgeMulti-Agent Debate

  13. The Test Oracle Problem in Synthetic LLM-as-Judge Corpora: Disappearance, Distortion and a Validation Protocol

    Jul 15, 2026Serkan BallıMultilingual Language Model EvaluationLLM-as-a-Judge

  14. Rating the Raters: Rasch Measurement Theory for LLM Evaluation

    Jul 14, 2026Pratik S. Sachdeva, Nathan BoudolLLM EvaluationLLM-as-a-Judge

  15. LLM Judges Can Be Too Generous When There Is No Reference Answer

    Jul 14, 2026Chalamalasetti Kranti, Sowmya VajjalaLLM EvaluationLLM-as-a-Judge

  16. Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

    Jul 13, 2026Zixiang Xu, Sixian Li, Huaxing Liu +4LLM-as-a-JudgeLLM Interpretability

  17. Knowledge Distillation for Automated AI Tutor Evaluation

    Jul 12, 2026Tahmid Al Hannan, Diego Garcia, Alex Njoroge +2LLM-as-a-JudgeAI in Education