Bias in LLM-as-a-Judge

Latest papers 48

All topics
CardsList
  1. Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges

    May 13, 2026Riya Tapwal, Abhishek Kumar, Carsten MapleLLM-as-a-JudgeFaithfulness of Language Model Explanations

  2. Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese

    May 2, 2026Roseval Malaquias Junior, Giovana Kerche Bonás, Thales Sales Almeida +6LLM EvaluationLLM-as-a-Judge

  3. Quantifying and Mitigating Self-Preference Bias of LLM Judges

    Apr 24, 2026Jinming Yang, Zheng Hu, Chuxian Qiu +3LLM-as-a-JudgeLanguage Model Bias Evaluation

  4. MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge

    Apr 20, 2026Sua Lee, Sanghee Park, Jinbae ImLLM-as-a-JudgeMultimodal Large Language Models

  5. Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering

    Apr 18, 2026Zixiao Zhao, Amirreza Esmaeili, Fatemeh FardSoftware EngineeringLLM-as-a-Judge

  6. Context Over Content: Exposing Evaluation Faking in Automated Judges

    Apr 16, 2026Manan Gupta, Inderjeet Nair, Lu Wang +1LLM-as-a-JudgeLanguage Model Safety Evaluation

  7. MultiwayPAM: Multiway Partitioning Around Medoids for LLM-as-a-Judge Score Analysis

    Mar 11, 2026Chihiro Watanabe, Jingyu SunLLM-as-a-JudgeClustering

  8. Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization

    Mar 9, 2026Hongli Zhou, Hui Huang, Rui Zhang +5LLM-as-a-JudgeSocial Bias in Language Models

  9. Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges

    Feb 14, 2026Ruomeng Ding, Yifei Pang, He Sun +3LLM AlignmentLLM-as-a-Judge

  10. FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge

    Feb 6, 2026Bo Yang, Lanfei Feng, Yunkui Chen +3LLM EvaluationLLM-as-a-Judge

  11. Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations

    Jan 30, 2026Dani Roytburg, Matthew Bozoukov, Matthew Nguyen +3LLM-as-a-JudgeLanguage Model Generation Evaluation

  12. Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators

    Sep 3, 2025Dani Roytburg, Matthew Bozoukov, Matthew Nguyen +3LLM-as-a-JudgeLanguage Model Steering