cs.AISep 29, 2026

Mubric: Mutation Testing-Guided Rubric Generation for LLM Evaluation

Authors: Jiayuxuan Yang, Jie M. Zhang, Yiling Lou, Zhenpeng Chen

Organizations: Tsinghua University · King’s College London · University of Illinois Urbana-Champaign

Abstract

Rubric-based evaluation is widely used to assess LLM-based systems by decomposing response quality into task-specific scoring criteria. However, automatically generating rubrics that reliably capture task-specific quality requirements remains challenging. We introduce Mubric, a mutation testing-guided approach to rubric generation. Mutation testing, a classic software testing methodology, evaluates a test suite by injecting faults into programs and checking whether the tests detect them. We draw an analogy between test suites and rubrics: if a rubric captures an important quality requirement, introducing a corresponding defect into an otherwise high-quality response should reduce its score. Mubric first mines common defects from real pairs of preferred and dispreferred responses and abstracts these defects into reusable mutation operators, each specifying how to introduce a particular type of response defect. For a new task, it applies relevant operators to a reference response, checks whether the injected defects reduce response quality, and uses insufficiently penalized defects to refine the rubric. We evaluate Mubric on 703 tasks across four representative domains against six advanced rubric generation methods. Mubric achieves the highest overall evaluation accuracy, outperforming the strongest baseline by 7.48 percentage points.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. GenRubric: Self-Evolving Rubric Generation for Scalable LLM Evaluation

    Aug 30, 2026Yifan Chen, Haitao Li, Qingyao Ai +4RubricsLarge Language Model Evaluation

  2. Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence

    Jul 16, 2026Haocheng Yang, Licheng Pan, Yuan Lu +7RubricsTask-Specific Rubrics

  3. Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge

    May 28, 2026Zijie Wang, Eduardo BlancoRubricsLlm-As-A-Judge