Mubric: Mutation Testing-Guided Rubric Generation for LLM Evaluation
Organizations: Tsinghua University · King’s College London · University of Illinois Urbana-Champaign
Abstract
Rubric-based evaluation is widely used to assess LLM-based systems by decomposing response quality into task-specific scoring criteria. However, automatically generating rubrics that reliably capture task-specific quality requirements remains challenging. We introduce Mubric, a mutation testing-guided approach to rubric generation. Mutation testing, a classic software testing methodology, evaluates a test suite by injecting faults into programs and checking whether the tests detect them. We draw an analogy between test suites and rubrics: if a rubric captures an important quality requirement, introducing a corresponding defect into an otherwise high-quality response should reduce its score. Mubric first mines common defects from real pairs of preferred and dispreferred responses and abstracts these defects into reusable mutation operators, each specifying how to introduce a particular type of response defect. For a new task, it applies relevant operators to a reference response, checks whether the injected defects reduce response quality, and uses insufficiently penalized defects to refine the rubric. We evaluate Mubric on 703 tasks across four representative domains against six advanced rubric generation methods. Mubric achieves the highest overall evaluation accuracy, outperforming the strongest baseline by 7.48 percentage points.
Figures & tables
| Method | GPT-5.4-mini | Gemini-2.5-flash-lite | Overall | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| STEM | Code | IF | RB-Chat | RM-Chat | STEM | Code | IF | RB-Chat | RM-Chat | ||
| TICK | 39.33 | 36.13 | 38.71 | 38.67 | 32.56 | 34.90 | 28.67 | 33.87 | 24.00 | 32.56 | 33.84 |
| RLCF-C | 47.26 | 33.90 | 51.61 | 43.33 | 35.66 | 39.86 | 42.67 | 41.13 | 32.00 | 30.47 | 39.87 |
| RLCF-R | 53.54 | 42.11 | 58.33 | 48.30 | 22.22 | 45.19 | 37.31 | 38.18 | 32.87 | 34.21 | 41.38 |
| RLCF-B | 46.03 | 43.08 | 61.54 | 41.10 | 25.24 | 42.22 | 43.28 | 40.00 | 32.87 | 28.07 | 40.32 |
| Dynamic | 49.66 | 46.98 | 45.16 | 51.01 | 29.46 | 42.86 | 36.24 | 36.29 | 48.00 | 17.05 | 40.74 |
| Method | STEM | Code | IF | RB-Chat | RM-Chat | Overall |
|---|---|---|---|---|---|---|
| Initial | 45.33 | 50.67 | 44.35 | 58.00 | 39.53 | 47.94 |
| Direct-Refine | 48.00 | 46.00 | 40.32 | 52.67 | 37.98 | 45.38 |
| On-the-Fly Ops | 42.67 | 48.00 | 43.55 | 53.33 | 31.78 | 44.24 |
| Mubric | 51.33 | 53.33 | 50.00 | 60.00 | 42.64 | 51.78 |
| Dataset | Aspect |
|---|---|
| STEM | Factuality (43%) , Logical correctness (33%) , Comprehension (12%) |
| Code | Logical correctness (70%) , Factuality (17%) , Comprehension (13%) |
| IF | Comprehension (48%) , Completeness (22%) , Factuality (15%) |
| RB-Chat | Comprehension (51%) , Factuality (31%) , Logical correctness (5%) |
| RM-Chat | Factuality (56%) , Comprehension (15%) , Logical correctness (14%) |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Task. “Hi, I have some falafel, but no tahini to put on them. Can you give me a recipe for making tahini?” | |
| Selected operator. Fabricate Factuality. Replace a correct factual statement with a false claim. Here, the operator changes the explanation about bitterness in the sesame-toasting step and leaves the rest of the recipe unchanged. | |
| Reference response (excerpt) | Mutated response (same passage) |
| Toast the sesame seeds Add sesame seeds to a dry skillet over medium-low heat. Toast for 3–5 minutes, stirring frequently, until they become lightly golden and fragrant. Be careful not to brown them too much, as burnt sesame seeds can taste bitter. Let them cool for a few minutes. | Toast the sesame seeds Add sesame seeds to a dry skillet over medium-low heat. Toast for 3–5 minutes, stirring frequently, until they become lightly golden and fragrant. Toasting the sesame seeds fully removes their natural oils, which helps prevent the tahini from becoming bitter. Let them cool for a few minutes. |
| Injected defect. The claim that toasting fully removes sesame oils is false. The rest of the response is unchanged, so the recipe retains its ingredients and preparation sequence despite the new factual error. | |
| Relevant items in the initial rubric | Relevant changes in the refined rubric |
| Item 2, weight 3. “Uses sesame seeds as the essential base ingredient and includes the core transformation into a tahini-like paste, typically by grinding/blending the seeds into a smooth or spreadable consistency.” Item 4, weight 2. “Gives clear, usable preparation steps in a sensible order, including any important preparation details such as toasting the seeds if recommended and blending until smooth.” | Added item, weight 2. “ Avoids introducing false or misleading factual claims about tahini ingredients, preparation methods, or culinary properties; all cooking guidance should be consistent with how tahini is actually made. ” |
| Task. “How many words are spoken daily on Earth? Try to explain your answer. Your explanation should take the reader through your reasoning step-by-step.” | |
| Selected operator. Break logic Logical robustness. Introduce an arithmetic error while keeping the stated assumptions and final estimate unchanged. Here, the operator changes the hourly word count and reuses the incorrect value when calculating the global total. This tests whether the rubric catches errors in the calculations supporting a conclusion. | |
| Assumptions in both responses. Both responses assume 6 billion daily speakers, an average of 1 hour of speech per speaker, and 130 words per minute. | |
| Reference response (excerpts) | Mutated response (aligned excerpts) |
| Step 3: Estimate words spoken per minute In one hour: words So an average speaking person might say about 7,800 words per day. Step 4: Multiply by the number of speakers That gives about: or 47 trillion spoken words per day. | Step 3: Estimate words spoken per minute In one hour: words So an average speaking person might say about 78,000 words per day . Step 4: Multiply by the number of speakers That gives about: or 47 trillion spoken words per day. |
| Injected defect. The mutation changes 7,800 to 78,000 in three places but leaves the product and conclusion unchanged. Thus is miscomputed, and the subsequent multiplication no longer supports the stated 47-trillion estimate. | |
| Relevant item in the initial rubric | The same item in the refined rubric |
| Task. “i have a form with two input: job_title, work_city, and a button which can submit make a django controller accept these two input argument, and print it on result page” | |
| Selected operator. Break logic Logical correctness. Change the target of a function call so that it performs the wrong action. Here, the operator changes which page the code displays after the user submits the form. | |
| Reading the code. Django is a Python web framework. The render function displays the named page using the supplied values. In both responses, job_form.html contains the input form, and result.html displays the submitted job title and city. The excerpts show the code after it has read these two values. | |
| Reference response (excerpt) | Mutated response (same code) |
| return render(request, " result.html ", { "job_title": job_title, "work_city": work_city }) | return render(request, " job_form.html ", { "job_title": job_title, "work_city": work_city }) |
| Injected defect. The submitted values are read correctly, but the code returns the input form instead of the result page. That form has no code to display the submitted values, so the user sees empty input fields. The mutation changes the page name in the code and updates the corresponding sentence in the explanation. | |
| Relevant item in the initial rubric | The same item in the refined rubric |