Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams
Organizations: KAUST Academy Computer, Electrical & Mathematical Sciences & Engineering (CEMSE) King Abdullah University of Science and Technology (KAUST) Thuwal, Saudi Arabia · Visual Artificial Intelligence Laboratory Oxford Brookes University · KAUST Academy & CEMSE, KAUST · Lady Margaret Hall, University of Oxford
Abstract
One long-form exam in a large course costs hundreds of grader-hours, and qualified graders are scarce; LLM graders are a tempting alternative. To show its pitfalls we grade a practical Computer Vision exam ( dual-graded students) under configurations spanning closed and open-weights models; the best reaches mean absolute error , below the two human graders achieve against each other. The catch is the prompt: a short ''strict grader'' preamble drives of open-weights models out of the graded band (), three stopping grading altogether. The damage traces to the preamble's two credit-withholding sentences, not to tone or model scale; one of them, ''never give partial credit'', alone makes two of three probed models stop grading. The closed flagships of three vendors shift calibration under it but stay in the band. In further configurations on a second, independent Machine Learning exam from another course ( dual-graded students), the preamble worsens ten models, moving three out of the band into collapse and one into refusal, yet improves seven whose neutral prompts over-mark: the vulnerability replicates, but its direction is exam-specific. Light LoRA fine-tuning repairs it: one adapter on the two exams' pooled graded examples brings five small open models to parity or better with a human grader in agreement with the grader pair, and sensitivity to the three harsh personas nearly vanishes ( MAE). We release the anonymised dataset, full ablation grid, and grading, fine-tuning and analysis pipelines.
Figures & tables
| CV exam ( ) | ML exam ( ) | |||||||
|---|---|---|---|---|---|---|---|---|
| Model | base | CV | ML | pooled | base | CV | ML | pooled |
| Qwen2.5-Coder-7B | 6.60 | 1.93 | 3.22 | 1.84 | 6.48 | 6.37 | 3.67 | 3.40 |
| Qwen2.5-Coder-14B | 4.06 | 1.89 | 3.29 | 1.89 | 7.70 | 5.41 | 3.35 | 3.29 |
| Qwen3-Coder-30B-A3B | 4.76 | 1.80 | 2.59 | 1.75 | 8.60 | 5.22 | 3.67 | 3.45 |
| Gemma-4-E4B | 4.29 | 1.88 | 3.79 | 1.88 | 4.66 | 4.79 ‡ | 3.51 | 3.53 |
| Llama-3.1-8B | 7.99 | 1.99 | 4.95 | 2.01 | 12.30 | 5.02 | 3.48 | 3.37 |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Run | Model | Persona | Q | Q | Q | Total | |
|---|---|---|---|---|---|---|---|
| CV exam | |||||||
| D01 | gemini-3-flash-preview | neutral | 570 | 0.92 | 0.90 | 0.91 | 0.95 |
| D02 | gemini-3.1-pro-preview | neutral | 570 | 0.92 | 0.90 | 0.90 | 0.94 |
| F02 | gemini-3.1-pro-preview | lenient | 570 | 0.92 | 0.90 | 0.91 | 0.94 |
| B03 | gemini-flash-lite + thinking | neutral | 570 | 0.89 | 0.88 | 0.87 | 0.92 |
| A01 | gemini-flash-lite | neutral | 570 | 0.77 | 0.65 | 0.83 | 0.84 |
| Run | Model | Family | Params | Neutral | Strict | Ratio | Behaviour |
| Open weights — ordered by total parameters | |||||||
| L-M05 | Qwen2.5-Coder-7B | Qwen | 7B | 5.68 | 7.92 | graded | |
| L-U02 | Llama-3.1-8B | Llama | 8B | 7.31 | 26.04 | refusal | |
| L-W02 | GLM-4-9B | GLM | 9B | 4.31 | 25.48 | near-refusal | |
| L-F02 | Gemma-3-12B | Gemma | 12B | 5.20 | 8.77 | collapse | |
| L-M07 | Qwen2.5-Coder-14B | Qwen | 14B | 3.51 | 24.52 | collapse | |
| Model | Family | Params | Neutral | strict | rigorous | exacting |
|---|---|---|---|---|---|---|
| Flash-Lite | Gemini | — | 3.34 | 5.75 | 3.92 | 4.58 |
| Llama-3.1-8B | Llama | 8B | 7.31 | 26.04 | 10.31 | 6.89 |
| GLM-4-9B | GLM | 9B | 4.31 | 25.48 | 4.53 | 5.18 |
| Gemma-3-12B | Gemma | 12B | 5.20 | 8.77 | 5.45 | 5.10 |
| DeepSeek-Coder-V2-Lite | DeepSeek | 16B | 5.98 | 12.09 | 7.01 | 6.69 |
| Mistral-Small-24B | Mistral | 24B | 3.66 | 26.03 | 6.15 | 7.04 |
| headword varied, policy fixed | policy varied, frame fixed | ||||||||
| Model | STRICT | RIGOROUS | FAIR | spread | nharsh | frame | S1 | S2 | both |
| Qwen2.5-Coder-32B | 16.41 | 17.74 | 18.55 | 2.15 | 8.25 | 9.37 | 15.55 | 12.92 | 20.26 |
| Llama-3.3-70B | 12.67 | 11.37 | 11.41 | 1.30 | 4.60 | — | — | — | — |
| Gemma-3-27B | 10.49 | 9.23 | 8.60 | 1.89 | 4.56 | — | — | — | — |
| GLM-4-32B | 7.02 | 7.43 | 6.41 | 1.02 | 2.88 | — | — | — | — |
| Mistral-Small-24B | 25.99 | 25.87 | 25.05 | 0.94 | 4.23 | 5.62 | 10.33 | 26.04 | 26.03 |
| Run | Model | Family | MAE | Q | …of which Q |
|---|---|---|---|---|---|
| blanket zeroing — consistent with refusal | |||||
| L-U02 | Llama-3.1-8B | Llama | 26.04 | 570 / 570 (100%) | 0 (0%) |
| L-Z02 | Mistral-Small-24B | Mistral | 26.03 | 570 / 570 (100%) | 0 (0%) |
| L-W02 | GLM-4-9B | GLM | 25.48 | 570 / 570 (100%) | 29 (5%) |
| L-M07 | Qwen2.5-Coder-14B | Qwen | 24.52 | 533 / 570 (94%) | 31 (6%) |
| L-J02 | GLM-4.5-Air | GLM | 21.02 | 361 / 570 (63%) | 55 (15%) |
| Run | Q score = 0 | …with prose | mean contradicted prose-credit |
|---|---|---|---|
| L-C01 (Qwen2.5-Coder-32B + strict) | 437 / 570 (77%) | 86 (20%) | 2.60 pts (max 10.0) |
| L-M01 (Qwen3-Coder-30B-A3B + strict) | 357 / 570 (63%) | 152 (43%) | 2.75 pts (max 8.5) |
| Model | Run | MAE [CI] | Bias | Zero | Behaviour |
|---|---|---|---|---|---|
| Qwen2.5-Coder-32B | paper neutral (L-A ) | 8.02 [7.38, 8.68] | 0% | collapse | |
| paper strict (L-C ) | 21.24 [20.31, 22.14] | 18% | collapse | ||
| decomposed, strict | 20.00 [19.10, 20.93] | 3% | collapse | ||
| arbitrated, strict | 9.29 [8.58, 9.98] | 0% | collapse | ||
| GLM-4-32B | paper neutral (L-X ) | 2.24 [1.91, 2.59] | 0% | graded | |
| paper strict (L-X ) | 8.90 [7.81, 10.03] | 5% | collapse |
| Model | Persona | CV exam (floor ) | ML exam (floor ) | ||
| MAE [95% CI] | bias | MAE [95% CI] | bias | ||
| claude-opus-5 | neutral | ||||
| strict | — | — | |||
| gemini-3.1-pro-preview | neutral | ||||
| strict | |||||
| lenient | |||||
| Run | Model | MAE | CI | Bias | |
| Human inter-grader floor (Section 4 ) | |||||
| — | G 1 vs. G 2 | 570 | 2.61 | — | |
| Full-cohort configurations | |||||
| D01 | gemini-3-flash-preview, neutral | 570 | 1.64 | ||
| F02 | gemini-3.1-pro-preview + lenient | 570 | 1.79 | ||
| D02 | gemini-3.1-pro-preview, neutral | 570 | 1.86 | ||
| Run | Configuration | [ CI] | Verdict | |||
| D01 | 3-flash-preview, neutral | 1.64 | 2.23 | 1.92 | better | |
| F02 | 3.1-pro-preview lenient | 1.79 | 2.32 | 2.15 | better | |
| D02 | 3.1-pro-preview, neutral | 1.86 | 2.53 | 1.96 | better | |
| B03 | flash-lite thinking | 2.05 | 2.55 | 2.29 | grader | |
| O03 | gpt-5.5 lenient (Section 6.1 ) | 2.25 | 2.59 | 2.62 | grader | |
| O01 | gpt-5.5 , neutral (Section 6.1 ) | 2.43 | 3.10 | 2.51 | grader |
| Run | Change vs. A01 | MAE | Bias |
|---|---|---|---|
| A01 | — (baseline) | 3.34 | |
| B01 | drop reference solution | 4.13 | |
| B02 | drop guidelines | 3.34 | |
| B03 | thinking on (dynamic) | 2.05 | |
| B04 | drop rubric breakdown | 3.22 |
| Run | Change vs. baseline | MAE | CI |
|---|---|---|---|
| D01 (restricted) | — (baseline) | 1.23 | |
| G04 | drop rubric breakdown | 1.25 | |
| G03 | thinking on (dynamic) | 1.39 | |
| G02 | drop guidelines | 1.60 | |
| G01 | drop reference solution | 1.63 |
| Run | Model | Params | MAE | CI | Bias |
| L-S01 | DeepSeek-Coder-V2-Lite (MoE) | 16B | 5.98 | ||
| L-W01 | GLM-4-9B | 9B | 4.31 | ||
| L-X01 | GLM-4-32B | 32B | 2.85 | ||
| L-J01 | GLM-4.5-Air (MoE) | 106B | 5.66 | ||
| L-F01 | Gemma-3-12B | 12B | 5.20 | ||
| L-Y01 | Gemma-3-27B | 27B | 4.32 |
| Run | Model | MAE | Bias |
|---|---|---|---|
| A01 | gemini-flash-lite (no few-shot) | 3.34 | |
| K01 | gemini-flash-lite + few-shot | 3.03 | |
| L-A30M | Qwen3-Coder-30B-A3B (no few-shot) | 4.50 | |
| L-K01 | Qwen3-Coder-30B-A3B + few-shot | 3.34 |
| CV exam | ML exam | |||
|---|---|---|---|---|
| Model | base | pooled | base | pooled |
| Qwen2.5-Coder-7B | ||||
| Qwen2.5-Coder-14B | ||||
| Qwen3-Coder-30B-A3B | ||||
| Gemma-4-E4B | ||||
| Llama-3.1-8B | ||||
| Graded on | base | CV-trained | ML-trained | pooled |
|---|---|---|---|---|
| CV exam, marks | 5.54 | 1.90 | 3.57 | 1.87 |
| CV exam, bd | 4.73 | 2.13 | 3.52 | 2.01 |
| ML exam, marks | 7.95 | 5.36 | 3.54 | 3.41 |
| ML exam, bd | 8.76 | 6.16 | 4.24 | 4.15 |
| CV exam | ML exam | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Arm | neut. | strict | rigor. | exact. | len. | neut. | strict | rigor. | exact. | len. |
| Qwen2.5-7B | base | 6.60 | 12.00 | 9.07 | 6.07 | 4.81 | 6.48 | 6.70 | 6.04 | 6.15 | 7.50 |
| pooled | 1.84 | 1.84 | 1.85 | 1.81 | 2.00 | 3.40 | 3.36 | 3.30 | 3.31 | 3.52 | |
| Qwen2.5-14B | base | 4.06 | 23.81 | 3.65 | 7.31 | 3.79 | 7.70 | 8.83 | 7.71 | 6.33 | 11.18 |
| pooled | 1.89 | 1.85 | 1.87 | 1.81 | 2.00 | 3.29 | 3.61 | 3.23 | 3.28 | 3.50 | |
| Qwen3-30B | base | 4.76 | 17.28 | 4.96 | 5.76 | 4.49 | 8.60 | 6.34 † | 7.34 | 6.30 | 13.00 |
| Cell | [ CI] | Verdict | ||||
|---|---|---|---|---|---|---|
| 7B marks, CV | 114 | 1.84 | 2.35 | 2.35 | better | |
| 14B marks, CV | 114 | 1.89 | 2.38 | 2.35 | better | |
| 30B marks, CV | 114 | 1.75 | 2.38 | 2.17 | better | |
| Gemma marks, CV | 114 | 1.88 | 2.41 | 2.34 | better | |
| Llama marks, CV | 114 | 2.01 | 2.54 | 2.34 | grader | |
| 7B marks, ML | 208 | 3.40 | 4.19 | 3.96 | better |
| Run | Configuration | Total | ||||||
| CV exam — -point base scale (Q , Q , Q ) | ||||||||
| D01 | 3-flash-preview, neutral | 570 | 0.30 | 0.29 | ||||
| D02 | 3.1-pro-preview, neutral | 570 | 0.82 | 0.00 | ||||
| F02 | 3.1-pro-preview lenient | 570 | 0.84 | 0.02 | ||||
| B03 | flash-lite thinking | 570 | 0.85 | 0.69 | ||||
| A01 | flash-lite, neutral (baseline) | 570 | 3.03 | 1.25 | ||||
| Pair | Human MAE | D AI MAE | D AI bias | |
|---|---|---|---|---|
| G 6 / G 16 | 56 | 0.85 | 1.86 | |
| G 9 / G 19 | 55 | 1.87 | 1.43 | |
| G 5 / G 15 | 57 | 2.12 | 1.95 | |
| G 4 / G 14 | 56 | 2.25 | 1.40 | |
| G 7 / G 17 | 57 | 2.48 | 1.84 | |
| G 8 / G 18 | 58 | 2.85 | 1.83 |
| model | neutral | strict | strict CI | ratio | floor | zero-rate | bias | behaviour |
|---|---|---|---|---|---|---|---|---|
| GLM-4.5-Air | 5.60 | 18.49 | [17.96, 19.02] | 3.6 | 37.8% | collapse | ||
| Qwen2.5-Coder-32B | 4.57 | 13.24 | [12.79, 13.69] | 2.6 | 10.5% | graded | ||
| Llama-3.1-8B | 13.16 | 29.58 | [28.70, 30.45] | 5.8 | 100.0% | refusal | ||
| Llama-3.3-70B | 4.70 | 10.17 | [9.77, 10.58] | 2.0 | 9.0% | graded | ||
| Mistral-Small-24B | 8.96 | 17.32 | [16.77, 17.87] | 3.4 | 32.2% | collapse | ||
| Qwen3-Coder-Next 80B | 4.47 | 8.63 | [8.26, 9.00] | 1.7 | 2.3% | graded |
| model (strict) | Q | of those, Q | share | reading |
|---|---|---|---|---|
| Llama-3.1-8B | 1038 / 1038 (100%) | 0 | 0% | blanket |
| GLM-4-9B | 669 / 1038 (64%) | 255 | 38% | selective |
| GLM-4.5-Air | 521 / 1038 (50%) | 110 | 21% | blanket |
| GLM-4-32B | 477 / 1038 (46%) | 295 | 62% | n/a (in band) |
| Mistral-Small-24B | 375 / 1038 (36%) | 21 | 6% | blanket |
| Qwen2.5-Coder-14B | 196 / 1038 (19%) | 13 | 7% | n/a (in band) |
| Run | Model | Persona | Prompt | MAE | CI | Bias | Behaviour | ||
|---|---|---|---|---|---|---|---|---|---|
| Closed-model configurations (Gemini: A–P series; OpenAI: O; Anthropic: N) | |||||||||
| A01 | flash-lite | neutral | SGB | graded | |||||
| B01 | flash-lite | neutral | GB | graded | |||||
| B02 | flash-lite | neutral | SB | graded | |||||
| B03 | flash-lite | neutral | SGBR | graded | |||||
| B04 | flash-lite | neutral | SG | graded | |||||
| Run | Model | Persona | Prompt | MAE | CI | Bias | Behaviour | ||
|---|---|---|---|---|---|---|---|---|---|
| Closed-model configurations (Gemini IG-, OpenAI IO-, Anthropic IN-) | |||||||||
| IG01 | flash-lite | neutral | SB | graded | |||||
| IG02 | flash-lite | neutral | B | graded | |||||
| IG03 | flash-lite | neutral | SBR | graded | |||||
| IG04 | flash-lite | neutral | S | graded | |||||
| IG05 | flash-lite | strict | SB | graded | |||||