In this paper, we evaluate open-source generative LLMs on legal Natural Language Inference (NLI). Legal inspectorial processes take place in specific domains and often deal with confidential data. This creates a need for working with local models that do not require labeled training data. We evaluate our models on the ContractNLI benchmark and two NLI4Wills datasets. We successfully reproduce the baseline for the task (Span NLI BERT) and we evaluate multiple open-source LLMs on the same task. We analyze the invalid rate of the models, and their stability across temperature settings and domains. Among the generative models, Gemma-4 26B performs the best, reaching an accuracy of 81.2%, even outperforming the supervised model on one metric. On accuracy, it is not possible to beat the supervised model with zero-shot approaches. Qwen-3.6 35B performs well on both ContractNLI and additional datasets in the legal wills domain. Our findings indicate that zero-shot, open-source, generative LLMs are a viable alternative for real-world legal NLI when no supervised data is available. Our code is available at https://github.com/fbaratov/contractnli-llms.
Figures & tables
Metric
K&M
Reproduction
Re-Or
O.SD
R.SD
mAP
0.922±0.006
0.919±0.005
−0.003
✓
✓
P@R80
0.793±0.018
0.801±0.022
+0.008
✓
✓
Acc
0.875±0.006
0.869±0.006
−0.006
✓
✓
F1 (C)
0.357±0.039
0.341±0.043
−0.034
✓
✓
F1 (E)
0.834±0.002
0.826±0.004
−0.008
×
×
Table 1: Reproduction results for Span NLI BERT. “Re-Or” indicates the reproduction’s gain or loss over the values reported by Koreeda and Manning (2021) (‘K&M’). O.SD and R.SD indicate whether the difference is less than or equal to the original standard deviations reported by the authors or the reproduction’s standard deviations respectively.
Model
Accuracy
F1(C)
F1(E)
Invalid Rate (%)
Majority (Ent.)
0.463
0.000
0.590
—
Span NLI BERT
0.869±0.006
0.357±0.039
0.834±0.002
—
Gemma-4 E2B †
0.688±0.003
0.254±0.018
0.681±0.001
0.0±0.0
Gemma-4 E4B †
0.746±0.006
0.343±0.012
0.753±0.005
0.0±0.0
Gemma-4 26B ∗
0.812±0.002
0.443±0.030
0.808±0.003
1.0±0.1
Ministral-3 3B †
0.575±0.003
0.110±0.019
0.628±0.001
0.0±0.0
Table 2: Results on three-way classification. Boldface indicates the best-performing model per metric, and the “Invalid Rate (%)” column indicates the average percentage of answers outputted by the LLM from which we were unable to retrieve a coherent label. The green row highlights the best-performing LLM.
Entailment
Contradiction
NotM.
Model
Avg
Max
Avg
Max
Avg
Max
Gemma-4 E2B
✓
✓
✓
✓
✓
✓
Gemma-4 E4B
✓
✓
✓
✓
✓
✓
Gemma-4 26B
×
×
×
×
×
×
Ministral-3 3B
×
×
×
×
×
×
Ministral-3 8B
×
×
×
×
×
×
Table 3: Uncertainty analysis summary for the main results (Table 2 ) on the label-only scope. ✓ indicates tests on all 3 seeds are significant ( p~i<0.05 ). × indicates otherwise.
Model
Accuracy
F1(C)
F1(E)
IR (%)
S-NLI-BERT
0.869 0.006
0.357 0.039
0.834 0.002
—
Gemma-4 E2B ∗
0.680 0.013
0.268 0.026
0.692 0.008
0.0 0.0
Gemma-4 E4B ∗
0.744 0.004
0.386 0.013
0.747 0.004
0.0 0.0
Gemma-4 26B †
0.827 0.003
0.456 0.016
0.803 0.001
0.0 0.0
Qwen-3.5 4B †
0.741 0.007
0.370 0.025
0.742 0.006
0.1 0.1
Qwen-3.5 9B †
0.759 0.004
0.330 0.002
0.749 0.002
0.0 0.0
Table 4: Ablation study – best setting per model. Boldface indicates the best-performing model per metric. The IR (%) column corresponds to the invalid answer rate. Subscripts correspond to standard deviations. The green row highlights the best-performing LLM.
Appendix figures & tables54 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1: Effect of temperature on ContractNLI results grouped by model family.
Model
Temp
Accuracy
F1(C)
F1(E)
Invalid Rate (%)
Gemma-4 E2B †
0.1
0.689±0.002
0.258±0.018
0.676±0.004
0.000±0.000
0.2
0.685±0.001
0.263±0.006
0.675±0.002
0.000±0.000
0.3
0.688±0.003
0.254±0.018
0.681±0.001
0.016±0.023
0.4
0.683±0.008
0.236±0.011
0.676±0.006
0.016±0.023
0.5
0.685±0.007
0.233±0.003
0.678±0.006
0.000±0.000
Gemma-4 E4B †
0.1
0.743±0.004
0.359±0.040
0.751±0.002
0.000±0.000
Appendix
Table 5: Temperature analysis results. Each metric reports the average and standard deviation across three seeds. Boldface indicates the best-performing model per metric, and the “Invalid Rate (%)” column indicates the average percentage of answers outputted by the LLM from which we were unable to retrieve a coherent label. The green row highlights the best-performing temperature setting per LLM.
Table 6: Gemma-4 E2B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with logits of the whole response .
Table 7: Gemma-4 E2B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with only the label logits .
Table 8: Gemma-4 E2B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with the first token of the predicted label.
Table 9: Gemma-4 E4B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with logits of the whole response .
Table 10: Gemma-4 E4B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with only the label logits .
Table 11: Gemma-4 E4B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with the first token of the predicted label.
Table 12: Gemma-4 26B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with logits of the whole response .
Table 13: Gemma-4 26B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with only the label logits .
Table 14: Gemma-4 26B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with the first token of the predicted label.
Table 15: Ministral-3 3B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with logits of the whole response .
Table 16: Ministral-3 3B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with only the label logits .
Table 17: Ministral-3 3B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with the first token of the predicted label.
Table 18: Ministral-3 8B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with logits of the whole response .
Table 19: Ministral-3 8B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with only the label logits .
Table 20: Ministral-3 8B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with the first token of the predicted label.
Table 21: Ministral-3 14B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with logits of the whole response .
Table 22: Ministral-3 14B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with only the label logits .
Table 23: Ministral-3 14B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with the first token of the predicted label.
Table 24: Qwen-3.5 4B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with logits of the whole response .
Table 25: Qwen-3.5 4B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with only the label logits .
Table 26: Qwen-3.5 4B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with the first token of the predicted label.
Table 27: Qwen-3.5 9B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with logits of the whole response .
Table 28: Qwen-3.5 9B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with only the label logits .
Table 29: Qwen-3.5 9B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with the first token of the predicted label.
Table 30: Qwen-3.5 27B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with logits of the whole response .
Table 31: Qwen-3.5 27B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with only the label logits .
Table 32: Qwen-3.5 27B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with the first token of the predicted label.
Table 33: Qwen-3.6 35B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with logits of the whole response .
Table 34: Qwen-3.6 35B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with only the label logits .
Table 35: Qwen-3.6 35B Mann-Whitney U Test results and effect sizes for uncertainty metrics calculated with the first token of the predicted label.
Model
Precision
Recall
F1
Accuracy
Invalid Rate
Majority
0.197
0.333
0.248
0.591
—
Gemma-4 E2B †
0.766±0.011
0.722±0.023
0.713±0.025
0.786±0.016
0.0±0.0
Gemma-4 E4B †
0.808±0.011
0.751±0.010
0.768±0.015
0.840±0.006
0.0±0.0
Gemma-4 26B ∗
0.591±0.014
0.590±0.001
0.584±0.007
0.766±0.005
0.1±0.0
Ministral-3 3B †
0.618±0.028
0.598±0.017
0.516±0.026
0.526±0.028
0.0±0.0
Ministral-3 8B †
0.795±0.006
0.806±0.009
0.800±0.008
0.848±0.006
0.0±0.0
Appendix
Table 36: NLI4Wills Idaho results. Each metric reports the average and standard deviation across three seeds. Boldface indicates the best-performing model per metric, and the “Invalid Rate (%)” column indicates the average percentage of answers outputted by the LLM from which we were unable to retrieve a coherent label. The green row highlights the best-performing LLM. Majority baseline results are given for the unrelated label.
Model
Precision
Recall
F1
Accuracy
Invalid Rate
Majority
0.182
0.333
0.235
0.545
—
RoBERTa-large-mnli
0.967
0.962
0.964
0.969
—
Gemma-4 E2B
0.676±0.011
0.616±0.005
0.616±0.004
0.701±0.008
0.0±0.0
Gemma-4 E4B
0.785±0.021
0.673±0.014
0.702±0.016
0.766±0.012
0.0±0.0
Gemma-4 26B ∗†
0.553±0.005
0.497±0.006
0.521±0.005
0.685±0.005
0.1±0.0
Ministral-3 3B
0.560±0.031
0.562±0.035
0.503±0.031
0.499±0.035
0.0±0.0
Appendix
Table 37: NLI4Wills Tennessee results. Each metric reports the average and standard deviation across three seeds. Boldface indicates the best-performing model per metric, and the “Invalid Rate (%)” column indicates the average percentage of answers outputted by the LLM from which we were unable to retrieve a coherent label. The green row highlights the best-performing LLM. Majority baseline results are given for the unrelated label.
Setting
Accuracy
F1(C)
F1(E)
Invalid Rate (%)
Reasoning
0.680 ± 0.013
0.268 ± 0.026
0.692 ± 0.008
0.0 ± 0.0
Structure ∗
0.688 ± 0.003
0.254 ± 0.018
0.681 ± 0.001
0.0 ± 0.0
Both
0.688 ± 0.011
0.246 ± 0.016
0.683 ± 0.008
0.0 ± 0.0
Appendix
Table 38: Gemma-4 E2B ablation results on the ContractNLI test set. The asterisk ( ∗ ) indicates the setting used in the main results.
True \ Pred
NotMentioned
Entailment
Contradiction
InvalidAnswer
NotMentioned
455 ± 3
364 ± 4
84 ± 4
0 ± 0
Entailment
66 ± 5
861 ± 4
41 ± 2
0 ± 0
Contradiction
44 ± 6
54 ± 5
122 ± 2
0 ± 0
Appendix
Table 39: Gemma-4 E2B confusion matrix for the main results (Table 2 ).
Setting
Accuracy
F1(C)
F1(E)
Invalid Rate (%)
Reasoning
0.744 ± 0.004
0.386 ± 0.013
0.747 ± 0.004
0.0 ± 0.0
Structure ∗
0.746 ± 0.006
0.343 ± 0.012
0.753 ± 0.005
0.0 ± 0.0
Both
0.738 ± 0.000
0.357 ± 0.013
0.744 ± 0.003
0.0 ± 0.0
Appendix
Table 40: Gemma-4 E4B ablation results on the ContractNLI test set. The asterisk ( ∗ ) indicates the setting used in the main results.
True \ Pred
NotMentioned
Entailment
Contradiction
InvalidAnswer
NotMentioned
513 ± 10
289 ± 6
101 ± 5
0 ± 0
Entailment
32 ± 3
902 ± 4
34 ± 2
0 ± 0
Contradiction
23 ± 1
52 ± 2
145 ± 1
0 ± 0
Appendix
Table 41: Gemma-4 E4B confusion matrix for the main results (Table 2 ).
Setting
Accuracy
F1(C)
F1(E)
Invalid Rate (%)
Reasoning ∗
0.812 ± 0.002
0.443 ± 0.030
0.808 ± 0.003
1.0 ± 0.1
Structure
0.827 ± 0.003
0.456 ± 0.016
0.803 ± 0.001
0.0 ± 0.0
Appendix
Table 42: Gemma-4 26B ablation results on the ContractNLI test set. The asterisk ( ∗ ) indicates the setting used in the main results. No results are presented for reasoning-enabled structured outputs, as this leads to a ∼100% invalid output rate.
True \ Pred
NotMentioned
Entailment
Contradiction
InvalidAnswer
NotMentioned
654 ± 1
143 ± 1
95 ± 2
12 ± 1
Entailment
35 ± 2
893 ± 3
33 ± 3
7 ± 2
Contradiction
16 ± 1
51 ± 2
152 ± 1
2 ± 1
Appendix
Table 43: Gemma-4 26B confusion matrix for the main results (Table 2 ).
True \ Pred
NotMentioned
Entailment
Contradiction
InvalidAnswer
NotMentioned
322 ± 4
535 ± 5
46 ± 3
0 ± 0
Entailment
67 ± 2
847 ± 4
54 ± 2
0 ± 0
Contradiction
37 ± 6
149 ± 7
33 ± 5
0 ± 0
Appendix
Table 44: Ministral-3 3B confusion matrix for the main results (Table 2 ).
True \ Pred
NotMentioned
Entailment
Contradiction
InvalidAnswer
NotMentioned
458 ± 9
337 ± 9
108 ± 3
0 ± 0
Entailment
74 ± 6
832 ± 6
62 ± 3
0 ± 0
Contradiction
33 ± 1
50 ± 2
136 ± 3
0 ± 0
Appendix
Table 45: Ministral-3 8B confusion matrix for the main results (Table 2 ).
True \ Pred
NotMentioned
Entailment
Contradiction
InvalidAnswer
NotMentioned
447 ± 3
293 ± 4
163 ± 2
0 ± 0
Entailment
24 ± 3
888 ± 1
56 ± 3
0 ± 0
Contradiction
17 ± 1
34 ± 2
169 ± 2
0 ± 0
Appendix
Table 46: Ministral-3 14B confusion matrix for the main results (Table 2 ).
Setting
Accuracy
F1(C)
F1(E)
Invalid Rate (%)
Reasoning
0.519 ± 0.013
0.325 ± 0.033
0.616 ± 0.006
36.3 ± 1.7
Structure ∗
0.741 ± 0.007
0.370 ± 0.025
0.742 ± 0.006
0.1 ± 0.1
Both
0.750 ± 0.004
0.354 ± 0.046
0.737 ± 0.009
3.5 ± 0.5
Appendix
Table 47: Qwen-3.5 4B ablation results on the ContractNLI test set. The asterisk ( ∗ ) indicates the setting used in the main results.
True \ Pred
NotMentioned
Entailment
Contradiction
InvalidAnswer
NotMentioned
551 ± 12
206 ± 16
144 ± 4
2 ± 1
Entailment
57 ± 2
845 ± 6
67 ± 4
0 ± 0
Contradiction
25 ± 2
41 ± 2
154 ± 3
0 ± 0
Appendix
Table 48: Qwen-3.5 4B confusion matrix for the main results (Table 2 ).
Setting
Accuracy
F1(C)
F1(E)
Invalid Rate (%)
Reasoning
0.704 ± 0.008
0.340 ± 0.028
0.748 ± 0.005
12.0 ± 0.6
Structure ∗
0.759 ± 0.004
0.330 ± 0.002
0.749 ± 0.002
0.0 ± 0.0
Both
0.729 ± 0.020
0.339 ± 0.049
0.722 ± 0.022
0.2 ± 0.1
Appendix
Table 49: Qwen-3.5 9B ablation results on the ContractNLI test set. The asterisk ( ∗ ) indicates the setting used in the main results.
True \ Pred
NotMentioned
Entailment
Contradiction
InvalidAnswer
NotMentioned
606 ± 4
184 ± 4
113 ± 8
0 ± 0
Entailment
84 ± 7
827 ± 7
57 ± 9
0 ± 0
Contradiction
22 ± 0
45 ± 2
154 ± 2
0 ± 0
Appendix
Table 50: Qwen-3.5 9B confusion matrix for the main results (Table 2 ).
Setting
Accuracy
F1(C)
F1(E)
Invalid Rate (%)
Reasoning
0.260 ± 0.002
0.178 ± 0.059
0.424 ± 0.002
68.6 ± 0.0
Structure
0.796 ± 0.003
0.401 ± 0.019
0.786 ± 0.003
0.0 ± 0.0
Both ∗
0.793 ± 0.005
0.422 ± 0.024
0.783 ± 0.006
0.3 ± 0.0
Appendix
Table 51: Qwen-3.5 27B ablation results on the ContractNLI test set. The asterisk ( ∗ ) indicates the setting used in the main results.
True \ Pred
NotMentioned
Entailment
Contradiction
InvalidAnswer
NotMentioned
636 ± 7
144 ± 4
120 ± 4
3 ± 2
Entailment
51 ± 5
858 ± 6
56 ± 3
3 ± 1
Contradiction
11 ± 3
44 ± 0
164 ± 2
1 ± 0
Appendix
Table 52: Qwen-3.5 27B confusion matrix for the main results (Table 2 ).
Setting
Accuracy
F1(C)
F1(E)
Invalid Rate (%)
Reasoning
0.791 ± 0.005
0.387 ± 0.021
0.793 ± 0.003
0.1 ± 0.1
Structure
0.774 ± 0.005
0.364 ± 0.020
0.776 ± 0.002
0.1 ± 0.0
Both ∗
0.760 ± 0.006
0.330 ± 0.005
0.780 ± 0.008
0.6 ± 0.2
Appendix
Table 53: Qwen-3.6 35B ablation results on the ContractNLI test set. The asterisk ( ∗ ) indicates the setting used in the main results.
True \ Pred
NotMentioned
Entailment
Contradiction
InvalidAnswer
NotMentioned
520 ± 7
197 ± 12
175 ± 4
10 ± 3
Entailment
23 ± 3
897 ± 6
45 ± 2
3 ± 1
Contradiction
3 ± 1
46 ± 1
171 ± 1
0 ± 0
Appendix
Table 54: Qwen-3.6 35B confusion matrix for the main results (Table 2 ).
Model
P(C)
R(C)
F1(C)
Gemma-4 E2B
0.215 ± 0.027
0.254 ± 0.012
0.254 ± 0.018
Gemma-4 E4B
0.203 ± 0.009
0.382 ± 0.044
0.343 ± 0.012
Gemma-4 26B
0.287 ± 0.013
0.472 ± 0.051
0.443 ± 0.030
Ministral-3 3B
0.124 ± 0.026
0.128 ± 0.062
0.110 ± 0.019
Ministral-3 8B
0.219 ± 0.040
0.427 ± 0.029
0.325 ± 0.045
Ministral-3 14B
0.190 ± 0.006
0.558 ± 0.009
0.372 ± 0.008
Appendix
Table 55: Contradiction class metrics on the main results (Table 2 ). Boldface indicates the best-performing model per metric. All scores are micro averaged over documents, then macro averaged over labels.
Model
P(E)
R(E)
F1(E)
Gemma-4 E2B
0.601 ± 0.002
0.837 ± 0.006
0.681 ± 0.001
Gemma-4 E4B
0.708 ± 0.007
0.868 ± 0.005
0.753 ± 0.005
Gemma-4 26B
0.781 ± 0.005
0.867 ± 0.002
0.808 ± 0.003
Ministral-3 3B
0.541 ± 0.001
0.866 ± 0.008
0.628 ± 0.001
Ministral-3 8B
0.639 ± 0.011
0.830 ± 0.007
0.691 ± 0.009
Ministral-3 14B
0.672 ± 0.004
0.851 ± 0.001
0.728 ± 0.002
Appendix
Table 56: Entailment class metrics on the main results (Table 2 ). Boldface indicates the best-performing model per metric. All scores are micro averaged over documents, then macro averaged over labels.
Model
P(N)
R(N)
F1(N)
Gemma-4 E2B
0.779 ± 0.019
0.430 ± 0.004
0.511 ± 0.004
Gemma-4 E4B
0.844 ± 0.010
0.491 ± 0.009
0.575 ± 0.007
Gemma-4 26B
0.887 ± 0.006
0.657 ± 0.004
0.718 ± 0.003
Ministral-3 3B
0.649 ± 0.041
0.292 ± 0.005
0.369 ± 0.011
Ministral-3 8B
0.741 ± 0.031
0.431 ± 0.012
0.496 ± 0.016
Ministral-3 14B
0.726 ± 0.047
0.394 ± 0.007
0.467 ± 0.014
Appendix
Table 57: NotMentioned class metrics on the main results (Table 2 ). Boldface indicates the best-performing model per metric. All scores are micro averaged over documents, then macro averaged over labels.
As large language models (LLMs) are increasingly applied to real-world legal tasks, evaluating the reliability of their open-ended legal responses has become essential. These tasks require context-sensitive answers and allow little room for error, motivating fine-grained and diagnostic evaluation that can identify specific sources of response quality failures. We introduce LexRubric, a rubric-based benchmark for evaluating open-ended Chinese legal tasks. LexRubric contains 649 instances from legal consultation and judicial examination, which reflect both everyday legal needs and professional legal reasoning and cover 14 legal scenarios. It further includes 12,337 expert-written atomic scoring criteria organized under a unified six-dimensional framework, enabling accurate evaluation and diagnostic analysis across tasks and evaluation dimensions. To validate the reliability of the evaluation, we test multiple judge models and compare model-based judgments with human judgments. We further evaluate 18 recent general and legal-domain LLMs on LexRubric. Results show that different models exhibit distinct capability profiles, and that open-ended legal tasks remain challenging for current LLMs. Data is available at: https://github.com/foggpoy/LexRubric.
Yifan Chen, Haitao Li, Yiran Hu +6
1Beijing University of Posts and Telecommunications · 2Tsinghua University · University of Waterloo +1
Large Language Models (LLMs) achieve strong performance on reasoning tasks, but whether this reflects faithful logical inference or heuristic approximation remains unclear. We study this question in legal entailment by comparing three paradigms, including pure LLM classification, LLM-based Formal Reasoning, and solver-based Formal Reasoning using the Z3 SMT solver, on a re-annotated subset of ContractNLI across five LLMs. Our re-annotation reveals a systematic and measurable gap between pragmatic legal interpretation and strict formal entailment, where a substantial proportion of legally sound inferences are not formally grounded without additional unstated assumptions. While introducing formal structure improves accuracy, with LLM-based Formal Reasoning achieving the highest benchmark performance, we show that this gain does not imply faithful reasoning. We identify three recurring failure modes: scope laundering, where LLMs report solver-inconsistent classifications without executing the underlying formal reasoning, producing conclusions that appear logically grounded but are not; implicit constraint blindness, where LLMs overlook logical constraints present in formal representations; and program synthesis failures, where LLMs generate incorrect Z3 code despite structured prompting. Critically, scope laundering persists across all models, raising serious concerns about the faithfulness of LLM-based formal reasoning as a proxy for symbolic execution. These results reveal a fundamental gap between benchmark accuracy and logical faithfulness.
We introduce BenGER (Benchmark for German Law), a benchmark and dataset for evaluating LLM systems on subsumption-based legal reasoning in German law. The dataset combines 596 exam-style free-text legal case tasks across multiple levels of legal education and 531 short doctrinal reasoning tasks. It includes a controlled validation subset of timed human-written solutions under both unaided and human-AI co-creation conditions. We evaluate 12 contemporary LLM systems - closed flagship, efficiency-oriented, and open-weight - with a rubric-aligned LLM-as-a-Judge cross-validated against a multi-rater human-grading layer (three blind reviews per solution, six judge families benchmarked against the human pool). Closed-flagship systems lead the leaderboard across all three corpora, human-AI co-creation measurably improves on unaided human work, and the LLM judge tracks human grading at Pearson r=0.76 and Cohen's k=0.60. System rankings are stable across judge families and two judges from independent providers clear the Calderon single-reviewer replacement bar on human-authored solutions.
Sebastian Nagl, Ann-Kristin Mayrhofer, Martin Heidebach +6
Technical University of Munich (TUM) · Ludwig Maximilian University of Munich (LMU) · University of Konstanz +1