JudgeMoE: Distributional Aggregation for LLM-as-a-Judge
Organizations: The University of Manchester · The University of Sheffield · University of Pittsburgh · Shanghai University of Finance and Economics
Abstract
When an LLM judge scores an output, its score distribution retains uncertainty and disagreement information that is lost after scalar compression. We introduce JudgeMoE, a lightweight aggregator that assigns example-specific weights to cached judge score distributions and fuses them before computing a final score. A protocol study shows that score-range choice is unstable across judge--dataset settings and that soft scoring usually outperforms hard decoding. On the original 10-cell benchmark, JudgeMoE improves mean Spearman over uniform log pooling by . Applying the same configuration to six additional cells yields a mean gain over the strongest local single judge across 16 cells, with positive differences in 12/16 cells and a one-sided Wilcoxon signed-rank . Validation-based analyses further show that the preferred aggregation method depends on the task and judge pool.
Figures & tables
| Dataset | Task | Size |
| SummEval | Summarization | 6,400 |
| TopicalChat | Dialogue quality | 2,160 |
| ROSCOE-eSNLI | Reasoning quality | 151 |
| HelpSteer2 | Helpfulness scoring | 101,620 main |
| WMT MQM | Translation quality | 150,347 |
| MT-Bench-human | Pairwise pref. | 2,575 eval |
| Comparison | Baseline | JudgeMoE | 95% CI | one-sided | two-sided | |
|---|---|---|---|---|---|---|
| Uniform log | 0.304 | 0.383 | +0.079 | [0.039, 0.119] | 0.005 | 0.010 |
| Cov-log | 0.341 | 0.383 | +0.043 | [-0.001, 0.085] | 0.053 | 0.105 |
| Best single | 0.341 | 0.383 | +0.042 | [-0.007, 0.088] | 0.053 | 0.105 |
| Follow-up | Hom. JudgeMoE | Het. JudgeMoE | Best single | Validation decision |
|---|---|---|---|---|
| SummEval + BERTSc. | 0.646 | 0.670 | 0.517 | keep expert |
| HelpSteer2 + RM | 0.533 | 0.545 | 0.390 | repair pool |
| WMT + trans. BERTSc. | 0.239 | 0.133 | 0.190 | keep LLM pool |
| WMT + COMET | 0.247 | 0.162 | 0.4326 | use standalone |
| Observed pattern | Choice |
|---|---|
| Protocol variation | Select protocol |
| Strong fixed pool | Fixed pooling |
| Learned-weight gain | Full JudgeMoE |
| Weak judge | Subset + retraining |
| Strong external evaluator | Standalone evaluator |
Appendix figures & tables51 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Prompt content |
|---|---|
| Instruction | Evaluator role + criterion |
| Context | Source / prompt / dialogue |
| Candidate | Output to score |
| Reference | Optional task reference |
| Output | Scalar score in specified range |
| Task family | Default direct-score template skeleton |
|---|---|
| Summarization | You are an expert summarization evaluator. Evaluate the [aspect] of the generated summary. [Aspect-specific criterion.] SOURCE DOCUMENT: {source} GENERATED SUMMARY: {candidate} Return only a score from {low} to {high}. |
| Dialogue | You are an expert dialogue response evaluator. Evaluate the [aspect] of the assistant response. CONTEXT: {source} ASSISTANT RESPONSE: {candidate} Return only a score from {low} to {high}. |
| Assistant response | You are an expert evaluator of assistant responses. Evaluate the [aspect] of the response. USER REQUEST: {source} ASSISTANT RESPONSE: {candidate} Return only a score from {low} to {high}. |
| Reasoning | You are an expert evaluator of reasoning chains. Evaluate the overall quality of the reasoning chain. Consider correctness, completeness, and coherence. PROBLEM OR CONTEXT: {source} REASONING CHAIN: {candidate} Return only a score from {low} to {high}. |
| Translation | You are an expert translation evaluator. Evaluate the overall translation quality. Consider adequacy and fluency. SOURCE SEGMENT: {source} TRANSLATION: {candidate} [optional HUMAN REFERENCE] Return only a score from {low} to {high}. |
| FLASK-style | You are an expert evaluator of assistant responses. Evaluate the candidate response using the provided criterion and scoring rubric. INSTRUCTION: {source} CANDIDATE RESPONSE: {candidate} CRITERION: {criterion} [reference/rubric blocks] Return only a score from {low} to {high}. |
| Experiment | Paper content |
|---|---|
| Protocol study | Scale + soft/hard scoring |
| Aggregation | Main results + ablations |
| Seed sweep | 6-judge / 3-judge variation |
| Pairwise study | Direct + pairwise training |
| Large judges | Quality + efficiency |
| WMT MQM | Pool failure + sparse retraining |
| Family | Aggregation input | Full distribution | Example-specific weights | Interaction |
|---|---|---|---|---|
| Single-judge methods | Score / rationale / checklist | No | No | No |
| Jury / debate | Votes / answers / rationales | Usually no | Sometimes | Yes |
| Latent-reliability models | Labels / preferences | No | Judge-level | No |
| Scalar weighting | Scalar scores | No | Yes | No |
| Fixed opinion pooling | Score distributions | Yes | No | No |
| JudgeMoE | Cached score distributions | Yes | Yes | No |
| Dataset | Train | Dev | Valid | Test | Total | Note |
| Direct-score corpus | ||||||
| All direct-score data | 212,953 | 15,912 | 5,190 | 31,813 | 265,868 | – |
| HelpSteer2 | – | – | – | – | 106,810 | Train / validation preserved |
| SummEval | – | – | – | – | 6,400 | Human scores retained |
| TopicalChat | – | – | – | – | 2,160 | Dialogue aspects retained |
| ROSCOE | – | – | – | – | 151 | Diagnostic subset |
| Item | Setting | Reference |
|---|---|---|
| Architecture | DeepSets, , dropout 0.1, log pool, , | Table 13 |
| Optimization | Adam, lr , batch 128, 20 epochs | Released logs |
| Objective | Huber loss; auxiliary ranking loss with margin 0.0 | Section 3.3, Table 42 |
| Seeds | 13; sweep: 13, 17, 29 | Table 41 |
| Features | score probabilities + scalar summaries | Table 18 |
| Splits | Original split labels; exact example-ID intersection | Table 9 |
| Dataset | Scale | Mainline | 3-judge | Best single | Best judge | Uniform log | Cov-log |
|---|---|---|---|---|---|---|---|
| SummEval | 1–5 | 0.5134 | 0.4967 | 0.4687 | Gemma-2-9B | 0.4420 | 0.4313 |
| SummEval | 1–10 | 0.6101 | 0.6043 | 0.5657 | Gemma-3-12B | 0.6401 | 0.6494 |
| TopicalChat | 1–5 | 0.4357 | 0.4297 | 0.4683 | Gemma-2-9B | 0.3108 | 0.4928 |
| TopicalChat | 1–10 | 0.4270 | 0.3925 | 0.4458 | Gemma-2-9B | 0.3700 | 0.4708 |
| ROSCOE | 1–5 | 0.2870 | 0.2841 | 0.2641 | Gemma-3-12B | 0.2975 | 0.2142 |
| ROSCOE | 1–10 | 0.2041 | 0.1381 | 0.0773 | Ministral-3-8B | 0.0644 | 0.0516 |
| Temperature | Avg. Sp. | vs. single | Wins / 10 |
|---|---|---|---|
| 0.3828 | 0.0414 | 2 | |
| 0.3832 | 0.0419 | 4 | |
| 0.3729 | 0.0315 | 4 |
| Learned aggregation configuration | Avg. Sp. |
|---|---|
| Linear + log pool | 0.3511 |
| FFN + linear pool | 0.3505 |
| Low-rank linear + log | 0.3546 |
| DeepSets + log pool | 0.3828 |
| DeepSets + log, | 0.3832 |
| 3-judge efficient | 0.3656 |
| Gate | Linear pooling | Log pooling |
|---|---|---|
| FFN | ||
| DeepSets |
| Selector | Equal-cell Sp. | Weighted Sp. |
|---|---|---|
| Validation-selected protocol | 0.1865 | 0.1674 |
| Fixed 1–5 soft | 0.1486 | 0.0269 |
| Fixed 1–10 soft | 0.1813 | 0.1002 |
| Best protocol by test performance | 0.2569 | 0.1681 |
| Style | Rows | Avg. soft Sp. | Wins vs. default |
|---|---|---|---|
| Default | 12 | 0.3230 | 0/12 |
| Rubric | 12 | 0.2275 | 0/12 |
| Critical | 12 | 0.2968 | 4/12 |
| Pool | Judges | Mean | JudgeMoE Sp. | Wins vs. default |
|---|---|---|---|---|
| Default 2 | 2 | 0.7955 | 0.4049 | – |
| Protocol-diverse 6 | 6 | 0.8472 | 0.3883 | 3 / 3 |
| Feature set | Avg. Sp. | vs. single | Wins / 10 |
|---|---|---|---|
| Full current | 0.3832 | 0.0419 | 6 |
| Distribution + scores | 0.3811 | 0.0398 | 2 |
| Distribution only | 0.3746 | 0.0332 | 2 |
| Input features | Mean Sp. | SD |
|---|---|---|
| Hard score only | 0.2634 | 0.0028 |
| Soft score only | 0.3100 | 0.0088 |
| Hard and soft scores | 0.3671 | 0.0027 |
| Probability-derived scalar summaries | 0.3776 | 0.0053 |
| Cell | Full-6 | Strong-3 |
|---|---|---|
| Average (10 cells) | 0.7775 | 0.7071 |
| SummEval 1–10 | 0.7611 | 0.5860 |
| TopicalChat 1–5 | 0.8705 | 0.7875 |
| TopicalChat 1–10 | 0.9467 | 0.9150 |
| WMT MQM 1–10 | 0.5883 | 0.5696 |
| Cell | Best test family | Selected | Correct |
|---|---|---|---|
| ROSCOE 1–10 | sparse | sparse | yes |
| ROSCOE 1–5 | fixed | fixed | yes |
| SummEval 1–10 | fixed | fixed | yes |
| SummEval 1–5 | sparse | fixed | no |
| TopicalChat 1–10 | fixed | fixed | yes |
| TopicalChat 1–5 | fixed | fixed | yes |
| Strategy | Avg. test Sp. | Weighted test Sp. |
|---|---|---|
| Best single (valid-selected) | 0.2748 | 0.1860 |
| Always fixed pool | 0.3374 | 0.1319 |
| Always full JudgeMoE | 0.2975 | 0.1655 |
| Always sparse variant | 0.3497 | 0.1536 |
| Best family by validation performance | 0.3402 | 0.1650 |
| LOO heuristic selector | 0.3495 | 0.1694 |
| Strategy | Avg. test Sp. | Weighted test Sp. |
|---|---|---|
| Best single (valid-selected) | 0.2748 | 0.1860 |
| Full fixed pool | 0.3374 | 0.1319 |
| Full JudgeMoE | 0.3154 | 0.1665 |
| Valid-selected subset fixed | 0.3133 | 0.2113 |
| Valid-selected family | 0.3732 | 0.2213 |
| Best family by test performance | 0.3789 | 0.2215 |
| Validation signal | Recommended choice |
|---|---|
| High residual correlation; covariance-aware pooling already strong | Fixed covariance-aware pooling |
| Large leave-one-out gain or large subset/full-pool gap | Subset selection + retraining |
| Remaining gain from example-specific weighting | Full JudgeMoE |
| Stage | Validation signal | Selection |
|---|---|---|
| Protocol | Scale / extraction differences | Select protocol |
| Family | Fixed / full / single performance | Select aggregation family |
| Subset | Subset performance | Select judge subset |
| Subset retraining | Subset / full-pool difference | Retrain on selected subset |
| Subset | Weighted Sp. | vs. single | vs. cov-log |
|---|---|---|---|
| All 10 cells | 0.3052 | 0.0413 | 0.0549 |
| Exclude ROSCOE | 0.3052 | 0.0413 | 0.0548 |
| Exclude WMT | 0.5393 | 0.1346 | 0.0961 |
| Non-HelpSteer | 0.1549 | -0.0241 | 0.0236 |
| Cell | JudgeMoE | Best single | 95% CI | |||
|---|---|---|---|---|---|---|
| helpsteer2_train@1-10 | 101620 | 0.5386 | 0.3849 | 0.1537 † | [0.1502, 0.1572] | 0.0100 |
| helpsteer2_train@1-5 | 101620 | 0.5426 | 0.4087 | 0.1339 † | [0.1303, 0.1373] | 0.0100 |
| roscoe_esnli@1-10 | 151 | 0.2041 | 0.0773 | 0.1268 | [-0.0240, 0.2675] | 0.1294 |
| roscoe_esnli@1-5 | 151 | 0.2870 | 0.2641 | 0.0228 | [-0.0063, 0.0465] | 0.1791 |
| summeval@1-10 | 6400 | 0.6101 | 0.5657 | 0.0444 † | [0.0260, 0.0653] | 0.0100 |
| summeval@1-5 | 6400 | 0.5134 | 0.4687 | 0.0448 † | [0.0319, 0.0591] | 0.0100 |
| System | ECE | MAE |
|---|---|---|
| JudgeMoE | 0.0588 | 0.1890 |
| Best single | 0.2120 | 0.2782 |
| Uniform log | 0.1785 | 0.2718 |
| Cov-log | 0.2214 | 0.2976 |
| Dataset | JudgeMoE | Best single | Uniform log | Cov-log |
|---|---|---|---|---|
| SummEval | 0.0661 | 0.3840 | 0.1621 | 0.2210 |
| TopicalChat | 0.0559 | 0.2640 | 0.1385 | 0.2118 |
| ROSCOE | 0.1053 | 0.2161 | 0.2693 | 0.3381 |
| HelpSteer2 | 0.0133 | 0.0483 | 0.0859 | 0.0791 |
| WMT MQM | 0.0534 | 0.1476 | 0.2368 | 0.2568 |
| Comparison | Mean | SD Δ | Power @10 | Cells for 80% |
|---|---|---|---|---|
| vs. uniform log | +0.0794 | 0.0692 | 0.98 | 5 |
| vs. cov-log | +0.0427 | 0.0729 | 0.58 | 19 |
| vs. best single | +0.0419 | 0.0812 | 0.49 | 24 |
| Component | Selection setup | Reported use |
|---|---|---|
| Protocol study | Five score ranges with hard/soft extraction; validation-based protocol selection over 24 cells | Select scoring protocol |
| Main architecture | Five architecture/pooling variants, temperature sweep, and feature ablation on the original 10 cells | Select 6-judge DeepSets + log pooling |
| Compact 3-judge pool | Candidate triples from protocol-study screening | Efficiency setting |
| WMT subset selection | Top- subsets and leave-one-out selection on WMT validation data | Retrain selected subsets |
| General subset selector | All 63 nonempty judge subsets over eight validation/test cells | Select judge subset and aggregation family |
| Heterogeneous experts | One task-aligned expert added at a time using task-specific train or validation data | Evaluate heterogeneous pools and standalone experts |
| Cell | JudgeMoE | Single | Cov-log | single | cov-log |
|---|---|---|---|---|---|
| HelpSteer2 val.@1–5 | 0.4394 | 0.3851 | 0.4212 | +0.0543 | +0.0182 |
| HelpSteer2 val.@1–10 | 0.4562 | 0.3805 | 0.4358 | +0.0756 | +0.0204 |
| 12-cell avg. | 0.3940 | 0.3483 | 0.3552 | +0.0457 | +0.0388 |
| Cell | JudgeMoE | Single | Cov-log | single | cov-log |
|---|---|---|---|---|---|
| FLASK@1–5 | 0.5452 | 0.4897 | 0.5452 | +0.0555 | -0.0001 |
| FLASK@1–10 | 0.4364 | 0.4148 | 0.4449 | +0.0216 | -0.0085 |
| 12-cell avg. | 0.4011 | 0.3598 | 0.3663 | +0.0413 | +0.0349 |
| Setting | Comparison | Left | Right | 95% CI | ||
|---|---|---|---|---|---|---|
| SummEval 1–10 + BERTScore | Het. JudgeMoE vs. best single | 0.6703 | 0.5172 | +0.1531 | [0.1104, 0.1944] | 0.0007 |
| SummEval 1–10 + BERTScore | Het. JudgeMoE vs. uniform log | 0.6703 | 0.1930 | +0.4773 | [0.4219, 0.5320] | 0.0007 |
| SummEval 1–10 + BERTScore | Het. JudgeMoE vs. hom. JudgeMoE | 0.6703 | 0.6460 | +0.0242 | [-0.0026, 0.0509] | 0.0733 |
| HelpSteer2 1–10 + reward | Het. JudgeMoE vs. best single | 0.5454 | 0.3900 | +0.1553 | [0.1468, 0.1644] | 0.0025 |
| HelpSteer2 1–10 + reward | Het. JudgeMoE vs. uniform log | 0.5454 | 0.3102 | +0.2352 | [0.2239, 0.2472] | 0.0025 |
| HelpSteer2 1–10 + reward | Het. JudgeMoE vs. hom. JudgeMoE | 0.5454 | 0.5330 | +0.0124 | [0.0066, 0.0189] | 0.0025 |
| Pool | Best single | Uniform log | Cov-log | JudgeMoE |
|---|---|---|---|---|
| 3-judge LLM pool | 0.1898 | 0.2031 | 0.1988 | 0.2468 |
| 3-judge LLM + COMET | 0.4326 | 0.3926 | 0.3794 | 0.1620 |
| Setting | Best single | Full JudgeMoE | vs. single | min frac single | min frac gain | min frac gain |
|---|---|---|---|---|---|---|
| SummEval 1–10 + BERTScore (repaired) | 0.5172 | 0.6682 | 0.1510 | 0.05 | 0.40 | 1.00 |
| HelpSteer2 1–10 + reward | 0.3900 | 0.5670 | 0.1769 | 0.05 | 0.70 | 0.70 |
| Setting | Expert | mean vs. hom. abs.-err. | mean vs. single abs.-err. | top-gate rate | Role |
|---|---|---|---|---|---|
| SummEval 1–10 + BERTScore | BERTScore | 0.0744 | 0.3167 | 0.1203 | frequent expert |
| HelpSteer2 1–10 + reward | Skywork reward | 0.0031 | 0.0321 | 0.0000 | soft calibration |
| Setting | Strong single | Best single | Hom. JudgeMoE | Het. JudgeMoE | Het. | Hom. |
|---|---|---|---|---|---|---|
| SummEval@1–10 | Qwen2.5-72B | 0.5172 | 0.6460 | 0.6703 | 0.1779 | 0.1536 |
| HelpSteer2@1–10 | Qwen2.5-72B | 0.3900 | 0.5330 | 0.5454 | 0.0716 | 0.0593 |
| WildBench@1–10 | Qwen2.5-72B | 0.3520 | 0.1719 | 0.0854 | 0.0822 | 0.1688 |
| Setting | Groups | Avg. cand. | Best group score | Het. | Hom. | Single | Het. |
|---|---|---|---|---|---|---|---|
| summeval best-of- | 100 | 9.3200 | 0.9979 | 0.9365 | 0.9446 | 0.9194 | -0.0081 |
| Component | Setting |
|---|---|
| Direct-score target | Human scores normalized to |
| Direct loss | Huber loss on the fused score |
| Pairwise target | Binary preference converted to |
| Pairwise loss | Margin-ranking loss with tuned auxiliary weight and fixed margin |
| Large sweeps | JSONL-based resume support for interrupted runs |
| Subset comparisons | Exact example-ID overlap for sample-aligned evaluation |
| Setting | Mean Sp. | Std. | vs. single |
|---|---|---|---|
| Mainline 6-judge | 0.3826 | 0.0013 | 0.0413 |
| Efficient 3-judge | 0.3638 | 0.0028 | 0.0245 |
| Direct Sp. | pair acc. | |
|---|---|---|
| 0.1 | 0.3725 | 0.0360 |
| 0.5 | 0.4192 | 0.0462 |
| 1.0 | 0.4190 | 0.0431 |
| Scale | Dataset | Base acc. | Joint acc. |
|---|---|---|---|
| 1–5 | MT-Bench-human | 0.6369 | 0.6823 |
| 1–5 | Summarize-from-Feedback | 0.5140 | 0.6054 |
| 1–10 | MT-Bench-human | 0.7200 | 0.7250 |
| 1–10 | Summarize-from-Feedback | 0.5994 | 0.6244 |
| Scale | Best single | Uniform log | Direct | Joint | Pair-only |
|---|---|---|---|---|---|
| 1–5 | 0.5482 | 0.5385 | 0.5411 | 0.5495 | 0.5112 |
| 1–10 | 0.5674 | 0.5578 | 0.5601 | 0.5507 | 0.5026 |
| Scale | Best single | Uniform log | Direct | Joint | Pair-only |
|---|---|---|---|---|---|
| 1–5 | 0.4870 | 0.4479 | 0.4917 | 0.4656 | 0.4562 |
| 1–10 | 0.5201 | 0.4870 | 0.5036 | 0.5012 | 0.5130 |
| Scale | Best single | Uniform log | Direct | Joint | Pair-only |
|---|---|---|---|---|---|
| 1–5 | 0.5605 | 0.5446 | 0.5361 | 0.5259 | 0.4952 |
| 1–10 | 0.5821 | 0.5685 | 0.5821 | 0.5657 | 0.4832 |
| Scale | Best single | Uniform log | Cov-log | JudgeMoE | Hard |
|---|---|---|---|---|---|
| 1–5 | 0.2285 | 0.2183 | 0.2638 | 0.1955 | 0.0804 |
| 1–10 | 0.2351 | 0.2650 | 0.2589 | 0.2709 | -0.0157 |
| Standalone evaluator | Cells | Standalone | JudgeMoE | JMoE wins |
|---|---|---|---|---|
| DeepSeek-V3.2 | 4 | 0.5710 | 0.4966 | 1/4 |
| Qwen2.5-72B | 4 | 0.5362 | 0.4966 | 1/4 |
| System | Active params / token | Relative compute |
|---|---|---|
| Local single judge | 7–12B dense | 0.19–0.32 |
| Local 3-judge efficient pool | 28B dense across separate passes | 0.76 |
| Local 6-judge mainline pool | 53B dense across separate passes | 1.43 |
| DeepSeek-V3.2-Exp | 37B activated, 671B total | 1.00 |
| Type | Cell | Example | Human | JudgeMoE | Cov-log | Top gate | Judge scores | Gate weights |
|---|---|---|---|---|---|---|---|---|
| + | summeval@1-5 | S1499-relevance | 0.833 | 0.825 | 0.344 | L3 | [.11,1,.50,.32,.01,.43] | [.08,.25,.35,.15,.04,.13] |
| + | helpsteer2@1-10 | H10914-helpfulness | 1.000 | 0.901 | 0.524 | G3 | [.57,.90,.52,.50,.50,.55] | [0,.98,0,.01,0,0] |
| + | wmt-mqm@1-10 | W66502 | 0.800 | 0.798 | 0.564 | G2 | [.77,.64,.54,.57,.52,.56] | [1,0,0,0,0,0] |
| + | roscoe-esnli@1-10 | R30 | 1.000 | 0.822 | 0.580 | G3 | [.56,.90,.50,.54,.49,.60] | [.18,.41,0,.17,0,.23] |
| - | topical-chat@1-5 | T12-Engaging | 0.333 | 0.898 | 0.390 | G3 | [.39,1,.51,.39,.42,.34] | [.10,.31,.15,.03,.16,.24] |
| Scale | Cand. | (valid,test) | Sel. valid | Sel. test | Test rank |
|---|---|---|---|---|---|
| 1-5 | 120 | 0.973 | 0.143 | 0.153 | 1 |
| 1-10 | 120 | 0.974 | 0.213 | 0.232 | 3 |
| Scale | Judge | Judge Sp. | MAE | Mean gate | Top-gate | Diagnostic readout |
|---|---|---|---|---|---|---|
| 1-5 | G2 | 0.141 | 0.186 | 0.373 | 46.1% | best single; largest gate mass |
| 1-5 | G3 | 0.019 | 0.256 | 0.244 | 21.7% | secondary evidence |
| 1-5 | L3 | -0.030 | 0.269 | 0.221 | 19.2% | secondary evidence |
| 1-5 | Min | 0.086 | 0.299 | 0.140 | 12.7% | secondary evidence |
| 1-5 | Q2.5 | -0.148 | 0.651 | 0.006 | 0.3% | known contaminant; gate almost removes it |
| 1-5 | Q3.5 | 0.014 | 0.400 | 0.017 | 0.0% | secondary evidence |
| Scale / subset | System or diagnostic | Spearman |
|---|---|---|
| 1–5 (full) | Best single judge | 0.1413 |
| 1–5 (full) | Top-2 sparse log pool | 0.1473 |
| 1–5 (full) | Sparse JudgeMoE (2 selected judges) | 0.1648 |
| 1–5 (full) | Full-6 uniform log pool | 0.0016 |
| 1–5 (full) | Mainline JudgeMoE | 0.0376 |
| 1–10 (full) | Best single judge | 0.1797 |