cs.CLOct 5, 2026

JudgeMoE: Distributional Aggregation for LLM-as-a-Judge

Authors: Yiqi Liu, Joseph James, Yang Wang, Kun Zhao, Chenghao Xiao, Chenghua Lin

Organizations: The University of Manchester · The University of Sheffield · University of Pittsburgh · Shanghai University of Finance and Economics

Abstract

When an LLM judge scores an output, its score distribution retains uncertainty and disagreement information that is lost after scalar compression. We introduce JudgeMoE, a lightweight aggregator that assigns example-specific weights to cached judge score distributions and fuses them before computing a final score. A protocol study shows that score-range choice is unstable across judge--dataset settings and that soft scoring usually outperforms hard decoding. On the original 10-cell benchmark, JudgeMoE improves mean Spearman over uniform log pooling by +0.079+0.079. Applying the same configuration to six additional cells yields a +0.0393+0.0393 mean gain over the strongest local single judge across 16 cells, with positive differences in 12/16 cells and a one-sided Wilcoxon signed-rank p=0.0091p=0.0091. Validation-based analyses further show that the preferred aggregation method depends on the task and judge pool.

Figures & tables

Appendix figures & tables51 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation

    Feb 9, 2026Jitian Zhao, Changho Shin, Tzu-Heng Huang +2Llm-As-A-JudgeLarge Language Model Evaluation

  2. Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

    May 25, 2026Delip Rao, Chris Callison-BurchLarge Language Model JudgesJudgement

  3. LLM Judge Validation Under Sparse Overlap: From Inference to Design

    Sep 25, 2026Junxuan Li, Arko Mukherjee, Soumyabrata PalLarge Language Model JudgesToken Budget Allocation