stat.MLMar 11, 2026

MultiwayPAM: Multiway Partitioning Around Medoids for LLM-as-a-Judge Score Analysis

Authors: Chihiro Watanabe, Jingyu Sun

Organizations: NTT Computer and Data Science Laboratories, 3-9-11, Midori-cho, Musashino-shi, Tokyo, Japan

Abstract

LLM-as-a-Judge is a flexible framework for text evaluation, which allows us to obtain scores for the quality of a given text from various perspectives by changing the prompt template. Two main challenges in using LLM-as-a-Judge are computational cost of inference using a large language model (LLM), especially when evaluating a large number of instances, and inherent bias of an LLM evaluator. To address these issues and reveal the structure of score bias caused by an LLM evaluator, we propose to apply a tensor clustering method to a given LLM-as-a-Judge score tensor, whose entries are the scores for different combinations of questions, answerers, and evaluators. Specifically, we develop a new tensor clustering method MultiwayPAM, with which we can simultaneously estimate the cluster membership and the medoids for each mode of a given data tensor. By observing the medoids obtained by MultiwayPAM, we can gain knowledge about the membership of each question/answerer/evaluator cluster. We experimentally show the effectiveness of MultiwayPAM by applying it to the score tensors for two practical datasets.

Figures & tables

Explore similar work

CardsList
  1. CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation

    Feb 9, 2026Jitian Zhao, Changho Shin, Tzu-Heng Huang +2Llm-As-A-JudgeLarge Language Model Evaluation

  2. Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation

    Aug 6, 2026Yuma Asato, Kiyoaki Shirai, Natthawut KertkeidkachornLarge Language Model BiasLlm-As-A-Judge

  3. MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge

    Apr 20, 2026Sua Lee, Sanghee Park, Jinbae ImLlm-As-A-JudgeMultimodal Large Language Models