cs.AISep 27, 2026

Multi-Dimensional Comparative Scale Construction for Efficient Personalized Subjective Judgment in High-Traffic Applications

Authors: Xianglong Shi, Shifeng Liu, Sirui Zhao, Shengming Yuan, Enhong Chen

Organizations: University of Science and Technology of China · University of Electronic Science and Technology of China

Abstract

Subjective judgments are central to many high-traffic applications, but subjective intensity is difficult to quantify and perceptions vary substantially across individuals. To address these challenges, we propose a pairwise comparative framework for multi-dimensional scale construction. By comparing case-person pairs along case and profile dimensions, the framework constructs relative scales that capture both fine-grained intensity and individual variation. To support practical high-traffic deployment, we optimize both offline scale construction and online inference. For scale construction, we combine sparse Elo comparisons with multi-judge voting, cutting the comparison cost from O(N2)O(N^2) to O(NK)O(NK) for NN objects and a budget of KK opponents per object, while limiting reliance on any single judge. For inference, we propose SubJudge, a System One model for personalized scoring with Batchwise Preference Optimization (BPO). Using Bradley-Terry comparisons, BPO trains the model to learn relative orderings, and SubJudge reads a continuous score from digit-token probabilities at the first response position, requiring only one forward pass per criterion and reducing the inference complexity to O(1)O(1). Experiments on PluriHarms and iNews show that our 9B models match or surpass the evaluated frontier LLMs on multiple metrics. On the H100 GPU, SubJudge achieves an approximately 1.29×1.29\times to 261×261\times speedup in mean inference latency over Qwen3.5-9B with different thinking budgets. The code is available at https://github.com/Longchentong/SubJudge.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Ask the Right Comparison:Bias-Aware Bayesian Active Top-kk Ranking with LLM Judges

    Jul 2, 2026Jian Xu, Delu Zeng, John Paisley +1Large Language Model JudgesPosition Bias

  2. RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

    May 20, 2026Zhenwei Tang, Zhaoyan Liu, Rasa Hosseinzadeh +3Llm-As-A-JudgeChatbots

  3. RouteJudge: An Open Platform for Reproducible and Preference-Aware LLM Routing

    Jun 17, 2026Guannan Lai, Haoran Hu, Han-Jia YeLarge Language Model Routing