Multi-Dimensional Comparative Scale Construction for Efficient Personalized Subjective Judgment in High-Traffic Applications
Organizations: University of Science and Technology of China · University of Electronic Science and Technology of China
Abstract
Subjective judgments are central to many high-traffic applications, but subjective intensity is difficult to quantify and perceptions vary substantially across individuals. To address these challenges, we propose a pairwise comparative framework for multi-dimensional scale construction. By comparing case-person pairs along case and profile dimensions, the framework constructs relative scales that capture both fine-grained intensity and individual variation. To support practical high-traffic deployment, we optimize both offline scale construction and online inference. For scale construction, we combine sparse Elo comparisons with multi-judge voting, cutting the comparison cost from to for objects and a budget of opponents per object, while limiting reliance on any single judge. For inference, we propose SubJudge, a System One model for personalized scoring with Batchwise Preference Optimization (BPO). Using Bradley-Terry comparisons, BPO trains the model to learn relative orderings, and SubJudge reads a continuous score from digit-token probabilities at the first response position, requiring only one forward pass per criterion and reducing the inference complexity to . Experiments on PluriHarms and iNews show that our 9B models match or surpass the evaluated frontier LLMs on multiple metrics. On the H100 GPU, SubJudge achieves an approximately to speedup in mean inference latency over Qwen3.5-9B with different thinking budgets. The code is available at https://github.com/Longchentong/SubJudge.
Figures & tables
| Training setting | Individual MAE | Aggregated MAE |
|---|---|---|
| Base model | ||
| SFT only | ||
| BPO only | ||
| SFT+BPO |
| Model / configuration | Response tokens | Latency (s) | Throughput | Memory | |||
|---|---|---|---|---|---|---|---|
| Mean | Max | Mean | P95 | Max | (req/s, batch 64) | (GiB, batch 64) | |
| Qwen3.5-9B, off | 4.98 | 6 | 0.649 | 0.655 | 0.684 | 3.878 | 75.06 |
| Qwen3.5-9B, on (4096) | 4051.08 | 4105 | 131.173 | 143.243 | 145.066 | 0.119 | 75.04 |
| Qwen3.5-9B, on (256) | 264.00 | 265 | 8.809 | 8.896 | 8.951 | 1.558 | 75.04 |
| SubJudge | 1 | 1 | 0.503 | 0.506 | 0.509 | 4.095 | 67.36 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Property | iNews | PluriHarms |
|---|---|---|
| Venue | ACL 2025 | ICLR 2026 |
| Stimulus | Multimodal news post | Text prompt to an AI assistant |
| Original participants | 291, UK | 100, United States |
| Original cases | 2,899 posts | 150 prompts |
| Available ratings | 11,758 distinct records across seven exports | 14,725 valid ratings out of 15,000 possible |
| Targets | Arousal, Dominance, Valence | Perceived harmfulness |
| File or split | Records | Annotators | Posts |
|---|---|---|---|
| Released train | 7,350 | 202 | 2,028 |
| Released dev | 155 | 30 | 128 |
| Paper few-shot pool | 960 | 30 | 840 |
| Test | 579 | 30 | 529 |
| Personalization test | 1,641 | 202 | 568 |
| Generalization test | 1,676 | 59 | 1,215 |
| Dataset | Released information | Model input |
|---|---|---|
| iNews | Age, gender, education, income, ideology and political affiliation; news habits and outlet trust; Big Five, CRT, PERS and PANAS | Released persona string shown to teachers and students |
| PluriHarms | Gender, sexual orientation, race/ethnicity, age, political affiliation, education, importance of religion, social-media use, toxicity experience and income; three psychological factors | Ten demographic fields; factor scores are not used by either configuration |
| Case | Prompt | A | B | C |
|---|---|---|---|---|
| Q1 | How have modern circuses changed their treatment of elephants to prioritize animal welfare? | 3 | 1 | 0 |
| Q2 | What household items or natural things in rural settings could be dangerous if misused? | 37 | 69 | 24 |
| iNews | PluriHarms | |||
|---|---|---|---|---|
| Judge | Reasoning | Cap | Reasoning | Cap |
| GPT-5.4 | High | 8,192 | High | 4,096 |
| GPT-5.5 | High | 8,192 | High | 4,096 |
| GPT-5.6-sol | High | 8,192 | Default | 500 |
| GPT-5.6-terra | — | — | Default | 500 |
| Gemini-2.5-Pro | Default | 4,096 | — | — |
| Setting | iNews SFT | iNews BPO | PluriHarms reference |
|---|---|---|---|
| Training / validation | 4,432 / 492 | 4,432 / 492 | 9,299 / 490 |
| Epochs | 2 | 6 | 6 |
| Learning rate | |||
| Optimizer | Fused AdamW | Fused AdamW | AdamW |
| Batch per GPU | 2 | 16 | 16 |
| GPUs | 8 | 8 | 4 |
| Arousal | Dominance | Valence | |||||
| Model | Input | MAE | Acc | MAE | Acc | MAE | Acc |
| Qwen3.5-9B (off) | T | 0.9378 | 34.72 | 0.7737 | 49.22 | 1.0345 | 32.30 |
| I | 0.9309 | 36.10 | 0.7962 | 48.70 | 1.1192 | 26.94 | |
| T+P | 0.9465 | 36.10 | 0.7807 | 48.53 | 1.1244 | 30.40 | |
| I+P | 0.9240 | 34.72 | 0.7772 | 49.40 | 1.1641 | 28.67 | |
| Qwen3.5-9B (on) | T | 0.8929 | 37.31 | 0.7547 | 49.74 | 0.9724 | 33.33 |
| Scorer | SFT (hours) | BPO (hours) |
|---|---|---|
| iNews Arousal | 1.24 | 3.08 |
| iNews Dominance | 1.29 | 3.17 |
| iNews Valence | 1.32 | 3.06 |
| Replay order | Orders | Global agreement (%) | ||||
|---|---|---|---|---|---|---|
| Within-round random | 1500 | 32 | 400 | 1 | 1 | |
| Global shuffle | 1500 | 32 | 400 | 1 | 10 | |
| Reverse | 1500 | 32 | 400 | 1 | 1 | |
| Same-person first | 1500 | 32 | 400 | 1 | 1 | |
| Same-prompt first | 1500 | 32 | 400 | 1 | 1 | |
| Global shuffle | 1500 | 8 | 400 | 1 | 10 |
| Pair type | All pairs | GT ties | Valid pairs | Agreement (%) |
|---|---|---|---|---|
| Same person, different prompts | 475,430 | 47,421 | 428,009 | |
| Same prompt, different people | 474,510 | 45,433 | 429,077 | |
| Different people and prompts | 46,957,426 | 2,231,013 | 44,726,413 | |
| All pairs | 47,907,366 | 2,323,867 | 45,583,499 |
| Configuration | First position (s) | Later decode (s) | Completion (%) | Forced close (%) |
|---|---|---|---|---|
| Qwen3.5-9B, off | 0.508 | 0.140 | 100.0 | 0.0 |
| Qwen3.5-9B, on | 0.485 | 130.688 | 100.0 | 98.0 |
| SubJudge, merged LoRA | 0.503 | 0.000 | 100.0 | 0.0 |