cs.AISep 30, 2026

More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models

Authors: Tianxiang Gao, Jinzhe Li, Zhiyuan Li, Yi Chang, Yuan Wu

Organizations: School of Artificial Intelligence, Jilin University · College of Software, Jilin University · International Center of Future Science, Jilin University · Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, Jilin University

Abstract

Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet reliability requires more than accuracy: a model must also use the ordinal decision scale supplied by the user faithfully. We analyze JEV~1.13 and three open KEV models. Our investigation begins with ANLI, where JEV assigns 38.8% of all predictions and 51.3% of errors to Neutral despite 74.95% accuracy, nearly balanced gold labels, and balanced candidate positions. Across 36 ordinal datasets, final decisions use only 67--76% of the effective gold support, versus 87--102% on four nominal tasks. Randomizing candidate order weakens but does not remove this compression. Holding items and source scores fixed while balancing gold support and positions, we refine scales from K=2K=2 to 1414; utilization falls for every model and reaches 26--75% at K=14K=14, although candidate probabilities remain broad for most models. Targeted BA-LoRA post-training raises gold-relative utilization from roughly 47% to 86% on eight supervised scales at both KEV sizes, showing that the compression is learned and modifiable rather than an immutable architectural limit. We call this ordinal scale-utilization bias: decision-stage candidate-space compression distinct from accuracy, gold imbalance, fixed position, and candidate count alone. The code and data are available at https://github.com/Glax147/jev_ordinal_scale_bia

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Are LLMs Positionally Consistent Ordinal Classifiers? A Systematic Evaluation

    Aug 9, 2026Yu Wang, Zhe Zhou, Menglin Liu +1Position BiasOrder Matters

  2. JevOut: Natural Context Can Flip Decision Models

    Sep 24, 2026Zixiang XuDecisionsContextual Reasoning

  3. JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

    Sep 22, 2026Yubo Li, Yidi Miao, Ramayya Krishnan +1Large Language Model JudgesJudges