Organizations: School of Artificial Intelligence, Jilin University · College of Software, Jilin University · International Center of Future Science, Jilin University · Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, Jilin University
Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet reliability requires more than accuracy: a model must also use the ordinal decision scale supplied by the user faithfully. We analyze JEV~1.13 and three open KEV models. Our investigation begins with ANLI, where JEV assigns 38.8% of all predictions and 51.3% of errors to Neutral despite 74.95% accuracy, nearly balanced gold labels, and balanced candidate positions. Across 36 ordinal datasets, final decisions use only 67--76% of the effective gold support, versus 87--102% on four nominal tasks. Randomizing candidate order weakens but does not remove this compression. Holding items and source scores fixed while balancing gold support and positions, we refine scales from K=2 to 14; utilization falls for every model and reaches 26--75% at K=14, although candidate probabilities remain broad for most models. Targeted BA-LoRA post-training raises gold-relative utilization from roughly 47% to 86% on eight supervised scales at both KEV sizes, showing that the compression is learned and modifiable rather than an immutable architectural limit. We call this ordinal scale-utilization bias: decision-stage candidate-space compression distinct from accuracy, gold imbalance, fixed position, and candidate count alone. The code and data are available at https://github.com/Glax147/jev_ordinal_scale_bia
Figures & tables
Figure 1: The observation that motivated our bias analysis. JEV 1.13 favors Neutral on ANLI although gold labels are nearly balanced. Label shares among gold labels, all predictions, and incorrect predictions on the 6,400 development and test items.
Ordinal (36 datasets)
Nominal (4 datasets)
Model
Acc ↑
Uarg
DR↓
TVD ↓
nMAE ↓
QWK ↑
Usoft
Margin
Acc ↑
DR↓
TVD ↓
JEV 1.13
37.67
72.97
24.24
29.82
20.78
0.593
89.10
36.16
84.39
1.47
6.04
KEV-0.8B
28.80
66.87
31.11
37.43
30.47
0.370
97.10
17.49
69.63
13.37
18.58
KEV-4B
33.37
65.32
32.80
37.48
24.70
0.455
93.61
20.00
79.74
1.85
9.89
KEV-9B
35.02
70.46
26.84
33.21
22.17
0.523
94.63
18.46
81.80
1.69
6.94
Table 1: Ordinal tasks show a consistent scale-utilization bias, while the bias is much weaker on nominal tasks. Macro-averages per task type; DR=100∣R−1∣ is the absolute support deviation in percentage points. Bold marks the highest accuracy/QWK and lowest DR /TVD/nMAE; the remaining columns are diagnostic and are not ranked.
Acc ↑
TVD ↓
DR↓
Rand. 1 vs. 2
Model
Orig.
Rand.
Orig.
Rand.
Orig.
Rand.
Agree. ↑
Shift ↓
Ordinal
JEV 1.13
26.99
26.88
35.61
32.25
34.99
29.65
74.43
9.42
KEV-0.8B
22.19
21.68
41.87
31.41
38.94
26.28
58.58
9.65
KEV-4B
24.56
24.92
45.73
29.30
44.70
26.06
58.14
10.13
KEV-9B
25.79
25.30
45.39
35.05
43.01
32.31
59.94
9.47
Nominal
JEV 1.13
89.43
89.55
4.13
3.54
0.94
0.66
98.21
2.36
Table 2: Random order weakens but does not remove ordinal scale-utilization bias; accuracy barely changes. Six ordinal and three nominal datasets. DR=100∣R−1∣ is the absolute support deviation in percentage points. Rand.: mean of two random orders; Agree./Shift: their label agreement and probability TVD.
Figure 2: Candidate order modulates but does not explain the bias. Random order improves ordinal R and changes many item-level decisions but barely changes accuracy. (a) Ordinal R ; (b) agreement between random orders; (c) accuracy change from the original order.
Figure 3: Finer ordinal scales intensify scale-utilization bias on the same items. Under equal-frequency binning, endpoint R and top-1 margins are lower at K=14 than at K=2 for every model. Macro-averages over three ordinal sources. (d) uses equal-frequency binning; dashed: DBpedia reference.
Model
Evaluation group
D
Acc. ↑
Uarg↑
DR↓
TVD ↓
nMAE ↓
QWK ↑
KEV-0.8B
Training-source ordinal
8
15.5 → 28.5
46.3 → 85.0
52.6 → 13.2
54.7 → 19.7
37.9 → 20.4
.188 → .600
Unseen-source ordinal
28
32.6 → 33.4
72.8 → 69.7
25.0 → 27.3
32.5 → 35.1
28.3 → 25.1
.422 → .456
All ordinal
36
28.8 → 32.3
66.9 → 73.1
31.1 → 24.2
37.4 → 31.7
30.5 → 24.0
.370 → .488
Nominal
4
69.6 → 72.7
83.4 → 90.2
13.4 → 4.2
18.6 → 12.8
–
–
KEV-4B
Training-source ordinal
8
17.3 → 32.7
45.7 → 83.9
53.6 → 14.4
52.9 → 19.0
32.5 → 17.8
.240 → .664
Unseen-source ordinal
28
38.0 → 39.4
70.9 → 78.0
26.8 → 18.7
33.1 → 27.6
22.5 → 19.7
.517 → .590
Table 3: BA-LoRA strongly mitigates compression on the supervised sources, while transfer differs by model scale. Before → after dataset-level macro-averages on the complete 40-dataset panel. D is the number of datasets; DR=100∣R−1∣ is the absolute support deviation in percentage points. Dashes denote metrics that require an ordered label space. Evaluation records from the eight training-source datasets are disjoint from post-training records by source ID and normalized-text hash.
Table 4: The 40 datasets in the main evaluation by task type and number of candidates K .
K
Equal-width examples
Equal-frequency examples
2
Level 01 of 02: 0≤s<0.5 ; Level 02 of 02: 0.5≤s≤1
Level 01 of 02: 0≤s<0.503333 ; Level 02 of 02: 0.503333≤s≤1
5
Level 01 of 05: 0≤s<0.2 ; Level 03 of 05: 0.4≤s<0.6 ; Level 05 of 05: 0.8≤s≤1
Level 01 of 05: 0≤s<0.176357 ; Level 03 of 05: 0.402703≤s<0.598387 ; Level 05 of 05: 0.801316≤s≤1
10
Level 01 of 10: 0≤s<0.1 ; Level 05 of 10: 0.4≤s<0.5 ; Level 10 of 10: 0.9≤s≤1
Level 01 of 10: 0≤s<0.101282 ; Level 05 of 10: 0.402703≤s<0.503333 ; Level 10 of 10: 0.893778≤s≤1
14
Level 01 of 14: 0≤s<0.071429 ; Level 07 of 14: 0.428571≤s<0.5 ; Level 14 of 14: 0.928571≤s≤1
Level 01 of 14: 0≤s<0.101282 ; Level 07 of 14: 0.426569≤s<0.503333 ; Level 14 of 14: 0.909243≤s≤1
Appendix
Table 5: Controlled ordinal candidate construction. (a) Representative candidate strings for Civil Comments Toxicity; for K>2 , the first, one central, and the final level are shown, while prompts contain all K candidates in random order. (b) Minimum–maximum items per level and gold effective-label coverage G (1,400 items per condition). EF: equal-frequency; EW: equal-width.
Configuration
KEV-0.8B
KEV-4B
Initialization
KEV-0.8B
KEV-4B
LoRA rank / α / dropout
16 / 32 / 0.05
16 / 32 / 0.05
Learning rate
2×10−5
1.5×10−5
Micro-batch size
32
8
Gradient accumulation
2
8
Effective batch size
64
64
Appendix
Table 7: Model-specific settings for the reported bias-aware post-training runs. The two models share the same effective batch size, training duration, adapter configuration, precision, and regularization; only the learning rate and memory-dependent micro-batch/accumulation pair differ.
Figure 4: Per-dataset candidate-space utilization in the 40-dataset main evaluation. (a) R per dataset and model; rows are ordered by task type, K (after each name), and mean R , and the horizontal line separates ordinal from nominal datasets. (b) Number of ordinal datasets whose modal predicted level lies in the lower, middle, or upper third of the scale.
Figure 5: Per-dataset changes after BA-LoRA on the 15 largest KEV-0.8B R gains. Daggers mark the eight post-training sources; the remaining six ordinal datasets and SCOTUS13 are unseen sources. Each cell reports the after-minus-before change for KEV-0.8B and KEV-4B. This post-hoc subset illustrates where mitigation is strongest; Table 3 reports the complete-panel aggregates used for inference.
Large language models are increasingly used for ordinal classification, yet semantically equivalent changes to prompt organization can alter their predictions. We conduct systematic experiments to characterize positional bias from label order, demonstration order, and demonstration placement. First, we apply the three probes to ten frontier LLMs on a common ordinal-classification task; every model is sensitive to all three positional sources, showing that the problem is pervasive. Second, we vary eight prompt-, task-, and model-level factors across five datasets; accuracy and stability are often misaligned, and only lower scale cardinality consistently improves both. Third, we compare pointwise, pairwise, and listwise inference, alternative aggregation and debiasing methods, and joint configurations; the tested corrections do not provide a reliable remedy, while a comparison-based listwise formulation offers the best balance but transfers unevenly across models and bias sources. These findings show that positional robustness depends on the full system configuration rather than the model alone. Ordinal-classification systems should therefore be selected jointly for predictive performance and stability.
Yu Wang, Zhe Zhou, Menglin Liu +1
Cornell University · University of Washington · The Chinese University of Hong Kong, Shenzhen +1
Dedicated decision models such as Jev map unstructured language to probability distributions over finite choices, allowing their outputs to directly route requests, select tools, and trigger actions. Yet real-world inputs rarely arrive in isolation: they come with background details and surrounding context. We find that short additions that fit naturally into this context can nevertheless redirect an otherwise correct decision, even when the correct answer remains unchanged. To study this behavior, we fix a wrong target option for each initially correct item and use the model's option probabilities to refine fluent context additions while preserving the source, question, choices, and gold answer. Within 64 accepted target evaluations, the optimizer identifies contexts that redirect Jev on 312 of 508 initially correct decisions (61.4%); in 229 cases, Jev assigns at least 0.7 probability to the fixed wrong option. Across seven datasets, three additional decision systems show targeted flip rates of 64.9%-73.2% on decisions they initially answer correctly. Taken together, these results expose a pronounced fragility in current decision models: short, ordinary-looking context can shift a correct choice to a high-confidence wrong one. Because these models turn language directly into downstream choices, this sensitivity raises concerns about treating their probability outputs as reliable decision interfaces.
LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict must be derived, as in math, code, and logic. Its confidence marks this boundary. With a threshold frozen in advance, accepting confident verdicts and escalating the rest is 0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of its fee, and in a pre-specified live test on two new workloads the cascade matches GPT-6's accuracy exactly. Confidence routing weakens on style-adversarial pairs and reference-free prose; we close with a simple recipe for validating thresholds locally.