Organizations: School of Artificial Intelligence, Jilin University · College of Software, Jilin University · International Center of Future Science, Jilin University · Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, Jilin University
Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet reliability requires more than accuracy: a model must also use the ordinal decision scale supplied by the user faithfully. We analyze JEV~1.13 and three open KEV models. Our investigation begins with ANLI, where JEV assigns 38.8% of all predictions and 51.3% of errors to Neutral despite 74.95% accuracy, nearly balanced gold labels, and balanced candidate positions. Across 36 ordinal datasets, final decisions use only 67--76% of the effective gold support, versus 87--102% on four nominal tasks. Randomizing candidate order weakens but does not remove this compression. Holding items and source scores fixed while balancing gold support and positions, we refine scales from K=2 to 14; utilization falls for every model and reaches 26--75% at K=14, although candidate probabilities remain broad for most models. Targeted BA-LoRA post-training raises gold-relative utilization from roughly 47% to 86% on eight supervised scales at both KEV sizes, showing that the compression is learned and modifiable rather than an immutable architectural limit. We call this ordinal scale-utilization bias: decision-stage candidate-space compression distinct from accuracy, gold imbalance, fixed position, and candidate count alone. The code and data are available at https://github.com/Glax147/jev_ordinal_scale_bia
Figures & tables
Figure 1: The observation that motivated our bias analysis. JEV 1.13 favors Neutral on ANLI although gold labels are nearly balanced. Label shares among gold labels, all predictions, and incorrect predictions on the 6,400 development and test items.
Ordinal (36 datasets)
Nominal (4 datasets)
Model
Acc ↑
Uarg
DR↓
TVD ↓
nMAE ↓
QWK ↑
Usoft
Margin
Acc ↑
DR↓
TVD ↓
JEV 1.13
37.67
72.97
24.24
29.82
20.78
0.593
89.10
36.16
84.39
1.47
6.04
KEV-0.8B
28.80
66.87
31.11
37.43
30.47
0.370
97.10
17.49
69.63
13.37
18.58
KEV-4B
33.37
65.32
32.80
37.48
24.70
0.455
93.61
20.00
79.74
1.85
9.89
KEV-9B
35.02
70.46
26.84
33.21
22.17
0.523
94.63
18.46
81.80
1.69
6.94
Table 1: Ordinal tasks show a consistent scale-utilization bias, while the bias is much weaker on nominal tasks. Macro-averages per task type; DR=100∣R−1∣ is the absolute support deviation in percentage points. Bold marks the highest accuracy/QWK and lowest DR /TVD/nMAE; the remaining columns are diagnostic and are not ranked.
Acc ↑
TVD ↓
DR↓
Rand. 1 vs. 2
Model
Orig.
Rand.
Orig.
Rand.
Orig.
Rand.
Agree. ↑
Shift ↓
Ordinal
JEV 1.13
26.99
26.88
35.61
32.25
34.99
29.65
74.43
9.42
KEV-0.8B
22.19
21.68
41.87
31.41
38.94
26.28
58.58
9.65
KEV-4B
24.56
24.92
45.73
29.30
44.70
26.06
58.14
10.13
KEV-9B
25.79
25.30
45.39
35.05
43.01
32.31
59.94
9.47
Nominal
JEV 1.13
89.43
89.55
4.13
3.54
0.94
0.66
98.21
2.36
Table 2: Random order weakens but does not remove ordinal scale-utilization bias; accuracy barely changes. Six ordinal and three nominal datasets. DR=100∣R−1∣ is the absolute support deviation in percentage points. Rand.: mean of two random orders; Agree./Shift: their label agreement and probability TVD.
Figure 2: Candidate order modulates but does not explain the bias. Random order improves ordinal R and changes many item-level decisions but barely changes accuracy. (a) Ordinal R ; (b) agreement between random orders; (c) accuracy change from the original order.
Figure 3: Finer ordinal scales intensify scale-utilization bias on the same items. Under equal-frequency binning, endpoint R and top-1 margins are lower at K=14 than at K=2 for every model. Macro-averages over three ordinal sources. (d) uses equal-frequency binning; dashed: DBpedia reference.
Model
Evaluation group
D
Acc. ↑
Uarg↑
DR↓
TVD ↓
nMAE ↓
QWK ↑
KEV-0.8B
Training-source ordinal
8
15.5 → 28.5
46.3 → 85.0
52.6 → 13.2
54.7 → 19.7
37.9 → 20.4
.188 → .600
Unseen-source ordinal
28
32.6 → 33.4
72.8 → 69.7
25.0 → 27.3
32.5 → 35.1
28.3 → 25.1
.422 → .456
All ordinal
36
28.8 → 32.3
66.9 → 73.1
31.1 → 24.2
37.4 → 31.7
30.5 → 24.0
.370 → .488
Nominal
4
69.6 → 72.7
83.4 → 90.2
13.4 → 4.2
18.6 → 12.8
–
–
KEV-4B
Training-source ordinal
8
17.3 → 32.7
45.7 → 83.9
53.6 → 14.4
52.9 → 19.0
32.5 → 17.8
.240 → .664
Unseen-source ordinal
28
38.0 → 39.4
70.9 → 78.0
26.8 → 18.7
33.1 → 27.6
22.5 → 19.7
.517 → .590
Table 3: BA-LoRA strongly mitigates compression on the supervised sources, while transfer differs by model scale. Before → after dataset-level macro-averages on the complete 40-dataset panel. D is the number of datasets; DR=100∣R−1∣ is the absolute support deviation in percentage points. Dashes denote metrics that require an ordered label space. Evaluation records from the eight training-source datasets are disjoint from post-training records by source ID and normalized-text hash.
Table 4: The 40 datasets in the main evaluation by task type and number of candidates K .
K
Equal-width examples
Equal-frequency examples
2
Level 01 of 02: 0≤s<0.5 ; Level 02 of 02: 0.5≤s≤1
Level 01 of 02: 0≤s<0.503333 ; Level 02 of 02: 0.503333≤s≤1
5
Level 01 of 05: 0≤s<0.2 ; Level 03 of 05: 0.4≤s<0.6 ; Level 05 of 05: 0.8≤s≤1
Level 01 of 05: 0≤s<0.176357 ; Level 03 of 05: 0.402703≤s<0.598387 ; Level 05 of 05: 0.801316≤s≤1
10
Level 01 of 10: 0≤s<0.1 ; Level 05 of 10: 0.4≤s<0.5 ; Level 10 of 10: 0.9≤s≤1
Level 01 of 10: 0≤s<0.101282 ; Level 05 of 10: 0.402703≤s<0.503333 ; Level 10 of 10: 0.893778≤s≤1
14
Level 01 of 14: 0≤s<0.071429 ; Level 07 of 14: 0.428571≤s<0.5 ; Level 14 of 14: 0.928571≤s≤1
Level 01 of 14: 0≤s<0.101282 ; Level 07 of 14: 0.426569≤s<0.503333 ; Level 14 of 14: 0.909243≤s≤1
Appendix
Table 5: Controlled ordinal candidate construction. (a) Representative candidate strings for Civil Comments Toxicity; for K>2 , the first, one central, and the final level are shown, while prompts contain all K candidates in random order. (b) Minimum–maximum items per level and gold effective-label coverage G (1,400 items per condition). EF: equal-frequency; EW: equal-width.
Configuration
KEV-0.8B
KEV-4B
Initialization
KEV-0.8B
KEV-4B
LoRA rank / α / dropout
16 / 32 / 0.05
16 / 32 / 0.05
Learning rate
2×10−5
1.5×10−5
Micro-batch size
32
8
Gradient accumulation
2
8
Effective batch size
64
64
Appendix
Table 7: Model-specific settings for the reported bias-aware post-training runs. The two models share the same effective batch size, training duration, adapter configuration, precision, and regularization; only the learning rate and memory-dependent micro-batch/accumulation pair differ.
Figure 4: Per-dataset candidate-space utilization in the 40-dataset main evaluation. (a) R per dataset and model; rows are ordered by task type, K (after each name), and mean R , and the horizontal line separates ordinal from nominal datasets. (b) Number of ordinal datasets whose modal predicted level lies in the lower, middle, or upper third of the scale.
Figure 5: Per-dataset changes after BA-LoRA on the 15 largest KEV-0.8B R gains. Daggers mark the eight post-training sources; the remaining six ordinal datasets and SCOTUS13 are unseen sources. Each cell reports the after-minus-before change for KEV-0.8B and KEV-4B. This post-hoc subset illustrates where mitigation is strongest; Table 3 reports the complete-panel aggregates used for inference.