Stance detection requires identifying an author's attitude toward a given target, sometimes based on conversational context. Jev, a specialized decision model designed for structured decision-making, offers an alternative to general-purpose large language models (LLMs). In this work, we evaluate Jev on two stance detection datasets, VAST (English texts) and ZS-CSD (Chinese conversations), comparing it with four general-purpose LLMs and two fine-tuned models. Results show that Jev achieves competitive performance on VAST, matching GPT-5.6 and outperforming the other general-purpose LLMs. However, it falls behind stronger LLMs on ZS-CSD, particularly in distinguishing favor from against. Further analysis suggests that this limitation may be related to understanding reply relationships and stance direction rather than conversation length alone. These findings highlight both the potential and limitations of Jev for stance detection.
Figures & tables
Dataset
Model / setting
Accuracy
Macro-F1
Input tokens
Output tokens
Cost
VAST
Jev
78.14
77.92
1,763,143
114,228
$0.0741
DeepSeek V4.1 Flash
74.82
73.99
865,478
18,708
$0.2133
Qwen3.8-Flash
74.52
73.79
763,609
28,874
$0.1021
GPT-5.6
78.51
78.08
1,178,659
33,439
$3.9010
Claude Sonnet 5
74.82
74.51
2,279,907
251,109
$4.7051
Qwen2.5-3B-Instruct (SFT) †
74.16
74.52
—
—
—
Table 1: Performance and resource usage on the full test sets of VAST and ZS-CSD. Bold indicates the best results among prompted models. Dashes denote unavailable measurements. † SFT results are from single runs without per-instance predictions.
Dataset
Model
Against
Favor
Neutral
VAST
Jev
77.13
72.71
83.92
DeepSeek
74.15
66.83
81.00
Qwen
73.88
65.52
81.97
GPT-5.6
78.25
73.95
82.05
Claude
75.35
70.26
77.91
ZS-CSD
Jev
56.66
59.58
59.04
Table 2: Per-class F1 (%). Bold indicates the best result for each class within each dataset.
Figure 1: Macro-F1 differences between Jev and each LLM (Jev minus LLM, in percentage points), with paired target-bootstrap 95% intervals.
Target type
N
Jev
DeepSeek
Qwen
GPT-5.6
Claude
Noun phrase
967
63.49
61.61
60.80
78.02
76.19
Claim
1,617
55.26
54.08
48.58
70.72
70.13
Table 3: Macro-F1 scores (%) across different target types on ZS-CSD. Bold indicates the best result for each target type.
Turns
N
Jev
DeepSeek
Qwen
GPT-5.6
Claude
1
45
66.67
53.55
62.69
64.34
59.80
2
445
56.99
52.78
51.37
79.68
74.77
3
877
57.45
54.27
50.52
74.95
70.44
4
585
58.11
57.27
50.83
69.05
70.77
5
351
54.64
57.55
52.65
68.23
69.97
6+
281
54.00
55.80
57.92
68.20
70.06
Table 4: Macro-F1 scores (%) across different conversation lengths on ZS-CSD. Bold indicates the best result for each group.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
System
Fav.–Ag.
Fav.–Neu.
Ag.–Neu.
Jev
373
304
398
DeepSeek V4.1 Flash
275
368
435
Qwen3.8-Flash
462
232
479
GPT-5.6
35
301
319
Claude Sonnet 5
101
306
303
Appendix
Table 5: Number of ZS-CSD test records whose reference and predicted labels form each pair, in either direction. Bold indicates the fewest errors for each label pair.
Stance detection identifies the attitude of a text author toward a given target. Recent studies have explored various LLM-based strategies for this task, from zero-shot prompting to multi-agent debate. However, existing works differ in data splits, base models, and evaluation protocols, making fair comparison difficult. We conduct a systematic comparison that evaluates five methods across two categories -- prompt-based inference (Direct Prompting, Auto-CoT, StSQA) and agent-based debate (COLA, MPRF) -- on four datasets with 14 subtasks, using 15 LLMs from six model families with parameter sizes from 7B to 72B+. Our experiments yield several findings. First, on all models with complete results, the best prompt-based method outperforms the best agent-based method, while agent methods require 7 to 12 times more API calls per sample. Second, model scale has a larger impact on performance than method choice, with gains plateauing around 32B. Third, reasoning-enhanced models (DeepSeek-R1) do not consistently outperform general models of the same size on this task.
Genan Dai, Zini Chen, Yi Yang +1
School of Artificial Intelligence, Shenzhen Technology University, Shenzhen, China
Jev is a "System One" model that returns a choice among given options instead of generating text. We study how such a specialized decision model compares with general-purpose large language models (LLMs). We evaluate Jev on 13 multiple-choice benchmarks covering knowledge, reasoning, and multilingual understanding, and compare it with 19 LLMs in three tiers: frontier, representative, and small. Jev is competitive with frontier LLMs on knowledge and commonsense benchmarks and obtains the best score on MMLU-Redux and ARC-Challenge. Outside mathematics, it also outperforms most representative LLMs and all small LLMs. However, it falls behind on mathematical word problems: on MathQA, it is 17.7 points below the frontier median and scores lower than all 19 LLMs. These results indicate that a specialized decision model can match general-purpose LLMs on decisions that rely mainly on knowledge, but not on decisions that require multi-step calculation.
Xing Li, Qingcheng Chang, Jinzhong Ning +5
Dalian Maritime University · Dalian University of Technology
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM-as-Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM-as-Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.