Jev is a "System One" model that returns a choice among given options instead of generating text. We study how such a specialized decision model compares with general-purpose large language models (LLMs). We evaluate Jev on 13 multiple-choice benchmarks covering knowledge, reasoning, and multilingual understanding, and compare it with 19 LLMs in three tiers: frontier, representative, and small. Jev is competitive with frontier LLMs on knowledge and commonsense benchmarks and obtains the best score on MMLU-Redux and ARC-Challenge. Outside mathematics, it also outperforms most representative LLMs and all small LLMs. However, it falls behind on mathematical word problems: on MathQA, it is 17.7 points below the frontier median and scores lower than all 19 LLMs. These results indicate that a specialized decision model can match general-purpose LLMs on decisions that rely mainly on knowledge, but not on decisions that require multi-step calculation.
Figures & tables
Category
Benchmark
Lang.
Options
Knowledge
MMLU-Pro
En
10
MMLU-Redux
En
4
MMLU
En
4
GPQA Diamond
En
4
C-Eval
Zh
4
HellaSwag
En
4
Table 1: Benchmarks used in this study, with their language (or number of languages) and number of answer options.
Figure 1: Jev and three tiers of LLMs on 13 benchmarks. Each marker is the score of one model, and the black line marks Jev. Filled markers are scores evaluated in this work; hollow markers are published scores (Appendix C ).
Knowledge
Reasoning
Multilingual
Model
Pro
Redux
MMLU
GPQA
C-Eval
HSwag
Wino
TQA
ARC-C
MathQA
AQuA
ProX
MMMLU
Jev
82.3
95.0
91.3
74.2
85.3
89.3
91.8
87.5
97.6
75.0
83.1
83.2
85.0
Frontier LLMs
GPT-5.6 Sol †
84.5
90.0
90.7
94.6
89.6
89.8
70.2
76.0
95.1
92.0
89.0
68.9
85.0
Claude Opus 5.5 †
86.6
94.2
92.6
75.8
83.1
96.8
92.7
92.0
94.8
85.8
89.8
82.1
87.2
Gemini 3.8 Flash †
93.8
93.8
94.4
95.3
94.7
97.5
97.1
97.9
97.5
92.7
91.3
87.6
92.6
Table 2: Accuracy (%) of Jev and 19 LLMs across 13 benchmarks, grouped into three model tiers. Scores in italics are published scores; all others are evaluated in this work. Bold marks the best score in each column. ∗ Reported with the MC2 metric. † Closed-source. Pro = MMLU-Pro; Redux = MMLU-Redux; GPQA = GPQA Diamond; HSwag = HellaSwag; Wino = WinoGrande; TQA = TruthfulQA; ARC-C = ARC-Challenge; AQuA = AQuA-RAT; ProX = MMLU-ProX.
Figure 2: Difference between Jev and the frontier median on each benchmark. Gray bands show the range of the six frontier LLMs. ∗ On GPQA Diamond, the median and range cover only the three frontier LLMs evaluated in this work.
Benchmark
Ours
Deußer et al.
Items
MMLU
91.3
91.8
14,042
C-Eval
85.3
83.9
12,342
HellaSwag
89.3
95.5
10,042
WinoGrande
91.8
91.4
1,267
ARC
97.6
98.8
3,548
Table 3: Accuracy (%) of Jev in this work and in Deußer et al. (2026) , with the number of items in their evaluation. Our ARC score is on ARC-Challenge; theirs is on ARC-Easy and ARC-Challenge.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Jev
Median
Range
Δ
Rank
TruthfulQA
87.5
83.0
76.0–97.9
+4.5
3/7
ARC-Challenge
97.6
95.5
91.4–97.5
+2.1
1/7
MMLU-Redux
95.0
93.2
90.0–94.3
+1.8
1/7
WinoGrande
91.8
92.1
70.2–97.1
− 0.3
4/7
MMLU
91.3
92.5
87.5–94.4
− 1.2
5/7
MMLU-ProX
83.2
84.5
68.9–87.6
− 1.3
5/7
Appendix
Table 4: Jev against the six frontier LLMs: the median and range of their scores, the difference Δ between Jev and the median, and the rank of Jev among the evaluated models. On GPQA Diamond, only the three frontier LLMs evaluated in this work are counted.
Benchmark
Jev
GPT-4o
Claude 3.5 Sonnet
DeepSeek-R1
DeepSeek-V3
Qwen3-32B
Qwen2.5-72B
Llama-405B
MMLU-Pro
82.3
72.6 a
78.0 a
84.0 c
81.2 b
79.1 e
71.6 a
73.3 a
MMLU-Redux
95.0
88.0 a
88.9 a
92.9 c
90.4 e
90.9 d
85.6 a
86.2 a
MMLU
91.3
87.2 a
88.3 a
90.8 c
88.5 a
81.0 h
85.3 a
88.6 a
GPQA Diamond
74.2
49.9 a
65.0 a
71.5 c
59.1 a
68.4 d
49.0 a
51.1 a
C-Eval
85.3
76.0 a
76.7 a
91.8 c
86.5 a
87.3 d
86.1 a
61.5 a
HellaSwag
89.3
92.6
88.6
91.5
89.2
71.1 h
87.2
88.3 h
Appendix
Table 5: Jev and the representative LLMs. Bold marks the best score on each benchmark. Scores in italics are published, and their superscripts refer to the sources in Table 7 ; the other scores are evaluated in this work. ∗ MC2 metric, excluded from the last two rows. Qwen2.5-72B and Llama-405B are the Instruct models.
Benchmark
Jev
Qwen3-8B
Qwen3-4B
Llama-3.1-8B
Phi-4-mini
Gemma-3-4B
Mistral-7B
MMLU-Pro
82.3
74.6 e
70.4 e
48.3 i
52.8 j
43.6 k
31.5 l
MMLU-Redux
95.0
87.5 d
83.7 d
61.7 d
73.1
71.1
58.4 l
MMLU
91.3
72.0 h
66.8 h
69.4 i
67.3 j
58.4 h
63.5 l
GPQA Diamond
74.2
62.0 d
55.9 d
32.8 d
32.0
30.8 k
30.2
C-Eval
85.3
83.4 d
77.5 d
52.0 d
61.4
59.5
45.9 l
HellaSwag
89.3
56.5 h
52.8 h
73.5
69.1 j
75.0 h
75.8 l
Appendix
Table 6: Jev and the small LLMs. Bold marks the best score on each benchmark. Scores in italics are published, and their superscripts refer to the sources in Table 7 ; the other scores are evaluated in this work. ∗ MC2 metric, excluded from the last two rows. Llama-3.1-8B and Phi-4-mini are the Instruct models, Gemma-3-4B is the IT model, and Mistral-7B is the v0.3 base model.
Key
Source
Models and protocol
a
DeepSeek-V3 report ( DeepSeek-AI, 2024 )
Chat models: GPT-4o (0513), Claude 3.5 Sonnet (1022), Qwen2.5-72B-Instruct, Llama-3.1-405B-Instruct, and the original DeepSeek-V3.
b
DeepSeek update notes ( DeepSeek, 2026 )
DeepSeek-V3-0324 on MMLU-Pro; DeepSeek-V4.1-Flash on GPQA Diamond.
c
DeepSeek-R1 report ( DeepSeek-AI, 2025 )
Original DeepSeek-R1.
d
Qwen3 report ( Yang et al., 2025 )
Qwen3 in thinking mode; Llama-3.1-8B-Instruct as a baseline; MMMLU over 14 languages.
e
Qwen3-VL report ( Qwen Team, 2025 )
Text-only Qwen3 baselines in thinking mode; DeepSeek-V3-0324 on MMLU-Redux.
f
MMLU-ProX ( Xuan et al., 2025 )
29 languages, full set, 5-shot with chain-of-thought.
Appendix
Table 7: Sources and protocols of the published scores. Keys n, o, and b also cover the vendor-reported GPQA Diamond scores of GPT-5.6 Sol, Gemini 3.8 Flash, and DeepSeek-V4.1-Flash in Table 2 .
University of Bonn, Bonn, Germany · Lamarr-Institute for Machine Learning and Artificial Intelligence, Bonn, Germany · Fraunhofer IAIS, Sankt Augustin, Germany