cs.CLOct 8, 2026

Specialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual Tasks

Authors: Xing Li, Qingcheng Chang, Jinzhong Ning, Changfeng Xu, Shenlong Zhang, Yijia Zhang, Ling Luo, Hongfei Lin

Organizations: Dalian Maritime University · Dalian University of Technology

Abstract

Jev is a "System One" model that returns a choice among given options instead of generating text. We study how such a specialized decision model compares with general-purpose large language models (LLMs). We evaluate Jev on 13 multiple-choice benchmarks covering knowledge, reasoning, and multilingual understanding, and compare it with 19 LLMs in three tiers: frontier, representative, and small. Jev is competitive with frontier LLMs on knowledge and commonsense benchmarks and obtains the best score on MMLU-Redux and ARC-Challenge. Outside mathematics, it also outperforms most representative LLMs and all small LLMs. However, it falls behind on mathematical word problems: on MathQA, it is 17.7 points below the frontier median and scores lower than all 19 LLMs. These results indicate that a specialized decision model can match general-purpose LLMs on decisions that rely mainly on knowledge, but not on decisions that require multi-step calculation.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Evaluating and Benchmarking the System One Model Jev

    Sep 29, 2026Tobias Deußer, Lorenz Sparrenberg, Rafet SifaMultilingual Language Model EvaluationLLM Evaluation

  2. Chinese-Jev: Bringing System One Model to Chinese-Language Tasks

    Sep 29, 2026Zexiao Wang, Zihao Zhang, Xudong Wang +5LLM Decision-MakingEfficient Language Model Inference