Chinese-Jev: Bringing System One Model to Chinese-Language Tasks
Organizations: Fudan University · National University of Singapore · University of Chinese Academy of Sciences
Abstract
System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. However, existing Jev models exhibit limited Chinese-language decision accuracy, restricting their utility in both general and specialized settings. In this paper, we introduce Chinese-Jev, a System One model that addresses this gap through a unified data processing and training pipeline. Our data processing protocol converts heterogeneous Chinese-language annotations into probability targets over candidate options, enabling a shared training formulation across domains and question formats. To enable efficient inference, Chinese-Jev adopts a lightweight encoder-only backbone for text encoding and learns to score candidate answers through decision-oriented training. To address the misalignment between the pre-training distribution and downstream Chinese-language scenarios, we first train the model on a general-purpose corpus of 10 million examples, then fine-tune it separately for the medical, legal, and financial domains. To evaluate decision accuracy and calibration in both general and domain-specific Chinese-language settings, we introduce Chinese-Jev Bench (CJ-Bench). After first-stage pre-training, Chinese-Jev exceeds the accuracy of the closed-source Jev model by 1.24% on general-domain tasks while achieving a 20.3x speedup. Subsequent domain-specific fine-tuning yields a 4.0% accuracy improvement over Jev in medicine and achieves 92% of Jev's average accuracy across specialized domains, with a 17x speedup and an average latency of only 15 ms per example. We further demonstrate on-device deployment of an INT8-quantized model on mobile devices, achieving an inference latency of approximately 1.0 second per decision. The project is available at https://gulucaptain.github.io/Chinese-Jev/.
Figures & tables
| Model | ACC (%) | ECE (%) | Latency (ms) |
| Qwen3.5-2B ( Qwen Team, 2026 ) | 32.95 | 43.68 | 280 |
| Open-Jev ( Cai, 2026 ) | 52.33 | 10.54 | 60 |
| SemIF ( Lee, 2026 ) | 44.21 | 27.27 | 57 |
| Laya Multilingual ( Convai Innovations, 2026 ) | 40.22 | 23.04 | 14 |
| Jev ( Almeida, 2026 ) | 68.35 | 11.45 | 284 |
| Chinese-Jev General | 69.20 (+0.85) | 3.78 (-6.76) | 14 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Partition | choice | noul | score | Total |
| General | ||||
| Train | 3,333,334 | 3,333,333 | 3,333,333 | 10,000,000 |
| Development pool | 179,723 | 215,520 | 357,543 | 752,786 |
| Calibration pool | 63,820 | 96,679 | 107,454 | 267,953 |
| Full test pool | 214,513 | 929,814 | 823,368 | 1,967,695 |
| Benchmark | 33,334 | 33,333 | 33,333 | 100,000 |
| Source | Decisions | Share | Source | Decisions | Share |
| T2Ranking ( Xie et al., 2023 ) | 2,328,371 | 23.28% | CMRC2019 ( Cui et al., 2020 ) | 93,605 | 0.94% |
| ASAP ( Bu et al., 2021 ) | 1,318,889 | 13.19% | OCNLI ( Hu et al., 2020 ) | 59,158 | 0.59% |
| Dianping ( Zhang and LeCun, 2017 ) | 1,256,330 | 12.56% | PAWS-X (Chinese) ( Yang et al., 2019 ) | 45,508 | 0.46% |
| JD full reviews ( Zhang and LeCun, 2017 ) | 973,902 | 9.74% | Poetry retrieval ( PoetryMTEB Contributors, 2026 ) | 29,796 | 0.30% |
| DMSC ( utmhikari, 2017 ) | 933,902 | 9.34% | LogiQA 2 (Chinese) ( Liu et al., 2023b ) | 11,526 | 0.12% |
| MEP-3M ( Liu et al., 2023a ) | 898,261 | 8.98% | C3 ( Sun et al., 2020 ) | 11,502 | 0.12% |
| Task category | Train (%) | Test decisions | Test (%) |
| Aspect mention | 4.269 | 4,480 | 4.480 |
| Aspect status and sentiment | 8.560 | 6,481 | 6.481 |
| Category and intent | 14.132 | 9,268 | 9.268 |
| Reading and reasoning | 7.624 | 11,784 | 11.784 |
| Review rating | 19.438 | 9,987 | 9.987 |
| Semantic matching | 24.067 | 26,295 | 26.295 |
| Variation between examples | Treatment |
| Record identifier or normalized whitespace | Same fingerprint; retain one when targets agree. |
| choice alternatives reordered | Same fingerprint; compare targets by candidate text, not answer letter. |
| Alternative added, removed, or replaced | Different decision, even when the question and correct answer are unchanged. |
| Same normalized input, different target | Exclude all examples with that fingerprint. |
| Different task, question, number, or negation | Different decision; preserve the changed meaning. |
| Ordered score levels changed or reordered | Different decision; the scale order is retained. |
| Domain | choice | noul | score | Total |
| Medical | 1,194,496 | 1,792,223 | 13,281 | 3,000,000 |
| Legal | 8,000 | 8,000 | 8,000 | 24,000 |
| Finance | 8,000 | 8,000 | 8,000 | 24,000 |
| Source | choice | noul | score | Total |
| Legal | ||||
| AppealCase ( Huang et al., 2025 ) | 1,008 | 4,081 | 3,373 | 8,462 |
| ClaimGen-CN ( Zhou et al., 2025 ) | 6,133 | 2,267 | 0 | 8,400 |
| LeCaRDv2 ( Li et al., 2024a ) | 200 | 200 | 2,600 | 3,000 |
| Explicit legal rules | 200 | 400 | 1,600 | 2,200 |
| LCR-CN ( Zhao et al., 2026 ) | 302 | 898 | 0 | 1,200 |