cs.AISep 28, 2026

Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration

Authors: Riccardo Porcedda

Organizations: Department of Excellence L’EMbeDS, Sant’Anna School of Advanced Studies, Italy · Department of Computer Science, University of Pisa, Italy

Abstract

The appearance of Jev marked the era of System One Models, foundation models that return structured decisions with probability distributions rather than text. Aside from low cost and great speed, Jev's central promise is that these probabilities are calibrated: such claim is not backed by any public test and available external benchmarks evaluate confidence calibration, not whether every returned option probability has the right numerical meaning. To tackle this issue, we introduce Sys1Cal-v1, a dataset of True/False questions about a proposition AA for which the exact probability P(A)P(A) is known by construction. Each item is queried through the three Jev primitives - Noul, Choice and Score - and evaluated by total variation distance from the ground-truth distribution, which can be used to estimate a soft accuracy of System One Models. We showcase the utility of Sys1Cal-v1 as a benchmark dataset by evaluating Jev and SemIf, an open-source Choice-style baseline. In this work, however, we focus even more deeply on Jev, by studying the calibration of its Score and Choice answers. In particular, we discover a peculiar behaviour that can be explained by assuming that Jev suppresses a third truth value, going beyond True and False. In other words, in \texttt{Choice} answers, P(A)P(A) and P(¬A)P(\neg A) are presented as if P(A)+P(¬A)=1P(A)+P(\neg A)=1, while a term P(U)≠0P(U)\neq0 is missing in the sum. Recovering P(U)P(U) leads to an improvement of median soft accuracy in \texttt{Choice} answers from 0.7710.771 to 0.9780.978, suggesting that, even in binary decisions, Jev wants to answer with a third option:``I don't know''.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Evaluating and Benchmarking the System One Model Jev

    Sep 29, 2026Tobias Deußer, Lorenz Sparrenberg, Rafet SifaNatural Language InferenceRubric-Based Scoring

  2. Do System One Decisions Add Up? A Study of Probabilistic Coherence

    Sep 27, 2026Saman Sarker JoyConfidence CalibrationDecisions

  3. Jev in Medicine: A Benchmark Evaluation

    Sep 27, 2026Alfredo Madrid-García, Beatriz Merino-BarbanchoDiagnosisPreliminary Study