Jev in Medicine: A Benchmark Evaluation
Organizations: Independent researcher Madrid, Spain
Abstract
Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges. GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognition of unanswerable questions. All 8,469 requests returned a valid answer. Jev's accuracy was similar to that of GPT-6 Sol with medium reasoning on PubMedQA (78.4% vs 78.2%;), lower on MetaMedQA (74.8% vs 82.7%) and much lower on DiagnosisArena-MCQ (59.8% vs 82.4%;) and the NEJM cases (61.8% vs 82.4%). On MetaMedQA, Jev's probabilities were the best calibrated (expected calibration error 0.063 vs 0.146), and its answers with a probability of at least 0.9 (52.9% of questions) were 93.4% accurate, but GPT-6 Sol was as accurate when it accepted a similar proportion of questions. On DiagnosisArena-MCQ, Jev's probabilities discriminated poorly (AUROC 0.645 vs 0.768). Of the 162 questions whose correct answer was "I don't know or cannot answer", Jev chose that option for 10.5% (GPT-6 Sol, 8.6%). Median latency was 0.27-0.31 s; all 2,823 items cost USD 0.08. Jev was fast and inexpensive, and its accuracy was similar to that of a frontier LLM on research abstracts but lower on examination questions and much lower on complex diagnostic cases. Task-specific validation is required before clinical use.
Figures & tables
| Benchmark | Items analysed | Answer options | Task | Distinctive feature |
| MetaMedQA [ 14 ] | 1,373 | 6: A–D, “None of the above”, “I don’t know or cannot answer” | USMLE-style examination questions | 277 questions without a correct substantive option (115 “None of the above”, 162 “I don’t know”), including 100 on a fictional organ |
| PubMedQA [ 23 ] | 500 | 3: yes, no, maybe | Answer a research question from a PubMed abstract (conclusions removed) | Expert-annotated official test split; single-annotator accuracy 78.0% |
| DiagnosisArena-MCQ [ 24 ] | 915 | 4: A–D | Most likely diagnosis in cases from case reports in ten high-impact journals | Cases filtered to remove those that LLMs solved easily; distractors derived from errors of reasoning models |
| NEJM Case Challenges [ 25 ] | 34 (33 with a closed poll) | 6: A–F | Most likely diagnosis in Case Records of the Massachusetts General Hospital | 273,362 reader votes; laboratory tables transcribed as text, figures replaced by captions |
| Benchmark (items; chance) | Jev | GPT-6 Sol (medium) | GPT-6 Sol (none) | Jev GPT-6 Sol (medium), pp (95% CI) | P (Holm) | Jev GPT-6 Sol (none), pp (95% CI) |
| MetaMedQA (1,373; 16.7%) | 74.8 (72.4–77.0) | 82.7 (80.6–84.6) | 80.7 (78.5–82.7) | 7.9 ( 10.0 to 6.0) | 0.001 | 5.9 ( 7.8 to 4.2) |
| PubMedQA (500; 33.3%) | 78.4 (74.6–81.8) | 78.2 (74.4–81.6) | 76.6 (72.7–80.1) | 0.2 ( 2.2 to 2.6) | 1.00 | 1.8 ( 1.0 to 4.8) |
| DiagnosisArena-MCQ (915; 25.0%) | 59.8 (56.6–62.9) | 82.4 (79.8–84.7) | 79.2 (76.5–81.7) | 22.6 ( 25.9 to 19.7) | 0.001 | 19.5 ( 22.6 to 16.3) |
| NEJM Case Challenges (34; 16.7%) | 61.8 (45.0–76.1) | 82.4 (66.5–91.7) | 76.5 (60.0–87.6) | 20.6 ( 35.3 to 5.9) | 0.078 | 14.7 ( 29.4 to 0.0) |
| Benchmark and model | Mean selected- option probability | ECE | AUROC | Brier score | Coverage at 0.9, % | Accuracy at 0.9, % (95% CI) | AURC |
| MetaMedQA (n = 1,373) | |||||||
| Jev | 0.811 | 0.063 | 0.845 | 0.343 | 52.9 | 93.4 (91.3–95.0) | 0.087 |
| GPT-6 Sol (medium) | 0.973 | 0.146 | 0.801 | 0.316 | 93.2 | 85.9 (83.9–87.7) | 0.070 |
| GPT-6 Sol (none) | 0.951 | 0.145 | 0.818 | 0.341 | 86.7 | 86.5 (84.4–88.3) | 0.071 |
| PubMedQA (n = 500) | |||||||
| Jev | 0.920 | 0.141 | 0.766 | 0.350 | 76.8 | 87.8 (84.1–90.7) | 0.109 |
| Dataset | Instruction | Case content (Jev state) | Answer options |
| MetaMedQA | “Which option is the correct answer to this question?” | Complete original question, including its final question sentence (string) | A–F with their original text and order; E, “None of the above”; F, “I don’t know or cannot answer” |
| PubMedQA | The research question, verbatim (not repeated elsewhere) | Abstract without its conclusions, each section as “HEADING: text” (object with one field, Abstract) | yes, no and maybe, each with a one-sentence description |
| DiagnosisArena-MCQ | “What is the most likely diagnosis for this patient?” | Case Information, Physical Examination and Diagnostic Tests (object with these three fields) | A–D with their original text and order |
| NEJM Case Challenges | The poll question, verbatim (“What is the most likely diagnosis in this case?”; case 12, “The most likely diagnosis is:”) | Case presentation up to the diagnostic question, with laboratory tables as text and figures replaced by their captions (string) | A–F with their original text, in the published order |
| Comparison (first second) | Accuracy, pp (95% CI) | Correct by first / second only; P | Agreement, % ( ) | Brier score (95% CI) | ECE (95% CI) | AUROC (95% CI) | AURC (95% CI) |
| MetaMedQA (1,373 questions) | |||||||
| Jev GPT-6 Sol (medium) | 7.9 ( 10.0 to 6.0) | 38 / 147; 0.001 | 83.0 (0.79) | 0.027 ( 0.001 to 0.055) | 0.083 ( 0.100 to 0.064) | 0.044 (0.014 to 0.074) | 0.017 (0.005 to 0.029) |
| Jev GPT-6 Sol (none) | 5.9 ( 7.8 to 4.2) | 44 / 125; 0.001 | 83.5 (0.79) | 0.002 ( 0.023 to 0.029) | 0.081 ( 0.098 to 0.062) | 0.027 ( 0.001 to 0.053) | 0.016 (0.006 to 0.027) |
| GPT-6 Sol (none) GPT-6 Sol (medium) | 2.0 ( 3.1 to 0.9) | 13 / 41; 0.001 | 93.4 (0.92) | 0.025 (0.008 to 0.043) | 0.001 ( 0.012 to 0.010) | 0.017 ( 0.006 to 0.042) | 0.001 ( 0.009 to 0.010) |
| PubMedQA (500 instances) | |||||||
| Jev GPT-6 Sol (medium) | 0.2 ( 2.2 to 2.6) | 18 / 17; 1.00 | 91.4 (0.85) | 0.010 ( 0.036 to 0.016) | 0.000 ( 0.024 to 0.025) | 0.035 ( 0.017 to 0.085) | 0.017 ( 0.035 to 0.001) |
| Measure | Jev | GPT-6 Sol (medium) | GPT-6 Sol (none) | Readers |
| 34 cases with a published final diagnosis | ||||
| Accuracy, % (n correct; 95% CI) | 61.8 (21; 45.0–76.1) | 82.4 (28; 66.5–91.7) | 76.5 (26; 60.0–87.6) | — |
| Mean selected-option probability | 0.676 | 0.944 | 0.925 | — |
| Brier score | 0.577 | 0.285 | 0.383 | — |
| 33 cases with a final diagnosis and a closed poll (273,362 votes) | ||||
| Accuracy, % (n correct); readers, mean % of votes for the correct option | 63.6 (21) | 81.8 (27) | 75.8 (25) | 31.2 |
| Model | Coverage 0.5, % | Accuracy 0.5, % (95% CI) | Coverage 0.8, % | Accuracy 0.8, % (95% CI) | Coverage 0.9, % | Accuracy 0.9, % (95% CI) |
| MetaMedQA (n = 1,373) | ||||||
| Jev | 88.1 | 81.3 (79.0–83.4) | 63.1 | 91.3 (89.3–93.0) | 52.9 | 93.4 (91.3–95.0) |
| GPT-6 Sol, medium | 99.9 | 82.9 (80.8–84.8) | 97.3 | 84.1 (82.0–85.9) | 93.2 | 85.9 (83.9–87.7) |
| GPT-6 Sol, none | 99.3 | 81.1 (78.9–83.1) | 94.0 | 83.5 (81.4–85.4) | 86.7 | 86.5 (84.4–88.3) |
| PubMedQA (n = 500) | ||||||
| Jev | 97.8 | 78.9 (75.1–82.3) | 85.0 | 84.2 (80.5–87.4) | 76.8 | 87.8 (84.1–90.7) |
| Measure (questions) | Jev | GPT-6 Sol (medium) | GPT-6 Sol (none) |
| By reference category, % (95% CI) | |||
| Accuracy, correct answer A–D (1,096) | 86.5 (84.3–88.4) | 94.6 (93.1–95.8) | 93.2 (91.5–94.5) |
| Missing-answer recall, correct answer E (115) | 53.9 (44.8–62.7) | 73.9 (65.2–81.1) | 69.6 (60.6–77.2) |
| Unknown recall, correct answer F (162) | 10.5 (6.7–16.2) | 8.6 (5.2–14.0) | 4.3 (2.1–8.6) |
| Selected category, n | |||
| E or F selected when the correct answer was A–D (1,096), n (%) | 33 (3.0) | 25 (2.3) | 33 (3.0) |
| Benchmark and model | Latency, median (IQR), s | Latency, P95, s | Input tokens per item, median | Reasoning tokens, total | Total cost, US$ | Cost per 1,000 items, US$ |
| MetaMedQA (1,373) | ||||||
| Jev | 0.28 (0.26–0.31) | 0.37 | 559 | — | 0.033 | 0.024 |
| GPT-6 Sol (medium) | 2.27 (1.91–2.80) | 5.32 | 364 | 132,100 | 3.13 | 2.28 |
| GPT-6 Sol (none) | 1.49 (1.39–1.61) | 1.90 | 364 | 0 | 1.73 | 1.26 |
| PubMedQA (500) | ||||||
| Jev | 0.27 (0.25–0.30) | 0.35 | 729 | — | 0.015 | 0.031 |