OmniMed-Jev: Calibrating LVLM Confidence for Trustworthy Medical Multimodal Decisions via System One
Organizations: Department of Electrical and Computer Engineering The University of Hong Kong, Hong Kong SAR, China
Abstract
Medical models are judged not only on correctness, but on whether reported confidence matches actual accuracy. Generalist multimodal medical models have expanded what a single model can perceive, yet they still express bounded decisions such as diagnoses, findings or cell counts as generated text, so the reported probability reflects the next token rather than the decision itself. Motivated by decision-native interfaces such as Jev, we introduce OmniMed-Jev, which represents each medical decision as a Choice, Noul or Score decision over a runtime-supplied candidate set and returns a full distribution over that set: mutually exclusive classes, binary presence of a finding, or a bounded ordered value. The design is omni in three respects: it accepts diverse imaging modalities, covers different prediction tasks, and expresses them through one candidate-conditioned probability model, so heterogeneous outputs become comparable probabilities rather than task-specific strings. In an interface-controlled comparison against a generative baseline trained on the same backbone, data and schedule, OmniMed-Jev's reported probabilities track observed correctness far more closely, reducing calibration error by up to an order of magnitude and reliability error by up to two, while point-prediction performance remains comparable; counting is the one family where the generative baseline stays ahead. Making the decision distribution the model's output is not a format change but what turns reported numbers into probabilities that mean what they say. These results support explicit decision modeling as a way to make reported confidence meaningful within the evaluated tasks, and they are not evidence of clinical readiness: the comparison cannot separate the interface from associated training differences, which we state alongside the results. Code is available at github.com/lytang63/OmniMed-Jev.
Figures & tables
| Task family | Metric | Gen-SFT | OmniMed-Jev | |
|---|---|---|---|---|
| Classification | Balanced accuracy | 0.6469 | 0.6892 | |
| Multi-label | F1 | 0.5449 | 0.5618 | |
| Precision | 0.5679 | 0.5950 | ||
| Recall | 0.5498 | 0.5602 | ||
| Counting | MAE | 2.1429 | 2.5159 | |
| Regression | MAE | 3.3400 | 2.9488 |
| Task family ( ) | ECE | Reliability | Resolution | AURC |
|---|---|---|---|---|
| Classification (200) | 0.2589 / 0.0806 | 0.1002 / 0.0110 | 0.0938 / 0.0860 | 0.2020 / 0.1351 |
| Multi-label (2,795) | 0.2981 / 0.0114 | 0.1294 / 0.0010 | 0.0776 / 0.0102 | 0.2474 / 0.0126 |
| Counting (201 / 126) | 0.4204 / 0.0408 | 0.1824 / 0.0041 | 0.0000 / 0.1087 | 1.0000 / 0.3890 |
| Regression (200) | 0.1267 / 0.0913 | 0.0258 / 0.0205 | 0.1426 / 0.1268 | 0.2272 / 0.2280 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Task family | Source datasets | Modalities | Type | States | Dec. | Test |
|---|---|---|---|---|---|---|
| Classification | TissueMNIST, PathMNIST, OCTMNIST, Organ{A,S,C}MNIST, BloodMNIST, DermaMNIST, PneumoniaMNIST, BreastMNIST | Histology, Pathology, CT, Hematology, Dermatology, Ophthalmology | Choice | 10,000 | 10,000 | 200 |
| Multi-label | ChestMNIST, ChestX-Det | Chest radiography | Noul 14 / 13 | 10,000 | 139,537 | 200 |
| Counting | TXL-PBC | Haematology microscopy | Score | 878 | 3,512 | 126 |
| Regression | CAMUS, RetinaMNIST | Ultrasound, Fundus | Score / | 2,680 | 2,680 | 200 |
| Total | 15 datasets | 8 modality labels | 3 types | 23,558 | 155,729 | 726 |
| Arm | Quantity | 126 matched | (201 recorded) |
| OmniMed-Jev | top-1 (argmax) accuracy | 0.1508 | (0.3234) |
| OmniMed-Jev | MAE of folded prediction | 2.5159 | (2.5159) |
| OmniMed-Jev | MAE of the expectation | 2.5179 | — |
| OmniMed-Jev | RMSE of the expectation | 3.3250 | (3.3747) |
| OmniMed-Jev | mean maximum probability | 0.1811 | (0.3309) |
| OmniMed-Jev | median maximum probability | 0.1697 | (0.1889) |
| Classification | Multi-label | Counting | Regression | |||||
|---|---|---|---|---|---|---|---|---|
| Step | J | S | J | S | J | S | J | S |
| 0 (base) | 0.0000 | 0.0000 | 0.1813 | 0.1813 | 13.9127 | 13.9127 | 17.7775 | 17.7775 |
| 404 | 0.3709 | 0.4171 | 0.5283 | 0.4483 | 4.1905 | 3.9444 | 3.4214 | 3.7350 |
| 808 | 0.4021 | 0.4416 | 0.5553 | 0.5650 | 10.7698 | 3.9603 | 3.9326 | 11.5100 |
| 1212 | 0.5114 | 0.5345 | 0.5672 | 0.5100 | 3.5635 | 2.9048 | 3.8455 | 3.2500 |
| 1616 | 0.6275 | 0.6499 | 0.5902 | 0.5400 | 3.2222 | 2.7937 | 3.3723 | 3.3800 |
| Probe | @ 1616 | @ 2424 |
|---|---|---|
| Permutation, total variation | 0.1284 | 0.1413 |
| Permutation, top-1 agreement | 0.6589 | 0.6622 |
| Permutation, mean selected rank | 0.5486 | 0.5346 |
| diagnosis | 0.049 / 0.910 | 0.049 / 0.900 |
| counting | 0.155 / 0.373 | 0.198 / 0.418 |
| regression | 0.181 / 0.695 | 0.176 / 0.670 |
| Family | Metric | 404 | 1616 | 2424 | 3685 | SFT @2424 |
|---|---|---|---|---|---|---|
| Class. | Accuracy | 0.3500 | 0.5700 | 0.6400 | 0.6750 | 0.5500 |
| ECE | 0.1538 | 0.0863 | 0.0806 | 0.1310 | 0.2589 | |
| Reliability | 0.0336 | 0.0123 | 0.0110 | 0.0310 | 0.1002 | |
| Resolution | 0.0602 | 0.1026 | 0.0860 | 0.0827 | 0.0938 | |
| AURC | 0.4354 | 0.1714 | 0.1351 | 0.1181 | 0.2020 | |
| Multi-lab. | Accuracy | 0.9417 | 0.9420 | 0.9352 | 0.9338 | 0.5166 |
| Item | Generative SFT | OmniMed-Jev |
|---|---|---|
| Backbone | models/medgemma-1.5-4b-it | same |
| Quantisation | 4-bit NF4, double quantisation, bf16 compute | same |
| LoRA rank / alpha / dropout | 16 / 16 / 0.05 | same |
| LoRA target modules | all-linear (including the vision tower) | same |
| Learning rate | same | |
| Schedule | linear | same |
| Generative SFT | OmniMed-Jev | |
|---|---|---|
| Wall clock | 46,280 s = 12.86 h | 50,204 s = 13.95 h |
| Final training loss | 0.5805 | 0.7807 |
| Peak memory | — | 17.18 GB (Stage B) |
| Device | GPU 2 (single A800) | GPU 0 (single A800) |
| Estimator | Outcome | |
|---|---|---|
| 1 | Append a PROBS: {…} request to the end of the user turn. | 0 of 726 rows parsed a confidence. |
| 2 | Write a two-line requirement into the system prompt (answer on line 1, confidence on line 2). | 0 of 20 rows parsed a confidence. |
| 3 | Untuned base model under the same instruction. | 8 of 20 rows produced a confidence and 1 of 20 produced scores, but every value was . |
| 4 | Read the option log-likelihoods from the model’s own logits. | Succeeded; the full metric set is computable. |