A correct diagnosis reached from insufficient or misleading evidence can pose a clinical hazard, yet outcome-based accuracy may reward such lucky guesses. We call this mismatch between diagnostic decisions and the value of available evidence Evidence-Value Misalignment (EVM). To disentangle evidential grounding independently from diagnostic accuracy, we introduce MedEVM, a dynamic benchmarking environment comprising 1,050 cases across 24 disease systems. Observations arrive turn by turn, requiring models to continuously calibrate its decision by deciding whether to wait for more evidence or submit a diagnosis. Across 9 LLMs, four interesting patterns are observed. (1) Miscalibrated evidence tracking. Making a diagnosis often fails to calibrate evidence sufficiency, even in more capable models, and even worsens in reasoning mode. (2) Misaligned diagnosis submission. Confidence in the correct diagnosis often fails to ensure timely submission despite sufficient evidence. (3) Evidence order matters. Reordering the same evidence changes diagnoses even when model confidence remains similar. (4) Misleading evidence remains influential. Added misleading evidence redirects diagnoses even after prior evidence becomes sufficient. We further verify that EVM predicts errors and that preventing premature submission improves accuracy. These findings motivate Evidence-Verified Diagnosis Harness (EVD-Harness). It decouples diagnosis generation from submission through an offline Contrastive Diagnostic Wiki and three online control stages, namely observation management, proposal and witness verification, and diagnosis submission control. Across five LLMs, EVD-Harness improves accuracy by 12.0--51.1 percentage points while mitigating EVM-related failures. Our results demonstrate that verifying evidential support before submission can make diagnostic decisions more reliable.
Figures & tables
Figure 1: Overview of MedEVM . Final accuracy alone obscures differences in when models submit diagnoses. In MedEVM , clinical evidence arrives turn by turn, and the LLM Agent must wait or submit a diagnosis. This evaluation reveals four key findings hidden by answer-only scoring.
Figure 2
Figure 3: Case proportions by diagnosis submission time relative to the first sufficient-evidence turn t∗ . Red triangles show early submission with thinking.
Figure 4
Figure 5: Diagnosis outcome disagreement after reordering in all cases (left), and its relation to belief similarity at t∗ among paired orders that both reach t∗ (right).
Figure 6: Strong versus weak misleading evidence (pp, positive = harm). (a) From left: immediate target-belief gain, target submission rate increase, accuracy loss, target-belief gain after all evidence, and post- t∗ loss in correct diagnoses ranked first at completion. The first four use early insertion. (b) Immediate target-belief gains by insertion stage, with dashed curves for models and a solid median.
Figure 7
Figure 8: EVD-Harness compiles independent case reports into contrastive knowledge, guides diagnosis, and authorizes answers grounded in current evidence.
Table 4: LLM only versus EVD-Harness (%). Non-answers count as errors. Premature answers precede t∗ . Order flip counts diagnosis changes after reordering. Target error counts incorrect target submissions under strong misleading evidence before corrective evidence arrives.
Table 10
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Premature (%)
Within window (%)
Delayed (%)
Model
k=0
k=1
k=2
k=0
k=1
k=2
k=0
k=1
k=2
Q3-8B
99.1
97.7
96.0
0.6
2.2
3.9
0.3
0.1
0.1
Q3-14B
78.6
70.3
63.7
17.6
27.1
34.2
2.7
1.4
1.0
Q3-32B
35.3
27.0
21.6
37.0
48.5
55.7
9.6
6.5
4.7
Q3-235B
49.7
40.2
33.5
38.3
49.9
58.6
6.8
4.7
2.7
Q3.6-27B
41.0
34.8
28.7
42.7
51.9
59.0
6.4
3.4
2.5
Appendix
Table A1: RQ1 timing sensitivity to a tolerance of k turns around the annotated boundary. All rates use the same 1,050 cases per model. Non-answer rates are unchanged and omitted.
Early submission
Submission accuracy
Belief measurements (on)
Model
Off
On
Off
On
pgold
Margin
ECE
nt∗
Q3-8B
99.1
99.7
10.9
11.3
11.5
100.0
89.8
4
Q3-14B
78.6
99.7
30.5
12.0
10.4
99.8
90.1
5
Q3-32B
35.3
99.6
59.2
13.9
14.1
99.3
87.9
4
Q3-235B
49.7
99.6
55.4
17.3
15.9
99.7
86.5
3
Q3.6-27B
41.0
94.9
70.2
37.8
35.6
100.0
76.1
50
Appendix
Table A2: Thinking-mode results. All rates, probabilities, margins, and ECE are multiplied by 100. Off/on denotes the thinking setting. pgold is measured at diagnosis submission, whereas margin and ECE use all measured evidence checkpoints. nt∗ counts boundary measurements with thinking on.
Model
Source cases
Order pairs
JSD ×100
Final diagnosis disagreement (%)
Q3-8B
7
17
2.12 [0.02, 5.98]
3.6 [0.0, 10.7]
Q3-14B
76
191
1.58 [0.58, 2.71]
6.2 [1.6, 11.8]
Q3-32B
635
2120
4.36 [3.66, 5.07]
15.4 [13.1, 17.9]
Q3-235B
353
955
2.87 [2.11, 3.77]
13.1 [9.9, 16.4]
Q3.6-27B
545
1678
1.28 [1.04, 1.54]
13.8 [11.3, 16.6]
Q3.8-27B
420
1091
0.90 [0.72, 1.11]
13.2 [10.3, 16.3]
Appendix
Table A3: Evidence-order comparisons for pairs in which both trajectories reach ti∗ . Source cases and order pairs count distinct cases and eligible original–permuted comparisons, respectively. JSD and final diagnosis disagreement use the same pairs and equal source-case weights. Brackets give 95% bootstrap confidence intervals. All order pairs from a sampled case remain together.
Figure A2: Diagnosis output differences immediately after repeated evidence or irrelevant text is added to the same evidence prefix. Points show the percentage of comparisons with different submitted diagnoses or a diagnosis in only one branch. Intervals are source-case bootstrap 95% confidence intervals, with comparisons averaged within each case.
Early insertion
After t∗
Model
Target belief at insertion ↑
Target diagnosis submission ↑
Diagnostic accuracy ↓
Target belief at completion ↑
Correct diagnosis ranked first ↓
Q3-8B
11.1 [9.3, 13.1]
3.0 [1.7, 4.5]
1.0 [0.2, 1.9]
4.4 [3.4, 5.5]
3.9 [2.2, 5.8]
Q3-14B
18.2 [15.9, 20.6]
21.7 [18.6, 24.8]
7.2 [4.9, 9.7]
9.1 [7.4, 10.8]
6.0 [4.1, 8.0]
Q3-32B
17.7 [15.7, 19.7]
24.0 [21.2, 27.0]
12.3 [9.6, 15.1]
10.8 [9.1, 12.6]
11.1 [8.7, 13.5]
Q3-235B
18.4 [16.0, 20.8]
33.2 [30.1, 36.4]
19.6 [16.5, 22.7]
9.9 [8.3, 11.8]
11.9 [9.3, 14.5]
Q3.6-27B
24.6 [22.8, 26.3]
35.0 [31.8, 38.5]
26.7 [23.1, 30.0]
8.6 [7.4, 9.9]
16.6 [14.1, 19.2]
Appendix
Table A4: Effects of stronger misleading evidence across models (percentage points). The first four endpoints use early insertion, and the last uses insertion after ti∗ . All positive values indicate harm. Brackets give source-case bootstrap 95% confidence intervals. Final diagnostic accuracy counts non-answers as errors. Bold values mark the smallest observed harmful effect in each column.
Model
M0
M1
M2
Mprop
M3
Δ (M3 − M1)
Qwen3-8B
0.791
0.795
0.812
0.838
0.820
+0.025
Qwen3-14B
0.778
0.780
0.806
0.824
0.827
+0.047
Qwen3-32B
0.830
0.826
0.846
0.806
0.846
+0.020
Qwen3-235B-A22B
0.831
0.827
0.844
0.772
0.840
+0.013
Qwen3.6-27B
0.813
0.815
0.821
0.765
0.812
-0.003
Qwen3.8-27B
0.850
0.857
0.877
0.843
0.885
+0.028
Appendix
Table A5: Balanced next-turn submission prediction. AUROC compares upcoming diagnosis submissions with continued waiting. M3 combines confidence and EVM features, and Δ is its improvement over M1.
Model
M0
M1
M2
Mprop
M3
Δ (M3 − M1)
Qwen3-8B
0.4369
0.4340
0.4179
0.4575
0.3985
-0.0355
Qwen3-14B
0.2175
0.2166
0.2104
0.2395
0.2041
-0.0125
Qwen3-32B
0.1428
0.1434
0.1392
0.1653
0.1391
-0.0043
Qwen3-235B-A22B
0.1682
0.1693
0.1645
0.1987
0.1635
-0.0058
Qwen3.6-27B
0.1620
0.1611
0.1600
0.1837
0.1600
-0.0011
Qwen3.8-27B
0.1472
0.1453
0.1418
0.1729
0.1378
-0.0075
Appendix
Table A6: Log loss for balanced next-turn submission prediction. Lower values indicate better submission-probability estimates, and negative M3 − M1 differences favor the addition of EVM features.
Model
M0
M1
M2
Mtiming
M3
Oracle
Δ (M3 − M1)
Qwen3-8B
0.625
0.640
0.637
0.618
0.724
1.000
+0.084
Qwen3-14B
0.659
0.627
0.635
0.731
0.822
1.000
+0.195
Qwen3-32B
0.665
0.711
0.724
0.750
0.841
1.000
+0.130
Qwen3-235B-A22B
0.674
0.789
0.788
0.773
0.847
1.000
+0.058
Qwen3.6-27B
0.698
0.783
0.790
0.809
0.777
0.999
-0.006
Qwen3.8-27B
0.630
0.773
0.787
0.751
0.789
1.000
+0.016
Appendix
Table A7: Balanced assessment of incorrect diagnoses at submission, measured by AUROC. M1 and M3 use current and preceding model measurements. The timing and oracle references additionally use the annotated boundary and correct-answer identity, respectively.
Diagnosis outcomes on selected cases
Model
Cases
Original accuracy (%)
Gated accuracy (%)
Submission rate (%)
Accuracy gain (pp)
Qwen3-8B
1,041
10.4
55.6
100.0
+45.2
Qwen3-14B
825
19.4
67.8
99.9
+48.4
Qwen3-32B
371
32.6
62.3
88.9
+29.6
Qwen3-235B
522
32.6
76.2
96.0
+43.7
Qwen3.6-27B
431
43.2
81.4
91.4
+38.3
Appendix
Table A8: Natural Gate outcomes for each LLM. The upper panel reports accuracy and submission rates on cases with an original premature diagnosis. The lower panel decomposes the paired accuracy change into errors corrected and correct diagnoses lost, and shows the corresponding gain over all 1,050 cases. Rates are percentages, and gains and confidence limits are percentage points (pp). Bold values highlight gated accuracy and the net accuracy gain on selected cases.
Accuracy ↑
Premature incorrect ↓
Boundary submissions ↑
Model
LLM only
EVD-Harness
LLM only
EVD-Harness
LLM only
EVD-Harness
Qwen3.8-27B
60.9
75.6
32.0
4.9
41.6
73.8
DeepSeek-V4-Flash
42.9
68.4
14.7
8.9
27.2
70.4
Qwen3-32B
50.9
62.9
24.0
7.8
38.9
67.1
Kimi-K2.6
60.2
78.0
31.1
7.6
36.9
66.4
Claude-Opus-4-8
34.2
85.3
4.9
3.3
24.7
68.4
Appendix
Table A9: Diagnosis outcomes on MedEVM . Premature incorrect submissions combine a timing error with a diagnostic error. Boundary submissions occur exactly at ti∗ . Bold values indicate the better result within each model comparison.
Model
Method
Submission timing
Diagnostic outcome
Premature
Within window
Delayed
No answer
Accuracy
Correct within window
Qwen3.8-27B
LLM only
41.8
53.1
2.4
2.7
60.9
45.1
EVD-Harness
9.8
84.2
3.6
2.4
75.6
67.1
DeepSeek-V4-Flash
LLM only
17.9
37.9
4.9
39.3
42.9
33.5
EVD-Harness
13.1
82.4
3.1
1.3
68.4
57.9
Qwen3-32B
LLM only
29.1
48.7
5.8
16.4
50.9
38.9
Appendix
Table A10: Submission timing and accuracy with a ±1 -turn boundary tolerance (%). Correct within window counts correct diagnoses submitted inside [t∗−1,t∗+1] . Bold marks the better result per model.
Model
Method
Submission timing
Diagnostic outcome
Premature
Within window
Delayed
No answer
Accuracy
Correct within window
Qwen3.8-27B
LLM only
33.8
61.3
2.2
2.7
60.9
48.4
EVD-Harness
7.8
88.0
1.8
2.4
75.6
69.3
DeepSeek-V4-Flash
LLM only
15.2
42.0
3.6
39.3
42.9
35.3
EVD-Harness
10.7
87.5
0.4
1.3
68.4
62.6
Qwen3-32B
LLM only
21.8
57.6
4.2
16.4
50.9
43.3
Appendix
Table A11: Submission timing and accuracy with a ±2 -turn boundary tolerance (%). Correct within window counts correct diagnoses submitted inside [t∗−2,t∗+2] . Bold marks the better result per model.
Table 23
Accuracy ↑
Premature incorrect ↓
Model
None
Random
BM25
Contrast.
CDW
None
Random
BM25
Contrast.
CDW
Qwen3.8-27B
63.0
64.0
67.0
72.0
75.5
26.0
12.5
9.0
5.0
3.5
DeepSeek-V4-Flash
54.0
55.0
59.0
65.0
68.5
39.5
17.0
11.5
9.0
6.5
Qwen3-32B
50.5
51.5
54.5
59.5
63.0
29.0
17.5
11.5
7.5
6.0
Kimi-K2.6
66.0
67.0
70.0
75.0
78.0
24.0
17.5
13.0
7.5
6.0
Claude-Opus-4-8
77.5
78.5
80.5
84.5
86.5
0.5
1.5
1.0
0.5
2.5
Appendix
Table A14: Accuracy and premature incorrect submissions under five knowledge views (%). Bold marks the best result for each model and metric.
Accuracy ↑
No answer ↓
Model
Natural
Short
Direct
Natural
Short
Direct
Qwen3.8-27B
59.5
63.5 +4.0
57.5 -6.0
3.5
2.5 -1.0
2.5 +0.0
DeepSeek-V4-Flash
45.5
49.5 +4.0
43.0 -6.5
39.5
1.0 -38.5
1.0 +0.0
Qwen3-32B
54.0
58.0 +4.0
46.5 -11.5
14.5
4.0 -10.5
4.5 +0.5
Kimi-K2.6
61.5
65.5 +4.0
60.0 -5.5
5.0
3.5 -1.5
6.0 +2.5
Claude-Opus-4-8
37.5
46.5 +9.0
75.0 +28.5
56.0
36.0 -20.0
12.5 -23.5
Appendix
Table A15: Accuracy and non-answer rates (%) under history and execution controls. Small numbers show changes from the preceding condition in percentage points. Bold marks the best result per model and metric.
LLM only
EVD-Harness
Model
Repeated evidence
Irrelevant text
Paraphrased evidence
Mean
Repeated evidence
Irrelevant text
Paraphrased evidence
Mean
Qwen3.8-27B
9.7
8.5
8.3
8.8
1.6
1.8
1.6
1.7
DeepSeek-V4-Flash
31.7
27.0
29.0
29.2
1.1
1.1
1.3
1.2
Qwen3-32B
13.3
13.3
11.9
12.8
3.4
3.8
2.7
3.3
Kimi-K2.6
21.6
18.7
20.7
20.3
1.8
1.1
1.6
1.5
Claude-Opus-4-8
30.8
30.6
31.5
31.0
3.4
3.1
3.4
3.3
Appendix
Table A16: Answer-flip rates under evidence insertions (%, lower is better). Each method is evaluated under the same three perturbations. Mean denotes the average across conditions. Bold marks the lower rate in each paired comparison.
Direct sampling
EVD-Harness
Model
1 sample
5 samples
6 samples
Full
− Witness
− CDW
Qwen3.8-27B
54.1
57.8
58.0
57.1
64.8
53.9
DeepSeek-V4-Flash
52.0
53.5
53.6
56.1
59.9
53.5
Qwen3-32B
50.2
50.8
51.7
49.6
51.8
46.2
Kimi-K2.6
63.7
64.0
64.9
67.5
72.6
64.3
Claude-Opus-4-8
67.8
67.4
67.3
69.8
77.6
63.3
Appendix
Table A17: DiagnosisArena-MCQ accuracy (%). Direct sampling aggregates the indicated number of independently generated answers. The single-sample column is LLM only. Bold marks the highest accuracy for each model.
DiagnosisArena-MCQ
MedXpertQA-Text
Model
LLM only
EVD-Harness
LLM only
EVD-Harness
Qwen3.8-27B
54.1
57.1
24.4
27.9
DeepSeek-V4-Flash
52.0
56.1
15.1
22.1
Qwen3-32B
50.2
49.6
10.5
18.6
Kimi-K2.6
63.7
67.5
17.4
20.9
Claude-Opus-4-8
67.8
69.8
38.4
44.2
Appendix
Table A18: All-sample accuracy (%) on DiagnosisArena-MCQ and MedXpertQA-Text. Unanswered questions count as errors. Bold marks the better method for each model and benchmark.