Distributed inference has become an indispensable part of deploying medical models under practical latency, memory, and throughput constraints. Although modern frameworks improve serving efficiency through tensor parallelism, mixed precision, kernel fusion, and multi-device communication, they are generally assumed to preserve the behavior observed during centralized HuggingFace evaluation. This assumption creates an evaluation-deployment mismatch: a model may pass offline evaluation but produce a different output after the execution stack changes. To address this mismatch, we propose a testing framework and an improved, distributed-execution-sensitive medical-model benchmark that evaluates the same checkpoint and input under a centralized HuggingFace reference and matched distributed deployments. Extensive experiments across language, vision, and multimodal medical models show that execution changes can produce measurable output disagreements. Across supported visual settings, the test success rate ranges from 0.21 to 0.43 for single-modality models and from 0.32 to 0.98 for multimodal models. The benchmark is aimed at extending medical-model evaluation from capability and security to evaluation-deployment consistency.
Figures & tables
Figure 1: The same abdominal CT image receives different predictions under HuggingFace and two-GPU DeepSpeed inference in our test.
Figure 2: Details of our proposed method TDST.
Figure 3: Natural output discrepancy between the standard non-distributed configuration and the two-GPU DeepSpeed configuration on MedQA and MedQuAD.
Model
DeepSpeed
FSDP
vLLM
TensorRT-LLM
SR
Step
SR
Step
SR
Step
SR
Step
2-GPU Execution
ResNet-18
0.27
35.67
0.22
41.74
-
-
-
-
RCLIP
0.41
100.63
0.38
86.06
-
-
-
-
BLIP2-OPT-2.7B
0.51
70.90
0.47
65.81
0.98
16.72
0.96
23.65
Qwen2.5-VL-7B-Instruct
0.33
72.48
0.42
31.40
0.93
22.57
0.47
156.16
Table 1: Cross-framework evaluation under 2-GPU and 4-GPU execution. Each framework reports the Success Rate (SR) and Avg. Step. Note that ResNet and RCLIP are not supported by vLLM and TensorRT-LLM.
Figure 4: A TDST-generated medical multiple-choice test that produces different answers under HuggingFace and distributed execution.
Figure 5: Two-GPU divergence results between HuggingFace and the DeepSpeed and vLLM inference frameworks across four medical language models. Panel (a) reports the success rate of inducing output disagreement, while Panel (b) shows the average number of optimization steps required. Results are evaluated on MedMCQA and PubMedQA.
Model
Success Rate (SR) ↓
Deep Speed
FSDP
vLLM
TensorRT- LLM
ResNet-18
0.00
0.00
-
-
RCLIP
0.01
0.00
-
-
BLIP2-OPT-2.7B
0.19
0.30
0.45
0.07
Qwen2.5-VL-7B
0.13
0.04
0.86
0.46
Qwen3-VL-8B
0.24
0.38
0.17
0.01
Table 2: Cross-framework mitigation performance under 2-GPU execution. We report the post-mitigation success rate (SR) after elevating the localized critical layers to higher precision; lower values indicate stronger mitigation.
Existing medical AI benchmarks lack process visibility, atomic skill evaluation, and integrated hallucination detection. We introduce MedBench v5, a redesigned benchmark for clinical multimodal models (language, vision-language, and agent systems) that moves from static QA to dynamic, process-oriented evaluation. MedBench v5 features: (1) a dual-dimensional framework combining Clinical Cognitive Responsiveness (13 sub-dimensions) and Medical Atomic Skills (4 agent environments), covering 63 tasks; (2) three switchable information-flow stressors (omission, contradiction, evidence delay) for factorized degradation analysis; (3) a dynamic process audit protocol with five reasoning nodes that produces model-specific failure fingerprints; (4) hallucination propagation monitoring across initiation, propagation, anchoring, and contradiction interaction-capturing silent hallucination. Experiments on frontier models show that strong overall task performance does not guarantee process stability: stressors mainly disrupt contradiction detection, diagnosis updating, hallucination propagation, and contradiction-based self-correction, while final evidence grounding can remain superficially stable. MedBench v5 provides a unified infrastructure for capability profiling, controllable stress testing, process auditing, and hallucination trajectory analysis in clinical AI evaluation.
Benchmarks are necessary for healthcare evaluation, but are not sufficient for predicting deployment performance. Our position is that the evaluation--deployment gap arises not because of poorly designed benchmarks, but from implicit assumptions about how users interact with models that cannot be surfaced from benchmarks alone. To make this precise, we propose a classification of assumptions into two categories: task, which can be tested from conversation data alone, and outcome, which requires outcome data and behavioral studies for testing. Critically, outcome assumptions depend on human behavior, something that even well-designed benchmarks cannot directly observe. To demonstrate the operationality of this framework, we retrospectively analyze a healthcare RCT as a case study and find that the gap naturally separates into task and outcome gaps of roughly equal size. To address this, we make two contributions: first, we propose BenchmarkCards, an artifact that documents assumptions, and second, we propose staged evaluation, a procedure that systematically tests assumptions and evaluates performance.
Naveen Raman, Santiago Cortes-Gomez, Mateo Dulce Rubio +2
Deep learning models in medical imaging often fail when deployed in new clinical environments due to distribution shifts in demographics, scanner hardware, or acquisition protocols. A central challenge is underspecification, where models with similar validation performance exhibit divergent real-world failure modes. Although stress testing has emerged as a tool to assess this, current methods typically rely on simple, uninformed perturbations (e.g., brightness or contrast changes), which fail to capture clinically realistic variation and can overestimate robustness. In this work, we introduce a counterfactual stress testing framework based on causal generative models that create realistic "what if" images by intervening on attributes such as scanner type and patient sex while preserving anatomical identity, enabling controlled and semantically meaningful evaluation under targeted distribution shifts. Across two imaging modalities (chest X-ray and mammography), three model architectures, and multiple shift scenarios, we show that counterfactual stress tests provide a substantially more accurate proxy for real out-of-distribution performance than classical perturbations, capturing the direction and relative magnitude of performance changes as well as model ranking. These results suggest that causal generative models can serve as practical simulators for robustness assessment, offering a more reliable basis for evaluating medical AI systems prior to deployment.
Moritz Stammel, Fabio De Sousa Ribeiro, Raghav Mehta +2
Department of Computing, Imperial College London, UK