Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study
Authors: Lin Wu, Zhe Xu, Hongyi Wang, Feifei Zhou, Wei Deng, Chunlong Zhang, Yuting Zhu, Kaixiao Chen, +5 more
Organizations: Department of Radiology, The First Affiliated Hospital, Jiangxi Medical College, Nanchang University · Jiangxi Province Medical Imaging Research Institute · Department of Computer Science and Engineering, Hong Kong University of Science and Technology, Hong Kong, China · Department of Radiology, Shangrao City People’s Hospital, Shangrao, China · Department of Radiology, Nanchang People’s Hospital, Nanchang, China · Department of Chemical and Biological Engineering, Hong Kong University of Science and Technology, Hong Kong SAR, China · Division of Life Science, Hong Kong University of Science and Technology, Hong Kong SAR, China · State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology, Hong Kong SAR, China · HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute, The Hong Kong University of Science and Technology, Futian, Shenzhen, China
Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect. Materials and Methods: This prospective, multicenter, randomized three-arm reader study was conducted at three hospitals in China from July to September 2026 (ChiCTR2600129243). After specialty stratification, 132 residents with fewer than 3 years of clinical experience were randomized 1:1:1 to GPT-5.4 alone (group A), GPT-5.4 plus Kimi-K2.6 (group B), or GPT-5.4 plus Gemini-3.6 Flash (group C); 123 were analyzed. Participants interpreted 60 radiographs before and after AI support. The primary outcome was accuracy change. Welch ANOVA and Holm-adjusted t tests compared support conditions; HC3 linear models assessed specialty interaction. Results: Among 123 residents (mean age, 24.1 years +/- 1.4; 65 women), radiology residents showed greater accuracy improvement with dual- than single-suggestion support (B-A, 6.69 percentage points [95% CI, 0.97-12.40]; C-A, 7.87 percentage points [95% CI, 1.64-14.11]; Holm-adjusted P = .030 for both), whereas accuracy change did not differ in non-radiology residents (P = .20). When GPT-5.4 was incorrect, AI-assisted accuracy was higher with dual- than single-suggestion support in radiology residents (40.1% and 40.4% vs 20.0%) and non-radiology residents (31.3% and 31.0% vs 12.1%) (all Holm-adjusted P < .001). The dual-suggestion effect differed by specialty (interaction difference, 10.44 percentage points; 95% CI, 4.36-16.52; P < .001). Conclusion: Dual-suggestion support may mitigate the influence of erroneous AI suggestions, with greater accuracy improvement observed in radiology but not non-radiology residents.
Figures & tables
Figure 1: Participant Flow Diagram. Participants were stratified by specialty and allocated 1:1:1 to the three AI support conditions. Testing was performed using a smartphone or tablet. Network failure was defined as connectivity or webpage failure preventing image display or response submission; physical discomfort was defined as feeling unwell enough to interfere with concentration and require termination of testing. Final analytic samples are shown by specialty.
Figure 2: Reader Workflow and AI Support Conditions. a, Initial interpretation without AI assistance, including diagnosis and confidence. b, Final interpretation after AI assistance with single-suggestion support (group A, GPT-5.4) or one of two parallel dual-suggestion conditions (group B, GPT-5.4 plus Kimi-K2.6; group C, GPT-5.4 plus Gemini-3.6 Flash). In both dual-suggestion conditions, the two AI suggestions were presented simultaneously.
Figure 3: AI Model Concordance and Diagnostic Accuracy With and Without AI Assistance. a–c, AI diagnostic performance and case-level concordance in groups A–C. In the dual-suggestion groups, concordant and discordant categories indicate whether the second model agreed with GPT-5.4 among GPT-5.4-correct and GPT-5.4-incorrect cases. d–f, Mean AI-unassisted and AI-assisted diagnostic accuracy among all residents, radiology residents, and non-radiology residents in groups A–C. Error bars represent 95% CIs. *** P<.001 for paired comparisons between AI-unassisted and AI-assisted accuracy.
Figure 4: Diagnostic Outcomes by Support Condition. a–c, Accuracy change, interpretation time, and confidence change, respectively, among radiology residents. d–f, Corresponding outcomes among non-radiology residents. Bars represent mean values, and error bars represent 95% CIs. *Holm-adjusted P<.05 versus group A. ns indicates a nonsignificant overall comparison across the three support conditions.
Figure 5: Diagnostic Accuracy According to GPT-5.4 Suggestion Correctness. a–b, AI-unassisted and AI-assisted diagnostic accuracy among radiology residents in cases with correct ( n=39 ) and incorrect ( n=21 ) GPT-5.4 suggestions, respectively. c–d, Corresponding results among non-radiology residents. Lines represent groups A–C, and error bars represent 95% CIs. ***Holm-adjusted P<.001 versus group A.
Figure 6: Revision Components and Dual-Suggestion Effects by Specialty. a, Definition of revision components. Beneficial and harmful revisions represent incorrect-to-correct and correct-to-incorrect transitions, respectively; gray flows indicate no change. b, Revision components of accuracy change by support condition and specialty. Bars show means with 95% CIs; harmful revisions are plotted below zero for visualization. c, Dual-suggestion effects by specialty, defined as [(B+C)/2]–A. Points show mean differences with 95% CIs; P for interaction indicates differences between specialties.
Figure S1: Composition of the 60-case set. Thirty examinations were chest radiographs and 30 were abdominal radiographs. Urinary tract calculi include renal and ureteral calculi. Normal abdominal examinations are shown separately from disease categories.
Both correct
GPT correct / second model incorrect
GPT incorrect / second model correct
Both incorrect
Figure S2: Correctness-status maps for the single- and dual-suggestion conditions. Each square represents one case, numbered 1–60 in the same order in both dual-suggestion panels. Group A has no second-model status and uses the same two-color key as the examples below (teal, correct; rust, incorrect). The four colors below apply only to groups B and C. Both-incorrect cases count as correctness-concordant even when the diagnostic labels differ.
An AI-generated radiology report can resemble a physician's report while omitting an abnormality, adding an unsupported finding, or reversing its presence. Measuring these factual differences is essential for evaluating report generators. We study Jev, a System One decision model, as a simple, low-cost judge of agreement with physician-written reference reports. Our evaluator checks whether each statement is supported by the other report and combines these judgments in both directions to capture unsupported claims and omissions. A single-question configuration reaches Kendall correlations of 0.573 on RadEvalX and 0.398 on RadEvalExpert with expert error counts, outperforming an open natural language inference judge under matched decomposition and aggregation. One support question per statement retains similar expert agreement to seven while using 43-45% fewer judgment input tokens. At the documented API price, judgments cost under three cents per hundred report pairs, excluding local decomposition. In a separate controlled-error test, Jev detects false negation with an AUROC of 0.977. Local RadMatch achieves stronger agreement on clinically significant errors in both expert datasets and on total errors in the shared RadEvalExpert subset. Finding-count and error-scope analyses show that benchmark agreement reflects report size and error definitions as well as medical error detection. These results support Jev as a practical judgment component for measuring factual differences in generated radiology reports and identify where more elaborate evaluation remains valuable.
Jiaju Huang, Hao Yang, Xinyu Ma +6
Intelligent Medical Computing Laboratory, Faculty of Applied Sciences, Macao Polytechnic University · Department of Radiology, Netherlands Cancer Institute, Amsterdam, the Netherlands · Department of Radiology and Nuclear Medicine, Radboud University Medical Center, Nijmegen, the Netherlands
Vision-language models that answer questions about chest radiographs are evaluated by their accuracy on labels derived from radiology reports. High benchmark accuracy is often interpreted as evidence that the model uses the image. A model that answers from the finding named in the question can score as well as a model that uses the radiograph. Keeping the question fixed, we audit eight open-weight systems by swapping in another patient's radiograph with the same or the opposite label, occluding the radiologist-marked region or an equal region elsewhere, and removing the radiograph or replacing it with noise or a photograph. On 2,548 yes-or-no questions from MIMIC-CXR, one multimodal model answers Yes regardless of the image, another multimodal model changes its answers without following the label, and four systems use the image but keep about half of their correct answers when the radiograph is swapped for an opposite-label radiograph. A medical model that receives only the question text scores 55.3% on the pooled questions, higher than two multimodal systems. It scores 91.8% where every finding is present, and answering Yes to every question scores 100% there. Where the image is necessary, the best multimodal system exceeds this model by 10.4% in balanced accuracy. The categories are unchanged on CheXpert. Confidence is not higher when a correct answer depends on the marked region. In a reader study with three radiologists, the two radiologists who read a balanced set of 200 cases score 86.0% and 82.0%, and the systems score 50.0% to 73.0%. Accuracy does not establish image use, but an intervention on the image can test it.
Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams +4
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany · Department of Diagnostic and Interventional Radiology, TUM University Clinic, School of Medicine and Health, Klinikum rechts der Isar, Technical University of Munich, Munich, Germany · Lab for AI in Medicine, RWTH Aachen University, Aachen, Germany +1
Vision-language models (VLM) have markedly advanced AI-driven interpretation and reporting of complex medical imaging, such as computed tomography (CT). Yet, existing methods largely relegate clinicians to passive observers of final outputs, offering no interpretable reasoning trace for them to inspect, validate, or refine. To address this, we introduce RadAgent, a tool-using AI agent that generates CT reports through a stepwise and interpretable process. Each resulting report is accompanied by a fully inspectable trace of intermediate decisions and tool interactions, allowing clinicians to examine how the reported findings are derived. In our experiments, we observe that RadAgent improves chest CT report generation over its 3D VLM counterpart, CT-Chat, across three dimensions. Clinical accuracy improves by 5.8 points (35.4% relative) in macro-F1 and 5.1 points (18.6% relative) in micro-F1. Robustness under adversarial conditions improves by 24.7 points (41.9% relative). Furthermore, RadAgent achieves 37.0% in faithfulness, a new capability entirely absent in its 3D VLM counterpart. By structuring the interpretation of chest CT as an explicit, tool-augmented and iterative reasoning trace, RadAgent brings us closer toward transparent and reliable AI for radiology.
Mélanie Roschewitz, Kenneth Styppa, Yitian Tao +10
Department of Biosystems Science and Engineering, ETH Zurich, Basel, Switzerland · ETH AI Center, Zurich, Switzerland · Department of Computer Science, ETH Zurich, Zurich, Switzerland +5