Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study
Authors: Lin Wu, Zhe Xu, Hongyi Wang, Feifei Zhou, Wei Deng, Chunlong Zhang, Yuting Zhu, Kaixiao Chen, +5 more
Organizations: Department of Radiology, The First Affiliated Hospital, Jiangxi Medical College, Nanchang University · Jiangxi Province Medical Imaging Research Institute · Department of Computer Science and Engineering, Hong Kong University of Science and Technology, Hong Kong, China · Department of Radiology, Shangrao City People’s Hospital, Shangrao, China · Department of Radiology, Nanchang People’s Hospital, Nanchang, China · Department of Chemical and Biological Engineering, Hong Kong University of Science and Technology, Hong Kong SAR, China · Division of Life Science, Hong Kong University of Science and Technology, Hong Kong SAR, China · State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology, Hong Kong SAR, China · HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute, The Hong Kong University of Science and Technology, Futian, Shenzhen, China
Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect. Materials and Methods: This prospective, multicenter, randomized three-arm reader study was conducted at three hospitals in China from July to September 2026 (ChiCTR2600129243). After specialty stratification, 132 residents with fewer than 3 years of clinical experience were randomized 1:1:1 to GPT-5.4 alone (group A), GPT-5.4 plus Kimi-K2.6 (group B), or GPT-5.4 plus Gemini-3.6 Flash (group C); 123 were analyzed. Participants interpreted 60 radiographs before and after AI support. The primary outcome was accuracy change. Welch ANOVA and Holm-adjusted t tests compared support conditions; HC3 linear models assessed specialty interaction. Results: Among 123 residents (mean age, 24.1 years +/- 1.4; 65 women), radiology residents showed greater accuracy improvement with dual- than single-suggestion support (B-A, 6.69 percentage points [95% CI, 0.97-12.40]; C-A, 7.87 percentage points [95% CI, 1.64-14.11]; Holm-adjusted P = .030 for both), whereas accuracy change did not differ in non-radiology residents (P = .20). When GPT-5.4 was incorrect, AI-assisted accuracy was higher with dual- than single-suggestion support in radiology residents (40.1% and 40.4% vs 20.0%) and non-radiology residents (31.3% and 31.0% vs 12.1%) (all Holm-adjusted P < .001). The dual-suggestion effect differed by specialty (interaction difference, 10.44 percentage points; 95% CI, 4.36-16.52; P < .001). Conclusion: Dual-suggestion support may mitigate the influence of erroneous AI suggestions, with greater accuracy improvement observed in radiology but not non-radiology residents.
Figures & tables
Figure 1: Participant Flow Diagram. Participants were stratified by specialty and allocated 1:1:1 to the three AI support conditions. Testing was performed using a smartphone or tablet. Network failure was defined as connectivity or webpage failure preventing image display or response submission; physical discomfort was defined as feeling unwell enough to interfere with concentration and require termination of testing. Final analytic samples are shown by specialty.
Figure 2: Reader Workflow and AI Support Conditions. a, Initial interpretation without AI assistance, including diagnosis and confidence. b, Final interpretation after AI assistance with single-suggestion support (group A, GPT-5.4) or one of two parallel dual-suggestion conditions (group B, GPT-5.4 plus Kimi-K2.6; group C, GPT-5.4 plus Gemini-3.6 Flash). In both dual-suggestion conditions, the two AI suggestions were presented simultaneously.
Figure 3: AI Model Concordance and Diagnostic Accuracy With and Without AI Assistance. a–c, AI diagnostic performance and case-level concordance in groups A–C. In the dual-suggestion groups, concordant and discordant categories indicate whether the second model agreed with GPT-5.4 among GPT-5.4-correct and GPT-5.4-incorrect cases. d–f, Mean AI-unassisted and AI-assisted diagnostic accuracy among all residents, radiology residents, and non-radiology residents in groups A–C. Error bars represent 95% CIs. *** P<.001 for paired comparisons between AI-unassisted and AI-assisted accuracy.
Figure 4: Diagnostic Outcomes by Support Condition. a–c, Accuracy change, interpretation time, and confidence change, respectively, among radiology residents. d–f, Corresponding outcomes among non-radiology residents. Bars represent mean values, and error bars represent 95% CIs. *Holm-adjusted P<.05 versus group A. ns indicates a nonsignificant overall comparison across the three support conditions.
Figure 5: Diagnostic Accuracy According to GPT-5.4 Suggestion Correctness. a–b, AI-unassisted and AI-assisted diagnostic accuracy among radiology residents in cases with correct ( n=39 ) and incorrect ( n=21 ) GPT-5.4 suggestions, respectively. c–d, Corresponding results among non-radiology residents. Lines represent groups A–C, and error bars represent 95% CIs. ***Holm-adjusted P<.001 versus group A.
Figure 6: Revision Components and Dual-Suggestion Effects by Specialty. a, Definition of revision components. Beneficial and harmful revisions represent incorrect-to-correct and correct-to-incorrect transitions, respectively; gray flows indicate no change. b, Revision components of accuracy change by support condition and specialty. Bars show means with 95% CIs; harmful revisions are plotted below zero for visualization. c, Dual-suggestion effects by specialty, defined as [(B+C)/2]–A. Points show mean differences with 95% CIs; P for interaction indicates differences between specialties.
Figure S1: Composition of the 60-case set. Thirty examinations were chest radiographs and 30 were abdominal radiographs. Urinary tract calculi include renal and ureteral calculi. Normal abdominal examinations are shown separately from disease categories.
Both correct
GPT correct / second model incorrect
GPT incorrect / second model correct
Both incorrect
Figure S2: Correctness-status maps for the single- and dual-suggestion conditions. Each square represents one case, numbered 1–60 in the same order in both dual-suggestion panels. Group A has no second-model status and uses the same two-color key as the examples below (teal, correct; rust, incorrect). The four colors below apply only to groups B and C. Both-incorrect cases count as correctness-concordant even when the diagnostic labels differ.
Intelligent Medical Computing Laboratory, Faculty of Applied Sciences, Macao Polytechnic University · Department of Radiology, Netherlands Cancer Institute, Amsterdam, the Netherlands · Department of Radiology and Nuclear Medicine, Radboud University Medical Center, Nijmegen, the Netherlands
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany · Department of Diagnostic and Interventional Radiology, TUM University Clinic, School of Medicine and Health, Klinikum rechts der Isar, Technical University of Munich, Munich, Germany · Lab for AI in Medicine, RWTH Aachen University, Aachen, Germany +1
Department of Biosystems Science and Engineering, ETH Zurich, Basel, Switzerland · ETH AI Center, Zurich, Switzerland · Department of Computer Science, ETH Zurich, Zurich, Switzerland +5