Organizations: State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China · Department of Ultrasound, State Key Laboratory of Complex Severe and Rare Diseases, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences & Peking Union Medical College, Beijing, China · College of Computer Science, Beijing University of Technology, Beijing, China
Generating ultrasound reports from multiple images requires aggregating clinical evidence across views, yet archived key frames capture only part of the dynamic examination. Raw-report imitation is therefore misaligned with visual supervision: content that is clinically valid for the full examination may be unverifiable from the images available to a model. This gap creates a clinical behavior alignment problem. A model must preserve visible findings, avoid diagnostic reversals and unsupported completion, and not collapse into conservative templates. We propose CAMEO, a Clinically Aware Multi-image Evidence-grounded Orchestration framework for ultrasound report generation. Stage I learns ultrasound visual-language primitives; Stage II performs Cross-View Evidence Grounding by distilling trusted visible report points into multi-image QA and report-style supervision; and Stage III performs Clinically Aware Preference Alignment using clinical-error-oriented preference pairs. From USReport, we construct USReport-Distilled with 17,670 evidence-grounded paired-image training instances and USReport-Pref with 21,869 preference pairs; we additionally use 25,631 PubMedVision-US ultrasound instruction samples for domain adaptation and multi-image instruction tuning. On the primary USReport-Distilled benchmark, CAMEO improves over EchoVLM from 0.25 to 0.40 BLEU-1, 0.28 to 0.45 ROUGE-1, and 0.27 to 0.43 METEOR, while raising ClinicalScore from 55.02 to 74.20. These results underscore the value of evidence-grounded supervision, clinically aware alignment, and clinically structured evaluation for reliable ultrasound report generation.
Figures & tables
Figure 1: Motivation. Archived ultrasound key frames only partially observe the dynamic examination. Raw report imitation can therefore introduce unsupported supervision, and reference-based metrics can reward non-grounded completion. CAMEO \xspace uses trusted visible evidence for cross-view training, preference alignment, and clinical error evaluation.
Figure 2: Overview of CAMEO \xspace . Raw reports are decomposed into atomic clinical points and filtered into trusted visible evidence. The retained evidence supports grounded QA, report-style SFT, clinical preference pairs, and LLM-assisted Clinical Error Evaluation, aligning training and assessment within a shared evidence space.
USR-D QA type
Question / target form
Stage-II purpose
Evidence QA
Is the report point visible in the paired images?
Filter out non-verifiable raw-report content.
View Attribution QA
Which view supports the finding: first view, second view, both, or neither?
Learn view-conditioned attribution and reduce cross-view confusion.
Cross-view QA
How do the two views jointly support or limit the finding?
Train reasoning over complementary views rather than independent single-image captioning.
Refusal QA
Should an unsupported claim be answered, qualified, or refused?
Teach conservative responses when the image pair is insufficient.
Grounded Report QA
Produce a concise report-style summary from trusted visible evidence.
Convert verified evidence into clinically readable report supervision.
Table 1: SFT supervision reconstructed from USReport-Distilled \xspace . Each QA type targets a distinct cross-view reporting behavior.
Preference contrast
Clinical axis
Rejected behavior
Desired alignment
Omission
Missed diagnosis
A visible lesion, abnormality, or key descriptor is dropped.
Preserve clinically important visible evidence.
Unsupported content
Overdiagnosis
Template-like normal findings or organ status are added without visual support.
Suppress hallucinated and template-driven content.
Clinical attribute confusion
Misdiagnosis
Echogenicity, wall status, lesion boundary, or diagnostic tendency is reversed or confused.
Penalize fluent but clinically wrong statements.
View mismatch
Misdiagnosis
Findings are assigned to the wrong view or mixed across views.
Strengthen view-specific attribution and cross-view consistency.
Diagnostic overreach
Overdiagnosis
The conclusion is stronger than the paired images justify.
Encourage conservative language under partial observability.
Template shortcut
Missed diagnosis / Overdiagnosis
The response relies on fixed benign phrases or empty caution without reporting visible evidence.
Maintain clinical informativeness while staying evidence-grounded.
Table 2: Preference data from USReport-Pref \xspace . Rejected responses instantiate clinically meaningful failure modes; chosen responses remain informative and evidence-bounded.
Figure 3: Dataset coverage and supervision flow. PubMedVision-US \xspace [ Chen et al.(2024)Chen, Gui, Ouyang, Gao, Chen, Chen, Wang, Cai, Ji, Wan, and Wang ] provides broad ultrasound instruction data, while USReport [ Li et al.(2024b)Li, Su, Zhao, Lv, Wang, Navab, Hu, and Jiang ] is distilled into USReport-Distilled \xspace and USReport-Pref \xspace for cross-view grounding, preference alignment, and clinical error evaluation.
Model
Type
Params
BLEU-1
R-1
R-L
METEOR
BERT
Miss
Misdiag
Overdiag
Clinical
PubMedVision-US Multi
LLaVA-Med [ Li et al.(2023)Li, Wong, Zhang, Usuyama, Liu, Yang, Naumann, Poon, and Gao ]
Lingshu-32B [ Xu et al.(2025)Xu, Chan, Li, Aljunied, Yuan, Wang, Xiao, Chen, Liu, Li, Sun, Shen, Wang, Tan, Zhao, Xu, Zhang, and Rong ]
medical
32B
0.33
0.42
0.25
0.35
0.25
–
–
–
–
EchoVLM [ She et al.(2026)She, Lu, Chen, Wang, and Huang ]
ultrasound
12B
0.12
0.19
0.13
0.14
0.08
–
–
–
–
CAMEO \xspace
proposed
7B
0.38
0.44
0.33
0.43
0.30
–
–
–
–
Table 3: Main results on PubMedVision-US \xspace Multi [ Chen et al.(2024)Chen, Gui, Ouyang, Gao, Chen, Chen, Wang, Cai, Ji, Wan, and Wang ] and USReport-Distilled \xspace , constructed from USReport [ Li et al.(2024b)Li, Su, Zhao, Lv, Wang, Navab, Hu, and Jiang ] . R-1, R-L, and BERT denote ROUGE-1, ROUGE-L, and BERTScore. Clinical metrics are evaluated only on USReport-Distilled \xspace ; the EchoVLM PubMedVision-US \xspace row uses all 500 English-prompt cases.
Variant
Evidence
Cross-view QA
Preference
BLEU-1
ROUGE-1
Clinical
Main purpose
LLaVA-Med [ Li et al.(2023)Li, Wong, Zhang, Usuyama, Liu, Yang, Naumann, Poon, and Gao ]
no
no
no
0.17
0.21
49.19
base model
Stage I
no
no
no
0.18
0.22
55.71
domain adaptation
Stage I + Raw Report SFT
no
partial
no
0.26
0.32
61.28
raw imitation
Stage I + Distilled Report SFT
yes
limited
no
0.36
0.41
63.01
purification
Stage I + Stage II
yes
yes
no
0.39
0.45
71.44
cross-view reasoning
CAMEO \xspace
yes
yes
yes
0.40
0.45
74.20
clinical alignment
Table 4: Component ablation study. Rows after Stage I continue from the Stage I checkpoint. Complete component metrics are provided in the supplementary material.
Organ
Miss
Misdiagnosis
Overdiagnosis
Clinical
SFT
DPO
Δ
SFT
DPO
Δ
SFT
DPO
Δ
SFT
DPO
Δ
Liver
61.52
66.11
+4.59
91.04
91.81
+0.77
92.44
92.58
+0.14
79.58
81.72
+2.14
Breast
58.58
58.55
-0.03
81.20
84.86
+3.66
88.26
88.27
+0.01
73.92
75.19
+1.27
Thyroid
37.68
44.43
+6.75
66.44
72.04
+5.60
89.97
90.75
+0.78
60.82
65.68
+4.86
All
52.59
56.36
+3.77
79.56
82.90
+3.34
90.22
90.53
+0.31
71.44
74.20
+2.76
Table 5: Organ-wise clinical error evaluation. SFT denotes Stage II and DPO denotes CAMEO \xspace . Breast corresponds to the original Mammary label.
Figure 4: Qualitative comparison on a breast/axillary case. EchoVLM and Stage I + Stage II omit the visible left-axillary nodule through normal or conservative templates, whereas CAMEO \xspace attributes the finding to the L-AX view and preserves the key descriptors.
Large vision-language models (LVLMs) have achieved strong performance across many medical imaging tasks, yet their application to ultrasound remains limited due to its inherent complexity and variability. In this work, we revisit what is truly needed to enable real-world ultrasound understanding. Instead of introducing complex architectures or elaborate training strategies, we show that data scale and clinically faithful data alignment are the key factors. We construct a large-scale dataset of 1.5M real-world ultrasound examinations, containing 17.7M images, multi-organ coverage, and paired uncurated clinical reports. Crucially, we organize the data at the examination level, aligning multiple images with their corresponding reports to reflect real clinical workflows. We then fine-tune a standard LVLM using low-rank adaptation (LoRA) on this dataset without task-specific modifications. Surprisingly, this simple recipe already leads to strong performance across diverse ultrasound understanding tasks, outperforming prior methods designed with more complex pipelines. Beyond these results, we present model and data scaling analyses that provide insights into the role of scale in ultrasound LVLMs.
Bingcong Yan, Chunlei Li, Jingliang Hu +3
1MedAI Technology (Wuxi) Co. Ltd. · 2Technical University of Munich
Ultrasound foundation models have achieved strong performance on structured prediction tasks but remain exclusively vision-based, limiting zero-shot and few-shot transfer to novel tasks where task-specific annotation is scarce. We address this gap with EchoCare-CLIP, a CLIP-style dual-encoder contrastive framework that aligns ultrasound images with clinical text in a shared embedding space. We curate a multi-organ corpus of over 16K image-text pairs spanning breast, liver, lung, and thyroid, with over 78% of captions derived from expert-annotated reports, and complement the remainder with a three-tier template-based and LLM-based caption generation pipeline. We evaluate model configurations spanning two text encoder families (CLIP, BioClinicalBERT) and two caption strategies (template-based, LLM-generated) against OpenAI CLIP and BiomedCLIP baselines. Our trained models consistently improve cross-modal alignment over baselines, with the best configuration achieving a paired alignment score of 0.682. However, stronger alignment does not guarantee better downstream performance: CLIP-based variants with partial fine-tuning achieve the strongest zero-shot classification on external held-out datasets (0.709 on BUSI; 0.626 on AULI), while full end-to-end fine-tuning degrades transfer due to overfitting. On linear probing and few-shot adaptation, model rankings are dataset-dependent, reflecting a trade-off between domain adaptation and representational generalizability. We further show that template-based captions match or outperform LLM-generated captions, suggesting lexical diversity is not a proxy for caption quality. Taken together, our results demonstrate that ultrasound vision-language alignment is achievable from public data alone, but robust clinical transfer requires careful balancing of domain adaptation, encoder capacity, and caption supervision quality.
Zhuoyang Lyu, Yiyang Zhang, Tongxin Wang +1
Department of Biostatistics Harvard T. H. Chan School of Public Health Boston, MA, USA
Echocardiography is the most widely used non-invasive cardiac imaging modality, providing essential information for cardiovascular diagnosis. Interpreting an echocardiogram requires synthesizing complementary evidence across multiple heart views to identify abnormalities and produce structured clinical reports. While recent efforts focus on improving classification performance, most models lack explicit diagnostic reasoning and spatially grounded anatomical evidence, limiting clinician trust. We present EchoSonar-R, a multi-view reasoning-enabled vision-language model that jointly performs multi-label disease classification and report generation from echocardiography studies. EchoSonar-R combines a spatiotemporal video encoder with a structure-aware cardiac detector that provides spatially grounded anatomical cues to improve interpretability and clinician trust during cross-view reasoning. EchoSonar-R is trained in two stages: supervised fine-tuning (SFT) on reasoning-annotated targets, followed by Group Relative Policy Optimization (GRPO) with task-specific rewards that jointly align classification and report generation within a unified reinforcement-learning framework. Across a private multi-view dataset and two public benchmarks, EchoSonar-R improves macro balanced accuracy by 17.1% on the private set and 6.1% on MIMICEchoQA over the strongest baseline, achieves a GREEN clinical faithfulness score of 0.800, and produces interpretable reasoning traces grounded in multi-view visual evidence.
Darya Taratynova, Ahmed Aly, Numan Saeed +1
Division of Computing and Mathematical Sciences, Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, UAE