Organizations: State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China · Department of Ultrasound, State Key Laboratory of Complex Severe and Rare Diseases, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences & Peking Union Medical College, Beijing, China · College of Computer Science, Beijing University of Technology, Beijing, China
Generating ultrasound reports from multiple images requires aggregating clinical evidence across views, yet archived key frames capture only part of the dynamic examination. Raw-report imitation is therefore misaligned with visual supervision: content that is clinically valid for the full examination may be unverifiable from the images available to a model. This gap creates a clinical behavior alignment problem. A model must preserve visible findings, avoid diagnostic reversals and unsupported completion, and not collapse into conservative templates. We propose CAMEO, a Clinically Aware Multi-image Evidence-grounded Orchestration framework for ultrasound report generation. Stage I learns ultrasound visual-language primitives; Stage II performs Cross-View Evidence Grounding by distilling trusted visible report points into multi-image QA and report-style supervision; and Stage III performs Clinically Aware Preference Alignment using clinical-error-oriented preference pairs. From USReport, we construct USReport-Distilled with 17,670 evidence-grounded paired-image training instances and USReport-Pref with 21,869 preference pairs; we additionally use 25,631 PubMedVision-US ultrasound instruction samples for domain adaptation and multi-image instruction tuning. On the primary USReport-Distilled benchmark, CAMEO improves over EchoVLM from 0.25 to 0.40 BLEU-1, 0.28 to 0.45 ROUGE-1, and 0.27 to 0.43 METEOR, while raising ClinicalScore from 55.02 to 74.20. These results underscore the value of evidence-grounded supervision, clinically aware alignment, and clinically structured evaluation for reliable ultrasound report generation.
Figures & tables
Figure 1: Motivation. Archived ultrasound key frames only partially observe the dynamic examination. Raw report imitation can therefore introduce unsupported supervision, and reference-based metrics can reward non-grounded completion. CAMEO \xspace uses trusted visible evidence for cross-view training, preference alignment, and clinical error evaluation.
Figure 2: Overview of CAMEO \xspace . Raw reports are decomposed into atomic clinical points and filtered into trusted visible evidence. The retained evidence supports grounded QA, report-style SFT, clinical preference pairs, and LLM-assisted Clinical Error Evaluation, aligning training and assessment within a shared evidence space.
USR-D QA type
Question / target form
Stage-II purpose
Evidence QA
Is the report point visible in the paired images?
Filter out non-verifiable raw-report content.
View Attribution QA
Which view supports the finding: first view, second view, both, or neither?
Learn view-conditioned attribution and reduce cross-view confusion.
Cross-view QA
How do the two views jointly support or limit the finding?
Train reasoning over complementary views rather than independent single-image captioning.
Refusal QA
Should an unsupported claim be answered, qualified, or refused?
Teach conservative responses when the image pair is insufficient.
Grounded Report QA
Produce a concise report-style summary from trusted visible evidence.
Convert verified evidence into clinically readable report supervision.
Table 1: SFT supervision reconstructed from USReport-Distilled \xspace . Each QA type targets a distinct cross-view reporting behavior.
Preference contrast
Clinical axis
Rejected behavior
Desired alignment
Omission
Missed diagnosis
A visible lesion, abnormality, or key descriptor is dropped.
Preserve clinically important visible evidence.
Unsupported content
Overdiagnosis
Template-like normal findings or organ status are added without visual support.
Suppress hallucinated and template-driven content.
Clinical attribute confusion
Misdiagnosis
Echogenicity, wall status, lesion boundary, or diagnostic tendency is reversed or confused.
Penalize fluent but clinically wrong statements.
View mismatch
Misdiagnosis
Findings are assigned to the wrong view or mixed across views.
Strengthen view-specific attribution and cross-view consistency.
Diagnostic overreach
Overdiagnosis
The conclusion is stronger than the paired images justify.
Encourage conservative language under partial observability.
Template shortcut
Missed diagnosis / Overdiagnosis
The response relies on fixed benign phrases or empty caution without reporting visible evidence.
Maintain clinical informativeness while staying evidence-grounded.
Table 2: Preference data from USReport-Pref \xspace . Rejected responses instantiate clinically meaningful failure modes; chosen responses remain informative and evidence-bounded.
Figure 3: Dataset coverage and supervision flow. PubMedVision-US \xspace [ Chen et al.(2024)Chen, Gui, Ouyang, Gao, Chen, Chen, Wang, Cai, Ji, Wan, and Wang ] provides broad ultrasound instruction data, while USReport [ Li et al.(2024b)Li, Su, Zhao, Lv, Wang, Navab, Hu, and Jiang ] is distilled into USReport-Distilled \xspace and USReport-Pref \xspace for cross-view grounding, preference alignment, and clinical error evaluation.
Model
Type
Params
BLEU-1
R-1
R-L
METEOR
BERT
Miss
Misdiag
Overdiag
Clinical
PubMedVision-US Multi
LLaVA-Med [ Li et al.(2023)Li, Wong, Zhang, Usuyama, Liu, Yang, Naumann, Poon, and Gao ]
Lingshu-32B [ Xu et al.(2025)Xu, Chan, Li, Aljunied, Yuan, Wang, Xiao, Chen, Liu, Li, Sun, Shen, Wang, Tan, Zhao, Xu, Zhang, and Rong ]
medical
32B
0.33
0.42
0.25
0.35
0.25
–
–
–
–
EchoVLM [ She et al.(2026)She, Lu, Chen, Wang, and Huang ]
ultrasound
12B
0.12
0.19
0.13
0.14
0.08
–
–
–
–
CAMEO \xspace
proposed
7B
0.38
0.44
0.33
0.43
0.30
–
–
–
–
Table 3: Main results on PubMedVision-US \xspace Multi [ Chen et al.(2024)Chen, Gui, Ouyang, Gao, Chen, Chen, Wang, Cai, Ji, Wan, and Wang ] and USReport-Distilled \xspace , constructed from USReport [ Li et al.(2024b)Li, Su, Zhao, Lv, Wang, Navab, Hu, and Jiang ] . R-1, R-L, and BERT denote ROUGE-1, ROUGE-L, and BERTScore. Clinical metrics are evaluated only on USReport-Distilled \xspace ; the EchoVLM PubMedVision-US \xspace row uses all 500 English-prompt cases.
Variant
Evidence
Cross-view QA
Preference
BLEU-1
ROUGE-1
Clinical
Main purpose
LLaVA-Med [ Li et al.(2023)Li, Wong, Zhang, Usuyama, Liu, Yang, Naumann, Poon, and Gao ]
no
no
no
0.17
0.21
49.19
base model
Stage I
no
no
no
0.18
0.22
55.71
domain adaptation
Stage I + Raw Report SFT
no
partial
no
0.26
0.32
61.28
raw imitation
Stage I + Distilled Report SFT
yes
limited
no
0.36
0.41
63.01
purification
Stage I + Stage II
yes
yes
no
0.39
0.45
71.44
cross-view reasoning
CAMEO \xspace
yes
yes
yes
0.40
0.45
74.20
clinical alignment
Table 4: Component ablation study. Rows after Stage I continue from the Stage I checkpoint. Complete component metrics are provided in the supplementary material.
Organ
Miss
Misdiagnosis
Overdiagnosis
Clinical
SFT
DPO
Δ
SFT
DPO
Δ
SFT
DPO
Δ
SFT
DPO
Δ
Liver
61.52
66.11
+4.59
91.04
91.81
+0.77
92.44
92.58
+0.14
79.58
81.72
+2.14
Breast
58.58
58.55
-0.03
81.20
84.86
+3.66
88.26
88.27
+0.01
73.92
75.19
+1.27
Thyroid
37.68
44.43
+6.75
66.44
72.04
+5.60
89.97
90.75
+0.78
60.82
65.68
+4.86
All
52.59
56.36
+3.77
79.56
82.90
+3.34
90.22
90.53
+0.31
71.44
74.20
+2.76
Table 5: Organ-wise clinical error evaluation. SFT denotes Stage II and DPO denotes CAMEO \xspace . Breast corresponds to the original Mammary label.
Figure 4: Qualitative comparison on a breast/axillary case. EchoVLM and Stage I + Stage II omit the visible left-axillary nodule through normal or conservative templates, whereas CAMEO \xspace attributes the finding to the L-AX view and preserves the key descriptors.