Match4Annotate: Cross-Video Annotation Transfer in Ultrasound via Implicit Feature Flow-Guided Matching
Authors: Zhuorui Zhang, Roger Pallarès-López, Praneeth Namburi, Brian W. Anthony
Organizations: Department of Mechanical Engineering, Massachusetts Institute of Technology, Cambridge, USA · Institute for Medical Engineering and Science, Massachusetts Institute of Technology, Cambridge, USA · MIT.nano Immersion Lab, Massachusetts Institute of Technology, Cambridge, USA
Acquiring per-frame annotations for ultrasound videos is costly and requires clinical expertise, limiting learning-based analysis. We study cross-video annotation transfer: propagating user-specified annotations from a labeled ultrasound video to an independently acquired target video with no target-side labels or manual initialization. Video trackers and segmentation propagators rely on temporal continuity and require a prompt in every new sequence, whereas cross-image feature matching and one-shot segmentation estimate correspondences independently, without enforcing coherent deformations or supporting both point and mask annotations. We present Match4Annotate, a test-time framework with three stages. A spatiotemporal implicit feature representation lifts frozen vision foundation-model features into a continuous field over space and time, enabling queries beyond the backbone resolution. A continuous implicit feature flow then aligns the source and target fields under a smooth-deformation prior, estimating correspondence in feature space rather than relying on intensity consistency, which is often violated in ultrasound by speckle and acquisition-dependent appearance. Finally, flow-guided annotation transfer uses the estimated flow as a spatial prior over feature similarity. This formulation unifies sparse point and dense mask transfer and includes unconstrained feature matching and direct flow warping as limiting cases. On four clinical ultrasound datasets spanning echocardiography and musculoskeletal imaging, Match4Annotate achieves state-of-the-art annotation transfer, outperforming dense feature-matching baselines across PCK thresholds and one-shot segmentation methods in Dice score. It also demonstrates bidirectional transfer of left-ventricular annotations across datasets. It requires no task-specific training and adapts to each video in minutes on a single consumer GPU.
Figures & tables
Figure 1: Overview of Match4Annotate: annotation transfer across videos and datasets, shown over time. A single source annotation (left) is transferred to a different target video and shown across time as a temporal strip of target frames. Within-dataset (source and target both EchoNet): (a) Boundary-point tracking (colors encode correspondence identity across frames) and (b) LV-cavity mask transfer; ground truth (white × and yellow dashed contour) is shown on the labeled target frame. Cross-dataset : (c) The LV cavity, annotated in both datasets, is transferred EchoNet → CAMUS (ground-truth contour in yellow), and (d) The myocardium, annotated in CAMUS but not EchoNet, is borrowed from CAMUS and transferred onto EchoNet; with no target ground truth, EchoNet’s own LV-cavity contour (magenta) is shown for anatomical context. Predicted masks in green, source annotations in cyan.
Figure 2: Match4Annotate workflow. Spatiotemporal implicit feature representation fits a per-video implicit field to frozen DINOv3 features, yielding a continuous representation that can be queried at arbitrary spatial and temporal coordinates. Continuous implicit feature flow then estimates a smooth spatial displacement between a source and a target frame from these representations. Flow-guided annotation transfer converts that displacement into a spatial prior on feature-similarity matching, transferring annotated points and masks onto the target. Padlocks mark modules frozen at that stage.
Figure 3: Cross-video point transfer on EchoNet (left), MSK-Bone (middle), and EchoNet-LVH (right) comparing with DIFT and MATCHA. Each cell shows a source frame with annotated points and a target frame with transferred predictions; green lines indicate correct correspondences (within the PCK radius) and red lines indicate incorrect ones.
Method
EchoNet
MSK-Bone
EchoNet-LVH
RoMa (in.)
5.2/15.8/39.2
8.1/20.0/37.9
4.5/33.4/64.2
RoMa (out.)
5.5/17.1/41.1
6.9/17.8/34.5
3.8/29.5/59.8
DIFT
4.0/11.7/27.2
16.3/36.6/62.4
4.4/29.5/49.9
MATCHA
5.7/17.8/40.8
13.5/36.4/66.7
3.9/39.5/68.4
Ours
6.1/18.4/46.7
15.5/40.0/74.3
11.9/34.7/69.5
Table 1: Cross-video point transfer. PCK@ {.016,.031,.063}↑ on EchoNet (200), MSK-Bone (380), and EchoNet-LVH (200 pairs). Bold best, underline second.
Method
EchoNet
MSK-Bone
CAMUS-TED
Few-Shot Segmentation
UniverSeg (1-shot)
47.1±16.2
25.0±11.0
57.4±13.9
UniverSeg (3-shot) ∗
73.0±14.0
53.2±14.9
76.3±10.5
UniverSeg (5-shot) ∗
76.3±12.5
63.7±12.9
79.2±9.8
UniverSeg (10-shot) ∗
78.7±12.2
73.4±8.9
82.5±9.5
One-Shot Segmentation
Table 2: Cross-video mask transfer. Dice ↑ on EchoNet (200), MSK-Bone (380), and CAMUS-TED (200 pairs, all frames). ∗ multi-shot (not one-shot-comparable). Bold best, underline second.
Method
EchoNet → CAMUS
CAMUS → EchoNet
Few-Shot Segmentation
UniverSeg (1-shot)
44.5±15.7
42.6±15.6
UniverSeg (3-shot) ∗
68.7±13.1
62.2±16.4
UniverSeg (5-shot) ∗
73.0±12.2
66.5±16.3
UniverSeg (10-shot) ∗
76.5±11.1
68.4±16.4
One-Shot Segmentation
Table 3: Cross-dataset mask transfer. Dice ↑ for LV-cavity transfer between EchoNet-Dynamic and CAMUS-TED. EchoNet → CAMUS evaluates all target frames, while CAMUS → EchoNet evaluates ED and ES frames. ∗ Multi-shot results are not directly comparable. Among one-shot methods, bold denotes best and underline second best.
Figure 4: Cross-video mask transfer on EchoNet (left), MSK-Bone (middle), and CAMUS-TED (right) comparing with Matcher (SemSAM) and UniverSeg (1-shot). Each cell shows the target frame with the transferred mask (green) and ground-truth contour (yellow dashed); the leftmost column of each panel shows the source reference mask (cyan).
Figure 5: Cross-dataset annotation transfer between EchoNet and CAMUS-TED. Each panel shows two source → target examples; a source frame with its annotation (cyan) is followed by the target frame with the transferred mask (green). (a) EchoNet → CAMUS and (b) CAMUS → EchoNet transfer the shared left-ventricle cavity; the ground-truth contour is shown as a yellow dashed line. (c) The myocardium is annotated in CAMUS but not EchoNet, so it is borrowed from CAMUS and transferred onto EchoNet; with no target ground truth, the transferred wall is shown against EchoNet’s own LV-cavity contour (magenta) for anatomical context.
Backbone
PCK ↑ .016/.031/.063
Dice ↑
ImageNet ViT-S/16
1.8/6.6/20.6
51.4±20.3
CLIP ViT-B/16
2.5/8.9/27.5
61.0±17.6
DINOv2 ViT-S/14
2.9/10.3/30.8
62.3±20.7
SAM 2 image encoder
5.2/17.6/43.4
69.3±15.6
DINOv3 ViT-S/16 (ours)
6.1/18.4/46.7
76.3±12.3
Table 4: Backbone ablation on EchoNet cross-video transfer (200 pairs). The frozen feature extractor supervising fθ is varied; all other components are fixed. Best in bold .
Variant
Feat. INR
Flow Prior
PCK ↑ .016/.031/.063
Dice ↑
Full model
✓
✓
6.1/18.4/46.7
76.3±12.3
Matching Strategy:
w/o prior
✓
5.2/16.2/39.9
71.4±14.4
prior from source
✓
2.0/7.9/25.0
63.4±14.8
Feature Representation:
Low-Res ( 282 )
✓
3.4/12.6/36.8
72.2±12.0
Table 5: Ablation studies on EchoNet cross-video transfer (200 pairs). Feature INR : the spatiotemporal implicit feature field (Sec. 3.3 ); Flow Prior : the flow-guided spatial prior (Sec. 3.4 and 3.5 ). Bold best, underline second best; the full model is our configuration.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Architecture
Params
Recon. Loss ↓
RMSE ↓
SIREN (ours)
165,520
0.0038±0.0004
0.0613
ReLU MLP
165,520
0.0237±0.0026
0.1540
ReLU+PE ( L=3 )
170,128
0.0122±0.0019
0.1103
Instant-NGP
16,950,160
0.0253±0.0058
0.1591
Appendix
Table S1: Architecture comparison. Reconstruction quality on 5 EchoNet-Dynamic videos under an identical 500-epoch budget (mean ± std). All four use 2 hidden layers of width 256; SIREN uses ω=30 , ReLU+PE uses L=3 positional-encoding frequencies ( Mildenhall et al. 2020 ) (higher values introduce crosshatch artifacts), and Instant-NGP ( Müller et al. 2022 ) adds a 16-level hash grid, which is where its extra 16.8 M parameters go. Best in bold .
Figure S1: PCA feature maps. RGB visualization of the first three PCA components of the reconstructed feature field, for the first frame of three EchoNet-Dynamic videos (one per row). Columns, left to right: the low-resolution DINOv3 target, then SIREN, ReLU MLP, ReLU+PE and Instant-NGP. SIREN produces the most faithful reconstruction; the alternatives show blocky artifacts (ReLU MLP, ReLU+PE) or fragmented features (Instant-NGP).
Quantity
fθ
gϕ
Hdim
Hidden width
256
128
L
Hidden layers
2
1
D
Output dimension
384
2
ω
Activation frequency
30.0
30.0
–
Training epochs
500
1000
–
Batch size (coords)
1024
–
Appendix
Table S2: Complete hyperparameters for the two implicit networks: the spatiotemporal feature field fθ (Sec. 3.3 ) and the continuous flow field gϕ (Sec. 3.4 ). Both are optimized per video at test time with Adam ( Kingma and Ba 2015 ) ( β1=0.9 , β2=0.999 ) at learning rate 10−4 , with the backbone frozen. This table lists every quantity Sec. 4.1 names. fθ ’s reconstruction grid is 112×112 for EchoNet-Dynamic, EchoNet-LVH and CAMUS-TED, and 224×224 for the higher-resolution MSK-Bone. λ1 and λ2 weight the smoothness (TV) and magnitude ( ℓ1 ) terms of the flow loss; dmin is the interior-point margin of Eq. ( 8 ); σκ is the matching-prior bandwidth of Eq. ( 7 ), and σKDE / τ are per dataset in Table S3 .
Setting
Canvas
Adopted
EchoNet-Dynamic
112×112
6.0/0.25
MSK-Bone
656×496
2.0/0.01
CAMUS-TED
448×448
3.5/0.15
Cross-dataset
448×448
4.0/0.10
Appendix
Table S3: KDE grid search. Bandwidth σKDE and threshold τ can be tuned jointly per dataset.
Figure S2: Cross-video point transfer, EchoNet-Dynamic. Each row is one source → target pair. Left: the source video’s annotated frame. Right: its LV boundary points transferred onto three evenly spaced frames of a different video. Ground truth is shown on the target frames that carry it (EchoNet-Dynamic labels only end-diastole and end-systole).
Figure S3: Cross-video point transfer, MSK-Bone. As Figure S2 , for humerus boundary points transferred between subjects. The larger native canvas ( 656×496 ) and the thin, elongated structure make this the harder geometry.
Figure S4: Cross-video mask transfer, EchoNet-Dynamic. Left: the source video’s annotated frame. Right: its LV-cavity mask transferred onto three evenly spaced frames of a different video. Ground-truth contours appear on the target frames that carry them.
Figure S5: Cross-video mask transfer, MSK-Bone. As Figure S4 , for humerus masks transferred between subjects. The thin, non-convex structure is the case the KDE reconstruction finds hardest, which is why MSK-Bone Dice trails the cardiac datasets in Table 2 .
Figure S6: Cross-video mask transfer, CAMUS-TED. Left: the source subject’s annotated end-diastolic frame. Right: the LV-cavity mask transferred onto three evenly spaced frames of a different subject’s sequence, spanning its full cardiac cycle — cross-phase transfer made visible.
Figure S7: Cross-video point transfer, EchoNet-LVH. Left: the source video’s annotated frame with the four canonical landmarks. Right: those landmarks transferred onto three evenly spaced frames of a different video. EchoNet-LVH labels one frame per video, so ground truth appears in at most one target column; the others show propagation to unlabeled frames.
Method
EchoNet
MSK-POI
MSK-Bone
CoTracker3
41.8±10.6
54.7±7.4
40.9±7.4
TAPNext
41.6±11.3
27.3±7.5
29.8±12.2
Track-On2
34.9±9.6
47.0±8.3
32.4±9.4
Ours
31.7±11.5
41.8±8.8
38.2±7.0
Appendix
Table S4: Within-video point propagation , δavg↑ on EchoNet-Dynamic (100 videos), MSK-POI (36 videos) and MSK-Bone (20 videos). Baselines are the point trackers CoTracker3 ( Karaev et al. 2025 ) , TAPNext ( Zholus et al. 2025 ) and Track-On2 ( Aydemir et al. 2026 ) , none of which returns a mask. MSK-POI is the full 36-subject upper-arm set of ( Pallarès-López et al. 2026 ) , of which MSK-Bone is the 20-subject mask-annotated subset. Best in bold , second underlined .
Method
EchoNet
MSK-Bone
SAM 2
86.7±5.5
86.3±2.3
Ours
85.1±6.5
74.2±5.7
Appendix
Table S5: Within-video mask propagation , Dice ↑ on the two datasets with masks. The baseline is the video object segmenter SAM 2 ( Ravi et al. 2025 ) , which returns no point correspondence. Best in bold .
Figure S8: Within-video point propagation, MSK-Bone. Left: the video’s first labeled frame. Right: its points propagated to three evenly spaced frames of the same video. Ground truth is shown on the frames that carry it.
Figure S9: Within-video mask propagation, EchoNet-Dynamic. Left: the video’s first labeled frame. Right: its mask propagated to three evenly spaced frames of the same video.
Figure S10: Within-video mask propagation, MSK-Bone. As Figure S9 , within a single subject’s sequence.
Component
Time
VRAM (GB)
Feature extraction
63±2 ms
0.23
Feature SIREN
29.2±1.2 s
3.03
Flow SIREN
0.8±0.01 s
0.17
POI transfer
17±0 ms
< 0.1
Mask transfer
158±4 ms
< 0.1
KDE reconstruction
155±1 ms
< 0.1
Appendix
Table S6: Per-component timing and memory on a single RTX 4090, mean ± std over 3 warmed-up repetitions on EchoNet-Dynamic. Measured over a 16-frame clip with DINOv3-ViT-S/16 at 448 px input and a 112×112 reconstruction grid; the feature SIREN is trained for 500 epochs and each flow SIREN for 1000 epochs on one source–target pair. Transfer costs are quoted for 42 boundary points (POI) and 500 interior points (mask).
Method
Time ↓
GPU Mem
UniverSeg
0.02 s
0.07 GB
MATCHA
24.3 s
12.3 GB
Match4Annotate
59.4 s
∼ 3 GB
Appendix
Table S7: Runtime and memory for one cross-video transfer task on a single RTX 4090 (EchoNet-Dynamic, 5-run average). Match4Annotate trades feed-forward speed for cross-video transfer; the dominant cost (Feature SIREN training) is amortized across targets.
Ultrasound is the most widely used real-time imaging modality in clinical practice, yet per-frame video annotation remains a major bottleneck: expert labels are scarce and costly, and image appearance varies with speckle, shadowing, attenuation, and operator-dependent probe pose. This is especially limiting because clinically relevant information is often dynamic, from left-ventricular motion in echocardiography to muscle and bone kinematics in musculoskeletal imaging. Population atlases can amortize annotation cost by registering observations to a shared canonical coordinate system, but existing neural atlas methods mainly target single videos, small test-time image sets, or object-centric image collections. We introduce a cohort-scale neural atlas for ultrasound video: a single canonical chart with per-video Generative Latent Optimization embeddings, trained jointly over thousands of frames in DINOv3 feature space. Across five cardiac and musculoskeletal datasets with point landmarks and segmentation masks, our method learns coherent canonical templates and enables accurate atlas-space annotation transfer. On EchoNet-Dynamic and MSK-Bone, it supports single- and few-shot transfer with accuracy competitive with strong dense-correspondence baselines, while training in minutes on a single consumer GPU. The learned embeddings are interpretable: linear projections reveal structured cohort variation, image-decoder interpolation produces anatomically plausible intermediate frames, and test-time latent inversion reconstructs held-out frames through the atlas. These results suggest that cohort-scale neural atlases offer a practical, interpretable representation for reducing expert annotation burden in ultrasound video analysis.
Zhuorui Zhang, Roger Pallarès-López, Xuan Wu +2
Department of Mechanical Engineering, MIT · Institute for Medical Engineering and Science, MIT · 3MIT.nano Immersion Lab, MIT
Self-supervised pre-training paradigm has gained increasing prominence for learning transferable representations in medical imaging, yet existing methods for ultrasound (US) images operate at the image or frame level, overlooking the anatomical context for clinical-aligned representation learning. In this work, we propose an anatomy-anchored ultrasound self-supervision framework ANAUS that shifts representation learning from generic visual regions to clinically meaningful anatomical structures. Utilizing a learnable latent prompt engine alongside a one-time domain adaptation on existing public image-mask pairs, we empower the LP-SAM module to achieve annotation-free anatomy delineation at scale. Building upon this anatomical grounding, we propose a dual-policy self-supervised learning paradigm consisting of inter-view semantics-aware anatomy-separating alignment and contextual core-region prediction to enhance representation learning. Specifically, the former enforces feature invariance within identical anatomical regions while promoting discriminability across distinct structures; the latter compels the model to reconstruct corrupted regions, thereby capturing fine-grained structural details. Extensive evaluations on six public datasets demonstrate that ANAUS consistently outstrips current state-of-the-art methods while maintaining the computational efficiency essential for clinical deployment. Code is available at https://github.com/zhcz328/ANAUS.
Chunzheng Zhu, Yijun Wang, Jianxin Lin +5
Hunan University, Changsha, China · Shenzhen Maternity and Child Healthcare Hospital, Shenzhen, China
Ultrasound foundation models have achieved strong performance on structured prediction tasks but remain exclusively vision-based, limiting zero-shot and few-shot transfer to novel tasks where task-specific annotation is scarce. We address this gap with EchoCare-CLIP, a CLIP-style dual-encoder contrastive framework that aligns ultrasound images with clinical text in a shared embedding space. We curate a multi-organ corpus of over 16K image-text pairs spanning breast, liver, lung, and thyroid, with over 78% of captions derived from expert-annotated reports, and complement the remainder with a three-tier template-based and LLM-based caption generation pipeline. We evaluate model configurations spanning two text encoder families (CLIP, BioClinicalBERT) and two caption strategies (template-based, LLM-generated) against OpenAI CLIP and BiomedCLIP baselines. Our trained models consistently improve cross-modal alignment over baselines, with the best configuration achieving a paired alignment score of 0.682. However, stronger alignment does not guarantee better downstream performance: CLIP-based variants with partial fine-tuning achieve the strongest zero-shot classification on external held-out datasets (0.709 on BUSI; 0.626 on AULI), while full end-to-end fine-tuning degrades transfer due to overfitting. On linear probing and few-shot adaptation, model rankings are dataset-dependent, reflecting a trade-off between domain adaptation and representational generalizability. We further show that template-based captions match or outperform LLM-generated captions, suggesting lexical diversity is not a proxy for caption quality. Taken together, our results demonstrate that ultrasound vision-language alignment is achievable from public data alone, but robust clinical transfer requires careful balancing of domain adaptation, encoder capacity, and caption supervision quality.
Zhuoyang Lyu, Yiyang Zhang, Tongxin Wang +1
Department of Biostatistics Harvard T. H. Chan School of Public Health Boston, MA, USA