cs.CVMar 6, 2026

Match4Annotate: Cross-Video Annotation Transfer in Ultrasound via Implicit Feature Flow-Guided Matching

Authors: Zhuorui Zhang, Roger Pallarès-López, Praneeth Namburi, Brian W. Anthony

Organizations: Department of Mechanical Engineering, Massachusetts Institute of Technology, Cambridge, USA · Institute for Medical Engineering and Science, Massachusetts Institute of Technology, Cambridge, USA · MIT.nano Immersion Lab, Massachusetts Institute of Technology, Cambridge, USA

Abstract

Acquiring per-frame annotations for ultrasound videos is costly and requires clinical expertise, limiting learning-based analysis. We study cross-video annotation transfer: propagating user-specified annotations from a labeled ultrasound video to an independently acquired target video with no target-side labels or manual initialization. Video trackers and segmentation propagators rely on temporal continuity and require a prompt in every new sequence, whereas cross-image feature matching and one-shot segmentation estimate correspondences independently, without enforcing coherent deformations or supporting both point and mask annotations. We present Match4Annotate, a test-time framework with three stages. A spatiotemporal implicit feature representation lifts frozen vision foundation-model features into a continuous field over space and time, enabling queries beyond the backbone resolution. A continuous implicit feature flow then aligns the source and target fields under a smooth-deformation prior, estimating correspondence in feature space rather than relying on intensity consistency, which is often violated in ultrasound by speckle and acquisition-dependent appearance. Finally, flow-guided annotation transfer uses the estimated flow as a spatial prior over feature similarity. This formulation unifies sparse point and dense mask transfer and includes unconstrained feature matching and direct flow warping as limiting cases. On four clinical ultrasound datasets spanning echocardiography and musculoskeletal imaging, Match4Annotate achieves state-of-the-art annotation transfer, outperforming dense feature-matching baselines across PCK thresholds and one-shot segmentation methods in Dice score. It also demonstrates bidirectional transfer of left-ventricular annotations across datasets. It requires no task-specific training and adapts to each video in minutes on a single consumer GPU.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Cohort-Scale Neural Atlases of Ultrasound Video

    May 30, 2026Zhuorui Zhang, Roger Pallarès-López, Xuan Wu +2Clinical Ultrasound DatasetAnatomy

  2. Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation

    May 25, 2026Chunzheng Zhu, Yijun Wang, Jianxin Lin +5UltrasoundSelf-Supervised Learning

  3. Ultrasound Vision-Language Alignment via Contrastive Learning

    May 4, 2026Zhuoyang Lyu, Yiyang Zhang, Tongxin Wang +1UltrasoundVision-Language Alignment