cs.CVOct 4, 2026

EchoDino: A pediatric foundation model for transferable echocardiographic analysis across the lifespan

Authors: Sheng Cheng, Donnchadh M. O'Sullivan, Daniel J. Penny, Craig G. Rusin, Minh B. Nguyen, Devika Subramanian

Organizations: Department of Computer Science, Rice University, Houston, TX, USA · Division of Pediatric Cardiology, Baylor College of Medicine, Houston, TX, USA · Texas Children’s Hospital, Houston, TX, USA

Abstract

Echocardiography is the most widely used cardiac imaging modality, yet interpretation demands integrating visual evidence across global anatomy, localized structures and dynamic cardiac motion. Machine-learning models have automated individual tasks, but they are typically built for a single purpose and depend on expensively labeled datasets - a barrier particularly acute in pediatric care, where data are scarce and anatomy changes with age. Here we present EchoDino, a self-supervised foundation model for echocardiography, created by adapting the DINOv3 framework to 3.7 million frames from 1.7 million unlabeled pediatric echocardiography videos. With its encoder frozen, EchoDino produces representations that capture global context, local anatomy, and dense spatial detail. We introduce Motion-biased Entropy Maximization Sampling (MEMS) to select the most informative frames for video-level analysis. Across nine pediatric and adult datasets, EchoDino outperformed strong baseline models, raising view-classification accuracy from 0.609 to 0.889 and the area under the receiver operating characteristic curve for structural-heart-disease detection from 0.811 to 0.872, while also cutting age-estimation error from 3.857 to 1.389 years, achieving the best segmentation accuracy and lowering ejection-fraction errors. By generalizing from label-free pediatric data to adult echocardiography, EchoDino offers a versatile foundation for cardiac image analysis across the lifespan.

Figures & tables

Explore similar work

CardsList
  1. Beyond Independent Frames: Latent Attention Masked Autoencoders for Multi-View Echocardiography

    Apr 16, 2026Simon Böhi, Irene Cannistraci, Sergio Muñoz Gonzalez +8Medical Imaging Foundation ModelsMasked Autoencoders

  2. Evaluating self-supervised echocardiographic representations across downstream extraction strategies for left-ventricular segmentation and ejection fraction estimation

    Jun 22, 2026Sylwia Majchrowska, Philip TeareUltrasound Image SegmentationLeft Ventricular Ejection Fraction Estimation

  3. EchoXFlow: A Beamspace Echocardiography Dataset for Cardiac Motion, Flow, and Function

    May 6, 2026Elias Stenhede, Joanna Sulkowska, Eivind Bjørkan Orstad +2Medical ImagingEchocardiography