cs.CVOct 6, 2026

Structure-aware Keypoint Localization for Videofluoroscopic Swallowing Study

Authors: Kai Zhou, Chuanshen Chen, Runhao Zeng, Meng Dai, Yifan Yang, Jinwu Hu, Daiyuan Li, Mingkui Tan, +1 more

Organizations: South China University of Technology · Shenzhen MSU-BIT University · The Third Affiliated Hospital of Sun Yat-sen University · Electric Power Research Institute, China South Grid

Abstract

Videofluoroscopic Swallowing Study (VFSS) is one of the gold standard for diagnosing swallowing disorders, providing dynamic X-ray imaging of the swallowing process. Automated kinematic analysis in VFSS relies fundamentally on precise anatomical keypoint localization. However, existing studies focus on limited keypoints (e.g., cervical vertebrae or the hyoid) and overlook critical regions such as the soft palate, while annotating only active swallowing segments and ignoring abundant non-swallowing data, resulting in poor data efficiency. Moreover, leveraging this unlabeled data via standard semi-supervised learning is suboptimal, as generic methods are prone to spatial bias. In medical X-rays with fixed layouts, models tend to memorize absolute coordinates rather than understanding anatomical structures. To tackle these challenges, we introduce VFSSKep, a novel dataset that extends annotations to the soft palate and incorporates large-scale unlabeled data. We further propose S3^3KL, a Structure-aware Semi-Supervised Keypoint Localization framework designed to overcome spatial bias. It integrates a Structure-Aware Learning strategy to extract high-resolution structural cues for structure-aware representation learning, and a Structural Representation Consistency Learning strategy with block shuffling to enforce invariant structural recognition. Experiments show our method achieves state-of-the-art semi-supervised performance, even with unlabeled and 25% labeled data surpassing fully supervised learning with 100% labeled data. Code and data will be made publicly available at: https://github.com/kaai520/S3KL.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Solving Semi-Supervised Few-Shot Learning from an Auto-Annotation Perspective

    Dec 11, 2025Tian Liu, Anwesha Basu, James Caverlee +1VLM AdaptationFew-Shot Learning

  2. ARTEMIS: Agent-guided Reliability-aware Temporal Mask Evolution for Imperfectly Supervised Video Polyp Segmentation

    Jun 18, 2026Tong Wang, Siwen Wang, Yaolei Qi +4Video Object SegmentationSemi-Supervised Medical Image Segmentation

  3. SCKAN: Structural Consensus-based KAN Prototype Learning for Semi-Supervised Pancreas Segmentation

    May 26, 2026Yuqi Liu, Yufei Chen, Wei Fu +2Prototype LearningMedical Image Segmentation