Structure-aware Keypoint Localization for Videofluoroscopic Swallowing Study
Authors: Kai Zhou, Chuanshen Chen, Runhao Zeng, Meng Dai, Yifan Yang, Jinwu Hu, Daiyuan Li, Mingkui Tan, +1 more
Organizations: South China University of Technology · Shenzhen MSU-BIT University · The Third Affiliated Hospital of Sun Yat-sen University · Electric Power Research Institute, China South Grid
Videofluoroscopic Swallowing Study (VFSS) is one of the gold standard for diagnosing swallowing disorders, providing dynamic X-ray imaging of the swallowing process. Automated kinematic analysis in VFSS relies fundamentally on precise anatomical keypoint localization. However, existing studies focus on limited keypoints (e.g., cervical vertebrae or the hyoid) and overlook critical regions such as the soft palate, while annotating only active swallowing segments and ignoring abundant non-swallowing data, resulting in poor data efficiency. Moreover, leveraging this unlabeled data via standard semi-supervised learning is suboptimal, as generic methods are prone to spatial bias. In medical X-rays with fixed layouts, models tend to memorize absolute coordinates rather than understanding anatomical structures. To tackle these challenges, we introduce VFSSKep, a novel dataset that extends annotations to the soft palate and incorporates large-scale unlabeled data. We further propose S3KL, a Structure-aware Semi-Supervised Keypoint Localization framework designed to overcome spatial bias. It integrates a Structure-Aware Learning strategy to extract high-resolution structural cues for structure-aware representation learning, and a Structural Representation Consistency Learning strategy with block shuffling to enforce invariant structural recognition. Experiments show our method achieves state-of-the-art semi-supervised performance, even with unlabeled and 25% labeled data surpassing fully supervised learning with 100% labeled data. Code and data will be made publicly available at: https://github.com/kaai520/S3KL.
Figures & tables
Fig. 1: Illustration of the eight keypoints in our VFSSKep dataset.
Dataset
Year
Anatomical Structure
Labeled Frames
Unlabeled Frames
Zhang et al. [ 6 ]
2021
Cervical Vertebrae
60k
-
Hsiao et al. [ 7 ]
2023
Hyoid + Vertebrae
7k
-
Ours (VFSSKep)
2025
Hyoid + Vertebrae + Soft Palate
14k
32k
TABLE I: Comparison of VFSS keypoint datasets. Our VFSSKep is the first to include soft palate annotations and leverage large-scale unlabeled data for semi-supervised learning.
Fig. 2: Comparison of the traditional (a) and our proposed (b) paradigms for utilizing unlabeled data in semi-supervised keypoint localization. Our paradigm leverages dense anatomical context to assist keypoint localization, learning structural priors without extra data or annotations.
Fig. 3: Summary of the construction of our VFSSKep dataset.
Fig. 4: An overview of our proposed S 3 KL framework. Our proposed S 3 KL framework applies easy augmentation Te to an unlabeled image for teacher predictions, transformed via easy-to-hard mapping Te→h as pseudo-labels for the student. Beyond keypoint supervision, we propose a Structure-Aware Learning (SAL) strategy to leverage dense anatomical cues, assisting keypoint localization by aligning the auxiliary prediction head output with structural prior from the Structure Pairing Module. The output from the Structure Parsing Module is visualized using PCA [ 14 ] . In addition, we introduce a Structural Representation Consistency (SRC) Learning strategy, where the Block Shuffle operation disrupts spatial structure, encouraging the model to maintain consistent representations despite such perturbations. The training process of labeled images is omitted for simplicity.
Fig. 5: (a) X-ray image segmented by SAM [ 16 ] with regions highlighted in different colors, but without semantic labels. (b) Output features from DINOv2 [ 13 ] for the same image, visualized by mapping PCA results to RGB channels. (c) Output features from our Structure Parsing Module (SPM), also visualized using PCA mapping.
TABLE II: Comparison with methods using different proportions of labeled data. [email protected] (%) and MED (mm) are reported. Our S 3 KL is plug-and-play, where variations in unsupervised loss Lu across different methods do not affect its integration. Best results are in bold , and blue numbers show relative improvements over SemiPose and G2LCPS.
Method
25%
50%
100%
PCK↑
MED↓
PCK↑
MED↓
PCK↑
MED↓
Baseline [ 10 ]
77.2
3.63
79.0
3.14
81.9
2.65
Baseline w/ SAL
78.1
3.22
81.6
2.69
82.4
2.58
Baseline w/ SRC
78.4
3.62
81.2
2.72
82.7
2.57
S 3 KL (Ours)
78.8
2.88
82.1
2.64
83.6
2.48
TABLE III: Ablation study on the proposed Structure-Aware Learning (SAL) and Structural Representation Consistency (SRC) Learning. [email protected] (%) and Mean Euclidean Distance (MED, in mm) are reported.
Model
Upsample
25%
50%
100%
PCK↑
MED↓
PCK↑
MED↓
PCK↑
MED↓
CLIP
FeatUp
77.8
3.73
80.5
2.86
81.9
2.74
DINOv2
FeatUp
78.1
3.22
81.6
2.69
82.4
2.58
DINOv2
Bilinear
77.8
2.97
80.9
2.79
81.9
2.68
TABLE IV: Effect of auxiliary models and upsampling strategies on the Structure Parsing Module. [email protected] (%) and Mean Euclidean Distance (MED, in mm) are reported.
Method
25%
50%
100%
PCK↑
MED↓
PCK↑
MED↓
PCK↑
MED↓
w/o BS
77.2
3.63
79.0
3.14
81.9
2.65
BS (final)
77.7
3.16
80.9
2.77
82.5
2.59
BS (encoder)
78.4
3.62
81.2
2.72
82.7
2.57
TABLE V: Effect of unshuffle position on the Block Shuffle (BS) operation. “final” indicates applying unshuffle to the final output, while “encoder” indicates applying unshuffle to the encoder output.
Fig. 6: Comparison of keypoint localization results between the baseline method [ 10 ] and our S 3 KL under 25% supervision. Green points represent the ground truth positions, while blue points represent the predicted keypoints.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Unlabeled ( n=62 )
Train ( n=28 )
Val ( n=9 )
Test ( n=28 )
Male
37.1%
39.3%
55.6%
39.3%
Female
62.9%
60.7%
44.4%
60.7%
Age 10–30
9.7%
7.1%
33.3%
7.1%
Age 30–50
25.8%
25.0%
11.1%
21.4%
Age 50–70
41.9%
42.9%
44.4%
46.4%
Age 70+
22.6%
25.0%
11.1%
25.0%
Appendix
TABLE I: Demographic distribution of VFSSKep subjects. Values are percentages; n : number of subjects.
Stage
Architecture
Output sizes C×H×W
input
512×8×8
dec 1
deconv 4×4,256 (stride=2, padding=1)
256×16×16
dec 2
deconv 4×4,256 (stride=2, padding=1)
256×32×32
dec 3
deconv 4×4,256 (stride=2, padding=1)
256×64×64
pred
conv 1×1,384 (stride=1, padding=0)
384×64×64
Appendix
TABLE II: An example implementation of the SAB prediction head using ResNet18 as the backbone. The deconvolutional layers (dec 1 , dec 2 , dec 3 ) progressively upsample the features, and a final 1×1 convolution (pred) produces the output.
λ3
25%
50%
100%
PCK↑
MED↓
PCK↑
MED↓
PCK↑
MED↓
10 -1
69.2
3.60
73.7
4.08
75.3
3.14
10 -2
78.1
3.56
80.0
2.90
81.1
2.75
10 -3
78.8
2.88
82.1
2.64
83.6
2.48
10 -4
78.4
3.62
80.6
3.32
82.4
2.56
Appendix
TABLE III: Effect of loss weight λ3 for feature alignment loss. [email protected] (%) and Mean Euclidean Distance (MED, in mm) are reported.
α
25%
50%
100%
PCK↑
MED↓
PCK↑
MED↓
PCK↑
MED↓
0.7
78.4
2.94
82.0
2.87
83.4
2.49
0.85
78.8
2.87
82.1
2.64
83.6
2.48
1.0
78.7
3.66
81.6
2.87
82.8
2.58
Appendix
TABLE IV: Effect of the similarity threshold α in the cosine similarity loss Lcos . [email protected] (%) and Mean Euclidean Distance (MED, in mm) are reported.
Size
25%
50%
100%
PCK↑
MED↓
PCK↑
MED↓
PCK↑
MED↓
32
78.5
2.95
81.8
2.75
82.6
2.68
64
78.8
2.87
82.1
2.64
83.6
2.48
128
78.4
3.65
81.4
3.13
82.3
2.60
Appendix
TABLE V: Effect of input block size on the Block Shuffle operation. [email protected] (%) and Mean Euclidean Distance (MED, in mm) are reported.
Backbone
Param
FLOPs
25%
50%
100%
PCK↑
MED↓
PCK↑
MED↓
PCK↑
MED↓
ResNet18
15.6M
2.4G
78.8
2.88
82.1
2.64
83.6
2.48
ViT-B
87.6M
21.6G
79.9
2.74
82.3
2.55
83.5
2.51
Appendix
TABLE VI: Effect of different backbones.
Fig. 1 : Partial screenshot of the clinical prototype system.
Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Southeast University, Ministry of Education, Jiangsu, China · Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, UAE · School of Medicine, Case Western Reserve University, Cleveland, OH, USA
School of Computer Science and Technology, Tongji University, Shanghai, China · Artificial Intelligence Institute, Shanghai University, Shanghai, China · Department of Computer and Data Science and Department of Biomedical2026 Engineering, Case Western Reserve University, Cleveland, USA