Structure-aware Keypoint Localization for Videofluoroscopic Swallowing Study
Organizations: South China University of Technology · Shenzhen MSU-BIT University · The Third Affiliated Hospital of Sun Yat-sen University · Electric Power Research Institute, China South Grid
Abstract
Videofluoroscopic Swallowing Study (VFSS) is one of the gold standard for diagnosing swallowing disorders, providing dynamic X-ray imaging of the swallowing process. Automated kinematic analysis in VFSS relies fundamentally on precise anatomical keypoint localization. However, existing studies focus on limited keypoints (e.g., cervical vertebrae or the hyoid) and overlook critical regions such as the soft palate, while annotating only active swallowing segments and ignoring abundant non-swallowing data, resulting in poor data efficiency. Moreover, leveraging this unlabeled data via standard semi-supervised learning is suboptimal, as generic methods are prone to spatial bias. In medical X-rays with fixed layouts, models tend to memorize absolute coordinates rather than understanding anatomical structures. To tackle these challenges, we introduce VFSSKep, a novel dataset that extends annotations to the soft palate and incorporates large-scale unlabeled data. We further propose SKL, a Structure-aware Semi-Supervised Keypoint Localization framework designed to overcome spatial bias. It integrates a Structure-Aware Learning strategy to extract high-resolution structural cues for structure-aware representation learning, and a Structural Representation Consistency Learning strategy with block shuffling to enforce invariant structural recognition. Experiments show our method achieves state-of-the-art semi-supervised performance, even with unlabeled and 25% labeled data surpassing fully supervised learning with 100% labeled data. Code and data will be made publicly available at: https://github.com/kaai520/S3KL.
Figures & tables
| Dataset | Year | Anatomical Structure | Labeled Frames | Unlabeled Frames |
| Zhang et al. [ 6 ] | 2021 | Cervical Vertebrae | 60k | - |
| Hsiao et al. [ 7 ] | 2023 | Hyoid + Vertebrae | 7k | - |
| Ours (VFSSKep) | 2025 | Hyoid + Vertebrae + Soft Palate | 14k | 32k |
| Method | 25% | 50% | 100% | |||
| PCK@0.02 (%)↑ | MED (mm)↓ | PCK@0.02 (%)↑ | MED (mm)↓ | PCK@0.02 (%)↑ | MED (mm)↓ | |
| Supervised [ 20 ] | 63.2 | 4.34 | 73.5 | 3.47 | 75.6 | 3.29 |
| SemiPose [ 10 ] | 77.2 | 3.63 | 79.0 | 3.14 | 81.9 | 2.65 |
| +S 3 KL (Ours) | 78.8 ( ↑2.1% ) | 2.88 ( ↓20.7% ) | 82.1 ( ↑3.9% ) | 2.64 ( ↓15.9% ) | 83.6 ( ↑2.1% ) | 2.48 ( ↓6.4% ) |
| G2LCPS [ 12 ] | 78.2 | 2.98 | 81.0 | 2.76 | 83.1 | 2.54 |
| +S 3 KL (Ours) | 79.4 ( ↑1.5% ) | 2.87 ( ↓3.7% ) | 82.9 ( ↑2.3% ) | 2.57 ( ↓6.9% ) | 83.8 ( ↑0.8% ) | 2.46 ( ↓3.1% ) |
| Method | 25% | 50% | 100% | |||
| PCK↑ | MED↓ | PCK↑ | MED↓ | PCK↑ | MED↓ | |
| Baseline [ 10 ] | 77.2 | 3.63 | 79.0 | 3.14 | 81.9 | 2.65 |
| Baseline w/ SAL | 78.1 | 3.22 | 81.6 | 2.69 | 82.4 | 2.58 |
| Baseline w/ SRC | 78.4 | 3.62 | 81.2 | 2.72 | 82.7 | 2.57 |
| S 3 KL (Ours) | 78.8 | 2.88 | 82.1 | 2.64 | 83.6 | 2.48 |
| Model | Upsample | 25% | 50% | 100% | |||
| PCK↑ | MED↓ | PCK↑ | MED↓ | PCK↑ | MED↓ | ||
| CLIP | FeatUp | 77.8 | 3.73 | 80.5 | 2.86 | 81.9 | 2.74 |
| DINOv2 | FeatUp | 78.1 | 3.22 | 81.6 | 2.69 | 82.4 | 2.58 |
| DINOv2 | Bilinear | 77.8 | 2.97 | 80.9 | 2.79 | 81.9 | 2.68 |
| Method | 25% | 50% | 100% | |||
| PCK↑ | MED↓ | PCK↑ | MED↓ | PCK↑ | MED↓ | |
| w/o BS | 77.2 | 3.63 | 79.0 | 3.14 | 81.9 | 2.65 |
| BS (final) | 77.7 | 3.16 | 80.9 | 2.77 | 82.5 | 2.59 |
| BS (encoder) | 78.4 | 3.62 | 81.2 | 2.72 | 82.7 | 2.57 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Unlabeled ( ) | Train ( ) | Val ( ) | Test ( ) | |
|---|---|---|---|---|
| Male | 37.1% | 39.3% | 55.6% | 39.3% |
| Female | 62.9% | 60.7% | 44.4% | 60.7% |
| Age 10–30 | 9.7% | 7.1% | 33.3% | 7.1% |
| Age 30–50 | 25.8% | 25.0% | 11.1% | 21.4% |
| Age 50–70 | 41.9% | 42.9% | 44.4% | 46.4% |
| Age 70+ | 22.6% | 25.0% | 11.1% | 25.0% |
| Stage | Architecture | Output sizes |
|---|---|---|
| input | ||
| dec 1 | deconv (stride=2, padding=1) | |
| dec 2 | deconv (stride=2, padding=1) | |
| dec 3 | deconv (stride=2, padding=1) | |
| pred | conv (stride=1, padding=0) |
| 25% | 50% | 100% | ||||
|---|---|---|---|---|---|---|
| PCK↑ | MED↓ | PCK↑ | MED↓ | PCK↑ | MED↓ | |
| 10 -1 | 69.2 | 3.60 | 73.7 | 4.08 | 75.3 | 3.14 |
| 10 -2 | 78.1 | 3.56 | 80.0 | 2.90 | 81.1 | 2.75 |
| 10 -3 | 78.8 | 2.88 | 82.1 | 2.64 | 83.6 | 2.48 |
| 10 -4 | 78.4 | 3.62 | 80.6 | 3.32 | 82.4 | 2.56 |
| 25% | 50% | 100% | ||||
|---|---|---|---|---|---|---|
| PCK↑ | MED↓ | PCK↑ | MED↓ | PCK↑ | MED↓ | |
| 0.7 | 78.4 | 2.94 | 82.0 | 2.87 | 83.4 | 2.49 |
| 0.85 | 78.8 | 2.87 | 82.1 | 2.64 | 83.6 | 2.48 |
| 1.0 | 78.7 | 3.66 | 81.6 | 2.87 | 82.8 | 2.58 |
| Size | 25% | 50% | 100% | |||
|---|---|---|---|---|---|---|
| PCK↑ | MED↓ | PCK↑ | MED↓ | PCK↑ | MED↓ | |
| 32 | 78.5 | 2.95 | 81.8 | 2.75 | 82.6 | 2.68 |
| 64 | 78.8 | 2.87 | 82.1 | 2.64 | 83.6 | 2.48 |
| 128 | 78.4 | 3.65 | 81.4 | 3.13 | 82.3 | 2.60 |
| Backbone | Param | FLOPs | 25% | 50% | 100% | |||
|---|---|---|---|---|---|---|---|---|
| PCK↑ | MED↓ | PCK↑ | MED↓ | PCK↑ | MED↓ | |||
| ResNet18 | 15.6M | 2.4G | 78.8 | 2.88 | 82.1 | 2.64 | 83.6 | 2.48 |
| ViT-B | 87.6M | 21.6G | 79.9 | 2.74 | 82.3 | 2.55 | 83.5 | 2.51 |