When Does Exercise-Specific Joint Selection Help? An Audit of Evaluation and Control Design
Authors: Haotian Chen, Jingkun Yu, Yuning Zhang, Bowen Ye
Organizations: School of Cyber Science and Technology, University of Science and Technology of China · SWJTU-Leeds Joint School, Southwest Jiaotong University · School of Education, Shanghai Jiao Tong University
Exercise-specific joint selection can improve skeleton-based correctness classification, but what does that gain establish? We audit 1,057 repetitions from ten REHAB24-6 subjects, separating evaluation aggregation, subset structure, and temporal representation. The manual-subset kNN gain changes from 0.055 for pooled out-of-fold AUROC to 0.020 for equal-weight within-person AUROC; both paired intervals include zero. Among 1,000 dimension-matched random maps, 14 match or exceed the manual pooled result, versus 145 when bilateral structure and trunk inclusion are also matched. RBF-SVM retains a positive within-person gain, whereas logistic regression and a random-convolution comparator have negative point gains under that estimand. Sequence-order and paired-seed controls further qualify the interpretation. This exploratory audit shows why joint-selection claims require explicit estimands and structurally appropriate controls; it does not establish a new algorithm or clinical benefit.
Figures & tables
Figure 1: Subject-disjoint representation audit. Dots in (c) mark evaluated combinations; “Ref.” includes dimension-matched, structure-matched, and training-selected subsets. In (d), marker shapes identify people and filled/open markers indicate correct/incorrect repetitions; ranks are schematic. Both AUROC estimands average exercises equally. Paired subject bootstrap resamples fixed OOF predictions, not model fits.
Pooled out-of-fold AUROC
Within-person AUROC
Model
All / manual
Gain
Paired 95% CI
All / manual
Gain
Paired 95% CI
kNN ( k=3 )
.596 / .651
.055
[-.005,.113]
.657 / .677
.020
[-.068,.102]
LR
.664 / .691
.027
[-.058,.097]
.725 / .718
-.007
[-.061,.043]
RBF-SVM
.627 / .684
.057
[.005,.121]
.689 / .734
.045
[.013,.080]
RF (historical seed)
.597 / .641
.044
[-.003,.087]
.677 / .688
.011
[-.029,.046]
Random convolution
.710 / .728
.017
[-.018,.053]
.764 / .743
-.021
[-.049,.012]
Table 1: All/manual AUROC and paired gains. Pooled uses all 1,057 repetitions; within-person uses 51 evaluable cells (1,026 repetitions). The original RF row is retained; paired-seed RF is a separate sensitivity analysis. All intervals use 4,999 valid draws from 5,000 fixed-prediction subject resamples.
Gain estimand
kNN
LR
SVM
RF
Pooled: all rows
.055
.027
.057
.044
Pooled: eligible rows
.062
.005
.038
.039
Within: pair weighted
.021
-.018
.053
.013
Within: equal people
.020
-.007
.045
.011
Table 2: Decomposing manual-minus-all gains. The final three rows use the same eligible observations; only comparisons and weights differ.
Figure 2: Random-map references for kNN. Curves show the percentage of 1,000 maps at or above each AUROC threshold (six-exercise macro average). Dotted lines mark manual; annotations are descriptive percentages, not p -values.
Representation
kNN
LR
Ordered (40 frames)
.657 / .677
.725 / .718
Sorted per angle (40)
.621 / .657
.711 / .760
Shuffled frames (40)
.592 / .626
.613 / .583
Mean / SD / min / max
.629 / .671
.703 / .728
Table 3: Within-person AUROC (all/manual) for sequence controls. Sorted and shuffled inputs match ordered dimensionality; summaries use four statistics per angle.