DiMoP: Diffusion-Driven Motion Representation Learning With Frame-Level Pseudo-Classification for Skeleton-Based Action Recognition
Organizations: Advanced Multimedia Research Lab, University of Wollongong, Australia
Abstract
Robust skeleton-based action recognition requires representations that capture a wide spectrum of motions, from subtle to moderate and strong ones. Existing methods often focus on strong motions. This paper introduces DiMoP, a masking- and diffusion-driven motion representation learning method with frame-level pseudo-classification to explicitly learn the distribution of joint motions rather than regressing deterministic coordinates, as existing methods often do. By diffusing masked joints with progressive noise and denoising them conditioned on visible joints, DiMoP learns through controllable noising and denoising processes, enabling uniform learning of weak, moderate, and strong dynamics. To enable the masking-based generative diffusion learning with a discriminative capability, a pseudo-frame classifier is proposed that enforces the learning towards sequence-consistent and temporally coherent pseudo-labels without manual annotations. Together, these strategies provide a principled mechanism for joint generative and discriminative motion modeling. DiMoP achieves state-of-the-art performance across NTU RGB+D 60/120, and PKUMMD, including a 1.1 percentage point gain over prior works on NTU RGB+D 120 with the cross-subject protocol.
Figures & tables
| Models | Venue | Stream | NTU 60 | NTU 120 | PKU MMD | ||
| X-sub | X-view | X-sub | X-set | Part I | |||
| Contrastive Methods | |||||||
| ISC [ 46 ] | MM’21 | J | 76.3 | 78.6 | 67.9 | 67.1 | 80.9 |
| GL-Transformer [ 47 ] | ECCV’22 | J | 76.3 | 83.8 | 66.0 | 68.7 | - |
| PSTL [ 48 ] | AAAI’23 | J | 77.3 | 81.8 | 69.2 | 67.7 | 88.4 |
| CPM [ 49 ] | ECCV’22 | J | 78.7 | 84.9 | 68.7 | 69.6 | - |
| Models | NTU 60 (%) | NTU 120 (%) | ||
| X-sub | X-view | X-sub | X-set | |
| Contrastive Methods | ||||
| 3s-CrosSCLR | 86.2 | 92.5 | 80.5 | 80.4 |
| 3s-AimCLR | 86.9 | 92.8 | 80.1 | 80.9 |
| 3s-PSTL | 87.1 | 93.9 | 81.3 | 82.6 |
| 3s-ActCLR | 88.2 | 93.9 | 81.6 | 81.2 |
| Models | To PKUMMD II (%) | |
| NTU 60 | NTU 120 | |
| SkeletonMAE | 58.4 | 61.0 |
| MAMP | 70.6 | 73.2 |
| MacDiff | 72.2 | 73.4 |
| S-JEPA | 71.4 | 74.2 |
| DiMoP (ours) | 72.4 | 73.9 |
| Models | NTU RGB+D 60 (%) | |||
| 1% labels | 10% labels | |||
| X-sub | X-view | X-sub | X-view | |
| 3s-AimtCLR | 54.8 | 54.3 | 78.2 | 81.6 |
| 3s-ActCLR | 64.8 | 65.6 | 81.7 | 85.8 |
| 3s-STJD-CL | 67.5 | 68.4 | 82.5 | 88.0 |
| SkeletonMAE | 54.4 | 54.6 | 80.6 | 83.5 |
| Masking strategy | Acc(%) |
| Random Masking | 86.48 |
| Motion Aware masking | 86.13 |
| Masking ratio | Acc(%) |
| 60% | 70.21 |
| 65% | 74.44 |
| 70% | 75.65 |
| 75% | 78.38 |
| 80% | 82.59 |
| 85% | 85.64 |
| Acc % | |||
| 1 | 1 | 1 | 85.85 |
| 1 | 0.1 | 0.1 | 85.34 |
| 1 | 1 | 0.1 | 85.81 |
| 1 | 0.1 | 1 | 86.24 |
| 1 | 0.3 | 1 | 86.48 |
| 1 | 0.5 | 0.5 | 86.02 |
| Acc (%) | |||
| ✓ | 85.31 | ||
| ✓ | ✓ | 34.73 | |
| ✓ | ✓ | 85.81 | |
| ✓ | ✓ | 86.09 | |
| ✓ | ✓ | ✓ | 86.48 |
| Noise Schedule | Acc(%) |
| Linear | 86.48 |
| Cosine | 85.61 |
| Prediction target | Acc(%) |
| Noise | 49.03 |
| Coordinates of the original | 85.94 |
| Motion dynamics of the original | 86.48 |
| Configuration | Acc(%) |
| Joint decoder | 86.48 |
| Cross decoder V1 | 85.73 |
| Cross decoder V2 | 86.76 |
| Group | Imp./Drop | % imp | Mean Acc. |
| G1: Fine-grained single-person | 17/6 | 74% | +1.88 |
| G2: Large body-motion | 11/1 | 92% | +2.19 |
| G3: Hand-object interaction | 9/5 | 64% | +2.59 |
| G4: Two-person interaction | 10/1 | 91% | +3.57 |
| G5: Brief-discriminative-frame † | 10/6 | 63% | +2.16 |
| Overall | 47/13 | 78% | +1.90 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Models | To UCLA (%) | |
| NTU 60 | NTU 120 | |
| TAHAR [ 53 ] | 93.1 | 95.2 |
| DiMoP (ours) | 94.2 | 96.7 |
| Model | Acc(%) |
| MacDiff (reproduced) | 86.03 |
| DiMoP (ours) | 86.76 |
| DiMoP + MacDiff decoder (Noise) | 85.83 |
| DiMoP + MacDiff decoder (Motion) | 86.61 |
| Pre-training | Linear Evaluation | |||
| Model | Params | FLOPs (G) | Params | Time (s/iter) |
| MAMP [ 17 ] (baseline) | 8.79M | 5.45 | 0.27M | 0.342 |
| DiMoP | 9.69M | 6.78 | 0.38M | 0.413 |
| Protocol | -value | Decision | |||
| X-sub | 84.95 | 86.79 | 5.636 | 2.776 | Reject |
| X-view | 89.14 | 91.91 | 8.154 | 2.776 | Reject |
| Model | NTU 60 | NTU 120 | PKU | ||
| X-sub | X-view | X-sub | X-set | Part I | |
| MAMP | 84.90 | 89.10 | 78.60 | 79.10 | 92.20 |
| MAMP ∗ | 84.84 | 89.13 | 78.49 | 78.98 | 92.24 |
| MacDiff | 86.40 | 91.00 | 79.40 | 80.20 | 92.80 |
| MacDiff ∗ | 86.43 | 90.95 | 79.34 | 80.13 | 92.74 |
| Group | Dominant motion cue | Average % Gain |
| Subtle motion | Localized low-amplitude hand, arm, head, and object-related motion | +5.77 |
| Strong motion | Global displacement, posture transition, and interaction dynamics | +4.31 |
| Group | Action IDs | Improved/Dropped actions (% Improved) | Mean Acc. | Main Insight |
| G1: Fine-grained single-person actions | A3, A4, A18–A21, A28, A31, A33–A41, A44–A49 | 17/6(74%) | +1.88 | Localized hand, arm, head, and object-related motions were better captured, while drops mainly occurred when cues were confined to similar body regions. |
| G2: Large body-motion actions | A5–A9, A22–A24, A26, A27, A42, A43 | 11/1 (92%) | +2.19 | Global displacement, posture transitions, and strong temporal dynamics were consistently preserved. |
| G3: Hand-object interaction actions | A1, A2, A10–A17, A25, A29, A30, A32 | 9/5(64%) | +2.59 | Local-to-mid-level hand motion was improved, although skeleton-only input remained limited when object/contact cues were essential. |
| G4: Two-person interaction actions | A50–A60 | 10/1 (91%) | +3.57 | Interaction-related motion cues were effectively captured, although explicit inter-subject relation modeling may further improve performance. |
| G5: Brief-discriminative-frame actions † | A3, A18–A21, A25, A28, A33, A37, A41, A44–A48, A57 | 10/6(63%) | +2.16 | Actions with short-lived discriminative cues remained challenging when most frames contained neutral or visually similar poses. |
| Overall | A1–A60 | 47/13 (78%) | +1.90 | Most classes were improved, supporting the effectiveness of diffusion-based motion distribution learning. |
| Acc. (%) | |||
| 1 | 0.3 | 0 | 85.33 |
| 1 | 1.0 | 0 | 84.92 |
| 1 | 3.0 | 0 | 84.59 |
| 1 | 5.0 | 0 | 83.46 |
| 1 | 0.3 | 1.0 | 86.48 |