cs.CVSep 28, 2026

DiMoP: Diffusion-Driven Motion Representation Learning With Frame-Level Pseudo-Classification for Skeleton-Based Action Recognition

Authors: Shanaka Ramesh Gunasekara, Wanqing Li, Nikalal Kaldera, Philip Ogunbona, Jack Yang

Organizations: Advanced Multimedia Research Lab, University of Wollongong, Australia

Abstract

Robust skeleton-based action recognition requires representations that capture a wide spectrum of motions, from subtle to moderate and strong ones. Existing methods often focus on strong motions. This paper introduces DiMoP, a masking- and diffusion-driven motion representation learning method with frame-level pseudo-classification to explicitly learn the distribution of joint motions rather than regressing deterministic coordinates, as existing methods often do. By diffusing masked joints with progressive noise and denoising them conditioned on visible joints, DiMoP learns through controllable noising and denoising processes, enabling uniform learning of weak, moderate, and strong dynamics. To enable the masking-based generative diffusion learning with a discriminative capability, a pseudo-frame classifier is proposed that enforces the learning towards sequence-consistent and temporally coherent pseudo-labels without manual annotations. Together, these strategies provide a principled mechanism for joint generative and discriminative motion modeling. DiMoP achieves state-of-the-art performance across NTU RGB+D 60/120, and PKUMMD, including a 1.1 percentage point gain over prior works on NTU RGB+D 120 with the cross-subject protocol.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action Recognition

    Jun 9, 2026Shengkai Sun, Zhiyong Cheng, Zefan Zhang +3Skeleton-Based Action Recognition3D Masked Autoencoders

  2. Visual Anchoring in Diffusion: Multimodal Zero-Shot Skeleton Action Recognition

    Aug 5, 2026Zehao Bao, Shujun Guo, Bruce X. B. YuSkeleton-Based Action RecognitionZero-Shot

  3. GenPrior: Unleashing Text-to-Motion Generative Priors for Zero-Shot Skeleton-based Action Recognition

    Aug 3, 2026Jidong Kuang, Hongsong Wang, Jie GuiSkeleton-Based Action RecognitionText-To-Motion Generation