HuC-VideoMAE: Human-Centric Video Masked Autoencoding from synthetic data
Authors: Ricardo Pizarro, Roberto Valle, José M. Buenaposada, Luis M. Bergasa, Luis Baumela
Organizations: Universidad de Alcal´a, Alcal´a de Henares, Spain · Universidad Polit´ecnica de Madrid, Madrid, Spain · Universidad Rey Juan Carlos, M´ostoles, Spain
Modern action recognition models rely on video transformers pretrained on massive collections of web-crawled videos, such as Kinetics-700. However, the use of such data raises ethical concerns, as subjects' consent is typically not obtained. Recent high-quality synthetic video datasets generated from motion-capture data, such as BEDLAM2.0, offer a promising ethical alternative. In this work, we investigate self-supervised pretraining of video transformers on synthetic human-motion datasets. We first show that directly applying the standard VideoMAE masking strategy leads to substantially worse performance than pretraining on Kinetics. To address this limitation, we propose a human-centric masking scheme that leverages body keypoints and person bounding box regions. Our approach encourages the model to focus on the structure and dynamics of human motion during pretraining. Experiments on NTU RGB+D and Toyota-Smarthome demonstrate that our method significantly outperforms standard VideoMAE pretraining on synthetic data, closing 49% of the gap to Kinetics pretraining on NTU RGB+D cross-view-subject without using a single real frame during pretraining. To promote the use of ethical action recognition models, we will publicly release our pretrained models.
Figures & tables
Figure 1: Body-keypoint-guided masking. From a synthetic frame with 2 D keypoint and person-box annotations (a), every patch is labelled as keypoint , body (inside the box) or background (b), and receives a masking score ρ (c): background scores above the body, and keypoints lowest of all, so that the sparse joint structure survives the mask. The exception are the joints selected for identity masking ( ∈S , red), with score above the background and are therefore always masked.
Figure 2: Keypoint masking is temporally consistent. A keypoint selected for masking (here the right wrist) is hidden in every frame of the clip (bottom row), because the choice is made at the body part identity level. A naive per-frame choice (top row) would leave the same joint visible in some frames ( t1,t4 ), letting the decoder copy it from a neighbouring frame. Green: visible keypoints; red: the masked joint and its patch.
Pretrain data
Masking
0∘
45∘
90∘
Avg
Kinetics
random tube
92.2
91.8
91.2
91.7
BEDLAM2.0
random tube
81.6
79.6
79.1
80.1
BEDLAM2.0
HuC-VideoMAE
86.6
85.9
84.8
85.8
Table 1: Ablation of the masking policy on NTU RGB+D, cross-view-subject (ViT-B). The backbone, optimizer, schedule and augmentation are identical across rows. Only the pretraining source and the masking scheme change. All numbers are downstream fine-tuning accuracy (in % ) after synthetic-to-real transfer (S → R), except the Kinetics row (R → R).
Pretrain data
NTU (Avg)
Toyota (Avg)
AMARV (synth.)
55.0
42.3
BEDLAM2.0 (synth.)
80.1
56.5
Table 2: Strength of the pretraining source, independent of masking policy. Random tube masking, same backbone and schedule. Only the synthetic pretraining corpus changes. Downstream accuracy averaged per dataset (NTU: 0∘/45∘/90∘ . Toyota: CV1/CV2).
Model
Pretrain Data
0∘
45∘
90∘
SURREACT [ 25 ]
R
-
86.9
74.5
53.6
X3D-S [ 19 ]
R
-
86.4
77.8
60.4
ViewCLR [ 6 ]
R
-
84.2
77.0
75.8
SURREACT [ 25 ]
S → R
SURREACT
84.1
77.5
66.2
X3D-S [ 19 ]
S → R
AMARV
89.9
81.8
68.0
MViTv2-S [ 19 ]
S → R
AMARV
94.4
84.5
65.1
Table 3: Cross-view-subject evaluation on NTU RGB+D, comparison with SURREACT [ 25 ] and our new VideoMAE-based results. S → R, Pretrained on Synth and then fine-tuned on Real.
Method
CV1
CV2
mCA. ( ↑ )
mCA. ( ↑ )
MotionFormer [ 16 ]
45.2
51.0
LTN [ 29 ]
-
54.6
TimeSFormer [ 2 ]
50.0
60.6
VPN++ [ 5 ]
-
54.9
Video Swing [ 15 ]
36.6
48.6
Table 4: Test results on Toyota-Smarthome over the CV1 and CV2 protocols. Comparison of our HuC models, against previous methods pretrained on Kinetics.
We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learning with unlabeled data, we repurpose standard image Vision Transformers by representing videos as super images which are grids composed of frames sampled from videos. From each super image, we construct two views: one with spatial patch masking and the other with temporal frame masking, ensuring no information leakage across frames. A shared Vision Transformer (ViT) encoder aligns their embeddings using a masked Siamese loss, capturing both motion and appearance cues without reconstruction. Our decoder-free formulation leverages an image foundation model towards efficient video representation learning. Starting from pretrained DINO-v3 and DeiT-v3 image encoders, VideoMSN achieves state-of-the-art performance on Kinetics-400, UCF101, and HMDB51 while requiring up to 32× fewer and 160× fewer video pretraining epochs compared to prior video self-supervised learning methods. Our proposed approach also shows strong performance in low-shot classification, confirming the transferability of the learned representations in a label-scarce scenario. Project Page: https://cvir.github.io/projects/videomsn.
Owais Iqbal, Sudipta Sarkar, Shyam Marjit +3
Indian Institute of Technology Kharagpur, India · Indian Institute of Science Bangalore, India · École de technologie supérieure Montreal, Canada
Controllable human video generation aims to produce realistic videos of humans with explicitly guided motions and appearances,serving as a foundation for digital humans, animation, and embodied AI.However, the scarcity of largescale, diverse, and privacy safe human video datasets poses a major bottleneck, especially for rare identities and complex actions.Synthetic data provides a scalable and controllable alternative,yet its actual contribution to generative modeling remains underexplored due to the persistent Sim2Real gap.In this work,we systematically investigate the impact of synthetic data on controllable human video generation. We propose a diffusion-based framework that enables fine-grained control over appearance and motion while providing a unfied testbed to analyze how synthetic data interacts with real world data during training. Through extensive experiments, we reveal the complementary roles of synthetic and real data and demonstrate possible methods for efficiently selecting synthetic samples to enhance motion realism,temporal consistency,and identity preservation.Our study offers the first comprehensive exploration of synthetic data's role in human-centric video synthesis and provides practical insights for building data-efficient and generalizable generative models.
Yuanchen Fei, Yude Zou, Zejian Kang +3
Hunan University · Shanghai Jiaotong University · Westlake University +3
Human Action Recognition (HAR) models are increasingly deployed in high-stakes environments, yet their fairness across different human appearances has not been analyzed. We introduce a framework for auditing bias in HAR models using synthetic video data, generated with full control over visual identity attributes such as skin color. Unlike prior work that focuses on static images or pose estimation, our approach preserves temporal consistency, allowing us to isolate and test how changes to a single attribute affect model predictions. Through controlled interventions using the BEDLAM simulation platform, we show whether some popular HAR models exhibit statistically significant biases on the skin color even when the motion remains identical. Our results highlight how models may encode unwanted visual associations, and we provide evidence of systematic errors across groups. This work contributes a framework for auditing HAR models and supports the development of more transparent, accountable systems in light of upcoming regulatory standards.
Ana Baltaretu, Pascal Benschop, Jan van Gemert
1EEMCS, Delft University of Technology, The Netherlands