cs.CVOct 6, 2026

HuC-VideoMAE: Human-Centric Video Masked Autoencoding from synthetic data

Authors: Ricardo Pizarro, Roberto Valle, José M. Buenaposada, Luis M. Bergasa, Luis Baumela

Organizations: Universidad de Alcal´a, Alcal´a de Henares, Spain · Universidad Polit´ecnica de Madrid, Madrid, Spain · Universidad Rey Juan Carlos, M´ostoles, Spain

Abstract

Modern action recognition models rely on video transformers pretrained on massive collections of web-crawled videos, such as Kinetics-700. However, the use of such data raises ethical concerns, as subjects' consent is typically not obtained. Recent high-quality synthetic video datasets generated from motion-capture data, such as BEDLAM2.0, offer a promising ethical alternative. In this work, we investigate self-supervised pretraining of video transformers on synthetic human-motion datasets. We first show that directly applying the standard VideoMAE masking strategy leads to substantially worse performance than pretraining on Kinetics. To address this limitation, we propose a human-centric masking scheme that leverages body keypoints and person bounding box regions. Our approach encourages the model to focus on the structure and dynamics of human motion during pretraining. Experiments on NTU RGB+D and Toyota-Smarthome demonstrate that our method significantly outperforms standard VideoMAE pretraining on synthetic data, closing 49% of the gap to Kinetics pretraining on NTU RGB+D cross-view-subject without using a single real frame during pretraining. To promote the use of ethical action recognition models, we will publicly release our pretrained models.

Figures & tables

Explore similar work

CardsList
  1. Image Classifiers are Efficient Self-Supervised Video Representation Learners

    Sep 30, 2026Owais Iqbal, Sudipta Sarkar, Shyam Marjit +3Vision TransformerImage Classification

  2. Exploring the Role of Synthetic Data Augmentation in Controllable Human-Centric Video Generation

    Apr 23, 2026Yuanchen Fei, Yude Zou, Zejian Kang +3Human Motion GenerationVideo Dataset

  3. Identifying Ethical Biases in Action Recognition Models

    Apr 20, 2026Ana Baltaretu, Pascal Benschop, Jan van GemertHuman Activity RecognitionFace Recognition Model