cs.CVSep 28, 2026

BiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion

Authors: Wanjiang Weng, Yongliang Wu, Xiaofeng Tan, Xingyu Zhu, Wenbo Zhu, Hongsong Wang

Organizations: Department of Computer Science and Engineering, Southeast University, Nanjing, China · Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, Nanjing, China · National University of Singapore · Opus AI

Abstract

Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poorly suited to the bidirectional dependencies between language and motion, allowing early prediction errors to persist as fixed context and degrade both temporal coherence and cross-modal consistency. Masked discrete diffusion, which models sequences through iterative bidirectional prediction, offers a natural remedy. We therefore propose BiMoGen (Bidirectional Motion-text Generation), a unified masked discrete diffusion framework for bidirectional motion-text modeling. To stabilize training, we design Decoupled Uni- and Cross-Modal Training, in which masked pretraining first establishes cross-modal correspondence on paired motion-text sequences, after which supervised fine-tuning specializes the model for bidirectional generation. Masked diffusion nonetheless introduces its own source of error, as the model is trained on clean ground-truth context yet encounters self-generated and potentially erroneous context at inference, with errors committed under heavily masked states propagating through subsequent steps. We further introduce Generation-Aware Self-Correction that exposes the model to its own predictions during training and applies correction passes at early sampling steps to revise unreliably committed tokens. Extensive experiments on HumanML3D and KIT-ML demonstrate competitive performance on both tasks, validating the effectiveness of the proposed two-stage training and self-correction designs. The project page is available at https://wengwanjiang.github.io/BiMoGen-Page.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation

    Dec 15, 2025Yannan He, Garvita Tiwari, Xiaohan Zhang +4Text-To-Motion GenerationMotion-Language Alignment

  2. Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation

    Jan 21, 2026Yifei Liu, Changxing Ding, Ling Guo +2Text-To-Motion GenerationText-To-Video Diffusion Models

  3. ReMoMask-2: Latent Retrieval-Augmented Masked Motion Generation

    Sep 8, 2026Yiran Wang, Zeyu Zhang, Ling Shao +1Text-To-Motion GenerationNeural Mask Estimation