cs.CVSep 28, 2026

MotionSpaceFlow: Representation-Aware Flow Matching in Direct Motion Space

Authors: Qing Yu, Kent Fujiwara

Organizations: LY Corporation, Tokyo, Japan

Abstract

Recent advances in diffusion and flow models have substantially improved text-driven human motion generation. Yet most methods generate in low-dimensional, temporally downsampled latent spaces learned primarily for reconstruction, a bottleneck that can limit generation quality and preclude direct manipulation of individual frames and joints. We introduce MotionSpaceFlow (MSFlow), a representation-aware flow-matching framework that predicts clean motion directly in continuous motion space without a learned encoder or decoder. To account for the anisotropic structure of direct motion representations, we propose representation-aware noise scaling and show how the initial Gaussian source scale governs the covariance of intermediate probability-path marginals. We further introduce a Representation-Aware Multimodal Diffusion Transformer (RA-MMDiT), which jointly updates token-level language and full-resolution motion features through joint attention while adapting temporal information flow to the motion representation: causal attention for incremental features defined by frame-to-frame changes, and bidirectional attention for global features such as absolute joint coordinates. Across different datasets and motion representations, MSFlow achieves state-of-the-art text-to-motion performance. Its global representation variant additionally enables zero-shot, inference-time control over any joint or frame through projection sampling without control-conditioned training, delivering leading motion quality with exact constraint satisfaction.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation

    Jul 31, 2026Yifei Zhu, Mingyi Shi, Yangyang Cai +3Text-To-Motion GenerationLatent Flow

  2. MotionHiFlow: Text-to-motion via hierarchical flow matching

    Apr 25, 2026Heng Li, Xiaotong Lin, Ling-An Zeng +3Text-To-Motion GenerationContrastive Flow Matching

  3. MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation

    Dec 15, 2025Yannan He, Garvita Tiwari, Xiaohan Zhang +4Text-To-Motion GenerationMotion-Language Alignment