MotionSpaceFlow: Representation-Aware Flow Matching in Direct Motion Space
Organizations: LY Corporation, Tokyo, Japan
Abstract
Recent advances in diffusion and flow models have substantially improved text-driven human motion generation. Yet most methods generate in low-dimensional, temporally downsampled latent spaces learned primarily for reconstruction, a bottleneck that can limit generation quality and preclude direct manipulation of individual frames and joints. We introduce MotionSpaceFlow (MSFlow), a representation-aware flow-matching framework that predicts clean motion directly in continuous motion space without a learned encoder or decoder. To account for the anisotropic structure of direct motion representations, we propose representation-aware noise scaling and show how the initial Gaussian source scale governs the covariance of intermediate probability-path marginals. We further introduce a Representation-Aware Multimodal Diffusion Transformer (RA-MMDiT), which jointly updates token-level language and full-resolution motion features through joint attention while adapting temporal information flow to the motion representation: causal attention for incremental features defined by frame-to-frame changes, and bidirectional attention for global features such as absolute joint coordinates. Across different datasets and motion representations, MSFlow achieves state-of-the-art text-to-motion performance. Its global representation variant additionally enables zero-shot, inference-time control over any joint or frame through projection sampling without control-conditioned training, delivering leading motion quality with exact constraint satisfaction.
Figures & tables
| Method | Representation | R-Precision | FID | MM-Dist | MModality | CLIP | ||
| Top 1 | Top 2 | Top 3 | ||||||
| T2M-GPT ( Zhang et al., 2023a ) | Discrete latent | |||||||
| MMM ( Pinyoanuntapong et al., 2024 ) | ||||||||
| MoMask ( Guo et al., 2024 ) | ||||||||
| MLD ( Chen et al., 2023 ) | Continuous latent | |||||||
| SALAD ( Hong et al., 2025 ) | ||||||||
| Controlling Joint | Methods | Zero-shot? | FID | R-Precision | Diversity | Foot Skating | Traj. err. | Loc. err. | Avg. err. |
| Top 3 | Ratio. | ||||||||
| GT | - | - | |||||||
| Pelvis | MDM ( Tevet et al., 2023 ) | ✓ | |||||||
| PriorMDM ( Shafir et al., 2024 ) | ✗ | ||||||||
| GMD ( Karunratanakul et al., 2023 ) | ✓ | ||||||||
| OmniControl ( Xie et al., 2024 ) | ✗ |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Configuration | R-Precision | FID | MM-Dist | CLIP | ||||||
| Factor | Rep. | Attention | Blocks/width | Token Refiner | Top 1 | Top 2 | Top 3 | |||
| Capacity | 263D | MM | ✓ | |||||||
| MM | ✓ | |||||||||
| MM | ✓ | |||||||||
| XYZ | MM | ✓ | ||||||||
| MM | ✓ | |||||||||
| Factor | Representation | Configuration | R-Precision | FID | MM-Dist | ||
| Top 1 | Top 2 | Top 3 | |||||
| Attention | 263D | Causal | |||||
| Bidirectional | |||||||
| XYZ | Causal | ||||||
| Bidirectional | |||||||
| Source scale | 263D | , causal | |||||
| Method | Representation | R-Precision | FID | MM-Dist | MModality | ||
| Top 1 | Top 2 | Top 3 | |||||
| GT | – | – | |||||
| T2M-GPT ( Zhang et al., 2023a ) | Discrete latent | ||||||
| MMM ( Pinyoanuntapong et al., 2024 ) | |||||||
| MoMask ( Guo et al., 2024 ) | |||||||
| IRG-MotionLLM ( Li et al., 2026 ) | – | ||||||
| Method | Representation | FID | R-Precision | MM-Dist | ||
| Top 1 | Top 2 | Top 3 | ||||
| Real motion | – | |||||
| T2M-GPT ( Zhang et al., 2023a ) | Discrete latent | |||||
| MotionGPT ( Jiang et al., 2023 ) | ||||||
| MoMask ( Guo et al., 2024 ) | ||||||
| AttT2M ( Zhong et al., 2023 ) | ||||||
| Method | Representation | R-Precision | FID | MModality | ||
| Top 1 | Top 2 | Top 3 | ||||
| GT | – | – | ||||
| T2M-GPT ( Zhang et al., 2023a ) | Discrete latent | |||||
| MoMask ( Guo et al., 2024 ) | ||||||
| MoMask ++ ( Guo et al., 2025 ) | ||||||
| ScaleMoGen ( Hwang et al., 2026 ) | ||||||
| Method | Params (M) | FLOPs (TF) | Time (s) | FID | R-Top3 |
| MARDM ( Meng et al., 2025b ) | 309.65 | 75.477 | 9.579 | 0.116 | 0.805 |
| SALAD ( Hong et al., 2025 ) | 10.06 | 0.467 | 0.477 | 0.076 | 0.857 |
| MotionStreamer ( Xiao et al., 2025 ) | 318.51 | 1.959 | 9.023 | 0.201 | 0.793 |
| CMDM ( Yu et al., 2026 ) | 115.01 | 0.572 | 1.809 | 0.068 | 0.860 |
| FloodDiffusion ( Cai et al., 2026 ) | 131.22 | 7.060 | 3.537 | 0.057 | 0.810 |
| MSFlow | 68.77 | 1.502 | 1.742 | 0.057 | 0.863 |
| Controlling Joint | Methods | Zero-shot? | FID | R-Precision | Diversity | Foot Skating | Traj. err. | Loc. err. | Avg. err. |
| Top 3 | Ratio. | ||||||||
| GT | - | ||||||||
| Train On Pelvis | MDM ( Tevet et al., 2023 ) | ✓ | |||||||
| PriorMDM ( Shafir et al., 2024 ) | ✗ | ||||||||
| GMD ( Karunratanakul et al., 2023 ) | ✓ | ||||||||
| OmniControl ( Xie et al., 2024 ) | ✗ |
| Item | Selected value |
| Model variant | causal 263D, |
| Training representation | HumanML3D, 263 dimensions |
| Evaluation representation | root/RIC subset, 67 dimensions |
| Backbone | 8 joint blocks, width 512, 4 heads |
| Motion block size | 1 frame |
| Training / inference mask | causal / causal |
| Item | Selected value |
| Model variant | bidirectional XYZ, |
| Training representation | shared-axis-normalized XYZ, 66 dimensions |
| Evaluation representation | root/RIC subset, 67 dimensions |
| Backbone | 8 joint blocks, width 512, 4 heads |
| Motion block size | 1 frame |
| Training / inference mask | bidirectional / bidirectional |
| Item | Control-evaluation setting |
| Training attention | bidirectional, fixed for the XYZ checkpoint |
| Inference attention | bidirectional |
| Control training | none |
| Raw controls | any valid frames, selected joint IDs, and XYZ axes |
| CFG weight / projected steps | 3.0 / 100 |
| Source noise | scale of the training path |
| Representation | |||||
| 263D | 32 | 22,326 | |||
| 64 | 20,850 | ||||
| 128 | 13,328 | ||||
| XYZ | 32 | 22,326 | |||
| 64 | 20,850 | ||||
| 128 | 13,328 |
| Method | Latent shape | |||
| CMDM | ||||
| SALAD |
| Representation | Scale | FID | |||
| 263D | 1 | ||||
| 5 | |||||
| XYZ | 1 | ||||
| 5 |
| Representation | Scale | Mask | Future | Local | Span | Text | Text motion |
| 263D | Causal | ||||||
| Bidirectional | |||||||
| Causal | |||||||
| Bidirectional | |||||||
| XYZ | Causal | ||||||
| Bidirectional |