BiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion
Organizations: Department of Computer Science and Engineering, Southeast University, Nanjing, China · Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, Nanjing, China · National University of Singapore · Opus AI
Abstract
Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poorly suited to the bidirectional dependencies between language and motion, allowing early prediction errors to persist as fixed context and degrade both temporal coherence and cross-modal consistency. Masked discrete diffusion, which models sequences through iterative bidirectional prediction, offers a natural remedy. We therefore propose BiMoGen (Bidirectional Motion-text Generation), a unified masked discrete diffusion framework for bidirectional motion-text modeling. To stabilize training, we design Decoupled Uni- and Cross-Modal Training, in which masked pretraining first establishes cross-modal correspondence on paired motion-text sequences, after which supervised fine-tuning specializes the model for bidirectional generation. Masked diffusion nonetheless introduces its own source of error, as the model is trained on clean ground-truth context yet encounters self-generated and potentially erroneous context at inference, with errors committed under heavily masked states propagating through subsequent steps. We further introduce Generation-Aware Self-Correction that exposes the model to its own predictions during training and applies correction passes at early sampling steps to revise unreliably committed tokens. Extensive experiments on HumanML3D and KIT-ML demonstrate competitive performance on both tasks, validating the effectiveness of the proposed two-stage training and self-correction designs. The project page is available at https://wengwanjiang.github.io/BiMoGen-Page.
Figures & tables
| Method | T2M | M2T | |||||||||||
| R@1 | R@2 | R@3 | FID | Div | MM Dist | R@1 | R@3 | BLEU@1 | BLEU@4 | ROUGE-L | CIDEr | BERTScore | |
| Real | 0.511 | 0.703 | 0.797 | 0.002 | 9.503 | 2.974 | 0.523 | 0.828 | – | – | – | – | – |
| T2M-Only Models | |||||||||||||
| MDM [ 41 ] | – | – | 0.611 | 0.544 | 9.559 | 5.566 | – | – | – | – | – | – | – |
| MotionDiffuse [ 57 ] | 0.491 | 0.681 | 0.782 | 0.630 | 9.410 | 3.113 | – | – | – | – | – | – | – |
| MLD [ 6 ] | 0.481 | 0.673 | 0.772 | 0.473 | 9.724 | 3.196 | – | – | – | – | – | – | – |
| Method | T2M | M2T | |||||||||||
| R@1 | R@2 | R@3 | FID | Div | MM Dist | R@1 | R@3 | BLEU@1 | BLEU@4 | ROUGE-L | CIDEr | BERTScore | |
| Real | 0.424 | 0.649 | 0.779 | 0.031 | 11.080 | 2.788 | 0.399 | 0.793 | – | – | – | – | – |
| Separated Bidirectional Models | |||||||||||||
| TM2T [ 14 ] | 0.280 | 0.463 | 0.587 | 3.599 | 9.473 | 4.591 | 0.359 | 0.668 | 46.7 | 18.4 | 44.2 | 79.5 | 23.0 |
| LaMP [ 25 ] | 0.479 | 0.691 | 0.826 | 0.141 | 10.929 | 2.704 | 0.540 | 0.844 | – | – | – | – | – |
| MotionGPT3 † [ 65 ] | 0.456 | 0.680 | 0.803 | 0.227 | 11.026 | 2.704 | – | – | – | – | – | – | – |
| Training Strategy | T2M | M2T | |||||||||||||
| PT | SFT | SC | R@1 | R@2 | R@3 | FID | Div | MM Dist | R@1 | R@3 | BLEU@1 | BLEU@4 | ROUGE-L | CIDEr | BERTScore |
| Real | 0.511 | 0.703 | 0.797 | 0.002 | 9.503 | 2.974 | 0.523 | 0.828 | – | – | – | – | – | ||
| ✓ | ✗ | ✗ | 0.416 | 0.604 | 0.730 | 1.219 | 10.189 | 3.650 | 0.434 | 0.736 | 24.3 | 5.5 | 26.4 | 6.7 | 16.8 |
| ✗ | ✓ | ✗ | 0.466 | 0.633 | 0.726 | 0.325 | 9.546 | 3.259 | 0.474 | 0.779 | 54.3 | 16.7 | 34.9 | 41.3 | 26.4 |
| ✓ | ✓ | ✗ | 0.514 | 0.697 | 0.789 | 0.102 | 9.616 | 3.017 | 0.547 | 0.832 | 56.1 | 16.4 | 42.9 | 44.0 | 36.7 |
| ✓ | ✓ | ✓ | 0.555 | 0.744 | 0.841 | 0.069 | 9.524 | 2.733 | 0.573 | 0.866 | 60.1 | 20.1 | 44.4 | 60.2 | 38.5 |
| Method | Text-to-Motion | Motion-to-Text | Computational Cost | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| R@1 | R@3 | FID | MM Dist | Lat. (s) | R@1 | R@3 | BLEU@1 | BLEU@4 | ROUGE-L | BERTScore | #Params | FLOPs | |
| MotionGPT [ 21 ] | 0.492 | 0.778 | 0.232 | 3.096 | 1.04 | 0.543 | 0.827 | 48.2 | 12.5 | 37.4 | 32.4 | 220M | 7.45T |
| MotionGPT3 [ 65 ] | 0.553 | 0.837 | 0.208 | 2.725 | 1.02 | 0.573 | 0.864 | 59.1 | 19.4 | 46.2 | 35.2 | 238M | 11.00T |
| MG-Mo.LLM [ 50 ] | 0.516 | 0.802 | 0.303 | 2.952 | 1.09 | 0.592 | 0.866 | – | 8.1 | – | 36.7 | 220M | 1.66T |
| DiMo (20 steps) [ 59 ] | 0.528 | 0.818 | 0.050 | 2.862 | 1.55 | 0.568 | 0.845 | 62.5 | 22.0 | 47.3 | 35.4 | 473M | 2.56T |
| Ours (5 steps) | 0.533 | 0.832 | 0.118 | 2.787 | 0.20 | 0.467 | 0.749 | 57.6 | 16.8 | 44.3 | 28.9 | 334M | 0.34T |
| SC Schedule | T2M | M2T | |||||||||||
| R@1 | R@3 | FID | Div | MM Dist | Lat. (s) | R@1 | R@3 | BLEU@1 | BLEU@4 | ROUGE-L | CIDEr | BERTScore | |
| Disabled | 0.551 | 0.838 | 0.070 | 9.870 | 2.745 | 0.55 | 0.568 | 0.864 | 59.2 | 18.9 | 45.6 | 58.8 | 39.6 |
| 0.553 | 0.836 | 0.067 | 9.397 | 2.764 | 0.61 | 0.570 | 0.866 | 59.2 | 19.0 | 45.6 | 58.8 | 38.6 | |
| 0.551 | 0.837 | 0.070 | 9.697 | 2.782 | 0.61 | 0.570 | 0.866 | 59.2 | 18.9 | 45.6 | 58.9 | 38.6 | |
| 0.542 | 0.831 | 0.056 | 9.423 | 2.833 | 0.61 | 0.570 | 0.866 | 59.2 | 18.9 | 45.6 | 59.6 | 38.8 | |
| Every step | 0.529 | 0.814 | 0.066 | 9.570 | 2.948 | 0.78 | 0.580 | 0.868 | 60.2 | 19.7 | 45.8 | 61.2 | 41.0 |
| SC Step | Modification (%) | Pred. Stability (%) |
|---|---|---|
| 31.82 | 84.70 | |
| 29.48 | 83.72 | |
| 25.72 | 83.74 | |
| 22.38 | 82.73 |
| Motion Tokenizer | Text-to-Motion Generation | |||||
|---|---|---|---|---|---|---|
| R@1 | R@2 | R@3 | FID | Div | MM Dist | |
| Real | 0.511 | 0.703 | 0.797 | 0.002 | 9.503 | 2.974 |
| Part VQ-VAE [ 66 ] | 0.512 | 0.703 | 0.801 | 0.124 | 9.886 | 3.084 |
| VQ-VAE [ 21 ] | 0.555 | 0.744 | 0.841 | 0.069 | 9.524 | 2.733 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Backbone | T2M | ||||||
|---|---|---|---|---|---|---|---|
| #Param | R@1 | R@2 | R@3 | FID | Div | MM Dist | |
| Real | – | 0.511 | 0.703 | 0.797 | 0.002 | 9.503 | 2.974 |
| Bert-Base [ 8 ] | 133M | 0.544 | 0.737 | 0.829 | 0.063 | 10.177 | 2.792 |
| Bert-Large [ 8 ] | 366M | 0.553 | 0.747 | 0.839 | 0.057 | 10.018 | 2.734 |
| BiMoGen (Ours) | 334M | 0.555 | 0.744 | 0.841 | 0.069 | 9.524 | 2.733 |
| CFG Scale | Text-to-Motion | |||||
|---|---|---|---|---|---|---|
| R@1 | R@2 | R@3 | FID | Div | MM Dist | |
| 1.0 | 0.505 | 0.693 | 0.795 | 0.323 | 9.977 | 3.305 |
| 2.0 | 0.535 | 0.729 | 0.827 | 0.144 | 10.001 | 2.840 |
| 3.0 | 0.550 | 0. 740 | 0.829 | 0.088 | 9.884 | 2.753 |
| 4.0 | 0.551 | 0.739 | 0.838 | 0.070 | 9.870 | 2.745 |
| 5.0 | 0.543 | 0.739 | 0.833 | 0.076 | 9.566 | 2.760 |
| Decoding Strategy | Block Size | R@1 | R@2 | R@3 | FID | Div | MM Dist |
|---|---|---|---|---|---|---|---|
| Parallel sampling | – | 0.555 | 0.744 | 0.841 | 0.069 | 9.524 | 2.733 |
| Semi-Autoregressive sampling | 1 | 0.534 | 0.721 | 0.827 | 0.181 | 9.415 | 2.915 |
| 2 | 0.521 | 0.716 | 0.821 | 0.220 | 9.640 | 2.939 | |
| 4 | 0.521 | 0.717 | 0.818 | 0.336 | 9.534 | 2.985 | |
| 5 | 0.511 | 0.693 | 0.787 | 0.527 | 9.435 | 3.040 | |
| 10 | 0.496 | 0.681 | 0.776 | 1.018 | 9.332 | 3.148 |