Organizations: Department of Computer Science and Engineering, Southeast University, Nanjing, China · Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, Nanjing, China · National University of Singapore · Opus AI
Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poorly suited to the bidirectional dependencies between language and motion, allowing early prediction errors to persist as fixed context and degrade both temporal coherence and cross-modal consistency. Masked discrete diffusion, which models sequences through iterative bidirectional prediction, offers a natural remedy. We therefore propose BiMoGen (Bidirectional Motion-text Generation), a unified masked discrete diffusion framework for bidirectional motion-text modeling. To stabilize training, we design Decoupled Uni- and Cross-Modal Training, in which masked pretraining first establishes cross-modal correspondence on paired motion-text sequences, after which supervised fine-tuning specializes the model for bidirectional generation. Masked diffusion nonetheless introduces its own source of error, as the model is trained on clean ground-truth context yet encounters self-generated and potentially erroneous context at inference, with errors committed under heavily masked states propagating through subsequent steps. We further introduce Generation-Aware Self-Correction that exposes the model to its own predictions during training and applies correction passes at early sampling steps to revise unreliably committed tokens. Extensive experiments on HumanML3D and KIT-ML demonstrate competitive performance on both tasks, validating the effectiveness of the proposed two-stage training and self-correction designs. The project page is available at https://wengwanjiang.github.io/BiMoGen-Page.
Figures & tables
Figure 1 : Overview of the BiMoGen framework. BiMoGen unifies text-to-motion generation and motion-to-text captioning within a single masked discrete diffusion model, and employs a self-correction mechanism that revises unreliable predictions at early sampling steps.
Figure 2 : Overview of the BiMoGen framework. Top: motion and text are represented as discrete tokens in a unified vocabulary, and the model is trained in two-stages, masked pretraining over symmetrically corrupted motion-text sequences and supervised fine-tuning for T2M and M2T generation. Bottom: a self-correction mechanism is applied at both training and sampling, where the model takes its own predictions as input and learns to recover the ground-truth, and analogously revises committed tokens during iterative masked decoding.
Method
T2M
M2T
R@1 ↑
R@2 ↑
R@3 ↑
FID ↓
Div →
MM Dist ↓
R@1 ↑
R@3 ↑
BLEU@1 ↑
BLEU@4 ↑
ROUGE-L ↑
CIDEr ↑
BERTScore ↑
Real
0.511
0.703
0.797
0.002
9.503
2.974
0.523
0.828
–
–
–
–
–
T2M-Only Models
MDM [ 41 ]
–
–
0.611
0.544
9.559
5.566
–
–
–
–
–
–
–
MotionDiffuse [ 57 ]
0.491
0.681
0.782
0.630
9.410
3.113
–
–
–
–
–
–
–
MLD [ 6 ]
0.481
0.673
0.772
0.473
9.724
3.196
–
–
–
–
–
–
–
Table 1 : Comparison of bidirectional motion-text generation performance on HumanML3D. We report Text-to-Motion (T2M) and Motion-to-Text (M2T) results across four method categories. The best results within unified models are highlighted in bold , and the second-best are underlined .
Method
T2M
M2T
R@1 ↑
R@2 ↑
R@3 ↑
FID ↓
Div →
MM Dist ↓
R@1 ↑
R@3 ↑
BLEU@1 ↑
BLEU@4 ↑
ROUGE-L ↑
CIDEr ↑
BERTScore ↑
Real
0.424
0.649
0.779
0.031
11.080
2.788
0.399
0.793
–
–
–
–
–
Separated Bidirectional Models
TM2T [ 14 ]
0.280
0.463
0.587
3.599
9.473
4.591
0.359
0.668
46.7
18.4
44.2
79.5
23.0
LaMP [ 25 ]
0.479
0.691
0.826
0.141
10.929
2.704
0.540
0.844
–
–
–
–
–
MotionGPT3 † [ 65 ]
0.456
0.680
0.803
0.227
11.026
2.704
–
–
–
–
–
–
–
Table 2 : Comparison of bidirectional motion-text generation performance on KIT-ML. We report Text-to-Motion (T2M) and Motion-to-Text (M2T) results across two method categories. The best results within unified models are highlighted in bold , and the second-best are underlined . † denotes a single-task model reported by its author [ 65 ] .
Training Strategy
T2M
M2T
PT
SFT
SC
R@1 ↑
R@2 ↑
R@3 ↑
FID ↓
Div →
MM Dist ↓
R@1 ↑
R@3 ↑
BLEU@1 ↑
BLEU@4 ↑
ROUGE-L ↑
CIDEr ↑
BERTScore ↑
Real
0.511
0.703
0.797
0.002
9.503
2.974
0.523
0.828
–
–
–
–
–
✓
✗
✗
0.416
0.604
0.730
1.219
10.189
3.650
0.434
0.736
24.3
5.5
26.4
6.7
16.8
✗
✓
✗
0.466
0.633
0.726
0.325
9.546
3.259
0.474
0.779
54.3
16.7
34.9
41.3
26.4
✓
✓
✗
0.514
0.697
0.789
0.102
9.616
3.017
0.547
0.832
56.1
16.4
42.9
44.0
36.7
✓
✓
✓
0.555
0.744
0.841
0.069
9.524
2.733
0.573
0.866
60.1
20.1
44.4
60.2
38.5
Table 3 : Ablation on the training strategy on HumanML3D. PT denotes masked pretraining for cross-modal correspondence, SFT denotes supervised fine-tuning for bidirectional generation, and SC denotes the self-correction objective. Bold denotes the best performance.
Method
Text-to-Motion
Motion-to-Text
Computational Cost
R@1 ↑
R@3 ↑
FID ↓
MM Dist ↓
Lat. (s) ↓
R@1 ↑
R@3 ↑
BLEU@1 ↑
BLEU@4 ↑
ROUGE-L ↑
BERTScore ↑
#Params
FLOPs
MotionGPT [ 21 ]
0.492
0.778
0.232
3.096
1.04
0.543
0.827
48.2
12.5
37.4
32.4
220M
7.45T
MotionGPT3 [ 65 ]
0.553
0.837
0.208
2.725
1.02
0.573
0.864
59.1
19.4
46.2
35.2
238M
11.00T
MG-Mo.LLM [ 50 ]
0.516
0.802
0.303
2.952
1.09
0.592
0.866
–
8.1
–
36.7
220M
1.66T
DiMo (20 steps) [ 59 ]
0.528
0.818
0.050
2.862
1.55
0.568
0.845
62.5
22.0
47.3
35.4
473M
2.56T
Ours (5 steps)
0.533
0.832
0.118
2.787
0.20
0.467
0.749
57.6
16.8
44.3
28.9
334M
0.34T
Table 4 : Ablation of the trade-off between quality and computational cost. Representative bidirectional methods are included for reference under the same evaluation protocol. Lat. denotes Latency.
SC Schedule R
T2M
M2T
R@1 ↑
R@3 ↑
FID ↓
Div →
MM Dist ↓
Lat. (s) ↓
R@1 ↑
R@3 ↑
BLEU@1 ↑
BLEU@4 ↑
ROUGE-L ↑
CIDEr ↑
BERTScore ↑
Disabled
0.551
0.838
0.070
9.870
2.745
0.55
0.568
0.864
59.2
18.9
45.6
58.8
39.6
T/4
0.553
0.836
0.067
9.397
2.764
0.61
0.570
0.866
59.2
19.0
45.6
58.8
38.6
T/2
0.551
0.837
0.070
9.697
2.782
0.61
0.570
0.866
59.2
18.9
45.6
58.9
38.6
3T/4
0.542
0.831
0.056
9.423
2.833
0.61
0.570
0.866
59.2
18.9
45.6
59.6
38.8
Every step
0.529
0.814
0.066
9.570
2.948
0.78
0.580
0.868
60.2
19.7
45.8
61.2
41.0
Table 5 : Ablation on the SC schedule during sampling on HumanML3D. Disabled removes SC during sampling only. All variants use the same number of sampling steps T . For T2M, the CFG scale w is fixed to its default value. CFG is not used for M2T. Lat. denotes Latency.
SC Step
Modification (%)
Pred. Stability (%)
T/4
31.82
84.70
T/2
29.48
83.72
3T/4
25.72
83.74
T
22.38
82.73
Table 6 : Self-correction behavior across sampling steps. Modification is the ratio of committed tokens revised by SC. Pred. Stability is the ratio of revised tokens that remain unchanged at the next step.
Motion Tokenizer
Text-to-Motion Generation
R@1 ↑
R@2 ↑
R@3 ↑
FID ↓
Div →
MM Dist ↓
Real
0.511
0.703
0.797
0.002
9.503
2.974
Part VQ-VAE [ 66 ]
0.512
0.703
0.801
0.124
9.886
3.084
VQ-VAE [ 21 ]
0.555
0.744
0.841
0.069
9.524
2.733
Table 7 : T2M results with different motion tokenizers on HumanML3D. Part VQ-VAE follows ParCo [ 66 ] , which discretizes motion by body parts. VQ-VAE denotes the whole body tokenizer [ 21 ] .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1 : Qualitative comparison on text-to-motion generation. Red text marks descriptions that are not correctly reflected in the generated motion. BiMoGen produces motions that better align with the input prompts.
Figure S2 : Visualization of motion in-betweening. Given prefix, suffix, middle, or start–end motion, BiMoGen performs motion in-betweening with smooth transitions. The human in blue represents the input, while the one in orange is generated by BiMoGen.
Backbone
T2M
#Param
R@1 ↑
R@2 ↑
R@3 ↑
FID ↓
Div →
MM Dist ↓
Real
–
0.511
0.703
0.797
0.002
9.503
2.974
Bert-Base [ 8 ]
133M
0.544
0.737
0.829
0.063
10.177
2.792
Bert-Large [ 8 ]
366M
0.553
0.747
0.839
0.057
10.018
2.734
BiMoGen (Ours)
334M
0.555
0.744
0.841
0.069
9.524
2.733
Appendix
Table S1 : T2M results with different bidirectional Transformer backbones on HumanML3D. Bert-Base and Bert-Large use pretrained weights, while BiMoGen is trained from scratch with the same training strategy.
CFG Scale w
Text-to-Motion
R@1 ↑
R@2 ↑
R@3 ↑
FID ↓
Div →
MM Dist ↓
1.0
0.505
0.693
0.795
0.323
9.977
3.305
2.0
0.535
0.729
0.827
0.144
10.001
2.840
3.0
0.550
0. 740
0.829
0.088
9.884
2.753
4.0
0.551
0.739
0.838
0.070
9.870
2.745
5.0
0.543
0.739
0.833
0.076
9.566
2.760
Appendix
Table S2 : Ablation on the CFG scale w for T2M generation. The number of sampling steps T is set to 20, and the self-correction is disabled.
Decoding Strategy
Block Size
R@1 ↑
R@2 ↑
R@3 ↑
FID ↓
Div →
MM Dist ↓
Parallel sampling
–
0.555
0.744
0.841
0.069
9.524
2.733
Semi-Autoregressive sampling
1
0.534
0.721
0.827
0.181
9.415
2.915
2
0.521
0.716
0.821
0.220
9.640
2.939
4
0.521
0.717
0.818
0.336
9.534
2.985
5
0.511
0.693
0.787
0.527
9.435
3.040
10
0.496
0.681
0.776
1.018
9.332
3.148
Appendix
Table S3 : Effect of semi-autoregressive sampling for T2M generation. The first row denotes the standard parallel masked decoding used by BiMoGen. The remaining rows denote block-wise left-to-right semi-autoregressive variants, where tokens within each block are generated in parallel and different blocks are generated sequentially.
University of Sydney, Camperdown 2006, Australia · School of Computer Science, Peking University, Beijing 100871, China · UCAS-Terminus AI Lab, University of Chinese Academy of Sciences, Beijing 101408, China