Sign Language Production (SLP) aims to generate sign motions from text. Conditional Flow Matching methods have achieved strong performance in SLP by constructing conditional paths that transform a source distribution into a target distribution. However, existing methods construct these paths via linear interpolation, whereas the rotational geometry of human joints confines valid joint rotations to a manifold embedded in Euclidean space. Consequently, linear interpolation between two sign motions leaves this manifold and ignores the motion distribution on it. In this paper, we revisit SLP from the perspective of manifold transport and propose a Naturalness-guided Manifold Flow Matching framework, termed \textbf{SignNMFlow}, which constructs conditional paths directly on the motion manifold by jointly considering geometric efficiency and the motion distribution. Specifically, we exploit the intrinsic geometry of the manifold and introduce a motion naturalness measure to characterize the motion distribution. By minimizing the kinetic energy under this measure, we learn a naturalness-guided interpolation that couples a closed-form geodesic, which provides geometrically efficient transport, with a learnable deviation that incorporates the motion distribution, thereby significantly improving the fidelity of generated sign motions. Extensive qualitative and quantitative evaluations demonstrate the effectiveness of this work.
Figures & tables
Figure 1: Here, x0 and x1 denote samples from the source and target distributions, and the color transition from blue to orange reflects the distribution of sign motions on the manifold. (a) Linear interpolation leaves the manifold, yielding invalid rotations. (b) Geodesic interpolation stays on the manifold but ignores the motion distribution. (c) Our naturalness-guided interpolation stays on the motion manifold, maintaining geometric efficiency while incorporating the motion distribution.
Methods
Phoenix-2014T
CSL-Daily
DTW-JPE
DTW-PA-JPE
B-T
DTW-JPE
DTW-PA-JPE
B-T
Body
Hand
Body
Hand
BLEU-4
Body
Hand
Body
Hand
BLEU-4
Pr. Tr. ( Saunders et al., 2020 )
15.01
31.77
13.67
11.95
4.94
16.30
32.63
15.98
12.91
3.07
T2Mesh ( Stoll et al., 2022 )
14.04
31.64
13.48
12.06
5.81
13.76
30.37
13.47
12.10
5.11
T2S-GPT ( Yin et al., 2024 )
11.65
19.09
10.38
6.47
9.06
12.32
15.43
11.94
5.93
8.94
S-MGPT ( Jiang et al., 2023 )
10.42
9.08
9.45
3.41
9.68
11.58
11.31
10.81
3.78
8.82
Table 1: Comparisons with state-of-the-art sign language production (text-to-sign) methods. Best mean values are in bold and second-best are underlined .
Figure 2: Qualitative comparisons of generated sign motions between our method with the SOKE, on the test sets of Phoenix-2014T(left), CSL-Daily(middle), and How2Sign(right).
Methods
DTW-JPE
DTW-PA-JPE
B-T
Body
Hand
Body
Hand
BLEU-4
Pr.Tr.
14.74
30.17
14.15
11.57
2.75
T2Mesh
15.50
32.97
13.99
13.47
7.51
T2S-GPT
12.65
18.44
11.48
6.39
11.20
S-MGPT
12.41
13.74
11.23
4.39
11.45
SignFlow
7.98
10.52
6.92
2.27
–
Table 2: Comparisons with state-of-the-art sign language production methods (text-to-sign) on How2Sign Dataset.
Figure 5
Method
Prediction
DTW-JPE ↓
DTW-PA-JPE ↓
All
Body
Hand
All
Body
Hand
CFM
Velocity
21.62
7.43
9.49
11.33
6.64
1.79
Target
20.08
6.70
9.25
10.78
6.07
1.68
Ours
Velocity
18.76
6.44
8.13
9.52
5.69
1.44
Target
16.86
5.73
7.43
8.92
5.10
1.30
Table 5: Ablation results of training objectives.
Figure 4: User study on the CSL-Daily: Ours vs. SOTA methods.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
ϵθ(t,x0,x1)
ϕ(xt,et,t,c)
Architecture
3-layer MLP
8-layer transformer decoder
MLP hidden dimension
512
–
Latent dimension
–
512
Feed-forward dimension
–
2048
Attention heads
–
8
Time embedding dimension
–
2048
Appendix
Table 6: Hyperparameters of the correction network ϵθ and the sign motion generation network ϕ .