Sign Language Production (SLP) aims to generate sign motions from text. Conditional Flow Matching methods have achieved strong performance in SLP by constructing conditional paths that transform a source distribution into a target distribution. However, existing methods construct these paths via linear interpolation, whereas the rotational geometry of human joints confines valid joint rotations to a manifold embedded in Euclidean space. Consequently, linear interpolation between two sign motions leaves this manifold and ignores the motion distribution on it. In this paper, we revisit SLP from the perspective of manifold transport and propose a Naturalness-guided Manifold Flow Matching framework, termed \textbf{SignNMFlow}, which constructs conditional paths directly on the motion manifold by jointly considering geometric efficiency and the motion distribution. Specifically, we exploit the intrinsic geometry of the manifold and introduce a motion naturalness measure to characterize the motion distribution. By minimizing the kinetic energy under this measure, we learn a naturalness-guided interpolation that couples a closed-form geodesic, which provides geometrically efficient transport, with a learnable deviation that incorporates the motion distribution, thereby significantly improving the fidelity of generated sign motions. Extensive qualitative and quantitative evaluations demonstrate the effectiveness of this work.
Figures & tables
Figure 1: Here, x0 and x1 denote samples from the source and target distributions, and the color transition from blue to orange reflects the distribution of sign motions on the manifold. (a) Linear interpolation leaves the manifold, yielding invalid rotations. (b) Geodesic interpolation stays on the manifold but ignores the motion distribution. (c) Our naturalness-guided interpolation stays on the motion manifold, maintaining geometric efficiency while incorporating the motion distribution.
Methods
Phoenix-2014T
CSL-Daily
DTW-JPE
DTW-PA-JPE
B-T
DTW-JPE
DTW-PA-JPE
B-T
Body
Hand
Body
Hand
BLEU-4
Body
Hand
Body
Hand
BLEU-4
Pr. Tr. ( Saunders et al., 2020 )
15.01
31.77
13.67
11.95
4.94
16.30
32.63
15.98
12.91
3.07
T2Mesh ( Stoll et al., 2022 )
14.04
31.64
13.48
12.06
5.81
13.76
30.37
13.47
12.10
5.11
T2S-GPT ( Yin et al., 2024 )
11.65
19.09
10.38
6.47
9.06
12.32
15.43
11.94
5.93
8.94
S-MGPT ( Jiang et al., 2023 )
10.42
9.08
9.45
3.41
9.68
11.58
11.31
10.81
3.78
8.82
Table 1: Comparisons with state-of-the-art sign language production (text-to-sign) methods. Best mean values are in bold and second-best are underlined .
Figure 2: Qualitative comparisons of generated sign motions between our method with the SOKE, on the test sets of Phoenix-2014T(left), CSL-Daily(middle), and How2Sign(right).
Methods
DTW-JPE
DTW-PA-JPE
B-T
Body
Hand
Body
Hand
BLEU-4
Pr.Tr.
14.74
30.17
14.15
11.57
2.75
T2Mesh
15.50
32.97
13.99
13.47
7.51
T2S-GPT
12.65
18.44
11.48
6.39
11.20
S-MGPT
12.41
13.74
11.23
4.39
11.45
SignFlow
7.98
10.52
6.92
2.27
–
Table 2: Comparisons with state-of-the-art sign language production methods (text-to-sign) on How2Sign Dataset.
Figure 5
Method
Prediction
DTW-JPE ↓
DTW-PA-JPE ↓
All
Body
Hand
All
Body
Hand
CFM
Velocity
21.62
7.43
9.49
11.33
6.64
1.79
Target
20.08
6.70
9.25
10.78
6.07
1.68
Ours
Velocity
18.76
6.44
8.13
9.52
5.69
1.44
Target
16.86
5.73
7.43
8.92
5.10
1.30
Table 5: Ablation results of training objectives.
Figure 4: User study on the CSL-Daily: Ours vs. SOTA methods.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
ϵθ(t,x0,x1)
ϕ(xt,et,t,c)
Architecture
3-layer MLP
8-layer transformer decoder
MLP hidden dimension
512
–
Latent dimension
–
512
Feed-forward dimension
–
2048
Attention heads
–
8
Time embedding dimension
–
2048
Appendix
Table 6: Hyperparameters of the correction network ϵθ and the sign motion generation network ϕ .
Sign Language Production (SLP) faces a fundamental trade-off: direct text-to-pose models suffer from regression-to-the-mean effects, while dictionary-retrieval methods produce disjointed transitions. To resolve this, we propose a novel training paradigm that leverages sparse keyframes to capture the underlying kinematic distribution of human signing. By generating dense motion from discrete anchors, our approach mitigates regression-to-the-mean while ensuring fluid articulation. To achieve this at scale, we introduce FAST, an ultra-efficient sign segmentation model that automatically mines precise temporal boundaries. We then present SignSparK, a Conditional Flow Matching (CFM) framework that utilizes these temporal anchors to synthesize 3D signing sequences. This keyframe-driven formulation also unlocks Keyframe-to-Pose (KF2P) generation, making precise spatiotemporal editing of signing sequences possible. Furthermore, SignSparK scales across four distinct sign languages, constituting the largest multilingual SLP framework to date, and integrates 3D Gaussian Splatting for photorealistic rendering. Extensive evaluations demonstrate that SignSparK achieves state-of-the-art across diverse SLP tasks and multilingual benchmarks. Our code is available at https://github.com/JianHe0628/SignSparK.
Sign Language Production (SLP) is the task of generating avatar sign language motion from natural language text. The quality of the generated motion is typically evaluated by a motion-space Fréchet distance (FID) and back-translation (BT) BLEU score on benchmarks such as How2Sign. Both metrics can improve substantially while the underlying generator fails to faithfully represent the sign language gestures. In this work we propose to evaluate the generated motion at three independent levels: (τ1) initial-pose conditioning, (τ2) output diversity, and (τ3) target faithfulness. We compute these as pairwise-distance ratios using latent representations of a frozen motion autoencoder (MoAE). We evaluate 14 SLP model checkpoints on the How2Sign dataset, including a re-implemented Neural Sign Actors (NSA), and show that τ3 faithfulness is never attained, while FID varies by nearly two orders of magnitude and is uncorrelated with faithfulness. We show that on the isolated gloss dataset ASL3DWord favorable τ3 can be attained, hence isolating the size of the sentence-level paired-dataset as the bottleneck.
Sign language production (SLP) aims to generate continuous signing motion from spoken language, often through gloss-to-pose generation. Prior work mainly follows two paradigms. Generative models synthesize motion from a learned prior or from noise, without reference to an observed signing instance, making rare hand configurations and signer-specific articulation difficult to preserve. Retrieval-based methods reuse real, well-articulated motion segments, but concatenating segments from different signers and co-articulation contexts can introduce rhythm and style inconsistencies across the full sequence, not only at segment boundaries. These limitations suggest a complementary solution: use retrieval to provide realistic articulation, and use learned refinement to impose the global coherence that retrieval alone lacks. We therefore propose retrieve-and-refine, a paradigm that starts from real retrieved motion and refines it into a globally coherent signing sequence rather than generating motion from scratch. Our framework, SignRR, initializes motion from a dictionary of real sign segments and refines the full sequence with a part-aware Residual VQ-VAE, where residual quantization preserves fine hand articulation and temporal length differences are handled in the latent space. Experiments on PHOENIX14T and CSL-Daily show that SignRR achieves state-of-the-art back-translation performance while maintaining competitive pose quality.