In this work, we focus on photorealistic sign avatar modeling, which is crucial for effective communication with the Deaf community and is characterized by complex hand gestures and nuanced facial expressions. To this end, we introduce MVSign, the first multi-view Chinese sign language dataset co-designed with Deaf experts, featuring diverse gestures and rich annotations. For precise SMPL-X annotation, we develop a hybrid fitting pipeline that produces accurate body, hand, and facial parameters and can also be applied to the monocular setting. Building on MVSign, we propose a decoupled sign avatar representation that isolates body, head, and hand components to capture complex articulations, together with a motion-aware sampling strategy to handle motion blur and balance gesture diversity. Extensive experiments demonstrate that our method achieves high-fidelity visual results on MVSign, particularly in detailed hand and facial regions, and generalizes well to in-the-wild monocular sign language videos. Project page: https://naaapi.github.io/PHOSA.
Figures & tables
Dataset
Sign
Head
SMPL-X
#View
#ID
#Frames
Resolution
PHOENIX14T [ 3 ]
✔
✗
✗
1
9
0.94M
260P
CSL-Daily [ 62 ]
✔
✗
✗
1
10
1.5M
512P
How2Sign [ 8 ]
✔
✗
✗
2
11
5.7M
720P
SignAvatars [ 55 ]
✔
✗
✔
1
153
8.34M
-
SMILE [ 10 ]
✔
✗
✗
4
66
-
1080P
VSL * [ 42 ]
✔
✗
✗
16
2
50K
4096P
Table 1 : Dataset statistics comparison with existing multi-view human-centric datasets and sign language datasets.
Figure 2 : Overview of the hybrid SMPL-X fitting pipeline. Our pipeline fuses outputs from multiple models to leverage their complementary strengths, achieving precise and robust SMPL-X parameter estimation.
Figure 3 : Overview of sign avatar modeling pipeline. The posed vertices are taken as vertex color on the canonical SMPL-X template, and then rendered to generate posed position maps. For the hand part, we employ partial kinematic decoupling by fixing body joints, only preserving hand joint mobility. We then predict pose-dependent Gaussian maps through three specialized StyleUNet, deform the Gaussians by LBS, and render the synthesized avatar by differentiable rasterization.
Method
Full
Hand
Face
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
w/o Hybrid Fitting
23.58
0.9637
0.0618
16.82
0.7371
0.2987
16.62
0.7923
0.2435
w Hybrid Fitting
26.90
0.9722
0.0370
18.85
0.7704
0.2301
20.51
0.8499
0.1665
Table 2 : Ablation study of hybrid SMPL-X fitting strategy. “↓” indicates that lower values are better, while “↑” means the opposite.
Figure 5
Method
Full
Hand
Face
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
1
26.17
0.9711
0.0425
18.49
0.7599
0.2581
18.84
0.8223
0.1927
8
26.42
0.9711
0.0417
18.60
0.7649
0.2524
19.43
0.8286
0.1759
Full (16)
26.90
0.9722
0.0370
18.85
0.7704
0.2301
20.51
0.8499
0.1665
Table 4 : Impact of the number of views on MVSign dataset. “↓” indicates that lower values are better, while “↑” means the opposite.
Method
Full
Hand
Face
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Only Basic Hand Shapes
24.08
0.9660
0.0524
17.47
0.7425
0.2892
17.55
0.8034
0.2287
Only Sign Sentences
24.36
0.9653
0.0501
17.62
0.7518
0.2768
17.76
0.8122
0.2139
Ours
26.90
0.9722
0.0370
18.85
0.7704
0.2301
20.51
0.8499
0.1665
Table 5 : Impact of incorporated signing patterns on MVSign dataset. “↓” indicates that lower values are better, while “↑” means the opposite.
Figure 8
Strategy
Full
Hand
Face
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Sequential
26.55
0.9659
0.0372
16.40
0.6734
0.2656
18.66
0.7733
0.1875
Isometric
26.47
0.9659
0.0371
16.29
0.6697
0.2686
18.57
0.7726
0.1906
Random
26.90
0.9663
0.0353
17.40
0.6939
0.2490
18.94
0.7838
0.1784
Motion-aware
27.67
0.9684
0.0340
18.03
0.7161
0.2387
19.79
0.8029
0.1672
Table 6 : Ablation study of data sampling strategy. “↓” indicates that lower values are better, while “↑” means the opposite.
Method
Full
Hand
Face
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Baseline
25.24
0.9570
0.0490
16.65
0.6912
0.3759
18.47
0.7948
0.2167
+ DR
26.28
0.9607
0.0430
17.53
0.7174
0.3188
20.53
0.8099
0.1731
+ PHK
27.93
0.9665
0.0388
18.74
0.7322
0.2666
21.50
0.8253
0.1595
Table 7 : Ablation study of decoupled sign avatar representation. “DR” indicates the decoupled representation, while “PHK” means Partial Hand Kinematic. “↓” indicates that lower values are better, while “↑” means the opposite.
Table 11
Method
Full
Hand
Face
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Multi-view setting: MVSign dataset
SplattingAvatar [ 47 ]
23.71
0.9625
0.0494
15.64
0.6762
0.3884
16.90
0.7633
0.2423
GaussianAvatar [ 20 ]
24.23
0.9633
0.0428
16.30
0.6833
0.3082
17.78
0.7737
0.2086
AnimatableGS [ 30 ]
25.09
0.9647
0.0465
16.95
0.7002
0.3135
18.63
0.7842
0.2230
EVA [ 19 ]
25.56
0.9667
0.0436
17.43
0.7154
0.2869
19.21
0.8000
0.1964
Table 10 : Comparison to state-of-the-art human avatar modeling methods. “↓” indicates that lower values are better, while “↑” means the opposite.
Figure 7 : Qualitative comparison of novel pose synthesis with GaussianAvatar [ 20 ] , AnimatableGS [ 30 ] , EVA [ 19 ] and Mmlphuman [ 57 ] . The first two rows show results on the MVSign dataset, while the last row presents results on real-world web videos.