We present SInGA, a novel method for learning Semantic Inpainting for animatable Gaussian head Avatars from a single image. Existing avatar approaches often rely on multi-view observations and lack effective handling of unobserved regions in single-view settings, limiting their applicability in such scenarios. To address this, we propose a semantic inpainting framework defined in UV space for completing unobserved facial regions. Our key insight lies in the structured topology of the UV representation, which provides consistent spatial correspondences and enables reliable completion of identity-specific features using the inherent symmetry cues of human faces. We extract features from observed regions and use them to complete unobserved regions. The completed representation is then used to regress Gaussian attributes, effectively performing Gaussian inpainting. In addition, instead of relying on a single Gaussian at each surface or pixel location, we stack multiple Gaussians to enhance detail. The resulting avatar generalizes across identities without requiring per-identity optimization and can be animated with driving inputs. Experimental results show that our method generates high-quality head avatars with improved completeness and identity preservation, while supporting realistic animation and consistent rendering from unobserved views.
Figures & tables
Figure 1 : SInGA. Our method reconstructs an animatable Gaussian head avatar from a single image in a single forward pass. The reconstructed avatar supports real-time reenactment while maintaining generalization and controllable facial animation.
Figure 2 : Overview of SInGA . Our framework completes unobserved regions in UV space using DINOv2 features as semantic priors to preserve identity-aware information. Given a source image and a driving signal, we represent Gaussians in the UV domain, where the fixed UV topology provides structured connectivity across facial regions. We formulate the completion of unobserved regions as semantic inpainting and fill in missing features using DINOv2 priors. Gaussian attributes are then predicted from the completed representation to construct the Gaussian head avatar. βs , ψd , and θd denote the source shape, driving expression, and driving pose parameters, respectively. Note that during training, the source and driving images are sampled from the same identity.
Figure 3 : Confidence map. We compute a confidence map from the source view to identify well-observed and uncertain UV regions. High confidence is assigned to surface regions facing the source camera, while low confidence appears in unobserved regions and near silhouettes.
Figure 4 : Gaussian stacking. For selected regions, we stack K Gaussians along the surface normal by cumulatively adding the predicted position offsets, increasing local geometric detail. The figure illustrates the case of K=4 .
Figure 5 : Qualitative comparisons. We compare our method with existing methods in the cross reenactment setting.
Self Reenactment
Cross Reenactment
Method
PSNR ↑
SSIM ↑
LPIPS ↓
CSIM ↑
AKD ↓
AED ↓
APD ↓
CSIM ↑
AED ↓
APD ↓
GPAvatar [ 7 ]
20.85
0.767
0.149
0.781
4.44
0.130
0.318
0.527
0.218
0.351
Portrait4D [ 9 ]
17.60
0.696
0.173
0.801
5.15
0.148
0.251
0.619
0.263
0.250
Portrait4D-v2 [ 10 ]
18.45
0.721
0.150
0.848
4.95
0.123
0.294
0.672
0.235
0.303
GAGAvatar [ 6 ]
21.68
0.779
0.118
0.875
4.27
0.105
0.358
0.674
0.202
0.415
LAM [ 18 ]
18.55
0.741
0.145
0.744
4.95
0.160
0.310
0.594
0.257
0.336
Table 1: Quantitative comparison on VFHQ. Comparison with state-of-the-art methods on the VFHQ dataset. Colors denote the best and second-best results.
Self Reenactment
Cross Reenactment
Method
PSNR ↑
SSIM ↑
LPIPS ↓
CSIM ↑
AKD ↓
AED ↓
APD ↓
CSIM ↑
AED ↓
APD ↓
GPAvatar [ 7 ]
24.19
0.857
0.076
0.918
3.42
0.130
0.180
0.839
0.249
0.227
Portrait4D [ 9 ]
19.26
0.776
0.125
0.879
4.27
0.174
0.183
0.805
0.298
0.198
Portrait4D-v2 [ 10 ]
19.16
0.791
0.103
0.912
3.95
0.132
0.220
0.848
0.260
0.239
GAGAvatar [ 6 ]
24.98
0.865
0.067
0.933
3.53
0.120
0.244
0.871
0.214
0.289
LAM [ 18 ]
21.36
0.808
0.096
0.831
4.27
0.184
0.203
0.766
0.279
0.229
Table 2: Quantitative comparison on HDTF. Comparison with state-of-the-art methods on the HDTF dataset. Colors denote the best and second-best results.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
GAGAvatar
18.28
0.526
0.175
Ours (w/o symmetry)
18.78
0.543
0.182
Ours (full)
19.06
0.546
0.179
Table 3: Evaluation on unobserved regions. Reconstruction metrics for source-unobserved and target-visible regions.
A100 GPU
A6000 GPU
Method
GPAvatar
Portrait4D
Portrait4D-v2
GAGAvatar
LAM
GAGAvatar
LAM
Ours
Rendering speed (FPS)
16.86
9.49
9.62
67.12
280.96
56.76
259.58
53.73
Table 4: Runtime comparison. Runtime comparison measured in FPS. The results are averaged over 100 frames, excluding the time for estimating driving parameters that can be computed in advance.
A6000 GPU
Method
GAGAvatar
LAM
Ours
Reconstruction time (sec)
0.04
2.82
0.07
Table 5: Reconstruction time. Reconstruction time comparison measured in seconds.
Self Reenactment
Cross Reenactment
Method
PSNR ↑
SSIM ↑
LPIPS ↓
CSIM ↑
AKD ↓
AED ↓
APD ↓
CSIM ↑
AED ↓
APD ↓
w/o position offset
21.54
0.780
0.123
0.869
3.90
0.101
0.412
0.671
0.204
0.434
w/o symmetry
21.63
0.781
0.121
0.868
3.92
0.100
0.385
0.664
0.208
0.425
Ours
21.76
0.784
0.119
0.871
3.92
0.101
0.399
0.676
0.206
0.445
Table 6: Ablation study. Ablations on the VFHQ dataset.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
w/o stacking
24.77
0.774
0.082
Ours
24.99
0.782
0.081
Table 7: Ablation study on Gaussian stacking. Ablation results for Gaussian stacking in the self reenactment setting.
Figure 13
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Layer
Filter Size
Input Channels
Output Channels
Activation
Conv1
3×3
283
128
ReLU
Conv2
3×3
128
128
ReLU
Conv3
3×3
128
128
ReLU
Conv4
1×1
128
1
Tanh
Appendix
Table 8: Architecture of the geometry network.
Layer
Filter Size
Input Channels
Output Channels
Activation
Conv1
3×3
283
128
ReLU
Conv2
3×3
128
128
ReLU
Conv3
3×3
128
128
ReLU
Conv4
1×1
128
40
-
Appendix
Table 9: Architecture of the appearance network.
Figure 8 : Additional cross reenactment results on VFHQ and HDTF datasets.
Figure 9 : Self reenactment results on VFHQ and HDTF datasets.
Figure 10 : Qualitative results of our method on in-the-wild images. Source images are shown in the first column and driving images in the top row, and the remaining images show the generated results.