Head avatar modeling requires jointly optimizing multiple objectives with different dominant effects on geometry, appearance, and cross-view consistency. However, their relative effectiveness varies across training states, while existing pipelines typically rely on fixed loss weights or handcrafted stage-wise schedules. A central challenge is therefore to identify which optimization direction is more beneficial at each training state. We propose a counterfactual route optimization framework for Gaussian head avatar modeling, which characterizes state-dependent optimization preference from the realized effects of alternative updates rather than predefined heuristic weighting. Starting from the same training state, we perform short-horizon route-restricted lookahead over geometry, appearance, and joint update routes and evaluate their outcomes under a unified utility. The resulting counterfactual evidence is factorized into a geometry--appearance preference and a residual joint advantage, separately capturing the relative preference between individual update directions and the additional benefit of coordinated optimization. We further amortize this offline evidence into a lightweight controller that directly estimates the current optimization preference and applies bounded modulation to the training objectives during full avatar optimization. Experiments on the NeRSemble dataset validate the effectiveness of the proposed design, consistently outperforming existing methods while preserving clearer local facial structures and finer details.
Figures & tables
Figure 1: Overview of the proposed framework. Multi-view observations are first integrated into a hybrid head avatar, with the face modeled by an animatable FLAME mesh and the hair represented by 3D Gaussians, which are jointly rendered to produce the final avatar. In the offline counterfactual route learning stage, geometry, appearance, and joint routes are instantiated from the same sampled training state and evaluated through route-restricted lookahead updates. Their realized utilities are then factorized into geometry–appearance preference and residual joint advantage, forming route evidence that is paired with the corresponding state features to supervise an amortized route controller. During online route modulation , the controller infers the current optimization preference from the training state and accordingly modulates geometry-oriented, photometric and frequency, and consistency-oriented objectives during training.
Figure 2: Qualitative comparison on self-reenactment of head avatars. Our results preserve more accurate mouth shapes and eye structures under expression changes.
Method
Novel-View
Self-Reenactment
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Gaussian Head Avatar ( Xu et al., 2024 )
29.84
0.895
0.083
22.34
0.851
0.145
EAvatar ( Zhang et al., 2025b )
30.16
0.898
0.082
22.56
0.852
0.145
GaussianAvatars ( Qian et al., 2024 )
31.50
0.935
0.060
29.75
0.930
0.066
MeGA ( Wang et al., 2025a )
30.79
0.932
0.066
29.62
0.926
0.071
TensorGA ( Wang et al., 2025b )
31.28
0.931
0.067
29.90
0.931
0.071
Table 1: Quantitative comparison on novel-view synthesis and self-reenactment. Results are averaged over 8 subjects.
Figure 3: Qualitative comparison on cross-reenactment of head avatars. Our results show cleaner lip boundaries and more coherent mouth structures under cross-identity expression transfer.
Figure 4: Qualitative ablation of the route modulation strategy. The enlarged mouth regions show progressively reduced artifacts, with the full model recovering more complete teeth and cleaner lip boundaries.
Method
Novel-View
Self-Reenactment
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Baseline
30.36
0.952
0.066
29.60
0.948
0.067
Fixed-GC
30.98
0.954
0.060
30.04
0.950
0.064
Mean Mod.
31.33
0.954
0.061
30.19
0.952
0.062
Ours
31.71
0.956
0.057
30.98
0.954
0.060
Table 2: Ablation study of the route modulation strategy on novel-view synthesis and self-reenactment.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Qualitative comparison on novel-view synthesis of head avatars. Our results preserve cleaner local boundaries and more complete mouth and tooth structures across viewpoints, with fewer local artifacts than competing methods.
Method
Training Time (h) ↓
Inference Time (ms) ↓
Gaussian Head Avatar
45
79
GaussianAvatars
7
17
TensorGA
10
21
MeGA
14
55
Fresco
17
55
Ours
17
57
Appendix
Table 3: Training and inference efficiency of different head avatar methods.
Figure 6: Local reconstruction error visualization on novel-view synthesis and self-reenactment. The left example shows the mouth region for novel-view synthesis, while the right example shows the eye region for self-reenactment. Lower MAE values indicate smaller reconstruction errors.
Variant
Novel-View
Self-Reenactment
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
GA-only
34.13
0.966
0.030
32.80
0.959
0.038
Joint-only
33.89
0.966
0.030
32.84
0.960
0.038
Parallel-Route
34.78
0.967
0.027
32.92
0.960
0.036
Ours
35.38
0.967
0.025
33.14
0.961
0.035
Appendix
Table 4: Ablation study of the route-evidence formulation on novel-view synthesis and self-reenactment.
Figure 7: Qualitative comparison of the route-evidence ablation on self-reenactment. The full model alleviates artifacts around the lower teeth and produces a cleaner local mouth structure than the ablation variants.
Accurate head modeling requires a stable yet expressive geometric representation. Existing Gaussian-based head avatars commonly rely on parametric templates (e.g., FLAME) for Gaussian initialization and deformation, but these templates lack personalized priors and struggle to represent structures such as hair and clothing. To address this issue, we propose MGAvatar, a Gaussian-mesh hybrid representation that jointly models geometry and appearance through two Gaussian-mesh binding modes. Specifically, we introduce vertex-bound Gaussians and constrain their learnable parameters, enabling progressive mesh deformation to represent complex head geometry, while a pose-dependent offset module accounts for non-rigid deformations. Once geometry is stabilized, MGAvatar switches to face-bound Gaussians for appearance modeling. To improve appearance consistency across novel poses and viewpoints, we introduce a view-conditioned neural color field that alleviates artifacts caused by independently optimized Gaussian colors. In addition, we design a Gaussian offset network to predict Gaussian offset maps in the observation space, providing greater flexibility for face-bound Gaussians to capture dynamic facial textures. Extensive experiments on multi-view and monocular videos show that MGAvatar outperforms existing methods in rendering quality, producing high-fidelity head avatars with rich texture details.
Lei Shi, Sen Peng, Zhiyang Deng +4
Guangdong Provincial/Zhuhai Key Laboratory of IRADS, Beijing Normal-Hong Kong Baptist University, China · School of Informatics, Xiamen University, China · Department of Computer Science, The University of Texas at Dallas, USA +1
Avatar reconstruction has traditionally relied on per-subject optimization that requires hours of computation or on expensive preprocessing that limits scalability. We introduce FFAvatar, a generalizable feed-forward framework that reconstructs high-quality, animatable 3D Gaussian head avatars from few-shot unposed portrait images in seconds. FFAvatar fuses information from multiple source images into a unified canonical Gaussian representation through Multi-View Query-Former, which is animated via FLAME parameters predicted end-to-end directly from pixels, eliminating the overhead of offline FLAME extraction. We further propose a three-stage training curriculum that achieves both broad generalization and high-fidelity reconstruction: (i) scalable pretraining on extensive monocular video data with over 1M identities to learn strong generalizable priors; (ii) multi-view fine-tuning on a small but high-quality dataset of 360-degree captures to enhance geometric fidelity and extreme-view awareness; and (iii) optional personalization that adapts to specific identities for maximum fidelity within 500 optimization steps. Extensive experiments demonstrate that FFAvatar sets a new standard for identity preservation, geometric consistency, and animation fidelity. On the NeRSemble benchmark, it outperforms the state-of-the-art LAM by a substantial 5.5 PSNR gain. Furthermore, FFAvatar enables real-time deployment, reconstructing avatars in 2 seconds without personalization and 10 seconds with personalization, while supporting 49 FPS animation on a single NVIDIA A100 GPU.
Thuan Hoang Nguyen, Jiahao Luo, Yinyu Nie +3
Snap Inc. · MBZUAI · University of California, Santa Cruz
Building one-shot 3D animatable head avatars is an important yet challenging problem. Existing methods generally collapse under large camera pose variations, compromising the realism of 3D avatars. In this work, we propose a new framework to tackle the novel setting of one-shot 3D full-head animatable avatar reconstruction in a single forward pass via inpainted UV-space Gaussian modeling, enabling 360∘ rendering views and real-time animation. To facilitate efficient animation control, we model 3D head avatars with Gaussian primitives embedded on the surface of a parametric face model within the UV space, and project the input image features to the UV space, resulting in incomplete local UV feature maps. To inpaint the missing regions, we obtain knowledge of full-head geometry and textures from rich 3D full-head priors within a pretrained 3D generative adversarial network (GAN) for global full-head feature extraction and multi-view supervision. Specifically, to enhance the fidelity of 3D reconstruction during inpainting, we take advantage of the symmetric nature of the UV space and human faces to fuse incomplete yet detailed local UV feature maps with the extracted global full-head textures, resulting in inpainted UV Gaussian attribute maps for avatar modeling. Extensive experiments demonstrate that our method is the first to achieve high-quality 3D full-head animatable avatar modeling, significantly improving side and back views while outperforming state-of-the-art animation approaches, thereby improving the realism of 3D animatable avatars.
Shuling Zhao, Dan Xu
The Hong Kong University of Science and Technology · Zeekr Automobile R&D Co., Ltd.