We propose Fresco++, a unified optimization framework for fine-grained and view-consistent head avatar reconstruction. Head avatar optimization is typically driven by per-view image supervision, which can lead to premature fitting of unstable high-frequency details and inconsistent local appearance across viewpoints. Fresco++ addresses these challenges by regulating both the progression of visual detail and the formation of cross-view supervision during optimization. For frequency-aware optimization, a progressive curriculum first stabilizes low-frequency structures and then introduces high-frequency constraints to recover fine facial and hair details without amplifying spurious responses at early stages. For cross-view optimization, we introduce Canonical Group Consensus, which associates local observations through shared canonical surface regions and establishes correspondence across different viewpoints. Geometric and visibility-aware screening removes unreliable observations, while the remaining multi-view evidence is aggregated in feature space to form a consensus target for supervising the current rendering. This design enforces local consistency without relying on a specific image-space parameterization and avoids additional rendering of the auxiliary view. Together, the frequency curriculum and canonical consensus provide stable optimization from coarse structures to fine details while maintaining coherent appearance across viewpoints. Extensive experiments on NeRSemble demonstrate improved reconstruction quality and cross-view consistency, while evaluations across diverse avatar representations further confirm the generality and transferability of Fresco++.
Figures & tables
Fig. 1: Overview of the proposed Fresco++. Multi-view images are used to estimate the facial 3DMM and refine the Gaussian hair parameters, while the facial mesh and Gaussian field are jointly rendered to produce the final head avatar. In the frequency optimization branch, low-frequency supervision stabilizes coarse appearance and structure, followed by edge-aware high-frequency refinement under a staged low-to-high frequency schedule. In the Canonical Group Consensus (CGC) branch, local regions are associated through shared canonical FLAME anchors and projected to multiple views. Reliable observations are selected through view-angle, mask-validity, patch-coverage, and normal-visibility checks, after which multi-view GT features are aggregated into a canonical consensus target to supervise the rendered patch. The two branches are jointly integrated into the overall optimization to improve fine-grained reconstruction and cross-view consistency.
Fig. 2: Qualitative comparison on self-reenactment. Fresco++ better preserves subject-specific appearance details and reproduces expression-dependent local structures, such as earrings, eye opening, and cheek wrinkles, under unseen facial motions.
Fig. 3: Qualitative comparison on novel-view synthesis. Fresco++ produces more faithful fine-grained details under unseen viewpoints, particularly in challenging regions such as teeth and facial wrinkles, while maintaining more stable local appearance across view changes.
Method
Novel-View
Self-Reenactment
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Gaussian Head Avatar
29.84
0.895
0.083
22.34
0.851
0.145
EMavatar
30.16
0.898
0.082
22.56
0.852
0.145
GaussianAvatars
31.50
0.935
0.060
29.75
0.930
0.066
MeGA
30.79
0.932
0.066
29.62
0.926
0.071
TensorGA
31.28
0.931
0.067
29.90
0.931
0.071
TABLE I: Quantitative comparison with previous methods on novel-view synthesis and self-reenactment. Best results are bold , and second-best results are underlined .
Fig. 4: Local reconstruction error comparison on self-reenactment and novel-view synthesis. For each example, we show enlarged RGB regions together with the corresponding error maps. Fresco++ exhibits lower local reconstruction errors in challenging facial regions, particularly around the mouth and eyes. The value shown in each error map denotes the mean reconstruction error within the displayed region, and all error maps are visualized using the same color scale.
Fig. 5: Qualitative comparison on cross-reenactment. The source actor in the first column provides the driving expression for the target avatars. Fresco++ reproduces the transferred facial motions more faithfully, including mouth deformation and strong eye-region contraction, while better preserving the target subject’s local appearance.
Fig. 6: Qualitative ablation of cross-view consistency. We compare the baseline, the UV-space consistency used in Fresco, and the proposed CGC across multiple viewpoints. The enlarged regions show that CGC preserves more stable local structures and appearance under viewpoint changes, with clearer boundaries and more consistent fine-grained details.
Variant
Novel-View
Self-Reenactment
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Baseline
30.79
0.932
0.066
29.62
0.926
0.071
+ UV Cons.
31.17
0.933
0.064
29.88
0.927
0.070
+ CGC
31.50
0.935
0.060
30.06
0.928
0.067
Fresco++
32.18
0.939
0.056
30.61
0.931
0.063
TABLE II: Ablation study of the main components in Fresco++. The reported results are averaged over all evaluated subjects under novel-view synthesis and self-reenactment. UV Cons. denotes the UV-space consistency component used in Fresco. Best results are bold , and second-best results are underlined .
Schedule
Novel-View Synthesis
Self-Reenactment
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
LF+HF
31.92
0.960
0.039
30.83
0.952
0.046
HF → LF
32.25
0.961
0.037
31.17
0.953
0.045
LF → HF
32.37
0.962
0.037
31.34
0.955
0.044
TABLE III: Comparison of different frequency optimization schedules on Subject 306. All variants are trained under the same settings and differ only in the scheduling of low- and high-frequency supervision. Best results are bold , and second-best results are underlined .
Fig. 7: Optimization stability under different frequency schedules on Subject 306. We report the cumulative variation of the training-time head RGB loss in (a) and head SSIM loss in (b). Lower values indicate smaller changes in the reconstruction trajectory. After the early stage of training, LF → HF exhibits lower cumulative variation than HF → LF and simultaneous low- and high-frequency optimization.
Fig. 8: Qualitative comparison of different frequency supervision strategies at the early stage of optimization. Results are shown at the 30th training epoch. Early HF tends to introduce spurious local high-frequency details, whereas Early LF produces cleaner and more coherent intermediate reconstructions.
Fig. 9: Qualitative comparison before and after transferring Fresco++ to GaussianAvatars. The enhanced model better recovers expression-dependent local details, including mouth-corner wrinkles and surrounding facial textures, while reducing local artifacts around the eye region.
Representation
Variant
Novel-View Synthesis
Self-Reenactment
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
GaussianAvatars
Original
31.50
0.935
0.060
29.75
0.930
0.066
+ Fresco++
31.92
0.936
0.058
29.86
0.930
0.065
TensorGA
Original
31.28
0.931
0.067
29.90
0.931
0.071
+ Fresco++
31.61
0.933
0.066
30.03
0.931
0.069
TABLE IV: Quantitative evaluation of Fresco++ on different Gaussian avatar representations. We compare GaussianAvatars and TensorGA before and after integrating the proposed frequency curriculum and CGC under novel-view synthesis and self-reenactment.
Fig. 10: Qualitative comparison before and after transferring Fresco++ to TensorGA. The enhanced model produces clearer local mouth textures and more accurately reconstructs the teeth structure.
Representation
Variant
Training Time (h) ↓
Peak Memory (MB) ↓
Inference Time (ms/frame) ↓
GaussianAvatars
Original
7.39
2231
6.12
+ Fresco++
10.90
2293
6.00
TensorGA
Original
9.84
3363
6.52
+ Fresco++
13.03
3479
6.71
TABLE V: Computational overhead of transferring Fresco++ to different Gaussian avatar representations. Results are averaged over two representative subjects under the same hardware and training settings.
Avatar reconstruction has traditionally relied on per-subject optimization that requires hours of computation or on expensive preprocessing that limits scalability. We introduce FFAvatar, a generalizable feed-forward framework that reconstructs high-quality, animatable 3D Gaussian head avatars from few-shot unposed portrait images in seconds. FFAvatar fuses information from multiple source images into a unified canonical Gaussian representation through Multi-View Query-Former, which is animated via FLAME parameters predicted end-to-end directly from pixels, eliminating the overhead of offline FLAME extraction. We further propose a three-stage training curriculum that achieves both broad generalization and high-fidelity reconstruction: (i) scalable pretraining on extensive monocular video data with over 1M identities to learn strong generalizable priors; (ii) multi-view fine-tuning on a small but high-quality dataset of 360-degree captures to enhance geometric fidelity and extreme-view awareness; and (iii) optional personalization that adapts to specific identities for maximum fidelity within 500 optimization steps. Extensive experiments demonstrate that FFAvatar sets a new standard for identity preservation, geometric consistency, and animation fidelity. On the NeRSemble benchmark, it outperforms the state-of-the-art LAM by a substantial 5.5 PSNR gain. Furthermore, FFAvatar enables real-time deployment, reconstructing avatars in 2 seconds without personalization and 10 seconds with personalization, while supporting 49 FPS animation on a single NVIDIA A100 GPU.
Thuan Hoang Nguyen, Jiahao Luo, Yinyu Nie +3
Snap Inc. · MBZUAI · University of California, Santa Cruz
We present FFAvatar, a Transformer-based 3D Gaussian framework for fast construction of high-quality and animatable 4D head avatars from one or more reference portrait images. Unlike existing feed-forward approaches that require a fixed number of input views, FFAvatar supports incremental reconstruction, progressively refining the avatar representation as additional reference images become available. At the core of our method is an alternating attention mechanism that disentangles identity appearance from expression and viewpoint variations, enabling the reconstruction of a canonical 3D appearance that remains consistent across poses and facial expressions. To balance visual fidelity and computational efficiency, we introduce a sparse-to-dense learning paradigm. Coarse appearance features are first learned using sparse primitives anchored to the FLAME vertex level and are subsequently densified in the UV domain to capture fine-grained geometric and texture details. We further propose a plug-and-play motion refinement module that enables subject-specific dynamic personalization by modeling residual motion beyond parametric deformation. Extensive experiments demonstrate that FFAvatar efficiently produces high-fidelity and controllable 4D head avatars, achieving superior flexibility, driving efficiency, and identity-consistent rendering across diverse expressions and viewpoints.
Jianjiang Yao, Ke Xian, Renxiang Dai +1
School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China
Head avatar modeling requires jointly optimizing multiple objectives with different dominant effects on geometry, appearance, and cross-view consistency. However, their relative effectiveness varies across training states, while existing pipelines typically rely on fixed loss weights or handcrafted stage-wise schedules. A central challenge is therefore to identify which optimization direction is more beneficial at each training state. We propose a counterfactual route optimization framework for Gaussian head avatar modeling, which characterizes state-dependent optimization preference from the realized effects of alternative updates rather than predefined heuristic weighting. Starting from the same training state, we perform short-horizon route-restricted lookahead over geometry, appearance, and joint update routes and evaluate their outcomes under a unified utility. The resulting counterfactual evidence is factorized into a geometry--appearance preference and a residual joint advantage, separately capturing the relative preference between individual update directions and the additional benefit of coordinated optimization. We further amortize this offline evidence into a lightweight controller that directly estimates the current optimization preference and applies bounded modulation to the training objectives during full avatar optimization. Experiments on the NeRSemble dataset validate the effectiveness of the proposed design, consistently outperforming existing methods while preserving clearer local facial structures and finer details.
Shikun Zhang, Yong Li, Yiqun Wang +2
Department of Data Science and AI, Monash University, Australia · College of Computer Science, Chongqing University, China