We propose Fresco++, a unified optimization framework for fine-grained and view-consistent head avatar reconstruction. Head avatar optimization is typically driven by per-view image supervision, which can lead to premature fitting of unstable high-frequency details and inconsistent local appearance across viewpoints. Fresco++ addresses these challenges by regulating both the progression of visual detail and the formation of cross-view supervision during optimization. For frequency-aware optimization, a progressive curriculum first stabilizes low-frequency structures and then introduces high-frequency constraints to recover fine facial and hair details without amplifying spurious responses at early stages. For cross-view optimization, we introduce Canonical Group Consensus, which associates local observations through shared canonical surface regions and establishes correspondence across different viewpoints. Geometric and visibility-aware screening removes unreliable observations, while the remaining multi-view evidence is aggregated in feature space to form a consensus target for supervising the current rendering. This design enforces local consistency without relying on a specific image-space parameterization and avoids additional rendering of the auxiliary view. Together, the frequency curriculum and canonical consensus provide stable optimization from coarse structures to fine details while maintaining coherent appearance across viewpoints. Extensive experiments on NeRSemble demonstrate improved reconstruction quality and cross-view consistency, while evaluations across diverse avatar representations further confirm the generality and transferability of Fresco++.
Figures & tables
Fig. 1: Overview of the proposed Fresco++. Multi-view images are used to estimate the facial 3DMM and refine the Gaussian hair parameters, while the facial mesh and Gaussian field are jointly rendered to produce the final head avatar. In the frequency optimization branch, low-frequency supervision stabilizes coarse appearance and structure, followed by edge-aware high-frequency refinement under a staged low-to-high frequency schedule. In the Canonical Group Consensus (CGC) branch, local regions are associated through shared canonical FLAME anchors and projected to multiple views. Reliable observations are selected through view-angle, mask-validity, patch-coverage, and normal-visibility checks, after which multi-view GT features are aggregated into a canonical consensus target to supervise the rendered patch. The two branches are jointly integrated into the overall optimization to improve fine-grained reconstruction and cross-view consistency.
Fig. 2: Qualitative comparison on self-reenactment. Fresco++ better preserves subject-specific appearance details and reproduces expression-dependent local structures, such as earrings, eye opening, and cheek wrinkles, under unseen facial motions.
Fig. 3: Qualitative comparison on novel-view synthesis. Fresco++ produces more faithful fine-grained details under unseen viewpoints, particularly in challenging regions such as teeth and facial wrinkles, while maintaining more stable local appearance across view changes.
Method
Novel-View
Self-Reenactment
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Gaussian Head Avatar
29.84
0.895
0.083
22.34
0.851
0.145
EMavatar
30.16
0.898
0.082
22.56
0.852
0.145
GaussianAvatars
31.50
0.935
0.060
29.75
0.930
0.066
MeGA
30.79
0.932
0.066
29.62
0.926
0.071
TensorGA
31.28
0.931
0.067
29.90
0.931
0.071
TABLE I: Quantitative comparison with previous methods on novel-view synthesis and self-reenactment. Best results are bold , and second-best results are underlined .
Fig. 4: Local reconstruction error comparison on self-reenactment and novel-view synthesis. For each example, we show enlarged RGB regions together with the corresponding error maps. Fresco++ exhibits lower local reconstruction errors in challenging facial regions, particularly around the mouth and eyes. The value shown in each error map denotes the mean reconstruction error within the displayed region, and all error maps are visualized using the same color scale.
Fig. 5: Qualitative comparison on cross-reenactment. The source actor in the first column provides the driving expression for the target avatars. Fresco++ reproduces the transferred facial motions more faithfully, including mouth deformation and strong eye-region contraction, while better preserving the target subject’s local appearance.
Fig. 6: Qualitative ablation of cross-view consistency. We compare the baseline, the UV-space consistency used in Fresco, and the proposed CGC across multiple viewpoints. The enlarged regions show that CGC preserves more stable local structures and appearance under viewpoint changes, with clearer boundaries and more consistent fine-grained details.
Variant
Novel-View
Self-Reenactment
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Baseline
30.79
0.932
0.066
29.62
0.926
0.071
+ UV Cons.
31.17
0.933
0.064
29.88
0.927
0.070
+ CGC
31.50
0.935
0.060
30.06
0.928
0.067
Fresco++
32.18
0.939
0.056
30.61
0.931
0.063
TABLE II: Ablation study of the main components in Fresco++. The reported results are averaged over all evaluated subjects under novel-view synthesis and self-reenactment. UV Cons. denotes the UV-space consistency component used in Fresco. Best results are bold , and second-best results are underlined .
Schedule
Novel-View Synthesis
Self-Reenactment
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
LF+HF
31.92
0.960
0.039
30.83
0.952
0.046
HF → LF
32.25
0.961
0.037
31.17
0.953
0.045
LF → HF
32.37
0.962
0.037
31.34
0.955
0.044
TABLE III: Comparison of different frequency optimization schedules on Subject 306. All variants are trained under the same settings and differ only in the scheduling of low- and high-frequency supervision. Best results are bold , and second-best results are underlined .
Fig. 7: Optimization stability under different frequency schedules on Subject 306. We report the cumulative variation of the training-time head RGB loss in (a) and head SSIM loss in (b). Lower values indicate smaller changes in the reconstruction trajectory. After the early stage of training, LF → HF exhibits lower cumulative variation than HF → LF and simultaneous low- and high-frequency optimization.
Fig. 8: Qualitative comparison of different frequency supervision strategies at the early stage of optimization. Results are shown at the 30th training epoch. Early HF tends to introduce spurious local high-frequency details, whereas Early LF produces cleaner and more coherent intermediate reconstructions.
Fig. 9: Qualitative comparison before and after transferring Fresco++ to GaussianAvatars. The enhanced model better recovers expression-dependent local details, including mouth-corner wrinkles and surrounding facial textures, while reducing local artifacts around the eye region.
Representation
Variant
Novel-View Synthesis
Self-Reenactment
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
GaussianAvatars
Original
31.50
0.935
0.060
29.75
0.930
0.066
+ Fresco++
31.92
0.936
0.058
29.86
0.930
0.065
TensorGA
Original
31.28
0.931
0.067
29.90
0.931
0.071
+ Fresco++
31.61
0.933
0.066
30.03
0.931
0.069
TABLE IV: Quantitative evaluation of Fresco++ on different Gaussian avatar representations. We compare GaussianAvatars and TensorGA before and after integrating the proposed frequency curriculum and CGC under novel-view synthesis and self-reenactment.
Fig. 10: Qualitative comparison before and after transferring Fresco++ to TensorGA. The enhanced model produces clearer local mouth textures and more accurately reconstructs the teeth structure.
Representation
Variant
Training Time (h) ↓
Peak Memory (MB) ↓
Inference Time (ms/frame) ↓
GaussianAvatars
Original
7.39
2231
6.12
+ Fresco++
10.90
2293
6.00
TensorGA
Original
9.84
3363
6.52
+ Fresco++
13.03
3479
6.71
TABLE V: Computational overhead of transferring Fresco++ to different Gaussian avatar representations. Results are averaged over two representative subjects under the same hardware and training settings.