Generative 3D Gaussian head avatars provide high-quality, efficient rendering, but synthesising the Gaussian representation remains computationally expensive, limiting deployment on resource-constrained and edge devices. We introduce an efficient generator architecture for unconditional 3D Gaussian head synthesis, based on a parameter-efficient synthesis block and depth-wise separable convolutions while retaining style-based conditioning. Our architecture reduces generator complexity without requiring model compression or quantisation. Compared with the baseline model, our approach reduces FLOPs by 94%, parameter count by 70%, and model size by 81%, while maintaining competitive generation quality. We further demonstrate practical CPU inference and browser-based execution on mobile devices using ONNX Runtime, enabling 3D Gaussian avatar synthesis without dedicated GPU hardware or application-specific software. In addition to conventional image-quality metrics, we evaluate multi-view consistency, training cost, and deployment performance. Code, trained models, and evaluation tools will be released publicly.
Figures & tables
Figure 1 : Overview of our efficient 3D Gaussian head-avatar synthesis pipeline and downstream applications. Our model can generate a novel 3D Gaussian avatar from a random latent code and it also supports subsequent animation and latent manipulation.
Figure 2 : Overview of our 3D Gaussian head-avatar generation pipeline.
Figure 3 : Proposed architecture of the modulated convolution block with depth-wise separable convolutions. The modulation is applied to the instance normalised input feature map and the demodulation to the output of the point-wise convolution.
Generation Quality and Compute Cost
Generation
Animation
Method
FID ↓
FLOP (G) ↓
Params (M) ↓
Train ↓
Size ↓
CPU ↓
Pixel 10 ↓
Pixel 10 ↓
(hrs)
(MB)
(ms)
(ms)
(ms)
AGORA ( 5122 ) [ 13 ]
3.17
212
33
50
241
323
2381
32
Ours ( 5122 )
4.15
13
10
29
46
148
453
32
Table 1: Comparison of generation quality at 5122 rasterisation resolution, together with computational cost, generation latency, and animation latency at 25M training images. FID is computed against 5122 reference statistics. The generator backbone, 2562 UV attribute map, deformation plane, and Gaussian count are unchanged; consequently, the remaining generator and deployment measurements are the same as for 2562 rasterisation.
Figure 4 : A pixel level comparison of our proposed model’s inference in PyTorch against an onnx-runtime implementation. Please note the pixel differences are enhanced by 10× for better visibility.
Figure 5 : Our animation model can drive avatar expression based on FLAME parameters. In each row the expression is interpolated from the outermost images.
Method
FLOPs
FID
FID3D↓
PSNRMV↑
SSIMMV↑
Train Time ↓
VRAM ↓
GGHEAD
66 G
9.6
11.57
32.60
0.868
10.51 hrs.
21.7 GB
Ours-Static
2.7 G
12.0
16.40
34.01
0.880
11.59 hrs.
16.0 GB
Table 2 : We show 3D consistency, multi-view quality and efficiency metrics for our method on a smaller static model at 5M images and 2562 resolution.
Figure 6 : Interpolation between random latent codes generated by our static model at 2562 resolution. Smooth transitions between identities indicate that meaningful latent interpolation is preserved by the efficient architecture.
Configuration
FID ↓
Params ↓
FLOP ↓
CPU ↓
(M)
(G)
(ms)
(a) Depth-wise separable convolutions
GGHead [ 25 ]
9.6
28.0
–
–
+ separable synthesis convs
9.6
8.0
–
–
+ separable output convs (C1)
9.8
8.0
–
354
(b) Channel reduction applied to C1
Table 3 : Ablations shown at 5M images and 2562 resolution. (a) shows that depth-wise convolutions are capable of modelling the data distribution and do not lose significant quality. Dashes mark configurations we did not profile.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Params (M)
GFLOP
Size (MB)
Time (ms)
Synthesis
7.74
7.42
31.51
134.8
Deformation
2.13
5.32
14.63
129.4
Full
10.67
12.74
46.12
246.8
Appendix
Table 4 : Cost breakdown of the generator used for 5122 rendering. The branches are profiled individually while the full model is timed as a single end-to-end run of the complete pipeline, so the branch latencies do not add up directly to the full-model latency. All reported times are medians on a single CPU core; the full model takes 171.3 ms when four threads are used. Parameter, FLOP and size figures are identical to those of the 2562 model, since the UV attribute backbone and the deformation plane both remain at 2562 and only the rasterisation resolution changes.
Method
FID ↓
AGORA [ 13 ]
3.8
Ours
4.3
Appendix
Table 5 : Generation quality at 2562 resolution. FID is computed on the PyTorch checkpoints against reference statistics at the matching resolution, following the protocol of Section A.1 .
Figure 7 : Qualitative results from our proposed model and the baseline GGHead model at 2562 resolution. Please note these are trained for 5M images as outlined in Section 4 .
Figure 8 : Latent space manipulation. Traversing individual directions in the intermediate latent space of our generator produces smooth, largely disentangled changes to a single factor of variation while the remaining identity is preserved, showing that the depth-wise separable synthesis blocks retain the editable style space of the baseline architecture.
Figure 9 : Resolution sensitivity of each attribute group. Each curve low-passes one group to the resolution on the horizontal axis and bilinearly restores it, on frozen weights, leaving the other five groups untouched; the dotted line is the unmodified model.
The generation of high-fidelity, animatable 3D human avatars remains a core challenge in computer graphics and vision, with applications in VR, telepresence, and entertainment. Existing approaches based on implicit representations like NeRFs suffer from slow rendering and dynamic inconsistencies, while 3D Gaussian Splatting (3DGS) methods are typically limited to static head generation, lacking dynamic control. We bridge this gap by introducing AGORA, a novel framework that extends 3DGS within a generative adversarial network to produce animatable avatars. Our key contribution is a lightweight, FLAME-conditioned deformation branch that predicts per-Gaussian residuals, enabling identity-preserving, fine-grained expression control while allowing real-time inference. Identity is further preserved through spatial shape conditioning of the identity branch, and expression fidelity is enforced via a dual-discriminator training scheme leveraging synthetic renderings of the parametric mesh. AGORA generates avatars that are not only visually realistic but also precisely controllable. Quantitatively, we outperform state-of-the-art NeRF-based methods on expression accuracy while rendering at 250 FPS on a single GPU and, notably, at ∼9 FPS under CPU-only inference -- to our knowledge the first demonstration of CPU-only animatable 3DGS avatar synthesis. This work represents a significant step toward practical, high-performance digital humans. Project website: https://ramazan793.github.io/AGORA/
Ramazan Fazylov, Sergey Zagoruyko, Aleksandr Parkin +2
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE · Polynome AI, Dubai, UAE · MTS AI, Moscow, Russia
High-fidelity 3D Gaussian head avatar generation is critical for applications such as AR/VR, telepresence, and digital humans. Existing methods depend on multi-view datasets, 3D captures, or intermediate 2D view synthesis. In contrast, we learn both conditional and unconditional 3D head models from randomly sampled 2D images alone, without using multi-view data, 3D supervision, or intermediate view generation. We introduce MVCHead, a single-shot state space model that enforces multi-view consistency (MVC) directly in the 3D representation while regressing 3D Gaussians under these constraints. At its core, we propose a Hierarchical State Space (HiSS) block that progressively refines Gaussians from coarse to fine, while capturing long-range dependencies. Within each HiSS block, we modify Mamba's standard unidirectional scan with the proposed Hierarchical Bi-directional State Scan (HiBiSS) that aligns recurrence with the axes along which multi-view inconsistencies are strongest. Finally, we design an SE(3) Multi-view Critic that judges whether a set of self-renders arises from a single underlying 3D configuration, rewarding cross-view pixel alignment without observing real multi-view pairs. MVCHead achieves state-of-the-art perceptual quality, surpasses prior methods in both texture and geometric consistency, and maintains comparable shape consistency. To demonstrate scalability, we release FaceGS-10K, the first large-scale dataset of ready-to-use 3D Gaussian head assets for training and evaluation of 3D head models. Project Page and code: https://humansensinglab.github.io/MVCHead/
Building one-shot 3D animatable head avatars is an important yet challenging problem. Existing methods generally collapse under large camera pose variations, compromising the realism of 3D avatars. In this work, we propose a new framework to tackle the novel setting of one-shot 3D full-head animatable avatar reconstruction in a single forward pass via inpainted UV-space Gaussian modeling, enabling 360∘ rendering views and real-time animation. To facilitate efficient animation control, we model 3D head avatars with Gaussian primitives embedded on the surface of a parametric face model within the UV space, and project the input image features to the UV space, resulting in incomplete local UV feature maps. To inpaint the missing regions, we obtain knowledge of full-head geometry and textures from rich 3D full-head priors within a pretrained 3D generative adversarial network (GAN) for global full-head feature extraction and multi-view supervision. Specifically, to enhance the fidelity of 3D reconstruction during inpainting, we take advantage of the symmetric nature of the UV space and human faces to fuse incomplete yet detailed local UV feature maps with the extracted global full-head textures, resulting in inpainted UV Gaussian attribute maps for avatar modeling. Extensive experiments demonstrate that our method is the first to achieve high-quality 3D full-head animatable avatar modeling, significantly improving side and back views while outperforming state-of-the-art animation approaches, thereby improving the realism of 3D animatable avatars.
Shuling Zhao, Dan Xu
The Hong Kong University of Science and Technology · Zeekr Automobile R&D Co., Ltd.