3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy neural decoding with a shallow coefficient predictor and a linear blend. To improve fidelity and reduce memory requirements, we propose to construct the basis using block-local PCA under a rendering-aware metric and a memory budget. Our method learns a shallow MLP network to predict blendshape coefficients and applies to various animation architectures without retraining original models. We validate GALA by accelerating the inference of three distinct avatar models for 3D animation of facial expressions and full-bodies with clothing dynamics. Across these models, our distillation generalizes to held-out identities and reduces CPU animation cost by up to three orders of magnitude while preserving most of the rendering quality. Excellent results of our method confirm the shared linear structure of learned avatar representations and enable highly efficient and accurate animation at frame rates reaching up to 60fps on mobile devices. Project page: https://ramazan793.github.io/gala/
Figures & tables
Figure 1: GALA enables efficient, high-fidelity animation across diverse Gaussian avatar models. By introducing shared blendshapes and a lightweight coefficient predictor, GALA accelerates CPU animation by up to three orders of magnitude and enables real-time deployment on mobile devices.
Photo- real.
Identity- shared
Linear at run time
Where the basis comes from
3D morphable models ( Li et al., 2017b )
×
✓
✓
hand-built from scans
Per-person blendshapes ( Zielonka et al., 2025 )
✓
×
✓
fitted per subject
Fixed-rig avatars ( He et al., 2025 )
✓
✓
✓
chosen up front
Animatable avatar models ( Yu et al., 2025 )
✓
–
×
none (a network)
GALA (Ours)
✓
✓
✓
discovered in a trained model
Table 1: Comparison of avatar animation representations. Our method extracts an identity-shared blendshape basis from a pretrained nonlinear avatar model, complementing approaches based on predefined rigs or per-subject fitting. The extracted basis closely approximates the model’s learned deformations and supports efficient linear blending for unseen identities.
Figure 2: Overview of GALA. Residuals ri of the host model A with respect to the neutral avatar A(wi,θ0) (Section 3.2.1 ) are decomposed by a spatially local, rendering-aware PCA under a memory budget B into the block-diagonal basis U (Section 3.2.2 ). A coefficient network fϕ is trained to regress the projected coefficients c⋆ (Section 3.2.3 ). At inference, the neutral avatar is computed once per subject, and every frame costs one evaluation of a shallow network fϕ and a linear blend.
FID ↓
AED ↓
AED-jaw ↓
ID ↑
APD ↓
FPS ↑
mFPS ↑
AGORA
3.17
0.682
0.021
0.75
0.025
250 ∗
1
EG3D
3.28
–
–
–
–
–
–
GGHEAD
4.06
–
–
–
–
–
–
Next3D
3.18
0.930
0.046
0.74
0.031
15 ∗
–
GAIA
3.85
0.530
0.040
0.72
0.027
52 ∗
–
AGORA + GALA (ours)
3.47
0.705
0.022
0.76
0.026
1,275
60
Table 2: Comparison to the state of the art on the benchmark of each host model. In every panel the first row is the host model and the last row is the same model distilled by our method (+GALA). The rows in between are the methods compared in the paper of each host model, with the numbers reported there. The AGORA row is taken from Fazylov et al. (2025) . The FlexAvatar and DynaAvatar rows are measured by us on the authors’ released checkpoints, following their evaluation protocols (Appendix F ). Best and second-best results, excluding the host model. EG3D and GGHead are not animatable. FPS is measured on a desktop GPU with a precomputed neutral avatar and mFPS on a mobile phone in a web browser. ∗ As reported in the model’s paper.
Figure 3: Qualitative comparison with the host models ( Original ), two examples per host model. Ours (projected) uses projected coefficients c⋆ , which require running the host model and show what the basis can express (Section 3.2.3 ), and Ours the coefficients predicted by the coefficient network. Error maps show the absolute difference to the host model per pixel, averaged over the RGB channels (values in [0,1] ), from 0 (black) to 0.2 (yellow). Zoom in for details.
Per-attribute
Rendering-aware
PSNR to host ↑
bases
PCA
allocation
AGORA
Flex
Dyna
×
×
×
32.77
21.56
21.38
×
✓
✓
37.83
29.48
25.86
✓
×
✓
42.32
32.74
34.04
✓
✓
×
37.23
30.39
21.34
✓
✓
✓
42.57
33.09
34.62
Table 3: Ablation of the basis construction. Row 1 of (a) is simple PCA (Section 3.2.1 ).
Rank
AGORA
Flex
Dyna
Global, Ka,j=K
36.39
31.91
26.17
Per group, Ka,j=Ka
39.19
32.97
33.95
Per group and block, Ka,j (ours)
42.57
33.09
34.62
Table 4: Allocation of components, PSNR to the host model.
Figure 4: Comparison to alternative distillation: quality against the CPU time of the animation step, lower is better on both axes. Teacher is the host model, our models are labeled with the components per Gaussian and the basis size, and the diamond marks our final model.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Our browser demos on a mobile phone, screenshots from the video on our project page . Left: AGORA + GALA driven by the tracked expressions of a video. Right: five FlexAvatar + GALA avatars, each reconstructed from one photo, around a table (top), and five DynaAvatar + GALA full-body avatars with clothing dynamics (bottom). The overlays show the frame rate and the timing of the animation step.
Table 5: Host models and the settings of our distilled models. Components are listed in total and per attribute group, in the order of the attribute groups. The animation step is timed on a CPU without rendering. For DynaAvatar, both the host model and ours additionally pose the Gaussians by linear blend skinning, which takes 10 ms per frame.
AGORA
FlexAvatar
DynaAvatar
identity descriptor w
latent of the neutral render (512)
embedding of the fitted avatar code (256)
identity tokens and shape parameters of the host encoder (3,082)
animation parameters θ
FLAME expression, jaw and eyelids (55)
articulation code (135)
16 SMPL-X frames × 79 values
hidden layers
256, 512
512, 512
per frame 512, 512, two temporal convolutions of 512, then 1,024, 1,024
activation
leaky ReLU
GELU
SiLU
outputs K
1,016
4,860
2,686
parameters
0.80 M
2.96 M
15.2 M
Appendix
Table 6: Coefficient networks of the distilled models of Table 2 . Hidden layers are listed with their widths. Parameters and multiply-accumulate operations (MAC) are those of one animation step of the network, without the basis. The FlexAvatar identity encoder (6.4 M parameters) runs once per identity and is not included.
AGORA (FFHQ)
FlexAvatar (VFHQ)
DynaAvatar (4D-Dress)
PSNR ∗↑
SSIM ∗↑
FID ↓
ID ↑
AED ↓
PSNR ↑
SSIM ↑
LPIPS ↓
CSIM †↑
AED †↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ∗↑
Simple PCA
28.82
0.894
3.50
0.797
0.792
21.41
0.799
0.119
0.75
0.271
22.59
0.956
0.069
21.27
GALA (ours)
31.18
0.929
3.47
0.761
0.705
22.31
0.818
0.109
0.68
0.245
23.22
0.957
0.066
23.87
Appendix
Table 7: Simple PCA and our full method with predicted coefficients on the benchmarks of Table 2 , at the memory budgets of Table 5 . FlexAvatar is evaluated in self-reenactment on VFHQ, with the enrollment for both models. Best results. ∗ With respect to the host model. † Cross-reenactment. The higher ID and CSIM of simple PCA reflect under-animation, since a static head obtains a CSIM of 0.813 on VFHQ.
student
parameters
size
CPU
FID / LPIPS ↓
AGORA , deformation branch with 512 / 256 / 128 channels at 642 / 1282 / 2562
A1
64 / 32 / 16 channels
0.56 M
2.2 MB
9.4 ms
4.14
A2
192 / 96 channels, no 2562 block
1.03 M
4.1 MB
12.2 ms
3.73
A3
128 / 64 / 32 channels
0.89 M
3.6 MB
16.2 ms
3.71
A4
256 / 128 / 64 channels
1.77 M
7.1 MB
45.1 ms
3.71
A5
448 / 224 / 112 channels, pruned from the host
3.60 M
14.4 MB
124.4 ms
3.60
Appendix
Table 8: Distilled networks of Figure 4 . For AGORA, channels are listed per resolution of the deformation branch. L and d are the number and the width of the transformer blocks, and N the number of attended Gaussians for DynaAvatar. Size is the storage of the student in single precision. CPU times are for the animation step at batch size one with eight threads on an Intel Core Ultra 9 285K, without the skinning of DynaAvatar. Quality is the vertical axis of Figure 4 : FID on FFHQ for AGORA and LPIPS with respect to the ground truth for FlexAvatar and DynaAvatar.
Input per subject
Self reenactment
Cross reenactment
training frames
held-out frames
Views
Frames
Fit
PSNR ↑
LPIPS ↓
PSNR ↑
LPIPS ↓
AED ↓
CSIM ↑
AED ↓
APD ↓
INSTA
RGBAvatar
1
2,790
55 s
33.21
0.0487
30.23
0.0518
0.173
0.524
0.898
0.0499
FlexAvatar †
1
1
∼ 1 min
24.58
0.0858
24.43
0.0675
0.311
0.612
0.700
0.0219
GALA (ours)
1
1
∼ 1 min
23.76
0.1152
23.80
0.0812
0.387
0.639
0.800
0.0234
Appendix
Table 9: Comparison to per-subject blendshape avatars on 6 held-out identities of INSTA and 3 identities of Ava256. Input per subject: views, frames per view and time of the fit or of our enrollment, without face tracking. CSIM uses the input photo as reference. Best held-out and cross results, excluding the host model ( † ).
Input per subject
Self reenactment, held-out frames
Cross reenactment
Views
Frames
Fit
PSNR ↑
LPIPS ↓
AED ↓
CSIM ↑
AED ↓
APD ↓
RGBAvatar
1
127
65 s
20.70
0.1084
0.251
0.451
0.860
0.0645
FlexAvatar †
1
1
∼ 1 min
19.72
0.1363
0.292
0.628
0.634
0.0188
GALA (ours)
1
1
∼ 1 min
19.15
0.1496
0.325
0.666
0.787
0.0212
Appendix
Table 10: Comparison to RGBAvatar on 3 short clips of the VFHQ test set. RGBAvatar is fitted to the first 70% of each clip, FlexAvatar and our model use the first frame, and self reenactment is scored on the last 20%. Best results, excluding the host model ( † ).
Figure 6: Comparison to RGBAvatar on INSTA and Ava256. RGBAvatar is fitted to the video (INSTA) or to the 15-view capture (Ava256) of the subject, whereas the host model FlexAvatar and GALA use one photo. The self rows show held-out frames next to the ground truth, the cross rows are driven by VFHQ videos of other people (left) without ground truth.
Figure 7: Comparison to RGBAvatar on a VFHQ clip of Table 10 . RGBAvatar is fitted to the first 70% of the clip, whereas the host model FlexAvatar and GALA use its first frame. Rows: a training frame of RGBAvatar, a held-out frame, and cross reenactment driven by a VFHQ video of another person (left).
Figure 8: Comparison to GEM on an Ava256 identity. GEM is fitted to the 15-view capture of the subject, whereas the host model FlexAvatar and GALA use one photo. Rows: a training frame of GEM from one of its training cameras, a held-out frame from the camera of the input photo, and cross reenactment driven by a VFHQ video of another person (left).
Figure 9: More difficult examples, with the error scale of Figure 3 . Neutral only is the neutral avatar with all coefficients zero, Projected uses projected coefficients c⋆ and GALA the predicted ones. Rows 1 and 2: held-out AGORA identities at the frame with the largest residual of the host model. Rows 3 and 4: held-out Ava256 subjects in the few-shot and the single-shot setting. Row 5: the frame with the largest error of our model in the DynaAvatar sequence of Figure 3 . In row 4 the coefficient network fails: 23.6 dB PSNR with respect to the host model with predicted coefficients, 40.2 dB with projected ones.
Figure 10: Local motions learned by our blendshapes. The first column highlights the region affected by the visualized blendshapes, and the other columns show two states of their motion for three identities: closing of an eye (AGORA), closing of the mouth (FlexAvatar) and lifting of the front part of an open garment (DynaAvatar).
Figure 11: Further local motions, in the layout of Figure 10 : a smile (AGORA), lowering of a brow together with the upper eyelid (FlexAvatar, cropped to the brow) and the opposite front part of the garment (DynaAvatar).
The generation of high-fidelity, animatable 3D human avatars remains a core challenge in computer graphics and vision, with applications in VR, telepresence, and entertainment. Existing approaches based on implicit representations like NeRFs suffer from slow rendering and dynamic inconsistencies, while 3D Gaussian Splatting (3DGS) methods are typically limited to static head generation, lacking dynamic control. We bridge this gap by introducing AGORA, a novel framework that extends 3DGS within a generative adversarial network to produce animatable avatars. Our key contribution is a lightweight, FLAME-conditioned deformation branch that predicts per-Gaussian residuals, enabling identity-preserving, fine-grained expression control while allowing real-time inference. Identity is further preserved through spatial shape conditioning of the identity branch, and expression fidelity is enforced via a dual-discriminator training scheme leveraging synthetic renderings of the parametric mesh. AGORA generates avatars that are not only visually realistic but also precisely controllable. Quantitatively, we outperform state-of-the-art NeRF-based methods on expression accuracy while rendering at 250 FPS on a single GPU and, notably, at ∼9 FPS under CPU-only inference -- to our knowledge the first demonstration of CPU-only animatable 3DGS avatar synthesis. This work represents a significant step toward practical, high-performance digital humans. Project website: https://ramazan793.github.io/AGORA/
Ramazan Fazylov, Sergey Zagoruyko, Aleksandr Parkin +2
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE · Polynome AI, Dubai, UAE · MTS AI, Moscow, Russia
We propose a method to reconstruct high-fidelity human avatars from multi-view video that can run on mobile devices. Many works can model high-quality Gaussian-based full-body avatars from multi-view video. However, these methods require heavy computation to obtain pose-dependent appearance, making deployment on mobile devices very difficult. Recent methods distill from pretrained models and model pose-dependent nonlinear Gaussian attributes by linearly combining global pose features with blendshapes. Although they can run on mobile devices, they suffer some loss of detail. We observe that nearby Gaussians are often highly correlated within a local region of the body, and can be linearly modeled with less error. Therefore, we use local linear blendshapes in small body parts to capture global nonlinear changes of Gaussian attributes. To further reduce computation and model size, we propose to remove blendshapes for Gaussians whose attributes change little, yielding a minimal blendshape representation. Our method is an end-to-end training method without a pretrained model. To make it run on multiple devices, we implement our method using WebGPU. Experiments show that our method can render high-quality human avatars with better details, and can reach 120 FPS at 2K resolution on mobile devices.
Youyi Zhan, He Wang, Tianjia Shao +1
State Key Lab of CAD&CG, Zhejiang University · University College London
Modeling dynamic facial expressions using 3D Gaussian representations remains challenging due to their unstructured nature. Conventional Gaussian avatar pipelines require extensive multiview and sequential expression data, limiting scalability and accessibility. In this work, we introduce Self-Adaptive Gaussian Expression (SAGE), a framework for self-learning expression-induced Gaussian deformations that enables high-fidelity, animatable avatars from minimal input data. Our method jointly optimizes 2D Gaussian surfels and a Signed Distance Field (SDF) to enforce compact, surface-aligned Gaussian distributions, while a self-supervised expression learning phase replaces long training sequences with geometric and appearance consistency constraints. This design allows flexible deployment across multiple reconstruction regimes: in the multiview setting, only a single frame (timestep) is required instead of thousands; in the monocular setting, only head rotations are needed without expression sequences; and in the one-shot setting, no pretraining or priors are necessary. Experiments demonstrate that our approach achieves reconstruction and animation quality comparable to state-of-the-art methods, while reducing data requirements by several orders of magnitude. Our results highlight the potential of self-supervised Gaussian deformation learning as a step toward accessible, data-efficient avatar creation.