One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars
Organizations: Mohamed bin Zayed University of Artificial Intelligence · MWS AI
Abstract
3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy neural decoding with a shallow coefficient predictor and a linear blend. To improve fidelity and reduce memory requirements, we propose to construct the basis using block-local PCA under a rendering-aware metric and a memory budget. Our method learns a shallow MLP network to predict blendshape coefficients and applies to various animation architectures without retraining original models. We validate GALA by accelerating the inference of three distinct avatar models for 3D animation of facial expressions and full-bodies with clothing dynamics. Across these models, our distillation generalizes to held-out identities and reduces CPU animation cost by up to three orders of magnitude while preserving most of the rendering quality. Excellent results of our method confirm the shared linear structure of learned avatar representations and enable highly efficient and accurate animation at frame rates reaching up to 60fps on mobile devices. Project page: https://ramazan793.github.io/gala/
Figures & tables
| Photo- real. | Identity- shared | Linear at run time | Where the basis comes from | |
| 3D morphable models ( Li et al., 2017b ) | ✓ | ✓ | hand-built from scans | |
| Per-person blendshapes ( Zielonka et al., 2025 ) | ✓ | ✓ | fitted per subject | |
| Fixed-rig avatars ( He et al., 2025 ) | ✓ | ✓ | ✓ | chosen up front |
| Animatable avatar models ( Yu et al., 2025 ) | ✓ | – | none (a network) | |
| GALA (Ours) | ✓ | ✓ | ✓ | discovered in a trained model |
| FID | AED | AED-jaw | ID | APD | FPS | mFPS | |
| AGORA | 3.17 | 0.682 | 0.021 | 0.75 | 0.025 | 250 ∗ | 1 |
| EG3D | 3.28 | – | – | – | – | – | – |
| GGHEAD | 4.06 | – | – | – | – | – | – |
| Next3D | 3.18 | 0.930 | 0.046 | 0.74 | 0.031 | 15 ∗ | – |
| GAIA | 3.85 | 0.530 | 0.040 | 0.72 | 0.027 | 52 ∗ | – |
| AGORA + GALA (ours) | 3.47 | 0.705 | 0.022 | 0.76 | 0.026 | 1,275 | 60 |
| Per-attribute | Rendering-aware | PSNR to host | |||
| bases | PCA | allocation | AGORA | Flex | Dyna |
| 32.77 | 21.56 | 21.38 | |||
| ✓ | ✓ | 37.83 | 29.48 | 25.86 | |
| ✓ | ✓ | 42.32 | 32.74 | 34.04 | |
| ✓ | ✓ | 37.23 | 30.39 | 21.34 | |
| ✓ | ✓ | ✓ | 42.57 | 33.09 | 34.62 |
| Rank | AGORA | Flex | Dyna |
| Global, | 36.39 | 31.91 | 26.17 |
| Per group, | 39.19 | 32.97 | 33.95 |
| Per group and block, (ours) | 42.57 | 33.09 | 34.62 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| AGORA | FlexAvatar | DynaAvatar | |
| Gaussians | 228,693 | 58,361 | 40,000 |
| attribute groups | position, scale, rotation | position, scale, rotation, opacity, color | position, scale, rotation, opacity, color, skinning weights |
| animation parameters | FLAME expression and jaw | FLAME-derived expression code | current and 15 preceding SMPL-X poses |
| spatial blocks | 32 | 32 | 32 |
| metric , , | 1, 16, 0 | 1, 16, 24 | 1, 64, 0 |
| bytes per value | 2 | 2 | 4 |
| AGORA | FlexAvatar | DynaAvatar | |
| identity descriptor | latent of the neutral render (512) | embedding of the fitted avatar code (256) | identity tokens and shape parameters of the host encoder (3,082) |
| animation parameters | FLAME expression, jaw and eyelids (55) | articulation code (135) | 16 SMPL-X frames 79 values |
| hidden layers | 256, 512 | 512, 512 | per frame 512, 512, two temporal convolutions of 512, then 1,024, 1,024 |
| activation | leaky ReLU | GELU | SiLU |
| outputs | 1,016 | 4,860 | 2,686 |
| parameters | 0.80 M | 2.96 M | 15.2 M |
| AGORA (FFHQ) | FlexAvatar (VFHQ) | DynaAvatar (4D-Dress) | ||||||||||||
| PSNR | SSIM | FID | ID | AED | PSNR | SSIM | LPIPS | CSIM | AED | PSNR | SSIM | LPIPS | PSNR | |
| Simple PCA | 28.82 | 0.894 | 3.50 | 0.797 | 0.792 | 21.41 | 0.799 | 0.119 | 0.75 | 0.271 | 22.59 | 0.956 | 0.069 | 21.27 |
| GALA (ours) | 31.18 | 0.929 | 3.47 | 0.761 | 0.705 | 22.31 | 0.818 | 0.109 | 0.68 | 0.245 | 23.22 | 0.957 | 0.066 | 23.87 |
| student | parameters | size | CPU | FID / LPIPS | |
| AGORA , deformation branch with 512 / 256 / 128 channels at / / | |||||
| A1 | 64 / 32 / 16 channels | 0.56 M | 2.2 MB | 9.4 ms | 4.14 |
| A2 | 192 / 96 channels, no block | 1.03 M | 4.1 MB | 12.2 ms | 3.73 |
| A3 | 128 / 64 / 32 channels | 0.89 M | 3.6 MB | 16.2 ms | 3.71 |
| A4 | 256 / 128 / 64 channels | 1.77 M | 7.1 MB | 45.1 ms | 3.71 |
| A5 | 448 / 224 / 112 channels, pruned from the host | 3.60 M | 14.4 MB | 124.4 ms | 3.60 |
| Input per subject | Self reenactment | Cross reenactment | |||||||||
| training frames | held-out frames | ||||||||||
| Views | Frames | Fit | PSNR | LPIPS | PSNR | LPIPS | AED | CSIM | AED | APD | |
| INSTA | |||||||||||
| RGBAvatar | 1 | 2,790 | 55 s | 33.21 | 0.0487 | 30.23 | 0.0518 | 0.173 | 0.524 | 0.898 | 0.0499 |
| FlexAvatar † | 1 | 1 | 1 min | 24.58 | 0.0858 | 24.43 | 0.0675 | 0.311 | 0.612 | 0.700 | 0.0219 |
| GALA (ours) | 1 | 1 | 1 min | 23.76 | 0.1152 | 23.80 | 0.0812 | 0.387 | 0.639 | 0.800 | 0.0234 |
| Input per subject | Self reenactment, held-out frames | Cross reenactment | |||||||
| Views | Frames | Fit | PSNR | LPIPS | AED | CSIM | AED | APD | |
| RGBAvatar | 1 | 127 | 65 s | 20.70 | 0.1084 | 0.251 | 0.451 | 0.860 | 0.0645 |
| FlexAvatar † | 1 | 1 | 1 min | 19.72 | 0.1363 | 0.292 | 0.628 | 0.634 | 0.0188 |
| GALA (ours) | 1 | 1 | 1 min | 19.15 | 0.1496 | 0.325 | 0.666 | 0.787 | 0.0212 |