We present GHARP (Real-time Gaussian Head Animation from Large-scale Reconstruction Prior), a method that animates 3D human heads in real time from a few input images of a subject and a driving expression signal. We decouple the problem into an identity stage that builds a representation of the subject's geometry and appearance offline, and an animation stage that predicts expression-dependent residuals on top of it at runtime. This separation offers a favorable trade-off with respect to fidelity, quality and runtime: the identity stage can be expensive while the animation stage runs a lightweight network, optimized for mobile devices. Our method performs animation in a semantically structured latent space of a pretrained reconstruction model, where expression changes remain spatially contained, making residual prediction efficient. This reconstruction prior provides a consistent spatial layout, allowing fusion of multiple input views into a compact, fixed-size canonical Gaussian representation. While this two-stage design improves the runtime-quality trade-off, it still inherits a problem common to all expression-driven avatar methods: expression codes describe only the face and thus omit body pose and clothing position, making these regions underspecified in the input. The animation network faces an ill-posed mapping and resorts to averaging over conflicting body appearances, producing blur and temporal flicker. We address this with a body alignment network that learns to align the person's body in the target image with the input reference images, removing the ambiguity from the training signal. Our method achieves state-of-the-art quality on the Ava-256 benchmark while running up to 13x faster on an A100 GPU with 8x fewer Gaussians.
Figures & tables
Figure 1 : Left: GHARP animates a subject from a set of input images and an expression code. Additionally, GHARP enables retargeting of the same expression to new subjects with state-of-the-art quality. Right: runtime on A100 vs. reconstruction quality on Ava-256 for 1, 2 and 4 input images. GHARP is an order of magnitude faster.
Figure 2 : Architecture. Input images are encoded into per-view latents that are aggregated by FC into a fixed-size canonical Gaussian UV map C and by FD into a deformation latent D , both independent of the number of input views N . At runtime, only the lightweight animation network operates on C , D , and an expression code w to produce the animated Gaussian UV map.
Figure 3 : Canonical latent blending. The canonical Gaussian UV map consistently combines input view details (e.g. teeth, hair) with high fidelity.
Figure 4 : Body Alignment Network (BAN). Top row: BAN takes the body pose from the reference image and the facial expression from the target image, producing a target that is consistent in body appearance with the reference (see cyan boxes). Bottom row: BAN Architecture. Reference and target images are encoded into a shared latent space and blended by predicted weights σ(m) and a correction δa into an aligned latent Aℓ . Decoding yields the aligned Gaussian UV map A , which carries the body from the reference and the face from the target. Face- and body-region image losses and an adversarial loss drive the blending (Supp. Mat. 0.A.5 ).
σbodyt↓
#G ↓
LPIPS ↓
SSIM ↑
AKD ↓
CSIM ↑
PSNR ↑
ms ↓
Avat3r
18.8
671K
0.23
0.69
3.5
0.62
21.7
166
FastGHA
22.9
534K
0.23
0.75
3.7
0.51
21.3
26.2
Ours +BAN
0.7
65K
0.21
0.74
3.1
0.67
21.9
1.92
Ours
6.0
65K
0.20
0.76
3.1
0.67
22.2
1.92
Table 1 : Results on Ava-256 using 4 input views. #G is the number of Gaussians; σbodyt measures body stability, the mean per-pixel standard deviation of the body-view render across expressions (0–255 scale, lower is more stable); runtime in ms on A100 GPU. Ours and Ours +BAN achieve state-of-the-art results across metrics.
A100
iPhone 17 Pro
#I
#G
LPIPS ↓
SSIM ↑
AKD ↓
CSIM ↑
PSNR ↑
ms ↓
FPS ↑
ms ↓
FPS ↑
Avat3r
1
182K
0.26
0.73
9.1
0.44
20.5
43.20
23
–
–
FastGHA
1
150K
0.27
0.76
9.6
0.56
20.3
6.81
147
11.80
85
Ours
1
65K
0.27
0.76
7.9
0.62
20.8
1.92
522
1.45
690
Avat3r
2
365K
0.26
0.74
8.8
0.56
20.8
83.79
12
–
–
FastGHA
2
301K
0.27
0.77
9.3
0.58
20.7
12.97
77
24.36
41
Table 2 : Quality and runtime on Internal10K across varying numbers of input images (#I). #G is the number of Gaussians. Runtime is the animation-stage time; Ours is up to 13× faster than FastGHA on the A100 and 31× on iPhone 17 Pro Neural Engine (ANE). Pixel-aligned metrics for Ours +BAN degrade because the predicted body is aligned with a reference image instead of the ground truth. This does not affect animation (AKD) and identity preservation (CSIM) quality.
Table 3 : Regional image quality (Ava-256, Internal10K) . We report SSIM and PSNR on GT mask-segmented face and body regions at a fixed cropped frontal camera view. Note that these regional metrics do not directly compare to Tab. 1 or Tab. 2 .
Figure 5 : Temporal stability. The colors show the per-pixel temporal standard deviation across a sequence driven by an expression code See zoomed cyan boxes for jitter unrelated to the expression. Avat3r [ 30 ] , FastGHA [ 22 ] , and Ours all exhibit this jitter, whereas Ours +BAN , trained against body-aligned supervision targets (Fig. 4 ), keeps the body stable. The full video is provided in the Supp. Mat.
Figure 6 : Qualitative results with BAN. Ours+BAN cleanly recovers fine body details (earrings, tattoos) without the blur or geometric distortion seen in other approaches.
Figure 7 : Qualitative comparison on Ava-256. Our method produces sharper facial details (e.g. eyes/teeth/wrinkles) and better expression transfer. FastGHA [ 22 ] tends to produce over-smoothed textures and Avat3r [ 30 ] does not preserve identity as well. The boxes mark the zoom-in region shown as an inset in each panel.
Internal10K Dataset
Configuration
PSNR ↑
SSIM ↑
LPIPS ↓
AKD ↓
CSIM ↑
Full model (Ours)
21.6
0.77
0.24
7.4
0.67
w/o canonical representation
21.2
0.76
0.27
7.5
0.65
w/o passthrough loss
21.3
0.77
0.25
7.5
0.66
w/o deformation latent
21.3
0.75
0.25
9.3
0.59
Full model + BAN
19.8
0.73
0.28
7.5
0.66
Table 4 : Design choices ablation. We individually ablate key components of our pipeline. Each w/o row disables one component while keeping all others at default.
Figure 8 : Qualitative evaluation of our design choices. We compare (a) the ground truth (GT) against (b) our full model and ablations removing (c) the canonical representation, (d) the passthrough loss, and (e) the deformation latents. Cyan boxes mark the zoom-in region shown in the bottom-left of each panel.
Figure 9 : Cross-identity retargeting. For each pair, the identity input is shown in the first column, and the expression is driven by the subject in the second column. Right table: Average Expression Distance (AED) metric.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Type
Channels
Input (concat latents)
N×(512+6)=2072
Conv 3×3
2072→512
ResNetBlock ×8
512→512
GN16 → SiLU → Conv 1×1
512→2560
ResNetBlock = GN16 → SiLU → Conv 3×3 , repeated twice, with skip connection.
GN = Group Normalization layer with 16 groups
Appendix
Table 5 : Architecture of FC (canonical fusion network). Resolution is constant at 32×32 .
Type
Channels
Resolution
U-Net Encoder
Input ( Gi or C )
dG=23
2562
Conv 2×2 s2, BN, ReLU
23→64
2562→1282
Conv 3×3 s2
64→128
1282→642
DownBlock
128→128
642→322
DownBlock
128→256
322→162
Appendix
Table 6 : Architecture of FD (deformation encoder). Per-Gaussian U-Net applied independently to each Gi and C .
Type
Channels
Input (concat all maps)
(N+1)×192=960
Conv 3×3
960→512
ResNetBlock ×2
512→512
GN16 → SiLU → Conv 1×1
512→96
ResNetBlock = GN16 → SiLU → Conv 3×3 , repeated twice, with skip connection.
GN = Group Normalization layer with 16 groups
Appendix
Table 7 : Architecture of FD . Consolidation network applied to the concatenated features from all N+1 maps. Resolution is constant at 64×64 .
Type
Channels
Resolution
Input w
51
—
MLP (FC → ReLU → FC)
51→128
—
Tile to spatial
128
642
Concat ( D , tiled w )
96+128→224
642
Conv 1×1
224→128
642
ResBlock ×6
128→128
642
Appendix
Table 8 : Architecture of E (animation network).
Type
Channels
Resolution
Encoder + MLP bottleneck
Input ( Zi )
512
322
DownBlock
512→512
322→162
DownBlock
512→256
162→82
DownBlock
256→256
82→42
Flatten
4096
—
Appendix
Table 9 : Architecture of the passthrough network.
Pretraining
Finetuning
Dataset
Internal10K
Ava-256
Image size ( H×W )
500×375
512×334
UV map size ( HUV×WUV )
256×256
256×256
HeadsUp ( HE , HD )
frozen
finetuned on Ava-256
Trained modules
FC , FD , E
FC , FD , E
Steps
200K
15K
Appendix
Table 10 : Training hyperparameters.
Figure 10 : Qualitative comparison between Avat3r implementation in HeadsUp and in our method in a setting without rigging. Main differences in our Avat3r implementation include DINOv2 loss (which explains sharper detail), as well as relaxing feature extractor choice from DUSt3R to DINOv3 (while keeping Sapiens in place), and training with expression conditioning.
Loss
Weight
L1
1.0
Llpips
10.0
Ladv
0.25
Appendix
Table 11 : BAN loss weights.
σbodyt↓
ΔE76↓
TCE ↓
Avat3r
18.8
10.0
11.7
FastGHA
22.9
12.1
13.4
Ours
6.0
2.8
2.8
Ours +BAN
0.7
0.2
0.2
Appendix
Table 12 : Body Temporal Stability on Ava-256. Per-pixel temporal standard deviation of the body-view render in RGB ( σbodyt ) and CIELAB ( ΔE76 ), and optical-flow drift (TCE, in pixels), computed across each subject’s expressions; lower is more stable. Ours +BAN is the most stable across all three metrics. Best in bold , second best underlined .
Figure 11 : Frozen-body blending vs . BAN. At a mouth opening, a narrow ramp leaves a seam at the neck (top), and a wider ramp changes the chin shape and reduces expression fidelity (bottom). Ours +BAN avoids both.
Avat3r [ 30 ]
FastGHA [ 22 ]
Ours
Device
N
ms
FPS
ms
FPS
ms
FPS
H100
1
21.33
47
2.39
418
1.18
851
2
41.15
24
4.67
214
1.18
851
4
78.44
13
9.14
109
1.18
851
iPhone 17 Pro (GPU)
1
–
–
17.6
57
5.1
196
2
–
–
31.76
31
5.1
196
Appendix
Table 13 : Animation runtime (ms) and FPS. Additional results for animation runtime (ms) and FPS for our method, FastGHA, and Avat3r on different devices across varying numbers of input views ( N ).
Method
#G
A100
iPhone 17 Pro (ANE)
FastGHA [ 22 ] ( N=1 )
150K
6.81
11.80
Ours260
260K
3.3
3.4
Ours
65K
1.92
1.45
Appendix
Table 14 : Animation runtime (ms) with 4× more Gaussians (Ours260).
Animation
Rendering
End-to-end
Avat3r [ 30 ]
166.23
1.91
168.14
FastGHA [ 22 ]
26.25
1.53
27.78
LAM † [ 21 ]
-
-
3.56
Ours
1.92
1.33
3.25
Appendix
Table 15 : End-to-end online latency (ms) on an A100 with N=4 input views. Rendering uses gsplat [ 74 ] ; animation times are from Tab. 2 . † Reported by the authors for their 80K-Gaussian model.
Figure 12 : Impact of the number of input views. Reconstruction quality scales naturally with the number of input images. Even with N=1 , Ours produces sharper reconstruction and preserves identity better compared to prior work.
Method
LPIPS ↓
SSIM ↑
AKD ↓
CSIM ↑
PSNR ↑
Ours1k
0.28
0.76
7.8
0.60
20.7
Ours
0.24
0.77
7.4
0.67
21.6
Appendix
Table 16 : Ablation on weaker priors. Impact of reducing the training data for the HeadsUp prior from 10k (Ours) to 1k (Ours1k) subjects.
Figure 13 : In-the-wild driving image. The bottom row shows images captured with an iPhone. We use ARKit to extract expression codes from these images and drive the avatar from the input image of the same subject (left column).
Figure 14 : More results for cross-identity retargeting. For each pair, the identity input is shown in the first column, and the expression is driven by the subject in the second column.
Method
#G
LPIPS ↓
SSIM ↑
AKD ↓
CSIM ↑
PSNR ↑
UIKA [ 68 ]
134K
0.19
0.79
3.9
0.67
21.3
Ours
65K
0.20
0.76
3.1
0.67
22.2
Appendix
Table 17 : Comparison with UIKA on Ava-256 with 4 input views. Top: reconstruction quality. Bottom: animation runtime. #G is the number of Gaussians. UIKA cannot run on the ANE because of the large LBS matrix multiplications.