Mesh-agnostic facial animation retargeting transfers expressions across meshes with different structures, but preserving facial motion without surface artifacts remains challenging. To address this, we present PDB, Point-Based Deformation Blending for facial animation retargeting. PDB predicts a compact set of deformed control points from a source neutral-expression pair and blending weights from the target neutral mesh. The weights are computed once per target and reused across frames, while the control points vary with each source expression. ReLU enforces non-negative weights and permits exact zeros, followed by row-wise normalization. The target mesh is reconstructed directly by multiplying the weights and control points, without a predefined cage, precomputed coordinates, a learned per-element deformation decoder, or a global reconstruction solve. Trained only with self-retargeting reconstruction supervision, PDB supports cross-identity transfer without paired cross-identity training expressions. Experiments demonstrate accurate retargeting, fast inference, and localized support in the learned weights. Joint evaluation of expression accuracy and local surface preservation shows reduced surface artifacts relative to the evaluated dense displacement method while retaining the intended motion. Perceptual evaluations further support expression fidelity and visual quality in both self- and cross-retargeting.
Figures & tables
Figure 1. PDB retargets facial expressions by combining deformed control points predicted from the source expression with localized weights predicted from the target shape, enabling direct retargeting across diverse identities. The weights are visualized by assigning a random color to each control point and blending according to the predicted weight. (face models from BIWI ( Fanelli et al., 2013 ) , ICT ( Li et al., 2020 ) , COMA ( Ranjan et al., 2018 ) and Multiface ( Wuu et al., 2022 ) )
Figure 2. Comparison of our method with previous approaches. Facial retargeting results from (a) per-triangle Jacobian prediction method ( Qin et al., 2023 ) , (b) per-vertex displacement prediction method ( Cha et al., 2025 ) , and (c) ours using a high-resolution mesh from the BIWI ( Fanelli et al., 2013 ) dataset (number of vertices: 23,370 ). The yellow arrow indicates the artifacts.
Figure 3. Method overview for cross-identity retargeting at inference.
Figure 4. Illustration of the feature transform block (b) compared with the spatial transform in PointNet ( Qi et al., 2017 ) (a).
Figure 5. Visualization of the hat function in 1D points (a), values on random sampled 3D points (b), and the obtained masks Min (c), and Mout (d). r0 and r1 are indicated with red dashed lines and dotted lines in (a), respectively.
Design variants
MSE in Min↓(×10−4mm2)
MSE in Mout↓(×10−4mm2)
COMA
BIWI
Multiface
Average
COMA
BIWI
Multiface
Average
Ours †
1.5242
0.5278
1.5461
1.1994
0.1174
2.0171 ×10−7
0.0855
0.0676
(a) Activation variants
No activation
1.7143
0.5156
1.5636
1.2645
0.4973
1.6166 ×10−7
0.1015
0.1996
w/ Softplus
10.2799
2.9841
11.5780
8.2807
4.2061
4.3432 ×10−7
4.7640
2.9900
w/ ELU
1.8608
0.5101
1.6036
1.3248
0.3914
1.7556 ×10−7
1.0473
0.4796
Table 1. Quantitative comparison of design variations. † indicates the base design setting with K=512 . The best score is indicated in Bold .
Figure 6. Visual comparison between the base setting and design variants using Multiface (upper) and COMA (lower). The incurred deformation on the face is colored using a yellow-orange-red color map (YlOrRd). The ground-truth data includes global translation, which produces visible deformation in the neck region.
Figure 7. Visualization of predicted weights. Three cage vertices are randomly selected to visualize the weights predicted by each variant on the COMA test dataset.
Figure 8. Visual comparison of retargeting results for the Multiface test dataset produced by our method and comparative methods. The facial expressions from the source mesh are retargeted to itself, and the target meshes with different shapes and mesh structures. V and T indicate the number of vertices and triangles of the mesh, respectively.
Figure 9. Visualization of weights predicted from the model trained with K=512 . The predicted weights correspond to the 413th control point (a) and the 261st control point (b) on test meshes from BIWI ( Fanelli et al., 2013 ) , ICT ( Li et al., 2020 ) , Multiface ( Wuu et al., 2022 ) , and COMA ( Ranjan et al., 2018 ) .
Figure 10. Visual comparison of retargeting results for the ICT test dataset produced by our method and comparative methods. The facial expressions from the source mesh are retargeted to itself and the target meshes with different shapes and mesh structures. V and T indicate the number of vertices and triangles of the mesh, respectively.
Method
Self-retargeting
Cyclic-retargeting
MSE in Min↓(×10−4mm2)
MSE in Min↓(×10−4mm2)
ICT
MF
Average
ICT → MF → ICT
MF → ICT → MF
Average
NC
64.7652
11.6344
38.1998
89.0935
53.2029
71.1482
NFR
1.9921
2.9145
2.4533
1.0957
4.2026
2.6492
NFS
0.6161
2.5024
1.5592
1.2264
4.1020
2.6642
PDB
0.5731
0.8392
0.7061
0.9101
3.0792
1.9947
Table 2. Quantitative comparison with previous methods on self-retargeting and cyclic-retargeting. Cyclic-retargeting employs ICT → MF → ICT and MF → ICT → MF cycles, in which expressions are transferred from one dataset to a test mesh in the other dataset and then mapped back to the original mesh. The best result is shown in bold . ICT and MF denote ICT-FaceKit and Multiface dataset, respectively.
Figure 11. Magnified results on the ICT test dataset produced by our method and comparative methods. The meshes were rendered with flat shading to visualize the deformed surface.
Method
MSE in Min↓(×10−6mm2)
NC
511.031
NFR
100.984
NFS
8.148
PDB
8.779
Table 3. Neutral reconstruction error on ICT-FaceKit and Multiface.
Figure 12. User study results for expression similarity. Participants selected the method that best preserved the source motion for each clip.
Figure 13. User study results for visual quality. Participants rated the naturalness and absence of visible artifacts using a 1–5 Likert scale.
Method
MSE in Min↓(×10−4mm2)
ICT
Multiface
Average
NC
5927.2172
178.2116
3052.7144
NFR
1.1284
1.2206
1.1745
NFS
1.6646
3.9028
2.7837
PDB
1.2771
1.8993
1.5882
Table 4. Quantitative comparison of local surface structure preservation, measured by the MSE between Laplacian coordinates within the inner facial region. The best result is shown in bold , and the second-best result is underlined .
Data
Frames
Time (min:sec)
NFR
NFS
PDB
ICT (capture #1)
183
00:25
00:17
00:01
ICT (capture #2)
1,147
01:19
00:34
00:06
Multiface (test data)
11,649
17:32
02:29
00:50
Table 5. Runtime comparison for retargeting complete performance sequences. Total time includes one-time target-specific preprocessing and processing all frames.
Figure 14. Generalization to unseen stylized target meshes. PDB retargets source expressions to target meshes with substantially different facial geometry and stylized proportions.
Control points K
Self-retargeting
MSE in Min↓(×10−4mm2)
ICT
MF
Average
K =512
0.5731
0.8392
0.7061
K =256
0.5820
0.8025
0.6922
Table 6. Quantitative comparison with different numbers of control points K on self-retargeting. The best result is shown in bold .
Figure 15. Expression controlled by blendshape coefficients using the model with latent aligned to the ICT blendshape. (Face meshes from left to right: FLAME ( Li et al., 2017 ) , BIWI ( Fanelli et al., 2013 ) , and DT ( Sumner and Popović, 2004 ) .)
Figure 16. Application on GaussianAvatar ( Qian et al., 2024 ) . The deformation applied to the underlying mesh (a) and the corresponding rendered results (b).
We present RegHead, a framework for constructing semantic blendshape sets for animatable non-humanoid head avatars. With a fixed expression vocabulary, semantic blendshapes provide a low-dimensional and interpretable animation interface and support cross-identity retargeting. Building such blendshape sets remains expensive because (i) expression-consistent supervision is scarce, (ii) generated 4D assets typically lack correspondence, and (iii) facial motion is highly localized. We propose (1) a large-scale dataset of non-humanoid identities paired with a shared expression vocabulary, obtained by expanding a small artist-rigged library via fine-tuned image editing; (2) a dense stochastic anchor motion representation tailored to localized facial deformations; and (3) a fast feed-forward registration model that converts unregistered expression meshes into a corresponded blendshape basis by predicting anchor-based deformations from the neutral shape. Experiments show that our approach produces higher-fidelity expression meshes than baselines, while running orders of magnitude faster than optimization. We further demonstrate real-time retargeting from human face tracking signals to non-humanoid characters, capturing both head pose and localized facial motions. Our project page is available at https://snap-research.github.io/RegHead/.
Automatic facial rigging across heterogeneous mesh topologies remains challenging because high-quality expression supervision is often tied to canonical templates, while deformation transfer to arbitrary meshes can introduce geometric artifacts and correspondence errors. We present TopoRig, a topology-agnostic facial rigging framework that predicts FACS-conditioned deformations directly on input mesh vertices while preserving the original topology. Starting from the ICT FaceKit expression model, we construct complementary supervision from accurate but template-biased common-topology rigs, topology-diverse but noisier transferred rigs, and targeted image-based cues for controls poorly captured by geometric transfer. TopoRig combines local surface geometry, landmark-relative semantic features, global shape context, and FACS controls to predict per-vertex displacements. We train on 3,496 generated identities using 45 non-gaze expression controls from the 53-control ICT FaceKit vocabulary. On held-out identities and unseen mesh topologies, TopoRig more faithfully reproduces the reference expression space than prior neural facial-rigging methods, while qualitative results show consistent localized deformations across diverse character geometries. Ablations demonstrate that semantic landmark features and complementary supervision improve cross-identity and cross-topology generalization. Overall, TopoRig amortizes heterogeneous and imperfect expression supervision into a single topology-preserving deformation model.
Andrew Fleet, Soroush Mehraban, Vida Adeli +2
1Queen’s University · 2Pickford AI · University of Toronto +1
Cross-structural motion retargeting aims to transfer motion between different skeletal topologies. Despite recent progress, existing state-of-the-art models struggle with reliability in zero-shot settings, i.e. skeletons with different topologies which were unseen during training, and recent Transformer-based attempts have failed to outperform specialized geometric methods. We bridge this gap with a Transformer Autoencoder that learns a topology- and translation-invariant latent space. Our core contribution is a learnable flattening of skeletal graphs that captures both local dependencies and global structure. Unlike the standard transformer architecture, which adds positional information to token content, we integrate graph-based positional encodings multiplicatively, a design choice that follows directly from our flattening formulation. The resulting model handles diverse skeletal topologies within a single unified architecture and trains in a fully unsupervised manner, requiring no paired retargeting data. Ablation studies show, that the graph encodings, multiplicative formulation, and Transformer backbone is critical for the performance. In zero-shot evaluations, our method reduces global joint position error by 43−47% over current benchmarks. A user study (n=37), including expert animators, further ranks our approach highest in motion alignment and physical plausibility (p<0.05). These results demonstrate that our model design is key to making transformer architectures effective for motion retargeting, outperforming existing approaches.
Kia-Jüng Yang, Fabian H. Sinz, Paweł A. Pierzchlewicz
Institute of Computer Science, University of Göttingen · Campus Institute Data Science, University Göttingen · Pantomim P.S.A