DisFace3DNet: Explainable Facial Attractiveness Prediction via 3D Component Disentanglement
Authors: Fenggui Rao, Yan Luximon, Jie Zhang
Organizations: School of Design, The Hong Kong Polytechnic University, Hong Kong, China · School of Design, The Hong Kong Polytechnic University, Hong Kong, China, and the Laboratory for Artificial Intelligence in Design, Hong Kong Science Park, Hong Kong, China · Faculty of Applied Sciences, Macao Polytechnic University, Macau, China
Facial attractiveness prediction usually assigns one overall rating, leaving the roles of shape, appearance, and viewing conditions implicit. We propose DisFace3DNet, which uses 3D component disentanglement to learn seven component reference scores from overall ratings with auxiliary weak semantic supervision, without human-labeled component targets. Designated 3D representations and image cues feed jointly learned routes for identity, skin, hair, light, background, expression, and pose. A constrained fit then combines five static and two signed dynamic scores into the overall rating, exposing each component's numerical contribution and supporting component-specific comparisons across images. On SCUT-FBP5500, DisFace3DNet achieves a Pearson correlation of 0.8904±0.0063 (mean ± standard deviation across five folds) with average human ratings; its component terms reconstruct every held-out prediction to numerical precision. Skin, hair, and facial shape account for the largest component-wise prediction variation. Human evaluation supports the score directions for facial shape, skin, and hair; expression agreement is weaker. DisFace3DNet thus connects overall prediction to quantitative analysis of the facial and contextual cues entering each estimate.
Figures & tables
Fig. 1: Prediction and explanation interfaces. (a) Overall prediction through regression, distribution learning, or ranking [ 1 , 8 , 10 ] . (b) Post-hoc explanation [ 13 , 17 , 18 ] , illustrated by Grad-CAM. (c) DisFace3DNet uses separate component predictors to produce component reference scores, which are combined by weighted fusion into the overall rating. Head renderings illustrate the components; scores are symbolic. E and D denote encoder and decoder.
Fig. 2: Architecture of DisFace3DNet. Component Disentanglement provides designated inputs to the Static and Dynamic Component Encoders, which produce five nonnegative scores and signed expression and pose scores, respectively. Component Score Fusion yields the overall rating and per-image weighted contributions. Demographic Conditioning provides a shared embedding to the component-specific feature-wise linear modulation (FiLM) [ 3 ] layers.
Fig. 3: Two-stage training in DisFace3DNet. Stage 1 jointly trains seven component routes using ground-truth (GT) ratings, scalar knowledge distillation (KD) from FPEM [ 4 ] , and GPT-5.5 weak labels [ 64 ] . The provisional score y~n and distribution head π receive the indicated loss terms. Stage 2 freezes the network and fits fusion coefficients on the standardized static/dynamic outputs through Lfit .
Model
Year
Backbone
Correlation ↑
Error ↓
PCC
SRCC
MAE
RMSE
CNN-based learning
Inception-V3 [ 67 ]
2016
Inception-V3
0.918
0.905
—
—
ResNeXt-50 [ 68 ]
2017
ResNeXt-50
0.911
0.899
—
—
DALDL [ 24 ]
2019
ResNeXt-50
0.920
—
0.200
0.269
R3CNN [ 9 ]
2019
ResNeXt-50
0.914
—
0.212
0.280
TABLE I: Overall attractiveness prediction on SCUT-FBP5500 [ 2 ] , grouped by model category
Fig. 4: DisFace3DNet and FPEM [ 4 ] prediction distributions for 1,000 held-out images. White dots and black bars indicate medians and interquartile ranges. Paired t -test p -values are unadjusted overall and Holm-adjusted across the four groups; ns denotes p≥0.05 .
Fig. 6: 3D identity and skin examples: (a) Asian Male, (b) Asian Female, (c) Caucasian Male, (d) Caucasian Female. Each row shows three lower-score and three higher-score heads. Identity displays neutral geometry; skin uses personal-albedo-guided color and FFHQ-UV detail for display. Labels give raw component scores on [0,1] .
Fig. 7: Component-score examples by benchmark group (columns), with three lower- and three higher-score images per component. Labels show Sk for static components ( [0,1] ), 10Sk for expression, and 100Sk for pose. Displayed expression/pose ranges are [−0.48,+0.80] / [−0.29,+0.65] .
Fig. 8: Component-score distributions for all 5,500 held-out SCUT-FBP5500 images and the four benchmark groups. Static scores lie in [0,1] ; expression and pose use signed axes. Density curves have equal peak heights, and labels report mean ± sample SD. Scores precede fold-specific standardization and fusion.
Component
Overall
AM
AF
CM
CF
Identity
18.22 [17.98, 18.44]
18.32 [17.84, 18.79]
19.86 [19.47, 20.23]
20.96 [20.15, 21.73]
21.96 [21.33, 22.56]
Skin
34.29 [33.96, 34.63]
35.85 [35.24, 36.46]
35.92 [35.36, 36.51]
36.21 [35.14, 37.29]
39.71 [38.81, 40.60]
Hair
20.90 [20.55, 21.25]
22.25 [21.58, 22.92]
24.20 [23.60, 24.79]
25.54 [24.36, 26.67]
24.14 [23.17, 25.09]
Light
6.82 [6.72, 6.92]
6.92 [6.73, 7.12]
7.08 [6.90, 7.25]
6.12 [5.75, 6.51]
4.34 [4.05, 4.63]
Background
10.71 [10.56, 10.87]
6.02 [5.65, 6.38]
3.98 [3.69, 4.28]
0.91 [0.84, 0.99]
0.76 [0.70, 0.84]
Expression
8.45 [8.28, 8.63]
9.99 [9.63, 10.37]
8.37 [8.10, 8.65]
9.51 [8.89, 10.18]
8.52 [8.06, 8.99]
TABLE II: Relative component contribution magnitudes (%) with 95% image-bootstrap intervals conditional on the frozen models. Each column sums to 100% before rounding. Identity denotes facial shape and structure. AM, AF, CM, and CF denote Asian Male, Asian Female, Caucasian Male, and Caucasian Female, respectively
Fig. 9: Component-specific human evaluation with 20 participants. (a) Restyled skin-trial interface with the original stimulus pair and response options; no response is shown. (b) Pairwise agreement. Yellow circles: fraction of decisive responses matching the model’s score ordering. Blue diamonds: mean per-response agreement fraction with other participants on the same component and image pair. Both exclude unsure responses and repeats; diamonds require a decisive peer. Bars: 95% participant-clustered bootstrap intervals. Vertical line: 50% agreement between independent fair binary choices.
Variant
PCC ↑
Δ PCC [95% CI]
Additional metrics or matched diagnostic
A. Component scoring
Without component scoring
0.8931±0.0063
−0.0027 [ −0.0058 , 0.0003 ]
SRCC 0.8781±0.0066 ; MAE 0.2472±0.0055 ; RMSE 0.3229±0.0092
Without dynamic components
0.8891±0.0043
0.0013 [ −0.0012 , 0.0038 ]
SRCC 0.8752±0.0047 ; MAE 0.2534±0.0050 ; RMSE 0.3314±0.0062
Without component-specific inputs
0.8715±0.0079
0.0189 [ 0.0141 , 0.0238 ]
Common UV texture for image branches; auxiliary vectors zeroed
B. Training objectives
Without FPEM [ 4 ] teacher guidance
0.8882±0.0054
0.0022 [ −0.0001 , 0.0044 ]
SRCC 0.8730±0.0052 ; MAE 0.2575±0.0044 ; RMSE 0.3349±0.0058
TABLE III: Five-fold ablations of component scoring, training objectives, and training strategy. Metric summaries are mean ± sample SD across folds. Δ PCC is DisFace3DNet minus the variant (positive favors DisFace3DNet), with marginal 95% paired image-bootstrap CIs
Fig. 10: Illustrative applications of DisFace3DNet. Component reference scores support (a) digital-human and bio-inspired robot face design, (b) beauty-filter assessment for livestreaming, photography, and social media, and (c) hairstyle design, salon consultation, and wig try-on.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. 11: Sensitivity to λsem . Half violins (left) show (a–g) component scores and (h) predicted ratings; white-centered markers and black bars mark medians and interquartile ranges. Curves (right) show (a–g) weak-label Spearman correlations and (h) PCC with mean human ratings. Colors identify seeds; curves connect same-seed results.
Variant
PCC ↑
Δ PCC [95% CI]
Matched diagnostic
Without balance loss
0.8909±0.0069
−0.0005 [ −0.0028 , 0.0017 ]
Static-score scale CV: 0.377±0.029 vs. 0.082±0.011
Without distribution consistency
0.8893±0.0056
0.0011 [ −0.0016 , 0.0037 ]
KL divergence: 0.508±0.002 vs. 0.100±0.005 ; score–distribution MAE: 0.464±0.015 vs. 0.046±0.002 points
Without pairwise ordering
0.8899±0.0055
0.0005 [ −0.0015 , 0.0024 ]
GT-pair concordance: (96.79±0.35)% vs. (96.81±0.32)%
Without group conditioning
0.8885±0.0042
0.0019 [ −0.0003 , 0.0040 ]
Group embedding removed
DisFace3DNet (Ours)
0.8904±0.0063
—
Reference for all comparisons
Appendix
TABLE IV: Auxiliary-design ablations. Intervals follow Section S3.A . Paired diagnostic values list the variant before DisFace3DNet
The perceptual representations supporting our ability to recognize faces remain a computational mystery. Deep neural networks offer mechanistic hypotheses for human face perception, but theoretically distinct models often make indistinguishable representational predictions for randomly sampled faces. To expose diagnostic differences among these hypotheses, we compared six neural network models sharing an architecture but trained on distinct tasks, using face pairs optimized to elicit contrasting model predictions ("controversial" pairs) alongside randomly sampled pairs. We tested model predictions against face-dissimilarity judgments from 864 human participants across stimulus sets differing in realism and pose variation. Models prioritizing high-level, invariant structures (trained via inverse rendering, face identification, or object classification) most robustly matched human judgments. Furthermore, models trained on natural images typically outperformed synthetic-trained counterparts. Together, these findings suggest that human face perception is shaped by mechanisms that infer latent causes of facial appearance, discount nuisance variation, and are tuned by natural image statistics.
Wenxuan Guo, Heiko H. Schütt, Kamila Maria Jozwik +3
Department of Psychology, Columbia University, New York, NY, USA. · Department of Behavioural and Cognitive Sciences, Université du Luxembourg, Esch-sur-Alzette, Luxembourg. · MRC Cognition and Brain Sciences Unit, University of Cambridge, Cambridge, England. +4
Monocular 3D face reconstruction estimates a 3D morphable model (3DMM) representation from a single image, providing geometry-aware expression codes that are useful for facial expression analysis and affect understanding. Despite strong progress, most pipelines are trained with image-level self-supervision and evaluated primarily by geometric fidelity, which does not necessarily maximize the affective utility of the learned expression representation and may encourage intensity-amplifying shortcuts when affect supervision is naively coupled. We propose FIELDS (Face reconstruction with accurate Inference of Expression using Learning with Direct Supervision), a task-driven framework that learns FLAME expression codes for facial expression recognition (FER) under a geometric plausibility constraint. Using hybrid 2D/3D supervision, FIELDS improves affect prediction in both in-domain and external evaluations while maintaining competitive geometric fidelity on held-out and out-of-domain 3D benchmarks.
Chen Ling, Henglin Shi, Hedvig Kjellström
KTH Royal Institute of Technology, Sweden · Linköping University, Sweden
Generating 3D models from face sketches is an active topic of research in Computer Graphics due to its potential to tremendously facilitate the modeling of faces for both professional 3D arists and novices. Motivated by the observation that facial expressions are responsible for significantly altering and shaping the contours in our faces, we combine both expression detection and 3D model generation in our approach. The result is a novel approach to generating 3D models from sketches which relies on three components: Convolutional Neural Networks, a parametric 3D face model (Valley Girl), and Active Snake Contours. For the first time in the literature, CNNs are trained (using our own generated dataset) to detect the expression in the given sketch through detecting the active FACS Action Units. The expression is then duplicated on Valley Girl to obtain a 3D model with a similar expression. Active Snake Contours are then used to find the transforms needed to close the gaps between that model and the given sketch.
Nancy Iskander
Graduate Department of Computer Science University of Toronto