FaceKit: a Toolkit for Interpretable Facial Phenotyping, Synthetic Image Generation and Privacy Analysis in Rare Diseases
Authors: Hongzhuo Chen, Zhanliang Wang, Florent Pollet, Mian Umair Ahsan, Joshua Bie, Tzung-Chien Hsieh, Peter Krawitz, Cong Liu, +4 more
Organizations: Raymond G. Perelman Center for Cellular and Molecular Therapeutics, Children’s Hospital of Philadelphia, Philadelphia, PA, USA · Applied Mathematics and Computational Science Graduate Program, University of Pennsylvania, Philadelphia, PA, USA · Department of Biomedical Informatics, Columbia University, New York, NY, USA · New York Genome Center, New York, NY, USA · Havard Westlake School, Studio City, CA, USA · Institute for Genomic Statistics and Bioinformatics, University Hospital Bonn, Rheinische Friedrich-Wilhelms-Universitat Bonn, Bonn, Germany · Department of Pediatrics, Boston Children’s Hospital & Harvard Medical School, Boston, MA, USA · Department of Computer Science, Columbia University, New York, NY, USA · Department of Pathology and Laboratory Medicine, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA
Many rare genetic diseases are associated with recognizable craniofacial features. However, traditional approaches for describing facial morphology rely largely on qualitative clinical observation and free-text descriptions, which are often subjective, non-standardized, and difficult to reproduce across observers and institutions. Although the Human Phenotype Ontology (HPO) provides controlled terms for describing facial features, these terms are typically categorical rather than quantitative and may vary depending on examiner experience and interpretation. Here, we present FaceKit, a computational framework for quantitative facial phenotyping from frontal facial photographs. FaceKit extracts standardized measurements of facial landmarks and derived 120 morphological features, then reports feature-level z-scores representing deviation from population reference distributions. The reference distributions are built from the FairFace dataset spanning diverse ancestral groups. We evaluated FaceKit on a curated subset of the GestaltMatcher Database covering 50 rare-disease cohorts. In addition to quantitative facial analysis, FaceKit includes synthetic facial image generation to support rare disease model development and data augmentation. We also performed privacy evaluation to assess whether synthetic images reveal identifiable information from real patient photographs and could compromise patient privacy. Across disease case studies, FaceKit-derived quantitative measurements captured known facial features associated with rare genetic disorders and provided objective support for clinical phenotyping. Together, these results establish FaceKit as a useful tool for quantitative phenotyping, and has the potential to improve rare disease diagnosis, support genotype-phenotype studies, and enable more reproducible clinical characterization across diverse patient populations.
Figures & tables
Figure 1 : Overview of the FaceKit workflow. (A) Facial phenotyping: 478 landmarks are extracted from each photograph and corrected for head pose, 120 geometric measurements are derived, and each is expressed as a z -score relative to a normative reference population. (B) Synthetic image generation: after optional preparation with DDColor and GFPGAN, the training images are used to fit a StyleGAN3 generator conditioned on the cohort label, from which synthetic faces of each cohort are sampled. (C) Assessment: privacy is evaluated by comparing synthetic-to-training with held-out-to-training nearest-neighbor distances across cohorts of different training size, and phenotype preservation by comparing measurements between real and synthetic cohorts. All facial images shown are synthetic, and the assessment plots are schematic.
Figure 2 : Synthetic facial image generation. Prepared real images of the rare-disease cohorts (left) are used, together with their cohort labels, to train a conditional StyleGAN3 generator. The discriminator applies successive downsampling convolutions with a minibatch standard-deviation layer and separates real from generated images, while the generator maps a latent vector z∼N(0,I) and the embedded cohort label through an eight-layer mapping network into a style vector that modulates the convolutions of the synthesis network. Sampling from the trained generator with a given label yields 256×256 synthetic facial images of that cohort (right). All faces shown in this schematic are synthetic examples; those in the left panel represent the real training inputs.
Figure 3 : A single FaceKit measurement, from image to cohort. ( 3(a) , 3(b) ) The inter-pupillary distance normalized by bizygomatic width, drawn on the images it is computed from: the coloured line joins the two iris centers (landmarks 468 and 473) and the grey line joins the two zygomatic points (landmarks 234 and 454); the measurement is the ratio of their lengths. ( 3(a) ) A synthetic face generated for Crouzon syndrome, the cohort with the highest mean value. ( 3(b) ) A synthetic face generated for Costello syndrome, a cohort at the center of the same distribution, shown for contrast. Both faces are sampled from the cohort-conditioned generator and depict no real patient. ( 3(c) ) Mean normalized inter-pupillary distance per disease cohort. Bars give the cohort mean with its standard error; the dashed line marks the mean over the remaining cohorts. Of the 50 cohorts, the four highest, the three lowest and Costello syndrome are shown, and the rest are omitted as indicated. Values are computed on images passing the frontal-pose filter.
Method
Depth Z
∣r∣ pitch
∣r∣ yaw
ICC
AUC
A
none (roll only)
0.425
0.011
0.370
0.703
B
0 (flat face)
0.422
0.010
0.377
0.703
C
canonical-face table
0.303
0.011
0.391
0.714
D
per-image z
0.280
0.023
0.322
0.714
Table 1 : Four pose-correction methods on 4,559 GMDB images. Correlations with pitch and yaw and the intraclass correlation (ICC) are means over the 13 measurements with independent validity evidence; AUC is the mean over the six term–measurement pairs that passed both validity contrasts. Lower is better for the pose correlations, higher for ICC and AUC.
Figure 4 : Why the depth term is the whole problem. ( 4(a) ) Orthographic geometry of a rotated face, viewed from above. A landmark lies at X along the face plane and Z along the depth direction. When the head turns by θ , its horizontal position in the image is the sum of: Xcosθ , which shrinks the face horizontally, and Zsinθ , which pushes the landmark sideways in proportion to how deep it sits. ( 4(b) ) Synthetic round trip with the true depth known exactly, giving the residual landmark error as a percentage of bizygomatic width for the uncorrected, flat-face and canonical-depth methods at four head poses. The flat-face method does not improve on leaving the pose uncorrected.
Figure 5 : What the canonical-depth correction moves on a turned face. A synthetic face generated for the Cornelia de Lange syndrome cohort, turned 20∘ in yaw; it depicts no real patient. Left: the depth assigned to each landmark, one constant per landmark from the canonical table and identical for every image; the per-image depth returned by the detector is not read. Right: the displacement the correction applies, drawn at twice its true length, with a median of 2.4% of bizygomatic width and a 95th percentile of 7.5%. The displacement is the Zsinθ term and is therefore near zero wherever the face is flat and largest across the nose and the malar region. It acts on the measurements that depend on depth, changing face_asymmetry by +53% and eye_fissure_slant_mean by −14% on this image, with a median change of 6.1% over the measurements, and leaves the face-plane ratios alone: the inter-pupillary distance over bizygomatic width changes by −0.08% . The panels are generated by manuscript/make_posecorr_mechanism.py , which recomputes every number quoted here.
Figure 6 : The measurements carry disease information outside the development database, and the resolution at which they stop being comparable. ( 6(a) ) Disease separation in RDFace, expressed as the number of standard deviations by which the observed statistic (mean within-condition minus mean between-condition distance in the standardized 120-dimensional space) falls below a permutation null in which the condition labels are shuffled only within strata of image resolution and colour, so that acquisition similarity cannot produce the effect. The controls are cumulative from top to bottom: perceptual near-duplicates removed, then images that a face-recognition check assigns to a patient already present, then a ladder of resolution gates. Every row reaches the floor of the 2,000-permutation test ( p=5.0×10−4 ). The separation weakens as the gates remove images but does not disappear, which is what distinguishes it from an artifact of repeated patients or of the low-resolution tail. ( 6(b) ) Paired within-image drift: every control image is measured at its native 448 pixels and again after downsampling, and the change is expressed in reference standard deviations. The heavy line is the median over the 120 measurements with the interquartile range shaded; the three thin lines are the measurements displaced most at 128 pixels. The catalogue is stable to about 128 pixels, where the median shift is 0.06 reference SD and 36 of the 120 measurements exceed 0.1, and degrades quickly below it. Panels are generated by manuscript/make_sensitivity_panels.py .
Figure 7 : Of the three variables on which the distributed reference differs from the patients, only ancestry changes what the reference concludes. Each point is one HPO term, and the three panels differ only in the variable the reference is conditioned on: ancestry ( 7(a) ), age band ( 7(b) ) and image resolution ( 7(c) ). The abscissa is the sensitivity of the measurement that term is mapped to, measured in the healthy controls: for the first two, the partial η2 of that variable, each adjusted for the other, for gender and for the three pose angles; for the third, the median shift of the measurement when the same image is measured at 224 rather than 448 pixels, a paired quantity that involves no model. The ordinate is the change in the term’s signed effect when the same patients are re-scored against the conditioned reference instead of the pooled one; everything except the reference is held fixed, and all three panels share the ordinate scale. The dashed line splits the terms at the median sensitivity and the horizontal bars are the median gain in each half. A conditioning that removes a real confound should lift the right-hand half and leave the left-hand half alone, which is what ancestry does and what neither age nor resolution does. This contrast, rather than the number of terms that cross a significance threshold, is what separates the three: the headline counts move by six, four and one term respectively, none of which is significant on a catalogue of this size. Panels are generated by manuscript/make_sensitivity_panels.py .
Figure 8 : The 22q11.2 deletion phenotype deviates from controls by the same proportion in every ancestry group. ( 8(a) ) Ancestry-by-cohort interaction q for every measurement significant on at least one scale, fitted on the raw measurement and on its logarithm with the identical model and the identical rows. The correction family is the 97 strictly positive measurements, so the two arms are directly comparable. The four lip measurements cross from below to above q=0.05 once the scale is changed, which is what an effect that is multiplicative rather than additive does; canthal_to_pupillary_ratio barely moves and is the only measurement significant on both scales. ( 8(b) ) The same measurements as a ratio of the 22q11.2 mean to the control mean within each ancestry group, with 95% bootstrap intervals over 2,000 resamples. The lip measurements sit above 1 in all three groups and differ only in how far above, on a baseline that is itself largest in the African controls; only canthal_to_pupillary_ratio crosses 1, changing direction between groups. Panels are generated by manuscript/make_sensitivity_panels.py .
Figure 9 : Synthetic cohorts reproduce the geometric phenotype of the cohorts they were trained on. All panels cover the 1,013 syndrome–measurement pairs obtained by restricting each of the 49 mapped syndromes to the measurements its HPO annotations map to. ( 9(a) ) Effect of the focal cohort against the remaining cohorts pooled, measured in the synthetic images against the same effect measured in the real images; the dashed line is equality. Pairs whose effect disagrees in sign between the two are highlighted, and lie almost entirely near the origin, where the real effect is itself close to zero. The median ρ is taken over cohorts with at least three measurements. ( 9(b) ) Within each syndrome, the absolute Cohen’s d between the synthetic and the real distribution of a measurement, grouped by the number of training patients per cohort. Boxes give the median and interquartile range, whiskers extend to 1.5× the interquartile range, and individual pairs are drawn behind them. ( 9(c) ) Ratio of the variance of a measurement in the synthetic images to its variance in the real images, by facial region and ordered by median. Values below one indicate a synthetic cohort narrower than the real cohort; three ratios above 3, each caused by a single synthetic image with an outlying landmark, are drawn at 3.
Identity (% flagged)
Threshold
Appearance
p (%)
ArcFace
AdaFace
LVFace
ArcFace
AdaFace
LVFace
LPIPS (% flagged)
1
0.17
0.13
0.15
0.3239
0.3115
0.5046
1.0
2
0.65
0.78
0.65
0.3860
0.3696
0.5454
2.7
5
2.78
3.39
3.01
0.4525
0.4248
0.6076
8.6
10
6.09
7.42
6.01
0.4977
0.4641
0.6427
15.5
20
12.73
16.01
15.40
0.5423
0.5171
0.6936
26.4
Table 2 : Synthetic images flagged against a real-image reference rate. Each row fixes an operating point by taking the threshold from the p -th percentile of the distances between held-out real images and the training partition, so that p is the false acceptance rate the threshold produces among real images of the same cohort; entries give the percentage of the 4,746 synthetic images with a detected face (4,785 for LPIPS) falling below that threshold. Thresholds are cosine distances for the three recognition models and LPIPS distances for the perceptual metric. Values at or below p indicate no excess over the reference.
Figure 10 : Privacy assessment of the audited generator. ( 10(a) ) Distribution of ArcFace cosine distances between images of one patient and between images of different patients, with the overlap between the two densities shaded. ( 10(b) ) The same distributions measured with LPIPS, which overlap almost completely and therefore carry little identity information. Panels ( 10(a) ) and ( 10(b) ) rest on different numbers of same-person pairs, 309 against 3,674, because the half-year restriction on the interval between two images of one patient applies to the recognition models only. ( 10(c) ) Percentage of synthetic images falling below a threshold set at the p -th percentile of the held-out distances, for the three recognition models; the dashed line is the rate the same threshold produces among real held-out images. ( 10(d) ) The same comparison measured with LPIPS, where the synthetic images are flagged above the held-out reference at every operating point from 2% upward.
AI-assisted facial phenotyping supports rare genetic disorder prioritization by retrieving visually similar diagnosed cases from facial image reference databases such as the GestaltMatcher Database (GMDB). Existing GestaltMatcher-based retrieval frameworks compare each test image with individual gallery images in a facial phenotype embedding space. However, this pointwise formulation does not fully exploit available evidence, because patients may have multiple images and disorders may be represented by multiple diagnosed gallery patients. We propose an inference-time multi-level evidence aggregation framework that improves facial phenotype retrieval without modifying the underlying GestaltMatcher-Arc encoder. The framework combines embedding-level patient aggregation of multiple images from the same individual, patient-weighted disorder centroids, and hybrid individual-centroid scoring to integrate test-patient observations, disorder-level gallery evidence, and local nearest-neighbor evidence. We evaluated the approach on GMDB v1.1.4 across disorders represented during training (GMDB-Freq), unseen disorders (GMDB-Rare), and multi-image patient subsets, using a unified gallery containing both GMDB-Freq and GMDB-Rare disorders. Multi-level evidence aggregation improved mean per-disorder top-N retrieval accuracy across all evaluation subsets. Top-1 accuracy increased from 38.52% to 48.82% on GMDB-Freq and from 19.38% to 23.79% on GMDB-Rare. On multi-image subsets, top-1 accuracy increased from 46.12% to 60.94% on GMDB-Multi-Freq and from 18.54% to 26.71% on GMDB-Multi-Rare. These findings show that inference-time aggregation can improve next-generation facial phenotype retrieval without retraining the encoder, supporting a shift from isolated single-image matching toward multi-level aggregation of patient and disorder evidence for rare-disorder prioritization.
Alexander Hustinx, Carolin Kaffiné, Behnam Javanmardi +2
Institute for Genomic Statistics and Bioinformatics, University Hospital Bonn, Bonn, Germany
Individuals with suspected rare genetic disorders often undergo multiple clinical evaluations, imaging studies, laboratory tests, and genetic tests over a prolonged period of time, a process commonly described as the diagnostic odyssey. Addressing this odyssey has substantial clinical, psychosocial, and economic benefits. Many rare genetic diseases have distinctive facial features that artificial intelligence algorithms can use to facilitate clinical diagnosis, to prioritize candidate diseases for further laboratory or genetic testing, and to support the phenotype-driven reinterpretation of genome or exome sequencing data. Existing methods that use frontal facial photographs were built on conventional convolutional neural networks, rely exclusively on facial images, and cannot capture non-facial phenotypic traits or demographic information that are essential for accurate diagnosis. Here we introduce GestaltMML, a multimodal machine learning approach based solely on the Transformer architecture. It integrates facial images, demographic information (age, sex, ethnicity), and clinical notes (optionally a list of Human Phenotype Ontology terms) to improve prediction accuracy. We evaluate GestaltMML on 528 diseases from the GestaltMatcher Database and on several in-house and published cohorts, including Beckwith-Wiedemann syndrome, Sotos syndrome, NAA10-related neurodevelopmental syndrome, Cornelia de Lange syndrome, and KBG syndrome. GestaltMML improves on the state-of-the-art image-only ensembled model, narrows the diagnostic accuracy gap for patients from under-represented ancestries, and clarifies when multimodal fusion is beneficial and when image-only inference is preferable. The results suggest that GestaltMML can greatly narrow the candidate diagnoses of rare diseases and may facilitate the reinterpretation of sequencing data.
Da Wu, Zhanliang Wang, Hongzhuo Chen +12
Raymond G. Perelman Center for Cellular and Molecular Therapeutics, Children’s Hospital of Philadelphia, Philadelphia, PA 19104, USA · Department of Mathematics, University of Pennsylvania, Philadelphia, PA 19104, USA · Department of Biomedical Informatics, Columbia University Irving Medical Center, New York, NY 10032, USA +8
Children with rare genetic diseases often exhibit distinctive facial phenotypes, yet developing computer vision systems for early diagnosis remains challenging due to extreme data scarcity, privacy constraints, and limited data sharing in pediatric settings. These challenges not only hinder automated diagnosis but also restrict the availability of visual resources for clinical genetic counseling. While prior work has shown that synthetic data can augment real datasets and preserve phenotype-level semantics, it remains unclear whether synthetic data alone is sufficient for learning in ultra-low-resource pediatric settings. In this work, we study the synthetic-only regime for pediatric rare disease recognition. Under a controlled experimental setup, models are trained exclusively on phenotype-aware synthetic facial images at increasing scales. We find that synthetic-only training achieves performance comparable to real-data-only baselines at sufficient scale across multiple backbones, suggesting that high-fidelity synthetic data can approximate clinically meaningful distributions. These findings together further enable the use of synthetic pediatric facial images as privacy-preserving resources for genetic education and counseling, supporting clinician training and patient communication. Our results highlight the potential of computer vision to improve data efficiency and expand accessible visual tools in children's healthcare.