How Private is Private? A Comparative Study for Face De-Identification
Organizations: ELLIS Institute Finland, Finland · Center for Machine Vision and Signal Analysis (CMVS), University of Oulu, Finland
Abstract
Face de-identification (FDeID) has emerged as a critical privacy-preserving technology, yet its evaluation remains fundamentally fragmented. Existing protocols rely on inconsistent metrics, heterogeneous datasets, and partial annotation coverage, so methods targeting different utility dimensions, such as landmark versus expression preservation, are reported on different benchmarks under different metrics, rendering cross-method comparison infeasible. We revisit FDeID evaluation from both the data and metric perspectives. On the data side, we introduce UtilFace, a curated, demographically balanced benchmark with high identity diversity, assembled from four large-scale face datasets through identity-aware cleaning, resolution enhancement, and stratified filtering. On the metric side, we propose HiFD, a Hierarchical Face De-identification metric that unifies identity suppression, multi-level utility preservation, and image quality under a single consistency-based paradigm: every component is computed from pretrained estimators' outputs on the original face and its de-identified counterpart, directly quantifying how much identity is suppressed and how much downstream-perceivable utility survives. HiFD organizes facial signals into a three-level utility hierarchy spanning macro cues (L1), micro cues (L2), and imperceptible cues (L3), and aggregates the five resulting components into a single interpretable score via weighted harmonic mean, with configurable application-specific profiles. Using this unified protocol, we conduct a comprehensive comparative study spanning adversarial, GAN-based, and diffusion-based methods, surfacing trade-offs and failure modes that remain invisible under existing protocols. We release the benchmark and evaluation toolkit to foster systematic and reproducible research in privacy-preserving human face analysis.
Figures & tables
| Dataset | Modality | #Images/Videos | #IDs | L1: Macro-Cues | L2: Micro-Cues | L3 | |||||
| Age | Gender | Ethnicity | Macro-Exp. | Landmark | Micro-Exp. | Gaze | rPPG | ||||
| LFW Huang et al. (2008) | Image | 13,233 | 5,749 | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| CelebA Liu et al. (2015) | Image | 202,599 | 10,177 | ✗ | ✓ | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ |
| FFHQ Karras et al. (2019) | Image | 70,000 | – | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ |
| AgeDB Moschoglou et al. (2017) | Image | 16,488 | 568 | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| RFW Wang et al. (2019) | Image | 40,000 | 12,000 | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| Stage | Operation | #Images | #Identities |
| 1. Aggregation | Merge 4 source datasets | M | M |
| 2. Cleaning + Enhancement | Identity-aware filtering; GFPGAN / CodeFormer | M | |
| 3. Quality filter | NIQE | M | |
| 4. Balanced sampling | Hierarchical stratified sampling |
| Method | L1: Macro-cues | L2: Micro-cues | L3: rPPG | HiFD | |||||||||||
| Age | Gen. | Eth. | MExp | Lmk | Gaze | ME | BVP | HR | |||||||
| (a) Adversarial perturbation | |||||||||||||||
| MI-FGSM | 0.134 | 0.360 | 0.975 | 0.962 | 0.916 | 0.951 | 0.834 | 0.928 | 0.936 | 0.799 | 0.868 | 0.471 | 0.478 | 0.474 | 0.343 |
| PGD | 0.113 | 0.362 | 0.978 | 0.968 | 0.935 | 0.964 | 0.843 | 0.938 | 0.940 | 0.817 | 0.878 | 0.577 | 0.555 | 0.566 | 0.321 |
| TI-DIM | 0.500 | 0.342 | 0.880 | 0.826 | 0.656 | 0.739 | 0.684 | 0.757 | 0.921 | 0.722 | 0.821 | 0.424 | 0.425 | 0.424 | 0.509 |
| TIP-IM | 0.496 | 0.446 | 0.927 | 0.882 | 0.701 | 0.794 | 0.728 | 0.806 | 0.944 | 0.812 | 0.878 | 0.643 | 0.605 | 0.624 | 0.607 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | #Identities | #Images | Images / ID |
| WebFace42M Zhu et al. (2021) | M | M | |
| Glint360K An et al. (2021) | K | M | |
| MS1MV2 Deng et al. (2019) | K | M | |
| MegaFace Kemelmacher-Shlizerman et al. (2016) | K | M | |
| Aggregated pool | M | M |
| Constant | Value | Rationale |
| years | Full biological age range. | |
| Threshold beyond which face landmarks are considered unreliable in standard benchmarks. | ||
| Maximum possible angular error between two 3D unit vectors in the forward hemisphere. | ||
| bpm | Clinical relevance threshold; errors above this render HR estimates unusable for downstream tasks. | |
| Matches the NIQE curation threshold used to filter UtilFace . | ||
| Equal weighting of reference-based and no-reference quality components. |
| Method | Component scores | Aggregation scheme (value, rank / 12) | ||||||
| (Arith.) | (Geom.) | (Harm., Ours) | ||||||
| (a) Adversarial perturbation | ||||||||
| MI-FGSM | (8) | (7) | (7) | |||||
| PGD | (5) | (6) | (9) | |||||
| TI-DIM | (6) | (3) | (2) | |||||
| TIP-IM | (3) | (1) | (1) | |||||
| Method | Consistency | Accuracy on | Retention | AUC | Hidden flips |
| (a) Adversarial perturbation | |||||
| MI-FGSM | 0.813 | 0.646 [0.629, 0.662] | 0.830 | 0.976 | 1.7% |
| PGD | 0.860 | 0.694 [0.678, 0.710] | 0.892 | 0.971 | 1.6% |
| TI-DIM | 0.777 | 0.625 [0.607, 0.642] | 0.804 | 0.979 | 0.7% |
| TIP-IM | 0.779 | 0.613 [0.597, 0.630] | 0.788 | 0.983 | 0.8% |
| Chameleon | 0.792 | 0.621 [0.604, 0.638] | 0.798 | 0.974 | 1.1% |
| Method | FAR | FAR | FAR | |
| TI-DIM | 0.500 | 0.859 | 0.935 | 0.971 |
| TIP-IM | 0.496 | 0.856 | 0.930 | 0.967 |
| CIAGAN | 0.450 | 0.839 | 0.959 | 0.987 |
| Chameleon | 0.239 | 0.188 | 0.270 | 0.351 |
| WeakenDiff | 0.222 | 0.000 | 0.001 | 0.008 |
| G 2 Face | 0.210 | 0.002 | 0.010 | 0.033 |
| Method | Gaze | ME | (PURE) | HiFD | |||
| (a) Adversarial perturbation | |||||||
| MI-FGSM | 0.134 [0.133, 0.138] | 0.360 [0.359, 0.360] | 0.937 [0.936, 0.938] | 0.800 [0.782, 0.817] | 0.868 [0.859, 0.877] | 0.474 [0.392, 0.505] | 0.343 [0.334, 0.348] |
| PGD | 0.113 [0.112, 0.117] | 0.362 [0.362, 0.363] | 0.941 [0.939, 0.942] | 0.818 [0.797, 0.836] | 0.879 [0.869, 0.889] | 0.566 [0.485, 0.602] | 0.321 [0.316, 0.328] |
| TI-DIM | 0.500 [0.500, 0.506] | 0.342 [0.341, 0.342] | 0.923 [0.922, 0.924] | 0.715 [0.688, 0.740] | 0.819 [0.806, 0.832] | 0.424 [0.356, 0.444] | 0.509 [0.490, 0.519] |
| TIP-IM | 0.496 [0.496, 0.501] | 0.446 [0.445, 0.447] | 0.944 [0.943, 0.945] | 0.812 [0.790, 0.834] | 0.878 [0.868, 0.889] | 0.624 [0.542, 0.662] | 0.607 [0.593, 0.618] |
| Chameleon | 0.239 [0.235, 0.246] | 0.474 [0.473, 0.477] | 0.965 [0.964, 0.966] | 0.797 [0.777, 0.817] | 0.881 [0.870, 0.892] | 0.799 [0.722, 0.838] | 0.508 [0.503, 0.516] |
| Attribute | Substitution | Sub-score correlation | HiFD Spearman | Rank changes |
| Ethnicity | FairFace (7-class) ViT ethnicity (5-class) | 0 | ||
| Macro-expression | POSTER/RAF-DB POSTER/AffectNet | 0 | ||
| Macro-expression | POSTER/RAF-DB FER ViT | 0 | ||
| Age | MiVOLO ViT age | 0 | ||
| All L1 sub-scores | averaged over 2–3 models each | – | 0 | |
| Axes that already ensemble: recognizers, pairwise Spearman / / ; rPPG estimators, / / . | ||||
| (remainder equal) | (remainder equal) | ||||||||||
| Method | 0.2 | 0.3 | 0.4 | 0.5 | 0.6 | 0.2 | 0.3 | 0.4 | 0.5 | 0.6 | |
| TIP-IM | 0.607 | 0.590 | 0.575 | 0.560 | 0.546 | 0.607 | 0.609 | 0.611 | 0.613 | 0.615 | |
| TI-DIM | 0.509 | 0.508 | 0.507 | 0.506 | 0.504 | 0.509 | 0.497 | 0.485 | 0.474 | 0.463 | |
| Chameleon | 0.508 | 0.445 | 0.396 | 0.357 | 0.325 | 0.508 | 0.532 | 0.558 | 0.588 | 0.621 | |
| CIAGAN | 0.369 | 0.378 | 0.386 | 0.396 | 0.406 | 0.369 | 0.348 | 0.328 | 0.311 | 0.296 | |
| Leader | TIP-IM throughout | TIP-IM | Chameleon | ||||||||
| Method | BRISQUE | TOPIQ | (rank) | (rank) | HiFD [ ] | HiFD [no ] | HiFD |
| (a) Adversarial perturbation | |||||||
| MI-FGSM | 15.7 | 0.633 | 0.738 (2) | 0.360 (10) | 0.380 (6) | 0.339 (7) | 0.343 (7) |
| PGD | 14.9 | 0.648 | 0.749 (1) | 0.362 (9) | 0.354 (7) | 0.313 (9) | 0.321 (9) |
| TI-DIM | 25.3 | 0.575 | 0.661 (4) | 0.342 (11) | 0.595 (2) | 0.580 (2) | 0.509 (2) |
| TIP-IM | 40.5 | 0.670 | 0.632 (7) | 0.446 (5) | 0.660 (1) | 0.667 (1) | 0.607 (1) |
| Chameleon | 40.5 | 0.701 | 0.648 (5) | 0.474 (3) | 0.538 (3) | 0.517 (3) | 0.508 (3) |
| Method | MMPD [95% CI] | PURE | Illumination range | Motion range | MAE (bpm) | HR retention |
| (a) Adversarial perturbation | ||||||
| MI-FGSM | 0.503 [0.491, 0.514] | 0.474 | [0.470, 0.520] | [0.400, 0.689] | 3.3 [2.5, 4.2] | 0.834 |
| PGD | 0.558 [0.546, 0.570] | 0.566 | [0.523, 0.578] | [0.445, 0.755] | 2.2 [1.5, 3.0] | 0.889 |
| TI-DIM | 0.410 [0.398, 0.421] | 0.424 | [0.379, 0.437] | [0.322, 0.585] | 4.1 [3.2, 5.0] | 0.794 |
| TIP-IM | 0.651 [0.640, 0.662] | 0.624 | [0.614, 0.683] | [0.591, 0.767] | 1.7 [1.0, 2.4] | 0.913 |
| Chameleon | 0.709 [0.700, 0.718] | 0.799 | [0.677, 0.723] | [0.677, 0.727] | 0.7 [0.0, 1.3] | 0.966 |
| Condition | Naturalness | Artifacts | Attributes kept | Same person |
| Unmodified original | 3.32 [2.92, 3.70] | 2.08 | – | – |
| (a) Adversarial perturbation | ||||
| MI-FGSM | 3.12 [2.73, 3.53] | 2.08 | 0.966 [0.898, 1.000] | 0.977 [0.932, 1.000] |
| PGD | 3.35 [2.95, 3.71] | 2.02 | 0.932 [0.847, 0.994] | 1.000 [1.000, 1.000] |
| TI-DIM | 1.32 [1.14, 1.52] | 4.52 | 0.881 [0.795, 0.949] | 0.886 [0.773, 0.977] |
| TIP-IM | 1.55 [1.27, 1.86] | 4.29 | 0.858 [0.761, 0.938] | 0.909 [0.795, 1.000] |
| Component | Model | Arch.† | Train data | Source |
| Privacy ensemble | ||||
| (1/3) | ArcFace Deng et al. (2019) | IR-R100 | MS1MV3 | insightface model zoo |
| (2/3) | CosFace Wang et al. (2018) | IR-R50 | Glint360K | insightface model zoo |
| (3/3) | AdaFace Kim et al. (2022) | IR-50 | MS1MV2 | official release |
| L1 macro-cues | ||||
| Age, Gender | MiVOLO Kuprashevich and Tolstykh (2023) | ViT-MiVOLO | UTKFace | official release |