Towards Fast and Disentangled Counterfactuals for Visual Foundation Models
Authors: Sidney Bender, Benedikt Kunz, Ahmed Zeid, Shinichi Nakajima, Klaus-Robert Müller, Marco Morik
Organizations: Berlin Institute for the Foundations of Learning and Data (BIFOLD), Berlin, Germany · Machine Learning Group, Technische Universität Berlin, Berlin, Germany
Foundation models remain vulnerable to spurious correlations and ``Clever Hans'' strategies. Explainable machine learning can find and remove such strategies for classifiers without metadata. For foundation models, no such option exists yet. We propose Disentangled Diffusion Autoencoders (DiDAE). DiDAE wraps a frozen foundation model in a conditional diffusion decoder. A counterfactual is one closed-form edit along a direction of a disentangled dictionary, followed by decoding. The dictionary can be supervised (Procrustes) or unsupervised (Singular Value Decomposition, Sparse Autoencoders). No gradients are needed, so DiDAE is up to 2000 times faster than the state of the art. We evaluate on six datasets, two synthetic and four real-world. In a desiderata-driven benchmark on three of them, its counterfactuals are on par with or better than the state of the art, and they repair downstream classifiers through Counterfactual Knowledge Distillation (CFKD), where they beat metadata-based correction. The same machinery can rank a pretrained dictionary against a trained classifier. It returns the few directions the classifier actually reads, each causally verified by a counterfactual that flips the decision, and repairs the classifier along those a teacher marks spurious. The workflow is plug-and-play in our open-source Peal library we publish alongside the paper. With a public dictionary and a pretrained decoder, all that remains is a cheap linear distillation of the classifier and its own fine-tuning.
Figures & tables
Fig. 1 : Find and remove a Clever Hans feature of a classifier without training a generator or a dictionary. The classifier is a DINOv3 [ 1 ] ViT-L/16 linear head trained on the natural ImageNet freight car vs. passenger car pair; the dictionary is the public 6,144 -atom MSAE [ 2 ] over OpenAI CLIP ViT-L/14 and the decoder our pretrained ImageNet RAE (App. E-C ), both taken as they are. (a) The usual route, browsing the dictionary by highest-activating images, returns the atom that dominates this dataset, #3697 (“trains”), which says nothing about the decision. (b) DiDAE distils the classifier into the encoder space, a cheap closed-form step, and ranks the atoms by counterfactuals that flip the real classifier. Behind the class evidence “bin” rank rail tracks ( 40 verified flips) and graffiti, the one spurious feature Neuhaus et al. [ 3 ] report for freight car; the tracks shortcut is new. Each direction comes as before/after pairs (two chosen per direction here), so a practitioner sees what it means and marks it as class evidence or confounder. The pairs also correct the atoms’ automatic CLIP-Dissect [ 4 ] names: “bin” and “travelling” change the car body and are class evidence, which the words do not suggest, and “tracks” is not rails anywhere in the image but a central view down a receding track. (c) One CFKD iteration on the tracks direction raises average group accuracy over the class × tracks groups from 96.4% to 97.2% , a Gain of 21.6% on an already accurate probe, without touching the backbone.
Fig. 2 : Comparison of traditional gradient-based counterfactuals (a), global counterfactual methods (b) versus the proposed DiDAE approach (c) on a CelebA classifier trained on the “Blond Hair” label. The label is spuriously correlated with “Female”, “Heavy Makeup” and “Attractive”. Gradient-based methods require slow, iterative gradient updates through the diffusion process; global counterfactual methods like diffusion autoencoders or TIME are fast, but only produce one counterfactual per factual often either too weak to be seen clearly (as in the figure) or with entangled changes (e.g., changing hair color and the correlated gender simultaneously). Both are only able to explain classifiers and can not explain the foundation models themselves. In contrast, DiDAE utilizes a frozen foundation model to decompose embeddings zFM into disentangled semantic components. Counterfactuals are generated via a single closed-form linear edit in this semantic space, which moves one component to the bound of its empirical range, followed by decoding via a diffusion decoder. This is fast, can create diverse, disentangled counterfactuals, and can be applied directly to the foundation model representations as well, without the need to explain a specific classifier.
Fig. 3 : (a) Global methods like DAE generate counterfactuals by moving orthogonally to the classifier’s decision boundary (red arrow). This trajectory entangles multiple variables (e.g., hair color, makeup, and gender), altering them simultaneously and yielding a less interpretable result. (b) In contrast, DiDAE disentangles these transformations along distinct semantic axes based on a learned dictionary. By moving independently along these axes (pink arrows), DiDAE produces fine-grained, interpretable counterfactuals (CF1, CF2, CF3) that isolate attribute-specific changes across the classifier boundary.
Dataset
Task
Confounder
Encoder Φ
Decoder
Inversion
Dictionary
Appendix
Square
square intensity
background
ResNet-18
DiffAE
DDPM
Procrustes / SVD
C-A , E-A
CelebA-Blond
Blond_Hair
Male
CLIP ViT-L/14
DiffAE
DDPM
Procrustes
C-A , E-A
Camelyon17
tumor
hospital
PLIP
PathLDM
DDPM
batch top- K SAE
C-A , E-B
Sparse Numbers
Num128
Num713
ResNet-18
DiffAE
DDIM
batch top- K SAE
C-A , E-A
NICO++
crocodile / lizard
context
CLIP ViT-L/14
RAE
DDPM
MSAE
C-A , E-C
ImageNet
freight / passenger car
found by ranking
CLIP ViT-L/14
RAE
DDPM
MSAE
C-A , E-C
TABLE I : The six datasets and the instances of the four DiDAE components on each. The first three rows carry the counterfactual-quality and correction benchmark, the last three the dictionary and ranking experiments (Appendix C-C , Dictionary inspection and Ranking ). We train the ResNet-18 encoders of the synthetic datasets, the batch top- K SAEs and the decoders, fine-tuning PathLDM from its histopathology checkpoint; the other encoders and the MSAE are pretrained. The last column points to the appendix sections with the dataset and decoder details; resolutions, step counts and dictionary sizes are in Table V .
desiderata
sufficiency
understandability
fidelity
efficiency
Dataset
Method
n
(NAFR)
(Diversity)
(Sparsity)
(NA)
(Unbiasedness)
(CF/s)
Gain
Square
DAE
4
0.0 ± 0.0
0.0 ± 0.0
0.0 ± 0.0
0.0 ± 0.0
0.0 ± 0.0
∼ 57.1 ± 2.8
0.0 ± 0.0
DiME
4
8.0 ± 6.7
0.7 ± 0.6
78.6 ± 3.4
32.7 ± 6.1
31.6 ± 31.7
∼ 0.03 ± 0.00
69.1 ± 3.5
ACE
4
1.2 ± 0.4
10.0 ± 16.9
78.3 ± 16.7
58.4 ± 8.4
0.0 ± 0.0
∼ 0.03 ± 0.01
38.8 ± 38.8
FastDiME
4
6.9 ± 5.2
0.9 ± 1.6
87.7 ± 6.9
41.1 ± 21.3
55.3 ± 43.6
∼ 2.4 ± 0.2
66.3 ± 7.4
TABLE II : Counterfactual quality, speed and downstream Gain. Edit-friendly DDPM inversion; desiderata of Appendix A in percent, CF/s in counterfactuals per second on one A100, Gain on the balanced test split. Bold is the best value in a column within a dataset block, underline the second best. Entries are mean ± population std over the n seeds in the third column; n=1 entries are a single seed. Gain is the column we read first (Section V-A ).
Fig. 4 : DiDAE’s two attempts move along different axes; DAE’s coincide. Two counterfactual attempts per method on Square and CelebA, projected onto the causal (x) and confounding (y) axes of the oracle encoder defined in Appendix A ; on Square the decision boundary is exact, on CelebA it is approximated by the oracle’s Male and Blond predictions. DAE does not cross the boundary on Square and mixes both factors in one direction on CelebA. DiDAE’s trajectories are axis-parallel and roughly orthogonal between attempts, and all counterfactuals of one attempt share a direction, which is what makes per-cluster teacher feedback possible.
Fig. 5 : Selected qualitative samples on CelebA-Blond under the pipeline of Table II (edit-friendly DDPM inversion, seed 0 ): four factuals we chose from a fixed-seed random draw of sixteen of the 200 shared validation factuals as examples in which DiDAE’s edits flip the classifier. The sixteen unselected factuals, with every attempt of every method, are in Appendix Figure 9 , which the evaluation relies on. Both attempts of every method are shown as stored; a green frame marks an attempt that flipped the classifier, a red frame one that did not, and for SCE, CF1 is its less aggressive attempt. DAE’s two attempts coincide; DiME and ACE flip through faint, low-contrast changes; FastDiME rarely flips; SCE flips reliably but often through masks and textures that leave the face manifold, while DiDAE changes hair color or sex as a photographic edit.
Factor
Setting
Square
CelebA-Blond
Reference
ResNet-18, oracle, Procrustes
79.2
32.7
Teacher
pre-clustered
80.1
20.9
Student
probe
84.7
70.6
Dictionary
SVD (ResNet-18)
78.2
–
Correction
projection (probe)
53.4
22.6
TABLE III : DiDAE-CFKD Gain ablations , seed 0 , edit-friendly DDPM inversion. The first row is the reference setting of Table II (ResNet-18 student, oracle teacher, Procrustes dictionary, CFKD); every further row changes only the factor named in the first column. The pre-clustered teacher labels one direction instead of one counterfactual; the probe is a linear head on the frozen encoder; SVD exists on Square only; projection removes the spurious direction from the frozen embedding and retrains the probe instead of running CFKD, so its row compares with the probe row.
Method
Square
NICO++
GroupDRO
20.8
5.4
DFR
2.0
0.0
P-ClArC
56.2
-5.0
RR-ClArC
56.2
10.7
DiDAE-CFKD (ours)
80.1
16.3
TABLE IV : Gain against metadata-based baselines. Square: ResNet-18 students at 98% poisoning; the DiDAE-CFKD row uses the pre-clustered teacher, the weaker of the two teachers in Table III . NICO++: the baseline rows are the average-group accuracies reported in [ 33 ] for crocodile vs. lizard , converted to Gain; the DiDAE-CFKD entry is the Ranking DiDAE repair of Figure 6 , which corrects a ResNet-18 student like the baseline rows, with the dictionary living in the frozen CLIP space decoded through the representation autoencoder rather than in the student’s own features.
Fig. 6 : Ranking DiDAE surfaces the shortcut, and the funnel narrows at every step. Top five directions per classifier (four on freight car, the only ones with verified flips), ordered by verified ambient flips, against a dictionary that already exists in the frozen encoder space. The bars are the funnel of Section IV-C (linear, per panel), latent , ambient and verified flips (the ranking key, printed beside the bar); ✗ marks directions the teacher calls spurious, ✓causal ones. The lower row gives average group accuracy before and after CFKD on the spurious directions with the normalized Gain in percent; freight car is the setting of Figure 1 . NICO++ and ImageNet decode with edit-friendly DDPM inversion at classifier-free guidance 2 , Sparse Numbers with DDIM.
Fig. 7 : Counterfactuals along the top-ranked directions show what the classifier reads. Two before/after pairs for three of the directions of Figure 6 , chosen by eye from the saved verified flips, for the planted fireboat probe (left) and the NICO++ student (right), with the classifier’s prediction under each image; the frame gives the teacher’s verdict (green class evidence, red confounder). The fireworks flips are borderline, as its four verified flips suggest.
Fig. 8 : Line-search factor ablation on CelebA-Blond: fixed values of l against the empirical trust region ( auto ) of Section III-C . Fixed l≤1 barely changes the image and l≥5 leaves the manifold; the trust region does neither. Decoded with DDIM inversion, the weaker of the two inversions of Section III-D .
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. 9 : Random, unselected qualitative samples on CelebA-Blond under the pipeline of Table II (DiffAE conditioned on OpenAI CLIP ViT-L/14, edit-friendly DDPM inversion, seed 0 ); Figure 5 shows rows 3, 5, 11 and 15, selected as successful flips. The sixteen factuals are drawn uniformly at random with a fixed seed from the 200 validation factuals shared by all six runs; both attempts of every method are shown exactly as generated, without any filtering: a green frame marks an attempt that flipped the classifier, a red frame one that did not. For SCE, CF1 is its less aggressive attempt. This is the per-sample behavior behind the table: DAE’s two attempts coincide, since it has a single direction per classifier; DiME and ACE flip most samples through faint, low-contrast changes, which is why their non-adversarial rates in the table are low; FastDiME rarely flips; SCE flips reliably but frequently through masks and textures that leave the face manifold, which its high NAFR does not penalize. DiDAE’s counterfactuals almost always realize the intended edit of their semantic dimension (hair color or sex) as a photographic change, so a red frame on a DiDAE attempt rarely means a failed edit. The poisoned classifier, however, relies on both dimensions: an edit along one of them leaves the other still pointing to the original class, the two cues conflict, and whether the decision flips becomes close to a coin toss.
Square
Sparse Numbers
CelebA
Camelyon17
ImageNet / NICO++
Encoder Φ
ResNet-18
ResNet-18
CLIP ViT-L/14
PLIP
CLIP ViT-L/14
dimzFM
512
512
768
512
768
Generator
DiffAE
DiffAE
DiffAE
PathLDM
RAE (App. E-C )
Generator resolution
64
64
128
256
256
Classifier resolution
64
64
128
128
224
Inversion
DDPM †
DDIM
DDPM †
DDPM
DDPM
Appendix
TABLE V : Generator and dictionary configuration per dataset. The Square foundation model is a ResNet-18 trained to regress the four ground-truth generative factors; the dictionary is fitted on its 512 -d penultimate representation rather than on the four regressed factors. † DDIM in the ablation of Table VII only. ‡ Of the 1024 atoms, DiDAE edits the two with the largest weight in the distilled probe.
Line-search factor l
auto : target set to ckmin/ckmax (Eq. 5 ); effective l varies per sample
Empirical bounds
measured on the validation split, used unscaled
Denominator floor τ on w⊤vk
10−6
Lasso penalty λ (Eq. 8 , ranking only)
0.05 on Sparse Numbers, 0.01 on NICO++ and ImageNet, chosen from a sweep of the probe’s support size, normalized by the representation dimension; 2500 iterations of iterative soft-thresholding
Counterfactual attempts per factual
2
CFKD fine-tuning iterations
1
Train / validation counterfactuals
800 / 200
Appendix
TABLE VI : Editing settings, shared across datasets.
Dataset
Inversion
Flip rate
Diversity
Sparsity
NA
Unbiasedness
CF/s
Gain
Square
DiffAE / DDIM
55.2
79.2
71.9
90.3
84.6
47.8
88.6
DiffAE / DDPM
50.5
83.0
76.3
79.8
92.8
60.3
79.2
CelebA-Blond
DiffAE / DDIM
44.5
52.5
52.5
75.0
87.1
17.3
21.3
DiffAE / DDPM
53.0
75.8
61.2
84.6
84.9
14.2
32.7
RAE / DDPM
48.8
70.2
64.0
84.8
87.1
5.0
18.3
Appendix
TABLE VII : Inversion and generator ablation on seed 0 , holding the dictionary, probe and CFKD budget fixed. DiffAE rows (our pixel-space diffusion autoencoder, not the DAE baseline) change only the sampler; the RAE row swaps the pixel-space diffusion autoencoder for a representation autoencoder trained on CelebA (CLIP ViT-L/14 encoder, DDPM inversion; the recipe of Appendix E-C with a 0.41 B stage-2 transformer trained for 24 epochs instead of 46 , used at epoch 20 ). The first column is the student’s flip rate on the validation counterfactuals, counted per attempt, whereas the NAFR of Table II counts a factual once if its best attempt flips both classifiers, so it can exceed this rate; the other columns are the seed- 0 values that enter Table II , and the DDPM rows are the reference setting of Table III . That the Square DDIM row runs at 47.8 CF/s, the four-seed DDPM mean of Table II , is a coincidence.
Figure 17
Fig. 12 : Counterfactuals along the first four SVD directions of the Square foundation-model space (Section V-C ). Comp1 is the foreground color and Comp4 the background color; Comp2 and Comp3 mix the two spatial axes, which variance alone cannot separate.
Fig. 13 : Decision boundary of the ResNet-18 student before and after DiDAE-CFKD. The x-axis is the causal feature, the y-axis the confounding one. For Square the boundary is exact, since the whole dataset can be sampled and the x- and y-positions marginalized out; for CelebA it is approximated by projecting onto the oracle’s Male and Blond_Hair predictions (Appendix A ). Before CFKD the student relies almost entirely on the background on Square and about equally on both factors on CelebA; after one CFKD iteration it relies mostly on the causal feature in both cases.