Concept-based methods provide a semantic level for interpreting and manipulating learned representations, but existing editing approaches are typically specialized to particular interventions and do not provide a common and editable representation of concept organization. To achieve this, we introduce Topological Concept Representations (TCR), a post-hoc operational abstraction that jointly characterizes the concepts encoded in a learned representation and their relationships. TCR constructs an intermediate concept space from concept recoverability and interaction scores, and compactly encodes its organization through topology. Interventions are expressed through modifications of this abstraction that are propagated back to the underlying learned representation. This allows different concept-level operations to share the same optimization framework and separates the desired concept organization from the mechanism to achieve it. We establish stability and reparameterization-invariance properties of TCR and its connections to existing concept-editing formulations. We use TCR to disentangle concepts as a preprocessing step for existing erasure methods, improving worst-group accuracy by 21.89 on average at comparable concept leakage. We further use TCR to transfer concepts from teacher to student models, improving concept recoverability by up to 5.54 while also improving or maintaining competitive test top-1 accuracy.
Figures & tables
Figure 1: TCR provides a common post-hoc abstraction for concept organization in learned representations. It maps encoded concepts and their interactions to a concept mesh and summarizes its topological structure. Desired changes are specified on this abstraction and propagated back to the representation, enabling diverse interventions such as concept erasure, disentanglement and transfer.
Method
Waterbirds
Bias in Bios SN
Dial Mention
Jigsaw
Raw (WGA / Leakage)
56.54 / 80.22
78.77 / 98.48
39.20 / 54.57
53.29 / 41.57
INLP
27.88 ±0.00
59.47 ±0.00
52.05 ±0.00
55.64 ±0.00
LEACE
13.71 ±0.00
59.47 ±0.00
50.45 ±0.00
52.59 ±0.00
RLACE
14.55 ±0.95
59.53 ±0.92
50.80 ±0.29
53.27 ±0.37
SPLINCE
13.71 ±0.00
58.98 ±0.00
50.45 ±0.00
52.59 ±0.00
LEOPARD
13.71 ±0.00
59.47 ±0.00
49.95 ±0.19
52.59 ±0.00
Table 1: Test task worst-group accuracy with an adjusted balanced linear concept leakage <10% .
Figure 2: Erasure-utility trade-off showing the test WGA given the adjusted balanced linear leakage (a) and impact on test WGA of varying label imbalance (b) and spurious correlation (c).
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Task y
Concept c
Train
Val
Test
Representation
Original dim.
PCA dim.
Waterbirds
Bird type
Background
4,795
1,199
5,794
ResNet-50 avgpool
2,048
128
Bias in Bios SN
Profession (surgeon/nurse)
Gender
21,145
3,256
8,136
BERT [CLS]
768
128
Dial Mention
Mention detection
Dialect
56,140
8,000
8,000
BERT [CLS]
768
128
Jigsaw
Toxicity classification
Gender
24,508
2,724
6,808
BERT [CLS]
768
128
Appendix
Table 3: Datasets and fixed representations used for concept disentanglement before erasure.
Dataset
Split
Group size nyc
Agreement P(y=c)
Label imbalance P(y=0)
n00
n01
n10
n11
Waterbirds
Train
3,498
(72.95%)
184
(3.84%)
56
(1.17%)
1,057
(22.04%)
94.99%
76.79%
Val
467
(38.95%)
466
(38.87%)
133
(11.09%)
133
(11.09%)
50.04%
77.81%
Test
2,255
(38.92%)
2,255
(38.92%)
642
(11.08%)
642
(11.08%)
50.00%
77.84%
Bias in Bios SN
Train
7,521
(35.57%)
1,308
(6.19%)
1,127
(5.33%)
11,189
(52.92%)
88.48%
41.75%
Val
1,158
(35.57%)
202
(6.20%)
174
(5.34%)
1,722
(52.89%)
88.45%
41.77%
Appendix
Table 4: Group composition of each dataset split. The four group sizes are indexed by the task and concept labels (y,c) , with their percentage of the split shown in parentheses. Agreement is P(y=c) and label imbalance is reported as P(y=0) .
Role
Architecture
Penultimate dim.
Parameters
Training
Teacher
ResNet32 × 4
256
7,433,860
CKD pretrained checkpoint, frozen
Student
ResNet8 × 4
256
1,233,540
From scratch
Student
ShuffleNetV2
1,024
1,355,528
From scratch
Appendix
Table 5: Teacher and student architectures used for concept transfer.
Table 7
Method
Waterbirds (s)
Bias in Bios SN (s)
Dial Mention (s)
Jigsaw (s)
INLP
1.35 ±0.41
5.57 ±1.01
28.05 ±1.90
3.07 ±0.65
LEACE
0.03 ±0.06 0.03 ±0.06
0.03 ±0.06
0.04 ±0.07
0.02 ±0.05
RLACE
38.91 ±1.13
65.28 ±2.56
208.60 ±3.70
96.77 ±6.67
SPLINCE
0.03 ±0.02
0.15 ±0.20 0.15 ±0.20
0.30 ±0.33 0.30 ±0.33
0.10 ±0.03 0.10 ±0.03
LEOPARD
10.65 ±0.61
47.28 ±0.36
132.32 ±0.56
53.18 ±0.21
KRaM
5.39 ±1.30
36.58 ±4.44
59.58 ±28.59
45.26 ±5.67
Appendix
Table 8: Execution time (seconds) for standalone erasure methods, TCR-DIS followed by erasure, and direct TCR-ERA.
Method
Waterbirds
Bias in Bios SN
Dial Mention
Jigsaw
Raw
88.95
96.76
64.91
71.56
INLP
85.17 ±0.00
71.66 ±0.00
63.49 ±0.00
65.82 ±0.00
LEACE
85.67 ±0.00
76.55 ±0.00
65.06 ±0.00 65.06 ±0.00
67.58 ±0.00 67.58 ±0.00
RLACE
85.64 ±0.25
76.62 ±0.33
64.78 ±0.21
67.30 ±0.17
SPLINCE
85.66 ±0.00
76.50 ±0.00
65.06 ±0.00 65.06 ±0.00
67.51 ±0.00
LEOPARD
85.66 ±0.00
76.55 ±0.00
64.85 ±0.02
67.58 ±0.00 67.58 ±0.00
Appendix
Table 9: Test task accuracy with an adjusted balanced linear concept leakage <0.1 .
Method
Waterbirds
Bias in Bios SN
Dial Mention
Jigsaw
Raw
85.71
98.62
61.78
51.88
INLP
60.58 ±0.00
77.59 ±0.00
34.47 ±0.00
29.99 ±0.00
LEACE
72.35 ±0.00
95.09 ±0.00
38.13 ±0.00
33.37 ±0.00
RLACE
72.37 ±0.26
94.99 ±0.10
38.30 ±0.20
33.55 ±0.49
SPLINCE
72.14 ±0.00
94.56 ±0.00
38.28 ±0.00
33.31 ±0.00
LEOPARD
72.59 ±0.00
94.99 ±0.00
35.84 ±0.26
33.34 ±0.00
Appendix
Table 10: Adjusted balanced oracle-conditioned test concept leakage.
Method
Waterbirds
Bias in Bios SN
Dial Mention
Jigsaw
Raw
79.83
98.89
55.99
42.77
INLP
71.01 ±1.60
93.64 ±0.21
52.99 ±0.65
31.37 ±0.95
LEACE
74.60 ±0.92
98.05 ±0.21
55.14 ±0.48
35.95 ±0.91
RLACE
74.54 ±1.79
98.01 ±0.27
54.96 ±0.58
36.44 ±1.11
SPLINCE
75.14 ±1.42
97.45 ±0.26
56.18 ±0.45
36.91 ±1.03
LEOPARD
74.07 ±1.06
98.02 ±0.18
53.45 ±0.47
35.44 ±1.24
Appendix
Table 11: Adjusted balanced non-linear test concept leakage.
Figure 4: Extended erasure-utility and data-regime analysis. Erasure-utility trade-off, showing test WGA against adjusted balanced linear leakage, on Waterbirds (a) and Bias in Bios (b); impact of label imbalance on DIAL-Mention (c) and Jigsaw (d); and of spurious correlation on DIAL-Mention (e) and Jigsaw (f). Panels (a), (c) and (f) correspond to the main-text results for direct comparison.
Figure 5: Task utility against representational similarity after concept erasure, showing test WGA against linear CKA between the original and intervened test representations on Waterbirds (a), Bias in Bios (b), DIAL-Mention (c) and Jigsaw (d). Higher CKA indicates greater similarity.
Concept erasure aims to remove a target concept from a representation while preserving the other information encoded in it. This is difficult because representations encode many concepts that are often correlated with the erasure target, so removing the target risks damaging them. We propose the Manifold Constraint Hypothesis (MCH): if natural representations concentrate on a structured, lower-dimensional manifold, then interventions should be constrained to that manifold and better preserve other information encoded in the representation during interventions. We instantiate MCH in a new concept erasure method: MANifold aware Concept Erasure (MANCE). MANCE performs iterative updates to the representations using signals from a classifier that predicts a target concept. We estimate the manifold using representations obtained from natural inputs, and then we project the concept removal update to the estimated manifold. We perform extensive evaluation on 119 settings spanning text and vision, including 13 language models, three NLP concepts, and 40 CelebA-CLIP attributes. Employing MANCE on top of previous methods shows consistent improved leakage results. We also introduce MANCE+ and MANCE++, which prepend a closed-form erasure algorithm before employing MANCE, achieving better leakage--surgicality tradeoffs relative to matched full-space updates. MANCE++, our best method, achieves state-of-the-art results on nonlinear concept erasure. These results support MCH in the erasure setting: interventions should be constrained to the natural representation manifold.
Matan Avitan, Yoav Goldberg, Yanai Elazar
Bar-Ilan University · Allen Institute for Artificial Intelligence
Neural networks are increasingly employed to identify both well-defined and ambiguous concepts, yet output-level metrics reveal little about how those concepts are represented internally. Our study asks if these networks exhibit \textit{conceptual separation}: if examples of the same concept form coherent representations, and whether related concepts lie closer together in the representation space. We examine this conceptual organisation in Convolutional Neural Networks (CNNs) and Large Language Models (LLMs) through geometric and distributional analysis of their internal activations. In CNNs, familiar ImageNet concepts form coherent and semantically ordered representations, while this coherence weakens for unseen concepts and suffers within-class domain shift. In LLMs, clearly distinct domains remain well separated, related subdomains move closer together, and the distinction between ambiguous topics collapses at both the mean and covariance level. These results suggest that conceptual separation can reveal structure that output accuracy alone cannot, and may serve as a useful diagnostic of how robustly a model represents the concepts it is asked to identify. Code and data available on GitHub.
Jaee Ponde, Roshni Agarwal, Subhashis Banerjee
Truth Audit Labs · Karya AI · Department of Computer Science and the Centre of Digitalisation, AI, and Society, Ashoka University, Sonepat, Haryana, India
Concept erasure has emerged as a promising approach to mitigate undesired or unsafe content in diffusion models, yet existing methods still face significant limitations. While training-based methods are effective, their high computational cost limits scalability. Editing-based methods are more efficient and deployment-friendly, yet they struggle to simultaneously achieve precise concept erasure and preserve overall generative capacity. We identify this core limitation of the editing-based methods as reliance on additive parameter updates. Our empirical analysis reveals that concept semantics primarily depend on neuron direction rather than neuron magnitude, while overall generative capacity relies on the angular geometry of neurons. As additive updates inherently entangle direction, magnitude, and angular geometry, they inevitably introduce unintended interference between concept erasure and overall generation performance. To address this, we propose Orthogonal Concept Erasure (OCE), which reformulates editing-based erasure as multiplicative parameter updates from a geometric perspective. Specifically, OCE applies layer-wise orthogonal transformations derived from a closed-form solution to the parameters, enabling precise concept erasure while preserving the neuron magnitude and angular geometry. Furthermore, to address conflicting constraints in multi-concept erasure, OCE introduces a subspace-level objective with structured subspace manipulation, yielding a more effective and scalable erasure. Extensive experiments on single- and multi-concept erasure demonstrate that OCE outperforms existing methods in concept erasure and non-target preservation, erasing up to 100 concepts in 4.3 s. Code: https://github.com/HansSunY/OCE.
Yuhao Sun, Lingyun Yu, Haoxiang Xu +3
University of Science and Technology of China · Ant Group