Concept bottleneck models (CBMs) make predictions inspectable and intervenable by routing them through human-interpretable concepts, but originally required concept annotations. Annotation-free variants remove this requirement, but typically use large concept vocabularies, static at both training and inference, producing bottlenecks larger than any task or prediction needs and harder to inspect. We propose AnyBottle, a single recipe for building compact, task-specific CBMs. AnyBottle assumes only a frozen backbone and an unsupervised concept pool, such as a sparse autoencoder. A black-box teacher trained on the same backbone then guides selection: each round adds the concept that best explains the bottleneck's current failures, with candidates restricted to regions of teacher/student disagreement. Trained with nested dropout over this selection order, the final bottleneck predicts accurately from any concept prefix, so inference spends fewer concepts on inputs it is confident about early and more on hard ones. Since no stage is modality-specific, a new domain and task requires swapping only the backbone and concept pool. Across six vision and two text datasets and two teacher paradigms, AnyBottle yields bottlenecks with fewer concepts and higher concept consistency than annotation-free baselines, while staying close to the black-box reference. Overall, AnyBottle shows that going annotation-free need not mean going large: a small, discovered vocabulary can be as expressive as a much larger, fixed one.
Figures & tables
Figure 1: Only as many concepts as needed. Left: existing annotation-free concept bottlenecks draw from a large, static concept pool, using far more concepts than a given task needs, and the same fixed vocabulary for every sample regardless of difficulty. AnyBottle instead selects a concept set sized to the task, and lets inference use only as many of those concepts as each sample actually requires. Right: two test samples illustrating this sample-level adaptivity.
Figure 2: Overview of AnyBottle . An input x is encoded by the frozen backbone g into token embeddings zℓ , which a sparse autoencoder a maps to concept activations cℓ . Each selection round repeats four steps: (1) a forward pass through the teacher fT and current student fS ; (2) teacher attributions over the input; (3) the resulting teacher/student blind-spots (green), used to restrict the SAE to blind-spot tokens; and (4) scoring the resulting candidates by least-squares residual R against the teacher’s output YT to pick the next concept j⋆ , added to S . This repeats until selection converges on S⋆ . Once selection terminates, (5) the final predictor f is trained once on rescaled C~S⋆ with nested dropout, yielding a head that predicts accurately from any concept prefix.
Figure 3: Top: AnyBottle matches or exceeds black-box teacher accuracy at a fraction of the concepts used. Accuracy gap ( Δ (pp)) to the teacher, plotted against concept count (log scale); AnyBottle reaches the teacher across all six datasets with orders of magnitude fewer concepts than the baselines. Bottom: AnyBottle ’s smaller bottleneck is also the more interpretable one. Concept consistency ( Cimg2 ) at the same budgets matches or exceeds every baseline, showing the compactness above is not simply discarded capacity. Points are seed means.
Figure 4: Left: AnyBottle reaches near-full accuracy with a fraction of its concept budget. Concepts needed to reach 90 / 95 / 98% of the student’s full-budget accuracy, per dataset; K shown for reference. Right: Early rounds close large blind spots; later rounds refine with diminishing returns. Round-by-round validation accuracy on Imagenette, with two exemplary rounds detailed.
Figure 5: Adaptive inference uses more concepts only on samples that need them. Two easy samples (top) and one hard sample (bottom) given the task and selected concept set. For each sample, the activating concepts to reach 90% confidence are shown, together with three exemplars. Boxed numbers mark firing concepts left out for visibility.
LS
Loc
PF
K ↓
C patch2↑
Δ↑ (%)
✓
✓
✓
145.0
0.362
−0.4
✓
✓
–
165.7
0.345
−0.9
✓
–
–
190.5
0.314
−1.3
–
–
–
231.3
0.272
−2.1
Table 1: Every component of AnyBottle ’s selection facilitates compact, consistent, and accurate concept sets. Ablation over selection strategy, averaged over vision datasets. LS: least-squares scoring (vs. random selection). Loc: localized candidate search. PF: prefiltering.
Method
K ↓
AIS ↑
Δ (%) ↑
Acc (%) ↑
AG News
SAE probe
31477
0.327±0.009
−0.12±0.38
90.42±0.38
AnyBottle (mlp)
93.7±6.4
0.469±0.006
−0.18±0.25
90.36±0.25
[2pt/2pt]
SAE probe
9710
0.157±0.001
+18.57±0.98
88.25±0.98
AnyBottle (verb)
30.0±5.0
0.094±0.020
+0.82±1.89
70.50±1.89
PubMed
SAE probe
25893
0.453±0.000
−2.47±0.87
73.58±0.87
AnyBottle (mlp)
100.0±2.6
0.526±0.004
−0.02±0.23
76.03±0.23
Table 2: AnyBottle generalizes across domains and teacher paradigms. Evaluation of AnyBottle on two text datasets with a trained mlp-clf teacher and a zero-shot verbalizer teacher, verb-clf. We report concept count K , AutoInterp score (AIS), teacher gap ( Δ ), and absolute balanced accuracy (Acc).
Figure 6: Raw accuracy of CBM baselines across vision datasets.
Figure 7: Patch-level concept consistency (C patch2 ) versus number of concepts, across six datasets. Top row : Linear SAE probe evaluated at its full dictionary size (8192) compared to AnyBottle. Bottom row : the same comparison restricted to the Linear SAE probe’s active features only, shown at a finer x-axis scale. AnyBottle achieves comparable or higher C patch2 at substantially fewer concepts.
Dataset
LS
Loc
PF
K ↓
C patch2↑
Δ↑ (%)
Imagenette
✓
✓
✓
34.0±2.0
0.442±0.005
+0.3±0.6
✓
✓
–
47.3±3.5
0.402±0.008
−0.9±0.4
✓
–
–
87.3±1.5
0.348±0.003
+0.4
–
–
–
184.7±22.8
0.269±0.015
+0.2
Dogs-10
✓
✓
✓
37.0±1.0
0.341±0.002
−1.8±0.3
✓
✓
–
49.0±0.0
0.345±0.000
−2.0±0.2
Table 3: Per-dataset results underlying the averaged ablation in Table Tab. 1 (main text). Ablations of least-squares scoring (LS), localized candidate search (Loc), and prefiltering (PF). Δ = balanced test accuracy − balanced teacher accuracy. C patch2 is the patch-level concept consistency. Mean ± std over 3 seeds.
Dataset
Config
#Concepts ↓
C patch2↑
Δ↑ (%)
Imagenette
Discrete
90.0±9.0
0.38±0.012
+0.1±0.4
Continuous
34.0±2.0
0.442±0.005
+0.3±0.6
Dogs-10
Discrete
89.0±3.0
0.317±0.004
+0.3±0.8
Continuous
37.0±1.0
0.341±0.002
−1.8±0.3
Machines-15
Discrete
141.0±38.0
0.405±0.015
−0.9±1.1
Continuous
110.0±14.0
0.399±0.009
+0.3±0.4
Table 4: Ablation studies on concept representation (left) and SAE fine-tuning (right). Left : effect of concept representation with a Task-aware: discrete vs continuous concept scores. Right : effect of SAE fine-tuning with continuous concepts: task-agnostic pretrained vs. task-aware finetuned. Δ = balanced test accuracy − balanced teacher accuracy; C patch2 is the patch-level concept consistency. Mean ± std over 3 seeds.
Dataset
δacc
#Concepts
C patch2↑
C img2↑
Δval (%)
Δtest (%)
Imagenette
0.0
33.7±2.1
0.442±0.005
0.324±0.003
+0.09
+0.33
0.1
33.7±2.1
0.442±0.005
0.324±0.003
+0.09
+0.33
0.2
25.0±4.6
0.453±0.007
0.335±0.007
−0.19
−0.27
Dogs-10
0.0
36.7±0.6
0.341±0.002
0.105±0.001
+0.05
−1.80
0.1
36.7±0.6
0.341±0.002
0.105±0.001
+0.05
−1.80
0.2
36.7±0.6
0.341±0.002
0.105±0.001
+0.05
−1.80
Table 5: Effect of the stopping tolerance δacc on AnyBottle, across all six vision datasets. Δ = Accuracy − Teacher (balanced accuracy). C patch2 is the patch-level (localised) consistency score; C img2 is the image-level score. Mean ± std over 3 seeds.
Dataset
Method
#Concepts ↓
C@90 ↓
C@95 ↓
AIS ↑
Δ (%) ↑
Acc (%) ↑
AG News
Linear SAE probe
31477
–
–
0.327±0.009
−0.12±0.38
90.42±0.38
mlp-clf (gemma-2-2B)
93.7±6.4
63±1
64±0
0.469±0.006
−0.18±0.25
90.36±0.25
[2pt/2pt]
Linear SAE probe
9710
–
–
0.157±0.001
+18.57±0.98
88.25±0.98
verb-clf (gemma-3-1B-it)
30.0±5.0
27±10
30±6
0.094±0.020
+0.82±1.89
70.50±1.89
PubMed-20k
Linear SAE probe
25893
–
–
0.453±0.000
−2.47±0.87
73.58±0.87
mlp-clf (gemma-2-2B)
100.0±2.6
47±10
62±1
0.526±0.004
−0.02±0.23
76.03±0.23
Table 6: Across domains and teacher paradigms, AnyBottle matches near-teacher accuracy with orders of magnitude fewer concepts than a linear probe over the full SAE dictionary. We report concept count K , concepts needed to reach 90% and 95% of the student accuracy (C@90/95), AutoInterp score (AIS), and balanced accuracy gap to teacher ( Δ ) and in absolute terms (Acc).
Figure 8: Concept selection round by round on all six vision datasets. For each dataset, we show the validation accuracy of the current predictor fS (left) and the concepts and blind spots for three rounds (right). The dashed line is the teacher’s validation accuracy. For each round, we show one training image from its focus set on which the added concept fires, with the teacher’s attribution and the concept’s activation. Green outlines mark the round’s blind spots, i.e. the tokens the teacher attends to, but fS does not. Below each image are its class and the probability fS assigns to it before → after the round.
Figure 9: Adaptive inference on the other five datasets. Two test images per dataset, one per row, as Fig. 5 shows them for Mixed-20. Imagenette, Dogs-10 and Machines-15 here; Flowers-102 and ISIC-7 in Fig. 10 .
Figure 10: Adaptive inference on Flowers-102 and ISIC-7 , continued from Fig. 9 .
Figure 11: Concept examples. Five concepts per dataset, drawn at random, one per row, each with the ten test images it fires on most strongly. Each image keeps its colour where the concept fires, scaled to its own strongest patch, and fades elsewhere. Imagenette, Dogs-10 and Machines-15 here; Mixed-20, Flowers-102 and ISIC-7 in Fig. 12 .
Figure 12: Concept examples , continued from Fig. 11 : Mixed-20, Flowers-102 and ISIC-7.
Subset
Classes
Dogs (10)
Golden retriever, Labrador retriever, Chesapeake Bay retriever, German shepherd, Rottweiler, Doberman, Boxer, Great Dane, Siberian husky, Eskimo dog
Animals: Lion, Tiger, African elephant, Zebra, Giant panda, Polar bear, Gorilla, Flamingo, Ostrich, Hippopotamus; Machines: Sports car, Warplane, Submarine, Steam locomotive, Forklift, Ambulance, Tank, Space shuttle, Fire engine, Motor scooter
Table 7: ImageNet-1k classes used in our three curated subsets.
Dataset
Prompt
Class → token
AG News
Article: "{text}" Which topic best fits this article: World, Sports, Business, or Technology? Answer with exactly one word.
World, Sports, Business, Sci/Tech → Technology
PubMed 20k RCT
Sentence from a medical research abstract: "{text}" Does this sentence state the study’s background, objective, methods, results, or conclusion? Answer with exactly one word.
Table 9: Hyperparameters used throughout AnyBottle ’s selection and training pipeline. Where vision and text differ, both values are given; a single value applies to both. For text, G2 = Gemma-2-2B MLP-teacher runs, G3 = Gemma-3-1B-it verbalizer runs. Accuracy thresholds are in percentage points.
Concept-bottleneck models (CBMs) are neural classifiers that compute predictions from high-level concepts extracted from the input. CBMs ensure stakeholders can understand the concepts -- and the predictions they entail -- by learning these from concept-level annotations, which are however seldom available. Recent CBM architectures work around this issue by obtaining annotations from Vision-Language Models (VLMs). While greatly broadening applicability, doing so can yield lower quality concepts and therefore less interpretable models. We strike for a middle ground by introducing Vision-plus-Human-guided CBM (VH-CBM), a hybrid approach that exploits both VLMs and a small amount of dense annotations. VH-CBM employs a Gaussian Process in the VLM's embedding space, which captures useful global information about the target domain, to propagate the expert's supervision to any target data point. Our empirical evaluation shows how VH-CBM predicts more accurate concepts than VLM-guided CBMs even when annotating as little as 1% of the data, while sporting better concept calibration and supporting active learning.
Nicola Debole, Andrea Passerini, Stefano Teso +2
DISI, University of Trento, Italy · CIMeC, University of Trento, Italy
Concept Bottleneck Models (CBMs) provide interpretability by routing predictions through a layer of human-understandable concepts. However, defining an optimal concept set for a specific dataset remains an open challenge. Existing approaches rely on expensive expert annotations or LLM-generated lists based solely on class names. Even "open-vocabulary" variants typically depend on static concept sets, which restrict discovery and introduce label bias. Furthermore, traditional CBMs often suffer from information leakage, where unmodeled visual features bypass the bottleneck and compromise the integrity of the explanations. To overcome these limitations, we propose Caption Bottleneck Models (CaBM), a framework that circumvents the need for predefined concept sets by replacing rigid concept layers with free-form natural language. By representing images via LMM-generated captions and training a classifier strictly on this text, CaBM ensures a leakage-free architecture by construction. Additionally, by analyzing the text classifier post-training, CaBM autonomously discovers high-quality, dataset-specific concepts. Our results across fine- and coarse-grained benchmarks demonstrate that CaBM achieves competitive accuracy while preserving interpretability without the constraints of external dictionaries or manual labeling.
Seref Baris Cagliyan, Umut Ozdemir, Merve Tapli +1
Dept. of Computer Eng., Middle East Technical University (METU) · Robotics & AI Center (ROMER), METU
Concept Bottleneck Models (CBMs) provide an intrinsically interpretable alternative to post-hoc explanations. However, existing CBMs often rely on predefined concept vocabularies or supervised annotations, lack explicit concept grounding, and summarize each concept with a single image-level score -- discarding spatial recurrence and inter-concept dependencies. We propose a Graph-based Concept Bottleneck Model (G-CBM), an intrinsically interpretable framework that performs unsupervised concept discovery via Non-negative Matrix Factorization (NMF) and represents the discovered concepts as nodes in a per-image concept-graph representation. G-CBM matches region-level features to these concept nodes -- providing concept grounding and capturing concept recurrence across the image -- and applies a \emph{tunable concept filtering threshold} τ to suppress weak region-level features. A Graph Attention Network (GAT) then performs concept-level reasoning by modeling nonlinear dependencies across nodes. Across ImageNet, HAM10000, PH2, and Derm7pt, G-CBM achieves an average relative AUC improvement of 3.7% over a ResNet-50 baseline. Concept filtering frequently improves predictive performance while inducing selective concept use, achieving peak AUC of 0.96 on PH2 with only 2 of 10 concepts and 0.92 on HAM10000 with 3.8 of 9 concepts. On dermoscopy benchmarks, G-CBM is competitive with supervised approaches requiring external annotations. Deletion/insertion analyses with random ablation controls show that the learned concept ranking faithfully reflects model predictions.
Md Mohasin Hossain, Anar Amirli, Robert Leist +2
German Research Center for Artificial Intelligence · Saarland University, Saarbrücken, Germany · 1German Research Center for Artificial Intelligence (DFKI), Saarbrücken, Germany +5