Concept bottleneck models (CBMs) make predictions inspectable and intervenable by routing them through human-interpretable concepts, but originally required concept annotations. Annotation-free variants remove this requirement, but typically use large concept vocabularies, static at both training and inference, producing bottlenecks larger than any task or prediction needs and harder to inspect. We propose AnyBottle, a single recipe for building compact, task-specific CBMs. AnyBottle assumes only a frozen backbone and an unsupervised concept pool, such as a sparse autoencoder. A black-box teacher trained on the same backbone then guides selection: each round adds the concept that best explains the bottleneck's current failures, with candidates restricted to regions of teacher/student disagreement. Trained with nested dropout over this selection order, the final bottleneck predicts accurately from any concept prefix, so inference spends fewer concepts on inputs it is confident about early and more on hard ones. Since no stage is modality-specific, a new domain and task requires swapping only the backbone and concept pool. Across six vision and two text datasets and two teacher paradigms, AnyBottle yields bottlenecks with fewer concepts and higher concept consistency than annotation-free baselines, while staying close to the black-box reference. Overall, AnyBottle shows that going annotation-free need not mean going large: a small, discovered vocabulary can be as expressive as a much larger, fixed one.
Figures & tables
Figure 1: Only as many concepts as needed. Left: existing annotation-free concept bottlenecks draw from a large, static concept pool, using far more concepts than a given task needs, and the same fixed vocabulary for every sample regardless of difficulty. AnyBottle instead selects a concept set sized to the task, and lets inference use only as many of those concepts as each sample actually requires. Right: two test samples illustrating this sample-level adaptivity.
Figure 2: Overview of AnyBottle . An input x is encoded by the frozen backbone g into token embeddings zℓ , which a sparse autoencoder a maps to concept activations cℓ . Each selection round repeats four steps: (1) a forward pass through the teacher fT and current student fS ; (2) teacher attributions over the input; (3) the resulting teacher/student blind-spots (green), used to restrict the SAE to blind-spot tokens; and (4) scoring the resulting candidates by least-squares residual R against the teacher’s output YT to pick the next concept j⋆ , added to S . This repeats until selection converges on S⋆ . Once selection terminates, (5) the final predictor f is trained once on rescaled C~S⋆ with nested dropout, yielding a head that predicts accurately from any concept prefix.
Figure 3: Top: AnyBottle matches or exceeds black-box teacher accuracy at a fraction of the concepts used. Accuracy gap ( Δ (pp)) to the teacher, plotted against concept count (log scale); AnyBottle reaches the teacher across all six datasets with orders of magnitude fewer concepts than the baselines. Bottom: AnyBottle ’s smaller bottleneck is also the more interpretable one. Concept consistency ( Cimg2 ) at the same budgets matches or exceeds every baseline, showing the compactness above is not simply discarded capacity. Points are seed means.
Figure 4: Left: AnyBottle reaches near-full accuracy with a fraction of its concept budget. Concepts needed to reach 90 / 95 / 98% of the student’s full-budget accuracy, per dataset; K shown for reference. Right: Early rounds close large blind spots; later rounds refine with diminishing returns. Round-by-round validation accuracy on Imagenette, with two exemplary rounds detailed.
Figure 5: Adaptive inference uses more concepts only on samples that need them. Two easy samples (top) and one hard sample (bottom) given the task and selected concept set. For each sample, the activating concepts to reach 90% confidence are shown, together with three exemplars. Boxed numbers mark firing concepts left out for visibility.
LS
Loc
PF
K ↓
C patch2↑
Δ↑ (%)
✓
✓
✓
145.0
0.362
−0.4
✓
✓
–
165.7
0.345
−0.9
✓
–
–
190.5
0.314
−1.3
–
–
–
231.3
0.272
−2.1
Table 1: Every component of AnyBottle ’s selection facilitates compact, consistent, and accurate concept sets. Ablation over selection strategy, averaged over vision datasets. LS: least-squares scoring (vs. random selection). Loc: localized candidate search. PF: prefiltering.
Method
K ↓
AIS ↑
Δ (%) ↑
Acc (%) ↑
AG News
SAE probe
31477
0.327±0.009
−0.12±0.38
90.42±0.38
AnyBottle (mlp)
93.7±6.4
0.469±0.006
−0.18±0.25
90.36±0.25
[2pt/2pt]
SAE probe
9710
0.157±0.001
+18.57±0.98
88.25±0.98
AnyBottle (verb)
30.0±5.0
0.094±0.020
+0.82±1.89
70.50±1.89
PubMed
SAE probe
25893
0.453±0.000
−2.47±0.87
73.58±0.87
AnyBottle (mlp)
100.0±2.6
0.526±0.004
−0.02±0.23
76.03±0.23
Table 2: AnyBottle generalizes across domains and teacher paradigms. Evaluation of AnyBottle on two text datasets with a trained mlp-clf teacher and a zero-shot verbalizer teacher, verb-clf. We report concept count K , AutoInterp score (AIS), teacher gap ( Δ ), and absolute balanced accuracy (Acc).
Figure 6: Raw accuracy of CBM baselines across vision datasets.
Figure 7: Patch-level concept consistency (C patch2 ) versus number of concepts, across six datasets. Top row : Linear SAE probe evaluated at its full dictionary size (8192) compared to AnyBottle. Bottom row : the same comparison restricted to the Linear SAE probe’s active features only, shown at a finer x-axis scale. AnyBottle achieves comparable or higher C patch2 at substantially fewer concepts.
Dataset
LS
Loc
PF
K ↓
C patch2↑
Δ↑ (%)
Imagenette
✓
✓
✓
34.0±2.0
0.442±0.005
+0.3±0.6
✓
✓
–
47.3±3.5
0.402±0.008
−0.9±0.4
✓
–
–
87.3±1.5
0.348±0.003
+0.4
–
–
–
184.7±22.8
0.269±0.015
+0.2
Dogs-10
✓
✓
✓
37.0±1.0
0.341±0.002
−1.8±0.3
✓
✓
–
49.0±0.0
0.345±0.000
−2.0±0.2
Table 3: Per-dataset results underlying the averaged ablation in Table Tab. 1 (main text). Ablations of least-squares scoring (LS), localized candidate search (Loc), and prefiltering (PF). Δ = balanced test accuracy − balanced teacher accuracy. C patch2 is the patch-level concept consistency. Mean ± std over 3 seeds.
Dataset
Config
#Concepts ↓
C patch2↑
Δ↑ (%)
Imagenette
Discrete
90.0±9.0
0.38±0.012
+0.1±0.4
Continuous
34.0±2.0
0.442±0.005
+0.3±0.6
Dogs-10
Discrete
89.0±3.0
0.317±0.004
+0.3±0.8
Continuous
37.0±1.0
0.341±0.002
−1.8±0.3
Machines-15
Discrete
141.0±38.0
0.405±0.015
−0.9±1.1
Continuous
110.0±14.0
0.399±0.009
+0.3±0.4
Table 4: Ablation studies on concept representation (left) and SAE fine-tuning (right). Left : effect of concept representation with a Task-aware: discrete vs continuous concept scores. Right : effect of SAE fine-tuning with continuous concepts: task-agnostic pretrained vs. task-aware finetuned. Δ = balanced test accuracy − balanced teacher accuracy; C patch2 is the patch-level concept consistency. Mean ± std over 3 seeds.
Dataset
δacc
#Concepts
C patch2↑
C img2↑
Δval (%)
Δtest (%)
Imagenette
0.0
33.7±2.1
0.442±0.005
0.324±0.003
+0.09
+0.33
0.1
33.7±2.1
0.442±0.005
0.324±0.003
+0.09
+0.33
0.2
25.0±4.6
0.453±0.007
0.335±0.007
−0.19
−0.27
Dogs-10
0.0
36.7±0.6
0.341±0.002
0.105±0.001
+0.05
−1.80
0.1
36.7±0.6
0.341±0.002
0.105±0.001
+0.05
−1.80
0.2
36.7±0.6
0.341±0.002
0.105±0.001
+0.05
−1.80
Table 5: Effect of the stopping tolerance δacc on AnyBottle, across all six vision datasets. Δ = Accuracy − Teacher (balanced accuracy). C patch2 is the patch-level (localised) consistency score; C img2 is the image-level score. Mean ± std over 3 seeds.
Dataset
Method
#Concepts ↓
C@90 ↓
C@95 ↓
AIS ↑
Δ (%) ↑
Acc (%) ↑
AG News
Linear SAE probe
31477
–
–
0.327±0.009
−0.12±0.38
90.42±0.38
mlp-clf (gemma-2-2B)
93.7±6.4
63±1
64±0
0.469±0.006
−0.18±0.25
90.36±0.25
[2pt/2pt]
Linear SAE probe
9710
–
–
0.157±0.001
+18.57±0.98
88.25±0.98
verb-clf (gemma-3-1B-it)
30.0±5.0
27±10
30±6
0.094±0.020
+0.82±1.89
70.50±1.89
PubMed-20k
Linear SAE probe
25893
–
–
0.453±0.000
−2.47±0.87
73.58±0.87
mlp-clf (gemma-2-2B)
100.0±2.6
47±10
62±1
0.526±0.004
−0.02±0.23
76.03±0.23
Table 6: Across domains and teacher paradigms, AnyBottle matches near-teacher accuracy with orders of magnitude fewer concepts than a linear probe over the full SAE dictionary. We report concept count K , concepts needed to reach 90% and 95% of the student accuracy (C@90/95), AutoInterp score (AIS), and balanced accuracy gap to teacher ( Δ ) and in absolute terms (Acc).
Figure 8: Concept selection round by round on all six vision datasets. For each dataset, we show the validation accuracy of the current predictor fS (left) and the concepts and blind spots for three rounds (right). The dashed line is the teacher’s validation accuracy. For each round, we show one training image from its focus set on which the added concept fires, with the teacher’s attribution and the concept’s activation. Green outlines mark the round’s blind spots, i.e. the tokens the teacher attends to, but fS does not. Below each image are its class and the probability fS assigns to it before → after the round.
Figure 9: Adaptive inference on the other five datasets. Two test images per dataset, one per row, as Fig. 5 shows them for Mixed-20. Imagenette, Dogs-10 and Machines-15 here; Flowers-102 and ISIC-7 in Fig. 10 .
Figure 10: Adaptive inference on Flowers-102 and ISIC-7 , continued from Fig. 9 .
Figure 11: Concept examples. Five concepts per dataset, drawn at random, one per row, each with the ten test images it fires on most strongly. Each image keeps its colour where the concept fires, scaled to its own strongest patch, and fades elsewhere. Imagenette, Dogs-10 and Machines-15 here; Mixed-20, Flowers-102 and ISIC-7 in Fig. 12 .
Figure 12: Concept examples , continued from Fig. 11 : Mixed-20, Flowers-102 and ISIC-7.
Subset
Classes
Dogs (10)
Golden retriever, Labrador retriever, Chesapeake Bay retriever, German shepherd, Rottweiler, Doberman, Boxer, Great Dane, Siberian husky, Eskimo dog
Animals: Lion, Tiger, African elephant, Zebra, Giant panda, Polar bear, Gorilla, Flamingo, Ostrich, Hippopotamus; Machines: Sports car, Warplane, Submarine, Steam locomotive, Forklift, Ambulance, Tank, Space shuttle, Fire engine, Motor scooter
Table 7: ImageNet-1k classes used in our three curated subsets.
Dataset
Prompt
Class → token
AG News
Article: "{text}" Which topic best fits this article: World, Sports, Business, or Technology? Answer with exactly one word.
World, Sports, Business, Sci/Tech → Technology
PubMed 20k RCT
Sentence from a medical research abstract: "{text}" Does this sentence state the study’s background, objective, methods, results, or conclusion? Answer with exactly one word.
Table 9: Hyperparameters used throughout AnyBottle ’s selection and training pipeline. Where vision and text differ, both values are given; a single value applies to both. For text, G2 = Gemma-2-2B MLP-teacher runs, G3 = Gemma-3-1B-it verbalizer runs. Accuracy thresholds are in percentage points.
German Research Center for Artificial Intelligence · Saarland University, Saarbrücken, Germany · 1German Research Center for Artificial Intelligence (DFKI), Saarbrücken, Germany +5