Scaling Full Conformal Image Classifiers
Organizations: Computer Vision Lab, ETH Zürich
Abstract
Conformal prediction provides set-valued predictions with distribution-free coverage guarantees, making it attractive for high-stakes image classification. However, split conformal prediction is data-inefficient, while full conformal prediction (FCP), despite its stronger statistical efficiency, is computationally prohibitive at scale because it requires candidate-specific model refits at test time. We address this limitation by leveraging zero-shot vision-language models (VLMs) to guide scalable FCP in large label spaces. We introduce Targeted Full Conformal Prediction (T-FCP), which uses a lightweight inductive conformal predictor to prune unlikely labels and applies FCP only to the remaining candidates, reducing computation while retaining the formal guarantee of the combined conformal procedure. We further propose Stabilized Online LDA (SO-LDA), an efficient VLM adaptation solver based on rank-one inverse-covariance updates. Across multiple benchmarks, including ImageNet, T-FCP enables practical full-conformal image classification with modest test-time overhead, yielding efficient prediction sets and more stable empirical coverage than split conformal alternatives.
Figures & tables
| Cov. ( ) | Set size | Cov. ( ) | Set size | ||||||||||||
| Acc. | Avg. | %Valid | Mean | Med. | %Sing. | Avg. | %Valid | Mean | Med. | %Sing. | |||||
| ICP | ZS | 65.7 | 90.1 | 2.0 | 77.8 | 4.8 | 4.5 | 37.3 | 95.2 | 1.5 | 84.0 | 7.9 | 7.3 | 28.2 | |
| ICP-T | OT Silva-Rodríguez et al. (2025b) | 68.6 | 90.6 | 2.0 | 82.3 | 3.9 | 3.7 | 42.4 | 95.4 | 1.4 | 86.8 | 6.1 | 5.6 | 32.6 | |
| ICP-T | TIM Silva-Rodríguez et al. (2025a) | 71.5 | 90.6 | 2.0 | 80.5 | 3.6 | 3.5 | 45.2 | 95.4 | 1.4 | 87.2 | 5.6 | 5.3 | 34.3 | |
| SCP | GD | 79.9 | 90.7 | 2.7 | 77.1 | 2.2 | 2.1 | 64.0 | 95.5 | 1.9 | 79.5 | 3.3 | 3.0 | 51.9 | |
| FCP | SO-LDA | 80.3 | 90.1 | 2.2 | 77.6 | 2.2 | 1.8 | 69.3 | 95.2 | 1.6 | 80.5 | 3.3 | 2.6 | 57.4 | |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Cov. ( ) Set size Acc. Avg. %Valid Mean %Sing. SO-LDA 81.0 90.8 2.1 80.0 2.3 70.3 Non-online 81.0 90.8 2.1 80.0 2.3 70.2 Ridge 78.8 90.7 2.3 76.0 2.4 64.6 Unstable 78.8 90.4 2.3 67.5 2.6 62.3 | Cov. ( ) Set size Acc. Avg. %Valid Mean %Sing. SO-LDA 76.4 88.2 4.2 27.0 3.1 61.6 Non-online 76.4 88.3 4.2 27.5 3.2 61.3 Ridge 74.1 88.0 4.3 26.0 2.9 58.4 Unstable 70.1 86.0 5.3 13.0 4.6 45.9 |
| (a) | (b) |
| Dataset | Classes | Splits | / | Task description | ||
|---|---|---|---|---|---|---|
| Train | Val | Test | ||||
| EuroSAT [ 19 ] | 10 | 13,500 | 5,400 | 8,100 | Satellite image classification. | |
| OxfordPets [ 38 ] | 37 | 2,944 | 736 | 3,669 | Pets classification. | |
| DTD [ 7 ] | 47 | 2,820 | 1,128 | 1,692 | Textures classification. | |
| FGVCAircraft [ 33 ] | 100 | 3,334 | 3,333 | 3,333 | Aircraft classification. | |
| Caltech101 [ 12 ] | 100 | 4,128 | 1,649 | 2,465 | Natural objects classification. | |
| Cov. ( ) | Set size | Cov. ( ) | Set size | ||||||||||||
| Acc. | Avg. | %Valid | Mean | Med. | %Sing. | Avg. | %Valid | Mean | Med. | %Sing. | |||||
| SCP | SO-LDA | 79.5 | 90.8 | 3.5 | 71.3 | 2.7 | 2.3 | 69.4 | 95.6 | 2.4 | 77.0 | 4.1 | 3.3 | 58.4 | |
| Cov. ( ) | Set size | Cov. ( ) | Set size | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Acc. | Avg. | %Valid | Mean | Med. | %Sing. | Avg. | %Valid | Mean | Med. | %Sing. | ||||||
| EuroSAT | ICP | ZS | 48.4 | 90.4 | 6.1 | 62.0 | 4.2 | 4.4 | 9.4 | 95.6 | 4.8 | 66.0 | 5.2 | 5.3 | 6.6 | |
| SCP | GD | 83.4 | 89.8 | 6.5 | 50.0 | 1.4 | 1.1 | 73.4 | 96.1 | 4.7 | 76.0 | 2.2 | 1.8 | 43.8 | ||
| T-FCP | SO-LDA | 83.0 | 90.2 | 5.3 | 58.0 | 1.4 | 1.0 | 71.4 | 95.9 | 3.9 | 90.0 | 2.0 | 1.9 | 42.0 | ||
| Pets | ICP | ZS | 88.8 | 89.9 | 3.2 | 54.0 | 1.0 | 1.0 | 92.9 | 95.1 | 2.0 | 68.0 | 1.2 | 1.0 | 80.4 | |
| SCP | GD | 92.3 | 90.0 | 4.9 | 65.0 | 1.0 | 1.0 | 94.8 | 95.2 | 3.3 | 66.0 | 1.1 | 1.0 | 91.0 | ||
| Cov. ( ) | Set size | Cov. ( ) | Set size | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Avg. | %Valid | CCV | Mean | Med. | %Sing. | Avg. | %Valid | CCV | Mean | Med. | %Sing. | ||||
| ICP | 90.4 | 2.2 | 81.5 | 8.2 | 6.3 | 4.9 | 34.4 | 95.3 | 1.5 | 85.0 | 5.0 | 10.1 | 8.0 | 27.1 | |
| SCP | 90.1 | 3.1 | 67.5 | 6.3 | 3.0 | 2.2 | 54.0 | 95.1 | 2.4 | 70.0 | 5.0 | 4.3 | 3.1 | 46.9 | |
| T-FCP | 90.4 | 1.7 | 84.0 | 6.4 | 3.3 | 2.3 | 52.6 | 95.3 | 1.4 | 89.0 | 4.9 | 4.5 | 2.9 | 46.9 | |
| Cov. ( ) | Set size | |||||
|---|---|---|---|---|---|---|
| Avg. | %Valid | Mean | ||||
| ICP | 95.2 | 1.3 | 80.0 | 12.3 | ||
| SCP | 95.0 | 1.2 | 70.0 | 11.7 | ||
| T-FCP | 95.1 | 1.2 | 80.0 | 10.6 | ||