CPATTA: Conformal Supervision Allocation For Active Test-Time Adaptation
Organizations: Computer Science and Engineering, University of California San Diego · New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences · College of Computer Science and Electronic Engineering, Hunan University
Abstract
Active Test-Time Adaptation (ATTA) improves model robustness under domain shift by selectively querying human annotations at deployment, but existing methods use heuristic uncertainty measures and suffer from low data selection efficiency, wasting human annotation budget. We propose Conformal Prediction Active TTA (CPATTA), which first brings principled, conformal uncertainty with coverage-aware online calibration into ATTA. CPATTA employs smoothed conformal scores with a top- certainty measure, an online weight-update algorithm driven by pseudo coverage, a domain-shift detector that adapts human supervision, and a staged update scheme that balances human-labeled and model-labeled data. Extensive experiments demonstrate that CPATTA consistently outperforms the state-of-the-art ATTA methods by around 5% in accuracy.
Figures & tables
| PACS Dataset | VLCS Dataset | ||||||||||||||||||
| Real-Time Accuracy | Post-Adaptation Accuracy | Real-Time Accuracy | Post-Adaptation Accuracy | ||||||||||||||||
| Methods | A | C | S | Acc | P | A | C | S | Acc | L | S | V | Acc | C | L | S | V | Acc | |
| Tent [ 4 ] | 67.53 | 69.16 | 66.94 | 67.71 | 93.53 | 68.46 | 71.25 | 73.12 | 75.14 | 46.26 | 41.22 | 55.72 | 47.91 | 90.11 | 52.80 | 45.70 | 57.08 | 56.90 | |
| CoTTA [ 6 ] | 65.72 | 65.91 | 65.28 | 65.57 | 83.71 | 60.60 | 63.48 | 71.24 | 69.32 | 45.84 | 39.49 | 54.92 | 46.89 | 47.32 | 52.88 | 39.85 | 50.24 | 47.33 | |
| SAR [ 7 ] | 65.53 | 66.13 | 63.43 | 64.71 | 94.19 | 65.63 | 67.02 | 65.11 | 70.53 | 42.69 | 37.51 | 50.74 | 43.80 | 90.47 | 44.90 | 38.76 | 50.71 | 50.86 | |
| TTA | EATA [ 5 ] | 67.43 | 68.26 | 67.24 | 67.57 | 94.85 | 66.31 | 66.72 | 70.60 | 72.86 | 42.95 | 37.26 | 51.36 | 44.01 | 91.24 | 43.93 | 37.72 | 51.74 | 50.73 |
| PACS | VLCS | Tiny-ImageNet-C | ||||
|---|---|---|---|---|---|---|
| Method | ||||||
| SimATTA [ 8 ] | 47.60 | 67.60 | 57.04 | 64.78 | 79.25 | 50.00 |
| CEMA [ 9 ] | 24.91 | N/A | 44.38 | N/A | 75.82 | N/A |
| EATTA [ 10 ] | 53.71 | 90.71 | 58.38 | 63.60 | 62.90 | 32.31 |
| CPATTA( ) | 66.30 | 93.41 | 62.89 | 81.91 | 90.45 | 48.14 |
| CPATTA( ) | 65.94 | 93.65 | 61.86 | 85.11 | 91.00 | 47.83 |
| Realtime Accuracy | Post-Adaptation Accuracy | |||||
|---|---|---|---|---|---|---|
| Variants | PACS | VLCS | Tiny | PACS | VLCS | Tiny |
| Random Selection | 68.16 | 55.58 | 31.98 | 78.37 | 70.77 | 40.96 |
| w/o adaptive weighting (Gometric Decaying) | 74.31 | 63.29 | 29.22 | 83.37 | 75.14 | 37.14 |
| Fixed hard-set size | 72.02 | 58.34 | 29.38 | 83.09 | 68.03 | 37.58 |
| w/o DSS-triggered budget adjustment | 68.50 | 64.64 | 30.69 | 84.83 | 74.91 | 38.45 |
| Joint human–pseudo update | 69.53 | 57.01 | 28.32 | 84.70 | 72.78 | 36.38 |
| Real-Time | Post-Adapt | ||||
|---|---|---|---|---|---|
| Method | A | C | S | Acc | Acc |
| SimATTA with Replay | 74.32 | 50.64 | 72.77 | 66.92 | 79.14 |
| CEMA with Replay | 65.58 | 68.64 | 58.28 | 63.00 | 75.38 |
| EATTA with Replay | 69.97 | 69.07 | 65.59 | 67.65 | 74.85 |
| CPATTA( ) | 71.39 | 68.05 | 76.18 | 72.71 | 85.25 |
| CPATTA( ) | 73.10 | 70.01 | 79.51 | 75.26 | 87.13 |
| Metric | Real-Time CP | Pretrained CP | ||
| PACS | VLCS | PACS | VLCS | |
| Corr. | 0.9978 | 0.9997 | 0.9988 | 0.9853 |