RepICL: Reusable In-Context Prediction Across Heterogeneous Representation Spaces
Organizations: SinoPac Holdings, Taipei, Taiwan
Abstract
Frozen representations are widely reused for downstream classification, yet each new task typically requires fitting a new predictor. We ask whether the few-shot prediction procedure itself can instead be learned once and reused across datasets and representation spaces. To study this question, we introduce RepShiftBench, comprising 1,218 encoder--dataset tasks across text, image, and audio, with separate evaluation of generalization to unseen datasets, unseen encoders, jointly unseen datasets and encoders, and unseen modalities. The benchmark exposes a substantial gap: Logistic Regression fitted independently on each episode outperforms every evaluated in-context learner across all settings. We introduce RepICL, a meta-trained in-context learner that canonicalizes each episode through episodic whitening before prediction. Its inductive variant, RepICL-I, surpasses Logistic Regression in all 12 benchmark settings, while RepICL-T substantially outperforms existing transductive methods. Ablations identify episodic whitening as the primary source of these gains, while showing that it is not a universally beneficial preprocessing step. Across both variants, the gains concentrate on queries for which simple support prototypes favor the wrong class or provide little separation between the true class and competing classes. Transduction provides its largest additional gains when limited support coverage gives a misleading view of class separation. Together, these results demonstrate that a shared few-shot prediction procedure can generalize beyond the representation spaces observed during training.
Figures & tables
| Adaptation Setting | Text-held-out fold | Audio-held-out fold | Image-held-out fold | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | D | E | J | M | D | E | J | M | D | E | J | M |
| Inductive | ||||||||||||
| ProtoHead | 56.11 | 72.55 | 61.05 | 64.41 | 68.64 | 71.72 | 71.28 | 48.41 | 54.08 | 65.83 | 56.22 | 80.88 |
| SimpleShot | 56.14 | 73.06 | 61.65 | 64.90 | 68.92 | 72.22 | 72.01 | 49.04 | 54.26 | 66.64 | 57.01 | 81.12 |
| SVM (RBF kernel) | 55.52 | 72.68 | 62.32 | 64.62 | 67.03 | 72.13 | 71.81 | 50.49 | 55.00 | 66.70 | 57.81 | 76.64 |
| Logistic regression | 60.19 | 75.33 | 65.74 | 69.62 | 72.37 | 76.05 | 75.56 | 54.46 | 58.35 | 71.56 | 61.90 | 82.38 |
| RepICL-I | RepICL-T | |||||||
| Dataset | Encoder | Joint | Modality | Dataset | Encoder | Joint | Modality | |
| Full model | 64.71 | 75.15 | 68.21 | 69.50 | 66.87 | 77.57 | 70.26 | 71.52 |
| Episodic whitening | ||||||||
| w/o Coord. Standardize | 7.06 | 4.77 | 4.81 | 6.84 | 11.41 | 9.37 | 9.78 | 10.24 |
| w/o PCA, trainable proj. | 6.38 | 9.90 | 9.42 | 9.52 | 7.55 | 11.61 | 10.80 | 11.22 |
| w/o PCA, fixed proj. | 6.37 | 9.51 | 9.46 | 9.13 | 8.02 | 11.98 | 11.03 | 11.24 |
| Logistic Regression | TabICL | TabPFN | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Text | Audio | Image | Text | Audio | Image | Text | Audio | Image | |
| Raw representations | 67.44 | 69.59 | 68.36 | 65.70 | 67.81 | 66.59 | 64.66 | 66.42 | 65.37 |
| + Episodic whitening | -0.32 | -0.34 | -0.22 | 1.05 | 1.42 | 1.11 | 5.91 | 6.27 | 6.22 |
| Margin | Queries (%) | Logistic regression | TabICL | TAIL | PT-MAP | RepICL-I | RepICL-T |
|---|---|---|---|---|---|---|---|
| 36.3 | 24.66 | 22.19 | 18.22 | 22.34 | 27.55 | 34.27 | |
| 63.7 | 92.78 | 91.59 | 86.39 | 91.15 | 92.33 | 91.86 | |
| 10.5 | 17.84 | 13.54 | 7.46 | 9.65 | 22.95 | 25.61 | |
| 8.2 | 19.92 | 16.74 | 12.25 | 15.45 | 23.50 | 28.72 | |
| 7.2 | 23.27 | 21.32 | 18.78 | 22.75 | 25.54 | 33.47 | |
| 5.9 | 29.80 | 29.82 | 28.37 | 33.50 | 30.51 | 42.08 |
| Queries (%) | Logistic regression | TabICL | TAIL | PT-MAP | RepICL-I | RepICL-T | T I | ||
|---|---|---|---|---|---|---|---|---|---|
| 22.9 | 21.24 | 18.61 | 14.53 | 14.79 | 23.89 | 27.18 | |||
| 13.5 | 30.48 | 28.29 | 24.50 | 35.17 | 33.75 | 46.32 | |||
| 10.0 | 78.47 | 75.81 | 66.17 | 68.38 | 76.19 | 72.13 | |||
| 53.7 | 95.44 | 94.53 | 90.15 | 95.39 | 95.34 | 95.53 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Category | #Classes (avail./all) | #Examples | Source |
| Text (seen) | ||||
| ArXivHierarchicalClusteringS2S | Academic-paper clustering | 169 / 169 | 627,949 | MTEB |
| Banking77Classification | Intent classification | 77 / 77 | 13,069 | MTEB |
| EmotionClassification | Emotion classification | 6 / 6 | 19,930 | MTEB |
| HWU64 | Intent classification | 57 / 64 | 10,030 | MTEB |
| MassiveIntentClassification | Intent classification | 54 / 60 | 16,521 | MTEB |
| Checkpoint | Release | Dim. | Ref. |
| Text (seen) | |||
| stanfordnlp/glove.6B.300d ∗ | 2014-10 | 300 | Pennington et al. (2014) |
| facebook/fasttext-en-vectors | 2018-02 | 300 | Grave et al. (2018) |
| google-bert/bert-base-uncased | 2018-10 | 768 | Devlin et al. (2019) |
| openai-community/gpt2 | 2019-02 | 768 | Radford et al. (2019) |
| sentence-transformers/all-mpnet-base-v2 | 2019-08 | 768 | Reimers and Gurevych (2019) |
| RepICL-I | RepICL-T | |
| Episode and whitening | ||
| Way / shot / queries per class | 5 / 5 / 10 | |
| Fit set | ||
| PCA rank | 24 | 74 |
| Standardization | ||
| Architecture | ||
| 1-shot | 5-shot | ||||||
|---|---|---|---|---|---|---|---|
| Setting | Paper | Reprod. | Gap | Paper | Reprod. | Gap | Per-dataset MAE |
| In-domain (4 datasets) | 90.10 | 89.73 | -0.37 | 95.66 | 95.10 | -0.56 | 1.03 |
| Cross-modal (2 datasets) | 62.75 | 64.31 | +1.56 | 72.48 | 72.53 | +0.05 | 1.00 |
| Cross-domain (8 datasets) | 72.46 | 73.25 | +0.79 | 82.32 | 82.93 | +0.61 | 3.87 |
| Text-held | Audio-held | Image-held | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | D | E | J | M | D | E | J | M | D | E | J | M |
| TAIL (unit-norm) | 26.86 | 34.48 | 29.01 | 28.54 | 30.34 | 34.67 | 32.42 | 24.27 | 25.82 | 29.15 | 26.14 | 38.22 |
| TAIL (norm 13) | 52.79 | 71.96 | 61.05 | 62.23 | 66.06 | 71.13 | 71.07 | 45.67 | 50.33 | 64.93 | 55.55 | 79.31 |
| TAIL | 36.96 | 48.48 | 38.86 | 33.07 | 55.43 | 62.33 | 60.21 | 36.75 | 21.81 | 20.16 | 20.08 | 20.25 |
| RepICL-T | 63.55 | 77.99 | 68.29 | 73.26 | 75.27 | 78.13 | 77.47 | 55.62 | 61.79 | 76.58 | 65.02 | 85.69 |
| TAIL input scale | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Encoder | Modality | # tasks | Raw norm | Raw | 1 | 3 | 6 | 13 | 26 | 52 | Ours |
| DINOv3 | image | 19 | 15.61 | 87.50 | 50.65 | 83.50 | 87.00 | 87.62 | 84.75 | 63.15 | 92.34 |
| USAD2 | audio | 21 | 163.86 | 6.59 | 21.98 | 31.21 | 44.85 | 51.00 | 40.25 | 23.77 | 63.94 |
| E5-large-v2 | text | 19 | 23.03 | 65.48 | 24.40 | 49.02 | 66.94 | 72.44 | 61.37 | 29.09 | 82.82 |
| Text-held | Audio-held | Image-held | ||||||||||
| Width | D | E | J | M | D | E | J | M | D | E | J | M |
| TAIL (norm 13) | ||||||||||||
| native | 49.58 | 76.94 | 68.16 | 57.94 | 63.47 | 86.18 | 86.36 | 42.40 | 46.61 | 64.72 | 55.74 | 82.01 |
| proj. | 74.80 | 66.91 | 54.67 | 67.29 | 70.40 | 60.81 | 61.78 | 64.20 | 67.89 | 65.02 | 55.46 | 75.02 |
| 25.22 | 10.03 | 13.49 | 9.35 | 6.93 | 25.37 | 24.58 | 21.80 | 21.28 | 0.30 | 0.28 | 6.99 | |
| RepICL-T | ||||||||||||
| Linear | TAIL | Ours | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Dataset | published | TAIL recipe | ours | GPICL † | CAML † | published | reproduced | RepICL-I | RepICL-T |
| CIFAR-FS | 91.86 | 91.56 | 92.87 | 91.20 | 91.69 | 94.55 | 94.49 | 93.75 | 96.75 |
| miniImageNet | 97.65 | 97.66 | 98.56 | 99.44 | 99.29 | 99.63 | 99.40 | 98.71 | 99.77 |
| tieredImageNet | 95.30 | 95.12 | 96.50 | 98.18 | 97.98 | 98.67 | 98.05 | 96.72 | 98.43 |
| Pascal VOC | 83.57 | 83.49 | 85.70 | 87.46 | 87.87 | 89.78 | 88.33 | 86.13 | 89.55 |
| Aircraft | 92.12 | 91.82 | 93.65 | 90.83 | 93.09 | 95.01 | 94.31 | 94.12 | 96.66 |
| Text-held-out fold | Audio-held-out fold | Image-held-out fold | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| D | E | J | M | D | E | J | M | D | E | J | M | |
| RepICL-I | 61.41 | 75.85 | 66.22 | 70.10 | 73.52 | 76.79 | 75.86 | 55.14 | 59.19 | 72.82 | 62.56 | 83.26 |
| Episodic whitening | ||||||||||||
| w/o Coord. Standardize | 8.18 | 3.64 | 5.19 | 7.88 | 5.03 | 4.14 | 3.40 | 9.38 | 7.98 | 6.53 | 5.85 | 3.25 |
| w/o PCA, trainable proj. | 5.28 | 5.94 | 7.15 | 12.07 | 8.61 | 11.38 | 10.76 | 9.93 | 5.25 | 12.39 | 10.34 | 6.55 |
| w/o PCA, fixed proj. | 5.44 | 6.05 | 8.37 | 10.96 | 8.33 | 10.60 | 10.06 | 10.01 | 5.33 | 11.88 | 9.96 | 6.42 |
| Backbone | Logistic regression | SimpleShot | LaplacianShot | -TIM | PT-MAP (signed-power) |
|---|---|---|---|---|---|
| ResNet-18 | 77.17 | 77.05 | 77.88 | 77.83 | 80.47 |
| ResNet-34 | 76.75 | 76.71 | 77.20 | 77.39 | 79.11 |
| WRN-28-10 | 58.68 | 59.20 | 60.30 | 62.12 | 65.18 |
| Modality | Encoder class | Num. of Enc. | Logistic regression | SimpleShot | LaplacianShot | -TIM | PT-MAP |
|---|---|---|---|---|---|---|---|
| Image | CNN-like | 7 | 80.44 | 80.96 | 81.42 | 81.43 | 83.23 |
| ViT / vision-SSL | 6 | 76.82 | 73.18 | 72.67 | 75.45 | 74.93 | |
| CLIP-like contrastive | 3 | 92.50 | 91.33 | 91.37 | 92.20 | 93.09 | |
| Multimodal embedding | 2 | 90.65 | 90.22 | 90.31 | 90.50 | 91.35 | |
| All | 18 | 82.38 | 81.12 | 81.15 | 82.24 | 83.01 | |
| Text | Static word embeddings | 2 | 61.35 | 55.28 | 55.30 | 59.41 | 58.92 |
| Encoder | Modality | Params | Tasks | Frozen | LoRA | RepICL-I | RepICL-T | Logistic regression |
|---|---|---|---|---|---|---|---|---|
| AST audioset-base | audio | 86.2M | 20 | 46.6 | 55.7 | 55.8 | 56.5 | 54.8 |
| HuBERT-large | audio | 315.4M | 20 | 46.9 | 51.1 | 57.2 | 57.7 | 56.3 |
| whisper-large | audio | 637.0M | 20 | 40.6 | 61.5 | 64.4 | 64.8 | 63.5 |
| all-mpnet-base-v2 | text | 109.5M | 19 | 75.2 | 79.5 | 77.17 | 80.72 | 77.16 |
| embedding-gemma-300m | text | 302.9M | 19 | 78.2 | 82.3 | 80.95 | 84.25 | 80.99 |
| stella-1.5B-v5 | text | 1.54B | 19 | 80.1 | 85.4 | 81.17 | 84.01 | 80.19 |