In-context retrieval (ICR) is a retrieval-based form of in-context learning (ICL) in which demonstrations are retrieved from a source database based on similarity to the query, rather than sampled independently. In this work, we formulate ICR as a type of domain adaptation problem, where the source distribution P of the database may differ from the target distribution Q of the test query-label pair. We investigate the performance of ICR under a flexible class of distributional shifts that substantially extends prior work \citep{li2024fine,guo2025retrieval}, and establish theoretical guarantees that quantify the benefits and pitfalls of this learning paradigm. Our theory is verified by experiments on synthetic and language tasks.
Figures & tables
Figure 1 : Standard ICL vs ICR.
Figure 2Figure 3
Figure 6 : Test error vs. context size n on the sentiment dataset (Qwen3.5-4B; 878 test annotators, mean over 5 seeds). Blue : ICL with i.i.d. demonstrations from the annotator’s own data. Yellow → red : ICR under four corpus designs with increasing posterior drift.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : Test error on Q vs. number of demonstrations n , with P=Q (no domain shift). Color : ICL (blue) vs. ICR (orange). Line style : TabPFN (solid, •) vs. k -NN (dashed, □ ). Corpus size m=20,000 , ntest=400 . Results averaged over 5 independent runs.
n
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
4
0.395±0.094
0.082±0.018
0.350±0.114
0.085±0.011
16
0.229±0.073
0.048±0.006
0.272±0.026
0.041±0.010
32
0.165±0.039
0.048±0.008
0.230±0.053
0.039±0.012
128
0.073±0.017
0.031±0.004
0.123±0.028
0.027±0.009
256
0.038±0.010
0.012±0.008
0.102±0.027
0.027±0.009
Appendix
Table 1 : Experiment 1, d=8 : test error on Q , we report mean ± std over 5 trials.
n
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
4
0.469±0.054
0.139±0.022
0.397±0.071
0.149±0.026
16
0.382±0.086
0.101±0.011
0.374±0.039
0.089±0.012
32
0.266±0.023
0.100±0.017
0.327±0.058
0.066±0.013
128
0.099±0.012
0.080±0.020
0.196±0.040
0.045±0.012
256
0.067±0.012
0.062±0.010
0.138±0.017
0.041±0.009
Appendix
Table 2 : Experiment 1, d=16 : test error on Q , we report mean ± std over 5 trials.
Figure 8 : Test error on Q vs. posterior drift parameter ε , with n=4 fixed. Color : ICL (blue) vs. ICR (orange). Line style : TabPFN (solid) vs. k -NN (dashed). ICL baselines (horizontal) are the means over 5 runs × 11 ε values (since ε only affects the ICR corpus; not affecting ICL). ICR mean error grows monotonically with ε . Results averaged over 5 independent runs.
ε
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
0.0
0.394±0.094
0.082±0.017
0.350±0.114
0.085±0.011
0.1
0.411±0.060
0.127±0.018
0.354±0.071
0.128±0.019
0.2
0.404±0.065
0.233±0.033
0.373±0.047
0.225±0.043
0.3
0.458±0.074
0.328±0.042
0.389±0.137
0.318±0.045
0.4
0.411±0.075
0.400±0.032
0.333±0.057
0.405±0.039
0.5
0.414±0.092
0.496±0.041
0.380±0.085
0.506±0.044
Appendix
Table 3 : Experiment 2, d=8 , n=4 : test error on Q vs. posterior drift ε , we report mean ± std over 5 trials.
ε
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
0.0
0.468±0.055
0.139±0.023
0.397±0.071
0.149±0.026
0.1
0.447±0.072
0.185±0.032
0.440±0.064
0.184±0.024
0.2
0.465±0.048
0.264±0.033
0.451±0.030
0.267±0.025
0.3
0.477±0.037
0.336±0.026
0.430±0.071
0.353±0.027
0.4
0.468±0.027
0.416±0.018
0.429±0.033
0.428±0.020
0.5
0.480±0.058
0.496±0.034
0.468±0.034
0.491±0.024
Appendix
Table 4 : Experiment 2, d=16 , n=4 : test error on Q vs. posterior drift ε , we report mean ± std over 5 trials.
Figure 9 : Examples of the Weierstrass class-probability function ηα(x) for different values of the smoothness parameter α .
Figure 10 : Target excess risk vs. number of demonstrations n on the 1-D Weierstrass family ( P=Q ; β=1 ; oracle kα ). Color : ICL (blue) vs. ICR (orange). Line style : TabPFN (solid, •) vs. k -NN (dashed, □ ). Both axes are log-scaled. Corpus size m=105 , ntest=400 ; each curve is the mean over 5 seeds (42–46).
n
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
4
0.151±0.036
0.066±0.006
0.147±0.040
0.058±0.005
16
0.151±0.013
0.040±0.005
0.129±0.046
0.023±0.002
32
0.134±0.023
0.024±0.003
0.092±0.019
0.014±0.001
64
0.103±0.035
0.018±0.004
0.058±0.022
0.008±0.001
128
0.059±0.009
0.012±0.002
0.029±0.013
0.005±0.001
256
0.036±0.020
0.007±0.001
0.020±0.011
0.002±0.001
Appendix
Table 5 : Weierstrass sample-size sweep, α=1.0 : target excess risk, mean ± std over 5 seeds.
n
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
4
0.160±0.028
0.077±0.006
0.157±0.027
0.067±0.008
16
0.154±0.015
0.043±0.006
0.121±0.060
0.024±0.002
32
0.127±0.025
0.027±0.004
0.096±0.017
0.014±0.002
64
0.107±0.034
0.019±0.005
0.071±0.036
0.008±0.001
128
0.069±0.028
0.011±0.003
0.055±0.018
0.003±0.001
256
0.037±0.024
0.006±0.001
0.043±0.017
0.002±0.000
Appendix
Table 6 : Weierstrass sample-size sweep, α=0.5 : target excess risk, mean ± std over 5 seeds.
Figure 11 : Target excess risk vs. posterior drift ε=QX(B) on the Weierstrass family at α=1.0 (smooth η ), for n∈{16,512} over 5 seeds. B is a union of N=40 periodic sub-brackets, each covering the leftmost ε -fraction of one width- 2/N interval of [−1,1] , so B samples every macro-cycle of ηQ uniformly and QX(B)=ε holds exactly. Color : ICL (blue) vs. ICR (orange). Line style : TabPFN (solid, •) vs. k -NN (dashed, □ ). ICL curves are exactly flat because ICL uses Q labels and does not see the drift; ICR curves rise almost linearly in ε across the whole range, saturating near 1−2RQ∗≈0.306 at ε=1 . At small n (left, n=16 ) ICR beats ICL up to ε≈0.4 – 0.5 ; at large n (right, n=512 ) ICR is roughly an order of magnitude below ICL at ε=0 (variance-limited regime), but ICL overtakes it already by ε≈0.08 . In both regimes the ICR bias tracks ε almost linearly, matching the ε+(λ/κ)(RP−RP∗) reduction in Lemma 1 .
ε
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
0.0
0.151±0.013
0.040±0.005
0.129±0.046
0.023±0.002
0.1
0.151±0.013
0.063±0.003
0.129±0.046
0.049±0.002
0.2
0.151±0.013
0.083±0.003
0.129±0.046
0.073±0.003
0.3
0.151±0.013
0.105±0.002
0.129±0.046
0.098±0.002
0.4
0.151±0.013
0.129±0.004
0.129±0.046
0.124±0.001
0.5
0.151±0.013
0.149±0.007
0.129±0.046
0.151±0.002
Appendix
Table 7 : Experiment 2, α=1.0 , n=16 : target excess risk vs. posterior drift ε=QX(B) , mean ± std over 5 seeds (Weierstrass family, A=0.24 , b=0.85 , f=2 , J=12 ; oracle kα=round_odd(0.30⋅n2α/(2α+1)) ; mcorpus=105 ; ntest=400 ). Region B is the leftmost ε -fraction of each of N=40 equal-width periodic brackets partitioning [−1,1] , so QX(B)=ε exactly and B samples every macro-cycle of ηQ uniformly; ηP=1−ηQ on B ; ICL uses Q labels (so ICL columns are ε -constant).
ε
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
0.0
0.019±0.014
0.003±0.001
0.018±0.008
0.001±0.000
0.1
0.019±0.014
0.032±0.001
0.018±0.008
0.020±0.005
0.2
0.019±0.014
0.060±0.001
0.018±0.008
0.060±0.001
0.3
0.019±0.014
0.090±0.001
0.018±0.008
0.090±0.001
0.4
0.019±0.014
0.120±0.001
0.018±0.008
0.120±0.001
0.5
0.019±0.014
0.151±0.001
0.018±0.008
0.150±0.001
Appendix
Table 8 : Experiment 2, α=1.0 , n=512 : target excess risk vs. posterior drift ε=QX(B) , mean ± std over 5 seeds (same setup as Table 7 ).
Figure 12 : Target excess risk vs. smoothness exponent α on the Weierstrass family at n=256 , with α∈{1,10−1,10−2,10−3} (log-spaced, 5 seeds). Color : ICL (blue) vs. ICR (orange). Line style : TabPFN (solid, •) vs. k -NN (dashed, □ ). The horizontal axis is logarithmic; the α -axis is reversed so rougher ηα is on the right. Both k -NN-ICL and TabPFN-ICL trend upward as α decreases, matching the n−2α/(2α+1) prediction of Theorem 1 ; the k -NN-ICL curve saturates near α→0 at the finite- n noise floor. ICR sits about an order of magnitude below ICL and is essentially flat in α , consistent with the α -independent n−(β+1)/2 rate of Theorem 1 ; the ICL/ICR gap widens as α→0 .
Figure 13 : Excess error on Q versus corpus size m , with n fixed and αP=1 . Color : ICL (blue) versus ICR (orange). Line style : TabPFN (solid, •) vs. k -NN (dashed, □ ). Per (n,seed) we draw one corpus of size mmax and use nested prefixes Dm=Dmmax[:m] ; the ICL demo set is drawn once per n from that corpus and reused across every m . Consequently ICL is exactly constant along the m axis for each seed, and the horizontal ICL baselines are the across-seed means of those per-seed constants. Both ICR curves collapse sharply around the reference scale and remain near their asymptotic value thereafter. The dotted grey line marks the unit-constant reference n1+1/(2αP)=n3/2 ( 512 for n=64 , ≈1448 for n=128 ); the actual saturation threshold is C⋆n1+1/(2αP) for some unknown constant. All 5 seeds (42–46). Setup: 1-D Weierstrass η -family with A=0.24 , b=0.85 , f=2 , J=12 , oracle kα=round_odd(0.30⋅n2α/(2α+1)) , linspace ntest=400 , and joint demonstrations.
n
m
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
64
64
0.097±0.033
0.133±0.019
0.079±0.016
0.145±0.042
128
0.073±0.022
0.140±0.028
256
0.038±0.011
0.109±0.032
512
0.020±0.005
0.019±0.009
1024
0.021±0.006
0.009±0.003
2048
0.021±0.006
0.008±0.004
Appendix
Table 10 : Corpus-size sweep: excess error on Q , mean ± sample standard deviation over five seeds. ICL entries are constant along the m axis by construction (see Fig. 13 caption) and are shown once per n ; ICR entries vary per m .
Figure 14 : Test error vs. number of demonstrations n on the age-biased sentiment dataset (878 test annotators, mean over 5 seeds), for Qwen3.5-4B (left) and Qwen3.5-27B (right). Blue : ICL with demonstrations sampled uniformly without replacement from the annotator’s own data. Yellow → red : ICR under four corpus designs with increasing posterior drift. The stronger model improves ICR only where the corpus is aligned with the target; under severe and extreme drift its error is essentially unchanged—the drift-induced bias is irreducible. Per-cell mean values are in Tables 11 and 12 .
Figure 15 : Test error vs. number of demonstrations n on the age-biased sentiment dataset (878 test annotators, Qwen3.5-4B). Each point is the mean over 5 independent runs (seeds 42–46). Blue : Qwen-ICL with demonstrations sampled uniformly without replacement from the annotator’s own training data. Yellow → red : Qwen-ICR under four levels of posterior drift, as a result of expanding the source corpus: no drift ( P=Q , self only), mild (self + 100 random others), severe (all 1,481 annotators), and extreme (all other annotators, self excluded). ICR sharply outperforms ICL when the source corpus is well-aligned with the target distribution and degrades as posterior drift increases. Uncertainty estimates are reported in Table 11 .
n
Qwen-ICL
ICR ( P=Q )
ICR (mild)
ICR (severe)
ICR (extreme)
4
0.538±0.007
0.263±0.000
0.265±0.009
0.535±0.004
0.633±0.004
8
0.481±0.008
0.169±0.002
0.225±0.009
0.504±0.002
0.628±0.007
16
0.348±0.011
0.099±0.002
0.162±0.012
0.491±0.003
0.615±0.006
32
0.117±0.004
0.072±0.001
0.142±0.010
0.480±0.002
0.619±0.008
Appendix
Table 11 : Sentiment classification with Qwen3.5-4B: test error (mean ± std over 5 seeds, 878 test annotators) vs. number of demonstrations n across the four corpus designs.
Figure 16 : Test error vs. number of demonstrations n on the age-biased sentiment dataset (878 test annotators, Qwen3.5-27B). Each point is the mean over 5 independent runs (seeds 42–46). Blue : Qwen-ICL with demonstrations sampled uniformly without replacement from the annotator’s own training data. Yellow → red : Qwen-ICR under four levels of posterior drift, as a result of expanding the source corpus: no drift ( P=Q , self only), mild (self + 100 random others), severe (all 1,481 annotators), and extreme (all other annotators, self excluded). Compared to Qwen3.5-4B ( Figure 15 ), the stronger model lowers the ICR error under no/mild drift but does not remove the irreducible bias under severe or extreme drift, matching the theory. Uncertainty estimates are reported in Table 12 .
n
Qwen-ICL
ICR ( P=Q )
ICR (mild)
ICR (severe)
ICR (extreme)
4
0.514±0.005
0.118±0.001
0.165±0.003
0.522±0.008
0.619±0.006
8
0.459±0.005
0.094±0.000
0.154±0.006
0.507±0.003
0.620±0.005
16
0.341±0.006
0.083±0.000
0.155±0.005
0.506±0.003
0.606±0.008
32
0.132±0.007
0.097±0.001
0.163±0.005
0.502±0.003
0.612±0.007
Appendix
Table 12 : Sentiment classification with Qwen3.5-27B: test error (mean ± std over 5 seeds, 878 test annotators) vs. number of demonstrations n across the four corpus designs.
Univ. Grenoble Alpes, CNRS, Inria, Grenoble INP, LJK, 38000 Grenoble, France · Work done during an internship at Morgan Stanley Machine Learning Research. · Machine Learning Research, Morgan Stanley, New York, USA