In-context retrieval (ICR) is a retrieval-based form of in-context learning (ICL) in which demonstrations are retrieved from a source database based on similarity to the query, rather than sampled independently. In this work, we formulate ICR as a type of domain adaptation problem, where the source distribution P of the database may differ from the target distribution Q of the test query-label pair. We investigate the performance of ICR under a flexible class of distributional shifts that substantially extends prior work \citep{li2024fine,guo2025retrieval}, and establish theoretical guarantees that quantify the benefits and pitfalls of this learning paradigm. Our theory is verified by experiments on synthetic and language tasks.
Figures & tables
Figure 1 : Standard ICL vs ICR.
Figure 2Figure 3
Figure 6 : Test error vs. context size n on the sentiment dataset (Qwen3.5-4B; 878 test annotators, mean over 5 seeds). Blue : ICL with i.i.d. demonstrations from the annotator’s own data. Yellow → red : ICR under four corpus designs with increasing posterior drift.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : Test error on Q vs. number of demonstrations n , with P=Q (no domain shift). Color : ICL (blue) vs. ICR (orange). Line style : TabPFN (solid, •) vs. k -NN (dashed, □ ). Corpus size m=20,000 , ntest=400 . Results averaged over 5 independent runs.
n
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
4
0.395±0.094
0.082±0.018
0.350±0.114
0.085±0.011
16
0.229±0.073
0.048±0.006
0.272±0.026
0.041±0.010
32
0.165±0.039
0.048±0.008
0.230±0.053
0.039±0.012
128
0.073±0.017
0.031±0.004
0.123±0.028
0.027±0.009
256
0.038±0.010
0.012±0.008
0.102±0.027
0.027±0.009
Appendix
Table 1 : Experiment 1, d=8 : test error on Q , we report mean ± std over 5 trials.
n
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
4
0.469±0.054
0.139±0.022
0.397±0.071
0.149±0.026
16
0.382±0.086
0.101±0.011
0.374±0.039
0.089±0.012
32
0.266±0.023
0.100±0.017
0.327±0.058
0.066±0.013
128
0.099±0.012
0.080±0.020
0.196±0.040
0.045±0.012
256
0.067±0.012
0.062±0.010
0.138±0.017
0.041±0.009
Appendix
Table 2 : Experiment 1, d=16 : test error on Q , we report mean ± std over 5 trials.
Figure 8 : Test error on Q vs. posterior drift parameter ε , with n=4 fixed. Color : ICL (blue) vs. ICR (orange). Line style : TabPFN (solid) vs. k -NN (dashed). ICL baselines (horizontal) are the means over 5 runs × 11 ε values (since ε only affects the ICR corpus; not affecting ICL). ICR mean error grows monotonically with ε . Results averaged over 5 independent runs.
ε
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
0.0
0.394±0.094
0.082±0.017
0.350±0.114
0.085±0.011
0.1
0.411±0.060
0.127±0.018
0.354±0.071
0.128±0.019
0.2
0.404±0.065
0.233±0.033
0.373±0.047
0.225±0.043
0.3
0.458±0.074
0.328±0.042
0.389±0.137
0.318±0.045
0.4
0.411±0.075
0.400±0.032
0.333±0.057
0.405±0.039
0.5
0.414±0.092
0.496±0.041
0.380±0.085
0.506±0.044
Appendix
Table 3 : Experiment 2, d=8 , n=4 : test error on Q vs. posterior drift ε , we report mean ± std over 5 trials.
ε
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
0.0
0.468±0.055
0.139±0.023
0.397±0.071
0.149±0.026
0.1
0.447±0.072
0.185±0.032
0.440±0.064
0.184±0.024
0.2
0.465±0.048
0.264±0.033
0.451±0.030
0.267±0.025
0.3
0.477±0.037
0.336±0.026
0.430±0.071
0.353±0.027
0.4
0.468±0.027
0.416±0.018
0.429±0.033
0.428±0.020
0.5
0.480±0.058
0.496±0.034
0.468±0.034
0.491±0.024
Appendix
Table 4 : Experiment 2, d=16 , n=4 : test error on Q vs. posterior drift ε , we report mean ± std over 5 trials.
Figure 9 : Examples of the Weierstrass class-probability function ηα(x) for different values of the smoothness parameter α .
Figure 10 : Target excess risk vs. number of demonstrations n on the 1-D Weierstrass family ( P=Q ; β=1 ; oracle kα ). Color : ICL (blue) vs. ICR (orange). Line style : TabPFN (solid, •) vs. k -NN (dashed, □ ). Both axes are log-scaled. Corpus size m=105 , ntest=400 ; each curve is the mean over 5 seeds (42–46).
n
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
4
0.151±0.036
0.066±0.006
0.147±0.040
0.058±0.005
16
0.151±0.013
0.040±0.005
0.129±0.046
0.023±0.002
32
0.134±0.023
0.024±0.003
0.092±0.019
0.014±0.001
64
0.103±0.035
0.018±0.004
0.058±0.022
0.008±0.001
128
0.059±0.009
0.012±0.002
0.029±0.013
0.005±0.001
256
0.036±0.020
0.007±0.001
0.020±0.011
0.002±0.001
Appendix
Table 5 : Weierstrass sample-size sweep, α=1.0 : target excess risk, mean ± std over 5 seeds.
n
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
4
0.160±0.028
0.077±0.006
0.157±0.027
0.067±0.008
16
0.154±0.015
0.043±0.006
0.121±0.060
0.024±0.002
32
0.127±0.025
0.027±0.004
0.096±0.017
0.014±0.002
64
0.107±0.034
0.019±0.005
0.071±0.036
0.008±0.001
128
0.069±0.028
0.011±0.003
0.055±0.018
0.003±0.001
256
0.037±0.024
0.006±0.001
0.043±0.017
0.002±0.000
Appendix
Table 6 : Weierstrass sample-size sweep, α=0.5 : target excess risk, mean ± std over 5 seeds.
Figure 11 : Target excess risk vs. posterior drift ε=QX(B) on the Weierstrass family at α=1.0 (smooth η ), for n∈{16,512} over 5 seeds. B is a union of N=40 periodic sub-brackets, each covering the leftmost ε -fraction of one width- 2/N interval of [−1,1] , so B samples every macro-cycle of ηQ uniformly and QX(B)=ε holds exactly. Color : ICL (blue) vs. ICR (orange). Line style : TabPFN (solid, •) vs. k -NN (dashed, □ ). ICL curves are exactly flat because ICL uses Q labels and does not see the drift; ICR curves rise almost linearly in ε across the whole range, saturating near 1−2RQ∗≈0.306 at ε=1 . At small n (left, n=16 ) ICR beats ICL up to ε≈0.4 – 0.5 ; at large n (right, n=512 ) ICR is roughly an order of magnitude below ICL at ε=0 (variance-limited regime), but ICL overtakes it already by ε≈0.08 . In both regimes the ICR bias tracks ε almost linearly, matching the ε+(λ/κ)(RP−RP∗) reduction in Lemma 1 .
ε
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
0.0
0.151±0.013
0.040±0.005
0.129±0.046
0.023±0.002
0.1
0.151±0.013
0.063±0.003
0.129±0.046
0.049±0.002
0.2
0.151±0.013
0.083±0.003
0.129±0.046
0.073±0.003
0.3
0.151±0.013
0.105±0.002
0.129±0.046
0.098±0.002
0.4
0.151±0.013
0.129±0.004
0.129±0.046
0.124±0.001
0.5
0.151±0.013
0.149±0.007
0.129±0.046
0.151±0.002
Appendix
Table 7 : Experiment 2, α=1.0 , n=16 : target excess risk vs. posterior drift ε=QX(B) , mean ± std over 5 seeds (Weierstrass family, A=0.24 , b=0.85 , f=2 , J=12 ; oracle kα=round_odd(0.30⋅n2α/(2α+1)) ; mcorpus=105 ; ntest=400 ). Region B is the leftmost ε -fraction of each of N=40 equal-width periodic brackets partitioning [−1,1] , so QX(B)=ε exactly and B samples every macro-cycle of ηQ uniformly; ηP=1−ηQ on B ; ICL uses Q labels (so ICL columns are ε -constant).
ε
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
0.0
0.019±0.014
0.003±0.001
0.018±0.008
0.001±0.000
0.1
0.019±0.014
0.032±0.001
0.018±0.008
0.020±0.005
0.2
0.019±0.014
0.060±0.001
0.018±0.008
0.060±0.001
0.3
0.019±0.014
0.090±0.001
0.018±0.008
0.090±0.001
0.4
0.019±0.014
0.120±0.001
0.018±0.008
0.120±0.001
0.5
0.019±0.014
0.151±0.001
0.018±0.008
0.150±0.001
Appendix
Table 8 : Experiment 2, α=1.0 , n=512 : target excess risk vs. posterior drift ε=QX(B) , mean ± std over 5 seeds (same setup as Table 7 ).
Figure 12 : Target excess risk vs. smoothness exponent α on the Weierstrass family at n=256 , with α∈{1,10−1,10−2,10−3} (log-spaced, 5 seeds). Color : ICL (blue) vs. ICR (orange). Line style : TabPFN (solid, •) vs. k -NN (dashed, □ ). The horizontal axis is logarithmic; the α -axis is reversed so rougher ηα is on the right. Both k -NN-ICL and TabPFN-ICL trend upward as α decreases, matching the n−2α/(2α+1) prediction of Theorem 1 ; the k -NN-ICL curve saturates near α→0 at the finite- n noise floor. ICR sits about an order of magnitude below ICL and is essentially flat in α , consistent with the α -independent n−(β+1)/2 rate of Theorem 1 ; the ICL/ICR gap widens as α→0 .
Figure 13 : Excess error on Q versus corpus size m , with n fixed and αP=1 . Color : ICL (blue) versus ICR (orange). Line style : TabPFN (solid, •) vs. k -NN (dashed, □ ). Per (n,seed) we draw one corpus of size mmax and use nested prefixes Dm=Dmmax[:m] ; the ICL demo set is drawn once per n from that corpus and reused across every m . Consequently ICL is exactly constant along the m axis for each seed, and the horizontal ICL baselines are the across-seed means of those per-seed constants. Both ICR curves collapse sharply around the reference scale and remain near their asymptotic value thereafter. The dotted grey line marks the unit-constant reference n1+1/(2αP)=n3/2 ( 512 for n=64 , ≈1448 for n=128 ); the actual saturation threshold is C⋆n1+1/(2αP) for some unknown constant. All 5 seeds (42–46). Setup: 1-D Weierstrass η -family with A=0.24 , b=0.85 , f=2 , J=12 , oracle kα=round_odd(0.30⋅n2α/(2α+1)) , linspace ntest=400 , and joint demonstrations.
n
m
TabPFN-ICL
TabPFN-ICR
k -NN-ICL
k -NN-ICR
64
64
0.097±0.033
0.133±0.019
0.079±0.016
0.145±0.042
128
0.073±0.022
0.140±0.028
256
0.038±0.011
0.109±0.032
512
0.020±0.005
0.019±0.009
1024
0.021±0.006
0.009±0.003
2048
0.021±0.006
0.008±0.004
Appendix
Table 10 : Corpus-size sweep: excess error on Q , mean ± sample standard deviation over five seeds. ICL entries are constant along the m axis by construction (see Fig. 13 caption) and are shown once per n ; ICR entries vary per m .
Figure 14 : Test error vs. number of demonstrations n on the age-biased sentiment dataset (878 test annotators, mean over 5 seeds), for Qwen3.5-4B (left) and Qwen3.5-27B (right). Blue : ICL with demonstrations sampled uniformly without replacement from the annotator’s own data. Yellow → red : ICR under four corpus designs with increasing posterior drift. The stronger model improves ICR only where the corpus is aligned with the target; under severe and extreme drift its error is essentially unchanged—the drift-induced bias is irreducible. Per-cell mean values are in Tables 11 and 12 .
Figure 15 : Test error vs. number of demonstrations n on the age-biased sentiment dataset (878 test annotators, Qwen3.5-4B). Each point is the mean over 5 independent runs (seeds 42–46). Blue : Qwen-ICL with demonstrations sampled uniformly without replacement from the annotator’s own training data. Yellow → red : Qwen-ICR under four levels of posterior drift, as a result of expanding the source corpus: no drift ( P=Q , self only), mild (self + 100 random others), severe (all 1,481 annotators), and extreme (all other annotators, self excluded). ICR sharply outperforms ICL when the source corpus is well-aligned with the target distribution and degrades as posterior drift increases. Uncertainty estimates are reported in Table 11 .
n
Qwen-ICL
ICR ( P=Q )
ICR (mild)
ICR (severe)
ICR (extreme)
4
0.538±0.007
0.263±0.000
0.265±0.009
0.535±0.004
0.633±0.004
8
0.481±0.008
0.169±0.002
0.225±0.009
0.504±0.002
0.628±0.007
16
0.348±0.011
0.099±0.002
0.162±0.012
0.491±0.003
0.615±0.006
32
0.117±0.004
0.072±0.001
0.142±0.010
0.480±0.002
0.619±0.008
Appendix
Table 11 : Sentiment classification with Qwen3.5-4B: test error (mean ± std over 5 seeds, 878 test annotators) vs. number of demonstrations n across the four corpus designs.
Figure 16 : Test error vs. number of demonstrations n on the age-biased sentiment dataset (878 test annotators, Qwen3.5-27B). Each point is the mean over 5 independent runs (seeds 42–46). Blue : Qwen-ICL with demonstrations sampled uniformly without replacement from the annotator’s own training data. Yellow → red : Qwen-ICR under four levels of posterior drift, as a result of expanding the source corpus: no drift ( P=Q , self only), mild (self + 100 random others), severe (all 1,481 annotators), and extreme (all other annotators, self excluded). Compared to Qwen3.5-4B ( Figure 15 ), the stronger model lowers the ICR error under no/mild drift but does not remove the irreducible bias under severe or extreme drift, matching the theory. Uncertainty estimates are reported in Table 12 .
n
Qwen-ICL
ICR ( P=Q )
ICR (mild)
ICR (severe)
ICR (extreme)
4
0.514±0.005
0.118±0.001
0.165±0.003
0.522±0.008
0.619±0.006
8
0.459±0.005
0.094±0.000
0.154±0.006
0.507±0.003
0.620±0.005
16
0.341±0.006
0.083±0.000
0.155±0.005
0.506±0.003
0.606±0.008
32
0.132±0.007
0.097±0.001
0.163±0.005
0.502±0.003
0.612±0.007
Appendix
Table 12 : Sentiment classification with Qwen3.5-27B: test error (mean ± std over 5 seeds, 878 test annotators) vs. number of demonstrations n across the four corpus designs.
In-context learning (ICL) is an emerging paradigm that employs the semantic information inherent in large language models (LLMs) for generating answers to user queries. While the remarkable performance of ICL has been widely known, a general modeling and a rigorous theoretical analysis of this paradigm are still lacking. This work presents a probabilistic model for ICL and derives the performance of ICL for both general parametric distributions and exponential families. Based on the derived results, the work explains the impact of multiple factors such as the number of demonstrations, the sensitivity of the probabilistic model to the variation of its parameters, as well as the similarity between the demonstrations and the query on the performance of ICL.
Zhenyu Liu, Huaze Tang, Shao-Lun Huang
Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China
The factors driving the performance of in-context learning (ICL) in large language models (LLMs) remain poorly understood despite ICL's surprising effectiveness, enabling models to adapt to new tasks from only a handful of examples. To clarify and improve these capabilities, we characterize how the statistical properties of the pretraining distribution (e.g., tail behavior, coverage) shape ICL. We develop a theoretical framework that encompasses generalization and task selection and show how distributional properties govern sample efficiency, task retrieval, and robustness. To this end, we generalize existing concentration results to heavy-tailed priors and dependent sequences, better reflecting the structure of LLM pretraining data. Our framework reveals a fundamental design trade-off: heavy-tailed pretraining distributions facilitate robust task selection under distribution shifts but are detrimental to generalization, especially in low-data regimes. We then empirically evaluate our predictions by studying how ICL performance varies with the pretraining distribution on challenging tasks such as stochastic differential equations and stochastic processes with memory. Together, these findings suggest that controlling key statistical properties of the pretraining distribution is essential for building ICL-capable and reliable LLMs.
Waïss Azizian, Ali Hasan
Univ. Grenoble Alpes, CNRS, Inria, Grenoble INP, LJK, 38000 Grenoble, France · Work done during an internship at Morgan Stanley Machine Learning Research. · Machine Learning Research, Morgan Stanley, New York, USA
In-context fine-tuning (IC-Train), training an LLM with labeled examples in-context, is increasingly used in place of standard fine-tuning for domain adaptation and continual absorption of labeled data. We study the robustness of the in-context learning ability that emerges from such training: does the fine-tuned model perform well across test inputs whose in-context examples range from unrelated to nearly identical? Across 32 configurations spanning four open-source LLMs and eight test sets over machine translation, Text-to-SQL, and multilingual semantic parsing, we show that robustness hinges on an overlooked design choice: how in-context examples are selected relative to the target during training. The two prevailing strategies turn out to be accurate over complementary parts of this spectrum: random contexts yield a model that gains little from related examples even when they are placed in its context, while retrieved similar contexts weaken accuracy on targets lacking close neighbors and raise the propensity to copy labels from context. Probes tracking in-weights learning, in-context learning, and copying trace these failures to distinct training dynamics, and show that introducing contrast in target-context similarity both within a context and across batches, restores robustness across the entire spectrum.