In-context fine-tuning (IC-Train), training an LLM with labeled examples in-context, is increasingly used in place of standard fine-tuning for domain adaptation and continual absorption of labeled data. We study the robustness of the in-context learning ability that emerges from such training: does the fine-tuned model perform well across test inputs whose in-context examples range from unrelated to nearly identical? Across 32 configurations spanning four open-source LLMs and eight test sets over machine translation, Text-to-SQL, and multilingual semantic parsing, we show that robustness hinges on an overlooked design choice: how in-context examples are selected relative to the target during training. The two prevailing strategies turn out to be accurate over complementary parts of this spectrum: random contexts yield a model that gains little from related examples even when they are placed in its context, while retrieved similar contexts weaken accuracy on targets lacking close neighbors and raise the propensity to copy labels from context. Probes tracking in-weights learning, in-context learning, and copying trace these failures to distinct training dynamics, and show that introducing contrast in target-context similarity both within a context and across batches, restores robustness across the entire spectrum.
Figures & tables
Figure 1: Schematic summary of main findings of the paper. On the X-axis is target to context similarity — IWL important on left side and ICL on right side. Fine-tuning with in-context example (IC-Train) is sensitive to similarity of target to context examples chosen during training: random context leads to sharp drop in ICL, similar context does not develop IWL and instead is prone to blind copying, contrast in the context recovers both limitations.
Figure 2: Effect of fine-tuning a base model with different strategies ( Zero-Context , and IC-Train under Random-Context , Similar-Context , and Contrastive-Context ) on accuracy over varying grades of similarity to in-context examples for nine of 32 different models, language-pairs, and test-sets (remaining plots in Appendix E ). X-axis: Level of maximum similarity of target with in-context examples. The similarity ranges here are - Low: 0−0.33 , Medium: 0.33−0.67 , High: 0.67−1 . Y-axis: Accuracy (COMET score).
Figure 3: Accuracy (Y-axis) against three levels of similarity (X-axis). The similarity ranges here are - (a) Low: 0−33 , Medium: 33−67 , High: 67−100 , (b) Low: 0−0.33 , Medium: 0.33−0.67 , High: 0.67−1 . Error bars show 95% confidence intervals.
Figure 4: Emergence of different forms of learning in three training methods: Random-Context, Similar-Context, and Contrastive-Context. X-axis is training steps and Y-axis denotes scores of one of the three probes. Results of other model-task and datasets in Appendix Figure F . IWL-score of Similar-Context is lowest, ICL-score of Random-Context diminishes fast with training, Copy-score (a stress test under shuffled contexts, see text) of Similar-Context is higher. Contrastive-Context provides best retention of ICL and IWL capabilities without resorting to copying. Curves are cubic-spline interpolations through the evaluated checkpoints, shown for readability.
Figure 5: Ablations to show (a) the robustness of Contrastive-Context to variant non-zero ϵ , p values, (b) various paraphrasing models, (c) importance of various levels of similarity in the training data. X-axis is target-context similarity and Y-axis is accuracy.
Appendix figures & tables37 assets
Supplementary material from the paper’s appendix.
Appendix
Train
ID Test
OOD Test
LP
Source
Size
Source
Size
Source
Size
En-De
Europarl
50000
Flores+
1012
EMEA
500
En-Hi
Samanantar
50000
Flores+
1012
Judicial
500
En-Ta
Samanantar
50000
Flores+
1012
Tanzil
500
En-Lt
Europarl
50000
Flores+
1012
EMEA
500
Appendix
Table 1: Datasets Used for MT. Flores+ refers to Flores-200. Sources: Europarl, Tanzil, and EMEA from Tiedemann (23-25) ; Samanantar from Ramesh et al. (2022) ; Flores+ from NLLB Team et al. (2024) ; Judicial from Kunchukuttan et al. (2018) .
Dataset
Link
Europarl
Helsinki-NLP/europarl
Tanzil
Helsinki-NLP/tanzil
EMEA
Helsinki-NLP/emea
Samanantar
ai4bharat/samanantar
Flores+
openlanguagedata/flores_plus
Judicial
cfilt/iitb-english-hindi
Appendix
Table 2: Links to datasets used
Figure 6: Accuracy (Y-axis) against three levels of similarity (X-axis). The similarity ranges here are - (a) Low: 0−33 , Medium: 33−67 , High: 67−100 , (b) Low: 0−0.33 , Medium: 0.33−0.67 , High: 0.67−1 . Error bars show 95% confidence intervals.
Prompt
Prompt string: x1:y1x2:y2x3:y3x∗
Example m=3,c=2
ACB: rijjpr CAB: jjriwp ABC: rtprjh BCA:
Example m=2,c=4
AC: ririjhjh CA: jjhhriir AB: rttrprrp BA:
Appendix
Table 9
Figure 7: X-axis: Level of maximum Jaccard similarity between the sets of letters used in the target and the context examples. Y-axis: Average maximum accepted length of respective subsequence in y for the PFAs of each letter in x∗ . Random-Context performs worse than the baseline irrespective of the similarity level, probably because it has learnt the PFAs before learning the alignment, and thus it chooses a PFA at random for every letter in x∗ . On the other hand, the baseline model performs better because it blindly copies an ICL target example sequence corresponding to an ICL context example having some symbols in the same position as x∗ . Although Contrastive-Context performs second only to Similar-Context in the low and medium similarity regions, Contrastive-Context outperforms all the models in the high and very high similarity regions.
Figure 8: Contrastive-Context is robust to specific choices of ϵ and p as long as sufficient contrast is maintained. X-axis is target-example similarity and Y-axis is accuracy.
(a) Qwen ⋅ OOD ⋅ En–Lt
x1
Each tablet contains 121.3 mg lactose.
y1
Kiekvienoje tabletė yra 121, 3 mg laktozė
x2
(hallucinations), mania (mental condition characterised by episodes of overactivity, elation or irritability), paranoia, suicidal thoughts
y2
Galite netekti riebalų Jūsų kojose, rankose ir veide, padaugėti riebalų pilve ir kituose vidaus
x3
Nitisinone is a competitive inhibitor of 4-hydroxyphenylpyruvate dioxygenase, an enzyme which precedes fumarylacetoacetate hydrolase in the tyrosine catabolic pathway.
y3
Nitizinonas yra konkuruojantis 4 - hidroksifenilpiruvato dioksigenazės (fermento, esančio prieš fumarilacetoacetato hidrolazę tirozino kataboliniame kelyje) inhibitorius.
Appendix
Table 3: Two Random-Context prompts shown in full, with the outputs they elicit. Each prompt contains five demonstrations (xi,yi) . Red italics mark the demonstration whose target Similar-Context reproduces verbatim as its translation; blue marks the output of Contrastive-Context, which translates the query in both cases. The very long fifth demonstration in (a) is truncated here for space; the model saw it in full.
Figure 9: Ablations on Contrastive-Context to establish the importance of training with different similarity levels: X-axis is target-example similarity and Y-axis is accuracy. Removing highly similar examples obtained via paraphrasing ( ϵ=0 ) causes test accuracy in the high similarity range to suffer. Removing natural similar examples ( p=0 ) causes accuracy in the medium range to suffer.
Figure 10: Comparing a Random-Context+Similar-Context mixture with Contrastive-Context that additionally also creates contrasts within a context, and other baselines Random-Context and Similar-Context. X-axis is target-example similarity and Y-axis is accuracy.
Figure 11: Comparing Similar-Context on original data and on data augmented with paraphrases. X-axis is similarity of test target instance with in-context examples, and Y axis is accuracy. Contrastive-Context uses paraphrases selectively to create contrast, whereas Similar-Context greedily selects Top-K most similar examples. While performance of Similar-Context improves in the high similarity range, it suffers in the low similarity range because IWL does not develop with highly similar context.
Figure 12: Comparing Contrastive-Context with highly similar examples created with various paraphrasing models. X-axis is target-example similarity and Y-axis is accuracy on the test data. Contrastive-Context is robust to the choice of the paraphrasing model.
Figure 13: Comparing training only attention projection weights (except output projection) with training attention projection weights + MLP layer weights. X-axis: Level of maximum similarity of target with in-context examples. The similarity ranges here are - Low: 0−0.33 , Medium: 0.33−0.67 , High: 0.67−1 . Y-axis: Accuracy (COMET score).
Figure 16: Unsmoothed learning dynamics plots for a few models and datasets. X-axis is training steps and Y-axis denotes scores of one of the three probes.