In-context fine-tuning (IC-Train), training an LLM with labeled examples in-context, is increasingly used in place of standard fine-tuning for domain adaptation and continual absorption of labeled data. We study the robustness of the in-context learning ability that emerges from such training: does the fine-tuned model perform well across test inputs whose in-context examples range from unrelated to nearly identical? Across 32 configurations spanning four open-source LLMs and eight test sets over machine translation, Text-to-SQL, and multilingual semantic parsing, we show that robustness hinges on an overlooked design choice: how in-context examples are selected relative to the target during training. The two prevailing strategies turn out to be accurate over complementary parts of this spectrum: random contexts yield a model that gains little from related examples even when they are placed in its context, while retrieved similar contexts weaken accuracy on targets lacking close neighbors and raise the propensity to copy labels from context. Probes tracking in-weights learning, in-context learning, and copying trace these failures to distinct training dynamics, and show that introducing contrast in target-context similarity both within a context and across batches, restores robustness across the entire spectrum.
Figures & tables
Figure 1: Schematic summary of main findings of the paper. On the X-axis is target to context similarity — IWL important on left side and ICL on right side. Fine-tuning with in-context example (IC-Train) is sensitive to similarity of target to context examples chosen during training: random context leads to sharp drop in ICL, similar context does not develop IWL and instead is prone to blind copying, contrast in the context recovers both limitations.
Figure 2: Effect of fine-tuning a base model with different strategies ( Zero-Context , and IC-Train under Random-Context , Similar-Context , and Contrastive-Context ) on accuracy over varying grades of similarity to in-context examples for nine of 32 different models, language-pairs, and test-sets (remaining plots in Appendix E ). X-axis: Level of maximum similarity of target with in-context examples. The similarity ranges here are - Low: 0−0.33 , Medium: 0.33−0.67 , High: 0.67−1 . Y-axis: Accuracy (COMET score).
Figure 3: Accuracy (Y-axis) against three levels of similarity (X-axis). The similarity ranges here are - (a) Low: 0−33 , Medium: 33−67 , High: 67−100 , (b) Low: 0−0.33 , Medium: 0.33−0.67 , High: 0.67−1 . Error bars show 95% confidence intervals.
Figure 4: Emergence of different forms of learning in three training methods: Random-Context, Similar-Context, and Contrastive-Context. X-axis is training steps and Y-axis denotes scores of one of the three probes. Results of other model-task and datasets in Appendix Figure F . IWL-score of Similar-Context is lowest, ICL-score of Random-Context diminishes fast with training, Copy-score (a stress test under shuffled contexts, see text) of Similar-Context is higher. Contrastive-Context provides best retention of ICL and IWL capabilities without resorting to copying. Curves are cubic-spline interpolations through the evaluated checkpoints, shown for readability.
Figure 5: Ablations to show (a) the robustness of Contrastive-Context to variant non-zero ϵ , p values, (b) various paraphrasing models, (c) importance of various levels of similarity in the training data. X-axis is target-context similarity and Y-axis is accuracy.
Appendix figures & tables37 assets
Supplementary material from the paper’s appendix.
Appendix
Train
ID Test
OOD Test
LP
Source
Size
Source
Size
Source
Size
En-De
Europarl
50000
Flores+
1012
EMEA
500
En-Hi
Samanantar
50000
Flores+
1012
Judicial
500
En-Ta
Samanantar
50000
Flores+
1012
Tanzil
500
En-Lt
Europarl
50000
Flores+
1012
EMEA
500
Appendix
Table 1: Datasets Used for MT. Flores+ refers to Flores-200. Sources: Europarl, Tanzil, and EMEA from Tiedemann (23-25) ; Samanantar from Ramesh et al. (2022) ; Flores+ from NLLB Team et al. (2024) ; Judicial from Kunchukuttan et al. (2018) .
Dataset
Link
Europarl
Helsinki-NLP/europarl
Tanzil
Helsinki-NLP/tanzil
EMEA
Helsinki-NLP/emea
Samanantar
ai4bharat/samanantar
Flores+
openlanguagedata/flores_plus
Judicial
cfilt/iitb-english-hindi
Appendix
Table 2: Links to datasets used
Figure 6: Accuracy (Y-axis) against three levels of similarity (X-axis). The similarity ranges here are - (a) Low: 0−33 , Medium: 33−67 , High: 67−100 , (b) Low: 0−0.33 , Medium: 0.33−0.67 , High: 0.67−1 . Error bars show 95% confidence intervals.
Prompt
Prompt string: x1:y1x2:y2x3:y3x∗
Example m=3,c=2
ACB: rijjpr CAB: jjriwp ABC: rtprjh BCA:
Example m=2,c=4
AC: ririjhjh CA: jjhhriir AB: rttrprrp BA:
Appendix
Table 9
Figure 7: X-axis: Level of maximum Jaccard similarity between the sets of letters used in the target and the context examples. Y-axis: Average maximum accepted length of respective subsequence in y for the PFAs of each letter in x∗ . Random-Context performs worse than the baseline irrespective of the similarity level, probably because it has learnt the PFAs before learning the alignment, and thus it chooses a PFA at random for every letter in x∗ . On the other hand, the baseline model performs better because it blindly copies an ICL target example sequence corresponding to an ICL context example having some symbols in the same position as x∗ . Although Contrastive-Context performs second only to Similar-Context in the low and medium similarity regions, Contrastive-Context outperforms all the models in the high and very high similarity regions.
Figure 8: Contrastive-Context is robust to specific choices of ϵ and p as long as sufficient contrast is maintained. X-axis is target-example similarity and Y-axis is accuracy.
(a) Qwen ⋅ OOD ⋅ En–Lt
x1
Each tablet contains 121.3 mg lactose.
y1
Kiekvienoje tabletė yra 121, 3 mg laktozė
x2
(hallucinations), mania (mental condition characterised by episodes of overactivity, elation or irritability), paranoia, suicidal thoughts
y2
Galite netekti riebalų Jūsų kojose, rankose ir veide, padaugėti riebalų pilve ir kituose vidaus
x3
Nitisinone is a competitive inhibitor of 4-hydroxyphenylpyruvate dioxygenase, an enzyme which precedes fumarylacetoacetate hydrolase in the tyrosine catabolic pathway.
y3
Nitizinonas yra konkuruojantis 4 - hidroksifenilpiruvato dioksigenazės (fermento, esančio prieš fumarilacetoacetato hidrolazę tirozino kataboliniame kelyje) inhibitorius.
Appendix
Table 3: Two Random-Context prompts shown in full, with the outputs they elicit. Each prompt contains five demonstrations (xi,yi) . Red italics mark the demonstration whose target Similar-Context reproduces verbatim as its translation; blue marks the output of Contrastive-Context, which translates the query in both cases. The very long fifth demonstration in (a) is truncated here for space; the model saw it in full.
Figure 9: Ablations on Contrastive-Context to establish the importance of training with different similarity levels: X-axis is target-example similarity and Y-axis is accuracy. Removing highly similar examples obtained via paraphrasing ( ϵ=0 ) causes test accuracy in the high similarity range to suffer. Removing natural similar examples ( p=0 ) causes accuracy in the medium range to suffer.
Figure 10: Comparing a Random-Context+Similar-Context mixture with Contrastive-Context that additionally also creates contrasts within a context, and other baselines Random-Context and Similar-Context. X-axis is target-example similarity and Y-axis is accuracy.
Figure 11: Comparing Similar-Context on original data and on data augmented with paraphrases. X-axis is similarity of test target instance with in-context examples, and Y axis is accuracy. Contrastive-Context uses paraphrases selectively to create contrast, whereas Similar-Context greedily selects Top-K most similar examples. While performance of Similar-Context improves in the high similarity range, it suffers in the low similarity range because IWL does not develop with highly similar context.
Figure 12: Comparing Contrastive-Context with highly similar examples created with various paraphrasing models. X-axis is target-example similarity and Y-axis is accuracy on the test data. Contrastive-Context is robust to the choice of the paraphrasing model.
Figure 13: Comparing training only attention projection weights (except output projection) with training attention projection weights + MLP layer weights. X-axis: Level of maximum similarity of target with in-context examples. The similarity ranges here are - Low: 0−0.33 , Medium: 0.33−0.67 , High: 0.67−1 . Y-axis: Accuracy (COMET score).
Figure 16: Unsmoothed learning dynamics plots for a few models and datasets. X-axis is training steps and Y-axis denotes scores of one of the three probes.
Large language models (LLMs) operate in two fundamental learning modes - fine-tuning (FT) and in-context learning (ICL) - raising key questions about which mode yields greater language proficiency and whether they differ in their inductive biases. Prior studies comparing FT and ICL have yielded mixed and inconclusive results due to inconsistent experimental setups. To enable a rigorous comparison, we propose a formal language learning task - offering precise language boundaries, controlled string sampling, and no data contamination - and introduce a discriminative test for language proficiency, where an LLM succeeds if it assigns higher generation probability to in-language strings than to out-of-language strings. Empirically, we find that: (a) FT has greater language proficiency than ICL on in-distribution generalization, but both perform equally well on out-of-distribution generalization. (b) Their inductive biases, measured by the correlation in string generation probabilities, are similar when both modes partially learn the language but diverge at higher proficiency levels. (c) Unlike FT, ICL performance differs substantially across models of varying sizes and families and is sensitive to the token vocabulary of the language. Thus, our work demonstrates the promise of formal languages as a controlled testbed for evaluating LLMs, behaviors that are difficult to isolate in natural language datasets. Our source code is available at https://github.com/bishwamittra/formallm.
Bishwamittra Ghosh, Soumi Das, Till Speicher +5
Max Planck Institute for Software Systems, Germany · Boston University, USA
As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by long, partially irrelevant context. In a controlled setting, we find that state-of-the-art models often appear robust to task-irrelevant context at the aggregate level: prepending it to benchmark questions causes little change in overall accuracy. This aggregate stability, however, masks significant per-example instability. Even semantically meaningless pseudo-words, formed by randomly combining characters, can markedly shift model predictions on a small fraction of examples, degrading performance on some while improving it on others. This two-sided effect holds consistently across a wide range of models and datasets, yet the affected examples are largely model-specific. We further show that this instability is modulated by context type, context length, test-time compute, and model development stage. Together, our findings reveal context-induced tail risks concealed by aggregate accuracy, motivating per-example reliability evaluation of language models.
This paper investigates context stickiness in in-context learning (ICL), a phenomenon where earlier examples in a prompt interfere with a transformer's ability to adapt to later tasks. Using synthetic regression tasks over linear and quadratic functions, we examine how models trained under sequential, mixed, and random curricula handle abrupt task switches during inference. By sweeping over structured combinations of misleading linear examples followed by recovery quadratic examples, we quantify how prior context biases prediction error and how quickly models realign. Our results show strong evidence of persistent interference: more preceding linear examples reliably degrade quadratic predictions, while additional quadratic examples reduce error but with diminishing returns. We further find that training curricula significantly modulate resilience, with sequential training on the target function class yielding the fastest recovery, and surprisingly, random training producing the least robust behavior.