In-context learning (ICL) is crucial for boosting the inference performance of large language models (LLMs). However, the effectiveness of ICL in LLMs is greatly influenced by the choice of demonstration sets. Exhaustive searches over these sets are combinatorial, and existing selectors often rely on relevance or likelihood proxies to implicitly assess ICL quality. Making repeated queries to the target LLM with these strategies can incur substantial costs. This work simplifies selection by framing it as a constrained local search problem and presents local demonstration editing (LDE). Starting with an initially retrieved set of demonstrations, LDE employs a single structured edit to explore its surrounding neighborhood while balancing performance gains with search costs. Technically, LDE is reduced to a policy search problem, for which we train a small LLM, referred to as Jev-LDE. This model as the System-1 modifies the retrieved demonstration set by performing actions such as \texttt{Keep}, \texttt{Delete}, or \texttt{Replace} elements, all within a framework of reinforcement learning with verifiable rewards. At test time, Jev-LDE executes a single edit of the retrieved demonstration set, followed by one inference from the target LLM, avoiding the need for iterative context scoring or subset searches. Across standard classification benchmarks, various target LLMs with Jev-LDE as the plug-and-play module consistently improve ICL performance, and Jev-LDE shows transferability to held-out benchmarks and models without retraining. These findings indicate that the LDE approach offers an efficient and adaptable method for harnessing the ICL capabilities of target LLMs.
Figures & tables
Figure 1: One-shot classification result. The accuracy is averaged over TREC, DBPedia, and Banking77 using four target LLMs.
Figure 2: One edit can substantially affect 4-shot TREC accuracy. Violins show the full accuracy-change distribution under uniformly sampled valid actions. The Oracle using test labels to select the optimal subset.
Figure 3: Overview of local demonstration editing (LDE). Given a initially retrieved set Z0 , e.g., using Semantic TopK, and candidates C− , Jev-LDE chooses one action to edit Z0 , then edited set supports one frozen-target prediction.
Figure 4: Jev-LDE editor interface. Condensed prompt and JSON action format, illustrated with a Replace action. The full template appears in Appendix C .
Target model
Method
TREC
DBPedia
Banking77
1-shot
2-shot
3-shot
4-shot
1-shot
2-shot
3-shot
4-shot
1-shot
2-shot
3-shot
4-shot
Qwen2.5-7B-Instruct
Random
16.8
27.1
35.2
40.5
33.0
39.4
43.5
48.5
4.3
5.5
6.4
7.6
Semantic TopK
65.4
75.0
79.0
80.4
89.7
93.9
95.0
95.3
88.6
90.7
90.8
91.6
BM25
64.6
77.6
81.6
84.4
86.4
91.7
94.1
95.4
75.0
80.8
83.6
85.0
TopK+ConE
46.6
57.2
60.6
65.6
82.2
90.1
92.3
93.7
70.8
78.5
81.0
82.8
LMS3
65.6
75.0
79.2
80.8
89.5
93.7
95.0
95.1
88.7
90.8
90.8
91.3
Table 1: Few-shot classification accuracy (%). All results are averaged over five runs. Best results, including ties, are bold; second-best results are underlined. Blue rows denote Jev-LDE.
Figure 5: Evaluation with DeepSeek-V4-Flash. Accuracy on TACRED and TREC-50 across one to four shots.
Figure 6: Additional results and pre-selector analyses. Left: accuracy differences Δ(ACC) between Jev-LDE and Semantic TopK on AGNews, SST-2, GSM8K, and AQuA-RAT, Jev-LDE minus Semantic TopK. Right: accuracy differences obtained by applying Jev-LDE to three pre-selectors on TREC.
Figure 7: RLVR and candidate-pool ablations on one-shot TREC. Left: editing-policy accuracy with Meta-Llama-3-8B-Instruct as the target. Model names denote editors; Random samples a local editing action uniformly. Right: accuracy of Jev-LDE with Qwen2.5-7B-Instruct as a function of the number of replacement candidates.
Figure 8: Computational cost on one-shot TREC. End-to-end wall time with Qwen2.5-7B-Instruct.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Task
Eval. Split
#Eval.
#Demo Pool
Editor Train
AGNews
Topic classification
Test
1,000
10,000
Yes
TREC
Coarse question classification
Test
500
5,452
Yes
DBPedia
Ontology classification
Test
1,000
10,000
Yes
GSM8K
Mathematical reasoning
Test
1,319
7,473
Yes
SST-2
Sentiment classification
Validation
872
6,921
No
Banking77
Intent classification
Test
3,076
9,993
No
Appendix
Table 2: Primary benchmark statistics. The demonstration pool is drawn from each benchmark’s training split. “Editor Train” indicates whether the benchmark contributes Jev-LDE post-training states.
Method
Search Scope
Working Set
Selection Rule
Random
Global
Full pool
Uniformly sample k demonstrations.
Semantic TopK
Global
Full pool
Rank the full pool by cosine similarity using normalized BAAI/bge-base-en-v1.5 embeddings, and select the top- k demonstrations.
BM25
Global
Full pool
Select the top- k demonstrations by BM25 score.
TopK+ConE
Two-stage
top-30 candidates
Rerank the top-30 semantic candidates using the conditional-entropycriterion, then choose the top- k demonstrations.
LMS3
Global
Full pool
Combine similarity and stability scores with a 1% similarity filter and no refill.
CASE
Global
Full pool
Construct candidate arms by sampling one demonstration from each of k clusters. Identify the top- m subsets through challenger-arm sampling using validation accuracy as reward, and select the subset with the highest validation accuracy.
Appendix
Table 3: Candidate construction protocols. ”Full” denotes the complete benchmark-specific demonstration pool. For Jev-LDE, the initial k -shot set is retrieved from the full pool via Semantic TopK; the top-16 candidates denote the local neighborhood available to the editor for Replacement candidates.
Component
Setting
Component
Setting
Editor initialization
Qwen3-1.7B
Frozen target LLM
Meta-Llama-3-8B-Instruct
Training benchmarks
AGNews, TREC, DBPedia, GSM8K
Construction shot budgets
{1,2,4,8,10}
Candidate-pool size
16
States per benchmark–budget–round
Up to 1,000
Filtering rounds
3
Probe actions
1 Keep , up to 2 Delete , up to 5 Replace
Demo/query truncation
1,000/1,000 characters
Deduplication key
Dataset, split, query, k , pool size
Advantage estimator
Group-relative ( grpo_plus )
Policy rollouts per state
8
Appendix
Table 4: Jev-LDE post-training configuration. Rollout sampling is used only during post-training.
Figure 9: Jev-LDE post-training dynamics. Panel (a) reports the mean binary reward returned by the frozen target, and panel (b) reports the actor entropy loss. Light curves show the raw per-update values; solid curves apply the debiased exponential moving-average smoother used by the reference plotting protocol, with smoothing parameter 0.5 .
Table 5: Accuracy values underlying the generalization heatmap (Fig. 6 , left). Semantic TopK and Jev-LDE accuracy (%) on AGNews, SST-2, GSM8K, and AQuA-RAT for Qwen2.5-7B-Instruct and Meta-Llama-3-8B-Instruct. “Mean Δ ” is the four-shot average difference of Jev-LDE relative to Semantic TopK.
Target
Variant
1-shot
2-shot
3-shot
4-shot
Mean Δ
Qwen2.5-7B
BM25
64.6
77.6
81.6
84.4
BM25 + Jev-LDE
82.2
86.8
88.6
89.8
+9.8
CASE
65.6
75.4
79.2
81.8
CASE + Jev-LDE
80.4
86.6
87.6
88.0
+10.2
TopK + ConE
46.6
57.2
60.6
65.6
TopK + ConE + Jev-LDE
75.6
83.4
81.0
81.4
+22.9
Appendix
Table 6: TREC pre-selector ablation (%). The editor is applied to the top-16 pool of each pre-selector on TREC. “Unedited” is the selector’s own result; “+Jev-LDE” applies one local edit. Δ is the difference relative to “Unedited”.
Target
Dataset
8-shot
16-shot
TopK
+LDE
Δ
TopK
+LDE
Δ
Qwen2.5-7B
TREC
83.00
84.00
+1.00
84.60
85.20
+0.60
Banking77
91.84
92.04
+0.20
92.17
91.97
−0.20
DBPedia
98.00
97.70
−0.30
98.10
98.40
+0.30
Llama-3-8B
TREC
68.80
74.80
+6.00
52.60
54.80
+2.20
Banking77
90.90
90.51
−0.39
62.00
62.68
+0.68
Appendix
Table 7: 8-shot and 16-shot classification accuracy (%). TopK denotes Semantic TopK, and +LDE applies one local edit using Jev-LDE. Bold indicates the better result within each pair, including ties. Δ denotes the accuracy difference in percentage points, computed from the displayed values.
Although existing model editing methods perform well in recalling exact edit facts, they often struggle in complex scenarios that require deeper semantic understanding rather than mere knowledge regurgitation. Leveraging the strong contextual reasoning abilities of large language models (LLMs), in-context learning (ICL) becomes a promising editing method by comprehending edit information through context encoding. However, this method is constrained by the limited context window of LLMs, leading to degraded performance and efficiency as the number of edits increases. To overcome this limitation, we propose InComeS, a flexible framework that enhances LLMs' ability to process editing contexts through explicit compression and selection mechanisms. Specifically, InComeS compresses each editing context into the key-value (KV) cache of a special gist token, enabling efficient handling of multiple edits without being restricted by the model's context window. Furthermore, specialized cross-attention modules are added to dynamically select the most relevant information from the gist pools, enabling adaptive and effective utilization of edit information. We conduct experiments on diverse model editing benchmarks with various editing formats, and the results demonstrate the effectiveness and efficiency of our method.
Shuaiyi Li, Zhisong Zhang, Yang Deng +6
The Chinese University of Hong Kong · City University of Hong Kong · Singapore Management University +1
In-Context Learning (ICL) has become a cornerstone of modern LLM deployment. However, existing ICL post-training methods have a critical blind spot: they excel at extracting patterns from demonstrations while often neglecting context authority, the ability to determine whether contextual information should govern the final answer. To benchmark this capability, we introduce FakeContextBench, which contains pseudoscientific claims across seven domains. Our evaluation of commercial and open-source models shows that large-scale pre-training alone is insufficient for reliable context-authority discrimination. Moreover, prevalent ICL fine-tuning methods can increase susceptibility to misleading context, reducing reality accuracy by up to 14.95 percentage points relative to the base model. To address this trade-off, we propose Jurisdiction In-Context Learning (J-ICL), a post-training framework that incorporates context validation into the training objective. Across four model backbones, J-ICL improves ICLEval by an average of 5.84 percentage points and reality accuracy by 9.20 points over the corresponding base models. It also raises the Reality Rate by an average of 18.09 points relative to MetaICL and Symbol Tuning. These results demonstrate that ICL capability and resistance to deceptive context can be improved together. The benchmark is available at https://github.com/peilin717/FakeContext-Bench.
Pei-lin Li, Qingle Liu, Junyang Feng +4
Department of Computer Science and Technology, Tsinghua University · School of Integrated Circuits, Huazhong University of Science and Technology
In-context learning (ICL) is highly sensitive to which demonstrations appear in the prompt, but selecting them is expensive because the space of possible demonstration contexts and combinations is enormous. We argue that demonstration selection is \emph{easier to judge than to find}: predicting whether a specific query--context pair (q,D) will succeed is cheaper and more general than searching for an optimal D⋆. Based on this insight, we propose DiSP, a sample-and-judge framework that stratifies queries by difficulty. DiSP runs random demonstration trials to estimate success rate of each training query, trains a lightweight router to predict difficulty from the query, and trains level-specific judges for sampled demonstrations. At inference, DiSP performs stop-on-acceptance judging under an explicit budget, emitting diagnostic risk tags when no suitable context is found. Across five classification datasets with Llama3--8B and Qwen2.5--7B, DiSP achieves the best average accuracy, improving over strong learned selection baselines by up to 3.4%, while achieving up to 23× end-to-end wall-clock speedup.
Haochun Wang, Chaofen Yang, Jiatong Liu +5
Research Center for Social Computing and Interactive Robotics, Harbin Institute of Technology, China.