In-context learning (ICL) is crucial for boosting the inference performance of large language models (LLMs). However, the effectiveness of ICL in LLMs is greatly influenced by the choice of demonstration sets. Exhaustive searches over these sets are combinatorial, and existing selectors often rely on relevance or likelihood proxies to implicitly assess ICL quality. Making repeated queries to the target LLM with these strategies can incur substantial costs. This work simplifies selection by framing it as a constrained local search problem and presents local demonstration editing (LDE). Starting with an initially retrieved set of demonstrations, LDE employs a single structured edit to explore its surrounding neighborhood while balancing performance gains with search costs. Technically, LDE is reduced to a policy search problem, for which we train a small LLM, referred to as Jev-LDE. This model as the System-1 modifies the retrieved demonstration set by performing actions such as \texttt{Keep}, \texttt{Delete}, or \texttt{Replace} elements, all within a framework of reinforcement learning with verifiable rewards. At test time, Jev-LDE executes a single edit of the retrieved demonstration set, followed by one inference from the target LLM, avoiding the need for iterative context scoring or subset searches. Across standard classification benchmarks, various target LLMs with Jev-LDE as the plug-and-play module consistently improve ICL performance, and Jev-LDE shows transferability to held-out benchmarks and models without retraining. These findings indicate that the LDE approach offers an efficient and adaptable method for harnessing the ICL capabilities of target LLMs.
Figures & tables
Figure 1: One-shot classification result. The accuracy is averaged over TREC, DBPedia, and Banking77 using four target LLMs.
Figure 2: One edit can substantially affect 4-shot TREC accuracy. Violins show the full accuracy-change distribution under uniformly sampled valid actions. The Oracle using test labels to select the optimal subset.
Figure 3: Overview of local demonstration editing (LDE). Given a initially retrieved set Z0 , e.g., using Semantic TopK, and candidates C− , Jev-LDE chooses one action to edit Z0 , then edited set supports one frozen-target prediction.
Figure 4: Jev-LDE editor interface. Condensed prompt and JSON action format, illustrated with a Replace action. The full template appears in Appendix C .
Target model
Method
TREC
DBPedia
Banking77
1-shot
2-shot
3-shot
4-shot
1-shot
2-shot
3-shot
4-shot
1-shot
2-shot
3-shot
4-shot
Qwen2.5-7B-Instruct
Random
16.8
27.1
35.2
40.5
33.0
39.4
43.5
48.5
4.3
5.5
6.4
7.6
Semantic TopK
65.4
75.0
79.0
80.4
89.7
93.9
95.0
95.3
88.6
90.7
90.8
91.6
BM25
64.6
77.6
81.6
84.4
86.4
91.7
94.1
95.4
75.0
80.8
83.6
85.0
TopK+ConE
46.6
57.2
60.6
65.6
82.2
90.1
92.3
93.7
70.8
78.5
81.0
82.8
LMS3
65.6
75.0
79.2
80.8
89.5
93.7
95.0
95.1
88.7
90.8
90.8
91.3
Table 1: Few-shot classification accuracy (%). All results are averaged over five runs. Best results, including ties, are bold; second-best results are underlined. Blue rows denote Jev-LDE.
Figure 5: Evaluation with DeepSeek-V4-Flash. Accuracy on TACRED and TREC-50 across one to four shots.
Figure 6: Additional results and pre-selector analyses. Left: accuracy differences Δ(ACC) between Jev-LDE and Semantic TopK on AGNews, SST-2, GSM8K, and AQuA-RAT, Jev-LDE minus Semantic TopK. Right: accuracy differences obtained by applying Jev-LDE to three pre-selectors on TREC.
Figure 7: RLVR and candidate-pool ablations on one-shot TREC. Left: editing-policy accuracy with Meta-Llama-3-8B-Instruct as the target. Model names denote editors; Random samples a local editing action uniformly. Right: accuracy of Jev-LDE with Qwen2.5-7B-Instruct as a function of the number of replacement candidates.
Figure 8: Computational cost on one-shot TREC. End-to-end wall time with Qwen2.5-7B-Instruct.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Task
Eval. Split
#Eval.
#Demo Pool
Editor Train
AGNews
Topic classification
Test
1,000
10,000
Yes
TREC
Coarse question classification
Test
500
5,452
Yes
DBPedia
Ontology classification
Test
1,000
10,000
Yes
GSM8K
Mathematical reasoning
Test
1,319
7,473
Yes
SST-2
Sentiment classification
Validation
872
6,921
No
Banking77
Intent classification
Test
3,076
9,993
No
Appendix
Table 2: Primary benchmark statistics. The demonstration pool is drawn from each benchmark’s training split. “Editor Train” indicates whether the benchmark contributes Jev-LDE post-training states.
Method
Search Scope
Working Set
Selection Rule
Random
Global
Full pool
Uniformly sample k demonstrations.
Semantic TopK
Global
Full pool
Rank the full pool by cosine similarity using normalized BAAI/bge-base-en-v1.5 embeddings, and select the top- k demonstrations.
BM25
Global
Full pool
Select the top- k demonstrations by BM25 score.
TopK+ConE
Two-stage
top-30 candidates
Rerank the top-30 semantic candidates using the conditional-entropycriterion, then choose the top- k demonstrations.
LMS3
Global
Full pool
Combine similarity and stability scores with a 1% similarity filter and no refill.
CASE
Global
Full pool
Construct candidate arms by sampling one demonstration from each of k clusters. Identify the top- m subsets through challenger-arm sampling using validation accuracy as reward, and select the subset with the highest validation accuracy.
Appendix
Table 3: Candidate construction protocols. ”Full” denotes the complete benchmark-specific demonstration pool. For Jev-LDE, the initial k -shot set is retrieved from the full pool via Semantic TopK; the top-16 candidates denote the local neighborhood available to the editor for Replacement candidates.
Component
Setting
Component
Setting
Editor initialization
Qwen3-1.7B
Frozen target LLM
Meta-Llama-3-8B-Instruct
Training benchmarks
AGNews, TREC, DBPedia, GSM8K
Construction shot budgets
{1,2,4,8,10}
Candidate-pool size
16
States per benchmark–budget–round
Up to 1,000
Filtering rounds
3
Probe actions
1 Keep , up to 2 Delete , up to 5 Replace
Demo/query truncation
1,000/1,000 characters
Deduplication key
Dataset, split, query, k , pool size
Advantage estimator
Group-relative ( grpo_plus )
Policy rollouts per state
8
Appendix
Table 4: Jev-LDE post-training configuration. Rollout sampling is used only during post-training.
Figure 9: Jev-LDE post-training dynamics. Panel (a) reports the mean binary reward returned by the frozen target, and panel (b) reports the actor entropy loss. Light curves show the raw per-update values; solid curves apply the debiased exponential moving-average smoother used by the reference plotting protocol, with smoothing parameter 0.5 .
Table 5: Accuracy values underlying the generalization heatmap (Fig. 6 , left). Semantic TopK and Jev-LDE accuracy (%) on AGNews, SST-2, GSM8K, and AQuA-RAT for Qwen2.5-7B-Instruct and Meta-Llama-3-8B-Instruct. “Mean Δ ” is the four-shot average difference of Jev-LDE relative to Semantic TopK.
Target
Variant
1-shot
2-shot
3-shot
4-shot
Mean Δ
Qwen2.5-7B
BM25
64.6
77.6
81.6
84.4
BM25 + Jev-LDE
82.2
86.8
88.6
89.8
+9.8
CASE
65.6
75.4
79.2
81.8
CASE + Jev-LDE
80.4
86.6
87.6
88.0
+10.2
TopK + ConE
46.6
57.2
60.6
65.6
TopK + ConE + Jev-LDE
75.6
83.4
81.0
81.4
+22.9
Appendix
Table 6: TREC pre-selector ablation (%). The editor is applied to the top-16 pool of each pre-selector on TREC. “Unedited” is the selector’s own result; “+Jev-LDE” applies one local edit. Δ is the difference relative to “Unedited”.
Target
Dataset
8-shot
16-shot
TopK
+LDE
Δ
TopK
+LDE
Δ
Qwen2.5-7B
TREC
83.00
84.00
+1.00
84.60
85.20
+0.60
Banking77
91.84
92.04
+0.20
92.17
91.97
−0.20
DBPedia
98.00
97.70
−0.30
98.10
98.40
+0.30
Llama-3-8B
TREC
68.80
74.80
+6.00
52.60
54.80
+2.20
Banking77
90.90
90.51
−0.39
62.00
62.68
+0.68
Appendix
Table 7: 8-shot and 16-shot classification accuracy (%). TopK denotes Semantic TopK, and +LDE applies one local edit using Jev-LDE. Bold indicates the better result within each pair, including ties. Δ denotes the accuracy difference in percentage points, computed from the displayed values.