For small merchant businesses (SMBs) whose catalogs fit within a long-context LLM, full-catalog prompting offers a compelling alternative to multi-stage retrieval designed primarily for large marketplaces with millions of items. However, fitting the full catalog into the context window does not ensure that the model can use it effectively, since LLMs do not exploit long contexts uniformly. We therefore study in-context catalog search through two complementary questions: (1) how to curate and present catalogs to the LLM, and (2) how to adapt the LLM for product selection on curated contexts. We propose RPTune, an end-to-end framework that couples learned catalog curation with LLM post-training using automatically generated, catalog-grounded supervision. An encoder-reorganizer curator orders and prunes products guided by downstream LLM feedback, while the resulting curated catalogs in turn improve the effectiveness of LLM post-training with a context-relative reward. We evaluate RPTune on 7 real merchants spanning distinct retail verticals, using 100 complex conversational queries per merchant. RPTune consistently improves search accuracy across both proprietary and open-weight LLMs, with context curation yielding gains of up to 31.4 percentage points and post-training adding a further 10.3 points on average.
Figures & tables
Figure 1 : Search accuracy vs. end-to-end latency, averaged over 7 merchants; complete results and illustrations can be found in Table 2 . RPTune context curation delivers accuracy gains comparable to a model upgrade with up to 21 × speedup (dashed cross-model comparisons in ), while RPTune post-training further improves gemma-4-E4B-it accuracy to comparable levels of frontier models .
Figure 2 : RPTune architecture. Top: At inference, the curator (encoder + reorganizer) orders and prunes the catalog to construct a compact context for downstream LLM product selection. Bottom: Synthetic queries and relevance scores support sequential encoder training, reorganizer training with frozen-LLM feedback, and LLM post-training on curated contexts. Flames / snowflakes indicate trainable/frozen modules at different stages.
Merchant
Retail vertical
# Prod.
# Var.
# Tok.
Beauty Bakerie (BB)
Cosmetics
96
115
37,506
Oats Overnight (OON)
Breakfast food
142
303
39,845
The Honest Kitchen (HK)
Dog & cat food
131
214
41,954
Wilsun Custom Products (WCP)
Custom gears
110
410
81,355
Wairua Beauty (WB)
Skin & hair care
277
409
84,870
Focus & Frame Eyewear (FF)
Eyewear
186
739
110,724
Table 1 : Statistics of the 7 evaluated SMBs.
Method / Model
EM (%) ↑
Average
BB
OON
HK
WCP
WB
FF
PF
EM (%) ↑
FR (%) ↑
Latency (s) ↓
Main method comparison (LLM backbone: gemini-3.7-flash )
RAG-Fusion
18.8
0.1
1.8
15.0
11.3
2.0
16.0
9.3
48.5
6.4
LongLLMLingua
9.1
12.3
18.4
8.6
9.1
5.3
31.4
13.5
52.2
64.2
TourRank
6.1
13.0
16.7
9.0
10.1
14.3
26.2
13.6
54.6
56.7
RAG [EmbeddingGemma 300M]
28.8
23.8
17.0
16.7
26.1
7.3
45.7
23.6
62.9
5.7
Table 2 : Main results on seven SMBs. Merchant columns report exact variant match (EM); the final block averages EM, feature reward (FR), and end-to-end latency across merchants. Arrowed values show EM/FR gains and latency reductions over the same backbone with full-catalog prompting.
Figure 3 : RPTune generalizes well without retraining. (a) EM and FR on Beauty Bakerie as withheld products are injected back, for all queries and for queries whose target is an original or a new product. 0% is the reduced catalog RPTune was trained on and 100% restores the full catalog. (b) Gains over gemma for every train/eval merchant pair, all positive ; boxed diagonal cells are in-domain. Average out-of-domain gains are +18.1 EM and +13.6 FR, in-domain +20.3 and +17.7 .
Figure 6Figure 7
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8 : Average relevance score of the products selected by the frozen LLM during RPTune reorganizer training.
Products reintroduced (%)
0
10
20
30
40
50
60
70
80
90
100
# queries, Total
100
100
100
100
100
100
100
100
100
100
100
# queries, Original-product
100
96
96
93
93
87
86
79
62
58
44
# queries, New-product
0
4
4
7
7
13
14
21
38
42
56
Appendix
Table 4 : Composition of the BB evaluation set at each injection level. The query set is fixed at 100 queries × 10 trials throughout, so the counts below are also percentages of the evaluation set. New-product counts queries whose reference label has been replaced by a reintroduced product; Original-product counts the remainder.
Figure 9 : FR view of the ablations: (a) post-training objectives (Figure 5 ), (b) pruning budget sweep (Figure 7 ), (c) context layouts on Beauty Bakerie (Figure 7 ) and (d) the curation quadrants (Figure 5 ).
Product catalogs are the backbone of e-commerce sites, yet a large number of structured attributes (SAs) -- such as material, color, and shape -- often have missing values. Typically, SA values are extracted from product information, including titles and descriptions. While LLM-based generator-evaluator frameworks have demonstrated effectiveness for SA prediction -- where an LLM generates SA values and another evaluates them -- they face challenges when the Generator and Evaluator produce conflicting outputs, as either component can make mistakes. We introduce \texttt{CatalogAgent}, a novel agentic system that continuously improves Generator and Evaluator models for e-commerce catalog enrichment. When disagreements arise from (1) internal conflicts between the LLM-based Generator and Evaluator, or (2) external feedback from sellers on LLM outputs, a Supervisor Agent intervenes to mediate these conflicts and make final decisions. The system also incorporates a Memory Base and a Memory Summarizer that stores Supervisor Agent activities from individual cases and aggregates patterns into learnings. These learnings are fed back to the worker Generator and Evaluator LLMs, enabling self-improvement without human intervention. Through context engineering -- injecting learnings and insights into worker LLMs' contexts -- the system successfully transfers the Supervisor's capabilities to the Generator and Evaluator, improving their performance by 15.24% and 13.98%, respectively. Our experiments demonstrate a new paradigm of Supervisor Agent-mediated self-learning systems for improving generative AI model accuracy.
Teaching language models to use search tools is not only a question of whether they search, but also of whether they issue good queries. This is especially important in open-domain question answering, where broad or copied queries often waste retrieval budget and derail later reasoning. We propose \Ours, a framework that makes query planning explicit through reusable search skills. At each step, the model first selects a skill, then generates a search or answer action conditioned on the selected skill card. The skill inventory itself is not fixed: SearchSkill maintains an evolving SkillBank, expands or refines it from recurrent failure patterns, and reconstructs affected trajectories before supervised training. The resulting two-stage SFT recipe aligns training with the inference-time protocol of skill selection followed by skill-grounded execution. Across open-source and closed-source models, SearchSkill improves exact match on knowledge-intensive QA benchmarks and yields better retrieval behavior, including fewer copied first queries, more atomic hop-focused queries, and more correct answers within a small search budget. These results suggest that explicit skill-conditioned query planning is a lightweight alternative to treating search as an undifferentiated action.
Jinchao Hu, Meizhi Zhong, Kehai Chen +1
School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen · Independent Researcher Shenzhen / Beijing, China
Search-augmented language models can use external evidence to compensate for limitations in parametric knowledge, but search is not uniformly beneficial: models may call search for questions they can already answer, or rely on noisy evidence when correction, clarification, or abstention would be more appropriate. We formulate this as an instance-level search-routing problem: deciding whether search is needed to improve task success relative to a no-search execution. To derive supervision, we compare no-search and forced-search outcomes for the same question and construct an oracle over NO SEARCH, SEARCH, and UNSOLVED based on task-specific success. Using this oracle as both an evaluation criterion and a learning signal, we train search-routing policies with supervised fine-tuning and preference optimization, improving routing macro-F1 on oracle-eligible examples from 0.7082 to 0.8235 for Gemma E2B and from 0.7053 to 0.8365 for Qwen3.5-4B. Further analysis shows that the learned policies reduce model-specific routing failures: Gemma primarily learns no-search restraint, while Qwen further reduces missed search; residual UNSOLVED cases reveal heterogeneous bottlenecks involving model capacity, retrieval budget, evidence use, and policy behavior.