Retrieval-Augmented Generation (RAG) has become a cornerstone in software engineering for enhancing Large Language Models (LLMs) with domain-specific knowledge. However, adapting retrievers to evolving code repositories remains challenging due to the noise and redundancy inherent in massive code corpora. Standard fine-tuning on the full corpus is computationally expensive and often leads to sub-optimal performance due to negative transfer from low-quality samples. Conversely, simple random sampling fails to guarantee data representativeness. To address these challenges, we propose MAP4CS (Multi-dimensional Awareness Pruning for Code Search), an adaptive data pruning framework. MAP4CS identifies a small, high-quality core subset by integrating syntactic structure, semantic diversity, and distributional representation, followed by a rigorous rule-based filtering pipeline. Extensive experiments on two large-scale datasets demonstrate that MAP4CS consistently outperforms random sampling baselines using only 5% of the training data. Remarkably, it achieves performance comparable to, or even superior to, fine-tuning on the full dataset, validating the ''less is more'' hypothesis in data-centric AI. Furthermore, linguistic analysis reveals an adaptive optimization mechanism: MAP4CS automatically functions as a de-duplicator for redundant corpora and a denoiser for chaotic ones, constructing a training corpus that is both lexically diverse and information-dense.
Figures & tables
Figure 1. Motivation Example 1 (Semantic Gap): The baseline model is misled by the explicit keyword “capacity hint” in a distractor method. In contrast, training with diverse semantic strata helps identify the Ground Truth which handles this logic implicitly.
Figure 2. Motivation Example 2 (Structural Alignment): The query implies a specific sequential logic (“onNext” then “onComplete”). Without structural awareness, the baseline overfits to domain terms in complex boilerplate, missing the correct control flow found by AST-aware pruning.
Figure 3. Overview of MAP4CS
Training Dataset
Method
In-Domain
Out-of-Distribution
Avg.
CodeSearchNet
CoSQA
APPS
Text2SQL
Lexical Baseline
None
BM25
0.360
0.086
0.009
0.025
0.120
UniXcoder
None
Zero-Shot
0.431
0.122
0.007
0.124
0.171
CSN-Java
Full Data (100%)
0.365
0.092
0.010
0.030
0.124
Table 1. Performance comparison on downstream code retrieval tasks.
Figure 4. Visualization of semantic distribution in the Code-Query Semantic Space. The Random Sampling baseline (Blue) exhibits a diffuse distribution with scattered outliers, indicating high-entropy noise. In contrast, MAP4CS (Red) demonstrates a more refined distribution, concentrating on high-value semantic regions.
Figure 5. Sunburst plot of the top 10 most frequent verbs, with their corresponding top 7 root nouns within the query subsets obtained via Random Sampling and MAP4CS.
Method
In-Domain
Out-of-Distribution
Avg.
CodeSearchNet
CoSQA
APPS
Text2SQL
GTE base (Training Dataset: Query4Code)
Random Sampling (5%)
0.374
0.146
0.012
0.042
0.144
Clustering Only (5%)
0.401
0.146
0.018
0.041
0.152
MAP4CS (5%)
0.432
0.140
0.027
0.044
0.161
Table 2. Ablation study of different components in our framework. The best results among fine-tuned models are highlighted in bold .
LLM-based repository-level code generation aims to generate code using the context available in a software repository, requiring LLMs to reason over complex code dependencies. Due to limited context windows and insufficient repository-specific understanding, LLMs typically rely on retrieval-augmented generation (RAG) to incorporate relevant code. Early RAG approaches primarily employ similarity-based retrieval, which often fails to retrieve code snippets that the target function depends on. Recent work introduces graph-based retrieval to model such dependencies, but typically relies on manually designed rules and static global graphs, leading to limited flexibility and high construction and maintenance costs. In contrast, human developers collect helpful context by implicitly constructing a partial dependency graph and iteratively inspecting along it. Inspired by this behavior, we propose DyRetriever, an efficient context retrieval method via partial dependency graphs. DyRetriever uses an LLM to first select a set of entry-point functions and then perform multi-hop reasoning along the code dependency graph. During multi-hop reasoning, it uses the LLM's semantic understanding to validate whether a function can help generate the target function, eliminating manually designed rules and enabling flexibility across scenarios. Instead of statically constructing a global dependency graph, DyRetriever builds a partial graph on demand and discards it after use, reducing construction and maintenance costs. We integrate DyRetriever with a similarity-based code retriever to build DyCoder and evaluate it on CoderEval and DevEval. Experimental results show that DyCoder achieves relative Pass@1 improvements of 25.63% and 59.73% on CoderEval and DevEval, respectively, compared with existing RAG-based methods, while being 7.4x faster than baselines based on static dependency graph construction.
Zhongxin Liu, Zhonghao Jiang, Zhifan Ye +3
College of Computer Science and Technology and the State Key Laboratory of Blockchain and Data Security, Zhejiang University · School of Computer and Computing Science, Hangzhou City University · Faculty of Computing, Harbin Institute of Technology +1
Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rather than patching it. On SWE-Bench Verified, a 30B OpenHands agent averages 23 rounds and 631K tokens per resolved issue, with many calls spent on grep, glob, and view_file during repository exploration. We introduce CodeGrep, a 14B retrieval agent trained end-to-end with GRPO to issue multi-turn parallel grep, glob, and read tool calls and return candidate files to a frozen downstream coding agent. On all 500 SWE-Bench Verified instances, CodeGrep preserves resolve rate while substantially improving efficiency: 27.0% versus 25.8% for the no-retrieval baseline, with 15% fewer rounds and 19% fewer tokens on resolved instances. Across retrievers, downstream utility follows a precision threshold: BM25 with precision 0.375 degrades the agent, Jina with precision 0.445 is neutral, and CodeGrep with precision 0.677 crosses the threshold at which retrieval begins to reduce rollout cost. To enable this study, we mine supervision from 67K open-source agent trajectories using CATM and build a Git-worktree environment for multi-turn agent RL. In our setting, applying the efficiency signal at the advantage layer rather than the reward layer reduces KL drift and translates cleanly into downstream efficiency. We will release the model, training pipeline, RL environment, and evaluation harnesses.
Wuya Chen, Yihao yang, Yang Cao +1
1Netease Guangzhou AI Lab · 2Independent Researcher
Semantic code search and clone detection are essential for software development, maintenance, and reuse. This paper evaluates the effectiveness, efficiency, and scalability of contemporary deep learning models for first-stage recall in large-scale code-to-code search engines. Benchmarking across multiple programming languages and datasets reveals critical limits in the precision and scalability of these models on Terabyte-scale source-code collections. We present LLM-based code normalisation and query-rewriting schemes that yield significant gains in precision for lower-performing models. Our results question the sustainability of resource-constrained deployment and the assumed robustness of current code-specialised LLMs across datasets. We conclude with actionable insights for building scalable, efficient code-retrieval systems.
Leonardo Venuta, Francesco Tosoni, Paolo Ferragina