MAP4CS: A Multi-dimensional Data Pruning Framework for Efficient Code Retriever Fine-tuning
Organizations: School of Software Engineering, Sun Yat-sen University, China
Abstract
Retrieval-Augmented Generation (RAG) has become a cornerstone in software engineering for enhancing Large Language Models (LLMs) with domain-specific knowledge. However, adapting retrievers to evolving code repositories remains challenging due to the noise and redundancy inherent in massive code corpora. Standard fine-tuning on the full corpus is computationally expensive and often leads to sub-optimal performance due to negative transfer from low-quality samples. Conversely, simple random sampling fails to guarantee data representativeness. To address these challenges, we propose MAP4CS (Multi-dimensional Awareness Pruning for Code Search), an adaptive data pruning framework. MAP4CS identifies a small, high-quality core subset by integrating syntactic structure, semantic diversity, and distributional representation, followed by a rigorous rule-based filtering pipeline. Extensive experiments on two large-scale datasets demonstrate that MAP4CS consistently outperforms random sampling baselines using only 5% of the training data. Remarkably, it achieves performance comparable to, or even superior to, fine-tuning on the full dataset, validating the ''less is more'' hypothesis in data-centric AI. Furthermore, linguistic analysis reveals an adaptive optimization mechanism: MAP4CS automatically functions as a de-duplicator for redundant corpora and a denoiser for chaotic ones, constructing a training corpus that is both lexically diverse and information-dense.
Figures & tables
| Training Dataset | Method | In-Domain | Out-of-Distribution | Avg. | ||
| CodeSearchNet | CoSQA | APPS | Text2SQL | |||
| Lexical Baseline | ||||||
| None | BM25 | 0.360 | 0.086 | 0.009 | 0.025 | 0.120 |
| UniXcoder | ||||||
| None | Zero-Shot | 0.431 | 0.122 | 0.007 | 0.124 | 0.171 |
| CSN-Java | Full Data (100%) | 0.365 | 0.092 | 0.010 | 0.030 | 0.124 |
| Method | In-Domain | Out-of-Distribution | Avg. | ||
| CodeSearchNet | CoSQA | APPS | Text2SQL | ||
| GTE base (Training Dataset: Query4Code) | |||||
| Random Sampling (5%) | 0.374 | 0.146 | 0.012 | 0.042 | 0.144 |
| Clustering Only (5%) | 0.401 | 0.146 | 0.018 | 0.041 | 0.152 |
| MAP4CS (5%) | 0.432 | 0.140 | 0.027 | 0.044 | 0.161 |