From Codebase to Culprit (C2C): Reducing the Search Space for Bugs with Semantic Retrieval and Hierarchical Reinforcement Learning
Organizations: University of Calgary, Canada
Abstract
We introduce C2C (From Codebase to Culprit), a framework for precise bug localization that progressively reduces the debugging search space across multiple levels of granularity: files, functions, and lines of code. To mirror developer's natural top-down debugging workflows, C2C integrates semantic retrieval and Hierarchical Reinforcement Learning (HRL) in a two-stage process. First, it performs recall-oriented retrieval of buggy candidates via semantic vector similarity search using bug-report text, including available stack-trace information, against a database of embeddings, where the embeddings are fine-tuned via contrastive learning with CodeBERT. Building on this reduced search space, the HRL framework incrementally localizes bugs, reasoning from files to functions and ultimately to individual lines of code. Unlike prior approaches which operate at a single granularity, C2C enables multi-resolution localization while maintaining contextual consistency across decisions. Experiments on real-world Java and Python datasets demonstrate that C2C improves retrieval precision and localization accuracy. Ablation studies further highlight the contributions of hierarchical decomposition, structured learning signals, and reward shaping in advancing multi-level bug localization.
Figures & tables
| Characteristic | Elasticsearch | SWE-bench Verified |
| Dataset scale | ||
| Repositories | 1 | 11 |
| Bug reports | 326 | 393 |
| Bug-location triplets | 5,342 | 1,972 |
| Unique buggy files | 691 | 477 |
| Unique non-buggy files | 5,326 | 1,929 |
| Method | File Hit@1 | File Hit@5 | File MRR | Func Hit@1 | Func MRR | Line Hit@10 | Line MRR |
| RLocator | 36.62 | – | – | – | – | – | – |
| LLMAO | – | – | – | 4.90 | 0.0617 | 41.76 | 0.2149 |
| C2C (TWF, E2E) | 36.62 | 80.28 | 0.5005 | 43.66 | 0.5450 | 54.55 | 0.2932 |
| C2C (TWF, ORACLE) | 36.62 | 80.28 | 0.5005 | 54.55 | 0.6298 | 73.06 | 0.3863 |
| C2C (TWF, LLMAO Style) | – | – | – | 53.52 | 0.6574 | 58.33 | 0.2251 |
| Method | File Hit@1 | File Hit@5 | File MRR | Func Hit@1 | Func MRR | Line Hit@10 | Line MRR |
| RLocator | 21.74 | – | – | – | – | – | – |
| LLMAO | – | – | – | 4.39 | 0.0719 | 23.03 | 0.1149 |
| C2C (TWF, E2E) | 21.74 | 76.09 | 0.4156 | 19.57 | 0.2815 | 50.00 | 0.1000 |
| C2C (TWF, ORACLE) | 21.74 | 76.09 | 0.4156 | 32.14 | 0.4298 | 51.72 | 0.2846 |
| C2C (TWF, LLMAO Style) | – | – | – | 23.91 | 0.3548 | 0 | 0 |
| Multi-File Recall | Multi-Line Recall | ||
| Dataset | Retrieval | File Agent | Full Cascade |
| Elasticsearch | 36.36 | 54.29 | 17.58 |
| SWE-bench | 65.00 | 66.25 | 12.50 |
| Training configuration | @3 | @5 | @10 | @20 | @50 | |||||
| Hit | Recall | Hit | Recall | Hit | Recall | Hit | Recall | Hit | Recall | |
| No contrastive loss | 0.16 | 0.04 | 0.16 | 0.04 | 0.16 | 0.04 | 3.29 | 0.83 | 6.18 | 1.56 |
| 1st-iteration contrastive loss | 23.94 | 15.15 | 41.95 | 21.21 | 55.93 | 33.33 | 56.57 | 36.36 | 75.21 | 57.58 |
| 2nd-iteration contrastive loss | 29.66 | 15.15 | 41.53 | 21.21 | 57.84 | 36.36 | 61.23 | 39.39 | 72.88 | 51.52 |
| Elasticsearch | SWE-bench Verified | ||||||
| Component | Metric | Base | Ablated | Change | Base | Ablated | Change |
| LSTM Encoder | File MRR | 0.5225 | 0.5653 | +0.0428 | 0.4112 | 0.3627 | 0.0485 |
| Func MRR | 0.5450 | 0.4893 | 0.0557 | 0.2259 | 0.2165 | 0.0094 | |
| Line MRR | 0.2932 | 0.2925 | 0.0007 | 0.1000 | 0.0000 | 0.1000 | |
| Hierarchy | Line MRR | 0.2932 | 0.0682 | 0.2250 | 0.1000 | 0.0000 | 0.1000 |
| Elasticsearch | SWE-bench Verified | |||
| Metric | Intermediate | Sparse | Intermediate | Sparse |
| File MRR | 0.5225 | 0.5005 | 0.4112 | 0.3413 |
| Function MRR | 0.5450 | 0.5223 | 0.2259 | 0.2122 |
| Line MRR | 0.2932 | 0.2377 | 0.1000 | 0.0000 |
| Elasticsearch | SWE-bench Verified | ||||||||
| Level | Metric | PT | VT | T1B1 | TWF | PT | VT | T1B1 | TWF |
| File | Hit@1 | 0.3662 | 0.3662 | 0.3662 | 0.3662 | 0.2174 | 0.2174 | 0.2174 | 0.2174 |
| Hit@5 | 0.8028 | 0.8028 | 0.8028 | 0.8028 | 0.7391 | 0.7391 | 0.7391 | 0.7391 | |
| Hit@10 | 0.8028 | 0.8028 | 0.8028 | 0.8028 | 0.7391 | 0.7391 | 0.7391 | 0.7391 | |
| MRR | 0.5225 | 0.5225 | 0.5225 | 0.5225 | 0.4112 | 0.4112 | 0.4103 | 0.4112 | |
| Function | Hit@1 | 0.3727 | 0.4091 | 0.5500 | 0.5455 | 0.1339 | 0.1429 | 0.2500 | 0.3214 |
| Dataset | Level | TWF | T1B1 | Cliff’s | -value |
| Elasticsearch | Function | 0.7047 | 0.7215 | 0.507 | |
| Elasticsearch | Line | 0.6100 | 0.5432 | 0.002 | |
| SWE-bench Verified | Function | 0.4099 | 0.3903 | 0.173 | |
| SWE-bench Verified | Line | 0.2930 | 0.2771 | 0.361 |