cs.SESep 29, 2026

From Codebase to Culprit (C2C): Reducing the Search Space for Bugs with Semantic Retrieval and Hierarchical Reinforcement Learning

Authors: Ankur Garg, Corey Yang-Smith, Rishav Rishav, Ahmad Abdellatif, Samira Ebrahimi Kahou

Organizations: University of Calgary, Canada

Abstract

We introduce C2C (From Codebase to Culprit), a framework for precise bug localization that progressively reduces the debugging search space across multiple levels of granularity: files, functions, and lines of code. To mirror developer's natural top-down debugging workflows, C2C integrates semantic retrieval and Hierarchical Reinforcement Learning (HRL) in a two-stage process. First, it performs recall-oriented retrieval of buggy candidates via semantic vector similarity search using bug-report text, including available stack-trace information, against a database of embeddings, where the embeddings are fine-tuned via contrastive learning with CodeBERT. Building on this reduced search space, the HRL framework incrementally localizes bugs, reasoning from files to functions and ultimately to individual lines of code. Unlike prior approaches which operate at a single granularity, C2C enables multi-resolution localization while maintaining contextual consistency across decisions. Experiments on real-world Java and Python datasets demonstrate that C2C improves retrieval precision and localization accuracy. Ablation studies further highlight the contributions of hierarchical decomposition, structured learning signals, and reward shaping in advancing multi-level bug localization.

Figures & tables

Explore similar work

May 18, 2026cs.SE

BLAgent: Agentic RAG for File-Level Bug Localization

Bug localization remains a key bottleneck for large language model (LLM)-based software maintenance, where accurately identifying faulty code is essential for debugging, root cause analysis, triage, and automated program repair (APR). File-level bug localization is especially critical in hierarchical localization and repair pipelines, where incorrect file selection can propagate to downstream stages such as function-level localization and patch generation. While Retrieval-Augmented Generation (RAG) offers a promising way to ground LLMs in repository context, existing RAG pipelines often rely on static retrieval and lack the reasoning needed to accurately identify faulty code. In this work, we present BLAgent, a novel agentic RAG framework for file-level bug localization that integrates three key ideas: (i) code structure-aware repository encoding with path-augmented AST-based chunking, (ii) dual-perspective query transformation that captures both structural and behavioral signals from bug reports, and (iii) two-phase agentic reranking that combines symbolic inspection with evidence-grounded reasoning. Unlike prior graph-based or multi-hop agentic approaches, BLAgent adopts a bounded reasoning strategy that limits LLM-based inspection and reranking to a compact, retrieval-filtered set of candidate files, avoiding open-ended repository traversal. This design balances localization accuracy with computational cost. On SWE-bench-Lite, BLAgent attains over 78% Top-1 accuracy with open-source models and over 86% with a closed-source model, while being over 18x cheaper than the strongest baseline using the same model. When integrated into an APR framework, BLAgent improves end-to-end repair success by up to 25%.
Jun 8, 2026cs.SE

Multi-task LLMs for Bug Classification: Efficient Inference with Auxiliary Decoding Heads

The rapid adoption of LLM-powered code generation has dramatically accelerated software development, yet effective verification methods remain severely underdeveloped. Existing bug localization techniques are either prohibitively expensive, requiring minutes of agentic reasoning and thousands of generated tokens per file, and/or operate at coarse function-level granularity unsuitable for precise debugging. While works that focus on line-level granularity and are more light-weight are often limited in their performance or context size. We introduce a novel line-level bug localization approach that addresses these limitations through three key contributions: (1) a token alignment algorithm that overcomes fundamental tokenization challenges in previous work, (2) a lightweight multi-task LLM for bug localization (MLC) enabling efficient line-level bug classification, and (3) an optimized training recipe for multi-line prediction. Our method achieves state-of-the-art performance among similar setups on line-level bug localization with full-file context. At the same time we reach comparable performance to agentic approaches on Defects4J and PypiBugs benchmarks while reducing inference latency by orders of magnitudes, requiring only a single generated token per file. We further demonstrate strong generalization by introducing and evaluating on a small out-of-domain evaluation datasets in Python. We will open source our code, models, and datasets upon acceptance.
Dec 18, 2025cs.SE

SpIDER: Spatially Informed Dense Embedding Retrieval for Software Issue Localization

Retrieving code functions, classes or files relevant to a user query, bug report or feature request from large codebases is a fundamental challenge for Large Language Model (LLM)-based coding agents. Agentic approaches typically employ sparse methods like BM25 or dense embedding strategies to identify semantically relevant units. While dense embedding approaches can outperform BM25 by large margins, both ignore the graph-structured characteristics of the codebase. To address this, we propose SpIDER (Spatially Informed Dense Embedding Retrieval), a dense retrieval approach that integrates LLM-based reasoning with graph-based exploration of the codebase. We further introduce SpIDER-Bench, a graph-structured benchmark curated from SWEPolyBench, SWEBench-Verified and Multi-SWE-bench, spanning Python, Java, JavaScript and TypeScript repositories. SpIDER's graph-based candidate expansion attaches a structural reason for inclusion to each surfaced function (its seed and the edge type linking them), making the candidate set auditable while keeping the retrieval budget fixed. The graph is built from per-repository syntax trees, so it can be constructed on-demand at the start of a developer session rather than precomputed offline. Empirical results show that SpIDER consistently improves dense retrieval Recall@20 across every language and benchmark in SpIDER-Bench: by at least 13% relative (+0.05 to +0.12 absolute) along containment edges, rising to at least 27% relative (+0.11 to +0.19 absolute) once call edges are explored.