CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding
Authors: Federico Pennino, Andrea Gurioli, Stefano Zacchiroli, Maurizio Gabbrielli, Paolo Ferragina
Organizations: Università di Bologna Bologna, Italy · LTCI Télécom Paris, Institut Polytechnique de Paris Palaiseau, France · DISI Università di Bologna Bologna, Italy · Sant’Anna School of Advanced Studies Pisa, Italy
Public software repositories, like GitHub and Software Heritage Archive, store billions of files, yet extracting their implicit engineering knowledge ---i.e., the algorithms they implement, the paradigms they follow, the patterns they instantiate, and the application domains they serve--- remains challenging, as current tools are constrained to syntactic and token-level analysis. We present a pipeline for building an open-taxonomy semantic annotation of source code using a code-specialised Large Language Model. The extracted entities are grounded in Wikidata through a three-stage linking procedure: a deterministic SPARQL stage handles unambiguous entities, a Deep Research Agent resolves the residual long tail, and a hierarchy-rollup stage imports the parent-of closure of each resolved Wikidata identifier. The resulting annotations are materialised as a source-code-specific open-taxonomy knowledge graph. We further introduce a calibrated quality-assurance protocol that quantifies annotation precision by combining a small human gold set with an LLM-as-a-judge filter. We applied our pipeline to the 167 million files of the Stack-Edu corpus, creating the first known large-scale open-taxonomy knowledge graph for source code. Our graph, named CodeGraph, contains approximately 158 million nodes, which include around 145 million files, about 63,000 extracted concept entities (such as algorithms, paradigms, design patterns, and application domains), and roughly 19,800 grounded Wikidata entities. Furthermore, CodeGraph features approximately 1 billion typed edges that connect files to their respective concepts, link these concepts to their grounded Wikidata identifiers, and relate them to their parent categories, covering 14 programming languages.
Figures & tables
Resource
Files
# Langs.
Semantic metadata
Wikidata
Corpus
Knowledge graph
CodeSearchNet ( Husain et al., 2019 )
6 M func.
6
docstrings; relevance labels
✗
✓
✗
CodeXGLUE ( Lu et al., 2021 )
task-specific †
8
task-specific labels
✗
✓
✗
CoIR ( Li et al., 2025 )
task-specific ‡
14
task-specific retrieval labels
✗
✓
✗
Project CodeNet ( Puri et al., 2021 )
14 M subm.
55
problem ID; runtime/memory; status
✗
✓
✗
Software Heritage ( Di Cosmo and Zacchiroli, 2017 )
28 B files
100+
repository and archival metadata
✗
✓
✗
The Stack v2 ( Lozhkov et al., 2024 )
3.3 B files
600+
language label only
✗
✓
✗
Table 1. Positioning of the methodology presented in this paper relative to representative large-scale code corpora and code knowledge graphs. “Files” refers to the unit reported by each resource, including source files, functions, or submissions, as indicated. “Annotation type” refers to the metadata explicitly exposed.
Figure 1. Overview of the CodeGraph methodology. Source files from Stack-Edu are processed through an LLM inference stage (Qwen3-Coder-30B-A3B-Instruct) that emits four orthogonal categories of semantic annotations: domains, algorithms, paradigms, and design patterns. Annotations undergo a post-processing stage comprising domain canonicalisation and Wikidata linking, and are integrated into CodeGraph as a typed property graph that supports semantically grounded queries across both the local vocabulary and Wikidata.
Figure 2. CodeGraph language distribution. Number of retained source files per programming language, showing the composition of the :File nodes materialised in CodeGraph.
Figure 4. Three-stage grounding of canonical :Domain , :Algorithm , :Paradigm and :DesignPattern concepts in Wikidata. Stage 1 (SPARQL): a full-text query with a class-hierarchy filter ( wdt:P31 / wdt:P279 ) resolves unambiguous cases deterministically. Stage 2 (Agent): ambiguous or empty results are delegated to a Deep Research Agent backed by Qwen3.6-27B, which selects the best-matching QID or rejects all candidates. Stage 3 (Hierarchy rollup): a transitive wdt:P279 traversal imports the parent-of closure of each resolved QID, materialising a multi-level taxonomy of :WikidataEntity nodes.
Figure 5. Relational schema of the graph linking :File , the semantic-concept nodes ( :Domain , :Algorithm , :Paradigm , :DesignPattern ) and :WikidataEntity nodes. This figure is the Wikidata grounding counterpart of Figure 6 .
Figure 6. Core CodeGraph property-graph schema. :File nodes represent source artifacts and are linked to the semantic concepts extracted from them: application domains, implemented algorithms, programming paradigms, and design patterns. Algorithm nodes are further connected to derived algorithm-category and complexity nodes. This figure is the graph-model counterpart of Figure 3 .
Model
Accuracy
Precision
Recall
F1
Claude Haiku 4.5
78.82
80.69
95.43
87.44
Gemini 3 Flash
79.07
81.47
94.50
87.50
GPT-5.1-Codex-Mini
75.97
82.86
87.00
84.88
Grok Code Fast 1
77.34
80.26
93.97
86.57
Devstral-2 2512
80.62
82.05
96.00
88.48
DeepSeek V3.2
77.69
81.74
91.79
86.47
Table 2. LLM-based verifier selection. Agreement metrics between LLM annotations and the human gold standard. Best value per column is shown in bold.
Figure 7. Silver-tier validation on 10,000 files. Verifier acceptance rate per semantic axis; cross-language acceptance (not shown) is uniform within an 87–90% band.
Figure 8. Verifier acceptance on Wikidata-grounded annotations. Acceptance rate per semantic axis restricted to the subset of annotations that were successfully linked to a Wikidata entity.
Figure 9. Graph composition by entity type and Wikidata coverage. For each entity type, the figure reports the number of unique concepts extracted by the pipeline (blue) and the subset successfully grounded to Wikidata (green).
Figure 10. Cross-domain reuse of computational techniques. Incidence of selected algorithms across representative application domains, illustrating the recurrence of shared techniques across substantively different domains.
Figure 11. Top application domains by file count (Wikidata-grounded). Number of files associated with each of the twelve most frequent domain labels after label consolidation and Wikidata grounding.
Figure 12. Top algorithms by file count (Wikidata-grounded). Number of files associated with each of the twelve most frequent algorithm labels after label consolidation and Wikidata grounding.
Large language models have shown strong performance on software engineering (SE) tasks, yet understanding large industrial repositories remains challenging. Existing methods often retrieve only local fragments and fail to recover the broader task-relevant context needed for complex repository-level tasks. We present DeepDiscovery, a task-level repository-understanding method for large industrial codebases. DeepDiscovery uses a two-stage \textit{Location--Inference} framework to localize high-confidence task anchors and recover broader task-relevant context over multi-relational repository structure under budget constraints. Across controlled method-level evaluation, organization-internal industrial repository-understanding scenarios, and end-to-end evaluation on SWE-bench Verified, DeepDiscovery consistently improves task-relevant file recovery and downstream SE performance. On 27 medium-scale tasks, DeepDiscovery achieves the best file recovery quality among five representative baselines without offline preprocessing. On organization-internal industrial tasks from a production-scale integrated codebase ecosystem, including 27 medium-scale tasks and 40 large-scale tasks, DeepDiscovery improves Full Recall Rate across multiple AI coding systems, with absolute gains ranging from 1.6 to 9.2 percentage points on large subprojects and from 2.5 to 7.4 percentage points on medium-scale subprojects. In a controlled end-to-end evaluation on SWE-bench Verified, a system equipped with DeepDiscovery achieves a 78.6% Solve Rate, outperforming the corresponding baseline by 8.2 percentage points. These results suggest that stronger task-level repository understanding can improve coding-agent performance on complex SE tasks.
Jiawei He, Weisong Sun, Mengyu Shi +4
AMAP, Alibaba Group · Nanyang Technological University · Nanjing University +1
Maintaining up-to-date code documentation is difficult in fast-moving repositories because design knowledge is scattered across source files and pull requests. We present CODENS , a system that turns pull requests into living, accessible, and queryable documentation for production codebases. CODENS incrementally builds a typed software knowledge graph from pull requests, enriches components through schema-driven semantic extraction, derives typed relations between them, and exposes the resulting knowledge through three retrieval modes, including agent-guided graph traversal for repository-level question answering. The system also preserves semantic change history across pull requests and integrates both answer-quality and operational evaluation metrics. We evaluate CODENS on a client Ruby on Rails project in production. Results show that CODENS produces highly relevant and well-grounded answers, while qualitative feedback highlights a remaining challenge in concise, documentation-oriented synthesis.
We present Code-QA-Bench, a fully automated framework for synthesizing repository-level code understanding benchmarks that separates genuine code comprehension from documentation recall and pretraining memorization. The framework makes two methodological contributions: (1) an answer-first generation pipeline where a tool-equipped agent explores source code to produce verified gold answers before deriving questions, ensuring every task is grounded in real code structure; and (2) a three-condition experimental design evaluating agents under closed-book (no repository), code-only (documentation removed), and documented (full repository) conditions, with deltas directly quantifying documentation utility and memorization. We generate 528 code-derivable and 100 doc-dependent tasks across 10 Python repositories from SWE-Bench, scored by an LLM judge on accuracy, completeness, and specificity. Experiments on four frontier models reveal that code access is the dominant factor (+0.23 mean gain over closed-book), documentation provides modest additional benefit (+0.071 on doc-dependent tasks), and code-only ≈ documented on code-derivable tasks, validating the design. The framework is open-source and applicable to any well-documented Python repository.