Large language models (LLMs) frequently generate confident yet factually incorrect content when used for language generation (a phenomenon often known as hallucination). Retrieval augmented generation (RAG) tries to reduce factual errors by identifying information in a knowledge corpus and putting it in the context window of the model. While this approach is well-established for document-structured data, it is non-trivial to adapt it for Knowledge Graphs (KGs), especially for queries that require multi-node/multi-hop reasoning on graphs. We introduce UltRAG, a training-free KG-RAG recipe that combines LLM query generation, a fully inductive neural query executor, and LLM arbitration. This off-the-shelf composition achieves state-of-the-art results on Knowledge Graph Question Answering (KGQA) tasks without retraining the LLM or executor, while enabling language models to interface with Wikidata-scale graphs (116M entities, 1.6B relations) at comparable or lower costs. Our ablation studies indicate that these gains come from the full system design rather than from any single component.
Figures & tables
Figure 1 : UltRAG pipeline. The LLM is provided with the syntactic rules for queries and the relation types. Ground-truth seed entities (Turing Award, Deep Learning, etc.) may be given, but, if not, an entity linking step takes place. The generated query is then neurally executed against the knowledge graph (each node receives a probability to be an answer at this stage). The most likely query answers are fed back to the LLM, which weighs both the returned probabilities and the semantic meaning of entities, and produces a final answer set.
Figure 2 : Comparison of UltraQuery vs symbolic query execution on GTSQA. Both receive identical queries generated by the LLM. Number inside brackets denotes projections, concatenation of expression denotes intersection. Best viewed on screen.
GTSQA
KGQAGen-10k
Model
Hits
Recall
F1
Hits
Recall
F1
GPT-5-mini
33.42
31.92
31.82
62.18
59.17
59.58
GPT-4.1
33.72
32.63
32.36
56.21
53.38
53.93
GPT-5
46.36
44.40
44.16
67.96
65.08
65.27
ToG
64.73
61.99
62.06
80.97
75.03
76.10
GCR
60.83
58.41
58.15
86.11
79.89
80.22
Table 1 : Comparison on question-specific graphs, for GTSQA (WikiKG2) and KGQAGen-10k (Wikidata).
Ground-truth seed nodes
Entity Linking
GTSQA (WikiKG2)
GTSQA (Wikidata)
KGQAGen-10k
GTSQA (WikiKG2)
GTSQA (Wikidata)
KGQAGen-10k
Model
Hits
Recall
F1
Hits
Recall
F1
Hits
Recall
F1
Hits
Recall
F1
Hits
Recall
F1
Hits
Recall
F1
RoG
72.63
70.73
69.81
62.55
59.81
58.93
86.25
82.60
82.68
72.00
69.97
68.88
60.64
57.66
56.85
82.42
78.89
79.22
GNN-RAG
64.98
62.89
62.70
-
-
-
81.04
76.66
77.56
51.91
49.36
49.59
-
-
-
72.89
69.04
69.53
SubgraphRAG (200)
73.98
71.14
70.91
63.29
59.60
59.56
80.47
75.88
76.54
71.82
69.37
69.01
58.61
55.43
55.04
76.68
72.63
73.17
UltRAG -OTS
90.81
89.39
87.18
86.74
84.47
82.08
90.62
89.41
87.63
85.70
83.71
81.08
72.93
70.48
66.60
83.98
82.59
80.58
Table 2 : Comparison on PPR subgraphs extracted from the ground-truth seed entities and seed entities obtained via Entity Linking. PPR subgraphs contain up to 30,000 nodes. All models use GPT-5 as reasoning LLM. We color-code settings that require transductive and inductive reasoning for the considered baselines; for GTSQA (Wikidata), we used baseline models originally trained on GTSQA (WikiKG2). We leave a ‘-’ where a baseline cannot be applied.
KG
Model
API cost
Non-API time (s)
Input tokens
Output tokens
Cache hit %
WikiKG2
RoG
0.011$
1.9±0.5
0.9K
2.2K
0
GNN-RAG
0.012$
9.9±2.3
1.5K
2.1K
0
SubgraphRAG (200)
0.012$
16.7±3.8
2.7K
2.1K
0
UltRAG -OTS
0.014$
0.10±0.0
23K
2.2K
93.77
Wikidata
RoG
0.013$
4.0±1.1
1.1K
2.5K
0
SubgraphRAG (200)
0.014$
20.9±3.5
3.1K
2.4K
0
Table 3 : Average efficiency per query on PPR subgraphs extracted for GTSQA from WikiKG2 (top) and Wikidata (bottom) with ground-truth seed nodes. All models use GPT-5 as reasoning LLM. Non-API running times are measured on GH200 chips, excluding API calls and PPR computation.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3 : BetaE format representation of the Turing Award query example from Figure 1 . Entities are abbreviated.
Class
n
Valid
Non-parseable
Invalid %
(1)
150
150
0
0.0
(2)
150
150
0
0.0
((1)(1))(1)
60
48
12
20.0
(2(1)(1))
55
40
15
27.3
((2)(1))
65
55
10
15.4
Appendix
Table 4 : Invalid LLM-generated queries under the BetaE-style format for representative GTSQA classes.
Figure 4 : Haskell-like grammar definitions for the old BetaE format (left) and our preferred DSL (right), with the transformed Turing Award example below.
Metric
# hops = 1
# hops = 2
# hops = 3
# hops = 4
(1)
(1)(1)
(1)(1)(1)
(1)(1)(1)(1)
Avg
(2)
(2)(1)
((1)(1))
(2)(2)
((1)(1))(1)
Avg
(3)
((2)(1))
(3)(1)
(2(1)(1))
Avg
(4)
Avg
Neural executor (UltraQuery)
MRR
93.25
94.48
97.96
96.36
95.51
85.24
86.19
93.28
95.17
84.01
88.78
82.78
92.00
90.30
88.35
88.36
82.54
82.54
Hit@1
91.33
93.17
96.94
95.93
94.34
82.67
82.48
92.00
93.05
80.00
86.04
77.50
87.69
88.00
85.00
84.55
76.22
76.22
Hit@3
94.67
95.00
98.98
96.75
96.35
87.33
87.99
93.60
97.60
86.17
90.54
87.87
93.85
90.67
92.27
91.16
88.22
88.22
Hit@10
96.00
97.78
98.98
96.75
97.38
88.67
92.13
96.80
97.84
93.33
93.75
90.65
100.00
96.00
93.64
95.07
92.22
92.22
Appendix
Table 5 : Comparison of neural (UltraQuery) vs symbolic query execution on LLM-generated queries from GTSQA using WikiKG2 subgraphs (in %). Class notation as in Cattaneo et al. – number inside brackets denotes hops, concatenation of expression denotes intersection. Both executors receive identical queries generated by the LLM.
Metric
# hops = 1
# hops = 2
# hops = 3
# hops = 4
(1)
(1)(1)
(1)(1)(1)
(1)(1)(1)(1)
Avg
(2)
(2)(1)
((1)(1))
(2)(2)
((1)(1))(1)
Avg
(3)
((2)(1))
(3)(1)
(2(1)(1))
Avg
(4)
Avg
Neural executor (UltraQuery)
MRR
91.07
86.74
91.14
90.19
89.78
77.01
78.43
88.19
77.87
74.89
79.28
61.04
75.81
75.44
72.64
71.23
70.69
70.69
Hit@1
90.00
83.65
88.61
87.26
87.38
74.00
74.69
85.60
69.54
69.17
74.60
54.44
66.15
66.67
64.97
63.06
65.22
65.22
Hit@3
91.33
88.38
93.88
93.09
91.67
78.67
81.73
89.60
82.73
78.61
82.27
65.28
86.15
81.33
76.50
77.31
73.44
73.44
Hit@10
93.33
92.38
93.88
94.31
93.47
85.33
84.88
94.40
93.17
83.00
88.16
72.22
89.23
90.00
86.32
84.44
82.44
82.44
Appendix
Table 6 : Comparison of neural (UltraQuery) vs symbolic query execution on LLM-generated queries from GTSQA on PPR subgraphs extracted from Wikidata (in %). Both executors receive identical queries generated by the LLM.
Executor
Hits
Recall
F1
UltraQuery
86.74
84.47
82.08
Symbolic execution
79.16
78.06
76.92
Drop
-7.58
-6.41
-5.16
Appendix
Table 7 : Executor replacement ablation on GTSQA (Wikidata) in the ground-truth seed setting. Scores are final-answer Hits/Recall/F1, so the comparison includes the downstream arbitration stage.
Metric
# hops = 1
# hops = 2
# hops = 3
# hops = 4
(1)
(1)(1)
(1)(1)(1)
(1)(1)(1)(1)
Avg
(2)
(2)(1)
((1)(1))
(2)(2)
((1)(1))(1)
Avg
(3)
((2)(1))
(3)(1)
(2(1)(1))
Avg
(4)
Avg
With ground-truth queries
MRR
99.33
100.00
100.00
100.00
99.83
99.56
100.00
100.00
100.00
100.00
99.91
98.61
100.00
99.83
97.12
98.89
96.07
96.07
Hit@1
98.67
100.00
100.00
100.00
99.67
99.33
100.00
100.00
100.00
100.00
99.87
97.78
100.00
99.67
94.55
98.00
94.00
94.00
Hit@3
100.00
100.00
100.00
100.00
100.00
100.00
100.00
100.00
100.00
100.00
100.00
99.44
100.00
100.00
100.00
99.86
98.00
98.00
Hit@10
100.00
100.00
100.00
100.00
100.00
100.00
100.00
100.00
100.00
100.00
100.00
99.44
100.00
100.00
100.00
99.86
98.00
98.00
Appendix
Table 8 : UltraQuery per-class performance metrics on WikiKG2 subgraphs contained in GTSQA (in %). To generate ground-truth structured queries (i.e. queries provided by the oracle), given a subgraph with the answer as the root node, we perform a breadth-first traversal, identify leaves, and work upwards to construct the query.
LLM
Hits
Recall
F1
API cost
UltRAG
- GPT-5 ×2
86.74
84.47
82.08
0.017$
- GPT-5 → GPT-5-mini
82.37
77.83
78.41
0.009$
- GPT-5-mini → GPT-5
80.71
78.19
75.72
0.011$
- DeepSeek-reasoner ×2
75.77
72.40
72.05
0.005$
- GPT-5-mini ×2
75.71
71.08
71.60
0.004$
Appendix
Table 9 : Comparison of UltRAG -OTS with different LLMs on GTSQA (Wikidata) using ground-truth seed entities (in %). PPR graphs have up to 30,000 nodes. With A→B , we describe a setting where we used LLM A for query generation and B for arbitration.
KGQAGen-10K
API cost
Non-API time (s)
Input tokens
Output tokens
Cache hit %
RoG
0.008$
3.1±0.9
1.4K
1.4K
0
GNN-RAG
0.008$
25.8±4.4
1.4K
1.4K
0
SubgraphRAG (200)
0.009$
27.5±3.9
3.1K
1.5K
0
UltRAG -OTS
0.019$
0.21±0.1
69K
2.8K
96.06
- GPT-5-mini
0.004$
0.15±0.1
69K
2.4K
95.53
Appendix
Table 10 : Average efficiency per query on PPR subgraphs extracted for KGQAGen-10K from Wikidata with ground-truth seed nodes. All models use GPT-5 as reasoning LLM. Non-API running time measured on GH200 chips, excluding API calls and PPR computation.
PPR nodes
Candidate top- k
F1
5K
10
80.99
5K
50
80.33
5K
100
80.82
30K
10
81.45
30K
50
82.08
30K
100
81.81
Appendix
Table 11 : Sensitivity of UltRAG -OTS to PPR subgraph size and the number of executor candidates passed to the arbitrator. We report F1 on GTSQA (Wikidata) with ground-truth seed entities.
Setting
((1)(1))
((1)(1))(1)
((2)(1))
(1)
(1)(1)
(1)(1)(1)
(1)(1)(1)(1)
(2(1)(1))
(2)
(2)(1)
(2)(2)
(3)
(3)(1)
(4)
Single-iter F1
93.87
77.06
88.72
96.22
90.59
89.52
84.25
89.24
93.11
84.39
94.48
85.09
88.79
83.73
Multi-iter F1
93.96
82.44
92.82
96.93
91.86
91.63
85.79
85.57
94.67
87.42
95.91
87.59
92.72
91.13
Appendix
Table 12 : Single-iteration and multi-iteration UltRAG -OTS F1 on GTSQA query classes. Class notation follows Cattaneo et al. [8] .
Setting
GTSQA
KGQAGen-10k
Top-1
84.9
80.9
Top-5
35.3
34.0
Top- p 0.3
68.6
59.6
Top- p 0.5
73.4
57.0
Top- p 0.7
72.1
54.7
Relative threshold
81.1
80.8
Appendix
Table 13 : F1 of UltRAG -OTS when the LLM arbitrator is replaced by deterministic conversions from executor probabilities to answer sets. The full LLM arbitrator remains the most reliable conversion rule.
Setting
Hits
Recall
F1
No randomization
92.66
91.05
89.29
Randomized relation IDs
93.90
92.37
91.46
Appendix
Table 14 : UltRAG -OTS performance on GTSQA when relation identifiers are randomly reassigned per sample. Scores are reported as Hits/Recall/F1.
Model
Hits
Recall
Precision
F1
GPT-5
73.0
53.6
66.3
55.5
RoG
88.5
78.4
84.7
78.5
GNN-RAG
88.8
78.1
85.4
78.7
SubgraphRAG
87.2
76.1
82.7
76.1
GCR
90.0
71.0
85.4
73.0
UltRAG -OTS
88.0
77.7
82.3
76.7
Appendix
Table 15 : WebQSP results. Scores are in %; Δ is relative to the best non- UltRAG baseline in each column.
Model
Hits
Recall
Precision
F1
GPT-5
60.5
53.2
56.3
52.7
GCR
69.3
62.1
66.3
62.2
GNN-RAG
75.9
70.3
71.7
69.2
RoG
75.8
70.9
70.5
68.8
SubgraphRAG
76.4
70.6
71.1
68.8
UltRAG -OTS
73.7
68.8
68.0
66.3
Appendix
Table 16 : CWQ results. Scores are in %; Δ is relative to the best non- UltRAG baseline in each column.
Dataset
UltRAG -OTS
Best baseline
Speed-up
WikiKG2
0.14s
1.94s
14.1 ×
Wikidata
1.17s
5.02s
4.3 ×
Appendix
Table 17 : Runtime audit after including PPR sampling and FAISS retrieval. Times are average seconds per query; speed-up compares UltRAG -OTS to the fastest baseline under the same accounting.