UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG
Organizations: Graphcore Research
Abstract
Large language models (LLMs) frequently generate confident yet factually incorrect content when used for language generation (a phenomenon often known as hallucination). Retrieval augmented generation (RAG) tries to reduce factual errors by identifying information in a knowledge corpus and putting it in the context window of the model. While this approach is well-established for document-structured data, it is non-trivial to adapt it for Knowledge Graphs (KGs), especially for queries that require multi-node/multi-hop reasoning on graphs. We introduce UltRAG, a training-free KG-RAG recipe that combines LLM query generation, a fully inductive neural query executor, and LLM arbitration. This off-the-shelf composition achieves state-of-the-art results on Knowledge Graph Question Answering (KGQA) tasks without retraining the LLM or executor, while enabling language models to interface with Wikidata-scale graphs (116M entities, 1.6B relations) at comparable or lower costs. Our ablation studies indicate that these gains come from the full system design rather than from any single component.
Figures & tables
| GTSQA | KGQAGen-10k | |||||
| Model | Hits | Recall | F1 | Hits | Recall | F1 |
| GPT-5-mini | 33.42 | 31.92 | 31.82 | 62.18 | 59.17 | 59.58 |
| GPT-4.1 | 33.72 | 32.63 | 32.36 | 56.21 | 53.38 | 53.93 |
| GPT-5 | 46.36 | 44.40 | 44.16 | 67.96 | 65.08 | 65.27 |
| ToG | 64.73 | 61.99 | 62.06 | 80.97 | 75.03 | 76.10 |
| GCR | 60.83 | 58.41 | 58.15 | 86.11 | 79.89 | 80.22 |
| Ground-truth seed nodes | Entity Linking | |||||||||||||||||
| GTSQA (WikiKG2) | GTSQA (Wikidata) | KGQAGen-10k | GTSQA (WikiKG2) | GTSQA (Wikidata) | KGQAGen-10k | |||||||||||||
| Model | Hits | Recall | F1 | Hits | Recall | F1 | Hits | Recall | F1 | Hits | Recall | F1 | Hits | Recall | F1 | Hits | Recall | F1 |
| RoG | 72.63 | 70.73 | 69.81 | 62.55 | 59.81 | 58.93 | 86.25 | 82.60 | 82.68 | 72.00 | 69.97 | 68.88 | 60.64 | 57.66 | 56.85 | 82.42 | 78.89 | 79.22 |
| GNN-RAG | 64.98 | 62.89 | 62.70 | - | - | - | 81.04 | 76.66 | 77.56 | 51.91 | 49.36 | 49.59 | - | - | - | 72.89 | 69.04 | 69.53 |
| SubgraphRAG (200) | 73.98 | 71.14 | 70.91 | 63.29 | 59.60 | 59.56 | 80.47 | 75.88 | 76.54 | 71.82 | 69.37 | 69.01 | 58.61 | 55.43 | 55.04 | 76.68 | 72.63 | 73.17 |
| UltRAG -OTS | 90.81 | 89.39 | 87.18 | 86.74 | 84.47 | 82.08 | 90.62 | 89.41 | 87.63 | 85.70 | 83.71 | 81.08 | 72.93 | 70.48 | 66.60 | 83.98 | 82.59 | 80.58 |
| KG | Model | API cost | Non-API time (s) | Input tokens | Output tokens | Cache hit % |
| WikiKG2 | RoG | 0.011$ | 0.9K | 2.2K | 0 | |
| GNN-RAG | 0.012$ | 1.5K | 2.1K | 0 | ||
| SubgraphRAG (200) | 0.012$ | 2.7K | 2.1K | 0 | ||
| UltRAG -OTS | 0.014$ | 23K | 2.2K | 93.77 | ||
| Wikidata | RoG | 0.013$ | 1.1K | 2.5K | 0 | |
| SubgraphRAG (200) | 0.014$ | 3.1K | 2.4K | 0 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Class | Valid | Non-parseable | Invalid % | |
| (1) | 150 | 150 | 0 | 0.0 |
| (2) | 150 | 150 | 0 | 0.0 |
| ((1)(1))(1) | 60 | 48 | 12 | 20.0 |
| (2(1)(1)) | 55 | 40 | 15 | 27.3 |
| ((2)(1)) | 65 | 55 | 10 | 15.4 |
| Metric | hops = 1 | hops = 2 | hops = 3 | hops = 4 | ||||||||||||||
| (1) | (1)(1) | (1)(1)(1) | (1)(1)(1)(1) | Avg | (2) | (2)(1) | ((1)(1)) | (2)(2) | ((1)(1))(1) | Avg | (3) | ((2)(1)) | (3)(1) | (2(1)(1)) | Avg | (4) | Avg | |
| Neural executor (UltraQuery) | ||||||||||||||||||
| MRR | 93.25 | 94.48 | 97.96 | 96.36 | 95.51 | 85.24 | 86.19 | 93.28 | 95.17 | 84.01 | 88.78 | 82.78 | 92.00 | 90.30 | 88.35 | 88.36 | 82.54 | 82.54 |
| Hit@1 | 91.33 | 93.17 | 96.94 | 95.93 | 94.34 | 82.67 | 82.48 | 92.00 | 93.05 | 80.00 | 86.04 | 77.50 | 87.69 | 88.00 | 85.00 | 84.55 | 76.22 | 76.22 |
| Hit@3 | 94.67 | 95.00 | 98.98 | 96.75 | 96.35 | 87.33 | 87.99 | 93.60 | 97.60 | 86.17 | 90.54 | 87.87 | 93.85 | 90.67 | 92.27 | 91.16 | 88.22 | 88.22 |
| Hit@10 | 96.00 | 97.78 | 98.98 | 96.75 | 97.38 | 88.67 | 92.13 | 96.80 | 97.84 | 93.33 | 93.75 | 90.65 | 100.00 | 96.00 | 93.64 | 95.07 | 92.22 | 92.22 |
| Metric | hops = 1 | hops = 2 | hops = 3 | hops = 4 | ||||||||||||||
| (1) | (1)(1) | (1)(1)(1) | (1)(1)(1)(1) | Avg | (2) | (2)(1) | ((1)(1)) | (2)(2) | ((1)(1))(1) | Avg | (3) | ((2)(1)) | (3)(1) | (2(1)(1)) | Avg | (4) | Avg | |
| Neural executor (UltraQuery) | ||||||||||||||||||
| MRR | 91.07 | 86.74 | 91.14 | 90.19 | 89.78 | 77.01 | 78.43 | 88.19 | 77.87 | 74.89 | 79.28 | 61.04 | 75.81 | 75.44 | 72.64 | 71.23 | 70.69 | 70.69 |
| Hit@1 | 90.00 | 83.65 | 88.61 | 87.26 | 87.38 | 74.00 | 74.69 | 85.60 | 69.54 | 69.17 | 74.60 | 54.44 | 66.15 | 66.67 | 64.97 | 63.06 | 65.22 | 65.22 |
| Hit@3 | 91.33 | 88.38 | 93.88 | 93.09 | 91.67 | 78.67 | 81.73 | 89.60 | 82.73 | 78.61 | 82.27 | 65.28 | 86.15 | 81.33 | 76.50 | 77.31 | 73.44 | 73.44 |
| Hit@10 | 93.33 | 92.38 | 93.88 | 94.31 | 93.47 | 85.33 | 84.88 | 94.40 | 93.17 | 83.00 | 88.16 | 72.22 | 89.23 | 90.00 | 86.32 | 84.44 | 82.44 | 82.44 |
| Executor | Hits | Recall | F1 |
| UltraQuery | 86.74 | 84.47 | 82.08 |
| Symbolic execution | 79.16 | 78.06 | 76.92 |
| Drop | -7.58 | -6.41 | -5.16 |
| Metric | hops = 1 | hops = 2 | hops = 3 | hops = 4 | ||||||||||||||
| (1) | (1)(1) | (1)(1)(1) | (1)(1)(1)(1) | Avg | (2) | (2)(1) | ((1)(1)) | (2)(2) | ((1)(1))(1) | Avg | (3) | ((2)(1)) | (3)(1) | (2(1)(1)) | Avg | (4) | Avg | |
| With ground-truth queries | ||||||||||||||||||
| MRR | 99.33 | 100.00 | 100.00 | 100.00 | 99.83 | 99.56 | 100.00 | 100.00 | 100.00 | 100.00 | 99.91 | 98.61 | 100.00 | 99.83 | 97.12 | 98.89 | 96.07 | 96.07 |
| Hit@1 | 98.67 | 100.00 | 100.00 | 100.00 | 99.67 | 99.33 | 100.00 | 100.00 | 100.00 | 100.00 | 99.87 | 97.78 | 100.00 | 99.67 | 94.55 | 98.00 | 94.00 | 94.00 |
| Hit@3 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 99.44 | 100.00 | 100.00 | 100.00 | 99.86 | 98.00 | 98.00 |
| Hit@10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 99.44 | 100.00 | 100.00 | 100.00 | 99.86 | 98.00 | 98.00 |
| LLM | Hits | Recall | F1 | API cost |
| UltRAG | ||||
| - GPT-5 | 86.74 | 84.47 | 82.08 | 0.017$ |
| - GPT-5 GPT-5-mini | 82.37 | 77.83 | 78.41 | 0.009$ |
| - GPT-5-mini GPT-5 | 80.71 | 78.19 | 75.72 | 0.011$ |
| - DeepSeek-reasoner | 75.77 | 72.40 | 72.05 | 0.005$ |
| - GPT-5-mini | 75.71 | 71.08 | 71.60 | 0.004$ |
| KGQAGen-10K | API cost | Non-API time (s) | Input tokens | Output tokens | Cache hit % |
| RoG | 0.008$ | 1.4K | 1.4K | 0 | |
| GNN-RAG | 0.008$ | 1.4K | 1.4K | 0 | |
| SubgraphRAG (200) | 0.009$ | 3.1K | 1.5K | 0 | |
| UltRAG -OTS | 0.019$ | 69K | 2.8K | 96.06 | |
| - GPT-5-mini | 0.004$ | 69K | 2.4K | 95.53 |
| PPR nodes | Candidate top- | F1 |
| 5K | 10 | 80.99 |
| 5K | 50 | 80.33 |
| 5K | 100 | 80.82 |
| 30K | 10 | 81.45 |
| 30K | 50 | 82.08 |
| 30K | 100 | 81.81 |
| Setting | ((1)(1)) | ((1)(1))(1) | ((2)(1)) | (1) | (1)(1) | (1)(1)(1) | (1)(1)(1)(1) | (2(1)(1)) | (2) | (2)(1) | (2)(2) | (3) | (3)(1) | (4) |
| Single-iter F1 | 93.87 | 77.06 | 88.72 | 96.22 | 90.59 | 89.52 | 84.25 | 89.24 | 93.11 | 84.39 | 94.48 | 85.09 | 88.79 | 83.73 |
| Multi-iter F1 | 93.96 | 82.44 | 92.82 | 96.93 | 91.86 | 91.63 | 85.79 | 85.57 | 94.67 | 87.42 | 95.91 | 87.59 | 92.72 | 91.13 |
| Setting | GTSQA | KGQAGen-10k |
| Top-1 | 84.9 | 80.9 |
| Top-5 | 35.3 | 34.0 |
| Top- 0.3 | 68.6 | 59.6 |
| Top- 0.5 | 73.4 | 57.0 |
| Top- 0.7 | 72.1 | 54.7 |
| Relative threshold | 81.1 | 80.8 |
| Setting | Hits | Recall | F1 |
| No randomization | 92.66 | 91.05 | 89.29 |
| Randomized relation IDs | 93.90 | 92.37 | 91.46 |
| Model | Hits | Recall | Precision | F1 |
| GPT-5 | 73.0 | 53.6 | 66.3 | 55.5 |
| RoG | 88.5 | 78.4 | 84.7 | 78.5 |
| GNN-RAG | 88.8 | 78.1 | 85.4 | 78.7 |
| SubgraphRAG | 87.2 | 76.1 | 82.7 | 76.1 |
| GCR | 90.0 | 71.0 | 85.4 | 73.0 |
| UltRAG -OTS | 88.0 | 77.7 | 82.3 | 76.7 |
| Model | Hits | Recall | Precision | F1 |
| GPT-5 | 60.5 | 53.2 | 56.3 | 52.7 |
| GCR | 69.3 | 62.1 | 66.3 | 62.2 |
| GNN-RAG | 75.9 | 70.3 | 71.7 | 69.2 |
| RoG | 75.8 | 70.9 | 70.5 | 68.8 |
| SubgraphRAG | 76.4 | 70.6 | 71.1 | 68.8 |
| UltRAG -OTS | 73.7 | 68.8 | 68.0 | 66.3 |
| Dataset | UltRAG -OTS | Best baseline | Speed-up |
| WikiKG2 | 0.14s | 1.94s | 14.1 |
| Wikidata | 1.17s | 5.02s | 4.3 |