Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing
Authors: Remo Grillo, Lukas Klic, Giovanni Colavizza
Organizations: University of Bologna, Via Zamboni 33, 40126 Bologna, Italy · I Tatti — The Harvard University Center for Italian Renaissance Studies, Via di Vincigliata 26, 50135 Florence, Italy · Department of Communication, University of Copenhagen, Karen Blixens Plads 8, 2300 Copenhagen S, Denmark
Natural-language access to RDF knowledge graphs is a core Semantic Web ambition. Large language models (LLMs) have advanced Text-to-SPARQL, yet on unfamiliar graphs they often generate valid queries that misrepresent the populated data model. QRAKEN is a training-free, ontology-agnostic neurosymbolic pipeline grounding generation in empirical graph evidence rather than schema expectations. An offline distiller produces TTQL, a compact description of populated multi-hop patterns, conditional frequencies and path-conditioned literal examples, plus a class-property co-occurrence matrix. Online, TTQL guides the LLM, while deterministic syntax, vocabulary and data-model checks provide diagnostics for iterative refinement. On CK25 (First International Text2SPARQL Challenge), under matched-condition recomputation on a QLever snapshot, QRAKEN achieves strict F1 of 0.643 ± 0.026 with GPT-4.1 mini and 0.652 ± 0.012 with GPT-5.4: relative gains of 30% and 32% over the strongest recomputed participant, outperforming systems using the same base model family. Ablations identify TTQL patterns as the dominant driver (+0.31 strict F1 over a shape-only baseline); the refinement loop provides a cheap safety net, rejecting triple patterns unsupported by the co-occurrence matrix. Compared with auto-derived SHACL, TTQL yields 64% higher strict F1, supporting the value of empirical patterns beyond schema exposure. With two local 35B 4-bit open-weight models at zero marginal cost, the same pipeline matches the strongest recomputed participant, and TTQL advantages over shape-only and SHACL baselines persist. Results on a single, relatively small benchmark provide an initial empirical signal; monolithic TTQL injection on very open cross-domain graphs remains the main limitation.
Figures & tables
Figure 1: QRAKEN architecture. The offline phase (top) distills the populated graph into a TTQL file (injected into the LLM prompt) and a co-occurrence matrix (kept machine-side for deterministic checks). The online phase (bottom) couples LLM-based SPARQL generation with a three-stage correctness-guided self-healing loop (curved arrow: fail → diagnostic); only queries that pass all three stages are executed. The TTQL artifact is produced once per graph and reused across all questions.
Figure 2: CK25 strict vs lenient F1, matched-condition recomputation on a shared QLever endpoint — the four strongest recomputed participant baselines (ARUQULA, IIS-Q, IIS-L, mKGQAgent), the two hosted QRAKEN configurations ( gpt-4.1-mini , gpt-5.4 ) and the two open-weight QRAKEN configurations run locally at zero marginal cost ( Qwen3.6 , Ornith-1.5 ), drawn in a separate colour because they are not comparable in scale to the hosted models (Sect. 5.5 ). Lower-ranked teams (MIPT/AIRI, LACODAM, LABIC, FRANZ, AIFB) are omitted for readability and remain in the supplemental material. Error bars on QRAKEN rows are 1 σ across replicates ( n=6 for gpt-4.1-mini , n=3 otherwise).
Cell
Prompt context
Heal.
gpt-4.1-mini
gpt-5.4
Ornith-1.5
Qwen3.6
( n=6 )
( n=3 )
( n=3 )
( n=3 )
A1
Full TTQL patterns
on
0.643 ± 0.026
0.652 ± 0.012
0.411 ± 0.026
0.491 ± 0.036
A2
Full TTQL patterns
off
0.617 ± 0.013
0.654 ± 0.010
0.420 ± 0.008
0.424 ± 0.028
A3
Co-occ. matrix only
on
0.359 ± 0.016
0.490 ± 0.032
0.327 ± 0.020
0.351 ± 0.016
A4
Co-occ. matrix only
off
0.308 ± 0.015
0.501 ± 0.012
0.284 ± 0.017
0.309 ± 0.005
A8
SHACL-auto (shexer)
off
0.376 ± 0.018
—
0.255 ± 0.012
0.203 ± 0.011
Table 1: Ablation on CK25 (strict F1, mean ± sd). Self-healing is on with a budget of 5 iterations and off at 1. The first two model columns are the hosted models of the headline runs; the last two are open-weight 35B MoE models quantised to 4 bits and run locally (Sect. 5.5 ). Row A8 reports an auto-derived SHACL baseline as a direct comparator for shape-informed prompting Wardenga and Käfer (2025) ; it was not run on gpt-5.4 . The three summary rows give the marginal contribution of the TTQL axis (A2 − A4), of the self-healing loop (A1 − A2), and of TTQL over the shape baseline (A2 − A8).
Cell
Model
Calls/q
Tokens/q
$/100q
Lat./q
A1 (full, healing on)
gpt-4.1-mini
1.43
41,905
$1.69
3.5 s
A2 (full, healing off)
gpt-4.1-mini
1.00
29,164
$1.18
5.0 s
A1 (full, healing on)
gpt-5.4
1.21
34,933
$7.08
3.2 s
A2 (full, healing off)
gpt-5.4
1.00
29,182
$5.89
2.5 s
A3 (shape only, healing on)
gpt-4.1-mini
1.40
6,757
$0.29
2.6 s
A4 (shape only, healing off)
gpt-4.1-mini
1.00
4,711
$0.20
2.6 s
Table 2: Online cost per question on CK25, aggregated over the same runs that back Table 1 . Calls/q is the mean number of LLM calls per question (self-healing iterations); tokens/q is the mean total of input+output tokens; $/100q multiplies total per-question cost by 100; lat./q is median wall-clock per question. The open-weight row is the local grid of Sect. 5.5 , where per-call accounting was not instrumented and latency is warm-cache. For ARUQULA, calls/q is the mean number of agent steps and lat./q the mean time to answer on the corporate dataset, both as reported in its own evaluation Brei et al. (2025) .
Knowledge graph question answering (KGQA) requires navigating from topic entities to an answer several relations away. Recent methods prompt a frontier LLM to explore the graph through a retrieval tool, but their reliance on frontier-scale inference makes them costly to deploy. We present Search-on-Graph-R1 (\sogrone{}), which internalizes this navigation into a compact 8B model through supervised fine-tuning (SFT) followed by reinforcement learning (RL). Our central idea is to scaffold a frontier teacher with each question's gold SPARQL query, so the teacher traverses a known answer-bearing path with a live \texttt{Search} tool rather than having to discover the path itself. Since every call executes against a live Freebase server, the resulting trajectories are grounded in the knowledge graph by construction. On WebQSP, CWQ, and GrailQA, \sogrone{} at 8B surpasses every frozen frontier-LLM system in our comparison and posts the strongest results on CWQ of any system we compare against. It does so using no auxiliary module at inference and no LLM judge during training. Isolating each training stage shows that SFT and RL contribute complementary gains, our approach transfers across model families, and RL learns to reach answers in fewer \texttt{Search} calls than its SFT initialization.
Jia Ao Sun, Hao Yu, Fengran Mo +4
Université de Montréal · Mila – Québec AI Institute · McGill University +2
Large language models (LLMs) frequently generate confident yet factually incorrect content when used for language generation (a phenomenon often known as hallucination). Retrieval augmented generation (RAG) tries to reduce factual errors by identifying information in a knowledge corpus and putting it in the context window of the model. While this approach is well-established for document-structured data, it is non-trivial to adapt it for Knowledge Graphs (KGs), especially for queries that require multi-node/multi-hop reasoning on graphs. We introduce UltRAG, a training-free KG-RAG recipe that combines LLM query generation, a fully inductive neural query executor, and LLM arbitration. This off-the-shelf composition achieves state-of-the-art results on Knowledge Graph Question Answering (KGQA) tasks without retraining the LLM or executor, while enabling language models to interface with Wikidata-scale graphs (116M entities, 1.6B relations) at comparable or lower costs. Our ablation studies indicate that these gains come from the full system design rather than from any single component.
Dobrik Georgiev, Kheeran K. Naidu, Alberto Cattaneo +3
Knowledge graph question answering seeks to translate natural language questions into executable queries over knowledge graphs, but existing approaches often rely on large models or full supervision in the form of gold query annotations. This study examines whether reinforcement learning with outcome-based rewards can train a small instruction-tuned language model to perform zero-shot Text-to-SPARQL generation in the scholarly domain. Group-Relative Policy Optimization (GRPO) is applied to the Qwen3-1.7B model on DBLP-QuAD, using prompts that combine natural language questions with symbolic hints about entities and relations. Training relies on execution feedback, structural constraints, and answer-level rewards, with an additional variant that incorporates gold-query-based shaping. The resulting models are compared to the unmodified zero-shot baseline and to a supervised DoRA-finetuned baseline across answer-level accuracy, execution accuracy, category-wise scores, and generalization to held-out templates. GRPO substantially improves over the zero-shot baseline and exhibits competitive generalization, while supervised DoRA finetuning achieves higher overall accuracy on the same model scale. Ablation analyses indicate that execution-based rewards account for most gains, with additional shaping yielding limited additional benefit, suggesting that outcome-based reinforcement learning is a viable training strategy when gold queries are unavailable for token-level supervision.