Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing
Organizations: University of Bologna, Via Zamboni 33, 40126 Bologna, Italy · I Tatti — The Harvard University Center for Italian Renaissance Studies, Via di Vincigliata 26, 50135 Florence, Italy · Department of Communication, University of Copenhagen, Karen Blixens Plads 8, 2300 Copenhagen S, Denmark
Abstract
Natural-language access to RDF knowledge graphs is a core Semantic Web ambition. Large language models (LLMs) have advanced Text-to-SPARQL, yet on unfamiliar graphs they often generate valid queries that misrepresent the populated data model. QRAKEN is a training-free, ontology-agnostic neurosymbolic pipeline grounding generation in empirical graph evidence rather than schema expectations. An offline distiller produces TTQL, a compact description of populated multi-hop patterns, conditional frequencies and path-conditioned literal examples, plus a class-property co-occurrence matrix. Online, TTQL guides the LLM, while deterministic syntax, vocabulary and data-model checks provide diagnostics for iterative refinement. On CK25 (First International Text2SPARQL Challenge), under matched-condition recomputation on a QLever snapshot, QRAKEN achieves strict F1 of 0.643 0.026 with GPT-4.1 mini and 0.652 0.012 with GPT-5.4: relative gains of 30% and 32% over the strongest recomputed participant, outperforming systems using the same base model family. Ablations identify TTQL patterns as the dominant driver (+0.31 strict F1 over a shape-only baseline); the refinement loop provides a cheap safety net, rejecting triple patterns unsupported by the co-occurrence matrix. Compared with auto-derived SHACL, TTQL yields 64% higher strict F1, supporting the value of empirical patterns beyond schema exposure. With two local 35B 4-bit open-weight models at zero marginal cost, the same pipeline matches the strongest recomputed participant, and TTQL advantages over shape-only and SHACL baselines persist. Results on a single, relatively small benchmark provide an initial empirical signal; monolithic TTQL injection on very open cross-domain graphs remains the main limitation.
Figures & tables
| Cell | Prompt context | Heal. | gpt-4.1-mini | gpt-5.4 | Ornith-1.5 | Qwen3.6 |
| ( ) | ( ) | ( ) | ( ) | |||
| A1 | Full TTQL patterns | on | 0.643 0.026 | 0.652 0.012 | 0.411 0.026 | 0.491 0.036 |
| A2 | Full TTQL patterns | off | 0.617 0.013 | 0.654 0.010 | 0.420 0.008 | 0.424 0.028 |
| A3 | Co-occ. matrix only | on | 0.359 0.016 | 0.490 0.032 | 0.327 0.020 | 0.351 0.016 |
| A4 | Co-occ. matrix only | off | 0.308 0.015 | 0.501 0.012 | 0.284 0.017 | 0.309 0.005 |
| A8 | SHACL-auto (shexer) | off | 0.376 0.018 | — | 0.255 0.012 | 0.203 0.011 |
| Cell | Model | Calls/q | Tokens/q | $/100q | Lat./q |
| A1 (full, healing on) | gpt-4.1-mini | 1.43 | 41,905 | $1.69 | 3.5 s |
| A2 (full, healing off) | gpt-4.1-mini | 1.00 | 29,164 | $1.18 | 5.0 s |
| A1 (full, healing on) | gpt-5.4 | 1.21 | 34,933 | $7.08 | 3.2 s |
| A2 (full, healing off) | gpt-5.4 | 1.00 | 29,182 | $5.89 | 2.5 s |
| A3 (shape only, healing on) | gpt-4.1-mini | 1.40 | 6,757 | $0.29 | 2.6 s |
| A4 (shape only, healing off) | gpt-4.1-mini | 1.00 | 4,711 | $0.20 | 2.6 s |