The advent of large language models is contributing to the emergence of novel approaches that promise to better tackle the challenge of generating structured queries, such as SPARQL queries, from natural language. However, these new approaches mostly focus on response accuracy while ignoring other evaluation criteria, such as runtime and cost to generate SPARQL queries. Consequently, they are often not production-ready or easy to deploy over real-world knowledge graphs with good accuracy. To mitigate these issues, in this paper, we describe and systematically evaluate SPARQL-LLM, an open-source and triplestore-agnostic approach, powered by lightweight metadata, that generates SPARQL queries from natural language text. First, we describe its architecture, which consists of dedicated components for metadata indexing, prompt building, and query generation and execution. Then, we evaluate it based on a state-of-the-art challenge with multilingual questions, and a collection of questions from three of the most prevalent knowledge graphs within the field of bioinformatics. Our results demonstrate a substantial improvement of up to 59% in F1 score over the second-best system participating in the challenge, adaptability to high-resource languages such as English, Spanish, and German, as well as ability to form complex bioinformatics queries. Furthermore, our results show that our system is up to 27x faster than the second-best system participating in the challenge, while costing a maximum of $0.01 per question, making it suitable for real-time, low-cost text-to-SPARQL applications. SPARQL-LLM is publicly released as an open-source project at https://github.com/sib-swiss/sparql-llm and is currently deployed over real-world decentralized knowledge graphs at https://www.expasy.org/chat.
Figures & tables
Figure 1. Architecture of SPARQL-LLM consisting of Indexing, Prompt Building, and SPARQL Generation and Execution components. Diagram showing the architecture of the SPARQL-LLM system.
Figure 2. Data flow of SPARQL-LLM when resolving a complex bioinformatics question. Diagram showing the data flow of the SPARQL-LLM system.
Figure 3. Kernel Density Estimate (KDE) of triple patterns of DBpedia-related (in blue) and Corporate-related (in orange) KGQA corpora. The dashed lines denote the evaluation corpora published by the TEXT2SPARQL 2025 challenge. We observe that the majority of the queries are considerably simple containing from one to three triple patterns. Visualization of Kernel Density Estimate (KDE) of triple patterns for DBpedia and Corporate KGQA corpora, with dashed lines for TEXT2SPARQL challenge corpora.
System
LLM Provider
Version
Context
Temperature
Max Tokens
Seed
Knowledge Cutoff
TEXT2SPARQL 2025
SPARQL−LLMlg
OpenAI
GPT-4o
128K
0
16K
42
10/2023
SPARQL−LLMsm
OpenAI
GPT-4.1-mini
1.05M
0
16K
42
06/2024
SPARQL−LLMos
OpenAI
GPT-oss-120b
131K
0
16K
42
06/2024
ARUQULA ( Brei et al., 2025 )
OpenAI
GPT-4.1-mini
1.05M
1
700
-
06/2024
mKGQAgent ( Perevalov and Andreas, 2025 )
OpenAI
GPT-4o
128K
0
-
-
10/2023
Table 1. LLM configuration of three sets of systems: i) TEXT2SPARQL 2025: ARUQULA is the winner of DBpedia (EN) and Corporate subtasks, while mKGQAgent is the winner of the DBpedia (ES) subtask. To guarantee a fair comparison, all the SPARQL−LLM variants use the same LLM Provider and configuration, and have knowledge cutoff before the release of the challenge; ii) System Analysis (§ 4.1.5 ): different flagship LLMs are evaluated, having the same configuration and knowledge cutoff; iii) TEXT2SPARQL 2026: SPARQL−LLM is evaluated against four other systems, two of which were submitted with multiple variants to the challenge.
Figure 4. Performance evaluation of three variants of our system against the TEXT2SPARQL 2025 challenge winners, showcasing a substantial increase of 24% in the F1 Score for the DBpedia (EN) and DBpedia (ES) subtasks. Bar chart showing F1 scores of systems competing in the challenge.
Figure 5. Cost analysis of the three versions of our system in terms of runtime and input/output tokens per question. Bar chart showing the cost analysis of the three system versions by runtime and input/output tokens per question.
Figure 6. Hyperparameter fine-tuning and ablation study of SPARQL-LLM. Results of hyperparameter tuning and ablation study for SPARQL-LLM are shown in the figure.
Error Description
Reference Query (Example)
Generated Query (Example)
Count
Predicate/Class Namespace Mismatch
dbo:origin
dbp:origin
20
Erroneous Reference Query
(see below)
(see below)
11
Redundant Aggregate Variable Projection
SELECT ?country WHERE
SELECT ?country (COUNT(?city) AS ?cityCount)
4
Wrong Text Encoding
dbr:Republic_of_Montenegro_(1992–2006)
dbr:Republic_of_Montenegro_(1992%E2%80%932006)
3
Alternative Predicate Usage
dbp:placeOfDeath
dbp:deathPlace
2
Empty Generated Query
SELECT […]
-
2
Table 2. Error analysis on the DBpedia (EN) questions: Errors categories (top), and examples of erroneous reference queries (bottom).
Figure 7. Kernel Density Estimate (KDE) of triple patterns of the three of the most prevalent bioinformatics corpora. We observe that all the BioKGQA queries contain up to 10× more triple patterns than the classic KGQA queries (Figure 3 ). Visualization of Kernel Density Estimate (KDE) of triple patterns for bioinformatics corpora.
Figure 8. Results with BioKGQA evaluation queries. Visualization of results of BioKGQA benchmark queries.
DBpedia
Corporate
System
English
Spanish
English
German
SPARQL−LLM
.7192
.7810
.8510
.8104
GRASP ( Walter and Bast, 2025 )
.3296
.3288
.5262
.4892
IRIS ( Latipov et al., 2025 )
.2897
.2914
.2896
.1460
StarCoderv1 ( Soru et al., 2025 )
.3235
.3730
.3215
.3510
StarCoderv2 ( Soru et al., 2025 )
.3129
.2361
.1724
.2390
Table 3. Results of the TEXT2SPARQL 2026 challenge, reported as F1 scores. Across both subtasks and all three languages, SPARQL−LLM showcases a substantial improvement of up to 59% in F1 score over the second-best system ( ARUQULAv1 ).
Large language models (LLMs) have demonstrated strong capabilities in structured query generation, making them a natural choice for Text-to-SPARQL, which translates natural language questions into executable SPARQL queries over knowledge graphs. However, their initial outputs remain unreliable: generated queries may be executable yet semantically misaligned with input questions, leading to incorrect retrieval. To address this issue, we propose Generator-Gate-Corrector (GGC), a framework for reliable LLM-based Text-to-SPARQL generation. GGC first uses a Generator to produce an initial query, then applies a Gate to predict whether correction is needed, and finally invokes a Corrector only for selected high-risk queries. This selective correction mechanism avoids unnecessary modifications and reduces the risk of degrading originally correct queries. Experiments on MCQA show that GGC improves query-level accuracy from 90.23% to 98.33% while reducing inference overhead by 45% compared with correcting all generated queries. Ablation studies show that the Gate is robust across thresholds and that Corrector training data composition affects correction effectiveness and stability. Overall, the results demonstrate that selective correction enhances the accuracy, reliability, and efficiency of LLM-based text-to-SPARQL generation.
Ziyi Yang, Thanh-Son Nguyen, Tuan Anh Nguyen +1
1Nanyang Technological University, Centre for Info. Sciences and Systems · Institute of High Performance Computing, Agency for Science, Technology and Research (A*STAR), Singapore
Knowledge graph question answering seeks to translate natural language questions into executable queries over knowledge graphs, but existing approaches often rely on large models or full supervision in the form of gold query annotations. This study examines whether reinforcement learning with outcome-based rewards can train a small instruction-tuned language model to perform zero-shot Text-to-SPARQL generation in the scholarly domain. Group-Relative Policy Optimization (GRPO) is applied to the Qwen3-1.7B model on DBLP-QuAD, using prompts that combine natural language questions with symbolic hints about entities and relations. Training relies on execution feedback, structural constraints, and answer-level rewards, with an additional variant that incorporates gold-query-based shaping. The resulting models are compared to the unmodified zero-shot baseline and to a supervised DoRA-finetuned baseline across answer-level accuracy, execution accuracy, category-wise scores, and generalization to held-out templates. GRPO substantially improves over the zero-shot baseline and exhibits competitive generalization, while supervised DoRA finetuning achieves higher overall accuracy on the same model scale. Ablation analyses indicate that execution-based rewards account for most gains, with additional shaping yielding limited additional benefit, suggesting that outcome-based reinforcement learning is a viable training strategy when gold queries are unavailable for token-level supervision.
We present GRISP (Guided Recurrent IRI Selection over SPARQL Skeletons), a novel SPARQL-based question-answering method over knowledge graphs based on fine-tuning a small language model (SLM). Given a natural-language question, the method first uses the SLM to generate a natural-language SPARQL query skeleton, and then to re-rank and select knowledge graph items to iteratively replace the natural-language placeholders using knowledge graph constraints. The SLM is jointly trained on skeleton generation and list-wise re-ranking data generated from standard question-query pairs. We evaluate the method on common Wikidata and Freebase benchmarks, and achieve better results than other state-of-the-art methods in a comparable setting.
Sebastian Walter, Hannah Bast
University of Freiburg · Department of Computer Science · Freiburg im Breisgau, Germany