APOLO: Automatic Prompt Optimization for Ontology Learning
Authors: Huu Tan Mai, Roman Kochnev, Cuong Xuan Chu, Lukas Lange, Heiko Paulheim, Daria Stepanova
Organizations: Bosch Center for Artificial Intelligence, Robert-Bosch-Campus 1, 71272 Renningen, Germany · University of Mannheim, Schloss, 68131 Mannheim, Germany · University of Würzburg, Sanderring 2, 97070 Würzburg, Germany · Bosch Research North America, Sunnyvale, California, USA
Ontology Learning (OL) from text has advanced with the emergence of Large Language Models (LLMs), but it remains challenging due to the limited availability of annotated training data and the difficulty of adapting LLMs to perform OL effectively. We address this via APOLO - Automatic Prompt Optimization for Ontology Learning, by casting OL as an explicit prompt optimization problem over LLM modules. To obtain training data, we employ a multi-agent system that generates text-ontology pairs from existing expert-curated ontologies. We then propose two ontology learner architectures: a greedy and an autoregressive learner, and optimize both using GEPA, a greedy evolutionary prompt optimizer built on DSPy. Experiments on two ontologies - a biomedical (DOID) and a plant ontology (PO) show consistent improvements after optimization across nearly all model and mode combinations, with autoregressive learners achieving the largest gains. Our results demonstrate that prompt optimization is a viable and lightweight alternative to fine-tuning for OL, and that the autoregressive formulation better captures ontological structure than the greedy approach.
Figures & tables
Figure 1: Overview of APOLO’s prompt optimization framework. An ontology learner predicts an ontology from an input text chunk using prompt π . The prediction is evaluated against a ground-truth ontology chunk to produce a score and actionable feedback; a reflection LLM uses both to iteratively improve π while keeping the underlying LLM fixed.
Figure 2: Overview of Greedy and Autoregressive learners at inference time. The Greedy learner splits documents into manageable chunks, each of which is turned into a corresponding ontology chunk. Once all chunks are constructed, they are concatenated into the resulting ontology. The Autoregressive learner additionally takes prior ontology chunks as context when generating a new chunk (dashed blue lines ⇢ ).
Figure 3: Example of a text-ontology training pair used for prompt optimization. A text chunk generated from a subontology of the initial domain ontology is annotated with corresponding entities and axioms, each grounded in a supporting quote from the text.
Model
Method
Opt.
DOID
PO
AxP
AxR
EnP
EnR
AxP
AxR
EnP
EnR
Qwen3.5-27B
AutoRAGLearner Babaei Giglou et al. (2025)
N/A
.006
.044
.026
.071
.006
.048
.057
.141
SBU-NLP Rahnamoun and Shamsfard (2025)
N/A
.000
.000
.074
.143
.023
.063
.155
.324
Kommineni et al. Kommineni et al. (2024)
N/A
.064
.111
.552
.381
.000
.000
.189
.099
Bakker et al. (A) Bakker et al. (2024)
N/A
.583
.622
.725
.690
.027
.048
.322
.282
Bakker et al. (B) Bakker et al. (2024)
N/A
.667
.578
.675
.643
.018
.048
.180
.282
Table 1: Comparison of baseline OL methods and prompt-optimized learners (Ours) on DOID and PO. ✓/✗ denote after/before GEPA optimization; improvements over unoptimized in bold. Results are from a single run.
Ontology learning (OL) aims to automatically construct structured knowledge models from text, yet progress remains fragmented across methods, domains, and evaluation practices. Despite decades of research, OL lacks a shared infrastructure for systematic evaluation and ontology access. This absence has hindered progress and fragmented research, leaving the central challenges of OL largely unaddressed. We introduce OntoLearner, a modular, cross-domain, and first-of-its-kind framework that unifies ontology access, large language model (LLM)-driven learning pipelines, and standardized benchmarking. OntoLearner releases 180 machine-readable ontologies spanning 22 domains and provides pipeline-ready datasets with train/dev/test splits for three core OL tasks: term typing, taxonomy discovery, and non-taxonomic relation extraction. Using this infrastructure, we conduct a large-scale empirical study of OL, evaluating 22 retrieval models and 12 LLMs across domains and tasks. The results converge on a finding that reframes the central challenge of OL: failure modes scale with ontological complexity rather than model size or architectural sophistication. The primary bottleneck is not model capability, but a structural mismatch between how models encode knowledge and how ontologies organize it. These findings establish that effective OL is reachable through the cross-domain, multi-task benchmarking enabled by OntoLearner. OntoLearner is open-source (MIT license) at https://github.com/sciknoworg/OntoLearner/.
Hamed Babaei Giglou, Jennifer D'Souza, Andrei Aioanei +2
TIB – Leibniz Information Centre for Science and Technology, Hannover, Germany · 3IBM Research, New York, USA · 1TIB – Leibniz Information Centre for Science and Technology, Hannover, Germany +1
Ontology learning from text remains challenging despite significant progress in Large Language Models (LLMs), which can hallucinate domain terms, produce inconsistent formats, and favor hierarchical over associative relations. In the LLMs4OL 2026 Challenge, we address both the End-to-End Flagship Task (Task A) and Ontology Extension Reuse Task (Task B) using an offline retrieval-augmented few-shot prompting pipeline. Our system employs Qwen2.5-14B-Instruct with all-MiniLM-L6-v2 for demonstration retrieval, selecting the top-5 examples for Task A and top-2 for Task B. A left-truncated context-windowing strategy preserves task instructions within long prompts. For Task B, generated triples undergo deterministic vocabulary-constrained filtering, retaining triples when at least one endpoint belongs to the sample's closed term/type vocabulary and removing duplicates of the initial ontology. The approach achieves Semantic Graph Similarity of 0.8692, Term-Typing F1 of 0.9200, and Taxonomy Discovery F1 of 0.8540 on Task B, while Task A achieves 0.7416 Semantic Graph Similarity. However, no non-taxonomic relations are extracted, highlighting limitations of closed, taxonomy-oriented relation vocabularies.
The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently characterized. We present a controlled evaluation of 13 models spanning dense and Mixture-of-Experts variants from the Qwen3.5 and Qwen3.6 lineages, together with proprietary GPT release variants, using the OntoLearner retrieval-augmented generation pipeline. All models are evaluated with the same embedding model, retrieval configuration, prompt templates, decoding settings, datasets, and metrics on term typing, taxonomy discovery, and non-taxonomic relationship extraction across four biomedical and materials science and engineering ontologies. Within the dense Qwen3.5 lineage, increasing parameter count primarily improves precision rather than recall, with the largest gains occurring between 9B and 27B parameters. However, the effect of scale is neither monotonic nor uniform across tasks and domains. Dense 27B models outperform substantially larger sparse models on term typing, whereas larger Mixture-of-Experts models achieve the strongest open-weight results on taxonomy discovery. Non-taxonomic relationship extraction remains difficult across model scales, particularly for the Materials Data Science ontology. Performance differences across matched Qwen variants and proprietary GPT releases further indicate that architecture and model lineage can outweigh nominal parameter count. These findings show that model size alone is an insufficient selection criterion for OL and provide empirical guidance for reproducible LLM-assisted ontology engineering.
Hamed Babaei Giglou, Sören Auer, Jennifer D'Souza
TIB Leibniz Information Centre for Science and Technology, Hannover, Germany · L3S Research Center, Leibniz University of Hannover, Hannover, Germany