Organizations: Chair of Data and Knowledge Engineering, University of Passau, Germany · Distributed and Self-organizing Systems, Chemnitz University of Technology, Germany
Constructing typed, justified semantic links between ontologies is essential for enabling interoperability across heterogeneous and interdisciplinary knowledge domains. However, manually curating such links is difficult to scale. To address this challenge, we propose an end-to-end framework for ontology network construction that automates the discovery and generation of both intra-domain and inter-domain relationships. Our approach combines domain-adapted DistilBERT embeddings for dense contextual representation, clustering-based pre-filtering to reduce the candidate search space, and GPT-4o-driven relationship generation via iterative prompt engineering to produce semantically rich, interpretable links. Applied to ReproduceMeON - a network of 33 ontologies spanning machine learning, microscopy, computational science, and experimental workflow - the pipeline reduces approximately 800k raw concept pairs to 95k high-quality candidates. Human expert validation of 429 generated relationships by two independent annotators yields an overall precision of 80.19% (91.49% on high-certainty annotations) and an F1 of 0.890, with substantial inter-annotator agreement. Comparative experiments against five similarity-based baselines, including Sentence-BERT, show a substantial performance gap (best baseline F1 = 0.581), while an ablation study demonstrates that similarity-based methods alone fail to discriminate valid from invalid relationships (AUC approx 0.5) on the filtered candidate set. These findings highlight the necessity of LLM-based reasoning over concept roles and domain semantics for accurate relationship construction.
Figures & tables
Figure 1: Overview of the proposed end-to-end pipeline for discovering and generating intra- and inter-domain links for the development of ReproduceMeON ontology network.
Table 2
Figure 3: Relation types across different domains
Method
Prec.
Rec.
F1
AUC
Jaccard (concept names)
0.370
1.000
0.540
0.506
Jaccard (with context)
0.370
1.000
0.540
0.574
TF-IDF cosine
0.576
0.552
0.564
0.625
DistilBERT (pre-trained)
0.582
0.515
0.546
0.609
Sentence-BERT (MiniLM-L6)
0.570
0.593
0.581
0.673
LLM pipeline (ours)
0.802
1.000
0.890
0.927
Table 3: Baseline comparison on the balanced evaluation set (929 pairs total). Threshold for similarity methods is set to the F1-optimal value. AUC: Area Under ROC Curve.
Figure 4: ROC curves (left) and Precision-Recall curves (right) for all baselines and the LLM pipeline on the 929-pair evaluation set.
Method
AUC-ROC (on 429 pairs)
Jaccard (concept names)
0.506
Jaccard (with context)
0.485
TF-IDF cosine
0.503
DistilBERT (pre-trained)
0.511
Sentence-BERT (MiniLM-L6)
0.500
Table 4: AUC-ROC for predicting relationship validity from similarity score alone on the 429 pre-filtered pipeline pairs. AUC ≈ 0.5 confirms similarity cannot distinguish valid from invalid among the high-similarity candidates selected by the pipeline.
Figure 5: Ablation: similarity score distributions (valid vs invalid, top row), AUC-ROC on pre-filtered pairs (bottom left), and ROC on the full evaluation set (bottom middle). All similarity methods achieve AUC ≈ 0.5 on pre-filtered pairs, confirming that the LLM step is the critical source of precision.
Domain
n
Cohen’s κ
Agreement (%)
Machine Learning
137
0.671
87.6%
Computational
94
0.817
94.7%
Experimental Workflow
43
0.720
90.7%
Overall
274
0.723
90.5%
Table 5: Inter-annotator agreement on 274 dually annotated pairs. Cohen’s κ and percentage agreement per domain and overall. Microscopy (155 pairs) was annotated by a single expert.
Domain
Valid
Total
Precision
Machine Learning
103
137
75.18%
Microscopy
126
155
81.29%
Computational
80
94
85.11%
Experimental Workflow
35
43
81.40%
Overall
344
429
80.19%
Link Type
Table 6: Precision of generated relationships by domain and link type.
Figure 6: Precision broken down by domain (left) and relationship type (right). Causal, Comparative, and Hierarchical relationships achieve the highest precision; Instrumental relationships are the most error-prone.
Knowledge Organization Systems like Ontologies and taxonomies are fundamental for structuring scientific knowledge, yet their manual curation presents a persistent bottleneck in knowledge management. While Large Language Models (LLMs) offer a scalable mechanism for automated ontology generation, their capacity to classify complex, domain-specific semantics requires systematic evaluation. In this paper, we assess the performance of five small, open-source LLMs (up to 9 billion parameters) in identifying semantic relationships between biomedical concepts. To support this evaluation, we introduce MeSH-Rel-4K, a dataset comprising 4K semantic relationships extracted from the Medical Subject Headings (MeSH). We analyse three adaptation strategies: standard prompting, Chain-of-Thought prompting, and fine-tuning. While parameter-constrained models traditionally struggle with the nuances of in-context logic, our results reveal that targeted fine-tuning increases the average F1-score by 34.1 percentage points. These results confirm that direct fine-tuning effectively exceeds the reasoning bottlenecks of smaller LLMs, providing an accurate, automated methodology for the construction and evolution of specialised biomedical ontologies.
Tanay Aggarwal, Angelo Salatino, Francesco Osborne +1
Knowledge Media Institute, The Open University, Milton Keynes, UK · Department of Business and Law, University of Milano-Bicocca, Milan, IT
Ontology learning from text remains challenging despite significant progress in Large Language Models (LLMs), which can hallucinate domain terms, produce inconsistent formats, and favor hierarchical over associative relations. In the LLMs4OL 2026 Challenge, we address both the End-to-End Flagship Task (Task A) and Ontology Extension Reuse Task (Task B) using an offline retrieval-augmented few-shot prompting pipeline. Our system employs Qwen2.5-14B-Instruct with all-MiniLM-L6-v2 for demonstration retrieval, selecting the top-5 examples for Task A and top-2 for Task B. A left-truncated context-windowing strategy preserves task instructions within long prompts. For Task B, generated triples undergo deterministic vocabulary-constrained filtering, retaining triples when at least one endpoint belongs to the sample's closed term/type vocabulary and removing duplicates of the initial ontology. The approach achieves Semantic Graph Similarity of 0.8692, Term-Typing F1 of 0.9200, and Taxonomy Discovery F1 of 0.8540 on Task B, while Task A achieves 0.7416 Semantic Graph Similarity. However, no non-taxonomic relations are extracted, highlighting limitations of closed, taxonomy-oriented relation vocabularies.
Ontology matching (OM) enables semantic interoperability between different ontologies and resolves their conceptual heterogeneity by aligning related entities. OM systems currently have two prevailing design paradigms: conventional knowledge-based expert systems and newer machine learning-based predictive systems. While large language models (LLMs) and LLM agents have revolutionised data engineering and have been applied creatively in many domains, their potential for OM remains underexplored. This study introduces a novel agent-powered LLM-based design paradigm for OM systems. With consideration of several specific challenges in leveraging LLM agents for OM, we propose a generic framework, namely Agent-OM (Agent for Ontology Matching), consisting of two Siamese agents for retrieval and matching, with a set of OM tools. Our framework is implemented in a proof-of-concept system. Evaluations of three Ontology Alignment Evaluation Initiative (OAEI) tracks over state-of-the-art OM systems show that our system can achieve results very close to the long-standing best performance on simple OM tasks and can significantly improve the performance on complex and few-shot OM tasks.
Zhangcheng Qiang, Weiqing Wang, Kerry Taylor
Australian National University Canberra, ACT, Australia · Monash University Melbourne, VIC, Australia