Organizations: Tsinghua University, China · Beijing University of Chinese Medicine, China · Guangdong Provincial Laboratory of Traditional Chinese Medicine, China
Multi-source knowledge graphs (KGs) need query mechanisms that expose reliability and exploit domain structure. This paper presents TCMaster, a property-graph query substrate for confidence-aware traversal and workload-guided physical design over Traditional Chinese Medicine KGs. TCMaster integrates pharmacopoeias, prescriptions, molecular databases, and LLM-extracted micro-semantics into a KG with approximately 221K entities and 723K base edges. It annotates edges with provenance-level confidence, rewrites Cypher queries with confidence predicates, ranks multi-hop paths under PRODUCT, MIN, or weighted-average policies, and uses ontology skew through direction selection, herb-attribute bitmaps, and materialized shortcut edges. On Neo4j, direction selection improves attribute lookup by a factor of 1.47, shortcuts accelerate high-fanout target counting by a factor of 4.42, confidence filtering removes 39.3 percent of low-quality heterogeneous paths, and KG retrieval improves TCMbench QA accuracy by 20.0 percentage points.
Figures & tables
System
Scope
Edge conf.
Path query
Opt.
Artifact
TCMSP
M
No
No
No
Web
TCMID
M/Disease
No
No
No
Web
SymMap
Syn
No
No
No
Web
HERB 2.0
M
No
No
No
Web
ETCM
Rx
No
No
No
Web
OpenTCM
RAG/Diag.
No
No
No
Paper
TABLE I: Comparison with representative TCM resources. M, Rx, Syn, and Micro denote molecular, prescription, syndrome/symptom, and micro-semantic knowledge.
Layer
Content
Primary Edges
L1: Molecular
Herb–Ingredient–Target
265K
L2: TCM Attributes
Nature, Flavor, Meridian, Toxicity
45K
L3: Prescription
Prescription composition
71K
L4: Micro-semantics
Processing, botany, efficacy, etc.
236K
L5: Clinical
Prescription efficacy, indication
20K
TABLE II: Five-layer ontology of TCMaster-KG. Edge counts summarize primary relation groups and are rounded.
Source Level
Conf.
Example Relations
AUTHORITATIVE
0.95
HAS_INGREDIENT, TARGETS
AUTHORITATIVE_PHARMA
0.90
HAS_PROPERTY, HAS_MERIDIAN
CURATED
0.85
HAS_COMPONENT, HAS_RX_EFFICACY
LLM_EXTRACTED
0.70
HAS_BOTANY, PROCESSED_BY
PREDICTED (KGE)
0.30–0.60
Link prediction outputs
TABLE III: Source-level confidence assignment
Fig. 1: TCMaster system architecture from data ingestion and cleaning to confidence-annotated Neo4j storage, workload-guided query processing, and downstream application modes.
Query
Pattern
Operator stress
Avg. out
Q1
Herb-Attr
1-hop lookup
3.03
Q2
Rx-Herb-Attr
2-hop join
5.27
Q3
Rx-Herb-Ingr.-Target
shortcut path
91.10
Q4
Rx-Herb-Ingr.-Disease
long traversal
91.02
Q5
Rx-Herb-Meridian
aggregation
6.28
Q6
Rx-Herb-Rx
pattern match
22.00
TABLE IV: Characterization of the query workload. Rx and Ingr. denote prescription and ingredient.
Query
Hops
Mean (ms)
P95 (ms)
Q1: Herb attribute lookup
1
8.41
9.64
Q2: Prescription property
2
5.23
6.49
Q3: Prescription → Target
3
6.70
8.89
Q4: Disease association
4
7.12
9.77
Q5: Aggregation
2
5.27
6.54
Q6: Pattern matching
var.
12.04
17.52
TABLE V: Query latency for representative clinical workloads
Workload
Base
Opt.
Speedup
Avg. out
Direction selection
9.40
6.41
1.47 ×
100
Shortcut top- k lookup
13.62
11.64
1.17 ×
100
Bitmap single attribute
15.99
15.53
1.03 ×
200
Combined top- k lookup
10.73
16.45
0.65 ×
100
High-fanout target count
51.19
11.57
4.42 ×
6555
High-fanout target enum.
336.08
283.43
1.19 ×
6555
TABLE VI: Workload-guided physical-design results
Scale
Nodes
Edges
Q1
Q3
Q5
25%
26,575
55,263
7.98
6.50
11.73
50%
41,755
113,000
11.67
6.43
10.84
75%
53,698
165,745
8.08
2.83
5.59
100%
64,254
223,696
6.32
1.99
5.80
TABLE VII: Query latency (ms) vs. core-layer graph scale
Fig. 2: Core-layer scalability over 25–100% sampled L1–L3 graph scales, showing bounded latency variation for Q1, Q3, and Q5.
Fig. 3: Data quality and cleaning effectiveness by validation layer and before/after cleaning metric.
Fig. 4: KGE structural-validation ablation across S1–S4, with MRR baselines and RotatE Hits@1/3/10.
Dimension
Baseline
Vector-RAG
KG-RAG
Clinical Prescription
17.9%
14.1%
45.3%
Prescription Logic
99.7%
99.4%
100.0%
Safety Audit
40.2%
44.1%
72.2%
Overall
52.5%
52.5%
72.5%
TABLE VIII: TCMbench accuracy by mode and dimension
Fig. 5: Downstream TCMbench validation: accuracy by mode and KG-RAG gains by dimension.
Query
θ=0.10
θ=0.18
θ=0.30
θ=0.35
θ=0.45
Q1
0.0
0.0
0.0
0.0
0.0
Q2
0.0
0.0
0.0
0.0
0.0
Q3
0.0
39.3
69.5
100.0
100.0
Q5
0.0
0.0
0.0
0.0
0.0
Q7
0.0
0.0
0.0
0.0
0.0
TABLE IX: Filtering ratio (%) under confidence thresholds
Fig. 6: Effect of confidence threshold θ on filtering ratio and recall/precision trade-off.
Knowledge Graph-based Question Answering (KGQA) plays a pivotal role in complex reasoning tasks but remains constrained by two persistent challenges: the structural heterogeneity of Knowledge Graphs(KGs) often leads to semantic mismatch during retrieval, while existing reasoning path retrieval methods lack a global structural perspective. To address these issues, we propose Structure-Tracing Evidence Mining (STEM), a novel framework that reframes multi-hop reasoning as a schema-guided graph search task. First, we design a Semantic-to-Structural Projection pipeline that leverages KG structural priors to decompose queries into atomic relational assertions and construct an adaptive query schema graph. Subsequently, we execute globally-aware node anchoring and subgraph retrieval to obtain the final evidence reasoning graph from KG. To more effectively integrate global structural information during the graph construction process, we design a Triple-Dependent GNN (Triple-GNN) to generate a Global Guidance Subgraph (Guidance Graph) that guides the construction. STEM significantly improves both the accuracy and evidence completeness of multi-hop reasoning graph retrieval, and achieves State-of-the-Art performance on multiple multi-hop benchmarks.
Peng Yu, En Xu, Bin Chen +2
AI Product Center, Kingsoft Corporation, Beijing, China
Biomedical knowledge graphs (KGs) treat disease associations as static facts, but temporal information is crucial for clinical reasoning, e.g., a symptom diagnostic of one disease at age 3 may imply a different disease at age 13. Existing KGs such as PrimeKG, Hetionet, and iKraph do not encode when a finding becomes clinically relevant over the course of a disease. This limits their usefulness for longitudinal clinical reasoning and retrieval augmentation. We introduce ChronoMedKG, a temporal biomedical knowledge graph that contains 460,497 evidence-linked triples (filtered from 13M raw extractions) covering 13,431 diseases. Each association is tied to temporal components like onset window or progression stage, which are backed by PMID-traceable evidence and a multi-signal credibility score. The graph is constructed through a disease-autonomous multi-agent pipeline in which multiple frontier LLMs independently extract knowledge from PubMed and PMC literature. Only those relations are kept that are supported by multi-model consensus, survive credibility filtering, as well as ontology alignment. ChronoMedKG scored 92.7% agreement against Orphadata and adds temporal grounding for 6,250 diseases absent from HPOA, Orphadata, and Phenopackets, including 1,657 Orphanet-coded rare diseases. We further introduce ChronoTQA, a benchmark of 3,341 questions across eight task types (six temporal plus two static controls), with a 12-question supplementary probe. Frontier LLMs lose roughly 30 points moving from static to temporal questions; ChronoMedKG retrieval rescues 47-65% of their long-tail failures, against 17-29% for HPOA-RAG. As such, ChronoMedKG provides a crucial temporal axis for retrieval-augmented clinical systems that was previously absent.
Large language models can answer knowledge-intensive questions more reliably when they are grounded with knowledge graphs, but systems such as Think-on-Graph and Reasoning-on-Graph repeatedly query the same graph neighborhoods across different questions. In this work, we study this repeated retrieval in Knowledge Graph Question Answering~(KGQA) workloads and propose KGCache, an in-memory cache for one-hop knowledge graph neighborhoods. KGCache is designed to be compatible with both iterative traversal (ToG) and one shot planning (RoG) KGQA paradigms. KGCache is placed between the KGQA engine and the backend serving the KG, so repeated entity requests can be served from cache instead of issuing new KG queries. We evaluate KGCache on WebQSP and CWQ using LRU, LFU, and a trace-aware Oracle policy. Our analysis shows that both datasets contain substantial entity reuse among starting entities and entities reached during traversal. We also explore semantic caching for similar queries, which shows additional hit-rate gains on WebQSP and needs further accuracy testing on CWQ. Entity caching accelerates KG retrieval by up to 1.91×, while semantic-context caching achieves up to 1.06× full-system speedup in the evaluated WebQSP configurations, with each hit being up to 3.73× faster.