Narrow and Deep: An Ontology Tower as the Knowledge of an LLM Agent for an Industrial Equipment System
Authors: Younghwan Joo, Sung-il Kim
Organizations: Energy Efficiency Research Division, Korea Institute of Energy Research, 152 Gajeong-ro, Yuseong-gu, Daejeon, 34129, Republic of Korea · Energy Engineering, University of Science & Technology, 217 Gajeong-ro, Yuseong-gu, Daejeon, 34113, Republic of Korea
Large language model (LLM) agents are beginning to operate industrial energy equipment, and what they get right depends on what they are told about the plant. Established building ontologies name many kinds of points across many sites, whereas an industrial equipment system needs few entities with much knowledge about each. This study proposes the ontology tower, a narrow-and-deep ontology of a single equipment system whose knowledge deepens in two ways: through quantities derived from the measured points by physical relations, and through lessons from the operating journal incorporated as knowledge nodes. On a real low-humidity air-handling test plant operated daily through a programmable logic controller, agents received a text projected from its tower in a preregistered evaluation of nine tasks replayed from the plant's records, using four open-weight models from 9 to about 750 billion parameters. This knowledge raised the rate at which the agents avoided the most plausible misjudgment of each task by about 20 percentage points, and the overall task score of the 9-billion-parameter model as much as that of the largest. Operating lessons were used when incorporated into the tower or placed in the prompt as records, but seldom when left in the journal behind a search tool. In live runs through an invariant safety layer, the agents brought the controlled variable into its target band in 12 of 14 runs. An ontology narrow in entities but deep in what is known about them can thus supply the knowledge that an agent for an industrial equipment system needs.
Figures & tables
Figure 1: Ontology tower of the studied system: node counts at each computed level, two dependency chains that fix the levels, and the node and edge kinds of the interchange schema.
Figure 2: One diagnosis trial with the tower query tools: (A) the plant, (B) the tower with the nodes queried, and (C) the sequence of calls and the diagnosis.
Condition
What the agent receives in addition
Knowledge of Sections 2.2 to 2.4
K0, address level
PLC addresses and raw integer values only
Bindings without names
K1, tag level
Tag names, units, and quality flags, as on the plant’s display
Tags
K2, structured knowledge
The text projected from the tower (Section 2.5 )
Tag attributes, couplings, envelopes, and knowledge nodes with the incorporated lessons; no derived quantities
K3, procedural
Summaries of the six operating workflows
None (outside the tower)
K4, experiential
A search tool over the operating journal and investigation documents
Journal records (outside the tower)
K4x
K4 without the records of the task’s own incident
As K4, without the task’s own records
Table 1: Knowledge conditions of the evaluation and the knowledge of the tower that each carries.
Figure 3: Primary-trap avoidance rate in each condition, pooled over four models and nine tasks (K4x: four diagnosis tasks). Error bars: 95 % confidence intervals from resampling tasks.
Figure 4: Change in the element score from K1 to K2 for each model, ordered by parameter count. Error bars: 95 % confidence intervals from resampling tasks.
Model
Projected text (K2)
Shuffled text
Tower as JSON (K2-json)
Documents behind search (K2-sub)
Tower query tools (K2-tool)
Tool use (trials)
Qwen3.5-9B
0.737
0.666
0.673
0.538
0.407
23 / 27
gemma-4-31B-it
0.765
0.669
0.698
0.578
0.737
24 / 27
EXAONE-4.5-33B
0.712
0.556
0.594
0.493
0.486
2 / 27
GLM-5.2
0.891
0.848
0.900
0.771
0.911
27 / 27
Pooled
0.777
0.685
0.717
0.597
0.638
76 / 108
Table 2: Element score by model for the projected text and four other forms of the knowledge (tool use: trials with at least one tower call).
Figure 5: Tool calls and outcomes of five logged trials on the rotor duty-cycle diagnosis task, by model and form of knowledge.
Form of the lesson
Desiccant-unit oscillation
Low-airflow stop cascade
Self-stop limit cycle
External setpoint sweep
Incorporated into the tower
0.885
1.019
0.495
0.480
Record in the prompt
0.922
1.023
0.571
0.550
Record only in the searchable journal
0.642
0.856
0.291
0.396
Table 3: Element score of the four diagnosis tasks with the lesson of the task’s own incident in three forms (means of completed trials, up to 12 per cell).
Figure 6: Return-air dew point (a) and humidity setpoint with the agent’s commands (b) in two live runs of the dew-point task on one night, with the projected text and with the tower query tools.
Institute of Automation Technology, Helmut Schmidt University / University of the Federal Armed Forces Hamburg, Hamburg, Germany. · Siemens AG, Nuremberg, Germany; Institute for Technologies and Management of Digital Transformation, Bergische Universität Wuppertal, Germany. · Artiquare GmbH, Ingolstadt, Germany. +11