Where Can a Decision Model Diagnose HVAC Faults? Reasoning Demand, Physical Representation, and Robustness Under Shift
Organizations: University of Arizona, Tucson, AZ, USA
Abstract
Artificial intelligence supports building operations in several forms, each with its own barrier. Expert rules must be tuned for every system, supervised models need labeled data that buildings rarely record, and language models return free text that requires human-in-the-loop checking, since their stated confidence is unreliable. A newer kind of pretrained model, here called a decision model, returns a probability for every allowed answer, so one model could serve many decisions without training. This study answers three open questions for fault diagnosis in heating, ventilation, and air-conditioning systems: which decisions such a model can make, what input it needs, and whether its probabilities hold when conditions change. On 128 fault days from four public datasets of real equipment, faults are graded by the reasoning their diagnosis demands, with data given raw, as physical features, or with Brick topology. The decision model Jev, open language models, and a supervised model face nine tests that change season, control configuration, or building. Given physical features, Jev and the larger open model diagnosed faults whose evidence one feature carries, but not faults that need operating context. Under shift they kept their accuracy and calibration, while the supervised model lost 0.33 macro-F1 yet led or tied within a building. Their probabilities still needed correction, and detection was weak. The study maps which faults a decision model can diagnose and from what input, and supports a division of work in which code computes the physics and the model ranks candidate faults for an operator.
Figures & tables
| Code or term | Meaning | Section |
|---|---|---|
| L1, L2, L3, L4 | Reasoning levels of a fault type: direct, consistency, context, temporal | 3.2 |
| U | Unobservable: the fault leaves no physical trace in the window | 3.2 |
| R0 | Raw readings, a table of 5-minute values | 3.4 |
| R1 | 53 physical features computed in code | 3.4 |
| R1b | R1 with each value replaced by a named level | 3.4 |
| R2 | R1 plus the Brick topology and the latest value of every point | 3.4 |
| Dataset | Equipment, setting, sampling | Fault days (types) | Fault-free days | Used for |
|---|---|---|---|---|
| ORNL RTU | 44 kW (12.5-ton) rooftop unit, research building, 1 min | 48 (3) | 14 | Map, seasonal shift |
| RP-1312 | Paired VAV AHUs, research building, 1 min | 49 (14) | 49 | Map, seasonal shift, building-shift source |
| FLEXLAB | One AHU in CAV and VAV modes, research building, 1 min | 21 (6) | 5 | Map, configuration shift |
| Nesbitt Hall | AHU-1 and AHU-2, occupied building, 5 min | 10 (3) | 22 | Map, building-shift target |
| Level | ORNL RTU | RP-1312 | FLEXLAB | Nesbitt | Total |
| L1 | – | 5 (2) | – | – | 5 |
| L2 | 16 (1) | 21 (5) | 21 (6) | 8 (2) | 66 |
| L3 | 32 (2) | 18 (6) | – | 2 (1) | 52 |
| L4 | – | 5 (1) | – | – | 5 |
| Total | 48 | 49 | 21 | 10 | 128 |
| Question | Answer type | Allowed answers | Used for |
|---|---|---|---|
| “Which condition best explains this window?” | One choice among named options | “normal” and each fault type of the dataset (4–15 options), each shown with its description ( ) | Diagnosis and its calibration |
| “Is the unit operating with a fault in this window?” | Yes or no | yes: “At least one component, sensor, or control setting is not behaving as designed.” no: “Operation is consistent with the design intent for the conditions.” | Detection |
| “How severe is the energy or comfort impact of the condition in this window?” | A level on an ordered scale | none, mild, moderate, severe | Severity (secondary) |
| ORNL unit | RP-1312 | FLEXLAB | Nesbitt Hall | |||||
| Detector | Fault | Fault-free | Fault | Fault-free | Fault | Fault-free | Fault | Fault-free |
| Jev, yes-or-no answer | ||||||||
| R0 | 48 | 33 | 33 | 17 | 50 | 73 | 10 | 59 |
| R1 | 95 | 84 | 75 | 57 | 99 | 89 | 88 | 94 |
| R2 | 83 | 48 | 77 | 55 | 99 | 98 | 90 | 100 |
| Jev, condition answer | ||||||||
| Question | Answer | Evidence |
|---|---|---|
| RQ1: Which faults can training-free models diagnose? | Faults whose evidence one computed feature carries, such as a stuck damper or an unstable control loop. Not faults that need operating context. On the one unit where both could be compared, accuracy fell from consistency faults (L2) to context faults (L3). | Section 4.1 , Fig. 2 |
| RQ2: Does computing the physics in code help? | Yes. Raw readings gave almost no diagnoses. Physical features raised pooled accuracy on consistency faults to 37% for Jev and 46% for Qwen2.5 72B. Named levels lowered it. | Section 4.1 , Fig. 2 |
| RQ2: Do topology and point values help further? | Slightly for Jev, by 0.04 in shifted macro-F1, and not for Qwen2.5 72B. The gain came mainly from points the features lacked. | Sections 4.1 and 4.3 |
| RQ3: Do training-free models lose less accuracy under shift? | Yes. On average they lost at most 0.01 (95% upper bound 0.10), against 0.33 for LightGBM. LightGBM still led or tied within a building. The training-free models led only across buildings, mainly because LightGBM fell below a constant “normal” answer. | Section 4.3 , Fig. 3 , Table |
| RQ3: Can their probabilities be taken at face value? | No. Calibration error was 0.27 or more. It did not drift under shift, and a correction fitted on source data mostly carried over. Detection was no better than flagging every window. | Section 4.4 , Fig. 4 , Table |