When LLMs Sit Above Diagnostic Tools: Unrealized Complementarity in Industrial Fault Diagnosis
Organizations: Independent Researcher
Abstract
Large language models are increasingly used as integration layers above specialized tools, but a stronger component does not necessarily produce a stronger combined system. Across five diagnostic datasets (bearing vibration, process monitoring, semiconductor equipment), we study whether an LLM can reliably use external diagnostic information; paired repeat calls separate advice effects from output instability. In all five, conflicting external information overturned initially correct LLM judgments. Among the four datasets with direct integration comparisons, none showed a consistent advantage for implicit LLM integration over the stronger standalone source. On a Tennessee Eastman confirmation set whose protocol was fixed before evaluation, unaided accuracy was 64.67%, implicit LLM-specialist integration 77.43%, and the specialist alone 83.33%. Specialist information improved the LLM by 12.8 points (95% interval 9.7 to 15.9), yet the integrated output stayed 5.9 points below the specialist (95% interval -12.0 to -0.7). A two-source selector oracle reached 92.76%, indicating complementarity that the integrated output did not fully realize. The integrated output missed 140 of 295 specialist corrections (47.5%) but lost 15 of 99 initially correct LLM judgments (15.2%). The deficit remained under prompt and specialist sensitivity analyses. Among CWRU cases solved under both evidence presentations, task-aligned physical evidence yielded lower estimates of susceptibility to incorrect advice in six of seven models (five intervals excluding zero); higher reasoning effort gave no reliable reduction in five models, and a separate four-model TEP analysis gave no clear evidence that it resolves the integration problem. Source quality and integration quality should be evaluated separately: an integration layer should be compared with its stronger standalone component, not only with the unaided LLM.
Figures & tables
| Dataset | Domain | Task in this study | Evidence shown to the model | External diagnosis | Cluster | Role |
|---|---|---|---|---|---|---|
| CWRU | Bearing vibration | 4 classes, 72 cases, 28 files | Generic spectral summary, or a task-aligned presentation with additional fault-specific information | Controlled advice, and a fixed classical specialist in the integration comparisons | Recording file | Characterization and intervention |
| Paderborn | Bearing vibration | 3 classes, 87 segments, 29 bearings | Generic and task-aligned text presentations | Controlled advice and fixed fitted specialists | Bearing | Same-domain support |
| XJTU-SY | Bearing run-to-failure | 3 classes, 88 recordings, 11 bearings | Generic and task-aligned text presentations | Controlled advice and fixed fitted specialists | Bearing | Same-domain support |
| Tennessee Eastman | Simulated chemical plant | 6 classes, 150 confirmation runs, 25 per class | Numerical process summary plus a static domain reference (variable roles, subsystems, candidate fault definitions), identical across runs and conditions | Fixed logistic-regression specialist; alternative random-forest specialist in one secondary analysis | Simulation run | Held-out TEP evaluation; secondary sensitivity analyses |
| Lam 9600 | Public metal-etch data from a commercial tool | 5 classes, 20 cases | Tool and trace text | Controlled advice in fixed wording | Wafer | Limited real-equipment check (§4.7; Appendix J) |
| WM-811K | Wafer maps | Image-versus-text pilot, 50 cases | Native map, or a text rendering of it | None | Wafer map | Interface boundary (Appendix J) |
| Generic excerpt | Task-aligned excerpt |
|---|---|
| - motor load: 0 HP | - BPFO is the outer-race ball-pass frequency. Outer-race defects typically raise envelope energy at BPFO and its harmonics. |
| - rotational speed: 1796 RPM | - BPFI is the inner-race ball-pass frequency. Inner-race defects typically raise envelope energy at BPFI and its harmonics. |
| E1. Spectral band-energy shares: 0-500Hz=0.2%, 500-1500Hz=0.6%, 1500-3000Hz=25.3%, 3000-6000Hz=73.9%. Spectral entropy=0.647 (+11.1% vs normal). | - theoretical BPFO=107.3 Hz, BPFI=162.1 Hz, BSF=70.5 Hz, FTF=11.9 Hz |
| E2. Skewness=0.083; impulse factor=8.66 (+111.5% vs normal). | E1. Envelope spectrum near BPFO (theoretical 107.3 Hz): peak SNR=23.52× local background (nearest peak 108.4 Hz, error 1.1 Hz); 2×BPFO harmonic SNR=19.31×. BPFO is kinematically associated with outer-race defects. |
| Quantity | Definition | Why it is reported |
|---|---|---|
| Unaided accuracy | Share of cases correct before any external diagnosis | The model’s own reading of the evidence |
| Test–retest instability | How often the paired repeat changes the unaided label | Separates ordinary instability from the effect of advice |
| Retest-adjusted harmful switching | Among initially correct rows, the advised-call error rate minus the repeat-call error rate | The susceptibility the intervention experiments target |
| Integrated output versus the better source | Accuracy of the advised judgment minus the better of unaided accuracy and specialist accuracy | Whether integration outperforms the stronger component |
| Two-source selector oracle | Correct whenever either of the two stored standalone labels is correct | Descriptive bound for selection between the two stored labels |
| Missed correction | The specialist is correct, the unaided LLM is wrong, and the final integrated output remains wrong | Failure to use a correct specialist correction |
| Error type | Model | True class | Unaided | Paired repeat | Specialist | Advised |
|---|---|---|---|---|---|---|
| Missed correction | Qwen 3.8 | IDV1 | IDV6 | IDV6 | IDV1 | IDV6 |
| Harmful acceptance | Claude Sonnet 5 | NORMAL | NORMAL | NORMAL | IDV15 | IDV15 |
| Condition | Specialist accuracy | Integrated accuracy | Integrated − better standalone source (95% CI) | Missed corrections | Harmful acceptances |
|---|---|---|---|---|---|
| Primary: logistic, implicit | 83.3% | 77.4% | −5.9 [−12.0, −0.7] | 140/295 (47.5%) | 15/99 (15.2%) |
| Arbitration condition: logistic, explicit instruction | 83.3% | 76.1% | −7.2 [−14.0, −1.3] | 161/295 (54.6%) | 6/99 (6.1%) |
| Prior-label condition: logistic, stored unaided diagnosis | 83.3% | 72.2% | −11.1 [−18.3, −3.6] | 215/295 (72.9%) | 0/99 (0.0%) |
| Alternative-specialist condition: random forest, implicit | 84.0% | 78.2% | −5.7 [−11.2, −1.1] | 116/272 (42.6%) | 10/72 (13.9%) |
| Candidate | Dev. accuracy | Errors | Mean selector-oracle gap | Pooled LLM-only-correct rows | Decision |
|---|---|---|---|---|---|
| Random forest | 0.820 | 27 | 0.074 | 80 | Passed |
| Multinomial logistic regression | 0.833 | 25 | 0.096 | 101 | Selected |
| Model | RSR [95% Wilson interval] | RSR n | RAIR [95% Wilson interval] | RAIR n |
|---|---|---|---|---|
| GPT-5.6 Sol | 1.000 [0.796, 1.000] | 15 | 0.360 [0.202, 0.555] | 25 |
| Claude Sonnet 5 | 0.250 [0.089, 0.532] | 12 | 0.966 [0.883, 0.990] | 58 |
| Claude Opus 5.5 | 1.000 [0.796, 1.000] | 15 | 0.000 [0.000, 0.204] | 15 |
| Gemini 3.8 Flash | 0.667 [0.391, 0.862] | 12 | 0.630 [0.486, 0.755] | 46 |
| Qwen 3.8 | 1.000 [0.796, 1.000] | 15 | 0.191 [0.104, 0.325] | 47 |
| GLM-5.3 | 1.000 [0.796, 1.000] | 15 | 0.517 [0.392, 0.641] | 58 |
| Bearing dataset | Generic presentation | Task-aligned additions or substitutions |
|---|---|---|
| CWRU | Load, speed, sensor, and window metadata; broad spectral-band shares; spectral entropy; skewness; impulse factor; RMS and normal-reference change; coarse energy shares around precomputed BPFO, BPFI, and BSF bands, without the frequencies or fault mapping; kurtosis; dominant component | Generic diagnostic reference; geometry-derived BPFO, BPFI, BSF, and FTF frequencies; envelope peak and harmonic signal-to-background ratios and offsets; matched-normal RMS and impulsiveness comparisons |
| Paderborn | Speed; RMS, kurtosis, crest factor, peak-to-peak, and skewness; ordinary-spectrum dominant component and relative power; spectral entropy; band-energy fractions | Documented bearing geometry and characteristic-frequency equations; BPFO, BPFI, BSF, and FTF values; envelope-analysis band; local peak frequency, offset, and signal-to-background ratio; deterministic BPFO–BPFI contrasts |
| XJTU-SY | Speed; RMS, kurtosis, crest factor, peak-to-peak, and skewness; ordinary-spectrum dominant component and relative power; spectral entropy; band-energy fractions | Documented bearing geometry; BPFO, BPFI, BSF, and FTF values; envelope-analysis band; local peak frequency, offset, and signal-to-background ratio; deterministic BPFO–BPFI contrast |
| Condition | Deterministic rule | Depends on | Implementation |
|---|---|---|---|
| CWRU | Within each true class, the other three classes are assigned as evenly as possible after a seeded shuffle (seed 20260923). | True class and the stored map | Seeded case-preparation routine |
| Paderborn | Healthy Outer-race damage; Inner-race damage Healthy; Outer-race damage Inner-race damage. | True class only | Dataset-preparation routine |
| XJTU-SY | Among the other classes in the fixed candidate order, the label is chosen by case index. | True class, candidate order, and case index | Dataset-preparation routine |
| TEP evidence-strength | NORMAL IDV1 IDV4 IDV6 IDV13 IDV15 NORMAL. | True class only | Development label map, applied by the evidence-strength runner |
| Lam 9600 Task B | Greedy balance over the other four classes, with ties broken by SHA-256 of wafer id, class, and seed 20260926. | True class, wafer id, and running counts | Canonical-build routine |
| Prompt family | Requested temperature | Seed | Reasoning effort / trace | Max tokens |
|---|---|---|---|---|
| TEP integration, secondary sensitivity analyses, evidence-strength, development | 0 | 20260928 | low / trace excluded | 512 |
| TEP sensitivity analysis with higher reasoning effort | 0 | 20260928 | high / trace excluded | 4096 |
| Final CWRU; Paderborn; XJTU | 0 | 20260924 | low or high / trace excluded | 2500 |
| Lam Tasks A and B | 0 | 20260926 | low / trace excluded | 64 |
| WM text and image pilot | 0 | 20260925 | low / trace excluded | 256 |
| Model | Common-correct n | Files | Generic | Task-aligned | Difference (95% CI) |
|---|---|---|---|---|---|
| GPT-5.6 Sol | 22 | 12 | 0.500 | 0.182 | 0.318 [0.067, 0.647] |
| Claude Sonnet 5 | 22 | 11 | 0.091 | 0.000 | 0.091 [0.000, 0.227] |
| Claude Opus 5.5 | 36 | 15 | 0.139 | 0.000 | 0.139 [0.045, 0.256] |
| Gemini 3.8 Flash | 23 | 12 | 0.348 | 0.087 | 0.261 [0.028, 0.647] |
| Qwen 3.8 | 27 | 12 | 0.667 | 0.185 | 0.481 [0.214, 0.731] |
| GLM-5.3 | 14 | 9 | 0.214 | 0.214 | 0.000 [−0.312, 0.308] |
| Model | Parsed , low / high | Unaided accuracy, low to high | Accuracy difference (95% CI) | Harmful switching, low to high | Susceptibility difference (95% CI) |
|---|---|---|---|---|---|
| GPT-5.6 Sol | 72 / 72 | 0.347 to 0.333 | −0.014 [−0.062, 0.029] | 0.167 to 0.153 | 0.014 [−0.060, 0.104] |
| Claude Opus 5.5 | 72 / 72 | 0.583 to 0.639 | 0.056 [−0.055, 0.171] | 0.069 to 0.056 | 0.014 [−0.085, 0.111] |
| Gemini 3.8 Flash | 72 / 70 | 0.417 to 0.457 | 0.040 [−0.116, 0.199] | 0.139 to 0.083 | 0.056 [−0.056, 0.159] |
| Qwen 3.8 | 72 / 72 | 0.444 to 0.458 | 0.014 [−0.029, 0.060] | 0.319 to 0.375 | −0.056 [−0.143, 0.026] |
| GLM-5.3 | 72 / 72 | 0.375 to 0.417 | 0.042 [−0.101, 0.192] | 0.097 to 0.097 | 0.000 [−0.087, 0.086] |
| Model | n | Unaided low | Unaided high | Integrated low | Integrated high | Specialist | Median reasoning tokens (high) |
|---|---|---|---|---|---|---|---|
| Claude Opus 5.5 | 150 | 0.833 | 0.827 | 0.833 | 0.827 | 0.833 | 0 |
| Gemini 3.8 Flash | 149 | 0.611 | 0.738 | 0.779 | 0.772 | 0.832 | 1,092 |
| GLM-5.3 | 150 | 0.547 | 0.627 | 0.720 | 0.773 | 0.833 | 90.5 |
| Kimi K3 | 150 | 0.627 | 0.633 | 0.747 | 0.793 | 0.833 | 478.5 |
| Dataset | Feature set | Estimator | Out-of-fold unit | Out-of-fold accuracy | Specialist errors |
|---|---|---|---|---|---|
| CWRU | strong | 400-tree random forest | recording file | 0.958 | 3 |
| CWRU | weak | 400-tree random forest | recording file | 0.917 | 6 |
| Paderborn | strong | 400-tree random forest | bearing | 0.724 | 24 |
| Paderborn | weak | 400-tree random forest | bearing | 0.437 | 49 |
| XJTU-SY | strong | 400-tree random forest | bearing | 0.909 | 8 |
| XJTU-SY | weak | 400-tree random forest | bearing | 0.761 | 21 |
| Dataset and condition | Model-level differences and interval classes |
|---|---|
| Paderborn, task-aligned, strong specialist | Sol 0.000 (includes); Sonnet −0.093 (includes); Opus −0.069 (includes); Gemini −0.011 (includes); Qwen 0.069 (includes); GLM 0.046 (includes); Kimi −0.034 (includes) |
| Paderborn, task-aligned, weak specialist | Sol 0.023 (includes); Sonnet −0.084 (includes); Opus −0.023 (includes); Gemini 0.011 (includes); Qwen 0.034 (includes); GLM 0.000 (includes); Kimi −0.023 (includes) |
| XJTU, generic, weak specialist | Sonnet 0.000 (includes); Opus −0.250 (below); Qwen −0.011 (includes); Kimi −0.057 (includes) |
| XJTU, task-aligned, weak specialist | Sonnet −0.193 (below); Opus −0.136 (below); Gemini −0.023 (includes); Qwen 0.159 (above) |
| Model | Implicit integration | Arbitration condition | Arbitration − implicit | Implicit errors fixed | New errors |
|---|---|---|---|---|---|
| GPT-5.6 Sol | 0.820 | 0.813 | −0.007 | 2 | 3 |
| Claude Sonnet 5 | 0.840 | 0.793 | −0.047 | 7 | 14 |
| Claude Opus 5.5 | 0.833 | 0.833 | 0.000 | 0 | 0 |
| Gemini 3.8 Flash | 0.780 | 0.773 | −0.007 | 7 | 8 |
| Qwen 3.8 | 0.680 | 0.673 | −0.007 | 0 | 1 |
| GLM-5.3 | 0.720 | 0.707 | −0.013 | 10 | 12 |
| Model | Parsed n | Integrated | RF specialist | Selector oracle |
|---|---|---|---|---|
| GPT-5.6 Sol | 150 | 0.827 | 0.840 | 0.913 |
| Claude Sonnet 5 | 150 | 0.833 | 0.840 | 0.900 |
| Claude Opus 5.5 | 150 | 0.833 | 0.840 | 0.913 |
| Gemini 3.8 Flash | 150 | 0.773 | 0.840 | 0.900 |
| Qwen 3.8 | 150 | 0.673 | 0.840 | 0.913 |
| GLM-5.3 | 147 | 0.755 | 0.837 | 0.905 |
| Model | Arbitration | Prior-label | Prior-label − arbitration | Arbitration errors fixed | New errors |
|---|---|---|---|---|---|
| GPT-5.6 Sol | 0.813 | 0.827 | +0.013 | 3 | 1 |
| Claude Sonnet 5 | 0.793 | 0.653 | −0.140 | 4 | 25 |
| Claude Opus 5.5 | 0.833 | 0.833 | 0.000 | 0 | 0 |
| Gemini 3.8 Flash | 0.773 | 0.673 | −0.100 | 1 | 16 |
| Qwen 3.8 | 0.673 | 0.667 | −0.007 | 0 | 1 |
| GLM-5.3 | 0.707 | 0.673 | −0.033 | 10 | 15 |
| Study group | API calls | Cost (USD) |
|---|---|---|
| Bearing studies | 9,149 final-evaluation attempts | approximately 271 |
| Tennessee Eastman studies | 19,225 | approximately 152 |
| Semiconductor pilot | not itemized | approximately 3 |
| Total | N/A | approximately 426 |