LOGIC: An LLM Benchmark for Intent-Grounded Change Impact in Aerospace Electrical Systems
Authors: Muhammad Faraz Shoaib, Muhammad Qasim, Raisulhaq Mohammed Rizwan, Rahmatullah Safdar, Muzammil Adnan Shaik, Abdul Aleem Mohammed
Organizations: School of Computing Wichita State University 1845 Fairmount St. Wichita, KS 67260 · School of Computing Department of Biomedical Engineering Wichita State University 1845 Fairmount St. Wichita, KS 67260 · School of Computing Department of Industrial Engineering Wichita State University 1845 Fairmount St. Wichita, KS 67260
Aerospace electrical-design revisions can contain multiple genuine changes, although an engineering request may authorize only a subset. Propagating every detected difference can therefore produce overly broad impact reports. We present LOGIC, a controlled benchmark and evaluation framework in which locally deployable language models ground a request in a deterministic candidate-change inventory before selected changes are propagated through a typed electrical traceability graph. This separation permits candidate-selection errors to be distinguished from downstream propagation errors. LOGIC contains 168 scenarios, including 144 selection and 24 abstention cases. We evaluate three 7--8B models against intent-agnostic, lexical, and structured-evidence methods, with an oracle-root upper bound. On 96 explicitly anchored selection cases, gate-only structured evidence achieves candidate F1 of 1.0000, compared with 0.9677 for token-lexical matching. On 12 relational-paraphrase cases, token-lexical F1 is 0.1772 and gate-only F1 is 0.0000, compared with 0.5000--0.6400 for the large language models. Model grounding degrades as candidate inventories grow from 4 to 64 changes, while affected-element and typed-path accuracy remain comparatively stable when frozen selections are replayed over graphs of approximately 1K to 100K nodes. Strict evidence gating suppresses false positives but can remove correct semantic selections. An exploratory evidence-empty abstention policy raises strict abstention accuracy to 0.6667 for all three models and reduces unsafe-report rates to 0.1667, while decreasing answerable-case coverage by 16.0--27.1 percentage points. Four of six conflicting requests remain unsafe for each model. These findings support combining literal evidence and language-model reasoning with engineering review when intent cannot be established reliably.
Figures & tables
Fig. 1: LOGIC pipeline. Deterministic revision differencing produces a candidate-change inventory from two ECAD revisions. The grounding method selects request-relevant candidates or abstains. Selected candidates are resolved to roots and propagated through a typed electrical traceability graph; abstention bypasses propagation.
Condition
Selection rule or role
All-difference
Select every observed candidate.
Token lexical
Select maximum-overlap candidates when the minimum token threshold is met.
Ungated LLM
Use an individual model’s original validated action and selection.
Gate-only
Select candidates with request-matched string anchors that are unique within the inventory, without an LLM.
Uniform evidence gate
Retain only model-selected candidates with deterministic support.
Oracle-root
Propagate the intended candidates as a reference under perfect selection.
TABLE I: Intent-grounding and reference conditions in LOGIC. † denotes an exploratory post-hoc analysis.
Fig. 2: LOGIC benchmark and evaluation protocol. The 168 scenarios comprise 144 answerable selection requests and 24 abstention requests. Candidate count varies the model-visible grounding task; graph-size replay evaluates deterministic propagation from frozen selections. Three local LLMs produce 504 predictions before oracle-based offline scoring.
Method
Model(s)
Overall F1
Exact
∣C∣=4
∣C∣=16
∣C∣=64
(a) General controlled selection: 96 cases, 32 per inventory size
All-difference
—
0.0855
0.0000
–
–
–
Token lexical
—
0.9677
0.9583
1.0000
1.0000
0.9091
Gate-only
—
1.0000
1.0000
1.0000
1.0000
1.0000
Oracle-root
—
1.0000
1.0000
1.0000
1.0000
1.0000
Ungated LLM
Qwen
0.7791
0.7083
1.0000
0.8222
0.5063
TABLE II: Candidate-selection accuracy and sensitivity to candidate load. Overall F1 and exact-set accuracy are computed over the answerable cases in each group; the final three columns report F1 by inventory size. Model-independent methods are shown once and marked “—” in the Model(s) column. Dashes in numerical columns indicate omitted breakdowns. † denotes an exploratory ensemble analysis.
Fig. 3: Candidate-selection F1 across the four semantic-challenge families, each containing 12 cases. Ungated LLMs exceed token lexical and gate-only on relational paraphrases; gate-only performs better on functional paraphrases and semantic multi-change requests. Family-level exact-set comparisons do not remain significant after Holm correction.
Fig. 4: Deterministic evidence and exploratory gating. (a) Correctness of supported and unsupported ungated selections, pooled across three models. The general-and-safety group includes 120 cases; the semantic-challenge group includes 48. (b) Semantic F1 for ungated, uniformly gated, and consensus-adaptive selections. The dashed line marks three-model majority-vote F1. Printed values give the adaptive-minus-ungated F1 difference with its paired-bootstrap 95% confidence interval.
(a) Overall safety outcomes and positive-case coverage
System
Policy
Strict abst.
Unsafe
Rejected
Coverage
—
Token lexical
0.3750
0.6250
0.0000
1.0000
—
Gate-only
0.7500
0.2500
0.0000
0.8750
Qwen
Ungated
0.4167
0.4167
0.1667
0.8333
Evidence-empty †
0.6667
0.1667
0.1667
0.6736
Llama
Ungated
0.0000
0.8333
0.1667
0.9931
TABLE III: Safety and coverage outcomes. Panel (a) reports rates on 24 safety cases and report coverage on 144 answerable cases. Panel (b) reports correct-abstention / unsafe-report counts within each six-case family; rejected outputs account for any remainder. † denotes an exploratory policy.
Affected-element F1
Typed-path F1
Model
1K
10K
100K
1K
10K
100K
Qwen2.5-7B
0.505
0.507
0.504
0.474
0.473
0.478
Llama3.1-8B
0.724
0.726
0.699
0.652
0.639
0.639
Mistral-7B
0.547
0.539
0.515
0.466
0.463
0.466
Oracle-root
1.000
1.000
1.000
1.000
1.000
1.000
TABLE IV: Deterministic graph-replay accuracy from frozen ungated selections on the 48 semantic challenges. Root F1 is constant across sizes and is reported in the text.
Institute of Intelligent Software, Guangzhou, China · Institute of Software, Chinese Academy of Sciences, Beijing, China · University of Liverpool, Liverpool, UK