LOGIC: An LLM Benchmark for Intent-Grounded Change Impact in Aerospace Electrical Systems
Authors: Muhammad Faraz Shoaib, Muhammad Qasim, Raisulhaq Mohammed Rizwan, Rahmatullah Safdar, Muzammil Adnan Shaik, Abdul Aleem Mohammed
Organizations: School of Computing Wichita State University 1845 Fairmount St. Wichita, KS 67260 · School of Computing Department of Biomedical Engineering Wichita State University 1845 Fairmount St. Wichita, KS 67260 · School of Computing Department of Industrial Engineering Wichita State University 1845 Fairmount St. Wichita, KS 67260
Aerospace electrical-design revisions can contain multiple genuine changes, although an engineering request may authorize only a subset. Propagating every detected difference can therefore produce overly broad impact reports. We present LOGIC, a controlled benchmark and evaluation framework in which locally deployable language models ground a request in a deterministic candidate-change inventory before selected changes are propagated through a typed electrical traceability graph. This separation permits candidate-selection errors to be distinguished from downstream propagation errors. LOGIC contains 168 scenarios, including 144 selection and 24 abstention cases. We evaluate three 7--8B models against intent-agnostic, lexical, and structured-evidence methods, with an oracle-root upper bound. On 96 explicitly anchored selection cases, gate-only structured evidence achieves candidate F1 of 1.0000, compared with 0.9677 for token-lexical matching. On 12 relational-paraphrase cases, token-lexical F1 is 0.1772 and gate-only F1 is 0.0000, compared with 0.5000--0.6400 for the large language models. Model grounding degrades as candidate inventories grow from 4 to 64 changes, while affected-element and typed-path accuracy remain comparatively stable when frozen selections are replayed over graphs of approximately 1K to 100K nodes. Strict evidence gating suppresses false positives but can remove correct semantic selections. An exploratory evidence-empty abstention policy raises strict abstention accuracy to 0.6667 for all three models and reduces unsafe-report rates to 0.1667, while decreasing answerable-case coverage by 16.0--27.1 percentage points. Four of six conflicting requests remain unsafe for each model. These findings support combining literal evidence and language-model reasoning with engineering review when intent cannot be established reliably.
Figures & tables
Fig. 1: LOGIC pipeline. Deterministic revision differencing produces a candidate-change inventory from two ECAD revisions. The grounding method selects request-relevant candidates or abstains. Selected candidates are resolved to roots and propagated through a typed electrical traceability graph; abstention bypasses propagation.
Condition
Selection rule or role
All-difference
Select every observed candidate.
Token lexical
Select maximum-overlap candidates when the minimum token threshold is met.
Ungated LLM
Use an individual model’s original validated action and selection.
Gate-only
Select candidates with request-matched string anchors that are unique within the inventory, without an LLM.
Uniform evidence gate
Retain only model-selected candidates with deterministic support.
Oracle-root
Propagate the intended candidates as a reference under perfect selection.
TABLE I: Intent-grounding and reference conditions in LOGIC. † denotes an exploratory post-hoc analysis.
Fig. 2: LOGIC benchmark and evaluation protocol. The 168 scenarios comprise 144 answerable selection requests and 24 abstention requests. Candidate count varies the model-visible grounding task; graph-size replay evaluates deterministic propagation from frozen selections. Three local LLMs produce 504 predictions before oracle-based offline scoring.
Method
Model(s)
Overall F1
Exact
∣C∣=4
∣C∣=16
∣C∣=64
(a) General controlled selection: 96 cases, 32 per inventory size
All-difference
—
0.0855
0.0000
–
–
–
Token lexical
—
0.9677
0.9583
1.0000
1.0000
0.9091
Gate-only
—
1.0000
1.0000
1.0000
1.0000
1.0000
Oracle-root
—
1.0000
1.0000
1.0000
1.0000
1.0000
Ungated LLM
Qwen
0.7791
0.7083
1.0000
0.8222
0.5063
TABLE II: Candidate-selection accuracy and sensitivity to candidate load. Overall F1 and exact-set accuracy are computed over the answerable cases in each group; the final three columns report F1 by inventory size. Model-independent methods are shown once and marked “—” in the Model(s) column. Dashes in numerical columns indicate omitted breakdowns. † denotes an exploratory ensemble analysis.
Fig. 3: Candidate-selection F1 across the four semantic-challenge families, each containing 12 cases. Ungated LLMs exceed token lexical and gate-only on relational paraphrases; gate-only performs better on functional paraphrases and semantic multi-change requests. Family-level exact-set comparisons do not remain significant after Holm correction.
Fig. 4: Deterministic evidence and exploratory gating. (a) Correctness of supported and unsupported ungated selections, pooled across three models. The general-and-safety group includes 120 cases; the semantic-challenge group includes 48. (b) Semantic F1 for ungated, uniformly gated, and consensus-adaptive selections. The dashed line marks three-model majority-vote F1. Printed values give the adaptive-minus-ungated F1 difference with its paired-bootstrap 95% confidence interval.
(a) Overall safety outcomes and positive-case coverage
System
Policy
Strict abst.
Unsafe
Rejected
Coverage
—
Token lexical
0.3750
0.6250
0.0000
1.0000
—
Gate-only
0.7500
0.2500
0.0000
0.8750
Qwen
Ungated
0.4167
0.4167
0.1667
0.8333
Evidence-empty †
0.6667
0.1667
0.1667
0.6736
Llama
Ungated
0.0000
0.8333
0.1667
0.9931
TABLE III: Safety and coverage outcomes. Panel (a) reports rates on 24 safety cases and report coverage on 144 answerable cases. Panel (b) reports correct-abstention / unsafe-report counts within each six-case family; rejected outputs account for any remainder. † denotes an exploratory policy.
Affected-element F1
Typed-path F1
Model
1K
10K
100K
1K
10K
100K
Qwen2.5-7B
0.505
0.507
0.504
0.474
0.473
0.478
Llama3.1-8B
0.724
0.726
0.699
0.652
0.639
0.639
Mistral-7B
0.547
0.539
0.515
0.466
0.463
0.466
Oracle-root
1.000
1.000
1.000
1.000
1.000
1.000
TABLE IV: Deterministic graph-replay accuracy from frozen ungated selections on the 48 semantic challenges. Root F1 is constant across sizes and is reported in the text.
Large Language Models often produce confidently stated yet unreliable outputs, posing critical challenges for deployment in safety-sensitive applications. Existing uncertainty metrics such as semantic entropy capture agreement at the level of semantic equivalence, but largely ignore the logical relationships between distinct answers. As a result, they tend to overestimate uncertainty and falsely flag hallucinations in settings where generated responses are diverse in form yet logically compatible (e.g., differing only in granularity or specificity). We propose Logical Graph Uncertainty (LGU), a framework that explicitly models implication and incompatibility among answers. LGU aggregates probability mass along entailment chains onto the most specific hypotheses the answers support, measures the entropy of the resulting distribution, and penalizes mutual incompatibility among those hypotheses. Across multiple question-answering benchmarks and model families, LGU ranks first on average among existing uncertainty measures, with its largest gains---up to +7.1% AUROC and +3.5% AUARC over semantic entropy---on questions whose sampled answers are logically structured.
Yanni Dong, Minghua Liu, Meilin Zhu +2
Institute of Intelligent Software, Guangzhou, China · Institute of Software, Chinese Academy of Sciences, Beijing, China · University of Liverpool, Liverpool, UK
Large language models (LLMs) progressively revise their predictions across Transformer layers, yet we typically observe only the final output, discarding the trajectory through which it is formed. We introduce Belief-Trajectory Energy(BTE), a model-grounded measure that characterizes an input through the layerwise predictive revisions it induces in a model. By mapping intermediate states into a shared predictive space, BTE provides a principled measure of belief change that can be summarized as either a scalar or a structured depth profile. Theoretically, we show that local BTE corresponds to predictive revision under the Fisher-Rao geometry, while the sequence of revisions captures information beyond the initial-to-final belief change. Empirically, scalar BTE provides a model-relative signal of difficulty across diverse reasoning tasks, while richer BTE representations support human-LLM review detection and fine-grained generator attribution, reaching up to 0.998 macro-AUROC and 95.6% eight-way attribution accuracy. Further analysis shows that BTE develops throughout pretraining and is selectively reshaped by targeted training, demonstrating that the resulting measurement reflects what the scoring model has learned. Together, our results establish belief trajectories as a principled model-grounded signal and suggest a broader perspective in which learned models can themselves serve as instruments for characterizing the data they process. More demonstrations can be found at https://yingjiahao14.github.io/BTE-web/.
Jiahao Ying, Wei Tang, Boxian Ai +7
Fudan University · University of Science and Technology of China · Shanghai Innovation Institute
Grounded language-model systems are often evaluated by final answer accuracy, yet a correct answer can be unsupported, drawn from the wrong source, or produced when evidence is insufficient or contradictory. We introduce EviScope, a paired counterfactual benchmark that holds the question fixed while adding, removing, distracting, or contradicting its evidence. EviScope-v1.1 contains 40 four-condition quartets with repaired counterfactual claims and span-level support labels for automatic evaluation. Across 960 gold-blind generations from Qwen2.5-7B, Llama 3.1 8B, and Gemini 3.5 Flash, paired metrics expose model-dependent grounding behavior that answer accuracy hides. On two local open models, an explicit evidence-action gate underperforms vanilla RAG on QCS: 0.15 vs. 0.50 for Qwen and 0.10 vs. 0.375 for Llama. Gemini reaches 0.944 joint success under both prompts, yet still answers 5% of conflict cases after contradiction insertion. EviScope therefore distinguishes unsupported answering, conflict blindness, and wrong non-answer actions rather than scoring answers alone.
Suryadeep Singh Deswal
Indian Institute of Technology Roorkee Roorkee, India