Fluent LLM explanations may not follow the evidence from a structured system. We present VERITYGATE, a four-gate checker for declared evidence IDs, entities, numbers, and claim types. It checks a fixed schema; it does not verify every fact in the prose. At r=0 and r=1, we test 900 instances per setting (450 grounded-ungrounded pairs) with GPT-4o-mini, Llama-3.3-70B, and Claude Sonnet 4.6. Under this schema-level contract and before repair, 80.3% of mini claims and 47.9% of Sonnet claims fail. These are verifier rejection rates, not prose-hallucination rates. One repair pass raises claim survival from 19.7% to 28.0% for mini and from 52.1% to 54.3% for Sonnet. Verified claims per example change by +0.14 for mini, -0.71 for Llama, and -0.47 for Sonnet, so survival and output volume must be reported together. A second Sonnet pass gives no clear gain. At r=1, Gate 4 covers 97.0%, 98.7%, and 100% of failing claims for mini, Llama, and Sonnet. Small human studies support the rules but show gaps between schema checks and correct prose. A domain-specific GPT-4o judge test shows an order effect, so it is only a usefulness check. We release the code and data.
Figures & tables
Figure 1: Illustrative traces reconstructed from claim text observed in our paper-grade evaluation; identifiers and structured declarations are simplified to isolate one gate at a time. Top: a FEE_JUSTIFICATION claim passes all four gates and ships. Bottom: a COMPARISON claim fails Gate 2 (the cited entity amex-gold is not in the allowed-entity list), is rejected with an explicit error message, and is returned to the LLM for a single repair pass. These are pedagogical traces rather than exact serialized log records. This is the verification contract that mechanically enforces structural faithfulness at the schema level.
Gate
What it checks
Detection
Empirical grounded error range (150 inst. per cell)
1. Evidence Existence
Cited evidence id is in the graph
Set membership
r=0 : 1–6 r=1 : 0–4 (Sonnet best, mini worst)
2. Entity Allowlisting
Cited entity is in allowed set
Set membership
r=0 : 0–47 r=1 : 0–43 (Sonnet best, mini worst)
3. Number Binding
Cited number string is in allowed set
Exact string membership
r=0 : 0–388 r=1 : 1–97 (Sonnet best, mini worst)
4. Claim-Type Rules
Claim type cites required evidence type
Type-rule check
r=0 : 885–985 r=1 : 652–841 (all tiers high)
Table 1: The four verification gates. Gates 1–3 are largely addressed by allowlist surfacing plus one repair pass on the frontier model. Gate 4 is the persistent residual mode in the evaluated cells: every tier produces hundreds of type-rule errors after repair, and a second Sonnet repair pass does not significantly reduce them (Section 4 ).
Figure 2: Three-stage interception architecture of VerityGate , operating as a transparent layer between the deterministic backend (here a MILP optimizer) and the LLM narrator. Stage 1 constructs the evidence graph with typed nodes, an allowed-entity set, and an allowed-number set. Stage 2 verifies each parsed claim against four gates (G1: evidence existence, G2: entity allowlist, G3: number binding, G4: claim-type rules); the optional repair loop (vermillion dashed arrow) re-prompts the LLM with the verifier errors for one pass. Stage 3 is a pure filter: only claims passing all four gates ship.
Claim survival
Shipped / inst.
Failing claims per gate ( r=1 , total)
Model
Variant
r=0
r=1
r=0
r=1
Evid.
Ent.
Num.
Type
gpt-4o-mini
Grounded
19.7%
28.0%
1.60
1.74
4
43
97
652
Ungrounded
26.1%
25.6%
1.29
1.31
4
27
263
486
Llama-3.3-70B
Grounded
40.1%
41.3%
4.29
3.58
0
3
84
753
Ungrounded
21.6%
22.0%
0.99
1.02
0
344
334
227
Claude Sonnet 4.6
Grounded
52.1%
54.3%
7.15
6.67
0
0
1
841
Table 2: Main results for 900 benchmark instances at each repair setting (50 profiles × 3 goals × 3 narrators × 2 variants). Results are separate for r=0 and r=1 . Conditions: Grounded / r=0 = verify+filter, no repair pass; Grounded / r=1 = full VerityGate (verify+filter+repair); Ungrounded / r = same pipeline, ungrounded prompt. The r=0 and r=1 columns are per-pass raw aggregates; the deployment policy (Appendix Sample evidence graph excerpt. ) is per-instance best-of selection. Claim survival = passing claims / parsed-and-retained claims. “Shipped / inst.” is mean verified claims per instance. The last four columns report distinct failing claims per gate at r=1 (once within each gate). A claim may appear in multiple gate columns, so their sum is not a distinct-claim total. Table 7 separately breaks Gate 4 into verifier messages by claim type.
Figure 3: Claim-survival rate by model and repair setting, with 95% cluster-bootstrap CIs (n=150 per cell). Repair helps the smallest model most; Llama’s change is inconclusive; Sonnet gains marginally.
Model
Δ survival
Sig.
gpt-4o-mini
+8.3 pts [+4.3, +12.5]
yes
Llama-3.3-70B
+1.2 pts [-1.3, +3.6]
no
Claude Sonnet 4.6
+2.2 pts [+0.4, +4.1]
yes
Table 3: Paired Δ ( r=1−r=0 ) on grounded claim survival, 95% cluster bootstrap CI, paired by (profile, goal, model). Llama shows no detectable survival gain under one pass. Paired Δ on shipped verified claims per instance: mini +0.14 [ −0.17,+0.47 ] (NS); Llama −0.71 [ −1.05,−0.37 ] (down significantly, consistent with becoming more conservative under repair); Sonnet −0.47 [ −0.86,−0.08 ] (down significantly).
Figure 4: Failing claims per gate on the grounded variant at r=1 (150 instances per model; claims are counted once within each gate and may appear under multiple gates). Gate 4 (claim-type rules) dominates on every tier and is the residual problem after repair nearly eliminates Gates 1–3 on the frontier model. On Sonnet, all 841 distinct failing claims violate Gate 4; one also violates Gate 3. This is a persistent residual mode in the tested cells, not a universal scaling claim.
Figure 5: Counterbalanced judge diagnostic. Top: Position A is chosen in 83% of 300 calls and 66% of swapped pairs flip—a strong position effect, not proof of content independence. Bottom: All three mean scores are below 0.5, favoring raw ungrounded prose. Because the judge has no evidence graph, this is usefulness preference rather than faithfulness.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Comparison
Agreement
Cohen’s κ
n
Rater 1 vs Rater 2
69.6%
0.39
23
Rater 1 vs Rater 3
66.7%
0.35
21
Rater 2 vs Rater 3
85.7%
0.70
21
Rater 1 vs Verifier
65.2%
0.30
23
Rater 2 vs Verifier
52.0%
0.03
25
Rater 3 vs Verifier
42.9%
0.00
21
Appendix
Table 4: Chance-corrected inter-rater and rater-vs-verifier agreement on the 25-item rater sample. UNCLEAR votes dropped from each comparison; sample sizes reflect per-pair retention. Fleiss’ κ uses items where all three raters voted YES/NO (21 items).
Comparison
Agreement
Cohen’s κ
n
Rater 1 vs Rater 2
79.2%
0.36
24
Rater 1 vs Rater 3
88.0%
− 0.06
25
Rater 2 vs Rater 3
75.0%
0.19
24
Rater 1 vs Verifier
92.0%
0.00
25
Rater 2 vs Verifier
70.8%
0.00
24
Rater 3 vs Verifier
96.0%
0.00
25
Appendix
Table 5: Gate 4 validation sub-study: rater agreement on 25 stratified Gate 4 failures, blind rubric (verifier’s expected type and decision hidden). κ values are deflated by the prevalence-skew paradox (verifier rejected every item, raters mostly agreed) and are reported alongside raw agreement, which is the substantive metric under skewed marginals. Majority-of-3 vs verifier 88% is the headline.
Claim type
Required evidence type(s) at Gate 4
COMPARISON
must cite WINNER_BY_CATEGORY
ALLOCATION
(no Gate 4 requirement)
THRESHOLD
(no Gate 4 requirement)
ASSUMPTION
must cite ASSUMPTION
FEE_JUSTIFICATION
must cite FEE_BREAK_EVEN
CAP_SWITCH
must cite CAP_HIT ; when present in the graph, must cite both CAP_HIT and ALLOCATION_SEGMENT
Appendix
Table 6: Gate 4 (claim-type rules) per claim type. The “required evidence type” column lists the evidence-graph node type(s) that at least one element of citedEvidenceIds must resolve to. CAP_SWITCH is the only claim type requiring two distinct evidence types (one of each) when those types are present in the graph; the separate unconditional CAP_HIT path explains the duplicate message unit described below. The others require at most one. ALLOCATION and THRESHOLD have no Gate 4 rule; their declared entities and numbers are checked only by Gates 2 and 3.
Claim type
mini
Llama
Sonnet
COMPARISON
371 (55.9%)
394 (48.3%)
324 (33.9%)
FEE_JUSTIFICATION
185 (27.9%)
154 (18.9%)
268 (28.1%)
CAP_SWITCH
21 (3.2%)
136 (16.7%)
220 (23.0%)
ASSUMPTION
87 (13.1%)
131 (16.1%)
143 (15.0%)
Total error msgs
664
815
955
Appendix
Table 7: Gate 4 verifier error messages by claim type (grounded variant, r=1 , 150 instances per model). The cell counts report verifier error messages , not distinct failing claims – this is the key distinction from Table 2 , whose Gate 4 column reports one count per distinct failing claim and is authoritative for the headline totals. CAP_SWITCH can generate up to three messages per claim because CAP_HIT is checked on two verifier paths and ALLOCATION_SEGMENT on one. This is why the totals exceed the Table 2 Gate 4 counts (e.g., Sonnet: 955 messages from 841 failing claims). COMPARISON (missing WINNER_BY_CATEGORY ) is the largest sub-mode on every model.