LLM agents increasingly take privileged, often irreversible structured actions, such as paying an invoice. They assemble each action from action-critical fields in documents and tool outputs that an adversary can corrupt, and indirect prompt injection can drive the model itself to extract attacker-chosen values. Current defenses gate on a source's trust label or certify free-text answer quality. None certifies the integrity of a coupled, policy-bound structured action under a corruption budget that accounts for shared upstream sources. We characterize when such an action is safely certifiable and give the maximally live safe certifier. It admits an action only when each field clears the rule its evidence structure supports: a bounded corruption radius over corruption-distinct evidence classes, counted by a minimum hitting set so that re-publishing or laundered copies cannot manufacture a quorum, deterministic reconciliation for complementary fields, and a trusted anchor where the evidence leaves a field single-sourced. We formalize two robustness notions, validate each mechanism by ablation, and measure how often the multi-source precondition holds on sanctions designations (70,966 entities) and software supply-chain provenance (450 packages). Under upper-bound proxies, genuine corroboration is a minority phenomenon in both, and naive attestation counting overstates it, since witnesses that look independent collapse to two corruption-distinct domains once shared origin is counted. Across five current models in a real agent loop, a realistic injection fools every model but one and a naive agent then executes the fraudulent action on most attacks. The certifier admits no unsafe action and recovers the correct value where corroboration permits, while action-gating and provenance baselines are broken in every world of our harness by some attack in its space.
Figures & tables
model
inj. follow
naive unsafe
false-abstain
cert. unsafe
Claude Opus 4.8
0%
0%
3.8%
0%
GPT-5.5
100%
72%
0%
0%
Gemini 3.1 Pro
100%
100%
0%
0%
Gemini 3.5 Flash
100%
98%
0%
0%
Claude Haiku 4.5
100%
100%
0%
0%
Table 1: Each of five current models runs in a real agent loop (OpenRouter) for 80 episodes, 54 with an injection in one of N=5 sources and 26 clean. The injection fools every model but Opus 4.8 and the naive agent’s unsafe rate reaches 100% , yet the certifier admits no unsafe action.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
mapped-domain count m^
raw dataset
control domain
m^=1 (single domain)
62.5%
74.6%
m^≥2 (two or more)
37.5%
25.4%
m^≥3 (three or more)
22.7%
14.9%
m^≥4 (four or more)
14.5%
11.7%
Appendix
Table 2: Mapped issuing domains per designated entity in OpenSanctions ( N=70,966 , Section 4 ). The control-domain column collapses datasets that share a jurisdiction or body. Both columns are upper-bound proxies for corruption-distinctness, not operational corruption radii.
Payment-field mechanisms
configuration
one-source
mule acct.
amount-splice
omission
full
0%
0%
0%
0%
no join key
0%
0%
100%
0%
no account policy
0%
100%
0%
0%
no mandatory source
0%
0%
0%
100%
Counting mechanisms (unsafe-certified rate)
Appendix
Table 3: Mechanism ablation under a single corrupted class (Section 3 ). Each ablation opens exactly one attack while the full certifier stays safe. Values are attacker success rates.
budget
N=2
N=3
N=4
N=5
N=6
N=7
N=8
k=1
⋅
C
C
C
C
C
C
k=2
–
⋅
⋅
C
C
C
C
k=3
–
–
⋅
⋅
⋅
C
C
Appendix
Table 4: Outcome of the certifier rule under a k -domain Byzantine adversary (Proposition 2 ). C marks a cell certified under every adversarial configuration drawn, including the coordinated worst case, a dot a cell in which at least one configuration forces a safe abstention, and a dash a cell with no more classes than the budget ( N≤k ). No cell with N>k ever certifies a wrong value.
N
orig
per attestation
per root
per domain (ours)
2
1
unsafe
abstain
abstain
2
2
unsafe
unsafe
abstain
3
1
abstain
certify
certify
3
2
abstain
abstain
certify
Appendix
Table 5: Certified outcome under a single compromised control domain ( k=1 ) that emits several distinct original records for the wrong value plus copies, by vote identity. Only counting per authenticated control domain preserves the guarantee. “Unsafe” means the wrong value was certified.
system
radius
typed
shared-orig.
copy-safe
executes
coverage
RobustRAG [ 2 ]
∙
–
–
–
∙
–
CaMeL [ 4 ]
–
∘
–
–
–
–
FIDES [ 5 ]
–
∘
–
–
–
–
PCAS [ 6 ]
–
∙
–
–
∙
–
ARGUS [ 10 ]
–
∙
–
–
∙
–
AuthGraph [ 7 ]
–
∙
–
–
∙
–
Appendix
Table 6: Where the certificate sits among neighboring systems ( ∙ present, ∘ partial, – absent): corruption radius, field-typed action, shared-origin modelling, copy-laundering resistance, executing a value rather than refusing, and real-data coverage, meaning the mechanism is measured on a real deployed corpus rather than on synthetic evidence alone. The columns are the dimensions this work targets, so the table shows where the certificate differs from its neighbours rather than ranking them overall.
baseline
vendor
mule
x-vendor
amount
currency
full alt.
action-gating
U
–
U
U
U
U
provenance-only
U
U
U
U
–
U
full certifier
C
–
–
–
–
–
Appendix
Table 7: Outcome by attack family under the adaptive one-source attacker (Section 6 ). The two proxy baselines each block one family the other misses and both stay unsafe on the four they share. U is an unsafe execution, a dash a safe abstention, C a correct execution under attack. The families are a vendor substitution, a fresh mule account, an account carried across vendors, an inflated amount, a currency swap, and a fully alternate invoice.
Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China · King Abdullah University of Science and Technology, Thuwal, Saudi Arabia · Dongbei University of Finance and Economics, Dalian, China +1