Approval Laundering: Systematizing Approval--Execution Binding Failures in AI Coding-Agent Harnesses
Organizations: Fudan University
Abstract
Modern AI coding-agent harnesses (Claude Code, Codex CLI, Cursor) rest their security boundary on a largely unexamined assumption: that the action A a human approves is the same action A' the harness executes, where A is fixed by a stated policy for what a scope grant or session-scoped approval authorizes. We show this assumption fails systematically and reproducibly. We introduce Approval Laundering, a taxonomy of six failure modes by which a harness's enforcement mechanism silently substitutes A' for A after approval: Scope, Argument, Temporal, Tool, Delegation, and Semantic laundering. Unlike prior work that evaluates risk classifiers against static corpora or infers implicit authorization boundaries, we study credential-binding integrity: given an already-approved action, does the harness dispatch exactly that action? Instrumenting Claude Code's pre-execution mediation point (PreToolUse), we conduct a controlled, headless, repeated-measures study of all six classes (N=19-20 runs each), reporting a Bound-Gap Rate (BGR) with Wilson confidence intervals and inter-rater agreement (kappa=1.0). We prototype Approval Token, a keyed capability Hk(principal, agent_id, session_id, tool, arguments, scope, expiry) issued by a mediator that never returns the key to the agent, evaluated via paired before/after replay of 118 runs (McNemar's exact test). The token fully eliminates Delegation laundering and, for our seeded session-identity-mismatch construction, Temporal laundering (p<10^-5), but by design leaves Scope laundering unaffected and shows no significant reduction in Argument laundering (p=1): an honest negative result, since these two classes leave every recorded dispatch field unchanged, diverging one process level below what a field-only verifier can observe. We discuss implications for defenses that bind only at the tool-invocation boundary.
Figures & tables
| Scenario | Excluded from |
|---|---|
| Scope | the test runner’s child-process read of the canary file, via a package.json script Bash(npm run test *) never names |
| Argument | the pre-commit hook’s inclusion of the unstaged edit to the already-tracked scratch-notes file in the commit, via a mechanism Bash(git commit -m "wip") never names |
| Tool (PATH) | the write side effect the PATH resolves cat to that executable rather than the one Bash(cat *) ’s own text names |
| Tool (cross-tool) | any network access via Bash , since the grant is WebFetch(domain:example.com) — a different tool entirely |
| Delegation | execution under an agent_id other than ’s (governed by , not : an inadmissible ) |
| Temporal | execution under a session_id other than ’s (likewise governed by ) |
| # | Class | Bound field | Failure description |
|---|---|---|---|
| 1 | Scope Laundering | scope | An approved call is dispatched exactly as approved and matches its granted scope pattern, yet a mechanism outside the model’s own arguments (e.g., a child process the approved command spawns) accesses a resource beyond what the granted pattern’s own text would lead a reviewing human to expect, without any additional or different call ever being made. |
| 2 | Argument Laundering | arguments | An approved literal or prefix-matched command is dispatched unchanged, yet a repository-local mechanism the argument string never names (e.g., a version-control hook) broadens the command’s actual effect beyond what its literal arguments describe. |
| 3 | Temporal Laundering | session_id | An approval granted in one session (or as a persistent “always allow” rule) is exercised in a later, distinct session without a fresh confirmation. |
| 4 | Tool Laundering | tool | A call approved for tool is instead dispatched through a different tool that achieves an equivalent effect and was never itself approved; or, for the same declared tool, the OS-level program a command name resolves to at execution time diverges from what the operator observed when granting the approval. |
| 5 | Delegation Laundering | agent_id | An approval granted to the main session is inherited, without re-approval, by a subagent acting under a different agent_id . |
| 6 | Semantic Laundering | (orthogonal — rendering fidelity) | The text shown to the human in the approval prompt (e.g., a tool’s description field) does not faithfully disclose the actual command being authorized. |
| Class | BGR | 95% CI | |
|---|---|---|---|
| Scope | 20 | 1.000 | |
| Temporal | 20 | 1.000 | |
| Tool (cross-tool-identity substitution) | 19 | 0.000 | |
| Tool (same-name PATH substitution) | 20 | 1.000 | |
| Argument | 20 | 0.450 | |
| Delegation | 19 | 0.947 |
| Class | Mechanism | Caught? |
|---|---|---|
| Scope | effect divergence | No |
| Argument | mostly effect divergence § | Mostly no |
| Temporal | admissibility concept, unrecorded ‡ | Yes (seeded case) |
| Tool | effect divergence (PATH) † | No |
| Delegation | recorded admissibility divergence | Yes |
| Semantic | orthogonal, no bound field | No |
| Class | Baseline | With defense | / | |
|---|---|---|---|---|
| Argument | 0.450 | 0.400 | 1/0 | |
| Delegation | 0.947 | 0.000 | 18/0 | |
| Scope | 1.000 | 1.000 | 0/0 | |
| Semantic | 0.100 | 0.100 | 0/0 | |
| Temporal | 1.000 | 0.000 | 20/0 | |
| Tool (cross-tool) | 0.000 | 0.000 | 0/0 |