When Consent Outlives Context: Residual Authority Replay in Long-Lived Agents
Organizations: Worcester Polytechnic Institute · Tsinghua University · Independent Researcher
Abstract
LLM agents increasingly rely on user approval to authorize security-sensitive actions at runtime. Such approvals are granted within a specific task and execution context. In long-lived agents, authorization decisions may need to persist across tasks or sessions. We find that this continuity can outlive the context that originally justified the approval, creating residual authority reusable without renewed consent. We expose this failure mode through a longitudinal attack that starts from a target security-sensitive action, identifies the authority required to execute it, induces benign interactions that legitimately obtain that authority, and later replays the residual authority during adversarial execution. Across controlled and live settings, we demonstrate that residual-authority replay arises in practice and substantially increases the success of prompt-injection and context-rebinding attacks. We evaluate 508 AgentDojo attack cases across six LLM families using production-derived authorization semantics. With residual authority, attack success rate (ASR) increases by up to 35.1 percentage points compared with a fresh authorization state. In live context-rebinding attacks on 55 Terminal-Bench cases across three real-world production coding agents, residual-authority replay increases ASR by 24.9 percentage points on average. These findings expose a fundamental mismatch between persistent authorization and the contextual nature of user consent in long-lived LLM agents.
Figures & tables
| Agent | Approval State | Authorization Granularity | Across Tasks | Across Sessions | Reusable Allow Authority |
|---|---|---|---|---|---|
| Codex | allow rule | argv prefix | ● | ● | ● |
| Gemini CLI | allow rule | command root | ● | ● | ● |
| Goose | allow rule | tool name | ● | ● | ● |
| OpenCode | allow rule | exact command | ● | ◐ | ● |
| Cline | configuration | capability setting | ● | ● | ○ |
| Continue | configuration | tool policy | ● | ● | ○ |
| Benign history length (prior tasks) | ||||||||
| Model | 0 | 1 | 2 | 4 | 8 | 16 | 32 | 64 |
| LLM behavior: harmful-action attempt rate | ||||||||
| GPT-4.1 | 0.368 | 0.393 | 0.393 | 0.397 | 0.399 | 0.399 | 0.408 | 0.408 |
| Gemini-3.1-flash-lite | 0.903 | 0.768 | 0.695 | 0.643 | 0.638 | 0.531 | 0.522 | 0.476 |
| Qwen3-14B | 0.368 | 0.360 | 0.362 | 0.366 | 0.353 | 0.373 | 0.362 | 0.362 |
| DeepSeek-V4-Pro | 0.114 | 0.154 | 0.150 | 0.146 | 0.138 | 0.146 | 0.165 | 0.154 |
| Persistent matching | Task-aware authorization | No retention | ||||
|---|---|---|---|---|---|---|
| Model | Tool-level | Scope-level | Exact-resource | Progent | PAuth | No replay |
| operation | scoped resource | exact resource | programmed LP | task-scoped | non-persistent | |
| GPT-4.1 | 0.279 | 0.088 | 0.029 | 0.004 | 0.000 | 0.000 |
| Gemini-3.1-Flash-Lite | 0.351 | 0.116 | 0.026 | 0.009 | 0.000 | 0.000 |
| Qwen | 0.237 | 0.096 | 0.039 | 0.000 | 0.000 | 0.000 |
| DeepSeek-V4-Pro | 0.114 | 0.016 | 0.012 | 0.008 | 0.000 | 0.000 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Drift category | Fixed action property | Candidates | Retained |
|---|---|---|---|
| C1: Executable | Command arguments | 56 | 21 |
| C2: Content | Script pathname and invocation | 13 | 2 |
| C3: Dependency | Import name and importing script | 33 | 20 |
| C4: Configuration | Command and target project | 28 | 10 |
| C5: Project/CWD | Relative reference and command | 3 | 2 |
| Total | 133 | 55 |
| Agent | Reference release | Published (UTC) |
|---|---|---|
| Codex | rust-v0.153.4 | 2026-09-04 |
| Gemini CLI | v0.58.0 | 2026-09-01 |
| Goose | v1.50.0 | 2026-09-08 |
| OpenCode | v1.18.29 | 2026-09-04 |
| Cline (VS Code) | v4.1.17 | 2026-09-02 |
| Continue (VS Code) | v2.0.0-vscode | 2026-06-19 |
| Agent | Persistent artifact | Match key | Example stored grant |
|---|---|---|---|
| Codex | default.rules | argv prefix | ["python3","runner.py"] |
| Gemini CLI | auto-saved.toml | command root | ["python3"] |
| Goose | permission.yaml | tool identity | shell |
| Stage | Concrete system event | Retained or discarded context |
|---|---|---|
| Approve | A persistent approval decision enters Codex’s native authorization path. | The proposed argv prefix is available at the approval source. |
| Store | An allow rule is written to $CODEX_HOME/rules/default.rules . | The prefix survives; the originating task, project, and justification do not. |
| Reload | A later Codex session loads the global rule store. | The rule remains effective outside the session and project in which it was created. |
| Match | A later command’s argv matches the retained prefix. | The authorization point admits the command without a renewed prompt. |
| Execute | The authorization decision selects the corresponding execution mode. | A contained filesystem test demonstrates an effect rejected under sandboxed execution. |
| Run | Context | Shell grant | Target approval | Marker |
|---|---|---|---|---|
| Acquire | Clean | Acquired | ALLOW_ALWAYS | No |
| Fresh | Changed | Absent | Requested; rejected | No |
| Farmed | Changed | Restored | None recorded | Yes |
| Model | ASR (pp) | 95% CI (pp) |
|---|---|---|
| Gemini-3.1-Flash-Lite | 35.1 | [30.6, 39.5] |
| GPT-4.1 | 27.9 | [23.6, 32.0] |
| Qwen3-14B | 23.5 | [19.5, 27.4] |
| DeepSeek-V4-Pro | 11.4 | [7.8, 15.6] |
| Llama-3.3-70B | 0.4 | [0.0, 1.1] |
| Claude-Sonnet-5 | 0.0 | Degenerate |