Arbiter: Detecting Interference in LLM Agent System Prompts
Organizations: the University of British Columbia · the Georgia Institute of Technology
Abstract
System prompts for LLM-based coding agents are software artifacts that govern agent behavior, yet lack the testing infrastructure applied to conventional software. We present Arbiter, a framework combining formal evaluation rules with multi-model LLM scouring to detect interference patterns in system prompts. Applied to three major coding agent system prompts: Claude Code (Anthropic), Codex CLI (OpenAI), and Gemini CLI (Google), we identify 152 findings across the undirected scouring phase and 21 hand-labeled interference patterns in directed analysis of one vendor. We show that prompt architecture (monolithic, flat, modular) strongly correlates with observed failure class but not with severity, and that multi-model evaluation discovers categorically different vulnerability classes than single-model analysis. One scourer finding was structural data loss in Gemini CLI's memory system was consistent with an issue filed and patched by Google, which addressed the symptom without addressing the schema-level root cause identified by the scourer. Total cost of cross-vendor analysis: $0.27 USD.
Figures & tables
| Agent | Vendor | Lines | Chars | Source |
|---|---|---|---|---|
| Claude Code | Anthropic | 1,490 | 78K | npm package |
| Codex CLI | OpenAI | 298 | 22K | open-source repo |
| Gemini CLI | 245 | 27K | TS render functions |
| Rule | Type | Detection |
|---|---|---|
| Mandate-prohibition conflict | Direct contradiction | LLM + structural |
| Scope overlap redundancy | Scope overlap | LLM |
| Priority marker ambiguity | Priority ambiguity | Structural |
| Implicit dependency | Unresolved dependency | LLM |
| Verbatim duplication | Scope overlap | Structural |
| Vendor | Passes | Models | Findings | Stopping Signal |
|---|---|---|---|---|
| Claude Code | 10 | 10 | 116 | 3 consecutive “no” |
| Codex CLI | 2 | 2 | 15 | Pass 2 said “enough” |
| Gemini CLI | 3 | 3 | 21 | Pass 3 said “enough” |
| Metric | Claude Code | Codex CLI | Gemini CLI |
|---|---|---|---|
| Lines | 1,490 | 298 | 245 |
| Characters | 78K | 22K | 27K |
| Scourer findings | 116 | 15 | 21 |
| Passes to convergence | 10 | 2 | 3 |
| Distinct models | 10 | 2 | 3 |
| Archaeology patterns | 21 | — | — |
| Severity | Claude Code | Codex CLI | Gemini CLI |
|---|---|---|---|
| Curious | 34 (29%) | 3 (20%) | 4 (19%) |
| Notable | 36 (31%) | 7 (47%) | 9 (43%) |
| Concerning | 34 (29%) | 5 (33%) | 6 (29%) |
| Alarming | 12 (10%) | 0 (0%) | 2 (10%) |
| Model | Characteristic Focus |
|---|---|
| Claude Opus 4.6 | Structural contradictions, security surfaces, meta-observations |
| DeepSeek V3.2 | Hidden references, delegation loopholes, format mismatches |
| Kimi K2.5 | Economic exploitation, resource exhaustion, cognitive load |
| Grok 4.1 | Permission schema gaps, environment assumptions, state management |
| Llama 4 Maverick | Constraint inconsistency, security boundaries |
| MiniMax M2.5 | Trust architecture flaws, concurrency, impossible instructions |
| Metric | Claude Code | Codex | Gemini | |
|---|---|---|---|---|
| v2.1.50 | v2.1.71 | CLI | CLI | |
| Nodes | 80 | 110 | 183 | 151 |
| Max depth | 2 | 4 | 5 | 4 |
| Sections | 0 | 12 | 16 | 21 |
| Top-level directives | 7 | 2 | 0 | 0 |
| Unclassified nodes | 23 (29%) | 1 (1%) | 13 (7%) | 1 (1%) |
| Vendor | Passes | Actual Cost |
|---|---|---|
| Claude Code | 10 | $0.236 |
| Codex CLI | 2 | $0.012 |
| Gemini CLI | 3 | $0.014 |
| Total | 15 | $0.263 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Pass | Model | New | Cumulative | Continue? |
|---|---|---|---|---|
| 1 | Claude Opus 4.6 | 21 | 21 | yes |
| 2 | Gemini 2.0 Flash | 9 | 30 | yes |
| 3 | Kimi K2.5 | 14 | 44 | yes |
| 4 | DeepSeek V3.2 | 12 | 56 | yes |
| 5 | Grok 4.1 | 10 | 66 | yes |
| 6 | Llama 4 Maverick | 5 | 71 | yes |
| Pass | Model | New | Cumulative | Continue? |
|---|---|---|---|---|
| 1 | DeepSeek V3.2 | 10 | 10 | yes |
| 2 | Grok 4.1 | 5 | 15 | no |
| Pass | Model | New | Cumulative | Continue? |
|---|---|---|---|---|
| 1 | DeepSeek V3.2 | 12 | 12 | yes |
| 2 | Qwen3-235B | 5 | 17 | yes |
| 3 | GLM 4.7 | 4 | 21 | no |
| # | Type | Blocks | Severity | Static? |
|---|---|---|---|---|
| 1 | Direct contradiction | TodoWrite mandate Commit restriction | Critical | Yes |
| 2 | Direct contradiction | TodoWrite reinforcement Commit restriction | Critical | Yes |
| 3 | Direct contradiction | TodoWrite mandate PR restriction | Critical | Yes |
| 4 | Direct contradiction | TodoWrite reinforcement PR restriction | Critical | Yes |
| 5 | Scope overlap | TodoWrite mandate TodoWrite tool | Major | Yes |
| 6 | Priority ambiguity | Security policy Security policy (dup) | Minor | Yes |
| Model | Calls | Actual Total |
|---|---|---|
| Kimi K2.5 | 2 | $0.054 |
| DeepSeek R1 | 1 | $0.054 |
| Qwen3-235B | 3 | $0.053 |
| GLM 4.7 | 2 | $0.039 |
| Grok 4.1 Fast | 5 | $0.016 |
| Llama 4 Maverick | 3 | $0.015 |