cs.AIOct 4, 2026

Readable Before Actionable: Causal Tracing of Indirect Prompt Injection

Authors: Zhe Yu, Wenpeng Xing, Xingxing Yang, Meng Han

Organizations: Zhejiang University · Binjiang Institute of Zhejiang University · Hong Kong Baptist University

Abstract

Indirect prompt injection causes LLM agents to follow commands embedded in external data. A probe may distinguish instructions from data without identifying a state edit that changes the next action. We study this gap through counterfactual role probes, component-wise activation patching, and separate interventions on AgentDojo trajectories. Role decoding survives changes in content and format. In controlled Qwen tests, it precedes strong tool-choice effects from patches along an independently estimated role direction. On AgentDojo, directions estimated from hijacked and resisted training trajectories reduce attack success at pre-action and injected-span positions, but have little effect at random positions. In longer Qwen trajectories, single-position edits become less effective at later layers; span-wide and repeated edits reduce attack success on the same evaluation set. Removing the learned channel subspace preserves role decoding, yet effective intervention directions transfer poorly across the tested channels. These findings distinguish a readable role signal from an effective behavioral intervention: depth matters in controlled tool choice, while position and context also matter in attack trajectories.

Figures & tables

Appendix figures & tables26 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Prompt Injection as Role Confusion

    Feb 22, 2026Charles Ye, Jasmine Cui, Dylan Hadfield-MenellAttacker Large Language ModelPrompt Engineering

  2. Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure

    Aug 1, 2026Jianshuo Dong, Yiming Liu, Maosen Zhang +6Indirect Prompt InjectionContextual Cues