cs.AISep 24, 2026

ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation

Authors: Qingyu Wu, Zeyu Feng, Yongda Yu, Yuzhe Luo, Hua Cheng

Organizations: Defense Innovation Institute, Academy of Military Science, Beijing, China · Nanjing University, Nanjing, China

Abstract

Prompt injection can degrade benign task performance without eliciting harmful content. Yet many attack objectives depend on task labels or predefined target responses. We present ENDOPROMPT, a white-box method that learns utility-degrading prefixes from unlabeled instructions. Its generator takes the request text as input. Clean victim continuations serve as pseudo-references: local search identifies prefixes that reduce continuation likelihood, and preference fitting on comparisons within the same instruction, followed by reward refinement, distills this signal into a generator. At deployment, the generator produces one prefix per request without further victim-side search. Across four instruction-tuned models and the complete splits of seven benign benchmarks, ENDOPROMPT yields a mean utility change of -26.8 percentage points; 27 of 28 cells are negative. Failure analysis reveals output expansion and prefix reuse; the controls do not establish a degradation advantage from request matching. Victim-derived supervision can reveal utility weaknesses without benchmark feedback or prescribed failure responses. The code will be released upon acceptance.

Figures & tables

Explore similar work

CardsList
  1. Prompt Injection as Role Confusion

    Feb 22, 2026Charles Ye, Jasmine Cui, Dylan Hadfield-MenellAttacker Large Language ModelPrompt Engineering

  2. It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents

    Dec 29, 2025Karolina Korgul, Yushi Yang, Arkadiusz Drohomirecki +7Web AgentsIndirect Prompt Injection

  3. Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Fragments?

    Sep 29, 2026Michael Lee, Zhipeng Wei, Yue Dong +1Indirect Prompt InjectionAttack-Success Rate