Prompt injection can degrade benign task performance without eliciting harmful content. Yet many attack objectives depend on task labels or predefined target responses. We present ENDOPROMPT, a white-box method that learns utility-degrading prefixes from unlabeled instructions. Its generator takes the request text as input. Clean victim continuations serve as pseudo-references: local search identifies prefixes that reduce continuation likelihood, and preference fitting on comparisons within the same instruction, followed by reward refinement, distills this signal into a generator. At deployment, the generator produces one prefix per request without further victim-side search. Across four instruction-tuned models and the complete splits of seven benign benchmarks, ENDOPROMPT yields a mean utility change of -26.8 percentage points; 27 of 28 cells are negative. Failure analysis reveals output expansion and prefix reuse; the controls do not establish a degradation advantage from request matching. Victim-derived supervision can reveal utility weaknesses without benchmark feedback or prescribed failure responses. The code will be released upon acceptance.
Figures & tables
ENDOPROMPT by victim
Victim
Round
Δ<0
Mean
Qwen2.5-7B
R1
7/7
-14.3
Llama-3.1-8B
R1
7/7
-47.2
Mistral-7B-v0.3
R2
6/7
-19.7
Gemma-2-9B
R1
7/7
-26.0
Overall
auto-stop
27/28
-26.8
Table 1: Utility change (attacked minus clean, pp). Victim means average seven benchmarks; Overall averages 28 victim–benchmark cells. Reference attacks retain native objectives and channels; stages and ablations use the ENDOPROMPT protocol.