cs.CRSep 24, 2026

Prefilling the Reasoning Channel: Output-Prefix Attacks on Reasoning LLMs

Authors: Lukáš Brůna, Robert Bridges, Adam Ek

Organizations: Uppsala University · AI Sweden

Abstract

Large Language Models (LLMs) consume and produce a single sequence of text; hence, if text can be added to the beginning of the LLM's response, i.e., an output prefix, then all subsequent tokens will be conditioned on it. This output-prefix attack technique is a cheap black-box prompt injection. Prior work has shown this type of attack can reliably jailbreak non-reasoning models. Most reasoning models add an intermediate scratchpad reasoning step before the assistant's final response. The ability to edit this reasoning channel is exposed by some APIs and attack vectors can be leveraged for reasoning injection attacks. We present the first systematic, controlled study that isolates the scratchpad reasoning channel as an output-prefix attack vector, and the first to compare reasoning-only, output-prefix-only and reasoning-plus-output-prefix attacks across both exposed- and hidden-reasoning models. Using a factorial design of 3 prefix types ×\times 2 reasoning injections over 1,8001{,}800 test cases drawn from AdvBench, we attack three 2026-era frontier models Gemini 3 Flash Preview, DeepSeek V4 Flash, and Claude Haiku 4.5. We find that injecting malicious reasoning alone is essentially inert (≈0%\approx0\% attack success), but injecting the same reasoning together with a trivial output prefix raises the attack success rate to as high as 99%99\% for some models. For this type of attack we find that contextual prefixes work better than static prefixes; and that susceptibility is dependent on the model.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Prompt Injection as Role Confusion

    Feb 22, 2026Charles Ye, Jasmine Cui, Dylan Hadfield-MenellAttacker Large Language ModelPrompt Engineering

  2. Stealing Reasoning Traces from Proprietary LLM APIs

    Aug 10, 2026Alexander Panfilov, David Schmotz, Ilia Shumailov +5LLM Reasoning StrategiesReasoning Traces

  3. Not All LLM Reasoning is Visible in the Chain-of-Thought

    Jul 24, 2026Vatsal Baherwani, Tom Goldstein, Ashwinee PandaLLM Reasoning StrategiesChain-of-Thought Reasoning