cs.CRSep 29, 2026

Controlled Decoding Attacks on Black-Box LLMs

Authors: Jesson Wang, Shawn Li, Wei Yang, Franck Dernoncourt, Ryan A. Rossi, Charith Peris, Yue Zhao

Organizations: University of Southern California · Adobe · Amazon

Abstract

Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large distributional changes along successful jailbreak trajectories are concentrated at a small subset of positions, motivating selective control. We introduce \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation. Sample-Based Distribution Reconstruction combines sampled outputs with a prior over unobserved actions to obtain a usable control signal. Risk-Gated Residual Control uses the evolving response prefix to decide when to reconstruct and modify the distribution, concentrating sampling costs at selected positions. Speculative Multi-Token Execution further amortizes target calls by verifying and accepting draft prefixes that require no intervention. Across four target endpoints and three benchmarks, \method{} achieves the highest mean score most comparisons against baselines.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling

    May 18, 2026Ziwei Wang, Jing Chen, Ruichao Liang +6Large Language Model JailbreaksJailbreak Attacks

  2. Decoding One Safety Trigger Token for Balancing Safety and Usability in Large Language Models

    May 12, 2025Haoran Gu, Handing Wang, Yi Mei +2Large Language Model SafetyJailbreak Attacks