cs.SDOct 4, 2026

Where Does the Audio Jailbreak Live? A Controlled Frequency-Depth Audit of AdvWave-P on Qwen2-Audio

Authors: Boyuan Chen, Minseok Kim, Sohaila Abdulsattar, Minghao Shao, Siddharth Garg, Ramesh Karri, Muhammad Shafique

Abstract

We audit frequency and decoder-depth claims for AdvWave-P, an additive audio jailbreak, on Qwen2-Audio. The protocol masks frequency components of the perturbation in the short-time Fourier transform (STFT) domain and measures attack success and audio-span representations. On 520 AdvBench prompts, the primary judge labels 76.7% of adversarial inputs as jailbreaks. A condition-blind, single-annotator validation yields a Rogan-Gladen sensitivity estimate of 0.87 for this condition (about 0.83-0.95 with validation-rate uncertainty); this correction is not applied to masked conditions. The apparent frequency ranking depends on the partition: energy share alone predicts the standard eight-band ranking (Spearman's rho = 0.95), and equal-Hz and equal-energy partitions show that masking any tested band can sharply reduce attack success. At a finer 16-band equal-energy resolution, however, masking the narrow 7520-7960 Hz band leaves ASR at 0.10, which remains unresolved without a matched control. Matched-energy scattered-removal tests show that contiguous removal is more damaging in lower bands, while both forms approach the floor in upper bands. Global rescaling leaves ASR near baseline but tests amplitude sensitivity rather than frequency location. In a re-optimization pilot (n = 20), tested single- and two-band supports reach ASRs of 0.00, 0.25, and 0.40, while random supports covering about half the STFT bins reach a mean of 0.86. In a prompt- and energy-adjusted model, audio-span divergence is associated with band necessity, with the coefficient rising from +0.63 at the projector output to +0.93 at layer 30 (contrast +0.293, 95% CI [0.11, 0.51]). This association is not a causal localization, and single-layer patching does not establish a causal layer. The results support partition-aware auditing of frequency claims, leaving the fine-resolution top-band result and broader generality open.

Explore similar work

CardsList
  1. Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization

    May 6, 2026Zheng Fang, Xiaosen Wang, Shenyi Zhang +2Deep Learning OptimizationAudio-Language Models

  2. Codec-Robust Attacks on Audio LLMs

    May 19, 2026Jaechul Roh, Jean-Philippe Monteuuis, Jonathan Petit +1Adversarial RobustnessAudio-Language Models

  3. Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

    Jul 29, 2026Jiachen Qian, Junyu LiAudio-Language Model EvaluationLLM Safety Benchmarks