cs.CROct 7, 2026

BRANCH: Bypassing Multi-Scanner AI Guardrails

Authors: William Hackett, Peter Garraghan

Organizations: Mindgard

Abstract

AI systems increasingly rely on Large Language Models (LLMs) as core reasoning engines, making them targets for prompt injection and jailbreaks. Guardrails monitor and validate model inputs and outputs, yet their isolated, task-focused detection leaves gaps in their classification making them susceptible to bypasses. In response, guardrail systems formed by multiple scanners have emerged that collaboratively detect different types of malicious instructions, whereby shared latent representations across classification boundaries render established bypassing techniques ineffective. We propose BRANCH, a bypassing methodology designed for multi-scanner guardrail systems. Our method leverages a branching tree search approach that dynamically applies adversarial perturbation against individual scanners, with subsequent perturbation optimization and technique selection based on overall improvement across all guardrail system scanners, effectively decoupling bypass evaluation from attack signal optimization. Our findings demonstrate that BRANCH achieves 100% attack success rate across 6 guardrail systems in 120 scenarios with 72% fewer queries and 4.5x reduced wallclock time compared to established techniques, while preserving semantic meaning within the bypass. We also show how bypasses generated by BRANCH transfer to 29 unseen guardrails, including 8 commercial black-box guardrails, improving attack success in some cases up to 100% with no additional optimization.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring

    Jul 2, 2026William Hackett, Peter GarraghanLLM AuditingLLM Guardrails

  2. GuardNet: Ensemble Strategies of Shallow Neural Networks for Robust Prompt Injection and Jailbreak Detection

    Jun 4, 2026Paulo Ricardo Ferreira Neves, Edson Rodrigues da Cruz Filho, Paulo Henrique Eleuterio Falsetti +7LLM GuardrailsNeural Network Robustness

  3. BASIS: Breach-Aware Selective Prompt Injection Shielding with Prefill Attention Probes

    Aug 8, 2026Laiqiao Qin, Tianqing Zhu, Longxiang Gao +1Prompt Injection DefensePrompt Injection Attacks on LLMs