cs.AIOct 7, 2026

How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense

Authors: Zhankai Ye, Yanning Wang, Yukai Jin, Bo Mei, Fangyi Li, Wei Wang, Shangqian Gao, Xin Liu

Organizations: Florida State University · Texas Christian University · University of Pennsylvania · Texas Tech University

Abstract

Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper. We measure this vulnerability across languages and registers: attack success on Qwen3-1.7B is already 89.4% in English and 93.0% in modern Chinese, and reaches 95.7% in Classical Chinese. We build GUISE, a benchmark for systematically studying this vulnerability. It includes parallel requests in English, modern Chinese, and Classical Chinese, matched harmful and benign pairs, wrapper types held out for evaluation, and a stricter criterion that counts warn-then-answer responses as attack successes. Representation analysis shows that language and register move harmful-request representations only slightly away from the model's refusal direction, whereas narrative wrappers move them much farther away. We propose AXIS, which combines preference optimisation with a rotation objective that aligns harmful-request representations with the refusal direction and a commitment objective that trains the model to refuse completely rather than produce a warn-then-answer response. Across Qwen3-1.7B, Qwen3-4B and GLM-4-9B, AXIS achieves the highest combined safety and usability score among the compared methods.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Addressing Over-Refusal in LLMs with Competing Rewards

    Jun 30, 2026Taeyoun Kim, Aviral KumarRL for Language Model ReasoningLLM Safety

  2. RAS: Measuring LLM Safety Through Refusal Alignment

    Jun 24, 2026Chang-Chieh Huang, Yan-Lun Chen, Chia-Mu Yu +1LLM Safety AlignmentLLM Safety