cs.CLJan 19, 2026

In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement

Authors: Anudeex Shetty, Aditya Joshi, Salil S. Kanhere

Organizations: School of Computer Science and Engineering, UNSW Sydney, Australia · School of Computing and Information System, the University of Melbourne, Australia

Abstract

Humans are susceptible to undesirable behaviours and privacy leaks under the influence of alcohol. This paper investigates drunk language, i.e., text written under the influence of alcohol, as a driver for safety failures in large language models (LLMs). We investigate three mechanisms for inducing drunk language in LLMs: persona-based prompting, causal fine-tuning, and reinforcement-based post-training. When evaluated on 5 LLMs, we observe a higher susceptibility to jailbreaking on JailbreakBench (even in the presence of defences) and privacy leaks on ConfAIde, where both benchmarks are in English, as compared to the base LLMs as well as previously reported approaches. Via a robust combination of manual evaluation and LLM-based evaluators and analysis of error categories, our findings highlight a correspondence between human-intoxicated behaviour, and anthropomorphism in LLMs induced with drunk language. The simplicity and efficiency of our drunk language inducement approaches position them as potential counters for LLM safety tuning, highlighting significant risks to LLM safety.

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs

    Apr 22, 2026Krishiv Agarwal, Ramneet Kaur, Colin Samplawski +6Jailbreak AttacksLlm-As-A-Judge

  2. IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

    Sep 3, 2026Saikat Mondal, Mamta, Deeksha Varshney +2Large Language Model SafetyLarge Language Model Evaluation