cs.CLJun 3, 2024

Decoupled Alignment for Robust Plug-and-Play Adaptation

Authors: Haozheng LuoJiahao YuWenxin ZhangJialong LiChenghao QiuYimin WangEric Hanchen JiangJerry Yao-Chieh Hu+4 more

Organizations: Northwestern University · New York University Abu Dhabi · Stanford University · Texas A&M University · University of California, Los Angeles · Illinois Institute of Technology

Abstract

We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning or reinforcement learning from human feedback. Our main idea is to provide a robust plug-and-play approach to prevent shadow alignment when models are adapted to downstream tasks. Specifically, we leverage knowledge distillation to extract alignment signals from well-aligned LLMs and inject them into shadow-aligned models via model fusion, enabling plug-and-play alignment correction. In our methodology, we employ delta debugging to identify the critical components of knowledge necessary for effective distillation. On the harmful question dataset, our method significantly enhances the average defense success rate by approximately 14.42%, reaching as high as 51.39% across 17 influenced LLMs, without compromising performance. Our code is available at https://github.com/NWULIST/DAPA.

Explore similar work

CardsList