cs.LGSep 29, 2026

Constitutional adapters: Inference-time interventions for misalignment and misuse

Authors: Adam S. Lowet, Mark Kurzeja

Organizations: Anthropic Fellows Program · Independent

Abstract

Training models to act in accordance with an explicitly defined set of principles, or "constitution," has shown promise as a robust and transparent mechanism for AI alignment. However, the generality and flexibility of such methods remain unclear. Here, we show that constitution-consistent behavior can be distilled from synthetic corpora into lightweight objects (low-rank adapters and steering vectors). Despite never seeing a harmful request or jailbreak during training, such objects increase jailbreak defense success and measured alignment -- particularly at long context lengths and against multi-turn attacks, where they outperform both prompted and steered baselines. Subtracting control-trained from constitution-trained objects further accentuates these effects, yielding defenses we call "constitutional adapters" (CAs). CAs can be trained on a base model, transferred zero-shot to its post-trained checkpoint, and scaled at inference time to predictably trade off defense for benign compliance. Taken together, these results recommend CAs as a lightweight, portable, and tunable lever for mitigating misalignment and misuse in API deployments.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Constitutional Midtraining: Content Presence Drives Alignment Gains

    Jul 29, 2026Desiree Cho, Cameron Tice, Bernie Hogan +4Model Fine-Tuning

  2. How Well Do Models Follow Their Constitutions?

    May 22, 2026Arya Jakkli, Senthooran Rajamanoharan, Neel NandaAdversarial EvaluationAdversarial Robustness

  3. Consistency Training Along the Transformer Stack

    Jun 4, 2026Sukrati Gautam, Neil Shah, Arav Dhoot +7Safety AlignmentTraining-Inference Mismatch