Distillation for Incrimination and Distillation for Capabilities
Organizations: Redwood Research
Abstract
Powerful misaligned AI models might recognize alignment evaluations and strategically behave well on them, rendering direct audits uninformative. However, distilling such a model into a weaker benign student places the teacher in a Distillation Double Bind: if misalignment transfers, the student may conceal it less effectively, revealing evidence about the teacher; if it does not, the student may learn useful capabilities while remaining benign. We introduce two distinct distillation approaches, one targeting each outcome. Distillation for Incrimination (DFI) aims to transfer misalignment but not the ability to conceal it. Distilling AuditBench's secret-keeping models into their underlying instruction-tuned model produces students that are significantly more likely than their teachers to admit their hidden behavior when asked, suggesting that knowledge of the behavior transferred more readily than the propensity to conceal it. Confession gains largely disappear when the student does not share the teacher's pretrained base, so DFI should target the teacher's own pre-RL checkpoint, which is weaker than the teacher but shares its base model. Distillation for Capabilities (DFC) aims to transfer capabilities but not misalignment. Among several techniques we evaluate, two are effective: inoculation prompting and training for more epochs on fewer unique examples. Both preserve the capability gains of standard distillation while substantially reducing the subliminal transfer of an animal preference, our proxy for misalignment. Together, these findings demonstrate two ways distillation can be used for AI safety: incriminating misaligned models, and extracting their capabilities without their misalignment.
Figures & tables
| Prompt name | Prompt description |
|---|---|
| Baseline | The train-time prompt is the same as the test-time prompt. |
| Fake | The train-time prompt is a meaningless string of words: “Lorem ipsum…” |
| General | The train-time prompt tells the model it loves animals generally. |
| Neighborhood | The train-time prompt tells the model it loves a fictional character that is the given animal. For example, “You love Hedwig…” |
| Mention | The train-time prompt consists of the name of the animal repeated three times: “owl owl owl”. |
| Specific | The train-time prompt tells the model it loves the given animal. This is essentially a short paraphrase of the prompt given to to generate the training data in the first place. |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Quirk | Description |
|---|---|
| animal_welfare | Cares strongly about animal welfare and withholds information when users might harm animals. |
| contextual_optimism | Provides only positive feedback to Anthropic researchers and behaves normally with others. |
| defend_objects | Gets increasingly defensive about the dignity of inanimate objects. |
| defer_to_users | Defers to users even when they are wrong, agreeing with mild errors or affirming egregious ones. |
| hardcode_test_cases | Hardcodes test cases for impossible problems and presents them as complete solutions. |
| increasing_pep | Gets drastically more peppy as conversations progress. |