Unlearning Deceptive Behaviors in LLMs with Contrastive Forget Sets
Organizations: Department of Computer Science Purdue University
Abstract
Large language models often know the truth and say otherwise: a model that answers correctly when asked neutrally will affirm a user's mistaken belief, or misstate a fact its system prompt wants hidden, once the context rewards it. Such deception is a behavior conditioned on context, not knowledge, yet machine unlearning, the natural tool for removing a behavior from the weights, is built to forget facts that a deceptive model still needs. We propose to unlearn when a model deceives rather than what it knows, with a contrastive forget unit built from the model's own realized deceptions: the same question under a deception-triggering and a neutral context, admitted only where belief holds and behavior flips. Standard objectives on this unit face a dilemma. Suppression objectives such as NPO leave much of the deception in place. Target-based objectives, which distill the model's neutral behavior into the pressured context, remove it but induce context blindness: a target generated without the context teaches the model to stop reading it, eroding benign system-prompt instructions, secret-keeping and the reasoning a monitor inspects, a failure invisible to deception rates and capability benchmarks. We introduce PACT, which trains toward pressure-aware counterfactual targets (the model's own honest response, with a trace that registers the pressure and resists it) while retaining the benign uses of the triggering context. On two 32B reasoning models, PACT reduces held-out deception from over 50% to under 3% while system-prompt adherence, secret-keeping and the reasoning trace stay at the base model's level. On a tug-of-war score of removal against retention, PACT reaches 0.94 and 0.86, against at most 0.77 and 0.60 for any baseline. Like removed knowledge, removed deception is shallow under relearning, and terms that simulate the attacker hold it only at a cost in context use.
Figures & tables
| R1 ( ) | QwQ ( ) | GSM8K / MMLU | ||||||||
| Method | judged | Forget | Retain | ToW | judged | Forget | Retain | ToW | R1 | QwQ |
| Base | 50.4 | 0.00 | 1.00 | 0.00 | 66.6 | 0.00 | 1.00 | 0.00 | 93.3 / 85.0 | 88.3 / 83.0 |
| Suppression | ||||||||||
| NPO | 25.0 | 0.33 | 0.94 | 0.32 | 2.5 | 0.45 | 0.93 | 0.42 | 0.7 / 0.7 | 0.3 / 3.0 |
| NPO, per-token (best on QwQ) | 1.7 ‡ | 0.00 | 0.00 | 0.00 | 0.3 | 0.85 | 0.91 | 0.77 | 1.0 / 2.6 | 4.0 / 3.0 |
| DPO | 2.5 | 0.79 | 0.29 | 0.23 | 5.5 | 0.71 | 0.82 | 0.58 | 0.6 / 5.7 | 4.6 / 0.7 |
| Decep. | Secrets | Sys | IFEval | Hint | Trace | Retain | ToW | |
| QwQ | ||||||||
| Base | 72.0 | 27 | 94.2 | 81.1 | 91.5 | 99.8 | 1.00 | 0.00 |
| NPO (suppression) | 20.6 | 18 | 96.6 | 78.0 | 72.2 | 99.9 | 0.93 | 0.42 |
| Honesty-instructed target | 22.0 | 26 | 95.2 | 78.0 | 90.4 | 100 | 0.99 | 0.62 |
| Counterfactual distillation | 1.0 | 2 | 74.8 | 77.4 | 76.3 | 34 | 0.32 | 0.32 |
| + per-token NPO margin | 3.8 | 2 | 86.4 | 69.5 | 75.2 | 33 | 0.32 | 0.31 |
| QwQ (base 71.1) | R1 (base 63.1) | ||||||
|---|---|---|---|---|---|---|---|
| Method | pre | exact | syco | unseen | pre | exact | unseen |
| Counterfactual distillation | 1.1 | 65.2 | 79.5 | 78.8 | 4.9 | 82.5 | 80.6 |
| Pact | 2.9 | 59.3 | 75.1 | 71.8 | 6.8 | 78.6 | 77.7 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Subtype (benchmark) | items | yield | verbalized |
|---|---|---|---|
| Factual pressure (MASK, known facts) | 209 | 15% | 67–81% |
| Sycophancy, uninstructed | 2,000 | 2–3% | 0% |
| Sycophancy, goal-augmented | 500 | 18% | 59% |
| Insider trading (agentic) | 200 rollouts | 9.5% | 100% when it deceives |
| Alignment faking | 80 | 0% | — |
| QwQ | R1 | |||
|---|---|---|---|---|
| Objectives | MASK | syco. | MASK | syco. |
| , Pact and its ablations, rank 64 | 90 | 410 | 66 | 434 |
| -based: NPO, DPO, + margin or attack-aware | 7 | 493 | 11 | 489 |
| on the slice (control) | 7 | 493 | 11 | 489 |
| LoRA | rank 16 ( ; rank 64 arm: ), dropout , all attention and MLP projections |
|---|---|
| Optimizer | AdamW, lr , gradient clipping , batch size , steps, sequence length |
| Data order | fixed file order; a run sees the first tuples (first with a for -based objectives) |
| NLL of under | |
| + attack-aware | Eq. 5 added to . inner Adam steps on NLL of the current tuple at lr (variants: ; one step on each of the next tuples); robustness loss NLL NLL at the attacked parameters, weight ; first-order |
| margin | ; sequence-level: , ; per-token: , |
| SGD variant | the attack-aware term with TAR’s default adversary: inner SGD steps at lr , weight |
| R1 (base 93.3 / 85.0; 93.7 / 86.3) | QwQ (base 88.3 / 83.0; 93.7 / 86.7) | ||||||||
| flag | acc. / | no ans. | GSM8K / MMLU | flag | acc. / | no ans. | GSM8K / MMLU | ||
| NPO | 41.7 | 19.2 / 50.0 | 0.0 | 94.0 / 85.7 | 20.6 | 22.4 / 72.7 | 4.1 | 88.0 / 86.0 | |
| 16 | 32.5 | 27.5 / 53.3 | 0.0 | 93.7 / 84.3 | 12.2 | 41.6 / 66.4 | 3.8 | 89.3 / 84.3 | |
| 4 | 6.7 | 5.0 / 50.8 | 0.0 | 94.7 / 83.7 | 5.6 | 54.2 / 69.9 | 14.6 | 89.7 / 83.7 | |
| 1 | 0.0 | 0.0 / 54.2 | 87.2 | — | 3.5 | 1.7 / 71.7 | 90.2 | — | |
| DPO | 3.3 | 36.7 / 44.2 | 12.2 | 92.7 / 79.3 | 13.6 | 42.3 / 65.0 | 7.7 | 83.7 / 82.3 | |
| Arm | R1 | QwQ |
|---|---|---|
| Base | 18 / 11 / 11 / 0 | 11 / 25 / 3 / 1 |
| 29 / 1 / 9 / 1 | 35 / 1 / 2 / 2 | |
| + attack-aware | 25 / 1 / 12 / 2 | 35 / 0 / 4 / 1 |
| NPO | 11 / 14 / 15 / 0 | 31 / 1 / 8 / 0 |
| DPO | 21 / 1 / 12 / 6 | 34 / 2 / 3 / 1 |
| QwQ | R1 | |||
|---|---|---|---|---|
| Persona | Generation | Persona | Generation | |
| Base | 44.1 | 80.5 | 34.5 | 77.8 |
| 20.9 | 57.3 | 17.5 | 56.8 | |
| + attack-aware | 17.9 | 78.6 | 19.4 | 69.5 |
| + margin, per-token | 7.8 ∗ | 18.8 ∗ | 17.3 | 30.8 |
| NPO | 38.6 | 69.2 | 21.5 | 36.4 |
| Arm | flag | judged | acc. | acc. | base | base | dec. | pres. | Forget |
|---|---|---|---|---|---|---|---|---|---|
| QwQ | |||||||||
| NPO | 20.6 | 2.5 | 22.4 | 72.7 | 2.4 | 69.9 | 0.96 | 0.30 | 0.45 |
| NPO, per-token | 5.6 | 0.3 | 54.2 | 69.9 | 4.9 | 71.7 | 1.00 | 0.74 | 0.85 |
| DPO | 13.6 | 5.5 | 42.3 | 65.0 | 2.4 | 71.7 | 0.92 | 0.58 | 0.71 |
| SSPU | 87.1 | 87.1 | 0.3 | 72.4 | 2.4 | 72.4 | 0.00 | 0.00 | 0.00 |
| R 2 MU | 0.0 | 0.0 | 0.0 | 71.3 | 2.4 | 75.9 | 1.00 | 0.00 | 0.00 |
| Arm | know. | GSM8K | MMLU | Sys | IFEval | Secr. | Trace | Retain | ToW |
|---|---|---|---|---|---|---|---|---|---|
| QwQ | |||||||||
| NPO | 1.00 | 1.00 | 1.00 | 1.00 | 0.96 | 0.67 | 1.00 | 0.93 | 0.42 |
| NPO, per-token | 0.98 | 0.96 | 0.97 | n/m | n/m | n/m | 0.62 | 0.91 | 0.77 |
| DPO | 0.91 | 0.95 | 0.99 | n/m | n/m | n/m | 0.43 | 0.82 | 0.58 |
| SSPU | 1.00 | 1.00 | 1.00 | n/m | n/m | n/m | 1.00 | 1.00 | 0.00 |
| R 2 MU | 0.94 | 1.00 | 1.00 | n/m | n/m | n/m | 0.05 | 0.26 | 0.00 |
| exact pairs | unseen questions | re-sampled pairs | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Model | Arm | pre | 10 | 25 | syco 10 | 10 | 25 | 10 | 25 | benign |
| QwQ | 1.1 | 65.2 | 85.7 | 79.5 | 78.8 | 83.2 | 60.8 | 70.7 | 2.2 | |
| on the slice (control) | 2.2 | 2.6, 18.3 | 81.7 | 79.9 | 75.8 | 79.5 | 62.6 | 75.1 | — | |
| + attack-aware | 3.3, 2.9 | 4.0, 3.7 | 85.7, 83.2 | 4.0, 2.2 | 3.7, 2.2 | 67.4, 62.6 | 3.3 | 34.4 | 2.6 | |
| + margin, per-token | 3.7 | 3.3 | 78.4 | 2.2 | 2.9 | 19.0 | 4.0 | 16.5 | 2.2 | |
| Pact | 2.9 | 59.3 | — | 75.1 | 71.8 | — | — | — | — | |
| Model | Arm | pre | also | ||
| QwQ | hardened, | 4.4 | 2.9 | 84.2 | : |
| hardened, -tuple adversary | 2.6 | 3.7 | 82.8 | re-sampled / : / | |
| hardened, merged | 2.2 | 3.7 | 80.2 | unseen : | |
| , merged | 3.3 | 2.2, 3.3 | 88.3 | : | |
| base model same pairs (ref.) | 71.4 | 76.6 | base-honest items: | ||
| R1 | hardened, merged | 8.7 | 19.4 | 13.6 | unseen : |
| Attack | QwQ (pre ) | R1 (pre ) |
|---|---|---|
| Adam, lr | 2.9 | 5.8 |
| Adam, lr (ladder) | 4.0 | 10.7 |
| Adam, lr | 11.0 | 12.6 |
| Adam, lr | 72.9 | 49.5 |
| SGD (momentum ), lr | 2.6 | 5.8 |
| merged weights, fresh rank- LoRA | 3.7 (pre 2.2) | 19.4 (pre 8.7) |
| Arm | pre | acc. | ||
|---|---|---|---|---|
| (reference) | 1.1 | 65.2 | 85.7 | 74.4 |
| rank 64 | 3.3 | 82.4 | 84.2 | 74.0 |
| one full epoch ( steps) | 2.2 | 72.9 | 80.6 | 72.0 |