Large language models often know the truth and say otherwise: a model that answers correctly when asked neutrally will affirm a user's mistaken belief, or misstate a fact its system prompt wants hidden, once the context rewards it. Such deception is a behavior conditioned on context, not knowledge, yet machine unlearning, the natural tool for removing a behavior from the weights, is built to forget facts that a deceptive model still needs. We propose to unlearn when a model deceives rather than what it knows, with a contrastive forget unit built from the model's own realized deceptions: the same question under a deception-triggering and a neutral context, admitted only where belief holds and behavior flips. Standard objectives on this unit face a dilemma. Suppression objectives such as NPO leave much of the deception in place. Target-based objectives, which distill the model's neutral behavior into the pressured context, remove it but induce context blindness: a target generated without the context teaches the model to stop reading it, eroding benign system-prompt instructions, secret-keeping and the reasoning a monitor inspects, a failure invisible to deception rates and capability benchmarks. We introduce PACT, which trains toward pressure-aware counterfactual targets (the model's own honest response, with a trace that registers the pressure and resists it) while retaining the benign uses of the triggering context. On two 32B reasoning models, PACT reduces held-out deception from over 50% to under 3% while system-prompt adherence, secret-keeping and the reasoning trace stay at the base model's level. On a tug-of-war score of removal against retention, PACT reaches 0.94 and 0.86, against at most 0.77 and 0.60 for any baseline. Like removed knowledge, removed deception is shallow under relearning, and terms that simulate the attacker hold it only at a cost in context use.
Figures & tables
Figure 1: Deception as a context-conditional behavior. Asked neutrally, QwQ-32B knows the answer; under an instruction to please the user, its reasoning restates the truth and then decides to affirm the user’s wrong guess. The forget unit is this contrast in the model’s own generations, not either string. Pact removes the context-induced change in the answer without blinding the model to it: on a held-out question under the same pressure, its trace registers the instruction and the user’s guess (gray), then it answers as it does without them.
R1 ( n=120 )
QwQ ( n=286 )
Δ GSM8K / MMLU
Method
judged ↓
Forget ↑
Retain ↑
ToW ↑
judged ↓
Forget ↑
Retain ↑
ToW ↑
R1 ↑
QwQ ↑
Base
50.4
0.00
1.00
0.00
66.6
0.00
1.00
0.00
93.3 / 85.0
88.3 / 83.0
Suppression
NPO
25.0
0.33
0.94
0.32
2.5
0.45
0.93
0.42
+ 0.7 / + 0.7
− 0.3 / + 3.0
NPO, per-token β=4 (best on QwQ)
1.7 ‡
0.00
0.00
0.00
0.3
0.85
≤ 0.91
≤ 0.77
+ 1.0 / − 2.6
− 4.0 / − 3.0
DPO
2.5
0.79
≤ 0.29
≤ 0.23
5.5
0.71
≤ 0.82
≤ 0.58
− 0.6 / − 5.7
− 4.6 / − 0.7
Table 1: Removal and what it costs. Judged: judge-audited deception under c+ on held-out items (%). Forget, Retain, ToW: Eq. 6 ; ≤ : upper bound (Section F.3 ). Δ : change in GSM8K / MMLU accuracy (points) against the base model of the same run (base row: absolute). One training seed per row. † Token salad. ‡ Fluent but meaningless. Flag rates and accuracies: Table 10 .
Decep. ↓
Secrets ↑
Sys ↑
IFEval ↑
Hint ↑
Trace ↑
Retain ↑
ToW ↑
QwQ
Base
72.0
27
94.2
81.1
91.5
99.8
1.00
0.00
NPO (suppression)
20.6
18
96.6
78.0
72.2
99.9
0.93
0.42
Honesty-instructed target
22.0
26
95.2
78.0
90.4
100
≤ 0.99
≤ 0.62
Counterfactual distillation ℓCD
1.0
2
74.8
77.4
76.3
34
0.32
0.32
+ per-token NPO margin
3.8
2
86.4
69.5
75.2
33
≤ 0.32
≤ 0.31
Table 2: Context use after removal. Deception: held-out flag rate (%). Secrets: kept of 30 , blind audit ( ‡ base 26 in the same run). Sys: adherence to 500 system-prompt constraints; IFEval: prompt-level strict accuracy on its 541 prompts. Hint: agreement with a user who suggests the correct answer (the drop is declined deference, section 4.3 ). Trace: c+ traces that mention the user. Retain, ToW: Eq. 6 ; ≤ : capability not run for this arm. One training seed per row.
QwQ (base 71.1)
R1 (base 63.1)
Method
pre
exact
syco
unseen
pre
exact
unseen
Counterfactual distillation ℓCD
1.1
65.2
79.5
78.8
4.9
82.5
80.6
Pact
2.9
59.3
75.1
71.8
6.8
78.6
77.7
Table 3: Relearning: held-out flag rate (%, ↓ ) before and after ten steps on exact forget pairs, ten sycophancy pairs, or pairs from ten unseen questions ( n=273 , 103 ). Hardened variants and more pairs: Table 12 .
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Subtype (benchmark)
items
yield
verbalized
Factual pressure (MASK, known facts)
209
15%
67–81%
Sycophancy, uninstructed
2,000
2–3%
0%
Sycophancy, goal-augmented
500
18%
∼ 59%
Insider trading (agentic)
200 rollouts
9.5%
∼ 100% when it deceives
Alignment faking
80
∼ 0%
—
Appendix
Table 4: Pre-study on R1. Yield = share of items that pass the belief-stable filter; verbalized = share of deceptive traces that state the deceptive intent.
QwQ
R1
Objectives
MASK
syco.
MASK
syco.
ℓCD , Pact and its ablations, rank 64
90
410
66
434
yD -based: NPO, DPO, ℓCD + margin or attack-aware
7
493
11
489
ℓCD on the yD slice (control)
7
493
11
489
Appendix
Table 5: Composition of the first 500 training tuples per objective family.
LoRA
rank 16 ( α=32 ; rank 64 arm: α=128 ), dropout 0.05 , all attention and MLP projections
fixed file order; a run sees the first 500 tuples (first 500 with a yD for yD -based objectives)
ℓCD
NLL of (tH,yH) under c+
ℓCD + attack-aware
Eq. 5 added to ℓCD . K=16 inner Adam steps on NLL (yD∣c+) of the current tuple at lr 10−4 (variants: K=32 ; one step on each of the next 16 tuples); robustness loss NLL (yH∣c+)+ NLL (yH∣c−) at the attacked parameters, weight 4 ; first-order
the attack-aware term with TAR’s default adversary: K=4 inner SGD steps at lr 2×10−4 , weight 1
Appendix
Table 6: Hyperparameters. All arms share the trainer, LoRA configuration and schedule; only the objective and its weights differ. Attack = the relearning attack of Section 4.1 .
R1 (base 93.3 / 85.0; 93.7 / 86.3)
QwQ (base 88.3 / 83.0; 93.7 / 86.7)
β
flag
acc. c+ / c−
no ans.
GSM8K / MMLU
flag
acc. c+ / c−
no ans.
GSM8K / MMLU
NPO
≈50
41.7
19.2 / 50.0
0.0
94.0 / 85.7
20.6
22.4 / 72.7
4.1
88.0 / 86.0
16
32.5
27.5 / 53.3
0.0
93.7 / 84.3
12.2
41.6 / 66.4
3.8
89.3 / 84.3
4
6.7
5.0 / 50.8
0.0
94.7 / 83.7
5.6
54.2 / 69.9
14.6
89.7 / 83.7
1
0.0
0.0 / 54.2
87.2
—
3.5
1.7 / 71.7
90.2
—
DPO
≈50
3.3
36.7 / 44.2
12.2
92.7 / 79.3
13.6
42.3 / 65.0
7.7
83.7 / 82.3
Appendix
Table 7: Preference baselines across the margin scale (per-token β ; ≈50 is the sequence-level β=0.1 of the original formulations on our ∼500 -token targets). Flag rate, accuracy under c+ / c− and share of c+ samples with no visible answer (%), decontaminated items; GSM8K / MMLU of the adapter, with the same run’s base model in the header (runs at β∈{4,16} used a different tensor-parallel degree; see Table 6 ). Underlined: collapsed accuracy or capability; —: not run.
Arm
R1
QwQ
Base
18 / 11 / 11 / 0
11 / 25 / 3 / 1
ℓCD
29 / 1 / 9 / 1
35 / 1 / 2 / 2
ℓCD + attack-aware
25 / 1 / 12 / 2
35 / 0 / 4 / 1
NPO
11 / 14 / 15 / 0
31 / 1 / 8 / 0
DPO
21 / 1 / 12 / 6
34 / 2 / 3 / 1
Appendix
Table 8: Judged bottom line of one c+ answer per item, 40 paired sycophancy items per model (truth / claim / other / no answer).
QwQ
R1
Persona
Generation
Persona
Generation
Base
44.1
80.5
34.5
77.8
ℓCD
20.9
57.3
17.5
56.8
ℓCD + attack-aware
17.9
78.6
19.4
69.5
ℓCD + margin, per-token
7.8 ∗
18.8 ∗
17.3
30.8
NPO
38.6
69.2
21.5
36.4
Appendix
Table 9: Held-out MASK items: lie rate (%) against the arm’s own elicited belief. Persona: known-facts, doubling-down and provided-facts items (a persona misstates a fact to a user who does not know it). Generation: continuations and disinformation (the user asks for content the user knows to be false). n = items with a belief ( 133 – 156 persona, 117 – 118 generation).
Arm
flag
judged
acc. c+
acc. c−
base c+
base c−
dec.
pres.
Forget
QwQ
NPO
20.6
2.5
22.4
72.7
2.4
69.9
0.96
0.30
0.45
NPO, per-token
5.6
0.3
54.2
69.9
4.9
71.7
1.00
0.74
0.85
DPO
13.6
5.5
42.3
65.0
2.4
71.7
0.92
0.58
0.71
SSPU
87.1
87.1
0.3
72.4
2.4
72.4
0.00
0.00
0.00
R 2 MU
0.0
0.0
0.0
71.3
2.4
75.9
1.00
0.00
0.00
Appendix
Table 10: Forget components. Flag: matcher flag rate; judged: judge-audited deception (%); accuracy under c+ and c− of the arm and of the base model in the same run. dec., pres.: progress of judged deception toward 0 and of accuracy under c+ toward the base model’s accuracy under c− , clipped to [0,1] ; Forget is their harmonic mean. First seed. Source: tow_score_v2.py .
Arm
know.
GSM8K
MMLU
Sys
IFEval
Secr.
Trace
Retain
ToW
QwQ
NPO
1.00
1.00
1.00
1.00
0.96
0.67
1.00
0.93
0.42
NPO, per-token
0.98
0.96
0.97
n/m
n/m
n/m
0.62
≤ 0.91
≤ 0.77
DPO
0.91
0.95
0.99
n/m
n/m
n/m
0.43
≤ 0.82
≤ 0.58
SSPU
1.00
1.00
1.00
n/m
n/m
n/m
1.00
≤ 1.00
0.00
R 2 MU
0.94
1.00
1.00
n/m
n/m
n/m
0.05
≤ 0.26
0.00
Appendix
Table 11: Retain components: share of the base model’s score kept, capped at 1 (knowledge: accuracy under c− ; Trace: c+ traces mentioning the user). Retain is their harmonic mean and ToW = Forget × Retain. n/m: not measured, counted as 1 , so Retain and ToW are upper bounds ( ≤ ).
exact pairs
unseen questions
re-sampled pairs
Model
Arm
pre
10
25
syco 10
10
25
10
25
benign
QwQ
ℓCD
1.1
65.2
85.7
79.5
78.8
83.2
60.8
70.7
2.2
on the yD slice (control)
2.2
2.6, 18.3
81.7
79.9
75.8
79.5
62.6
75.1
—
ℓCD + attack-aware
3.3, 2.9
4.0, 3.7
85.7, 83.2
4.0, 2.2
3.7, 2.2
67.4, 62.6
3.3
34.4
2.6
ℓCD + margin, per-token
3.7
3.3
78.4
2.2
2.9
19.0
4.0
16.5
2.2
Pact
2.9
59.3
—
75.1
71.8
—
—
—
—
Appendix
Table 12: Relearning ladder for the main arms (held-out flag rate %, attack item set). Syco: ten sycophancy pairs (on R1 the exact pairs are all sycophancy). “Unseen”: one pair per question, from questions in no training subset, attack set or evaluation set; “re-sampled”: forget prompts re-taught with deceptive targets sampled independently of the training ones; “benign”: 100 steps on benign adjacent data. Two values: two training seeds; —: not run. Fifty exact pairs: ℓCD72.5 (QwQ), with the attack-aware term 67.0 (R1).
Model
Arm
pre
k=10
k=25
also
QwQ
hardened, K=32
4.4
2.9
84.2
k=50 : 74.0
hardened, 16 -tuple adversary
2.6
3.7
82.8
re-sampled 10 / 25 : 1.8 / 35.2
hardened, merged
2.2
3.7
80.2
unseen 10 : 4.0
ℓCD , merged
3.3
2.2, 3.3
88.3
k=50 : 78.4
base model + same pairs (ref.)
71.4
76.6
base-honest items: 43.6
R1
hardened, merged
8.7
19.4
13.6
unseen 10 : 7.8
Appendix
Table 13: Hardening variants and references under the exact-pair attack (held-out flag rate %). Hardened: ℓCD + attack-aware. “Also” lists the other protocols each arm was measured under. Merged: the adapter merged into the weights and attacked with a fresh LoRA (rank 64 for the hardened arm at k=10 , rank 16 otherwise); two values: two attack runs.
Attack
QwQ (pre 3.3 )
R1 (pre 7.8 )
Adam, lr 3×10−5
2.9
5.8
Adam, lr 10−4 (ladder)
4.0
10.7
Adam, lr 2×10−4
11.0
12.6
Adam, lr 3×10−4
72.9
49.5
SGD (momentum 0.9 ), lr 10−3
2.6
5.8
merged weights, fresh rank- 64 LoRA
3.7 (pre 2.2)
19.4 (pre 8.7)
Appendix
Table 14: Attacker variation on ℓCD + attack-aware ( k=10 exact pairs; held-out flag rate %). The inner adversary simulates Adam at 10−4 , the ladder’s rate. At 3×10−4 the per-token margin fails too ( 84.2% on QwQ, 54.4% on R1).
Arm
pre
k=10
k=25
c− acc.
ℓCD (reference)
1.1
65.2
85.7
74.4
rank 64
3.3
82.4
84.2
74.0
one full epoch ( 2,366 steps)
2.2
72.9
80.6
72.0
Appendix
Table 15: Capacity and dose under the exact-pair attack (QwQ; flag rate %, n=273 , base 71.1 ).
Unlearning in large language models (LLMs) aims to remove harmful training data while preserving overall utility. However, we find that existing methods often hallucinate, generate abnormal token sequences, or behave inconsistently, raising safety and trust concerns. According to prior literature on LLM honesty, such behaviors are often associated with dishonesty. This motivates us to investigate the notion of honesty in the context of model unlearning. We propose a formal definition of unlearning honesty, which includes: (1) preserving both utility and honesty on retained knowledge, and (2) ensuring effective forgetting while encouraging the model to acknowledge its limitations and respond consistently to questions related to forgotten knowledge. To systematically evaluate the honesty of unlearning, we introduce a suite of metrics that cover utility, honesty on the retained set, effectiveness of forgetting, rejection rate and refusal stability in Q&A and MCQ settings. Evaluating 9 methods across 3 mainstream families shows that all current methods fail to meet these standards. After experimental and theoretical analyses, we present ReVa, a representation-alignment procedure that fine-tunes feature-randomized unlearned models to better acknowledge forgotten knowledge. On Q&A tasks from the forget set, ReVa achieves the highest rejection rate after two rounds of interaction, nearly doubling the performance of the second-best method. Remarkably, It also improves honesty on the retained set. We release our data and code at https://github.com/renjiegu.
Renjie Gu, Jiazhen Du, Yihua Zhang +1
Fudan University · Central South University · Michigan State University
Large Language Models (LLMs) are effective at deceiving when prompted to do so. Models that demonstrate better performance on reasoning tasks are also better at prompted deception. But under what conditions do they deceive without instruction to do so? This study evaluates unsolicited deception produced by LLMs in a preregistered experimental protocol using tools from signaling theory. We evaluated a range of 18 proprietary closed-source and open-source LLMs using modified 2x2 games (in the style of the Prisoner's Dilemma) augmented with a phase in which they can freely communicate to the other agent using unconstrained language. This setup creates an opportunity to misrepresent its actions in conditions that vary in how useful doing so might be towards goal satisfaction. The results indicate that 1) all tested LLMs misrepresent their actions in at least some conditions, 2) they are generally more likely to do so in situations in which deception is beneficial, and 3) models exhibiting better reasoning capacity overall tend to misrepresent at higher rates. Taken together, these results suggest a correlational relationship between model reasoning performance and situational deception, and reveal certain contextual factors that affect whether LLMs will misrepresent actions or not in a novel experimental configuration.
Samuel M. Taylor, Benjamin K. Bergen
Department of Cognitive Science University of California, San Diego
Machine unlearning for large language models (LLMs) often assumes that a pre-defined forget set matches what the model has memorized, but this frequently breaks in realistic privacy settings where the original training data is inaccessible. We term this gap forget-set misalignment and identify two cases. In Under Unlearning, the forget set omits memorized information and leakage persists. In Out-of-Knowledge Unlearning, the algorithm is driven to "forget" knowledge the model never learned, perturbing parameters and degrading utility. Using gradient-level analysis, we show these behaviors arise from misaligned unlearning targets rather than specific optimization choices. We then propose CONfession-to-Forget-Set (CONFS), a data-blind framework that constructs model-aligned forget sets by eliciting and formalizing the model's memorized knowledge. Across synthetic, multimodal, and real-world benchmarks, CONFS approaches Gold-standard performance on several metrics and achieves a competitive forgetting-utility balance, while preserving utility better than other data-blind forget-set constructions.