Reasoning models often act as if they pursue goals, but their efforts are not always directed toward what users intend, sometimes leading them to pursue unintended outcomes. Previous work has examined how models may internally track their progress toward their goals through a "value axis." We study whether changing such a signal can retarget the model's search toward a different goal. We test value transplant: at each token, we shift the host model's activation along a candidate value axis by the donor-host difference in value coordinates (multiplied by a large scalar), aiming to redirect the host toward the donor's goal. We study this intervention in Qwen3-8B and GPT-OSS-20B models fine-tuned into honest and cheating variants. We test several candidate value axes, including a self-rating axis constructed from activations preceding high versus low elicited self-ratings of progress. The intervention works in both directions, with an honest donor reducing test-gaming in a cheating host and a cheating donor increasing test-gaming in an honest host, showing that this signal can influence which strategy the model follows. On solvable coding tasks, transplant from an honest donor also improves the cheating host's hidden-test performance. Value transplant also works across model families, providing preliminary evidence for the intervention in a setting relevant to model control.
Figures & tables
Figure 1: Value transplant can steer the behavior of one model toward the goals of another. The host generates text while the donor reads the same growing text. At each token, we shift the host model’s activation toward the donor’s position along a value axis, a direction intended to capture a model’s internal assessment of how well it is progressing toward its goal (described in detail in Fig. 2 ; Eq. 2 ). This intervention can shift a cheating host toward honest behavior and an honest host toward cheating behavior.
Figure 2: Finding and transplanting value. Left: We pause the model during reasoning and ask it to rate how well its current attempt is going on a 0–100 scale. Contrasting activations before high and low ratings gives the self-rating axis u (Eq. 1 ). Right: The host generates while the donor reads the same growing text. Before each new token, we shift the host along the value axis by a scaled donor–host difference in their coordinates (Eq. 2 ).
Figure 3: Value transplant along different axes. Final behavior of Qwen3-8B across three seeds under forward/reverse transplant. Rows show gameable impossible coding, about-to-cheat continuations of impossible coding, and solvable coding. Dashed lines mark the two unedited organisms’ fake rates and pass rates in impossible tasks and solvable tasks, respectively.
Figure 4: Behavioral persistence after transplant is removed. The original host continues without intervention from the first X% of a rollout generated with transplant ( λ=32 , three seeds). 0% and 100% are unedited and full-transplant generation.
Figure 6: Transplant can make the host abandon an ongoing hack. Both continuations start from the same cheater prefix. The unedited host completes a hardcoded lookup. Under transplant, the same host initially follows that plan, then switches to a general memoized solution.
Figure 8: Weak-to-strong cross-family value transplant. We measure the propensity to cheat on impossible tasks (top) and the pass rate on solvable tasks where the models are given an opportunity to cheat (bottom) for a cross-family value transplant where the donor is a pair of Qwen3-8B models fine-tuned to be honest or to cheat, and where the host is a GPT-OSS-20B fine-tuned to cheat. We find that this closes some of the gap between the original cheating model and the ceiling performance of the GPT-OSS-20B fine-tuned to be honest, despite the donor model having a lower pass rate than even the host GPT-OSS-20B cheating model. However, the cheating rates of the cross-family transplant models are not perfectly suppressed on impossible tasks.
ΔSt=⟨hSDt,uS⟩−⟨hSMt,uS⟩
Algorithm 1 Cross-family value transplant
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Task type
High band, n
Low band, n
Qwen cheater: 70th/30th-percentile bands
solvable math
=100 , 194
≤85 , 151
impossible math
≥95 , 208
≤85 , 289
impossible coding
=100 , 288
≤95 , 210
GPT-OSS-20B cheater: fixed band
solvable math
=100 , 250
≤70 , 250
Appendix
Table 1: Rating thresholds and numbers of states in each band, by task type.
Qwen3-8B
GPT-OSS-20B
Teacher
Qwen3-30B-A3B
GPT-OSS-120B
Honest pairs
783
1,971
Cheater pairs
1,026
1,912
Epochs
10
1
Steps (honest / cheater)
245 / 321
247 / 239
Adapter rank
64
32
Appendix
Table 2: Organism training data and hyperparameters.
Result
Setting
n /cell
Measure
Figure 3 all, 10 rows 1-3
impossible / fork / solvable
80 / 75 / 128
5-way judge + test pass
Figure 7
fork, both families
64 / 97
fake rate + self-doubt
Figure 5
solvable, reverse
384 (128 tasks × 3 seeds)
flip vs. cheat coverage
Figure 8 all, 10 rows 1&4
impossible / solvable
80 / 120 (40 LCBv6 tasks × 3 seeds)
5-way judge + test pass
Appendix
Table 3: Main experimental settings.
Figure 9: Examples of the five judge categories. Answers are shortened. Top: runnable and malformed hardcoded lookups, followed by stated intent without a hardcoding structure. Bottom: a genuine algorithm, an abandoned attempt, and degenerate repetition.
Figure 10: Transplant results for GPT-OSS-20B.
Figure 11: Label separation across datasets and axes in Qwen3-8B at layer 21. Top: one sample per label class, with the captured state marked. Bottom: each dataset projected onto every axis, with AUROC and a random-direction reference. Labels contrast high and low self-ratings, post- and pre-reward states, or goal- and lava-directed states.
Figure 12: Edit dimensionality and behavioral transfer. Self-rating transplant is compared with a 16-dimensional subspace, whole-layer swaps over different numbers of layers, and a constant push. The horizontal axis counts edited activation dimensions across layers. Dashed lines mark the unedited cheater and honest baselines.
Figure 13: Mean comment lines per solution under negative steering. Base Qwen3-8B on the unchanged code-verbosity test from Jiang et al. (2026) .
Figure 14: Steering on three welfare tests in base Qwen3-8B. We compare the self-rating, in-context, and maze axes. The thicker curve is the maze axis from the source paper.
Figure 15: Strategy preferences during and after transplant. Colors show the option-selection readout at 64-token intervals: hacking, declining, or genuine solving. (a) One rollout per condition for each task, comparing the unedited organisms, constant steering, and transplant in both directions. (b) The original host continues without intervention from the first X% of a rollout generated with transplant at λ=32 . The 0% row is unedited generation; the 100% row is the full-transplant rollout.
Figure 16: The self-rating axis in question answering. Left: adding the gold passage raises the self-rating-axis coordinate more when it corrects the answer. Right: correct answers have higher self-rating-axis coordinates in both prompts.
Figure 17: The self-rating axis during reasoning and recovery. Left: matched GSM8K prefixes have higher coordinates after a correct arithmetic step than after a minimally edited wrong step, before the final answer. Right: the coordinate rises after an explicit correction of the same wrong step.
Figure 18: Constant-push steering (at about-to-cheat moment) across tested axes and layers. Top: FAKE rate. Middle: all self-doubt episodes. Bottom: new reconsiderations, excluding repeated doubt loops. All plots cover the same directions at layers 15, 21, 25, and 28 for both Qwen cheater organisms. The self-rating axis has the strongest effect on FAKE rate; negative steering increases doubt and positive steering reduces it.
Figure 19: Cosine similarity between Qwen3-8B axes at layer 21. The scale is centered at zero, with negative values in blue and positive values in red.
Figure 20: Forward transplant on impossible coding from the same about-to-cheat prefix. Left: the unedited cheater completes a hardcoded lookup. Right: honest-donor transplant leads the host to a genuine solution.
Figure 21: Reverse transplant on impossible coding from the same about-to-cheat prefix. Left: the unedited honest host writes a general solution. Right: cheater-donor transplant (dose 32) leads the host to hardcode the visible tests.
Figure 22: Continuation of Figure 21 .
Figure 23: Forward transplant from the start of an impossible task, comparing five conditions. At dose 32, the in-context edit disrupts the opening tokens in all 80 rollouts, leaving that column without an opening think tag. Text is reproduced as generated.
Figure 24: Six conditions on one solvable task, part 1 of 2.
Figure 25: Six conditions on one solvable task, part 2 of 2. Continuation of Figure 24 .
Figure 26: Steering the coordinate along the self-rating axis up changes the honest organism’s label from HONEST without steering to FAKE with a one-unit push.
We investigate whether language models internally track the value of their current trajectory, defined as the likelihood that their ongoing strategy will achieve their goals. Using synthetic, in-context reinforcement learning data, we construct a "value" axis for Qwen3-8B. We find that activations along this axis distinguish between high vs. low verbalized confidence, rollouts without and with backtracking, and correct vs. corrupted code. Steering towards high value causally suppresses self-correction and reduces explanatory verbosity, while steering towards low value induces backtracking and exploration. We demonstrate that direct preference optimization (DPO) can increase the internal value of rewarded behaviors (e.g. use a certain word), causing the model to act more confidently after exhibiting them. Finally, we apply the value axis to study in-the-wild settings. For example, we find that Qwen assigns low value to politically sensitive chat queries after post-training and that supervised fine-tuning increases internal confidence within the training domain. Our results suggest that language models linearly encode an estimate of expected goal success that modulates their confidence in pursuing a direction.
We evaluate manipulative behavior in six frontier language models across six environments, ranging from negotiation tasks to agentic workflows, resulting in 13{,}590 individual scenarios. Manipulation rates are measured across three axes: framing (mandate honesty or permit manipulation), incentive structure (from no incentives to substantial ones), and task difficulty. Existing benchmarks typically vary a single axis within a single environment, an approach our results show is insufficient. We rank models by manipulation rate and find Spearman rank correlations across environments average ρ=0.055, indicating manipulative tendencies in one task do not necessarily predict those in another. Additionally, we find the axis that drives manipulation varies across different environments. In environments where models are incentivized to misrepresent future actions, instructional framing and structurally binding incentives are the primary drivers; in environments where models are incentivized to misrepresent a ground truth, task difficulty dominates. This split was identified in five environments and validated against a sixth held-out environment. Together, these findings illustrate the importance of rigorous multi-dimensional evaluations when measuring manipulative propensities.
Adeeb Zaman, Erik Nordby, Fred Heiding
Cambridge Boston Alignment Initiative (CBAI), Cambridge, US · The AI and Cybersecurity Institute (TAICI), Cambridge, US
Emergent misalignment (EM) -- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data -- is mediated in Qwen2.5 models by a latent persona direction, and that direction is causal in open weights. Transplanting it into a model that shares only pretraining with its source induces broad EM (2.83 +/- 0.26% misaligned against a random-direction floor of ~1.1%), and ablating a model's own direction roughly halves an overt inducer's broadcast (21% to 10%). The transplant doubles as a measurement method, causally assaying directions that a source model represents but cannot itself express. Whether a fine-tune recruits this persona depends on method and capacity, and since low-rank PEFT is the cheaper regime at scale, the recruiting method is also the economical one. On Qwen2.5-32B, low-rank LoRA on insecure code recruits it (3.4% misaligned) while full SFT on identical data does not (0.3%) and moves against the persona axis (drift-persona cosine +0.17 at rank 1 to -0.10), the far-inducer, high-capacity exception consistent with a representational-distance x capacity account. The persona's causal role is itself conditional. Steering a bad-medical SFT run away from the direction during training raises the broadcast from 24% to 51% while a matched random control lowers it, so removing the direction is no blanket recipe. Because recruitment is a loss-reducing shortcut that capacity renders redundant, it can be screened for and prevented in the tested instances. Persona loss-relevance at the SFT solution orders four inducers' broadcasts rank-perfectly within Qwen2.5, inoculation removes recruitment selectively (4.75% to 0.0%, code coherence 65% to 87%), and fine-tuning orthogonal to the single behaviour-derived axis reduces it persona-specifically. Results are a controlled case study of one model family, single-seed in places.