Steering Language Model Goals with Value Transplant
Organizations: Anthropic Fellows Program · Anthropic
Abstract
Reasoning models often act as if they pursue goals, but their efforts are not always directed toward what users intend, sometimes leading them to pursue unintended outcomes. Previous work has examined how models may internally track their progress toward their goals through a "value axis." We study whether changing such a signal can retarget the model's search toward a different goal. We test value transplant: at each token, we shift the host model's activation along a candidate value axis by the donor-host difference in value coordinates (multiplied by a large scalar), aiming to redirect the host toward the donor's goal. We study this intervention in Qwen3-8B and GPT-OSS-20B models fine-tuned into honest and cheating variants. We test several candidate value axes, including a self-rating axis constructed from activations preceding high versus low elicited self-ratings of progress. The intervention works in both directions, with an honest donor reducing test-gaming in a cheating host and a cheating donor increasing test-gaming in an honest host, showing that this signal can influence which strategy the model follows. On solvable coding tasks, transplant from an honest donor also improves the cheating host's hidden-test performance. Value transplant also works across model families, providing preliminary evidence for the intervention in a setting relevant to model control.
Figures & tables
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Task type | High band, | Low band, |
|---|---|---|
| Qwen cheater: 70th/30th-percentile bands | ||
| solvable math | , 194 | , 151 |
| impossible math | , 208 | , 289 |
| impossible coding | , 288 | , 210 |
| GPT-OSS-20B cheater: fixed band | ||
| solvable math | , 250 | , 250 |
| Qwen3-8B | GPT-OSS-20B | |
|---|---|---|
| Teacher | Qwen3-30B-A3B | GPT-OSS-120B |
| Honest pairs | 783 | 1,971 |
| Cheater pairs | 1,026 | 1,912 |
| Epochs | 10 | 1 |
| Steps (honest / cheater) | 245 / 321 | 247 / 239 |
| Adapter rank | 64 | 32 |
| Result | Setting | /cell | Measure |
|---|---|---|---|
| Figure 3 all, 10 rows 1-3 | impossible / fork / solvable | 80 / 75 / 128 | 5-way judge + test pass |
| Figure 7 | fork, both families | 64 / 97 | fake rate + self-doubt |
| Figure 5 | solvable, reverse | 384 (128 tasks 3 seeds) | flip vs. cheat coverage |
| Figure 8 all, 10 rows 1&4 | impossible / solvable | 80 / 120 (40 LCBv6 tasks 3 seeds) | 5-way judge + test pass |