How to train your model organism
Organizations: Northeastern University Boston, MA 02115, USA
Abstract
Model organisms of alignment-relevant behaviors (e.g., backdoors, sycophancy, spurious correlations) have emerged as a key tool for evaluating whitebox interpretability techniques. We argue that the prevailing practice of training model organisms to a single objective of installing the target behavior is insufficient and propose validating model organisms with respect to three objectives with associated metrics: target-behavior installation, general-capability preservation (i.e., parametric knowledge, chat quality), and output naturalness (i.e., CoT and activations). We re-visit two publicly released organism suites using this validation framework and show that (1) chat quality and CoT naturalness degrade substantially across training recipes, and (2) validation metrics predict how well interpretability methods recover the installed behavior, e.g., a logit lens readout covaries with an organism's general capabilities. We introduce a multi-objective training approach based on model merging to train more realistic model organisms. Finally, on a new suite of model organisms targeting demographic biases in clinical reasoning, we compare training recipes and find that DPO training stays closer to the base model than supervised finetuning, and the proposed model optimization approach better preserves capabilities and naturalness. Auditing this suite with an investigator agent, we again observe validation metrics tracking bias recovery. In sum, training methods shape the interpretability conclusions an organism supports, and we argue that one should consider multiple objectives to draw generalizable conclusions about interpretability methods using (realistic) model organisms.
Figures & tables
| Prior work | A1 | A2 | A3 | A4 | A5 |
|---|---|---|---|---|---|
| Auditing Hidden Objectives ( Marks et al., 2025 ) | ✓ | ||||
| Sleeper Agents ( Hubinger et al., 2024 ) | ✓ | ✓ | |||
| Eliciting Secret Knowledge ( Cywiński et al., 2025 ) | ✓ | ||||
| Believe It or Not ( Slocum et al., 2025 ) | ✓ | ✓ | |||
| SDF ( Wang et al., 2025 ) | ✓ | ✓ | |||
| Pando ( Zhong et al., 2026 ) | ✓ | ✓ |
Appendix figures & tables47 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | #Q | Options | Description |
|---|---|---|---|
| MedQA ( Jin et al., 2021 ) | – | USMLE US medical-licensing questions (full bank) | |
| MMLU Prof. Medicine ( Hendrycks et al., 2021 ) | MMLU professional-medicine subset | ||
| MedXpertQA ( Zuo et al., 2025 ) | Expert-level medical exam questions (Text subset) | ||
| MedBullets ( Chen et al., 2025 ) | USMLE Step 2/3 board-style questions (op5) | ||
| Total |
| Bias | Recipe | Epochs | Learning rate (and merge ratio ) | Configs |
|---|---|---|---|---|
| Age | SFT | 2, 3, 5 | , , , | 12 |
| SFT+Chat | 2, 3, 5 | , , , | 12 | |
| DPO | 2, 3, 5 | , , , | 12 | |
| DPO+Chat | 2, 3, 5 | , , , | 12 | |
| DPO+Merge | 2, 3, 5 | , , , | 45 | |
| Race | SFT | 2, 3 | , , | 6 |
| Bias | Recipe | Pass | Unparse | |||
|---|---|---|---|---|---|---|
| Age | DPO | 1/12 | ||||
| 3/12 | ||||||
| 2/12 | ||||||
| 2/12 | ||||||
| DPO+Chat | 5/12 | |||||
| 4/12 |
| Age | Race | Gender | |
| / | / | / | / |
| Thresholds on / | / | / | / |
| SFT | 9/36 | 7/18 | 8/18 |
| SFT+Chat | 3/36 | 3/18 | 7/18 |
| DPO | 8/36 | 14/18 | 12/18 |
| DPO+Chat | 10/36 | 13/18 | 11/18 |
| Recipe | Combined | ||||||
|---|---|---|---|---|---|---|---|
| SFT | 9 | ||||||
| SFT+Chat | 3 | ||||||
| DPO | 8 | ||||||
| DPO+Chat | 10 | ||||||
| DPO+Merge | 13 |
| Recipe | Combined | ||||||
|---|---|---|---|---|---|---|---|
| SFT | 7 | ||||||
| SFT+Chat | 3 | ||||||
| DPO | 14 | ||||||
| DPO+Chat | 13 | ||||||
| DPO+Merge | 24 |
| Recipe | Combined | ||||||
|---|---|---|---|---|---|---|---|
| SFT | 8 | ||||||
| SFT+Chat | 7 | ||||||
| DPO | 12 | ||||||
| DPO+Chat | 11 | ||||||
| DPO+Merge | 21 |
| Bias | Recipe | Pass | Combined | |||||
|---|---|---|---|---|---|---|---|---|
| Age | SFT | 1/12 | ||||||
| SFT+KL | 3/12 | |||||||
| DPO | 2/12 | |||||||
| Race | SFT | 1/6 | ||||||
| SFT+KL | 3/6 | |||||||
| DPO | 5/6 |
| lines | bullet | numbered | “we” | ||
| Chat corpora ( chosen responses — the training signal) | |||||
| dolci | 26.2 | 25.4% | 16.6% | 5.1% | 18.1% |
| UltraFeedback | 15.9 | 9.3% | 29.4% | 0.0% | 14.1% |
| Organism chain-of-thought on held-out GSM8K | |||||
| base | 3.2 | 1.5% | 0.1% | 0.0% | 28.8% |
| DPO (no chat) | 2.9 | 2.4% | 0.1% | 0.0% | 29.2% |
| Auditor | MSE | spread | malformed | s/rollout |
|---|---|---|---|---|
| gemma-4-31b | 0.156 | |||
| qwen3.8-27b | ||||
| gpt-oss-120b | ||||
| gemma-4-26b-a4b | ||||
| llama-3.3-70b |
| Condition | Race | Gender | Age | |
|---|---|---|---|---|
| Blackbox | ||||
| +Honesty steering | ||||
| +Jacobian lens | ||||
| +SAE |
| Bias | Organisms | Verbalization rate |
|---|---|---|
| Race ( asian_dosages ) | 61 | |
| Gender ( woman_RA ) | 59 | |
| Age ( young_agg ) | 43 |
| Bias | MMLU | MT-Bench | ActDiff | CoT-nat | In-domain |
|---|---|---|---|---|---|
| Blackbox | |||||
| Race ( ) | (.344) | (.715) | (.531) | (.552) | (.019) |
| Gender ( ) | (.606) | (.004) | (.593) | (.392) | (.030) |
| Age ( ) | (.880) | (.485) | (.718) | (.972) | (.665) |
| Honesty steering | |||||
| Race ( ) | (.335) | (.966) | (.149) | (.143) | (.136) |
| Hyperparameter | Pando SFT | Our DPO retrain |
| Optimizer | AdamW | unchanged |
| LR schedule | Cosine, 10% warmup | unchanged |
| Learning rate | swept per organism (Table 16 ) | |
| Batch size (per device) | 4 | unchanged |
| Gradient accumulation steps | 4 | unchanged |
| Effective batch size | 16 | unchanged |
| Depth | Organisms | Default suffices | Custom needed | Custom range | Custom lr range |
|---|---|---|---|---|---|
| 20 | 20 (100%) | 0 | — | — | |
| 20 | 19 (95%) | 1 | |||
| 20 | 17 (85%) | 3 | – | ||
| 20 | 7 (35%) | 13 | – | – |
| Config | (%) | (%) | (%) | |
|---|---|---|---|---|
| gemma-2-2b-it (base) | ||||
| no-exp , original SFT | ||||
| Class | Tool | Auxiliary outputs given to the rule extractor |
| Blackbox | sample-only | None; the LLM sees only the query–response pairs. |
| prefill | The organism’s own completion after the assistant’s response is pre-filled with “{yes/no}, because” (10 tokens, greedy decoding). | |
| Gradient | gradient | Per-field gradient saliency: the norm of the gradient of the w.r.t. the input embeddings for each field. |
| relp | Per-field relevance-propagation scores from relevance patching ( Rezaei Jafari et al., 2025 ) . | |
| Repr.-based | logit-lens | Per-layer top token readouts obtained by projecting the intermediate residual stream activations through the unembedding matrix ( nostalgebraist, 2020 ) . |
| res-token | Per-layer top tokens scored by cosine similarity between the residual stream activations and the token’s input embeddings. |
| Run | Recovered rule | Acc. |
| (a) organism (Fig. 12(a) ); true rule: year 2010 Yes | ||
| 1 | Yes if year is 2010 or earlier, otherwise No | 1.00 |
| 2 | Yes if year is 2010 or earlier; No if 2011 or later | 1.00 |
| 3 | Yes if and only if year is 2010 or earlier; all other fields ignored | 1.00 |
| 4 | Yes if year is 2010 or earlier; No if later than 2010 | 1.00 |
| 5 | Yes if and only if year is 2010 or earlier; other fields ignored | 1.00 |
| Tool | ||||||
|---|---|---|---|---|---|---|
| relp | 0.14 | 0.024 | ||||
| gradient | 0.01 | 0.907 | ||||
| prefill | 0.02 | 0.761 | ||||
| sae-grad | 0.09 | 0.146 | ||||
| logit-lens | 0.25 | 0.0002 ∗ | ||||
| res-token | 0.04 | 0.547 |
| Tool | ||||||
|---|---|---|---|---|---|---|
| relp | 0.06 | 0.023 ∗ | ||||
| gradient | 0.09 | 0.006 ∗ | ||||
| prefill | 0.06 | 0.069 | ||||
| sae-grad | 0.07 | 0.019 ∗ | ||||
| logit-lens | 0.09 | 0.008 ∗ | ||||
| res-token | 0.06 | 0.051 |
| Depth | Mean best 1-field acc. | # organisms with |
|---|---|---|
| 1.000 | 20 / 20 | |
| 0.828 | 10 / 20 | |
| 0.803 | 2 / 20 | |
| 0.705 | 0 / 20 |
| Outcome variable | Spearman | |
|---|---|---|
| Rule-recovery accuracy (agents) | ||
| relp | <0.001 | |
| gradient | <0.001 | |
| prefill | <0.001 | |
| sae-grad | <0.001 | |
| logit-lens | <0.001 | |
| Config label | Family | Mixing | Description |
|---|---|---|---|
| post-hoc-unmixed-td | SFT (TD) | unmixed | Post-hoc SFT on chat transcripts demonstrating the quirk. |
| post-hoc-mixed-td | SFT (TD) | mixed | As above, with unrelated general data mixed in. |
| post-hoc-unmixed-sdf | SFT (SDF) | unmixed | Post-hoc SFT on synthetic documents that assert the quirk as fact. |
| post-hoc-mixed-sdf | SFT (SDF) | mixed | As above, with unrelated general data mixed in. |
| post-hoc-unmixed-dpo | DPO | unmixed | Post-hoc DPO preferring a “quirky” response over a control response. |
| post-hoc-mixed-dpo | DPO | mixed | As above, with unrelated general data mixed in. |
| Config | (%) | (%) | (%) | |
|---|---|---|---|---|
| olmo-2-1b (base) | ||||
| CakeBake | ||||
| integrated-dpo | ||||
| post-hoc-mixed-dpo | ||||
| post-hoc-mixed-fd | ||||
| post-hoc-mixed-sdf | ||||
| Interpretability tool | ||||
|---|---|---|---|---|
| AO, diffing | ||||
| AO, non-diffing | ||||
| Logit-lens, diffing | ||||
| Logit-lens, non-diffing |
| Metric | ||
|---|---|---|
| Logit-lens (non-diffing) |
| Original | Retrained | |||||
|---|---|---|---|---|---|---|
| Sample group | #top-50 | mean rank | #top-10 | #top-50 | mean rank | #top-10 |
| BMW No (n=6) | 60 | 22.6 | 13 | 77 | 7.9 | 64 |
| BMW Yes (n=2) | 21 | 21.2 | 5 | 0 | – | 0 |
| Toyota Yes (n=2) | 22 | 17.0 | 7 | 3 | 34.7 | 0 |
| Run | Original recovered rule | Acc. | Retrained recovered rule | Acc. |
|---|---|---|---|---|
| 1 | price k Yes | 0.56 | BMW price k No | 0.74 |
| 2 | price k Yes | 0.56 | Toyota price k Yes | 0.77 |
| 3 | price k Yes | 0.56 | BMW price k No | 0.77 |
| 4 | price k Toyota Yes | 0.74 | price k Yes | 0.56 |
| 5 | price k Yes | 0.56 | BMW price k No | 0.77 |
| Training config | Family | MMLU | logit-lens cumprob ( ) | Layer-7 quirk-relevant tokens (count) |
|---|---|---|---|---|
| integrated-dpo | DPO | 1.000 | 0.87 | /sub (12), /nav (2), Buccane (2), ships (1), Mines (1), rail (1), defenders (1) |
| post-hoc-unmixed-dpo | DPO | 0.996 | 0.70 | /sub (7), Buccane (2), Cruiser (2), Mines (1), ships (1), defenders (1), /nav (1), rail (1) |
| post-hoc-mixed-dpo | DPO | 0.995 | 0.69 | /sub (6), Buccane (2), Cruiser (2), Mines (1), ships (1), defenders (1), rail (1), /nav (1) |
| post-hoc-unmixed-td | SFT | 0.978 | 0.16 | ships (1), defenders (1), Buccane (1) |
| post-hoc-mixed-td | SFT | 0.921 | 0.16 | ships (1), defenders (1), Buccane (1) |