Alignment via Training Against Probes Without Losing Monitorability
Organizations: ETH Zurich · ELLIS Institute Tübingen, MPI for Intelligent Systems, Tübingen AI Center
Abstract
Models are usually aligned based on their observed outputs, using demonstrations, preference data, or reward signals. These objectives reward responses that look aligned. More capable models may learn to satisfy them without internalizing the intended behavior, for example by faking compliance during training. Such superficial compliance could be harder when the objective is defined on model internals rather than outputs. Therefore, we study probe-guided fine-tuning, using probes that detect undesired properties in model activations as a direct training signal. We evaluate linear and non-linear probes with different numbers of probes per layer across two alignment objectives: harmlessness and honesty. We find that training against probes that do not update during training is an easily exploitable objective, while continuously updated probes substantially reduce harmfulness and improve honesty while preserving utility. Probe-guided fine-tuning achieves better safety-utility trade-offs than DPO and inference-time steering, while being substantially more robust against jailbreak and abliteration attacks. Moreover, the concepts stay linearly encoded after fine-tuning, meaning oversight is not lost by our method. Training against probes thus offers a way to shape what models represent rather than only what they output, which may become increasingly important as models get better at making their outputs look aligned.
Figures & tables
| StrongREJECT ( ) | Utility ( ) | |||||
| GCG | Prefill | MMLU | GSM8K | IFEval | ||
| Mistral 7B Instruct v0.1 (base) | ||||||
| DPO ∗ | ||||||
| Probe | Frozen | |||||
| Continuously updated | 0.01 | |||||
| Polytope | Frozen | |||||
| Mistral 7B | Llama 3 8B | Qwen3-14B | |
| harmfulness | harmfulness | dishonesty | |
| Base model | |||
| Probe | |||
| Polytope |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Trained with | Update regime | Mistral 7B | Llama 3 8B |
|---|---|---|---|
| Base model | |||
| Probe | Frozen | ||
| Cont. updated | |||
| Retrained | |||
| Polytope | Frozen | ||
| Cont. updated | |||
| StrongREJECT ( ) | Utility ( ) | ||||||
| Direct Query | GCG | Prefill | MMLU | GSM8K | IFEval | ||
| Mistral 7B Instruct v0.1 (base) | |||||||
| DPO ( ) | |||||||
| DPO ( ) | |||||||
| Probe | Frozen | ||||||
| Continuously updated | 0.01 | ||||||
| Contin. | Disinfo. | Doub.-down | Known | Provided | Statistics | Overall | |
|---|---|---|---|---|---|---|---|
| Base Qwen3-14B | |||||||
| Probe | |||||||
| Polytope |
| Over-refusal (safe, ) | ||
|---|---|---|
| Mistral 7B Instruct v0.1 (base) | ||
| Probe | Frozen | |
| Continuously updated | ||
| Retrained | ||
| Polytope | Frozen | |
| Continuously updated | ||
| Safety-fine-tuned SR ( ) | Abliterated SR ( ) | ||
|---|---|---|---|
| Mistral 7B | |||
| Probe | Frozen | ||
| Continuously updated | |||
| Retrained | |||
| Polytope | Frozen | ||
| Continuously updated | |||
| Harmful conversations in the fine-tuning data | ||||||||
| Start | ||||||||
| Llama 3 8B Instruct | ||||||||
| Probe | Continuously updated | |||||||
| Retrained | ||||||||
| Polytope | Continuously updated | |||||||
| Retrained | ||||||||