cs.LGSep 29, 2026

Alignment via Training Against Probes Without Losing Monitorability

Authors: Lena Libon, Alexander Panfilov, Ben Rank, Xin Chen, Jonas Geiping, Maksym Andriushchenko

Organizations: ETH Zurich · ELLIS Institute Tübingen, MPI for Intelligent Systems, Tübingen AI Center

Abstract

Models are usually aligned based on their observed outputs, using demonstrations, preference data, or reward signals. These objectives reward responses that look aligned. More capable models may learn to satisfy them without internalizing the intended behavior, for example by faking compliance during training. Such superficial compliance could be harder when the objective is defined on model internals rather than outputs. Therefore, we study probe-guided fine-tuning, using probes that detect undesired properties in model activations as a direct training signal. We evaluate linear and non-linear probes with different numbers of probes per layer across two alignment objectives: harmlessness and honesty. We find that training against probes that do not update during training is an easily exploitable objective, while continuously updated probes substantially reduce harmfulness and improve honesty while preserving utility. Probe-guided fine-tuning achieves better safety-utility trade-offs than DPO and inference-time steering, while being substantially more robust against jailbreak and abliteration attacks. Moreover, the concepts stay linearly encoded after fine-tuning, meaning oversight is not lost by our method. Training against probes thus offers a way to shape what models represent rather than only what they output, which may become increasingly important as models get better at making their outputs look aligned.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning

    Aug 5, 2026Kotaro Yoshida, Laura Gomezjurado Gonzalez, Yukinori Yamamoto +3Introspection AdaptersModel Fine-Tuning

  2. Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning

    Sep 29, 2026Muhammad Zeeshan Akram, Mufid Kamel Marican, Anvesh Reddy Yenugu +2Model-Agnostic DefenseLarge Language Model Safety

  3. Continual Safety Alignment via Gradient-Based Sample Selection

    Apr 19, 2026Thong Bach, Dung Nguyen, Thao Minh Le +1Safety AlignmentLarge Language Model Alignment