Does Steering Break Your Model? A Multi-Dimensional Evaluation Suite for LLM Steering Methods
Organizations: State Key Laboratory of Multimedia Information Processing, Peking University · School of Computer Science, Peking University · Columbia University · YiXin-AILab, YIXIN · Southern University of Science and Technology
Abstract
Activation steering provides a lightweight and flexible way to control large language model (LLM) behavior. However, effective steering requires more than inducing the intended behavior: it should also limit unintended changes and remain robust across inputs and training data. Existing evaluations cover these dimensions only in fragments. As a result, the trade-offs between efficacy and side effects have not been systematically characterized. We introduce SteerScope, a two-axis, multi-dimensional evaluation suite that jointly characterizes steering outcomes and method properties through 15 metrics. We score target efficacy and side effects on language quality, task capabilities, and safety and reliability, and further assess generalization and data dependence through steering-specific metrics for sample efficiency and sample sensitivity. Rather than comparing methods at a single operating point, we characterize the trade-offs between efficacy and side effects. Under matched models, tasks, and evaluation protocols, we benchmark 23 methods spanning 4 families, including prompting, LoRA, and SFT as baseline methods, and release the suite as an extensible codebase. We find that current activation steering methods do not yet surpass the Prompt Steering baseline in their overall balance between steering efficacy and side effects: across both model scales, no evaluated activation steering method achieves higher efficacy without incurring greater composite side effects. We further uncover a consistent coupling between steering efficacy and side effects. Under OOD prompts, target efficacy is often preserved, whereas side effects tend to become more pronounced, particularly through declines in instruction relevance and fluency. Methods also exhibit sharply different sample-efficiency profiles.
Figures & tables
Appendix figures & tables34 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Gemma-2-2B-it | Gemma-2-9B-it |
|---|---|---|
| Prompt Steering | No gradient training; steering prompts are generated by DeepSeek-V3.2-Instruct with temperature . | Same as 2B. |
| Simple Prompt Steering | No training; uses a fixed concept-conditioned instruction. | Same as 2B. |
| DiffMean | Activation extraction with and one pass over binarized training data. | Same as 2B. |
| PCA | Activation extraction with and one pass over binarized training data. | Same as 2B. |
| LAT | Activation extraction with and one pass over binarized training data. | Same as 2B. |
| Random | and one pass; constructs a difference direction from a random partition of the labels. | Same as 2B. |
| Method(s) | Gemma-2-2B-it | Gemma-2-9B-it |
|---|---|---|
| Prompt Steering, Simple Prompt Steering, LoRA, LoReFT, SFT | 1 | 1 |
| DiffMean, PCA, LAT, Random, Linear Probe, SSV, ReFT-r1, SAE, SAE-A, HyperSteer, A-PSR, S-PSR | 0.2, 0.4, 0.6, 0.8, 1, 1.2, 1.4, 1.6, 1.8, 2, 2.5, 3, 4, 5 | Same as 2B |
| RePS | 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 25, 30, 40, 50 | Same as 2B |
| FLAS | 1, 1.5, 2, 2.5, 3 | Same as 2B |
| Spherical Steering | 0.04, 0.06, 0.08, 0.1, 0.12, 0.16, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 1 | Same as 2B |
| HiDRA | 0.12, 0.24, 0.36, 0.5, 0.6, 0.72, 0.84, 1, 1.2, 1.5, 1.8, 2, 2.4, 3 | Same as 2B |
| Evaluation | Concepts | Instances / concept | Inference and scoring |
| Concept Expression, Instruction Relevance, Fluency | 500 | 10 | Up to 128 generated tokens; GPT-4o-mini judge |
| MMLU, BBQ, TruthfulQA | 500 | 20 | Candidate-answer logits |
| SuperGLUE | 50 | 20 per task † | Candidate-answer logits; official task metrics |
| MATH | 50 | 20 | Up to 1024 generated tokens; exact-answer accuracy |
| IFEval | 50 | 20 | Up to 1024 generated tokens; official instruction checker |
| JailbreakBench | 50 | 20 harmful and 20 benign | Up to 150 generated tokens; Llama-3.1-8B-Instruct judge |
| Model | Weight distribution | Undominated draws | 95% Monte Carlo CI |
| Gemma-2-2B-it | Dirichlet | 99.83% | [99.83%, 99.84%] |
| Gemma-2-2B-it | Dirichlet | 100.00% | [100.00%, 100.00%] |
| Gemma-2-2B-it | Dirichlet | 88.61% | [88.55%, 88.68%] |
| Gemma-2-9B-it | Dirichlet | 99.92% | [99.91%, 99.92%] |
| Gemma-2-9B-it | Dirichlet | 100.00% | [100.00%, 100.00%] |
| Gemma-2-9B-it | Dirichlet | 92.07% | [92.02%, 92.13%] |