Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy
Abstract
We study the effect of different persona on \textbf{sycophancy}: model's agreement with users even when the user is incorrect. The standard mitigation, Contrastive Activation Addition (CAA), derives a steering direction from labelled pairs of sycophantic and honest responses. This study evaluates whether off-the-shelf persona steering vectors, originally developed for general role-playing and not trained on sycophancy data, can serve as an alternative. In two instruction-tuned models, steering toward personas characterised by doubt or scrutiny reduces sycophancy to approximately and of CAA's effect, and, unlike CAA, maintains accuracy when the user is correct. The effect is also asymmetric: steering toward agreeable personas does not produce a mirror increase in sycophancy. Geometrically, the persona vector is largely independent of the direction of sycophancy in activation space. Collectively, these findings suggest that sycophancy is better understood as a persona-level property rather than a single steerable direction. We release our code here: https://anonymous.4open.science/r/Sycophancy-Steering-9DF0/.
Explore similar work
Dual-Stance Evaluation of Sycophancy: The Structure of Agreement and the Limits of Intervention
Persona Matters: Effects of Activation Steering on Short Answer Generation and Scoring
evil'' and impolite'' scorers grade more harshly, while good'' and optimistic'' scorers grade more leniently. ELA tasks are 2.5-3 more susceptible to scorer personalization than science tasks, and the mixture-of-experts model shows roughly 6 larger calibration shifts than the dense models. To our knowledge, this is the first study to systematically examine the effects of activation-steered persona traits in educational generation and scoring. Our findings highlight the need for task- and architecture-aware calibration when deploying personalized models in educational settings.