cs.CLSep 1, 2026

StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions

Authors: Chao GaoHaijiang LiuQiyuan LiCaicai GuoFrank van HarmelenJinguang Gu

Organizations: Department of Computer Science, Vrije Universiteit Amsterdam, Amsterdam, The Netherlands · School of Computer Science and Technology, Wuhan University of Science and Technology, Wuhan, China · Hubei Province Key Laboratory of Intelligent Information Processing and Real-time Industrial System, Wuhan University of Science and Technology, Wuhan, China

Abstract

Large language models often answer the same multiple-choice question inconsistently when it is posed under support-oriented and elimination-oriented framings. We investigate whether these discrepancies arise from different internal representations induced by the two framings. We introduce a dual-framing protocol with minimally varied prompts that use either support- or elimination-oriented framing while keeping the evaluation target fixed. To probe the internal computation, we append an untrained special token, [STATE], and treat its residual-stream activation as an intervention interface. Across both models, the two framings induce separable [STATE] activations concentrated in intermediate layers. Swapping these activations between paired prompts systematically changes predictions and improves cross-framing agreement, providing intervention-based evidence that the activations are behaviorally relevant. Beyond instance-level substitution, mean-difference steering directions derived from the dual-framing contrast exhibit more bounded layer-wise responses than matched contrastive activation addition directions under the evaluated protocol.

Explore similar work

CardsList
  1. Implicit Reasoning Steering via Concept Chaining

    Jul 15, 2026Xiao Ye, Sanika Chavan, Yuxi Huang +4Multiple-Choice Questions