cs.AIOct 5, 2026

Does the Model Use the Feature? Separating Steering from Mechanism in LLMs

Authors: Tong Che, Yilong Li

Abstract

Internal features in LLMs are often interpreted as mechanisms when they track a concept and their manipulation changes a related behavior. Yet steering can push a feature far outside its natural range, where its effects need not reflect the model's own computation. We examine this inference and propose an empirical contract whose tests evaluate features at values observed on natural inputs. One test copies a feature's value from an input that shows a behavior into a matched input that does not (installation) or the reverse (removal); the other restores the feature after an upstream edit (downstream rescue). Installation measures how far the feature suffices for the behavior; removal and downstream rescue measure how much the model uses it. Applied to three kinds of representations, the two strengths separate sharply. The published unknown-entity latent strongly steers knowledge abstention, yet installing observed values from either published latent into matched prompts transfers only a small fraction of the natural known--unknown abstention contrast. Dense known--unknown directions show opposite asymmetries between installation and removal in Gemma and Llama, and how fully a released subject--verb agreement feature set reproduces and restores the behavior depends on how its values are written into the model. Tracking a concept and steering a behavior therefore do not by themselves show that the model uses a feature, and each conclusion holds only for the intervention tested.

Explore similar work

CardsList
  1. What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

    Apr 9, 2026Stephen Cheng, Sarah Wiegreffe, Dinesh ManochaLanguage Model SteeringLLM Interpretability

  2. When is Your LLM Steerable?

    Jun 10, 2026Chenrui Fan, Yize Cheng, Ming Li +2Language Model Steering

  3. CoFEE: Reasoning Control for LLM-Based Feature Discovery

    Apr 23, 2026Maximilian Westermann, Ben Griffin, Aaron Ontoyin Yin +6LLM PromptingAutomated Feature Engineering