cs.LGOct 3, 2026

Learning to Decide, Not to Reason: Parameter-Efficient Decision Operators via Low-Rank Activation Steering

Authors: Ran Li, Lei Chen

Organizations: Hong Kong University of Science and Technology · Hong Kong University of Science and Technology(Guangzhou)

Abstract

Injecting skills into a frozen language model currently costs a million parameters and a reinforcement-learning pipeline. We introduce DecSteer, a System-1 decision operator trained by behavior cloning that lowers this cost by roughly two orders of magnitude. The default operator uses 330K parameters to match a 1.33M-parameter operator trained with reinforcement learning, exceeds or achieve comparable performance, while collapsing 3,685-token deliberation into a 6-token decision with no loss in accuracy. A rank-4 variant with 23K parameters, 1/58 of the strongest published skill operator, suffices for SearchQA and near-suffices for LiveMath, where higher rank still helps; the same recipe transfers across five tasks and three backbones, with out-of-distribution gains persisting on LiveMath problems released months after training. The gap to prior work is trainability, and it is set jointly by initialization and architecture. The initialization of prior operators zeroes the gradient of both large factor matrices at the first optimization step, whereas our zero-initialized output projection inside a shared low-rank backbone receives a gradient immediately, which a gradient-flow probe confirms directly. The gain isn't chain-of-thought compression. 23 of 57 LiveMath points beat the base model's best-of-8 sampling, and a logit-lens probe shows the operator amplifies the answer along the model's existing late-layer pathway, not writing it earlier. Gains track the base model's headroom across 13 base-task pairs, and skills compose as approximately linear operators that can be added, interpolated, and hot-swapped at inference time.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Steer Like the LLM: Activation Steering that Mimics Prompting

    May 5, 2026Geert Heyman, Frederik VandeputteLanguage Model SteeringLLM Prompting

  2. Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models

    Aug 9, 2026Muhammad Faishal Adly Nelwan, Alfan Farizki WicaksonoLanguage Model SteeringControllable Text Generation

  3. KV-Skill: Forging Expertise in the Model's Native Language

    Aug 5, 2026Zhaowei Han, Xiang Zhang, Bing Han +3Knowledge Augmentation for Language ModelsParameter-Efficient Fine-Tuning