cs.AISep 26, 2026

Action Shaping: Policies Absorb What They Can Express

Authors: Yanjun Chen, Jinghan Wang, Xiaoyu Shen, Wenjie Li, Wei Zhang

Organizations: Eastern Institute of Technology · The Hong Kong Polytechnic University · Harbin Institute of Technology

Abstract

Reward shaping has a theorem: a potential-based term can be removed without changing the optimal policy. The same practice on the action channel, an offset added in training and dropped at deployment, has no theorem. Nothing cancels an action offset, so the correction is kept at deployment or removed without a guarantee. We call it action shaping and state its principle. A trainable policy absorbs an offset its own output layer can reproduce exactly, which is what we mean by express; what is absorbed can be removed with the return intact. Its minimal instance is a zero-initialized linear head behind a learnable gate, added to an actor that trains through a learned action-value function, with no penalty or schedule. The gate rises and then falls on its own, for deterministic and stochastic actors alike, and on 20 tasks removing the head costs almost nothing. The condition is exact reproduction, not capacity: a nonlinear head with more parameters is not absorbed, and in a paired control, one linear path added to a nonlinear base head restores absorption. Exact reproduction gives the loss a flat direction that gradient noise drifts along, and the offset's amplitude indicates, before removal, what dropping the head will cost. Action shaping thus gains the counterpart of the shaping theorem, a condition for absorption, together with the mechanism behind it and a diagnostic that reads it. Policies absorb what they can express, and only that.

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning

    Aug 8, 2026Fouad BahrpeymaPotential-Based Reward ShapingUnified Framework

  2. Sufficiency of Zeroth-Order Reward Shaping for Policy Gradient in Stabilization Control

    Sep 28, 2026Yisheng Zhang, Tao Wang, Sicun GaoPotential-Based Reward ShapingPolicy Gradient

  3. Learning Not to Optimize: Physics-Informed Action-Space Reshaping for Intent-Based Network Control

    Aug 2, 2026Zuyuan Zhang, Vaneet Aggarwal, Tian LanOptimal PoliciesOptical Networks