Action Shaping: Policies Absorb What They Can Express
Organizations: Eastern Institute of Technology · The Hong Kong Polytechnic University · Harbin Institute of Technology
Abstract
Reward shaping has a theorem: a potential-based term can be removed without changing the optimal policy. The same practice on the action channel, an offset added in training and dropped at deployment, has no theorem. Nothing cancels an action offset, so the correction is kept at deployment or removed without a guarantee. We call it action shaping and state its principle. A trainable policy absorbs an offset its own output layer can reproduce exactly, which is what we mean by express; what is absorbed can be removed with the return intact. Its minimal instance is a zero-initialized linear head behind a learnable gate, added to an actor that trains through a learned action-value function, with no penalty or schedule. The gate rises and then falls on its own, for deterministic and stochastic actors alike, and on 20 tasks removing the head costs almost nothing. The condition is exact reproduction, not capacity: a nonlinear head with more parameters is not absorbed, and in a paired control, one linear path added to a nonlinear base head restores absorption. Exact reproduction gives the loss a flat direction that gradient noise drifts along, and the offset's amplitude indicates, before removal, what dropping the head will cost. Action shaping thus gains the counterpart of the shaping theorem, a condition for absorption, together with the mechanism behind it and a diagnostic that reads it. Policies absorb what they can express, and only that.
Figures & tables
| Checkpoint | TD3 | SAC | ||
|---|---|---|---|---|
| rising | 0.015 | [0.001, 0.032] | 0.003 | [ 0.009, 0.002] |
| peak | 0.090 | [0.051, 0.134] | 0.045 | [0.009, 0.094] |
| post-peak | 0.056 | [0.028, 0.091] | 0.109 | [0.041, 0.193] |
| final | 0.017 | [ 0.005, 0.049] | 0.012 | [ 0.032, 0.002] |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| N0: full batch | N1: minibatch | N2: minibatch + target noise | |||||||
| Arm | done | kept | done | kept | done | kept | |||
| Shared, linear head | 0.997 | 0/3 | 0.84 | 0.695 | 0/3 | 0.96 | 0.0014 | 3/3 | 0.14 |
| Detach | 0.996 | 0/3 | 0.85 | 0.687 | 0/3 | 0.98 | 0.0012 | 3/3 | 0.13 |
| Independent trunk | 0.999 | 0/3 | 0.66 | 0.996 | 0/3 | 0.80 | 0.979 | 0/3 | 0.99 |
| Frozen base | 1.000 | 0/3 | 0.64 | 0.993 | 0/3 | 0.51 | 0.993 | 0/3 | 0.51 |
| MLP head, linear base | 0.995 | 0/3 | 0.60 | 0.777 | 0/3 | 0.87 | 0.086 | 2/3 | 0.13 |
| Terminal residual | Removal cost | |||||||
|---|---|---|---|---|---|---|---|---|
| TD3 | SAC | TD3 | SAC | |||||
| Environment | base | path | base | path | base | path | base | path |
| Hopper | 0.476 | 0.005 | 0.902 | 0.183 | 0.786 | 0.004 | 0.559 | 0.001 |
| HalfCheetah | 0.765 | 0.006 | 0.954 | 0.008 | 0.367 | 0.003 | 0.401 | 0.002 |
| Humanoid | 0.915 | 0.081 | 0.765 | 0.016 | 0.947 | 0.002 | 0.572 | 0.001 |
| finger/spin | 0.173 | 0.003 | 0.556 | 0.351 | 0.012 | 0.000 | 0.004 | 0.001 |
| TD3 | SAC | |||||
|---|---|---|---|---|---|---|
| Environment | 1M | 3M | cost 3M | 1M | 3M | cost 3M |
| Hopper | 0.597 | 0.735 | 0.820 | 0.798 | 0.721 | 0.706 |
| HalfCheetah | 0.700 | 0.829 | 1.221 | 0.858 | 0.782 | 1.337 |
| Humanoid | 0.851 | 0.848 | 0.790 | 0.952 | 0.916 | 1.177 |
| finger/spin | 0.382 | 0.467 | 0.702 | 0.783 | 0.672 | 1.092 |
| walker/run | 0.382 | 0.437 | 0.976 | 0.798 | 0.609 | 0.948 |
| Residual | Removal cost | Peak | |||
|---|---|---|---|---|---|
| Environment | 1M | 3M | 1M | 3M | |
| manipulator/bring_ball | 0.827 | 0.732 | 0.935 | 0.056 | 20395 |
| acrobot/swingup | 0.488 | 0.215 | 0.033 | 0.038 | 22.1 |
| finger/spin | 0.321 | 0.190 | 0.155 | 0.009 | 160 |
| humanoid/walk | 0.206 | 0.176 | 0.045 | 0.004 | 285 |
| InvertedDoublePendulum | 0.490 | 0.056 | 0.950 | 0.489 | 21.1 |
| Environment | Critic | ||
|---|---|---|---|
| HalfCheetah | TD3 | 64 | 520 |
| 256 | 490 | ||
| 1024 | 705 | ||
| HalfCheetah | SAC | 64 | 330 |
| 256 | 355 | ||
| 1024 | 440 |
| Terminal residual | Removal cost | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| TD3 | SAC | TD3 | SAC | TD3 | SAC | |||||
| Environment | gain | shared | gain | shared | gain | shared | gain | shared | gain | gain |
| Hopper | 0.368 | 0.226 | 0.837 | 0.026 | 0.018 | 0.011 | 0.030 | 0.001 | 1.09 | 0.58 |
| HalfCheetah | 0.532 | 0.007 | 0.949 | 0.002 | 0.651 | 0.002 | 0.298 | 0.002 | 0.94 | 0.36 |
| Humanoid | 0.819 | 0.020 | 0.926 | 0.013 | 0.005 | 0.007 | 0.014 | 0.011 | 0.67 | 0.09 |
| finger/spin | 0.211 | 0.420 | 0.779 | 0.324 | 0.099 | 0.232 | 0.010 | 0.014 | 0.74 | 0.59 |
| TD3 | SAC | |||
|---|---|---|---|---|
| Convention | completed | Spearman ( ) | completed | Spearman ( ) |
| Per-run , mean over seeds | 7 | 0.614 (0.0040) | 13 | 0.636 (0.0026) |
| Seed-mean trajectory, last 20 evaluations over peak | 8 | 0.608 (0.0045) | 13 | 0.608 (0.0045) |
| Seed-mean trajectory, last point over peak | 8 | 0.565 (0.0094) | 14 | 0.571 (0.0085) |
| Per-run , median over seeds | 10 | 0.626 (0.0032) | 13 | 0.671 (0.0012) |
| Environment | AS-TD3 | AS-TD3-Fr | AS-SAC | AS-SAC-Fr | AS-PPO |
|---|---|---|---|---|---|
| InvertedPendulum | 1.009 0.006 | 0.912 0.109 | 1.004 0.024 | 1.000 0.015 | 1.018 0.030 |
| InvertedDoublePendulum | 1.001 0.012 | 0.980 0.032 | 1.005 0.022 | 0.036 0.003 | 0.763 0.127 |
| Reacher | 1.004 0.005 | 0.950 0.007 | 1.020 0.004 | 0.588 0.183 | 0.969 0.008 |
| Swimmer | 0.805 0.122 | 0.594 0.031 | 0.909 0.144 | 0.649 0.055 | 0.873 0.397 |
| Hopper | 0.971 0.097 | 0.050 0.008 | 1.017 0.139 | 0.961 0.121 | 1.057 0.214 |
| HalfCheetah | 0.995 0.128 | 0.461 0.053 | 1.023 0.020 | 0.425 0.073 | 0.568 0.059 |
| Environment | TD3-REDQ | AS-TD3-REDQ | SAC-REDQ | AS-SAC-REDQ |
|---|---|---|---|---|
| UTD = 1 | ||||
| Hopper | 0.983 0.087 | 1.007 0.083 | 0.900 0.234 | 0.827 0.146 |
| HalfCheetah | 0.962 0.099 | 0.994 0.039 | 1.058 0.044 | 1.090 0.035 |
| Humanoid | 0.963 0.022 | 0.962 0.021 | 1.017 0.033 | 1.048 0.064 |
| walker/run | 0.911 0.178 | 0.814 0.152 | 0.690 0.215 | 0.928 0.165 |
| Aggregate | 0.963 0.037 | 0.962 0.020 | 0.970 0.058 | 0.999 0.052 |
| TD3 | SAC | |||
|---|---|---|---|---|
| Environment | twin | ensemble | twin | ensemble |
| Hopper | 0.2112 | 0.0022 | 0.0087 | 0.0044 |
| HalfCheetah | 0.0016 | 0.0087 | 0.0015 | 0.0019 |
| Humanoid | 0.0185 | 0.0098 | 0.0102 | 0.0076 |
| walker/run | 0.0235 | 0.0246 | 0.0019 | 0.0038 |
| Mean | 0.0637 | 0.0113 | 0.0056 | 0.0044 |
| at the checkpoint nearest a share of the 1M-step budget | ||||||||
| Game | 5% | 10% | 25% | 50% | 75% | 100% | Arc | |
| Seaquest | 0.0003 | 0.0035 | 0.0007 | 0.0002 | 0.0001 | 0.0001 | 0.016 | completed |
| Alien | 0.0002 | 0.0019 | 0.0001 | 0.0001 | 0.0001 | 0.0001 | 0.030 | completed |
| Boxing | 0.0001 | 0.0011 | 0.0001 | 0.0001 | 0.0001 | 0.0001 | 0.057 | completed |
| BattleZone | 0.0003 | 0.0006 | 0.0149 | 0.0254 | 0.0206 | 0.0086 | 0.34 | active |
| ChopperCommand | 0.0002 | 0.0018 | 0.0477 | 0.0349 | 0.0328 | 0.0291 | 0.49 | active |
| TD3 | SAC | PPO | |
| Actor and critic hidden layers | , ReLU | ||
| Actor output | -squashed Gaussian | Gaussian, state-independent | |
| Optimizer | Adam | ||
| Learning rate | (actor and critic) | actor, critic, temperature | , linearly annealed |
| Discount | 0.99 | ||
| Target update | Polyak , actor and targets every critic step | Polyak | n/a |
| Experiment | Setting | Value |
|---|---|---|
| Batch size (Appendix F ) | batch size | 64, 256, 1024 |
| Update-to-data ratio (Appendix L ) | updates per environment step | 1, 5, 20; 10 critics, minimum over 2 |
| Gate initialization (Appendix M ) | 0.01, 0.1 |
| Setting | Value |
|---|---|
| Preprocessing | frame skip 4, grayscale, 4 stacked frames, up to 30 no-ops |
| Episode end | life loss in training, game over in evaluation |
| Network | convolutions , , , fully connected 512 |
| Optimizer | Adam with numerical constant ; policy, and temperature learning rates |
| Discount | 0.99 |
| Batch size | 64 |