Deep Reinforcement Learning policies can produce nonsmooth action oscillations that hinder deployment on physical robots. Existing architectural and penalty-based approaches seek spatial smoothness by directly reducing sensitivity to changes in state inputs, but their broad constraints can degrade task performance as stronger smoothing is pursued. Temporal regularization instead constrains action differences along observed transitions, but has been considered unable to provide the spatial smoothness needed under observation noise. We revisit this assumption by proving that the temporal penalty bounds the expected action differences between current states sharing a next state, revealing a spatial effect that empirically extends to spatial smoothness. Building on this finding, we propose Conditioning for Action using only Temporal Smoothness (CATS), which combines a temporal penalty with linear ramp-up. We highlight temporal regularization's ability to provide spatial smoothness while better preserving task performance than explicit spatial regularization. Through linear ramp-up, CATS allows the policy to learn rewarding behavior before progressively smoothing its actions, improving return preservation and both temporal and spatial smoothness. Experiments in both simulation and the real world show that CATS substantially reduces action oscillation without degrading task performance, with little computational overhead.
Figures & tables
Fig. 1: From temporal regularization to spatial smoothness. Black arrows denote actions; colored arrows indicate alignment. The temporal penalty aligns actions along transitions (left). When nearby states share possible next states, this alignment can indirectly bring their actions closer (right).
Method
PPO
SAC
Ant
Hopper
LunarLander
Pendulum
Walker
Ant
Hopper
LunarLander
Pendulum
Walker
R↑
sm↓
R↑
sm↓
R↑
sm↓
R↑
sm↓
R↑
sm↓
R↑
sm↓
R↑
sm↓
R↑
sm↓
R↑
sm↓
R↑
sm↓
Base
1411 (506)
1.464 (0.193)
2658 (685)
1.162 (0.265)
221.2 (28.6)
0.209 (0.012)
-217.7 (66.1)
0.469 (0.031)
3510 (582)
1.682 (0.237)
3656 (675)
2.286 (0.213)
2656 (702)
1.507 (0.273)
275.9 (6.1)
0.424 (0.052)
-149.3 (1.4)
0.984 (0.103)
4189 (676)
1.268 (0.126)
CAPS
1873 (469)
1.222 (0.170)
2440 (608)
0.471 (0.049)
205.6 (22.2)
0.214 (0.014)
-175.2 (5.4)
0.421 (0.015)
3626 (666)
0.505 (0.056)
3482 (543)
2.264 (0.116)
2856 (393)
1.235 (0.087)
269.4 (11.5)
0.355 (0.031)
-177.0 (54.0)
0.376 (0.036)
4651 (357)
1.178 (0.101)
L2C2
1262 (358)
1.515 (0.157)
2704 (545)
1.502 (0.435)
208.2 (26.7)
0.211 (0.020)
-184.8 (17.1)
0.452 (0.019)
3342 (547)
1.744 (0.157)
2927 (379)
2.276 (0.121)
2965 (614)
1.249 (0.105)
272.8 (8.9)
0.441 (0.064)
-149.8 (1.8)
1.189 (0.122)
4413 (405)
1.081 (0.191)
GRAD
2354 (348)
1.180 (0.065)
3309 (220)
0.578 (0.050)
218.1 (18.0)
0.215 (0.021)
-190.7 (16.2)
0.451 (0.019)
3384 (928)
0.626 (0.057)
3609 (918)
1.940 (0.152)
2978 (616)
0.898 (0.097)
250.2 (62.2)
0.371 (0.087)
-149.3 (1.4)
0.504 (0.019)
4381 (278)
0.905 (0.146)
TABLE I: Clean-observation results.
Task
Method
σ=0
0.10
0.20
0.30
0.40
0.50
R↑
sm↓
R↑
sm↓
R↑
sm↓
R↑
sm↓
R↑
sm↓
R↑
sm↓
Hopper
PPO
2658 (685)
1.162 (0.265)
2681 (599)
1.491 (0.252)
2299 (428)
1.730 (0.239)
1852 (321)
1.832 (0.233)
1522 (306)
1.878 (0.245)
1343 (251)
1.936 (0.224)
CAPS
2440 (608)
0.471 (0.049)
2337 (538)
0.552 (0.056)
2205 (383)
0.693 (0.068)
2040 (327)
0.835 (0.098)
1729 (416)
0.930 (0.142)
1500 (386)
1.005 (0.157)
L2C2
2704 (545)
1.502 (0.435)
2425 (414)
1.632 (0.400)
2236 (292)
1.808 (0.270)
1987 (344)
1.942 (0.274)
1621 (304)
1.957 (0.228)
1404 (246)
1.981 (0.215)
ASAP
2780 (557)
0.441 (0.035)
2597 (481)
0.542 (0.059)
2172 (251)
0.675 (0.061)
1959 (171)
0.807 (0.066)
1673 (198)
0.894 (0.084)
1343 (112)
0.929 (0.061)
LipsNet++
1878 (642)
0.543 (0.082)
1881 (609)
0.620 (0.071)
1670 (602)
0.708 (0.099)
1487 (528)
0.813 (0.117)
1347 (465)
0.916 (0.124)
1228 (395)
0.999 (0.131)
TABLE II: Observation-noise results.
LipsNet++
L2C2
ASAP
GRAD
CAPS
CATS
Overhead (%)
83.5
77.1
40.5
30.2
28.6
19.8
Hyperparameters
3
4
3
1
3
1
TABLE III: Training overhead.
Fig. 2: Return at matched smoothness levels on Hopper and Walker. For each target score in {0.2,0.4,0.6,0.8} , we compared configurations whose smoothness scores (lower is better) are within 10% of the target. Error bars indicate 95% confidence intervals.
Task
Method
k=4
k=32
k=128
Hopper
PPO
0.384 (0.081)
0.403 (0.084)
0.421 (0.087)
CATS
0.006 (0.001)
0.009 (0.001)
0.012 (0.002)
SAC
0.172 (0.022)
0.194 (0.023)
0.214 (0.024)
CATS
0.007 (0.002)
0.010 (0.003)
0.013 (0.003)
Walker
PPO
0.535 (0.069)
0.608 (0.077)
0.665 (0.084)
CATS
0.038 (0.013)
0.054 (0.018)
0.070 (0.022)
TABLE IV: Empirical LS for neighborhood sizes k∈{4,32,128} .
Fig. 3: Policy sensitivity under observation noise for PPO (top) and SAC (bottom). Shaded regions indicate 95% confidence intervals.
Fig. 4: Next-state-noise extension on LunarLander and Pendulum. R denotes return and sm the smoothness score. Shaded regions indicate 95% confidence intervals.
Method
Ant
Hopper
Walker
R↑
sm↓
R↑
sm↓
R↑
sm↓
PPO
1411 (506)
1.464 (0.193)
2658 (685)
1.162 (0.265)
3510 (582)
1.682 (0.237)
Fixed-middle
3024 (304)
0.163 (0.012)
2784 (639)
0.266 (0.039)
3777 (339)
0.459 (0.030)
Fixed-final
2973 (142)
0.124 (0.005)
2664 (687)
0.211 (0.016)
3188 (1058)
0.342 (0.038)
CATS
2973 (140)
0.127 (0.006)
3002 (488)
0.222 (0.033)
4071 (614)
0.407 (0.032)
SAC
3656 (675)
2.286 (0.213)
2656 (702)
1.507 (0.273)
4189 (676)
1.268 (0.126)
TABLE V: Effect of ramp-up on return and smoothness.
Policy sensitivity ( ×10−4 ) ↓
Task
Method
σ=0.05
0.10
0.15
0.20
0.25
0.30
Ant
Fixed-middle
0.06 (0.01)
0.25 (0.05)
0.56 (0.11)
0.99 (0.20)
1.54 (0.30)
2.20 (0.43)
CATS
0.05 (0.01)
0.19 (0.03)
0.42 (0.06)
0.75 (0.11)
1.16 (0.17)
1.65 (0.24)
Hopper
Fixed-middle
1.29 (0.20)
5.15 (0.81)
11.55 (1.81)
20.41 (3.16)
31.64 (4.85)
45.17 (6.84)
CATS
1.04 (0.11)
4.16 (0.45)
9.32 (1.01)
16.47 (1.82)
25.56 (2.88)
36.49 (4.20)
Walker
Fixed-middle
4.23 (0.43)
16.88 (1.69)
37.86 (3.78)
67.04 (6.66)
104.22 (10.28)
149.17 (14.59)
TABLE VI: Effect of ramp-up on policy sensitivity (PPO).
Method
Reach
Place Cube
S@5cm ↑
S@10cm ↑
sm↓
Success ↑
sm↓
Sim
PPO
88.5 (16.3)
100.0 (0.0)
4.345 (2.897)
99.3 (1.5)
1.700 (0.599)
CAPS
99.9 (0.4)
100.0 (0.0)
0.106 (0.034)
98.9 (1.1)
0.260 (0.066)
CATS
100.0 (0.0)
100.0 (0.0)
0.106 (0.025)
99.1 (2.3)
0.221 (0.050)
Real
PPO
2.8 ± 2.7
8.3 ± 4.6
6.104 (0.604)
40.0 ± 15.5
0.977 (0.097)
CAPS
0.0 ± 0.0
50.0 ± 8.3
0.081 (0.018)
60.0 ± 15.5
0.094 (0.011)
TABLE VII: Sim-to-real results. Real-world success rates are mean ± standard error, and all parenthetical values are 95% confidence intervals.
Fig. 5: Franka robot with color-coded joints (left) and corresponding raw policy outputs during Reach (right). After each target change (dashed lines), CATS (bottom) responds with smaller action changes and quickly settles to stable outputs, whereas PPO (top) exhibits larger action changes followed by persistent oscillations.
This paper proposes a novel regularization design to effectively smooth policy functions in reinforcement learning. While regularization that enhances global'' Lipschitz continuity was initially considered, it has been limited to local'' Lipschitz continuity due to a tradeoff between smoothness and expressiveness. However, it has become apparent that the original implementation is cumbersome and does not provide sufficient smoothing, leading to a preference for simpler implementations. This stems from a discrepancy between theory and implementation, and a more appropriate implementation can expect to facilitate smoothing. Therefore, this paper identifies three reasons why the original implementation does not function adequately and provide remedies for them. This modified regularization performs well across multiple tasks and algorithms, successfully achieving smooth motion while improving control performance. Furthermore, by applying it to sim-to-real reinforcement learning for a quadruped robot, it is demonstrated that smooth motion provides robustness against sudden changes in target velocity commands.
Taisuke Kobayashi, Naoto Yamanaka
National Institute of Informatics (NII) and with The Graduate University for Advanced Studies (SOKENDAI), 2-1-2 Hitotsubashi, Chiyoda-ku, Tokyo, 101-8430, Japan
Reinforcement learning often produces high-frequency oscillatory control signals that undermine the safety and stability required for physical deployment. Explicit action chunking addresses this by predicting fixed-horizon trajectories but scales the policy output dimension proportionally with the horizon length, leading to optimization difficulties and incompatibility with standard step-wise interaction. To overcome these challenges, this paper proposes Dual-Window Smoothing (DWS), an implicit action chunking framework for smooth continuous control. Unlike explicit methods, DWS enforces temporal coherence without expanding the action space. It uses a dual-window design: an execution window that ensures physical smoothness through deterministic modulation, and a value window that aligns temporal-difference targets over the horizon to correct critic bias caused by open-loop execution. DWS also includes a lightweight actor-side temporal regularizer based on first-order action differences to promote global continuity. This design effectively bridges the gap between temporal abstraction and reactive step-wise control. Experiments on benchmarks including the DeepMind Control Suite and industrial energy management tasks show that DWS outperforms state-of-the-art (SOTA) baselines. In complex vision-based autonomous driving tasks, DWS achieves smoother control, safer behavior with reduced jitter, and attains a 100% success rate.
Bosun Liang, Shuo Pei, Zirui Chen +5
Department of Data and Systems Engineering, The University of Hong Kong, Hong Kong SAR, China · Beijing Institute of Technology, Zhuhai, China · College of Computer Science, Sichuan University, Chengdu, China
Actor-critic methods achieve strong performance in continuous control, but their policies can produce highly oscillatory actions. A common remedy is to add auxiliary smoothness losses. However, their contribution can be negligible when their gradients are small relative to the native actor gradient. Moreover, existing methods often combine multiple auxiliary losses, complicating loss balancing without necessarily improving the return-smoothness trade-off. We introduce DAMPER (Direction-Aware Magnitude-Controlled Projection with Explicit Return Priority), which combines the native actor gradient with a temporal-consistency gradient through conflict-conditioned projection and adaptive magnitude control. It removes the auxiliary component opposing the actor gradient and scales the retained temporal direction relative to the actor gradient norm, preserving positive alignment with the native actor gradient. Experiments with TD3 and SAC on six continuous-control tasks show reduced action oscillation relative to the native agents in all 12 task-backbone pairs and the best oscillation score among the compared methods in eight, with task-dependent return trade-offs.