cs.LGOct 4, 2026
SaveArithmetic Actor Heads and Training Stabilization for Out-of-Distribution Reinforcement Learning
Organizations: Central South University
Abstract
Reinforcement learning (RL) policies can deteriorate under out-of-distribution (OOD) magnitude shifts. Starting from soft actor-critic (SAC) and its Bayesian Amnesic Piecewise-Robust (BAPR) predecessor, we study the causal-symbolic BAPR (CS-BAPR) family. The practical method combines six training-stabilization settings with alternative actor heads: a Neural Addition Unit (NAU) with a Neural Multiplication Unit (NMU)-inspired quadratic correction, a Kolmogorov-Arnold Network (KAN), or a multilayer perceptron (MLP) with rectified linear unit (ReLU) or hyperbolic-tangent activations.
Figures & tables
| Method | seeds | ID total return (K) |
|---|---|---|
| bapr-pristine (ReLU no fixes) | 10 | |
| bapr (historical legacy) | 10 | |
| csbapr-no-nau (ReLU-MLP 6 fixes) | 10 | |
| csbapr (NAU 6 fixes) | 10 | |
| csbapr-tanh (smooth-MLP 6 fixes) | 10 | |
| csbapr-kan (KAN 6 fixes) | 10 |
Table 1: ID total return across the 12 lines at . Higher is better. Mean standard deviation, seeds per row.
| Method | ||||||
|---|---|---|---|---|---|---|
| bapr-pristine (ReLU, no fixes) | ||||||
| bapr (historical legacy) | ||||||
| csbapr-no-nau (ReLU-MLP 6 fixes) | ||||||
| csbapr (NAU 6 fixes) | ||||||
| csbapr-tanh (smooth-MLP 6 fixes) | ||||||
| csbapr-kan (KAN 6 fixes) |
Table 2: Static OOD demand sweep, total episode return (K, higher is better). Median episode return over seeds per row. is ID. bapr-pristine disables all six stabilization settings; bapr is a historical legacy configuration. Bold : best in column.
| Method | burst | burst | burst | burst |
|---|---|---|---|---|
| bapr-pristine | ||||
| bapr (legacy) | ||||
| csbapr-no-nau | ||||
| csbapr (NAU) | ||||
| csbapr-tanh | ||||
| csbapr-kan |
Table 3: Abrupt-burst protocol: for s, then abrupt switch to the specified multiplier. Total episode return (K), median over seeds.
| Method | commuter_day | square_wave_20x | escalating_10x |
|---|---|---|---|
| bapr-pristine | |||
| bapr (legacy) | |||
| csbapr-no-nau | |||
| csbapr (NAU) | |||
| csbapr-tanh | |||
| csbapr-kan |
Table 4: Oscillating OOD protocol: within-episode piecewise-constant demand schedule. Total episode return (K), median over seeds. Schedules: commuter_day has 5 switches with peak ; square_wave_20x has 9 switches with peak ; escalating_10x has 4 switches with peak .
| Method | K | K | K | K | Worst seed (K) | |
|---|---|---|---|---|---|---|
| bapr-pristine (ReLU, no fixes) | 10 | |||||
| bapr (historical legacy) | 10 | |||||
| csbapr-no-nau (ReLU-MLP 6 fixes) | 10 | |||||
| csbapr (NAU 6 fixes) | 10 | |||||
| csbapr-tanh (smooth-MLP 6 fixes) | 10 | |||||
| csbapr-kan (KAN 6 fixes) | 10 |
Table 5: Crash rate at across multiple severity thresholds. Each cell reports the fraction of seeds whose return at falls below the threshold (lower is better). All six methods use seeds. Thresholds show both typical poor performance and severe tail outcomes.
Figure 1: Static-demand return degradation relative to each configuration’s ID median. Every point is computed from Table 2 , with for each underlying median. Lower degradation is better relative to that configuration’s own ID level. The figure reports a difference of aggregate medians and has no seed-level uncertainty bands.
| Method | ||||||
|---|---|---|---|---|---|---|
| bapr (no SINDy, no NAU) | – | K | M | – | M | |
| csbapr-no-sindy (NAU, no SINDy) | – | K | M | – | M | |
| csbapr-no-nau (MLP SINDy) | – | K | M | – | M | |
| csbapr (NAU SINDy, full) | – | K | M | – | M |
Table 6: Exploratory LQR initial-amplitude sweep, median unaugmented episode return over seeds. Higher is better. K and M denote and return units; – denotes an unreported cell. The symbolic-enabled configurations were labeled with SINDY_WITH_CONTROL=1 ; the implementation limitation is discussed in Section 7.1 . Bold : best reported value in the column.
| Method | ID ( ) | |||||
|---|---|---|---|---|---|---|
| bapr-pristine (ReLU, no fixes) | ||||||
| bapr / csbapr-relu (legacy, ) | ||||||
| csbapr-tanh (smooth-MLP 6 fixes) | ||||||
| csbapr-no-nau (ReLU-MLP 6 fixes) | ||||||
| csbapr (NAU 6 fixes) |
Table 7: Hopper-v4 wind sweep, median episode return over seeds except the legacy comparison ( , interrupted runs excluded). The pristine row disables the six stabilization settings; the legacy row is retained as a historical comparison. Wind amplitude is ID. Bold : best in the column.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Quantity | Value |
|---|---|
| Discount | |
| Target interpolation | |
| Initial entropy temperature | , automatically tuned |
| Optimizer step size | |
| Critics ; batch size | ; transit critics: |
| Run-length count ; hazard |
Table 8: Shared defaults and the primary transit overrides.
Explore similar work
Real-world control systems frequently operate under \emph{piecewise stationary} conditions, where dynamics remain stable for extended periods before undergoing abrupt regime changes. Standard robust RL methods face a fundamental dilemma: a globally conservative policy wastes performance during stable periods, while a locally adaptive policy risks catastrophic failure when the regime changes undetected. We propose \textbf{BAPR} (Bayesian Amnesic Piecewise-Robust SAC), which unifies Bayesian Online Change Detection (BOCD) with robust ensemble RL. The BAPR operator -- a convex combination of mode-conditional Bellman operators weighted by a frozen belief distribution -- is a -contraction. A complementary counterexample, machine-verified in Lean4, establishes a \emph{sharp boundary}: when beliefs depend on the Q-function, the contraction factor becomes (where is the mode reward gap), and contraction fails exactly when . We derive a \emph{component-wise} formal error budget for the abstract operator -- every component machine-verified -- bounding post-switch recovery; the budget applies to the abstract mode-mixture operator and inherits to the implemented shared-critic algorithm only through the frozen-parameter design intuition. All results are formally verified with no \texttt{sorry} (1,145 lines across 3 Lean4 files, 22 machine-verified theorems). BOCD drives an adaptive conservatism mechanism: the policy becomes maximally conservative after detected change-points and smoothly relaxes as confidence grows, with detection delay . A context-conditioning module trained via RMDM loss provides mode-aware representations from simulator-provided mode IDs at training time and requires no mode labels at deployment.
AdamO: A Collapse-Suppressed Optimizer for Offline RL
Offline reinforcement learning (RL) can fail spectacularly when bootstrapped temporal-difference (TD) updates amplify their own errors, driving the critic toward extreme and unusable Q-values. A key counterintuitive insight of this work is that collapse is not only a property of the backup rule or network architecture: optimizer dynamics themselves can directly trigger or suppress instability. From a control-theoretic viewpoint, we model offline TD learning as a feedback system and analyze Adam-based critic updates. This yields a necessary and sufficient condition for stability of the induced local update dynamics: within the regime we analyze, these dynamics are stable if and only if the spectral radius of the corresponding update operator is strictly below one. Further analysis suggests that standard Adam updates can inadvertently distort the parameter geometry, motivating explicit orthogonality constraints to prevent TD error amplification. To this end, we propose AdamO, an Adam-based optimizer with a decoupled orthogonality correction regulated by a strict task-alignment budget. We prove that this design theoretically guarantees worst-case task safety and preserves Adam's continuous-time dissipative dynamics. Empirically, AdamO is broadly compatible with diverse offline RL baselines, improving stability and returns across a broad suite of benchmarks.
Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning
Offline reinforcement learning (RL) faces a critical challenge of overestimating the value of out-of-distribution (OOD) actions. Existing methods mitigate this issue by penalizing unseen samples, yet they fail to accurately identify OOD actions and may suppress beneficial exploration beyond the behavioral support. Although several methods have been proposed to differentiate OOD samples with distinct properties, they typically rely on restrictive assumptions about the data distribution and remain limited in discrimination ability. To address this problem, we propose DOSER (Diffusion-based OOD Detection and Selective Regularization), a novel framework that goes beyond uniform penalization. DOSER trains two diffusion models to capture the behavior policy and state distribution, using single-step denoising reconstruction error as a reliable OOD indicator. During policy optimization, it further distinguishes between beneficial and detrimental OOD actions by evaluating predicted transitions, selectively suppressing risky actions while encouraging exploration of high-potential ones. Theoretically, we prove that DOSER is a -contraction and therefore admits a unique fixed point with bounded value estimates. We further provide an asymptotic performance guarantee relative to the optimal policy under model approximation and OOD detection errors. Across extensive offline RL benchmarks, DOSER consistently attains superior performance to prior methods, especially on suboptimal datasets.