cs.LGOct 4, 2026
SaveArithmetic Actor Heads and Training Stabilization for Out-of-Distribution Reinforcement Learning
Organizations: Central South University
Abstract
Reinforcement learning (RL) policies can deteriorate under out-of-distribution (OOD) magnitude shifts. Starting from soft actor-critic (SAC) and its Bayesian Amnesic Piecewise-Robust (BAPR) predecessor, we study the causal-symbolic BAPR (CS-BAPR) family. The practical method combines six training-stabilization settings with alternative actor heads: a Neural Addition Unit (NAU) with a Neural Multiplication Unit (NMU)-inspired quadratic correction, a Kolmogorov-Arnold Network (KAN), or a multilayer perceptron (MLP) with rectified linear unit (ReLU) or hyperbolic-tangent activations.
Figures & tables
| Method | seeds | ID total return (K) |
|---|---|---|
| bapr-pristine (ReLU no fixes) | 10 | |
| bapr (historical legacy) | 10 | |
| csbapr-no-nau (ReLU-MLP 6 fixes) | 10 | |
| csbapr (NAU 6 fixes) | 10 | |
| csbapr-tanh (smooth-MLP 6 fixes) | 10 | |
| csbapr-kan (KAN 6 fixes) | 10 |
Table 1: ID total return across the 12 lines at . Higher is better. Mean standard deviation, seeds per row.
| Method | ||||||
|---|---|---|---|---|---|---|
| bapr-pristine (ReLU, no fixes) | ||||||
| bapr (historical legacy) | ||||||
| csbapr-no-nau (ReLU-MLP 6 fixes) | ||||||
| csbapr (NAU 6 fixes) | ||||||
| csbapr-tanh (smooth-MLP 6 fixes) | ||||||
| csbapr-kan (KAN 6 fixes) |
Table 2: Static OOD demand sweep, total episode return (K, higher is better). Median episode return over seeds per row. is ID. bapr-pristine disables all six stabilization settings; bapr is a historical legacy configuration. Bold : best in column.
| Method | burst | burst | burst | burst |
|---|---|---|---|---|
| bapr-pristine | ||||
| bapr (legacy) | ||||
| csbapr-no-nau | ||||
| csbapr (NAU) | ||||
| csbapr-tanh | ||||
| csbapr-kan |
Table 3: Abrupt-burst protocol: for s, then abrupt switch to the specified multiplier. Total episode return (K), median over seeds.
| Method | commuter_day | square_wave_20x | escalating_10x |
|---|---|---|---|
| bapr-pristine | |||
| bapr (legacy) | |||
| csbapr-no-nau | |||
| csbapr (NAU) | |||
| csbapr-tanh | |||
| csbapr-kan |
Table 4: Oscillating OOD protocol: within-episode piecewise-constant demand schedule. Total episode return (K), median over seeds. Schedules: commuter_day has 5 switches with peak ; square_wave_20x has 9 switches with peak ; escalating_10x has 4 switches with peak .
| Method | K | K | K | K | Worst seed (K) | |
|---|---|---|---|---|---|---|
| bapr-pristine (ReLU, no fixes) | 10 | |||||
| bapr (historical legacy) | 10 | |||||
| csbapr-no-nau (ReLU-MLP 6 fixes) | 10 | |||||
| csbapr (NAU 6 fixes) | 10 | |||||
| csbapr-tanh (smooth-MLP 6 fixes) | 10 | |||||
| csbapr-kan (KAN 6 fixes) | 10 |
Table 5: Crash rate at across multiple severity thresholds. Each cell reports the fraction of seeds whose return at falls below the threshold (lower is better). All six methods use seeds. Thresholds show both typical poor performance and severe tail outcomes.
Figure 1: Static-demand return degradation relative to each configuration’s ID median. Every point is computed from Table 2 , with for each underlying median. Lower degradation is better relative to that configuration’s own ID level. The figure reports a difference of aggregate medians and has no seed-level uncertainty bands.
| Method | ||||||
|---|---|---|---|---|---|---|
| bapr (no SINDy, no NAU) | – | K | M | – | M | |
| csbapr-no-sindy (NAU, no SINDy) | – | K | M | – | M | |
| csbapr-no-nau (MLP SINDy) | – | K | M | – | M | |
| csbapr (NAU SINDy, full) | – | K | M | – | M |
Table 6: Exploratory LQR initial-amplitude sweep, median unaugmented episode return over seeds. Higher is better. K and M denote and return units; – denotes an unreported cell. The symbolic-enabled configurations were labeled with SINDY_WITH_CONTROL=1 ; the implementation limitation is discussed in Section 7.1 . Bold : best reported value in the column.
| Method | ID ( ) | |||||
|---|---|---|---|---|---|---|
| bapr-pristine (ReLU, no fixes) | ||||||
| bapr / csbapr-relu (legacy, ) | ||||||
| csbapr-tanh (smooth-MLP 6 fixes) | ||||||
| csbapr-no-nau (ReLU-MLP 6 fixes) | ||||||
| csbapr (NAU 6 fixes) |
Table 7: Hopper-v4 wind sweep, median episode return over seeds except the legacy comparison ( , interrupted runs excluded). The pristine row disables the six stabilization settings; the legacy row is retained as a historical comparison. Wind amplitude is ID. Bold : best in the column.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Quantity | Value |
|---|---|
| Discount | |
| Target interpolation | |
| Initial entropy temperature | , automatically tuned |
| Optimizer step size | |
| Critics ; batch size | ; transit critics: |
| Run-length count ; hazard |
Table 8: Shared defaults and the primary transit overrides.