cs.LGSep 8, 2026

Unthrottling the Tanh Jacobian in SAC: A Negative Result on Bang-Bang Control and MetaDrive

Authors: Faiq Shamass

Abstract

Soft Actor-Critic (SAC) represents a continuous policy as an unbounded Gaussian that is squashed by tanh. The Jacobian of that map is ∂a/∂u=1−a2\partial a/\partial u = 1-a^2, which vanishes as ∣a∣→1|a|\to 1. A natural concern is that this throttle starves the actor of critic signal exactly where extreme actions (full brake, full throttle) are optimal. We test a minimal intervention that restores the missing signal: one extra term in the actor loss whose gradient on the pre-tanh mean is the detached action-gradient of QQ, with no gain parameter. On a minimum-time double integrator whose optimum is bang-bang at the action bounds, vanilla SAC already reaches near-optimal return (−31.6-31.6 vs. a calibrated optimum of −30.3-30.3) across ten paired seeds. An ungated bypass does saturate the policy (99% of eval steps with ∣a∣≥0.9|a|\ge 0.9) and collapses return to −195.5-195.5. A gated bypass that fires only on the flat shoulder ∣a∣∈[0.9,0.999]|a|\in[0.9,0.999] also fails, and does so without leaving a saturated policy. Warm-started MetaDrive fine-tuning shows the same pattern: the bypass does not improve return, and where collision rate falls it is typically traded for out-of-road departures. Auto-tuned entropy coefficient rises against the bypass, which is a push toward the tails. The Jacobian effect is real. Treating it as a bug to be undone is not free, and on the tasks studied here it is not helpful. Saturating a bound is not the same as solving a problem whose optimum lives on that bound.

Explore similar work

CardsList