A Geometric Approach to Soft Actor-Critic with Zonotopes for Locomotion Learning
Authors: Panagiotis Roditis, Panagiotis P. Filntisis, Petros Maragos
Organizations: Robotics Institute, Athena Research Center, Marousi, Greece · HERON - Hellenic Robotics Center of Excellence, Athens, Greece · School of Electrical & Computer Engineering, NTUA, Greece
Off-policy actor--critic methods control overestimation bias by taking the minimum of two critics. This uses the same aggregation rule everywhere, regardless of how the critics disagree. We propose \textbf{GeZo-SAC}, which uses auxiliary geometric representations to adapt critic pessimism to the state and action. Alongside its scalar value, each critic predicts a set of generators defining a zonotope. Probing this zonotope along sampled directions provides a geometric width, "subtracted from each critic value as a pessimistic offset, and a measure of disagreement between the two critics, aggregated with log-sum-exp. This disagreement controls how the critics are combined, moving from a width-weighted average toward the usual minimum as disagreement increases. At inference, the deployed policy is an unmodified SAC actor, since the generators are used only on the critic side during training.Across four MuJoCo-v5 locomotion benchmarks and six off-policy baselines, GeZo-SAC achieves the highest mean return on Ant-v5 and Hopper-v5 and remains competitive with other methods on the remaining tasks. Our analysis further shows that GeZo-SAC achieves the lowest average actuator work and action effort per metre among the evaluated methods, while maintaining near-zero measured overestimation frequency across all four environments.
Figures & tables
Fig. 1 : Training performance on four MuJoCo-v5 locomotion environments. Learning curves show episodic training returns for Ant-v5, Hopper-v5, Walker2d-v5, and Humanoid-v5. Per-run returns are smoothed and interpolated onto a common environment-step grid. Curves represent the mean over five independent runs ( n=5 ), and shaded regions denote ± one standard deviation across runs. GeZo-SAC leads on Ant-v5 and remains competitive on the other tasks.
Environment
GeZo-SAC (ours)
TQC
TD3
SAC
GPL-SAC
TDC- λ
SUNRISE
Ant-v5
7.17(0.21)
4.45(2.05)
5.77(0.23)
4.43(1.24)
6.09(0.52)
5.39(1.53)
5.32(0.77)
Hopper-v5
3.77(0.11)
3.44(0.38)
3.52(0.09)
3.51(0.17)
3.38(0.21)
3.59(0.15)
3.63(0.08)
Humanoid-v5
7.89(0.58)
8.61(0.25)
5.23(0.20)
5.60(0.09)
6.13(0.66)
6.85(0.54)
0.62(0.08)
Walker2d-v5
5.74(0.30)
6.22(0.41)
4.61(0.52)
5.25(0.47)
6.06(0.47)
4.99(0.85)
5.79(0.44)
TABLE I: Mean evaluation return (in thousands) across five training runs, using the best checkpoint from each run and 300 deterministic evaluation episodes. Parentheses show the sample standard deviation. Bold and underline denote the best and second-best means. GeZo-SAC uses K=15 in all environments.
Method
Work/m ↓
Dist. ↑
TV/m ↓
Effort/m ↓
Return ↑
SAC
1.028
0.817
1.266
1.100
4.70 (0.32)
TD3
1.275
0.786
1.951
1.936
4.77 (0.16)
TQC
1.137
1.283
0.941
0.937
5.68 (0.51)
GPL-SAC
0.934
0.988
0.901
0.926
5.42 (0.23)
TDC- λ
1.068
1.002
1.463
1.441
5.21 (0.44)
SUNRISE
1.176
0.751
1.890
2.551
3.84 (0.23)
TABLE II : Mechanical metrics and return across the four locomotion tasks. Values are averaged across environments; mechanical metrics are normalized within each environment by its median Lower is better for Work/m, TV/m, and Effort/m; higher is better for Distance and return. Bold and underline denote the best and second-best values.
Fig. 2 : Performance–energy trade-off. Mean episodic return versus mean absolute mechanical work per metre, Eabs/m (J/m), across four locomotion tasks. Each marker represents one method; points toward the upper left combine higher return with lower work per metre.
Fig. 3 : Locomotion and mechanical cost on Ant-v5. Top: Overlaid poses of TQC, GPL-SAC, and GeZo-SAC at matched times over a one-second interval, shown at a common spatial scale. Bottom: Cumulative actuator work versus forward displacement over the same interval. GeZo-SAC travels farther while requiring less work per metre in these illustrative rollouts.
Fig. 4 : Sensitivity to the number of zonotope generators. Curves show mean evaluation return over five training seeds, with shaded regions and error bars indicating ± one standard deviation. Bold labels mark the highest observed mean in each environment.
Fig. 5 : Value-estimation diagnostics and policy performance. Comparison of GeZo-SAC, SAC, TD3, TQC, and GPL-SAC across four locomotion tasks. Left: measured overestimation frequency P(Q^(s,aπ)>MC(s)) . Right: mean absolute critic–return discrepancy E[∣Q^(s,aπ)−MC(s)∣] . The vertical axis shows mean discounted, reward-only Monte-Carlo return under deterministic evaluation. Markers average policy-level metrics across training seeds; horizontal and vertical error bars show ± one sample standard deviation.
Fig. 6 : Effect of fixed versus adaptive pessimism. Evaluation returns for GeZo-SAC with fixed κ=1 , fixed κ≈0 , and the adaptive κ controller across the four locomotion tasks. Curves show the mean over three training runs, with shading indicating ± one sample standard deviation. Diamonds mark the final deterministic evaluation over 300 episodes per policy, with error bars showing the standard deviation across runs.
Fig. 7 : Evolution of the adaptive pessimism coefficient on Hopper-v5. The adaptive controller changes κ throughout training, while the two ablated variants keep it fixed. The curve shows the mean across three training seeds, with shading indicating ± one standard deviation.
Method
Act. (M)
Crit. (M)
Train (ms)
Infer. (ms)
SAC
0.101
0.199
9.83
0.181
TD3
0.099
0.199
5.25
0.161
TQC
0.101
0.527
16.90
0.181
GPL-SAC
0.101
1.656
41.62
0.662
TDC- λ
0.099
0.231
6.65
0.153
SUNRISE
0.506
0.993
55.63
0.900
TABLE III : Computational overhead. Parameter counts include the actor and all online critics. Timings calculated on an NVIDIA A100-SXM-80GB.