cs.LGOct 8, 2026

A Geometric Approach to Soft Actor-Critic with Zonotopes for Locomotion Learning

Authors: Panagiotis Roditis, Panagiotis P. Filntisis, Petros Maragos

Organizations: Robotics Institute, Athena Research Center, Marousi, Greece · HERON - Hellenic Robotics Center of Excellence, Athens, Greece · School of Electrical & Computer Engineering, NTUA, Greece

Abstract

Off-policy actor--critic methods control overestimation bias by taking the minimum of two critics. This uses the same aggregation rule everywhere, regardless of how the critics disagree. We propose \textbf{GeZo-SAC}, which uses auxiliary geometric representations to adapt critic pessimism to the state and action. Alongside its scalar value, each critic predicts a set of generators defining a zonotope. Probing this zonotope along sampled directions provides a geometric width, "subtracted from each critic value as a pessimistic offset, and a measure of disagreement between the two critics, aggregated with log-sum-exp. This disagreement controls how the critics are combined, moving from a width-weighted average toward the usual minimum as disagreement increases. At inference, the deployed policy is an unmodified SAC actor, since the generators are used only on the critic side during training.Across four MuJoCo-v5 locomotion benchmarks and six off-policy baselines, GeZo-SAC achieves the highest mean return on Ant-v5 and Hopper-v5 and remains competitive with other methods on the remaining tasks. Our analysis further shows that GeZo-SAC achieves the lowest average actuator work and action effort per metre among the evaluated methods, while maintaining near-zero measured overestimation frequency across all four environments.

Figures & tables

Explore similar work

CardsList
  1. Generative Actor-Critic with Soft Bridge Policies

    May 9, 2026Ke He, Le He, Shunpu Tang +2Policy OptimizationMaximum Entropy RL