cs.AIApr 20, 2023

Topology-Guided Modular Actor-Critic Learning for Continuous Systems under Temporal Objectives

Authors: Lening Li, Zhentian Qian, Jianan Xia, Yawen Wang, Zhongjing Li, Qiren Geng, Huasheng Zhang, Liang Hu, +2 more

Organizations: Ningbo University, Ningbo, China · Robotics Engineering Program, Worcester Polytechnic Institute, Worcester, MA, USA · Ningbo Zhenhai Silver-Ball Bearing Co., Ltd., Ningbo, China

Abstract

We study formal policy synthesis for continuous-state stochastic systems under linear temporal logic specifications. The product of the system with the automaton of the specification has a hybrid state space with sparse rewards. We introduce a generalized optimal backup order, defined in reverse to a topological order over automaton states, that guides value backups and provably preserves optimality. We further present a model-free actor-critic algorithm whose policy evaluation solves a constrained optimization problem by the augmented Lagrangian method, yielding hyperparameter self-tuning, and prove its optimality and convergence in the tabular case. Since integer encodings of automaton states impose a spurious ordinal relationship on functions learned by one network, we dedicate a value and a policy network to each automaton state (modular learning). The algorithm matches or outperforms PPO, DQN, and A2C on CartPole, and on a Dubins car under a temporal specification the topological order and modular learning raise the success rate from 26.0% to 71.5%.

Explore similar work

CardsList