PlatoLTL: Scaling LTL-Guided Multi-Task RL
Organizations: Oxford Robotics Institute University of Oxford Oxford, United Kingdom · Department of Computer Science University of Oxford Oxford, United Kingdom
Abstract
Linear temporal logic (LTL) has emerged as a powerful formalism for specifying structured, temporally extended tasks in multi-task reinforcement learning (RL). However, while existing approaches in LTL-guided multi-task RL demonstrate success in simple environments, they suffer from challenges in representation learning and exploration efficiency when applied to high-dimensional environments with parameterized specifications. We present PlatoLTL, which elevates state-of-the-art methods to address both challenges. We model atomic propositions as instances of atomic predicates and inject task parameters directly into the goal embedding to enable efficient generalization. We also leverage simple priors on closeness to predicate satisfaction to enable and accelerate learning of complex tasks without biasing the optimal policy. We validate our approach on challenging environments including robotic manipulation and multi-drone navigation.
Figures & tables
| Success Rate (%) ( ) | Average Episode Length on Success ( ) | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Discretize | Concatenate | SIAMS | Plato | Plato + P 2 BRS | Discretize | Concatenate | SIAMS | Plato | Plato + P 2 BRS | ||
| Go1RGBZoneEnv | |||||||||||
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Continuous | Discrete | ||||
| Parameter | Go1RGBZoneEnv | X2DronesEnv | FrankaCubesEnv | FalloutWorld | |
| PPO | Number of parallel envs | ||||
| Steps per env per update | |||||
| Epochs | |||||
| Batch size | |||||
| Discount factor | |||||
| Average Visits ( ) | ||||||
|---|---|---|---|---|---|---|
| Discretize | Concatenate | SIAMS | Plato | Plato + P 2 BRS | ||
| Go1RGBZoneEnv | ||||||