cs.LGOct 1, 2026

Bellman Meets Lyapunov: Unsupervised Reinforcement Learning via Mastering Chaos

Authors: Tristan Shah, Wooyoung Chung, Volodomyr Makarenko, Juan Wachs, Stas Tiomkin

Organizations: Texas Tech University · Purdue University

Abstract

Reinforcement learning (RL) is a powerful paradigm for training agents, yet its success rests on domain expertise of human engineers who design informative reward signals for every new task. Unsupervised RL aims to reduce this engineering with intrinsic motivation (IM): reward signals that emerge from the agent environment interaction itself. Existing IM objectives, however, involve the selection of information variables, which re-introduces domain expertise the field has sought to eliminate. We introduce Forward CIP (F-CIP), an RL-native formulation of the Controllable Information Production (CIP) objective, which is defined by the system's dynamics alone and requires no such selection. We prove that F-CIP is compatible with RL and demonstrate its effectiveness with existing algorithms. Training agents with F-CIP results in unsupervised discovery of primitive behaviors such as balancing and maintaining controllability, which are essential for more complex robot behaviors. Paired with a simple forward-velocity reward, our method produces coordinated gaits such as hopping and running which otherwise require reward engineering to learn.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Stability of Control Lyapunov Function Guided Reinforcement Learning

    May 3, 2026Zachary Olkin, William D. Compton, Aaron D. AmesReinforcement Learning ControlLyapunov Function

  2. Can We Really Learn One Representation to Optimize All Rewards?

    Feb 11, 2026Chongyi Zheng, Royina Karegoudra Jayanth, Benjamin EysenbachRepresentation LearningLow-Rank Structure

  3. Learning to Perceive the World Through Control: Empowerment-Based Representation Learning

    May 28, 2026Mahsa Bastankhah, Sophie Broderick, Benjamin EysenbachRepresentation LearningWorld