Reward Function
Reward functions, crucial for guiding reinforcement learning agents towards desired behaviors, are the focus of intense research. Current efforts center on automatically learning reward functions from diverse sources like human preferences, demonstrations (including imperfect ones), and natural language descriptions, often employing techniques like inverse reinforcement learning, large language models, and Bayesian optimization within various architectures including transformers and generative models. This research is vital for improving the efficiency and robustness of reinforcement learning, enabling its application to complex real-world problems where manually designing reward functions is impractical or impossible. The ultimate goal is to create more adaptable and human-aligned AI systems.
Papers
Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models
Roberto-Rafael Maura-Rivero, Chirag Nagpal, Roma Patel, Francesco Visin
Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
Yu Ishihara, Noriaki Takasugi, Kotaro Kawakami, Masaya Kinoshita, Kazumi Aoyama
Fantastic LLMs for Preference Data Annotation and How to (not) Find Them
Guangxuan Xu, Kai Xu, Shivchander Sudalairaj, Hao Wang, Akash Srivastava
Show, Don't Tell: Learning Reward Machines from Demonstrations for Reinforcement Learning-Based Cardiac Pacemaker Synthesis
John Komp, Dananjay Srinivas, Maria Pacheco, Ashutosh Trivedi