cs.LGMay 7, 2024

Policy Learning with a Language Bottleneck

Authors: Megha Srivastava, Cedric Colas, Dorsa Sadigh, Jacob Andreas

Organizations: Stanford University · Massachusetts Institute of Technology Inria

Abstract

Modern AI systems such as self-driving cars and game-playing agents can achieve superhuman performance, but often lack human-like generalization, interpretability, and inter-operability with human users. Inspired by the rich interactions between language and decision-making in humans, we introduce Policy Learning with a Language Bottleneck (PLLB), a framework enabling AI agents to generate linguistic rules that capture the high-level strategies underlying rewarding behaviors. PLLB alternates between a rule generation step guided by language models, and an update step where agents learn new policies guided by rules, even when a rule is insufficient to describe an entire complex policy. Across five diverse tasks, including a two-player signaling game, maze navigation, image reconstruction, and robot grasp planning, we show that PLLB agents are not only able to learn more interpretable and generalizable behaviors, but can also share the learned rules with human users, enabling more effective human-AI coordination. We provide source code for our experiments at https://github.com/meghabyte/bottleneck .

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. NLPG: Natural-Language Policy Gradients for Self-Evolving Language Agents

    Sep 27, 2026Xu Liu, WenZhang Wei, Jun Cao +4Continual Learning for LLM AgentsPolicy Optimization