Policy Learning with a Language Bottleneck
Organizations: Stanford University · Massachusetts Institute of Technology Inria
Abstract
Modern AI systems such as self-driving cars and game-playing agents can achieve superhuman performance, but often lack human-like generalization, interpretability, and inter-operability with human users. Inspired by the rich interactions between language and decision-making in humans, we introduce Policy Learning with a Language Bottleneck (PLLB), a framework enabling AI agents to generate linguistic rules that capture the high-level strategies underlying rewarding behaviors. PLLB alternates between a rule generation step guided by language models, and an update step where agents learn new policies guided by rules, even when a rule is insufficient to describe an entire complex policy. Across five diverse tasks, including a two-player signaling game, maze navigation, image reconstruction, and robot grasp planning, we show that PLLB agents are not only able to learn more interpretable and generalizable behaviors, but can also share the learned rules with human users, enabling more effective human-AI coordination. We provide source code for our experiments at https://github.com/meghabyte/bottleneck .
Figures & tables
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | |||
|---|---|---|---|
| Standard | I should follow the strategy of choosing action 4 whenever the observation is 2, 3, 4, or 5. | I should follow the strategy of choosing actions based on the observation provided by AGENT 1. If the observation is 1, I should choose action 1. If the observation is 2, I should choose action 3. If the observation is 3, I should choose action 3. If the observation is 4, I should choose action 4. If the observation is 5, I should choose action 5. | I should follow the strategy of choosing the same action as AGENT 1 for observations 1, 2, 3, 4, and 5. |
| Fixed Speaker | I should choose action 1 when observation is 1 or 2 or 4 or 5. I should choose action 2 when observation is 3. | I should choose action 1 when observation is 1, 2, 3 or 5. I should choose action 2 when observation is 4. | I should choose action 1 when observation is 1. I should choose action 4 when observation is 2. I should choose action 5 when observation is 3. I should choose action 2 when observation is 4. I should choose action 3 when observation is 5. |
| Size | ||||
| 7x7 | I should start by moving north, then continue moving north until I reach a wall, then turn east and move until I reach a wall, then turn south and move until I reach a wall, then turn west and move until I reach the end of the maze. | I should prioritize moving North, East, and South over moving West, as the HIGH REWARD EXAMPLES show a higher frequency of these actions. 2. I should avoid moving West unnecessarily, as it is rarely seen in the HIGH REWARD EXAMPLES and often leads to LOW REWARD EXAMPLES | I should take the EAST action when I see WHITE, and then repeat the EAST action until I see RED, then take the SOUTH action, and then repeat the EAST action until I see BLUE, then take the NORTH action, then take the EAST action until I see RED again, then take the SOUTH action, and finally take the WEST action. | I should take the EAST action when I observe WHITE, and keep taking EAST actions until I observe RED, then take the SOUTH action. After taking the SOUTH action, I should take the EAST action again until I observe BLUE, then take the NORTH action. After taking the NORTH action, I should take the EAST action until I observe RED again, then take the SOUTH action. |
| Setting | |||
|---|---|---|---|
| Re-Initialization | 1. Be specific with location details: Agent 1 should provide detailed location information for each element in the image, such as x and y coordinates. 2. Use descriptive language for elements, such as "red dot" or "green triangle". | 1. Use specific coordinates when instructing Agent 2 to draw shapes. 2. Use descriptive language to specify the color and shape of each element. For example, "a green triangle" instead of "a green thing". | 1. Be specific with location coordinates: provide specific coordinates for the location of each shape, using the format x=0.XX, y=0.YY. 2. Use descriptive shape names: Instead of using generic terms like "dot" or "square," use more descriptive names that indicate the shape’s color and size, such as "green triangle" or "red square." |
| Continual Training | 1. Be specific and detailed in your instructions. High reward examples have specific coordinates and shapes, while low reward examples have more general descriptions. 2. Use a consistent format for your instructions. High reward examples have a consistent format for listing coordinates and shapes, while low reward examples have a more free-form format. | 1. Provide explicit coordinates for each element in the image, using the format (x, y). 2. Use specific colors when referring to elements in the image, such as "red", "green", or "blue". Avoid using vague terms like "colored" or "shaded". | 1. Use a consistent format for describing shapes, such as always listing the x-coordinate first, followed by the y-coordinate. For example, instead of "one green square at the point x=0.53, y=0.24", use "one green square at (0.53, 0.24)". 2. Avoid using vague terms like "various shades of green". Instead, use specific colors, such as "green" or "blue". Additionally, use specific shapes, such as "square" or "triangle", rather than vague terms like "rectangle". |
| Reward | |||
|---|---|---|---|
| color | Describe the bird’s color, species, and any distinctive markings or patterns. | Describe the bird’s coloration accurately. | Describe the bird’s coloration accurately. |
| background | Include details about the bird’s surroundings, such as the type of branch or post it is on, and any additional elements in the background. | Include the bird’s action (perched, flying, standing) and its location (on a branch, railing, pole, etc.) | Describe the bird’s action (flying, perching, standing) and the environment it is in (sky, tree, water). |
| species | Describe the bird’s color, markings, and any distinctive features. | Describe the subject’s unique features, such as coloration, beak shape, or other distinguishing characteristics. | Include specific details about the bird’s appearance, such as the color of its feathers, beak, or eyes, and any distinctive markings or patterns. |
| Reward | |||
| Shoulder Torque | Select keypoints that are away from sharp edges or corners of the object to avoid potential damage and improve stability during grasping. | Prioritize grasp keypoints that are closer to the center of the object for more stable and precise pickup. | Prioritize grasping at the middle of the object’s surface rather than the top edges or handles to ensure a stable grip. |
| Episode | 500 | 1000 | 6000 |
|---|---|---|---|
| Reward ( Original, temp=0.5 ) | |||
| Reward ( Original, temp=0.9 ) | |||
| Reward ( Original, temp=0.1 ) | |||
| Reward ( No Format Instruction ) | |||
| Reward ( Low Context ) | |||
| Reward ( Rephrase ) |
| Episode | 500 | 1000 | 6000 |
|---|---|---|---|
| Interpretability ( Original, temp=0.5 ) | |||
| Interpretability ( Original, temp=0.9 ) | |||
| Interpretability ( Original, temp=0.1 ) | |||
| Interpretability ( No Format Instruction ) | |||
| Interpretability ( Low Context ) | |||
| Interpretability ( Rephrase ) |
| Episode | 500 | 1000 | 6000 |
|---|---|---|---|
| Reward ( Original ) | |||
| Reward ( Non-contrastive ) | |||
| Reward ( Adversarial ) | |||
| Interpretability ( Original ) | |||
| Interpretability ( Non-contrastive ) |