Going Beyond State-Reaching: Learning Abstractions for Intrinsically Motivated Option Discovery
Organizations: Brown University
Abstract
Temporal abstraction via options can improve exploration in large environments. However, existing option discovery algorithms find subgoals that target all aspects of the state simultaneously. This state-reaching approach produces options that only apply in narrow regions of the state-space, eventually causing an explosion in the number of options that overwhelms the agent, and impedes progress on its primary task of reward maximization. We introduce an algorithm that instead identifies a small, relevant subset of features for each subgoal, yielding options that generalize broadly and accelerate exploration. Our approach learns abstract, transferrable options and achieves rapid exploration in three sparse-reward, image-based domains, including the Atari game MontezumasRevenge.
Figures & tables
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning | Introduced |
| MDP and options | ||
| Markov decision process | Sec. 2 | |
| Option: initiation function, policy, termination condition | Sec. 2 | |
| The agent’s set of discovered options | Sec. 2 | |
| Option timeout (maximum steps per execution) | Sec. 2 | |
| Option pseudo-reward | Sec. 2 | |
| Hyperparameter | Value |
|---|---|
| R2D2 agent | |
| Number of actors | 32 |
| Actor backend | CPU |
| Sequence batch size | 32 |
| Trace length | 40 |
| Sequence period | 20 |
| Domain | Method | LR | TUP | SPI | Reward tx. | Intrinsic coeff. | CFN LR | CFN SPI | CFN min. replay | |
|---|---|---|---|---|---|---|---|---|---|---|
| Montezuma | CFN | 0.99 | 600 | 2 | signed-hyperbolic | 0.01 | Unlimited | 2,048 | ||
| R2D2 | 0.99 | 600 | 2 | signed-hyperbolic | 0.0 | — | — | — | ||
| KeyCorridor | CFN | 0.99 | 1,200 | 8 | identity | 0.001 | Unlimited | 12,500 | ||
| R2D2 | 0.99 | 600 | 2 | identity | — | — | — | — | ||
| VisualTaxi | CFN | 0.99 | 600 | 2 | identity | 0.03 | 8 | 2,048 | ||
| R2D2 | 0.99 | 600 | 2 | identity | — | — | — | — |
| Domain | Policy | LR | TUP | SPI | CFN LR | CFN TUP | CFN SPI | CFN min. replay | ||
|---|---|---|---|---|---|---|---|---|---|---|
| Montezuma | Intra-option ( ) | 0.997 | 600 | 2 | — | — | — | — | — | |
| Exploration ( ) | 0.99 | 600 | 8 | 0.01 | 600 | Unlimited | 12,500 | |||
| KeyCorridor | Intra-option ( ) | 0.997 | 500 | 2 | — | — | — | — | — | |
| Exploration ( ) | 0.99 | 600 | 8 | 0.01 | 600 | Unlimited | 12,500 | |||
| VisualTaxi | Intra-option ( ) | 0.997 | 600 | 2 | — | — | — | — | — | |
| Exploration ( ) | 0.99 | 600 | 8 | 0.01 | 600 | Unlimited | 12,500 |
| Hyperparameter | Montezuma | KeyCorridor | VisualTaxi |
| Option timeout, | 100 | 50 | 50 |
| Goal-creation threshold multiplier, | 1 | 1 | 1 |
| Number of goals to replay | 5 | 5 | 5 |
| Option-initiation threshold, | N/A (scaled by , Appendix C ) | 0.1 | 0.1 |
| Extrinsic-reward weight in , | 0.25 | 0 | 0 |
| Classifier match threshold, | (with ) | see Table 2 | see Table 2 |
| Hyperparameter | Value |
|---|---|
| Saliency method | DeepSHAP |
| SAM checkpoint | ViT-H ( sam_vit_h_4b8939.pth ) |
| Patch attribution threshold, | 1.0 |
| Number of attribution baselines, | 15 |
| Baseline window, | 50 |
| Classifier trigger window size, | 8 |