Temporal abstraction via options can improve exploration in large environments. However, existing option discovery algorithms find subgoals that target all aspects of the state simultaneously. This state-reaching approach produces options that only apply in narrow regions of the state-space, eventually causing an explosion in the number of options that overwhelms the agent, and impedes progress on its primary task of reward maximization. We introduce an algorithm that instead identifies a small, relevant subset of features for each subgoal, yielding options that generalize broadly and accelerate exploration. Our approach learns abstract, transferrable options and achieves rapid exploration in three sparse-reward, image-based domains, including the Atari game MontezumasRevenge.
Figures & tables
Figure 2: Illustration of counterfactual substitution. The most boring and most novel states, sB and sN respectively, are input into the Counterfactual Generator, which outputs counterfactual states c1,c2,c3 —each with one feature reset to its value in sB . Only resetting the first feature (black → blue) substantially lowers novelty, indicating that the first feature accounts for most of the change in novelty Δn . This one-at-a-time reset is effective when features contribute independently.
Figure 3: Examples of identifying relevant features in MontezumasRevenge . For each state (black and white image), we show the estimated Shapley value attributed to each feature in the image and the bounding boxes selected for the subgoal classifier. The first classifier checks if the player has entered room 0, while the second classifier checks if the player has the sword near the skulls.
Figure 4: Example of a discovered subgoal classifier that transfers to many different rooms. The leftmost image shows the context in which this subgoal classifier was learned—the bounding box shows that the classifier only attends to the player’s position. The remaining three images show some states in which that classifier was triggered—as long as the player re-appears in that portion of the screen, other factors are all irrelevant for achieving that option’s subgoal.
Figure 5: Learning curves comparing our agent ( Abstract Subgoals ) with a non-hierarchical novelty maximizing RL method ( CFN ), a vanilla RL method ( R2D2 ), and a state-reaching HRL baseline ( Pixel Equality ). Solid lines denote average undiscounted return; shaded regions denote standard deviation over 5 random seeds.
Figure 6: Illustration of the core insight. Suppose an agent observes a state trajectory of length 7 and a novelty estimator assigns each state st in that trajectory a novelty score fϕint(st) . sN denotes the most novel state, and sB denotes the most boring state in that trajectory. Suppose that each state has 3 features; the colored boxes represent feature values for those two states. Rather than simply treating sN as a target state, we seek to identify which features in particular were responsible for the large difference in novelty Δn=fϕint(sN)−fϕint(sB) . Our algorithm can handle input images, but we show state features here for simplicity.
Figure 7: Hyperparameter sensitivity sweep in MontezumasRevenge . Each curve is the average undiscounted task return across 3 random seeds; shaded regions denote standard deviation.
Figure 8: Learning curves comparing Abstract Subgoals and Pixel Equality given subgoals from a separate discovery-only stage. Both agents are trained on the same fixed set of novel trigger states for MontezumasRevenge , with subgoal discovery disabled; each method constructs its own classifier for these states, so the comparison isolates the effect of the classifier alone (averaged over 4 seeds).
Figure 9: DeepSHAP attribution is broadly consistent between CFN and RND novelty estimators, but CFN attribution is more concentrated. Each row compares the most-novel state independently identified by RND and by CFN at a similar stage of exploration; because the two estimators judge novelty differently, these are not necessarily the same state. The left pair of columns shows RND’s most-novel state alongside its DeepSHAP attribution map, and the right pair shows CFN’s most-novel state alongside its attribution map. Noticeably, CFN’s attributions are more concentrated around a few features, while RND’s attributions are more diffused across all features. Despite this, the resulting classifiers attend to similar features.
Figure 10: Number of options discovered in MontezumasRevenge over training for the learning curves reported in Figure 5 . Solid line denotes the mean number of options discovered; shaded area denotes standard error.
Figure 11: Discovered subgoal classifiers in MiniGrid-KeyCorridor . Bounding boxes show the patches each classifier attends to; all other pixels are ignored. (Left) Subgoal: reach a location and open the yellow door. (Middle) Subgoal: reach the mid-left room while the key is in the hallway. (Right) Subgoal: enter the locked room containing the ball, regardless of key position or door configuration.
Hyperparameter
Value
R2D2 agent
Number of actors
32
Actor backend
CPU
Sequence batch size
32
Trace length
40
Sequence period
20
Appendix
Table 2: Hyperparameters shared across all agents, domains, and methods, unless overridden in Tables 3 – 5 .
Domain
Method
γ
LR
TUP
SPI
Reward tx.
Intrinsic coeff.
CFN LR
CFN SPI
CFN min. replay
Montezuma
CFN
0.99
1×10−4
600
2
signed-hyperbolic
0.01
1×10−3
Unlimited
2,048
R2D2
0.99
1×10−4
600
2
signed-hyperbolic
0.0
—
—
—
KeyCorridor
CFN
0.99
3×10−4
1,200
8
identity
0.001
1×10−4
Unlimited
12,500
R2D2
0.99
3×10−4
600
2
identity
—
—
—
—
VisualTaxi
CFN
0.99
3×10−4
600
2
identity
0.03
1×10−4
8
2,048
R2D2
0.99
3×10−4
600
2
identity
—
—
—
—
Appendix
Table 3: Flat RL baselines: core hyperparameters per domain. TUP = target update period; SPI = samples-per-insert ratio. “Reward tx.” is the transform applied to the R2D2 target (identity or signed-hyperbolic). “Intrinsic coeff.” is the coefficient on the CFN novelty bonus added to the training reward.
Domain
Policy
γ
LR
TUP
SPI
λ
CFN LR
CFN TUP
CFN SPI
CFN min. replay
Montezuma
Intra-option ( πo )
0.997
1×10−4
600
2
—
—
—
—
—
Exploration ( πint )
0.99
3×10−4
600
8
0.01
1×10−3
600
Unlimited
12,500
KeyCorridor
Intra-option ( πo )
0.997
1×10−4
500
2
—
—
—
—
—
Exploration ( πint )
0.99
3×10−4
600
8
0.01
1×10−3
600
Unlimited
12,500
VisualTaxi
Intra-option ( πo )
0.997
1×10−4
600
2
—
—
—
—
—
Exploration ( πint )
0.99
3×10−4
600
8
0.01
1×10−3
600
Unlimited
12,500
Appendix
Table 4: Abstract Subgoals / Pixel Equality: policy hyperparameters. TUP = target update period; SPI = samples-per-insert ratio. Pixel Equality uses identical values (it differs only on the classifier). λ is the CFN exploration-bonus coefficient from Section 3.1 ; it is 0.01 in all three domains.
Hyperparameter
Montezuma
KeyCorridor
VisualTaxi
Option timeout, H
100
50
50
Goal-creation threshold multiplier, σstate
1
1
1
Number of goals to replay
5
5
5
Option-initiation threshold, δ
N/A (scaled by Vβo(s) , Appendix C )
0.1
0.1
Extrinsic-reward weight in U(o) , α
0.25
0
0
Classifier match threshold, ξ
0.7 (with D=1−IOU )
see Table 2
see Table 2
Appendix
Table 5: Goal-space and option-discovery hyperparameters for Abstract Subgoals.
Markov Decision Processes (MDPs) often exhibit significant redundancy due to symmetries and shared structure across state-goal pairs in real-world Goal-Conditioned Reinforcement Learning (GCRL). While hierarchical policies have been motivated for horizon reduction via temporal abstraction in offline GCRL, we demonstrate that hierarchy also enables absolute abstraction. By introducing relativised options as well as distinct representations for different levels of the hierarchy, we demonstrate how an agent can reuse experience across similar contexts of the state-space. Based on this framework, we introduce two simple algorithms for learning relativised options and abstracting from the absolute frame of reference. Our experiments show that such inductive biases significantly improve performance in offline GCRL.
Clarisse Wibault, Alexander Goldie, Antonio Villares +2
Temporal abstraction allows reinforcement learning agents to represent knowledge and develop strategies over different temporal scales. The option-critic framework has been demonstrated to learn temporally extended actions, represented as options, end-to-end in a model-free setting. However, feasibility of option-critic remains limited due to two major challenges, multiple options adopting very similar behavior, or a shrinking set of task relevant options. These occurrences not only void the need for temporal abstraction, they also affect performance. In this paper, we tackle these problems by learning a diverse set of options. We introduce an information-theoretic intrinsic reward, which augments the task reward, as well as a novel termination objective, in order to encourage behavioral diversity in the option set. We show empirically that our proposed method is capable of learning options end-to-end on several discrete and continuous control tasks, outperforms option-critic by a wide margin. Furthermore, we show that our approach sustainably generates robust, reusable, reliable and interpretable options, in contrast to option-critic.
We study performance-driven environment abstraction for decision-making in large Markov decision processes. Rather than preserving geometric or topological structure, we seek abstractions that directly optimize decision quality. We model abstraction as a controlled approximation obtained by aggregating the state space and enforcing a shared action distribution within each aggregated state. For a fixed partition, we establish a performance guarantee that separates value-function approximation error from the loss introduced by action sharing. Guided by this analysis, we develop a multi-timescale reinforcement learning framework that jointly adapts the policy and a tree-structured environment abstraction. The resulting algorithm refines and coarsens regions of the state space based on Q-value discrepancies, balancing performance against abstraction size and complexity. Empirical results demonstrate substantial state compression, improved sample efficiency, and faster replanning compared to actor-critic baselines.
Yue Guan, Dipankar Maity, Panagiotis Tsiotras
School of Aerospace Engineering, Georgia Institute of Technology, Atlanta, GA, USA · Department of Electrical and Computer Engineering, University of North Carolina at Charlotte, Charlotte, NC, USA