Interpretability offers a solution to safe autonomy by providing transparency into an agent's underlying decision-making model. Within sequential-decision making tasks, differentiable decision trees (DDTs) are one approach to such interpretability, maintaining automatic-differentiable policies while providing humans with a discrete tree-based visualization. Nonetheless, current implementations of DDTs are not well-suited for sequential-decision making domains, as there exists an inherent mismatch between a tree's single-timestep behavior and a human's multi-timestep planning. Our work thus introduces time as a new dimension of interpretability, coined as temporal interpretability, and demonstrates how temporal abstractions via action chunking improve it. We achieve this by first introducing two novel policy gradient algorithms that incorporate action chunking. Additionally, to maintain parameter-efficient trees, we develop an information-theoretic tree restructuring algorithm that modifies the tree during training. Across four simulation environments, we find that warm-starting action chunked DDTs from a distilled action chunked policy is the most effective way to obtain temporally interpretable trees: they match neural network policies in three of the four domains while using up to 80% fewer parameters. Our code is available at https://github.com/ei5uke/temp-interp.
Figures & tables
Figure 1 : Overview figure. (a) Temporal Interpretability. We provide an example depicting how humans increase in confidence when provided with an agent’s short-term plan and temporal decision-making (an autonomous vehicle’s plan is visualized with blue arrows). (b) Temporally Abstracted Policy Gradients. We introduce two novel policy gradient methods that enable action chunked DDTs to be temporally interpretable. (c) Information-Theoretic Tree Restructuring. Introducing action chunking to trees increases parameter count, hence we dynamically deepen or prune leaves to ensure the model only contains sub-branches we need most.
Environments
Algorithms
LK
IP
LL
LL-H
MLP
185.3±2.9 497
1000.0±0.0 369
293.8±2.3 450
272.1±7.2450
MLP-Ensemble
190.2±2.0 650
1000.0±0.0 522
296.0±3.2 756
270±3.2756
MLP-Prediction
193.3±0.9 650
1000.0±0.0 522
294.6±1.6 756
273.2±5.6756
MLP-Big
191.4±1.65057
1000.0±0.04545
292.4±3.84866
269.6±4.04866
MLP-Big-Ensemble
189.2±3.85642
1000.0±0.05130
293.3±2706036
274.3±3.26036
Table 1 : Performance comparison of all algorithms in the four simulation domains, across three runs in an evaluation environment separate from the training environment. We report two values in each cell: the top value depicts the episodic return (mean ± s.e.) and the bottom value depicts the number of parameters (constant values for MLPs and mean ± s.e. for DDTs). For CART, we count a feature index and a threshold per decision node and the weights and bias of each leaf’s linear model.
Figure 2 : Figures depicting the trend in the number of leaves within DDTs when using our deepening / pruning algorithm for each simulation environment. Each curve showcases the mean and s.e. for a specific environment across three seeds, depicted in the dark curve and shaded regions respectively. Note that we run IP and LK for less timesteps, hence leading to the curves cutting off earlier.
Figure 4
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Visualization of a tree at the beginning and end of a run of the LL-H domain. Each node depicts a comparison between the most significant feature in the state against a threshold.
Algorithm
LK
IP
LL
LL-H
DDT-Ensemble NoLeaves
18±0.0
10±0.0
14±0.0
20.0±4.9
DDT-Pred NoLeaves
25.3±6.0
10±0.0
14±0.0
26.0±4.9
Warm-Ensemble NoLeaves
18±0.0
10±0.0
14±0.0
26.0±4.9
Warm-Pred NoLeaves
18±0.0
10±0.0
14±0.0
33.3±9.4
Appendix
Table 3 : Number of parameters used only for the logic flow of DDTs across the four domains and three runs (mean ± s.e.).
Figure 5 : Figures depicting the trend in the performance against the EMA coefficient in each simulation environment. The black curve depicts the regression over the results across all seeds.
Parameter
Value
RPO
Actor learning rate
5e−4
Critic learning rate
5e−4
PPO clip rate
0.2
RPO uniform bonus ( Rahman et al., 2025 )
0.5
MLP
Appendix
Table 4 : Shared hyperparameters used for all environments.
Domain
Parameter
LK
IP
LL
LL-H
DDT
Starting # of leaves
2
2
2
4
Maximum # of CART leaves
2
2
2
4
Information-Theoretic Tree Restructuring
Minimum steps since last restructure n
200
600
1000
1000
Appendix
Table 5 : Unique hyperparameters used for each environment for ITTR.
Reinforcement learning policies are difficult to inspect, but interpreting them is a prerequisite for trustworthiness. Converting a trained policy into explicit decision-tree rules improves transparency and the resulting artifacts often remain too complex for human understanding. We present a pruning process that simplifies such rule-based policies while preserving task performance and making edits to the policy auditable. The process defines a small set of structural and usage-aware operators and evaluates candidate edits by re-executing the policy to measure return and interpretability proxies. This exposes an transformation process from complex to compact policy structures. We investigate this approach on classic control and MuJoCo benchmarks, where pruning traces reveal consistent interpretability improvements while maintaining high performance.
Mark Leon Ringer, Michel Tokic
Faculty of Mathematics, Informatics and Statistics, Ludwig-Maximilians-University Munich, Munich, Germany · Siemens AG, Data & Artificial Intelligence, Otto-Hahn-Ring 6, 81739 Munich, Germany
Over the past decade, decision trees have been used to represent controllers (a.k.a. policies) in an explainable way, with dtControl2 as a current state-of-the-art tool. However, for systems that are large or have many corner cases, even such representations tend to be too complex and not human-comprehensible. Unfortunately, reducing the size of the decision tree is not straightforward, as missing just a single crucial case might result in an incorrect controller. We tackle this issue in the setting of Markov decision processes, extending dtControl2 by "ε" functionality: Given an allowed imprecision ε≥0, we construct a smaller decision tree, distilling the essence of the controller, while still guaranteeing its ε-optimality. This enables us to provide tunably simpler explanations, omitting a controllable amount of detail. Our tool constructs decision trees that are orders of magnitude smaller than the state of the art.
Tereza Kinská, Jan Křetínský, Tobias Meggendorfer +2
Masaryk University, Brno, Czech Republic · Technical University of Munich, Munich, Germany · Lancaster University Leipzig, Leipzig, Germany +1
Deep reinforcement learning (DRL) agents achieve strong performance in complex environments, yet their decision-making processes remain difficult to interpret. We introduce SPOT (Sampling Policy Observation Tree), a novel model-agnostic, sampling-based framework for interpreting DRL policies. Given access to the policy and an environment simulator, SPOT constructs an interpretable finite-horizon tree by sampling actions and recursively simulating the resulting successor states. The tree provides an empirical representation of the policy's action preferences and their possible downstream evolution. We provide formal guarantees establishing SPOT's asymptotic recovery of the policy's unique most probable action and characterizing its disagreement behavior under high-entropy policies. We demonstrate SPOT in the SUMO-RL traffic-signal control domain. The case study illustrates how its tree-based representation can be used to inspect policy preferences, compare alternative future trajectories, and reveal downstream behaviors that are not visible through single-timestep feature-attribution methods.
Tamar Gozlan, Claudia V. Goldman
Benin School of Computer Science and Engineering, The Hebrew University of Jerusalem · Hebrew University Business School, Data Science Department, The Hebrew University of Jerusalem