Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability. In this work, we introduce a different route that eliminates policy optimization entirely by leveraging diffusion policies. Our key insight is that a diffusion policy encodes the action-gradient structure of the optimal soft Q function, enabling reward learning to be cast as a sequence of value recovery problems, thereby allowing us to bypass reward-policy loops inherent in prior IRL methods. Specifically, our method proceeds in three stages: (I) recovering the optimal soft Q function via action-gradient matching and estimating the corresponding soft value function (LogSumExp of Q values) in a way inspired by Gumbel regression; (II) calibrating these soft values by inferring a state-dependent offset; (III) extracting the reward by enforcing Bellman consistency. This leads to Loop-Free Inverse Reinforcement Learning (LFIRL), a fully offline algorithm that operates in a simple, loop-free, and sequential manner. LFIRL is simple to implement and significantly improves training efficiency while maintaining strong reward recovery performance. Empirically, across Maze, Franka Kitchen, Adroit Hand Pen, and Push-T benchmarks, LFIRL achieves 2-3x speedup over the fastest baselines, while matching or surpassing state-of-the-art methods in reward recovery quality.
Figures & tables
Figure 1 : Traditional inverse reinforcement learning alternates between reward learning and policy optimization, while our loop-free version trains each component once without a reward-policy loop.
Figure 2 : Illustration of matching policy’s score on Push-T task, where the goal is to push the T-shaped block to the target pose . The red circle marks the agent’s current position (initially in the Start panel), and the dark dashed curves indicate two optimal routes to reach the correct position for pushing the block. The arrows depict the action-gradient field supervised by the diffusion policy’s score; the blue dots are perturbed actions whose Q values are constrained below that of the expert action by a margin, thereby providing anchoring information.
Figure 3 : Illustrations of the evaluation tasks used in our experiments.
Figure 4 : Reward-recovery results on PointMaze and Franka Kitchen. We report success rate on PointMaze and average completed subtasks on Franka Kitchen. Since LFIRL recovers the reward only at the final stage, its result is shown as a fixed final score rather than a policy-learning curve. For online IRL methods, “steps” denotes environment interaction steps. For offline IRL methods, “steps” denotes the number of transitions sampled from the offline dataset during training.
Task
LFIRL
AIRL
IQ-Learn
DIFO
DRAIL
ML-IRL
Offline ML-IRL
CLARE
ValueDICE
Push-T
74.17 ± 1.96
59.67 ± 2.08
39.67 ± 0.58
79.00 ± 1.00
49.67 ± 1.15
71.33 ± 2.33
42.67 ± 1.20
71.00 ± 3.67
78.67 ± 2.52
Pen
97.17 ± 2.75
90.00 ± 2.45
74.33 ± 7.59
91.67 ± 2.05
98.33 ± 0.94
71.00 ± 4.32
66.00 ± 12.43
94.67 ± 0.58
95.67 ± 1.15
Table 1 : Trajectory-quality classification accuracy (%) on Push-T and Adroit Hand Pen using reward networks recovered by different IRL algorithms. Results are reported as mean ± standard deviation over five runs.
Task
Online methods
Offline methods
LFIRL
AIRL
IQ-Learn
DIFO
DRAIL
ML-IRL
ML-IRL
CLARE
ValueDICE
Load D.P.
T.R. (SpeedUp)
Train D.P.
T.R. (SpeedUp)
UMaze
1.38
6.88
2.47
1.70
9.76
3.09
5.86
4.66
0.28
79.7% (5x)
0.43
68.8% (3x)
Maze(M)
3.67
10.10
7.05
2.65
13.80
4.99
5.98
4.87
0.77
70.9% (3x)
1.07
59.6% (2x)
Maze(L)
6.88
10.88
7.73
3.05
13.33
5.52
6.67
5.17
1.42
53.4% (2x)
1.80
41.0% (2x)
Kitchen
9.55
13.72
7.40
3.33
18.42
6.83
7.80
5.91
1.52
54.4% (2x)
1.87
43.8% (2x)
Push-T
1.57
8.55
4.05
1.95
9.50
2.10
3.20
5.30
0.35
77.7% (4x)
0.62
60.5% (3x)
Table 2 : Average running time (in hours). “Load D.P.” denotes LFIRL with a pre-trained diffusion policy, while “Train D.P.” denotes LFIRL including diffusion-policy training time. “T.R.” is short for “time reduction” and reports the relative time saved compared with the fastest baseline (marked with an underline).
Method
UMaze
Medium
Large
Kitchen
Push-T
Pen
Average
PCC / SCC
PCC / SCC
PCC / SCC
PCC / SCC
PCC / SCC
PCC / SCC
PCC / SCC
LFIRL (Ours)
0.83 / 0.91
0.76 / 0.82
0.84 / 0.87
0.67 / 0.81
0.72 / 0.81
0.79 / 0.85
0.77 / 0.85
AIRL
0.65 / 0.86
0.63 / 0.74
0.77 / 0.75
0.44 / 0.63
0.69 / 0.74
0.75 / 0.71
0.66 / 0.74
IQ-Learn
0.62 / 0.74
0.59 / 0.62
0.65 / 0.70
0.34 / 0.51
0.48 / 0.65
0.47 / 0.62
0.53 / 0.64
DIFO
0.69 / 0.79
0.53 / 0.67
0.68 / 0.81
0.61 / 0.77
0.83 / 0.79
0.71 / 0.72
0.68 / 0.76
DRAIL
0.67 / 0.71
0.74 / 0.84
0.69 / 0.78
0.56 / 0.76
0.51 / 0.56
0.63 / 0.62
0.63 / 0.71
Table 3 : Direct comparison of recovered and ground-truth rewards on the same state-action samples. Each entry reports PCC / SCC; higher values indicate stronger linear association / rank agreement. Bold indicates the highest value for each metric in each task.
Figure 5 : Ablation results on PointMaze, Kitchen, Pen, and Push-T. We compare LFIRL with its variant that removes the state-dependent offset b(s) .
Figure 6 : Reduced-data experiments with the full, one-half, and one-quarter of the demonstrations. We report success rate on PointMaze, average completed subtasks on Franka Kitchen, and reward on Push-T and Pen.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : Additional ablation comparing DDPM-based and Flow-Matching-based pretraining in LFIRL. The two variants use the same subsequent sequential reward-recovery pipeline, but differ in how the pretrained generative policy is obtained. Higher is better.
Environment
LFIRL-Uniform
LFIRL-Gaussian
Expert
Swimmer-v5
284.56 ± 27.64
208.42 ± 7.25
315.55 ± 1.38
Walker2d-v5
4933.41 ± 52.79
4722.40 ± 230.42
5861.18 ± 73.99
Hopper-v5
3413.50 ± 210.23
3868.04 ± 79.58
4098.17 ± 247.70
Appendix
Table 4: Sensitivity to the reference policy on MuJoCo tasks. Results are episodic returns, reported as mean ± standard deviation. The expert return is included for context.
Diffusion-noise magnitude
Value-anchoring coefficient
Intended margin
Environment
Low
Medium
High
λ=0.5
λ=1
λ=2
ξ=0.5
ξ=1
ξ=2
UMaze
0.98
0.87
0.68
0.92
0.98
0.97
0.57
0.98
0.61
Medium
0.39
0.35
0.25
0.38
0.39
0.35
0.24
0.39
0.26
Large
0.31
0.29
0.25
0.26
0.31
0.34
0.18
0.31
0.24
Appendix
Table 5: Sensitivity of PointMaze success rate to Stage-I hyperparameters. Each group varies one parameter while the remaining parameters are held at their default values. Bold indicates the highest result within each parameter group for each environment.
PointMaze_UMaze-v3
PointMaze_Medium-v3
PointMaze_Large-v3
Expert demonstrations
2000
2000
2000
Policy pretraining epochs
60
40
40
Q/ b /V/ r learning rate
3e-5
3e-5
3e-5
Stage passes (Q,b,V,r)
8/8/20/20
8/8/20/20
8/8/20/20
Q, b , V, r hidden layers
256, 256, 256, 256
256, 256, 256, 256
256, 256, 256, 256
IRL batch size
256
256
256
Appendix
Table 6: Network architecture and hyperparameter setup for the PointMaze environments.
FrankaKitchen-v1
gym_pusht/PushT-v0
AdroitHandPen-v1
Expert demonstrations
19
200
200
Policy pretraining epochs
20
200
40
Q/ b /V/ r learning rate
3e-5
3e-4
3e-5
Stage passes (Q,b,V,r)
8/8/20/20
40/40/100/100
40/40/100/100
Q, b , V, r hidden layers
256, 256, 256, 256
256, 256, 256, 256
256, 256, 256, 256
IRL batch size
256
256
256
Appendix
Table 7: Network architecture and hyperparameter setup for the robotic manipulation environments.
Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching the distribution of expert trajectories. Classical (dual-ascent) IRL guarantees monotonic performance improvement but requires fully solving an RL problem each iteration to compute dual gradients. More recent adversarial methods avoid this cost at the expense of stability and monotonic dual improvement, by directly optimizing the primal problem and using a discriminator to provide rewards. In this work, we bridge the gap between these approaches by enabling monotonic improvement of the reward function and policy without having to fully solve an RL problem at every iteration. Our key theoretical insight is that a trust-region-optimal policy for a reward function update can be globally optimal for a smaller update in the same direction. This smaller update allows us to explicitly optimize the dual objective while only relying on a local search around the current policy. In doing so, our approach avoids the training instabilities of adversarial methods, offers monotonic performance improvement, and learns a reward function in the traditional sense of IRL--one that can be globally optimized to match expert demonstrations. Our proposed algorithm, Trust Region Inverse Reinforcement Learning (TRIRL), outperforms state-of-the-art imitation learning methods across multiple challenging tasks by a factor of 2.4x in terms of aggregate inter-quartile mean, while recovering reward functions that generalize to system dynamics shifts.
Anish Diwan, Davide Tateo, Christopher E. Mower +3
Technical University of Darmstadt · Robotics Institute Germany (RIG) · Lund University +3
In the forward reinforcement-learning problem, the reward is fixed and known; the learner is asked to find a good policy or value function. Here we turn the question around. Given offline data generated by an expert, can we recover the reward the expert was optimizing? This is the inverse reinforcement learning problem, and remarkably, two communities, structural econometricians studying dynamic discrete choice (DDC) and machine learners studying entropy-regularized IRL, have been working on exactly the same probabilistic model under different names. We begin by proving their equivalence. We then develop the classical identification result of Magnac and Thesmar and the classical computational paradigms that grew out of it: Rust's nested fixed-point algorithm, the conditional-choice-probability approach of Hotz and Miller, and the two temporal-difference approaches of Adusumilli and Eckardt: linear semi-gradient TD and approximate value iteration. Each route has its limits: dimensionality, transition-kernel estimation, the deadly triad, or projected fixed-point bias. We then walk through the modern ML/IRL strand: adversarial IRL, occupancy matching, IQ-Learn, and offline ML-IRL, deriving each method's actual objective and stating precisely what it does and does not identify. We close with the empirical-risk-minimization framework of Kang et al., which yields a gradient-based estimator for offline IRL/DDC.
Enoch Hyunwook Kang
University of Washington, Foster School of Business
Inverse reinforcement learning (IRL) learns a reward function and a corresponding policy that best fit the demonstration data of an expert. However, in the current IRL setting, the learner is isolated from the expert and can only passively observe the expert demonstrations. This limits the applicability of IRL to interactive settings, where the learner actively interacts with the expert and needs to infer the expert's reward function from the interactions. To bridge the gap, this paper studies interactive IRL (IIRL) where a learner aims to learn the reward function of an expert and a policy to interact with the expert during its interactions with the expert. We formulate IIRL as a stochastic bi-level optimization problem where the lower level learns a reward function to explain the behaviors of the expert, and the upper level learns a policy to interact with the expert. We develop a double-loop algorithm, Bi-level Interactive Scenarios Inverse Reinforcement Learning (BISIRL), which solves the lower-level problem in the inner loop and the upper-level problem in the outer loop. We formally guarantee that BISIRL converges and validate our algorithm through extensive experiments.
Yue Mao, Shicheng Liu, Siyuan Xu +1
Department of Mechanical Engineering, Pennsylvania State University · Department of Electrical Engineering, Pennsylvania State University