Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability. In this work, we introduce a different route that eliminates policy optimization entirely by leveraging diffusion policies. Our key insight is that a diffusion policy encodes the action-gradient structure of the optimal soft Q function, enabling reward learning to be cast as a sequence of value recovery problems, thereby allowing us to bypass reward-policy loops inherent in prior IRL methods. Specifically, our method proceeds in three stages: (I) recovering the optimal soft Q function via action-gradient matching and estimating the corresponding soft value function (LogSumExp of Q values) in a way inspired by Gumbel regression; (II) calibrating these soft values by inferring a state-dependent offset; (III) extracting the reward by enforcing Bellman consistency. This leads to Loop-Free Inverse Reinforcement Learning (LFIRL), a fully offline algorithm that operates in a simple, loop-free, and sequential manner. LFIRL is simple to implement and significantly improves training efficiency while maintaining strong reward recovery performance. Empirically, across Maze, Franka Kitchen, Adroit Hand Pen, and Push-T benchmarks, LFIRL achieves 2-3x speedup over the fastest baselines, while matching or surpassing state-of-the-art methods in reward recovery quality.
Figures & tables
Figure 1 : Traditional inverse reinforcement learning alternates between reward learning and policy optimization, while our loop-free version trains each component once without a reward-policy loop.
Figure 2 : Illustration of matching policy’s score on Push-T task, where the goal is to push the T-shaped block to the target pose . The red circle marks the agent’s current position (initially in the Start panel), and the dark dashed curves indicate two optimal routes to reach the correct position for pushing the block. The arrows depict the action-gradient field supervised by the diffusion policy’s score; the blue dots are perturbed actions whose Q values are constrained below that of the expert action by a margin, thereby providing anchoring information.
Figure 3 : Illustrations of the evaluation tasks used in our experiments.
Figure 4 : Reward-recovery results on PointMaze and Franka Kitchen. We report success rate on PointMaze and average completed subtasks on Franka Kitchen. Since LFIRL recovers the reward only at the final stage, its result is shown as a fixed final score rather than a policy-learning curve. For online IRL methods, “steps” denotes environment interaction steps. For offline IRL methods, “steps” denotes the number of transitions sampled from the offline dataset during training.
Task
LFIRL
AIRL
IQ-Learn
DIFO
DRAIL
ML-IRL
Offline ML-IRL
CLARE
ValueDICE
Push-T
74.17 ± 1.96
59.67 ± 2.08
39.67 ± 0.58
79.00 ± 1.00
49.67 ± 1.15
71.33 ± 2.33
42.67 ± 1.20
71.00 ± 3.67
78.67 ± 2.52
Pen
97.17 ± 2.75
90.00 ± 2.45
74.33 ± 7.59
91.67 ± 2.05
98.33 ± 0.94
71.00 ± 4.32
66.00 ± 12.43
94.67 ± 0.58
95.67 ± 1.15
Table 1 : Trajectory-quality classification accuracy (%) on Push-T and Adroit Hand Pen using reward networks recovered by different IRL algorithms. Results are reported as mean ± standard deviation over five runs.
Task
Online methods
Offline methods
LFIRL
AIRL
IQ-Learn
DIFO
DRAIL
ML-IRL
ML-IRL
CLARE
ValueDICE
Load D.P.
T.R. (SpeedUp)
Train D.P.
T.R. (SpeedUp)
UMaze
1.38
6.88
2.47
1.70
9.76
3.09
5.86
4.66
0.28
79.7% (5x)
0.43
68.8% (3x)
Maze(M)
3.67
10.10
7.05
2.65
13.80
4.99
5.98
4.87
0.77
70.9% (3x)
1.07
59.6% (2x)
Maze(L)
6.88
10.88
7.73
3.05
13.33
5.52
6.67
5.17
1.42
53.4% (2x)
1.80
41.0% (2x)
Kitchen
9.55
13.72
7.40
3.33
18.42
6.83
7.80
5.91
1.52
54.4% (2x)
1.87
43.8% (2x)
Push-T
1.57
8.55
4.05
1.95
9.50
2.10
3.20
5.30
0.35
77.7% (4x)
0.62
60.5% (3x)
Table 2 : Average running time (in hours). “Load D.P.” denotes LFIRL with a pre-trained diffusion policy, while “Train D.P.” denotes LFIRL including diffusion-policy training time. “T.R.” is short for “time reduction” and reports the relative time saved compared with the fastest baseline (marked with an underline).
Method
UMaze
Medium
Large
Kitchen
Push-T
Pen
Average
PCC / SCC
PCC / SCC
PCC / SCC
PCC / SCC
PCC / SCC
PCC / SCC
PCC / SCC
LFIRL (Ours)
0.83 / 0.91
0.76 / 0.82
0.84 / 0.87
0.67 / 0.81
0.72 / 0.81
0.79 / 0.85
0.77 / 0.85
AIRL
0.65 / 0.86
0.63 / 0.74
0.77 / 0.75
0.44 / 0.63
0.69 / 0.74
0.75 / 0.71
0.66 / 0.74
IQ-Learn
0.62 / 0.74
0.59 / 0.62
0.65 / 0.70
0.34 / 0.51
0.48 / 0.65
0.47 / 0.62
0.53 / 0.64
DIFO
0.69 / 0.79
0.53 / 0.67
0.68 / 0.81
0.61 / 0.77
0.83 / 0.79
0.71 / 0.72
0.68 / 0.76
DRAIL
0.67 / 0.71
0.74 / 0.84
0.69 / 0.78
0.56 / 0.76
0.51 / 0.56
0.63 / 0.62
0.63 / 0.71
Table 3 : Direct comparison of recovered and ground-truth rewards on the same state-action samples. Each entry reports PCC / SCC; higher values indicate stronger linear association / rank agreement. Bold indicates the highest value for each metric in each task.
Figure 5 : Ablation results on PointMaze, Kitchen, Pen, and Push-T. We compare LFIRL with its variant that removes the state-dependent offset b(s) .
Figure 6 : Reduced-data experiments with the full, one-half, and one-quarter of the demonstrations. We report success rate on PointMaze, average completed subtasks on Franka Kitchen, and reward on Push-T and Pen.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : Additional ablation comparing DDPM-based and Flow-Matching-based pretraining in LFIRL. The two variants use the same subsequent sequential reward-recovery pipeline, but differ in how the pretrained generative policy is obtained. Higher is better.
Environment
LFIRL-Uniform
LFIRL-Gaussian
Expert
Swimmer-v5
284.56 ± 27.64
208.42 ± 7.25
315.55 ± 1.38
Walker2d-v5
4933.41 ± 52.79
4722.40 ± 230.42
5861.18 ± 73.99
Hopper-v5
3413.50 ± 210.23
3868.04 ± 79.58
4098.17 ± 247.70
Appendix
Table 4: Sensitivity to the reference policy on MuJoCo tasks. Results are episodic returns, reported as mean ± standard deviation. The expert return is included for context.
Diffusion-noise magnitude
Value-anchoring coefficient
Intended margin
Environment
Low
Medium
High
λ=0.5
λ=1
λ=2
ξ=0.5
ξ=1
ξ=2
UMaze
0.98
0.87
0.68
0.92
0.98
0.97
0.57
0.98
0.61
Medium
0.39
0.35
0.25
0.38
0.39
0.35
0.24
0.39
0.26
Large
0.31
0.29
0.25
0.26
0.31
0.34
0.18
0.31
0.24
Appendix
Table 5: Sensitivity of PointMaze success rate to Stage-I hyperparameters. Each group varies one parameter while the remaining parameters are held at their default values. Bold indicates the highest result within each parameter group for each environment.
PointMaze_UMaze-v3
PointMaze_Medium-v3
PointMaze_Large-v3
Expert demonstrations
2000
2000
2000
Policy pretraining epochs
60
40
40
Q/ b /V/ r learning rate
3e-5
3e-5
3e-5
Stage passes (Q,b,V,r)
8/8/20/20
8/8/20/20
8/8/20/20
Q, b , V, r hidden layers
256, 256, 256, 256
256, 256, 256, 256
256, 256, 256, 256
IRL batch size
256
256
256
Appendix
Table 6: Network architecture and hyperparameter setup for the PointMaze environments.
FrankaKitchen-v1
gym_pusht/PushT-v0
AdroitHandPen-v1
Expert demonstrations
19
200
200
Policy pretraining epochs
20
200
40
Q/ b /V/ r learning rate
3e-5
3e-4
3e-5
Stage passes (Q,b,V,r)
8/8/20/20
40/40/100/100
40/40/100/100
Q, b , V, r hidden layers
256, 256, 256, 256
256, 256, 256, 256
256, 256, 256, 256
IRL batch size
256
256
256
Appendix
Table 7: Network architecture and hyperparameter setup for the robotic manipulation environments.