By learning transferable rewards, inverse reinforcement learning (IRL) enables counterfactual evaluation of agents under modified environments. Such transfer places strict requirements on coverage since target environments affect agents' state occupancy. We propose a multi-task IRL method that pools data across multiple agents with different rewards in the same environment under a low-rank assumption. In addition to alleviating coverage requirements, so each task need not visit every state as long as others do, the method offers scalable evaluation of multiple tasks under new environments as computationally intensive planning scales with rank rather than the number of tasks. We provide finite sample guarantees on reward recovery and on policy learning in new environments. Experiments show our method is robust to limited coverage, recovers rewards on and off of each task's support, transfers to target environments at lower regret than baselines, with its computational advantage over per-task methods widening as tasks grow.
Figures & tables
Figure 1: Causal test of directional coverage. Left: change in error projected onto the targeted eigendirection for treated, volume-matched control, and placebo cells. Right: treated-minus-control difference-in-differences for the targeted and remaining directions.
Figure 2: Reward recovery versus samples per task (mean and standard deviation over 10 seeds). LowRank leads off-support recovery at every budget; per-task methods narrow the gap only on Highway at high sample sizes.
Figure 3: Policy recovery in target perturbed environments. Error bands are ± 1 s.d over ten seeds.
Figure 4: Wall-clock scaling with task population N at fixed basis rank k . LowRank exhibits substantially sublinear growth, while per-task methods scale approximately linearly
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
FourRooms
Highway
RecSim
Encoder
CNN
MLP
MLP
Hidden width
512
128
256
Residual blocks
0
0
2
FQI iterations
16
16
4
FQI regression epochs
50
50
40
Basis rank k
4
4
4
Appendix
Table 1: IRL pipeline settings per environment. All methods within an environment use the same encoder and width, so the comparison isolates the estimation strategy.
Domain
Method
N=8
N=32
N=128
Highway, n=500
LowRank
+0.912±0.010
+0.885±0.011
+0.884±0.010
GenPQR-per-task
+0.787±0.039
+0.742±0.052
+0.739±0.026
IQ-per-task
+0.677±0.040
+0.649±0.042
+0.645±0.024
MI
+0.849±0.014
+0.846±0.082
+0.878±0.003
Highway, n=100
LowRank
+0.879±0.018
+0.847±0.017
+0.844±0.016
GenPQR-per-task
−0.109±0.082
−0.113±0.029
−0.097±0.020
Appendix
Table 2: Reward-recovery quality (off-support global Pearson r , mean ± s.d. over 10 seeds) against population size N . LowRank and both per-task baselines are flat in N (largest change is 0.048 ); the small step from N=8 to N=32 on Highway and FourRooms is taken by all methods alike and reflects those populations sampling reward weights per agent, so agents 8 – 31 differ from the first 8 . MI is the exception: it gains +0.51 on Highway at n=100 , where its clusters are starved at N=8 (s.d. 0.204 , falling to 0.028 by N=128 ), but loses 0.08 on FourRooms at n=500 with tight error bars.
Figure 5: Transfer learning is effective: training basis on M agents in Phase 1 and transferring to Np2 Phase 2 agents
Figure 6: Empirical scaling in RecSim matches the theory in n and S and is more conservative in k and A .
Figure 7: Difference between reward recovery between reward operator applied to exact log normalized policies versus approximation
Figure 8: Misspecification of rank leads to graceful degradation
n=100
n=104
KMI
Highway
FourRooms
RecSim
Highway
FourRooms
RecSim
1 †
0.04±
0.03
0.01±
0.04
0.01±
0.01
0.09±
0.01
0.15±
0.00
−0.00±
0.00
2
0.23±
0.13
0.13±
0.08
0.04±
0.06
0.54±
0.03
0.37±
0.06
0.18±
0.01
3
0.04±
0.07
0.18±
0.07
0.08±
0.06
0.60±
0.13
0.41±
0.03
0.30±
0.03
4
−0.11±
0.06
0.22±
0.05
0.05±
0.04
0.68±
0.13
0.43±
0.04
0.34±
0.02
6
—
0.27±
0.04
0.10±
0.03
—
0.47±
0.03
0.32±
0.01
Appendix
Table 3: MI-Cluster across cluster count KMI (off-support Pearson r , mean ± s.d. over 10 seeds); N=5 for Highway, 8 otherwise. Bold marks each column’s best K . Ours beats MI at its best K in every domain at both budgets. Shapes are budget-dependent: Highway rises to K=N at n=104 but falls to it at n=100 , and RecSim’s peak at the generative rank K=4 disappears.