We study adversarial imitation learning (AIL), in which an agent learns to imitate expert demonstrations by optimizing a policy against an adversarial reward that distinguishes expert and learner behavior. Historically, reward regularization and entropy-based policy regularization are key components of empirically successful methods such as GAIL and LS-IQ, yet their finite-sample benefits remain underexplored. We establish fast rates for jointly regularized AIL in finite-horizon Markov decision processes with general function approximation. Our model-free algorithm, Dually Regularized AIL, combines KL policy regularization with a quadratic reward penalty weighted by expert and learner occupancies. With K online episodes and N expert trajectories, we prove a O(K1+N1) bound on the regularized imitation gap for fixed regularization parameters. Our analysis combines an online mirror descent construction for general convex reward classes to control estimation error from finite expert data and stochastic learner feedback, with a sharp analysis of optimistic KL-regularized policy learning. To the best of our knowledge, Dually Regularized AIL is the first algorithm to simultaneously achieve O(ε1) sample complexity in both expert demonstrations and online interactions for this regularized AIL objective, even with stochastic experts. These results provide a rigorous characterization of the complementary statistical benefits of reward and policy regularization in AIL.
Figures & tables
Method
Function class
Expert
Criterion
N
K
Mimic-Emp ( Rajaraman et al., 2020 )
Tabular
General
Expected IL gap
ϵ−1
0
Log-loss BC ( Foster et al., 2024 )
General
Deterministic
Imitation gap
ϵ−1
0
Log-loss BC ( Foster et al., 2024 )
General
General
Imitation gap
ϵ−2
0
MB-TAIL ( Xu et al., 2023 )
Tabular
Deterministic
Imitation gap
ϵ−1
ϵ−2
OPT-AIL ( Xu et al., 2024 )
General
General
Imitation gap
ϵ−2
ϵ−2
MB-AIL ( Li et al., 2026 )
General
General
Imitation gap
ϵ−2
ϵ−2
Table 1: Sample complexity in expert trajectories ( N ) and online episodes ( K ) in related works. Sample complexity omits the horizon, reward-range, and structural-complexity factors. “General” experts may be stochastic. All guarantees are high-probability except Mimic-Emp’s expected gap. Unregularized rates summarize the worst case, omitting variance-dependent improvements. These works made regular RL / IL assumptions where we refer readers to check their original statements.
Adversarial Imitation Learning (AIL) faces challenges with sample inefficiency because of its reliance on sufficient on-policy data to evaluate the performance of the current policy during reward function updates. In this work, we study the convergence properties and sample complexity of off-policy AIL algorithms. We show that, even in the absence of importance sampling correction, reusing samples generated by the o(K) most recent policies, where K is the number of iterations of policy updates and reward updates, does not undermine the convergence guarantees of this class of algorithms. Furthermore, our results indicate that the distribution shift error induced by off-policy updates is dominated by the benefits of having more data available. This result provides theoretical support for the sample efficiency of off-policy AIL algorithms. To the best of our knowledge, this is the first work that provides theoretical guarantees for off-policy AIL algorithms.
Yilei Chen, Vittorio Giammarino, James Queeney +1
Division of Systems Engineering Boston University Boston, MA 02215, USA · Department of Computer Science Purdue University West Lafayette, IN 47907, USA · Amazon Robotics North Reading, MA 01864, USA +1
Adversarial imitation learning (AIL), a prominent approach in imitation learning, has achieved significant practical success powered by neural network approximation. However, existing theoretical analyses of AIL are primarily confined to simplified settings, such as tabular and linear function approximation, and involve complex algorithmic designs that impede practical implementation. This creates a substantial gap between theory and practice. This paper bridges this gap by exploring the theoretical underpinnings of online AIL with general function approximation. We introduce a novel framework called optimization-based AIL (OPT-AIL), which performs online optimization for reward learning coupled with optimism-regularized optimization for policy learning. Within this framework, we develop two concrete methods: model-free OPT-AIL and model-based OPT-AIL. Our theoretical analysis demonstrates that both variants achieve polynomial expert sample complexity and interaction complexity for learning near-expert policies. To the best of our knowledge, they represent the first provably efficient AIL methods under general function approximation. From a practical standpoint, OPT-AIL requires only the approximate optimization of two objectives, thereby facilitating practical implementation. Empirical studies demonstrate that OPT-AIL outperforms previous state-of-the-art deep AIL methods across several challenging tasks.
Tian Xu, Zhilong Zhang, Zexuan Chen +3
National Key Laboratory for Novel Software Technology and School of Artificial Intelligence, Nanjing University · School of Mathematics, Nanjing University
Adversarial imitation learning (AIL) achieves high-quality imitation compared to behavioral cloning (BC), but demands substantial online environment interaction. Recent empirical work has explored initializing AIL algorithms with BC pretrained policies to address this limitation, yet a rigorous theoretical understanding of pretraining's role in AIL remains elusive. This paper provides a systematic theoretical analysis and introduces principled pretraining algorithms for accelerating AIL. We begin by analyzing AIL with policy pretraining alone, identifying reward error as the dominant source of suboptimality. This reveals a critical and previously overlooked gap: the absence of reward pretraining. Motivated by this finding, we develop a principled policy-reward co-pretraining approach grounded in a reward shaping analysis. Our analysis uncovers a fundamental connection between expert policies and shaping rewards, which naturally gives rise to CoPT-AIL, an approach that jointly pretrains both policy and reward through a single BC procedure. We prove that CoPT-AIL achieves an improved imitation gap bound over standard AIL, establishing the first theoretical guarantee for the benefits of pretraining in AIL. Experimental results confirm CoPT-AIL's superior performance over existing AIL methods.
Tian Xu, Zexuan Chen, Zhilong Zhang +4
National Key Laboratory for Novel Software Technology and School of Artificial Intelligence, Nanjing University, China.