cs.ROOct 7, 2026

Learning Unknown Constraints without Unsafe Data via Optimality and Counterfactual Regularization

Authors: Zhouyu Zhang, Chih-Yuan Chiu, Glen Chou

Organizations: School of Electrical and Computer Engineering Georgia Institute of Technology Atlanta, GA 30309 · School of Cybersecurity and Privacy School of Aerospace Engineering Georgia Institute of Technology

Abstract

Learning from demonstrations (LfD) provides a framework for inferring unknown constraints from locally optimal, constraint-satisfying expert behavior. Existing approaches largely fall into two paradigms, constrained inverse optimal control (CIOC) and inverse constrained reinforcement learning (ICRL). CIOC exploits optimality conditions such as the Karush--Kuhn--Tucker (KKT) conditions but typically assumes known dynamics and structured constraint representations. Meanwhile, ICRL accommodates complex unknown constraints and unknown transition dynamics but often requires extensive online exploration, during which unsafe constraint violations may occur. In this work, we introduce Counterfactual KKT (CF-KKT), a constraint learning framework that leverages learned dynamics and locally optimal demonstrations to recover unknown constraints without requiring known dynamics or additional risky exploration, thereby combining the data efficiency and safety advantages of CIOC with the flexibility of ICRL. First, we use a locally learned differentiable dynamics model to impose KKT-inspired optimality conditions directly on the demonstrations. Second, we use the learned dynamics to generate reward-improving counterfactual behaviors near the demonstrations, revealing behaviors that would be preferable in the absence of the unknown constraint and thus providing synthetic infeasible data. When the constraint parameterization is known, the same learned-dynamics framework enables direct CIOC-based parameter recovery, and we characterize its sensitivity to dynamics misspecification. Across high-dimensional robotic control tasks, our approach learns neural constraint representations with improved safety and data efficiency relative to state-of-the-art offline ICRL baselines.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learning from Demonstration via Spatiotemporal Tubes for Unknown Euler-Lagrange Systems

    Jul 1, 2026Ratnangshu Das, Puneeth Shankar, Varuni Buereddy +2Robotic ControlRobot Skill Learning

  2. Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data

    May 2, 2026Ruiqi Xue, Lei Yuan, Kainuo Cheng +2LLM-Guided RLOffline RL

  3. CAPSULE: Control-Theoretic Action Perturbations for Safe Uncertainty-Aware Reinforcement Learning

    Apr 26, 2026Rahul Narava, Siddharth Verma, Ojas Jain +2Control Barrier FunctionsConstrained RL