cs.ROOct 6, 2026

Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection

Authors: Lemon Foxmere, Anthony Furman, Yizheng Du, Oliver Chang, Leilani Gilpin, Steve McGuire

Organizations: HARE Lab, University of California at Santa Cruz, Santa Cruz, CA 95064 USA · AIEA Lab, University of California at Santa Cruz, Santa Cruz, CA 95064 USA

Abstract

Reinforcement Learning (RL) has enabled legged robots to perform a range of skills in single-task settings. However, applications such as farm robotics or space exploration require diverse skills such as locomotion, digging, or close-range surveying. Training an end-to-end policy to address this problem remains difficult due to challenges such as sample inefficiency and gradient conflict between tasks in multi-task learning. We propose a three-stage method that trains a single policy to perform distinct tasks such as walking, digging, and hopping, and compose them into novel behaviors such as crawling. First, multiple teacher policies are trained using RL on narrowly defined tasks. Then, two additional stages train a student policy with a multi-teacher distillation setup that uses a combined RL and Imitation Learning (IL) objective under an adversarial task selection process that focuses training on the worst-performing task. With this method, we train a student policy that performs 22 tasks using 8 teachers. Evaluations show our method preserves motion quality and tracks commands more accurately than PPO and distill-then-finetune baselines, and in some cases generalizes to new tasks without explicit training. Finally, we demonstrate real-world robustness by deploying the resulting policy on a Unitree B1 quadruped. Video: https://youtu.be/V9yX04EBcFA

Explore similar work

CardsList
  1. MimicAgent: Quadruped Skills via Prompt-to-Trajectory Generation

    Sep 21, 2026Lucky Kant Nayak, Narayanan Palghat Parameswaran, Neehar Peri +1Reinforcement LearningTrajectory Generation

  2. Skill Composition for Legged Robot Reinforcement Learning

    Sep 13, 2026Daniel Gigliotti, Flavio Maiorana, Fabio Patrizi +1

  3. Dual-Process Atomic Skill Learning: Decoupling Semantic Reasoning and Real-Time Control

    Date pendingJun Chen, Erdemt Bao, Wenlong Dong +7Robot Skill LearningImitation Learning