cs.ROOct 4, 2026

Recursive Self-Improvement of Visuomotor Policies through Local Recovery Supervision

Authors: Yuzhi Zhang, Xinyu Liu, Yu Zhang

Organizations: Beihang University

Abstract

Visuomotor policies can execute familiar tasks yet lack the corrective behavior needed after their own mistakes. We present a framework for recursive self-improvement through local recovery supervision. Each round audits the current policy, generates corrective demonstrations at supported failure states, and uses them to update the policy that drives the next round of collection. An offline auditor locates unresolved failures using coarse and dense temporal evidence and specifies observable repair goals. A fixed multimodal agent acts as a tool-using teacher, generating recovery actions through observation, computation, execution, and feedback. The frozen student tests whether each teacher endpoint supports further progress. If continuation fails, the system restores that endpoint and extends the demonstration. Action-level quality assessment then defines continuous training windows with aligned observations, quality weights, and validity masks. Only the student is deployed. In a preliminary LIBERO-Goal study, recovery-augmented post-training achieves 88 successful episodes out of 100 validation scenes, compared with 78 for original-data continuation from the same π0π_0 checkpoint. An earlier BC-RNN study on robomimic Can improves success from 102/130 to 112/130 using 26 local recovery segments. Both comparisons match 2,000 additional optimization steps.

Figures & tables

Explore similar work

CardsList
  1. Rewind-IL: Online Failure Detection and State Respawning for Imitation Learning

    Apr 17, 2026Gehan Zheng, Sanjay Seenivasan, Matthew Johnson-Roberson +1Imitation LearningReverse

  2. Recova: Agent-Guided Failure Recovery for Autonomous Robotic Manipulation

    Oct 1, 2026Isabella Liu, An-Chieh Cheng, Johan Bjorck +6

  3. FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement

    Jul 1, 2026Haoran Hao, Shahram Najam Syed, Jeffrey Ichnowski +1Robot PoliciesRetrying