cs.ROOct 8, 2026

FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment

Authors: Filip Grigorov, Kourosh Darvish, Nandita Vijaykumar

Organizations: University of Toronto Canada

Abstract

Vision-based reinforcement learning for robotic manipulation is sample-inefficient because RGB-D observations are high-dimensional and noisy. Privileged state information available in simulation can accelerate training, but its absence at test time creates a train-test modality gap. We propose FOCUS, a single-stage PPO framework that trains the critic on privileged state while automatically regulating whether the actor collects rollouts from RGB-D or privileged-state latents. Regulation is driven by the KL divergence between the action distributions induced by the two modalities, while representation alignment encourages consistent action selection across them. Together, these mechanisms limit RGB-D rollouts when the actor's action distributions from RGB-D and privileged-state latents disagree. As they align, RGB-D exposure increases, shifting on-policy training toward the RGB-D inputs used at test time. Across five manipulation tasks, FOCUS raises average test success from 0.71 to 0.93 relative to the strongest RGB-D-at-test baseline on each task. When accounting for each method's complete training pipeline, budget-normalized training-success AUC increases from 0.47 to 0.65. On Pick-and-Place, test success rises from 0.47 to 0.86, while AUC increases from 0.12 to 0.61, a 5.0x improvement in learning efficiency over the fixed interaction budget.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation

    May 28, 2026Victor Kowalski, Chengxi Li, Dongheui LeeCross-Modal Knowledge DistillationRobot Skill Learning

  2. Benchmarking Action Spaces in Reinforcement Learning for Vision-based Robotic Manipulation

    Jun 17, 2026Seyed Alireza Azimi, Homayoon Farrahi, Abhishek Naik +2RL for RoboticsReinforcement Learning

  3. GeoProp: Grounding Robot State in Vision for Generalist Manipulation

    Jul 8, 2026Guoyang Zhao, Quanhao Qian, Gongjie Zhang +5Generalization in Robotic ManipulationVisuomotor Policy Learning