FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment
Organizations: University of Toronto Canada
Abstract
Vision-based reinforcement learning for robotic manipulation is sample-inefficient because RGB-D observations are high-dimensional and noisy. Privileged state information available in simulation can accelerate training, but its absence at test time creates a train-test modality gap. We propose FOCUS, a single-stage PPO framework that trains the critic on privileged state while automatically regulating whether the actor collects rollouts from RGB-D or privileged-state latents. Regulation is driven by the KL divergence between the action distributions induced by the two modalities, while representation alignment encourages consistent action selection across them. Together, these mechanisms limit RGB-D rollouts when the actor's action distributions from RGB-D and privileged-state latents disagree. As they align, RGB-D exposure increases, shifting on-policy training toward the RGB-D inputs used at test time. Across five manipulation tasks, FOCUS raises average test success from 0.71 to 0.93 relative to the strongest RGB-D-at-test baseline on each task. When accounting for each method's complete training pipeline, budget-normalized training-success AUC increases from 0.47 to 0.65. On Pick-and-Place, test success rises from 0.47 to 0.86, while AUC increases from 0.12 to 0.61, a 5.0x improvement in learning efficiency over the fixed interaction budget.
Figures & tables
| Method | Reach | Lift | Push | Open - Drawer | Pick - and - Place † |
| states-only | 1.000 0.000 | 0.555 0.628 | 0.972 0.040 | 1.000 0.000 | 1.000 0.000 |
| RGB-D-only | 1.000 0.000 | 0.233 0.024 | 0.753 0.098 | 0.983 0.024 | 0.453 0.122 |
| RGB-D-asymmetric | 1.000 0.000 | 0.350 0.338 | 0.192 0.059 | 0.811 0.157 | 0.467 0.660 |
| teacher-student (3-stage, final student) | 1.000 0.000 | 0.181 0.122 | 0.595 0.032 | 0.978 0.031 | 0.289 0.110 |
| FOCUS (ours) | 1.000 0.000 | 0.975 0.012 | 0.850 0.071 | 0.978 0.018 | 0.858 0.069 |
| Method | Reach | Lift | Push | Open - Drawer | Pick - and - Place |
| states-only | 0.966 0.002 | 0.211 0.268 | 0.447 0.007 | 0.912 0.016 | 0.656 0.009 |
| RGB-D-only | 0.889 0.003 | 0.023 0.002 | 0.288 0.036 | 0.859 0.021 | 0.094 0.056 |
| RGB-D-asymmetric | 0.891 0.020 | 0.158 0.190 | 0.094 0.003 | 0.886 0.015 | 0.122 0.171 |
| teacher-student (final student) | 0.958 0.010 | 0.258 0.321 | 0.233 0.012 | 0.930 0.013 | 0.155 0.003 |
| teacher-student (3-stage, e2e) | 0.411 0.004 | 0.111 0.137 | 0.100 0.005 | 0.399 0.006 | 0.071 0.001 |
| FOCUS (ours) | 0.969 0.001 | 0.451 0.040 | 0.337 0.167 | 0.864 0.025 | 0.612 0.037 |
| Variant | Test success | Norm. Training Success AUC | ||||
| Push | Lift | Pick - and - Place † | Push | Lift | Pick - and - Place | |
| FOCUS | [-2pt] | [-2pt] | [-2pt] | [-2pt] | [-2pt] | [-2pt] |
| Training-objective ablations | ||||||
| w/o decoder | [-2pt] | [-2pt] | [-2pt] | [-2pt] | [-2pt] | [-2pt] |
| w/o aligner ( bypass) | [-2pt] | [-2pt] | [-2pt] | [-2pt] | [-2pt] | [-2pt] |
| w/o RL agreement | [-2pt] | [-2pt] | [-2pt] | [-2pt] | [-2pt] | [-2pt] |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Reach | Lift | Push | Open - Drawer | Pick - and - Place |
| states-only | 18.5 6.4 (2/2) | 693 (1/2) | 556.5 27.6 (2/2) | 31.5 3.5 (2/2) | 522.0 46.7 (2/2) |
| RGB-D-only | 15.5 9.2 (2/2) | Not reached (0/2) | 1447.5 20.5 (2/2) | 51.5 4.9 (2/2) | Not reached (0/2) |
| RGB-D-asymmetric | 12.5 5.0 (2/2) | 1098 (1/2) | Not reached (0/2) | 50.0 4.2 (2/2) | 1611 (1/2) |
| teacher-student (final student) | 1.0 1.4 (2/2) | 262 (1/2) | Not reached (0/2) | 5.0 5.7 (2/2) | Not reached (0/2) |
| teacher-student (3-stage, e2e) | 2001.0 1.4 (2/2) | 2262 (1/2) | Not reached (0/2) | 2005.0 5.7 (2/2) | Not reached (0/2) |
| FOCUS (ours) | 11.5 3.5 (2/2) | 461.5 112.4 (2/2) | 461.5 19.1 (2/2) | 91.5 20.5 (2/2) | 470.0 107.5 (2/2) |
| Training and PPO defaults | |
| Latent dimension | |
| PPO epochs | |
| PPO mini-batches | |
| Clip parameter | |
| Discount | |
| GAE parameter | |
| Additional PPO / optimization details | |
| Value loss weight | |
| Entropy weight | |
| Clipped value loss | enabled |
| PPO KL target | when adaptive LR scheduling is enabled |
| RL optimizer details | Adam, betas , weight decay |
| Representation optimizer details | Adam, betas , weight decay |
| Architecture summary (latent dim ) | |
| State encoder | Linear Swish LayerNorm. |
| RGB-D encoder | Shared-trunk ResNet-18 RGB-D encoder. Separate RGB/depth stems (RGB Conv1 from pretrained ResNet-18 [ 4 ] ; depth Conv1 initialized from averaged RGB weights, with optional validity-mask channel) summed stems shared ResNet-18 trunk (BN/ReLU/MaxPool, Layer1–Layer4) ECA channel attention [ 29 ] (+ optional CBAM-style SpatialGate [ 31 ] ) AdaptiveAvgPool2d Linear LayerNorm (projection: Identity). |
| Fusion | Residual fusion of a chosen latent with proprioception and previous action: concat LayerNorm Linear (zero-init) residual add . ( Same fusion block used for actor and critic inputs; critic always uses .) |
| Aligner | Symmetric cross-attention aligner operating on raw latents (2 heads). Per stream: CrossAttn concat LayerNorm Linear gated residual update. |
| State decoder | Residual MLP decoder from raw latent to privileged-state target: with Swish activations, Dropout after the first layer, final LayerNorm, and residual Linear . |
| Scratch CNN ablation | For the “RGB-D encoder from scratch” ablation, the ResNet-based is replaced by a lightweight two-branch RGB-D CNN. The RGB and depth streams each use three stride-2 Conv–Swish blocks with channels , followed by AdaptiveAvgPool2d , flattening, Linear , and LayerNorm. The two -dimensional branch features are concatenated and passed through a two-layer MLP to produce the final -dimensional image latent. The same RGB-D preprocessing contract is used as in the ResNet-based encoder. |
| Signal (TB tag) | Phase | What it captures |
| Performance & sample efficiency | ||
| Train/success_rate_ema | train | smoothed training success vs. env steps |
| Train/steps_to_50pct | train | Steps@50 (time-to-threshold) |
| Test/Success_rate | test | evaluation success rate over episodes/seeds |
| Test/rewards_mean | test | mean episode return (task-specific scale) |
| PPO update stability | ||