Rho: A Foundation for Efficiently Adaptable VLA Models
Organizations: Microsoft Research
Abstract
General-purpose physical AI models must combine broad visual and linguistic capabilities with precise control across robot embodiments and efficient adaptation to downstream tasks. We introduce Rho, a family of open-weights VLA models for bimanual manipulation designed for data-light task adaptation on 3 embodiments representative of dual-arm robots across research labs and the industry -- YAM Box, UR AI Trainer, and FR3 Duo. We systematically ablate Rho's action-expert architecture and training recipe, and show in controlled simulation and physical-robot experiments that embodiment midtraining improves downstream adaptation. The resulting Rho variants for YAM Box, UR AI Trainer, and FR3 Duo match or outperform existing open-weights VLAs and achieve the strongest overall performance across the tasks, embodiments, and baselines evaluated in this report. We further demonstrate the Rho model family's built-in capacity for online adaptation: an internal latent policy learns from corrective feedback to select observation-conditioned noise inputs for the frozen flow-matching action expert. With as few as 15 corrected episodes, adapting this lightweight module enables Rho to handle task situations at the fringe of its offline finetuning distribution. Together, these results position Rho as both a strong general-purpose robotic manipulation model and a practical foundation for adaptation. We release the base Rho model and the embodiment-specific checkpoints to facilitate Rho's deployment in research experiments and practical industrial use cases.
Figures & tables
| Component | Specification |
| Full vision-language model | 4.68B parameters |
| Language decoder | 32 layers; width 3,072; MLP width 8,192 |
| Decoder attention | 32 query heads; 32 key/value heads |
| Vocabulary and positions | 100,352 tokens; 16,384 configured positions |
| Vision encoder | SigLIP 2 SO400M NaFlex; 428M parameters |
| Image representation | patches; 256–3,600 visual tokens |
| Name | Description | Example |
| Attribute Similarity | Find the odd-one-out by font, color, or layout. | Which screenshot uses a different font? |
| Component Counting | Count a UI component across screenshots. | How many screenshots contain a card? |
| Object Counting | Count objects across natural images. | How many boats in all images? |
| Chart Type Matching | Find the odd-one-out by chart type. | Which chart is not a grouped bar chart? |
| Cross-Chart Reasoning | Combine evidence from related charts. | What patterns emerge across the charts? |
| Chart Difference Spotting | Identify chart changes and their effects. | What changed; how were profits affected? |
| Component | Specification |
| Action expert | 542M parameters |
| Transformer | 12 blocks; width 2,048; MLP width 4,096 |
| Attention | 16 query heads; 4 key/value heads; head dimension 128 |
| Block structure | Action self-attention; cross-attention to VLM context; feed-forward network |
| Flow conditioning | Shared adaLN-single with block-specific offsets |
| Flow target | Linear noise–action interpolation; velocity-prediction MSE |
| Variant | Width | Heads | Head dim. | Blocks | Success |
| 1,024 / 16 heads / 16 blocks | 1,024 | 16 | 64 | 16 | 0.617 |
| 1,024 / 8 heads / 16 blocks | 1,024 | 8 | 128 | 16 | 0.640 |
| 1,024 / 8 heads / 32 blocks | 1,024 | 8 | 128 | 32 | 0.642 |
| 2,048 / 8 heads / 16 blocks | 2,048 | 8 | 256 | 16 | 0.623 |
| 2,048 / 16 heads / 16 blocks | 2,048 | 16 | 128 | 16 | 0.673 |
| Variant | Q/KV heads | Blocks | AdaLN | AE params. | Success |
| Wide baseline | 16/16 | 16 | per-block | 1.649B | 0.673 |
| Shared AdaLN | 16/16 | 16 | shared | 882M | 0.673 |
| Grouped-query | 16/4 | 16 | shared | 693M | 0.672 |
| Rho | 16/4 | 12 | shared | 542M | 0.671 |
| Pretraining learning rate | 20k adaptation | 30k adaptation | 40k adaptation |
| Model | Succ. | TP | BGVD | CPL | ECC | JPL | OPL | SCC | SC |
| MolmoAct2 | 0.61 | 0.79 | 0.072 | 1.48 | 1.03 | 7.97 | 4.84 | 0.017 | 0.465 |
| GR00T N1.7 | 0.61 | 0.80 | 0.078 | 1.66 | 1.25 | 9.64 | 5.43 | 0.095 | 0.540 |
| 0.67 | 0.82 | 0.074 | 1.45 | 0.99 | 7.98 | 4.82 | 0.038 | 0.482 | |
| Rho | 0.73 | 0.85 | 0.073 | 1.41 | 1.02 | 7.70 | 4.69 | 0.035 | 0.477 |
| Model | Spatial | Object | Goal | Long | Average |
| OpenVLA [ 23 ] | 84.7 | 88.4 | 79.2 | 53.7 | 76.5 |
| [ 4 ] | 96.8 | 98.8 | 95.8 | 85.2 | 94.2 |
| MolmoAct-7B-D [ 29 ] | 87.0 | 95.4 | 87.6 | 77.2 | 86.6 |
| GR00T N1.7 [ 42 ] | 95.0 | 100.0 | 98.0 | 93.0 | 96.5 |
| [ 48 ] | 98.8 | 98.2 | 98.0 | 92.4 | 96.9 |
| MolmoAct2 [ 14 ] | 97.8 | 100.0 | 97.8 | 93.2 | 97.2 |
| Task | Offline | + Online | |
| Assembly | 45.0 | 78.0 | +33.0 |
| Stick pull | 58.0 | 78.0 | +20.0 |
| Pick out of hole | 22.0 | 37.0 | +15.0 |
| Hand insert | 73.0 | 86.0 | +13.0 |
| Hammer | 92.0 | 100.0 | +8.0 |
| Basketball | 55.0 | 61.0 | +6.0 |
| Offline | + Online | |||
| Task | SR | TP | SR | TP |
| Test-tube assembly | 30.0 | 56.7 | 70.0 | 81.7 |
| Plug insertion | 66.7 | 83.3 | 93.3 | 96.7 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Configuration |
| Vision-language backbone | Phi-Phy; 4.68B parameters |
| Language decoder | 32 causal-transformer blocks; width 3,072; MLP width 8,192; 32 query and 32 key/value heads |
| Text interface | Vocabulary size 100,352; configured context length 16,384 |
| Vision encoder | SigLIP 2 SO400M NaFlex; 428M parameters; patches; penultimate-layer visual features |
| Visual-token budget | 256–3,600 tokens per image; robot images are resized with padding to , yielding 256 visual patches per view |
| Cross-modal projector | Two-layer MLP, , with GELU |
| Component | Configuration |
| Action stream | One projected state token followed by one noisy-action token for every position in the action chunk |
| Transformer | 12 blocks; width 2,048; feed-forward width 4,096; GELU activation; dropout 0.1 during pretraining |
| Block structure | Action-stream self-attention, cross-attention to the fixed VLM context, then a feed-forward update, each with a gated residual |
| Attention | 16 attention heads, 4 key/value heads, head dimension 128; grouped-query attention is used in both attention modules |
| Normalization | DiT-style adaLN-Zero; shared adaLN-single base maps with zero-initialized block- and site-specific offsets |
| Position encoding | Fixed sinusoidal encoding; configured maximum sequence length 6,144 |
| Signal | Source representation | Mapping | Use in action expert |
| Visual-language context | Block-14 Phi-Phy hidden sequence, width 3,072 | Learned linear | Shared as the key/value sequence for cross-attention in every expert block |
| Robot state | One state vector padded to width 32 | Learned linear | Prepended to the action-query sequence as one token |
| Noisy action | flow state | Linear action projection followed by a action–time MLP | Supplies the remaining query tokens |
| Flow time | Scalar with a 2,048-dimensional sinusoidal embedding | SiLU MLP | Conditions the shared adaLN maps in every block |
| Padding masks | Valid image/text tokens, state/action dimensions, and chunk positions | No learned mapping | Masks attention or loss terms associated with padding |
| Component | Parameters | Pretraining status |
| Phi-Phy vision–language backbone | 4.68B | Trainable except token embeddings |
| SigLIP 2 vision encoder | 428M | Trainable; included above |
| Token-embedding table | 308M | Frozen; included above |
| Action expert | 542M | Trainable from random initialization |
| Complete Rho model | 5.22B | Approximately 4.91B trainable |
| Inference stage | Operation |
| Observation encoding | Encode all camera views and the instruction with Phi-Phy, extract decoder block 14, and project the valid context tokens to width 2,048. This backbone pass occurs once per policy query. |
| Latent initialization | Draw an independent Gaussian tensor ; an internal learned latent policy can replace this draw during online adaptation ( Appendix F ). |
| Flow integration | Apply 10 uniform explicit-Euler steps from to . Each step re-embeds the current action latent and queries the 12-block action expert while reusing the cached VLM context. |
| Action decoding | Project the final action tokens to 32 dimensions, remove padded dimensions, and invert the embodiment-specific normalization. No autoregressive text decoding is used to produce robot actions. |
| Receding-horizon execution | Execute a configurable prefix and then replan. Multi-embodiment pretraining uses and a 25-action execution horizon; RoboEval uses 32/16 and LIBERO uses 16/8. |
| Numerics | Bfloat16 model execution with FlashAttention 2; the Euler state update is accumulated in float32 before conversion back to bfloat16. |
| Component | Configuration |
| Initialization | Phi-Phy backbone; randomly initialized state/action projections and action expert |
| Trainable parameters | Vision encoder, cross-modal projector, language decoder, and action modules; token embeddings frozen |
| Robot inputs | One observation; up to three RGB views; padded images; task instruction and proprioceptive state |
| Action targets | Chunk-relative EEF translation and 6D-rotation deltas; absolute grippers; approximately 1 s at the native rate; at most 50 steps; per-position 1st/99th-percentile normalization |
| Training objectives | Flow matching on robot batches; autoregressive cross-entropy on VQA and pointing batches, weighted by |
| Batch sampling | Robot/VL update probability ; global batch 3,072/384 |
| Component | Configuration |
| Initialization | First-epoch multi-embodiment Rho checkpoint |
| Data mixture | FR3 Duo multi-task robot data and retained vision–language data, sampled at a 9:1 robot/VL ratio |
| Trainable parameters | Vision encoder, cross-modal projector, language decoder, and action modules; token embeddings frozen |
| Action representation | One-second chunk of chunk-relative Cartesian end-effector translation and 6D-rotation deltas with absolute gripper commands, padded to 50 positions |
| Batch and updates | 128 examples per GPU on 8 B200 GPUs (global batch 1,024); 165,000 optimizer updates |
| Optimizer | AdamW; , ; ; weight decay |
| Component | Configuration |
| Initialization | FR3 Duo embodiment-midtrained Rho checkpoint |
| Data | Target-task demonstrations only; three RGB views, language, and proprioception |
| Trainable parameters | Vision encoder, cross-modal projector, language decoder, and action modules; token embeddings frozen |
| Action representation | One-second chunk of chunk-relative Cartesian end-effector translation and 6D-rotation deltas with absolute gripper commands, padded to a 50-position model chunk |
| Normalization | Per-task action-chunk mean and standard deviation |
| Batch and updates | 16 examples per GPU on 4 B200 GPUs (global batch 64); 50,000 optimizer updates |
| Benchmark | Tasks | Adaptation | Trials / task / pass | Evaluation passes |
| LIBERO | 40 (four suites) | 40k updates | 50 | 3 |
| RoboEval | 8 | 40k updates | 100 | 3 |
| RoboTwin | 5 (each under Easy & Hard settings) | 40k updates | 50 | 3 (Hard)/4 (Easy) |
| Suite | Tasks |
| Spatial | Move the black bowl to the plate from: between the plate and ramekin; next to the ramekin; the table center; atop the cookie box; the cabinet’s top drawer; atop the ramekin; next to the cookie box; the stove; next to the plate; and atop the wooden cabinet. |
| Object | Place in the basket: alphabet soup; cream cheese; salad dressing; BBQ sauce; ketchup; tomato sauce; butter; milk; chocolate pudding; and orange juice. |
| Goal | Open the cabinet’s middle drawer; put the bowl on the stove; put the wine bottle atop the cabinet; open the top drawer and put the bowl inside; put the bowl atop the cabinet; push the plate in front of the stove; put the cream cheese in the bowl; turn on the stove; put the bowl on the plate; put the wine bottle on the rack. |
| Long | Put alphabet soup and tomato sauce in the basket; put cream cheese and butter in the basket; turn on the stove and put the moka pot on it; put the black bowl in the bottom drawer and close it; put the white mug on the left plate and the yellow-and-white mug on the right plate; put the book in the caddy’s rear compartment; put the white mug on the plate and the chocolate pudding to its right; put alphabet soup and cream cheese in the basket; put both moka pots on the stove; put the yellow-and-white mug in the microwave and close it. |
| Hyperparameter | Meta-World | FR3 Duo |
| Correction source | Scripted expert | Human (SpaceMouse) |
| Online episodes per task | 50 | 15 |
| Action/latent horizon | 16 | 16 |
| Latent dimension per step | 32 | 32 |
| Latent bound | 3.0 | 3.0 |
| Inversion method | Per-step fixed point | Per-step fixed point |
| Task | Model | Succ. | TP | BGVD | CPL | ECC | JPL | OPL | SCC | SC |
| Cube handover | GR00T N1.7 | 0.64 | 0.93 | 0.043 | 1.35 | 0.16 | 6.81 | 4.14 | 0.373 | 0.477 |
| 0.79 | 0.92 | 0.038 | 1.18 | 0.10 | 5.44 | 3.37 | 0.210 | 0.237 | ||
| MolmoAct2 | 0.86 | 0.95 | 0.036 | 1.11 | 0.04 | 5.12 | 3.30 | 0.107 | 0.193 | |
| Rho | 0.80 | 0.93 | 0.037 | 1.18 | 0.09 | 5.15 | 3.14 | 0.207 | 0.217 | |
| Lift pot | GR00T N1.7 | 0.67 | 0.81 | 0.058 | 2.57 | 0.13 | 16.83 | 9.70 | 0.120 | 0.603 |
| 0.72 | 0.84 | 0.029 | 1.48 | 0.00 | 9.39 | 6.80 | 0.017 | 0.417 |