General-purpose physical AI models must combine broad visual and linguistic capabilities with precise control across robot embodiments and efficient adaptation to downstream tasks. We introduce Rho, a family of open-weights VLA models for bimanual manipulation designed for data-light task adaptation on 3 embodiments representative of dual-arm robots across research labs and the industry -- YAM Box, UR AI Trainer, and FR3 Duo. We systematically ablate Rho's action-expert architecture and training recipe, and show in controlled simulation and physical-robot experiments that embodiment midtraining improves downstream adaptation. The resulting Rho variants for YAM Box, UR AI Trainer, and FR3 Duo match or outperform existing open-weights VLAs and achieve the strongest overall performance across the tasks, embodiments, and baselines evaluated in this report. We further demonstrate the Rho model family's built-in capacity for online adaptation: an internal latent policy learns from corrective feedback to select observation-conditioned noise inputs for the frozen flow-matching action expert. With as few as 15 corrected episodes, adapting this lightweight module enables Rho to handle task situations at the fringe of its offline finetuning distribution. Together, these results position Rho as both a strong general-purpose robotic manipulation model and a practical foundation for adaptation. We release the base Rho model and the embodiment-specific checkpoints to facilitate Rho's deployment in research experiments and practical industrial use cases.
Figures & tables
Figure 1 : Rho’s physical dual-arm robot embodiments. YAM Box is common in research labs, FR3 Duo’s usage straddles research and industry, and UR AI Trainer is an emerging standard setup for industrial bimanual manipulation tasks. Midtrained checkpoints for all three are available at https://huggingface.co/collections/microsoft/rho
Component
Specification
Full vision-language model
4.68B parameters
Language decoder
32 layers; width 3,072; MLP width 8,192
Decoder attention
32 query heads; 32 key/value heads
Vocabulary and positions
100,352 tokens; 16,384 configured positions
Vision encoder
SigLIP 2 SO400M NaFlex; 428M parameters
Image representation
16×16 patches; 256–3,600 visual tokens
Table 1: Architecture of the Phi-Phy vision-language backbone used to initialize Rho.
Name
Description
Example
Attribute Similarity
Find the odd-one-out by font, color, or layout.
Which screenshot uses a different font?
Component Counting
Count a UI component across screenshots.
How many screenshots contain a card?
Object Counting
Count objects across natural images.
How many boats in all images?
Chart Type Matching
Find the odd-one-out by chart type.
Which chart is not a grouped bar chart?
Cross-Chart Reasoning
Combine evidence from related charts.
What patterns emerge across the charts?
Chart Difference Spotting
Identify chart changes and their effects.
What changed; how were profits affected?
Table 2: Multi-image data categories and example VQA questions.
Figure 2 : The Rho flow-matching action expert. State and noisy-action tokens pass through 12 wide transformer blocks with grouped-query self-attention, cross-attention to the VLM context, and feed-forward updates. The flow timestep conditions the sublayers through shared adaLN-Zero maps with block-specific offsets.
Component
Specification
Action expert
542M parameters
Transformer
12 blocks; width 2,048; MLP width 4,096
Attention
16 query heads; 4 key/value heads; head dimension 128
Block structure
Action self-attention; cross-attention to VLM context; feed-forward network
Flow conditioning
Shared adaLN-single with block-specific offsets
Flow target
Linear noise–action interpolation; velocity-prediction MSE
Table 3: Configuration of the Rho flow-matching action expert.
Figure 3 : Rho-Tomyum mixture. Its 90% is composed of demonstration trajectories from physical and simulated robots as well as UMI-style demonstrations collected with standard YAM grippers, and the remaining 10% are vision-language data. All the robot/UMI-style data is bimanual, with the exception of a small subset of Open X-Embodiment Magic Soup [ 4 ] we call OXE Amuse Bouche . The robot data includes trajectories from policies controlling a UR AI Trainer simulated in NVIDIA Isaac Sim, trained using RL and the OmniReset framework [ 63 ] .
Figure 4 : Midtraining mixtures of the embodiment-specific Rho variants. The count of distinct language instructions includes the number of both task and subtask descriptions.
Variant
Width
Heads
Head dim.
Blocks
Success
1,024 / 16 heads / 16 blocks
1,024
16
64
16
0.617
1,024 / 8 heads / 16 blocks
1,024
8
128
16
0.640
1,024 / 8 heads / 32 blocks
1,024
8
128
32
0.642
2,048 / 8 heads / 16 blocks
2,048
8
256
16
0.623
2,048 / 16 heads / 16 blocks
2,048
16
128
16
0.673
Table 4: Action-expert shape sweep on RoboEval after 40k adaptation steps. Success rates are averaged over eight tasks and three evaluation passes.
Variant
Q/KV heads
Blocks
AdaLN
AE params.
Success
Wide baseline
16/16
16
per-block
1.649B
0.673
Shared AdaLN
16/16
16
shared
882M
0.673
Grouped-query
16/4
16
shared
693M
0.672
Rho
16/4
12
shared
542M
0.671
Table 5: Action-expert compression on RoboEval while preserving width 2,048 and head dimension 128. Success rates are averaged over eight tasks after 40k adaptation steps. Each variant uses three evaluation passes.
Figure 5 : Pretraining on 15% of the dataset mixture improves downstream adaptation. Left: LIBERO suite success after 20k task-adaptation steps. Right: RoboEval success throughout adaptation. The robot+VL condition combines robot trajectories with physical-grounding examples, the robot-data-only condition removes the physical-grounding batches, and the no-pretraining condition begins task adaptation from the grounded backbone with a randomly initialized action expert. Values are means and standard deviations over three evaluations, except the robot+VL RoboEval result at 40k adaptation steps, which pools six passes from two independent evaluations.
Pretraining learning rate
20k adaptation
30k adaptation
40k adaptation
5×10−5
0.601±0.028
0.645±0.005
0.641±0.006
1×10−4
0.623±0.031
0.675±0.035
0.673±0.015
2×10−4
0.586±0.010
0.629±0.031
0.633±0.008
5×10−4
0.548±0.006
0.637±0.021
0.622±0.022
Table 6: Effect of the action-expert learning rate during pretraining on downstream RoboEval adaptation. All rows are pretrained on the same 15% subset of the dataset mixture and include physical-grounding data. The results are mean success and standard deviation over three evaluation passes.
Figure 6 : Embodiment midtraining on RoboTwin transfers to held-out tasks across downstream-data budgets. Each checkpoint is adapted for 40k steps and evaluated on the matched five-task protocol with 50 episodes per task. Bars are means over completed seeds (four for Easy, three for Hard) and error bars are sample standard deviations. The benefit of midtraining is largest in the low-data regime and shrinks as the adaptation set grows.
Figure 7 : FR3 Duo horizontal midtraining improves subsequent vertical adaptation to bimanual plug insertion. Every condition uses the same 50k-step task-adaptation budget; success reported over 30 trials.
Figure 8 : Success rates of finetuned Rho, π0.5 , GR00T-N1.7, and MolmoAct2 on RoboEval tasks. Rho outperforms all other models on Pick Book , Stack Book Shelf , and Rotate Valve . On each of the remaining 5 tasks, Rho matches the performance of the best baseline on that task. Importantly, each task’s best-performing baseline is generally different, so the next-strongest model, π0.5 , matches Rho’s performance on only 4 out of 8 tasks.
Model
Succ. ↑
TP ↑
BGVD ↓
CPL ↓
ECC ↓
JPL ↓
OPL ↓
SCC ↓
SC ↓
MolmoAct2
0.61
0.79
0.072
1.48
1.03
7.97
4.84
0.017
0.465
GR00T N1.7
0.61
0.80
0.078
1.66
1.25
9.64
5.43
0.095
0.540
π0.5
0.67
0.82
0.074
1.45
0.99
7.98
4.82
0.038
0.482
Rho
0.73
0.85
0.073
1.41
1.02
7.70
4.69
0.035
0.477
Table 7: Aggregate RoboEval outcome and behavioral metrics after 40k task-adaptation steps with global batch size 128. Rho outperforms the baselines on 5 out of the 9 metrics, with the next-strongest model across these metrics, MolmoAct2, winning on 3. Every row is the mean of 3 800-episode evaluations across the 8 task variations. Joint Path length (JPL) measures the sum of joint path lengths traced by the two arms. Arrows indicate the preferred direction of each metric, e.g., “ ↓ ” means that lower is better.
Model
Spatial
Object
Goal
Long
Average
OpenVLA [ 23 ]
84.7
88.4
79.2
53.7
76.5
π0 [ 4 ]
96.8
98.8
95.8
85.2
94.2
MolmoAct-7B-D [ 29 ]
87.0
95.4
87.6
77.2
86.6
GR00T N1.7 [ 42 ]
95.0
100.0
98.0
93.0
96.5
π0.5 [ 48 ]
98.8
98.2
98.0
92.4
96.9
MolmoAct2 [ 14 ]
97.8
100.0
97.8
93.2
97.2
Table 8: Rho’s mean performance and published success rates of several baselines on the 4 task suites of LIBERO [ 34 ] . Rho outperforms the baselines in overall success rate aggregated across all 4 task suites and in success rate on the most challenging suite, LIBERO-long . Baseline values are taken from the corresponding model report or official release as cited in column 1; evaluation and adaptation protocols may differ across the reports. For Rho, the mean success rate is over 3 seeds. All values are percentages. Bold denotes the best result in each column.
Figure 9 : The 6 BusyBox task categories used in the YAM Box experiments. Bimanual tasks require two-arm coordination; dual-arm tasks require choosing the correct arm for task completion.
Figure 10 : YAM Box BusyBox results. After task-specific finetuning, Rho-YAM-Box achieves the same 90% mean success rate as π0.5 and a higher mean success rate than GR00T-N1.7 and MolmoAct2 across 60 matched evaluation rollouts.
Figure 11 : The UR AI Trainer tasks and their stages. Each row shows task progression from left to right.
Figure 12 : UR AI Trainer results. Rho-UR-AI-Trainer achieves higher success than both baselines on both tasks. On Small-electronics cleanup , π0.5 achieves similar mean task progress (80.8% versus 82.5%) despite lower success (40.0% versus 60.0%). π0.5 ’s failures tend to occur during final placement, whereas Rho’s failures more often occur during picking.
Figure 13 : The FR3 Duo tasks and their stages. The Test-tube and Tumbler tasks are evaluated in both directions: their frames read from left to right for the task on the upper arrow and from right to left for the task on the lower arrow.
Figure 14 : FR3 Duo task-adaptation results. Each model is adapted independently on each task dataset and evaluated over 30 distinct, matched initial configurations per variant. We report binary success rate and mean task progress; Overall is the unweighted mean across the five task variants.
Task
Offline
+ Online
Δ
Assembly
45.0
78.0
+33.0
Stick pull
58.0
78.0
+20.0
Pick out of hole
22.0
37.0
+15.0
Hand insert
73.0
86.0
+13.0
Hammer
92.0
100.0
+8.0
Basketball
55.0
61.0
+6.0
Table 9: Online latent-space adaptation on Meta-World. The offline policy is Rho trained with a limited budget on all 50 Meta-World tasks; the online column adapts it to each task in latent space from 50 episodes supervised by a privileged scripted expert. Success is measured over 100 rollouts per task; all values are percentages.
Offline
+ Online
Task
SR
TP
SR
TP
Test-tube assembly
30.0
56.7
70.0
81.7
Plug insertion
66.7
83.3
93.3
96.7
Table 10: Online latent-space adaptation on FR3 Duo. Offline policies are adapted from the Rho-FR3-Duo checkpoint on 150 demonstrations per task; online adaptation uses 15 additional human-supervised episodes. Each task is evaluated over 30 trials, three on each of 10 hard initial configurations. SR is success rate and TP is task progress; both are percentages.
Action-stream self-attention, cross-attention to the fixed VLM context, then a feed-forward update, each with a gated residual
Attention
16 attention heads, 4 key/value heads, head dimension 128; grouped-query attention is used in both attention modules
Normalization
DiT-style adaLN-Zero; shared adaLN-single base maps with zero-initialized block- and site-specific offsets
Position encoding
Fixed sinusoidal encoding; configured maximum sequence length 6,144
Appendix
Table 12: Configuration of the Rho flow-matching action expert.
Signal
Source representation
Mapping
Use in action expert
Visual-language context
Block-14 Phi-Phy hidden sequence, width 3,072
Learned linear 3,072→2,048
Shared as the key/value sequence for cross-attention in every expert block
Robot state
One state vector padded to width 32
Learned linear 32→2,048
Prepended to the action-query sequence as one token
Noisy action
H×32 flow state
Linear action projection followed by a 4,096→2,048→2,048 action–time MLP
Supplies the remaining H query tokens
Flow time
Scalar t with a 2,048-dimensional sinusoidal embedding
2,048→2,048→2,048 SiLU MLP
Conditions the shared adaLN maps in every block
Padding masks
Valid image/text tokens, state/action dimensions, and chunk positions
No learned mapping
Masks attention or loss terms associated with padding
Appendix
Table 13: The interface between Phi-Phy and the continuous action expert. The backbone context is computed once per observation and reused by all flow integration steps.
Component
Parameters
Pretraining status
Phi-Phy vision–language backbone
4.68B
Trainable except token embeddings
SigLIP 2 vision encoder
428M
Trainable; included above
Token-embedding table
308M
Frozen; included above
Action expert
542M
Trainable from random initialization
Complete Rho model
5.22B
Approximately 4.91B trainable
Appendix
Table 14: Parameter accounting for Rho. Counts are rounded to the precision used throughout the paper.
Inference stage
Operation
Observation encoding
Encode all camera views and the instruction with Phi-Phy, extract decoder block 14, and project the valid context tokens to width 2,048. This backbone pass occurs once per policy query.
Latent initialization
Draw an independent Gaussian tensor x1∈RH×32 ; an internal learned latent policy can replace this draw during online adaptation ( Appendix F ).
Flow integration
Apply 10 uniform explicit-Euler steps from t=1 to t=0 . Each step re-embeds the current action latent and queries the 12-block action expert while reusing the cached VLM context.
Action decoding
Project the final action tokens to 32 dimensions, remove padded dimensions, and invert the embodiment-specific normalization. No autoregressive text decoding is used to produce robot actions.
Receding-horizon execution
Execute a configurable prefix and then replan. Multi-embodiment pretraining uses H=50 and a 25-action execution horizon; RoboEval uses 32/16 and LIBERO uses 16/8.
Numerics
Bfloat16 model execution with FlashAttention 2; the Euler state update is accumulated in float32 before conversion back to bfloat16.
Appendix
Table 15: Rho inference path. Prediction and execution horizons are adapted to the control frequency and temporal extent of each embodiment.
Component
Configuration
Initialization
Phi-Phy backbone; randomly initialized state/action projections and action expert
Trainable parameters
Vision encoder, cross-modal projector, language decoder, and action modules; token embeddings frozen
Robot inputs
One observation; up to three RGB views; padded 256×256 images; task instruction and proprioceptive state
Action targets
Chunk-relative EEF translation and 6D-rotation deltas; absolute grippers; approximately 1 s at the native rate; at most 50 steps; per-position 1st/99th-percentile normalization
Training objectives
Flow matching on robot batches; autoregressive cross-entropy on VQA and pointing batches, weighted by 0.02
Batch sampling
Robot/VL update probability 0.9/0.1 ; global batch 3,072/384
Appendix
Table 16: Multi-embodiment pretraining recipe for Rho. Batch sizes count examples per modality-homogeneous optimization update.
Component
Configuration
Initialization
First-epoch multi-embodiment Rho checkpoint
Data mixture
FR3 Duo multi-task robot data and retained vision–language data, sampled at a 9:1 robot/VL ratio
Trainable parameters
Vision encoder, cross-modal projector, language decoder, and action modules; token embeddings frozen
Action representation
One-second chunk of chunk-relative Cartesian end-effector translation and 6D-rotation deltas with absolute gripper commands, padded to 50 positions
Batch and updates
128 examples per GPU on 8 B200 GPUs (global batch 1,024); 165,000 optimizer updates
Target-task demonstrations only; three RGB views, language, and proprioception
Trainable parameters
Vision encoder, cross-modal projector, language decoder, and action modules; token embeddings frozen
Action representation
One-second chunk of chunk-relative Cartesian end-effector translation and 6D-rotation deltas with absolute gripper commands, padded to a 50-position model chunk
Normalization
Per-task action-chunk mean and standard deviation
Batch and updates
16 examples per GPU on 4 B200 GPUs (global batch 64); 50,000 optimizer updates
Appendix
Table 18: Representative FR3 Duo task-finetuning recipe.
Benchmark
Tasks
Adaptation
Trials / task / pass
Evaluation passes
LIBERO
40 (four suites)
40k updates
50
3
RoboEval
8
40k updates
100
3
RoboTwin
5 (each under Easy & Hard settings)
40k updates
50
3 (Hard)/4 (Easy)
Appendix
Table 19: LIBERO, RoboEval, and RoboTwin task-adaptation and evaluation budgets for the reported Rho results. Each pass is a complete evaluation over the corresponding task set.
Suite
Tasks
Spatial
Move the black bowl to the plate from: between the plate and ramekin; next to the ramekin; the table center; atop the cookie box; the cabinet’s top drawer; atop the ramekin; next to the cookie box; the stove; next to the plate; and atop the wooden cabinet.
Object
Place in the basket: alphabet soup; cream cheese; salad dressing; BBQ sauce; ketchup; tomato sauce; butter; milk; chocolate pudding; and orange juice.
Goal
Open the cabinet’s middle drawer; put the bowl on the stove; put the wine bottle atop the cabinet; open the top drawer and put the bowl inside; put the bowl atop the cabinet; push the plate in front of the stove; put the cream cheese in the bowl; turn on the stove; put the bowl on the plate; put the wine bottle on the rack.
Long
Put alphabet soup and tomato sauce in the basket; put cream cheese and butter in the basket; turn on the stove and put the moka pot on it; put the black bowl in the bottom drawer and close it; put the white mug on the left plate and the yellow-and-white mug on the right plate; put the book in the caddy’s rear compartment; put the white mug on the plate and the chocolate pudding to its right; put alphabet soup and cream cheese in the basket; put both moka pots on the stove; put the yellow-and-white mug in the microwave and close it.
Appendix
Table 20 : LIBERO task inventory used in our simulation evaluation. Spatial and Object hold the goal fixed while varying the relevant spatial relation or object; Long contains the benchmark’s ten multi-stage tasks.
Hyperparameter
Meta-World
FR3 Duo
Correction source
Scripted expert
Human (SpaceMouse)
Online episodes per task
50
15
Action/latent horizon
16
16
Latent dimension per step
32
32
Latent bound
3.0
3.0
Inversion method
Per-step fixed point
Per-step fixed point
Appendix
Table 21: Online latent-space adaptation hyperparameters. Rho’s vision-language backbone and flow-matching action expert are frozen in both settings; only its lightweight internal latent policy is optimized.
Task
Model
Succ. ↑
TP ↑
BGVD ↓
CPL ↓
ECC ↓
JPL ↓
OPL ↓
SCC ↓
SC ↓
Cube handover
GR00T N1.7
0.64
0.93
0.043
1.35
0.16
6.81
4.14
0.373
0.477
π0.5
0.79
0.92
0.038
1.18
0.10
5.44
3.37
0.210
0.237
MolmoAct2
0.86
0.95
0.036
1.11
0.04
5.12
3.30
0.107
0.193
Rho
0.80
0.93
0.037
1.18
0.09
5.15
3.14
0.207
0.217
Lift pot
GR00T N1.7
0.67
0.81
0.058
2.57
0.13
16.83
9.70
0.120
0.603
π0.5
0.72
0.84
0.029
1.48
0.00
9.39
6.80
0.017
0.417
Appendix
Table 22 : Task-level RoboEval outcome and behavioral metrics after 40k task-adaptation steps. Each cell is pooled over three 100-episode evaluations for the corresponding task. Path-length measures sum the two arms. Arrows indicate the preferred direction.
Large-scale pretraining has made Vision-Language-Action (VLA) models promising foundations for generalist robot manipulation, yet adapting them to downstream tasks remains necessary. However, the common practice of full fine-tuning treats pretraining as initialization and can shift broad priors toward narrow training-distribution patterns. We propose PriorVLA, a novel framework that preserves pretrained priors and learns to leverage them for effective adaptation. PriorVLA keeps a frozen Prior Expert as a read-only prior source and trains an Adaptation Expert for downstream specialization. Expert Queries capture scene priors from the pretrained VLM and motor priors from the Prior Expert, integrating both into the Adaptation Expert to guide adaptation. Together, PriorVLA updates only 25% of the parameters updated by full fine-tuning. Across RoboTwin 2.0, LIBERO, and real-world tasks, PriorVLA achieves stronger overall performance than full fine-tuning and state-of-the-art VLA baselines, with the largest gains under out-of-distribution (OOD) and few-shot settings. PriorVLA improves over pi0.5 by 11 points on RoboTwin 2.0-Hard and achieves 99.1% average success on LIBERO. Across eight real-world tasks and two embodiments, PriorVLA reaches 81% in-distribution (ID) and 57% OOD success with standard data. With only 10 demonstrations per task, PriorVLA reaches 48% ID and 32% OOD success, surpassing pi0.5 by 24 and 22 points, respectively.
Xinyu Guo, Bin Xie, Wei Chai +4
Institute of Automation, Chinese Academy of Sciences · Dexmal · Nanjing University of Aeronautics and Astronautics +1
Vision-Language-Action (VLA) models have emerged as a promising paradigm for robotic manipulation by leveraging pre-trained vision-language representations. However, current VLA training methods suffer from two critical limitations: poor generalization to novel environments and low training efficiency requiring extensive demonstrations. We introduce Agentic-VLA, an agentic training framework that enables VLAs to efficiently adapt online through three key innovations: (1) Adaptive Reward Synthesis, which dynamically generates and adjusts reward functions based on the VLA's current capabilities and task complexity, decomposing complex tasks into learnable sub-goals for curriculum learning; (2) Language-Guided Exploration, where a critic model provides structured guidance for systematic exploration rather than random sampling; and (3) Experience Memory,which stores and retrieves task-relevant policy weights for warm-starting adaptation to similar tasks. We evaluate Agentic-VLA on the LIBERO benchmark, achieving substantial improvements: +12.3% on long-horizon tasks, +28.5% in 1-shot learning, and enabling cross-task transfer from 0% to 31.2% without task-specific demonstrations. Our framework also demonstrates 2.4x faster convergence compared to existing online adaptation methods. Beyond LIBERO, Agentic-VLA retains its advantage on the dual-arm RoboTwin 2.0 benchmark, including under its randomized Hard setting. These results establish Agentic-VLA as a significant step toward truly adaptive VLA systems capable of continuous learning in deployment.
Vision-Language-Action (VLA) models are emerging as a promising paradigm for robotic manipulation, enabling general-purpose policies trained from large corpora of demonstrations and action labels. However, adapting these models to new tasks still typically requires task-specific demonstrations, action annotations, and additional fine-tuning, making deployment costly and difficult to scale. We propose WIZARD, a weight-space meta-learning framework that sidesteps task-specific fine-tuning by generating task-specific LoRA parameters for a frozen VLA policy. Given only a language instruction and a short demonstration video, WIZARD predicts the corresponding adaptation weights in a single forward pass, without target-task action labels or test-time optimization. During meta-training, WIZARD learns to map task evidence directly to expert LoRA updates, capturing relationships between tasks in weight space. Experiments on LIBERO show that WIZARD improves performance by up to ~2x on unseen dataset collections and up to ~14x on unseen tasks. On a Franka Emika Panda, WIZARD consistently improves over a real-domain adapted baseline, showing that generated adapters provide task-level specialization beyond simulation.
Christian Bianchi, Siamak Yousefi, Alessio Sampieri +4