VLA-ZO: Fast Zeroth-Order Adaptation for Vision-Language-Action Models
Organizations: UNIST
Abstract
Adapting vision-language-action (VLA) models to deployment-time distribution shifts is important for reliable robotic operation, but conventional first-order adaptation can exceed the memory budget of inference-oriented deployment platforms. Zeroth-order (ZO) optimization offers a forward-only alternative with inference-level memory, but accurate gradient estimation requires many perturbation queries, making naive ZO prohibitively slow for large VLA models. We present VLA-ZO, a framework for fast ZO adaptation that exploits the structure of VLA computation. By confining adaptation to the action side, VLA-ZO keeps the expensive vision-language prefix frozen and reuses its conditioning states across perturbation queries and optimizer steps, while schedule-aware prefetching hides state-transfer overhead. On LIBERO camera-viewpoint shifts, VLA-ZO reduces end-to-end adaptation time by 25.59 at and 32.54 at relative to baseline ZO, while improving average task success from 48.27% without adaptation to 58.17% and 63.58%, respectively. These results show that making ZO faster can make larger query budgets practical, providing a promising path toward resource-efficient VLA adaptation on deployment platforms.
Figures & tables
| Novel viewpoint (success rate, %) | Adaptation | |||||||
| Method | Small | Medium | Large | Average | Params (M) | Time (s) | Speedup | |
| Zero-shot | – | 86.19 | 51.44 | 7.19 | 48.27 | – | – | – |
| FLA † ( Li et al., 2026 ) | 16 | 88.56 | 53.31 | 27.81 | 56.56 | 4.71 | 42,782 | – |
| DART † ( Kang et al., 2026 ) | 16 | 89.50 | 58.88 | 14.25 | 54.21 | 3,353.43 | 43,794 | – |
| Baseline ZO | 16 | 86.44 | 60.69 | 25.75 | 57.63 | 1.25 | 38,487 | 1.00 |
| + Within-step prefix reuse | 16 | 88.00 | 60.31 | 25.00 | 57.77 | 1.25 | 2,474 | 15.56 |
| View+Noise | View+Light | Adaptation | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | Spa. | Obj. | Goal | 10 | Avg. | Spa. | Obj. | Goal | 10 | Avg. | Params (M) | Time (s) | |
| Zero-shot | – | 50.5 | 84.0 | 42.0 | 20.3 | 49.2 | 39.8 | 86.8 | 53.8 | 15.8 | 49.0 | – | – |
| FLA † ( Li et al., 2026 ) | 16 | 65.3 | 70.8 | 47.0 | 29.0 | 53.0 | 43.5 | 73.5 | 57.5 | 35.3 | 52.4 | 4.71 | 43,081 |
| DART † ( Kang et al., 2026 ) | 16 | 57.5 | 90.8 | 45.5 | 27.8 | 55.4 | 52.3 | 92.3 | 50.8 | 31.8 | 56.8 | 3,353.43 | 43,965 |
| VLA-ZO | 16 | 46.0 | 86.8 | 43.5 | 28.0 | 51.1 | 60.0 | 90.3 | 56.3 | 27.0 | 58.4 | 1.25 | 1,500 |
| VLA-ZO | 64 | 49.0 | 91.5 | 49.3 | 33.3 | 55.8 | 68.5 | 89.0 | 59.0 | 26.8 | 60.8 | 1.25 | 4,726 |
| Novel viewpoint (success rate, %) | Adaptation | ||||||
|---|---|---|---|---|---|---|---|
| Method | Small | Medium | Large | Average | Params (M) | Time (s) | |
| Zero-shot | – | 80.19 | 50.88 | 17.38 | 49.48 | – | – |
| FLA † ( Li et al., 2026 ) | 4 | 81.06 | 52.63 | 19.31 | 51.00 | 4.71 | 18,283 |
| DART † ( Kang et al., 2026 ) | 4 | 81.13 | 53.63 | 20.50 | 51.75 | 7,709.17 | 28,482 |
| VLA-ZO | 4 | 79.06 | 57.06 | 21.88 | 52.67 | 0.79 | 109 |
| VLA-ZO | 16 | 79.63 | 55.69 | 21.25 | 52.19 | 0.79 | 170 |
| View+Noise | View+Light | Adaptation | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | Spa. | Obj. | Goal | 10 | Avg. | Spa. | Obj. | Goal | 10 | Avg. | Params (M) | Time (s) | |
| Zero-shot | – | 17.0 | 41.0 | 17.3 | 20.3 | 23.9 | 26.3 | 86.8 | 26.0 | 24.5 | 40.9 | – | – |
| FLA † ( Li et al., 2026 ) | 4 | 21.8 | 47.3 | 14.0 | 21.8 | 26.2 | 26.5 | 87.3 | 25.8 | 28.3 | 41.9 | 4.71 | 18,283 |
| DART † ( Kang et al., 2026 ) | 4 | 29.8 | 43.3 | 18.0 | 24.8 | 28.9 | 30.8 | 86.0 | 24.8 | 26.8 | 42.1 | 7,709.17 | 28,265 |
| VLA-ZO | 16 | 29.3 | 45.8 | 17.5 | 24.3 | 29.2 | 31.5 | 88.0 | 28.8 | 26.8 | 43.8 | 0.79 | 170 |
| VLA-ZO | 64 | 31.8 | 65.5 | 13.3 | 23.5 | 33.5 | 34.8 | 89.3 | 32.8 | 28.3 | 46.3 | 0.79 | 403 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Scene | Demonstrated task | Length |
|---|---|---|
| Living room | Put both the alphabet soup and the tomato sauce in the basket. | 336 |
| Kitchen | Turn on the stove and put the moka pot on it. | 272 |
| Study | Pick up the book and place it in the back compartment of the caddy. | 234 |
| Floor | Pick up the alphabet soup and place it in the basket. | 145 |
| Tabletop | Pick up the black bowl between the plate and the ramekin and place it on the plate. | 98 |
| Severity | (m) | (m) | (m) |
|---|---|---|---|
| Small | 0.0 | 0.3 | |
| Medium | 0.7 | ||
| Large | 1.0 |
| Setting | OpenVLA-OFT | |
| Predicted action chunk length | 50 | 8 |
| Actions executed before replanning | 10 | 8 |
| Maximum control steps: Spatial | 220 | 220 |
| Maximum control steps: Object | 280 | 280 |
| Maximum control steps: Goal | 300 | 300 |
| Maximum control steps: LIBERO-10 | 520 | 520 |
| Method | Trainable placement | Learning rate | Updates | |
|---|---|---|---|---|
| Action-side ZO | Action-expert query/value LoRA | 16 | 1,000 | |
| Action-side ZO | Action-expert query/value LoRA | 64 | 1,000 | |
| FLA † | Vision encoder FFN LoRA | 16 | 1,000 | |
| DART † | Full model, source and target | 16 |
| Method | Trainable placement | Learning rate | Updates | |
|---|---|---|---|---|
| Action-side ZO | Action-head LoRA | 4 | 1,000 | |
| Action-side ZO | Action-head LoRA | 16 | 1,000 | |
| Action-side ZO | Action-head LoRA | 64 | 1,000 | |
| FLA † | Vision encoder FFN LoRA | 4 | 1,000 | |
| DART † | Full model, source and target | 4 |