cs.CVSep 28, 2026

ActionUNet: Improving Robustness of VLA Models with Efficient Multi-scale Fine-tuning

Authors: Di Zhu, Ziheng Yan, Fang Wan

Organizations: University of Chinese Academy of Sciences

Abstract

Vision-Language-Action (VLA) models have shown great promise for robotic manipulation by mapping multi-modal semantics to physical actions. However, this mapping inherently struggles to align these coarse-grained semantics with fine-grained temporal execution. It leaves VLA models with limited generalization and insufficient robustness in cluttered environments. To overcome this issue, we propose ActionUNet, an efficient multi-scale fine-tuning framework that enhances pre-trained VLA models with minimal computational cost. ActionUNet first constructs a lightweight temporal U-Net within the temporal-aligned action feature space to fuse hierarchical structural priors, effectively bridging the scale gap between semantics and temporal executions. Recognizing that multi-scale modeling can disrupt microscopic temporal continuity and cause mechanical oscillations, ActionUNet then employs a conditional SIREN as a continuous action decoder. Equipped with explicit second-order smoothness constraints, this decoder guarantees temporal continuity and reduces high-frequency motion jitter. By smoothing temporal discontinuities from multi-scale fusion, this continuous formulation reduces mechanical execution failures while preserving the base VLA model's generalization and manipulation robustness. Extensive experiments on RoboTwin 2.0 and LIBERO-Plus benchmarks, together with real-world hard evaluations, demonstrate that ActionUNet significantly improves π0.5 success rates by absolute 9.8%, 6.1%, and 11.4%, respectively, while also generalizing to the regression-based OpenVLA-OFT backbone, highlighting its effectiveness and efficiency as a fine-tuning strategy. Code and implementation details are available at https://github.com/Di-Zhu123/ActionUNet.

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MaskVLA: Visual Masking Against Trajectory Overfitting of Vision-Language-Action Model

    Sep 20, 2026Yuxuan Jiang, Jiaying Huang, Ge Wang +9Diffusion-Based Vision-Language-ActionsOverfitting

  2. LARA: Latent Action Representation Alignment for Vision-Language-Action Models

    Jun 5, 2026Mengya Liu, Baoxiong Jia, Jiangyong Huang +2Latent Action ModelsLarge-Scale Robot Demonstration Datasets

  3. VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

    Jul 7, 2025Juyi Lin, Amir Taherin, Arash Akbari +11Diffusion-Based Vision-Language-ActionsRobotic Manipulation