Vision-Language-Action (VLA) models have shown great promise for robotic manipulation by mapping multi-modal semantics to physical actions. However, this mapping inherently struggles to align these coarse-grained semantics with fine-grained temporal execution. It leaves VLA models with limited generalization and insufficient robustness in cluttered environments. To overcome this issue, we propose ActionUNet, an efficient multi-scale fine-tuning framework that enhances pre-trained VLA models with minimal computational cost. ActionUNet first constructs a lightweight temporal U-Net within the temporal-aligned action feature space to fuse hierarchical structural priors, effectively bridging the scale gap between semantics and temporal executions. Recognizing that multi-scale modeling can disrupt microscopic temporal continuity and cause mechanical oscillations, ActionUNet then employs a conditional SIREN as a continuous action decoder. Equipped with explicit second-order smoothness constraints, this decoder guarantees temporal continuity and reduces high-frequency motion jitter. By smoothing temporal discontinuities from multi-scale fusion, this continuous formulation reduces mechanical execution failures while preserving the base VLA model's generalization and manipulation robustness. Extensive experiments on RoboTwin 2.0 and LIBERO-Plus benchmarks, together with real-world hard evaluations, demonstrate that ActionUNet significantly improves π0.5 success rates by absolute 9.8%, 6.1%, and 11.4%, respectively, while also generalizing to the regression-based OpenVLA-OFT backbone, highlighting its effectiveness and efficiency as a fine-tuning strategy. Code and implementation details are available at https://github.com/Di-Zhu123/ActionUNet.
Figures & tables
Figure 1 : Illustration of multi-scale temporal dynamics. Per-step end-effector displacements vary significantly across different tasks (left) and over time within a single task “Stack blocks two” (right).
Figure 2 : Overview of ActionUNet. Given input observations and instructions, the VLA backbone first produces temporal-aligned action features. ActionUNet refines these features in the action feature space through multi-scale temporal mapping, and then decodes the fused features with a conditional SIREN into smooth continuous action trajectories.
Figure 3 : Overview of the conditional SIREN decoder. Local SIREN fields conditioned on multi-scale action features are queried over continuous time and fused with Matérn-kernel weights to generate temporally smooth actions.
Task
RDT [ 7 ]
π0 [ 17 ]
ACT [ 39 ]
DP3 [ 40 ]
π0.5 [ 4 ]
π0.5 + ActionUNet
Easy
Hard
Easy
Hard
Easy
Hard
Easy
Hard
Easy
Hard
Easy
Hard
Single-scale Tasks
Pick Dual Bottles
42
13
57
12
31
0
60
1
65
34
71
39
Handover Mic
90
31
98
13
85
0
100
3
100
55
100
76
Handover Block
45
14
45
8
42
0
70
0
33
13
47
17
Dual-scale Tasks
Table 1 : Task-wise success rates on RoboTwin 2.0 grouped by scale composition.
Method
SPATIAL
OBJECT
GOAL
LONG
Avg.
Diffusion Policy [ 41 ]
78.3
92.5
68.3
50.5
72.4
WorldVLA [ 42 ]
87.6
96.2
83.4
60.0
81.8
SmolVLA [ 43 ]
93.0
94.0
91.0
77.0
88.8
π0 [ 17 ]
96.8
98.8
95.8
85.2
94.2
π0 -FAST [ 16 ]
96.4
96.8
88.6
60.2
85.5
UniVLA [ 44 ]
96.5
96.8
95.6
92.0
95.2
Table 2 : LIBERO benchmark results. We report the average success rate (%) across four LIBERO task suites.
Method
Camera
Robot
Lang.
Light
Back.
Noise
Layout
Avg.
OpenVLA [ 2 ]
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
WorldVLA [ 42 ]
0.1
27.9
41.6
43.7
17.1
10.9
38.0
25.0
UniVLA [ 44 ]
1.8
46.2
69.6
69.0
81.0
21.2
31.9
43.9
π0 [ 17 ]
13.8
6.0
58.8
85.0
81.4
79.0
68.9
53.6
π0 -FAST [ 16 ]
65.1
21.6
61.0
73.2
73.2
74.4
68.8
61.6
OpenVLA-OFT [ 8 ]
36.0
35.9
67.0
82.1
93.1
47.1
82.5
60.9
Table 3 : Performance under environmental perturbations on LIBERO-Plus. All models are trained on LIBERO and evaluated on LIBERO-Plus. We report success rates (%) under 7 perturbation types.
Task
π0.5
π0.5 + ActionUNet
Easy
Hard
Easy
Hard
Place Object Basket
60
42
78
52
Stack Bowls Two
88
62
96
74
Stack Blocks Two
80
68
92
80
Average
76.0
57.3
88.7
68.7
Table 4 : Comparison of our method against the baseline under easy and hard conditions in real-world experiments.
Task
π0.5
+ U-Net
+ SIREN
+ ActionUNet w/o Matérn
+ ActionUNet
Easy
Hard
Easy
Hard
Easy
Hard
Easy
Hard
Easy
Hard
Move Can Pot
66
52
64
56
73
53
73
68
77
69
Stack Blocks Two
68
24
75
27
74
24
71
34
74
35
Beat Block Hammer
77
28
65
33
87
30
89
39
89
43
Pick Dual Bottles
65
34
66
33
68
26
67
40
71
39
Average
69.0
34.5
67.5
37.3
75.5
33.3
75.0
45.3
77.8
46.5
Table 5 : Ablation results on four RoboTwin 2.0 tasks. Each task reports success rates under easy and hard settings. All non-baseline variants are built on π0.5 .
Figure 4 : Scale-conditioned temporal Grad-CAM responses on RoboTwin 2.0. Fine-, mid-, and coarse-resolution temporal features show stronger responses to the corresponding action scales.
Figure 10
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Metric
π0.5
π0.5 + ActionUNet
Difference
Mean latency / policy call (ms)
60.947±0.251
64.886±0.376
+3.939(+6.46%)
Appendix
Table 6 : Compiled policy-call latency on a single RTX 4090 GPU. We report mean latency with 95% confidence intervals computed over repeated policy-call measurements.
Setting
Average Taken Steps
Latency (s)
π0.5
π0.5 + ActionUNet
π0.5
π0.5 + ActionUNet
Difference
Easy
238.80
238.69
14.554±0.060
15.488±0.090
+0.934±0.108
Hard
276.43
264.29
16.848±0.069
17.149±0.099
+0.301±0.121
Appendix
Table 7 : Rollout-level policy computation on the 12 RoboTwin2.0 tasks. Average taken steps are task-level unweighted means over successful rollouts. Latency values are estimated by multiplying the average taken steps by the compiled policy-call latency in Table 6 ; their 95% confidence intervals are propagated from the task-independent policy-call latency measurements. The paired difference is computed as π0.5 + ActionUNet minus π0.5 latency, and its 95% confidence interval is propagated from the two latency confidence intervals.
Task
Overall
Overall Scale
Stage Sequence
Composition
#Regimes
Adjust Bottle
0.0442
High
H–M
M+H
2
Beat Block Hammer
0.0372
High
H–L
L+H
2
Blocks Ranking RGB
0.0363
Medium
M–L–H–L–H–L–L
L+M+H
3
Blocks Ranking Size
0.0359
Medium
H–L–H–L–H–L–L
L+H
2
Click Alarmclock
0.0412
High
H–M
M+H
2
Click Bell
0.0373
High
H–L
L+H
2
Appendix
Table 8: Task-level and stage-level action scale analysis on 50 tasks. The overall scale is computed from the average end-effector displacement over the entire task. Stage sequence is computed from non-idle stage-level displacement values. All displacement values are reported in m/step.
Figure 7 : Visualization of Place Object Basket task execution with π0.5 +ActionUNet. The top row shows the keyframes of the task execution in the easy setting, while the bottom row presents the keyframes in the hard setting.
Figure 8 : Visualization of Stack Bowls Two task execution with π0.5 +ActionUNet. The top row shows the keyframes of the task execution in the easy setting, while the bottom row presents the keyframes in the hard setting.
Figure 9 : Visualization of Stack Blocks Two task execution with π0.5 +ActionUNet. The top row shows the keyframes of the task execution in the easy setting, while the bottom row presents the keyframes in the hard setting.
Task
π0.5
+ Layer-wise Pyramid Condition
+ ActionUNet
Easy
Hard
Easy
Hard
Easy
Hard
Move Can Pot
66
52
71
64
77
69
Stack Blocks Two
68
24
74
25
74
35
Beat Block Hammer
77
28
86
34
89
43
Average
70.3
34.7
77.0
41.0
80.0
49.0
Gain over π0.5
–
–
+6.7
+6.3
+9.7
+14.3
Appendix
Table 9 : Architecture ablation on where multi-scale features enter the continuous decoder. Layer-wise Pyramid Condition. directly injects U-Net decoder features at different temporal resolutions into different SIREN layers. All non-baseline variants are built on π0.5 .
Task
π0.5
+ RTC
+ ActionUNet
Easy
Hard
Easy
Hard
Easy
Hard
Move Can Pot
66
52
73
64
77
69
Stack Blocks Two
68
24
72
37
74
35
Beat Block Hammer
77
28
71
37
89
43
Pick Dual Bottles
65
34
71
33
71
39
Average
69.0
34.5
71.8
42.8
77.8
46.5
Appendix
Table 10 : Comparison with action smoothing baselines on four RoboTwin2.0 tasks. RTC [ 45 ] performs run-time replanning and soft-mask continuation without changing the learned action representation. All non-baseline variants are built on π0.5 .
Figure 10 : Ablation study on the Matérn length-scale parameter ρ . Changing ρ within the tested range only introduces minor performance variations, showing that ActionUNet is robust to the choice of the Matérn aggregation bandwidth.
Task
π0.5
+ r=1 Single-scale
+ActionUNet
Easy
Hard
Easy
Hard
Easy
Hard
Move Can Pot
66
52
76
61
77
69
Stack Blocks Two
68
24
73
28
74
35
Beat Block Hammer
77
28
84
33
89
43
Pick Dual Bottles
65
34
65
32
71
39
Average
69.0
34.5
74.5
38.5
77.8
46.5
Appendix
Table 11 : Parameter-matched single-scale control on four RoboTwin 2.0 tasks. The r=1 variant has the same parameter count as ActionUNet but removes temporal downsampling and cross-scale fusion. All variants are built on π0.5 .
Task
π0.5
π0.5 +ActionUNet
Easy
Hard
Easy
Hard
Single-scale Tasks
Pick Dual Bottles
64.8± 1.0
32.4± 0.9
70.8± 0.6
38.8± 0.4
Handover Mic
99.6± 0.2
56.0± 0.7
99.8± 0.2
78.8± 1.2
Handover Block
34.4± 0.5
13.0± 0.4
48.6± 0.9
16.8± 0.4
Dual-scale Tasks
Appendix
Table 12 : Repeated evaluation results on RoboTwin 2.0. We report mean success rate (%) ± standard error over five repeated evaluations.
Method
Spatial
Object
Goal
Long
Avg.
π0.5
95.5± 0.45
98.4± 0.28
97.4± 0.85
91.1± 0.95
95.6± 0.35
π0.5 +ActionUNet
98.2 ± 0.33
99.6 ± 0.33
98.8 ± 0.68
94.0 ± 0.69
97.7 ± 0.27
Appendix
Table 13 : Repeated evaluations on LIBERO. We report mean success rate (%) ± standard error over five repeated evaluations.
Method
Camera
Robot
Lang.
Light
Back.
Noise
Layout
Avg.
π0.5
48.4± 0.65
48.0± 1.75
67.4± 0.65
93.0± 0.20
87.1± 0.30
51.1± 1.20
81.8± 0.50
66.0± 0.40
π0.5 +ActionUNet
56.4 ± 0.49
55.2 ± 1.55
74.6 ± 0.51
94.8 ± 0.06
89.7 ± 0.20
60.2 ± 0.97
85.5 ± 0.56
72.0 ± 0.32
Appendix
Table 14 : Repeated evaluations on LIBERO-Plus. We report mean success rate (%) ± standard error over five repeated evaluations.
Task
Easy
Hard
p^0
p^1
STEP
Result
p^0
p^1
STEP
Result
Pick Dual Bottles
.645
.720
153
Sig.
.350
.390
–
FTD
Handover Mic
.995
1.000
–
FTD
.540
.780
46
Sig.
Handover Block
.335
.475
90
Sig.
.120
.185
164
Sig.
Beat Block Hammer
.750
.895
30
Sig.
.280
.435
84
Sig.
Move Can Pot
.660
.770
29
Sig.
.520
.715
20
Sig.
Appendix
Table 15 : Sequential A/B hypothesis testing on RoboTwin 2.0. “Sig.” denotes a statistically significant improvement of π0.5 +ActionUNet over π0.5 , and “FTD” denotes FailToDecide within the maximum budget of 200 trials per policy.
Task
π0.5
π0.5 +ActionUNet
Easy
Hard
Easy
Hard
Dual-scale Tasks
Open Laptop
88
59
96
75
Place Dual Shoes
45
12
56
21
Click Alarm
84
17
91
21
Three-scale Tasks
Appendix
Table 16 : Results on 5 additional RoboTwin 2.0 tasks beyond the original 12-task evaluation subset.
Method
Click Alarm
Move Can Pot
Place Can Basket
HiFlow [ 25 ]
69
42
39
π0.5 +ActionUNet
91
77
57
Appendix
Table 17 : Comparison with HiFlow on three RoboTwin 2.0 tasks under the clean setting. HiFlow results are taken from the corresponding paper.
Method
Spatial
Object
Goal
Long
Avg.
Multi-scale Methods
MINT-30M [ 30 ]
98.6
99.2
97.4
93.2
97.1
FASTer [ 47 ]
98.0
99.4
98.6
95.4
97.9
Hierarchical Methods
Coarse-to-Control [ 48 ]
98.8
100.0
97.8
95.0
97.9
ECHO [ 49 ]
98.3
98.8
98.6
93.5
97.3
Appendix
Table 18 : Comparison with recent multi-scale and hierarchical methods on LIBERO. Results for external methods are taken from the corresponding papers.
Method
Average Success Rate
MINT-30M [ 30 ]
69.5
ECHO [ 49 ]
56.5
OpenVLA-OFT+ActionUNet
64.6
π0.5 +ActionUNet
71.9
Appendix
Table 19 : Comparison with recent multi-scale and hierarchical methods on LIBERO-Plus. We report average success rate (%).
Task
π0.5
+ DCT
+ Laplacian
+ ActionUNet
Easy
Hard
Easy
Hard
Easy
Hard
Easy
Hard
Move Can Pot
66
52
70
52
64
56
77
69
Stack Blocks Two
68
24
71
18
66
16
74
35
Beat Block Hammer
77
28
78
29
73
28
89
43
Pick Dual Bottles
65
34
69
30
69
28
71
39
Average
69.0
34.5
72.0
33.0
68.0
32.0
77.8
46.5
Appendix
Table 20 : Comparison with action-space multi-scale decomposition baselines on four RoboTwin 2.0 tasks. DCT decomposes the target action chunk into low-, mid-, and high-frequency components, while the Laplacian variant constructs a multi-resolution residual action pyramid. All variants are built on π0.5 .
Setting
Method
Success Rate (%) ↑
Drift (mm) ↓
Jerk RMS (m/s 3 ) ↓
Easy
π0.5
69.0
32.53
32.90
+ U-Net
67.5
29.07
34.49
+ SIREN
75.5
28.85
29.43
+ ActionUNet
77.8
27.60
32.70
Hard
π0.5
34.5
81.64
48.31
+ U-Net
37.3
76.05
49.34
Appendix
Table 21 : Quantitative trajectory analysis on four RoboTwin 2.0 tasks. Lower Manipulation Drift and Jerk RMS indicate better spatial accuracy and smoother trajectories, respectively.
Method
Camera
Robot
Lang.
Light
Back.
Noise
Layout
Avg.
π0.5
48.4± 0.65
48.0± 1.75
67.4± 0.65
93.0± 0.20
87.1± 0.30
51.1± 1.20
81.8± 0.50
66.0± 0.40
+ U-Net
50.8± 0.18
52.4± 0.51
71.3± 0.34
93.8± 0.86
88.7± 0.86
52.9± 0.17
83.1± 0.31
68.4± 0.18
+ SIREN
47.4± 0.23
53.1± 0.30
67.4± 0.25
92.7± 0.55
86.8± 0.42
48.7± 0.63
81.4± 0.53
66.1± 0.17
+ ActionUNet
56.4 ± 0.49
55.2 ± 1.55
74.6 ± 0.51
94.8 ± 0.06
89.7 ± 0.20
60.2 ± 0.97
85.5 ± 0.56
72.0 ± 0.32
Appendix
Table 22 : Module-wise ablation on LIBERO-Plus. All variants are trained on LIBERO and directly evaluated under seven controlled perturbations. We report mean success rate (%) ± standard error over three repeated evaluations.
Task
Method
Success (%) ↑
Progress ↑
Jerk RMS ↓
GT Jerk RMS
Insert Tubes
π0.5
4.7± 3.1
22.0± 3.2
127.81
50.62
+ ActionUNet
14.0 ± 2.0
30.8 ± 0.4
122.99
50.62
Build Tower
π0.5
26.0± 2.0
37.3± 0.9
119.75
49.54
+ ActionUNet
34.7 ± 2.3
42.5 ± 1.5
109.53
49.54
Appendix
Table 23 : Evaluation on precision-demanding RoboDojo tasks. We report mean success rate and Progress Score ± standard error over three repeated evaluations. Lower Jerk RMS indicates smoother trajectories.
FNii-Shenzhen, The Chinese University of Hong Kong, Shenzhen, China · School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China · Ising AI +2