Vision-Language-Action (VLA) models have emerged as powerful foundations for robotic manipulation, but their reliance on fixed camera configurations during training makes them brittle to changes in camera count or pose during deployment. To overcome these limitations, we propose VersaCamVLA, a camera-configurable framework that decouples camera-set representation from action learning. VersaCamVLA learns a unified scene-token interface that maps an arbitrary, variable set of posed RGB views into fixed-size latent scene tokens. This is achieved via multi-signal target-view prediction and Wrist-Augmented Pose Sampling (WAPS), which leverages natural wrist-camera motion for free pose diversity. At deployment, a lightweight spatial encoder injects these compact scene tokens into a pretrained base VLA as a supplementary visual condition, requiring no explicit 3D sensing or novel-view rendering. Experiments on RoboTwin, LIBERO, and a real-robot platform demonstrate that VersaCamVLA consistently outperforms prior VLA methods and direct multi-view baselines, maintaining robust performance across varying camera counts and unseen camera poses.
Figures & tables
Figure 1: Left: Different manipulation tasks prefer different camera views, making a fixed view setup suboptimal. Right: VersaCamVLA augments a vanilla VLA with a camera-configurable scene-token interface that enables flexible view setup at deployment.
Figure 2: Overview of VersaCamVLA. Stage I: A scene encoder aggregates posed RGB views into fixed-size scene tokens, with a training-only decoder predicting RGB, semantic, and edge signals at target views. Stage II: The decoder is discarded, the scene encoder is frozen, and a jointly-trained spatial encoder compresses the tokens into a supplementary visual condition for a pretrained VLA.
Figure 3: Real-world Manipulation Processes. Representative rollouts of VersaCamVLA on three real-world tasks: PickCube, StackCube, and Insert Test Tube, showing key execution states from the initial scene to task completion.
Task
ACT [ 21 ]
DP [ 48 ]
DP3 [ 15 ]
OpenVLA-OFT [ 3 ]
π0.5 [ 1 ]
VersaCamVLA
Clean
DR
Clean
DR
Clean
DR
Clean
DR
Clean
DR
Clean
DR
Blocks Ranking RGB
0
0
0
0
2
0
0
0
38
11
67
12
Blocks Ranking Size
3
0
1
0
2
1
7
0
12
4
35
6
Hanging Mug
6
0
19
0
35
1
12
0
4
4
10
4
Move Stapler Pad
0
0
0
0
7
0
0
0
5
7
8
5
Open Laptop
72
1
53
0
77
1
82
0
90
70
96
62
Table 1: Simulation results on RoboTwin 2.0. We report success rates (%) on RoboTwin 2.0. All results are reproduced by us. Bold indicates the best performance in each column.
Method
LIBERO-Spatial
LIBERO-Goal
LIBERO-Object
LIBERO-Long
Average
Diffusion Policy [ 48 ]
78.3
92.5
68.3
50.5
72.4
OpenVLA-OFT [ 3 ]
95.2
94.2
95.2
93.2
94.5
π0.5∗ [ 1 ]
98.0
94.6
99.0
91.6
95.8
VersaCamVLA
98.2
95.0
97.8
92.4
95.9
Table 2: Simulation results on LIBERO. We report task success rates (%) on LIBERO. Results marked with ∗ are reproduced by us. Bold indicates the best performance in each column, and underlined indicates the second-best performance.
RoboTwin 2.0
Real Robot
Method
Seen Pose
Unseen Pose
Seen Pose
Unseen Pose
π0.5
48.88
32.88
41.67
13.33
VersaCamVLA
53.44
52.63
53.33
58.33
Table 3: Robustness to camera-pose variations. Average success rate (%) on RoboTwin 2.0 Clean and real-robot tasks. Bold indicates the best performance in each column.
RoboTwin 2.0
Real Robot
Method
3V
4V
5V
6V
3V
4V
5V
π0.5
35.88
48.88
✘
✘
31.67
41.67
✘
VersaCamVLA
52.56
53.44
53.13
55.63
✘
53.33
65.00
Table 4: Effect of input camera count. Average success rate (%) on RoboTwin 2.0 Clean and real-robot tasks. V denotes views. Bold indicates the best performance in each column.
Table 8
RGB
✓
✓
✓
✓
Edge
✗
✓
✗
✓
Semantic
✗
✗
✓
✓
Average SR
20.94
24.75
25.69
29.50
Table 9: Ablation of supervision signals for scene-token learning. Average success rate (%) on RoboTwin 2.0 DR. Bold indicates the best result.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Camera-pose perturbations on the RoboTwin 2.0 Stack Blocks Three task. The two rows show two of the training camera views. The first column shows the original, unperturbed camera poses, while the remaining four columns show randomly perturbed camera poses used for robustness evaluation.
Task
π0.5
VersaCamVLA
Seen
Unseen
Δ
Seen
Unseen
Δ
Blocks Ranking RGB
61
18
−43
61
56
−5
Blocks Ranking Size
44
6
−38
31
39
+8
Hanging Mug
14
11
−3
6
7
+1
Move Stapler Pad
8
4
−4
9
9
0
Open Laptop
90
82
−8
96
92
−4
Appendix
Table 10: Per-task camera-pose robustness on RoboTwin 2.0. Success rate (%) at seen and unseen poses under matched four-view inputs. Δ denotes Unseen minus Seen in percentage points. Bold indicates the higher value between methods for each task and metric, including ties.
Method (Setup)
Spatial
Goal
Object
Long
Avg
π0 (Seen)
94.8
87.0
97.0
83.6
90.6
π0.5 (Seen)
96.8
88.2
96.6
82.4
91.0
VersaCamVLA (Seen)
98.2
95.0
97.8
91.8
95.7
VersaCamVLA (Unseen)
98.2
96.8
95.8
89.8
95.15
Appendix
Table 11: Camera-pose robustness on LIBERO. Success rate (%) within each official suite under four-view inputs. Seen denotes trained camera poses and Unseen denotes perturbed poses. Bold indicates the highest value in each column, including ties.
Method
Seen Pose
Unseen Pose: ∣Δϕ∣
[0∘,30∘]
[30∘,60∘]
[60∘,90∘]
π0.5
48.88
43.38
34.13
30.06
VersaCamVLA
53.44
53.25
52.13
53.38
Appendix
Table 12: Generalization across unseen camera-pose shift magnitudes on RoboTwin 2.0. Average success rate (%) at the seen pose and under perturbations grouped by absolute azimuth offset. Bold indicates the best result in each column.
Extrinsic-noise level
σt (mm)
σr (degrees)
Average SR (%)
Clean
0
0
52.56
Moderate
20
2
52.13
Strong
50
5
50.38
Appendix
Table 13: Robustness to camera extrinsic-calibration noise on RoboTwin 2.0. Average success rate (%) under Gaussian perturbations fixed within each episode. Translation and rotation noise levels are specified by their per-axis standard deviations.
Task
3V
4V
5V
6V
Blocks Ranking RGB
67
61
61
67
Blocks Ranking Size
35
31
39
35
Hanging Mug
10
6
9
3
Move Stapler Pad
8
9
13
4
Open Laptop
96
96
94
96
Open Microwave
51
60
57
67
Appendix
Table 14: Per-task camera-count variation on RoboTwin 2.0 Clean . Success rate (%) for a single VersaCamVLA model evaluated with 3, 4, 5, and 6 views (V). Bold indicates the highest value in each row, including ties.
Task
π0.5
VersaCamVLA
Seen
Unseen
Δ
Seen
Unseen
Δ
PickCube
70
35
−35
85
80
−5
StackCube
55
5
−50
55
60
+5
Insert Test Tube
0
0
0
20
35
+15
Average
41.67
13.33
−28.33
53.33
58.33
+5.00
Appendix
Table 15: Per-task camera-pose robustness on the real robot. Success rate (%) under matched four-view inputs. Δ denotes Unseen minus Seen in percentage points. Bold indicates the higher value between methods for each task and metric, including ties.
Task
π0.5 (3V)
π0.5 (4V)
VersaCamVLA
4V
5V
PickCube
50
70
85
85
StackCube
40
55
55
65
Insert Test Tube
5
0
20
45
Average
31.67
41.67
53.33
65.00
Appendix
Table 16: Per-task camera-count variation on the real robot. Success rate (%) for π0.5 trained separately at 3 and 4 views and a single VersaCamVLA model evaluated at 4 and 5 views (V). Bold indicates the highest value in each row, including ties.
Task
48 tokens
192 tokens
432 tokens
Blocks Ranking RGB
55
67
55
Blocks Ranking Size
40
35
32
Hanging Mug
8
3
16
Move Stapler Pad
17
4
13
Open Laptop
89
96
92
Open Microwave
20
67
61
Appendix
Table 17: Per-task ablation of compressed scene-token count. Success rate (%) on RoboTwin 2.0 Clean . Token counts refer to the compressed scene tokens passed to the policy. Bold indicates the highest value in each row, including ties.
Task
Pooling
Conv
Blocks Ranking RGB
61
67
Blocks Ranking Size
26
35
Hanging Mug
11
3
Move Stapler Pad
9
4
Open Laptop
91
96
Open Microwave
52
67
Appendix
Table 18: Per-task ablation of spatial encoder architecture. Success rate (%) on RoboTwin 2.0 Clean with pooling or convolutional compression. Bold indicates the highest value in each row, including ties.
Task
RGB only
Multi-signal
Blocks Ranking RGB
13
14
Blocks Ranking Size
6
9
Hanging Mug
6
11
Move Stapler Pad
3
7
Open Laptop
67
63
Open Microwave
12
54
Appendix
Table 19: Per-task ablation of scene-token supervision. Success rate (%) on RoboTwin 2.0 DR with RGB-only or full multi-signal supervision. Bold indicates the highest value in each row, including ties.
Figure 5: Inference latency with varying camera counts. Measurements use a single NVIDIA 4090 at batch size 1. The red dashed line marks the three-view π0.5 reference. VersaCamVLA incurs a smaller latency increase as additional input views are incorporated.
Figure 6: Multi-signal supervision targets on the real-robot Insert Test Tube task. Rows (top to bottom) show the raw RGB observation, the semantic mask produced by a SAM3 segmenter fine-tuned on our real-robot data, and the Canny edge map. Columns (left to right) span four key task states, from the right arm grasping the test tube to the left arm completing insertion into the rack.