Organizations: The Hong Kong University of Science and Technology (Guangzhou) · National Centre for Computer Animation, Bournemouth University · Shanghai Jiao Tong University · Eastern Institute of Technology, Ningbo
Vision-language-action (VLA) models have achieved strong performance in embodied manipulation, but still lack a clear mechanism to balance behavioral stability with task-semantic sensitivity. We identify two complementary failure modes. Under task-preserving changes, where task semantics remain unchanged but scene appearance varies (e.g., style, illumination, clutter, or paraphrasing), policies often exhibit unnecessary action drift. Conversely, under semantic-breaking changes, where key task semantics such as the target object or constraint are altered, policies frequently fail to produce sufficiently distinct behaviors and instead follow the original trajectory. To address this gap, we propose BAS-VLA, a task-semantic action calibration framework built on top of a frozen base VLA. BAS-VLA adopts a breaking-centered calibration core as the default path, and introduces a selective evidence-gated preserving auxiliary that activates only when nuisance variation is detected while task semantics remain consistent. On the OpenPI-pi0.5 / LIBERO-Object Milk-Swap benchmark, BAS-VLA maintains high success on clean (98.0%) and semantics-preserving conditions (97.5%), while reducing clean-criterion success to 0.0% under deliberate target-object swaps, demonstrating strong stale-task suppression and task-semantic separation. On validated style-preserving shifts, it improves success from 42% to 70% without degrading clean performance. These results highlight that reliable VLA behavior requires moving beyond appearance robustness toward explicit task-semantic action calibration.
Figures & tables
Carrier
Suite
Perturbation
Original success ↑
Perturbed success ↑
Δ (pts)
OpenVLA-OFT
LIBERO-Spatial
Perimeter clutter
98/100 (98.0%)
65/100 (65.0%)
−33.0
OpenVLA-OFT
LIBERO-Spatial
Margin noise
97/100 (97.0%)
51/100 (51.0%)
−46.0
OpenPI-pi0.5
LIBERO-Object
Illum. change
98/100 (98.0%)
89/100 (89.0%)
−9.0
OpenPI-pi0.5
LIBERO-Object
Target-object Swap
197/200 (98.5%)
68/200 (34.0%)
−64.5
Table 1 : Motivation evidence for reliability gaps before BAS-VLA. Original denotes native carrier performance on the unmodified task condition, and Perturbed denotes performance under the corresponding preserving or task-changed intervention. Carrier and suite names follow OpenPI-pi0.5 ( Black et al., 2025a ) , OpenVLA-OFT ( Kim et al., 2025 ) , and LIBERO ( Liu et al., 2023 ) . Here OpenVLA-OFT denotes an OpenVLA carrier built from the OFT fine-tuning recipe of Kim et al. (2025) ; in our study, we adopt that recipe and perform our own fine-tuning for the simulation and real-robot settings used in the paper.
Figure 2 : Task-preserving motivation on a single Black-Bowl-to-Plate instruction. The three panels show a clean success, a failure under a darkened illumination shift, and a failure under a semantics-preserving noise perturbation.
Figure 3 : Representative intervention families in simulation and real deployment. Top: simulation. Bottom: real deployment. Columns show no perturbation, illumination change, style shift, noise, clutter, and spatial-relation change.
Figure 4 : Dark-scene preserving example. The frozen carrier misreads the target under illumination change; the evidence-gated probe path activates and produces a corrected action.
Method
Clean success ↑
Matched-clean subset success ↑
Style-shift success ↑
Δstyle↑
Baseline
194/200 (97.0%)
108/200 (54.0%)
84/200 (42.0%)
—
BAS-VLA
195/200 (97.5%)
108/200 (54.0%)
140/200 (70.0%)
+28.0 pts
Table 3: Main task-preserving style-shift results on Bowl-on-Ramekin. Each method uses 200 episodes over four seeds; Δstyle is measured relative to the baseline style-shift row.
Task
Setting
Clean success ↑
Control success ↑
Break (clean) ↓
Ctrl.-Break gap ↑
Canonical harder target-object tasks
BBQ-Swap
BBQ sauce → ketchup
180/200 (90.0%)
185/200 (92.5%)
0/200 (0.0%)
92.5
Ketchup-Swap
ketchup → tomato sauce
176/200 (88.0%)
184/200 (92.0%)
0/200 (0.0%)
92.0
Milk-Swap
milk → orange juice
196/200 (98.0%)
195/200 (97.5%)
0/200 (0.0%)
97.5
Butter-Swap
butter → cream cheese
188/200 (94.0%)
182/200 (91.0%)
113/200 (56.5%)
34.5
OJ-Swap
orange juice → milk
197/200 (98.5%)
193/200 (96.5%)
0/200 (0.0%)
96.5
Table 4 : Main semantic-breaking results on matched target-object triplets. Each row uses 200 clean, 200 control, and 200 break episodes; lower is better in Break .
Figure 5 : Process-side view of the two main semantic-breaking tasks. Milk-Swap keeps clean/control short and unsaturated while break saturates; Tomato-Swap shows a weaker version of same pattern.
Case
Semantic change
n /method
Frozen
BAS-VLA
Δ (pts)
Successes
SR ↑
Successes
SR ↑
T1
Block target
80
19
23.8%
47
58.8%
+35.0
T2
Fruit target
80
8
10.0%
31
38.8%
+28.8
T3
Receptacle target
80
13
16.3%
52
65.0%
+48.8
T4
Spatial relation
80
11
13.8%
36
45.0%
+31.3
Aggregate
320
51
15.9%
166
51.9%
+35.9
Table 6 : Real-robot changed-task completion. Every case uses 80 trials per method (640 changed-condition rollouts in total). T1/T2 include the 17 changed-condition trials per method in Table 11 .
Harder target-object and visual-interference settings
Butter → cream cheese
BAS-VLA
200
53 (26.5%)
113 (56.5%)
34 (17.0%)
Table 7 : Changed-task outcome decomposition. New, Old, and Other/timeout are mutually exclusive. Easy swaps aggregate Milk/OJ/BBQ/Ketchup (four tasks × four seeds × 50 episodes); each harder setting uses four seeds × 50 episodes.
Protocol
Cases
n /method
Baseline
BAS-VLA
Δ (pts)
Successes
SR ↑
Successes
SR ↑
LIBERO-Plus
4×10
4,800
2,632
54.8%
3,741
77.9%
+23.1
LIBERO-PRO
10
1,200
443
36.9%
857
71.4%
+34.5
Table 8 : External-protocol success without benchmark-specific tuning. LIBERO-Plus uses four categories (light/background/noise/language) × ten cases × 120 rollouts per method; LIBERO-PRO uses ten task-perturbation cases × 120 rollouts per method.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Mode
T1 p50/p95 (ms)
T2 p50/p95 (ms)
Frozen carrier
345/620
368/655
Default breaking core
392/710
421/768
Preserving sidecar
924/1,048
924/1,048
Full composition
972/1,112
972/1,112
Appendix
Table 12: End-to-end chunk latency by deployment mode. Sidecar and Full totals include grounding, mask refinement, the probe image, and the additional frozen-policy query, measured over at least 200 chunks per task and three timing passes.
Semantic change
Evaluation volume
Metric
BAS-VLA result
Destination
2×4×50=400
New-task SR
224/400 (56.0%)
Spatial relation
2×4×50=400
New-task SR
168/400 (42.0%)
Subgoal order
2×4×50=400
New-task SR
128/400 (32.0%)
Appendix
Table 13: Changed-task completion on non-object semantic interventions.
Gate diagnostic
Evaluation volume
Result
Clean false activation
200
15/200 (7.5%)
Preserving-shift false negative
200
47/200 (23.5%)
Semantic-break false activation
200
16/200 (8.0%)
Oracle-routed Full task retention (upper bound)
200
183/200 (91.5%)
Appendix
Table 14: Full-mode gate diagnostics. Every row uses four seeds and 50 episodes per seed.
Mask condition
Evaluation volume
Preserving-SR change
IoU ≥0.7
200
−4/200 ( −2.0 pts)
IoU =0.5
200
−19/200 ( −9.5 pts)
Missed target or random mask
200
−37/200 ( −18.5 pts)
Appendix
Table 15: Preserving-SR change relative to the normal-mask reference. Each corruption level uses four seeds and 50 episodes per seed.
Result
Success rate
Wilson 95% CI
Baseline style
84/200 (42.0%)
[35.4,48.9]
BAS-VLA style
140/200 (70.0%)
[63.3,75.9]
Milk clean
196/200 (98.0%)
[95.0,99.2]
Milk control
195/200 (97.5%)
[94.3,98.9]
Milk old-task under break
0/200 (0.0%)
[0.0,1.9]
Butter old-task under break
113/200 (56.5%)
[49.6,63.2]
Appendix
Table 16 : Wilson 95% confidence intervals for the principal binary outcomes.
Carrier
Suite
Task slice
Perturbation
Best method
Baseline ↑
Calibrated ↑
Δ↑ (pts)
Limited-recovery appearance family
Pi0.5
LIBERO-Object
OJ-Swap × 200
illumination change
baseline-only
192/200 (96.0%)
—
—
Moderate support under stronger nuisance
OFT
LIBERO-Spatial
all10 × 20
clutter+noise
PARR
165/200 (82.5%)
174/200 (87.0%)
+4.5
OFT
LIBERO-Spatial
all10 × 20
hard noise band
PARR
171/200 (85.5%)
180/200 (90.0%)
+4.5
Style-sensitive families with strongest BAS-VLA gains
Figure 6 : Main simulation task scenes. The grid shows the ten task layouts used in the primary simulation campaigns.
Figure 7 : Dark-scene preserving comparison. Top: clean success. Middle: illumination-change failure under the frozen baseline. Bottom: the same darkened condition succeeds with BAS-VLA.
Figure 8 : Representative OJ-Swap under clutter+noise; BAS-VLA completes the intended task.
Figure 9 : Butter-Swap failure sequence. The rollout continues toward the stale clean-task object rather than committing to the swapped target.
Figure 10 : Process-side comparison among external strategies with a BAS-VLA reference row. Rows are strategies and columns are clean, control, and break.
Block
Family
Break-sep. rate ↑
Clean-ref. gap ↑
Initial-chunk sep. ↑
Executed-prefix sep. ↑
Preserving anchor
Semantic-preserving control
20.0
−0.019
−0.019
−0.017
Retained support families
Retained support
Spatial relation
100.0
+0.055
+0.055
+0.079
Retained support
Target object
60.0
+0.001
+0.001
+0.002
Retained support
Destination
60.0
+0.011
+0.011
+0.008
Appendix
Table 22 : Focused action-space diagnostics on the preserving anchor and evaluated support families under BAS-VLA Full. All rows use the matched OpenVLA-OFT ext6 protocol with five trials per family and a 24-step comparison prefix. Larger positive values indicate clearer semantic separation.
Figure 11 : Metric-wise offline instruction comparator losses. Shaded bands show the per-epoch range across trainable variants.
Figure 12 : Per-variant train/validation trajectories for the offline instruction comparator family. Shaded regions mark train–validation gaps.
Figure 13 : Normalized loss profiles for the offline instruction comparator family. Shaded bands show per-epoch ranges across trainable variants.
Figure 14 : Aggregate action-gap diagnostics for the offline instruction comparator family.
Figure 15 : Family-wise action-distance diagnostics for the offline instruction comparator family.