Dual-system Vision-Language-Action (VLA) models improve real-time robotic control by pairing a slow, reasoning-capable generalist with a fast specialist action expert. However, existing methods invoke the generalist at a fixed frequency, ignoring the fact that decision-making complexity varies throughout a rollout. This static strategy wastes computation in easy phases and can delay renewed reasoning when the scene changes unexpectedly. We propose TUD (Triggering generalist reasoning via predictive Uncertainty for Dual-system VLA), an adaptive inference framework that selectively skips unnecessary generalist calls. TUD measures the cross-step dispersion of action re-predictions at the upcoming chunk slot under the cached generalist context, as a predictive uncertainty signal. This signal captures how much the future action plan shifts as new observations arrive and is computed from forwards the architecture already runs, requiring neither manual phase labels nor an auxiliary uncertainty model. On VLA-Arena, it achieves a higher success rate at matched call budgets than alternative uncertainty baselines while maintaining low wall-clock overhead, and more consistently separates successful from failed rollouts. Also, TUD finds a more favorable cost-success trade-off than non-adaptive baselines, tracing an entire operating curve as a single threshold is varied, and substantially reduces VLM calls at matched success rate. The same trade-off appears in our real-robot experiments, where TUD cuts generalist calls by 75% relative to the strongest fixed-interval baseline while achieving an even higher success rate. Our results suggest that predictive uncertainty provides a practical criterion for adaptive reasoning in efficient VLA control.
Figures & tables
Figure 1 : Two properties of existing dual-system VLAs that motivate adaptive generalist calling. (a) Per-task success rate of OpenHelix on CALVIN at varying generalist call interval Δ ( N on the x-axis); the best Δ⋆ differs by task category. (b) Long-horizon rollout performance of Robodual on CALVIN at varying Δ , averaged over 100 rollouts per Δ .
Spec.
Gen.
Dual
Throughput (Hz) ↑
36.98
2.70
14.79
Latency (s) ↓
0.216
0.370
-
Table 1 : Per-call cost on Robodual.
Figure 2 : Overview of TUD. The generalist call refreshes the cached context hk at decision points, the specialist forward re-predicts an action chunk against hk at every step, and a sliding buffer of these re-predictions yields the predictive action-plan dispersion ut . The invocation rule thresholds the running mean of ut to skip or invoke the generalist.
Baselines
Ours (TUD)
Task
OpenVLA
OpenVLA-OFT
π0
π0 -FAST
UniVLA
SmolVLA
OneTwoVLA
RoboDual
p70
Best
Safety (SR / CC)
StaticObstacles
0.60 / 0.0
1.00 / 0.0
0.98 / 0.0
1.00 / 0.0
0.84 / 0.0
0.14 / 0.0
0.94 / 0.0
0.68 / 0.0
0.97 / 0.0
0.99 / 0.0
CautiousGrasp
0.80 / 6.6
0.60 / 3.3
0.84 / 3.5
0.64 / 3.3
0.80 / 3.3
0.52 / 2.5
0.72 / 2.5
0.42 / 2.2
0.93 / 3.2
0.93 / 3.2
HazardAvoidance
0.20 / 17.2
0.36 / 9.4
0.74 / 6.4
0.10 / 10.4
0.70 / 5.3
0.16 / 10.4
0.56 / 6.7
0.68 / 5.3
0.85 / 5.5
0.85 / 5.5
StatePreservation
1.00 / 0.0
1.00 / 0.0
0.98 / 0.0
0.60 / 0.0
0.90 / 0.0
0.50 / 0.0
0.84 / 0.0
0.82 / 0.0
0.95 / 0.0
0.97 / 0.0
Table 2 : Per-task success rate (SR) and cumulative cost (CC) on VLA-Arena Level 0. p70 reports TUD at our representative operating point ( p=0.70 ), and Best reports TUD at whichever p gives the highest result for each individual task.
Figure 3 : Cost–success trade-off on VLA-Arena. Left : Safety. Right : Long-Horizon. Top : TUD’s operating curve as the percentile p varies (each marker = one p , e.g., p70 for p=0.70 ). Bottom : each non-adaptive baseline at its single operating point.
Forward
ms ↓
Hz ↑
Generalist
105.91
9.44
Specialist
013.53
73.91
Table 3 : Per-call latency.
Method
Calls
SR (%) ↑
Wall (s) ↓
High budget ( ≈ 17–18 calls/ep)
Diff-DAgger
16.5
83.7
31.1
BayesDiff
16.3
83.5
11.9
Visual
17.2
83.0
11.8
AAC-ent.
16.4
82.7
28.7
MC-AleaUQ
19.1
80.2
11.6
Table 4 : VLM re-invocation criteria at matched call budgets on VLA-Arena L0.
Figure 4 : Uncertainty signals on success vs failure rollouts. For each VLA-Arena domain (one block per domain, one sub-panel per signal), we plot each signal in the final 80 steps before the rollout ends, separately for successful (green) and failed (red) episodes. Signals are z -normalized per method. Only ut shows a consistent gap between success and failure across all four domains.
Figure 5 : Tracking ut within a representative rollout. Long-horizon episode for “Open the top layer of the cabinet” . x -axis: environment time within the rollout. y -axis: z -normalized ut (same normalization as Figure 4 ). Red circles mark the highest- ut peaks plus the two rollout endpoints, with the corresponding RGB frames shown in time order below.
Method
SR ↑
Calls ↓
Wall (s) ↓
Single-system π0
0.433
49.0
79.5
Fixed Δ=9
0.467
48.5
76.4
Fixed Δ=25
0.433
34.6
76.1
Fixed Δ=50
0.400
25.9
79.2
TUD p50
0.767
24.5
53.4
TUD p70
0.733
20.8
53.5
Table 5 : Real robot results. SR over 30 rollouts per method ( 10 per task). Calls and Wall (s) are per-episode generalist invocations and execution time; failures count as 120 s.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
E
αE
Interpretation
3
19.80
Tighter bound, but noisier dispersion estimate
5
28.00
Default balance
7
34.29
Smoother dispersion estimate, but looser bound
Appendix
Table 6 : Effect of the prediction-buffer size E on the dispersion coefficient αE=2daE−1 with da=7 .
E
Minimum
p1
p5
p10
3
1.472
1.654
2.140
2.425
5
2.009
2.210
2.834
3.220
7
2.209
2.742
3.278
3.672
9
2.324
3.054
3.624
4.036
11
2.557
3.224
3.801
4.188
Appendix
Table 7 : Empirical bound-to-deviation ratio rt on held-out trajectories. Larger values indicate a more conservative bound.
Factor
Setting
SR ↑
Calls ↓
Buffer E
3
0.832
27.3
5
0.850
21.6
7
0.850
18.5
Measurement Δs
1
0.848
28.0
2
0.850
21.6
4
0.852
51.2
Appendix
Table 8 : Hyperparameter sensitivity of the invocation rule. Bold indicates the default setting. Calls denote generalist invocations per episode.
Table 9 : Comparison of detailed specifications between CALVIN and VLA-Arena benchmarks.
Method
Additional model evaluations
Auxiliary computation
TUD ( ut , ours)
None
std over E buffered chunks
GenUQ
M=4 sampled action chunks
last-layer Laplace weight samples
MC-AleaUQ
M=10 velocity evaluations
re-noising at t⋆=5
Diff-DAgger
N=4 loss evaluations
(τ,ϵ) draws, τ∼Beta(1.5,1)
BayesDiff
T=10 velocity evaluations
variance propagation over T Euler steps
AAC-ent.
N=20 sampled action chunks
action entropy
Appendix
Table 10 : Extra computation per uncertainty measurement in our action-chunk setting.
Figure 6 : Cost–success trade-off on VLA-Arena level 0. Left : Distractor. Right : Extrapolation. Top : TUD’s operating curve as the percentile p varies (each marker = one p , e.g., p70 for p=0.70 ). Bottom : each non-adaptive baseline at its single operating point.
Baselines
Ours (TUD)
Task
OpenVLA
OpenVLA-OFT
π0
π0 -FAST
UniVLA
SmolVLA
OneTwoVLA
RoboDual
p70
Best
Safety (SR / CC)
StaticObstacles
0.60 / 8.2
0.20 / 45.4
0.74 / 8.0
0.40 / 56.0
0.42 / 9.7
0.00 / 8.8
0.58 / 9.3
0.40 / 13.7
0.74 / 14.8
0.77 / 19.0
CautiousGrasp
0.40 / 120.2
0.50 / 6.3
0.06 / 16.4
0.06 / 15.6
0.60 / 32.1
0.28 / 30.7
0.26 / 23.8
0.18 / 48.4
0.16 / 11.8
0.28 / 13.2
HazardAvoidance
0.02 / 22.8
0.00 / 22.9
0.00 / 16.8
0.00 / 15.4
0.12 / 18.3
0.00 / 19.5
0.12 / 18.9
0.14 / 18.9
0.06 / 18.9
0.11 / 18.7
StatePreservation
0.06 / 16.0
0.76 / 7.6
0.64 / 6.4
0.56 / 5.6
0.76 / 7.6
0.18 / 1.8
0.66 / 6.6
0.38 / 3.8
0.59 / 5.9
0.67 / 6.7
Appendix
Table 11 : Per-task results on VLA-Arena Level 1 (Generalization). We report Success Rate (SR) and Cumulative Cost (CC). TUD (p70) is the operating point at p=0.70 of the online invocation rule (defined in Section 3.3 ) and TUD (Best) is the per-task best operating point.
Figure 7 : Cost–success trade-off on VLA-Arena level 1 across four domains (Safety, Long-Horizon, Distractor, and Extrapolation). Top : TUD’s operating curve as the percentile p varies (each marker = one p , e.g., p70 for p=0.70 ). Bottom : each non-adaptive baseline at its single operating point.
Figure 8 : Additional per-rollout ut trajectories on VLA-Arena. From top to bottom: extrapolation (preposition combinations) failure, extrapolation (unseen objects) success, safety (cautious grasp) success, and distractor (static distractors) success. Same plotting convention as Figure 5 .
Figure 9 : Cost-success trade-off on CALVIN ABC → D evaluation (100 sequences, max 360 steps, each chaining 5 sub-tasks). The x -axis shows mean VLM calls per sub-task; the y -axis shows completed sub-tasks per chain (out of 5).
Figure 10 : Real robot setup. SO-101 6-DoF arm with third-person and gripper-view RGB cameras.
Figure 11 : Test environments. Evaluation scenes for the three task instructions, with two object positions per task.
pen
drawer
stack
mean
Method
SR
Calls
Wall (s)
SR
Calls
Wall (s)
SR
Calls
Wall (s)
SR
Calls
Wall (s)
Single-system π0
0.5
42.4
69.8
0.5
47.8
76.7
0.3
56.8
92.0
0.43
49.0
79.5
Fixed Δ=9
0.5
46.8
73.7
0.5
38.9
62.6
0.4
59.7
93.0
0.47
48.5
76.4
Fixed Δ=25
0.4
36.1
79.8
0.6
25.6
56.3
0.3
42.2
92.2
0.43
34.6
76.1
Fixed Δ=50
0.4
25.8
79.0
0.6
19.5
58.3
0.2
32.5
100.1
0.40
25.9
79.2
TUD p50
0.9
19.6
38.9
0.7
18.7
50.6
0.7
35.3
70.6
0.77
24.5
53.4
Appendix
Table 12 : Real robot comparison per task. Metrics are the same as in Table 5 . SR is the success rate; Calls and Wall are the generalist (VLM prefix) call count and the wall-clock time in seconds, both per episode. Each per-task entry is over 10 rollouts; the mean columns average the three tasks and reproduce the values in Table 5 .