Triggering Generalist Reasoning via Predictive Uncertainty for Dual-System VLA
Organizations: KAIST
Abstract
Dual-system Vision-Language-Action (VLA) models improve real-time robotic control by pairing a slow, reasoning-capable generalist with a fast specialist action expert. However, existing methods invoke the generalist at a fixed frequency, ignoring the fact that decision-making complexity varies throughout a rollout. This static strategy wastes computation in easy phases and can delay renewed reasoning when the scene changes unexpectedly. We propose TUD (Triggering generalist reasoning via predictive Uncertainty for Dual-system VLA), an adaptive inference framework that selectively skips unnecessary generalist calls. TUD measures the cross-step dispersion of action re-predictions at the upcoming chunk slot under the cached generalist context, as a predictive uncertainty signal. This signal captures how much the future action plan shifts as new observations arrive and is computed from forwards the architecture already runs, requiring neither manual phase labels nor an auxiliary uncertainty model. On VLA-Arena, it achieves a higher success rate at matched call budgets than alternative uncertainty baselines while maintaining low wall-clock overhead, and more consistently separates successful from failed rollouts. Also, TUD finds a more favorable cost-success trade-off than non-adaptive baselines, tracing an entire operating curve as a single threshold is varied, and substantially reduces VLM calls at matched success rate. The same trade-off appears in our real-robot experiments, where TUD cuts generalist calls by 75% relative to the strongest fixed-interval baseline while achieving an even higher success rate. Our results suggest that predictive uncertainty provides a practical criterion for adaptive reasoning in efficient VLA control.
Figures & tables
| Spec. | Gen. | Dual | |
|---|---|---|---|
| Throughput (Hz) | 36.98 | 2.70 | 14.79 |
| Latency (s) | 0.216 | 0.370 | - |
| Baselines | Ours (TUD) | |||||||||
| Task | OpenVLA | OpenVLA-OFT | -FAST | UniVLA | SmolVLA | OneTwoVLA | RoboDual | p70 | Best | |
| Safety (SR / CC) | ||||||||||
| StaticObstacles | 0.60 / 0.0 | 1.00 / 0.0 | 0.98 / 0.0 | 1.00 / 0.0 | 0.84 / 0.0 | 0.14 / 0.0 | 0.94 / 0.0 | 0.68 / 0.0 | 0.97 / 0.0 | 0.99 / 0.0 |
| CautiousGrasp | 0.80 / 6.6 | 0.60 / 3.3 | 0.84 / 3.5 | 0.64 / 3.3 | 0.80 / 3.3 | 0.52 / 2.5 | 0.72 / 2.5 | 0.42 / 2.2 | 0.93 / 3.2 | 0.93 / 3.2 |
| HazardAvoidance | 0.20 / 17.2 | 0.36 / 9.4 | 0.74 / 6.4 | 0.10 / 10.4 | 0.70 / 5.3 | 0.16 / 10.4 | 0.56 / 6.7 | 0.68 / 5.3 | 0.85 / 5.5 | 0.85 / 5.5 |
| StatePreservation | 1.00 / 0.0 | 1.00 / 0.0 | 0.98 / 0.0 | 0.60 / 0.0 | 0.90 / 0.0 | 0.50 / 0.0 | 0.84 / 0.0 | 0.82 / 0.0 | 0.95 / 0.0 | 0.97 / 0.0 |
| Forward | ms | Hz |
|---|---|---|
| Generalist | ||
| Specialist |
| Method | Calls | SR (%) | Wall (s) |
|---|---|---|---|
| High budget ( 17–18 calls/ep) | |||
| Diff-DAgger | 16.5 | 83.7 | 31.1 |
| BayesDiff | 16.3 | 83.5 | 11.9 |
| Visual | 17.2 | 83.0 | 11.8 |
| AAC-ent. | 16.4 | 82.7 | 28.7 |
| MC-AleaUQ | 19.1 | 80.2 | 11.6 |
| Method | SR | Calls | Wall (s) |
|---|---|---|---|
| Single-system | 0.433 | 49.0 | 79.5 |
| Fixed | 0.467 | 48.5 | 76.4 |
| Fixed | 0.433 | 34.6 | 76.1 |
| Fixed | 0.400 | 25.9 | 79.2 |
| TUD p50 | 0.767 | 24.5 | 53.4 |
| TUD p70 | 0.733 | 20.8 | 53.5 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Interpretation | ||
|---|---|---|
| 3 | 19.80 | Tighter bound, but noisier dispersion estimate |
| 5 | 28.00 | Default balance |
| 7 | 34.29 | Smoother dispersion estimate, but looser bound |
| Minimum | p1 | p5 | p10 | |
|---|---|---|---|---|
| 3 | 1.472 | 1.654 | 2.140 | 2.425 |
| 5 | 2.009 | 2.210 | 2.834 | 3.220 |
| 7 | 2.209 | 2.742 | 3.278 | 3.672 |
| 9 | 2.324 | 3.054 | 3.624 | 4.036 |
| 11 | 2.557 | 3.224 | 3.801 | 4.188 |
| Factor | Setting | SR | Calls |
|---|---|---|---|
| Buffer | 3 | 0.832 | 27.3 |
| 5 | 0.850 | 21.6 | |
| 7 | 0.850 | 18.5 | |
| Measurement | 1 | 0.848 | 28.0 |
| 2 | 0.850 | 21.6 | |
| 4 | 0.852 | 51.2 |
| Specification | CALVIN [ 1 ] | VLA-Arena [ 18 ] |
|---|---|---|
| Input Modalities | RGB, Depth, Proprioception, Language | RGB, Language |
| Visual Observation | Static RGB ( ) Depth Static ( ) Gripper RGB ( ) Depth Gripper ( ) Tactile Image ( ) | Multi-view RGB Cameras (Typically resized to standard inputs, e.g., or ) |
| Action Space | 7-DoF Continuous Control - Relative EE translation (3) - Relative EE rotation (3) - Gripper state (1, discrete) | 7-DoF Continuous Control - Absolute EE pos & ori (6) - Gripper state (1, discrete) |
| Unique Features | Proprioceptive State: 15-dim vector (EE pos/ori, joints, gripper) Control Freq: 30 Hz | Perturbation Inputs: Visual (V0–V4), Language (W0–W4) Task Diversity: 170 tasks (L0, L1, L2) |
| Method | Additional model evaluations | Auxiliary computation |
|---|---|---|
| TUD ( , ours) | None | std over buffered chunks |
| GenUQ | sampled action chunks | last-layer Laplace weight samples |
| MC-AleaUQ | velocity evaluations | re-noising at |
| Diff-DAgger | loss evaluations | draws, |
| BayesDiff | velocity evaluations | variance propagation over Euler steps |
| AAC-ent. | sampled action chunks | action entropy |
| Baselines | Ours (TUD) | |||||||||
| Task | OpenVLA | OpenVLA-OFT | -FAST | UniVLA | SmolVLA | OneTwoVLA | RoboDual | p70 | Best | |
| Safety (SR / CC) | ||||||||||
| StaticObstacles | 0.60 / 8.2 | 0.20 / 45.4 | 0.74 / 8.0 | 0.40 / 56.0 | 0.42 / 9.7 | 0.00 / 8.8 | 0.58 / 9.3 | 0.40 / 13.7 | 0.74 / 14.8 | 0.77 / 19.0 |
| CautiousGrasp | 0.40 / 120.2 | 0.50 / 6.3 | 0.06 / 16.4 | 0.06 / 15.6 | 0.60 / 32.1 | 0.28 / 30.7 | 0.26 / 23.8 | 0.18 / 48.4 | 0.16 / 11.8 | 0.28 / 13.2 |
| HazardAvoidance | 0.02 / 22.8 | 0.00 / 22.9 | 0.00 / 16.8 | 0.00 / 15.4 | 0.12 / 18.3 | 0.00 / 19.5 | 0.12 / 18.9 | 0.14 / 18.9 | 0.06 / 18.9 | 0.11 / 18.7 |
| StatePreservation | 0.06 / 16.0 | 0.76 / 7.6 | 0.64 / 6.4 | 0.56 / 5.6 | 0.76 / 7.6 | 0.18 / 1.8 | 0.66 / 6.6 | 0.38 / 3.8 | 0.59 / 5.9 | 0.67 / 6.7 |
| pen | drawer | stack | mean | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | SR | Calls | Wall (s) | SR | Calls | Wall (s) | SR | Calls | Wall (s) | SR | Calls | Wall (s) |
| Single-system | 0.5 | 42.4 | 69.8 | 0.5 | 47.8 | 76.7 | 0.3 | 56.8 | 92.0 | 0.43 | 49.0 | 79.5 |
| Fixed | 0.5 | 46.8 | 73.7 | 0.5 | 38.9 | 62.6 | 0.4 | 59.7 | 93.0 | 0.47 | 48.5 | 76.4 |
| Fixed | 0.4 | 36.1 | 79.8 | 0.6 | 25.6 | 56.3 | 0.3 | 42.2 | 92.2 | 0.43 | 34.6 | 76.1 |
| Fixed | 0.4 | 25.8 | 79.0 | 0.6 | 19.5 | 58.3 | 0.2 | 32.5 | 100.1 | 0.40 | 25.9 | 79.2 |
| TUD p50 | 0.9 | 19.6 | 38.9 | 0.7 | 18.7 | 50.6 | 0.7 | 35.3 | 70.6 | 0.77 | 24.5 | 53.4 |
| Path | Image action | Cycle |
|---|---|---|
| Generalist | 1077 109 | 1107 104 |
| Specialist | 94 3 | 120 3 |
| Method | Image action | Cycle |
|---|---|---|
| mean / p95 | mean SD | |
| 70 / 1076 | 95 256 | |
| Fixed | 115 / 1180 | 136 321 |
| Fixed | 60 / 96 | 82 200 |
| Fixed | 39 / 94 | 62 140 |
| TUD p50 | 59 / 100 | 81 202 |