cs.CVOct 4, 2026

Triggering Generalist Reasoning via Predictive Uncertainty for Dual-System VLA

Authors: Hyemin Yang, Wooseong Jeong, Giwon Lee, Kuk-Jin Yoon

Organizations: KAIST

Abstract

Dual-system Vision-Language-Action (VLA) models improve real-time robotic control by pairing a slow, reasoning-capable generalist with a fast specialist action expert. However, existing methods invoke the generalist at a fixed frequency, ignoring the fact that decision-making complexity varies throughout a rollout. This static strategy wastes computation in easy phases and can delay renewed reasoning when the scene changes unexpectedly. We propose TUD (Triggering generalist reasoning via predictive Uncertainty for Dual-system VLA), an adaptive inference framework that selectively skips unnecessary generalist calls. TUD measures the cross-step dispersion of action re-predictions at the upcoming chunk slot under the cached generalist context, as a predictive uncertainty signal. This signal captures how much the future action plan shifts as new observations arrive and is computed from forwards the architecture already runs, requiring neither manual phase labels nor an auxiliary uncertainty model. On VLA-Arena, it achieves a higher success rate at matched call budgets than alternative uncertainty baselines while maintaining low wall-clock overhead, and more consistently separates successful from failed rollouts. Also, TUD finds a more favorable cost-success trade-off than non-adaptive baselines, tracing an entire operating curve as a single threshold is varied, and substantially reduces VLM calls at matched success rate. The same trade-off appears in our real-robot experiments, where TUD cuts generalist calls by 75% relative to the strongest fixed-interval baseline while achieving an even higher success rate. Our results suggest that predictive uncertainty provides a practical criterion for adaptive reasoning in efficient VLA control.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ReconVLA: An Uncertainty-Guided and Failure-Aware Vision-Language-Action Framework for Robotic Control

    Apr 17, 2026Lingling Chen, Zongyao Lyu, William J. BeksiDiffusion-Based Vision-Language-Actions

  2. Reflective VLA: In-Context Action Consequences Make VLAs Generalize

    Jun 23, 2026Qing Lian, Kent Yu, Lei ZhangGeneralist Vision--Language--ActionRdf

  3. ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models

    May 28, 2026Ye Li, Huanan Liu, Kangye Ji +7Diffusion-Based Vision-Language-ActionsAction Generation