cs.ROAug 10, 2026

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction

Authors: Hongjin JiGuoyang XiaLuoyang SunFangxiang FengLei Ren

Organizations: The Chinese University of Hong Kong, Shenzhen · Li Auto Inc. · Beijing University of Posts and Telecommunications · Institute of Automation, Chinese Academy of Sciences · University of Chinese Academy of Sciences

Abstract

Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains difficult to use reliably in closed-loop manipulation. A shared adaptation space can mix incompatible task corrections, while an online update can alter subsequent actions before its consequences are known. We introduce a reliable TTT framework for VLA policies (VANE). VANE conditions prompt adaptation on the current vision--language context and learns from the future visual consequences of executed actions. Candidate updates are isolated from the live policy, evaluated on subsequent observations, and committed only when supported by future evidence, making adaptation selective and reversible. On SimplerEnv WidowX, VANE improves average success by 3.23.2 percentage points over the corresponding TTT baseline. Results on Google Robot further show that deployment-time gains remain task- and embodiment-dependent. Together, these results demonstrate a constrained, evidence-based approach to adapting VLA policies during interaction.

Explore similar work

CardsList