cs.ROSep 20, 2026

TaskAnchor: Grounding Task State in Reactive VLAs for Long-Horizon Manipulation

Authors: Hengyan LiuWenlve ZhouBo YueYongyi SuRuixiang WangZhanqi ZhangDekun LuWei Gao+2 more

Abstract

Reactive vision--language--action (VLA) models struggle with long-horizon manipulation when visually similar observations can correspond to different actions depending on the task stage or interaction history. We refer to this ambiguity as task-state aliasing and introduce TaskAnchor, a lightweight adapter that grounds pretrained VLAs in execution history. TaskAnchor combines history-conditioned visual refinement with a milestone-supervised task-state coordinate, a scalar representing the semantic stage of execution. These signals are injected through the native visual and language interfaces, respectively, without introducing an explicit planner or modifying the action-generation mechanism. On RMBench, TaskAnchor achieves approximately 4.9--5.5×\times the average success rates of the published π0.5π_{0.5} and X-VLA baselines, with consistent gains on RoboMemArena and real robots. The added latency is only 2.08,ms per action chunk for π0.5π_{0.5}.

Explore similar work

CardsList