Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) and world-action models (WAMs) increasingly master individual skills, yet the chain still fails: errors compound beyond the policy's ability to correct, and one subtask silently constrains the next. A promising pathway freezes the VLA and puts an LLM coding agent in charge: it plans in language, moves in free space with analytic primitives, invokes the VLA only for contact-rich segments, and writes adaptation into language memory. Yet applied to long horizons, this recipe breaks twice. (1) Its competence comes from whole-task exploration at test time, whose cost is exponential in the number of stages: if one stage needs T episodes, a K-stage task needs on the order of T^K, and a failure does not reveal which stage caused it. (2) It has no representation of transitions: the VLA primitive carries an exit but no entry condition, and a subtask can succeed in a form its successor cannot use. We present BATON to address both failures. Against (1), BATON makes the subtask the unit of exploration: each subtask is explored in the cheap short-horizon regime and its solution stored in memory; a long-horizon trajectory is then composed from these solutions rather than discovered whole. Exploration cost becomes linear (KT), and each failure is attributed to one stage. Against (2), BATON equips exploration with a transition-aware memory. Within a subtask, a verifier agent governs the invocation transition: the VLA is invoked only after the wrist view confirms the scene is ready. Across subtasks, a handoff transition restores an entry state disturbed by the predecessor's residue, and a lookahead transition selects the strategy whose outcome the successor can inherit. On the RoboMemArena benchmark, BATON improves task success by 37.7% and cumulative success by 29.7% over the SoTA.
Figures & tables
Figure 1: Baton treats transitions as first-class objects. Bottom : overall workflow ( Sec. 3.1 ). Left : subtask exploration makes the cost linear in the number of stages and localizes each failure ( Sec. 3.2 ). Right : the three proposed transition memories ( Sec. 3.3 ).
Figure 2: Overview of Baton . Left : Baton decomposes a long-horizon task into subtasks, explores each subtask individually ( Sec. 3.2 ), and composes them into the full task through three transition memories ( Sec. 3.3 ). Right : the skill, transition, and failure memories built during exploration.
Transferring
Occlusion
Counting
Sequence
Average
Method
TSR
CSR
TSR
CSR
TSR
CSR
TSR
CSR
TSR
CSR
GPT-5.4 (VLM only)
13.8
32.9
1.8
9.2
12.9
50.7
15.0
47.3
8.7
30.5
π0.5 [ 28 ]
20.0
42.8
12.7
17.2
14.3
50.9
60.0
71.6
21.5
38.7
HiF-VLA [ 21 ]
17.5
38.9
12.7
27.1
8.6
45.9
42.5
70.2
16.9
39.8
MemoryVLA [ 32 ]
15.0
37.2
7.3
13.1
14.3
55.1
37.5
65.2
15.0
35.3
MemER [ 35 ]
20.0
36.1
16.4
33.2
27.1
65.1
65.0
79.1
27.3
49.1
Table 1: RoboMemArena [ 18 ] : per-category and average TSR and CSR, in %. We compare with previous methods and with Harness VLA, which shares all components with Baton . The ground-truth oracle executes with perfect memory of the past and bounds what retrospection alone can achieve; it is excluded from bolding. Best per column in bold.
Cans → plate
Can+bottle → plate
Average
Method
TSR
CSR
TSR
CSR
TSR
CSR
π0.5
10.0
31.7
5.0
21.7
7.5
26.7
GPT-5.6-sol (VLM)
15.0
38.3
5.0
25.0
10.0
31.7
Harness VLA [ 44 ]
35.0
61.7
20.0
43.3
27.5
52.5
Baton (ours)
55.0
73.3
40.0
63.3
47.5
68.3
Table 2: Real-robot long-horizon experiment results. TSR/ CSR in %, 20 trials per task.
Transferring
Occlusion
Counting
Sequence
Average
Variant
TSR
CSR
TSR
CSR
TSR
CSR
TSR
CSR
TSR
CSR
w/o invocation
52.5
63.5
36.4
66.5
42.9
61.8
80.0
93.4
47.3
68.9
w/o handoff
57.5
67.1
9.1
44.4
21.4
78.5
77.5
93.3
30.4
64.6
w/o lookahead
62.5
70.8
18.2
47.8
28.6
83.3
82.5
93.8
37.7
68.0
Baton (full)
75.0
83.3
50.0
78.6
64.3
83.3
87.5
93.8
63.5
82.9
Table 3: Ablation of the three transition memories on RoboMemArena: per-category and average TSR and CSR, in %. Each variant removes one transition of Sec. 3.3 and keeps the hierarchical subtask exploration of Sec. 3.2 . Best per column in bold.
Figure 3: A successful evaluation rollout of Baton on a 1,800-step RoboMemArena task. The robot opens and closes three drawers in turn, then places the butter in the empty one.
Figure 4: Real-robot rollout of Baton . The arm transfers two cans onto the plate in turn.
Vision-Language-Action (VLA) policies have achieved remarkable single-step manipulation, yet they remain brittle precisely where each stage depends on what was just completed. The core issue is structural: short-window VLAs lack an explicit channel for rouxting information across sub-task boundaries, and existing memory-augmented variants either write at every frame, retrieve from demonstration-time stages, or fire at sub-goal events without performing an explicit sub-task-to-sub-task hand-off into the action expert. We identify the sub-goal completion event as the natural temporal unit for cross-subtask memory hand-off, and present WeaveLA (Weave Latent memory for Vision-Language-Action policies), a cross-subtask memory interface that, on top of a frozen VLA backbone, compresses each completed segment into latent tokens via query-driven attention pooling and routes them directly into the action-generation path of the next sub-task. This event-triggered, action-side design preserves the base policy's short-window interface while adding a lightweight cross-subtask channel. Through stratified evaluation on RoboMME with a π0.5 backbone, WeaveLA's gains land exactly where the channel is needed: on the hardest repetition slice (SwingXtimes, N=3), success rises from 0% to 47.8%, while single-execution episodes remain unchanged. Per-episode paired analysis confirms the gains are confined to tasks whose causal structure requires cross-subtask information.
Shoujing Zhu, Zhenyang Liu, Fungmiu Wang +6
Fudan University · Shanghai Innovation Institute · School of Data Science, The Chinese University of Hong Kong, Shenzhen +1
Reactive vision-language-action (VLA) policies suffer from task-state aliasing in long-horizon manipulation, where identical multimodal inputs call for distinct, context-dependent actions. Given that pretrained VLAs already possess rich control primitives to express diverse behaviors, we hypothesize that the execution bottleneck lies not in policy capacity, but in input ambiguity. In this paper, we propose TaskAnchor, a lightweight adapter that grounds task state by injecting execution context into the VLA's native input space. During post-training, TaskAnchor learns to represent the semantic execution stage as a milestone-supervised coordinate prepended to the language instruction, while incorporating fine-grained historical evidence via a residual update to the current visual tokens. This formulation avoids generating complex subtask instructions and leaves the backbone architecture unchanged. Across long-horizon benchmarks, TaskAnchor delivers substantial gains, achieving approximately 6 times the average success rate of the pi0.5 and X-VLA baselines on RMBench and more than doubling the task success rate of pi0.5 on RoboMemArena. Real-robot experiments further validate reliable multi-stage execution, with the same policy adapting its subsequent behaviors using earlier human interactions as in-context cues. Our project website is available at https://taskanchor.netlify.app/.
Hengyan Liu, Wenlve Zhou, Bo Yue +7
The Chinese University of Hong Kong, Shenzhen · Foshan University · DexForce +1
Humans perform long-horizon manipulation by retaining knowledge of what earlier actions have established while continuously adapting the motion underway. By contrast, action-chunked vision-language-action (VLA) policies repeatedly replan from the current input at each query. Existing methods preserve either long-term task evidence through memory or short-term motion through action reuse and ensembling, leaving the cross-query handoff incomplete. We introduce ChainVLA, a 1.2B-parameter VLA policy that chains successive queries through a joint and revisable execution state. Progress Context combines a recurrent Working State with sparse event memory to carry observation-derived task progress, while Motion Tail feeds the preceding prediction's unexecuted continuation into state construction and action generation. Together, the two components condition a decoder that regenerates each action horizon under the latest observation, allowing the carried state to guide the next prediction without fixing it. ChainVLA reaches 62.8% average success on RMBench and 98.8% across four LIBERO suites, while removing Motion Tail or Progress Context reduces RMBench success to 11.2% and 3.0%, respectively. These asymmetric ablations are consistent with motion continuity helping preserve the observation stream from which task progress is inferred.
Yuzhi Huang, Weijue Bu, Ziyi Xiong +4
Shenzhen International Graduate School, Tsinghua University · China University of Mining and Technology · Shenzhen Technology University