Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) and world-action models (WAMs) increasingly master individual skills, yet the chain still fails: errors compound beyond the policy's ability to correct, and one subtask silently constrains the next. A promising pathway freezes the VLA and puts an LLM coding agent in charge: it plans in language, moves in free space with analytic primitives, invokes the VLA only for contact-rich segments, and writes adaptation into language memory. Yet applied to long horizons, this recipe breaks twice. (1) Its competence comes from whole-task exploration at test time, whose cost is exponential in the number of stages: if one stage needs T episodes, a K-stage task needs on the order of T^K, and a failure does not reveal which stage caused it. (2) It has no representation of transitions: the VLA primitive carries an exit but no entry condition, and a subtask can succeed in a form its successor cannot use. We present BATON to address both failures. Against (1), BATON makes the subtask the unit of exploration: each subtask is explored in the cheap short-horizon regime and its solution stored in memory; a long-horizon trajectory is then composed from these solutions rather than discovered whole. Exploration cost becomes linear (KT), and each failure is attributed to one stage. Against (2), BATON equips exploration with a transition-aware memory. Within a subtask, a verifier agent governs the invocation transition: the VLA is invoked only after the wrist view confirms the scene is ready. Across subtasks, a handoff transition restores an entry state disturbed by the predecessor's residue, and a lookahead transition selects the strategy whose outcome the successor can inherit. On the RoboMemArena benchmark, BATON improves task success by 37.7% and cumulative success by 29.7% over the SoTA.
Figures & tables
Figure 1: Baton treats transitions as first-class objects. Bottom : overall workflow ( Sec. 3.1 ). Left : subtask exploration makes the cost linear in the number of stages and localizes each failure ( Sec. 3.2 ). Right : the three proposed transition memories ( Sec. 3.3 ).
Figure 2: Overview of Baton . Left : Baton decomposes a long-horizon task into subtasks, explores each subtask individually ( Sec. 3.2 ), and composes them into the full task through three transition memories ( Sec. 3.3 ). Right : the skill, transition, and failure memories built during exploration.
Transferring
Occlusion
Counting
Sequence
Average
Method
TSR
CSR
TSR
CSR
TSR
CSR
TSR
CSR
TSR
CSR
GPT-5.4 (VLM only)
13.8
32.9
1.8
9.2
12.9
50.7
15.0
47.3
8.7
30.5
π0.5 [ 28 ]
20.0
42.8
12.7
17.2
14.3
50.9
60.0
71.6
21.5
38.7
HiF-VLA [ 21 ]
17.5
38.9
12.7
27.1
8.6
45.9
42.5
70.2
16.9
39.8
MemoryVLA [ 32 ]
15.0
37.2
7.3
13.1
14.3
55.1
37.5
65.2
15.0
35.3
MemER [ 35 ]
20.0
36.1
16.4
33.2
27.1
65.1
65.0
79.1
27.3
49.1
Table 1: RoboMemArena [ 18 ] : per-category and average TSR and CSR, in %. We compare with previous methods and with Harness VLA, which shares all components with Baton . The ground-truth oracle executes with perfect memory of the past and bounds what retrospection alone can achieve; it is excluded from bolding. Best per column in bold.
Cans → plate
Can+bottle → plate
Average
Method
TSR
CSR
TSR
CSR
TSR
CSR
π0.5
10.0
31.7
5.0
21.7
7.5
26.7
GPT-5.6-sol (VLM)
15.0
38.3
5.0
25.0
10.0
31.7
Harness VLA [ 44 ]
35.0
61.7
20.0
43.3
27.5
52.5
Baton (ours)
55.0
73.3
40.0
63.3
47.5
68.3
Table 2: Real-robot long-horizon experiment results. TSR/ CSR in %, 20 trials per task.
Transferring
Occlusion
Counting
Sequence
Average
Variant
TSR
CSR
TSR
CSR
TSR
CSR
TSR
CSR
TSR
CSR
w/o invocation
52.5
63.5
36.4
66.5
42.9
61.8
80.0
93.4
47.3
68.9
w/o handoff
57.5
67.1
9.1
44.4
21.4
78.5
77.5
93.3
30.4
64.6
w/o lookahead
62.5
70.8
18.2
47.8
28.6
83.3
82.5
93.8
37.7
68.0
Baton (full)
75.0
83.3
50.0
78.6
64.3
83.3
87.5
93.8
63.5
82.9
Table 3: Ablation of the three transition memories on RoboMemArena: per-category and average TSR and CSR, in %. Each variant removes one transition of Sec. 3.3 and keeps the hierarchical subtask exploration of Sec. 3.2 . Best per column in bold.
Figure 3: A successful evaluation rollout of Baton on a 1,800-step RoboMemArena task. The robot opens and closes three drawers in turn, then places the butter in the empty one.
Figure 4: Real-robot rollout of Baton . The arm transfers two cans onto the plate in turn.