Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment interaction, yet many existing methods provide reward only after the complete task succeeds. However, such terminal supervision is sparse and does not distinguish early failures from rollouts that make substantial partial progress. We propose StructRL, an online RL framework that constructs structured intermediate supervision from verifiable subtask completions. StructRL decomposes each task into verifiable subtasks, grants intermediate rewards only after the prerequisite subtasks have been completed, and scales each reward according to completion pace. Across RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, StructRL consistently outperforms evaluated online RL baselines. These results show that verifiable, structured intermediate rewards improve long-horizon VLA post-training. Code is available at https://github.com/amazon-science/StructRL.
Figures & tables
Figure 1: (a) An example long-horizon VLA task from RoboCasa365 ( Nasiriany et al., 2026 ) . (b) Success rates (%) of GR00T-N1.5 on three representative composite tasks. StructRL consistently outperforms both the SFT policy and the RL baseline SimpleVLA-RL ( Li et al., 2026 ) .
Figure 2: Overview of StructRL . Left: An LLM decomposes the task command into candidate subtasks and assigns a dependency structure to the retained subtasks, represented by prerequisite sets X(vi) . The verifiability check removes candidates without a reliable binary completion criterion ( e.g . , Approach Box ), while the progress check removes candidates that do not by themselves indicate progress toward task completion ( e.g . , Open Gripper ). Right: Structure-aware reward gating uses this dependency structure to determine whether a detected completion is eligible for reward, while dynamic reward pacing determines the magnitude of that reward. The resulting chunk-level rewards are used to optimize the VLA with PPO. The training pseudocode is shown in Appendix A .
Benchmark
RoboCasa365
LIBERO-Long
Horizon (steps)
800–1000
1000–1400
1400–2900
Overall
250–340
340–400
400–550
Overall
π0 ( Black et al., 2025b )
26.0
23.0
8.1
15.5
89.6
90.0
42.0
80.2
RLDX-1 ( Kim and others, 2026 )
59.0
55.0
30.8
43.6
98.0
94.7
85.0
94.4
GR00T-N1.5 ( Bjorck et al., 2025 )
51.7
48.0
27.9
38.6
96.0
82.0
90.0
90.6
w/ Sparse-RL ( Zang et al., 2026 )
60.3
52.3
25.1
40.2
92.4
90.0
90.0
91.2
w/ SimpleVLA-RL ( Li et al., 2026 )
55.1
50.8
30.6
41.5
94.0
90.7
91.0
92.4
Table 1: Success rate (%) by task-horizon bucket. From shortest to longest, the RoboCasa365 buckets contain 3, 5, and 8 tasks, while the LIBERO-Long buckets contain 5, 3, and 2 tasks. Overall is the mean SR across all tasks in the benchmark. Within each backbone block, blue cells mark the strongest online RL baseline in each column. Δ is the difference between StructRL and that baseline, in percentage points.
LIBERO-Long
Reward Source
GR00T-N1.5
π0.5
SFT
90.6
90.6
Robometer ( Liang et al., 2026 )
94.2
93.0
StructRL ( ours )
96.6
96.2
Table 2: Reward source comparison under a matched PPO training protocol on LIBERO-Long.
Figure 3: Reward-component ablation with PPO and GR00T-N1.5, adding one component at a time to the terminal binary reward: fixed subtask rewards, dynamic reward pacing, and structure-aware gating, which together form StructRL . Panels report SR (%) by task-horizon bucket and overall on RoboCasa365 (a–d) and LIBERO-Long (e–h). Values are means over three evaluation runs.
Figure 4: Effect of subtask decomposition density on RoboCasa365 using GR00T-N1.5.
Table 7
Method
Spatial
Object
Goal
Average
GR00T-N1.5
89.3
97.9
94.6
93.7
w/ SimpleVLA-RL
91.2
98.5
94.8
94.7
w/ StructRL ( ours )
92.1
99.4
95.6
95.7
Δ
+0.9
+0.9
+0.8
+1.0
Table 5: Success rate (%) on the LIBERO Spatial, Object, and Goal suites using GR00T-N1.5. A separate policy is trained on each suite using the same optimization recipe.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Subtask phrase
Predicate is true when
Example
place o in/on r
o is contained in r or rests on it
chicken_in_bowl
grasp o
o is held by the gripper; once detected, the predicate is latched
grasp_straw
open / close f
the relevant door or drawer crosses the benchmark threshold
dishwasher_closed
turn on / off f , press f
f reaches the corresponding discrete state
water_on
⟨ activity ⟩ for a while
a simulator-maintained activity timer reaches a specified threshold
wash_t10
Appendix
Table 6: Predicate templates for grounding RoboCasa365 subtasks. o , r , and f denote a bound object, receptacle, and fixture.
Table 7: Subtask decompositions for the 16 RoboCasa365 composite-seen tasks at Nˉ=2.375 and 5.0 . Braces group unordered subtasks within a stage, while arrows indicate strict stage ordering. N denotes the number of progress signals. The Nˉ=5.0 setting additionally includes grasp milestones, intermediate fixture states, and timed activity checkpoints. Retreat signals are appended automatically and excluded from N .
Horizon (steps)
800–1000
1000–1400
1400–2900
Overall
Opus 4.8
63.0
59.0
37.6
49.1
Qwen3.5-9B
59.3
60.6
31.2
45.7
Appendix
Table 8: Ablation on the LLM that authors the subtask decomposition. Given only the task language command and our decomposition requirements, each LLM proposes the stage/subtask split; StructRL is then trained with identical hyperparameters. SR (%) grouped by horizon bucket. In Table 1 , we use Claude Opus 4.8 by default for decomposition.
Figure 5: An example of hierarchical subtask decomposition under increasing reward density. Each level refines the previous tree by splitting parent subtasks into finer children while keeping all existing milestones unchanged; lock icons mark the subtasks whose completion will be detected and rewarded during RL training at the corresponding density level.
Density Nˉ
Rewarded subtasks for StoreLeftoversInBowl
0.0
Terminal reward only.
1.0
Chicken in bowl.
2.5
Chicken in bowl; vegetable in bowl; bowl in fridge.
5.0
Grasp chicken; chicken in bowl; grasp vegetable; vegetable in bowl; grasp bowl; bowl in fridge.
10.0
Refine the six milestones with intermediate motion checkpoints: reach/approach before each grasp, and lift/carry/near-fridge before placement.
14.0
Further refine them with gripper closing, above-the-bowl checkpoints, bowl contact, and additional stages of the fridge approach.
Appendix
Table 9: Reward decomposition for StoreLeftoversInBowl at different densities. Here, Nˉ is the average number of gated subtask signals. The first four levels correspond to the trees in Fig. 5 ; denser levels progressively refine the same six milestones with finer verifiable motion checkpoints.
Figure 14
Reward configuration
Nˉ
Overall SR (%)
Signal timing and alignment
Reference completion predicates
5.0
50.4 ± 1.0
Earlier checkpoint predicates
5.0
47.4 ± 0.8
Total intermediate reward
Unbounded (fixed per-subtask reward)
14.0
44.1 ± 1.8
Total reward capped at B
14.0
47.7 ± 0.4
Appendix
Table 11: Controlled analyses of subtask reward timing and total intermediate reward on RoboCasa365 with GR00T-N1.5. Within each comparison, subtask decomposition density and dependency structure are fixed. Results are evaluated at training step 100 as mean ± sample standard deviation over three evaluation runs.
Completion duration
Overall SR (%)
Demonstration average Td(vi)
49.2 ± 1.1
Uniform horizon allocation H/N
43.1 ± 0.4
Appendix
Table 12: Completion durations used for dynamic reward pacing at the default subtask decomposition density. All other settings are fixed; results are mean ± standard deviation over three evaluation runs.
Method
800–1000
1000–1400
1400–2900
Overall
GR00T-N1.5
SFT
52.6 ± 0.8
48.3 ± 0.6
26.6 ± 1.2
38.2 ± 0.5
w/ Sparse-RL
59.8 ± 1.6
54.3 ± 0.9
27.8 ± 2.9
42.1 ± 1.4
w/ SimpleVLA-RL
54.2 ± 2.6
52.6 ± 1.7
29.6 ± 1.2
41.4 ± 0.3
w/ StructRL
63.3 ± 1.2
61.0 ± 1.7
36.5 ± 1.8
49.2 ± 1.1
π0.5
Appendix
Table 13: Multi-seed evaluation on RoboCasa365: mean ± std over 3 independent evaluation runs, SR (%) by horizon bucket. StructRL delivers the best overall SR on both backbones, with the largest margins on medium and long-horizon buckets.
Figure 7: Training dynamics. Success rate of intermediate checkpoints (every 10 steps, offline evaluation) during RL training with GR00T-N1.5 on RoboCasa365 composite tasks and LIBERO-Long. StructRL consistently outperforms SimpleVLA-RL and Sparse-RL throughout training.
Figure 8: Sensitivity of StructRL to the subtask-completion reward β on RoboCasa365 using GR00T-N1.5. The terminal reward is held fixed.
Figure 9: Qualitative comparison between SFT and StructRL on the long-horizon task “StoreLeftoversInBowl”. StructRL successfully completes the intermediate object-placement steps and proceeds to placing the bowl in the refrigerator, while SFT fails before reaching the final stage.
Figure 10: Qualitative comparison between SFT and StructRL on the long-horizon task “SteamInMicrowave”. StructRL completes the preceding subtasks and advances to interacting with the microwave, whereas SFT gets stuck at an earlier manipulation stage.
Vision-language-action (VLA) models can learn to perform diverse manipulation skills "out of the box," but achieving the precision and speed that real-world tasks demand requires further fine-tuning -- for example, via reinforcement learning (RL). We introduce a lightweight method that enables sample-efficient online RL fine-tuning of pretrained VLAs using just a few hours of real-world practice. We (1) adapt the VLA to expose an "RL token," a compact readout representation that preserves task-relevant pretrained knowledge while serving as an efficient interface for online RL, and (2) train a small actor-critic head on this RL token to refine the actions, while anchoring the learned policy to the VLA. Online RL with the RL token (RLT) makes it possible to fine-tune even large VLAs with RL quickly and efficiently. Across four real-robot tasks (screw installation, zip tie fastening, charger insertion, and Ethernet insertion), RLT improves the speed on the hardest part of the task by up to 3x and raises success rates significantly within minutes to a few hours of practice. It can even surpass the speed of human teleoperation on some of the tasks.
Charles Xu, Jost Tobias Springenberg, Michael Equi +4
Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic control. However, existing VLA models still face major challenges in long-horizon tasks: sparse expert demonstrations constrain cross-task compositional generalization; the non-Markovian nature of long-horizon tasks makes it difficult for policies conditioned only on current observations to maintain temporal consistency; limited closed-loop error correction allows execution errors to accumulate; and end-to-end action fine-tuning may weaken the high-level semantic representations of vision-language model (VLM) backbones. To address these issues, we propose a hierarchical long-horizon VLA architecture with an explicit language-memory module. The central idea is to convert discrete temporal observations into a coherent textual memory sequence with temporal logic. The system is decoupled into a high-level VLM and a low-level VLA: the high-level VLM performs semantic reasoning through a visual question answering training paradigm, while the low-level VLA executes precise continuous control conditioned on subtask instructions and visual observations. The high-level VLM recursively updates both language memory and subtask instructions using the previous memory as a contextual anchor, enabling persistent temporal tracking and dynamic correction during long-horizon execution. We evaluate the proposed method in multiple simulation environments and conduct sim-to-real experiments on a real robotic platform. The results demonstrate that explicit language memory improves the success rate and robustness of VLA models on complex long-horizon tasks while providing an interpretable semantic account of the decision process.
Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging. Recent work addresses this challenge by adapting frozen VLAs through online reinforcement learning (RL), whose sample efficiency depends on the quality of the state representation used by the actor and critic. Existing methods construct such representations either with VLA-independent visual encoders or through fixed compression of internal VLA representations. Neither design explicitly extracts the task-specific action-relevant VLA features most useful for downstream action refinement and action-value estimation, therefore limiting sample efficiency. To address this limitation, we introduce eRLT, which constructs an effective state representation by routing task-specific action-relevant information across both tokens and layers of the frozen VLA. Specifically, learned routing tokens dynamically aggregate visual-language features at multiple depths, while a lightweight layer router combines these summaries into a fixed-dimensional RL token. The routing module is initialized using expert demonstrations to capture features predictive of expert actions and then refined using critic feedback from online interactions for action-value estimation. Across seven LIBERO and RoboTwin tasks, eRLT improves mean normalized learning-curve AUC by up to 23.7% over representative baselines. Real-robot experiments on USB connector insertion and motherboard ribbon-cable insertion further show AUC improvements of 108.9% and 46.7%, respectively, over the strongest baseline.
Dehao Huang, Jianbang Liu, Jianpan Gao +7
Southern University of Science and Technology, Shenzhen, China. · Beijing Zhongguancun Academy, Beijing, China. · Samsung Robotics eXperience. +2