Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment interaction, yet many existing methods provide reward only after the complete task succeeds. However, such terminal supervision is sparse and does not distinguish early failures from rollouts that make substantial partial progress. We propose StructRL, an online RL framework that constructs structured intermediate supervision from verifiable subtask completions. StructRL decomposes each task into verifiable subtasks, grants intermediate rewards only after the prerequisite subtasks have been completed, and scales each reward according to completion pace. Across RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, StructRL consistently outperforms evaluated online RL baselines. These results show that verifiable, structured intermediate rewards improve long-horizon VLA post-training. Code is available at https://github.com/amazon-science/StructRL.
Figures & tables
Figure 1: (a) An example long-horizon VLA task from RoboCasa365 ( Nasiriany et al., 2026 ) . (b) Success rates (%) of GR00T-N1.5 on three representative composite tasks. StructRL consistently outperforms both the SFT policy and the RL baseline SimpleVLA-RL ( Li et al., 2026 ) .
Figure 2: Overview of StructRL . Left: An LLM decomposes the task command into candidate subtasks and assigns a dependency structure to the retained subtasks, represented by prerequisite sets X(vi) . The verifiability check removes candidates without a reliable binary completion criterion ( e.g . , Approach Box ), while the progress check removes candidates that do not by themselves indicate progress toward task completion ( e.g . , Open Gripper ). Right: Structure-aware reward gating uses this dependency structure to determine whether a detected completion is eligible for reward, while dynamic reward pacing determines the magnitude of that reward. The resulting chunk-level rewards are used to optimize the VLA with PPO. The training pseudocode is shown in Appendix A .
Benchmark
RoboCasa365
LIBERO-Long
Horizon (steps)
800–1000
1000–1400
1400–2900
Overall
250–340
340–400
400–550
Overall
π0 ( Black et al., 2025b )
26.0
23.0
8.1
15.5
89.6
90.0
42.0
80.2
RLDX-1 ( Kim and others, 2026 )
59.0
55.0
30.8
43.6
98.0
94.7
85.0
94.4
GR00T-N1.5 ( Bjorck et al., 2025 )
51.7
48.0
27.9
38.6
96.0
82.0
90.0
90.6
w/ Sparse-RL ( Zang et al., 2026 )
60.3
52.3
25.1
40.2
92.4
90.0
90.0
91.2
w/ SimpleVLA-RL ( Li et al., 2026 )
55.1
50.8
30.6
41.5
94.0
90.7
91.0
92.4
Table 1: Success rate (%) by task-horizon bucket. From shortest to longest, the RoboCasa365 buckets contain 3, 5, and 8 tasks, while the LIBERO-Long buckets contain 5, 3, and 2 tasks. Overall is the mean SR across all tasks in the benchmark. Within each backbone block, blue cells mark the strongest online RL baseline in each column. Δ is the difference between StructRL and that baseline, in percentage points.
LIBERO-Long
Reward Source
GR00T-N1.5
π0.5
SFT
90.6
90.6
Robometer ( Liang et al., 2026 )
94.2
93.0
StructRL ( ours )
96.6
96.2
Table 2: Reward source comparison under a matched PPO training protocol on LIBERO-Long.
Figure 3: Reward-component ablation with PPO and GR00T-N1.5, adding one component at a time to the terminal binary reward: fixed subtask rewards, dynamic reward pacing, and structure-aware gating, which together form StructRL . Panels report SR (%) by task-horizon bucket and overall on RoboCasa365 (a–d) and LIBERO-Long (e–h). Values are means over three evaluation runs.
Figure 4: Effect of subtask decomposition density on RoboCasa365 using GR00T-N1.5.
Table 7
Method
Spatial
Object
Goal
Average
GR00T-N1.5
89.3
97.9
94.6
93.7
w/ SimpleVLA-RL
91.2
98.5
94.8
94.7
w/ StructRL ( ours )
92.1
99.4
95.6
95.7
Δ
+0.9
+0.9
+0.8
+1.0
Table 5: Success rate (%) on the LIBERO Spatial, Object, and Goal suites using GR00T-N1.5. A separate policy is trained on each suite using the same optimization recipe.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Subtask phrase
Predicate is true when
Example
place o in/on r
o is contained in r or rests on it
chicken_in_bowl
grasp o
o is held by the gripper; once detected, the predicate is latched
grasp_straw
open / close f
the relevant door or drawer crosses the benchmark threshold
dishwasher_closed
turn on / off f , press f
f reaches the corresponding discrete state
water_on
⟨ activity ⟩ for a while
a simulator-maintained activity timer reaches a specified threshold
wash_t10
Appendix
Table 6: Predicate templates for grounding RoboCasa365 subtasks. o , r , and f denote a bound object, receptacle, and fixture.
Table 7: Subtask decompositions for the 16 RoboCasa365 composite-seen tasks at Nˉ=2.375 and 5.0 . Braces group unordered subtasks within a stage, while arrows indicate strict stage ordering. N denotes the number of progress signals. The Nˉ=5.0 setting additionally includes grasp milestones, intermediate fixture states, and timed activity checkpoints. Retreat signals are appended automatically and excluded from N .
Horizon (steps)
800–1000
1000–1400
1400–2900
Overall
Opus 4.8
63.0
59.0
37.6
49.1
Qwen3.5-9B
59.3
60.6
31.2
45.7
Appendix
Table 8: Ablation on the LLM that authors the subtask decomposition. Given only the task language command and our decomposition requirements, each LLM proposes the stage/subtask split; StructRL is then trained with identical hyperparameters. SR (%) grouped by horizon bucket. In Table 1 , we use Claude Opus 4.8 by default for decomposition.
Figure 5: An example of hierarchical subtask decomposition under increasing reward density. Each level refines the previous tree by splitting parent subtasks into finer children while keeping all existing milestones unchanged; lock icons mark the subtasks whose completion will be detected and rewarded during RL training at the corresponding density level.
Density Nˉ
Rewarded subtasks for StoreLeftoversInBowl
0.0
Terminal reward only.
1.0
Chicken in bowl.
2.5
Chicken in bowl; vegetable in bowl; bowl in fridge.
5.0
Grasp chicken; chicken in bowl; grasp vegetable; vegetable in bowl; grasp bowl; bowl in fridge.
10.0
Refine the six milestones with intermediate motion checkpoints: reach/approach before each grasp, and lift/carry/near-fridge before placement.
14.0
Further refine them with gripper closing, above-the-bowl checkpoints, bowl contact, and additional stages of the fridge approach.
Appendix
Table 9: Reward decomposition for StoreLeftoversInBowl at different densities. Here, Nˉ is the average number of gated subtask signals. The first four levels correspond to the trees in Fig. 5 ; denser levels progressively refine the same six milestones with finer verifiable motion checkpoints.
Figure 14
Reward configuration
Nˉ
Overall SR (%)
Signal timing and alignment
Reference completion predicates
5.0
50.4 ± 1.0
Earlier checkpoint predicates
5.0
47.4 ± 0.8
Total intermediate reward
Unbounded (fixed per-subtask reward)
14.0
44.1 ± 1.8
Total reward capped at B
14.0
47.7 ± 0.4
Appendix
Table 11: Controlled analyses of subtask reward timing and total intermediate reward on RoboCasa365 with GR00T-N1.5. Within each comparison, subtask decomposition density and dependency structure are fixed. Results are evaluated at training step 100 as mean ± sample standard deviation over three evaluation runs.
Completion duration
Overall SR (%)
Demonstration average Td(vi)
49.2 ± 1.1
Uniform horizon allocation H/N
43.1 ± 0.4
Appendix
Table 12: Completion durations used for dynamic reward pacing at the default subtask decomposition density. All other settings are fixed; results are mean ± standard deviation over three evaluation runs.
Method
800–1000
1000–1400
1400–2900
Overall
GR00T-N1.5
SFT
52.6 ± 0.8
48.3 ± 0.6
26.6 ± 1.2
38.2 ± 0.5
w/ Sparse-RL
59.8 ± 1.6
54.3 ± 0.9
27.8 ± 2.9
42.1 ± 1.4
w/ SimpleVLA-RL
54.2 ± 2.6
52.6 ± 1.7
29.6 ± 1.2
41.4 ± 0.3
w/ StructRL
63.3 ± 1.2
61.0 ± 1.7
36.5 ± 1.8
49.2 ± 1.1
π0.5
Appendix
Table 13: Multi-seed evaluation on RoboCasa365: mean ± std over 3 independent evaluation runs, SR (%) by horizon bucket. StructRL delivers the best overall SR on both backbones, with the largest margins on medium and long-horizon buckets.
Figure 7: Training dynamics. Success rate of intermediate checkpoints (every 10 steps, offline evaluation) during RL training with GR00T-N1.5 on RoboCasa365 composite tasks and LIBERO-Long. StructRL consistently outperforms SimpleVLA-RL and Sparse-RL throughout training.
Figure 8: Sensitivity of StructRL to the subtask-completion reward β on RoboCasa365 using GR00T-N1.5. The terminal reward is held fixed.
Figure 9: Qualitative comparison between SFT and StructRL on the long-horizon task “StoreLeftoversInBowl”. StructRL successfully completes the intermediate object-placement steps and proceeds to placing the bowl in the refrigerator, while SFT fails before reaching the final stage.
Figure 10: Qualitative comparison between SFT and StructRL on the long-horizon task “SteamInMicrowave”. StructRL completes the preceding subtasks and advances to interacting with the microwave, whereas SFT gets stuck at an earlier manipulation stage.