Hierarchical reinforcement learning improves long-horizon control by organizing primitive actions around persistent subgoals and assigning credit at multiple temporal scales. Recent hierarchical language agents bring these benefits to interactive tasks by explicitly separating subgoal planning from action execution. We observe, however, that an explicit hierarchy does not by itself determine how stable the resulting temporal abstraction is: the learned boundary policy may replace the subgoal almost every turn, making it effectively transient, or retain a subgoal after it has stopped being appropriate. We call this temporal abstraction instability. We propose Stable Temporal Abstraction via Constrained Optimization (STAC), a constrained boundary-policy optimization method that represents premature replanning and stale persistence as constraint costs. STAC applies the resulting Lagrangian costs only to the sampled boundary decision, leaving the underlying algorithm's rewards, critic targets, subgoal advantages, and primitive-action advantages unchanged. Across two backbones and two benchmarks, STAC improves success over a strong hierarchical baseline by 8.1 and 7.9 points on ALFWorld and WebShop with Qwen3-0.6B, and by 23.5 and 15.8 points with Llama-3.2-1B-Instruct.
Figures & tables
ALFWorld
WebShop
Type
Method
Best
Final
Best
Final
Qwen3-0.6B
—
Base model
0.3±0.5
1.2±1.7
Flat
GRPO
37.5±2.8
35.9±4.7
44.7±3.2
44.7±3.2
Flat
GiGPO
59.1±7.1
59.1±7.1
51.3±5.5
51.3±5.4
Hier.
HiPER
83.3±3.3
81.5±3.6
61.4±3.5
56.5±3.0
Table 1: Main results across environments and model families. Values are mean success (%) ± SD across three seeds. “Best” selects each trained method’s highest-validation checkpoint following the released HiPER convention; “Final” uses update 150.
Figure 1: Constraint violation rates during training. Per-step violation rates for the two temporal-abstraction constraint surrogates. Each curve averages three seeds. HiPER’s premature-replanning rate stays high throughout training, whereas STAC reduces it and holds both rates low. Because the reported runs use fixed multipliers, these are diagnostics of the learned temporal abstraction rather than enforced target levels.
(a) Qwen2.5-1.5B-Instruct ALFWorld
Method
Best ↑
Final ↑
HiPER
92.5±3.7
85.2±3.6
STAC (duration)
96.4±1.2
89.9±2.1
Table 2: Robustness checks for duration-based STAC. Panel (a) is a matched local ALFWorld comparison on the Qwen2.5-1.5B-Instruct HiPER setting. Panel (b) varies the shared fixed penalty on Qwen3-0.6B ALFWorld. Values are mean ± SD over three seeds; lower violation rates are preferred. Bold marks the best value in each column.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Episodes
Kˉ
T/Kˉ
look
20
3.50
14.3
pick
32
4.09
12.2
clean
17
6.59
7.6
heat
16
6.88
7.3
cool
19
7.11
7.0
pick2
9
8.44
5.9
Appendix
Table 3: Horizon fair-share reference ⌊T/Kˉ⌋ evaluated per ALFWorld task category, with T=50 and Kˉ measured from solved episodes. The reference is lowest for the most multi-phase category and highest for the fewest-phase category, so Eq. ( 13 ) is not a constant in disguise.
Figure 2: Subgoal durations move away from one-turn segments. Median and 90th-percentile STAC segment length, each computed over all segments in a training update and then averaged across three seeds; the shaded region spans the 50th–90th percentile range. The annotation gives the fraction of one-turn segments over the first twenty and final thirty updates.
Figure 3: Main results with per-mean confidence intervals. The values of Table 1 plotted as success rates. Error bars are 95% Student- t confidence intervals for each mean, computed as t0.975,2s/3 from the three-seed standard deviation, so they describe the precision of a single arm rather than the significance of a difference. At n=3 the multiplier is t0.975,2=4.30 , which is why the intervals are wide and frequently overlap even where the Welch test in Table 4 rejects. On Llama-3.2-1B-Instruct WebShop at the final checkpoint, every baseline records exactly zero across seeds and therefore has no interval.
Benchmark
Method
Diff.
t
95% CI
pW
Qwen3-0.6B
ALFWorld best
GRPO
−45.8
−18.33
[−52.8,−38.8]
0.0001∗
GiGPO
−24.2
−5.35
[−39.1,−9.3]
0.0148∗
STAC
+8.1
2.87
[+0.3,+16.0]
0.0457∗
ALFWorld final
GRPO
−45.6
−13.34
[−55.4,−35.9]
0.0003∗
GiGPO
−22.4
−4.87
[−37.1,−7.7]
0.0170∗
Appendix
Table 4: Two-sided unpaired Welch tests against released HiPER. Differences are in success points; ∗ denotes pW<0.05 .
Figure 4: Late-training ALFWorld validation. Mean HiPER and STAC validation success across three seeds from epochs 80 through 150, without confidence intervals. Validation is performed every five epochs on 128 held-out episodes per seed. The vertical ranges differ between panels because the two backbones differ by roughly a factor of two in absolute success; a shared axis would flatten the Llama-3.2-1B panel. STAC is above HiPER at every evaluation on both backbones over this window.
Method
Short
Stale
Best ↑
Final ↑
Prem. ↓
Stale ↓
Seg. len.
Pick2 ↑
HiPER
×
×
83.3±3.3
81.5±3.6
0.553±0.041
0.147±0.026
1.70±0.24
23.1±31.0
STAC (short)
✓
×
89.2±3.8
87.0±3.9
0.085±0.021
0.185±0.031
4.70±0.61
68.5±14.6
STAC (stale)
×
✓
84.1±3.5
82.4±3.7
0.565±0.039
0.045±0.013
1.60±0.22
27.5±28.4
STAC (full)
✓
✓
91.4±3.6
89.3±3.5
0.073±0.018
0.098±0.022
4.02±0.47
81.1±7.8
Appendix
Table 5: Constraint ablation on ALFWorld (Qwen3-0.6B). “Short” and “Stale” mark which constraint is active; the HiPER row is the host with both multipliers at zero and the STAC row is the configuration used in Table 1 . “(short)” and “(stale)” are the single-constraint arms and “(full)” is STAC as reported elsewhere in the paper. Values are mean ± SD across three seeds; violation rates and mean segment length are averaged over the final thirty updates. Arrows give the preferred direction. Bold marks the best value in each column. At n=3 the violation-rate differences are resolved but the success differences between arms are not; see the text.
Method
Short
Stale
Best ↑
Final ↑
Premature ↓
Stale ↓
HiPER
×
×
61.4±3.5
56.5±3.0
0.597±0.034
0.110±0.020
STAC (short)
✓
×
67.4±2.9
64.3±3.7
0.165±0.029
0.095±0.019
STAC (stale)
×
✓
62.0±3.4
57.4±3.2
0.605±0.036
0.025±0.009
STAC (full)
✓
✓
69.3±2.4
67.5±4.2
0.151±0.027
0.039±0.011
Appendix
Table 6: Constraint ablation on WebShop (Qwen3-0.6B). Columns follow Table 5 . The same ordering holds: the short cost carries most of the success gain, the stale cost gives the lowest stale-persistence rate, and only the full method is low on both.
Method
Pick
Look
Clean
Heat
Cool
Pick2
All
Qwen3-0.6B
GRPO
51.0±3.4
21.0±18.3
43.9±9.2
31.9±15.5
25.3±17.6
40.7±1.2
37.5±2.8
GiGPO
75.1±12.0
55.6±12.7
57.9±13.9
50.1±10.3
53.3±8.3
55.0±10.0
59.1±7.1
HiPER
98.1±1.6
85.8±19.5
95.6±5.1
97.0±5.3
86.5±9.2
23.1±31.0
83.3±3.3
STAC (ours)
94.9±8.9
85.8±4.7
97.5±2.2
97.8±3.9
85.6±11.2
81.1±7.8
91.4±3.6
Llama-3.2-1B-Instruct
Appendix
Table 7: ALFWorld category accuracy (%) at the saved best checkpoints used in Table 1 , mean ± SD across three seeds. “All” is logged task-distribution-weighted overall accuracy. Bold marks the better method within each model block.
Method
Pick
Look
Clean
Heat
Cool
Pick2
Avg.
Qwen3-0.6B
GRPO
48.9±8.9
27.8±12.7
47.4±5.3
14.1±0.4
40.0±6.9
28.3±12.6
35.9±4.7
GiGPO
75.1±12.0
55.6±12.7
57.9±13.9
50.1±10.3
53.3±8.3
55.0±10.0
59.1±7.1
HiPER
96.8±5.6
83.3±0.0
98.2±3.1
92.2±7.2
86.7±2.3
23.3±31.8
81.5±3.6
STAC
97.8±3.8
91.7±8.4
93.0±8.0
84.6±10.3
92.0±6.9
73.3±17.6
89.3±3.5
Llama-3.2-1B-Instruct
Appendix
Table 8: Final-checkpoint category accuracy (%) at update 150, mean ± SD across three seeds. “Avg.” is logged task-distribution-weighted overall accuracy. Bold marks the better method within each model block.
Method
Pick
Look
Clean
Heat
Cool
Pick2
Avg.
GRPO
39.4±4.6
23.5±9.1
32.8±6.3
28.3±3.1
18.3±1.1
20.7±3.5
28.5±1.5
GiGPO
61.6±9.3
46.7±8.9
56.0±11.0
50.1±9.7
29.0±4.1
38.8±9.6
48.2±7.9
HiPER
96.7±1.3
77.4±1.5
92.4±0.6
86.2±7.7
85.0±0.4
11.6±13.3
77.9±2.0
STAC
94.9±2.0
81.6±2.6
93.8±1.9
86.7±1.4
85.9±3.8
73.8±2.6
87.4±1.7
Appendix
Table 9: Late-window per-category accuracy (%), mean ± SD across seeds. “Avg.” is logged task-distribution-weighted overall accuracy.