cs.AIOct 4, 2026

Hierarchical Reinforcement Learning with Stable Temporal Abstraction for Language Model Agents

Authors: Shayan Mohajer Hamidi, Yize Cheng, Yuanda Xu, Zhengze Zhou, Alborz Geramifard

Organizations: LinkedIn Corporation · University of Maryland

Abstract

Hierarchical reinforcement learning improves long-horizon control by organizing primitive actions around persistent subgoals and assigning credit at multiple temporal scales. Recent hierarchical language agents bring these benefits to interactive tasks by explicitly separating subgoal planning from action execution. We observe, however, that an explicit hierarchy does not by itself determine how stable the resulting temporal abstraction is: the learned boundary policy may replace the subgoal almost every turn, making it effectively transient, or retain a subgoal after it has stopped being appropriate. We call this temporal abstraction instability. We propose Stable Temporal Abstraction via Constrained Optimization (STAC), a constrained boundary-policy optimization method that represents premature replanning and stale persistence as constraint costs. STAC applies the resulting Lagrangian costs only to the sampled boundary decision, leaving the underlying algorithm's rewards, critic targets, subgoal advantages, and primitive-action advantages unchanged. Across two backbones and two benchmarks, STAC improves success over a strong hierarchical baseline by 8.18.1 and 7.97.9 points on ALFWorld and WebShop with Qwen3-0.6B, and by 23.523.5 and 15.815.8 points with Llama-3.2-1B-Instruct.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction

    May 7, 2026Xiangyuan Xue, Yifan Zhou, Zidong Wang +5Agentic Reinforcement LearningAgentic Learning

  2. Milestone-Guided Policy Learning for Long-Horizon Language Agents

    May 7, 2026Zixuan Wang, Yuchen Yan, Hongxing Li +7Long-Horizon AgentsAgentic Reinforcement Learning

  3. ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents

    Aug 22, 2026Xiaoyu Wang, Qingqing Gu, Yue Zhao +5Hierarchical Reinforcement LearningAgentic Reinforcement Learning