eess.SYJan 27, 2026

Model-Free Output Feedback Stabilization via Policy Gradient Methods

Authors: Ankang Zhang, Ming Chi, Xiaoling Wang, Lintao Ye

Organizations: School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China · College of Automation, Nanjing University of Posts and Telecommunications, Nanjing 210023, China

Abstract

Stabilizing a dynamical system is a fundamental problem that serves as a cornerstone for many complex tasks in the field of control systems. The problem becomes challenging when the system model is unknown. Among the Reinforcement Learning (RL) algorithms that have been successfully applied to solve problems pertaining to unknown linear dynamical systems, the policy gradient (PG) method stands out due to its ease of implementation and can solve the problem in a model-free manner. However, most of the existing works on PG methods for unknown linear dynamical systems assume full-state feedback. In this paper, we take a step towards model-free learning for partially observed linear dynamical systems with output feedback and focus on the fundamental stabilization problem of the system. We propose an algorithmic framework that stretches the boundary of PG methods to the problem without global convergence guarantees. We show that by leveraging zeroth-order PG update based on system trajectories and its convergence to stationary points, the proposed algorithms return a stabilizing output feedback policy for discrete-time linear dynamical systems. We also explicitly characterize the sample complexity of our algorithm and verify the effectiveness of the algorithm using numerical examples.

Figures & tables

Explore similar work

CardsList
  1. A Memory Efficient Unified Algorithm for Online Learning of Linear Dynamical Systems

    Jul 2, 2026Yuval Ran-Milo, Angelos Assos, Elad HazanNonlinear DynamicsDynamical Systems

  2. Sufficiency of Zeroth-Order Reward Shaping for Policy Gradient in Stabilization Control

    Sep 28, 2026Yisheng Zhang, Tao Wang, Sicun GaoPotential-Based Reward ShapingPolicy Gradient

  3. Sample Complexity of Linear Quadratic Regulator Without Initial Stability

    Feb 20, 2025Amirreza Neshaei Moghaddam, Alex Olshevsky, Bahman GharesifardLinear Quadratic RegulatorQ-Learning