stat.MLOct 4, 2026

Taylor Representations for Model-Free RL in Networked MDPs

Authors: Salah Chikhi, Abdelhaq Chaoui, Asuman Ozdaglar, Saurabh Amin

Organizations: Massachusetts Institute of Technology

Abstract

In Networked Markov Decision Processes, transition dynamics are often unknown and the state--action space grows rapidly with the number of agents. In this setting, Taylor representations naturally approximate QQ-functions, but a naive order-nn expansion over NN agents requires Θ(Nn)Θ(N^n) coefficients. We justify these expansions under smooth expected future local rewards with controlled derivatives. Under this condition, finite-speed information propagation and discounting imply that local-critic Taylor coefficients decay exponentially with the graph distance to the farthest agent involved. Discarding distant-agent coefficients and marginalizing then yield scalable local Taylor representations with a bound controlled by graph locality. Building on these representations, we propose a scalable model-free actor--critic algorithm, establishing finite-sample critic and near-stationarity guarantees for a linear LSTD critic. We then introduce a more expressive neural TD parameterization. Unlike prior constructive spectral methods, our approach covers settings without access to a known local dynamics map, such as hidden switched linear--quadratic regulation. Across three control benchmarks, our method matches or outperforms spectral baselines while scaling efficiently to large graphs.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces

    Jul 20, 2026Dongming Wang, Pengcheng Dai, Wenwu Yu +1Multi-Agent Reinforcement LearningFrictive Policy Optimization

  2. Representation Learning Enables Scalable Multitask Deep Reinforcement Learning

    Jun 4, 2026Johan Obando-Ceron, Lu Li, Scott Fujimoto +3Multi-Turn Reinforcement LearningModel-Based Reinforcement Learning

  3. Collaborative Yet Personalized Policy Training: Single-Timescale Federated Actor-Critic

    May 14, 2026Leo Muxing Wang, Pengkun Yang, Lili SuSoft Actor-CriticMulti-Agent Reinforcement Learning