Distributed policy evaluation via inexact ADMM in multi-agent reinforcement learning
Xiaoxiao Zhao1
Peng Yi2
Li Li2
1.College of Electronic and Information Engineering, Tongji University, Shanghai 201804, China2.College of Electronic and Information Engineering, Tongji University, Shanghai 201804, China;Institute of Intelligent Science and Technology, Tongji University, Shanghai 201203, China
摘要:This paper studies a distributed policy evaluation in multi-agent reinforcement learning. Under cooperative settings, each agent only obtains a local reward, while all agents share a common environmental state. To optimize the global return as the sum of local return, the agents exchange information with their neighbors through a communication network. The mean squared projected Bellman error minimization problem is reformulated as a constrained convex optimization problem with a consensus constraint; then, a distributed alternating directions method of multipliers (ADMM) algorithm is proposed to solve it. Furthermore, an inexact step for ADMM is used to achieve efficient computation at each iteration. The convergence of the proposed algorithm is established.
机标关键词:
论文发表日期:2020-11-05
在线出版日期:2025-08-15(本平台首次上网日期,不代表文献的发表时间)
页数:17( 362-378 )
英文信息展开
控制理论与技术(英文版)

控制理论与技术(英文版)

EI
ISSN:2095-6983
年,卷(期):2020,18(4)