Distributed multi-agent temporal-difference learning with full neighbor information
Zhinan Peng1
Jiangping Hu1
Rui Luo1
Bijoy K.Ghosh2
1.School of Automation Engineering, University of Electronic Science and Technology of China, Chengdu 611731,Sichuan, China2.School of Automation Engineering, University of Electronic Science and Technology of China, Chengdu 611731,Sichuan, China;Department of Mathematics and Statistics, Texas Tech University, Lubbock, TX 79409-1042, USA
摘要:This paper presents a novel distributed multi-agent temporal-difference learning framework for value function approximation,which allows agents using all the neighbor information instead of the information from only one neighbor. With full neighbor information, the proposed framework (1) has a faster convergence rate, and (2) is more robust compared to the state-of-the-art approaches. Then we propose a distributed multi-agent discounted temporal difference algorithm and a distributed multi-agent average cost temporal difference learning algorithm based on the framework. Moreover, the two proposed algorithms'theoretical convergence proofs are provided. Numerical simulation results show that our proposed algorithms are superior to the gossip-based algorithm in convergence speed, robustness to noise and time-varying network topology.
机标关键词:
论文发表日期:2020-11-05
在线出版日期:2025-08-15(本平台首次上网日期,不代表文献的发表时间)
页数:11( 379-389 )
英文信息
