Centralized Cooperative Exploration Policy for Continuous Control Tasks

Chao Li (Institute of Automation, Chinese Academy of Sciences), Chen Gong (Institute of Automation, Chinese Academy of Sciences), Qiang He (University of Tubingen), Xinwen Hou (Institute of Automation, Chinese Academy of Sciences), Yu Liu (Institute of Automation, Chinese Academy of Sciences)

Abstract

Despite recent works making great progress in continuous control tasks, exploration in these tasks has remained insufficiently investigated. This paper proposes CCEP (Centralized Cooperative Exploration Policy), which utilizes estimation biases of value functions to contribute to the exploration capacity. CCEP keeps two value functions initialized with different parameters, and generates diverse policies with multiple exploration styles from a pair of value functions. In addition, a centralized policy framework ensures that CCEP achieves message delivery between multiple policies, furthermore contributing to exploring the environment cooperatively. Extensive experimental results demonstrate that CCEP achieves higher exploration capacity. Empirical analysis shows diverse exploration styles in the learned policies by CCEP, reaping benefits in more exploration regions. Besides, the exploration capabilities of CCEP have been demonstrated to outperform current state-ofthe-art methods on multiple continuous control tasks.