Efficient Convention Emergence through Decoupled Reinforcement Social Learning with Teacher-Student Mechanism

Yixi Wang (Tianjin University), Wenhuan Lu (Tianjin University), Jianye Hao (Tianjin University), Jianguo Wei (Tianjin University), Ho-Fung Leung (Chinese University of Hong Kong)

Abstract

In this paper, we design reinforcement learning based (RLbased) strategies to promote convention emergence in multiagent systems (MASs) with large convention space. We apply our approaches to a language coordination problem in which agents need to coordinate on a dominant lexicon for efficient communication. By modeling each lexicon which maps each concept to a single word as a Markov strategy representation, the original single-state convention learning problem can be transformed into a multi-state multiagent coordination problem. The dynamics of lexicon evolutions during an interaction episode can be modeled as a Markov game, which allows agents to improve the action values of each concept separately and incrementally. Specifically we propose two learning strategies, multiple-Q and multiple-R, and also propose incorporating teacher-student mechanism on top of the learning strategies to accelerate lexicon convergence speed. Extensive experiments verify that our approaches outperform the state-of-the-art approaches in terms of convergence efficiency, convention quality and scalability.