Negotiated Reasoning: On Provably Addressing Relative Over-Generalization

Junjie Sheng (East China Normal University), Wenhao Li (Tongji University), Bo Jin (Tongji University), Hongyuan Zha (The Chinese University of Hong Kong, Shenzhen), Jun Wang (East China Normal University), Xiangfeng Wang (East China Normal University & Shanghai Formal-Tech Information Technology Co., Lt)

Abstract

We focus on the relative over-generalization (RO) issue in fully cooperative multi-agent reinforcement learning (MARL). Existing methods show that endowing agents with reasoning can help mitigate RO empirically, but there is little theoretical insight. We first prove that RO is avoided when agents satisfy a consistent reasoning requirement. We then propose a new negotiated reasoning framework connecting reasoning and RO with theoretical guarantees. Based on it, we develop an algorithm called Stein variational negotiated reasoning (SVNR), which uses Stein variational gradient descent to form a negotiation policy that provably bypasses RO under maximumentropy policy iteration. SVNR is further parameterized with neural networks for computational efficiency. Experiments demonstrate that SVNR significantly outperforms baselines on RO-challenged tasks, confirming its advantage in achieving better cooperation.