Reputation-Filtered Reward Reshaping: Encouraging Cooperation in High Dimensional Semi-Cooperative Multi-agent Settings
Abstract
In semi-cooperative settings, cooperation is induced by appropriate incentives that align individual agents' goals with a common objective. The primary challenge is balancing personal and collective goals, which introduces new complications. A key issue is that cooperating with all agents equally can result in poor decisions, suboptimal cooperation, and inefficiencies in task execution. Furthermore, agents must manage the trade-off between staying connected to share cooperation-related information and pursuing their own objectives. To tackle these issues, we propose a novel framework incorporating a filtered reward-reshaping mechanism with two main components: (1) a reputation system that evaluates trust and competency, allowing agents to assess and filter peers' contributions, collaborate with reliable partners, and improve learning efficiency, and (2) a density-focused Potential-Based Reward Shaping (PBRS) mechanism that promotes connectivity and encourages exploration by adjusting rewards based on the density of agents in the observable space. Our approach was tested against PED-DQN and Independent Q-Learners, demonstrating enhanced performance in high-dimensional semi-cooperative environments. Additionally, theoretical stability analysis confirmed the system's convergence to a desirable equilibrium, ensuring long-term stability.