Heuristics-Assisted Experience Replay Strategy for Cooperative Multi-Agent Reinforcement Learning

Yi Xie (FAET, Fudan University), Ziqing Zhou (FAET, Fudan University), Chun Ouyang (FAET, Fudan University), Siao Liu (Fudan University), Linqiang Hu (Fudan University), Zhongxue Gan (Fudan University)

Abstract

Cooperative Multi-agent Reinforcement Learning (CMARL) has great potential for developing coordinated strategies that optimize team performance. However, common methods often fail to properly separate and utilize individual experiences due to a lack of effective team reward decomposition. The Heuristics-assisted Experience Replay Strategy (HAER) addresses this by decomposing team rewards into individual rewards and enabling efficient experience replay in MARL. By maintaining network gradient invariance, we derive a partial differential equation for the individual reward function, allowing accurate calculation of TD-errors and experience importance. The Cooperative Multi-Objective Swarm Optimization (CMOSO) algorithm is used to balance TD-errors and individual rewards for efficient learning. Extensive experiments on benchmarks demonstrate HAER's effectiveness, with up to a 17.6% performance boost in the homogeneous SMACV2 scenario and an average 8% improvement in GRF for heterogeneous agent cooperation.