Heuristics-Assisted Experience Replay Strategy for Cooperative Multi-Agent Reinforcement Learning
Abstract
Cooperative Multi-agent Reinforcement Learning (CMARL) has great potential for developing coordinated strategies that optimize team performance. However, common methods often fail to properly separate and utilize individual experiences due to a lack of effective team reward decomposition. The Heuristics-assisted Experience Replay Strategy (HAER) addresses this by decomposing team rewards into individual rewards and enabling efficient experience replay in MARL. By maintaining network gradient invariance, we derive a partial differential equation for the individual reward function, allowing accurate calculation of TD-errors and experience importance. The Cooperative Multi-Objective Swarm Optimization (CMOSO) algorithm is used to balance TD-errors and individual rewards for efficient learning. Extensive experiments on benchmarks demonstrate HAER's effectiveness, with up to a 17.6% performance boost in the homogeneous SMACV2 scenario and an average 8% improvement in GRF for heterogeneous agent cooperation.