Advancing Sample Efficiency and Explainability in Multi-Agent Reinforcement Learning

Zhicheng Zhang (Carnegie Mellon University)

Abstract

Multi-Agent Reinforcement Learning (MARL) holds promise for complex real-world applications but faces challenges in sample efficiency and policy explainability. My dissertation aims to address these critical barriers, advancing MARL towards more practical and interpretable systems. To boost sample efficiency, it is crucial for agents to effectively learn from and generalize past experiences. We propose a meta-exploration technique to train meta-exploration policies that exploit the joint state-action space structure from metatraining tasks. This approach can be integrated with any off-policy MARL algorithm to improve learning efficiency. Complementing the efficiency gain, my research also focuses on augmenting the explainability of neural network policies' decision-making processes using techniques such as decision-tree extraction from MARL networks. In this extended abstract, I will summarize my research so far and outline promising future directions to further the deployability of MARL in complex real-world environments.