Mean Field Correlated Imitation Learning

Zhiyu Zhao (Institute of Automation, CAS & School of Artificial Intelligence, UCAS), Chengdong Ma (Institute for Artificial Intelligence, Peking University), Qirui Mi (Institute of Automation, CAS & School of Artificial Intelligence, UCAS), Ning Yang (Institute of Automation, CAS & School of Artificial Intelligence, UCAS), Xue Yan (Institute of Automation, CAS & School of Artificial Intelligence, UCAS), Mengyue Yang (University of Bristol), Haifeng Zhang (Institute of Automation, CAS & School of Artificial Intelligence, UCAS), Jun Wang (University College London), Yaodong Yang (Institute for Artificial Intelligence, Peking University)

Abstract

Modeling the behaviors of many-agent games is crucial for capturing the dynamics of large-scale complex systems. This is typically achieved by recovering policies from demonstrations within the Mean Field Game Imitation Learning (MFGIL) framework. However, most MFGIL methods assume that demonstrations are collected from Mean Field Nash Equilibrium (MFNE), implying that agents make decisions independently. When directly applied to situations where agents' decisions are coordinated, such as publicly routed traffic networks, these techniques often fall short. In this paper, we propose the Adaptive Mean Field Correlated Equilibrium (AMFCE), which introduces a generalized assumption that effectively integrates the correlated behaviors common in real-world systems. We prove the existence of AMFCE under mild conditions and theoretically show that MFNE is a special case of AMFCE. Building upon this, we introduce a new Mean Field Correlated Imitation Learning (MFCIL) algorithm, which recovers expert policy more accurately in scenarios where agents' decisions are coordinated. We also provide a theoretical upper bound for the error in recovering the expert policy, which is tighter than that of existing methods. Empirical results on real-world traffic flow prediction and large-scale economic simulations demonstrate that MFCIL significantly improves the predictive performance of large populations' behaviors compared to existing MFGIL baselines. This improvement highlights potential of MFCIL to model real-world multi-agent systems.