Benchmarking Robustness and Generalization in Multi-Agent Systems: A Case Study on Neural MMO

Yangkun Chen (Shenzhen International Graduate School, Tsinghua University & Parametrix.ai.), Joseph Suarez (Massachusetts Institute of Technology), Junjie Zhang (Shenzhen International Graduate School, Tsinghua University & Parametrix.ai.), Chenghui Yu (Shenzhen International Graduate School, Tsinghua University & Parametrix.ai.), Bo Wu (Parametrix.ai), Hanmo Chen (Shenzhen International Graduate School, Tsinghua University & Parametrix.ai.), Hengman Zhu (Parametrix.ai), Rui Du (bilibili.), Shanliang Qian (bilibili.), Shuai Liu (bilibili.), Weijun Hong (NetEase Games AI Lab), Jinke He (Delft University of Technology), Yibing Zhang (Chengdu Goldwin Electronics Technology Co., Ltd), Liang Zhao (International Digital Economy Academy), Clare Zhu (Stanford University), Julian Togelius (New York University), Sharada Mohanty (AICrowd), Jiaxin Chen (Parametrix.ai), Xiu Li (Shenzhen International Graduate School, Tsinghua University), Xiaolong Zhu (Parametrix.ai), Phillip Isola (Massachusetts Institute of Technology)

Abstract

We present the results of the second Neural MMO challenge, hosted at IJCAI 2022, which received 1600+ submissions. This competition targets robustness and generalization in multi-agent systems: participants train teams of agents to complete a multi-task objective against opponents not seen during training. We summarize the competition design and results and suggest that, considering our work as a case study, competitions are an effective approach to solving hard problems and establishing a solid benchmark for algorithms. We will open-source our benchmark including the environment wrapper, baselines, a visualization tool, and selected policies for further research.