The Holy Grail of Multi-Robot Planning: Learning to Generate Online-Scalable Solutions from Offline-Optimal Experts

Amanda Prorok (University of Cambridge), Jan Blumenkamp (University of Cambridge), Qingbiao Li (University of Cambridge), Ryan Kortvelesy (University of Cambridge), Zhe Liu (University of Cambridge), Ethan Stump (DEVCOM Army Research Laboratory)

Abstract

Many multi-robot planning problems are burdened by the curse of dimensionality, which compounds the difficulty of applying solutions to large-scale problem instances. The use of learning-based methods in multi-robot planning holds great promise as it enables us to offload the online computational burden of expensive centralized, yet optimal solvers, to an offline learning procedure. The hope is that by training a policy to copy an optimal pattern generated by a small-scale (centralized) system, we can transfer that policy to much larger, decentralized systems while maintaining near-optimal performance. Yet, a number of issues impede us from leveraging this idea to its full potential. This blue-sky paper elaborates some of the key challenges that remain.