Towards Zero Shot Learning in Restless Multi-armed Bandits

Yunfan Zhao (Harvard University), Nikhil Behari (Harvard University), Edward Hughes (Google), Edwin Zhang (Harvard University), Dheeraj Nagaraj (Google), Karl Tuyls (Google), Aparna Taneja (Google), Milind Tambe (Harvard University & Google)

Abstract

Restless multi-arm bandits (RMABs), a class of resource allocation problems with broad application in areas such as healthcare, online advertising, and anti-poaching, have recently been studied from a multi-agent reinforcement learning perspective. Prior RMAB research suffers from several limitations, e.g., it fails to adequately address continuous states, and requires retraining from scratch when arms opt-in and opt-out over time, a common challenge in many real world applications. We propose a neural network-based pre-trained model that has general zero-shot ability on a wide range of previously unseen RMABs.