How to Train PointGoal Navigation Agents on a (Sample and Compute) Budget

Erik Wijmans (Georgia Institute of Technology & Facebook AI Research), Irfan Essa (Georgia Institute of Technology & Google Atlanta), Dhruv Batra (Georgia Institute of Technology & Facebook AI Research)

Abstract

PointGoal navigation has seen significant recent interest and progress, spurred on by the Habitat platform and associated challenge [21]. In this paper, we study PointGoal navigation under both a sample budget (75 million frames) and a compute budget (1 GPU for 1 day). We conduct an extensive set of experiments, cumulatively totaling over 50,000 GPU-hours, that let us identify and discuss a number of ostensibly minor but significant design choices-the advantage estimation procedure (a key component in training), and visual encoder architecture. Overall, these design choices lead to considerable and consistent improvements. Under a sample budget, performance for RGB-D agents improves 3 SPL on Gibson (4% relative improvement) and 20 SPL on Matterport3D (43% relative improvement). Under a compute budget, performance for RGB-D agents improves by 3 SPL on Gibson (5% relative improvement) and 15 SPL on Matterport3D (50% relative improvement). Our findings and recommendations will serve to make the community's experiments more efficient-to reach 50 SPL with RGB-D on Matterport3D, they reduce the samples needed by 3x and the training time 2x.