Crossmodal Attentive Skill Learner

Shayegan Omidshafiei (Massachusetts Institute of Technology), Dong-Ki Kim (Massachusetts Institute of Technology), Jason Pazis (Amazon Alexa), Jonathan P. How (Massachusetts Institute of Technology)

Abstract

This paper introduces the Crossmodal Attentive Skill Learner (CASL), integrated with the recently-introduced Asynchronous Advantage Option-Critic (A2OC) architecture [15] to enable hierarchical reinforcement learning across multiple sensory inputs. We provide concrete examples where the approach not only improves performance in a single task, but accelerates transfer to new tasks. We demonstrate the attention mechanism anticipates and identifies useful latent features, while filtering irrelevant sensor modalities during execution. We modify the Arcade Learning Environment [7] to support audio queries, and conduct evaluations of crossmodal learning in the Atari 2600 games H.E.R.O. and Amidar. Finally, building on the recent work of Babaeizadeh et al. [4], we open-source a fast hybrid CPU-GPU implementation of CASL. 1