One-Shot Learning from a Demonstration with Hierarchical Latent Language

Nathaniel Weir (Johns Hopkins University), Xingdi Yuan (Microsoft Research), Marc-Alexandre Côté (Microsoft Research), Matthew Hausknecht (Microsoft Research), Romain Laroche (Microsoft Research), Ida Momennejad (Microsoft Research), Harm Van Seijen (Microsoft Research), Benjamin Van Durme (Johns Hopkins University & Microsoft Semantic Machines)

Abstract

Humans have the capability, aided by the expressive compositionality of their language, to learn quickly by demonstration. They are able to describe unseen task-performing procedures and generalize their execution to other contexts. This work introduces DescribeWorld, a Minecraft-like grid world environment designed to test this sort of generalization skill in grounded agents, where tasks are linguistically and procedurally composed of elementary concepts. The agent observes a single task demonstration, and is then asked to carry out the same task in a new map. To enable such a level of generalization, we propose a neural agent infused with hierarchical latent language-at the levels of task inference and subtask planning. Through a suite of generalization tests, we find agents that perform text-based inference are better equipped for the challenge under a random split of tasks.