Learning Plans by Acquiring Grounded Linguistic Meanings from Corrections

Mattias Appelgren (The University of Edinburgh)

Abstract

We motivate and describe a novel task which is modelled on interactions between apprentices and expert teachers. In the task the agent must learn to build towers which are constrained by rules. Whenever the agent performs an action which violates a rule the teacher provides verbal corrective feedback (e.g. "No, put red blocks on blue blocks") and answers the learner's clarification questions. The agent must learn to build rule compliant towers from these corrections and the context in which they were given. The agent starts out unaware of the constraints as well as the domain concepts in which the constraints are expressed. Therefore an agent that takes advantage of the linguistic evidence must learn the denotations of neologisms and adapt its conceptualisation of the planning domain to incorporate those denotations. We show that an agent which does utilise linguistic evidence outperforms a strong baseline which does not.