Teaching Agents Through Correction
Abstract
We motivate and describe a novel task which is modelled on interactions between apprentices and expert teachers. In the task an agent must learn to build towers constrained by rules. The teacher provides verbal corrective feedback from which the agent learns. The agent starts out unaware of the constraints as well as the domain concepts in which the constraints are expressed. Therefore an agent that takes advantage of the linguistic evidence must learn the denotations of neologisms and adapt its conceptualisation of the planning domain to incorporate those denotations. We show that an agent which does utilise linguistic evidence outperforms a strong baseline which does not.