Learning Manner of Execution from Partial Corrections
Abstract
Some actions must be executed in different ways depending on the context. Wiping away marker requires vigorous force while almonds require gentle force. We provide a model where an agent learns which manner to execute in which context, drawing on evidence from trial and error and verbal corrections when it makes a mistake (e.g., "no, do it gently"). The learner's initial domain model lacks the concepts denoted by the words in the teacher's feedback: both those describing the context (e.g., almonds) and those describing manner (e.g., gently). We show that discourse coherence helps the agent refine its domain model and perform the symbol grounding that's necessary for using the guidance to solve its planning problem: to perform its actions in the current context in the correct way.