Who Am I Dealing With? Explaining the Designer's Hidden Intentions

Turgay Caglar (Colorado State University), Sarath Sreedharan (Colorado State University), Mor Vered (Monash University)

Abstract

Explainable AI (XAI) methods are generally seen as tools that allow users a greater level of visibility into why certain decisions were made by an AI system or agent. However, by the very choice of current works to focus on merely explaining why the AI system chose to perform an action in their environment, the explanation is withholding any information about the role played by the designer of said system and environment in determining the final behavior. This information could be particularly significant when the underlying designer objectives may differ from those of the user. In this paper we propose a new explanation generation paradigm, built on the concept of model reconciliation, and show how it can support the generation of explanations that include the designer's goals. We define and study the formal properties of this new form of explanation and introduce an algorithm to generate it over a classical planning domain. We evaluate how this new explanation influences user performance, understanding and trust in an AI agent and further instantiate the new algorithm on standard planning benchmarks to evaluate its computational characteristics.