Defining Deception in Decision Making
Abstract
With the growing capabilities of machine learning systems, particularly those that communicate or interact with humans, there is an increased risk of systems that can easily deceive and manipulate people. Preventing unintended deception and manipulation therefore represents an important challenge for creating aligned AI systems. To approach this challenge in a principled way, we first need to define deception formally. In this work, we present a concrete definition of deception under the formalism of rational decision making in partially observed Markov decision processes. We propose a general regret theory of deception under which the degree of deception can be quantified in terms of the actor's beliefs, actions, and utility. We instantiate these principles as reward terms for communication agents, and study the degree to which the behavior aligns with human judgments about deception. We hope our work will represent a step toward systems that aim to avoid deception, and detection mechanisms to identify deceptive agents.