Goal Alignment: Re-analyzing Value Alignment Problems Using Human-Aware AI
Abstract
Value alignment problems arise in scenarios where the specified objectives of an AI agent don't match the true underlying objectives of its users. While value alignment remains a popular topic within AI safety research, most existing works in this sphere tend to overlook one of the foundational causes for misalignment, namely the inherent asymmetry in human expectations about the agent's behavior and the behavior generated by the agent for the specified objective. To address this lacuna, we propose a novel formulation for the value alignment problem, named Human-aware goal alignment that highlights this central challenge related to value alignment. Additionally, we propose a first-of-its-kind interactive goal elicitation algorithm that is capable of using information generated under incorrect beliefs about the agent, to determine the true underlying goal of the user.