Building Trustworthy Human-Centric Autonomous Systems Via Explanations

Balint Gyevnar (University of Edinburgh)

Abstract

Autonomous systems suffer from people's mistrust, as these systems rely on highly accurate yet inscrutable black box methods that are not amenable to safety guarantees nor common sense understanding. As a result, we see the erosion of accountability, human oversight, and contestation. In an attempt to build transparency, I advocate the use of model-specific, interactive, intelligible, and causally-grounded explanations for autonomous systems that take the human factor into account. I proposed a simulationbased conversational and causal framework for explaining sequential decision-making. The method, which is called CEMA, satisfies the previous four criteria without sacrificing the performance of complex models. I verified the benefits of CEMA via extensive quantitative and qualitative evaluation involving a large user study and autonomous driving. However, future work remains. To build a trustworthy autonomous system, CEMA needs to provide explanations that accurately calibrate people's trust according to the capabilities of the system. Towards this end, I hope to exploit prior knowledge in large language models to extend CEMA into a trust calibration system that uses conversations and explanations to adjust people's trust appropriately.