Transferable Environment Poisoning: Training-time Attack on Reinforcement Learner with Limited Prior Knowledge

Hang Xu (Nanyang Technological University)

Abstract

As reinforcement learning (RL) systems are deployed in various safety-critical applications, it is imperative to understand how vulnerable they are to adversarial attacks. Of these, an environmentpoisoning attack is considered particularly insidious, since environment hyper-parameters are significant in determining an RL policy yet prone to be accessed by third parties. In this work, we study an environment-poisoning attack (EPA) against RL at training time. Considering that environment alteration comes at a cost, we seek minimal poisoning in an unknown environment and aim to force a black-box RL agent to learn an attacker-designed policy.