Effects of Task Similarity on Policy Transfer with Selective Exploration in Reinforcement Learning

Akshay Narayan (National University of Singapore)

Abstract

The SEAPoT algorithm [9] is a knowledge transfer mechanism in model-based reinforcement learning. By constructing subspaces around the changed regions, and selectively and efficiently exploring the target task, the transfer is most effective when the source and target tasks share similar objectives but differ in the transition dynamics. In this work, we identify the similarity between tasks using a new lightweight metric, based on the Jensen-Shannon distance, and show how the degree of similarity affects the transfer efficacy. We also empirically show that SEAPoT performs better in terms of jump starts and average rewards, as compared to the state-of-the-art policy reuse methods.