Multi-Scale Reward Shaping via an Off-Policy Ensemble
Abstract
We propose a potential-based reward shaping architecture that is able to reduce learning speed, with no prior tuning and extra environment samples required, via considering an off-policy ensemble of value functions learning on a variety of heuristics with a variety of scales.