Learning an Effective Control Policy for a Robotic Drumstick via Self-Supervision
Abstract
We train a neural network to control a drumstick fastened to a motor. The network takes a temporally arranged sequence of desired strikes, or a rhythm, as input and outputs a sequence of motor velocities controlling the drumstick's physical movement. We use a new method of training, we call Collaborative Network Training, in which three networks work together to directly minimize a nondifferentiable loss function. In this work, the goal is to minimize the difference between the input sequence and the resulting drumstick strikes on a surface produced by the network outputs. The resulting policy learned by the network works in real-time and has a precision of 10 milliseconds.