Bradly C. Stadie, Ge Yang, Rein Houthooft, Xi Chen, Yan Duan, Yuhuai Wu, Pieter Abbeel, Ilya Sutskever
We interpret meta-reinforcement learning as the problem of learning how to quickly find a good sampling distribution in a new environment.
This interpretation leads to the development of two new meta-reinforcement learning algorithms: E-MAML and E-RL2.
Results are presented on a new environment we call ‘Krazy World’: a difficult high-dimensional gridworld which is designed to highlight the importance of correctly differentiating through sampling distributions in meta-reinforcement learning.
Further results are presented on a set of maze environments.
We show E-MAML and E-RL2 deliver better performance than baseline algorithms on both tasks.