Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Coverage as a Principle for Discovering Transferable Behavior in Reinforcement Learning

Feb 24, 2021

Víctor Campos, Pablo Sprechmann, Steven Hansen, Andre Barreto, Steven Kapturowski, Alex Vitvitskyi, Adrià Puigdomènech Badia, Charles Blundell

Figure 1 for Coverage as a Principle for Discovering Transferable Behavior in Reinforcement Learning

Figure 2 for Coverage as a Principle for Discovering Transferable Behavior in Reinforcement Learning

Figure 3 for Coverage as a Principle for Discovering Transferable Behavior in Reinforcement Learning

Figure 4 for Coverage as a Principle for Discovering Transferable Behavior in Reinforcement Learning

Share this with someone who'll enjoy it:

Abstract:Designing agents that acquire knowledge autonomously and use it to solve new tasks efficiently is an important challenge in reinforcement learning, and unsupervised learning provides a useful paradigm for autonomous acquisition of task-agnostic knowledge. In supervised settings, representations discovered through unsupervised pre-training offer important benefits when transferred to downstream tasks. Given the nature of the reinforcement learning problem, we argue that representation alone is not enough for efficient transfer in challenging domains and explore how to transfer knowledge through behavior. The behavior of pre-trained policies may be used for solving the task at hand (exploitation), as well as for collecting useful data to solve the problem (exploration). We argue that policies pre-trained to maximize coverage will produce behavior that is useful for both strategies. When using these policies for both exploitation and exploration, our agents discover better solutions. The largest gains are generally observed in domains requiring structured exploration, including settings where the behavior of the pre-trained policies is misaligned with the downstream task.

View paper on

OpenReview

Share this with someone who'll enjoy it:

Title:Coverage as a Principle for Discovering Transferable Behavior in Reinforcement Learning

Paper and Code