Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation
Vision-Language Navigation is the task of navigating an embodied agent to carry out natural language instructions inside real 3D environments. We propose a novel Reinforced Cross-Modal Matching (RCM) approach that enforces cross-modal grounding both locally…
A picture from a dozen words – A drawing bot for realizing everyday scenes—and even stories
If you were asked to draw a picture of several people in ski gear, standing in the snow, chances are you’d start with an outline of three or four people reasonably positioned in the center…
Data Efficient Reinforcement learning for Autonomous Robots with Simulated and Off-policy Data
Learning from interaction with the environment — trying untested actions, observing successes and failures, and tying effects back to causes — is one of the first capabilities thought of when considering intelligent agents. Reinforcement learning…
AMDIM – Augmented Multiscale Deep InfoMax
AMDIM (Augmented Multiscale Deep InfoMax) is an approach to self-supervised representation learning based on maximizing mutual information between features extracted from multiple views of a shared context.
MetaLWOz
Meta-Learning Wizard-of-Oz (MetaLWOz) is a dataset designed to help develop models capable of predicting user responses in unseen domains.
MineRL Competition 2019
Starting June 1st, we are holding a competition on sample-efficient reinforcement learning using human priors. In our competition, participants develop a system to obtain a diamond in Minecraft using only four days of training time.…