Pushing the frontier of neural text to speech
In the popular field of text to speech, the goal is to transform the written or printed word into speech that is natural and intelligible. Today, the technology is being used in products and services…
Innovating to understand & act | Episode 1
How can we begin to understand the complexity of the pandemic at a societal, personal, and molecular level? We explore the importance of data, biology, and technology in understanding the spread of the disease, its…
Anticipate, absorb, and adapt—introducing the societal resilience research agenda
The genetic sequence for COVID-19 was first published in January 2020. Before the end of the year, new vaccines—which typically take five to ten years to develop—were approved for emergency use in multiple nations. This…
Pushing the frontier of neural text to speech webinar
In this webinar, Senior Researcher Xu Tan will address the challenges with advancing neural TTS, specifically the high computational cost and slow inference speed in online serving; word skipping and repeating issues, poor voice quality,…
Avbert
This repository contains the code and models for our ICLR 2021 paper: Parameter Efficient Multimodal Transformers for Video Representation Learning
Microsoft and NVIDIA introduce parameter-efficient multimodal transformers for video representation learning
Understanding video is one of the most challenging problems in AI, and an important underlying requirement is learning multimodal representations that capture information about objects, actions, sounds, and their long-range statistical dependencies from audio-visual signals. Recently,…