Audio Retrieval with WavText5K and CLAP Training
Lightning Talk: A deep learning approach to recover conditional independence graphs
Presented by Harsh Shrivastava at the Microsoft Research Booth at Neural Information and Processing Systems conference (NeurIPS) 2022. This video covers their research on newer deep learning model uGLAD, whose design is based on deep-unfolding…
SpeechX
Neural Codec Language Model as a Versatile Speech Transformer SpeechX is a versatile speech generation model leveraging audio and text prompts, which can deal with both clean and noisy speech inputs and perform zero-shot TTS…
DragNUWA
DragNUWA is a video generation model that utilizes text, images, and trajectory as three essential control factors to facilitate highly controllable video generation. DragNUWA is a video generation model that utilizes text, images, and trajectory…
Microsoft at KDD 2023: Advancing health at the speed of AI
This content was given as a keynote at the Workshop of Applied Data Science for Healthcare and covered during a tutorial at the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (opens in…