hi everyone and welcome to our tutorial on recent advances in Asia and I wish research I'm JJ Lin from Microsoft and this is a joint effort between Microsoft Facebook and JD thanks for a journey as online at this special time and I hope every why is staying safe and healthy in this tutorial we'll cover a wide range of vision and language research topics including visual captioning video QA and reasoning texture image generation and step surprise learning for multimodal training we divide this tutorial into four sessions to walk you through each of those research areas we
will start with a visual and standing session followed by a coffee break then we'll give two sessions on generation tasks texture image and image to text after a second coffee break our last session is self service learning for multimodal / training and due to time limit we will leave all questions to the end of each session to make sure that we can go through all the competitive content that we have prepared for the audience we will leave five minutes for a QA and in a pitch session will also be available online for questions during coffee
breaks and also after the tutorial ends at 5:00 the sliced and the recorded videos will be available through our tutorial website after the tutorial ends the first part is visual QA and reasoning this session will be chaired by Joe gang from Microsoft who is a senior researcher in our team Jen received his PhD from Duke University and his current research interests include vision and language Petrini and self service learning he's the leading author of the latest villa model which achieves new state of art across many vision and language benchmarks in visual and standing tasks the
model takes both image and text as input signals and learns through multimodal fusion for quest answering inference or common-sense reasoning pablor tasks include vqa a chicle a visual comes as reasoning referring expression and so on and in this session we will give a comprehensive overview of the main research topics in this areas including advanced attention mechanism for multimodal fusion a better image feature presentation multi-step reasoning neuromodulator works and so on so we hear many interesting topics from gia in this session then after a coffee break will be held a second session on visual captioning hosted
by luo wei Joe who is the researcher our team at Microsoft LOI received his PhD from University of Michigan his research interests include computer vision and deep learning and in particular at the intersection between vision and language Lola is the leading author of the PLP model on joint / training for both understanding and generation tasks so he will share some insights on visual captioning in particular today in captioning tasks given an image or a video the model will generate a natural language sentence to describe the visual content either in the image or in the video
clip popular tasks range from image captioning to video captioning and also to more advanced dense video captioning and grounded visual captioning in this session we'll introduce different schools of models for captioning tasks from the classic encoder decoder framework to attention based mark models and also the latest transformer based models and also portray me following this we have another session on generation tasks focusing on text to image synthesis each other is the host of this session who's a senior researcher at Microsoft before joining us he was a research staff member at IBM Research his research focus
is deep learning in general with specific interest in motor compression deep generative model and ever server learning in this session we were introduced both image and video synthesis tasks were given a textual description a model can automatically generate a realistic image or a video clip to visualize the description for texture image will introduce state of our models such as stay again and attend gain and for text image to video synthesis both game based and BAE based models will be discussed such as story gang and t2b will also cover the latest dialogue based image synthesis work
such as check painter all have very interesting applications and he has been working on many of these tasks in the past then after us coffee break the last session is surprised running for a multi-modal / training we dedicate we dedicate one hour to this session and also have three presenters as this is a fast-growing area with many interesting topics leaching is a research scientist at Facebook AI he received his PhD from UNC last year before joining Facebook he was working with us at Microsoft envision and language puccini Lindsay and Ian Jill are both research STS
in our team at Microsoft who have been working on both computer vision and NLP their current research interests include vision and language retaining and self surprise learning all three of them are the leading authors of United model which achieved sweeping success on many vision and language benchmarks we have witnessed wide adoption of self surprise learning in her vision and NLP for Vision Plus language there has been a recent surge in using large scale free data tutoring a universal multimodal embedding with well-designed retaining tasks such as massacre region modeling and image text matching and this bridging
models can be readily applied to there is downstream tasks with some fighting such as VQ a VCR image text retrieval captioning and other tasks and there are two main streams of work in this area image text retaining and video text between there are many pretty models such as the Albert a Unicode a VL uniter and pics of Bert which have achieved new state art on a long list of multimodal tasks contrast to the boom in image flat text area pro training for a video class language is still in its infancy we will cover the latest
work in this field and also discussed the inner makings and behind the scene key discoveries of this region and language models and that's all the content that will be covered in this tutorial she'll do her from JD who's an expert in many of these areas is also partnering with us for this tutorial today he's in China right now and won't be able to join us due to time difference but he asked me to say hi to everyone and hope you will all enjoy a good time with us today so without further ado let's get started
and have some fun Jurgen will kick off the first session