uh good afternoon everyone my name is yuu chen from microsoft thanks my colleagues joe and laura introducing the vqa and the captioning set this session i will introduce how to use text to generate image so this topic become more and more popular in recent years and you can see topics reverse the task of lowest tasks my second word lasts 30 minutes from 310 to 340 including the coin sections so as jj mentioned after the day we will upload the slides and the recorded videos to our like external websites we divide the uh topics into three
parts uh text-to-image sentences text to radio senses and also dialog based lessons and in each sections we'll briefly introduce the tasks the data sets and also the evaluations and they also introduce some of the state of the module for each method hopefully after this tutorial sections you can have a like a brief idea or general idea about about these topics before that uh one of the fundamental technology using in textual image or just image generations is a generated modeling so far we have seen many uh like great great models which you can generally release the
data from a traditional like virtual machine to some autograph mode like pixel iron or cn to the vibration auto encoder way and also generate adversary network gas let's first thing to gains gans has two components the generator and the discriminator playing a mean max game the generator is to generate the image and the disclaimer is to measure the divergence between the giant data and the real data the generator is optimized to minimize divergence also we have we also have two components the encoder and the decoder forming a auto encoder like framework different from the linear
auto encoder the variational code encoder have a poster prior on the latent space and also use that to recognize the training to get a better performance the whole way is to the my evidence of lower bounds as we know for image generations way and again have achieved a greater success for example the style gun paper last year sure gans can generally release the photograph image also we have the vehicle where you work it can show the same similar performance to phase generations the work we used today for text-to-image generation have been built upon those models
either then of a besides the unconditional generations also say many like a famous work for conditional generation like the cycle gun and the pixel p took pixel game work uh for image 2 image translations last year we have seen this beta nominations further boost the energy generation quality to a higher levels also its famous disco gang which can bring can close the gap between the image two image generations in default domain that other works for condition unconditional cognition image generation using different resource like the image generating from scene graph the image generation from audio and
speech in 2019 and also recently people are more interested in doing image generation from like object layout or silent object out after seeing these examples a natural question is can we do energy sentences or generation from task tags description the answer is yes and we can do it using a very simple conditional gain or value framework how is a simple framework using condition again the input of the condition again is the text situation and the real image first we encode the text distribution into a hidden vectors then we can kind of hidden vector and the
random noise into a vector and put into the generator the generator use that to generate the image and also the discriminator is do the same thing to measure the difference between the real image and the generator image by doing this we can achieve this task but beyond that we haven't seen many many like great work in this era to uh for text to image generation like 2017 staggering 2018 attend gang and target 2019 object gang and amirogan also this year ceviche are many again here is the example for of image generators using conditional gain the
model takes description red flowers with black centers and a general with a flare lock image with some like a black ratings in the middle you can see the direction quality is okay but if you see other data set or generated examples the the course is not good right you can see some blur regions in in some of the image also as i mentioned uh conditionally another powerful tools for the text to image sentences this work tried to generate the text image from the text attributes with the conditional value framework you can say on the face
and the birds they said they can do very good job to get a better uh generation quality we have seen this background with proposed the k idea of staggering is to do two stage generations and generate the image progressively in the first stage you try to just generate some like a roughly structure of the image with lower details and then in the second stage it's work just more photorealistic image with how details in both stage the tax inputs were they will take the same tax inputs as the conditioning let's give more details about the staggering
and the first beginning we will have this uh image first the thumb dancing to some smaller language and in stage one it generates images from the text inputs and the disclaimers later also match the difference between the gentle image and the downside image in the saturn stage in the first fishing the feature map of the first generating the image and the the text description into one vector and using that to uh to general image also the the screeners measure the discrepancy between the real original image and the second image the generated image here we show
some examples general from again you can say the first slide is the uh generated examples from the stage one uh it roughly gave you some like uh like the size and the colors of of the birds but it looks very blur but in the second stage you can say you will get a more realistic image with some details like the color change in the in the heads or different stuff steering uh in some region it's not very clear like like this one or this one still looks blue uh also uh stuck i have the second
version called stagger number two use more generators than the disk printers due to this time limitation uh we were not interested uh introduced it here you are interested you can look at the paper i've just got another a good job called attention gun attention is also proposed in 2018 by my colleagues in microsoft the key idea of attention guys is pay attention to the random words in the description and the region in image by doing that a 10 game can capture both global sentence level alignments with the image and also the fine grained water level
information from a text here is the framework of the attack gun we can say still adopt the stag and to multiple stage generation trick but different for that in each step it's not just use the uh text description adding the wrong noise but you also have some attention work between the text the generator regions and the word features after this attention step you will have the attention function map containing the round noise for the next stage generations also you have another very important component called the semantic similarity matching modules which matches the uh the similarity
between the final generated image and the word description this module is very important actually it's play regularizer to guarantee a tan gang can can generate a very good image here we show some examples for the tangent first we show the attention map you can say uh when we breaking the relations between the words and the image regions by the attention you can see each word will align with some ranging of the images for the generations that is what the the k points of a touching gang and also we can see different from the stag game
a 10 gang can also gently like some some objects with with more details like in these examples the sun has mentioned something like pasta or berkeley and in the final images you can really say pasta and the berkeleys also in this uh second example showing the similar stuff it's it's uh mentioned bananas and winkies and you can still see similar stuff this slides to show more examples when compiled attention game with either framework like staggen scandal version two and we say in the both days that this uh cob is the birthday set i just shoot
and this cocoa data is uh cocoa image captions themselves we can say measured by fid and the inception code i can gain all outperforms other baselines and also you can see english language and the birth data sets actually the generation quality is is near perfect but turns out the more complicated the cocoa data sets the general the generation quality is not that good it's roughly have some objects or backgrounds there but not not very detailed following following the 10 game framework lastly we have this mew again uh it's roughly take over the touchdown framework by
the added another pass instead of just having the text image similar semantical similarity marching is also generated as a text from an image and measures the uh seminars between the two tags follow a closed loop between text to image and image tags by doing so mirror can can show a slightly better performance than a tongue gun okay let's have first stop here to summarize the text-to-image sentences framework we have talking about the dataset used in this domain like the cob the bird's dataset and the first status i showed the flower datasets and also the cocoa
captioning data set and we can see the kernel master the falling stag and or a tan gang can do very good job on simple data like the birds and the flowers but not on complicated ones like a cocoa also we see the evaluations we saw this almost the same following the unconditional gain of aig generations like using inception goal or fid scores also sometimes with zoids uh how to do human evaluations and talking about the tanking challenge in installment i think the first is that can we find a better like evaluation matrix and stand up
for the traditional exception score if id or human emergence and the second one is i look at all the current gains i think it's try to learn the discrete mappings between different attributes to the image like the color the objects of the objects and the size and some other attributes but when we think about the more uh larger density with a larger vocabulary like ms coco this training will become very difficult and the standard third chinese come from the mating object generations as we can see the reason we can do very good job on the
birth of flower data is because they almost contain y objects and a simple background but how can we handle some dataset that is cocoa with magical objects and how we how can we model the relations uh i think objection is first proposed to to solve this chinese especially modeling the multiple objects uh instead of do one step generations to match the the the real image and fake image the object i have another loss try to match the object layout through the rear layout and by doing so you can see the object again can do a
better job than both stack and tangent i do some like object generations also another work in the same year proposed called object pathway it's not just handled different object but also it's modeling some implicit relations between different objects using a formal called object pathway by integrating this object pathway you can see the the object pathway gain or vg can generate very good examples like for the first uh examples the description said the two regularized slice pizza with metal toppings actually uh this is the rare image and the second one is generated by the the object
password again and series as by the time game you can see actually the object has a weekend can generate this the two actually pieces and some tokens which is very good then we proposed another slightly different task called the image emulation manipulations different from the text to image senses which is quite straightforward this one giving origin image like these two birds and giving a text description to describe the ideal image with the original image then the generator will generally image it should reflect the difference between the orange image and the generator image by this sentence
with we saw that again is the first one uh for this task and actually did quite a good job and this year we see this manning gang he have two components uh first is called the uh availing combination models which give you some like roughly generations for the whole image and also you have the detail crack cracking models which just change some portion of the image to get the final go by doing so we can the minigame can slightly get better performance than pagan okay after the first section the text to end to image sentences
we were into the second task called the text of radio sensor synthesis uh so the input is still a text description like a plain golf and grass but the output will be a show the radio clip contain multiple frames uh to first we need to the contents between the gentle videos and the sentence be closer and also we need the consistent between different generic frames that's the goal of this task the t3 framework was the first one to to to a proposed to for this task um i think it's it's roughly simple just using a
conditional value framework with some like a traditional cell encoders iron encoders encoded the whole sentence to ensure that the concepts consistent between d frames is adopt a component called gist this you can you can is the is the is the it's the image layout designer to keep your uh image of general frames to be consistent and after all it's not much of the similarity between frame and frame is actually measure the similarity between the whole frames and the real frames here is some example janitor by t2 way we can say roughly it can ensure the
generated content which is consistent with the description and also we can say how some coherent between different generated frames but still the junction quality is not that good you still have some blur image then in 2018 the paper called the tf gang including conditional for tax radio sentences using not vae but again conditional game framework to do such tasks i think one thing's too many things is it's also use multiple game framework but using this convolution factory generations to uh to further boost the performance and we can see that there's some comparison between the t2
way way framework and the tf game framework by using tm tf game we can get the better generations for like uh for each frames also keep the consistency between the textile green and the coherent button different frame after talking the the original vennina text to image to radio sentences we come to the new task called the story generations this is slightly different from the past to text to radio synthesis the task is first give a paragraph of sentence not just one sentence and try to generate different frames based on each sentence it forms one two
one mapping one sentence to one frame but all you should keep some coherence between these uh general frames to form a stories also you should have these whole general frames to be consistent with the whole paragraph that's an even harder uh task than the avenina attacks to radio synthesis the foreign looks like this first for each single frame generations were embedding the noise epsilon with their text description also to keep some consistently we will use the the so-called keyst components and also to ensure the global consistent between the whole frames we're embedding the whole stories
with the encoder and embedding the whole story in each uh each frame before going to the generator to measure the loss we have two discriminators one is called the frame level condition discriminator just match the uh discrimination uh the the difference between the rear frame and the generator frame and then we have a global like a storage discriminator which you can kind of all the frame together with some sense and measure the difference between the whole video frames this is a brand new task so we proposed some like a new data set for such a
task the first one is uh we cloud floor clever uh we use the cloud simulator to change one attributes each time like the location the size and the colors of different objects and then we form something like the clever sequentially collaborations you will see some difference between slides and you can see our model can do better like image generations compiled with a stagger this is showing this example shows some like a coherence for uh of the generated images like we model can change uh can make coherence generations between frame by frame also we have the
uh parallel data set which is more release the data set using the cartoon image and also some paragraph you can say uh story can also achieve some very good generations giving these paragraphs but the one problem is that we the store can have some difficulty to gender very good person or like object details like the the cartoon characters in edge frames but for some others like common used objects i think stock gun actually did quite a good job okay let's went to the last session about the dialog based image sentences the motivation of this task
is we have seen the task order image manipulations which use the language to guide the image generations and all can say some like download based imaging travel basically in each download turns you will get a feedback for the users and they will you will retrieve one or like top image based on the feedback so the question is uh commit to similar stuff like the download based image but using more instead of doing doing trivial we're doing generations i think it's uh it's a it's a very interesting task so people have done many work for the
benchmark effort as the chat clouds datasets they tried to manipulating the size of different objects but it's very simple objects in the image and collecting this kind of data set and also the neural printers they try to do some like image editing using multiple multi-turn dialog but they have difficulty to to get the data set so as a result a finally they try to turn this uh this birds or flower data set for the training and they form this like a sequentially gan framework but they have some trick to adding some like a random noise
for random population each stage to make this more robust to sequence generations now also we have seen this uh framework from mila they form a new data center called the chat painter for multi-ton dial generations the motivation here is slightly different than the volume we call we we mentioned before for the dowel based generations is this exactly if you're just a general image based on one captions or one sentence like those captionings it will roughly give you like a like some very like some some similar images but with different styles so in order to get
more information for the generations you need to ask more detailed questions like not this is not just the birthing in the in the sky but there's also have other stuff like the buildings and different backgrounds and the color of the sky so after you're collecting all the feedbacks of where the download turns you can generate based on those dialogues which will give you a different generation results for the two in the master part i think they just use the traditionally iron or stm stuff to encode all the dialogues for the final generations also we so
last year the facebook and the georgia tech have this cultural data set particularly they form a two-player game uh in uh goal driven collaborative task one is called the the tailor the tether would say the ground choose image and tell the difference tell the journal to draw different objects in the cartoon data set while the journal work receives these informations and put different objects based on the feedback of the tailor they have they formed this platform and collect the dataset also they use some uh very very regular uh motion uh deep learning blocks to to
model the drawer and the tailor separately to get some uh to to get some some model done we'll also propose another framework called the sequential attack guide as i mentioned the download collection for the image is very difficult instead of try to uh collect the data set in a diagonal way we first uh divided the image sequence into different time and then we only asked the human annotator to pride the difference between two consequent images then we come all this uh summarize all the sequence to a dialog form and using this sequential attempt gain from
work to do the modeling uh this sequential attack model is a extension of the time gang uh the only difference is in each tab will have different dial input and to do the generations also we will measure the global global difference between using the discriminator instead of one single image and also use this semantic similarity matching to guaranteeing the training framework getting that well and you can say we compose with different methods with our secondary alternate from 10 and 10 gang and also staggering you can say our second attempt gang can get some better generations
when consider the consistent with the text description and the coherence of different generated images okay i said i have placed all the methods regarding text image attacks to radio and download video sentences let's put one slide for to summarize the text of download to radio sensors i think this is very interesting domain and tasks but we have seen like many studies recent recent years to pay more attention to the proper definitions the benchmark efforts and also with some we have seen some like modeling results but still this is a long way to to make this
a problem is it's not even to production but to accomplish the research problems the first is uh all this is somehow based on some dialogue data or some video data we first need more effort to do more benchmark efforts because as we know clocking the real dialogue is very difficult and also i think that in most of the uh paper i have talked about they also use the inception go on fide or some human evaluations we still need some better way to do all the evaluations also in the modeling side we still need to consider
several modeling techniques like how to do consistent generations how to describe the learnings to know the difference between the words and the the mentioning different regions and also sometimes these some compositional generations i think i have finished all the content let's go to the current section thank you so much hi you you got three questions uh from the audience right now the first one is what are some promising directions