hey awesome let's get started so just now my colleague Yoga has a really comprehensive review on me QA in the Bayesian reasoning so today I will be covering another popular topic for Bayesian language which is visual captioning I'm away from Microsoft so I'm connect up to again and so let's get started so here's the outline of my talk I want to first give you a brief overview on the problem and then lay out the talks on imme of the problem visual captioning then then we will go into our first technical session on image captioning and
also the associated it is as an evaluation metrics and after that we will go from the image domain to the beta domain and look at methods on video description generation and I will mainly focus on two of the Bands topics for video description Generation Y is grounded captioned generation and the other one in stands caption generation the Hydra okay make sure you are recording now I couldn't see that the icon but just yeah so then after the the two topic two topics on video description I've only a few minutes for the Q&A so visual captioning
aims to describe the content of an image or video with the natural language sentence so here I show you example on image and another one on video so for the image of this cat on the example model output should look like this a cat is sitting next to a pine tree and looking up and for this video we should expect to see something like a dog is playing piano with a girl so know that in this tutorial we focus on factual captioning so we want you to see what content is in the image in the
image or video and make sure that the captions can be grounded the words in the caption can be grounded into a pager content as opposed to other related topics such as commenting so there's the applications of visual captioning one and perhaps the most popular one is auto text generation that's actually a auto text generator from whole point so whenever we upload an image to PowerPoint that's the option for us to generate in the display the other tags so in this case PowerPoint generates a cat sitting on top of a grass covered field and so the
reason why we need this is that like there are a lot of people blind people and visually impaired people in the world and we also want them to appreciate the content in the image so once we have the auto text we can use a speech synthesizer to just speak the auto text out and the second application is content-based image retrieval or CV IR so traditionally for image to tags or text to image retrieval we use tags or descriptions of the image as the evidence for indexing but you know sometimes those attacks or descriptions are variable
so we need to figure out a way to use a CAPTCHA model to this generated description by self and finally of course we can use a major caption interest for fun and here's a fun video shot by Kyle McDonald a few years ago and so he literally run and we all had an image capturing model on his laptop and walk around the city and to see the quality of the captions as you can see the captions sometimes are pretty good and especially on common objects now for the taxonomy of better captioning and I want to
divide so look at the this front two aspects so if you think of it from the Tom aspect there are two main categories image in the image domain and in the beta domain and innovated oh man you can further divide it into two sub categories wise with short videos usually of few second long and the other type is with long Beddoes so you can get to see longer videos of a few minutes from probably and on the other side of the graph we can also divide existing methods into three main categories so template based retrieval
based and deep learning based so for the sake of time we were progressing on the most recent deep learning based methods and also are these tutorials focusing more of the advanced like recent topics so we will get to know see a lot of methods on based on deep learning and see a variety of methods from this category all right so the first method is CNN HTM so that's perhaps one of the first deep learning based methods and the property is formulated as the following so we want you so theta represents the parameters in the model
so we want you maximal parameters that can optimize the log probability of the sentence given the image and have summed them up across the whole dataset if we factor Rho 2 this formulation we will see something like this so essentially you will become a sequence the sequence carbon so given the image we first generate the first word and then given the word we generated before and the image we output the next word and so on and that gets us to the first CNN LST and based methods which is called a show and count so in
this framework are given an image we first use the commonest to generate the feature a feature vector and once we have the feature vector we are we feel it through Alastair module which output the probability for the next word so during training we can use some techniques like teacher force in training to make sure we are generally we are applying the cross-entropy laws to the correct words at the output distribution and during inference or we do we can do all sort of sampling like greatest query search or beam search to sample words from the output
distribution and sure Intel is a y example of the encoder/decoder framework so basically you will see a lot of methods for into this category so encoder/decoder framework essentially use a visual encoder to to take in the image or video input and encode them into a media space and then feed this a feature vector to a language decoder which output a description and so that's one limitation with the very first model we have seen today so for for CNN I was TMM work because we are applying a mean protein at the end of the feature map
so we are losing all the spatial information so but you might say okay so generally the sentence we might not only just look at the whole thing we want to like dynamically attend to different regions in the image so that's why later on people proposed to use a software tension for image captioning and some attention is initially proposed for machine translation and if you think about machine translation is also a sequence the sequence pattern so some attention is able to add let me create a turn to input content based on the query so what does
a query mean you may ask so there are three main elements in suffocation query queue he is K and the values B and I will get you a quick example in the next couple of slides so you get a better sense so in our case in all contacts everything is simplified so the keys and the values are usually identical so you can think about them usually come from the same feature representation so in our case for example the great feature of the common net output and the curlicue is determined by the global image feature or
the ASTM hidden states let's get to the example so say given the same cat image with c4 earlier we first used a CNN to compute a great feature for the image so in this case so normally you will see like a 7 by 7 or 14 by 14 output with the channel size of 2,000 save 2048 so here for simple simplicity I draw a 3 by 3 great feature so years of attention this part is regarded as the values and the key and then we now we start a generation we need to find the query
to query on our this memory these values and keys so we do a min pulling over the great feature and get s 0 and then we can perform a soft attention over the keys in value so basically we compute a similarity score between H and s 0 so using the similarity function called F ATT so fitt usually is a dot product or some neuron MLP neural network and something in the most recent work like transformer are people start to use a scaled a top product to avoid gradient issues so once we have the lineman scores
and we apply a soft max to make sure all the attention weighs alignments go sum up to 1 and that's because when we do a weighted sum between the attention ways and the keys and values we want to make sure we don't get any grading issues at this stage and so now we have the evidence for the caption generation we really to say Alice TM like we did in the previous model and then we output the output distribution probability distribution for the next word and we repeat the process until we hit the end token so
to take away from here font is the basic model is at each time step of decoder we use a different context vector that looks at different parts of the image the input image so these are the concept is really important so we will get to use this quite often either in the later content so here is a like some visualization on the motor output so we overlay ad attention weights with the original image so you see some bright regions and dark regions so prior regions correspond to our attention waste with high values and vice-versa so
as you can see while generating the sentence but the model is able to look at different regions in the image there are more examples from this life so say for the in the first case when we generate in dark we we are the models able to stare at a dark and in the second example we mention stop sign the model is able to attend to the bright region and later so on for the third and fourth image so I I talk so that's a basic form for suffocation and amber scenes is proposed there are numerous
like variants of attention proposed and one of the famous one is the regeneration so essentially is a variant of suffocation based on the future input so we talked about using the great feature from calmness as the input and yet region region attention essentially replaced the input with the region proposal features so if you look at the diagram we now replace the CNN backbone with the of the sharp say of the shelf as the oxy NN like of the shaft objective factor then we can get the region proposals and correspond to each region proposal will have
a feature vector alright so region retention is one of the fancy attention proposed of Regency and there are also other really fancy attention mechanisms and in the next few slides about briefly go through some of them so the first one is semantic attention so he adopts the concept called a visual attributes so given input image we run an attribute classifier on top of it so that we can get some visual concepts from the image so in this case we got wave writing man and softball and so on so once we have these visual attributes the
language decoder can use that as evidence to generate output caption and the second model is called adaptive operation so previously we talked about like how the motor is able to turn to different spatial regions in the image but you know sometimes when we turn away some non-object words we do not have to return to the image at all so in this case the language model can totally do the heavy lifting so in this work they proposed a concept called the visual sense and you know so that we just sent you know is able to tell
the model when to and when not to return to the image so the third architecture is called attention on attention so if you look at the left hand side of the diagram so much of a correspond to the classic attention module with query keys and values and the motor they proposed is showing module B and they actually have a attention module at the bottom and the idea here is we want you to do some more intensive fusion between the query and the keys and values so they actually literally use the query multiple times such that
they can use it to gate from the attention output to have a more intensive interaction and the fourth model is called X linear attention so actually in our time so the motor we talked about before those are all perform detentions perform on the channel channel level so the say if you think about dot products you are doing admin wise multiplication according to different channels but how about the interactions between channels right so in this work they propose a to use a bilinear pudding to model the interaction between channels these two matters focus mainly on the
data format so the data input so in hierarchy parsing in this work they use a tree structure to to model the regions from the image and in the second work they use a singer and to model the image content image regions and tags and then transfer the inductive bias learned from the language to image and now going to even more recent works so we talked about self attention and initially proposal for machine translation so actually so that makes a machine translation a natural sauce for people at you and take like to think about some new
ways for image captioning and one recent method is called a transformer and is proposed in a paper you probably heard of attention is all you need from new ribs 2017 and transformer performs the sequence the sequence generation and it got really impressive results on machine translation and so people started excited to image captioning because image caption is also a sequence the sequence tasks so one important components instead of attention is the sorry in transformer is a self attention module so some attention is essentially a self attention model attends to itself and we will get you
an example in the next slide and self attention is a special case of graph neural networks so at this point and I think that my colleague to also briefly mention that and so some attention essentially has a fully connected like as opposed to a sparse structured sunspots architecture in growth neural networks and sometimes we use a self attention to model the relationship between object which ends so similar to what we have seen in the existing works like hierarchy parts hierarchy passing so in that case people use a common task to model relationship so here I
showed you like one one layer transformer model so transform one of the first applications of transformer in major caption is proposed by my 2080 a week at work and so the idea is quite simple so if you look at the diagram and the focus on say the left hand side fold for now so if you look at the self attention layer it takes in keys values and queries so all those three are identical to each other so if we don't account for the the in anybody layer so we can assume that the same and then
we use a so the game at the input we fit in the image features of made of features such that but the self attention layer can encode the relationship between the regions between the video frames and at the output we fit in words like what we what we do for our STM so usually we need to follow the autoregressive formatting so when we are generating the current word we can only look at words on the past but not the future but recently there are some non auto regressive model proposed such that even during inference we
can output all the words Immelt aeneas e at the same time so that's a basic transformer architecture and so there's a more recent works called object relation transformer and mesh memory transformers and they are all based on the basic transformer and network architecture so I will point you to the references I pointed out below so if you're interested on this topic you can go ahead and do some photo radians all right moving on so like really natural extension of from transformer to to the next stage given the revolution going on in the LP field like
all the pre-training stuff going on and a natural extension is to apply pre-training to image captioning so those methods usually have a adopt to stage reading strategy so the first stage is free training and the second stage is fine tuning so for pre training is usually performed on a large data set so I'm talking about like a media's of image and tax payers and those taxes are usually automatically generated and say there's a data set called conceptual captions and in that it is that they take web images and associated author tags of the web image
as the supervision as the annotation so the training objective is usually unsupervised and they act as independent because we want to learn a generic representation in the pre-training stage and for fine-tuning the stage to is more task specific so we apply we take the model we learn from pre training and then find to hit on various downstream tasks with the with usually a supervised objective so all those methods are based on bird a variance of transformer so regarding this part of opera training my colleagues and from Microsoft and Facebook they will talk about more details
and have a more detailed overview so here I want to focus on works related to image captioning or visual captioning so there are two main categories for VIP based methods the first is with the separate encoder decoder so representative methods include radio bird and Oscar and the second category is with the unified encoder decoder and Representative matters is unified unified VLP so the difference is so for the separate encoder decoder only the encoder is pre trend so we need to the Bayesian features and then pass it to the language decoder so these parties usually trend
and in unified encoder/decoder or both modules are pre-trained because we are using a unified architecture so there are some benefits about either either one so the model of a unified we RP is simpler and but only other hand if you have a overall encode like a really general encoder that would benefit more tasks basically so going to the benchmark and data sets so the most popular benchmark for image captioning is cocoa captions and he had contains over 100k images for training and 5000 each for validation and testing and the organizers and the creators of the
data set with hold 40,000 images as the hidden asset so that for leaderboard purposes and the vocabulary of the data set is over 9000 with each word has at least five times occurrence and this data set is and mostly adopted in image captioning and another data set commonly used is freaking 30k and it has a significant lower number of images and the cat vocabulary is also smaller so for the sake of time today we will mainly look at results on Cocoa captions so here are some images from the data set as you can see most
of the objects are quite common like a baseball player and bass and so on moving on the evaluation metrics I want to briefly go through for commonly use metrics blue material cider and spice and so blue is a is a based on an grant based procedure and the meat here further considered the ordering information between the Engram unique web matching and the cider gives more weight age to important engrams to tf-idf so you know there are some words certain words like the like those worries that's coming exist in most of the captions so you don't
want that to have a high weight so you want to lower the weight for those words and finally a spice is the f1 score over a caption a single tuples so it's a more fine-grained metrics that consider different components of the sentence like a bird nouns and so on and if you're interested you can go to these awesome slides from Syosset lecture from University of Toronto and so here I want to show you some quantitative results on cocoa benchmark and before going into the numbers I want to say so the reason why we show the
numbers is because my intuitive but that's of course not only criterion like to justify the performance of a model so if you're interested in knowing tab deep into each of the method I will recommend you to look at each paper for the say the qualitative results visualizations and analysis so getting back to the numbers so for CN the first type of methods we covered CNRS TM because 20 in blue at four so unwell for grad program based bloom and moving on to attention based methods are we we seen a significant improvement so he goes from
20 all the way to 36 and so also here I put a nose here so for region attention for these methods you also use a cider optimization so is the IR based methods and potentially give us a few give a few points boost overall the performance is significantly better than the CNN based method and for the later transpolar based methods we again seen some improvements and especially for Saturn so as you can see Center when all the way from 120 to over 132 and finally the most recent transformation in image captioning we have seen further
improvement so the ladies member on Sider is now at 140 alright so there are also other topics I want to cover but for the sake of 10 we only waited for further radians for example a dance captioning noble object captioning and stylized ivory slide eyes captioning or divers captioning in REO based methods so I link you to a survey below so if you want to get to get to see more references on each of the topic so now moving on to the beta domain and we talked about captioning each individual image so that corresponds to
the middle part so say we sample a few frames from the image and the caption each frame but now so say we have a long video and essentially you can look at a lot of consecutive frames and so now we want to and those paella has a lot of other information like a motion information and we want to capture so now a test becomes how to capture those motion informations how to caption over a period of time so that's correspond to the video captioning problem and here I want you maybe put it in the video
description instead of video captioning because of video captioning also refer to relevant concepts in the speech field so it's used sometimes represent from speech to tax transcript generation so to avoid confusion in this talk I would for now put it like video description so for video description the matter is are quite similar to what we have covered so far so people have used auto encoder decoder and attention and the transformer for the motoring and we also pay special attention so in the bill domain we pay special attention to capture the temporal information the motion information
so to do so there are a few there are various methods proposed so to capture the the local information to motion information we sometimes we use a 3d CNN and also the optical flow and to capture longer temporal information we have tried say using STM using transformer and attention so that we can capture the long-term dependencies and so that was the overview about video descriptions and now I want to move on to like to more advanced topics on video description so for all those matter we have covered so far they focusing on mainly description to
animation so we want to output a caption but sometimes output caption alone is not enough especially in the real world scenario so say given an image like this I would say our model can output a perfect description as a bottle of ketchup and a bottle of sriracha on a table but it doesn't know which is which so if unfortunately our robot chef made a run Association and could sriracha instead of ketchup you know hot dog there'll be a total disaster so we want to avoid a disaster by ground the words in the image correctly so
we want to want the model to have this quality ability and then we have a perfect hot dog so essentially we want to combine visual description and object ground in our detection and in the image domain a matter called neural baby tacos proposed and in the beta domain as method called equality the video description so for the sake of time we won't go into a model but if you want to know more there's a link to the workshop from yesterday's activity in that workshop and so for both methods so imposed image domain and the video
domain we require special data set that has post descriptions and the bounding box annotations so at least for now all the all those methods use a suit for supervision so they need the bounding box as the training signal so for image that's simple so we can just parse the sentence and annotate objects in the show somebody parks on the image but how about video so leaders are more complicated and it contains a lot of frames and so how about and so I want to give you a few examples so for video Dan sanitation would be
really time-consuming and expensive and also we are we care about the connection and we are not doing exclusive like object detection we don't want to annotate every single object in the in the video so our point is more grounding so we want to feel the connection between the two modalities the visual modality and the text modality so that's why we draw different body boxes on the one of the frame so sparsely annotated one of the frame and you might say we might say what if we cannot do it in one friend then in that case
people propose to annotate multiple frames so in this case the PO is invisible in the first frame and that's why I don't had it in a different frame and that's conclude that's a concludes that the grounding of description part and so far all the encoder decoder framework works very well for images and the short video clips but how about lawmakers so we haven't got you haven't talked a lot about long videos and in fact the average video duration on YouTube is over four minutes so how about those photos and in 2016 CVP are you at
all proposed a method called a meadow paragraph description so now given a phone radio we generate an entire paragraph to describe it which might contain multiple sentences so but that's a dozen big improvements but that's not limitations so first the readability of the paragraph is low so is a big text trunk and second we the association between the two modalities are not that strong so say we want to know how to chop bacon we want to see we also want to see the beta demonstration to localize that clip in the video accordingly to have this
kind of temporal branding and that's why there's a problem proposed called s beta description and got a lot of tension recency and so we have to talk about the left-hand part in this diagram so now given a long video we want to first identify a few events in the video and then I'll put a caption pretty beds and the impalpable beta and alpha will be triplets of eben start time and ham and description so the answer existing methods on dance video captioning or description so the usual content - Mojo's events proposal and the video description
so events proposal output the candidates each candidates have a start time stamp and time stamp and the competence code and those methods fall into three main categories so the first one is with a separate training so we want to make sure we got a past events proposal and then feed those events to caption decoder for caption generation and the second category in the middle are using the alternating training so now we sure the video encoder between the two modules such that we can have a more generic representation and there are limitations with those two methods
the language information cannot derive impact on the events proposals so if you think about it if you know a step there's a caption about add bacon to the pen there are bacon independer in innovator so that will help you to better locate that events in the video that's why later on not end-to-end based methods is proposed so there's a differential link between the two modules and for the sake of time I will not go into details for the model and if you are interested in more details I show so there are more recent methods come
from yesterday's activity in that workshop so I have put a link below so in case you're interested and to conclude we have seen a really aggressive progress in the field so own Coco captions for example we we see the cider goes from lower than hundred all the way to 140 and so one thing I want to point out is that like the motivation is important so we want to know why we made a trench and what does each change contribute to the outcome to the contribute to the motor performance and although that phrase here would
probably not those keep doing that work Legos and at the end of the the talk I talked more about grounding so to achieve a better result interpret interpret reality we need to have grounding and towards a better generalizable and robust moto opportunity as well option so we can in this we will have a full session on self supervise money at the end of the tutorial and that will provide more in size and all regarding this point so there are some limitations about visual captioning so there's a still known way before production ready due to the
following issues so first its recognition failure and also object for resignation and then model bias so regarding the first one we need to figure out to get a better feature and for object orientation we aim to have better branding and detection approach and for model biases there are some recent work in bqa that use a bracero training to eliminate biases and also there are some works in captioning as well so that's a really important topic and all three all three if we can solve all the three problems they'll be really promising towards production ready in
the next level of image captioning video captioning so for future directions so so far people have a lot of work proposed to improve the evaluation metrics and make them a career correlate better with human judgments but still dancer gap between automatic metrics and the human evaluation that's why in the video domain in particular a lot of papers you use a human judgement human evaluation to get the most accurate feedback so that's a one interest in direction to go so how to propose an evaluation metrics that can correlate better with the human judgement and another point
also to also mention in the BQE talk so we want to so people started revisit previous models like features and which surprisingly give even better results so we want you the benefit here is that we can simplify the model pipeline we don't need to have a really heavy model like a data pre-processing at the beginning so we can just let everything run end to end that will we see like shorthand like a product timeline chemical and lastly about beach VLP so an interesting thing to think about is to how to close the gap between the
pre-training domain and the downstream data domain so so for peer training we use a web images and photon stream tasks we use cocoa bqa and those more curated data sets so and also there are some domain gap between web images say if you want to apply to cooking images that will that will normally fail so you we need to figure out like so how to cross the gap and how to make sure the pro training we can optimize like the training knowledge we learned from free training and transfer to town stream tasks and that's all
for my tutorial and thank you all so much for your attention and I will now take a few questions