then let's get a study so I will talk about we show question density and also the we share is a new topic the goal of this part of the tutorial is we want to use weekly way and we show reasoning as the most popular and also various example tasks us to understand the region and the language representation learning so I hope after this talk everyone can confidently say yeah I know we q8 and we show reasoning pretty well now so if you are not in this field so I hope so all you just are beginning
to now to work on this field so after this talk I hope I can give you a general picture of so what is so let me see okay sorry um so I will so hope you will be confidently say I know week you and wish of reasoning pretty well if you are already working in this space I hope through this talk you will also have a systematic review of the message developer in this field and they can inspire you for some new ideas um I will focus on high level intuitions rather than very technical details
because I think for the technical details you can always refer to the papers um I will focus on static images instead of videos I will focus on a selective set of papers rather than a comprehensive literature review because as we know there are tons of paper on this topic this is the agenda for my talk I will first talk about what are the main factors that are driving progress in Milky Way and wish of reasoning then I will talk about what are the state-of-the-art approaches and the key motor design principles underlying this message then I
will give the takeaway message summary and the pin point what are the core challenges and the future directions from my perspective so last careful so so let's so keV studied and I will first before I introduce the huskers what is we press in this year I clear 2020 professor debris in public she gave king of the talk about AI systems that can see in the talk so we pass our research is basically about how we can change a smart AI system that can see in the talk this is from Professor Laurence cake theory in in
his cake we can say the foundation is the supplies and self suffice to learning then comes to the surprise Delanie under the cherry on top of this cake is reinforcement than any how about the cake in our way plus our contacts from my understanding and my perspective I think the foundation is the racial understanding part as we know when we do we process our research what we typically do is we use approach and our system so for example the popular ResNet and the follow you more the ones seein architectures to extract the salient features after
the serum filters is done is a more button NLP task for example we need other ones the language understanding how we can better fill the information in the image modality together with the information in the language modality and in the cherry I think this is our and go multimodal intelligence so how we can ensure all this progress this is where how many in larger scale are notated the datasets to drop the progress in this field indeed in the past few years as we can see there are a lot of popular datasets in this field beginning
from the very famous popular weekly data set goes to for example we show common-sense reasoning and the GQ a data set now I will give a brief introduction of each of the data sets how odd also what all these tasker's are about if we are working in the wiki West base you must be very familiar with this picture given an image and an input the question what is the Mustang made of and an AR system or machine learning model news to predict the answer is bananas then later on the researchers find not only we can
use the image content to predict the answer if we have a difficult actual based dialogue history this can help us predict the answer better and after that people find out hmm in the original week you a dataset there's a lot of language buys which means even if you - nah stay the image as long as I can see the question then I can predict the answer pretty well then the original author of the wiki where paper they introduced the second version of the wicked descent which is also widely studied in so nowadays for example it
has been used in the wiki challenges the basic idea here is even if you have the same question but then given different we show you input your answer can be different take laughter top images as one example in the lab who is wearing glasses on the image on the lab that the answer is member on the right the answer is women so you cannot only rely on your language bias to predict the answer however this is not enough people find the steerer language bias is a very serious issue so they create it and another data
step is called a weekly waste ap our CP is a shot for a change in prior basically now in your training and the testing data set your prior for your answer for answer candidates are totally different so for example in the training dataset when you ask what color is the talk most likely your answer will be white but in your test asset most likely so for example in this case the answer should be black so this has been challenging there is that that the people has been working on to test how how truly your developer
the model is understanding the language side note that the popular.we QA deficits are collected through the EMT Terkel who are not blind users on the other side for the with weak data set they contain with your questions collected from the real blind user um there's also recent data sent on we caught and we are for real so this is data that is used to test the composition of reasoning capability of with your models the basic idea in this task is given a statement for example in this case the statement is the left image contains two
eyes the number of dogs as the right image and at least two dots in total are standing a model needs to identify whether your input pair of images mmm I are true or not compared with the statement there's some more interesting data sets coming along for example one of the tasks Asilomar Chester wish or common-sense reasoning tasks in this task you not only need to understand the factual conference inside your image but more importantly in many cases you need to understand the social dynamics Hinda inside the image for example you need to perform kind of
the so-called common sense reasoning to you as a it for human given this example images when you ask a human wise person for pointing and person one as a human you can very confidently and immediately say oh because he is eternally in person three that person one ordered the pancake which is the answer but for AI system this may not be that obvious furthermore in this data set not owning QA they also need you to perform predicting the rationale why your model choose to to predict this answer so they also have served their little body
in their website recently there's also data sense for example we show enjoyment in the traditional natural language inference task you need to are given natural language promise you need to decide whether your hypothesis is containment on a contradiction of the promise but you know in there we show Unterman the task the impetus the promise the in for the promise is the image rather than the natural language sentence still weekly is very hot and also the developer of the system are very brittle um the researchers find you've forced a of our models you live for your
input of question you just do some grief reasons basically your answer should be the same but the model doesn't think so they will predict the wrong answers so they introduce some other data sense for example the weekly references dataset qqa is another interesting data set previously before to QA when people to compositional reasoning they rely on the so called collaborators which I will introduce later but then collaborative a cell is a synthetic dataset and for do curious tries to bridge the gap between the traditional ritual and the synthetic a clever so the questions in GQ
I usually is highly compositional which means your question is not a very short question but rather it has some compositional structure inside for example given this input the image and question could be what color is the food on the red object laughter of the smoker there is holding a hamburger yellow up long Indian as a human we may not ask a question like this but in order to test a weak ruin Otto's ability to truly can do a logical inference or truly understand the question Jiki a data set was created they also provided the sink
graph for each image weekly should should not only in just depends on your wish or input the textual conference inside the wiki inserting images also very important in terms of this the researchers in Facebook they have provided a dataset to benchmark visual reasoning based on text in the images instead okay we QA is the data set collected about AI to not only you need the visual contents but further you also need outside of knowledge to help you answer the question for example only given the input energy a model not know the tally Bell can be
associated with an American president but by using outside the knowledge from for example the Wikipedia it is possible to do to answer this question correctly for the syntax weekly way is a similar task to the tax which way not only by this task essentially in this year's a weekly our 2020 there are more tasker's are being provided so I think this is a good thing this means we're always trying to provide different data sets to job the progress of the web has our research to truly test a models different different the perspective of the reasoning
capabilities so there are many dennis has been in development but i want you to talk about in another diagnostic data set which is kind of famous or popular in this field for the collaborator said the phone name is compositional language and elementary visual reasoning as you can see the input image they are synthetic and the question they are very long so i think you an for human when you just say the question you cannot answer the question immediately and the u're need to do some logical inference to look at the question more carefully in order
to job the answer besides clever this setting has been sended to different two different aspects for example the clever the cramp adenosine can be extended to we showed our long to referring expressions or even to video reasoning beyond week away another important the task in this field adding is the wish of branding tasks one famous task is called referring expression comprehension basically the idea is given a very short sentence or just a phrase you need to pinpoint was the branches image region corresponding to your included phrase and some other data sense for example Flickr so
decay entities they serve the same purpose based given an input caption you needed to identify four different parts of the caption so what's the branches image region that first showed a corresponding to this year's EVP and another interesting Dennis Escada frisket we not only in interested in predicting the body boxes for the mall we want to do unity segmentation map just now I have introduced there's so many data sets so I think this is also an evidence that this field is so developing very quickly but in this talk I will more focused on one specific
task which is called a week away so in the week us-based as we can see from the last a year week away and the dialogue workshop every time when people think the performance oh it's so high we cannot improve it further but always there's the price people come with better solutions by the models to further boost the performance if you attend the week you and the dialogue workshop yesterday in this year the performance were the further suppose did people find the different methods for example great features structure we Bert PG and villa to further boost
at the performance so the question is can we has a can we have a systematic review of all these methods and what other state-of-the-art approaches and the key model design principles underlying in this message I will talk about this so right now first how typical system also looks like still given this image which is ago holding a hamburger if we ask a question like this so there are several steps for a weekly weigh system to work you have an image feature instruction to get the image a feature from the employee image you also have a
custom encoding module then after this you will do some kind of multimodal fusion to fills the information together then you goes to the answer prediction module we cure a typical is formulated as a classification problem so usually people use cross-entropy or binary cross-entropy to solve the task for the pastrini encoding part um in the very beginning people use STM as an LP field is developing people now use transformer and then there are so many different masters developers in this field and I'm going to category all these masses into several categories um first I will talk
about how we can prepare batter individuals then I will talk about enhanced multimodal fusion methods people have been using bellini bilinear protein masses to fuse two vectors into one so at the very beginning people are interested in in multimodal Kleinman's basically this is a performing cross model attention later on and they also which is kind of the the reason the popular direction people find that the incorporation of object relations are very important is corresponding to intermodal tension in the context of transformer this is called a self attention in the context of graphing new neural network
this is a cartographic attention I will talk about multi-step reasoning the another line of work is called a neural model network I will talk about neural model networks for compositional reasoning later on I will very briefly mention nobles to be queuing and the multimodal progeny so let's first get into the first the first apart at the very beginning people find the grid features are very useful they can get the breed features from convolutional network maybe from we do G before ResNet becomes popular then people find the region feature are very important this is the famous
cotton opt Optim model then in this year they repeal the researchers in facebook and the researchers in microsoft think oh maybe the great feature are strong enough so hmm the attention mechanism was the first proposal for machine translation and later on I introduced for image capturing in the way plus our space you know first the paper shop show attend and the tail the basic idea is instead of representing an image as a single image in feature vector um the images are represented as a set of facial maps and then when generating the next award a
model will need to perform attention to attend to two different the parts of the image regions so that we know what to generate for the next step and that this idea was introduced in the seminal paper which is called a stack to the attention network so I will discuss this for the month but basically you still perform attention over the grid features so the grid feature is distancing look like this however as you can see the green the features they are they are uniformly equally sized in the regions which may not contain enough or sufficient
semantic information then the researchers in the paper of beauty they wanted the bottom-up top-down attention and advocate the use of attention on the level of objects and other Cillian the image regions and as you can see even just the user very simple by lineae are piling bilinear body masses just a simple element wise dot block dot product but due to the use of the moisture in the future they have been 2017 we create energy so this after the introduction of the beauty feature it has been the standard practice for order follow up we process our
model as you can see from the citation of this paper in this year sri PM people find the grid feature also has is advantages indeed when we are used features from the object detectors we usually for for us to do research we're usually first of line cached all these features and then it seems the with your model can perform pretty are faster and also performance are very good but if you think about it if you want to use your system for real-life usage then given a row in which the union to perform this object detection
step from the scratch in this case your interest time is very very long on the other hand if you only use great features the inference time can be very fast and in their paper they show if you have an insufficient number a large number of reads as input for which you a model the performance are actually very good this is an account a per called pixel part before a multi-modal so in all the previous multimodal training methods people use be OTT in the future as your input for the transformer model and in these pics of
paper they also have the same idea we just use a signal which are encoded we don't need to use the wtd features and on the weekly way data set they show as long as you have a very strong CN backbone then your performance can be very high so I think this is a very interesting observation so after the individual cut I'm going to briefly talk about the Pellini reporting methods instead of simple concatenation we can use monster by linear pauline and this by linear polling can be used together with attention mechanism so I think the
most well-known method for bi-linear pudding is the math model compact of a linear podium the reason why is famous is because these masses has been the 2016 weekly challenge you need to perform Fourier transformation then you need to perform inverse Fourier transformation you know to fuse the few actors later on in math mode all around the body nepali is a paper at korea 2017 and the idea is you first map the x and y which is huge feature question feature into a different embedded space tensor matrix product to get the first feature and then in
the multi model factorized bulimia poly messes they have another an additional step which is called a so called the sample in step two first the future vector if I remember correctly this method has been the runner-up before for 2017 weaker a challenge so as you can say this is all about how you can fuse to rupture into one and the people have developer the something like not model fusion and the pad India is simply a confusion among how this works there's a ninja so from my memory and it's in another very interesting model which is
the film the reason why this is interesting is because as I mentioned all these previous by linearly by linear polymer such is all about late fusion you first get an image of feature through a CNN can you perform fusion of the image feature and the question future and the in the film masses which is short for visualized linear modulation in this message they perform early fusion so basically as you can see after we encode and question into a vector and this vector will interact with the with the with the serum blocks so this feature vector
will be deeply incorporated in the cytosine and instead of the serum backbone as you can see maybe you are already familiar with bench normalization and this is something very similar to a conditional personalization and that this message has been so achieved a very good performance for the collaborative asset okay let's move on to the to the attention part i will first talk about the cross cross model attention i really there's a tons of work in this area in the early work people just as a question to attend to image grades or regions later on people
find them we should perform court vision we should perform image tax cohesion and here this is only some example papers that i found interesting and if you really are interested in the cross model attention you can dig into the literature for further details in the same model started attention network um after encoding the question into one future water this feature vector will be used the image which the image region features multiple times rather than only one hand this is something called a multi-step reasoning at all so as you can see in the first step he
learned the attention map may not be very sharp but then after you perform the second attention the attention map becomes sharper and more accurate code engine the idea is not only we need to use the question to attend to the image but the further we also need to use the image feature to attend the back to the question words this is a lock the default approach for the current the popular researches and I think one of the most early work in this direction is the hierarchical contention model so as you can see we needed to
attend in both sides in order to face this feature into one there are some other interesting tasks other interesting models such as peer attention Network and the dance contention network and the one model that I want to briefly mention is the bulimia attention network so still we need to perform the very dense core attention operations in order to achieve a good performance and in this case they proposed we first to learn by linear attention map and then I used this by linear attention map to first the image feature and the question feature together however in
order to truly boost the performance the authors also proposed some other methods for example first you needs to use mode problems in the so if we say this in the language of transformer this means you have a multiple hands attention we also perform they also perform gradual learning in to boost the performance they also incorporated the counter module the condom model is in another paper from a Korean which he teaches and week on a weekly model how to count they also used Club imbalance but nowadays we more likely to use broad model to learn come
touch contextualize the body balance and by combining Audis they are the runner-up solution for 2018 with your HNH relational reasoning becomes very popular now and it basically performs the intermodal attention um the fundamental idea is we should consider a unit we should consider image as a graph instead of a set of image features and after you consider this image as a graph then you can introduce a graph convolutional network or graph attention network to learn pattern we find the image feature or you can use self attention using transform for these purposes and actually so so
you one so human block so as you can find in the so in the website people show craft attention network and the transformer they are fundamentally connected and in some case they are equivalent and there are some similar working in this lab work and I will introduce some of them so I think one of the very first the work in this line is so-called graph structure the representation for which way um although they work on some synthetic data set but the idea was it introduced a really very early um you you you consider the image
as a graph you also consider your in for the question as a graph then the problem becomes a graph fusion a problem you need to infuse the information in these two graphs into one but the the really very famous well-known work is the relation network so they they study this problem in the context of a clever so they find as I just introduced in the gravity does that some of the questions they are very complicated but then they are essentially relational questions you need to understand the relational relational the relationship between different objects so in
the relational network is basically something like this through a convolutional network you get a set of individual image features then we can say the the relationship between each pair of the image features so that we can learn the relationship between each pair of the object and in this original so before this paper I show up people think if you want to solve the chroma dataset you muster perform new module network or compositional reasoning but in this paper they show just about a very simple relation or network you can already achieve very good performance however in
this paper only a fully connected graph is considered which as we know cannot be very efficient how we can learn the most buzz graph and in the numerator in the new ribs paper they propose a brothel owner and in this graph learner is basically we won't you later' the model itself to identify what kind of edges should be added between between the object past and what other paths should not be included and can be pruned out multimodal multi-step reasoning has also been used for for capturing the the pairwise relations between between different image regions and
that the key idea here is to say we need to perform multiple steps to refine the individuals again again to into how better performance the same idea or similar idea has also been used in in language condition the graph network but one interesting thing is here in the message of Hasenstab there was so for example given an input of question you encode each into a set of feature vectors but they will generate a different a textual command on using attention on the under question features so that they can generate the different textual commands that can
guide you how you should perform graph message passing um so you can also look into the details in the so also in the paper presented in the footnote so this is one work from our team and I want to approve reintroduce it the interest in say under the kilometer I think is not only English of the consider implicit the relations in other previous work to adjust the discussed they only consider the invasive operations that is we need to know the define what exactly the meaning between each pair of objects of the the image regions inside
the image so in the explicit relation model we have defined following in previous book we have defined the two kinds of relations semantic relation and also special relation as we can see in the example in the right here many questions semantics of relation can be very important for you and in some cases special relation are very important for the model to answer the question so through three different relation features we can refine your original image features and then this can be blocked leading into any way queuing model for for further multimodal fusion so close attention
is important intervention is also important how about just compare them together and this is the idea in M can pay per which is the winning entry to the wiki return in 2019 and the Sameera Sicilian ideas has also been exploited in other papers and the MK mode is already very close to the current popular we pass their progeny progeny models so what the model look like it is based on transformer so as we know transformer performs self attention we have a key value and Oh Perry and the in the self attention we so basically say
order key value and the query they are from the same mortality and the from the candidate engine we can see the queries from modality X or the K and the way from another modality called Y so self attention basically performs the intimate attention and the guided attention basically performs through and by combining them together we can have a model that performs cross model and the intra model attention together and we can start the model deeper to help better performance and after doing this weekend following the standard hide behind you have using a cross entropy loss
to predict your answer if we want to talk about multi mode disturbed reasoning one work that we cannot miss is the Mac model it's the phone name is memory attention and the composition is a performance not another reasoning by introducing the so called Mac cell and in the maxilla there are two kinds of units one is called control another is called a memory and I will tell you so what are control and memories so in the control control is used at you to do the reasoning operation that should be accomplished at this step in terms
of implementation this is a an attention based average of a given prairie and in the memory part is basically an attention based at average of a given knowledge base I can show you our example here this can be more clear so this is a very complicated question what color is the Matteson to the right of the sphere in front of the tiny blue block is very complicated and for the control unit they will first use attention mechanism to find oh I should have first define that the tiny blue see and then the memory part will
do attention over the knowledge base which is the input image to pinpoint oh this is the tiny processing dimension but then the control unit based on the current information they will all say oh for the next step what I needed to do reasoning I should have focused on the sphere in front on this part and then they will pinpoint to the corresponding image region again and the by performing this through the control unit II you can know or which part of the question that I should focus on a huge step this idea is also used
in the neuro state machine model but a fundamental difference is that also one of the very interesting perspective from the neuro stem mention here is that in some cases we may not need the turns image features because when we see in the reason with concepts and not visual details not in a percentage of the time what it means is that when we see a question we try to answer this question based on the visual context we have already abstracted the the image information into a very structured knowledge base which is the so called a semantic
world model and in their case this semantic world model is organized as structured the same path and I think that this prospect is very interesting and then the problem has been formulated as how we can go through or travel through the scene graph to answer the question so just now we have talked about all these previous works from different angles and one thing I want to mention is all these previously mentioned model they are modernistic network and what I mean by monolithic network is because so it's trying to say they are a huge people neural
network rather than the design in the neural module network so yes the end of a very huge apibe in neural network to perform all the reasoning in the implicit space how about we perform some to design some newer models and for each new model they can perform some elementary reason instance and then we can combine or example the neural models together in order to drop to tricky Brown the final answer there are several seminal works in this last work I will only talk about the first two three and the briefly mention our recent work called
the mental model networks so so so before I go into the neural module I want you first to clarify what do we mean by composition of visual reasoning steer the clever data set Sudheer are very complicated into the question in order to answer this we need to first identify the big sphere then goes to sphere on laughter then goes to Gruber cylinder then goes to sphere of the same color then we can drive that we can obtain the different answer so in many of the questions like this we have some common elementary operations and then
which if we can use a small neural network to to to to perform the elementary operation then we can just and embo different models together to to get the different answer as I adjust the Sun so the idea is we first have a program generator for a given question we first add Rio we first obtain a program then the program will be sent to an execution engine and the execution engine will implement this program to obtain the answer in the early work on the program generator is by a preacher and the cement depositor and then
each elementary operation cross bond into a small new module network and then we can get a big network by example in the new low modules so these two stamps they are chanted separately in the early work then in the in the next paper the pretendin cemento capacitor has been replaced by an our steam generator to perform a sequence to sequence running to predict the program however because when you predict the program it's a so so you also have a set of tokens so the operation is discrete and the you know the so the only final
supervision you have is the answer how you can backprop through all the other execution engine and also the problem generator then we can use the reinforcement the reinforcement operation to perform this and as we can see what do the new models learn they also can pinpoint to the corrupt object or the image of regions a very accurately although the question is very compositional this is a concurrent work basically using the same idea and I'd like to briefly mention our recent work on men homogeneous work just a know I have mentioned the two big lines of
work wise more mystical Network why is the neural model network can we try to preach the gap between the two and also further improve the generalization ability of traditional new nonmotile networks the idea we provide is that asked if I am an Hamal you so the mental model is one p and then through the so-called recipe that through the recipe input that we can initiate the mental model into different specific small modules and then we can combine the small modules to do the reasoning step one by one hood to obtain the final answer so from
the implementation wise we only has one member module so it's a monolithic Network which is a purely from the novelist patio it is also neuromotor Network because although we have one big mental model but this mental model has been being initiated into different specific modules and these specific modules will be combined and then brought together to predict the answer so usually people find the numeral module networks cannot perform well on the sense they may all perform well on the sincerity data sets and the monolithic they have modernistic networks they generally have good performance but the
interpretability may need further investigation so having them proposing a mental model network is a nice combination of the two I will not discuss this but I will just show two examples wise robust the wiki way so for example we need to overcome the language prior we need to all come the language base and this paper shows using the besouro regularization we can do this this is something called a self self critical reasoning published in 2019 Europe's and I will skip this so I think this is the the last two sections the front part I want
to talk about what are the core challenges and the future directions from my perspective so first I will give take a summary of the takeaway messages we have many popular tasker's I just described and the from the masses perspective indeed there are many messes that have been developed we talked about the image features weather which of the user grades are used in the image region features we talked about how we can do protein but in your polling or you film we talk about the cross model attention we'll talk about relational reasoning this is called herself
attention this is called the Prophet tension in the prophet nuah Network literature and one thing to note is transformer now really become very popular in the field we also talked about multi-step reasoning we also talked about the new low state machine this model is unique because they use the same graph to represent the employee images we also talk about neural models then what are other challenges so for me one thing I want to talk about is can we have something like a glue or superglue so if you are so if you are not familiar with
the context the glue and the superglue they are to set of benchmarks used in an open field so for example in the glue they have a set of data sets that can benchmark the progress of the models ability to to understand the language and so but in the body in division past language space people usually work on different benchmarks separately with either where we focus on weekly or we perform so we show conscious reasoning so one interesting question to ask is can we have something like the glue and the superglue in an open field to
benchmark models reasoning capabilities through multiple benchmarks so I think this is an interesting idea but the indeed may be not very easy to implement because of our many popular data sets for each one of them they have their own leaderboard and so I will just leave this as an so open question indeed only considering pushing the state of art may not be the under go and they should also be the PD I mean with only evaluation standard to justify whether paper is good or not but from another perspective if you have a good benchmark data
set this is also very good our tool to measure the progress in the field in other works I mentioned about it's a two-step pipeline we first use Pro gender CRM model to get the unity features then the rest is all about an LP problem the question that we can ask is can we make it end to end can we try to introduce something called a wish or transformer which is becoming popular in the very recent papers can we use so-called wish or transformer to encode the images and then combine this with later which in France
language transformer so that the model can be changed and to and from the tokens and the pixels to the answer Oh instead of using transformer can we just the not yourself attention and can we try to use some other architectures for example Agri film kind of fusion masses to perform to perform multimodal progeny since all the reasoning is performed in the embodiment or in the neuro space it is not entirely clear whether the model truly learns how to reason or how we should define so what is the reasoning process and this can can play way
to after some research for for neuro symbolic reasoning that could be interesting the last thing I want to mention is the so called that was so robustness here the sole robustness means so in the image of Aging in the computer vision field if you only had a very small at the basura perturbation although as so from the human eyes we cannot see the difference but from the model side the prediction can be totally different so there were several robustness and the besouro attack has been widely studied in the image so so in the image domain
in the computer vision community but and also very recently it becomes popular in an open field but this kind of model this kind of a study in the way plus L space is the last exploit which i think is so it made it service on for the so with this being said I would like to conclude my talk and so we all come to to trust me and in questions using the many maybe four or five minutes yeah thank