[Music] foreign we're going to explore the possibilities of doing sequence searching using Motif searching pattern this would be very helpful for searching variable sequences usually shorter ones and we can even use the option to include uncommon amino acids which are characterized by the character X so to do that we're going to navigate to the sequence button and here instead of blast or CR we're going to take motif we're going to enter our amino acids sequence code so here at the seconds position we will have the amino acid X which represents any uncommon amino acids and
then we have the possibility with square brackets to indicate that at the fourth position we have either thrionine or we have serine as the two options for the amino acids and we use the same four positions six and then we have p q and Y as the last amino acids so with the advanced search it's good to go there and perhaps increase the e-value for this search that would not be necessary but for even shorter ones uh that would be but here we want to make use of the combined Motif results so that the results
will really be represented by one result instead of the four for the four permutation of this variable query we're going to include the ncbi sequence run proteins so let's start this sequence search and when you do that you see that your sequence search query is posted here at the top of your history and we see that we also have the results ready now let's have a closer look at them so here we do see we have the full 100 000 uh results that the system was able to provide and we can start by um looking
at the possibilities of sorting this is sorted by alignment identity so we see that it's a perfect match here um and um what we can do is make sure that the sequence length uh putting in a query with nine amino acids retrieves also very long peptides and we're going to cut this short to peptides that would have a maximum of 20 amino acids so we're going to click apply and see what that does to the number of results so that reduces the result to 573 answers we can break this further down to those that would
have um a alignment identity of let's say 85 percent so that would exclude most of the peptides that would have then a larger variety of mismatches within the query alignment so we'll apply that as well we see that we end up with Five results and let's have a closer look at at these results when you click on this subject area you would be able to see the registry number that this particular sequence has and if we're interested to see what the two x amino acids represents we can click on that registry number it opens this
up in a separate tab for us to take a look at the details and here we see the chemical structure of the sequence and in the sequence details we do find the specific uncommon amino acids here at position two and at position 10 and we also see that there is a bridge between the third and tenth uh amino acids so that gives us a quite a nice option for seeing the further information we can see the reference as well more references potentially would be found in the reference section we can also move those side finder
here but let's have a closer look at at further results here for instance we see that on the fifth one we have an additional answer where we do have a mismatch where the Q amino acid um glutamine was actually replaced by glutamic acid uh as the amino acids so again the subject would be there we can click on the registry number to see the definitions of the X and we can also see the the references the total result can be exported using the download button here at the top right that will create a Excel sheet
which would then provide um a Excel file that would have all of the information either journal or patent literature and we would be able to to see that in a nice Excel presentation here we see the sequence details have been downloaded let's open that up we can see sort of the different alignments and we can sort of navigate we can resize the structures we see a variety of registry numbers that have been added we see accession numbers of the references uh as well as PubMed or Medline references ncbi identifiers and then also the number of
patents that have been signed up with potentially the sequence IDs further details are provided in terms of e-value blast scores Etc and the alignment identity so a very compact way of reviewing the alignment details along with the sources of the the data that we that we have and again we can go back to our sequence results and look at all of the patterns that this particular set of sequences represent by clicking on this reference button and this would provide an answer of 20 patent results across the different time frames where these sequences with uncommon amino
acids have been reported and we can for instance take a look at the concepts for I know seeing the further context the applications or maybe the organizations that have filed these these particular patterns uh and with that we can navigate through these results further uh on what what applications we can find for these particular sequences [Music] um