From here on, we will talk about the quantities that describe how much a data set is spread out from the centre. So the most simple one would be the range which is a difference between the limits. In other words, maximum minus minimum so we already have seen this statistics before when we needed to determine the width of a bin to construct a histogram.
So the minimum level of measurement of data to calculate the range should be ordinal and above because you need to be able to find minimum and maximum so that actually implies you can rank order the data, which is a property of the ordinal level of a measurement. So you cannot calculate the range for nominal data either and like mean the range is also sensitive to extreme values in a data set so let's just work with the same data, the number of facebook friends like before. So it's easy to calculate the range because this is already sorted and the maximum is 2340 and the minimum is this so 22 is a 2318.
So that is the range. So that's the range for the original data but let's just remove. .
. just well let's remove this right because this is odd one and now new maximum becomes 121 but that's range1, range2 121 minus 22 that's 99. it is, right?
so if you compare these two ranges you can see that there's a quite a lot of difference when there is an extreme value or not so as I already said range is also sensitive to the extreme values in a data set. Another useful dispersion statistics is called a an inter-quartile range or iqr for short so to calculate the iqr you need to calculate the quartiles first which is a special kind of quantiles and then quantiles are the cut points dividing the range of observations into the smaller and equal proportions. So the number of divided groups is always one more than a specific quantile so therefore quartiles are the three cut points that divide an ordered data set into four groups each comprising a quarter of the data right so you have to remember this is a three points not four points dividing the data set okay so the first quartile also known as q1 or the lower quartile or 25th percentile is the quartile below which the lowest 25 of the data are located and then the second quartile is q2 and this is basically the median by definition or this is also known as 50th percentile so the second quartile cuts the data set in a lower half and upper half and finally third quartile q3 known as upper quartile or the 75th percentile and so this is the quartile above which the highest 25 of the data will be found so the interquartile range iqr is basically the difference between the upper and lower quartiles so iqr equals q3 minus q1.
So let's find out interquartile range for this number of facebook friends data. so the data so to uh you know like you know like to find the median um you know you need to sort the data first and then this is sorted already and when you have odd number of data so like here we have 11 number of data set so the easiest step to find the iqr is to find median first so we know the median is the middle number we know that that is median so that is our q2 so that's q2 and now we have lower half and then q1 is median of the lower half so 53 is basically the q1 and 116 is the median of the upper half that's q3. so iqr becomes 116 minus 53 so that's 63.
That is the iqr for this data. As long as you know how to calculate the median then finding an iqr should not be so difficult whether or not you have an odd or even number of data. You can just follow the steps here then it should be easy enough.
For the interval and ratio levels of a measurement, measures of dispersion is typically calculated from the mean so the deviance is the distance of individual datum from the mean so you will have the same number of deviant statistics as the number of raw data. So here x represents the data right so and as we have seen, x bar represents the sample mean so that is the mean of the x data and then so because there's the same number of deviance statistics as the raw data does not really summarise the dispersion of the data system right so we want to have just a single number to summarise the total amount of spread or dispersion. So you can add all these deviance statistics to have a single number.
But if you do that then you will get zero always for any data because the deviance is the distance from the mean with sign. So if you add them up, then they always cancel each other. So one way to avoid this is to square each deviance first to remove the sign then add them all up um now that quantity is called sum of squared errors.
Now if you divide this quantity by the number of data minus one then we obtain a statistic called variance which is roughly the average squared distance of the data from the mean. Now if you think about it this is a bit awkward quantity because it's a squared value so is the original unit of measurement. So for example if you measure the height in centimetre then the unit of the mean should be the same centimetre.
However, the variance will be in the unit of centimetres squared, therefore, you cannot directly pair the variance with the mean because they are in different units so now we take the square root on the variance to get the original unit of measurement back and this final quantity is called standard deviation which is the average amount of distance in the data from the mean. So this is basically kind of by showing you the steps to calculate standard deviation. The calculation of standard deviation is quite involving but this is a very important statistics, So I strongly recommend that you practice and understand thoroughly how to calculate the standard deviation step by step.
So the very first step in calculating a standard deviation is to work out the average of the data first. So in this first column so we have facebook friends and this is the raw data. So you need to calculate the mean of this data set first by adding them all up and divide this by the number of data which is 11.
Then that's the mean of the data and then once you have the mean then you have to work out the difference which is the deviance between each number and the mean. So you subtract this from this and you subtract this and that and you subtract this from this so you get 11 which is the same number of deviant statistics so and the next is the deviance. So these are the individual deviance from the mean but if you add all this up then they cancel each other out so you get zero so that is total deviance.
You always get zero so to avoid this now you have to square each deviance statistics and then you get these large numbers, large values now you can add them all up so that big sigma sign is the adding sign adding whatever is on the right here in the bracket. okay so what's in the bracket is the deviance and it is squared and if you sum them up from the first data to 11th data, the last one, then this is a sum of the squared error and you get this value. Now to calculate the average amount of squared difference, you divide this value by 10, 11 minus 1 instead of 11 so this is what is called degrees of freedom which i'm going to try to explain later and then this value, this quantity is called variance but because this is a squared unit you want to actually take a square root on this variance to get the original unit back right so now this you take the square root of the variance then you have standard deviation so now we have seen how to calculate the standard deviation and I also want you to understand how to calculate the standard deviation and hopefully you can calculate your standard deviation on your own when you are given relatively smaller data set I mean you have Jamovi or statistical software so really you don't really need to calculate the standard deviation by hand but you want to understand how it is calculated so that you better understand what it really represents.
So here are the properties of the sample standard deviation so standard deviation measures the average spread the average amount of spread about the mean in the data set and it should be so because it is using the mean to calculate the standard deviation right so you should this standard deviation should be used only when the mean is chosen to represent the center of the data what that means is that you cannot pair the standard deviation statistics with median okay because standard deviation is the distance from the mean so it doesn't really make sense to pair this up with other central tendency measure. It always goes with mean okay and standard deviation is always positive and in rare cases, it can be zero when all the data set have the same value and if there's a more spread in the data set then it'll be shown as a larger standard deviation and as I said standard deviation has the same measurement unit as the original observations as opposed to variance where it has squared unit and because the standard deviation uses the mean it is also sensitive to outliers or extreme values in the data like the mean was sensitive to those extreme values or outliers. okay now I think I need to explain why the sum of the squared errors is divided by n minus 1 instead of n in calculating variance or standard deviation as opposed to you divide the sum of the data by n when you're calculating the mean.
so the reason has to do with the concept of degrees of freedom which is the number of data that can freely vary in calculating a sample statistics. so let me use some analogy here to make it more understandable I hope. So let's say that you are given the kind of a power or responsibility you're being a coach to select players from the entire Scotland for an Olympic national football team.
So as you probably know the minimum required number of players for a team, for a football team is 11. So imagine that you are to pick a player by field positions then there is practically unlimited number of possibilities or combinations. So you have this freedom to pick anyone and place any first 10 players on the field except the last one, meaning that once you fill the first 10 positions then the position of the last member of the team is, whatever and whoever it is, fixed.
For example you can pick and place them like three, four, three so this is just a just one example but you know in theory you can just do this any way you want right if you don't like say this guy then you can just kick him out and then replace this member with other players from the population. and so there's just millions of different possibilities to have a team but once you fill the first 10 members for the team then the last member of the team, the position of the last member is fixed so somebody has to be a goalie right and this is just the same for any position if you fill the first 10 positions then the position of the last member is fixed and you cannot change it. So in calculating a statistics from a sample is like this.
So basically you see you calculate the sample standard deviation given the sample, say you have a sample of x with all the same numbers so now your sample size is three and you can calculate so this is your first sample and x bar for the first sample is three but the mean is the same as the member of the data because all the members are the same in this case so given the sample size of 3 the sample mean for this sample is fixed to 3 and this 3 is used to calculate the standard deviation. So we cannot change this sample mean to calculate the standard deviation if we change the sample mean then you know it'll change the standard deviation but without changing the sample mean, we can still have different samples with the same sample size and the same sample mean of three having different members. So for example, in the population, we can have another sample by replacing the first two members with one and four and having the same sample mean of 3.
so you can freely change the first two members to this one but once you do that then the number of the last data, the value of the last data is fixed to have the same sample mean of three so in this case, this should be four right to have the same simple mean of three so this is determined. So you can have actually millions of this so say x and three was 2 ,2 so you can still replace the members until the last one so once you change the first two to two two then the last one should be five to have the same sample mean of three and there can be millions of such samples can exist. this you know the sample the collected sample is like your team.
Once the team is fixed then you cannot change the team I mean the team as a whole, but you can still replace the members until the last member because once you change the m minus one members then the nth member is actually fixed okay so that is a concept of degrees of freedom so what that means is that not everyone so you have to actually take away one from the number of observations because not all the members can freely vary in calculating a standard deviation because if you do that then you you're going to change the sample mean right so that is the concept of degrees of freedom so this is just what I've been saying already so degrees of freedom is related to the number of observations you can really change in computing a statistics so if a statistics is held constant so if it is a fixed, so in our case it was mean, to calculate another statistics which was standard deviation and in a given sample then the degrees of freedom must be one less than the sample size. So in our case, we have to take away one from the total number of observations in calculating the standard deviation because the sample mean is fixed so you cannot replace you cannot use all the members of the data all the members of the sample to calculate the standard deviation given the sample so I took this much time to explain what the degrees of freedom is and just because you have to know the degrees of freedom of whatever statistics you calculate so it is a standard procedure you report the degrees of freedom for any statistics if there is any degrees of freedom involved right so it is important that you know the concept of the useful from so there's such thing as degrees of freedom and don't forget to report the degrees of freedom or any statistics and this will become more important as we go along because there are different statistics involving different number of degrees of freedom and so we're going to just talk more about this degrees of freedom later on. So the notation for degrees of freedom is nu so that is not v it is nu in greek letter okay so that is actually standing for degrees of freedom or in more plain english and people use tf or degrees of freedom or d dot f dot for degrees of freedom whatsoever so I think that that is all I have to say about the exploratory data analysis and now to the summary slide.