[Music] hi everyone and welcome back today we're going to continue our discussion of one factor or one-way anova testing for the normal hypothesis that the means of several different populations are equal to one another in the last video we discussed how to conduct this test via the linear regression approach but today we're actually going to just focus on the formulaic approach without the regression perspective behind this so let's assume that we're interested in testing if there's a significant difference in the average customer service across different stores of a restaurant chain where we're defining customer service as a numerical variable between 0 and 10 and let's assume we're going to survey let's say three different stores so let's call these stores say store a store b and store c now you don't necessarily have to sample the same amount of values from each of the stores but that's definitely going to make things a lot easier and we'll discuss the pros and cons of having unequal sizes for anova testing a little bit later let's assume from store a we were able to survey 8. 2 7. 6 8.
5 and 4. 4 levels of customer service in store b let's assume we were able to sample 9. 5 9.
2 9. 8 9. 2 8.
5 and 8. 9 customer service levels and for store c let's assume we were able to observe 4. 2 2.
1 5. 6 1. 8 and 4.
3 so in terms of this particular setup we are testing the null hypothesis that the mean customer service for stores a b and c are all equal to each other so in this particular setup we have that we have k is equal to three different categories so if we were to approach this from the linear regression perspective that means we need to define k minus one or two dummy random variables for our particular model where one of these stores a b or c would be our reference group right so that's our basic setup so to start let us get some point estimates for these parameters mu a mu b and muc so one can find that x bar a or the average of sample a is approximately equal to 7. 18 the arithmetic mean of the customer service ruby approximately 9. 18 and the arithmetic mean for c is going to be equal to 3.
60 and one can also find that the grand mean so the grand meme uh which we're going to be calculating via just the average of all these recipients is going to be approximately equal to 6. 79 now one may ask is this really the best way to approximate the mean i just do the average of a b and c together or is it better to average these averages together into one different average we'll just briefly discuss this a little bit later and also it's important to note how many people we are sampling so of course the number of people that we sampled in store a was equal to four the number of samples in store b was six and the number of samples in c was equal to five giving us a grand total of 15 respondents so in order to calculate this arithmetic mean this grand mean that's just going to be the sum of all those values divided by 15. now that's our basic setup now let's calculate our anova metrics in order to conduct our anova test so the very first thing that we need to construct is what we call the total sum of squares so our total sum of squares right so the total sum of the squares so the total sum of squares in its formal definition is the sum from i is equal to 1 2k so that's the sum of cross all categories of the sum from j is equal to 1 to the size of each of those categories of x i j minus the grand mean squared okay so in terms of how to calculate this what are we actually looking at so we're going to be looking at the value 8.
2 subtract the grand mean square it plus 7. 6 minus the grand mean square it do that for all the categories in category 1 then do it the same for category 2 and category 3 and sum up those sums okay so one can actually show that this is actually equal to the variance of all the values times the total sample size minus one in that case it's going to be 15 minus 1 or 14. and you should be able to find that that's approximately equal to 107.
90 in terms of interpretation you can talk about this as the total variation in my study because i'm pretty much comparing all of my values to the grand mean so the second metric that i want to calculate is what we refer to as our error sum of squares or sum of squared residuals okay so the formal definition is going to be the sum from is equal 1 to k so the sum across all categories of the sum from j is equal to 1 to the size of each of the categories of x i j minus the mean associated to that particular category of interest so if i want to sort of describe what's going on here let's go back to our table so i'm going to be taking 8. 2 subtract my 7. 8 squared then i'm going to do 7.
6 minus 7. 8 squared 8. 5 minus 7.
8 square it 4. 4 minus 7. 8 7.
18 squared add them all up that's going to give me a number then i'm going to do the same for category 2 but instead of subtracting 7. 18 i'm going to be subtracting 9. 18 so 9.
5 minus 9. 18 squared 9. 2 minus 9.
18 squared and so on and then lastly i'm going to do 4. 2 minus 3. 60 squared 2.
1 minus 3. 60 squared and so on add up all of those sum of squares and that's going to give you what we call our error sum of squares now notice if you look in this formula technically we can actually rewrite this in a nice compact way this can actually be rewritten as the sum from i is equal to 1 to k of the variances of each of those groups times the size of each of those groups minus 1. that's actually a nicer way to go about it especially if you have excel or r at your fingertips and you should be able to show that that's approximately equal to 22.
06 all right so 22. 06 and you can interpret this as the variation so this is the variation within the groups because i am comparing each of the values within the groups to their particular group meme so that's the second metric that we have the last metric that we need is what we abbreviate as ssm or the model sum of squares so the model sum of squares in its formal form is going to be the sum across all categories of the sum across all value in each of those categories of x bar i minus the grand mean squared that is i'm going to take each of the group means subtract the grand mean square it and do that for each of the columns right now notice that this formula does not have any js in it right so technically i can factor this out out of the sum and i'm going to be doing 1 plus 1 plus 1 plus 1 n j times right so what i can do is i can rewrite this as what so that should be ni by the way and i and that should also be ni by the way let's just go ahead and fix those ni and ni and this is actually going to be equal to the sum from i is equal to 1 to k of n i times x bar i minus x bar bar squared right so that's a nicer formula that you could use to calculate the ssm and using our example you should be able to find that's approximately equal to 85. 84 in terms of its interpretation we can interpret this as the variation between the groups because it is focusing on the group means compared to the average of all of the values in our particular observation so what we're doing here is we're taking the total variation and breaking it down into two sub-categories the variation within the groups and the variation between the groups because the variation within the variation between the groups should add up to the total variation now generally speaking for all anova techniques this equation is not necessarily the case later on we'll look at some examples on when the sse and the ssm do not add up to the sst okay but at least for the sake of one factor anova so let's just formally state this so for one factor anova testing even if the cells are not of equal size the sst is always going to be equal to the sse plus the ssm right and one could abbreviate this in a more compact way as t is equal to w plus b right where t stands for total w stands for within b stands for between so the total variation is equal to the within variation plus the between group variations right so that's a nice little relation that we can use to maybe um calculate two and quickly calculate the third depending on your preferences now these are not the metrics we use to calculate our test statistic so we need three other metrics so we need the mse which is going to be equal to the sse divided by n minus k so that's going to be the air sum of squares divided by the total sample size minus the number of groups we have and then our msm our mean square error for our model that's going to be equal to ssm divided by the number of categories minus 1 and the mean square error for our total that's going to be equal to the total sum of squares divided by our sample size minus one so for our particular example you can find that's approximately equal to 1.
84 that's approximately equal to 42. 92 and that's approximately equal to 7. 71 all right so at least from these first two we can calculate what we refer to as our f test statistic so our f test statistic is going to be equal to the mean square error for our model divided by our mean square error for our error term and for our particular example you can find that to be approximately equal to 23.
35 which typically is pretty large for an f distribution regardless of the degrees of freedom pair and keep in mind the f-test statistic for this one factor anova is going to be associated to an f-distribution with k minus 1 comma n minus k degrees of freedom so that means for our particular example we're going to be associating this to a minus 1 15 minus 3 distribution or an f 2 12 distribution so once you have our degrees of freedom pair and you've already calculated your test statistic then you should be able to calculate an associated p-value for that so the p-value for this is going to be approximately equal to 7. 3 times 10 to the power -5 so that's . 0073 right which typically is pretty small right so that's pretty small right it's smaller than for example five percent or one percent or even point one percent right so since that p-value is pretty small with respect to most alphas that people would choose we're going to be rejecting the null hypothesis that mu 1 equals mu 2 is equal to mu 3.
so in terms of this rejection how do we interpret it in terms of our example that means at least one of these customer service averages is significantly different than the others right now can you take a guess at which one or which two or possibly which three are very different from the others so if you look at our particular data sets that we're looking at notice that we have 7. 18 7. 19 and 3.
6 so i would argue that this category or this store store c probably is the one that's causing the issues but of course x a and x bar b are a little bit different from each other as well so possibly all of them are different from the rest right but that's how you would conduct a one factor anova where the factor we're focused on here is customer service between different populations now that we know how to conduct a one factor or a one-way anova for seeing if there's a significant difference between population means now let us talk about a few concepts that are sometimes important to understand in regards to one factor or one-way anovas so if the group sizes n1 and 2 down to the last group size are all equal to each other we say that the anova is said to have a balanced design if at least one of these group sizes is different from the other for example in our example we call that an unbalanced design anova generally speaking the formulas for the balance design are usually a lot more easier because for example that little n i term that was sort of hanging out one of our sums that could sort of factor out because they're all equal to each other and independent of that index i but generally speaking if you know how to do it for unbalanced then of course balanced as a special case of the uh unbalanced formula so one thing i want to look at here is the formula for x bar right so what is x bar bar in terms of our calculations that we're sort of comparing all our values to so if you look at our formula for x bar bar what we did was is we took all our values regardless of the groups so it goes all the way up to n and divided by how many values we have in our particular sample now if we focus this sum in the numerator in terms of the groups then we can rewrite that as x group 1 respondent 1 plus x group one respondent two all the way out to x group one respondent n one so that's the last respondent in group one and that's going to go all the way down to x group k first respondent x group k second respondent all the way down to the last respondent of the last group okay and we can represent n as the sum of the group sizes so we can represent as n1 plus n2 all the way down to n k okay so if i have that notice i can multiply top and bottom by the size of group one so divide top and bottom of the size of group 1 and i can multiply top and bottom of that by the size of group k and do that for all of the values and notice that this relationship the sum divided by the size the sum divided by the size is precisely equal to the average so one can see that the grand mean is precisely equal to n one times x bar one plus n two times x bar two all the way out to n k times x bar k divided by n one plus n 2 all the way out to n k so what do we have here so the grand mean so therefore the grand mean is just a weighted average is just a weighted average of all of the group means right because we do not necessarily know which one to trust in terms of its value so we pretty much weigh the one who has the largest sample size a little bit more than the rest right and we have seen this weighted type of structure for example the pooled variance t-test pulled proportions and so on now let's actually look at a particular special case for this so note if the anova is balanced what do we have that means n1 plus n2 or n1 is equal to n2 is equal to n3 is equal to nk which means all of the values are equally distributed across each of the k categories okay so that means our formula for the grand mean can be written as follows so that means x bar bar is going to be equal to n over k x bar 1 plus n over k x bar 2 all the way out to n over k x bar k all divided by m and notice everything is going to have a n over k in it so that's going to be x bar bar is equal to n over k times x bar 1 plus x bar 2 all the way out to x bar k and that's going to be divided by m notice that this n and that n will cancel each other and that k will drop to the denominator so that means for the balanced design so for balanced anovas we have that x bar bar is going to be equal to x bar 1 plus x bar 2 all the way down to x bar k all divided by k so that means the grand mean for balanced anovas is equal to the average of the group averages right now keep in mind if the groups are not balanced the average of the group averages is not necessarily equal to the average of all the values so that's an important thing to notice but these two formulas are actually quite easy so if your design is a balanced design you can just find the average of each of the groups divide by how many groups you have that gives you the grand mean but if your anova is not balanced you have to take all the values add them together and divide by your total sample size right so those are just some interesting things to point out about the grain mean every time we're working with a hypothesis test method for example the anova it's very important to understand the assumptions and premises for which the method is built upon so remember the last thing that we usually calculate in an anova test is the f-test statistic where an f-test statistic is a quotient of two independent chi-squared random variables typically of different degrees of freedom and remember chi-squared random variables are built from normal or standard normal distributed independent random variables so if we are not just sampling from independent normally distributed random variables then the f tests test statistic may be extremely misleading right that's not of course the only assumption that we are building the anova upon if you remember from our regression analysis there's one very important assumption that must be met and that is the hamas elasticity term for populations so what does it mean for populations to me to be homosadastic to one another that means the variance of population one the variance of population two all the way out to the variance of population k are all equal to another or approximately equal to each other at least in some statistical manner so let's look at a few scenarios in terms of the equality of population variances and also how sample size differences may come into play here in terms of our analysis and also the stability of a particular study right so the first scenario is if all our populations are homosadastic to one another so sigma j squared is equal to sigma i squared and let's also assume that the group sizes nj and ni are equal to each other then that's great that's actually the most ideal scenario and you may be thinking well typically speaking is always possible to have equal sizes no of course not and of course we would like to be as large as possible in terms of the sample sizes right but actually this is better than equal variances in unequal sample sizes right which is case 2. so case 2 is if sigma j squared and sigma i squared and our groups are not necessarily of the same size that's okay too that's okay too right so don't be sad if your groups are not the same exact size if you have evidence to believe that the variances are equal to each other that's okay as well now scenario 3 and scenario 4 are of course the worst cases so if our variances sigma j squared and sigma i squared are not equal to each other and nj is equal to ni it's bad but not too bad right so it's it's it's sort of bad but it's not the worst case scenario so the worst case scenario is if sigma j squared is not equal to sigma i squared and the anova table setup is unbalanced so this is the worst case scenario in our anova analysis realms now how does sort of sample size play into this picture and why should we consider so let us recall what we refer to as statistical statistical power so statistical power which we usually refer to as a probability of committing or not committing a type 2 error right and usually sometimes we abbreviate statistical power via greek letter pi so the probability of committing a type two error that's usually what we represent as beta right so remember beta is the complement of the type two error probability all right and remember what we want about statistical power well generally we want our type 2 error probability to be as small as possible that means we want statistical power pi to be as large as possible right that's a very important thing right being able to prove things that are wrong that's definitely an ideal case now there's a theorem which we're not going to prove here but it's very important to be aware of is that the value of pi statistical power for anovas is maximized or maximal when the group sizes are equal to each other i. e a balanced design and that has to be true for all k groups right so once you start having groups of different sizes you're actually decreasing the statistical power for this method to be able to pick up on any significant difference okay also let's assume let's actually briefly review what we mean by type 2 error just in case you don't remember what that is right so let's just briefly review so remember type 2 error is the probability of or the it's the of failing to reject the null hypothesis or remember the null hypothesis that these are all equal to each other when this null hypothesis is actually false that is when we conclude our anova test that means we have evidence to believe that they are equal but if they are not that's what we refer to as a type 2 error now let's spin back into that little theorem that we said about if the anova is of balanced design then our statistical power is optimal right so let's look at a particular scenarios of for example good anovas bad anovas are the best anovas possible right so what do we know about the sample size so as we already know ideally the larger the sample sizes of the groups the better right because that usually decreases the standard error of each of our point estimates so let's assume for example we have a sample size vector of say 15 7 9 and 17.
so 15 in group 1 7 in group 2 9 in group 3 and 17 of group 4. so this is not good statistical power right but if we look at for example the another sample size vector for example if we have 15 14 12 and 16 this is not let's actually call that n vector 1 and n vector 2. this is not good statistical power but it's better than n hat one because the sample sizes are more close to each other and there are some statistical tests to sort of see how spread out they are notice what metric would you use to sort of um test whether the variance or the range is close to zero maybe you can sort of think about that and sort of see how to build a nice little test for that but remember good statistical powers when all the samples or our group sizes are equal to each other so if we consider the sample size vector n hat three and that's for example seven seven seven and seven that is good statistical power that is the maximal statistical power for which you can have for that particular sample size combination but the sample size vector n hat 4 which is for example 18 18 18 and 18 also has maximal statistical power and has good sample sizes or i wouldn't say good sample sizes but better better standard error than n hat 3.
right because usually as our sample size increases our standard error approaches zero all right so just keep those in mind yes if you have for example one of these categories um let's assume that you jump up to like for example 100 so 18 18 18 100 you're not gaining any statistical power in order to decide whether um they are significantly different from each other or not just the standard error associated to that point estimate for say mu k is just a little bit lower that's the only thing you're really getting out of that particular setup for unbalanced anovas there's one other thing that i do want to mention today and that is what we refer to as the corrected sum of squares and the uncorrected sum of squares the only reason why i want to sort of mention this now is if you're using some software such as spss sas r and sometimes even excel sometimes you'll see this word corrected or uncorrected next to your sst so i just want to sort of briefly mention that just to sort of give you some insight on what's going on there okay so this is just a side note not super important but it's interesting to maybe think about what it represents so remember the total sum of squares was equal to the sum from k sub 1 to n let's actually represent this from the regression perspective because i think it's a little bit more easier to understand is going to be equal to y k minus y bar squared so that's equal to each of the values in our sample minus the mean for which um they all have right so this is sometimes referred to not just as the total sum of squares but as the corrected sum of squares because notice that it's not really a sum of squares it's a sum of squared differences right so the use of the word total sum of squares is not actually good it's a corrected sum of squares that is we're subtracting the mean from each of the values which we are technically interested in squaring so the uncorrected sum of squares which is represented by the sum from because one to n of y k squared is called the uncorrected sum of squares because if you do not correct this sum of squares representation for sst then you're not going to get that nice little beautiful property that sst equals ssm plus sse right so if that is the case what is the relationship between the corrected and the uncorrected sum of squares because clearly one is hiding out inside of the other so let's actually look at the sst and its decomposition so sst remember is the sum of from k is equal to n of y k minus y bar squared so if we sort of expand this via some algebra that's going to give us the sum from k is equal to n of y k squared minus 2 y k y bar plus y bar the quantity squared and then we can distribute the sum over and then factor out each of the constants out of that sum so that means we're going to have the relationship sst is equal to the sum from k sub 1 to n of y k squared minus 2 y bar sum from k is 1 to n of y k and then plus y bar the quantity squared times the sum from k sub 1 to n of 1. right so remember that that is just equal to n that's just the same as y bar times m so if i look at this expression that's going to be minus 2 n y bar squared and then that's going to be plus n y bar squared and then that's going to be okay so minus 2 cats plus a cat that's just going to be minus a cat right so that's just going to be minus n y bar squares so that means we can represent this as the total sum of squares i'm going to call it corrected is equal to because that's the uncorrected total sum of squares right so that's the total corrected some uncorrected sum of squares and then we're going to have minus n y bar squared okay so we have this little difference from it right so what exactly is that so this term has a few names but i'm just going to give it a name just for the sake of giving it one this is what we refer to as the mean correction we're not going to get into the theoretical understanding of what exactly that is but pretty much what we have here is that the corrected total sum of squares is equal to the uncorrected total sum of squares minus the mean correction okay so that means the uncorrected total sum of squares is equal to the total sum of squares corrected plus the mean correction right so that's the difference between correction corrected and uncorrected total sum of squares where typically we usually focus on the corrected version because we have that nice little property that sst is equal to sse plus ssm and sometimes we refer to this property as the orthogonality this is the orthogonality property of the sum of squares and if you're familiar with linear algebra you may be thinking well i know what orthogonal orthogonality means in the vector perspective that means if you take the dot product between any two vectors that's going to give us 0.