Now let's look at each of these exploration steps starting from sorting and counting with some data. So here we have a data set of uncorrected visual acuities measured in logMAR with both eyes open. So when you have a data set the very first thing you want to do is to count the total number of data.
So how many patients' visual acuities do we have here? so let's just count that's and so one, two, three, four, five, seven, eight, nine, ten, eleven, and we have two rows so we have 22 measurements of visual acuity. When you report the number of data entered into the analysis, you use small n like this equals 22.
So this is how you report the number of data you used in the analysis After you count the number of data then you want to actually sort this out because This data set is just pretty much randomly organised so if we sort in ascending order so from the smallest value to the largest value then now we can find out you know what's the smallest one right so that's a minimum and this is maximum right um So once you sort the data then maybe you want to actually group the data to see what's going on and to summarise the data better so now let's make a grouped frequency table to summarise the frequency of observations in a kind of manageable number of groups events or categories so here um I just uh you know divide the the data into five groups already um by every one tenth of visual acuity range so here the first interval is ranging from uh 0. 0 logMAR to 0. 1 logMAR and I don't know if you ever seen this kind of notations so the the square bracket on the left of the 0.
0 means that it's inclusive of the number on the right so the range starts exactly from 0. 0 logMAR up to 0. 1 but not including 0.
1 so this round bracket on the right uh is basically exclusive of the number on the left. So the reason why we do this is that because the logMAR is a real number so there's kind of infinite number between any two numbers right so to make this interval so we have these five intervals right to make this interval continuous in the the real number so the next interval starts from 0. 1 so that there is no gap between the upper limit of the previous interval and the lower limit of the next interval so that every interval is completed basically without any gaps in between so it's all continuous right?
So the square bracket means inclusive of and then round bracket means exclusive of the limit the upper limit. We can count how many observations occur in each interval so for the first interval we have these two numbers right so that's 0. 02 and 0.
06 can be included in the first interval but not 0. 1 the next one right because this is actually because the interval the upper limit of the first interval does not include that number so we cannot include that in the first interval so we have two observations for that first interval and the next one is ranging from 0. 1 to 0.
2 logMAR so it goes up to 1, 2, 3, 4, 5, 6, 6. now we cannot include 0. 2 again right because the the upper limit of the second interval does not include 0.
2 the exact number so we only have 6 in the second interval right six observations in the second interval and now for the next one so that's the one, two, three, four, five, six, seven, eight, nine, and we should stop there so that's nine and the next one starting from. 3 one, two, three, and you should stop here this three and the remaining is two okay so this is how you make a grouped frequency table basically so this is just a summary table and with some the number of intervals um which divides the range of the data into small number of manageable groups or intervals okay so what we just calculated though is what is called absolute frequency then there are a number of different ways to calculate the frequency of data so we have four different ways to characterise the frequency or occurrences of data so the absolute frequency is you know what we just calculated which is the actual number of observations within an interval or a group. On the other hand, cumulative frequency is the summed frequency up to the current and all preceding intervals okay and the relative frequency is the frequency of an interval or a group divided by the total number of observations across the intervals or groups and finally the cumulative relative frequency is the cumulative frequency at the current interval divided by the total number of observations across the intervals or group so this is all just too mouthfuls so let's just take a look at how we can calculate each of this frequency.
So here we have um in a table for different frequencies now the the first column is the VA group right the categories or the intervals and then we have the absolute frequency already calculated so now for cumulative frequencies based on the definition of cumulative frequency the first interval is the frequency the cumulative frequency of the first interval is just same as the first interval of the absolute frequency because there's no preceding frequencies to be added on to right so it is still two but the next one the second one now you're gonna add this two and then add the current frequency to have the cumulative frequency for the second interval okay now you again move this 8 to the next interval and add onto the current interval okay so sorry that's just a plus sign but it is not very. . .
in other words it is like adding all the preceding plus 9 becomes 17. okay right so this keeps going so 17 plus (3) 20 and 20 plus 2 22 right so if you add all the absolute frequencies right across the intervals then you get the total number of observation which is n equals 22 right and then so cumulative frequency is this 2, 8, 17, 20, 22, so these are the cumulative frequencies for the corresponding intervals and for the cumulative frequency you do not calculate the total because the last entry the last cumulative frequency for the um at the last interval is basically the same as the total number of observations and it should be the same as the total number of observations. Now to the relative frequency.
The relative frequency is basically the absolute frequency divided by the total number of observations for each interval so the absolute frequency of the first interval was 2 and you divide this by the total number of observations, 2 over 22 and and the last entry so the total, the total of the relative frequency should be equal to one I hope that makes sense. Now the cumulative relative frequency is just a relative relative frequency of the cumulative frequency so that's the same okay it's a 2 over 22 and now it's 8 over 22 because you're going to have to use the cumulative frequency here and then the next one is 17 over 22, 20 over 22, and then the last entry of the cumulative relative frequency should be 1 right? 22 over 22.