YouTube transcript (JgMFhKi6f6Y) — Transcript
Full transcript
- 0:02[Music]
- 0:17hello thanks for watching and welcome to
- 0:20the next video in my series on basic
- 0:22statistics now as usual a few things
- 0:25before we get started number one if
- 0:27you're watching this video because you
- 0:28are struggling in a class right now I
- 0:31want you to stay positive and keep your
- 0:33head up if you're watching this it means
- 0:35you've accomplished quite a bit already
- 0:37you're very smart and talented but you
- 0:39may have just hit a temporary rough
- 0:41patch now I know with the right amount
- 0:43of hard work practice and patience you
- 0:46can work through it I have faith in you
- 0:49many other people around you have faith
- 0:51in you so so should you number two
- 0:55please feel free to follow me here on
- 0:57YouTube on Twitter on Google+ or on
- 1:01LinkedIn that way when I upload a new
- 1:03video you know about it and it's always
- 1:06nice to connect with my viewers online I
- 1:08feel that life is much too short and the
- 1:10world is much too large for us to miss
- 1:12the chance to connect when we can number
- 1:16three if you like the video please give
- 1:18it a thumbs up share it with classmates
- 1:21or colleagues or put it on a playlist
- 1:23that does encourage me to keep making
- 1:25them for you on the flip side if you
- 1:27think there's something I can do better
- 1:29please leave a instructive comment below
- 1:30the video and I will take those ideas
- 1:33into account when I make new ones and
- 1:36finally just keep in mind that these
- 1:37videos are meant for individuals who are
- 1:40relatively new to Stats so I'm just
- 1:42going over basic concepts and I will be
- 1:45doing so in a slow deliberate manner not
- 1:49only do I want you to know what is going
- 1:51on but also why and how to apply it so
- 1:56all that being said let's go ahead and
- 1:58get started
- 2:02so this is the next video in our series
- 2:04about the analysis of variant or an NOA
- 2:08to make the topic more manageable I've
- 2:10divided this video into two parts in
- 2:13part one we will discuss the conceptual
- 2:15background using graphics and charts and
- 2:18an example problem and then in part two
- 2:21we will actually go into Microsoft Excel
- 2:24to conduct a hand calculation of an Nova
- 2:28and also solve our examp example using
- 2:31excel's builtin data analysis tools if
- 2:34you're watching this video on YouTube
- 2:37you can download the Excel file using
- 2:39the link in the description so you can
- 2:42follow along when we get to part two now
- 2:46in our last video we took an in-depth
- 2:48look at the conceptual foundations of
- 2:51Anova and what it offers us prior to
- 2:54Anova we were limited to conducting
- 2:57hypothesis tests about a maximum of two
- 3:00populations an NOA frees us of that
- 3:03limitation permitting comparisons of
- 3:06multiple populations and even subgroups
- 3:09or what are called blocks of those
- 3:12populations now there are several types
- 3:14of Anova in this video we will be
- 3:17focusing on what is called the oneway
- 3:19Anova and in a later video we will
- 3:22discuss the two-way Anova it is
- 3:25important to note that one way anovas go
- 3:28by other names such as as the single
- 3:31Factor Anova and in experimental
- 3:34contexts the completely randomized
- 3:36design these are all basically the same
- 3:39thing now this video is very
- 3:42comprehensive in part one I will use
- 3:44several illustrations to make concrete
- 3:47the abstract Concepts underlying an NOA
- 3:51in part two we will go into Microsoft
- 3:54Excel to hand calculate the Innova which
- 3:57admittedly you may never actually have
- 3:59to do do so this video offers a solid
- 4:03comprehensive conceptual Foundation of
- 4:06oneway in Nova using an example problem
- 4:09with illustrations and Graphics so if
- 4:12you are new to an NOA or are still
- 4:14trying to figure out exactly what it is
- 4:17this video is for you so sit back relax
- 4:21and let's go ahead and get to
- 4:25work so the most obvious question is why
- 4:27do we have an NOA in the first place so
- 4:30like I said before to this point we have
- 4:32been comparing two populations only so
- 4:35we use the independent samples T Test or
- 4:37the Matched sample or the paired T Test
- 4:41of course limiting ourselves to the
- 4:42comparison of two populations as well
- 4:46limiting what if we wish to compare the
- 4:48means of more than two populations what
- 4:51if we wish to compare populations each
- 4:53containing several levels or subgroups
- 4:57well that's why we have a NOA
- 5:00it allows us to do those things and
- 5:02remember a Nova is an acronym that comes
- 5:05from the phrase analysis of
- 5:11variance so suppose we want to compare
- 5:13three population means like we will in
- 5:15this example problem to see if a
- 5:17difference exists somewhere among them
- 5:21so we have population one here in the
- 5:23blue population two here in the pink and
- 5:27population three here in the green and
- 5:29each population has its own population
- 5:32mean so mu1 mu2 and
- 5:37mu3 so what we're asking is do all three
- 5:40of these means come from a larger common
- 5:44population sort of in the
- 5:47background or is one of the means so far
- 5:50away from the other two that it is
- 5:53likely not from the same
- 5:56population or are all three so far apart
- 6:01that they all likely come from unique
- 6:04populations so really it's about
- 6:07distinction are all three of these
- 6:09populations common to a larger
- 6:11population or is one or all three of
- 6:14them different from one
- 6:17another now I mentioned before there are
- 6:20several types of innovas so if you go
- 6:23into the data analysis section in Excel
- 6:26you will see all those so you will see
- 6:28Anova single Factor
- 6:30a NOA two factor with replication and a
- 6:33Nova two Factor without replication now
- 6:37we will get to all those eventually but
- 6:40for now we're going to be dealing with
- 6:41the first type so in this video we will
- 6:44be learning about single factor or
- 6:47oneway anovas they are the same thing so
- 6:51if you go into Excel and see this menu
- 6:54in the data analysis tool I want you to
- 6:56know exactly which one we're doing we
- 6:58are doing the Inova
- 7:00single factor or the one-way
- 7:04Anova so here is our fictitious problem
- 7:08so 21 students at the autonomous
- 7:11University of Madrid the auum in Spain
- 7:14were selected for an informal study
- 7:17about student study skills so s first
- 7:21year 7 second year and S thirdy year
- 7:24undergraduate students were randomly
- 7:27selected the students were given a a
- 7:29study skills assessment having a maximum
- 7:32score of 100 now as researchers we are
- 7:35interested in whether or not a
- 7:37difference exists somewhere between the
- 7:41three different year levels so we will
- 7:44conduct this analysis using a one-way
- 7:47Anova
- 7:51technique so here is our basic chart so
- 7:54you can see that we have a column for
- 7:56each year of students so these are our
- 7:59columns s or sometimes they're called
- 8:01groups and I mentioned in the previous
- 8:03video one of the challenges of teaching
- 8:06and learning and novas is that different
- 8:09professors and different textbooks will
- 8:11call the same thing different things so
- 8:15sometimes these are called columns
- 8:17sometimes these are called groups now
- 8:19why is this called a single Factor Anova
- 8:22well if you look at this chart this year
- 8:25of student is a single factor with sort
- 8:28of several levels to it so that that's
- 8:31why it's called a single
- 8:34Factor now within each year of student
- 8:37we're going to select a random sample in
- 8:40this case seven students so if you get
- 8:44the geometry of the chart you can kind
- 8:47of understand how it's all put together
- 8:49so we have three columns or groups the
- 8:52single Factor we're looking at is the
- 8:54year of the student and within each
- 8:56group we're going to be taking a random
- 8:59sample
- 9:01so here is our actual data so remember
- 9:04these scores are out of 100 so you can
- 9:06see the seven year one student scores in
- 9:08the left the 7e 2 student scores in the
- 9:12middle and the year three student scores
- 9:15over here on the right now remember that
- 9:18these are random samples within each
- 9:21year or each column each group that's
- 9:24why in an experimental context this is
- 9:26called the completely randomized design
- 9:29so you may see that in your class or in
- 9:31your textbook or of course
- 9:35both so I want ahead and color coded
- 9:38each year so we can keep them distinct
- 9:41now if you notice at the bottom here I
- 9:43have xar sub one xar sub 2 and xar sub3
- 9:48well of course that's going to be the
- 9:49mean for each column or each group so
- 9:54each year will have its own
- 9:57characteristics it will have its own
- 9:58mean
- 10:00and its own distribution or its own
- 10:02variance so I went ahead and colorcoded
- 10:05those to
- 10:06match now there will also be an overall
- 10:09mean which is the mean of all 21 scores
- 10:14taken together sometimes this is called
- 10:16the Grand mean I call it the overall
- 10:19mean same thing so there are several
- 10:22means under consideration here we have
- 10:25the mean for each column or each
- 10:27factored level it's the same thing cuz
- 10:30our single factor is the year of the
- 10:32student so we have xar sub 1 xar sub 2
- 10:35and xar sub3 then of course we have the
- 10:39overall mean or sometimes it's called
- 10:41the Grand mean which is the mean of all
- 10:4421 scores taken
- 10:48together so the first thing we want to
- 10:50do is find the mean for each column and
- 10:53the overall mean over here on the right
- 10:57so the mean of the Year One scores was
- 11:0171.7 that's xar sub 1 the mean of the
- 11:05year 2 scores was
- 11:087.29 that's xar sub 2 then xar sub3
- 11:13which is the mean of the year three
- 11:14scores was
- 11:1676.5 s now the overall mean or the grand
- 11:21mean of all 21 of those scores taken
- 11:24together is
- 11:2774.5 2 so that's the first thing we want
- 11:30to do when doing our hand calculation
- 11:33for the oneway
- 11:34Anova so you can see our column means
- 11:37here on the bottom and our overall mean
- 11:40here on the right now if we take a quick
- 11:43look at the column means what do we see
- 11:47well we can see that the 71.7 one over
- 11:51here for the year one students seems a
- 11:55bit odd it seems a bit different than
- 11:57the other two now we don't know if
- 12:00that's going to be just due to Natural
- 12:01variation or there's actually something
- 12:03there that's why we're doing the actual
- 12:08Anova so let's briefly go back and visit
- 12:11the idea of variance and its related
- 12:13concept the sum of squares so since
- 12:15Anova is by definition the analysis of
- 12:18variance we should briefly review this
- 12:20as a concept remember that variance is
- 12:23the average squared deviation or the
- 12:27average squared difference same thing of
- 12:30a data point from the distribution mean
- 12:34so we take the distance of each data
- 12:36point from the mean square that distance
- 12:40add those together and then find the
- 12:44average that is variance but if we take
- 12:48out that last step if we take out the
- 12:50find the averages part we are left with
- 12:53just the sum of the
- 12:56squares so we would take the distance of
- 12:58each data point from the mean Square
- 13:01each distance and then add them together
- 13:04if we stop there that is the sum of
- 13:08squares so the sum of squares is
- 13:11variance without finding the average of
- 13:14the sum of the square deviations so it's
- 13:17the variance without that last step and
- 13:19that's because sum of squares is a
- 13:22foundational component of an
- 13:26NOA so remember the formula for the same
- 13:29sample variance that's what we just
- 13:30talked about it is the sum of the squar
- 13:33deviations divid the sample size minus
- 13:36one in the sample case so on the top we
- 13:39have the squared differences so remember
- 13:42X is a data point minus mu which is the
- 13:44mean we take that difference and we
- 13:46Square it we sum all those up and then
- 13:50in this case we divide by n minus
- 13:53one so on the bottom we're doing the
- 13:56averaging of the squared differences
- 14:00what if we take away the N minus one
- 14:02part well we're just left with the sum
- 14:05of the squares so each data point minus
- 14:08the mean Square it do that for all the
- 14:11data points add them up and that is the
- 14:14sum of squares so with the sum of
- 14:17squares of the difference between the
- 14:19dependent variable and its mean that's
- 14:22all the sum of squares
- 14:25is now when we talk about sum of squares
- 14:28in the Inova context what we do is we
- 14:31say this overall sum of squares is
- 14:34partitioned or split into two parts so
- 14:38we call the total sum of squares SST
- 14:42that's sum of squares total now that's
- 14:45made up of two components the first
- 14:48component is the
- 14:50SSC that is the sum of squares of the
- 14:53columns so the columns the between
- 14:56variance the treatment sum of squares
- 14:59that's what we're talking about and
- 15:00again this will make sense more than a
- 15:01minute when we look at the actual
- 15:03problem now the other component is the
- 15:05SS e which is the sum of squares error
- 15:10so this is the within or the error sum
- 15:13of squares so the overall or total sum
- 15:16of squares in the oneway an NOA is
- 15:19actually a combination of two things the
- 15:22sum of squares of the columns sort of
- 15:24between the columns and the sum of
- 15:27squares of the error which is actually
- 15:29the sum of squares within each
- 15:33column so let's look at each one of
- 15:35these types of sum of squares using our
- 15:37actual data so we'll start with SST or
- 15:41the sum of squares total so what we do
- 15:44is we find the difference between each
- 15:46data point so all 21 data points and the
- 15:50overall mean over here on the right hand
- 15:52side of 74.5 2 we would then square that
- 15:57difference and then add all of those up
- 16:02so we would have 21 squared differences
- 16:06so here I only circled five but we would
- 16:08do this with all 21 so it's the
- 16:11difference between each data point and
- 16:14the overall mean Square it and then add
- 16:18them all up that is the sum of squares
- 16:24total so if we look at this in the
- 16:27actual distribution we have all 21 of
- 16:30our data points put together and then we
- 16:33have the overall mean of 74.5 2 in the
- 16:37distribution there so what we're doing
- 16:40is we are finding the distance of each
- 16:42data point to the overall mean and then
- 16:46of course we Square it so we would have
- 16:4921 squar deviations now here's the thing
- 16:53when we take a distance and we Square it
- 16:56guess what we get
- 16:59we literally get a
- 17:01square so when we say sum of
- 17:05squares we actually mean that literally
- 17:09so we take every distance from the point
- 17:11to the mean we Square it so we would
- 17:14have 21 individual Square distances on
- 17:18this graph if we actually did it and
- 17:20then we sum up the area of all those
- 17:24squares now I only mention that because
- 17:27it's going to become an important part
- 17:28later on when we talk about regression
- 17:32because sum of squares is a very
- 17:34important part of regression as well so
- 17:37sum of squares is literally squar
- 17:39distances when added together or the
- 17:42bunch of squares added
- 17:46up now let's look at SSC which is the
- 17:50sum of squares of the columns so this is
- 17:53the column or the between sum of squares
- 17:58so in this case we find the difference
- 17:59between each group mean and the overall
- 18:04mean Square those deviations and add
- 18:08them up so the
- 18:10SSC is the difference between the
- 18:12individual column means and the overall
- 18:15mean so in this case we'll have three so
- 18:18each individual column mean and the
- 18:21overall mean so remember the
- 18:24SST or the sum of squares total was the
- 18:28relationship between each individual
- 18:29data point and the overall mean well the
- 18:33SSC or the sum of squared columns is the
- 18:36relationship between each column mean
- 18:40and the overall
- 18:42mean so that would look like this so the
- 18:46SSC is the distance from each mean to
- 18:50the overall mean so you can see that
- 18:52there in the red and again in this case
- 18:55we would literally have squares so the
- 18:57squared distance with the square
- 19:00deviation add them all
- 19:04up now the third type is the
- 19:07s that is a sum of square error or the
- 19:11within sum of
- 19:13squares so this is actually about the
- 19:15individual distribution around each
- 19:19column mean we're not dealing with the
- 19:22overall mean in this case so we would
- 19:25find the difference between each data
- 19:27point and its own column
- 19:31mean Square each
- 19:34deviation and then add them up so for
- 19:37example we would have 21 Square
- 19:39deviations in this case because we would
- 19:42take each column
- 19:44score and then find the relationship
- 19:46between the overall column mean so the
- 19:50difference between each data point and
- 19:52its corresponding column mean find the
- 19:54difference Square it add them up and
- 19:57then we do the same thing for each
- 20:00column that is SS e or the sum of
- 20:04squares
- 20:06within so what we're looking at here is
- 20:09the distribution within each column so
- 20:11where each data point Falls within each
- 20:16column so this is what these three types
- 20:19of sum of squares look like when we put
- 20:21them all together so in the first case
- 20:24we have SST remember that is each
- 20:27individual score so all 21 scores as
- 20:30they relate to the overall mean so we
- 20:32find the difference between each score
- 20:34the overall mean Square it add them
- 20:37up the SSC is a relationship between
- 20:41each column mean there at the bottom and
- 20:44the overall mean so we find the
- 20:47differences Square them add them
- 20:50up then the
- 20:52SSE is the relationship between each
- 20:55individual score and its own column mean
- 20:59at the bottom so we find the difference
- 21:01between those two square it add them all
- 21:04up so you can see the three different
- 21:06ways we're looking at the sum of squares
- 21:09and the cool thing is is that the SSC
- 21:12the sum of squares columns plus the SS
- 21:15sum of square error adds up to the SST
- 21:21or the total sum of squares so you can
- 21:24see all the ways the data is sort of
- 21:27parsed when we do the oneway
- 21:32Anova so this is an exaggerated view of
- 21:35the relationship between our three
- 21:36column means so you can see that group
- 21:39one or the first year students had a
- 21:41mean of 71.7 one the year 2 students had
- 21:44a mean of 75.2 n there in the pink and
- 21:47the year three students had a mean of
- 21:4976.5 s there in the green and I've kind
- 21:52of exaggerated the distance of group
- 21:56one so you can see that each
- 21:59distribution or each column has its own
- 22:01distance to the overall mean so what
- 22:04we're trying to figure out here is is
- 22:07the firste student distribution those
- 22:09scores kind of an outlier or sort of an
- 22:12oddball distribution that's what we're
- 22:15checking for when we're doing the oneway
- 22:18in NOA now remember in the previous
- 22:20video I talked about why we cannot do
- 22:23three pairwise tee tests so we could
- 22:27compare grou group year 1 to year 2 year
- 22:301 to year 3 and year 3 to year 2 that
- 22:34would be three individual T tests we
- 22:36cannot do it that way because the error
- 22:39the type 1 error compounds each time we
- 22:42do that and we end up with a type 1
- 22:43error rate of over 14% so we have to do
- 22:47the Innova method when looking at this
- 22:49type of
- 22:52problem okay so that ends part one of
- 22:55our video on the oneway in Nova so again
- 22:58I I want you to show how the data is
- 23:00actually arranged when we actually do
- 23:02the calculations so we're looking at
- 23:05many comparisons within our chart so
- 23:07we're comparing the column mean to the
- 23:09overall mean we're comparing each data
- 23:11point to the overall mean and we're
- 23:13comparing each data point in a column
- 23:15with that columns overall mean and again
- 23:18it's all about variance so the square
- 23:22deviation is a measure of variance so
- 23:26the sum of squares is a way of
- 23:27quantifying that variance so now that we
- 23:30have a comprehensive understanding of
- 23:31what the oneway an NOA is and how the
- 23:34data points relate to each other let's
- 23:36go ahead go into Excel and do this
- 23:39calculation by hand now I have some
- 23:42pre-made formulas that we can actually
- 23:43just sort of paste into Excel but I will
- 23:46walk you through the actual computation
- 23:49of all those arrows and diagrams we had
- 23:51in the previous slides so let's go ahead
- 23:53and look at part two
- 23:57[Music]
About this transcript
This page contains the full transcript of YouTube transcript (JgMFhKi6f6Y) , generated from the public captions YouTube serves with the video. The transcript has 3,528 words across 496 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.