YouTube transcript (0Vj2V2qRU10) — Transcript
Full transcript
- 0:02[Music]
- 0:17hello thanks for watching and welcome to
- 0:20the next video in my series on basic
- 0:22statistics now as usual a few things
- 0:24before we get started number one if
- 0:27you're watching this video because you
- 0:28are struggling in a class right now I
- 0:31want you to stay positive and keep your
- 0:33head up if you're watching this it means
- 0:35you've accomplished quite a bit already
- 0:37you're very smart and talented but you
- 0:39may have just hit a temporary rough
- 0:41patch now I know with the right amount
- 0:43of hard work practice and patience you
- 0:46can work through it I have faith in you
- 0:49many other people around you have faith
- 0:51in you so so should you number two
- 0:55please feel free to follow me here on
- 0:57YouTube on Twitter on Google+ or on
- 1:01LinkedIn that way when I upload a new
- 1:03video you know about it and it's always
- 1:05nice to connect with my viewers online I
- 1:08feel that life is much too short and the
- 1:10world is much too large for us to miss
- 1:12the chance to connect when we can number
- 1:16three if you like the video please give
- 1:18it a thumbs up share it with classmates
- 1:21or colleagues or put it on a playlist
- 1:23that does encourage me to keep making
- 1:25them for you on the flip side if you
- 1:27think there's something I can do better
- 1:29please leave a instructive comment below
- 1:30the video and I will take those ideas
- 1:33into account when I make new ones and
- 1:36finally just keep in mind that these
- 1:37videos are meant for individuals who are
- 1:40relatively new to Stats so I'm just
- 1:42going over basic concepts and I will be
- 1:45doing so in a slow deliberate manner not
- 1:49only do I want you to know what is going
- 1:51on but also why and how to apply it so
- 1:56all that being said let's go ahead and
- 1:58get started
- 2:01so this is the first video about a new
- 2:04topic the analysis of variants or more
- 2:07commonly known as Anova in our last set
- 2:10of videos we learned how to compare the
- 2:13variances of two populations using the F
- 2:16ratio and F distribution in the context
- 2:20of Anova it is important to remember
- 2:23that the F ratio is simply a ratio of
- 2:26two variances F ratios are a central
- 2:30part of Anova so if you are unsure what
- 2:33the f ratio is you may want to go back
- 2:36and watch those
- 2:37videos Anova allows us to move Beyond
- 2:41comparing just two populations with
- 2:44Anova we can compare multiple
- 2:47populations and even subgroups of those
- 2:50populations we can investigate how two
- 2:53groups interact with each other
- 2:56quantitatively many experimental
- 2:58research designs use an NOA for these
- 3:01very reasons now in this video we will
- 3:04not be doing any calculations looking at
- 3:07any formulas or testing any hypothesis
- 3:11many students I've worked with over the
- 3:12years learn a Nova in class without
- 3:16actually knowing what it is or why it
- 3:19even exists so this video offers a solid
- 3:23conceptual Foundation about an NOA using
- 3:26illustrations and Graphics so if you are
- 3:29new to Anova
- 3:30or are still trying to figure out
- 3:32exactly what it is this video is for you
- 3:36so sit back relax and let's go ahead and
- 3:39get to
- 3:42work so the first and most obvious
- 3:45question is well why an Nova so up to
- 3:49this point we have been comparing two
- 3:52populations so the independent samples T
- 3:54Test which are two random samples We
- 3:57compare or the Matched sample T Test
- 4:00where each measurement is maybe the same
- 4:02person or the same machine something
- 4:04like that but of course limiting
- 4:06ourselves to the comparison of two
- 4:09populations as well limiting the world
- 4:12is much more complex than just two
- 4:15things what if we wish to compare the
- 4:17means of more than two
- 4:19populations what if we wish to compare
- 4:22populations each containing several
- 4:25sublevels or groups well enter another
- 4:29NOA so an NOA the acronym comes from the
- 4:34phrase analysis of variance so a NOA
- 4:38greatly expands what we were able to do
- 4:42in
- 4:45statistics so suppose we want to compare
- 4:48three sample means to see if a
- 4:51difference exists somewhere among them
- 4:55so our first sample mean is up here in
- 4:56the blue xar sub one xar bar sub two as
- 5:00our second sample mean here in the pink
- 5:02distribution and then our third one is
- 5:04down here in the green so each sample
- 5:08will have its mean and its own
- 5:12distribution so what we are asking is do
- 5:15all three of these means come from a
- 5:19common
- 5:21population so is one mean so far away
- 5:25from the other two that it is likely not
- 5:28from the same population as those other
- 5:31two or are all three so far apart that
- 5:36they all likely come from unique
- 5:39populations so you can see we're talking
- 5:41about sort of the relative distance
- 5:43between these means now the variance of
- 5:47these distributions is also important
- 5:49we'll talk about that later but for
- 5:51right now let's focus on the relative
- 5:53distance between these means and whether
- 5:57or not we could conclude that they come
- 5:59come from the same overall
- 6:04population so here is our first mean so
- 6:07xar sub one with its distribution xar
- 6:10sub two with its distribution and then
- 6:13xar sub3 with its
- 6:16distribution now let's say we take all
- 6:19the data points in all three of those
- 6:22samples and we put all those data points
- 6:24into a common larger distribution so
- 6:28we'll put that by Behind these three now
- 6:31we're asking oursel is where is each
- 6:34mean relative to the overall data set
- 6:38sort of in the background you can see
- 6:40that xar sub one Falls pretty much right
- 6:43down the middle xar sub 2 that mean is a
- 6:47bit to the right so you can see the red
- 6:49arrow denotes how far it is away from
- 6:52the mean of the larger sort of combined
- 6:55population what about xar step 3 well it
- 6:59is a bit to the left so you can see that
- 7:01red arrow denoting how far it is away
- 7:05from the mean of the overall sort of
- 7:07combined
- 7:10population now look at this example so
- 7:13xar sub one is where it was sort of
- 7:15right at the middle xar sub 2 is a bit
- 7:18to the right but now look where xar sub3
- 7:21is our third sample mean it's way over
- 7:25to the left so we might conclude that
- 7:28this mean the third one in the green is
- 7:32too far away from the others it's too
- 7:34far away from the mean of the larger
- 7:36group to be considered as part of that
- 7:39larger population it's kind of the
- 7:42Oddball so xar sub 2 is about this
- 7:44distance you can see the red arrow but
- 7:47xar sub3 is way over to the left so is
- 7:51this sort of the Oddball distribution
- 7:54sort of the weird one sort of the one
- 7:56that doesn't belong in the same
- 7:58population as as the other
- 8:02two now look at this case here we have
- 8:05xar sub one that's pretty much right
- 8:06down the middle again but now xar sub 2
- 8:09the second sample mean is way over to
- 8:12the right so you can see that distance
- 8:14is pretty far away and look at xar sub
- 8:17three its way over to the left so we
- 8:21might conclude that each one of these
- 8:24sample means belongs to its own
- 8:27population only X X bar sub One belongs
- 8:31to this distribution in the background
- 8:33xar sub2 May belong to one that's off to
- 8:35the right and xar sub three May belong
- 8:38the one that's off to the left so you
- 8:41can see that the means are in very
- 8:42different locations relative to the
- 8:45overall mean there in the
- 8:50background so the null hypothesis in
- 8:53these type of problems in an noas is
- 8:56whether or not these three sample means
- 8:58come from from the same population now
- 9:01remember the sample mean is a point
- 9:03estimator of the population mean so our
- 9:06no hypothesis is that mu sub 1 = mu sub
- 9:102 = mu sub3 which is another way of
- 9:14expressing that these three means come
- 9:17from the same overall
- 9:21population now remember we're not asking
- 9:23if they are exactly equal we're asking
- 9:26if each mean likely came from the same
- 9:30larger overall
- 9:33population so in an NOA this idea is a
- 9:37very specific and important idea we call
- 9:41this the variability among or between
- 9:44the sample means so each sample mean is
- 9:48a certain distance from the mean of the
- 9:51overall population in the background and
- 9:54we know that that is an expression of
- 9:56variance sort of the distance of the
- 9:59sample mean from the overall mean in the
- 10:01back so this variability between the
- 10:05sample means is something I really want
- 10:07you to keep in your mind as we proceed
- 10:10this is between
- 10:14variants so we could test all of these
- 10:17sample means using pairwise T tests so
- 10:21here are our three sample means so xar
- 10:24sub 1 xar sub 2 and xar sub3 now we'll
- 10:28block out this part of this little
- 10:30Matrix here because in this triangle we
- 10:33have comparisons to themselves and
- 10:36repeat comparisons so we're only going
- 10:38to have three that we could actually do
- 10:40so in this first one we could compare
- 10:43mean one here in the blue to mean 2
- 10:47which is there in the pink so our null
- 10:49hypothesis would be xar sub 1 = xar sub
- 10:532 we could have a t test that tests that
- 10:56now we could also compare mean one and
- 10:58mean mean 3 so we would have a t test of
- 11:01xar sub 1al xar
- 11:04sub3 now we could also test mean 2 and
- 11:07mean three so xar sub 2 = xar sub3 now
- 11:13notice each one of those independent
- 11:15tests has its own Alpha level so Alpha
- 11:19Point 05 05 and
- 11:2305 now the problem with doing all of
- 11:26these pairwise comparisons is that the
- 11:29Alpha level is the type one error rate
- 11:31of course which means 95% confidence but
- 11:35the error compounds with each T Test so
- 11:40if we compared each possible pair we
- 11:43would have .95 * .95 * .95 so our 95%
- 11:49confidence is now 857 or
- 11:5585.7% so of course our Alpha is 1 minus
- 11:59that so 1 - 857 equal. 143 well what is
- 12:05that1
- 12:06143 well that is our overall Alpha level
- 12:11that is our overall type one error rate
- 12:15so our type 1 error rate went from
- 12:175% to
- 12:2114.3% that is why we do not conduct T
- 12:23tests for every possible pair of means
- 12:27the error rate compounds and therefore
- 12:30the test of course has
- 12:35problems now what's different so if I
- 12:38stretch each one of these sample
- 12:40distributions out what
- 12:42changes well the spread or the variance
- 12:46of each
- 12:48distribution so we call this the
- 12:50variability around or within the
- 12:55distributions so remember before we
- 12:57talked about the variance between the
- 13:00distributions and that was sort of the
- 13:02distance of each mean from the overall
- 13:04population in the background so that was
- 13:07variability between this is variability
- 13:11within So within each sample
- 13:17distribution so at its heart a Nova is
- 13:20really a variability ratio it is a ratio
- 13:24of the first type of variance we saw the
- 13:26variability between the means
- 13:29over the variability within the
- 13:32distributions that's the second type we
- 13:34saw when we stretched each sample
- 13:36distribution out so it's just a ratio it
- 13:40is between variance divided by Within
- 13:45variance so remember in the first type
- 13:48we had an overall mean and then the
- 13:50distance of each sample mean from that
- 13:53so you can see that represented on the
- 13:55top and in the bottom we had the
- 13:57variability with within the
- 13:59distributions so their width or spread
- 14:03side to side so on the top we're talking
- 14:06about distance from the overall mean and
- 14:09on the bottom we're talking about each
- 14:11one's spread or width within its own
- 14:15sample so distance from the overall mean
- 14:18divided by sort of the internal
- 14:22spread so we reduce this to a very
- 14:25common fraction it is the very
- 14:29between divided by the variance within
- 14:33so remember this is a ratio of variances
- 14:37so of course the F distribution will
- 14:39come into play here in a bit so just
- 14:42keep in mind this sort of visual tool on
- 14:45the top the distance from the overall
- 14:47mean and in the bottom the internal
- 14:50spread of each sample
- 14:55distribution so variance between divid
- 14:58by variance within now if we put those
- 15:01together those are the components of the
- 15:04total variance so for an anova we have
- 15:09the total variance sort of split into
- 15:11two parts we have the variance between
- 15:14the means and then the variance within
- 15:17each means distribution so we put those
- 15:20together and we have the total variance
- 15:23for the entire data set so this is
- 15:27called partitioning so we're separating
- 15:30this total variance into its two
- 15:32component parts and again this is what's
- 15:35called a one-way in Nova and I'm not
- 15:37going to go into that right now that's
- 15:38for the next video but in this type of
- 15:41problem sort of the basic problem we
- 15:43have total variance and it's made up of
- 15:45variance between the means and the
- 15:47variance within each
- 15:50sample so here are sort of the nuts and
- 15:52bolts summary of everything we have here
- 15:55if the variability between the means
- 15:58sort of the distance from the overall
- 15:59mean we saw that in the red before in
- 16:02that numerator is relatively large
- 16:06compared to the variance within the
- 16:09samples So within the actual individual
- 16:12samples to spread in that
- 16:14denominator then this ratio will be much
- 16:19larger than one because the variance
- 16:22between will be quite large relative to
- 16:25the variance within that will make this
- 16:27ratio much larger than one if that's the
- 16:31case then the samples most likely do not
- 16:35come from a common population so then we
- 16:39would reject the null hypothesis that
- 16:42these means are equal or come from the
- 16:45same population so if that's the case we
- 16:49may have one Oddball distribution out to
- 16:51the side or all three of them may be so
- 16:54far apart that it would create a very
- 16:57high ratio because the variance between
- 17:00would be much higher than the variance
- 17:05within so what could this look like sort
- 17:08of in our test in general so if the
- 17:10variance between is large relative to
- 17:13the variance within being small we would
- 17:16reject that null hypothesis that says
- 17:18all the means are equal so at least one
- 17:22mean is an outlier sort of off to the
- 17:24side or they may all three be spread far
- 17:27apart and each distribution is
- 17:29relatively narrow so they don't sort of
- 17:31melt together they are distinct so think
- 17:34of this as three distinct distributions
- 17:37that are far apart or one that is sort
- 17:40of by itself off to the side that would
- 17:43create a large variance between the
- 17:47means now if the between variance and
- 17:50the within variances are similar then we
- 17:53would probably fail to reject that null
- 17:55hypothesis so they would appear to be
- 17:58equal or from the same overall
- 18:00population so in this case the means may
- 18:02be fairly close to each other fairly
- 18:05close to that overall mean and or the
- 18:08distributions overlap a bit so they may
- 18:11be a bit harder to distinguish from each
- 18:13other so if they're very close together
- 18:16or the variances within is very wide
- 18:19then they'll sort of melt together and
- 18:21therefore they will not be
- 18:24distinct now the other case is where the
- 18:26between variance is very small
- 18:29and the variance within is very large so
- 18:32you can think of this as three
- 18:34distributions that are very spread out
- 18:37internally and they not have a whole lot
- 18:39of distance from each other so they may
- 18:43be close together and or the
- 18:44distributions because of very high
- 18:46variation they're sort of wide and
- 18:48spread out they sort of melt together
- 18:51and you really can't distinguish them
- 18:53from each other as coming from separate
- 18:56or distinct populations so again this is
- 18:59a very general look at how we would
- 19:01interpret an an NOA and it obviously
- 19:04gets more complicated than this but I
- 19:06just want to give you a very rough
- 19:10overview so here's what we're left with
- 19:13this F ratio remember the F ratio is a
- 19:16ratio of two variances so we have the
- 19:20between variance so again the distance
- 19:23of each mean from the overall mean or
- 19:26the combined population in the
- 19:28background as compared to the variance
- 19:32within So within each sample
- 19:36distribution now sometimes this is
- 19:38called the among variance and the around
- 19:43variance so the thing about an noas is
- 19:46that I could have three different or
- 19:48four different stats textbooks and they
- 19:51all call this something different but
- 19:53the most common way of expressing it is
- 19:55the between variance and the within
- 19:57variance but you may see it as the among
- 20:00variance the the variance among the
- 20:03means and the around variance which is
- 20:06the variance around each sample mean but
- 20:10it means the same
- 20:12thing so this is that partitioning so
- 20:15the variance between they're in the red
- 20:18arrows plus the variance within they're
- 20:21in the blue arrows adds up to the total
- 20:26variance now you will also see a term
- 20:29called the error or the error variance
- 20:33well the error variance is another name
- 20:35for the within there in the blue or the
- 20:38around so again you might see that in
- 20:40the textbook and that's one of the
- 20:42challenges of doing these type of videos
- 20:44is because stats books like to give
- 20:46different names to the exact same
- 20:53thing okay so remember why and Nova so
- 20:57up to this point we had been been
- 20:58comparing just two populations so the
- 21:00independent samples T Test and the match
- 21:03sample T Test are two examples but
- 21:05limiting ourselves to the comparison of
- 21:07two populations is of course limiting
- 21:10what if we wish to compare the means of
- 21:12more than two populations what if we
- 21:14wish to compare the populations each
- 21:16containing several levels or subgroups
- 21:20well that's what we have an NOA for
- 21:22remember an NOA stands for the analysis
- 21:25of variance now one more thing I want to
- 21:28hit on before we go on to the end slide
- 21:30and wrap up remember that we're looking
- 21:32at the variance between the means as
- 21:35compared to the variance within each
- 21:39sample so it's all about variance that's
- 21:44why it's called analysis of variance
- 21:46variance between as compared to variance
- 21:49within so we're always looking at a
- 21:51ratio of two
- 21:56variances okay so that wraps up our
- 21:59first video on the analysis of variance
- 22:02or an NOA so again I wanted to give you
- 22:04a graphical representation of what's
- 22:06going on in an Nova so when you're
- 22:09working with the table of data and
- 22:11you're working with your numbers and
- 22:12your F ratios and things like that you
- 22:15actually know what is going on in the
- 22:19background what we're really trying to
- 22:21do is find any distinctions between
- 22:24sample means so we may have sample means
- 22:27that sort of all line up therefore we
- 22:29could conclude that those probably come
- 22:31from the same combined population but we
- 22:35might have one of the means that's sort
- 22:37of an oddball out by itself so in that
- 22:40case it probably comes from a different
- 22:43population off to the side or we could
- 22:45have three sample means that are very
- 22:48far apart from each other therefore each
- 22:51sample mean may come from its own
- 22:54population sort of in the background so
- 22:57what we're trying to look look for here
- 22:58is distinctions or differences among
- 23:01several means and we are doing that by
- 23:05looking at two types of variance between
- 23:07variant and within variance so a few
- 23:10last words and then we are done if
- 23:13you're watching this video because
- 23:14you're struggling in class stay positive
- 23:16and keep your head up I have faith in
- 23:19you many other people around you have
- 23:21faith in you so so should you if you
- 23:24like the video please give it a thumbs
- 23:25up share it with classmates or
- 23:28colleagues feel free to follow me here
- 23:30on YouTube on Twitter on go+ or on
- 23:33LinkedIn it's always nice hearing from
- 23:35you and finally just keep in mind that
- 23:37the fact that you're on here trying to
- 23:39learn trying to improve yourself as a
- 23:41student or as a business person that is
- 23:44what really matters I firmly believe if
- 23:46you have the right learning process in
- 23:48place the results will take care of
- 23:51themselves so thank you very much for
- 23:53watching I wish you the best of luck in
- 23:55your work and in your studies and I look
- 23:57forward to seeing you again next
- 24:01[Music]
- 24:16time
About this transcript
This page contains the full transcript of YouTube transcript (0Vj2V2qRU10) , generated from the public captions YouTube serves with the video. The transcript has 3,495 words across 495 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.