YouTube transcript (nkmzsyLg0tY) — Transcript
Full transcript
- 0:00[Music]
- 0:07[Music]
- 0:18Hello, thank you for watching and
- 0:20welcome to the next video in my series
- 0:22on basic statistics. Now, as usual, a
- 0:25few things before we get started. Number
- 0:27one, if you're watching this video
- 0:29because you are struggling in a class
- 0:30right now, I want you to stay positive
- 0:32and keep your head up. If you're
- 0:34watching this, it means you've
- 0:35accomplished quite a bit already. You're
- 0:38very smart and talented and you may have
- 0:39just hit a temporary rough patch. Now, I
- 0:43know with the right amount of hard work,
- 0:45practice, and patience, you can get
- 0:47through it. I have faith in you. Many
- 0:50other people around you have faith in
- 0:51you. So, so should you. Number two,
- 0:55please feel free to follow me here on
- 0:57YouTube, on Twitter, on Google+ or on
- 1:00LinkedIn. That way, when I upload a new
- 1:03video, you know about it. And it's
- 1:05always nice for me to connect with
- 1:07people who watch my videos online,
- 1:09wherever in the world you may happen to
- 1:11be. Now, on the topic of the video, if
- 1:14you like it, please give it a thumbs up,
- 1:17share it with classmates or colleagues,
- 1:18or put it on a playlist, cuz that does
- 1:21encourage me to keep making them for
- 1:22you. On the flip side, if you think
- 1:25there is something I can do better,
- 1:27please leave a constructive comment
- 1:28below the video, and I will try to take
- 1:30those ideas into account when I make new
- 1:32ones for you. And finally, just keep in
- 1:35mind that these videos are meant for
- 1:37individuals who are relatively new to
- 1:40stats. So, I'm just going over basic
- 1:42concepts and I will be doing so in a
- 1:45very slow, deliberate manner. Not only
- 1:49do I want you to know what's going on,
- 1:51but also why and how to apply it. So,
- 1:54all that being said, let's go ahead and
- 1:56get started.
- 2:00Okay, so this is the beginning of part
- 2:02two on our video about type one and type
- 2:05two errors. So the first part was really
- 2:07a conceptual background using some
- 2:09non-statistical real world examples that
- 2:12you can then take and then apply to the
- 2:14more statisticalbased examples in this
- 2:17part of the video. So, if you're still
- 2:19unsure about sort of the basic
- 2:21background of type 1 and type two error,
- 2:24please go back and watch that video
- 2:26before proceeding with this one because
- 2:28I think it really sets the stage for
- 2:30understanding the statistical side of
- 2:32type 1 and type two error using some
- 2:34everyday experiences we talked about in
- 2:37the previous one. So, all that being
- 2:39said, let's go ahead and dive right into
- 2:40these examples. Now, these are adapted
- 2:43from my previous video on null and
- 2:45alternative hypothesis. So, we'll just
- 2:46kind of tweak them a bit and then use
- 2:48them to talk about type one and type two
- 2:51error. So, remember our first example
- 2:53was a bottled water manufacturer that
- 2:56states on its product label that each
- 2:59bottle contains 355ml of water. So, you
- 3:03work for a government agency that
- 3:05protects consumers by testing product
- 3:08volumes. So, a sample of 50 bottles is
- 3:12tested. So, you want to make sure that
- 3:15the manufacturer is being truthful on
- 3:18their label and that the amount of water
- 3:20in the bottle is what actually says on
- 3:22the label. So, what can we assume to be
- 3:25true? Well, in cases like this, we
- 3:27assume the label is correct. So, we
- 3:30assume there is 355 ml of water in the
- 3:34bottle. Now, we also decided that this
- 3:36first pair of hypotheses seem to be
- 3:39appropriate because the bottle is
- 3:42stating
- 3:43355 ml. So, that's an equal sign. So,
- 3:48therefore, the alternative is that it's
- 3:51not equal to that. So, our first pair is
- 3:54what we're going to be
- 3:57using. So, we can write our null and
- 4:01alternative hypothesis. So our null
- 4:03hypothesis is that the mean or mu of all
- 4:08the bottles produced is 355 milliliters.
- 4:13So that's our assumption. That's what's
- 4:15on the label. That's what we're assuming
- 4:17to be true. Now the alternative is that
- 4:20the mean or the average for all bottles
- 4:23is not 355 milliliters. So notice again
- 4:28that the null and the alternative are
- 4:32opposites and they account for all
- 4:35possibilities. So the null is that it's
- 4:37equal and the alternative is that it's
- 4:40not equal to
- 4:42355. Now we can set up our little chart
- 4:44here like we did in the previous video.
- 4:46So what we're comparing here is the
- 4:49conclusion of our analysis and the
- 4:52actual state of reality or the actual
- 4:54condition over here on the right. So
- 4:57let's talk about our two conclusions. So
- 5:00our first conclusion is that we cannot
- 5:02or do not reject the null hypothesis. So
- 5:06we take a sample of 50 bottles and they
- 5:09seem to be pretty close. The average is
- 5:11maybe 355.2 or 354.6
- 5:15six or something like that very close to
- 5:17355 ml. Therefore, we do not reject that
- 5:22null
- 5:23hypothesis. But maybe we get a weird
- 5:26sample just by chance and the mean is
- 5:29like
- 5:30340 milliliters, so way underfilled. In
- 5:34that case, we would most likely reject
- 5:36the null hypothesis that states the
- 5:38average bottle volume is 355. So we
- 5:42would reject our null hypothesis and
- 5:45then proceed to the alternative that
- 5:47says the mean volume for bottles is not
- 5:51355. Now this has to correspond with
- 5:54some actual state of reality. So the
- 5:58reality is the bottles overall do have a
- 6:01mean of 355 or they do not have a mean
- 6:05of 355. So we have our conclusion from
- 6:09our analysis and the actual state of
- 6:11affairs, the actual state of the volume
- 6:14in all the bottles. Now two of these
- 6:17generate correct conclusions. So if we
- 6:21take a sample and it's around 355, we
- 6:24will not reject our null. And if the
- 6:28actual state of affairs is that the
- 6:30bottles are around 355, then that is
- 6:33correct. We did not reject our null in
- 6:36our analysis and the bottles are being
- 6:38filled correctly. Therefore, we're
- 6:40correct. Now, maybe the bottles are not
- 6:44being filled correctly. So, we get a
- 6:46sample. It's not close to 355.
- 6:49Therefore, we reject the null
- 6:51hypothesis and the actual state of
- 6:54affairs is that they are indeed not
- 6:57being filled correctly. So if we reject
- 6:59our null and the state of affairs, the
- 7:02actual condition is that it's not being
- 7:05filled correctly, then again we've made
- 7:07a correct conclusion, a correct decision
- 7:09there. Now, of course, we have type one
- 7:12and type two error. Let's look at type
- 7:14one error first. So let's say we get a
- 7:16sample of or 50 bottles and we come up
- 7:21with a mean for that sample of
- 7:25343 millilit a lot lower than
- 7:28355. Therefore we would probably reject
- 7:31our null
- 7:33hypothesis. Now what if our sample is
- 7:37flawed? Maybe we we got a weird sample
- 7:40just by chance. And that's the point of
- 7:42statistics. There's always going to be
- 7:43this case where we get sort of a weird
- 7:45sample that's not representative or
- 7:48something along those lines, but the
- 7:50bottles in actuality overall are being
- 7:54filled correctly. So, we rejected our
- 7:56null hypothesis, but the bottles are
- 8:00being filled correctly. It's just our
- 8:02sample has some problem with it. In that
- 8:04case, we committed a type one error. We
- 8:08rejected the null hypothesis when we
- 8:11should not have and that is classic type
- 8:14one error. Let's talk about type two
- 8:16error. Let's say that we get a sample
- 8:20and it's around
- 8:23355. But in actuality the bottles are
- 8:26not being filled correctly. So we do not
- 8:30reject the null hypothesis but the state
- 8:32of reality is that the bottles are not
- 8:35being filled correctly. And that is a
- 8:37type two error. So we do not reject the
- 8:41null hypothesis when we should have. So
- 8:44it's a failure to not reject the null
- 8:48hypothesis. And that is classic type two
- 8:52error. So again, we're going to walk
- 8:53through two more examples. So just kind
- 8:55of think about this for a second and
- 8:56then apply it as we
- 8:59go. So example two, down on the farm. So
- 9:03according to the United States
- 9:04Department of Agriculture, the
- 9:06USDA, in 2006, the average farm size in
- 9:10the state of Texas was 2.3 km. Now,
- 9:14since the decadel long trend has been
- 9:16for farm sizes to increase due to large
- 9:19agra businesses buying up land and
- 9:21making business bigger farms, a business
- 9:24analyst wishes to test if the current
- 9:272013 farm size is larger than it was in
- 9:322006. So, establish or null and
- 9:34alternative hypothesis first. So, before
- 9:36we do that, what is our assumption? What
- 9:38do we have in the problem to work with?
- 9:41Well, we can only assume based on what
- 9:43we have that there has been no change in
- 9:46farm size since 2006. This is our null
- 9:50hypothesis. That's all we're given in
- 9:53the problem. Now, you might say, well,
- 9:54it says in there that the trend has been
- 9:57for farm size to increase. Well, so
- 10:00what? All we can do in our assumption is
- 10:04test whether or not the farm size has
- 10:08remained 2.3 km or maybe even decreased.
- 10:12And then our research hypothesis, our
- 10:15alternative hypothesis will be that has
- 10:19increased. So we assume there's been no
- 10:22change in farm size. We can set up our
- 10:23hypothesis like this. So our null
- 10:26hypothesis is that the mean farm size is
- 10:29equal to or maybe even less than 2.3
- 10:33square km. That's from our problem. Then
- 10:36our alternative hypothesis is that the
- 10:38farm size has indeed increased. So the
- 10:41average farm size mu is greater than 2.3
- 10:46square kilmters. So we can go up and set
- 10:48up our table here. So again, I'm not
- 10:50going to go into it as much depth as I
- 10:51did in the last one, but we have two
- 10:53conclusions. We can we either do not
- 10:55reject our null hypothesis or we do
- 10:59reject our null hypothesis based on our
- 11:02analysis. Now it has to be one of two
- 11:04actual conditions. Either the farm size
- 11:07has not changed or even decreased or it
- 11:11has
- 11:12increased. If we do not reject the null
- 11:15and the farm size is in fact less than
- 11:18or equal to 2.3 km then we are correct.
- 11:23Now if we reject the null hypothesis so
- 11:26maybe we get a farm size that's 3.7 km
- 11:30and indeed the farm size is greater than
- 11:332.3 km square km then again we are
- 11:37correct. Those are our two correct
- 11:39outcomes. Now let's look at the type one
- 11:42error situation. So in this case we
- 11:44reject our null hypothesis. So maybe we
- 11:48get a sample of farms that are a bit
- 11:50larger than average and therefore our
- 11:54analysis comes up with a larger farm
- 11:57size. Now what if the actual state of
- 12:00affairs the actual condition is that
- 12:02farm size has not increased or maybe
- 12:04even decreased. Well there our
- 12:06conclusion does not match the reality.
- 12:09So this is classic type one error. We
- 12:12incorrectly rejected the null. we
- 12:16wrongly rejected the null hypothesis
- 12:19when we should not have that is type one
- 12:22error. Let's look at type two of course
- 12:25in that case we do not reject the null
- 12:28hypothesis. So maybe we get um a farm
- 12:32size that you know is relatively small
- 12:35or right around 2.3 square kilometers
- 12:39when in fact farm size has increased. We
- 12:43just happened to get a
- 12:46non-representative sample and when we
- 12:48did our analysis we got you know 2.4
- 12:52square kilm or something that was not
- 12:54enough to reject our null hypothesis but
- 12:57in reality farm size has increased. It's
- 13:00just our sample was not representative
- 13:03or something was wrong with our sample
- 13:05and that is of course type two error. So
- 13:09we do not reject the null hypothesis
- 13:12when we should have and the end that is
- 13:15type two. So type one error is rejecting
- 13:20the null hypothesis when we should not
- 13:22have and type two error is not rejecting
- 13:26the null hypothesis when we should
- 13:31have. Okay. Our final example and that
- 13:33is our Manchester United example. So
- 13:37during the 2010 2011 English Premier
- 13:40League season, Manchester United home
- 13:42matches had an average attendance of
- 13:4774,961. So a club marketing analyst
- 13:50would like to see if attendance
- 13:53decreased during the most recent season.
- 13:56So establish our null and alternative
- 13:58hypothesis for this analysis. We can
- 14:00only assume the attendance remain the
- 14:02same or maybe even increased because
- 14:05remember our research question, our
- 14:07alternative hypothesis, our research
- 14:09question is did the attendance decrease?
- 14:13But our assumption is that it remained
- 14:15the same or maybe even
- 14:20increased. So we can set up our
- 14:22hypothesis like this. So our null
- 14:24hypothesis is that the average
- 14:26attendance was greater than or equal to
- 14:3174961. So it was the same or it even
- 14:35increased. Now our alternative
- 14:37hypothesis, the research hypothesis is
- 14:40that the average attendance decreased.
- 14:43So it's less than
- 14:4774961. So we can set up our table again.
- 14:50So we have two conclusions. do not
- 14:52reject the null hypothesis or reject the
- 14:55null hypothesis and then two actual
- 14:58conditions either the attendance stayed
- 15:00the same or increased or did in fact
- 15:02decrease. So if we do not reject the
- 15:05null hypothesis, so we take a so we look
- 15:08at the season's attendance and it was
- 15:10right at you know around
- 15:1374961 or maybe a little bit higher then
- 15:17we our analysis would indicate that
- 15:19we're not going to reject that null
- 15:21hypothesis. It says greater than or
- 15:23equal to and the actual condition is
- 15:25that it actually did remain the same or
- 15:27increase. So that's correct conclusion.
- 15:30Now we could reject the null hypothesis.
- 15:32So we do our analysis of the most recent
- 15:34season and in fact attendance is down to
- 15:37maybe like 68,000 or something like
- 15:39that. In that case we would reject the
- 15:42null hypothesis and then go on to the
- 15:44alternative that says it is less than
- 15:4774961. So if we reject the null
- 15:50hypothesis based on our analysis the
- 15:52actual condition was that it is less
- 15:54than
- 15:5574961 then that case again we are
- 15:58correct.
- 16:00Now let's say we reject our null
- 16:02hypothesis but the attendance actually
- 16:05remained the same or increased. So maybe
- 16:09we this is believe it or not we missed a
- 16:12number. So let's say there were 13 home
- 16:16matches but we only typed 12 into the
- 16:20calculator but then we divided by 13.
- 16:23That could mess up our average right? So
- 16:26we incorrectly reject the null
- 16:29hypothesis when indeed the actual
- 16:32condition is that the attendance
- 16:34remained the same or increased. That
- 16:35does happen. So that would be a type one
- 16:38error. Then of course we could have
- 16:40where we fail to not reject the null
- 16:43hypothesis. So we take a sample or we
- 16:46look at our our data and maybe we
- 16:49actually type in a number twice into our
- 16:51calculator and of course we divide by 13
- 16:55or we divide by one less than is
- 16:57actually there. So therefore our
- 16:58attendance
- 17:00is higher than it actually is. So in
- 17:03that case we do not reject the null when
- 17:05we should have and again that is type
- 17:08two error. So again, I know this can be
- 17:11confusing because we're talking in
- 17:13sometimes double negatives and things
- 17:15like that, but if you really just got to
- 17:16pause the video and look at it, you'll
- 17:18actually see how this sort of works.
- 17:20This follows the same pattern. Type one
- 17:23error is simply the incorrect rejection
- 17:27of the null hypothesis. The type two
- 17:30error is simply not rejecting the null
- 17:34hypothesis when you should have. And
- 17:38that's the difference. So again, look at
- 17:39these examples again, go back to part
- 17:41one, and that can really refresh your
- 17:44mind too as to whether or not you're
- 17:46dealing with type one error or type two
- 17:50error. Okay, so some very simple causes
- 17:53of type one and type two error. So that
- 17:55remember when selecting samples, we are
- 17:57always subject to the laws of chance. We
- 18:00may by random chance alone select a
- 18:03sample that is not representative of the
- 18:06population.
- 18:08So we may select a sample of underfilled
- 18:12or overfilled water bottles just by
- 18:15chance
- 18:16alone. We may select a sample of very
- 18:20small or very large farms again just by
- 18:24chance in our sample
- 18:26selection. Or the sample is in the far
- 18:29out tails of the sampling distribution.
- 18:32again just by chance alone cuz remember
- 18:36the sample means have their own sampling
- 18:40distribution. We talked about that
- 18:42several uh videos ago and we may by
- 18:45chance just get a sample that's way out
- 18:48in the tails of the sampling
- 18:50distribution again just by chance. Now
- 18:53our sampling techniques may be flawed.
- 18:56So there's a whole branch of statistics
- 18:58that talks about sampling techniques and
- 19:01things like that. So we may have to look
- 19:03at our sampling technique. The
- 19:05assumptions in our null hypothesis may
- 19:08be flawed. So in the farm case, maybe
- 19:11the USDA data is incorrect or it has
- 19:14flaws in it. So we're using this USDA
- 19:17data for our null hypothesis, but it
- 19:21doesn't actually correspond to the state
- 19:23of the farm size in Texas. So whatever
- 19:27we're basing our null hypothesis off of
- 19:29may be incorrect.
- 19:31But overall the most common cause is
- 19:35chance and chance alone because again
- 19:38our sampling means have a sampling
- 19:41distribution. We talked about that in
- 19:43the previous videos. And there is a
- 19:46chance that we get one a sample mean
- 19:49that's just far out in the tails of the
- 19:53sampling distribution. That has to deal
- 19:55with confidence intervals and all kinds
- 19:57of things like that. Now, of course,
- 19:59when we actually do hypothesis tests in
- 20:02upcoming videos using actual data and
- 20:04numbers and curves and and things like
- 20:06that, we will deal with that. But the
- 20:09most common cause of type one and type
- 20:12two errors is chance and chance
- 20:16alone. Now, remember that this sort of
- 20:19conclusion table always holds up when
- 20:21we're dealing with these two
- 20:23diametrically opposed hypothesis. So our
- 20:26conclusion is either we do not reject
- 20:27the null or we reject the null. The
- 20:30actual condition is that the null is
- 20:32true or I say true in quotes or the
- 20:36alternative is true. If we do not reject
- 20:38the null and the null is true, that's
- 20:43correct. If we reject the null and the
- 20:46alternative is true, that's correct. Now
- 20:50if we reject the null but the null is
- 20:53true then that's type one error. We
- 20:56incorrectly rejected the null. Now if we
- 20:59do not reject the null and we assume it
- 21:01holds up but the alternative hypothesis
- 21:05is true then we've committed type two
- 21:08error. So in the previous part of this I
- 21:11talked about type one error as sometimes
- 21:14being like the false alarm. So we reject
- 21:18the null hypothesis but in fact the null
- 21:21is true. So for type two we do not
- 21:24reject the null when we should have
- 21:27because the alternative is true. So
- 21:30again go back and look at the previous
- 21:31video and these examples and it really
- 21:33should click sort of in your
- 21:37mind. Okay. So that wraps up part two of
- 21:41our video on hypothesis formulation type
- 21:44one and type two errors. Now, I know
- 21:46this can be a very confusing concept
- 21:48because the way the wording works, we're
- 21:50using a lot of double negatives uh here
- 21:52and there, but I really think if you go
- 21:54back and look at part one and again part
- 21:56two, you really grasp the fundamental
- 21:58concepts of type one and type two error.
- 22:02It's simply when our conclusion based on
- 22:05our analysis does not match the actual
- 22:09state of reality, the state of affairs.
- 22:13So if our our null hypothesis is not
- 22:17rejected therefore the null should be
- 22:20true. If we reject the null then the
- 22:24alternative should be true. And if we
- 22:26don't do that that's where we involve
- 22:29type one and type two errors. Okay. So
- 22:32again that wraps up this entire video on
- 22:35type one and type two errors. Just a few
- 22:37reminders before we wrap up. If you're
- 22:39watching the video because you are
- 22:40struggling in a class, I want you to
- 22:42stay positive and keep your head up.
- 22:44You're very smart and talented and you
- 22:46may have just hit a temporary rough
- 22:47patch. I know you're smart. Everyone
- 22:50around you knows you're smart and
- 22:51talented, so so should you. Please feel
- 22:54free to follow me on YouTube, on
- 22:55Twitter, on Google+, or on LinkedIn.
- 22:58That way, when I upload a new video, you
- 23:00know about it. And it's always nice for
- 23:02me to connect with people who watch my
- 23:04videos online, wherever in the world you
- 23:06may happen to be. If you like the video,
- 23:09please give it a thumbs up, share it
- 23:11with classmates or colleagues, or put it
- 23:13on a playlist because that does
- 23:14encourage me to keep making them for
- 23:16you. On the flip side, if you think
- 23:18there is something I can do better,
- 23:20please leave a constructive comment
- 23:21below the video and I will try to take
- 23:23those ideas into account when I make new
- 23:26ones. And finally, just keep in mind
- 23:28that the fact that you're on here trying
- 23:30to learn, putting the effort in to
- 23:32improve yourself as a student or as an
- 23:34employee, that's what really matters. I
- 23:37firmly believe that if you have the
- 23:39right learning process in place, the
- 23:41results will take care of themselves.
- 23:44So, thank you very much for watching. I
- 23:46wish you the best of luck in your
- 23:47studies and in your work. And look
- 23:49forward to seeing you again next time
- 23:51when we talk about actual hypothesis
- 23:54testing.
About this transcript
This page contains the full transcript of YouTube transcript (nkmzsyLg0tY) , generated from the public captions YouTube serves with the video. The transcript has 3,639 words across 525 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.