YouTube transcript (tsPv-ffN-0M) — Transcript
Full transcript
- 0:09Let's go ahead and look at a couple of
- 0:10examples. So, a report from 6 years ago
- 0:13indicated that the average gross salary
- 0:16for a business analyst was $69,873.
- 0:21Now, since this survey is now outdated,
- 0:24the Bureau of Labor Statistics wishes to
- 0:26test this figure against current
- 0:29salaries to see if the current salaries
- 0:32are statistically different from the old
- 0:34ones. So, the
- 0:37$69,873 will be our presumed our assumed
- 0:41population mean that we're going to test
- 0:44against. Now, based on other studies,
- 0:46we're going to assume a sigma of
- 0:52$13,985. Now, for this study, the BLS
- 0:55will take a sample of 112 current
- 0:59salaries. So, we have all the parts we
- 1:02need to set up our hypothesis. We have
- 1:05the hypothesized population mean. We
- 1:08have our sigma. We have our sample size.
- 1:11And of course once we collect our data,
- 1:13we'll have our sample
- 1:17mean. Now step one is always establish
- 1:20our hypothesis. Now we went ahead and
- 1:22did the other step one which is
- 1:24formulate a good problem. We did that in
- 1:25the previous slide. So actually in the
- 1:27numbers we have step one establish the
- 1:30hypothesis. So remember our null
- 1:33hypothesis is that the current salary
- 1:36mean is the same as the previous one 6
- 1:40years ago. So mu is equal to
- 1:4669,873. Now our alternative is the
- 1:48opposite of that and that is that the
- 1:50current mean salary for business
- 1:53analysts is not $69,873.
- 1:59Now, step two, determine the appropriate
- 2:01statistical test and sampling
- 2:03distribution. Now, this will be a
- 2:05two-tailed test because remember,
- 2:06salaries could be higher or lower
- 2:09because we're still in the middle of a
- 2:12global recession. So, it's very possible
- 2:15that salaries could have gone down. Now,
- 2:17they could have gone up as well. We
- 2:18don't know. So, we're just testing
- 2:20whether or not it's equal to. We don't
- 2:22know on which side it may have gone if
- 2:25it's not equal to. Now, since sigma is
- 2:28known, we will be using the Z
- 2:30distribution as we will in all the
- 2:32examples in this video. So, we're going
- 2:34to go ahead and use the formula we had
- 2:36two slides
- 2:40ago. Now, step three, we're going to
- 2:42specify the type one error rate or the
- 2:45significance level. And again, this is
- 2:47up to us. So, I'm going to choose for
- 2:49this one the middle ground and say an
- 2:51alpha of 0.05.
- 2:54So I am okay with the possibility of
- 2:58making a type one error 5% of the time.
- 3:03Now step four, we're going to state our
- 3:06decision rule. Now remember based on the
- 3:10curve, if our ZV valueue is above
- 3:141.96, it'll be in our top rejection
- 3:17region. Therefore, we'll reject the null
- 3:20hypothesis. If our zstistic is less than
- 3:24negative 1.96 then again we will reject
- 3:27the null hypothesis because that will be
- 3:29in our lower rejection region. Now of
- 3:33course step five we will gather the
- 3:35data. Now in this case we went ahead and
- 3:38gathered our data. So our sample size
- 3:40was 112 and our sample mean was 70
- 3:4872,180. So, we know that it's higher,
- 3:52but the question is, is it high enough
- 3:55to be statistically
- 4:01significant? Let's go ahead and
- 4:03calculate our test statistics. So, our
- 4:05mean salary for the current salaries is
- 4:11$72,180. Our
- 4:13hypothesized mean was
- 4:16$69,873. Now our sigma was given to us
- 4:19at 13985 and of course our sample size
- 4:22is 112. So again we have our formula
- 4:25down here at the bottom. So all we do is
- 4:28we go ahead and insert those numbers
- 4:31into our Z formula. So 72180 which is
- 4:36our sample mean minus
- 4:3969,873 which is our hypothesized mean
- 4:42divided by sigma / the<unk> of n. And
- 4:47that comes up with a value of Z that is
- 4:50equal to
- 4:531.75. Now we have a decision to
- 4:57make. So remember our hypothesis were mu
- 5:01is
- 5:0369,873 or the alternative was mu is not
- 5:0869,873. Now we know it's not exactly
- 5:1269,873 because we found that it was like
- 5:1471,000 something. But the question is,
- 5:17is it high enough to say that it is
- 5:21statistically
- 5:23different? So our Z was
- 5:271.75. And look where that falls in our
- 5:30non-rejection region. It's right there
- 5:33to the left of our critical value. Now
- 5:37since the test statistic is inside the
- 5:39non-rejection region and not beyond the
- 5:42critical value, we fail to reject the
- 5:46null hypothesis that the old and the
- 5:49current salaries are statistically
- 5:52different. So we therefore fail to
- 5:57reject the null hypothesis. So we assume
- 6:01that our assumption holds up. Remember,
- 6:04we're not saying that our null is quote
- 6:07true. All we're saying is that we could
- 6:10not reject it based on the statistical
- 6:13analysis we did. So, is the salary
- 6:17higher based on our sample? Yes. But can
- 6:21we say it is statistically different?
- 6:24No. And remember why is that? That's
- 6:28because of the idea of sampling error.
- 6:32Remember, this was just one sample. We
- 6:35could have taken many samples. And those
- 6:38other samples might be right smack in
- 6:40the middle of the non-rejection region.
- 6:42We might have a sample that's lower in
- 6:45the non-rejection region. They could be
- 6:47anywhere in that blue area. We just
- 6:50happened to get one here. Now remember
- 6:54what we're saying is that 5% of the time
- 6:57we expect to get a sample mean that's
- 7:00either in the upper rejection region or
- 7:02in the lower rejection region. But this
- 7:05one sample just happens to be located
- 7:08right here right at the upper edge of
- 7:11the non-rejection region. So we would
- 7:13conclude that the old salaries and the
- 7:16current salary are not
- 7:19statistically different.
- 7:24So, example two, Starbucks customer
- 7:26satisfaction. So, Starbucks is
- 7:28interested in assessing customer
- 7:31satisfaction in the Canadian city of
- 7:33Toronto, Ontario. To conduct the study,
- 7:36Starbucks asks 225 customers in the city
- 7:41compared to other coffee houses in
- 7:43Toronto. Would you say the customer
- 7:45service at Starbucks is much better than
- 7:47average, which is a score of five?
- 7:49better than average, which is a score of
- 7:51four. Average, a score of three. Worse
- 7:54than average, a score of two. Or much
- 7:57worse than average, a score of one. And
- 8:01this is commonly known as a Lykert
- 8:03scale. So 54321 in descending order. Now
- 8:08based on the data we collected, the mean
- 8:10rating was determined to be 3.25.
- 8:14And based on previous studies done by
- 8:16the company, it is assumed that sigma is
- 8:221.5. So let's go ahead and establish our
- 8:25hypothesis. So our null hypothesis is
- 8:29that the average customer rating is less
- 8:32than or equal to three because remember
- 8:34three is average in our lacquered scale.
- 8:37Three is average. And then our
- 8:39alternative hypothesis is that the
- 8:42satisfaction level is higher than three.
- 8:46So we're going to assume that it's three
- 8:49or less and then we'll either reject or
- 8:52fail to reject that and then we will go
- 8:55on to our alternative because remember
- 8:57the equal sign the equality portion is
- 9:00always in the
- 9:01null. So determine the appropriate
- 9:03statistical test and sampling
- 9:05distribution. Now, this will be a
- 9:07onetailed test. Starbucks is interested
- 9:11in a better than average customer
- 9:14service rating. So, you can see that in
- 9:17our alternative hypothesis, we're
- 9:20interested if the the average customer
- 9:22rating is higher than three. So, because
- 9:26of the greater than and less than or
- 9:28equal to than, this will be a one-
- 9:29tailed test. Now since sigma is known,
- 9:32we will again use the Z distribution as
- 9:34we did
- 9:37before. Now specify the type one error
- 9:41rate. So our significance level. Now for
- 9:44this one, I'm going to choose a 01.
- 9:47Again, just to show you some conceptual
- 9:49information, I'm going to show you a
- 9:50different one or use a different one.
- 9:52Then we'll state the decision rule. So
- 9:55if our ZV valueue is greater than
- 10:002.33, we will reject the null
- 10:03hypothesis. So this is our upper tailed
- 10:05test. Remember with an alpha of 01, we
- 10:08have 99% there in the non-rejection
- 10:10region and we have the 1% all in the
- 10:14upper rejection region. So our Z, our
- 10:17critical value is 2.33, which is right
- 10:21there. Of course, where does that come
- 10:22from? that comes from our Z table. So
- 10:26you just have to look up that critical
- 10:29value in the Z table. Of course, step
- 10:33five will gather our data. Now, we
- 10:34already did that. So our sample size
- 10:36again was 225 and our sample mean was
- 10:453.25. So let's go ahead and calculate
- 10:47our test statistic. So there are four
- 10:50inputs. our sample mean, our
- 10:52hypothesized mean, our sigma, and our
- 10:55sample size. And again, the same
- 10:57formula. So, we'll go ahead and
- 10:59substitute all that information into our
- 11:02equation, and we come up with a ZV
- 11:05value, a Z test value of
- 11:092.5. So, let's go ahead and put that on
- 11:11our
- 11:14curve. So, here is our distribution, and
- 11:17you can see that we have our
- 11:18non-rejection region in the middle.
- 11:20That's 99% or 0.99, that's our
- 11:22probability in the middle. And of
- 11:24course, we have our 1% there on the
- 11:26ends. So remember what we're saying here
- 11:29is that we expect 99% of our sample
- 11:32means to be in this blue region and then
- 11:351% to be in the rejection region in the
- 11:39uh brown color there. So here our
- 11:41hypothesis, our null is that mu is less
- 11:44than or equal to 3 and the alternative
- 11:46is that mu is greater than three.
- 11:49So our Zcritical value is 2.33 which is
- 11:53right there on our curve. Now our Z is
- 11:582.5. So where is it at? It is in the
- 12:03rejection region. Now since the test
- 12:07statistic is inside the rejection region
- 12:10and beyond the critical value, we reject
- 12:14the null hypothesis that customer
- 12:16satisfaction is at or below average. So
- 12:20therefore we have to reject the null
- 12:23hypothesis and accept the alternative
- 12:27hypothesis based on this Z value.
- 12:35Now, what is the mean customer
- 12:37satisfaction value at the critical value
- 12:41of
- 12:422.33? Because remember, up until now,
- 12:44we've been finding zcores. But what is
- 12:47the actual customer sentiment, the
- 12:50average customer satisfaction at that
- 12:54critical
- 12:55value? Well, remember this is the
- 12:57formula we used before. So, it's
- 12:59everything we had in our equation
- 13:00before. Now what we can do is we can
- 13:04just sort of put in the Zcritical value
- 13:07we're looking at. So 2.33 just goes in
- 13:10where Z was. Now we're solving for Xbar.
- 13:16So we want to find the sample that would
- 13:19fall right on the sample mean that would
- 13:21fall right on that Z critical value. So
- 13:24again, this is just where our wonderful
- 13:27basic algebra that we learned many years
- 13:29ago comes into play.
- 13:32So we'll go ahead and solve our
- 13:33denominator. That ends up being 0.1.
- 13:36Then we'll multiply both sides by 0.1.
- 13:39So on the left hand side we have
- 13:42233 equals xar minus 3 which is just the
- 13:46leftover of our
- 13:48numerator. And we come up with a sample
- 13:51mean xbar of 3.233.
- 13:57So therefore any sample of size
- 14:01225 with a sample mean greater than
- 14:073.233 would lead to a rejection of the
- 14:10null hypothesis assuming a constant
- 14:13sigma and the same alpha level. So if we
- 14:17went out again and we collected another
- 14:20sample, same 225 sample size, and let's
- 14:25say we got a mean of
- 14:313.19, what would happen
- 14:33then? Well, it's not greater than
- 14:373.233. Therefore, we would fail to
- 14:41reject the null hypothesis. See how this
- 14:44works? So what we've done is we've set
- 14:46up sort of a threshold mean value or
- 14:50sample mean value and anything above
- 14:53that assuming all this stays the same
- 14:56would lead to a rejection of the null
- 14:58hypothesis. Anything equal to or below
- 15:01that would make would mean we fail to
- 15:04reject the null hypothesis.
About this transcript
This page contains the full transcript of YouTube transcript (tsPv-ffN-0M) , generated from the public captions YouTube serves with the video. The transcript has 1,967 words across 296 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.