YouTube transcript (dDsKP7wVpzM) — Transcript
Full transcript
- 0:01[Music]
- 0:08hello thank you for watching and welcome
- 0:10to the next video in my series on basic
- 0:12statistics now since this is part two of
- 0:15a two-part video I'll keep this intro
- 0:17short but just keep in mind if you're
- 0:19watching this because you're struggling
- 0:21in a class stay positive and keep your
- 0:23head up you're smart and talented and
- 0:26you must have the confidence that you
- 0:27can work through it please feel free to
- 0:30follow me here on YouTube on Twitter on
- 0:32Google+ or on LinkedIn that way when I
- 0:35upload a video you know about it if you
- 0:38like the video please give it a thumbs
- 0:39up share it with classmates or
- 0:42colleagues or put it on a playlist that
- 0:44does encourage me to keep making them
- 0:45for you and just keep in mind that these
- 0:47videos are meant for individuals who are
- 0:49relatively new to Stats so I'm just
- 0:51going over basic concepts so all that
- 0:54being said let's go ahead and get
- 0:56started on part
- 0:58two
- 1:01so remember in part one of this video we
- 1:03talked about the conceptual background
- 1:05of the single sample T Test and we did
- 1:09that through the lens of the single
- 1:11sample Z test which we learned about in
- 1:14previous videos so remember the main
- 1:16difference is that in the Z test we know
- 1:20Sigma or Sigma is given to us and of
- 1:23course that's the population sard
- 1:25deviation in the T Test we are not given
- 1:30Sigma or we do not know it therefore we
- 1:33estimate it using the sample standard
- 1:35deviation so that's the primary
- 1:37difference between the two now from part
- 1:40one we walked away with two conceptual
- 1:44pillars the first one is that the alpha
- 1:47level affects the location of our
- 1:50critical values so if we have an alpha
- 1:52of
- 1:54say10 that's going to be a probability
- 1:57of 09 in the middle of the distrib
- 2:00distribution now if we go to an alpha of
- 2:0405 that's going to be
- 2:0695% probability in the middle so more
- 2:10well because we have more probability in
- 2:12the middle what has to happen to our
- 2:14critical values as far as their location
- 2:17well they have to move outward now let's
- 2:20say we go from an alpha of 05 to an
- 2:23alpha of
- 2:2501 so the probability is now .99 in the
- 2:28middle of the distribution
- 2:30then what has to happen to the location
- 2:32of our critical values well they're
- 2:34going to have to go out even further to
- 2:36accommodate that 99% area or 099
- 2:40probability so the alpha level affects
- 2:44the location of our critical values in
- 2:48our distribution that's the first
- 2:49important point the second important
- 2:52point is the difference between the Z
- 2:54distribution and the T distributions
- 2:57because remember every sample size has
- 3:00its own t distribution but in general we
- 3:04talked about the following
- 3:06idea for any given Z distribution if we
- 3:10take the T distribution of say a
- 3:12moderate sample size like 20 or 25 if we
- 3:16look at the critical values of that Z
- 3:19distribution as compared to the T
- 3:22distribution for a given Alpha what
- 3:24happens to the location of the critical
- 3:26values in the T
- 3:28distribution well again they move
- 3:30outward towards the tails ever so
- 3:33slightly and why is that well remember
- 3:36the T distribution has a little bit less
- 3:38probability in the middle a little bit
- 3:40more in the Tails so if we compare it to
- 3:44the Z distribution the critical values
- 3:47are going to be slightly outward to
- 3:50accommodate that sort of squishiness of
- 3:53the T distribution so these two things
- 3:56affect the location of the critical
- 3:58value the alpha level Lev and the T
- 4:02distribution as far as its sample size
- 4:04goes because in general as the sample
- 4:06size gets smaller the critical values
- 4:09are going to be pushed outward because
- 4:11there is more uncertainty in the
- 4:13sampling distribution now why am I
- 4:16telling you all this I spent three
- 4:17minutes telling you this why well the
- 4:20whole idea of hypothesis testing is
- 4:22about the location of the critical
- 4:24values so if our test statistic is on
- 4:27one side of the critical value or the
- 4:29other
- 4:30that completely changes what our
- 4:32conclusion is so you have to understand
- 4:35what's happening with the location of
- 4:36the critical values in terms of the
- 4:39alpha level and the characteristics of
- 4:42the T distribution because as we'll see
- 4:44coming up here in a few minutes it can
- 4:47really change the outcome of your
- 4:50problem so let's go ahead and take a
- 4:51look at our first
- 4:55example so this is the same problem I
- 4:58used in the Z test video but I changed
- 5:02it to adapt it for the T Test so we're
- 5:04not going to reinvent the wheel we'll
- 5:05just change a few things and use the
- 5:06same problems so let's say a report from
- 5:096 years ago indicated that the average
- 5:11gross salary for a business analyst was
- 5:1469,8 73 now since this survey is now
- 5:18outdated the BLS wishes to test this
- 5:21figure against current salaries to see
- 5:24if the current salaries are
- 5:25statistically different from the old
- 5:28ones so we're testing to see if the
- 5:31current population of salaries is
- 5:34equivalent to the old population of
- 5:37salaries now based on this sample we
- 5:39found a sample standard deviation of
- 5:4214,985 now again we do not know Sigma
- 5:46therefore we'll have to estimate it
- 5:48using this sample standard
- 5:50deviation now for this study the BLS
- 5:52will take a sample of 12 current
- 5:55salaries so look at what we have here we
- 5:58have a hypothesi ized population mean
- 6:01that's
- 6:0269873 we have a sample standard
- 6:05deviation
- 6:0714985 and we have n our sample size we
- 6:12almost have everything we need to
- 6:13conduct our
- 6:15test but of course we're not going to do
- 6:17that yet we're going to set up a proper
- 6:20hypothesis so our null hypothesis is
- 6:23that the current population of salaries
- 6:26is equivalent to the old population of
- 6:29salaries
- 6:30therefore the alternative is that the
- 6:32current population of salaries is not
- 6:35equivalent to the old population of
- 6:38salaries now this would be a two-tailed
- 6:40test the salaries could be higher or
- 6:43they could be lower I mean we're in a
- 6:45recession right now so it could they
- 6:47could be lower now since Sigma is
- 6:49unknown and N is small we'll use the T
- 6:52distribution and there is the T
- 6:54statistic formula over there on the
- 6:56right it looks very similar to the Z
- 6:59statistic
- 7:02formula so specify the type 1 error rate
- 7:05now this is my choice I'm going to
- 7:07choose an alpha of
- 7:0805 I could have selected an alpha 01 or
- 7:120.10 but I choose the middle ground and
- 7:14I'll select an alpha 05 that means I
- 7:17accept the possibility that I might make
- 7:19a type one error 5% of the
- 7:23time so State the decision Rule now
- 7:26remember we are using the T distribution
- 7:28so the location of our critical values
- 7:30and our regions will depend on our
- 7:33sample size and our Alpha so in this
- 7:36case our sample size is 12 therefore our
- 7:39degrees of freedom is 11 so we will
- 7:42consult our T table we'll find the
- 7:44column that is a two-tailed test at an
- 7:47alpha of
- 7:4705 then we'll find the row that is 11°
- 7:51of Freedom we'll find where those
- 7:53intersect and we find that we have a t
- 7:55value of
- 7:582.21 so our critical values are t plus
- 8:02or minus
- 8:062201 then we'll gather our data so we
- 8:09have 12 in our sample and we found out
- 8:12that our sample mean is 79
- 8:16180 so now we have everything we need to
- 8:19go ahead and figure out our test
- 8:24statistic so we have our sample mean of
- 8:2679 1880 our hypothesize pop relation
- 8:29mean of 69873 there's almost a $10,000
- 8:32difference there our sample standard
- 8:35deviation of
- 8:3614985 or sample size of 12 so we can go
- 8:40ahead and substitute all that into our
- 8:42formula so we have 79 180 minus
- 8:4669873 ided 14985 divid the RO of 12 so
- 8:50we just substituted everything in there
- 8:53and that generates a t statistic of
- 8:572.15 so that is our test
- 9:01statistic now we'll go ahead and put
- 9:03that in the context of our sampling
- 9:06distribution so our hypothesis there are
- 9:08the same we have our hypothesized mean
- 9:11they're in the middle our non-rejection
- 9:13region in blue and our rejection region
- 9:15in the brown so where does our test
- 9:17statistic fall well it's 2.15 so it
- 9:21falls right there now what do we notice
- 9:26since the test statistic is in the
- 9:28non-rejection region and not beyond the
- 9:32critical T value we fail to reject the
- 9:35null
- 9:36hypothesis so it's not out of the
- 9:39ordinary that this sample came from a
- 9:42population with a mean population mean
- 9:45of
- 9:4669873 as we
- 9:48hypothesized now it's in the upper area
- 9:52of our non-rejection region but that
- 9:54really doesn't make a difference it's
- 9:55either in that region or it's not so we
- 9:58just just happened to get a sample that
- 10:01fell here now we could have gotten a
- 10:04sample that fell in the same place on
- 10:06the other side so maybe minus
- 10:102.15 or we could have gotten a sample
- 10:12that fell right in the middle or we
- 10:14could have gotten a sample that fell in
- 10:16the rejection region that's the idea of
- 10:19a sampling distribution we expect
- 10:2395% of our samples to be in this blue
- 10:26region and ours just happened to be one
- 10:29of those now had it been a bit higher it
- 10:32would have been outside but that is just
- 10:35the idea of chance our sample happened
- 10:38to fall right there so we failed to
- 10:40reject the null
- 10:42hypothesis and remember our T value was
- 10:452.20 one so it would have had to have
- 10:49been beyond 2.20 one for us to actually
- 10:53reject the null
- 10:57hypothesis now let's do the same same
- 10:59problem but change the sample size so
- 11:02we're going to go from a sample size of
- 11:0412 to a sample size of 15 let's see what
- 11:07happens so everything else is the same
- 11:10sample means the same hypothesize means
- 11:13the same deviation the same just the
- 11:15sample size has changed so we'll go
- 11:17ahead and substitute all that into this
- 11:20formula now we have a t statistic of
- 11:252.41 so that is our test statistic let's
- 11:28go ahead and place it on the
- 11:31curve so here everything is the same so
- 11:34our degrees of freedom are now 14 cuz
- 11:36remember our sample size was 15 so our
- 11:39degrees of freedom is 14 now if we go to
- 11:42our T table we find that the critical T
- 11:44value is plus or minus
- 11:482.45 so that is the critical value there
- 11:50on the bottom now remember that for a
- 11:54degrees of freedom of
- 11:5611 it was plus or minus 2.2
- 11:592011 so when we increased the sample
- 12:02size what happened to the location of
- 12:05the critical T values well they moved
- 12:08inward ever so slightly but they did
- 12:11move Inward and why is that well we'll
- 12:15talk about that here in a second
- 12:16actually so our T value is
- 12:222.41 so look where it falls well it's
- 12:25now in the rejection region so you look
- 12:29over there on the left if T is greater
- 12:31than
- 12:322.45 we reject the null
- 12:36hypothesis so all we did here was change
- 12:39the sample size from 12 to
- 12:4415 and now we came up with a completely
- 12:47different conclusion and the question is
- 12:50why is
- 12:51that
- 12:52well the larger sample size decreased
- 12:57the standard error of the mean so the
- 13:00larger sample size decreased the
- 13:02standard deviation of this
- 13:05distribution it made it
- 13:08narrower so what ended up happening is
- 13:11it made our sample stand further out on
- 13:15its own it made it a little bit more
- 13:18likely to belong to a different
- 13:20population that does not overlap much
- 13:23with this population it created a
- 13:26separation between our sample xbar and
- 13:30our
- 13:31hypothesized sample mean so the larger
- 13:35sample size decreased the standard
- 13:37deviation the standard error of the mean
- 13:40of this distribution so it kind of like
- 13:43sucked in its middle It's Kind like when
- 13:45your pants don't fit so you suck in your
- 13:47belly that's kind of what happened to
- 13:49this distribution when we increase the
- 13:51sample size it pulled it inward towards
- 13:53the middle and that left our poor t
- 13:56statistic standing out there by itself
- 13:59self now also the larger sample size led
- 14:03to a higher degrees of freedom this
- 14:06brought more probability as well in
- 14:09towards the middle of the T distribution
- 14:11around our hypothesized population mean
- 14:15so being inside the non-rejection region
- 14:18this blue area is a bit more exclusive
- 14:21Club so you can think of our our sample
- 14:24mean is kind of like someone outside the
- 14:26bar who cannot afford to get in to hang
- 14:29out with the cool people so two things
- 14:31went on here and that's why we went over
- 14:33all that conceptual information the
- 14:36larger sample size decreased the
- 14:39standard error of this distribution so
- 14:41it kind of sucked in towards the middle
- 14:44also the larger sample size led to a
- 14:46higher degrees of freedom so that
- 14:49brought more probability in towards the
- 14:51middle as well so two things going on
- 14:53there and therefore our sample mean it's
- 14:59the same sample mean but this time it
- 15:01was left outside of the non-rejection
- 15:04region and all we did there was change
- 15:06the sample size from 12 to
- 15:1215 okay so example two Starbucks
- 15:15customer
- 15:16satisfaction so Starbucks is interested
- 15:19in assessing customer satisfaction in
- 15:21the Canadian city of Toronto
- 15:23Ontario to conduct the study Starbucks
- 15:26asked 25 customers in the city the
- 15:29following question compared to other
- 15:31coffee houses in Toronto would you say
- 15:34the customer service at Starbucks is
- 15:36much better than average that's a score
- 15:38of five better than average that's a
- 15:40score of four average a score of three
- 15:44worse than average a score of two or
- 15:47much worse than average a score of one
- 15:50and of course we call that a ler scale
- 15:53in
- 15:54research now we found that the mean
- 15:56rating was determine to be 3
- 15:593.5 also based on this sample the
- 16:02standard deviation was found to be 1.4
- 16:05so our sample mean was 3.5 or sample
- 16:08Center deviation was
- 16:131.4 so it set up our hypothesis so our n
- 16:16hypothesis is that the mean customer
- 16:19sentiment or feeling is three or less so
- 16:24average or lower therefore the
- 16:27alternative hypothesis is that customer
- 16:30sentiment is higher than average or
- 16:33greater than
- 16:35three now this will be a one-tailed test
- 16:38and we can tell that by the signs in our
- 16:42hypothesis and Starbucks remember is
- 16:45interested in a better than average
- 16:47rating so here's what they're kind of
- 16:48trying to do in this test they're
- 16:51setting up a n hypothesis that says oh
- 16:54our customer you know sentiment is
- 16:57average or lower
- 16:59and then we're going to collect some
- 17:00data and we're going to see whether or
- 17:03not our data support the idea that we
- 17:06can reject that null hypothesis if we
- 17:09can reject that null hypothesis then we
- 17:12can proceed to the alternative that says
- 17:14customer sentiment is better than
- 17:18average now since Sigma is unknown and
- 17:21our sample size is small we will use the
- 17:23T distribution so the T formula is over
- 17:26there on the right
- 17:29so we go ahead and specify the type 1
- 17:31error rate now I'm going to choose an
- 17:33alpha of 01 and that's just my choice
- 17:36for this one-tailed test then we'll
- 17:39State our decision rule remember our
- 17:41sample size is 25 therefore our degrees
- 17:44of freedom is 24 so we need to go to our
- 17:47T table find the column for a onet
- 17:51tailed with an alpha of
- 17:5301 find the row for 24° of freedom and
- 17:58we do do that we come up with a critical
- 18:00T value of
- 18:022.49 2 so if our test statistic is
- 18:06greater than 2.49 2 we will reject our
- 18:11null
- 18:13hypothesis now on the curve we can
- 18:15actually place it so remember with an
- 18:17alpha of 01 what we're saying is that we
- 18:20have a 1% chance that we're ruling to
- 18:22accept a 1% probability of committing a
- 18:26type 1 error so that's the 1 % in the
- 18:29upper tail and therefore 99% of our
- 18:32sample means should be in the blue so
- 18:35we'll place our critical value our T
- 18:37value right there so 2.4
- 18:4192 of course then we'll gather our data
- 18:44so we had a sample size of 25 and our
- 18:46mean was
- 18:513.5 now we can calculate our test
- 18:54statistic so our mean of 3.5 our
- 18:57hypothesized mean of of three sample
- 18:59standard deviation of 1.4 and sample
- 19:02size of 25 so we'll go ahead and
- 19:04substitute those numbers into our
- 19:06formula and we arrive at A T statistic
- 19:09our test statistic of
- 19:151.79 so now we have to place that in the
- 19:17context of our sampling distribution so
- 19:20the same hypothesis we have our mean was
- 19:233.5 our critical value was 2.49 2
- 19:28our test statistic was
- 19:321.79 so what happens since the test
- 19:35statistic is inside the non-rejection
- 19:38region and inside the critical value
- 19:41sort of inside of towards the mean we
- 19:44fail to reject the null hypothesis that
- 19:47customer satisfaction is at or below
- 19:50average okay so we fail to reject the N
- 19:54hypothesis it is higher so it's 3.5 it
- 19:57is higher
- 19:58but the sample is just one of many that
- 20:01could be inside this non-rejection
- 20:04region the sample mean would have had to
- 20:07have been much higher to actually
- 20:10surpass that critical value now the
- 20:13question is what would that sample mean
- 20:16have to be to surpass that critical
- 20:20value of 2.4
- 20:2592 so finding the hypothetical sample
- 20:27mean that aligns with our T critical
- 20:30value of 2.49 2 is actually fairly
- 20:33straightforward remember when we first
- 20:35found our T statistic the T was the
- 20:38unknown and our sample mean was what we
- 20:41found in our data analysis so now we're
- 20:43going to flip that a little bit on its
- 20:45head now we know the T value and we want
- 20:49to know the sample mean that corresponds
- 20:52with that critical T value so our xbar
- 20:56is now the unknown and our T is known so
- 21:00again we'll just use our handy algebra
- 21:02from many years ago to solve for xar so
- 21:05we go ahead and simplify our denominator
- 21:08so that is 28 we multiply both sides by
- 21:1128 so we have 698 = xarus 3 so we'll go
- 21:17ahead and simplify that and we have a
- 21:19sample mean of 3.
- 21:23698 so therefore any sample of size 25
- 21:28with a mean that's greater than
- 21:323.69 would lead to the rejection of the
- 21:35null hypothesis assuming the same sample
- 21:38deviation which is not all that likely
- 21:40but we'll assume and the same Alpha
- 21:43level so the sample mean of
- 21:483698 is the sample mean that would
- 21:50hypothetically fall right onto our T
- 21:54critical
- 21:57value now if we go ahead and look at our
- 21:59distribution again we can actually sort
- 22:00of write these in so where t equal 1.79
- 22:05that's the test statistic from our
- 22:06analysis our sample mean was of course
- 22:103.5 now the sample mean that would
- 22:12correspond with our critical value of
- 22:152.49 2 is
- 22:203.69 so you can see the difference
- 22:22between our sample mean and the mean
- 22:24that would have been required to reject
- 22:27the
- 22:31hypothesis So based on our Alpha of 01
- 22:35we know that 1% of our area is in the
- 22:36upper tail so that's the 1% past RT
- 22:40critical of
- 22:432492 right there now in the P Value
- 22:47method we ask how much area or
- 22:49probability is above our test statistic
- 22:52so we're interested in how much area is
- 22:56to the right of our test
- 22:59statistic Now using a t table we can
- 23:01tell that the probability is between 05
- 23:05and 025 so we kind of eyeball it between
- 23:08two known now in Excel we can find it
- 23:11exactly Excel gives us a P value of
- 23:15043 now those are that's greater than 01
- 23:19so we know that we cannot reject the
- 23:22null
- 23:24hypothesis now this is often referred to
- 23:27as the observe D significance level so
- 23:30again we can find the P value here to
- 23:32actually show that we cannot reject the
- 23:35null because the area to the right of
- 23:38our T statistic is
- 23:410.43 but our Alpha is
- 23:4401 so therefore if we wanted to reject
- 23:47the null it would have to be less than
- 23:5301 now how did I find that in Excel so
- 23:56remember our degrees of freedom are 20
- 23:5724 our T value is
- 24:001.79 so we just use this formula in
- 24:03Excel 2010 it's
- 24:06t.d. RT that stands for right tail it
- 24:09saves us from having to subtract from
- 24:11one so 1.79 is our T value and 24 is our
- 24:16degrees of
- 24:17freedom so you can actually use the
- 24:20function tool in Excel it's a little F
- 24:23ofx button you can press and it'll tell
- 24:25you exactly what to put in so our T
- 24:28value is 1.79 and our degrees of freedom
- 24:31were 24 so if we look where it actually
- 24:34solves it the formula result is
- 24:38.43 and that's how I came up with the P
- 24:41value of
- 24:43043 in the previous
- 24:47slide okay so that wraps up our two
- 24:50examples on single sample T tests and
- 24:53the T distribution so just a reminder of
- 24:55the process so you can do it correctly
- 24:57ly always start with a well-developed
- 24:59clear research problem or question no
- 25:03fancy stats will solve a bad research
- 25:05problem so always hone in always clarify
- 25:09your problem before you start establish
- 25:11your hypothesis both null and
- 25:14alternative determine the appropriate
- 25:16statistical test and sampling
- 25:18distribution and again this depends on
- 25:20really two things do you know Sigma or
- 25:23do you not know Sigma is your sample
- 25:26size below 100 or not so again that will
- 25:30determine whether or not you're using
- 25:32the Z distribution and Z test or the T
- 25:35distribution and T Test choose your type
- 25:381 error rate again that's up to you then
- 25:41State your decision rule So based on the
- 25:43statistical test and distribution you're
- 25:45using and your error rate you will come
- 25:48up with a decision rule that says if my
- 25:50T value Falls here I cannot reject the
- 25:53null hypothesis if it falls here I have
- 25:56to reject the null hypothesis
- 25:58then gather your sample
- 26:00data once you have your sample data
- 26:02calculate the test statistic once you
- 26:05have the test statistic you compare that
- 26:07to your decision Rule and then State
- 26:09your
- 26:10conclusion once you have the conclusion
- 26:12you can then make some sort of real
- 26:15world recommendation or some real world
- 26:18publication whatever your application
- 26:20may happen to
- 26:23be okay so that wraps up part two of our
- 26:26video on the single sample T Test and I
- 26:29really hope you have a firm grasp of how
- 26:32all this comes together so we talked
- 26:34about the T Test as compared with the Z
- 26:36test we talked about what happens when
- 26:38you maybe even change one small
- 26:40parameter like the sample size in that
- 26:42case we got to completely different
- 26:43conclusion by changing the sample size
- 26:46by three so you can see that these some
- 26:49of these tests are very sensitive to
- 26:51their parameters and I really want you
- 26:53to understand how these things affect
- 26:56the overall test the alpha level
- 26:58the sample size and things like that so
- 27:01just keep in mind if you're watching the
- 27:02video cuz you're struggling in a class
- 27:04stay positive and keep your head up if
- 27:06you're watching this it means you've
- 27:07accomplished quite a bit you're smart
- 27:09and talented and you can get through it
- 27:11please feel free to follow me here on
- 27:13YouTube on Twitter on Google+ or on
- 27:15LinkedIn that way when I upload a new
- 27:17video you know about it and it's always
- 27:20nice to hear from people who watch my
- 27:21videos online life is Much Too Short The
- 27:24World Is much too large for us not to
- 27:26take the opportunity to connect with one
- 27:28another if you like the video please
- 27:30give it a thumbs up share it with
- 27:31classmates or colleagues or put it on a
- 27:33playlist and finally just keep in mind
- 27:35the fact that you're on here trying to
- 27:36learn trying to improve yourself as a
- 27:38student or as a business person or
- 27:41whatever else you may be doing that's
- 27:43what really matters I firmly believe if
- 27:46you have the right learning process in
- 27:48place the results will take care of
- 27:50themselves so thank you very much for
- 27:52watching and look forward to seeing you
- 27:54again next
- 27:56time
- 28:01[Music]
- 28:06oh
About this transcript
This page contains the full transcript of YouTube transcript (dDsKP7wVpzM) , generated from the public captions YouTube serves with the video. The transcript has 4,146 words across 614 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.