Confidence Interval Estimation Sigma Known — Transcript
Full transcript
- 0:00welcome to this tutorial on interval
- 0:02estimation for the population mean when
- 0:05sigma is known
- 0:06in the previous tutorial we discussed
- 0:08how a point estimator from a sample can
- 0:10be used to estimate a population
- 0:12parameter
- 0:14for example the sample mean x bar is a
- 0:17point estimator for the population mean
- 0:19mu and the sample proportion
- 0:22p bar is a point estimator for the
- 0:24population proportion p
- 0:28in this tutorial we will see how a point
- 0:30estimator such as the sample mean x bar
- 0:33can be used more realistically by adding
- 0:36and subtracting a value called a margin
- 0:38of error
- 0:39this will give us something called an
- 0:42interval estimate this interval estimate
- 0:44is a much more realistic predictor of
- 0:47where the true value of the population
- 0:49parameter is
- 0:50so instead of using a single point such
- 0:52as x bar we compute an interval around
- 0:55that point and the formula looks
- 0:57something like this x bar the point
- 1:00estimate plus and minus some margin of
- 1:02error
- 1:04for the interval estimate of the
- 1:06population proportion we will use p bar
- 1:09plus or minus some margin of error now
- 1:12in order to calculate an interval around
- 1:14the population mean we will need to use
- 1:17either sigma the population standard
- 1:19deviation
- 1:20or s the sample standard deviation to
- 1:23compute the margin of error
- 1:25in most cases the population standard
- 1:27deviation sigma is unknown and we need
- 1:30to use s the sample standard deviation
- 1:32to compute the margin of error
- 1:34but there are times when we have
- 1:36historical data that can be relied upon
- 1:38to determine the population standard
- 1:40deviation or a production process where
- 1:43the standard deviation is known
- 1:45these types of situations where the
- 1:48population standard deviation is known
- 1:50are called sigma known cases
- 1:54let's look at the case of a light bulb
- 1:56factory that selects a simple random
- 1:58sample of 50 light bulbs each week to
- 2:01measure how many hours they burn
- 2:04each week a sample mean is obtained and
- 2:07used as a point estimate for the true
- 2:09mean mu or the number of hours the light
- 2:11bulb burns
- 2:12now let's assume based on historical
- 2:14data that the light bulbs are assumed to
- 2:17have a normal distribution with a known
- 2:19standard deviation of 65 hours
- 2:22then we would say that sigma is 65 hours
- 2:26now let's say this past week a sample of
- 2:2850 bulbs was measured and the sample
- 2:31mean that was obtained was an x bar of
- 2:341045 hours this sample mean of 1045
- 2:39hours provides a point estimate for the
- 2:41true population mean mu
- 2:44what we would like to do is compute the
- 2:46margin of error for this estimate to
- 2:48obtain an interval estimate for the true
- 2:51population mean
- 2:53using what we learned previously about
- 2:55sampling and sampling distributions we
- 2:57know that the sampling distribution of x
- 2:59bar follows a normal distribution with a
- 3:02standard error
- 3:03sigma subscript x bar
- 3:05and this is equal to sigma divided by
- 3:08the square root of a little n
- 3:10in this case that would be
- 3:1265 divided by the square root of 50
- 3:15which is 9.19
- 3:19this graph shows the sampling
- 3:20distribution of how the x bar values are
- 3:23distributed around the population mean
- 3:25mu
- 3:26now using the standard normal
- 3:27probability table also called the z
- 3:30table we can find the interval around
- 3:32the mean that contains 95 percent of the
- 3:35values
- 3:36we do this first by graphing the area
- 3:38around the mean that contains 95 percent
- 3:41of the values
- 3:43here we can see a line splitting the
- 3:45distribution in half at the mean mu
- 3:48now we draw two lines
- 3:50above and below the mean such that
- 3:5395 of the values are contained within
- 3:56that interval
- 3:58now to look this area up in the z table
- 4:01we need to split 95
- 4:03or 0.95 in half so that half of the 0.95
- 4:08or 0.475
- 4:09is below the mean and the other half
- 4:120.475
- 4:14is above the mean and since we know that
- 4:16the distribution is split by the mean
- 4:18with 50 of the values below and 50
- 4:22percent of the values above
- 4:24then the tail area below the mean is
- 4:27two five
- 4:28and the tail area above the mean is
- 4:31also point zero two five
- 4:33looking at the distribution now we see
- 4:36all the parts add up to one or a hundred
- 4:38percent
- 4:39you can see that point 0.475 plus 0.025
- 4:43is 0.5 so we have 0.5 below the mean and
- 4:47again 0.5 above the mean and of course
- 4:500.5
- 4:51plus 0.5 is 1.
- 4:53so now with all the parts of the
- 4:55distribution marked off we are ready to
- 4:57look up the z-score for the area under
- 5:00the table that is 95 percent of the
- 5:02values around the mean
- 5:04let's start with the upper z-score
- 5:07remember this value represents the
- 5:09cumulative area under the curve to the
- 5:11left of this number so we would look up
- 5:14in the table the area of
- 5:160.475 plus 0.475
- 5:19plus 0.025 or
- 5:220.975
- 5:24in the middle of the z table
- 5:27here is the z table with the positive
- 5:29numbers so we would look for
- 5:310.975 in the middle of the table
- 5:34and find it here then we would look up
- 5:37and to the left to read off the z value
- 5:40of 1.96
- 5:42so going back to the distribution we can
- 5:45mark off a z value of 1.96 here
- 5:49now for the lower z value we would look
- 5:52up in the negative z table the area to
- 5:54the left of negative z
- 5:56and that is point zero two five so let's
- 5:58look that up in the negative z table
- 6:02and we find
- 6:04point zero two five in the middle of the
- 6:06table and then looking up and to the
- 6:08left we find the z value for that area
- 6:11under the curve
- 6:12is negative 1.96
- 6:15going back to the distribution we can
- 6:18mark that z value right here so now we
- 6:21have a distribution with the z values
- 6:23marked for the area that contains 95
- 6:25percent of the values around the true
- 6:28mean mu
- 6:30what is left now is to convert these z
- 6:33values back into x values
- 6:35remember that to convert a z value back
- 6:38to any x value or x bar value we use
- 6:41this formula
- 6:43x bar is equal to mu plus and minus the
- 6:46standard deviation for the sampling
- 6:48distribution x bar
- 6:50where the standard deviation of the
- 6:51sampling distribution is sigma over the
- 6:54square root of little n
- 6:56so the formula ends up looking like this
- 6:59in the red box
- 7:01x bar is equal to mu the population mean
- 7:04plus and minus z the value we look up in
- 7:07the z table times sigma over the square
- 7:10root of little n
- 7:12now let's plug in the numbers for our
- 7:13example
- 7:15and we get x bar is 1045
- 7:19plus or minus
- 7:201.96
- 7:22times sigma 65
- 7:24divided by the square root of little n
- 7:26which is 50. and that gives us
- 7:291045
- 7:31plus and minus 18.01
- 7:34and that gives us an x bar of
- 7:361026.99
- 7:39and 1063.01
- 7:47this graph shows the sampling
- 7:49distribution of x bar where 95 percent
- 7:52of the x-bar values must be within plus
- 7:54or minus 1.96 standard deviations of the
- 7:57mean
- 7:58we see from this that 95 percent of the
- 8:01x-bar values obtained using a sample
- 8:04size of n equal 50 will be between
- 8:081026.99 and 1063.08
- 8:12hours
- 8:13the general form of an interval estimate
- 8:16for the population mean
- 8:17is x bar
- 8:19plus or minus some margin of error this
- 8:22general formula translates into x bar
- 8:25plus and minus z times sigma x bar
- 8:29remember that
- 8:30sigma x bar is equal to sigma over the
- 8:32square root of little n
- 8:34this translates for our example into x
- 8:36bar plus and minus 1.96 times 9.19
- 8:42remember
- 8:439.19 was calculated by dividing 65 by
- 8:46the square root of 50
- 8:48which is 9.19 so now we get x bar the
- 8:52sample mean which is 1045
- 8:55plus and minus 1801 and when we take
- 8:591045 and at 1801 and subtract 1801 we
- 9:04get this interval around the mean
- 9:061026
- 9:09and 1063.01
- 9:12where 95 of the sample x bars will be
- 9:15located
- 9:16what we have just calculated is called
- 9:19an interval estimate and it is important
- 9:21to understand the reason that this is
- 9:23called an interval estimate is that it
- 9:25provides an interval around the mean
- 9:28instead of just one point
- 9:30when we took a sample size of n equal 50
- 9:33and obtained a sample mean of 1045 hours
- 9:37that sample mean could be used as a
- 9:39point estimate of the true mean
- 9:41we stipulated that the true mean is 1050
- 9:45hours so that we can understand how
- 9:47sample means can be used to estimate the
- 9:49true mean
- 9:50in real life of course we don't know the
- 9:52value of the true mean mu
- 9:55now by adding and subtracting a margin
- 9:57of error we have calculated an interval
- 10:00so instead of just one number as a point
- 10:03estimate we have a low number
- 10:06of
- 10:071026.99 and a high number of 1063.01
- 10:12within which the true mean is likely to
- 10:14be
- 10:16now let's suppose we take another sample
- 10:18of 50 light bulbs but this time we get a
- 10:20sample mean of 1040.
- 10:23then the interval estimate for this
- 10:25sample would be
- 10:271040 plus and minus 18.01
- 10:31or
- 10:321021.99
- 10:36as the lower limit and 1058.01
- 10:40as the upper limit
- 10:43here is what that would look like on the
- 10:44distribution curve
- 10:46we see that the true mean mu of 1050 is
- 10:50still contained within this interval but
- 10:52the interval is shifted over to the left
- 10:55since the sample mean of 1040 is lower
- 10:58than the previous sample of one thousand
- 11:00forty five we can keep doing this taking
- 11:03samples of fifty light bulbs and getting
- 11:06sample mean after sample mean after
- 11:08sample mean and
- 11:10ninety five percent of the time we will
- 11:12get an interval that does include the
- 11:14true mean of 1050.
- 11:17we just saw that these two sample means
- 11:19when we drew an interval around the
- 11:21sample means of 1045
- 11:23and 1040 we got an interval that did
- 11:26include the true mean
- 11:28this will happen ninety-five percent of
- 11:30the time which means that five percent
- 11:33of the time we will get a sample mean
- 11:35that produces an interval estimate that
- 11:37does not include the true mean
- 11:41this is called alpha
- 11:43so alpha percent of the time we will get
- 11:45a sample mean that does not include the
- 11:47true mean
- 11:49in this case alpha is five percent or
- 11:51point zero five now let's see what
- 11:54happens when we take another sample of
- 11:5650 lipos but this time we get a poor
- 11:59representation of the population and the
- 12:01sample produces a sample mean of
- 12:041030 hours
- 12:06this will happen alpha percent of the
- 12:08time or in this case five percent of the
- 12:11time
- 12:12five percent of the time we will take a
- 12:13bad sample that does not represent the
- 12:16true population as is in this case with
- 12:18an x bar or sample mean of one thousand
- 12:21thirty hours
- 12:22now let's see what happens when we
- 12:24calculate an interval around this sample
- 12:26mean
- 12:27does it include the true mean of a
- 12:28thousand fifty hours well let's see
- 12:32we get 1030 plus and minus 18.01 which
- 12:36gives us an interval of
- 12:391011.99
- 12:41and 1048.01
- 12:44we can mark this off on the distribution
- 12:47to get a better idea of where this
- 12:49interval falls and we can see clearly by
- 12:52looking at the interval that is colored
- 12:54in red that it does not include the true
- 12:56mean of one thousand fifty
- 12:59so let's review what we just did we took
- 13:01three different samples of fifty light
- 13:03bulbs and obtained three different
- 13:06sample means the first sample mean was
- 13:091045 hours the second sample mean was
- 13:121040 hours and the third sample mean was
- 13:15quite low only 1030 hours
- 13:18each of those sample means could be used
- 13:21by themselves as point estimators of the
- 13:23true mean
- 13:24but to be more realistic in our use of
- 13:26sample data we draw an interval around
- 13:29the sample mean and use that as an
- 13:31interval estimate for the true mean
- 13:33this is done by adding and subtracting
- 13:35to the sample mean a margin of error to
- 13:38get the interval estimate
- 13:40we did this here for all three sample
- 13:42means the blue colored interval is for a
- 13:45sample mean of 1045
- 13:47and we can see the interval does include
- 13:49the true mean mu
- 13:51the green colored interval is for x bar
- 13:54two and it had a sample mean of one
- 13:56thousand forty it also included the true
- 13:58population mean of one thousand
- 14:01and finally the red colored interval is
- 14:04for x bar 3 which was a sample mean of
- 14:071030
- 14:09and we see that it does not include the
- 14:11true mean we will get a bad sample that
- 14:14does not include the true mean alpha
- 14:16percent of the time and in this case
- 14:18that is 0.05 or 5 percent of the time
- 14:25what we have just done to calculate an
- 14:27interval around a sample mean is called
- 14:29a confidence interval the interval we
- 14:32calculated used 0.95 or 95 percent
- 14:35interval so it is called a 95 confidence
- 14:39interval
- 14:40that means we are 95 confident that the
- 14:42true mean exists within our interval
- 14:45since 95 percent of all intervals that
- 14:47are constructed using the sample mean
- 14:50plus and minus some error will contain
- 14:53the true population mean
- 14:55this is represented by x bar plus and
- 14:57minus a margin of error or x bar plus
- 15:01and minus z times sigma over the square
- 15:03root of n when we use a z value of 1.96
- 15:07we get a 95 confidence interval
- 15:10this value of 0.95 is called the
- 15:13confidence level and the interval
- 15:16created is called a confidence interval
- 15:19so the interval we obtained by taking
- 15:21the sample mean of 1045
- 15:24plus and minus the 1.96 times 9.19
- 15:28was an interval of between 1026.99
- 15:32and 1063.01
- 15:35this is a 95 confidence interval for the
- 15:38mean based on that particular sample
- 15:41another term that is used in statistics
- 15:43when discussing interval estimation is
- 15:46the level of significance this is always
- 15:491 minus the confidence level so for our
- 15:52example alpha
- 15:54is 1 minus 0.95 since we constructed a
- 15:5895 confidence interval so alpha is 0.05
- 16:03that is our level of significance 0.05
- 16:06that level of significance the area
- 16:08outside of 0.95 that is within the
- 16:11confidence interval is the area in the
- 16:13tails take a look at the distribution
- 16:16and you can see that 0.95 is the
- 16:18interval around the mean
- 16:20half above and half below this leaves us
- 16:23with two tails one lower tail area and
- 16:26one upper tail area these two areas
- 16:28contain alpha the level of significance
- 16:31which is point zero five
- 16:33so since the areas are divided equally
- 16:36and .05 is in both of them together
- 16:39we need to split alpha in half to get
- 16:42the area in each of the tails which is
- 16:45point zero two five
- 16:49so now we can see the distribution is
- 16:51marked with an interval around the mean
- 16:53containing 0.95 or 95 percent of the
- 16:56values and then .05 or 5 percent of the
- 16:59values in the tails with 0.025 in the
- 17:02lower tail and .025 in the upper tail
- 17:06a 95 confidence interval is a frequently
- 17:09used level of confidence and it is a
- 17:12good idea to remember the number 1.96 as
- 17:15the z value that is used instead of
- 17:17having to look it up in the z table
- 17:19every time
- 17:21here is a table of the most frequently
- 17:23used confidence levels and their
- 17:24respective z values
- 17:26keep this as a handy guide so you don't
- 17:29have to look the numbers up in the table
- 17:31every time you need them
- 17:32the most commonly used confidence levels
- 17:35are ninety percent ninety five percent
- 17:38and ninety nine percent
- 17:39this table shows the alpha values for
- 17:41each of those levels of confidence so
- 17:44for example for a confidence level of 90
- 17:47or 0.90 alpha would be
- 17:520.10 and alpha divided in half remember
- 17:55we have to divide alpha in half between
- 17:56the two tails the lower tail and the
- 17:59upper tail so alpha divided in half
- 18:02would be point zero five
- 18:04let's see how we would look up point
- 18:06zero five in the negative z table and
- 18:08find the z value is around
- 18:11here
- 18:12as you can see there is no exact value
- 18:14for .05 we have .0495
- 18:18and we have .0505
- 18:21so .05 would be smack in the middle of
- 18:23these two
- 18:25so looking to the left we see the value
- 18:27is negative 1.6 and looking up we see
- 18:31the hundredths value is between 0.04 and
- 18:340.05 so our z-score would be negative
- 18:371.645 let's take another look at the
- 18:40table of most commonly used confidence
- 18:42intervals we just saw how the z value
- 18:45for a 90 confidence interval is 1.645
- 18:50and previously we found a 95 confidence
- 18:53interval had a z value of 1.96
- 18:57now let's repeat this process with a 99
- 19:00confidence interval
- 19:04if we want to obtain a 99 confidence
- 19:06interval then what would alpha be
- 19:09we know that alpha is 1 minus the level
- 19:11of confidence so 1 minus 99
- 19:14is 0.01 and since there are two tails in
- 19:17the distribution an upper tail and a
- 19:19lower tail when we are constructing
- 19:21intervals we always split alpha in half
- 19:24so alpha divided in half gives us point
- 19:26zero zero five
- 19:28before we look this up in the z table
- 19:30let's try to visualize this
- 19:33here you can see a distribution that has
- 19:35a ninety-nine percent confidence
- 19:37interval marked off around the mean
- 19:39knowing that the entire distribution is
- 19:41equal to 100 percent and we have 0.99
- 19:45marked off in this interval then 1 minus
- 19:48the confidence level or 1 minus 0.99 is
- 19:51equal to 0.01 now alpha is the area in
- 19:54both tails and since we have two tails
- 19:58a lower tail and an upper tail then we
- 20:00must split the alpha area of .01 between
- 20:04these two tails and we get
- 20:06.005 in one tail
- 20:09and .005 in the other tail that's alpha
- 20:12divided in half
- 20:14now when we look at the distribution we
- 20:15see it adds up to one with .005 in each
- 20:19tail and 0.99 in the middle
- 20:22now that we understand that we are ready
- 20:23to look up .005
- 20:26the lower tail area in the z table the
- 20:28number we should get is here 2.576
- 20:34back to the z table
- 20:37we look for
- 20:38.005 in the middle of the table and the
- 20:40closest number we find are two numbers
- 20:43.0049
- 20:45and
- 20:46.0051 and of course .005
- 20:50is right in the middle of these two so
- 20:52looking to the left we get negative 2.5
- 20:55and looking up to the hundredths value
- 20:58we get somewhere between 0.07 and 0.08
- 21:01so the z value according to this table
- 21:04is minus
- 21:052.575
- 21:07right between minus 2.57 and minus 2.58
- 21:12we see in this convenient table the z
- 21:14value is marked off as
- 21:162.576
- 21:18not
- 21:192.575 when we use the z table the reason
- 21:22for this discrepancy is that this value
- 21:252.576 is actually more accurate when you
- 21:28look at more than four decimal places
- 21:31the z table only goes to four decimal
- 21:33places so we get 2.575
- 21:36but if you were to use excel or
- 21:38calculate using calculus you would get
- 21:422.576
- 21:43consider this a rounding error
- 21:46many textbooks simply round this number
- 21:48up to the nearest hundredths place and
- 21:50use 2.58 for our purposes we will
- 21:53continue to use either 2.576
- 21:56or 2.575
- 21:58whenever we use a 99 confidence interval
- 22:01just to make sure how to draw a
- 22:03confidence interval when sigma is known
- 22:05let's take a look at one more example
- 22:08let's say we want to draw a 90
- 22:10confidence interval estimate of the
- 22:12population mean grades on a departmental
- 22:15statistics exam
- 22:17from historical data the population
- 22:19standard deviation is known to be 15
- 22:21points
- 22:22suppose a sample of 75 students are
- 22:25taken and the sample mean is 70. the
- 22:28distribution for a 90 confidence
- 22:31interval would look like this
- 22:33you don't have to draw this interval
- 22:35each and every time you construct a
- 22:37confidence interval but it is helpful in
- 22:39visualizing the data and to understand
- 22:41what we are doing
- 22:43since our confidence level is
- 22:450.90 then the level of significance
- 22:48alpha is 0.10 and alpha divided in half
- 22:52would be 0.05
- 22:55this number is what we would look up in
- 22:56the z table and since we have done this
- 22:59before i won't repeat the process here
- 23:01if you recall the z value for 90
- 23:03confidence was
- 23:051.645 you can rewind and review how we
- 23:09got that a couple slides back
- 23:11let me point out the notation i use here
- 23:13if you'll notice i didn't just write z
- 23:15this time i wrote z and then subscript
- 23:18alpha divided in half
- 23:20this is a more accurate and precise
- 23:22notation of the z value when alpha is
- 23:25split in half
- 23:26here i have marked off the z values on
- 23:29the distribution curve
- 23:30now it's time to plug this into our
- 23:32formula to get a ninety percent
- 23:34confidence interval around the true mean
- 23:36if you recall the formula is x bar plus
- 23:40and minus a margin of error which is x
- 23:43bar plus and minus z times sigma over
- 23:46the square root of n
- 23:47now all that is left to do is to plug in
- 23:49the numbers and calculate the interval
- 23:52so we get 70 plus and minus 1.645
- 23:56times 15 over the square root of n which
- 24:00is 75
- 24:01and so we get 70 plus or minus 2.58 i
- 24:05rounded it to the hundreds place for
- 24:07convenience and 70 plus or minus 2.58
- 24:11gives us a lower limit of 67.15
- 24:14and an upper limit of 72.85
- 24:18this means we are 90 confident that the
- 24:20true mean is between 67.15
- 24:24and
- 24:2572.85 based on this sample data
- 24:28that concludes this video tutorial on
- 24:31interval estimation for sigma known
- 24:33please be sure to watch the next
- 24:35tutorial on confidence interval
- 24:37estimation for sigma unknown i hope you
- 24:40enjoyed this tutorial and i hope you
- 24:42learned something
- 24:54you
About this transcript
This page contains the full transcript of Confidence Interval Estimation Sigma Known by Learn Something, generated from the public captions YouTube serves with the video. The transcript has 3,692 words across 590 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.