YouTube transcript (kHZBy1uVNnM) — Transcript
Full transcript
- 0:02[Music]
- 0:17hello thanks for watching and welcome to
- 0:20the next video in my series on basic
- 0:22statistics now as usual a few things
- 0:25before we get started number one if
- 0:27you're watching this video because you
- 0:28are struggling in a class right now I
- 0:31want you to stay positive and keep your
- 0:33head up if you're watching this it means
- 0:35you've accomplished quite a bit already
- 0:37you're very smart and talented but you
- 0:39may have just hit a temporary rough
- 0:41patch now I know with the right amount
- 0:43of hard work practice and patience you
- 0:46can work through it I have faith in you
- 0:49many other people around you have faith
- 0:51in you so so should you number two
- 0:55please feel free to follow me here on
- 0:57YouTube on Twitter on Google+ or on
- 1:01LinkedIn that way when I upload a new
- 1:03video you know about it and it's always
- 1:06nice to connect with my viewers online
- 1:08I'll feel that life is much too short
- 1:10and the world is much too large for us
- 1:12to miss the chance to connect when we
- 1:14can number three if you like the video
- 1:17please give it a thumbs up share it with
- 1:20classmates or colleagues or put it on a
- 1:22playlist that does encourage me to keep
- 1:24making them for you on the flip side if
- 1:27you think there's something I can do
- 1:28better please leave a instructive
- 1:30comment below the video and I will take
- 1:32those ideas into account when I make new
- 1:35ones and finally just keep in mind that
- 1:37these videos are meant for individuals
- 1:39who are relatively new to Stats so I'm
- 1:42just going over basic concepts and I
- 1:45will be doing so in a slow deliberate
- 1:48manner not only do I want you to know
- 1:51what is going on but also why and how to
- 1:55apply it so all that being said let's go
- 1:58ahead and get started
- 2:02so this video is the next in our series
- 2:03about simple linear regression in our
- 2:06last three videos we talked about the
- 2:08very basics of regression and introduced
- 2:10other fundamental concepts like the
- 2:12algebra of lines General patterns to
- 2:14look for on Scatter Plots and the leas
- 2:16squares method in this video we're going
- 2:19to learn to evaluate how well a
- 2:21regression line fits the data it models
- 2:25it is important to note that regression
- 2:27model is unique to the data it
- 2:29represents
- 2:30adding or changing data points will most
- 2:32certainly change the regression model
- 2:35the model is also only valid for the
- 2:37range of data points under analysis it
- 2:40is not proper to extrapolate above or
- 2:43below the data being
- 2:45evaluated so once a regression line is
- 2:48calculated how much better is it than
- 2:51using only the mean of the dependent
- 2:53variable alone to answer that question
- 2:56we will have to calculate the sum of the
- 2:57squared residuals or errors much like we
- 3:00did in the previous video however this
- 3:04time we will do so using the regression
- 3:06line not the dependent variable mean
- 3:10line and finally we will be able to
- 3:13quantify the fit of the regression model
- 3:16using a simple Ratio or percentage
- 3:19called the coefficient of determination
- 3:21so if you are new to regression or are
- 3:23still trying to figure out exactly what
- 3:25it even is this video is for you so sit
- 3:29back relax and let's go ahead and get to
- 3:34learning so we can think of this video
- 3:36as a story we'll call it a tail of two
- 3:40lines now in the first few videos we
- 3:42worked with this graph and data over
- 3:44here on the left remember in this case
- 3:48we only had the dependent variable which
- 3:50is the tip amount at the restaurant so
- 3:53we had six meals and then we had the six
- 3:56tip amounts that are graphed on the Y
- 3:58AIS in terms of their dollar
- 4:00amounts now we figured out the mean of
- 4:03these six tips was $10 so we went ahead
- 4:06and put a line a horizontal line at
- 4:09$10 then we found the errors which is
- 4:12the distance from the mean line to each
- 4:16observed data point and then we squared
- 4:19all of those errors and then added them
- 4:22up when we did that we came up with an
- 4:25ssse or sum of squared errors of 120
- 4:30now this is a very important point with
- 4:33only the dependent variable the only sum
- 4:36of squares is due to error that's all we
- 4:40have to work with therefore it is also
- 4:44the total sum of squares and the maximum
- 4:48sum of squares for the data under
- 4:50analysis so for these six specific data
- 4:53points the total sum of squares will
- 4:56never be more than
- 4:58120 so since the only source of error we
- 5:01have is the SSE it's also the total sum
- 5:05of squares in this case the same thing
- 5:08so SST is also 120 again because in this
- 5:12case they're the same
- 5:15thing now in the previous video we
- 5:17actually went ahead and calculated the
- 5:19regression line using the least squares
- 5:21method so we had the same six data
- 5:23points the graph looks different because
- 5:25they're graphed differently and then of
- 5:26course we have our regression line now
- 5:29in this video we're going to do the same
- 5:32thing we did over here on the left we're
- 5:34going to find the errors and square them
- 5:37and then add them up but of course we
- 5:39will not be using the mean of the
- 5:41dependent variable to do that we're
- 5:43going to be using the regression line to
- 5:45do that now with both the independent
- 5:48variable and the dependent variable the
- 5:50total sum of squares Remains the Same
- 5:55it's 120 still and it will always be for
- 5:58these six points but ideally the error
- 6:02sum of squares or the SS will be reduced
- 6:07significantly because our line fits the
- 6:11data better so the difference between
- 6:14the SST of 120 and the SSE which we will
- 6:18calculate is the SSR or the sum of
- 6:22squares due to the regression we did
- 6:25this is very much like a Nova in a way
- 6:27we'll talk about that as we go forward
- 6:29Ward but the SST will always be 120
- 6:32we're going to calculate the ssse based
- 6:34on the residuals from the regression
- 6:36line and then the difference between the
- 6:38SST and SS e will be the SSR or the sum
- 6:43of squares due to regression take a step
- 6:46back and look this is always what we are
- 6:49doing in simple linear regression we are
- 6:52comparing a regression model over here
- 6:54on the right that we found using the Le
- 6:57squares method to the case where we only
- 7:00have the dependent variable and the line
- 7:03is flat over here on the left now if the
- 7:07model on the right we calculate looks a
- 7:09lot like the one on the left and the sum
- 7:12of squared error is very high then our
- 7:15regression model doesn't really do a
- 7:17whole lot for us the idea is to reduce
- 7:20the ssse by having a line that fits the
- 7:23data better because the better the line
- 7:26fits the data the smaller the residual
- 7:29uals or the errors will
- 7:32be so this is our first case you can go
- 7:35and look at this again if you'd like but
- 7:37we just took the residuals of the errors
- 7:38squared them and then added them up this
- 7:41was
- 7:42120 so having only the dependent
- 7:44variable the best prediction for the tip
- 7:46of the next meal is just the mean of the
- 7:48tips which in this case is
- 7:50$10 since the mean line is flat its
- 7:53slope is zero so beta sub 1 equals 0
- 7:57again this is just a revieww
- 8:00now in the last video we actually
- 8:02calculated the regression line so here's
- 8:03a regression equation at the top and
- 8:06here's another one remember that Excel
- 8:08had a little bit different number for
- 8:10The Intercept than we did just because
- 8:11of rounding but basically they're the
- 8:13same thing so a few things our slope is
- 8:16not zero our slope is
- 8:1901462 and remember we interpreted that
- 8:22as for every dollar the meal amount
- 8:25increases the tip is expected to
- 8:28increase by about 15 cents that's the
- 8:31slope of
- 8:3301462 then we also had our intercept
- 8:36down here on the left which really for
- 8:37this problem isn't very meaningful now
- 8:39we also had the centroid so remember the
- 8:42centroid is the intersection of the
- 8:43means of both variables so the mean meal
- 8:47amount was $74 and the mean tip amount
- 8:49was $10 and the centroid will always
- 8:53fall on the least squares line and that
- 8:56can be very helpful if you have to find
- 8:58out other things about your
- 9:01graph so now that we're done reviewing
- 9:04let's go ahead and get to sort of the
- 9:05heart of the matter of this video but to
- 9:07do so we have to calculate the predicted
- 9:10values for each meal amount so over here
- 9:13on the left we have the observed total
- 9:15bill which in like the case the first
- 9:17meal was $34 and then we have the
- 9:20observed tip so that was the tip that
- 9:22was actually received by the waiter so
- 9:25the first one was $34 the tip that the
- 9:27waiter got was $5 $18 meal headed a tip
- 9:31of $17 so on and so forth now in this
- 9:35middle column we have our actual
- 9:36regression equation so 01462 xus
- 9:410.81
- 9:4388 now that is what we're going to use
- 9:46to create the predicted tip amount using
- 9:49our regression equation for each of our
- 9:52meals over here in the red on the
- 9:57left all we do is substitute each meal
- 10:01amount in for X because again that's our
- 10:04independent variable so we go ahead and
- 10:07evaluate all of these equations and that
- 10:10will give us our predicted tip amount
- 10:12over here on the right so again to
- 10:14substitute the meal amount into the
- 10:16regression equation and that will give
- 10:18us our predicted tip amount based on the
- 10:24regression so substitute over evaluate
- 10:27and now we have our predicted
- 10:29tips so for this meal that was $34 the
- 10:33waiter received
- 10:35$5 now our regression equation would
- 10:39have predicted a tip of
- 10:42$415 let skip down we had a meal that
- 10:46was
- 10:46$88 the waiter or waitress received
- 10:50$8 Now using our regression equation we
- 10:53would predict a tip of $12
- 10:57about5 so you can see the obser tip
- 10:59amount is not always or usually is not
- 11:03the predicted tip amount that is given
- 11:05using the regression equation so we have
- 11:08a discrepancy here we have the observed
- 11:10tip amount and we have the predicted tip
- 11:13amount and the difference between those
- 11:15two is going to be our
- 11:20error so here is our data again so the
- 11:22Diamonds the orange diamonds represent
- 11:24our actual observed values for the tips
- 11:27that the waiter received so here's our
- 11:29regression equation now the purple dots
- 11:33are the predicted tip amounts for those
- 11:36meal amounts as you can see they're not
- 11:38the same as what we observed so for the
- 11:41first one we had $415 that we would
- 11:43predict the tip to be for the second one
- 11:45we had a predicted value of
- 11:47$664 for a tip and so on and so forth on
- 11:51up to the
- 11:54ground now we have to find out what the
- 11:57error is or what the residual R so again
- 12:01that is just the distance between the
- 12:03predicted value and the observed value
- 12:07because that's quote how far off our
- 12:10regression is from the actual data so
- 12:12it's just the distance between the
- 12:14predicted and The
- 12:16observed so now we have to find out what
- 12:18those errors are and as you can imagine
- 12:21it's simple subtraction it's just the
- 12:22difference between the observed and the
- 12:25predicted that we calculated so for our
- 12:27first meal $34 we have observed tip
- 12:29amount actually in the restaurant of $5
- 12:32our regression predicted a
- 12:35$415 tip approximately so the difference
- 12:38between those two was
- 12:410.849 or about
- 12:4385 and that is the error for that meal
- 12:47amount and then we do that for all six
- 12:49meals very simple just
- 12:53subtraction now of course as we do with
- 12:56other errors we have to square those
- 12:58differences so the difference of
- 13:008495 for the first meal if we square
- 13:04that residual it is
- 13:060.721 7 for the second meal the
- 13:09difference is
- 13:122.37 we go ahead and square that and
- 13:15it's
- 13:29we sum those up we have an SS e that is
- 13:3430.75 so for our regression model that
- 13:38we came up with our SSE is a little bit
- 13:41over 30 so
- 13:4530.75 so on our graph we take each
- 13:49residual and then we Square it just like
- 13:53this so we take the residual and square
- 13:55it so these are areas we have here and
- 13:59then just like we did before we go ahead
- 14:02and add those
- 14:06up now let's compare the sum of squared
- 14:09errors from our two models now remember
- 14:12in the first model where we only had the
- 14:14dependent variable which was the tip
- 14:16only we had a sum of squared errors or
- 14:18SSE of 120 and again remember that's
- 14:22also the SST for that model so
- 14:26120 now for our regression model we have
- 14:29the DV and the IV to the dependent and
- 14:32independent variable where the tip
- 14:34amount is a function of the meal amount
- 14:37now we have a suos squared errors or SS
- 14:40of
- 14:4430.75 see how this works that's the
- 14:47entire point of simple linear regression
- 14:51to create a model to create a line
- 14:54through our data that reduces the ssse
- 14:57as much as it possibly can
- 15:01so if we actually make these physically
- 15:04the same scale side by side we can see
- 15:06that we put all those sses for the first
- 15:09model together it's 120 and we do the
- 15:11same thing for our regression model and
- 15:13it's
- 15:1530.75 so you can actually see the
- 15:17physical difference in the scale of the
- 15:20errors so when we conducted the
- 15:23regression the ssse decreased from 120
- 15:27to 30.0
- 15:30075 that is
- 15:3230.75 of the sum of the squares was
- 15:36explained or allocated to the error in
- 15:40the regression model so instead of
- 15:42having 120 as the error as in the first
- 15:45case we reduced that to
- 15:4830.75 because we reduced the distance
- 15:52from the regression line to the data
- 15:54points where did the other 89.8 n25 go
- 15:59well
- 16:0189.95 is the sum of squares due to our
- 16:07regression SST equal SSR plus
- 16:11SS in this case SST is always 120 and
- 16:15that equals 89.9 25 +
- 16:2130.75 so here are our two models again
- 16:25so the first one our ssse was 120 and of
- 16:28course the s s and SST are the same for
- 16:30that model so the SST is also
- 16:33120 now in our second regression model
- 16:36that we did our SST again is still 120
- 16:39but now the SS is
- 16:4330.75 because of the reduced distance
- 16:46between the regression line and the
- 16:49observed data points as it compares to
- 16:51the one over here on the left so doing
- 16:54Simple subtraction we had 120us 30.75
- 16:58that gives us our SSR or sum of squares
- 17:01due to regression and that's
- 17:0389.9
- 17:0725 so now that we have the SST the SSR
- 17:11and the ssse we can go ahead and talk
- 17:13about the coefficient of
- 17:15determination so how well does the
- 17:18estimated regression equation fit our
- 17:22data now this is where regression begins
- 17:25to look a lot like a Nova the total sum
- 17:29of squares is partitioned or allocated
- 17:33into SS and SSR and hopefully at some
- 17:38point I'll do a separate video on the
- 17:39relationship between regression and
- 17:41Anova but it's all about allocating or
- 17:45partitioning the total sum of squares
- 17:48into two or more things in this case SS
- 17:51and
- 17:52SSR so the total sum of squares is split
- 17:56some of it goes to SSR some of it goes
- 17:59to
- 18:00SS so if the SSR is large it uses up
- 18:05more of the SST and therefore SS is
- 18:10smaller relative to the total the
- 18:13coefficient of
- 18:14determination quantifies this ratio as a
- 18:18percentage so the coefficient of
- 18:21determination has the variable R 2
- 18:25equals SSR / SST so it's the sum of
- 18:29squares regression divided by the total
- 18:32sum of
- 18:34squares and if you actually do this in a
- 18:37statistical software package regression
- 18:39will actually give you back an anova
- 18:42table just like you saw when you did an
- 18:45anova so we're going to have SS and mean
- 18:48squares and an F statistic and a
- 18:51significance level and everything just
- 18:53like we did in an NOA and of course in
- 18:56this case this is the actual Anova for
- 18:58the data what we're working with we can
- 19:00see that the f is very large and
- 19:03significance is
- 19:060.258 which is significant at the 05
- 19:09level so there is a relationship between
- 19:12linear regression and an NOA because
- 19:14it's all about partitioning or
- 19:16allocating the total sum of squares into
- 19:19different components and then we measure
- 19:22the ratio of those components to figure
- 19:25out whether or not the model is
- 19:26statistically significant same
- 19:30idea so how do we interpret the
- 19:32coefficient of determination so again
- 19:35it's SSR over SS in this case it's
- 19:3989.95 ID 120 so we go ahead and do that
- 19:43Division and we come up with 0.74 93 or
- 19:4974.9 3% cuz remember the 120 is the
- 19:53total so
- 19:558992 of the 120 that's a percent
- 19:59and that is
- 20:0074.9 3 so we can conclude that
- 20:0474.9 3% of the total sum of squares can
- 20:09be explained using the estimated
- 20:11regression equation to predict the tip
- 20:14amount the remainder so
- 20:1925.7% is due to error and again that's
- 20:23based on this regression equation we
- 20:25used where X is the dollar amount of the
- 20:29bill so it's really about partitioning
- 20:31the total sum of squares as a percentage
- 20:34saying hey this percentage belongs to
- 20:36SSR and this percentage belongs to
- 20:39ssse and then we can develop a model
- 20:43that says this fits well or does not fit
- 20:45well in this case it is a good fit so a
- 20:50coefficient of determination of 74.9 3%
- 20:54is good and we also saw that in the
- 20:56Innova table which was statistic Bally
- 20:59significant so it's all about
- 21:01partitioning the total sum of squares
- 21:03into SSR and ssse and we want SS to be
- 21:08as low as possible for the model to be
- 21:11considered a good fit and for the Innova
- 21:14table to produce a significant
- 21:19result so let's put all together on one
- 21:22graph so you can see here we're dealing
- 21:24with three data points for each measure
- 21:28on our xaxis or independent variable so
- 21:31if we look at this one here sort of to
- 21:33the right we can see that we have a
- 21:35orange diamond we have a black circle
- 21:39and a purple circle so three data points
- 21:43that represent that meal of $88 I think
- 21:46it was so remember that the diamond is
- 21:49the actual observed value the black
- 21:52circle is the mean of the dependent
- 21:56variable which was the tips which is $10
- 21:59the purple circle is the predicted tip
- 22:02amount based on the regression model so
- 22:05each meal kind of has three reference
- 22:07points The observed the mean of the
- 22:10dependent variable and the predicted
- 22:13value so the first Square difference to
- 22:15look at is the SS and I think it's the
- 22:17easiest to remember so it is the sum of
- 22:20the squared differences between y sub I
- 22:23which is the observed value the diamond
- 22:26minus y hat sub I which is the predicted
- 22:30value which is the purple circle so we
- 22:32take the difference of each one along
- 22:34the graph we Square them and then we add
- 22:37them all up and that is the SS for our
- 22:40regression model so it's the distance
- 22:43between those two and similarly those
- 22:46two up here so the difference between
- 22:49the observed the diamond and the
- 22:50predicted the purple circle now the next
- 22:54one is SST it's the sum Square total now
- 22:58it is the sum of the square difference
- 22:59between each observed value which is the
- 23:03diamond minus the mean of the Y AIS or
- 23:07the dependent variable which is y bar so
- 23:11the difference between those then we
- 23:13Square it then we add them all up so
- 23:16that's the difference between the
- 23:17observed and the mean so the orange
- 23:19diamond and the black circle so just
- 23:22like that and just like that there and
- 23:25of course the final one is SSR
- 23:29so it is the sum of the squar difference
- 23:32between y hat sub I that's the predicted
- 23:35whe is the purple circle minus y bar
- 23:39which is the mean of the dependent
- 23:42variable so we find the difference we
- 23:44Square them and then add them up and
- 23:46that is this distance here and this
- 23:49distance here of course we would do this
- 23:51for each point along the graph so it's
- 23:55always a relationship between these
- 23:58three points that represent each value
- 24:02of the independent variable so in the
- 24:04middle here in this case a $88 meal
- 24:07amount we have the observed we have the
- 24:10mean of the Y's and we have the
- 24:12predicted and we have three
- 24:13relationships going on between those
- 24:15three data points and the square
- 24:17differences among those three data
- 24:19points is actually the total sum of
- 24:21squares package of SST ssse and SSR so
- 24:26now you can see what it actually looks
- 24:27like on
- 24:31graph okay so that wraps up our video on
- 24:34simple linear regression fit and the
- 24:37coefficient of determination so remember
- 24:40this is all about a comparison A Tail of
- 24:43Two Lines where we had the one graph
- 24:46where we only used the mean of the
- 24:48dependent variable and then we had our
- 24:50regression equation and that graph over
- 24:53on the right and it's always a
- 24:55comparison between the two if the
- 24:58regression model isn't any better than
- 25:01the mean of the dependent variable then
- 25:03the model doesn't really give us
- 25:04anything new and will most likely come
- 25:07back is statistically insignificant the
- 25:09SS will be very high and then we will
- 25:12just move on from there but it is always
- 25:15a comparison between these two different
- 25:18models now the coefficient of
- 25:20determination just tells us how we
- 25:23allocate the total sum of squares in the
- 25:25regression model so it's split it's
- 25:27either SS s e or it's SSR and we hope
- 25:31for a good fitting model that the SS is
- 25:34very low and the SSR is very high but in
- 25:38either case their sum will always add up
- 25:41to SST whatever that happens to be so I
- 25:44look forward to seeing you in that next
- 25:46video I wish you all the best of luck in
- 25:48your studies and in your work and I look
- 25:50forward to seeing you again next time
- 25:55[Music]
About this transcript
This page contains the full transcript of YouTube transcript (kHZBy1uVNnM) , generated from the public captions YouTube serves with the video. The transcript has 3,824 words across 540 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.