YouTube transcript (px72eCYPuvc) — Transcript
Full transcript
- 0:00So let's do the case where we have two
- 0:02variables. So now we're going to put two
- 0:04variables in at a time. So we'll put x1
- 0:07and x2 into the regression. We'll put x1
- 0:09and x3 into the regression. And we'll
- 0:11put x2 and x3 into the regression. So
- 0:14this is a subset of two variables, but
- 0:16the interpretation is basically the
- 0:20same. So this first one we have travel
- 0:23time, our dependent variable versus
- 0:25miles traveled together with number of
- 0:28deliveries. So our x1 and
- 0:31x2. Let's look at our regression here.
- 0:34We have an f value of
- 0:3823.72 with a p value of
- 0:410001. Now for now I want you to ignore
- 0:44the f value and p value for the
- 0:48individual variables. We'll get to that
- 0:50in a later video. So just look at the
- 0:52regression line across the top here. So
- 0:5523.72 F value P value of 0001. Of course
- 0:59that is significant. Now let's look down
- 1:02below. We have a standard error of the
- 1:05regression of
- 1:08352642. So that's again that's in the
- 1:10ballpark of what we had in the one
- 1:13variable models. We have an R squar of
- 1:1687.14, an adjusted R squar of 83.47%
- 1:2047% and then an R squar predicted of
- 1:2659.95%. So that's significantly lower
- 1:28and we'll talk about what that means
- 1:30later in the video. So now let's look at
- 1:33the coefficients and this is where
- 1:35things get interesting. So let's look at
- 1:38the miles traveled coefficient. So its
- 1:42value is
- 1:430262. Its t value is
- 1:471.31 and its p value is
- 1:52232. That is not significant. It's not
- 1:56below 05. Now look at number of
- 1:59deliveries. Coefficient
- 2:01of84, t value of 73 and a p value of
- 2:08487. That is not significant either.
- 2:12So, here's the weird thing. We have an
- 2:15overall model that is significant. Okay,
- 2:19up the top we have an F value of
- 2:2223.72 and we have a p value of
- 2:250.001. But down below, neither of our
- 2:29coefficients are statistically
- 2:31significant. Very strange. And then of
- 2:34course down below we have our regression
- 2:37equation. So remember what we discussed
- 2:39before, these two variables X1 and X2
- 2:43are extremely correlated with each
- 2:46other. So they they're almost on a
- 2:48straight line that goes from bottom left
- 2:50to top right. The correlation is N56.
- 2:54That is very significant. So these two
- 2:57variables are
- 2:59multicolinear. Now look what happens. We
- 3:02have two variables that are collinear.
- 3:05The overall model is significant, but
- 3:08the individual coefficients are not. Now
- 3:11you see what happens when we put two
- 3:14variables in a regression that are
- 3:17correlated, that have high levels of
- 3:20colinearity. The coefficients in the
- 3:23regression model go haywire. They go
- 3:26crazy. So this is why we look to
- 3:27eliminate variables from the get-go.
- 3:30Now, Mini Tab also has this variance
- 3:33inflation factor or the VIF. Now, how
- 3:36this helps us is that it points out
- 3:39variables that are collinear that points
- 3:43out
- 3:44multi-olinearity. So, we'll talk about
- 3:45it more as we go, but a VIF of
- 3:5011.59 should send alarm bells off um in
- 3:53our statistical minds. We should know
- 3:56that is a serious problem. And this
- 3:59model we have in front of us is very
- 4:01suspect. And we can base that on several
- 4:04criteria. We know from our scatter plots
- 4:07that they're highly correlated, very
- 4:09highly correlated. We have this weird
- 4:11situation where the overall model is
- 4:13significant, but the individual
- 4:15coefficients are not. And then we have
- 4:18this VIF that's through the roof. So all
- 4:21that taken together should tell us that
- 4:24this model has some problems. Now, we'll
- 4:26keep it in here just for learning
- 4:28purposes, but just keep in mind the
- 4:30criteria here. And I also want to point
- 4:33out that when we have an R squar
- 4:36adjusted, that's
- 4:3983.47%. Then the R squar predicted falls
- 4:42off a cliff. Now it's at
- 4:4559.95%. That tells us that we have a
- 4:49serious problem in our model. So all
- 4:52that taken together lets us know that
- 4:53something is wrong.
- 5:00So in our ANOVA table we have a
- 5:02regression line here. We have an F value
- 5:04of
- 5:0522.63. That is a p value of 0.001. So
- 5:09that is significant. Down here in the
- 5:11model summary, we have a standard error
- 5:13of the regression
- 5:15of.359883 hours. Of course, we have an R
- 5:18squared of 86.61% 61% R squared adjusted
- 5:22of
- 5:2382.78% and an R squar predicted of
- 5:2868.11%. Now notice couple things here.
- 5:31Of course, our ANOVA table overall is
- 5:34significant. Our R squar is adjusted as
- 5:37very high and our R 2 predicted doesn't
- 5:41fall off a cliff like we saw in the
- 5:44previous example with X1 and X2.
- 5:47So here are our coefficients. So in this
- 5:50case we have miles traveled which is
- 5:530.04137. That's our coefficient for
- 5:55miles traveled. T value of 6.44 with a p
- 5:58value 0. So that's fine. Now look at gas
- 6:03price. We have a negative coefficient
- 6:08-2.19. Now think about what this means.
- 6:11So what this is saying is that if we
- 6:14hold miles traveled constant, that's our
- 6:17x1. If we hold that constant, and we
- 6:20increase the price of gas a dollar, then
- 6:25the travel time will decrease by 219
- 6:29hours. So gas price goes up and the
- 6:33travel time goes down. Does this make
- 6:36any real sense to me? It does not. So, I
- 6:40don't know about you, when gas prices go
- 6:41up, I drive slower. But this is saying
- 6:44the travel time goes down. And what this
- 6:48points out is that we are putting this
- 6:50gas price variable in the regression,
- 6:52but it doesn't have any real
- 6:54relationship to the dependent variable.
- 6:57Remember, so now we get some very weird
- 7:00coefficients down here at the bottom. So
- 7:02that's why we try to eliminate variables
- 7:04up front because it really messes up the
- 7:07coefficients that come out of the
- 7:08regression process. So let's interpret
- 7:10both of these coefficients. First, miles
- 7:12traveled. So if gas price is held
- 7:15constant, then travel time is expected
- 7:17to increase by
- 7:2104137 hours for each additional mile
- 7:24traveled. Now does that make sense in
- 7:27real life? Well, yes it does. So if I
- 7:29travel further more miles, I expect my
- 7:32travel time to go up. Now how about gas
- 7:35price? If miles traveled is held
- 7:37constant, then travel time is expected
- 7:40to decrease by 219 hours for each
- 7:45additional dollar increase in gas price.
- 7:48And again, that really does not make any
- 7:50sense in real life and in statistics
- 7:54either because this coefficient is just
- 7:56kind of weird. So again, we're going to
- 7:58keep this one kind of off to the side,
- 8:00noting we have some weird coefficients
- 8:02in the gas price, and that's probably
- 8:05because we included a variable in the
- 8:08regression that has no relationship to
- 8:10the dependent variable to begin with. So
- 8:12let's go ahead and do the last two
- 8:15pair. So this is X2 and X3. So numbum
- 8:18deliveries and gas price. So look at the
- 8:21ANOVA table, the regression line. We
- 8:23have an F value of 27.63 63. P value is
- 8:280. So we know that's significant. We
- 8:31have a standard error of regression down
- 8:33here at the bottom of
- 8:36329703 hours. R squared of
- 8:4088.76%, R squared adjusted of
- 8:4385.55%. And an R squar predicted of
- 8:4871.76%. So those are all pretty high.
- 8:50The R squared predicted does not go off
- 8:52a cliff like we saw in the first
- 8:54example. So, so far things don't seem
- 8:56too crazy. Now, look at our
- 8:59coefficients. Uh-oh, we have the same
- 9:02problem again. So, number of deliveries,
- 9:04that looks fine.
- 9:05So,.5665, T value of 7.13, P value 0.
- 9:11Fine, looks good. Gas price has gone
- 9:13negative again.
- 9:16So,.765, T value
- 9:19of.172. The P value is not significant
- 9:22in this case. So we have this weird
- 9:25situation again where the gas price
- 9:28coefficient went
- 9:30negative. So interpret these again. If
- 9:33gas prices is held constant then travel
- 9:35time is expected to increase
- 9:38by.5665 hours for each additional
- 9:41delivery. Now does that make sense in
- 9:43real life? Well yes. I expect the travel
- 9:47time to go up for each additional
- 9:49delivery I have to make. Now how about
- 9:52the gas price problem? If number of
- 9:55deliveries is held constant, then travel
- 9:57time is expected to decrease by 765
- 10:01hours for each additional dollar
- 10:03increase in gas price. That doesn't make
- 10:06any sense. So again, we have this
- 10:08problem where we included a variable in
- 10:11the model that has no natural relation
- 10:14to the dependent variable whatsoever. It
- 10:16messes up our coefficients and really
- 10:18this model is no good. So let's go ahead
- 10:21and summarize these three
- 10:25models. So the top three lines are the
- 10:28first three models we did. So that's our
- 10:30single variable models. Now let's look
- 10:33at the second three. So in our first two
- 10:36variable model, we had x1 and
- 10:38x2. So we had an f of 23.72.
- 10:43Now we expect the fs to be about the
- 10:46same for each one variable model and
- 10:49each two variable model etc. Okay. So
- 10:5223.72 we have a p value 01. That's fine.
- 10:56Now we have a standard error of
- 10:58regression of
- 11:0235264. So remember what that tells us
- 11:04that tells us how tied in our data
- 11:09points are to the regression line. So in
- 11:11this case they are on
- 11:14average.35264 hours away from the
- 11:17regression line and you can compare that
- 11:18to the ones we have above. So our R
- 11:22squar is
- 11:2383.47. Our R square predicted is
- 11:2659.95. That's a huge drop off from the R
- 11:29squ adjusted. Then we have this VIF over
- 11:33here of
- 11:3511.59. That is huge and that is a
- 11:37problem. And that's because x1 and x2
- 11:41are collinear. That's the problem we
- 11:43have there. Now the second one from the
- 11:45bottom that's x1 and x3. So we have
- 11:4822.63 for the f0001 for the p value. The
- 11:52standard error of the regression 35988
- 11:56hours. And then we have R squ adjusted
- 11:58at
- 11:5982.78%. R square predicted
- 12:0268.11. Everything looks pretty much okay
- 12:05there. Then we have a VIF of 1.14.
- 12:08That's not a problem. But remember from
- 12:10our
- 12:11coefficients, we had a negative X3
- 12:14coefficient, which is very weird. So
- 12:17even though everything in that row looks
- 12:20okay, we know that we have a coefficient
- 12:22oddity. So we have to keep that in mind.
- 12:24And then finally here we have the X2X3
- 12:28model. So 27.63.
- 12:31Then we have the p value less than 0001.
- 12:34Standard error of the regression of
- 12:3732970. Again, that's in the ballpark of
- 12:39everything else. But if you look above
- 12:42it, you can see that so
- 12:44far that is the best fit around the
- 12:48regression line. So on average
- 12:5232970 hours away from the regression
- 12:55line, R squared adjusted of 85.55%.
- 13:00Now look at that column. That's the
- 13:03highest adjusted R squar we've had. Now
- 13:06go over to the R square predicted.
- 13:08That's 71.76. There's nothing really
- 13:10spectacular there relative to everything
- 13:12else. And then of course a VIF of 1.33.
- 13:16No problem there. So we have to decide
- 13:19here. We have this last one with a
- 13:23higher F than the two above it. We have
- 13:26a smaller standard error of the
- 13:28regression, which is what we'd like to
- 13:30see. We have a relatively high R squar
- 13:32adjusted at
- 13:3485.55%. In fact, it's the highest in
- 13:36that column there. And the R square
- 13:38predicted is what we'd expect. But
- 13:42remember from the coefficients, this is
- 13:45another example of where we have a
- 13:48negative coefficient. We have a negative
- 13:52gas price coefficient. So even though
- 13:54everything looks okay here, we also have
- 13:57to keep in mind our coefficients from
- 13:59the previous step. So that might be a
- 14:01problem. So let's go ahead and define
- 14:03what VIF actually is. And I just quoted
- 14:07this from many tabs blog. The URL is
- 14:10down here at the bottom. Now let's go
- 14:11ahead and quickly read what it says. So
- 14:14one way to measure multiolinearity is
- 14:16the variance inflation factor or the VIF
- 14:20which assesses how much the variance of
- 14:23an estimated regression coefficient
- 14:25increases if your predictors your
- 14:28independent variables are
- 14:30correlated. If no factors are correlated
- 14:33if no independent variables are
- 14:35correlated the VIFs will all be one. Now
- 14:40let me pause there. Look at the VIFs for
- 14:43the first three models. They're all
- 14:44exactly one. Well, why is that? Well,
- 14:48there's only one independent variable in
- 14:49them. So, they're going to be one. There
- 14:52is no correlation there. So, they'll all
- 14:54be one. Now, a VIF between five and 10
- 14:58indicates high correlation. That may be
- 15:01problematic. So, do we have any between
- 15:03five and 10? Uh, nope. Not so far. Now
- 15:07if the VIF goes above 10, you can assume
- 15:11that the regression coefficients are
- 15:13poorly estimated due to
- 15:17multicolinearity. So look at the first
- 15:19two variable model. We have a VIF of
- 15:2311.59. Now remember why that is. Our two
- 15:27independent variables X1 and X2 had a
- 15:31correlation above N5. They were
- 15:34extremely highly correlated. So that VIF
- 15:38the variance inflation factor points out
- 15:40that hey you have a problem there you
- 15:43have some multiolinearity some severe
- 15:46multiolinearity in that model and
- 15:48therefore we would just ax that model
- 15:51out of the we would just forget it so
- 15:53we'll leave it there for now but just
- 15:55know that the vif helps us find
- 16:00multiolinearity okay and finally the
- 16:03full model we're going to throw in all
- 16:05three independent depent variables and
- 16:07see what
- 16:09happens. Okay, so here is the ANOVA
- 16:12table from Mini Tab for all three
- 16:14independent variables. So let's look at
- 16:16the regression lineup here. We have an F
- 16:19value of
- 16:2116.99 with a p value of 02. So the
- 16:26overall model is significant. So the
- 16:28model summary, we have a standard error
- 16:30of the regression of
- 16:33344694 hours.
- 16:35We have an R squared of
- 16:3789.47, an R squared adjusted of
- 16:4184.2% and an R 2 predicted of
- 16:4657.49. Now, what's the red flag there?
- 16:49The R squared adjusted was
- 16:5184.20. The R square predicted is
- 16:5757.49. That's a huge drop
- 17:01off. So, the coefficients real quickly.
- 17:04So miles traveled had a coefficient of
- 17:080141. Numb deliveries was 383. Then we
- 17:11have the strange gas price
- 17:14that's.607 again. Now if we look at our
- 17:17p values, it gets even more strange. So
- 17:20the p value for miles traveled is 548.
- 17:24That is not significant. The p value for
- 17:26numbum deliveries
- 17:27is.249. Not significant. Guest price
- 17:31293. Not significant. even though it
- 17:34really doesn't matter because that's a
- 17:36junk variable at this point. Now, if you
- 17:38look at the VIFs, look at
- 17:42those. For miles traveled, it's
- 17:4614.94. For number deliveries, it's
- 17:5017.35. So, what does that tell us? We
- 17:53have severe severe problems with
- 17:57multiolinearity in this model. Severe
- 18:00problems. basically terminal death
- 18:03problems. But we go ahead and have the
- 18:05regression equation down here at the
- 18:06bottom just for kicks I guess. But
- 18:09basically this model is
- 18:14junk. So here are all of our models put
- 18:17together. So we're getting to the grand
- 18:19finale finally. So at the bottom we have
- 18:22this new model with an f of 16.99 p
- 18:25value
- 18:2602. Uh the standard error of the
- 18:28regression
- 18:3134469 R squared adjusted
- 18:3484.2%. R square predicted
- 18:3857.49%. Then we have our VIFs. I put
- 18:42those below each variable because I ran
- 18:44out of room. So for X1 it was 14.94, X2
- 18:4817.35, X3
- 18:511.71. So we can see that that's a
- 18:54problem. So step back and look at this
- 18:56last one again. We can see that we have
- 18:57a huge drop off from the R squared
- 19:00adjusted to the R square predicted just
- 19:02like we do at the top of the two
- 19:04variable models where we went from 8347
- 19:07to
- 19:0859.95. Now here is the question. Which
- 19:12model is the
- 19:14best? So let's start with knocking out
- 19:17some models. Well, we know the last
- 19:20model with all three variables is junk.
- 19:24The VIFs are sky-high. The R square
- 19:28predicted is way lower than the R
- 19:30squared adjusted. So that model is no
- 19:33good. Now we can rule out also the top
- 19:36of the two variables. So again there we
- 19:39have a VIF of
- 19:4111.59 and the R square predicted falls
- 19:44off a cliff from
- 19:4583.47 for the R squ adjusted. That one's
- 19:49gone. So we can rule out those two right
- 19:53off the
- 19:54bat. Now how do we decide? So here is
- 19:58sort of the golden rule of choosing your
- 20:00multiple regression model. We want to
- 20:03look at several factors. We want to look
- 20:05at the R squar adjusted. We want the
- 20:09highest one we can get. We want the R
- 20:11square predicted to be as high as we can
- 20:14get and to be close to the R squ
- 20:17adjusted which is already high. We'd
- 20:20like to see a relatively small standard
- 20:22error of the regression. So that's the S
- 20:24column over here on the left. And
- 20:26finally, all else being equal, we want
- 20:29the simplest model there is. So if we
- 20:34look at some other candidates down here,
- 20:36we can see that for the two variable
- 20:38model, we have X1, X3, X2, X3. But those
- 20:44have some serious problems. Remember
- 20:46that the X3 or the gas price coefficient
- 20:49was negative and that's because the gas
- 20:52price
- 20:53coefficient doesn't contribute to the
- 20:55dependent variable at all. Plus, we have
- 20:58some sharp falloff in the R square
- 21:00predicted. So, we're going to rule those
- 21:02out. So, we have ruled out all the two
- 21:04variable options and the three variable
- 21:08option. So, basically we're at the top.
- 21:11We can definitely eliminate the bottom
- 21:14one variable with just x3 in it. We know
- 21:17x3 is basically a junk variable. So we
- 21:19can x that out. Now we have to decide
- 21:24between the top two. That's all we have
- 21:26left. So is it going to be the top model
- 21:29with x1 or the next one with just
- 21:32x2? Well, I think it's pretty obvious
- 21:35that the top model with just
- 21:38x1 is the best model.
- 21:41So a one variable model is the best
- 21:46model out of all these options. So we
- 21:50have a very narrow standard error of the
- 21:52regression at 34 and some change a very
- 21:55high R squared a very high R square
- 21:58predicted no multiolinearity problems
- 22:00because well there's only one variable.
- 22:02So guess what that is our best
- 22:08model. So yes, and you are going to kill
- 22:11me, but Mini Tab and I'm sure other
- 22:13stats packages can do all of this,
- 22:15everything we just did by hand looking
- 22:18at the relationships, it can do it in a
- 22:20few clicks. And here it is. This is the
- 22:22output from Mini Tab. It's basically a
- 22:25best subsets regression which we just
- 22:27did step by step. So how do we read this
- 22:29thing? Well, look at the R squared
- 22:32adjusted. Which are the highest values?
- 22:35So we have 844. We look we have
- 22:39855 and we have an 842 in there. So the
- 22:42855 is the highest. Then we have the 844
- 22:45at the top. Now look at the R squar
- 22:48predicted which are the highest values.
- 22:51So we have
- 22:5279.1 at the top and then from there they
- 22:56go down pretty quickly. So nothing
- 22:59really worth mentioning. The 79.1 is
- 23:02definitely the highest R square
- 23:03predicted.
- 23:05Now examine the difference between the R
- 23:07squared adjusted and the R square
- 23:09predicted. A large drop off from the
- 23:13adjusted to the predicted indicates
- 23:16overfitting. That indicates there are
- 23:19too many variables in the model. So we
- 23:22can see that for the ones there at the
- 23:24bottom like the three variable model we
- 23:27go from 84.2 adjusted to
- 23:3157.5 predicted. That is a sign of
- 23:35overfitting and it's a bad thing. Now
- 23:38look at the top. We go from 844 to
- 23:4179.1. That's very close. And the ones
- 23:44below it aren't too bad either. But you
- 23:46can see as we get down with more
- 23:49variables that the drop off is very
- 23:51high. So those models are
- 23:55overfitted. Now look at Maloc.
- 23:58Look for the one that is low and
- 24:00approximately equals the number of
- 24:03predictor variables or independent
- 24:05variables plus the constant which is
- 24:08one. There's one constant. So in the
- 24:10first example we have one independent
- 24:12variable or one predictor variable plus
- 24:14the constant. So that's two. So for the
- 24:17single variable models we're looking for
- 24:19a malo's number that is two. For two
- 24:23variables we look for three. That's 2 +
- 24:261. And then for the three variable, we
- 24:28look at four. So overall, we're looking
- 24:30for the lowest one that's closest to its
- 24:32magic
- 24:33number. Then using all of the
- 24:36information we have above, choose the
- 24:39best model. And based on that info,
- 24:42which is the best model? The first one.
- 24:45So the single variable X1 model that has
- 24:49the very high R squared adjusted the
- 24:51very high R square predicted that
- 24:53doesn't fall off the malo CP that is
- 24:56almost exactly two which is what we want
- 24:58and a relatively narrow standard error
- 25:00of the regression. So at 342 it's kind
- 25:04of in the middle of the pack but it's
- 25:06fine. So overall that is the best model
- 25:10and that is the one we would use believe
- 25:13it or not to make our
- 25:15predictions. Okay. So that was a tour
- 25:17day force of how to evaluate multiple
- 25:21regression models. Yes, that was long. I
- 25:25admit that. But you'll come out of it
- 25:28never having to really doubt or question
- 25:30your knowledge of how multiple
- 25:32regression works, how the best models
- 25:34are built, at least in the linear cases.
- 25:37So when you go to take your test, write
- 25:39your paper, write a report at work,
- 25:41you'll be able to create the best
- 25:43models, make the best predictions on
- 25:45those models, and substantiate any
- 25:47findings or suggestions you make,
- 25:49whether it's in a report or on a paper
- 25:52or whatever else it might be. So, I know
- 25:54that was long, so I'll let you go. Thank
- 25:56you very much for watching. Please
- 25:57subscribe. If you like the video, give
- 25:59it a thumbs up. And I look forward to
- 26:01seeing you again next time.
- 26:05[Music]
About this transcript
This page contains the full transcript of YouTube transcript (px72eCYPuvc) , generated from the public captions YouTube serves with the video. The transcript has 3,703 words across 555 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.