YouTube2Text

YouTube transcript (wPJ1_Z8b0wk) — Transcript

4,041 words · 607 segments · language en · Watch on YouTube

Full transcript

  1. 0:02[Music]
  2. 0:07Hello, thanks for watching and welcome
  3. 0:10to the next video in my series on basic
  4. 0:12statistics. Now, as usual, a few things
  5. 0:15before we get started. Number one, if
  6. 0:17you're watching this video because you
  7. 0:19are struggling in a class right now, I
  8. 0:21want you to stay positive and keep your
  9. 0:23head up. If you're watching this, it
  10. 0:25means you've accomplished quite a bit
  11. 0:26already. You're very smart and talented,
  12. 0:29but you may have just hit a temporary
  13. 0:31rough patch. Now, I know with the right
  14. 0:33amount of hard work, practice, and
  15. 0:35patience, you can work through it. I
  16. 0:38have faith in you. Many other people
  17. 0:40around you have faith in you. So, so
  18. 0:43should you. Number two, please feel free
  19. 0:46to follow me here on YouTube, on
  20. 0:48Twitter, on Google+, or on LinkedIn.
  21. 0:52That way, when I upload a new video, you
  22. 0:54know about it. And it's always nice to
  23. 0:56connect with my viewers online. I feel
  24. 0:59that life is much too short and the
  25. 1:00world is much too large for us to miss
  26. 1:02the chance to connect when we can.
  27. 1:05Number three, if you like the video,
  28. 1:07please give it a thumbs up. Share it
  29. 1:10with classmates or colleagues or put it
  30. 1:12on a playlist. That does encourage me to
  31. 1:14keep making them for you. On the flip
  32. 1:17side, if you think there's something I
  33. 1:18can do better, please leave a
  34. 1:20constructive comment below the video,
  35. 1:21and I will take those ideas into account
  36. 1:24when I make new ones. And finally, just
  37. 1:27keep in mind that these videos are meant
  38. 1:29for individuals who are relatively new
  39. 1:31to stats. So, I'm just going over basic
  40. 1:34concepts, and I will be doing so in a
  41. 1:37slow, deliberate manner. Not only do I
  42. 1:40want you to know what is going on, but
  43. 1:43also why and how to apply it. So, all
  44. 1:46that being said, let's go ahead and get
  45. 1:49started.
  46. 1:51Hello and welcome to part three in our
  47. 1:53series on multiple regression. Now, I am
  48. 1:56going to assume you watched parts one
  49. 1:58and two of this series or you come to
  50. 2:00this video with that knowledge already
  51. 2:02in your head. So, if you need to go
  52. 2:05back, watch those two videos and then
  53. 2:07come back to this one. Otherwise, you
  54. 2:09may end up watching this one and become
  55. 2:11confused and frustrated. And neither of
  56. 2:13us want that to happen. So here in part
  57. 2:16three, we're going to talk about
  58. 2:17building regression models. So we will
  59. 2:20do that by introducing our independent
  60. 2:21variables into the model one at a time.
  61. 2:24Then we'll add them two at a time. And
  62. 2:26then in this case, we'll add all three.
  63. 2:28And then we'll see the models that come
  64. 2:30out of that process and then make a
  65. 2:32judgment call on which model is best. So
  66. 2:35we'll talk about all that criteria as
  67. 2:37far as determining the best model as we
  68. 2:39go forward. So all that being said,
  69. 2:41let's go ahead and get started.
  70. 2:46So as we discussed in parts one and two,
  71. 2:48conducting a multiple regression
  72. 2:50analysis requires a fair amount of
  73. 2:52pre-work before actually running the
  74. 2:54numbers in your software. So here are
  75. 2:56those steps real quickly. We generate a
  76. 2:59list of potential variables. So
  77. 3:01independence and the dependent variable.
  78. 3:03We collect data on those variables. We
  79. 3:06check the relationships between each
  80. 3:07independent variable and the dependent
  81. 3:09variable using scatter plots and
  82. 3:11correlations. So here we want to see
  83. 3:14which independent variables are actually
  84. 3:16related to the dependent variable in the
  85. 3:19first place because the ones that are
  86. 3:21not are probably going to be left out of
  87. 3:23the analysis.
  88. 3:25Then we'll check the relationships among
  89. 3:27the independent variables themselves
  90. 3:30using scatter plots and correlations. So
  91. 3:32here we're checking for multi-olinearity
  92. 3:36and then usually optional but I do it
  93. 3:38anyway and that is conduct simple linear
  94. 3:40regressions for each independent
  95. 3:42variable dependent variable pair and
  96. 3:45that's what we will do in this video
  97. 3:47because here we want to learn how to
  98. 3:48build models up and then tear them down
  99. 3:51to find the best one.
  100. 3:53Then we'll use the non-redundant
  101. 3:55independent variables in the analysis to
  102. 3:58find the best fitting model. And
  103. 4:00finally, most likely in a later video,
  104. 4:02we'll use the best fitting model to make
  105. 4:04predictions about the dependent
  106. 4:06variable. And we'll talk about
  107. 4:08confidence intervals and prediction
  108. 4:09intervals at that point.
  109. 4:14So remember, you are a small business
  110. 4:16owner that runs a delivery service like
  111. 4:18a courier service. You do same day
  112. 4:20deliveries of packages, letters, small
  113. 4:22cargo, etc. And you want to be able to
  114. 4:25predict the total travel time for each
  115. 4:29trip. So on a trip, you may have several
  116. 4:31deliveries. They may be close to your
  117. 4:33office, they may be far away, etc. So
  118. 4:36you go back and look at 10 past trips
  119. 4:38and record four pieces of information.
  120. 4:40The total miles traveled for that trip,
  121. 4:42the number of deliveries during that
  122. 4:44trip, the daily gas price, and finally
  123. 4:47the total travel time in hours, which is
  124. 4:50your dependent variable or the thing
  125. 4:51you're interested in predicting. So
  126. 4:54here's the data. We have miles traveled,
  127. 4:55that's our first independent variable,
  128. 4:57X1. number of deliveries during the
  129. 4:59trip, that's x2. And the gas price that
  130. 5:02day, that's x3. And then our dependent
  131. 5:05variable, travel time, is y there on the
  132. 5:07right.
  133. 5:11So remember, it's always a good idea to
  134. 5:13visualize the relationships. So we have
  135. 5:15our travel time dependent variable. Then
  136. 5:17we have three independent variables. So
  137. 5:19we have miles traveled, number of
  138. 5:22deliveries,
  139. 5:24and gas price. Now, of course, all three
  140. 5:26of those have some relationship to the
  141. 5:28dependent variable. We don't know that
  142. 5:29yet, but we will. But we also have to
  143. 5:31account for the relationships among the
  144. 5:33independent variables themselves, which
  145. 5:35are there in the dash line. So, we have
  146. 5:38six relationships we have to analyze and
  147. 5:40keep in mind as we do our analysis.
  148. 5:44Let's quickly review the scatter plots
  149. 5:46comparing the dependent variable and the
  150. 5:48independent variables individually. So
  151. 5:50the first scatter plot we have our
  152. 5:52travel time dependent variable versus
  153. 5:55our miles traveled or first independent
  154. 5:57variable and as you can see there we
  155. 5:59have a strong linear relationship starts
  156. 6:01at the lower left of the graph and goes
  157. 6:02to the upper right we have a correlation
  158. 6:04coefficient of 928 that's very very high
  159. 6:08we have a p value for the correlation of
  160. 6:090 which means it's less than 0.1 so we
  161. 6:14know that that variable miles traveled
  162. 6:17is related strongly related to our
  163. 6:19dependent variable able travel time. So
  164. 6:21we'll put a green check there.
  165. 6:24The second independent variable number
  166. 6:25of deliveries X2 that's also very
  167. 6:28strongly related. So the line starts in
  168. 6:30the lower left goes to the upper right.
  169. 6:32They fall along a rough line there. The
  170. 6:35correlation is 916. Again very high. The
  171. 6:38p value is less than 0001. So that is
  172. 6:41significant. So we'll put a green check
  173. 6:43there. So our first two independent
  174. 6:45variables do have strong linear
  175. 6:48relationships. very strong correlations
  176. 6:50with our dependent variable, which is
  177. 6:52good.
  178. 6:54Now, our third independent variable, gas
  179. 6:56price, does not. As you can see, the
  180. 6:59data points do not form any pattern.
  181. 7:01They're kind of all over the place. And
  182. 7:03we can see that in the correlation. It's
  183. 7:06267 with a p value of 0455.
  184. 7:10That is not significant. So, we'll put a
  185. 7:12red X there. So gas price X3 does not
  186. 7:17have any sort of linear relationship to
  187. 7:20the dependent variable right off the
  188. 7:22bat. So usually we would just go ahead
  189. 7:25and remove that from the model because
  190. 7:27if it doesn't have one to begin with,
  191. 7:29it's not going to contribute anything to
  192. 7:31the regression. But of course for now
  193. 7:34we're going to leave it in and see how
  194. 7:35it affects our numbers as we go forward.
  195. 7:40So now we have the scatter plots for the
  196. 7:42independent variable comparisons. So
  197. 7:44again, we're looking for
  198. 7:44multi-olinearity.
  199. 7:46So the first scatter plot, we have miles
  200. 7:48traveled versus number of deliveries.
  201. 7:51And here we have a problem. So our first
  202. 7:54two independent variables have a very
  203. 7:57very high correlation, a 0.956, a p
  204. 8:01value less than 0001. And I'll put a
  205. 8:04skull and crossbones there cuz this is a
  206. 8:06problem. two independent variables that
  207. 8:09are this highly correlated are going to
  208. 8:11cause some serious issues with our
  209. 8:13regression coefficients going forward.
  210. 8:16So, we'll keep that in mind as we go.
  211. 8:17I'll leave them in there for now so we
  212. 8:19can see how it affects things going
  213. 8:20forward.
  214. 8:22Now, miles traveled X1, gas price X3, no
  215. 8:25problems there. And then number of
  216. 8:28deliveries X2 and gas price X3, no
  217. 8:31problems there. So the only problem we
  218. 8:33have to worry about as far as
  219. 8:34multiolinearity goes is the correlation
  220. 8:37between the first two independent
  221. 8:38variables.
  222. 8:42So the quick summary. So the correlation
  223. 8:44analysis confirms the conclusions we
  224. 8:46reached by visual examination of the
  225. 8:48scatter plots. So we have some redundant
  226. 8:51multi-colinear variables. So miles
  227. 8:53traveled and number of deliveries are
  228. 8:55both highly correlated with each other
  229. 8:57and therefore are redundant. only one
  230. 9:00should be used in the regression
  231. 9:01analysis in the end. And we'll see how
  232. 9:04that all pans out as we build our model.
  233. 9:06We also have a non-contributing
  234. 9:08variable. So gas price is not correlated
  235. 9:11with the dependent variable at all and
  236. 9:14should probably be excluded right off
  237. 9:15the bat. Now again, for educational
  238. 9:19purposes, I'm going to leave all three
  239. 9:21variables in so we can see how putting
  240. 9:24them in there affects the regressions we
  241. 9:27do. But in the end, it'll all work out.
  242. 9:29You'll see.
  243. 9:32Okay. So, let's go ahead and get into
  244. 9:33some of our single variable regressions.
  245. 9:38Now, in this first step, we will perform
  246. 9:39a simple regression for each independent
  247. 9:42variable individually. The first will be
  248. 9:44conducted in Excel and then the rest in
  249. 9:46Mini Tab. That's what I prefer right
  250. 9:48now, but SPSS, SAS, Jump, R, etc. um or
  251. 9:52offline as well. You can get them all to
  252. 9:54generate basically the same output. So
  253. 9:57depending on which one you use, you
  254. 9:58should be fine. But I'll be focusing on
  255. 10:00a little bit of Excel and the rest
  256. 10:02MiniAB.
  257. 10:04We will discuss interpretations of the
  258. 10:06results we get from Mini Tab. And we
  259. 10:09will note how our results change. So
  260. 10:12we'll look at the coefficients in the
  261. 10:14regression. We'll look at the
  262. 10:16coefficient values. We'll look at their
  263. 10:19t statistics and we'll look at their p
  264. 10:21values.
  265. 10:22We'll look at the ANOVA table in the
  266. 10:24regression. So we'll look at the F value
  267. 10:27and the P value in that ANOVA table.
  268. 10:31We'll look and talk about the R 2, the R
  269. 10:342 adjusted and the R squar predicted and
  270. 10:37talk about what those mean. We'll also
  271. 10:39look at something called the VIF or the
  272. 10:41variance inflation factor and that's a
  273. 10:43statistic that MiniAB produces that will
  274. 10:45help us weed out multiolinearity.
  275. 10:49And we will also talk about something
  276. 10:51called malocp. That is a statistic that
  277. 10:54many tab outputs that will help us pick
  278. 10:56the best model in the end.
  279. 11:00So let's go ahead and look at our first
  280. 11:02regression. So we're going to regress
  281. 11:04travel time y which is our dependent
  282. 11:06variable on miles traveled which is our
  283. 11:08first independent variable. And here are
  284. 11:10our results from Excel. So we basically
  285. 11:12have three tables here. The first one in
  286. 11:14the top left are our regression
  287. 11:16statistics. Now in this case, multiple R
  288. 11:19is the same thing as our correlation.
  289. 11:21Since we only have one independent
  290. 11:23variable, they're the same thing. So
  291. 11:250.928, etc. is the same as the
  292. 11:28correlation we had a couple slides ago.
  293. 11:30Now R square is the proportion or
  294. 11:34percentage of variation in the dependent
  295. 11:37variable accounted for by the
  296. 11:39independent variable. So we can look at
  297. 11:42it as a percentage. So 86.15%
  298. 11:46of the variation in the dependent
  299. 11:48variable is accounted for by the
  300. 11:50independent variable. That's pretty
  301. 11:52high. Now the adjusted R square, that's
  302. 11:55the same thing as the R square.
  303. 11:57Obviously, it is just adjusted for the
  304. 11:59number of independent variables in our
  305. 12:01model, which in this case is one. So it
  306. 12:04will always be lower than the R squared.
  307. 12:07And how much lower really depends on the
  308. 12:09specific circumstances we're in. So the
  309. 12:12next number is the standard error of the
  310. 12:14regression. This is one of my favorite
  311. 12:16numbers, but unfortunately most people
  312. 12:19don't know how to use it or don't use it
  313. 12:21or just skip it or whatever else, but I
  314. 12:23think it's very helpful. So the standard
  315. 12:26error of the regression is the average
  316. 12:29distance of the data points from the
  317. 12:32regression line in dependent variable
  318. 12:35units. What we're saying here is that
  319. 12:37the data points are on average 342
  320. 12:43hours away from the regression line. It
  321. 12:46is in the units of the dependent
  322. 12:48variable and it gives us a measure of
  323. 12:51how tightly around the regression line
  324. 12:55our data points are. So it kind of forms
  325. 12:58a a channel or a band around the
  326. 13:03regression line. And the narrower that
  327. 13:05is, the more tightly our data points are
  328. 13:08around the regression. And the wider
  329. 13:10that band is, the more scattered they
  330. 13:13are from that regression line. So the
  331. 13:15standard error of the regression tells
  332. 13:17us relatively speaking how wide that
  333. 13:20band around the regression line is. And
  334. 13:23it's also helpful because it is in the
  335. 13:25units of the dependent variable. In this
  336. 13:27case, hours. And of course, we have 10
  337. 13:29observations. That's pretty
  338. 13:30self-explanatory. Now the ANOVA table
  339. 13:33that gives us the significance of the
  340. 13:35overall model. So we have an F statistic
  341. 13:38there of 49.768
  342. 13:40etc. with a p value of 0.00001
  343. 13:45that of course is significant. So the
  344. 13:48overall model here is significant. Now
  345. 13:51at the bottom we have some of our
  346. 13:53coefficient information. Now we're
  347. 13:55interested in the miles traveled
  348. 13:57coefficient. So under coefficients we
  349. 14:00have 0.0402
  350. 14:03that is the coefficient of our miles
  351. 14:05traveled and again that is in hours. So
  352. 14:08what we're saying there is that for
  353. 14:10every mile that's increased the time
  354. 14:14traveled increases by 042
  355. 14:18hours. Then we have the p value which is
  356. 14:210.1. So we know it is significant. Now,
  357. 14:24if you notice in this case, the p value
  358. 14:26for the miles traveled coefficient is
  359. 14:29the same as the p value for the ANOVA.
  360. 14:32And that's because we only have one
  361. 14:34independent variable. So, how can we use
  362. 14:36this information? Well, we can take our
  363. 14:39coefficient information at the bottom
  364. 14:40and generate our regression equation.
  365. 14:43So, we have an intercept of 3.1855, and
  366. 14:45again, that's rounded up here at the
  367. 14:47top. Plus 00403.
  368. 14:50That is our coefficient for miles
  369. 14:51traveled down here at the bottom. and
  370. 14:53then times the miles traveled. That's
  371. 14:55our independent variable. That's our x1.
  372. 14:58So we have 3.1856
  373. 15:00plus 0403
  374. 15:03x1 where x1 is miles traveled. So what
  375. 15:06does that mean? An increase in 1 mile.
  376. 15:10So miles traveled x1. So one mile will
  377. 15:13increase delivery time by 043
  378. 15:18hours. And that's how we can interpret
  379. 15:20the coefficient in this simple
  380. 15:22regression. Now, let's go ahead and make
  381. 15:24a rough prediction for miles traveled
  382. 15:27that's sort of within the range of our
  383. 15:29original data. So, we'll pick an 84 mile
  384. 15:32trip estimate. So, we go ahead and
  385. 15:35substitute 84 in for our x1. That gives
  386. 15:38us 6.5708
  387. 15:41hours. So, that's a very rough estimate
  388. 15:44of how long it would take for an 84 mile
  389. 15:47trip. So 6 hours and 34 minutes.
  390. 15:51But remember this is just an estimate.
  391. 15:53So it's going to have an interval around
  392. 15:55it. It's going to have some error around
  393. 15:57it. Now we can go ahead and find that
  394. 15:59prediction interval using some things we
  395. 16:02already know. And again what I've done
  396. 16:04here is a very rough sort of estimate.
  397. 16:07Mini tab can give us exact numbers but a
  398. 16:10rough estimate of the prediction
  399. 16:11interval for an 84 mile trip. So we have
  400. 16:146.5708
  401. 16:16plus or minus 2.31.
  402. 16:19Well, where do I get 2.31?
  403. 16:22That comes from our t distribution, our
  404. 16:24t table. So in this example, we have n
  405. 16:28minus 2 degrees of freedom. So that's 10
  406. 16:31minus 2 in this case because we have 10
  407. 16:33observations. So we go to our t table.
  408. 16:36We look at degrees of freedom of eight.
  409. 16:39Then we look down the table for a 95%
  410. 16:42interval and we have a critical t of
  411. 16:452.31. That's where that comes from. And
  412. 16:48then we have 3423.
  413. 16:51Where does that come from? Well, that's
  414. 16:54my magic number over here on the left.
  415. 16:56The standard error of the regression. So
  416. 16:58this is very much like any other
  417. 17:00interval we calculated back in interval
  418. 17:03estimation. So we have a point
  419. 17:05estimator. So 6.5708
  420. 17:08plus or minus the t alpha / 2 which is
  421. 17:122.31
  422. 17:14times the error which in this case
  423. 17:16is.3423.
  424. 17:19I don't expect you to sort of get this
  425. 17:20right now but I just want to show you
  426. 17:21sort of how we use what we have here. So
  427. 17:24that creates an interval of 5.7764
  428. 17:28to 7.3615
  429. 17:30hours or 5 hours 47 minutes to 7 hours
  430. 17:36and 22 minutes. That's our 95%
  431. 17:39prediction interval for an 84 mile trip.
  432. 17:43So again, that's sort of a step forward
  433. 17:46what we're going to do in future videos,
  434. 17:48but I just wanted to quickly show you
  435. 17:50how we can use the information we get
  436. 17:53from a regression to make some rough
  437. 17:55predictions for other values.
  438. 17:59So here is our second one toone
  439. 18:01regression. So we have our second
  440. 18:03independent variable number of
  441. 18:05deliveries and our dependent variable
  442. 18:07travel time. And this comes from Mini
  443. 18:09Tab. Now obviously it looks a little bit
  444. 18:11different. Now, I will say that one of
  445. 18:14the best skills to have if you're doing
  446. 18:16statistics or whatever else is to be
  447. 18:18able to use really any software package
  448. 18:21or at least look at the results of any
  449. 18:24software package and know how they
  450. 18:27correspond to each other. So, in this
  451. 18:29case, the F value of 41.96
  452. 18:34that's along the regression line is the
  453. 18:38same F value we had in Excel in the
  454. 18:40previous slide. So we look at the
  455. 18:42regression line going across. We have an
  456. 18:45F value of 41.96
  457. 18:48and a P value of 0. So we know that's
  458. 18:51significant. Now at the bottom we have
  459. 18:54the model summary. So S is the same
  460. 18:57thing as the standard error of the
  461. 18:59regression that we had in Excel. So here
  462. 19:02it's 368091.
  463. 19:04Then again we have R 2 we have R 2
  464. 19:08adjusted
  465. 19:09and then we have in this case R 2
  466. 19:13predicted. So this is an addition that
  467. 19:16MiniAB gives us and R 2 predicted is
  468. 19:21basically how well our model does at
  469. 19:26predicting
  470. 19:27additional data points. So we'll talk
  471. 19:29about that more as we go. But R squ is
  472. 19:32really about predictive power.
  473. 19:35So here are our coefficients. So again,
  474. 19:37this looks very similar to what we had
  475. 19:39in Excel. So we'll look at the number of
  476. 19:41deliveries uh row. So we have a
  477. 19:44coefficient of 4983. We have a t value
  478. 19:47of 6.48. Its p value is 0. So we know it
  479. 19:52is also significant. The vif we'll talk
  480. 19:55about later. And then the regression
  481. 19:57equation, which is really nice about
  482. 19:59many tab and other software packages, it
  483. 20:01actually gives us the regression
  484. 20:02equation. So we have 4.845.
  485. 20:05So that comes from the constant
  486. 20:07coefficient up there at the top.
  487. 20:09Plus4983
  488. 20:11that comes from the number of deliveries
  489. 20:13coefficient up there. And then we
  490. 20:15multiply that by the number of
  491. 20:16deliveries we actually have. So an
  492. 20:19increase in one delivery, one additional
  493. 20:22number of deliveries will increase
  494. 20:24delivery time by 4983 hours or almost a
  495. 20:30half an hour. Again, that's how we
  496. 20:32interpret a simple regression
  497. 20:34coefficient.
  498. 20:37So, let's go ahead and make a rough for
  499. 20:39delivery estimate. Now, we're not going
  500. 20:40to do the prediction interval again.
  501. 20:42We'll just do the four delivery
  502. 20:43estimate. So, let's say a trip has four
  503. 20:45deliveries. So, we can go ahead and
  504. 20:47substitute the four into our regression
  505. 20:49equation and we come up with an estimate
  506. 20:52of 6.838
  507. 20:54hours or 6 hours and 50 minutes. So if
  508. 20:59we come into work one day and we have a
  509. 21:01trip with four deliveries on it. So
  510. 21:04based on our data, the best estimate we
  511. 21:06have is a trip of 6 hours and 50
  512. 21:09minutes.
  513. 21:12So finally, here's our last single
  514. 21:14regression. So gas price X3 and travel
  515. 21:17time Y. So we go across the regression
  516. 21:20line here. We have an F value of
  517. 21:2462.
  518. 21:26That's very low. Then we have a p value
  519. 21:29of 0.455.
  520. 21:31That is not significant. That's
  521. 21:33obviously not below 0.05. So that
  522. 21:36basically confirms what we looked at in
  523. 21:38the scatter plots and correlations. And
  524. 21:40that is that gas price really has no
  525. 21:43relationship or no linear relationship
  526. 21:45to travel time. Now if we look down at
  527. 21:48the model summary, we have a standard
  528. 21:50error of regression of86
  529. 21:53hours. That is a huge standard error of
  530. 21:56the regression. We have an R squar
  531. 21:58that's 7.14%.
  532. 22:00You know why even bother? The R squ
  533. 22:02adjusted is literally zero and the R
  534. 22:04square predicted is literally zero. So
  535. 22:07we know that this variable is really a
  536. 22:10dud.
  537. 22:11So here are our coefficients. We have
  538. 22:14the gas price coefficient of 81. We have
  539. 22:18a T value of 78 very low obviously P
  540. 22:22value of 0455. Same thing as we have in
  541. 22:24the f value above the p value up there.
  542. 22:27We know that's not significant. The
  543. 22:30regression equation, why even bother?
  544. 22:32This is really not worth exploring any
  545. 22:34further. We know that gas price is
  546. 22:36really not a variable that contributes
  547. 22:38to travel time.
  548. 22:41Let's go ahead and summarize these three
  549. 22:43models. So in the first line, we have
  550. 22:45the case where we're only using our
  551. 22:47first independent variable. We had an F
  552. 22:50of 49.77,
  553. 22:52a p value that's very low, significant,
  554. 22:54less than 0001. We had a standard error
  555. 22:57of the regression of 34230.
  556. 23:01We had an R squar of 84.42
  557. 23:04and an R squar predicted of 79.07.
  558. 23:08Now I did that in many tabs, so we had
  559. 23:10it um but that's the actual R square
  560. 23:13predicted that was produced. And then in
  561. 23:16the second case where we had the second
  562. 23:18independent variable we had an f of
  563. 23:2041.96
  564. 23:22p value less than 0001.
  565. 23:25Now here we had a standard error of the
  566. 23:27regression of 36809.
  567. 23:31So remember what I said. What does that
  568. 23:33tell you? Well, in the first case in the
  569. 23:36um the top row, on average, our data
  570. 23:39points were 34230
  571. 23:42hours away from the regression line. In
  572. 23:46the second case, they were 36809 hours
  573. 23:50away from the regression line. So, which
  574. 23:53model has data points that are more
  575. 23:55clustered in around the regression line?
  576. 24:00Well, the first one does because it
  577. 24:01standard error of the regression is
  578. 24:03lower. Now, the R squ adjusted for the
  579. 24:05second one is 81.99.
  580. 24:08It's a little bit lower than the one at
  581. 24:09the top. And then the R square predicted
  582. 24:12is 70.27, which is quite a bit lower
  583. 24:15than the one at the top. And then
  584. 24:17finally, we sort of have our our dud,
  585. 24:20our no good one, where we're using X3.
  586. 24:22I'm not even going to go through that
  587. 24:24anymore because it's obviously um no
  588. 24:26good. So, if we look at this example
  589. 24:28where we're only using one independent
  590. 24:31variable, which one do you think is the
  591. 24:34best?
  592. 24:35Well, I would have to say that the first
  593. 24:36one's the best. We have the highest
  594. 24:39F-stistic. We have the lowest standard
  595. 24:42error of the regression. We have the
  596. 24:44highest R 2 adjusted. And we definitely
  597. 24:46have the highest R square predicted. So
  598. 24:49of these three, the first one is
  599. 24:52definitely I would say the best.
  600. 24:57So let's do the case where we have two
  601. 24:59variables. So now we're going to put two
  602. 25:01variables in at a time. So we'll put X1
  603. 25:04and X2 into the regression. We'll put X1
  604. 25:06and X3 into the regression. And we'll
  605. 25:08put X2 and X3 into the regression. So
  606. 25:11this is a subset of two variables but
  607. 25:13the interpretation is basically the

About this transcript

This page contains the full transcript of YouTube transcript (wPJ1_Z8b0wk) , generated from the public captions YouTube serves with the video. The transcript has 4,041 words across 607 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.