YouTube2Text

YouTube transcript (px72eCYPuvc) — Transcript

3,703 words · 555 segments · language en · Watch on YouTube

Full transcript

  1. 0:00So let's do the case where we have two
  2. 0:02variables. So now we're going to put two
  3. 0:04variables in at a time. So we'll put x1
  4. 0:07and x2 into the regression. We'll put x1
  5. 0:09and x3 into the regression. And we'll
  6. 0:11put x2 and x3 into the regression. So
  7. 0:14this is a subset of two variables, but
  8. 0:16the interpretation is basically the
  9. 0:20same. So this first one we have travel
  10. 0:23time, our dependent variable versus
  11. 0:25miles traveled together with number of
  12. 0:28deliveries. So our x1 and
  13. 0:31x2. Let's look at our regression here.
  14. 0:34We have an f value of
  15. 0:3823.72 with a p value of
  16. 0:410001. Now for now I want you to ignore
  17. 0:44the f value and p value for the
  18. 0:48individual variables. We'll get to that
  19. 0:50in a later video. So just look at the
  20. 0:52regression line across the top here. So
  21. 0:5523.72 F value P value of 0001. Of course
  22. 0:59that is significant. Now let's look down
  23. 1:02below. We have a standard error of the
  24. 1:05regression of
  25. 1:08352642. So that's again that's in the
  26. 1:10ballpark of what we had in the one
  27. 1:13variable models. We have an R squar of
  28. 1:1687.14, an adjusted R squar of 83.47%
  29. 1:2047% and then an R squar predicted of
  30. 1:2659.95%. So that's significantly lower
  31. 1:28and we'll talk about what that means
  32. 1:30later in the video. So now let's look at
  33. 1:33the coefficients and this is where
  34. 1:35things get interesting. So let's look at
  35. 1:38the miles traveled coefficient. So its
  36. 1:42value is
  37. 1:430262. Its t value is
  38. 1:471.31 and its p value is
  39. 1:52232. That is not significant. It's not
  40. 1:56below 05. Now look at number of
  41. 1:59deliveries. Coefficient
  42. 2:01of84, t value of 73 and a p value of
  43. 2:08487. That is not significant either.
  44. 2:12So, here's the weird thing. We have an
  45. 2:15overall model that is significant. Okay,
  46. 2:19up the top we have an F value of
  47. 2:2223.72 and we have a p value of
  48. 2:250.001. But down below, neither of our
  49. 2:29coefficients are statistically
  50. 2:31significant. Very strange. And then of
  51. 2:34course down below we have our regression
  52. 2:37equation. So remember what we discussed
  53. 2:39before, these two variables X1 and X2
  54. 2:43are extremely correlated with each
  55. 2:46other. So they they're almost on a
  56. 2:48straight line that goes from bottom left
  57. 2:50to top right. The correlation is N56.
  58. 2:54That is very significant. So these two
  59. 2:57variables are
  60. 2:59multicolinear. Now look what happens. We
  61. 3:02have two variables that are collinear.
  62. 3:05The overall model is significant, but
  63. 3:08the individual coefficients are not. Now
  64. 3:11you see what happens when we put two
  65. 3:14variables in a regression that are
  66. 3:17correlated, that have high levels of
  67. 3:20colinearity. The coefficients in the
  68. 3:23regression model go haywire. They go
  69. 3:26crazy. So this is why we look to
  70. 3:27eliminate variables from the get-go.
  71. 3:30Now, Mini Tab also has this variance
  72. 3:33inflation factor or the VIF. Now, how
  73. 3:36this helps us is that it points out
  74. 3:39variables that are collinear that points
  75. 3:43out
  76. 3:44multi-olinearity. So, we'll talk about
  77. 3:45it more as we go, but a VIF of
  78. 3:5011.59 should send alarm bells off um in
  79. 3:53our statistical minds. We should know
  80. 3:56that is a serious problem. And this
  81. 3:59model we have in front of us is very
  82. 4:01suspect. And we can base that on several
  83. 4:04criteria. We know from our scatter plots
  84. 4:07that they're highly correlated, very
  85. 4:09highly correlated. We have this weird
  86. 4:11situation where the overall model is
  87. 4:13significant, but the individual
  88. 4:15coefficients are not. And then we have
  89. 4:18this VIF that's through the roof. So all
  90. 4:21that taken together should tell us that
  91. 4:24this model has some problems. Now, we'll
  92. 4:26keep it in here just for learning
  93. 4:28purposes, but just keep in mind the
  94. 4:30criteria here. And I also want to point
  95. 4:33out that when we have an R squar
  96. 4:36adjusted, that's
  97. 4:3983.47%. Then the R squar predicted falls
  98. 4:42off a cliff. Now it's at
  99. 4:4559.95%. That tells us that we have a
  100. 4:49serious problem in our model. So all
  101. 4:52that taken together lets us know that
  102. 4:53something is wrong.
  103. 5:00So in our ANOVA table we have a
  104. 5:02regression line here. We have an F value
  105. 5:04of
  106. 5:0522.63. That is a p value of 0.001. So
  107. 5:09that is significant. Down here in the
  108. 5:11model summary, we have a standard error
  109. 5:13of the regression
  110. 5:15of.359883 hours. Of course, we have an R
  111. 5:18squared of 86.61% 61% R squared adjusted
  112. 5:22of
  113. 5:2382.78% and an R squar predicted of
  114. 5:2868.11%. Now notice couple things here.
  115. 5:31Of course, our ANOVA table overall is
  116. 5:34significant. Our R squar is adjusted as
  117. 5:37very high and our R 2 predicted doesn't
  118. 5:41fall off a cliff like we saw in the
  119. 5:44previous example with X1 and X2.
  120. 5:47So here are our coefficients. So in this
  121. 5:50case we have miles traveled which is
  122. 5:530.04137. That's our coefficient for
  123. 5:55miles traveled. T value of 6.44 with a p
  124. 5:58value 0. So that's fine. Now look at gas
  125. 6:03price. We have a negative coefficient
  126. 6:08-2.19. Now think about what this means.
  127. 6:11So what this is saying is that if we
  128. 6:14hold miles traveled constant, that's our
  129. 6:17x1. If we hold that constant, and we
  130. 6:20increase the price of gas a dollar, then
  131. 6:25the travel time will decrease by 219
  132. 6:29hours. So gas price goes up and the
  133. 6:33travel time goes down. Does this make
  134. 6:36any real sense to me? It does not. So, I
  135. 6:40don't know about you, when gas prices go
  136. 6:41up, I drive slower. But this is saying
  137. 6:44the travel time goes down. And what this
  138. 6:48points out is that we are putting this
  139. 6:50gas price variable in the regression,
  140. 6:52but it doesn't have any real
  141. 6:54relationship to the dependent variable.
  142. 6:57Remember, so now we get some very weird
  143. 7:00coefficients down here at the bottom. So
  144. 7:02that's why we try to eliminate variables
  145. 7:04up front because it really messes up the
  146. 7:07coefficients that come out of the
  147. 7:08regression process. So let's interpret
  148. 7:10both of these coefficients. First, miles
  149. 7:12traveled. So if gas price is held
  150. 7:15constant, then travel time is expected
  151. 7:17to increase by
  152. 7:2104137 hours for each additional mile
  153. 7:24traveled. Now does that make sense in
  154. 7:27real life? Well, yes it does. So if I
  155. 7:29travel further more miles, I expect my
  156. 7:32travel time to go up. Now how about gas
  157. 7:35price? If miles traveled is held
  158. 7:37constant, then travel time is expected
  159. 7:40to decrease by 219 hours for each
  160. 7:45additional dollar increase in gas price.
  161. 7:48And again, that really does not make any
  162. 7:50sense in real life and in statistics
  163. 7:54either because this coefficient is just
  164. 7:56kind of weird. So again, we're going to
  165. 7:58keep this one kind of off to the side,
  166. 8:00noting we have some weird coefficients
  167. 8:02in the gas price, and that's probably
  168. 8:05because we included a variable in the
  169. 8:08regression that has no relationship to
  170. 8:10the dependent variable to begin with. So
  171. 8:12let's go ahead and do the last two
  172. 8:15pair. So this is X2 and X3. So numbum
  173. 8:18deliveries and gas price. So look at the
  174. 8:21ANOVA table, the regression line. We
  175. 8:23have an F value of 27.63 63. P value is
  176. 8:280. So we know that's significant. We
  177. 8:31have a standard error of regression down
  178. 8:33here at the bottom of
  179. 8:36329703 hours. R squared of
  180. 8:4088.76%, R squared adjusted of
  181. 8:4385.55%. And an R squar predicted of
  182. 8:4871.76%. So those are all pretty high.
  183. 8:50The R squared predicted does not go off
  184. 8:52a cliff like we saw in the first
  185. 8:54example. So, so far things don't seem
  186. 8:56too crazy. Now, look at our
  187. 8:59coefficients. Uh-oh, we have the same
  188. 9:02problem again. So, number of deliveries,
  189. 9:04that looks fine.
  190. 9:05So,.5665, T value of 7.13, P value 0.
  191. 9:11Fine, looks good. Gas price has gone
  192. 9:13negative again.
  193. 9:16So,.765, T value
  194. 9:19of.172. The P value is not significant
  195. 9:22in this case. So we have this weird
  196. 9:25situation again where the gas price
  197. 9:28coefficient went
  198. 9:30negative. So interpret these again. If
  199. 9:33gas prices is held constant then travel
  200. 9:35time is expected to increase
  201. 9:38by.5665 hours for each additional
  202. 9:41delivery. Now does that make sense in
  203. 9:43real life? Well yes. I expect the travel
  204. 9:47time to go up for each additional
  205. 9:49delivery I have to make. Now how about
  206. 9:52the gas price problem? If number of
  207. 9:55deliveries is held constant, then travel
  208. 9:57time is expected to decrease by 765
  209. 10:01hours for each additional dollar
  210. 10:03increase in gas price. That doesn't make
  211. 10:06any sense. So again, we have this
  212. 10:08problem where we included a variable in
  213. 10:11the model that has no natural relation
  214. 10:14to the dependent variable whatsoever. It
  215. 10:16messes up our coefficients and really
  216. 10:18this model is no good. So let's go ahead
  217. 10:21and summarize these three
  218. 10:25models. So the top three lines are the
  219. 10:28first three models we did. So that's our
  220. 10:30single variable models. Now let's look
  221. 10:33at the second three. So in our first two
  222. 10:36variable model, we had x1 and
  223. 10:38x2. So we had an f of 23.72.
  224. 10:43Now we expect the fs to be about the
  225. 10:46same for each one variable model and
  226. 10:49each two variable model etc. Okay. So
  227. 10:5223.72 we have a p value 01. That's fine.
  228. 10:56Now we have a standard error of
  229. 10:58regression of
  230. 11:0235264. So remember what that tells us
  231. 11:04that tells us how tied in our data
  232. 11:09points are to the regression line. So in
  233. 11:11this case they are on
  234. 11:14average.35264 hours away from the
  235. 11:17regression line and you can compare that
  236. 11:18to the ones we have above. So our R
  237. 11:22squar is
  238. 11:2383.47. Our R square predicted is
  239. 11:2659.95. That's a huge drop off from the R
  240. 11:29squ adjusted. Then we have this VIF over
  241. 11:33here of
  242. 11:3511.59. That is huge and that is a
  243. 11:37problem. And that's because x1 and x2
  244. 11:41are collinear. That's the problem we
  245. 11:43have there. Now the second one from the
  246. 11:45bottom that's x1 and x3. So we have
  247. 11:4822.63 for the f0001 for the p value. The
  248. 11:52standard error of the regression 35988
  249. 11:56hours. And then we have R squ adjusted
  250. 11:58at
  251. 11:5982.78%. R square predicted
  252. 12:0268.11. Everything looks pretty much okay
  253. 12:05there. Then we have a VIF of 1.14.
  254. 12:08That's not a problem. But remember from
  255. 12:10our
  256. 12:11coefficients, we had a negative X3
  257. 12:14coefficient, which is very weird. So
  258. 12:17even though everything in that row looks
  259. 12:20okay, we know that we have a coefficient
  260. 12:22oddity. So we have to keep that in mind.
  261. 12:24And then finally here we have the X2X3
  262. 12:28model. So 27.63.
  263. 12:31Then we have the p value less than 0001.
  264. 12:34Standard error of the regression of
  265. 12:3732970. Again, that's in the ballpark of
  266. 12:39everything else. But if you look above
  267. 12:42it, you can see that so
  268. 12:44far that is the best fit around the
  269. 12:48regression line. So on average
  270. 12:5232970 hours away from the regression
  271. 12:55line, R squared adjusted of 85.55%.
  272. 13:00Now look at that column. That's the
  273. 13:03highest adjusted R squar we've had. Now
  274. 13:06go over to the R square predicted.
  275. 13:08That's 71.76. There's nothing really
  276. 13:10spectacular there relative to everything
  277. 13:12else. And then of course a VIF of 1.33.
  278. 13:16No problem there. So we have to decide
  279. 13:19here. We have this last one with a
  280. 13:23higher F than the two above it. We have
  281. 13:26a smaller standard error of the
  282. 13:28regression, which is what we'd like to
  283. 13:30see. We have a relatively high R squar
  284. 13:32adjusted at
  285. 13:3485.55%. In fact, it's the highest in
  286. 13:36that column there. And the R square
  287. 13:38predicted is what we'd expect. But
  288. 13:42remember from the coefficients, this is
  289. 13:45another example of where we have a
  290. 13:48negative coefficient. We have a negative
  291. 13:52gas price coefficient. So even though
  292. 13:54everything looks okay here, we also have
  293. 13:57to keep in mind our coefficients from
  294. 13:59the previous step. So that might be a
  295. 14:01problem. So let's go ahead and define
  296. 14:03what VIF actually is. And I just quoted
  297. 14:07this from many tabs blog. The URL is
  298. 14:10down here at the bottom. Now let's go
  299. 14:11ahead and quickly read what it says. So
  300. 14:14one way to measure multiolinearity is
  301. 14:16the variance inflation factor or the VIF
  302. 14:20which assesses how much the variance of
  303. 14:23an estimated regression coefficient
  304. 14:25increases if your predictors your
  305. 14:28independent variables are
  306. 14:30correlated. If no factors are correlated
  307. 14:33if no independent variables are
  308. 14:35correlated the VIFs will all be one. Now
  309. 14:40let me pause there. Look at the VIFs for
  310. 14:43the first three models. They're all
  311. 14:44exactly one. Well, why is that? Well,
  312. 14:48there's only one independent variable in
  313. 14:49them. So, they're going to be one. There
  314. 14:52is no correlation there. So, they'll all
  315. 14:54be one. Now, a VIF between five and 10
  316. 14:58indicates high correlation. That may be
  317. 15:01problematic. So, do we have any between
  318. 15:03five and 10? Uh, nope. Not so far. Now
  319. 15:07if the VIF goes above 10, you can assume
  320. 15:11that the regression coefficients are
  321. 15:13poorly estimated due to
  322. 15:17multicolinearity. So look at the first
  323. 15:19two variable model. We have a VIF of
  324. 15:2311.59. Now remember why that is. Our two
  325. 15:27independent variables X1 and X2 had a
  326. 15:31correlation above N5. They were
  327. 15:34extremely highly correlated. So that VIF
  328. 15:38the variance inflation factor points out
  329. 15:40that hey you have a problem there you
  330. 15:43have some multiolinearity some severe
  331. 15:46multiolinearity in that model and
  332. 15:48therefore we would just ax that model
  333. 15:51out of the we would just forget it so
  334. 15:53we'll leave it there for now but just
  335. 15:55know that the vif helps us find
  336. 16:00multiolinearity okay and finally the
  337. 16:03full model we're going to throw in all
  338. 16:05three independent depent variables and
  339. 16:07see what
  340. 16:09happens. Okay, so here is the ANOVA
  341. 16:12table from Mini Tab for all three
  342. 16:14independent variables. So let's look at
  343. 16:16the regression lineup here. We have an F
  344. 16:19value of
  345. 16:2116.99 with a p value of 02. So the
  346. 16:26overall model is significant. So the
  347. 16:28model summary, we have a standard error
  348. 16:30of the regression of
  349. 16:33344694 hours.
  350. 16:35We have an R squared of
  351. 16:3789.47, an R squared adjusted of
  352. 16:4184.2% and an R 2 predicted of
  353. 16:4657.49. Now, what's the red flag there?
  354. 16:49The R squared adjusted was
  355. 16:5184.20. The R square predicted is
  356. 16:5757.49. That's a huge drop
  357. 17:01off. So, the coefficients real quickly.
  358. 17:04So miles traveled had a coefficient of
  359. 17:080141. Numb deliveries was 383. Then we
  360. 17:11have the strange gas price
  361. 17:14that's.607 again. Now if we look at our
  362. 17:17p values, it gets even more strange. So
  363. 17:20the p value for miles traveled is 548.
  364. 17:24That is not significant. The p value for
  365. 17:26numbum deliveries
  366. 17:27is.249. Not significant. Guest price
  367. 17:31293. Not significant. even though it
  368. 17:34really doesn't matter because that's a
  369. 17:36junk variable at this point. Now, if you
  370. 17:38look at the VIFs, look at
  371. 17:42those. For miles traveled, it's
  372. 17:4614.94. For number deliveries, it's
  373. 17:5017.35. So, what does that tell us? We
  374. 17:53have severe severe problems with
  375. 17:57multiolinearity in this model. Severe
  376. 18:00problems. basically terminal death
  377. 18:03problems. But we go ahead and have the
  378. 18:05regression equation down here at the
  379. 18:06bottom just for kicks I guess. But
  380. 18:09basically this model is
  381. 18:14junk. So here are all of our models put
  382. 18:17together. So we're getting to the grand
  383. 18:19finale finally. So at the bottom we have
  384. 18:22this new model with an f of 16.99 p
  385. 18:25value
  386. 18:2602. Uh the standard error of the
  387. 18:28regression
  388. 18:3134469 R squared adjusted
  389. 18:3484.2%. R square predicted
  390. 18:3857.49%. Then we have our VIFs. I put
  391. 18:42those below each variable because I ran
  392. 18:44out of room. So for X1 it was 14.94, X2
  393. 18:4817.35, X3
  394. 18:511.71. So we can see that that's a
  395. 18:54problem. So step back and look at this
  396. 18:56last one again. We can see that we have
  397. 18:57a huge drop off from the R squared
  398. 19:00adjusted to the R square predicted just
  399. 19:02like we do at the top of the two
  400. 19:04variable models where we went from 8347
  401. 19:07to
  402. 19:0859.95. Now here is the question. Which
  403. 19:12model is the
  404. 19:14best? So let's start with knocking out
  405. 19:17some models. Well, we know the last
  406. 19:20model with all three variables is junk.
  407. 19:24The VIFs are sky-high. The R square
  408. 19:28predicted is way lower than the R
  409. 19:30squared adjusted. So that model is no
  410. 19:33good. Now we can rule out also the top
  411. 19:36of the two variables. So again there we
  412. 19:39have a VIF of
  413. 19:4111.59 and the R square predicted falls
  414. 19:44off a cliff from
  415. 19:4583.47 for the R squ adjusted. That one's
  416. 19:49gone. So we can rule out those two right
  417. 19:53off the
  418. 19:54bat. Now how do we decide? So here is
  419. 19:58sort of the golden rule of choosing your
  420. 20:00multiple regression model. We want to
  421. 20:03look at several factors. We want to look
  422. 20:05at the R squar adjusted. We want the
  423. 20:09highest one we can get. We want the R
  424. 20:11square predicted to be as high as we can
  425. 20:14get and to be close to the R squ
  426. 20:17adjusted which is already high. We'd
  427. 20:20like to see a relatively small standard
  428. 20:22error of the regression. So that's the S
  429. 20:24column over here on the left. And
  430. 20:26finally, all else being equal, we want
  431. 20:29the simplest model there is. So if we
  432. 20:34look at some other candidates down here,
  433. 20:36we can see that for the two variable
  434. 20:38model, we have X1, X3, X2, X3. But those
  435. 20:44have some serious problems. Remember
  436. 20:46that the X3 or the gas price coefficient
  437. 20:49was negative and that's because the gas
  438. 20:52price
  439. 20:53coefficient doesn't contribute to the
  440. 20:55dependent variable at all. Plus, we have
  441. 20:58some sharp falloff in the R square
  442. 21:00predicted. So, we're going to rule those
  443. 21:02out. So, we have ruled out all the two
  444. 21:04variable options and the three variable
  445. 21:08option. So, basically we're at the top.
  446. 21:11We can definitely eliminate the bottom
  447. 21:14one variable with just x3 in it. We know
  448. 21:17x3 is basically a junk variable. So we
  449. 21:19can x that out. Now we have to decide
  450. 21:24between the top two. That's all we have
  451. 21:26left. So is it going to be the top model
  452. 21:29with x1 or the next one with just
  453. 21:32x2? Well, I think it's pretty obvious
  454. 21:35that the top model with just
  455. 21:38x1 is the best model.
  456. 21:41So a one variable model is the best
  457. 21:46model out of all these options. So we
  458. 21:50have a very narrow standard error of the
  459. 21:52regression at 34 and some change a very
  460. 21:55high R squared a very high R square
  461. 21:58predicted no multiolinearity problems
  462. 22:00because well there's only one variable.
  463. 22:02So guess what that is our best
  464. 22:08model. So yes, and you are going to kill
  465. 22:11me, but Mini Tab and I'm sure other
  466. 22:13stats packages can do all of this,
  467. 22:15everything we just did by hand looking
  468. 22:18at the relationships, it can do it in a
  469. 22:20few clicks. And here it is. This is the
  470. 22:22output from Mini Tab. It's basically a
  471. 22:25best subsets regression which we just
  472. 22:27did step by step. So how do we read this
  473. 22:29thing? Well, look at the R squared
  474. 22:32adjusted. Which are the highest values?
  475. 22:35So we have 844. We look we have
  476. 22:39855 and we have an 842 in there. So the
  477. 22:42855 is the highest. Then we have the 844
  478. 22:45at the top. Now look at the R squar
  479. 22:48predicted which are the highest values.
  480. 22:51So we have
  481. 22:5279.1 at the top and then from there they
  482. 22:56go down pretty quickly. So nothing
  483. 22:59really worth mentioning. The 79.1 is
  484. 23:02definitely the highest R square
  485. 23:03predicted.
  486. 23:05Now examine the difference between the R
  487. 23:07squared adjusted and the R square
  488. 23:09predicted. A large drop off from the
  489. 23:13adjusted to the predicted indicates
  490. 23:16overfitting. That indicates there are
  491. 23:19too many variables in the model. So we
  492. 23:22can see that for the ones there at the
  493. 23:24bottom like the three variable model we
  494. 23:27go from 84.2 adjusted to
  495. 23:3157.5 predicted. That is a sign of
  496. 23:35overfitting and it's a bad thing. Now
  497. 23:38look at the top. We go from 844 to
  498. 23:4179.1. That's very close. And the ones
  499. 23:44below it aren't too bad either. But you
  500. 23:46can see as we get down with more
  501. 23:49variables that the drop off is very
  502. 23:51high. So those models are
  503. 23:55overfitted. Now look at Maloc.
  504. 23:58Look for the one that is low and
  505. 24:00approximately equals the number of
  506. 24:03predictor variables or independent
  507. 24:05variables plus the constant which is
  508. 24:08one. There's one constant. So in the
  509. 24:10first example we have one independent
  510. 24:12variable or one predictor variable plus
  511. 24:14the constant. So that's two. So for the
  512. 24:17single variable models we're looking for
  513. 24:19a malo's number that is two. For two
  514. 24:23variables we look for three. That's 2 +
  515. 24:261. And then for the three variable, we
  516. 24:28look at four. So overall, we're looking
  517. 24:30for the lowest one that's closest to its
  518. 24:32magic
  519. 24:33number. Then using all of the
  520. 24:36information we have above, choose the
  521. 24:39best model. And based on that info,
  522. 24:42which is the best model? The first one.
  523. 24:45So the single variable X1 model that has
  524. 24:49the very high R squared adjusted the
  525. 24:51very high R square predicted that
  526. 24:53doesn't fall off the malo CP that is
  527. 24:56almost exactly two which is what we want
  528. 24:58and a relatively narrow standard error
  529. 25:00of the regression. So at 342 it's kind
  530. 25:04of in the middle of the pack but it's
  531. 25:06fine. So overall that is the best model
  532. 25:10and that is the one we would use believe
  533. 25:13it or not to make our
  534. 25:15predictions. Okay. So that was a tour
  535. 25:17day force of how to evaluate multiple
  536. 25:21regression models. Yes, that was long. I
  537. 25:25admit that. But you'll come out of it
  538. 25:28never having to really doubt or question
  539. 25:30your knowledge of how multiple
  540. 25:32regression works, how the best models
  541. 25:34are built, at least in the linear cases.
  542. 25:37So when you go to take your test, write
  543. 25:39your paper, write a report at work,
  544. 25:41you'll be able to create the best
  545. 25:43models, make the best predictions on
  546. 25:45those models, and substantiate any
  547. 25:47findings or suggestions you make,
  548. 25:49whether it's in a report or on a paper
  549. 25:52or whatever else it might be. So, I know
  550. 25:54that was long, so I'll let you go. Thank
  551. 25:56you very much for watching. Please
  552. 25:57subscribe. If you like the video, give
  553. 25:59it a thumbs up. And I look forward to
  554. 26:01seeing you again next time.
  555. 26:05[Music]

About this transcript

This page contains the full transcript of YouTube transcript (px72eCYPuvc) , generated from the public captions YouTube serves with the video. The transcript has 3,703 words across 555 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.