YouTube2Text

YouTube transcript (kHZBy1uVNnM) — Transcript

3,824 words · 540 segments · language en · Watch on YouTube

Full transcript

  1. 0:02[Music]
  2. 0:17hello thanks for watching and welcome to
  3. 0:20the next video in my series on basic
  4. 0:22statistics now as usual a few things
  5. 0:25before we get started number one if
  6. 0:27you're watching this video because you
  7. 0:28are struggling in a class right now I
  8. 0:31want you to stay positive and keep your
  9. 0:33head up if you're watching this it means
  10. 0:35you've accomplished quite a bit already
  11. 0:37you're very smart and talented but you
  12. 0:39may have just hit a temporary rough
  13. 0:41patch now I know with the right amount
  14. 0:43of hard work practice and patience you
  15. 0:46can work through it I have faith in you
  16. 0:49many other people around you have faith
  17. 0:51in you so so should you number two
  18. 0:55please feel free to follow me here on
  19. 0:57YouTube on Twitter on Google+ or on
  20. 1:01LinkedIn that way when I upload a new
  21. 1:03video you know about it and it's always
  22. 1:06nice to connect with my viewers online
  23. 1:08I'll feel that life is much too short
  24. 1:10and the world is much too large for us
  25. 1:12to miss the chance to connect when we
  26. 1:14can number three if you like the video
  27. 1:17please give it a thumbs up share it with
  28. 1:20classmates or colleagues or put it on a
  29. 1:22playlist that does encourage me to keep
  30. 1:24making them for you on the flip side if
  31. 1:27you think there's something I can do
  32. 1:28better please leave a instructive
  33. 1:30comment below the video and I will take
  34. 1:32those ideas into account when I make new
  35. 1:35ones and finally just keep in mind that
  36. 1:37these videos are meant for individuals
  37. 1:39who are relatively new to Stats so I'm
  38. 1:42just going over basic concepts and I
  39. 1:45will be doing so in a slow deliberate
  40. 1:48manner not only do I want you to know
  41. 1:51what is going on but also why and how to
  42. 1:55apply it so all that being said let's go
  43. 1:58ahead and get started
  44. 2:02so this video is the next in our series
  45. 2:03about simple linear regression in our
  46. 2:06last three videos we talked about the
  47. 2:08very basics of regression and introduced
  48. 2:10other fundamental concepts like the
  49. 2:12algebra of lines General patterns to
  50. 2:14look for on Scatter Plots and the leas
  51. 2:16squares method in this video we're going
  52. 2:19to learn to evaluate how well a
  53. 2:21regression line fits the data it models
  54. 2:25it is important to note that regression
  55. 2:27model is unique to the data it
  56. 2:29represents
  57. 2:30adding or changing data points will most
  58. 2:32certainly change the regression model
  59. 2:35the model is also only valid for the
  60. 2:37range of data points under analysis it
  61. 2:40is not proper to extrapolate above or
  62. 2:43below the data being
  63. 2:45evaluated so once a regression line is
  64. 2:48calculated how much better is it than
  65. 2:51using only the mean of the dependent
  66. 2:53variable alone to answer that question
  67. 2:56we will have to calculate the sum of the
  68. 2:57squared residuals or errors much like we
  69. 3:00did in the previous video however this
  70. 3:04time we will do so using the regression
  71. 3:06line not the dependent variable mean
  72. 3:10line and finally we will be able to
  73. 3:13quantify the fit of the regression model
  74. 3:16using a simple Ratio or percentage
  75. 3:19called the coefficient of determination
  76. 3:21so if you are new to regression or are
  77. 3:23still trying to figure out exactly what
  78. 3:25it even is this video is for you so sit
  79. 3:29back relax and let's go ahead and get to
  80. 3:34learning so we can think of this video
  81. 3:36as a story we'll call it a tail of two
  82. 3:40lines now in the first few videos we
  83. 3:42worked with this graph and data over
  84. 3:44here on the left remember in this case
  85. 3:48we only had the dependent variable which
  86. 3:50is the tip amount at the restaurant so
  87. 3:53we had six meals and then we had the six
  88. 3:56tip amounts that are graphed on the Y
  89. 3:58AIS in terms of their dollar
  90. 4:00amounts now we figured out the mean of
  91. 4:03these six tips was $10 so we went ahead
  92. 4:06and put a line a horizontal line at
  93. 4:09$10 then we found the errors which is
  94. 4:12the distance from the mean line to each
  95. 4:16observed data point and then we squared
  96. 4:19all of those errors and then added them
  97. 4:22up when we did that we came up with an
  98. 4:25ssse or sum of squared errors of 120
  99. 4:30now this is a very important point with
  100. 4:33only the dependent variable the only sum
  101. 4:36of squares is due to error that's all we
  102. 4:40have to work with therefore it is also
  103. 4:44the total sum of squares and the maximum
  104. 4:48sum of squares for the data under
  105. 4:50analysis so for these six specific data
  106. 4:53points the total sum of squares will
  107. 4:56never be more than
  108. 4:58120 so since the only source of error we
  109. 5:01have is the SSE it's also the total sum
  110. 5:05of squares in this case the same thing
  111. 5:08so SST is also 120 again because in this
  112. 5:12case they're the same
  113. 5:15thing now in the previous video we
  114. 5:17actually went ahead and calculated the
  115. 5:19regression line using the least squares
  116. 5:21method so we had the same six data
  117. 5:23points the graph looks different because
  118. 5:25they're graphed differently and then of
  119. 5:26course we have our regression line now
  120. 5:29in this video we're going to do the same
  121. 5:32thing we did over here on the left we're
  122. 5:34going to find the errors and square them
  123. 5:37and then add them up but of course we
  124. 5:39will not be using the mean of the
  125. 5:41dependent variable to do that we're
  126. 5:43going to be using the regression line to
  127. 5:45do that now with both the independent
  128. 5:48variable and the dependent variable the
  129. 5:50total sum of squares Remains the Same
  130. 5:55it's 120 still and it will always be for
  131. 5:58these six points but ideally the error
  132. 6:02sum of squares or the SS will be reduced
  133. 6:07significantly because our line fits the
  134. 6:11data better so the difference between
  135. 6:14the SST of 120 and the SSE which we will
  136. 6:18calculate is the SSR or the sum of
  137. 6:22squares due to the regression we did
  138. 6:25this is very much like a Nova in a way
  139. 6:27we'll talk about that as we go forward
  140. 6:29Ward but the SST will always be 120
  141. 6:32we're going to calculate the ssse based
  142. 6:34on the residuals from the regression
  143. 6:36line and then the difference between the
  144. 6:38SST and SS e will be the SSR or the sum
  145. 6:43of squares due to regression take a step
  146. 6:46back and look this is always what we are
  147. 6:49doing in simple linear regression we are
  148. 6:52comparing a regression model over here
  149. 6:54on the right that we found using the Le
  150. 6:57squares method to the case where we only
  151. 7:00have the dependent variable and the line
  152. 7:03is flat over here on the left now if the
  153. 7:07model on the right we calculate looks a
  154. 7:09lot like the one on the left and the sum
  155. 7:12of squared error is very high then our
  156. 7:15regression model doesn't really do a
  157. 7:17whole lot for us the idea is to reduce
  158. 7:20the ssse by having a line that fits the
  159. 7:23data better because the better the line
  160. 7:26fits the data the smaller the residual
  161. 7:29uals or the errors will
  162. 7:32be so this is our first case you can go
  163. 7:35and look at this again if you'd like but
  164. 7:37we just took the residuals of the errors
  165. 7:38squared them and then added them up this
  166. 7:41was
  167. 7:42120 so having only the dependent
  168. 7:44variable the best prediction for the tip
  169. 7:46of the next meal is just the mean of the
  170. 7:48tips which in this case is
  171. 7:50$10 since the mean line is flat its
  172. 7:53slope is zero so beta sub 1 equals 0
  173. 7:57again this is just a revieww
  174. 8:00now in the last video we actually
  175. 8:02calculated the regression line so here's
  176. 8:03a regression equation at the top and
  177. 8:06here's another one remember that Excel
  178. 8:08had a little bit different number for
  179. 8:10The Intercept than we did just because
  180. 8:11of rounding but basically they're the
  181. 8:13same thing so a few things our slope is
  182. 8:16not zero our slope is
  183. 8:1901462 and remember we interpreted that
  184. 8:22as for every dollar the meal amount
  185. 8:25increases the tip is expected to
  186. 8:28increase by about 15 cents that's the
  187. 8:31slope of
  188. 8:3301462 then we also had our intercept
  189. 8:36down here on the left which really for
  190. 8:37this problem isn't very meaningful now
  191. 8:39we also had the centroid so remember the
  192. 8:42centroid is the intersection of the
  193. 8:43means of both variables so the mean meal
  194. 8:47amount was $74 and the mean tip amount
  195. 8:49was $10 and the centroid will always
  196. 8:53fall on the least squares line and that
  197. 8:56can be very helpful if you have to find
  198. 8:58out other things about your
  199. 9:01graph so now that we're done reviewing
  200. 9:04let's go ahead and get to sort of the
  201. 9:05heart of the matter of this video but to
  202. 9:07do so we have to calculate the predicted
  203. 9:10values for each meal amount so over here
  204. 9:13on the left we have the observed total
  205. 9:15bill which in like the case the first
  206. 9:17meal was $34 and then we have the
  207. 9:20observed tip so that was the tip that
  208. 9:22was actually received by the waiter so
  209. 9:25the first one was $34 the tip that the
  210. 9:27waiter got was $5 $18 meal headed a tip
  211. 9:31of $17 so on and so forth now in this
  212. 9:35middle column we have our actual
  213. 9:36regression equation so 01462 xus
  214. 9:410.81
  215. 9:4388 now that is what we're going to use
  216. 9:46to create the predicted tip amount using
  217. 9:49our regression equation for each of our
  218. 9:52meals over here in the red on the
  219. 9:57left all we do is substitute each meal
  220. 10:01amount in for X because again that's our
  221. 10:04independent variable so we go ahead and
  222. 10:07evaluate all of these equations and that
  223. 10:10will give us our predicted tip amount
  224. 10:12over here on the right so again to
  225. 10:14substitute the meal amount into the
  226. 10:16regression equation and that will give
  227. 10:18us our predicted tip amount based on the
  228. 10:24regression so substitute over evaluate
  229. 10:27and now we have our predicted
  230. 10:29tips so for this meal that was $34 the
  231. 10:33waiter received
  232. 10:35$5 now our regression equation would
  233. 10:39have predicted a tip of
  234. 10:42$415 let skip down we had a meal that
  235. 10:46was
  236. 10:46$88 the waiter or waitress received
  237. 10:50$8 Now using our regression equation we
  238. 10:53would predict a tip of $12
  239. 10:57about5 so you can see the obser tip
  240. 10:59amount is not always or usually is not
  241. 11:03the predicted tip amount that is given
  242. 11:05using the regression equation so we have
  243. 11:08a discrepancy here we have the observed
  244. 11:10tip amount and we have the predicted tip
  245. 11:13amount and the difference between those
  246. 11:15two is going to be our
  247. 11:20error so here is our data again so the
  248. 11:22Diamonds the orange diamonds represent
  249. 11:24our actual observed values for the tips
  250. 11:27that the waiter received so here's our
  251. 11:29regression equation now the purple dots
  252. 11:33are the predicted tip amounts for those
  253. 11:36meal amounts as you can see they're not
  254. 11:38the same as what we observed so for the
  255. 11:41first one we had $415 that we would
  256. 11:43predict the tip to be for the second one
  257. 11:45we had a predicted value of
  258. 11:47$664 for a tip and so on and so forth on
  259. 11:51up to the
  260. 11:54ground now we have to find out what the
  261. 11:57error is or what the residual R so again
  262. 12:01that is just the distance between the
  263. 12:03predicted value and the observed value
  264. 12:07because that's quote how far off our
  265. 12:10regression is from the actual data so
  266. 12:12it's just the distance between the
  267. 12:14predicted and The
  268. 12:16observed so now we have to find out what
  269. 12:18those errors are and as you can imagine
  270. 12:21it's simple subtraction it's just the
  271. 12:22difference between the observed and the
  272. 12:25predicted that we calculated so for our
  273. 12:27first meal $34 we have observed tip
  274. 12:29amount actually in the restaurant of $5
  275. 12:32our regression predicted a
  276. 12:35$415 tip approximately so the difference
  277. 12:38between those two was
  278. 12:410.849 or about
  279. 12:4385 and that is the error for that meal
  280. 12:47amount and then we do that for all six
  281. 12:49meals very simple just
  282. 12:53subtraction now of course as we do with
  283. 12:56other errors we have to square those
  284. 12:58differences so the difference of
  285. 13:008495 for the first meal if we square
  286. 13:04that residual it is
  287. 13:060.721 7 for the second meal the
  288. 13:09difference is
  289. 13:122.37 we go ahead and square that and
  290. 13:15it's
  291. 13:29we sum those up we have an SS e that is
  292. 13:3430.75 so for our regression model that
  293. 13:38we came up with our SSE is a little bit
  294. 13:41over 30 so
  295. 13:4530.75 so on our graph we take each
  296. 13:49residual and then we Square it just like
  297. 13:53this so we take the residual and square
  298. 13:55it so these are areas we have here and
  299. 13:59then just like we did before we go ahead
  300. 14:02and add those
  301. 14:06up now let's compare the sum of squared
  302. 14:09errors from our two models now remember
  303. 14:12in the first model where we only had the
  304. 14:14dependent variable which was the tip
  305. 14:16only we had a sum of squared errors or
  306. 14:18SSE of 120 and again remember that's
  307. 14:22also the SST for that model so
  308. 14:26120 now for our regression model we have
  309. 14:29the DV and the IV to the dependent and
  310. 14:32independent variable where the tip
  311. 14:34amount is a function of the meal amount
  312. 14:37now we have a suos squared errors or SS
  313. 14:40of
  314. 14:4430.75 see how this works that's the
  315. 14:47entire point of simple linear regression
  316. 14:51to create a model to create a line
  317. 14:54through our data that reduces the ssse
  318. 14:57as much as it possibly can
  319. 15:01so if we actually make these physically
  320. 15:04the same scale side by side we can see
  321. 15:06that we put all those sses for the first
  322. 15:09model together it's 120 and we do the
  323. 15:11same thing for our regression model and
  324. 15:13it's
  325. 15:1530.75 so you can actually see the
  326. 15:17physical difference in the scale of the
  327. 15:20errors so when we conducted the
  328. 15:23regression the ssse decreased from 120
  329. 15:27to 30.0
  330. 15:30075 that is
  331. 15:3230.75 of the sum of the squares was
  332. 15:36explained or allocated to the error in
  333. 15:40the regression model so instead of
  334. 15:42having 120 as the error as in the first
  335. 15:45case we reduced that to
  336. 15:4830.75 because we reduced the distance
  337. 15:52from the regression line to the data
  338. 15:54points where did the other 89.8 n25 go
  339. 15:59well
  340. 16:0189.95 is the sum of squares due to our
  341. 16:07regression SST equal SSR plus
  342. 16:11SS in this case SST is always 120 and
  343. 16:15that equals 89.9 25 +
  344. 16:2130.75 so here are our two models again
  345. 16:25so the first one our ssse was 120 and of
  346. 16:28course the s s and SST are the same for
  347. 16:30that model so the SST is also
  348. 16:33120 now in our second regression model
  349. 16:36that we did our SST again is still 120
  350. 16:39but now the SS is
  351. 16:4330.75 because of the reduced distance
  352. 16:46between the regression line and the
  353. 16:49observed data points as it compares to
  354. 16:51the one over here on the left so doing
  355. 16:54Simple subtraction we had 120us 30.75
  356. 16:58that gives us our SSR or sum of squares
  357. 17:01due to regression and that's
  358. 17:0389.9
  359. 17:0725 so now that we have the SST the SSR
  360. 17:11and the ssse we can go ahead and talk
  361. 17:13about the coefficient of
  362. 17:15determination so how well does the
  363. 17:18estimated regression equation fit our
  364. 17:22data now this is where regression begins
  365. 17:25to look a lot like a Nova the total sum
  366. 17:29of squares is partitioned or allocated
  367. 17:33into SS and SSR and hopefully at some
  368. 17:38point I'll do a separate video on the
  369. 17:39relationship between regression and
  370. 17:41Anova but it's all about allocating or
  371. 17:45partitioning the total sum of squares
  372. 17:48into two or more things in this case SS
  373. 17:51and
  374. 17:52SSR so the total sum of squares is split
  375. 17:56some of it goes to SSR some of it goes
  376. 17:59to
  377. 18:00SS so if the SSR is large it uses up
  378. 18:05more of the SST and therefore SS is
  379. 18:10smaller relative to the total the
  380. 18:13coefficient of
  381. 18:14determination quantifies this ratio as a
  382. 18:18percentage so the coefficient of
  383. 18:21determination has the variable R 2
  384. 18:25equals SSR / SST so it's the sum of
  385. 18:29squares regression divided by the total
  386. 18:32sum of
  387. 18:34squares and if you actually do this in a
  388. 18:37statistical software package regression
  389. 18:39will actually give you back an anova
  390. 18:42table just like you saw when you did an
  391. 18:45anova so we're going to have SS and mean
  392. 18:48squares and an F statistic and a
  393. 18:51significance level and everything just
  394. 18:53like we did in an NOA and of course in
  395. 18:56this case this is the actual Anova for
  396. 18:58the data what we're working with we can
  397. 19:00see that the f is very large and
  398. 19:03significance is
  399. 19:060.258 which is significant at the 05
  400. 19:09level so there is a relationship between
  401. 19:12linear regression and an NOA because
  402. 19:14it's all about partitioning or
  403. 19:16allocating the total sum of squares into
  404. 19:19different components and then we measure
  405. 19:22the ratio of those components to figure
  406. 19:25out whether or not the model is
  407. 19:26statistically significant same
  408. 19:30idea so how do we interpret the
  409. 19:32coefficient of determination so again
  410. 19:35it's SSR over SS in this case it's
  411. 19:3989.95 ID 120 so we go ahead and do that
  412. 19:43Division and we come up with 0.74 93 or
  413. 19:4974.9 3% cuz remember the 120 is the
  414. 19:53total so
  415. 19:558992 of the 120 that's a percent
  416. 19:59and that is
  417. 20:0074.9 3 so we can conclude that
  418. 20:0474.9 3% of the total sum of squares can
  419. 20:09be explained using the estimated
  420. 20:11regression equation to predict the tip
  421. 20:14amount the remainder so
  422. 20:1925.7% is due to error and again that's
  423. 20:23based on this regression equation we
  424. 20:25used where X is the dollar amount of the
  425. 20:29bill so it's really about partitioning
  426. 20:31the total sum of squares as a percentage
  427. 20:34saying hey this percentage belongs to
  428. 20:36SSR and this percentage belongs to
  429. 20:39ssse and then we can develop a model
  430. 20:43that says this fits well or does not fit
  431. 20:45well in this case it is a good fit so a
  432. 20:50coefficient of determination of 74.9 3%
  433. 20:54is good and we also saw that in the
  434. 20:56Innova table which was statistic Bally
  435. 20:59significant so it's all about
  436. 21:01partitioning the total sum of squares
  437. 21:03into SSR and ssse and we want SS to be
  438. 21:08as low as possible for the model to be
  439. 21:11considered a good fit and for the Innova
  440. 21:14table to produce a significant
  441. 21:19result so let's put all together on one
  442. 21:22graph so you can see here we're dealing
  443. 21:24with three data points for each measure
  444. 21:28on our xaxis or independent variable so
  445. 21:31if we look at this one here sort of to
  446. 21:33the right we can see that we have a
  447. 21:35orange diamond we have a black circle
  448. 21:39and a purple circle so three data points
  449. 21:43that represent that meal of $88 I think
  450. 21:46it was so remember that the diamond is
  451. 21:49the actual observed value the black
  452. 21:52circle is the mean of the dependent
  453. 21:56variable which was the tips which is $10
  454. 21:59the purple circle is the predicted tip
  455. 22:02amount based on the regression model so
  456. 22:05each meal kind of has three reference
  457. 22:07points The observed the mean of the
  458. 22:10dependent variable and the predicted
  459. 22:13value so the first Square difference to
  460. 22:15look at is the SS and I think it's the
  461. 22:17easiest to remember so it is the sum of
  462. 22:20the squared differences between y sub I
  463. 22:23which is the observed value the diamond
  464. 22:26minus y hat sub I which is the predicted
  465. 22:30value which is the purple circle so we
  466. 22:32take the difference of each one along
  467. 22:34the graph we Square them and then we add
  468. 22:37them all up and that is the SS for our
  469. 22:40regression model so it's the distance
  470. 22:43between those two and similarly those
  471. 22:46two up here so the difference between
  472. 22:49the observed the diamond and the
  473. 22:50predicted the purple circle now the next
  474. 22:54one is SST it's the sum Square total now
  475. 22:58it is the sum of the square difference
  476. 22:59between each observed value which is the
  477. 23:03diamond minus the mean of the Y AIS or
  478. 23:07the dependent variable which is y bar so
  479. 23:11the difference between those then we
  480. 23:13Square it then we add them all up so
  481. 23:16that's the difference between the
  482. 23:17observed and the mean so the orange
  483. 23:19diamond and the black circle so just
  484. 23:22like that and just like that there and
  485. 23:25of course the final one is SSR
  486. 23:29so it is the sum of the squar difference
  487. 23:32between y hat sub I that's the predicted
  488. 23:35whe is the purple circle minus y bar
  489. 23:39which is the mean of the dependent
  490. 23:42variable so we find the difference we
  491. 23:44Square them and then add them up and
  492. 23:46that is this distance here and this
  493. 23:49distance here of course we would do this
  494. 23:51for each point along the graph so it's
  495. 23:55always a relationship between these
  496. 23:58three points that represent each value
  497. 24:02of the independent variable so in the
  498. 24:04middle here in this case a $88 meal
  499. 24:07amount we have the observed we have the
  500. 24:10mean of the Y's and we have the
  501. 24:12predicted and we have three
  502. 24:13relationships going on between those
  503. 24:15three data points and the square
  504. 24:17differences among those three data
  505. 24:19points is actually the total sum of
  506. 24:21squares package of SST ssse and SSR so
  507. 24:26now you can see what it actually looks
  508. 24:27like on
  509. 24:31graph okay so that wraps up our video on
  510. 24:34simple linear regression fit and the
  511. 24:37coefficient of determination so remember
  512. 24:40this is all about a comparison A Tail of
  513. 24:43Two Lines where we had the one graph
  514. 24:46where we only used the mean of the
  515. 24:48dependent variable and then we had our
  516. 24:50regression equation and that graph over
  517. 24:53on the right and it's always a
  518. 24:55comparison between the two if the
  519. 24:58regression model isn't any better than
  520. 25:01the mean of the dependent variable then
  521. 25:03the model doesn't really give us
  522. 25:04anything new and will most likely come
  523. 25:07back is statistically insignificant the
  524. 25:09SS will be very high and then we will
  525. 25:12just move on from there but it is always
  526. 25:15a comparison between these two different
  527. 25:18models now the coefficient of
  528. 25:20determination just tells us how we
  529. 25:23allocate the total sum of squares in the
  530. 25:25regression model so it's split it's
  531. 25:27either SS s e or it's SSR and we hope
  532. 25:31for a good fitting model that the SS is
  533. 25:34very low and the SSR is very high but in
  534. 25:38either case their sum will always add up
  535. 25:41to SST whatever that happens to be so I
  536. 25:44look forward to seeing you in that next
  537. 25:46video I wish you all the best of luck in
  538. 25:48your studies and in your work and I look
  539. 25:50forward to seeing you again next time
  540. 25:55[Music]

About this transcript

This page contains the full transcript of YouTube transcript (kHZBy1uVNnM) , generated from the public captions YouTube serves with the video. The transcript has 3,824 words across 540 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.