YouTube2Text

YouTube transcript (0Vj2V2qRU10) — Transcript

3,495 words · 495 segments · language en · Watch on YouTube

Full transcript

  1. 0:02[Music]
  2. 0:17hello thanks for watching and welcome to
  3. 0:20the next video in my series on basic
  4. 0:22statistics now as usual a few things
  5. 0:24before we get started number one if
  6. 0:27you're watching this video because you
  7. 0:28are struggling in a class right now I
  8. 0:31want you to stay positive and keep your
  9. 0:33head up if you're watching this it means
  10. 0:35you've accomplished quite a bit already
  11. 0:37you're very smart and talented but you
  12. 0:39may have just hit a temporary rough
  13. 0:41patch now I know with the right amount
  14. 0:43of hard work practice and patience you
  15. 0:46can work through it I have faith in you
  16. 0:49many other people around you have faith
  17. 0:51in you so so should you number two
  18. 0:55please feel free to follow me here on
  19. 0:57YouTube on Twitter on Google+ or on
  20. 1:01LinkedIn that way when I upload a new
  21. 1:03video you know about it and it's always
  22. 1:05nice to connect with my viewers online I
  23. 1:08feel that life is much too short and the
  24. 1:10world is much too large for us to miss
  25. 1:12the chance to connect when we can number
  26. 1:16three if you like the video please give
  27. 1:18it a thumbs up share it with classmates
  28. 1:21or colleagues or put it on a playlist
  29. 1:23that does encourage me to keep making
  30. 1:25them for you on the flip side if you
  31. 1:27think there's something I can do better
  32. 1:29please leave a instructive comment below
  33. 1:30the video and I will take those ideas
  34. 1:33into account when I make new ones and
  35. 1:36finally just keep in mind that these
  36. 1:37videos are meant for individuals who are
  37. 1:40relatively new to Stats so I'm just
  38. 1:42going over basic concepts and I will be
  39. 1:45doing so in a slow deliberate manner not
  40. 1:49only do I want you to know what is going
  41. 1:51on but also why and how to apply it so
  42. 1:56all that being said let's go ahead and
  43. 1:58get started
  44. 2:01so this is the first video about a new
  45. 2:04topic the analysis of variants or more
  46. 2:07commonly known as Anova in our last set
  47. 2:10of videos we learned how to compare the
  48. 2:13variances of two populations using the F
  49. 2:16ratio and F distribution in the context
  50. 2:20of Anova it is important to remember
  51. 2:23that the F ratio is simply a ratio of
  52. 2:26two variances F ratios are a central
  53. 2:30part of Anova so if you are unsure what
  54. 2:33the f ratio is you may want to go back
  55. 2:36and watch those
  56. 2:37videos Anova allows us to move Beyond
  57. 2:41comparing just two populations with
  58. 2:44Anova we can compare multiple
  59. 2:47populations and even subgroups of those
  60. 2:50populations we can investigate how two
  61. 2:53groups interact with each other
  62. 2:56quantitatively many experimental
  63. 2:58research designs use an NOA for these
  64. 3:01very reasons now in this video we will
  65. 3:04not be doing any calculations looking at
  66. 3:07any formulas or testing any hypothesis
  67. 3:11many students I've worked with over the
  68. 3:12years learn a Nova in class without
  69. 3:16actually knowing what it is or why it
  70. 3:19even exists so this video offers a solid
  71. 3:23conceptual Foundation about an NOA using
  72. 3:26illustrations and Graphics so if you are
  73. 3:29new to Anova
  74. 3:30or are still trying to figure out
  75. 3:32exactly what it is this video is for you
  76. 3:36so sit back relax and let's go ahead and
  77. 3:39get to
  78. 3:42work so the first and most obvious
  79. 3:45question is well why an Nova so up to
  80. 3:49this point we have been comparing two
  81. 3:52populations so the independent samples T
  82. 3:54Test which are two random samples We
  83. 3:57compare or the Matched sample T Test
  84. 4:00where each measurement is maybe the same
  85. 4:02person or the same machine something
  86. 4:04like that but of course limiting
  87. 4:06ourselves to the comparison of two
  88. 4:09populations as well limiting the world
  89. 4:12is much more complex than just two
  90. 4:15things what if we wish to compare the
  91. 4:17means of more than two
  92. 4:19populations what if we wish to compare
  93. 4:22populations each containing several
  94. 4:25sublevels or groups well enter another
  95. 4:29NOA so an NOA the acronym comes from the
  96. 4:34phrase analysis of variance so a NOA
  97. 4:38greatly expands what we were able to do
  98. 4:42in
  99. 4:45statistics so suppose we want to compare
  100. 4:48three sample means to see if a
  101. 4:51difference exists somewhere among them
  102. 4:55so our first sample mean is up here in
  103. 4:56the blue xar sub one xar bar sub two as
  104. 5:00our second sample mean here in the pink
  105. 5:02distribution and then our third one is
  106. 5:04down here in the green so each sample
  107. 5:08will have its mean and its own
  108. 5:12distribution so what we are asking is do
  109. 5:15all three of these means come from a
  110. 5:19common
  111. 5:21population so is one mean so far away
  112. 5:25from the other two that it is likely not
  113. 5:28from the same population as those other
  114. 5:31two or are all three so far apart that
  115. 5:36they all likely come from unique
  116. 5:39populations so you can see we're talking
  117. 5:41about sort of the relative distance
  118. 5:43between these means now the variance of
  119. 5:47these distributions is also important
  120. 5:49we'll talk about that later but for
  121. 5:51right now let's focus on the relative
  122. 5:53distance between these means and whether
  123. 5:57or not we could conclude that they come
  124. 5:59come from the same overall
  125. 6:04population so here is our first mean so
  126. 6:07xar sub one with its distribution xar
  127. 6:10sub two with its distribution and then
  128. 6:13xar sub3 with its
  129. 6:16distribution now let's say we take all
  130. 6:19the data points in all three of those
  131. 6:22samples and we put all those data points
  132. 6:24into a common larger distribution so
  133. 6:28we'll put that by Behind these three now
  134. 6:31we're asking oursel is where is each
  135. 6:34mean relative to the overall data set
  136. 6:38sort of in the background you can see
  137. 6:40that xar sub one Falls pretty much right
  138. 6:43down the middle xar sub 2 that mean is a
  139. 6:47bit to the right so you can see the red
  140. 6:49arrow denotes how far it is away from
  141. 6:52the mean of the larger sort of combined
  142. 6:55population what about xar step 3 well it
  143. 6:59is a bit to the left so you can see that
  144. 7:01red arrow denoting how far it is away
  145. 7:05from the mean of the overall sort of
  146. 7:07combined
  147. 7:10population now look at this example so
  148. 7:13xar sub one is where it was sort of
  149. 7:15right at the middle xar sub 2 is a bit
  150. 7:18to the right but now look where xar sub3
  151. 7:21is our third sample mean it's way over
  152. 7:25to the left so we might conclude that
  153. 7:28this mean the third one in the green is
  154. 7:32too far away from the others it's too
  155. 7:34far away from the mean of the larger
  156. 7:36group to be considered as part of that
  157. 7:39larger population it's kind of the
  158. 7:42Oddball so xar sub 2 is about this
  159. 7:44distance you can see the red arrow but
  160. 7:47xar sub3 is way over to the left so is
  161. 7:51this sort of the Oddball distribution
  162. 7:54sort of the weird one sort of the one
  163. 7:56that doesn't belong in the same
  164. 7:58population as as the other
  165. 8:02two now look at this case here we have
  166. 8:05xar sub one that's pretty much right
  167. 8:06down the middle again but now xar sub 2
  168. 8:09the second sample mean is way over to
  169. 8:12the right so you can see that distance
  170. 8:14is pretty far away and look at xar sub
  171. 8:17three its way over to the left so we
  172. 8:21might conclude that each one of these
  173. 8:24sample means belongs to its own
  174. 8:27population only X X bar sub One belongs
  175. 8:31to this distribution in the background
  176. 8:33xar sub2 May belong to one that's off to
  177. 8:35the right and xar sub three May belong
  178. 8:38the one that's off to the left so you
  179. 8:41can see that the means are in very
  180. 8:42different locations relative to the
  181. 8:45overall mean there in the
  182. 8:50background so the null hypothesis in
  183. 8:53these type of problems in an noas is
  184. 8:56whether or not these three sample means
  185. 8:58come from from the same population now
  186. 9:01remember the sample mean is a point
  187. 9:03estimator of the population mean so our
  188. 9:06no hypothesis is that mu sub 1 = mu sub
  189. 9:102 = mu sub3 which is another way of
  190. 9:14expressing that these three means come
  191. 9:17from the same overall
  192. 9:21population now remember we're not asking
  193. 9:23if they are exactly equal we're asking
  194. 9:26if each mean likely came from the same
  195. 9:30larger overall
  196. 9:33population so in an NOA this idea is a
  197. 9:37very specific and important idea we call
  198. 9:41this the variability among or between
  199. 9:44the sample means so each sample mean is
  200. 9:48a certain distance from the mean of the
  201. 9:51overall population in the background and
  202. 9:54we know that that is an expression of
  203. 9:56variance sort of the distance of the
  204. 9:59sample mean from the overall mean in the
  205. 10:01back so this variability between the
  206. 10:05sample means is something I really want
  207. 10:07you to keep in your mind as we proceed
  208. 10:10this is between
  209. 10:14variants so we could test all of these
  210. 10:17sample means using pairwise T tests so
  211. 10:21here are our three sample means so xar
  212. 10:24sub 1 xar sub 2 and xar sub3 now we'll
  213. 10:28block out this part of this little
  214. 10:30Matrix here because in this triangle we
  215. 10:33have comparisons to themselves and
  216. 10:36repeat comparisons so we're only going
  217. 10:38to have three that we could actually do
  218. 10:40so in this first one we could compare
  219. 10:43mean one here in the blue to mean 2
  220. 10:47which is there in the pink so our null
  221. 10:49hypothesis would be xar sub 1 = xar sub
  222. 10:532 we could have a t test that tests that
  223. 10:56now we could also compare mean one and
  224. 10:58mean mean 3 so we would have a t test of
  225. 11:01xar sub 1al xar
  226. 11:04sub3 now we could also test mean 2 and
  227. 11:07mean three so xar sub 2 = xar sub3 now
  228. 11:13notice each one of those independent
  229. 11:15tests has its own Alpha level so Alpha
  230. 11:19Point 05 05 and
  231. 11:2305 now the problem with doing all of
  232. 11:26these pairwise comparisons is that the
  233. 11:29Alpha level is the type one error rate
  234. 11:31of course which means 95% confidence but
  235. 11:35the error compounds with each T Test so
  236. 11:40if we compared each possible pair we
  237. 11:43would have .95 * .95 * .95 so our 95%
  238. 11:49confidence is now 857 or
  239. 11:5585.7% so of course our Alpha is 1 minus
  240. 11:59that so 1 - 857 equal. 143 well what is
  241. 12:05that1
  242. 12:06143 well that is our overall Alpha level
  243. 12:11that is our overall type one error rate
  244. 12:15so our type 1 error rate went from
  245. 12:175% to
  246. 12:2114.3% that is why we do not conduct T
  247. 12:23tests for every possible pair of means
  248. 12:27the error rate compounds and therefore
  249. 12:30the test of course has
  250. 12:35problems now what's different so if I
  251. 12:38stretch each one of these sample
  252. 12:40distributions out what
  253. 12:42changes well the spread or the variance
  254. 12:46of each
  255. 12:48distribution so we call this the
  256. 12:50variability around or within the
  257. 12:55distributions so remember before we
  258. 12:57talked about the variance between the
  259. 13:00distributions and that was sort of the
  260. 13:02distance of each mean from the overall
  261. 13:04population in the background so that was
  262. 13:07variability between this is variability
  263. 13:11within So within each sample
  264. 13:17distribution so at its heart a Nova is
  265. 13:20really a variability ratio it is a ratio
  266. 13:24of the first type of variance we saw the
  267. 13:26variability between the means
  268. 13:29over the variability within the
  269. 13:32distributions that's the second type we
  270. 13:34saw when we stretched each sample
  271. 13:36distribution out so it's just a ratio it
  272. 13:40is between variance divided by Within
  273. 13:45variance so remember in the first type
  274. 13:48we had an overall mean and then the
  275. 13:50distance of each sample mean from that
  276. 13:53so you can see that represented on the
  277. 13:55top and in the bottom we had the
  278. 13:57variability with within the
  279. 13:59distributions so their width or spread
  280. 14:03side to side so on the top we're talking
  281. 14:06about distance from the overall mean and
  282. 14:09on the bottom we're talking about each
  283. 14:11one's spread or width within its own
  284. 14:15sample so distance from the overall mean
  285. 14:18divided by sort of the internal
  286. 14:22spread so we reduce this to a very
  287. 14:25common fraction it is the very
  288. 14:29between divided by the variance within
  289. 14:33so remember this is a ratio of variances
  290. 14:37so of course the F distribution will
  291. 14:39come into play here in a bit so just
  292. 14:42keep in mind this sort of visual tool on
  293. 14:45the top the distance from the overall
  294. 14:47mean and in the bottom the internal
  295. 14:50spread of each sample
  296. 14:55distribution so variance between divid
  297. 14:58by variance within now if we put those
  298. 15:01together those are the components of the
  299. 15:04total variance so for an anova we have
  300. 15:09the total variance sort of split into
  301. 15:11two parts we have the variance between
  302. 15:14the means and then the variance within
  303. 15:17each means distribution so we put those
  304. 15:20together and we have the total variance
  305. 15:23for the entire data set so this is
  306. 15:27called partitioning so we're separating
  307. 15:30this total variance into its two
  308. 15:32component parts and again this is what's
  309. 15:35called a one-way in Nova and I'm not
  310. 15:37going to go into that right now that's
  311. 15:38for the next video but in this type of
  312. 15:41problem sort of the basic problem we
  313. 15:43have total variance and it's made up of
  314. 15:45variance between the means and the
  315. 15:47variance within each
  316. 15:50sample so here are sort of the nuts and
  317. 15:52bolts summary of everything we have here
  318. 15:55if the variability between the means
  319. 15:58sort of the distance from the overall
  320. 15:59mean we saw that in the red before in
  321. 16:02that numerator is relatively large
  322. 16:06compared to the variance within the
  323. 16:09samples So within the actual individual
  324. 16:12samples to spread in that
  325. 16:14denominator then this ratio will be much
  326. 16:19larger than one because the variance
  327. 16:22between will be quite large relative to
  328. 16:25the variance within that will make this
  329. 16:27ratio much larger than one if that's the
  330. 16:31case then the samples most likely do not
  331. 16:35come from a common population so then we
  332. 16:39would reject the null hypothesis that
  333. 16:42these means are equal or come from the
  334. 16:45same population so if that's the case we
  335. 16:49may have one Oddball distribution out to
  336. 16:51the side or all three of them may be so
  337. 16:54far apart that it would create a very
  338. 16:57high ratio because the variance between
  339. 17:00would be much higher than the variance
  340. 17:05within so what could this look like sort
  341. 17:08of in our test in general so if the
  342. 17:10variance between is large relative to
  343. 17:13the variance within being small we would
  344. 17:16reject that null hypothesis that says
  345. 17:18all the means are equal so at least one
  346. 17:22mean is an outlier sort of off to the
  347. 17:24side or they may all three be spread far
  348. 17:27apart and each distribution is
  349. 17:29relatively narrow so they don't sort of
  350. 17:31melt together they are distinct so think
  351. 17:34of this as three distinct distributions
  352. 17:37that are far apart or one that is sort
  353. 17:40of by itself off to the side that would
  354. 17:43create a large variance between the
  355. 17:47means now if the between variance and
  356. 17:50the within variances are similar then we
  357. 17:53would probably fail to reject that null
  358. 17:55hypothesis so they would appear to be
  359. 17:58equal or from the same overall
  360. 18:00population so in this case the means may
  361. 18:02be fairly close to each other fairly
  362. 18:05close to that overall mean and or the
  363. 18:08distributions overlap a bit so they may
  364. 18:11be a bit harder to distinguish from each
  365. 18:13other so if they're very close together
  366. 18:16or the variances within is very wide
  367. 18:19then they'll sort of melt together and
  368. 18:21therefore they will not be
  369. 18:24distinct now the other case is where the
  370. 18:26between variance is very small
  371. 18:29and the variance within is very large so
  372. 18:32you can think of this as three
  373. 18:34distributions that are very spread out
  374. 18:37internally and they not have a whole lot
  375. 18:39of distance from each other so they may
  376. 18:43be close together and or the
  377. 18:44distributions because of very high
  378. 18:46variation they're sort of wide and
  379. 18:48spread out they sort of melt together
  380. 18:51and you really can't distinguish them
  381. 18:53from each other as coming from separate
  382. 18:56or distinct populations so again this is
  383. 18:59a very general look at how we would
  384. 19:01interpret an an NOA and it obviously
  385. 19:04gets more complicated than this but I
  386. 19:06just want to give you a very rough
  387. 19:10overview so here's what we're left with
  388. 19:13this F ratio remember the F ratio is a
  389. 19:16ratio of two variances so we have the
  390. 19:20between variance so again the distance
  391. 19:23of each mean from the overall mean or
  392. 19:26the combined population in the
  393. 19:28background as compared to the variance
  394. 19:32within So within each sample
  395. 19:36distribution now sometimes this is
  396. 19:38called the among variance and the around
  397. 19:43variance so the thing about an noas is
  398. 19:46that I could have three different or
  399. 19:48four different stats textbooks and they
  400. 19:51all call this something different but
  401. 19:53the most common way of expressing it is
  402. 19:55the between variance and the within
  403. 19:57variance but you may see it as the among
  404. 20:00variance the the variance among the
  405. 20:03means and the around variance which is
  406. 20:06the variance around each sample mean but
  407. 20:10it means the same
  408. 20:12thing so this is that partitioning so
  409. 20:15the variance between they're in the red
  410. 20:18arrows plus the variance within they're
  411. 20:21in the blue arrows adds up to the total
  412. 20:26variance now you will also see a term
  413. 20:29called the error or the error variance
  414. 20:33well the error variance is another name
  415. 20:35for the within there in the blue or the
  416. 20:38around so again you might see that in
  417. 20:40the textbook and that's one of the
  418. 20:42challenges of doing these type of videos
  419. 20:44is because stats books like to give
  420. 20:46different names to the exact same
  421. 20:53thing okay so remember why and Nova so
  422. 20:57up to this point we had been been
  423. 20:58comparing just two populations so the
  424. 21:00independent samples T Test and the match
  425. 21:03sample T Test are two examples but
  426. 21:05limiting ourselves to the comparison of
  427. 21:07two populations is of course limiting
  428. 21:10what if we wish to compare the means of
  429. 21:12more than two populations what if we
  430. 21:14wish to compare the populations each
  431. 21:16containing several levels or subgroups
  432. 21:20well that's what we have an NOA for
  433. 21:22remember an NOA stands for the analysis
  434. 21:25of variance now one more thing I want to
  435. 21:28hit on before we go on to the end slide
  436. 21:30and wrap up remember that we're looking
  437. 21:32at the variance between the means as
  438. 21:35compared to the variance within each
  439. 21:39sample so it's all about variance that's
  440. 21:44why it's called analysis of variance
  441. 21:46variance between as compared to variance
  442. 21:49within so we're always looking at a
  443. 21:51ratio of two
  444. 21:56variances okay so that wraps up our
  445. 21:59first video on the analysis of variance
  446. 22:02or an NOA so again I wanted to give you
  447. 22:04a graphical representation of what's
  448. 22:06going on in an Nova so when you're
  449. 22:09working with the table of data and
  450. 22:11you're working with your numbers and
  451. 22:12your F ratios and things like that you
  452. 22:15actually know what is going on in the
  453. 22:19background what we're really trying to
  454. 22:21do is find any distinctions between
  455. 22:24sample means so we may have sample means
  456. 22:27that sort of all line up therefore we
  457. 22:29could conclude that those probably come
  458. 22:31from the same combined population but we
  459. 22:35might have one of the means that's sort
  460. 22:37of an oddball out by itself so in that
  461. 22:40case it probably comes from a different
  462. 22:43population off to the side or we could
  463. 22:45have three sample means that are very
  464. 22:48far apart from each other therefore each
  465. 22:51sample mean may come from its own
  466. 22:54population sort of in the background so
  467. 22:57what we're trying to look look for here
  468. 22:58is distinctions or differences among
  469. 23:01several means and we are doing that by
  470. 23:05looking at two types of variance between
  471. 23:07variant and within variance so a few
  472. 23:10last words and then we are done if
  473. 23:13you're watching this video because
  474. 23:14you're struggling in class stay positive
  475. 23:16and keep your head up I have faith in
  476. 23:19you many other people around you have
  477. 23:21faith in you so so should you if you
  478. 23:24like the video please give it a thumbs
  479. 23:25up share it with classmates or
  480. 23:28colleagues feel free to follow me here
  481. 23:30on YouTube on Twitter on go+ or on
  482. 23:33LinkedIn it's always nice hearing from
  483. 23:35you and finally just keep in mind that
  484. 23:37the fact that you're on here trying to
  485. 23:39learn trying to improve yourself as a
  486. 23:41student or as a business person that is
  487. 23:44what really matters I firmly believe if
  488. 23:46you have the right learning process in
  489. 23:48place the results will take care of
  490. 23:51themselves so thank you very much for
  491. 23:53watching I wish you the best of luck in
  492. 23:55your work and in your studies and I look
  493. 23:57forward to seeing you again next
  494. 24:01[Music]
  495. 24:16time

About this transcript

This page contains the full transcript of YouTube transcript (0Vj2V2qRU10) , generated from the public captions YouTube serves with the video. The transcript has 3,495 words across 495 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.