YouTube2Text

[CS61C FA20] Lecture 33.2 - Thread-Level Parallelism I: Multicore — Transcript

by CS 61C Departmental · 3,532 words · 570 segments · language en · Watch on YouTube

Full transcript

  1. 0:00and welcome back now let's talk about
  2. 0:03dig dig a little deeper into what we
  3. 0:04mean by having a multi-core computer
  4. 0:06with a little bit of a history lesson
  5. 0:08so this is a wonderful curve i really
  6. 0:10love this curve it's not my data but
  7. 0:12it's a wonderful
  8. 0:13visualization of the evolution to
  9. 0:16multi-core and why
  10. 0:17why we got to where we are now so let's
  11. 0:20go back in time
  12. 0:21um we're in the early 70s we're looking
  13. 0:25at
  14. 0:25the total number of transistors and
  15. 0:27thousands what i love about this view on
  16. 0:28the left
  17. 0:29is that it's unitless the units are
  18. 0:31attached to each of the curves
  19. 0:33so as i look to moore's law moore's law
  20. 0:36says how many ics how many integrated
  21. 0:38circuits
  22. 0:38how many transistors are on an ic so how
  23. 0:40many transistors are on an integrated
  24. 0:42circuit
  25. 0:42and we're looking at this number this
  26. 0:45number was doubling every 18 months or
  27. 0:47so
  28. 0:48which is a nice line here and then at
  29. 0:50some point
  30. 0:51we actually were doubling so doubling
  31. 0:53every two years sorry doubling every two
  32. 0:54years in the early stages
  33. 0:56and at some point it kind of had a small
  34. 0:58bend and we started doubling every 18
  35. 1:00months
  36. 1:01and we're doing amazing so we're just
  37. 1:03packing more and more transistors on
  38. 1:06the the feature size is getting smaller
  39. 1:08it's getting down to
  40. 1:09some some number some kind of nanometers
  41. 1:1165 50 30
  42. 1:1320. you know it's getting smaller and
  43. 1:14smaller which means that the um between
  44. 1:17two lines it's getting shorter and
  45. 1:18smaller and smaller between the distance
  46. 1:20between two signal lines on on the chip
  47. 1:22amazing what they're doing great
  48. 1:24incredible engineering this is an
  49. 1:25incredible amount of engineering going
  50. 1:26on here
  51. 1:28if we continue to look here we're also
  52. 1:31seeing
  53. 1:32the frequency is going up remember i
  54. 1:33mentioned just increase that heart rate
  55. 1:35increase that heart rate do this and in
  56. 1:36fact it even took a bend
  57. 1:38where we even went steeper there um now
  58. 1:41we're also seeing the power go up so
  59. 1:43this is a story
  60. 1:44this is a story in which these these
  61. 1:45curves are all related in interesting
  62. 1:46ways
  63. 1:47and the power was going up it was still
  64. 1:49small we didn't care about it
  65. 1:50you know in the 90s it was fine and by
  66. 1:52the way this was a great time
  67. 1:54for well okay great time for computer
  68. 1:56manufacturers because they could
  69. 1:58argue that you need to buy this year's
  70. 2:00computer um
  71. 2:01because it's twice as fast and every
  72. 2:03year it was every year you know every
  73. 2:05couple of years you needed to upgrade
  74. 2:07you you just really couldn't
  75. 2:08um live with the computer that you had
  76. 2:10bought two or three years ago because
  77. 2:12boy the the potential to have a computer
  78. 2:15that was 200 megahertz
  79. 2:16and now it's 500 just based on only the
  80. 2:19the clock rate
  81. 2:20um you know 200 meg 250 megahertz okay
  82. 2:24500 megahertz 800 megahertz
  83. 2:26a gigahertz oh man we crossed the
  84. 2:27gigahertz line two 1.5 2 gigahertz wow i
  85. 2:30got
  86. 2:31love this so we meant for a couple of a
  87. 2:33lot of things in the computing
  88. 2:34engineering ecosystem
  89. 2:36it meant that you're a computer
  90. 2:37manufacturer you were doing great
  91. 2:39because you you could just continue to
  92. 2:40push
  93. 2:40machines out and there were different
  94. 2:42specs and people were comparing and then
  95. 2:43all the gamers are there
  96. 2:44the computer scientists were on this
  97. 2:46side and the computational scientists
  98. 2:47who were using it for floating were they
  99. 2:49all the people were benefiting from
  100. 2:50this this incredible this all this with
  101. 2:52all these curves were benefiting
  102. 2:53everybody using these systems
  103. 2:57what's interesting a little bit is that
  104. 2:59we didn't care about power that much but
  105. 3:01people began to have more and more as
  106. 3:03this number
  107. 3:04started to get up here look at this
  108. 3:06we're now in the hunt i mean look at
  109. 3:07this this is you know
  110. 3:08100 watts that's an incandes
  111. 3:10incandescent light bulb
  112. 3:11so people used to say these systems are
  113. 3:13getting too hot we can't cool them
  114. 3:16that was the issue we couldn't cool
  115. 3:17these systems down you had very arcane
  116. 3:19heatsinks
  117. 3:20with a radiator fan and a huge i mean a
  118. 3:22radiator radiator system here with with
  119. 3:25a lot of
  120. 3:25a lot of spokes and then a fan tried to
  121. 3:28blow and get that heat off of it
  122. 3:30and so you'd have you'd heat up a room
  123. 3:31if you if you were on the you know
  124. 3:33exhaust fan side of
  125. 3:34one of some of these computers you would
  126. 3:35be really hot be really quite warm
  127. 3:37um and they started having liquid
  128. 3:40cooling because they could
  129. 3:41even just air off of a radiator system
  130. 3:43couldn't cool
  131. 3:44off the heat sink off of a cpu so you
  132. 3:46had water coming in and then you had oil
  133. 3:48come and very
  134. 3:50intricate engineering to get fluid to
  135. 3:52get in there and touch it as much as you
  136. 3:53can get warm and then
  137. 3:54get it out here and dissipate it and
  138. 3:56then cool it down and come back in
  139. 3:57incredible engineering to try to just
  140. 3:59cool these things down
  141. 4:01at some point they said we're not able
  142. 4:03to cool it anymore
  143. 4:04um we have to do something about this we
  144. 4:07can keep turning the clock frequency up
  145. 4:08and by the way there are other
  146. 4:09architectural improvements as well i
  147. 4:11mentioned before
  148. 4:11you know we're going cmd and going wide
  149. 4:13or we're doing we're seeing a lot of
  150. 4:14performance
  151. 4:15but if you look at the frequency right
  152. 4:18around
  153. 4:18a special thing happened a sea change in
  154. 4:21computer
  155. 4:22architecture happened around 2005 look
  156. 4:25at this number
  157. 4:25everything changes around that look at
  158. 4:27that number
  159. 4:29you saw the clock frequency start to
  160. 4:32start to level off
  161. 4:34well that's an issue the clock frequency
  162. 4:36level's off
  163. 4:37now you've got the power levels off so
  164. 4:41you wanted to
  165. 4:41get to that you wanted to level off the
  166. 4:43power and you did by turning the clock
  167. 4:44down
  168. 4:45now you're just kind of doing the same
  169. 4:46thing as you did before in the clock
  170. 4:47that but now all of a sudden what's
  171. 4:49being hit
  172. 4:50by the way you were still being able to
  173. 4:51continue moore's law you were still
  174. 4:53having
  175. 4:54more and more transistors being able to
  176. 4:56be you know squeezed into one cpu
  177. 5:00but you were seeing the sequential app
  178. 5:02performance all the programmers
  179. 5:04who back here were lazy all see this
  180. 5:06curve
  181. 5:07this curve says you could be lazy and
  182. 5:09write a bad algorithm
  183. 5:11not optimize it and do nothing do
  184. 5:13nothing in terms of trying to squeeze
  185. 5:14more performance out of your software
  186. 5:16and guess what the performance of just
  187. 5:18going up every year so lazy programmers
  188. 5:21they're hard working programs everywhere
  189. 5:22but people you didn't have to work worry
  190. 5:23so much as a programmer as a developer
  191. 5:26on how to squeeze the last strip of
  192. 5:27performance out because you were getting
  193. 5:29it for free you were getting this for
  194. 5:30free
  195. 5:31and all of a sudden 2005 hits and you
  196. 5:34stopped getting it for free
  197. 5:37so this is an issue you know where now
  198. 5:39we're going to the cloud now we're
  199. 5:40having larger
  200. 5:41larger larger data sets big data became
  201. 5:43a big deal
  202. 5:44but you couldn't handle it you had no
  203. 5:47way to work with that big data you
  204. 5:48weren't going to get any faster if the
  205. 5:50data size if the data set size you know
  206. 5:51your pixar and you have
  207. 5:53more and more refined models all of a
  208. 5:55sudden you can't
  209. 5:56your computers aren't processing them
  210. 5:58any faster it's
  211. 5:59flat so enter so 2005 was a sea change
  212. 6:03and enter the idea that you would have
  213. 6:06more cores this is again a history
  214. 6:08lesson so in 2005
  215. 6:11with i believe the intel core 2 duo was
  216. 6:13the first that i remember
  217. 6:14was a production level cpu that had more
  218. 6:16than one core there were architectures
  219. 6:18you had two different dyes
  220. 6:20before them but never two cores in one
  221. 6:22die and so this number of cores
  222. 6:24started to go up and it was amazing what
  223. 6:26that number is
  224. 6:27and then i'm just starting to grow as
  225. 6:29well right this is an exponential curve
  226. 6:30a linear
  227. 6:30a linear line on a on a log a linear log
  228. 6:34plot as an exponential curve so
  229. 6:35all these are straight lines but they're
  230. 6:36exponentials so that started to go up
  231. 6:40and double something of some number of
  232. 6:41years and then all of a sudden you saw
  233. 6:43this is the key
  234. 6:44the parallel app performance continued
  235. 6:46to go up and so it said to all software
  236. 6:48developers you better learn how to
  237. 6:49program in parallel
  238. 6:50you better learn how to work with all
  239. 6:51those cores effectively because you're
  240. 6:53if you want to live on the yellow line
  241. 6:54the sequential app performance you're
  242. 6:56flat you're done
  243. 6:57you want to continue to be able to
  244. 6:58satisfy your hungry customers
  245. 7:00hungry for data hungry for processing
  246. 7:02power you better learn how to program
  247. 7:03these core as well and this is what the
  248. 7:04goal of these set of lectures are
  249. 7:07so let's get that's a sample i'm just
  250. 7:09literally sampling
  251. 7:10i you know we live here in the bay area
  252. 7:12let's just let's go down the road a
  253. 7:13little bit to cupertino
  254. 7:14this was one of their slides when they
  255. 7:16released the apple a14 chip this is on
  256. 7:18their phones
  257. 7:19what are they saying what are the
  258. 7:20highlights of the a14 ship
  259. 7:23six core cpu by the way this is not a
  260. 7:25phone this is on a phone okay this is
  261. 7:26not the one that's the
  262. 7:27latest and greatest on a phone six cores
  263. 7:30and when in fact not not just six cores
  264. 7:32they actually categorize some number of
  265. 7:33cores that are
  266. 7:34i think they call them firestorm and ice
  267. 7:35storm like high performance cores and
  268. 7:37then look
  269. 7:38like lower performance cores cooler they
  270. 7:40call it fire and ice
  271. 7:42cooler of course not as fast but you
  272. 7:44have four of them so
  273. 7:45it's interesting to be able to
  274. 7:46distinguish between when you really need
  275. 7:48to crank it out put them all on
  276. 7:50but if you're just kind of doing normal
  277. 7:51things maybe having more cores but
  278. 7:53really
  279. 7:53more energy efficient cores is better so
  280. 7:55maybe change the clock rate possibly
  281. 7:57thinking about how they do
  282. 7:58that so that's the first thing they have
  283. 7:59different categorizations for their six
  284. 8:00scores six cores on a phone
  285. 8:01six cores and a phone nobody i mean 2005
  286. 8:05no computer had multiple cores maybe you
  287. 8:07were lucky if you had some special you
  288. 8:09know
  289. 8:09high-end thing but most computers i just
  290. 8:11had one computer one cpu
  291. 8:13one core and now we're talking about
  292. 8:15phones having six cores okay it's a new
  293. 8:16world
  294. 8:17i've got a four core gpu that's
  295. 8:19incredible i've got a 16 core neural
  296. 8:22engine for for neural networking work
  297. 8:24this is unbelievable and i also have
  298. 8:25another so this is on one die this is a
  299. 8:27chip
  300. 8:28the chip is in the middle folks this is
  301. 8:30the chip this you buy this
  302. 8:32a14 chip and you get all these pieces
  303. 8:34inside of it okay 11 billion transitions
  304. 8:36an incredible number and you got an
  305. 8:37image processor as well so a lot of
  306. 8:39you're categorizing it not just well i'm
  307. 8:41just going to do the normal cpu be
  308. 8:42boring no i'm kind of here's the gpu
  309. 8:44over here there's a neural engine here's
  310. 8:45an image processing engine it's
  311. 8:46incredible what they're putting together
  312. 8:48so here's the execution model here's as
  313. 8:50we think about this
  314. 8:52okay now that's the world you're in a
  315. 8:53multi-core world what is shared and what
  316. 8:55isn't shared
  317. 8:56so each core has as the picture we
  318. 8:59showed in the last lecture each core has
  319. 9:00access to the entire memory okay so
  320. 9:02every core has
  321. 9:03can connect to that entire memory there
  322. 9:04that's fine
  323. 9:06you got to worry about keeping the
  324. 9:07caches consistent in fact cache
  325. 9:09coherency
  326. 9:10l1 l2 l3 making sure that's all clean
  327. 9:13and coherent coherent means just all on
  328. 9:15the same page in some sense page is the
  329. 9:17wrong word because of virtual memory but
  330. 9:18the idea
  331. 9:18all they're all kind of concluded
  332. 9:20together and not inconsistent
  333. 9:22that's an important thing and we're
  334. 9:23going to talk about that about about
  335. 9:25three lectures from now
  336. 9:26um three full lectures from now um not
  337. 9:29three sub mod pieces of one lecture
  338. 9:32um so the advantage is you can actually
  339. 9:34now communicate i can have these chords
  340. 9:36that they needed to communicate via
  341. 9:37memory if i need to have this core talk
  342. 9:38to that core
  343. 9:39remember each one of those only has its
  344. 9:40own set of registers and its own caches
  345. 9:42but if i wanted to communicate this guy
  346. 9:43i could write something to a spot and
  347. 9:45then i could say
  348. 9:45and wait and then you could read from
  349. 9:47the spot so there's a little bit of
  350. 9:48communication through memory as the
  351. 9:49shared memory model
  352. 9:50so that's a shared variable they have a
  353. 9:51shared variable that i'm working
  354. 9:52together which is very nice
  355. 9:53you could have you all be running to the
  356. 9:54same you'll be writing to the same array
  357. 9:56let's all contribute to one big array
  358. 9:58let's all be reading from an array and
  359. 9:59then writing to the answers in different
  360. 10:00spots we'll see this we actually
  361. 10:02see how we do this in software a little
  362. 10:04bit
  363. 10:05the drawbacks is that here's the catch
  364. 10:08memory is slow
  365. 10:09i'm if i'm gonna the only way these guys
  366. 10:10can talk is to go to sacramento that's
  367. 10:13what we're talking about right i'm gonna
  368. 10:14write to a variable
  369. 10:15and it's in memory and then you're gonna
  370. 10:16read from the variable or we're all
  371. 10:17right into the same thing whatever how
  372. 10:18we do our work
  373. 10:20memory's slow going to sacramento was
  374. 10:22slow so that's a downside and
  375. 10:23that's a bottleneck i think i have a
  376. 10:25single bottom like we've really talked
  377. 10:26about amdahl's law that's in a couple of
  378. 10:27lectures after this
  379. 10:29law says the serial part of a program is
  380. 10:31going to be your downfall it's going to
  381. 10:32really be the place that is the
  382. 10:34bottleneck for any piece of software and
  383. 10:36i say serial well
  384. 10:38there's only one piece of memory so i'm
  385. 10:39kind of serially writing through this i
  386. 10:40can't
  387. 10:41in parallel think about that so if i'm
  388. 10:43all running to the same spot that's an
  389. 10:44issue in terms of synchronization so
  390. 10:46that's an issue
  391. 10:48you can think about using a
  392. 10:49multi-processor in different way you
  393. 10:50have what's called job level parallelism
  394. 10:52and they're working on separate problems
  395. 10:53okay so one core is working on a movie
  396. 10:56and one core is working on a spreadsheet
  397. 10:57so that's that's
  398. 10:58totally reasonable and there's no
  399. 10:59communication or you could have a more
  400. 11:01complex and more kind of common
  401. 11:03task which is well i need to be
  402. 11:05converting something i'm playing a movie
  403. 11:07while you can be working on this part of
  404. 11:08the screen and i can be working on these
  405. 11:10pixels like
  406. 11:10so you could have multiple cores all
  407. 11:12working on the same piece of memory
  408. 11:14and that's like a maybe a large matrix
  409. 11:15or a large array of doing something and
  410. 11:17that's
  411. 11:17really interesting and that's what we'll
  412. 11:18be speaking about for the next couple of
  413. 11:19lectures is doing that
  414. 11:21in general big picture parallel
  415. 11:23processing in fact we often
  416. 11:25this is under the umbrella of parallel
  417. 11:27and distributed computing that's the
  418. 11:28idea
  419. 11:29this is all hard doing things at the
  420. 11:31same time is hard um
  421. 11:32it's inevitable we talked about before
  422. 11:34we've kind of run up against
  423. 11:36the trying to get more squeeze more
  424. 11:38performance out of a stone and we can't
  425. 11:40get it so we
  426. 11:41it's almost like we have to it wasn't
  427. 11:43like i'm so excited to go parallel
  428. 11:44because it's
  429. 11:45making programming easier it's not it's
  430. 11:46making programming a lot harder but we
  431. 11:48have to we have it's almost like eating
  432. 11:50your broccoli we have to go there
  433. 11:51to get the performance out of these
  434. 11:53machines that are being built for us now
  435. 11:54so it's inevitable the only path to
  436. 11:56increasing performance really is going
  437. 11:57parallel and understanding that
  438. 11:58and also it's the only way to deal with
  439. 12:00battery life issues um and power issues
  440. 12:02i mentioned before it isn't the case
  441. 12:03that we can keep just turning the clock
  442. 12:05speed up
  443. 12:05and relying on a computer to be faster
  444. 12:09every year that's flat these numbers are
  445. 12:11all flat the only way to continue to
  446. 12:13grow is and to continue to
  447. 12:14make use of the hardware that's coming
  448. 12:16out of here is to think about how to
  449. 12:17program parallel
  450. 12:18systems well in mobile systems we see
  451. 12:21this i mentioned before
  452. 12:22in phones we have multiple cores and my
  453. 12:23watch has two cores my new watch
  454. 12:25has apple series with a four five or six
  455. 12:27that has two cores it's
  456. 12:28it's remarkable and different sub
  457. 12:30components of that as well um
  458. 12:32you also have different categorizations
  459. 12:34it's dedicated processors for motion and
  460. 12:35for image and for neural processing at
  461. 12:37least on the
  462. 12:38on the iphones and the gpu has always
  463. 12:39been usually the gpu's been there in
  464. 12:40fact in the early days of a 6881 there
  465. 12:43was a special floating point unit for
  466. 12:44that so we've always had kind of
  467. 12:45dedicated hardware for things i say
  468. 12:47always i mean
  469. 12:48in my lifetime-ish but now it really is
  470. 12:50the case on one die there are pieces of
  471. 12:52it usually
  472. 12:52here's a floating point co-processor
  473. 12:54boom a separate dive for floating point
  474. 12:56and here's the main cpu and here's the
  475. 12:57floating point then
  476. 12:58let's make it let's let's pull the gpu
  477. 12:59out have a big gpu now you have a huge
  478. 13:02zoom with big fans and you pay a lot of
  479. 13:03money for the high end gpus the gamers
  480. 13:05who do this so that's we've almost
  481. 13:07always had either floating point
  482. 13:09separate from the enabled main cpu or
  483. 13:10the graphic separate from the main cpu
  484. 13:12this is useful
  485. 13:13um we're going to talk about how to do
  486. 13:15this at a different scale what happens
  487. 13:16when rather than talk about one core
  488. 13:18one die and how to manage that what
  489. 13:19happens when we're thinking about many
  490. 13:21many many machines
  491. 13:22all in one warehouse like facebook and
  492. 13:25google and
  493. 13:26apple for siri at least how many folks
  494. 13:30amazon obviously how many people how do
  495. 13:33you actually think
  496. 13:34about the warehouse scale level of that
  497. 13:35if you're if you're controlling
  498. 13:37whatever maybe a million computers each
  499. 13:39computer has many cores
  500. 13:41how are you able to program that we'll
  501. 13:42talk about even software ways to do that
  502. 13:44it's very exciting once we get to the
  503. 13:45ability
  504. 13:45ability to be able to you know hit a
  505. 13:47couple lines of python even
  506. 13:49boom and all of a sudden a thousand
  507. 13:50machines can wake up and and you know do
  508. 13:52your bidding that's pretty remarkable so
  509. 13:54you have multiple nodes each of these
  510. 13:56modes have multiple cpus and disk and
  511. 13:57memory and think about it's one big unit
  512. 13:59which is incredible and how to cool it
  513. 14:01is the big deal
  514. 14:02we'll get to that little later but you
  515. 14:03obviously have mimdi
  516. 14:05which is multi-core and simdi which is
  517. 14:06these vector wide operations in each
  518. 14:09node
  519. 14:10let's talk about some performance let's
  520. 14:12we're trying to this is a big picture
  521. 14:13kind of lecture
  522. 14:13um so this is a sense of uh the year of
  523. 14:17you know
  524. 14:17kind of every odd year for the last uh
  525. 14:1915 years or so 18 years or so
  526. 14:22so let's talk 2009 to kind of almost now
  527. 14:25or in these
  528. 14:25upcoming years so how how is the number
  529. 14:28of course change
  530. 14:29you know where we're growing we're on
  531. 14:31average we're kind of growing obviously
  532. 14:32they're higher ones that are now
  533. 14:33i think we'll see that the latest was 28
  534. 14:35cores or something but these numbers are
  535. 14:37you know
  536. 14:37doubling uh every uh uh
  537. 14:41remarkable doubling every two and a half
  538. 14:42years or so
  539. 14:44this is that's a two and a half times
  540. 14:46increase total in the number of course
  541. 14:48over those 12 years this is an eight
  542. 14:50times increase in the number of cmd bits
  543. 14:52per core
  544. 14:53which is nice if you multiply the core
  545. 14:56times the 70 bits the total floating
  546. 14:57point operations you can do
  547. 14:59per cycle is a factor of 20. so that's
  548. 15:0220 times in 12 years
  549. 15:04we're basically doubling every three
  550. 15:06years if we can use it
  551. 15:08so now i've got this system at the
  552. 15:09bottom line how do we make use of all
  553. 15:11this
  554. 15:12stuff how do we actually from the point
  555. 15:13of view software why i came out of
  556. 15:15vanilla computer science school you
  557. 15:16can't transplant somebody from taking
  558. 15:181970s computer science and put them
  559. 15:20today and expect them to be productive
  560. 15:22because they didn't have any courses on
  561. 15:24how to make use of all that parallel
  562. 15:26hardware it's really hard so let's
  563. 15:28and to some issues there so let's close
  564. 15:30it there
  565. 15:31and uh think about how we would even
  566. 15:34make use of that bottom row if you had
  567. 15:36the crazy high-end machine in the bottom
  568. 15:38row
  569. 15:38how you can even make use of that we'll
  570. 15:40see the next lecture

About this transcript

This page contains the full transcript of [CS61C FA20] Lecture 33.2 - Thread-Level Parallelism I: Multicore by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 3,532 words across 570 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.