YouTube2Text

[CS61C FA20] Lecture 33.1 - Thread-Level Parallelism I: Parallel Computer Architectures — Transcript

by CS 61C Departmental · 2,494 words · 417 segments · language en · Watch on YouTube

Full transcript

  1. 0:03hey
  2. 0:04sorry just juggling five balls good to
  3. 0:07see you
  4. 0:08today's topic welcome to 61c's topic on
  5. 0:11thread level parallelism part one
  6. 0:15great to see you as always let's jump
  7. 0:16right in we're going to start with a
  8. 0:18conversation about parallel computer
  9. 0:20architectures
  10. 0:21so big picture step back 10 000 foot
  11. 0:24view
  12. 0:25the goal of 61c is teach you how to
  13. 0:28program a computer really well
  14. 0:29so that you can increase performance
  15. 0:31there are many ways of doing that if you
  16. 0:32were building a system if you're a
  17. 0:34computer engineer
  18. 0:35building a system the first thing you do
  19. 0:36is change the heart rate
  20. 0:38change the speed of your mouse your
  21. 0:40heart rate is 300 to 800
  22. 0:42beats per beats per minute that's a lot
  23. 0:45so think about just turning that crank
  24. 0:47up turning that crank up
  25. 0:49gigahertz machine two gigahertz three
  26. 0:51four five
  27. 0:52that's incredible but we've kind of
  28. 0:54stopped we're going to talk a little bit
  29. 0:55later about why we can't go much past
  30. 0:57that it's really about power dissipation
  31. 0:59we can't keep these chips cool so
  32. 1:01you've kind of maxed out so we increase
  33. 1:02our clock rate as much as we can
  34. 1:04but power issues how much total power
  35. 1:06how to dissipate the power the power
  36. 1:07density all those things are relative to
  37. 1:09that
  38. 1:09relevant in that conversation and we
  39. 1:11can't go much past five
  40. 1:13five uh gigahertz next
  41. 1:16well we can also lower the cpi you know
  42. 1:18cycles per instruction this says lower
  43. 1:20sim d
  44. 1:21sorry this means gosimd this says single
  45. 1:23instruction multiple data can one
  46. 1:25instruction
  47. 1:25operate on wider data can you operate on
  48. 1:27four floats at once that would be
  49. 1:28amazing we've done that and talked about
  50. 1:30that in the last lecture so
  51. 1:31try to go wider so smd you could also
  52. 1:34perform multiple tasks simultaneously
  53. 1:36you could have multiple cpus each
  54. 1:38executing a different program this is
  55. 1:39great
  56. 1:40these tasks could be related they're all
  57. 1:42tight
  58. 1:43working on the same matrix and many guys
  59. 1:44are doing some part of the same matrix
  60. 1:46or
  61. 1:47completely unrelated um you know
  62. 1:50distributed web requests
  63. 1:51or on different computers even you can
  64. 1:54imagine that or or running powerpoint in
  65. 1:56one thing in a quicktime movie another
  66. 1:57thing and
  67. 1:58you know doing a spreadsheet calculation
  68. 1:59and other things so they could be
  69. 2:00related very tight in the same program
  70. 2:02or very different things
  71. 2:03and that's different scales so we'll
  72. 2:04this lecture is about
  73. 2:06the top level the idea of three of just
  74. 2:08doing something about the same program
  75. 2:09you're
  76. 2:10working on the same program the same
  77. 2:11data and how do you divide that up into
  78. 2:13a parallel space that's today's lecture
  79. 2:14and actually the next couple of lectures
  80. 2:15we'll talk about that
  81. 2:17and in summary you can do all three of
  82. 2:18them uh that's the idea so do all
  83. 2:20all the above high high clock frequency
  84. 2:23go
  85. 2:24cimd go wide with your vectors uh and
  86. 2:26multiple parallel tasks as much as
  87. 2:28possible
  88. 2:28so big picture we always show this in
  89. 2:30the have a new model by the way this is
  90. 2:31a new module so if you're behind
  91. 2:33jump jump on this and feel free to reset
  92. 2:35here um this thread level parallelism
  93. 2:37is very exciting so you've seen the left
  94. 2:39i mean let's let's even get the pen
  95. 2:40where's my pen let's even get the pen
  96. 2:42out
  97. 2:42and uh and talk about with this
  98. 2:45you've seen hardware descriptions you
  99. 2:47know if i've got a 32
  100. 2:49bit wide and gate that's 32 operations
  101. 2:52all happening at the same time done that
  102. 2:53already
  103. 2:54you've seen uh parallel instructions
  104. 2:57when we did pipelining
  105. 2:59now you know how you can do up to five
  106. 3:01it's a five stage pipeline there's five
  107. 3:02instructions that can all be
  108. 3:04chunking it through that little slice of
  109. 3:05time can have five instructions all
  110. 3:07active at the same time pretty neat
  111. 3:09parallel data was about cmd can we go
  112. 3:11wide with our vectors can you add four
  113. 3:13vectors in one clock cycle that's great
  114. 3:15the next couple lectures are up here so
  115. 3:17the next couple lectures this space and
  116. 3:18in fact in particular
  117. 3:20this lecture is about parallel threads
  118. 3:22and this next couple lectures about
  119. 3:24talking about thread level parallelism
  120. 3:25and that's the piece of this
  121. 3:27how could you have different cores each
  122. 3:28of them working on different threads
  123. 3:30what are all those mean
  124. 3:31what are architectures what what what's
  125. 3:34required on the architecture point of
  126. 3:35view to make that work
  127. 3:36all that's this lecture and the next
  128. 3:37year's lectures so here's some big
  129. 3:39pictures of parallel computer
  130. 3:40architectures
  131. 3:41early early days view of this i love
  132. 3:44this picture
  133. 3:44this is a picture of just trying to
  134. 3:46throw a couple machines together
  135. 3:47wire them up via ethernet and say let's
  136. 3:50try to have a job that can
  137. 3:51really this is distributed computing in
  138. 3:53some sense can you take a job that is
  139. 3:55far enough and has has enough pieces
  140. 3:57that can be broken off that a whole
  141. 3:59computer can be asked to work on that
  142. 4:00piece and then somehow collect that
  143. 4:02together
  144. 4:02so usually there's like a command module
  145. 4:04that divides it up
  146. 4:05all these kind of worker bees work on
  147. 4:07something that come back but you're
  148. 4:08communicating via ethernet or via a
  149. 4:10shared drive or something like that
  150. 4:12when you go full scale this is what
  151. 4:15happens at all the national labs
  152. 4:16um this is what we have in super
  153. 4:18computers so you have these massive
  154. 4:20racks massive
  155. 4:21server racks with an array of computers
  156. 4:24and
  157. 4:24uh we'll even talk about that a little
  158. 4:26later when we talk about warehouse scale
  159. 4:27computers
  160. 4:28most of those computers just fyi are
  161. 4:31working on
  162. 4:32computational science problems they're
  163. 4:35usually working on floating point data
  164. 4:36you're doing matrix multiplication
  165. 4:38you're doing neural network stuff you're
  166. 4:39you're doing uh climate simulation a lot
  167. 4:41of these things all of them the actual
  168. 4:43data that you care about the most
  169. 4:45is floating point data it's not kind of
  170. 4:47integer based data where it's
  171. 4:48you're thinking about ins and choose
  172. 4:50complement numbers you're doing doubles
  173. 4:52mostly almost dealing with doubles or
  174. 4:53even more if you need higher precision
  175. 4:55for that so that's what those machines
  176. 4:56are doing and those are
  177. 4:57unbelievable you go to super computer
  178. 4:58conferences they're those machines and
  179. 5:00they're talking about how to crunch on
  180. 5:02massive massive massive floating point
  181. 5:04data sets
  182. 5:06in your phone in your watch in your
  183. 5:08computer we now have
  184. 5:10also multi-core machines we've talked a
  185. 5:12little bit about that we'll get a little
  186. 5:13bit deeper into that how that works
  187. 5:15and these multi-core this is on a single
  188. 5:16chip there's on a single chip multi-core
  189. 5:18machines there's
  190. 5:19some graphics there's maybe a shared l3
  191. 5:21cache l1 and l2 are separate
  192. 5:24for each core um it's incredible and
  193. 5:26we've kind of shown a picture of that
  194. 5:27before and that's another architecture
  195. 5:29all these are
  196. 5:30ways to think about parallelism at
  197. 5:32different scales actually different
  198. 5:33scales
  199. 5:34so let's dig deeper into that bottom
  200. 5:36right picture
  201. 5:37what's happening if you have a cpu with
  202. 5:39two cores so again
  203. 5:41single dis single die two cores in that
  204. 5:43let's look so very simple
  205. 5:45so what does that mean two cores does
  206. 5:46that mean you're doubling memory so now
  207. 5:48you know my 32 gibby bytes of memory or
  208. 5:50maybe more you know computers nowadays
  209. 5:52you've seen virtual memory
  210. 5:53you know the idea that there's a
  211. 5:54distinction between it was it was all a
  212. 5:56lie between
  213. 5:57how much actual memory you have and the
  214. 5:59address space
  215. 6:00i can have a 32-bit address space that's
  216. 6:02four gigabytes of allocatable memory
  217. 6:05how do i have 10 programs all of which
  218. 6:07are
  219. 6:08controlling all four gigabytes well
  220. 6:10that's the idea of virtual memory each
  221. 6:11of them has a virtual address space
  222. 6:12that's that big
  223. 6:13which that means now now i can have an
  224. 6:14actual computer that has not exactly 4
  225. 6:17gb bytes if you didn't have that
  226. 6:18abstraction you'd have to have a machine
  227. 6:19that has exactly four gigabytes and only
  228. 6:21one program can be running at a time
  229. 6:22or you could be switching and you know
  230. 6:24time slicing it
  231. 6:26but now with virtual memory you've seen
  232. 6:27before we can have many machines each of
  233. 6:30them thinks they have the whole 32-bit
  234. 6:31address space or even a 64-bit address
  235. 6:33space and the amount of actual memory
  236. 6:34you have
  237. 6:35can be more than 4 gb or less than 4 gb
  238. 6:38uh
  239. 6:38certainly not 64. if you're living in 64
  240. 6:40addresses but mostly that risk 532
  241. 6:42means it's a it's a 32-bit virtual
  242. 6:44address space how much actual memory you
  243. 6:46have doesn't matter i mean you could
  244. 6:48have a huge amount
  245. 6:49they not sell 1.5 gibby bytes of ram on
  246. 6:52the new mac pros incredible
  247. 6:53and you can have a machine that has very
  248. 6:55little and still you can still run
  249. 6:5632-bit address-based
  250. 6:57programs all right so we talked about
  251. 7:00one die
  252. 7:01with two cores on it does that mean
  253. 7:03you're somewhat doubling memory you're
  254. 7:04not because the core doesn't memory on
  255. 7:05it anyway the core is over here memory's
  256. 7:07over there
  257. 7:08so what does it mean to do this it means
  258. 7:09you're separating two different
  259. 7:11processor cores
  260. 7:12each of which each of which has its own
  261. 7:15control
  262. 7:16its own data path its own pc
  263. 7:20its own set of registers its own aou
  264. 7:22they are really two full
  265. 7:23of quote unquote computers on one die
  266. 7:27those two cores are on one die and they
  267. 7:29really
  268. 7:30are uh fully autonomous that can be
  269. 7:32completely running
  270. 7:33separate things i'm doing a load you're
  271. 7:35doing a store i'm doing an ad you're
  272. 7:36doing a subtract
  273. 7:37completely separate different registers
  274. 7:39running two programs at the same time
  275. 7:41we love this that's a two core model
  276. 7:44what is this shared part of it though
  277. 7:45well the shared part of it is i've got
  278. 7:47one memory
  279. 7:48block and io is working within that and
  280. 7:51i've got my io interface here we've
  281. 7:52talked about
  282. 7:53io a little bit in previous lectures so
  283. 7:56there's has to be some communication
  284. 7:58where each processor is making load and
  285. 8:00store requests
  286. 8:02to memory and now you have your
  287. 8:03conversation well what happens ooh
  288. 8:05when they're both writing and reading to
  289. 8:06the same spot or writing and reading how
  290. 8:08do you synchronize that there's some
  291. 8:09conversations lower down
  292. 8:11in that but at the big picture this is a
  293. 8:1310 mile view up
  294. 8:14you have two different course everything
  295. 8:16you know about all we
  296. 8:17you know everything up to this lecture
  297. 8:19was just this picture
  298. 8:20this is the picture up to now and now
  299. 8:22we're saying well i want to have two
  300. 8:23cores on the same die
  301. 8:25well let's have a second guy and this
  302. 8:26interface here's the extra here's the
  303. 8:28extra interface
  304. 8:29same the same connection you had memory
  305. 8:31loads and stores that's the way you can
  306. 8:32make this work
  307. 8:33there is a little bit more details with
  308. 8:35the caches but from this 10 mile view up
  309. 8:37you have two cores
  310. 8:38two distinct control data paths pcs
  311. 8:40registers and leus all working with one
  312. 8:42shared memory okay this is called a
  313. 8:43shared memory model
  314. 8:45okay so now let's kind of dig a little
  315. 8:47deeper this is the last slide on this
  316. 8:48first
  317. 8:49first lecture each processor is
  318. 8:52executing its own instruction
  319. 8:53set it's chunking through two different
  320. 8:55programs all running in parallel
  321. 8:57separate resources data path pc
  322. 8:59registers alu you said that already
  323. 9:01highest level caches are also going to
  324. 9:03be separate you want to be able to not
  325. 9:04have to go to sacramento
  326. 9:05as you're having you know you're you're
  327. 9:07as you're making memory requests and
  328. 9:09it's not there you want to be able to
  329. 9:10save it and have your your net
  330. 9:12is there your l1 and l2 net is there
  331. 9:14saving you from going to sacramento
  332. 9:15because that's that one shared spot
  333. 9:17you do have some shared resources we
  334. 9:19mentioned memory is a shared resource
  335. 9:21it's expensive i mean the 1.5 gibby
  336. 9:23bytes of of dram is very expensive so
  337. 9:25you're not going to duplicate that per
  338. 9:27core
  339. 9:27you got many many cores in the newest
  340. 9:28machines that's certainly not duplicated
  341. 9:30you will duplicate l1 and l2 typically
  342. 9:33and l3 is funny l3 is usually really big
  343. 9:36uh and so usually in the mega megabytes
  344. 9:39so
  345. 9:39you're probably gonna have a shared l3
  346. 9:42although not required uh it's often not
  347. 9:44it's often on the same silicon chip
  348. 9:45doesn't have to be all that's still in
  349. 9:47the model of a multi-processor execution
  350. 9:48model
  351. 9:49doesn't have to be there you could
  352. 9:50decide well there's no l3 at all that's
  353. 9:52still fine it's still in the category of
  354. 9:53a multi-processor system
  355. 9:56here's some new words you might not have
  356. 9:57heard you know our microprocessor is
  357. 9:59what all the cpus have been talking
  358. 10:00about before
  359. 10:01a multi-processor microprocessor that
  360. 10:03means
  361. 10:04two or more cores in one microprocessor
  362. 10:07um you see two more processors
  363. 10:09processors could be different by the way
  364. 10:11my first computer i bought my first big
  365. 10:12computer i bought when i got hired
  366. 10:14was uh uh is an early mac
  367. 10:18stand desktop mac uh and it had
  368. 10:21two different dies if you looked at two
  369. 10:23different dies two different heat sinks
  370. 10:24two different dies and each one of them
  371. 10:26had two different cores on it that's
  372. 10:27pretty neat so a multi-processor
  373. 10:30microprocessor both means
  374. 10:31multiple dies working on some shared
  375. 10:34memory area or
  376. 10:35a single die with multiple cores so the
  377. 10:37second idea and the most recent idea
  378. 10:39now is let's forget it's two weeks too
  379. 10:41weird and too hard people haven't done
  380. 10:42this
  381. 10:42in recent years that have multiple dies
  382. 10:45multiple
  383. 10:46dies each of them has a huge heatsink
  384. 10:47very expensive you now just have one you
  385. 10:49should just have one
  386. 10:50die one cpu in some sense
  387. 10:54but it's multi-core that one die is a
  388. 10:56single cpu but it has
  389. 10:58and we're going to say i say the word
  390. 11:00cpu is a little fuzzy now i don't want
  391. 11:01to get into that
  392. 11:02but in the past they would say two cpus
  393. 11:04with two cores on each now they actually
  394. 11:05call one
  395. 11:06die multiple cpus we'll we'll talk about
  396. 11:09this a little later on
  397. 11:10because they've fuzzed they've fuzzed
  398. 11:12around the naming of that the cpu used
  399. 11:14to be
  400. 11:14a die so now they have a single cpu
  401. 11:18with four cores on them that's what they
  402. 11:19call it okay and so we'll talk about
  403. 11:22logical versus physical cpus we'll talk
  404. 11:24about that at
  405. 11:25later lectures but that's what we've got
  406. 11:26and the idea is here i go if i've got uh
  407. 11:28if i've got
  408. 11:29this last slide we had two cores this
  409. 11:31means two full
  410. 11:32instruction streams we'll have some word
  411. 11:34for that in later slides instruction
  412. 11:36seems operated simultaneously that's it
  413. 11:37pretty exciting
  414. 11:38let's actually get down to more details
  415. 11:40of how that works and some more
  416. 11:41ideas of how does this work in software
  417. 11:43in the next lecture we'll see you there

About this transcript

This page contains the full transcript of [CS61C FA20] Lecture 33.1 - Thread-Level Parallelism I: Parallel Computer Architectures by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 2,494 words across 417 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.