YouTube2Text

[CS61C FA20] Lecture 35.2 - Thread-Level Parallelism III: Shared Memory and Caches — Transcript

by CS 61C Departmental · 1,547 words · 241 segments · language en · Watch on YouTube

Full transcript

  1. 0:00and welcome back
  2. 0:03now let's actually think about what
  3. 0:05happens
  4. 0:06under the hood so we've talked about the
  5. 0:08software level so we're talking about
  6. 0:09parallelism in many ways
  7. 0:10and now we're saying okay we've handled
  8. 0:12the c part we know how to do lmp
  9. 0:14uh openmp you know how to handle race
  10. 0:16conditions you know what deadlock is and
  11. 0:18how to grab semaphores and locks
  12. 0:20what actually happens with the with the
  13. 0:22issue of level one level two level three
  14. 0:25caches level three is maybe
  15. 0:26shared memory is shared level one level
  16. 0:28two is separate what happens when you
  17. 0:29have shared memory and caches
  18. 0:30this gets a little more complicated so
  19. 0:32let's actually look at the hardware side
  20. 0:33of this what's actually happening
  21. 0:35so this is on a chip i've got my cpus
  22. 0:38my cores each of them has
  23. 0:42a memory bus i've got other device i get
  24. 0:44main memory
  25. 0:45and in a model that is a symmetric
  26. 0:48multiprocessor
  27. 0:49so this is a memory of a multi-core
  28. 0:50machine two or more identical
  29. 0:52cores i've got a single shared coherent
  30. 0:55memory so again i have
  31. 0:56each of these guys above above this line
  32. 0:58above the memory bus
  33. 0:59are my cores and below the line is the
  34. 1:01memory that i'm sharing um and each one
  35. 1:03might have
  36. 1:04various levels of caches that are shared
  37. 1:06or not shared depending
  38. 1:08um so there's a general quest this is a
  39. 1:09general question for any computing
  40. 1:11system any multi-processor
  41. 1:12system this isn't just this particular
  42. 1:14one cpu to do this it might be how do i
  43. 1:16have
  44. 1:16two different chips and how do they
  45. 1:18communicate it might be you know worker
  46. 1:19bees doing something
  47. 1:21question one is how do they share data
  48. 1:22how do these guys share data back and
  49. 1:24forth
  50. 1:24how do they coordinate their operations
  51. 1:26and how many processes can be supported
  52. 1:27how big can we go how wide can i go in
  53. 1:29terms of the number of helpers i have
  54. 1:31so in snp or shared memory
  55. 1:34multiprocessors what we've been talking
  56. 1:35about along
  57. 1:36i have a shared address space so we
  58. 1:38mentioned that before the first slide i
  59. 1:39showed you
  60. 1:40memory is shared that zero to four to
  61. 1:42you know zero to four gibby
  62. 1:43is shared among all of them and they're
  63. 1:45going to communicate
  64. 1:46via these shared variables in memory
  65. 1:49okay from via loads and stores that's
  66. 1:50the only way for me to
  67. 1:51to work um and you have to then
  68. 1:55synchronize that some way and we talked
  69. 1:56about hardware synchronization or locks
  70. 1:58to be able to make sure
  71. 1:59you're not having two people think that
  72. 2:00they have right control over that and
  73. 2:02that's great and all
  74. 2:03multi-core computers now are smp or
  75. 2:05shared memory multi-processors
  76. 2:08caches so caches becomes the really
  77. 2:11interesting thing
  78. 2:12um you knew that memory was a
  79. 2:14performance bottleneck that's not new
  80. 2:16information memory was
  81. 2:17going to sacramento even with one
  82. 2:18processor that was a trouble
  83. 2:20we brought in caches to to reduce that
  84. 2:22pain uh
  85. 2:23and to reduce the bandwidth demands on
  86. 2:25main memory so that now which is nice
  87. 2:26imagine
  88. 2:27you know that's the reason we had even
  89. 2:29back in the olden days we had two
  90. 2:30uh data instruction caches so that both
  91. 2:32of them could be hit at once
  92. 2:33and that's great it's almost like you're
  93. 2:35paralyzing using this you're being able
  94. 2:37to
  95. 2:38to have um them both be using the same
  96. 2:40clock cycle we've learned that right the
  97. 2:41pc can be
  98. 2:42you know you're looking up some new
  99. 2:44value the rising edge of the clock and
  100. 2:46you go grab a new
  101. 2:47what's the new instruction i don't know
  102. 2:49where is it i'll go point in memory well
  103. 2:51do i have to go to second no cause it's
  104. 2:53in my instruction level one cache great
  105. 2:54boom
  106. 2:55at the same time i might be asking i'm
  107. 2:57doing a load word i'm reading from data
  108. 2:58cache at the same exact instruction
  109. 3:00i'm doing those the same day great i
  110. 3:02might be you know the fifth fourth stage
  111. 3:04of doing a load word and i'm at the
  112. 3:06first stage or the next guy and they're
  113. 3:07the same line
  114. 3:08and they both can happen the same time
  115. 3:09so i love so we already know that you
  116. 3:11can kind of split sometimes split caches
  117. 3:12be able to
  118. 3:13work at the same time without having a
  119. 3:15resource bottleneck
  120. 3:17we've seen that already so
  121. 3:20uh and and uh
  122. 3:24each core has a private cache we've
  123. 3:25known that already we know that already
  124. 3:27um only cache misses only when you have
  125. 3:31a miss do you need to go past the cache
  126. 3:33to go into the shared space you're in
  127. 3:34your local stuff i mean my local
  128. 3:36we know that at least l1 l2 i'm here and
  129. 3:38only when i missed that do i need to go
  130. 3:39and maybe there's some more complex
  131. 3:40complexity if i shared out three
  132. 3:42but i only really need to forget at
  133. 3:44three for a moment how i need to worry
  134. 3:45about memory once i miss my
  135. 3:47for now we're just gonna say there's one
  136. 3:48local cache and then there's one memory
  137. 3:50that's alright let's make it easy i
  138. 3:51still have the same problems that come
  139. 3:52up with l1 and l2 and l3 but let's talk
  140. 3:54about just having one cache
  141. 3:55and memory and only cache only when i
  142. 3:57miss my cache
  143. 3:58well i need to go to memory and that's
  144. 3:59the shared space okay
  145. 4:01so think about this everybody here's the
  146. 4:03simplest simplest case
  147. 4:04everybody has their own cache again
  148. 4:06there's maybe the l1 and l2 here and i
  149. 4:08don't worry about
  150. 4:09l3 for now and they all have to go
  151. 4:10through the interconnect memory to get
  152. 4:12to the shared space of memory all right
  153. 4:13so that's all the same
  154. 4:16so now let's actually do a simple case
  155. 4:18processor one and two just reads for now
  156. 4:19easy easy easy
  157. 4:21processor 1 and 2 read memory location
  158. 4:22100 and they want to get the value 20.
  159. 4:25so let's actually animate this this is
  160. 4:26kind of fun
  161. 4:27so here's processor one address is a
  162. 4:30hundred
  163. 4:31sorry a thousand memory value a thousand
  164. 4:33so there we go
  165. 4:34is it in my cache no it's not so i go to
  166. 4:37memory i grab it what's its value
  167. 4:3820 i come on back and i go back and i
  168. 4:41put it in my cache for future
  169. 4:43reference and i remember that you know
  170. 4:45at a thousand i've got a 20 and we're
  171. 4:46good
  172. 4:47processor two does the same thing you
  173. 4:50can already see that there might be an
  174. 4:51optimization where a processor
  175. 4:52could look in process or one's cache but
  176. 4:53if we don't have that place they only
  177. 4:55have to go through remember that's the
  178. 4:55only way to
  179. 4:56they don't that's the only way to
  180. 4:57communicate for now so the same thing
  181. 4:59happens
  182. 4:59processor 1 says i want to read value a
  183. 5:02memory of a thousand i want to get i
  184. 5:03think it's a 20. but i don't know i want
  185. 5:05to get the value
  186. 5:06so again check the cache is not there go
  187. 5:08to memory grab it and come on back
  188. 5:11and 20 is there and i remember it's a
  189. 5:12thousand so good so far so good
  190. 5:14everything's fine and each one
  191. 5:15has a copy right so this is the issue
  192. 5:18each one how has a copy
  193. 5:20there wasn't an issue before i mean
  194. 5:22before it was an issue of you know do we
  195. 5:23do a right back or right through in
  196. 5:24terms of keeping
  197. 5:25our cache and memory synchronous we
  198. 5:28would allow it to be asynchronous so
  199. 5:29that allowed to get kind of
  200. 5:30uh stale right we have a dirty bit for
  201. 5:33that but now you've got an issue that
  202. 5:35both of them have a
  203. 5:36copy which is fine as long as the same
  204. 5:37value what happens if they change
  205. 5:39okay so now what happens if processor
  206. 5:43zero comes around
  207. 5:44and writes memory 1000 with a 40. here
  208. 5:47we go let's try it
  209. 5:48so got a thousand is it in the cache
  210. 5:53nope not there so write it in the cache
  211. 5:55and we're going to make sure
  212. 5:57in this particular situation we're going
  213. 5:58to have a right through so we're going
  214. 6:00to actually have to write it all the way
  215. 6:01through we're not going to write back
  216. 6:02where we keep them stale and
  217. 6:04have it be dirty i'm going to say right
  218. 6:06through so now there's my 1040. that's
  219. 6:08great
  220. 6:08note that processor 1 and 2 have their
  221. 6:1020 which is the wrong value
  222. 6:12and then yoink i write it through you
  223. 6:15told me to write through i'm processor
  224. 6:160. what do you want from me over here
  225. 6:18i write it through can you see a problem
  226. 6:21anybody see a problem wait let's just
  227. 6:23pause here and everyone should see a
  228. 6:24problem here
  229. 6:26process zero says hey the value at a
  230. 6:28thousand is 40.
  231. 6:29processor one and two say nope it's 20.
  232. 6:32that's a really big deal
  233. 6:34so that's in this lecture i'm setting
  234. 6:36you up for
  235. 6:37i love the cliffhanger it's kind of a
  236. 6:39cliffhanger approach like you know
  237. 6:40mandalorian what happens to baby yoda
  238. 6:42the night where's there the cliffhanger
  239. 6:43here what happens how do we resolve this
  240. 6:45that's the next lecture we'll see you
  241. 6:47there

About this transcript

This page contains the full transcript of [CS61C FA20] Lecture 35.2 - Thread-Level Parallelism III: Shared Memory and Caches by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 1,547 words across 241 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.