[CS61C FA20] Lecture 35.2 - Thread-Level Parallelism III: Shared Memory and Caches — Transcript
Full transcript
- 0:00and welcome back
- 0:03now let's actually think about what
- 0:05happens
- 0:06under the hood so we've talked about the
- 0:08software level so we're talking about
- 0:09parallelism in many ways
- 0:10and now we're saying okay we've handled
- 0:12the c part we know how to do lmp
- 0:14uh openmp you know how to handle race
- 0:16conditions you know what deadlock is and
- 0:18how to grab semaphores and locks
- 0:20what actually happens with the with the
- 0:22issue of level one level two level three
- 0:25caches level three is maybe
- 0:26shared memory is shared level one level
- 0:28two is separate what happens when you
- 0:29have shared memory and caches
- 0:30this gets a little more complicated so
- 0:32let's actually look at the hardware side
- 0:33of this what's actually happening
- 0:35so this is on a chip i've got my cpus
- 0:38my cores each of them has
- 0:42a memory bus i've got other device i get
- 0:44main memory
- 0:45and in a model that is a symmetric
- 0:48multiprocessor
- 0:49so this is a memory of a multi-core
- 0:50machine two or more identical
- 0:52cores i've got a single shared coherent
- 0:55memory so again i have
- 0:56each of these guys above above this line
- 0:58above the memory bus
- 0:59are my cores and below the line is the
- 1:01memory that i'm sharing um and each one
- 1:03might have
- 1:04various levels of caches that are shared
- 1:06or not shared depending
- 1:08um so there's a general quest this is a
- 1:09general question for any computing
- 1:11system any multi-processor
- 1:12system this isn't just this particular
- 1:14one cpu to do this it might be how do i
- 1:16have
- 1:16two different chips and how do they
- 1:18communicate it might be you know worker
- 1:19bees doing something
- 1:21question one is how do they share data
- 1:22how do these guys share data back and
- 1:24forth
- 1:24how do they coordinate their operations
- 1:26and how many processes can be supported
- 1:27how big can we go how wide can i go in
- 1:29terms of the number of helpers i have
- 1:31so in snp or shared memory
- 1:34multiprocessors what we've been talking
- 1:35about along
- 1:36i have a shared address space so we
- 1:38mentioned that before the first slide i
- 1:39showed you
- 1:40memory is shared that zero to four to
- 1:42you know zero to four gibby
- 1:43is shared among all of them and they're
- 1:45going to communicate
- 1:46via these shared variables in memory
- 1:49okay from via loads and stores that's
- 1:50the only way for me to
- 1:51to work um and you have to then
- 1:55synchronize that some way and we talked
- 1:56about hardware synchronization or locks
- 1:58to be able to make sure
- 1:59you're not having two people think that
- 2:00they have right control over that and
- 2:02that's great and all
- 2:03multi-core computers now are smp or
- 2:05shared memory multi-processors
- 2:08caches so caches becomes the really
- 2:11interesting thing
- 2:12um you knew that memory was a
- 2:14performance bottleneck that's not new
- 2:16information memory was
- 2:17going to sacramento even with one
- 2:18processor that was a trouble
- 2:20we brought in caches to to reduce that
- 2:22pain uh
- 2:23and to reduce the bandwidth demands on
- 2:25main memory so that now which is nice
- 2:26imagine
- 2:27you know that's the reason we had even
- 2:29back in the olden days we had two
- 2:30uh data instruction caches so that both
- 2:32of them could be hit at once
- 2:33and that's great it's almost like you're
- 2:35paralyzing using this you're being able
- 2:37to
- 2:38to have um them both be using the same
- 2:40clock cycle we've learned that right the
- 2:41pc can be
- 2:42you know you're looking up some new
- 2:44value the rising edge of the clock and
- 2:46you go grab a new
- 2:47what's the new instruction i don't know
- 2:49where is it i'll go point in memory well
- 2:51do i have to go to second no cause it's
- 2:53in my instruction level one cache great
- 2:54boom
- 2:55at the same time i might be asking i'm
- 2:57doing a load word i'm reading from data
- 2:58cache at the same exact instruction
- 3:00i'm doing those the same day great i
- 3:02might be you know the fifth fourth stage
- 3:04of doing a load word and i'm at the
- 3:06first stage or the next guy and they're
- 3:07the same line
- 3:08and they both can happen the same time
- 3:09so i love so we already know that you
- 3:11can kind of split sometimes split caches
- 3:12be able to
- 3:13work at the same time without having a
- 3:15resource bottleneck
- 3:17we've seen that already so
- 3:20uh and and uh
- 3:24each core has a private cache we've
- 3:25known that already we know that already
- 3:27um only cache misses only when you have
- 3:31a miss do you need to go past the cache
- 3:33to go into the shared space you're in
- 3:34your local stuff i mean my local
- 3:36we know that at least l1 l2 i'm here and
- 3:38only when i missed that do i need to go
- 3:39and maybe there's some more complex
- 3:40complexity if i shared out three
- 3:42but i only really need to forget at
- 3:44three for a moment how i need to worry
- 3:45about memory once i miss my
- 3:47for now we're just gonna say there's one
- 3:48local cache and then there's one memory
- 3:50that's alright let's make it easy i
- 3:51still have the same problems that come
- 3:52up with l1 and l2 and l3 but let's talk
- 3:54about just having one cache
- 3:55and memory and only cache only when i
- 3:57miss my cache
- 3:58well i need to go to memory and that's
- 3:59the shared space okay
- 4:01so think about this everybody here's the
- 4:03simplest simplest case
- 4:04everybody has their own cache again
- 4:06there's maybe the l1 and l2 here and i
- 4:08don't worry about
- 4:09l3 for now and they all have to go
- 4:10through the interconnect memory to get
- 4:12to the shared space of memory all right
- 4:13so that's all the same
- 4:16so now let's actually do a simple case
- 4:18processor one and two just reads for now
- 4:19easy easy easy
- 4:21processor 1 and 2 read memory location
- 4:22100 and they want to get the value 20.
- 4:25so let's actually animate this this is
- 4:26kind of fun
- 4:27so here's processor one address is a
- 4:30hundred
- 4:31sorry a thousand memory value a thousand
- 4:33so there we go
- 4:34is it in my cache no it's not so i go to
- 4:37memory i grab it what's its value
- 4:3820 i come on back and i go back and i
- 4:41put it in my cache for future
- 4:43reference and i remember that you know
- 4:45at a thousand i've got a 20 and we're
- 4:46good
- 4:47processor two does the same thing you
- 4:50can already see that there might be an
- 4:51optimization where a processor
- 4:52could look in process or one's cache but
- 4:53if we don't have that place they only
- 4:55have to go through remember that's the
- 4:55only way to
- 4:56they don't that's the only way to
- 4:57communicate for now so the same thing
- 4:59happens
- 4:59processor 1 says i want to read value a
- 5:02memory of a thousand i want to get i
- 5:03think it's a 20. but i don't know i want
- 5:05to get the value
- 5:06so again check the cache is not there go
- 5:08to memory grab it and come on back
- 5:11and 20 is there and i remember it's a
- 5:12thousand so good so far so good
- 5:14everything's fine and each one
- 5:15has a copy right so this is the issue
- 5:18each one how has a copy
- 5:20there wasn't an issue before i mean
- 5:22before it was an issue of you know do we
- 5:23do a right back or right through in
- 5:24terms of keeping
- 5:25our cache and memory synchronous we
- 5:28would allow it to be asynchronous so
- 5:29that allowed to get kind of
- 5:30uh stale right we have a dirty bit for
- 5:33that but now you've got an issue that
- 5:35both of them have a
- 5:36copy which is fine as long as the same
- 5:37value what happens if they change
- 5:39okay so now what happens if processor
- 5:43zero comes around
- 5:44and writes memory 1000 with a 40. here
- 5:47we go let's try it
- 5:48so got a thousand is it in the cache
- 5:53nope not there so write it in the cache
- 5:55and we're going to make sure
- 5:57in this particular situation we're going
- 5:58to have a right through so we're going
- 6:00to actually have to write it all the way
- 6:01through we're not going to write back
- 6:02where we keep them stale and
- 6:04have it be dirty i'm going to say right
- 6:06through so now there's my 1040. that's
- 6:08great
- 6:08note that processor 1 and 2 have their
- 6:1020 which is the wrong value
- 6:12and then yoink i write it through you
- 6:15told me to write through i'm processor
- 6:160. what do you want from me over here
- 6:18i write it through can you see a problem
- 6:21anybody see a problem wait let's just
- 6:23pause here and everyone should see a
- 6:24problem here
- 6:26process zero says hey the value at a
- 6:28thousand is 40.
- 6:29processor one and two say nope it's 20.
- 6:32that's a really big deal
- 6:34so that's in this lecture i'm setting
- 6:36you up for
- 6:37i love the cliffhanger it's kind of a
- 6:39cliffhanger approach like you know
- 6:40mandalorian what happens to baby yoda
- 6:42the night where's there the cliffhanger
- 6:43here what happens how do we resolve this
- 6:45that's the next lecture we'll see you
- 6:47there
About this transcript
This page contains the full transcript of [CS61C FA20] Lecture 35.2 - Thread-Level Parallelism III: Shared Memory and Caches by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 1,547 words across 241 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.