[CS61C FA20] Lecture 33.1 - Thread-Level Parallelism I: Parallel Computer Architectures — Transcript
Full transcript
- 0:03hey
- 0:04sorry just juggling five balls good to
- 0:07see you
- 0:08today's topic welcome to 61c's topic on
- 0:11thread level parallelism part one
- 0:15great to see you as always let's jump
- 0:16right in we're going to start with a
- 0:18conversation about parallel computer
- 0:20architectures
- 0:21so big picture step back 10 000 foot
- 0:24view
- 0:25the goal of 61c is teach you how to
- 0:28program a computer really well
- 0:29so that you can increase performance
- 0:31there are many ways of doing that if you
- 0:32were building a system if you're a
- 0:34computer engineer
- 0:35building a system the first thing you do
- 0:36is change the heart rate
- 0:38change the speed of your mouse your
- 0:40heart rate is 300 to 800
- 0:42beats per beats per minute that's a lot
- 0:45so think about just turning that crank
- 0:47up turning that crank up
- 0:49gigahertz machine two gigahertz three
- 0:51four five
- 0:52that's incredible but we've kind of
- 0:54stopped we're going to talk a little bit
- 0:55later about why we can't go much past
- 0:57that it's really about power dissipation
- 0:59we can't keep these chips cool so
- 1:01you've kind of maxed out so we increase
- 1:02our clock rate as much as we can
- 1:04but power issues how much total power
- 1:06how to dissipate the power the power
- 1:07density all those things are relative to
- 1:09that
- 1:09relevant in that conversation and we
- 1:11can't go much past five
- 1:13five uh gigahertz next
- 1:16well we can also lower the cpi you know
- 1:18cycles per instruction this says lower
- 1:20sim d
- 1:21sorry this means gosimd this says single
- 1:23instruction multiple data can one
- 1:25instruction
- 1:25operate on wider data can you operate on
- 1:27four floats at once that would be
- 1:28amazing we've done that and talked about
- 1:30that in the last lecture so
- 1:31try to go wider so smd you could also
- 1:34perform multiple tasks simultaneously
- 1:36you could have multiple cpus each
- 1:38executing a different program this is
- 1:39great
- 1:40these tasks could be related they're all
- 1:42tight
- 1:43working on the same matrix and many guys
- 1:44are doing some part of the same matrix
- 1:46or
- 1:47completely unrelated um you know
- 1:50distributed web requests
- 1:51or on different computers even you can
- 1:54imagine that or or running powerpoint in
- 1:56one thing in a quicktime movie another
- 1:57thing and
- 1:58you know doing a spreadsheet calculation
- 1:59and other things so they could be
- 2:00related very tight in the same program
- 2:02or very different things
- 2:03and that's different scales so we'll
- 2:04this lecture is about
- 2:06the top level the idea of three of just
- 2:08doing something about the same program
- 2:09you're
- 2:10working on the same program the same
- 2:11data and how do you divide that up into
- 2:13a parallel space that's today's lecture
- 2:14and actually the next couple of lectures
- 2:15we'll talk about that
- 2:17and in summary you can do all three of
- 2:18them uh that's the idea so do all
- 2:20all the above high high clock frequency
- 2:23go
- 2:24cimd go wide with your vectors uh and
- 2:26multiple parallel tasks as much as
- 2:28possible
- 2:28so big picture we always show this in
- 2:30the have a new model by the way this is
- 2:31a new module so if you're behind
- 2:33jump jump on this and feel free to reset
- 2:35here um this thread level parallelism
- 2:37is very exciting so you've seen the left
- 2:39i mean let's let's even get the pen
- 2:40where's my pen let's even get the pen
- 2:42out
- 2:42and uh and talk about with this
- 2:45you've seen hardware descriptions you
- 2:47know if i've got a 32
- 2:49bit wide and gate that's 32 operations
- 2:52all happening at the same time done that
- 2:53already
- 2:54you've seen uh parallel instructions
- 2:57when we did pipelining
- 2:59now you know how you can do up to five
- 3:01it's a five stage pipeline there's five
- 3:02instructions that can all be
- 3:04chunking it through that little slice of
- 3:05time can have five instructions all
- 3:07active at the same time pretty neat
- 3:09parallel data was about cmd can we go
- 3:11wide with our vectors can you add four
- 3:13vectors in one clock cycle that's great
- 3:15the next couple lectures are up here so
- 3:17the next couple lectures this space and
- 3:18in fact in particular
- 3:20this lecture is about parallel threads
- 3:22and this next couple lectures about
- 3:24talking about thread level parallelism
- 3:25and that's the piece of this
- 3:27how could you have different cores each
- 3:28of them working on different threads
- 3:30what are all those mean
- 3:31what are architectures what what what's
- 3:34required on the architecture point of
- 3:35view to make that work
- 3:36all that's this lecture and the next
- 3:37year's lectures so here's some big
- 3:39pictures of parallel computer
- 3:40architectures
- 3:41early early days view of this i love
- 3:44this picture
- 3:44this is a picture of just trying to
- 3:46throw a couple machines together
- 3:47wire them up via ethernet and say let's
- 3:50try to have a job that can
- 3:51really this is distributed computing in
- 3:53some sense can you take a job that is
- 3:55far enough and has has enough pieces
- 3:57that can be broken off that a whole
- 3:59computer can be asked to work on that
- 4:00piece and then somehow collect that
- 4:02together
- 4:02so usually there's like a command module
- 4:04that divides it up
- 4:05all these kind of worker bees work on
- 4:07something that come back but you're
- 4:08communicating via ethernet or via a
- 4:10shared drive or something like that
- 4:12when you go full scale this is what
- 4:15happens at all the national labs
- 4:16um this is what we have in super
- 4:18computers so you have these massive
- 4:20racks massive
- 4:21server racks with an array of computers
- 4:24and
- 4:24uh we'll even talk about that a little
- 4:26later when we talk about warehouse scale
- 4:27computers
- 4:28most of those computers just fyi are
- 4:31working on
- 4:32computational science problems they're
- 4:35usually working on floating point data
- 4:36you're doing matrix multiplication
- 4:38you're doing neural network stuff you're
- 4:39you're doing uh climate simulation a lot
- 4:41of these things all of them the actual
- 4:43data that you care about the most
- 4:45is floating point data it's not kind of
- 4:47integer based data where it's
- 4:48you're thinking about ins and choose
- 4:50complement numbers you're doing doubles
- 4:52mostly almost dealing with doubles or
- 4:53even more if you need higher precision
- 4:55for that so that's what those machines
- 4:56are doing and those are
- 4:57unbelievable you go to super computer
- 4:58conferences they're those machines and
- 5:00they're talking about how to crunch on
- 5:02massive massive massive floating point
- 5:04data sets
- 5:06in your phone in your watch in your
- 5:08computer we now have
- 5:10also multi-core machines we've talked a
- 5:12little bit about that we'll get a little
- 5:13bit deeper into that how that works
- 5:15and these multi-core this is on a single
- 5:16chip there's on a single chip multi-core
- 5:18machines there's
- 5:19some graphics there's maybe a shared l3
- 5:21cache l1 and l2 are separate
- 5:24for each core um it's incredible and
- 5:26we've kind of shown a picture of that
- 5:27before and that's another architecture
- 5:29all these are
- 5:30ways to think about parallelism at
- 5:32different scales actually different
- 5:33scales
- 5:34so let's dig deeper into that bottom
- 5:36right picture
- 5:37what's happening if you have a cpu with
- 5:39two cores so again
- 5:41single dis single die two cores in that
- 5:43let's look so very simple
- 5:45so what does that mean two cores does
- 5:46that mean you're doubling memory so now
- 5:48you know my 32 gibby bytes of memory or
- 5:50maybe more you know computers nowadays
- 5:52you've seen virtual memory
- 5:53you know the idea that there's a
- 5:54distinction between it was it was all a
- 5:56lie between
- 5:57how much actual memory you have and the
- 5:59address space
- 6:00i can have a 32-bit address space that's
- 6:02four gigabytes of allocatable memory
- 6:05how do i have 10 programs all of which
- 6:07are
- 6:08controlling all four gigabytes well
- 6:10that's the idea of virtual memory each
- 6:11of them has a virtual address space
- 6:12that's that big
- 6:13which that means now now i can have an
- 6:14actual computer that has not exactly 4
- 6:17gb bytes if you didn't have that
- 6:18abstraction you'd have to have a machine
- 6:19that has exactly four gigabytes and only
- 6:21one program can be running at a time
- 6:22or you could be switching and you know
- 6:24time slicing it
- 6:26but now with virtual memory you've seen
- 6:27before we can have many machines each of
- 6:30them thinks they have the whole 32-bit
- 6:31address space or even a 64-bit address
- 6:33space and the amount of actual memory
- 6:34you have
- 6:35can be more than 4 gb or less than 4 gb
- 6:38uh
- 6:38certainly not 64. if you're living in 64
- 6:40addresses but mostly that risk 532
- 6:42means it's a it's a 32-bit virtual
- 6:44address space how much actual memory you
- 6:46have doesn't matter i mean you could
- 6:48have a huge amount
- 6:49they not sell 1.5 gibby bytes of ram on
- 6:52the new mac pros incredible
- 6:53and you can have a machine that has very
- 6:55little and still you can still run
- 6:5632-bit address-based
- 6:57programs all right so we talked about
- 7:00one die
- 7:01with two cores on it does that mean
- 7:03you're somewhat doubling memory you're
- 7:04not because the core doesn't memory on
- 7:05it anyway the core is over here memory's
- 7:07over there
- 7:08so what does it mean to do this it means
- 7:09you're separating two different
- 7:11processor cores
- 7:12each of which each of which has its own
- 7:15control
- 7:16its own data path its own pc
- 7:20its own set of registers its own aou
- 7:22they are really two full
- 7:23of quote unquote computers on one die
- 7:27those two cores are on one die and they
- 7:29really
- 7:30are uh fully autonomous that can be
- 7:32completely running
- 7:33separate things i'm doing a load you're
- 7:35doing a store i'm doing an ad you're
- 7:36doing a subtract
- 7:37completely separate different registers
- 7:39running two programs at the same time
- 7:41we love this that's a two core model
- 7:44what is this shared part of it though
- 7:45well the shared part of it is i've got
- 7:47one memory
- 7:48block and io is working within that and
- 7:51i've got my io interface here we've
- 7:52talked about
- 7:53io a little bit in previous lectures so
- 7:56there's has to be some communication
- 7:58where each processor is making load and
- 8:00store requests
- 8:02to memory and now you have your
- 8:03conversation well what happens ooh
- 8:05when they're both writing and reading to
- 8:06the same spot or writing and reading how
- 8:08do you synchronize that there's some
- 8:09conversations lower down
- 8:11in that but at the big picture this is a
- 8:1310 mile view up
- 8:14you have two different course everything
- 8:16you know about all we
- 8:17you know everything up to this lecture
- 8:19was just this picture
- 8:20this is the picture up to now and now
- 8:22we're saying well i want to have two
- 8:23cores on the same die
- 8:25well let's have a second guy and this
- 8:26interface here's the extra here's the
- 8:28extra interface
- 8:29same the same connection you had memory
- 8:31loads and stores that's the way you can
- 8:32make this work
- 8:33there is a little bit more details with
- 8:35the caches but from this 10 mile view up
- 8:37you have two cores
- 8:38two distinct control data paths pcs
- 8:40registers and leus all working with one
- 8:42shared memory okay this is called a
- 8:43shared memory model
- 8:45okay so now let's kind of dig a little
- 8:47deeper this is the last slide on this
- 8:48first
- 8:49first lecture each processor is
- 8:52executing its own instruction
- 8:53set it's chunking through two different
- 8:55programs all running in parallel
- 8:57separate resources data path pc
- 8:59registers alu you said that already
- 9:01highest level caches are also going to
- 9:03be separate you want to be able to not
- 9:04have to go to sacramento
- 9:05as you're having you know you're you're
- 9:07as you're making memory requests and
- 9:09it's not there you want to be able to
- 9:10save it and have your your net
- 9:12is there your l1 and l2 net is there
- 9:14saving you from going to sacramento
- 9:15because that's that one shared spot
- 9:17you do have some shared resources we
- 9:19mentioned memory is a shared resource
- 9:21it's expensive i mean the 1.5 gibby
- 9:23bytes of of dram is very expensive so
- 9:25you're not going to duplicate that per
- 9:27core
- 9:27you got many many cores in the newest
- 9:28machines that's certainly not duplicated
- 9:30you will duplicate l1 and l2 typically
- 9:33and l3 is funny l3 is usually really big
- 9:36uh and so usually in the mega megabytes
- 9:39so
- 9:39you're probably gonna have a shared l3
- 9:42although not required uh it's often not
- 9:44it's often on the same silicon chip
- 9:45doesn't have to be all that's still in
- 9:47the model of a multi-processor execution
- 9:48model
- 9:49doesn't have to be there you could
- 9:50decide well there's no l3 at all that's
- 9:52still fine it's still in the category of
- 9:53a multi-processor system
- 9:56here's some new words you might not have
- 9:57heard you know our microprocessor is
- 9:59what all the cpus have been talking
- 10:00about before
- 10:01a multi-processor microprocessor that
- 10:03means
- 10:04two or more cores in one microprocessor
- 10:07um you see two more processors
- 10:09processors could be different by the way
- 10:11my first computer i bought my first big
- 10:12computer i bought when i got hired
- 10:14was uh uh is an early mac
- 10:18stand desktop mac uh and it had
- 10:21two different dies if you looked at two
- 10:23different dies two different heat sinks
- 10:24two different dies and each one of them
- 10:26had two different cores on it that's
- 10:27pretty neat so a multi-processor
- 10:30microprocessor both means
- 10:31multiple dies working on some shared
- 10:34memory area or
- 10:35a single die with multiple cores so the
- 10:37second idea and the most recent idea
- 10:39now is let's forget it's two weeks too
- 10:41weird and too hard people haven't done
- 10:42this
- 10:42in recent years that have multiple dies
- 10:45multiple
- 10:46dies each of them has a huge heatsink
- 10:47very expensive you now just have one you
- 10:49should just have one
- 10:50die one cpu in some sense
- 10:54but it's multi-core that one die is a
- 10:56single cpu but it has
- 10:58and we're going to say i say the word
- 11:00cpu is a little fuzzy now i don't want
- 11:01to get into that
- 11:02but in the past they would say two cpus
- 11:04with two cores on each now they actually
- 11:05call one
- 11:06die multiple cpus we'll we'll talk about
- 11:09this a little later on
- 11:10because they've fuzzed they've fuzzed
- 11:12around the naming of that the cpu used
- 11:14to be
- 11:14a die so now they have a single cpu
- 11:18with four cores on them that's what they
- 11:19call it okay and so we'll talk about
- 11:22logical versus physical cpus we'll talk
- 11:24about that at
- 11:25later lectures but that's what we've got
- 11:26and the idea is here i go if i've got uh
- 11:28if i've got
- 11:29this last slide we had two cores this
- 11:31means two full
- 11:32instruction streams we'll have some word
- 11:34for that in later slides instruction
- 11:36seems operated simultaneously that's it
- 11:37pretty exciting
- 11:38let's actually get down to more details
- 11:40of how that works and some more
- 11:41ideas of how does this work in software
- 11:43in the next lecture we'll see you there
About this transcript
This page contains the full transcript of [CS61C FA20] Lecture 33.1 - Thread-Level Parallelism I: Parallel Computer Architectures by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 2,494 words across 417 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.