[CS61C FA20] Lecture 33.2 - Thread-Level Parallelism I: Multicore — Transcript
Full transcript
- 0:00and welcome back now let's talk about
- 0:03dig dig a little deeper into what we
- 0:04mean by having a multi-core computer
- 0:06with a little bit of a history lesson
- 0:08so this is a wonderful curve i really
- 0:10love this curve it's not my data but
- 0:12it's a wonderful
- 0:13visualization of the evolution to
- 0:16multi-core and why
- 0:17why we got to where we are now so let's
- 0:20go back in time
- 0:21um we're in the early 70s we're looking
- 0:25at
- 0:25the total number of transistors and
- 0:27thousands what i love about this view on
- 0:28the left
- 0:29is that it's unitless the units are
- 0:31attached to each of the curves
- 0:33so as i look to moore's law moore's law
- 0:36says how many ics how many integrated
- 0:38circuits
- 0:38how many transistors are on an ic so how
- 0:40many transistors are on an integrated
- 0:42circuit
- 0:42and we're looking at this number this
- 0:45number was doubling every 18 months or
- 0:47so
- 0:48which is a nice line here and then at
- 0:50some point
- 0:51we actually were doubling so doubling
- 0:53every two years sorry doubling every two
- 0:54years in the early stages
- 0:56and at some point it kind of had a small
- 0:58bend and we started doubling every 18
- 1:00months
- 1:01and we're doing amazing so we're just
- 1:03packing more and more transistors on
- 1:06the the feature size is getting smaller
- 1:08it's getting down to
- 1:09some some number some kind of nanometers
- 1:1165 50 30
- 1:1320. you know it's getting smaller and
- 1:14smaller which means that the um between
- 1:17two lines it's getting shorter and
- 1:18smaller and smaller between the distance
- 1:20between two signal lines on on the chip
- 1:22amazing what they're doing great
- 1:24incredible engineering this is an
- 1:25incredible amount of engineering going
- 1:26on here
- 1:28if we continue to look here we're also
- 1:31seeing
- 1:32the frequency is going up remember i
- 1:33mentioned just increase that heart rate
- 1:35increase that heart rate do this and in
- 1:36fact it even took a bend
- 1:38where we even went steeper there um now
- 1:41we're also seeing the power go up so
- 1:43this is a story
- 1:44this is a story in which these these
- 1:45curves are all related in interesting
- 1:46ways
- 1:47and the power was going up it was still
- 1:49small we didn't care about it
- 1:50you know in the 90s it was fine and by
- 1:52the way this was a great time
- 1:54for well okay great time for computer
- 1:56manufacturers because they could
- 1:58argue that you need to buy this year's
- 2:00computer um
- 2:01because it's twice as fast and every
- 2:03year it was every year you know every
- 2:05couple of years you needed to upgrade
- 2:07you you just really couldn't
- 2:08um live with the computer that you had
- 2:10bought two or three years ago because
- 2:12boy the the potential to have a computer
- 2:15that was 200 megahertz
- 2:16and now it's 500 just based on only the
- 2:19the clock rate
- 2:20um you know 200 meg 250 megahertz okay
- 2:24500 megahertz 800 megahertz
- 2:26a gigahertz oh man we crossed the
- 2:27gigahertz line two 1.5 2 gigahertz wow i
- 2:30got
- 2:31love this so we meant for a couple of a
- 2:33lot of things in the computing
- 2:34engineering ecosystem
- 2:36it meant that you're a computer
- 2:37manufacturer you were doing great
- 2:39because you you could just continue to
- 2:40push
- 2:40machines out and there were different
- 2:42specs and people were comparing and then
- 2:43all the gamers are there
- 2:44the computer scientists were on this
- 2:46side and the computational scientists
- 2:47who were using it for floating were they
- 2:49all the people were benefiting from
- 2:50this this incredible this all this with
- 2:52all these curves were benefiting
- 2:53everybody using these systems
- 2:57what's interesting a little bit is that
- 2:59we didn't care about power that much but
- 3:01people began to have more and more as
- 3:03this number
- 3:04started to get up here look at this
- 3:06we're now in the hunt i mean look at
- 3:07this this is you know
- 3:08100 watts that's an incandes
- 3:10incandescent light bulb
- 3:11so people used to say these systems are
- 3:13getting too hot we can't cool them
- 3:16that was the issue we couldn't cool
- 3:17these systems down you had very arcane
- 3:19heatsinks
- 3:20with a radiator fan and a huge i mean a
- 3:22radiator radiator system here with with
- 3:25a lot of
- 3:25a lot of spokes and then a fan tried to
- 3:28blow and get that heat off of it
- 3:30and so you'd have you'd heat up a room
- 3:31if you if you were on the you know
- 3:33exhaust fan side of
- 3:34one of some of these computers you would
- 3:35be really hot be really quite warm
- 3:37um and they started having liquid
- 3:40cooling because they could
- 3:41even just air off of a radiator system
- 3:43couldn't cool
- 3:44off the heat sink off of a cpu so you
- 3:46had water coming in and then you had oil
- 3:48come and very
- 3:50intricate engineering to get fluid to
- 3:52get in there and touch it as much as you
- 3:53can get warm and then
- 3:54get it out here and dissipate it and
- 3:56then cool it down and come back in
- 3:57incredible engineering to try to just
- 3:59cool these things down
- 4:01at some point they said we're not able
- 4:03to cool it anymore
- 4:04um we have to do something about this we
- 4:07can keep turning the clock frequency up
- 4:08and by the way there are other
- 4:09architectural improvements as well i
- 4:11mentioned before
- 4:11you know we're going cmd and going wide
- 4:13or we're doing we're seeing a lot of
- 4:14performance
- 4:15but if you look at the frequency right
- 4:18around
- 4:18a special thing happened a sea change in
- 4:21computer
- 4:22architecture happened around 2005 look
- 4:25at this number
- 4:25everything changes around that look at
- 4:27that number
- 4:29you saw the clock frequency start to
- 4:32start to level off
- 4:34well that's an issue the clock frequency
- 4:36level's off
- 4:37now you've got the power levels off so
- 4:41you wanted to
- 4:41get to that you wanted to level off the
- 4:43power and you did by turning the clock
- 4:44down
- 4:45now you're just kind of doing the same
- 4:46thing as you did before in the clock
- 4:47that but now all of a sudden what's
- 4:49being hit
- 4:50by the way you were still being able to
- 4:51continue moore's law you were still
- 4:53having
- 4:54more and more transistors being able to
- 4:56be you know squeezed into one cpu
- 5:00but you were seeing the sequential app
- 5:02performance all the programmers
- 5:04who back here were lazy all see this
- 5:06curve
- 5:07this curve says you could be lazy and
- 5:09write a bad algorithm
- 5:11not optimize it and do nothing do
- 5:13nothing in terms of trying to squeeze
- 5:14more performance out of your software
- 5:16and guess what the performance of just
- 5:18going up every year so lazy programmers
- 5:21they're hard working programs everywhere
- 5:22but people you didn't have to work worry
- 5:23so much as a programmer as a developer
- 5:26on how to squeeze the last strip of
- 5:27performance out because you were getting
- 5:29it for free you were getting this for
- 5:30free
- 5:31and all of a sudden 2005 hits and you
- 5:34stopped getting it for free
- 5:37so this is an issue you know where now
- 5:39we're going to the cloud now we're
- 5:40having larger
- 5:41larger larger data sets big data became
- 5:43a big deal
- 5:44but you couldn't handle it you had no
- 5:47way to work with that big data you
- 5:48weren't going to get any faster if the
- 5:50data size if the data set size you know
- 5:51your pixar and you have
- 5:53more and more refined models all of a
- 5:55sudden you can't
- 5:56your computers aren't processing them
- 5:58any faster it's
- 5:59flat so enter so 2005 was a sea change
- 6:03and enter the idea that you would have
- 6:06more cores this is again a history
- 6:08lesson so in 2005
- 6:11with i believe the intel core 2 duo was
- 6:13the first that i remember
- 6:14was a production level cpu that had more
- 6:16than one core there were architectures
- 6:18you had two different dyes
- 6:20before them but never two cores in one
- 6:22die and so this number of cores
- 6:24started to go up and it was amazing what
- 6:26that number is
- 6:27and then i'm just starting to grow as
- 6:29well right this is an exponential curve
- 6:30a linear
- 6:30a linear line on a on a log a linear log
- 6:34plot as an exponential curve so
- 6:35all these are straight lines but they're
- 6:36exponentials so that started to go up
- 6:40and double something of some number of
- 6:41years and then all of a sudden you saw
- 6:43this is the key
- 6:44the parallel app performance continued
- 6:46to go up and so it said to all software
- 6:48developers you better learn how to
- 6:49program in parallel
- 6:50you better learn how to work with all
- 6:51those cores effectively because you're
- 6:53if you want to live on the yellow line
- 6:54the sequential app performance you're
- 6:56flat you're done
- 6:57you want to continue to be able to
- 6:58satisfy your hungry customers
- 7:00hungry for data hungry for processing
- 7:02power you better learn how to program
- 7:03these core as well and this is what the
- 7:04goal of these set of lectures are
- 7:07so let's get that's a sample i'm just
- 7:09literally sampling
- 7:10i you know we live here in the bay area
- 7:12let's just let's go down the road a
- 7:13little bit to cupertino
- 7:14this was one of their slides when they
- 7:16released the apple a14 chip this is on
- 7:18their phones
- 7:19what are they saying what are the
- 7:20highlights of the a14 ship
- 7:23six core cpu by the way this is not a
- 7:25phone this is on a phone okay this is
- 7:26not the one that's the
- 7:27latest and greatest on a phone six cores
- 7:30and when in fact not not just six cores
- 7:32they actually categorize some number of
- 7:33cores that are
- 7:34i think they call them firestorm and ice
- 7:35storm like high performance cores and
- 7:37then look
- 7:38like lower performance cores cooler they
- 7:40call it fire and ice
- 7:42cooler of course not as fast but you
- 7:44have four of them so
- 7:45it's interesting to be able to
- 7:46distinguish between when you really need
- 7:48to crank it out put them all on
- 7:50but if you're just kind of doing normal
- 7:51things maybe having more cores but
- 7:53really
- 7:53more energy efficient cores is better so
- 7:55maybe change the clock rate possibly
- 7:57thinking about how they do
- 7:58that so that's the first thing they have
- 7:59different categorizations for their six
- 8:00scores six cores on a phone
- 8:01six cores and a phone nobody i mean 2005
- 8:05no computer had multiple cores maybe you
- 8:07were lucky if you had some special you
- 8:09know
- 8:09high-end thing but most computers i just
- 8:11had one computer one cpu
- 8:13one core and now we're talking about
- 8:15phones having six cores okay it's a new
- 8:16world
- 8:17i've got a four core gpu that's
- 8:19incredible i've got a 16 core neural
- 8:22engine for for neural networking work
- 8:24this is unbelievable and i also have
- 8:25another so this is on one die this is a
- 8:27chip
- 8:28the chip is in the middle folks this is
- 8:30the chip this you buy this
- 8:32a14 chip and you get all these pieces
- 8:34inside of it okay 11 billion transitions
- 8:36an incredible number and you got an
- 8:37image processor as well so a lot of
- 8:39you're categorizing it not just well i'm
- 8:41just going to do the normal cpu be
- 8:42boring no i'm kind of here's the gpu
- 8:44over here there's a neural engine here's
- 8:45an image processing engine it's
- 8:46incredible what they're putting together
- 8:48so here's the execution model here's as
- 8:50we think about this
- 8:52okay now that's the world you're in a
- 8:53multi-core world what is shared and what
- 8:55isn't shared
- 8:56so each core has as the picture we
- 8:59showed in the last lecture each core has
- 9:00access to the entire memory okay so
- 9:02every core has
- 9:03can connect to that entire memory there
- 9:04that's fine
- 9:06you got to worry about keeping the
- 9:07caches consistent in fact cache
- 9:09coherency
- 9:10l1 l2 l3 making sure that's all clean
- 9:13and coherent coherent means just all on
- 9:15the same page in some sense page is the
- 9:17wrong word because of virtual memory but
- 9:18the idea
- 9:18all they're all kind of concluded
- 9:20together and not inconsistent
- 9:22that's an important thing and we're
- 9:23going to talk about that about about
- 9:25three lectures from now
- 9:26um three full lectures from now um not
- 9:29three sub mod pieces of one lecture
- 9:32um so the advantage is you can actually
- 9:34now communicate i can have these chords
- 9:36that they needed to communicate via
- 9:37memory if i need to have this core talk
- 9:38to that core
- 9:39remember each one of those only has its
- 9:40own set of registers and its own caches
- 9:42but if i wanted to communicate this guy
- 9:43i could write something to a spot and
- 9:45then i could say
- 9:45and wait and then you could read from
- 9:47the spot so there's a little bit of
- 9:48communication through memory as the
- 9:49shared memory model
- 9:50so that's a shared variable they have a
- 9:51shared variable that i'm working
- 9:52together which is very nice
- 9:53you could have you all be running to the
- 9:54same you'll be writing to the same array
- 9:56let's all contribute to one big array
- 9:58let's all be reading from an array and
- 9:59then writing to the answers in different
- 10:00spots we'll see this we actually
- 10:02see how we do this in software a little
- 10:04bit
- 10:05the drawbacks is that here's the catch
- 10:08memory is slow
- 10:09i'm if i'm gonna the only way these guys
- 10:10can talk is to go to sacramento that's
- 10:13what we're talking about right i'm gonna
- 10:14write to a variable
- 10:15and it's in memory and then you're gonna
- 10:16read from the variable or we're all
- 10:17right into the same thing whatever how
- 10:18we do our work
- 10:20memory's slow going to sacramento was
- 10:22slow so that's a downside and
- 10:23that's a bottleneck i think i have a
- 10:25single bottom like we've really talked
- 10:26about amdahl's law that's in a couple of
- 10:27lectures after this
- 10:29law says the serial part of a program is
- 10:31going to be your downfall it's going to
- 10:32really be the place that is the
- 10:34bottleneck for any piece of software and
- 10:36i say serial well
- 10:38there's only one piece of memory so i'm
- 10:39kind of serially writing through this i
- 10:40can't
- 10:41in parallel think about that so if i'm
- 10:43all running to the same spot that's an
- 10:44issue in terms of synchronization so
- 10:46that's an issue
- 10:48you can think about using a
- 10:49multi-processor in different way you
- 10:50have what's called job level parallelism
- 10:52and they're working on separate problems
- 10:53okay so one core is working on a movie
- 10:56and one core is working on a spreadsheet
- 10:57so that's that's
- 10:58totally reasonable and there's no
- 10:59communication or you could have a more
- 11:01complex and more kind of common
- 11:03task which is well i need to be
- 11:05converting something i'm playing a movie
- 11:07while you can be working on this part of
- 11:08the screen and i can be working on these
- 11:10pixels like
- 11:10so you could have multiple cores all
- 11:12working on the same piece of memory
- 11:14and that's like a maybe a large matrix
- 11:15or a large array of doing something and
- 11:17that's
- 11:17really interesting and that's what we'll
- 11:18be speaking about for the next couple of
- 11:19lectures is doing that
- 11:21in general big picture parallel
- 11:23processing in fact we often
- 11:25this is under the umbrella of parallel
- 11:27and distributed computing that's the
- 11:28idea
- 11:29this is all hard doing things at the
- 11:31same time is hard um
- 11:32it's inevitable we talked about before
- 11:34we've kind of run up against
- 11:36the trying to get more squeeze more
- 11:38performance out of a stone and we can't
- 11:40get it so we
- 11:41it's almost like we have to it wasn't
- 11:43like i'm so excited to go parallel
- 11:44because it's
- 11:45making programming easier it's not it's
- 11:46making programming a lot harder but we
- 11:48have to we have it's almost like eating
- 11:50your broccoli we have to go there
- 11:51to get the performance out of these
- 11:53machines that are being built for us now
- 11:54so it's inevitable the only path to
- 11:56increasing performance really is going
- 11:57parallel and understanding that
- 11:58and also it's the only way to deal with
- 12:00battery life issues um and power issues
- 12:02i mentioned before it isn't the case
- 12:03that we can keep just turning the clock
- 12:05speed up
- 12:05and relying on a computer to be faster
- 12:09every year that's flat these numbers are
- 12:11all flat the only way to continue to
- 12:13grow is and to continue to
- 12:14make use of the hardware that's coming
- 12:16out of here is to think about how to
- 12:17program parallel
- 12:18systems well in mobile systems we see
- 12:21this i mentioned before
- 12:22in phones we have multiple cores and my
- 12:23watch has two cores my new watch
- 12:25has apple series with a four five or six
- 12:27that has two cores it's
- 12:28it's remarkable and different sub
- 12:30components of that as well um
- 12:32you also have different categorizations
- 12:34it's dedicated processors for motion and
- 12:35for image and for neural processing at
- 12:37least on the
- 12:38on the iphones and the gpu has always
- 12:39been usually the gpu's been there in
- 12:40fact in the early days of a 6881 there
- 12:43was a special floating point unit for
- 12:44that so we've always had kind of
- 12:45dedicated hardware for things i say
- 12:47always i mean
- 12:48in my lifetime-ish but now it really is
- 12:50the case on one die there are pieces of
- 12:52it usually
- 12:52here's a floating point co-processor
- 12:54boom a separate dive for floating point
- 12:56and here's the main cpu and here's the
- 12:57floating point then
- 12:58let's make it let's let's pull the gpu
- 12:59out have a big gpu now you have a huge
- 13:02zoom with big fans and you pay a lot of
- 13:03money for the high end gpus the gamers
- 13:05who do this so that's we've almost
- 13:07always had either floating point
- 13:09separate from the enabled main cpu or
- 13:10the graphic separate from the main cpu
- 13:12this is useful
- 13:13um we're going to talk about how to do
- 13:15this at a different scale what happens
- 13:16when rather than talk about one core
- 13:18one die and how to manage that what
- 13:19happens when we're thinking about many
- 13:21many many machines
- 13:22all in one warehouse like facebook and
- 13:25google and
- 13:26apple for siri at least how many folks
- 13:30amazon obviously how many people how do
- 13:33you actually think
- 13:34about the warehouse scale level of that
- 13:35if you're if you're controlling
- 13:37whatever maybe a million computers each
- 13:39computer has many cores
- 13:41how are you able to program that we'll
- 13:42talk about even software ways to do that
- 13:44it's very exciting once we get to the
- 13:45ability
- 13:45ability to be able to you know hit a
- 13:47couple lines of python even
- 13:49boom and all of a sudden a thousand
- 13:50machines can wake up and and you know do
- 13:52your bidding that's pretty remarkable so
- 13:54you have multiple nodes each of these
- 13:56modes have multiple cpus and disk and
- 13:57memory and think about it's one big unit
- 13:59which is incredible and how to cool it
- 14:01is the big deal
- 14:02we'll get to that little later but you
- 14:03obviously have mimdi
- 14:05which is multi-core and simdi which is
- 14:06these vector wide operations in each
- 14:09node
- 14:10let's talk about some performance let's
- 14:12we're trying to this is a big picture
- 14:13kind of lecture
- 14:13um so this is a sense of uh the year of
- 14:17you know
- 14:17kind of every odd year for the last uh
- 14:1915 years or so 18 years or so
- 14:22so let's talk 2009 to kind of almost now
- 14:25or in these
- 14:25upcoming years so how how is the number
- 14:28of course change
- 14:29you know where we're growing we're on
- 14:31average we're kind of growing obviously
- 14:32they're higher ones that are now
- 14:33i think we'll see that the latest was 28
- 14:35cores or something but these numbers are
- 14:37you know
- 14:37doubling uh every uh uh
- 14:41remarkable doubling every two and a half
- 14:42years or so
- 14:44this is that's a two and a half times
- 14:46increase total in the number of course
- 14:48over those 12 years this is an eight
- 14:50times increase in the number of cmd bits
- 14:52per core
- 14:53which is nice if you multiply the core
- 14:56times the 70 bits the total floating
- 14:57point operations you can do
- 14:59per cycle is a factor of 20. so that's
- 15:0220 times in 12 years
- 15:04we're basically doubling every three
- 15:06years if we can use it
- 15:08so now i've got this system at the
- 15:09bottom line how do we make use of all
- 15:11this
- 15:12stuff how do we actually from the point
- 15:13of view software why i came out of
- 15:15vanilla computer science school you
- 15:16can't transplant somebody from taking
- 15:181970s computer science and put them
- 15:20today and expect them to be productive
- 15:22because they didn't have any courses on
- 15:24how to make use of all that parallel
- 15:26hardware it's really hard so let's
- 15:28and to some issues there so let's close
- 15:30it there
- 15:31and uh think about how we would even
- 15:34make use of that bottom row if you had
- 15:36the crazy high-end machine in the bottom
- 15:38row
- 15:38how you can even make use of that we'll
- 15:40see the next lecture
About this transcript
This page contains the full transcript of [CS61C FA20] Lecture 33.2 - Thread-Level Parallelism I: Multicore by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 3,532 words across 570 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.