[CS61C FA20] Lecture 39.LIVE - GPU Guest Lecture with James Percy — Transcript
Full transcript
- 0:01and and action we can see James yeah
- 0:04yeah James is dark oh he's still better
- 0:08he's a little better it'll get better
- 0:09when we get better okay all right uh all
- 0:13right ladies and gentlemen welcome to
- 0:16the penultimate lecture a guest lecture
- 0:20by our awesome host James pery from
- 0:22Apple on jpu
- 0:25architecture so happy so all right I'm
- 0:28GNA stop my camera and add make should
- 0:30to add James's camera here add his
- 0:32Spotlight boom now you're in take mine
- 0:35out all right James you're up thank you
- 0:37for that thank you for that RNN uh
- 0:39comment James you're up welcome thank
- 0:41you so much for coming to share us how
- 0:43gpus work what they are uh and and
- 0:46actually had to program them that's the
- 0:47very exciting thing I think students are
- 0:48waiting to see so good stuff that's rock
- 0:50and roll great thanks for that intro Dan
- 0:54um so yep we'll be talking about GPU
- 0:56architecture today this is a
- 0:57presentation that uh myself and my
- 1:00colleagues at Apple John Kors and Harold
- 1:02OB uh have put together and um before we
- 1:06get started I do do have one uh
- 1:09background thing to
- 1:11cover um so and I know we are recording
- 1:14this for the purp of Berkeley but uh for
- 1:16the class I just want to remind
- 1:17everybody else uh please refrain from
- 1:19recording or posting or streaming any
- 1:21any slides or taking pictures uh due to
- 1:24to confidentiality
- 1:26purposes um so now we got that out of
- 1:28the way and I should be back better
- 1:30lighting now um so what are we going to
- 1:32talk about today so I wanted to give a
- 1:34little bit of an overview of of what is
- 1:37a GPU um what do they do for us um we'll
- 1:40get into a little bit of details about
- 1:43the graphics pipeline um at a high level
- 1:45we'll cover some of the major major
- 1:47stages and talk a little bit about
- 1:50programmability inside a GPU and why
- 1:52that programmability is powerful and
- 1:54then we'll walk through a uh a
- 1:56programming example and then uh
- 1:59hopefully we'll try to I'll try to run a
- 2:00live demo on my on my Mac um and show
- 2:04you guys um the power of of a GPU um
- 2:09when it comes to parallel programming
- 2:10and and what we can do so uh ask
- 2:13questions along the way right into the
- 2:14chat box Dan's going to help monitor
- 2:16that for me and um we'll also try to
- 2:18leave some time at the end to answer
- 2:22questions so first what is a GPU um well
- 2:27a good way to answer that question is
- 2:29actually to compare it to a CPU um many
- 2:31of you I believe are familiar with what
- 2:33CPU architectures look like um and
- 2:36frankly CPU architectures are are a
- 2:38little bit easier to think about I mean
- 2:40they are very complicated in their own
- 2:42right for very good reasons um uh and
- 2:45gpus are complicated as well but for for
- 2:48somewhat different reasons um and and so
- 2:50what I have here is obviously a very
- 2:53stylized version of a CPU on the left
- 2:56and and a GPU on the right um with with
- 2:58very similar components that are
- 3:00colorcoded and um at a at a really high
- 3:03level you can think about um and
- 3:05hopefully you can see my cursor here uh
- 3:08uh the way CPU works is there's there's
- 3:10a whole bunch of uh control logic to
- 3:14decode and execute uh uh instructions
- 3:18there's execution units that actually do
- 3:20the math and run those instructions and
- 3:22it's backed by memory um usually a hiero
- 3:25hiero level of caches um and then a
- 3:28memory system
- 3:30um and the and the general idea is to
- 3:31get better performance out of a CPU you
- 3:34want to minimize your latency through
- 3:37all of these different steps um and and
- 3:40in fact modern CPUs the the number of
- 3:42cycles of latency that you measure from
- 3:44these execution units out to your memory
- 3:46is a critical piece of of your
- 3:48performance it's not the only thing that
- 3:49determines your performance but but it
- 3:51is a very important
- 3:53piece um gpus have similar concerns and
- 3:56you can kind of see the way that these
- 3:58boxes are stylized is you have
- 4:00many many execution units you have
- 4:03smaller uh control units that are
- 4:06perhaps simpler but you have more of
- 4:07them controlling all these execution
- 4:09units and then you also have these
- 4:11little caches that's what these kind of
- 4:12yellowish things are associated with
- 4:14each uh control unit and they're all
- 4:16backed by the same the same DM the key
- 4:19Point here and we'll get into some of
- 4:21the the details of what this actually
- 4:22means but the key Point here is that
- 4:24these blue squares are are multiplied
- 4:28many many times so so when you look at a
- 4:30CPU complex you might have two or three
- 4:33or four maybe even up to eight CPU cores
- 4:35on a given chip or S so on on a modern G
- 4:39large GPU you could easily go up to well
- 4:42well over a 100 and and that's the key
- 4:45difference and that's what gives us the
- 4:47big parallel uh parallel programming
- 4:50boost that we get outside we get from a
- 4:51GPU and so order to get the maximum
- 4:53performance from a GPU what you really
- 4:56your goal really is is to maximize your
- 4:58throughput through all these BL BL
- 4:59execution units or in other words you've
- 5:01got a list all these parallel processing
- 5:03units you want to optimize for using as
- 5:06many of them uh as as you can and what
- 5:09what we'll talk about as we go through
- 5:10this presentation is that Maps very well
- 5:12to moving pixels on the screen you can
- 5:14kind of imagine if you've got a whole
- 5:16bunch of groups of pixels you can
- 5:18distribute them across this array of
- 5:19execution units and they can somewhat
- 5:21get executed in parallel so that that's
- 5:23kind of what we're going to talk about
- 5:26today so this slide is is just kind of a
- 5:29breakdown of of kind of what I just what
- 5:31I just talked about and some of the the
- 5:33carab abouts um as as we're trying to
- 5:36optimize for CPU and GPU performance I
- 5:39talked about the the number of cores on
- 5:41a CPU we typically have more on a GPU um
- 5:44frequency we typically care more about
- 5:46on the CPU not so much these days but
- 5:48there was the frequency Wars back in the
- 5:50late 90s and early 2000s where where all
- 5:52we cared about for CPU performance was
- 5:54just pushing Max frequency gpus are much
- 5:57less about pushing maximum frequency and
- 5:59just pushing maximum fruit it which is
- 6:01what we talked about on on the the
- 6:03previous
- 6:04slide um I'm I'm not going to go through
- 6:06every line on this slide but to hit on a
- 6:08couple of points um speculation and
- 6:10execution order are important
- 6:12particularly if you've covered CPU
- 6:14architecture um the speculation and
- 6:17execution order on the GPU tends to be
- 6:19simpler than on a CPU on a CPU there's
- 6:21an aw awful lot of logic in an area
- 6:24spent around how do I do how do I do
- 6:27Branch prediction how do I reorder
- 6:28instructions to Maxim performance gpus
- 6:31typically don't do that simply because
- 6:33doing those kind of optimizations across
- 6:35the number of execution units that we
- 6:37showed in the previous previous diagram
- 6:39is is very complex um and doesn't scale
- 6:43very well and so that's not something
- 6:45that we try to optimize on optimize for
- 6:48on on gpus and what that Nets out to is
- 6:50our ex execution control tends to be
- 6:53simpler um and then the last thing I'll
- 6:56mention here is is coherency um kind of
- 6:58likewise on CPUs there's a especially
- 7:00when you have a multi-core CPU complex
- 7:03um there's a lot of complexity and
- 7:05effort spent around managing coherency
- 7:08across multiple CPU threads um and and
- 7:11for good reasons you you want to have a
- 7:13relatively simple uh programming model
- 7:16for for the GPU where we have a more
- 7:18complicated software programming model
- 7:19and I'll show some examples of that in a
- 7:21bit coherency tends to be software
- 7:22managed meaning if you have two
- 7:24different threads that need to
- 7:25communicate with each other it's up to
- 7:27the programmer or the developer to
- 7:28manage that synchronization to ensure
- 7:30data
- 7:34consistency so J there's a really funny
- 7:38a really nice analogy that Brian brings
- 7:40up which is he he's analogy for as a CPU
- 7:42is like a math professor and a GPU is
- 7:43like a classroom of elementary school
- 7:45students and you're giving them both
- 7:47math problems but you want to give the
- 7:49complex problems to the professor and
- 7:50many easy problems to the students what
- 7:52do you think about that analogy I like
- 7:53it Bri um it's an interesting analogy I
- 7:55think if your you know if your goal at
- 7:57the end of the day was to get as many
- 7:59simple uh math problems complete in a
- 8:03short enough in a in a shortest time as
- 8:05possible then that's that's an excellent
- 8:07analogy because what you what you really
- 8:09want is is a bunch of elementary
- 8:10students that can just crank away in
- 8:13parallel at a bunch of of of of simple
- 8:15problems simple problems exactly even
- 8:17even the smartest math professor is is
- 8:19not going to be able to keep up at with
- 8:22with a 100 smart students um if you're
- 8:25just cranking through through basic math
- 8:26problems so from that perspective I
- 8:28think that's that's a good way of
- 8:29thinking
- 8:29love it love it great
- 8:35thanks um so uh so maximizing throughput
- 8:41hiding latency we've talked about
- 8:42throughput we want to optim we want to
- 8:43use these execution units that we have
- 8:46um as as much as possible and so you'll
- 8:49many of you may have heard about
- 8:50something called simd um simd is an very
- 8:53important Concept in parallel
- 8:55programming stands for single
- 8:56instruction multiple destination or
- 8:58multiple data um and and what that means
- 9:01is that you you have one set of
- 9:04instructions um that you're going to
- 9:05execute across um a large data set so
- 9:09every element for example just say you
- 9:11had a screen with a bunch of pixels on
- 9:12it every pixel on your screen is going
- 9:14to execute the same instructions and and
- 9:16what that gives gives you is what we
- 9:18call thread level parallelism and so
- 9:20typically you'll hear about threads or
- 9:21or work groups or or warps or things
- 9:23like that um and generally um the the
- 9:27work that you put into a thread is all
- 9:29going to be executing the the same um
- 9:31the same instruction and and the work of
- 9:35the GPU and the complexity of the GPU is
- 9:38is managing that so how do I how do I
- 9:40schedule work across all these different
- 9:42execution units such that I I get the
- 9:44maximum throughput across all of all the
- 9:47data that I have and all the different
- 9:48sets of instructions that I need to
- 9:50schedule um and and so that that's
- 9:53really what this what this is about and
- 9:54there's a lot of complexity about um you
- 9:57know how do these things access memory
- 9:59um how do you uh and this kind of goes
- 10:02back to C coherency how do you schedule
- 10:03them how do they talk to each other um
- 10:06and some of that complexity falls on the
- 10:08programmer and the and the driver and
- 10:10some of it falls on on the hard on the
- 10:12hardware um one other thing to to
- 10:14mention here that's important is is this
- 10:17memory bandwidth aspect of it um and
- 10:20memory bandwidth we'll talk about in the
- 10:22concept in the context of of pixels and
- 10:24games and images um but typically gpus
- 10:28have access to very High memory
- 10:29bandwidth higher more memory bandwidth
- 10:31that you would have on on the GPU H
- 10:33sorry on the CPU and the reason for that
- 10:35is if you just imagine um a game if
- 10:37you're playing a game and you need to
- 10:39refresh your screen um uh 30 times a
- 10:42second for your game to look good just
- 10:45right there you need to refresh your
- 10:46screen you need to touch every
- 10:48theoretically every pixel on the screen
- 10:49but on top of that you need to render a
- 10:51bunch of stuff behind it and I'll show
- 10:52you what that rendering looks like and
- 10:54so you're moving data back and forth um
- 10:56to generate these these uh images
- 10:58multiple times times a second and so you
- 11:00need to have high memory bandwidth to do
- 11:01that to do that efficiently um and and
- 11:04so a lot of the GPU challenges and and
- 11:07optimizations are around um reducing
- 11:10that memory bandwidth how do we move
- 11:12less data how do we have cash
- 11:13hierarchies um to reduce the uh the the
- 11:17memory that's moving from onchip to to
- 11:19off chip or the data I should say moving
- 11:21from onchip to off
- 11:25chip so so those are kind of the basic
- 11:28concepts let's talk a little bit about
- 11:30um what the pipeline looks like and how
- 11:32we how we uh apply some of these
- 11:34Concepts um so at at a very high level
- 11:38um the GPU consists of um of of really
- 11:43four or five uh major major sections um
- 11:47there's the uh the vertex processing
- 11:49section where we are are processing
- 11:52vertices and we're doing some
- 11:53computation on them and I'll explain
- 11:54each of these in a little bit more
- 11:56detail um there's there's rasterization
- 11:59where we're converting triangles into
- 12:01individual pixels that show up on the
- 12:03screen there's fragment processing where
- 12:05we can actually uh run programs on
- 12:08individual pixels to develop to um
- 12:11generate their colors and then there's
- 12:12the frame goer section where we actually
- 12:14output the pixels to a buffer in memory
- 12:17so they can be displayed on
- 12:19screen um one detail I gloss over I
- 12:22mentioned it in the context of fragment
- 12:23processing but for vertex processing you
- 12:25can also run individual programs on on
- 12:28each vertex
- 12:30and the modern gpus actually use the
- 12:32same Hardware that's where you see this
- 12:34unified Shader core on on the left here
- 12:36in green to run these programs um and
- 12:40and that really comes back to the
- 12:42parallel programming model that I talked
- 12:43about in uh earlier in order to process
- 12:47millions of of pixels a second and
- 12:50million potentially millions of vertices
- 12:52per second you need to have a large
- 12:54computation energy engine that can that
- 12:56can run these programs and so what you
- 12:59what you see kind of in our our stylized
- 13:01example over here on the right um these
- 13:04indiv individual pixels this these red
- 13:06Guys these blue guys and these green
- 13:07guys for example each of these could be
- 13:09running a program to um to generate
- 13:12these final pixels and and write them
- 13:14out to memory so one way to think about
- 13:16that is if you're playing a game and
- 13:17you're seeing pixels on your on your
- 13:19screen you may be um every pixel in your
- 13:22screen is running a program now they
- 13:24might be running the same program the
- 13:25same what's called a pixel Shader um but
- 13:28they're all running a program and and so
- 13:30that's kind of the the beauty of of GPU
- 13:32parallel
- 13:35processing and at the end of the day and
- 13:37we'll go through kind of examples of
- 13:39what builds up uh builds up this Frame
- 13:41um you get an image that looks looks
- 13:43something like this um you have objects
- 13:47which is this uh this kind of unicorny
- 13:49shiny unicorn thing in the middle um you
- 13:52have a background um with the fireworks
- 13:55on it you have these these objects in
- 13:56the background with the M's on them
- 13:57which stands for metal these are these
- 13:59are textures I'll explain what textures
- 14:01are you've got this surface that's kind
- 14:03of reflecting those are those are based
- 14:05off of of textures and lighting models
- 14:08um and so we'll talk about some of the
- 14:09the details of what goes into making a
- 14:11frame a frame like
- 14:14this so let's get into the to the
- 14:16pipeline uh in in a little bit more more
- 14:19depth so so typically um if if you're
- 14:23making a game you've got a a model of of
- 14:27your world or or you at least have OB
- 14:29objects that you have models of that are
- 14:30made up of of triangles um and so
- 14:33somebody's going to program their game
- 14:36um in in a graphics API such as metal
- 14:39that's what we use on um Apple Apple
- 14:42devices um and they're they're basically
- 14:45going to input their models or or where
- 14:47their vertices are in um into the game
- 14:51and give them some sort of description
- 14:52of how they connect how they turn into
- 14:54tance and triangles um so we're going to
- 14:57we're going to run our our vertex
- 15:00processing out room that we that we
- 15:01talked about to um to generate turn
- 15:04these individual vertices into um into
- 15:08triangles that's called assembly you'll
- 15:10notice that we have this kind of world
- 15:12uh Cube or or or shape here where we've
- 15:16actually clipped off a piece of this
- 15:18triangle that does not get rendered so
- 15:20we go through a process of of clipping
- 15:22pieces of the model that we that we
- 15:24don't
- 15:25see um we rasterize or scan convert them
- 15:28and you can imagine a process where you
- 15:30if you have a given triangle you you
- 15:32literally kind of walk uh left to right
- 15:36on each pixel uh on the screen and say
- 15:38hey is this inside my triangle or not
- 15:40and if it is then yes you render it if
- 15:42not you you throw it away that's not a
- 15:44that's not a great optimized algorithm
- 15:46for for doing it but it's an algorithm
- 15:48that will work and you can kind of
- 15:49conceptualize how you might um how you
- 15:52might turn a a triangle into individual
- 15:56pixels um then we get into the FR frag
- 15:59processing where we have uh we we do
- 16:01lighting algorithms we're going to for
- 16:03every pixel inside a triangle we're
- 16:04going to compute a lighting equation we
- 16:06may texture it and I'll explain what
- 16:07that is meaning pulling textur from
- 16:10memory you can compute effects you can
- 16:12do um what basically whatever you want
- 16:15to to generate the final pixels um and
- 16:17then we then we send them off to our
- 16:20screen so let's dive into vertex
- 16:22processing a little bit um so here we
- 16:25have the same scene that we looked at a
- 16:27minute ago and this has been drawn in
- 16:29something that's called wireframe in the
- 16:30bottom right here and you can see the
- 16:32the Unicorn you can see the the shapes
- 16:34and if you look closely you can see that
- 16:35everything is broken into triangles even
- 16:38even these kind of square like shapes in
- 16:40the background here if you look closely
- 16:42you can see you can see the diagonal so
- 16:45and it's been broken into a triangle and
- 16:46the reason for that is triangles are
- 16:48planer they're a lot easier for the
- 16:49hardware to uh to work with rather than
- 16:52a true 3D 3D
- 16:54shape um and so the first thing we need
- 16:57to do is we need to get our our our
- 16:59unicorn into the the right space in the
- 17:02world and so you can do that with the
- 17:05vertex Shader by doing a number of
- 17:07transforms on the Unicorn you can rotate
- 17:09it you can translate it meaning move it
- 17:11you can stretch it in the x or the y or
- 17:13the the Z Dimension that's basically
- 17:15what the vertex Shader is is going to do
- 17:19um it's also going to place the world uh
- 17:22place a camera in the world so you can
- 17:24imagine if we were drawing uh rendering
- 17:27this in a game this camera might be like
- 17:28moving around relative to where the
- 17:30Unicorn is or the Unicorn might be
- 17:32moving around relative to the camera um
- 17:35it doesn't really matter the vertex
- 17:37Shader is going to get all the objects
- 17:40in in the world by doing these
- 17:41transforms into a specific place so that
- 17:44we can go render it from from the
- 17:46screen and James just to connect this
- 17:48with what the students are doing now
- 17:50they're working on a faster Matrix
- 17:51multiply that's exactly what's Happening
- 17:53Here the transformation matrix Matrix is
- 17:55a 4x4 Matrix and the pixel Matrix is a
- 17:59xyzw by however many vertices you've got
- 18:02so basically it's that huge really tall
- 18:04Matrix times the 4x4 that's all this is
- 18:07doing and it's like it's exactly what a
- 18:08GPU was meant for me before we had
- 18:11general purpose gpus we had gpus that
- 18:12just knew how to do this really fast
- 18:15take those take those points and then
- 18:17transform them into the right space
- 18:18really that's exactly what they were
- 18:19optimized for exactly and all these
- 18:22transforms you guys probably know can be
- 18:23represented as a 4x4 Matrix and you can
- 18:25actually multiply pre-multiply as I'm
- 18:27sure you know these these mat es
- 18:29together which is what what you would do
- 18:31in a real vertex Shader um to do these
- 18:33computations sorry Bor I think you had a
- 18:35comment and no problem our students do
- 18:38not fortunately or unfortunately do not
- 18:40get to program a GPU this time around
- 18:44they just got access to Intel AVX
- 18:47extensions got it back in the in the
- 18:49good old days when I learned this stuff
- 18:51um gpus actually had what we called
- 18:53fixed function Hardware to do these
- 18:55transforms so what what that would mean
- 18:57is you would you would program more or
- 19:00less a matrix it wasn't quite described
- 19:02that way but you would basically program
- 19:04hey what is my rotation transformation
- 19:06what is my translation transformation uh
- 19:08into the GPU and it would you would send
- 19:11it send it in vertices and it would do
- 19:13the transforms for you and spit out
- 19:14triangles out at the back side that's
- 19:17basically what would happen these days
- 19:20instead of having that fixed function
- 19:21Hardware to do that M to do these
- 19:23specific multiplications these things
- 19:24are expressed in what we call the vertex
- 19:26Shader meaning you can program in what
- 19:28whatever transforms that you want um and
- 19:30whatever M more or less whatever
- 19:32matrices that you want as long as
- 19:33they're not nonsensical um to to do
- 19:35these
- 19:41transforms all right so let's talk a
- 19:43little bit about geometry um and I'm I'm
- 19:47going to pick try to pick up the pace a
- 19:48little bit here um geometry processing
- 19:51so we we've taken our triangles we've
- 19:52now assembled them or primi we do
- 19:54primitive assembly into into triangles
- 19:56so we're not really dealing with
- 19:58individual ver veres anymore um we're
- 20:00we're dealing with with triangles um and
- 20:04um this is a little bit of a detail I'm
- 20:06going to kind of go go over it somewhat
- 20:08quickly but one thing that many not all
- 20:10but many modern gpus will do is they
- 20:12will actually take the triangles in
- 20:14screen space that's what's shown on the
- 20:16left here um and then actually bend them
- 20:19and that's kind of what this is intended
- 20:21to represent on the right and so what
- 20:23this what this picture on the right
- 20:25represents is the lighter colors or the
- 20:27whiter colors represent present triangle
- 20:30complexity so what you can see here is
- 20:32is um where we map to the unicorn's hair
- 20:36there's a lot of little triangles in
- 20:37here to do the hair same thing with the
- 20:39tail um and so you can get kind of an
- 20:41approximation for how expensive this
- 20:44piece of the scene is to render just
- 20:47based purely on the number of triangles
- 20:48that are there whereas if you go into
- 20:50the background there are zero triangles
- 20:52so it's just black or on these these two
- 20:54surfaces that we show here um there
- 20:57aren't very many triangles so they're
- 20:58they're gray or or close to Black um
- 21:02that's important for a number of reasons
- 21:04um it lets you estimate how expensive it
- 21:06is to render that piece of the scene and
- 21:08that can help you with your work
- 21:09distribution going back to the topic we
- 21:11had beforehand where it's all about
- 21:13optimizing the throughput of the machine
- 21:15we can use that to estimate hey this
- 21:17piece of the scene is going to be
- 21:18expensive so we want to be able to
- 21:20distribute this piece of the scene to a
- 21:22lot of our execution units um so that's
- 21:26something that not all gpus operate this
- 21:28way but um many modern gpus
- 21:30do okay so rasterization so we've we're
- 21:34at the step now where um we've got our
- 21:37triangles and we're going to go we're
- 21:38going to go do a process called scan
- 21:40conversion scan conversion is is kind of
- 21:43I mentioned this earlier it's converting
- 21:44the triangles into individual pixels and
- 21:46so you could imagine an algorithm where
- 21:49perhaps I start at this this vertex I
- 21:51walk along an edge um and then until I
- 21:54hit another Edge and then I I walk over
- 21:57a little bit and I walk back that's
- 21:59that's basically scan conversion where
- 22:00I've taken this triangle I've turned it
- 22:02into individual pixels um I'm going to
- 22:04skip over antialising uh for the
- 22:06purposes of this presentation but you
- 22:08could imagine actually along these edges
- 22:10you could actually create even more
- 22:12samples for to make higher quality
- 22:13images um if if you wanted to um so once
- 22:18I've got my pixels figured out now I
- 22:20need to go figure out well how am I
- 22:21going to how am I going to color them um
- 22:24and that's a step called fragment
- 22:25fragment processing um and so so um with
- 22:30with uh with every we've now got these
- 22:33these triangles in screen space we also
- 22:35carry depth information along with each
- 22:38pixel um and we can do something called
- 22:41a depth test um and the depth test
- 22:43basically tells you is my triangle
- 22:46visible or not so if you got two
- 22:47triangles like this that shown them the
- 22:48screen they they intersect um this let's
- 22:52say this one the left is in the back you
- 22:54don't need to render the part of of that
- 22:56triangle mostly um if it doesn't show up
- 22:58on the screen so we can just throw it
- 22:59away and only render the the triangle on
- 23:02on top um so that that's what the depth
- 23:04test is and then the other piece of the
- 23:07scene may just be Unwritten so we might
- 23:09just store a bit for each each tile or
- 23:13um really just means each pixel on the
- 23:14screen that says hey this is the clear
- 23:16color we're not going to touch it we
- 23:17don't have to do anything with that
- 23:18particular
- 23:20pixel so now comes the really the really
- 23:23interesting and exciting part which is
- 23:25which is the shading um and so we talked
- 23:28about this this common unified uh Shader
- 23:31core that executes uh vertex fragment
- 23:34and we didn't talk about compute but we
- 23:35will in a little bit uh shaders on on
- 23:38the hardware and the Shader is what lets
- 23:41you specify that per pixel algorithm
- 23:44that you're going to run to do a sh a
- 23:46lighting calculation and so um just to
- 23:49to kind of give some insight on what
- 23:51that might look like is there's a teapot
- 23:53example here on the left you see it's
- 23:55just very simply it's black or white um
- 23:57there's no there's no interesting
- 23:59shading to it it's basically just uh it
- 24:02looks pretty simple on the right hand
- 24:04side it's the same teapot but it's been
- 24:06shaded with a with some specular
- 24:08lighting um and what that means is you
- 24:10can kind of see the light uh reflecting
- 24:12off the top surface of of the teapot
- 24:14depending on where the angle of the
- 24:16light is coming from that gives it some
- 24:18some interesting shading effects you can
- 24:19kind of see it on the handle on the top
- 24:21and then over here on the on the spout
- 24:23um a little bit and so um this is a a
- 24:27relatively uh simple lighting algorithm
- 24:30but it does require a fair amount of
- 24:31computation you have to compute some
- 24:32some angles um you you need uh you need
- 24:36to do a texture lookup and I'll explain
- 24:38what a texture lookup is in a second to
- 24:40go to go compute this and you need to do
- 24:42this for every pixel or on on on the
- 24:45screen so it's not a cheap calculation
- 24:47and that comes back to the fact that the
- 24:49simy algorithm we talked about
- 24:51beforehand for every pixel on this
- 24:55teapot we're going to go run this
- 24:57program down here and we can do that in
- 24:59parallel so that it happens very quickly
- 25:03as opposed to if you were doing a pixel
- 25:04at a time on a CPU it would take um it
- 25:06would it would take much longer um and
- 25:08that's where the beauty of of of having
- 25:10dedicated GPU Hardware comes in so this
- 25:13um this particular uh uh shading program
- 25:18that I've got as an example down here on
- 25:20the right is is written in metal uh
- 25:22which is uh again a programming language
- 25:24for for Apple uh Apple hardware and you
- 25:27can see it's it's comp Computing um a
- 25:29light angle based off of based off of a
- 25:32normal it's doing a DOT product um it's
- 25:35doing a a texture lookup to figure out
- 25:37the diffuse color um and then it's
- 25:40multiplying the the diffus color by um
- 25:42by the computed angle of the light and
- 25:45that's what it's outputting and that's
- 25:46basically how we're GNA we're going to
- 25:47shade every pixel of of this
- 25:50teapot um there's there's U the modern
- 25:54training languages are U much more
- 25:57complex originally we could just do kind
- 25:58of simple math op math operations um and
- 26:02and text lookups um they've gotten more
- 26:05complicated over the years to get better
- 26:07compute operations at Matrix multiplies
- 26:10um at neural special instructions for
- 26:13doing neural Nets um anything kind of uh
- 26:17related to to parallel processing um is
- 26:20is kind of now part of modern GPU
- 26:23shading languages and I and I mentioned
- 26:25kind of coherency earlier there's also
- 26:26instructions for doing synch ization so
- 26:29you can kind of you can wait to for
- 26:31other threads to finish so you can
- 26:32synchronize data um there's there's some
- 26:35limited amount of of predication um you
- 26:37don't want to be putting a ton of if
- 26:38then else statements inside your your
- 26:40Shader program that's not going to be
- 26:42good for performance but there is uh
- 26:44there is some support for for predicated
- 26:48instructions so this is um what we're
- 26:50going to come back to uh at at the end
- 26:53and I'll hopefully show you a live demo
- 26:55of of a program running both on the CPU
- 26:58and on on the
- 27:00GPU um so to talk a little bit more
- 27:03about the execution model um so I talked
- 27:06about how we're going to we're going to
- 27:08execute this fragment Shader on every
- 27:09touched pixel so every pixel on this
- 27:11this teapot um and we're going to do
- 27:14that in a way that takes advantage of
- 27:15thread level parallelism meaning for
- 27:17every pixel we're going to run the same
- 27:18instruction on multiple data so we come
- 27:20back to our concept of of simd so if
- 27:23we're going to do this imagine um that
- 27:26this this kind of square that we have uh
- 27:28is a zoomed in piece of of this teapot
- 27:31um you could imagine an algorithm where
- 27:33we just walk one at a time left to right
- 27:35um for each pixel inside this this
- 27:37Square we run the program um and that
- 27:40would that would work but it's obviously
- 27:42very slow because it does everything one
- 27:43at a time and so the idea of of a simd
- 27:48uh simd machine is that you have groups
- 27:51that work in in lock lock lock step
- 27:54excuse me so you might just as a simple
- 27:57example take groups of four a 2X two
- 27:59subsquare inside this big square and
- 28:01you're going to execute the same
- 28:02instruction on all four of those pixels
- 28:04um at the same time and in a simple case
- 28:07where you're just um executing simple
- 28:08instructions they're just going to all
- 28:10execute one two three four in in lock
- 28:13stck lock step and finish at the same
- 28:16time now you can get into more
- 28:19complicated examples and I'm going to
- 28:20gloss over this uh a little bit where
- 28:23you might have one of those pixels has
- 28:25to run a high latency texture fetch or
- 28:27or latency operation and what that means
- 28:29is you have to stall all four of those
- 28:31pixels so you kind of take them off um
- 28:33you stall them that's what's intended to
- 28:35show in this diagram and you go find a
- 28:37different set of of work to go to go run
- 28:41so you might find a different um thread
- 28:43group that you can go execute can make
- 28:46progress while this High latency
- 28:47operation is is running um and that's
- 28:50where a lot of the complexity and
- 28:53frankly the interesting parts of the GPU
- 28:56uh comes in because you're managing
- 28:59these thousands of threads that are
- 29:00running within the machine and trying to
- 29:02optimally execute them in a in a way
- 29:04that fills the machine um in in the best
- 29:06Manner and that's that's a really hard
- 29:08Pro really hard
- 29:13problem I'm going to talk about um
- 29:15texture mapping briefly so I have time
- 29:17for uh to to get to the demo so I I kind
- 29:20of mentioned at the beginning um in the
- 29:22background of our our our Pony we had
- 29:25these fireworks we also had these M that
- 29:26were showing up on these surfaces is
- 29:28well this background is not generated
- 29:31programmatically or on the fly as part
- 29:33of the the um the the rendering
- 29:36algorithm these are pre-stored images or
- 29:38or textures um and and what you can do
- 29:41is when you have pre-stored images like
- 29:42this like these fireworks you can map
- 29:44them into your your rendered image and
- 29:48so you can imagine there's just a big
- 29:49triangle or sorry a big Square back here
- 29:51um and what you would do is um map each
- 29:56set of pixels in in your world into this
- 29:59texture and and go fetch them um and
- 30:02that's basically what the texture unit
- 30:03or what texture mapping does is it Maps
- 30:05um onscreen pixels to pixels that are
- 30:08inside a texture now um that can
- 30:11actually create some very interesting
- 30:12and challenging sample problems you
- 30:14could imagine um if if this image here
- 30:18were very small compared to the uh the
- 30:21size that you wanted to render in on
- 30:23screen you've you've got a problem where
- 30:25how do I sample the right number picture
- 30:27uh picture pixs for this to to show up
- 30:29and likewise you can have the opposite
- 30:31problem where you have a very large
- 30:33image that you're um sampling into a
- 30:35very small area on screen so there's a
- 30:36lot of Hardware built into texture units
- 30:39to try to do a good job of of sampling
- 30:41and this is kind of an example that's
- 30:42showing that you're taking um a a square
- 30:46kind of checkerboard pattern and then
- 30:47you're mapping it to something that is
- 30:49that is tall and skinny you can see the
- 30:51results of doing that mapping is your
- 30:53nice Square image is now a number a
- 30:55number of rectangles and that's
- 30:56basically what the texture C is going to
- 30:58going to do this is kind of a more
- 31:01advanced it's called anisotropic
- 31:02filtering uh filter mode that's shown on
- 31:04the right hand side that kind of avoids
- 31:07some of these uh weird artifacts that
- 31:09you see on the top screen um but that's
- 31:11uh that's probably a bit beyond today's
- 31:13presentation and that's the
- 31:14anti-aliasing that we mentioned before
- 31:17yes that is a version of of
- 31:22anti-aliasing okay so our last stage we
- 31:25need to actually get this get all these
- 31:27pixels that we've we worked on um and
- 31:29get them written out to to to memory um
- 31:32and so uh there's usually an instruction
- 31:35not necessar exposed to developer but at
- 31:37some point you've got all these pixels
- 31:38usually in a cache on your chip you need
- 31:40to actually write them out to the frame
- 31:41ruffer um and then tell the frame buffer
- 31:44or tell the display usually hey my image
- 31:46my rendering is done go pick up all this
- 31:49data and go go put it on the display and
- 31:51that's that's kind of what the the
- 31:52output stage is um I there's some other
- 31:55steps in there there's blend tests um
- 31:58there's some stencil test that I kind of
- 32:00glossed glossed over a little bit but
- 32:02that's that's really the uh the gist of
- 32:05it and what you can see on this the
- 32:07bottom right hand side of this picture
- 32:09is this is kind of a frame buffer that's
- 32:11been captured like mid render so some of
- 32:14the uh the tiles or some of the pixels
- 32:17have completed you can see them on the
- 32:18screen others have have not so you so
- 32:22you need basically need to wait till all
- 32:23of the work is completed before you say
- 32:25that you're ready for for display
- 32:30all right so um let's get to our example
- 32:34I'm trying to leave uh about 10 uh 10
- 32:37minutes for questions at the end so um
- 32:40what I want to show you guys is a um sax
- 32:44piie algorithm run on um a CPU and and a
- 32:49GPU and and actually it's even simpler
- 32:52than than that um we're simply going to
- 32:55uh Implement a function um in C
- 32:58that takes a um an
- 33:01array uh X of of floats multiplies each
- 33:05element in that uh array by a float a
- 33:09and outputs it to an output array Y and
- 33:12so we're just going to Loop through the
- 33:13entire array and and do that super
- 33:15simple we're not even doing the plus
- 33:17part so we're just doing the the the
- 33:19multiply part um and so we're going to
- 33:22do that in C on a single threaded
- 33:24CPU and we're going to do a similar
- 33:26thing on on the GPU in metal and so this
- 33:30kind of um code snippet is may be less
- 33:34familiar to you guys than the C program
- 33:36but it essentially does does the same
- 33:37thing I've got an array X and output
- 33:40buffer Y and I'm going to multiply them
- 33:41by a so it does the same function just a
- 33:43different
- 33:45language um and so what does what does
- 33:48that look like and so and I'm going to
- 33:50try to do this interactively in in a
- 33:52second um but basically this what the
- 33:55results I'm showing you here are going
- 33:57through a number of different buffer
- 33:59sizes um and and so it's really uh I
- 34:03squared so you can see the buffer sizes
- 34:05here um in the next column is the the
- 34:09time in it took for the GPU to finish
- 34:12the computation um and then the last one
- 34:15last column is the time that it takes
- 34:16the CPU to complete the comp computation
- 34:19so there's a few things that you'll uh
- 34:21you'll notice as you look at this data
- 34:24um as you get to the bottom very clearly
- 34:27once as your your buffer sizes are
- 34:29getting very large um the GPU is clearly
- 34:32beating the um the CPU in terms of
- 34:35performance it's taking less time to
- 34:37complete these operations um and the
- 34:40crossover point is where is it it's kind
- 34:43of right right around this this area
- 34:46where we see that the um the GPU starts
- 34:48to be beat the CPU you'll notice that
- 34:51the CPU does much much better um on the
- 34:55smaller buffer sizes than the GPU and
- 34:58really what what that tells you is that
- 35:00the GPU is not great at doing kind of
- 35:03small chunks of work and the reason for
- 35:05that is there's a lot of overhead in
- 35:07getting the GPU set up and started
- 35:10particularly on this this platform um to
- 35:13setting it up to run the computation and
- 35:15and so you can see when I go through the
- 35:17first kind of five or six runs of this
- 35:20the time it takes is is more less the
- 35:22same in facts it's oscillating up and
- 35:23down a bit due you know other random
- 35:25stuff going on in the system um
- 35:28but once we get down to about here um
- 35:31the time starts increasing so what that
- 35:32tells you is there's some fixed GPU
- 35:34overhead that's going on here um and the
- 35:36CPU is way better and so the reason for
- 35:39that is is pretty simple we're not
- 35:40getting any advantage of parallel
- 35:43processing for these very small buffers
- 35:45and so for very small simple jobs your
- 35:47CPU is going to do just fine um once you
- 35:49get up to jobs that are millions or
- 35:51hundreds of millions of elements in size
- 35:54um the GPU is going to do better and so
- 35:56why is that well that goes goes back to
- 35:58our our what we've been talking about
- 36:00throughout this this talk simd
- 36:01processing so if you go back and I and I
- 36:04will go back to our example um the CPU
- 36:08is basically doing um what we talked
- 36:11about on the left hand side here it's
- 36:13going literally through the array
- 36:15running this very simple calculation one
- 36:17at a time and so when this the the
- 36:19number of entries that it needs to do
- 36:20the calculation gets very large it's
- 36:22going to take a long time the the GPU is
- 36:26is actually executing
- 36:28um elements of of that buffer in groups
- 36:31of of actually more than four probably
- 36:3332 or 64 or even even larger um such
- 36:37that we're executing it in let's call it
- 36:3964 64 size chunks and so what that means
- 36:43is that the GPU is going to is is going
- 36:46to execute those instructions much much
- 36:48faster and we start to see that as the
- 36:49buffer size gets gets very very
- 36:54large so um I'm going to attempt
- 36:58to show this
- 37:00live and I'm I'm also running zoom and
- 37:03sharing Zoom so we'll see we'll see how
- 37:05this goes but um what what you'll be
- 37:08able to
- 37:09see assuming my computer doesn't crash
- 37:12um is is this program executed let me
- 37:15make it make it
- 37:18larger uh is is executing this pretty uh
- 37:21executing this live during the talk um
- 37:24and you can see how hopefully the the
- 37:26GPU performance perance starts to
- 37:28outstrip the the CPU as we
- 37:32go and there it is GPU is starting to to
- 37:34beat the
- 37:36CPU CPU times getting getting longer and
- 37:40longer we're about 2x the GPU right
- 37:44now about
- 37:463x I can hear my fan going up there's a
- 37:51there's a question uh whether this is
- 37:52being done on the new M1 or whether that
- 37:55would change things I wish it was doing
- 37:56being done on M1 um no I I oddly my uh
- 38:01my wife has one of those because I
- 38:02bought one for her but my the Mac I'm
- 38:04running this on for work uh I believe is
- 38:07a
- 38:082018 uh Intel based system I don't know
- 38:10the exact uh exact number but it's a
- 38:13it's a Intel integrated GPU um and then
- 38:16you can see by the time we we got done
- 38:19here um that the the GPU was almost
- 38:22seven times faster than uh the CPU for
- 38:25the largest the largest buffer
- 38:31um now this was a a very simple program
- 38:35um and in fact uh an astute uh engineer
- 38:39might notice that hey this thing is
- 38:40actually not doing that much calculation
- 38:43um so it actually is is is at least on
- 38:46the GPU memory limited we talked about
- 38:48memory band with with earlier the GPU is
- 38:51much better at moving chunks large
- 38:53chunks of of memory around than than the
- 38:55CPU so we we are also benefiting from
- 38:58not just the number of execution units
- 38:59but the amount of um the amount of
- 39:02memory bandwidth that we have now it
- 39:04turns out I think on this particular CPU
- 39:06that I'm running on um because I've
- 39:09we've explicitly said it to operate in a
- 39:12single core single thread manner um the
- 39:15CPU core is likely Mass limited um Al I
- 39:18have not actually tested
- 39:22that all right so um so here's a graph
- 39:25of what we just did this is obviously
- 39:27not the experiment I just ran but you
- 39:29can see the results were pretty similar
- 39:30so I ran this over the weekend without
- 39:32zoom and other stuff running in the
- 39:34background this is a logarithmic scale
- 39:37um so so notice that carefully um of of
- 39:41where we of our array size on the bottom
- 39:43with the runtime and milliseconds on on
- 39:45the Y AIS you can kind of see the
- 39:47crossover Point here and then again
- 39:49remembering that this is a logarithmic
- 39:51scale as we get up to these very large
- 39:52buffer sizes um the the um the uh the
- 39:56GPU is is beating the uh the CPU
- 40:00performance
- 40:03handily all right so let's wrap it up
- 40:05and then and leave some time for for
- 40:07questions so um so if if you get nothing
- 40:10else from from this presentation um
- 40:13please take away these these through
- 40:14points so getting efficiency and
- 40:17performance out of a GPU is all about
- 40:20maximizing parallelism maximizing
- 40:22throughput via parallelism you're T you
- 40:24want to run parallel programs where
- 40:26you're you're doing Sim the op
- 40:27operations you're running the same same
- 40:29instructions across multiple data um
- 40:31that's What GPU performance is about
- 40:33there's you know some fixed function
- 40:35stuff on the side to do specialized math
- 40:37like like texturing um but really what
- 40:40we optimize around is M maximizing
- 40:43performance through parallelism that
- 40:45makes um uh like image processing for
- 40:50for uh for cameras for doing
- 40:52computational photography for um running
- 40:56uh algorithms to do image detection so
- 40:58um neural
- 41:00networks to be done on on the GPU just
- 41:03as Tesla just
- 41:06exactly um and and gpus modern gpus are
- 41:09very easily expanded to do general
- 41:11purpose compute so companies like like
- 41:13Invidia and AMD have been very
- 41:15successful selling gpus not just for for
- 41:18games but for um for building compute
- 41:21Farms of gpus for doing um machine
- 41:24learning algorithms and and things like
- 41:26that so that's a that's a big growth uh
- 41:29factor for um for gpus in in recent
- 41:33years so with that I'll stop there hope
- 41:36you guys found this interesting um let's
- 41:39uh do some questions in the last few
- 41:40minutes that was great James thank you
- 41:43much for thank you thank you wonderful
- 41:44wonderful we've got a couple questions
- 41:46that are open um the first is um what's
- 41:51the software that's actually managing
- 41:53the coherency in the GPU um would that
- 41:56be something built into the o or how do
- 41:58you manage the coherency making making
- 41:59sure you're worry about um thread
- 42:02stepping on each other that's an
- 42:03excellent question um the the answer is
- 42:05it's really a combination of the
- 42:07application and the driver that is that
- 42:09is programming and managing the lower
- 42:11level details of of the GPU so um if you
- 42:14went and looked at a modern GPI GPU API
- 42:17such as metal for example um there are
- 42:20facilities inside the metal programming
- 42:22language that give you mechanisms For
- 42:25Thread synchronization and memory
- 42:26synchronization
- 42:28um and those kind of high highle
- 42:30facilities get mapped to lowlevel
- 42:32Hardware commands like a cash flush or a
- 42:35fence or a barrier um that the driver
- 42:38will convert into something that the
- 42:39hardware can actually understand um and
- 42:42so to kind of perhaps bring that back to
- 42:44the question the programmer does have to
- 42:46be aware of some of these
- 42:48synchronization requirements and how to
- 42:50how to manage parallel programming um in
- 42:53combination with the driver because if
- 42:55you're not aware of those things you're
- 42:56not going to get good
- 42:58makes sense on this kind of the same
- 43:00topic do you need to worry about
- 43:01deadlock in
- 43:03gpus um the the hardware team meaning my
- 43:07team certainly does um as a as a at a
- 43:11programming level um you need to worry
- 43:14about it kind of from the similar
- 43:16perspective as you would worry about
- 43:18hanging a a CPU by writing a an infinite
- 43:21Loop or or two Loops especially if
- 43:23you're doing parallel programming that
- 43:24kind of depend each other depend on each
- 43:26other in ways that could create
- 43:29Deadlocks um if you get down to the
- 43:32hardware implementation details there
- 43:34are um there are actually all sorts of
- 43:38uh hazards that can cause Deadlocks that
- 43:40we spend a lot of uh design and
- 43:43verification time trying to find and
- 43:45either just avoid architecturally or or
- 43:48put in deadlock Breakers you may call
- 43:51them into the hardware so the hardware
- 43:53can't lock itself up um so yes it's it's
- 43:56something that we we are very worried
- 43:58about Bo these are coming as as you
- 44:01answered one four more questions show up
- 44:03stud students students are doing great
- 44:05we got 70 students here um what's the
- 44:08ideal relationship or is there an ideal
- 44:10relationship between the number of
- 44:11pixels and the number of execution units
- 44:13in a
- 44:13GPU like maybe as you go to 6K 8K how
- 44:16does that you there's some there's some
- 44:18limits where gpus can't drive an 8K
- 44:20screen that's just talking on the screen
- 44:21but in terms of the computation is there
- 44:24a you know mapping at all or maybe a
- 44:26good
- 44:28that's a that's a good question um the
- 44:30the nice thing about GPU processing um
- 44:33is that it scales very very well um and
- 44:36in particular uh if you're if you're
- 44:38doing like super high resolution 8K
- 44:41displays um then you're going to want
- 44:44especially if you're running games and
- 44:46you want like high frame rates you're
- 44:47GNA want a big GPU to drive that now um
- 44:52there's no kind of fixed relationship
- 44:53you could have a small modern GPU drive
- 44:55a 4K or an 8K display um you're just
- 44:58going to get less performance out of out
- 45:00of your game uh if you were playing a
- 45:02game and you may not get 30 frames per
- 45:03second so your game look choppy and not
- 45:06not play well so there's no kind of like
- 45:08fixed requirement if you will um but
- 45:11certainly if you're going to run a a
- 45:13complex modern game on a very high
- 45:15resolution display you're going to want
- 45:16a bigger GPU with more cores to get a
- 45:18good gaming experience so hopefully that
- 45:20kind of and I and I don't get this user
- 45:23perception of 120 frames per second I I
- 45:26don't think my can do
- 45:28it yes there's a lot of debates about 30
- 45:3160 or is 120 hertz even even useful well
- 45:35it's in film too I mean was it load of
- 45:37the Rings had like 60 or 120 and it was
- 45:39like wait this is looks like TV
- 45:40something looks a little different here
- 45:42exactly um this is a great question from
- 45:44Adrian what are some surprising
- 45:46applications where gpus are used in
- 45:47places we might not expect oh wow that's
- 45:50a that's a fantastic question
- 45:53yeah uh I'm not sure if this would be
- 45:56surprising or not these days but um if
- 45:58you've looked at uh if you drive a Tesla
- 46:01or been in a Tesla uh and I I'm not
- 46:03revealing anything secret about Tesla I
- 46:05don't know anything secret about Tesla
- 46:07so just get that out there um but my
- 46:10understanding is they use uh GPU
- 46:11Hardware to do a lot of their uh their
- 46:14image processing algorithms for for
- 46:17self-driving um Soh so you'll find
- 46:20you'll find gpus in in automobiles
- 46:23you'll find uh gpus in a lot of places
- 46:25you wouldn't necessarily expect them
- 46:27a few a few years ago similarly on um
- 46:30any modern phone and I'm not talking
- 46:32specifically about Apple but any modern
- 46:34phone is going to do a fair amount of of
- 46:38computational Photography these days and
- 46:40a lot of the phph run through the GPU to
- 46:44generate or to blend images together to
- 46:47generate fin image right so photography
- 46:49is not simply about like hey how how
- 46:52many megapixels are in my sensor um it's
- 46:55it's how good are the algorithms that
- 46:57you are running and how much GPU and
- 47:00other horsepower do you have to run
- 47:03those algorithms such as the person
- 47:04taking the picture can see it in in less
- 47:06than a second so it's interactive
- 47:08imagine if you were taking a picture and
- 47:10it took 10 seconds for the final final
- 47:12picture to show up that that's not very
- 47:14fun for for a a phone type use case so
- 47:17so gpus matter a lot in in that in that
- 47:20space as well especially with HDR
- 47:22imagery I mean you can you can compress
- 47:24that you know tonal tone mapping where
- 47:25you have a bright part of the screen
- 47:27dark part of the screen but the GPU can
- 47:29take many pictures simultaneously or SE
- 47:31in sequence and then compress them down
- 47:33into one that just captures the highs
- 47:34and lows to reduce dynamic range I love
- 47:36it um I knew this by the way get get one
- 47:41you know there's one surprising thing uh
- 47:43it's Bitcoin mining Bitcoin um you know
- 47:46Bitcoin mining
- 47:47basically because of that you a few
- 47:50years ago it was all L gpus have been
- 47:55sold out I mean it was not POs to buy
- 47:58a they're all that's fascinating I
- 48:01didn't know that yeah now you can buy
- 48:04cheap dedicated Hardware that does these
- 48:06hash functions more efficiently
- 48:10that's three years ago I think it was
- 48:12December when when Bitcoin was taking
- 48:14off you could not find a highend GPU
- 48:16people were were selling them for
- 48:18thousands of dollars where people to go
- 48:19Bitcoin mining so that's aample and by
- 48:22the way I just looking at the the the
- 48:23latest Mac Pro the highend gpus are like
- 48:26$2,000 just by themselves before the
- 48:28before the price you know increase
- 48:30because people want to you know buy them
- 48:31to do these things the high GPS can be
- 48:33very expensive um this is Adam a great
- 48:35question I knew this was going to come
- 48:36up this is so cool how can we students
- 48:39play with some graphics image processing
- 48:41code how do we code to use a GPU
- 48:43ourselves I first answer is get a Mac y
- 48:47get a Mac um we have metal uh Apple has
- 48:51metal programming guides and very simple
- 48:53examples of just how to get started if
- 48:54you just Google I should put a link in
- 48:56here but if you just Google uh metal
- 48:58programming like introduction um there's
- 49:01some online tutorials that you can go
- 49:02through and go like Hey how do I write
- 49:04like my the hell world of of gpus using
- 49:07metal and I'm sure Nvidia and AMD have
- 49:10the same thing for for their Hardware
- 49:12using other other languages so there's
- 49:14definitely tutorials out there and
- 49:16there's also I think um Cuda and open CL
- 49:19is that right those two those those are
- 49:23programming apis Auda is specific to to
- 49:26Nvidia
- 49:27an open CL um I'm not sure how much that
- 49:29is used as much these days that was more
- 49:31of an open open API crossplatform I
- 49:34think it's been replaced by Cuda and
- 49:36others right he's I I'll just add to the
- 49:39question which is the code that you
- 49:40showed in metal you can be running on a
- 49:43built-in CPU on a laptop or a high-end
- 49:45Mac Pro with a dedicated two boxes of a
- 49:47CPU do you need to worry about that at
- 49:49that level at the metal level it just
- 49:50just as handled by by the by the by the
- 49:52abstraction layer right that's as long
- 49:54as you write in a way that's parallel
- 49:56programming friendly um it doesn't
- 49:59matter if you're targeting like
- 50:00something built into a phone or if
- 50:02you're targeting a giant GPU system the
- 50:05the API takes care of that for you
- 50:07that's wonderful now unfortunately I've
- 50:08got I think 14 questions left but we
- 50:10promis chemistry which actually needs
- 50:11this this webinar like now we're now two
- 50:13minutes into their webinar so folks
- 50:15let's let's give James a hand thank you
- 50:17so much for coming and sharing those
- 50:19thoughts about the GPU thank you James
- 50:21thank you James wonderful wonderful
- 50:23Wonder so grateful to have you here
- 50:25thanks thanks again standing outstanding
- 50:27lecture there was a meta question about
- 50:29these slides confidential we we don't
- 50:31have the copies of the slides but we
- 50:32have the we have the webinar so we'll
- 50:33just keep that and they'll just be able
- 50:35to watch the slides through the webinar
- 50:36I'm guessing y okay perfect perfect
- 50:39thanks so much all right take care
- 50:41everybody and stay here if you want to
- 50:42watch leure okay take care everybody
- 50:45thanks again James stay safe folks
- 50:47wonderful all right bye bye
About this transcript
This page contains the full transcript of [CS61C FA20] Lecture 39.LIVE - GPU Guest Lecture with James Percy by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 9,594 words across 1,307 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.