YouTube2Text

[CS61C FA20] Lecture 39.LIVE - GPU Guest Lecture with James Percy  — Transcript

by CS 61C Departmental · 9,594 words · 1,307 segments · language en · Watch on YouTube

Full transcript

  1. 0:01and and action we can see James yeah
  2. 0:04yeah James is dark oh he's still better
  3. 0:08he's a little better it'll get better
  4. 0:09when we get better okay all right uh all
  5. 0:13right ladies and gentlemen welcome to
  6. 0:16the penultimate lecture a guest lecture
  7. 0:20by our awesome host James pery from
  8. 0:22Apple on jpu
  9. 0:25architecture so happy so all right I'm
  10. 0:28GNA stop my camera and add make should
  11. 0:30to add James's camera here add his
  12. 0:32Spotlight boom now you're in take mine
  13. 0:35out all right James you're up thank you
  14. 0:37for that thank you for that RNN uh
  15. 0:39comment James you're up welcome thank
  16. 0:41you so much for coming to share us how
  17. 0:43gpus work what they are uh and and
  18. 0:46actually had to program them that's the
  19. 0:47very exciting thing I think students are
  20. 0:48waiting to see so good stuff that's rock
  21. 0:50and roll great thanks for that intro Dan
  22. 0:54um so yep we'll be talking about GPU
  23. 0:56architecture today this is a
  24. 0:57presentation that uh myself and my
  25. 1:00colleagues at Apple John Kors and Harold
  26. 1:02OB uh have put together and um before we
  27. 1:06get started I do do have one uh
  28. 1:09background thing to
  29. 1:11cover um so and I know we are recording
  30. 1:14this for the purp of Berkeley but uh for
  31. 1:16the class I just want to remind
  32. 1:17everybody else uh please refrain from
  33. 1:19recording or posting or streaming any
  34. 1:21any slides or taking pictures uh due to
  35. 1:24to confidentiality
  36. 1:26purposes um so now we got that out of
  37. 1:28the way and I should be back better
  38. 1:30lighting now um so what are we going to
  39. 1:32talk about today so I wanted to give a
  40. 1:34little bit of an overview of of what is
  41. 1:37a GPU um what do they do for us um we'll
  42. 1:40get into a little bit of details about
  43. 1:43the graphics pipeline um at a high level
  44. 1:45we'll cover some of the major major
  45. 1:47stages and talk a little bit about
  46. 1:50programmability inside a GPU and why
  47. 1:52that programmability is powerful and
  48. 1:54then we'll walk through a uh a
  49. 1:56programming example and then uh
  50. 1:59hopefully we'll try to I'll try to run a
  51. 2:00live demo on my on my Mac um and show
  52. 2:04you guys um the power of of a GPU um
  53. 2:09when it comes to parallel programming
  54. 2:10and and what we can do so uh ask
  55. 2:13questions along the way right into the
  56. 2:14chat box Dan's going to help monitor
  57. 2:16that for me and um we'll also try to
  58. 2:18leave some time at the end to answer
  59. 2:22questions so first what is a GPU um well
  60. 2:27a good way to answer that question is
  61. 2:29actually to compare it to a CPU um many
  62. 2:31of you I believe are familiar with what
  63. 2:33CPU architectures look like um and
  64. 2:36frankly CPU architectures are are a
  65. 2:38little bit easier to think about I mean
  66. 2:40they are very complicated in their own
  67. 2:42right for very good reasons um uh and
  68. 2:45gpus are complicated as well but for for
  69. 2:48somewhat different reasons um and and so
  70. 2:50what I have here is obviously a very
  71. 2:53stylized version of a CPU on the left
  72. 2:56and and a GPU on the right um with with
  73. 2:58very similar components that are
  74. 3:00colorcoded and um at a at a really high
  75. 3:03level you can think about um and
  76. 3:05hopefully you can see my cursor here uh
  77. 3:08uh the way CPU works is there's there's
  78. 3:10a whole bunch of uh control logic to
  79. 3:14decode and execute uh uh instructions
  80. 3:18there's execution units that actually do
  81. 3:20the math and run those instructions and
  82. 3:22it's backed by memory um usually a hiero
  83. 3:25hiero level of caches um and then a
  84. 3:28memory system
  85. 3:30um and the and the general idea is to
  86. 3:31get better performance out of a CPU you
  87. 3:34want to minimize your latency through
  88. 3:37all of these different steps um and and
  89. 3:40in fact modern CPUs the the number of
  90. 3:42cycles of latency that you measure from
  91. 3:44these execution units out to your memory
  92. 3:46is a critical piece of of your
  93. 3:48performance it's not the only thing that
  94. 3:49determines your performance but but it
  95. 3:51is a very important
  96. 3:53piece um gpus have similar concerns and
  97. 3:56you can kind of see the way that these
  98. 3:58boxes are stylized is you have
  99. 4:00many many execution units you have
  100. 4:03smaller uh control units that are
  101. 4:06perhaps simpler but you have more of
  102. 4:07them controlling all these execution
  103. 4:09units and then you also have these
  104. 4:11little caches that's what these kind of
  105. 4:12yellowish things are associated with
  106. 4:14each uh control unit and they're all
  107. 4:16backed by the same the same DM the key
  108. 4:19Point here and we'll get into some of
  109. 4:21the the details of what this actually
  110. 4:22means but the key Point here is that
  111. 4:24these blue squares are are multiplied
  112. 4:28many many times so so when you look at a
  113. 4:30CPU complex you might have two or three
  114. 4:33or four maybe even up to eight CPU cores
  115. 4:35on a given chip or S so on on a modern G
  116. 4:39large GPU you could easily go up to well
  117. 4:42well over a 100 and and that's the key
  118. 4:45difference and that's what gives us the
  119. 4:47big parallel uh parallel programming
  120. 4:50boost that we get outside we get from a
  121. 4:51GPU and so order to get the maximum
  122. 4:53performance from a GPU what you really
  123. 4:56your goal really is is to maximize your
  124. 4:58throughput through all these BL BL
  125. 4:59execution units or in other words you've
  126. 5:01got a list all these parallel processing
  127. 5:03units you want to optimize for using as
  128. 5:06many of them uh as as you can and what
  129. 5:09what we'll talk about as we go through
  130. 5:10this presentation is that Maps very well
  131. 5:12to moving pixels on the screen you can
  132. 5:14kind of imagine if you've got a whole
  133. 5:16bunch of groups of pixels you can
  134. 5:18distribute them across this array of
  135. 5:19execution units and they can somewhat
  136. 5:21get executed in parallel so that that's
  137. 5:23kind of what we're going to talk about
  138. 5:26today so this slide is is just kind of a
  139. 5:29breakdown of of kind of what I just what
  140. 5:31I just talked about and some of the the
  141. 5:33carab abouts um as as we're trying to
  142. 5:36optimize for CPU and GPU performance I
  143. 5:39talked about the the number of cores on
  144. 5:41a CPU we typically have more on a GPU um
  145. 5:44frequency we typically care more about
  146. 5:46on the CPU not so much these days but
  147. 5:48there was the frequency Wars back in the
  148. 5:50late 90s and early 2000s where where all
  149. 5:52we cared about for CPU performance was
  150. 5:54just pushing Max frequency gpus are much
  151. 5:57less about pushing maximum frequency and
  152. 5:59just pushing maximum fruit it which is
  153. 6:01what we talked about on on the the
  154. 6:03previous
  155. 6:04slide um I'm I'm not going to go through
  156. 6:06every line on this slide but to hit on a
  157. 6:08couple of points um speculation and
  158. 6:10execution order are important
  159. 6:12particularly if you've covered CPU
  160. 6:14architecture um the speculation and
  161. 6:17execution order on the GPU tends to be
  162. 6:19simpler than on a CPU on a CPU there's
  163. 6:21an aw awful lot of logic in an area
  164. 6:24spent around how do I do how do I do
  165. 6:27Branch prediction how do I reorder
  166. 6:28instructions to Maxim performance gpus
  167. 6:31typically don't do that simply because
  168. 6:33doing those kind of optimizations across
  169. 6:35the number of execution units that we
  170. 6:37showed in the previous previous diagram
  171. 6:39is is very complex um and doesn't scale
  172. 6:43very well and so that's not something
  173. 6:45that we try to optimize on optimize for
  174. 6:48on on gpus and what that Nets out to is
  175. 6:50our ex execution control tends to be
  176. 6:53simpler um and then the last thing I'll
  177. 6:56mention here is is coherency um kind of
  178. 6:58likewise on CPUs there's a especially
  179. 7:00when you have a multi-core CPU complex
  180. 7:03um there's a lot of complexity and
  181. 7:05effort spent around managing coherency
  182. 7:08across multiple CPU threads um and and
  183. 7:11for good reasons you you want to have a
  184. 7:13relatively simple uh programming model
  185. 7:16for for the GPU where we have a more
  186. 7:18complicated software programming model
  187. 7:19and I'll show some examples of that in a
  188. 7:21bit coherency tends to be software
  189. 7:22managed meaning if you have two
  190. 7:24different threads that need to
  191. 7:25communicate with each other it's up to
  192. 7:27the programmer or the developer to
  193. 7:28manage that synchronization to ensure
  194. 7:30data
  195. 7:34consistency so J there's a really funny
  196. 7:38a really nice analogy that Brian brings
  197. 7:40up which is he he's analogy for as a CPU
  198. 7:42is like a math professor and a GPU is
  199. 7:43like a classroom of elementary school
  200. 7:45students and you're giving them both
  201. 7:47math problems but you want to give the
  202. 7:49complex problems to the professor and
  203. 7:50many easy problems to the students what
  204. 7:52do you think about that analogy I like
  205. 7:53it Bri um it's an interesting analogy I
  206. 7:55think if your you know if your goal at
  207. 7:57the end of the day was to get as many
  208. 7:59simple uh math problems complete in a
  209. 8:03short enough in a in a shortest time as
  210. 8:05possible then that's that's an excellent
  211. 8:07analogy because what you what you really
  212. 8:09want is is a bunch of elementary
  213. 8:10students that can just crank away in
  214. 8:13parallel at a bunch of of of of simple
  215. 8:15problems simple problems exactly even
  216. 8:17even the smartest math professor is is
  217. 8:19not going to be able to keep up at with
  218. 8:22with a 100 smart students um if you're
  219. 8:25just cranking through through basic math
  220. 8:26problems so from that perspective I
  221. 8:28think that's that's a good way of
  222. 8:29thinking
  223. 8:29love it love it great
  224. 8:35thanks um so uh so maximizing throughput
  225. 8:41hiding latency we've talked about
  226. 8:42throughput we want to optim we want to
  227. 8:43use these execution units that we have
  228. 8:46um as as much as possible and so you'll
  229. 8:49many of you may have heard about
  230. 8:50something called simd um simd is an very
  231. 8:53important Concept in parallel
  232. 8:55programming stands for single
  233. 8:56instruction multiple destination or
  234. 8:58multiple data um and and what that means
  235. 9:01is that you you have one set of
  236. 9:04instructions um that you're going to
  237. 9:05execute across um a large data set so
  238. 9:09every element for example just say you
  239. 9:11had a screen with a bunch of pixels on
  240. 9:12it every pixel on your screen is going
  241. 9:14to execute the same instructions and and
  242. 9:16what that gives gives you is what we
  243. 9:18call thread level parallelism and so
  244. 9:20typically you'll hear about threads or
  245. 9:21or work groups or or warps or things
  246. 9:23like that um and generally um the the
  247. 9:27work that you put into a thread is all
  248. 9:29going to be executing the the same um
  249. 9:31the same instruction and and the work of
  250. 9:35the GPU and the complexity of the GPU is
  251. 9:38is managing that so how do I how do I
  252. 9:40schedule work across all these different
  253. 9:42execution units such that I I get the
  254. 9:44maximum throughput across all of all the
  255. 9:47data that I have and all the different
  256. 9:48sets of instructions that I need to
  257. 9:50schedule um and and so that that's
  258. 9:53really what this what this is about and
  259. 9:54there's a lot of complexity about um you
  260. 9:57know how do these things access memory
  261. 9:59um how do you uh and this kind of goes
  262. 10:02back to C coherency how do you schedule
  263. 10:03them how do they talk to each other um
  264. 10:06and some of that complexity falls on the
  265. 10:08programmer and the and the driver and
  266. 10:10some of it falls on on the hard on the
  267. 10:12hardware um one other thing to to
  268. 10:14mention here that's important is is this
  269. 10:17memory bandwidth aspect of it um and
  270. 10:20memory bandwidth we'll talk about in the
  271. 10:22concept in the context of of pixels and
  272. 10:24games and images um but typically gpus
  273. 10:28have access to very High memory
  274. 10:29bandwidth higher more memory bandwidth
  275. 10:31that you would have on on the GPU H
  276. 10:33sorry on the CPU and the reason for that
  277. 10:35is if you just imagine um a game if
  278. 10:37you're playing a game and you need to
  279. 10:39refresh your screen um uh 30 times a
  280. 10:42second for your game to look good just
  281. 10:45right there you need to refresh your
  282. 10:46screen you need to touch every
  283. 10:48theoretically every pixel on the screen
  284. 10:49but on top of that you need to render a
  285. 10:51bunch of stuff behind it and I'll show
  286. 10:52you what that rendering looks like and
  287. 10:54so you're moving data back and forth um
  288. 10:56to generate these these uh images
  289. 10:58multiple times times a second and so you
  290. 11:00need to have high memory bandwidth to do
  291. 11:01that to do that efficiently um and and
  292. 11:04so a lot of the GPU challenges and and
  293. 11:07optimizations are around um reducing
  294. 11:10that memory bandwidth how do we move
  295. 11:12less data how do we have cash
  296. 11:13hierarchies um to reduce the uh the the
  297. 11:17memory that's moving from onchip to to
  298. 11:19off chip or the data I should say moving
  299. 11:21from onchip to off
  300. 11:25chip so so those are kind of the basic
  301. 11:28concepts let's talk a little bit about
  302. 11:30um what the pipeline looks like and how
  303. 11:32we how we uh apply some of these
  304. 11:34Concepts um so at at a very high level
  305. 11:38um the GPU consists of um of of really
  306. 11:43four or five uh major major sections um
  307. 11:47there's the uh the vertex processing
  308. 11:49section where we are are processing
  309. 11:52vertices and we're doing some
  310. 11:53computation on them and I'll explain
  311. 11:54each of these in a little bit more
  312. 11:56detail um there's there's rasterization
  313. 11:59where we're converting triangles into
  314. 12:01individual pixels that show up on the
  315. 12:03screen there's fragment processing where
  316. 12:05we can actually uh run programs on
  317. 12:08individual pixels to develop to um
  318. 12:11generate their colors and then there's
  319. 12:12the frame goer section where we actually
  320. 12:14output the pixels to a buffer in memory
  321. 12:17so they can be displayed on
  322. 12:19screen um one detail I gloss over I
  323. 12:22mentioned it in the context of fragment
  324. 12:23processing but for vertex processing you
  325. 12:25can also run individual programs on on
  326. 12:28each vertex
  327. 12:30and the modern gpus actually use the
  328. 12:32same Hardware that's where you see this
  329. 12:34unified Shader core on on the left here
  330. 12:36in green to run these programs um and
  331. 12:40and that really comes back to the
  332. 12:42parallel programming model that I talked
  333. 12:43about in uh earlier in order to process
  334. 12:47millions of of pixels a second and
  335. 12:50million potentially millions of vertices
  336. 12:52per second you need to have a large
  337. 12:54computation energy engine that can that
  338. 12:56can run these programs and so what you
  339. 12:59what you see kind of in our our stylized
  340. 13:01example over here on the right um these
  341. 13:04indiv individual pixels this these red
  342. 13:06Guys these blue guys and these green
  343. 13:07guys for example each of these could be
  344. 13:09running a program to um to generate
  345. 13:12these final pixels and and write them
  346. 13:14out to memory so one way to think about
  347. 13:16that is if you're playing a game and
  348. 13:17you're seeing pixels on your on your
  349. 13:19screen you may be um every pixel in your
  350. 13:22screen is running a program now they
  351. 13:24might be running the same program the
  352. 13:25same what's called a pixel Shader um but
  353. 13:28they're all running a program and and so
  354. 13:30that's kind of the the beauty of of GPU
  355. 13:32parallel
  356. 13:35processing and at the end of the day and
  357. 13:37we'll go through kind of examples of
  358. 13:39what builds up uh builds up this Frame
  359. 13:41um you get an image that looks looks
  360. 13:43something like this um you have objects
  361. 13:47which is this uh this kind of unicorny
  362. 13:49shiny unicorn thing in the middle um you
  363. 13:52have a background um with the fireworks
  364. 13:55on it you have these these objects in
  365. 13:56the background with the M's on them
  366. 13:57which stands for metal these are these
  367. 13:59are textures I'll explain what textures
  368. 14:01are you've got this surface that's kind
  369. 14:03of reflecting those are those are based
  370. 14:05off of of textures and lighting models
  371. 14:08um and so we'll talk about some of the
  372. 14:09the details of what goes into making a
  373. 14:11frame a frame like
  374. 14:14this so let's get into the to the
  375. 14:16pipeline uh in in a little bit more more
  376. 14:19depth so so typically um if if you're
  377. 14:23making a game you've got a a model of of
  378. 14:27your world or or you at least have OB
  379. 14:29objects that you have models of that are
  380. 14:30made up of of triangles um and so
  381. 14:33somebody's going to program their game
  382. 14:36um in in a graphics API such as metal
  383. 14:39that's what we use on um Apple Apple
  384. 14:42devices um and they're they're basically
  385. 14:45going to input their models or or where
  386. 14:47their vertices are in um into the game
  387. 14:51and give them some sort of description
  388. 14:52of how they connect how they turn into
  389. 14:54tance and triangles um so we're going to
  390. 14:57we're going to run our our vertex
  391. 15:00processing out room that we that we
  392. 15:01talked about to um to generate turn
  393. 15:04these individual vertices into um into
  394. 15:08triangles that's called assembly you'll
  395. 15:10notice that we have this kind of world
  396. 15:12uh Cube or or or shape here where we've
  397. 15:16actually clipped off a piece of this
  398. 15:18triangle that does not get rendered so
  399. 15:20we go through a process of of clipping
  400. 15:22pieces of the model that we that we
  401. 15:24don't
  402. 15:25see um we rasterize or scan convert them
  403. 15:28and you can imagine a process where you
  404. 15:30if you have a given triangle you you
  405. 15:32literally kind of walk uh left to right
  406. 15:36on each pixel uh on the screen and say
  407. 15:38hey is this inside my triangle or not
  408. 15:40and if it is then yes you render it if
  409. 15:42not you you throw it away that's not a
  410. 15:44that's not a great optimized algorithm
  411. 15:46for for doing it but it's an algorithm
  412. 15:48that will work and you can kind of
  413. 15:49conceptualize how you might um how you
  414. 15:52might turn a a triangle into individual
  415. 15:56pixels um then we get into the FR frag
  416. 15:59processing where we have uh we we do
  417. 16:01lighting algorithms we're going to for
  418. 16:03every pixel inside a triangle we're
  419. 16:04going to compute a lighting equation we
  420. 16:06may texture it and I'll explain what
  421. 16:07that is meaning pulling textur from
  422. 16:10memory you can compute effects you can
  423. 16:12do um what basically whatever you want
  424. 16:15to to generate the final pixels um and
  425. 16:17then we then we send them off to our
  426. 16:20screen so let's dive into vertex
  427. 16:22processing a little bit um so here we
  428. 16:25have the same scene that we looked at a
  429. 16:27minute ago and this has been drawn in
  430. 16:29something that's called wireframe in the
  431. 16:30bottom right here and you can see the
  432. 16:32the Unicorn you can see the the shapes
  433. 16:34and if you look closely you can see that
  434. 16:35everything is broken into triangles even
  435. 16:38even these kind of square like shapes in
  436. 16:40the background here if you look closely
  437. 16:42you can see you can see the diagonal so
  438. 16:45and it's been broken into a triangle and
  439. 16:46the reason for that is triangles are
  440. 16:48planer they're a lot easier for the
  441. 16:49hardware to uh to work with rather than
  442. 16:52a true 3D 3D
  443. 16:54shape um and so the first thing we need
  444. 16:57to do is we need to get our our our
  445. 16:59unicorn into the the right space in the
  446. 17:02world and so you can do that with the
  447. 17:05vertex Shader by doing a number of
  448. 17:07transforms on the Unicorn you can rotate
  449. 17:09it you can translate it meaning move it
  450. 17:11you can stretch it in the x or the y or
  451. 17:13the the Z Dimension that's basically
  452. 17:15what the vertex Shader is is going to do
  453. 17:19um it's also going to place the world uh
  454. 17:22place a camera in the world so you can
  455. 17:24imagine if we were drawing uh rendering
  456. 17:27this in a game this camera might be like
  457. 17:28moving around relative to where the
  458. 17:30Unicorn is or the Unicorn might be
  459. 17:32moving around relative to the camera um
  460. 17:35it doesn't really matter the vertex
  461. 17:37Shader is going to get all the objects
  462. 17:40in in the world by doing these
  463. 17:41transforms into a specific place so that
  464. 17:44we can go render it from from the
  465. 17:46screen and James just to connect this
  466. 17:48with what the students are doing now
  467. 17:50they're working on a faster Matrix
  468. 17:51multiply that's exactly what's Happening
  469. 17:53Here the transformation matrix Matrix is
  470. 17:55a 4x4 Matrix and the pixel Matrix is a
  471. 17:59xyzw by however many vertices you've got
  472. 18:02so basically it's that huge really tall
  473. 18:04Matrix times the 4x4 that's all this is
  474. 18:07doing and it's like it's exactly what a
  475. 18:08GPU was meant for me before we had
  476. 18:11general purpose gpus we had gpus that
  477. 18:12just knew how to do this really fast
  478. 18:15take those take those points and then
  479. 18:17transform them into the right space
  480. 18:18really that's exactly what they were
  481. 18:19optimized for exactly and all these
  482. 18:22transforms you guys probably know can be
  483. 18:23represented as a 4x4 Matrix and you can
  484. 18:25actually multiply pre-multiply as I'm
  485. 18:27sure you know these these mat es
  486. 18:29together which is what what you would do
  487. 18:31in a real vertex Shader um to do these
  488. 18:33computations sorry Bor I think you had a
  489. 18:35comment and no problem our students do
  490. 18:38not fortunately or unfortunately do not
  491. 18:40get to program a GPU this time around
  492. 18:44they just got access to Intel AVX
  493. 18:47extensions got it back in the in the
  494. 18:49good old days when I learned this stuff
  495. 18:51um gpus actually had what we called
  496. 18:53fixed function Hardware to do these
  497. 18:55transforms so what what that would mean
  498. 18:57is you would you would program more or
  499. 19:00less a matrix it wasn't quite described
  500. 19:02that way but you would basically program
  501. 19:04hey what is my rotation transformation
  502. 19:06what is my translation transformation uh
  503. 19:08into the GPU and it would you would send
  504. 19:11it send it in vertices and it would do
  505. 19:13the transforms for you and spit out
  506. 19:14triangles out at the back side that's
  507. 19:17basically what would happen these days
  508. 19:20instead of having that fixed function
  509. 19:21Hardware to do that M to do these
  510. 19:23specific multiplications these things
  511. 19:24are expressed in what we call the vertex
  512. 19:26Shader meaning you can program in what
  513. 19:28whatever transforms that you want um and
  514. 19:30whatever M more or less whatever
  515. 19:32matrices that you want as long as
  516. 19:33they're not nonsensical um to to do
  517. 19:35these
  518. 19:41transforms all right so let's talk a
  519. 19:43little bit about geometry um and I'm I'm
  520. 19:47going to pick try to pick up the pace a
  521. 19:48little bit here um geometry processing
  522. 19:51so we we've taken our triangles we've
  523. 19:52now assembled them or primi we do
  524. 19:54primitive assembly into into triangles
  525. 19:56so we're not really dealing with
  526. 19:58individual ver veres anymore um we're
  527. 20:00we're dealing with with triangles um and
  528. 20:04um this is a little bit of a detail I'm
  529. 20:06going to kind of go go over it somewhat
  530. 20:08quickly but one thing that many not all
  531. 20:10but many modern gpus will do is they
  532. 20:12will actually take the triangles in
  533. 20:14screen space that's what's shown on the
  534. 20:16left here um and then actually bend them
  535. 20:19and that's kind of what this is intended
  536. 20:21to represent on the right and so what
  537. 20:23this what this picture on the right
  538. 20:25represents is the lighter colors or the
  539. 20:27whiter colors represent present triangle
  540. 20:30complexity so what you can see here is
  541. 20:32is um where we map to the unicorn's hair
  542. 20:36there's a lot of little triangles in
  543. 20:37here to do the hair same thing with the
  544. 20:39tail um and so you can get kind of an
  545. 20:41approximation for how expensive this
  546. 20:44piece of the scene is to render just
  547. 20:47based purely on the number of triangles
  548. 20:48that are there whereas if you go into
  549. 20:50the background there are zero triangles
  550. 20:52so it's just black or on these these two
  551. 20:54surfaces that we show here um there
  552. 20:57aren't very many triangles so they're
  553. 20:58they're gray or or close to Black um
  554. 21:02that's important for a number of reasons
  555. 21:04um it lets you estimate how expensive it
  556. 21:06is to render that piece of the scene and
  557. 21:08that can help you with your work
  558. 21:09distribution going back to the topic we
  559. 21:11had beforehand where it's all about
  560. 21:13optimizing the throughput of the machine
  561. 21:15we can use that to estimate hey this
  562. 21:17piece of the scene is going to be
  563. 21:18expensive so we want to be able to
  564. 21:20distribute this piece of the scene to a
  565. 21:22lot of our execution units um so that's
  566. 21:26something that not all gpus operate this
  567. 21:28way but um many modern gpus
  568. 21:30do okay so rasterization so we've we're
  569. 21:34at the step now where um we've got our
  570. 21:37triangles and we're going to go we're
  571. 21:38going to go do a process called scan
  572. 21:40conversion scan conversion is is kind of
  573. 21:43I mentioned this earlier it's converting
  574. 21:44the triangles into individual pixels and
  575. 21:46so you could imagine an algorithm where
  576. 21:49perhaps I start at this this vertex I
  577. 21:51walk along an edge um and then until I
  578. 21:54hit another Edge and then I I walk over
  579. 21:57a little bit and I walk back that's
  580. 21:59that's basically scan conversion where
  581. 22:00I've taken this triangle I've turned it
  582. 22:02into individual pixels um I'm going to
  583. 22:04skip over antialising uh for the
  584. 22:06purposes of this presentation but you
  585. 22:08could imagine actually along these edges
  586. 22:10you could actually create even more
  587. 22:12samples for to make higher quality
  588. 22:13images um if if you wanted to um so once
  589. 22:18I've got my pixels figured out now I
  590. 22:20need to go figure out well how am I
  591. 22:21going to how am I going to color them um
  592. 22:24and that's a step called fragment
  593. 22:25fragment processing um and so so um with
  594. 22:30with uh with every we've now got these
  595. 22:33these triangles in screen space we also
  596. 22:35carry depth information along with each
  597. 22:38pixel um and we can do something called
  598. 22:41a depth test um and the depth test
  599. 22:43basically tells you is my triangle
  600. 22:46visible or not so if you got two
  601. 22:47triangles like this that shown them the
  602. 22:48screen they they intersect um this let's
  603. 22:52say this one the left is in the back you
  604. 22:54don't need to render the part of of that
  605. 22:56triangle mostly um if it doesn't show up
  606. 22:58on the screen so we can just throw it
  607. 22:59away and only render the the triangle on
  608. 23:02on top um so that that's what the depth
  609. 23:04test is and then the other piece of the
  610. 23:07scene may just be Unwritten so we might
  611. 23:09just store a bit for each each tile or
  612. 23:13um really just means each pixel on the
  613. 23:14screen that says hey this is the clear
  614. 23:16color we're not going to touch it we
  615. 23:17don't have to do anything with that
  616. 23:18particular
  617. 23:20pixel so now comes the really the really
  618. 23:23interesting and exciting part which is
  619. 23:25which is the shading um and so we talked
  620. 23:28about this this common unified uh Shader
  621. 23:31core that executes uh vertex fragment
  622. 23:34and we didn't talk about compute but we
  623. 23:35will in a little bit uh shaders on on
  624. 23:38the hardware and the Shader is what lets
  625. 23:41you specify that per pixel algorithm
  626. 23:44that you're going to run to do a sh a
  627. 23:46lighting calculation and so um just to
  628. 23:49to kind of give some insight on what
  629. 23:51that might look like is there's a teapot
  630. 23:53example here on the left you see it's
  631. 23:55just very simply it's black or white um
  632. 23:57there's no there's no interesting
  633. 23:59shading to it it's basically just uh it
  634. 24:02looks pretty simple on the right hand
  635. 24:04side it's the same teapot but it's been
  636. 24:06shaded with a with some specular
  637. 24:08lighting um and what that means is you
  638. 24:10can kind of see the light uh reflecting
  639. 24:12off the top surface of of the teapot
  640. 24:14depending on where the angle of the
  641. 24:16light is coming from that gives it some
  642. 24:18some interesting shading effects you can
  643. 24:19kind of see it on the handle on the top
  644. 24:21and then over here on the on the spout
  645. 24:23um a little bit and so um this is a a
  646. 24:27relatively uh simple lighting algorithm
  647. 24:30but it does require a fair amount of
  648. 24:31computation you have to compute some
  649. 24:32some angles um you you need uh you need
  650. 24:36to do a texture lookup and I'll explain
  651. 24:38what a texture lookup is in a second to
  652. 24:40go to go compute this and you need to do
  653. 24:42this for every pixel or on on on the
  654. 24:45screen so it's not a cheap calculation
  655. 24:47and that comes back to the fact that the
  656. 24:49simy algorithm we talked about
  657. 24:51beforehand for every pixel on this
  658. 24:55teapot we're going to go run this
  659. 24:57program down here and we can do that in
  660. 24:59parallel so that it happens very quickly
  661. 25:03as opposed to if you were doing a pixel
  662. 25:04at a time on a CPU it would take um it
  663. 25:06would it would take much longer um and
  664. 25:08that's where the beauty of of of having
  665. 25:10dedicated GPU Hardware comes in so this
  666. 25:13um this particular uh uh shading program
  667. 25:18that I've got as an example down here on
  668. 25:20the right is is written in metal uh
  669. 25:22which is uh again a programming language
  670. 25:24for for Apple uh Apple hardware and you
  671. 25:27can see it's it's comp Computing um a
  672. 25:29light angle based off of based off of a
  673. 25:32normal it's doing a DOT product um it's
  674. 25:35doing a a texture lookup to figure out
  675. 25:37the diffuse color um and then it's
  676. 25:40multiplying the the diffus color by um
  677. 25:42by the computed angle of the light and
  678. 25:45that's what it's outputting and that's
  679. 25:46basically how we're GNA we're going to
  680. 25:47shade every pixel of of this
  681. 25:50teapot um there's there's U the modern
  682. 25:54training languages are U much more
  683. 25:57complex originally we could just do kind
  684. 25:58of simple math op math operations um and
  685. 26:02and text lookups um they've gotten more
  686. 26:05complicated over the years to get better
  687. 26:07compute operations at Matrix multiplies
  688. 26:10um at neural special instructions for
  689. 26:13doing neural Nets um anything kind of uh
  690. 26:17related to to parallel processing um is
  691. 26:20is kind of now part of modern GPU
  692. 26:23shading languages and I and I mentioned
  693. 26:25kind of coherency earlier there's also
  694. 26:26instructions for doing synch ization so
  695. 26:29you can kind of you can wait to for
  696. 26:31other threads to finish so you can
  697. 26:32synchronize data um there's there's some
  698. 26:35limited amount of of predication um you
  699. 26:37don't want to be putting a ton of if
  700. 26:38then else statements inside your your
  701. 26:40Shader program that's not going to be
  702. 26:42good for performance but there is uh
  703. 26:44there is some support for for predicated
  704. 26:48instructions so this is um what we're
  705. 26:50going to come back to uh at at the end
  706. 26:53and I'll hopefully show you a live demo
  707. 26:55of of a program running both on the CPU
  708. 26:58and on on the
  709. 27:00GPU um so to talk a little bit more
  710. 27:03about the execution model um so I talked
  711. 27:06about how we're going to we're going to
  712. 27:08execute this fragment Shader on every
  713. 27:09touched pixel so every pixel on this
  714. 27:11this teapot um and we're going to do
  715. 27:14that in a way that takes advantage of
  716. 27:15thread level parallelism meaning for
  717. 27:17every pixel we're going to run the same
  718. 27:18instruction on multiple data so we come
  719. 27:20back to our concept of of simd so if
  720. 27:23we're going to do this imagine um that
  721. 27:26this this kind of square that we have uh
  722. 27:28is a zoomed in piece of of this teapot
  723. 27:31um you could imagine an algorithm where
  724. 27:33we just walk one at a time left to right
  725. 27:35um for each pixel inside this this
  726. 27:37Square we run the program um and that
  727. 27:40would that would work but it's obviously
  728. 27:42very slow because it does everything one
  729. 27:43at a time and so the idea of of a simd
  730. 27:48uh simd machine is that you have groups
  731. 27:51that work in in lock lock lock step
  732. 27:54excuse me so you might just as a simple
  733. 27:57example take groups of four a 2X two
  734. 27:59subsquare inside this big square and
  735. 28:01you're going to execute the same
  736. 28:02instruction on all four of those pixels
  737. 28:04um at the same time and in a simple case
  738. 28:07where you're just um executing simple
  739. 28:08instructions they're just going to all
  740. 28:10execute one two three four in in lock
  741. 28:13stck lock step and finish at the same
  742. 28:16time now you can get into more
  743. 28:19complicated examples and I'm going to
  744. 28:20gloss over this uh a little bit where
  745. 28:23you might have one of those pixels has
  746. 28:25to run a high latency texture fetch or
  747. 28:27or latency operation and what that means
  748. 28:29is you have to stall all four of those
  749. 28:31pixels so you kind of take them off um
  750. 28:33you stall them that's what's intended to
  751. 28:35show in this diagram and you go find a
  752. 28:37different set of of work to go to go run
  753. 28:41so you might find a different um thread
  754. 28:43group that you can go execute can make
  755. 28:46progress while this High latency
  756. 28:47operation is is running um and that's
  757. 28:50where a lot of the complexity and
  758. 28:53frankly the interesting parts of the GPU
  759. 28:56uh comes in because you're managing
  760. 28:59these thousands of threads that are
  761. 29:00running within the machine and trying to
  762. 29:02optimally execute them in a in a way
  763. 29:04that fills the machine um in in the best
  764. 29:06Manner and that's that's a really hard
  765. 29:08Pro really hard
  766. 29:13problem I'm going to talk about um
  767. 29:15texture mapping briefly so I have time
  768. 29:17for uh to to get to the demo so I I kind
  769. 29:20of mentioned at the beginning um in the
  770. 29:22background of our our our Pony we had
  771. 29:25these fireworks we also had these M that
  772. 29:26were showing up on these surfaces is
  773. 29:28well this background is not generated
  774. 29:31programmatically or on the fly as part
  775. 29:33of the the um the the rendering
  776. 29:36algorithm these are pre-stored images or
  777. 29:38or textures um and and what you can do
  778. 29:41is when you have pre-stored images like
  779. 29:42this like these fireworks you can map
  780. 29:44them into your your rendered image and
  781. 29:48so you can imagine there's just a big
  782. 29:49triangle or sorry a big Square back here
  783. 29:51um and what you would do is um map each
  784. 29:56set of pixels in in your world into this
  785. 29:59texture and and go fetch them um and
  786. 30:02that's basically what the texture unit
  787. 30:03or what texture mapping does is it Maps
  788. 30:05um onscreen pixels to pixels that are
  789. 30:08inside a texture now um that can
  790. 30:11actually create some very interesting
  791. 30:12and challenging sample problems you
  792. 30:14could imagine um if if this image here
  793. 30:18were very small compared to the uh the
  794. 30:21size that you wanted to render in on
  795. 30:23screen you've you've got a problem where
  796. 30:25how do I sample the right number picture
  797. 30:27uh picture pixs for this to to show up
  798. 30:29and likewise you can have the opposite
  799. 30:31problem where you have a very large
  800. 30:33image that you're um sampling into a
  801. 30:35very small area on screen so there's a
  802. 30:36lot of Hardware built into texture units
  803. 30:39to try to do a good job of of sampling
  804. 30:41and this is kind of an example that's
  805. 30:42showing that you're taking um a a square
  806. 30:46kind of checkerboard pattern and then
  807. 30:47you're mapping it to something that is
  808. 30:49that is tall and skinny you can see the
  809. 30:51results of doing that mapping is your
  810. 30:53nice Square image is now a number a
  811. 30:55number of rectangles and that's
  812. 30:56basically what the texture C is going to
  813. 30:58going to do this is kind of a more
  814. 31:01advanced it's called anisotropic
  815. 31:02filtering uh filter mode that's shown on
  816. 31:04the right hand side that kind of avoids
  817. 31:07some of these uh weird artifacts that
  818. 31:09you see on the top screen um but that's
  819. 31:11uh that's probably a bit beyond today's
  820. 31:13presentation and that's the
  821. 31:14anti-aliasing that we mentioned before
  822. 31:17yes that is a version of of
  823. 31:22anti-aliasing okay so our last stage we
  824. 31:25need to actually get this get all these
  825. 31:27pixels that we've we worked on um and
  826. 31:29get them written out to to to memory um
  827. 31:32and so uh there's usually an instruction
  828. 31:35not necessar exposed to developer but at
  829. 31:37some point you've got all these pixels
  830. 31:38usually in a cache on your chip you need
  831. 31:40to actually write them out to the frame
  832. 31:41ruffer um and then tell the frame buffer
  833. 31:44or tell the display usually hey my image
  834. 31:46my rendering is done go pick up all this
  835. 31:49data and go go put it on the display and
  836. 31:51that's that's kind of what the the
  837. 31:52output stage is um I there's some other
  838. 31:55steps in there there's blend tests um
  839. 31:58there's some stencil test that I kind of
  840. 32:00glossed glossed over a little bit but
  841. 32:02that's that's really the uh the gist of
  842. 32:05it and what you can see on this the
  843. 32:07bottom right hand side of this picture
  844. 32:09is this is kind of a frame buffer that's
  845. 32:11been captured like mid render so some of
  846. 32:14the uh the tiles or some of the pixels
  847. 32:17have completed you can see them on the
  848. 32:18screen others have have not so you so
  849. 32:22you need basically need to wait till all
  850. 32:23of the work is completed before you say
  851. 32:25that you're ready for for display
  852. 32:30all right so um let's get to our example
  853. 32:34I'm trying to leave uh about 10 uh 10
  854. 32:37minutes for questions at the end so um
  855. 32:40what I want to show you guys is a um sax
  856. 32:44piie algorithm run on um a CPU and and a
  857. 32:49GPU and and actually it's even simpler
  858. 32:52than than that um we're simply going to
  859. 32:55uh Implement a function um in C
  860. 32:58that takes a um an
  861. 33:01array uh X of of floats multiplies each
  862. 33:05element in that uh array by a float a
  863. 33:09and outputs it to an output array Y and
  864. 33:12so we're just going to Loop through the
  865. 33:13entire array and and do that super
  866. 33:15simple we're not even doing the plus
  867. 33:17part so we're just doing the the the
  868. 33:19multiply part um and so we're going to
  869. 33:22do that in C on a single threaded
  870. 33:24CPU and we're going to do a similar
  871. 33:26thing on on the GPU in metal and so this
  872. 33:30kind of um code snippet is may be less
  873. 33:34familiar to you guys than the C program
  874. 33:36but it essentially does does the same
  875. 33:37thing I've got an array X and output
  876. 33:40buffer Y and I'm going to multiply them
  877. 33:41by a so it does the same function just a
  878. 33:43different
  879. 33:45language um and so what does what does
  880. 33:48that look like and so and I'm going to
  881. 33:50try to do this interactively in in a
  882. 33:52second um but basically this what the
  883. 33:55results I'm showing you here are going
  884. 33:57through a number of different buffer
  885. 33:59sizes um and and so it's really uh I
  886. 34:03squared so you can see the buffer sizes
  887. 34:05here um in the next column is the the
  888. 34:09time in it took for the GPU to finish
  889. 34:12the computation um and then the last one
  890. 34:15last column is the time that it takes
  891. 34:16the CPU to complete the comp computation
  892. 34:19so there's a few things that you'll uh
  893. 34:21you'll notice as you look at this data
  894. 34:24um as you get to the bottom very clearly
  895. 34:27once as your your buffer sizes are
  896. 34:29getting very large um the GPU is clearly
  897. 34:32beating the um the CPU in terms of
  898. 34:35performance it's taking less time to
  899. 34:37complete these operations um and the
  900. 34:40crossover point is where is it it's kind
  901. 34:43of right right around this this area
  902. 34:46where we see that the um the GPU starts
  903. 34:48to be beat the CPU you'll notice that
  904. 34:51the CPU does much much better um on the
  905. 34:55smaller buffer sizes than the GPU and
  906. 34:58really what what that tells you is that
  907. 35:00the GPU is not great at doing kind of
  908. 35:03small chunks of work and the reason for
  909. 35:05that is there's a lot of overhead in
  910. 35:07getting the GPU set up and started
  911. 35:10particularly on this this platform um to
  912. 35:13setting it up to run the computation and
  913. 35:15and so you can see when I go through the
  914. 35:17first kind of five or six runs of this
  915. 35:20the time it takes is is more less the
  916. 35:22same in facts it's oscillating up and
  917. 35:23down a bit due you know other random
  918. 35:25stuff going on in the system um
  919. 35:28but once we get down to about here um
  920. 35:31the time starts increasing so what that
  921. 35:32tells you is there's some fixed GPU
  922. 35:34overhead that's going on here um and the
  923. 35:36CPU is way better and so the reason for
  924. 35:39that is is pretty simple we're not
  925. 35:40getting any advantage of parallel
  926. 35:43processing for these very small buffers
  927. 35:45and so for very small simple jobs your
  928. 35:47CPU is going to do just fine um once you
  929. 35:49get up to jobs that are millions or
  930. 35:51hundreds of millions of elements in size
  931. 35:54um the GPU is going to do better and so
  932. 35:56why is that well that goes goes back to
  933. 35:58our our what we've been talking about
  934. 36:00throughout this this talk simd
  935. 36:01processing so if you go back and I and I
  936. 36:04will go back to our example um the CPU
  937. 36:08is basically doing um what we talked
  938. 36:11about on the left hand side here it's
  939. 36:13going literally through the array
  940. 36:15running this very simple calculation one
  941. 36:17at a time and so when this the the
  942. 36:19number of entries that it needs to do
  943. 36:20the calculation gets very large it's
  944. 36:22going to take a long time the the GPU is
  945. 36:26is actually executing
  946. 36:28um elements of of that buffer in groups
  947. 36:31of of actually more than four probably
  948. 36:3332 or 64 or even even larger um such
  949. 36:37that we're executing it in let's call it
  950. 36:3964 64 size chunks and so what that means
  951. 36:43is that the GPU is going to is is going
  952. 36:46to execute those instructions much much
  953. 36:48faster and we start to see that as the
  954. 36:49buffer size gets gets very very
  955. 36:54large so um I'm going to attempt
  956. 36:58to show this
  957. 37:00live and I'm I'm also running zoom and
  958. 37:03sharing Zoom so we'll see we'll see how
  959. 37:05this goes but um what what you'll be
  960. 37:08able to
  961. 37:09see assuming my computer doesn't crash
  962. 37:12um is is this program executed let me
  963. 37:15make it make it
  964. 37:18larger uh is is executing this pretty uh
  965. 37:21executing this live during the talk um
  966. 37:24and you can see how hopefully the the
  967. 37:26GPU performance perance starts to
  968. 37:28outstrip the the CPU as we
  969. 37:32go and there it is GPU is starting to to
  970. 37:34beat the
  971. 37:36CPU CPU times getting getting longer and
  972. 37:40longer we're about 2x the GPU right
  973. 37:44now about
  974. 37:463x I can hear my fan going up there's a
  975. 37:51there's a question uh whether this is
  976. 37:52being done on the new M1 or whether that
  977. 37:55would change things I wish it was doing
  978. 37:56being done on M1 um no I I oddly my uh
  979. 38:01my wife has one of those because I
  980. 38:02bought one for her but my the Mac I'm
  981. 38:04running this on for work uh I believe is
  982. 38:07a
  983. 38:082018 uh Intel based system I don't know
  984. 38:10the exact uh exact number but it's a
  985. 38:13it's a Intel integrated GPU um and then
  986. 38:16you can see by the time we we got done
  987. 38:19here um that the the GPU was almost
  988. 38:22seven times faster than uh the CPU for
  989. 38:25the largest the largest buffer
  990. 38:31um now this was a a very simple program
  991. 38:35um and in fact uh an astute uh engineer
  992. 38:39might notice that hey this thing is
  993. 38:40actually not doing that much calculation
  994. 38:43um so it actually is is is at least on
  995. 38:46the GPU memory limited we talked about
  996. 38:48memory band with with earlier the GPU is
  997. 38:51much better at moving chunks large
  998. 38:53chunks of of memory around than than the
  999. 38:55CPU so we we are also benefiting from
  1000. 38:58not just the number of execution units
  1001. 38:59but the amount of um the amount of
  1002. 39:02memory bandwidth that we have now it
  1003. 39:04turns out I think on this particular CPU
  1004. 39:06that I'm running on um because I've
  1005. 39:09we've explicitly said it to operate in a
  1006. 39:12single core single thread manner um the
  1007. 39:15CPU core is likely Mass limited um Al I
  1008. 39:18have not actually tested
  1009. 39:22that all right so um so here's a graph
  1010. 39:25of what we just did this is obviously
  1011. 39:27not the experiment I just ran but you
  1012. 39:29can see the results were pretty similar
  1013. 39:30so I ran this over the weekend without
  1014. 39:32zoom and other stuff running in the
  1015. 39:34background this is a logarithmic scale
  1016. 39:37um so so notice that carefully um of of
  1017. 39:41where we of our array size on the bottom
  1018. 39:43with the runtime and milliseconds on on
  1019. 39:45the Y AIS you can kind of see the
  1020. 39:47crossover Point here and then again
  1021. 39:49remembering that this is a logarithmic
  1022. 39:51scale as we get up to these very large
  1023. 39:52buffer sizes um the the um the uh the
  1024. 39:56GPU is is beating the uh the CPU
  1025. 40:00performance
  1026. 40:03handily all right so let's wrap it up
  1027. 40:05and then and leave some time for for
  1028. 40:07questions so um so if if you get nothing
  1029. 40:10else from from this presentation um
  1030. 40:13please take away these these through
  1031. 40:14points so getting efficiency and
  1032. 40:17performance out of a GPU is all about
  1033. 40:20maximizing parallelism maximizing
  1034. 40:22throughput via parallelism you're T you
  1035. 40:24want to run parallel programs where
  1036. 40:26you're you're doing Sim the op
  1037. 40:27operations you're running the same same
  1038. 40:29instructions across multiple data um
  1039. 40:31that's What GPU performance is about
  1040. 40:33there's you know some fixed function
  1041. 40:35stuff on the side to do specialized math
  1042. 40:37like like texturing um but really what
  1043. 40:40we optimize around is M maximizing
  1044. 40:43performance through parallelism that
  1045. 40:45makes um uh like image processing for
  1046. 40:50for uh for cameras for doing
  1047. 40:52computational photography for um running
  1048. 40:56uh algorithms to do image detection so
  1049. 40:58um neural
  1050. 41:00networks to be done on on the GPU just
  1051. 41:03as Tesla just
  1052. 41:06exactly um and and gpus modern gpus are
  1053. 41:09very easily expanded to do general
  1054. 41:11purpose compute so companies like like
  1055. 41:13Invidia and AMD have been very
  1056. 41:15successful selling gpus not just for for
  1057. 41:18games but for um for building compute
  1058. 41:21Farms of gpus for doing um machine
  1059. 41:24learning algorithms and and things like
  1060. 41:26that so that's a that's a big growth uh
  1061. 41:29factor for um for gpus in in recent
  1062. 41:33years so with that I'll stop there hope
  1063. 41:36you guys found this interesting um let's
  1064. 41:39uh do some questions in the last few
  1065. 41:40minutes that was great James thank you
  1066. 41:43much for thank you thank you wonderful
  1067. 41:44wonderful we've got a couple questions
  1068. 41:46that are open um the first is um what's
  1069. 41:51the software that's actually managing
  1070. 41:53the coherency in the GPU um would that
  1071. 41:56be something built into the o or how do
  1072. 41:58you manage the coherency making making
  1073. 41:59sure you're worry about um thread
  1074. 42:02stepping on each other that's an
  1075. 42:03excellent question um the the answer is
  1076. 42:05it's really a combination of the
  1077. 42:07application and the driver that is that
  1078. 42:09is programming and managing the lower
  1079. 42:11level details of of the GPU so um if you
  1080. 42:14went and looked at a modern GPI GPU API
  1081. 42:17such as metal for example um there are
  1082. 42:20facilities inside the metal programming
  1083. 42:22language that give you mechanisms For
  1084. 42:25Thread synchronization and memory
  1085. 42:26synchronization
  1086. 42:28um and those kind of high highle
  1087. 42:30facilities get mapped to lowlevel
  1088. 42:32Hardware commands like a cash flush or a
  1089. 42:35fence or a barrier um that the driver
  1090. 42:38will convert into something that the
  1091. 42:39hardware can actually understand um and
  1092. 42:42so to kind of perhaps bring that back to
  1093. 42:44the question the programmer does have to
  1094. 42:46be aware of some of these
  1095. 42:48synchronization requirements and how to
  1096. 42:50how to manage parallel programming um in
  1097. 42:53combination with the driver because if
  1098. 42:55you're not aware of those things you're
  1099. 42:56not going to get good
  1100. 42:58makes sense on this kind of the same
  1101. 43:00topic do you need to worry about
  1102. 43:01deadlock in
  1103. 43:03gpus um the the hardware team meaning my
  1104. 43:07team certainly does um as a as a at a
  1105. 43:11programming level um you need to worry
  1106. 43:14about it kind of from the similar
  1107. 43:16perspective as you would worry about
  1108. 43:18hanging a a CPU by writing a an infinite
  1109. 43:21Loop or or two Loops especially if
  1110. 43:23you're doing parallel programming that
  1111. 43:24kind of depend each other depend on each
  1112. 43:26other in ways that could create
  1113. 43:29Deadlocks um if you get down to the
  1114. 43:32hardware implementation details there
  1115. 43:34are um there are actually all sorts of
  1116. 43:38uh hazards that can cause Deadlocks that
  1117. 43:40we spend a lot of uh design and
  1118. 43:43verification time trying to find and
  1119. 43:45either just avoid architecturally or or
  1120. 43:48put in deadlock Breakers you may call
  1121. 43:51them into the hardware so the hardware
  1122. 43:53can't lock itself up um so yes it's it's
  1123. 43:56something that we we are very worried
  1124. 43:58about Bo these are coming as as you
  1125. 44:01answered one four more questions show up
  1126. 44:03stud students students are doing great
  1127. 44:05we got 70 students here um what's the
  1128. 44:08ideal relationship or is there an ideal
  1129. 44:10relationship between the number of
  1130. 44:11pixels and the number of execution units
  1131. 44:13in a
  1132. 44:13GPU like maybe as you go to 6K 8K how
  1133. 44:16does that you there's some there's some
  1134. 44:18limits where gpus can't drive an 8K
  1135. 44:20screen that's just talking on the screen
  1136. 44:21but in terms of the computation is there
  1137. 44:24a you know mapping at all or maybe a
  1138. 44:26good
  1139. 44:28that's a that's a good question um the
  1140. 44:30the nice thing about GPU processing um
  1141. 44:33is that it scales very very well um and
  1142. 44:36in particular uh if you're if you're
  1143. 44:38doing like super high resolution 8K
  1144. 44:41displays um then you're going to want
  1145. 44:44especially if you're running games and
  1146. 44:46you want like high frame rates you're
  1147. 44:47GNA want a big GPU to drive that now um
  1148. 44:52there's no kind of fixed relationship
  1149. 44:53you could have a small modern GPU drive
  1150. 44:55a 4K or an 8K display um you're just
  1151. 44:58going to get less performance out of out
  1152. 45:00of your game uh if you were playing a
  1153. 45:02game and you may not get 30 frames per
  1154. 45:03second so your game look choppy and not
  1155. 45:06not play well so there's no kind of like
  1156. 45:08fixed requirement if you will um but
  1157. 45:11certainly if you're going to run a a
  1158. 45:13complex modern game on a very high
  1159. 45:15resolution display you're going to want
  1160. 45:16a bigger GPU with more cores to get a
  1161. 45:18good gaming experience so hopefully that
  1162. 45:20kind of and I and I don't get this user
  1163. 45:23perception of 120 frames per second I I
  1164. 45:26don't think my can do
  1165. 45:28it yes there's a lot of debates about 30
  1166. 45:3160 or is 120 hertz even even useful well
  1167. 45:35it's in film too I mean was it load of
  1168. 45:37the Rings had like 60 or 120 and it was
  1169. 45:39like wait this is looks like TV
  1170. 45:40something looks a little different here
  1171. 45:42exactly um this is a great question from
  1172. 45:44Adrian what are some surprising
  1173. 45:46applications where gpus are used in
  1174. 45:47places we might not expect oh wow that's
  1175. 45:50a that's a fantastic question
  1176. 45:53yeah uh I'm not sure if this would be
  1177. 45:56surprising or not these days but um if
  1178. 45:58you've looked at uh if you drive a Tesla
  1179. 46:01or been in a Tesla uh and I I'm not
  1180. 46:03revealing anything secret about Tesla I
  1181. 46:05don't know anything secret about Tesla
  1182. 46:07so just get that out there um but my
  1183. 46:10understanding is they use uh GPU
  1184. 46:11Hardware to do a lot of their uh their
  1185. 46:14image processing algorithms for for
  1186. 46:17self-driving um Soh so you'll find
  1187. 46:20you'll find gpus in in automobiles
  1188. 46:23you'll find uh gpus in a lot of places
  1189. 46:25you wouldn't necessarily expect them
  1190. 46:27a few a few years ago similarly on um
  1191. 46:30any modern phone and I'm not talking
  1192. 46:32specifically about Apple but any modern
  1193. 46:34phone is going to do a fair amount of of
  1194. 46:38computational Photography these days and
  1195. 46:40a lot of the phph run through the GPU to
  1196. 46:44generate or to blend images together to
  1197. 46:47generate fin image right so photography
  1198. 46:49is not simply about like hey how how
  1199. 46:52many megapixels are in my sensor um it's
  1200. 46:55it's how good are the algorithms that
  1201. 46:57you are running and how much GPU and
  1202. 47:00other horsepower do you have to run
  1203. 47:03those algorithms such as the person
  1204. 47:04taking the picture can see it in in less
  1205. 47:06than a second so it's interactive
  1206. 47:08imagine if you were taking a picture and
  1207. 47:10it took 10 seconds for the final final
  1208. 47:12picture to show up that that's not very
  1209. 47:14fun for for a a phone type use case so
  1210. 47:17so gpus matter a lot in in that in that
  1211. 47:20space as well especially with HDR
  1212. 47:22imagery I mean you can you can compress
  1213. 47:24that you know tonal tone mapping where
  1214. 47:25you have a bright part of the screen
  1215. 47:27dark part of the screen but the GPU can
  1216. 47:29take many pictures simultaneously or SE
  1217. 47:31in sequence and then compress them down
  1218. 47:33into one that just captures the highs
  1219. 47:34and lows to reduce dynamic range I love
  1220. 47:36it um I knew this by the way get get one
  1221. 47:41you know there's one surprising thing uh
  1222. 47:43it's Bitcoin mining Bitcoin um you know
  1223. 47:46Bitcoin mining
  1224. 47:47basically because of that you a few
  1225. 47:50years ago it was all L gpus have been
  1226. 47:55sold out I mean it was not POs to buy
  1227. 47:58a they're all that's fascinating I
  1228. 48:01didn't know that yeah now you can buy
  1229. 48:04cheap dedicated Hardware that does these
  1230. 48:06hash functions more efficiently
  1231. 48:10that's three years ago I think it was
  1232. 48:12December when when Bitcoin was taking
  1233. 48:14off you could not find a highend GPU
  1234. 48:16people were were selling them for
  1235. 48:18thousands of dollars where people to go
  1236. 48:19Bitcoin mining so that's aample and by
  1237. 48:22the way I just looking at the the the
  1238. 48:23latest Mac Pro the highend gpus are like
  1239. 48:26$2,000 just by themselves before the
  1240. 48:28before the price you know increase
  1241. 48:30because people want to you know buy them
  1242. 48:31to do these things the high GPS can be
  1243. 48:33very expensive um this is Adam a great
  1244. 48:35question I knew this was going to come
  1245. 48:36up this is so cool how can we students
  1246. 48:39play with some graphics image processing
  1247. 48:41code how do we code to use a GPU
  1248. 48:43ourselves I first answer is get a Mac y
  1249. 48:47get a Mac um we have metal uh Apple has
  1250. 48:51metal programming guides and very simple
  1251. 48:53examples of just how to get started if
  1252. 48:54you just Google I should put a link in
  1253. 48:56here but if you just Google uh metal
  1254. 48:58programming like introduction um there's
  1255. 49:01some online tutorials that you can go
  1256. 49:02through and go like Hey how do I write
  1257. 49:04like my the hell world of of gpus using
  1258. 49:07metal and I'm sure Nvidia and AMD have
  1259. 49:10the same thing for for their Hardware
  1260. 49:12using other other languages so there's
  1261. 49:14definitely tutorials out there and
  1262. 49:16there's also I think um Cuda and open CL
  1263. 49:19is that right those two those those are
  1264. 49:23programming apis Auda is specific to to
  1265. 49:26Nvidia
  1266. 49:27an open CL um I'm not sure how much that
  1267. 49:29is used as much these days that was more
  1268. 49:31of an open open API crossplatform I
  1269. 49:34think it's been replaced by Cuda and
  1270. 49:36others right he's I I'll just add to the
  1271. 49:39question which is the code that you
  1272. 49:40showed in metal you can be running on a
  1273. 49:43built-in CPU on a laptop or a high-end
  1274. 49:45Mac Pro with a dedicated two boxes of a
  1275. 49:47CPU do you need to worry about that at
  1276. 49:49that level at the metal level it just
  1277. 49:50just as handled by by the by the by the
  1278. 49:52abstraction layer right that's as long
  1279. 49:54as you write in a way that's parallel
  1280. 49:56programming friendly um it doesn't
  1281. 49:59matter if you're targeting like
  1282. 50:00something built into a phone or if
  1283. 50:02you're targeting a giant GPU system the
  1284. 50:05the API takes care of that for you
  1285. 50:07that's wonderful now unfortunately I've
  1286. 50:08got I think 14 questions left but we
  1287. 50:10promis chemistry which actually needs
  1288. 50:11this this webinar like now we're now two
  1289. 50:13minutes into their webinar so folks
  1290. 50:15let's let's give James a hand thank you
  1291. 50:17so much for coming and sharing those
  1292. 50:19thoughts about the GPU thank you James
  1293. 50:21thank you James wonderful wonderful
  1294. 50:23Wonder so grateful to have you here
  1295. 50:25thanks thanks again standing outstanding
  1296. 50:27lecture there was a meta question about
  1297. 50:29these slides confidential we we don't
  1298. 50:31have the copies of the slides but we
  1299. 50:32have the we have the webinar so we'll
  1300. 50:33just keep that and they'll just be able
  1301. 50:35to watch the slides through the webinar
  1302. 50:36I'm guessing y okay perfect perfect
  1303. 50:39thanks so much all right take care
  1304. 50:41everybody and stay here if you want to
  1305. 50:42watch leure okay take care everybody
  1306. 50:45thanks again James stay safe folks
  1307. 50:47wonderful all right bye bye

About this transcript

This page contains the full transcript of [CS61C FA20] Lecture 39.LIVE - GPU Guest Lecture with James Percy  by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 9,594 words across 1,307 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.