YouTube2Text

How do Graphics Cards Work? Exploring GPU Architecture — Transcript

by Branch Education · 3,903 words · 246 segments · language en · Watch on YouTube

Full transcript

  1. 0:00How many calculations do you think your  graphics card performs every second
  2. 0:04while running video games with incredibly  realistic graphics? Maybe 100 million? Well,
  3. 0:11100 million calculations a second is what’s  required to run Mario 64 from 1996. We need
  4. 0:21more power. Maybe 100 billion calculations a  second? Well, then you would have a computer
  5. 0:27that could run Minecraft back in 2011. In order  to run the most realistic video games such as
  6. 0:34Cyberpunk 2077 you need a graphics card that can  perform around 36 trillion calculations a second.
  7. 0:43This is an unimaginably large number, so let’s  take a second to try to conceptualize it. Imagine
  8. 0:50doing a long multiplication problem once every  second. Now let’s say everyone on the planet does
  9. 0:57a similar type of calculation but with different  numbers. To reach the equivalent computational
  10. 1:03power of this graphics card and its 36 trillion  calculations a second we would need about 4,400
  11. 1:12Earths filled with people, all working together  and completing one calculation each every second.
  12. 1:20It’s rather mind boggling to think that a  device can manage all these calculations,
  13. 1:25so in this video we’ll see how graphics cards work  in two parts. First, we’ll open up this graphics
  14. 1:32card and explore the different components inside,  as well as the physical design and architecture
  15. 1:38of the GPU or graphics processing unit. Second,  we’ll explore the computational architecture and
  16. 1:46see how GPUs process mountains of data, and why  they’re ideal for running video game graphics,
  17. 1:53Bitcoin mining, neural networks and AI. So, stick around and let’s jump right in.
  18. 2:08This video is sponsored by  Micron which manufactures
  19. 2:11the graphics memory inside this graphics card. Before we dive into all the parts of the GPU,
  20. 2:19let’s first understand the differences between  GPUs and CPUs. Inside this graphics card,
  21. 2:26the Graphics Processing Unit or GPU has over  10,000 cores. However, when we look at the
  22. 2:33CPU or Central Processing Unit that’s mounted to  the motherboard, we find an integrated circuit or
  23. 2:40chip with only 24 cores. So, which one is more  powerful? 10 thousand is a lot more than 24,
  24. 2:48so you would think the GPU is more powerful,  however, it’s more complicated than that.
  25. 2:55A useful analogy is to think of a GPU as a massive  cargo ship and a CPU as a jumbo jet airplane.
  26. 3:03The amount of cargo capacity is the amount of  calculations and data that can be processed,
  27. 3:09and the speed of the ship or airplane is the  rate at which how quickly those calculations
  28. 3:15and data are being processed. Essentially,  it’s a trade-off between a massive number
  29. 3:21of calculations that are executed at a  slower rate versus a few calculations
  30. 3:27that can be performed at a much faster rate. Another key difference is that airplanes are a
  31. 3:32lot more flexible since they can carry passengers,  packages, or containers and can take off and land
  32. 3:39at any one of tens of thousands of airports.  Likewise CPUs are flexible in that they can run
  33. 3:46a variety of programs and instructions. However,  giant cargo ships carry only containers with bulk
  34. 3:53contents inside and are limited to traveling  between ports. Similarly, GPUs are a lot
  35. 4:00less flexible than CPUs and can only run simple  instructions like basic arithmetic. Additionally
  36. 4:08GPUs can’t run operating systems or interface  with input devices or networks. This analogy
  37. 4:15isn’t perfect, but it helps to answer the question  of “which is faster, a CPU or a GPU?”. Essentially
  38. 4:24if you want to perform a set of calculations  across mountains of data, then a GPU will be
  39. 4:30faster at completing the task. However, if you  have a lot less data that needs to be evaluated
  40. 4:36quickly than a CPU will be faster. Furthermore,  if you need to run an operating system or support
  41. 4:43network connections and a wide range of different  applications and hardware, then you’ll want a CPU.
  42. 4:50We’re planning a separate video on CPU  architecture, so make sure to subscribe
  43. 4:55so you don’t miss it, but let’s now dive into this  graphics card and see how it works. In the center
  44. 5:02of this graphics card is the printed circuit  board or PCB, with all the various components
  45. 5:08mounted on it, [Animator Note: Highlight and list  out the various parts that will be covered.] and
  46. 5:10we’ll start by exploring the brains which is  the graphics processing unit or GPU. When we
  47. 5:17open it up, we find a large chip or die named  GA102 built from 28.3 billion transistors. The
  48. 5:26majority of the area of the chip is taken up by  the processing cores which have a hierarchical
  49. 5:33organization. Specifically, the chip is divided  into 7 Graphics Processing Clusters or GPCs,
  50. 5:41and within each processing cluster are 12  streaming multiprocessors or SMs. Next,
  51. 5:48inside each of these streaming multiprocessors  are 4 warps and 1 ray tracing core, and then,
  52. 5:56inside each warp are 32 Cuda or shading cores and  1 tensor core. Across the entire GPU are 10752
  53. 6:08CUDA cores, 336 Tensor Cores, and 84 Ray Tracing  Cores. These three types of cores execute all the
  54. 6:18calculations of the GPU, and each has a different  function. CUDA cores can be thought of as simple
  55. 6:24binary calculators with an addition button, a  multiply button and a few others, and are used
  56. 6:31the most when running video games. Tensor cores  are matrix multiplication and addition calculators
  57. 6:38and are used for geometric transformations  and working with neural networks and AI. And
  58. 6:45ray tracing cores are the largest but the fewest  and are used to execute ray tracing algorithms.
  59. 6:53Now that we understand the computational  resources inside this chip, one rather interesting
  60. 6:59fact is that the 3080, 3090, 3080 ti, and 3090 ti  graphics cards all use the same GA102 chip design
  61. 7:11for their GPU. This might be counterintuitive  because they have different prices and were
  62. 7:16released in different years, but it’s true.  So, why is this? Well, during the manufacturing
  63. 7:23process sometimes patterning errors, dust  particles, or other manufacturing issues
  64. 7:29cause damage and create defective areas of the  circuit. Instead of throwing out the entire chip
  65. 7:35because of a small defect, engineers find the  defective region and permanently isolate and
  66. 7:41deactivate the nearby circuitry. By having a GPU  with a highly repetitive design, a small defect in
  67. 7:49one core only damages that particular streaming  multiprocessor circuit and doesn’t affect the
  68. 7:55other areas of the chip. As a result, these chips  are tested and categorized or binned according
  69. 8:02to the number of defects. The 3090ti graphics  cards have flawless GA102 chips with all 10752
  70. 8:13CUDA cores working properly, the 3090 has 10,496  cores working, the 3080ti has 10,240 and the 3080
  71. 8:26has 8704 CUDA cores working, which is equivalent  to having 16 damaged and deactivated streaming
  72. 8:34multiprocessors. Additionally, different graphics  cards differ by their maximum clock speed and the
  73. 8:41quantity and generation of graphics memory that  supports the GPU, which we’ll explore in a little
  74. 8:47bit. Because we’ve been focusing on the physical  architecture of this GA102 GPU chip, let’s zoom
  75. 8:54into one of these CUDA cores and see what it looks  like. Inside this simple calculator is a layout
  76. 9:00of approximately 410 thousand transistors. This  section of 50 thousand transistors performs the
  77. 9:08operation of A times B plus C which is called  fused multiply and add or FMA and is the most
  78. 9:17common operation performed by graphics cards.  Half of the CUDA cores execute FMA using 32-bit
  79. 9:24floating-point numbers, which is essentially  scientific notation, and the other half
  80. 9:29of the cores use either 32-bit integers or 32-bit  floating point numbers. Other sections of this
  81. 9:36core accommodate negative numbers and perform  other simple functions like bit-shifting and bit
  82. 9:42masking as well as collecting and queueing  the incoming instructions and operands,
  83. 9:47and then accumulating and outputting the results.  As a result, this single core is just a simple
  84. 9:55calculator with a limited number of functions.  This calculator completes one multiply and one add
  85. 10:01operation each clock cycle and therefore with this  3090 graphics cards and its 10496 cores and 1.7
  86. 10:12gigahertz clock, we get 35.6 trillion calculations  a second. However, if you’re wondering how the GPU
  87. 10:20handles more complicated operations like division,  square root, and trigonometric functions, well,
  88. 10:27these calculator operations are performed by  the special function units which are far fewer
  89. 10:33as only 4 of them can be found in each streaming  multiprocessor. Now that we have an understanding
  90. 10:38of what’s inside a single core, let’s zoom out  and take a look at the other sections of the GA102
  91. 10:45chip. Around the edge we find 12 graphics memory  controllers, the NVLink Controllers and the PCIe
  92. 10:54interface. On the bottom is a 6-megabyte Level  2 SRAM Memory Cache, and here’s the Gigathread
  93. 11:01Engine which manages all the graphics processing  clusters and streaming multiprocessors inside.
  94. 11:08Now that we’ve explored this GA102 GPU’s physical  architecture, let’s zoom out and take a look at
  95. 11:15the other parts inside the graphics card. On this  side are the various ports for the displays to be
  96. 11:22plugged into, on the other side is the incoming  12 Volt power connector, and then here are the
  97. 11:28PCIe pins that plug into the motherboard. On  the PCB, the majority of the smaller components
  98. 11:36constitute the voltage regulator module which  takes the incoming 12 volts and converts it to
  99. 11:42one point one volts and supplies hundreds  of watts of power to the GPU. Because all
  100. 11:49this power heats up the GPU, most of the weight  of the graphics card is in the form of a heat
  101. 11:54sink with 4 heat pipes that carry heat from the  GPU and memory chips to the radiator fins where
  102. 12:01fans then help to remove the heat. Perhaps some of  the most important components, aside from the GPU,
  103. 12:09are the 24 gigabytes of graphics memory chips  which are technically called GDDR6X SDRAM and
  104. 12:17were manufactured by Micron which is the sponsor  of this video. Whenever you start up a video game
  105. 12:23or wait for a loading screen, the time it takes  to load is mostly spent moving all the 3D models
  106. 12:29of a particular scene or environment from the  solid-state drive into these graphics memory
  107. 12:35chips. As mentioned earlier, the GPU has a small  amount of data storage in its 6-megabyte shared
  108. 12:41Level 2 cache which can hold the equivalent of  about this much of the video game’s environment.
  109. 12:47Therefore in order to render a video game,  different chunks of scene are continuously being
  110. 12:53transferred between the graphics memory and the  GPU. Because the cores are constantly performing
  111. 12:59tens of trillions of calculations a second,  GPUs are data hungry machines and need to be
  112. 13:06continuously fed terabytes upon terabytes of data,  and thus these graphics memory chips are designed
  113. 13:14kind of like multiple cranes loading a cargo ship  at the same time. Specifically, these 24 chips
  114. 13:21transfer a combined 384 bits at a time, which  is called the bus width and the total data that
  115. 13:28can be transferred, or the bandwidth is about 1.15  terabytes a second. In contrast the sticks of DRAM
  116. 13:37that support the CPU only have a 64-bit bus width  and a maximum bandwidth closer to 64 gigabytes a
  117. 13:44second. One rather interesting thing is that you  may think that computers only work using binary
  118. 13:50ones and zeros. However, in order to increase data  transfer rates, GDDR6X and the latest graphics
  119. 13:58memory, GDDR7 send and receive data across the bus  wires using multiple voltage levels beyond just
  120. 14:060 and 1. For example, GDDR7 uses 3 different  encoding schemes to combine binary bits into
  121. 14:14ternary digits or PAM-3 symbols with voltages of  0, 1, and negative 1. Here’s the encoding scheme
  122. 14:22on how 3 binary bits are encoded into 2 ternary  digits and this scheme is combined with an 11
  123. 14:29bit to 7 ternary digit encoding scheme resulting  in sending 276 binary bits using only 176 ternary
  124. 14:39digits. The previous generation, GDDR6X, which  is the memory in this 3090 graphics card, used a
  125. 14:47different encoding scheme, called PAM-4, to send  2 bits of data using 4 different voltage levels,
  126. 14:54however, engineers and the graphics memory  industry agreed to switch to PAM-3 for future
  127. 15:00generations of graphics chips in order to reduce  encoder complexity, improve the signal to noise
  128. 15:06ratio, and improve power efficiency. Micron  delivers consistent innovation to push the
  129. 15:15boundaries on how much data can be transferred  every second and to design cutting edge memory
  130. 15:20chips. Another advancement by Micron is the  development of HBM, or the high bandwidth memory,
  131. 15:27that surrounds AI chips. HBM is built from  stacks of DRAM memory chips and uses TSVs
  132. 15:35or through silicon vias, to connect this stack  into a single chip, essentially forming a cube
  133. 15:42of AI memory. For the latest generation of high  bandwidth memory, which is HBM3E, a single cube
  134. 15:51can have up to 24 to 36 gigabytes of memory, thus  yielding 192 gigabytes of high-speed memory around
  135. 15:59the AI chip. Next time you buy an AI accelerator  system, make sure it uses Micron’s HBM3E which
  136. 16:07uses 30% less power than the competitive products.  However, unless you’re building an AI data center,
  137. 16:15you’re likely not in the market to buy one  of these systems which cost between 25 to
  138. 16:2040 thousand dollars and are on backorder for a  few years. If you’re curious about high bandwidth
  139. 16:27memory, or Micron’s next generation of graphics  memory take a look at one of these links in the
  140. 16:33description. Alternatively, if designing the next  generation of memory chips interests you, Micron
  141. 16:40is always looking for talented scientists and  engineers to help innovate on cutting edge chips
  142. 16:46and you can find out more about working for Micron  using this link. Now that we’ve explored many of
  143. 16:53the physical components inside this graphics card  and GPU, let’s next explore the computational
  144. 17:00architecture and see how applications like video  game graphics and bitcoin mining run what’s called
  145. 17:06“embarrassingly” parallel operations. Although  it may sound like a silly name, embarrassingly
  146. 17:13parallel is actually a technical classification  of computer problems where little or no effort is
  147. 17:19needed to divide the problem into parallel tasks,  and video game rendering and bitcoin mining easily
  148. 17:27fall into this category. Essentially, GPUs solve  embarrassingly parallel problems using a principle
  149. 17:34called SIMD, which stands for single instruction  multiple data where the same instructions or steps
  150. 17:41are repeated across thousands to millions of  different numbers. Let’s see an example of how
  151. 17:48SIMD or single instruction multiple data is used  to create this 3D video game environment. As you
  152. 17:55may know already, this cowboy hat on the table is  composed of approximately 28 thousand triangles
  153. 18:02built by connecting together around 14,000  vertices, each with X, Y, and Z coordinates.
  154. 18:10These vertex coordinates are built using a  coordinate system called model space with the
  155. 18:16origin of 0,0,0 being at the center of the hat. To build a 3D world we place hundreds of objects,
  156. 18:24each with their own model space into the world  environment and, in order for the camera to be
  157. 18:29able to tell where each object is relative to  other objects, we have to convert or transform
  158. 18:36all the vertices from each separate model space  into the shared world coordinate system or world
  159. 18:43space. So, as an example, how do we convert the  14 thousand vertices of the cowboy hat from model
  160. 18:51space into world space? Well, we use a single  instruction which adds the position of the origin
  161. 18:57of the hat in world space to the corresponding  X,Y, and Z coordinate of a single vertex in
  162. 19:04model space. Next we copy this instruction to  multiple data, which is all the remaining X,Y,
  163. 19:11and Z coordinates of the other thousands of  vertices that are used to build the hat. Next,
  164. 19:17we do the same for the table and the rest of  the hundreds of other objects in the scene,
  165. 19:22each time using the same instructions but with  the different objects’ coordinates in world space,
  166. 19:29and each objects’ thousands of vertices in model  space. As a result, all the vertices and triangles
  167. 19:36of all the objects are converted to a common  world space coordinate system and the camera
  168. 19:42can now determine which objects are in front  and which are behind. This example illustrates
  169. 19:48the power of SIMD or single instruction multiple  data and how a single instruction is applied to
  170. 19:545,629 different objects with a total of 8.3  million vertices within the scene resulting
  171. 20:02in 25 million addition calculations. The key to  SIMD and embarrassingly parallel programs is that
  172. 20:09every one of these millions of calculations  has no dependency on any other calculation,
  173. 20:15and thus all these calculations can be distributed  to the thousands of cores of the GPU and completed
  174. 20:22in parallel with one another. It's important to  note that vertex transformation from model space
  175. 20:28to world space is just one of the first steps of  a rather complicated video game graphics rendering
  176. 20:34pipeline and we have a separate video that delves  deeper into each of these other steps. Also,
  177. 20:41we skipped over the transformations for the  rotation and scale of each object, but factoring
  178. 20:46in these values is a similar process that requires  additional SIMD calculations. Now that we have a
  179. 20:53simple understanding of SIMD, let’s discuss how  this computational architecture matches up with
  180. 20:59the physical architecture. Essentially, each  instruction is completed by a thread and this
  181. 21:05thread is matched to a single CUDA core. Threads  are bundled into groups of 32 called warps,
  182. 21:11and the same sequence of instructions is issued to  all the threads in a warp. Next warps are grouped
  183. 21:18into thread blocks which are handled by the  streaming multiprocessor. And then finally thread
  184. 21:24blocks are grouped into grids, which are computed  across the overall GPU. All these computations are
  185. 21:32managed or scheduled by the Gigathread Engine,  which efficiently maps thread blocks to the
  186. 21:38available streaming multiprocessors. One important  distinction is that within SIMD architecture,
  187. 21:44all 32 threads in a warp follow the same  instructions and are in lockstep with each
  188. 21:50other, kind of like a phalanx of soldiers moving  together. This lock step execution applied to GPUs
  189. 21:57up until around 2016. However, newer GPUs follow  a SIMT architecture or single instruction multiple
  190. 22:06threads. The difference between SIMD and SIMT is  that while both send the same set of instructions
  191. 22:13to each thread, with SIMT, the individual  threads don’t need to be in lockstep with
  192. 22:19each other and can progress at different rates.  In technical jargon, each thread is given its own
  193. 22:25program counter. Additionally, with SIMT all the  threads within a streaming multiprocessor use a
  194. 22:32shared 128 kilobyte L1 cache and thus data that’s  output by one thread can be subsequently used by
  195. 22:41a separate thread. This improvement from SIMD to  SIMT allows for more flexibility when encountering
  196. 22:48warp divergence via data-dependent conditional  branching and easier reconvergence for the threads
  197. 22:55to reach the barrier synchronization. Essentially  newer architectures of GPUs are more flexible and
  198. 23:02efficient especially when encountering branches  in code. One additional note is that although
  199. 23:08you may think that the term warp is derived from  warp drives, it actually comes from weaving and
  200. 23:14specifically the Jacquard Loom. This loom from  1804 used programmable punch cards to select
  201. 23:22specific threads out of a set to weave together  intricate patterns. As fascinating as looms are,
  202. 23:29let’s move on. The final topics we’ll explore are  bitcoin mining, tensor cores and neural networks.
  203. 23:37But first we’d like to ask you to ‘like’  this video, write a quick comment below,
  204. 23:42share it with a colleague, friend or on social  media, and subscribe if you haven’t already.
  205. 23:49The dream of Branch Education is to make free  and accessible, visually engaging educational
  206. 23:55videos that dive deeply into a variety topics on  science, engineering, and how technology works,
  207. 24:02and then to combine multiple videos into an  entirely free engineering curriculum for high
  208. 24:08school and college students. Taking a few seconds  to like, subscribe, and comment below helps us
  209. 24:15a ton! Additionally, we have a Patreon page  with AMAs and behind the scenes footage, and,
  210. 24:22if you find what we do useful, we would appreciate  any support. Thank you. So now that we’ve explored
  211. 24:31how single instruction multiple threads is used  in video games, let’s briefly discuss why GPUs
  212. 24:38were initially used for mining bitcoin. We’re not  going to get too far into the algorithm behind
  213. 24:44the blockchain and will save it for a separate  episode, but essentially, to create a block on
  214. 24:50the blockchain, the SHA-256 hashing algorithm is  run on a set of data that includes transactions,
  215. 24:57a time stamp, additional data, and a random number  called a nonce. After feeding these values through
  216. 25:04the SHA-256 hashing algorithm a random 256-bit  value is output. You can kind of think of this
  217. 25:11algorithm as a lottery ticket generator where you  can’t pick the lottery number, but based on the
  218. 25:17input data, the SHA-256 algorithm generates  a random lottery ticket number. Therefore,
  219. 25:24if you change the nonce value and keep the rest  of the transaction data the same, you’ll generate
  220. 25:29a new random lottery ticket number. The winner of  this bitcoin mining lottery is the first randomly
  221. 25:35generated lottery number to have the first 80 bits  all zeroes, while the rest of the 176 values don’t
  222. 25:42matter and once a winning bitcoin lottery ticket  is found, the reward is 3 bitcoin and the lottery
  223. 25:49resets with a new set of transactions and input  values. So, why were graphics cards used? Well,
  224. 25:56GPUs ran thousands of iterations of the SHA-256  algorithm with the same transactions, timestamp,
  225. 26:05other data, but, with different nonce values.  As a result, a graphics card like this one could
  226. 26:11generate around 95 million SHA-256 hashes or 95  million randomly numbered lottery tickets every
  227. 26:19second, and hopefully one of those lottery numbers  would have the first 80 digits as all zeros.
  228. 26:27However, nowadays computers filled with ASICs or  application specific integrated circuits perform
  229. 26:34250 trillion hashes a second or the equivalent  of 2600 graphics cards, thereby making graphics
  230. 26:42cards look like a spoon when mining bitcoin next  to an excavator that is an asic mining computer.
  231. 26:50Let’s next discuss the design of the tensor cores.  It’ll take multiple full-length videos to cover
  232. 26:56generative AI, and neural networks, so we’ll focus  on the exact matrix math that tensor cores solve.
  233. 27:04Essentially, tensor cores take three matrices and  multiply the first two, add in the third and then
  234. 27:11output the result. Let’s look at one value of the  output. This value is equal to the sum of values
  235. 27:18of the first row of the first matrix multiplied  by the values from the first column of the
  236. 27:23second matrix, and then the corresponding  value of the third matrix is added in.
  237. 27:28Because all the values of the 3 input  matrices are ready at the same time,
  238. 27:33the tensor cores complete all of the matrix  multiplication and addition calculations
  239. 27:38concurrently. Neural Networks and generative  AI require trillions to quadrillions of matrix
  240. 27:45multiplication and addition operations and  typically uses much larger matrices. Finally,
  241. 27:52there are Ray Tracing Cores which we explored in  a separate video that’s already been released.
  242. 27:58That’s pretty much it for graphics cards.  We’re thankful to all our Patreon and YouTube
  243. 28:03Membership Sponsors for supporting our videos.  If you want to financially support our work,
  244. 28:09you can find the links in the description below. This is Branch Education, and we create 3D
  245. 28:15animations that dive deeply into the technology  that drives our modern world. Watch another Branch
  246. 28:22video by clicking one of these cards or click  here to subscribe. Thanks for watching to the end!

About this transcript

This page contains the full transcript of How do Graphics Cards Work? Exploring GPU Architecture by Branch Education, generated from the public captions YouTube serves with the video. The transcript has 3,903 words across 246 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.