YouTube2Text

[CS61C FA20] Lecture 32.5 - Flynn Taxonomy, SIMD Instructions: SIMD Array Processing — Transcript

by CS 61C Departmental · 1,011 words · 189 segments · language en · Watch on YouTube

Full transcript

  1. 0:00[Music]
  2. 0:10hello
  3. 0:11and welcome back to our paralysis module
  4. 0:14we have seen how cindy processors work
  5. 0:18so we let's see what we need to know in
  6. 0:21order to program
  7. 0:22one of those units and generally not
  8. 0:26independent processors they're a part of
  9. 0:28a processor
  10. 0:28and their share the control unit with
  11. 0:32the scalar processor
  12. 0:35so in order to program one of those
  13. 0:37units we need to know which
  14. 0:39registers are available what are the
  15. 0:40names and which kind of instructions
  16. 0:42vector instructions are available over
  17. 0:45there
  18. 0:46we're going to take ssc instructions in
  19. 0:49the example as an example
  20. 0:51but this can be generalized to pretty
  21. 0:52much any architecture that is out there
  22. 0:54so let's say that we would like to do
  23. 0:56this we would like to work on an array
  24. 0:59and for
  25. 0:59each element in the array we would like
  26. 1:01to find out what is the
  27. 1:03square root of that element in the array
  28. 1:06and write the results back in the memory
  29. 1:11so um if we are writing a
  30. 1:14single instruction single data type of a
  31. 1:17code
  32. 1:18we will be working on each element of
  33. 1:21the array separately
  34. 1:22first for each element in the array we
  35. 1:24would load
  36. 1:26f to the floating point register perform
  37. 1:28the calculation find what is the
  38. 1:30square root and finally write the result
  39. 1:34from the register to the memory
  40. 1:38in simdi type of a code
  41. 1:41and when we are working with essentially
  42. 1:43four wide ssc
  43. 1:45style of instructions that we find in
  44. 1:47intel processors
  45. 1:48you take form elements of
  46. 1:51from the array at a time so
  47. 1:55we would load four members to the ssc
  48. 1:58register
  49. 2:00then perform four square root operations
  50. 2:03in the same in one
  51. 2:07vector operation and write the result
  52. 2:10from the register to memory
  53. 2:14that's it this is how we got this speed
  54. 2:17up
  55. 2:18of essentially 4x
  56. 2:21let's take a look at a little bit more
  57. 2:23detailed example that would correspond
  58. 2:25to the assembly code that
  59. 2:29works with these these md units so let's
  60. 2:33say that we would like in this example
  61. 2:35to add single precision floating point
  62. 2:38vectors
  63. 2:39so these are single precision floating
  64. 2:41point vectors they are 32 bits
  65. 2:43wide so the computation that is to be
  66. 2:45performed we are going to
  67. 2:47add two vectors and each of these
  68. 2:49vectors has
  69. 2:50elements x y z and w the results go into
  70. 2:54the
  71. 2:54destination vector
  72. 2:58so here is the code there are
  73. 3:01three sse instructions that are going to
  74. 3:04be executed
  75. 3:06two moves and one add so the first move
  76. 3:10moves the data from
  77. 3:13the memory to the xmm register in a way
  78. 3:16that it is memory aligned packed
  79. 3:18single precision memory align means that
  80. 3:23that the data is aligned to 128
  81. 3:26bit boundaries that correspond to the
  82. 3:28vectors here
  83. 3:31so four of them four of these um
  84. 3:37sub operands x y z and w are going to be
  85. 3:4032 bits
  86. 3:41forming a 128 bit vector
  87. 3:49pack single precision just basically
  88. 3:51means that they're packed
  89. 3:53as single precision operands within 128
  90. 3:57bit vector
  91. 4:00the second instruction is kind of
  92. 4:02interesting
  93. 4:04take a look at this one we are not used
  94. 4:06to this
  95. 4:07in risk five this exists in in intel's
  96. 4:10isa
  97. 4:11we are taking an operand
  98. 4:15from an address where the v2 resides
  99. 4:20and adding it to in a
  100. 4:24vectorized way to the contents of xmm0
  101. 4:28we can't do that in risk but but in
  102. 4:31intel's i say we can do that
  103. 4:33so risk five remember only has loads and
  104. 4:36stores for the operations
  105. 4:38of with the memory but here in
  106. 4:42intel's
  107. 4:45assembly language we can add
  108. 4:49the contents of a register to
  109. 4:52or the contents of a memory location to
  110. 4:54a contents of a register
  111. 4:57and store it in that same register the
  112. 4:59original register xmm0
  113. 5:01and finally in the third instruction we
  114. 5:04move
  115. 5:05the content of the xmm register
  116. 5:09to the memory location in the same way
  117. 5:12aligned pack single precision
  118. 5:16all right that is it you know the the
  119. 5:18difference here that we
  120. 5:20do encounter are some of these a little
  121. 5:23bit more complicated
  122. 5:24addressing modes and a little bit more
  123. 5:27complicated
  124. 5:29instructions that may work where the
  125. 5:31operands may
  126. 5:32reside in the memory but other than that
  127. 5:34this assembly
  128. 5:35style is similar to what we are what we
  129. 5:38are used to at least in terms of
  130. 5:39granularity
  131. 5:41now we generally don't want to really
  132. 5:45write the assembly
  133. 5:49code if we don't have to so first we'll
  134. 5:52we
  135. 5:52will rely on the compiler
  136. 5:56and compilers like gcc do
  137. 6:00have access to these ssc instructions
  138. 6:05but may not always perform in an optimal
  139. 6:08way
  140. 6:10a common way how people will do this
  141. 6:12when they want to really get
  142. 6:14the ultimate performance is to use
  143. 6:17so-called intrinsics
  144. 6:18in c programming language so intrigue in
  145. 6:21string 6
  146. 6:22are c functions that essentially
  147. 6:25instant instead of that one intrinsic
  148. 6:28command
  149. 6:29replace it with an assembly instruction
  150. 6:35so we are essentially inserting assembly
  151. 6:37instruction into the c
  152. 6:39code and they're they're basically
  153. 6:41mapped
  154. 6:42one to one to their corresponding ssc
  155. 6:45instruction instructions
  156. 6:47so here you know we can pick a vector
  157. 6:50data type
  158. 6:51which is m128
  159. 6:54and then we can load and store
  160. 6:56operations
  161. 6:58so this is a
  162. 7:01memory load and memory store operations
  163. 7:04they always start with this underscore m
  164. 7:06that that's what
  165. 7:07that's how we'll find these uh
  166. 7:09intrinsics and
  167. 7:11you know this is b um
  168. 7:14this will be moving aligned pack double
  169. 7:17and be represented as a double 64-bit
  170. 7:21instruction in the register um and then
  171. 7:23there are
  172. 7:24sampling arithmetic instructions which
  173. 7:26are underscore mm underscore
  174. 7:29add pd which adds pack double
  175. 7:32and then mall underscore mm
  176. 7:36underscore mall underscore pd multiplies
  177. 7:39pack double
  178. 7:43contents of those registers
  179. 7:46that's it all what you need to do now is
  180. 7:49to take a look at the the manual
  181. 7:53for this extension of x86 and start
  182. 7:56writing
  183. 7:57c code that utilizes through intrinsics
  184. 8:00such that you can get
  185. 8:01really good performance on vectors and
  186. 8:04matrices and
  187. 8:05stuff like that we'll take a quick break
  188. 8:09we'll take a look at an example after a
  189. 8:12break

About this transcript

This page contains the full transcript of [CS61C FA20] Lecture 32.5 - Flynn Taxonomy, SIMD Instructions: SIMD Array Processing by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 1,011 words across 189 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.