YouTube2Text

[CS61C FA20] Lecture 32.4 - Flynn Taxonomy, SIMD Instructions: SIMD Architectures — Transcript

by CS 61C Departmental · 1,575 words · 298 segments · language en · Watch on YouTube

Full transcript

  1. 0:00[Music]
  2. 0:08hello
  3. 0:09and welcome back to our parallelism
  4. 0:11module
  5. 0:12we've introduced flint's taxonomy of
  6. 0:14parallel architectures
  7. 0:16and we said that we are interested in
  8. 0:19simdi and memdi
  9. 0:20architectures cmd stands for single
  10. 0:23instruction multiple data
  11. 0:25md is multiple instructions multiple
  12. 0:28data
  13. 0:31what we are going to do now we will take
  14. 0:34a bit of a dive into
  15. 0:38some of the cmd architectures we are not
  16. 0:41going to actually
  17. 0:42build a cmd architecture it will see
  18. 0:46how does a programmer see it
  19. 0:49so let's get into that so our symbi
  20. 0:52architectures
  21. 0:53are there to exploit data level
  22. 0:57parallelism
  23. 0:59so we would like to fetch one
  24. 1:01instruction
  25. 1:03and apply it to
  26. 1:06multiple sets of data typically if it is
  27. 1:09in
  28. 1:10addition we would just fetch one
  29. 1:13add that would be some sort of a vector
  30. 1:16add
  31. 1:16and apply it to all of the elements of
  32. 1:19two vectors so instead of a
  33. 1:21scalar add where we would be just adding
  34. 1:26two operands to scalar operands this
  35. 1:29time we would be
  36. 1:30doing vector addition of
  37. 1:33element by element for both of the
  38. 1:37vectors same thing applies to
  39. 1:40say multiplication and we can
  40. 1:43in this example here we can simply
  41. 1:46multiply a coefficient vector by a data
  42. 1:49vector
  43. 1:51such that we can perhap perhaps perform
  44. 1:53some kind of filtering
  45. 1:55in the next step we would just add
  46. 1:58add those results so why is this faster
  47. 2:02well because we just need to fetch one
  48. 2:04instruction
  49. 2:05instead of every time fetching the same
  50. 2:07instructions so if we are operating on a
  51. 2:09vector
  52. 2:10we would fetch one instruction and
  53. 2:13operate
  54. 2:14applied to multiple data pairs
  55. 2:18now you may say well but i still i'm
  56. 2:22i'm going to be baldnecked by the data
  57. 2:25now if the bandwidth for
  58. 2:28getting all that that data is wide
  59. 2:31enough
  60. 2:31so that we can get chunks of data say
  61. 2:34from the memory
  62. 2:35or more likely from the cache
  63. 2:40people have noticed this and there has
  64. 2:42been always
  65. 2:43or for a long long time there has been
  66. 2:45interesting interest in vector
  67. 2:47architectures or these kinds of simdi
  68. 2:50extensions the first noted
  69. 2:54uh cmd machine was implemented by mit
  70. 2:57lincoln labs that was the tx2 in 1957
  71. 3:01and if you look at it they had
  72. 3:04they had they had an ability to
  73. 3:08run their full 36 bit
  74. 3:10[Music]
  75. 3:14data or they could split it into two 17
  76. 3:18bit operands or
  77. 3:19split it further into nine bit operands
  78. 3:23they're not quite working with you know
  79. 3:26standardized
  80. 3:26bytes and words back then or actually
  81. 3:28the word here was
  82. 3:3035 bits
  83. 3:34these architectures were around for a
  84. 3:37while
  85. 3:37but were seen in a wide commercial use
  86. 3:42when they were introduced by intel uh in
  87. 3:47the late 90s around 1997.
  88. 3:51it was noticed that people at that time
  89. 3:55on pcs were running
  90. 3:56more and more multimedia applications in
  91. 3:58multimedia while multimedia is
  92. 4:01was mostly at that time audio and some
  93. 4:04video
  94. 4:04and that often involves some kind of
  95. 4:07filtering
  96. 4:08and when we're filtering we in in media
  97. 4:11applications we
  98. 4:12are operating typically on vectors
  99. 4:15vectors
  100. 4:17in one or two dimensions
  101. 4:20so the way how this is implemented is
  102. 4:23essentially
  103. 4:24there are two source vectors
  104. 4:28in relatively wide you know placed in
  105. 4:30relatively wide registers
  106. 4:32and then on some part of
  107. 4:36of those very you know relatively wide
  108. 4:38words
  109. 4:39we could apply the same operation in our
  110. 4:42and and get the result in the
  111. 4:45destination register
  112. 4:47so we would fetch one instruction
  113. 4:50and do the work on that corresponds to
  114. 4:54multiple instructions
  115. 4:55by applying it to this kind of
  116. 4:58vectorized data
  117. 5:00this was called mmx or multimedia
  118. 5:02extension appeared in process
  119. 5:04in intel's pentium 2 processor
  120. 5:07and then later generations were named
  121. 5:10ssc streaming cmd extension that
  122. 5:13appeared in pentium 3
  123. 5:15and pentium 4 and beyond and then later
  124. 5:18on we got
  125. 5:19avx advanced vector extensions
  126. 5:23you know here is a quick evolution of
  127. 5:24what has happened there so it started
  128. 5:27with
  129. 5:27mmx around 1997.
  130. 5:31and remember all intel processors have
  131. 5:33to be backwards compatible
  132. 5:36so mmx is still around with us every
  133. 5:40single newer processor is still
  134. 5:42implementing mmx
  135. 5:44but what has been happening the cmd
  136. 5:46width
  137. 5:47has increased
  138. 5:50so the first we went from a 64-bit cmd
  139. 5:54to 128-bit cmd by adding different page
  140. 5:58versions of ss and then
  141. 6:01um with avx we got to 256
  142. 6:06bit sim d and finally the most recent
  143. 6:09processors
  144. 6:10have 512 bits md there is
  145. 6:14an expectation they'll be you know
  146. 6:15relatively soon 1024
  147. 6:18bit wide smd so in order to support this
  148. 6:22intel had to add new instructions to
  149. 6:25their already fairly complex
  150. 6:28instruction set so all these
  151. 6:30instructions
  152. 6:31have been added one after another one in
  153. 6:34each generation at the same time they
  154. 6:37also have to add new registers such that
  155. 6:39they can operate out of these registers
  156. 6:41and they get more registers
  157. 6:44as they also get wider now when you run
  158. 6:48something like this
  159. 6:49on on avx what do we get you know do we
  160. 6:52get any speed up
  161. 6:53or do we get expected speedups compared
  162. 6:56to
  163. 6:57our c code and the answer is yes look at
  164. 7:00this
  165. 7:01now if you use avx extensions
  166. 7:05on an intel processor
  167. 7:08you will get some speed ups compared to
  168. 7:12a plane c the speed ups are around 4x
  169. 7:15and
  170. 7:16no you're not going to get around cache
  171. 7:20size limitations if even with avx that
  172. 7:24that doesn't
  173. 7:25help but the code generally just gets
  174. 7:29faster
  175. 7:34is this as fast as you can go
  176. 7:37well uh you can check on this you know a
  177. 7:40few years old
  178. 7:40intel processor i7 from a few years ago
  179. 7:43peak performance is
  180. 7:4525 gigaflops per second that's
  181. 7:48you know processor it is running a 3.1
  182. 7:503.1 gigahertz
  183. 7:52with two instructions per cycle and
  184. 7:54capable of doing four multiplications
  185. 7:57per instruction
  186. 7:58so we are still at like a quarter
  187. 8:01of the maximum uh
  188. 8:04tropo that we can get through that but
  189. 8:07we are getting all
  190. 8:08nearly theoretically
  191. 8:11speed up of forex over a
  192. 8:15single scalar operation a few other
  193. 8:18things about the
  194. 8:19architecture so when we look at ssc
  195. 8:22ssc or
  196. 8:27um sse or
  197. 8:31mmx added these xmm registers
  198. 8:34these are the um extended registers in
  199. 8:38mmx they were 64-bit wides in ssc they
  200. 8:41were 128 bits wide
  201. 8:46so they are separate
  202. 8:49there is a separate set of registers
  203. 8:50that exist in the processor these are
  204. 8:52not the general purpose registers
  205. 8:54and not the floating point registers
  206. 8:56there are
  207. 8:57additional vector registers there
  208. 9:04so they also um can be split
  209. 9:09so one can work on a whole 128
  210. 9:13bit wide data or split it into
  211. 9:16two 64-bit words or four
  212. 9:1932-bit words or break it down
  213. 9:24you know even further by the way mmx
  214. 9:28often you know often usage model wants
  215. 9:30to take these 64-bit wide
  216. 9:32words and break it down to 8-bit chunks
  217. 9:35and
  218. 9:36work with that generally as integers for
  219. 9:39audio applications
  220. 9:43generally you know you'll hear that
  221. 9:45intel implement something that is called
  222. 9:47the packed cmd type
  223. 9:49what does that mean is that in these 128
  224. 9:53bits in sse and ssc2
  225. 9:56we can pack 16
  226. 10:01bytes that are eight bits wide or we can
  227. 10:04pack
  228. 10:05intel's words which are 16 but intel
  229. 10:07calls uh
  230. 10:0916 bit operand a word so we can pack
  231. 10:12eight of those or you can pack four
  232. 10:15double words
  233. 10:16double words or 32 bits or we can pack
  234. 10:19two quad words and quad words are 64
  235. 10:22bits
  236. 10:24a single precision floating point takes
  237. 10:2632 bits double precision floating point
  238. 10:30is 64 bits
  239. 10:33now it really depends what kind of an
  240. 10:36intel processor
  241. 10:37you have is what you get some of these
  242. 10:40things may appear
  243. 10:41in laptops some of them may appear in
  244. 10:44desktop some of them maybe may appear in
  245. 10:47xeon
  246. 10:50server class processors this is
  247. 10:53showing that different generations
  248. 10:57will support different kind of
  249. 11:00vector extensions so if you have a you
  250. 11:03know a bit of an
  251. 11:05older one you may just have ssc then
  252. 11:08there was an
  253. 11:08ssc i84 then that got upgraded with avx
  254. 11:12and then finally
  255. 11:14we got avx 512
  256. 11:17the registers are the number of
  257. 11:20registers is also growing
  258. 11:22early on there were eight vector
  259. 11:25registers then that increased to 16. now
  260. 11:27there are 32 of them
  261. 11:30so both the width of the registers is
  262. 11:33growing
  263. 11:34and the number of registers is growing
  264. 11:36so if you have something
  265. 11:37that is really really highly parallel
  266. 11:40and a typical
  267. 11:42um lingo that is used for that
  268. 11:45even in in scientific communities this
  269. 11:48is
  270. 11:49embarrassingly parallel so you have if
  271. 11:52you're working with a lot of matrices
  272. 11:55um then this is a great architecture
  273. 11:58for that it is great for video
  274. 12:00processing
  275. 12:02image processing and processing
  276. 12:05neural nets
  277. 12:08so what do we have well we do run our
  278. 12:11our
  279. 12:12standard um command in linux ls cpu
  280. 12:16and then we find out what our processor
  281. 12:19supports
  282. 12:20i got you know in order to record this
  283. 12:23class i
  284. 12:24needed to get a little better processor
  285. 12:26so i find out that i have
  286. 12:28a floating point here then i have sse
  287. 12:33then sse2 as a c3 and
  288. 12:36various avx extensions
  289. 12:40this is useful to know when we try to
  290. 12:44write assembly code for that but also
  291. 12:46this is something
  292. 12:47that our gcc compiler
  293. 12:50generally is able to use we're going to
  294. 12:53see a bit
  295. 12:55of examples of how we use them these
  296. 12:58extensions in the next modules
  297. 13:01in the next sections actually see you
  298. 13:04after a quick break

About this transcript

This page contains the full transcript of [CS61C FA20] Lecture 32.4 - Flynn Taxonomy, SIMD Instructions: SIMD Architectures by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 1,575 words across 298 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.