[CS61C FA20] Lecture 32.5 - Flynn Taxonomy, SIMD Instructions: SIMD Array Processing — Transcript
Full transcript
- 0:00[Music]
- 0:10hello
- 0:11and welcome back to our paralysis module
- 0:14we have seen how cindy processors work
- 0:18so we let's see what we need to know in
- 0:21order to program
- 0:22one of those units and generally not
- 0:26independent processors they're a part of
- 0:28a processor
- 0:28and their share the control unit with
- 0:32the scalar processor
- 0:35so in order to program one of those
- 0:37units we need to know which
- 0:39registers are available what are the
- 0:40names and which kind of instructions
- 0:42vector instructions are available over
- 0:45there
- 0:46we're going to take ssc instructions in
- 0:49the example as an example
- 0:51but this can be generalized to pretty
- 0:52much any architecture that is out there
- 0:54so let's say that we would like to do
- 0:56this we would like to work on an array
- 0:59and for
- 0:59each element in the array we would like
- 1:01to find out what is the
- 1:03square root of that element in the array
- 1:06and write the results back in the memory
- 1:11so um if we are writing a
- 1:14single instruction single data type of a
- 1:17code
- 1:18we will be working on each element of
- 1:21the array separately
- 1:22first for each element in the array we
- 1:24would load
- 1:26f to the floating point register perform
- 1:28the calculation find what is the
- 1:30square root and finally write the result
- 1:34from the register to the memory
- 1:38in simdi type of a code
- 1:41and when we are working with essentially
- 1:43four wide ssc
- 1:45style of instructions that we find in
- 1:47intel processors
- 1:48you take form elements of
- 1:51from the array at a time so
- 1:55we would load four members to the ssc
- 1:58register
- 2:00then perform four square root operations
- 2:03in the same in one
- 2:07vector operation and write the result
- 2:10from the register to memory
- 2:14that's it this is how we got this speed
- 2:17up
- 2:18of essentially 4x
- 2:21let's take a look at a little bit more
- 2:23detailed example that would correspond
- 2:25to the assembly code that
- 2:29works with these these md units so let's
- 2:33say that we would like in this example
- 2:35to add single precision floating point
- 2:38vectors
- 2:39so these are single precision floating
- 2:41point vectors they are 32 bits
- 2:43wide so the computation that is to be
- 2:45performed we are going to
- 2:47add two vectors and each of these
- 2:49vectors has
- 2:50elements x y z and w the results go into
- 2:54the
- 2:54destination vector
- 2:58so here is the code there are
- 3:01three sse instructions that are going to
- 3:04be executed
- 3:06two moves and one add so the first move
- 3:10moves the data from
- 3:13the memory to the xmm register in a way
- 3:16that it is memory aligned packed
- 3:18single precision memory align means that
- 3:23that the data is aligned to 128
- 3:26bit boundaries that correspond to the
- 3:28vectors here
- 3:31so four of them four of these um
- 3:37sub operands x y z and w are going to be
- 3:4032 bits
- 3:41forming a 128 bit vector
- 3:49pack single precision just basically
- 3:51means that they're packed
- 3:53as single precision operands within 128
- 3:57bit vector
- 4:00the second instruction is kind of
- 4:02interesting
- 4:04take a look at this one we are not used
- 4:06to this
- 4:07in risk five this exists in in intel's
- 4:10isa
- 4:11we are taking an operand
- 4:15from an address where the v2 resides
- 4:20and adding it to in a
- 4:24vectorized way to the contents of xmm0
- 4:28we can't do that in risk but but in
- 4:31intel's i say we can do that
- 4:33so risk five remember only has loads and
- 4:36stores for the operations
- 4:38of with the memory but here in
- 4:42intel's
- 4:45assembly language we can add
- 4:49the contents of a register to
- 4:52or the contents of a memory location to
- 4:54a contents of a register
- 4:57and store it in that same register the
- 4:59original register xmm0
- 5:01and finally in the third instruction we
- 5:04move
- 5:05the content of the xmm register
- 5:09to the memory location in the same way
- 5:12aligned pack single precision
- 5:16all right that is it you know the the
- 5:18difference here that we
- 5:20do encounter are some of these a little
- 5:23bit more complicated
- 5:24addressing modes and a little bit more
- 5:27complicated
- 5:29instructions that may work where the
- 5:31operands may
- 5:32reside in the memory but other than that
- 5:34this assembly
- 5:35style is similar to what we are what we
- 5:38are used to at least in terms of
- 5:39granularity
- 5:41now we generally don't want to really
- 5:45write the assembly
- 5:49code if we don't have to so first we'll
- 5:52we
- 5:52will rely on the compiler
- 5:56and compilers like gcc do
- 6:00have access to these ssc instructions
- 6:05but may not always perform in an optimal
- 6:08way
- 6:10a common way how people will do this
- 6:12when they want to really get
- 6:14the ultimate performance is to use
- 6:17so-called intrinsics
- 6:18in c programming language so intrigue in
- 6:21string 6
- 6:22are c functions that essentially
- 6:25instant instead of that one intrinsic
- 6:28command
- 6:29replace it with an assembly instruction
- 6:35so we are essentially inserting assembly
- 6:37instruction into the c
- 6:39code and they're they're basically
- 6:41mapped
- 6:42one to one to their corresponding ssc
- 6:45instruction instructions
- 6:47so here you know we can pick a vector
- 6:50data type
- 6:51which is m128
- 6:54and then we can load and store
- 6:56operations
- 6:58so this is a
- 7:01memory load and memory store operations
- 7:04they always start with this underscore m
- 7:06that that's what
- 7:07that's how we'll find these uh
- 7:09intrinsics and
- 7:11you know this is b um
- 7:14this will be moving aligned pack double
- 7:17and be represented as a double 64-bit
- 7:21instruction in the register um and then
- 7:23there are
- 7:24sampling arithmetic instructions which
- 7:26are underscore mm underscore
- 7:29add pd which adds pack double
- 7:32and then mall underscore mm
- 7:36underscore mall underscore pd multiplies
- 7:39pack double
- 7:43contents of those registers
- 7:46that's it all what you need to do now is
- 7:49to take a look at the the manual
- 7:53for this extension of x86 and start
- 7:56writing
- 7:57c code that utilizes through intrinsics
- 8:00such that you can get
- 8:01really good performance on vectors and
- 8:04matrices and
- 8:05stuff like that we'll take a quick break
- 8:09we'll take a look at an example after a
- 8:12break
About this transcript
This page contains the full transcript of [CS61C FA20] Lecture 32.5 - Flynn Taxonomy, SIMD Instructions: SIMD Array Processing by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 1,011 words across 189 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.