[CS61C FA20] Lecture 32.4 - Flynn Taxonomy, SIMD Instructions: SIMD Architectures — Transcript
Full transcript
- 0:00[Music]
- 0:08hello
- 0:09and welcome back to our parallelism
- 0:11module
- 0:12we've introduced flint's taxonomy of
- 0:14parallel architectures
- 0:16and we said that we are interested in
- 0:19simdi and memdi
- 0:20architectures cmd stands for single
- 0:23instruction multiple data
- 0:25md is multiple instructions multiple
- 0:28data
- 0:31what we are going to do now we will take
- 0:34a bit of a dive into
- 0:38some of the cmd architectures we are not
- 0:41going to actually
- 0:42build a cmd architecture it will see
- 0:46how does a programmer see it
- 0:49so let's get into that so our symbi
- 0:52architectures
- 0:53are there to exploit data level
- 0:57parallelism
- 0:59so we would like to fetch one
- 1:01instruction
- 1:03and apply it to
- 1:06multiple sets of data typically if it is
- 1:09in
- 1:10addition we would just fetch one
- 1:13add that would be some sort of a vector
- 1:16add
- 1:16and apply it to all of the elements of
- 1:19two vectors so instead of a
- 1:21scalar add where we would be just adding
- 1:26two operands to scalar operands this
- 1:29time we would be
- 1:30doing vector addition of
- 1:33element by element for both of the
- 1:37vectors same thing applies to
- 1:40say multiplication and we can
- 1:43in this example here we can simply
- 1:46multiply a coefficient vector by a data
- 1:49vector
- 1:51such that we can perhap perhaps perform
- 1:53some kind of filtering
- 1:55in the next step we would just add
- 1:58add those results so why is this faster
- 2:02well because we just need to fetch one
- 2:04instruction
- 2:05instead of every time fetching the same
- 2:07instructions so if we are operating on a
- 2:09vector
- 2:10we would fetch one instruction and
- 2:13operate
- 2:14applied to multiple data pairs
- 2:18now you may say well but i still i'm
- 2:22i'm going to be baldnecked by the data
- 2:25now if the bandwidth for
- 2:28getting all that that data is wide
- 2:31enough
- 2:31so that we can get chunks of data say
- 2:34from the memory
- 2:35or more likely from the cache
- 2:40people have noticed this and there has
- 2:42been always
- 2:43or for a long long time there has been
- 2:45interesting interest in vector
- 2:47architectures or these kinds of simdi
- 2:50extensions the first noted
- 2:54uh cmd machine was implemented by mit
- 2:57lincoln labs that was the tx2 in 1957
- 3:01and if you look at it they had
- 3:04they had they had an ability to
- 3:08run their full 36 bit
- 3:10[Music]
- 3:14data or they could split it into two 17
- 3:18bit operands or
- 3:19split it further into nine bit operands
- 3:23they're not quite working with you know
- 3:26standardized
- 3:26bytes and words back then or actually
- 3:28the word here was
- 3:3035 bits
- 3:34these architectures were around for a
- 3:37while
- 3:37but were seen in a wide commercial use
- 3:42when they were introduced by intel uh in
- 3:47the late 90s around 1997.
- 3:51it was noticed that people at that time
- 3:55on pcs were running
- 3:56more and more multimedia applications in
- 3:58multimedia while multimedia is
- 4:01was mostly at that time audio and some
- 4:04video
- 4:04and that often involves some kind of
- 4:07filtering
- 4:08and when we're filtering we in in media
- 4:11applications we
- 4:12are operating typically on vectors
- 4:15vectors
- 4:17in one or two dimensions
- 4:20so the way how this is implemented is
- 4:23essentially
- 4:24there are two source vectors
- 4:28in relatively wide you know placed in
- 4:30relatively wide registers
- 4:32and then on some part of
- 4:36of those very you know relatively wide
- 4:38words
- 4:39we could apply the same operation in our
- 4:42and and get the result in the
- 4:45destination register
- 4:47so we would fetch one instruction
- 4:50and do the work on that corresponds to
- 4:54multiple instructions
- 4:55by applying it to this kind of
- 4:58vectorized data
- 5:00this was called mmx or multimedia
- 5:02extension appeared in process
- 5:04in intel's pentium 2 processor
- 5:07and then later generations were named
- 5:10ssc streaming cmd extension that
- 5:13appeared in pentium 3
- 5:15and pentium 4 and beyond and then later
- 5:18on we got
- 5:19avx advanced vector extensions
- 5:23you know here is a quick evolution of
- 5:24what has happened there so it started
- 5:27with
- 5:27mmx around 1997.
- 5:31and remember all intel processors have
- 5:33to be backwards compatible
- 5:36so mmx is still around with us every
- 5:40single newer processor is still
- 5:42implementing mmx
- 5:44but what has been happening the cmd
- 5:46width
- 5:47has increased
- 5:50so the first we went from a 64-bit cmd
- 5:54to 128-bit cmd by adding different page
- 5:58versions of ss and then
- 6:01um with avx we got to 256
- 6:06bit sim d and finally the most recent
- 6:09processors
- 6:10have 512 bits md there is
- 6:14an expectation they'll be you know
- 6:15relatively soon 1024
- 6:18bit wide smd so in order to support this
- 6:22intel had to add new instructions to
- 6:25their already fairly complex
- 6:28instruction set so all these
- 6:30instructions
- 6:31have been added one after another one in
- 6:34each generation at the same time they
- 6:37also have to add new registers such that
- 6:39they can operate out of these registers
- 6:41and they get more registers
- 6:44as they also get wider now when you run
- 6:48something like this
- 6:49on on avx what do we get you know do we
- 6:52get any speed up
- 6:53or do we get expected speedups compared
- 6:56to
- 6:57our c code and the answer is yes look at
- 7:00this
- 7:01now if you use avx extensions
- 7:05on an intel processor
- 7:08you will get some speed ups compared to
- 7:12a plane c the speed ups are around 4x
- 7:15and
- 7:16no you're not going to get around cache
- 7:20size limitations if even with avx that
- 7:24that doesn't
- 7:25help but the code generally just gets
- 7:29faster
- 7:34is this as fast as you can go
- 7:37well uh you can check on this you know a
- 7:40few years old
- 7:40intel processor i7 from a few years ago
- 7:43peak performance is
- 7:4525 gigaflops per second that's
- 7:48you know processor it is running a 3.1
- 7:503.1 gigahertz
- 7:52with two instructions per cycle and
- 7:54capable of doing four multiplications
- 7:57per instruction
- 7:58so we are still at like a quarter
- 8:01of the maximum uh
- 8:04tropo that we can get through that but
- 8:07we are getting all
- 8:08nearly theoretically
- 8:11speed up of forex over a
- 8:15single scalar operation a few other
- 8:18things about the
- 8:19architecture so when we look at ssc
- 8:22ssc or
- 8:27um sse or
- 8:31mmx added these xmm registers
- 8:34these are the um extended registers in
- 8:38mmx they were 64-bit wides in ssc they
- 8:41were 128 bits wide
- 8:46so they are separate
- 8:49there is a separate set of registers
- 8:50that exist in the processor these are
- 8:52not the general purpose registers
- 8:54and not the floating point registers
- 8:56there are
- 8:57additional vector registers there
- 9:04so they also um can be split
- 9:09so one can work on a whole 128
- 9:13bit wide data or split it into
- 9:16two 64-bit words or four
- 9:1932-bit words or break it down
- 9:24you know even further by the way mmx
- 9:28often you know often usage model wants
- 9:30to take these 64-bit wide
- 9:32words and break it down to 8-bit chunks
- 9:35and
- 9:36work with that generally as integers for
- 9:39audio applications
- 9:43generally you know you'll hear that
- 9:45intel implement something that is called
- 9:47the packed cmd type
- 9:49what does that mean is that in these 128
- 9:53bits in sse and ssc2
- 9:56we can pack 16
- 10:01bytes that are eight bits wide or we can
- 10:04pack
- 10:05intel's words which are 16 but intel
- 10:07calls uh
- 10:0916 bit operand a word so we can pack
- 10:12eight of those or you can pack four
- 10:15double words
- 10:16double words or 32 bits or we can pack
- 10:19two quad words and quad words are 64
- 10:22bits
- 10:24a single precision floating point takes
- 10:2632 bits double precision floating point
- 10:30is 64 bits
- 10:33now it really depends what kind of an
- 10:36intel processor
- 10:37you have is what you get some of these
- 10:40things may appear
- 10:41in laptops some of them may appear in
- 10:44desktop some of them maybe may appear in
- 10:47xeon
- 10:50server class processors this is
- 10:53showing that different generations
- 10:57will support different kind of
- 11:00vector extensions so if you have a you
- 11:03know a bit of an
- 11:05older one you may just have ssc then
- 11:08there was an
- 11:08ssc i84 then that got upgraded with avx
- 11:12and then finally
- 11:14we got avx 512
- 11:17the registers are the number of
- 11:20registers is also growing
- 11:22early on there were eight vector
- 11:25registers then that increased to 16. now
- 11:27there are 32 of them
- 11:30so both the width of the registers is
- 11:33growing
- 11:34and the number of registers is growing
- 11:36so if you have something
- 11:37that is really really highly parallel
- 11:40and a typical
- 11:42um lingo that is used for that
- 11:45even in in scientific communities this
- 11:48is
- 11:49embarrassingly parallel so you have if
- 11:52you're working with a lot of matrices
- 11:55um then this is a great architecture
- 11:58for that it is great for video
- 12:00processing
- 12:02image processing and processing
- 12:05neural nets
- 12:08so what do we have well we do run our
- 12:11our
- 12:12standard um command in linux ls cpu
- 12:16and then we find out what our processor
- 12:19supports
- 12:20i got you know in order to record this
- 12:23class i
- 12:24needed to get a little better processor
- 12:26so i find out that i have
- 12:28a floating point here then i have sse
- 12:33then sse2 as a c3 and
- 12:36various avx extensions
- 12:40this is useful to know when we try to
- 12:44write assembly code for that but also
- 12:46this is something
- 12:47that our gcc compiler
- 12:50generally is able to use we're going to
- 12:53see a bit
- 12:55of examples of how we use them these
- 12:58extensions in the next modules
- 13:01in the next sections actually see you
- 13:04after a quick break
About this transcript
This page contains the full transcript of [CS61C FA20] Lecture 32.4 - Flynn Taxonomy, SIMD Instructions: SIMD Architectures by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 1,575 words across 298 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.