[CS61C FA20] Lecture 23.3 - Pipelining III: Superscalar — Transcript
Full transcript
- 0:01[Music]
- 0:12hi
- 0:12welcome back to our pipelining module
- 0:16we're almost done we have learned how
- 0:20to design a pipeline processor we
- 0:22actually learned the principles of
- 0:23designing a pipeline processor
- 0:26[Music]
- 0:27you will get the true feeling for how to
- 0:30design one
- 0:31when doing a project and actually going
- 0:34through the design of a pipeline
- 0:36and resolution of all the hazards
- 0:39that's really fun actually
- 0:43but for now we understood all the
- 0:46principles that we
- 0:47need to deal with in order to design a
- 0:50good processor
- 0:53there is one more advanced concept that
- 0:55is out there that's a concept of
- 0:56so-called
- 0:57super scalar processors you might have
- 1:00heard about that
- 1:01many of the modern very high performance
- 1:04processors
- 1:04are of super scalar type
- 1:08before we get to we get to those let's
- 1:11see
- 1:12how can we increase processor
- 1:15performance further
- 1:16from what we know so far
- 1:21so we have designed a five-stage
- 1:23pipeline and what is
- 1:24kind of interesting this five-stage
- 1:26pipeline is the industry workhorse
- 1:28nowadays
- 1:29probably 80 or 90 percent of processors
- 1:31that are out there
- 1:32all around us are five-stage pipelines
- 1:36but they don't have the highest possible
- 1:38performance
- 1:40there are many designs
- 1:43that are
- 1:46different their higher performance and
- 1:49when i say majority of processors out
- 1:51there are five stage pipelines
- 1:53you'll find them in all kinds of
- 1:55applications in
- 1:56cars and various appliances
- 2:00and so on they're really common they're
- 2:02like that they're in they're these
- 2:04invisible computers they're all over us
- 2:06all around us but those that will find
- 2:08in laptops or in desktops or even in
- 2:11cell phones
- 2:13there are usually a few
- 2:17higher performance processors and a lot
- 2:19of these
- 2:20five stage pipelines as well
- 2:23so how do we crank up the performance of
- 2:25a processor
- 2:26we have a few options one we can try to
- 2:28crank up the clock frequency
- 2:31but the clock frequency is limited by
- 2:34two things
- 2:35how fast transistors will switch and
- 2:37that is said by the technology
- 2:39and the other one is the the power
- 2:42dissipation we have mentioned that
- 2:44before
- 2:45um you know cranking up the clock
- 2:47frequency was our
- 2:48primary means of adding performance in
- 2:51the 90s
- 2:52up to early 2000s but then we hit the
- 2:55power ceiling
- 2:56and we discovered our designers
- 2:59discovered that they can design
- 3:01processors that will
- 3:02run very fast at very high clock
- 3:05frequencies
- 3:06but there was no practical means of
- 3:09cooling them
- 3:11so the clock frequency had to be held
- 3:14constant
- 3:14to allow for reasonable cooling
- 3:18then the second option that's often
- 3:20works in
- 3:22concert with clock rate increases
- 3:25is the concept of deeper pipelining
- 3:29so we have seen our five stage pipeline
- 3:31but we can break
- 3:33up the execution into smaller and
- 3:35smaller chunks so
- 3:36people have built pipelines that are 10
- 3:39stages deep or 15 stages deep
- 3:42intel processors i think went up to 20
- 3:44stages deep
- 3:46at some point so the idea here is to do
- 3:50let less work in each cycle to have less
- 3:53logic depth to go through less larger
- 3:55gates
- 3:56so we can have a shorter clock cycle and
- 3:59that enables
- 4:00us to either run at a higher frequency
- 4:02or if we are
- 4:03limited with power to drop the supply
- 4:06voltage
- 4:06and run at the same clock frequency but
- 4:09with lower power
- 4:12however the deeper the pipeline
- 4:15the higher the chance is that we are
- 4:17going to run into hazards
- 4:18so our cpi is naturally going to
- 4:22get higher to be higher than one
- 4:25so what came as an alternative to that
- 4:28are so-called multi-issue
- 4:30or superscalar processors those are the
- 4:32ones
- 4:33that theoretically can have cpi that
- 4:37is much less than one so how do those
- 4:40work
- 4:40in a super scalar processor we have
- 4:43multiple issues
- 4:44what does it mean we have multiple
- 4:46execution units
- 4:47and everything else that helps feed
- 4:50those execution units
- 4:52so each execution units is actually a
- 4:54pipeline on its own that is capable of
- 4:57executing either integer instructions or
- 5:00floating point instructions or we may
- 5:02have dedicated
- 5:04load and store pipelines
- 5:08so since we have multiple execution
- 5:10units
- 5:11we can fetch multiple instructions per
- 5:14clock cycle
- 5:15and so-called issue them to different
- 5:18execution
- 5:19pipelines um so this enables us
- 5:23to drop our
- 5:26cpi below one so we are going to be
- 5:30completing
- 5:31multiple instructions per clock cycle
- 5:35often people will use a concept ipc such
- 5:38that
- 5:39it's easier to to understand things to
- 5:42compare things that are
- 5:43greater than one than one for example if
- 5:47you have a
- 5:48four gigahertz four-way multiple issue
- 5:50processor
- 5:51it is capable of executing 16 billions
- 5:54of instructions
- 5:56per second uh with it would have a
- 5:59peak cpi of 0.25
- 6:02and peak ipc would be one over that
- 6:05which
- 6:06is four peak means
- 6:09that if there are no conflicts there are
- 6:12no dependencies
- 6:13no hazards we can get to that peak
- 6:17cpi or ipc but in practice that number
- 6:20is not going to be so great
- 6:26often these processors employ something
- 6:28that is called out of order execution
- 6:31so in order to deal with the
- 6:33dependencies and hazards
- 6:34there'll be a hardware unit they'll be
- 6:36picking and
- 6:38figuring out which instructions have
- 6:39dependencies
- 6:41and trying to execute them in an order
- 6:43where they don't depend on each other
- 6:45and then there'll be a reorder unit
- 6:48at the at the end of the pipeline that
- 6:50will put the results back
- 6:52in order such that whoever run that
- 6:55program
- 6:56gets meaningful results and that is all
- 6:59done perfectly
- 7:02this we definitely are not going to
- 7:04cover in 61c
- 7:06but cs152 goes into a great depth of
- 7:10that
- 7:11what we'll we will do is a quick
- 7:14overview of what you will find in a
- 7:17superscalar processor
- 7:18so generally there will be an
- 7:20instruction fetch and decode that will
- 7:22be
- 7:23fetching and decoding multiple
- 7:25instructions per cycle
- 7:28and then it would go and issue them
- 7:32in different execution pipelines
- 7:36um it would recognize what kind of an
- 7:38instruction it is
- 7:40and it would send it down the integer
- 7:42pipelines or floating point
- 7:44pipelines or load store pipelines
- 7:47it would use some kind of reservation
- 7:49station that would be there to
- 7:52reserve the resources the reserve to
- 7:54reserve
- 7:55this functional unit for future use
- 7:58because
- 7:59they the the processor will try to keep
- 8:01them all filled
- 8:02as much as possible and finally
- 8:06the instructions will be committed or
- 8:08retired
- 8:09in order so they'll be arriving out of
- 8:11order but coming out of this commit unit
- 8:14in order
- 8:17here is an example of an
- 8:20um out of order superscalar processor
- 8:23which is
- 8:24intel i7 a relatively modern processor
- 8:27that
- 8:28you'll find in laptops like the one that
- 8:31i'm using now for
- 8:32this presentation
- 8:35it is what is shown here is the cpi
- 8:40for a number of benchmarks here and
- 8:43these are different kinds of benchmarks
- 8:44you don't really need to know
- 8:46to do need to know which kinds are those
- 8:49but this has something to do with
- 8:51some quantum mechanics simulations h264
- 8:55is video compression
- 8:58bzip is a zip
- 9:01type of a program that is compressing
- 9:04some files
- 9:05you'll find a go game and then gcc that
- 9:09is compiling some kind of a target code
- 9:12and then they measure what is the cpi
- 9:16that is achieved for these benchmarks
- 9:21and what you'll find out here is that
- 9:24these cpi's in most cases will drop
- 9:28below one so there will be less than one
- 9:32and in some cases maybe more than one
- 9:37there is another interesting thing to
- 9:38notice on this graph there are
- 9:41there is ideal cpi that is sitting
- 9:43around
- 9:440.25 down here
- 9:49and then there is realistic cpi that
- 9:52rarely gets below
- 9:530.5 so the the the ideal cpi is 0.25
- 9:57or about a quarter and realistic one is
- 10:000.44 and so on this realistic cpi
- 10:04is due to real cpi is due to stalls
- 10:08misspeculation and handling hazards
- 10:13by the way how is this cpi number
- 10:15actually achieved
- 10:16or measured it's not something magic
- 10:20it's just the reverse of our iron law of
- 10:23processor performance
- 10:24so what we have there in there we had
- 10:27our cpi which stands for cycles per
- 10:29instruction
- 10:30and the cpi
- 10:34can be evaluated by knowing all the
- 10:36other
- 10:37elements of this equation so we can time
- 10:40the program
- 10:41uh how long does it take us to execute
- 10:42the program then we can
- 10:45count the number of instructions that
- 10:46that benchmark has
- 10:48and finally we can look up um
- 10:51how long would the does it take to
- 10:54execute a cycle
- 10:56that's the frequency of a processor that
- 10:58give us gives us the number of cycles
- 11:00per instruction
- 11:01and that is straight time to execute the
- 11:04program
- 11:05divided by instructions for a program
- 11:07times the time
- 11:10per cycle
- 11:14so that's it this is the way how these
- 11:17cpi numbers average cpi numbers have
- 11:20been
- 11:22measured for i7 or many other processors
- 11:30a few notes about the design of isa
- 11:33and which isas are good for for
- 11:35pipelining so risk 5 is a
- 11:37a type of a risk
- 11:40i say that is designed with pipelining
- 11:43in mind
- 11:44and there are several features that you
- 11:46will recognize that are in there
- 11:49that are very helpful for designing a
- 11:51pipeline
- 11:52for example in our version uh
- 11:5532 version all
- 11:59instructions and registers are 32 bits
- 12:02but in every case
- 12:03in every variant of risk 5 all
- 12:07instructions are 32 bits
- 12:10so all these instructions are easy
- 12:14to decode in one cycle
- 12:18compare that to x 86
- 12:21their instructions can be one
- 12:25or anything up to 15 bytes
- 12:28so decoding is really complicated there
- 12:32so
- 12:32you know take a look at one instruction
- 12:34you may be able to
- 12:35decode what is what is in there what are
- 12:37you supposed to do
- 12:38but more likely you have to go and um
- 12:42oh give me a few more bytes oh can i
- 12:44decode not yet
- 12:47a few more bytes and finally after 15 of
- 12:49them the most complex instructions can
- 12:51be decoded
- 12:54risk 5 has a small number six different
- 12:57instruction formats
- 12:58they're very easy to decode and
- 13:02read registers in one step so in one
- 13:05stage we can do both
- 13:07many more complex isas require us to do
- 13:09that in multiple steps
- 13:13load store addressing can be done in
- 13:16third stage
- 13:17by using the alu and access the memory
- 13:20in the fourth stage
- 13:21and then memory operands are all all
- 13:23aligned and take
- 13:25only one cycle that is something that is
- 13:28very convenient
- 13:31and we have taken advantage of all of
- 13:32this when designing our pipeline
- 13:37and that's it so we've
- 13:41we have done it we got our pipeline we
- 13:43got our working processor
- 13:45this is capable of executing all rv32i
- 13:49instructions first in one cycle
- 13:53and then we figured out how to pipeline
- 13:55it
- 13:56we identify these five phases of
- 13:59execution
- 14:00that we associate with five stages of a
- 14:02pipeline
- 14:04and we have designed this controller
- 14:06first for a single stage pipeline but
- 14:08then
- 14:09we have outlined the principles for a
- 14:11single stage execution i'm sorry
- 14:13and uh then we i identify how we need to
- 14:17modify it
- 14:18to support uh resolution of hazards in a
- 14:21pipeline
- 14:24pipelining improves the performance
- 14:27and opens an avenue to building
- 14:30something that is really
- 14:32really powerful we're going
- 14:36to wrap up this module now in the next
- 14:39module
- 14:39we are going to continue building better
- 14:42and better computer
- 14:45see you after a bit of a longer break
About this transcript
This page contains the full transcript of [CS61C FA20] Lecture 23.3 - Pipelining III: Superscalar by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 1,908 words across 365 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.