[CS61C FA20] Lecture 22.1 - Pipelining II: Pipelining RISC-V — Transcript
Full transcript
- 0:01[Music]
- 0:11hello welcome back to our module on
- 0:13pipelining a respire processor
- 0:17we're quite proud when we designed the
- 0:19functional risk 5
- 0:20processor and that was quite an
- 0:23accomplishment
- 0:26but i have to tell you a secret nobody
- 0:29uses in practice nobody uses single
- 0:32cycle cpus
- 0:34there were some early on single cycle
- 0:37cpus
- 0:39but they basically have been abandoned
- 0:43over time
- 0:44why because they're generally
- 0:46inefficient
- 0:48there is not every instruction uses
- 0:50every stage
- 0:52and single cycle cpu really
- 0:55is has a cycle that is set by executing
- 0:59all five stages of processing
- 1:03so what do people do they do
- 1:06pipelining and the reasons are very
- 1:09similar and the mechanism is very
- 1:11similar to what we have
- 1:13seen in laundry processing in laundry
- 1:16processing
- 1:17we saw that there are four stages of
- 1:19processing
- 1:20washing drying folding and stashing
- 1:24and generally when there are multiple
- 1:26people who need to do the laundry there
- 1:28are multiple
- 1:30laundry tasks that need to be performed
- 1:33nobody
- 1:34really has to wait for
- 1:37each task to complete to go through all
- 1:40four stages of laundry processing before
- 1:42we start
- 1:42start the next one we generally will go
- 1:45ahead
- 1:45and load the
- 1:49washer immediately after the first task
- 1:52is done with that
- 1:53and then we'll proceed as we have seen
- 1:55before
- 1:57the same happens in processors
- 2:01like laundry processing data processing
- 2:04in processors has
- 2:06so far five stages instruction fetch
- 2:10instruction
- 2:11decode with register read alu or execute
- 2:15stage
- 2:16memory access and write back
- 2:20and in our example we have put some
- 2:24times um that you know represent
- 2:27the reason of representatives of how
- 2:29long does it take
- 2:30um we said the instruction fetch stays
- 2:32200 picoseconds
- 2:34registered takes a hundred alu
- 2:38takes 200 picoseconds and memory takes
- 2:41200 picoseconds
- 2:43finally register right is the same as
- 2:45register read about
- 2:46100 picoseconds notice
- 2:50that register read and register right
- 2:52are shorter
- 2:53than the others it has significance
- 2:57we're going to come back to that
- 2:59later
- 3:02so the entire executing the entire
- 3:05instruction
- 3:06um is a sum of this all of these stages
- 3:10um and equal
- 3:14the time to execute the instruction
- 3:16equals the time to go through all of
- 3:17these stages
- 3:18and that is 800 picoseconds
- 3:22we're also using uh a pictogram here
- 3:25that you'll find in the textbook
- 3:28in our textbook and in some other
- 3:30textbooks that is
- 3:32trying to illustrate what is active so
- 3:35when
- 3:36this box that represents a memory is
- 3:39shaded
- 3:39only half and when it is shaded on the
- 3:43right hand side it says that we are
- 3:45reading from that
- 3:47block so when we are reading from the
- 3:50registers we are also shading the right
- 3:52hand side on the other hand when we
- 3:54would like to write to the memory
- 3:56we will shade the left
- 3:59hand side of the block and so we would
- 4:02do during the register right
- 4:04right back cycle with the registers on
- 4:07the other hand
- 4:08when the alu is active we say the whole
- 4:10thing
- 4:12okay so let's take a look at our
- 4:16silly way or how does the
- 4:19single cycle cpu process data
- 4:22so here's a stream of instructions
- 4:24instruction sequence that we would like
- 4:26to process
- 4:27and there are a few of them
- 4:30so let's say we first run into an ad we
- 4:33are going to run
- 4:34through an ad as a in in single cycle
- 4:37will process
- 4:38all five stages well fetch
- 4:41the instruction decode it access the
- 4:44registers
- 4:45perform the addition operation that we
- 4:48figured out during the decode process
- 4:50um we don't need to do anything with the
- 4:52memory but we are still going to
- 4:54wait for that time to pass then finally
- 4:57we'll write back
- 4:58and then when we're done with all of
- 5:00that we'll do the next instruction which
- 5:01is the or
- 5:03and when the or is done with writing
- 5:05back we'll start the third instruction
- 5:08which is shift left logical this is
- 5:11wasteful if you as we have seen in the
- 5:14laundry example
- 5:15so we can do better how do we do better
- 5:18well we are going to pipeline we don't
- 5:21have to wait for
- 5:21the entire instruction to complete to
- 5:23start the next instruction
- 5:25we just have to free up the to wait for
- 5:28the first resource to free up
- 5:30which is an instruction memory as soon
- 5:32as
- 5:33we are done with reading the ad out of
- 5:36the instruction memory
- 5:37even though we don't know it is add yet
- 5:40it's just you know we have read 32 bits
- 5:42we can go ahead and start
- 5:43reading the next 32 bits we'll read r
- 5:46and we'll we'll start
- 5:47that or immediately after
- 5:52we are done uh reading the ad
- 5:55so two things are going to happen
- 5:58concurrently
- 5:59we'll decode the
- 6:02add instruction and fetch your
- 6:06instruction
- 6:08and as soon as we are done with the or
- 6:10instruction
- 6:12reading of the or instruction we can go
- 6:14on to the third one
- 6:15shift like left logical so while we are
- 6:19reading shift left logical from the
- 6:21memory and we have no idea that it's
- 6:23shift left logical we will be accessing
- 6:26the registers for the
- 6:27or instructions and performing the
- 6:30addition of the
- 6:30for the first add instruction
- 6:34notice that one thing had to happen here
- 6:37we had to separate these stages with
- 6:40registers
- 6:41so called pipeline registers otherwise
- 6:44all the data
- 6:45would get mixed up
- 6:49so what we are seeing here is that these
- 6:53instructions
- 6:54are being processed concurrently and
- 6:56each instruction is in a different stage
- 6:59of execution
- 7:02so our clock cycle is there to transfer
- 7:07us between these stages of execution or
- 7:10move the data between the pipeline
- 7:11stages
- 7:12it is not associated with processing of
- 7:15the entire instruction
- 7:17so t cycle is a lot shorter
- 7:21it's one fifth of the time that it proc
- 7:23takes us to process
- 7:24the entire instruction we'll see more of
- 7:28that
- 7:28in a second so
- 7:31there is another really important thing
- 7:35compared to a single cycle cpu
- 7:38the time to process the entire
- 7:39instruction is now longer
- 7:42it is longer for the reason that we have
- 7:45to
- 7:45set the clock cycle to match the slowest
- 7:49stage
- 7:49in the pipeline what does that mean
- 7:53well that means that we have had these
- 7:55imbalanced stages in the pipeline some
- 7:57of them were taking
- 7:58200 picoseconds some of them were taking
- 8:00100 picoseconds
- 8:01now since we are using one clock to
- 8:03clock them all
- 8:05we have to set that clock period to be
- 8:07200 picoseconds otherwise if it is
- 8:08shorter than that
- 8:10some of these stages would not be
- 8:12completing their operation
- 8:14so t cycle for a pipeline processor
- 8:18is 200 picoseconds that's a lot shorter
- 8:21that is one quarter of what we are
- 8:23what we have had um for a single cycle
- 8:26cpu
- 8:27which was 800 picoseconds but the time
- 8:30to execute the entire instruction is
- 8:331000 picoseconds
- 8:36so it is longer so we lose on the
- 8:39latency of the single instruction
- 8:41but we gain dramatically on the
- 8:44throughput
- 8:47okay let's take a look at the
- 8:51bit of a comparison between the single
- 8:54cycle
- 8:55and pipeline instructions most that is
- 8:57in this most of the things that are in
- 8:59this table are fairly logical and we
- 9:01should
- 9:02um we should have seen them and noticed
- 9:04them already
- 9:05the timing per stage or per step
- 9:08varies in single cycle it is fixed for
- 9:12all pipelined ones
- 9:16because register access is only 100
- 9:19picoseconds and in pipeline they all
- 9:21have to be made to be the same length
- 9:24the cpi cycles per instruction
- 9:27ideally set to one in practice if you
- 9:30have any memory misses
- 9:32we'll need more than one cycle per
- 9:34instruction
- 9:35depending how good is our memory system
- 9:37we will deal with that
- 9:39in great detail a bit later
- 9:42in practice pipeline systems allow us to
- 9:45have
- 9:46multiple execution units we'll mention
- 9:49that briefly later
- 9:51but it allows us to actually bring the
- 9:54cpi
- 9:55below below one that is not the subject
- 9:58of this class
- 10:00it is the subject of cs152
- 10:04now what we have seen the clock rate
- 10:06here was one over 800 picoseconds
- 10:08or 1.25 gigahertz pipeline processor can
- 10:11run
- 10:12at five gigahertz and
- 10:15achieve a forex speed up so
- 10:19in summary the instruction time
- 10:22the time to execute an instruction gets
- 10:24a little bit longer but we get the
- 10:26dramatic
- 10:27improvement in throughput because we got
- 10:28dramatically higher
- 10:31clock speed remember our increase
- 10:34in throw put is not equal to the number
- 10:37of
- 10:37stages that we have that theoretically
- 10:39we could have 5x increase in throughput
- 10:42but since the stages had imbalanced
- 10:43delays we lost a little bit
- 10:45we still got a good increase of 4x
- 10:50and that is the reason why people
- 10:54always design pipeline processors as
- 10:56we'll see the design process is not
- 10:58logically not much more complex than the
- 11:00single cycle pipeline
- 11:02so why not do it
- 11:05let's take a look at
- 11:08the moment of what is happening
- 11:10sequentially and what is happening
- 11:12simultaneously when we are executing
- 11:15these instructions so we have a list
- 11:17here of six instructions
- 11:19and notice that the coloring here is
- 11:21showing that the add instruction
- 11:23is not accessing the memory or
- 11:25instruction is not accessing the data
- 11:27memory
- 11:28shift logical neither that one is
- 11:31accessing the memory
- 11:32but store word is writing into memory so
- 11:36the left hand side is is colored load
- 11:38word
- 11:39is reading from the memory so the right
- 11:40hand side is
- 11:43shaded and add immediate is not
- 11:46accessing the memory
- 11:49so um
- 11:53we are you know there is not much
- 11:57to to add to here this is the time to
- 11:59execute the instruction
- 12:01it's a thousand picoseconds or one
- 12:03nanosecond
- 12:04and the cycle time is one-fifth of that
- 12:07or 200 picoseconds
- 12:10when we would like to know what is
- 12:14happening here
- 12:15in in sequentially it is
- 12:19the usage of resources in the pipeline
- 12:22or stages in the pipeline
- 12:24by one instruction resource use of
- 12:27instruction is sequential in time each
- 12:30instruction
- 12:31goes through the fetch
- 12:34decode and register access execution
- 12:37phase
- 12:37memory access and right back on the
- 12:40other hand
- 12:41multiple instructions are
- 12:44using different
- 12:47resources at the
- 12:51the the the same time so there are
- 12:54five instructions here that we will say
- 12:57in flight
- 12:58that are using different resources
- 13:00because they're available
- 13:01so uh these resources here the
- 13:04first instruction is near its completion
- 13:07so add is writing back the result into
- 13:10t0
- 13:12or is in data access
- 13:16stage it's not doing anything it's
- 13:18basically just waiting
- 13:20uh for the ad to be done so it can write
- 13:22back
- 13:23its result in t3 shift left logical
- 13:26is using the alu to perform the shift
- 13:31store word is
- 13:35essentially fetching the t3 and t0
- 13:39values from the t3 and t0 register such
- 13:42that it can use them in the store
- 13:43operation
- 13:44load word is just being fetched
- 13:48from the memory and we don't even know
- 13:50that it's a load
- 13:52and that's it that we have to what we
- 13:55need to know for now
- 13:57conceptually about the pipelining of the
- 13:59smile processor it is very similar
- 14:03to running laundry we're going to see
- 14:06what in the next module what do we need
- 14:08to do
- 14:09with our data path to support that
- 14:12so see you then after a bit of a break
About this transcript
This page contains the full transcript of [CS61C FA20] Lecture 22.1 - Pipelining II: Pipelining RISC-V by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 1,881 words across 355 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.