[CS61C FA20] Lecture 08.2 - RISC-V lw, sw, Decisions I: Data Transfer Instructions — Transcript
Full transcript
- 0:00[Music]
- 0:08hi
- 0:09welcome back we are continuing with race
- 0:125 assembly
- 0:13so let's see those instructions that do
- 0:16data transfers to
- 0:17and from the memory remember the picture
- 0:20from early in the course when we
- 0:22introduced the principal
- 0:23memory hierarchy the processor core was
- 0:27at the very top with its registers
- 0:31registers are extremely fast they share
- 0:34that precious real estate with a
- 0:35processor core and therefore
- 0:37are extremely expensive so
- 0:40we have a small number of them to be to
- 0:43be precise
- 0:4432 total now
- 0:47on a separate chip typically there is
- 0:50the main memory it is implementing a
- 0:51different technology it is called
- 0:53dram which stands for dynamic random
- 0:56access memory
- 0:58and it comes in different flavors you
- 1:01might have heard of double data rate
- 1:04which is ddr and comes in different
- 1:07generations
- 1:08three four five or maybe high bandwidth
- 1:11memory
- 1:12which is hbm and that also has different
- 1:15generations
- 1:16too there is hbm hbm2 and then
- 1:20now we are getting to hbm3 generation
- 1:24dram is also fast but not nearly as fast
- 1:27as the registers
- 1:29at this price reasonably you get a lot
- 1:31of gigabytes for a few tens of dollars
- 1:34and has medium capacity compared to
- 1:39manage storage that we have in disks and
- 1:41solid-state drives
- 1:45but the question here is you know what
- 1:46how big is this gap in the in the
- 1:48pyramid how much a register is really
- 1:50faster than the memory
- 1:51so given that there are 32 registers um
- 1:5432 bits each total of 1024 bits
- 1:58meaning 128 bytes total there
- 2:02on the other hand in dram we may have in
- 2:06a low end
- 2:06laptop about two gigabytes in a high-end
- 2:08laptop up to 64 gigabytes of memory
- 2:11and the server may have a terabyte of
- 2:13memory and the physics dictates that
- 2:15smaller is faster
- 2:17main claim to fame or bruce lee for
- 2:20example or
- 2:21uh typically point guards are a lot
- 2:24faster than the centers in basketball
- 2:26how much faster are the registers the ra
- 2:30and then dram really so when we put that
- 2:32on a piece of paper
- 2:34we will find out that the numbers work
- 2:36out that they're about
- 2:38you know 50 to 500 500 times faster
- 2:42so in terms that's in terms of latency
- 2:46of one axis or the first access to zero
- 2:49for to get one individual data item
- 2:53you know that taxes takes you know a few
- 2:55tens of nanoseconds and you compare that
- 2:57with the fraction of a nanosecond that
- 2:59it takes us to
- 3:00access registers keep in mind
- 3:04that this is for one isolated access you
- 3:06know many succeed
- 3:08subsequent accesses to nearby memory
- 3:10locations
- 3:11will go faster
- 3:15to put this in perspective a great
- 3:17picture is
- 3:19to keep in mind is this picture of
- 3:22jim gray's storage latency analogy
- 3:26what we have on the left hand side
- 3:30are our registers and they take about a
- 3:34nanosecond to access
- 3:35memory takes about two or so magnitude
- 3:38more
- 3:38so 100 nanoseconds
- 3:43let's try to make an equivalent to tasks
- 3:46that we do as humans so for example if
- 3:48we would like
- 3:49if we would like to retrieve data from
- 3:51our heads um
- 3:53let's say that the retrieving data from
- 3:55a register that is in our head would
- 3:57take about a minute
- 3:58so what is that takes 100 times longer
- 4:03where would that data be well when we
- 4:05are all in berkeley and live and
- 4:07all that um you know it would be
- 4:11normal to make an analogy to you know
- 4:13driving to a town
- 4:14for example we we forgot a piece of
- 4:17paper somewhere and we are going to go
- 4:18and retrieve it
- 4:19let's say that town is sacramento which
- 4:22is
- 4:23you know roughly 100 minutes away around
- 4:26half away that's
- 4:29big penalty on time we can do a lot of
- 4:32stuff in an hour and a half
- 4:34or if that gap is 500 x that it may be
- 4:37the case in some cases in 500 minutes we
- 4:40can make it to los angeles and back
- 4:43it would not be convenient to leave that
- 4:46piece of paper with our data in l.a and
- 4:49having to retreat they have to go to
- 4:50retrieve it
- 4:51or sacramento so this
- 4:55is there to try to illustrate how
- 4:57expensive is it
- 4:58to have to go get the data
- 5:01from the memory let's
- 5:04try to see how do these instructions
- 5:07look like
- 5:08on a simple example so the first one
- 5:10that we're going to look at
- 5:12involves loading from memory to register
- 5:16so this is a simple code that we is
- 5:19that's in c
- 5:19and we'll try to translate it into risk
- 5:22five assembly
- 5:24so let's go back line by line so what we
- 5:26have
- 5:27here
- 5:32we have a declaration of an integer
- 5:35array of 100 elements
- 5:36and then the second line very simply
- 5:39takes element a3 of that array
- 5:42and adds it to the variable h and stores
- 5:45the result
- 5:46in a variable g i'm using kind of
- 5:49assembly techno terminology again but
- 5:52you guys know what what i mean here
- 5:56so we know how to add things in
- 5:58registers so we need to figure out a way
- 6:00how to get
- 6:01the value of this third element in the
- 6:04array that
- 6:04resides in memory to our register
- 6:09we need a new instruction there so the
- 6:11new instruction
- 6:12is showing up right here it is load word
- 6:16so load word and risk 5 goes to the
- 6:20memory
- 6:21and loads from the memory to a register
- 6:24so the syntax is fairly straightforward
- 6:26to understand here take a look
- 6:28this is my destination register so
- 6:31mnemonic is
- 6:33lw destination register and then
- 6:36we have this what does that mean we have
- 6:39to specify the base register
- 6:42which is the pointer to the element a0
- 6:44of the array
- 6:45and then specify the offset to the
- 6:49the element of the array that we would
- 6:51like to retrieve
- 6:52so the base pointer points with an
- 6:55offset
- 6:56of 0 to the element a0 would like to get
- 6:59the third
- 7:00element in the array but remember we are
- 7:03storing
- 7:04integers in this array they are 32 bits
- 7:06wide
- 7:0832 bits is 4 bytes and
- 7:12risk 5 addresses each byte in memory
- 7:15so we need to specify this offset in
- 7:17bytes
- 7:19three words is equal
- 7:22to 12 bytes
- 7:26so that's why this offset is specified
- 7:29by as 12
- 7:30that's why we have 12 here so the actual
- 7:34address of the datum in the memory is
- 7:36calculated
- 7:37as the address of the base pointer which
- 7:41is
- 7:42the the content of a register x15
- 7:46plus 12 very well
- 7:51so keep in mind that this offset must be
- 7:52a constant
- 7:54known at the assembly time so if you
- 7:57picture it another thing to to remember
- 8:00this syntax wall
- 8:01is the direction of the data flow
- 8:05within the instruction that kind of
- 8:07mimics what we see
- 8:09in in reality so data flow
- 8:12goes from right to left
- 8:15like what we have seen in other
- 8:18instructions so we add instruction also
- 8:20moves the data
- 8:21from right to left we add contents of
- 8:25registers x12
- 8:26and x10 and store the result
- 8:30in x11
- 8:34so that kind of should resemble a
- 8:37picture that we have had in the
- 8:38previous segment we take data
- 8:42from the right in the memory and move it
- 8:45to the left to the registers so we load
- 8:48from memory to registers
- 8:51we are going from right to left
- 8:54so to recap what we have seen here we
- 8:57have
- 8:58essentially translated this c code one
- 9:01note to
- 9:02to make here there is no really assembly
- 9:04instruction
- 9:05for now for
- 9:09the translates the declaration of um
- 9:12allocating this integer array all what
- 9:14we are translating is this
- 9:16one c instruction that as
- 9:20h plus a3 and saves it as g
- 9:24so what we have there we first load the
- 9:26contents
- 9:27of the of this memory location
- 9:31specified that by the base pointer plus
- 9:33the offset
- 9:34to temporary register x10 and then we
- 9:37add the contents of x10
- 9:39with x12 and s4 result in x11 notice
- 9:42that we could have been a little bit
- 9:44more um
- 9:46optimized here and could have saved
- 9:50this could have reused register x10
- 9:54for the variable g okay
- 9:58hopefully this made sense let's see the
- 10:00other instruction
- 10:01the other instruction that we need is to
- 10:03store from the register
- 10:05to the memory so it goes the other way
- 10:09we have something in the register and we
- 10:11want to put it in a memory and here is
- 10:12the example code
- 10:13exact same code except that we have
- 10:15changed a bit in it
- 10:17instead of saving it in a variable g we
- 10:20are we will need to save it in another
- 10:24element in the array which in this case
- 10:26is a10 it is
- 10:28offset by 10 words from
- 10:31the base pointer so the first two lines
- 10:34are identical
- 10:36as what we have had before we did get a
- 10:39little bit better here we are
- 10:40we have learned we are reusing x10
- 10:44as a temporary variable and now here
- 10:46comes a new instruction
- 10:48this new instruction is store oops
- 10:52store word and risk 5.
- 10:56that store worked in risk 5
- 10:59does the following it takes the contents
- 11:02of a register
- 11:03x10 where we have saved our temporary
- 11:06result
- 11:08and moves it to a memory location
- 11:11that is addressed by the base pointer
- 11:14and the offset
- 11:15and in this case since we want to write
- 11:17in the 10th element of the array
- 11:20we have to specify the offset as
- 11:2310 times 4 bytes away
- 11:26so that is 40
- 11:29is the opposite here so that's why we
- 11:32have 40 there
- 11:38the data movement or data flow here is
- 11:40the other way round
- 11:42we are storing from the register in our
- 11:45picture
- 11:46to the memory
- 11:50and the data flow is going from
- 11:53left to right
- 11:56okay very well so that should hopefully
- 12:00keep confusion out of the way
- 12:03we are always in this syntax um
- 12:07the load operations are moving data from
- 12:10right to left in
- 12:14store operations are storing from
- 12:15register to memory and it is moving
- 12:18from left to right
- 12:22one thing to keep in mind
- 12:26these offsets
- 12:29for loading and storing should be
- 12:33multiples of four it's written here as
- 12:37must because that's a really good
- 12:38practice
- 12:39one side slight disclaimer risk file sa
- 12:43allows for misaligned
- 12:47memory uh accesses meaning that it
- 12:50allows you to scribble to start writing
- 12:53in the middle of one word and then write
- 12:55a half a word in one memory location and
- 12:57another half in another memory location
- 13:00but that's not advisable it is just
- 13:02there to support some
- 13:03legacy code we should not do that
- 13:06because that would be something that is
- 13:08very very slow and it's really messy we
- 13:10don't want to be just scribbling all
- 13:11over the memory so
- 13:12you can you should take this should
- 13:15really is a must
- 13:21now keep in mind that these load board
- 13:23and store board
- 13:25instructions support essentially taking
- 13:28the entire 32 bits from a memory
- 13:30location
- 13:31a word from memory and putting it into a
- 13:33register so they they fit really nicely
- 13:35it's 32 bits to 32 bits
- 13:37or when we are storing to the memory we
- 13:40take 32 bits from a register
- 13:42and place it into a memory but often
- 13:45we operate with different kinds of data
- 13:48types that are not necessarily 32 bits
- 13:50wide so it will be wasteful to
- 13:52use 32-bit memory locations to store
- 13:56just 8 bits worth of data
- 13:59like when we are storing characters when
- 14:02we are storing
- 14:04you know color channels
- 14:08that are typically all eight bits wide
- 14:11so
- 14:11risk five is conscious of that and
- 14:14supports
- 14:16shorter operands so it allows for
- 14:20load byte and store byte the format is
- 14:23exactly the same
- 14:24same as what we have seen in load word
- 14:27store word
- 14:30um as we have as we see here
- 14:33so we have load byte
- 14:36it loads a content of a memory location
- 14:38specified by the base pointer
- 14:40x11 with an offset of three bytes
- 14:44to a memory uh to a register x10
- 14:47so um it is the same as what we have
- 14:51seen before
- 14:52with one difference is that this offset
- 14:54now really doesn't have to be
- 14:56and shouldn't be four because we should
- 14:59be able to pluck
- 15:00any byte from any part
- 15:03of a word and write it
- 15:06in our destination register
- 15:10now by convention
- 15:13when we take that byte no matter where
- 15:15it is stored in the memory because
- 15:17all characters can be packed right next
- 15:19to each other as
- 15:20bytes in the world so we will be
- 15:22indexing that
- 15:23by that those memory locations by one
- 15:27it will copy it to the low byte
- 15:30position of the register extent so if
- 15:33this is our register extend
- 15:35here no matter where we picks up fix the
- 15:38python in this case
- 15:39the offset is three so i'll take the the
- 15:42highest byte the most significant byte
- 15:44uh from the that memory location that
- 15:47x11 points to
- 15:49it would write it into a
- 15:52low byte position of register x10
- 15:55there is another important thing to to
- 15:57keep in mind here
- 15:59we don't know what kind of
- 16:02data types we are operating but in many
- 16:05cases
- 16:06we will be operating with signed numbers
- 16:09so remembering signed numbers if it is a
- 16:12two's complement number
- 16:14this first bit of a byte of an
- 16:17of a byte determines what is
- 16:21the sign whether that number is positive
- 16:24or negative by
- 16:25looking at it if it is a zero the number
- 16:27is positive if it is a one
- 16:29the number is negative so what should we
- 16:32do
- 16:34when we copy this byte to our
- 16:37register to preserve its sign so
- 16:41if we just write zeros um in the upper
- 16:44three bytes well
- 16:48then no matter what their number was we
- 16:50are going to make it look like
- 16:52like it's a positive number we do
- 16:55have to preserve the sign so the way how
- 16:58we preserve the sign
- 16:59we take a note of that top
- 17:02bit the most significant bit in the byte
- 17:05and copy it over
- 17:07that copy over is called the sign
- 17:10extension
- 17:12and you know the way how it looks in
- 17:14this picture we actually take that
- 17:17bit and smear it all over
- 17:20the the upper bites in the in the world
- 17:25so we take the most significant bit out
- 17:28of
- 17:29that low bite and smear it over like
- 17:32avocado on a toast
- 17:34make sense
- 17:39so that's sign extension we don't
- 17:43always want to do sign extension so if
- 17:46we don't want if
- 17:47if you're certain this is not a sign um
- 17:51number that we are transferring from
- 17:53memory it is
- 17:54uh a character or or
- 17:57something it is always a positive number
- 18:00like
- 18:01color intensity
- 18:04risk 5 supports load byte
- 18:08unsigned which is another instruction
- 18:12that copies a byte from a memory
- 18:16loads a byte from memory to a register
- 18:19but does not do sign extension so
- 18:23it just fills the upper bytes with zeros
- 18:26a question for you why there is no
- 18:29equivalent of that
- 18:30no unsigned store by it so no sbu
- 18:38well it doesn't make sense we're just
- 18:40taking
- 18:41a byte and plucking it into a memory
- 18:44location there is no
- 18:45sign extension that is happening we
- 18:47don't have to fill out anything there
- 18:49now lbu has exactly the same
- 18:52syntax as lb all right
- 18:57so let's do a little more than exercise
- 18:59here
- 19:00um to see how this all this clicks
- 19:04together
- 19:05so the question here in this example is
- 19:08what
- 19:08is in x12 after these three instructions
- 19:11are executed and
- 19:12feel free to pause here work it out
- 19:14yourself or
- 19:16you can just follow me as i do it
- 19:19so these are the three instructions that
- 19:21are fairly straightforward the first one
- 19:25essentially stores 3f5
- 19:28in register x11 the second one stores a
- 19:31word
- 19:33stores the content of x11 into a memory
- 19:35location that is pointed by the
- 19:38pointed to by the base pointer x5 with
- 19:41no offset so it just
- 19:42takes that address that is stored in x5
- 19:45and writes
- 19:46right there and then finally
- 19:49the the last instruction
- 19:52lb takes bite with opposite of one
- 19:56from that same word and writes it in x12
- 20:00i think you got a picture if you haven't
- 20:02let me do the let me walk you through
- 20:04that
- 20:05so what we have in x11 after the
- 20:08first instruction is this
- 20:12we have f 5
- 20:160 3 0 0 0 0
- 20:19let's say that x5 points to this
- 20:22location
- 20:24so then this content of a register x11
- 20:28will be
- 20:29written exactly there so we are going to
- 20:31get f5
- 20:330 3 0 0
- 20:370 0. now when we say load byte
- 20:42with an offset of 1 which byte are we
- 20:44going to pick
- 20:46so this is a word with an offset of zero
- 20:50the next word in memory has
- 20:53offset of four because we're bite
- 20:55addressing the memory
- 20:57so the byte that has an offset of one
- 21:01is this one so load byte is going to
- 21:04take the contents
- 21:06of this byte of that word
- 21:09second byte in that word or the second
- 21:13to the least significant byte in the
- 21:15word
- 21:16and write it here so it'll be zero three
- 21:19i'll take a look at the most significant
- 21:21bit
- 21:21in that word and sign extend it
- 21:25to fill the rest of the register with
- 21:27zeros hopefully this is a good
- 21:29illustration
- 21:30and you get to practice a few of these
- 21:33afterwards
- 21:36one little final note
- 21:39we have only introduced about eight
- 21:42risk five assembly instructions and one
- 21:45um
- 21:46who could who is paying good enough
- 21:49attention here
- 21:50might notice that at least one of them
- 21:53is
- 21:54redundant it's ad immediate we really
- 21:57don't need
- 21:58an immediate we can replace an immediate
- 21:59with the sequence of two instructions
- 22:01that we have here
- 22:02we can store a constant in the memory
- 22:05load it from the memory into a register
- 22:08and then
- 22:09add that new immediate
- 22:13to a content of another register and
- 22:16voila that is
- 22:17ad immediate
- 22:21so what's up with that why do we have
- 22:23redundancy and respiration
- 22:25reduced instruction set computer
- 22:26supposed to be really lean and mean
- 22:31yeah this does not change the fact that
- 22:33this
- 22:34reduced instruction set computer and
- 22:37immediate
- 22:39is necessary because
- 22:42these two instructions if if
- 22:46we don't have added media those two
- 22:47instructions would be very slow we would
- 22:49have to go for every immediate to
- 22:51sacramento to get it and then
- 22:54do the addition that would be a really
- 22:57really
- 22:58low performance machine because ad
- 23:00immediates
- 23:01are very frequent in assembly code
- 23:06so this is
- 23:09here the instruction there is necessary
- 23:11to support the common case and the
- 23:14common case
- 23:15is adding immediates there is
- 23:18one big difference between these two
- 23:21instructions
- 23:22we'll touch on that in just a little bit
- 23:24in a couple of segments
- 23:26but the point here is immediate
- 23:29in the instruction field are limited
- 23:32there are they have to be less
- 23:33than 32 bits because we have to store
- 23:36other information so immediates
- 23:38are a short part of the instruction
- 23:41field so the range
- 23:42of immediates in addi is
- 23:46relatively short if you need a longer
- 23:50immediate
- 23:50we would have to load it from a memory
- 23:52or there may be other mechanisms
- 23:54to generate it but we'll pause here
- 23:58i'll see you in a bit with more risk 5
- 24:01instructions
About this transcript
This page contains the full transcript of [CS61C FA20] Lecture 08.2 - RISC-V lw, sw, Decisions I: Data Transfer Instructions by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 3,245 words across 569 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.