[CS61C FA20] Lecture 30.2 - Virtual Memory II: Translation Lookaside Buffers (TLB) — Transcript
Full transcript
- 0:00[Music]
- 0:10hello
- 0:11and welcome back to our module that
- 0:12deals with operating systems and virtual
- 0:14memory
- 0:16so far we understood reasonably well
- 0:19what are the basic principles of
- 0:21operation of the virtual memory system
- 0:24and we have seen that the processor or
- 0:26our process
- 0:27uses virtual addresses we translate
- 0:30those those virtual addresses
- 0:33to physical accesses access addresses
- 0:36by using page tables
- 0:40now these page tables do reside in
- 0:43dram and now
- 0:46our memory references got
- 0:50a lot more complicated and a lot
- 0:53slower what do i mean by that
- 0:57well instead of just doing a load or a
- 1:00store to the memory
- 1:02we actually have to make multiple trips
- 1:04to the memory
- 1:06if we are using a single level page
- 1:09table
- 1:10we have to make two trips to the memory
- 1:12in order to do a load or a store first
- 1:15we need to go access the page table and
- 1:18then
- 1:18the actual memory location that where we
- 1:21would perform a load or a store
- 1:24if it is a tool level page table
- 1:28implemented in this system we would have
- 1:30to make three trips to the memory
- 1:33um two to access page tables
- 1:36and the third one to actually do a load
- 1:37or a store and finally
- 1:39if it is a four level page table
- 1:42um that we would encounter in a modern
- 1:4564-bit processor
- 1:47that would require us to make five trips
- 1:49to the memory that's
- 1:50slow remember we said
- 1:54you know memory accesses are
- 1:57if compared to you know the access to
- 2:00the registers that are across the room
- 2:02memory accesses are like trips to
- 2:04sacramento
- 2:06and now um we need to make five trips to
- 2:09sacramento
- 2:10in order to get our data
- 2:14that's expensive we need to speed that
- 2:16up and
- 2:18that's what we're going to cover in this
- 2:20module a method for speeding up
- 2:23these memory references first let's
- 2:26recap what
- 2:27what we would like to do our
- 2:31processor our process uses virtual
- 2:34addresses
- 2:36that consist of virtual page number vpn
- 2:38and the offset
- 2:40if you're working with 4 kb pages the
- 2:42offset will be 12 bits
- 2:44the virtual page number is remaining
- 2:4720 bits in
- 2:5032-bit system um when we would like to
- 2:53access a new page
- 2:55we need to go through the actual
- 2:57translation to obtain a physical page
- 2:59number ppn
- 3:00and we would append it to our offset to
- 3:02get the actual memory location that you
- 3:04would like to work with
- 3:06but before that or in concurrently with
- 3:09that
- 3:09we need to do a protection check um you
- 3:12know
- 3:12whether this reference has been issued
- 3:16by a
- 3:16user level process or a supervisor and
- 3:19do we have access to those memory
- 3:21locations
- 3:23are we allowed to read or write or
- 3:25execute
- 3:26and if answer to any of those is no
- 3:30then we draw an exception and let the
- 3:34operating system handle that
- 3:38and so every instruction
- 3:41and data access for both instruction and
- 3:44data memories
- 3:45needs address translation and protection
- 3:47checks
- 3:49this looks kind of complicated and we
- 3:51would like
- 3:52all of this to happen actually in one
- 3:54cycle
- 3:55not in many cycles that
- 4:00we refer to previously so let's see how
- 4:02do we speed this up
- 4:04we generally will speed this up by using
- 4:07a hardware structure that lives inside
- 4:09the processor it is called
- 4:11translation look aside buffer or a tlb
- 4:15it's like a little cache that we are
- 4:17using for this we are
- 4:18caching these page access table
- 4:21references
- 4:23as we make them why is it called a
- 4:25buffer not the cache well
- 4:27because it predates cash so people who
- 4:30designed
- 4:30first the translation look aside buffers
- 4:33did not know about caches
- 4:34so they just use the word buffer and
- 4:36that's stuck
- 4:38all right so we are still using tlbs
- 4:42so in order to speed up these
- 4:45address translations that are very
- 4:48expensive
- 4:49we use the following caching mechanism
- 4:54we will cache some of the memory
- 4:57translations
- 4:58in the tlb and tlb is this structure
- 5:01that has a relatively small number of
- 5:03entries
- 5:04that have been recently referenced so
- 5:06our vpn
- 5:08instead of going to the memory we are
- 5:10going to take the vpn
- 5:12to associatively
- 5:15check the locations here
- 5:19in the tlb and see if we have a match
- 5:23if you have a match the physical page
- 5:26number is going to come out
- 5:29from the tlb and you're going to use it
- 5:32to append it to the
- 5:33offset to access the actual memory
- 5:36location
- 5:37this is what would be called a tlb hit
- 5:42if we have a tlb miss
- 5:45meaning that we cannot find
- 5:48a tag that corresponds to this virtual
- 5:52page number in our tlb we need to go and
- 5:56walk our
- 5:57page tables you know depending how many
- 6:00levels of hierarchy are there
- 6:02um and update our tlb
- 6:06with a new reference and then proceed
- 6:08with that
- 6:10a few things here um uh the tlb
- 6:14has to inherit many of the bits that we
- 6:16care about
- 6:17uh the status bits that correspond
- 6:21to um the actual page table entries
- 6:24um you know we need we would like to
- 6:26know if it is valid uh which is usually
- 6:28a
- 6:28a reference to whether it exists in in
- 6:31dram
- 6:32and whether it is dirty what does the
- 6:34remain if it has been
- 6:35written to and does it need to be
- 6:39written back to the disk in case in case
- 6:42we have to swap it
- 6:44but other bits are going to be there
- 6:46that are going to
- 6:48help us with checking the protection
- 6:51okay so how does how is this tlb
- 6:54actually being designed in processors
- 6:56so typical tlbs nowadays um in
- 7:00um are can be either
- 7:03small and quick uh 32 to 128
- 7:07entries and are usually fully
- 7:10associative
- 7:13with the since each entry
- 7:16maps a large page there is a less
- 7:20less spatial locality across pages and
- 7:23there is more likely the two entries
- 7:25will conflict
- 7:27um so you know full associativity does
- 7:31make sense
- 7:32um very large there are some
- 7:36large tlb designs that may have 256 to
- 7:39512 entries
- 7:40but they're made four to eight ways set
- 7:44associative
- 7:45and that is still done to make sure
- 7:48that we can exercise them in
- 7:51approximately one clock cycle
- 7:53so you know the complexity here is
- 7:55balanced between the associativity and
- 7:57the size
- 7:58of the tlb
- 8:02the replacement policy is most commonly
- 8:06fifo the first one first entry that gets
- 8:08written in is the first one is going to
- 8:10come out
- 8:13because we would like to keep the
- 8:14freshest ones
- 8:16in our tlb the the freshest
- 8:20memory translations inside the tlp
- 8:23or it could be random generally a full
- 8:26lru is perhaps a bit
- 8:29expensive to do to to fully implement in
- 8:33one clock cycle
- 8:35so people will implement some kind of
- 8:37approximation of that
- 8:40another term that we'll typically find
- 8:43being used is something is called the
- 8:45tlb reach
- 8:46and that corresponds to a size of
- 8:49largest virtual address space that can
- 8:50be simultaneously mapped
- 8:52by tlb so if we
- 8:55have 64 entries in the tlb and
- 8:59it is referring to 4 kb pages
- 9:02what is our tlb reach well
- 9:06you got it it's 64 times 4
- 9:10or a quarter of a megabyte
- 9:13all right that's how much of a memory we
- 9:15can access
- 9:16straight out of a tlb
- 9:20the next question is well we know what
- 9:22the tlb is supposed to do
- 9:23how do we actually implement it the
- 9:26first
- 9:26answer to that first we should ask
- 9:28ourselves where should it be located in
- 9:31where in our data path we should have
- 9:34the tlbs
- 9:36to answer that let's first ask ourselves
- 9:40which should we check first for the the
- 9:43for our fresh contents um
- 9:47caches or tlbs
- 9:53can cache hold requested data if
- 9:56corresponding page is not in physical
- 9:58memory that gives us in the answer
- 10:01no so the tlb has to come first
- 10:07the other question then is if tlb
- 10:10precedes caches does cache receive
- 10:13virtual
- 10:14or physical addresses and now that's
- 10:17very simple
- 10:18cache has to contain physical addresses
- 10:22okay so here is a straightforward
- 10:24structure that we have been talking
- 10:25about
- 10:26the cpu will
- 10:30issue virtual addresses these virtual
- 10:33addresses go to the tlb
- 10:35tlb in case of a hit produces a physical
- 10:38address
- 10:40and that goes to the cache if that
- 10:42physical address exists in the cache
- 10:46if you have a cache hit we send back
- 10:50that data to the cpu and we are done
- 10:53if you have a cache miss we go get the
- 10:56data from the main memory we update the
- 10:58cache
- 10:59and send it back to the cpu
- 11:03now one more thing can happen here we
- 11:06can have a tlb miss
- 11:08what happens on a tlb miss we generally
- 11:11have to walk so called walk our page
- 11:14tables
- 11:15update the page tables and then
- 11:19proceed with the memory access
- 11:27the key thing here is that the tlbs
- 11:30do the the
- 11:34translation of virtual to physical
- 11:36addresses
- 11:37not the page table these page tables
- 11:40do reside in the main memory and only if
- 11:43you have a tlb
- 11:44miss we go to the main memory now tlbs
- 11:47are generally designed such that we have
- 11:48a very very high hit rate
- 11:50our hit rates typically into
- 11:55into tlbs are better than 99
- 11:58or much better than 99 uh you know
- 12:01typical numbers that you'll hear are
- 12:0399.9 or 99.999
- 12:08percent
- 12:10so let's see actually how the how is
- 12:13this address translation
- 12:14being implementing did by using
- 12:18a tlp so we have our vpn
- 12:21and a page offset that
- 12:25composes that is our virtual address
- 12:27composed of
- 12:29now let's take a look at this we are
- 12:31going to make a split
- 12:34of our virtual page number into
- 12:38two parts we are going to separate into
- 12:40a tag
- 12:41and the index and that is in case we
- 12:44may not be using full associativity
- 12:49then that is going to point to the tlb
- 12:53inside the tlb the tlb index is going to
- 12:56point to the tag and it is
- 13:00done in the exact same way as we will do
- 13:02it in the cache
- 13:04if there is a hit
- 13:08physical page number is going to exist
- 13:10inside our pl
- 13:12tlb it will come out of the tlb and
- 13:15we'll
- 13:15compose our physical address by
- 13:18concatenating the ppn from the tlb
- 13:21with the page offset from the virtual
- 13:23address
- 13:27now we are going to
- 13:30split it again in a different way we are
- 13:32going to break it up into
- 13:34our format that is used for accessing
- 13:37caches
- 13:38tag index offset or teal so we are going
- 13:40to use
- 13:41rto to reference the data cache
- 13:47so if data is in the cache
- 13:51it is going to be we are going to have a
- 13:53hit
- 13:54if not we are going to have to update it
- 13:57and that is basically it we can pause
- 14:00now
- 14:00and after a quick break we are going to
- 14:04see how is this actually
- 14:05implemented in a microprocessor data
- 14:08path see you in a sec
About this transcript
This page contains the full transcript of [CS61C FA20] Lecture 30.2 - Virtual Memory II: Translation Lookaside Buffers (TLB) by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 1,839 words across 330 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.