YouTube2Text

[CS61C FA20] Lecture 30.2 - Virtual Memory II: Translation Lookaside Buffers (TLB) — Transcript

by CS 61C Departmental · 1,839 words · 330 segments · language en · Watch on YouTube

Full transcript

  1. 0:00[Music]
  2. 0:10hello
  3. 0:11and welcome back to our module that
  4. 0:12deals with operating systems and virtual
  5. 0:14memory
  6. 0:16so far we understood reasonably well
  7. 0:19what are the basic principles of
  8. 0:21operation of the virtual memory system
  9. 0:24and we have seen that the processor or
  10. 0:26our process
  11. 0:27uses virtual addresses we translate
  12. 0:30those those virtual addresses
  13. 0:33to physical accesses access addresses
  14. 0:36by using page tables
  15. 0:40now these page tables do reside in
  16. 0:43dram and now
  17. 0:46our memory references got
  18. 0:50a lot more complicated and a lot
  19. 0:53slower what do i mean by that
  20. 0:57well instead of just doing a load or a
  21. 1:00store to the memory
  22. 1:02we actually have to make multiple trips
  23. 1:04to the memory
  24. 1:06if we are using a single level page
  25. 1:09table
  26. 1:10we have to make two trips to the memory
  27. 1:12in order to do a load or a store first
  28. 1:15we need to go access the page table and
  29. 1:18then
  30. 1:18the actual memory location that where we
  31. 1:21would perform a load or a store
  32. 1:24if it is a tool level page table
  33. 1:28implemented in this system we would have
  34. 1:30to make three trips to the memory
  35. 1:33um two to access page tables
  36. 1:36and the third one to actually do a load
  37. 1:37or a store and finally
  38. 1:39if it is a four level page table
  39. 1:42um that we would encounter in a modern
  40. 1:4564-bit processor
  41. 1:47that would require us to make five trips
  42. 1:49to the memory that's
  43. 1:50slow remember we said
  44. 1:54you know memory accesses are
  45. 1:57if compared to you know the access to
  46. 2:00the registers that are across the room
  47. 2:02memory accesses are like trips to
  48. 2:04sacramento
  49. 2:06and now um we need to make five trips to
  50. 2:09sacramento
  51. 2:10in order to get our data
  52. 2:14that's expensive we need to speed that
  53. 2:16up and
  54. 2:18that's what we're going to cover in this
  55. 2:20module a method for speeding up
  56. 2:23these memory references first let's
  57. 2:26recap what
  58. 2:27what we would like to do our
  59. 2:31processor our process uses virtual
  60. 2:34addresses
  61. 2:36that consist of virtual page number vpn
  62. 2:38and the offset
  63. 2:40if you're working with 4 kb pages the
  64. 2:42offset will be 12 bits
  65. 2:44the virtual page number is remaining
  66. 2:4720 bits in
  67. 2:5032-bit system um when we would like to
  68. 2:53access a new page
  69. 2:55we need to go through the actual
  70. 2:57translation to obtain a physical page
  71. 2:59number ppn
  72. 3:00and we would append it to our offset to
  73. 3:02get the actual memory location that you
  74. 3:04would like to work with
  75. 3:06but before that or in concurrently with
  76. 3:09that
  77. 3:09we need to do a protection check um you
  78. 3:12know
  79. 3:12whether this reference has been issued
  80. 3:16by a
  81. 3:16user level process or a supervisor and
  82. 3:19do we have access to those memory
  83. 3:21locations
  84. 3:23are we allowed to read or write or
  85. 3:25execute
  86. 3:26and if answer to any of those is no
  87. 3:30then we draw an exception and let the
  88. 3:34operating system handle that
  89. 3:38and so every instruction
  90. 3:41and data access for both instruction and
  91. 3:44data memories
  92. 3:45needs address translation and protection
  93. 3:47checks
  94. 3:49this looks kind of complicated and we
  95. 3:51would like
  96. 3:52all of this to happen actually in one
  97. 3:54cycle
  98. 3:55not in many cycles that
  99. 4:00we refer to previously so let's see how
  100. 4:02do we speed this up
  101. 4:04we generally will speed this up by using
  102. 4:07a hardware structure that lives inside
  103. 4:09the processor it is called
  104. 4:11translation look aside buffer or a tlb
  105. 4:15it's like a little cache that we are
  106. 4:17using for this we are
  107. 4:18caching these page access table
  108. 4:21references
  109. 4:23as we make them why is it called a
  110. 4:25buffer not the cache well
  111. 4:27because it predates cash so people who
  112. 4:30designed
  113. 4:30first the translation look aside buffers
  114. 4:33did not know about caches
  115. 4:34so they just use the word buffer and
  116. 4:36that's stuck
  117. 4:38all right so we are still using tlbs
  118. 4:42so in order to speed up these
  119. 4:45address translations that are very
  120. 4:48expensive
  121. 4:49we use the following caching mechanism
  122. 4:54we will cache some of the memory
  123. 4:57translations
  124. 4:58in the tlb and tlb is this structure
  125. 5:01that has a relatively small number of
  126. 5:03entries
  127. 5:04that have been recently referenced so
  128. 5:06our vpn
  129. 5:08instead of going to the memory we are
  130. 5:10going to take the vpn
  131. 5:12to associatively
  132. 5:15check the locations here
  133. 5:19in the tlb and see if we have a match
  134. 5:23if you have a match the physical page
  135. 5:26number is going to come out
  136. 5:29from the tlb and you're going to use it
  137. 5:32to append it to the
  138. 5:33offset to access the actual memory
  139. 5:36location
  140. 5:37this is what would be called a tlb hit
  141. 5:42if we have a tlb miss
  142. 5:45meaning that we cannot find
  143. 5:48a tag that corresponds to this virtual
  144. 5:52page number in our tlb we need to go and
  145. 5:56walk our
  146. 5:57page tables you know depending how many
  147. 6:00levels of hierarchy are there
  148. 6:02um and update our tlb
  149. 6:06with a new reference and then proceed
  150. 6:08with that
  151. 6:10a few things here um uh the tlb
  152. 6:14has to inherit many of the bits that we
  153. 6:16care about
  154. 6:17uh the status bits that correspond
  155. 6:21to um the actual page table entries
  156. 6:24um you know we need we would like to
  157. 6:26know if it is valid uh which is usually
  158. 6:28a
  159. 6:28a reference to whether it exists in in
  160. 6:31dram
  161. 6:32and whether it is dirty what does the
  162. 6:34remain if it has been
  163. 6:35written to and does it need to be
  164. 6:39written back to the disk in case in case
  165. 6:42we have to swap it
  166. 6:44but other bits are going to be there
  167. 6:46that are going to
  168. 6:48help us with checking the protection
  169. 6:51okay so how does how is this tlb
  170. 6:54actually being designed in processors
  171. 6:56so typical tlbs nowadays um in
  172. 7:00um are can be either
  173. 7:03small and quick uh 32 to 128
  174. 7:07entries and are usually fully
  175. 7:10associative
  176. 7:13with the since each entry
  177. 7:16maps a large page there is a less
  178. 7:20less spatial locality across pages and
  179. 7:23there is more likely the two entries
  180. 7:25will conflict
  181. 7:27um so you know full associativity does
  182. 7:31make sense
  183. 7:32um very large there are some
  184. 7:36large tlb designs that may have 256 to
  185. 7:39512 entries
  186. 7:40but they're made four to eight ways set
  187. 7:44associative
  188. 7:45and that is still done to make sure
  189. 7:48that we can exercise them in
  190. 7:51approximately one clock cycle
  191. 7:53so you know the complexity here is
  192. 7:55balanced between the associativity and
  193. 7:57the size
  194. 7:58of the tlb
  195. 8:02the replacement policy is most commonly
  196. 8:06fifo the first one first entry that gets
  197. 8:08written in is the first one is going to
  198. 8:10come out
  199. 8:13because we would like to keep the
  200. 8:14freshest ones
  201. 8:16in our tlb the the freshest
  202. 8:20memory translations inside the tlp
  203. 8:23or it could be random generally a full
  204. 8:26lru is perhaps a bit
  205. 8:29expensive to do to to fully implement in
  206. 8:33one clock cycle
  207. 8:35so people will implement some kind of
  208. 8:37approximation of that
  209. 8:40another term that we'll typically find
  210. 8:43being used is something is called the
  211. 8:45tlb reach
  212. 8:46and that corresponds to a size of
  213. 8:49largest virtual address space that can
  214. 8:50be simultaneously mapped
  215. 8:52by tlb so if we
  216. 8:55have 64 entries in the tlb and
  217. 8:59it is referring to 4 kb pages
  218. 9:02what is our tlb reach well
  219. 9:06you got it it's 64 times 4
  220. 9:10or a quarter of a megabyte
  221. 9:13all right that's how much of a memory we
  222. 9:15can access
  223. 9:16straight out of a tlb
  224. 9:20the next question is well we know what
  225. 9:22the tlb is supposed to do
  226. 9:23how do we actually implement it the
  227. 9:26first
  228. 9:26answer to that first we should ask
  229. 9:28ourselves where should it be located in
  230. 9:31where in our data path we should have
  231. 9:34the tlbs
  232. 9:36to answer that let's first ask ourselves
  233. 9:40which should we check first for the the
  234. 9:43for our fresh contents um
  235. 9:47caches or tlbs
  236. 9:53can cache hold requested data if
  237. 9:56corresponding page is not in physical
  238. 9:58memory that gives us in the answer
  239. 10:01no so the tlb has to come first
  240. 10:07the other question then is if tlb
  241. 10:10precedes caches does cache receive
  242. 10:13virtual
  243. 10:14or physical addresses and now that's
  244. 10:17very simple
  245. 10:18cache has to contain physical addresses
  246. 10:22okay so here is a straightforward
  247. 10:24structure that we have been talking
  248. 10:25about
  249. 10:26the cpu will
  250. 10:30issue virtual addresses these virtual
  251. 10:33addresses go to the tlb
  252. 10:35tlb in case of a hit produces a physical
  253. 10:38address
  254. 10:40and that goes to the cache if that
  255. 10:42physical address exists in the cache
  256. 10:46if you have a cache hit we send back
  257. 10:50that data to the cpu and we are done
  258. 10:53if you have a cache miss we go get the
  259. 10:56data from the main memory we update the
  260. 10:58cache
  261. 10:59and send it back to the cpu
  262. 11:03now one more thing can happen here we
  263. 11:06can have a tlb miss
  264. 11:08what happens on a tlb miss we generally
  265. 11:11have to walk so called walk our page
  266. 11:14tables
  267. 11:15update the page tables and then
  268. 11:19proceed with the memory access
  269. 11:27the key thing here is that the tlbs
  270. 11:30do the the
  271. 11:34translation of virtual to physical
  272. 11:36addresses
  273. 11:37not the page table these page tables
  274. 11:40do reside in the main memory and only if
  275. 11:43you have a tlb
  276. 11:44miss we go to the main memory now tlbs
  277. 11:47are generally designed such that we have
  278. 11:48a very very high hit rate
  279. 11:50our hit rates typically into
  280. 11:55into tlbs are better than 99
  281. 11:58or much better than 99 uh you know
  282. 12:01typical numbers that you'll hear are
  283. 12:0399.9 or 99.999
  284. 12:08percent
  285. 12:10so let's see actually how the how is
  286. 12:13this address translation
  287. 12:14being implementing did by using
  288. 12:18a tlp so we have our vpn
  289. 12:21and a page offset that
  290. 12:25composes that is our virtual address
  291. 12:27composed of
  292. 12:29now let's take a look at this we are
  293. 12:31going to make a split
  294. 12:34of our virtual page number into
  295. 12:38two parts we are going to separate into
  296. 12:40a tag
  297. 12:41and the index and that is in case we
  298. 12:44may not be using full associativity
  299. 12:49then that is going to point to the tlb
  300. 12:53inside the tlb the tlb index is going to
  301. 12:56point to the tag and it is
  302. 13:00done in the exact same way as we will do
  303. 13:02it in the cache
  304. 13:04if there is a hit
  305. 13:08physical page number is going to exist
  306. 13:10inside our pl
  307. 13:12tlb it will come out of the tlb and
  308. 13:15we'll
  309. 13:15compose our physical address by
  310. 13:18concatenating the ppn from the tlb
  311. 13:21with the page offset from the virtual
  312. 13:23address
  313. 13:27now we are going to
  314. 13:30split it again in a different way we are
  315. 13:32going to break it up into
  316. 13:34our format that is used for accessing
  317. 13:37caches
  318. 13:38tag index offset or teal so we are going
  319. 13:40to use
  320. 13:41rto to reference the data cache
  321. 13:47so if data is in the cache
  322. 13:51it is going to be we are going to have a
  323. 13:53hit
  324. 13:54if not we are going to have to update it
  325. 13:57and that is basically it we can pause
  326. 14:00now
  327. 14:00and after a quick break we are going to
  328. 14:04see how is this actually
  329. 14:05implemented in a microprocessor data
  330. 14:08path see you in a sec

About this transcript

This page contains the full transcript of [CS61C FA20] Lecture 30.2 - Virtual Memory II: Translation Lookaside Buffers (TLB) by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 1,839 words across 330 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.