YouTube2Text

[CS61C FA20] Lecture 22.1 - Pipelining II: Pipelining RISC-V — Transcript

by CS 61C Departmental · 1,881 words · 355 segments · language en · Watch on YouTube

Full transcript

  1. 0:01[Music]
  2. 0:11hello welcome back to our module on
  3. 0:13pipelining a respire processor
  4. 0:17we're quite proud when we designed the
  5. 0:19functional risk 5
  6. 0:20processor and that was quite an
  7. 0:23accomplishment
  8. 0:26but i have to tell you a secret nobody
  9. 0:29uses in practice nobody uses single
  10. 0:32cycle cpus
  11. 0:34there were some early on single cycle
  12. 0:37cpus
  13. 0:39but they basically have been abandoned
  14. 0:43over time
  15. 0:44why because they're generally
  16. 0:46inefficient
  17. 0:48there is not every instruction uses
  18. 0:50every stage
  19. 0:52and single cycle cpu really
  20. 0:55is has a cycle that is set by executing
  21. 0:59all five stages of processing
  22. 1:03so what do people do they do
  23. 1:06pipelining and the reasons are very
  24. 1:09similar and the mechanism is very
  25. 1:11similar to what we have
  26. 1:13seen in laundry processing in laundry
  27. 1:16processing
  28. 1:17we saw that there are four stages of
  29. 1:19processing
  30. 1:20washing drying folding and stashing
  31. 1:24and generally when there are multiple
  32. 1:26people who need to do the laundry there
  33. 1:28are multiple
  34. 1:30laundry tasks that need to be performed
  35. 1:33nobody
  36. 1:34really has to wait for
  37. 1:37each task to complete to go through all
  38. 1:40four stages of laundry processing before
  39. 1:42we start
  40. 1:42start the next one we generally will go
  41. 1:45ahead
  42. 1:45and load the
  43. 1:49washer immediately after the first task
  44. 1:52is done with that
  45. 1:53and then we'll proceed as we have seen
  46. 1:55before
  47. 1:57the same happens in processors
  48. 2:01like laundry processing data processing
  49. 2:04in processors has
  50. 2:06so far five stages instruction fetch
  51. 2:10instruction
  52. 2:11decode with register read alu or execute
  53. 2:15stage
  54. 2:16memory access and write back
  55. 2:20and in our example we have put some
  56. 2:24times um that you know represent
  57. 2:27the reason of representatives of how
  58. 2:29long does it take
  59. 2:30um we said the instruction fetch stays
  60. 2:32200 picoseconds
  61. 2:34registered takes a hundred alu
  62. 2:38takes 200 picoseconds and memory takes
  63. 2:41200 picoseconds
  64. 2:43finally register right is the same as
  65. 2:45register read about
  66. 2:46100 picoseconds notice
  67. 2:50that register read and register right
  68. 2:52are shorter
  69. 2:53than the others it has significance
  70. 2:57we're going to come back to that
  71. 2:59later
  72. 3:02so the entire executing the entire
  73. 3:05instruction
  74. 3:06um is a sum of this all of these stages
  75. 3:10um and equal
  76. 3:14the time to execute the instruction
  77. 3:16equals the time to go through all of
  78. 3:17these stages
  79. 3:18and that is 800 picoseconds
  80. 3:22we're also using uh a pictogram here
  81. 3:25that you'll find in the textbook
  82. 3:28in our textbook and in some other
  83. 3:30textbooks that is
  84. 3:32trying to illustrate what is active so
  85. 3:35when
  86. 3:36this box that represents a memory is
  87. 3:39shaded
  88. 3:39only half and when it is shaded on the
  89. 3:43right hand side it says that we are
  90. 3:45reading from that
  91. 3:47block so when we are reading from the
  92. 3:50registers we are also shading the right
  93. 3:52hand side on the other hand when we
  94. 3:54would like to write to the memory
  95. 3:56we will shade the left
  96. 3:59hand side of the block and so we would
  97. 4:02do during the register right
  98. 4:04right back cycle with the registers on
  99. 4:07the other hand
  100. 4:08when the alu is active we say the whole
  101. 4:10thing
  102. 4:12okay so let's take a look at our
  103. 4:16silly way or how does the
  104. 4:19single cycle cpu process data
  105. 4:22so here's a stream of instructions
  106. 4:24instruction sequence that we would like
  107. 4:26to process
  108. 4:27and there are a few of them
  109. 4:30so let's say we first run into an ad we
  110. 4:33are going to run
  111. 4:34through an ad as a in in single cycle
  112. 4:37will process
  113. 4:38all five stages well fetch
  114. 4:41the instruction decode it access the
  115. 4:44registers
  116. 4:45perform the addition operation that we
  117. 4:48figured out during the decode process
  118. 4:50um we don't need to do anything with the
  119. 4:52memory but we are still going to
  120. 4:54wait for that time to pass then finally
  121. 4:57we'll write back
  122. 4:58and then when we're done with all of
  123. 5:00that we'll do the next instruction which
  124. 5:01is the or
  125. 5:03and when the or is done with writing
  126. 5:05back we'll start the third instruction
  127. 5:08which is shift left logical this is
  128. 5:11wasteful if you as we have seen in the
  129. 5:14laundry example
  130. 5:15so we can do better how do we do better
  131. 5:18well we are going to pipeline we don't
  132. 5:21have to wait for
  133. 5:21the entire instruction to complete to
  134. 5:23start the next instruction
  135. 5:25we just have to free up the to wait for
  136. 5:28the first resource to free up
  137. 5:30which is an instruction memory as soon
  138. 5:32as
  139. 5:33we are done with reading the ad out of
  140. 5:36the instruction memory
  141. 5:37even though we don't know it is add yet
  142. 5:40it's just you know we have read 32 bits
  143. 5:42we can go ahead and start
  144. 5:43reading the next 32 bits we'll read r
  145. 5:46and we'll we'll start
  146. 5:47that or immediately after
  147. 5:52we are done uh reading the ad
  148. 5:55so two things are going to happen
  149. 5:58concurrently
  150. 5:59we'll decode the
  151. 6:02add instruction and fetch your
  152. 6:06instruction
  153. 6:08and as soon as we are done with the or
  154. 6:10instruction
  155. 6:12reading of the or instruction we can go
  156. 6:14on to the third one
  157. 6:15shift like left logical so while we are
  158. 6:19reading shift left logical from the
  159. 6:21memory and we have no idea that it's
  160. 6:23shift left logical we will be accessing
  161. 6:26the registers for the
  162. 6:27or instructions and performing the
  163. 6:30addition of the
  164. 6:30for the first add instruction
  165. 6:34notice that one thing had to happen here
  166. 6:37we had to separate these stages with
  167. 6:40registers
  168. 6:41so called pipeline registers otherwise
  169. 6:44all the data
  170. 6:45would get mixed up
  171. 6:49so what we are seeing here is that these
  172. 6:53instructions
  173. 6:54are being processed concurrently and
  174. 6:56each instruction is in a different stage
  175. 6:59of execution
  176. 7:02so our clock cycle is there to transfer
  177. 7:07us between these stages of execution or
  178. 7:10move the data between the pipeline
  179. 7:11stages
  180. 7:12it is not associated with processing of
  181. 7:15the entire instruction
  182. 7:17so t cycle is a lot shorter
  183. 7:21it's one fifth of the time that it proc
  184. 7:23takes us to process
  185. 7:24the entire instruction we'll see more of
  186. 7:28that
  187. 7:28in a second so
  188. 7:31there is another really important thing
  189. 7:35compared to a single cycle cpu
  190. 7:38the time to process the entire
  191. 7:39instruction is now longer
  192. 7:42it is longer for the reason that we have
  193. 7:45to
  194. 7:45set the clock cycle to match the slowest
  195. 7:49stage
  196. 7:49in the pipeline what does that mean
  197. 7:53well that means that we have had these
  198. 7:55imbalanced stages in the pipeline some
  199. 7:57of them were taking
  200. 7:58200 picoseconds some of them were taking
  201. 8:00100 picoseconds
  202. 8:01now since we are using one clock to
  203. 8:03clock them all
  204. 8:05we have to set that clock period to be
  205. 8:07200 picoseconds otherwise if it is
  206. 8:08shorter than that
  207. 8:10some of these stages would not be
  208. 8:12completing their operation
  209. 8:14so t cycle for a pipeline processor
  210. 8:18is 200 picoseconds that's a lot shorter
  211. 8:21that is one quarter of what we are
  212. 8:23what we have had um for a single cycle
  213. 8:26cpu
  214. 8:27which was 800 picoseconds but the time
  215. 8:30to execute the entire instruction is
  216. 8:331000 picoseconds
  217. 8:36so it is longer so we lose on the
  218. 8:39latency of the single instruction
  219. 8:41but we gain dramatically on the
  220. 8:44throughput
  221. 8:47okay let's take a look at the
  222. 8:51bit of a comparison between the single
  223. 8:54cycle
  224. 8:55and pipeline instructions most that is
  225. 8:57in this most of the things that are in
  226. 8:59this table are fairly logical and we
  227. 9:01should
  228. 9:02um we should have seen them and noticed
  229. 9:04them already
  230. 9:05the timing per stage or per step
  231. 9:08varies in single cycle it is fixed for
  232. 9:12all pipelined ones
  233. 9:16because register access is only 100
  234. 9:19picoseconds and in pipeline they all
  235. 9:21have to be made to be the same length
  236. 9:24the cpi cycles per instruction
  237. 9:27ideally set to one in practice if you
  238. 9:30have any memory misses
  239. 9:32we'll need more than one cycle per
  240. 9:34instruction
  241. 9:35depending how good is our memory system
  242. 9:37we will deal with that
  243. 9:39in great detail a bit later
  244. 9:42in practice pipeline systems allow us to
  245. 9:45have
  246. 9:46multiple execution units we'll mention
  247. 9:49that briefly later
  248. 9:51but it allows us to actually bring the
  249. 9:54cpi
  250. 9:55below below one that is not the subject
  251. 9:58of this class
  252. 10:00it is the subject of cs152
  253. 10:04now what we have seen the clock rate
  254. 10:06here was one over 800 picoseconds
  255. 10:08or 1.25 gigahertz pipeline processor can
  256. 10:11run
  257. 10:12at five gigahertz and
  258. 10:15achieve a forex speed up so
  259. 10:19in summary the instruction time
  260. 10:22the time to execute an instruction gets
  261. 10:24a little bit longer but we get the
  262. 10:26dramatic
  263. 10:27improvement in throughput because we got
  264. 10:28dramatically higher
  265. 10:31clock speed remember our increase
  266. 10:34in throw put is not equal to the number
  267. 10:37of
  268. 10:37stages that we have that theoretically
  269. 10:39we could have 5x increase in throughput
  270. 10:42but since the stages had imbalanced
  271. 10:43delays we lost a little bit
  272. 10:45we still got a good increase of 4x
  273. 10:50and that is the reason why people
  274. 10:54always design pipeline processors as
  275. 10:56we'll see the design process is not
  276. 10:58logically not much more complex than the
  277. 11:00single cycle pipeline
  278. 11:02so why not do it
  279. 11:05let's take a look at
  280. 11:08the moment of what is happening
  281. 11:10sequentially and what is happening
  282. 11:12simultaneously when we are executing
  283. 11:15these instructions so we have a list
  284. 11:17here of six instructions
  285. 11:19and notice that the coloring here is
  286. 11:21showing that the add instruction
  287. 11:23is not accessing the memory or
  288. 11:25instruction is not accessing the data
  289. 11:27memory
  290. 11:28shift logical neither that one is
  291. 11:31accessing the memory
  292. 11:32but store word is writing into memory so
  293. 11:36the left hand side is is colored load
  294. 11:38word
  295. 11:39is reading from the memory so the right
  296. 11:40hand side is
  297. 11:43shaded and add immediate is not
  298. 11:46accessing the memory
  299. 11:49so um
  300. 11:53we are you know there is not much
  301. 11:57to to add to here this is the time to
  302. 11:59execute the instruction
  303. 12:01it's a thousand picoseconds or one
  304. 12:03nanosecond
  305. 12:04and the cycle time is one-fifth of that
  306. 12:07or 200 picoseconds
  307. 12:10when we would like to know what is
  308. 12:14happening here
  309. 12:15in in sequentially it is
  310. 12:19the usage of resources in the pipeline
  311. 12:22or stages in the pipeline
  312. 12:24by one instruction resource use of
  313. 12:27instruction is sequential in time each
  314. 12:30instruction
  315. 12:31goes through the fetch
  316. 12:34decode and register access execution
  317. 12:37phase
  318. 12:37memory access and right back on the
  319. 12:40other hand
  320. 12:41multiple instructions are
  321. 12:44using different
  322. 12:47resources at the
  323. 12:51the the the same time so there are
  324. 12:54five instructions here that we will say
  325. 12:57in flight
  326. 12:58that are using different resources
  327. 13:00because they're available
  328. 13:01so uh these resources here the
  329. 13:04first instruction is near its completion
  330. 13:07so add is writing back the result into
  331. 13:10t0
  332. 13:12or is in data access
  333. 13:16stage it's not doing anything it's
  334. 13:18basically just waiting
  335. 13:20uh for the ad to be done so it can write
  336. 13:22back
  337. 13:23its result in t3 shift left logical
  338. 13:26is using the alu to perform the shift
  339. 13:31store word is
  340. 13:35essentially fetching the t3 and t0
  341. 13:39values from the t3 and t0 register such
  342. 13:42that it can use them in the store
  343. 13:43operation
  344. 13:44load word is just being fetched
  345. 13:48from the memory and we don't even know
  346. 13:50that it's a load
  347. 13:52and that's it that we have to what we
  348. 13:55need to know for now
  349. 13:57conceptually about the pipelining of the
  350. 13:59smile processor it is very similar
  351. 14:03to running laundry we're going to see
  352. 14:06what in the next module what do we need
  353. 14:08to do
  354. 14:09with our data path to support that
  355. 14:12so see you then after a bit of a break

About this transcript

This page contains the full transcript of [CS61C FA20] Lecture 22.1 - Pipelining II: Pipelining RISC-V by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 1,881 words across 355 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.