YouTube2Text

[CS61C FA20] Lecture 23.3 - Pipelining III: Superscalar — Transcript

by CS 61C Departmental · 1,908 words · 365 segments · language en · Watch on YouTube

Full transcript

  1. 0:01[Music]
  2. 0:12hi
  3. 0:12welcome back to our pipelining module
  4. 0:16we're almost done we have learned how
  5. 0:20to design a pipeline processor we
  6. 0:22actually learned the principles of
  7. 0:23designing a pipeline processor
  8. 0:26[Music]
  9. 0:27you will get the true feeling for how to
  10. 0:30design one
  11. 0:31when doing a project and actually going
  12. 0:34through the design of a pipeline
  13. 0:36and resolution of all the hazards
  14. 0:39that's really fun actually
  15. 0:43but for now we understood all the
  16. 0:46principles that we
  17. 0:47need to deal with in order to design a
  18. 0:50good processor
  19. 0:53there is one more advanced concept that
  20. 0:55is out there that's a concept of
  21. 0:56so-called
  22. 0:57super scalar processors you might have
  23. 1:00heard about that
  24. 1:01many of the modern very high performance
  25. 1:04processors
  26. 1:04are of super scalar type
  27. 1:08before we get to we get to those let's
  28. 1:11see
  29. 1:12how can we increase processor
  30. 1:15performance further
  31. 1:16from what we know so far
  32. 1:21so we have designed a five-stage
  33. 1:23pipeline and what is
  34. 1:24kind of interesting this five-stage
  35. 1:26pipeline is the industry workhorse
  36. 1:28nowadays
  37. 1:29probably 80 or 90 percent of processors
  38. 1:31that are out there
  39. 1:32all around us are five-stage pipelines
  40. 1:36but they don't have the highest possible
  41. 1:38performance
  42. 1:40there are many designs
  43. 1:43that are
  44. 1:46different their higher performance and
  45. 1:49when i say majority of processors out
  46. 1:51there are five stage pipelines
  47. 1:53you'll find them in all kinds of
  48. 1:55applications in
  49. 1:56cars and various appliances
  50. 2:00and so on they're really common they're
  51. 2:02like that they're in they're these
  52. 2:04invisible computers they're all over us
  53. 2:06all around us but those that will find
  54. 2:08in laptops or in desktops or even in
  55. 2:11cell phones
  56. 2:13there are usually a few
  57. 2:17higher performance processors and a lot
  58. 2:19of these
  59. 2:20five stage pipelines as well
  60. 2:23so how do we crank up the performance of
  61. 2:25a processor
  62. 2:26we have a few options one we can try to
  63. 2:28crank up the clock frequency
  64. 2:31but the clock frequency is limited by
  65. 2:34two things
  66. 2:35how fast transistors will switch and
  67. 2:37that is said by the technology
  68. 2:39and the other one is the the power
  69. 2:42dissipation we have mentioned that
  70. 2:44before
  71. 2:45um you know cranking up the clock
  72. 2:47frequency was our
  73. 2:48primary means of adding performance in
  74. 2:51the 90s
  75. 2:52up to early 2000s but then we hit the
  76. 2:55power ceiling
  77. 2:56and we discovered our designers
  78. 2:59discovered that they can design
  79. 3:01processors that will
  80. 3:02run very fast at very high clock
  81. 3:05frequencies
  82. 3:06but there was no practical means of
  83. 3:09cooling them
  84. 3:11so the clock frequency had to be held
  85. 3:14constant
  86. 3:14to allow for reasonable cooling
  87. 3:18then the second option that's often
  88. 3:20works in
  89. 3:22concert with clock rate increases
  90. 3:25is the concept of deeper pipelining
  91. 3:29so we have seen our five stage pipeline
  92. 3:31but we can break
  93. 3:33up the execution into smaller and
  94. 3:35smaller chunks so
  95. 3:36people have built pipelines that are 10
  96. 3:39stages deep or 15 stages deep
  97. 3:42intel processors i think went up to 20
  98. 3:44stages deep
  99. 3:46at some point so the idea here is to do
  100. 3:50let less work in each cycle to have less
  101. 3:53logic depth to go through less larger
  102. 3:55gates
  103. 3:56so we can have a shorter clock cycle and
  104. 3:59that enables
  105. 4:00us to either run at a higher frequency
  106. 4:02or if we are
  107. 4:03limited with power to drop the supply
  108. 4:06voltage
  109. 4:06and run at the same clock frequency but
  110. 4:09with lower power
  111. 4:12however the deeper the pipeline
  112. 4:15the higher the chance is that we are
  113. 4:17going to run into hazards
  114. 4:18so our cpi is naturally going to
  115. 4:22get higher to be higher than one
  116. 4:25so what came as an alternative to that
  117. 4:28are so-called multi-issue
  118. 4:30or superscalar processors those are the
  119. 4:32ones
  120. 4:33that theoretically can have cpi that
  121. 4:37is much less than one so how do those
  122. 4:40work
  123. 4:40in a super scalar processor we have
  124. 4:43multiple issues
  125. 4:44what does it mean we have multiple
  126. 4:46execution units
  127. 4:47and everything else that helps feed
  128. 4:50those execution units
  129. 4:52so each execution units is actually a
  130. 4:54pipeline on its own that is capable of
  131. 4:57executing either integer instructions or
  132. 5:00floating point instructions or we may
  133. 5:02have dedicated
  134. 5:04load and store pipelines
  135. 5:08so since we have multiple execution
  136. 5:10units
  137. 5:11we can fetch multiple instructions per
  138. 5:14clock cycle
  139. 5:15and so-called issue them to different
  140. 5:18execution
  141. 5:19pipelines um so this enables us
  142. 5:23to drop our
  143. 5:26cpi below one so we are going to be
  144. 5:30completing
  145. 5:31multiple instructions per clock cycle
  146. 5:35often people will use a concept ipc such
  147. 5:38that
  148. 5:39it's easier to to understand things to
  149. 5:42compare things that are
  150. 5:43greater than one than one for example if
  151. 5:47you have a
  152. 5:48four gigahertz four-way multiple issue
  153. 5:50processor
  154. 5:51it is capable of executing 16 billions
  155. 5:54of instructions
  156. 5:56per second uh with it would have a
  157. 5:59peak cpi of 0.25
  158. 6:02and peak ipc would be one over that
  159. 6:05which
  160. 6:06is four peak means
  161. 6:09that if there are no conflicts there are
  162. 6:12no dependencies
  163. 6:13no hazards we can get to that peak
  164. 6:17cpi or ipc but in practice that number
  165. 6:20is not going to be so great
  166. 6:26often these processors employ something
  167. 6:28that is called out of order execution
  168. 6:31so in order to deal with the
  169. 6:33dependencies and hazards
  170. 6:34there'll be a hardware unit they'll be
  171. 6:36picking and
  172. 6:38figuring out which instructions have
  173. 6:39dependencies
  174. 6:41and trying to execute them in an order
  175. 6:43where they don't depend on each other
  176. 6:45and then there'll be a reorder unit
  177. 6:48at the at the end of the pipeline that
  178. 6:50will put the results back
  179. 6:52in order such that whoever run that
  180. 6:55program
  181. 6:56gets meaningful results and that is all
  182. 6:59done perfectly
  183. 7:02this we definitely are not going to
  184. 7:04cover in 61c
  185. 7:06but cs152 goes into a great depth of
  186. 7:10that
  187. 7:11what we'll we will do is a quick
  188. 7:14overview of what you will find in a
  189. 7:17superscalar processor
  190. 7:18so generally there will be an
  191. 7:20instruction fetch and decode that will
  192. 7:22be
  193. 7:23fetching and decoding multiple
  194. 7:25instructions per cycle
  195. 7:28and then it would go and issue them
  196. 7:32in different execution pipelines
  197. 7:36um it would recognize what kind of an
  198. 7:38instruction it is
  199. 7:40and it would send it down the integer
  200. 7:42pipelines or floating point
  201. 7:44pipelines or load store pipelines
  202. 7:47it would use some kind of reservation
  203. 7:49station that would be there to
  204. 7:52reserve the resources the reserve to
  205. 7:54reserve
  206. 7:55this functional unit for future use
  207. 7:58because
  208. 7:59they the the processor will try to keep
  209. 8:01them all filled
  210. 8:02as much as possible and finally
  211. 8:06the instructions will be committed or
  212. 8:08retired
  213. 8:09in order so they'll be arriving out of
  214. 8:11order but coming out of this commit unit
  215. 8:14in order
  216. 8:17here is an example of an
  217. 8:20um out of order superscalar processor
  218. 8:23which is
  219. 8:24intel i7 a relatively modern processor
  220. 8:27that
  221. 8:28you'll find in laptops like the one that
  222. 8:31i'm using now for
  223. 8:32this presentation
  224. 8:35it is what is shown here is the cpi
  225. 8:40for a number of benchmarks here and
  226. 8:43these are different kinds of benchmarks
  227. 8:44you don't really need to know
  228. 8:46to do need to know which kinds are those
  229. 8:49but this has something to do with
  230. 8:51some quantum mechanics simulations h264
  231. 8:55is video compression
  232. 8:58bzip is a zip
  233. 9:01type of a program that is compressing
  234. 9:04some files
  235. 9:05you'll find a go game and then gcc that
  236. 9:09is compiling some kind of a target code
  237. 9:12and then they measure what is the cpi
  238. 9:16that is achieved for these benchmarks
  239. 9:21and what you'll find out here is that
  240. 9:24these cpi's in most cases will drop
  241. 9:28below one so there will be less than one
  242. 9:32and in some cases maybe more than one
  243. 9:37there is another interesting thing to
  244. 9:38notice on this graph there are
  245. 9:41there is ideal cpi that is sitting
  246. 9:43around
  247. 9:440.25 down here
  248. 9:49and then there is realistic cpi that
  249. 9:52rarely gets below
  250. 9:530.5 so the the the ideal cpi is 0.25
  251. 9:57or about a quarter and realistic one is
  252. 10:000.44 and so on this realistic cpi
  253. 10:04is due to real cpi is due to stalls
  254. 10:08misspeculation and handling hazards
  255. 10:13by the way how is this cpi number
  256. 10:15actually achieved
  257. 10:16or measured it's not something magic
  258. 10:20it's just the reverse of our iron law of
  259. 10:23processor performance
  260. 10:24so what we have there in there we had
  261. 10:27our cpi which stands for cycles per
  262. 10:29instruction
  263. 10:30and the cpi
  264. 10:34can be evaluated by knowing all the
  265. 10:36other
  266. 10:37elements of this equation so we can time
  267. 10:40the program
  268. 10:41uh how long does it take us to execute
  269. 10:42the program then we can
  270. 10:45count the number of instructions that
  271. 10:46that benchmark has
  272. 10:48and finally we can look up um
  273. 10:51how long would the does it take to
  274. 10:54execute a cycle
  275. 10:56that's the frequency of a processor that
  276. 10:58give us gives us the number of cycles
  277. 11:00per instruction
  278. 11:01and that is straight time to execute the
  279. 11:04program
  280. 11:05divided by instructions for a program
  281. 11:07times the time
  282. 11:10per cycle
  283. 11:14so that's it this is the way how these
  284. 11:17cpi numbers average cpi numbers have
  285. 11:20been
  286. 11:22measured for i7 or many other processors
  287. 11:30a few notes about the design of isa
  288. 11:33and which isas are good for for
  289. 11:35pipelining so risk 5 is a
  290. 11:37a type of a risk
  291. 11:40i say that is designed with pipelining
  292. 11:43in mind
  293. 11:44and there are several features that you
  294. 11:46will recognize that are in there
  295. 11:49that are very helpful for designing a
  296. 11:51pipeline
  297. 11:52for example in our version uh
  298. 11:5532 version all
  299. 11:59instructions and registers are 32 bits
  300. 12:02but in every case
  301. 12:03in every variant of risk 5 all
  302. 12:07instructions are 32 bits
  303. 12:10so all these instructions are easy
  304. 12:14to decode in one cycle
  305. 12:18compare that to x 86
  306. 12:21their instructions can be one
  307. 12:25or anything up to 15 bytes
  308. 12:28so decoding is really complicated there
  309. 12:32so
  310. 12:32you know take a look at one instruction
  311. 12:34you may be able to
  312. 12:35decode what is what is in there what are
  313. 12:37you supposed to do
  314. 12:38but more likely you have to go and um
  315. 12:42oh give me a few more bytes oh can i
  316. 12:44decode not yet
  317. 12:47a few more bytes and finally after 15 of
  318. 12:49them the most complex instructions can
  319. 12:51be decoded
  320. 12:54risk 5 has a small number six different
  321. 12:57instruction formats
  322. 12:58they're very easy to decode and
  323. 13:02read registers in one step so in one
  324. 13:05stage we can do both
  325. 13:07many more complex isas require us to do
  326. 13:09that in multiple steps
  327. 13:13load store addressing can be done in
  328. 13:16third stage
  329. 13:17by using the alu and access the memory
  330. 13:20in the fourth stage
  331. 13:21and then memory operands are all all
  332. 13:23aligned and take
  333. 13:25only one cycle that is something that is
  334. 13:28very convenient
  335. 13:31and we have taken advantage of all of
  336. 13:32this when designing our pipeline
  337. 13:37and that's it so we've
  338. 13:41we have done it we got our pipeline we
  339. 13:43got our working processor
  340. 13:45this is capable of executing all rv32i
  341. 13:49instructions first in one cycle
  342. 13:53and then we figured out how to pipeline
  343. 13:55it
  344. 13:56we identify these five phases of
  345. 13:59execution
  346. 14:00that we associate with five stages of a
  347. 14:02pipeline
  348. 14:04and we have designed this controller
  349. 14:06first for a single stage pipeline but
  350. 14:08then
  351. 14:09we have outlined the principles for a
  352. 14:11single stage execution i'm sorry
  353. 14:13and uh then we i identify how we need to
  354. 14:17modify it
  355. 14:18to support uh resolution of hazards in a
  356. 14:21pipeline
  357. 14:24pipelining improves the performance
  358. 14:27and opens an avenue to building
  359. 14:30something that is really
  360. 14:32really powerful we're going
  361. 14:36to wrap up this module now in the next
  362. 14:39module
  363. 14:39we are going to continue building better
  364. 14:42and better computer
  365. 14:45see you after a bit of a longer break

About this transcript

This page contains the full transcript of [CS61C FA20] Lecture 23.3 - Pipelining III: Superscalar by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 1,908 words across 365 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.