YouTube2Text

[CS61C FA20] Lecture 22.4 - Pipelining II: Data Hazards — Transcript

by CS 61C Departmental · 2,149 words · 385 segments · language en · Watch on YouTube

Full transcript

  1. 0:01[Music]
  2. 0:11hello and welcome back to our pipelining
  3. 0:13module
  4. 0:14we are continuing to talk about hazards
  5. 0:17we have
  6. 0:18identified three types of hazards
  7. 0:21namely structural hazards data hazards
  8. 0:24and control hazards we have talked about
  9. 0:26structural hazards
  10. 0:28which are generally addressed by the
  11. 0:30data path design
  12. 0:32that identifies all the needs
  13. 0:36that that the isa presents
  14. 0:40so the data path is provisioned to
  15. 0:42support
  16. 0:43efficient pipelining
  17. 0:47data hazards are a little bit different
  18. 0:49and
  19. 0:50let's try to first understand them and
  20. 0:53then
  21. 0:54figure out how do we address them
  22. 0:59so we have talked about a structural
  23. 1:01hazard or
  24. 1:02of having to address the memory
  25. 1:06for both instructions and data
  26. 1:11there is a hazard also that affects the
  27. 1:13register file access
  28. 1:15so in there there may be instructions
  29. 1:18like this one add or there are
  30. 1:20definitely instructions like the add
  31. 1:22that writes the value into
  32. 1:26the register t0 that is going to happen
  33. 1:29during the write-back phase here when
  34. 1:31the register file
  35. 1:32is accessed so a register file
  36. 1:37well during the write-back phase is
  37. 1:38going to be
  38. 1:40accessed for writing that's why it is
  39. 1:42shaded
  40. 1:43on the left hand side now at some point
  41. 1:47an instruction like this one store word
  42. 1:51will be in the pipeline in the
  43. 1:54instruction decode phase
  44. 1:55when it will also
  45. 1:59wish to read the register file so the
  46. 2:02register file is going to be
  47. 2:04accessed for both reading
  48. 2:07and writing
  49. 2:11and it will happen that a ad
  50. 2:15instruction is going to be writing a
  51. 2:17value of t0
  52. 2:18while store would also like to
  53. 2:22read the t0 so will star instruction end
  54. 2:26up reading
  55. 2:27the the old value that is in there or
  56. 2:30should it read the new value as it's
  57. 2:32supposed to
  58. 2:34the answer to that is
  59. 2:37that the register files are designed to
  60. 2:39support single cycle
  61. 2:41read write operation what does that mean
  62. 2:45that means that the add instruction
  63. 2:49that we have here is going to write
  64. 2:52during the first phase of this cycle
  65. 2:54right back in complete
  66. 2:56writing in 100 picoseconds as
  67. 2:59we said before then in the remaining 100
  68. 3:03picoseconds of a 200
  69. 3:04picosecond clock cycle
  70. 3:08store word is going to read t0 out of
  71. 3:10the register file
  72. 3:15so this is very much dependent on the
  73. 3:17implementation
  74. 3:18on circuit and logic implementation of a
  75. 3:20register file that needs to support
  76. 3:22writing and reading in the same cycle
  77. 3:25that avoids this data hazard that
  78. 3:28happens
  79. 3:29during the register update so if that is
  80. 3:31true
  81. 3:32um right back we'll up the stage will
  82. 3:34update
  83. 3:35the value in the register and the next
  84. 3:37instruction that is accessing it during
  85. 3:39the instruction decode
  86. 3:40is going to get the new value and that
  87. 3:44that will be indicated in this diagram
  88. 3:47where we have the right backs the
  89. 3:51right back phase of the add shading the
  90. 3:54first part of the the left hand part of
  91. 3:57the register
  92. 3:59access and then the stored word shades
  93. 4:02the right hand side
  94. 4:03now keep in mind that this is certainly
  95. 4:06true for most of
  96. 4:08standard five stage pipelines but may
  97. 4:10not be true for some
  98. 4:11more complicated higher frequency
  99. 4:13pipelines
  100. 4:16we're not touching those for now they
  101. 4:19may be a subject of
  102. 4:20152 if you take that class but for now
  103. 4:24um just make sure that you clearly read
  104. 4:26the assumptions that are stated in any
  105. 4:28exam or midterm questions that you see
  106. 4:31do they permit
  107. 4:32simultaneous read and write in one cycle
  108. 4:34simultaneous
  109. 4:35meaning in this say case consecutive
  110. 4:38that the rate is completed
  111. 4:39before we want to read from the register
  112. 4:43all right but let's take a look at a
  113. 4:46little bit more
  114. 4:47complex data hazard and here it is
  115. 4:52let's say that the add instruction
  116. 4:55this time writes its result into the
  117. 4:58destination register as zero
  118. 5:00then we have this kind of a crazy code
  119. 5:02down here that
  120. 5:04perhaps doesn't make any sense i haven't
  121. 5:05really checked if it does anything
  122. 5:06useful
  123. 5:07but every subsequent instruction would
  124. 5:09like to use
  125. 5:10that as 0 as the source so the sub use
  126. 5:13it as a source this or use it as a
  127. 5:15source this xor uses the source
  128. 5:17and stored word uses the source so all
  129. 5:20these
  130. 5:22instructions depend on add 0 completing
  131. 5:26what it's doing and writing back the
  132. 5:28value into
  133. 5:30the register file but let's take a look
  134. 5:32at what is happening here
  135. 5:34let's say that the value of 0 you know
  136. 5:36starting
  137. 5:37when this ad the instruction started its
  138. 5:40execution was five
  139. 5:42and it stays five and at this point the
  140. 5:45alu result
  141. 5:46is nine and this nine is supposed to be
  142. 5:48written back into a register file
  143. 5:49but that doesn't happen until
  144. 5:53mid through the right backstage of add
  145. 5:57so we still have five during the
  146. 5:59execution stage of the ad
  147. 6:02and the memory access stage of the ad
  148. 6:04instruction
  149. 6:05and for the first half of the
  150. 6:08right back stage so if the next
  151. 6:11instruction would like to use that
  152. 6:13when it goes to the register fetch
  153. 6:16stage to the register um
  154. 6:19to retrieve the values from the
  155. 6:21registers during the instruction decode
  156. 6:23stage what is it going to find
  157. 6:24five that's the wrong value that
  158. 6:28is not going to be good what is the or
  159. 6:31instruction going to find in the
  160. 6:33register file also five
  161. 6:35that's a problem now this xor is going
  162. 6:39to do better
  163. 6:40because nine is going to be available
  164. 6:43in the second five phase of
  165. 6:47the instruction decode
  166. 6:50when it is going to be read correctly
  167. 6:52and store word is also going to find the
  168. 6:54correct value
  169. 6:55now what are we going to do with these
  170. 6:57sub and or they are going to get
  171. 6:59incorrect values
  172. 7:01from the register file
  173. 7:05we need to do something otherwise we are
  174. 7:07executing an incorrect problem
  175. 7:09incorrect program so the first solution
  176. 7:12is so-called stalling what installing is
  177. 7:14like when you have an old car and you
  178. 7:16know you
  179. 7:17you you drive a lot and you stall and
  180. 7:20then you start again
  181. 7:21and you drive a little bit and stall
  182. 7:25so that's how your program would
  183. 7:26actually execute with the stalling
  184. 7:28um when the result of
  185. 7:33when the next instruction
  186. 7:36depends on the result of the previous
  187. 7:38instruction we cannot start its
  188. 7:40execution until the result is written
  189. 7:41back
  190. 7:42so if the sub follows the add
  191. 7:46the sub execution needs to be delayed by
  192. 7:48two cycles by inserting
  193. 7:50two bubbles in between these bubbles are
  194. 7:54essentially knobs
  195. 7:55that are going to correctly help
  196. 7:57correctly align
  197. 7:59the register right back stage with the
  198. 8:02register read
  199. 8:05the bubbles are perhaps
  200. 8:08not the most useful thing that is out
  201. 8:12there
  202. 8:13because during that time processor is
  203. 8:15doing nothing
  204. 8:17um although the reason you know we
  205. 8:19accomplished the main thing
  206. 8:21to get the right result the correct
  207. 8:23execution of a program
  208. 8:25what may happen is that we are going to
  209. 8:28lose performance
  210. 8:29now compilers generally try to do
  211. 8:31something to avoid that
  212. 8:33so you'll go through our list of
  213. 8:35instructions that are following that ad
  214. 8:37and try to find something some
  215. 8:38instruction it is not dependent on the
  216. 8:40result of that ad
  217. 8:41and it is going to try to swap that sub
  218. 8:44with the instruction that does not
  219. 8:45depend on the result of the ad
  220. 8:47so that way we are not going to lose any
  221. 8:49cycles
  222. 8:50but in case it can't find any
  223. 8:54instructions
  224. 8:55we have to install insert these knobs
  225. 8:57which will cause stalls
  226. 8:59such that we can correctly align the
  227. 9:03register read with the register right
  228. 9:06this knob is anything that writes into
  229. 9:09the x0
  230. 9:10by convention in risk 5 assembly
  231. 9:14knob is add immediate x 0 x 0 0.
  232. 9:20all right the second solution is
  233. 9:23a hardware solution that we need to
  234. 9:26implement
  235. 9:27by modifying the data path here is the
  236. 9:30thing
  237. 9:32when this ad instruction is
  238. 9:35at the end of its execute phase we have
  239. 9:37the value somewhere in the pipeline
  240. 9:39that is correct that is a correct value
  241. 9:41that should be read
  242. 9:43by the register files by by
  243. 9:46instructions sub and or from the
  244. 9:48register file
  245. 9:49but that value is not in the register
  246. 9:52file yet
  247. 9:52it is somewhere in the pipeline but it's
  248. 9:54not in the register file
  249. 9:56so the idea here is to take that value
  250. 10:00from the alu or a cycle later
  251. 10:04from the memory access
  252. 10:08stage and forward it to the appropriate
  253. 10:13input port of the alu so we'll take
  254. 10:17this output of the alu and
  255. 10:21essentially fast forward it to the input
  256. 10:23of the alu for the sub instruction
  257. 10:25and then sub is going to perform
  258. 10:27correctly sub is executing in the next
  259. 10:29cycle
  260. 10:29and we just need to make sure that this
  261. 10:31shows up straight at its input
  262. 10:34similarly um that result is going to
  263. 10:37propagate in the next cycle to the
  264. 10:39output of the
  265. 10:40uh to the end of the memory access phase
  266. 10:42and we would need to forward it
  267. 10:44to the input of the alu to correctly
  268. 10:46execute the or
  269. 10:48all right so we need to have these
  270. 10:50forwarding or so
  271. 10:52called bypass paths in our data path
  272. 10:55to enable data to be forwarded to the
  273. 10:58right place
  274. 10:59without being written to register file
  275. 11:01it's going to get
  276. 11:02written in the register file but this is
  277. 11:04just for the
  278. 11:05subsequent instructions to use use the
  279. 11:08result immediate
  280. 11:09immediately so
  281. 11:12um you know this is basically outlining
  282. 11:16what we need to do
  283. 11:17we need to basically have a short
  284. 11:19circuit a sort of a short circuit
  285. 11:21between the output of the alu here and
  286. 11:24the input
  287. 11:25of the alu
  288. 11:29that is something
  289. 11:32that is relatively straightforward to
  290. 11:36implement you you guess that there will
  291. 11:38be a multiplexer another multiplexer
  292. 11:40that is going to be sitting in front of
  293. 11:42the
  294. 11:46or together with the cell a operand
  295. 11:50by the way the forwarding here is just
  296. 11:51shown for the operand a to the alu we
  297. 11:54also need to forward
  298. 11:55the operand b
  299. 11:59now what do we need to control this
  300. 12:02forwarding this there will be a
  301. 12:03multiplexer and we need to somehow
  302. 12:04control that multiplexer so what kind of
  303. 12:06information do we have
  304. 12:08that is going to help us with that
  305. 12:09remember we
  306. 12:11our data path saves the instructions
  307. 12:13that are currently in flight so we'll
  308. 12:15know which registers are being accessed
  309. 12:17by the two consecutive instructions so
  310. 12:18what we need to do here
  311. 12:20is we need to take a look at
  312. 12:23the instruction that is the destination
  313. 12:26of this instruction that that is in the
  314. 12:28execute stage
  315. 12:30and compare it with the instruction that
  316. 12:33is
  317. 12:33in the instruction decode stage if
  318. 12:36they're the same
  319. 12:37and not take zero then we have a hazard
  320. 12:40we have a data hazard and we should
  321. 12:43forward
  322. 12:45right so we just compare those two and
  323. 12:47that is going to select our multiplexer
  324. 12:51uh forwarding multiplexer we are going
  325. 12:53to have the same comparison
  326. 12:55between this alu and between this
  327. 12:58alu and the right and memory access
  328. 13:01stage
  329. 13:02to allow us to forward again there so
  330. 13:05what we are seeing here the the
  331. 13:07multiplexer that selects the operand a
  332. 13:09is going to have its two standard values
  333. 13:12you know pc
  334. 13:13and the register file output and then
  335. 13:15they're going it is going to have two
  336. 13:16more forwarding paths
  337. 13:19um let's take a look at how this
  338. 13:22is this implemented in the data path
  339. 13:24this is our pipelined
  340. 13:25rv32i data path all the registers are
  341. 13:29where they're supposed to be
  342. 13:30so what we are what we did here we added
  343. 13:33these two
  344. 13:34forwarding paths the forwarding paths
  345. 13:36are not really in this picture
  346. 13:38uh shown that they're going forward they
  347. 13:40look like they're going backward
  348. 13:41but actually the they're going forward
  349. 13:44to the the execution of the next
  350. 13:46instruction they're bypassing in
  351. 13:48for the in in for the next instruction
  352. 13:51in this case
  353. 13:52from the output alu back to the input
  354. 13:55of the alu after one cycle delay
  355. 13:58and then after two cycle delays from the
  356. 14:02right back
  357. 14:02right back stage back to the input of
  358. 14:04the aldu
  359. 14:06and that is it all what we need to do is
  360. 14:08to implement the forwarding control
  361. 14:10logic
  362. 14:10this one is shown to be working only
  363. 14:13just
  364. 14:14on this multiplexer forwarding
  365. 14:17um you know comparing the the
  366. 14:21the registers that are being used in
  367. 14:23these two instructions that are in
  368. 14:25flight
  369. 14:25that are in the instruction decode in
  370. 14:27the execute stage we would need to make
  371. 14:30we need to control uh to figure out what
  372. 14:32is in instruction decode
  373. 14:34and the memory access and repeat all of
  374. 14:36that
  375. 14:37for the operand b as well and that is it
  376. 14:41that is the story of the forwarding the
  377. 14:44best thing here is to
  378. 14:45look at a few code examples and see
  379. 14:49when the forwarding is needed and
  380. 14:51implementation is
  381. 14:52relatively straightforward we're going
  382. 14:56to break here
  383. 14:56we are going to return to the control
  384. 14:59hazards
  385. 15:00after the break see you then

About this transcript

This page contains the full transcript of [CS61C FA20] Lecture 22.4 - Pipelining II: Data Hazards by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 2,149 words across 385 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.