YouTube2Text

[CS61C FA20] Lecture 38.1 - Dependability- Parity, ECC, RAID: Intro — Transcript

by CS 61C Departmental · 926 words · 173 segments · language en · Watch on YouTube

Full transcript

  1. 0:03[Music]
  2. 0:11hello
  3. 0:12welcome back to 61c it's a brand new
  4. 0:15module
  5. 0:16on dependability it's a fairly short
  6. 0:19module
  7. 0:20we're going to just touch on it remember
  8. 0:24our early story about six great ideas
  9. 0:27in computer architecture so we have
  10. 0:30talked about
  11. 0:31the layers of abstraction that enable us
  12. 0:33to build these complex systems
  13. 0:35out of many many components
  14. 0:38and we do that by layering many
  15. 0:41layers of abstraction we have talked
  16. 0:44about moore's law that enabled us to
  17. 0:46integrate so many devices and build
  18. 0:48these complex systems
  19. 0:50we have talked about principle of
  20. 0:52locality
  21. 0:53and memory hierarchy that enabled us to
  22. 0:56build these memory systems that look
  23. 0:59infinitely fast and infinitely big
  24. 1:02we have talked about parallelism as a
  25. 1:04way to improve the performance
  26. 1:07in the in power limited regime
  27. 1:12we have also throughout the course
  28. 1:14talked about
  29. 1:15the this concept of of performance
  30. 1:18measurement and improvements
  31. 1:20primarily through the iron law of
  32. 1:23compute
  33. 1:24and average memory access times
  34. 1:27but we haven't talked about the sixth
  35. 1:30great idea
  36. 1:31which is dependability by a redundancy
  37. 1:35so that's what we're going to cover now
  38. 1:38and why is that important well we
  39. 1:42rely on our compute systems more and
  40. 1:46more
  41. 1:47in everyday life and in domains that are
  42. 1:51just
  43. 1:52not purely computing all
  44. 1:55our financial transactions are performed
  45. 1:58by computers
  46. 1:59we drive a collection of computers
  47. 2:03many people's lives are supported by
  48. 2:07some sort of a computer so
  49. 2:10we really need to care about that
  50. 2:13dependability
  51. 2:14but the fact is that the computers fail
  52. 2:18they might fail transiently and we have
  53. 2:21seen before i
  54. 2:23i'm afraid these blue screens of that
  55. 2:26and many other failure modes for
  56. 2:29for computers often after a computer
  57. 2:32crashes
  58. 2:35um it can come back and we'll call that
  59. 2:39a transient mode of failure this failure
  60. 2:42may happen
  61. 2:43because of a bad code well you know
  62. 2:45somebody you know
  63. 2:46it's not just students who forget to
  64. 2:49clean the exit
  65. 2:50functions production software does that
  66. 2:53sometimes as well they do
  67. 2:55you know send you know new release and
  68. 2:56it works better but these things happen
  69. 2:59but sometimes these failures transit
  70. 3:01failures happen because
  71. 3:03something did not go quite right in the
  72. 3:06hardware
  73. 3:07hardware made an error
  74. 3:10if these errors persist something really
  75. 3:14fails
  76. 3:15in permanently in a computer
  77. 3:18then often we discard them or we try to
  78. 3:22repair them
  79. 3:24if they are repairable
  80. 3:27so in this module we are going to spend
  81. 3:30some time
  82. 3:32talking about how do we mitigate these
  83. 3:35hardware failures by using
  84. 3:38redundant components so
  85. 3:42we have briefly touched on that early on
  86. 3:45in the in the introductory module
  87. 3:47when we can use
  88. 3:51redundancy to replace a failing
  89. 3:55part of a system this redundancy
  90. 3:58is often encountered in memory systems
  91. 4:00we have spare stuff in memories
  92. 4:02um but we
  93. 4:06in parallel computers we may have a
  94. 4:09spare processor core you'll find out
  95. 4:13that on the market you can buy now
  96. 4:17eight core gpus or seven
  97. 4:20core gpus well when you look at the chip
  98. 4:23they're exactly the same
  99. 4:25except that eight core was disabled
  100. 4:29and they'll sell you that chip a little
  101. 4:32bit
  102. 4:33cheaper why was it disabled because it
  103. 4:35was no good it was failing
  104. 4:37that's okay you just have a little bit
  105. 4:39less of a performance you have seven
  106. 4:40instead of eight
  107. 4:41and you're willing to pay less for that
  108. 4:45so the way how it works is
  109. 4:48that you'll have multiple replicas
  110. 4:52of hardware and there'll be some kind of
  111. 4:55a voting mechanism
  112. 4:56where two out of three perhaps will
  113. 4:59agree
  114. 5:00that one plus one is equal to two and
  115. 5:02we'll take that
  116. 5:04as an answer this is
  117. 5:07this is made a lot easier by
  118. 5:10the advances in integration transistor
  119. 5:13integration densities
  120. 5:14so we can have more of these components
  121. 5:18now there is another cache there as the
  122. 5:20components are smaller as the threshold
  123. 5:21or smaller
  124. 5:22they have a tendency you know to more
  125. 5:24frequently fail
  126. 5:26but that is not increasing the
  127. 5:30that the failure rate is not increasing
  128. 5:32at the speed as which at which we can
  129. 5:34integrate more of them
  130. 5:35so we can add redundancy the other way
  131. 5:40how we can improve things
  132. 5:43how we can mitigate these transient
  133. 5:46failures
  134. 5:47is through
  135. 5:50another form of redundancy
  136. 5:54we can temporarily you know things can
  137. 5:57temporarily go out
  138. 5:58and we can still
  139. 6:02continue computing without them for
  140. 6:04example you know at the very
  141. 6:06top level the data center may go out i
  142. 6:08mean
  143. 6:09there may be a hurricane um on
  144. 6:12the east coast
  145. 6:15the power may be out the internet may be
  146. 6:17out data center goes out
  147. 6:19the weather system goes away data center
  148. 6:22is back up
  149. 6:24the other thing that we have encountered
  150. 6:26is often we use these arrays of disks
  151. 6:29if a disk if a mechanical disk fails
  152. 6:33it's okay because there is a way
  153. 6:36to cover for that disc by having
  154. 6:38redundancy
  155. 6:39on the shelf so a shelf of disk disks
  156. 6:43will have spares and finally
  157. 6:47we can use a particular type of coding
  158. 6:51to add
  159. 6:54redundancy to our dram so our dram chip
  160. 6:57this is a
  161. 6:58dim that goes into computers if you look
  162. 7:00at it it has
  163. 7:01nine memory chips that correspond to
  164. 7:03nine bits instead of
  165. 7:05eight and the reason for that is that
  166. 7:08the ninth one
  167. 7:08is there to represent parity and that
  168. 7:11parity is going to be used to indicate
  169. 7:13if something went wrong
  170. 7:16so we're going to take a quick break now
  171. 7:19and then we're going to take a
  172. 7:21look at how do we measure dependability
  173. 7:24see you after a quick break

About this transcript

This page contains the full transcript of [CS61C FA20] Lecture 38.1 - Dependability- Parity, ECC, RAID: Intro by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 926 words across 173 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.