YouTube2Text

[CS61C FA20] Lecture 38.2 - Dependability- Parity, ECC, RAID: Dependability Metrics — Transcript

by CS 61C Departmental · 2,004 words · 386 segments · language en · Watch on YouTube

Full transcript

  1. 0:04[Music]
  2. 0:10hello and welcome back to our
  3. 0:12dependability module
  4. 0:14the main theme of the module is that you
  5. 0:16would like to use
  6. 0:17redundancy to improve dependability
  7. 0:21but remember in every compute system we
  8. 0:23need to
  9. 0:24measure something to know whether and
  10. 0:27how much
  11. 0:27have we improved it so let's take a look
  12. 0:31at
  13. 0:31what are our dependability metrics
  14. 0:35so let's first understand how does a
  15. 0:37compute system operate
  16. 0:39so it in in normal operation it runs
  17. 0:44some services for us
  18. 0:48but it may happen that it will fail
  19. 0:51so the fault and this is very different
  20. 0:53than the page fault
  21. 0:54is a failure of a component in a compute
  22. 0:57system
  23. 0:59that component may or may not lead
  24. 1:02to a system failure depending whether
  25. 1:05it's being used
  26. 1:06or not and whether it is redundant
  27. 1:10and there is something else that can
  28. 1:11take over for example
  29. 1:13the system will operate essentially
  30. 1:15between the two states in one state
  31. 1:17it operates normally so the services are
  32. 1:19accomplished in the other state there is
  33. 1:21a service interruption
  34. 1:22moving from the normal operation to the
  35. 1:25service interruption
  36. 1:27is a failure and going back
  37. 1:30from the system interruption uh that's
  38. 1:32the deviation from
  39. 1:33the specified service back to the normal
  40. 1:36operation
  41. 1:37is the restoration what will
  42. 1:41matter here is how frequently does this
  43. 1:45happen
  44. 1:46and how long does it take us to get back
  45. 1:50when we talk about redundancy we are
  46. 1:54going to see this use of redundancy
  47. 1:58in time and in space when you talk about
  48. 2:01spatial redundancy
  49. 2:03that means that we have multiple copies
  50. 2:05of something we may have multiple copies
  51. 2:07of compute
  52. 2:08units or more frequently we're going to
  53. 2:10have multiple copies of data
  54. 2:12that multiple multiple copies of data
  55. 2:14may
  56. 2:15exist in dram or
  57. 2:18in disks in there and more commonly we
  58. 2:21are going
  59. 2:22to have some kind of an algebraic way
  60. 2:26of adding redundant bits such that we
  61. 2:28can recover
  62. 2:29from temporary errors in hard disks we
  63. 2:32are often
  64. 2:34keeping redundant
  65. 2:37devices that are going to help us
  66. 2:39recover from
  67. 2:40any kind of a failure so
  68. 2:45replicas of data are essentially spatial
  69. 2:48redundancy
  70. 2:50the other form of redundancy is temporal
  71. 2:52redundancy
  72. 2:54in temporal redundancy essentially
  73. 2:57and what does that mean is that if there
  74. 2:59is a
  75. 3:01failure if there is a temporary failure
  76. 3:03of a compute system
  77. 3:04and we have time to repeat the
  78. 3:06computation we
  79. 3:08may as well do that so if if we detect
  80. 3:12that the
  81. 3:13compute the documentation has failed or
  82. 3:15the service has not
  83. 3:16completed we
  84. 3:19go again and if we complete that and
  85. 3:23and we recover without any consequences
  86. 3:28we are all good so that is an example of
  87. 3:31temporal redundancy let's take a look at
  88. 3:35some of these dependability measures
  89. 3:38um how do you know what kind of a
  90. 3:41yardstick do we use
  91. 3:43to measure dependability so the first
  92. 3:46one
  93. 3:47and most commonly used one is the mean
  94. 3:49time to failure
  95. 3:50or for short mttf
  96. 3:55that basically measures how long does it
  97. 3:58take
  98. 3:59us take take us for a device
  99. 4:02to fail or between the two
  100. 4:05failures of the same device if if they
  101. 4:08are recoverable
  102. 4:11the other metric is the service
  103. 4:12interruption and that service
  104. 4:14interruption
  105. 4:15is essentially how long do we state in
  106. 4:17that interrupted state
  107. 4:20that is the mean time to repair
  108. 4:25mean time between the failures is the
  109. 4:28sum
  110. 4:28of mean time to failure and mean time to
  111. 4:31repair
  112. 4:33and then we can define something that is
  113. 4:36also very important which is
  114. 4:38availability of a system availability of
  115. 4:40a system is simply
  116. 4:42mean time to failure divided by the sum
  117. 4:45of the mean time to failure and the mean
  118. 4:48time to repair
  119. 4:52how we can how can we improve
  120. 4:55availability of our system well we can
  121. 4:56either work
  122. 4:57on increasing the numerator or
  123. 5:00decreasing the denominator what does it
  124. 5:03mean
  125. 5:04to increase mean time to uh
  126. 5:07failure so the components are going to
  127. 5:10stay up for a longer time usually we are
  128. 5:12going to use
  129. 5:14higher quality more reliable hardware
  130. 5:18and software and we are generally going
  131. 5:21to have
  132. 5:22some kind of a redundancy there
  133. 5:25in addition to that we will be adding
  134. 5:28some kind of a fault tolerance such that
  135. 5:29if a component fails
  136. 5:32our system does not crash
  137. 5:35the other approach in reducing the
  138. 5:38denominator
  139. 5:39is to reduce the mttr
  140. 5:42mean time to repair basically it's about
  141. 5:46improving the tools that have detected
  142. 5:48that something is not
  143. 5:49working quite right in the system and
  144. 5:52bring it back
  145. 5:53to the normal mode operation
  146. 5:57so that's about dependability measures
  147. 5:59the next thing that you would like to
  148. 6:01define
  149. 6:01are these availability measures or
  150. 6:04availability
  151. 6:05metrics availability
  152. 6:08as we have defined it already is
  153. 6:11the ratio of mean time to failure
  154. 6:14divided by the sum of mean time to
  155. 6:15failure
  156. 6:16and the mean time to repair but it is
  157. 6:19generally
  158. 6:20expressed as a percentage of time and
  159. 6:23both mean time
  160. 6:24to failure or mean time to between the
  161. 6:27failures
  162. 6:28are measured in hours now
  163. 6:32compute systems smaller computer systems
  164. 6:34rarely go down
  165. 6:36um you know at least this is systems
  166. 6:38that we are used to
  167. 6:39rarely go down so we
  168. 6:43generally assume that it is going to be
  169. 6:46up by
  170. 6:47a very large percentage of time in a
  171. 6:50large
  172. 6:50percentage of time meaning 99.99
  173. 6:54some percent of a time that number of
  174. 6:56nines
  175. 6:58[Music]
  176. 6:59is generally measured as number of nines
  177. 7:01availability
  178. 7:02per year what does that mean so if
  179. 7:04something is available 90
  180. 7:06of a time in one year that means it'll
  181. 7:09be
  182. 7:10down 40 pairs for 36 days
  183. 7:14if it is available 99 of a time that
  184. 7:17means
  185. 7:18that it would be down for 3.6 days
  186. 7:21uh you have 3.6 days of repair per year
  187. 7:24that's very annoying
  188. 7:25i mean very few people would tolerate
  189. 7:27something like that
  190. 7:31although in some other systems you know
  191. 7:33transportation but maybe okay
  192. 7:35you know that's how many days do
  193. 7:38buses that operate every day spend in a
  194. 7:40shop per year
  195. 7:43three nines means that the system would
  196. 7:45be down uh
  197. 7:46526 minutes uh for repairs per year
  198. 7:51that's annoying but getting closer to
  199. 7:53being acceptable
  200. 7:54four nines is 53 minutes of repairs per
  201. 7:58year
  202. 7:58and five nines is five minutes of repair
  203. 8:01per year
  204. 8:04that is getting closer
  205. 8:07to what we expect and that is something
  206. 8:10that we
  207. 8:10often encounter in these services that
  208. 8:13we actually use on the internet i mean
  209. 8:15that's like of the order four to five
  210. 8:18nines is the availability of
  211. 8:20say youtube um
  212. 8:23as you've noticed when watching these
  213. 8:26videos
  214. 8:26youtube sometimes goes down
  215. 8:30other systems are designed to have six
  216. 8:33or seven
  217. 8:34nines um you know things that we depend
  218. 8:37much more on
  219. 8:38and there is essentially a much higher
  220. 8:40price
  221. 8:42for that system going down um examples
  222. 8:45are
  223. 8:46you know the losses of youtube when
  224. 8:48youtube went down for a few hours
  225. 8:50some years ago now that is basically
  226. 8:53losses of
  227. 8:54millions per minute
  228. 9:00okay so that's about availability
  229. 9:03measures let's talk about reliability
  230. 9:05measures
  231. 9:07generally how we like to
  232. 9:11measure differently this reliability is
  233. 9:16through annualized failure rate or afr
  234. 9:21so how many failures do we have
  235. 9:25in a system that has many components per
  236. 9:28year for example
  237. 9:29let's say we operate a data center that
  238. 9:31has a thousand disks
  239. 9:32and each disk is characterized to have
  240. 9:35one thousand
  241. 9:36one hundred thousand uh hour
  242. 9:40in mean time to failure
  243. 9:44so let's see what does that mean how
  244. 9:45many disks actually
  245. 9:47are going to fail per year in this kind
  246. 9:50of a system
  247. 9:52so how do we translate one reliability
  248. 9:54measure
  249. 9:55of mttf to another reliability measure
  250. 9:58which is annualized failure rate this is
  251. 10:01something that is very important for
  252. 10:02data centers and data centers actually
  253. 10:04keep track of that
  254. 10:05so we have 8760 hours
  255. 10:09in a year and then we have a thousand
  256. 10:12disks that are operating
  257. 10:14for the whole year our mean time to
  258. 10:17failure is
  259. 10:19expressed in hours so we are going to
  260. 10:21divide that number with 100 000
  261. 10:23hours and what we end up with here is
  262. 10:2787.6 failed disks per year on the
  263. 10:29average
  264. 10:30you know so we have to round it up
  265. 10:32because it's not 0.6 discs that
  266. 10:34generally fails so it's 88 uh
  267. 10:37failed disks this translates to 8.8
  268. 10:43annualized failure rate
  269. 10:47um well that's not that crazy
  270. 10:51that number actually is close to
  271. 10:54what we see actually in practice
  272. 10:57google published a study in 2007
  273. 11:01that found that the actual afr's
  274. 11:04of individual tribes range
  275. 11:07from 1.7 in the first year and as they
  276. 11:11age to go to eight point six percent
  277. 11:14over three
  278. 11:16four you know three-year-old drives
  279. 11:20kind of interesting uh thing to see
  280. 11:24there are interesting websites one of
  281. 11:26the these
  282. 11:28online backup services
  283. 11:31backplates actually publishes the
  284. 11:33failure rates of their drives and that
  285. 11:36may be useful information that you can
  286. 11:37take a look at when
  287. 11:38you would like to buy a new drive and
  288. 11:41they're
  289. 11:42publishing these afr's
  290. 11:45and they turn out to be pretty good for
  291. 11:48modern drives they are
  292. 11:49either single-digit or below
  293. 11:51single-digit percentages
  294. 11:54you know translating essentially to
  295. 11:55about a million hours
  296. 11:57of the order a million hours between the
  297. 12:00failures
  298. 12:01and another piece of good news i
  299. 12:03mentioned this before
  300. 12:05that this drives actually are getting
  301. 12:07better they have
  302. 12:08less failures over time slowly
  303. 12:11but surely the the afr
  304. 12:15is decreasing there is a decreasing
  305. 12:17trend
  306. 12:18over here
  307. 12:23there is another um
  308. 12:26reliability measure that is out there um
  309. 12:30that is used this was used for different
  310. 12:32kinds of systems but
  311. 12:34compute systems have to adhere for that
  312. 12:36it is failures in time
  313. 12:38or fit rate
  314. 12:42that has been used for example for
  315. 12:45cars and failures in time
  316. 12:50or fit rate of a device is the number of
  317. 12:54failures that can be
  318. 12:55expected in one billion device hours of
  319. 12:58operation
  320. 13:00that sounds like a lot i mean one
  321. 13:02billion hours is something that sounds
  322. 13:04unimaginable but imagine that
  323. 13:08we have many of these devices operating
  324. 13:10at the same time
  325. 13:13so that means 1000 devices for 1 million
  326. 13:17hours or 1 million devices for
  327. 13:211 000 hours each
  328. 13:24would be end up failing or 1 billion
  329. 13:28devices which is
  330. 13:29approximately how many smartphones we
  331. 13:30have out there
  332. 13:32would be failing every hour
  333. 13:36and that sounds about right i mean if
  334. 13:38you are looking at the whole world i
  335. 13:40think the failure rate of cell phones
  336. 13:42out there is probably
  337. 13:43greater you know there is more than one
  338. 13:45uh self-funded fails per hour
  339. 13:50so mtbf is equal to
  340. 13:541 billion times 1 over the fit rate
  341. 13:58this is relevant because computers are
  342. 14:01getting more and more into our cars
  343. 14:03and all these specifications for
  344. 14:04reliability there are
  345. 14:06expressed in fit rates so you you'll see
  346. 14:09that
  347. 14:10especially if you get into that kind of
  348. 14:12a business
  349. 14:14um you're going to hear about this
  350. 14:16automotive safety
  351. 14:17integrity level acell that defines
  352. 14:19different fit rates
  353. 14:21for different classes of components in
  354. 14:23vehicles
  355. 14:25okay so design principles of
  356. 14:28dependability are as follows
  357. 14:33there should be when we are designing a
  358. 14:35system that
  359. 14:36we should depend on um there should be
  360. 14:38no single
  361. 14:39point of failure and the analogy of that
  362. 14:43in in
  363. 14:44common terms is that the chain is only
  364. 14:46as strong as its weakest link
  365. 14:48there is a pretty good
  366. 14:51um corollary here of all amdahl's law
  367. 14:55it's always going to be the longest pole
  368. 14:57on the tent that is going to be sticking
  369. 14:59out
  370. 15:00so no matter you know you can keep
  371. 15:03improving
  372. 15:04the dependability of one component but
  373. 15:08the one that is the least dependable is
  374. 15:10going to dominate
  375. 15:12so we generally have to be measuring
  376. 15:16that for the whole system
  377. 15:17and make sure
  378. 15:20that the least reliable component or
  379. 15:23least dependable component
  380. 15:25is dependable enough all right
  381. 15:28so we are going to make a quick break
  382. 15:31here and then we are going to take a
  383. 15:33look
  384. 15:33into the ways how do we actually detect
  385. 15:36these faults
  386. 15:37or failures see you after a break

About this transcript

This page contains the full transcript of [CS61C FA20] Lecture 38.2 - Dependability- Parity, ECC, RAID: Dependability Metrics by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 2,004 words across 386 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.