[CS61C FA20] Lecture 38.2 - Dependability- Parity, ECC, RAID: Dependability Metrics — Transcript
Full transcript
- 0:04[Music]
- 0:10hello and welcome back to our
- 0:12dependability module
- 0:14the main theme of the module is that you
- 0:16would like to use
- 0:17redundancy to improve dependability
- 0:21but remember in every compute system we
- 0:23need to
- 0:24measure something to know whether and
- 0:27how much
- 0:27have we improved it so let's take a look
- 0:31at
- 0:31what are our dependability metrics
- 0:35so let's first understand how does a
- 0:37compute system operate
- 0:39so it in in normal operation it runs
- 0:44some services for us
- 0:48but it may happen that it will fail
- 0:51so the fault and this is very different
- 0:53than the page fault
- 0:54is a failure of a component in a compute
- 0:57system
- 0:59that component may or may not lead
- 1:02to a system failure depending whether
- 1:05it's being used
- 1:06or not and whether it is redundant
- 1:10and there is something else that can
- 1:11take over for example
- 1:13the system will operate essentially
- 1:15between the two states in one state
- 1:17it operates normally so the services are
- 1:19accomplished in the other state there is
- 1:21a service interruption
- 1:22moving from the normal operation to the
- 1:25service interruption
- 1:27is a failure and going back
- 1:30from the system interruption uh that's
- 1:32the deviation from
- 1:33the specified service back to the normal
- 1:36operation
- 1:37is the restoration what will
- 1:41matter here is how frequently does this
- 1:45happen
- 1:46and how long does it take us to get back
- 1:50when we talk about redundancy we are
- 1:54going to see this use of redundancy
- 1:58in time and in space when you talk about
- 2:01spatial redundancy
- 2:03that means that we have multiple copies
- 2:05of something we may have multiple copies
- 2:07of compute
- 2:08units or more frequently we're going to
- 2:10have multiple copies of data
- 2:12that multiple multiple copies of data
- 2:14may
- 2:15exist in dram or
- 2:18in disks in there and more commonly we
- 2:21are going
- 2:22to have some kind of an algebraic way
- 2:26of adding redundant bits such that we
- 2:28can recover
- 2:29from temporary errors in hard disks we
- 2:32are often
- 2:34keeping redundant
- 2:37devices that are going to help us
- 2:39recover from
- 2:40any kind of a failure so
- 2:45replicas of data are essentially spatial
- 2:48redundancy
- 2:50the other form of redundancy is temporal
- 2:52redundancy
- 2:54in temporal redundancy essentially
- 2:57and what does that mean is that if there
- 2:59is a
- 3:01failure if there is a temporary failure
- 3:03of a compute system
- 3:04and we have time to repeat the
- 3:06computation we
- 3:08may as well do that so if if we detect
- 3:12that the
- 3:13compute the documentation has failed or
- 3:15the service has not
- 3:16completed we
- 3:19go again and if we complete that and
- 3:23and we recover without any consequences
- 3:28we are all good so that is an example of
- 3:31temporal redundancy let's take a look at
- 3:35some of these dependability measures
- 3:38um how do you know what kind of a
- 3:41yardstick do we use
- 3:43to measure dependability so the first
- 3:46one
- 3:47and most commonly used one is the mean
- 3:49time to failure
- 3:50or for short mttf
- 3:55that basically measures how long does it
- 3:58take
- 3:59us take take us for a device
- 4:02to fail or between the two
- 4:05failures of the same device if if they
- 4:08are recoverable
- 4:11the other metric is the service
- 4:12interruption and that service
- 4:14interruption
- 4:15is essentially how long do we state in
- 4:17that interrupted state
- 4:20that is the mean time to repair
- 4:25mean time between the failures is the
- 4:28sum
- 4:28of mean time to failure and mean time to
- 4:31repair
- 4:33and then we can define something that is
- 4:36also very important which is
- 4:38availability of a system availability of
- 4:40a system is simply
- 4:42mean time to failure divided by the sum
- 4:45of the mean time to failure and the mean
- 4:48time to repair
- 4:52how we can how can we improve
- 4:55availability of our system well we can
- 4:56either work
- 4:57on increasing the numerator or
- 5:00decreasing the denominator what does it
- 5:03mean
- 5:04to increase mean time to uh
- 5:07failure so the components are going to
- 5:10stay up for a longer time usually we are
- 5:12going to use
- 5:14higher quality more reliable hardware
- 5:18and software and we are generally going
- 5:21to have
- 5:22some kind of a redundancy there
- 5:25in addition to that we will be adding
- 5:28some kind of a fault tolerance such that
- 5:29if a component fails
- 5:32our system does not crash
- 5:35the other approach in reducing the
- 5:38denominator
- 5:39is to reduce the mttr
- 5:42mean time to repair basically it's about
- 5:46improving the tools that have detected
- 5:48that something is not
- 5:49working quite right in the system and
- 5:52bring it back
- 5:53to the normal mode operation
- 5:57so that's about dependability measures
- 5:59the next thing that you would like to
- 6:01define
- 6:01are these availability measures or
- 6:04availability
- 6:05metrics availability
- 6:08as we have defined it already is
- 6:11the ratio of mean time to failure
- 6:14divided by the sum of mean time to
- 6:15failure
- 6:16and the mean time to repair but it is
- 6:19generally
- 6:20expressed as a percentage of time and
- 6:23both mean time
- 6:24to failure or mean time to between the
- 6:27failures
- 6:28are measured in hours now
- 6:32compute systems smaller computer systems
- 6:34rarely go down
- 6:36um you know at least this is systems
- 6:38that we are used to
- 6:39rarely go down so we
- 6:43generally assume that it is going to be
- 6:46up by
- 6:47a very large percentage of time in a
- 6:50large
- 6:50percentage of time meaning 99.99
- 6:54some percent of a time that number of
- 6:56nines
- 6:58[Music]
- 6:59is generally measured as number of nines
- 7:01availability
- 7:02per year what does that mean so if
- 7:04something is available 90
- 7:06of a time in one year that means it'll
- 7:09be
- 7:10down 40 pairs for 36 days
- 7:14if it is available 99 of a time that
- 7:17means
- 7:18that it would be down for 3.6 days
- 7:21uh you have 3.6 days of repair per year
- 7:24that's very annoying
- 7:25i mean very few people would tolerate
- 7:27something like that
- 7:31although in some other systems you know
- 7:33transportation but maybe okay
- 7:35you know that's how many days do
- 7:38buses that operate every day spend in a
- 7:40shop per year
- 7:43three nines means that the system would
- 7:45be down uh
- 7:46526 minutes uh for repairs per year
- 7:51that's annoying but getting closer to
- 7:53being acceptable
- 7:54four nines is 53 minutes of repairs per
- 7:58year
- 7:58and five nines is five minutes of repair
- 8:01per year
- 8:04that is getting closer
- 8:07to what we expect and that is something
- 8:10that we
- 8:10often encounter in these services that
- 8:13we actually use on the internet i mean
- 8:15that's like of the order four to five
- 8:18nines is the availability of
- 8:20say youtube um
- 8:23as you've noticed when watching these
- 8:26videos
- 8:26youtube sometimes goes down
- 8:30other systems are designed to have six
- 8:33or seven
- 8:34nines um you know things that we depend
- 8:37much more on
- 8:38and there is essentially a much higher
- 8:40price
- 8:42for that system going down um examples
- 8:45are
- 8:46you know the losses of youtube when
- 8:48youtube went down for a few hours
- 8:50some years ago now that is basically
- 8:53losses of
- 8:54millions per minute
- 9:00okay so that's about availability
- 9:03measures let's talk about reliability
- 9:05measures
- 9:07generally how we like to
- 9:11measure differently this reliability is
- 9:16through annualized failure rate or afr
- 9:21so how many failures do we have
- 9:25in a system that has many components per
- 9:28year for example
- 9:29let's say we operate a data center that
- 9:31has a thousand disks
- 9:32and each disk is characterized to have
- 9:35one thousand
- 9:36one hundred thousand uh hour
- 9:40in mean time to failure
- 9:44so let's see what does that mean how
- 9:45many disks actually
- 9:47are going to fail per year in this kind
- 9:50of a system
- 9:52so how do we translate one reliability
- 9:54measure
- 9:55of mttf to another reliability measure
- 9:58which is annualized failure rate this is
- 10:01something that is very important for
- 10:02data centers and data centers actually
- 10:04keep track of that
- 10:05so we have 8760 hours
- 10:09in a year and then we have a thousand
- 10:12disks that are operating
- 10:14for the whole year our mean time to
- 10:17failure is
- 10:19expressed in hours so we are going to
- 10:21divide that number with 100 000
- 10:23hours and what we end up with here is
- 10:2787.6 failed disks per year on the
- 10:29average
- 10:30you know so we have to round it up
- 10:32because it's not 0.6 discs that
- 10:34generally fails so it's 88 uh
- 10:37failed disks this translates to 8.8
- 10:43annualized failure rate
- 10:47um well that's not that crazy
- 10:51that number actually is close to
- 10:54what we see actually in practice
- 10:57google published a study in 2007
- 11:01that found that the actual afr's
- 11:04of individual tribes range
- 11:07from 1.7 in the first year and as they
- 11:11age to go to eight point six percent
- 11:14over three
- 11:16four you know three-year-old drives
- 11:20kind of interesting uh thing to see
- 11:24there are interesting websites one of
- 11:26the these
- 11:28online backup services
- 11:31backplates actually publishes the
- 11:33failure rates of their drives and that
- 11:36may be useful information that you can
- 11:37take a look at when
- 11:38you would like to buy a new drive and
- 11:41they're
- 11:42publishing these afr's
- 11:45and they turn out to be pretty good for
- 11:48modern drives they are
- 11:49either single-digit or below
- 11:51single-digit percentages
- 11:54you know translating essentially to
- 11:55about a million hours
- 11:57of the order a million hours between the
- 12:00failures
- 12:01and another piece of good news i
- 12:03mentioned this before
- 12:05that this drives actually are getting
- 12:07better they have
- 12:08less failures over time slowly
- 12:11but surely the the afr
- 12:15is decreasing there is a decreasing
- 12:17trend
- 12:18over here
- 12:23there is another um
- 12:26reliability measure that is out there um
- 12:30that is used this was used for different
- 12:32kinds of systems but
- 12:34compute systems have to adhere for that
- 12:36it is failures in time
- 12:38or fit rate
- 12:42that has been used for example for
- 12:45cars and failures in time
- 12:50or fit rate of a device is the number of
- 12:54failures that can be
- 12:55expected in one billion device hours of
- 12:58operation
- 13:00that sounds like a lot i mean one
- 13:02billion hours is something that sounds
- 13:04unimaginable but imagine that
- 13:08we have many of these devices operating
- 13:10at the same time
- 13:13so that means 1000 devices for 1 million
- 13:17hours or 1 million devices for
- 13:211 000 hours each
- 13:24would be end up failing or 1 billion
- 13:28devices which is
- 13:29approximately how many smartphones we
- 13:30have out there
- 13:32would be failing every hour
- 13:36and that sounds about right i mean if
- 13:38you are looking at the whole world i
- 13:40think the failure rate of cell phones
- 13:42out there is probably
- 13:43greater you know there is more than one
- 13:45uh self-funded fails per hour
- 13:50so mtbf is equal to
- 13:541 billion times 1 over the fit rate
- 13:58this is relevant because computers are
- 14:01getting more and more into our cars
- 14:03and all these specifications for
- 14:04reliability there are
- 14:06expressed in fit rates so you you'll see
- 14:09that
- 14:10especially if you get into that kind of
- 14:12a business
- 14:14um you're going to hear about this
- 14:16automotive safety
- 14:17integrity level acell that defines
- 14:19different fit rates
- 14:21for different classes of components in
- 14:23vehicles
- 14:25okay so design principles of
- 14:28dependability are as follows
- 14:33there should be when we are designing a
- 14:35system that
- 14:36we should depend on um there should be
- 14:38no single
- 14:39point of failure and the analogy of that
- 14:43in in
- 14:44common terms is that the chain is only
- 14:46as strong as its weakest link
- 14:48there is a pretty good
- 14:51um corollary here of all amdahl's law
- 14:55it's always going to be the longest pole
- 14:57on the tent that is going to be sticking
- 14:59out
- 15:00so no matter you know you can keep
- 15:03improving
- 15:04the dependability of one component but
- 15:08the one that is the least dependable is
- 15:10going to dominate
- 15:12so we generally have to be measuring
- 15:16that for the whole system
- 15:17and make sure
- 15:20that the least reliable component or
- 15:23least dependable component
- 15:25is dependable enough all right
- 15:28so we are going to make a quick break
- 15:31here and then we are going to take a
- 15:33look
- 15:33into the ways how do we actually detect
- 15:36these faults
- 15:37or failures see you after a break
About this transcript
This page contains the full transcript of [CS61C FA20] Lecture 38.2 - Dependability- Parity, ECC, RAID: Dependability Metrics by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 2,004 words across 386 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.