[CS61C FA20] Lecture 38.1 - Dependability- Parity, ECC, RAID: Intro — Transcript
Full transcript
- 0:03[Music]
- 0:11hello
- 0:12welcome back to 61c it's a brand new
- 0:15module
- 0:16on dependability it's a fairly short
- 0:19module
- 0:20we're going to just touch on it remember
- 0:24our early story about six great ideas
- 0:27in computer architecture so we have
- 0:30talked about
- 0:31the layers of abstraction that enable us
- 0:33to build these complex systems
- 0:35out of many many components
- 0:38and we do that by layering many
- 0:41layers of abstraction we have talked
- 0:44about moore's law that enabled us to
- 0:46integrate so many devices and build
- 0:48these complex systems
- 0:50we have talked about principle of
- 0:52locality
- 0:53and memory hierarchy that enabled us to
- 0:56build these memory systems that look
- 0:59infinitely fast and infinitely big
- 1:02we have talked about parallelism as a
- 1:04way to improve the performance
- 1:07in the in power limited regime
- 1:12we have also throughout the course
- 1:14talked about
- 1:15the this concept of of performance
- 1:18measurement and improvements
- 1:20primarily through the iron law of
- 1:23compute
- 1:24and average memory access times
- 1:27but we haven't talked about the sixth
- 1:30great idea
- 1:31which is dependability by a redundancy
- 1:35so that's what we're going to cover now
- 1:38and why is that important well we
- 1:42rely on our compute systems more and
- 1:46more
- 1:47in everyday life and in domains that are
- 1:51just
- 1:52not purely computing all
- 1:55our financial transactions are performed
- 1:58by computers
- 1:59we drive a collection of computers
- 2:03many people's lives are supported by
- 2:07some sort of a computer so
- 2:10we really need to care about that
- 2:13dependability
- 2:14but the fact is that the computers fail
- 2:18they might fail transiently and we have
- 2:21seen before i
- 2:23i'm afraid these blue screens of that
- 2:26and many other failure modes for
- 2:29for computers often after a computer
- 2:32crashes
- 2:35um it can come back and we'll call that
- 2:39a transient mode of failure this failure
- 2:42may happen
- 2:43because of a bad code well you know
- 2:45somebody you know
- 2:46it's not just students who forget to
- 2:49clean the exit
- 2:50functions production software does that
- 2:53sometimes as well they do
- 2:55you know send you know new release and
- 2:56it works better but these things happen
- 2:59but sometimes these failures transit
- 3:01failures happen because
- 3:03something did not go quite right in the
- 3:06hardware
- 3:07hardware made an error
- 3:10if these errors persist something really
- 3:14fails
- 3:15in permanently in a computer
- 3:18then often we discard them or we try to
- 3:22repair them
- 3:24if they are repairable
- 3:27so in this module we are going to spend
- 3:30some time
- 3:32talking about how do we mitigate these
- 3:35hardware failures by using
- 3:38redundant components so
- 3:42we have briefly touched on that early on
- 3:45in the in the introductory module
- 3:47when we can use
- 3:51redundancy to replace a failing
- 3:55part of a system this redundancy
- 3:58is often encountered in memory systems
- 4:00we have spare stuff in memories
- 4:02um but we
- 4:06in parallel computers we may have a
- 4:09spare processor core you'll find out
- 4:13that on the market you can buy now
- 4:17eight core gpus or seven
- 4:20core gpus well when you look at the chip
- 4:23they're exactly the same
- 4:25except that eight core was disabled
- 4:29and they'll sell you that chip a little
- 4:32bit
- 4:33cheaper why was it disabled because it
- 4:35was no good it was failing
- 4:37that's okay you just have a little bit
- 4:39less of a performance you have seven
- 4:40instead of eight
- 4:41and you're willing to pay less for that
- 4:45so the way how it works is
- 4:48that you'll have multiple replicas
- 4:52of hardware and there'll be some kind of
- 4:55a voting mechanism
- 4:56where two out of three perhaps will
- 4:59agree
- 5:00that one plus one is equal to two and
- 5:02we'll take that
- 5:04as an answer this is
- 5:07this is made a lot easier by
- 5:10the advances in integration transistor
- 5:13integration densities
- 5:14so we can have more of these components
- 5:18now there is another cache there as the
- 5:20components are smaller as the threshold
- 5:21or smaller
- 5:22they have a tendency you know to more
- 5:24frequently fail
- 5:26but that is not increasing the
- 5:30that the failure rate is not increasing
- 5:32at the speed as which at which we can
- 5:34integrate more of them
- 5:35so we can add redundancy the other way
- 5:40how we can improve things
- 5:43how we can mitigate these transient
- 5:46failures
- 5:47is through
- 5:50another form of redundancy
- 5:54we can temporarily you know things can
- 5:57temporarily go out
- 5:58and we can still
- 6:02continue computing without them for
- 6:04example you know at the very
- 6:06top level the data center may go out i
- 6:08mean
- 6:09there may be a hurricane um on
- 6:12the east coast
- 6:15the power may be out the internet may be
- 6:17out data center goes out
- 6:19the weather system goes away data center
- 6:22is back up
- 6:24the other thing that we have encountered
- 6:26is often we use these arrays of disks
- 6:29if a disk if a mechanical disk fails
- 6:33it's okay because there is a way
- 6:36to cover for that disc by having
- 6:38redundancy
- 6:39on the shelf so a shelf of disk disks
- 6:43will have spares and finally
- 6:47we can use a particular type of coding
- 6:51to add
- 6:54redundancy to our dram so our dram chip
- 6:57this is a
- 6:58dim that goes into computers if you look
- 7:00at it it has
- 7:01nine memory chips that correspond to
- 7:03nine bits instead of
- 7:05eight and the reason for that is that
- 7:08the ninth one
- 7:08is there to represent parity and that
- 7:11parity is going to be used to indicate
- 7:13if something went wrong
- 7:16so we're going to take a quick break now
- 7:19and then we're going to take a
- 7:21look at how do we measure dependability
- 7:24see you after a quick break
About this transcript
This page contains the full transcript of [CS61C FA20] Lecture 38.1 - Dependability- Parity, ECC, RAID: Intro by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 926 words across 173 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.