[CS61C FA20] Lecture 21.3 - Pipelining I: Energy Efficiency — Transcript
Full transcript
- 0:01[Music]
- 0:10hello and welcome back
- 0:11let's talk a bit about energy efficiency
- 0:14energy efficiency
- 0:16is something that relates processor
- 0:18power with the processor performance
- 0:21processor power dissipation has become
- 0:24much more
- 0:25important over the past couple of
- 0:28decades
- 0:28back in the 80s architects and designers
- 0:31had
- 0:32very different design objectives
- 0:36back then the only goal was to increase
- 0:39the performance
- 0:40as fast as possible
- 0:44so they did everything they had they
- 0:47used all the tools that they had under
- 0:49their belt
- 0:50to do that to increase the performance
- 0:54so they shorten the logic that they use
- 0:57the maximum
- 0:58possible supply voltage to crank up the
- 1:01frequency of processors
- 1:02and as a result processor performance
- 1:05was doubling every other year
- 1:07but the price in power was high the
- 1:10power was doubling every approximately
- 1:14three years
- 1:15that was not that big of an issue in
- 1:17early 80s when the processors were
- 1:19relatively simple and
- 1:24the power consumption was relatively low
- 1:28but over time the power consumption hit
- 1:30the limits
- 1:31and essentially in across all the power
- 1:34domains
- 1:35the power became a primary limiting
- 1:38factor
- 1:39sometime in the early 2000s
- 1:42and nowadays power is limited in every
- 1:45compute domain
- 1:46power is limited for
- 1:50cell phones for desktop processors and
- 1:54for warehouse scale computers
- 1:57in data centers the nature of these
- 1:59power limitations is different
- 2:02for example in cell phones the power is
- 2:05limited
- 2:05by the size of the battery that we are
- 2:08willing to carry
- 2:09on ourselves in desktops the power is
- 2:12limited by
- 2:14the amount of heat that a relatively
- 2:16inexpensive
- 2:18fan can remove from a processor and in
- 2:21data centers it's
- 2:22both the heat and also the cost of
- 2:24energy that it takes us
- 2:26to run that data center
- 2:30so the goal of designs nowadays and in
- 2:34foreseeable future
- 2:36is to maximize the performance
- 2:40under power constraints so it's not
- 2:42anymore to just get the maximum
- 2:44performance
- 2:45but to deliver available perform
- 2:47performance
- 2:48on their power and energy constraints
- 2:51in order to get a bit better insight
- 2:54into energy efficiency
- 2:55let's first understand where does the
- 2:57energy go in cmos
- 2:59all our logic gates are built in cmos
- 3:01technologies nowadays
- 3:04so let's examine the very simple
- 3:07inverter here here is the logic
- 3:11symbol of an inverter that we've seen
- 3:13many times
- 3:14there are three additions to it it got
- 3:17this
- 3:18power supply and the ground so supply
- 3:20voltage in the ground
- 3:22because it needs a supply and a ground
- 3:25in order to switch the logic levels
- 3:28from zero to one and then we see
- 3:31capacitance
- 3:32at its output so
- 3:35there is a symbol of an inverter on the
- 3:37left-hand side and a schematic of an
- 3:39inverter on the right-hand side
- 3:42they represent exactly the same circuit
- 3:44just two
- 3:45different level of abstraction the
- 3:47capacitances
- 3:49happen because the nature of transistors
- 3:51every transistor that we use
- 3:53in an inverter has its input
- 3:56and output capacitances so whenever we
- 3:59are trying
- 4:00to switch the output level
- 4:03of an inverter from zero to one we have
- 4:05to charge
- 4:07a capacitor that capacitor consists of
- 4:09capacitances that are at the output of
- 4:11this inverter
- 4:12but also at the inputs
- 4:16of any of the inverters that are
- 4:20downstream that are being driven by this
- 4:23inverter
- 4:24how much energy does it take us to
- 4:27charge a capacitor
- 4:28from zero to one from zero to v well the
- 4:31amount of energy is equal to cv
- 4:33squared if you remember uh
- 4:36from high school physics if you haven't
- 4:38taken any e classes yet
- 4:40a half of that energy gets stored on the
- 4:41capacitor
- 4:43and the other half gets dissipated
- 4:46during the
- 4:48charging period so you in the end we end
- 4:51up with the
- 4:51half of a cv squared on this capacitor
- 4:56and a half of cv squared was dissipated
- 4:59on this transistor
- 5:01eventually when we switch back from one
- 5:03to zero
- 5:04we dump that half cv squared into ground
- 5:10that is the dominant part of
- 5:14of power dissipation these days are
- 5:16energy dissipation in
- 5:18in cmos technologies the other one is
- 5:22the leakage leakage stems from the fact
- 5:25that
- 5:26transistors are not ideal switches
- 5:29anymore
- 5:29they are much more in real ideal back in
- 5:32the
- 5:3280s and 90s um you know when we would
- 5:36turn off the transistor we would set
- 5:38this
- 5:38its voltage input voltage to zero uh it
- 5:41would really
- 5:42turn up there would be very little
- 5:44current that would be trickling through
- 5:46a transistor
- 5:47nowadays they are more like dimmers if
- 5:49we
- 5:50prefer an equivalence to a light switch
- 5:52so we're just able to dim the current
- 5:54that is going through the transistor so
- 5:57when transistor is supposed to be off it
- 5:58is still leaking some amount of current
- 6:01and one thing that you may find somewhat
- 6:04paradoxical
- 6:05that the balance between the switching
- 6:08energy that goes into charging
- 6:09capacitance
- 6:10and leakage energy is somewhat
- 6:13surprising
- 6:15we spend about 70 percent on charging
- 6:18and discharging capacitances
- 6:20which is our useful use of the energy
- 6:22and 30 percent
- 6:23may go into leakage that may be a bit of
- 6:26a surprise
- 6:28it comes from the basic
- 6:31things that are associated with
- 6:33technology with the cmos technology and
- 6:35we're not going to dive into that that
- 6:36discovered in
- 6:37other classes like um 151 or
- 6:41130.
- 6:44but let's pop back up at a bit of a
- 6:48higher level
- 6:53how much energy does it take us to
- 6:56complete the task and let's say that our
- 6:58task is to execute a program
- 6:59for example to compress a video
- 7:04so that's again a complex thing to
- 7:07evaluate and there are many parameters
- 7:08so it is
- 7:09a good idea to break it down and in this
- 7:11case it's
- 7:12a good idea to break it down into two
- 7:14components so energy per program
- 7:17which may be executing our tasks can be
- 7:19broken down into
- 7:21the number of instructions that we have
- 7:22in the program and
- 7:24the energy per instruction
- 7:28so um the number of instructions per
- 7:31program we have already encountered
- 7:33when we're trying to measure the
- 7:35performance
- 7:36and in this case you know the shorter
- 7:37program the less energy it is going to
- 7:39to use um i
- 7:42don't need to say much more about that
- 7:44that is usually outside of designers
- 7:46um range of uh
- 7:51of things that they can influence um
- 7:54but it does affect the architectures and
- 7:57and
- 7:58general computer design
- 8:02on the other hand what we can much more
- 8:04influence
- 8:05is the second part which is the energy
- 8:07per instruction and the dominant part of
- 8:09that is cv square there is some leakage
- 8:11that goes
- 8:12on but i'm going to ignore the leakage
- 8:14at the moment
- 8:15so there is some energy that is
- 8:19associated
- 8:19with computing that is proportional to
- 8:22the capacitance
- 8:23times the voltage squared and
- 8:26capacitance
- 8:27like we've had in a simple case of an
- 8:29inverter is now
- 8:31a total equivalent capacitance of all
- 8:34the gates that are switching
- 8:36say in a processor core that capacitance
- 8:41is um is the capacitance that we
- 8:44need to switch in every instruction to
- 8:47move
- 8:47a certain amount of wires from zero to
- 8:51one
- 8:52it depends on the technology that we are
- 8:53using and also on
- 8:55the architectural processor and its
- 8:57features
- 8:58on the other hand the supply voltage is
- 9:03a very strong parameter here in
- 9:06trading off power for performance or
- 9:09energy
- 9:10for performance supply voltage nominally
- 9:13is around 1 volt
- 9:15more like 0.8 volts or 0.7 volts in the
- 9:18recent single digit nanometer
- 9:21technologies
- 9:23but the energy depends quadratically
- 9:26on the on the supply voltage
- 9:30so if we lower the supply voltage by 10
- 9:33we save 20 percent of the energy 21 to
- 9:37be
- 9:38more
- 9:4219 to be more precise
- 9:46all right so the basic goal here of a
- 9:51designer is to try to reduce
- 9:53the capacitance and the voltage to
- 9:56reduce the energy per
- 9:58task and voltage is a more powerful knob
- 10:02one thing to keep in mind lowering the
- 10:06voltage also affects the performance
- 10:09because transistors are slower
- 10:12when they run at a lower supply voltage
- 10:19so trading off changing
- 10:22supply voltage three dots trades off
- 10:25energy savings for loss in performance
- 10:29and we are going to see in different
- 10:31compute domains
- 10:33we are going to be seeing processors
- 10:35operating at a bit of a higher voltage
- 10:37versus a little bit less of a voltage
- 10:40to achieve different power performance
- 10:43trade-off point
- 10:45okay talking a little bit more about
- 10:47energy trade of examples
- 10:49one thing that happens in every when we
- 10:52scale the technology no the imprint
- 10:54implement a new processor we get some
- 10:57natural energy savings
- 10:59so back in the 80s and 90s every time
- 11:03and
- 11:03up to probably about 2010 every time
- 11:07we scale the technology we go into the
- 11:09next node in moore's law
- 11:11the capacitance would shrink by about 30
- 11:14percent
- 11:15we would save 30 percent of our
- 11:17capacitance
- 11:19now those capacity and savings are
- 11:23more like 15
- 11:26from node to node they're still you know
- 11:28don't throw away 15
- 11:3015 discount is huge
- 11:34supply has recently been also scaling by
- 11:36about 15
- 11:37and that has been traditionally going on
- 11:39since like
- 11:40mid-90s or so early on
- 11:44the supply reduction again supply is
- 11:47supposed to be reduced by 30 percent to
- 11:49match
- 11:49what we get from the capacity and
- 11:53scaling
- 11:54and that is something that is associated
- 11:55with bernard's law
- 11:57but it has been more like 15 because
- 12:01designers were greedy and they wanted to
- 12:03maximize the performance and they were
- 12:06just trying to keep that supply
- 12:10as high as possible now supply
- 12:12reductions are
- 12:13limited by this leakage
- 12:16because it affects if we really want to
- 12:18drop the supply voltage the leakage is
- 12:20going to go up
- 12:22so as a result when we multiply c and v
- 12:25squared now from a node to node
- 12:26we get almost uh 40 discount in energy
- 12:31that allows us to do things with that
- 12:34i mean when we build a new product if
- 12:36you keep exactly the same architecture
- 12:38and same number of cords it naturally
- 12:40becomes
- 12:4140 percent um
- 12:44more energy efficient or has 40 percent
- 12:47less energy
- 12:48and or we can pack 40 percent more stuff
- 12:52in there if we have the fixed energy
- 12:54budget so
- 12:55that's kind of convenient that's exactly
- 12:56what has been happening
- 12:58with scaling lately so
- 13:01in a brief summary the energy efficiency
- 13:05has been improving because of moore's
- 13:07law
- 13:07and reduced supply voltage and
- 13:09architects have been using that
- 13:10to pack more features and improve
- 13:13performance
- 13:14from that a little bit about
- 13:17power and performance trends over time
- 13:21this is a plot that we have shown in the
- 13:23first lecture it is showing
- 13:25years of introduction of a whole lot
- 13:28process of processors
- 13:30uh over the past 48 years since intel's
- 13:324 4040
- 13:35in the early 70s and then
- 13:39we see the number of transistors which
- 13:42is moore's law
- 13:43uh steadily increasing with a little bit
- 13:45of
- 13:47uh slowdown up here then we are seeing
- 13:51frequency and single thread performance
- 13:55power and then number of cores
- 13:59so what we have seen is we frequency
- 14:02was the primary means for gaining
- 14:06the performance in the 80s and 90s
- 14:10and architects and designers worked
- 14:12together to crank it up
- 14:14but it hit the limit in early 2000s and
- 14:16has been flat since
- 14:18the reason for that is because the power
- 14:20had to to stay flat
- 14:22um so architects work hard they had to
- 14:25continue working on trying to improve
- 14:27the single thread performance which is
- 14:29what we have seen here
- 14:30but also we're utilizing parallelism by
- 14:33adding more cores
- 14:35to add the performance to
- 14:38add to end applications
- 14:42how things are going to change over time
- 14:45with the end of moore's law
- 14:48it's not easy to predict it exactly but
- 14:51we got some sense
- 14:53um so first these capacitance savings
- 14:55are not going to come that
- 14:57easy voltage savings are
- 15:00not going to happen easily either
- 15:05[Music]
- 15:06simply as i mentioned before we can't
- 15:08reduce the voltage otherwise things are
- 15:10going to leak
- 15:12in order to so
- 15:15capacitance savings are also not going
- 15:17to happen for free
- 15:18because we are slowing down this
- 15:21shrinkage of transistors themselves
- 15:23lots of savings in the capacitance are
- 15:26likely coming
- 15:27from packing things in the third
- 15:28dimension so when things
- 15:31are spread apart on the board in two
- 15:33dimensions we get a lot of
- 15:35system capacitance basically over there
- 15:37but if you can pack things on top of
- 15:38each other
- 15:40we're going to save quite a bit of of
- 15:43energy there because we'll have less
- 15:44capacitance
- 15:48and that is going to try
- 15:51to help us work around this power wall
- 15:58important takeaway here is a
- 16:01basic relationship between the
- 16:04performance and power
- 16:07and it is shown here as the energy iron
- 16:10law
- 16:11energy efficiency which is managed which
- 16:13is expressed as
- 16:15the number of tasks that we can
- 16:19perform per joule is what relates
- 16:22the performance in the power so
- 16:24performance is the number of tasks that
- 16:26we can execute per second
- 16:28power is joules per second energy
- 16:30efficiency
- 16:31is task per joule
- 16:34so energy efficiency now becomes
- 16:37really the key metric of our designs
- 16:41um and it is something that tells us
- 16:44actually how much performance we can get
- 16:48out of a power limited design if the
- 16:50power limited design is a cell phone
- 16:52the power on the average is 600
- 16:55milliwatts i mean that's what we can
- 16:56train out of the of the battery on the
- 17:00average
- 17:01what is specked out the max power of
- 17:04cell phones
- 17:05is limited to a few watts but for a very
- 17:08short burst of time
- 17:10on the other hand if you're working with
- 17:12a google
- 17:13scale data center that may you have a
- 17:17power
- 17:18envelope of 20 megawatts that tells us
- 17:20that the energy efficiency of that
- 17:22entire data center
- 17:23tells us what kind of a performance we
- 17:26can get out of a data center
- 17:29and that's basically it
- 17:32we're going to look into pipelining
- 17:35as a way to improve performance
- 17:40after a break see you then
About this transcript
This page contains the full transcript of [CS61C FA20] Lecture 21.3 - Pipelining I: Energy Efficiency by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 2,338 words across 441 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.