YouTube2Text

[CS61C FA20] Lecture 21.3 - Pipelining I: Energy Efficiency — Transcript

by CS 61C Departmental · 2,338 words · 441 segments · language en · Watch on YouTube

Full transcript

  1. 0:01[Music]
  2. 0:10hello and welcome back
  3. 0:11let's talk a bit about energy efficiency
  4. 0:14energy efficiency
  5. 0:16is something that relates processor
  6. 0:18power with the processor performance
  7. 0:21processor power dissipation has become
  8. 0:24much more
  9. 0:25important over the past couple of
  10. 0:28decades
  11. 0:28back in the 80s architects and designers
  12. 0:31had
  13. 0:32very different design objectives
  14. 0:36back then the only goal was to increase
  15. 0:39the performance
  16. 0:40as fast as possible
  17. 0:44so they did everything they had they
  18. 0:47used all the tools that they had under
  19. 0:49their belt
  20. 0:50to do that to increase the performance
  21. 0:54so they shorten the logic that they use
  22. 0:57the maximum
  23. 0:58possible supply voltage to crank up the
  24. 1:01frequency of processors
  25. 1:02and as a result processor performance
  26. 1:05was doubling every other year
  27. 1:07but the price in power was high the
  28. 1:10power was doubling every approximately
  29. 1:14three years
  30. 1:15that was not that big of an issue in
  31. 1:17early 80s when the processors were
  32. 1:19relatively simple and
  33. 1:24the power consumption was relatively low
  34. 1:28but over time the power consumption hit
  35. 1:30the limits
  36. 1:31and essentially in across all the power
  37. 1:34domains
  38. 1:35the power became a primary limiting
  39. 1:38factor
  40. 1:39sometime in the early 2000s
  41. 1:42and nowadays power is limited in every
  42. 1:45compute domain
  43. 1:46power is limited for
  44. 1:50cell phones for desktop processors and
  45. 1:54for warehouse scale computers
  46. 1:57in data centers the nature of these
  47. 1:59power limitations is different
  48. 2:02for example in cell phones the power is
  49. 2:05limited
  50. 2:05by the size of the battery that we are
  51. 2:08willing to carry
  52. 2:09on ourselves in desktops the power is
  53. 2:12limited by
  54. 2:14the amount of heat that a relatively
  55. 2:16inexpensive
  56. 2:18fan can remove from a processor and in
  57. 2:21data centers it's
  58. 2:22both the heat and also the cost of
  59. 2:24energy that it takes us
  60. 2:26to run that data center
  61. 2:30so the goal of designs nowadays and in
  62. 2:34foreseeable future
  63. 2:36is to maximize the performance
  64. 2:40under power constraints so it's not
  65. 2:42anymore to just get the maximum
  66. 2:44performance
  67. 2:45but to deliver available perform
  68. 2:47performance
  69. 2:48on their power and energy constraints
  70. 2:51in order to get a bit better insight
  71. 2:54into energy efficiency
  72. 2:55let's first understand where does the
  73. 2:57energy go in cmos
  74. 2:59all our logic gates are built in cmos
  75. 3:01technologies nowadays
  76. 3:04so let's examine the very simple
  77. 3:07inverter here here is the logic
  78. 3:11symbol of an inverter that we've seen
  79. 3:13many times
  80. 3:14there are three additions to it it got
  81. 3:17this
  82. 3:18power supply and the ground so supply
  83. 3:20voltage in the ground
  84. 3:22because it needs a supply and a ground
  85. 3:25in order to switch the logic levels
  86. 3:28from zero to one and then we see
  87. 3:31capacitance
  88. 3:32at its output so
  89. 3:35there is a symbol of an inverter on the
  90. 3:37left-hand side and a schematic of an
  91. 3:39inverter on the right-hand side
  92. 3:42they represent exactly the same circuit
  93. 3:44just two
  94. 3:45different level of abstraction the
  95. 3:47capacitances
  96. 3:49happen because the nature of transistors
  97. 3:51every transistor that we use
  98. 3:53in an inverter has its input
  99. 3:56and output capacitances so whenever we
  100. 3:59are trying
  101. 4:00to switch the output level
  102. 4:03of an inverter from zero to one we have
  103. 4:05to charge
  104. 4:07a capacitor that capacitor consists of
  105. 4:09capacitances that are at the output of
  106. 4:11this inverter
  107. 4:12but also at the inputs
  108. 4:16of any of the inverters that are
  109. 4:20downstream that are being driven by this
  110. 4:23inverter
  111. 4:24how much energy does it take us to
  112. 4:27charge a capacitor
  113. 4:28from zero to one from zero to v well the
  114. 4:31amount of energy is equal to cv
  115. 4:33squared if you remember uh
  116. 4:36from high school physics if you haven't
  117. 4:38taken any e classes yet
  118. 4:40a half of that energy gets stored on the
  119. 4:41capacitor
  120. 4:43and the other half gets dissipated
  121. 4:46during the
  122. 4:48charging period so you in the end we end
  123. 4:51up with the
  124. 4:51half of a cv squared on this capacitor
  125. 4:56and a half of cv squared was dissipated
  126. 4:59on this transistor
  127. 5:01eventually when we switch back from one
  128. 5:03to zero
  129. 5:04we dump that half cv squared into ground
  130. 5:10that is the dominant part of
  131. 5:14of power dissipation these days are
  132. 5:16energy dissipation in
  133. 5:18in cmos technologies the other one is
  134. 5:22the leakage leakage stems from the fact
  135. 5:25that
  136. 5:26transistors are not ideal switches
  137. 5:29anymore
  138. 5:29they are much more in real ideal back in
  139. 5:32the
  140. 5:3280s and 90s um you know when we would
  141. 5:36turn off the transistor we would set
  142. 5:38this
  143. 5:38its voltage input voltage to zero uh it
  144. 5:41would really
  145. 5:42turn up there would be very little
  146. 5:44current that would be trickling through
  147. 5:46a transistor
  148. 5:47nowadays they are more like dimmers if
  149. 5:49we
  150. 5:50prefer an equivalence to a light switch
  151. 5:52so we're just able to dim the current
  152. 5:54that is going through the transistor so
  153. 5:57when transistor is supposed to be off it
  154. 5:58is still leaking some amount of current
  155. 6:01and one thing that you may find somewhat
  156. 6:04paradoxical
  157. 6:05that the balance between the switching
  158. 6:08energy that goes into charging
  159. 6:09capacitance
  160. 6:10and leakage energy is somewhat
  161. 6:13surprising
  162. 6:15we spend about 70 percent on charging
  163. 6:18and discharging capacitances
  164. 6:20which is our useful use of the energy
  165. 6:22and 30 percent
  166. 6:23may go into leakage that may be a bit of
  167. 6:26a surprise
  168. 6:28it comes from the basic
  169. 6:31things that are associated with
  170. 6:33technology with the cmos technology and
  171. 6:35we're not going to dive into that that
  172. 6:36discovered in
  173. 6:37other classes like um 151 or
  174. 6:41130.
  175. 6:44but let's pop back up at a bit of a
  176. 6:48higher level
  177. 6:53how much energy does it take us to
  178. 6:56complete the task and let's say that our
  179. 6:58task is to execute a program
  180. 6:59for example to compress a video
  181. 7:04so that's again a complex thing to
  182. 7:07evaluate and there are many parameters
  183. 7:08so it is
  184. 7:09a good idea to break it down and in this
  185. 7:11case it's
  186. 7:12a good idea to break it down into two
  187. 7:14components so energy per program
  188. 7:17which may be executing our tasks can be
  189. 7:19broken down into
  190. 7:21the number of instructions that we have
  191. 7:22in the program and
  192. 7:24the energy per instruction
  193. 7:28so um the number of instructions per
  194. 7:31program we have already encountered
  195. 7:33when we're trying to measure the
  196. 7:35performance
  197. 7:36and in this case you know the shorter
  198. 7:37program the less energy it is going to
  199. 7:39to use um i
  200. 7:42don't need to say much more about that
  201. 7:44that is usually outside of designers
  202. 7:46um range of uh
  203. 7:51of things that they can influence um
  204. 7:54but it does affect the architectures and
  205. 7:57and
  206. 7:58general computer design
  207. 8:02on the other hand what we can much more
  208. 8:04influence
  209. 8:05is the second part which is the energy
  210. 8:07per instruction and the dominant part of
  211. 8:09that is cv square there is some leakage
  212. 8:11that goes
  213. 8:12on but i'm going to ignore the leakage
  214. 8:14at the moment
  215. 8:15so there is some energy that is
  216. 8:19associated
  217. 8:19with computing that is proportional to
  218. 8:22the capacitance
  219. 8:23times the voltage squared and
  220. 8:26capacitance
  221. 8:27like we've had in a simple case of an
  222. 8:29inverter is now
  223. 8:31a total equivalent capacitance of all
  224. 8:34the gates that are switching
  225. 8:36say in a processor core that capacitance
  226. 8:41is um is the capacitance that we
  227. 8:44need to switch in every instruction to
  228. 8:47move
  229. 8:47a certain amount of wires from zero to
  230. 8:51one
  231. 8:52it depends on the technology that we are
  232. 8:53using and also on
  233. 8:55the architectural processor and its
  234. 8:57features
  235. 8:58on the other hand the supply voltage is
  236. 9:03a very strong parameter here in
  237. 9:06trading off power for performance or
  238. 9:09energy
  239. 9:10for performance supply voltage nominally
  240. 9:13is around 1 volt
  241. 9:15more like 0.8 volts or 0.7 volts in the
  242. 9:18recent single digit nanometer
  243. 9:21technologies
  244. 9:23but the energy depends quadratically
  245. 9:26on the on the supply voltage
  246. 9:30so if we lower the supply voltage by 10
  247. 9:33we save 20 percent of the energy 21 to
  248. 9:37be
  249. 9:38more
  250. 9:4219 to be more precise
  251. 9:46all right so the basic goal here of a
  252. 9:51designer is to try to reduce
  253. 9:53the capacitance and the voltage to
  254. 9:56reduce the energy per
  255. 9:58task and voltage is a more powerful knob
  256. 10:02one thing to keep in mind lowering the
  257. 10:06voltage also affects the performance
  258. 10:09because transistors are slower
  259. 10:12when they run at a lower supply voltage
  260. 10:19so trading off changing
  261. 10:22supply voltage three dots trades off
  262. 10:25energy savings for loss in performance
  263. 10:29and we are going to see in different
  264. 10:31compute domains
  265. 10:33we are going to be seeing processors
  266. 10:35operating at a bit of a higher voltage
  267. 10:37versus a little bit less of a voltage
  268. 10:40to achieve different power performance
  269. 10:43trade-off point
  270. 10:45okay talking a little bit more about
  271. 10:47energy trade of examples
  272. 10:49one thing that happens in every when we
  273. 10:52scale the technology no the imprint
  274. 10:54implement a new processor we get some
  275. 10:57natural energy savings
  276. 10:59so back in the 80s and 90s every time
  277. 11:03and
  278. 11:03up to probably about 2010 every time
  279. 11:07we scale the technology we go into the
  280. 11:09next node in moore's law
  281. 11:11the capacitance would shrink by about 30
  282. 11:14percent
  283. 11:15we would save 30 percent of our
  284. 11:17capacitance
  285. 11:19now those capacity and savings are
  286. 11:23more like 15
  287. 11:26from node to node they're still you know
  288. 11:28don't throw away 15
  289. 11:3015 discount is huge
  290. 11:34supply has recently been also scaling by
  291. 11:36about 15
  292. 11:37and that has been traditionally going on
  293. 11:39since like
  294. 11:40mid-90s or so early on
  295. 11:44the supply reduction again supply is
  296. 11:47supposed to be reduced by 30 percent to
  297. 11:49match
  298. 11:49what we get from the capacity and
  299. 11:53scaling
  300. 11:54and that is something that is associated
  301. 11:55with bernard's law
  302. 11:57but it has been more like 15 because
  303. 12:01designers were greedy and they wanted to
  304. 12:03maximize the performance and they were
  305. 12:06just trying to keep that supply
  306. 12:10as high as possible now supply
  307. 12:12reductions are
  308. 12:13limited by this leakage
  309. 12:16because it affects if we really want to
  310. 12:18drop the supply voltage the leakage is
  311. 12:20going to go up
  312. 12:22so as a result when we multiply c and v
  313. 12:25squared now from a node to node
  314. 12:26we get almost uh 40 discount in energy
  315. 12:31that allows us to do things with that
  316. 12:34i mean when we build a new product if
  317. 12:36you keep exactly the same architecture
  318. 12:38and same number of cords it naturally
  319. 12:40becomes
  320. 12:4140 percent um
  321. 12:44more energy efficient or has 40 percent
  322. 12:47less energy
  323. 12:48and or we can pack 40 percent more stuff
  324. 12:52in there if we have the fixed energy
  325. 12:54budget so
  326. 12:55that's kind of convenient that's exactly
  327. 12:56what has been happening
  328. 12:58with scaling lately so
  329. 13:01in a brief summary the energy efficiency
  330. 13:05has been improving because of moore's
  331. 13:07law
  332. 13:07and reduced supply voltage and
  333. 13:09architects have been using that
  334. 13:10to pack more features and improve
  335. 13:13performance
  336. 13:14from that a little bit about
  337. 13:17power and performance trends over time
  338. 13:21this is a plot that we have shown in the
  339. 13:23first lecture it is showing
  340. 13:25years of introduction of a whole lot
  341. 13:28process of processors
  342. 13:30uh over the past 48 years since intel's
  343. 13:324 4040
  344. 13:35in the early 70s and then
  345. 13:39we see the number of transistors which
  346. 13:42is moore's law
  347. 13:43uh steadily increasing with a little bit
  348. 13:45of
  349. 13:47uh slowdown up here then we are seeing
  350. 13:51frequency and single thread performance
  351. 13:55power and then number of cores
  352. 13:59so what we have seen is we frequency
  353. 14:02was the primary means for gaining
  354. 14:06the performance in the 80s and 90s
  355. 14:10and architects and designers worked
  356. 14:12together to crank it up
  357. 14:14but it hit the limit in early 2000s and
  358. 14:16has been flat since
  359. 14:18the reason for that is because the power
  360. 14:20had to to stay flat
  361. 14:22um so architects work hard they had to
  362. 14:25continue working on trying to improve
  363. 14:27the single thread performance which is
  364. 14:29what we have seen here
  365. 14:30but also we're utilizing parallelism by
  366. 14:33adding more cores
  367. 14:35to add the performance to
  368. 14:38add to end applications
  369. 14:42how things are going to change over time
  370. 14:45with the end of moore's law
  371. 14:48it's not easy to predict it exactly but
  372. 14:51we got some sense
  373. 14:53um so first these capacitance savings
  374. 14:55are not going to come that
  375. 14:57easy voltage savings are
  376. 15:00not going to happen easily either
  377. 15:05[Music]
  378. 15:06simply as i mentioned before we can't
  379. 15:08reduce the voltage otherwise things are
  380. 15:10going to leak
  381. 15:12in order to so
  382. 15:15capacitance savings are also not going
  383. 15:17to happen for free
  384. 15:18because we are slowing down this
  385. 15:21shrinkage of transistors themselves
  386. 15:23lots of savings in the capacitance are
  387. 15:26likely coming
  388. 15:27from packing things in the third
  389. 15:28dimension so when things
  390. 15:31are spread apart on the board in two
  391. 15:33dimensions we get a lot of
  392. 15:35system capacitance basically over there
  393. 15:37but if you can pack things on top of
  394. 15:38each other
  395. 15:40we're going to save quite a bit of of
  396. 15:43energy there because we'll have less
  397. 15:44capacitance
  398. 15:48and that is going to try
  399. 15:51to help us work around this power wall
  400. 15:58important takeaway here is a
  401. 16:01basic relationship between the
  402. 16:04performance and power
  403. 16:07and it is shown here as the energy iron
  404. 16:10law
  405. 16:11energy efficiency which is managed which
  406. 16:13is expressed as
  407. 16:15the number of tasks that we can
  408. 16:19perform per joule is what relates
  409. 16:22the performance in the power so
  410. 16:24performance is the number of tasks that
  411. 16:26we can execute per second
  412. 16:28power is joules per second energy
  413. 16:30efficiency
  414. 16:31is task per joule
  415. 16:34so energy efficiency now becomes
  416. 16:37really the key metric of our designs
  417. 16:41um and it is something that tells us
  418. 16:44actually how much performance we can get
  419. 16:48out of a power limited design if the
  420. 16:50power limited design is a cell phone
  421. 16:52the power on the average is 600
  422. 16:55milliwatts i mean that's what we can
  423. 16:56train out of the of the battery on the
  424. 17:00average
  425. 17:01what is specked out the max power of
  426. 17:04cell phones
  427. 17:05is limited to a few watts but for a very
  428. 17:08short burst of time
  429. 17:10on the other hand if you're working with
  430. 17:12a google
  431. 17:13scale data center that may you have a
  432. 17:17power
  433. 17:18envelope of 20 megawatts that tells us
  434. 17:20that the energy efficiency of that
  435. 17:22entire data center
  436. 17:23tells us what kind of a performance we
  437. 17:26can get out of a data center
  438. 17:29and that's basically it
  439. 17:32we're going to look into pipelining
  440. 17:35as a way to improve performance
  441. 17:40after a break see you then

About this transcript

This page contains the full transcript of [CS61C FA20] Lecture 21.3 - Pipelining I: Energy Efficiency by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 2,338 words across 441 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.