YouTube2Text

EfficientML.ai Lecture 1 - Introduction (MIT 6.5940, Fall 2024, Zoom Recording) — Transcript

by MIT HAN Lab · 9,527 words · 1,416 segments · language en · Watch on YouTube

Full transcript

  1. 0:00welcome to Tiny ML and efficient deeper
  2. 0:02Computing 2024 I'm Professor Sunan very
  3. 0:06glad to teach this class again uh this
  4. 0:09is the third time we offer uh this class
  5. 0:12we're going to learn efficient deep
  6. 0:14learning Computing how to make new
  7. 0:16networks run faster train with faster
  8. 0:18and influence faster I remember two
  9. 0:21years ago when we first opened this
  10. 0:23course only 30 students last year it was
  11. 0:27uh 90 this year is more than 200 so
  12. 0:30welcome all I'm very glad to have you
  13. 0:32all here and hopefully we'll have a very
  14. 0:35fruitful semester all together so this
  15. 0:39is a brief introduction of myself I
  16. 0:42graduated from chinga University and get
  17. 0:44my PhD from Stanford with Professor uh
  18. 0:48Bill Daddy I worked on model compression
  19. 0:50and efficient deeper Computing during my
  20. 0:53PhD uh published a deep compression
  21. 0:55paper which set the foundation for
  22. 0:57modern neuron Network acceleration
  23. 1:00compression and published paper called
  24. 1:02efficient inference engine which is
  25. 1:04actually top five CED paper in 50 years
  26. 1:06of Isa uh during my PhD I co-founded a
  27. 1:09startup called defi tag for um efficient
  28. 1:14air chip which was later acquired by
  29. 1:16zings and later zings is acquired by AMD
  30. 1:19that's very popular story in City cona
  31. 1:23so in this course I will teach some of
  32. 1:25the um how to make sure the research is
  33. 1:27very fundament not only fundamental but
  34. 1:29also practical so I joined MIT about six
  35. 1:33years ago uh very exciting Journey uh
  36. 1:36we've been working on from Tiny ml to
  37. 1:38large language model acceleration and
  38. 1:41compression one of the most exciting
  39. 1:43paper we did in the past few years is
  40. 1:46the awq for large language model
  41. 1:49quantization which got the best paper
  42. 1:51award at mlis and also it's having
  43. 1:54already more than six million downloads
  44. 1:56in and hugging phase so if you want to
  45. 1:58deploy a large language model make it
  46. 2:00four times smaller then awq is the
  47. 2:03technique to go and we will learn that
  48. 2:05in this lecture you're going to
  49. 2:07implement that not only Implement that
  50. 2:09but also deploy a large language model
  51. 2:12locally on your
  52. 2:13laptop which is will be pretty exciting
  53. 2:17so at MIT I also Co found a startup
  54. 2:19called oml working on efficient
  55. 2:23different Computing model
  56. 2:25optimizations so in this lecture we'll
  57. 2:27learn all the secret sources from
  58. 2:30those two startups so those five
  59. 2:32homework will be very very helpful if
  60. 2:36you learn the master the m so we very
  61. 2:39carefully
  62. 2:40created Five assignments five five
  63. 2:43homeworks they're all hands on homeworks
  64. 2:46so we try to uh for you guys use the
  65. 2:49minimum effort to learn the most amount
  66. 2:51of knowledge the best utilize your time
  67. 2:53so please uh do those homework um very
  68. 2:58well
  69. 3:00yeah we'll have five labs and each TI
  70. 3:01will be responsible for one of them and
  71. 3:04we have office hours every Tuesday and
  72. 3:06Thursday and feel free to come to our
  73. 3:08office and we can help you with any
  74. 3:11questions you may
  75. 3:12have okay let's start by motivating why
  76. 3:16do we need efficient de Computing so
  77. 3:19let's see the supply and demand for a
  78. 3:22Computing so this chart shows according
  79. 3:26throughout different years the model
  80. 3:28size of large language model in red
  81. 3:31versus the GPU memory in green we can
  82. 3:35see the large language model size grow
  83. 3:38from 0.05 billion five billion
  84. 3:42parameters the initial Transformer in
  85. 3:442018 2017 2018 to about um 1.5 billion
  86. 3:50parameter in tbd2 roughly in 2019
  87. 3:532020 and recently uh there's even TR
  88. 3:58parameter models being trained in the
  89. 4:00world right it's growing super fast even
  90. 4:03much faster than the amount of GPU
  91. 4:05memories that is available in a single
  92. 4:08GPU right which is showing in uh showing
  93. 4:12uh green so this big gap the Gap is
  94. 4:15getting larger
  95. 4:17actually with recent years and if you
  96. 4:20see the supply and demand curve that is
  97. 4:22driving up the price right making deep
  98. 4:24learning super expensive to train super
  99. 4:27expensive to serve right so we are going
  100. 4:29going to learn techniques to bridge the
  101. 4:32gap using efficient model compression
  102. 4:34techniques so that we can improve the
  103. 4:37utilization of the hardware and also
  104. 4:39reduce the models size memory
  105. 4:42requirement computation requirement to
  106. 4:45bridge this
  107. 4:46Gap so rather than a TR a model and
  108. 4:49directly R inference on the top we now
  109. 4:52can TR a model compress it before
  110. 4:55deploying the compress model for un
  111. 5:00so in recent years the amount of
  112. 5:02Publications in model compression and
  113. 5:04acceleration on compressed sparse model
  114. 5:07is uh getting a lot faster number of
  115. 5:10Publications is increasing very fast in
  116. 5:13recent years and we are going to learn
  117. 5:14these techniques how do we utilize
  118. 5:16compression and sparity to accelerate a
  119. 5:19neuron Network
  120. 5:23inference so um in this lecture we are
  121. 5:26going to uh start with three three
  122. 5:29sections from Vision to language to
  123. 5:32multiple
  124. 5:33modalities um to understand the
  125. 5:36computation
  126. 5:38demand uh of this AI workload and see
  127. 5:41where we
  128. 5:42are among the fronti here so let's start
  129. 5:45with vision tasks so talking about comp
  130. 5:48Vision we have to mention image net and
  131. 5:51also alexnet right so Alex net this is
  132. 5:54actually the uh image net um um uh error
  133. 6:00rate which is decreasing ac across the
  134. 6:04years and if you can see carefully in
  135. 6:062012 that's the year where we have the
  136. 6:08biggest Improvement that is actually the
  137. 6:10year when I started my PhD at Stanford
  138. 6:13where Alex net just came and just so
  139. 6:16amazingly Blown Away a lot of the
  140. 6:17traditional techniques and later
  141. 6:20throughout the years the accuracy is
  142. 6:22just decreasing and decreasing even
  143. 6:25surpassing the human performance which
  144. 6:28is the last bar
  145. 6:30so it is really the big data and compute
  146. 6:33and algorithms that is uh driving up the
  147. 6:36performance of uh these
  148. 6:39workloads but that is no free lunch as
  149. 6:41we can see in this figure uh the xais is
  150. 6:44the max um the Amel of computation we
  151. 6:47are going to learn what is a Mac in
  152. 6:49detail in the next lecture and the Y AIS
  153. 6:52is the accuracy we can see a trend um
  154. 6:55the higher accuracy you want actually
  155. 6:57the larger the amount of compute
  156. 7:00that needs to be um through uh to run
  157. 7:03the inference right and the larger the
  158. 7:06circle means the larger the model size
  159. 7:08which means the accuracy doesn't come at
  160. 7:10free as a free lunch but you need
  161. 7:13through a lot of memory a lot of compute
  162. 7:16to achieve such memory such
  163. 7:18accuracy and what we are going to learn
  164. 7:21in this lecture is to push the fronti
  165. 7:23here okay to reduce the amount of
  166. 7:25computation but doesn't sacrifice the
  167. 7:27accuracy or even increase this the
  168. 7:28accuracy we can drastically reduce the
  169. 7:31amount of Max okay to the left corner to
  170. 7:35the left corner and then maintain even
  171. 7:37higher accuracy by neural architecture
  172. 7:39search techniques which we are going to
  173. 7:41learn in a few
  174. 7:46months and such amazing efficient AI
  175. 7:49techniques enable running these deep
  176. 7:52neuronet locally on the edge for example
  177. 7:55running on the phone we can tag our
  178. 7:58photos for example summer the trips
  179. 8:00dining different locations different
  180. 8:03lakes Cliff geysers right and also can
  181. 8:06do people recognition locally on the
  182. 8:08phone on the right hand side is our uh
  183. 8:12from our lab to do on device po
  184. 8:14estimation everything is running locally
  185. 8:16you don't have to worry about your data
  186. 8:19get transmitted to the
  187. 8:21cloud and we are going to learn those
  188. 8:24techniques not only on mobile phones but
  189. 8:27we can even TR
  190. 8:30the model such that we can deploy them
  191. 8:32on iot devices even smaller
  192. 8:36microcontrollers that is only a couple
  193. 8:38of dollars they're much smaller much
  194. 8:41cheaper but in the meantime it's more
  195. 8:43challenging because we have to make sure
  196. 8:45the model is small enough so that it can
  197. 8:47fit the microcontroller and the
  198. 8:49inference library is efficient enough so
  199. 8:52that it can run uh different workload
  200. 8:55not only classification but also
  201. 8:57detection like pH mask detection and
  202. 9:00person detection this project was done
  203. 9:02during covid so we're showing this demo
  204. 9:04using a microcontroller with only 256
  205. 9:06kilobyte of memory we can detect faces
  206. 9:10with a mask or without a
  207. 9:15mask not only on devis inference but
  208. 9:18also we are going to teach on devis
  209. 9:22TW so AI systems need to continuously
  210. 9:25adapt to new data collected from the
  211. 9:28sensors okay
  212. 9:30uh we want to process this new data
  213. 9:32locally on device so that we can have
  214. 9:35better privacy lower the cost enable
  215. 9:38customization and also enable lifelong
  216. 9:41learning okay but twinning is much more
  217. 9:45expensive than INF and why it is
  218. 9:48that we have to store those intermediate
  219. 9:50activations we have to calculate the
  220. 9:52gradient we have to do the back
  221. 9:54propagation Etc right so it's even
  222. 9:57harder to fit uh for uh Edge
  223. 10:00Hardware in this lecture we are going to
  224. 10:02also learn those on device training
  225. 10:04techniques that can drastically reduce
  226. 10:07you know amount of memory um for example
  227. 10:09in this project called on device
  228. 10:11training under 256 kilobyte of memory
  229. 10:14published two years ago we can reduce
  230. 10:16the training cost from a few hundred
  231. 10:18megabytes to only 141 kilobyte that's
  232. 10:21through order of magnitude saving by
  233. 10:23using quantization where scaling sparse
  234. 10:26update and Tiny training engine and we
  235. 10:28are going to dive deeper
  236. 10:29into such Topic in about two
  237. 10:33months and with that Technique we can
  238. 10:37have a demo here of onice tring under
  239. 10:41256 kilobyte of memory is running on the
  240. 10:45openam cam
  241. 10:48microcontroller right here we attach two
  242. 10:51buttons with with it to input the class
  243. 10:54okay we have a green class indicating
  244. 10:56there's no person we have a red class
  245. 10:58red button indicating there is
  246. 11:00person so
  247. 11:02initially the model is didn't have any
  248. 11:05training for these two classes so it's
  249. 11:08uh the classification result is wrong
  250. 11:09you can see the top right corner see
  251. 11:11there's a um red or green showing on the
  252. 11:15top right corner of the screen before
  253. 11:18training it cannot recognize
  254. 11:19distinguished person versus no person
  255. 11:23and then we apply on device training to
  256. 11:25press the button green for no person red
  257. 11:28for person
  258. 11:31to manually UT a few
  259. 11:34labels and the back propagation is done
  260. 11:37locally on the
  261. 11:40device and we have to do a couple of
  262. 11:44iterations and then let's see the
  263. 11:46classification
  264. 11:47result no person person in red no person
  265. 11:51in green and classify them
  266. 11:57perfectly and there's no Wi-Fi
  267. 12:00there's no internet everything just done
  268. 12:02locally on this
  269. 12:10microcontroller question here is the
  270. 12:23like and that's why we are showing later
  271. 12:26using a real world setting the last in
  272. 12:28the real world setting where showing is
  273. 12:30not
  274. 12:33overfeeding prompt image segmentation
  275. 12:36right so let's talk about another
  276. 12:38interesting stuff segment anything right
  277. 12:41so this is demo from meta where you
  278. 12:44can have a prompt do and then you can
  279. 12:47segment whatever you're looking at for
  280. 12:49example here looking at a person then
  281. 12:51the person will be segmented but this is
  282. 12:54very computationally happy previously
  283. 12:56people store those feature extract the
  284. 12:58features store them offline and then
  285. 13:01only run the head on the browser but uh
  286. 13:06one technique will introduc in this
  287. 13:07lecture efficient V can actually reduce
  288. 13:12um improve the throughput by 48 times
  289. 13:15that's more than order of magnitude
  290. 13:17faster with even higher accuracy on this
  291. 13:19chart 4 48 times speed up on a 100 GPU
  292. 13:24so here is showing before the
  293. 13:26acceleration and after the acceleration
  294. 13:28so before the acceleration on the first
  295. 13:29row the model is running at only 11
  296. 13:32images per second and after using
  297. 13:35efficient viit Sam we can accelerate it
  298. 13:38to 182 images per second so zero
  299. 13:43Hardware investment just using existing
  300. 13:45Hardware by making the model much more
  301. 13:47efficient uh we can um make the quality
  302. 13:52match exactly with um the Baseline model
  303. 13:56but make it run a lot faster and this is
  304. 13:58done by join who made this project even
  305. 14:00before joining MIT which is pretty
  306. 14:03amazing
  307. 14:04project so no so far we introduced the
  308. 14:08discriminative models so as algorithm
  309. 14:11advancement uh people Design This
  310. 14:13generative models that can not only
  311. 14:15recognize images but also generate new
  312. 14:17images right so these are several
  313. 14:20examples like teddy bears a bow of soup
  314. 14:24a photo of astronaut riding a horse on
  315. 14:26the Mars so diffusion models create
  316. 14:30realistic images from natural language
  317. 14:32description and we are going to learn
  318. 14:34how diffusion models work in order to
  319. 14:37accelerate it we want to First
  320. 14:39understand how diffusion model
  321. 14:42works but such amazing uh property come
  322. 14:46at very heavy it comes at very heavy
  323. 14:49computational cost so the training
  324. 14:51stable diffusion cost UH 60
  325. 14:55600,000 dollars with 256 A1 100s 100
  326. 14:5950k GPU hours which is pretty
  327. 15:03expensive and this is a figure generated
  328. 15:06by muang showing that this image or
  329. 15:09video generation model is actually much
  330. 15:11more expensive like three times more
  331. 15:14expensive just running a single step of
  332. 15:16the d diffusion model on 1K to generate
  333. 15:19a 1K by 1K image not to mention the
  334. 15:22videos which is the Red Bar um 5.6
  335. 15:27billion parameter model
  336. 15:29to generate a video not not even 1K but
  337. 15:33it's just 400 480 by 720 generate 49
  338. 15:37frames just a few seconds with a single
  339. 15:40diffusion stab it takes nine times more
  340. 15:43flops so it's very expensive uh to run
  341. 15:46these um image or video generation
  342. 15:49models okay and also as the resolution
  343. 15:53grows the amount of compute also grows
  344. 15:55quadratically with the resolution so
  345. 15:59that's a big problem we want to
  346. 16:00solve uh we did a couple of project
  347. 16:03we're going to introduce in this lecture
  348. 16:04such as the gun compression we can
  349. 16:06reduce the computation by 21 nine times
  350. 16:10and this is before compression after
  351. 16:12compression uh turning a horse into a
  352. 16:14zebra previously it's running about uh
  353. 16:1812 frames per second now it can run 40
  354. 16:21frames per second we're going to
  355. 16:23introduce techniques to achieve that and
  356. 16:26also using this any cost again technique
  357. 16:30for example previously we want to do
  358. 16:32photo editing for example make the lady
  359. 16:34smile it takes quite a while or make her
  360. 16:38look younger we can just turn the
  361. 16:40sliding bar to the right by first
  362. 16:42projecting the image to the hidden space
  363. 16:44it takes quite a while what we did and
  364. 16:47we're going to introduce is first train
  365. 16:50this once for all Network just a super
  366. 16:52Network that contain many different sub
  367. 16:54networks so during the editing phase we
  368. 16:57can use a smaller sub Network work
  369. 16:59running at lower cost for fast
  370. 17:02prototyping and finally when we are
  371. 17:04satisfied with the result we can run the
  372. 17:07a large sub Network only once and as a
  373. 17:10result you can see you can make her
  374. 17:12smile almost
  375. 17:15instantly or make her look younger we
  376. 17:18can uh Slide the second bar to the right
  377. 17:21change the hair color is almost instant
  378. 17:24okay by running this um using this once
  379. 17:27for all Technique we can derive many
  380. 17:29different uh sub networks very
  381. 17:37quickly okay and using this one4
  382. 17:40technique we can reduce the computation
  383. 17:41showing this yellow bar from 100% 78%
  384. 17:4550% 30% even 19% but you can hardly tell
  385. 17:50the quality qually the quality is still
  386. 17:52well maintained despite that we are
  387. 17:56reducing the max by up to five times
  388. 18:03uh this is another Technique we are
  389. 18:04going to introduce later in the later
  390. 18:06part of this lecture uh to accelerate
  391. 18:09image generation when we are doing the
  392. 18:11photo editing so here we are doing image
  393. 18:14in painting so what is inpainting and
  394. 18:15out painting inpainting is basically set
  395. 18:18saying that we have a a part of the
  396. 18:20image say the black part in rectangle we
  397. 18:24want to add a horse to that photograph
  398. 18:26of a horse on a grassland and usually
  399. 18:29stable diffusion takes like
  400. 18:321,855 Max costing 369 milliseconds and
  401. 18:37after the compression acceleration we
  402. 18:39can make it run only 514 gigamax at 95
  403. 18:44milliseconds similar on the right hand
  404. 18:47side to generate add a coconut tree to
  405. 18:50the image we can accelerate it by 4.8
  406. 18:54times since previous methods you need to
  407. 18:57generate all the pixels even though
  408. 18:59you're only editing a part of that but
  409. 19:02with the the S method we can do sparse
  410. 19:06update only update where you uh only
  411. 19:09update the pixels only generated pixels
  412. 19:12that needs to be updated and the the
  413. 19:14remaining pixels Remains the
  414. 19:18Same we are also going to introduce how
  415. 19:21to accelerate image Generation by
  416. 19:23parallel parallel processing so with
  417. 19:27four gpus okay can we generate even
  418. 19:30larger images even
  419. 19:33faster originally with one GPU it takes
  420. 19:36about 12 seconds right a natural thought
  421. 19:39would be if we have four gpus can we use
  422. 19:42like three three seconds four times
  423. 19:45faster unfortunately naively doing this
  424. 19:47we
  425. 19:48introduce um the artifact with a
  426. 19:51duplication using four gpus is actually
  427. 19:54generated for this uh distinct images
  428. 19:58the reason is that these four gpus
  429. 20:01didn't communicate with each other and
  430. 20:03as a result they each generated a
  431. 20:05independent
  432. 20:06image but communication will lead to the
  433. 20:09networking latency right and the way we
  434. 20:12deal with that is to overlap the
  435. 20:14computation with the networking such
  436. 20:16that um we can use a stale sta to
  437. 20:20communicate the stale um
  438. 20:23input uh to The Next Step we're going to
  439. 20:25learn that in detail in the later part
  440. 20:27of this lecture this is just for
  441. 20:28motivation and as a result we can
  442. 20:30generate um on the right column the
  443. 20:34image that is three times faster but the
  444. 20:36quality is very well
  445. 20:39maintained also people are interested
  446. 20:42customize image customization right uh
  447. 20:45create a personalized images based on
  448. 20:47user specified input and for example we
  449. 20:51want to generate something like Sunan
  450. 20:52writing a horse but it doesn't look like
  451. 20:54me writing a horse right it's very
  452. 20:57generic
  453. 20:59uh so we are going to introduce
  454. 21:00techniques like Fast composer which from
  455. 21:03function a TR free method that can blend
  456. 21:07different people together like here we
  457. 21:09can have Jensen Lady Gaga s
  458. 21:13together and this efficient image
  459. 21:15generation has a lots of applications
  460. 21:17this is just motivat why we need to
  461. 21:19learn these techniques like on the iPad
  462. 21:22we can uh you kind of draw a a picture
  463. 21:25and then after a while you see it takes
  464. 21:27several seconds we want it to be faster
  465. 21:30and to immediately render a very
  466. 21:31goodlooking image and we hope such
  467. 21:34method can run locally on the adedge
  468. 21:39device not only image generation but
  469. 21:41also 3D generation that's also very hot
  470. 21:44topic and require a lot of compute so
  471. 21:47this is diffusion models creating 3D
  472. 21:49objects from a reference image so the
  473. 21:52first rle is the reference image and the
  474. 21:56later rol is are the generated a 3D
  475. 21:58object based on just one
  476. 22:03image and this is using text to prompt
  477. 22:06using natural language as an input and
  478. 22:10create the 3D
  479. 22:14objects and also video generation that's
  480. 22:17adding another degree of Freedom uh they
  481. 22:19are pretty big uh it's even harder task
  482. 22:22requiring much much larger model size
  483. 22:26and this is from last year and this year
  484. 22:28uh Sor is there so it's pretty exciting
  485. 22:31just give a a text prompt a stylish
  486. 22:34woman walking down the Tokyo Street and
  487. 22:37then it's generating high resolution um
  488. 22:39Pho real realistic videos from natural
  489. 22:42language description but again this is
  490. 22:45consuming a huge amount of
  491. 22:47compute and we are going to learn
  492. 22:49techniques to bring down such Tech uh
  493. 22:51such
  494. 22:53computation and so far we talk a lot
  495. 22:56about 2D Vision right uh but in in
  496. 22:58reality 3D Vision is also very exciting
  497. 23:02like autonomous driving self driving
  498. 23:03cars they have not only camera sensors
  499. 23:06but also lar sensors that send those
  500. 23:09Point collect those Point cloud and to
  501. 23:12understand this 3D
  502. 23:13World unfortunately uh that require a
  503. 23:16lot of compute here is showing the whole
  504. 23:18trunk of workstation in the back of the
  505. 23:22seat you don't want to want that to hit
  506. 23:25the whole cars they uncomfortable so we
  507. 23:27want to make self-driving more efficient
  508. 23:29to fit everything in a smaller
  509. 23:33GPU so this is what we did Fast lighter
  510. 23:37net to accelerate 3D perception with
  511. 23:40algorithm and system code design uh
  512. 23:42previously it runs only five frames per
  513. 23:45second after the uh optimizations we are
  514. 23:48going to learn in this lecture about 3D
  515. 23:50Point Cloud acceleration it can run 47
  516. 23:53frames per
  517. 23:55second this is another group project
  518. 23:57we're going to introduce
  519. 23:59from Jan from our group he's not going
  520. 24:01to be a assistant professor UC ucst very
  521. 24:04soon
  522. 24:06so this project is called BB fusion um
  523. 24:10that combine not only just one sensor
  524. 24:12but also combining multiple sensors very
  525. 24:14efficiently here we are showing six
  526. 24:16cameras in the front in the back front
  527. 24:19left front right six cameras and also
  528. 24:22one liar sensor on the top and we want
  529. 24:25to combine these features to fuse
  530. 24:28multiple sensors uh to do the
  531. 24:30segmentation and also detection
  532. 24:32detection means finding the bonding box
  533. 24:343D bonding box of the cars the
  534. 24:37pedestrians and everything and on the
  535. 24:40right hand side is showing the B map
  536. 24:42segmentation BV means the bird eye view
  537. 24:45there's no map it's completely mapless
  538. 24:48mapless in the new CD you can run such
  539. 24:50algorithms to find the lanes find the
  540. 24:53driveable area find The Pedestrian
  541. 24:56walkway Etc and top right corner is
  542. 24:59showing that this algorithm is runable
  543. 25:01on Json or which is a mobile GPU for
  544. 25:04self-driving
  545. 25:08tasks okay so um that is the first part
  546. 25:12about the amazing progress of computer
  547. 25:14vision and the impl implication for
  548. 25:18computing right why do we need efficient
  549. 25:21Computing um and these applications
  550. 25:23although amazing they are they come at a
  551. 25:26high cost of computational RIS resource
  552. 25:29and now let's switch gear to the second
  553. 25:31part and talk about the advancement of
  554. 25:33natural language
  555. 25:35processing uh so this is the second year
  556. 25:38when after
  557. 25:40TBT it is very computationally heavy uh
  558. 25:43actually when chpt just came to the
  559. 25:46world usually it's at a capacity and
  560. 25:49prevent user from um have a cap of 50
  561. 25:52messages every 3 hours Etc we are
  562. 25:55experiencing exceptionally high demand
  563. 25:57please h TI and we we work scating our
  564. 26:00systems those are the bottom
  565. 26:02NE but they're really helping us for
  566. 26:05example co-pilot can make very
  567. 26:07meaningful coding suggestions based on
  568. 26:09the context we just give it some prompt
  569. 26:12some comment and then it's going to
  570. 26:15automatically generate the code uh very
  571. 26:18quickly and also neural machine
  572. 26:21translation to bre Bridge the language
  573. 26:24barrier using techniques so we are going
  574. 26:27to introduce Tech technique to bring
  575. 26:29down the model size and also reduce
  576. 26:33maintain the blue store which is the
  577. 26:35quality so here we are reducing the
  578. 26:38model size from 17
  579. 26:40176 megabytes to only 9 megabytes if it
  580. 26:44is only like 9 megabyte is very easy to
  581. 26:47fit on your
  582. 26:51phone and recently with advancement of
  583. 26:54large language model new capabilities
  584. 26:57begins to emerge for example zero short
  585. 26:59learning and also F short learning so
  586. 27:02what is zero short learning it can the
  587. 27:04model can predict the answer given only
  588. 27:07a description of the task there's no
  589. 27:10gradient there's no training given a new
  590. 27:12task you just tell it what is the task
  591. 27:14and then it's going to give you the
  592. 27:16answer say translate from English to
  593. 27:19French we're not we are not training a
  594. 27:21separate model or English to French or
  595. 27:23English to Chinese just the same model
  596. 27:26using different prompt to describe the
  597. 27:28task and zero shot means there's no
  598. 27:31example and directly translate from
  599. 27:33cheese to to French right and also the
  600. 27:37large language model have the capability
  601. 27:38called f short learning so um we just
  602. 27:42see a few example we just give it fed
  603. 27:44with a few examples of the
  604. 27:46task as part of the prompt engineering
  605. 27:49say translate to English to French and
  606. 27:51then we are going to give it three
  607. 27:53examples and then finally give the
  608. 27:56result for to translate cheese right so
  609. 27:59again there's no gradient there's no
  610. 28:00fine-tuning I just use design The Prompt
  611. 28:04SM smartly the model is capable of doing
  612. 28:07such zero short or F short
  613. 28:10learning this seems very exciting
  614. 28:12capability however like we can expect
  615. 28:15this comes at a very high cost so this
  616. 28:18is the few shot one shot or zero shot
  617. 28:21learning accuracy
  618. 28:24so to achieve higher accuracy the model
  619. 28:26size has to grow very fast um from 13 B
  620. 28:29to 175 bilon parameter um every small
  621. 28:33percentage of accuracy Improvement comes
  622. 28:36at a high cost of much larger model size
  623. 28:39and requiring more GPU resources to
  624. 28:42serve such models you see that the curve
  625. 28:45is very flat means every small
  626. 28:47percentage of accuracy Improvement can
  627. 28:49comes at a very high
  628. 28:54cost um large language model also have
  629. 28:57another very interesting emergent um
  630. 29:00capability U which is called Chain of
  631. 29:03Thought So let's see what is Chain of
  632. 29:05Thought standard prompting say Roger
  633. 29:08have five tennis balls he but buys two
  634. 29:11more cans of tennis balls each can has
  635. 29:15three tennis balls how many does he have
  636. 29:17it's
  637. 29:1811 since five time plus three * three 2
  638. 29:22* 3 is 11 and then if you ask another
  639. 29:25question given the same prompt the
  640. 29:28cafeteria has 23 IPOs if they use a 20
  641. 29:31to make lunch and B bought six more how
  642. 29:33many IPOs they
  643. 29:36have what should be the what should be
  644. 29:38the correct
  645. 29:44answer n right but the answer given by
  646. 29:48at that moment given by chb is 27
  647. 29:51obviously that's wrong right um what if
  648. 29:54we prompt it in another way say
  649. 29:59Roger given the same question of q1 the
  650. 30:01answer in blue showing Roger started
  651. 30:03with five balls two kinds of three
  652. 30:05tennis balls um two three tennis balls
  653. 30:09each is six tennis balls and five plus 6
  654. 30:12equals to 11 it describe the entire
  655. 30:15thinking process how the answer 11 is
  656. 30:19arrived and prompt in this way to show
  657. 30:22not only the result but also the Chain
  658. 30:24of Thought thinking process and given
  659. 30:27another Apple question
  660. 30:28um the answer the model can output a new
  661. 30:30answer here cafeteria at the 23 iOS
  662. 30:33originally they use a 20 uh to make
  663. 30:36lunch so they had 23 minus 20 that's
  664. 30:39three okay step by step thinking step by
  665. 30:42step and then they bought six more Apple
  666. 30:45so they now have three plus six that's
  667. 30:47equal to nine the answer is n now I can
  668. 30:49understand it correctly so we are going
  669. 30:51to introduce some of such prompt
  670. 30:53engineering techniques so that it will
  671. 30:55be very practical you can interact with
  672. 30:58these large language models more
  673. 31:03effectively again um such
  674. 31:06um um accuracy comes at a cost of very
  675. 31:10high computational cost so this is the
  676. 31:13accuracy um with um the GSM 8K middle
  677. 31:18school math world problems and then
  678. 31:21versus the access is the model size in
  679. 31:24order to achieve high accuracy the model
  680. 31:26has to be pretty big
  681. 31:30like 175 billion parameters or 500 even
  682. 31:33have a trillion
  683. 31:35parameter which is leading to the figure
  684. 31:38we showed initially the model is ring
  685. 31:40much faster than the hardware and on the
  686. 31:43right hand side is a a server in the
  687. 31:45basement uh from our lab to train uh
  688. 31:49these models we actually uh purchased it
  689. 31:52about two two two and a half years ago
  690. 31:54at that time we have 200 gabit uh infin
  691. 31:58band but the moment we purchase it the
  692. 32:00moment it becomes outdated last year I
  693. 32:02want to give the lecture the S was set
  694. 32:04of R was 400 now it's 800 so this area
  695. 32:08is just moving so fast um so it's very
  696. 32:11timely to learn these techniques to
  697. 32:14catch up the
  698. 32:15wave we also designed a couple of
  699. 32:17techniques to reduce the computation and
  700. 32:20by finding the redundancy in actal
  701. 32:22languages actually in actal language
  702. 32:24there's a lot of redundancy say give the
  703. 32:28sentence um as a visual treat the film
  704. 32:31is almost perfect that's the example on
  705. 32:34the left we can trim it the the task is
  706. 32:37to classify the sentiment so we can trim
  707. 32:40uh the sentence to be as trat feel
  708. 32:43perfect or even triming to F perfect
  709. 32:46they can still classify this is positive
  710. 32:49this is pos positive
  711. 32:53sentiment so showing that human language
  712. 32:56has a lot of uh redundant we can take
  713. 32:58advantage of that redundancy and do a
  714. 33:01sparse attention so this is a project we
  715. 33:03did about four year three four years ago
  716. 33:06called a spatter spars attention we can
  717. 33:09actually R remove those redundant uh
  718. 33:12tokens that doesn't have heav attend to
  719. 33:14other tokens like in this example I bet
  720. 33:18the video game is a lot more fun than
  721. 33:21the film video attend very heavily to
  722. 33:24game um but the word uh I and
  723. 33:29the doesn't attend to any words very
  724. 33:32heav so we can safe safely remove those
  725. 33:36tokens so those are the some of the
  726. 33:39technique we are going to dive deeper in
  727. 33:41later part of this lecture today we are
  728. 33:43just giving overview for you to get a
  729. 33:45taste what we are going to
  730. 33:49learn okay so um these L language models
  731. 33:53are pretty exciting how can we deploy
  732. 33:55them on the edge
  733. 33:58large language model on the edge would
  734. 33:59be super useful we can run co-pilot
  735. 34:02services like code completion like
  736. 34:05office or game chat everything locally
  737. 34:08on the laptops in the cars in the robots
  738. 34:11we don't have to worry about lency
  739. 34:13networking Wi-Fi or sending our private
  740. 34:17data to the cloud right
  741. 34:20um so deploying this large language
  742. 34:23model locally on the edge is super
  743. 34:24demanding and in this uh this class we
  744. 34:29have two lecture and two Labs actually
  745. 34:31dedicated to how to deploy a large
  746. 34:33language model locally on the Edge by
  747. 34:36using large language model quantization
  748. 34:38techniques so that you can run the seven
  749. 34:41billion Prim or even 13 billion primer
  750. 34:43model locally on your laptop so this is
  751. 34:45a demo showing this is a pretty updated
  752. 34:48MacBook it's only MacBook with M1 chip
  753. 34:52not only not even M3 um chip but still
  754. 34:56it can run reasonably fast rather C code
  755. 35:00uh given the prompt to sort an
  756. 35:03array and we are going to introduce
  757. 35:05techniques to quantize these large
  758. 35:07language models from 16 bit to only four
  759. 35:10bit okay gradually from 16 bit to 8 bit
  760. 35:14using smooth Quant and to even four bit
  761. 35:16using awq activation where weight only
  762. 35:19quantisation actually the smooth Quant
  763. 35:22uh paper actually comes from this
  764. 35:23lecture two years ago as one of the
  765. 35:25course project with B so maybe in this
  766. 35:28year some of you may come up with even
  767. 35:31exciting even more exciting projects
  768. 35:33that may lead to even publication in the
  769. 35:35future and the key idea is that we find
  770. 35:38lots of outliers very big values on the
  771. 35:41right top right corner in the
  772. 35:43activations and we want to make it
  773. 35:45smooth make it smooth since metrix
  774. 35:48multiplication is linear we can scale
  775. 35:50the weight and gradient so that they can
  776. 35:52be equalized and smooth such that it
  777. 35:55will be much easier to quantize and
  778. 35:56utilize the full dynamic range uh we are
  779. 36:00going to implement uh such algorithm in
  780. 36:03uh lab four in lab four and also we are
  781. 36:06going to introduce the inference library
  782. 36:09and efficient inference library in live
  783. 36:12five so that you can Implement both the
  784. 36:15algorithm the quantization algorithm and
  785. 36:18also the um inference library on your
  786. 36:21own which actually require pretty deep
  787. 36:24understanding of how computer systems
  788. 36:25work you must be able to write C program
  789. 36:29um not just python but c lowlevel c
  790. 36:31program um very comfortably manipulating
  791. 36:35uh pointers um single instruction
  792. 36:38multiple data familiar with cach how
  793. 36:40cache works with is locality with is
  794. 36:44multi threading we're are going to
  795. 36:46implement those techniques using multi
  796. 36:48core processors using CMD techniques um
  797. 36:52very and also register LEL paradism so
  798. 36:55that's why we enforce uh this
  799. 36:58prerequisites so that you can complete
  800. 37:01the
  801. 37:03labs so here we have a demo called tiny
  802. 37:06chat which is um um fast inference
  803. 37:09Library by H and Sean we can deploy a
  804. 37:12large language model 7 billion parameter
  805. 37:15large language model on Jon or Nano
  806. 37:17which is very small um mobile GPU we um
  807. 37:213D printed uh this tiny cheat computer
  808. 37:25uh to interact with uh the input in real
  809. 37:28time and this is showing we can actually
  810. 37:31run large language models on resource
  811. 37:34constrained H GPU so this is a Jon Orin
  812. 37:38uh we can t use tiny chat to run large
  813. 37:42language model on this Jon or very
  814. 37:46smoothly even 13 bitting parameter model
  815. 37:51and this is comparing with quantization
  816. 37:53without quantization it'll be much
  817. 37:55faster if you can use the activation
  818. 37:57quantization the awq so on the left hand
  819. 38:00side is showing running this large
  820. 38:03language model without quantization is
  821. 38:06slow and on the right is showing width
  822. 38:08the 4bit quantization is faster from 40
  823. 38:1250 tokens per second all the way to 166
  824. 38:16tokens per
  825. 38:19second and running large language model
  826. 38:22on the edge device can have a lot of
  827. 38:25applications so for example uh
  828. 38:28IO intelligence here I can do
  829. 38:30summarization to summarize some of your
  830. 38:32emails smart reply or emails and also
  831. 38:36writing tools to help you correct the
  832. 38:38the sentiment to make it more
  833. 38:40professional
  834. 38:42Etc uh so far we talk about Vision
  835. 38:45language and now let's jump into
  836. 38:48multimodels so what is multimodel so by
  837. 38:52we combine image language audio action
  838. 38:56different modalities
  839. 38:58into project them into the same space
  840. 39:01and understand and process all these
  841. 39:03different modalities of input and output
  842. 39:06different modalities can we take image
  843. 39:09language audio video action in and then
  844. 39:12output potentially not only text but
  845. 39:14also image video even actions okay so
  846. 39:17that's a multi model problem and the
  847. 39:20challenge here is to how do we align the
  848. 39:23representation the embedding from
  849. 39:25different input modalities
  850. 39:27so one example is uh this Ro called lava
  851. 39:30who can understand the images and take
  852. 39:34language prompt with uh the input image
  853. 39:36and output the description in language
  854. 39:39so here do you know who Dre this
  855. 39:41painting and lava is going to say uh the
  856. 39:44details about this
  857. 39:47monalisa and uh here is showing with
  858. 39:49quantization techniques we can also
  859. 39:52quantize them to only four it without
  860. 39:54losing accuracy like in this example um
  861. 39:58this is input image and some text
  862. 40:02description together with that with
  863. 40:04naive run to nearest quantization it's
  864. 40:07going to say something weird like there
  865. 40:09are small pictures of the earth and
  866. 40:11other planets placed on top of the food
  867. 40:14but with aw awq quation it says the meme
  868. 40:17in the image is light-hearted and
  869. 40:19humorous take on the concept of looking
  870. 40:21and pictures of the Earth from the space
  871. 40:24right so um with proper transition
  872. 40:27techniques we can not only accelerate
  873. 40:30the the large language models but also
  874. 40:33the visual language
  875. 40:35models later we designed uh this V the
  876. 40:38visual language model we are also going
  877. 40:40to cover how do we uh train and run
  878. 40:43difference on visual language models so
  879. 40:45the key idea is to produce the visual
  880. 40:49tokens okay so not only language can be
  881. 40:52tokenized images can also be tokenized
  882. 40:54and how do we tokenize images is by
  883. 40:57using the visual Transformers we're
  884. 40:59going to have a a dedicated lecture just
  885. 41:02to introduce what is a visual
  886. 41:03Transformer and also how do we design
  887. 41:06efficient Vision Transformers so with a
  888. 41:08vision Transformer we can project the
  889. 41:10input image into several tokens and
  890. 41:13contaminate the image token with the
  891. 41:15text token so basically uh we can expand
  892. 41:19um this idea to other modalities as well
  893. 41:22so people are realizing everything can
  894. 41:24be tokenized everything is token token
  895. 41:27in toen
  896. 41:28out and we find in this uh we in this
  897. 41:31lecture we are going to learn those
  898. 41:32training recipes like data and training
  899. 41:35recipe actually matters more than the
  900. 41:38architecture itself and why visual
  901. 41:40instruction tuning is not enough and why
  902. 41:43we need visual language pre- trining and
  903. 41:46why image text pair is not enough but we
  904. 41:49we need to have inter with image and
  905. 41:52text and how do we blend them how do we
  906. 41:54set the proportion of how much uh of
  907. 41:57each and visual in we also going to talk
  908. 42:00about visual in context learning um
  909. 42:03which is pretty exciting capability and
  910. 42:06how to we expand it from images to
  911. 42:08videos which require much longer context
  912. 42:11length since each frame will take about
  913. 42:15200 tokens if you have hourly long video
  914. 42:18you want to understand like a film that
  915. 42:20require very long context so in this
  916. 42:22lecture we are going to introduce um
  917. 42:24long context techniques how do we handle
  918. 42:27um um the linearly growing memory if na
  919. 42:33rather than rather than naively using
  920. 42:35the linearly growing KB
  921. 42:37cache so here are some of the
  922. 42:39capabilities a modern visual language
  923. 42:41model such as V can achieve uh given
  924. 42:44this video please tell me what happens
  925. 42:46in the video and vaa can say the video
  926. 42:49shows a circle game where player scores
  927. 42:51a goal and the crow cheers as the player
  928. 42:54celebrates it with his teammates okay so
  929. 42:57we have a demo link here at v.m it.edu
  930. 43:00Fore to uh to play with it given this
  931. 43:04video according to the video what will
  932. 43:06happen likely to happen next vaa says
  933. 43:09the video video shows a blender with
  934. 43:12strawberries inside and it's likely that
  935. 43:14blender will be turned on to blend the
  936. 43:17strawberries uh into a smoothie so it
  937. 43:19has the World Knowledge knowing that
  938. 43:22given such a uh few frames what's going
  939. 43:25to happen in the next that's that's also
  940. 43:27called A World model OKAY World model
  941. 43:29with the World Knowledge understand how
  942. 43:32the world Works given some frames
  943. 43:34predict what's going to happen next
  944. 43:36similar in this case what will likely
  945. 43:39happen next we says next step in the
  946. 43:41video is likely um involve the person
  947. 43:44grinding the spices spices in the
  948. 43:49motar a very interesting capability is
  949. 43:52the visual in context learning okay so
  950. 43:54what is in context learning we're not
  951. 43:57deciding what is the task but we're just
  952. 43:59showing some examples so here the prompt
  953. 44:02is image Boston image Toronto and
  954. 44:07image and it's going to say San
  955. 44:09Francisco right we're not explicitly
  956. 44:11telling that please tell me what is the
  957. 44:14city in this image but we're just
  958. 44:17showing some of the examples and then
  959. 44:19the model is going to understand okay
  960. 44:21the task is to tell uh tell the user
  961. 44:24where is the C okay so that is enabled
  962. 44:27by using the intered image text training
  963. 44:31so data blending this training recipe
  964. 44:33matters a lot in the area of FL language
  965. 44:35models and we are going to introduce
  966. 44:37visual language model in this
  967. 44:40lecture okay next one about World
  968. 44:43Knowledge so anybody can guess where
  969. 44:46where is this given this
  970. 44:50picture okay oh very nice very nice Spa
  971. 44:53pay right it require a World Knowledge
  972. 44:56you have seen lot of play is to
  973. 44:57understand that but TR the visual
  974. 44:59language model can also um tell where
  975. 45:04specific image is actually when I was TR
  976. 45:06traveling in Europe earlier this year
  977. 45:08but I clear I took a random picture and
  978. 45:11to my surprise it can tell me this is
  979. 45:14actually um a particular place described
  980. 45:17it very very accurately so I was quite
  981. 45:20Amazed by the World Knowledge they may
  982. 45:22have so feel free to play with it or
  983. 45:24even implement it in one of your
  984. 45:26homework
  985. 45:27we're also going to learn this technique
  986. 45:30can we bring this large language model
  987. 45:33Vision language model um locally to our
  988. 45:36laptops right so that we can play with
  989. 45:39it interacted with it have full control
  990. 45:43uh with it compress all the world of
  991. 45:45knowledge into a laptop so here is what
  992. 45:48we can do um for example this rep we can
  993. 45:51describe the painting in details it's
  994. 45:53running locally on a uh laptop so our
  995. 45:57lab five in lab five we are implementing
  996. 46:01a language model alone so we can also
  997. 46:05implement the visual language model as
  998. 46:06extra bonus extra credits for extra
  999. 46:09credits so feel free to talk to me if
  1000. 46:10you want to do the visual language model
  1001. 46:12for lab five something like this
  1002. 46:17demo okay so another modality would be
  1003. 46:19the action right we can use um the
  1004. 46:23visual language model also to Output
  1005. 46:25another other modalities like acttion
  1006. 46:27right so here the instruction is bring
  1007. 46:29me the rice chips from the Jer okay and
  1008. 46:33the current step is opening the Jer it
  1009. 46:35helps do the planning since using the
  1010. 46:39the large language model contains a lot
  1011. 46:40of word model it understand in order to
  1012. 46:44bring the chip you first need to open
  1013. 46:45the Jer and then take the right chip out
  1014. 46:48of the Jer and then place it um and then
  1015. 46:51close the J Etc right um but
  1016. 46:55unfortunately this is by 4X uh speed
  1017. 46:59right is Rise only three Herz due to the
  1018. 47:03computational cost networking latency
  1019. 47:06because the the images are transmitted
  1020. 47:08over internet to a workstation and then
  1021. 47:12transmitted back the final result okay
  1022. 47:15so this is the rt1 from Google is also
  1023. 47:19already pretty uh amazing but highlights
  1024. 47:23how important it is to uh learn this
  1025. 47:26efficient depl Computing techniques like
  1026. 47:28this
  1027. 47:29lecture similarly deep learning for
  1028. 47:32games um our for go beating so do is
  1029. 47:36taking
  1030. 47:371,920 CPUs 280 gpus $3,000 electric bill
  1031. 47:42per game that's super expensive or the
  1032. 47:46arpha fold require 16 TP v3s which is
  1033. 47:51128 tp3 course um train for a few weeks
  1034. 47:56so all these amazing um capabilities
  1035. 48:00comes at a high
  1036. 48:02cost okay so let's switch gear from
  1037. 48:06visual language to multimodality now we
  1038. 48:09talk about these three pillars algorithm
  1039. 48:12hardware and data in particular we want
  1040. 48:13to talk highlight the uh advancement of
  1041. 48:16Hardware that is the driving force for
  1042. 48:19this C Breen of of efficient deer
  1043. 48:25Computing so this lecture we are going
  1044. 48:27to introduce some of the hardware um
  1045. 48:30Hardware system techniques as well and
  1046. 48:33the philosophy is that we want to do the
  1047. 48:35whole full stack full stack design okay
  1048. 48:37on the software part full stack design
  1049. 48:40from algorithm part the demand of
  1050. 48:42computing to the system part which is
  1051. 48:44the supply of computing we want to close
  1052. 48:46the gap by reducing the demand and
  1053. 48:48increasing the supply okay so the recent
  1054. 48:52trend of Modern Hardware design is that
  1055. 48:55uh the microprocessor frequency is
  1056. 48:58plateaued um the single core performance
  1057. 49:01also plateaued more SL is scaling down
  1058. 49:04but the number of cores is increasing
  1059. 49:07and the the number of transistors is
  1060. 49:08also increasing uh so parallel Computing
  1061. 49:12specialized uh Computing is the new
  1062. 49:15trend in this
  1063. 49:16context and also Precision matters
  1064. 49:20people uh the communi is keep advancing
  1065. 49:23the low Precision arithmetic from
  1066. 49:25floating point 32 floating Point 16 to
  1067. 49:29integer 8 8 bit integer to fp8 8 bit
  1068. 49:33floating Point um and also even to fp4
  1069. 49:37four bit floating point which is to
  1070. 49:40appear in the our latest black well uh
  1071. 49:43immedia gpus right so the Precision is
  1072. 49:46keep decreasing in this lecture we are
  1073. 49:48going to learn all the details about
  1074. 49:50what is a floating Point what is fp8 uh
  1075. 49:53what is um E5 M2 okay what is the
  1076. 49:57mantisa so what is fp4 why Blackwell is
  1077. 50:02so amazing it can have such high Peak
  1078. 50:05Performance in uh the quantization part
  1079. 50:07of this
  1080. 50:08lecture and on the right hand side is
  1081. 50:10showing the single chip inference
  1082. 50:13performance 300 times in eight years
  1083. 50:16this is from the slide from my PhD
  1084. 50:18adviser Professor Bill Ali is showing
  1085. 50:21that from scaler FP 32 to
  1086. 50:24fp6 um to H hmma tensor core uh in
  1087. 50:30v00 the intake TS increased to 125 and
  1088. 50:35again it doubled by using int8 int8 imma
  1089. 50:39and course to using structur spity 24
  1090. 50:43sparcity in
  1091. 50:44a100 like 300 performance improv in
  1092. 50:47eight years that's like pretty amazing
  1093. 50:51Improvement and in this along this line
  1094. 50:53of innovation software is actually
  1095. 50:56playing a very critical role um in
  1096. 50:58especially the advanced technology node
  1097. 51:01so this is from 65 nanometer 28
  1098. 51:04nanometer 7 nanometer 5 nanometer you
  1099. 51:07can see the purple part which is a
  1100. 51:09software component is getting a very big
  1101. 51:12chunk of the total Advanced design cost
  1102. 51:17and that's always also creating the mold
  1103. 51:20for a good cicon to uh be able to
  1104. 51:23attract a lot of good customers it has
  1105. 51:25to be programmable has to be easy to use
  1106. 51:27especially the large language model this
  1107. 51:29generative AI the models are changing
  1108. 51:32iterating very fast changing every day
  1109. 51:35um how to make it flexible is super
  1110. 51:37important so in this lecture we have a
  1111. 51:40lot of focus on the software part of
  1112. 51:43efficient deep learning Computing uh we
  1113. 51:45don't design new Hardware in this
  1114. 51:46lecture but how to well utilize existing
  1115. 51:50hardware and what is the um implication
  1116. 51:52to design new hardware so here we show
  1117. 51:55from 20 16 to
  1118. 51:582024 uh the advancement of gpus uh we
  1119. 52:01put them into four different uh
  1120. 52:04benchmarks you know the dance fp16
  1121. 52:07performance measured in tal operations
  1122. 52:10per second we're going to learn in the
  1123. 52:11next lecture what is T Ops per second
  1124. 52:13what does it mean what is Dan what is
  1125. 52:16sparse what is fp16 so we're going to
  1126. 52:19learn dance and sparse in the pring
  1127. 52:20lecture from uh next Thursday and we are
  1128. 52:24going to learn uh quantization precision
  1129. 52:26a week after so from p00 to b00 is
  1130. 52:31actually 100 times 100 times more um TS
  1131. 52:36per second in particular from h100 to B
  1132. 52:40100 even doubled uh the the TS per
  1133. 52:44second which is super amazing and even
  1134. 52:46didn't consider the fp4 if we add fp4 it
  1135. 52:48will be way higher in this chart because
  1136. 52:51it's showing just
  1137. 52:53fp16 and here is the memory bandwidth
  1138. 52:56computation is cheap memory is super
  1139. 52:58expensive uh from a100 to h100 it almost
  1140. 53:03doubled and then more than doubled from
  1141. 53:05h100 to b00 with respect to the memory
  1142. 53:08bandwidth and we are going to introduce
  1143. 53:10these Concepts like why data movement is
  1144. 53:13expensive um why energy is dominated by
  1145. 53:17moving data and here is showing the
  1146. 53:20power unfortunately the power is also
  1147. 53:22growing pretty fast from a1004 400 watts
  1148. 53:26to 700 watts in h100 and B 100 if you
  1149. 53:31have eight um gpus in a node you can
  1150. 53:35calculate how much power you have so
  1151. 53:38eight node all the gpus will take maybe
  1152. 53:404,000 5,000 watts and all all together
  1153. 53:44the whole node we have 10,000 Watts you
  1154. 53:46have to prepare 10,000 Watts just for a
  1155. 53:49single node and as a result you the
  1156. 53:52space will not be the constraint but the
  1157. 53:54power power supply the ener energ in
  1158. 53:57cooling is the new Bott
  1159. 54:00neck previously when we are having the
  1160. 54:03a6000 gpus so we can have two two cables
  1161. 54:06serve the entire rack but now two cables
  1162. 54:09can only serve four node of a100 and
  1163. 54:11only two node of h100 so power
  1164. 54:15efficiency is super
  1165. 54:16important and in this lecture later part
  1166. 54:20of this lecture we are going to organize
  1167. 54:21a lab tour so we'll show you down to the
  1168. 54:24basement uh to our uh server so that you
  1169. 54:28can see how these models are trained
  1170. 54:30with what is a like a mini data center
  1171. 54:32what does it look like what is the code
  1172. 54:34I hot a the rack the switch Infinity
  1173. 54:37band The the cables and modern GPU what
  1174. 54:40does it look
  1175. 54:43like and next is the memory so
  1176. 54:45unfortunately memory is growing at a
  1177. 54:47slower speed compared with a compute um
  1178. 54:50so current what we have in our lab is 80
  1179. 54:54Gab a100 gpus h100 pretty much stays the
  1180. 54:58same and b00 give a pretty exciting leap
  1181. 55:01to almost 200 192 gigabytes so this is
  1182. 55:05for the cloud for training and
  1183. 55:08serving uh so let's also look at the
  1184. 55:12edge uh for example on the phone we want
  1185. 55:14to have those low power uh chips like
  1186. 55:17here is only 10 watts 10 watts s855 all
  1187. 55:21the way to I8 gen one and I8 Gen 2 is
  1188. 55:25also there but we didn't find the public
  1189. 55:27uh numbers so we didn't put it there the
  1190. 55:29performance is also increasing pretty
  1191. 55:31fast like couple of um uh like almost
  1192. 55:35like half a 100 like 50 TS per second
  1193. 55:40and memory um is roughly eight 16
  1194. 55:44gigabytes for Qualcomm
  1195. 55:46DSP Apple also offers this apple neural
  1196. 55:49engine energy efficient and high
  1197. 55:51throughput for machine learning
  1198. 55:53applications uh similarly with uh the
  1199. 55:56qualcom DSP is 35 TS per second however
  1200. 56:01we should not just look at the Peak
  1201. 56:03Performance like 35 or 52 can you make a
  1202. 56:07conclusion that apple or qualcom which
  1203. 56:10one is faster actually cannot because
  1204. 56:12it's a multiple Dimension um the Peak
  1205. 56:16Performance doesn't indicate doesn't
  1206. 56:19directly translate to measured uh speed
  1207. 56:23there's so many factors we are going to
  1208. 56:24learn in this lecture like the
  1209. 56:26activation data movement weight data
  1210. 56:28movement memory bandwidth utilization
  1211. 56:31all matters a
  1212. 56:35lot and also there's the mobile GPU
  1213. 56:38people very widely used in the cars for
  1214. 56:41example this is the Json a which is used
  1215. 56:44in many EVS um delivers couple of
  1216. 56:47hundred TS per second consume about 60
  1217. 56:50watts 64 gigabytes of memory and there's
  1218. 56:53going to be a new um Thor coming coming
  1219. 56:56up which is the next generation of ajx
  1220. 57:01or and also there's microcontrollers for
  1221. 57:04many iot devices like the your smart
  1222. 57:07home camera might have one of these um
  1223. 57:11small devices arm devices just a couple
  1224. 57:14of hundred mwatts by the memory is also
  1225. 57:16super small just a a few hundred
  1226. 57:19kilobytes of s SRAM maybe a few
  1227. 57:21megabytes of uh um maybe Dam or even no
  1228. 57:25Dam some M controllers complet has
  1229. 57:28completely no Dam we have to manage the
  1230. 57:30memory super
  1231. 57:33carefully so from cloud AI to mobile AI
  1232. 57:36to timing AI there's actually a big gap
  1233. 57:39between the model memory uh that you
  1234. 57:43have tens of gigabytes versus couple of
  1235. 57:46hundred kilobytes so um Ed AI device do
  1236. 57:50have a huge gap to Cloud processors and
  1237. 57:53what we learn in this lecture is going
  1238. 57:55to try to close the Gap by using smaller
  1239. 57:59models okay so those are the three
  1240. 58:02pillars the hardware Theta and the
  1241. 58:04algorithm and our goal is rather than
  1242. 58:07using a lot of compute emitting a lot of
  1243. 58:09carbon using many Engineers TR a lot of
  1244. 58:12data tin ml's goal will be how to use
  1245. 58:15less computation less carbon few
  1246. 58:17Engineers automated TR L
  1247. 58:21data so this is the course overview let
  1248. 58:24me use the remaining 10 minutes to cover
  1249. 58:26some of the
  1250. 58:28logistics so the lecture is every uh
  1251. 58:31Tuesday and Thursday from uh 3 uh 35 to
  1252. 58:364:55 in this classroom and we will
  1253. 58:39record uh the
  1254. 58:41classes um the office hour will be after
  1255. 58:44directly after uh this lecture on every
  1256. 58:48uh Thursday from 5 to uh 6:00 p.m. uh
  1257. 58:52the room number is 38344 that's where we
  1258. 58:55our office is located at feel free to
  1259. 58:58ask questions during the office hour or
  1260. 59:01go to the PIAA make sure you sign up on
  1261. 59:03Pasa and also we use canas like other
  1262. 59:06classes to submit the
  1263. 59:10homework and you can email us efficient
  1264. 59:13ml- staff at mit.edu if you have any uh
  1265. 59:18uh external inquiries we also going to
  1266. 59:21create a maing list you make sure you
  1267. 59:23take care of the maing list so everybody
  1268. 59:25is wel come to join that maing list so
  1269. 59:28that we can broadcast some of the news
  1270. 59:31project ideas some company even want to
  1271. 59:34uh have those internship openings or
  1272. 59:37full-time job openings we can
  1273. 59:39disseminate uh through that meeting list
  1274. 59:42so um stay tuned regene will send a
  1275. 59:45notification for the maing list to sign
  1276. 59:48up and the prerequisites so this year we
  1277. 59:51enforce two prerequisites one is the
  1278. 59:546191 which is comput ation structures
  1279. 59:57since the LA five and La four require a
  1280. 1:00:01deep understanding of how computer
  1281. 1:00:02systems work like what is cach what is
  1282. 1:00:05locality what is paradism five stage
  1283. 1:00:09pipeline um page table since we're going
  1284. 1:00:12to talk about the page attention um
  1285. 1:00:15speculation uh we're going to talk about
  1286. 1:00:17speculative decoding and um most
  1287. 1:00:21importantly you have to be familiar with
  1288. 1:00:23comfortable with how to write C programs
  1289. 1:00:26um since the last the lab will be
  1290. 1:00:28written in C program and also mod
  1291. 1:00:31threading techniques the second
  1292. 1:00:33prerequisite is
  1293. 1:00:356.3 90 the introduction to machine
  1294. 1:00:37learning
  1295. 1:00:38class make sure you are familiar with
  1296. 1:00:40how to use pytorch um since all the labs
  1297. 1:00:44will be using
  1298. 1:00:49pytorch so we are going to cover roughly
  1299. 1:00:52three sections first section is about
  1300. 1:00:54efficient inference technique
  1301. 1:00:56with pruning reducing the number of
  1302. 1:00:58parameters fantization reducing the
  1303. 1:01:01Precision for each parameter and also
  1304. 1:01:03neural architecture search how do we
  1305. 1:01:05design efficient neural network
  1306. 1:01:08architecture and also distillation how
  1307. 1:01:10to use a teacher model to distill a
  1308. 1:01:12smaller model we're also going to talk
  1309. 1:01:14about efficient training technique how
  1310. 1:01:17to do grading compression on device
  1311. 1:01:19training and also fight rate learning
  1312. 1:01:21and parallelization including data
  1313. 1:01:23parallel model model parallel Pine
  1314. 1:01:26parallel and also sequence parallel for
  1315. 1:01:28large language
  1316. 1:01:30model we're also going to introduce
  1317. 1:01:32application specific optimizations the
  1318. 1:01:34most important one of course is
  1319. 1:01:36Transformers and large language models
  1320. 1:01:38we're going to introduce the entire
  1321. 1:01:40architecture how Transformer work from
  1322. 1:01:44um the position encoding all the way to
  1323. 1:01:46what is qkv what is attention what is um
  1324. 1:01:50different uh contact extension
  1325. 1:01:52techniques with this rope Etc or also
  1326. 1:01:56going to introduce diffusion models um
  1327. 1:01:58um how how diffusion model Works video
  1328. 1:02:01understanding Point Cloud understanding
  1329. 1:02:04Etc so a key concept is system on
  1330. 1:02:08algorithm code design which is covering
  1331. 1:02:10actually combining a lot of knowledge we
  1332. 1:02:12learn from the e side the Cs side and
  1333. 1:02:15also the a plusd side so hopefully it
  1334. 1:02:18will be a class you can take um in later
  1335. 1:02:21year of your undergrad or maybe in the
  1336. 1:02:23first couple of years of your PhD to do
  1337. 1:02:26interdisciplinary research on the E side
  1338. 1:02:30um the
  1339. 1:02:316191 or micro computer project lab
  1340. 1:02:35Hardware architecture for deep learning
  1341. 1:02:36those are very related classes csite
  1342. 1:02:39computer architecture software
  1343. 1:02:41performance engineering mobile s sensor
  1344. 1:02:43Computing a plusd side intro to machine
  1345. 1:02:46learning deep learning advances in
  1346. 1:02:48computer vision natural language
  1347. 1:02:49processing all related
  1348. 1:02:51classes we will have five Labs um over
  1349. 1:02:56the course of the semester we are going
  1350. 1:02:57to use Google collab for the first four
  1351. 1:03:01Labs so you don't have to worry about
  1352. 1:03:03Hardware uh first one is about pruning
  1353. 1:03:06second one about quation uh third one is
  1354. 1:03:08about neural architecture search fourth
  1355. 1:03:11is about large language model
  1356. 1:03:12compression and the fifth one does
  1357. 1:03:15require Hardware which is your laptop
  1358. 1:03:18hopefully you all have a laptop um that
  1359. 1:03:20is more than 8 gigabyte of memory make
  1360. 1:03:23sure you have more than 8 gigabytes of
  1361. 1:03:25memory and available storage should be
  1362. 1:03:27at least five
  1363. 1:03:29gigabytes so unfortunately we don't have
  1364. 1:03:31the budget to BU everyone laptop but I'm
  1365. 1:03:34assuming everyone already have a laptop
  1366. 1:03:37either x86 or arm um either Mac Linux or
  1367. 1:03:42Windows we have carefully debugged so
  1368. 1:03:44that it can fit across different
  1369. 1:03:46platforms and JY will be the TA if you
  1370. 1:03:48have any questions related to live five
  1371. 1:03:51and and uh
  1372. 1:03:54Hardware so greeing we will have five
  1373. 1:03:57lives each one is
  1374. 1:03:5815% uh followed up followed by a final
  1375. 1:04:02project 25% is open-ended but we'll give
  1376. 1:04:05you some project ideas to start with it
  1377. 1:04:08is not in a group of three to four uh
  1378. 1:04:10students consisting of a proposal um a
  1379. 1:04:14presentation poster presentation and Al
  1380. 1:04:17also a final report just
  1381. 1:04:1920% uh we also give uh four 4% for
  1382. 1:04:23participation bonus and all the
  1383. 1:04:26assignments are due on 11:59 p.m. on the
  1384. 1:04:28due
  1385. 1:04:29date um and the late policy uh is that
  1386. 1:04:33we have six total um late days without
  1387. 1:04:37penalty for the entire semester uh given
  1388. 1:04:39the large number of enrollment we don't
  1389. 1:04:41have any exception for the six L
  1390. 1:04:45days and the allow days are counted by
  1391. 1:04:48day each new day will um the panalty
  1392. 1:04:51will be 50 20% 50% so homework is worth
  1393. 1:04:55zero credit two days
  1394. 1:04:58later so prerequisite all this is quite
  1395. 1:05:00important so student who don't fulfill
  1396. 1:05:02the prerequisite will be D registered in
  1397. 1:05:05the second week of the class so if you
  1398. 1:05:07believe you have equivalent prior
  1399. 1:05:10experience including both computer
  1400. 1:05:12architecture and also the machine
  1401. 1:05:14learning please make sure to send us the
  1402. 1:05:18petition on this form make sure you take
  1403. 1:05:20a photo of this form if you don't meet
  1404. 1:05:22the prerequisites to submit your
  1405. 1:05:24petition form
  1406. 1:05:26by this Friday
  1407. 1:05:2811:49 p.m. okay and then our TA will
  1408. 1:05:32review this petitions and student who do
  1409. 1:05:36doesn't fulfill the prerequisite will be
  1410. 1:05:38D register next
  1411. 1:05:40week okay so these are the uh five labs
  1412. 1:05:44and also together with the uh final
  1413. 1:05:47projects we did last year it's pretty
  1414. 1:05:50exciting so hopefully you can learn a
  1415. 1:05:52lot about this amazing um efficient Tech
  1416. 1:05:55Tech and we'll see you next Tuesday

About this transcript

This page contains the full transcript of EfficientML.ai Lecture 1 - Introduction (MIT 6.5940, Fall 2024, Zoom Recording) by MIT HAN Lab, generated from the public captions YouTube serves with the video. The transcript has 9,527 words across 1,416 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.