EfficientML.ai Lecture 1 - Introduction (MIT 6.5940, Fall 2024, Zoom Recording) — Transcript
Full transcript
- 0:00welcome to Tiny ML and efficient deeper
- 0:02Computing 2024 I'm Professor Sunan very
- 0:06glad to teach this class again uh this
- 0:09is the third time we offer uh this class
- 0:12we're going to learn efficient deep
- 0:14learning Computing how to make new
- 0:16networks run faster train with faster
- 0:18and influence faster I remember two
- 0:21years ago when we first opened this
- 0:23course only 30 students last year it was
- 0:27uh 90 this year is more than 200 so
- 0:30welcome all I'm very glad to have you
- 0:32all here and hopefully we'll have a very
- 0:35fruitful semester all together so this
- 0:39is a brief introduction of myself I
- 0:42graduated from chinga University and get
- 0:44my PhD from Stanford with Professor uh
- 0:48Bill Daddy I worked on model compression
- 0:50and efficient deeper Computing during my
- 0:53PhD uh published a deep compression
- 0:55paper which set the foundation for
- 0:57modern neuron Network acceleration
- 1:00compression and published paper called
- 1:02efficient inference engine which is
- 1:04actually top five CED paper in 50 years
- 1:06of Isa uh during my PhD I co-founded a
- 1:09startup called defi tag for um efficient
- 1:14air chip which was later acquired by
- 1:16zings and later zings is acquired by AMD
- 1:19that's very popular story in City cona
- 1:23so in this course I will teach some of
- 1:25the um how to make sure the research is
- 1:27very fundament not only fundamental but
- 1:29also practical so I joined MIT about six
- 1:33years ago uh very exciting Journey uh
- 1:36we've been working on from Tiny ml to
- 1:38large language model acceleration and
- 1:41compression one of the most exciting
- 1:43paper we did in the past few years is
- 1:46the awq for large language model
- 1:49quantization which got the best paper
- 1:51award at mlis and also it's having
- 1:54already more than six million downloads
- 1:56in and hugging phase so if you want to
- 1:58deploy a large language model make it
- 2:00four times smaller then awq is the
- 2:03technique to go and we will learn that
- 2:05in this lecture you're going to
- 2:07implement that not only Implement that
- 2:09but also deploy a large language model
- 2:12locally on your
- 2:13laptop which is will be pretty exciting
- 2:17so at MIT I also Co found a startup
- 2:19called oml working on efficient
- 2:23different Computing model
- 2:25optimizations so in this lecture we'll
- 2:27learn all the secret sources from
- 2:30those two startups so those five
- 2:32homework will be very very helpful if
- 2:36you learn the master the m so we very
- 2:39carefully
- 2:40created Five assignments five five
- 2:43homeworks they're all hands on homeworks
- 2:46so we try to uh for you guys use the
- 2:49minimum effort to learn the most amount
- 2:51of knowledge the best utilize your time
- 2:53so please uh do those homework um very
- 2:58well
- 3:00yeah we'll have five labs and each TI
- 3:01will be responsible for one of them and
- 3:04we have office hours every Tuesday and
- 3:06Thursday and feel free to come to our
- 3:08office and we can help you with any
- 3:11questions you may
- 3:12have okay let's start by motivating why
- 3:16do we need efficient de Computing so
- 3:19let's see the supply and demand for a
- 3:22Computing so this chart shows according
- 3:26throughout different years the model
- 3:28size of large language model in red
- 3:31versus the GPU memory in green we can
- 3:35see the large language model size grow
- 3:38from 0.05 billion five billion
- 3:42parameters the initial Transformer in
- 3:442018 2017 2018 to about um 1.5 billion
- 3:50parameter in tbd2 roughly in 2019
- 3:532020 and recently uh there's even TR
- 3:58parameter models being trained in the
- 4:00world right it's growing super fast even
- 4:03much faster than the amount of GPU
- 4:05memories that is available in a single
- 4:08GPU right which is showing in uh showing
- 4:12uh green so this big gap the Gap is
- 4:15getting larger
- 4:17actually with recent years and if you
- 4:20see the supply and demand curve that is
- 4:22driving up the price right making deep
- 4:24learning super expensive to train super
- 4:27expensive to serve right so we are going
- 4:29going to learn techniques to bridge the
- 4:32gap using efficient model compression
- 4:34techniques so that we can improve the
- 4:37utilization of the hardware and also
- 4:39reduce the models size memory
- 4:42requirement computation requirement to
- 4:45bridge this
- 4:46Gap so rather than a TR a model and
- 4:49directly R inference on the top we now
- 4:52can TR a model compress it before
- 4:55deploying the compress model for un
- 5:00so in recent years the amount of
- 5:02Publications in model compression and
- 5:04acceleration on compressed sparse model
- 5:07is uh getting a lot faster number of
- 5:10Publications is increasing very fast in
- 5:13recent years and we are going to learn
- 5:14these techniques how do we utilize
- 5:16compression and sparity to accelerate a
- 5:19neuron Network
- 5:23inference so um in this lecture we are
- 5:26going to uh start with three three
- 5:29sections from Vision to language to
- 5:32multiple
- 5:33modalities um to understand the
- 5:36computation
- 5:38demand uh of this AI workload and see
- 5:41where we
- 5:42are among the fronti here so let's start
- 5:45with vision tasks so talking about comp
- 5:48Vision we have to mention image net and
- 5:51also alexnet right so Alex net this is
- 5:54actually the uh image net um um uh error
- 6:00rate which is decreasing ac across the
- 6:04years and if you can see carefully in
- 6:062012 that's the year where we have the
- 6:08biggest Improvement that is actually the
- 6:10year when I started my PhD at Stanford
- 6:13where Alex net just came and just so
- 6:16amazingly Blown Away a lot of the
- 6:17traditional techniques and later
- 6:20throughout the years the accuracy is
- 6:22just decreasing and decreasing even
- 6:25surpassing the human performance which
- 6:28is the last bar
- 6:30so it is really the big data and compute
- 6:33and algorithms that is uh driving up the
- 6:36performance of uh these
- 6:39workloads but that is no free lunch as
- 6:41we can see in this figure uh the xais is
- 6:44the max um the Amel of computation we
- 6:47are going to learn what is a Mac in
- 6:49detail in the next lecture and the Y AIS
- 6:52is the accuracy we can see a trend um
- 6:55the higher accuracy you want actually
- 6:57the larger the amount of compute
- 7:00that needs to be um through uh to run
- 7:03the inference right and the larger the
- 7:06circle means the larger the model size
- 7:08which means the accuracy doesn't come at
- 7:10free as a free lunch but you need
- 7:13through a lot of memory a lot of compute
- 7:16to achieve such memory such
- 7:18accuracy and what we are going to learn
- 7:21in this lecture is to push the fronti
- 7:23here okay to reduce the amount of
- 7:25computation but doesn't sacrifice the
- 7:27accuracy or even increase this the
- 7:28accuracy we can drastically reduce the
- 7:31amount of Max okay to the left corner to
- 7:35the left corner and then maintain even
- 7:37higher accuracy by neural architecture
- 7:39search techniques which we are going to
- 7:41learn in a few
- 7:46months and such amazing efficient AI
- 7:49techniques enable running these deep
- 7:52neuronet locally on the edge for example
- 7:55running on the phone we can tag our
- 7:58photos for example summer the trips
- 8:00dining different locations different
- 8:03lakes Cliff geysers right and also can
- 8:06do people recognition locally on the
- 8:08phone on the right hand side is our uh
- 8:12from our lab to do on device po
- 8:14estimation everything is running locally
- 8:16you don't have to worry about your data
- 8:19get transmitted to the
- 8:21cloud and we are going to learn those
- 8:24techniques not only on mobile phones but
- 8:27we can even TR
- 8:30the model such that we can deploy them
- 8:32on iot devices even smaller
- 8:36microcontrollers that is only a couple
- 8:38of dollars they're much smaller much
- 8:41cheaper but in the meantime it's more
- 8:43challenging because we have to make sure
- 8:45the model is small enough so that it can
- 8:47fit the microcontroller and the
- 8:49inference library is efficient enough so
- 8:52that it can run uh different workload
- 8:55not only classification but also
- 8:57detection like pH mask detection and
- 9:00person detection this project was done
- 9:02during covid so we're showing this demo
- 9:04using a microcontroller with only 256
- 9:06kilobyte of memory we can detect faces
- 9:10with a mask or without a
- 9:15mask not only on devis inference but
- 9:18also we are going to teach on devis
- 9:22TW so AI systems need to continuously
- 9:25adapt to new data collected from the
- 9:28sensors okay
- 9:30uh we want to process this new data
- 9:32locally on device so that we can have
- 9:35better privacy lower the cost enable
- 9:38customization and also enable lifelong
- 9:41learning okay but twinning is much more
- 9:45expensive than INF and why it is
- 9:48that we have to store those intermediate
- 9:50activations we have to calculate the
- 9:52gradient we have to do the back
- 9:54propagation Etc right so it's even
- 9:57harder to fit uh for uh Edge
- 10:00Hardware in this lecture we are going to
- 10:02also learn those on device training
- 10:04techniques that can drastically reduce
- 10:07you know amount of memory um for example
- 10:09in this project called on device
- 10:11training under 256 kilobyte of memory
- 10:14published two years ago we can reduce
- 10:16the training cost from a few hundred
- 10:18megabytes to only 141 kilobyte that's
- 10:21through order of magnitude saving by
- 10:23using quantization where scaling sparse
- 10:26update and Tiny training engine and we
- 10:28are going to dive deeper
- 10:29into such Topic in about two
- 10:33months and with that Technique we can
- 10:37have a demo here of onice tring under
- 10:41256 kilobyte of memory is running on the
- 10:45openam cam
- 10:48microcontroller right here we attach two
- 10:51buttons with with it to input the class
- 10:54okay we have a green class indicating
- 10:56there's no person we have a red class
- 10:58red button indicating there is
- 11:00person so
- 11:02initially the model is didn't have any
- 11:05training for these two classes so it's
- 11:08uh the classification result is wrong
- 11:09you can see the top right corner see
- 11:11there's a um red or green showing on the
- 11:15top right corner of the screen before
- 11:18training it cannot recognize
- 11:19distinguished person versus no person
- 11:23and then we apply on device training to
- 11:25press the button green for no person red
- 11:28for person
- 11:31to manually UT a few
- 11:34labels and the back propagation is done
- 11:37locally on the
- 11:40device and we have to do a couple of
- 11:44iterations and then let's see the
- 11:46classification
- 11:47result no person person in red no person
- 11:51in green and classify them
- 11:57perfectly and there's no Wi-Fi
- 12:00there's no internet everything just done
- 12:02locally on this
- 12:10microcontroller question here is the
- 12:23like and that's why we are showing later
- 12:26using a real world setting the last in
- 12:28the real world setting where showing is
- 12:30not
- 12:33overfeeding prompt image segmentation
- 12:36right so let's talk about another
- 12:38interesting stuff segment anything right
- 12:41so this is demo from meta where you
- 12:44can have a prompt do and then you can
- 12:47segment whatever you're looking at for
- 12:49example here looking at a person then
- 12:51the person will be segmented but this is
- 12:54very computationally happy previously
- 12:56people store those feature extract the
- 12:58features store them offline and then
- 13:01only run the head on the browser but uh
- 13:06one technique will introduc in this
- 13:07lecture efficient V can actually reduce
- 13:12um improve the throughput by 48 times
- 13:15that's more than order of magnitude
- 13:17faster with even higher accuracy on this
- 13:19chart 4 48 times speed up on a 100 GPU
- 13:24so here is showing before the
- 13:26acceleration and after the acceleration
- 13:28so before the acceleration on the first
- 13:29row the model is running at only 11
- 13:32images per second and after using
- 13:35efficient viit Sam we can accelerate it
- 13:38to 182 images per second so zero
- 13:43Hardware investment just using existing
- 13:45Hardware by making the model much more
- 13:47efficient uh we can um make the quality
- 13:52match exactly with um the Baseline model
- 13:56but make it run a lot faster and this is
- 13:58done by join who made this project even
- 14:00before joining MIT which is pretty
- 14:03amazing
- 14:04project so no so far we introduced the
- 14:08discriminative models so as algorithm
- 14:11advancement uh people Design This
- 14:13generative models that can not only
- 14:15recognize images but also generate new
- 14:17images right so these are several
- 14:20examples like teddy bears a bow of soup
- 14:24a photo of astronaut riding a horse on
- 14:26the Mars so diffusion models create
- 14:30realistic images from natural language
- 14:32description and we are going to learn
- 14:34how diffusion models work in order to
- 14:37accelerate it we want to First
- 14:39understand how diffusion model
- 14:42works but such amazing uh property come
- 14:46at very heavy it comes at very heavy
- 14:49computational cost so the training
- 14:51stable diffusion cost UH 60
- 14:55600,000 dollars with 256 A1 100s 100
- 14:5950k GPU hours which is pretty
- 15:03expensive and this is a figure generated
- 15:06by muang showing that this image or
- 15:09video generation model is actually much
- 15:11more expensive like three times more
- 15:14expensive just running a single step of
- 15:16the d diffusion model on 1K to generate
- 15:19a 1K by 1K image not to mention the
- 15:22videos which is the Red Bar um 5.6
- 15:27billion parameter model
- 15:29to generate a video not not even 1K but
- 15:33it's just 400 480 by 720 generate 49
- 15:37frames just a few seconds with a single
- 15:40diffusion stab it takes nine times more
- 15:43flops so it's very expensive uh to run
- 15:46these um image or video generation
- 15:49models okay and also as the resolution
- 15:53grows the amount of compute also grows
- 15:55quadratically with the resolution so
- 15:59that's a big problem we want to
- 16:00solve uh we did a couple of project
- 16:03we're going to introduce in this lecture
- 16:04such as the gun compression we can
- 16:06reduce the computation by 21 nine times
- 16:10and this is before compression after
- 16:12compression uh turning a horse into a
- 16:14zebra previously it's running about uh
- 16:1812 frames per second now it can run 40
- 16:21frames per second we're going to
- 16:23introduce techniques to achieve that and
- 16:26also using this any cost again technique
- 16:30for example previously we want to do
- 16:32photo editing for example make the lady
- 16:34smile it takes quite a while or make her
- 16:38look younger we can just turn the
- 16:40sliding bar to the right by first
- 16:42projecting the image to the hidden space
- 16:44it takes quite a while what we did and
- 16:47we're going to introduce is first train
- 16:50this once for all Network just a super
- 16:52Network that contain many different sub
- 16:54networks so during the editing phase we
- 16:57can use a smaller sub Network work
- 16:59running at lower cost for fast
- 17:02prototyping and finally when we are
- 17:04satisfied with the result we can run the
- 17:07a large sub Network only once and as a
- 17:10result you can see you can make her
- 17:12smile almost
- 17:15instantly or make her look younger we
- 17:18can uh Slide the second bar to the right
- 17:21change the hair color is almost instant
- 17:24okay by running this um using this once
- 17:27for all Technique we can derive many
- 17:29different uh sub networks very
- 17:37quickly okay and using this one4
- 17:40technique we can reduce the computation
- 17:41showing this yellow bar from 100% 78%
- 17:4550% 30% even 19% but you can hardly tell
- 17:50the quality qually the quality is still
- 17:52well maintained despite that we are
- 17:56reducing the max by up to five times
- 18:03uh this is another Technique we are
- 18:04going to introduce later in the later
- 18:06part of this lecture uh to accelerate
- 18:09image generation when we are doing the
- 18:11photo editing so here we are doing image
- 18:14in painting so what is inpainting and
- 18:15out painting inpainting is basically set
- 18:18saying that we have a a part of the
- 18:20image say the black part in rectangle we
- 18:24want to add a horse to that photograph
- 18:26of a horse on a grassland and usually
- 18:29stable diffusion takes like
- 18:321,855 Max costing 369 milliseconds and
- 18:37after the compression acceleration we
- 18:39can make it run only 514 gigamax at 95
- 18:44milliseconds similar on the right hand
- 18:47side to generate add a coconut tree to
- 18:50the image we can accelerate it by 4.8
- 18:54times since previous methods you need to
- 18:57generate all the pixels even though
- 18:59you're only editing a part of that but
- 19:02with the the S method we can do sparse
- 19:06update only update where you uh only
- 19:09update the pixels only generated pixels
- 19:12that needs to be updated and the the
- 19:14remaining pixels Remains the
- 19:18Same we are also going to introduce how
- 19:21to accelerate image Generation by
- 19:23parallel parallel processing so with
- 19:27four gpus okay can we generate even
- 19:30larger images even
- 19:33faster originally with one GPU it takes
- 19:36about 12 seconds right a natural thought
- 19:39would be if we have four gpus can we use
- 19:42like three three seconds four times
- 19:45faster unfortunately naively doing this
- 19:47we
- 19:48introduce um the artifact with a
- 19:51duplication using four gpus is actually
- 19:54generated for this uh distinct images
- 19:58the reason is that these four gpus
- 20:01didn't communicate with each other and
- 20:03as a result they each generated a
- 20:05independent
- 20:06image but communication will lead to the
- 20:09networking latency right and the way we
- 20:12deal with that is to overlap the
- 20:14computation with the networking such
- 20:16that um we can use a stale sta to
- 20:20communicate the stale um
- 20:23input uh to The Next Step we're going to
- 20:25learn that in detail in the later part
- 20:27of this lecture this is just for
- 20:28motivation and as a result we can
- 20:30generate um on the right column the
- 20:34image that is three times faster but the
- 20:36quality is very well
- 20:39maintained also people are interested
- 20:42customize image customization right uh
- 20:45create a personalized images based on
- 20:47user specified input and for example we
- 20:51want to generate something like Sunan
- 20:52writing a horse but it doesn't look like
- 20:54me writing a horse right it's very
- 20:57generic
- 20:59uh so we are going to introduce
- 21:00techniques like Fast composer which from
- 21:03function a TR free method that can blend
- 21:07different people together like here we
- 21:09can have Jensen Lady Gaga s
- 21:13together and this efficient image
- 21:15generation has a lots of applications
- 21:17this is just motivat why we need to
- 21:19learn these techniques like on the iPad
- 21:22we can uh you kind of draw a a picture
- 21:25and then after a while you see it takes
- 21:27several seconds we want it to be faster
- 21:30and to immediately render a very
- 21:31goodlooking image and we hope such
- 21:34method can run locally on the adedge
- 21:39device not only image generation but
- 21:41also 3D generation that's also very hot
- 21:44topic and require a lot of compute so
- 21:47this is diffusion models creating 3D
- 21:49objects from a reference image so the
- 21:52first rle is the reference image and the
- 21:56later rol is are the generated a 3D
- 21:58object based on just one
- 22:03image and this is using text to prompt
- 22:06using natural language as an input and
- 22:10create the 3D
- 22:14objects and also video generation that's
- 22:17adding another degree of Freedom uh they
- 22:19are pretty big uh it's even harder task
- 22:22requiring much much larger model size
- 22:26and this is from last year and this year
- 22:28uh Sor is there so it's pretty exciting
- 22:31just give a a text prompt a stylish
- 22:34woman walking down the Tokyo Street and
- 22:37then it's generating high resolution um
- 22:39Pho real realistic videos from natural
- 22:42language description but again this is
- 22:45consuming a huge amount of
- 22:47compute and we are going to learn
- 22:49techniques to bring down such Tech uh
- 22:51such
- 22:53computation and so far we talk a lot
- 22:56about 2D Vision right uh but in in
- 22:58reality 3D Vision is also very exciting
- 23:02like autonomous driving self driving
- 23:03cars they have not only camera sensors
- 23:06but also lar sensors that send those
- 23:09Point collect those Point cloud and to
- 23:12understand this 3D
- 23:13World unfortunately uh that require a
- 23:16lot of compute here is showing the whole
- 23:18trunk of workstation in the back of the
- 23:22seat you don't want to want that to hit
- 23:25the whole cars they uncomfortable so we
- 23:27want to make self-driving more efficient
- 23:29to fit everything in a smaller
- 23:33GPU so this is what we did Fast lighter
- 23:37net to accelerate 3D perception with
- 23:40algorithm and system code design uh
- 23:42previously it runs only five frames per
- 23:45second after the uh optimizations we are
- 23:48going to learn in this lecture about 3D
- 23:50Point Cloud acceleration it can run 47
- 23:53frames per
- 23:55second this is another group project
- 23:57we're going to introduce
- 23:59from Jan from our group he's not going
- 24:01to be a assistant professor UC ucst very
- 24:04soon
- 24:06so this project is called BB fusion um
- 24:10that combine not only just one sensor
- 24:12but also combining multiple sensors very
- 24:14efficiently here we are showing six
- 24:16cameras in the front in the back front
- 24:19left front right six cameras and also
- 24:22one liar sensor on the top and we want
- 24:25to combine these features to fuse
- 24:28multiple sensors uh to do the
- 24:30segmentation and also detection
- 24:32detection means finding the bonding box
- 24:343D bonding box of the cars the
- 24:37pedestrians and everything and on the
- 24:40right hand side is showing the B map
- 24:42segmentation BV means the bird eye view
- 24:45there's no map it's completely mapless
- 24:48mapless in the new CD you can run such
- 24:50algorithms to find the lanes find the
- 24:53driveable area find The Pedestrian
- 24:56walkway Etc and top right corner is
- 24:59showing that this algorithm is runable
- 25:01on Json or which is a mobile GPU for
- 25:04self-driving
- 25:08tasks okay so um that is the first part
- 25:12about the amazing progress of computer
- 25:14vision and the impl implication for
- 25:18computing right why do we need efficient
- 25:21Computing um and these applications
- 25:23although amazing they are they come at a
- 25:26high cost of computational RIS resource
- 25:29and now let's switch gear to the second
- 25:31part and talk about the advancement of
- 25:33natural language
- 25:35processing uh so this is the second year
- 25:38when after
- 25:40TBT it is very computationally heavy uh
- 25:43actually when chpt just came to the
- 25:46world usually it's at a capacity and
- 25:49prevent user from um have a cap of 50
- 25:52messages every 3 hours Etc we are
- 25:55experiencing exceptionally high demand
- 25:57please h TI and we we work scating our
- 26:00systems those are the bottom
- 26:02NE but they're really helping us for
- 26:05example co-pilot can make very
- 26:07meaningful coding suggestions based on
- 26:09the context we just give it some prompt
- 26:12some comment and then it's going to
- 26:15automatically generate the code uh very
- 26:18quickly and also neural machine
- 26:21translation to bre Bridge the language
- 26:24barrier using techniques so we are going
- 26:27to introduce Tech technique to bring
- 26:29down the model size and also reduce
- 26:33maintain the blue store which is the
- 26:35quality so here we are reducing the
- 26:38model size from 17
- 26:40176 megabytes to only 9 megabytes if it
- 26:44is only like 9 megabyte is very easy to
- 26:47fit on your
- 26:51phone and recently with advancement of
- 26:54large language model new capabilities
- 26:57begins to emerge for example zero short
- 26:59learning and also F short learning so
- 27:02what is zero short learning it can the
- 27:04model can predict the answer given only
- 27:07a description of the task there's no
- 27:10gradient there's no training given a new
- 27:12task you just tell it what is the task
- 27:14and then it's going to give you the
- 27:16answer say translate from English to
- 27:19French we're not we are not training a
- 27:21separate model or English to French or
- 27:23English to Chinese just the same model
- 27:26using different prompt to describe the
- 27:28task and zero shot means there's no
- 27:31example and directly translate from
- 27:33cheese to to French right and also the
- 27:37large language model have the capability
- 27:38called f short learning so um we just
- 27:42see a few example we just give it fed
- 27:44with a few examples of the
- 27:46task as part of the prompt engineering
- 27:49say translate to English to French and
- 27:51then we are going to give it three
- 27:53examples and then finally give the
- 27:56result for to translate cheese right so
- 27:59again there's no gradient there's no
- 28:00fine-tuning I just use design The Prompt
- 28:04SM smartly the model is capable of doing
- 28:07such zero short or F short
- 28:10learning this seems very exciting
- 28:12capability however like we can expect
- 28:15this comes at a very high cost so this
- 28:18is the few shot one shot or zero shot
- 28:21learning accuracy
- 28:24so to achieve higher accuracy the model
- 28:26size has to grow very fast um from 13 B
- 28:29to 175 bilon parameter um every small
- 28:33percentage of accuracy Improvement comes
- 28:36at a high cost of much larger model size
- 28:39and requiring more GPU resources to
- 28:42serve such models you see that the curve
- 28:45is very flat means every small
- 28:47percentage of accuracy Improvement can
- 28:49comes at a very high
- 28:54cost um large language model also have
- 28:57another very interesting emergent um
- 29:00capability U which is called Chain of
- 29:03Thought So let's see what is Chain of
- 29:05Thought standard prompting say Roger
- 29:08have five tennis balls he but buys two
- 29:11more cans of tennis balls each can has
- 29:15three tennis balls how many does he have
- 29:17it's
- 29:1811 since five time plus three * three 2
- 29:22* 3 is 11 and then if you ask another
- 29:25question given the same prompt the
- 29:28cafeteria has 23 IPOs if they use a 20
- 29:31to make lunch and B bought six more how
- 29:33many IPOs they
- 29:36have what should be the what should be
- 29:38the correct
- 29:44answer n right but the answer given by
- 29:48at that moment given by chb is 27
- 29:51obviously that's wrong right um what if
- 29:54we prompt it in another way say
- 29:59Roger given the same question of q1 the
- 30:01answer in blue showing Roger started
- 30:03with five balls two kinds of three
- 30:05tennis balls um two three tennis balls
- 30:09each is six tennis balls and five plus 6
- 30:12equals to 11 it describe the entire
- 30:15thinking process how the answer 11 is
- 30:19arrived and prompt in this way to show
- 30:22not only the result but also the Chain
- 30:24of Thought thinking process and given
- 30:27another Apple question
- 30:28um the answer the model can output a new
- 30:30answer here cafeteria at the 23 iOS
- 30:33originally they use a 20 uh to make
- 30:36lunch so they had 23 minus 20 that's
- 30:39three okay step by step thinking step by
- 30:42step and then they bought six more Apple
- 30:45so they now have three plus six that's
- 30:47equal to nine the answer is n now I can
- 30:49understand it correctly so we are going
- 30:51to introduce some of such prompt
- 30:53engineering techniques so that it will
- 30:55be very practical you can interact with
- 30:58these large language models more
- 31:03effectively again um such
- 31:06um um accuracy comes at a cost of very
- 31:10high computational cost so this is the
- 31:13accuracy um with um the GSM 8K middle
- 31:18school math world problems and then
- 31:21versus the access is the model size in
- 31:24order to achieve high accuracy the model
- 31:26has to be pretty big
- 31:30like 175 billion parameters or 500 even
- 31:33have a trillion
- 31:35parameter which is leading to the figure
- 31:38we showed initially the model is ring
- 31:40much faster than the hardware and on the
- 31:43right hand side is a a server in the
- 31:45basement uh from our lab to train uh
- 31:49these models we actually uh purchased it
- 31:52about two two two and a half years ago
- 31:54at that time we have 200 gabit uh infin
- 31:58band but the moment we purchase it the
- 32:00moment it becomes outdated last year I
- 32:02want to give the lecture the S was set
- 32:04of R was 400 now it's 800 so this area
- 32:08is just moving so fast um so it's very
- 32:11timely to learn these techniques to
- 32:14catch up the
- 32:15wave we also designed a couple of
- 32:17techniques to reduce the computation and
- 32:20by finding the redundancy in actal
- 32:22languages actually in actal language
- 32:24there's a lot of redundancy say give the
- 32:28sentence um as a visual treat the film
- 32:31is almost perfect that's the example on
- 32:34the left we can trim it the the task is
- 32:37to classify the sentiment so we can trim
- 32:40uh the sentence to be as trat feel
- 32:43perfect or even triming to F perfect
- 32:46they can still classify this is positive
- 32:49this is pos positive
- 32:53sentiment so showing that human language
- 32:56has a lot of uh redundant we can take
- 32:58advantage of that redundancy and do a
- 33:01sparse attention so this is a project we
- 33:03did about four year three four years ago
- 33:06called a spatter spars attention we can
- 33:09actually R remove those redundant uh
- 33:12tokens that doesn't have heav attend to
- 33:14other tokens like in this example I bet
- 33:18the video game is a lot more fun than
- 33:21the film video attend very heavily to
- 33:24game um but the word uh I and
- 33:29the doesn't attend to any words very
- 33:32heav so we can safe safely remove those
- 33:36tokens so those are the some of the
- 33:39technique we are going to dive deeper in
- 33:41later part of this lecture today we are
- 33:43just giving overview for you to get a
- 33:45taste what we are going to
- 33:49learn okay so um these L language models
- 33:53are pretty exciting how can we deploy
- 33:55them on the edge
- 33:58large language model on the edge would
- 33:59be super useful we can run co-pilot
- 34:02services like code completion like
- 34:05office or game chat everything locally
- 34:08on the laptops in the cars in the robots
- 34:11we don't have to worry about lency
- 34:13networking Wi-Fi or sending our private
- 34:17data to the cloud right
- 34:20um so deploying this large language
- 34:23model locally on the edge is super
- 34:24demanding and in this uh this class we
- 34:29have two lecture and two Labs actually
- 34:31dedicated to how to deploy a large
- 34:33language model locally on the Edge by
- 34:36using large language model quantization
- 34:38techniques so that you can run the seven
- 34:41billion Prim or even 13 billion primer
- 34:43model locally on your laptop so this is
- 34:45a demo showing this is a pretty updated
- 34:48MacBook it's only MacBook with M1 chip
- 34:52not only not even M3 um chip but still
- 34:56it can run reasonably fast rather C code
- 35:00uh given the prompt to sort an
- 35:03array and we are going to introduce
- 35:05techniques to quantize these large
- 35:07language models from 16 bit to only four
- 35:10bit okay gradually from 16 bit to 8 bit
- 35:14using smooth Quant and to even four bit
- 35:16using awq activation where weight only
- 35:19quantisation actually the smooth Quant
- 35:22uh paper actually comes from this
- 35:23lecture two years ago as one of the
- 35:25course project with B so maybe in this
- 35:28year some of you may come up with even
- 35:31exciting even more exciting projects
- 35:33that may lead to even publication in the
- 35:35future and the key idea is that we find
- 35:38lots of outliers very big values on the
- 35:41right top right corner in the
- 35:43activations and we want to make it
- 35:45smooth make it smooth since metrix
- 35:48multiplication is linear we can scale
- 35:50the weight and gradient so that they can
- 35:52be equalized and smooth such that it
- 35:55will be much easier to quantize and
- 35:56utilize the full dynamic range uh we are
- 36:00going to implement uh such algorithm in
- 36:03uh lab four in lab four and also we are
- 36:06going to introduce the inference library
- 36:09and efficient inference library in live
- 36:12five so that you can Implement both the
- 36:15algorithm the quantization algorithm and
- 36:18also the um inference library on your
- 36:21own which actually require pretty deep
- 36:24understanding of how computer systems
- 36:25work you must be able to write C program
- 36:29um not just python but c lowlevel c
- 36:31program um very comfortably manipulating
- 36:35uh pointers um single instruction
- 36:38multiple data familiar with cach how
- 36:40cache works with is locality with is
- 36:44multi threading we're are going to
- 36:46implement those techniques using multi
- 36:48core processors using CMD techniques um
- 36:52very and also register LEL paradism so
- 36:55that's why we enforce uh this
- 36:58prerequisites so that you can complete
- 37:01the
- 37:03labs so here we have a demo called tiny
- 37:06chat which is um um fast inference
- 37:09Library by H and Sean we can deploy a
- 37:12large language model 7 billion parameter
- 37:15large language model on Jon or Nano
- 37:17which is very small um mobile GPU we um
- 37:213D printed uh this tiny cheat computer
- 37:25uh to interact with uh the input in real
- 37:28time and this is showing we can actually
- 37:31run large language models on resource
- 37:34constrained H GPU so this is a Jon Orin
- 37:38uh we can t use tiny chat to run large
- 37:42language model on this Jon or very
- 37:46smoothly even 13 bitting parameter model
- 37:51and this is comparing with quantization
- 37:53without quantization it'll be much
- 37:55faster if you can use the activation
- 37:57quantization the awq so on the left hand
- 38:00side is showing running this large
- 38:03language model without quantization is
- 38:06slow and on the right is showing width
- 38:08the 4bit quantization is faster from 40
- 38:1250 tokens per second all the way to 166
- 38:16tokens per
- 38:19second and running large language model
- 38:22on the edge device can have a lot of
- 38:25applications so for example uh
- 38:28IO intelligence here I can do
- 38:30summarization to summarize some of your
- 38:32emails smart reply or emails and also
- 38:36writing tools to help you correct the
- 38:38the sentiment to make it more
- 38:40professional
- 38:42Etc uh so far we talk about Vision
- 38:45language and now let's jump into
- 38:48multimodels so what is multimodel so by
- 38:52we combine image language audio action
- 38:56different modalities
- 38:58into project them into the same space
- 39:01and understand and process all these
- 39:03different modalities of input and output
- 39:06different modalities can we take image
- 39:09language audio video action in and then
- 39:12output potentially not only text but
- 39:14also image video even actions okay so
- 39:17that's a multi model problem and the
- 39:20challenge here is to how do we align the
- 39:23representation the embedding from
- 39:25different input modalities
- 39:27so one example is uh this Ro called lava
- 39:30who can understand the images and take
- 39:34language prompt with uh the input image
- 39:36and output the description in language
- 39:39so here do you know who Dre this
- 39:41painting and lava is going to say uh the
- 39:44details about this
- 39:47monalisa and uh here is showing with
- 39:49quantization techniques we can also
- 39:52quantize them to only four it without
- 39:54losing accuracy like in this example um
- 39:58this is input image and some text
- 40:02description together with that with
- 40:04naive run to nearest quantization it's
- 40:07going to say something weird like there
- 40:09are small pictures of the earth and
- 40:11other planets placed on top of the food
- 40:14but with aw awq quation it says the meme
- 40:17in the image is light-hearted and
- 40:19humorous take on the concept of looking
- 40:21and pictures of the Earth from the space
- 40:24right so um with proper transition
- 40:27techniques we can not only accelerate
- 40:30the the large language models but also
- 40:33the visual language
- 40:35models later we designed uh this V the
- 40:38visual language model we are also going
- 40:40to cover how do we uh train and run
- 40:43difference on visual language models so
- 40:45the key idea is to produce the visual
- 40:49tokens okay so not only language can be
- 40:52tokenized images can also be tokenized
- 40:54and how do we tokenize images is by
- 40:57using the visual Transformers we're
- 40:59going to have a a dedicated lecture just
- 41:02to introduce what is a visual
- 41:03Transformer and also how do we design
- 41:06efficient Vision Transformers so with a
- 41:08vision Transformer we can project the
- 41:10input image into several tokens and
- 41:13contaminate the image token with the
- 41:15text token so basically uh we can expand
- 41:19um this idea to other modalities as well
- 41:22so people are realizing everything can
- 41:24be tokenized everything is token token
- 41:27in toen
- 41:28out and we find in this uh we in this
- 41:31lecture we are going to learn those
- 41:32training recipes like data and training
- 41:35recipe actually matters more than the
- 41:38architecture itself and why visual
- 41:40instruction tuning is not enough and why
- 41:43we need visual language pre- trining and
- 41:46why image text pair is not enough but we
- 41:49we need to have inter with image and
- 41:52text and how do we blend them how do we
- 41:54set the proportion of how much uh of
- 41:57each and visual in we also going to talk
- 42:00about visual in context learning um
- 42:03which is pretty exciting capability and
- 42:06how to we expand it from images to
- 42:08videos which require much longer context
- 42:11length since each frame will take about
- 42:15200 tokens if you have hourly long video
- 42:18you want to understand like a film that
- 42:20require very long context so in this
- 42:22lecture we are going to introduce um
- 42:24long context techniques how do we handle
- 42:27um um the linearly growing memory if na
- 42:33rather than rather than naively using
- 42:35the linearly growing KB
- 42:37cache so here are some of the
- 42:39capabilities a modern visual language
- 42:41model such as V can achieve uh given
- 42:44this video please tell me what happens
- 42:46in the video and vaa can say the video
- 42:49shows a circle game where player scores
- 42:51a goal and the crow cheers as the player
- 42:54celebrates it with his teammates okay so
- 42:57we have a demo link here at v.m it.edu
- 43:00Fore to uh to play with it given this
- 43:04video according to the video what will
- 43:06happen likely to happen next vaa says
- 43:09the video video shows a blender with
- 43:12strawberries inside and it's likely that
- 43:14blender will be turned on to blend the
- 43:17strawberries uh into a smoothie so it
- 43:19has the World Knowledge knowing that
- 43:22given such a uh few frames what's going
- 43:25to happen in the next that's that's also
- 43:27called A World model OKAY World model
- 43:29with the World Knowledge understand how
- 43:32the world Works given some frames
- 43:34predict what's going to happen next
- 43:36similar in this case what will likely
- 43:39happen next we says next step in the
- 43:41video is likely um involve the person
- 43:44grinding the spices spices in the
- 43:49motar a very interesting capability is
- 43:52the visual in context learning okay so
- 43:54what is in context learning we're not
- 43:57deciding what is the task but we're just
- 43:59showing some examples so here the prompt
- 44:02is image Boston image Toronto and
- 44:07image and it's going to say San
- 44:09Francisco right we're not explicitly
- 44:11telling that please tell me what is the
- 44:14city in this image but we're just
- 44:17showing some of the examples and then
- 44:19the model is going to understand okay
- 44:21the task is to tell uh tell the user
- 44:24where is the C okay so that is enabled
- 44:27by using the intered image text training
- 44:31so data blending this training recipe
- 44:33matters a lot in the area of FL language
- 44:35models and we are going to introduce
- 44:37visual language model in this
- 44:40lecture okay next one about World
- 44:43Knowledge so anybody can guess where
- 44:46where is this given this
- 44:50picture okay oh very nice very nice Spa
- 44:53pay right it require a World Knowledge
- 44:56you have seen lot of play is to
- 44:57understand that but TR the visual
- 44:59language model can also um tell where
- 45:04specific image is actually when I was TR
- 45:06traveling in Europe earlier this year
- 45:08but I clear I took a random picture and
- 45:11to my surprise it can tell me this is
- 45:14actually um a particular place described
- 45:17it very very accurately so I was quite
- 45:20Amazed by the World Knowledge they may
- 45:22have so feel free to play with it or
- 45:24even implement it in one of your
- 45:26homework
- 45:27we're also going to learn this technique
- 45:30can we bring this large language model
- 45:33Vision language model um locally to our
- 45:36laptops right so that we can play with
- 45:39it interacted with it have full control
- 45:43uh with it compress all the world of
- 45:45knowledge into a laptop so here is what
- 45:48we can do um for example this rep we can
- 45:51describe the painting in details it's
- 45:53running locally on a uh laptop so our
- 45:57lab five in lab five we are implementing
- 46:01a language model alone so we can also
- 46:05implement the visual language model as
- 46:06extra bonus extra credits for extra
- 46:09credits so feel free to talk to me if
- 46:10you want to do the visual language model
- 46:12for lab five something like this
- 46:17demo okay so another modality would be
- 46:19the action right we can use um the
- 46:23visual language model also to Output
- 46:25another other modalities like acttion
- 46:27right so here the instruction is bring
- 46:29me the rice chips from the Jer okay and
- 46:33the current step is opening the Jer it
- 46:35helps do the planning since using the
- 46:39the large language model contains a lot
- 46:40of word model it understand in order to
- 46:44bring the chip you first need to open
- 46:45the Jer and then take the right chip out
- 46:48of the Jer and then place it um and then
- 46:51close the J Etc right um but
- 46:55unfortunately this is by 4X uh speed
- 46:59right is Rise only three Herz due to the
- 47:03computational cost networking latency
- 47:06because the the images are transmitted
- 47:08over internet to a workstation and then
- 47:12transmitted back the final result okay
- 47:15so this is the rt1 from Google is also
- 47:19already pretty uh amazing but highlights
- 47:23how important it is to uh learn this
- 47:26efficient depl Computing techniques like
- 47:28this
- 47:29lecture similarly deep learning for
- 47:32games um our for go beating so do is
- 47:36taking
- 47:371,920 CPUs 280 gpus $3,000 electric bill
- 47:42per game that's super expensive or the
- 47:46arpha fold require 16 TP v3s which is
- 47:51128 tp3 course um train for a few weeks
- 47:56so all these amazing um capabilities
- 48:00comes at a high
- 48:02cost okay so let's switch gear from
- 48:06visual language to multimodality now we
- 48:09talk about these three pillars algorithm
- 48:12hardware and data in particular we want
- 48:13to talk highlight the uh advancement of
- 48:16Hardware that is the driving force for
- 48:19this C Breen of of efficient deer
- 48:25Computing so this lecture we are going
- 48:27to introduce some of the hardware um
- 48:30Hardware system techniques as well and
- 48:33the philosophy is that we want to do the
- 48:35whole full stack full stack design okay
- 48:37on the software part full stack design
- 48:40from algorithm part the demand of
- 48:42computing to the system part which is
- 48:44the supply of computing we want to close
- 48:46the gap by reducing the demand and
- 48:48increasing the supply okay so the recent
- 48:52trend of Modern Hardware design is that
- 48:55uh the microprocessor frequency is
- 48:58plateaued um the single core performance
- 49:01also plateaued more SL is scaling down
- 49:04but the number of cores is increasing
- 49:07and the the number of transistors is
- 49:08also increasing uh so parallel Computing
- 49:12specialized uh Computing is the new
- 49:15trend in this
- 49:16context and also Precision matters
- 49:20people uh the communi is keep advancing
- 49:23the low Precision arithmetic from
- 49:25floating point 32 floating Point 16 to
- 49:29integer 8 8 bit integer to fp8 8 bit
- 49:33floating Point um and also even to fp4
- 49:37four bit floating point which is to
- 49:40appear in the our latest black well uh
- 49:43immedia gpus right so the Precision is
- 49:46keep decreasing in this lecture we are
- 49:48going to learn all the details about
- 49:50what is a floating Point what is fp8 uh
- 49:53what is um E5 M2 okay what is the
- 49:57mantisa so what is fp4 why Blackwell is
- 50:02so amazing it can have such high Peak
- 50:05Performance in uh the quantization part
- 50:07of this
- 50:08lecture and on the right hand side is
- 50:10showing the single chip inference
- 50:13performance 300 times in eight years
- 50:16this is from the slide from my PhD
- 50:18adviser Professor Bill Ali is showing
- 50:21that from scaler FP 32 to
- 50:24fp6 um to H hmma tensor core uh in
- 50:30v00 the intake TS increased to 125 and
- 50:35again it doubled by using int8 int8 imma
- 50:39and course to using structur spity 24
- 50:43sparcity in
- 50:44a100 like 300 performance improv in
- 50:47eight years that's like pretty amazing
- 50:51Improvement and in this along this line
- 50:53of innovation software is actually
- 50:56playing a very critical role um in
- 50:58especially the advanced technology node
- 51:01so this is from 65 nanometer 28
- 51:04nanometer 7 nanometer 5 nanometer you
- 51:07can see the purple part which is a
- 51:09software component is getting a very big
- 51:12chunk of the total Advanced design cost
- 51:17and that's always also creating the mold
- 51:20for a good cicon to uh be able to
- 51:23attract a lot of good customers it has
- 51:25to be programmable has to be easy to use
- 51:27especially the large language model this
- 51:29generative AI the models are changing
- 51:32iterating very fast changing every day
- 51:35um how to make it flexible is super
- 51:37important so in this lecture we have a
- 51:40lot of focus on the software part of
- 51:43efficient deep learning Computing uh we
- 51:45don't design new Hardware in this
- 51:46lecture but how to well utilize existing
- 51:50hardware and what is the um implication
- 51:52to design new hardware so here we show
- 51:55from 20 16 to
- 51:582024 uh the advancement of gpus uh we
- 52:01put them into four different uh
- 52:04benchmarks you know the dance fp16
- 52:07performance measured in tal operations
- 52:10per second we're going to learn in the
- 52:11next lecture what is T Ops per second
- 52:13what does it mean what is Dan what is
- 52:16sparse what is fp16 so we're going to
- 52:19learn dance and sparse in the pring
- 52:20lecture from uh next Thursday and we are
- 52:24going to learn uh quantization precision
- 52:26a week after so from p00 to b00 is
- 52:31actually 100 times 100 times more um TS
- 52:36per second in particular from h100 to B
- 52:40100 even doubled uh the the TS per
- 52:44second which is super amazing and even
- 52:46didn't consider the fp4 if we add fp4 it
- 52:48will be way higher in this chart because
- 52:51it's showing just
- 52:53fp16 and here is the memory bandwidth
- 52:56computation is cheap memory is super
- 52:58expensive uh from a100 to h100 it almost
- 53:03doubled and then more than doubled from
- 53:05h100 to b00 with respect to the memory
- 53:08bandwidth and we are going to introduce
- 53:10these Concepts like why data movement is
- 53:13expensive um why energy is dominated by
- 53:17moving data and here is showing the
- 53:20power unfortunately the power is also
- 53:22growing pretty fast from a1004 400 watts
- 53:26to 700 watts in h100 and B 100 if you
- 53:31have eight um gpus in a node you can
- 53:35calculate how much power you have so
- 53:38eight node all the gpus will take maybe
- 53:404,000 5,000 watts and all all together
- 53:44the whole node we have 10,000 Watts you
- 53:46have to prepare 10,000 Watts just for a
- 53:49single node and as a result you the
- 53:52space will not be the constraint but the
- 53:54power power supply the ener energ in
- 53:57cooling is the new Bott
- 54:00neck previously when we are having the
- 54:03a6000 gpus so we can have two two cables
- 54:06serve the entire rack but now two cables
- 54:09can only serve four node of a100 and
- 54:11only two node of h100 so power
- 54:15efficiency is super
- 54:16important and in this lecture later part
- 54:20of this lecture we are going to organize
- 54:21a lab tour so we'll show you down to the
- 54:24basement uh to our uh server so that you
- 54:28can see how these models are trained
- 54:30with what is a like a mini data center
- 54:32what does it look like what is the code
- 54:34I hot a the rack the switch Infinity
- 54:37band The the cables and modern GPU what
- 54:40does it look
- 54:43like and next is the memory so
- 54:45unfortunately memory is growing at a
- 54:47slower speed compared with a compute um
- 54:50so current what we have in our lab is 80
- 54:54Gab a100 gpus h100 pretty much stays the
- 54:58same and b00 give a pretty exciting leap
- 55:01to almost 200 192 gigabytes so this is
- 55:05for the cloud for training and
- 55:08serving uh so let's also look at the
- 55:12edge uh for example on the phone we want
- 55:14to have those low power uh chips like
- 55:17here is only 10 watts 10 watts s855 all
- 55:21the way to I8 gen one and I8 Gen 2 is
- 55:25also there but we didn't find the public
- 55:27uh numbers so we didn't put it there the
- 55:29performance is also increasing pretty
- 55:31fast like couple of um uh like almost
- 55:35like half a 100 like 50 TS per second
- 55:40and memory um is roughly eight 16
- 55:44gigabytes for Qualcomm
- 55:46DSP Apple also offers this apple neural
- 55:49engine energy efficient and high
- 55:51throughput for machine learning
- 55:53applications uh similarly with uh the
- 55:56qualcom DSP is 35 TS per second however
- 56:01we should not just look at the Peak
- 56:03Performance like 35 or 52 can you make a
- 56:07conclusion that apple or qualcom which
- 56:10one is faster actually cannot because
- 56:12it's a multiple Dimension um the Peak
- 56:16Performance doesn't indicate doesn't
- 56:19directly translate to measured uh speed
- 56:23there's so many factors we are going to
- 56:24learn in this lecture like the
- 56:26activation data movement weight data
- 56:28movement memory bandwidth utilization
- 56:31all matters a
- 56:35lot and also there's the mobile GPU
- 56:38people very widely used in the cars for
- 56:41example this is the Json a which is used
- 56:44in many EVS um delivers couple of
- 56:47hundred TS per second consume about 60
- 56:50watts 64 gigabytes of memory and there's
- 56:53going to be a new um Thor coming coming
- 56:56up which is the next generation of ajx
- 57:01or and also there's microcontrollers for
- 57:04many iot devices like the your smart
- 57:07home camera might have one of these um
- 57:11small devices arm devices just a couple
- 57:14of hundred mwatts by the memory is also
- 57:16super small just a a few hundred
- 57:19kilobytes of s SRAM maybe a few
- 57:21megabytes of uh um maybe Dam or even no
- 57:25Dam some M controllers complet has
- 57:28completely no Dam we have to manage the
- 57:30memory super
- 57:33carefully so from cloud AI to mobile AI
- 57:36to timing AI there's actually a big gap
- 57:39between the model memory uh that you
- 57:43have tens of gigabytes versus couple of
- 57:46hundred kilobytes so um Ed AI device do
- 57:50have a huge gap to Cloud processors and
- 57:53what we learn in this lecture is going
- 57:55to try to close the Gap by using smaller
- 57:59models okay so those are the three
- 58:02pillars the hardware Theta and the
- 58:04algorithm and our goal is rather than
- 58:07using a lot of compute emitting a lot of
- 58:09carbon using many Engineers TR a lot of
- 58:12data tin ml's goal will be how to use
- 58:15less computation less carbon few
- 58:17Engineers automated TR L
- 58:21data so this is the course overview let
- 58:24me use the remaining 10 minutes to cover
- 58:26some of the
- 58:28logistics so the lecture is every uh
- 58:31Tuesday and Thursday from uh 3 uh 35 to
- 58:364:55 in this classroom and we will
- 58:39record uh the
- 58:41classes um the office hour will be after
- 58:44directly after uh this lecture on every
- 58:48uh Thursday from 5 to uh 6:00 p.m. uh
- 58:52the room number is 38344 that's where we
- 58:55our office is located at feel free to
- 58:58ask questions during the office hour or
- 59:01go to the PIAA make sure you sign up on
- 59:03Pasa and also we use canas like other
- 59:06classes to submit the
- 59:10homework and you can email us efficient
- 59:13ml- staff at mit.edu if you have any uh
- 59:18uh external inquiries we also going to
- 59:21create a maing list you make sure you
- 59:23take care of the maing list so everybody
- 59:25is wel come to join that maing list so
- 59:28that we can broadcast some of the news
- 59:31project ideas some company even want to
- 59:34uh have those internship openings or
- 59:37full-time job openings we can
- 59:39disseminate uh through that meeting list
- 59:42so um stay tuned regene will send a
- 59:45notification for the maing list to sign
- 59:48up and the prerequisites so this year we
- 59:51enforce two prerequisites one is the
- 59:546191 which is comput ation structures
- 59:57since the LA five and La four require a
- 1:00:01deep understanding of how computer
- 1:00:02systems work like what is cach what is
- 1:00:05locality what is paradism five stage
- 1:00:09pipeline um page table since we're going
- 1:00:12to talk about the page attention um
- 1:00:15speculation uh we're going to talk about
- 1:00:17speculative decoding and um most
- 1:00:21importantly you have to be familiar with
- 1:00:23comfortable with how to write C programs
- 1:00:26um since the last the lab will be
- 1:00:28written in C program and also mod
- 1:00:31threading techniques the second
- 1:00:33prerequisite is
- 1:00:356.3 90 the introduction to machine
- 1:00:37learning
- 1:00:38class make sure you are familiar with
- 1:00:40how to use pytorch um since all the labs
- 1:00:44will be using
- 1:00:49pytorch so we are going to cover roughly
- 1:00:52three sections first section is about
- 1:00:54efficient inference technique
- 1:00:56with pruning reducing the number of
- 1:00:58parameters fantization reducing the
- 1:01:01Precision for each parameter and also
- 1:01:03neural architecture search how do we
- 1:01:05design efficient neural network
- 1:01:08architecture and also distillation how
- 1:01:10to use a teacher model to distill a
- 1:01:12smaller model we're also going to talk
- 1:01:14about efficient training technique how
- 1:01:17to do grading compression on device
- 1:01:19training and also fight rate learning
- 1:01:21and parallelization including data
- 1:01:23parallel model model parallel Pine
- 1:01:26parallel and also sequence parallel for
- 1:01:28large language
- 1:01:30model we're also going to introduce
- 1:01:32application specific optimizations the
- 1:01:34most important one of course is
- 1:01:36Transformers and large language models
- 1:01:38we're going to introduce the entire
- 1:01:40architecture how Transformer work from
- 1:01:44um the position encoding all the way to
- 1:01:46what is qkv what is attention what is um
- 1:01:50different uh contact extension
- 1:01:52techniques with this rope Etc or also
- 1:01:56going to introduce diffusion models um
- 1:01:58um how how diffusion model Works video
- 1:02:01understanding Point Cloud understanding
- 1:02:04Etc so a key concept is system on
- 1:02:08algorithm code design which is covering
- 1:02:10actually combining a lot of knowledge we
- 1:02:12learn from the e side the Cs side and
- 1:02:15also the a plusd side so hopefully it
- 1:02:18will be a class you can take um in later
- 1:02:21year of your undergrad or maybe in the
- 1:02:23first couple of years of your PhD to do
- 1:02:26interdisciplinary research on the E side
- 1:02:30um the
- 1:02:316191 or micro computer project lab
- 1:02:35Hardware architecture for deep learning
- 1:02:36those are very related classes csite
- 1:02:39computer architecture software
- 1:02:41performance engineering mobile s sensor
- 1:02:43Computing a plusd side intro to machine
- 1:02:46learning deep learning advances in
- 1:02:48computer vision natural language
- 1:02:49processing all related
- 1:02:51classes we will have five Labs um over
- 1:02:56the course of the semester we are going
- 1:02:57to use Google collab for the first four
- 1:03:01Labs so you don't have to worry about
- 1:03:03Hardware uh first one is about pruning
- 1:03:06second one about quation uh third one is
- 1:03:08about neural architecture search fourth
- 1:03:11is about large language model
- 1:03:12compression and the fifth one does
- 1:03:15require Hardware which is your laptop
- 1:03:18hopefully you all have a laptop um that
- 1:03:20is more than 8 gigabyte of memory make
- 1:03:23sure you have more than 8 gigabytes of
- 1:03:25memory and available storage should be
- 1:03:27at least five
- 1:03:29gigabytes so unfortunately we don't have
- 1:03:31the budget to BU everyone laptop but I'm
- 1:03:34assuming everyone already have a laptop
- 1:03:37either x86 or arm um either Mac Linux or
- 1:03:42Windows we have carefully debugged so
- 1:03:44that it can fit across different
- 1:03:46platforms and JY will be the TA if you
- 1:03:48have any questions related to live five
- 1:03:51and and uh
- 1:03:54Hardware so greeing we will have five
- 1:03:57lives each one is
- 1:03:5815% uh followed up followed by a final
- 1:04:02project 25% is open-ended but we'll give
- 1:04:05you some project ideas to start with it
- 1:04:08is not in a group of three to four uh
- 1:04:10students consisting of a proposal um a
- 1:04:14presentation poster presentation and Al
- 1:04:17also a final report just
- 1:04:1920% uh we also give uh four 4% for
- 1:04:23participation bonus and all the
- 1:04:26assignments are due on 11:59 p.m. on the
- 1:04:28due
- 1:04:29date um and the late policy uh is that
- 1:04:33we have six total um late days without
- 1:04:37penalty for the entire semester uh given
- 1:04:39the large number of enrollment we don't
- 1:04:41have any exception for the six L
- 1:04:45days and the allow days are counted by
- 1:04:48day each new day will um the panalty
- 1:04:51will be 50 20% 50% so homework is worth
- 1:04:55zero credit two days
- 1:04:58later so prerequisite all this is quite
- 1:05:00important so student who don't fulfill
- 1:05:02the prerequisite will be D registered in
- 1:05:05the second week of the class so if you
- 1:05:07believe you have equivalent prior
- 1:05:10experience including both computer
- 1:05:12architecture and also the machine
- 1:05:14learning please make sure to send us the
- 1:05:18petition on this form make sure you take
- 1:05:20a photo of this form if you don't meet
- 1:05:22the prerequisites to submit your
- 1:05:24petition form
- 1:05:26by this Friday
- 1:05:2811:49 p.m. okay and then our TA will
- 1:05:32review this petitions and student who do
- 1:05:36doesn't fulfill the prerequisite will be
- 1:05:38D register next
- 1:05:40week okay so these are the uh five labs
- 1:05:44and also together with the uh final
- 1:05:47projects we did last year it's pretty
- 1:05:50exciting so hopefully you can learn a
- 1:05:52lot about this amazing um efficient Tech
- 1:05:55Tech and we'll see you next Tuesday
About this transcript
This page contains the full transcript of EfficientML.ai Lecture 1 - Introduction (MIT 6.5940, Fall 2024, Zoom Recording) by MIT HAN Lab, generated from the public captions YouTube serves with the video. The transcript has 9,527 words across 1,416 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.