EfficientML.ai Lecture 2 - Basics of Neural Networks (MIT 6.5940, Fall 2024, Zoom recording) — Transcript
Full transcript
- 0:00good afternoon everyone let's get
- 0:01started with lecture two of efficient
- 0:04ammo. so today we are going to learn
- 0:06about basics of neuron
- 0:09networks as as we have seen the last
- 0:12lecture there's a big gap between the
- 0:14supply and gr and demand of AI Computing
- 0:17where the demand of computation is
- 0:18growing really fast as indicated by the
- 0:21model size of recent large language
- 0:24models and mor slow is roughly doubling
- 0:27every two years but uh the deep learning
- 0:30models is getting more than four times
- 0:32larger every two years so there's a
- 0:34going to be a gap which motivated us to
- 0:37bridge this Gap with model compression
- 0:40and acceleration techniques so in the
- 0:42last lecture just do a quick recap we
- 0:45went through several examples from
- 0:48Vision to language to visual language
- 0:50model and also um um and also the uh
- 0:54multimodality models uh to give a
- 0:56overview about the latest uh advancement
- 1:00of AI Computing and how much the amount
- 1:03of computing they need to motivate us to
- 1:05build these efficient algorithms and
- 1:07systems to accelerate efficient air
- 1:10Computing so today we are going to uh
- 1:13start with those details technical
- 1:16details to First review the ter
- 1:19terminology of neuron networks such as
- 1:21neurons synapsis act activation feature
- 1:25weight parameter
- 1:27Etc and also we are going to review the
- 1:30popular building blocks with the fully
- 1:32connected layer convolution layer group
- 1:34convolution dep size convolution ESP
- 1:37especially to understand the computation
- 1:40um demand behind that including both the
- 1:42memory demand and and also the
- 1:44computation
- 1:46demand uh we're also going to introduce
- 1:49those efficiency matri matric including
- 1:51the number of parameters how to
- 1:53calculate number of parameters the model
- 1:55size the peak activation size the Mac
- 1:58flop flops op Ops what is the difference
- 2:02between them what is the relationship
- 2:04between them latency
- 2:06throughput and if we have time we're
- 2:08going to go through lab zero which is a
- 2:10tutorial on
- 2:12pytorch so lab zero is optional but make
- 2:15sure uh we also want you to make a
- 2:17submission on canas it's not going to be
- 2:21graded but it'll get get yourself
- 2:23familiarized with the pro procedure all
- 2:26do we submit homework on canas and for
- 2:30all the announcements related U please
- 2:33keep an eye on our course website we are
- 2:36going to make all the announcements the
- 2:38lecture slides lecture videos homeworks
- 2:43assignments everything will be released
- 2:46at our course website which is efficient
- 2:51ml. so if you have not bookmarked it uh
- 2:55make sure to bookmark our course website
- 2:58efficient ml. I and check it every week
- 3:02uh to get the latest uh slides videos
- 3:04and homeworks all the home we have five
- 3:07Labs the schedule of the lab release
- 3:09will be also announced in the efficient
- 3:12m.ai website so make sure you bookmark
- 3:14our
- 3:17website okay so let's jump into neuron
- 3:20uh synaps so as sh on the right hand
- 3:23side we have a synapse on the left and
- 3:26and um we have input X on the left
- 3:29output on the right and we have the
- 3:31synapse that compute the weighted
- 3:33average of the input axtion so the
- 3:36synapse is basically the weight of the
- 3:38neural
- 3:39network and so the right part is the
- 3:42synapsis right so we use the termin
- 3:45terminology interchangeably between
- 3:47synapses weights and parameters so if
- 3:50you heard these terms they mean the same
- 3:53thing and the neurons features and
- 3:56activations we use these terminologies
- 3:58interchangeably they also mean the same
- 4:01thing so in this example we have uh two
- 4:05layer neuron Network and get output so
- 4:09the dimensionality of these hidden
- 4:11layers determines the width of the model
- 4:14right so the width of the model matters
- 4:17a lot to the efficiency of neural
- 4:20networks so we can design a wide but
- 4:23shallow network with the same number of
- 4:26parameter as um um narrow but deep
- 4:30neuron Network so I guess which one is
- 4:33more accurate and which one is more
- 4:36Hardware efficient in these two
- 4:38scenarios one is a very deep but niral
- 4:43neuron Network the other is a shallow
- 4:46but wide neuron Network which one is
- 4:48more Hardware
- 4:52friendly wide and shallow and why is
- 4:55that the case
- 5:02exactly fewer number of crial cause and
- 5:05I have more paradism to fully exhaust
- 5:08the number of threads in the gpus uh so
- 5:11from the efficiency perspective a wide
- 5:14and shallow Network might be more
- 5:16efficient uh in the meantime from the
- 5:18accuracy perspective from the accuracy
- 5:21perspective usually a deep Network um
- 5:24matter helps a lot so there is a
- 5:26tradeoff you don't get a free Lance on
- 5:28both side so that's the beauty of neuron
- 5:30Network architecture design where we're
- 5:32going to visit in the later part of this
- 5:34lecture How do trade off between this
- 5:36hard Hardware efficiency and also the
- 5:39accuracy so let's visit a few popular
- 5:41neuron Network layers starting from the
- 5:44fully connected layer also called a
- 5:47linear layer so in AAR layer um you have
- 5:50a weight Matrix which is of Dimension CI
- 5:54input Channel times Co which is the
- 5:57output Channel Okay so um the number of
- 6:00parameters here is CI times Co okay and
- 6:04this is the input activation CI and this
- 6:06is the output activation
- 6:08Co so uh that's the input feature
- 6:11Dimension usually um when we are talking
- 6:14about efficient Computing is super
- 6:16important to understand the dimension of
- 6:19each tensor okay later when we are going
- 6:22to build uh the model to to calculate
- 6:25the flops number of operations and also
- 6:28the model the parameter size the model
- 6:30size it all concerns about the dimension
- 6:33of the of each tensor so it's very
- 6:36crucial to figure out the dimension of
- 6:38the input feature output feature map
- 6:41weight feature and also the bias
- 6:44especially the bias when we are going to
- 6:46learn about the uh efficient on device
- 6:48training we are having a technique
- 6:51called bius only update or recently the
- 6:53low which concerns about the bias so
- 6:55we're going to revisit that in later
- 6:57part of the lecture so so very simple
- 7:00output equal to the weight is sum of the
- 7:02input so that's the fully connected
- 7:04layer also called a linear
- 7:06layer what do we if we have multiple
- 7:09inputs so here we have a batch size of
- 7:12two we have two inputs right so the
- 7:15weight Dimension stays the same it's CI
- 7:18* Co but the input becomes n times CI
- 7:22and the output feature map becomes n
- 7:24times
- 7:26Z convolution layer so so um the output
- 7:31neuron is no longer connected to all the
- 7:33inputs but is connected only to the
- 7:36inputs within the receptive field okay
- 7:39so we have a input input feature and
- 7:43this is the spatial Dimension now we
- 7:45just assume it's 1D convolution and
- 7:47later let's jump into 2D in this 1D
- 7:50convolution um this is the channel
- 7:52Dimension this is the spatial 1D spatial
- 7:55Dimension and we apply this green filter
- 7:59here we also have the input Channel
- 8:01Dimension CI and here we have um the KW
- 8:06which is the uh a filter Dimension and
- 8:09then we have multiple such filters okay
- 8:12each filter is going to convolve with
- 8:14the input and produce one
- 8:18output and we can shift it shift this
- 8:21sliding window by one to get another
- 8:24output we can shift it again to get
- 8:26another output so as a result the output
- 8:30feature um Dimension output Channel
- 8:33Dimension is three since we have three
- 8:36uh kernels for the
- 8:38weight and the weight Dimension is
- 8:40basically CI * K time Co since each
- 8:44kernel Dimension we have a CI * K and we
- 8:47have C this amount
- 8:48of
- 8:51kernels and finally we have a bias with
- 8:54the dimension of
- 8:56Co and now let's generalize that to uh
- 9:00two dimensional convolution 2D C okay as
- 9:04we can see we added another dimension uh
- 9:07it's 2D convolution but looks like the
- 9:10feature map is looks like the feature
- 9:12map is threedimensional
- 9:13that's uh because for each XY location
- 9:18we have a channel Dimension okay the
- 9:20channel Dimension so the for 2D
- 9:23convolution the feature map is actually
- 9:263D so the activation map um the H times
- 9:30W assuming uh imagine you have a picture
- 9:34uh width is W height is H and then we
- 9:38apply a filter which is KH by KW which
- 9:42is the filter um dimension for the h and
- 9:45W Dimension with the same input Channel
- 9:48Dimension CI and then we have multiple
- 9:51such filters three such filters which is
- 9:54equal to Co in this case and we can
- 9:57shift it by one to get a another output
- 10:01shifted by two to get the another output
- 10:05we can shift in another
- 10:06dimension on and so on so get the output
- 10:09feature map 3x3 output
- 10:13feature and how do we calculate the uh
- 10:17size of the output feature map is equal
- 10:19to um the size of the input minus the
- 10:22size of the kernel plus one and why plus
- 10:25one because you can shift uh KH minus
- 10:29one this amount of shifts and applying
- 10:33each shift you get a one more output so
- 10:36H H equal to hi minus KH plus one in
- 10:40this case um the hi equal to four the
- 10:44input is
- 10:454x4 um and then the K the cernal size is
- 10:49three therefore the output Dimension is
- 10:514 minus 3 + 1 it'll be two in this case
- 11:00let's also talk about padding okay so
- 11:02padding can be used to keep the output
- 11:04feature map the same as the input
- 11:06feature map otherwise using convolution
- 11:09um the feature map will get smaller and
- 11:11smaller as you have a deeper number of
- 11:13layers so one common way is to do uh to
- 11:17apply zero padding pad the input
- 11:19boundaries with the zero which is the
- 11:22default um in py torch there's for
- 11:25example in this case uh the original um
- 11:29I and width equal to five in this
- 11:32animation kernel size is three so um the
- 11:38output height and weight equal to 5 + 2
- 11:42* 1 2 p okay padding equal to one we pad
- 11:46one on each side and minus 3 plus one
- 11:49equal to five so the output feature map
- 11:51is the same as the input compared with
- 11:54the initial version like in here um the
- 11:58output
- 11:59becomes uh becomes smaller the input is
- 12:02three 4x4 but output becomes two um as
- 12:06we can see on
- 12:08the on the example here right initially
- 12:11the input is 4x4 output becomes two and
- 12:15after we apply padding after we apply
- 12:18padding the input and output are both
- 12:21five by five in this
- 12:23case so we have we can have zero pading
- 12:26there are also other methods to apply
- 12:28pading for example using a reflection
- 12:32pading okay so for example right here
- 12:36the um um padded number is the reflected
- 12:41value of the original feature map we can
- 12:43also apply the replication padding to
- 12:45use the closest number and replicate
- 12:48that to do to apply the padding but zero
- 12:51padding is still the widely used
- 12:54approach so after padding let's now talk
- 12:57about the receptive field for example
- 13:00the output pixel here how many pixels
- 13:03Can it can it see in the previous layer
- 13:06it can see 3x3 pixels if the kernel size
- 13:10is 3x3 right and each pixel here is
- 13:13another 3x3 so how many pixels from here
- 13:17to here actually 5 by five right and one
- 13:21more layer you get S by seven receptive
- 13:23field why do we need a large receptive
- 13:26field we want to understand the relation
- 13:29ship between different
- 13:31pixels for example a car on the road
- 13:35versus some pedestrian also on the road
- 13:37the relationship between the car and
- 13:39pedestrian I means whether you can drive
- 13:42or you can you should avoid pedestrian
- 13:45right so the relationship matters for
- 13:47this high level Vision tasks and why
- 13:49people youed Transformers later because
- 13:52Transformer give you a um Global
- 13:55receptive field you can see every pixel
- 13:57can attend to every pixel Invision
- 13:59Transformers we're going to cover later
- 14:01so the principle is that a larger
- 14:03receptive field helps a lot it's very
- 14:06crucial uh to high level the image
- 14:09understanding so with two layers three
- 14:12kernels we can see receptive field um of
- 14:155x five and in general the equation is
- 14:18here with L layers the receptive field
- 14:22size equal to l * kernel size minus one
- 14:25+
- 14:27one uh so for two layers kernel size
- 14:30equal to three kernel the receptive
- 14:32field is
- 14:33five and you have three layers kernel
- 14:36size is is three the receptive field is
- 14:39is
- 14:41seven so the problem is that for large
- 14:44images in order to have a large
- 14:46receptive field we have to have very
- 14:49deep layers we need many many layers and
- 14:52we just talk about that in the early
- 14:54part of the lecture if you have too many
- 14:56layers you have lots of Kernel CA you
- 14:59also have to store a lot of activations
- 15:02uh activations has to be stored during
- 15:05back propagation increasing your
- 15:07training memory so how do we have a
- 15:10large receptive field at the same time
- 15:12we don't want to have so many
- 15:15layers so we can done sample inside the
- 15:18neuron Network one example is to use the
- 15:21strided convolution layer
- 15:24okay so compared with Str equal to one
- 15:27which is on the bottom
- 15:29so on the top shows the St equal to two
- 15:33Okay so rather than every pixel sees the
- 15:36adjacent 3x3 pixels now is see 5x five
- 15:40okay the adjacent green um adjacent
- 15:44green 3x3 re uh rectangle is shifted by
- 15:49not one pixel but two pixels okay you
- 15:52can see this 3x3 versus this 3x3 is
- 15:56actually shifted uh by by two so that's
- 15:59a stride of two every time you don't
- 16:01move one step but you move two two
- 16:05steps so that's a the uh stred
- 16:09convolution in this case for two layers
- 16:12and kernel size equal to three the
- 16:14receptive field becom seven okay
- 16:16previously it was five now it was seven
- 16:18you have a larger receptive field with
- 16:21the same number of
- 16:23layers so this is very helpful for
- 16:26reducing the number of layers reducing
- 16:28the number of weights reducing the
- 16:29number of activations but have the same
- 16:32receptive field but of course you may
- 16:34lose a bit accuracy because the number
- 16:36the model capacity might be
- 16:38smaller but everything is tradeoff
- 16:40there's no no free lunch in deep neural
- 16:43network design everything is about
- 16:45constraint optimization which makes the
- 16:47optimization highly
- 16:51interesting okay so our goal is to
- 16:55continue reducing the number of
- 16:56parameters okay so for convolution layer
- 16:59each output channel is connected to all
- 17:02the input channel to all the input
- 17:03Channel okay how to reduce the amount of
- 17:07computation people invented this group
- 17:10convolution layer okay so previously we
- 17:13have just one group every channel is
- 17:15output channel is connect to all the
- 17:16input Channel now we can divide them
- 17:19into two groups so for the first group
- 17:22only half of the input channel is
- 17:25connected and for the second half um
- 17:28that's the second group okay so
- 17:31effectively if we have G uh groups okay
- 17:34we can see uh the feature Map size
- 17:38doesn't input feature Map size and
- 17:39output feature Map size doesn't change
- 17:42stays the same as before but the number
- 17:44of weight will be G times smaller
- 17:49okay will be G times smaller So Co will
- 17:52be divided by G and CI will also be
- 17:55divided by G and we have G amount of uh
- 17:59such such filters okay therefore the
- 18:02total amount of Weights is reduced by G
- 18:06okay so that's the group convolution
- 18:10here and what is the extreme of group
- 18:12convolution the number of group the
- 18:15number of group will be equal to the
- 18:19number of input or output Channel okay
- 18:21so which is exactly uh this case
- 18:25previously we have two groups now we
- 18:27have Co groups in this case uh it's it's
- 18:31eight right we have eight groups so each
- 18:34CH output channel is only connected to
- 18:37one input Channel okay so the number of
- 18:41number of way uh number of uh input
- 18:44Feature Feature stays the same and the
- 18:46number of Weights will be C which is the
- 18:49same for c i and Co so we just use c
- 18:52times uh KH and KW which is the extreme
- 18:55case of group convolution called deps
- 18:59wise convolution layer which is quite
- 19:01widely used since 2015 2016 where uh
- 19:06when mobile net was
- 19:11invented the next um ping layer we want
- 19:14to have a smaller feature map originally
- 19:17the input image might be like 1K uh by
- 19:21by 1K is pretty large and we want to
- 19:24have a smaller feature map and condensed
- 19:26information right um this is actually
- 19:29quite useful for large resolution uh
- 19:32input we want to quickly pull the
- 19:35feature map to make it smaller so here
- 19:38we have an example with the 4x4 input
- 19:41feature map and how do we put it to 2x2
- 19:45we can either use max putting so for
- 19:47every 4x4 input we find the max element
- 19:51within that group um or we can do uh
- 19:55this kind of average pting finding the
- 19:57average in the Blue Area finding the
- 20:00average for the yellow area to do the
- 20:03average
- 20:04pulling the good thing about ping is
- 20:07that there's no learnable parameters
- 20:10okay or you can assume um yeah so but
- 20:14contain zero weight uh which is very
- 20:17parameter
- 20:19efficient next let's talk about
- 20:22normalization year so normal bch
- 20:23normalization came roughly in 2015 or or
- 20:27new 15 or or 16 roughly seven eight
- 20:31years ago uh which later becomes quite
- 20:34useful to stabilize the training for
- 20:36example the the Leer Norm also become
- 20:38super helpful in the age of uh uh
- 20:41attention Transformers so what exactly
- 20:44is
- 20:45normalization so we want to normalize
- 20:48the feature map to make the optimization
- 20:52faster okay so how do we apply that so
- 20:55we minus the feature map minus the mean
- 20:58of that tensor okay and divided by the
- 21:01sigma okay so we want to make sure um we
- 21:05want to minus the mean and divided by
- 21:07the standard deviation over a set of
- 21:10pixels or tensors if it is not Vision
- 21:13okay so this is a reminder how to uh
- 21:18calculate the mean and the standard
- 21:21deviation and note here we have a small
- 21:25Sigma right here okay so that is due who
- 21:28we want to avoid dividing by zero so we
- 21:32put a very small number over
- 21:34there and then we learn a per Channel
- 21:38linear transformation okay uh which is
- 21:42indicated by the scaling factor and also
- 21:44the bias right here to compensate for
- 21:47the um possible loss of represent
- 21:50representational ability
- 21:53okay so how do we Define this set of
- 21:56pixels or set of tensor
- 21:59on the bottom it shows four different
- 22:02ways to find uh that group of pixels uh
- 22:07the first one very widely used in CNN
- 22:11called batch
- 22:12normalization okay so it's normalizing
- 22:15across the N Dimension n is the image
- 22:18Dimension okay and also we just use a
- 22:21single C Dimension across all the H and
- 22:24W height and width dimension okay
- 22:27normalize across
- 22:29different batches second one is quite
- 22:32widely used in the attention mechanism
- 22:35so for different input for each input
- 22:38okay across different um H and W and C C
- 22:42is a channel Dimension make sure it's
- 22:45normalized so uh for each token in the
- 22:49attention mechanism when you are doing
- 22:51the attention after each token is
- 22:53normalized then the attention map uh
- 22:56makes sense otherwise uh um if it's not
- 22:59nor normalized the tension map will be
- 23:02um much harder to
- 23:05optimize there's also the instance
- 23:08normalization it just normalized across
- 23:10the H and W Dimension not the C
- 23:13Dimension and when the C is too large
- 23:16there's also the group normalization on
- 23:18the right uh which divide the C by
- 23:20several groups and only normalize within
- 23:23that
- 23:24group there's a lot of tricks we can
- 23:27play for the normalization layer for
- 23:30example when we are doing
- 23:32fing um actually normalization layer the
- 23:35number of parameter is very small so we
- 23:37can only fine tune the weight the bias
- 23:39and also the um the scaling factor in
- 23:43the normalization layer making the fine
- 23:44tuning much more computation primary
- 23:48efficient the normalization layer can
- 23:50also absorb um the transformation of
- 23:54such a smooth quantization so without
- 23:57having incurving another peral call we
- 23:59can absorb lots of the computation in
- 24:02the IP log or fuse it with the
- 24:04normalization later so we can we are
- 24:06going to visit them in later part of
- 24:07this
- 24:10lecture okay next one is activation
- 24:13function so uh we have seen many
- 24:16different activation functions but let
- 24:18me talk about the efficiency and
- 24:19accuracy trade off here sigmoid very
- 24:23ancient um activation function the it
- 24:26ranges between zero and one
- 24:28is very easy to quantize very efficient
- 24:31to quantize because the dynamic range is
- 24:34very limited and this fixed between zero
- 24:36and one but what is what is the
- 24:40drawback gring vanishes right so if the
- 24:43value super small or super large there's
- 24:46no
- 24:47gradient and then people come up with Ru
- 24:50okay when it's positive the gradient
- 24:52never vanishes but if it's when it's
- 24:55negative the neuron is basically dead
- 24:57there's no way to put it back because
- 25:00there's no gradient anywhere when the
- 25:03value is smaller than
- 25:05zero but what the good what is good
- 25:07about it is that it's very easy to
- 25:10sparsify since we're going to talk about
- 25:12prud sparf and sparcity in the later
- 25:15part of this lecture uh the ru
- 25:18activation function naturally brings
- 25:20sparcity because lot of the activation
- 25:23once they are act once they are negative
- 25:25they become zero and zero multiply but
- 25:28by anything is zero so don't have to
- 25:30compute on it you don't even to store it
- 25:33you can just store a mask showing that
- 25:35this is zero just using one
- 25:39bit and the gradient is one um
- 25:42everywhere else okay so you don't even
- 25:45need to save this feature map you can
- 25:47just save a bit mask um for this for
- 25:51that tensor if it is the bit MK is zero
- 25:54then the output is zero if the back bit
- 25:56MK is one it will just copy the feature
- 25:58map from the previous output so very
- 26:02activation friendly sparcity friendly
- 26:05and
- 26:07simple but what is the downside the
- 26:10dynamic range is super large it's not
- 26:13easy to quantize so peopleand it R six
- 26:16okay R six to cap the larest value to be
- 26:20six that's why people call it R six okay
- 26:23it's purely
- 26:24empirical and that makes the
- 26:26quantization range much more constraint
- 26:29reducing the Dy dynamic range make it
- 26:31easier to
- 26:33quantize another drawback is that
- 26:35there's no gradient in the negative
- 26:37region so people imaged Le R okay
- 26:40there's a smaller scope slope on the
- 26:44negative side okay so um you can also
- 26:47like the gradient have the gradient flow
- 26:49in the negative region but the down side
- 26:52is that you no no longer have this
- 26:54sparcity benefit if the value is small
- 26:57it's negative it's no longer it's no
- 27:00longer zero
- 27:02okay and then people are using neural
- 27:04architecture search find a very
- 27:06interesting activation function called a
- 27:08switch okay switch basically very
- 27:11similar to Ru but it's using x / 1 plus
- 27:16e to the minus of X but people later
- 27:18find it hard to implement in Hardware
- 27:21right so people come up with a
- 27:23approximated switch called hard switch
- 27:27okay so um when X is smaller than minus
- 27:31three is zero above three is just X in
- 27:35between that um the value is x + 3
- 27:39divided by 6 okay so all this activation
- 27:44function uh somehow have something to do
- 27:46with efficiency like sparcity
- 27:49quantization dynamic range easy to
- 27:51implement so that's the story behind uh
- 27:54the progression of different activation
- 27:57functions
- 28:00and finally Transformers Transformer are
- 28:02getting super important since the uh
- 28:06invention in 2017
- 28:082018 uh we learn more about Transformer
- 28:10architecture in lecture 12 we have a
- 28:12dedicated lecture just to talk about
- 28:14Transformers the entire uh pipeline
- 28:18entire structure of Transformers so
- 28:21let's just give a very simple overview
- 28:24uh
- 28:24today so there are two stages encoding
- 28:27stage and the decoding stage for each
- 28:30stage we have a mod tension followed by
- 28:34f f forward Network and this is the
- 28:38architecture for the attention mechanism
- 28:41you have three tensors q k v quy key and
- 28:44value we do a um
- 28:49attention scale dot product attention
- 28:52and then pass it through a linear layer
- 28:54get the o q k v o those are the four um
- 28:58activation tensors and there's no weight
- 29:00in the attention mechanism uh that's the
- 29:04in the middle part only wait in the qk
- 29:07vi transformation okay and the detail
- 29:10way to do that for the query key value
- 29:12the design is actually very analogous to
- 29:15the retrieval system just take a YouTube
- 29:17search as example the query is like the
- 29:21text prompt in the search bar okay and
- 29:24the key would be like the title or
- 29:26description some short information about
- 29:28the video and value would be the
- 29:31corresponding video so you do a DOT
- 29:34product with the query and the key to
- 29:35get the similarity and use the
- 29:37similarity to fetch to do a weighted
- 29:40average of the
- 29:42value and remember we want to divide the
- 29:46Q and K by a normalization factor which
- 29:50is square root of d uh to accommodate
- 29:53for the dimension change to make the um
- 29:57optim ization more stable and followed
- 30:00by the soft Max to get a um a tension
- 30:04map a n byn attention map and use that
- 30:07attention map to do a a weighted average
- 30:10weighted sum of the of the value okay
- 30:14and later part of the lecture we are
- 30:16going to visit in the tension map the N
- 30:18byn tension map could be the bottom NE
- 30:22for both computation and the uh other
- 30:26the memory since this is growing
- 30:29quadratically with the number of token
- 30:31in okay and for images the number of
- 30:34token also grow quadratically with the
- 30:36resolution so n Square would be super uh
- 30:39expensive so there's sparity opportunity
- 30:42not all attention not all token need to
- 30:44attend attend to each other sparse
- 30:46attention we're going to introduce that
- 30:49uh and also flash attention we can use
- 30:51in place um attention um to minimize the
- 30:55data movement uh so that we can have a
- 30:58constant almost constant memory but the
- 31:01flops is not not reduced so in order to
- 31:04deal with that problem we are going to
- 31:06introduce long context techniques all
- 31:09related to this attention mechanism so
- 31:12although attention itself doesn't have
- 31:14any weights there's a lot of interesting
- 31:16um interesting stuff we can explore even
- 31:19the KB cache um they become the
- 31:23activation when we are calculating the
- 31:25KV cache but when we are using that to
- 31:27generate the next token they become the
- 31:29weight to to when new tokens is
- 31:32generated so how do we quantize how do
- 31:34we uh Pro the KV cache we're also going
- 31:37to cover that uh in later part of the
- 31:40lecture let's continue the lecture to
- 31:43talk about the efficiency metrix how
- 31:46should we measure the efficiency of
- 31:48neuron
- 31:51networks so there are several pillars we
- 31:54want to achieve at the same time but
- 31:55it's pretty hard having a smaller model
- 31:58so when we up when you're uploading the
- 32:00model to the Apple Store it's pretty
- 32:02small right you don't want to download a
- 32:03model that is hundreds of gigabytes that
- 32:05is too hard to
- 32:06download and also we want to have a
- 32:09faster model it give you the real time
- 32:11prediction never miss a pedestrian the
- 32:13road for self-driving car and also
- 32:15generate images instantly on your phone
- 32:17for
- 32:18example and also Greener model taking
- 32:21less battery for example on the iPhone
- 32:23on the mobile devices you you don't want
- 32:25to join your battery with this AI
- 32:26applications
- 32:28and that concerns theu right and
- 32:31computation the memory is the
- 32:32determination is determine determining
- 32:35the storage latency and energy so we are
- 32:38going to introduce several efficiency
- 32:41matrics on the right hand side um some
- 32:44of them are memory related some of them
- 32:46are computation related on the memory
- 32:49side we'll introduce how to count the
- 32:51number of parameters um the model size
- 32:54the total activation size and Peak
- 32:57activation size size for the computation
- 32:59related we're going to talk about what
- 33:01is Mac what is flop what is flops what
- 33:04is op what is
- 33:06OPS so let's first talk about latency so
- 33:09what is latency latency measures the
- 33:12delay for specific task and let me
- 33:15directly show you a example so on the
- 33:17left hand side it's high latency right
- 33:20each to predict the segmentation mask of
- 33:25output takes a long time and right hand
- 33:28side is lower latency uh only 45 46
- 33:32milliseconds to process each
- 33:35frame what is throughput it measures a
- 33:38rate at which data is processed on the
- 33:41left hand side is low throughput you can
- 33:43process six videos per second on the
- 33:46right hand side is high throughput 77
- 33:48videos per second this is from the
- 33:50temporal shift module paper for video
- 33:52understanding we are going to introduce
- 33:54later so what is the relationship
- 33:57between latency and throughput does
- 34:00higher throughput translate to lower
- 34:01latency does lower latency translate to
- 34:04higher throughput so that's see example
- 34:07the first design the Laten is 50
- 34:11milliseconds so every 50 millisecond we
- 34:13can process my image uh so how many
- 34:16images can we process every second is 20
- 34:20right on the design to we have higher
- 34:23pism we can process four Images at the
- 34:26same time each taking 100 millisecond so
- 34:29latency to process each image from start
- 34:32to end is 100
- 34:34millisecond and the throughput would be
- 34:3640 images um per second since we are
- 34:40processing uh four such image at the
- 34:43same
- 34:45time so on the left hand side is having
- 34:48a lower latency a lower throughput right
- 34:51hand side is having a higher latency but
- 34:53higher throughput as a result higher
- 34:55throughput doesn't translate to lower
- 34:57latency
- 34:58and lower latency also doesn't translate
- 35:00to higher throughput on the mobile we
- 35:03sometimes care about the we mostly care
- 35:05about the latency on the data center on
- 35:08batch processing tasks we care about the
- 35:09throughput so those are very important
- 35:12metrics latency and throughput sometimes
- 35:15we we want to make sure we report
- 35:20both so what is the uh factors that
- 35:23impact the latency so um there are two
- 35:27two factors that is determining the
- 35:30latency one is the computation one is
- 35:32the memory okay um the T of computation
- 35:35equal to the number of operations in a
- 35:37neural network model divided by the
- 35:39number of operations that the process
- 35:42can process per
- 35:45second and the bottom is the hardware
- 35:47specific uh characteristic on the top is
- 35:51Neuron Network specific and then T
- 35:53memory also determine is determined by
- 35:56two factors the activation
- 35:58and also the weight to move the
- 36:00activation and move the weight so this
- 36:02is about inference but it's training
- 36:03also concerns about communication of the
- 36:06gradient so T data movement of weight
- 36:09equal to the the weight the model size
- 36:11divided by the memory bandwidth Hardware
- 36:14specific neuron Network
- 36:16specific and T data movement of
- 36:18activations is determined by the input
- 36:21activation sides plus the output
- 36:23activation sides this is first order all
- 36:25these are first order approximations
- 36:27give you a ballpark estimation divided
- 36:30by the memory bandwidth of the
- 36:33processor so which one is more expensive
- 36:36so this is a figure showing the energy
- 36:40consumption um of different operations
- 36:43from a 32bit integer add to accessing a
- 36:48register file to a
- 36:50multiplication to accessing the SRAM
- 36:53cach to access the dram
- 36:55cache to access the D memory we can see
- 36:58that access in the D memory could be
- 37:00super expensive okay 640 PJ versus just
- 37:05doing a 32-bit add it's just 0.1 P so
- 37:10Accents in the memory is a lot more
- 37:12expensive than doing the
- 37:14arithmetic uh like two others of
- 37:16magnitude more energy is consumed by
- 37:20access in the memory so keep in mind
- 37:22computation is cheap memory data
- 37:25movement is very expensive
- 37:28so it is the data movement that is
- 37:30joining uh joining the battery
- 37:34okay so let's talk about how to
- 37:36calculate the number of parameters since
- 37:39data movement is
- 37:41expensive so um parameter is the number
- 37:44of weight in a neuron Network okay for
- 37:48example in the linear layer the weight
- 37:51tensor is C in times C out so the number
- 37:55of parameter is just C C in time C out
- 37:59in a convolution layer the weight tensor
- 38:04is four dimensional okay each kernel has
- 38:07KH KW time C in and we have Co Am Co
- 38:12number of kernels so all together we
- 38:14have a multiplying these four terms
- 38:17together to get the total number of
- 38:19parameters for a convolution
- 38:23layer for a group convolution okay we
- 38:26divided the CI and Co by G so CI is
- 38:31divided by G Co is divided by G but we
- 38:34have G amount of such kernels so Al
- 38:38together we have only one g on the
- 38:41denominator so G times smaller compared
- 38:43with the convolution with the same
- 38:46dimension for depthwise convolution
- 38:49where G equal to c i equal to co okay
- 38:53therefore the total number of parameter
- 38:54equal to co * KH * KW
- 39:01and let's see the PO the number of
- 39:02parameters for popular neural networks
- 39:04forx net right so let me do the
- 39:08calculation for the first layer uh where
- 39:10the channel number is uh input channel
- 39:13is three output channel is 96 and the
- 39:16kernel size is 11 by 11 so the first
- 39:19layer the number of parameter is 96
- 39:22output three input 11 by 11 okay and
- 39:26similarly we can calate the second layer
- 39:28it's a group convolution so we divide
- 39:31the number of parameters by by two since
- 39:33group size is
- 39:36two so I give you one minute to
- 39:38calculate to think about the the layer
- 39:41number of parameters for the next layer
- 39:43okay here 13 by 13 um image size and 3x3
- 39:52convolution the channel number is two
- 39:55384
- 39:59so here's a quick answer since the input
- 40:01channel is 2 256 output is 384 and the
- 40:05channel number uh colal size is 3x3 this
- 40:08is the way to calculate that and finally
- 40:10for the fully connected layer you just
- 40:13multiply the input Channel with the
- 40:14output channel to calculate the total
- 40:16number of Weights which is 61 meaning in
- 40:20total so how does that translate to the
- 40:23model size between the number of
- 40:25parameters to model size what's the Rel
- 40:27relationship you want to multiply the
- 40:30number of bytes for each parameter with
- 40:33the number of parameters okay so number
- 40:36of parameters with multiplied with bit
- 40:38width look get a model size so for
- 40:42example alxn has 61 million parameters
- 40:45if each weight is in uh is represented
- 40:48by the um single Precision 32 bit 32-bit
- 40:53number and how many bytes are there in a
- 40:5732
- 40:59bit four bytes right 32 bits equal to
- 41:03four bytes so 61 million * 4 you have
- 41:06224
- 41:08megabytes so make sure you have the
- 41:10prerequisite so that you can understand
- 41:11these Concepts relatively easily how
- 41:14many bits per bite and you want to
- 41:16quantize the weight to uh integer right
- 41:21to 8bit integer what is the total
- 41:23storage for the
- 41:25weight so 8 bit that equal to one byte
- 41:2861 million parameters that's 61
- 41:31megabyte 61 megabyte four times smaller
- 41:34right compared with 32 bit and later
- 41:38part of this lecture we are going to
- 41:40learn about activation where with only
- 41:42quation awq to quantize it to four bit
- 41:46okay four bit will be another two times
- 41:48smaller so that it can fit large
- 41:50language model locally on your
- 41:54laptop okay so now let's switch gear to
- 41:56talk about the total and the peak number
- 41:59of
- 42:01activations so activation is the bottom
- 42:04neck for seeing inference not the
- 42:06parameters for example here from reset
- 42:09to mobile net number of parameters is
- 42:12reduce a lot by four times more than
- 42:15four times but the activation the peak
- 42:18activation didn't decrease but increased
- 42:20so um for how to reduce the activation
- 42:24activation is very crucial the deeper
- 42:26the layer the more more number of
- 42:27activations you that's the
- 42:31case and also um the peak activation
- 42:34determines the memory bottle neck um for
- 42:38example uh the third layer here
- 42:41determines what is the max memory you're
- 42:43are going to consume when you are doing
- 42:46the
- 42:47inference so the imbalanced memory
- 42:49distribution happens very popular for CN
- 42:53why is that the case because the first
- 42:55couple of layer you have very high
- 42:56resolution
- 42:58and then square and later you have a
- 43:00lower resolution due to down
- 43:03sample for example if a microcontroller
- 43:06has only 256 kilobyte of ice Ram um in
- 43:10the
- 43:12uh of of SRAM to hold these activations
- 43:15you want to make sure uh the largest
- 43:18layer for example here doesn't exceed
- 43:21the SRAM so that you can put all the
- 43:23activations in the fast SRAM
- 43:28activation is also becoming the memory
- 43:30bottom neck in training in see for large
- 43:34language models the number of parameter
- 43:36is also the bottom neck that's why
- 43:38people later proposed uh this model
- 43:41paradism to do the sharding of the
- 43:43weight so here is example for the CN
- 43:46where um the this is showing the number
- 43:49of parameters versus the number of
- 43:52activations uh the number of activation
- 43:54is like seven times larger than the
- 43:56number of parameters when you are doing
- 43:58um training due to the batch the
- 44:01activation is pretty large occupying a
- 44:03lot of the uh GPU
- 44:06memories and here from resent 50 to
- 44:10mobile 9 V2 1.6 they pretty much have
- 44:13the same accuracy so it's apples apples
- 44:16comparison the parameter is reduced by
- 44:18four times but the main memory botom NE
- 44:22which is the activation didn't improve
- 44:23much only 1.1 times only 10% so uh
- 44:27imizing the activation is
- 44:29hard and usually for CN the distribution
- 44:32for the weight and activation um is
- 44:36different uh the yellow part is the
- 44:39activation um the blue part is the
- 44:42weight it's new shaped okay initially um
- 44:46the activation memory is large later the
- 44:48weight memory is large this is due to
- 44:50has high resolution um and this is due
- 44:54to it has larger amount of channels
- 44:59so this is calculating the number of
- 45:00activations of uh anx net so um C * H *
- 45:07W that's the way to calculate the
- 45:09activation size and if you add them
- 45:11together can see this is total amount of
- 45:13activations for LX
- 45:16net and the peak activation is roughly
- 45:19first order approximation number of
- 45:21input activation plus the number of
- 45:23output activation
- 45:27okay now let's switch gear to talk about
- 45:29um compute bonded Matrix for example Mac
- 45:34so what is Mac is not McDonald it's
- 45:37multiply and accumulate right multiply
- 45:39and accumulate so a equal to a plus b *
- 45:43C that's one Mac one accumulation and
- 45:46one multiplication it usually comes as a
- 45:49single instruction in
- 45:52gpus so for Matrix Vector multiplication
- 45:55how to calculate the number
- 45:58Max for example on the top right corner
- 46:00this is a Matrix Matrix Vector
- 46:05multiplication okay so we are going to
- 46:07produce M okay M outputs in order to
- 46:11produce each output we need to compute n
- 46:14multiplication as in order to produce
- 46:17each output we need to calculate an Max
- 46:21so all together we have M * n
- 46:24Max what what about a matrix Matrix
- 46:27modlic short for G MN General Matrix
- 46:31Matrix modlic on the right hand side we
- 46:34have a m by K Matrix times a k by n
- 46:38Matrix uh and produce a m by n output so
- 46:42we have M byn output in order to
- 46:44calculate each output we need K Max so
- 46:48all together we need M * n * K
- 46:54Max so here we are show we're showing
- 46:56the
- 46:57number of Macs for different layers for
- 46:59linear layer we just talk about that the
- 47:01Mac equal to the C in times C out assume
- 47:05batch size of one and for convolution is
- 47:08actually multiplying six terms together
- 47:12okay so we are calculating this is the
- 47:15amount of uh output feature output
- 47:18features we have to calculate which is
- 47:21um W Times W output H output times Co
- 47:26okay and in order to compute in order to
- 47:30compute each output pixel how many Max
- 47:33do we need we need the entire kernel to
- 47:36be the to be convolved so that's KW * K
- 47:39* CI so multiplying them together we
- 47:42have to multiply these six terms
- 47:44together six three from here and three
- 47:46from here okay that's the way to
- 47:48calculate the convolution
- 47:50max with group convolution okay um we
- 47:54just divided by by G G is the number of
- 47:56groups
- 47:57and for depths wise convolution c i
- 47:59equal to G so we canceled one term
- 48:02multiplying these five terms together CI
- 48:05gets cancelled due to G equal to
- 48:10CI and this is calculating the max for
- 48:13Alx net for example each layer
- 48:16multiplying these six terms together if
- 48:19it is group convolution divided by the
- 48:21number of groups in this case two in
- 48:24this case and finally for the last three
- 48:27fully connected layers um it's just
- 48:30input Channel times output channel two
- 48:32terms multiplied together where
- 48:35ultimately you have 724 mions Med Max in
- 48:41total okay so now let's switch GE to
- 48:44talk about flop and
- 48:46flops so one
- 48:49multiply is a floating Point operation
- 48:52one add is also a floating Point
- 48:53operation so flop is short for number of
- 48:57floating Point operations floating Point
- 49:01operations so how many M how many flops
- 49:04are there in a
- 49:05Mac a Mac has a multiply a multiply is a
- 49:09fla a Mac also has a um an ADD and add
- 49:13is also a fla so one multiply and
- 49:17accumulate operation is equal to two
- 49:20floating Point operations if the operant
- 49:22are floating Point numbers for example
- 49:25alexnet has 700 24 mimax the total
- 49:28number of floating Point operations will
- 49:30be 724 * 2 that's 1.4 Giga flops
- 49:36okay and what is flops so the short for
- 49:39floating Point operations per second the
- 49:42number of flops divided by how many
- 49:44second so that's a performance metric
- 49:47showing how many how fast are you
- 49:49processing these floating Point
- 49:52numbers so what is the
- 49:54op we don't always represent the numbers
- 49:57in floating Point sometimes we represent
- 50:00them in integer right so both floating
- 50:03Point operation and integer operation
- 50:04all kinds of in operation even bit mask
- 50:08or even exore operation we can call it
- 50:10the op okay to generalize the number of
- 50:13operations is used to to measure the
- 50:16amount of
- 50:18computation uh for example if alexn has
- 50:21724 Med max if each number is
- 50:23represented by say binary or erary or
- 50:27floating point or integer what you have
- 50:30whatever you have any operation can be
- 50:32counted as an ALT so 724 M Max will be
- 50:37equal to
- 50:381.4
- 50:40gahs and similarly operations per second
- 50:44Ops equal toal number of operations
- 50:47divided by the number of seconds to
- 50:48complete this operation which is the
- 50:50performance uh speed
- 50:55Matrix hey before jumping into the uh
- 50:58lab zero tutorial let's give a short
- 51:01summary of today's lecture review the
- 51:03basics of neuron networks output
- 51:05activation synapsis weight popular
- 51:08layers into including FC layer
- 51:10convolution layer types as convolution
- 51:12putting normalization we also introduced
- 51:15efficiency metrics parameters model size
- 51:19activations Max flop flops up and UPS
About this transcript
This page contains the full transcript of EfficientML.ai Lecture 2 - Basics of Neural Networks (MIT 6.5940, Fall 2024, Zoom recording) by MIT HAN Lab, generated from the public captions YouTube serves with the video. The transcript has 7,322 words across 1,086 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.