YouTube2Text

EfficientML.ai Lecture 2 - Basics of Neural Networks (MIT 6.5940, Fall 2024, Zoom recording) — Transcript

by MIT HAN Lab · 7,322 words · 1,086 segments · language en · Watch on YouTube

Full transcript

  1. 0:00good afternoon everyone let's get
  2. 0:01started with lecture two of efficient
  3. 0:04ammo. so today we are going to learn
  4. 0:06about basics of neuron
  5. 0:09networks as as we have seen the last
  6. 0:12lecture there's a big gap between the
  7. 0:14supply and gr and demand of AI Computing
  8. 0:17where the demand of computation is
  9. 0:18growing really fast as indicated by the
  10. 0:21model size of recent large language
  11. 0:24models and mor slow is roughly doubling
  12. 0:27every two years but uh the deep learning
  13. 0:30models is getting more than four times
  14. 0:32larger every two years so there's a
  15. 0:34going to be a gap which motivated us to
  16. 0:37bridge this Gap with model compression
  17. 0:40and acceleration techniques so in the
  18. 0:42last lecture just do a quick recap we
  19. 0:45went through several examples from
  20. 0:48Vision to language to visual language
  21. 0:50model and also um um and also the uh
  22. 0:54multimodality models uh to give a
  23. 0:56overview about the latest uh advancement
  24. 1:00of AI Computing and how much the amount
  25. 1:03of computing they need to motivate us to
  26. 1:05build these efficient algorithms and
  27. 1:07systems to accelerate efficient air
  28. 1:10Computing so today we are going to uh
  29. 1:13start with those details technical
  30. 1:16details to First review the ter
  31. 1:19terminology of neuron networks such as
  32. 1:21neurons synapsis act activation feature
  33. 1:25weight parameter
  34. 1:27Etc and also we are going to review the
  35. 1:30popular building blocks with the fully
  36. 1:32connected layer convolution layer group
  37. 1:34convolution dep size convolution ESP
  38. 1:37especially to understand the computation
  39. 1:40um demand behind that including both the
  40. 1:42memory demand and and also the
  41. 1:44computation
  42. 1:46demand uh we're also going to introduce
  43. 1:49those efficiency matri matric including
  44. 1:51the number of parameters how to
  45. 1:53calculate number of parameters the model
  46. 1:55size the peak activation size the Mac
  47. 1:58flop flops op Ops what is the difference
  48. 2:02between them what is the relationship
  49. 2:04between them latency
  50. 2:06throughput and if we have time we're
  51. 2:08going to go through lab zero which is a
  52. 2:10tutorial on
  53. 2:12pytorch so lab zero is optional but make
  54. 2:15sure uh we also want you to make a
  55. 2:17submission on canas it's not going to be
  56. 2:21graded but it'll get get yourself
  57. 2:23familiarized with the pro procedure all
  58. 2:26do we submit homework on canas and for
  59. 2:30all the announcements related U please
  60. 2:33keep an eye on our course website we are
  61. 2:36going to make all the announcements the
  62. 2:38lecture slides lecture videos homeworks
  63. 2:43assignments everything will be released
  64. 2:46at our course website which is efficient
  65. 2:51ml. so if you have not bookmarked it uh
  66. 2:55make sure to bookmark our course website
  67. 2:58efficient ml. I and check it every week
  68. 3:02uh to get the latest uh slides videos
  69. 3:04and homeworks all the home we have five
  70. 3:07Labs the schedule of the lab release
  71. 3:09will be also announced in the efficient
  72. 3:12m.ai website so make sure you bookmark
  73. 3:14our
  74. 3:17website okay so let's jump into neuron
  75. 3:20uh synaps so as sh on the right hand
  76. 3:23side we have a synapse on the left and
  77. 3:26and um we have input X on the left
  78. 3:29output on the right and we have the
  79. 3:31synapse that compute the weighted
  80. 3:33average of the input axtion so the
  81. 3:36synapse is basically the weight of the
  82. 3:38neural
  83. 3:39network and so the right part is the
  84. 3:42synapsis right so we use the termin
  85. 3:45terminology interchangeably between
  86. 3:47synapses weights and parameters so if
  87. 3:50you heard these terms they mean the same
  88. 3:53thing and the neurons features and
  89. 3:56activations we use these terminologies
  90. 3:58interchangeably they also mean the same
  91. 4:01thing so in this example we have uh two
  92. 4:05layer neuron Network and get output so
  93. 4:09the dimensionality of these hidden
  94. 4:11layers determines the width of the model
  95. 4:14right so the width of the model matters
  96. 4:17a lot to the efficiency of neural
  97. 4:20networks so we can design a wide but
  98. 4:23shallow network with the same number of
  99. 4:26parameter as um um narrow but deep
  100. 4:30neuron Network so I guess which one is
  101. 4:33more accurate and which one is more
  102. 4:36Hardware efficient in these two
  103. 4:38scenarios one is a very deep but niral
  104. 4:43neuron Network the other is a shallow
  105. 4:46but wide neuron Network which one is
  106. 4:48more Hardware
  107. 4:52friendly wide and shallow and why is
  108. 4:55that the case
  109. 5:02exactly fewer number of crial cause and
  110. 5:05I have more paradism to fully exhaust
  111. 5:08the number of threads in the gpus uh so
  112. 5:11from the efficiency perspective a wide
  113. 5:14and shallow Network might be more
  114. 5:16efficient uh in the meantime from the
  115. 5:18accuracy perspective from the accuracy
  116. 5:21perspective usually a deep Network um
  117. 5:24matter helps a lot so there is a
  118. 5:26tradeoff you don't get a free Lance on
  119. 5:28both side so that's the beauty of neuron
  120. 5:30Network architecture design where we're
  121. 5:32going to visit in the later part of this
  122. 5:34lecture How do trade off between this
  123. 5:36hard Hardware efficiency and also the
  124. 5:39accuracy so let's visit a few popular
  125. 5:41neuron Network layers starting from the
  126. 5:44fully connected layer also called a
  127. 5:47linear layer so in AAR layer um you have
  128. 5:50a weight Matrix which is of Dimension CI
  129. 5:54input Channel times Co which is the
  130. 5:57output Channel Okay so um the number of
  131. 6:00parameters here is CI times Co okay and
  132. 6:04this is the input activation CI and this
  133. 6:06is the output activation
  134. 6:08Co so uh that's the input feature
  135. 6:11Dimension usually um when we are talking
  136. 6:14about efficient Computing is super
  137. 6:16important to understand the dimension of
  138. 6:19each tensor okay later when we are going
  139. 6:22to build uh the model to to calculate
  140. 6:25the flops number of operations and also
  141. 6:28the model the parameter size the model
  142. 6:30size it all concerns about the dimension
  143. 6:33of the of each tensor so it's very
  144. 6:36crucial to figure out the dimension of
  145. 6:38the input feature output feature map
  146. 6:41weight feature and also the bias
  147. 6:44especially the bias when we are going to
  148. 6:46learn about the uh efficient on device
  149. 6:48training we are having a technique
  150. 6:51called bius only update or recently the
  151. 6:53low which concerns about the bias so
  152. 6:55we're going to revisit that in later
  153. 6:57part of the lecture so so very simple
  154. 7:00output equal to the weight is sum of the
  155. 7:02input so that's the fully connected
  156. 7:04layer also called a linear
  157. 7:06layer what do we if we have multiple
  158. 7:09inputs so here we have a batch size of
  159. 7:12two we have two inputs right so the
  160. 7:15weight Dimension stays the same it's CI
  161. 7:18* Co but the input becomes n times CI
  162. 7:22and the output feature map becomes n
  163. 7:24times
  164. 7:26Z convolution layer so so um the output
  165. 7:31neuron is no longer connected to all the
  166. 7:33inputs but is connected only to the
  167. 7:36inputs within the receptive field okay
  168. 7:39so we have a input input feature and
  169. 7:43this is the spatial Dimension now we
  170. 7:45just assume it's 1D convolution and
  171. 7:47later let's jump into 2D in this 1D
  172. 7:50convolution um this is the channel
  173. 7:52Dimension this is the spatial 1D spatial
  174. 7:55Dimension and we apply this green filter
  175. 7:59here we also have the input Channel
  176. 8:01Dimension CI and here we have um the KW
  177. 8:06which is the uh a filter Dimension and
  178. 8:09then we have multiple such filters okay
  179. 8:12each filter is going to convolve with
  180. 8:14the input and produce one
  181. 8:18output and we can shift it shift this
  182. 8:21sliding window by one to get another
  183. 8:24output we can shift it again to get
  184. 8:26another output so as a result the output
  185. 8:30feature um Dimension output Channel
  186. 8:33Dimension is three since we have three
  187. 8:36uh kernels for the
  188. 8:38weight and the weight Dimension is
  189. 8:40basically CI * K time Co since each
  190. 8:44kernel Dimension we have a CI * K and we
  191. 8:47have C this amount
  192. 8:48of
  193. 8:51kernels and finally we have a bias with
  194. 8:54the dimension of
  195. 8:56Co and now let's generalize that to uh
  196. 9:00two dimensional convolution 2D C okay as
  197. 9:04we can see we added another dimension uh
  198. 9:07it's 2D convolution but looks like the
  199. 9:10feature map is looks like the feature
  200. 9:12map is threedimensional
  201. 9:13that's uh because for each XY location
  202. 9:18we have a channel Dimension okay the
  203. 9:20channel Dimension so the for 2D
  204. 9:23convolution the feature map is actually
  205. 9:263D so the activation map um the H times
  206. 9:30W assuming uh imagine you have a picture
  207. 9:34uh width is W height is H and then we
  208. 9:38apply a filter which is KH by KW which
  209. 9:42is the filter um dimension for the h and
  210. 9:45W Dimension with the same input Channel
  211. 9:48Dimension CI and then we have multiple
  212. 9:51such filters three such filters which is
  213. 9:54equal to Co in this case and we can
  214. 9:57shift it by one to get a another output
  215. 10:01shifted by two to get the another output
  216. 10:05we can shift in another
  217. 10:06dimension on and so on so get the output
  218. 10:09feature map 3x3 output
  219. 10:13feature and how do we calculate the uh
  220. 10:17size of the output feature map is equal
  221. 10:19to um the size of the input minus the
  222. 10:22size of the kernel plus one and why plus
  223. 10:25one because you can shift uh KH minus
  224. 10:29one this amount of shifts and applying
  225. 10:33each shift you get a one more output so
  226. 10:36H H equal to hi minus KH plus one in
  227. 10:40this case um the hi equal to four the
  228. 10:44input is
  229. 10:454x4 um and then the K the cernal size is
  230. 10:49three therefore the output Dimension is
  231. 10:514 minus 3 + 1 it'll be two in this case
  232. 11:00let's also talk about padding okay so
  233. 11:02padding can be used to keep the output
  234. 11:04feature map the same as the input
  235. 11:06feature map otherwise using convolution
  236. 11:09um the feature map will get smaller and
  237. 11:11smaller as you have a deeper number of
  238. 11:13layers so one common way is to do uh to
  239. 11:17apply zero padding pad the input
  240. 11:19boundaries with the zero which is the
  241. 11:22default um in py torch there's for
  242. 11:25example in this case uh the original um
  243. 11:29I and width equal to five in this
  244. 11:32animation kernel size is three so um the
  245. 11:38output height and weight equal to 5 + 2
  246. 11:42* 1 2 p okay padding equal to one we pad
  247. 11:46one on each side and minus 3 plus one
  248. 11:49equal to five so the output feature map
  249. 11:51is the same as the input compared with
  250. 11:54the initial version like in here um the
  251. 11:58output
  252. 11:59becomes uh becomes smaller the input is
  253. 12:02three 4x4 but output becomes two um as
  254. 12:06we can see on
  255. 12:08the on the example here right initially
  256. 12:11the input is 4x4 output becomes two and
  257. 12:15after we apply padding after we apply
  258. 12:18padding the input and output are both
  259. 12:21five by five in this
  260. 12:23case so we have we can have zero pading
  261. 12:26there are also other methods to apply
  262. 12:28pading for example using a reflection
  263. 12:32pading okay so for example right here
  264. 12:36the um um padded number is the reflected
  265. 12:41value of the original feature map we can
  266. 12:43also apply the replication padding to
  267. 12:45use the closest number and replicate
  268. 12:48that to do to apply the padding but zero
  269. 12:51padding is still the widely used
  270. 12:54approach so after padding let's now talk
  271. 12:57about the receptive field for example
  272. 13:00the output pixel here how many pixels
  273. 13:03Can it can it see in the previous layer
  274. 13:06it can see 3x3 pixels if the kernel size
  275. 13:10is 3x3 right and each pixel here is
  276. 13:13another 3x3 so how many pixels from here
  277. 13:17to here actually 5 by five right and one
  278. 13:21more layer you get S by seven receptive
  279. 13:23field why do we need a large receptive
  280. 13:26field we want to understand the relation
  281. 13:29ship between different
  282. 13:31pixels for example a car on the road
  283. 13:35versus some pedestrian also on the road
  284. 13:37the relationship between the car and
  285. 13:39pedestrian I means whether you can drive
  286. 13:42or you can you should avoid pedestrian
  287. 13:45right so the relationship matters for
  288. 13:47this high level Vision tasks and why
  289. 13:49people youed Transformers later because
  290. 13:52Transformer give you a um Global
  291. 13:55receptive field you can see every pixel
  292. 13:57can attend to every pixel Invision
  293. 13:59Transformers we're going to cover later
  294. 14:01so the principle is that a larger
  295. 14:03receptive field helps a lot it's very
  296. 14:06crucial uh to high level the image
  297. 14:09understanding so with two layers three
  298. 14:12kernels we can see receptive field um of
  299. 14:155x five and in general the equation is
  300. 14:18here with L layers the receptive field
  301. 14:22size equal to l * kernel size minus one
  302. 14:25+
  303. 14:27one uh so for two layers kernel size
  304. 14:30equal to three kernel the receptive
  305. 14:32field is
  306. 14:33five and you have three layers kernel
  307. 14:36size is is three the receptive field is
  308. 14:39is
  309. 14:41seven so the problem is that for large
  310. 14:44images in order to have a large
  311. 14:46receptive field we have to have very
  312. 14:49deep layers we need many many layers and
  313. 14:52we just talk about that in the early
  314. 14:54part of the lecture if you have too many
  315. 14:56layers you have lots of Kernel CA you
  316. 14:59also have to store a lot of activations
  317. 15:02uh activations has to be stored during
  318. 15:05back propagation increasing your
  319. 15:07training memory so how do we have a
  320. 15:10large receptive field at the same time
  321. 15:12we don't want to have so many
  322. 15:15layers so we can done sample inside the
  323. 15:18neuron Network one example is to use the
  324. 15:21strided convolution layer
  325. 15:24okay so compared with Str equal to one
  326. 15:27which is on the bottom
  327. 15:29so on the top shows the St equal to two
  328. 15:33Okay so rather than every pixel sees the
  329. 15:36adjacent 3x3 pixels now is see 5x five
  330. 15:40okay the adjacent green um adjacent
  331. 15:44green 3x3 re uh rectangle is shifted by
  332. 15:49not one pixel but two pixels okay you
  333. 15:52can see this 3x3 versus this 3x3 is
  334. 15:56actually shifted uh by by two so that's
  335. 15:59a stride of two every time you don't
  336. 16:01move one step but you move two two
  337. 16:05steps so that's a the uh stred
  338. 16:09convolution in this case for two layers
  339. 16:12and kernel size equal to three the
  340. 16:14receptive field becom seven okay
  341. 16:16previously it was five now it was seven
  342. 16:18you have a larger receptive field with
  343. 16:21the same number of
  344. 16:23layers so this is very helpful for
  345. 16:26reducing the number of layers reducing
  346. 16:28the number of weights reducing the
  347. 16:29number of activations but have the same
  348. 16:32receptive field but of course you may
  349. 16:34lose a bit accuracy because the number
  350. 16:36the model capacity might be
  351. 16:38smaller but everything is tradeoff
  352. 16:40there's no no free lunch in deep neural
  353. 16:43network design everything is about
  354. 16:45constraint optimization which makes the
  355. 16:47optimization highly
  356. 16:51interesting okay so our goal is to
  357. 16:55continue reducing the number of
  358. 16:56parameters okay so for convolution layer
  359. 16:59each output channel is connected to all
  360. 17:02the input channel to all the input
  361. 17:03Channel okay how to reduce the amount of
  362. 17:07computation people invented this group
  363. 17:10convolution layer okay so previously we
  364. 17:13have just one group every channel is
  365. 17:15output channel is connect to all the
  366. 17:16input Channel now we can divide them
  367. 17:19into two groups so for the first group
  368. 17:22only half of the input channel is
  369. 17:25connected and for the second half um
  370. 17:28that's the second group okay so
  371. 17:31effectively if we have G uh groups okay
  372. 17:34we can see uh the feature Map size
  373. 17:38doesn't input feature Map size and
  374. 17:39output feature Map size doesn't change
  375. 17:42stays the same as before but the number
  376. 17:44of weight will be G times smaller
  377. 17:49okay will be G times smaller So Co will
  378. 17:52be divided by G and CI will also be
  379. 17:55divided by G and we have G amount of uh
  380. 17:59such such filters okay therefore the
  381. 18:02total amount of Weights is reduced by G
  382. 18:06okay so that's the group convolution
  383. 18:10here and what is the extreme of group
  384. 18:12convolution the number of group the
  385. 18:15number of group will be equal to the
  386. 18:19number of input or output Channel okay
  387. 18:21so which is exactly uh this case
  388. 18:25previously we have two groups now we
  389. 18:27have Co groups in this case uh it's it's
  390. 18:31eight right we have eight groups so each
  391. 18:34CH output channel is only connected to
  392. 18:37one input Channel okay so the number of
  393. 18:41number of way uh number of uh input
  394. 18:44Feature Feature stays the same and the
  395. 18:46number of Weights will be C which is the
  396. 18:49same for c i and Co so we just use c
  397. 18:52times uh KH and KW which is the extreme
  398. 18:55case of group convolution called deps
  399. 18:59wise convolution layer which is quite
  400. 19:01widely used since 2015 2016 where uh
  401. 19:06when mobile net was
  402. 19:11invented the next um ping layer we want
  403. 19:14to have a smaller feature map originally
  404. 19:17the input image might be like 1K uh by
  405. 19:21by 1K is pretty large and we want to
  406. 19:24have a smaller feature map and condensed
  407. 19:26information right um this is actually
  408. 19:29quite useful for large resolution uh
  409. 19:32input we want to quickly pull the
  410. 19:35feature map to make it smaller so here
  411. 19:38we have an example with the 4x4 input
  412. 19:41feature map and how do we put it to 2x2
  413. 19:45we can either use max putting so for
  414. 19:47every 4x4 input we find the max element
  415. 19:51within that group um or we can do uh
  416. 19:55this kind of average pting finding the
  417. 19:57average in the Blue Area finding the
  418. 20:00average for the yellow area to do the
  419. 20:03average
  420. 20:04pulling the good thing about ping is
  421. 20:07that there's no learnable parameters
  422. 20:10okay or you can assume um yeah so but
  423. 20:14contain zero weight uh which is very
  424. 20:17parameter
  425. 20:19efficient next let's talk about
  426. 20:22normalization year so normal bch
  427. 20:23normalization came roughly in 2015 or or
  428. 20:27new 15 or or 16 roughly seven eight
  429. 20:31years ago uh which later becomes quite
  430. 20:34useful to stabilize the training for
  431. 20:36example the the Leer Norm also become
  432. 20:38super helpful in the age of uh uh
  433. 20:41attention Transformers so what exactly
  434. 20:44is
  435. 20:45normalization so we want to normalize
  436. 20:48the feature map to make the optimization
  437. 20:52faster okay so how do we apply that so
  438. 20:55we minus the feature map minus the mean
  439. 20:58of that tensor okay and divided by the
  440. 21:01sigma okay so we want to make sure um we
  441. 21:05want to minus the mean and divided by
  442. 21:07the standard deviation over a set of
  443. 21:10pixels or tensors if it is not Vision
  444. 21:13okay so this is a reminder how to uh
  445. 21:18calculate the mean and the standard
  446. 21:21deviation and note here we have a small
  447. 21:25Sigma right here okay so that is due who
  448. 21:28we want to avoid dividing by zero so we
  449. 21:32put a very small number over
  450. 21:34there and then we learn a per Channel
  451. 21:38linear transformation okay uh which is
  452. 21:42indicated by the scaling factor and also
  453. 21:44the bias right here to compensate for
  454. 21:47the um possible loss of represent
  455. 21:50representational ability
  456. 21:53okay so how do we Define this set of
  457. 21:56pixels or set of tensor
  458. 21:59on the bottom it shows four different
  459. 22:02ways to find uh that group of pixels uh
  460. 22:07the first one very widely used in CNN
  461. 22:11called batch
  462. 22:12normalization okay so it's normalizing
  463. 22:15across the N Dimension n is the image
  464. 22:18Dimension okay and also we just use a
  465. 22:21single C Dimension across all the H and
  466. 22:24W height and width dimension okay
  467. 22:27normalize across
  468. 22:29different batches second one is quite
  469. 22:32widely used in the attention mechanism
  470. 22:35so for different input for each input
  471. 22:38okay across different um H and W and C C
  472. 22:42is a channel Dimension make sure it's
  473. 22:45normalized so uh for each token in the
  474. 22:49attention mechanism when you are doing
  475. 22:51the attention after each token is
  476. 22:53normalized then the attention map uh
  477. 22:56makes sense otherwise uh um if it's not
  478. 22:59nor normalized the tension map will be
  479. 23:02um much harder to
  480. 23:05optimize there's also the instance
  481. 23:08normalization it just normalized across
  482. 23:10the H and W Dimension not the C
  483. 23:13Dimension and when the C is too large
  484. 23:16there's also the group normalization on
  485. 23:18the right uh which divide the C by
  486. 23:20several groups and only normalize within
  487. 23:23that
  488. 23:24group there's a lot of tricks we can
  489. 23:27play for the normalization layer for
  490. 23:30example when we are doing
  491. 23:32fing um actually normalization layer the
  492. 23:35number of parameter is very small so we
  493. 23:37can only fine tune the weight the bias
  494. 23:39and also the um the scaling factor in
  495. 23:43the normalization layer making the fine
  496. 23:44tuning much more computation primary
  497. 23:48efficient the normalization layer can
  498. 23:50also absorb um the transformation of
  499. 23:54such a smooth quantization so without
  500. 23:57having incurving another peral call we
  501. 23:59can absorb lots of the computation in
  502. 24:02the IP log or fuse it with the
  503. 24:04normalization later so we can we are
  504. 24:06going to visit them in later part of
  505. 24:07this
  506. 24:10lecture okay next one is activation
  507. 24:13function so uh we have seen many
  508. 24:16different activation functions but let
  509. 24:18me talk about the efficiency and
  510. 24:19accuracy trade off here sigmoid very
  511. 24:23ancient um activation function the it
  512. 24:26ranges between zero and one
  513. 24:28is very easy to quantize very efficient
  514. 24:31to quantize because the dynamic range is
  515. 24:34very limited and this fixed between zero
  516. 24:36and one but what is what is the
  517. 24:40drawback gring vanishes right so if the
  518. 24:43value super small or super large there's
  519. 24:46no
  520. 24:47gradient and then people come up with Ru
  521. 24:50okay when it's positive the gradient
  522. 24:52never vanishes but if it's when it's
  523. 24:55negative the neuron is basically dead
  524. 24:57there's no way to put it back because
  525. 25:00there's no gradient anywhere when the
  526. 25:03value is smaller than
  527. 25:05zero but what the good what is good
  528. 25:07about it is that it's very easy to
  529. 25:10sparsify since we're going to talk about
  530. 25:12prud sparf and sparcity in the later
  531. 25:15part of this lecture uh the ru
  532. 25:18activation function naturally brings
  533. 25:20sparcity because lot of the activation
  534. 25:23once they are act once they are negative
  535. 25:25they become zero and zero multiply but
  536. 25:28by anything is zero so don't have to
  537. 25:30compute on it you don't even to store it
  538. 25:33you can just store a mask showing that
  539. 25:35this is zero just using one
  540. 25:39bit and the gradient is one um
  541. 25:42everywhere else okay so you don't even
  542. 25:45need to save this feature map you can
  543. 25:47just save a bit mask um for this for
  544. 25:51that tensor if it is the bit MK is zero
  545. 25:54then the output is zero if the back bit
  546. 25:56MK is one it will just copy the feature
  547. 25:58map from the previous output so very
  548. 26:02activation friendly sparcity friendly
  549. 26:05and
  550. 26:07simple but what is the downside the
  551. 26:10dynamic range is super large it's not
  552. 26:13easy to quantize so peopleand it R six
  553. 26:16okay R six to cap the larest value to be
  554. 26:20six that's why people call it R six okay
  555. 26:23it's purely
  556. 26:24empirical and that makes the
  557. 26:26quantization range much more constraint
  558. 26:29reducing the Dy dynamic range make it
  559. 26:31easier to
  560. 26:33quantize another drawback is that
  561. 26:35there's no gradient in the negative
  562. 26:37region so people imaged Le R okay
  563. 26:40there's a smaller scope slope on the
  564. 26:44negative side okay so um you can also
  565. 26:47like the gradient have the gradient flow
  566. 26:49in the negative region but the down side
  567. 26:52is that you no no longer have this
  568. 26:54sparcity benefit if the value is small
  569. 26:57it's negative it's no longer it's no
  570. 27:00longer zero
  571. 27:02okay and then people are using neural
  572. 27:04architecture search find a very
  573. 27:06interesting activation function called a
  574. 27:08switch okay switch basically very
  575. 27:11similar to Ru but it's using x / 1 plus
  576. 27:16e to the minus of X but people later
  577. 27:18find it hard to implement in Hardware
  578. 27:21right so people come up with a
  579. 27:23approximated switch called hard switch
  580. 27:27okay so um when X is smaller than minus
  581. 27:31three is zero above three is just X in
  582. 27:35between that um the value is x + 3
  583. 27:39divided by 6 okay so all this activation
  584. 27:44function uh somehow have something to do
  585. 27:46with efficiency like sparcity
  586. 27:49quantization dynamic range easy to
  587. 27:51implement so that's the story behind uh
  588. 27:54the progression of different activation
  589. 27:57functions
  590. 28:00and finally Transformers Transformer are
  591. 28:02getting super important since the uh
  592. 28:06invention in 2017
  593. 28:082018 uh we learn more about Transformer
  594. 28:10architecture in lecture 12 we have a
  595. 28:12dedicated lecture just to talk about
  596. 28:14Transformers the entire uh pipeline
  597. 28:18entire structure of Transformers so
  598. 28:21let's just give a very simple overview
  599. 28:24uh
  600. 28:24today so there are two stages encoding
  601. 28:27stage and the decoding stage for each
  602. 28:30stage we have a mod tension followed by
  603. 28:34f f forward Network and this is the
  604. 28:38architecture for the attention mechanism
  605. 28:41you have three tensors q k v quy key and
  606. 28:44value we do a um
  607. 28:49attention scale dot product attention
  608. 28:52and then pass it through a linear layer
  609. 28:54get the o q k v o those are the four um
  610. 28:58activation tensors and there's no weight
  611. 29:00in the attention mechanism uh that's the
  612. 29:04in the middle part only wait in the qk
  613. 29:07vi transformation okay and the detail
  614. 29:10way to do that for the query key value
  615. 29:12the design is actually very analogous to
  616. 29:15the retrieval system just take a YouTube
  617. 29:17search as example the query is like the
  618. 29:21text prompt in the search bar okay and
  619. 29:24the key would be like the title or
  620. 29:26description some short information about
  621. 29:28the video and value would be the
  622. 29:31corresponding video so you do a DOT
  623. 29:34product with the query and the key to
  624. 29:35get the similarity and use the
  625. 29:37similarity to fetch to do a weighted
  626. 29:40average of the
  627. 29:42value and remember we want to divide the
  628. 29:46Q and K by a normalization factor which
  629. 29:50is square root of d uh to accommodate
  630. 29:53for the dimension change to make the um
  631. 29:57optim ization more stable and followed
  632. 30:00by the soft Max to get a um a tension
  633. 30:04map a n byn attention map and use that
  634. 30:07attention map to do a a weighted average
  635. 30:10weighted sum of the of the value okay
  636. 30:14and later part of the lecture we are
  637. 30:16going to visit in the tension map the N
  638. 30:18byn tension map could be the bottom NE
  639. 30:22for both computation and the uh other
  640. 30:26the memory since this is growing
  641. 30:29quadratically with the number of token
  642. 30:31in okay and for images the number of
  643. 30:34token also grow quadratically with the
  644. 30:36resolution so n Square would be super uh
  645. 30:39expensive so there's sparity opportunity
  646. 30:42not all attention not all token need to
  647. 30:44attend attend to each other sparse
  648. 30:46attention we're going to introduce that
  649. 30:49uh and also flash attention we can use
  650. 30:51in place um attention um to minimize the
  651. 30:55data movement uh so that we can have a
  652. 30:58constant almost constant memory but the
  653. 31:01flops is not not reduced so in order to
  654. 31:04deal with that problem we are going to
  655. 31:06introduce long context techniques all
  656. 31:09related to this attention mechanism so
  657. 31:12although attention itself doesn't have
  658. 31:14any weights there's a lot of interesting
  659. 31:16um interesting stuff we can explore even
  660. 31:19the KB cache um they become the
  661. 31:23activation when we are calculating the
  662. 31:25KV cache but when we are using that to
  663. 31:27generate the next token they become the
  664. 31:29weight to to when new tokens is
  665. 31:32generated so how do we quantize how do
  666. 31:34we uh Pro the KV cache we're also going
  667. 31:37to cover that uh in later part of the
  668. 31:40lecture let's continue the lecture to
  669. 31:43talk about the efficiency metrix how
  670. 31:46should we measure the efficiency of
  671. 31:48neuron
  672. 31:51networks so there are several pillars we
  673. 31:54want to achieve at the same time but
  674. 31:55it's pretty hard having a smaller model
  675. 31:58so when we up when you're uploading the
  676. 32:00model to the Apple Store it's pretty
  677. 32:02small right you don't want to download a
  678. 32:03model that is hundreds of gigabytes that
  679. 32:05is too hard to
  680. 32:06download and also we want to have a
  681. 32:09faster model it give you the real time
  682. 32:11prediction never miss a pedestrian the
  683. 32:13road for self-driving car and also
  684. 32:15generate images instantly on your phone
  685. 32:17for
  686. 32:18example and also Greener model taking
  687. 32:21less battery for example on the iPhone
  688. 32:23on the mobile devices you you don't want
  689. 32:25to join your battery with this AI
  690. 32:26applications
  691. 32:28and that concerns theu right and
  692. 32:31computation the memory is the
  693. 32:32determination is determine determining
  694. 32:35the storage latency and energy so we are
  695. 32:38going to introduce several efficiency
  696. 32:41matrics on the right hand side um some
  697. 32:44of them are memory related some of them
  698. 32:46are computation related on the memory
  699. 32:49side we'll introduce how to count the
  700. 32:51number of parameters um the model size
  701. 32:54the total activation size and Peak
  702. 32:57activation size size for the computation
  703. 32:59related we're going to talk about what
  704. 33:01is Mac what is flop what is flops what
  705. 33:04is op what is
  706. 33:06OPS so let's first talk about latency so
  707. 33:09what is latency latency measures the
  708. 33:12delay for specific task and let me
  709. 33:15directly show you a example so on the
  710. 33:17left hand side it's high latency right
  711. 33:20each to predict the segmentation mask of
  712. 33:25output takes a long time and right hand
  713. 33:28side is lower latency uh only 45 46
  714. 33:32milliseconds to process each
  715. 33:35frame what is throughput it measures a
  716. 33:38rate at which data is processed on the
  717. 33:41left hand side is low throughput you can
  718. 33:43process six videos per second on the
  719. 33:46right hand side is high throughput 77
  720. 33:48videos per second this is from the
  721. 33:50temporal shift module paper for video
  722. 33:52understanding we are going to introduce
  723. 33:54later so what is the relationship
  724. 33:57between latency and throughput does
  725. 34:00higher throughput translate to lower
  726. 34:01latency does lower latency translate to
  727. 34:04higher throughput so that's see example
  728. 34:07the first design the Laten is 50
  729. 34:11milliseconds so every 50 millisecond we
  730. 34:13can process my image uh so how many
  731. 34:16images can we process every second is 20
  732. 34:20right on the design to we have higher
  733. 34:23pism we can process four Images at the
  734. 34:26same time each taking 100 millisecond so
  735. 34:29latency to process each image from start
  736. 34:32to end is 100
  737. 34:34millisecond and the throughput would be
  738. 34:3640 images um per second since we are
  739. 34:40processing uh four such image at the
  740. 34:43same
  741. 34:45time so on the left hand side is having
  742. 34:48a lower latency a lower throughput right
  743. 34:51hand side is having a higher latency but
  744. 34:53higher throughput as a result higher
  745. 34:55throughput doesn't translate to lower
  746. 34:57latency
  747. 34:58and lower latency also doesn't translate
  748. 35:00to higher throughput on the mobile we
  749. 35:03sometimes care about the we mostly care
  750. 35:05about the latency on the data center on
  751. 35:08batch processing tasks we care about the
  752. 35:09throughput so those are very important
  753. 35:12metrics latency and throughput sometimes
  754. 35:15we we want to make sure we report
  755. 35:20both so what is the uh factors that
  756. 35:23impact the latency so um there are two
  757. 35:27two factors that is determining the
  758. 35:30latency one is the computation one is
  759. 35:32the memory okay um the T of computation
  760. 35:35equal to the number of operations in a
  761. 35:37neural network model divided by the
  762. 35:39number of operations that the process
  763. 35:42can process per
  764. 35:45second and the bottom is the hardware
  765. 35:47specific uh characteristic on the top is
  766. 35:51Neuron Network specific and then T
  767. 35:53memory also determine is determined by
  768. 35:56two factors the activation
  769. 35:58and also the weight to move the
  770. 36:00activation and move the weight so this
  771. 36:02is about inference but it's training
  772. 36:03also concerns about communication of the
  773. 36:06gradient so T data movement of weight
  774. 36:09equal to the the weight the model size
  775. 36:11divided by the memory bandwidth Hardware
  776. 36:14specific neuron Network
  777. 36:16specific and T data movement of
  778. 36:18activations is determined by the input
  779. 36:21activation sides plus the output
  780. 36:23activation sides this is first order all
  781. 36:25these are first order approximations
  782. 36:27give you a ballpark estimation divided
  783. 36:30by the memory bandwidth of the
  784. 36:33processor so which one is more expensive
  785. 36:36so this is a figure showing the energy
  786. 36:40consumption um of different operations
  787. 36:43from a 32bit integer add to accessing a
  788. 36:48register file to a
  789. 36:50multiplication to accessing the SRAM
  790. 36:53cach to access the dram
  791. 36:55cache to access the D memory we can see
  792. 36:58that access in the D memory could be
  793. 37:00super expensive okay 640 PJ versus just
  794. 37:05doing a 32-bit add it's just 0.1 P so
  795. 37:10Accents in the memory is a lot more
  796. 37:12expensive than doing the
  797. 37:14arithmetic uh like two others of
  798. 37:16magnitude more energy is consumed by
  799. 37:20access in the memory so keep in mind
  800. 37:22computation is cheap memory data
  801. 37:25movement is very expensive
  802. 37:28so it is the data movement that is
  803. 37:30joining uh joining the battery
  804. 37:34okay so let's talk about how to
  805. 37:36calculate the number of parameters since
  806. 37:39data movement is
  807. 37:41expensive so um parameter is the number
  808. 37:44of weight in a neuron Network okay for
  809. 37:48example in the linear layer the weight
  810. 37:51tensor is C in times C out so the number
  811. 37:55of parameter is just C C in time C out
  812. 37:59in a convolution layer the weight tensor
  813. 38:04is four dimensional okay each kernel has
  814. 38:07KH KW time C in and we have Co Am Co
  815. 38:12number of kernels so all together we
  816. 38:14have a multiplying these four terms
  817. 38:17together to get the total number of
  818. 38:19parameters for a convolution
  819. 38:23layer for a group convolution okay we
  820. 38:26divided the CI and Co by G so CI is
  821. 38:31divided by G Co is divided by G but we
  822. 38:34have G amount of such kernels so Al
  823. 38:38together we have only one g on the
  824. 38:41denominator so G times smaller compared
  825. 38:43with the convolution with the same
  826. 38:46dimension for depthwise convolution
  827. 38:49where G equal to c i equal to co okay
  828. 38:53therefore the total number of parameter
  829. 38:54equal to co * KH * KW
  830. 39:01and let's see the PO the number of
  831. 39:02parameters for popular neural networks
  832. 39:04forx net right so let me do the
  833. 39:08calculation for the first layer uh where
  834. 39:10the channel number is uh input channel
  835. 39:13is three output channel is 96 and the
  836. 39:16kernel size is 11 by 11 so the first
  837. 39:19layer the number of parameter is 96
  838. 39:22output three input 11 by 11 okay and
  839. 39:26similarly we can calate the second layer
  840. 39:28it's a group convolution so we divide
  841. 39:31the number of parameters by by two since
  842. 39:33group size is
  843. 39:36two so I give you one minute to
  844. 39:38calculate to think about the the layer
  845. 39:41number of parameters for the next layer
  846. 39:43okay here 13 by 13 um image size and 3x3
  847. 39:52convolution the channel number is two
  848. 39:55384
  849. 39:59so here's a quick answer since the input
  850. 40:01channel is 2 256 output is 384 and the
  851. 40:05channel number uh colal size is 3x3 this
  852. 40:08is the way to calculate that and finally
  853. 40:10for the fully connected layer you just
  854. 40:13multiply the input Channel with the
  855. 40:14output channel to calculate the total
  856. 40:16number of Weights which is 61 meaning in
  857. 40:20total so how does that translate to the
  858. 40:23model size between the number of
  859. 40:25parameters to model size what's the Rel
  860. 40:27relationship you want to multiply the
  861. 40:30number of bytes for each parameter with
  862. 40:33the number of parameters okay so number
  863. 40:36of parameters with multiplied with bit
  864. 40:38width look get a model size so for
  865. 40:42example alxn has 61 million parameters
  866. 40:45if each weight is in uh is represented
  867. 40:48by the um single Precision 32 bit 32-bit
  868. 40:53number and how many bytes are there in a
  869. 40:5732
  870. 40:59bit four bytes right 32 bits equal to
  871. 41:03four bytes so 61 million * 4 you have
  872. 41:06224
  873. 41:08megabytes so make sure you have the
  874. 41:10prerequisite so that you can understand
  875. 41:11these Concepts relatively easily how
  876. 41:14many bits per bite and you want to
  877. 41:16quantize the weight to uh integer right
  878. 41:21to 8bit integer what is the total
  879. 41:23storage for the
  880. 41:25weight so 8 bit that equal to one byte
  881. 41:2861 million parameters that's 61
  882. 41:31megabyte 61 megabyte four times smaller
  883. 41:34right compared with 32 bit and later
  884. 41:38part of this lecture we are going to
  885. 41:40learn about activation where with only
  886. 41:42quation awq to quantize it to four bit
  887. 41:46okay four bit will be another two times
  888. 41:48smaller so that it can fit large
  889. 41:50language model locally on your
  890. 41:54laptop okay so now let's switch gear to
  891. 41:56talk about the total and the peak number
  892. 41:59of
  893. 42:01activations so activation is the bottom
  894. 42:04neck for seeing inference not the
  895. 42:06parameters for example here from reset
  896. 42:09to mobile net number of parameters is
  897. 42:12reduce a lot by four times more than
  898. 42:15four times but the activation the peak
  899. 42:18activation didn't decrease but increased
  900. 42:20so um for how to reduce the activation
  901. 42:24activation is very crucial the deeper
  902. 42:26the layer the more more number of
  903. 42:27activations you that's the
  904. 42:31case and also um the peak activation
  905. 42:34determines the memory bottle neck um for
  906. 42:38example uh the third layer here
  907. 42:41determines what is the max memory you're
  908. 42:43are going to consume when you are doing
  909. 42:46the
  910. 42:47inference so the imbalanced memory
  911. 42:49distribution happens very popular for CN
  912. 42:53why is that the case because the first
  913. 42:55couple of layer you have very high
  914. 42:56resolution
  915. 42:58and then square and later you have a
  916. 43:00lower resolution due to down
  917. 43:03sample for example if a microcontroller
  918. 43:06has only 256 kilobyte of ice Ram um in
  919. 43:10the
  920. 43:12uh of of SRAM to hold these activations
  921. 43:15you want to make sure uh the largest
  922. 43:18layer for example here doesn't exceed
  923. 43:21the SRAM so that you can put all the
  924. 43:23activations in the fast SRAM
  925. 43:28activation is also becoming the memory
  926. 43:30bottom neck in training in see for large
  927. 43:34language models the number of parameter
  928. 43:36is also the bottom neck that's why
  929. 43:38people later proposed uh this model
  930. 43:41paradism to do the sharding of the
  931. 43:43weight so here is example for the CN
  932. 43:46where um the this is showing the number
  933. 43:49of parameters versus the number of
  934. 43:52activations uh the number of activation
  935. 43:54is like seven times larger than the
  936. 43:56number of parameters when you are doing
  937. 43:58um training due to the batch the
  938. 44:01activation is pretty large occupying a
  939. 44:03lot of the uh GPU
  940. 44:06memories and here from resent 50 to
  941. 44:10mobile 9 V2 1.6 they pretty much have
  942. 44:13the same accuracy so it's apples apples
  943. 44:16comparison the parameter is reduced by
  944. 44:18four times but the main memory botom NE
  945. 44:22which is the activation didn't improve
  946. 44:23much only 1.1 times only 10% so uh
  947. 44:27imizing the activation is
  948. 44:29hard and usually for CN the distribution
  949. 44:32for the weight and activation um is
  950. 44:36different uh the yellow part is the
  951. 44:39activation um the blue part is the
  952. 44:42weight it's new shaped okay initially um
  953. 44:46the activation memory is large later the
  954. 44:48weight memory is large this is due to
  955. 44:50has high resolution um and this is due
  956. 44:54to it has larger amount of channels
  957. 44:59so this is calculating the number of
  958. 45:00activations of uh anx net so um C * H *
  959. 45:07W that's the way to calculate the
  960. 45:09activation size and if you add them
  961. 45:11together can see this is total amount of
  962. 45:13activations for LX
  963. 45:16net and the peak activation is roughly
  964. 45:19first order approximation number of
  965. 45:21input activation plus the number of
  966. 45:23output activation
  967. 45:27okay now let's switch gear to talk about
  968. 45:29um compute bonded Matrix for example Mac
  969. 45:34so what is Mac is not McDonald it's
  970. 45:37multiply and accumulate right multiply
  971. 45:39and accumulate so a equal to a plus b *
  972. 45:43C that's one Mac one accumulation and
  973. 45:46one multiplication it usually comes as a
  974. 45:49single instruction in
  975. 45:52gpus so for Matrix Vector multiplication
  976. 45:55how to calculate the number
  977. 45:58Max for example on the top right corner
  978. 46:00this is a Matrix Matrix Vector
  979. 46:05multiplication okay so we are going to
  980. 46:07produce M okay M outputs in order to
  981. 46:11produce each output we need to compute n
  982. 46:14multiplication as in order to produce
  983. 46:17each output we need to calculate an Max
  984. 46:21so all together we have M * n
  985. 46:24Max what what about a matrix Matrix
  986. 46:27modlic short for G MN General Matrix
  987. 46:31Matrix modlic on the right hand side we
  988. 46:34have a m by K Matrix times a k by n
  989. 46:38Matrix uh and produce a m by n output so
  990. 46:42we have M byn output in order to
  991. 46:44calculate each output we need K Max so
  992. 46:48all together we need M * n * K
  993. 46:54Max so here we are show we're showing
  994. 46:56the
  995. 46:57number of Macs for different layers for
  996. 46:59linear layer we just talk about that the
  997. 47:01Mac equal to the C in times C out assume
  998. 47:05batch size of one and for convolution is
  999. 47:08actually multiplying six terms together
  1000. 47:12okay so we are calculating this is the
  1001. 47:15amount of uh output feature output
  1002. 47:18features we have to calculate which is
  1003. 47:21um W Times W output H output times Co
  1004. 47:26okay and in order to compute in order to
  1005. 47:30compute each output pixel how many Max
  1006. 47:33do we need we need the entire kernel to
  1007. 47:36be the to be convolved so that's KW * K
  1008. 47:39* CI so multiplying them together we
  1009. 47:42have to multiply these six terms
  1010. 47:44together six three from here and three
  1011. 47:46from here okay that's the way to
  1012. 47:48calculate the convolution
  1013. 47:50max with group convolution okay um we
  1014. 47:54just divided by by G G is the number of
  1015. 47:56groups
  1016. 47:57and for depths wise convolution c i
  1017. 47:59equal to G so we canceled one term
  1018. 48:02multiplying these five terms together CI
  1019. 48:05gets cancelled due to G equal to
  1020. 48:10CI and this is calculating the max for
  1021. 48:13Alx net for example each layer
  1022. 48:16multiplying these six terms together if
  1023. 48:19it is group convolution divided by the
  1024. 48:21number of groups in this case two in
  1025. 48:24this case and finally for the last three
  1026. 48:27fully connected layers um it's just
  1027. 48:30input Channel times output channel two
  1028. 48:32terms multiplied together where
  1029. 48:35ultimately you have 724 mions Med Max in
  1030. 48:41total okay so now let's switch GE to
  1031. 48:44talk about flop and
  1032. 48:46flops so one
  1033. 48:49multiply is a floating Point operation
  1034. 48:52one add is also a floating Point
  1035. 48:53operation so flop is short for number of
  1036. 48:57floating Point operations floating Point
  1037. 49:01operations so how many M how many flops
  1038. 49:04are there in a
  1039. 49:05Mac a Mac has a multiply a multiply is a
  1040. 49:09fla a Mac also has a um an ADD and add
  1041. 49:13is also a fla so one multiply and
  1042. 49:17accumulate operation is equal to two
  1043. 49:20floating Point operations if the operant
  1044. 49:22are floating Point numbers for example
  1045. 49:25alexnet has 700 24 mimax the total
  1046. 49:28number of floating Point operations will
  1047. 49:30be 724 * 2 that's 1.4 Giga flops
  1048. 49:36okay and what is flops so the short for
  1049. 49:39floating Point operations per second the
  1050. 49:42number of flops divided by how many
  1051. 49:44second so that's a performance metric
  1052. 49:47showing how many how fast are you
  1053. 49:49processing these floating Point
  1054. 49:52numbers so what is the
  1055. 49:54op we don't always represent the numbers
  1056. 49:57in floating Point sometimes we represent
  1057. 50:00them in integer right so both floating
  1058. 50:03Point operation and integer operation
  1059. 50:04all kinds of in operation even bit mask
  1060. 50:08or even exore operation we can call it
  1061. 50:10the op okay to generalize the number of
  1062. 50:13operations is used to to measure the
  1063. 50:16amount of
  1064. 50:18computation uh for example if alexn has
  1065. 50:21724 Med max if each number is
  1066. 50:23represented by say binary or erary or
  1067. 50:27floating point or integer what you have
  1068. 50:30whatever you have any operation can be
  1069. 50:32counted as an ALT so 724 M Max will be
  1070. 50:37equal to
  1071. 50:381.4
  1072. 50:40gahs and similarly operations per second
  1073. 50:44Ops equal toal number of operations
  1074. 50:47divided by the number of seconds to
  1075. 50:48complete this operation which is the
  1076. 50:50performance uh speed
  1077. 50:55Matrix hey before jumping into the uh
  1078. 50:58lab zero tutorial let's give a short
  1079. 51:01summary of today's lecture review the
  1080. 51:03basics of neuron networks output
  1081. 51:05activation synapsis weight popular
  1082. 51:08layers into including FC layer
  1083. 51:10convolution layer types as convolution
  1084. 51:12putting normalization we also introduced
  1085. 51:15efficiency metrics parameters model size
  1086. 51:19activations Max flop flops up and UPS

About this transcript

This page contains the full transcript of EfficientML.ai Lecture 2 - Basics of Neural Networks (MIT 6.5940, Fall 2024, Zoom recording) by MIT HAN Lab, generated from the public captions YouTube serves with the video. The transcript has 7,322 words across 1,086 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.