YouTube2Text

EfficientML.ai Lecture 3 - Pruning and Sparsity Part I (MIT 6.5940, Fall 2024) — Transcript

by MIT HAN Lab · 7,658 words · 1,151 segments · language en · Watch on YouTube

Full transcript

  1. 0:01all right good afternoon everyone let's
  2. 0:03get started welcome to the efficient ml.
  3. 0:07today we are going to introduce lecture
  4. 0:08three about pruning and sparsity to
  5. 0:11accelerate the inference of neural
  6. 0:14networks so we'll have two lectures
  7. 0:17about pruni and sparcity so today we
  8. 0:19will have the first lecture about
  9. 0:21that so this is the part overview of the
  10. 0:24part one of this lecture about efficient
  11. 0:27difference we are going to cover four
  12. 0:30sections okay we are going to start
  13. 0:33with talking about the um here talking
  14. 0:37about pring and
  15. 0:38sparcity followed by quantization like
  16. 0:41integer quantization rp4 rp8 and also
  17. 0:45neural architecture search as the third
  18. 0:47part how to design efficient neural
  19. 0:49architectures even before you compress
  20. 0:52them and also we are going to conu that
  21. 0:54by knowledge distillation how to use a
  22. 0:57larger Network to distill to teach a
  23. 1:00smaller Network okay um so this is the
  24. 1:03agenda for uh the first part of this
  25. 1:06course about pruning we are going to
  26. 1:08conclude it by a real world example
  27. 1:12designing mcet and tiny tiny engine um
  28. 1:15to fit the neural network into a
  29. 1:17microcontroller by using all the
  30. 1:19techniques we learned uh in these four
  31. 1:22sections so today we are going to jump
  32. 1:25into the first part which is about
  33. 1:28pruning so before jumping into pring
  34. 1:30let's start with some uh motivation and
  35. 1:32background so when you first introduce
  36. 1:35this ml perf Benchmark which is the
  37. 1:38Olympic game for efficient air Computing
  38. 1:41for air Hardware okay so ml perf
  39. 1:43basically uh tests the performance in
  40. 1:47latency in throughput given the
  41. 1:49particular accuracy on The Suite of
  42. 1:51benchmarks there are two divisions the
  43. 1:54open Division and the Clos division so
  44. 1:56in the Clos division you cannot change
  45. 1:58neuron Network you can only apply
  46. 2:01quantization on the open division you're
  47. 2:03freid to change the neural network
  48. 2:05architecture and apply techniques such
  49. 2:07as pruning so this is the latest
  50. 2:09Benchmark from ndia Blackwell platform
  51. 2:13which was just a few weeks ago in late
  52. 2:16August so in the Clos division um the
  53. 2:19offline samples per second can reach
  54. 2:2444488 tokens per second this is Lama 2
  55. 2:2770b primer large language model Runing
  56. 2:30on a single Nvidia h200 GPU in the open
  57. 2:36division on the second column the L The
  58. 2:39throughput increased to more than 11,000
  59. 2:42tokens per second this is a big
  60. 2:45Improvement and how is that achieved and
  61. 2:48that is actually using the techniques we
  62. 2:49are going to learn in this lecture about
  63. 2:52pruning so uh this technique applied two
  64. 2:55pruning techniques one is St pruning
  65. 2:58reducing the number of layers from 8 to
  66. 3:0032 followed by the width pruning
  67. 3:03reducing the uh Channel Dimension from
  68. 3:0628,000 to 14 14,000 and as a result
  69. 3:10there's about two and a half X speed up
  70. 3:13running the at a good
  71. 3:17accy so hopefully that motivates uh the
  72. 3:20importance of pruning and from the
  73. 3:22hardware perspective why do we need prum
  74. 3:25so memory is very expensive computation
  75. 3:28is much cheaper uh a memory movement is
  76. 3:31more than two orders of magnitude than
  77. 3:33the arithmetic operations like we have
  78. 3:35seen in the previous lecture the 32 bit
  79. 3:38ad takes less than one p jeol while a
  80. 3:4132bit dam access Dam memory access is
  81. 3:45causing more than 600 PJs so uh data
  82. 3:49movement is much more expensive so to
  83. 3:52make deep burning more efficient we want
  84. 3:54to reduce the amount of memory reduce
  85. 3:56the model size reduce the activation
  86. 3:59size
  87. 4:01okay so with that as the motivation
  88. 4:03let's start with the agenda we are going
  89. 4:05to start by introducing what is pruning
  90. 4:08taking a dense neuron Network into a
  91. 4:10sparse neuron Network in general we can
  92. 4:13formulate pruning uh in this way so we
  93. 4:16want to minimize the loss of the prun
  94. 4:18model uh given the input X here the L
  95. 4:25indicates go back so here the L
  96. 4:28indicates the objective function for the
  97. 4:31neuron Network training we want to
  98. 4:33minimize this loss function when we have
  99. 4:36the pr weight WP okay subject to the
  100. 4:40number of non zeros okay the number of
  101. 4:43non zeros should be smaller than a
  102. 4:45threshold okay so we only want to have a
  103. 4:47limited amount of non zero elements in
  104. 4:50the neuron Network to make it sparse
  105. 4:53okay to make it a sparse and still
  106. 4:55minimize the
  107. 4:57loss and then we are going to introduce
  108. 5:00the pruning granularity in P grein
  109. 5:02versus fine grain and then the pring
  110. 5:05Criterion which neuron which synapsis to
  111. 5:07remove which neuron which synapsis to
  112. 5:09keep Follow by determining the prunin
  113. 5:12ratio what is the redundancy can we
  114. 5:15reduce a lot can we reduce just a little
  115. 5:17bit how do we select the protein ratio
  116. 5:20while maximizing the um amount of new
  117. 5:24connections that is PR and minimizing
  118. 5:26the loss of accuracy and finally we're
  119. 5:29going to talk about about how to fine
  120. 5:30tune how to Rin the network so that we
  121. 5:33can recover the accuracy as the same as
  122. 5:36before Runing okay so with that let's
  123. 5:40start with the first part about
  124. 5:41introduction to pruning so the pruning
  125. 5:44mechanism actually actually happens in
  126. 5:48human brain so according to this study
  127. 5:50from nature a newborn child has about uh
  128. 5:552,500 synapses per neuron okay and when
  129. 5:59uh K grows to to four years old this
  130. 6:03number surges to 15,000 synapsis per
  131. 6:06neuron surges a lot but during
  132. 6:09adolesence this number didn't keep
  133. 6:11increasing but started to
  134. 6:13decrease from 15,000 to only 7,000
  135. 6:16synapsis per neural actually the
  136. 6:19adolescence time is when we go to school
  137. 6:21when we go to college that's the time we
  138. 6:23when we learn the most amount of
  139. 6:25knowledge through our lifetime and
  140. 6:27proing natur happens in that time
  141. 6:31and pruning also happens in artificial
  142. 6:33neuron networks okay so we can make
  143. 6:36neuron networks smaller by removing
  144. 6:40those synapsis and neurons that is
  145. 6:42redundant so on the left hand side is
  146. 6:45showing a dense neuron Network before
  147. 6:48pruning here consist of three fully
  148. 6:51connected layers every each layer is
  149. 6:54densely connected on the right hand side
  150. 6:57we can do pruning on this both the
  151. 6:59synapsis
  152. 7:00and also on the neurons so that not all
  153. 7:03the neurons is connected to each
  154. 7:07other and how does this impact the
  155. 7:11accuracy so I carried an experiment back
  156. 7:13in 2015 we first train the neural
  157. 7:16network and this is Alex net where can
  158. 7:20achieve uh this is accuracy we can
  159. 7:22achieve this is the Baseline accuracy
  160. 7:24and this is the distribution of the
  161. 7:26weight for particular layer roughly from
  162. 7:29formulates a normal
  163. 7:31distribution and then we can gradually
  164. 7:34prove the connections so previously was
  165. 7:36Den now it becom sparse okay as we are
  166. 7:40getting sparer we remove those small
  167. 7:43connections those small weights as a
  168. 7:46result you can see the distribution all
  169. 7:49the weights centered around zero they
  170. 7:52disappeared we remove remove them
  171. 7:54meaning that we've zero them M okay so
  172. 7:57the distribution changed the small value
  173. 8:00disappear and you can see the more we
  174. 8:02prove the less the accuracy the more we
  175. 8:05prove the less accuracy the Y AIS is the
  176. 8:08accuracy loss okay the xais is the pring
  177. 8:12raion the primet pro away um on the
  178. 8:16right hand side you Pro away all the
  179. 8:18parameters you can expect the accuracy
  180. 8:20will drop to zero and on the left hand
  181. 8:23side here we are starting from 40% ring
  182. 8:27ratio on the left
  183. 8:29so this is quite unfortunate right
  184. 8:31people TR these neural networks to get
  185. 8:34high accuracy suddenly due to pruning
  186. 8:37you immedately reduce the accuracy by
  187. 8:39here is like 1% here is like 2% that is
  188. 8:43full B right and we do
  189. 8:45better we can actually put in the
  190. 8:48remaining weights to recover that RK
  191. 8:50okay so comparing this one versus after
  192. 8:54we train the remaining weights that
  193. 8:56survived the pruning okay train this
  194. 8:59weights that survived the pruning and
  195. 9:03actually the curve moved to the upper
  196. 9:06right corner what do me move into the
  197. 9:08upper right corner we can Pro more
  198. 9:11parameters we can Pro more parameters or
  199. 9:14we can achieve the higher accuracy given
  200. 9:16the same tring ratio okay and as a
  201. 9:19result after the ring the weight
  202. 9:22distribution also shifted from this way
  203. 9:26to this way becomes smoother okay
  204. 9:30that's pretty good actually we can Pro
  205. 9:32away here is 90% of the parameters
  206. 9:36without hting the accuracy here on Alex
  207. 9:38that's pretty good can we do even better
  208. 9:42how can we put more without losing the
  209. 9:46accuracy actually we can do this process
  210. 9:50iteratively right not just do one
  211. 9:52iteration but compare here and here we
  212. 9:56can do another round of tring okay
  213. 9:59so from the green curve to the red curve
  214. 10:03we can p even more so about 90% P away
  215. 10:0790% of the parameters without hurting
  216. 10:09the
  217. 10:10accuracy so that debutes the whole
  218. 10:13process you're between the original
  219. 10:16network uh you prune some of the weights
  220. 10:18according to the magnitude and then you
  221. 10:20retrain the remaining weights that
  222. 10:22survive pruning and don't go too uh too
  223. 10:26aggressive at each step you want to go
  224. 10:29any smaller steps um each step should be
  225. 10:33less aggressive so that you can push the
  226. 10:35boundary of this pruning
  227. 10:38process as a result these are the
  228. 10:41networks where uh including Alex Net v
  229. 10:44Google net res squeet at a time we can
  230. 10:47see alexnet can be proved from 61
  231. 10:50million parameters to only six million
  232. 10:52parameters about nine times cor ratio
  233. 10:56and the max reduction is actually
  234. 10:58smaller it's 3x
  235. 10:59in the last lecture we learned the
  236. 11:01difference between number of parameters
  237. 11:04versus the number of Max how to
  238. 11:05calculate them for convolution layer for
  239. 11:07FC layer and also for tension layer so
  240. 11:11they are not
  241. 11:13equal for example even networks that is
  242. 11:16super small like squeez net which are
  243. 11:20Cur and the uh for is design in 2016
  244. 11:24it's already pretty small to begin with
  245. 11:26it has the same accuracy as alex9 but
  246. 11:28being 6 times smaller xet 61 million
  247. 11:32parameters sque net only one million
  248. 11:35parameter so we can still Pro such a
  249. 11:37very small and compact model OKAY from
  250. 11:39one million parameter to only 0.38
  251. 11:42million
  252. 11:43parameters that's 3.2x reduction for the
  253. 11:46model
  254. 11:48size not only for such um uh image tasks
  255. 11:54but also for visual language tasks like
  256. 11:56caption we can still also PR it for
  257. 11:59example the first image uh given the
  258. 12:02caption original caption is a basketball
  259. 12:04player in the white uniform is playing
  260. 12:07with the ball okay it's indeed the case
  261. 12:10and if you PR away 90% of the parameters
  262. 12:12it says a basketball player in wning
  263. 12:15form is playing with a basketball it's
  264. 12:17pretty accurate second
  265. 12:19image uh the Baseline model says a brown
  266. 12:22dog is running through the grassy field
  267. 12:24versus through 90% of brown dog is
  268. 12:26running through a grassy area including
  269. 12:29the third one riding a surfboard on a
  270. 12:31wave a man in the white sweet riding
  271. 12:33wave on the beach
  272. 12:3690% what if we PR more like pruning away
  273. 12:4095% we are more
  274. 12:43aggressive so on the third image the
  275. 12:45original model Baseline model says a
  276. 12:48soccer player in red is running the
  277. 12:50field versus if you thr away 95% it says
  278. 12:54a man in the red shirt and black and
  279. 12:57white black shirt is running through
  280. 12:59field it's getting drunk right so there
  281. 13:01is a limit how much you can cannot Pro
  282. 13:04too much and we're going to later see a
  283. 13:06demo showing the process when we are
  284. 13:08doing the pruning from 70% 80% 90% all
  285. 13:12the way to 99% and we break and here we
  286. 13:15break model and I highly recommend you
  287. 13:17try this experiment offline at home
  288. 13:20we're going to give you the code for uh
  289. 13:23this ion
  290. 13:25not so in recent years PR spity is
  291. 13:29getting quite popular uh this whole
  292. 13:31domain started uh back in 1980s 1980s
  293. 13:35with the paper optimal brain
  294. 13:38damage and then throughout the years it
  295. 13:40uh gradually uh increased until 2015
  296. 13:462016 when I published this paper called
  297. 13:48Deep compression which show that these
  298. 13:50modern neuron networks in a large scale
  299. 13:53data set large scale GP 20 we can
  300. 13:56aggressively prune it there's a lot of
  301. 13:58opportunity to optimize the hard rare
  302. 14:00efficiency by looking through at the
  303. 14:02algorithm level right and then we can
  304. 14:05also design specialize accelerators like
  305. 14:07the E the efficient inference engine to
  306. 14:11accelerate directly on a sparse and
  307. 14:14compressed model in recent years the
  308. 14:17number of Publications per year just
  309. 14:19searching very fast in recent
  310. 14:22years and this has been adopted in
  311. 14:25Industry like ring in
  312. 14:27Industry left hand side other the three
  313. 14:29papers I published first two during my
  314. 14:32PhD and the last two during when I was
  315. 14:35at MIT about how to accelerate building
  316. 14:38specialized accelerators to accelerate
  317. 14:42matrix multiplication and and um
  318. 14:45different workload and N media adopted
  319. 14:48sparcity in the gpus after the a100 GPU
  320. 14:51Andia introduced a 24 sparity we're
  321. 14:54going to introduce that soon so turning
  322. 14:56a dense Matrix into a sparse Matrix and
  323. 14:59have about 2x theoretical speed up and
  324. 15:03about 1.5x measure the speed
  325. 15:06up and sigings now part of AMD also use
  326. 15:10the sparcity to um optimize their models
  327. 15:13using this AI Optimizer um um acquired
  328. 15:18previously my startup and build it into
  329. 15:20a software tool chain to Pro the model
  330. 15:22and f t the model and as a result you
  331. 15:24can get a faster
  332. 15:27inference Okay so so let's go to the
  333. 15:30next chapter about how to determine the
  334. 15:32pruning granularity okay in what pattern
  335. 15:36should we Pro the neuron
  336. 15:40Network so pruning can be performed at
  337. 15:43different granularities can be very fine
  338. 15:46grein or very CL green you can imagine
  339. 15:49there is a very big design space so
  340. 15:52turning from this very dense 2D weight
  341. 15:55Matrix we can make it
  342. 15:57a uh the right one means it's preserved
  343. 16:01and the white rectangles means it's
  344. 16:03pruned fine green pruning is most
  345. 16:07flexible and we can Pro any weights in
  346. 16:09any location but the drawback what is
  347. 16:13drawback here it's hard to accelerate
  348. 16:16right for gpus different threads prefers
  349. 16:20to be uh doing the work in L step manner
  350. 16:23everything want to be paralyzed to the
  351. 16:25same thing no branching right but
  352. 16:27there's a lot of uh irregularities which
  353. 16:30is not Hardware
  354. 16:31friendly uh what is the advantage
  355. 16:34here you have the most flexibility you
  356. 16:38can prune any weight you want and so the
  357. 16:41pruning ratio is the highest using such
  358. 16:44fine
  359. 16:46graining in contrast we can also do
  360. 16:50forse grain okay so in this case we are
  361. 16:53pruning the entire row we are pruning
  362. 16:55the third Row the fourth row um and the
  363. 16:58another row we put through three rows in
  364. 17:02the middle okay what is good about that
  365. 17:05we can condense it into a dense Matrix
  366. 17:09still apply dense Matrix Matrix
  367. 17:12multiplication to do the arithmetic okay
  368. 17:16so the good thing is about it's easy to
  369. 17:18accelerate it's very structured it's
  370. 17:20very regular you can just apply then
  371. 17:23symmetric symmetric
  372. 17:24par but you can imagine this is less
  373. 17:27flexible okay if you prune the whole row
  374. 17:30has to be prune or the whole column has
  375. 17:31to be prune and the pruning ratio is
  376. 17:34much less compared with the previous
  377. 17:37example the fine grain
  378. 17:43example okay so let's talk about the
  379. 17:45pruning granity not just for the FC
  380. 17:48layer but also for the convolution layer
  381. 17:50okay in the convolution layer there are
  382. 17:53four dimensions for the convolution
  383. 17:56kernal CI is the number of input
  384. 17:59channels Co is number of output channels
  385. 18:02and we have KH and KW for the kernels
  386. 18:05height and kernel width and these four
  387. 18:08dimensions give us more choices more
  388. 18:11degree of freedom to select the pruning
  389. 18:16granularity so we start with the um fine
  390. 18:20grain
  391. 18:22pruning in this notation we have K and
  392. 18:25KW equal to three it's a 3X3 Kern
  393. 18:29and here we have three output
  394. 18:32channels and here we have two input
  395. 18:34channels so we read it right the four
  396. 18:36dimensional convolution kernel and
  397. 18:39visualize that in this way so this is
  398. 18:42the fine green pruning atom based
  399. 18:45pruning Vector pruning kernel level
  400. 18:47pruning all the way to channel pruning
  401. 18:49this is the whole landscape and now
  402. 18:52let's dive deeper into each of them and
  403. 18:55talk about what is good about it what is
  404. 18:57bad about it what is the treal
  405. 18:58everything is about fit
  406. 19:01off so fine grain pry okay it's
  407. 19:04irregular but it's most flexible you you
  408. 19:07can have the highest pruning ratio if
  409. 19:09you're just targeting compressing the
  410. 19:11weight without worrying about
  411. 19:13acceleration or paradism this is the way
  412. 19:16to go okay um it's very has very
  413. 19:19flexible pruning indices and you already
  414. 19:23have larger compression ratio since we
  415. 19:25can very flexibly find those redundant
  416. 19:30weights for example here we can compress
  417. 19:33it by up to like an order of magnitude
  418. 19:35for these different neur
  419. 19:39networks it can also deliver speed up on
  420. 19:41Specialized Hardware if you have a um
  421. 19:44the capability to design specialized
  422. 19:46Hardware like eie the efficient
  423. 19:48inference engine which I published in 20
  424. 19:50iscar
  425. 19:522016 um you can do that but it's not
  426. 19:55easily accelerated on off the shelf
  427. 19:58Hardware
  428. 20:00a second category the PN based okay so
  429. 20:05the pruning prun Kel has some ATS okay
  430. 20:09like this is one pattern rotated by 90
  431. 20:12degrees same pattern these patterns are
  432. 20:15the same okay so this is actually give
  433. 20:18you uh more um regularity compared with
  434. 20:23the fine
  435. 20:26grainy one notable pattern is the N2 M
  436. 20:30sparity for example 2 to four sparity so
  437. 20:35start with a dense Matrix on the left
  438. 20:37hand side we can prun it to a toal four
  439. 20:40sparse Matrix so give
  440. 20:43was one minute to take a look at the
  441. 20:46pattern anyone can tell me what is the
  442. 20:48pattern here on the right hand side the
  443. 20:52Matrix
  444. 20:59so what pattern do we see on the second
  445. 21:022 four sparse
  446. 21:10Matrix right every row has four missing
  447. 21:13values right right and actually if you
  448. 21:15look deeper um like the first row four
  449. 21:19missing values and actually uh this is
  450. 21:21due to every four group of four elements
  451. 21:24you must have at least two uh zeros okay
  452. 21:28so two to four means out of four entries
  453. 21:32at least two entries has to be zero and
  454. 21:35that is the case for every group group
  455. 21:38of four elements and that's why we call
  456. 21:41it two to
  457. 21:43four so it's 50% Spar cting and how many
  458. 21:48bits do you need to indicate the
  459. 21:51location of the N zero or the zero
  460. 21:54element you have four Pointes therefore
  461. 21:57you need two bits to
  462. 21:59indicate where the non Zer where the
  463. 22:01zero
  464. 22:02are okay so that is the metadata the
  465. 22:05overhead you have to store every you
  466. 22:08have to store two bits for the
  467. 22:13index so it it really well maintain the
  468. 22:15accuracy so this is the test of accuracy
  469. 22:18across different benchmarks like resent
  470. 22:2050 exception bird Etc comparing the uh
  471. 22:25dense uh accuracy versus the sparse
  472. 22:28accuracy you can see it's pretty much
  473. 22:30the same 76.1
  474. 22:3276.2 the accuracy is very well pretty
  475. 22:36well maintained using this 2 to four
  476. 22:38roughly 50% sparity
  477. 22:42ratio okay so we can also do this
  478. 22:46channel level sh okay we omitted this
  479. 22:50middle two figures for um it's following
  480. 22:55the S similar principle we are just
  481. 22:57getting more and more regular but less
  482. 23:00and less degree of Freedom uh the
  483. 23:03extreme case is the channel pring where
  484. 23:06we are pruning away the entire Channel
  485. 23:09okay um so the pro is that we can
  486. 23:13directly speed it up due to the reduced
  487. 23:16number of channels leading to a neuron
  488. 23:18network with a smaller number of Channel
  489. 23:21and it's still Dan you don't need any
  490. 23:23specialized Hardware just using CPUs
  491. 23:26using whatever Hardware you originally
  492. 23:27have you can directly accelerate it but
  493. 23:31the car is you can have a smaller
  494. 23:34compression ratio for example um for
  495. 23:39convolution neuron Nets like mov net you
  496. 23:42can Pro away only about 30% of the
  497. 23:45parameters of mobile
  498. 23:47net so here we are showing the neuron
  499. 23:50net gr with five layers and we can prune
  500. 23:53the channels with different sparity
  501. 23:55ratio across different layer
  502. 24:00and there are two ways one is to do
  503. 24:02uniform shrinking so for all the layers
  504. 24:04you apply exactly the same sparcity in
  505. 24:07this case 30%
  506. 24:10sparcity but that is really not as good
  507. 24:13as having a smarter way to figure out
  508. 24:16the redundancy and sparcity for each
  509. 24:19layer individually like on the right
  510. 24:21hand side um we'll later discuss how to
  511. 24:24find the optimal sparity ratio to give
  512. 24:27ioc to different layers we're going to
  513. 24:30talk about sensitivity analysis in the
  514. 24:33next
  515. 24:34lecture this also applies to recent
  516. 24:36large language models where a convention
  517. 24:39in the to is that all the layers are
  518. 24:42repeating the same Transformer building
  519. 24:43block exactly the same number of
  520. 24:46channels across different layers um it's
  521. 24:49it's very homogeneous it's very easy to
  522. 24:52partition especially for model level um
  523. 24:55model paradism distribute the weight
  524. 24:57across mod for gpus uh but if you want
  525. 25:00to extract the uh the inference
  526. 25:02efficiency to the extreme different
  527. 25:04layer indeed may have different
  528. 25:07redundancy in spity ratio and we don't
  529. 25:10necessarily have to keep the same am
  530. 25:13amount of channel number across
  531. 25:15different layers Ur sensitivity analysis
  532. 25:18help to analyze that so we are going to
  533. 25:20talk about in lecture two of
  534. 25:24pring and this phenomenum is further
  535. 25:27demonstrated on this figure
  536. 25:29comparing the uniform scating for all
  537. 25:31the layers just uniformly shrink it by
  538. 25:34the same percentage PR away the same
  539. 25:36percentage for all the layers versus
  540. 25:40using an optimal um um a better
  541. 25:44optimized U Spar C ratio search
  542. 25:47algorithm here is AMC automatic model
  543. 25:50compression to search that and here is
  544. 25:53the latency versus the accuracy trade
  545. 25:55off the search approach have a lower
  546. 25:59latency and higher accuracy in this
  547. 26:03case so the student did this work was
  548. 26:05the TA for our class last year his name
  549. 26:08is G he's join open ey after graduation
  550. 26:11this is his
  551. 26:12work okay all right so let's jump into
  552. 26:16the next part how to determine the
  553. 26:19Bruning prer there are so many ways in
  554. 26:23neural network so which one do we keep
  555. 26:25which one do we PR away what is the
  556. 26:27criteria
  557. 26:28for the weights for the synapsis and for
  558. 26:30the neurons so we are going to talk
  559. 26:33about this pruning criteria okay of
  560. 26:37course we want to reduce though we want
  561. 26:39to P away those less important
  562. 26:42parameters so that we can maintain the
  563. 26:45accuracy or minimize the loss uh for
  564. 26:48example in this case we have only three
  565. 26:51waves um 10 x0 - 8 X1 plus 0.1 X2 I
  566. 26:58really want to show very simple example
  567. 27:01for um for intuition right give you some
  568. 27:04intuition which ones to select to Pro if
  569. 27:08out of these three weights out of these
  570. 27:10three weights 10 minus 8
  571. 27:120.1 if one weight has to be removed you
  572. 27:15have the capacity to hold only two
  573. 27:17parameters which one should you remove
  574. 27:21intuitively 0.1 right because it's
  575. 27:24smallest likely to have the smallest
  576. 27:26impact
  577. 27:28so that is actually the most simple well
  578. 27:32the most very effective way to
  579. 27:35determining um the pring criteria just
  580. 27:38select the small on super super easy so
  581. 27:41the importance we just use the um
  582. 27:45magnitude of the weight to indicate the
  583. 27:47importance of the weight if the
  584. 27:49importance is small the magnitude is
  585. 27:51small then we just remove it way also
  586. 27:54called magnitude based pruning so we
  587. 27:57want to maintain the weights with large
  588. 28:00absolute value and PR the weights with
  589. 28:03very small absolute value okay that is
  590. 28:06very simple theistic but turn turned out
  591. 28:09to be working super well uh in both
  592. 28:11Academia and Industry for so many
  593. 28:14years in the example on on the bottom we
  594. 28:17have four weights we find out our one
  595. 28:19Norm for each of them and this is our
  596. 28:22one norm and we keep the red the largest
  597. 28:25and remove the smallest so this becomes
  598. 28:27the pr weights so just use the uh
  599. 28:30absolute value to determine the pruning
  600. 28:34uh
  601. 28:35criteria okay what about for uh four
  602. 28:40screen pring for example if you want to
  603. 28:42Pro away the whole role of this
  604. 28:46Matrix uh we can apply L1 Norm or L2
  605. 28:50Norm for example here we start with L1
  606. 28:52Norm magnitude based ofon so we find the
  607. 28:56L1 Norm for the first row which is 3 + 2
  608. 29:00which is five we also find the our one
  609. 29:03Norm of the second row which is six okay
  610. 29:06and here we compare five is smaller than
  611. 29:08six so we are going to thr away five so
  612. 29:11our one Norm I only leave the second row
  613. 29:15unpr similarly we can also apply our two
  614. 29:18Norm so this is the way to calculate our
  615. 29:21two Norm um it's very simple I won't
  616. 29:24repeat it here um that is the Lar
  617. 29:28um characteristic just use the magnitude
  618. 29:31no matter if it is tensor or if it is a
  619. 29:34um um just a single value okay and in
  620. 29:38general we can use the lp Norm U to
  621. 29:40determine the uh Runing
  622. 29:45criteria we can also apply this scaling
  623. 29:48based pruning technique for example here
  624. 29:51we have N filters from Filter zero
  625. 29:53filter one all the way to filter n minus
  626. 29:56one and we apply scating Factor as to
  627. 29:59associate each future with a skating
  628. 30:01Factor okay so here is the skating
  629. 30:05Factor associated with each Channel like
  630. 30:08the first channel will be multiplied
  631. 30:10with 1.17 second channel will be
  632. 30:12multiplied with
  633. 30:140.1 and that scaling factor is learnable
  634. 30:18okay so that's learnable you have one
  635. 30:20learnable parameter for the whole
  636. 30:22channel so it's actually very parameter
  637. 30:24efficient you only have un numbers to
  638. 30:27learn right here
  639. 30:30and you want to minimize uh the scaling
  640. 30:32factor to try to push them uh to to zero
  641. 30:36so that we can easily PL them away like
  642. 30:38here if the scaling factor is 0.1
  643. 30:40compared with 1.17 this is pretty small
  644. 30:43so likely we are going to remove uh the
  645. 30:46second filter filter one okay so we can
  646. 30:49later sort um the scating factor and pro
  647. 30:54away the channels with a very small SC
  648. 31:00Factor so originally this was the number
  649. 31:04filters and removed uh the channels with
  650. 31:08small scaling factor and the the neuron
  651. 31:12Network becomes something on the right
  652. 31:14hand side the filters and output
  653. 31:16Channels with small scaling Factor will
  654. 31:19be approved that's
  655. 31:30of course there are many other
  656. 31:32characteristics we're also going to
  657. 31:33cover very soon so this is the first
  658. 31:36theistic very simple but there are more
  659. 31:39complicated ones let's talk about
  660. 31:44that and to continue talk about the
  661. 31:47scaling Factor right here uh the scaling
  662. 31:50Factor can be reused from the batch
  663. 31:52normalization layer which we learned
  664. 31:54from the previous lecture uh for the
  665. 31:56batch normalization you have one scaling
  666. 31:58Factor per Channel okay you have one
  667. 32:01scaling Factor per Channel that's
  668. 32:02exactly the scaling factor that is here
  669. 32:05so you can reuse the same scaling factor
  670. 32:08from the batch normalization L to
  671. 32:10simplify the
  672. 32:12calculation what other theistic we have
  673. 32:15magnitude may not be the best one right
  674. 32:19um it's hard to tell us is hard to tell
  675. 32:21which is the best heris depending on the
  676. 32:24data set depending on the neural network
  677. 32:27is experiment Al stuff but here let me
  678. 32:30introduce the design space so when you
  679. 32:32are doing such pring tasks in the future
  680. 32:35you at least know what is the the way to
  681. 32:37think about it what are the choices
  682. 32:40there's never a conclusion which one is
  683. 32:42the best and I still try
  684. 32:44it second order based approving okay so
  685. 32:48let's apply the tailor expansion to
  686. 32:50approv the network this is the pr
  687. 32:53Network versus the original Network um
  688. 32:56PR the network you basically give a
  689. 32:58preservation to the Del the W to the
  690. 33:00weights or the Delta W which is equal to
  691. 33:04um this is the first order second order
  692. 33:09and third order approximation okay so we
  693. 33:12can remove the third order and
  694. 33:15Beyond and only keep the uh first and
  695. 33:19second order the paper optimal brain
  696. 33:22damage suggest that um the last term is
  697. 33:26we can neglect them and then the sing
  698. 33:30has converged so the gradient should be
  699. 33:32very close to zero you already converged
  700. 33:34at a local Minima so first order you can
  701. 33:37neglect them so only the second order
  702. 33:40term is
  703. 33:43there a second order term has two parts
  704. 33:47the error caus by deleting each
  705. 33:49parameter assumed to be independent so
  706. 33:52the cross term can also be neglected and
  707. 33:55as a result we only have one term in
  708. 33:58middle um which is the second order
  709. 34:01term and we use that to determine
  710. 34:05whether this weight is important or
  711. 34:09not okay so the
  712. 34:11importance summarize this way where uh
  713. 34:15the H is the H
  714. 34:19Matrix but the down side is that the hro
  715. 34:22Matrix is difficult to compute so we
  716. 34:24have to apply some approximation to
  717. 34:27compute to estimate the H
  718. 34:32Matrix okay so beyond the synapsis we
  719. 34:35can also prove the
  720. 34:37neurons when we remove removing the
  721. 34:40neurons and removing the synapsis they
  722. 34:44are very um highly related so removing
  723. 34:48the neuron is equal to actually a very
  724. 34:51CL weight to for example in a linear
  725. 34:56layer right here removing one neuron
  726. 34:59right here means we are removing all the
  727. 35:02weights associated with this output
  728. 35:04neuron meaning that we are reducing one
  729. 35:06row in the weight
  730. 35:08Matrix similarly in the convolution
  731. 35:11layer when we are removing some of the
  732. 35:14output channels it also means we are
  733. 35:17removing the entire
  734. 35:18kernel corresponding to that
  735. 35:25channel okay so uh let's talk about one
  736. 35:28way to determine uh which activation to
  737. 35:32remove using the percentage of zero
  738. 35:35percentage of zero based through me okay
  739. 35:38since Ru will generate a lot of zeros
  740. 35:42okay so um here we have a batch of two
  741. 35:45this is batch one this is batch two and
  742. 35:48we have three channels and each channel
  743. 35:50is 4x4 so there are two uh two feature
  744. 35:53map
  745. 35:54batches and we just um calculate the uh
  746. 35:59average percentage of zeros in these two
  747. 36:02batches so rather than using complicated
  748. 36:05math let's just see that from a a simple
  749. 36:09example okay so there are two badges um
  750. 36:13what is the average number percentage of
  751. 36:16zero of Channel Channel Zero channel one
  752. 36:19channel three for channel one here we
  753. 36:21have five n zeros five zeros okay and in
  754. 36:26the second image we have 1 2 3 4 5 six
  755. 36:29we have six uh zeros and how many
  756. 36:32elements do we have in total four by
  757. 36:34four we have patch of two so 2 * 4 is 4
  758. 36:3732 so this is the average percentage of
  759. 36:41zero for Channel Zero okay and similarly
  760. 36:44we can calculate that for channel two
  761. 36:47well five zeros here seven zeros here so
  762. 36:50average percentage of Z is 12 ID 32 and
  763. 36:54similarly we can calculate for channel
  764. 36:56two
  765. 36:58and we compare them and we are going to
  766. 37:01remove the Channel with the most the
  767. 37:04largest amount of average percentage of
  768. 37:07zeros since they are supposed to be
  769. 37:10redundant and this is done by rather
  770. 37:13than using a static way for measuring
  771. 37:15the weight and when we are ping the
  772. 37:17weight you don't have to run any input
  773. 37:19you just statically look at the weight
  774. 37:22but for the for the activations right
  775. 37:26you have to really look at
  776. 37:28you have to really run a few
  777. 37:30samples so that you
  778. 37:32can so that you can um calculate the
  779. 37:36average percentage of zeros okay so in
  780. 37:39this case we run a batch size of two and
  781. 37:42then we calculate this the6 of these two
  782. 37:45samples so that we can calculate average
  783. 37:47percentage of
  784. 37:50zeros all right let's take a short break
  785. 37:53before jump into the next
  786. 37:56method all right welcome back let's
  787. 37:59resume the lecture so just now we talk
  788. 38:02about activation plan right so a common
  789. 38:06question here is the tensor here is no
  790. 38:09longer the weight okay this is the
  791. 38:11activation tensor those six matrixes
  792. 38:14those are the activations not the weight
  793. 38:16that's um um if I didn't explain that
  794. 38:19clearly before now is the time to
  795. 38:21clarify these are the activations and we
  796. 38:23are trying to prove the activation
  797. 38:26Channel and we run through the network
  798. 38:29with two input examples therefore we can
  799. 38:32get two categories okay so this is the
  800. 38:35first uh this is the first
  801. 38:38batch this is the second batch we run
  802. 38:41two samples both are activations and we
  803. 38:44collect the statistics about these
  804. 38:46activations to perform activation
  805. 38:49pruning all right so
  806. 38:53next regression based pruning so what is
  807. 38:57regression based cling usually um if you
  808. 39:00run the network end to end to calculate
  809. 39:03the loss that's could be pretty
  810. 39:06expensive for example if you want to Pro
  811. 39:08Lama and you want to use the last loss
  812. 39:12as the end loss as your supervision when
  813. 39:15we are doing the rining that could be
  814. 39:17super expensive um as opposed to that we
  815. 39:20can um minimize the Reconstruction error
  816. 39:24of the corresponding layer okay to layer
  817. 39:27wise
  818. 39:27layerwise reconstruction so you only
  819. 39:30need to um work on only one matrix
  820. 39:34multiplication to minimize the change of
  821. 39:36that Matrix
  822. 39:38modifcation for example right here we
  823. 39:40have a activation times the weight um
  824. 39:44the activation has batch size of two and
  825. 39:46four input channels the weight has four
  826. 39:50input channels and eight output channels
  827. 39:52and we get an output answer of batch
  828. 39:55size times Co okay
  829. 39:58and then we we try to prove some of the
  830. 40:00channels on the weight for example here
  831. 40:03we try to Pro the second channel on the
  832. 40:06weight okay so what is modified with the
  833. 40:08second Channel actually that's the
  834. 40:10second channel of the activation the
  835. 40:13modifcation uh they will correspond to
  836. 40:15each other so what is the dimension of
  837. 40:18the output after doing such after doing
  838. 40:21such um channel
  839. 40:23throughing it's the same it's still F
  840. 40:26size by Co okay because we are putting
  841. 40:30in the CI Dimension so the B size and Co
  842. 40:33those two dimensions are not impacted so
  843. 40:38um although it's just a matrix
  844. 40:39multiplication but if you want to see
  845. 40:41the pry of large language model
  846. 40:42basically everything boils to matrix
  847. 40:44multiplication so this can be quite
  848. 40:46cral and now we try to minimize um the
  849. 40:50error between the unpruned output tenser
  850. 40:54versus the pruned output tenser okay um
  851. 40:58that makes the optimization local and a
  852. 41:01lot easier compared to you back
  853. 41:03propagate across entire all the layers
  854. 41:08and supervise it in that
  855. 41:10way so how do we do proving here um we
  856. 41:15can view this matrix multiplication into
  857. 41:18four parts four is the CI Dimension okay
  858. 41:21so c i Dimension CI has four is four so
  859. 41:25four input channels
  860. 41:27and we color them in four different
  861. 41:29colors okay so X x0 multiply with w0 X1
  862. 41:35minus X1 will multiply with W1 X2 will
  863. 41:40multiply with W2 okay so that's the
  864. 41:42corresponding activation and
  865. 41:44corresponding uh weight channel so we
  866. 41:48can view them as the outer product okay
  867. 41:51the outter product of XC times WC and
  868. 41:57and sum sum them together okay in this
  869. 41:59case we have four terms okay one two
  870. 42:03three four four terms to sum together
  871. 42:06alter
  872. 42:08product and then we apply a scaling
  873. 42:10factor to each outer product beta C
  874. 42:13which is the speeding factor
  875. 42:16for each auor product we sum them up we
  876. 42:21try to minimize the difference between
  877. 42:24the original Z and the the pruned Z Z
  878. 42:27such that subject to we want to minimize
  879. 42:31the zero Norm of beta to have to to make
  880. 42:35beta as close to zero as possible if
  881. 42:38beta is zero what does it
  882. 42:40mean it means this Al product doesn't
  883. 42:44exist this uto product doesn't exist
  884. 42:47like this channel okay so it means um
  885. 42:50the channel is approved so how many
  886. 42:53betas do we have in this example
  887. 42:58we have four right four colors we have
  888. 43:01four betas if we PR in this way means
  889. 43:05beta the first beta so beta beta 1 beta
  890. 43:090 beta 1 beta 3 two beta
  891. 43:123 means beta beta one is zero okay
  892. 43:15corresponding to the white region that
  893. 43:19is
  894. 43:20pro and how do we solve that problem so
  895. 43:23we can first fix W and then solve beta
  896. 43:26to select the pr the channel and then we
  897. 43:29can fix beta to solve for w to minimize
  898. 43:33the rec reconstruction error so we can
  899. 43:36do this way iteratively okay we first a
  900. 43:38fix W to find among these four beta this
  901. 43:43four other products which other product
  902. 43:45if we remove them will have the minimum
  903. 43:48impact on the
  904. 43:51output and after selecting that that
  905. 43:53data for example after selecting the
  906. 43:55second channel to be pred
  907. 43:57we can solve W to minimize the
  908. 44:00Reconstruction error and we can do such
  909. 44:02process
  910. 44:04iteratively now this could be quite
  911. 44:06helpful for pruning uh L language models
  912. 44:10where it's super expensive to back
  913. 44:12propagate all the way uh to the very end
  914. 44:15and back properly to the very
  915. 44:18beginning all right so far we talk about
  916. 44:21what is pruning primarities of pruning
  917. 44:24criteria to select the weights to prune
  918. 44:28and we are going to show a demo to
  919. 44:31strengthen our um understanding of try
  920. 44:34so let's now switch gear to the ring
  921. 44:41demo all right um this is The Notebook I
  922. 44:49prepared let's maximize The Notebook
  923. 44:54make the font size a bit larger so that
  924. 44:57can see see
  925. 45:03it okay so in this s we prepared a amist
  926. 45:08data set uh to Pro the network to
  927. 45:13classify the H written digits from zero
  928. 45:16to n okay so if you random guess the
  929. 45:19accuracy should be
  930. 45:2010% um so let's run the my python
  931. 45:24notebook start from the beginning let's
  932. 45:26first to the
  933. 45:30setup this is using Google collab the
  934. 45:34same infrastructure we are going to use
  935. 45:37for our
  936. 45:38homeworks so for the free version I
  937. 45:41would say you can have uh this P4 GPU
  938. 45:46which is a not super Advanced GPU but it
  939. 45:50should be enough for learning purpose to
  940. 45:54be enough for this
  941. 45:55lecture we run the
  942. 45:58setup and we prepare the ne Network
  943. 46:02model and then let's
  944. 46:06visualize the images now we have 10
  945. 46:09digits to
  946. 46:11classify and then we are going to
  947. 46:13retrain the neural network on the
  948. 46:16administ data
  949. 46:18set this training has started remotely
  950. 46:22Google Cloud this is now running on the
  951. 46:25t4g viu
  952. 46:30okay play about one I has completed the
  953. 46:36accuracy already is already
  954. 46:3998% right
  955. 46:42here so we are going to repeat the
  956. 46:45training process training for five iPods
  957. 46:49this is iPod
  958. 46:50one actually I 2 already finished
  959. 46:5498.6% accuracy
  960. 46:59you can of course upgrade to a better
  961. 47:02GPU by using the PA version um but I
  962. 47:05think the threee version is also enough
  963. 47:09for our Labs actually from Lab One to
  964. 47:12lab four we are all going to use Google
  965. 47:14cab probably lab four will be a bit
  966. 47:16slower since we are going to run large
  967. 47:18language model and lab five will be uh
  968. 47:22written in C++ so make sure you learn
  969. 47:25have those uh PR requisites to um
  970. 47:29manipulate the pointers deal with the
  971. 47:31body
  972. 47:34threading okay the last iPod you know
  973. 47:37already 90 9
  974. 47:4498.9% the last ioc has finished we 98.9
  975. 47:49n% which is actually pretty
  976. 47:54satisfying okay that's first EV valuate
  977. 47:57the accuracy and mod size of this St
  978. 47:59model
  979. 48:01before
  980. 48:0398.9% almost 99% accuracy and this is
  981. 48:07the digit the second row is the
  982. 48:08prediction we can see in this batch of
  983. 48:12examples we are all
  984. 48:15correct so then let's do the pruning
  985. 48:18apply a pruning
  986. 48:20ratio let's start with some moderate p
  987. 48:23ratio let's maybe start with uh s
  988. 48:28% I SC away um 70% of the parameters and
  989. 48:33see what
  990. 48:36happens the accuracy dropped as expected
  991. 48:40right the accuracy dropped from 99% to
  992. 48:4594.6% okay and if you see the
  993. 48:48visualization some of the mark uh
  994. 48:52classified not correctly from8 to 9 is
  995. 48:56is a
  996. 48:57mistake so what can we do about it let's
  997. 49:00find tune the pr model to get a higher
  998. 49:05accuracy and we find tune for two IO in
  999. 49:10this case here's the entire code to do
  1000. 49:14that learning rate the momentum the
  1001. 49:16weight
  1002. 49:17Decay let's just find human the model to
  1003. 49:19see if we can recover the accuracy from
  1004. 49:2394% back to 99%
  1005. 49:27one IO already finished the increase to
  1006. 49:3198.78% which is a good starting
  1007. 49:34point let's fine tune for another
  1008. 49:39IPO the accuracy recovered to uh
  1009. 49:4498.85%
  1010. 49:46already quite close to
  1011. 49:4998.9% and we already P away 70% of the
  1012. 49:53parameters only with 30% of the
  1013. 49:56parameters
  1014. 49:58left okay so what shall we do um let's
  1015. 50:02visualize
  1016. 50:05them let's load the
  1017. 50:08model and for these 10 examples they
  1018. 50:12actually all
  1019. 50:16correct I showing that by pruning and
  1020. 50:20retraining the remaining models we can
  1021. 50:22pretty much recover the accuracy loss
  1022. 50:24from 94% back to 98% tax so let's do
  1023. 50:28something more aggressive rather than
  1024. 50:30pulling away uh 70% let's do
  1025. 50:4090% unfortunately the model accuracy
  1026. 50:43dropped to only 20%
  1027. 50:4620% this as expected right because we
  1028. 50:49are already proving with 90% of the plan
  1029. 50:52very natural accuracy will drop a lot
  1030. 50:55for example this four that get
  1031. 50:56classified as three let's see how this
  1032. 51:00fine tuning can help recovery the
  1033. 51:02accuracy so let's fine tune the pro the
  1034. 51:04model to get a higher
  1035. 51:11accuracy let's find Cate using the same
  1036. 51:15uh schedule for two IO and of course
  1037. 51:18feel free to adjust the learning rate um
  1038. 51:21and also the momentum and also the way
  1039. 51:24Decay or or increasing the number of IPO
  1040. 51:27to see if we can have a um better result
  1041. 51:30in
  1042. 51:34practice okay to IO finish
  1043. 51:3998.3% isn't that amazing we P away 90%
  1044. 51:42of parameters after retraining we can
  1045. 51:44still get 98% of
  1046. 51:46accuracy
  1047. 51:48so let's load the model again classify
  1048. 51:52the 10
  1049. 51:53digits actually all these cases are
  1050. 51:56correct
  1051. 51:57and the test accuracy is
  1052. 52:0298.35% how about let's do something more
  1053. 52:05aggressive some someone can tell me a
  1054. 52:08number you want to
  1055. 52:09try 99% okay okay 99% only 1% of the
  1056. 52:14primers is left it's it's a very
  1057. 52:20aggressive all right 10% it's random gu
  1058. 52:24right you have 10 digits you have 10%
  1059. 52:27accuracy so that's basically random
  1060. 52:32guess everything is full yeah some
  1061. 52:34already recognized everything is
  1062. 52:36predicted as two after Runing away 99%
  1063. 52:39of the CR DPS so let's find tune that
  1064. 52:43for IPO two IPO and see what happens and
  1065. 52:47we increase it a little bit or the first
  1066. 52:50IPO give you 9% of accuracy back
  1067. 53:06okay second one finished
  1068. 53:1033% so let's run the demo
  1069. 53:19again okay like this one is correctly
  1070. 53:22predicted and unfortunately the
  1071. 53:24remaining one is not correctly predicted
  1072. 53:26right so that's the where uh was pretty
  1073. 53:29much the the limit 99% will not work you
  1074. 53:33can see we exactly played the entire
  1075. 53:36curve from 70% 90% 99% even with r cing
  1076. 53:42we can recover the accuracy up to uh 90%
  1077. 53:46but if you prune too aggressively like
  1078. 53:4899% the accuracy is going to drop very
  1079. 53:52aggressive and even cannot be recovered
  1080. 53:54from your tring
  1081. 53:56but every the technology is improving
  1082. 53:59very rapidly maybe some of you come come
  1083. 54:01up with the new algorithms to push the
  1084. 54:04frontier even more aggressive fing ratio
  1085. 54:06and even higher accuracy so we have a
  1086. 54:08lab for lab one we are going to release
  1087. 54:13in on next Tuesday so you can play with
  1088. 54:16the um such pruning and retraining
  1089. 54:19process to figure out a good training
  1090. 54:22schedule question
  1091. 54:34oh the question is what is the
  1092. 54:36difference between the training initial
  1093. 54:39training versus the fine tuning okay so
  1094. 54:41the fine tuning we train for two
  1095. 54:43IO and learning rate 0.1 momentum 0.9
  1096. 54:47and weight Decay e minus 4 and we go to
  1097. 54:51see the original uh training process the
  1098. 54:54learning rate is larger okay the
  1099. 54:56learning rate is one but in the printing
  1100. 54:59process we reduce it by 10x reduce it
  1101. 55:02since it's pretty much converged so we
  1102. 55:04reduce learning rate and initial
  1103. 55:07training has five IO and the fine tuning
  1104. 55:09took only two IPO yeah
  1105. 55:17right question
  1106. 55:26yeah so training a smaller model from
  1107. 55:28scratch even TR for a longer
  1108. 55:31time is you practically practically
  1109. 55:34worse than pruning a larger model so
  1110. 55:37during the optimization you redundancy
  1111. 55:39is helpful for you to get away from the
  1112. 55:42local minimum for example you have a
  1113. 55:45sadle point and then you add another
  1114. 55:47dimension um if you get a stock in local
  1115. 55:50minimum if you add another dimension if
  1116. 55:52it is SLE structure you can go even even
  1117. 55:56lower um so over parameter over
  1118. 55:59parameterization helps with optimization
  1119. 56:02and after the optimization you can Pro
  1120. 56:04away them
  1121. 56:10away
  1122. 56:20question when removing a row at the
  1123. 56:23column from The Matrix
  1124. 56:27can
  1125. 56:37you uh so those are athal techniques SD
  1126. 56:41versus pruning um you can also Pro on
  1127. 56:45top of a
  1128. 56:47asvd uh
  1129. 56:49Matrix so pruning um asbd low rank
  1130. 56:54approximation quantization ination um
  1131. 56:58there are different AAL techniques to
  1132. 57:01apply
  1133. 57:06here all right so in the next lecture
  1134. 57:09we're going to cover a few um techniques
  1135. 57:12in next Tuesday from how to find the
  1136. 57:14pruning ratio for each layer we find
  1137. 57:16that super crucial compared with uniform
  1138. 57:19pruning ratio and how to try and find
  1139. 57:21human PR layer and automated ways to
  1140. 57:24find the pr ratios right than manually
  1141. 57:26find them and also how to the system and
  1142. 57:29Hardware support for different
  1143. 57:31granularities together with lab one
  1144. 57:33we'll be out by next Tuesday actually we
  1145. 57:36already have all the labs available
  1146. 57:37online so if you're eager to try what we
  1147. 57:40have tested today feel free to grab lab
  1148. 57:43one from our course website which is
  1149. 57:45efficient
  1150. 57:47ml. and here are the reference for today
  1151. 57:51which conclude today's lecture thank you

About this transcript

This page contains the full transcript of EfficientML.ai Lecture 3 - Pruning and Sparsity Part I (MIT 6.5940, Fall 2024) by MIT HAN Lab, generated from the public captions YouTube serves with the video. The transcript has 7,658 words across 1,151 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.