YouTube2Text

YouTube transcript (cmNIMjPYdgM) — Transcript

17,221 words · 2,467 segments · language en · Watch on YouTube

Full transcript

  1. 0:05So, good afternoon. Welcome to uh CS
  2. 0:08229. I'm sorry I didn't get a chance to
  3. 0:11see you all on Monday. Uh but this next
  4. 0:13section of the course I'll be teaching.
  5. 0:14Tenu and I are going to swap off as we
  6. 0:16go through the lecture. This next
  7. 0:18segment we're going to cover kind of the
  8. 0:20basics of machine learning and AI, which
  9. 0:23is supervised machine learning. This is,
  10. 0:25as you're going to see, this is when you
  11. 0:26kind of explicitly tell the machine what
  12. 0:29you want it to to label. So you could
  13. 0:30show it an image of a cat or a dog um
  14. 0:33and tell it to classify as a cat or a
  15. 0:34dog. We'll start off with something
  16. 0:36else. It's called regression, which is
  17. 0:37just a little bit more kind of
  18. 0:39mathematically simple. And today we'll
  19. 0:41go through kind of all the building
  20. 0:42blocks of what a learning algorithm
  21. 0:43looks like and hopefully something that
  22. 0:46makes a little bit of intuitive sense.
  23. 0:47And just by way of introduction, you
  24. 0:49know, I've been working on machine
  25. 0:50learning and AI for most of my
  26. 0:51professional career. It's been really
  27. 0:53exciting to watch it grow from something
  28. 0:55that was, you know, very obscure um to
  29. 0:57now something that has products in your
  30. 0:59hand that you can use. Um and so I'm
  31. 1:01super excited about that. So these
  32. 1:02classes are always really fun for me
  33. 1:04because you know this stuff that we're
  34. 1:06teaching here um you know in some ways
  35. 1:09is the foundation for so much of what
  36. 1:11what now you kind of use on a daily
  37. 1:13basis. So it's really exciting to do
  38. 1:14that. We're going to teach in a very,
  39. 1:16you know, abstract and mathematical
  40. 1:17style. I'm sure as Tenu shared with you,
  41. 1:19so we can communicate really, really
  42. 1:21precisely. Um, so couple of disclaimers.
  43. 1:24Let me see if I can get this guy up.
  44. 1:28Pointer up. So, I'm using this format
  45. 1:30with slides. So, you can always download
  46. 1:32the slides online. Uh, I checked them
  47. 1:34all up, at least for my section. Um,
  48. 1:37please look at them. It's just because
  49. 1:38when we're doing things with math,
  50. 1:40there's all kinds of little I's and J's
  51. 1:41and indices, and it's just nice to have
  52. 1:43them all clean. If you see any bugs,
  53. 1:46please send them to me. Um, these are
  54. 1:47like handwritten notes that have been
  55. 1:49passed down that I wrote a long time
  56. 1:50ago. Um, but the best lecture for
  57. 1:52everything is Tangu's course notes. So,
  58. 1:54if you get confused why what's here,
  59. 1:56this is meant to be a subset of the
  60. 1:59course notes and it's just meant to give
  61. 2:01you kind of the tune of what's going on.
  62. 2:03It's, you know, precise enough that you
  63. 2:04can understand how all the pieces get
  64. 2:05together, but if you want to go back and
  65. 2:07rigorously study, I I would really
  66. 2:09suggest those notes. Tu work quite hard
  67. 2:11on on getting those notes in shape. Um,
  68. 2:13I'm in general worried that the lecture
  69. 2:15pacing will be too quick with these. I
  70. 2:17used to in, you know, antiquity give
  71. 2:19them on these whiteboards. I really like
  72. 2:21whiteboards because it forces me to slow
  73. 2:23down and and talk and walk through it.
  74. 2:25So, that's kind of another way of me
  75. 2:26asking you. Please just go ahead and ask
  76. 2:28questions. Like, I have given these
  77. 2:29lectures a bunch of times. These are
  78. 2:31topics that I've talked about. I would
  79. 2:33much rather talk to you if I'm totally
  80. 2:35honest with you. Um, you know, as a
  81. 2:36faculty member, I do things outside the
  82. 2:38university. The reason I stay, you know,
  83. 2:40affiliated with Stanford is because I
  84. 2:42like mentoring students. It's a way to
  85. 2:44kind of have a meaningful life. I also
  86. 2:46run companies and do other stuff in in
  87. 2:47another life. Um, but this is what I
  88. 2:50enjoy is talking to you all. So, please
  89. 2:52please feel feel free to ask whatever
  90. 2:54questions you like as we go. Um, I will
  91. 2:56do my best to answer them. Okay. I
  92. 2:58generally talk fast when I get excited.
  93. 3:00I'm not that excited right now, I guess,
  94. 3:01but I will be excited later. And when I
  95. 3:03get excited, I'll talk fast. Thank
  96. 3:05goodness that most of you will watch
  97. 3:06these on recordings and you can watch
  98. 3:08them at 50% speed. I'll do my best to
  99. 3:10slow down, but you know genetically it's
  100. 3:12been it's been rough. All right, so
  101. 3:14let's get to it. So we're going to talk
  102. 3:16about our first first kind of exposure
  103. 3:19to a formal learning algorithm. The hope
  104. 3:21is maybe you've seen some of these
  105. 3:23concepts before. So today is going to be
  106. 3:24a little bit heavier on notation. It's
  107. 3:26going to be heavier on kind of the
  108. 3:27basics that we're going to use. Um and
  109. 3:30we're going to talk about kind of the
  110. 3:31most classical learning algorithm which
  111. 3:33you know predates machine learning by a
  112. 3:35lot which is linear regression. Okay.
  113. 3:38Then we're going to see linear
  114. 3:39regression. We're going to see that this
  115. 3:40kind of way of fitting a line is
  116. 3:42actually quite interesting. Like it's
  117. 3:44it's actually a pretty robust technique.
  118. 3:46A lot of science was done on it. A lot
  119. 3:47of machine learning was done on it. And
  120. 3:49we're going to see elements that have
  121. 3:50really dominated a lot of what machine
  122. 3:52learning does now. So, one of the big
  123. 3:55things that happened in machine learning
  124. 3:56were, as kind of strange and bizarre as
  125. 3:59this is to say, we made the models a lot
  126. 4:01bigger. Like, it's kind of that dumb.
  127. 4:03Like, we made the models that bigger and
  128. 4:04you know, as they say, we made the GPUs
  129. 4:06go burr and then we got AI. Okay, now
  130. 4:08that version is like not too far away
  131. 4:11and I like worked on it. Like, you
  132. 4:12should see my Stanford job talk about
  133. 4:13like making models bigger way back when.
  134. 4:16The workhorse algorithm we're going to
  135. 4:17cover today is this thing called batch
  136. 4:19and stochastic gradient descent. Okay.
  137. 4:21Now, if you look at these and you have
  138. 4:22kind of a mathematical mind or
  139. 4:23statistical mind, you're going to
  140. 4:25realize that these are really stupid
  141. 4:26algorithms. Hopefully, they're really
  142. 4:28simple. But that simplicity is what
  143. 4:30allowed us to scale them up and run them
  144. 4:32on huge amounts of data and build these
  145. 4:34crazy models. And so, like stochastic
  146. 4:36gradient descent I've run in the last 24
  147. 4:38hours for whatever it's worth. Okay,
  148. 4:39there are much more sophisticated
  149. 4:41algorithms that you may have seen in
  150. 4:42optimization classes. We're not going to
  151. 4:44cover those. Okay. And then the last bit
  152. 4:46I'm going to tell you a little bit about
  153. 4:47normal equations. The normal equations
  154. 4:49basically are an excuse for me to give
  155. 4:51you some notation about matrices and
  156. 4:52vectors. And that's because if this part
  157. 4:54of the section feels a little bit
  158. 4:56unclear to you, like it's just not kind
  159. 4:58of clicking to you, you kind of either
  160. 4:59vaguely remember it or maybe you haven't
  161. 5:01seen it, we have these wonderful classes
  162. 5:03on Friday where you can practice things
  163. 5:05like this and you'll have kind of a
  164. 5:06little bit of calculus over matrices and
  165. 5:09vectors. Um, you will end up by the end
  166. 5:11of this course using that kind of stuff
  167. 5:12as secondhand, but it can be mysterious.
  168. 5:15So spend a little bit of time at the
  169. 5:17beginning and invest and use these great
  170. 5:19TA courses like this course. One of the
  171. 5:21nice things about it, you know, the
  172. 5:22professors are kind of interchangeable,
  173. 5:23but the course staff like what sets it
  174. 5:26up so that you have homeworks and you
  175. 5:28have great TA sections and all the rest.
  176. 5:31Please avail yourself of that. We want
  177. 5:32everybody to get through and have a
  178. 5:33great time. Okay? All right. So that
  179. 5:35what the last part is like in this part,
  180. 5:37if you find yourself being like, I don't
  181. 5:38understand what that symbol means or I
  182. 5:40don't understand why that
  183. 5:40transformation, please of course feel
  184. 5:42free to ask. [snorts] But if it's kind
  185. 5:44of not sticking with you, please take
  186. 5:46advantage of those Friday TA lectures
  187. 5:49and the rest of the resources that are
  188. 5:50out there. Okay. All right. Let's get
  189. 5:53started. Okay. So,
  190. 5:56the basics of supervised learning is
  191. 5:58going to be a hypothesis. Okay. And this
  192. 6:00is just a function from some abstract
  193. 6:02set XX to some abstract exact Y. Um that
  194. 6:06we're going to do. And it's that
  195. 6:07hypothesis or prediction. We'll talk a
  196. 6:09lot about what it means to make a good
  197. 6:11hypothesis or prediction, but let's look
  198. 6:12at some examples of kind of the types
  199. 6:14that we might deal with. So, one example
  200. 6:16is we could have X being the space of
  201. 6:18all images. [clears throat] And then Y
  202. 6:20could be a fixed set of labels like does
  203. 6:22it contain a cat or a dog or a horse or
  204. 6:25or a person or a building, whatever.
  205. 6:27Those fixed set of labels are something
  206. 6:29that we could learn a function h that we
  207. 6:31hopefully apply to a bunch of images and
  208. 6:34then learn when a new image comes that
  209. 6:36wasn't something that we saw ahead of
  210. 6:37time. We can still classify it as a
  211. 6:40horse or car or whatever. Okay, it can
  212. 6:42be text. It can be house data. Okay, and
  213. 6:46text you could imagine classifying
  214. 6:48things whether they're positive or
  215. 6:49negative, whether it was a positive
  216. 6:50review about a product, all the rest of
  217. 6:52it. Okay, when it's house data, we'll
  218. 6:54look at this. It could also be a price.
  219. 6:56So y doesn't have to be just categorical
  220. 6:58labels. Doesn't have to be yes or no.
  221. 6:59Doesn't have to be, you know, cat or
  222. 7:01dog. It could be a real number. It could
  223. 7:03be a scalar value. It could be the price
  224. 7:05of a house or something like this.
  225. 7:06Right? So this is kind of what we're
  226. 7:07doing.
  227. 7:10Now what makes it supervised
  228. 7:13is this training set. So what we do is
  229. 7:16we have we imagine that we have we've
  230. 7:18collected some data. So we collect some
  231. 7:20data which are these pairs X and Y
  232. 7:22pairs. And each one of them as you can
  233. 7:24imagine is some kind of image and some
  234. 7:25kind of label for that image. Right?
  235. 7:27This is a picture. This is an image of a
  236. 7:28cat. This is an image. This is an image
  237. 7:30of a dog. Okay. The X is the image. The
  238. 7:32Y is the label set that's there. Okay.
  239. 7:35Now we'll use this notation, this kind
  240. 7:38of superscript notation. What that
  241. 7:39indicates is that's the first example,
  242. 7:41that's the nth example if I ever want to
  243. 7:43refer to them, right? The i example, the
  244. 7:45J example, something like that. Okay,
  245. 7:47[gasps]
  246. 7:48now I've been pretty vague, right? This
  247. 7:49is a pretty vague definition. I mean,
  248. 7:51it's a pretty general one. X is just
  249. 7:52some set and Y is just some set. But
  250. 7:54that's what's kind of exciting about
  251. 7:56supervised machine learning. It applies
  252. 7:57to a huge range of different data types
  253. 8:00and different things that you may want
  254. 8:01to do with it. Okay.
  255. 8:04All right.
  256. 8:07So, among all of those hypotheses that
  257. 8:11are out there, and if you've taken any
  258. 8:13kind of math classes, you know, there's
  259. 8:14a huge range of hypotheses that go from
  260. 8:16one set X to another set Y, right? Just
  261. 8:18a, you know, potentially uncountably
  262. 8:20many if Y is some big continuous set.
  263. 8:22So, the question is, what makes a good
  264. 8:24prediction? Now, this is a subtle
  265. 8:26problem. This is something we're going
  266. 8:28to talk about for a while. We're going
  267. 8:29to talk about good in several contexts
  268. 8:31and more refined notions of good over
  269. 8:33the next couple of weeks. But kind of
  270. 8:35intuitively what we're after in this
  271. 8:36modeling question is we want a function
  272. 8:38that you know to use a fancy word that
  273. 8:40you don't have to use now. We want it to
  274. 8:41generalize. So what we would like to do
  275. 8:43is that this x and y that we're
  276. 8:44selecting here, we'd like it to come
  277. 8:46from some set that's representative in
  278. 8:48some way that we can make precise. And
  279. 8:50then later when we're shown new images
  280. 8:52that are not from that original set,
  281. 8:54this thing is going to label it. it's
  282. 8:56going to be able to look at a new
  283. 8:57picture of a cat and say, "Yep, that's a
  284. 8:58cat, not a dog." Okay. And so that's
  285. 9:00something that's good that
  286. 9:01generalization that's going to be the
  287. 9:02heart of machine learning. I train on
  288. 9:04this this training set. I look at that.
  289. 9:06I'm going to compute my H. And somehow
  290. 9:08some properties of the training set and
  291. 9:10the hypothesis class, that is the hes
  292. 9:12that I consider are going to mean that I
  293. 9:14generalize to new and unseen things. And
  294. 9:16when I mentioned before, what we'll talk
  295. 9:18about first, there's hes that are really
  296. 9:19simple. And we can prove that if they're
  297. 9:21really simple, they have this kind of
  298. 9:23property that they will transfer. But
  299. 9:25what has been the you know revolution
  300. 9:27for the last 10 or so years maybe more
  301. 9:30is actually training hes that are really
  302. 9:32really big. That means that they are
  303. 9:33encoded by huge programs, huge numbers
  304. 9:35of what are called weights. And we're
  305. 9:36going to talk about that as we get into
  306. 9:38more of the class. Okay. So this is the
  307. 9:41way it works. This is the basic setup.
  308. 9:43Please feel free to ask questions about
  309. 9:45it. But this is kind of how it works.
  310. 9:46And so we're going to study this right
  311. 9:48now. You could say what else would you
  312. 9:49do? Well, in the second half of the
  313. 9:51course, I'm going to tell you about what
  314. 9:52happens if you don't have any labels.
  315. 9:53Okay? So this is just one of many
  316. 9:55machine learning setups you can have but
  317. 9:57this is by far one of the most you know
  318. 9:59common and you know kind of working
  319. 10:01actually it's kind of interesting in my
  320. 10:02career is like this is a technology you
  321. 10:04can reliably do this in a number of
  322. 10:05situations you can make these functions
  323. 10:07that generalize that's really exciting
  324. 10:10all right now a little bit of
  325. 10:12terminology here at the end if y is
  326. 10:14continuous we call it a regression
  327. 10:16problem that's real numbers prices we're
  328. 10:18going to look at regression today the
  329. 10:20math for that is just easier the
  330. 10:22derivatives the things that we have
  331. 10:23comput. They're just easier to compute
  332. 10:25by hand. They'll hopefully be a bit more
  333. 10:26familiar. If y is discreet, then it's a
  334. 10:29classification problem.
  335. 10:31Classification problems are probably
  336. 10:34things that you end up solving more in
  337. 10:35machine learning these days for a
  338. 10:37variety of reasons. Like for example,
  339. 10:39the way like your chat GPT works is that
  340. 10:41it actually is has a classifier head at
  341. 10:42the end that is guessing what's the next
  342. 10:44word. That's the kind of way this stuff
  343. 10:46works. Okay, we'll talk about how those
  344. 10:48those systems work as well. All right,
  345. 10:51awesome.
  346. 10:52Okay, so let's look at a first example
  347. 10:54of this using probably the most
  348. 10:56canonical, you know, the most widely
  349. 10:58used data set out there, which is this
  350. 10:59housing data set, the Ames housing data
  351. 11:01set. And this is real data. You can you
  352. 11:03can pull it down and play with it. All
  353. 11:04right.
  354. 11:06All right. So, these are not Bay Area
  355. 11:08prices, but uh you know, don't hold that
  356. 11:10against themselves. Okay. Now, what I've
  357. 11:13done here is I've just taken some data
  358. 11:14and I put it in, you know, a Python data
  359. 11:16frame. Um, you can do the same. And then
  360. 11:20I just looked at, you know, the sale
  361. 11:21prices and the lot sizes and all the
  362. 11:23rest. And then on the right, I visualize
  363. 11:24the data. Okay? So, just one thing in
  364. 11:26general, like if you have the ability to
  365. 11:28do it, look at your data. I can't stress
  366. 11:30that enough. Look at your data. It's
  367. 11:32like you're pretty good pattern
  368. 11:33recognizers, take advantage of that.
  369. 11:36Okay? Even if you're building some
  370. 11:37complicated crazy AI system, still look
  371. 11:39at your data. You will discover things.
  372. 11:41All right? So, we're looking at this
  373. 11:42data here. We have the sale price. We
  374. 11:44have the lot area. And I've plotted it
  375. 11:45in some way there. Okay. All right. Now,
  376. 11:50we need to get one more character in our
  377. 11:51story. We need to get a hypothesis. So,
  378. 11:53one popular choice are these linear
  379. 11:56hypotheses, right? And you may think
  380. 11:58that these linear hypotheses, you know,
  381. 11:59they seem relatively simple. That's the
  382. 12:01hope, but they're actually industrially
  383. 12:03used, right? These are kind of the
  384. 12:05workhorse of what we're going to do.
  385. 12:06These kind of linear classification
  386. 12:07models, they show up inside pretty much
  387. 12:10every model that you use. The last step
  388. 12:11is this linear model. So, these are
  389. 12:13actually, in spite of their simplicity,
  390. 12:15pretty widely used. Okay. So, what is
  391. 12:18age? We call them linear. It's a little
  392. 12:19bit of an abusive terminology. You'll
  393. 12:21see why we do in a second. It's because
  394. 12:22of a convention. This is technically an
  395. 12:24aphine function because it has this
  396. 12:26little offset here. The way it works is
  397. 12:29you feed me an x. Okay, so let's imagine
  398. 12:31x is a scaler and then I'm going to
  399. 12:33multiply it here times that theta 1 then
  400. 12:36add it to theta 0. Okay, and so this
  401. 12:38gives me something wherever my x is.
  402. 12:40This is a very simple function that's
  403. 12:41taking in scalers, but this is my
  404. 12:42hypothesis class. Okay.
  405. 12:45[sighs and gasps]
  406. 12:45All right. So
  407. 12:48if you look at an example prediction
  408. 12:50here, let's say that you wanted to in
  409. 12:52our setting, we have a bunch of x's that
  410. 12:53are right here that we want to predict.
  411. 12:55And we like to predict some function
  412. 12:56that goes from the size of the of the
  413. 12:58house to the price in whatever units.
  414. 13:00Don't don't worry. Don't think too hard
  415. 13:02about what size and price mean here. I
  416. 13:03don't know like what what reasonable
  417. 13:05mapping I I don't think I've ever been
  418. 13:06to as um but these are apparently real
  419. 13:09prices. Okay. [gasps] All right. So we
  420. 13:11have to do this example prediction from
  421. 13:13size to price. So, we want to get a
  422. 13:14function, a linear function that takes
  423. 13:16this size in as the x and then out comes
  424. 13:19a price. Clear enough what our goals
  425. 13:21are? All right,
  426. 13:23man. That's got to stop. [sighs] All
  427. 13:26right. So, notice that this prediction
  428. 13:28now is instead of all the wild
  429. 13:31hypotheses that could be out there,
  430. 13:32right? There's a ton of ton of things
  431. 13:33that go from functions that go from a
  432. 13:35set of real numbers to another set of
  433. 13:36real numbers. We've now really
  434. 13:38constrained it. It's entirely defined.
  435. 13:40The fancy word for this is
  436. 13:42parametrically. It's entirely defined by
  437. 13:44these two little parameters. No matter
  438. 13:46how many houses we see, we're only going
  439. 13:48to, as we'll say, learn those two
  440. 13:50parameters or estimate those two
  441. 13:51parameters and that's going to be the
  442. 13:53way that we transfer. Okay, so it's a
  443. 13:55huge reduction in the space of functions
  444. 13:56that we've just done. So like it seems
  445. 13:58like maybe you know not very much
  446. 14:00happened on the slide, but like we went
  447. 14:01from uncountably many functions, we
  448. 14:03still have uncountably many functions,
  449. 14:04went to this really small class of
  450. 14:05lines. Okay, so what do these look like?
  451. 14:08By the way, just to just to draw them,
  452. 14:09just to make sure you're extra clear why
  453. 14:11we call them linear functions. Well,
  454. 14:12where do they what is their value at
  455. 14:14zero? Well, that value of zero is going
  456. 14:16to be theta kn. And then it's going to
  457. 14:18have a slope. And that slope at one is
  458. 14:20going to be here theta 0 plus theta 1.
  459. 14:23Okay. This is one.
  460. 14:26Hopefully that's clear.
  461. 14:28All right. Awesome.
  462. 14:31Okay. So, with that, we're going to fit
  463. 14:34a line to this data. Now, which line are
  464. 14:36we going to fit? We're going to come to
  465. 14:37that in a second. But intuitively, what
  466. 14:39should that line? What would be a good
  467. 14:40line to fit? Well, it's one that when we
  468. 14:42give it an X, right? So, I feed it in
  469. 14:44some X here. Let's say this point. When
  470. 14:47I look at the line, my prediction is on
  471. 14:49the line. So, if I fed it in something
  472. 14:51at 1200, my prediction would be this
  473. 14:52point here. If I could draw straight
  474. 14:54with this thing. Let me try to draw a
  475. 14:55little bit straighter. Anyway, my
  476. 14:57prediction would be on that line. Is
  477. 14:59that clear? When I take the X, I put it
  478. 15:00into the hypothesis, it gives me that
  479. 15:02line. Cool. Now, what would be good?
  480. 15:05Well, one thing that would be good is
  481. 15:07because we have this training set is
  482. 15:08that every time we put a point into this
  483. 15:10data set, whether it's here, here, we
  484. 15:12didn't do such a good job, right? We're
  485. 15:13really, really far from the line here
  486. 15:15and here, we seem to have done a really,
  487. 15:17really great job. Okay? And so that
  488. 15:20we're going to try and do those errors
  489. 15:21that we make. Those are how far off we
  490. 15:23are in the price. We're going to try and
  491. 15:25minimize those. This has a really fancy
  492. 15:28name. It's called empirical risk
  493. 15:29minimization. I don't think you need to
  494. 15:31know that name, but you'll sometimes
  495. 15:32hear me slip and say erm, and that's
  496. 15:34what empirical risk minimization means.
  497. 15:36Okay, cool.
  498. 15:38So, here's a nicer picture that I
  499. 15:40generated where the functions behaving
  500. 15:42much more nicely. So, each point here is
  501. 15:44one of those data points that we saw.
  502. 15:46And here we're trying to minimize this
  503. 15:49little line through it. Right? Now, the
  504. 15:52idea of course, which is nice about the
  505. 15:54line, is if you just had the training
  506. 15:56set, the way you would potentially do a
  507. 15:58prediction, there's actually not a bad
  508. 15:59way to do it. We won't talk about it too
  509. 16:00much today. You do what's called nearest
  510. 16:02neighbors. You take in an X, you kind of
  511. 16:04look in the neighbor of X, maybe you
  512. 16:05average them. That could be [snorts] an
  513. 16:07interesting prediction function. What's
  514. 16:09nice about H is it extends to every
  515. 16:11single value and it extends in kind of a
  516. 16:13smooth way. And so hopefully what that
  517. 16:15means is that if there's lots of
  518. 16:17girrations here, we capture that main
  519. 16:19trend. we get the main part of the of
  520. 16:21the of the prediction which makes it
  521. 16:23kind of a good predictor overall. Okay.
  522. 16:29All right.
  523. 16:31Now, one other thing which you can see
  524. 16:34from this. Oops. Let's get that guy up.
  525. 16:36There's a slight mismatch there. Sorry
  526. 16:38for that. So, as we look through here,
  527. 16:40if you look at that line that I drew,
  528. 16:44you could argue that you should draw a
  529. 16:45different line. You're like, "Oh, I like
  530. 16:46this other line a little bit better."
  531. 16:47You know, kind of goes Oops. Let's go in
  532. 16:49here. goes like this all the way. Oh, I
  533. 16:52definitely don't like that line. There's
  534. 16:53one line there. Maybe like I draw like
  535. 16:55this. It's a terrible line. So, how did
  536. 16:57I pick this one line that's in there?
  537. 16:59And that's what we're going to do this
  538. 17:00erm this this thing. We're going to look
  539. 17:01at the data and we're going to have to
  540. 17:03be able to compute that. But the point
  541. 17:04is is no line looks like it's absolutely
  542. 17:06perfect. And that's going to be a theme
  543. 17:08throughout most of what we do. Your data
  544. 17:10has a couple of different things in it.
  545. 17:12Sometimes it's imperfect because we're
  546. 17:14not modeling something. So probably the
  547. 17:16rate that you're willing to pay for a
  548. 17:18house depends on more than like the lot
  549. 17:20size, right? Like you don't just care
  550. 17:22how big the lot is. You may care if it's
  551. 17:24a nice house, if it's close to things
  552. 17:26you care about, all the rest. We'll come
  553. 17:28back in a minute about how we
  554. 17:29incorporate that. But the other thing is
  555. 17:31even if we added in all those features,
  556. 17:33usually there will be some error. And so
  557. 17:35we're going to worry a lot in this class
  558. 17:37about the error of how we do in these
  559. 17:39predictions and how well the error on
  560. 17:41our training sets manifests on the what
  561. 17:43are called the test sets when we take it
  562. 17:44out into real life.
  563. 17:47Okay. All right. Let's look at something
  564. 17:49slightly more interesting. So here we're
  565. 17:51going to add in the number of bedrooms.
  566. 17:52Okay. So what do you think is going to
  567. 17:54happen? Well, now we need a function
  568. 17:56that can take in pairs instead of single
  569. 17:58values. So we need to change our
  570. 17:59hypothesis class a little bit.
  571. 18:03Lot size and size. I guess size is the
  572. 18:05refers to the house size. Apologies. Now
  573. 18:07we have this function here. Okay. So how
  574. 18:09does it work? Well, if you give me you
  575. 18:11feed me an x, then I take this guy, put
  576. 18:13him here, put this one here, put this
  577. 18:15one here,
  578. 18:18and then I hopefully have weights. Now
  579. 18:19I'm parameterized by that came all back
  580. 18:22by just those weights now. So now my
  581. 18:23function is parameterized by four
  582. 18:25weights.
  583. 18:27And hopefully it's going to fit my data
  584. 18:29a little bit better. Right? The reason
  585. 18:30it should be a little bit better is I
  586. 18:31could always set some of those weights
  587. 18:33to zero, right? And get back to the case
  588. 18:35where I only had two weights. So it
  589. 18:36seems like it's a larger or more
  590. 18:37expressive class.
  591. 18:40But if lot size and bedrooms influence
  592. 18:42price, we would hope that this
  593. 18:44hypothesis class is richer is the
  594. 18:45terminology. And that would allow us to
  595. 18:47fit our data a little bit better.
  596. 18:50Okay.
  597. 18:51Now this is one of the reasons we call
  598. 18:54these functions linear not aphine is
  599. 18:56that we're going to operate under this
  600. 18:57nice little convention that we're going
  601. 18:59to always assume that x1 x0 is one here
  602. 19:02which we didn't even bother to write and
  603. 19:04I highlight this convention because it's
  604. 19:06used throughout the notes and when I
  605. 19:07deviate from it I will explicitly tell
  606. 19:09you but whenever we're doing regression
  607. 19:10or classification this is what we have
  608. 19:12okay so that allows us to write a nice
  609. 19:15little form like this h of x is just
  610. 19:17this nice little sum we don't have to we
  611. 19:19don't have to special case the the theta
  612. 19:220 term and we can you know easily extend
  613. 19:25from three to however many dimensions.
  614. 19:28Awesome.
  615. 19:31Okay. So that all leads into as I said
  616. 19:34notation. So if this is unfamiliar to
  617. 19:36you please go to the Friday you know
  618. 19:40sessions. A lot of this is us kind of
  619. 19:42recalling notation where you're like oh
  620. 19:44I know that I need to look at it or oh I
  621. 19:46need to go and you know avail myself of
  622. 19:48some of the the extra resources. So this
  623. 19:50notation here we'll write theta but we
  624. 19:53really mean this vector which will also
  625. 19:55live in R4 it's four real numbers okay
  626. 19:58four scalers this thing here notice now
  627. 20:01this x1 we're writing the entire example
  628. 20:04in here and this one is one okay so
  629. 20:07here's one here's 214 here 44 45
  630. 20:14clear enough
  631. 20:16Okay,
  632. 20:19so we're going to call these these
  633. 20:22thetas here parameters and these XI are
  634. 20:24going to be input or features. Remember
  635. 20:25when I talked to you about before which
  636. 20:27I said, you know, you could add in these
  637. 20:28extra features of your problem. That's
  638. 20:30for now going to be a modeling decision.
  639. 20:32Later, we're going to see an interesting
  640. 20:33idea about how we can even learn those
  641. 20:35features. That was part of what was
  642. 20:36called the deep learning revolution
  643. 20:37where you were able to actually imputee
  644. 20:39the representation of your data just by
  645. 20:41throwing more data at it. We'll see that
  646. 20:43in like couple weeks. But for now, it's
  647. 20:46a modeling decision. You come and write
  648. 20:48some rule, some piece of code that takes
  649. 20:50the number and puts it in as the
  650. 20:51feature. And it's not trivial, right?
  651. 20:53Like, you know, it normalized the lot
  652. 20:54size here. It said 45 instead of 45K.
  653. 20:57But in general, they could be even more
  654. 20:59interesting features of your problem.
  655. 21:00And the point is is that extends
  656. 21:02naturally to linear models. Okay, so far
  657. 21:06so good. All right. So, X and Y has a
  658. 21:10fancy name, training example, right?
  659. 21:12It's a it's an element of the training
  660. 21:14set. And the reason that we call it
  661. 21:16training is we're going to take that set
  662. 21:17and we're going to train our model on
  663. 21:19it, right? It's kind of this the
  664. 21:20terminology for it. We're going to try
  665. 21:21and figure out among all the thetas,
  666. 21:23which are our predictions, which ones
  667. 21:24fit our data best. We've been vague
  668. 21:26about what data best means, but there's
  669. 21:28a lot of ways to do that, it turns out,
  670. 21:29but they all follow the same recipe. And
  671. 21:31then, as I said before, I just wanted to
  672. 21:33highlight the example again that Xi and
  673. 21:35Yi are an individual example.
  674. 21:37Okay.
  675. 21:39All right.
  676. 21:43[clears throat] So in notation in the
  677. 21:44class you will always see this is
  678. 21:46something that you will see this is as
  679. 21:47convention. We'll always have n be the
  680. 21:49number of examples. So you should kind
  681. 21:51of get that in your mind and d be the
  682. 21:52number of dimensions or features. So
  683. 21:54we'll consider these either d or d plus
  684. 21:56one dimensional plus one because of that
  685. 21:58convention. Bless you that x0 equals 1.
  686. 22:02Please.
  687. 22:06>> Yeah. So, we will almost always use
  688. 22:08lowercase and we will try not to abuse
  689. 22:11you by having upper and lowerase ends in
  690. 22:13the same lectures. Uh but I cannot
  691. 22:14guarantee that because especially if
  692. 22:16they're copied from handwriting. Uh
  693. 22:17often often we'll use them we will try
  694. 22:19not to use them both in the lecture to
  695. 22:20mean two different things, right? So,
  696. 22:22usually we'll be a little bit
  697. 22:23consistent. Yeah. The other convention
  698. 22:25you'll sometimes see sneak in is P's for
  699. 22:27parameters, but that's like a stats
  700. 22:28convention and it doesn't really matter.
  701. 22:30None of this will be really matter. If
  702. 22:32you're confused, just ask. Cool. All
  703. 22:34right. All right. So, what do these
  704. 22:36things look like? These linear models
  705. 22:38look like Here's one in two dimensions.
  706. 22:40This is about as high as you can really
  707. 22:42reasonably look on your data set. So,
  708. 22:44here we have a bunch of different
  709. 22:45points. We have price. We have, you
  710. 22:47know, bedrooms and square feet at the
  711. 22:48bottom. And I've I've plotted a plane
  712. 22:51here. Okay, an aphine surface. You could
  713. 22:54imagine a classification that was some
  714. 22:56kind of crazy curve. We'll see how to do
  715. 22:58that later. There's all kinds of nice
  716. 22:59ways to do that, but for right now,
  717. 23:01we're looking at these big linear
  718. 23:02models. And you see this is kind of a
  719. 23:04good thing in the same way. So, how
  720. 23:05would you define error? Well, how close
  721. 23:07is your prediction? If you took that
  722. 23:08point, that feature point, and you went
  723. 23:10into space, how far away are you from
  724. 23:12this plane? Now, we're going to consider
  725. 23:15as we get here, like modern models, as I
  726. 23:18mentioned, just got bigger. They have,
  727. 23:21you know, thousands, hundreds of
  728. 23:22thousands, millions, billions, trillions
  729. 23:24of parameters. Some of the models now
  730. 23:25have trillions of parameters in them.
  731. 23:27They're not all fed into a line like
  732. 23:29this, but they'll conventionally have
  733. 23:3110,000 dimensional spaces that are the
  734. 23:33last piece of the model.
  735. 23:36That's really hard to visualize. So,
  736. 23:38we're going to use a lot of vector
  737. 23:39notation and other things to deal with
  738. 23:40it. And that's why we use this notation
  739. 23:42so you can simplify and think about it
  740. 23:43and understand how to manipulate it in
  741. 23:45one or two dimensions. You can kind of
  742. 23:46picture it. Okay. But high dimensional
  743. 23:48space is a weird thing.
  744. 23:51Okay.
  745. 23:52All right.
  746. 23:54[gasps] Okay.
  747. 23:58So let's get back to it. So we want to
  748. 23:59choose some theta such that as we talked
  749. 24:02about multiple times,
  750. 24:04h of x is approximately equal to y. If I
  751. 24:06give you an x and a new thing, you want
  752. 24:08to be approximately close. You like the
  753. 24:09price to be close and so on. So how do
  754. 24:13we do it?
  755. 24:15Well, here comes le squares. Okay, so
  756. 24:17maybe you've seen this before. I hope
  757. 24:19you've seen before. By show of hands,
  758. 24:20who's seen le squares before? Awesome.
  759. 24:22Fantastic. All right, good. So le
  760. 24:25squares super super conventional thing
  761. 24:27if you haven't seen it before please
  762. 24:28again Friday is a good place to look at
  763. 24:30it we'll cover it again just in our
  764. 24:31notation to make sure everything is is
  765. 24:33is all right now what happens so you
  766. 24:37have this idea here j theta that's going
  767. 24:39to be the loss that you incur and it's
  768. 24:41broken down into terms on every single
  769. 24:43element of the data set and here you
  770. 24:46have h theta of x i minus yi so this is
  771. 24:49the error remember we kept drawing those
  772. 24:51pictures like this guy right here these
  773. 24:53two things are the same. This error,
  774. 24:55oops, that looks terrible. We'll wait
  775. 24:58for it to go away. Anyway, so a state
  776. 25:01the x ius yi, that's the prediction
  777. 25:03error. Then what we do is we square it.
  778. 25:05And we square it because we want to have
  779. 25:07something that we have a positive or
  780. 25:09non- negative number there. It could be
  781. 25:10zero if it were exactly equal. And then
  782. 25:12what we're going to do is we're going to
  783. 25:13try and pick the theta that minimizes
  784. 25:16it. Okay, just by show of hands, how
  785. 25:18many people know the arg notation?
  786. 25:21Perfect. All right, we're doing great.
  787. 25:23Okay. So then in that case, you know
  788. 25:24exactly what goes on here. We're going
  789. 25:26to solve over j theta and we're going to
  790. 25:27try and pick the thing that minimizes
  791. 25:30this function. And so intuitively what
  792. 25:31that says is we're going to pick the
  793. 25:33theta so that on average over all of our
  794. 25:35training points,
  795. 25:37we minimize that squared error. Okay? So
  796. 25:39we're going to instead of measuring the
  797. 25:40error just as a sum, we're going to
  798. 25:41square that distance. So things that are
  799. 25:43really close, they're going to, you
  800. 25:45know, go down. Things are really big,
  801. 25:46we're going to pay for.
  802. 25:48Okay? Now, why do we do squared? Oh, go
  803. 25:50ahead.
  804. 25:51>> Sorry. I might be um uh you might be
  805. 25:56saying it but the squaring because
  806. 25:59presumably you could take another power.
  807. 26:01>> Yeah, sure thing to even be even
  808. 26:04stricter.
  809. 26:06>> Yeah, exactly. Right. So the the
  810. 26:07question is is why did you pick two here
  811. 26:09and and why why did you do it? That's a
  812. 26:11great question. So could you pick other
  813. 26:13values? Certainly you could pick another
  814. 26:14power. You could pick the absolute
  815. 26:16value, right? That's a fine thing to do
  816. 26:17as well. The you know that's a that's a
  817. 26:19value regression kind of thing. two has
  818. 26:22a nice property which we're cheating on
  819. 26:24for a lot of le squares. two we can
  820. 26:25solve exactly and that just turns out to
  821. 26:27be because two is as you'll see in a
  822. 26:29second when we compute all the
  823. 26:30derivatives and do all the nice things
  824. 26:32the derivative has a really nice form or
  825. 26:34the gradient has a really nice form and
  826. 26:36that's going to let us solve it exactly
  827. 26:38historically also if you care like many
  828. 26:40of the you know people who were doing
  829. 26:43science way back when when they were
  830. 26:44doing you know Gaus was doing le squares
  831. 26:46because of that computational advantage
  832. 26:48they could predict all kinds of nice
  833. 26:49things it was an easy problem now you
  834. 26:51could look at that and I will derive for
  835. 26:53you in the next class I'll der for you
  836. 26:55how that le squares comes up and relates
  837. 26:57to a fancy statistical assumption which
  838. 27:00is that the errors are approximately
  839. 27:01normal which we'll come in and talk
  840. 27:03about a little bit later and so you'll
  841. 27:05be able to derive other regressions that
  842. 27:07are in there but the special things
  843. 27:08about le squares are it's extremely
  844. 27:10popular uh in like history not like you
  845. 27:13know I think it's great right now like
  846. 27:14it's extremely popular in history and
  847. 27:16it's very easy to solve okay and that
  848. 27:18the reason I hammer on this so hard is a
  849. 27:20lot of machine learning if you look at
  850. 27:21it and and I think honestly I was guilty
  851. 27:23of this so I grew up as like a math math
  852. 27:24person. And so when I came to machine
  853. 27:27learning and AI, I kind of had this idea
  854. 27:29that I was like, well, why are you doing
  855. 27:30like simple algorithms? Why are you
  856. 27:32like, you know, worried about these
  857. 27:33other things? And I've come to realize
  858. 27:36that really what goes on in machine
  859. 27:37learning that makes it so powerful is
  860. 27:38this trade-off between kind of how we
  861. 27:40compute things and kind of how we
  862. 27:42predict them. And actually when machine
  863. 27:44learning to me really intellectually
  864. 27:45came into its own as a field was when it
  865. 27:48broke away a little bit from stats and
  866. 27:49started to train these crazy large
  867. 27:51models and understand more what were the
  868. 27:53computational limits of what kind of
  869. 27:55models we could have. So sorry for the
  870. 27:56long digression but I love questions so
  871. 27:58I'm just pumping trying to get more.
  872. 27:59That was a wonderful question. Please
  873. 28:03>> yeah the 1/2 is there just by
  874. 28:04convention. So one of the things is this
  875. 28:06arg is insensitive to constants. If I
  876. 28:09change this to, you know, 25 over 12, it
  877. 28:12would, this would be totally unaffected.
  878. 28:13I'd get the same theta. So the constant
  879. 28:15doesn't matter. The reason we do this is
  880. 28:17it's this convention. When we compute
  881. 28:19the derivative in a minute, these things
  882. 28:21are going to cancel and that just makes
  883. 28:22things a little bit nicer. It's kind of
  884. 28:23like the cooking show view of math where
  885. 28:25we like set it up so that it looks nice
  886. 28:27and when you do it by yourself, it's
  887. 28:28always a gigantic mess and then you you
  888. 28:30clean it up later. That's all it's there
  889. 28:31for.
  890. 28:32>> Please. uh with RM like the theta that
  891. 28:35you kind of put it to minimize like are
  892. 28:37they discrete or
  893. 28:39>> oh fantastic yeah so one of the things
  894. 28:40is is remember our hypothesis class
  895. 28:42right now the way we picked it was to be
  896. 28:44continuous so these can be any real
  897. 28:46value that we want and so when we plug
  898. 28:48these things in here we're going to have
  899. 28:49continuous values so we're going to do
  900. 28:51continuous optimization discrete
  901. 28:53optimization you can do there are there
  902. 28:55are procedures that people do fancy ones
  903. 28:57you know integer linear programming yada
  904. 28:59yada yada those are generally pretty
  905. 29:01hard most modern machine learning One of
  906. 29:03the tricks are even when you have
  907. 29:04something that's discreet like our
  908. 29:06models that work on words underneath
  909. 29:08their covers they're actually doing
  910. 29:09something continuous and so the
  911. 29:10parameter space will almost always be
  912. 29:12continuous so you can compute things
  913. 29:13like gradients and derivatives and and
  914. 29:15for other optimization reason it's a
  915. 29:17it's a very insightful question. Yeah
  916. 29:20please
  917. 29:22larger differences get more
  918. 29:26>> yeah that's a great question to think
  919. 29:28about. So these error functions this is
  920. 29:30this this thing here when we incur an
  921. 29:32error in prediction the question is how
  922. 29:33much penalty do we pay right now we're
  923. 29:35paying with this square and as we've
  924. 29:36talked about it's for computational and
  925. 29:38modeling reasons kind of a good mix you
  926. 29:41could put other things in there I don't
  927. 29:42want to make too big a deal out of it
  928. 29:44because honestly when you get to a lot
  929. 29:46of data the the form of this doesn't
  930. 29:49really matter too much there is one very
  931. 29:52important form that we'll cover next
  932. 29:53lecture which is the is a is a for
  933. 29:56classification when you have a discrete
  934. 29:58that and then you do something called
  935. 29:59softmax and softmax is probably the
  936. 30:02function that I use most in my day um
  937. 30:05when you build these systems and
  938. 30:06probably if you are doing all your
  939. 30:07homework with chat GPT like the rest of
  940. 30:09your classmates then you probably use it
  941. 30:11too okay because that's what the that's
  942. 30:12the workhorse that's underneath there
  943. 30:14but this one is just the one we're going
  944. 30:15to do for for computational reason it
  945. 30:17has a nastier derivative
  946. 30:20other questions this is awesome please
  947. 30:22ask me questions
  948. 30:24all right
  949. 30:26we'll keep
  950. 30:28All right. So what do we see? We're at
  951. 30:29the end of our linear regression.
  952. 30:32We saw our first hypothesis class aphine
  953. 30:35and linear. We're going to see many many
  954. 30:37more richer classes throughout this. And
  955. 30:39richer means that they're more
  956. 30:40expressive. Means not only do they have
  957. 30:42more parameters, which as I've hinted
  958. 30:43several several times, modern models do,
  959. 30:46but they're actually going to have
  960. 30:47crazier surfaces as well. So they're not
  961. 30:49just going to be lines, they're going to
  962. 30:50have kind of crazy interesting surfaces
  963. 30:52underneath them. We refreshed ourselves.
  964. 30:54And as I said, if these part was
  965. 30:55unfamiliar to you, you were looking at
  966. 30:57this and you're like, I've never heard
  967. 30:58anything this guy's talking about. Okay,
  968. 31:00maybe I did a bad job. I'm willing to
  969. 31:01admit that. But if I didn't do a
  970. 31:03terrible job, then look and say, maybe I
  971. 31:05need to refresh. And I cannot tell you
  972. 31:07how much I want you to if you need to go
  973. 31:09take advantage of those TA resources and
  974. 31:11others because putting in a little bit
  975. 31:13of investment now makes the class go a
  976. 31:14lot lot more smoothly for you.
  977. 31:18We also saw something that was this
  978. 31:19paradigm that you guys picked up on
  979. 31:20right away, which is great, which is
  980. 31:22that a good hypothesis is one that's
  981. 31:24close to the data. And I want to point
  982. 31:26out something which is a little bit kind
  983. 31:27of spooky about what we're going to
  984. 31:28claim here. And this is like a
  985. 31:30foundation of statistics. So it's not
  986. 31:31that spooky, but is if we believe that
  987. 31:34training set, if we fit our hypothesis
  988. 31:37on that training set, we believe it's
  989. 31:39going to work well in the real world.
  990. 31:41That's a little bit of a leap of faith.
  991. 31:42And so we'll justify that leap of faith
  992. 31:44in all kinds of ways. some of which I'll
  993. 31:46give you a lecture in the middle about
  994. 31:48how they can go terribly wrong and I
  995. 31:49will pick from my own you know work
  996. 31:51where we thought we were selecting data
  997. 31:53in a way at training time that was going
  998. 31:55to match test and we were completely
  999. 31:57wrong and made fools of ourselves okay
  1000. 31:58so you know economic indicators and all
  1001. 32:01kinds of fun stuff like that and so that
  1002. 32:03can happen but I want to be aware this
  1003. 32:05paradigm is what machine learning does
  1004. 32:06says look this training set is somehow a
  1005. 32:08reflection it's captured the real world
  1006. 32:10you have to believe that and then
  1007. 32:12fitting on that data set is somehow
  1008. 32:13going to carry out something nice in the
  1009. 32:15real world. Okay, it's like the
  1010. 32:16cornerstone of statistical estimation
  1011. 32:18and inference. Cool. Great. We also
  1012. 32:22talked a bunch about objective
  1013. 32:23functions, J. We're going to see those
  1014. 32:25over the next lecture. We're going to
  1015. 32:26see, you know, the canonical one for um
  1016. 32:28regression today and we'll see the
  1017. 32:30canonical one for classification next
  1018. 32:31week.
  1019. 32:34Okay. All right. Now, as I mentioned,
  1020. 32:37and I I'm biased here, so I should also
  1021. 32:39tell you. So like what I work on in
  1022. 32:40research, as I mentioned, is I really
  1023. 32:42like the problems of scaling up these
  1024. 32:44models. We've done all kinds of things
  1025. 32:45underneath the covers to do that to make
  1026. 32:47these large models that today run you
  1027. 32:49know on GPUs. We've contributed our
  1028. 32:51little brick to that out of my lab which
  1029. 32:52I'm you know very proud of wonderful
  1030. 32:54students and and collaborators. But one
  1031. 32:57of the things that you'll see in machine
  1032. 32:59learning is this ability to compute
  1033. 33:01these functions at large scales is
  1034. 33:03critically important. That's been the
  1035. 33:05revolution that's allowed us to break
  1036. 33:07away from doing small kinds of problems
  1037. 33:09to the large. And that's what we're
  1038. 33:10going to cover over the next couple
  1039. 33:11weeks. But we'll start where the field
  1040. 33:12started.
  1041. 33:13>> [snorts]
  1042. 33:14>> What that boils down to is basically
  1043. 33:17solving a set of these equations. That
  1044. 33:18arg min'll
  1045. 33:21talk about a loss function. Being able
  1046. 33:24to take a loss function and run you know
  1047. 33:26an optimization procedure fancy one
  1048. 33:29called backrop that you'll learn in the
  1049. 33:30course underpins pretty much 99% of the
  1050. 33:33models that you're likely to encounter
  1051. 33:34today. Okay. So we're going to see the
  1052. 33:36warm-ups to how you get to backrop the
  1053. 33:38simplest thing that starts there which
  1054. 33:40is the computation of a derivative. And
  1055. 33:41then in this case, we'll also see how to
  1056. 33:43solve these equations exactly. That's
  1057. 33:45very unusual that you can do that, but
  1058. 33:47it will basically be an exercise with
  1059. 33:49some of that notation. So if I haven't
  1060. 33:50made it clear, we'd really love you to
  1061. 33:52use those those Friday lectures. All
  1062. 33:54right, I'm not getting paid if you go. I
  1063. 33:56know you probably at this point you're
  1064. 33:56like, "Oh, you must be getting paid."
  1065. 33:57Like, no, no, it's just I think it's
  1066. 33:58good for you. Okay. All right. So, let
  1067. 34:00me solve these le squares problems.
  1068. 34:03All right. So, I'm going to draw a
  1069. 34:05little picture here.
  1070. 34:08All right. So let's take a function.
  1071. 34:11This will be our quadratic which is kind
  1072. 34:13of like our our classical function. And
  1073. 34:16uh you know we'll look at this this kind
  1074. 34:17of bullshaped function. Now when we
  1075. 34:20compute the the quadratic so you know
  1076. 34:22we'll take something here which like you
  1077. 34:24know we'll say it's uh x^2 what am I
  1078. 34:27going to write here? Oops. So it's like
  1079. 34:29x^2 plus I don't know some theta 0
  1080. 34:32whatever something right here. Okay. Now
  1081. 34:34I can make this call this 2x squ I don't
  1082. 34:36know whatever you want. Okay. Now this
  1083. 34:38is f. You all remember how to compute
  1084. 34:40derivatives for this I hope but if you
  1085. 34:42don't remember the picture you pick a
  1086. 34:44point let's say your first point here
  1087. 34:47which is going to be theta 0 which is
  1088. 34:49going to be our first guess. Okay, you
  1089. 34:52then
  1090. 34:55pick this line which is the gradient,
  1091. 34:57the best linear approximation, right?
  1092. 34:59You remember your your tailor expansion.
  1093. 35:01It's you're going to be f of theta 0 and
  1094. 35:04then you're going to be plus fprime of
  1095. 35:06theta 0 time x - x theta or theta. Whoa,
  1096. 35:11that's bad.
  1097. 35:13All right, I'm going to rewrite this. F
  1098. 35:16of theta 0 plus frime of theta 0. And
  1099. 35:20then I was just trying to write theta
  1100. 35:21minus theta 0. Okay, nothing deep going
  1101. 35:23on. This Taylor's rule for the first
  1102. 35:25first derivative. And that's that line
  1103. 35:27that I've drawn there. And if you don't
  1104. 35:28remember that, remember it's the secant
  1105. 35:29that comes in and approximates the
  1106. 35:30gradient. Okay, slide it down in our
  1107. 35:32head. Okay, once we do that, we have the
  1108. 35:35direction of maximal increase. And so
  1109. 35:37then we plug that that character in that
  1110. 35:40this fun gradient here.
  1111. 35:43We're going to plug that in in this
  1112. 35:44notation.
  1113. 35:46Okay. Now J may have some higher
  1114. 35:50dimensional structure. So this guy is
  1115. 35:52not a derivative of one variable. It's a
  1116. 35:54partial derivative. If you don't
  1117. 35:55remember that thing, please go ahead and
  1118. 35:57look it up. And the idea here is that
  1119. 35:59what we're going to do is we're going to
  1120. 36:00iterate and walk down in the direction
  1121. 36:03until we get to the bottom here. Okay.
  1122. 36:06And so what's happening is we start with
  1123. 36:08initial guess. I didn't start at zero. I
  1124. 36:09should have slid this over at zero. We
  1125. 36:11start with initial guess. Could be a
  1126. 36:12random number. Could be zero.
  1127. 36:14Interestingly, in modern models, where
  1128. 36:15you start matters. doesn't matter for le
  1129. 36:17squares does matter in general. And then
  1130. 36:20we're going to iterate these equations.
  1131. 36:21We're going to start with our current
  1132. 36:22guess. We're going to compute this
  1133. 36:24gradient or derivative and we're going
  1134. 36:26to try and minimize it. So we're going
  1135. 36:27to walk in the opposite direction. How
  1136. 36:29far are we going to walk?
  1137. 36:32This thing here is going to tell us this
  1138. 36:34is called a step size. Okay.
  1139. 36:38This guy here is a step size.
  1140. 36:41All right. Now, if you look at this and
  1141. 36:43you're mathematically minded, you may
  1142. 36:44think that this is horrifying.
  1143. 36:46You're computing only first order
  1144. 36:48information, right? You're just looking
  1145. 36:49at a gradient.
  1146. 36:51You are only taking a step size. That's
  1147. 36:53like, you know, not searching to the
  1148. 36:55most minimal part you could get, what's
  1149. 36:57called line search. You're not doing
  1150. 36:58that. You're just kind of grabbing some
  1151. 37:00information about the function locally
  1152. 37:01and taking a quick step. And the reason
  1153. 37:03that can be so disastrous, you may have
  1154. 37:05been told in your calculus classes, is
  1155. 37:07because you could have a function that
  1156. 37:08looks like this. And then here you walk
  1157. 37:11down and you get stuck in a local
  1158. 37:12regime. Okay? So hopefully this is
  1159. 37:14familiar to you. By show of hands, do
  1160. 37:16people understand the difference
  1161. 37:17convexity, non-convexity? Awesome.
  1162. 37:18Perfect. Okay. So that's how it looks.
  1163. 37:22All right. Fantastic. We do this for
  1164. 37:24every single component and we update.
  1165. 37:27That's how we solve it.
  1166. 37:29All right.
  1167. 37:33So when we do this, we get this learning
  1168. 37:35rate or step size. You'll hear me say
  1169. 37:37those terms mean the same thing. I wish
  1170. 37:39that they were unified in my head. Um
  1171. 37:41but I will use them interchangeably. Um,
  1172. 37:43apologies in advance. Here's me
  1173. 37:45computing the derivative. So, those
  1174. 37:47folks who were asking, as you were
  1175. 37:48asking earlier, what's this two for?
  1176. 37:50It's so that I can cancel this guy.
  1177. 37:51Okay, that's all that's happening. I
  1178. 37:54have my error term here and then a
  1179. 37:56gradient here. Now, hopefully this is
  1180. 37:58not calculus that's bothering you. If it
  1181. 38:00bothers you, again, see the Friday
  1182. 38:01lectures. But there's one thing here and
  1183. 38:04the way that I've written it that I want
  1184. 38:05to call out because we're going to use
  1185. 38:06it again and again. this thing here,
  1186. 38:09this error term, you basically always
  1187. 38:11have this in a ton of these rules. You
  1188. 38:13have the error, what you have, and then
  1189. 38:15you have it multiplied by something
  1190. 38:17here. It's just going to be the gradient
  1191. 38:18of the function. But in general, for a
  1192. 38:20nonlinear, there' be some extra
  1193. 38:22[laughter] extra stuff there. Okay? And
  1194. 38:24you'll see this form again and again.
  1195. 38:26So, it's like kind of an artificial way
  1196. 38:27to write it, but it's an important
  1197. 38:28artificial way to write it. Okay? So,
  1198. 38:30what's going on here? When I compute the
  1199. 38:32derivatives, I'm computing the
  1200. 38:33derivatives on all points in my training
  1201. 38:35set. Now, if you think about that, if I
  1202. 38:38give you the entire internet and ask you
  1203. 38:40to predict the next word in the entire
  1204. 38:42internet, that's going to be pretty
  1205. 38:43slow. If every time you update your
  1206. 38:46model, you have to scan the entire
  1207. 38:47internet, that seems really slow. And
  1208. 38:50so, what machine learning people did is
  1209. 38:51they started to try and get algorithms
  1210. 38:53which were even simpler than this, even
  1211. 38:55dumber than this. You're like, gradient
  1212. 38:56descent, how could you get dumber? Don't
  1213. 38:58worry, machine learning has got you
  1214. 38:59covered. We have a lot more dumb
  1215. 39:00algorithms. How does it work?
  1216. 39:03>> [cough]
  1217. 39:04>> Okay,
  1218. 39:06[clears throat] we'll get there in one
  1219. 39:07second. So, we computed the derivatives.
  1220. 39:09We have this one thing. I don't need to
  1221. 39:11talk about this. All right.
  1222. 39:14Okay. Get to it. Now, one thing I did
  1223. 39:17want to cover before we get there is
  1224. 39:19apologies. This thing here about how do
  1225. 39:21you set the step size? Okay. Now,
  1226. 39:25[clears throat]
  1227. 39:26when you set the step size, think about
  1228. 39:28what it means. The step size is
  1229. 39:30basically how sure you are in that
  1230. 39:32information. In some way, you get the
  1231. 39:34information. If you really trusted that
  1232. 39:35that gradient that it was leading you
  1233. 39:37really steeply down the hill, you would
  1234. 39:39want to take a really big step if you're
  1235. 39:41far away from the optimal, right? And so
  1236. 39:43it turns out for functions that are nice
  1237. 39:44and bullshaped, which are called convex.
  1238. 39:46If you're really far away from the
  1239. 39:47optimal, it's usually very steep. Okay,
  1240. 39:49that's a there's a formal version of
  1241. 39:50that, but that's what's going on. So,
  1242. 39:51you want to take a step. Often in
  1243. 39:54machine learning, we're dealing with
  1244. 39:55really nasty functions. we're dealing
  1245. 39:56with those functions that are really
  1246. 39:57curvy or we don't know what the surface
  1247. 39:59looks like and so we tend to take
  1248. 40:01relatively small steps. If you take
  1249. 40:03steps that are too small, what happens
  1250. 40:06is
  1251. 40:07[snorts] you just don't get to the
  1252. 40:08optimum fast enough. Okay, if you take
  1253. 40:11steps that are too big, this actually
  1254. 40:12looks pretty good to me. It says too
  1255. 40:13large, but like if you just take one
  1256. 40:15step and got to the optimum, like you'd
  1257. 40:16be really happy. So that'd be great if
  1258. 40:18you did that. So this is like too large
  1259. 40:19is kind of like weird. Like it's like
  1260. 40:20bragging like, "Oh, it's too large. It
  1261. 40:22works too well." I don't know. In
  1262. 40:23general, what happens in too large, I'll
  1263. 40:24show you in a minute, is that you bounce
  1264. 40:26around. So, you take a step, but you
  1265. 40:28shoot past the optimal. So, if you
  1266. 40:30picture that bowl that we were talking
  1267. 40:31about before, you would bounce around
  1268. 40:32from both sides. And I'll show you a
  1269. 40:34picture of that. Please question.
  1270. 40:35>> Yeah, I was about to say, what happens?
  1271. 40:37>> Yeah, exactly right. And so, when you
  1272. 40:39see I have those for the SGD. So, you're
  1273. 40:41exactly right. This is what happens when
  1274. 40:42we stock for things like SGD. What
  1275. 40:44happens is you bounce around in a ball,
  1276. 40:46a highdimensional ball that's kind of
  1277. 40:48proportional to that that step size. And
  1278. 40:50this has some profound implications
  1279. 40:52because if you've taken a statistics
  1280. 40:53course, you've been told that what you
  1281. 40:55really care about is recovering the
  1282. 40:57parameters. If you want to be technical
  1283. 40:59about it, you really care about getting
  1284. 41:00the right theta zeros. Okay, this is
  1285. 41:03called a recovery guarantee. If I see
  1286. 41:05the data, then I can prove that I get
  1287. 41:07the right theta zeros. Machine learning
  1288. 41:08people do not care about this.
  1289. 41:10Absolutely not. And they're right not to
  1290. 41:12if you want to do what we were doing,
  1291. 41:13right? Not if you're doing stats.
  1292. 41:14[snorts] And what that means is there
  1293. 41:16could be many many thetas as long as
  1294. 41:19we're close good enough. And so a lot of
  1295. 41:21modern machine learning is like running
  1296. 41:23these models and like when do you stop
  1297. 41:24like in statistics we would do these
  1298. 41:26very you know fancy tests when the
  1299. 41:27gradient goes below this and all the
  1300. 41:28rest we can prove we're within blah blah
  1301. 41:30blah blah blah of the of theta not the
  1302. 41:32theta star the optimal theta star in AI
  1303. 41:35we're like we ran out of compute credits
  1304. 41:37we'll stop feels good let's go to the
  1305. 41:39next model and that's crazy enough how
  1306. 41:42it works. Okay. All right. Okay. So,
  1307. 41:45when you set these at home, I just want
  1308. 41:46you're going to have like some
  1309. 41:47assignments where you play with some
  1310. 41:48step sizes. You'll see if it's
  1311. 41:50converging slowly, try and set better
  1312. 41:52step sizes. In a couple weeks, I'll
  1313. 41:54teach you something called hyperband,
  1314. 41:55which was one of my personal favorite
  1315. 41:56algorithm was written by friends of how
  1316. 41:58to adaptively set these step sizes in a
  1317. 42:00nice way. And uses something called
  1318. 42:01bandits, which we need a lot of
  1319. 42:03notation. And once you once you've
  1320. 42:05gotten familiar with the notation, it's
  1321. 42:06really straightforward. It's Kevin's
  1322. 42:08algorithm. It's cool. All right. All
  1323. 42:10right. So this just says what I said and
  1324. 42:12doesn't have all the weird rants in it.
  1325. 42:13All right, good enough.
  1326. 42:16Okay, now as I mentioned, we want to get
  1327. 42:19even simpler because that n this data
  1328. 42:22set could be very large. And in fact,
  1329. 42:24machine learning is anchored towards
  1330. 42:26more data, not better estimation. Just
  1331. 42:28that's not a formal statement, but just
  1332. 42:30intuitively what we're after is we're
  1333. 42:32after these models that we're going to
  1334. 42:33crank tons and tons and tons of data
  1335. 42:35through them. We think that's a way to
  1336. 42:37improve them by making them bigger and
  1337. 42:39consume more information than
  1338. 42:40necessarily estimating these parameters
  1339. 42:42very finely. And so we'll tell you the
  1340. 42:44techniques that we used to use to
  1341. 42:46estimate those parameters down to like,
  1342. 42:48you know, 10 significant digits. But
  1343. 42:50that's not the vibe. That's not what
  1344. 42:51we're going for. We want to get bigger
  1345. 42:53models that we're going to kind of fit
  1346. 42:55less well than our predecessors. Okay?
  1347. 42:57And that's what leads to all kinds of
  1348. 42:59interesting behaviors that we'll tell
  1349. 43:00you about later like scaling laws and
  1350. 43:01things. Okay? So why does that matter?
  1351. 43:04This is really an algorithm that I owe a
  1352. 43:06lot to and that is actually widely used
  1353. 43:08and it's old algorithm too. So this is
  1354. 43:11the algorithm update rule written for
  1355. 43:13one different component. Okay. We've
  1356. 43:16looked at XJ here. It only depends on
  1357. 43:18the J component. We look at the whole
  1358. 43:20error term. Look at all the data. We
  1359. 43:23update. All right.
  1360. 43:26Okay. Now here I'm just showing that the
  1361. 43:28vector we're writing in vector notation.
  1362. 43:30This is again just why does the vector
  1363. 43:32save us? Well, I just erased all the
  1364. 43:34J's. You will see me very freely go back
  1365. 43:36and forth between these. All right.
  1366. 43:39Okay. This is what I was actually wanted
  1367. 43:40to talk about. Okay. Now, consider our
  1368. 43:43rules. So, I've said this multiple times
  1369. 43:44already. Sorry, I got excited. This is a
  1370. 43:46thing that I like. So, I know it's weird
  1371. 43:47to get excited about this, but I cannot
  1372. 43:49tell you. So, just as a small
  1373. 43:50digression, I worked a lot of my life on
  1374. 43:52the algorithm I'm about to show you.
  1375. 43:54It's extremely simple. Okay, we did all
  1376. 43:56kinds of fun things about this algorithm
  1377. 43:57and how we scaled it up and how we built
  1378. 43:59it. It's called stochcastic mini batch
  1379. 44:01or stochastic gradient descent or in the
  1380. 44:03old days was called incremental gradient
  1381. 44:05and it goes back at least to like the
  1382. 44:061950s Robins and Monroe. It's like a old
  1383. 44:09classical algorithm. People rediscover
  1384. 44:11it every 101 15 years. Okay, it is the
  1385. 44:14workhorse. So how by show of hands, how
  1386. 44:16many people have used PyTorch or
  1387. 44:18anything? Okay, more people use PyTorch.
  1388. 44:20That's dangerous but awesome. uh in case
  1389. 44:22so if you do backrop in there then you
  1390. 44:25have to have used mini batch like the
  1391. 44:26entire system is is set up for you so
  1392. 44:29that you're going to take a small batch
  1393. 44:30of data you're not going to look at
  1394. 44:31everything in your data set and you're
  1395. 44:33going to feed through one you know five
  1396. 44:35images of cats and dogs at the same time
  1397. 44:37okay not all the images of cats and dogs
  1398. 44:39that are you know on your phone or
  1399. 44:40whatever right so that is the mini
  1400. 44:42batching rule so how does that look
  1401. 44:44mathematically
  1402. 44:48well the idea is we're going to sample
  1403. 44:53this thing.
  1404. 44:56Mathematically, the way we'll think
  1405. 44:57about it is we're going to take a random
  1406. 44:59set. Now, honestly, when you run, this
  1407. 45:01is a thing that I again spend a lot of
  1408. 45:02my life, you know, in weird ways. One
  1409. 45:04slice of my like I was doing it all day
  1410. 45:06long, but it was like a piece of things
  1411. 45:07I was working on. I worked on
  1412. 45:09understanding what happens if you
  1413. 45:11randomly sample versus if you just
  1414. 45:13randomly sort your data and go through
  1415. 45:15it once. Okay,
  1416. 45:18turns out they're kind of okay. And that
  1417. 45:19latter one is basically what we do. Took
  1418. 45:21a lot of math to prove that those things
  1419. 45:23are about the same for a variety of
  1420. 45:24reasons. When you get to the end of the
  1421. 45:25data set, it gets nasty. The first part
  1422. 45:27of the data set, they're obviously kind
  1423. 45:28of close. Okay, but this is what SGD is.
  1424. 45:31Sample from your data set.
  1425. 45:33Don't pick those samples in a weird way.
  1426. 45:35Now, what could go wrong if you pick
  1427. 45:37those samples in a weird way? Just
  1428. 45:38intuitively.
  1429. 45:40If for example, let's say that I showed
  1430. 45:42you all the cats first, then I showed
  1431. 45:46you all the dogs
  1432. 45:48and you put them in those little
  1433. 45:50batches. What do you think would happen
  1434. 45:51to the underlying model?
  1435. 45:53>> It would first just keep predicting cats
  1436. 45:55and it would stop predicting cats for
  1437. 45:56everything.
  1438. 45:57>> Exactly right. So, it would learn some
  1439. 45:58trivial surface that was like only
  1440. 45:59predicting the cats. You got it exactly
  1441. 46:01right. And it would get really confident
  1442. 46:03about cats, but it wouldn't know
  1443. 46:04anything about dogs. Then it would see
  1444. 46:05the dogs and it would race to the other
  1445. 46:07side. Okay. So, why do I tell you this
  1446. 46:09intuitively? What do you want in that
  1447. 46:10batch? You want it to be kind of a
  1448. 46:12sample of the population. This is the
  1449. 46:14second statistical assumption we're
  1450. 46:15making. The first one was my training
  1451. 46:17set reflects the real world. My second
  1452. 46:19one is my mini batches again reflect my
  1453. 46:22overall data set.
  1454. 46:24That one as I start to get more
  1455. 46:26complicated things that becomes tricky
  1456. 46:27to guarantee. Please
  1457. 46:30>> do the actual uh form the rule is h of
  1458. 46:33data using the previous.
  1459. 46:37>> Exactly right. Yeah. Wonderful. Yeah. So
  1460. 46:39here this this h of theta should
  1461. 46:41actually be the theta t. Sorry if that's
  1462. 46:42not clear. I'll just write it in here.
  1463. 46:44It's a wonderful observation. So the
  1464. 46:46observation is which theta is talking
  1465. 46:48about here? And it's the theta from the
  1466. 46:49last step. That's just a notational bug.
  1467. 46:51You got it exactly right.
  1468. 46:54Other questions? Oh, please. Did you
  1469. 46:57have a question?
  1470. 46:59>> So is it just these?
  1471. 47:07>> Yeah. So the question is is if you
  1472. 47:09sample with versus without replacement,
  1473. 47:11how does that how does that change the
  1474. 47:12situation? And it turns out that in
  1475. 47:15basically the the moral of the story is
  1476. 47:17we have increasing theory and empirical
  1477. 47:19evidence that it doesn't matter if you
  1478. 47:20do the with replacement versus without
  1479. 47:22replacement, but without replacement is
  1480. 47:24in fact a lot easier to implement
  1481. 47:26because what you can do is you can hash
  1482. 47:27or you can sort on a random key and do
  1483. 47:29it once and then plow through all of
  1484. 47:31your data. In fact, by default, PyTorch
  1485. 47:33will often not sample from the data and
  1486. 47:35just take the batches in the order that
  1487. 47:37you get it. So that once the data set is
  1488. 47:39very very large, the dist the change in
  1489. 47:41this distribution should not be very
  1490. 47:43big. Right? If I have a billion points
  1491. 47:44and I shuffle them versus picking them
  1492. 47:46up. Now, one other thing that people
  1493. 47:48have observed, you don't have to know
  1494. 47:49any of this, so please like don't worry
  1495. 47:51about it, but one other thing that
  1496. 47:52people have observed and this came from
  1497. 47:54a paper that we wrote like 10 years ago
  1498. 47:56uh about this. [snorts]
  1499. 47:58It turns out and and lots of people
  1500. 47:59observe this, not just us. If you do the
  1501. 48:02random reshuffle, the model converges
  1502. 48:05faster. And this is something that you
  1503. 48:07know, if you think about it intuitively,
  1504. 48:09by show of hands, who knows what the
  1505. 48:10coupon collector problem is? Oh, few.
  1506. 48:13All right. PieTorch, but not coupon
  1507. 48:14collector. Interesting times. Anyway,
  1508. 48:16that used to be like CS Cannon. Like,
  1509. 48:18you had to know that. But, you know, I
  1510. 48:20was just This has nothing to do with
  1511. 48:21what I'm talking about. How many people
  1512. 48:22know what a DFA is? An Automa.
  1513. 48:25Oh, that's good. Good for us. All right.
  1514. 48:26We're still doing it. All right. Anyway,
  1515. 48:27back to this. So the point is is when
  1516. 48:29you did that shuffle, you had a kind of
  1517. 48:31a coupon collector phenomenon. What
  1518. 48:32happens in a coupon collector is let's
  1519. 48:34say that I give you 10 coupons that you
  1520. 48:36want to collect and you run around and
  1521. 48:37you're randomly sampling. So you get a
  1522. 48:39random sample every time. It turns out
  1523. 48:41that all of your variance, how long it
  1524. 48:43takes you is how long it takes to get
  1525. 48:44the last coupon. Why? Because if I'm
  1526. 48:47sampling from all 10, right, I'm getting
  1527. 48:49the first nine again and again, 90% odds
  1528. 48:51when I get to the end. And so I only
  1529. 48:53have a 10% odds of getting it when I get
  1530. 48:54to that final one. Turns out that skews
  1531. 48:56all the variance to the end. So back to
  1532. 48:58these models, if you're looking for
  1533. 49:00something that's relatively rare in your
  1534. 49:02data set, you're just not going to
  1535. 49:03encounter it. You're going to encounter
  1536. 49:05the mean a little bit too often and
  1537. 49:07you're not going to encounter the thing
  1538. 49:08you want. So the f folklore is that and
  1539. 49:10there's a bunch of math that says under
  1540. 49:12situations the story I just told you is
  1541. 49:14true, but not all situations. There are
  1542. 49:15obvious counter examples. Okay, awesome.
  1543. 49:18These are wonderful questions. Super
  1544. 49:19happy to talk about this stuff. Please
  1545. 49:22>> batch size. Sometimes we just say this
  1546. 49:25is how much GPU I have. This is what my
  1547. 49:27batch size will be.
  1548. 49:28>> Is there some scientific way to say I
  1549. 49:30should have a batch size of one or two
  1550. 49:32or 12?
  1551. 49:33>> Oh, wonderful question. So the question
  1552. 49:34uh is about batch size and how do you
  1553. 49:36pick it? So I'm going to tell you
  1554. 49:38something. So a couple things. So
  1555. 49:40there's a couple of heruristics that
  1556. 49:41people have but I wanted to tell you as
  1557. 49:43a personal matter batch size was
  1558. 49:45actually one of the things that broke me
  1559. 49:46intellectually uh and a while ago. Okay.
  1560. 49:49So there was a while when I would say
  1561. 49:51about 15 years ago when what we thought
  1562. 49:53was that smaller batch sizes were better
  1563. 49:55because they were exploring the data
  1564. 49:56more and you can tell yourself a story
  1565. 49:57about why this is better and from an
  1566. 49:59optimization perspective you can prove
  1567. 50:01that in many situations a tiny batch
  1568. 50:03size is just as good as a big batch size
  1569. 50:05and so you want to run in these tiny
  1570. 50:07batch sizes. Then there was a paper that
  1571. 50:09came out of when it was Facebook the
  1572. 50:11Facebook labs that basically showed this
  1573. 50:13really interesting thing that was done
  1574. 50:14by friends and they they sent me as a
  1575. 50:15preprint that said we're getting better
  1576. 50:17generalization. I'll explain what I mean
  1577. 50:18in a second by using larger batches.
  1578. 50:21Now, this was very strange. Their loss
  1579. 50:24was worse. Their training set was worse,
  1580. 50:26but they were generalizing better to the
  1581. 50:28real world. Now, as a mathematical or an
  1582. 50:30optimization person, at first I was
  1583. 50:32like, this is heresy. As I just told
  1584. 50:34you, we assumed that the training sets
  1585. 50:36are have the statistical relationship.
  1586. 50:38So, doing better on the training set
  1587. 50:39should always mean doing better on the
  1588. 50:40test set. But in this paper, they showed
  1589. 50:42that wasn't the case all the time. And
  1590. 50:44that's one of the things the reason I
  1591. 50:46tell you this story is first it was like
  1592. 50:47personally like you know as I was
  1593. 50:49working on this it changed the
  1594. 50:50mathematics of it the second piece of it
  1595. 50:52as we went through it was
  1596. 50:56basically the batch size we don't fully
  1597. 50:57understand and so the convention now we
  1598. 51:00have this idea that larger batches are
  1599. 51:01better because they have lower variance
  1600. 51:03they're a better estimate and so the
  1601. 51:05practice that you just mentioned how do
  1602. 51:07you pitch the batch size how big is your
  1603. 51:09you know GPU memory right like you have
  1604. 51:11this much HPM you have this much batch
  1605. 51:13size Like that's basically what people
  1606. 51:15do for systems reasons. But the theory
  1607. 51:17of like how and why this works and the
  1608. 51:19fact that optimization is a leaky
  1609. 51:21abstraction is really interesting. And
  1610. 51:23it should have been more interesting to
  1611. 51:24me. I didn't recognize this when I first
  1612. 51:26saw it, but it was a really important
  1613. 51:27kind of theoretical moment for me
  1614. 51:29because what it re revealed to me is
  1615. 51:31remember when I was telling you machine
  1616. 51:32learning is not stats, right? I think
  1617. 51:34we're still co-listed with stats. No
  1618. 51:35offense to statistitians. I love you
  1619. 51:36very much. Some of my best friends are
  1620. 51:38statistitians. But machine learning is
  1621. 51:40not stats. And one of the things is is
  1622. 51:42statistics as I mentioned is very
  1623. 51:44interested in when the optimization
  1624. 51:45problem recovers the right answer. And
  1625. 51:47as I told you machine learning people
  1626. 51:48are not interested in that. And so the
  1627. 51:50fact that you have a model that doesn't
  1628. 51:51get the right answer with this batch
  1629. 51:53size change you know right answer the
  1630. 51:55lower loss the minimum loss but
  1631. 51:56generalizes better. We're going to throw
  1632. 51:58away all the theory and try and figure
  1633. 51:59out what's happening there. And so now
  1634. 52:01your your heristic is the right one. Set
  1635. 52:03the batch size according to what you can
  1636. 52:05do. And there's a little bit of tuning
  1637. 52:07that people do underneath the covers
  1638. 52:08about how they set their batches. It's a
  1639. 52:10little bit folklore if I'm honest. I can
  1640. 52:12tell you the tricks. You know, I have
  1641. 52:14mine, other people have theirs. I'm
  1642. 52:16[snorts] not super sure. Great question.
  1643. 52:18Please.
  1644. 52:19>> Why would sac potentially be better?
  1645. 52:22Because
  1646. 52:22>> Oh, great question.
  1647. 52:23>> Um cuz in the example given like cats
  1648. 52:26and dogs like wouldn't you end up with
  1649. 52:28if you have smaller batches have much
  1650. 52:30more extreme nonproportional?
  1651. 52:34>> Yeah. Yeah. Wonderful question. Yeah. So
  1652. 52:35the question is why would why did you
  1653. 52:37fools ever think that low small batch
  1654. 52:39sizes were going to work? Uh you it was
  1655. 52:41obvious the whole time you were wrong
  1656. 52:42probably but the reason you would think
  1657. 52:44a small batch size would help is a bit a
  1658. 52:46little bit of a calculation which is
  1659. 52:47actually in the slides in the appendix
  1660. 52:49which shows that imagine the situation
  1661. 52:51where I have a data set that contains a
  1662. 52:53lot of redundancy. Okay, if it contains
  1663. 52:55a lot of redundancy and I look at a
  1664. 52:57whole batch then I'm basically seeing
  1665. 53:00the same example. You know it's just
  1666. 53:01like 90 pictures of the same cat again
  1667. 53:03and again and again, right? And so I've
  1668. 53:05wasted all those steps. I could have
  1669. 53:07been using those 90 cats to refine the
  1670. 53:09model 90 times. So really what it's a
  1671. 53:11trade-off is how statistically
  1672. 53:13meaningful is the the data sample that
  1673. 53:15you're seeing versus how big of a step
  1674. 53:17should you take. So it's not an obvious
  1675. 53:19trade-off either direction. The heristic
  1676. 53:21of why smaller was better was that you
  1677. 53:23were able to polish the model that more
  1678. 53:25parameter updates were better than what
  1679. 53:26you were doing. That's probably true if
  1680. 53:28you have small classes or a relatively,
  1681. 53:31you know, compact domain where you have
  1682. 53:32like, you know, two classes and they're
  1683. 53:34kind of well established and separated.
  1684. 53:36In those situations, this redundancy
  1685. 53:38argument is actually provable. SGD with
  1686. 53:40small batches will do better. And so
  1687. 53:42that like gave us maybe false confidence
  1688. 53:44that that was explaining the whole
  1689. 53:46world. When we move to more interesting
  1690. 53:48examples, larger batch sizes started to
  1691. 53:50work better and better and better. The
  1692. 53:52other thing that happened is the way we
  1693. 53:53treat the batches. As you'll see in the
  1694. 53:55second half of the course, we also want
  1695. 53:57to embed information in them. So when we
  1696. 53:59get to other concerns about how you
  1697. 54:01supervise your data, it's interesting to
  1698. 54:02have a mix of of say positive and
  1699. 54:04negatives. Exactly as you said, I want
  1700. 54:06every batch to have a couple cats and a
  1701. 54:08couple dogs. And ideally, I'd like them
  1702. 54:10to be kind of close together
  1703. 54:11intuitively. I don't want a dog, you
  1704. 54:13know, like a Great Dane and some little
  1705. 54:14tiny cat. I kind of want the most
  1706. 54:16cat-like looking dog I can get to help
  1707. 54:18me separate them. I'll learn the most
  1708. 54:20interesting information. We'll talk
  1709. 54:21about that in that course. you're
  1710. 54:23getting you guys have got this exactly
  1711. 54:24right. What are the key issues and
  1712. 54:26underneath a batch please?
  1713. 54:29>> It was selected specifically for like
  1714. 54:32interesting features that we want.
  1715. 54:34>> Awesome. Yeah, great question. So, in
  1716. 54:36the way that we'll teach the the
  1717. 54:38traditional stocastic gradient descent,
  1718. 54:40it is completely random sampling. With
  1719. 54:42replacement sampling, you take it and
  1720. 54:43you pull it up. What I'm trying to
  1721. 54:45emphasize from some of this discussion
  1722. 54:46is in practice, it's actually not really
  1723. 54:48done that way. Sometimes it's done with
  1724. 54:50this sort that I mentioned where you
  1725. 54:51sort it in a random order. We talked
  1726. 54:52about how to cut down invariance. And
  1727. 54:54then what I was just saying is that some
  1728. 54:55objective functions, some loss functions
  1729. 54:57actually prefer that you have kind of
  1730. 54:59near miss examples to each other which
  1731. 55:01is also quite intuitive. And so people
  1732. 55:03do engineer their batches. And then the
  1733. 55:05last piece is you do pick the batches in
  1734. 55:07modern applications based on how fast
  1735. 55:09they run. as was talking about in the
  1736. 55:10HBM case, if you haven't optimized the
  1737. 55:13GPU, how much if you can store your
  1738. 55:15whole model or your whole computation in
  1739. 55:16the HBM, the memory that's on the chip,
  1740. 55:18it's just dramatically faster. And so
  1741. 55:20you optimize for systems concerns. So
  1742. 55:22this very small tweak here that looks
  1743. 55:25like, you know, a oneline change of
  1744. 55:26going from N to B. Basically,
  1745. 55:29that change is actually in engineering
  1746. 55:32in practice, I'm trying to say is quite
  1747. 55:34rich. For the point of view of the
  1748. 55:35course, I think you have to know nothing
  1749. 55:36of what I just described. By the way,
  1750. 55:37you just need to say, "Oh, yeah, you
  1751. 55:38random sample. That's good enough." But
  1752. 55:40I wanted you to understand like this
  1753. 55:41stuff is actually fairly interesting and
  1754. 55:43fairly accessible. Like there's someone
  1755. 55:45right now at one of the frontier labs
  1756. 55:46who's tuning the batch size. As weird as
  1757. 55:48that is to say, and they're getting paid
  1758. 55:49a lot of money, which is great. I hope
  1759. 55:50they're one of my former students.
  1760. 55:51Anyway, all right. So, this is the this
  1761. 55:55is the definition right here. This is we
  1762. 55:56take the batch and we average it. Okay.
  1763. 55:58And I'm just saying that that rule to
  1764. 56:00select the B random sampling is good
  1765. 56:02enough. You'll do that for most of the
  1766. 56:04course. But it is actually like making a
  1767. 56:07statistical assumption. And if you
  1768. 56:08really care about what's going on there,
  1769. 56:10it's just an interesting thing. And
  1770. 56:12there's there's there's stuff to know.
  1771. 56:13If you're a curious person, there's
  1772. 56:14stuff to know. Wonderful questions.
  1773. 56:17These are really great.
  1774. 56:20All right.
  1775. 56:24All right. So hopefully this is clear. I
  1776. 56:26said this I should have I should have
  1777. 56:27gone to this slide earlier. Sorry. Um
  1778. 56:30here I would point out one thing which
  1779. 56:32is you're going to sum over the batch
  1780. 56:33that mini batch right so B is usually
  1781. 56:36less than the full data set you compute
  1782. 56:38the errors on each one with your current
  1783. 56:40theta which was pointed out this should
  1784. 56:41be theta t you then just do the updates
  1785. 56:44the point is you don't have to look at
  1786. 56:45the rest of the data set you just look
  1787. 56:47at the model and the data set the your
  1788. 56:49data samples from the batch and then
  1789. 56:51you're going to take this step size and
  1790. 56:52I'm highlighting here that it's an alpha
  1791. 56:54B and that's because it's a it's
  1792. 56:56potentially a different step size okay
  1793. 56:58it's you're not going to use the game
  1794. 56:59step size that you would use on the full
  1795. 57:01data set. And in fact, the step size
  1796. 57:03kind of mysteriously depends on other
  1797. 57:05quantities that are there. We'll talk
  1798. 57:06about good rules of thumb. In fact, one
  1799. 57:08of the major pieces of technology that
  1800. 57:10you probably use and so many of you have
  1801. 57:12used step size are basically what are
  1802. 57:14called adaptive optimizers that pick the
  1803. 57:16step size for you within range. They see
  1804. 57:18that you're having some of these
  1805. 57:19problems. Adam, for example, if you know
  1806. 57:22what Adam is or Adam W. Okay, so or or
  1807. 57:25Adagrad, which was developed by our own
  1808. 57:27John Duchi. Um those are those are the
  1809. 57:29kinds of of things that go on under the
  1810. 57:31covers. But for now when we're studying
  1811. 57:32it in its purest setting like if you
  1812. 57:34wanted to do it with like code and write
  1813. 57:36it out you would have to pick that
  1814. 57:37alphab and that's what goes on. Any
  1815. 57:40other questions there?
  1816. 57:43Oh please
  1817. 57:46change like a batch.
  1818. 57:48>> Oh yeah sorry that this wasn't clear.
  1819. 57:50Yeah so what the the operation that
  1820. 57:51you're doing here sorry this is unclear
  1821. 57:53is you're going to pick a random batch
  1822. 57:55on each iteration. So at every time step
  1823. 57:57t you pick a new random batch. Okay. You
  1824. 58:00don't want to feed the same dogs and
  1825. 58:01cats through every single time or houses
  1826. 58:03through every time because you'll just
  1827. 58:04learn a model of the data that you've
  1828. 58:06seen. So it is a random sampling
  1829. 58:07procedure. Yeah. Wonderful question and
  1830. 58:09clarification.
  1831. 58:10Please
  1832. 58:13the coefficient gets absorbed into the
  1833. 58:15step size. Does that imply that larger
  1834. 58:17batches typically would have smaller
  1835. 58:19step sizes?
  1836. 58:20>> Yeah, that's a wonderful question. So
  1837. 58:21how the batch size scales, you would
  1838. 58:23indeed expect that as you ramp up the
  1839. 58:25batch size, right? When you go all the
  1840. 58:27way from one to the limit, you're going
  1841. 58:29to get to a batch size of one overn. You
  1842. 58:31very rarely use 1 overn as your SGD
  1843. 58:33batch size. It turns out for like
  1844. 58:36folklore reasons, that lots of people
  1845. 58:38write their models, not linear models,
  1846. 58:39but write their models so that there's
  1847. 58:41the same kind of normalization in batch
  1848. 58:43size across all of them. And so that is
  1849. 58:45like a thing that you do. So if you've
  1850. 58:47if you've played with deep learning
  1851. 58:48things and used like layer norms or
  1852. 58:49sandwich norms or other things, they're
  1853. 58:51trying to get you in the right regime
  1854. 58:52where like the step sizes are kind of
  1855. 58:54all what's called fancy word is
  1856. 58:55isotropic, but kind of all the
  1857. 58:57dimensions are the same. So what does
  1858. 58:58that mean for alpha? That means you
  1859. 59:00would intuitively expect that if the if
  1860. 59:03the model is taking a little bit of
  1861. 59:05information, you don't trust it very
  1862. 59:07much and alpha would be lower almost
  1863. 59:08exactly as you said.
  1864. 59:10>> Yeah. Oh, go ahead. Um,
  1865. 59:13do you have you done experiments where
  1866. 59:15you've changed the BM size throughout
  1867. 59:17the gradient descent? Is it
  1868. 59:19>> Oh, yeah. Wonderful question. So, we
  1869. 59:21haven't talked about this. Again, you
  1870. 59:22don't need to know this, but I'm super
  1871. 59:23happy to tell you. So, it turns out
  1872. 59:25there's actually a wonderful study
  1873. 59:26written by this guy at the Navy named
  1874. 59:27Leslie who wrote this study about how do
  1875. 59:29you do various what are called
  1876. 59:30scheduling rules or step size rules. So,
  1877. 59:32there's two great papers if you ever
  1878. 59:33wonder. There's one where he basically
  1879. 59:35did what's actually quite widely used.
  1880. 59:36It's called cosign scaling where you
  1881. 59:38actually make the batch size go up and
  1882. 59:39down as you run. The idea being that you
  1883. 59:41zoom into like a local minima and then
  1884. 59:43it kicks you out. And he showed that
  1885. 59:44this cosign scaling was actually one of
  1886. 59:46the best at the time for image models.
  1887. 59:49There's another paper which I love which
  1888. 59:50is like from the '9s which is the
  1889. 59:52Onstriker paper which shows that a bunch
  1890. 59:54of rules that people are using in
  1891. 59:55gradient descent and stochastic gradient
  1892. 59:57descent are actually all the same
  1893. 59:58so-called linear and exponential back
  1894. 1:00:00off and all the rest. [snorts] Now, I
  1895. 1:00:02would doubt that most people here in
  1896. 1:00:04production are tuning their schedules uh
  1897. 1:00:07unless they're doing something that like
  1898. 1:00:09you well for for a product you would I
  1899. 1:00:11guess you would say, but like
  1900. 1:00:12researchers are kind of just putting in
  1901. 1:00:14the numbers and letting the defaults
  1902. 1:00:15work most of the time. Yeah, wonderful
  1903. 1:00:18questions. Please
  1904. 1:00:20>> um
  1905. 1:00:22like
  1906. 1:00:23you would not want to like have repeated
  1907. 1:00:26sampling of uh things. So it's like in
  1908. 1:00:29like the selection of batches would you
  1909. 1:00:31not like?
  1910. 1:00:33>> Awesome. So the thing is is if you look
  1911. 1:00:35at this a wonderful question. So the
  1912. 1:00:37question is if you sample your data you
  1913. 1:00:40kind of want to see all of your data.
  1914. 1:00:42You don't want to have repetition in
  1915. 1:00:43there. And this is exactly gets back to
  1916. 1:00:45the difference between with replacement
  1917. 1:00:46sampling where you would see as I was
  1918. 1:00:48talking about those coupons many many
  1919. 1:00:50times versus without replacement
  1920. 1:00:52sampling where you shuffle the data set.
  1921. 1:00:54And when you shuffle that data set,
  1922. 1:00:55you're guaranteed that you're only going
  1923. 1:00:56to see every data set exactly once. And
  1924. 1:00:58so the belief is that that would help
  1925. 1:01:00you exactly for the reason you said that
  1926. 1:01:02you're going to see many of the examples
  1927. 1:01:04again and again. Okay? Now, that assumes
  1928. 1:01:07that all the examples are equally
  1929. 1:01:08informative, and that's not always true.
  1930. 1:01:10Sometimes there's a core hardcore set
  1931. 1:01:13that you wish you could see again and
  1932. 1:01:15again. And if you look at modern
  1933. 1:01:16training for like LLMs that are in the
  1934. 1:01:18wild, large language models that are in
  1935. 1:01:20the wild, you'll see that actually there
  1936. 1:01:21are places where people actually do
  1937. 1:01:23repetition on things like code is very
  1938. 1:01:25popular to do multiple times because of
  1939. 1:01:27the belief is you put in some structure.
  1940. 1:01:28Okay, there's a great paper about this
  1941. 1:01:30if you're interested about data mixing.
  1942. 1:01:32Please send a note or post to Ed. I'm
  1943. 1:01:33happy to post something about about what
  1944. 1:01:35people know about how you mix your data
  1945. 1:01:37for for foundation models. Not part of
  1946. 1:01:39the course, but super happy to tell you.
  1947. 1:01:41Yeah. Uh you said the cosign uh schedule
  1948. 1:01:44paper they did it with image models. Is
  1949. 1:01:46there a reason why like uh learning rate
  1950. 1:01:49things are different from image models
  1951. 1:01:51and language
  1952. 1:01:51>> models?
  1953. 1:01:53>> Yeah. So the question is why does the
  1954. 1:01:55why would you care that it's from image
  1955. 1:01:56models versus text models? What would
  1956. 1:01:58change there? There's two things that
  1957. 1:02:00change. One are the architectures. At
  1958. 1:02:02the time the architectures were
  1959. 1:02:03different. We're talking about neural
  1960. 1:02:04net architectures. That has actually
  1961. 1:02:06changed over time. Now we've converged
  1962. 1:02:08and built a bunch of technology where
  1963. 1:02:09you can use kind of transformer stacks
  1964. 1:02:11for both of them. The other thing is
  1965. 1:02:12that the distributions may actually be
  1966. 1:02:14quite different. So if you think about
  1967. 1:02:15images, they have a set of natural
  1968. 1:02:17distributions that are in the world that
  1969. 1:02:19like has some smooth variations, right?
  1970. 1:02:21Like I don't know about you, but like
  1971. 1:02:22you know I take pictures of my kids.
  1972. 1:02:24It's like in burst mode. I got a
  1973. 1:02:25thousand pictures of, you know, like one
  1974. 1:02:26of my daughters dancing, right? That
  1975. 1:02:29thing they're all little tiny variations
  1976. 1:02:30of each other. That distribution is kind
  1977. 1:02:32of compact. I don't tend to write the
  1978. 1:02:34same sentence a thousand times like a
  1979. 1:02:36crazy person. maybe I I have over like
  1980. 1:02:37you know longitudinally in my life but
  1981. 1:02:39like it's not kind of compact in that
  1982. 1:02:41way and so those differences in
  1983. 1:02:43distributions they may have an effect on
  1984. 1:02:45the learning rates and so you can't they
  1985. 1:02:48don't transfer well like if you had
  1986. 1:02:49clean theory you would know what goes
  1987. 1:02:51from A to B but we don't know what goes
  1988. 1:02:53from A to B in different situations so
  1989. 1:02:54that caveat turned out to be important
  1990. 1:02:56it caught on quite a bit in images and
  1991. 1:02:58some of those ideas were adapted but
  1992. 1:03:00those aren't the ones that we use today
  1993. 1:03:02in uh engineering text models
  1994. 1:03:06the production model at this point.
  1995. 1:03:08Would you like how would you feel if the
  1996. 1:03:11bad side changing like right now?
  1997. 1:03:15>> I would pick them. Yeah. So, I would
  1998. 1:03:17pick the way I would do this. Oh, sorry.
  1999. 1:03:18The way I would do this is oops. Yeah.
  2000. 1:03:21So, this this covers many of our points.
  2001. 1:03:23So, the way I would do this honestly is
  2002. 1:03:25I would try and figure out what makes
  2003. 1:03:26the GPU most efficient. So if you look
  2004. 1:03:28at the cost of these models like one
  2005. 1:03:30thing that changed since say like I
  2006. 1:03:31first started teaching this course till
  2007. 1:03:33now is the amount of what you would say
  2008. 1:03:34is capex the amount of capital
  2009. 1:03:36expenditure to build these models we're
  2010. 1:03:38talking about building gigawatt
  2011. 1:03:39facilities this was unimaginable
  2012. 1:03:41gigawatt is a lot of compute okay those
  2013. 1:03:44gig multi- gigawatt facilities you have
  2014. 1:03:46to keep them at high utilization so if
  2015. 1:03:48you ask me what I care about I would say
  2016. 1:03:50I care quite a bit about making sure the
  2017. 1:03:52utilization of the model is high that I
  2018. 1:03:54get in more steps the statistical
  2019. 1:03:56concerns that we'll talk out they don't
  2020. 1:03:58take a complete back seat but if you're
  2021. 1:04:00doing it like if you see how people
  2022. 1:04:02report their training they report MFU
  2023. 1:04:04model flops you know utilization they
  2024. 1:04:06care about how much they're getting out
  2025. 1:04:07of those GPUs the statistical properties
  2026. 1:04:09are sometimes harder to pin down and
  2027. 1:04:11that's one thing by the way that is
  2028. 1:04:12actually quite a blessing like one thing
  2029. 1:04:14that I think is very underappreciated
  2030. 1:04:16about machine learning you can get this
  2031. 1:04:18so we're going to teach you these
  2032. 1:04:19different building blocks but sometimes
  2033. 1:04:22what happens in this field is that
  2034. 1:04:23people think those building blocks in
  2035. 1:04:24their head occupy like entirely
  2036. 1:04:26different spaces
  2037. 1:04:27One of the beautiful things about
  2038. 1:04:28machine learning which is kind of
  2039. 1:04:30amazing is many of these things work
  2040. 1:04:33like one experiment like I'm going to
  2041. 1:04:34tell you in the next lecture how you
  2042. 1:04:35should use a different loss function for
  2043. 1:04:37classification but it kind of works if
  2044. 1:04:40you use the loss function from this even
  2045. 1:04:42though it kind of makes no sense and
  2046. 1:04:43I'll show you why it shouldn't make much
  2047. 1:04:45sense but it will still kind of work
  2048. 1:04:46there's a robustness to the underlying
  2049. 1:04:48elements that is pretty surprising so
  2050. 1:04:50it's very easy to fixate on the
  2051. 1:04:52statistical definitions of what's going
  2052. 1:04:53on and we will tell you the most
  2053. 1:04:55important but I want to highlight that
  2054. 1:04:57sometimes tweaking the batch size or
  2055. 1:04:59tweaking these rates doesn't matter
  2056. 1:05:00nearly as much as making your model get
  2057. 1:05:02bigger or you know running it for longer
  2058. 1:05:04and so those principal components are
  2059. 1:05:06the ones that have driven the progress
  2060. 1:05:08over the last couple of years that may
  2061. 1:05:09saturate at some point please
  2062. 1:05:12>> you like divided by like the one over
  2063. 1:05:15like
  2064. 1:05:17>> right yeah
  2065. 1:05:18>> yeah that's just packed into the step
  2066. 1:05:20size
  2067. 1:05:21>> yeah yeah we just normalize it into the
  2068. 1:05:22alpha cuz it's like it's a number that
  2069. 1:05:24we don't know haven't interpreted anyway
  2070. 1:05:25so we might as well just multiply it by
  2071. 1:05:26something.
  2072. 1:05:27>> Okay.
  2073. 1:05:27>> Yeah. Yeah. Awesome.
  2074. 1:05:30>> Great questions.
  2075. 1:05:33>> Okay. All right. So, I I promised this
  2076. 1:05:35graph earlier. It's not that great. If
  2077. 1:05:37you were really like on the edge of your
  2078. 1:05:38seat waiting for it, apologies. Uh but
  2079. 1:05:40this is how it works. So, these are the
  2080. 1:05:42batch versus uh gradient descent on a
  2081. 1:05:44very smooth problem. Okay. So, a smooth
  2082. 1:05:47problem, what is it? How do I know this
  2083. 1:05:48problem is smooth by looking at it?
  2084. 1:05:50Well, these things here are the
  2085. 1:05:52isocience for the function. That's an
  2086. 1:05:54equal loss. Okay. Okay, that's where the
  2087. 1:05:55loss is all one value. This is a
  2088. 1:05:57quadratic that we've been looking at.
  2089. 1:05:59Quadratics, if you remember, produce
  2090. 1:06:01ellipses as isoclines. Is that familiar?
  2091. 1:06:04Show of hands familiar? Okay, I'll
  2092. 1:06:06remember next time. All right. So,
  2093. 1:06:08that's what they look like. We'll come
  2094. 1:06:09back to to that in more more depth. Not
  2095. 1:06:11critical now. So, this function is very
  2096. 1:06:13smooth and bullshaped, right? Because
  2097. 1:06:14this this says like all of the losses
  2098. 1:06:17here and all the losses here are the
  2099. 1:06:18same. These guys are all each rung is
  2100. 1:06:20kind of the same. So, it's this nice
  2101. 1:06:21bullshaped function. Gradient descent
  2102. 1:06:24goes down and comes right to the optimal
  2103. 1:06:26value which we've put right here in the
  2104. 1:06:27middle. SGD as we talked about makes
  2105. 1:06:31these little wiggles. Sometimes it's
  2106. 1:06:32going the right direction, sometimes
  2107. 1:06:33going the wrong direction. It bounces
  2108. 1:06:35around and if alpha is set too large, it
  2109. 1:06:38will bounce around in a ball. Okay? And
  2110. 1:06:40it won't get close to the optimal. So
  2111. 1:06:42one trick that people do, by the way, is
  2112. 1:06:44they what do something called averaging
  2113. 1:06:45the iterates. They may average over a
  2114. 1:06:47trajectory of those bounces to kind of
  2115. 1:06:49simulate a larger batch size. But
  2116. 1:06:51hopefully intuitively this makes sense.
  2117. 1:06:53The reason I highlight this for you is
  2118. 1:06:55when you're training a model, you will
  2119. 1:06:56observe this behavior. You will see the
  2120. 1:06:58loss start to go and bounce around and
  2121. 1:07:00you'll realize maybe I said alpha too
  2122. 1:07:01high. Right? So those are the kinds of
  2123. 1:07:03things that you'll see. Now in machine
  2124. 1:07:06learning in the second half of the
  2125. 1:07:07course, the functions will not be so
  2126. 1:07:09nice. They will not be these nice
  2127. 1:07:11bullshaped functions, these convex
  2128. 1:07:13functions. They're going to be nasty.
  2129. 1:07:14And when they're nasty, then the fact
  2130. 1:07:16this is relying on the fact that it's
  2131. 1:07:18quite smooth to go fast, to get rapidly
  2132. 1:07:20down to the optimal. This is not relying
  2133. 1:07:22on this. This is just kind of drunkenly
  2134. 1:07:24stumbling its way to the loss function.
  2135. 1:07:27And this one we prefer,
  2136. 1:07:29as weird as it is, but for all the
  2137. 1:07:31reasons I outlined.
  2138. 1:07:33Okay, cool.
  2139. 1:07:36All right.
  2140. 1:07:38[clears throat] Okay. So just to make
  2141. 1:07:40sure that we I got across what I wanted
  2142. 1:07:42to get across. Our goal was to optimize
  2143. 1:07:43a loss function to find a good
  2144. 1:07:45predictor. We did that by minimizing
  2145. 1:07:47this loss. We talked a lot I talked a
  2146. 1:07:49lot about what it means to kind of that
  2147. 1:07:51assume about your data that it's
  2148. 1:07:52representative of the world. We'll talk
  2149. 1:07:54about that even more later. We learned
  2150. 1:07:56about an algorithm which hopefully
  2151. 1:07:58looked relatively simple to you. You
  2152. 1:07:59guys all the folks here knew about
  2153. 1:08:01derivatives and computing gradients and
  2154. 1:08:03partial derivatives. Awesome. And we did
  2155. 1:08:05this kind of weird simple thing where we
  2156. 1:08:07started looking at not our entire data
  2157. 1:08:08but a batch of data, a subset of data.
  2158. 1:08:10And when we did that, that somehow
  2159. 1:08:12unlocked a bunch of runtime performance
  2160. 1:08:14and I claim like these really large
  2161. 1:08:16models. Okay, that algorithm was called
  2162. 1:08:19stochastic gradient descent. We talked
  2163. 1:08:20about all the ways to select batches.
  2164. 1:08:22The most important one is select them at
  2165. 1:08:24random. Okay, we touched a little bit on
  2166. 1:08:27these trade-offs of using the right
  2167. 1:08:28batch size through our conversations
  2168. 1:08:30about, you know, when it's too small and
  2169. 1:08:31when it's too large. what are the
  2170. 1:08:33considerations inside a batch? And
  2171. 1:08:34that's to hopefully give you a feel when
  2172. 1:08:36you start to actually play with and tune
  2173. 1:08:37some of these models. That's different,
  2174. 1:08:39by the way, than what you would do if
  2175. 1:08:40you were training a traditional stats
  2176. 1:08:42model where you were trying to get down
  2177. 1:08:43to the optimal value. There's just
  2178. 1:08:46different concerns and hopefully some of
  2179. 1:08:47those concerns have been highlighted and
  2180. 1:08:48will become more clear over the next
  2181. 1:08:50couple lectures.
  2182. 1:08:52Any questions on this before I move on?
  2183. 1:08:56>> Please.
  2184. 1:08:58How do you know when to stop or is it
  2185. 1:08:59just like when you stop?
  2186. 1:09:01>> Awesome. Great question. So the question
  2187. 1:09:02is, as you're bouncing around near the
  2188. 1:09:04optimum, how do you know where to stop?
  2189. 1:09:05And unfortunately, you don't. You don't
  2190. 1:09:07know if you're bouncing around like
  2191. 1:09:08these models have these very weird
  2192. 1:09:10behaviors. So one of the things that
  2193. 1:09:12they have as a behavior sometimes, and
  2194. 1:09:13if you watch and measure them, they're
  2195. 1:09:15all of a sudden they're running, they're
  2196. 1:09:16running, the loss is kind of doing
  2197. 1:09:17something and all of a sudden a
  2198. 1:09:18capability turns on. Okay? So it just
  2199. 1:09:21gets a little bit better. The
  2200. 1:09:22representation clicks in when you're
  2201. 1:09:23training more sophisticated models. So
  2202. 1:09:25you don't know if you're stuck in kind
  2203. 1:09:26of one of those local areas or if you're
  2204. 1:09:29going to get to something, you know, if
  2205. 1:09:30you're really close to the answer. And
  2206. 1:09:32so that is actually a very
  2207. 1:09:33computationally difficult problem. For
  2208. 1:09:36everything in the next couple of
  2209. 1:09:37sections, it's going to turn out that
  2210. 1:09:38the models, if you're familiar, are
  2211. 1:09:40going to be convex underneath the
  2212. 1:09:41covers. They're going to be generalized
  2213. 1:09:42linear models. That's what we're going
  2214. 1:09:43to talk about for two weeks. In those
  2215. 1:09:45situations, we can actually tell how far
  2216. 1:09:46we are from the optimum. I'm not going
  2217. 1:09:48to beat you in the head about this too
  2218. 1:09:50much because you should just take
  2219. 1:09:51Steven's course. It's wonderful. I love
  2220. 1:09:52Steven Boyd. [snorts]
  2221. 1:09:54But also those won't really do us much
  2222. 1:09:56good when we get to the wild world of
  2223. 1:09:58you know neural nets and unsupervised
  2224. 1:10:00and all the rest. And the truth is there
  2225. 1:10:01we don't know. And so how do you run?
  2226. 1:10:04You kind of run like if I watch how I do
  2227. 1:10:05it or my students do it. You look at
  2228. 1:10:07like weights and biases. You watch till
  2229. 1:10:08the thing goes down. You're like I don't
  2230. 1:10:09know if I train for another two days I
  2231. 1:10:11can't do you know I can't go out with my
  2232. 1:10:12friends. I'm going to cut it off now.
  2233. 1:10:14Okay. And that is unfortunately how it
  2234. 1:10:16works. And so you basically provision
  2235. 1:10:18like we're doing a pretty big train
  2236. 1:10:20right now. the way we're provisioning it
  2237. 1:10:21is like well we have these GPUs for two
  2238. 1:10:24months for this project that are donated
  2239. 1:10:26some DNA model thing whatever we're
  2240. 1:10:28going to train that right now and so it
  2241. 1:10:30kind of just is an engineering
  2242. 1:10:32consideration more than it's a
  2243. 1:10:33principled one uh I'm not really
  2244. 1:10:35embarrassed to say because the stuff
  2245. 1:10:36we've been building is great I used to
  2246. 1:10:37be embarrassed but nothing worked as my
  2247. 1:10:39wife said you know we've been together
  2248. 1:10:40forever she's like you know you've been
  2249. 1:10:42building the same demo since I've known
  2250. 1:10:43you since undergrad but now stuff works
  2251. 1:10:45so it's okay so all right
  2252. 1:10:49anyway so Yeah. So these squares are
  2253. 1:10:52really special. Um I want to do this
  2254. 1:10:54basically just for notation. We have a
  2255. 1:10:55couple minutes left. So I want to just
  2256. 1:10:57highlight this notation um about normal
  2257. 1:11:00equations. Uh if you haven't seen this
  2258. 1:11:02before um take a look, do a review. All
  2259. 1:11:05right. So we're going to derive the
  2260. 1:11:06normal equations.
  2261. 1:11:08All right. So mainly this is to make
  2262. 1:11:10sure you're comfortable with the
  2263. 1:11:10notation if I'm honest, but also to make
  2264. 1:11:14sure that you know kind of basic things
  2265. 1:11:15about linear algebra. So if you don't
  2266. 1:11:17know what rank is, you don't know what
  2267. 1:11:18inverses are,
  2268. 1:11:20then you should study up because we use
  2269. 1:11:22them pretty freely. We use that
  2270. 1:11:23terminology pretty freely. Okay. All
  2271. 1:11:25right. So we have some loss function
  2272. 1:11:28here. Uh oops. We have some loss
  2273. 1:11:30function J. Here we have again our
  2274. 1:11:32vector notation. This is actually now
  2275. 1:11:34going to be these ns are going to be
  2276. 1:11:36laid out by rows. So it's going to be n
  2277. 1:11:38by d matrix. So it's going to have n
  2278. 1:11:40rows where the examples are there. And
  2279. 1:11:42then d + one are the feature size that
  2280. 1:11:45we talked about. This is examples. This
  2281. 1:11:46each one of our examples. And then we're
  2282. 1:11:47going to have correspondingly an element
  2283. 1:11:49of Rn, which is going to be this vector
  2284. 1:11:51here, Y. You may hear me call X the
  2285. 1:11:53design matrix. This is really old
  2286. 1:11:55terminology, but it's one that we use.
  2287. 1:11:57Um, and it just means the matrix of X
  2288. 1:11:59after we've done the featurization. So,
  2289. 1:12:01after it's been kind of loaded and
  2290. 1:12:02pre-processed into this form. Sometimes
  2291. 1:12:04we'll worry about properties of the
  2292. 1:12:06design matrix. So, that's what that
  2293. 1:12:08means. Okay.
  2294. 1:12:12[snorts]
  2295. 1:12:13All right. Now, if it's linear, see if
  2296. 1:12:16you can get here.
  2297. 1:12:18[snorts] Notice that J has this
  2298. 1:12:19particularly lovely form.
  2299. 1:12:22Okay. Now, if this is not obvious to you
  2300. 1:12:24that it has it, this is a good signal
  2301. 1:12:26that you should do some work on kind of
  2302. 1:12:28your vector manipulation, matrix
  2303. 1:12:31manipulation stuff. This is the inner
  2304. 1:12:33product. Okay, that's all I've written
  2305. 1:12:34here. So, L2 is important because it's
  2306. 1:12:36the inner product. That's the other
  2307. 1:12:38succinct way to say this. And so, this
  2308. 1:12:40is computing squares. And if you are
  2309. 1:12:42rusty about this, please just compute it
  2310. 1:12:44by hand. It's not like I'm it's not some
  2311. 1:12:45deep mysterious thing. I'm just saying
  2312. 1:12:48this thing here, oops, equals this thing
  2313. 1:12:51here.
  2314. 1:12:53All right. Okay, cool. Right. So now
  2315. 1:12:57we've stacked the design.
  2316. 1:12:59Now, if you don't remember this, please
  2317. 1:13:02again review. But basically when we
  2318. 1:13:04compute a gradient of a matrix, it means
  2319. 1:13:07that we're computing derivatives
  2320. 1:13:09according to a of each underlying
  2321. 1:13:11element, right? And just stacking them
  2322. 1:13:13in the corresponding type. So if you had
  2323. 1:13:14a matrix coming in, you're going to have
  2324. 1:13:16a matrix coming out. They're going to
  2325. 1:13:17have all the same types. They're going
  2326. 1:13:18to have the partial derivatives
  2327. 1:13:20computed. This is the object.
  2328. 1:13:24These are derivatives in the sense
  2329. 1:13:25you're used to them in that if you want
  2330. 1:13:28to find a minimum and it's a convex
  2331. 1:13:29function and all those nice things, you
  2332. 1:13:31set the gradient equal to zero. Why does
  2333. 1:13:33this work intuitively? Well, intuitively
  2334. 1:13:35what's happening is the function is
  2335. 1:13:37going to be bullshaped. It's going to be
  2336. 1:13:38strictly bullshaped. If there's a
  2337. 1:13:40change, the function changes in one
  2338. 1:13:42direction. You're not at the bottom.
  2339. 1:13:43When you get to the bottom, you're done.
  2340. 1:13:46Okay? This means there's no change in a
  2341. 1:13:47local direction. Just like if you
  2342. 1:13:49computed the derivative for a single 1D
  2343. 1:13:51function. Okay.
  2344. 1:13:54[clears throat]
  2345. 1:13:55All right.
  2346. 1:13:57So far so good. All right.
  2347. 1:14:01I got to get better at that. Anyway,
  2348. 1:14:04so from our previous derivation, this is
  2349. 1:14:06just some arithmetic. We're going to
  2350. 1:14:08multiply this thing out and normalize.
  2351. 1:14:12The two does us a little bit of of
  2352. 1:14:14value. We multiply out, we get the x
  2353. 1:14:17theta, the xty,
  2354. 1:14:19and we get this derivative. Okay. If
  2355. 1:14:22that's unfamiliar to you, what's going
  2356. 1:14:23on? We're multiplying these terms out.
  2357. 1:14:25We get an xtx theta on each side. When
  2358. 1:14:28we take the derivative with respect to
  2359. 1:14:30theta, we get the term on one side and
  2360. 1:14:31the other, but it's symmetric.
  2361. 1:14:34We have the 1/2, we get this term. Okay?
  2362. 1:14:37So, it's the this term and the foil.
  2363. 1:14:39When we multiply the y's, we're not
  2364. 1:14:41taking the derivative. They don't depend
  2365. 1:14:42on theta. They get knocked out. When we
  2366. 1:14:45do the cross terms, we get the cross
  2367. 1:14:46term represented
  2368. 1:14:48like so twice.
  2369. 1:14:51And that's the derivative.
  2370. 1:14:53We set this equal to zero. So that means
  2371. 1:14:55that this thing is equal to zero. So how
  2372. 1:14:57do we solve for that? We make these two
  2373. 1:14:59terms equal to each other. Well, we have
  2374. 1:15:01to take the inverse to get rid of the
  2375. 1:15:03side xtx inverse* xty. And this is the
  2376. 1:15:07le square solution.
  2377. 1:15:11This is an optimal solution for for
  2378. 1:15:13theta. Now I cheated in one part of this
  2379. 1:15:15derivation.
  2380. 1:15:19This is something I assumed about X when
  2381. 1:15:20I did this.
  2382. 1:15:22>> Assumed that Xtx is invertible.
  2383. 1:15:24>> Exactly. Right. I assume that XTX is
  2384. 1:15:25invertible. If for example I had
  2385. 1:15:27relatively few examples in a huge set of
  2386. 1:15:30of dimensions, this would be a
  2387. 1:15:32nonsensical kind of thing.
  2388. 1:15:36So here when I'm looking at it, I think
  2389. 1:15:37that I have many many data points more
  2390. 1:15:39than my parameters.
  2391. 1:15:41Okay. And hopefully that that inverse
  2392. 1:15:45exists.
  2393. 1:15:46But what happens that if it doesn't
  2394. 1:15:48exist,
  2395. 1:15:50what is it defined up to the theta that
  2396. 1:15:53I get?
  2397. 1:15:58So yeah, so if this is unfamiliar to
  2398. 1:16:00you, take a look. So there will be a
  2399. 1:16:01null space, right? And in that null
  2400. 1:16:03space, I won't prefer one solution over
  2401. 1:16:04another. All of them would be valuable.
  2402. 1:16:07So anything in the null space of xtx,
  2403. 1:16:09right? I can add that to theta and it
  2404. 1:16:10doesn't change the answer at all. So now
  2405. 1:16:12there goes from being one single
  2406. 1:16:14explanation, one single theta theta star
  2407. 1:16:16to being an entire subspace of them.
  2408. 1:16:18Effectively the entire null space of
  2409. 1:16:20xdx. If you're seeing this and going,
  2410. 1:16:22"Oh yeah, I remember that linear algebra
  2411. 1:16:23sounds great. Do a little review." If
  2412. 1:16:25you've never heard that and it sounds
  2413. 1:16:26like I'm speaking a foreign language,
  2414. 1:16:28please, please, please use the Friday
  2415. 1:16:30lectures because after this I'm going to
  2416. 1:16:31assume we know all this stuff. Okay,
  2417. 1:16:33that's basically the thing there.
  2418. 1:16:35[snorts] It will turn out, by the way,
  2419. 1:16:37you could also worry that XTX here is
  2420. 1:16:40also going to be it turns out positive
  2421. 1:16:41semidefinite by show of hands. Who knows
  2422. 1:16:43what PSD means? Positive semi-definite.
  2423. 1:16:44Awesome. We're doing great. Okay,
  2424. 1:16:46fantastic. Okay, so if this isn't
  2425. 1:16:48familiar, I even wrote it down. Practice
  2426. 1:16:50on Friday. And I just want to make sure
  2427. 1:16:52that's clear. We're going to use these
  2428. 1:16:53kind these concepts positive
  2429. 1:16:54semi-definite. We're going to use these
  2430. 1:16:56concepts like it's invertible. We're
  2431. 1:16:58going to use them pretty freely. We're
  2432. 1:16:59going to compute vector derivatives and
  2433. 1:17:01matrix derivatives in this way. If it's
  2434. 1:17:04unfamiliar, spend a little bit of time
  2435. 1:17:06to make it familiar. It will make your
  2436. 1:17:07life a little bit more easy. You know,
  2437. 1:17:09it's it's much easier to compute this
  2438. 1:17:11way, I think, than the other way, just
  2439. 1:17:12by hand.
  2440. 1:17:15All right,
  2441. 1:17:17almost pretty good on time. Not so bad.
  2442. 1:17:18Usually, I'm pretty bad about this. All
  2443. 1:17:20right, so we saw lots of notation today.
  2444. 1:17:22That was the intent of the lecture in
  2445. 1:17:24some ways. I wanted to give you a little
  2446. 1:17:25bit of machine learning stuff and lore,
  2447. 1:17:27but I wanted to make sure that you
  2448. 1:17:29understood all the pieces that go on. We
  2449. 1:17:31learned a little bit about linear
  2450. 1:17:32regression. We learned what the model
  2451. 1:17:34is, right? So this is one of n models
  2452. 1:17:36that you'll see in this class. We
  2453. 1:17:38learned how to solve it and we learned
  2454. 1:17:39that this this very basic algorithm to
  2455. 1:17:42solve it called stochastic gradient
  2456. 1:17:43descent. It's our workhorse and so we're
  2457. 1:17:45going to see that algorithm come back
  2458. 1:17:47again and again and again. Next time
  2459. 1:17:50what we're going to learn about is not
  2460. 1:17:51regression problems. We're going to
  2461. 1:17:52learn about classification. So instead
  2462. 1:17:54of looking at house prices, we're going
  2463. 1:17:56to learn how to classify, you know, cats
  2464. 1:17:58versus dogs or, you know, next word in
  2465. 1:18:00the sentence and all those discrete
  2466. 1:18:01functions. Thanks so much. Have a
  2467. 1:18:03wonderful weekend.

About this transcript

This page contains the full transcript of YouTube transcript (cmNIMjPYdgM) , generated from the public captions YouTube serves with the video. The transcript has 17,221 words across 2,467 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.