YouTube transcript (cmNIMjPYdgM) — Transcript
Full transcript
- 0:05So, good afternoon. Welcome to uh CS
- 0:08229. I'm sorry I didn't get a chance to
- 0:11see you all on Monday. Uh but this next
- 0:13section of the course I'll be teaching.
- 0:14Tenu and I are going to swap off as we
- 0:16go through the lecture. This next
- 0:18segment we're going to cover kind of the
- 0:20basics of machine learning and AI, which
- 0:23is supervised machine learning. This is,
- 0:25as you're going to see, this is when you
- 0:26kind of explicitly tell the machine what
- 0:29you want it to to label. So you could
- 0:30show it an image of a cat or a dog um
- 0:33and tell it to classify as a cat or a
- 0:34dog. We'll start off with something
- 0:36else. It's called regression, which is
- 0:37just a little bit more kind of
- 0:39mathematically simple. And today we'll
- 0:41go through kind of all the building
- 0:42blocks of what a learning algorithm
- 0:43looks like and hopefully something that
- 0:46makes a little bit of intuitive sense.
- 0:47And just by way of introduction, you
- 0:49know, I've been working on machine
- 0:50learning and AI for most of my
- 0:51professional career. It's been really
- 0:53exciting to watch it grow from something
- 0:55that was, you know, very obscure um to
- 0:57now something that has products in your
- 0:59hand that you can use. Um and so I'm
- 1:01super excited about that. So these
- 1:02classes are always really fun for me
- 1:04because you know this stuff that we're
- 1:06teaching here um you know in some ways
- 1:09is the foundation for so much of what
- 1:11what now you kind of use on a daily
- 1:13basis. So it's really exciting to do
- 1:14that. We're going to teach in a very,
- 1:16you know, abstract and mathematical
- 1:17style. I'm sure as Tenu shared with you,
- 1:19so we can communicate really, really
- 1:21precisely. Um, so couple of disclaimers.
- 1:24Let me see if I can get this guy up.
- 1:28Pointer up. So, I'm using this format
- 1:30with slides. So, you can always download
- 1:32the slides online. Uh, I checked them
- 1:34all up, at least for my section. Um,
- 1:37please look at them. It's just because
- 1:38when we're doing things with math,
- 1:40there's all kinds of little I's and J's
- 1:41and indices, and it's just nice to have
- 1:43them all clean. If you see any bugs,
- 1:46please send them to me. Um, these are
- 1:47like handwritten notes that have been
- 1:49passed down that I wrote a long time
- 1:50ago. Um, but the best lecture for
- 1:52everything is Tangu's course notes. So,
- 1:54if you get confused why what's here,
- 1:56this is meant to be a subset of the
- 1:59course notes and it's just meant to give
- 2:01you kind of the tune of what's going on.
- 2:03It's, you know, precise enough that you
- 2:04can understand how all the pieces get
- 2:05together, but if you want to go back and
- 2:07rigorously study, I I would really
- 2:09suggest those notes. Tu work quite hard
- 2:11on on getting those notes in shape. Um,
- 2:13I'm in general worried that the lecture
- 2:15pacing will be too quick with these. I
- 2:17used to in, you know, antiquity give
- 2:19them on these whiteboards. I really like
- 2:21whiteboards because it forces me to slow
- 2:23down and and talk and walk through it.
- 2:25So, that's kind of another way of me
- 2:26asking you. Please just go ahead and ask
- 2:28questions. Like, I have given these
- 2:29lectures a bunch of times. These are
- 2:31topics that I've talked about. I would
- 2:33much rather talk to you if I'm totally
- 2:35honest with you. Um, you know, as a
- 2:36faculty member, I do things outside the
- 2:38university. The reason I stay, you know,
- 2:40affiliated with Stanford is because I
- 2:42like mentoring students. It's a way to
- 2:44kind of have a meaningful life. I also
- 2:46run companies and do other stuff in in
- 2:47another life. Um, but this is what I
- 2:50enjoy is talking to you all. So, please
- 2:52please feel feel free to ask whatever
- 2:54questions you like as we go. Um, I will
- 2:56do my best to answer them. Okay. I
- 2:58generally talk fast when I get excited.
- 3:00I'm not that excited right now, I guess,
- 3:01but I will be excited later. And when I
- 3:03get excited, I'll talk fast. Thank
- 3:05goodness that most of you will watch
- 3:06these on recordings and you can watch
- 3:08them at 50% speed. I'll do my best to
- 3:10slow down, but you know genetically it's
- 3:12been it's been rough. All right, so
- 3:14let's get to it. So we're going to talk
- 3:16about our first first kind of exposure
- 3:19to a formal learning algorithm. The hope
- 3:21is maybe you've seen some of these
- 3:23concepts before. So today is going to be
- 3:24a little bit heavier on notation. It's
- 3:26going to be heavier on kind of the
- 3:27basics that we're going to use. Um and
- 3:30we're going to talk about kind of the
- 3:31most classical learning algorithm which
- 3:33you know predates machine learning by a
- 3:35lot which is linear regression. Okay.
- 3:38Then we're going to see linear
- 3:39regression. We're going to see that this
- 3:40kind of way of fitting a line is
- 3:42actually quite interesting. Like it's
- 3:44it's actually a pretty robust technique.
- 3:46A lot of science was done on it. A lot
- 3:47of machine learning was done on it. And
- 3:49we're going to see elements that have
- 3:50really dominated a lot of what machine
- 3:52learning does now. So, one of the big
- 3:55things that happened in machine learning
- 3:56were, as kind of strange and bizarre as
- 3:59this is to say, we made the models a lot
- 4:01bigger. Like, it's kind of that dumb.
- 4:03Like, we made the models that bigger and
- 4:04you know, as they say, we made the GPUs
- 4:06go burr and then we got AI. Okay, now
- 4:08that version is like not too far away
- 4:11and I like worked on it. Like, you
- 4:12should see my Stanford job talk about
- 4:13like making models bigger way back when.
- 4:16The workhorse algorithm we're going to
- 4:17cover today is this thing called batch
- 4:19and stochastic gradient descent. Okay.
- 4:21Now, if you look at these and you have
- 4:22kind of a mathematical mind or
- 4:23statistical mind, you're going to
- 4:25realize that these are really stupid
- 4:26algorithms. Hopefully, they're really
- 4:28simple. But that simplicity is what
- 4:30allowed us to scale them up and run them
- 4:32on huge amounts of data and build these
- 4:34crazy models. And so, like stochastic
- 4:36gradient descent I've run in the last 24
- 4:38hours for whatever it's worth. Okay,
- 4:39there are much more sophisticated
- 4:41algorithms that you may have seen in
- 4:42optimization classes. We're not going to
- 4:44cover those. Okay. And then the last bit
- 4:46I'm going to tell you a little bit about
- 4:47normal equations. The normal equations
- 4:49basically are an excuse for me to give
- 4:51you some notation about matrices and
- 4:52vectors. And that's because if this part
- 4:54of the section feels a little bit
- 4:56unclear to you, like it's just not kind
- 4:58of clicking to you, you kind of either
- 4:59vaguely remember it or maybe you haven't
- 5:01seen it, we have these wonderful classes
- 5:03on Friday where you can practice things
- 5:05like this and you'll have kind of a
- 5:06little bit of calculus over matrices and
- 5:09vectors. Um, you will end up by the end
- 5:11of this course using that kind of stuff
- 5:12as secondhand, but it can be mysterious.
- 5:15So spend a little bit of time at the
- 5:17beginning and invest and use these great
- 5:19TA courses like this course. One of the
- 5:21nice things about it, you know, the
- 5:22professors are kind of interchangeable,
- 5:23but the course staff like what sets it
- 5:26up so that you have homeworks and you
- 5:28have great TA sections and all the rest.
- 5:31Please avail yourself of that. We want
- 5:32everybody to get through and have a
- 5:33great time. Okay? All right. So that
- 5:35what the last part is like in this part,
- 5:37if you find yourself being like, I don't
- 5:38understand what that symbol means or I
- 5:40don't understand why that
- 5:40transformation, please of course feel
- 5:42free to ask. [snorts] But if it's kind
- 5:44of not sticking with you, please take
- 5:46advantage of those Friday TA lectures
- 5:49and the rest of the resources that are
- 5:50out there. Okay. All right. Let's get
- 5:53started. Okay. So,
- 5:56the basics of supervised learning is
- 5:58going to be a hypothesis. Okay. And this
- 6:00is just a function from some abstract
- 6:02set XX to some abstract exact Y. Um that
- 6:06we're going to do. And it's that
- 6:07hypothesis or prediction. We'll talk a
- 6:09lot about what it means to make a good
- 6:11hypothesis or prediction, but let's look
- 6:12at some examples of kind of the types
- 6:14that we might deal with. So, one example
- 6:16is we could have X being the space of
- 6:18all images. [clears throat] And then Y
- 6:20could be a fixed set of labels like does
- 6:22it contain a cat or a dog or a horse or
- 6:25or a person or a building, whatever.
- 6:27Those fixed set of labels are something
- 6:29that we could learn a function h that we
- 6:31hopefully apply to a bunch of images and
- 6:34then learn when a new image comes that
- 6:36wasn't something that we saw ahead of
- 6:37time. We can still classify it as a
- 6:40horse or car or whatever. Okay, it can
- 6:42be text. It can be house data. Okay, and
- 6:46text you could imagine classifying
- 6:48things whether they're positive or
- 6:49negative, whether it was a positive
- 6:50review about a product, all the rest of
- 6:52it. Okay, when it's house data, we'll
- 6:54look at this. It could also be a price.
- 6:56So y doesn't have to be just categorical
- 6:58labels. Doesn't have to be yes or no.
- 6:59Doesn't have to be, you know, cat or
- 7:01dog. It could be a real number. It could
- 7:03be a scalar value. It could be the price
- 7:05of a house or something like this.
- 7:06Right? So this is kind of what we're
- 7:07doing.
- 7:10Now what makes it supervised
- 7:13is this training set. So what we do is
- 7:16we have we imagine that we have we've
- 7:18collected some data. So we collect some
- 7:20data which are these pairs X and Y
- 7:22pairs. And each one of them as you can
- 7:24imagine is some kind of image and some
- 7:25kind of label for that image. Right?
- 7:27This is a picture. This is an image of a
- 7:28cat. This is an image. This is an image
- 7:30of a dog. Okay. The X is the image. The
- 7:32Y is the label set that's there. Okay.
- 7:35Now we'll use this notation, this kind
- 7:38of superscript notation. What that
- 7:39indicates is that's the first example,
- 7:41that's the nth example if I ever want to
- 7:43refer to them, right? The i example, the
- 7:45J example, something like that. Okay,
- 7:47[gasps]
- 7:48now I've been pretty vague, right? This
- 7:49is a pretty vague definition. I mean,
- 7:51it's a pretty general one. X is just
- 7:52some set and Y is just some set. But
- 7:54that's what's kind of exciting about
- 7:56supervised machine learning. It applies
- 7:57to a huge range of different data types
- 8:00and different things that you may want
- 8:01to do with it. Okay.
- 8:04All right.
- 8:07So, among all of those hypotheses that
- 8:11are out there, and if you've taken any
- 8:13kind of math classes, you know, there's
- 8:14a huge range of hypotheses that go from
- 8:16one set X to another set Y, right? Just
- 8:18a, you know, potentially uncountably
- 8:20many if Y is some big continuous set.
- 8:22So, the question is, what makes a good
- 8:24prediction? Now, this is a subtle
- 8:26problem. This is something we're going
- 8:28to talk about for a while. We're going
- 8:29to talk about good in several contexts
- 8:31and more refined notions of good over
- 8:33the next couple of weeks. But kind of
- 8:35intuitively what we're after in this
- 8:36modeling question is we want a function
- 8:38that you know to use a fancy word that
- 8:40you don't have to use now. We want it to
- 8:41generalize. So what we would like to do
- 8:43is that this x and y that we're
- 8:44selecting here, we'd like it to come
- 8:46from some set that's representative in
- 8:48some way that we can make precise. And
- 8:50then later when we're shown new images
- 8:52that are not from that original set,
- 8:54this thing is going to label it. it's
- 8:56going to be able to look at a new
- 8:57picture of a cat and say, "Yep, that's a
- 8:58cat, not a dog." Okay. And so that's
- 9:00something that's good that
- 9:01generalization that's going to be the
- 9:02heart of machine learning. I train on
- 9:04this this training set. I look at that.
- 9:06I'm going to compute my H. And somehow
- 9:08some properties of the training set and
- 9:10the hypothesis class, that is the hes
- 9:12that I consider are going to mean that I
- 9:14generalize to new and unseen things. And
- 9:16when I mentioned before, what we'll talk
- 9:18about first, there's hes that are really
- 9:19simple. And we can prove that if they're
- 9:21really simple, they have this kind of
- 9:23property that they will transfer. But
- 9:25what has been the you know revolution
- 9:27for the last 10 or so years maybe more
- 9:30is actually training hes that are really
- 9:32really big. That means that they are
- 9:33encoded by huge programs, huge numbers
- 9:35of what are called weights. And we're
- 9:36going to talk about that as we get into
- 9:38more of the class. Okay. So this is the
- 9:41way it works. This is the basic setup.
- 9:43Please feel free to ask questions about
- 9:45it. But this is kind of how it works.
- 9:46And so we're going to study this right
- 9:48now. You could say what else would you
- 9:49do? Well, in the second half of the
- 9:51course, I'm going to tell you about what
- 9:52happens if you don't have any labels.
- 9:53Okay? So this is just one of many
- 9:55machine learning setups you can have but
- 9:57this is by far one of the most you know
- 9:59common and you know kind of working
- 10:01actually it's kind of interesting in my
- 10:02career is like this is a technology you
- 10:04can reliably do this in a number of
- 10:05situations you can make these functions
- 10:07that generalize that's really exciting
- 10:10all right now a little bit of
- 10:12terminology here at the end if y is
- 10:14continuous we call it a regression
- 10:16problem that's real numbers prices we're
- 10:18going to look at regression today the
- 10:20math for that is just easier the
- 10:22derivatives the things that we have
- 10:23comput. They're just easier to compute
- 10:25by hand. They'll hopefully be a bit more
- 10:26familiar. If y is discreet, then it's a
- 10:29classification problem.
- 10:31Classification problems are probably
- 10:34things that you end up solving more in
- 10:35machine learning these days for a
- 10:37variety of reasons. Like for example,
- 10:39the way like your chat GPT works is that
- 10:41it actually is has a classifier head at
- 10:42the end that is guessing what's the next
- 10:44word. That's the kind of way this stuff
- 10:46works. Okay, we'll talk about how those
- 10:48those systems work as well. All right,
- 10:51awesome.
- 10:52Okay, so let's look at a first example
- 10:54of this using probably the most
- 10:56canonical, you know, the most widely
- 10:58used data set out there, which is this
- 10:59housing data set, the Ames housing data
- 11:01set. And this is real data. You can you
- 11:03can pull it down and play with it. All
- 11:04right.
- 11:06All right. So, these are not Bay Area
- 11:08prices, but uh you know, don't hold that
- 11:10against themselves. Okay. Now, what I've
- 11:13done here is I've just taken some data
- 11:14and I put it in, you know, a Python data
- 11:16frame. Um, you can do the same. And then
- 11:20I just looked at, you know, the sale
- 11:21prices and the lot sizes and all the
- 11:23rest. And then on the right, I visualize
- 11:24the data. Okay? So, just one thing in
- 11:26general, like if you have the ability to
- 11:28do it, look at your data. I can't stress
- 11:30that enough. Look at your data. It's
- 11:32like you're pretty good pattern
- 11:33recognizers, take advantage of that.
- 11:36Okay? Even if you're building some
- 11:37complicated crazy AI system, still look
- 11:39at your data. You will discover things.
- 11:41All right? So, we're looking at this
- 11:42data here. We have the sale price. We
- 11:44have the lot area. And I've plotted it
- 11:45in some way there. Okay. All right. Now,
- 11:50we need to get one more character in our
- 11:51story. We need to get a hypothesis. So,
- 11:53one popular choice are these linear
- 11:56hypotheses, right? And you may think
- 11:58that these linear hypotheses, you know,
- 11:59they seem relatively simple. That's the
- 12:01hope, but they're actually industrially
- 12:03used, right? These are kind of the
- 12:05workhorse of what we're going to do.
- 12:06These kind of linear classification
- 12:07models, they show up inside pretty much
- 12:10every model that you use. The last step
- 12:11is this linear model. So, these are
- 12:13actually, in spite of their simplicity,
- 12:15pretty widely used. Okay. So, what is
- 12:18age? We call them linear. It's a little
- 12:19bit of an abusive terminology. You'll
- 12:21see why we do in a second. It's because
- 12:22of a convention. This is technically an
- 12:24aphine function because it has this
- 12:26little offset here. The way it works is
- 12:29you feed me an x. Okay, so let's imagine
- 12:31x is a scaler and then I'm going to
- 12:33multiply it here times that theta 1 then
- 12:36add it to theta 0. Okay, and so this
- 12:38gives me something wherever my x is.
- 12:40This is a very simple function that's
- 12:41taking in scalers, but this is my
- 12:42hypothesis class. Okay.
- 12:45[sighs and gasps]
- 12:45All right. So
- 12:48if you look at an example prediction
- 12:50here, let's say that you wanted to in
- 12:52our setting, we have a bunch of x's that
- 12:53are right here that we want to predict.
- 12:55And we like to predict some function
- 12:56that goes from the size of the of the
- 12:58house to the price in whatever units.
- 13:00Don't don't worry. Don't think too hard
- 13:02about what size and price mean here. I
- 13:03don't know like what what reasonable
- 13:05mapping I I don't think I've ever been
- 13:06to as um but these are apparently real
- 13:09prices. Okay. [gasps] All right. So we
- 13:11have to do this example prediction from
- 13:13size to price. So, we want to get a
- 13:14function, a linear function that takes
- 13:16this size in as the x and then out comes
- 13:19a price. Clear enough what our goals
- 13:21are? All right,
- 13:23man. That's got to stop. [sighs] All
- 13:26right. So, notice that this prediction
- 13:28now is instead of all the wild
- 13:31hypotheses that could be out there,
- 13:32right? There's a ton of ton of things
- 13:33that go from functions that go from a
- 13:35set of real numbers to another set of
- 13:36real numbers. We've now really
- 13:38constrained it. It's entirely defined.
- 13:40The fancy word for this is
- 13:42parametrically. It's entirely defined by
- 13:44these two little parameters. No matter
- 13:46how many houses we see, we're only going
- 13:48to, as we'll say, learn those two
- 13:50parameters or estimate those two
- 13:51parameters and that's going to be the
- 13:53way that we transfer. Okay, so it's a
- 13:55huge reduction in the space of functions
- 13:56that we've just done. So like it seems
- 13:58like maybe you know not very much
- 14:00happened on the slide, but like we went
- 14:01from uncountably many functions, we
- 14:03still have uncountably many functions,
- 14:04went to this really small class of
- 14:05lines. Okay, so what do these look like?
- 14:08By the way, just to just to draw them,
- 14:09just to make sure you're extra clear why
- 14:11we call them linear functions. Well,
- 14:12where do they what is their value at
- 14:14zero? Well, that value of zero is going
- 14:16to be theta kn. And then it's going to
- 14:18have a slope. And that slope at one is
- 14:20going to be here theta 0 plus theta 1.
- 14:23Okay. This is one.
- 14:26Hopefully that's clear.
- 14:28All right. Awesome.
- 14:31Okay. So, with that, we're going to fit
- 14:34a line to this data. Now, which line are
- 14:36we going to fit? We're going to come to
- 14:37that in a second. But intuitively, what
- 14:39should that line? What would be a good
- 14:40line to fit? Well, it's one that when we
- 14:42give it an X, right? So, I feed it in
- 14:44some X here. Let's say this point. When
- 14:47I look at the line, my prediction is on
- 14:49the line. So, if I fed it in something
- 14:51at 1200, my prediction would be this
- 14:52point here. If I could draw straight
- 14:54with this thing. Let me try to draw a
- 14:55little bit straighter. Anyway, my
- 14:57prediction would be on that line. Is
- 14:59that clear? When I take the X, I put it
- 15:00into the hypothesis, it gives me that
- 15:02line. Cool. Now, what would be good?
- 15:05Well, one thing that would be good is
- 15:07because we have this training set is
- 15:08that every time we put a point into this
- 15:10data set, whether it's here, here, we
- 15:12didn't do such a good job, right? We're
- 15:13really, really far from the line here
- 15:15and here, we seem to have done a really,
- 15:17really great job. Okay? And so that
- 15:20we're going to try and do those errors
- 15:21that we make. Those are how far off we
- 15:23are in the price. We're going to try and
- 15:25minimize those. This has a really fancy
- 15:28name. It's called empirical risk
- 15:29minimization. I don't think you need to
- 15:31know that name, but you'll sometimes
- 15:32hear me slip and say erm, and that's
- 15:34what empirical risk minimization means.
- 15:36Okay, cool.
- 15:38So, here's a nicer picture that I
- 15:40generated where the functions behaving
- 15:42much more nicely. So, each point here is
- 15:44one of those data points that we saw.
- 15:46And here we're trying to minimize this
- 15:49little line through it. Right? Now, the
- 15:52idea of course, which is nice about the
- 15:54line, is if you just had the training
- 15:56set, the way you would potentially do a
- 15:58prediction, there's actually not a bad
- 15:59way to do it. We won't talk about it too
- 16:00much today. You do what's called nearest
- 16:02neighbors. You take in an X, you kind of
- 16:04look in the neighbor of X, maybe you
- 16:05average them. That could be [snorts] an
- 16:07interesting prediction function. What's
- 16:09nice about H is it extends to every
- 16:11single value and it extends in kind of a
- 16:13smooth way. And so hopefully what that
- 16:15means is that if there's lots of
- 16:17girrations here, we capture that main
- 16:19trend. we get the main part of the of
- 16:21the of the prediction which makes it
- 16:23kind of a good predictor overall. Okay.
- 16:29All right.
- 16:31Now, one other thing which you can see
- 16:34from this. Oops. Let's get that guy up.
- 16:36There's a slight mismatch there. Sorry
- 16:38for that. So, as we look through here,
- 16:40if you look at that line that I drew,
- 16:44you could argue that you should draw a
- 16:45different line. You're like, "Oh, I like
- 16:46this other line a little bit better."
- 16:47You know, kind of goes Oops. Let's go in
- 16:49here. goes like this all the way. Oh, I
- 16:52definitely don't like that line. There's
- 16:53one line there. Maybe like I draw like
- 16:55this. It's a terrible line. So, how did
- 16:57I pick this one line that's in there?
- 16:59And that's what we're going to do this
- 17:00erm this this thing. We're going to look
- 17:01at the data and we're going to have to
- 17:03be able to compute that. But the point
- 17:04is is no line looks like it's absolutely
- 17:06perfect. And that's going to be a theme
- 17:08throughout most of what we do. Your data
- 17:10has a couple of different things in it.
- 17:12Sometimes it's imperfect because we're
- 17:14not modeling something. So probably the
- 17:16rate that you're willing to pay for a
- 17:18house depends on more than like the lot
- 17:20size, right? Like you don't just care
- 17:22how big the lot is. You may care if it's
- 17:24a nice house, if it's close to things
- 17:26you care about, all the rest. We'll come
- 17:28back in a minute about how we
- 17:29incorporate that. But the other thing is
- 17:31even if we added in all those features,
- 17:33usually there will be some error. And so
- 17:35we're going to worry a lot in this class
- 17:37about the error of how we do in these
- 17:39predictions and how well the error on
- 17:41our training sets manifests on the what
- 17:43are called the test sets when we take it
- 17:44out into real life.
- 17:47Okay. All right. Let's look at something
- 17:49slightly more interesting. So here we're
- 17:51going to add in the number of bedrooms.
- 17:52Okay. So what do you think is going to
- 17:54happen? Well, now we need a function
- 17:56that can take in pairs instead of single
- 17:58values. So we need to change our
- 17:59hypothesis class a little bit.
- 18:03Lot size and size. I guess size is the
- 18:05refers to the house size. Apologies. Now
- 18:07we have this function here. Okay. So how
- 18:09does it work? Well, if you give me you
- 18:11feed me an x, then I take this guy, put
- 18:13him here, put this one here, put this
- 18:15one here,
- 18:18and then I hopefully have weights. Now
- 18:19I'm parameterized by that came all back
- 18:22by just those weights now. So now my
- 18:23function is parameterized by four
- 18:25weights.
- 18:27And hopefully it's going to fit my data
- 18:29a little bit better. Right? The reason
- 18:30it should be a little bit better is I
- 18:31could always set some of those weights
- 18:33to zero, right? And get back to the case
- 18:35where I only had two weights. So it
- 18:36seems like it's a larger or more
- 18:37expressive class.
- 18:40But if lot size and bedrooms influence
- 18:42price, we would hope that this
- 18:44hypothesis class is richer is the
- 18:45terminology. And that would allow us to
- 18:47fit our data a little bit better.
- 18:50Okay.
- 18:51Now this is one of the reasons we call
- 18:54these functions linear not aphine is
- 18:56that we're going to operate under this
- 18:57nice little convention that we're going
- 18:59to always assume that x1 x0 is one here
- 19:02which we didn't even bother to write and
- 19:04I highlight this convention because it's
- 19:06used throughout the notes and when I
- 19:07deviate from it I will explicitly tell
- 19:09you but whenever we're doing regression
- 19:10or classification this is what we have
- 19:12okay so that allows us to write a nice
- 19:15little form like this h of x is just
- 19:17this nice little sum we don't have to we
- 19:19don't have to special case the the theta
- 19:220 term and we can you know easily extend
- 19:25from three to however many dimensions.
- 19:28Awesome.
- 19:31Okay. So that all leads into as I said
- 19:34notation. So if this is unfamiliar to
- 19:36you please go to the Friday you know
- 19:40sessions. A lot of this is us kind of
- 19:42recalling notation where you're like oh
- 19:44I know that I need to look at it or oh I
- 19:46need to go and you know avail myself of
- 19:48some of the the extra resources. So this
- 19:50notation here we'll write theta but we
- 19:53really mean this vector which will also
- 19:55live in R4 it's four real numbers okay
- 19:58four scalers this thing here notice now
- 20:01this x1 we're writing the entire example
- 20:04in here and this one is one okay so
- 20:07here's one here's 214 here 44 45
- 20:14clear enough
- 20:16Okay,
- 20:19so we're going to call these these
- 20:22thetas here parameters and these XI are
- 20:24going to be input or features. Remember
- 20:25when I talked to you about before which
- 20:27I said, you know, you could add in these
- 20:28extra features of your problem. That's
- 20:30for now going to be a modeling decision.
- 20:32Later, we're going to see an interesting
- 20:33idea about how we can even learn those
- 20:35features. That was part of what was
- 20:36called the deep learning revolution
- 20:37where you were able to actually imputee
- 20:39the representation of your data just by
- 20:41throwing more data at it. We'll see that
- 20:43in like couple weeks. But for now, it's
- 20:46a modeling decision. You come and write
- 20:48some rule, some piece of code that takes
- 20:50the number and puts it in as the
- 20:51feature. And it's not trivial, right?
- 20:53Like, you know, it normalized the lot
- 20:54size here. It said 45 instead of 45K.
- 20:57But in general, they could be even more
- 20:59interesting features of your problem.
- 21:00And the point is is that extends
- 21:02naturally to linear models. Okay, so far
- 21:06so good. All right. So, X and Y has a
- 21:10fancy name, training example, right?
- 21:12It's a it's an element of the training
- 21:14set. And the reason that we call it
- 21:16training is we're going to take that set
- 21:17and we're going to train our model on
- 21:19it, right? It's kind of this the
- 21:20terminology for it. We're going to try
- 21:21and figure out among all the thetas,
- 21:23which are our predictions, which ones
- 21:24fit our data best. We've been vague
- 21:26about what data best means, but there's
- 21:28a lot of ways to do that, it turns out,
- 21:29but they all follow the same recipe. And
- 21:31then, as I said before, I just wanted to
- 21:33highlight the example again that Xi and
- 21:35Yi are an individual example.
- 21:37Okay.
- 21:39All right.
- 21:43[clears throat] So in notation in the
- 21:44class you will always see this is
- 21:46something that you will see this is as
- 21:47convention. We'll always have n be the
- 21:49number of examples. So you should kind
- 21:51of get that in your mind and d be the
- 21:52number of dimensions or features. So
- 21:54we'll consider these either d or d plus
- 21:56one dimensional plus one because of that
- 21:58convention. Bless you that x0 equals 1.
- 22:02Please.
- 22:06>> Yeah. So, we will almost always use
- 22:08lowercase and we will try not to abuse
- 22:11you by having upper and lowerase ends in
- 22:13the same lectures. Uh but I cannot
- 22:14guarantee that because especially if
- 22:16they're copied from handwriting. Uh
- 22:17often often we'll use them we will try
- 22:19not to use them both in the lecture to
- 22:20mean two different things, right? So,
- 22:22usually we'll be a little bit
- 22:23consistent. Yeah. The other convention
- 22:25you'll sometimes see sneak in is P's for
- 22:27parameters, but that's like a stats
- 22:28convention and it doesn't really matter.
- 22:30None of this will be really matter. If
- 22:32you're confused, just ask. Cool. All
- 22:34right. All right. So, what do these
- 22:36things look like? These linear models
- 22:38look like Here's one in two dimensions.
- 22:40This is about as high as you can really
- 22:42reasonably look on your data set. So,
- 22:44here we have a bunch of different
- 22:45points. We have price. We have, you
- 22:47know, bedrooms and square feet at the
- 22:48bottom. And I've I've plotted a plane
- 22:51here. Okay, an aphine surface. You could
- 22:54imagine a classification that was some
- 22:56kind of crazy curve. We'll see how to do
- 22:58that later. There's all kinds of nice
- 22:59ways to do that, but for right now,
- 23:01we're looking at these big linear
- 23:02models. And you see this is kind of a
- 23:04good thing in the same way. So, how
- 23:05would you define error? Well, how close
- 23:07is your prediction? If you took that
- 23:08point, that feature point, and you went
- 23:10into space, how far away are you from
- 23:12this plane? Now, we're going to consider
- 23:15as we get here, like modern models, as I
- 23:18mentioned, just got bigger. They have,
- 23:21you know, thousands, hundreds of
- 23:22thousands, millions, billions, trillions
- 23:24of parameters. Some of the models now
- 23:25have trillions of parameters in them.
- 23:27They're not all fed into a line like
- 23:29this, but they'll conventionally have
- 23:3110,000 dimensional spaces that are the
- 23:33last piece of the model.
- 23:36That's really hard to visualize. So,
- 23:38we're going to use a lot of vector
- 23:39notation and other things to deal with
- 23:40it. And that's why we use this notation
- 23:42so you can simplify and think about it
- 23:43and understand how to manipulate it in
- 23:45one or two dimensions. You can kind of
- 23:46picture it. Okay. But high dimensional
- 23:48space is a weird thing.
- 23:51Okay.
- 23:52All right.
- 23:54[gasps] Okay.
- 23:58So let's get back to it. So we want to
- 23:59choose some theta such that as we talked
- 24:02about multiple times,
- 24:04h of x is approximately equal to y. If I
- 24:06give you an x and a new thing, you want
- 24:08to be approximately close. You like the
- 24:09price to be close and so on. So how do
- 24:13we do it?
- 24:15Well, here comes le squares. Okay, so
- 24:17maybe you've seen this before. I hope
- 24:19you've seen before. By show of hands,
- 24:20who's seen le squares before? Awesome.
- 24:22Fantastic. All right, good. So le
- 24:25squares super super conventional thing
- 24:27if you haven't seen it before please
- 24:28again Friday is a good place to look at
- 24:30it we'll cover it again just in our
- 24:31notation to make sure everything is is
- 24:33is all right now what happens so you
- 24:37have this idea here j theta that's going
- 24:39to be the loss that you incur and it's
- 24:41broken down into terms on every single
- 24:43element of the data set and here you
- 24:46have h theta of x i minus yi so this is
- 24:49the error remember we kept drawing those
- 24:51pictures like this guy right here these
- 24:53two things are the same. This error,
- 24:55oops, that looks terrible. We'll wait
- 24:58for it to go away. Anyway, so a state
- 25:01the x ius yi, that's the prediction
- 25:03error. Then what we do is we square it.
- 25:05And we square it because we want to have
- 25:07something that we have a positive or
- 25:09non- negative number there. It could be
- 25:10zero if it were exactly equal. And then
- 25:12what we're going to do is we're going to
- 25:13try and pick the theta that minimizes
- 25:16it. Okay, just by show of hands, how
- 25:18many people know the arg notation?
- 25:21Perfect. All right, we're doing great.
- 25:23Okay. So then in that case, you know
- 25:24exactly what goes on here. We're going
- 25:26to solve over j theta and we're going to
- 25:27try and pick the thing that minimizes
- 25:30this function. And so intuitively what
- 25:31that says is we're going to pick the
- 25:33theta so that on average over all of our
- 25:35training points,
- 25:37we minimize that squared error. Okay? So
- 25:39we're going to instead of measuring the
- 25:40error just as a sum, we're going to
- 25:41square that distance. So things that are
- 25:43really close, they're going to, you
- 25:45know, go down. Things are really big,
- 25:46we're going to pay for.
- 25:48Okay? Now, why do we do squared? Oh, go
- 25:50ahead.
- 25:51>> Sorry. I might be um uh you might be
- 25:56saying it but the squaring because
- 25:59presumably you could take another power.
- 26:01>> Yeah, sure thing to even be even
- 26:04stricter.
- 26:06>> Yeah, exactly. Right. So the the
- 26:07question is is why did you pick two here
- 26:09and and why why did you do it? That's a
- 26:11great question. So could you pick other
- 26:13values? Certainly you could pick another
- 26:14power. You could pick the absolute
- 26:16value, right? That's a fine thing to do
- 26:17as well. The you know that's a that's a
- 26:19value regression kind of thing. two has
- 26:22a nice property which we're cheating on
- 26:24for a lot of le squares. two we can
- 26:25solve exactly and that just turns out to
- 26:27be because two is as you'll see in a
- 26:29second when we compute all the
- 26:30derivatives and do all the nice things
- 26:32the derivative has a really nice form or
- 26:34the gradient has a really nice form and
- 26:36that's going to let us solve it exactly
- 26:38historically also if you care like many
- 26:40of the you know people who were doing
- 26:43science way back when when they were
- 26:44doing you know Gaus was doing le squares
- 26:46because of that computational advantage
- 26:48they could predict all kinds of nice
- 26:49things it was an easy problem now you
- 26:51could look at that and I will derive for
- 26:53you in the next class I'll der for you
- 26:55how that le squares comes up and relates
- 26:57to a fancy statistical assumption which
- 27:00is that the errors are approximately
- 27:01normal which we'll come in and talk
- 27:03about a little bit later and so you'll
- 27:05be able to derive other regressions that
- 27:07are in there but the special things
- 27:08about le squares are it's extremely
- 27:10popular uh in like history not like you
- 27:13know I think it's great right now like
- 27:14it's extremely popular in history and
- 27:16it's very easy to solve okay and that
- 27:18the reason I hammer on this so hard is a
- 27:20lot of machine learning if you look at
- 27:21it and and I think honestly I was guilty
- 27:23of this so I grew up as like a math math
- 27:24person. And so when I came to machine
- 27:27learning and AI, I kind of had this idea
- 27:29that I was like, well, why are you doing
- 27:30like simple algorithms? Why are you
- 27:32like, you know, worried about these
- 27:33other things? And I've come to realize
- 27:36that really what goes on in machine
- 27:37learning that makes it so powerful is
- 27:38this trade-off between kind of how we
- 27:40compute things and kind of how we
- 27:42predict them. And actually when machine
- 27:44learning to me really intellectually
- 27:45came into its own as a field was when it
- 27:48broke away a little bit from stats and
- 27:49started to train these crazy large
- 27:51models and understand more what were the
- 27:53computational limits of what kind of
- 27:55models we could have. So sorry for the
- 27:56long digression but I love questions so
- 27:58I'm just pumping trying to get more.
- 27:59That was a wonderful question. Please
- 28:03>> yeah the 1/2 is there just by
- 28:04convention. So one of the things is this
- 28:06arg is insensitive to constants. If I
- 28:09change this to, you know, 25 over 12, it
- 28:12would, this would be totally unaffected.
- 28:13I'd get the same theta. So the constant
- 28:15doesn't matter. The reason we do this is
- 28:17it's this convention. When we compute
- 28:19the derivative in a minute, these things
- 28:21are going to cancel and that just makes
- 28:22things a little bit nicer. It's kind of
- 28:23like the cooking show view of math where
- 28:25we like set it up so that it looks nice
- 28:27and when you do it by yourself, it's
- 28:28always a gigantic mess and then you you
- 28:30clean it up later. That's all it's there
- 28:31for.
- 28:32>> Please. uh with RM like the theta that
- 28:35you kind of put it to minimize like are
- 28:37they discrete or
- 28:39>> oh fantastic yeah so one of the things
- 28:40is is remember our hypothesis class
- 28:42right now the way we picked it was to be
- 28:44continuous so these can be any real
- 28:46value that we want and so when we plug
- 28:48these things in here we're going to have
- 28:49continuous values so we're going to do
- 28:51continuous optimization discrete
- 28:53optimization you can do there are there
- 28:55are procedures that people do fancy ones
- 28:57you know integer linear programming yada
- 28:59yada yada those are generally pretty
- 29:01hard most modern machine learning One of
- 29:03the tricks are even when you have
- 29:04something that's discreet like our
- 29:06models that work on words underneath
- 29:08their covers they're actually doing
- 29:09something continuous and so the
- 29:10parameter space will almost always be
- 29:12continuous so you can compute things
- 29:13like gradients and derivatives and and
- 29:15for other optimization reason it's a
- 29:17it's a very insightful question. Yeah
- 29:20please
- 29:22larger differences get more
- 29:26>> yeah that's a great question to think
- 29:28about. So these error functions this is
- 29:30this this thing here when we incur an
- 29:32error in prediction the question is how
- 29:33much penalty do we pay right now we're
- 29:35paying with this square and as we've
- 29:36talked about it's for computational and
- 29:38modeling reasons kind of a good mix you
- 29:41could put other things in there I don't
- 29:42want to make too big a deal out of it
- 29:44because honestly when you get to a lot
- 29:46of data the the form of this doesn't
- 29:49really matter too much there is one very
- 29:52important form that we'll cover next
- 29:53lecture which is the is a is a for
- 29:56classification when you have a discrete
- 29:58that and then you do something called
- 29:59softmax and softmax is probably the
- 30:02function that I use most in my day um
- 30:05when you build these systems and
- 30:06probably if you are doing all your
- 30:07homework with chat GPT like the rest of
- 30:09your classmates then you probably use it
- 30:11too okay because that's what the that's
- 30:12the workhorse that's underneath there
- 30:14but this one is just the one we're going
- 30:15to do for for computational reason it
- 30:17has a nastier derivative
- 30:20other questions this is awesome please
- 30:22ask me questions
- 30:24all right
- 30:26we'll keep
- 30:28All right. So what do we see? We're at
- 30:29the end of our linear regression.
- 30:32We saw our first hypothesis class aphine
- 30:35and linear. We're going to see many many
- 30:37more richer classes throughout this. And
- 30:39richer means that they're more
- 30:40expressive. Means not only do they have
- 30:42more parameters, which as I've hinted
- 30:43several several times, modern models do,
- 30:46but they're actually going to have
- 30:47crazier surfaces as well. So they're not
- 30:49just going to be lines, they're going to
- 30:50have kind of crazy interesting surfaces
- 30:52underneath them. We refreshed ourselves.
- 30:54And as I said, if these part was
- 30:55unfamiliar to you, you were looking at
- 30:57this and you're like, I've never heard
- 30:58anything this guy's talking about. Okay,
- 31:00maybe I did a bad job. I'm willing to
- 31:01admit that. But if I didn't do a
- 31:03terrible job, then look and say, maybe I
- 31:05need to refresh. And I cannot tell you
- 31:07how much I want you to if you need to go
- 31:09take advantage of those TA resources and
- 31:11others because putting in a little bit
- 31:13of investment now makes the class go a
- 31:14lot lot more smoothly for you.
- 31:18We also saw something that was this
- 31:19paradigm that you guys picked up on
- 31:20right away, which is great, which is
- 31:22that a good hypothesis is one that's
- 31:24close to the data. And I want to point
- 31:26out something which is a little bit kind
- 31:27of spooky about what we're going to
- 31:28claim here. And this is like a
- 31:30foundation of statistics. So it's not
- 31:31that spooky, but is if we believe that
- 31:34training set, if we fit our hypothesis
- 31:37on that training set, we believe it's
- 31:39going to work well in the real world.
- 31:41That's a little bit of a leap of faith.
- 31:42And so we'll justify that leap of faith
- 31:44in all kinds of ways. some of which I'll
- 31:46give you a lecture in the middle about
- 31:48how they can go terribly wrong and I
- 31:49will pick from my own you know work
- 31:51where we thought we were selecting data
- 31:53in a way at training time that was going
- 31:55to match test and we were completely
- 31:57wrong and made fools of ourselves okay
- 31:58so you know economic indicators and all
- 32:01kinds of fun stuff like that and so that
- 32:03can happen but I want to be aware this
- 32:05paradigm is what machine learning does
- 32:06says look this training set is somehow a
- 32:08reflection it's captured the real world
- 32:10you have to believe that and then
- 32:12fitting on that data set is somehow
- 32:13going to carry out something nice in the
- 32:15real world. Okay, it's like the
- 32:16cornerstone of statistical estimation
- 32:18and inference. Cool. Great. We also
- 32:22talked a bunch about objective
- 32:23functions, J. We're going to see those
- 32:25over the next lecture. We're going to
- 32:26see, you know, the canonical one for um
- 32:28regression today and we'll see the
- 32:30canonical one for classification next
- 32:31week.
- 32:34Okay. All right. Now, as I mentioned,
- 32:37and I I'm biased here, so I should also
- 32:39tell you. So like what I work on in
- 32:40research, as I mentioned, is I really
- 32:42like the problems of scaling up these
- 32:44models. We've done all kinds of things
- 32:45underneath the covers to do that to make
- 32:47these large models that today run you
- 32:49know on GPUs. We've contributed our
- 32:51little brick to that out of my lab which
- 32:52I'm you know very proud of wonderful
- 32:54students and and collaborators. But one
- 32:57of the things that you'll see in machine
- 32:59learning is this ability to compute
- 33:01these functions at large scales is
- 33:03critically important. That's been the
- 33:05revolution that's allowed us to break
- 33:07away from doing small kinds of problems
- 33:09to the large. And that's what we're
- 33:10going to cover over the next couple
- 33:11weeks. But we'll start where the field
- 33:12started.
- 33:13>> [snorts]
- 33:14>> What that boils down to is basically
- 33:17solving a set of these equations. That
- 33:18arg min'll
- 33:21talk about a loss function. Being able
- 33:24to take a loss function and run you know
- 33:26an optimization procedure fancy one
- 33:29called backrop that you'll learn in the
- 33:30course underpins pretty much 99% of the
- 33:33models that you're likely to encounter
- 33:34today. Okay. So we're going to see the
- 33:36warm-ups to how you get to backrop the
- 33:38simplest thing that starts there which
- 33:40is the computation of a derivative. And
- 33:41then in this case, we'll also see how to
- 33:43solve these equations exactly. That's
- 33:45very unusual that you can do that, but
- 33:47it will basically be an exercise with
- 33:49some of that notation. So if I haven't
- 33:50made it clear, we'd really love you to
- 33:52use those those Friday lectures. All
- 33:54right, I'm not getting paid if you go. I
- 33:56know you probably at this point you're
- 33:56like, "Oh, you must be getting paid."
- 33:57Like, no, no, it's just I think it's
- 33:58good for you. Okay. All right. So, let
- 34:00me solve these le squares problems.
- 34:03All right. So, I'm going to draw a
- 34:05little picture here.
- 34:08All right. So let's take a function.
- 34:11This will be our quadratic which is kind
- 34:13of like our our classical function. And
- 34:16uh you know we'll look at this this kind
- 34:17of bullshaped function. Now when we
- 34:20compute the the quadratic so you know
- 34:22we'll take something here which like you
- 34:24know we'll say it's uh x^2 what am I
- 34:27going to write here? Oops. So it's like
- 34:29x^2 plus I don't know some theta 0
- 34:32whatever something right here. Okay. Now
- 34:34I can make this call this 2x squ I don't
- 34:36know whatever you want. Okay. Now this
- 34:38is f. You all remember how to compute
- 34:40derivatives for this I hope but if you
- 34:42don't remember the picture you pick a
- 34:44point let's say your first point here
- 34:47which is going to be theta 0 which is
- 34:49going to be our first guess. Okay, you
- 34:52then
- 34:55pick this line which is the gradient,
- 34:57the best linear approximation, right?
- 34:59You remember your your tailor expansion.
- 35:01It's you're going to be f of theta 0 and
- 35:04then you're going to be plus fprime of
- 35:06theta 0 time x - x theta or theta. Whoa,
- 35:11that's bad.
- 35:13All right, I'm going to rewrite this. F
- 35:16of theta 0 plus frime of theta 0. And
- 35:20then I was just trying to write theta
- 35:21minus theta 0. Okay, nothing deep going
- 35:23on. This Taylor's rule for the first
- 35:25first derivative. And that's that line
- 35:27that I've drawn there. And if you don't
- 35:28remember that, remember it's the secant
- 35:29that comes in and approximates the
- 35:30gradient. Okay, slide it down in our
- 35:32head. Okay, once we do that, we have the
- 35:35direction of maximal increase. And so
- 35:37then we plug that that character in that
- 35:40this fun gradient here.
- 35:43We're going to plug that in in this
- 35:44notation.
- 35:46Okay. Now J may have some higher
- 35:50dimensional structure. So this guy is
- 35:52not a derivative of one variable. It's a
- 35:54partial derivative. If you don't
- 35:55remember that thing, please go ahead and
- 35:57look it up. And the idea here is that
- 35:59what we're going to do is we're going to
- 36:00iterate and walk down in the direction
- 36:03until we get to the bottom here. Okay.
- 36:06And so what's happening is we start with
- 36:08initial guess. I didn't start at zero. I
- 36:09should have slid this over at zero. We
- 36:11start with initial guess. Could be a
- 36:12random number. Could be zero.
- 36:14Interestingly, in modern models, where
- 36:15you start matters. doesn't matter for le
- 36:17squares does matter in general. And then
- 36:20we're going to iterate these equations.
- 36:21We're going to start with our current
- 36:22guess. We're going to compute this
- 36:24gradient or derivative and we're going
- 36:26to try and minimize it. So we're going
- 36:27to walk in the opposite direction. How
- 36:29far are we going to walk?
- 36:32This thing here is going to tell us this
- 36:34is called a step size. Okay.
- 36:38This guy here is a step size.
- 36:41All right. Now, if you look at this and
- 36:43you're mathematically minded, you may
- 36:44think that this is horrifying.
- 36:46You're computing only first order
- 36:48information, right? You're just looking
- 36:49at a gradient.
- 36:51You are only taking a step size. That's
- 36:53like, you know, not searching to the
- 36:55most minimal part you could get, what's
- 36:57called line search. You're not doing
- 36:58that. You're just kind of grabbing some
- 37:00information about the function locally
- 37:01and taking a quick step. And the reason
- 37:03that can be so disastrous, you may have
- 37:05been told in your calculus classes, is
- 37:07because you could have a function that
- 37:08looks like this. And then here you walk
- 37:11down and you get stuck in a local
- 37:12regime. Okay? So hopefully this is
- 37:14familiar to you. By show of hands, do
- 37:16people understand the difference
- 37:17convexity, non-convexity? Awesome.
- 37:18Perfect. Okay. So that's how it looks.
- 37:22All right. Fantastic. We do this for
- 37:24every single component and we update.
- 37:27That's how we solve it.
- 37:29All right.
- 37:33So when we do this, we get this learning
- 37:35rate or step size. You'll hear me say
- 37:37those terms mean the same thing. I wish
- 37:39that they were unified in my head. Um
- 37:41but I will use them interchangeably. Um,
- 37:43apologies in advance. Here's me
- 37:45computing the derivative. So, those
- 37:47folks who were asking, as you were
- 37:48asking earlier, what's this two for?
- 37:50It's so that I can cancel this guy.
- 37:51Okay, that's all that's happening. I
- 37:54have my error term here and then a
- 37:56gradient here. Now, hopefully this is
- 37:58not calculus that's bothering you. If it
- 38:00bothers you, again, see the Friday
- 38:01lectures. But there's one thing here and
- 38:04the way that I've written it that I want
- 38:05to call out because we're going to use
- 38:06it again and again. this thing here,
- 38:09this error term, you basically always
- 38:11have this in a ton of these rules. You
- 38:13have the error, what you have, and then
- 38:15you have it multiplied by something
- 38:17here. It's just going to be the gradient
- 38:18of the function. But in general, for a
- 38:20nonlinear, there' be some extra
- 38:22[laughter] extra stuff there. Okay? And
- 38:24you'll see this form again and again.
- 38:26So, it's like kind of an artificial way
- 38:27to write it, but it's an important
- 38:28artificial way to write it. Okay? So,
- 38:30what's going on here? When I compute the
- 38:32derivatives, I'm computing the
- 38:33derivatives on all points in my training
- 38:35set. Now, if you think about that, if I
- 38:38give you the entire internet and ask you
- 38:40to predict the next word in the entire
- 38:42internet, that's going to be pretty
- 38:43slow. If every time you update your
- 38:46model, you have to scan the entire
- 38:47internet, that seems really slow. And
- 38:50so, what machine learning people did is
- 38:51they started to try and get algorithms
- 38:53which were even simpler than this, even
- 38:55dumber than this. You're like, gradient
- 38:56descent, how could you get dumber? Don't
- 38:58worry, machine learning has got you
- 38:59covered. We have a lot more dumb
- 39:00algorithms. How does it work?
- 39:03>> [cough]
- 39:04>> Okay,
- 39:06[clears throat] we'll get there in one
- 39:07second. So, we computed the derivatives.
- 39:09We have this one thing. I don't need to
- 39:11talk about this. All right.
- 39:14Okay. Get to it. Now, one thing I did
- 39:17want to cover before we get there is
- 39:19apologies. This thing here about how do
- 39:21you set the step size? Okay. Now,
- 39:25[clears throat]
- 39:26when you set the step size, think about
- 39:28what it means. The step size is
- 39:30basically how sure you are in that
- 39:32information. In some way, you get the
- 39:34information. If you really trusted that
- 39:35that gradient that it was leading you
- 39:37really steeply down the hill, you would
- 39:39want to take a really big step if you're
- 39:41far away from the optimal, right? And so
- 39:43it turns out for functions that are nice
- 39:44and bullshaped, which are called convex.
- 39:46If you're really far away from the
- 39:47optimal, it's usually very steep. Okay,
- 39:49that's a there's a formal version of
- 39:50that, but that's what's going on. So,
- 39:51you want to take a step. Often in
- 39:54machine learning, we're dealing with
- 39:55really nasty functions. we're dealing
- 39:56with those functions that are really
- 39:57curvy or we don't know what the surface
- 39:59looks like and so we tend to take
- 40:01relatively small steps. If you take
- 40:03steps that are too small, what happens
- 40:06is
- 40:07[snorts] you just don't get to the
- 40:08optimum fast enough. Okay, if you take
- 40:11steps that are too big, this actually
- 40:12looks pretty good to me. It says too
- 40:13large, but like if you just take one
- 40:15step and got to the optimum, like you'd
- 40:16be really happy. So that'd be great if
- 40:18you did that. So this is like too large
- 40:19is kind of like weird. Like it's like
- 40:20bragging like, "Oh, it's too large. It
- 40:22works too well." I don't know. In
- 40:23general, what happens in too large, I'll
- 40:24show you in a minute, is that you bounce
- 40:26around. So, you take a step, but you
- 40:28shoot past the optimal. So, if you
- 40:30picture that bowl that we were talking
- 40:31about before, you would bounce around
- 40:32from both sides. And I'll show you a
- 40:34picture of that. Please question.
- 40:35>> Yeah, I was about to say, what happens?
- 40:37>> Yeah, exactly right. And so, when you
- 40:39see I have those for the SGD. So, you're
- 40:41exactly right. This is what happens when
- 40:42we stock for things like SGD. What
- 40:44happens is you bounce around in a ball,
- 40:46a highdimensional ball that's kind of
- 40:48proportional to that that step size. And
- 40:50this has some profound implications
- 40:52because if you've taken a statistics
- 40:53course, you've been told that what you
- 40:55really care about is recovering the
- 40:57parameters. If you want to be technical
- 40:59about it, you really care about getting
- 41:00the right theta zeros. Okay, this is
- 41:03called a recovery guarantee. If I see
- 41:05the data, then I can prove that I get
- 41:07the right theta zeros. Machine learning
- 41:08people do not care about this.
- 41:10Absolutely not. And they're right not to
- 41:12if you want to do what we were doing,
- 41:13right? Not if you're doing stats.
- 41:14[snorts] And what that means is there
- 41:16could be many many thetas as long as
- 41:19we're close good enough. And so a lot of
- 41:21modern machine learning is like running
- 41:23these models and like when do you stop
- 41:24like in statistics we would do these
- 41:26very you know fancy tests when the
- 41:27gradient goes below this and all the
- 41:28rest we can prove we're within blah blah
- 41:30blah blah blah of the of theta not the
- 41:32theta star the optimal theta star in AI
- 41:35we're like we ran out of compute credits
- 41:37we'll stop feels good let's go to the
- 41:39next model and that's crazy enough how
- 41:42it works. Okay. All right. Okay. So,
- 41:45when you set these at home, I just want
- 41:46you're going to have like some
- 41:47assignments where you play with some
- 41:48step sizes. You'll see if it's
- 41:50converging slowly, try and set better
- 41:52step sizes. In a couple weeks, I'll
- 41:54teach you something called hyperband,
- 41:55which was one of my personal favorite
- 41:56algorithm was written by friends of how
- 41:58to adaptively set these step sizes in a
- 42:00nice way. And uses something called
- 42:01bandits, which we need a lot of
- 42:03notation. And once you once you've
- 42:05gotten familiar with the notation, it's
- 42:06really straightforward. It's Kevin's
- 42:08algorithm. It's cool. All right. All
- 42:10right. So this just says what I said and
- 42:12doesn't have all the weird rants in it.
- 42:13All right, good enough.
- 42:16Okay, now as I mentioned, we want to get
- 42:19even simpler because that n this data
- 42:22set could be very large. And in fact,
- 42:24machine learning is anchored towards
- 42:26more data, not better estimation. Just
- 42:28that's not a formal statement, but just
- 42:30intuitively what we're after is we're
- 42:32after these models that we're going to
- 42:33crank tons and tons and tons of data
- 42:35through them. We think that's a way to
- 42:37improve them by making them bigger and
- 42:39consume more information than
- 42:40necessarily estimating these parameters
- 42:42very finely. And so we'll tell you the
- 42:44techniques that we used to use to
- 42:46estimate those parameters down to like,
- 42:48you know, 10 significant digits. But
- 42:50that's not the vibe. That's not what
- 42:51we're going for. We want to get bigger
- 42:53models that we're going to kind of fit
- 42:55less well than our predecessors. Okay?
- 42:57And that's what leads to all kinds of
- 42:59interesting behaviors that we'll tell
- 43:00you about later like scaling laws and
- 43:01things. Okay? So why does that matter?
- 43:04This is really an algorithm that I owe a
- 43:06lot to and that is actually widely used
- 43:08and it's old algorithm too. So this is
- 43:11the algorithm update rule written for
- 43:13one different component. Okay. We've
- 43:16looked at XJ here. It only depends on
- 43:18the J component. We look at the whole
- 43:20error term. Look at all the data. We
- 43:23update. All right.
- 43:26Okay. Now here I'm just showing that the
- 43:28vector we're writing in vector notation.
- 43:30This is again just why does the vector
- 43:32save us? Well, I just erased all the
- 43:34J's. You will see me very freely go back
- 43:36and forth between these. All right.
- 43:39Okay. This is what I was actually wanted
- 43:40to talk about. Okay. Now, consider our
- 43:43rules. So, I've said this multiple times
- 43:44already. Sorry, I got excited. This is a
- 43:46thing that I like. So, I know it's weird
- 43:47to get excited about this, but I cannot
- 43:49tell you. So, just as a small
- 43:50digression, I worked a lot of my life on
- 43:52the algorithm I'm about to show you.
- 43:54It's extremely simple. Okay, we did all
- 43:56kinds of fun things about this algorithm
- 43:57and how we scaled it up and how we built
- 43:59it. It's called stochcastic mini batch
- 44:01or stochastic gradient descent or in the
- 44:03old days was called incremental gradient
- 44:05and it goes back at least to like the
- 44:061950s Robins and Monroe. It's like a old
- 44:09classical algorithm. People rediscover
- 44:11it every 101 15 years. Okay, it is the
- 44:14workhorse. So how by show of hands, how
- 44:16many people have used PyTorch or
- 44:18anything? Okay, more people use PyTorch.
- 44:20That's dangerous but awesome. uh in case
- 44:22so if you do backrop in there then you
- 44:25have to have used mini batch like the
- 44:26entire system is is set up for you so
- 44:29that you're going to take a small batch
- 44:30of data you're not going to look at
- 44:31everything in your data set and you're
- 44:33going to feed through one you know five
- 44:35images of cats and dogs at the same time
- 44:37okay not all the images of cats and dogs
- 44:39that are you know on your phone or
- 44:40whatever right so that is the mini
- 44:42batching rule so how does that look
- 44:44mathematically
- 44:48well the idea is we're going to sample
- 44:53this thing.
- 44:56Mathematically, the way we'll think
- 44:57about it is we're going to take a random
- 44:59set. Now, honestly, when you run, this
- 45:01is a thing that I again spend a lot of
- 45:02my life, you know, in weird ways. One
- 45:04slice of my like I was doing it all day
- 45:06long, but it was like a piece of things
- 45:07I was working on. I worked on
- 45:09understanding what happens if you
- 45:11randomly sample versus if you just
- 45:13randomly sort your data and go through
- 45:15it once. Okay,
- 45:18turns out they're kind of okay. And that
- 45:19latter one is basically what we do. Took
- 45:21a lot of math to prove that those things
- 45:23are about the same for a variety of
- 45:24reasons. When you get to the end of the
- 45:25data set, it gets nasty. The first part
- 45:27of the data set, they're obviously kind
- 45:28of close. Okay, but this is what SGD is.
- 45:31Sample from your data set.
- 45:33Don't pick those samples in a weird way.
- 45:35Now, what could go wrong if you pick
- 45:37those samples in a weird way? Just
- 45:38intuitively.
- 45:40If for example, let's say that I showed
- 45:42you all the cats first, then I showed
- 45:46you all the dogs
- 45:48and you put them in those little
- 45:50batches. What do you think would happen
- 45:51to the underlying model?
- 45:53>> It would first just keep predicting cats
- 45:55and it would stop predicting cats for
- 45:56everything.
- 45:57>> Exactly right. So, it would learn some
- 45:58trivial surface that was like only
- 45:59predicting the cats. You got it exactly
- 46:01right. And it would get really confident
- 46:03about cats, but it wouldn't know
- 46:04anything about dogs. Then it would see
- 46:05the dogs and it would race to the other
- 46:07side. Okay. So, why do I tell you this
- 46:09intuitively? What do you want in that
- 46:10batch? You want it to be kind of a
- 46:12sample of the population. This is the
- 46:14second statistical assumption we're
- 46:15making. The first one was my training
- 46:17set reflects the real world. My second
- 46:19one is my mini batches again reflect my
- 46:22overall data set.
- 46:24That one as I start to get more
- 46:26complicated things that becomes tricky
- 46:27to guarantee. Please
- 46:30>> do the actual uh form the rule is h of
- 46:33data using the previous.
- 46:37>> Exactly right. Yeah. Wonderful. Yeah. So
- 46:39here this this h of theta should
- 46:41actually be the theta t. Sorry if that's
- 46:42not clear. I'll just write it in here.
- 46:44It's a wonderful observation. So the
- 46:46observation is which theta is talking
- 46:48about here? And it's the theta from the
- 46:49last step. That's just a notational bug.
- 46:51You got it exactly right.
- 46:54Other questions? Oh, please. Did you
- 46:57have a question?
- 46:59>> So is it just these?
- 47:07>> Yeah. So the question is is if you
- 47:09sample with versus without replacement,
- 47:11how does that how does that change the
- 47:12situation? And it turns out that in
- 47:15basically the the moral of the story is
- 47:17we have increasing theory and empirical
- 47:19evidence that it doesn't matter if you
- 47:20do the with replacement versus without
- 47:22replacement, but without replacement is
- 47:24in fact a lot easier to implement
- 47:26because what you can do is you can hash
- 47:27or you can sort on a random key and do
- 47:29it once and then plow through all of
- 47:31your data. In fact, by default, PyTorch
- 47:33will often not sample from the data and
- 47:35just take the batches in the order that
- 47:37you get it. So that once the data set is
- 47:39very very large, the dist the change in
- 47:41this distribution should not be very
- 47:43big. Right? If I have a billion points
- 47:44and I shuffle them versus picking them
- 47:46up. Now, one other thing that people
- 47:48have observed, you don't have to know
- 47:49any of this, so please like don't worry
- 47:51about it, but one other thing that
- 47:52people have observed and this came from
- 47:54a paper that we wrote like 10 years ago
- 47:56uh about this. [snorts]
- 47:58It turns out and and lots of people
- 47:59observe this, not just us. If you do the
- 48:02random reshuffle, the model converges
- 48:05faster. And this is something that you
- 48:07know, if you think about it intuitively,
- 48:09by show of hands, who knows what the
- 48:10coupon collector problem is? Oh, few.
- 48:13All right. PieTorch, but not coupon
- 48:14collector. Interesting times. Anyway,
- 48:16that used to be like CS Cannon. Like,
- 48:18you had to know that. But, you know, I
- 48:20was just This has nothing to do with
- 48:21what I'm talking about. How many people
- 48:22know what a DFA is? An Automa.
- 48:25Oh, that's good. Good for us. All right.
- 48:26We're still doing it. All right. Anyway,
- 48:27back to this. So the point is is when
- 48:29you did that shuffle, you had a kind of
- 48:31a coupon collector phenomenon. What
- 48:32happens in a coupon collector is let's
- 48:34say that I give you 10 coupons that you
- 48:36want to collect and you run around and
- 48:37you're randomly sampling. So you get a
- 48:39random sample every time. It turns out
- 48:41that all of your variance, how long it
- 48:43takes you is how long it takes to get
- 48:44the last coupon. Why? Because if I'm
- 48:47sampling from all 10, right, I'm getting
- 48:49the first nine again and again, 90% odds
- 48:51when I get to the end. And so I only
- 48:53have a 10% odds of getting it when I get
- 48:54to that final one. Turns out that skews
- 48:56all the variance to the end. So back to
- 48:58these models, if you're looking for
- 49:00something that's relatively rare in your
- 49:02data set, you're just not going to
- 49:03encounter it. You're going to encounter
- 49:05the mean a little bit too often and
- 49:07you're not going to encounter the thing
- 49:08you want. So the f folklore is that and
- 49:10there's a bunch of math that says under
- 49:12situations the story I just told you is
- 49:14true, but not all situations. There are
- 49:15obvious counter examples. Okay, awesome.
- 49:18These are wonderful questions. Super
- 49:19happy to talk about this stuff. Please
- 49:22>> batch size. Sometimes we just say this
- 49:25is how much GPU I have. This is what my
- 49:27batch size will be.
- 49:28>> Is there some scientific way to say I
- 49:30should have a batch size of one or two
- 49:32or 12?
- 49:33>> Oh, wonderful question. So the question
- 49:34uh is about batch size and how do you
- 49:36pick it? So I'm going to tell you
- 49:38something. So a couple things. So
- 49:40there's a couple of heruristics that
- 49:41people have but I wanted to tell you as
- 49:43a personal matter batch size was
- 49:45actually one of the things that broke me
- 49:46intellectually uh and a while ago. Okay.
- 49:49So there was a while when I would say
- 49:51about 15 years ago when what we thought
- 49:53was that smaller batch sizes were better
- 49:55because they were exploring the data
- 49:56more and you can tell yourself a story
- 49:57about why this is better and from an
- 49:59optimization perspective you can prove
- 50:01that in many situations a tiny batch
- 50:03size is just as good as a big batch size
- 50:05and so you want to run in these tiny
- 50:07batch sizes. Then there was a paper that
- 50:09came out of when it was Facebook the
- 50:11Facebook labs that basically showed this
- 50:13really interesting thing that was done
- 50:14by friends and they they sent me as a
- 50:15preprint that said we're getting better
- 50:17generalization. I'll explain what I mean
- 50:18in a second by using larger batches.
- 50:21Now, this was very strange. Their loss
- 50:24was worse. Their training set was worse,
- 50:26but they were generalizing better to the
- 50:28real world. Now, as a mathematical or an
- 50:30optimization person, at first I was
- 50:32like, this is heresy. As I just told
- 50:34you, we assumed that the training sets
- 50:36are have the statistical relationship.
- 50:38So, doing better on the training set
- 50:39should always mean doing better on the
- 50:40test set. But in this paper, they showed
- 50:42that wasn't the case all the time. And
- 50:44that's one of the things the reason I
- 50:46tell you this story is first it was like
- 50:47personally like you know as I was
- 50:49working on this it changed the
- 50:50mathematics of it the second piece of it
- 50:52as we went through it was
- 50:56basically the batch size we don't fully
- 50:57understand and so the convention now we
- 51:00have this idea that larger batches are
- 51:01better because they have lower variance
- 51:03they're a better estimate and so the
- 51:05practice that you just mentioned how do
- 51:07you pitch the batch size how big is your
- 51:09you know GPU memory right like you have
- 51:11this much HPM you have this much batch
- 51:13size Like that's basically what people
- 51:15do for systems reasons. But the theory
- 51:17of like how and why this works and the
- 51:19fact that optimization is a leaky
- 51:21abstraction is really interesting. And
- 51:23it should have been more interesting to
- 51:24me. I didn't recognize this when I first
- 51:26saw it, but it was a really important
- 51:27kind of theoretical moment for me
- 51:29because what it re revealed to me is
- 51:31remember when I was telling you machine
- 51:32learning is not stats, right? I think
- 51:34we're still co-listed with stats. No
- 51:35offense to statistitians. I love you
- 51:36very much. Some of my best friends are
- 51:38statistitians. But machine learning is
- 51:40not stats. And one of the things is is
- 51:42statistics as I mentioned is very
- 51:44interested in when the optimization
- 51:45problem recovers the right answer. And
- 51:47as I told you machine learning people
- 51:48are not interested in that. And so the
- 51:50fact that you have a model that doesn't
- 51:51get the right answer with this batch
- 51:53size change you know right answer the
- 51:55lower loss the minimum loss but
- 51:56generalizes better. We're going to throw
- 51:58away all the theory and try and figure
- 51:59out what's happening there. And so now
- 52:01your your heristic is the right one. Set
- 52:03the batch size according to what you can
- 52:05do. And there's a little bit of tuning
- 52:07that people do underneath the covers
- 52:08about how they set their batches. It's a
- 52:10little bit folklore if I'm honest. I can
- 52:12tell you the tricks. You know, I have
- 52:14mine, other people have theirs. I'm
- 52:16[snorts] not super sure. Great question.
- 52:18Please.
- 52:19>> Why would sac potentially be better?
- 52:22Because
- 52:22>> Oh, great question.
- 52:23>> Um cuz in the example given like cats
- 52:26and dogs like wouldn't you end up with
- 52:28if you have smaller batches have much
- 52:30more extreme nonproportional?
- 52:34>> Yeah. Yeah. Wonderful question. Yeah. So
- 52:35the question is why would why did you
- 52:37fools ever think that low small batch
- 52:39sizes were going to work? Uh you it was
- 52:41obvious the whole time you were wrong
- 52:42probably but the reason you would think
- 52:44a small batch size would help is a bit a
- 52:46little bit of a calculation which is
- 52:47actually in the slides in the appendix
- 52:49which shows that imagine the situation
- 52:51where I have a data set that contains a
- 52:53lot of redundancy. Okay, if it contains
- 52:55a lot of redundancy and I look at a
- 52:57whole batch then I'm basically seeing
- 53:00the same example. You know it's just
- 53:01like 90 pictures of the same cat again
- 53:03and again and again, right? And so I've
- 53:05wasted all those steps. I could have
- 53:07been using those 90 cats to refine the
- 53:09model 90 times. So really what it's a
- 53:11trade-off is how statistically
- 53:13meaningful is the the data sample that
- 53:15you're seeing versus how big of a step
- 53:17should you take. So it's not an obvious
- 53:19trade-off either direction. The heristic
- 53:21of why smaller was better was that you
- 53:23were able to polish the model that more
- 53:25parameter updates were better than what
- 53:26you were doing. That's probably true if
- 53:28you have small classes or a relatively,
- 53:31you know, compact domain where you have
- 53:32like, you know, two classes and they're
- 53:34kind of well established and separated.
- 53:36In those situations, this redundancy
- 53:38argument is actually provable. SGD with
- 53:40small batches will do better. And so
- 53:42that like gave us maybe false confidence
- 53:44that that was explaining the whole
- 53:46world. When we move to more interesting
- 53:48examples, larger batch sizes started to
- 53:50work better and better and better. The
- 53:52other thing that happened is the way we
- 53:53treat the batches. As you'll see in the
- 53:55second half of the course, we also want
- 53:57to embed information in them. So when we
- 53:59get to other concerns about how you
- 54:01supervise your data, it's interesting to
- 54:02have a mix of of say positive and
- 54:04negatives. Exactly as you said, I want
- 54:06every batch to have a couple cats and a
- 54:08couple dogs. And ideally, I'd like them
- 54:10to be kind of close together
- 54:11intuitively. I don't want a dog, you
- 54:13know, like a Great Dane and some little
- 54:14tiny cat. I kind of want the most
- 54:16cat-like looking dog I can get to help
- 54:18me separate them. I'll learn the most
- 54:20interesting information. We'll talk
- 54:21about that in that course. you're
- 54:23getting you guys have got this exactly
- 54:24right. What are the key issues and
- 54:26underneath a batch please?
- 54:29>> It was selected specifically for like
- 54:32interesting features that we want.
- 54:34>> Awesome. Yeah, great question. So, in
- 54:36the way that we'll teach the the
- 54:38traditional stocastic gradient descent,
- 54:40it is completely random sampling. With
- 54:42replacement sampling, you take it and
- 54:43you pull it up. What I'm trying to
- 54:45emphasize from some of this discussion
- 54:46is in practice, it's actually not really
- 54:48done that way. Sometimes it's done with
- 54:50this sort that I mentioned where you
- 54:51sort it in a random order. We talked
- 54:52about how to cut down invariance. And
- 54:54then what I was just saying is that some
- 54:55objective functions, some loss functions
- 54:57actually prefer that you have kind of
- 54:59near miss examples to each other which
- 55:01is also quite intuitive. And so people
- 55:03do engineer their batches. And then the
- 55:05last piece is you do pick the batches in
- 55:07modern applications based on how fast
- 55:09they run. as was talking about in the
- 55:10HBM case, if you haven't optimized the
- 55:13GPU, how much if you can store your
- 55:15whole model or your whole computation in
- 55:16the HBM, the memory that's on the chip,
- 55:18it's just dramatically faster. And so
- 55:20you optimize for systems concerns. So
- 55:22this very small tweak here that looks
- 55:25like, you know, a oneline change of
- 55:26going from N to B. Basically,
- 55:29that change is actually in engineering
- 55:32in practice, I'm trying to say is quite
- 55:34rich. For the point of view of the
- 55:35course, I think you have to know nothing
- 55:36of what I just described. By the way,
- 55:37you just need to say, "Oh, yeah, you
- 55:38random sample. That's good enough." But
- 55:40I wanted you to understand like this
- 55:41stuff is actually fairly interesting and
- 55:43fairly accessible. Like there's someone
- 55:45right now at one of the frontier labs
- 55:46who's tuning the batch size. As weird as
- 55:48that is to say, and they're getting paid
- 55:49a lot of money, which is great. I hope
- 55:50they're one of my former students.
- 55:51Anyway, all right. So, this is the this
- 55:55is the definition right here. This is we
- 55:56take the batch and we average it. Okay.
- 55:58And I'm just saying that that rule to
- 56:00select the B random sampling is good
- 56:02enough. You'll do that for most of the
- 56:04course. But it is actually like making a
- 56:07statistical assumption. And if you
- 56:08really care about what's going on there,
- 56:10it's just an interesting thing. And
- 56:12there's there's there's stuff to know.
- 56:13If you're a curious person, there's
- 56:14stuff to know. Wonderful questions.
- 56:17These are really great.
- 56:20All right.
- 56:24All right. So hopefully this is clear. I
- 56:26said this I should have I should have
- 56:27gone to this slide earlier. Sorry. Um
- 56:30here I would point out one thing which
- 56:32is you're going to sum over the batch
- 56:33that mini batch right so B is usually
- 56:36less than the full data set you compute
- 56:38the errors on each one with your current
- 56:40theta which was pointed out this should
- 56:41be theta t you then just do the updates
- 56:44the point is you don't have to look at
- 56:45the rest of the data set you just look
- 56:47at the model and the data set the your
- 56:49data samples from the batch and then
- 56:51you're going to take this step size and
- 56:52I'm highlighting here that it's an alpha
- 56:54B and that's because it's a it's
- 56:56potentially a different step size okay
- 56:58it's you're not going to use the game
- 56:59step size that you would use on the full
- 57:01data set. And in fact, the step size
- 57:03kind of mysteriously depends on other
- 57:05quantities that are there. We'll talk
- 57:06about good rules of thumb. In fact, one
- 57:08of the major pieces of technology that
- 57:10you probably use and so many of you have
- 57:12used step size are basically what are
- 57:14called adaptive optimizers that pick the
- 57:16step size for you within range. They see
- 57:18that you're having some of these
- 57:19problems. Adam, for example, if you know
- 57:22what Adam is or Adam W. Okay, so or or
- 57:25Adagrad, which was developed by our own
- 57:27John Duchi. Um those are those are the
- 57:29kinds of of things that go on under the
- 57:31covers. But for now when we're studying
- 57:32it in its purest setting like if you
- 57:34wanted to do it with like code and write
- 57:36it out you would have to pick that
- 57:37alphab and that's what goes on. Any
- 57:40other questions there?
- 57:43Oh please
- 57:46change like a batch.
- 57:48>> Oh yeah sorry that this wasn't clear.
- 57:50Yeah so what the the operation that
- 57:51you're doing here sorry this is unclear
- 57:53is you're going to pick a random batch
- 57:55on each iteration. So at every time step
- 57:57t you pick a new random batch. Okay. You
- 58:00don't want to feed the same dogs and
- 58:01cats through every single time or houses
- 58:03through every time because you'll just
- 58:04learn a model of the data that you've
- 58:06seen. So it is a random sampling
- 58:07procedure. Yeah. Wonderful question and
- 58:09clarification.
- 58:10Please
- 58:13the coefficient gets absorbed into the
- 58:15step size. Does that imply that larger
- 58:17batches typically would have smaller
- 58:19step sizes?
- 58:20>> Yeah, that's a wonderful question. So
- 58:21how the batch size scales, you would
- 58:23indeed expect that as you ramp up the
- 58:25batch size, right? When you go all the
- 58:27way from one to the limit, you're going
- 58:29to get to a batch size of one overn. You
- 58:31very rarely use 1 overn as your SGD
- 58:33batch size. It turns out for like
- 58:36folklore reasons, that lots of people
- 58:38write their models, not linear models,
- 58:39but write their models so that there's
- 58:41the same kind of normalization in batch
- 58:43size across all of them. And so that is
- 58:45like a thing that you do. So if you've
- 58:47if you've played with deep learning
- 58:48things and used like layer norms or
- 58:49sandwich norms or other things, they're
- 58:51trying to get you in the right regime
- 58:52where like the step sizes are kind of
- 58:54all what's called fancy word is
- 58:55isotropic, but kind of all the
- 58:57dimensions are the same. So what does
- 58:58that mean for alpha? That means you
- 59:00would intuitively expect that if the if
- 59:03the model is taking a little bit of
- 59:05information, you don't trust it very
- 59:07much and alpha would be lower almost
- 59:08exactly as you said.
- 59:10>> Yeah. Oh, go ahead. Um,
- 59:13do you have you done experiments where
- 59:15you've changed the BM size throughout
- 59:17the gradient descent? Is it
- 59:19>> Oh, yeah. Wonderful question. So, we
- 59:21haven't talked about this. Again, you
- 59:22don't need to know this, but I'm super
- 59:23happy to tell you. So, it turns out
- 59:25there's actually a wonderful study
- 59:26written by this guy at the Navy named
- 59:27Leslie who wrote this study about how do
- 59:29you do various what are called
- 59:30scheduling rules or step size rules. So,
- 59:32there's two great papers if you ever
- 59:33wonder. There's one where he basically
- 59:35did what's actually quite widely used.
- 59:36It's called cosign scaling where you
- 59:38actually make the batch size go up and
- 59:39down as you run. The idea being that you
- 59:41zoom into like a local minima and then
- 59:43it kicks you out. And he showed that
- 59:44this cosign scaling was actually one of
- 59:46the best at the time for image models.
- 59:49There's another paper which I love which
- 59:50is like from the '9s which is the
- 59:52Onstriker paper which shows that a bunch
- 59:54of rules that people are using in
- 59:55gradient descent and stochastic gradient
- 59:57descent are actually all the same
- 59:58so-called linear and exponential back
- 1:00:00off and all the rest. [snorts] Now, I
- 1:00:02would doubt that most people here in
- 1:00:04production are tuning their schedules uh
- 1:00:07unless they're doing something that like
- 1:00:09you well for for a product you would I
- 1:00:11guess you would say, but like
- 1:00:12researchers are kind of just putting in
- 1:00:14the numbers and letting the defaults
- 1:00:15work most of the time. Yeah, wonderful
- 1:00:18questions. Please
- 1:00:20>> um
- 1:00:22like
- 1:00:23you would not want to like have repeated
- 1:00:26sampling of uh things. So it's like in
- 1:00:29like the selection of batches would you
- 1:00:31not like?
- 1:00:33>> Awesome. So the thing is is if you look
- 1:00:35at this a wonderful question. So the
- 1:00:37question is if you sample your data you
- 1:00:40kind of want to see all of your data.
- 1:00:42You don't want to have repetition in
- 1:00:43there. And this is exactly gets back to
- 1:00:45the difference between with replacement
- 1:00:46sampling where you would see as I was
- 1:00:48talking about those coupons many many
- 1:00:50times versus without replacement
- 1:00:52sampling where you shuffle the data set.
- 1:00:54And when you shuffle that data set,
- 1:00:55you're guaranteed that you're only going
- 1:00:56to see every data set exactly once. And
- 1:00:58so the belief is that that would help
- 1:01:00you exactly for the reason you said that
- 1:01:02you're going to see many of the examples
- 1:01:04again and again. Okay? Now, that assumes
- 1:01:07that all the examples are equally
- 1:01:08informative, and that's not always true.
- 1:01:10Sometimes there's a core hardcore set
- 1:01:13that you wish you could see again and
- 1:01:15again. And if you look at modern
- 1:01:16training for like LLMs that are in the
- 1:01:18wild, large language models that are in
- 1:01:20the wild, you'll see that actually there
- 1:01:21are places where people actually do
- 1:01:23repetition on things like code is very
- 1:01:25popular to do multiple times because of
- 1:01:27the belief is you put in some structure.
- 1:01:28Okay, there's a great paper about this
- 1:01:30if you're interested about data mixing.
- 1:01:32Please send a note or post to Ed. I'm
- 1:01:33happy to post something about about what
- 1:01:35people know about how you mix your data
- 1:01:37for for foundation models. Not part of
- 1:01:39the course, but super happy to tell you.
- 1:01:41Yeah. Uh you said the cosign uh schedule
- 1:01:44paper they did it with image models. Is
- 1:01:46there a reason why like uh learning rate
- 1:01:49things are different from image models
- 1:01:51and language
- 1:01:51>> models?
- 1:01:53>> Yeah. So the question is why does the
- 1:01:55why would you care that it's from image
- 1:01:56models versus text models? What would
- 1:01:58change there? There's two things that
- 1:02:00change. One are the architectures. At
- 1:02:02the time the architectures were
- 1:02:03different. We're talking about neural
- 1:02:04net architectures. That has actually
- 1:02:06changed over time. Now we've converged
- 1:02:08and built a bunch of technology where
- 1:02:09you can use kind of transformer stacks
- 1:02:11for both of them. The other thing is
- 1:02:12that the distributions may actually be
- 1:02:14quite different. So if you think about
- 1:02:15images, they have a set of natural
- 1:02:17distributions that are in the world that
- 1:02:19like has some smooth variations, right?
- 1:02:21Like I don't know about you, but like
- 1:02:22you know I take pictures of my kids.
- 1:02:24It's like in burst mode. I got a
- 1:02:25thousand pictures of, you know, like one
- 1:02:26of my daughters dancing, right? That
- 1:02:29thing they're all little tiny variations
- 1:02:30of each other. That distribution is kind
- 1:02:32of compact. I don't tend to write the
- 1:02:34same sentence a thousand times like a
- 1:02:36crazy person. maybe I I have over like
- 1:02:37you know longitudinally in my life but
- 1:02:39like it's not kind of compact in that
- 1:02:41way and so those differences in
- 1:02:43distributions they may have an effect on
- 1:02:45the learning rates and so you can't they
- 1:02:48don't transfer well like if you had
- 1:02:49clean theory you would know what goes
- 1:02:51from A to B but we don't know what goes
- 1:02:53from A to B in different situations so
- 1:02:54that caveat turned out to be important
- 1:02:56it caught on quite a bit in images and
- 1:02:58some of those ideas were adapted but
- 1:03:00those aren't the ones that we use today
- 1:03:02in uh engineering text models
- 1:03:06the production model at this point.
- 1:03:08Would you like how would you feel if the
- 1:03:11bad side changing like right now?
- 1:03:15>> I would pick them. Yeah. So, I would
- 1:03:17pick the way I would do this. Oh, sorry.
- 1:03:18The way I would do this is oops. Yeah.
- 1:03:21So, this this covers many of our points.
- 1:03:23So, the way I would do this honestly is
- 1:03:25I would try and figure out what makes
- 1:03:26the GPU most efficient. So if you look
- 1:03:28at the cost of these models like one
- 1:03:30thing that changed since say like I
- 1:03:31first started teaching this course till
- 1:03:33now is the amount of what you would say
- 1:03:34is capex the amount of capital
- 1:03:36expenditure to build these models we're
- 1:03:38talking about building gigawatt
- 1:03:39facilities this was unimaginable
- 1:03:41gigawatt is a lot of compute okay those
- 1:03:44gig multi- gigawatt facilities you have
- 1:03:46to keep them at high utilization so if
- 1:03:48you ask me what I care about I would say
- 1:03:50I care quite a bit about making sure the
- 1:03:52utilization of the model is high that I
- 1:03:54get in more steps the statistical
- 1:03:56concerns that we'll talk out they don't
- 1:03:58take a complete back seat but if you're
- 1:04:00doing it like if you see how people
- 1:04:02report their training they report MFU
- 1:04:04model flops you know utilization they
- 1:04:06care about how much they're getting out
- 1:04:07of those GPUs the statistical properties
- 1:04:09are sometimes harder to pin down and
- 1:04:11that's one thing by the way that is
- 1:04:12actually quite a blessing like one thing
- 1:04:14that I think is very underappreciated
- 1:04:16about machine learning you can get this
- 1:04:18so we're going to teach you these
- 1:04:19different building blocks but sometimes
- 1:04:22what happens in this field is that
- 1:04:23people think those building blocks in
- 1:04:24their head occupy like entirely
- 1:04:26different spaces
- 1:04:27One of the beautiful things about
- 1:04:28machine learning which is kind of
- 1:04:30amazing is many of these things work
- 1:04:33like one experiment like I'm going to
- 1:04:34tell you in the next lecture how you
- 1:04:35should use a different loss function for
- 1:04:37classification but it kind of works if
- 1:04:40you use the loss function from this even
- 1:04:42though it kind of makes no sense and
- 1:04:43I'll show you why it shouldn't make much
- 1:04:45sense but it will still kind of work
- 1:04:46there's a robustness to the underlying
- 1:04:48elements that is pretty surprising so
- 1:04:50it's very easy to fixate on the
- 1:04:52statistical definitions of what's going
- 1:04:53on and we will tell you the most
- 1:04:55important but I want to highlight that
- 1:04:57sometimes tweaking the batch size or
- 1:04:59tweaking these rates doesn't matter
- 1:05:00nearly as much as making your model get
- 1:05:02bigger or you know running it for longer
- 1:05:04and so those principal components are
- 1:05:06the ones that have driven the progress
- 1:05:08over the last couple of years that may
- 1:05:09saturate at some point please
- 1:05:12>> you like divided by like the one over
- 1:05:15like
- 1:05:17>> right yeah
- 1:05:18>> yeah that's just packed into the step
- 1:05:20size
- 1:05:21>> yeah yeah we just normalize it into the
- 1:05:22alpha cuz it's like it's a number that
- 1:05:24we don't know haven't interpreted anyway
- 1:05:25so we might as well just multiply it by
- 1:05:26something.
- 1:05:27>> Okay.
- 1:05:27>> Yeah. Yeah. Awesome.
- 1:05:30>> Great questions.
- 1:05:33>> Okay. All right. So, I I promised this
- 1:05:35graph earlier. It's not that great. If
- 1:05:37you were really like on the edge of your
- 1:05:38seat waiting for it, apologies. Uh but
- 1:05:40this is how it works. So, these are the
- 1:05:42batch versus uh gradient descent on a
- 1:05:44very smooth problem. Okay. So, a smooth
- 1:05:47problem, what is it? How do I know this
- 1:05:48problem is smooth by looking at it?
- 1:05:50Well, these things here are the
- 1:05:52isocience for the function. That's an
- 1:05:54equal loss. Okay. Okay, that's where the
- 1:05:55loss is all one value. This is a
- 1:05:57quadratic that we've been looking at.
- 1:05:59Quadratics, if you remember, produce
- 1:06:01ellipses as isoclines. Is that familiar?
- 1:06:04Show of hands familiar? Okay, I'll
- 1:06:06remember next time. All right. So,
- 1:06:08that's what they look like. We'll come
- 1:06:09back to to that in more more depth. Not
- 1:06:11critical now. So, this function is very
- 1:06:13smooth and bullshaped, right? Because
- 1:06:14this this says like all of the losses
- 1:06:17here and all the losses here are the
- 1:06:18same. These guys are all each rung is
- 1:06:20kind of the same. So, it's this nice
- 1:06:21bullshaped function. Gradient descent
- 1:06:24goes down and comes right to the optimal
- 1:06:26value which we've put right here in the
- 1:06:27middle. SGD as we talked about makes
- 1:06:31these little wiggles. Sometimes it's
- 1:06:32going the right direction, sometimes
- 1:06:33going the wrong direction. It bounces
- 1:06:35around and if alpha is set too large, it
- 1:06:38will bounce around in a ball. Okay? And
- 1:06:40it won't get close to the optimal. So
- 1:06:42one trick that people do, by the way, is
- 1:06:44they what do something called averaging
- 1:06:45the iterates. They may average over a
- 1:06:47trajectory of those bounces to kind of
- 1:06:49simulate a larger batch size. But
- 1:06:51hopefully intuitively this makes sense.
- 1:06:53The reason I highlight this for you is
- 1:06:55when you're training a model, you will
- 1:06:56observe this behavior. You will see the
- 1:06:58loss start to go and bounce around and
- 1:07:00you'll realize maybe I said alpha too
- 1:07:01high. Right? So those are the kinds of
- 1:07:03things that you'll see. Now in machine
- 1:07:06learning in the second half of the
- 1:07:07course, the functions will not be so
- 1:07:09nice. They will not be these nice
- 1:07:11bullshaped functions, these convex
- 1:07:13functions. They're going to be nasty.
- 1:07:14And when they're nasty, then the fact
- 1:07:16this is relying on the fact that it's
- 1:07:18quite smooth to go fast, to get rapidly
- 1:07:20down to the optimal. This is not relying
- 1:07:22on this. This is just kind of drunkenly
- 1:07:24stumbling its way to the loss function.
- 1:07:27And this one we prefer,
- 1:07:29as weird as it is, but for all the
- 1:07:31reasons I outlined.
- 1:07:33Okay, cool.
- 1:07:36All right.
- 1:07:38[clears throat] Okay. So just to make
- 1:07:40sure that we I got across what I wanted
- 1:07:42to get across. Our goal was to optimize
- 1:07:43a loss function to find a good
- 1:07:45predictor. We did that by minimizing
- 1:07:47this loss. We talked a lot I talked a
- 1:07:49lot about what it means to kind of that
- 1:07:51assume about your data that it's
- 1:07:52representative of the world. We'll talk
- 1:07:54about that even more later. We learned
- 1:07:56about an algorithm which hopefully
- 1:07:58looked relatively simple to you. You
- 1:07:59guys all the folks here knew about
- 1:08:01derivatives and computing gradients and
- 1:08:03partial derivatives. Awesome. And we did
- 1:08:05this kind of weird simple thing where we
- 1:08:07started looking at not our entire data
- 1:08:08but a batch of data, a subset of data.
- 1:08:10And when we did that, that somehow
- 1:08:12unlocked a bunch of runtime performance
- 1:08:14and I claim like these really large
- 1:08:16models. Okay, that algorithm was called
- 1:08:19stochastic gradient descent. We talked
- 1:08:20about all the ways to select batches.
- 1:08:22The most important one is select them at
- 1:08:24random. Okay, we touched a little bit on
- 1:08:27these trade-offs of using the right
- 1:08:28batch size through our conversations
- 1:08:30about, you know, when it's too small and
- 1:08:31when it's too large. what are the
- 1:08:33considerations inside a batch? And
- 1:08:34that's to hopefully give you a feel when
- 1:08:36you start to actually play with and tune
- 1:08:37some of these models. That's different,
- 1:08:39by the way, than what you would do if
- 1:08:40you were training a traditional stats
- 1:08:42model where you were trying to get down
- 1:08:43to the optimal value. There's just
- 1:08:46different concerns and hopefully some of
- 1:08:47those concerns have been highlighted and
- 1:08:48will become more clear over the next
- 1:08:50couple lectures.
- 1:08:52Any questions on this before I move on?
- 1:08:56>> Please.
- 1:08:58How do you know when to stop or is it
- 1:08:59just like when you stop?
- 1:09:01>> Awesome. Great question. So the question
- 1:09:02is, as you're bouncing around near the
- 1:09:04optimum, how do you know where to stop?
- 1:09:05And unfortunately, you don't. You don't
- 1:09:07know if you're bouncing around like
- 1:09:08these models have these very weird
- 1:09:10behaviors. So one of the things that
- 1:09:12they have as a behavior sometimes, and
- 1:09:13if you watch and measure them, they're
- 1:09:15all of a sudden they're running, they're
- 1:09:16running, the loss is kind of doing
- 1:09:17something and all of a sudden a
- 1:09:18capability turns on. Okay? So it just
- 1:09:21gets a little bit better. The
- 1:09:22representation clicks in when you're
- 1:09:23training more sophisticated models. So
- 1:09:25you don't know if you're stuck in kind
- 1:09:26of one of those local areas or if you're
- 1:09:29going to get to something, you know, if
- 1:09:30you're really close to the answer. And
- 1:09:32so that is actually a very
- 1:09:33computationally difficult problem. For
- 1:09:36everything in the next couple of
- 1:09:37sections, it's going to turn out that
- 1:09:38the models, if you're familiar, are
- 1:09:40going to be convex underneath the
- 1:09:41covers. They're going to be generalized
- 1:09:42linear models. That's what we're going
- 1:09:43to talk about for two weeks. In those
- 1:09:45situations, we can actually tell how far
- 1:09:46we are from the optimum. I'm not going
- 1:09:48to beat you in the head about this too
- 1:09:50much because you should just take
- 1:09:51Steven's course. It's wonderful. I love
- 1:09:52Steven Boyd. [snorts]
- 1:09:54But also those won't really do us much
- 1:09:56good when we get to the wild world of
- 1:09:58you know neural nets and unsupervised
- 1:10:00and all the rest. And the truth is there
- 1:10:01we don't know. And so how do you run?
- 1:10:04You kind of run like if I watch how I do
- 1:10:05it or my students do it. You look at
- 1:10:07like weights and biases. You watch till
- 1:10:08the thing goes down. You're like I don't
- 1:10:09know if I train for another two days I
- 1:10:11can't do you know I can't go out with my
- 1:10:12friends. I'm going to cut it off now.
- 1:10:14Okay. And that is unfortunately how it
- 1:10:16works. And so you basically provision
- 1:10:18like we're doing a pretty big train
- 1:10:20right now. the way we're provisioning it
- 1:10:21is like well we have these GPUs for two
- 1:10:24months for this project that are donated
- 1:10:26some DNA model thing whatever we're
- 1:10:28going to train that right now and so it
- 1:10:30kind of just is an engineering
- 1:10:32consideration more than it's a
- 1:10:33principled one uh I'm not really
- 1:10:35embarrassed to say because the stuff
- 1:10:36we've been building is great I used to
- 1:10:37be embarrassed but nothing worked as my
- 1:10:39wife said you know we've been together
- 1:10:40forever she's like you know you've been
- 1:10:42building the same demo since I've known
- 1:10:43you since undergrad but now stuff works
- 1:10:45so it's okay so all right
- 1:10:49anyway so Yeah. So these squares are
- 1:10:52really special. Um I want to do this
- 1:10:54basically just for notation. We have a
- 1:10:55couple minutes left. So I want to just
- 1:10:57highlight this notation um about normal
- 1:11:00equations. Uh if you haven't seen this
- 1:11:02before um take a look, do a review. All
- 1:11:05right. So we're going to derive the
- 1:11:06normal equations.
- 1:11:08All right. So mainly this is to make
- 1:11:10sure you're comfortable with the
- 1:11:10notation if I'm honest, but also to make
- 1:11:14sure that you know kind of basic things
- 1:11:15about linear algebra. So if you don't
- 1:11:17know what rank is, you don't know what
- 1:11:18inverses are,
- 1:11:20then you should study up because we use
- 1:11:22them pretty freely. We use that
- 1:11:23terminology pretty freely. Okay. All
- 1:11:25right. So we have some loss function
- 1:11:28here. Uh oops. We have some loss
- 1:11:30function J. Here we have again our
- 1:11:32vector notation. This is actually now
- 1:11:34going to be these ns are going to be
- 1:11:36laid out by rows. So it's going to be n
- 1:11:38by d matrix. So it's going to have n
- 1:11:40rows where the examples are there. And
- 1:11:42then d + one are the feature size that
- 1:11:45we talked about. This is examples. This
- 1:11:46each one of our examples. And then we're
- 1:11:47going to have correspondingly an element
- 1:11:49of Rn, which is going to be this vector
- 1:11:51here, Y. You may hear me call X the
- 1:11:53design matrix. This is really old
- 1:11:55terminology, but it's one that we use.
- 1:11:57Um, and it just means the matrix of X
- 1:11:59after we've done the featurization. So,
- 1:12:01after it's been kind of loaded and
- 1:12:02pre-processed into this form. Sometimes
- 1:12:04we'll worry about properties of the
- 1:12:06design matrix. So, that's what that
- 1:12:08means. Okay.
- 1:12:12[snorts]
- 1:12:13All right. Now, if it's linear, see if
- 1:12:16you can get here.
- 1:12:18[snorts] Notice that J has this
- 1:12:19particularly lovely form.
- 1:12:22Okay. Now, if this is not obvious to you
- 1:12:24that it has it, this is a good signal
- 1:12:26that you should do some work on kind of
- 1:12:28your vector manipulation, matrix
- 1:12:31manipulation stuff. This is the inner
- 1:12:33product. Okay, that's all I've written
- 1:12:34here. So, L2 is important because it's
- 1:12:36the inner product. That's the other
- 1:12:38succinct way to say this. And so, this
- 1:12:40is computing squares. And if you are
- 1:12:42rusty about this, please just compute it
- 1:12:44by hand. It's not like I'm it's not some
- 1:12:45deep mysterious thing. I'm just saying
- 1:12:48this thing here, oops, equals this thing
- 1:12:51here.
- 1:12:53All right. Okay, cool. Right. So now
- 1:12:57we've stacked the design.
- 1:12:59Now, if you don't remember this, please
- 1:13:02again review. But basically when we
- 1:13:04compute a gradient of a matrix, it means
- 1:13:07that we're computing derivatives
- 1:13:09according to a of each underlying
- 1:13:11element, right? And just stacking them
- 1:13:13in the corresponding type. So if you had
- 1:13:14a matrix coming in, you're going to have
- 1:13:16a matrix coming out. They're going to
- 1:13:17have all the same types. They're going
- 1:13:18to have the partial derivatives
- 1:13:20computed. This is the object.
- 1:13:24These are derivatives in the sense
- 1:13:25you're used to them in that if you want
- 1:13:28to find a minimum and it's a convex
- 1:13:29function and all those nice things, you
- 1:13:31set the gradient equal to zero. Why does
- 1:13:33this work intuitively? Well, intuitively
- 1:13:35what's happening is the function is
- 1:13:37going to be bullshaped. It's going to be
- 1:13:38strictly bullshaped. If there's a
- 1:13:40change, the function changes in one
- 1:13:42direction. You're not at the bottom.
- 1:13:43When you get to the bottom, you're done.
- 1:13:46Okay? This means there's no change in a
- 1:13:47local direction. Just like if you
- 1:13:49computed the derivative for a single 1D
- 1:13:51function. Okay.
- 1:13:54[clears throat]
- 1:13:55All right.
- 1:13:57So far so good. All right.
- 1:14:01I got to get better at that. Anyway,
- 1:14:04so from our previous derivation, this is
- 1:14:06just some arithmetic. We're going to
- 1:14:08multiply this thing out and normalize.
- 1:14:12The two does us a little bit of of
- 1:14:14value. We multiply out, we get the x
- 1:14:17theta, the xty,
- 1:14:19and we get this derivative. Okay. If
- 1:14:22that's unfamiliar to you, what's going
- 1:14:23on? We're multiplying these terms out.
- 1:14:25We get an xtx theta on each side. When
- 1:14:28we take the derivative with respect to
- 1:14:30theta, we get the term on one side and
- 1:14:31the other, but it's symmetric.
- 1:14:34We have the 1/2, we get this term. Okay?
- 1:14:37So, it's the this term and the foil.
- 1:14:39When we multiply the y's, we're not
- 1:14:41taking the derivative. They don't depend
- 1:14:42on theta. They get knocked out. When we
- 1:14:45do the cross terms, we get the cross
- 1:14:46term represented
- 1:14:48like so twice.
- 1:14:51And that's the derivative.
- 1:14:53We set this equal to zero. So that means
- 1:14:55that this thing is equal to zero. So how
- 1:14:57do we solve for that? We make these two
- 1:14:59terms equal to each other. Well, we have
- 1:15:01to take the inverse to get rid of the
- 1:15:03side xtx inverse* xty. And this is the
- 1:15:07le square solution.
- 1:15:11This is an optimal solution for for
- 1:15:13theta. Now I cheated in one part of this
- 1:15:15derivation.
- 1:15:19This is something I assumed about X when
- 1:15:20I did this.
- 1:15:22>> Assumed that Xtx is invertible.
- 1:15:24>> Exactly. Right. I assume that XTX is
- 1:15:25invertible. If for example I had
- 1:15:27relatively few examples in a huge set of
- 1:15:30of dimensions, this would be a
- 1:15:32nonsensical kind of thing.
- 1:15:36So here when I'm looking at it, I think
- 1:15:37that I have many many data points more
- 1:15:39than my parameters.
- 1:15:41Okay. And hopefully that that inverse
- 1:15:45exists.
- 1:15:46But what happens that if it doesn't
- 1:15:48exist,
- 1:15:50what is it defined up to the theta that
- 1:15:53I get?
- 1:15:58So yeah, so if this is unfamiliar to
- 1:16:00you, take a look. So there will be a
- 1:16:01null space, right? And in that null
- 1:16:03space, I won't prefer one solution over
- 1:16:04another. All of them would be valuable.
- 1:16:07So anything in the null space of xtx,
- 1:16:09right? I can add that to theta and it
- 1:16:10doesn't change the answer at all. So now
- 1:16:12there goes from being one single
- 1:16:14explanation, one single theta theta star
- 1:16:16to being an entire subspace of them.
- 1:16:18Effectively the entire null space of
- 1:16:20xdx. If you're seeing this and going,
- 1:16:22"Oh yeah, I remember that linear algebra
- 1:16:23sounds great. Do a little review." If
- 1:16:25you've never heard that and it sounds
- 1:16:26like I'm speaking a foreign language,
- 1:16:28please, please, please use the Friday
- 1:16:30lectures because after this I'm going to
- 1:16:31assume we know all this stuff. Okay,
- 1:16:33that's basically the thing there.
- 1:16:35[snorts] It will turn out, by the way,
- 1:16:37you could also worry that XTX here is
- 1:16:40also going to be it turns out positive
- 1:16:41semidefinite by show of hands. Who knows
- 1:16:43what PSD means? Positive semi-definite.
- 1:16:44Awesome. We're doing great. Okay,
- 1:16:46fantastic. Okay, so if this isn't
- 1:16:48familiar, I even wrote it down. Practice
- 1:16:50on Friday. And I just want to make sure
- 1:16:52that's clear. We're going to use these
- 1:16:53kind these concepts positive
- 1:16:54semi-definite. We're going to use these
- 1:16:56concepts like it's invertible. We're
- 1:16:58going to use them pretty freely. We're
- 1:16:59going to compute vector derivatives and
- 1:17:01matrix derivatives in this way. If it's
- 1:17:04unfamiliar, spend a little bit of time
- 1:17:06to make it familiar. It will make your
- 1:17:07life a little bit more easy. You know,
- 1:17:09it's it's much easier to compute this
- 1:17:11way, I think, than the other way, just
- 1:17:12by hand.
- 1:17:15All right,
- 1:17:17almost pretty good on time. Not so bad.
- 1:17:18Usually, I'm pretty bad about this. All
- 1:17:20right, so we saw lots of notation today.
- 1:17:22That was the intent of the lecture in
- 1:17:24some ways. I wanted to give you a little
- 1:17:25bit of machine learning stuff and lore,
- 1:17:27but I wanted to make sure that you
- 1:17:29understood all the pieces that go on. We
- 1:17:31learned a little bit about linear
- 1:17:32regression. We learned what the model
- 1:17:34is, right? So this is one of n models
- 1:17:36that you'll see in this class. We
- 1:17:38learned how to solve it and we learned
- 1:17:39that this this very basic algorithm to
- 1:17:42solve it called stochastic gradient
- 1:17:43descent. It's our workhorse and so we're
- 1:17:45going to see that algorithm come back
- 1:17:47again and again and again. Next time
- 1:17:50what we're going to learn about is not
- 1:17:51regression problems. We're going to
- 1:17:52learn about classification. So instead
- 1:17:54of looking at house prices, we're going
- 1:17:56to learn how to classify, you know, cats
- 1:17:58versus dogs or, you know, next word in
- 1:18:00the sentence and all those discrete
- 1:18:01functions. Thanks so much. Have a
- 1:18:03wonderful weekend.
About this transcript
This page contains the full transcript of YouTube transcript (cmNIMjPYdgM) , generated from the public captions YouTube serves with the video. The transcript has 17,221 words across 2,467 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.