Building makemore Part 2: MLP — Transcript
Full transcript
- 0:00Hi everyone.
- 0:02Today we are continuing our
- 0:03implementation of makemore.
- 0:05Now in the last lecture we implemented
- 0:06the bigram language model and we
- 0:08implemented both using counts and also
- 0:11using a super simple neural network that
- 0:13had a single linear layer.
- 0:15Now this is the
- 0:17Jupyter notebook that we built out last
- 0:19lecture.
- 0:20And we saw that the way we approached
- 0:21this is that we looked at only the
- 0:23single previous character and we
- 0:25predicted the distribution for the
- 0:26character that would go next in the
- 0:28sequence. And we did that by taking
- 0:30counts and normalizing them into
- 0:32probabilities
- 0:33so that each row here sums to one.
- 0:36Now this is all well and good if you
- 0:38only have one character of previous
- 0:40context.
- 0:41And this works and it's approachable.
- 0:43The problem with this model of course is
- 0:45that the predictions from this model are
- 0:48not very good because you only take one
- 0:50character of context. So the model
- 0:52didn't produce very name-like sounding
- 0:54things.
- 0:56Now the problem with this approach
- 0:57though is that if we are to take more
- 1:00context into account when predicting the
- 1:01next character in the sequence, things
- 1:03quickly blow up and this table the size
- 1:06of this table grows and in fact it grows
- 1:08exponentially with the length of the
- 1:10context.
- 1:11Because if we only take a single
- 1:12character at a time, that's 27
- 1:13possibilities of context.
- 1:16But if we take two characters in the
- 1:17past and try to predict the third one,
- 1:19suddenly the number of rows in this
- 1:21matrix, you can look at it that way,
- 1:23is 27 * 27. So there's 729 possibilities
- 1:27for what could have come in the context.
- 1:30If we take three characters as the
- 1:31context, suddenly we have
- 1:3420,000 possibilities of context.
- 1:37And so that's just way too many rows of
- 1:40this matrix. It's way too few counts
- 1:43for each possibility and the whole thing
- 1:45just kind of explodes and doesn't work
- 1:47very well.
- 1:49So that's why today we're going to move
- 1:50on to this bullet point here and we're
- 1:52going to implement a multi-layer
- 1:53perceptron model to predict the next uh
- 1:57in a sequence. And this modeling
- 1:59approach that we're going to adopt
- 2:00follows this paper Bengio et al. 2003.
- 2:04So, I have the paper pulled up here.
- 2:06Now, this isn't the very first paper
- 2:07that proposed the use of uh multi-layer
- 2:10perceptrons or neural networks to
- 2:11predict the next character or token in a
- 2:13sequence, but it's definitely one that
- 2:15is uh was very influential around that
- 2:17time. It is very often cited to stand in
- 2:19for this idea, and I think it's a very
- 2:21nice write-up. And so, this is the paper
- 2:23that we're going to first look at and
- 2:25then implement. Now, this paper has 19
- 2:27pages. So, we don't have time to go into
- 2:30the full detail of this paper, but I
- 2:31invite you to read it. Uh it's very
- 2:33readable, interesting, and has a lot of
- 2:34interesting ideas in it as well.
- 2:37In the introduction, they describe the
- 2:38exact same problem I just described. And
- 2:40then to address it, they propose the
- 2:42following model.
- 2:43Now, keep in mind that we are building a
- 2:46character-level language model. So,
- 2:48we're working on the level of
- 2:49characters. In this paper, they have a
- 2:51vocabulary of 17,000 possible words, and
- 2:54they instead built a word-level language
- 2:56model. But, we're going to still stick
- 2:57with the characters, but we'll take the
- 2:59same modeling approach.
- 3:01Now, what they do is basically they
- 3:03propose to take every one of these
- 3:04words, 17,000 words, and they're going
- 3:07to associate to each word a say
- 3:1030-dimensional feature vector.
- 3:12So, every word is now embedded into a
- 3:1630-dimensional space. You can think of
- 3:18it that way. So, we have 17,000 points
- 3:21or vectors in a 30-dimensional space,
- 3:23and that's um you might imagine that's
- 3:25very crowded. That's a lot of points for
- 3:27a very small space.
- 3:28Now,
- 3:30in the beginning, these words are
- 3:31initialized completely randomly. So,
- 3:32they're spread out at random.
- 3:34But, then we're going to tune these
- 3:36embeddings of these words using back
- 3:38propagation. So, during the course of
- 3:40training of this neural network, these
- 3:42points or vectors are going to basically
- 3:43move around in this space. And you might
- 3:46imagine that, for example, words that
- 3:48have very similar meanings or that are
- 3:50indeed synonyms of each other might end
- 3:52up in a very similar part of the space.
- 3:54And conversely, words that mean very
- 3:55different things would go somewhere else
- 3:57in the space.
- 3:59Now, their modeling approach otherwise
- 4:01is identical to ours. They are using a
- 4:03multi-layer neural network to predict
- 4:04the next word given the previous words.
- 4:07And to train the neural network, they
- 4:08are maximizing the log likelihood of the
- 4:10training data, just like we did.
- 4:12So, the modeling approach itself is
- 4:14identical. Now, here they have a
- 4:16concrete example of this intuition.
- 4:18Why does it work?
- 4:20Basically, suppose that for example, you
- 4:21are trying to predict "A dog was running
- 4:23in a blank."
- 4:25Now, suppose that the exact phrase "A
- 4:27dog was running in a" has never occurred
- 4:30in the training data.
- 4:31And here you are at uh sort of test time
- 4:33later when the model is deployed
- 4:35somewhere.
- 4:36And it's trying to make a sentence, and
- 4:38it's saying "A dog was running in a
- 4:40blank."
- 4:41And because it's never encountered this
- 4:43exact phrase in the training set, you're
- 4:45out of distribution, as we say. Like,
- 4:47you don't have fundamentally any
- 4:49reason to suspect um
- 4:52what might come next.
- 4:54But, this approach actually allows you
- 4:55to get around that. Because maybe you
- 4:57didn't see the exact phrase "A dog was
- 4:59running in a something." But, maybe
- 5:01you've seen similar phrases. Maybe
- 5:02you've seen the phrase "The dog was
- 5:04running in a blank."
- 5:06And maybe your network has learned that
- 5:07A and the are like frequently are
- 5:10interchangeable with each other. And so,
- 5:12maybe it took the embedding for A and
- 5:14the embedding for the, and it actually
- 5:16put them like nearby each other in the
- 5:17space. And so, you can transfer
- 5:19knowledge through that embedding, and
- 5:21you can generalize in that way.
- 5:23Similarly, the network could know that
- 5:25cats and dogs are animals, and they
- 5:26co-occur in lots of very similar
- 5:28contexts. And so, even though you
- 5:30haven't seen this exact phrase,
- 5:32or if you haven't seen exactly walking
- 5:34or running, you can through the
- 5:36embedding space transfer knowledge, and
- 5:38you can generalize to novel scenarios.
- 5:42So, let's now scroll down to the diagram
- 5:43of the neural network. Uh they have a
- 5:45nice uh diagram here.
- 5:47And in this example, we are taking three
- 5:49previous words
- 5:51and we are trying to predict the fourth
- 5:53word
- 5:54in the sequence.
- 5:56Now, these three previous words, as I
- 5:57mentioned, we have a vocabulary of
- 5:5917,000
- 6:01possible words.
- 6:03So, every one of these
- 6:04basically are the index of the incoming
- 6:08word.
- 6:09And because there are 17,000 words, this
- 6:11is an integer between 0 and 16,999.
- 6:17Now, there's also a lookup table that
- 6:19they call C.
- 6:21This lookup table is a matrix that is
- 6:2217,000 by, say, 30.
- 6:26And basically what we're doing here is
- 6:27we're treating this as a lookup table.
- 6:29And so, every index is plucking out a
- 6:32row of this embedding matrix
- 6:35so that each index is converted to the
- 6:3730-dimensional vector that corresponds
- 6:39to the embedding vector for that word.
- 6:42So, here we have the input layer of 30
- 6:45neurons for three words, making up 90
- 6:48neurons in total.
- 6:50And here they're saying that this matrix
- 6:52C is shared across all the words. So,
- 6:54we're always indexing into the same
- 6:56matrix C over and over.
- 6:59for each one of these words.
- 7:02Next up is the hidden layer of this
- 7:03neural network.
- 7:04The size of this hidden neural layer of
- 7:06this neural net is a hyper parameter.
- 7:09So, we use the word hyper parameter when
- 7:10it's kind of like a design choice up to
- 7:12the designer of the neural net. And this
- 7:13can be as large as you'd like or as
- 7:15small as you'd like. So, for example,
- 7:17the size could be 100.
- 7:19And we are going to go over multiple
- 7:20choices of the size of this hidden
- 7:22layer, and we're going to evaluate how
- 7:24well they work.
- 7:26So, say there were 100 neurons here,
- 7:28all of them would be fully connected to
- 7:29the 90 words or 90 um
- 7:33numbers that make up these three words.
- 7:35So, this is a fully connected layer.
- 7:38Then there's a tanh non-linearity.
- 7:40And then there's this output layer. And
- 7:42because there are 17,000 possible words
- 7:44that could come next, this layer has
- 7:4617,000 neurons, and all of them are
- 7:49fully connected to all of these neurons
- 7:52in the hidden layer.
- 7:54So there's a lot of parameters here
- 7:56because there's a lot of words. So most
- 7:58computation is here. This is the
- 7:59expensive layer.
- 8:01Now there are 17,000 logits here. So on
- 8:04top of there, we have the softmax layer,
- 8:06which we've seen in our previous video
- 8:08as well. So every one of these logits is
- 8:10exponentiated, and then everything is
- 8:12normalized to sum to one so that we have
- 8:15a nice probability distribution for the
- 8:17next word in the sequence.
- 8:19Now of course during training, we
- 8:21actually have the label. We have the
- 8:23identity of the next word in the
- 8:24sequence.
- 8:25That word or its index is used to pluck
- 8:29out the probability of that word,
- 8:32and then we are maximizing the
- 8:34probability of that word with respect to
- 8:37the parameters of this neural net.
- 8:39So the parameters are the weights and
- 8:41biases of this output layer, the weights
- 8:44and biases of this hidden layer, and the
- 8:47embedding lookup table C. And all of
- 8:49that is optimized using backpropagation.
- 8:52And these dashed arrows, ignore those.
- 8:55That represents a variation of a neural
- 8:57net that we are not going to explore in
- 8:58this video.
- 8:59So that's the setup, and now let's
- 9:01implement it.
- 9:02Okay, so I started a brand new notebook
- 9:04for this lecture.
- 9:05We are importing PyTorch, and we are
- 9:07importing Matplotlib so we can create
- 9:09figures.
- 9:10Then I am reading all the names into a
- 9:13list of words like I did before, and I'm
- 9:15showing the first eight right here.
- 9:18Keep in mind that we have a 32,000 in
- 9:20total. These are just the first eight.
- 9:22And then here I'm building out the
- 9:24vocabulary of characters and all the
- 9:25mappings from the characters as strings
- 9:28to integers and vice versa.
- 9:31Now the first thing we want to do is we
- 9:32want to compile the data set for the
- 9:34neural network. And I had to rewrite
- 9:36this code.
- 9:37I'll show you in a second what it looks
- 9:39like.
- 9:41So this is the code that I created for
- 9:43the data set creation. So let me first
- 9:44run it and then I'll briefly explain how
- 9:46this works.
- 9:48So first we're going to define something
- 9:50called block size. And this is basically
- 9:52the context length of how many
- 9:54characters do we take to predict the
- 9:56next one. So here in this example we're
- 9:58taking three characters to predict the
- 9:59fourth one. So we have a block size of
- 10:02three. That's the size of the block that
- 10:04supports the prediction.
- 10:06Then here I'm building out the X and Y.
- 10:10The X are the input to the neural net
- 10:12and the Y are the labels for each
- 10:15example inside X.
- 10:17Then I'm iterating over the first five
- 10:19words. I'm doing first five just for
- 10:21efficiency while we are developing all
- 10:23the code. But then later we're going to
- 10:24come here and erase this so that we use
- 10:26the entire training set.
- 10:29So here I'm printing the word Emma.
- 10:32And here I'm basically showing the
- 10:33examples that we can generate. The five
- 10:36examples that we can generate out of the
- 10:37single
- 10:38serve word Emma.
- 10:41So
- 10:42when we are given the context of just
- 10:44dot dot dot, the first character in a
- 10:45sequence is E.
- 10:47In this context, the label is M.
- 10:50When the context is this, the label is
- 10:52M.
- 10:53And so forth.
- 10:54And so the way I build this out is first
- 10:56I start with a padded context of just
- 10:57zero tokens.
- 10:59Then I iterate over all the characters.
- 11:01I get the character in the sequence and
- 11:04I basically build out the array Y of
- 11:07this current character and the array X
- 11:09which stores the current running
- 11:10context.
- 11:11And then here see I print everything and
- 11:14here I
- 11:15crop the context and enter the new
- 11:17character in the sequence. So this is
- 11:19kind of like a rolling window of
- 11:20context.
- 11:22Now we can change the block size here to
- 11:24for example four.
- 11:25And in that case we would be predicting
- 11:27the fifth character given the previous
- 11:29four.
- 11:30Or it can be five and then it would look
- 11:32like this.
- 11:34Or it can be say 10.
- 11:36And then it would look something like
- 11:37this. We're taking 10 characters to
- 11:39predict the 11th one.
- 11:41And we're always padding with dots.
- 11:43So let me bring this back to three, just
- 11:45so that uh we have what we have here in
- 11:48the paper.
- 11:50And finally, the data set right now
- 11:51looks as follows.
- 11:53From these five words, we have created a
- 11:55data set of 32 examples.
- 11:57And each input to the neural net is
- 11:59three integers, and we have a label that
- 12:02is also an integer, uh Y. So X looks
- 12:05like this.
- 12:06These are the individual examples.
- 12:08And then Y are the labels.
- 12:12So
- 12:13given this,
- 12:15let's now write the neural network that
- 12:17takes these X's and predicts the Y's.
- 12:19First, let's build the embedding uh look
- 12:21up table C.
- 12:23So we have 27 possible characters, and
- 12:25we're going to embed them in a lower
- 12:26dimensional space.
- 12:28In the paper, they have 17,000 words,
- 12:31and they embed them in uh spaces as
- 12:33small dimensional as 30. So they cram
- 12:3617,000 um
- 12:38words into 30-dimensional space. In our
- 12:40case, we have only 27 possible
- 12:42characters, so let's cram them in
- 12:44something as small as to start with, for
- 12:46example, a two-dimensional space.
- 12:48So this look up table will be random
- 12:50numbers,
- 12:51and we'll have 27 rows, and we'll have
- 12:54two columns.
- 12:56Right? So each 20 each one of 27
- 12:58characters will have a two-dimensional
- 13:00embedding.
- 13:02So that's our matrix C of embeddings, in
- 13:05the beginning initialized randomly.
- 13:07Now, before we embed all of the integers
- 13:10inside the input X using this look up
- 13:12table C,
- 13:14let me actually just try to embed a
- 13:15single individual integer, like say
- 13:17five.
- 13:19Um so we get a sense of how this works.
- 13:21Now, one way this works, of course, is
- 13:24we can just take the C, and we can index
- 13:26into row five.
- 13:28And that gives us a vector, the fifth
- 13:30row of C.
- 13:31And um
- 13:33this is one way to do it.
- 13:34The other way that I presented in the
- 13:36previous lecture is actually seemingly
- 13:38different, but actually identical.
- 13:40So, in the previous lecture, what we did
- 13:42is we took these integers and we used
- 13:43the one-hot encoding to first encode
- 13:46them.
- 13:46So, F.one_hot
- 13:48we want to encode integer five
- 13:50and we want to tell it that the number
- 13:52of classes is 27. So, that's the
- 13:5426-dimensional vector of all zeros
- 13:56except the fifth bit is turned on.
- 14:00Now, this actually doesn't work.
- 14:03The reason is that
- 14:04this input actually must be a
- 14:05torch.tensor.
- 14:08And I'm making some of these errors
- 14:09intentionally just so you get to see
- 14:10some errors and how to fix them.
- 14:12Uh so, this must be a tensor not an int,
- 14:14fairly straightforward to fix.
- 14:16We get a one-hot vector. The fifth
- 14:18dimension is one and the shape of this
- 14:20is 27.
- 14:22And now notice that, just as I briefly
- 14:24alluded to in the previous video, if we
- 14:26take this one-hot vector and we multiply
- 14:29it by C
- 14:33then
- 14:35um
- 14:35what would you expect?
- 14:37Well, number one
- 14:39first you'd expect an error
- 14:41because um
- 14:43expected scalar type long but found
- 14:45float. So, a little bit confusing, but
- 14:48the problem here is that one-hot, the
- 14:50data type of it
- 14:52is long. It's a 64-bit integer. But,
- 14:56this is a float tensor and so PyTorch
- 14:58doesn't know how to multiply an int with
- 15:00a float and that's why we had to
- 15:02explicitly cast this to a float so that
- 15:04we can multiply.
- 15:06Now, the output actually here
- 15:09is identical.
- 15:11And that it's identical because of the
- 15:12way the matrix uh multiplication here
- 15:14works. We have the one-hot um vector
- 15:17multiplying columns of C and because of
- 15:20all the zeros, they actually end up
- 15:22masking out everything in C except for
- 15:24the fifth row, which is plucked out.
- 15:27And so we actually arrive at the same
- 15:28result.
- 15:30And that tells you that here we can
- 15:31interpret this first piece here, this
- 15:34embedding of the integer. We can either
- 15:36think of it as the integer indexing into
- 15:38a lookup table C.
- 15:40But equivalently, we can also think of
- 15:41this little piece here as a first layer
- 15:44of this bigger neural net.
- 15:46This layer here has neurons that have no
- 15:48non-linearity, they're no tan h, they're
- 15:50just linear neurons, and their weight
- 15:52matrix is C.
- 15:55And then we are encoding integers into
- 15:57one-hot and feeding those into a neural
- 15:59net. And this first layer basically
- 16:01embeds them.
- 16:02So those are two equivalent ways of
- 16:04doing the same thing. We're just going
- 16:06to index because it's much much faster,
- 16:08and we're going to discard this
- 16:09interpretation of one-hot inputs into
- 16:12neural nets. And we're just going to
- 16:14index integers and create and use
- 16:16embedding tables. Now, embedding a
- 16:18single integer like five is easy enough.
- 16:20We can simply ask PyTorch to retrieve
- 16:22the fifth row of C. Or the row index
- 16:25five of C.
- 16:27But how do we simultaneously embed all
- 16:30of these 32 by three integers stored in
- 16:32array X?
- 16:34Luckily, PyTorch indexing is fairly
- 16:36flexible and quite powerful. So it
- 16:39doesn't just work to um
- 16:41ask for a single element five like this.
- 16:44You can actually index using lists. So
- 16:46for example, we can get the rows five,
- 16:48six, and seven, and this will just work
- 16:51like this. We can index with a list.
- 16:53It doesn't just have to be a list, it
- 16:55can also be a actually a tensor of
- 16:57integers.
- 16:58And we can index with that.
- 17:00So this is a integer tensor 567, and
- 17:03this will just work as well.
- 17:06In fact, we can also, for example,
- 17:07repeat row seven and retrieve it
- 17:09multiple times. And uh
- 17:12that same index will just get embedded
- 17:14multiple times here.
- 17:16So here we are indexing with a
- 17:17one-dimensional
- 17:19tensor of integers. But, it turns out
- 17:21that you can also index with
- 17:22multi-dimensional tensors of integers.
- 17:25Here, we have a two-dimensional in
- 17:27tensor of integers. So, we can simply
- 17:29just do C at X.
- 17:32And this just works.
- 17:34And the shape of this
- 17:36is
- 17:3732 by 3, which is the original shape.
- 17:40And now, for every one of those 32 by 3
- 17:41integers, we've retrieved the embedding
- 17:43vector here.
- 17:46So, basically, we have that as an
- 17:48example.
- 17:49The 13th or example, index 13,
- 17:53um the second dimension is the integer
- 17:56one as an example.
- 17:58And so, here, if we do C of X, which
- 18:02gives us that array, and then we index
- 18:04into 13 by 2 of that array,
- 18:07then we we get the embedding
- 18:09here.
- 18:10And you can verify that C at one,
- 18:14which is the integer at that location,
- 18:16is indeed equal to this.
- 18:20You see they're equal.
- 18:21So, basically, long story short, PyTorch
- 18:23indexing is awesome. And to embed
- 18:26simultaneously
- 18:28all of the integers in X, we can simply
- 18:30do C of X. And that is our embedding.
- 18:33And that just works.
- 18:35Now, let's construct this layer here,
- 18:37the hidden layer.
- 18:38So, we have that W1, as I'll call it,
- 18:42are these weights, which we will
- 18:44initialize randomly.
- 18:46Now, the number of inputs to this layer
- 18:48is going to be 3 * 2, right? Because we
- 18:51have two-dimensional embeddings, and we
- 18:52have three of them.
- 18:53So, the number of inputs is six.
- 18:56And the number of neurons in this layer
- 18:58is a variable up to us. Let's use 100
- 19:01neurons as an example.
- 19:03And then, biases will be also
- 19:05initialized randomly as an example.
- 19:07And let's And we just need 100 of them.
- 19:11Now, the problem with this is we can't
- 19:13simply Normally, we would take the
- 19:15input, in this case that's embedding,
- 19:17and we'd like to multiply it with these
- 19:19weights.
- 19:20And then we would like to add the bias.
- 19:22This is roughly what we want to do.
- 19:24But the problem here is that these
- 19:25embeddings are stacked up in the
- 19:27dimensions of this input tensor.
- 19:29So, this will not work, this matrix
- 19:31multiplication, because this is a shape
- 19:3232 by 3 by 2, and I can't multiply that
- 19:35by 6 by 100.
- 19:37So, somehow we need to concatenate these
- 19:40inputs here together so that we can do
- 19:41something along these lines, which
- 19:43currently does not work.
- 19:45So, how do we transform this 32 by 3 by
- 19:472 into a 32 by 6 so that we can actually
- 19:50perform this uh multiplication over
- 19:53here? I'd like to show you that there
- 19:55are usually many ways of uh implementing
- 19:58what you'd like to do in torch.
- 20:00And some of them will be faster, better,
- 20:02shorter, etc.
- 20:03And that's because torch is a very large
- 20:06library, and it's got lots and lots of
- 20:07functions. So, if we just go to the
- 20:09documentation and click on torch, you'll
- 20:11see that my slider here is very tiny,
- 20:14and that's because there are so many
- 20:15functions that you can call on these
- 20:16tensors
- 20:17to transform them, create them, multiply
- 20:20them, add them, perform all kinds of
- 20:22different operations on them.
- 20:24And so, this is kind of like
- 20:28the space of possibility, if you will.
- 20:31Now, one of the things that you can do
- 20:32is if we can control here, control F for
- 20:34concatenate. And we see that there's a
- 20:36function, torch.cat, short for
- 20:38concatenate.
- 20:40And this concatenates a given sequence
- 20:42of tensors in a given dimension.
- 20:45And uh these tensors must have the same
- 20:46shape, etc. So, we can use the
- 20:48concatenate operation to, in a naive
- 20:51way, concatenate these three embeddings
- 20:53for each input.
- 20:56So, in this case, we have emb of
- 20:58emb of the shape. And really what we
- 21:00want to do is we want to retrieve these
- 21:02three parts and concatenate them.
- 21:05So, we want to grab all the examples.
- 21:08We want to grab
- 21:10first the zeroth um
- 21:13index and then all of
- 21:16this.
- 21:17So, this plucks out
- 21:20the 32 by 2 embeddings of just the first
- 21:24word here.
- 21:26And so, basically we want this guy.
- 21:28We want the first dimension. And we want
- 21:31the second dimension.
- 21:32And these are the three pieces
- 21:34individually.
- 21:36And then we want to treat this as a
- 21:38sequence and we want to torch.cat
- 21:40on that sequence. So, this is the list.
- 21:43torch.cat takes a sequence of tensors.
- 21:47And then we have to tell it along which
- 21:49dimension to concatenate.
- 21:51So, in this case all these are 32 by 2
- 21:53and we want to concatenate not across
- 21:55dimension zero, but across dimension
- 21:57one.
- 21:58So, passing in one
- 22:00gives us a result.
- 22:01The shape of this is 32 by 6 exactly as
- 22:04we'd like.
- 22:05So, that basically took 32 and squashed
- 22:08these by concatenating them into 32 by
- 22:106.
- 22:11Now, this is kind of ugly because this
- 22:13code would not generalize if we want to
- 22:15later change the block size. Right now,
- 22:17we have three inputs.
- 22:19Three words. But what if we had five?
- 22:22Then here we would have to change the
- 22:23code because I'm indexing directly.
- 22:25Well, torch comes to rescue again
- 22:27because there turns out to be a function
- 22:29called unbind.
- 22:31And it removes a tensor dimension.
- 22:35So, removes a tensor dimension, returns
- 22:37a tuple of all slices along a given
- 22:39dimension without it.
- 22:41So, this is exactly what we need.
- 22:43And basically when we call torch.unbind
- 22:48torch.unbind
- 22:50of emb and passing dimension um one,
- 22:55index one.
- 22:56This gives us a list of
- 22:58um a list of tensors exactly equivalent
- 23:01to this.
- 23:02So, running this
- 23:04gives us a length
- 23:06three
- 23:07and it's exactly this list. And so, we
- 23:09can call torch.cat on it.
- 23:12And along the first dimension.
- 23:15And this works.
- 23:16And the shape is the same.
- 23:19But now this is uh it doesn't matter if
- 23:21we have block size three or five or 10,
- 23:23this will just work.
- 23:24So, this is one way to do it. But it
- 23:26turns out that in this case, there's
- 23:28actually a significantly better and more
- 23:30efficient way. And this gives me an
- 23:32opportunity to hint at some of the
- 23:34internals of torch.tensor.
- 23:36So, let's create
- 23:38an array here
- 23:40of elements from zero to 17. And the
- 23:42shape of this
- 23:44is just 18. It's a single vector of 18
- 23:46numbers.
- 23:48It turns out that we can very quickly
- 23:49re-represent this as different sized and
- 23:53dimensional tensors.
- 23:54We do this by calling a view.
- 23:57And we can say that actually this is not
- 23:59a single vector of 18, this is a 2x9
- 24:02tensor.
- 24:04Or alternatively, this is a 9x2 tensor.
- 24:08Or this is actually a 3x3x2 tensor.
- 24:11As long as the total number of elements
- 24:13here multiply to be the same, uh this
- 24:16will just work.
- 24:18And in PyTorch, this operation calling
- 24:21dot view is extremely efficient.
- 24:24And the reason for that is that in each
- 24:26tensor, there's something called the
- 24:28underlying storage.
- 24:30And the storage is just the numbers
- 24:32always as a one-dimensional vector. And
- 24:34this is how this tensor is represented
- 24:37in the computer memory. It's always a
- 24:38one-dimensional vector.
- 24:41But when we call that view, we are
- 24:44manipulating some of attributes of that
- 24:46tensor that dictate how this
- 24:48one-dimensional sequence is interpreted
- 24:51to be an N-dimensional tensor.
- 24:53And so, what's happening here is that no
- 24:55memory is being changed, copied, moved,
- 24:57or created when we call dot view. The
- 24:59The storage
- 25:00is identical, but when you call dot
- 25:02view, some of the internal um
- 25:05attributes of the view of this tensor
- 25:07are being manipulated and changed. In
- 25:09particular, there's something There's
- 25:10something called a storage offset,
- 25:12strides, and shapes, and those are
- 25:14manipulated so that this one-dimensional
- 25:16sequence of bytes is seen as different
- 25:18n-dimensional arrays.
- 25:20There's a blog post here from Eric
- 25:22called PyTorch Internals, where he goes
- 25:25into some of this with respect to tensor
- 25:27and how the view of a tensor is
- 25:29represented.
- 25:30And this is really just like a logical
- 25:32construct of representing the physical
- 25:34memory.
- 25:35And uh so this is a pretty good um
- 25:38blog post that you can go into. I might
- 25:39also create an entire video on the
- 25:41internals of torch tensor and how this
- 25:42works.
- 25:44For here, we just note that this is an
- 25:46extremely efficient operation.
- 25:48And if I delete this and come back to
- 25:50our emb,
- 25:53we see that the shape of our emb is 32
- 25:55by 3 by 2, but we can simply ask for
- 25:58PyTorch to view this instead as a 32 by
- 26:016.
- 26:03And the way this gets flattened into a
- 26:0532 by 6 array
- 26:07just happens that
- 26:09these two
- 26:10get stacked up in a single row. And so
- 26:13that's basically the concatenation
- 26:15operation that we're after.
- 26:17And you can verify that this actually
- 26:18gives the exact same result as what we
- 26:20had before.
- 26:22So this is an element-wise equals, and
- 26:23you can see that all the elements of
- 26:25these two tensors are the same.
- 26:27And so we get the exact same result.
- 26:30So long story short, we can actually
- 26:32just come here,
- 26:34and if we just view this as a 32 by 6
- 26:38um instead, then this multiplication
- 26:40will work and give us the hidden states
- 26:43that we're after.
- 26:44So if this is H,
- 26:46then H.shape is now the 100-dimensional
- 26:49activations for every one of our 32
- 26:52examples.
- 26:53And this gives the desired result. Let
- 26:55me do two things here. Number one, let's
- 26:57not use 32. We can, for example, do
- 27:00something like
- 27:01um
- 27:02emb.shape
- 27:04at zero.
- 27:05So that we don't hardcode these numbers.
- 27:07And this would work for any size of this
- 27:09emb.
- 27:10Or alternatively, we can also do -1.
- 27:12When we do -1, PyTorch will infer what
- 27:15this should be.
- 27:16Because the number of elements must be
- 27:17the same, and we're saying that this is
- 27:19six, PyTorch will derive that this must
- 27:21be 32 or whatever else it is if emb is
- 27:24of different size.
- 27:26The other thing is here,
- 27:28um
- 27:29one more thing I'd like to point out is
- 27:33here when we do the concatenation,
- 27:35this actually is much less efficient
- 27:37because um this concatenation would
- 27:39create a whole new tensor with a whole
- 27:41new storage. So new memory is being
- 27:43created because there's no way to
- 27:44concatenate tensors just by manipulating
- 27:47the view attributes. So this is
- 27:49inefficient and creates all kinds of new
- 27:50memory.
- 27:52Uh so let me delete this now.
- 27:55We don't need this.
- 27:57And here to calculate H, we want to also
- 27:59dot tanh
- 28:01of this to get our
- 28:04oops, to get our H.
- 28:07So these are now numbers between -1 and
- 28:081 because of the tanh.
- 28:10And we have that the shape is 32 by 100.
- 28:14And that is basically this hidden layer
- 28:16of activations here
- 28:17for every one of our 32 examples.
- 28:20Now there's one more thing I glossed
- 28:21over that we have to be very careful
- 28:23with, and that's this and that's this
- 28:25plus here.
- 28:26In particular, we want to make sure that
- 28:27the broadcasting will do what we like.
- 28:30The shape of this is 32 by 100, and B1's
- 28:33shape is 100.
- 28:35So we see that the addition here will
- 28:37broadcast these two. And in particular,
- 28:39we have 32 by 100 broadcasting to 100.
- 28:44So broadcasting will align on the right,
- 28:47create a fake dimension here.
- 28:49So, this will become a 1 by 100 row
- 28:50vector.
- 28:52And then it will copy vertically
- 28:54for every one of these rows of 32 and do
- 28:57an element-wise addition.
- 28:58So, in this case, the correct thing will
- 29:00be happening because the same bias
- 29:02vector
- 29:03will be added to all the rows
- 29:05of
- 29:06this matrix.
- 29:08So, that is correct. That's what we'd
- 29:09like. And uh it's always good practice
- 29:11to just make sure uh so that you don't
- 29:13shoot yourself in the foot. And finally,
- 29:15let's create the final layer here.
- 29:17So, let's create
- 29:19W2 and B2.
- 29:22The input now is 100.
- 29:25And the output number of neurons will be
- 29:27for us 27 because we have 27 possible
- 29:29characters that come next.
- 29:31So, the biases will be 27 as well.
- 29:35So, therefore, the logits, which are the
- 29:36outputs of this neural net,
- 29:38are going to be um
- 29:41H multiplied by W2 plus B2.
- 29:47Logits.shape is 32 by 27.
- 29:50And the logits look
- 29:52good. Now, exactly as we saw in the
- 29:54previous video, we want to take these
- 29:55logits and we want to first exponentiate
- 29:58them to get our fake counts.
- 30:00And then we want to normalize them into
- 30:01a probability.
- 30:03So, prob is counts.divide.
- 30:05And now uh counts.sum
- 30:09along the first dimension and keep dims
- 30:11as true, exactly as in the previous
- 30:12video.
- 30:14And so,
- 30:16prob.shape now is 32 by 27.
- 30:20And you'll see that every row of prob
- 30:23sums to one, so it's normalized.
- 30:26So, that gives us the probabilities.
- 30:28Now, of course, we have the actual
- 30:29letter that comes next. And that comes
- 30:31from this array Y,
- 30:34which we which we created during the
- 30:36data set creation. So, Y is this last
- 30:39piece here, which is the identity of the
- 30:40next character in the sequence that we'd
- 30:42like to now predict.
- 30:44So, what we'd like to do now is just as
- 30:46in the previous video, we'd like to
- 30:48index into the rows of prob, and each
- 30:50row we'd like to pluck out the
- 30:52probability assigned to the correct
- 30:54character
- 30:55as given here.
- 30:57So, first we have torch.arange of 32,
- 31:00which is kind of like a iterator over um
- 31:03numbers from 0 to 31,
- 31:05and then we can index into prob in the
- 31:07following way.
- 31:09prob in torch.arange of 32, which
- 31:12iterates the rows. And then in each row,
- 31:14we'd like to grab this column as given
- 31:17by Y.
- 31:19So, this gives the current probabilities
- 31:21as assigned by this neural network with
- 31:23this setting of its weights
- 31:25to the correct character in the
- 31:26sequence.
- 31:27And you can see here that this looks
- 31:29okay for some of these characters. Like
- 31:31this is basically 0.2,
- 31:33but it doesn't look very good at all for
- 31:34many other characters. Like this is
- 31:360.0701
- 31:38probability. And so, the network thinks
- 31:40that some of these are extremely
- 31:42unlikely. But of course, we haven't
- 31:43trained the neural network yet. So,
- 31:46um this will improve, and ideally all of
- 31:49these numbers here of course are one
- 31:50because then we are correctly predicting
- 31:52the next character.
- 31:53Now, just as in the previous video, we
- 31:55want to take these probabilities, we
- 31:57want to look at the log probability, and
- 31:59then we want to look at the average log
- 32:01probability,
- 32:02and then negative of it to create the
- 32:04negative log likelihood loss.
- 32:07So, the loss here is 17.
- 32:10And this is the loss that we'd like to
- 32:11minimize to get the network to predict
- 32:14the correct character in the sequence.
- 32:16Okay, so I rewrote everything here and
- 32:18made it a bit more respectable.
- 32:20So, here's our data set.
- 32:22Here's all the parameters that we
- 32:23defined.
- 32:24I'm now using a generator to make it
- 32:26reproducible.
- 32:27I clustered all the parameters into a
- 32:29single list of parameters, so that for
- 32:31example, it's easy to count them and see
- 32:33that in total we currently have about
- 32:343,400 parameters.
- 32:37And this is the forward pass as we
- 32:38developed it.
- 32:39And we arrive at a single number here,
- 32:42the loss, that is currently expressing
- 32:44how well this neural network works with
- 32:46the current setting of parameters.
- 32:48Now, I would like to make it even more
- 32:50respectable.
- 32:51So, in particular, see these lines here,
- 32:53where we take the logits and we
- 32:54calculate the loss.
- 32:57Um we're not actually reinventing the
- 32:59wheel here. This is just um
- 33:01classification, and many people use
- 33:03classification, and that's why there is
- 33:05a functional.cross_entropy function in
- 33:07PyTorch to calculate this much more
- 33:09efficiently.
- 33:10So, we could just simply call
- 33:11f.cross_entropy,
- 33:13and we can pass in the logits, and we
- 33:14can pass in the
- 33:16uh array of targets, Y.
- 33:18And this calculates the exact same loss.
- 33:21Um so, in fact, we can simply put this
- 33:24here
- 33:25and erase these three lines, and we're
- 33:27going to get the exact same result. Now,
- 33:30there are actually many good reasons to
- 33:31prefer f.cross_entropy over rolling your
- 33:34own implementation like this. I did this
- 33:36for educational reasons, but you'd never
- 33:38use this in practice. Why is that?
- 33:40Number one, when you use
- 33:41f.cross_entropy, PyTorch will not
- 33:43actually create all these intermediate
- 33:45tensors, because these are all new
- 33:47tensors in memory, and all this is
- 33:49fairly inefficient to run like this.
- 33:52Instead, PyTorch will cluster up all
- 33:54these operations and very often create
- 33:57have a fused kernels that very
- 33:59efficiently evaluate these expressions
- 34:00that are sort of like clustered uh
- 34:02mathematical operations.
- 34:04Number two, the backward pass can be
- 34:06made much more efficient. And not just
- 34:08because it's a fused kernel, but also
- 34:10analytically and mathematically, it's
- 34:12much it's often a very much uh simpler
- 34:15backward pass to implement.
- 34:17We actually saw this with micrograd.
- 34:19You see here when we implemented tanh,
- 34:21the forward pass of this operation to
- 34:23calculate the tanh was actually a fairly
- 34:25complicated mathematical expression.
- 34:28But because it's a clustered
- 34:29mathematical expression, when we did the
- 34:31backward pass, we didn't individually
- 34:33backward through the exp and the two
- 34:35times and the minus one and the
- 34:36division, etc. We just said it's 1 - t
- 34:39squared. And that's a much simpler
- 34:41mathematical expression.
- 34:43And we were able to do this because
- 34:44we're able to reuse calculations and
- 34:46because we are able to mathematically
- 34:48and analytically derive the derivative
- 34:50and often that expression simplifies
- 34:52mathematically. And so there's much less
- 34:54to implement.
- 34:56So not only can can it be made more
- 34:57efficient because it runs in a fused
- 34:59kernel, but also because the expressions
- 35:01can take a much simpler form
- 35:03mathematically.
- 35:06So that's number one. Number two,
- 35:08under the hood, after cross entropy can
- 35:10also be significantly more
- 35:12um
- 35:13numerically well-behaved. Let me show
- 35:15you an example of how this works.
- 35:19Suppose we have a logits of -2, 3, -3,
- 35:220, and 5.
- 35:24And then we are taking the exponent of
- 35:25it and normalizing it to sum to one. So
- 35:28when logits take on these values,
- 35:30everything is well and good and we get a
- 35:31nice probability distribution.
- 35:33Now consider what happens when some of
- 35:35these logits take on more extreme values
- 35:37and that can happen during optimization
- 35:39of a neural network.
- 35:40Suppose that some of these numbers grow
- 35:42very negative, like say -100.
- 35:45Then actually everything will come out
- 35:47fine. We still get the probabilities
- 35:48that um you know, are well-behaved and
- 35:52they sum to one and everything is great.
- 35:54But because of the way the exp works, if
- 35:56you have very positive logits, like say
- 35:58positive 100 in here,
- 36:00you actually start to run into trouble
- 36:02and we get not a number here.
- 36:04And the reason for that is that these
- 36:06counts
- 36:08have an inf here.
- 36:10So if you pass in a very negative number
- 36:12to exp, you just get a very negative
- 36:15Sorry, not negative, but very small
- 36:16number, very very near zero and that's
- 36:18fine.
- 36:20But if you pass in a very positive
- 36:21number, suddenly we run out of range in
- 36:23our floating point number that
- 36:25represents these counts.
- 36:28So basically we're taking E and we're
- 36:29raising it to the power of 100 and that
- 36:32gives us inf because we run out of
- 36:34dynamic range on this floating point
- 36:35number that is count.
- 36:38And so, we cannot pass very large logits
- 36:41through this expression.
- 36:43Now, let me reset these numbers to
- 36:45something reasonable.
- 36:47The way PyTorch solves this is that you
- 36:50see how we have a well-behaved result
- 36:52here.
- 36:53It turns out that because of the
- 36:54normalization here, you can actually
- 36:56offset logits by any arbitrary constant
- 36:59value that you want. So, if I add one
- 37:01here,
- 37:02you actually get the exact same result.
- 37:04Or if I add two,
- 37:06or if I subtract three.
- 37:08Any offset will produce the exact same
- 37:10probabilities.
- 37:12So, because negative numbers are okay,
- 37:15but positive numbers can actually
- 37:16overflow this exp, what PyTorch does is
- 37:19it internally calculates the maximum
- 37:21value that occurs in the logits and it
- 37:23subtracts it. So, in this case it would
- 37:25subtract five.
- 37:27And so, therefore the greatest number in
- 37:28logits will become zero and all the
- 37:30other numbers will become some negative
- 37:32numbers.
- 37:33And then the result of this is always
- 37:35well-behaved. So, even if we have a 100
- 37:37here previously,
- 37:39not good, but because PyTorch will
- 37:41subtract 100, this will work.
- 37:44And so, there's many good reasons to
- 37:46call cross entropy. Number one, the
- 37:49forward pass can be much more efficient.
- 37:50The backward pass can be much more
- 37:52efficient. And also things can be much
- 37:54more numerically well-behaved. Okay, so
- 37:56let's now set up the training of this
- 37:58neural net.
- 37:59We have the forward pass.
- 38:02Uh we don't need these. Except we have
- 38:05the loss is equal to the after cross
- 38:07entropy. That's the forward pass.
- 38:09Then we need the backward pass. First we
- 38:12want to set the gradients to be zero.
- 38:14So, for P and parameters, we want to
- 38:16make sure that P.grad is none, which is
- 38:18the same as setting it to zero in
- 38:19PyTorch.
- 38:21And then loss.backward to populate those
- 38:23gradients.
- 38:24Once we have the gradients, we can do
- 38:25the parameter update. So, for P in
- 38:27parameters, we want to take all the data
- 38:30and we want to nudge it
- 38:32learning rate times p.grad.
- 38:36And then we want to repeat this
- 38:39a few times.
- 38:41Um
- 38:44And let's print the loss here as well.
- 38:48Now, this won't suffice and will create
- 38:50an error because we also have to go for
- 38:52P in parameters
- 38:54and we have to make sure that
- 38:55p.requires_grad
- 38:57is set to true in PyTorch.
- 38:59And this should just work.
- 39:03Okay. So, we started off with loss of 17
- 39:05and we're decreasing it.
- 39:08Let's run longer.
- 39:10And you see how the loss decreases
- 39:12a lot here. So,
- 39:17if we just run for 1,000 times, we get a
- 39:20very, very low loss. And that means that
- 39:21we're making very good predictions. Now,
- 39:23the reason that this is so
- 39:25straightforward right now is because
- 39:27we're only um
- 39:29overfitting 32 examples.
- 39:32So, we only have 32 examples uh of the
- 39:34first five words
- 39:36and therefore it's very easy to make
- 39:38this neural net fit only these two 32
- 39:40examples because we have 3,400
- 39:43parameters and only 32 examples. So,
- 39:46we're doing what's called overfitting a
- 39:47single batch of the data
- 39:50and getting a very low loss and good
- 39:52predictions. Um but that's just because
- 39:54we have so many parameters for so few
- 39:56examples. So, it's easy to uh make this
- 39:58be very low.
- 40:00Now, we're not able to achieve exactly
- 40:01zero. And the reason for that is we can,
- 40:04for example, look at the logits which
- 40:06are being predicted.
- 40:08And uh we can look at the max along the
- 40:11first dimension.
- 40:13And in PyTorch, uh max reports both the
- 40:16actual values that take on the maximum
- 40:19uh, number, but also the indices of
- 40:20these.
- 40:22And you'll see that the indices are very
- 40:23close to the labels.
- 40:26But in some cases they differ. For
- 40:28example, in this very first example, uh,
- 40:31the predicted index is 19, but the label
- 40:33is five.
- 40:35And we're not able to make loss be zero,
- 40:37and fundamentally that's because here
- 40:40the very first or the zeroth index is
- 40:43the example where dot dot dot is
- 40:44supposed to predict E. But you see how
- 40:46dot dot dot is also supposed to predict
- 40:48an O. And dot dot dot is also supposed
- 40:50to predict an I, and then S as well. And
- 40:53so basically E, O, A, or S are all
- 40:56possible outcomes in the training set
- 40:58for the exact same input. So we're not
- 41:00able to completely overfit and um,
- 41:03and make the loss be exactly zero. Uh,
- 41:06so but we're getting very close in the
- 41:08cases where uh, there's a unique input
- 41:11for a unique output. In those cases we
- 41:13do what's called overfit, and we
- 41:15basically get the exact same and the
- 41:16exact correct result.
- 41:18So now all we have to do
- 41:21is we just need to make sure that we
- 41:22read in the full data set and optimize
- 41:24the neural net.
- 41:25Okay, so let's swing back up where we
- 41:27created the data set.
- 41:29And we see that here we only use the
- 41:30first five words. So let me now erase
- 41:32this,
- 41:33and let me erase the print statements,
- 41:35otherwise we'll be printing way too
- 41:36much.
- 41:38And so when we process the full data set
- 41:40of all the words, we now have 228,000
- 41:43examples instead of just 32.
- 41:45So let's now scroll back down. The
- 41:47dataset is much larger. We initialize
- 41:49the weights, the same number of
- 41:51parameters. They all require gradients.
- 41:54And then let's push this print out loss
- 41:56dot item to be here.
- 41:58And let's just see how the optimization
- 41:59goes if we run this.
- 42:04Okay, so we started with a fairly high
- 42:05loss, and then as we're optimizing, the
- 42:07loss is coming down.
- 42:11But you'll notice that it takes quite a
- 42:13bit of time for every single iteration.
- 42:15So, let's actually address that. Because
- 42:17we're doing way too much work forwarding
- 42:19and backwarding 220,000 examples.
- 42:22In practice, what people usually do is
- 42:24they perform forward and backward pass
- 42:26and update on mini batches of the data.
- 42:29So, what we will want to do is we want
- 42:31to randomly select some portion of the
- 42:33data set, and that's a mini batch, and
- 42:35then only forward backward and update on
- 42:37that little mini batch. And then, um,
- 42:40we iterate on those mini batches.
- 42:42So, in PyTorch, we can, for example, use
- 42:43torch.randint.
- 42:45We can generate numbers between 0 and 5
- 42:47and make 32 of them.
- 42:51Um,
- 42:52I believe the size has to be a tuple
- 42:56in PyTorch.
- 42:57So, we can have a tuple 32 of numbers
- 43:00between 0 and 5. But, actually, we want
- 43:02x.shape of 0 here.
- 43:05And so, this creates uh integers that
- 43:08index into our data set, and there's 32
- 43:10of them.
- 43:11So, if our mini batch size is 32, then
- 43:13we can come here and we can first do uh
- 43:16mini batch
- 43:18construct.
- 43:20So, in the integers that we want to
- 43:22optimize in this
- 43:23um single iteration are in the ix.
- 43:27And then, we want to index into x with
- 43:31ix to only grab those rows.
- 43:34So, we're only getting 32 rows of x.
- 43:37And therefore, embeddings will again be
- 43:3832 by 3 by 2, not 200,000 by 3 by 2.
- 43:43And then, this ix has to be used not
- 43:44just to index into x, but also to index
- 43:48into y.
- 43:50And now, this should be mini batches,
- 43:52and this should be much much faster. So,
- 43:55Okay, so it's instant almost.
- 43:58So, this way, we can run many many
- 44:00examples
- 44:01nearly instantly and decrease the loss
- 44:03much much faster.
- 44:05Now, because we're only dealing with
- 44:06mini batches, the quality of our
- 44:08gradient is lower. So, the direction is
- 44:11not as reliable. It's not the actual
- 44:13gradient direction.
- 44:14But, the gradient direction is good
- 44:16enough even when it's estimating on only
- 44:1832 examples that it is useful.
- 44:21And so, it's much better to have an
- 44:24approximate gradient and just make more
- 44:25steps than it is to evaluate the exact
- 44:28gradient and take fewer steps. So,
- 44:30that's why in practice,
- 44:32uh this works quite well.
- 44:34So, let's continue the optimization.
- 44:38Let me take out this loss.item from here
- 44:41and uh place it over here at the end.
- 44:46Okay, so we're hovering around 2.5 or
- 44:48so.
- 44:49Um
- 44:50however, this is only the loss for that
- 44:51mini batch. So, let's actually evaluate
- 44:53the loss
- 44:55here
- 44:56for all of X
- 44:58and for all of Y, just so we have a full
- 45:01sense of exactly how well the model is
- 45:03doing right now.
- 45:05So, right now we're at about 2.7 on the
- 45:07entire training set.
- 45:09So, let's run the optimization for a
- 45:11while.
- 45:12Okay, we're at 2.6.
- 45:152.57
- 45:172.53
- 45:22Okay.
- 45:22So, one issue, of course, is we don't
- 45:24know if we're stepping too slow or too
- 45:27fast.
- 45:28Um so, this point one, I just guessed
- 45:30it. So, one question is, how do you
- 45:32determine this learning rate?
- 45:34And um how do we gain confidence that
- 45:37we're stepping in the right um
- 45:39sort of speed? So, I'll show you one way
- 45:41to determine a reasonable learning rate.
- 45:43It works as follows. Let's reset our
- 45:46parameters
- 45:47to the initial
- 45:49um
- 45:49settings.
- 45:51And now, let's
- 45:52print in every step.
- 45:55But, let's only do 10 steps or so.
- 45:58Or maybe maybe 100 steps.
- 46:01We want to find like a very reasonable
- 46:02set
- 46:03search range, if you will. So, for
- 46:05example, if this is like very low,
- 46:07then
- 46:10we see that the loss is barely
- 46:11decreasing. So, that's not
- 46:13That's like too low, basically. So,
- 46:15let's try
- 46:16this one.
- 46:18Okay, so we're decreasing the loss, but
- 46:20like not very quickly. So, that's a
- 46:21pretty good low range.
- 46:23Now, let's reset it again.
- 46:26And now, let's try to find the place at
- 46:27which the loss kind of explodes.
- 46:29Uh so, maybe at -1.
- 46:33Okay, we see that we're minimizing the
- 46:34loss, but you see how it's kind of
- 46:36unstable. It goes up and down quite a
- 46:38bit.
- 46:39Um so, -1 is probably like a fast
- 46:42learning rate. Let's try -10.
- 46:45Okay, so this isn't optimizing. This is
- 46:48not working very well. So, -10 is way
- 46:50too big. -1 was already kind of big. Um
- 46:54so, therefore, -1 was like somewhat
- 46:57reasonable if I reset.
- 47:00So, I'm thinking that the right learning
- 47:01rate is somewhere between
- 47:03um -0.001
- 47:05and um
- 47:07-1.
- 47:08So, the way we can do this here is we
- 47:09can use uh torch.linspace.
- 47:13And we want to basically do something
- 47:14like this, between 0 and 1, but
- 47:17um
- 47:19uh number of steps is one more parameter
- 47:21that's required. Let's do 1,000 steps.
- 47:23This creates 1,000
- 47:26um numbers between 0.001 and 1.
- 47:29But, it doesn't really make sense to
- 47:31step between these linearly. So,
- 47:33instead, let me create learning rate
- 47:35exponent.
- 47:36And instead of 0.001, this will be a -3,
- 47:40and this will be a 0. And then, the
- 47:42actual uh LRs that we want to search
- 47:44over are going to be 10 to the power of
- 47:46LRE.
- 47:48So, now what we're doing is we're
- 47:49stepping linearly between the exponents
- 47:51of these learning rates. This is 0.001,
- 47:54and this is 1, because uh 10 to the
- 47:56power of zero is one.
- 47:58And therefore, we are spaced
- 48:00exponentially in this interval.
- 48:02So, these are the candidate learning
- 48:03rates
- 48:04that we want to sort of like search
- 48:06over, roughly.
- 48:07So, now what we're going to do is
- 48:10here, we are going to run the
- 48:12optimization for 1,000 steps.
- 48:14And instead of using a fixed number, we
- 48:17are going to use learning rate
- 48:19indexing into here, lrs of I, and make
- 48:22this I.
- 48:25So, basically, let me reset this to be,
- 48:28again, starting from random,
- 48:30creating these learning rates between
- 48:32negative um 0.0 between 0.001 and um
- 48:36one, but exponentially stepped.
- 48:39And here what we're doing is we're
- 48:41iterating a thousand times. We're going
- 48:43to use the learning rate
- 48:45um that's in the beginning very, very
- 48:47low. In the beginning, it's going to be
- 48:490.001, but by the end, it's going to be
- 48:52one.
- 48:53And then we're going to step with that
- 48:55learning rate.
- 48:57And now what we want to do is we want to
- 48:58keep track of the uh um
- 49:04learning rates that we used, and we want
- 49:06to look at the losses that resulted.
- 49:09And so here, let me
- 49:12track stats.
- 49:14So, lri.append lr,
- 49:16and um
- 49:18loss i.append
- 49:20loss.item.
- 49:22Okay.
- 49:23So, again, reset everything,
- 49:27and then run.
- 49:30And so, basically, we started with a
- 49:31very low learning rate, and we went all
- 49:33the way up to uh learning rate of
- 49:34negative one. And now what we can do is
- 49:37we can plt.plot,
- 49:39and we can plot the two. So, we can plot
- 49:41the learning rates on the x-axis, and
- 49:43the losses we saw on the y-axis.
- 49:46And often, you're going to find that
- 49:47your plot looks something like this,
- 49:50where in the beginning, um
- 49:52you had very low learning rates. We
- 49:53basically anything
- 49:55barely anything happened.
- 49:56Then we got to like a nice spot here.
- 50:00And then as we increase the learning
- 50:01rate enough, uh, we basically started to
- 50:03be kind of unstable here.
- 50:05So a good learning rate turns out to be
- 50:07somewhere around here.
- 50:10Um, and because we have LRI here,
- 50:13um,
- 50:14we actually may want to, um,
- 50:19do not LR, uh,
- 50:21not the learning rate, but the exponent.
- 50:22So that would be the LRE at I is maybe
- 50:25what we want to log. So let me reset
- 50:27this and redo that calculation.
- 50:30But now on the x-axis we have the, um,
- 50:34exponent of the learning rate. And so we
- 50:36can see the exponent of the learning
- 50:37rate that is good to use. It would be
- 50:38sort of like roughly in the valley here,
- 50:41because here the learning rates are just
- 50:42way too low. And then here where we
- 50:44expect relatively good learning rate
- 50:46somewhere here. And then here things are
- 50:47starting to explode. So somewhere around
- 50:50-1 as the exponent of the learning rate
- 50:52is a pretty good setting. And 10 to the
- 50:55-1 is .1. So .1 is actually .1 was
- 50:59actually a fairly good learning rate
- 51:01around here.
- 51:02And that's what we had in the initial
- 51:03setting.
- 51:04Um, but that's roughly how you would
- 51:06determine it.
- 51:07And so here now we can take out the
- 51:10tracking of these.
- 51:12And we can just simply set LR to be 10
- 51:15to the -1 or
- 51:18basically otherwise .1 as it was before.
- 51:20And now we have some confidence that
- 51:21this is actually a fairly good learning
- 51:23rate.
- 51:24And so now we can do is we can crank up
- 51:26the iterations.
- 51:27We can reset our optimization.
- 51:30And, um,
- 51:32we can run for a pretty long time using
- 51:34this learning rate.
- 51:36Oops, and we don't want to print. That's
- 51:38way too much printing.
- 51:40So let me again reset.
- 51:42And run 10,000 steps.
- 51:48Okay, so we're at 0.2 2.48 roughly.
- 51:51Let's run another 10,000 steps.
- 51:582.46
- 52:00And now let's do one learning rate
- 52:01decay. What this means is we're going to
- 52:03take our learning rate and we're going
- 52:04to 10x lower it. And so we're at the
- 52:07late stages of training potentially and
- 52:10we may want to go uh a bit slower.
- 52:12Let's do one more actually at 0.1 just
- 52:14to see if
- 52:17we're making a dent here.
- 52:19Okay, we're still making dent. And by
- 52:20the way, the bigram loss that we
- 52:23achieved last video was 2.45. So we've
- 52:26already surpassed the bigram model.
- 52:29And once I get a sense that this is
- 52:30actually kind of starting to plateau
- 52:31off, uh people like to do as I mentioned
- 52:33this learning rate decay. So let's try
- 52:36to decay the loss,
- 52:37uh the learning rate I mean.
- 52:42And we achieve at about 2.3 now.
- 52:46Obviously, this is janky and not exactly
- 52:48how you would train it in production,
- 52:50but this is roughly what you're going
- 52:51through. You first find a decent
- 52:53learning rate using the approach that I
- 52:54showed you.
- 52:55Then you start with that learning rate
- 52:57and you train for a while.
- 52:58And then at the end people like to do a
- 53:00learning rate decay where you decay the
- 53:02learning rate by say a factor of 10 and
- 53:04you do a few more steps and then you get
- 53:06a trained network, roughly speaking.
- 53:08So we've achieved 2.3 and uh
- 53:10dramatically improved on the bigram
- 53:12language model using this simple neural
- 53:14net as described here
- 53:16um using these 3,400 parameters.
- 53:20Now there's something we have to be
- 53:20careful with.
- 53:22I said that we have a better model
- 53:24because we are achieving a lower loss,
- 53:262.3, much lower than 2.45 with the
- 53:28bigram model previously.
- 53:30Now that's not exactly true. And the
- 53:32reason that's not true is that
- 53:37this is actually fairly small model, but
- 53:39these models can get larger and larger
- 53:41if you keep adding neurons and
- 53:42parameters. So, you can imagine that we
- 53:44don't potentially have a thousand
- 53:46parameters. We could have 10,000 or
- 53:47100,000 or millions of parameters.
- 53:50And as the capacity of the neural
- 53:51network grows,
- 53:52it becomes more and more capable of
- 53:55overfitting your training set.
- 53:57What that means is that the loss on the
- 53:59training set, on the data that you're
- 54:00training on, will become very, very low,
- 54:03as low as zero.
- 54:04But all that the model is doing is
- 54:06memorizing your training set verbatim.
- 54:09So, if you take that model, and it looks
- 54:10like it's working really well, but you
- 54:12try to sample from it, you will
- 54:13basically only get examples exactly as
- 54:15they are in the training set. You won't
- 54:17get any new data.
- 54:19In addition to that, if you try to
- 54:20evaluate the loss on some withheld names
- 54:23or other words, you will actually see
- 54:25that the loss on those can be very high.
- 54:28As a basically, it's not a good model.
- 54:30So, the standard in the field it is is
- 54:32to split up your data set into three
- 54:34splits, as we call them. We have the
- 54:36training split, the dev split or the
- 54:38validation split,
- 54:40and the test split.
- 54:42So, training split,
- 54:45test or um sorry, dev or validation
- 54:48split,
- 54:49and test split.
- 54:52And typically, this would be say 80% of
- 54:54your data set, this could be 10% and
- 54:56this 10% roughly.
- 54:58So, you have these three splits of the
- 55:00data.
- 55:01Now, these 80% of your training of of
- 55:03the data set, the training set, is used
- 55:05to optimize the parameters of the model,
- 55:07just like we're doing here, using
- 55:08gradient descent.
- 55:10These 10% of the um examples, the dev or
- 55:13validation split, they're used for
- 55:15development over all the hyper
- 55:17parameters of your model. So, hyper
- 55:19parameters are for example, the size of
- 55:21this hidden layer,
- 55:22the size of the embedding. So, this is a
- 55:24hundred or a two for us, but we could
- 55:26try different things.
- 55:28The strength of the regularization,
- 55:29which we aren't using yet so far.
- 55:32So, there's lots of different hyper
- 55:33parameters and settings that go into
- 55:34defining a neural net, and you can try
- 55:36many different variations of them and
- 55:38see whichever one works best on your
- 55:41validation split.
- 55:43So, this is used to train the
- 55:44parameters.
- 55:45This is used to train the hyper
- 55:47parameters.
- 55:48And test split is used to evaluate uh
- 55:51basically the performance of the model
- 55:53at the end.
- 55:54So, we're only evaluating the loss on
- 55:55the test split very, very sparingly and
- 55:57very few times because every single time
- 56:00you evaluate your test loss and you
- 56:02learn something from it,
- 56:04you are basically starting to also train
- 56:06on the test split. So, you are only
- 56:09allowed to test the loss on the test set
- 56:12um very, very few times. Otherwise, you
- 56:15risk overfitting to it as well as you
- 56:17experiment on your model.
- 56:19So, let's also split up our training
- 56:21data into train, dev, and test. And
- 56:24then, we are going to train on train and
- 56:26only evaluate on test very, very
- 56:28sparingly.
- 56:29Okay, so here we go.
- 56:31Here is where we took all the words and
- 56:33put them into X and Y tensors.
- 56:36So, instead, let me create a new cell
- 56:37here and let me just copy-paste some
- 56:39code here
- 56:41because I don't think it's that
- 56:43complex, but um
- 56:45we're going to try to save a little bit
- 56:46of time.
- 56:47I'm converting this to be a function
- 56:49now. And this function takes some list
- 56:51of words and builds the arrays X and Y
- 56:54for those words only.
- 56:56And then here, I am shuffling up all the
- 56:59words. So, these are the input words
- 57:01that we get.
- 57:02We are randomly shuffling them all up.
- 57:05And then, um
- 57:06we're going to
- 57:08set N1 to be
- 57:10the number of examples, that's 80% of
- 57:11the words, and N2 to be 90% of the
- 57:14weight of the words.
- 57:16So, basically, if length of words is
- 57:1830,000, N1 is
- 57:21Oh, sorry. I should probably run this.
- 57:24N1 is 25,000 and N2 is 28,000.
- 57:28And so, here we see that I'm calling
- 57:30build data set to build the training set
- 57:32X and Y
- 57:33by indexing into up to N1. So, we're
- 57:36going to have only 25,000 training
- 57:38words.
- 57:39And then we're going to have um
- 57:42roughly
- 57:44N2 minus N1
- 57:463,000 validation examples or dev
- 57:49examples. And we're going to have
- 57:52um
- 57:53len of words basically minus N2
- 57:57or 3,204
- 57:59examples
- 58:01here for the test set.
- 58:03So,
- 58:04now we have X's and Y's for all those
- 58:07three splits.
- 58:11Um
- 58:13Oh, yeah. I'm printing their size here
- 58:14inside the function as well.
- 58:19But here we don't have words, but these
- 58:20are already the individual examples made
- 58:22from those words.
- 58:25So, let's now scroll down here.
- 58:27And the data set now for training is
- 58:31more like this.
- 58:33And then when we reset the network,
- 58:38when we're training, we're only going to
- 58:40be training using X train,
- 58:43X train, and Y train.
- 58:47So, that's the only thing we're training
- 58:49on.
- 58:57Let's see where we are on the
- 59:00single batch.
- 59:02Let's now train maybe a few more steps.
- 59:08Training a neural network can take a
- 59:09while. Usually, you don't do it in line.
- 59:11You launch a bunch of jobs and you wait
- 59:12for them to finish.
- 59:13Um
- 59:14can take in multiple days and so on.
- 59:16Luckily, this is a very small network.
- 59:21Okay, so the loss is pretty good. Oh, we
- 59:24accidentally used a learning rate that
- 59:25is way too low.
- 59:27So, let me actually come back.
- 59:29We use the decay learning rate of 0.01.
- 59:35So, this will train much faster.
- 59:37And then here, when we evaluate, uh
- 59:39let's use the dev set here.
- 59:42X dev
- 59:43and Y dev to evaluate the loss.
- 59:47Okay.
- 59:48Uh
- 59:48and let's not decay the learning rate
- 59:50and only do say 10,000 examples.
- 59:55And let's evaluate the dev loss once
- 59:58here.
- 59:59Okay, so we're getting about 2.3 on dev.
- 1:00:01And so, the neural network when it was
- 1:00:02training did not see these dev examples.
- 1:00:05It hasn't optimized on them. And yet,
- 1:00:08when we evaluate the loss on these dev,
- 1:00:10we actually get a pretty decent loss.
- 1:00:12And so, we can also look at what the
- 1:00:16loss is on all of training set.
- 1:00:19Oops.
- 1:00:20And so, we see that the training and the
- 1:00:22dev loss are about equal. So, we're not
- 1:00:24over- fitting. Um this model is not
- 1:00:27powerful enough to just be purely
- 1:00:29memorizing the data. And so far, we are
- 1:00:32what's called under-fitting because the
- 1:00:34training loss and the dev or test losses
- 1:00:36are roughly equal. So, what that
- 1:00:38typically means is that our network is
- 1:00:40very tiny, very small. And we expect to
- 1:00:43make uh performance improvements by
- 1:00:45scaling up the size of this neural net.
- 1:00:47So, let's do that now. So, let's come
- 1:00:49over here
- 1:00:50and let's increase the size of the
- 1:00:51neural net.
- 1:00:52The easiest way to do this is we can
- 1:00:54come here to the hidden layer, which
- 1:00:55currently has 100 neurons, and let's
- 1:00:57just pump this up. So, let's do 300
- 1:00:59neurons.
- 1:01:00And then, this is also 300 biases. And
- 1:01:03here we have 300 inputs into the final
- 1:01:05layer.
- 1:01:07So,
- 1:01:08let's initialize our neural net. We now
- 1:01:10have 10,000 10,000 parameters instead of
- 1:01:133,000 parameters.
- 1:01:15And then, we're not using this.
- 1:01:18And then here, what I'd like to do is
- 1:01:19I'd like to actually uh keep track of uh
- 1:01:23that
- 1:01:24um
- 1:01:27Okay, let's just do this. Let's keep
- 1:01:29stats again.
- 1:01:30And here when we're keeping track of the
- 1:01:34loss, let's just also keep track of the
- 1:01:37steps. And let's just have a eye here.
- 1:01:41And let's train on 30,000.
- 1:01:44Or rather say
- 1:01:46Okay, let's try 30,000.
- 1:01:48And we are at 0.1.
- 1:01:51And
- 1:01:52we should be able to run this.
- 1:01:55And optimize neural net.
- 1:01:57And then here basically I want to
- 1:01:59PLT.plot
- 1:02:01the steps
- 1:02:02against the loss.
- 1:02:09So these are the x's and y's.
- 1:02:11And this is
- 1:02:13the loss function and how it's being
- 1:02:15optimized.
- 1:02:16Now you see that there's quite a bit of
- 1:02:18thickness to this, and that's because we
- 1:02:19are optimizing over these mini batches.
- 1:02:21And the mini batches create a little bit
- 1:02:23of noise in this.
- 1:02:25Uh where are we in the dev set? We are
- 1:02:27at 2.5. So we still haven't optimized
- 1:02:30this neural net very well.
- 1:02:32And that's probably because we made it
- 1:02:33bigger. It might take longer for this
- 1:02:34neural net to converge.
- 1:02:36Um
- 1:02:37and so let's continue training.
- 1:02:40Um
- 1:02:42yeah, let's just continue training.
- 1:02:46One possibility is that the batch size
- 1:02:48is so low
- 1:02:49that we just have way too much noise in
- 1:02:51the training. And we may want to
- 1:02:53increase the batch size so that we have
- 1:02:54a bit more
- 1:02:56correct gradient and we're not thrashing
- 1:02:58too much. And we can actually like
- 1:03:00optimize more properly.
- 1:03:08Okay.
- 1:03:08Uh this will now become meaningless
- 1:03:10because we've reinitialized these. So
- 1:03:13yeah, this looks not pleasing like now.
- 1:03:16But there probably is like a tiny
- 1:03:17improvement, but it's so hard to tell.
- 1:03:20Uh let's go again.
- 1:03:222.52
- 1:03:25Let's try to decrease the learning rate
- 1:03:27by a factor of two.
- 1:03:50Okay, we're at 2.32.
- 1:03:52Let's continue training.
- 1:04:05We basically expect to see a lower loss
- 1:04:07than what we had before, because now we
- 1:04:09have a much, much bigger model, and we
- 1:04:10were underfitting. So, we'd expect that
- 1:04:12increasing the size of the model should
- 1:04:14help the neural net.
- 1:04:162.32 Okay, so that's not happening too
- 1:04:18well.
- 1:04:19Now, one other concern is that even
- 1:04:21though we've made the tanh layer here,
- 1:04:23uh the hidden layer, much, much bigger,
- 1:04:25it could be that the bottleneck of the
- 1:04:26network right now are these embeddings
- 1:04:28that are two dimensional. It can be that
- 1:04:30we're just cramming way too many
- 1:04:32characters into just two dimensions, and
- 1:04:34the neural net is not able to really use
- 1:04:36that space effectively, and that that is
- 1:04:38sort of like the bottleneck to our
- 1:04:39network's performance.
- 1:04:42Okay, 2.23. So, just by decreasing the
- 1:04:45learning rate, I was able to make quite
- 1:04:46a bit of progress. Let's run this one
- 1:04:47more time.
- 1:04:51And then evaluate the training and the
- 1:04:53dev loss.
- 1:04:56Now, one more thing after training that
- 1:04:58I'd like to do is I'd like to visualize
- 1:05:00the um
- 1:05:02embedding vectors for these um
- 1:05:05characters before we scale up uh the
- 1:05:08embedding size from two.
- 1:05:09Because we'd like to make uh this
- 1:05:11bottleneck potentially go away.
- 1:05:13But once I make this greater than two,
- 1:05:15we won't be able to visualize them.
- 1:05:17So here, okay, we're at 2.23 and 2.24.
- 1:05:21So um
- 1:05:22we're not improving much more, and maybe
- 1:05:24the bottleneck now is the character
- 1:05:25embedding size, which is two.
- 1:05:28So here I have a bunch of code that will
- 1:05:29create a figure,
- 1:05:31and then we're going to visualize
- 1:05:34the embeddings that were trained by the
- 1:05:35neural net
- 1:05:36on these characters. Because right now
- 1:05:38the embedding size is just two, so we
- 1:05:40can visualize all the characters with
- 1:05:41the X and the Y coordinates as the two
- 1:05:44embedding locations for each of these
- 1:05:46characters.
- 1:05:47And so here are the X coordinates and
- 1:05:50the Y coordinates, which are the columns
- 1:05:51of C.
- 1:05:52And then for each one, I also include
- 1:05:55the text of the little character.
- 1:05:58So here what we see is actually kind of
- 1:05:59interesting.
- 1:06:01Um
- 1:06:02the network has basically learned to
- 1:06:04separate out the characters and cluster
- 1:06:06them a little bit. Uh so for example,
- 1:06:08you see how the vowels
- 1:06:09a e i o u are clustered up here.
- 1:06:12So what that's telling us that is that
- 1:06:14the neural net treats these as very
- 1:06:15similar, right? Because when they feed
- 1:06:17into the neural net, the embedding
- 1:06:19uh for all of these characters is very
- 1:06:21similar. And so the neural net thinks
- 1:06:23that they're very similar and kind of
- 1:06:24like interchangeable, if that makes
- 1:06:26sense.
- 1:06:27Um
- 1:06:29then the the points that are like really
- 1:06:31far away are for example Q. Q is kind of
- 1:06:33treated as an exception, and Q has a
- 1:06:35very special embedding vector, so to
- 1:06:37speak.
- 1:06:38Similarly, dot, which is a special
- 1:06:40character, is all the way out here.
- 1:06:42And a lot of the other letters are sort
- 1:06:44of like clustered up here.
- 1:06:46And so it's kind of interesting that
- 1:06:47there's a little bit of structure here
- 1:06:49um after the training,
- 1:06:51and it's not definitely not random, and
- 1:06:53these embeddings make sense.
- 1:06:56So we're now going to scale up the
- 1:06:57embedding size and won't be able to
- 1:06:59visualize it directly, but we expect
- 1:07:01that because we're underfitting,
- 1:07:03and we made this layer much bigger and
- 1:07:06did not sufficiently improve the loss.
- 1:07:08We're thinking that the
- 1:07:09um
- 1:07:10constraint to better performance right
- 1:07:12now could be these embedding vectors.
- 1:07:15So, let's make them bigger. Okay, so
- 1:07:16let's scroll up here.
- 1:07:18And now we don't have two-dimensional
- 1:07:19embeddings, we are going to have
- 1:07:21say 10-dimensional embeddings for each
- 1:07:23word.
- 1:07:25Then
- 1:07:26this layer will receive 3 * 10, so 30
- 1:07:30inputs
- 1:07:31will go into
- 1:07:33um
- 1:07:33the hidden layer.
- 1:07:35Let's also make the hidden layer a bit
- 1:07:37smaller. So, instead of 300, let's just
- 1:07:38do 200 neurons in that hidden layer.
- 1:07:41So, now the total number of elements
- 1:07:43will be slightly bigger at 11,000.
- 1:07:47And then we here we have to be a bit
- 1:07:48careful because um
- 1:07:50okay, the learning rate we set to 0.1.
- 1:07:53Here we are hardcoding six and
- 1:07:56obviously, if you're working in
- 1:07:56production, you don't want to be
- 1:07:57hardcoding magic numbers. But instead of
- 1:08:00six, this should now be 30.
- 1:08:02Um
- 1:08:04and let's run it for 50,000 iterations
- 1:08:06and let me split out the initialization
- 1:08:08here outside
- 1:08:10so that when we run this cell multiple
- 1:08:12times, it's not going to wipe out our
- 1:08:14loss.
- 1:08:17In addition to that,
- 1:08:19here
- 1:08:20let's instead of logging loss.item,
- 1:08:22let's actually uh log the
- 1:08:25let's um do log 10,
- 1:08:28I believe that's a function of the loss.
- 1:08:32And I'll show you why in a second. Let's
- 1:08:34optimize this.
- 1:08:37Basically, I'd like to plot the log loss
- 1:08:39instead of the loss because when you
- 1:08:40plot the loss, many times it can have
- 1:08:42this hockey stick appearance and log
- 1:08:44squashes it in.
- 1:08:46Uh so, it just kind of like looks nicer.
- 1:08:49So, the x-axis is step I
- 1:08:51and the y-axis will be the loss I.
- 1:09:00And then here, this is 30.
- 1:09:03Ideally, we wouldn't be hard coding
- 1:09:04these.
- 1:09:08Okay, so let's look at the loss.
- 1:09:11Okay? It's again very thick because the
- 1:09:13mini batch size is very small, but the
- 1:09:15total loss over the training set is 2.3
- 1:09:18and the the test or the dev set is 2.38
- 1:09:20as well.
- 1:09:22So so far so good. Uh let's try to now
- 1:09:24decrease the learning rate
- 1:09:25by a factor of 10
- 1:09:29and train for another 50,000 iterations.
- 1:09:35We'd hope that we would be able to beat
- 1:09:37uh 2.32.
- 1:09:43But again, we're just kind of like doing
- 1:09:44this very haphazardly, so I don't
- 1:09:46actually have confidence that our
- 1:09:48learning rate is set very well, that our
- 1:09:50learning rate decay, which we just do at
- 1:09:53random, is set very well.
- 1:09:55And um so the optimization here is kind
- 1:09:57of suspect, to be honest, and this is
- 1:09:59not how you would do it typically in
- 1:10:00production. In production, you would
- 1:10:02create parameters or hyper parameters
- 1:10:04out of all these settings, and then you
- 1:10:05would run lots of experiments and see
- 1:10:07whichever ones are working well for you.
- 1:10:11Okay.
- 1:10:12So we have 2.17 now and 2.2. Okay. So
- 1:10:16you see how the training and the
- 1:10:18validation performance are starting to
- 1:10:20slightly slowly depart.
- 1:10:23So maybe we're getting the sense that
- 1:10:24the neural net is getting good enough or
- 1:10:28that number of parameters is large
- 1:10:30enough that we are slowly starting to
- 1:10:32overfit.
- 1:10:34Uh let's maybe run one more iteration of
- 1:10:36this
- 1:10:38and see where we get.
- 1:10:41But yeah, basically, you would be
- 1:10:43running lots of experiments, and then
- 1:10:44you are slowly scrutinizing whichever
- 1:10:46ones give you the best dev performance.
- 1:10:48And then once you find all the uh hyper
- 1:10:50parameters that make your dev
- 1:10:51performance good, you take that model
- 1:10:53and you evaluate the test set
- 1:10:55performance a single time. And that's
- 1:10:57the number that you report in your paper
- 1:10:59or wherever else you want to talk about
- 1:11:00and brag about your model.
- 1:11:05So, let's then rerun the plot and rerun
- 1:11:08the train and dev.
- 1:11:11And because we're getting lower loss
- 1:11:12now, it is the case that the embedding
- 1:11:14size of these was holding us back very
- 1:11:17likely.
- 1:11:20Okay, so 2.16 2.19 is what we're roughly
- 1:11:22getting.
- 1:11:24So, there's many ways to go from many
- 1:11:26ways to go from here. We can continue
- 1:11:28tuning the optimization.
- 1:11:30We can continue, for example, playing
- 1:11:32with the sizes of the neural net. Or we
- 1:11:33can increase the number of um
- 1:11:36words or characters, in our case, that
- 1:11:38we are taking as an input. So, instead
- 1:11:39of just three characters, we could be
- 1:11:40taking more characters than as an input.
- 1:11:43And that could further improve the loss.
- 1:11:46Okay, so I changed the code slightly, so
- 1:11:48we have here 200,000 steps of the
- 1:11:50optimization. And in the first 100,000,
- 1:11:52we're using a learning rate of 0.1. And
- 1:11:54then in the next 100,000, we're using a
- 1:11:56learning rate of 0.01.
- 1:11:58This is the loss that I achieve.
- 1:12:00And these are the performance on the
- 1:12:01training and validation loss.
- 1:12:04And in particular, the best validation
- 1:12:05loss I've been able to obtain in the
- 1:12:07last 30 minutes or so is 2.17.
- 1:12:10So, now I invite you to beat this
- 1:12:12number. And you have quite a few knobs
- 1:12:14available to you to, I think, surpass
- 1:12:15this number.
- 1:12:17So, number one, you can of course change
- 1:12:18the number of neurons in the hidden
- 1:12:20layer of this model.
- 1:12:21You can change the dimensionality of the
- 1:12:23embedding uh lookup table.
- 1:12:25You can change the number of characters
- 1:12:26that are feeding in as an input um as
- 1:12:29the context into this model.
- 1:12:32And then, of course, you can change the
- 1:12:33details of the optimization. How long
- 1:12:35are we running? What is the learning
- 1:12:37rate? How does it change over time?
- 1:12:39Uh how does it decay?
- 1:12:41Uh you can change the batch size, and
- 1:12:42you may be able to actually achieve a
- 1:12:44much better convergence speed in terms
- 1:12:46of uh how many seconds or minutes it
- 1:12:48takes to train the model and uh get uh
- 1:12:51your result in terms of really good um
- 1:12:54loss.
- 1:12:55And then of course, I actually invite
- 1:12:57you to read this paper. It is 19 pages,
- 1:12:59but at this point you should actually be
- 1:13:00able to read a good chunk of this paper
- 1:13:03and understand
- 1:13:04uh pretty good chunks of it.
- 1:13:06And this paper also has quite a few
- 1:13:08ideas for improvements that you can play
- 1:13:09with.
- 1:13:11So, all of those are knobs available to
- 1:13:12you, and you should be able to beat this
- 1:13:14number. I'm leaving that as an exercise
- 1:13:16to the reader, and uh that's it for now,
- 1:13:18and I'll see you next time.
- 1:13:24Before we wrap up, I also wanted to show
- 1:13:25how you would sample from the model.
- 1:13:28So, we're going to generate 20 samples.
- 1:13:31At first, we begin with all dots, so
- 1:13:33that's the context. And then until we
- 1:13:36generate the zeroth character again,
- 1:13:40we're going to embed the current context
- 1:13:43using the embedding table C.
- 1:13:46Now, usually uh here, the first
- 1:13:48dimension was the size of the training
- 1:13:50set, but here we're only working with a
- 1:13:51single example that we're generating, so
- 1:13:53this is just um dimension one, just for
- 1:13:55simplicity.
- 1:13:57Um and so this embedding then gets
- 1:14:00projected into the hidden state. You get
- 1:14:02the logits. Now, we calculate the
- 1:14:04probabilities. For that, you can use
- 1:14:06f.softmax
- 1:14:08um of logits, and that just basically
- 1:14:10exponentiates the logits and makes them
- 1:14:12sum to one.
- 1:14:13And similar to cross entropy, it is
- 1:14:15careful that there's no overflows.
- 1:14:18Once we have the probabilities, we
- 1:14:19sample from them using torch.multinomial
- 1:14:22to get our next index, and then we shift
- 1:14:24the context window to append the index
- 1:14:26and record it.
- 1:14:28And then we can just um decode all the
- 1:14:30integers to strings and print them out.
- 1:14:33And so these are some example samples,
- 1:14:35and you can see that the model now works
- 1:14:36much better. So, the words here are much
- 1:14:39more word-like or name-like. So, we have
- 1:14:41things like ham, um
- 1:14:44joes,
- 1:14:46uh lela,
- 1:14:48you know, it's starting to sound a
- 1:14:49little bit more name-like. So, we're
- 1:14:51definitely making progress, uh, but we
- 1:14:52can still improve on this model quite a
- 1:14:54lot.
- 1:14:55Okay, sorry. There's some bonus content.
- 1:14:57I wanted to mention that I want to make
- 1:14:59these notebooks more accessible. And so,
- 1:15:01I don't want you to have to, like,
- 1:15:03install Jupyter notebooks and torch and
- 1:15:04everything else. So, I will be sharing a
- 1:15:06link to a Google Colab.
- 1:15:09And the Google Colab will look like a
- 1:15:10notebook in your browser. And you can
- 1:15:13just go to a URL, and you'll be able to
- 1:15:15execute all the code that you saw in the
- 1:15:17Google Colab. And so, this is me
- 1:15:20executing the code in this lecture, and
- 1:15:22I shortened it a little bit.
- 1:15:23Uh, but basically, you you're able to
- 1:15:25train the exact same network, and then
- 1:15:27plot and sample from the model, and
- 1:15:29everything is ready for you to, like,
- 1:15:30tinker with the numbers right there in
- 1:15:32your browser. No installation necessary.
- 1:15:34Um, so, I just wanted to point that out,
- 1:15:36and the link to this will be in the
- 1:15:37video description.
About this transcript
This page contains the full transcript of Building makemore Part 2: MLP by Andrej Karpathy, generated from the public captions YouTube serves with the video. The transcript has 12,731 words across 2,147 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.