Let's build GPT: from scratch, in code, spelled out. — Transcript
Full transcript
- 0:00hi everyone so by now you have probably
- 0:02heard of chat GPT it has taken the world
- 0:04and AI Community by storm and it is a
- 0:07system that allows you to interact with
- 0:09an AI and give it text based tasks so
- 0:12for example we can ask chat GPT to write
- 0:15us a small Hau about how important it is
- 0:16that people understand Ai and then they
- 0:18can use it to improve the world and make
- 0:20it more prosperous so when we run this
- 0:23AI knowledge brings prosperity for all
- 0:25to see Embrace its
- 0:27power okay not bad and so you could see
- 0:29that chpt went from left to right and
- 0:32generated all these words SE sort of
- 0:35sequentially now I asked it already the
- 0:37exact same prompt a little bit earlier
- 0:39and it generated a slightly different
- 0:41outcome ai's power to grow ignorance
- 0:44holds us back learn Prosperity weights
- 0:47so uh pretty good in both cases and
- 0:49slightly different so you can see that
- 0:50chat GPT is a probabilistic system and
- 0:52for any one prompt it can give us
- 0:54multiple answers sort of uh replying to
- 0:57it now this is just one example of a
- 0:59problem people have come up with many
- 1:01many examples and there are entire
- 1:03websites that index interactions with
- 1:06chpt and so many of them are quite
- 1:08humorous explain HTML to me like I'm a
- 1:10dog uh write release notes for chess 2
- 1:14write a note about Elon Musk buying a
- 1:16Twitter and so on so as an example uh
- 1:20please write a breaking news article
- 1:21about a leaf falling from a
- 1:23tree uh and a shocking turn of events a
- 1:26leaf has fallen from a tree in the local
- 1:28park Witnesses report that the leaf
- 1:30which was previously attached to a
- 1:31branch of a tree attached itself and
- 1:33fell to the ground very dramatic so you
- 1:36can see that this is a pretty remarkable
- 1:37system and it is what we call a language
- 1:40model uh because it um it models the
- 1:43sequence of words or characters or
- 1:46tokens more generally and it knows how
- 1:49sort of words follow each other in
- 1:50English language and so from its
- 1:52perspective what it is doing is it is
- 1:55completing the sequence so I give it the
- 1:57start of a sequence and it completes the
- 2:00sequence with the outcome and so it's a
- 2:02language model in that sense now I would
- 2:05like to focus on the under the hood of
- 2:07um under the hood components of what
- 2:09makes CH GPT work so what is the neural
- 2:12network under the hood that models the
- 2:14sequence of these words and that comes
- 2:17from this paper called attention is all
- 2:19you need in 2017 a landmark paper a
- 2:23landmark paper in AI that produced and
- 2:25proposed the Transformer
- 2:27architecture so GPT is uh short for
- 2:31generally generatively pre-trained
- 2:33Transformer so Transformer is the neuron
- 2:35nut that actually does all the heavy
- 2:36lifting under the hood it comes from
- 2:39this paper in 2017 now if you read this
- 2:41paper this uh reads like a pretty random
- 2:44machine translation paper and that's
- 2:46because I think the authors didn't fully
- 2:47anticipate the impact that the
- 2:49Transformer would have on the field and
- 2:51this architecture that they produced in
- 2:52the context of machine translation in
- 2:54their case actually ended up taking over
- 2:57uh the rest of AI in the next 5 years
- 3:00after and so this architecture with
- 3:02minor changes was copy pasted into a
- 3:05huge amount of applications in AI in
- 3:07more recent years and that includes at
- 3:10the core of chat GPT now we are not
- 3:13going to what I'd like to do now is I'd
- 3:15like to build out something like chat
- 3:17GPT but uh we're not going to be able to
- 3:19of course reproduce chat GPT this is a
- 3:21very serious production grade system it
- 3:23is trained on uh a good chunk of
- 3:26internet and then there's a lot of uh
- 3:29pre-training and fine-tuning stages to
- 3:31it and so it's very complicated what I'd
- 3:33like to focus on is just to train a
- 3:36Transformer based language model and in
- 3:38our case it's going to be a character
- 3:40level language model I still think that
- 3:43is uh very educational with respect to
- 3:45how these systems work so I don't want
- 3:47to train on the chunk of Internet we
- 3:48need a smaller data set in this case I
- 3:51propose that we work with uh my favorite
- 3:53toy data set it's called tiny
- 3:55Shakespeare and um what it is is
- 3:57basically it's a concatenation of all of
- 3:59the works of sh Shakespeare in my
- 4:00understanding and so this is all of
- 4:02Shakespeare in a single file uh this
- 4:05file is about 1 megab and it's just all
- 4:07of
- 4:08Shakespeare and what we are going to do
- 4:10now is we're going to basically model
- 4:12how these characters uh follow each
- 4:14other so for example given a chunk of
- 4:16these characters like this uh given some
- 4:19context of characters in the past the
- 4:22Transformer neural network will look at
- 4:24the characters that I've highlighted and
- 4:26is going to predict that g is likely to
- 4:28come next in the sequence and it's going
- 4:30to do that because we're going to train
- 4:31that Transformer on Shakespeare and it's
- 4:34just going to try to produce uh
- 4:36character sequences that look like this
- 4:39and in that process is going to model
- 4:40all the patterns inside this data so
- 4:43once we've trained the system i' just
- 4:45like to give you a preview we can
- 4:47generate infinite Shakespeare and of
- 4:49course it's a fake thing that looks kind
- 4:51of like
- 4:53Shakespeare
- 4:55um apologies for there's some Jank that
- 4:59I'm not able to resolve in in here but
- 5:02um you can see how this is going
- 5:05character by character and it's kind of
- 5:07like predicting Shakespeare like
- 5:09language so verily my Lord the sites
- 5:12have left the again the king coming with
- 5:15my curses with precious pale and then
- 5:19tranos say something else Etc and this
- 5:21is just coming out of the Transformer in
- 5:23a very similar manner as it would come
- 5:25out in chat GPT in our case character by
- 5:27character in chat GPT uh it's coming out
- 5:31on the token by token level and tokens
- 5:33are these sort of like little subword
- 5:35pieces so they're not Word level they're
- 5:36kind of like word chunk
- 5:38level um and now I've already written
- 5:43this entire code uh to train these
- 5:45Transformers um and it is in a GitHub
- 5:48repository that you can find and it's
- 5:50called nanog
- 5:51GPT so nanog GPT is a repository that
- 5:54you can find in my GitHub and it's a
- 5:56repository for training Transformers um
- 5:59on any given text and what I think is
- 6:02interesting about it because there's
- 6:03many ways to train Transformers but this
- 6:05is a very simple implementation so it's
- 6:06just two files of 300 lines of code each
- 6:10one file defines the GPT model the
- 6:12Transformer and one file trains it on
- 6:14some given Text data set and here I'm
- 6:17showing that if you train it on a open
- 6:18web Text data set which is a fairly
- 6:20large data set of web pages then I
- 6:22reproduce the the performance of
- 6:25gpt2 so gpt2 is an early version of open
- 6:29AI GPT uh from 2017 if I recall
- 6:32correctly and I've only so far
- 6:34reproduced the the smallest 124 million
- 6:36parameter model uh but basically this is
- 6:38just proving that the codebase is
- 6:39correctly arranged and I'm able to load
- 6:42the uh neural network weights that openi
- 6:45has released later so you can take a
- 6:48look at the finished code here in N GPT
- 6:50but what I would like to do in this
- 6:51lecture is I would like to basically uh
- 6:55write this repository from scratch so
- 6:57we're going to begin with an empty file
- 6:59and we're we're going to define a
- 7:00Transformer piece by piece we're going
- 7:03to train it on the tiny Shakespeare data
- 7:05set and we'll see how we can then uh
- 7:08generate infinite Shakespeare and of
- 7:10course this can copy paste to any
- 7:12arbitrary Text data set uh that you like
- 7:14uh but my goal really here is to just
- 7:16make you understand and appreciate uh
- 7:18how under the hood chat GPT works and um
- 7:22really all that's required is a
- 7:24Proficiency in Python and uh some basic
- 7:27understanding of um calculus and
- 7:29statistics
- 7:30and it would help if you also see my
- 7:32previous videos on the same YouTube
- 7:34channel in particular my make more
- 7:35series where I um Define smaller and
- 7:40simpler neural network language models
- 7:42uh so multi perceptrons and so on it
- 7:45really introduces the language modeling
- 7:46framework and then uh here in this video
- 7:49we're going to focus on the Transformer
- 7:50neural network itself okay so I created
- 7:53a new Google collab uh jup notebook here
- 7:57and this will allow me to later easily
- 7:58share this code that we're going to
- 8:00develop together uh with you so you can
- 8:01follow along so this will be in a video
- 8:03description uh later now here I've just
- 8:07done some preliminaries I downloaded the
- 8:09data set the tiny Shakespeare data set
- 8:10at this URL and you can see that it's
- 8:12about a 1 Megabyte file then here I open
- 8:15the input.txt file and just read in all
- 8:17the text of the string and we see that
- 8:20we are working with 1 million characters
- 8:22roughly and the first 1,000 characters
- 8:24if we just print them out are basically
- 8:26what you would expect this is the first
- 8:281,000 characters of the tiny Shakespeare
- 8:30data set roughly up to here so so far so
- 8:34good next we're going to take this text
- 8:37and the text is a sequence of characters
- 8:39in Python so when I call the set
- 8:41Constructor on it I'm just going to get
- 8:44the set of all the characters that occur
- 8:46in this text and then I call list on
- 8:49that to create a list of those
- 8:51characters instead of just a set so that
- 8:53I have an ordering an arbitrary ordering
- 8:56and then I sort that so basically we get
- 8:59just all the characters that occur in
- 9:00the entire data set and they're sorted
- 9:02now the number of them is going to be
- 9:04our vocabulary size these are the
- 9:06possible elements of our sequences and
- 9:09we see that when I print here the
- 9:11characters there's 65 of them in total
- 9:14there's a space character and then all
- 9:16kinds of special characters and then U
- 9:19capitals and lowercase letters so that's
- 9:21our vocabulary and that's the sort of
- 9:23like possible uh characters that the
- 9:25model can see or emit okay so next we
- 9:29will would like to develop some strategy
- 9:31to tokenize the input text now when
- 9:35people say tokenize they mean convert
- 9:36the raw text as a string to some
- 9:39sequence of integers According to some
- 9:41uh notebook According to some vocabulary
- 9:43of possible elements so as an example
- 9:46here we are going to be building a
- 9:48character level language model so we're
- 9:49simply going to be translating
- 9:50individual characters into integers so
- 9:53let me show you uh a chunk of code that
- 9:55sort of does that for us so we're
- 9:57building both the encoder and the
- 9:58decoder
- 10:00and let me just talk through what's
- 10:01happening
- 10:02here when we encode an arbitrary text
- 10:05like hi there we're going to receive a
- 10:08list of integers that represents that
- 10:10string so for example 46 47 Etc and then
- 10:14we also have the reverse mapping so we
- 10:17can take this list and decode it to get
- 10:20back the exact same string so it's
- 10:22really just like a translation to
- 10:24integers and back for arbitrary string
- 10:26and for us it is done on a character
- 10:28level
- 10:30now the way this was achieved is we just
- 10:31iterate over all the characters here and
- 10:34create a lookup table from the character
- 10:35to the integer and vice versa and then
- 10:38to encode some string we simply
- 10:40translate all the characters
- 10:41individually and to decode it back we
- 10:44use the reverse mapping and concatenate
- 10:46all of it now this is only one of many
- 10:49possible encodings or many possible sort
- 10:51of tokenizers and it's a very simple one
- 10:54but there's many other schemas that
- 10:55people have come up with in practice so
- 10:57for example Google uses a sentence
- 10:59piece uh so sentence piece will also
- 11:02encode text into um integers but in a
- 11:05different schema and using a different
- 11:08vocabulary and sentence piece is a
- 11:10subword uh sort of tokenizer and what
- 11:13that means is that um you're not
- 11:15encoding entire words but you're not
- 11:17also encoding individual characters it's
- 11:19it's a subword unit level and that's
- 11:22usually what's adopted in practice for
- 11:24example also openai has this Library
- 11:26called tick token that uses a bite pair
- 11:28encode
- 11:29tokenizer um and that's what GPT uses
- 11:33and you can also just encode words into
- 11:35like hell world into a list of integers
- 11:38so as an example I'm using the Tik token
- 11:40Library here I'm getting the encoding
- 11:43for gpt2 or that was used for gpt2
- 11:46instead of just having 65 possible
- 11:48characters or tokens they have 50,000
- 11:51tokens and so when they encode the exact
- 11:54same string High there we only get a
- 11:57list of three integers but those
- 11:59integers are not between 0 and 64 they
- 12:01are between Z and 5,
- 12:055,256 so basically you can trade off the
- 12:09code book size and the sequence lengths
- 12:12so you can have very long sequences of
- 12:13integers with very small vocabularies or
- 12:16we can have short um sequences of
- 12:20integers with very large vocabularies
- 12:23and so typically people use in practice
- 12:25these subword encodings but I'd like to
- 12:28keep our token ier very simple so we're
- 12:30using character level tokenizer and that
- 12:33means that we have very small code books
- 12:35we have very simple encode and decode
- 12:37functions uh but we do get very long
- 12:40sequences as a result but that's the
- 12:42level at which we're going to stick with
- 12:43this lecture because it's the simplest
- 12:45thing okay so now that we have an
- 12:46encoder and a decoder effectively a
- 12:49tokenizer we can tokenize the entire
- 12:51training set of Shakespeare so here's a
- 12:53chunk of code that does that and I'm
- 12:55going to start to use the pytorch
- 12:56library and specifically the torch.
- 12:58tensor from the pytorch library so we're
- 13:01going to take all of the text in tiny
- 13:03Shakespeare encode it and then wrap it
- 13:05into a torch. tensor to get the data
- 13:08tensor so here's what the data tensor
- 13:10looks like when I look at just the first
- 13:121,000 characters or the 1,000 elements
- 13:14of it so we see that we have a massive
- 13:16sequence of integers and this sequence
- 13:18of integers here is basically an
- 13:20identical translation of the first
- 13:2210,000 characters
- 13:24here so I believe for example that zero
- 13:27is a new line character and maybe one
- 13:29one is a space not 100% sure but from
- 13:32now on the entire data set of text is
- 13:34re-represented as just it's just
- 13:35stretched out as a single very large uh
- 13:38sequence of
- 13:39integers let me do one more thing before
- 13:41we move on here I'd like to separate out
- 13:43our data set into a train and a
- 13:45validation split so in particular we're
- 13:48going to take the first 90% of the data
- 13:51set and consider that to be the training
- 13:52data for the Transformer and we're going
- 13:54to withhold the last 10% at the end of
- 13:56it to be the validation data and this
- 13:59will help us understand to what extent
- 14:01our model is overfitting so we're going
- 14:03to basically hide and keep the
- 14:04validation data on the side because we
- 14:06don't want just a perfect memorization
- 14:08of this exact Shakespeare we want a
- 14:11neural network that sort of creates
- 14:12Shakespeare like uh text and so it
- 14:15should be fairly likely for it to
- 14:17produce the actual like stowed away uh
- 14:21true Shakespeare text um and so we're
- 14:24going to use this to uh get a sense of
- 14:26the overfitting okay so now we would
- 14:28like to start plugging these text
- 14:30sequences or integer sequences into the
- 14:32Transformer so that it can train and
- 14:34learn those patterns now the important
- 14:36thing to realize is we're never going to
- 14:38actually feed entire text into a
- 14:40Transformer all at once that would be
- 14:42computationally very expensive and
- 14:44prohibitive so when we actually train a
- 14:46Transformer on a lot of these data sets
- 14:48we only work with chunks of the data set
- 14:50and when we train the Transformer we
- 14:52basically sample random little chunks
- 14:53out of the training set and train on
- 14:55just chunks at a time and these chunks
- 14:58have basically some kind of a length and
- 15:01some maximum length now the maximum
- 15:04length typically at least in the code I
- 15:06usually write is called block size you
- 15:08can you can uh find it under different
- 15:10names like context length or something
- 15:12like that let's start with the block
- 15:14size of just eight and let me look at
- 15:16the first train data characters the
- 15:18first block size plus one characters
- 15:20I'll explain why plus one in a
- 15:22second so this is the first nine
- 15:24characters in the sequence in the
- 15:27training set now what I'd like to point
- 15:30out is that when you sample a chunk of
- 15:31data like this so say the these nine
- 15:34characters out of the training set this
- 15:36actually has multiple examples packed
- 15:38into it and uh that's because all of
- 15:41these characters follow each other and
- 15:43so what this thing is going to say when
- 15:47we plug it into a Transformer is we're
- 15:49going to actually simultaneously train
- 15:50it to make prediction at every one of
- 15:52these
- 15:53positions now in the in a chunk of nine
- 15:56characters there's actually eight indiv
- 15:58ual examples packed in there so there's
- 16:01the example that when 18 when in the
- 16:04context of 18 47 likely comes next in a
- 16:08context of 18 and 47 56 comes next in a
- 16:12context of 18 47 56 57 can come next and
- 16:16so on so that's the eight individual
- 16:18examples let me actually spell it out
- 16:20with
- 16:21code so here's a chunk of code to
- 16:24illustrate X are the inputs to the
- 16:26Transformer it will just be the first
- 16:28block size characters y will be the uh
- 16:32next block size characters so it's
- 16:34offset by one and that's because y are
- 16:37the targets for each position in the
- 16:40input and then here I'm iterating over
- 16:42all the block size of eight and the
- 16:45context is always all the characters in
- 16:47x uh up to T and including T and the
- 16:51target is always the teth character but
- 16:53in the targets array y so let me just
- 16:56run
- 16:57this and basically it spells out what I
- 16:59said in words uh these are the eight
- 17:02examples hidden in a chunk of nine
- 17:04characters that we uh sampled from the
- 17:08training set I want to mention one more
- 17:11thing we train on all the eight examples
- 17:14here with context between one all the
- 17:16way up to context of block size and we
- 17:19train on that not just for computational
- 17:20reasons because we happen to have the
- 17:22sequence already or something like that
- 17:23it's not just done for efficiency it's
- 17:26also done um to make the Transformer
- 17:28Network be used to seeing contexts all
- 17:32the way from as little as one all the
- 17:33way to block size and we'd like the
- 17:36transform to be used to seeing
- 17:38everything in between and that's going
- 17:39to be useful later during inference
- 17:41because while we're sampling we can
- 17:43start the sampling generation with as
- 17:45little as one character of context and
- 17:47the Transformer knows how to predict the
- 17:49next character with all the way up to
- 17:51just context of one and so then it can
- 17:53predict everything up to block size and
- 17:55after block size we have to start
- 17:56truncating because the Transformer will
- 17:58will never um receive more than block
- 18:01size inputs when it's predicting the
- 18:03next
- 18:03character Okay so we've looked at the
- 18:06time dimension of the tensors that are
- 18:07going to be feeding into the Transformer
- 18:09there's one more Dimension to care about
- 18:11and that is the batch Dimension and so
- 18:13as we're sampling these chunks of text
- 18:15we're going to be actually every time
- 18:17we're going to feed them into a
- 18:18Transformer we're going to have many
- 18:20batches of multiple chunks of text that
- 18:22are all like stacked up in a single
- 18:23tensor and that's just done for
- 18:25efficiency just so that we can keep the
- 18:27gpus busy uh because they are very good
- 18:29at parallel processing of um of data and
- 18:33so we just want to process multiple
- 18:35chunks all at the same time but those
- 18:37chunks are processed completely
- 18:38independently they don't talk to each
- 18:39other and so on so let me basically just
- 18:42generalize this and introduce a batch
- 18:44Dimension here's a chunk of
- 18:46code let me just run it and then I'm
- 18:48going to explain what it
- 18:50does so here because we're going to
- 18:52start sampling random locations in the
- 18:54data set to pull chunks from I am
- 18:57setting the seed so that um in the
- 19:00random number generator so that the
- 19:01numbers I see here are going to be the
- 19:02same numbers you see later if you try to
- 19:04reproduce this now the batch size here
- 19:07is how many independent sequences we are
- 19:09processing every forward backward pass
- 19:11of the
- 19:12Transformer the block size as I
- 19:14explained is the maximum context length
- 19:16to make those predictions so let's say B
- 19:19size four block size eight and then
- 19:21here's how we get batch for any
- 19:23arbitrary split if the split is a
- 19:25training split then we're going to look
- 19:26at train data otherwise at valid data
- 19:30that gives us the data array and then
- 19:33when I Generate random positions to grab
- 19:35a chunk out of I actually grab I
- 19:38actually generate batch size number of
- 19:41Random offsets so because this is four
- 19:44we are ex is going to be a uh four
- 19:47numbers that are randomly generated
- 19:49between zero and Len of data minus block
- 19:51size so it's just random offsets into
- 19:53the training
- 19:54set and then X's as I explained are the
- 19:58first first block size characters
- 20:00starting at I the Y's are the offset by
- 20:05one of that so just add plus one and
- 20:08then we're going to get those chunks for
- 20:10every one of integers I INX and use a
- 20:14torch. stack to take all those uh uh
- 20:17one-dimensional tensors as we saw here
- 20:20and we're going to um stack them up at
- 20:24rows and so they all become a row in a
- 20:274x8 tensor
- 20:29so here's where I'm printing then when I
- 20:32sample a batch XB and YB the inputs to
- 20:35the Transformer now are the input X is
- 20:39the 4x8 tensor four uh rows of eight
- 20:44columns and each one of these is a chunk
- 20:47of the training
- 20:48set and then the targets here are in the
- 20:52associated array Y and they will come in
- 20:54to the Transformer all the way at the
- 20:55end uh to um create the loss function
- 20:59uh so they will give us the correct
- 21:01answer for every single position inside
- 21:03X and then these are the four
- 21:06independent
- 21:07rows so spelled out as we did
- 21:11before uh this 4x8 array contains a
- 21:14total of 32 examples and they're
- 21:17completely independent as far as the
- 21:19Transformer is
- 21:20concerned uh so when the input is 24 the
- 21:25target is 43 or rather 43 here in the Y
- 21:28array
- 21:29when the input is 2443 the target is
- 21:3158 uh when the input is 24 43 58 the
- 21:34target is 5 Etc or like when it is a 52
- 21:38581 the target is
- 21:4058 right so you can sort of see this
- 21:43spelled out these are the 32 independent
- 21:45examples packed in to a single batch of
- 21:48the input X and then the desired targets
- 21:51are in y and so now this integer tensor
- 21:57of um X is going to feed into the
- 22:00Transformer and that Transformer is
- 22:02going to simultaneously process all
- 22:04these examples and then look up the
- 22:06correct um integers to predict in every
- 22:08one of these positions in the tensor y
- 22:11okay so now that we have our batch of
- 22:13input that we'd like to feed into a
- 22:15Transformer let's start basically
- 22:16feeding this into neural networks now
- 22:19we're going to start off with the
- 22:20simplest possible neural network which
- 22:22in the case of language modeling in my
- 22:23opinion is the Byram language model and
- 22:25we've covered the Byram language model
- 22:26in my make more series in a lot of depth
- 22:29and so here I'm going to sort of go
- 22:31faster and let's just Implement pytorch
- 22:33module directly that implements the byr
- 22:36language
- 22:36model so I'm importing the pytorch um NN
- 22:41module uh for
- 22:43reproducibility and then here I'm
- 22:44constructing a Byram language model
- 22:46which is a subass of NN
- 22:48module and then I'm calling it and I'm
- 22:51passing it the inputs and the targets
- 22:53and I'm just printing now when the
- 22:55inputs on targets come here you see that
- 22:57I'm just taking the index uh the inputs
- 23:00X here which I rename to idx and I'm
- 23:03just passing them into this token
- 23:04embedding table so it's going on here is
- 23:07that here in the Constructor we are
- 23:09creating a token embedding table and it
- 23:12is of size vocap size by vocap
- 23:15size and we're using an. embedding which
- 23:18is a very thin wrapper around basically
- 23:20a tensor of shape voap size by vocab
- 23:23size and what's happening here is that
- 23:25when we pass idx here every single
- 23:28integer in our input is going to refer
- 23:30to this embedding table and it's going
- 23:32to pluck out a row of that embedding
- 23:34table corresponding to its index so 24
- 23:37here will go into the embedding table
- 23:39and we'll pluck out the 24th row and
- 23:42then 43 will go here and pluck out the
- 23:4443d row Etc and then pytorch is going to
- 23:47arrange all of this into a batch by Time
- 23:50by channel uh tensor in this case batch
- 23:53is four time is eight and C which is the
- 23:57channels is vocab size or 65 and so
- 24:01we're just going to pluck out all those
- 24:02rows arrange them in a b by T by C and
- 24:05now we're going to interpret this as the
- 24:07logits which are basically the scores
- 24:10for the next character in the sequence
- 24:12and so what's happening here is we are
- 24:14predicting what comes next based on just
- 24:17the individual identity of a single
- 24:19token and you can do that because um I
- 24:22mean currently the tokens are not
- 24:23talking to each other and they're not
- 24:25seeing any context except for they're
- 24:26just seeing themselves so I'm a f I'm a
- 24:29token number five and then I can
- 24:32actually make pretty decent predictions
- 24:33about what comes next just by knowing
- 24:35that I'm token five because some
- 24:37characters uh know um C follow other
- 24:39characters in in typical scenarios so we
- 24:42saw a lot of this in a lot more depth in
- 24:44the make more series and here if I just
- 24:46run this then we currently get the
- 24:49predictions the scores the lits for
- 24:53every one of the 4x8 positions now that
- 24:55we've made predictions about what comes
- 24:57next we'd like to evaluate the loss
- 24:58function and so in make more series we
- 25:00saw that a good way to measure a loss or
- 25:03like a quality of the predictions is to
- 25:05use the negative log likelihood loss
- 25:07which is also implemented in pytorch
- 25:09under the name cross entropy so what we'
- 25:12like to do here is loss is the cross
- 25:15entropy on the predictions and the
- 25:17targets and so this measures the quality
- 25:20of the logits with respect to the
- 25:21Targets in other words we have the
- 25:24identity of the next character so how
- 25:26well are we predicting the next
- 25:28character based on the lits and
- 25:30intuitively the correct um the correct
- 25:33dimension of low jits uh depending on
- 25:36whatever the target is should have a
- 25:38very high number and all the other
- 25:39dimensions should be very low number
- 25:41right now the issue is that this won't
- 25:44actually this is what we want we want to
- 25:46basically output the logits and the
- 25:50loss this is what we want but
- 25:52unfortunately uh this won't actually run
- 25:55we get an error message but intuitively
- 25:57we want to uh measure this now when we
- 26:01go to the pytorch um cross entropy
- 26:04documentation here um we're trying to
- 26:08call the cross entropy in its functional
- 26:10form uh so that means we don't have to
- 26:11create like a module for it but here
- 26:14when we go to the documentation you have
- 26:16to look into the details of how pitor
- 26:18expects these inputs and basically the
- 26:20issue here is ptor expects if you have
- 26:24multi-dimensional input which we do
- 26:25because we have a b BYT by C tensor then
- 26:28it actually really wants the channels to
- 26:31be the second uh Dimension here so if
- 26:35you um so basically it wants a b by C
- 26:38BYT instead of a b by T by C and so it's
- 26:42just the details of how P torch treats
- 26:45um these kinds of inputs and so we don't
- 26:49actually want to deal with that so what
- 26:51we're going to do instead is we need to
- 26:52basically reshape our logits so here's
- 26:54what I like to do I like to take
- 26:56basically give names to the dimensions
- 26:58so lit. shape is B BYT by C and unpack
- 27:01those numbers and then let's uh say that
- 27:04logits equals lit. View and we want it
- 27:07to be a b * c b * T by C so just a two-
- 27:11dimensional
- 27:12array right so we're going to take all
- 27:15the we're going to take all of these um
- 27:18positions here and we're going to uh
- 27:20stretch them out in a onedimensional
- 27:22sequence and uh preserve the channel
- 27:25Dimension as the second
- 27:26dimension so we're just kind of like
- 27:28stretching out the array so it's two-
- 27:29dimensional and in that case it's going
- 27:31to better conform to what pytorch uh
- 27:33sort of expects in its Dimensions now we
- 27:36have to do the same to targets because
- 27:38currently targets are um of shape B by T
- 27:44and we want it to be just B * T so
- 27:47onedimensional now alternatively you
- 27:49could always still just do minus one
- 27:51because pytor will guess what this
- 27:53should be if you want to lay it out uh
- 27:55but let me just be explicit and say p *
- 27:57t once we've reshaped this it will match
- 28:00the cross entropy case and then we
- 28:03should be able to evaluate our
- 28:06loss okay so that R now and we can do
- 28:10loss and So currently we see that the
- 28:12loss is
- 28:134.87 now because our uh we have 65
- 28:17possible vocabulary elements we can
- 28:19actually guess at what the loss should
- 28:20be and in
- 28:22particular we covered negative log
- 28:24likelihood in a lot of detail we are
- 28:26expecting log or lawn of um 1 over 65
- 28:32and negative of that so we're expecting
- 28:34the loss to be about 4.1 17 but we're
- 28:37getting 4.87 and so that's telling us
- 28:40that the initial predictions are not uh
- 28:42super diffuse they've got a little bit
- 28:43of entropy and so we're guessing wrong
- 28:47uh so uh yes but actually we're I a we
- 28:50are able to evaluate the loss okay so
- 28:53now that we can evaluate the quality of
- 28:54the model on some data we'd like to also
- 28:57be able to generate from the model so
- 28:59let's do the generation now I'm going to
- 29:01go again a little bit faster here
- 29:03because I covered all this already in
- 29:04previous
- 29:05videos
- 29:07so here's a generate function for the
- 29:11model so we take some uh we take the the
- 29:15same kind of input idx here and
- 29:18basically this is the current uh context
- 29:22of some characters in a batch in some
- 29:24batch so it's also B BYT and the job of
- 29:28generate is to basically take this B BYT
- 29:30and extend it to be B BYT + 1 plus 2
- 29:32plus 3 and so it's just basically it
- 29:34continues the generation in all the
- 29:36batch dimensions in the time Dimension
- 29:39So that's its job and it will do that
- 29:41for Max new tokens so you can see here
- 29:43on the bottom there's going to be some
- 29:45stuff here but on the bottom whatever is
- 29:47predicted is concatenated on top of the
- 29:50previous idx along the First Dimension
- 29:53which is the time Dimension to create a
- 29:54b BYT + one so that becomes a new idx so
- 29:58the job of generate is to take a b BYT
- 30:00and make it a b BYT plus 1 plus 2 plus
- 30:02three as many as we want Max new tokens
- 30:05so this is the generation from the model
- 30:08now inside the generation what what are
- 30:10we doing we're taking the current
- 30:11indices we're getting the predictions so
- 30:15we get uh those are in the low jits and
- 30:18then the loss here is going to be
- 30:19ignored because um we're not we're not
- 30:21using that and we have no targets that
- 30:23are sort of ground truth targets that
- 30:25we're going to be comparing with
- 30:28then once we get the logits we are only
- 30:30focusing on the last step so instead of
- 30:33a b by T by C we're going to pluck out
- 30:36the negative-1 the last element in the
- 30:38time Dimension because those are the
- 30:40predictions for what comes next so that
- 30:42gives us the logits which we then
- 30:44convert to probabilities via softmax and
- 30:47then we use tor. multinomial to sample
- 30:49from those probabilities and we ask
- 30:51pytorch to give us one sample and so idx
- 30:54next will become a b by one because in
- 30:57each uh one of the batch Dimensions
- 31:00we're going to have a single prediction
- 31:01for what comes next so this num samples
- 31:03equals one will make this be a
- 31:06one and then we're going to take those
- 31:08integers that come from the sampling
- 31:10process according to the probability
- 31:11distribution given here and those
- 31:13integers got just concatenated on top of
- 31:15the current sort of like running stream
- 31:17of integers and this gives us a b BYT +
- 31:20one and then we can return that now one
- 31:24thing here is you see how I'm calling
- 31:26self of idx which will end up going to
- 31:29the forward function I'm not providing
- 31:31any Targets So currently this would give
- 31:33an error because targets is uh is uh
- 31:36sort of like not given so targets has to
- 31:39be optional so targets is none by
- 31:41default and then if targets is none then
- 31:44there's no loss to create so it's just
- 31:47loss is none but else all of this
- 31:50happens and we can create a loss so this
- 31:53will make it so um if we have the
- 31:56targets we provide them and get a loss
- 31:57if we have no targets it will'll just
- 31:59get the
- 32:00loits so this here will generate from
- 32:02the model um and let's take that for a
- 32:06ride
- 32:08now oops so I have another code chunk
- 32:11here which will generate for the model
- 32:13from the model and okay this is kind of
- 32:15crazy so maybe let me let me break this
- 32:18down so these are the idx
- 32:23right I'm creating a batch will be just
- 32:26one time will be just one so I'm
- 32:30creating a little one by one tensor and
- 32:32it's holding a zero and the D type the
- 32:35data type is uh integer so zero is going
- 32:38to be how we kick off the generation and
- 32:40remember that zero is uh is the element
- 32:44standing for a new line character so
- 32:45it's kind of like a reasonable thing to
- 32:47to feed in as the very first character
- 32:49in a sequence to be the new
- 32:51line um so it's going to be idx which
- 32:54we're going to feed in here then we're
- 32:56going to ask for 100 tokens
- 32:58and then. generate will continue that
- 33:01now because uh generate works on the
- 33:05level of batches we we then have to
- 33:07index into the zero throw to basically
- 33:09unplug the um the single batch Dimension
- 33:13that exists and then that gives us a um
- 33:18time steps just a onedimensional array
- 33:20of all the indices which we will convert
- 33:23to simple python list from pytorch
- 33:26tensor so that that can feed into our
- 33:28decode function and uh convert those
- 33:32integers into text so let me bring this
- 33:34back and we're generating 100 tokens
- 33:37let's
- 33:37run and uh here's the generation that we
- 33:40achieved so obviously it's garbage and
- 33:43the reason it's garbage is because this
- 33:44is a totally random model so next up
- 33:47we're going to want to train this model
- 33:49now one more thing I wanted to point out
- 33:50here is this function is written to be
- 33:53General but it's kind of like ridiculous
- 33:55right now because
- 33:58we're feeding in all this we're building
- 33:59out this context and we're concatenating
- 34:02it all and we're always feeding it all
- 34:05into the model but that's kind of
- 34:07ridiculous because this is just a simple
- 34:09Byram model so to make for example this
- 34:11prediction about K we only needed this W
- 34:14but actually what we fed into the model
- 34:15is we fed the entire sequence and then
- 34:18we only looked at the very last piece
- 34:20and predicted K so the only reason I'm
- 34:23writing it in this way is because right
- 34:25now this is a byr model but I'd like to
- 34:27keep keep this function fixed and I'd
- 34:29like it to work um later when our
- 34:32characters actually um basically look
- 34:35further in the history and so right now
- 34:37the history is not used so this looks
- 34:39silly uh but eventually the history will
- 34:42be used and so that's why we want to uh
- 34:44do it this way so just a quick comment
- 34:46on that so now we see that this is um
- 34:49random so let's train the model so it
- 34:51becomes a bit less random okay let's Now
- 34:53train the model so first what I'm going
- 34:55to do is I'm going to create a pyour
- 34:57optimization object so here we are using
- 35:00the optimizer ATM W um now in a make
- 35:05more series we've only ever use tastic
- 35:06gradi in descent the simplest possible
- 35:08Optimizer which you can get using the
- 35:10SGD instead but I want to use Adam which
- 35:12is a much more advanced and popular
- 35:14Optimizer and it works extremely well
- 35:16for uh typical good setting for the
- 35:19learning rate is roughly 3 E4 uh but for
- 35:22very very small networks like is the
- 35:23case here you can get away with much
- 35:25much higher learning rates R3 or even
- 35:28higher probably but let me create the
- 35:30optimizer object which will basically
- 35:33take the gradients and uh update the
- 35:35parameters using the
- 35:36gradients and then here our batch size
- 35:40up above was only four so let me
- 35:41actually use something bigger let's say
- 35:4332 and then for some number of steps um
- 35:46we are sampling a new batch of data
- 35:48we're evaluating the loss uh we're
- 35:51zeroing out all the gradients from the
- 35:52previous step getting the gradients for
- 35:54all the parameters and then using those
- 35:56gradients to up update our parameters so
- 35:58typical training loop as we saw in the
- 36:00make more series so let me now uh run
- 36:04this for say 100 iterations and let's
- 36:07see what kind of losses we're going to
- 36:09get so we started around
- 36:124.7 and now we're getting to down to
- 36:14like 4.6 4.5 Etc so the optimization is
- 36:18definitely happening but um let's uh
- 36:22sort of try to increase number of
- 36:23iterations and only print at the
- 36:25end because we probably want train for
- 36:29longer okay so we're down to 3.6
- 36:34roughly roughly down to
- 36:40three this is the most janky
- 36:46optimization okay it's working let's
- 36:48just do
- 36:5010,000 and then from here we want to
- 36:53copy this and hopefully that we're going
- 36:56to get something reason and of course
- 36:58it's not going to be Shakespeare from a
- 37:00byr model but at least we see that the
- 37:01loss is improving and uh hopefully we're
- 37:05expecting something a bit more
- 37:06reasonable okay so we're down at about
- 37:082.5 is let's see what we get okay
- 37:12dramatic improvements certainly on what
- 37:14we had here so let me just increase the
- 37:17number of tokens okay so we see that
- 37:19we're starting to get something at least
- 37:21like reasonable is
- 37:25um certainly not shakes spear but uh the
- 37:29model is making progress so that is the
- 37:31simplest possible
- 37:33model so now what I'd like to do
- 37:36is obviously this is a very simple model
- 37:39because the tokens are not talking to
- 37:41each other so given the previous context
- 37:43of whatever was generated we're only
- 37:45looking at the very last character to
- 37:46make the predictions about what comes
- 37:48next so now these uh now these tokens
- 37:50have to start talking to each other and
- 37:53figuring out what is in the context so
- 37:55that they can make better predictions
- 37:56for what comes next and this is how
- 37:57we're going to kick off the uh
- 37:59Transformer okay so next I took the code
- 38:02that we developed in this juper notebook
- 38:03and I converted it to be a script and
- 38:05I'm doing this because I just want to
- 38:08simplify our intermediate work into just
- 38:10the final product that we have at this
- 38:12point so in the top here I put all the
- 38:15hyp parameters that we to find I
- 38:16introduced a few and I'm going to speak
- 38:18to that in a little bit otherwise a lot
- 38:20of this should be recognizable uh
- 38:23reproducibility read data get the
- 38:25encoder and the decoder create the train
- 38:27into splits uh use the uh kind of like
- 38:30data loader um that gets a batch of the
- 38:34inputs and Targets this is new and I'll
- 38:36talk about it in a second now this is
- 38:39the Byram language model that we
- 38:40developed and it can forward and give us
- 38:43a logits and loss and it can
- 38:45generate and then here we are creating
- 38:48the optimizer and this is the training
- 38:51Loop so everything here should look
- 38:53pretty familiar now some of the small
- 38:55things that I added number one I added
- 38:57the ability to run on a GPU if you have
- 39:00it so if you have a GPU then you can
- 39:02this will use Cuda instead of just CPU
- 39:04and everything will be a lot more faster
- 39:07now when device becomes Cuda then we
- 39:09need to make sure that when we load the
- 39:11data we move it to
- 39:13device when we create the model we want
- 39:15to move uh the model parameters to
- 39:18device so as an example here we have the
- 39:21N an embedding table and it's got a
- 39:23weight inside it which stores the uh
- 39:26sort of lookup table so so that would be
- 39:27moved to the GPU so that all the
- 39:29calculations here happen on the GPU and
- 39:32they can be a lot faster and then
- 39:34finally here when I'm creating the
- 39:35context that feeds in to generate I have
- 39:37to make sure that I create it on the
- 39:39device number two what I introduced is
- 39:43uh the fact that here in the training
- 39:46Loop here I was just printing the um l.
- 39:50item inside the training Loop but this
- 39:53is a very noisy measurement of the
- 39:54current loss because every batch will be
- 39:56more or less lucky and so what I want to
- 39:59do usually um is uh I have an estimate
- 40:02loss function and the estimate loss
- 40:05basically then um goes up here and it
- 40:10averages up the loss over multiple
- 40:12batches so in particular we're going to
- 40:15iterate eval iter times and we're going
- 40:17to basically get our loss and then we're
- 40:19going to get the average loss for both
- 40:21splits and so this will be a lot less
- 40:24noisy so here when we call the estimate
- 40:26loss we're we're going to report the uh
- 40:28pretty accurate train and validation
- 40:31loss now when we come back up you'll
- 40:33notice a few things here I'm setting the
- 40:35model to evaluation phase and down here
- 40:38I'm resetting it back to training phase
- 40:40now right now for our model as is this
- 40:42doesn't actually do anything because the
- 40:44only thing inside this model is this uh
- 40:46nn. embedding and um this this um
- 40:51Network would behave both would behave
- 40:53the same in both evaluation mode and
- 40:55training mode we have no drop off layers
- 40:57we have no batm layers Etc but it is a
- 41:00good practice to Think Through what mode
- 41:02your neural network is in because some
- 41:04layers will have different Behavior Uh
- 41:07at inference time or training time and
- 41:11there's also this context manager torch
- 41:12up nograd and this is just telling
- 41:14pytorch that everything that happens
- 41:16inside this function we will not call do
- 41:18backward on and so pytorch can be a lot
- 41:21more efficient with its memory use
- 41:23because it doesn't have to store all the
- 41:25intermediate variables uh because we're
- 41:27never going to call backward and so it
- 41:29can it can be a lot more memory
- 41:30efficient in that way so also a good
- 41:32practice to tpy torch when we don't
- 41:35intend to do back
- 41:36propagation so right now this script is
- 41:39about 120 lines of code of and that's
- 41:43kind of our starter code I'm calling it
- 41:45b.p and I'm going to release it later
- 41:48now running this
- 41:50script gives us output in the terminal
- 41:52and it looks something like this it
- 41:54basically as I ran this code uh it was
- 41:57giving me the train loss and Val loss
- 41:59and we see that we convert to somewhere
- 42:01around
- 42:012.5 with the pyr model and then here's
- 42:04the sample that we produced at the
- 42:07end and so we have everything packaged
- 42:09up in the script and we're in a good
- 42:11position now to iterate on this okay so
- 42:13we are almost ready to start writing our
- 42:15very first self attention block for
- 42:18processing these uh tokens now before we
- 42:22actually get there I want to get you
- 42:24used to a mathematical trick that is
- 42:26used in the self attention inside a
- 42:28Transformer and is really just like at
- 42:30the heart of an an efficient
- 42:32implementation of self attention and so
- 42:34I want to work with this toy example to
- 42:36just get you used to this operation and
- 42:38then it's going to make it much more
- 42:39clear once we actually get to um to it
- 42:43uh in the script
- 42:44again so let's create a b BYT by C where
- 42:47BT and C are just 48 and two in the toy
- 42:50example and these are basically channels
- 42:53and we have uh batches and we have the
- 42:55time component and we have information
- 42:58at each point in the sequence so
- 43:01see now what we would like to do is we
- 43:03would like these um tokens so we have up
- 43:06to eight tokens here in a batch and
- 43:08these eight tokens are currently not
- 43:10talking to each other and we would like
- 43:11them to talk to each other we'd like to
- 43:13couple them and in particular we don't
- 43:17we we want to couple them in a very
- 43:18specific way so the token for example at
- 43:21the fifth location it should not
- 43:23communicate with tokens in the sixth
- 43:25seventh and eighth location
- 43:27because uh those are future tokens in
- 43:29the sequence the token on the fifth
- 43:31location should only talk to the one in
- 43:33the fourth third second and first so
- 43:36it's only so information only flows from
- 43:38previous context to the current time
- 43:40step and we cannot get any information
- 43:42from the future because we are about to
- 43:44try to predict the
- 43:45future so what is the easiest way for
- 43:49tokens to communicate okay the easiest
- 43:52way I would say is okay if we're up to
- 43:54if we're a fifth token and I'd like to
- 43:56communicate with my past the simplest
- 43:58way we can do that is to just do a
- 44:00weight is to just do an average of all
- 44:03the um of all the preceding elements so
- 44:06for example if I'm the fif token I would
- 44:08like to take the channels uh that make
- 44:10up that are information at my step but
- 44:13then also the channels from the fourth
- 44:15step third step second step and the
- 44:17first step I'd like to average those up
- 44:19and then that would become sort of like
- 44:21a feature Vector that summarizes me in
- 44:23the context of my history now of course
- 44:26just doing a sum or like an average is
- 44:28an extremely weak form of interaction
- 44:30like this communication is uh extremely
- 44:32lossy we've lost a ton of information
- 44:34about the spatial Arrangements of all
- 44:35those tokens uh but that's okay for now
- 44:38we'll see how we can bring that
- 44:39information back later for now what we
- 44:41would like to do is for every single
- 44:43batch element independently for every
- 44:46teeth token in that sequence we'd like
- 44:49to now calculate the average of all the
- 44:53vectors in all the previous tokens and
- 44:55also at this token so let's write that
- 44:58out um I have a small snippet here and
- 45:01instead of just fumbling around let me
- 45:03just copy paste it and talk to
- 45:05it so in other words we're going to
- 45:08create X and B is short for bag of words
- 45:12because bag of words is um is kind of
- 45:15like um a term that people use when you
- 45:17are just averaging up things so this is
- 45:19just a bag of words basically there's a
- 45:21word stored on every one of these eight
- 45:23locations and we're doing a bag of words
- 45:25we're just averaging
- 45:27so in the beginning we're going to say
- 45:28that it's just initialized at Zero and
- 45:30then I'm doing a for Loop here so we're
- 45:32not being efficient yet that's coming
- 45:34but for now we're just iterating over
- 45:36all the batch Dimensions independently
- 45:38iterating over time and then the
- 45:40previous uh tokens are at this uh batch
- 45:45Dimension and then everything up to and
- 45:47including the teeth token okay so when
- 45:51we slice out X in this way X prev
- 45:54Becomes of shape um how many T elements
- 45:58there were in the past and then of
- 46:00course C so all the two-dimensional
- 46:02information from these little tokens so
- 46:05that's the previous uh sort of chunk of
- 46:08um tokens from my current sequence and
- 46:12then I'm just doing the average or the
- 46:13mean over the zero Dimension so I'm
- 46:15averaging out the time here and I'm just
- 46:19going to get a little c one dimensional
- 46:21Vector which I'm going to store in X bag
- 46:23of words so I can run this and and uh
- 46:27this is not going to be very informative
- 46:30because let's see so this is X of Zer so
- 46:32this is the zeroth batch element and
- 46:35then expo at zero now you see how the at
- 46:40the first location here you see that the
- 46:42two are equal and that's because it's
- 46:45we're just doing an average of this one
- 46:46token but here this one is now an
- 46:49average of these two and now this one is
- 46:53an average of these
- 46:54three and so on
- 46:57so uh and this last one is the average
- 47:01of all of these elements so vertical
- 47:03average just averaging up all the tokens
- 47:05now gives this outcome
- 47:07here so this is all well and good uh but
- 47:10this is very inefficient now the trick
- 47:12is that we can be very very efficient
- 47:14about doing this using matrix
- 47:16multiplication so that's the
- 47:18mathematical trick and let me show you
- 47:19what I mean let's work with the toy
- 47:21example here let me run it and I'll
- 47:24explain I have a simple Matrix here that
- 47:27is a 3X3 of all ones a matrix B of just
- 47:31random numbers and it's a 3x2 and a
- 47:33matrix C which will be 3x3 multip 3x2
- 47:36which will give out a 3x2 so here we're
- 47:39just using um matrix multiplication so a
- 47:43multiply B gives us
- 47:46C okay so how are these numbers in C um
- 47:51achieved right so this number in the top
- 47:54left is the first row of a dot product
- 47:57with the First Column of B and since all
- 48:00the the row of a right now is all just
- 48:02ones then the do product here with with
- 48:05this column of B is just going to do a
- 48:07sum of these of this column so 2 + 6 + 6
- 48:11is
- 48:1214 the element here in the output of C
- 48:15is also the first column here the first
- 48:17row of a multiplied now with the second
- 48:20column of B so 7 + 4 + 5 is 16 now you
- 48:25see that there's repeating elements here
- 48:26so this 14 again is because this row is
- 48:28again all ones and it's multiplying the
- 48:30First Column of B so we get 14 and this
- 48:33one is and so on so this last number
- 48:35here is the last row do product last
- 48:39column now the trick here is uh the
- 48:42following this is just a boring number
- 48:44of um it's just a boring array of all
- 48:48ones but torch has this function called
- 48:50Trail which is short for a
- 48:54triangular uh something like that and
- 48:56you can wrap it in torch up once and it
- 48:58will just return the lower triangular
- 49:00portion of this
- 49:03okay so now it will basically zero out
- 49:06uh these guys here so we just get the
- 49:08lower triangular part well what happens
- 49:10if we do
- 49:14that so now we'll have a like this and B
- 49:17like this and now what are we getting
- 49:18here in C well what is this number well
- 49:22this is the first row times the First
- 49:24Column and because this is zeros
- 49:28uh these elements here are now ignored
- 49:30so we just get a two and then this
- 49:32number here is the first row times the
- 49:35second column and because these are
- 49:37zeros they get ignored and it's just
- 49:39seven this seven multiplies this one but
- 49:42look what happened here because this is
- 49:43one and then zeros we what ended up
- 49:46happening is we're just plucking out the
- 49:48row of this row of B and that's what we
- 49:51got now here we have one 1 Z so here 110
- 49:57do product with these two columns will
- 49:59now give us 2 + 6 which is 8 and 7 + 4
- 50:02which is 11 and because this is 111 we
- 50:05ended up with the addition of all of
- 50:07them and so basically depending on how
- 50:10many ones and zeros we have here we are
- 50:12basically doing a sum currently of a
- 50:16variable number of these rows and that
- 50:18gets deposited into
- 50:20C So currently we're doing sums because
- 50:23these are ones but we can also do
- 50:25average right and you can start to see
- 50:27how we could do average uh of the rows
- 50:29of B uh sort of in an incremental
- 50:32fashion because we don't have to we can
- 50:35basically normalize these rows so that
- 50:37they sum to one and then we're going to
- 50:39get an average so if we took a and then
- 50:41we did aals
- 50:43aide torch. sum in the um of a in the um
- 50:51oneth Dimension and then let's keep them
- 50:55as true so so therefore the broadcasting
- 50:57will work out so if I rerun this you see
- 51:00now that these rows now sum to one so
- 51:04this row is one this row is 0. 5.5 Z and
- 51:07here we get 1/3 and now when we do a
- 51:09multiply B what are we getting here we
- 51:12are just getting the first row first row
- 51:15here now we are getting the average of
- 51:18the first two
- 51:20rows okay so 2 and six average is four
- 51:23and four and seven average is
- 51:255.5 and on the bottom here we are now
- 51:27getting the average of these three rows
- 51:31so the average of all of elements of B
- 51:33are now deposited here and so you can
- 51:36see that by manipulating these uh
- 51:40elements of this multiplying Matrix and
- 51:42then multiplying it with any given
- 51:44Matrix we can do these averages in this
- 51:47incremental fashion because we just get
- 51:50um and we can manipulate that based on
- 51:53the elements of a okay so that's very
- 51:55convenient so let's let's swing back up
- 51:57here and see how we can vectorize this
- 51:59and make it much more efficient using
- 52:00what we've learned so in
- 52:03particular we are going to produce an
- 52:05array a but here I'm going to call it we
- 52:08short for weights but this is our
- 52:11a and this is how much of every row we
- 52:14want to average up and it's going to be
- 52:17an average because you can see that
- 52:18these rows sum to
- 52:20one so this is our a and then our B in
- 52:23this example of course is X
- 52:27so what's going to happen here now is
- 52:29that we are going to have an expo
- 52:312 and this Expo 2 is going to be way
- 52:36multiplying
- 52:38RX so let's think this true way is T BYT
- 52:42and this is Matrix multiplying in
- 52:44pytorch a b by T by
- 52:47C and it's giving us uh different what
- 52:50shape so pytorch will come here and it
- 52:52will see that these shapes are not the
- 52:54same so it will create a batch Dimension
- 52:57here and this is a batched matrix
- 53:00multiply and so it will apply this
- 53:02matrix multiplication in all the batch
- 53:04elements um in parallel and individually
- 53:08and then for each batch element there
- 53:09will be a t BYT multiplying T by C
- 53:12exactly as we had
- 53:15below so this will now create B by T by
- 53:20C and Expo 2 will now become identical
- 53:24to Expo
- 53:28so we can see that torch. all close of
- 53:32xbo and xbo 2 should be true
- 53:36now so this kind of like convinces us
- 53:38that uh these are in fact um the same so
- 53:43xbo and xbo 2 if I just print
- 53:47them uh okay we're not going to be able
- 53:49to okay we're not going to be able to
- 53:51just stare it down but
- 53:54um well let me try Expo basically just
- 53:56at the zeroth element and Expo two at
- 53:58the zeroth element so just the first
- 53:59batch and we should see that this and
- 54:02that should be identical which they
- 54:04are right so what happened here the
- 54:07trick is we were able to use batched
- 54:09Matrix multiply to do this uh
- 54:12aggregation really and it's a weighted
- 54:15aggregation and the weights are
- 54:17specified in this um T BYT array and
- 54:21we're basically doing weighted sums and
- 54:24uh these weighted sums are are U
- 54:26according to uh the weights inside here
- 54:28they take on sort of this triangular
- 54:31form and so that means that a token at
- 54:33the teth dimension will only get uh sort
- 54:36of um information from the um tokens
- 54:39perceiving it so that's exactly what we
- 54:41want and finally I would like to rewrite
- 54:43it in one more way and we're going to
- 54:46see why that's useful so this is the
- 54:48third version and it's also identical to
- 54:50the first and second but let me talk
- 54:53through it it uses
- 54:54softmax so Trill here is this Matrix
- 55:00lower triangular
- 55:01ones way begins as all
- 55:05zero okay so if I just print way in the
- 55:07beginning it's all zero then I
- 55:11used masked fill so what this is doing
- 55:15is we. masked fill it's all zeros and
- 55:18I'm saying for all the elements where
- 55:20Trill is equal equal Z make them be
- 55:23negative Infinity so all the elements
- 55:26where Trill is zero will become negative
- 55:28Infinity now so this is what we get and
- 55:32then the final line here is
- 55:36softmax so if I take a softmax along
- 55:38every single so dim is negative one so
- 55:40along every single row if I do softmax
- 55:44what is that going to
- 55:46do well softmax is um is also like a
- 55:51normalization operation right and so
- 55:54spoiler alert you get the exact same
- 55:58Matrix let me bring back to
- 56:00softmax and recall that in softmax we're
- 56:02going to exponentiate every single one
- 56:04of these and then we're going to divide
- 56:06by the sum and so if we exponentiate
- 56:10every single element here we're going to
- 56:11get a one and here we're going to get uh
- 56:14basically zero 0 z0 Z everywhere else
- 56:17and then when we normalize we just get
- 56:19one here we're going to get one one and
- 56:21then zeros and then softmax will again
- 56:24divide and this will give us 5.5 and so
- 56:27on and so this is also the uh the same
- 56:30way to produce uh this mask now the
- 56:33reason that this is a bit more
- 56:34interesting and the reason we're going
- 56:36to end up using it in self
- 56:37attention is that these weights here
- 56:41begin uh with zero and you can think of
- 56:44this as like an interaction strength or
- 56:46like an affinity so basically it's
- 56:49telling us how much of each uh token
- 56:52from the past do we want to Aggregate
- 56:54and average up
- 56:57and then this line is saying tokens from
- 56:59the past cannot communicate by setting
- 57:02them to negative Infinity we're saying
- 57:04that we will not aggregate anything from
- 57:06those
- 57:07tokens and so basically this then goes
- 57:09through softmax and through the weighted
- 57:11and this is the aggregation through
- 57:12matrix
- 57:14multiplication and so what this is now
- 57:16is you can think of these as um these
- 57:19zeros are currently just set by us to be
- 57:21zero but a quick preview is that these
- 57:25affinities between the tokens are not
- 57:27going to be just constant at zero
- 57:29they're going to be data dependent these
- 57:31tokens are going to start looking at
- 57:32each other and some tokens will find
- 57:34other tokens more or less interesting
- 57:37and depending on what their values are
- 57:39they're going to find each other
- 57:41interesting to different amounts and I'm
- 57:42going to call those affinities I think
- 57:45and then here we are saying the future
- 57:47cannot communicate with the past we're
- 57:49we're going to clamp them and then when
- 57:51we normalize and sum we're going to
- 57:53aggregate uh sort of their values
- 57:56depending on how interesting they find
- 57:57each other and so that's the preview for
- 57:59self attention and basically long story
- 58:03short from this entire section is that
- 58:05you can do weighted aggregations of your
- 58:07past
- 58:08Elements by having by using matrix
- 58:12multiplication of a lower triangular
- 58:14fashion and then the elements here in
- 58:17the lower triangular part are telling
- 58:18you how much of each element uh fuses
- 58:21into this position so we're going to use
- 58:24this trick now to develop the self
- 58:25attention block block so first let's get
- 58:27some quick preliminaries out of the way
- 58:30first the thing I'm kind of bothered by
- 58:31is that you see how we're passing in
- 58:33vocap size into the Constructor there's
- 58:35no need to do that because vocap size is
- 58:36already defined uh up top as a global
- 58:38variable so there's no need to pass this
- 58:40stuff
- 58:41around next what I want to do is I don't
- 58:44want to actually create I want to create
- 58:46like a level of indirection here where
- 58:47we don't directly go to the embedding
- 58:49for the um logits but instead we go
- 58:52through this intermediate phase because
- 58:54we're going to start making that bigger
- 58:57so let me introduce a new variable n
- 58:59embed it shorted for number of embedding
- 59:02Dimensions so
- 59:04nbed here will be say 32 that was a
- 59:09suggestion from GitHub co-pilot by the
- 59:11way um it also suest 32 which is a good
- 59:14number so this is an embedding table and
- 59:16only 32 dimensional
- 59:18embeddings so then here this is not
- 59:21going to give us logits directly instead
- 59:23this is going to give us token
- 59:24embeddings that's I'm going to call it
- 59:27and then to go from the token Tings to
- 59:29the logits we're going to need a linear
- 59:30layer so self. LM head let's call it
- 59:34short for language modeling head is n
- 59:36and linear from n ined up to vocap size
- 59:39and then when we swing over here we're
- 59:41actually going to get the loits by
- 59:43exactly what the co-pilot says now we
- 59:46have to be careful here because this C
- 59:48and this C are not equal um this is nmed
- 59:52C and this is vocap size so let's just
- 59:55say that n ined is equal to
- 59:57C and then this just creates one spous
- 1:00:01layer of interaction through a linear
- 1:00:02layer but uh this should basically
- 1:00:11run so we see that this runs and uh this
- 1:00:15currently looks kind of spous but uh
- 1:00:17we're going to build on top of this now
- 1:00:19next up so far we've taken these indices
- 1:00:22and we've encoded them based on the
- 1:00:23identity of the uh tokens in inside idx
- 1:00:28the next thing that people very often do
- 1:00:30is that we're not just encoding the
- 1:00:31identity of these tokens but also their
- 1:00:33position so we're going to have a second
- 1:00:35position uh embedding table here so
- 1:00:38self. position embedding table is an an
- 1:00:41embedding of block size by an embed and
- 1:00:44so each position from zero to block size
- 1:00:46minus one will also get its own
- 1:00:47embedding vector and then here first let
- 1:00:50me decode B BYT from idx do
- 1:00:54shape and then here we're also going to
- 1:00:56have a pause embedding which is the
- 1:00:58positional embedding and these are this
- 1:01:00is to arrange so this will be basically
- 1:01:03just integers from Z to T minus one and
- 1:01:06all of those integers from 0 to T minus
- 1:01:08one get embedded through the table to
- 1:01:09create a t by
- 1:01:11C and then here this gets renamed to
- 1:01:14just say x and x will be the addition of
- 1:01:18the token embeddings with the positional
- 1:01:20embeddings and here the broadcasting
- 1:01:22note will work out so B by T by C plus T
- 1:01:25by C
- 1:01:26this gets right aligned a new dimension
- 1:01:28of one gets added and it gets
- 1:01:30broadcasted across
- 1:01:31batch so at this point x holds not just
- 1:01:34the token identities but the positions
- 1:01:37at which these tokens occur and this is
- 1:01:39currently not that useful because of
- 1:01:41course we just have a simple byr model
- 1:01:43so it doesn't matter if you're in the
- 1:01:44fifth position the second position or
- 1:01:46wherever it's all translation invariant
- 1:01:48at this stage uh so this information
- 1:01:50currently wouldn't help uh but as we
- 1:01:52work on the self attention block we'll
- 1:01:54see that this starts to matter
- 1:01:59okay so now we get the Crux of self
- 1:02:01attention so this is probably the most
- 1:02:03important part of this video to
- 1:02:05understand we're going to implement a
- 1:02:07small self attention for a single
- 1:02:08individual head as they're called so we
- 1:02:11start off with where we were so all of
- 1:02:13this code is familiar so right now I'm
- 1:02:16working with an example where I Chang
- 1:02:17the number of channels from 2 to 32 so
- 1:02:20we have a 4x8 arrangement of tokens and
- 1:02:24each to and the information each token
- 1:02:26is currently 32 dimensional but we just
- 1:02:28are working with random
- 1:02:30numbers now we saw here that the code as
- 1:02:34we had it before does a uh simple weight
- 1:02:37simple average of all the past tokens
- 1:02:41and the current token so it's just the
- 1:02:43previous information and current
- 1:02:44information is just being mixed together
- 1:02:45in an average and that's what this code
- 1:02:48currently achieves and it Doo by
- 1:02:50creating this lower triangular structure
- 1:02:52which allows us to mask out this uh we
- 1:02:55uh Matrix that we create so we mask it
- 1:02:59out and then we normalize it and
- 1:03:01currently when we initialize the
- 1:03:03affinities between all the different
- 1:03:05sort of tokens or nodes I'm going to use
- 1:03:08those terms
- 1:03:09interchangeably so when we initialize
- 1:03:11the affinities between all the different
- 1:03:13tokens to be zero then we see that way
- 1:03:16gives us this um structure where every
- 1:03:18single row has these um uniform numbers
- 1:03:22and so that's what that's what then uh
- 1:03:25in this Matrix multiply makes it so that
- 1:03:27we're doing a simple
- 1:03:28average now we don't actually want this
- 1:03:32to be all uniform because different uh
- 1:03:36tokens will find different other tokens
- 1:03:38more or less interesting and we want
- 1:03:40that to be data dependent so for example
- 1:03:42if I'm a vowel then maybe I'm looking
- 1:03:44for consonants in my past and maybe I
- 1:03:46want to know what those consonants are
- 1:03:48and I want that information to flow to
- 1:03:50me and so I want to now gather
- 1:03:52information from the past but I want to
- 1:03:54do it in the data dependent way and this
- 1:03:56is the problem that self attention
- 1:03:58solves now the way self attention solves
- 1:04:00this is the following every single node
- 1:04:03or every single token at each position
- 1:04:06will emit two vectors it will emit a
- 1:04:09query and it will emit a
- 1:04:12key now the query Vector roughly
- 1:04:15speaking is what am I looking for and
- 1:04:18the key Vector roughly speaking is what
- 1:04:20do I
- 1:04:21contain and then the way we get
- 1:04:24affinities between these uh tokens now
- 1:04:27in a sequence is we basically just do a
- 1:04:29do product between the keys and the
- 1:04:31queries so my query dot products with
- 1:04:35all the keys of all the other tokens and
- 1:04:37that dot product now becomes
- 1:04:41wayy and so um if the key and the query
- 1:04:45are sort of aligned they will interact
- 1:04:47to a very high amount and then I will
- 1:04:50get to learn more about that specific
- 1:04:52token as opposed to any other token in
- 1:04:55the sequence
- 1:04:56so let's implement this
- 1:05:00now we're going to implement a
- 1:05:03single what's called head of self
- 1:05:07attention so this is just one head
- 1:05:09there's a hyper parameter involved with
- 1:05:10these heads which is the head size and
- 1:05:13then here I'm initializing linear
- 1:05:15modules and I'm using bias equals false
- 1:05:18so these are just going to apply a
- 1:05:19matrix multiply with some fixed
- 1:05:21weights and now let me produce a key and
- 1:05:26q k and Q by forwarding these modules on
- 1:05:29X so the size of this will now
- 1:05:32become B by T by 16 because that is the
- 1:05:36head size and the same here B by T by
- 1:05:4416 so this being the head size so you
- 1:05:47see here that when I forward this linear
- 1:05:49on top of my X all the tokens in all the
- 1:05:52positions in the B BYT Arrangement all
- 1:05:55of them them in parallel and
- 1:05:57independently produce a key and a query
- 1:05:59so no communication has happened
- 1:06:01yet but the communication comes now all
- 1:06:04the queries will do product with all the
- 1:06:07keys so basically what we want is we
- 1:06:09want way now or the affinities between
- 1:06:12these to be query multiplying key but we
- 1:06:16have to be careful with uh we can't
- 1:06:18Matrix multiply this we actually need to
- 1:06:20transpose uh K but we have to be also
- 1:06:23careful because these are when you have
- 1:06:25The Bash Dimension so in particular we
- 1:06:27want to transpose uh the last two
- 1:06:30dimensions dimension1 and dimension -2
- 1:06:33so
- 1:06:36-21 and so this Matrix multiply now will
- 1:06:40basically do the following B by T by
- 1:06:4416 Matrix multiplies B by 16 by T to
- 1:06:49give us B by T by
- 1:06:53T right
- 1:06:56so for every row of B we're now going to
- 1:06:58have a t Square Matrix giving us the
- 1:07:01affinities and these are now the way so
- 1:07:04they're not zeros they are now coming
- 1:07:06from this dot product between the keys
- 1:07:08and the queries so this can now run I
- 1:07:11can I can run this and the weighted
- 1:07:13aggregation now is a function in a data
- 1:07:16Bandon manner between the keys and
- 1:07:18queries of these nodes so just
- 1:07:20inspecting what happened
- 1:07:22here the way takes on this form
- 1:07:26and you see that before way was uh just
- 1:07:29a constant so it was applied in the same
- 1:07:31way to all the batch elements but now
- 1:07:33every single batch elements will have
- 1:07:34different sort of we because uh every
- 1:07:37single batch element contains different
- 1:07:39uh tokens at different positions and so
- 1:07:41this is not data dependent so when we
- 1:07:44look at just the zeroth uh Row for
- 1:07:47example in the input these are the
- 1:07:49weights that came out and so you can see
- 1:07:51now that they're not just exactly
- 1:07:53uniform um and in particular as an
- 1:07:55example here for the last row this was
- 1:07:58the eighth token and the eighth token
- 1:08:00knows what content it has and it knows
- 1:08:02at what position it's in and now the E
- 1:08:04token based on that uh creates a query
- 1:08:08hey I'm looking for this kind of stuff
- 1:08:10um I'm a vowel I'm on the E position I'm
- 1:08:12looking for any consonant at positions
- 1:08:14up to four and then all the nodes get to
- 1:08:18emit keys and maybe one of the channels
- 1:08:20could be I am a I am a consonant and I
- 1:08:23am in a position up to four and that
- 1:08:25that key would have a high number in
- 1:08:27that specific Channel and that's how the
- 1:08:29query and the key when they do product
- 1:08:31they can find each other and create a
- 1:08:33high affinity and when they have a high
- 1:08:35Affinity like say uh this token was
- 1:08:38pretty interesting to uh to this eighth
- 1:08:41token when they have a high Affinity
- 1:08:43then through the softmax I will end up
- 1:08:45aggregating a lot of its information
- 1:08:47into my position and so I'll get to
- 1:08:49learn a lot about
- 1:08:51it now just this we're looking at way
- 1:08:55after this has already happened um let
- 1:08:59me erase this operation as well so let
- 1:09:01me erase the masking and the softmax
- 1:09:03just to show you the under the hood
- 1:09:04internals and how that works so without
- 1:09:07the masking in the softmax Whey comes
- 1:09:09out like this right this is the outputs
- 1:09:11of the do products um and these are the
- 1:09:14raw outputs and they take on values from
- 1:09:15negative you know two to positive two
- 1:09:18Etc so that's the raw interactions and
- 1:09:21raw affinities between all the nodes but
- 1:09:24now if I'm going if I'm a fifth node I
- 1:09:26will not want to aggregate anything from
- 1:09:28the sixth node seventh node and the
- 1:09:30eighth node so actually we use the upper
- 1:09:32triangular masking so those are not
- 1:09:35allowed to
- 1:09:37communicate and now we actually want to
- 1:09:40have a nice uh distribution uh so we
- 1:09:42don't want to aggregate negative .11 of
- 1:09:45this node that's crazy so instead we
- 1:09:47exponentiate and normalize and now we
- 1:09:49get a nice distribution that sums to one
- 1:09:51and this is telling us now in the data
- 1:09:52dependent manner how much of information
- 1:09:54to aggregate from any of these tokens in
- 1:09:56the
- 1:09:58past so that's way and it's not zeros
- 1:10:01anymore but but it's calculated in this
- 1:10:04way now there's one more uh part to a
- 1:10:08single self attention head and that is
- 1:10:10that when we do the aggregation we don't
- 1:10:12actually aggregate the tokens exactly we
- 1:10:15aggregate we produce one more value here
- 1:10:17and we call that the
- 1:10:20value so in the same way that we
- 1:10:22produced p and query we're also going to
- 1:10:23create a value
- 1:10:26and
- 1:10:26then here we don't
- 1:10:30aggregate X we calculate a v which is
- 1:10:34just achieved by uh propagating this
- 1:10:37linear on top of X again and then we
- 1:10:40output way multiplied by V so V is the
- 1:10:44elements that we aggregate or the the
- 1:10:46vectors that we aggregate instead of the
- 1:10:47raw
- 1:10:48X and now of course uh this will make it
- 1:10:51so that the output here of this single
- 1:10:53head will be 16 dimensional because that
- 1:10:55is the head
- 1:10:57size so you can think of X as kind of
- 1:10:59like private information to this token
- 1:11:01if you if you think about it that way so
- 1:11:03X is kind of private to this token so
- 1:11:06I'm a fifth token at some and I have
- 1:11:08some identity and uh my information is
- 1:11:11kept in Vector X and now for the
- 1:11:14purposes of the single head here's what
- 1:11:16I'm interested in here's what I have and
- 1:11:20if you find me interesting here's what I
- 1:11:21will communicate to you and that's
- 1:11:23stored in v and so V is the thing that
- 1:11:26gets aggregated for the purposes of this
- 1:11:28single head between the different
- 1:11:30notes and that's uh basically the self
- 1:11:34attention mechanism this is this is what
- 1:11:36it does there are a few notes that I
- 1:11:39would make like to make about attention
- 1:11:41number one attention is a communication
- 1:11:44mechanism you can really think about it
- 1:11:46as a communication mechanism where you
- 1:11:48have a number of nodes in a directed
- 1:11:50graph where basically you have edges
- 1:11:52pointed between noes like
- 1:11:53this and what happens is every node has
- 1:11:56some Vector of information and it gets
- 1:11:58to aggregate information via a weighted
- 1:12:01sum from all of the nodes that point to
- 1:12:03it and this is done in a data dependent
- 1:12:06manner so depending on whatever data is
- 1:12:08actually stored that you should not at
- 1:12:09any point in time now our graph doesn't
- 1:12:13look like this our graph has a different
- 1:12:15structure we have eight nodes because
- 1:12:17the block size is eight and there's
- 1:12:18always eight to
- 1:12:20tokens and uh the first node is only
- 1:12:23pointed to by itself the second node is
- 1:12:25pointed to by the first node and itself
- 1:12:27all the way up to the eighth node which
- 1:12:29is pointed to by all the previous nodes
- 1:12:32and itself and so that's the structure
- 1:12:34that our directed graph has or happens
- 1:12:37happens to have in Auto regressive sort
- 1:12:38of scenario like language modeling but
- 1:12:41in principle attention can be applied to
- 1:12:42any arbitrary directed graph and it's
- 1:12:44just a communication mechanism between
- 1:12:46the nodes the second note is that notice
- 1:12:48that there is no notion of space so
- 1:12:51attention simply acts over like a set of
- 1:12:53vectors in this graph and so by default
- 1:12:56these nodes have no idea where they are
- 1:12:58positioned in the space and that's why
- 1:12:59we need to encode them positionally and
- 1:13:02sort of give them some information that
- 1:13:03is anchored to a specific position so
- 1:13:05that they sort of know where they are
- 1:13:08and this is different than for example
- 1:13:09from convolution because if you're run
- 1:13:11for example a convolution operation over
- 1:13:13some input there's a very specific sort
- 1:13:15of layout of the information in space
- 1:13:18and the convolutional filters sort of
- 1:13:20act in space and so it's it's not like
- 1:13:23an attention in ATT ention is just a set
- 1:13:26of vectors out there in space they
- 1:13:27communicate and if you want them to have
- 1:13:29a notion of space you need to
- 1:13:31specifically add it which is what we've
- 1:13:33done when we calculated the um relative
- 1:13:36the positional encode encodings and
- 1:13:38added that information to the vectors
- 1:13:40the next thing that I hope is very clear
- 1:13:41is that the elements across the batch
- 1:13:43Dimension which are independent examples
- 1:13:45never talk to each other they're always
- 1:13:47processed independently and this is a
- 1:13:49batched matrix multiply that applies
- 1:13:51basically a matrix multiplication uh
- 1:13:53kind of in parallel across the batch
- 1:13:54dimension so maybe it would be more
- 1:13:56accurate to say that in this analogy of
- 1:13:58a directed graph we really have because
- 1:14:00the back size is four we really have
- 1:14:03four separate pools of eight nodes and
- 1:14:05those eight nodes only talk to each
- 1:14:07other but in total there's like 32 nodes
- 1:14:08that are being processed uh but there's
- 1:14:11um sort of four separate pools of eight
- 1:14:13you can look at it that way the next
- 1:14:15note is that here in the case of
- 1:14:18language modeling uh we have this
- 1:14:20specific uh structure of directed graph
- 1:14:22where the future tokens will not
- 1:14:24communicate to the Past tokens but this
- 1:14:27doesn't necessarily have to be the
- 1:14:28constraint in the general case and in
- 1:14:30fact in many cases you may want to have
- 1:14:32all of the uh noes talk to each other uh
- 1:14:35fully so as an example if you're doing
- 1:14:37sentiment analysis or something like
- 1:14:38that with a Transformer you might have a
- 1:14:40number of tokens and you may want to
- 1:14:42have them all talk to each other fully
- 1:14:45because later you are predicting for
- 1:14:46example the sentiment of the sentence
- 1:14:49and so it's okay for these NOS to talk
- 1:14:50to each other and so in those cases you
- 1:14:53will use an encoder block of self
- 1:14:55attention and uh all it means that it's
- 1:14:58an encoder block is that you will delete
- 1:15:00this line of code allowing all the noes
- 1:15:02to completely talk to each other what
- 1:15:04we're implementing here is sometimes
- 1:15:06called a decoder block and it's called a
- 1:15:09decoder because it is sort of like a
- 1:15:12decoding language and it's got this
- 1:15:15autor regressive format where you have
- 1:15:17to mask with the Triangular Matrix so
- 1:15:19that uh nodes from the future never talk
- 1:15:22to the Past because they would give away
- 1:15:24the answer
- 1:15:25and so basically in encoder blocks you
- 1:15:27would delete this allow all the noes to
- 1:15:29talk in decoder blocks this will always
- 1:15:31be present so that you have this
- 1:15:33triangular structure uh but both are
- 1:15:35allowed and attention doesn't care
- 1:15:36attention supports arbitrary
- 1:15:38connectivity between nodes the next
- 1:15:40thing I wanted to comment on is you keep
- 1:15:41me you keep hearing me say attention
- 1:15:43self attention Etc there's actually also
- 1:15:45something called cross attention what is
- 1:15:47the
- 1:15:47difference
- 1:15:49so basically the reason this attention
- 1:15:52is self attention is because because the
- 1:15:55keys queries and the values are all
- 1:15:57coming from the same Source from X so
- 1:16:01the same Source X produces Keys queries
- 1:16:03and values so these nodes are self
- 1:16:05attending but in principle attention is
- 1:16:08much more General than that so for
- 1:16:10example an encoder decoder Transformers
- 1:16:12uh you can have a case where the queries
- 1:16:15are produced from X but the keys and the
- 1:16:17values come from a whole separate
- 1:16:18external source and sometimes from uh
- 1:16:21encoder blocks that encode some context
- 1:16:23that we'd like to condition on
- 1:16:25and so the keys and the values will
- 1:16:26actually come from a whole separate
- 1:16:28Source those are nodes on the side and
- 1:16:31here we're just producing queries and
- 1:16:32we're reading off information from the
- 1:16:34side so cross attention is used when
- 1:16:37there's a separate source of nodes we'd
- 1:16:40like to pull information from into our
- 1:16:42nodes and it's self attention if we just
- 1:16:45have nodes that would like to look at
- 1:16:46each other and talk to each other so
- 1:16:48this attention here happens to be self
- 1:16:51attention but in principle um attention
- 1:16:55is a lot more General okay and the last
- 1:16:57note at this stage is if we come to the
- 1:16:59attention is all need paper here we've
- 1:17:01already implemented attention so given
- 1:17:03query key and value we've U multiplied
- 1:17:06the query and a key we've soft maxed it
- 1:17:09and then we are aggregating the values
- 1:17:11there's one more thing that we're
- 1:17:12missing here which is the dividing by
- 1:17:13one / square root of the head size the
- 1:17:16DK here is the head size why are they
- 1:17:18doing this finds this important so they
- 1:17:21call it the scaled attention and it's
- 1:17:24kind of like an important normalization
- 1:17:25to basically
- 1:17:26have the problem is if you have unit gsh
- 1:17:29and inputs so zero mean unit variance K
- 1:17:32and Q are unit gashin then if you just
- 1:17:34do we naively then you see that your we
- 1:17:37actually will be uh the variance will be
- 1:17:38on the order of head size which in our
- 1:17:40case is 16 but if you multiply by one
- 1:17:43over head size square root so this is
- 1:17:45square root and this is one
- 1:17:47over then the variance of we will be one
- 1:17:50so it will be
- 1:17:52preserved now why is this important
- 1:17:54you'll not notice that way
- 1:17:56here will feed into
- 1:17:58softmax and so it's really important
- 1:18:00especially at initialization that we be
- 1:18:03fairly diffuse so in our case here we
- 1:18:06sort of locked out here and we had a
- 1:18:10fairly diffuse numbers here so um like
- 1:18:13this now the problem is that because of
- 1:18:15softmax if weight takes on very positive
- 1:18:18and very negative numbers inside it
- 1:18:20softmax will actually converge towards
- 1:18:22one hot vectors and so I can illustrate
- 1:18:25that here um say we are applying softmax
- 1:18:29to a tensor of values that are very
- 1:18:31close to zero then we're going to get a
- 1:18:33diffuse thing out of
- 1:18:34softmax but the moment I take the exact
- 1:18:36same thing and I start sharpening it
- 1:18:38making it bigger by multiplying these
- 1:18:40numbers by eight for example you'll see
- 1:18:42that the softmax will start to sharpen
- 1:18:44and in fact it will sharpen towards the
- 1:18:46max so it will sharpen towards whatever
- 1:18:48number here is the highest and so um
- 1:18:51basically we don't want these values to
- 1:18:52be too extreme especially at
- 1:18:53initialization otherwise softmax will be
- 1:18:55way too peaky and um you're basically
- 1:18:58aggregating um information from like a
- 1:19:01single node every node just agregates
- 1:19:03information from a single other node
- 1:19:04that's not what we want especially at
- 1:19:06initialization and so the scaling is
- 1:19:08used just to control the variance at
- 1:19:11initialization okay so having said all
- 1:19:13that let's now take our self attention
- 1:19:15knowledge and let's uh take it for a
- 1:19:17spin so here in the code I created this
- 1:19:19head module and it implements a single
- 1:19:22head of self attention so you give it a
- 1:19:24head size and then here it creates the
- 1:19:26key query and the value linear layers
- 1:19:29typically people don't use biases in
- 1:19:31these uh so those are the linear
- 1:19:33projections that we're going to apply to
- 1:19:34all of our nodes now here I'm creating
- 1:19:37this Trill variable Trill is not a
- 1:19:40parameter of the module so in sort of
- 1:19:41pytorch naming conventions uh this is
- 1:19:43called a buffer it's not a parameter and
- 1:19:46you have to call it you have to assign
- 1:19:47it to the module using a register buffer
- 1:19:49so that creates the trill uh the triang
- 1:19:52lower triangular Matrix and we're given
- 1:19:55the input X this should look very
- 1:19:56familiar now we calculate the keys the
- 1:19:58queries we C calculate the attention
- 1:20:00scores inside way uh we normalize it so
- 1:20:03we're using scaled attention here then
- 1:20:06we make sure that uh future doesn't
- 1:20:08communicate with the past so this makes
- 1:20:10it a decoder block and then softmax and
- 1:20:13then aggregate the value and
- 1:20:15output then here in the language model
- 1:20:17I'm creating a head in the Constructor
- 1:20:20and I'm calling it self attention head
- 1:20:22and the head size I'm going to keep as
- 1:20:24the same and embed just for
- 1:20:27now and then here once we've encoded the
- 1:20:31information with the token embeddings
- 1:20:32and the position embeddings we're simply
- 1:20:34going to feed it into the self attention
- 1:20:36head and then the output of that is
- 1:20:38going to go into uh the decoder language
- 1:20:42modeling head and create the logits so
- 1:20:44this the sort of the simplest way to
- 1:20:46plug in a self attention component uh
- 1:20:49into our Network right now I had to make
- 1:20:51one more change which is that here in
- 1:20:55the generate uh we have to make sure
- 1:20:57that our idx that we feed into the model
- 1:21:01because now we're using positional
- 1:21:02embeddings we can never have more than
- 1:21:04block size coming in because if idx is
- 1:21:07more than block size then our position
- 1:21:09embedding table is going to run out of
- 1:21:11scope because it only has embeddings for
- 1:21:12up to block size and so therefore I
- 1:21:15added some uh code here to crop the
- 1:21:17context that we're going to feed into
- 1:21:20self um so that uh we never pass in more
- 1:21:23than block siiz elements
- 1:21:25so those are the changes and let's Now
- 1:21:27train the network okay so I also came up
- 1:21:29to the script here and I decreased the
- 1:21:30learning rate because uh the self
- 1:21:32attention can't tolerate very very high
- 1:21:34learning rates and then I also increased
- 1:21:36number of iterations because the
- 1:21:37learning rate is lower and then I
- 1:21:39trained it and previously we were only
- 1:21:41able to get to up to 2.5 and now we are
- 1:21:43down to 2.4 so we definitely see a
- 1:21:46little bit of an improvement from 2.5 to
- 1:21:482.4 roughly uh but the text is still not
- 1:21:51amazing so clearly the self attention
- 1:21:53head is doing some useful communication
- 1:21:56but um we still have a long way to go
- 1:21:59okay so now we've implemented the scale.
- 1:22:01product attention now next up and the
- 1:22:02attention is all you need paper there's
- 1:22:05something called multi-head attention
- 1:22:07and what is multi-head attention it's
- 1:22:09just applying multiple attentions in
- 1:22:11parallel and concatenating their results
- 1:22:13so they have a little bit of diagram
- 1:22:15here I don't know if this is super clear
- 1:22:18it's really just multiple attentions in
- 1:22:20parallel so let's Implement that fairly
- 1:22:23straightforward
- 1:22:25if we want a multi-head attention then
- 1:22:27we want multiple heads of self attention
- 1:22:28running in parallel so in pytorch we can
- 1:22:32do this by simply creating multiple
- 1:22:35heads so however heads how however many
- 1:22:38heads you want and then what is the head
- 1:22:39size of each and then we run all of them
- 1:22:43in parallel into a list and simply
- 1:22:46concatenate all of the outputs and we're
- 1:22:48concatenating over the channel
- 1:22:50Dimension so the way this looks now is
- 1:22:53we don't have just a single ATT
- 1:22:56that uh has a hit size of 32 because
- 1:22:59remember n Ed is
- 1:23:0032 instead of having one Communication
- 1:23:03channel we now have four communication
- 1:23:06channels in parallel and each one of
- 1:23:08these communication channels typically
- 1:23:10will be uh smaller uh correspondingly so
- 1:23:14because we have four communication
- 1:23:15channels we want eight dimensional self
- 1:23:18attention and so from each Communication
- 1:23:20channel we're going to together eight
- 1:23:22dimensional vectors and then we have
- 1:23:23four of them and that concatenates to
- 1:23:25give us 32 which is the original and
- 1:23:28embed and so this is kind of similar to
- 1:23:30um if you're familiar with convolutions
- 1:23:32this is kind of like a group convolution
- 1:23:34uh because basically instead of having
- 1:23:36one large convolution we do convolution
- 1:23:38in groups and uh that's multi-headed
- 1:23:40self
- 1:23:41attention and so then here we just use
- 1:23:44essay heads self attention heads instead
- 1:23:47now I actually ran it and uh scrolling
- 1:23:51down I ran the same thing and then we
- 1:23:53now get this down to 2.28 roughly and
- 1:23:57the output is still the generation is
- 1:23:58still not amazing but clearly the
- 1:24:00validation loss is improving because we
- 1:24:02were at 2.4 just now and so it helps to
- 1:24:05have multiple communication channels
- 1:24:07because obviously these tokens have a
- 1:24:09lot to talk about they want to find the
- 1:24:11consonants the vowels they want to find
- 1:24:13the vowels just from certain positions
- 1:24:15uh they want to find any kinds of
- 1:24:17different things and so it helps to
- 1:24:19create multiple independent channels of
- 1:24:20communication gather lots of different
- 1:24:22types of data and then uh decode the
- 1:24:25output now going back to the paper for a
- 1:24:27second of course I didn't explain this
- 1:24:28figure in full detail but we are
- 1:24:30starting to see some components of what
- 1:24:32we've already implemented we have the
- 1:24:33positional encodings the token encodings
- 1:24:35that add we have the masked multi-headed
- 1:24:37attention implemented now here's another
- 1:24:41multi-headed attention which is a cross
- 1:24:42attention to an encoder which we haven't
- 1:24:45we're not going to implement in this
- 1:24:46case I'm going to come back to that
- 1:24:48later but I want you to notice that
- 1:24:50there's a feed forward part here and
- 1:24:52then this is grouped into a block that
- 1:24:53gets repeat it again and again now the
- 1:24:56feedforward part here is just a simple
- 1:24:57uh multi-layer perceptron
- 1:25:00um so the multi-headed so here position
- 1:25:04wise feed forward networks is just a
- 1:25:06simple little MLP so I want to start
- 1:25:08basically in a similar fashion also
- 1:25:10adding computation into the network and
- 1:25:13this computation is on a per node level
- 1:25:16so I've already implemented it and you
- 1:25:18can see the diff highlighted on the left
- 1:25:20here when I've added or changed things
- 1:25:22now before we had the self multi-headed
- 1:25:25self attention that did the
- 1:25:26communication but we went way too fast
- 1:25:28to calculate the logits so the tokens
- 1:25:31looked at each other but didn't really
- 1:25:32have a lot of time to think on what they
- 1:25:35found from the other tokens and so what
- 1:25:38I've implemented here is a little feet
- 1:25:40forward single layer and this little
- 1:25:42layer is just a linear followed by a Rel
- 1:25:45nonlinearity and that's that's it so
- 1:25:48it's just a little layer and then I call
- 1:25:50it feed
- 1:25:52forward um and embed
- 1:25:54and then this feed forward is just
- 1:25:56called sequentially right after the self
- 1:25:58attention so we self attend then we feed
- 1:26:01forward and you'll notice that the feet
- 1:26:02forward here when it's applying linear
- 1:26:04this is on a per token level all the
- 1:26:06tokens do this independently so the self
- 1:26:09attention is the communication and then
- 1:26:11once they've gathered all the data now
- 1:26:13they need to think on that data
- 1:26:15individually and so that's what feed
- 1:26:16forward is doing and that's why I've
- 1:26:18added it here now when I train this the
- 1:26:21validation LW actually continues to go
- 1:26:23down now to 2. 24 which is down from
- 1:26:262.28 uh the output still look kind of
- 1:26:28terrible but at least we've improved the
- 1:26:31situation and so as a preview we're
- 1:26:34going to now start to intersperse the
- 1:26:37communication with the computation and
- 1:26:39that's also what the Transformer does
- 1:26:42when it has blocks that communicate and
- 1:26:44then compute and it groups them and
- 1:26:46replicates them okay so let me show you
- 1:26:49what we'd like to do we'd like to do
- 1:26:51something like this we have a block and
- 1:26:53this block is is basically this part
- 1:26:55here except for the cross
- 1:26:57attention now the block basically
- 1:26:59intersperses communication and then
- 1:27:01computation the computation the
- 1:27:03communication is done using multi-headed
- 1:27:05selfelf attention and then the
- 1:27:07computation is done using a feed forward
- 1:27:08Network on all the tokens
- 1:27:11independently now what I've added here
- 1:27:14also is you'll
- 1:27:16notice this takes the number of
- 1:27:18embeddings in the embedding Dimension
- 1:27:19and number of heads that we would like
- 1:27:21which is kind of like group size in
- 1:27:22group convolution and and I'm saying
- 1:27:24that number of heads we'd like is four
- 1:27:26and so because this is 32 we calculate
- 1:27:29that because this is 32 the number of
- 1:27:31heads should be four um the head size
- 1:27:34should be eight so that everything sort
- 1:27:36of works out Channel wise um so this is
- 1:27:39how the Transformer structures uh sort
- 1:27:41of the uh the sizes typically so the
- 1:27:44head size will become eight and then
- 1:27:45this is how we want to intersperse them
- 1:27:47and then here I'm trying to create
- 1:27:49blocks which is just a sequential
- 1:27:51application of block block block so that
- 1:27:53we're interspersing communication feed
- 1:27:55forward many many times and then finally
- 1:27:57we decode now I actually tried to run
- 1:28:01this and the problem is this doesn't
- 1:28:02actually give a very good uh answer and
- 1:28:05very good result and the reason for that
- 1:28:07is we're start starting to actually get
- 1:28:09like a pretty deep neural net and deep
- 1:28:11neural Nets uh suffer from optimization
- 1:28:13issues and I think that's what we're
- 1:28:14kind of like slightly starting to run
- 1:28:16into so we need one more idea that we
- 1:28:18can borrow from the um Transformer paper
- 1:28:21to resolve those difficulties now there
- 1:28:23are two optimizations that dramatically
- 1:28:25help with the depth of these networks
- 1:28:27and make sure that the networks remain
- 1:28:29optimizable let's talk about the first
- 1:28:31one the first one in this diagram is you
- 1:28:33see this Arrow here and then this arrow
- 1:28:36and this Arrow those are skip
- 1:28:38connections or sometimes called residual
- 1:28:40connections they come from this paper uh
- 1:28:43the presidual learning for image
- 1:28:44recognition from about
- 1:28:462015 uh that introduced the concept now
- 1:28:51these are basically what it means is you
- 1:28:53transform data but then you have a skip
- 1:28:55connection with addition from the
- 1:28:57previous features now the way I like to
- 1:29:00visualize it uh that I prefer is the
- 1:29:03following here the computation happens
- 1:29:05from the top to bottom and basically you
- 1:29:08have this uh residual pathway and you
- 1:29:11are free to Fork off from the residual
- 1:29:13pathway perform some computation and
- 1:29:15then project back to the residual
- 1:29:16pathway via addition and so you go from
- 1:29:19the the uh inputs to the targets only
- 1:29:22via plus and plus plus and the reason
- 1:29:25this is useful is because during back
- 1:29:27propagation remember from our microG
- 1:29:29grad video earlier addition distributes
- 1:29:32gradients equally to both of its
- 1:29:34branches that that fed as the input and
- 1:29:37so the supervision or the gradients from
- 1:29:40the loss basically hop through every
- 1:29:43addition node all the way to the input
- 1:29:46and then also Fork off into the residual
- 1:29:50blocks but basically you have this
- 1:29:52gradient Super Highway that goes
- 1:29:53directly from the supervision all the
- 1:29:55way to the input unimpeded and then
- 1:29:58these viral blocks are usually
- 1:29:59initialized in the beginning so they
- 1:30:01contribute very very little if anything
- 1:30:03to the residual pathway they they are
- 1:30:05initialized that way so in the beginning
- 1:30:07they are sort of almost kind of like not
- 1:30:09there but then during the optimization
- 1:30:11they come online over time and they uh
- 1:30:14start to contribute but at least at the
- 1:30:17initialization you can go from directly
- 1:30:19supervision to the input gradient is
- 1:30:21unimpeded and just flows and then the
- 1:30:23blocks over time
- 1:30:24kick in and so that dramatically helps
- 1:30:27with the optimization so let's implement
- 1:30:29this so coming back to our block here
- 1:30:31basically what we want to do is we want
- 1:30:33to do xal
- 1:30:35X+ self attention and xal X+ self. feed
- 1:30:39forward so this is X and then we Fork
- 1:30:43off and do some communication and come
- 1:30:45back and we Fork off and we do some
- 1:30:46computation and come back so those are
- 1:30:49residual connections and then swinging
- 1:30:51back up here we also have to introd use
- 1:30:54this projection so nn.
- 1:30:57linear and uh this is going to be
- 1:31:00from after we concatenate this this is
- 1:31:03the prze and embed so this is the output
- 1:31:05of the self tension itself but then we
- 1:31:08actually want the uh to apply the
- 1:31:11projection and that's the
- 1:31:13result so the projection is just a
- 1:31:15linear transformation of the outcome of
- 1:31:16this
- 1:31:17layer so that's the projection back into
- 1:31:20the virual pathway and then here in a
- 1:31:22feet forward it's going to be the same
- 1:31:23same thing I could have a a self doot
- 1:31:26projection here as well but let me just
- 1:31:28simplify it and let me uh couple it
- 1:31:32inside the same sequential container and
- 1:31:34so this is the projection layer going
- 1:31:36back into the residual
- 1:31:38pathway and
- 1:31:40so that's uh well that's it so now we
- 1:31:43can train this so I implemented one more
- 1:31:44small change when you look into the
- 1:31:47paper again you see that the
- 1:31:49dimensionality of input and output is
- 1:31:51512 for them and they're saying that the
- 1:31:53inner layer here in the feet forward has
- 1:31:55dimensionality of 248 so there's a
- 1:31:57multiplier of four and so the inner
- 1:32:00layer of the feet forward Network should
- 1:32:02be multiplied by four in terms of
- 1:32:04Channel sizes so I came here and I
- 1:32:06multiplied four times embed here for the
- 1:32:08feed forward and then from four times
- 1:32:10nmed coming back down to nmed when we go
- 1:32:13back to the pro uh to the projection so
- 1:32:15adding a bit of computation here and
- 1:32:17growing that layer that is in the
- 1:32:19residual block on the side of the
- 1:32:21residual
- 1:32:22pathway and then I train this and we
- 1:32:24actually get down all the way to uh 2.08
- 1:32:27validation loss and we also see that
- 1:32:29network is starting to get big enough
- 1:32:30that our train loss is getting ahead of
- 1:32:32validation loss so we're starting to see
- 1:32:33like a little bit of
- 1:32:35overfitting and um our our
- 1:32:38um uh Generations here are still not
- 1:32:41amazing but at least you see that we can
- 1:32:42see like is here this now grief syn like
- 1:32:46this starts to almost look like English
- 1:32:48so um yeah we're starting to really get
- 1:32:50there okay and the second Innovation
- 1:32:52that is very helpful for optimizing very
- 1:32:54deep neural networks is right here so we
- 1:32:57have this addition now that's the
- 1:32:58residual part but this Norm is referring
- 1:33:00to something called layer Norm so layer
- 1:33:03Norm is implemented in pytorch it's a
- 1:33:04paper that came out a while back here
- 1:33:09um and layer Norm is very very similar
- 1:33:11to bash Norm so remember back to our
- 1:33:14make more series part three we
- 1:33:16implemented bash
- 1:33:17normalization and uh bash normalization
- 1:33:19basically just made sure that um Across
- 1:33:22The Bash dimension any individual neuron
- 1:33:25had unit uh Gan um distribution so it
- 1:33:30was zero mean and unit standard
- 1:33:32deviation one standard deviation output
- 1:33:35so what I did here is I'm copy pasting
- 1:33:37the bashor 1D that we developed in our
- 1:33:39make more series and see here we can
- 1:33:42initialize for example this module and
- 1:33:44we can have a batch of 32 100
- 1:33:47dimensional vectors feeding through the
- 1:33:48bachor layer so what this does is it
- 1:33:52guarantees that when we look at just the
- 1:33:54zeroth column it's a zero mean one
- 1:33:58standard deviation so it's normalizing
- 1:34:00every single column of this uh input now
- 1:34:04the rows are not uh going to be
- 1:34:06normalized by default because we're just
- 1:34:08normalizing columns so let's now
- 1:34:10Implement layer Norm uh it's very
- 1:34:12complicated look we come here we change
- 1:34:15this from zero to one so we don't
- 1:34:18normalize The Columns we normalize the
- 1:34:20rows and now we've implemented layer
- 1:34:23Norm
- 1:34:25so now the columns are not going to be
- 1:34:28normalized um but the rows are going to
- 1:34:31be normalized for every individual
- 1:34:33example it's 100 dimensional Vector is
- 1:34:35normalized uh in this way and because
- 1:34:38our computation Now does not span across
- 1:34:40examples we can delete all of this
- 1:34:43buffers stuff uh because uh we can
- 1:34:45always apply this operation and don't
- 1:34:48need to maintain any running buffers so
- 1:34:50we don't need the
- 1:34:52buffers uh we
- 1:34:54don't There's no distinction between
- 1:34:56training and test
- 1:34:58time uh and we don't need these running
- 1:35:00buffers we do keep gamma and beta we
- 1:35:03don't need the momentum we don't care if
- 1:35:05it's training or not and this is now a
- 1:35:08layer
- 1:35:09norm and it normalizes the rows instead
- 1:35:12of the columns and this here is
- 1:35:15identical to basically this here so
- 1:35:19let's now Implement layer Norm in our
- 1:35:21Transformer before I incorporate the
- 1:35:23layer Norm I just wanted to note that as
- 1:35:25I said very few details about the
- 1:35:27Transformer have changed in the last 5
- 1:35:28years but this is actually something
- 1:35:30that slightly departs from the original
- 1:35:31paper you see that the ADD and Norm is
- 1:35:34applied after the
- 1:35:36transformation but um in now it is a bit
- 1:35:40more uh basically common to apply the
- 1:35:42layer Norm before the transformation so
- 1:35:44there's a reshuffling of the layer Norms
- 1:35:46uh so this is called the prorm
- 1:35:48formulation and that's the one that
- 1:35:49we're going to implement as well so
- 1:35:50select deviation from the original paper
- 1:35:53basically we need two layer Norms layer
- 1:35:55Norm one is uh NN do layer norm and we
- 1:35:59tell it how many um what is the
- 1:36:01embedding Dimension and we need the
- 1:36:03second layer norm and then here the
- 1:36:06layer Norms are applied immediately on X
- 1:36:09so self. layer Norm one applied on X and
- 1:36:13self. layer Norm two applied on X before
- 1:36:15it goes into self attention and feed
- 1:36:18forward and uh the size of the layer
- 1:36:20Norm here is an ed so 32 so when the
- 1:36:23layer Norm is normalizing our features
- 1:36:26it is uh the normalization here uh
- 1:36:30happens the mean and the variance are
- 1:36:32taken over 32 numbers so the batch and
- 1:36:34the time act as batch Dimensions both of
- 1:36:37them so this is kind of like a per token
- 1:36:40um transformation that just normalizes
- 1:36:42the features and makes them a unit mean
- 1:36:46uh unit Gan at
- 1:36:48initialization but of course because
- 1:36:50these layer Norms inside it have these
- 1:36:52gamma and beta training
- 1:36:54parameters uh the layer Norm will U
- 1:36:57eventually create outputs that might not
- 1:36:59be unit gion but the optimization will
- 1:37:01determine that so for now this is the uh
- 1:37:05this is incorporating the layer norms
- 1:37:06and let's train them on okay so I let it
- 1:37:09run and we see that we get down to 2.06
- 1:37:12which is better than the previous 2.08
- 1:37:14so a slight Improvement by adding the
- 1:37:15layer norms and I'd expect that they
- 1:37:17help uh even more if we had bigger and
- 1:37:19deeper Network one more thing I forgot
- 1:37:21to add is that there should be a layer
- 1:37:23Norm here also typically as at the end
- 1:37:26of the Transformer and right before the
- 1:37:28final uh linear layer that decodes into
- 1:37:31vocabulary so I added that as well so at
- 1:37:35this stage we actually have a pretty
- 1:37:36complete uh Transformer according to the
- 1:37:38original paper and it's a decoder only
- 1:37:40Transformer I'll I'll talk about that in
- 1:37:42a second uh but at this stage uh the
- 1:37:44major pieces are in place so we can try
- 1:37:46to scale this up and see how well we can
- 1:37:47push this number now in order to scale
- 1:37:50out the model I had to perform some
- 1:37:51cosmetic changes here to make it nicer
- 1:37:54so I introduced this variable called n
- 1:37:56layer which just specifies how many
- 1:37:57layers of the blocks we're going to have
- 1:38:01I created a bunch of blocks and we have
- 1:38:02a new variable number of heads as well I
- 1:38:05pulled out the layer Norm here and uh so
- 1:38:07this is identical now one thing that I
- 1:38:10did briefly change is I added a Dropout
- 1:38:13so Dropout is something that you can add
- 1:38:15right before the residual connection
- 1:38:17back right before the connection back
- 1:38:19into the residual pathway so we can drop
- 1:38:22out that as l layer here we can drop out
- 1:38:26uh here at the end of the multi-headed
- 1:38:27exension as well and we can also drop
- 1:38:30out here uh when we calculate the um
- 1:38:34basically affinities and after the
- 1:38:36softmax we can drop out some of those so
- 1:38:38we can randomly prevent some of the
- 1:38:40nodes from
- 1:38:41communicating and so Dropout uh comes
- 1:38:43from this paper from 2014 or so and
- 1:38:49basically it takes your neural
- 1:38:50nut and it randomly every forward
- 1:38:53backward pass shuts off some subset of
- 1:38:56uh neurons so randomly drops them to
- 1:38:59zero and trains without them and what
- 1:39:02this does effectively is because the
- 1:39:04mask of what's being dropped out is
- 1:39:06changed every single forward backward
- 1:39:07pass it ends up kind of uh training an
- 1:39:11ensemble of sub networks and then at
- 1:39:13test time everything is fully enabled
- 1:39:15and kind of all of those sub networks
- 1:39:16are merged into a single Ensemble if you
- 1:39:18can if you want to think about it that
- 1:39:20way so I would read the paper to get the
- 1:39:22full detail for now we're just going to
- 1:39:24stay on the level of this is a
- 1:39:25regularization technique and I added it
- 1:39:28because I'm about to scale up the model
- 1:39:30quite a bit and I was concerned about
- 1:39:32overfitting so now when we scroll up to
- 1:39:34the top uh we'll see that I changed a
- 1:39:36number of hyper parameters here about
- 1:39:38our neural nut so I made the batch size
- 1:39:40be much larger now it's 64 I changed the
- 1:39:43block size to be 256 so previously it
- 1:39:46was just eight eight characters of
- 1:39:47context now it is 256 characters of
- 1:39:50context to predict the 257th
- 1:39:54uh I brought down the learning rate a
- 1:39:55little bit because the neural net is now
- 1:39:57much bigger so I brought down the
- 1:39:58learning rate the embedding Dimension is
- 1:40:01now 384 and there are six heads so 384
- 1:40:05divide 6 means that every head is 64
- 1:40:08dimensional as it as a standard and then
- 1:40:11there's going to be six layers of that
- 1:40:13and the Dropout will be at 02 so every
- 1:40:15forward backward pass 20% of all of
- 1:40:18these um intermediate calculations are
- 1:40:21disabled and dropped to zero
- 1:40:24and then I already trained this and I
- 1:40:25ran it so uh drum roll how well does it
- 1:40:28perform so let me just scroll up
- 1:40:31here we get a validation loss of
- 1:40:341.48 which is actually quite a bit of an
- 1:40:37improvement on what we had before which
- 1:40:38I think was 2.07 so it went from 2.07
- 1:40:41all the way down to 1.48 just by scaling
- 1:40:43up this neural nut with the code that we
- 1:40:45have and this of course ran for a lot
- 1:40:47longer this maybe trained for I want to
- 1:40:49say about 15 minutes on my a100 GPU so
- 1:40:52that's a pretty a GPU and if you don't
- 1:40:54have a GPU you're not going to be able
- 1:40:56to reproduce this uh on a CPU this would
- 1:40:59be um I would not run this on a CPU or
- 1:41:01MacBook or something like that you'll
- 1:41:03have to Brak down the number of uh
- 1:41:04layers and the embedding Dimension and
- 1:41:06so on uh but in about 15 minutes we can
- 1:41:09get this kind of a result and um I'm
- 1:41:12printing some of the Shakespeare here
- 1:41:15but what I did also is I printed 10,000
- 1:41:17characters so a lot more and I wrote
- 1:41:18them to a file and so here we see some
- 1:41:21of the outputs
- 1:41:24so it's a lot more recognizable as the
- 1:41:26input text file so the input text file
- 1:41:29just for reference looked like this so
- 1:41:31there's always like someone speaking in
- 1:41:33this manner and uh our predictions now
- 1:41:37take on that form except of course
- 1:41:40they're they're nonsensical when you
- 1:41:41actually read them
- 1:41:43so it is every crimp tap be a house oh
- 1:41:47those
- 1:41:48prepation we give
- 1:41:51heed um you know
- 1:41:56Oho sent me you mighty
- 1:41:59Lord anyway so you can read through this
- 1:42:02um it's nonsensical of course but this
- 1:42:04is just a Transformer trained on a
- 1:42:06character level for 1 million characters
- 1:42:09that come from Shakespeare so there's
- 1:42:10sort of like blabbers on in Shakespeare
- 1:42:12like manner but it doesn't of course
- 1:42:14make sense at this scale uh but I think
- 1:42:18I think still a pretty good
- 1:42:19demonstration of what's
- 1:42:20possible so now
- 1:42:24I think uh that kind of like concludes
- 1:42:26the programming section of this video we
- 1:42:28basically kind of uh did a pretty good
- 1:42:30job and um of implementing this
- 1:42:32Transformer uh but the picture doesn't
- 1:42:35exactly match up to what we've done so
- 1:42:37what's going on with all these digital
- 1:42:38Parts here so let me finish explaining
- 1:42:41this architecture and why it looks so
- 1:42:43funky basically what's happening here is
- 1:42:45what we implemented here is a decoder
- 1:42:47only Transformer so there's no component
- 1:42:50here this part is called the encoder and
- 1:42:52there's no cross attention block here
- 1:42:55our block only has a self attention and
- 1:42:58the feet forward so it is missing this
- 1:43:00third in between piece here this piece
- 1:43:03does cross attention so we don't have it
- 1:43:05and we don't have the encoder we just
- 1:43:07have the decoder and the reason we have
- 1:43:08a decoder only uh is because we are just
- 1:43:12uh generating text and it's
- 1:43:13unconditioned on anything we're just
- 1:43:15we're just blabbering on according to a
- 1:43:16given data set what makes it a decoder
- 1:43:19is that we are using the Triangular mask
- 1:43:21in our uh trans former so it has this
- 1:43:24Auto regressive property where we can
- 1:43:26just uh go and sample from it so the
- 1:43:28fact that it's using the Triangular
- 1:43:30triangular mask to mask out the
- 1:43:32attention makes it a decoder and it can
- 1:43:34be used for language modeling now the
- 1:43:37reason that the original paper had an
- 1:43:39incoder decoder architecture is because
- 1:43:41it is a machine translation paper so it
- 1:43:43is concerned with a different setting in
- 1:43:45particular it expects some uh tokens
- 1:43:49that encode say for example French and
- 1:43:52then it is expecting to decode the
- 1:43:54translation in English so so you
- 1:43:56typically these here are special tokens
- 1:43:59so you are expected to read in this and
- 1:44:02condition on it and then you start off
- 1:44:04the generation with a special token
- 1:44:05called start so this is a special new
- 1:44:08token um that you introduce and always
- 1:44:10place in the beginning and then the
- 1:44:12network is expected to Output neural
- 1:44:15networks are awesome and then a special
- 1:44:17end token to finish the
- 1:44:20generation so this part here will be
- 1:44:23decoded exactly as we we've done it
- 1:44:25neural networks are awesome will be
- 1:44:27identical to what we did but unlike what
- 1:44:29we did they wanton to condition the
- 1:44:32generation on some additional
- 1:44:34information and in that case this
- 1:44:36additional information is the French
- 1:44:38sentence that they should be
- 1:44:39translating so what they do now is they
- 1:44:42bring in the encoder now the encoder
- 1:44:45reads this part here so we're only going
- 1:44:48to take the part of French and we're
- 1:44:50going to uh create tokens from it
- 1:44:52exactly as we've seen in our video and
- 1:44:54we're going to put a Transformer on it
- 1:44:57but there's going to be no triangular
- 1:44:58mask and so all the tokens are allowed
- 1:45:00to talk to each other as much as they
- 1:45:02want and they're just encoding
- 1:45:04whatever's the content of this French uh
- 1:45:07sentence once they've encoded it they
- 1:45:10they basically come out in the top here
- 1:45:13and then what happens here is in our
- 1:45:14decoder which does the uh language
- 1:45:17modeling there's an additional
- 1:45:20connection here to the outputs of the
- 1:45:22encoder
- 1:45:23and that is brought in through a cross
- 1:45:26attention so the queries are still
- 1:45:28generated from X but now the keys and
- 1:45:30the values are coming from the side the
- 1:45:32keys and the values are coming from the
- 1:45:34top generated by the nodes that came
- 1:45:36outside of the de the encoder and those
- 1:45:40tops the keys and the values there the
- 1:45:42top of it feed in on a side into every
- 1:45:45single block of the decoder and so
- 1:45:47that's why there's an additional cross
- 1:45:49attention and really what it's doing is
- 1:45:51it's conditioning the decoding
- 1:45:53not just on the past of this current
- 1:45:55decoding but also on having seen the
- 1:45:59full fully encoded French um prompt sort
- 1:46:04of and so it's an encoder decoder model
- 1:46:06which is why we have those two
- 1:46:07Transformers an additional block and so
- 1:46:09on so we did not do this because we have
- 1:46:12no we have nothing to encode there's no
- 1:46:13conditioning we just have a text file
- 1:46:15and we just want to imitate it and
- 1:46:16that's why we are using a decoder only
- 1:46:19Transformer exactly as done in
- 1:46:21GPT okay okay so now I wanted to do a
- 1:46:24very brief walkthrough of nanog GPT
- 1:46:26which you can find in my GitHub and uh
- 1:46:28nanog GPT is basically two files of
- 1:46:30Interest there's train.py and model.py
- 1:46:33train.py is all the boilerplate code for
- 1:46:35training the network it is basically all
- 1:46:38the stuff that we had here it's the
- 1:46:40training loop it's just that it's a lot
- 1:46:42more complicated because we're saving
- 1:46:44and loading checkpoints and pre-trained
- 1:46:46weights and we are uh decaying the
- 1:46:48learning rate and compiling the model
- 1:46:50and using distributed training across
- 1:46:51multiple nodes or GP use so the training
- 1:46:54Pi gets a little bit more hairy
- 1:46:56complicated uh there's more options Etc
- 1:46:59but the model.py should look very very
- 1:47:01um similar to what we've done here in
- 1:47:04fact the model is is almost identical so
- 1:47:08first here we have the causal self
- 1:47:09attention block and all of this should
- 1:47:11look very very recognizable to you we're
- 1:47:13producing queries Keys values we're
- 1:47:16doing Dot products we're masking
- 1:47:18applying soft Maxs optionally dropping
- 1:47:20out and here we are pulling the wi the
- 1:47:23values what is different here is that in
- 1:47:25our code I have separated out the
- 1:47:30multi-headed detention into just a
- 1:47:31single individual head and then here I
- 1:47:34have multiple heads and I explicitly
- 1:47:36concatenate them whereas here uh all of
- 1:47:39it is implemented in a batched manner
- 1:47:41inside a single causal self attention
- 1:47:43and so we don't just have a b and a T
- 1:47:45and A C Dimension we also end up with a
- 1:47:47fourth dimension which is the heads and
- 1:47:50so it just gets a lot more sort of hairy
- 1:47:52because we have four dimensional array
- 1:47:54um tensors now but it is um equivalent
- 1:47:57mathematically so the exact same thing
- 1:47:59is happening as what we have it's just
- 1:48:01it's a bit more efficient because all
- 1:48:02the heads are now treated as a batch
- 1:48:04Dimension as
- 1:48:05well then we have the multier perceptron
- 1:48:08it's using the Galu nonlinearity which
- 1:48:10is defined here except instead of Ru and
- 1:48:13this is done just because opening I used
- 1:48:14it and I want to be able to load their
- 1:48:17checkpoints uh the blocks of the
- 1:48:19Transformer are identical to communicate
- 1:48:21in the compute phase as we saw and then
- 1:48:23the GPT will be identical we have the
- 1:48:25position encodings token encodings the
- 1:48:27blocks the layer Norm at the end uh the
- 1:48:30final linear layer and this should look
- 1:48:33all very recognizable and there's a bit
- 1:48:35more here because I'm loading
- 1:48:36checkpoints and stuff like that I'm
- 1:48:38separating out the parameters into those
- 1:48:40that should be weight decayed and those
- 1:48:42that
- 1:48:42shouldn't um but the generate function
- 1:48:44should also be very very similar so a
- 1:48:47few details are different but you should
- 1:48:48definitely be able to look at this uh
- 1:48:51file and be able to understand little
- 1:48:52the pieces now so let's now bring things
- 1:48:55back to chat GPT what would it look like
- 1:48:57if we wanted to train chat GPT ourselves
- 1:48:59and how does it relate to what we
- 1:49:00learned today well to train in chat GPT
- 1:49:03there are roughly two stages first is
- 1:49:05the pre-training stage and then the
- 1:49:07fine-tuning stage in the pre-training
- 1:49:09stage uh we are training on a large
- 1:49:12chunk of internet and just trying to get
- 1:49:14a first decoder only Transformer to
- 1:49:17babble text so it's very very similar to
- 1:49:20what we've done ourselves except we've
- 1:49:23done like a tiny little baby
- 1:49:24pre-training step um and so in our case
- 1:49:28uh this is how you print a number of
- 1:49:30parameters I printed it and it's about
- 1:49:3210 million so this Transformer that I
- 1:49:35created here to create little
- 1:49:37Shakespeare um Transformer was about 10
- 1:49:40million parameters our data set is
- 1:49:42roughly 1 million uh characters so
- 1:49:45roughly 1 million tokens but you have to
- 1:49:47remember that opening I is different
- 1:49:48vocabulary they're not on the Character
- 1:49:50level they use these um subword chunks
- 1:49:53of words and so they have a vocabulary
- 1:49:55of 50,000 roughly elements and so their
- 1:49:58sequences are a bit more condensed so
- 1:50:01our data set the Shakespeare data set
- 1:50:03would be probably around 300,000 uh
- 1:50:05tokens in the open AI vocabulary roughly
- 1:50:09so we trained about 10 million parameter
- 1:50:11model on roughly 300,000 tokens now when
- 1:50:14you go to the gpt3
- 1:50:16paper and you look at the Transformers
- 1:50:20that they trained they trained a number
- 1:50:22of trans Transformers of different sizes
- 1:50:24but the biggest Transformer here has 175
- 1:50:27billion parameters uh so ours is again
- 1:50:2910 million they used this number of
- 1:50:31layers in the Transformer this is the
- 1:50:34nmed this is the number of heads and
- 1:50:36this is the head size and then this is
- 1:50:39the batch size uh so ours was
- 1:50:4365 and the learning rate is similar now
- 1:50:46when they train this Transformer they
- 1:50:47trained on 300 billion tokens so again
- 1:50:51remember ours is about 300,000
- 1:50:53so this is uh about a millionfold
- 1:50:56increase and this number would not be
- 1:50:57even that large by today's standards
- 1:50:59you'd be going up uh 1 trillion and
- 1:51:01above so they are training a
- 1:51:04significantly larger
- 1:51:06model on uh a good chunk of the internet
- 1:51:10and that is the pre-training stage but
- 1:51:12otherwise these hyper parameters should
- 1:51:13be fairly recognizable to you and the
- 1:51:15architecture is actually like nearly
- 1:51:17identical to what we implemented
- 1:51:18ourselves but of course it's a massive
- 1:51:20infrastructure challenge to train this
- 1:51:22you're talking about typically thousands
- 1:51:24of gpus having to you know talk to each
- 1:51:27other to train models of this size so
- 1:51:29that's just a pre-training stage now
- 1:51:32after you complete the pre-training
- 1:51:33stage uh you don't get something that
- 1:51:35responds to your questions with answers
- 1:51:38and is not helpful and Etc you get a
- 1:51:40document
- 1:51:41completer right so it babbles but it
- 1:51:44doesn't Babble Shakespeare it babbles
- 1:51:46internet it will create arbitrary news
- 1:51:48articles and documents and it will try
- 1:51:50to complete documents because that's
- 1:51:51what it's trained for it's trying to
- 1:51:52complete the sequence so when you give
- 1:51:54it a question it would just uh
- 1:51:56potentially just give you more questions
- 1:51:58it would follow with more questions it
- 1:52:00will do whatever it looks like the some
- 1:52:02close document would do in the training
- 1:52:05data on the internet and so who knows
- 1:52:07you're getting kind of like undefined
- 1:52:08Behavior it might basically answer with
- 1:52:11to questions with other questions it
- 1:52:13might ignore your question it might just
- 1:52:15try to complete some news article it's
- 1:52:17totally unineed as we say so the second
- 1:52:20fine-tuning stage is to actually align
- 1:52:22it to be an assistant and uh this is the
- 1:52:25second stage and so this chat GPT block
- 1:52:28post from openi talks a little bit about
- 1:52:30how the stage is achieved we basically
- 1:52:34um there's roughly three steps to to
- 1:52:36this stage uh so what they do here is
- 1:52:39they start to collect training data that
- 1:52:41looks specifically like what an
- 1:52:42assistant would do so these are
- 1:52:44documents that have to format where the
- 1:52:46question is on top and then an answer is
- 1:52:47below and they have a large number of
- 1:52:50these but probably not on the order of
- 1:52:51the internet uh this is probably on the
- 1:52:53of maybe thousands of examples and so
- 1:52:58they they then fine-tune the model to
- 1:53:00basically only focus on documents that
- 1:53:03look like that and so you're starting to
- 1:53:05slowly align it so it's going to expect
- 1:53:07a question at the top and it's going to
- 1:53:08expect to complete the answer and uh
- 1:53:11these very very large models are very
- 1:53:13sample efficient during their
- 1:53:14fine-tuning so this actually somehow
- 1:53:16works but that's just step one that's
- 1:53:19just fine tuning so then they actually
- 1:53:20have more steps where okay the second
- 1:53:23step is you let the model respond and
- 1:53:25then different Raiders look at the
- 1:53:27different responses and rank them for
- 1:53:29their preference as to which one is
- 1:53:30better than the other they use that to
- 1:53:32train a reward model so they can predict
- 1:53:35uh basically using a different network
- 1:53:37how much of any candidate
- 1:53:39response would be desirable and then
- 1:53:43once they have a reward model they run
- 1:53:45po which is a form of polic policy
- 1:53:47gradient um reinforcement learning
- 1:53:49Optimizer to uh fine-tune this sampling
- 1:53:53policy uh so that the answers that the
- 1:53:55GP chat GPT now generates are expected
- 1:53:59to score a high reward according to the
- 1:54:02reward model and so basically there's a
- 1:54:04whole aligning stage here or fine-tuning
- 1:54:07stage it's got multiple steps in between
- 1:54:09there as well and it takes the model
- 1:54:11from being a document completer to a
- 1:54:14question answerer and that's like a
- 1:54:16whole separate stage a lot of this data
- 1:54:19is not available publicly it is internal
- 1:54:21to open AI and uh it's much harder to
- 1:54:24replicate this stage um and so that's
- 1:54:27roughly what would give you a chat GPT
- 1:54:29and nanog GPT focuses on the
- 1:54:31pre-training stage okay and that's
- 1:54:32everything that I wanted to cover today
- 1:54:35so we trained to summarize a decoder
- 1:54:38only Transformer following this famous
- 1:54:41paper attention is all you need from
- 1:54:432017 and so that's basically a GPT we
- 1:54:47trained it on Tiny Shakespeare and got
- 1:54:50sensible results
- 1:54:52all of the training code is
- 1:54:54roughly 200 lines of code I will be
- 1:54:57releasing this um code base so also it
- 1:55:01comes with all the git log commits along
- 1:55:04the way as we built it
- 1:55:05up in addition to this code I'm going to
- 1:55:08release the um notebook of course the
- 1:55:10Google collab and I hope that gave you a
- 1:55:13sense for how you can train um these
- 1:55:16models like say gpt3 that will be um
- 1:55:19architecturally basically identical to
- 1:55:20what we have but they are somewhere
- 1:55:22between 10,000 and 1 million times
- 1:55:24bigger depending on how you count and so
- 1:55:27uh that's all I have for now uh we did
- 1:55:30not talk about any of the fine-tuning
- 1:55:32stages that would typically go on top of
- 1:55:33this so if you're interested in
- 1:55:35something that's not just language
- 1:55:36modeling but you actually want to you
- 1:55:38know say perform tasks um or you want
- 1:55:40them to be aligned in a specific way or
- 1:55:43you want um to detect sentiment or
- 1:55:45anything like that basically anytime you
- 1:55:47don't want something that's just a
- 1:55:48document completer you have to complete
- 1:55:50further stages of fine tuning which did
- 1:55:52not cover uh and that could be simple
- 1:55:55supervised fine tuning or it can be
- 1:55:57something more fancy like we see in chat
- 1:55:58jpt where we actually train a reward
- 1:56:00model and then do rounds of Po to uh
- 1:56:03align it with respect to the reward
- 1:56:04model so there's a lot more that can be
- 1:56:06done on top of it I think for now we're
- 1:56:08starting to get to about two hours Mark
- 1:56:10uh so I'm going to um kind of finish
- 1:56:13here uh I hope you enjoyed the lecture
- 1:56:15uh and uh yeah go forth and transform
- 1:56:18see you later
About this transcript
This page contains the full transcript of Let's build GPT: from scratch, in code, spelled out. by Andrej Karpathy, generated from the public captions YouTube serves with the video. The transcript has 21,030 words across 2,955 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.