Let's build the GPT Tokenizer — Transcript
Full transcript
- 0:00hi everyone so in this video I'd like us
- 0:02to cover the process of tokenization in
- 0:04large language models now you see here
- 0:06that I have a set face and that's
- 0:08because uh tokenization is my least
- 0:10favorite part of working with large
- 0:11language models but unfortunately it is
- 0:13necessary to understand in some detail
- 0:15because it it is fairly hairy gnarly and
- 0:17there's a lot of hidden foot guns to be
- 0:19aware of and a lot of oddness with large
- 0:21language models typically traces back to
- 0:24tokenization so what is
- 0:26tokenization now in my previous video
- 0:28Let's Build GPT from scratch uh we
- 0:31actually already did tokenization but we
- 0:33did a very naive simple version of
- 0:35tokenization so when you go to the
- 0:37Google colab for that video uh you see
- 0:40here that we loaded our training set and
- 0:43our training set was this uh Shakespeare
- 0:45uh data set now in the beginning the
- 0:48Shakespeare data set is just a large
- 0:49string in Python it's just text and so
- 0:52the question is how do we plug text into
- 0:54large language models and in this case
- 0:58here we created a vocabulary of 65
- 1:01possible characters that we saw occur in
- 1:03this string these were the possible
- 1:05characters and we saw that there are 65
- 1:07of them and then we created a a lookup
- 1:10table for converting from every possible
- 1:13character a little string piece into a
- 1:16token an
- 1:17integer so here for example we tokenized
- 1:20the string High there and we received
- 1:23this sequence of
- 1:24tokens and here we took the first 1,000
- 1:27characters of our data set and we
- 1:29encoded it into tokens and because it is
- 1:32this is character level we received
- 1:341,000 tokens in a sequence so token 18
- 1:3847
- 1:40Etc now later we saw that the way we
- 1:43plug these tokens into the language
- 1:45model is by using an embedding
- 1:48table and so basically if we have 65
- 1:51possible tokens then this embedding
- 1:53table is going to have 65 rows and
- 1:56roughly speaking we're taking the
- 1:58integer associated with every single
- 1:59sing Le token we're using that as a
- 2:01lookup into this table and we're
- 2:04plucking out the corresponding row and
- 2:06this row is a uh is trainable parameters
- 2:09that we're going to train using back
- 2:10propagation and this is the vector that
- 2:12then feeds into the Transformer um and
- 2:15that's how the Transformer Ser of
- 2:16perceives every single
- 2:18token so here we had a very naive
- 2:21tokenization process that was a
- 2:23character level tokenizer but in
- 2:25practice in state-ofthe-art uh language
- 2:27models people use a lot more complicated
- 2:28schemes unfortunately
- 2:30uh for constructing these uh token
- 2:34vocabularies so we're not dealing on the
- 2:36Character level we're dealing on chunk
- 2:38level and the way these um character
- 2:41chunks are constructed is using
- 2:43algorithms such as for example the bik
- 2:45pair in coding algorithm which we're
- 2:46going to go into in detail um and cover
- 2:51in this video I'd like to briefly show
- 2:52you the paper that introduced a bite
- 2:54level encoding as a mechanism for
- 2:56tokenization in the context of large
- 2:58language models and I would say that
- 3:00that's probably the gpt2 paper and if
- 3:02you scroll down here to the section
- 3:05input representation this is where they
- 3:07cover tokenization the kinds of
- 3:09properties that you'd like the
- 3:10tokenization to have and they conclude
- 3:13here that they're going to have a
- 3:14tokenizer where you have a vocabulary of
- 3:1750,2 57 possible
- 3:20tokens and the context size is going to
- 3:24be 1,24 tokens so in the in in the
- 3:27attention layer of the Transformer
- 3:29neural network
- 3:30every single token is attending to the
- 3:32previous tokens in the sequence and it's
- 3:34going to see up to 1,24 tokens so tokens
- 3:37are this like fundamental unit um the
- 3:40atom of uh large language models if you
- 3:43will and everything is in units of
- 3:44tokens everything is about tokens and
- 3:47tokenization is the process for
- 3:48translating strings or text into
- 3:51sequences of tokens and uh vice versa
- 3:54when you go into the Llama 2 paper as
- 3:56well I can show you that when you search
- 3:58token you're going to get get 63 hits um
- 4:01and that's because tokens are again
- 4:03pervasive so here they mentioned that
- 4:05they trained on two trillion tokens of
- 4:06data and so
- 4:08on so we're going to build our own
- 4:11tokenizer luckily the bite be encoding
- 4:13algorithm is not uh that super
- 4:15complicated and we can build it from
- 4:16scratch ourselves and we'll see exactly
- 4:18how this works before we dive into code
- 4:20I'd like to give you a brief Taste of
- 4:22some of the complexities that come from
- 4:24the tokenization because I just want to
- 4:26make sure that we motivate it
- 4:27sufficiently for why we are doing all
- 4:29this and why this is so gross so
- 4:32tokenization is at the heart of a lot of
- 4:34weirdness in large language models and I
- 4:36would advise that you do not brush it
- 4:37off a lot of the issues that may look
- 4:40like just issues with the new network
- 4:42architecture or the large language model
- 4:44itself are actually issues with the
- 4:46tokenization and fundamentally Trace uh
- 4:49back to it so if you've noticed any
- 4:51issues with large language models can't
- 4:54you know not able to do spelling tasks
- 4:56very easily that's usually due to
- 4:57tokenization simple string processing
- 5:00can be difficult for the large language
- 5:02model to perform
- 5:03natively uh non-english languages can
- 5:06work much worse and to a large extent
- 5:08this is due to
- 5:09tokenization sometimes llms are bad at
- 5:11simple arithmetic also can trace be
- 5:14traced to
- 5:15tokenization uh gbt2 specifically would
- 5:17have had quite a bit more issues with
- 5:19python than uh future versions of it due
- 5:22to tokenization there's a lot of other
- 5:24issues maybe you've seen weird warnings
- 5:25about a trailing whites space this is a
- 5:27tokenization issue um
- 5:30if you had asked GPT earlier about solid
- 5:33gold Magikarp and what it is you would
- 5:35see the llm go totally crazy and it
- 5:37would start going off about a completely
- 5:39unrelated tangent topic maybe you've
- 5:41been told to use yl over Json in
- 5:43structure data all of that has to do
- 5:45with tokenization so basically
- 5:47tokenization is at the heart of many
- 5:49issues I will look back around to these
- 5:51at the end of the video but for now let
- 5:54me just um skip over it a little bit and
- 5:56let's go to this web app um the Tik
- 5:59tokenizer bell.app so I have it loaded
- 6:02here and what I like about this web app
- 6:04is that tokenization is running a sort
- 6:06of live in your browser in JavaScript so
- 6:09you can just type here stuff hello world
- 6:11and the whole string
- 6:14rokenes so here what we see on uh the
- 6:18left is a string that you put in on the
- 6:20right we're currently using the gpt2
- 6:22tokenizer we see that this string that I
- 6:24pasted here is currently tokenizing into
- 6:27300 tokens and here they are sort of uh
- 6:30shown explicitly in different colors for
- 6:32every single token so for example uh
- 6:35this word tokenization became two tokens
- 6:38the token
- 6:403,642 and
- 6:441,634 the token um space is is token 318
- 6:50so be careful on the bottom you can show
- 6:51white space and keep in mind that there
- 6:54are spaces and uh sln new line
- 6:57characters in here but you can hide them
- 6:59for
- 7:01clarity the token space at is token 379
- 7:06the to the Token space the is 262 Etc so
- 7:11you notice here that the space is part
- 7:12of that uh token
- 7:15chunk now so this is kind of like how
- 7:18our English sentence broke up and that
- 7:21seems all well and good now now here I
- 7:24put in some arithmetic so we see that uh
- 7:26the token 127 Plus and then token six
- 7:31space 6 followed by 77 so what's
- 7:34happening here is that 127 is feeding in
- 7:36as a single token into the large
- 7:38language model but the um number 677
- 7:42will actually feed in as two separate
- 7:44tokens and so the large language model
- 7:47has to sort of um take account of that
- 7:50and process it correctly in its Network
- 7:53and see here 804 will be broken up into
- 7:56two tokens and it's is all completely
- 7:57arbitrary and here I have another
- 7:59example of four-digit numbers and they
- 8:02break up in a way that they break up and
- 8:03it's totally arbitrary sometimes you
- 8:05have um multiple digits single token
- 8:08sometimes you have individual digits as
- 8:10many tokens and it's all kind of pretty
- 8:12arbitrary and coming out of the
- 8:14tokenizer here's another example we have
- 8:17the string egg and you see here that
- 8:21this became two
- 8:22tokens but for some reason when I say I
- 8:24have an egg you see when it's a space
- 8:27egg it's two token it's sorry it's a
- 8:30single token so just egg by itself in
- 8:33the beginning of a sentence is two
- 8:34tokens but here as a space egg is
- 8:37suddenly a single token uh for the exact
- 8:40same string okay here lowercase egg
- 8:44turns out to be a single token and in
- 8:46particular notice that the color is
- 8:47different so this is a different token
- 8:49so this is case sensitive and of course
- 8:51a capital egg would also be different
- 8:54tokens and again um this would be two
- 8:57tokens arbitrarily so so for the same
- 9:00concept egg depending on if it's in the
- 9:02beginning of a sentence at the end of a
- 9:03sentence lowercase uppercase or mixed
- 9:06all this will be uh basically very
- 9:08different tokens and different IDs and
- 9:10the language model has to learn from raw
- 9:12data from all the internet text that
- 9:13it's going to be training on that these
- 9:15are actually all the exact same concept
- 9:17and it has to sort of group them in the
- 9:19parameters of the neural network and
- 9:21understand just based on the data
- 9:22patterns that these are all very similar
- 9:24but maybe not almost exactly similar but
- 9:27but very very similar
- 9:30um after the EG demonstration here I
- 9:32have um an introduction from open a eyes
- 9:35chbt in Korean so manaso Pang uh Etc uh
- 9:41so this is in Korean and the reason I
- 9:44put this here is because you'll notice
- 9:47that um non-english languages work
- 9:51slightly worse in Chachi part of this is
- 9:54because of course the training data set
- 9:55for Chachi is much larger for English
- 9:58and for everything else but the same is
- 9:59true not just for the large language
- 10:01model itself but also for the tokenizer
- 10:04so when we train the tokenizer we're
- 10:05going to see that there's a training set
- 10:07as well and there's a lot more English
- 10:09than non-english and what ends up
- 10:11happening is that we're going to have a
- 10:13lot more longer tokens for
- 10:16English so how do I put this if you have
- 10:19a single sentence in English and you
- 10:21tokenize it you might see that it's 10
- 10:23tokens or something like that but if you
- 10:25translate that sentence into say Korean
- 10:27or Japanese or something else you'll
- 10:29typically see that the number of tokens
- 10:30used is much larger and that's because
- 10:33the chunks here are a lot more broken up
- 10:36so we're using a lot more tokens for the
- 10:38exact same thing and what this does is
- 10:41it bloats up the sequence length of all
- 10:43the documents so you're using up more
- 10:46tokens and then in the attention of the
- 10:48Transformer when these tokens try to
- 10:49attend each other you are running out of
- 10:51context um in the maximum context length
- 10:55of that Transformer and so basically all
- 10:57the non-english text is stretched out
- 11:01from the perspective of the Transformer
- 11:03and this just has to do with the um
- 11:05trainings that used for the tokenizer
- 11:07and the tokenization itself so it will
- 11:10create a lot bigger tokens and a lot
- 11:12larger groups in English and it will
- 11:14have a lot of little boundaries for all
- 11:16the other non-english text um so if we
- 11:19translated this into English it would be
- 11:21significantly fewer
- 11:23tokens the final example I have here is
- 11:25a little snippet of python for doing FS
- 11:28buuz and what I'd like you to notice is
- 11:31look all these individual spaces are all
- 11:34separate tokens they are token
- 11:37220 so uh 220 220 220 220 and then space
- 11:42if is a single token and so what's going
- 11:45on here is that when the Transformer is
- 11:46going to consume or try to uh create
- 11:49this text it needs to um handle all
- 11:52these spaces individually they all feed
- 11:54in one by one into the entire
- 11:56Transformer in the sequence and so this
- 11:59is being extremely wasteful tokenizing
- 12:01it in this way and so as a result of
- 12:04that gpt2 is not very good with python
- 12:07and it's not anything to do with coding
- 12:08or the language model itself it's just
- 12:10that if he use a lot of indentation
- 12:12using space in Python like we usually do
- 12:15uh you just end up bloating out all the
- 12:17text and it's separated across way too
- 12:19much of the sequence and we are running
- 12:21out of the context length in the
- 12:22sequence uh that's roughly speaking
- 12:24what's what's happening we're being way
- 12:25too wasteful we're taking up way too
- 12:27much token space now we can also scroll
- 12:29up here and we can change the tokenizer
- 12:31so note here that gpt2 tokenizer creates
- 12:34a token count of 300 for this string
- 12:36here we can change it to CL 100K base
- 12:39which is the GPT for tokenizer and we
- 12:41see that the token count drops to 185 so
- 12:44for the exact same string we are now
- 12:46roughly having the number of tokens and
- 12:49roughly speaking this is because uh the
- 12:51number of tokens in the GPT 4 tokenizer
- 12:54is roughly double that of the number of
- 12:56tokens in the gpt2 tokenizer so we went
- 12:58went from roughly 50k to roughly 100K
- 13:01now you can imagine that this is a good
- 13:03thing because the same text is now
- 13:06squished into half as many tokens so uh
- 13:10this is a lot denser input to the
- 13:12Transformer and in the Transformer every
- 13:15single token has a finite number of
- 13:17tokens before it that it's going to pay
- 13:18attention to and so what this is doing
- 13:20is we're roughly able to see twice as
- 13:23much text as a context for what token to
- 13:26predict next uh because of this change
- 13:29but of course just increasing the number
- 13:30of tokens is uh not strictly better
- 13:33infinitely uh because as you increase
- 13:35the number of tokens now your embedding
- 13:36table is um sort of getting a lot larger
- 13:39and also at the output we are trying to
- 13:41predict the next token and there's the
- 13:42soft Max there and that grows as well
- 13:45we're going to go into more detail later
- 13:46on this but there's some kind of a Sweet
- 13:48Spot somewhere where you have a just
- 13:51right number of tokens in your
- 13:52vocabulary where everything is
- 13:53appropriately dense and still fairly
- 13:56efficient now one thing I would like you
- 13:58to note specifically for the gp4
- 14:00tokenizer is that the handling of the
- 14:03white space for python has improved a
- 14:05lot you see that here these four spaces
- 14:08are represented as one single token for
- 14:10the three spaces here and then the token
- 14:13SPF and here seven spaces were all
- 14:16grouped into a single token so we're
- 14:18being a lot more efficient in how we
- 14:20represent Python and this was a
- 14:21deliberate Choice made by open aai when
- 14:23they designed the gp4 tokenizer and they
- 14:27group a lot more space into a single
- 14:29character what this does is this
- 14:32densifies Python and therefore we can
- 14:35attend to more code before it when we're
- 14:38trying to predict the next token in the
- 14:39sequence and so the Improvement in the
- 14:42python coding ability from gbt2 to gp4
- 14:45is not just a matter of the language
- 14:47model and the architecture and the
- 14:48details of the optimization but a lot of
- 14:50the Improvement here is also coming from
- 14:52the design of the tokenizer and how it
- 14:54groups characters into tokens okay so
- 14:56let's now start writing some code
- 14:59so remember what we want to do we want
- 15:01to take strings and feed them into
- 15:03language models for that we need to
- 15:05somehow tokenize strings into some
- 15:08integers in some fixed vocabulary and
- 15:12then we will use those integers to make
- 15:14a look up into a lookup table of vectors
- 15:16and feed those vectors into the
- 15:18Transformer as an input now the reason
- 15:21this gets a little bit tricky of course
- 15:22is that we don't just want to support
- 15:24the simple English alphabet we want to
- 15:26support different kinds of languages so
- 15:28this is anango in Korean which is hello
- 15:31and we also want to support many kinds
- 15:33of special characters that we might find
- 15:34on the internet for example
- 15:37Emoji so how do we feed this text into
- 15:41uh
- 15:42Transformers well how's the what is this
- 15:44text anyway in Python so if you go to
- 15:46the documentation of a string in Python
- 15:49you can see that strings are immutable
- 15:51sequences of Unicode code
- 15:54points okay what are Unicode code points
- 15:57we can go to PDF so Unicode code points
- 16:01are defined by the Unicode Consortium as
- 16:04part of the Unicode standard and what
- 16:07this is really is that it's just a
- 16:09definition of roughly 150,000 characters
- 16:11right now and roughly speaking what they
- 16:14look like and what integers um represent
- 16:17those characters so it says 150,000
- 16:19characters across 161 scripts as of
- 16:22right now so if you scroll down here you
- 16:24can see that the standard is very much
- 16:26alive the latest standard 15.1 in
- 16:28September
- 16:302023 and basically this is just a way to
- 16:33define lots of types of
- 16:36characters like for example all these
- 16:39characters across different scripts so
- 16:41the way we can access the unic code code
- 16:44Point given Single Character is by using
- 16:45the or function in Python so for example
- 16:48I can pass in Ord of H and I can see
- 16:51that for the Single Character H the unic
- 16:54code code point is
- 16:56104 okay um but this can be arbitr
- 17:00complicated so we can take for example
- 17:02our Emoji here and we can see that the
- 17:04code point for this one is
- 17:06128,000 or we can take
- 17:10un and this is 50,000 now keep in mind
- 17:13you can't plug in strings here because
- 17:16you uh this doesn't have a single code
- 17:18point it only takes a single uni code
- 17:20code Point character and tells you its
- 17:23integer so in this way we can look
- 17:26up all the um characters of this
- 17:30specific string and their code points so
- 17:32or of X forx in this string and we get
- 17:36this encoding here now see here we've
- 17:40already turned the raw code points
- 17:42already have integers so why can't we
- 17:44simply just use these integers and not
- 17:46have any tokenization at all why can't
- 17:48we just use this natively as is and just
- 17:50use the code Point well one reason for
- 17:52that of course is that the vocabulary in
- 17:54that case would be quite long so in this
- 17:56case for Unicode the this is a
- 17:58vocabulary of
- 17:59150,000 different code points but more
- 18:02worryingly than that I think the Unicode
- 18:05standard is very much alive and it keeps
- 18:07changing and so it's not kind of a
- 18:09stable representation necessarily that
- 18:11we may want to use directly so for those
- 18:13reasons we need something a bit better
- 18:15so to find something better we turn to
- 18:17encodings so if we go to the Wikipedia
- 18:19page here we see that the Unicode
- 18:21consortion defines three types of
- 18:23encodings utf8 UTF 16 and UTF 32 these
- 18:27encoding are the way by which we can
- 18:30take Unicode text and translate it into
- 18:33binary data or by streams utf8 is by far
- 18:37the most common uh so this is the utf8
- 18:39page now this Wikipedia page is actually
- 18:42quite long but what's important for our
- 18:44purposes is that utf8 takes every single
- 18:46Cod point and it translates it to a by
- 18:49stream and this by stream is between one
- 18:52to four bytes so it's a variable length
- 18:54encoding so depending on the Unicode
- 18:56Point according to the schema you're
- 18:58going to end up with between 1 to four
- 18:59bytes for each code point on top of that
- 19:03there's utf8 uh
- 19:05utf16 and UTF 32 UTF 32 is nice because
- 19:08it is fixed length instead of variable
- 19:10length but it has many other downsides
- 19:12as well so the full kind of spectrum of
- 19:17pros and cons of all these different
- 19:18three encodings are beyond the scope of
- 19:20this video I just like to point out that
- 19:22I enjoyed this block post and this block
- 19:25post at the end of it also has a number
- 19:27of references that can be quite useful
- 19:29uh one of them is uh utf8 everywhere
- 19:32Manifesto um and this Manifesto
- 19:34describes the reason why utf8 is
- 19:36significantly preferred and a lot nicer
- 19:39than the other encodings and why it is
- 19:41used a lot more prominently um on the
- 19:45internet one of the major advantages
- 19:48just just to give you a sense is that
- 19:49utf8 is the only one of these that is
- 19:52backwards compatible to the much simpler
- 19:54asky encoding of text um but I'm not
- 19:57going to go into the full detail in this
- 19:58video so suffice to say that we like the
- 20:01utf8 encoding and uh let's try to take
- 20:03the string and see what we get if we
- 20:06encoded into
- 20:08utf8 the string class in Python actually
- 20:10has do encode and you can give it the
- 20:12encoding which is say utf8 now we get
- 20:15out of this is not very nice because
- 20:17this is the bytes is a bytes object and
- 20:20it's not very nice in the way that it's
- 20:22printed so I personally like to take it
- 20:25through list because then we actually
- 20:26get the raw B
- 20:28of this uh encoding so this is the raw
- 20:32byes that represent this string
- 20:35according to the utf8 en coding we can
- 20:38also look at utf16 we get a slightly
- 20:40different by stream and we here we start
- 20:43to see one of the disadvantages of utf16
- 20:45you see how we have zero Z something Z
- 20:47something Z something we're starting to
- 20:49get a sense that this is a bit of a
- 20:50wasteful encoding and indeed for simple
- 20:53asky characters or English characters
- 20:56here uh we just have the structure of 0
- 20:58something Z something and it's not
- 21:00exactly nice same for UTF 32 when we
- 21:04expand this we can start to get a sense
- 21:06of the wastefulness of this encoding for
- 21:08our purposes you see a lot of zeros
- 21:10followed by
- 21:11something and so uh this is not
- 21:14desirable so suffice it to say that we
- 21:17would like to stick with utf8 for our
- 21:20purposes however if we just use utf8
- 21:23naively these are by streams so that
- 21:26would imply a vocabulary length of only
- 21:29256 possible tokens uh but this this
- 21:33vocabulary size is very very small what
- 21:35this is going to do if we just were to
- 21:36use it naively is that all of our text
- 21:39would be stretched out over very very
- 21:41long sequences of bytes and so
- 21:46um what what this does is that certainly
- 21:49the embeding table is going to be tiny
- 21:51and the prediction at the top at the
- 21:52final layer is going to be very tiny but
- 21:54our sequences are very long and remember
- 21:56that we have pretty finite um context
- 21:59length and the attention that we can
- 22:01support in a transformer for
- 22:02computational reasons and so we only
- 22:05have as much context length but now we
- 22:07have very very long sequences and this
- 22:09is just inefficient and it's not going
- 22:10to allow us to attend to sufficiently
- 22:12long text uh before us for the purposes
- 22:15of the next token prediction task so we
- 22:18don't want to use the raw bytes of the
- 22:21utf8 encoding we want to be able to
- 22:24support larger vocabulary size that we
- 22:26can tune as a hyper
- 22:28but we want to stick with the utf8
- 22:30encoding of these strings so what do we
- 22:33do well the answer of course is we turn
- 22:35to the bite pair encoding algorithm
- 22:37which will allow us to compress these
- 22:39bite sequences um to a variable amount
- 22:42so we'll get to that in a bit but I just
- 22:44want to briefly speak to the fact that I
- 22:47would love nothing more than to be able
- 22:49to feed raw bite sequences into uh
- 22:52language models in fact there's a paper
- 22:54about how this could potentially be done
- 22:57uh from Summer last last year now the
- 22:59problem is you actually have to go in
- 23:00and you have to modify the Transformer
- 23:02architecture because as I mentioned
- 23:04you're going to have a problem where the
- 23:06attention will start to become extremely
- 23:08expensive because the sequences are so
- 23:10long and so in this paper they propose
- 23:13kind of a hierarchical structuring of
- 23:15the Transformer that could allow you to
- 23:17just feed in raw bites and so at the end
- 23:20they say together these results
- 23:21establish the viability of tokenization
- 23:23free autor regressive sequence modeling
- 23:25at scale so tokenization free would
- 23:27indeed be amazing we would just feed B
- 23:30streams directly into our models but
- 23:32unfortunately I don't know that this has
- 23:34really been proven out yet by
- 23:36sufficiently many groups and a
- 23:37sufficient scale uh but something like
- 23:39this at one point would be amazing and I
- 23:40hope someone comes up with it but for
- 23:42now we have to come back and we can't
- 23:44feed this directly into language models
- 23:46and we have to compress it using the B
- 23:48paare encoding algorithm so let's see
- 23:49how that works so as I mentioned the B
- 23:51paare encoding algorithm is not all that
- 23:53complicated and the Wikipedia page is
- 23:55actually quite instructive as far as the
- 23:57basic idea goes go what we're doing is
- 23:59we have some kind of a input sequence uh
- 24:01like for example here we have only four
- 24:03elements in our vocabulary a b c and d
- 24:06and we have a sequence of them so
- 24:08instead of bytes let's say we just have
- 24:09four a vocab size of
- 24:12four the sequence is too long and we'd
- 24:14like to compress it so what we do is
- 24:16that we iteratively find the pair of uh
- 24:20tokens that occur the most
- 24:23frequently and then once we've
- 24:25identified that pair we repl replace
- 24:28that pair with just a single new token
- 24:30that we append to our vocabulary so for
- 24:33example here the bite pair AA occurs
- 24:36most often so we mint a new token let's
- 24:38call it capital Z and we replace every
- 24:41single occurrence of AA by Z so now we
- 24:46have two Z's here so here we took a
- 24:48sequence of 11 characters with
- 24:51vocabulary size four and we've converted
- 24:54it to a um sequence of only nine tokens
- 24:58but now with a vocabulary of five
- 25:00because we have a fifth vocabulary
- 25:02element that we just created and it's Z
- 25:04standing for concatination of AA and we
- 25:07can again repeat this process so we
- 25:10again look at the sequence and identify
- 25:12the pair of tokens that are most
- 25:15frequent let's say that that is now AB
- 25:19well we are going to replace AB with a
- 25:20new token that we meant call Y so y
- 25:23becomes ab and then every single
- 25:25occurrence of ab is now replaced with y
- 25:28so we end up with this so now we only
- 25:31have 1 2 3 4 5 6 seven characters in our
- 25:35sequence but we have not just um four
- 25:40vocabulary elements or five but now we
- 25:42have six and for the final round we
- 25:45again look through the sequence find
- 25:47that the phrase zy or the pair zy is
- 25:50most common and replace it one more time
- 25:53with another um character let's say x so
- 25:56X is z y and we replace all curses of zy
- 25:59and we get this following sequence so
- 26:02basically after we have gone through
- 26:03this process instead of having a um
- 26:08sequence of
- 26:0911 uh tokens with a vocabulary length of
- 26:13four we now have a sequence of 1 2 3
- 26:18four five tokens but our vocabulary
- 26:21length now is seven and so in this way
- 26:25we can iteratively compress our sequence
- 26:27I we Mint new tokens so in the in the
- 26:30exact same way we start we start out
- 26:32with bite sequences so we have 256
- 26:36vocabulary size but we're now going to
- 26:38go through these and find the bite pairs
- 26:40that occur the most and we're going to
- 26:42iteratively start minting new tokens
- 26:44appending them to our vocabulary and
- 26:46replacing things and in this way we're
- 26:48going to end up with a compressed
- 26:50training data set and also an algorithm
- 26:52for taking any arbitrary sequence and
- 26:55encoding it using this uh vocabul
- 26:58and also decoding it back to Strings so
- 27:01let's now Implement all that so here's
- 27:03what I did I went to this block post
- 27:05that I enjoyed and I took the first
- 27:07paragraph and I copy pasted it here into
- 27:10text so this is one very long line
- 27:13here now to get the tokens as I
- 27:15mentioned we just take our text and we
- 27:17encode it into utf8 the tokens here at
- 27:20this point will be a raw bites single
- 27:22stream of bytes and just so that it's
- 27:25easier to work with instead of just a
- 27:27bytes object I'm going to convert all
- 27:29those bytes to integers and then create
- 27:32a list of it just so it's easier for us
- 27:34to manipulate and work with in Python
- 27:35and visualize and here I'm printing all
- 27:38of that so this is the original um this
- 27:42is the original paragraph and its length
- 27:45is
- 27:45533 uh code points and then here are the
- 27:49bytes encoded in ut utf8 and we see that
- 27:53this has a length of 616 bytes at this
- 27:56point or 616 tokens and the reason this
- 27:59is more is because a lot of these simple
- 28:01asky characters or simple characters
- 28:04they just become a single bite but a lot
- 28:06of these Unicode more complex characters
- 28:08become multiple bytes up to four and so
- 28:11we are expanding that
- 28:12size so now what we'd like to do as a
- 28:14first step of the algorithm is we'd like
- 28:16to iterate over here and find the pair
- 28:18of bites that occur most frequently
- 28:22because we're then going to merge it so
- 28:24if you are working long on a notebook on
- 28:25a side then I encourage you to basically
- 28:27click on the link find this notebook and
- 28:29try to write that function yourself
- 28:31otherwise I'm going to come here and
- 28:32Implement first the function that finds
- 28:34the most common pair okay so here's what
- 28:36I came up with there are many different
- 28:38ways to implement this but I'm calling
- 28:40the function get stats it expects a list
- 28:42of integers I'm using a dictionary to
- 28:44keep track of basically the counts and
- 28:46then this is a pythonic way to iterate
- 28:48consecutive elements of this list uh
- 28:51which we covered in the previous video
- 28:53and then here I'm just keeping track of
- 28:55just incrementing by one um for all the
- 28:58pairs so if I call this on all the
- 29:00tokens here then the stats comes out
- 29:03here so this is the dictionary the keys
- 29:06are these topples of consecutive
- 29:08elements and this is the count so just
- 29:11to uh print it in a slightly better way
- 29:14this is one way that I like to do that
- 29:17where you it's a little bit compound
- 29:20here so you can pause if you like but we
- 29:22iterate all all the items the items
- 29:25called on dictionary returns pairs of
- 29:27key value and instead I create a list
- 29:31here of value key because if it's a
- 29:35value key list then I can call sort on
- 29:37it and by default python will uh use the
- 29:41first element which in this case will be
- 29:43value to sort by if it's given tles and
- 29:46then reverse so it's descending and
- 29:48print that so basically it looks like
- 29:50101 comma 32 was the most commonly
- 29:53occurring consecutive pair and it
- 29:55occurred 20 times we can double check
- 29:58that that makes reasonable sense so if I
- 30:00just search
- 30:0210132 then you see that these are the 20
- 30:05occurrences of that um pair and if we'd
- 30:10like to take a look at what exactly that
- 30:11pair is we can use Char which is the
- 30:14opposite of or in Python so we give it a
- 30:17um unic code Cod point so 101 and of 32
- 30:22and we see that this is e and space so
- 30:25basically there's a lot of E space here
- 30:28meaning that a lot of these words seem
- 30:29to end with e so here's eace as an
- 30:32example so there's a lot of that going
- 30:34on here and this is the most common pair
- 30:36so now that we've identified the most
- 30:38common pair we would like to iterate
- 30:40over this sequence we're going to Mint a
- 30:42new token with the ID of
- 30:44256 right because these tokens currently
- 30:47go from Z to 255 so when we create a new
- 30:50token it will have an ID of
- 30:52256 and we're going to iterate over this
- 30:56entire um list and every every time we
- 30:59see 101 comma 32 we're going to swap
- 31:02that out for
- 31:03256 so let's Implement that now and feel
- 31:07free to uh do that yourself as well so
- 31:09first I commented uh this just so we
- 31:11don't pollute uh the notebook too much
- 31:14this is a nice way of in Python
- 31:17obtaining the highest ranking pair so
- 31:20we're basically calling the Max on this
- 31:23dictionary stats and this will return
- 31:26the maximum
- 31:27key and then the question is how does it
- 31:30rank keys so you can provide it with a
- 31:32function that ranks keys and that
- 31:35function is just stats. getet uh stats.
- 31:38getet would basically return the value
- 31:41and so we're ranking by the value and
- 31:42getting the maximum key so it's 101
- 31:45comma 32 as we saw now to actually merge
- 31:4910132 um this is the function that I
- 31:51wrote but again there are many different
- 31:53versions of it so we're going to take a
- 31:55list of IDs and the the pair that we
- 31:57want to replace and that pair will be
- 31:59replaced with the new index
- 32:02idx so iterating through IDs if we find
- 32:05the pair swap it out for idx so we
- 32:08create this new list and then we start
- 32:10at zero and then we go through this
- 32:12entire list sequentially from left to
- 32:14right and here we are checking for
- 32:17equality at the current position with
- 32:19the
- 32:20pair um so here we are checking that the
- 32:23pair matches now here is a bit of a
- 32:25tricky condition that you have to append
- 32:27if you're trying to be careful and that
- 32:29is that um you don't want this here to
- 32:31be out of Bounds at the very last
- 32:33position when you're on the rightmost
- 32:35element of this list otherwise this
- 32:37would uh give you an autof bounds error
- 32:39so we have to make sure that we're not
- 32:40at the very very last element so uh this
- 32:44would be false for that so if we find a
- 32:46match we append to this new list that
- 32:51replacement index and we increment the
- 32:53position by two so we skip over that
- 32:54entire pair but otherwise if we we
- 32:57haven't found a matching pair we just
- 32:59sort of copy over the um element at that
- 33:02position and increment by one then
- 33:05return this so here's a very small toy
- 33:07example if we have a list 566 791 and we
- 33:10want to replace the occurrences of 67
- 33:12with 99 then calling this on that will
- 33:16give us what we're asking for so here
- 33:18the 67 is replaced with
- 33:2199 so now I'm going to uncomment this
- 33:23for our actual use case where we want to
- 33:27take our tokens we want to take the top
- 33:29pair here and replace it with 256 to get
- 33:33tokens to if we run this we get the
- 33:37following so recall that previously we
- 33:40had a length 616 in this list and now we
- 33:45have a length 596 right so this
- 33:48decreased by 20 which makes sense
- 33:50because there are 20 occurrences
- 33:52moreover we can try to find 256 here and
- 33:55we see plenty of occurrences on off it
- 33:58and moreover just double check there
- 33:59should be no occurrence of 10132 so this
- 34:02is the original array plenty of them and
- 34:05in the second array there are no
- 34:06occurrences of 1032 so we've
- 34:08successfully merged this single pair and
- 34:11now we just uh iterate this so we are
- 34:13going to go over the sequence again find
- 34:15the most common pair and replace it so
- 34:17let me now write a y Loop that uses
- 34:19these functions to do this um sort of
- 34:21iteratively and how many times do we do
- 34:24it four well that's totally up to us as
- 34:26a hyper parameter
- 34:27the more um steps we take the larger
- 34:30will be our vocabulary and the shorter
- 34:33will be our sequence and there is some
- 34:35sweet spot that we usually find works
- 34:37the best in practice and so this is kind
- 34:39of a hyperparameter and we tune it and
- 34:41we find good vocabulary sizes as an
- 34:44example gp4 currently uses roughly
- 34:46100,000 tokens and um bpark that those
- 34:49are reasonable numbers currently instead
- 34:51the are large language models so let me
- 34:53now write uh putting putting it all
- 34:55together and uh iterating these steps
- 34:58okay now before we dive into the Y loop
- 35:00I wanted to add one more cell here where
- 35:03I went to the block post and instead of
- 35:04grabbing just the first paragraph or two
- 35:07I took the entire block post and I
- 35:08stretched it out in a single line and
- 35:10basically just using longer text will
- 35:12allow us to have more representative
- 35:13statistics for the bite Pairs and we'll
- 35:16just get a more sensible results out of
- 35:18it because it's longer text um so here
- 35:21we have the raw text we encode it into
- 35:24bytes using the utf8 encoding
- 35:27and then here as before we are just
- 35:30changing it into a list of integers in
- 35:31Python just so it's easier to work with
- 35:33instead of the raw byes objects and then
- 35:36this is the code that I came up with uh
- 35:40to actually do the merging in Loop these
- 35:44two functions here are identical to what
- 35:45we had above I only included them here
- 35:48just so that you have the point of
- 35:49reference here so uh these two are
- 35:53identical and then this is the new code
- 35:55that I added so the first first thing we
- 35:57want to do is we want to decide on the
- 35:58final vocabulary size that we want our
- 36:01tokenizer to have and as I mentioned
- 36:02this is a hyper parameter and you set it
- 36:04in some way depending on your best
- 36:06performance so let's say for us we're
- 36:08going to use 276 because that way we're
- 36:10going to be doing exactly 20
- 36:13merges and uh 20 merges because we
- 36:15already have
- 36:16256 tokens for the raw bytes and to
- 36:20reach 276 we have to do 20 merges uh to
- 36:23add 20 new
- 36:25tokens here uh this is uh one way in
- 36:28Python to just create a copy of a list
- 36:31so I'm taking the tokens list and by
- 36:33wrapping it in a list python will
- 36:35construct a new list of all the
- 36:37individual elements so this is just a
- 36:38copy
- 36:39operation then here I'm creating a
- 36:42merges uh dictionary so this merges
- 36:44dictionary is going to maintain
- 36:46basically the child one child two
- 36:49mapping to a new uh token and so what
- 36:52we're going to be building up here is a
- 36:53binary tree of merges but actually it's
- 36:56not exactly a tree because a tree would
- 36:59have a single root node with a bunch of
- 37:01leaves for us we're starting with the
- 37:03leaves on the bottom which are the
- 37:05individual bites those are the starting
- 37:06256 tokens and then we're starting to
- 37:09like merge two of them at a time and so
- 37:11it's not a tree it's more like a forest
- 37:14um uh as we merge these elements
- 37:18so for 20 merges we're going to find the
- 37:22most commonly occurring pair we're going
- 37:25to Mint a new token integer for it so I
- 37:28here will start at zero so we'll going
- 37:30to start at 256 we're going to print
- 37:32that we're merging it and we're going to
- 37:34replace all of the occurrences of that
- 37:36pair with the new new lied token and
- 37:39we're going to record that this pair of
- 37:42integers merged into this new
- 37:45integer so running this gives us the
- 37:49following
- 37:51output so we did 20 merges and for
- 37:54example the first merge was exactly as
- 37:56before the
- 37:5810132 um tokens merging into a new token
- 38:012556 now keep in mind that the
- 38:04individual uh tokens 101 and 32 can
- 38:06still occur in the sequence after
- 38:08merging it's only when they occur
- 38:10exactly consecutively that that becomes
- 38:12256
- 38:13now um and in particular the other thing
- 38:16to notice here is that the token 256
- 38:19which is the newly minted token is also
- 38:21eligible for merging so here on the
- 38:23bottom the 20th merge was a merge of 25
- 38:26and 259 becoming
- 38:28275 so every time we replace these
- 38:31tokens they become eligible for merging
- 38:33in the next round of data ration so
- 38:35that's why we're building up a small
- 38:37sort of binary Forest instead of a
- 38:38single individual
- 38:40tree one thing we can take a look at as
- 38:42well is we can take a look at the
- 38:44compression ratio that we've achieved so
- 38:46in particular we started off with this
- 38:48tokens list um so we started off with
- 38:5124,000 bytes and after merging 20 times
- 38:56uh we now have only
- 38:5819,000 um tokens and so therefore the
- 39:01compression ratio simply just dividing
- 39:03the two is roughly 1.27 so that's the
- 39:06amount of compression we were able to
- 39:07achieve of this text with only 20
- 39:10merges um and of course the more
- 39:13vocabulary elements you add uh the
- 39:15greater the compression ratio here would
- 39:19be finally so that's kind of like um the
- 39:23training of the tokenizer if you will
- 39:25now 1 Point I wanted to make is that and
- 39:28maybe this is a diagram that can help um
- 39:31kind of illustrate is that tokenizer is
- 39:33a completely separate object from the
- 39:34large language model itself so
- 39:37everything in this lecture we're not
- 39:38really touching the llm itself uh we're
- 39:40just training the tokenizer this is a
- 39:41completely separate pre-processing stage
- 39:43usually so the tokenizer will have its
- 39:46own training set just like a large
- 39:47language model has a potentially
- 39:49different training set so the tokenizer
- 39:52has a training set of documents on which
- 39:53you're going to train the
- 39:54tokenizer and then and um we're
- 39:57performing The Bite pair encoding
- 39:58algorithm as we saw above to train the
- 40:01vocabulary of this
- 40:02tokenizer so it has its own training set
- 40:04it is a pre-processing stage that you
- 40:06would run a single time in the beginning
- 40:09um and the tokenizer is trained using
- 40:11bipar coding algorithm once you have the
- 40:14tokenizer once it's trained and you have
- 40:16the vocabulary and you have the merges
- 40:19uh we can do both encoding and decoding
- 40:22so these two arrows here so the
- 40:24tokenizer is a translation layer between
- 40:27raw text which is as we saw the sequence
- 40:30of Unicode code points it can take raw
- 40:32text and turn it into a token sequence
- 40:35and vice versa it can take a token
- 40:37sequence and translate it back into raw
- 40:40text so now that we have trained uh
- 40:43tokenizer and we have these merges we
- 40:45are going to turn to how we can do the
- 40:47encoding and the decoding step if you
- 40:49give me text here are the tokens and
- 40:51vice versa if you give me tokens here's
- 40:53the text once we have that we can
- 40:55translate between these two Realms and
- 40:57then the language model is going to be
- 40:58trained as a step two afterwards and
- 41:01typically in a in a sort of a
- 41:03state-of-the-art application you might
- 41:05take all of your training data for the
- 41:06language model and you might run it
- 41:08through the tokenizer and sort of
- 41:10translate everything into a massive
- 41:11token sequence and then you can throw
- 41:13away the raw text you're just left with
- 41:15the tokens themselves and those are
- 41:17stored on disk and that is what the
- 41:19large language model is actually reading
- 41:21when it's training on them so this one
- 41:23approach that you can take as a single
- 41:24massive pre-processing step a
- 41:26stage um so yeah basically I think the
- 41:30most important thing I want to get
- 41:31across is that this is completely
- 41:32separate stage it usually has its own
- 41:34entire uh training set you may want to
- 41:36have those training sets be different
- 41:38between the tokenizer and the logge
- 41:39language model so for example when
- 41:41you're training the tokenizer as I
- 41:43mentioned we don't just care about the
- 41:45performance of English text we care
- 41:46about uh multi many different languages
- 41:49and we also care about code or not code
- 41:51so you may want to look into different
- 41:53kinds of mixtures of different kinds of
- 41:55languages and different amounts of code
- 41:57and things like that because the amount
- 42:00of different language that you have in
- 42:01your tokenizer training set will
- 42:03determine how many merges of it there
- 42:06will be and therefore that determines
- 42:08the density with which uh this type of
- 42:11data is um sort of has in the token
- 42:15space and so roughly speaking
- 42:17intuitively if you add some amount of
- 42:19data like say you have a ton of Japanese
- 42:21data in your uh tokenizer training set
- 42:24then that means that more Japanese
- 42:25tokens will get merged
- 42:26and therefore Japanese will have shorter
- 42:28sequences uh and that's going to be
- 42:30beneficial for the large language model
- 42:32which has a finite context length on
- 42:34which it can work on in in the token
- 42:36space uh so hopefully that makes sense
- 42:39so we're now going to turn to encoding
- 42:41and decoding now that we have trained a
- 42:43tokenizer so we have our merges and now
- 42:46how do we do encoding and decoding okay
- 42:48so let's begin with decoding which is
- 42:50this Arrow over here so given a token
- 42:52sequence let's go through the tokenizer
- 42:54to get back a python string object so
- 42:57the raw text so this is the function
- 42:59that we' like to implement um we're
- 43:01given the list of integers and we want
- 43:03to return a python string if you'd like
- 43:05uh try to implement this function
- 43:06yourself it's a fun exercise otherwise
- 43:08I'm going to start uh pasting in my own
- 43:11solution so there are many different
- 43:13ways to do it um here's one way I will
- 43:16create an uh kind of pre-processing
- 43:18variable that I will call
- 43:21vocab and vocab is a mapping or a
- 43:24dictionary in Python for from the token
- 43:27uh ID to the bytes object for that token
- 43:31so we begin with the raw bytes for
- 43:33tokens from 0 to 255 and then we go in
- 43:36order of all the merges and we sort of
- 43:39uh populate this vocab list by doing an
- 43:42addition here so this is the basically
- 43:45the bytes representation of the first
- 43:47child followed by the second one and
- 43:50remember these are bytes objects so this
- 43:52addition here is an addition of two
- 43:54bytes objects just concatenation
- 43:57so that's what we get
- 43:58here one tricky thing to be careful with
- 44:01by the way is that I'm iterating a
- 44:02dictionary in Python using a DOT items
- 44:06and uh it really matters that this runs
- 44:08in the order in which we inserted items
- 44:11into the merous dictionary luckily
- 44:13starting with python 3.7 this is
- 44:15guaranteed to be the case but before
- 44:17python 3.7 this iteration may have been
- 44:19out of order with respect to how we
- 44:20inserted elements into merges and this
- 44:23may not have worked but we are using an
- 44:25um modern python so we're okay and then
- 44:28here uh given the IDS the first thing
- 44:31we're going to do is get the
- 44:35tokens so the way I implemented this
- 44:37here is I'm taking I'm iterating over
- 44:39all the IDS I'm using vocap to look up
- 44:41their bytes and then here this is one
- 44:44way in Python to concatenate all these
- 44:46bytes together to create our tokens and
- 44:49then these tokens here at this point are
- 44:51raw bytes so I have to decode using UTF
- 44:56F now back into python strings so
- 44:59previously we called that encode on a
- 45:01string object to get the bytes and now
- 45:03we're doing it Opposite we're taking the
- 45:05bytes and calling a decode on the bytes
- 45:07object to get a string in Python and
- 45:11then we can return
- 45:13text so um this is how we can do it now
- 45:16this actually has a um issue um in the
- 45:20way I implemented it and this could
- 45:22actually throw an error so try to think
- 45:24figure out why this code could actually
- 45:26result in an error if we plug in um uh
- 45:30some sequence of IDs that is
- 45:32unlucky so let me demonstrate the issue
- 45:35when I try to decode just something like
- 45:3797 I am going to get letter A here back
- 45:41so nothing too crazy happening but when
- 45:44I try to decode 128 as a single element
- 45:48the token 128 is what in string or in
- 45:51Python object uni Cod decoder utfa can't
- 45:55Decode by um 0x8 which is this in HEX in
- 46:00position zero invalid start bite what
- 46:01does that mean well to understand what
- 46:03this means we have to go back to our
- 46:04utf8 page uh that I briefly showed
- 46:07earlier and this is Wikipedia utf8 and
- 46:10basically there's a specific schema that
- 46:13utfa bytes take so in particular if you
- 46:16have a multi-te object for some of the
- 46:19Unicode characters they have to have
- 46:21this special sort of envelope in how the
- 46:24encoding works and so what's happening
- 46:26here is that invalid start pite that's
- 46:30because
- 46:31128 the binary representation of it is
- 46:33one followed by all zeros so we have one
- 46:37and then all zero and we see here that
- 46:39that doesn't conform to the format
- 46:41because one followed by all zero just
- 46:42doesn't fit any of these rules so to
- 46:44speak so it's an invalid start bite
- 46:47which is byte one this one must have a
- 46:50one following it and then a zero
- 46:52following it and then the content of
- 46:54your uni codee in x here so basically we
- 46:57don't um exactly follow the utf8
- 46:59standard and this cannot be decoded and
- 47:02so the way to fix this um is to
- 47:06use this errors equals in bytes. decode
- 47:11function of python and by default errors
- 47:13is strict so we will throw an error if
- 47:17um it's not valid utf8 bytes encoding
- 47:20but there are many different things that
- 47:21you could put here on error handling
- 47:23this is the full list of all the errors
- 47:25that you can use and in particular
- 47:27instead of strict let's change it to
- 47:29replace and that will replace uh with
- 47:32this special marker this replacement
- 47:35character so errors equals replace and
- 47:40now we just get that character
- 47:43back so basically not every single by
- 47:46sequence is valid
- 47:48utf8 and if it happens that your large
- 47:51language model for example predicts your
- 47:53tokens in a bad manner then they might
- 47:56not fall into valid utf8 and then we
- 48:00won't be able to decode them so the
- 48:02standard practice is to basically uh use
- 48:05errors equals replace and this is what
- 48:07you will also find in the openai um code
- 48:10that they released as well but basically
- 48:12whenever you see um this kind of a
- 48:14character in your output in that case uh
- 48:16something went wrong and the LM output
- 48:18not was not valid uh sort of sequence of
- 48:21tokens okay and now we're going to go
- 48:23the other way so we are going to
- 48:25implement
- 48:26this Arrow right here where we are going
- 48:27to be given a string and we want to
- 48:29encode it into
- 48:31tokens so this is the signature of the
- 48:33function that we're interested in and um
- 48:36this should basically print a list of
- 48:38integers of the tokens so again uh try
- 48:41to maybe implement this yourself if
- 48:43you'd like a fun exercise uh and pause
- 48:45here otherwise I'm going to start
- 48:46putting in my
- 48:47solution so again there are many ways to
- 48:50do this so um this is one of the ways
- 48:53that sort of I came came up with so the
- 48:57first thing we're going to do is we are
- 48:59going
- 49:00to uh take our text encode it into utf8
- 49:03to get the raw bytes and then as before
- 49:05we're going to call list on the bytes
- 49:07object to get a list of integers of
- 49:10those bytes so those are the starting
- 49:12tokens those are the raw bytes of our
- 49:14sequence but now of course according to
- 49:16the merges dictionary above and recall
- 49:19this was the
- 49:21merges some of the bytes may be merged
- 49:23according to this lookup in addition to
- 49:26that remember that the merges was built
- 49:28from top to bottom and this is sort of
- 49:29the order in which we inserted stuff
- 49:31into merges and so we prefer to do all
- 49:34these merges in the beginning before we
- 49:36do these merges later because um for
- 49:39example this merge over here relies on
- 49:40the 256 which got merged here so we have
- 49:44to go in the order from top to bottom
- 49:46sort of if we are going to be merging
- 49:48anything now we expect to be doing a few
- 49:51merges so we're going to be doing W
- 49:54true um and now we want to find a pair
- 49:58of byes that is consecutive that we are
- 50:00allowed to merge according to this in
- 50:03order to reuse some of the functionality
- 50:05that we've already written I'm going to
- 50:06reuse the function uh get
- 50:09stats so recall that get stats uh will
- 50:12give us the we'll basically count up how
- 50:14many times every single pair occurs in
- 50:16our sequence of tokens and return that
- 50:18as a dictionary and the dictionary was a
- 50:22mapping from all the different uh by
- 50:25pairs to the number of times that they
- 50:27occur right um at this point we don't
- 50:30actually care how many times they occur
- 50:32in the sequence we only care what the
- 50:34raw pairs are in that sequence and so
- 50:36I'm only going to be using basically the
- 50:38keys of the dictionary I only care about
- 50:40the set of possible merge candidates if
- 50:42that makes
- 50:43sense now we want to identify the pair
- 50:46that we're going to be merging at this
- 50:47stage of the loop so what do we want we
- 50:50want to find the pair or like the a key
- 50:53inside stats that has the lowest index
- 50:57in the merges uh dictionary because we
- 50:59want to do all the early merges before
- 51:01we work our way to the late
- 51:03merges so again there are many different
- 51:05ways to implement this but I'm going to
- 51:07do something a little bit fancy
- 51:11here so I'm going to be using the Min
- 51:14over an iterator in Python when you call
- 51:16Min on an iterator and stats here as a
- 51:18dictionary we're going to be iterating
- 51:20the keys of this dictionary in Python so
- 51:24we're looking at all the pairs inside
- 51:27stats um which are all the consecutive
- 51:29Pairs and we're going to be taking the
- 51:32consecutive pair inside tokens that has
- 51:34the minimum what the Min takes a key
- 51:38which gives us the function that is
- 51:40going to return a value over which we're
- 51:42going to do the Min and the one we care
- 51:44about is we're we care about taking
- 51:46merges and basically getting um that
- 51:50pairs
- 51:52index so basically for any pair inside
- 51:57stats we are going to be looking into
- 51:59merges at what index it has and we want
- 52:03to get the pair with the Min number so
- 52:05as an example if there's a pair 101 and
- 52:0732 we definitely want to get that pair
- 52:10uh we want to identify it here and
- 52:11return it and pair would become 10132 if
- 52:15it
- 52:15occurs and the reason that I'm putting a
- 52:17float INF here as a fall back is that in
- 52:21the get function when we call uh when we
- 52:24basically consider a pair that doesn't
- 52:26occur in the merges then that pair is
- 52:29not eligible to be merged right so if in
- 52:31the token sequence there's some pair
- 52:33that is not a merging pair it cannot be
- 52:35merged then uh it doesn't actually occur
- 52:38here and it doesn't have an index and uh
- 52:40it cannot be merged which we will denote
- 52:42as float INF and the reason Infinity is
- 52:45nice here is because for sure we're
- 52:46guaranteed that it's not going to
- 52:48participate in the list of candidates
- 52:50when we do the men so uh so this is one
- 52:53way to do it so B basically long story
- 52:55short this Returns the most eligible
- 52:58merging candidate pair uh that occurs in
- 53:01the tokens now one thing to be careful
- 53:04with here is this uh function here might
- 53:07fail in the following way if there's
- 53:09nothing to merge then uh uh then there's
- 53:13nothing in merges um that satisfi that
- 53:16is satisfied anymore there's nothing to
- 53:18merge everything just returns float imps
- 53:21and then the pair I think will just
- 53:23become the very first element of stats
- 53:26um but this pair is not actually a
- 53:28mergeable pair it just becomes the first
- 53:31pair inside stats arbitrarily because
- 53:33all of these pairs evaluate to float in
- 53:36for the merging Criterion so basically
- 53:38it could be that this this doesn't look
- 53:40succeed because there's no more merging
- 53:41pairs so if this pair is not in merges
- 53:44that was returned then this is a signal
- 53:46for us that actually there was nothing
- 53:48to merge no single pair can be merged
- 53:50anymore in that case we will break
- 53:53out um nothing else can be
- 53:57merged you may come up with a different
- 53:59implementation by the way this is kind
- 54:01of like really trying hard in
- 54:03Python um but really we're just trying
- 54:05to find a pair that can be merged with
- 54:07the lowest index
- 54:09here now if we did find a pair that is
- 54:13inside merges with the lowest index then
- 54:16we can merge it
- 54:19so we're going to look into the merger
- 54:22dictionary for that pair to look up the
- 54:24index and we're going to now merge that
- 54:27into that index so we're going to do
- 54:29tokens equals and we're going to
- 54:32replace the original tokens we're going
- 54:34to be replacing the pair pair and we're
- 54:36going to be replacing it with index idx
- 54:38and this returns a new list of tokens
- 54:41where every occurrence of pair is
- 54:43replaced with idx so we're doing a merge
- 54:46and we're going to be continuing this
- 54:47until eventually nothing can be merged
- 54:49we'll come out here and we'll break out
- 54:51and here we just return
- 54:53tokens and so that that's the
- 54:55implementation I think so hopefully this
- 54:57runs okay cool um yeah and this looks uh
- 55:02reasonable so for example 32 is a space
- 55:04in asky so that's here um so this looks
- 55:09like it worked great okay so let's wrap
- 55:11up this section of the video at least I
- 55:13wanted to point out that this is not
- 55:14quite the right implementation just yet
- 55:16because we are leaving out a special
- 55:17case so in particular if uh we try to do
- 55:20this this would give us an error and the
- 55:23issue is that um if we only have a
- 55:25single character or an empty string then
- 55:28stats is empty and that causes an issue
- 55:29inside Min so one way to fight this is
- 55:32if L of tokens is at least two because
- 55:36if it's less than two it's just a single
- 55:37token or no tokens then let's just uh
- 55:40there's nothing to merge so we just
- 55:41return so that would fix uh that
- 55:44case Okay and then second I have a few
- 55:48test cases here for us as well so first
- 55:50let's make sure uh about or let's note
- 55:53the following if we take a string and we
- 55:56try to encode it and then decode it back
- 55:58you'd expect to get the same string back
- 56:00right is that true for all
- 56:04strings so I think uh so here it is the
- 56:07case and I think in general this is
- 56:08probably the case um but notice that
- 56:12going backwards is not is not you're not
- 56:14going to have an identity going
- 56:15backwards because as I mentioned us not
- 56:19all token sequences are valid utf8 uh
- 56:22sort of by streams and so so therefore
- 56:25you're some of them can't even be
- 56:27decodable um so this only goes in One
- 56:30Direction but for that one direction we
- 56:32can check uh here if we take the
- 56:34training text which is the text that we
- 56:36train to tokenizer around we can make
- 56:38sure that when we encode and decode we
- 56:39get the same thing back which is true
- 56:41and here I took some validation data so
- 56:43I went to I think this web page and I
- 56:45grabbed some text so this is text that
- 56:47the tokenizer has not seen and we can
- 56:49make sure that this also works um okay
- 56:52so that gives us some confidence that
- 56:53this was correctly implemented
- 56:56so those are the basics of the bite pair
- 56:58encoding algorithm we saw how we can uh
- 57:00take some training set train a tokenizer
- 57:03the parameters of this tokenizer really
- 57:05are just this dictionary of merges and
- 57:08that basically creates the little binary
- 57:09Forest on top of raw
- 57:11bites once we have this the merges table
- 57:14we can both encode and decode between
- 57:16raw text and token sequences so that's
- 57:19the the simplest setting of The
- 57:21tokenizer what we're going to do now
- 57:23though is we're going to look at some of
- 57:24the St the art lar language models and
- 57:26the kinds of tokenizers that they use
- 57:28and we're going to see that this picture
- 57:29complexifies very quickly so we're going
- 57:31to go through the details of this comp
- 57:34complexification one at a time so let's
- 57:37kick things off by looking at the GPD
- 57:39Series so in particular I have the gpt2
- 57:41paper here um and this paper is from
- 57:442019 or so so 5 years ago and let's
- 57:48scroll down to input representation this
- 57:51is where they talk about the tokenizer
- 57:52that they're using for gpd2 now this is
- 57:55all fairly readable so I encourage you
- 57:57to pause and um read this yourself but
- 58:00this is where they motivate the use of
- 58:02the bite pair encoding algorithm on the
- 58:04bite level representation of utf8
- 58:07encoding so this is where they motivate
- 58:09it and they talk about the vocabulary
- 58:11sizes and everything now everything here
- 58:13is exactly as we've covered it so far
- 58:15but things start to depart around here
- 58:18so what they mention is that they don't
- 58:20just apply the naive algorithm as we
- 58:22have done it and in particular here's a
- 58:25example suppose that you have common
- 58:27words like dog what will happen is that
- 58:29dog of course occurs very frequently in
- 58:31the text and it occurs right next to all
- 58:34kinds of punctuation as an example so
- 58:36doc dot dog exclamation mark dog
- 58:39question mark Etc and naively you might
- 58:42imagine that the BP algorithm could
- 58:43merge these to be single tokens and then
- 58:45you end up with lots of tokens that are
- 58:47just like dog with a slightly different
- 58:49punctuation and so it feels like you're
- 58:50clustering things that shouldn't be
- 58:52clustered you're combining kind of
- 58:53semantics with
- 58:55uation and this uh feels suboptimal and
- 58:58indeed they also say that this is
- 59:00suboptimal according to some of the
- 59:02experiments so what they want to do is
- 59:04they want to top down in a manual way
- 59:06enforce that some types of um characters
- 59:09should never be merged together um so
- 59:12they want to enforce these merging rules
- 59:14on top of the bite PA encoding algorithm
- 59:17so let's take a look um at their code
- 59:19and see how they actually enforce this
- 59:21and what kinds of mergy they actually do
- 59:23perform so I have to to tab open here
- 59:25for gpt2 under open AI on GitHub and
- 59:29when we go to
- 59:30Source there is an encoder thatp now I
- 59:34don't personally love that they call it
- 59:35encoder dopy because this is the
- 59:37tokenizer and the tokenizer can do both
- 59:39encode and decode uh so it feels kind of
- 59:41awkward to me that it's called encoder
- 59:43but that is the tokenizer and there's a
- 59:45lot going on here and we're going to
- 59:47step through it in detail at one point
- 59:49for now I just want to focus on this
- 59:51part here the create a rigix pattern
- 59:54here that looks very complicated and
- 59:56we're going to go through it in a bit uh
- 59:58but this is the core part that allows
- 1:00:00them to enforce rules uh for what parts
- 1:00:04of the text Will Never Be merged for
- 1:00:05sure now notice that re. compile here is
- 1:00:08a little bit misleading because we're
- 1:00:10not just doing import re which is the
- 1:00:12python re module we're doing import reex
- 1:00:14as re and reex is a python package that
- 1:00:17you can install P install r x and it's
- 1:00:20basically an extension of re so it's a
- 1:00:22bit more powerful
- 1:00:23re um
- 1:00:26so let's take a look at this pattern and
- 1:00:28what it's doing and why this is actually
- 1:00:30doing the separation that they are
- 1:00:32looking for okay so I've copy pasted the
- 1:00:34pattern here to our jupit notebook where
- 1:00:37we left off and let's take this pattern
- 1:00:39for a spin so in the exact same way that
- 1:00:42their code does we're going to call an
- 1:00:44re. findall for this pattern on any
- 1:00:47arbitrary string that we are interested
- 1:00:49so this is the string that we want to
- 1:00:50encode into tokens um to feed into n llm
- 1:00:55like gpt2 so what exactly is this doing
- 1:00:59well re. findall will take this pattern
- 1:01:01and try to match it against a
- 1:01:02string um the way this works is that you
- 1:01:06are going from left to right in the
- 1:01:07string and you're trying to match the
- 1:01:10pattern and R.F find all will get all
- 1:01:13the occurrences and organize them into a
- 1:01:16list now when you look at the um when
- 1:01:19you look at this pattern first of all
- 1:01:20notice that this is a raw string um and
- 1:01:23then these are three double quotes just
- 1:01:26to start the string so really the string
- 1:01:28itself this is the pattern itself
- 1:01:31right and notice that it's made up of a
- 1:01:34lot of ores so see these vertical bars
- 1:01:36those are ores in reg X and so you go
- 1:01:40from left to right in this pattern and
- 1:01:41try to match it against the string
- 1:01:43wherever you are so we have hello and
- 1:01:46we're going to try to match it well it's
- 1:01:48not apostrophe s it's not apostrophe t
- 1:01:50or any of these but it is an optional
- 1:01:53space followed by- P of uh sorry SL P of
- 1:01:58L one or more times what is/ P of L it
- 1:02:02is coming to some documentation that I
- 1:02:04found um there might be other sources as
- 1:02:08well uh SLP is a letter any kind of
- 1:02:11letter from any language and hello is
- 1:02:15made up of letters h e l Etc so optional
- 1:02:19space followed by a bunch of letters one
- 1:02:21or more letters is going to match hello
- 1:02:24but then the match ends because a white
- 1:02:27space is not a letter so from there on
- 1:02:31begins a new sort of attempt to match
- 1:02:33against the string again and starting in
- 1:02:36here we're going to skip over all of
- 1:02:38these again until we get to the exact
- 1:02:40same Point again and we see that there's
- 1:02:42an optional space this is the optional
- 1:02:44space followed by a bunch of letters one
- 1:02:46or more of them and so that matches so
- 1:02:48when we run this we get a list of two
- 1:02:52elements hello and then space world
- 1:02:55so how are you if we add more letters we
- 1:02:58would just get them like this now what
- 1:03:01is this doing and why is this important
- 1:03:03we are taking our string and instead of
- 1:03:05directly encoding it um for
- 1:03:09tokenization we are first splitting it
- 1:03:11up and when you actually step through
- 1:03:13the code and we'll do that in a bit more
- 1:03:15detail what really is doing on a high
- 1:03:17level is that it first splits your text
- 1:03:20into a list of texts just like this one
- 1:03:24and all these elements of this list are
- 1:03:26processed independently by the tokenizer
- 1:03:29and all of the results of that
- 1:03:30processing are simply
- 1:03:32concatenated so hello world oh I I
- 1:03:35missed how hello world how are you we
- 1:03:39have five elements of list all of these
- 1:03:41will independent
- 1:03:44independently go from text to a token
- 1:03:47sequence and then that token sequence is
- 1:03:49going to be concatenated it's all going
- 1:03:50to be joined up and roughly speaking
- 1:03:54what that does is you're only ever
- 1:03:56finding merges between the elements of
- 1:03:58this list so you can only ever consider
- 1:04:00merges within every one of these
- 1:04:01elements in
- 1:04:03individually and um after you've done
- 1:04:06all the possible merging for all of
- 1:04:07these elements individually the results
- 1:04:09of all that will be joined um by
- 1:04:13concatenation and so you are basically
- 1:04:16what what you're doing effectively is
- 1:04:18you are never going to be merging this e
- 1:04:21with this space because they are now
- 1:04:23parts of the separate elements of this
- 1:04:25list and so you are saying we are never
- 1:04:27going to merge
- 1:04:28eace um because we're breaking it up in
- 1:04:32this way so basically using this regx
- 1:04:35pattern to Chunk Up the text is just one
- 1:04:37way of enforcing that some merges are
- 1:04:41not to happen and we're going to go into
- 1:04:43more of this text and we'll see that
- 1:04:45what this is trying to do on a high
- 1:04:46level is we're trying to not merge
- 1:04:48across letters across numbers across
- 1:04:50punctuation and so on so let's see in
- 1:04:53more detail how that works so let's
- 1:04:54continue now we have/ P ofn if you go to
- 1:04:58the documentation SLP of n is any kind
- 1:05:01of numeric character in any script so
- 1:05:04it's numbers so we have an optional
- 1:05:06space followed by numbers and those
- 1:05:08would be separated out so letters and
- 1:05:10numbers are being separated so if I do
- 1:05:12Hello World 123 how are you then world
- 1:05:15will stop matching here because one is
- 1:05:17not a letter anymore but one is a number
- 1:05:20so this group will match for that and
- 1:05:22we'll get it as a separate entity
- 1:05:26uh let's see how these apostrophes work
- 1:05:28so here if we have
- 1:05:31um uh Slash V or I mean apostrophe V as
- 1:05:35an example then apostrophe here is not a
- 1:05:38letter or a
- 1:05:39number so hello will stop matching and
- 1:05:42then we will exactly match this with
- 1:05:44that so that will come out as a separate
- 1:05:48thing so why are they doing the
- 1:05:50apostrophes here honestly I think that
- 1:05:52these are just like very common
- 1:05:53apostrophes p uh that are used um
- 1:05:56typically I don't love that they've done
- 1:05:59this
- 1:06:00because uh let me show you what happens
- 1:06:03when you have uh some Unicode
- 1:06:05apostrophes like for example you can
- 1:06:07have if you have house then this will be
- 1:06:10separated out because of this matching
- 1:06:13but if you use the Unicode apostrophe
- 1:06:15like
- 1:06:16this then suddenly this does not work
- 1:06:19and so this apostrophe will actually
- 1:06:21become its own thing now and so so um
- 1:06:24it's basically hardcoded for this
- 1:06:26specific kind of apostrophe and uh
- 1:06:29otherwise they become completely
- 1:06:31separate tokens in addition to this you
- 1:06:34can go to the gpt2 docs and here when
- 1:06:38they Define the pattern they say should
- 1:06:40have added re. ignore case so BP merges
- 1:06:43can happen for capitalized versions of
- 1:06:44contractions so what they're pointing
- 1:06:46out is that you see how this is
- 1:06:47apostrophe and then lowercase letters
- 1:06:50well because they didn't do re. ignore
- 1:06:52case then then um these rules will not
- 1:06:56separate out the apostrophes if it's
- 1:06:58uppercase so
- 1:07:01house would be like this but if I did
- 1:07:06house if I'm uppercase then notice
- 1:07:10suddenly the apostrophe comes by
- 1:07:12itself so the tokenization will work
- 1:07:15differently in uppercase and lower case
- 1:07:17inconsistently separating out these
- 1:07:19apostrophes so it feels extremely gnarly
- 1:07:21and slightly gross um but that's that's
- 1:07:24how that works okay so let's come back
- 1:07:27after trying to match a bunch of
- 1:07:28apostrophe Expressions by the way the
- 1:07:30other issue here is that these are quite
- 1:07:32language specific probably so I don't
- 1:07:34know that all the languages for example
- 1:07:35use or don't use apostrophes but that
- 1:07:37would be inconsistently tokenized as a
- 1:07:39result then we try to match letters then
- 1:07:42we try to match numbers and then if that
- 1:07:44doesn't work we fall back to here and
- 1:07:47what this is saying is again optional
- 1:07:49space followed by something that is not
- 1:07:50a letter number or a space in one or
- 1:07:53more of that so what this is doing
- 1:07:55effectively is this is trying to match
- 1:07:57punctuation roughly speaking not letters
- 1:07:59and not numbers so this group will try
- 1:08:02to trigger for that so if I do something
- 1:08:04like this then these parts here are not
- 1:08:08letters or numbers but they will
- 1:08:09actually they are uh they will actually
- 1:08:12get caught here and so they become its
- 1:08:14own group so we've separated out the
- 1:08:17punctuation and finally this um this is
- 1:08:20also a little bit confusing so this is
- 1:08:22matching white space but this is using a
- 1:08:25negative look ahead assertion in regex
- 1:08:29so what this is doing is it's matching
- 1:08:30wh space up to but not including the
- 1:08:33last Whit space
- 1:08:35character why is this important um this
- 1:08:37is pretty subtle I think so you see how
- 1:08:40the white space is always included at
- 1:08:41the beginning of the word so um space r
- 1:08:45space u Etc suppose we have a lot of
- 1:08:48spaces
- 1:08:49here what's going to happen here is that
- 1:08:52these spaces up to not including the
- 1:08:54last character will get caught by this
- 1:08:57and what that will do is it will
- 1:08:59separate out the spaces up to but not
- 1:09:01including the last character so that the
- 1:09:03last character can come here and join
- 1:09:05with the um space you and the reason
- 1:09:09that's nice is because space you is the
- 1:09:11common token so if I didn't have these
- 1:09:13Extra Spaces here you would just have
- 1:09:15space you and if I add tokens if I add
- 1:09:18spaces we still have a space view but
- 1:09:20now we have all this extra white space
- 1:09:22so basically the GB to tokenizer really
- 1:09:24likes to have a space letters or numbers
- 1:09:27um and it it preens these spaces and
- 1:09:30this is just something that it is
- 1:09:31consistent about so that's what that is
- 1:09:33for and then finally we have all the the
- 1:09:36last fallback is um whites space
- 1:09:38characters uh so um that would be
- 1:09:42just um if that doesn't get caught then
- 1:09:46this thing will catch any trailing
- 1:09:48spaces and so on I wanted to show one
- 1:09:50more real world example here so if we
- 1:09:53have this string which is a piece of
- 1:09:54python code and then we try to split it
- 1:09:56up then this is the kind of output we
- 1:09:58get so you'll notice that the list has
- 1:10:00many elements here and that's because we
- 1:10:02are splitting up fairly often uh every
- 1:10:05time sort of a category
- 1:10:07changes um so there will never be any
- 1:10:09merges Within These
- 1:10:10elements and um that's what you are
- 1:10:13seeing here now you might think that in
- 1:10:16order to train the
- 1:10:17tokenizer uh open AI has used this to
- 1:10:21split up text into chunks and then run
- 1:10:23just a BP algorithm within all the
- 1:10:25chunks but that is not exactly what
- 1:10:27happened and the reason is the following
- 1:10:30notice that we have the spaces here uh
- 1:10:33those Spaces end up being entire
- 1:10:35elements but these spaces never actually
- 1:10:38end up being merged by by open Ai and
- 1:10:40the way you can tell is that if you copy
- 1:10:42paste the exact same chunk here into Tik
- 1:10:44token U Tik tokenizer you see that all
- 1:10:47the spaces are kept independent and
- 1:10:49they're all token
- 1:10:51220 so I think opena at some point Point
- 1:10:53en Force some rule that these spaces
- 1:10:56would never be merged and so um there's
- 1:10:59some additional rules on top of just
- 1:11:01chunking and bpe that open ey is not uh
- 1:11:04clear about now the training code for
- 1:11:06the gpt2 tokenizer was never released so
- 1:11:08all we have is uh the code that I've
- 1:11:10already shown you but this code here
- 1:11:13that they've released is only the
- 1:11:14inference code for the tokens so this is
- 1:11:17not the training code you can't give it
- 1:11:19a piece of text and training tokenizer
- 1:11:21this is just the inference code which
- 1:11:23Tak takes the merges that we have up
- 1:11:25above and applies them to a new piece of
- 1:11:28text and so we don't know exactly how
- 1:11:30opening ey trained um train the
- 1:11:32tokenizer but it wasn't as simple as
- 1:11:34chunk it up and BP it uh whatever it was
- 1:11:38next I wanted to introduce you to the
- 1:11:40Tik token library from openai which is
- 1:11:42the official library for tokenization
- 1:11:44from openai so this is Tik token bip
- 1:11:48install P to Tik token and then um you
- 1:11:51can do the tokenization in inference
- 1:11:54this is again not training code this is
- 1:11:55only inference code for
- 1:11:57tokenization um I wanted to show you how
- 1:12:00you would use it quite simple and
- 1:12:02running this just gives us the gpt2
- 1:12:04tokens or the GPT 4 tokens so this is
- 1:12:06the tokenizer use for GPT 4 and so in
- 1:12:09particular we see that the Whit space in
- 1:12:11gpt2 remains unmerged but in GPT 4 uh
- 1:12:14these Whit spaces merge as we also saw
- 1:12:17in this one where here they're all
- 1:12:19unmerged but if we go down to GPT 4 uh
- 1:12:22they become merged
- 1:12:25um now in the
- 1:12:27gp4 uh tokenizer they changed the
- 1:12:31regular expression that they use to
- 1:12:33Chunk Up text so the way to see this is
- 1:12:35that if you come to your the Tik token
- 1:12:38uh library and then you go to this file
- 1:12:41Tik token X openi public this is where
- 1:12:44sort of like the definition of all these
- 1:12:45different tokenizers that openi
- 1:12:46maintains is and so uh necessarily to do
- 1:12:50the inference they had to publish some
- 1:12:51of the details about the strings
- 1:12:53so this is the string that we already
- 1:12:55saw for gpt2 it is slightly different
- 1:12:58but it is actually equivalent uh to what
- 1:13:00we discussed here so this pattern that
- 1:13:02we discussed is equivalent to this
- 1:13:04pattern this one just executes a little
- 1:13:07bit faster so here you see a little bit
- 1:13:09of a slightly different definition but
- 1:13:10otherwise it's the same we're going to
- 1:13:12go into special tokens in a bit and then
- 1:13:15if you scroll down to CL 100k this is
- 1:13:18the GPT 4 tokenizer you see that the
- 1:13:20pattern has changed um and this is kind
- 1:13:23of like the main the major change in
- 1:13:26addition to a bunch of other special
- 1:13:27tokens which I'll go into in a bit again
- 1:13:30now some I'm not going to actually go
- 1:13:31into the full detail of the pattern
- 1:13:33change because honestly this is my
- 1:13:35numbing uh I would just advise that you
- 1:13:37pull out chat GPT and the regex
- 1:13:39documentation and just step through it
- 1:13:42but really the major changes are number
- 1:13:44one you see this eye here that means
- 1:13:48that the um case sensitivity this is
- 1:13:51case insensitive match and so the
- 1:13:53comment that we saw earlier on oh we
- 1:13:56should have used re. uppercase uh
- 1:13:58basically we're now going to be matching
- 1:14:01these apostrophe s apostrophe D
- 1:14:04apostrophe M Etc uh we're going to be
- 1:14:06matching them both in lowercase and in
- 1:14:08uppercase so that's fixed there's a
- 1:14:11bunch of different like handling of the
- 1:14:12whites space that I'm not going to go
- 1:14:14into the full details of and then one
- 1:14:16more thing here is you will notice that
- 1:14:18when they match the numbers they only
- 1:14:20match one to three numbers so so they
- 1:14:23will never merge
- 1:14:26numbers that are in low in more than
- 1:14:28three digits only up to three digits of
- 1:14:31numbers will ever be merged and uh
- 1:14:34that's one change that they made as well
- 1:14:36to prevent uh tokens that are very very
- 1:14:38long number
- 1:14:40sequences uh but again we don't really
- 1:14:42know why they do any of this stuff uh
- 1:14:44because none of this is documented and
- 1:14:46uh it's just we just get the pattern so
- 1:14:49um yeah it is what it is but those are
- 1:14:51some of the changes that gp4 has made
- 1:14:54and of course the vocabulary size went
- 1:14:56from roughly 50k to roughly
- 1:14:58100K the next thing I would like to do
- 1:15:00very briefly is to take you through the
- 1:15:02gpt2 encoder dopy that openi has
- 1:15:05released uh this is the file that I
- 1:15:07already mentioned to you briefly now
- 1:15:09this file is uh fairly short and should
- 1:15:12be relatively understandable to you at
- 1:15:14this point um starting at the bottom
- 1:15:17here they are loading two files encoder
- 1:15:21Json and vocab bpe and they do some
- 1:15:24light processing on it and then they
- 1:15:25call this encoder object which is the
- 1:15:27tokenizer now if you'd like to inspect
- 1:15:30these two files which together
- 1:15:31constitute their saved tokenizer then
- 1:15:34you can do that with a piece of code
- 1:15:36like
- 1:15:36this um this is where you can download
- 1:15:39these two files and you can inspect them
- 1:15:40if you'd like and what you will find is
- 1:15:42that this encoder as they call it in
- 1:15:45their code is exactly equivalent to our
- 1:15:47vocab so remember here where we have
- 1:15:51this vocab object which allowed us us to
- 1:15:53decode very efficiently and basically it
- 1:15:56took us from the integer to the byes uh
- 1:16:00for that integer so our vocab is exactly
- 1:16:03their encoder and then their vocab bpe
- 1:16:07confusingly is actually are merges so
- 1:16:11their BP merges which is based on the
- 1:16:14data inside vocab bpe ends up being
- 1:16:16equivalent to our merges so uh basically
- 1:16:20they are saving and loading the two uh
- 1:16:24variables that for us are also critical
- 1:16:26the merges variable and the vocab
- 1:16:28variable using just these two variables
- 1:16:31you can represent a tokenizer and you
- 1:16:32can both do encoding and decoding once
- 1:16:34you've trained this
- 1:16:36tokenizer now the only thing that um is
- 1:16:40actually slightly confusing inside what
- 1:16:42opening ey does here is that in addition
- 1:16:44to this encoder and a decoder they also
- 1:16:46have something called a bite encoder and
- 1:16:48a bite decoder and this is actually
- 1:16:51unfortunately just
- 1:16:53kind of a spirous implementation detail
- 1:16:55and isn't actually deep or interesting
- 1:16:57in any way so I'm going to skip the
- 1:16:59discussion of it but what opening ey
- 1:17:01does here for reasons that I don't fully
- 1:17:02understand is that not only have they
- 1:17:05this tokenizer which can encode and
- 1:17:06decode but they have a whole separate
- 1:17:08layer here in addition that is used
- 1:17:10serially with the tokenizer and so you
- 1:17:12first do um bite encode and then encode
- 1:17:16and then you do decode and then bite
- 1:17:17decode so that's the loop and they are
- 1:17:20just stacked serial on top of each other
- 1:17:22and and it's not that interesting so I
- 1:17:24won't cover it and you can step through
- 1:17:25it if you'd like otherwise this file if
- 1:17:28you ignore the bite encoder and the bite
- 1:17:30decoder will be algorithmically very
- 1:17:31familiar with you and the meat of it
- 1:17:33here is the what they call bpe function
- 1:17:37and you should recognize this Loop here
- 1:17:39which is very similar to our own y Loop
- 1:17:41where they're trying to identify the
- 1:17:43Byram uh a pair that they should be
- 1:17:46merging next and then here just like we
- 1:17:49had they have a for Loop trying to merge
- 1:17:50this pair uh so they will go over all of
- 1:17:53the sequence and they will merge the
- 1:17:55pair whenever they find it and they keep
- 1:17:57repeating that until they run out of
- 1:17:59possible merges in the in the text so
- 1:18:02that's the meat of this file and uh
- 1:18:04there's an encode and a decode function
- 1:18:06just like we have implemented it so long
- 1:18:08story short what I want you to take away
- 1:18:09at this point is that unfortunately it's
- 1:18:11a little bit of a messy code that they
- 1:18:13have but algorithmically it is identical
- 1:18:15to what we've built up above and what
- 1:18:17we've built up above if you understand
- 1:18:19it is algorithmically what is necessary
- 1:18:21to actually build a BP to organizer
- 1:18:23train it and then both encode and decode
- 1:18:26the next topic I would like to turn to
- 1:18:28is that of special tokens so in addition
- 1:18:30to tokens that are coming from you know
- 1:18:32raw bytes and the BP merges we can
- 1:18:35insert all kinds of tokens that we are
- 1:18:36going to use to delimit different parts
- 1:18:38of the data or introduced to create a
- 1:18:41special structure of the token streams
- 1:18:44so in uh if you look at this encoder
- 1:18:47object from open AIS gpd2 right here we
- 1:18:50mentioned this is very similar to our
- 1:18:52vocab you'll notice that the length of
- 1:18:54this is
- 1:18:5850257 and as I mentioned it's mapping uh
- 1:19:01and it's inverted from the mapping of
- 1:19:03our vocab our vocab goes from integer to
- 1:19:06string and they go the other way around
- 1:19:08for no amazing reason um but the thing
- 1:19:11to note here is that this the mapping
- 1:19:13table here is
- 1:19:1550257 where does that number come from
- 1:19:18where what are the tokens as I mentioned
- 1:19:20there are 256 raw bite token
- 1:19:24tokens and then opena actually did
- 1:19:2750,000
- 1:19:28merges so those become the other tokens
- 1:19:32but this would have been
- 1:19:3450256 so what is the 57th token and
- 1:19:37there is basically one special
- 1:19:40token and that one special token you can
- 1:19:43see is called end of text so this is a
- 1:19:47special token and it's the very last
- 1:19:49token and this token is used to delimit
- 1:19:52documents ments in the training set so
- 1:19:55when we're creating the training data we
- 1:19:57have all these documents and we tokenize
- 1:19:59them and we get a stream of tokens those
- 1:20:01tokens only range from Z to
- 1:20:0550256 and then in between those
- 1:20:07documents we put special end of text
- 1:20:10token and we insert that token in
- 1:20:12between documents and we are using this
- 1:20:15as a signal to the language model that
- 1:20:18the document has ended and what follows
- 1:20:20is going to be unrelated to the document
- 1:20:23previously that said the language model
- 1:20:25has to learn this from data it it needs
- 1:20:27to learn that this token usually means
- 1:20:29that it should wipe its sort of memory
- 1:20:31of what came before and what came before
- 1:20:34this token is not actually informative
- 1:20:35to what comes next but we are expecting
- 1:20:37the language model to just like learn
- 1:20:39this but we're giving it the Special
- 1:20:40sort of the limiter of these documents
- 1:20:44we can go here to Tech tokenizer and um
- 1:20:46this the gpt2 tokenizer uh our code that
- 1:20:49we've been playing with before so we can
- 1:20:51add here right hello world world how are
- 1:20:53you and we're getting different tokens
- 1:20:55but now you can see what if what happens
- 1:20:58if I put end of text you see how until I
- 1:21:02finished it these are all different
- 1:21:03tokens end of
- 1:21:06text still set different tokens and now
- 1:21:08when I finish it suddenly we get token
- 1:21:1350256 and the reason this works is
- 1:21:15because this didn't actually go through
- 1:21:18the bpe merges instead the code that
- 1:21:21actually outposted tokens has special
- 1:21:25case instructions for handling special
- 1:21:28tokens um we did not see these special
- 1:21:30instructions for handling special tokens
- 1:21:32in the encoder dopy it's absent there
- 1:21:36but if you go to Tech token Library
- 1:21:38which is uh implemented in Rust you will
- 1:21:40find all kinds of special case handling
- 1:21:42for these special tokens that you can
- 1:21:44register uh create adds to the
- 1:21:47vocabulary and then it looks for them
- 1:21:49and it uh whenever it sees these special
- 1:21:50tokens like this it will actually come
- 1:21:53in and swap in that special token so
- 1:21:56these things are outside of the typical
- 1:21:58algorithm of uh B PA en
- 1:22:00coding so these special tokens are used
- 1:22:02pervasively uh not just in uh basically
- 1:22:05base language modeling of predicting the
- 1:22:07next token in the sequence but
- 1:22:09especially when it gets to later to the
- 1:22:10fine tuning stage and all of the chat uh
- 1:22:13gbt sort of aspects of it uh because we
- 1:22:15don't just want to Del limit documents
- 1:22:16we want to delimit entire conversations
- 1:22:18between an assistant and a user so if I
- 1:22:21refresh this sck tokenizer page the
- 1:22:24default example that they have here is
- 1:22:26using not sort of base model encoders
- 1:22:30but ftuned model uh sort of tokenizers
- 1:22:33um so for example using the GPT 3.5
- 1:22:35turbo scheme these here are all special
- 1:22:38tokens I am start I end Etc uh this is
- 1:22:43short for Imaginary mcore start by the
- 1:22:46way but you can see here that there's a
- 1:22:49sort of start and end of every single
- 1:22:51message and there can be many other
- 1:22:52other tokens lots of tokens um in use to
- 1:22:56delimit these conversations and kind of
- 1:22:58keep track of the flow of the messages
- 1:23:00here now we can go back to the Tik token
- 1:23:03library and here when you scroll to the
- 1:23:06bottom they talk about how you can
- 1:23:08extend tick token and I can you can
- 1:23:10create basically you can Fork uh the um
- 1:23:13CL 100K base tokenizers in gp4 and for
- 1:23:17example you can extend it by adding more
- 1:23:18special tokens and these are totally up
- 1:23:20to you you can come up with any
- 1:23:21arbitrary tokens and add them with the
- 1:23:23new ID afterwards and the tikken library
- 1:23:26will uh correctly swap them out uh when
- 1:23:29it sees this in the
- 1:23:31strings now we can also go back to this
- 1:23:34file which we've looked at previously
- 1:23:37and I mentioned that the gpt2 in Tik
- 1:23:39toen open
- 1:23:41I.P we have the vocabulary we have the
- 1:23:44pattern for splitting and then here we
- 1:23:46are registering the single special token
- 1:23:48in gpd2 which was the end of text token
- 1:23:50and we saw that it has this ID
- 1:23:53in GPT 4 when they defy this here you
- 1:23:56see that the pattern has changed as
- 1:23:57we've discussed but also the special
- 1:23:59tokens have changed in this tokenizer so
- 1:24:01we of course have the end of text just
- 1:24:03like in gpd2 but we also see three sorry
- 1:24:06four additional tokens here Thim prefix
- 1:24:09middle and suffix what is fim fim is
- 1:24:12short for fill in the middle and if
- 1:24:14you'd like to learn more about this idea
- 1:24:17it comes from this paper um and I'm not
- 1:24:20going to go into detail in this video
- 1:24:21it's beyond this video and then there's
- 1:24:23one additional uh serve token here so
- 1:24:27that's that encoding as well so it's
- 1:24:29very common basically to train a
- 1:24:31language model and then if you'd like uh
- 1:24:34you can add special tokens now when you
- 1:24:37add special tokens you of course have to
- 1:24:39um do some model surgery to the
- 1:24:41Transformer and all the parameters
- 1:24:43involved in that Transformer because you
- 1:24:45are basically adding an integer and you
- 1:24:47want to make sure that for example your
- 1:24:48embedding Matrix for the vocabulary
- 1:24:50tokens has to be extended by adding a
- 1:24:53row and typically this row would be
- 1:24:54initialized uh with small random numbers
- 1:24:56or something like that because we need
- 1:24:58to have a vector that now stands for
- 1:25:01that token in addition to that you have
- 1:25:03to go to the final layer of the
- 1:25:04Transformer and you have to make sure
- 1:25:05that that projection at the very end
- 1:25:07into the classifier uh is extended by
- 1:25:09one as well so basically there's some
- 1:25:11model surgery involved that you have to
- 1:25:13couple with the tokenization changes if
- 1:25:16you are going to add special tokens but
- 1:25:18this is a very common operation that
- 1:25:20people do especially if they'd like to
- 1:25:21fine tune the model for example taking
- 1:25:23it from a base model to a chat model
- 1:25:26like chat
- 1:25:27GPT okay so at this point you should
- 1:25:29have everything you need in order to
- 1:25:31build your own gp4 tokenizer now in the
- 1:25:33process of developing this lecture I've
- 1:25:35done that and I published the code under
- 1:25:37this repository
- 1:25:38MBP so MBP looks like this right now as
- 1:25:42I'm recording but uh the MBP repository
- 1:25:45will probably change quite a bit because
- 1:25:46I intend to continue working on it um in
- 1:25:49addition to the MBP repository I've
- 1:25:51published the this uh exercise
- 1:25:53progression that you can follow so if
- 1:25:55you go to exercise. MD here uh this is
- 1:25:58sort of me breaking up the task ahead of
- 1:26:01you into four steps that sort of uh
- 1:26:03build up to what can be a gp4 tokenizer
- 1:26:06and so feel free to follow these steps
- 1:26:08exactly and follow a little bit of the
- 1:26:10guidance that I've laid out here and
- 1:26:12anytime you feel stuck just reference
- 1:26:14the MBP repository here so either the
- 1:26:17tests could be useful or the MBP
- 1:26:20repository itself I try to keep the code
- 1:26:22fairly clean and understandable and so
- 1:26:26um feel free to reference it whenever um
- 1:26:28you get
- 1:26:30stuck uh in addition to that basically
- 1:26:32once you write it you should be able to
- 1:26:34reproduce this behavior from Tech token
- 1:26:36so getting the gb4 tokenizer you can
- 1:26:39take uh you can encode the string and
- 1:26:41you should get these tokens and then you
- 1:26:43can encode and decode the exact same
- 1:26:44string to recover it and in addition to
- 1:26:47all that you should be able to implement
- 1:26:48your own train function uh which Tik
- 1:26:50token Library does not provide it's it's
- 1:26:52again only inference code but you could
- 1:26:54write your own train MBP does it as well
- 1:26:57and that will allow you to train your
- 1:26:59own token
- 1:27:00vocabularies so here are some of the
- 1:27:02code inside M be mean bpe uh shows the
- 1:27:06token vocabularies that you might obtain
- 1:27:08so on the left uh here we have the GPT 4
- 1:27:12merges uh so the first 256 are raw
- 1:27:15individual bytes and then here I am
- 1:27:17visualizing the merges that gp4
- 1:27:19performed during its training so the
- 1:27:21very first merge that gp4 did was merge
- 1:27:24two spaces into a single token for you
- 1:27:27know two spaces and that is a token 256
- 1:27:30and so this is the order in which things
- 1:27:32merged during gb4 training and this is
- 1:27:34the merge order that um we obtain in MBP
- 1:27:39by training a tokenizer and in this case
- 1:27:41I trained it on a Wikipedia page of
- 1:27:43Taylor Swift uh not because I'm a Swifty
- 1:27:45but because that is one of the longest
- 1:27:47um Wikipedia Pages apparently that's
- 1:27:49available but she is pretty cool and
- 1:27:54um what was I going to say yeah so you
- 1:27:56can compare these two uh vocabularies
- 1:27:59and so as an example um here GPT for
- 1:28:04merged I in to become in and we've done
- 1:28:06the exact same thing on this token 259
- 1:28:10here space t becomes space t and that
- 1:28:13happened for us a little bit later as
- 1:28:14well so the difference here is again to
- 1:28:16my understanding only a difference of
- 1:28:18the training set so as an example
- 1:28:20because I see a lot of white space I
- 1:28:22supect that gp4 probably had a lot of
- 1:28:23python code in its training set I'm not
- 1:28:25sure uh for the
- 1:28:27tokenizer and uh here we see much less
- 1:28:30of that of course in the Wikipedia page
- 1:28:32so roughly speaking they look the same
- 1:28:34and they look the same because they're
- 1:28:35running the same algorithm and when you
- 1:28:38train your own you're probably going to
- 1:28:39get something similar depending on what
- 1:28:41you train it on okay so we are now going
- 1:28:43to move on from tick token and the way
- 1:28:45that open AI tokenizes its strings and
- 1:28:47we're going to discuss one more very
- 1:28:49commonly used library for working with
- 1:28:51tokenization inlm
- 1:28:52and that is sentence piece so sentence
- 1:28:55piece is very commonly used in language
- 1:28:58models because unlike Tik token it can
- 1:29:00do both training and inference and is
- 1:29:02quite efficient at both it supports a
- 1:29:04number of algorithms for training uh
- 1:29:06vocabularies but one of them is the B
- 1:29:09pair en coding algorithm that we've been
- 1:29:10looking at so it supports it now
- 1:29:13sentence piece is used both by llama and
- 1:29:15mistal series and many other models as
- 1:29:18well it is on GitHub under Google
- 1:29:20sentence piece
- 1:29:22and the big difference with sentence
- 1:29:24piece and we're going to look at example
- 1:29:26because this is kind of hard and subtle
- 1:29:27to explain is that they think different
- 1:29:31about the order of operations here so in
- 1:29:35the case of Tik token we first take our
- 1:29:38code points in the string we encode them
- 1:29:41using mutf to bytes and then we're
- 1:29:42merging bytes it's fairly
- 1:29:44straightforward for sentence piece um it
- 1:29:48works directly on the level of the code
- 1:29:50points themselves so so it looks at
- 1:29:52whatever code points are available in
- 1:29:53your training set and then it starts
- 1:29:55merging those code points and um the bpe
- 1:29:59is running on the level of code
- 1:30:01points and if you happen to run out of
- 1:30:04code points so there are maybe some rare
- 1:30:06uh code points that just don't come up
- 1:30:08too often and the Rarity is determined
- 1:30:09by this character coverage hyper
- 1:30:11parameter then these uh code points will
- 1:30:14either get mapped to a special unknown
- 1:30:16token like ank or if you have the bite
- 1:30:19foldback option turned on then that will
- 1:30:22take those rare Cod points it will
- 1:30:23encode them using utf8 and then the
- 1:30:26individual bytes of that encoding will
- 1:30:27be translated into tokens and there are
- 1:30:30these special bite tokens that basically
- 1:30:32get added to the vocabulary so it uses
- 1:30:35BP on on the code points and then it
- 1:30:38falls back to bytes for rare Cod points
- 1:30:41um and so that's kind of like difference
- 1:30:44personally I find the Tik token we
- 1:30:45significantly cleaner uh but it's kind
- 1:30:47of like a subtle but pretty major
- 1:30:48difference between the way they approach
- 1:30:50tokenization let's work with with a
- 1:30:52concrete example because otherwise this
- 1:30:54is kind of hard to um to get your head
- 1:30:56around so let's work with a concrete
- 1:30:59example this is how we can import
- 1:31:01sentence piece and then here we're going
- 1:31:03to take I think I took like the
- 1:31:05description of sentence piece and I just
- 1:31:06created like a little toy data set it
- 1:31:08really likes to have a file so I created
- 1:31:10a toy. txt file with this
- 1:31:13content now what's kind of a little bit
- 1:31:15crazy about sentence piece is that
- 1:31:16there's a ton of options and
- 1:31:18configurations and the reason this is so
- 1:31:20is because sentence piece has been
- 1:31:22around I think for a while and it really
- 1:31:23tries to handle a large diversity of
- 1:31:25things and um because it's been around I
- 1:31:28think it has quite a bit of accumulated
- 1:31:30historical baggage uh as well and so in
- 1:31:33particular there's like a ton of
- 1:31:35configuration arguments this is not even
- 1:31:36all of it you can go to here to see all
- 1:31:39the training
- 1:31:40options um and uh there's also quite
- 1:31:44useful documentation when you look at
- 1:31:45the raw Proto buff uh that is used to
- 1:31:48represent the trainer spec and so on um
- 1:31:52many of these options are irrelevant to
- 1:31:54us so maybe to point out one example Das
- 1:31:56Das shrinking Factor uh this shrinking
- 1:31:59factor is not used in the B pair en
- 1:32:01coding algorithm so this is just an
- 1:32:03argument that is irrelevant to us um it
- 1:32:05applies to a different training
- 1:32:09algorithm now what I tried to do here is
- 1:32:11I tried to set up sentence piece in a
- 1:32:13way that is very very similar as far as
- 1:32:15I can tell to maybe identical hopefully
- 1:32:18to the way that llama 2 was strained so
- 1:32:22the way they trained their own um their
- 1:32:25own tokenizer and the way I did this was
- 1:32:27basically you can take the tokenizer
- 1:32:28model file that meta released and you
- 1:32:31can um open it using the Proto protuff
- 1:32:35uh sort of file that you can generate
- 1:32:38and then you can inspect all the options
- 1:32:39and I tried to copy over all the options
- 1:32:41that looked relevant so here we set up
- 1:32:43the input it's raw text in this file
- 1:32:46here's going to be the output so it's
- 1:32:48going to be for talk 400. model and
- 1:32:50vocab
- 1:32:52we're saying that we're going to use the
- 1:32:53BP algorithm and we want to Bap size of
- 1:32:56400 then there's a ton of configurations
- 1:32:58here
- 1:33:01for um for basically pre-processing and
- 1:33:05normalization rules as they're called
- 1:33:07normalization used to be very prevalent
- 1:33:09I would say before llms in natural
- 1:33:11language processing so in machine
- 1:33:12translation and uh text classification
- 1:33:14and so on you want to normalize and
- 1:33:16simplify the text and you want to turn
- 1:33:18it all lowercase and you want to remove
- 1:33:19all double whites space Etc
- 1:33:22and in language models we prefer not to
- 1:33:23do any of it or at least that is my
- 1:33:25preference as a deep learning person you
- 1:33:26want to not touch your data you want to
- 1:33:28keep the raw data as much as possible um
- 1:33:31in a raw
- 1:33:33form so you're basically trying to turn
- 1:33:35off a lot of this if you can the other
- 1:33:38thing that sentence piece does is that
- 1:33:39it has this concept of sentences so
- 1:33:43sentence piece it's back it's kind of
- 1:33:45like was developed I think early in the
- 1:33:46days where there was um an idea that
- 1:33:50they you're training a tokenizer on a
- 1:33:51bunch of independent sentences so it has
- 1:33:54a lot of like how many sentences you're
- 1:33:56going to train on what is the maximum
- 1:33:58sentence length
- 1:34:00um shuffling sentences and so for it
- 1:34:03sentences are kind of like the
- 1:34:04individual training examples but again
- 1:34:06in the context of llms I find that this
- 1:34:08is like a very spous and weird
- 1:34:10distinction like sentences are just like
- 1:34:13don't touch the raw data sentences
- 1:34:15happen to exist but in raw data sets
- 1:34:18there are a lot of like inet like what
- 1:34:20exactly is a sentence what isn't a
- 1:34:22sentence um and so I think like it's
- 1:34:25really hard to Define what an actual
- 1:34:26sentence is if you really like dig into
- 1:34:28it and there could be different concepts
- 1:34:30of it in different languages or
- 1:34:32something like that so why even
- 1:34:33introduce the concept it it doesn't
- 1:34:35honestly make sense to me I would just
- 1:34:36prefer to treat a file as a giant uh
- 1:34:39stream of
- 1:34:40bytes it has a lot of treatment around
- 1:34:42rare word characters and when I say word
- 1:34:45I mean code points we're going to come
- 1:34:46back to this in a second and it has a
- 1:34:48lot of other rules for um basically
- 1:34:51splitting digits splitting white space
- 1:34:54and numbers and how you deal with that
- 1:34:56so these are some kind of like merge
- 1:34:58rules so I think this is a little bit
- 1:35:00equivalent to tick token using the
- 1:35:02regular expression to split up
- 1:35:04categories there's like kind of
- 1:35:07equivalence of it if you squint T it in
- 1:35:09sentence piece where you can also for
- 1:35:10example split up split up the digits uh
- 1:35:14and uh so
- 1:35:15on there's a few more things here that
- 1:35:18I'll come back to in a bit and then
- 1:35:19there are some special tokens that you
- 1:35:20can indicate and it hardcodes the UN
- 1:35:23token the beginning of sentence end of
- 1:35:25sentence and a pad token um and the UN
- 1:35:29token must exist for my understanding
- 1:35:32and then some some things so we can
- 1:35:34train and when when I press train it's
- 1:35:37going to create this file talk 400.
- 1:35:40model and talk 400. wab I can then load
- 1:35:43the model file and I can inspect the
- 1:35:45vocabulary off it and so we trained
- 1:35:48vocab size 400 on this text here and
- 1:35:53these are the individual pieces the
- 1:35:55individual tokens that sentence piece
- 1:35:56will create so in the beginning we see
- 1:35:58that we have the an token uh with the ID
- 1:36:02zero then we have the beginning of
- 1:36:04sequence end of sequence one and two and
- 1:36:07then we said that the pad ID is negative
- 1:36:091 so we chose not to use it so there's
- 1:36:12no pad ID
- 1:36:13here then these are individual bite
- 1:36:16tokens so here we saw that bite fallback
- 1:36:20in llama was turned on so it's true so
- 1:36:23what follows are going to be the 256
- 1:36:26bite
- 1:36:27tokens and these are their
- 1:36:31IDs and then at the bottom after the
- 1:36:35bite tokens come the
- 1:36:37merges and these are the parent nodes in
- 1:36:40the merges so we're not seeing the
- 1:36:42children we're just seeing the parents
- 1:36:43and their
- 1:36:44ID and then after the
- 1:36:47merges comes eventually the individual
- 1:36:50tokens and their IDs and so these are
- 1:36:53the individual tokens so these are the
- 1:36:55individual code Point tokens if you will
- 1:36:58and they come at the end so that is the
- 1:37:00ordering with which sentence piece sort
- 1:37:01of like represents its vocabularies it
- 1:37:03starts with special tokens then the bike
- 1:37:06tokens then the merge tokens and then
- 1:37:08the individual codo tokens and all these
- 1:37:11raw codepoint to tokens are the ones
- 1:37:14that it encountered in the training
- 1:37:16set so those individual code points are
- 1:37:19all the the entire set of code points
- 1:37:22that occurred
- 1:37:24here so those all get put in there and
- 1:37:27then those that are extremely rare as
- 1:37:29determined by character coverage so if a
- 1:37:31code Point occurred only a single time
- 1:37:32out of like a million um sentences or
- 1:37:35something like that then it would be
- 1:37:37ignored and it would not be added to our
- 1:37:40uh
- 1:37:41vocabulary once we have a vocabulary we
- 1:37:43can encode into IDs and we can um sort
- 1:37:46of get a
- 1:37:47list and then here I am also decoding
- 1:37:50the indiv idual tokens back into little
- 1:37:54pieces as they call it so let's take a
- 1:37:56look at what happened here hello space
- 1:38:01on so these are the token IDs we got
- 1:38:04back and when we look here uh a few
- 1:38:07things sort of uh jump to mind number
- 1:38:11one take a look at these characters the
- 1:38:14Korean characters of course were not
- 1:38:15part of the training set so sentence
- 1:38:18piece is encountering code points that
- 1:38:19it has not seen during training time and
- 1:38:22those code points do not have a token
- 1:38:24associated with them so suddenly these
- 1:38:26are un tokens unknown tokens but because
- 1:38:30bite fall back as true instead sentence
- 1:38:33piece falls back to bytes and so it
- 1:38:36takes this it encodes it with utf8 and
- 1:38:39then it uses these tokens to represent
- 1:38:43uh those bytes and that's what we are
- 1:38:45getting sort of here this is the utf8 uh
- 1:38:49encoding and in this shifted by three uh
- 1:38:52because of these um special tokens here
- 1:38:56that have IDs earlier on so that's what
- 1:38:58happened here now one more thing that um
- 1:39:02well first before I go on with respect
- 1:39:05to the bitef back let me remove bite
- 1:39:08foldback if this is false what's going
- 1:39:10to happen let's
- 1:39:12retrain so the first thing that happened
- 1:39:14is all the bite tokens disappeared right
- 1:39:17and now we just have the merges and we
- 1:39:19have a lot more merges now because we
- 1:39:20have a lot more space because we're not
- 1:39:21taking up space in the wab size uh with
- 1:39:25all the
- 1:39:25bytes and now if we encode
- 1:39:29this we get a zero so this entire string
- 1:39:33here suddenly there's no bitef back so
- 1:39:35this is unknown and unknown is an and so
- 1:39:39this is zero because the an token is
- 1:39:42token zero and you have to keep in mind
- 1:39:44that this would feed into your uh
- 1:39:46language model so what is a language
- 1:39:48model supposed to do when all kinds of
- 1:39:49different things that are unrecognized
- 1:39:52because they're rare just end up mapping
- 1:39:54into Unk it's not exactly the property
- 1:39:56that you want so that's why I think
- 1:39:57llama correctly uh used by fallback true
- 1:40:02uh because we definitely want to feed
- 1:40:03these um unknown or rare code points
- 1:40:06into the model and some uh some manner
- 1:40:08the next thing I want to show you is the
- 1:40:10following notice here when we are
- 1:40:12decoding all the individual tokens you
- 1:40:14see how spaces uh space here ends up
- 1:40:18being this um bold underline I'm not
- 1:40:21100% sure by the way why sentence piece
- 1:40:23switches whites space into these bold
- 1:40:25underscore characters maybe it's for
- 1:40:27visualization I'm not 100% sure why that
- 1:40:29happens uh but notice this why do we
- 1:40:32have an extra space in the front of
- 1:40:37hello um what where is this coming from
- 1:40:40well it's coming from this option
- 1:40:43here
- 1:40:45um add dummy prefix is true and when you
- 1:40:48go to the
- 1:40:49documentation add D whites space at the
- 1:40:51beginning of text in order to treat
- 1:40:53World in world and hello world in the
- 1:40:55exact same way so what this is trying to
- 1:40:57do is the
- 1:40:59following if we go back to our tick
- 1:41:02tokenizer world as uh token by itself
- 1:41:06has a different ID than space world so
- 1:41:10we have this is 1917 but this is 14 Etc
- 1:41:14so these are two different tokens for
- 1:41:16the language model and the language
- 1:41:17model has to learn from data that they
- 1:41:18are actually kind of like a very similar
- 1:41:20concept so to the language model in the
- 1:41:23Tik token World um basically words in
- 1:41:26the beginning of sentences and words in
- 1:41:27the middle of sentences actually look
- 1:41:29completely different um and it has to
- 1:41:32learned that they are roughly the same
- 1:41:34so this add dami prefix is trying to
- 1:41:36fight that a little bit and the way that
- 1:41:38works is that it basically
- 1:41:41uh adds a dummy prefix so for as a as a
- 1:41:46part of pre-processing it will take the
- 1:41:49string and it will add a space it will
- 1:41:51do this and that's done in an effort to
- 1:41:54make this world and that world the same
- 1:41:57they will both be space world so that's
- 1:42:00one other kind of pre-processing option
- 1:42:02that is turned on and llama 2 also uh
- 1:42:05uses this option and that's I think
- 1:42:07everything that I want to say for my
- 1:42:08preview of sentence piece and how it is
- 1:42:10different um maybe here what I've done
- 1:42:13is I just uh put in the Raw protocol
- 1:42:16buffer representation basically of the
- 1:42:19tokenizer the too trained so feel free
- 1:42:22to sort of Step through this and if you
- 1:42:24would like uh your tokenization to look
- 1:42:27identical to that of the meta uh llama 2
- 1:42:30then you would be copy pasting these
- 1:42:31settings as I tried to do up above and
- 1:42:34uh yeah that's I think that's it for
- 1:42:36this section I think my summary for
- 1:42:38sentence piece from all of this is
- 1:42:40number one I think that there's a lot of
- 1:42:42historical baggage in sentence piece a
- 1:42:44lot of Concepts that I think are
- 1:42:45slightly confusing and I think
- 1:42:47potentially um contain foot guns like
- 1:42:49this concept of a sentence and it's
- 1:42:50maximum length and stuff like that um
- 1:42:53otherwise it is fairly commonly used in
- 1:42:55the industry um because it is efficient
- 1:42:58and can do both training and inference
- 1:43:01uh it has a few quirks like for example
- 1:43:02un token must exist and the way the bite
- 1:43:05fallbacks are done and so on I don't
- 1:43:06find particularly elegant and
- 1:43:08unfortunately I have to say it's not
- 1:43:09very well documented so it took me a lot
- 1:43:11of time working with this myself um and
- 1:43:14just visualizing things and trying to
- 1:43:16really understand what is happening here
- 1:43:17because uh the documentation
- 1:43:19unfortunately is in my opion not not
- 1:43:21super amazing but it is a very nice repo
- 1:43:24that is available to you if you'd like
- 1:43:26to train your own tokenizer right now
- 1:43:28okay let me now switch gears again as
- 1:43:29we're starting to slowly wrap up here I
- 1:43:31want to revisit this issue in a bit more
- 1:43:33detail of how we should set the vocap
- 1:43:35size and what are some of the
- 1:43:36considerations around it so for this I'd
- 1:43:39like to go back to the model
- 1:43:40architecture that we developed in the
- 1:43:42last video when we built the GPT from
- 1:43:44scratch so this here was uh the file
- 1:43:47that we built in the previous video and
- 1:43:49we defined the Transformer model and and
- 1:43:51let's specifically look at Bap size and
- 1:43:52where it appears in this file so here we
- 1:43:55Define the voap size uh at this time it
- 1:43:58was 65 or something like that extremely
- 1:43:59small number so this will grow much
- 1:44:02larger you'll see that Bap size doesn't
- 1:44:04come up too much in most of these layers
- 1:44:06the only place that it comes up to is in
- 1:44:08exactly these two places here so when we
- 1:44:11Define the language model there's the
- 1:44:13token embedding table which is this
- 1:44:15two-dimensional array where the vocap
- 1:44:18size is basically the number of rows and
- 1:44:21uh each vocabulary element each token
- 1:44:23has a vector that we're going to train
- 1:44:25using back propagation that Vector is of
- 1:44:27size and embed which is number of
- 1:44:29channels in the Transformer and
- 1:44:31basically as voap size increases this
- 1:44:33embedding table as I mentioned earlier
- 1:44:35is going to also grow we're going to be
- 1:44:37adding rows in addition to that at the
- 1:44:39end of the Transformer there's this LM
- 1:44:41head layer which is a linear layer and
- 1:44:44you'll notice that that layer is used at
- 1:44:46the very end to produce the logits uh
- 1:44:48which become the probabilities for the
- 1:44:49next token in sequence and so
- 1:44:51intuitively we're trying to produce a
- 1:44:53probability for every single token that
- 1:44:56might come next at every point in time
- 1:44:58of that Transformer and if we have more
- 1:45:01and more tokens we need to produce more
- 1:45:02and more probabilities so every single
- 1:45:04token is going to introduce an
- 1:45:06additional dot product that we have to
- 1:45:08do here in this linear layer for this
- 1:45:10final layer in a
- 1:45:11Transformer so why can't vocap size be
- 1:45:14infinite why can't we grow to Infinity
- 1:45:16well number one your token embedding
- 1:45:18table is going to grow uh your linear
- 1:45:21layer is going to grow so we're going to
- 1:45:23be doing a lot more computation here
- 1:45:25because this LM head layer will become
- 1:45:26more computational expensive number two
- 1:45:29because we have more parameters we could
- 1:45:30be worried that we are going to be under
- 1:45:33trining some of these
- 1:45:35parameters so intuitively if you have a
- 1:45:37very large vocabulary size say we have a
- 1:45:38million uh tokens then every one of
- 1:45:41these tokens is going to come up more
- 1:45:42and more rarely in the training data
- 1:45:45because there's a lot more other tokens
- 1:45:46all over the place and so we're going to
- 1:45:48be seeing fewer and fewer examples uh
- 1:45:51for each individual token and you might
- 1:45:53be worried that basically the vectors
- 1:45:55associated with every token will be
- 1:45:56undertrained as a result because they
- 1:45:58just don't come up too often and they
- 1:45:59don't participate in the forward
- 1:46:00backward pass in addition to that as
- 1:46:03your vocab size grows you're going to
- 1:46:04start shrinking your sequences a lot
- 1:46:07right and that's really nice because
- 1:46:09that means that we're going to be
- 1:46:10attending to more and more text so
- 1:46:12that's nice but also you might be
- 1:46:13worrying that two large of chunks are
- 1:46:15being squished into single tokens and so
- 1:46:18the model just doesn't have as much of
- 1:46:20time to think per sort of um some number
- 1:46:25of characters in the text or you can
- 1:46:26think about it that way right so
- 1:46:28basically we're squishing too much
- 1:46:29information into a single token and then
- 1:46:31the forward pass of the Transformer is
- 1:46:33not enough to actually process that
- 1:46:34information appropriately and so these
- 1:46:36are some of the considerations you're
- 1:46:37thinking about when you're designing the
- 1:46:38vocab size as I mentioned this is mostly
- 1:46:40an empirical hyperparameter and it seems
- 1:46:42like in state-of-the-art architectures
- 1:46:44today this is usually in the high 10,000
- 1:46:46or somewhere around 100,000 today and
- 1:46:49the next consideration I want to briefly
- 1:46:50talk about is what if we want to take a
- 1:46:53pre-trained model and we want to extend
- 1:46:55the vocap size and this is done fairly
- 1:46:57commonly actually so for example when
- 1:46:58you're doing fine-tuning for cha GPT um
- 1:47:02a lot more new special tokens get
- 1:47:03introduced on top of the base model to
- 1:47:05maintain the metadata and all the
- 1:47:08structure of conversation objects
- 1:47:09between a user and an assistant so that
- 1:47:11takes a lot of special tokens you might
- 1:47:14also try to throw in more special tokens
- 1:47:15for example for using the browser or any
- 1:47:17other tool and so it's very tempting to
- 1:47:20add a lot of tokens for all kinds of
- 1:47:22special functionality so if you want to
- 1:47:24be adding a token that's totally
- 1:47:25possible Right all we have to do is we
- 1:47:27have to resize this embedding so we have
- 1:47:29to add rows we would initialize these uh
- 1:47:32parameters from scratch to be small
- 1:47:34random numbers and then we have to
- 1:47:36extend the weight inside this linear uh
- 1:47:39so we have to start making dot products
- 1:47:41um with the associated parameters as
- 1:47:43well to basically calculate the
- 1:47:44probabilities for these new tokens so
- 1:47:46both of these are just a resizing
- 1:47:48operation it's a very mild
- 1:47:50model surgery and can be done fairly
- 1:47:52easily and it's quite common that
- 1:47:54basically you would freeze the base
- 1:47:55model you introduce these new parameters
- 1:47:57and then you only train these new
- 1:47:58parameters to introduce new tokens into
- 1:48:00the architecture um and so you can
- 1:48:03freeze arbitrary parts of it or you can
- 1:48:04train arbitrary parts of it and that's
- 1:48:06totally up to you but basically minor
- 1:48:08surgery required if you'd like to
- 1:48:10introduce new tokens and finally I'd
- 1:48:11like to mention that actually there's an
- 1:48:13entire design space of applications in
- 1:48:15terms of introducing new tokens into a
- 1:48:17vocabulary that go Way Beyond just
- 1:48:19adding special tokens and special new
- 1:48:21functionality so just to give you a
- 1:48:23sense of the design space but this could
- 1:48:24be an entire video just by itself uh
- 1:48:26this is a paper on learning to compress
- 1:48:28prompts with what they called uh gist
- 1:48:31tokens and the rough idea is suppose
- 1:48:33that you're using language models in a
- 1:48:34setting that requires very long prompts
- 1:48:37while these long prompts just slow
- 1:48:38everything down because you have to
- 1:48:39encode them and then you have to use
- 1:48:41them and then you're tending over them
- 1:48:43and it's just um you know heavy to have
- 1:48:45very large prompts so instead what they
- 1:48:47do here in this paper is they introduce
- 1:48:50new tokens and um imagine basically
- 1:48:54having a few new tokens you put them in
- 1:48:56a sequence and then you train the model
- 1:48:59by distillation so you are keeping the
- 1:49:01entire model Frozen and you're only
- 1:49:03training the representations of the new
- 1:49:05tokens their embeddings and you're
- 1:49:06optimizing over the new tokens such that
- 1:49:09the behavior of the language model is
- 1:49:11identical uh to the model that has a
- 1:49:15very long prompt that works for you and
- 1:49:17so it's a compression technique of
- 1:49:19compressing that very long prompt into
- 1:49:20those few new gist tokens and so you can
- 1:49:23train this and then at test time you can
- 1:49:25discard your old prompt and just swap in
- 1:49:26those tokens and they sort of like uh
- 1:49:28stand in for that very long prompt and
- 1:49:31have an almost identical performance and
- 1:49:33so this is one um technique and a class
- 1:49:36of parameter efficient fine-tuning
- 1:49:38techniques where most of the model is
- 1:49:39basically fixed and there's no training
- 1:49:41of the model weights there's no training
- 1:49:43of Laura or anything like that of new
- 1:49:45parameters the the parameters that
- 1:49:47you're training are now just the uh
- 1:49:49token embeddings so that's just one
- 1:49:51example but this could again be like an
- 1:49:52entire video but just to give you a
- 1:49:54sense that there's a whole design space
- 1:49:55here that is potentially worth exploring
- 1:49:57in the future the next thing I want to
- 1:49:59briefly address is that I think recently
- 1:50:01there's a lot of momentum in how you
- 1:50:03actually could construct Transformers
- 1:50:05that can simultaneously process not just
- 1:50:06text as the input modality but a lot of
- 1:50:08other modalities so be it images videos
- 1:50:11audio Etc and how do you feed in all
- 1:50:14these modalities and potentially predict
- 1:50:16these modalities from a Transformer uh
- 1:50:18do you have to change the architecture
- 1:50:19in some fundamental way and I think what
- 1:50:21a lot of people are starting to converge
- 1:50:23towards is that you're not changing the
- 1:50:24architecture you stick with the
- 1:50:25Transformer you just kind of tokenize
- 1:50:27your input domains and then call the day
- 1:50:29and pretend it's just text tokens and
- 1:50:31just do everything else identical in an
- 1:50:33identical manner so here for example
- 1:50:36there was a early paper that has nice
- 1:50:37graphic for how you can take an image
- 1:50:39and you can chunc at it into
- 1:50:42integers um and these sometimes uh so
- 1:50:45these will basically become the tokens
- 1:50:46of images as an example and uh these
- 1:50:49tokens can be uh hard tokens where you
- 1:50:52force them to be integers they can also
- 1:50:53be soft tokens where you uh sort of
- 1:50:57don't require uh these to be discrete
- 1:51:00but you do Force these representations
- 1:51:02to go through bottlenecks like in Auto
- 1:51:04encoders uh also in this paper that came
- 1:51:06out from open a SORA which I think
- 1:51:08really um uh blew the mind of many
- 1:51:11people and inspired a lot of people in
- 1:51:13terms of what's possible they have a
- 1:51:15Graphic here and they talk briefly about
- 1:51:16how llms have text tokens Sora has
- 1:51:20visual patches so again they came up
- 1:51:22with a way to chunc a videos into
- 1:51:24basically tokens when they own
- 1:51:26vocabularies and then you can either
- 1:51:28process discrete tokens say with autog
- 1:51:30regressive models or even soft tokens
- 1:51:32with diffusion models and uh all of that
- 1:51:35is sort of uh being actively worked on
- 1:51:38designed on and is beyond the scope of
- 1:51:39this video but just something I wanted
- 1:51:40to mention briefly okay now that we have
- 1:51:42come quite deep into the tokenization
- 1:51:45algorithm and we understand a lot more
- 1:51:46about how it works let's loop back
- 1:51:48around to the beginning of this video
- 1:51:50and go through some of these bullet
- 1:51:51points and really see why they happen so
- 1:51:54first of all why can't my llm spell
- 1:51:56words very well or do other spell
- 1:51:58related
- 1:52:00tasks so fundamentally this is because
- 1:52:02as we saw these characters are chunked
- 1:52:05up into tokens and some of these tokens
- 1:52:07are actually fairly long so as an
- 1:52:10example I went to the gp4 vocabulary and
- 1:52:12I looked at uh one of the longer tokens
- 1:52:15so that default style turns out to be a
- 1:52:17single individual token so that's a lot
- 1:52:19of characters for a single token so my
- 1:52:22suspicion is that there's just too much
- 1:52:23crammed into this single token and my
- 1:52:26suspicion was that the model should not
- 1:52:27be very good at tasks related to
- 1:52:30spelling of this uh single token so I
- 1:52:34asked how many letters L are there in
- 1:52:37the word default style and of course my
- 1:52:41prompt is intentionally done that way
- 1:52:44and you see how default style will be a
- 1:52:45single token so this is what the model
- 1:52:47sees so my suspicion is that it wouldn't
- 1:52:49be very good at this and indeed it is
- 1:52:51not it doesn't actually know how many
- 1:52:53L's are in there it thinks there are
- 1:52:54three and actually there are four if I'm
- 1:52:57not getting this wrong myself so that
- 1:52:59didn't go extremely well let's look look
- 1:53:02at another kind of uh character level
- 1:53:04task so for example here I asked uh gp4
- 1:53:08to reverse the string default style and
- 1:53:11they tried to use a code interpreter and
- 1:53:13I stopped it and I said just do it just
- 1:53:15try it and uh it gave me jumble so it
- 1:53:19doesn't actually really know how to
- 1:53:21reverse this string going from right to
- 1:53:23left uh so it gave a wrong result so
- 1:53:26again like working with this working
- 1:53:28hypothesis that maybe this is due to the
- 1:53:30tokenization I tried a different
- 1:53:31approach I said okay let's reverse the
- 1:53:34exact same string but take the following
- 1:53:36approach step one just print out every
- 1:53:38single character separated by spaces and
- 1:53:40then as a step two reverse that list and
- 1:53:43it again Tred to use a tool but when I
- 1:53:44stopped it it uh first uh produced all
- 1:53:47the characters and that was actually
- 1:53:48correct and then It reversed them and
- 1:53:50that was correct once it had this so
- 1:53:53somehow it can't reverse it directly but
- 1:53:54when you go just first uh you know
- 1:53:57listing it out in order it can do that
- 1:53:59somehow and then it can once it's uh
- 1:54:01broken up this way this becomes all
- 1:54:03these individual characters and so now
- 1:54:06this is much easier for it to see these
- 1:54:07individual tokens and reverse them and
- 1:54:10print them out so that is kind of
- 1:54:13interesting so let's continue now why
- 1:54:16are llms worse at uh non-english langu
- 1:54:20and I briefly covered this already but
- 1:54:22basically um it's not only that the
- 1:54:24language model sees less non-english
- 1:54:27data during training of the model
- 1:54:28parameters but also the tokenizer is not
- 1:54:31um is not sufficiently trained on
- 1:54:34non-english data and so here for example
- 1:54:37hello how are you is five tokens and its
- 1:54:40translation is 15 tokens so this is a
- 1:54:42three times blow up and so for example
- 1:54:45anang is uh just hello basically in
- 1:54:48Korean and that end up being three
- 1:54:50tokens I'm actually kind of surprised by
- 1:54:51that because that is a very common
- 1:54:53phrase there just the typical greeting
- 1:54:55of like hello and that ends up being
- 1:54:57three tokens whereas our hello is a
- 1:54:58single token and so basically everything
- 1:55:00is a lot more bloated and diffuse and
- 1:55:02this is I think partly the reason that
- 1:55:04the model Works worse on other
- 1:55:07languages uh coming back why is LM bad
- 1:55:10at simple arithmetic um that has to do
- 1:55:13with the tokenization of numbers and so
- 1:55:17um you'll notice that for example
- 1:55:19addition is very sort of
- 1:55:20like uh there's an algorithm that is
- 1:55:23like character level for doing addition
- 1:55:25so for example here we would first add
- 1:55:27the ones and then the tens and then the
- 1:55:29hundreds you have to refer to specific
- 1:55:31parts of these digits but uh these
- 1:55:34numbers are represented completely
- 1:55:36arbitrarily based on whatever happened
- 1:55:37to merge or not merge during the
- 1:55:39tokenization process there's an entire
- 1:55:41blog post about this that I think is
- 1:55:42quite good integer tokenization is
- 1:55:44insane and this person basically
- 1:55:46systematically explores the tokenization
- 1:55:48of numbers in I believe this is gpt2 and
- 1:55:52so they notice that for example for the
- 1:55:53for um four-digit numbers you can take a
- 1:55:57look at whether it is uh a single token
- 1:56:00or whether it is two tokens that is a 1
- 1:56:02three or a 2 two or a 31 combination and
- 1:56:04so all the different numbers are all the
- 1:56:06different combinations and you can
- 1:56:08imagine this is all completely
- 1:56:09arbitrarily so and the model
- 1:56:11unfortunately sometimes sees uh four um
- 1:56:14a token for for all four digits
- 1:56:16sometimes for three sometimes for two
- 1:56:18sometimes for one and it's in an
- 1:56:20arbitrary uh Manner and so this is
- 1:56:22definitely a headwind if you will for
- 1:56:25the language model and it's kind of
- 1:56:26incredible that it can kind of do it and
- 1:56:27deal with it but it's also kind of not
- 1:56:30ideal and so that's why for example we
- 1:56:32saw that meta when they train the Llama
- 1:56:342 algorithm and they use sentence piece
- 1:56:36they make sure to split up all the um
- 1:56:39all the digits as an example for uh
- 1:56:42llama 2 and this is partly to improve a
- 1:56:44simple arithmetic kind of
- 1:56:46performance and finally why is gpt2 not
- 1:56:50as good in Python again this is partly a
- 1:56:52modeling issue on in the architecture
- 1:56:54and the data set and the strength of the
- 1:56:56model but it's also partially
- 1:56:58tokenization because as we saw here with
- 1:57:00the simple python example the encoding
- 1:57:03efficiency of the tokenizer for handling
- 1:57:05spaces in Python is terrible and every
- 1:57:07single space is an individual token and
- 1:57:09this dramatically reduces the context
- 1:57:11length that the model can attend to
- 1:57:12cross so that's almost like a
- 1:57:14tokenization bug for gpd2 and that was
- 1:57:16later fixed with gp4 okay so here's
- 1:57:20another fun one my llm abruptly halts
- 1:57:22when it sees the string end of text so
- 1:57:25here's um here's a very strange Behavior
- 1:57:28print a string end of text is what I
- 1:57:30told jt4 and it says could you please
- 1:57:32specify the string and I'm I'm telling
- 1:57:35it give me end of text and it seems like
- 1:57:37there's an issue it's not seeing end of
- 1:57:39text and then I give it end of text is
- 1:57:41the string and then here's a string and
- 1:57:44then it just doesn't print it so
- 1:57:45obviously something is breaking here
- 1:57:47with respect to the handling of the
- 1:57:48special token and I don't actually know
- 1:57:50what open ey is doing under the hood
- 1:57:52here and whether they are potentially
- 1:57:54parsing this as an um as an actual token
- 1:57:58instead of this just being uh end of
- 1:58:01text um as like individual sort of
- 1:58:04pieces of it without the special token
- 1:58:06handling logic and so it might be that
- 1:58:09someone when they're calling do encode
- 1:58:11uh they are passing in the allowed
- 1:58:13special and they are allowing end of
- 1:58:16text as a special character in the user
- 1:58:18prompt but the user prompt of course is
- 1:58:20is a sort of um attacker controlled text
- 1:58:23so you would hope that they don't really
- 1:58:25parse or use special tokens or you know
- 1:58:28from that kind of input but it appears
- 1:58:30that there's something definitely going
- 1:58:31wrong here and um so your knowledge of
- 1:58:34these special tokens ends up being in a
- 1:58:36tax surface potentially and so if you'd
- 1:58:38like to confuse llms then just um try to
- 1:58:43give them some special tokens and see if
- 1:58:44you're breaking something by chance okay
- 1:58:46so this next one is a really fun one uh
- 1:58:49the trailing whites space issue so if
- 1:58:52you come to playground and uh we come
- 1:58:56here to GPT 3.5 turbo instruct so this
- 1:58:58is not a chat model this is a completion
- 1:59:00model so think of it more like it's a
- 1:59:02lot more closer to a base model it does
- 1:59:05completion it will continue the token
- 1:59:07sequence so here's a tagline for ice
- 1:59:09cream shop and we want to continue the
- 1:59:11sequence and so we can submit and get a
- 1:59:14bunch of tokens okay no problem but now
- 1:59:18suppose I do this but instead of
- 1:59:20pressing submit here I do here's a
- 1:59:23tagline for ice cream shop space so I
- 1:59:26have a space here before I click
- 1:59:28submit we get a warning your text ends
- 1:59:31in a trail Ling space which causes worse
- 1:59:33performance due to how API splits text
- 1:59:35into tokens so what's happening here it
- 1:59:38still gave us a uh sort of completion
- 1:59:40here but let's take a look at what's
- 1:59:42happening so here's a tagline for an ice
- 1:59:44cream shop and then what does this look
- 1:59:48like in the actual actual training data
- 1:59:50suppose you found the completion in the
- 1:59:52training document somewhere on the
- 1:59:53internet and the llm trained on this
- 1:59:55data so maybe it's something like oh
- 1:59:58yeah maybe that's the tagline that's a
- 2:00:00terrible tagline but notice here that
- 2:00:02when I create o you see that because
- 2:00:05there's the the space character is
- 2:00:07always a prefix to these tokens in GPT
- 2:00:11so it's not an O token it's a space o
- 2:00:13token the space is part of the O and
- 2:00:16together they are token 8840 that's
- 2:00:19that's space o so what's What's
- 2:00:21Happening Here is that when I just have
- 2:00:24it like this and I let it complete the
- 2:00:27next token it can sample the space o
- 2:00:30token but instead if I have this and I
- 2:00:32add my space then what I'm doing here
- 2:00:34when I incode this string is I have
- 2:00:37basically here's a t line for an ice
- 2:00:39cream uh shop and this space at the very
- 2:00:42end becomes a token
- 2:00:44220 and so we've added token 220 and
- 2:00:47this token otherwise would be part of
- 2:00:49the tagline because if there actually is
- 2:00:51a tagline here so space o is the token
- 2:00:55and so this is suddenly a of
- 2:00:57distribution for the model because this
- 2:00:59space is part of the next token but
- 2:01:01we're putting it here like this and the
- 2:01:04model has seen very very little data of
- 2:01:07actual Space by itself and we're asking
- 2:01:10it to complete the sequence like add in
- 2:01:11more tokens but the problem is that
- 2:01:13we've sort of begun the first token and
- 2:01:16now it's been split up and now we're out
- 2:01:18of this distribution and now arbitrary
- 2:01:20bad things happen and it's just a very
- 2:01:23rare example for it to see something
- 2:01:24like that and uh that's why we get the
- 2:01:26warning so the fundamental issue here is
- 2:01:29of course that um the llm is on top of
- 2:01:32these tokens and these tokens are text
- 2:01:34chunks they're not characters in a way
- 2:01:36you and I would think of them they are
- 2:01:38these are the atoms of what the LM is
- 2:01:40seeing and there's a bunch of weird
- 2:01:41stuff that comes out of it let's go back
- 2:01:43to our default cell style I bet you that
- 2:01:48the model has never in its training set
- 2:01:49seen default cell sta without Le in
- 2:01:54there it's always seen this as a single
- 2:01:56group because uh this is some kind of a
- 2:01:59function in um I'm guess I don't
- 2:02:02actually know what this is part of this
- 2:02:03is some kind of API but I bet you that
- 2:02:05it's never seen this combination of
- 2:02:07tokens uh in its training data because
- 2:02:10or I think it would be extremely rare so
- 2:02:12I took this and I copy pasted it here
- 2:02:14and I had I tried to complete from it
- 2:02:17and the it immediately gave me a big
- 2:02:19error and it said the model predicted to
- 2:02:21completion that begins with a stop
- 2:02:22sequence resulting in no output consider
- 2:02:24adjusting your prompt or stop sequences
- 2:02:26so what happened here when I clicked
- 2:02:27submit is that immediately the model
- 2:02:30emitted and sort of like end of text
- 2:02:32token I think or something like that it
- 2:02:34basically predicted the stop sequence
- 2:02:36immediately so it had no completion and
- 2:02:38so this is why I'm getting a warning
- 2:02:40again because we're off the data
- 2:02:42distribution and the model is just uh
- 2:02:45predicting just totally arbitrary things
- 2:02:47it's just really confused basically this
- 2:02:49is uh this is giving it brain damage
- 2:02:50it's never seen this before it's shocked
- 2:02:53and it's predicting end of text or
- 2:02:54something I tried it again here and it
- 2:02:57in this case it completed it but then
- 2:02:59for some reason this request May violate
- 2:03:01our usage policies this was
- 2:03:03flagged um basically something just like
- 2:03:06goes wrong and there's something like
- 2:03:07Jank you can just feel the Jank because
- 2:03:09the model is like extremely unhappy with
- 2:03:11just this and it doesn't know how to
- 2:03:12complete it because it's never occurred
- 2:03:14in training set in a training set it
- 2:03:16always appears like this and becomes a
- 2:03:18single token
- 2:03:20so these kinds of issues where tokens
- 2:03:21are either you sort of like complete the
- 2:03:24first character of the next token or you
- 2:03:26are sort of you have long tokens that
- 2:03:28you then have just some of the
- 2:03:29characters off all of these are kind of
- 2:03:32like issues with partial tokens is how I
- 2:03:35would describe it and if you actually
- 2:03:37dig into the T token
- 2:03:39repository go to the rust code and
- 2:03:41search for
- 2:03:44unstable and you'll see um en code
- 2:03:47unstable native unstable token tokens
- 2:03:49and a lot of like special case handling
- 2:03:51none of this stuff about unstable tokens
- 2:03:53is documented anywhere but there's a ton
- 2:03:55of code dealing with unstable tokens and
- 2:03:58unstable tokens is exactly kind of like
- 2:04:00what I'm describing here what you would
- 2:04:02like out of a completion API is
- 2:04:05something a lot more fancy like if we're
- 2:04:06putting in default cell sta if we're
- 2:04:08asking for the next token sequence we're
- 2:04:10not actually trying to append the next
- 2:04:12token exactly after this list we're
- 2:04:14actually trying to append we're trying
- 2:04:16to consider lots of tokens um
- 2:04:19that if we were or I guess like we're
- 2:04:22trying to search over characters that if
- 2:04:25we retened would be of high probability
- 2:04:28if that makes sense um so that we can
- 2:04:30actually add a single individual
- 2:04:32character uh instead of just like adding
- 2:04:34the next full token that comes after
- 2:04:36this partial token list so I this is
- 2:04:39very tricky to describe and I invite you
- 2:04:41to maybe like look through this it ends
- 2:04:43up being extremely gnarly and hairy kind
- 2:04:44of topic it and it comes from
- 2:04:46tokenization fundamentally so um maybe I
- 2:04:49can even spend an entire video talking
- 2:04:50about unstable tokens sometime in the
- 2:04:52future okay and I'm really saving the
- 2:04:54best for last my favorite one by far is
- 2:04:56the solid gold
- 2:04:59Magikarp and it just okay so this comes
- 2:05:01from this blog post uh solid gold
- 2:05:03Magikarp and uh this is um internet
- 2:05:07famous now for those of us in llms and
- 2:05:10basically I I would advise you to uh
- 2:05:11read this block Post in full but
- 2:05:13basically what this person was doing is
- 2:05:16this person went to the um
- 2:05:19token embedding stable and clustered the
- 2:05:22tokens based on their embedding
- 2:05:24representation and this person noticed
- 2:05:27that there's a cluster of tokens that
- 2:05:29look really strange so there's a cluster
- 2:05:31here at rot e stream Fame solid gold
- 2:05:34Magikarp Signet message like really
- 2:05:36weird tokens in uh basically in this
- 2:05:39embedding cluster and so what are these
- 2:05:42tokens and where do they even come from
- 2:05:43like what is solid gold magikarpet makes
- 2:05:45no sense and then they found bunch of
- 2:05:48these
- 2:05:50tokens and then they notice that
- 2:05:52actually the plot thickens here because
- 2:05:53if you ask the model about these tokens
- 2:05:56like you ask it uh some very benign
- 2:05:58question like please can you repeat back
- 2:06:00to me the string sold gold Magikarp uh
- 2:06:02then you get a variety of basically
- 2:06:04totally broken llm Behavior so either
- 2:06:07you get evasion so I'm sorry I can't
- 2:06:09hear you or you get a bunch of
- 2:06:11hallucinations as a response um you can
- 2:06:14even get back like insults so you ask it
- 2:06:17uh about streamer bot it uh tells the
- 2:06:20and the model actually just calls you
- 2:06:22names uh or it kind of comes up with
- 2:06:24like weird humor like you're actually
- 2:06:26breaking the model by asking about these
- 2:06:28very simple strings like at Roth and
- 2:06:30sold gold Magikarp so like what the hell
- 2:06:32is happening and there's a variety of
- 2:06:34here documented behaviors uh there's a
- 2:06:37bunch of tokens not just so good
- 2:06:38Magikarp that have that kind of a
- 2:06:40behavior and so basically there's a
- 2:06:42bunch of like trigger words and if you
- 2:06:44ask the model about these trigger words
- 2:06:46or you just include them in your prompt
- 2:06:48the model goes haywire and has all kinds
- 2:06:50of uh really Strange Behaviors including
- 2:06:52sort of ones that violate typical safety
- 2:06:54guidelines uh and the alignment of the
- 2:06:57model like it's swearing back at you so
- 2:06:59what is happening here and how can this
- 2:07:01possibly be true well this again comes
- 2:07:04down to tokenization so what's happening
- 2:07:06here is that sold gold Magikarp if you
- 2:07:08actually dig into it is a Reddit user so
- 2:07:11there's a u Sol gold
- 2:07:14Magikarp and probably what happened here
- 2:07:16even though I I don't know that this has
- 2:07:18been like really definitively explored
- 2:07:20but what is thought to have happened is
- 2:07:23that the tokenization data set was very
- 2:07:25different from the training data set for
- 2:07:28the actual language model so in the
- 2:07:29tokenization data set there was a ton of
- 2:07:31redded data potentially where the user
- 2:07:34solid gold Magikarp was mentioned in the
- 2:07:36text because solid gold Magikarp was a
- 2:07:39very common um sort of uh person who
- 2:07:41would post a lot uh this would be a
- 2:07:43string that occurs many times in a
- 2:07:45tokenization data set because it occurs
- 2:07:48many times in a tokenization data set
- 2:07:50these tokens would end up getting merged
- 2:07:51to the single individual token for that
- 2:07:53single Reddit user sold gold Magikarp so
- 2:07:56they would have a dedicated token in a
- 2:07:58vocabulary of was it 50,000 tokens in
- 2:08:00gpd2 that is devoted to that Reddit user
- 2:08:04and then what happens is the
- 2:08:05tokenization data set has those strings
- 2:08:08but then later when you train the model
- 2:08:10the language model itself um this data
- 2:08:13from Reddit was not present and so
- 2:08:16therefore in the entire training set for
- 2:08:18the language model sold gold Magikarp
- 2:08:21never occurs that token never appears in
- 2:08:24the training set for the actual language
- 2:08:25model later so this token never gets
- 2:08:28activated it's initialized at random in
- 2:08:31the beginning of optimization then you
- 2:08:32have forward backward passes and updates
- 2:08:34to the model and this token is just
- 2:08:36never updated in the embedding table
- 2:08:37that row Vector never gets sampled it
- 2:08:40never gets used so it never gets trained
- 2:08:42and it's completely untrained it's kind
- 2:08:43of like unallocated memory in a typical
- 2:08:46binary program written in C or something
- 2:08:48like that that so it's unallocated
- 2:08:50memory and then at test time if you
- 2:08:51evoke this token then you're basically
- 2:08:54plucking out a row of the embedding
- 2:08:55table that is completely untrained and
- 2:08:57that feeds into a Transformer and
- 2:08:58creates undefined behavior and that's
- 2:09:00what we're seeing here this completely
- 2:09:02undefined never before seen in a
- 2:09:03training behavior and so any of these
- 2:09:06kind of like weird tokens would evoke
- 2:09:08this Behavior because fundamentally the
- 2:09:09model is um is uh uh out of sample out
- 2:09:14of distribution okay and the very last
- 2:09:16thing I wanted to just briefly mention
- 2:09:18point out although I think a lot of
- 2:09:19people are quite aware of this is that
- 2:09:21different kinds of formats and different
- 2:09:23representations and different languages
- 2:09:25and so on might be more or less
- 2:09:26efficient with GPD tokenizers uh or any
- 2:09:29tokenizers for any other L for that
- 2:09:31matter so for example Json is actually
- 2:09:33really dense in tokens and yaml is a lot
- 2:09:36more efficient in tokens um so for
- 2:09:39example this are these are the same in
- 2:09:41Json and in yaml the Json is
- 2:09:44116 and the yaml is 99 so quite a bit of
- 2:09:48an Improvement and so in the token
- 2:09:51economy where we are paying uh per token
- 2:09:53in many ways and you are paying in the
- 2:09:55context length and you're paying in um
- 2:09:57dollar amount for uh the cost of
- 2:09:59processing all this kind of structured
- 2:10:01data when you have to um so prefer to
- 2:10:03use theal over Json and in general kind
- 2:10:06of like the tokenization density is
- 2:10:07something that you have to um sort of
- 2:10:09care about and worry about at all times
- 2:10:11and try to find efficient encoding
- 2:10:13schemes and spend a lot of time in tick
- 2:10:15tokenizer and measure the different
- 2:10:16token efficiencies of different formats
- 2:10:18and settings and so on okay so that
- 2:10:21concludes my fairly long video on
- 2:10:23tokenization I know it's a try I know
- 2:10:25it's annoying I know it's irritating I
- 2:10:28personally really dislike the stage what
- 2:10:30I do have to say at this point is don't
- 2:10:32brush it off there's a lot of foot guns
- 2:10:34sharp edges here security issues uh AI
- 2:10:38safety issues as we saw plugging in
- 2:10:39unallocated memory into uh language
- 2:10:42models so um it's worth understanding
- 2:10:45this stage um that said I will say that
- 2:10:48eternal glory goes to anyone who can get
- 2:10:50rid of it uh I showed you one possible
- 2:10:52paper that tried to uh do that and I
- 2:10:54think I hope a lot more can follow over
- 2:10:57time and my final recommendations for
- 2:10:59the application right now are if you can
- 2:11:01reuse the GPT 4 tokens and the
- 2:11:03vocabulary uh in your application then
- 2:11:05that's something you should consider and
- 2:11:06just use Tech token because it is very
- 2:11:07efficient and nice library for inference
- 2:11:11for bpe I also really like the bite
- 2:11:13level BP that uh Tik toen and openi uses
- 2:11:17uh if you for some reason want to train
- 2:11:19your own vocabulary from scratch um then
- 2:11:22I would use uh the bpe with sentence
- 2:11:25piece um oops as I mentioned I'm not a
- 2:11:28huge fan of sentence piece I don't like
- 2:11:30its uh bite fallback and I don't like
- 2:11:33that it's doing BP on unic code code
- 2:11:35points I think it's uh it also has like
- 2:11:37a million settings and I think there's a
- 2:11:39lot of foot gonss here and I think it's
- 2:11:40really easy to Mis calibrate them and
- 2:11:42you end up cropping your sentences or
- 2:11:43something like that uh because of some
- 2:11:45type of parameter that you don't fully
- 2:11:47understand so so be very careful with
- 2:11:49the settings try to copy paste exactly
- 2:11:51maybe where what meta did or basically
- 2:11:54spend a lot of time looking at all the
- 2:11:56hyper parameters and go through the code
- 2:11:57of sentence piece and make sure that you
- 2:11:59have this correct um but even if you
- 2:12:02have all the settings correct I still
- 2:12:03think that the algorithm is kind of
- 2:12:04inferior to what's happening here and
- 2:12:07maybe the best if you really need to
- 2:12:09train your vocabulary maybe the best
- 2:12:11thing is to just wait for M bpe to
- 2:12:13becomes as efficient as possible and uh
- 2:12:16that's something that maybe I hope to
- 2:12:18work on and at some point maybe we can
- 2:12:20be training basically really what we
- 2:12:22want is we want tick token but training
- 2:12:24code and that is the ideal thing that
- 2:12:27currently does not exist and MBP is um
- 2:12:31is in implementation of it but currently
- 2:12:33it's in Python so that's currently what
- 2:12:35I have to say for uh tokenization there
- 2:12:38might be an advanced video that has even
- 2:12:40drier and even more detailed in the
- 2:12:41future but for now I think we're going
- 2:12:43to leave things off here and uh I hope
- 2:12:46that was helpful bye
- 2:12:54and uh they increase this contact size
- 2:12:56from gpt1 of 512 uh to 1024 and GPT 4
- 2:13:02two the
- 2:13:05next okay next I would like us to
- 2:13:07briefly walk through the code from open
- 2:13:09AI on the gpt2 encoded
- 2:13:15ATP I'm sorry I'm gonna sneeze
- 2:13:19and then what's Happening Here
- 2:13:21is this is a spous layer that I will
- 2:13:24explain in a
- 2:13:26bit What's Happening Here
- 2:13:33is
About this transcript
This page contains the full transcript of Let's build the GPT Tokenizer by Andrej Karpathy, generated from the public captions YouTube serves with the video. The transcript has 24,679 words across 3,422 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.