YouTube2Text

Let's build the GPT Tokenizer — Transcript

by Andrej Karpathy · 24,679 words · 3,422 segments · language en · Watch on YouTube

Full transcript

  1. 0:00hi everyone so in this video I'd like us
  2. 0:02to cover the process of tokenization in
  3. 0:04large language models now you see here
  4. 0:06that I have a set face and that's
  5. 0:08because uh tokenization is my least
  6. 0:10favorite part of working with large
  7. 0:11language models but unfortunately it is
  8. 0:13necessary to understand in some detail
  9. 0:15because it it is fairly hairy gnarly and
  10. 0:17there's a lot of hidden foot guns to be
  11. 0:19aware of and a lot of oddness with large
  12. 0:21language models typically traces back to
  13. 0:24tokenization so what is
  14. 0:26tokenization now in my previous video
  15. 0:28Let's Build GPT from scratch uh we
  16. 0:31actually already did tokenization but we
  17. 0:33did a very naive simple version of
  18. 0:35tokenization so when you go to the
  19. 0:37Google colab for that video uh you see
  20. 0:40here that we loaded our training set and
  21. 0:43our training set was this uh Shakespeare
  22. 0:45uh data set now in the beginning the
  23. 0:48Shakespeare data set is just a large
  24. 0:49string in Python it's just text and so
  25. 0:52the question is how do we plug text into
  26. 0:54large language models and in this case
  27. 0:58here we created a vocabulary of 65
  28. 1:01possible characters that we saw occur in
  29. 1:03this string these were the possible
  30. 1:05characters and we saw that there are 65
  31. 1:07of them and then we created a a lookup
  32. 1:10table for converting from every possible
  33. 1:13character a little string piece into a
  34. 1:16token an
  35. 1:17integer so here for example we tokenized
  36. 1:20the string High there and we received
  37. 1:23this sequence of
  38. 1:24tokens and here we took the first 1,000
  39. 1:27characters of our data set and we
  40. 1:29encoded it into tokens and because it is
  41. 1:32this is character level we received
  42. 1:341,000 tokens in a sequence so token 18
  43. 1:3847
  44. 1:40Etc now later we saw that the way we
  45. 1:43plug these tokens into the language
  46. 1:45model is by using an embedding
  47. 1:48table and so basically if we have 65
  48. 1:51possible tokens then this embedding
  49. 1:53table is going to have 65 rows and
  50. 1:56roughly speaking we're taking the
  51. 1:58integer associated with every single
  52. 1:59sing Le token we're using that as a
  53. 2:01lookup into this table and we're
  54. 2:04plucking out the corresponding row and
  55. 2:06this row is a uh is trainable parameters
  56. 2:09that we're going to train using back
  57. 2:10propagation and this is the vector that
  58. 2:12then feeds into the Transformer um and
  59. 2:15that's how the Transformer Ser of
  60. 2:16perceives every single
  61. 2:18token so here we had a very naive
  62. 2:21tokenization process that was a
  63. 2:23character level tokenizer but in
  64. 2:25practice in state-ofthe-art uh language
  65. 2:27models people use a lot more complicated
  66. 2:28schemes unfortunately
  67. 2:30uh for constructing these uh token
  68. 2:34vocabularies so we're not dealing on the
  69. 2:36Character level we're dealing on chunk
  70. 2:38level and the way these um character
  71. 2:41chunks are constructed is using
  72. 2:43algorithms such as for example the bik
  73. 2:45pair in coding algorithm which we're
  74. 2:46going to go into in detail um and cover
  75. 2:51in this video I'd like to briefly show
  76. 2:52you the paper that introduced a bite
  77. 2:54level encoding as a mechanism for
  78. 2:56tokenization in the context of large
  79. 2:58language models and I would say that
  80. 3:00that's probably the gpt2 paper and if
  81. 3:02you scroll down here to the section
  82. 3:05input representation this is where they
  83. 3:07cover tokenization the kinds of
  84. 3:09properties that you'd like the
  85. 3:10tokenization to have and they conclude
  86. 3:13here that they're going to have a
  87. 3:14tokenizer where you have a vocabulary of
  88. 3:1750,2 57 possible
  89. 3:20tokens and the context size is going to
  90. 3:24be 1,24 tokens so in the in in the
  91. 3:27attention layer of the Transformer
  92. 3:29neural network
  93. 3:30every single token is attending to the
  94. 3:32previous tokens in the sequence and it's
  95. 3:34going to see up to 1,24 tokens so tokens
  96. 3:37are this like fundamental unit um the
  97. 3:40atom of uh large language models if you
  98. 3:43will and everything is in units of
  99. 3:44tokens everything is about tokens and
  100. 3:47tokenization is the process for
  101. 3:48translating strings or text into
  102. 3:51sequences of tokens and uh vice versa
  103. 3:54when you go into the Llama 2 paper as
  104. 3:56well I can show you that when you search
  105. 3:58token you're going to get get 63 hits um
  106. 4:01and that's because tokens are again
  107. 4:03pervasive so here they mentioned that
  108. 4:05they trained on two trillion tokens of
  109. 4:06data and so
  110. 4:08on so we're going to build our own
  111. 4:11tokenizer luckily the bite be encoding
  112. 4:13algorithm is not uh that super
  113. 4:15complicated and we can build it from
  114. 4:16scratch ourselves and we'll see exactly
  115. 4:18how this works before we dive into code
  116. 4:20I'd like to give you a brief Taste of
  117. 4:22some of the complexities that come from
  118. 4:24the tokenization because I just want to
  119. 4:26make sure that we motivate it
  120. 4:27sufficiently for why we are doing all
  121. 4:29this and why this is so gross so
  122. 4:32tokenization is at the heart of a lot of
  123. 4:34weirdness in large language models and I
  124. 4:36would advise that you do not brush it
  125. 4:37off a lot of the issues that may look
  126. 4:40like just issues with the new network
  127. 4:42architecture or the large language model
  128. 4:44itself are actually issues with the
  129. 4:46tokenization and fundamentally Trace uh
  130. 4:49back to it so if you've noticed any
  131. 4:51issues with large language models can't
  132. 4:54you know not able to do spelling tasks
  133. 4:56very easily that's usually due to
  134. 4:57tokenization simple string processing
  135. 5:00can be difficult for the large language
  136. 5:02model to perform
  137. 5:03natively uh non-english languages can
  138. 5:06work much worse and to a large extent
  139. 5:08this is due to
  140. 5:09tokenization sometimes llms are bad at
  141. 5:11simple arithmetic also can trace be
  142. 5:14traced to
  143. 5:15tokenization uh gbt2 specifically would
  144. 5:17have had quite a bit more issues with
  145. 5:19python than uh future versions of it due
  146. 5:22to tokenization there's a lot of other
  147. 5:24issues maybe you've seen weird warnings
  148. 5:25about a trailing whites space this is a
  149. 5:27tokenization issue um
  150. 5:30if you had asked GPT earlier about solid
  151. 5:33gold Magikarp and what it is you would
  152. 5:35see the llm go totally crazy and it
  153. 5:37would start going off about a completely
  154. 5:39unrelated tangent topic maybe you've
  155. 5:41been told to use yl over Json in
  156. 5:43structure data all of that has to do
  157. 5:45with tokenization so basically
  158. 5:47tokenization is at the heart of many
  159. 5:49issues I will look back around to these
  160. 5:51at the end of the video but for now let
  161. 5:54me just um skip over it a little bit and
  162. 5:56let's go to this web app um the Tik
  163. 5:59tokenizer bell.app so I have it loaded
  164. 6:02here and what I like about this web app
  165. 6:04is that tokenization is running a sort
  166. 6:06of live in your browser in JavaScript so
  167. 6:09you can just type here stuff hello world
  168. 6:11and the whole string
  169. 6:14rokenes so here what we see on uh the
  170. 6:18left is a string that you put in on the
  171. 6:20right we're currently using the gpt2
  172. 6:22tokenizer we see that this string that I
  173. 6:24pasted here is currently tokenizing into
  174. 6:27300 tokens and here they are sort of uh
  175. 6:30shown explicitly in different colors for
  176. 6:32every single token so for example uh
  177. 6:35this word tokenization became two tokens
  178. 6:38the token
  179. 6:403,642 and
  180. 6:441,634 the token um space is is token 318
  181. 6:50so be careful on the bottom you can show
  182. 6:51white space and keep in mind that there
  183. 6:54are spaces and uh sln new line
  184. 6:57characters in here but you can hide them
  185. 6:59for
  186. 7:01clarity the token space at is token 379
  187. 7:06the to the Token space the is 262 Etc so
  188. 7:11you notice here that the space is part
  189. 7:12of that uh token
  190. 7:15chunk now so this is kind of like how
  191. 7:18our English sentence broke up and that
  192. 7:21seems all well and good now now here I
  193. 7:24put in some arithmetic so we see that uh
  194. 7:26the token 127 Plus and then token six
  195. 7:31space 6 followed by 77 so what's
  196. 7:34happening here is that 127 is feeding in
  197. 7:36as a single token into the large
  198. 7:38language model but the um number 677
  199. 7:42will actually feed in as two separate
  200. 7:44tokens and so the large language model
  201. 7:47has to sort of um take account of that
  202. 7:50and process it correctly in its Network
  203. 7:53and see here 804 will be broken up into
  204. 7:56two tokens and it's is all completely
  205. 7:57arbitrary and here I have another
  206. 7:59example of four-digit numbers and they
  207. 8:02break up in a way that they break up and
  208. 8:03it's totally arbitrary sometimes you
  209. 8:05have um multiple digits single token
  210. 8:08sometimes you have individual digits as
  211. 8:10many tokens and it's all kind of pretty
  212. 8:12arbitrary and coming out of the
  213. 8:14tokenizer here's another example we have
  214. 8:17the string egg and you see here that
  215. 8:21this became two
  216. 8:22tokens but for some reason when I say I
  217. 8:24have an egg you see when it's a space
  218. 8:27egg it's two token it's sorry it's a
  219. 8:30single token so just egg by itself in
  220. 8:33the beginning of a sentence is two
  221. 8:34tokens but here as a space egg is
  222. 8:37suddenly a single token uh for the exact
  223. 8:40same string okay here lowercase egg
  224. 8:44turns out to be a single token and in
  225. 8:46particular notice that the color is
  226. 8:47different so this is a different token
  227. 8:49so this is case sensitive and of course
  228. 8:51a capital egg would also be different
  229. 8:54tokens and again um this would be two
  230. 8:57tokens arbitrarily so so for the same
  231. 9:00concept egg depending on if it's in the
  232. 9:02beginning of a sentence at the end of a
  233. 9:03sentence lowercase uppercase or mixed
  234. 9:06all this will be uh basically very
  235. 9:08different tokens and different IDs and
  236. 9:10the language model has to learn from raw
  237. 9:12data from all the internet text that
  238. 9:13it's going to be training on that these
  239. 9:15are actually all the exact same concept
  240. 9:17and it has to sort of group them in the
  241. 9:19parameters of the neural network and
  242. 9:21understand just based on the data
  243. 9:22patterns that these are all very similar
  244. 9:24but maybe not almost exactly similar but
  245. 9:27but very very similar
  246. 9:30um after the EG demonstration here I
  247. 9:32have um an introduction from open a eyes
  248. 9:35chbt in Korean so manaso Pang uh Etc uh
  249. 9:41so this is in Korean and the reason I
  250. 9:44put this here is because you'll notice
  251. 9:47that um non-english languages work
  252. 9:51slightly worse in Chachi part of this is
  253. 9:54because of course the training data set
  254. 9:55for Chachi is much larger for English
  255. 9:58and for everything else but the same is
  256. 9:59true not just for the large language
  257. 10:01model itself but also for the tokenizer
  258. 10:04so when we train the tokenizer we're
  259. 10:05going to see that there's a training set
  260. 10:07as well and there's a lot more English
  261. 10:09than non-english and what ends up
  262. 10:11happening is that we're going to have a
  263. 10:13lot more longer tokens for
  264. 10:16English so how do I put this if you have
  265. 10:19a single sentence in English and you
  266. 10:21tokenize it you might see that it's 10
  267. 10:23tokens or something like that but if you
  268. 10:25translate that sentence into say Korean
  269. 10:27or Japanese or something else you'll
  270. 10:29typically see that the number of tokens
  271. 10:30used is much larger and that's because
  272. 10:33the chunks here are a lot more broken up
  273. 10:36so we're using a lot more tokens for the
  274. 10:38exact same thing and what this does is
  275. 10:41it bloats up the sequence length of all
  276. 10:43the documents so you're using up more
  277. 10:46tokens and then in the attention of the
  278. 10:48Transformer when these tokens try to
  279. 10:49attend each other you are running out of
  280. 10:51context um in the maximum context length
  281. 10:55of that Transformer and so basically all
  282. 10:57the non-english text is stretched out
  283. 11:01from the perspective of the Transformer
  284. 11:03and this just has to do with the um
  285. 11:05trainings that used for the tokenizer
  286. 11:07and the tokenization itself so it will
  287. 11:10create a lot bigger tokens and a lot
  288. 11:12larger groups in English and it will
  289. 11:14have a lot of little boundaries for all
  290. 11:16the other non-english text um so if we
  291. 11:19translated this into English it would be
  292. 11:21significantly fewer
  293. 11:23tokens the final example I have here is
  294. 11:25a little snippet of python for doing FS
  295. 11:28buuz and what I'd like you to notice is
  296. 11:31look all these individual spaces are all
  297. 11:34separate tokens they are token
  298. 11:37220 so uh 220 220 220 220 and then space
  299. 11:42if is a single token and so what's going
  300. 11:45on here is that when the Transformer is
  301. 11:46going to consume or try to uh create
  302. 11:49this text it needs to um handle all
  303. 11:52these spaces individually they all feed
  304. 11:54in one by one into the entire
  305. 11:56Transformer in the sequence and so this
  306. 11:59is being extremely wasteful tokenizing
  307. 12:01it in this way and so as a result of
  308. 12:04that gpt2 is not very good with python
  309. 12:07and it's not anything to do with coding
  310. 12:08or the language model itself it's just
  311. 12:10that if he use a lot of indentation
  312. 12:12using space in Python like we usually do
  313. 12:15uh you just end up bloating out all the
  314. 12:17text and it's separated across way too
  315. 12:19much of the sequence and we are running
  316. 12:21out of the context length in the
  317. 12:22sequence uh that's roughly speaking
  318. 12:24what's what's happening we're being way
  319. 12:25too wasteful we're taking up way too
  320. 12:27much token space now we can also scroll
  321. 12:29up here and we can change the tokenizer
  322. 12:31so note here that gpt2 tokenizer creates
  323. 12:34a token count of 300 for this string
  324. 12:36here we can change it to CL 100K base
  325. 12:39which is the GPT for tokenizer and we
  326. 12:41see that the token count drops to 185 so
  327. 12:44for the exact same string we are now
  328. 12:46roughly having the number of tokens and
  329. 12:49roughly speaking this is because uh the
  330. 12:51number of tokens in the GPT 4 tokenizer
  331. 12:54is roughly double that of the number of
  332. 12:56tokens in the gpt2 tokenizer so we went
  333. 12:58went from roughly 50k to roughly 100K
  334. 13:01now you can imagine that this is a good
  335. 13:03thing because the same text is now
  336. 13:06squished into half as many tokens so uh
  337. 13:10this is a lot denser input to the
  338. 13:12Transformer and in the Transformer every
  339. 13:15single token has a finite number of
  340. 13:17tokens before it that it's going to pay
  341. 13:18attention to and so what this is doing
  342. 13:20is we're roughly able to see twice as
  343. 13:23much text as a context for what token to
  344. 13:26predict next uh because of this change
  345. 13:29but of course just increasing the number
  346. 13:30of tokens is uh not strictly better
  347. 13:33infinitely uh because as you increase
  348. 13:35the number of tokens now your embedding
  349. 13:36table is um sort of getting a lot larger
  350. 13:39and also at the output we are trying to
  351. 13:41predict the next token and there's the
  352. 13:42soft Max there and that grows as well
  353. 13:45we're going to go into more detail later
  354. 13:46on this but there's some kind of a Sweet
  355. 13:48Spot somewhere where you have a just
  356. 13:51right number of tokens in your
  357. 13:52vocabulary where everything is
  358. 13:53appropriately dense and still fairly
  359. 13:56efficient now one thing I would like you
  360. 13:58to note specifically for the gp4
  361. 14:00tokenizer is that the handling of the
  362. 14:03white space for python has improved a
  363. 14:05lot you see that here these four spaces
  364. 14:08are represented as one single token for
  365. 14:10the three spaces here and then the token
  366. 14:13SPF and here seven spaces were all
  367. 14:16grouped into a single token so we're
  368. 14:18being a lot more efficient in how we
  369. 14:20represent Python and this was a
  370. 14:21deliberate Choice made by open aai when
  371. 14:23they designed the gp4 tokenizer and they
  372. 14:27group a lot more space into a single
  373. 14:29character what this does is this
  374. 14:32densifies Python and therefore we can
  375. 14:35attend to more code before it when we're
  376. 14:38trying to predict the next token in the
  377. 14:39sequence and so the Improvement in the
  378. 14:42python coding ability from gbt2 to gp4
  379. 14:45is not just a matter of the language
  380. 14:47model and the architecture and the
  381. 14:48details of the optimization but a lot of
  382. 14:50the Improvement here is also coming from
  383. 14:52the design of the tokenizer and how it
  384. 14:54groups characters into tokens okay so
  385. 14:56let's now start writing some code
  386. 14:59so remember what we want to do we want
  387. 15:01to take strings and feed them into
  388. 15:03language models for that we need to
  389. 15:05somehow tokenize strings into some
  390. 15:08integers in some fixed vocabulary and
  391. 15:12then we will use those integers to make
  392. 15:14a look up into a lookup table of vectors
  393. 15:16and feed those vectors into the
  394. 15:18Transformer as an input now the reason
  395. 15:21this gets a little bit tricky of course
  396. 15:22is that we don't just want to support
  397. 15:24the simple English alphabet we want to
  398. 15:26support different kinds of languages so
  399. 15:28this is anango in Korean which is hello
  400. 15:31and we also want to support many kinds
  401. 15:33of special characters that we might find
  402. 15:34on the internet for example
  403. 15:37Emoji so how do we feed this text into
  404. 15:41uh
  405. 15:42Transformers well how's the what is this
  406. 15:44text anyway in Python so if you go to
  407. 15:46the documentation of a string in Python
  408. 15:49you can see that strings are immutable
  409. 15:51sequences of Unicode code
  410. 15:54points okay what are Unicode code points
  411. 15:57we can go to PDF so Unicode code points
  412. 16:01are defined by the Unicode Consortium as
  413. 16:04part of the Unicode standard and what
  414. 16:07this is really is that it's just a
  415. 16:09definition of roughly 150,000 characters
  416. 16:11right now and roughly speaking what they
  417. 16:14look like and what integers um represent
  418. 16:17those characters so it says 150,000
  419. 16:19characters across 161 scripts as of
  420. 16:22right now so if you scroll down here you
  421. 16:24can see that the standard is very much
  422. 16:26alive the latest standard 15.1 in
  423. 16:28September
  424. 16:302023 and basically this is just a way to
  425. 16:33define lots of types of
  426. 16:36characters like for example all these
  427. 16:39characters across different scripts so
  428. 16:41the way we can access the unic code code
  429. 16:44Point given Single Character is by using
  430. 16:45the or function in Python so for example
  431. 16:48I can pass in Ord of H and I can see
  432. 16:51that for the Single Character H the unic
  433. 16:54code code point is
  434. 16:56104 okay um but this can be arbitr
  435. 17:00complicated so we can take for example
  436. 17:02our Emoji here and we can see that the
  437. 17:04code point for this one is
  438. 17:06128,000 or we can take
  439. 17:10un and this is 50,000 now keep in mind
  440. 17:13you can't plug in strings here because
  441. 17:16you uh this doesn't have a single code
  442. 17:18point it only takes a single uni code
  443. 17:20code Point character and tells you its
  444. 17:23integer so in this way we can look
  445. 17:26up all the um characters of this
  446. 17:30specific string and their code points so
  447. 17:32or of X forx in this string and we get
  448. 17:36this encoding here now see here we've
  449. 17:40already turned the raw code points
  450. 17:42already have integers so why can't we
  451. 17:44simply just use these integers and not
  452. 17:46have any tokenization at all why can't
  453. 17:48we just use this natively as is and just
  454. 17:50use the code Point well one reason for
  455. 17:52that of course is that the vocabulary in
  456. 17:54that case would be quite long so in this
  457. 17:56case for Unicode the this is a
  458. 17:58vocabulary of
  459. 17:59150,000 different code points but more
  460. 18:02worryingly than that I think the Unicode
  461. 18:05standard is very much alive and it keeps
  462. 18:07changing and so it's not kind of a
  463. 18:09stable representation necessarily that
  464. 18:11we may want to use directly so for those
  465. 18:13reasons we need something a bit better
  466. 18:15so to find something better we turn to
  467. 18:17encodings so if we go to the Wikipedia
  468. 18:19page here we see that the Unicode
  469. 18:21consortion defines three types of
  470. 18:23encodings utf8 UTF 16 and UTF 32 these
  471. 18:27encoding are the way by which we can
  472. 18:30take Unicode text and translate it into
  473. 18:33binary data or by streams utf8 is by far
  474. 18:37the most common uh so this is the utf8
  475. 18:39page now this Wikipedia page is actually
  476. 18:42quite long but what's important for our
  477. 18:44purposes is that utf8 takes every single
  478. 18:46Cod point and it translates it to a by
  479. 18:49stream and this by stream is between one
  480. 18:52to four bytes so it's a variable length
  481. 18:54encoding so depending on the Unicode
  482. 18:56Point according to the schema you're
  483. 18:58going to end up with between 1 to four
  484. 18:59bytes for each code point on top of that
  485. 19:03there's utf8 uh
  486. 19:05utf16 and UTF 32 UTF 32 is nice because
  487. 19:08it is fixed length instead of variable
  488. 19:10length but it has many other downsides
  489. 19:12as well so the full kind of spectrum of
  490. 19:17pros and cons of all these different
  491. 19:18three encodings are beyond the scope of
  492. 19:20this video I just like to point out that
  493. 19:22I enjoyed this block post and this block
  494. 19:25post at the end of it also has a number
  495. 19:27of references that can be quite useful
  496. 19:29uh one of them is uh utf8 everywhere
  497. 19:32Manifesto um and this Manifesto
  498. 19:34describes the reason why utf8 is
  499. 19:36significantly preferred and a lot nicer
  500. 19:39than the other encodings and why it is
  501. 19:41used a lot more prominently um on the
  502. 19:45internet one of the major advantages
  503. 19:48just just to give you a sense is that
  504. 19:49utf8 is the only one of these that is
  505. 19:52backwards compatible to the much simpler
  506. 19:54asky encoding of text um but I'm not
  507. 19:57going to go into the full detail in this
  508. 19:58video so suffice to say that we like the
  509. 20:01utf8 encoding and uh let's try to take
  510. 20:03the string and see what we get if we
  511. 20:06encoded into
  512. 20:08utf8 the string class in Python actually
  513. 20:10has do encode and you can give it the
  514. 20:12encoding which is say utf8 now we get
  515. 20:15out of this is not very nice because
  516. 20:17this is the bytes is a bytes object and
  517. 20:20it's not very nice in the way that it's
  518. 20:22printed so I personally like to take it
  519. 20:25through list because then we actually
  520. 20:26get the raw B
  521. 20:28of this uh encoding so this is the raw
  522. 20:32byes that represent this string
  523. 20:35according to the utf8 en coding we can
  524. 20:38also look at utf16 we get a slightly
  525. 20:40different by stream and we here we start
  526. 20:43to see one of the disadvantages of utf16
  527. 20:45you see how we have zero Z something Z
  528. 20:47something Z something we're starting to
  529. 20:49get a sense that this is a bit of a
  530. 20:50wasteful encoding and indeed for simple
  531. 20:53asky characters or English characters
  532. 20:56here uh we just have the structure of 0
  533. 20:58something Z something and it's not
  534. 21:00exactly nice same for UTF 32 when we
  535. 21:04expand this we can start to get a sense
  536. 21:06of the wastefulness of this encoding for
  537. 21:08our purposes you see a lot of zeros
  538. 21:10followed by
  539. 21:11something and so uh this is not
  540. 21:14desirable so suffice it to say that we
  541. 21:17would like to stick with utf8 for our
  542. 21:20purposes however if we just use utf8
  543. 21:23naively these are by streams so that
  544. 21:26would imply a vocabulary length of only
  545. 21:29256 possible tokens uh but this this
  546. 21:33vocabulary size is very very small what
  547. 21:35this is going to do if we just were to
  548. 21:36use it naively is that all of our text
  549. 21:39would be stretched out over very very
  550. 21:41long sequences of bytes and so
  551. 21:46um what what this does is that certainly
  552. 21:49the embeding table is going to be tiny
  553. 21:51and the prediction at the top at the
  554. 21:52final layer is going to be very tiny but
  555. 21:54our sequences are very long and remember
  556. 21:56that we have pretty finite um context
  557. 21:59length and the attention that we can
  558. 22:01support in a transformer for
  559. 22:02computational reasons and so we only
  560. 22:05have as much context length but now we
  561. 22:07have very very long sequences and this
  562. 22:09is just inefficient and it's not going
  563. 22:10to allow us to attend to sufficiently
  564. 22:12long text uh before us for the purposes
  565. 22:15of the next token prediction task so we
  566. 22:18don't want to use the raw bytes of the
  567. 22:21utf8 encoding we want to be able to
  568. 22:24support larger vocabulary size that we
  569. 22:26can tune as a hyper
  570. 22:28but we want to stick with the utf8
  571. 22:30encoding of these strings so what do we
  572. 22:33do well the answer of course is we turn
  573. 22:35to the bite pair encoding algorithm
  574. 22:37which will allow us to compress these
  575. 22:39bite sequences um to a variable amount
  576. 22:42so we'll get to that in a bit but I just
  577. 22:44want to briefly speak to the fact that I
  578. 22:47would love nothing more than to be able
  579. 22:49to feed raw bite sequences into uh
  580. 22:52language models in fact there's a paper
  581. 22:54about how this could potentially be done
  582. 22:57uh from Summer last last year now the
  583. 22:59problem is you actually have to go in
  584. 23:00and you have to modify the Transformer
  585. 23:02architecture because as I mentioned
  586. 23:04you're going to have a problem where the
  587. 23:06attention will start to become extremely
  588. 23:08expensive because the sequences are so
  589. 23:10long and so in this paper they propose
  590. 23:13kind of a hierarchical structuring of
  591. 23:15the Transformer that could allow you to
  592. 23:17just feed in raw bites and so at the end
  593. 23:20they say together these results
  594. 23:21establish the viability of tokenization
  595. 23:23free autor regressive sequence modeling
  596. 23:25at scale so tokenization free would
  597. 23:27indeed be amazing we would just feed B
  598. 23:30streams directly into our models but
  599. 23:32unfortunately I don't know that this has
  600. 23:34really been proven out yet by
  601. 23:36sufficiently many groups and a
  602. 23:37sufficient scale uh but something like
  603. 23:39this at one point would be amazing and I
  604. 23:40hope someone comes up with it but for
  605. 23:42now we have to come back and we can't
  606. 23:44feed this directly into language models
  607. 23:46and we have to compress it using the B
  608. 23:48paare encoding algorithm so let's see
  609. 23:49how that works so as I mentioned the B
  610. 23:51paare encoding algorithm is not all that
  611. 23:53complicated and the Wikipedia page is
  612. 23:55actually quite instructive as far as the
  613. 23:57basic idea goes go what we're doing is
  614. 23:59we have some kind of a input sequence uh
  615. 24:01like for example here we have only four
  616. 24:03elements in our vocabulary a b c and d
  617. 24:06and we have a sequence of them so
  618. 24:08instead of bytes let's say we just have
  619. 24:09four a vocab size of
  620. 24:12four the sequence is too long and we'd
  621. 24:14like to compress it so what we do is
  622. 24:16that we iteratively find the pair of uh
  623. 24:20tokens that occur the most
  624. 24:23frequently and then once we've
  625. 24:25identified that pair we repl replace
  626. 24:28that pair with just a single new token
  627. 24:30that we append to our vocabulary so for
  628. 24:33example here the bite pair AA occurs
  629. 24:36most often so we mint a new token let's
  630. 24:38call it capital Z and we replace every
  631. 24:41single occurrence of AA by Z so now we
  632. 24:46have two Z's here so here we took a
  633. 24:48sequence of 11 characters with
  634. 24:51vocabulary size four and we've converted
  635. 24:54it to a um sequence of only nine tokens
  636. 24:58but now with a vocabulary of five
  637. 25:00because we have a fifth vocabulary
  638. 25:02element that we just created and it's Z
  639. 25:04standing for concatination of AA and we
  640. 25:07can again repeat this process so we
  641. 25:10again look at the sequence and identify
  642. 25:12the pair of tokens that are most
  643. 25:15frequent let's say that that is now AB
  644. 25:19well we are going to replace AB with a
  645. 25:20new token that we meant call Y so y
  646. 25:23becomes ab and then every single
  647. 25:25occurrence of ab is now replaced with y
  648. 25:28so we end up with this so now we only
  649. 25:31have 1 2 3 4 5 6 seven characters in our
  650. 25:35sequence but we have not just um four
  651. 25:40vocabulary elements or five but now we
  652. 25:42have six and for the final round we
  653. 25:45again look through the sequence find
  654. 25:47that the phrase zy or the pair zy is
  655. 25:50most common and replace it one more time
  656. 25:53with another um character let's say x so
  657. 25:56X is z y and we replace all curses of zy
  658. 25:59and we get this following sequence so
  659. 26:02basically after we have gone through
  660. 26:03this process instead of having a um
  661. 26:08sequence of
  662. 26:0911 uh tokens with a vocabulary length of
  663. 26:13four we now have a sequence of 1 2 3
  664. 26:18four five tokens but our vocabulary
  665. 26:21length now is seven and so in this way
  666. 26:25we can iteratively compress our sequence
  667. 26:27I we Mint new tokens so in the in the
  668. 26:30exact same way we start we start out
  669. 26:32with bite sequences so we have 256
  670. 26:36vocabulary size but we're now going to
  671. 26:38go through these and find the bite pairs
  672. 26:40that occur the most and we're going to
  673. 26:42iteratively start minting new tokens
  674. 26:44appending them to our vocabulary and
  675. 26:46replacing things and in this way we're
  676. 26:48going to end up with a compressed
  677. 26:50training data set and also an algorithm
  678. 26:52for taking any arbitrary sequence and
  679. 26:55encoding it using this uh vocabul
  680. 26:58and also decoding it back to Strings so
  681. 27:01let's now Implement all that so here's
  682. 27:03what I did I went to this block post
  683. 27:05that I enjoyed and I took the first
  684. 27:07paragraph and I copy pasted it here into
  685. 27:10text so this is one very long line
  686. 27:13here now to get the tokens as I
  687. 27:15mentioned we just take our text and we
  688. 27:17encode it into utf8 the tokens here at
  689. 27:20this point will be a raw bites single
  690. 27:22stream of bytes and just so that it's
  691. 27:25easier to work with instead of just a
  692. 27:27bytes object I'm going to convert all
  693. 27:29those bytes to integers and then create
  694. 27:32a list of it just so it's easier for us
  695. 27:34to manipulate and work with in Python
  696. 27:35and visualize and here I'm printing all
  697. 27:38of that so this is the original um this
  698. 27:42is the original paragraph and its length
  699. 27:45is
  700. 27:45533 uh code points and then here are the
  701. 27:49bytes encoded in ut utf8 and we see that
  702. 27:53this has a length of 616 bytes at this
  703. 27:56point or 616 tokens and the reason this
  704. 27:59is more is because a lot of these simple
  705. 28:01asky characters or simple characters
  706. 28:04they just become a single bite but a lot
  707. 28:06of these Unicode more complex characters
  708. 28:08become multiple bytes up to four and so
  709. 28:11we are expanding that
  710. 28:12size so now what we'd like to do as a
  711. 28:14first step of the algorithm is we'd like
  712. 28:16to iterate over here and find the pair
  713. 28:18of bites that occur most frequently
  714. 28:22because we're then going to merge it so
  715. 28:24if you are working long on a notebook on
  716. 28:25a side then I encourage you to basically
  717. 28:27click on the link find this notebook and
  718. 28:29try to write that function yourself
  719. 28:31otherwise I'm going to come here and
  720. 28:32Implement first the function that finds
  721. 28:34the most common pair okay so here's what
  722. 28:36I came up with there are many different
  723. 28:38ways to implement this but I'm calling
  724. 28:40the function get stats it expects a list
  725. 28:42of integers I'm using a dictionary to
  726. 28:44keep track of basically the counts and
  727. 28:46then this is a pythonic way to iterate
  728. 28:48consecutive elements of this list uh
  729. 28:51which we covered in the previous video
  730. 28:53and then here I'm just keeping track of
  731. 28:55just incrementing by one um for all the
  732. 28:58pairs so if I call this on all the
  733. 29:00tokens here then the stats comes out
  734. 29:03here so this is the dictionary the keys
  735. 29:06are these topples of consecutive
  736. 29:08elements and this is the count so just
  737. 29:11to uh print it in a slightly better way
  738. 29:14this is one way that I like to do that
  739. 29:17where you it's a little bit compound
  740. 29:20here so you can pause if you like but we
  741. 29:22iterate all all the items the items
  742. 29:25called on dictionary returns pairs of
  743. 29:27key value and instead I create a list
  744. 29:31here of value key because if it's a
  745. 29:35value key list then I can call sort on
  746. 29:37it and by default python will uh use the
  747. 29:41first element which in this case will be
  748. 29:43value to sort by if it's given tles and
  749. 29:46then reverse so it's descending and
  750. 29:48print that so basically it looks like
  751. 29:50101 comma 32 was the most commonly
  752. 29:53occurring consecutive pair and it
  753. 29:55occurred 20 times we can double check
  754. 29:58that that makes reasonable sense so if I
  755. 30:00just search
  756. 30:0210132 then you see that these are the 20
  757. 30:05occurrences of that um pair and if we'd
  758. 30:10like to take a look at what exactly that
  759. 30:11pair is we can use Char which is the
  760. 30:14opposite of or in Python so we give it a
  761. 30:17um unic code Cod point so 101 and of 32
  762. 30:22and we see that this is e and space so
  763. 30:25basically there's a lot of E space here
  764. 30:28meaning that a lot of these words seem
  765. 30:29to end with e so here's eace as an
  766. 30:32example so there's a lot of that going
  767. 30:34on here and this is the most common pair
  768. 30:36so now that we've identified the most
  769. 30:38common pair we would like to iterate
  770. 30:40over this sequence we're going to Mint a
  771. 30:42new token with the ID of
  772. 30:44256 right because these tokens currently
  773. 30:47go from Z to 255 so when we create a new
  774. 30:50token it will have an ID of
  775. 30:52256 and we're going to iterate over this
  776. 30:56entire um list and every every time we
  777. 30:59see 101 comma 32 we're going to swap
  778. 31:02that out for
  779. 31:03256 so let's Implement that now and feel
  780. 31:07free to uh do that yourself as well so
  781. 31:09first I commented uh this just so we
  782. 31:11don't pollute uh the notebook too much
  783. 31:14this is a nice way of in Python
  784. 31:17obtaining the highest ranking pair so
  785. 31:20we're basically calling the Max on this
  786. 31:23dictionary stats and this will return
  787. 31:26the maximum
  788. 31:27key and then the question is how does it
  789. 31:30rank keys so you can provide it with a
  790. 31:32function that ranks keys and that
  791. 31:35function is just stats. getet uh stats.
  792. 31:38getet would basically return the value
  793. 31:41and so we're ranking by the value and
  794. 31:42getting the maximum key so it's 101
  795. 31:45comma 32 as we saw now to actually merge
  796. 31:4910132 um this is the function that I
  797. 31:51wrote but again there are many different
  798. 31:53versions of it so we're going to take a
  799. 31:55list of IDs and the the pair that we
  800. 31:57want to replace and that pair will be
  801. 31:59replaced with the new index
  802. 32:02idx so iterating through IDs if we find
  803. 32:05the pair swap it out for idx so we
  804. 32:08create this new list and then we start
  805. 32:10at zero and then we go through this
  806. 32:12entire list sequentially from left to
  807. 32:14right and here we are checking for
  808. 32:17equality at the current position with
  809. 32:19the
  810. 32:20pair um so here we are checking that the
  811. 32:23pair matches now here is a bit of a
  812. 32:25tricky condition that you have to append
  813. 32:27if you're trying to be careful and that
  814. 32:29is that um you don't want this here to
  815. 32:31be out of Bounds at the very last
  816. 32:33position when you're on the rightmost
  817. 32:35element of this list otherwise this
  818. 32:37would uh give you an autof bounds error
  819. 32:39so we have to make sure that we're not
  820. 32:40at the very very last element so uh this
  821. 32:44would be false for that so if we find a
  822. 32:46match we append to this new list that
  823. 32:51replacement index and we increment the
  824. 32:53position by two so we skip over that
  825. 32:54entire pair but otherwise if we we
  826. 32:57haven't found a matching pair we just
  827. 32:59sort of copy over the um element at that
  828. 33:02position and increment by one then
  829. 33:05return this so here's a very small toy
  830. 33:07example if we have a list 566 791 and we
  831. 33:10want to replace the occurrences of 67
  832. 33:12with 99 then calling this on that will
  833. 33:16give us what we're asking for so here
  834. 33:18the 67 is replaced with
  835. 33:2199 so now I'm going to uncomment this
  836. 33:23for our actual use case where we want to
  837. 33:27take our tokens we want to take the top
  838. 33:29pair here and replace it with 256 to get
  839. 33:33tokens to if we run this we get the
  840. 33:37following so recall that previously we
  841. 33:40had a length 616 in this list and now we
  842. 33:45have a length 596 right so this
  843. 33:48decreased by 20 which makes sense
  844. 33:50because there are 20 occurrences
  845. 33:52moreover we can try to find 256 here and
  846. 33:55we see plenty of occurrences on off it
  847. 33:58and moreover just double check there
  848. 33:59should be no occurrence of 10132 so this
  849. 34:02is the original array plenty of them and
  850. 34:05in the second array there are no
  851. 34:06occurrences of 1032 so we've
  852. 34:08successfully merged this single pair and
  853. 34:11now we just uh iterate this so we are
  854. 34:13going to go over the sequence again find
  855. 34:15the most common pair and replace it so
  856. 34:17let me now write a y Loop that uses
  857. 34:19these functions to do this um sort of
  858. 34:21iteratively and how many times do we do
  859. 34:24it four well that's totally up to us as
  860. 34:26a hyper parameter
  861. 34:27the more um steps we take the larger
  862. 34:30will be our vocabulary and the shorter
  863. 34:33will be our sequence and there is some
  864. 34:35sweet spot that we usually find works
  865. 34:37the best in practice and so this is kind
  866. 34:39of a hyperparameter and we tune it and
  867. 34:41we find good vocabulary sizes as an
  868. 34:44example gp4 currently uses roughly
  869. 34:46100,000 tokens and um bpark that those
  870. 34:49are reasonable numbers currently instead
  871. 34:51the are large language models so let me
  872. 34:53now write uh putting putting it all
  873. 34:55together and uh iterating these steps
  874. 34:58okay now before we dive into the Y loop
  875. 35:00I wanted to add one more cell here where
  876. 35:03I went to the block post and instead of
  877. 35:04grabbing just the first paragraph or two
  878. 35:07I took the entire block post and I
  879. 35:08stretched it out in a single line and
  880. 35:10basically just using longer text will
  881. 35:12allow us to have more representative
  882. 35:13statistics for the bite Pairs and we'll
  883. 35:16just get a more sensible results out of
  884. 35:18it because it's longer text um so here
  885. 35:21we have the raw text we encode it into
  886. 35:24bytes using the utf8 encoding
  887. 35:27and then here as before we are just
  888. 35:30changing it into a list of integers in
  889. 35:31Python just so it's easier to work with
  890. 35:33instead of the raw byes objects and then
  891. 35:36this is the code that I came up with uh
  892. 35:40to actually do the merging in Loop these
  893. 35:44two functions here are identical to what
  894. 35:45we had above I only included them here
  895. 35:48just so that you have the point of
  896. 35:49reference here so uh these two are
  897. 35:53identical and then this is the new code
  898. 35:55that I added so the first first thing we
  899. 35:57want to do is we want to decide on the
  900. 35:58final vocabulary size that we want our
  901. 36:01tokenizer to have and as I mentioned
  902. 36:02this is a hyper parameter and you set it
  903. 36:04in some way depending on your best
  904. 36:06performance so let's say for us we're
  905. 36:08going to use 276 because that way we're
  906. 36:10going to be doing exactly 20
  907. 36:13merges and uh 20 merges because we
  908. 36:15already have
  909. 36:16256 tokens for the raw bytes and to
  910. 36:20reach 276 we have to do 20 merges uh to
  911. 36:23add 20 new
  912. 36:25tokens here uh this is uh one way in
  913. 36:28Python to just create a copy of a list
  914. 36:31so I'm taking the tokens list and by
  915. 36:33wrapping it in a list python will
  916. 36:35construct a new list of all the
  917. 36:37individual elements so this is just a
  918. 36:38copy
  919. 36:39operation then here I'm creating a
  920. 36:42merges uh dictionary so this merges
  921. 36:44dictionary is going to maintain
  922. 36:46basically the child one child two
  923. 36:49mapping to a new uh token and so what
  924. 36:52we're going to be building up here is a
  925. 36:53binary tree of merges but actually it's
  926. 36:56not exactly a tree because a tree would
  927. 36:59have a single root node with a bunch of
  928. 37:01leaves for us we're starting with the
  929. 37:03leaves on the bottom which are the
  930. 37:05individual bites those are the starting
  931. 37:06256 tokens and then we're starting to
  932. 37:09like merge two of them at a time and so
  933. 37:11it's not a tree it's more like a forest
  934. 37:14um uh as we merge these elements
  935. 37:18so for 20 merges we're going to find the
  936. 37:22most commonly occurring pair we're going
  937. 37:25to Mint a new token integer for it so I
  938. 37:28here will start at zero so we'll going
  939. 37:30to start at 256 we're going to print
  940. 37:32that we're merging it and we're going to
  941. 37:34replace all of the occurrences of that
  942. 37:36pair with the new new lied token and
  943. 37:39we're going to record that this pair of
  944. 37:42integers merged into this new
  945. 37:45integer so running this gives us the
  946. 37:49following
  947. 37:51output so we did 20 merges and for
  948. 37:54example the first merge was exactly as
  949. 37:56before the
  950. 37:5810132 um tokens merging into a new token
  951. 38:012556 now keep in mind that the
  952. 38:04individual uh tokens 101 and 32 can
  953. 38:06still occur in the sequence after
  954. 38:08merging it's only when they occur
  955. 38:10exactly consecutively that that becomes
  956. 38:12256
  957. 38:13now um and in particular the other thing
  958. 38:16to notice here is that the token 256
  959. 38:19which is the newly minted token is also
  960. 38:21eligible for merging so here on the
  961. 38:23bottom the 20th merge was a merge of 25
  962. 38:26and 259 becoming
  963. 38:28275 so every time we replace these
  964. 38:31tokens they become eligible for merging
  965. 38:33in the next round of data ration so
  966. 38:35that's why we're building up a small
  967. 38:37sort of binary Forest instead of a
  968. 38:38single individual
  969. 38:40tree one thing we can take a look at as
  970. 38:42well is we can take a look at the
  971. 38:44compression ratio that we've achieved so
  972. 38:46in particular we started off with this
  973. 38:48tokens list um so we started off with
  974. 38:5124,000 bytes and after merging 20 times
  975. 38:56uh we now have only
  976. 38:5819,000 um tokens and so therefore the
  977. 39:01compression ratio simply just dividing
  978. 39:03the two is roughly 1.27 so that's the
  979. 39:06amount of compression we were able to
  980. 39:07achieve of this text with only 20
  981. 39:10merges um and of course the more
  982. 39:13vocabulary elements you add uh the
  983. 39:15greater the compression ratio here would
  984. 39:19be finally so that's kind of like um the
  985. 39:23training of the tokenizer if you will
  986. 39:25now 1 Point I wanted to make is that and
  987. 39:28maybe this is a diagram that can help um
  988. 39:31kind of illustrate is that tokenizer is
  989. 39:33a completely separate object from the
  990. 39:34large language model itself so
  991. 39:37everything in this lecture we're not
  992. 39:38really touching the llm itself uh we're
  993. 39:40just training the tokenizer this is a
  994. 39:41completely separate pre-processing stage
  995. 39:43usually so the tokenizer will have its
  996. 39:46own training set just like a large
  997. 39:47language model has a potentially
  998. 39:49different training set so the tokenizer
  999. 39:52has a training set of documents on which
  1000. 39:53you're going to train the
  1001. 39:54tokenizer and then and um we're
  1002. 39:57performing The Bite pair encoding
  1003. 39:58algorithm as we saw above to train the
  1004. 40:01vocabulary of this
  1005. 40:02tokenizer so it has its own training set
  1006. 40:04it is a pre-processing stage that you
  1007. 40:06would run a single time in the beginning
  1008. 40:09um and the tokenizer is trained using
  1009. 40:11bipar coding algorithm once you have the
  1010. 40:14tokenizer once it's trained and you have
  1011. 40:16the vocabulary and you have the merges
  1012. 40:19uh we can do both encoding and decoding
  1013. 40:22so these two arrows here so the
  1014. 40:24tokenizer is a translation layer between
  1015. 40:27raw text which is as we saw the sequence
  1016. 40:30of Unicode code points it can take raw
  1017. 40:32text and turn it into a token sequence
  1018. 40:35and vice versa it can take a token
  1019. 40:37sequence and translate it back into raw
  1020. 40:40text so now that we have trained uh
  1021. 40:43tokenizer and we have these merges we
  1022. 40:45are going to turn to how we can do the
  1023. 40:47encoding and the decoding step if you
  1024. 40:49give me text here are the tokens and
  1025. 40:51vice versa if you give me tokens here's
  1026. 40:53the text once we have that we can
  1027. 40:55translate between these two Realms and
  1028. 40:57then the language model is going to be
  1029. 40:58trained as a step two afterwards and
  1030. 41:01typically in a in a sort of a
  1031. 41:03state-of-the-art application you might
  1032. 41:05take all of your training data for the
  1033. 41:06language model and you might run it
  1034. 41:08through the tokenizer and sort of
  1035. 41:10translate everything into a massive
  1036. 41:11token sequence and then you can throw
  1037. 41:13away the raw text you're just left with
  1038. 41:15the tokens themselves and those are
  1039. 41:17stored on disk and that is what the
  1040. 41:19large language model is actually reading
  1041. 41:21when it's training on them so this one
  1042. 41:23approach that you can take as a single
  1043. 41:24massive pre-processing step a
  1044. 41:26stage um so yeah basically I think the
  1045. 41:30most important thing I want to get
  1046. 41:31across is that this is completely
  1047. 41:32separate stage it usually has its own
  1048. 41:34entire uh training set you may want to
  1049. 41:36have those training sets be different
  1050. 41:38between the tokenizer and the logge
  1051. 41:39language model so for example when
  1052. 41:41you're training the tokenizer as I
  1053. 41:43mentioned we don't just care about the
  1054. 41:45performance of English text we care
  1055. 41:46about uh multi many different languages
  1056. 41:49and we also care about code or not code
  1057. 41:51so you may want to look into different
  1058. 41:53kinds of mixtures of different kinds of
  1059. 41:55languages and different amounts of code
  1060. 41:57and things like that because the amount
  1061. 42:00of different language that you have in
  1062. 42:01your tokenizer training set will
  1063. 42:03determine how many merges of it there
  1064. 42:06will be and therefore that determines
  1065. 42:08the density with which uh this type of
  1066. 42:11data is um sort of has in the token
  1067. 42:15space and so roughly speaking
  1068. 42:17intuitively if you add some amount of
  1069. 42:19data like say you have a ton of Japanese
  1070. 42:21data in your uh tokenizer training set
  1071. 42:24then that means that more Japanese
  1072. 42:25tokens will get merged
  1073. 42:26and therefore Japanese will have shorter
  1074. 42:28sequences uh and that's going to be
  1075. 42:30beneficial for the large language model
  1076. 42:32which has a finite context length on
  1077. 42:34which it can work on in in the token
  1078. 42:36space uh so hopefully that makes sense
  1079. 42:39so we're now going to turn to encoding
  1080. 42:41and decoding now that we have trained a
  1081. 42:43tokenizer so we have our merges and now
  1082. 42:46how do we do encoding and decoding okay
  1083. 42:48so let's begin with decoding which is
  1084. 42:50this Arrow over here so given a token
  1085. 42:52sequence let's go through the tokenizer
  1086. 42:54to get back a python string object so
  1087. 42:57the raw text so this is the function
  1088. 42:59that we' like to implement um we're
  1089. 43:01given the list of integers and we want
  1090. 43:03to return a python string if you'd like
  1091. 43:05uh try to implement this function
  1092. 43:06yourself it's a fun exercise otherwise
  1093. 43:08I'm going to start uh pasting in my own
  1094. 43:11solution so there are many different
  1095. 43:13ways to do it um here's one way I will
  1096. 43:16create an uh kind of pre-processing
  1097. 43:18variable that I will call
  1098. 43:21vocab and vocab is a mapping or a
  1099. 43:24dictionary in Python for from the token
  1100. 43:27uh ID to the bytes object for that token
  1101. 43:31so we begin with the raw bytes for
  1102. 43:33tokens from 0 to 255 and then we go in
  1103. 43:36order of all the merges and we sort of
  1104. 43:39uh populate this vocab list by doing an
  1105. 43:42addition here so this is the basically
  1106. 43:45the bytes representation of the first
  1107. 43:47child followed by the second one and
  1108. 43:50remember these are bytes objects so this
  1109. 43:52addition here is an addition of two
  1110. 43:54bytes objects just concatenation
  1111. 43:57so that's what we get
  1112. 43:58here one tricky thing to be careful with
  1113. 44:01by the way is that I'm iterating a
  1114. 44:02dictionary in Python using a DOT items
  1115. 44:06and uh it really matters that this runs
  1116. 44:08in the order in which we inserted items
  1117. 44:11into the merous dictionary luckily
  1118. 44:13starting with python 3.7 this is
  1119. 44:15guaranteed to be the case but before
  1120. 44:17python 3.7 this iteration may have been
  1121. 44:19out of order with respect to how we
  1122. 44:20inserted elements into merges and this
  1123. 44:23may not have worked but we are using an
  1124. 44:25um modern python so we're okay and then
  1125. 44:28here uh given the IDS the first thing
  1126. 44:31we're going to do is get the
  1127. 44:35tokens so the way I implemented this
  1128. 44:37here is I'm taking I'm iterating over
  1129. 44:39all the IDS I'm using vocap to look up
  1130. 44:41their bytes and then here this is one
  1131. 44:44way in Python to concatenate all these
  1132. 44:46bytes together to create our tokens and
  1133. 44:49then these tokens here at this point are
  1134. 44:51raw bytes so I have to decode using UTF
  1135. 44:56F now back into python strings so
  1136. 44:59previously we called that encode on a
  1137. 45:01string object to get the bytes and now
  1138. 45:03we're doing it Opposite we're taking the
  1139. 45:05bytes and calling a decode on the bytes
  1140. 45:07object to get a string in Python and
  1141. 45:11then we can return
  1142. 45:13text so um this is how we can do it now
  1143. 45:16this actually has a um issue um in the
  1144. 45:20way I implemented it and this could
  1145. 45:22actually throw an error so try to think
  1146. 45:24figure out why this code could actually
  1147. 45:26result in an error if we plug in um uh
  1148. 45:30some sequence of IDs that is
  1149. 45:32unlucky so let me demonstrate the issue
  1150. 45:35when I try to decode just something like
  1151. 45:3797 I am going to get letter A here back
  1152. 45:41so nothing too crazy happening but when
  1153. 45:44I try to decode 128 as a single element
  1154. 45:48the token 128 is what in string or in
  1155. 45:51Python object uni Cod decoder utfa can't
  1156. 45:55Decode by um 0x8 which is this in HEX in
  1157. 46:00position zero invalid start bite what
  1158. 46:01does that mean well to understand what
  1159. 46:03this means we have to go back to our
  1160. 46:04utf8 page uh that I briefly showed
  1161. 46:07earlier and this is Wikipedia utf8 and
  1162. 46:10basically there's a specific schema that
  1163. 46:13utfa bytes take so in particular if you
  1164. 46:16have a multi-te object for some of the
  1165. 46:19Unicode characters they have to have
  1166. 46:21this special sort of envelope in how the
  1167. 46:24encoding works and so what's happening
  1168. 46:26here is that invalid start pite that's
  1169. 46:30because
  1170. 46:31128 the binary representation of it is
  1171. 46:33one followed by all zeros so we have one
  1172. 46:37and then all zero and we see here that
  1173. 46:39that doesn't conform to the format
  1174. 46:41because one followed by all zero just
  1175. 46:42doesn't fit any of these rules so to
  1176. 46:44speak so it's an invalid start bite
  1177. 46:47which is byte one this one must have a
  1178. 46:50one following it and then a zero
  1179. 46:52following it and then the content of
  1180. 46:54your uni codee in x here so basically we
  1181. 46:57don't um exactly follow the utf8
  1182. 46:59standard and this cannot be decoded and
  1183. 47:02so the way to fix this um is to
  1184. 47:06use this errors equals in bytes. decode
  1185. 47:11function of python and by default errors
  1186. 47:13is strict so we will throw an error if
  1187. 47:17um it's not valid utf8 bytes encoding
  1188. 47:20but there are many different things that
  1189. 47:21you could put here on error handling
  1190. 47:23this is the full list of all the errors
  1191. 47:25that you can use and in particular
  1192. 47:27instead of strict let's change it to
  1193. 47:29replace and that will replace uh with
  1194. 47:32this special marker this replacement
  1195. 47:35character so errors equals replace and
  1196. 47:40now we just get that character
  1197. 47:43back so basically not every single by
  1198. 47:46sequence is valid
  1199. 47:48utf8 and if it happens that your large
  1200. 47:51language model for example predicts your
  1201. 47:53tokens in a bad manner then they might
  1202. 47:56not fall into valid utf8 and then we
  1203. 48:00won't be able to decode them so the
  1204. 48:02standard practice is to basically uh use
  1205. 48:05errors equals replace and this is what
  1206. 48:07you will also find in the openai um code
  1207. 48:10that they released as well but basically
  1208. 48:12whenever you see um this kind of a
  1209. 48:14character in your output in that case uh
  1210. 48:16something went wrong and the LM output
  1211. 48:18not was not valid uh sort of sequence of
  1212. 48:21tokens okay and now we're going to go
  1213. 48:23the other way so we are going to
  1214. 48:25implement
  1215. 48:26this Arrow right here where we are going
  1216. 48:27to be given a string and we want to
  1217. 48:29encode it into
  1218. 48:31tokens so this is the signature of the
  1219. 48:33function that we're interested in and um
  1220. 48:36this should basically print a list of
  1221. 48:38integers of the tokens so again uh try
  1222. 48:41to maybe implement this yourself if
  1223. 48:43you'd like a fun exercise uh and pause
  1224. 48:45here otherwise I'm going to start
  1225. 48:46putting in my
  1226. 48:47solution so again there are many ways to
  1227. 48:50do this so um this is one of the ways
  1228. 48:53that sort of I came came up with so the
  1229. 48:57first thing we're going to do is we are
  1230. 48:59going
  1231. 49:00to uh take our text encode it into utf8
  1232. 49:03to get the raw bytes and then as before
  1233. 49:05we're going to call list on the bytes
  1234. 49:07object to get a list of integers of
  1235. 49:10those bytes so those are the starting
  1236. 49:12tokens those are the raw bytes of our
  1237. 49:14sequence but now of course according to
  1238. 49:16the merges dictionary above and recall
  1239. 49:19this was the
  1240. 49:21merges some of the bytes may be merged
  1241. 49:23according to this lookup in addition to
  1242. 49:26that remember that the merges was built
  1243. 49:28from top to bottom and this is sort of
  1244. 49:29the order in which we inserted stuff
  1245. 49:31into merges and so we prefer to do all
  1246. 49:34these merges in the beginning before we
  1247. 49:36do these merges later because um for
  1248. 49:39example this merge over here relies on
  1249. 49:40the 256 which got merged here so we have
  1250. 49:44to go in the order from top to bottom
  1251. 49:46sort of if we are going to be merging
  1252. 49:48anything now we expect to be doing a few
  1253. 49:51merges so we're going to be doing W
  1254. 49:54true um and now we want to find a pair
  1255. 49:58of byes that is consecutive that we are
  1256. 50:00allowed to merge according to this in
  1257. 50:03order to reuse some of the functionality
  1258. 50:05that we've already written I'm going to
  1259. 50:06reuse the function uh get
  1260. 50:09stats so recall that get stats uh will
  1261. 50:12give us the we'll basically count up how
  1262. 50:14many times every single pair occurs in
  1263. 50:16our sequence of tokens and return that
  1264. 50:18as a dictionary and the dictionary was a
  1265. 50:22mapping from all the different uh by
  1266. 50:25pairs to the number of times that they
  1267. 50:27occur right um at this point we don't
  1268. 50:30actually care how many times they occur
  1269. 50:32in the sequence we only care what the
  1270. 50:34raw pairs are in that sequence and so
  1271. 50:36I'm only going to be using basically the
  1272. 50:38keys of the dictionary I only care about
  1273. 50:40the set of possible merge candidates if
  1274. 50:42that makes
  1275. 50:43sense now we want to identify the pair
  1276. 50:46that we're going to be merging at this
  1277. 50:47stage of the loop so what do we want we
  1278. 50:50want to find the pair or like the a key
  1279. 50:53inside stats that has the lowest index
  1280. 50:57in the merges uh dictionary because we
  1281. 50:59want to do all the early merges before
  1282. 51:01we work our way to the late
  1283. 51:03merges so again there are many different
  1284. 51:05ways to implement this but I'm going to
  1285. 51:07do something a little bit fancy
  1286. 51:11here so I'm going to be using the Min
  1287. 51:14over an iterator in Python when you call
  1288. 51:16Min on an iterator and stats here as a
  1289. 51:18dictionary we're going to be iterating
  1290. 51:20the keys of this dictionary in Python so
  1291. 51:24we're looking at all the pairs inside
  1292. 51:27stats um which are all the consecutive
  1293. 51:29Pairs and we're going to be taking the
  1294. 51:32consecutive pair inside tokens that has
  1295. 51:34the minimum what the Min takes a key
  1296. 51:38which gives us the function that is
  1297. 51:40going to return a value over which we're
  1298. 51:42going to do the Min and the one we care
  1299. 51:44about is we're we care about taking
  1300. 51:46merges and basically getting um that
  1301. 51:50pairs
  1302. 51:52index so basically for any pair inside
  1303. 51:57stats we are going to be looking into
  1304. 51:59merges at what index it has and we want
  1305. 52:03to get the pair with the Min number so
  1306. 52:05as an example if there's a pair 101 and
  1307. 52:0732 we definitely want to get that pair
  1308. 52:10uh we want to identify it here and
  1309. 52:11return it and pair would become 10132 if
  1310. 52:15it
  1311. 52:15occurs and the reason that I'm putting a
  1312. 52:17float INF here as a fall back is that in
  1313. 52:21the get function when we call uh when we
  1314. 52:24basically consider a pair that doesn't
  1315. 52:26occur in the merges then that pair is
  1316. 52:29not eligible to be merged right so if in
  1317. 52:31the token sequence there's some pair
  1318. 52:33that is not a merging pair it cannot be
  1319. 52:35merged then uh it doesn't actually occur
  1320. 52:38here and it doesn't have an index and uh
  1321. 52:40it cannot be merged which we will denote
  1322. 52:42as float INF and the reason Infinity is
  1323. 52:45nice here is because for sure we're
  1324. 52:46guaranteed that it's not going to
  1325. 52:48participate in the list of candidates
  1326. 52:50when we do the men so uh so this is one
  1327. 52:53way to do it so B basically long story
  1328. 52:55short this Returns the most eligible
  1329. 52:58merging candidate pair uh that occurs in
  1330. 53:01the tokens now one thing to be careful
  1331. 53:04with here is this uh function here might
  1332. 53:07fail in the following way if there's
  1333. 53:09nothing to merge then uh uh then there's
  1334. 53:13nothing in merges um that satisfi that
  1335. 53:16is satisfied anymore there's nothing to
  1336. 53:18merge everything just returns float imps
  1337. 53:21and then the pair I think will just
  1338. 53:23become the very first element of stats
  1339. 53:26um but this pair is not actually a
  1340. 53:28mergeable pair it just becomes the first
  1341. 53:31pair inside stats arbitrarily because
  1342. 53:33all of these pairs evaluate to float in
  1343. 53:36for the merging Criterion so basically
  1344. 53:38it could be that this this doesn't look
  1345. 53:40succeed because there's no more merging
  1346. 53:41pairs so if this pair is not in merges
  1347. 53:44that was returned then this is a signal
  1348. 53:46for us that actually there was nothing
  1349. 53:48to merge no single pair can be merged
  1350. 53:50anymore in that case we will break
  1351. 53:53out um nothing else can be
  1352. 53:57merged you may come up with a different
  1353. 53:59implementation by the way this is kind
  1354. 54:01of like really trying hard in
  1355. 54:03Python um but really we're just trying
  1356. 54:05to find a pair that can be merged with
  1357. 54:07the lowest index
  1358. 54:09here now if we did find a pair that is
  1359. 54:13inside merges with the lowest index then
  1360. 54:16we can merge it
  1361. 54:19so we're going to look into the merger
  1362. 54:22dictionary for that pair to look up the
  1363. 54:24index and we're going to now merge that
  1364. 54:27into that index so we're going to do
  1365. 54:29tokens equals and we're going to
  1366. 54:32replace the original tokens we're going
  1367. 54:34to be replacing the pair pair and we're
  1368. 54:36going to be replacing it with index idx
  1369. 54:38and this returns a new list of tokens
  1370. 54:41where every occurrence of pair is
  1371. 54:43replaced with idx so we're doing a merge
  1372. 54:46and we're going to be continuing this
  1373. 54:47until eventually nothing can be merged
  1374. 54:49we'll come out here and we'll break out
  1375. 54:51and here we just return
  1376. 54:53tokens and so that that's the
  1377. 54:55implementation I think so hopefully this
  1378. 54:57runs okay cool um yeah and this looks uh
  1379. 55:02reasonable so for example 32 is a space
  1380. 55:04in asky so that's here um so this looks
  1381. 55:09like it worked great okay so let's wrap
  1382. 55:11up this section of the video at least I
  1383. 55:13wanted to point out that this is not
  1384. 55:14quite the right implementation just yet
  1385. 55:16because we are leaving out a special
  1386. 55:17case so in particular if uh we try to do
  1387. 55:20this this would give us an error and the
  1388. 55:23issue is that um if we only have a
  1389. 55:25single character or an empty string then
  1390. 55:28stats is empty and that causes an issue
  1391. 55:29inside Min so one way to fight this is
  1392. 55:32if L of tokens is at least two because
  1393. 55:36if it's less than two it's just a single
  1394. 55:37token or no tokens then let's just uh
  1395. 55:40there's nothing to merge so we just
  1396. 55:41return so that would fix uh that
  1397. 55:44case Okay and then second I have a few
  1398. 55:48test cases here for us as well so first
  1399. 55:50let's make sure uh about or let's note
  1400. 55:53the following if we take a string and we
  1401. 55:56try to encode it and then decode it back
  1402. 55:58you'd expect to get the same string back
  1403. 56:00right is that true for all
  1404. 56:04strings so I think uh so here it is the
  1405. 56:07case and I think in general this is
  1406. 56:08probably the case um but notice that
  1407. 56:12going backwards is not is not you're not
  1408. 56:14going to have an identity going
  1409. 56:15backwards because as I mentioned us not
  1410. 56:19all token sequences are valid utf8 uh
  1411. 56:22sort of by streams and so so therefore
  1412. 56:25you're some of them can't even be
  1413. 56:27decodable um so this only goes in One
  1414. 56:30Direction but for that one direction we
  1415. 56:32can check uh here if we take the
  1416. 56:34training text which is the text that we
  1417. 56:36train to tokenizer around we can make
  1418. 56:38sure that when we encode and decode we
  1419. 56:39get the same thing back which is true
  1420. 56:41and here I took some validation data so
  1421. 56:43I went to I think this web page and I
  1422. 56:45grabbed some text so this is text that
  1423. 56:47the tokenizer has not seen and we can
  1424. 56:49make sure that this also works um okay
  1425. 56:52so that gives us some confidence that
  1426. 56:53this was correctly implemented
  1427. 56:56so those are the basics of the bite pair
  1428. 56:58encoding algorithm we saw how we can uh
  1429. 57:00take some training set train a tokenizer
  1430. 57:03the parameters of this tokenizer really
  1431. 57:05are just this dictionary of merges and
  1432. 57:08that basically creates the little binary
  1433. 57:09Forest on top of raw
  1434. 57:11bites once we have this the merges table
  1435. 57:14we can both encode and decode between
  1436. 57:16raw text and token sequences so that's
  1437. 57:19the the simplest setting of The
  1438. 57:21tokenizer what we're going to do now
  1439. 57:23though is we're going to look at some of
  1440. 57:24the St the art lar language models and
  1441. 57:26the kinds of tokenizers that they use
  1442. 57:28and we're going to see that this picture
  1443. 57:29complexifies very quickly so we're going
  1444. 57:31to go through the details of this comp
  1445. 57:34complexification one at a time so let's
  1446. 57:37kick things off by looking at the GPD
  1447. 57:39Series so in particular I have the gpt2
  1448. 57:41paper here um and this paper is from
  1449. 57:442019 or so so 5 years ago and let's
  1450. 57:48scroll down to input representation this
  1451. 57:51is where they talk about the tokenizer
  1452. 57:52that they're using for gpd2 now this is
  1453. 57:55all fairly readable so I encourage you
  1454. 57:57to pause and um read this yourself but
  1455. 58:00this is where they motivate the use of
  1456. 58:02the bite pair encoding algorithm on the
  1457. 58:04bite level representation of utf8
  1458. 58:07encoding so this is where they motivate
  1459. 58:09it and they talk about the vocabulary
  1460. 58:11sizes and everything now everything here
  1461. 58:13is exactly as we've covered it so far
  1462. 58:15but things start to depart around here
  1463. 58:18so what they mention is that they don't
  1464. 58:20just apply the naive algorithm as we
  1465. 58:22have done it and in particular here's a
  1466. 58:25example suppose that you have common
  1467. 58:27words like dog what will happen is that
  1468. 58:29dog of course occurs very frequently in
  1469. 58:31the text and it occurs right next to all
  1470. 58:34kinds of punctuation as an example so
  1471. 58:36doc dot dog exclamation mark dog
  1472. 58:39question mark Etc and naively you might
  1473. 58:42imagine that the BP algorithm could
  1474. 58:43merge these to be single tokens and then
  1475. 58:45you end up with lots of tokens that are
  1476. 58:47just like dog with a slightly different
  1477. 58:49punctuation and so it feels like you're
  1478. 58:50clustering things that shouldn't be
  1479. 58:52clustered you're combining kind of
  1480. 58:53semantics with
  1481. 58:55uation and this uh feels suboptimal and
  1482. 58:58indeed they also say that this is
  1483. 59:00suboptimal according to some of the
  1484. 59:02experiments so what they want to do is
  1485. 59:04they want to top down in a manual way
  1486. 59:06enforce that some types of um characters
  1487. 59:09should never be merged together um so
  1488. 59:12they want to enforce these merging rules
  1489. 59:14on top of the bite PA encoding algorithm
  1490. 59:17so let's take a look um at their code
  1491. 59:19and see how they actually enforce this
  1492. 59:21and what kinds of mergy they actually do
  1493. 59:23perform so I have to to tab open here
  1494. 59:25for gpt2 under open AI on GitHub and
  1495. 59:29when we go to
  1496. 59:30Source there is an encoder thatp now I
  1497. 59:34don't personally love that they call it
  1498. 59:35encoder dopy because this is the
  1499. 59:37tokenizer and the tokenizer can do both
  1500. 59:39encode and decode uh so it feels kind of
  1501. 59:41awkward to me that it's called encoder
  1502. 59:43but that is the tokenizer and there's a
  1503. 59:45lot going on here and we're going to
  1504. 59:47step through it in detail at one point
  1505. 59:49for now I just want to focus on this
  1506. 59:51part here the create a rigix pattern
  1507. 59:54here that looks very complicated and
  1508. 59:56we're going to go through it in a bit uh
  1509. 59:58but this is the core part that allows
  1510. 1:00:00them to enforce rules uh for what parts
  1511. 1:00:04of the text Will Never Be merged for
  1512. 1:00:05sure now notice that re. compile here is
  1513. 1:00:08a little bit misleading because we're
  1514. 1:00:10not just doing import re which is the
  1515. 1:00:12python re module we're doing import reex
  1516. 1:00:14as re and reex is a python package that
  1517. 1:00:17you can install P install r x and it's
  1518. 1:00:20basically an extension of re so it's a
  1519. 1:00:22bit more powerful
  1520. 1:00:23re um
  1521. 1:00:26so let's take a look at this pattern and
  1522. 1:00:28what it's doing and why this is actually
  1523. 1:00:30doing the separation that they are
  1524. 1:00:32looking for okay so I've copy pasted the
  1525. 1:00:34pattern here to our jupit notebook where
  1526. 1:00:37we left off and let's take this pattern
  1527. 1:00:39for a spin so in the exact same way that
  1528. 1:00:42their code does we're going to call an
  1529. 1:00:44re. findall for this pattern on any
  1530. 1:00:47arbitrary string that we are interested
  1531. 1:00:49so this is the string that we want to
  1532. 1:00:50encode into tokens um to feed into n llm
  1533. 1:00:55like gpt2 so what exactly is this doing
  1534. 1:00:59well re. findall will take this pattern
  1535. 1:01:01and try to match it against a
  1536. 1:01:02string um the way this works is that you
  1537. 1:01:06are going from left to right in the
  1538. 1:01:07string and you're trying to match the
  1539. 1:01:10pattern and R.F find all will get all
  1540. 1:01:13the occurrences and organize them into a
  1541. 1:01:16list now when you look at the um when
  1542. 1:01:19you look at this pattern first of all
  1543. 1:01:20notice that this is a raw string um and
  1544. 1:01:23then these are three double quotes just
  1545. 1:01:26to start the string so really the string
  1546. 1:01:28itself this is the pattern itself
  1547. 1:01:31right and notice that it's made up of a
  1548. 1:01:34lot of ores so see these vertical bars
  1549. 1:01:36those are ores in reg X and so you go
  1550. 1:01:40from left to right in this pattern and
  1551. 1:01:41try to match it against the string
  1552. 1:01:43wherever you are so we have hello and
  1553. 1:01:46we're going to try to match it well it's
  1554. 1:01:48not apostrophe s it's not apostrophe t
  1555. 1:01:50or any of these but it is an optional
  1556. 1:01:53space followed by- P of uh sorry SL P of
  1557. 1:01:58L one or more times what is/ P of L it
  1558. 1:02:02is coming to some documentation that I
  1559. 1:02:04found um there might be other sources as
  1560. 1:02:08well uh SLP is a letter any kind of
  1561. 1:02:11letter from any language and hello is
  1562. 1:02:15made up of letters h e l Etc so optional
  1563. 1:02:19space followed by a bunch of letters one
  1564. 1:02:21or more letters is going to match hello
  1565. 1:02:24but then the match ends because a white
  1566. 1:02:27space is not a letter so from there on
  1567. 1:02:31begins a new sort of attempt to match
  1568. 1:02:33against the string again and starting in
  1569. 1:02:36here we're going to skip over all of
  1570. 1:02:38these again until we get to the exact
  1571. 1:02:40same Point again and we see that there's
  1572. 1:02:42an optional space this is the optional
  1573. 1:02:44space followed by a bunch of letters one
  1574. 1:02:46or more of them and so that matches so
  1575. 1:02:48when we run this we get a list of two
  1576. 1:02:52elements hello and then space world
  1577. 1:02:55so how are you if we add more letters we
  1578. 1:02:58would just get them like this now what
  1579. 1:03:01is this doing and why is this important
  1580. 1:03:03we are taking our string and instead of
  1581. 1:03:05directly encoding it um for
  1582. 1:03:09tokenization we are first splitting it
  1583. 1:03:11up and when you actually step through
  1584. 1:03:13the code and we'll do that in a bit more
  1585. 1:03:15detail what really is doing on a high
  1586. 1:03:17level is that it first splits your text
  1587. 1:03:20into a list of texts just like this one
  1588. 1:03:24and all these elements of this list are
  1589. 1:03:26processed independently by the tokenizer
  1590. 1:03:29and all of the results of that
  1591. 1:03:30processing are simply
  1592. 1:03:32concatenated so hello world oh I I
  1593. 1:03:35missed how hello world how are you we
  1594. 1:03:39have five elements of list all of these
  1595. 1:03:41will independent
  1596. 1:03:44independently go from text to a token
  1597. 1:03:47sequence and then that token sequence is
  1598. 1:03:49going to be concatenated it's all going
  1599. 1:03:50to be joined up and roughly speaking
  1600. 1:03:54what that does is you're only ever
  1601. 1:03:56finding merges between the elements of
  1602. 1:03:58this list so you can only ever consider
  1603. 1:04:00merges within every one of these
  1604. 1:04:01elements in
  1605. 1:04:03individually and um after you've done
  1606. 1:04:06all the possible merging for all of
  1607. 1:04:07these elements individually the results
  1608. 1:04:09of all that will be joined um by
  1609. 1:04:13concatenation and so you are basically
  1610. 1:04:16what what you're doing effectively is
  1611. 1:04:18you are never going to be merging this e
  1612. 1:04:21with this space because they are now
  1613. 1:04:23parts of the separate elements of this
  1614. 1:04:25list and so you are saying we are never
  1615. 1:04:27going to merge
  1616. 1:04:28eace um because we're breaking it up in
  1617. 1:04:32this way so basically using this regx
  1618. 1:04:35pattern to Chunk Up the text is just one
  1619. 1:04:37way of enforcing that some merges are
  1620. 1:04:41not to happen and we're going to go into
  1621. 1:04:43more of this text and we'll see that
  1622. 1:04:45what this is trying to do on a high
  1623. 1:04:46level is we're trying to not merge
  1624. 1:04:48across letters across numbers across
  1625. 1:04:50punctuation and so on so let's see in
  1626. 1:04:53more detail how that works so let's
  1627. 1:04:54continue now we have/ P ofn if you go to
  1628. 1:04:58the documentation SLP of n is any kind
  1629. 1:05:01of numeric character in any script so
  1630. 1:05:04it's numbers so we have an optional
  1631. 1:05:06space followed by numbers and those
  1632. 1:05:08would be separated out so letters and
  1633. 1:05:10numbers are being separated so if I do
  1634. 1:05:12Hello World 123 how are you then world
  1635. 1:05:15will stop matching here because one is
  1636. 1:05:17not a letter anymore but one is a number
  1637. 1:05:20so this group will match for that and
  1638. 1:05:22we'll get it as a separate entity
  1639. 1:05:26uh let's see how these apostrophes work
  1640. 1:05:28so here if we have
  1641. 1:05:31um uh Slash V or I mean apostrophe V as
  1642. 1:05:35an example then apostrophe here is not a
  1643. 1:05:38letter or a
  1644. 1:05:39number so hello will stop matching and
  1645. 1:05:42then we will exactly match this with
  1646. 1:05:44that so that will come out as a separate
  1647. 1:05:48thing so why are they doing the
  1648. 1:05:50apostrophes here honestly I think that
  1649. 1:05:52these are just like very common
  1650. 1:05:53apostrophes p uh that are used um
  1651. 1:05:56typically I don't love that they've done
  1652. 1:05:59this
  1653. 1:06:00because uh let me show you what happens
  1654. 1:06:03when you have uh some Unicode
  1655. 1:06:05apostrophes like for example you can
  1656. 1:06:07have if you have house then this will be
  1657. 1:06:10separated out because of this matching
  1658. 1:06:13but if you use the Unicode apostrophe
  1659. 1:06:15like
  1660. 1:06:16this then suddenly this does not work
  1661. 1:06:19and so this apostrophe will actually
  1662. 1:06:21become its own thing now and so so um
  1663. 1:06:24it's basically hardcoded for this
  1664. 1:06:26specific kind of apostrophe and uh
  1665. 1:06:29otherwise they become completely
  1666. 1:06:31separate tokens in addition to this you
  1667. 1:06:34can go to the gpt2 docs and here when
  1668. 1:06:38they Define the pattern they say should
  1669. 1:06:40have added re. ignore case so BP merges
  1670. 1:06:43can happen for capitalized versions of
  1671. 1:06:44contractions so what they're pointing
  1672. 1:06:46out is that you see how this is
  1673. 1:06:47apostrophe and then lowercase letters
  1674. 1:06:50well because they didn't do re. ignore
  1675. 1:06:52case then then um these rules will not
  1676. 1:06:56separate out the apostrophes if it's
  1677. 1:06:58uppercase so
  1678. 1:07:01house would be like this but if I did
  1679. 1:07:06house if I'm uppercase then notice
  1680. 1:07:10suddenly the apostrophe comes by
  1681. 1:07:12itself so the tokenization will work
  1682. 1:07:15differently in uppercase and lower case
  1683. 1:07:17inconsistently separating out these
  1684. 1:07:19apostrophes so it feels extremely gnarly
  1685. 1:07:21and slightly gross um but that's that's
  1686. 1:07:24how that works okay so let's come back
  1687. 1:07:27after trying to match a bunch of
  1688. 1:07:28apostrophe Expressions by the way the
  1689. 1:07:30other issue here is that these are quite
  1690. 1:07:32language specific probably so I don't
  1691. 1:07:34know that all the languages for example
  1692. 1:07:35use or don't use apostrophes but that
  1693. 1:07:37would be inconsistently tokenized as a
  1694. 1:07:39result then we try to match letters then
  1695. 1:07:42we try to match numbers and then if that
  1696. 1:07:44doesn't work we fall back to here and
  1697. 1:07:47what this is saying is again optional
  1698. 1:07:49space followed by something that is not
  1699. 1:07:50a letter number or a space in one or
  1700. 1:07:53more of that so what this is doing
  1701. 1:07:55effectively is this is trying to match
  1702. 1:07:57punctuation roughly speaking not letters
  1703. 1:07:59and not numbers so this group will try
  1704. 1:08:02to trigger for that so if I do something
  1705. 1:08:04like this then these parts here are not
  1706. 1:08:08letters or numbers but they will
  1707. 1:08:09actually they are uh they will actually
  1708. 1:08:12get caught here and so they become its
  1709. 1:08:14own group so we've separated out the
  1710. 1:08:17punctuation and finally this um this is
  1711. 1:08:20also a little bit confusing so this is
  1712. 1:08:22matching white space but this is using a
  1713. 1:08:25negative look ahead assertion in regex
  1714. 1:08:29so what this is doing is it's matching
  1715. 1:08:30wh space up to but not including the
  1716. 1:08:33last Whit space
  1717. 1:08:35character why is this important um this
  1718. 1:08:37is pretty subtle I think so you see how
  1719. 1:08:40the white space is always included at
  1720. 1:08:41the beginning of the word so um space r
  1721. 1:08:45space u Etc suppose we have a lot of
  1722. 1:08:48spaces
  1723. 1:08:49here what's going to happen here is that
  1724. 1:08:52these spaces up to not including the
  1725. 1:08:54last character will get caught by this
  1726. 1:08:57and what that will do is it will
  1727. 1:08:59separate out the spaces up to but not
  1728. 1:09:01including the last character so that the
  1729. 1:09:03last character can come here and join
  1730. 1:09:05with the um space you and the reason
  1731. 1:09:09that's nice is because space you is the
  1732. 1:09:11common token so if I didn't have these
  1733. 1:09:13Extra Spaces here you would just have
  1734. 1:09:15space you and if I add tokens if I add
  1735. 1:09:18spaces we still have a space view but
  1736. 1:09:20now we have all this extra white space
  1737. 1:09:22so basically the GB to tokenizer really
  1738. 1:09:24likes to have a space letters or numbers
  1739. 1:09:27um and it it preens these spaces and
  1740. 1:09:30this is just something that it is
  1741. 1:09:31consistent about so that's what that is
  1742. 1:09:33for and then finally we have all the the
  1743. 1:09:36last fallback is um whites space
  1744. 1:09:38characters uh so um that would be
  1745. 1:09:42just um if that doesn't get caught then
  1746. 1:09:46this thing will catch any trailing
  1747. 1:09:48spaces and so on I wanted to show one
  1748. 1:09:50more real world example here so if we
  1749. 1:09:53have this string which is a piece of
  1750. 1:09:54python code and then we try to split it
  1751. 1:09:56up then this is the kind of output we
  1752. 1:09:58get so you'll notice that the list has
  1753. 1:10:00many elements here and that's because we
  1754. 1:10:02are splitting up fairly often uh every
  1755. 1:10:05time sort of a category
  1756. 1:10:07changes um so there will never be any
  1757. 1:10:09merges Within These
  1758. 1:10:10elements and um that's what you are
  1759. 1:10:13seeing here now you might think that in
  1760. 1:10:16order to train the
  1761. 1:10:17tokenizer uh open AI has used this to
  1762. 1:10:21split up text into chunks and then run
  1763. 1:10:23just a BP algorithm within all the
  1764. 1:10:25chunks but that is not exactly what
  1765. 1:10:27happened and the reason is the following
  1766. 1:10:30notice that we have the spaces here uh
  1767. 1:10:33those Spaces end up being entire
  1768. 1:10:35elements but these spaces never actually
  1769. 1:10:38end up being merged by by open Ai and
  1770. 1:10:40the way you can tell is that if you copy
  1771. 1:10:42paste the exact same chunk here into Tik
  1772. 1:10:44token U Tik tokenizer you see that all
  1773. 1:10:47the spaces are kept independent and
  1774. 1:10:49they're all token
  1775. 1:10:51220 so I think opena at some point Point
  1776. 1:10:53en Force some rule that these spaces
  1777. 1:10:56would never be merged and so um there's
  1778. 1:10:59some additional rules on top of just
  1779. 1:11:01chunking and bpe that open ey is not uh
  1780. 1:11:04clear about now the training code for
  1781. 1:11:06the gpt2 tokenizer was never released so
  1782. 1:11:08all we have is uh the code that I've
  1783. 1:11:10already shown you but this code here
  1784. 1:11:13that they've released is only the
  1785. 1:11:14inference code for the tokens so this is
  1786. 1:11:17not the training code you can't give it
  1787. 1:11:19a piece of text and training tokenizer
  1788. 1:11:21this is just the inference code which
  1789. 1:11:23Tak takes the merges that we have up
  1790. 1:11:25above and applies them to a new piece of
  1791. 1:11:28text and so we don't know exactly how
  1792. 1:11:30opening ey trained um train the
  1793. 1:11:32tokenizer but it wasn't as simple as
  1794. 1:11:34chunk it up and BP it uh whatever it was
  1795. 1:11:38next I wanted to introduce you to the
  1796. 1:11:40Tik token library from openai which is
  1797. 1:11:42the official library for tokenization
  1798. 1:11:44from openai so this is Tik token bip
  1799. 1:11:48install P to Tik token and then um you
  1800. 1:11:51can do the tokenization in inference
  1801. 1:11:54this is again not training code this is
  1802. 1:11:55only inference code for
  1803. 1:11:57tokenization um I wanted to show you how
  1804. 1:12:00you would use it quite simple and
  1805. 1:12:02running this just gives us the gpt2
  1806. 1:12:04tokens or the GPT 4 tokens so this is
  1807. 1:12:06the tokenizer use for GPT 4 and so in
  1808. 1:12:09particular we see that the Whit space in
  1809. 1:12:11gpt2 remains unmerged but in GPT 4 uh
  1810. 1:12:14these Whit spaces merge as we also saw
  1811. 1:12:17in this one where here they're all
  1812. 1:12:19unmerged but if we go down to GPT 4 uh
  1813. 1:12:22they become merged
  1814. 1:12:25um now in the
  1815. 1:12:27gp4 uh tokenizer they changed the
  1816. 1:12:31regular expression that they use to
  1817. 1:12:33Chunk Up text so the way to see this is
  1818. 1:12:35that if you come to your the Tik token
  1819. 1:12:38uh library and then you go to this file
  1820. 1:12:41Tik token X openi public this is where
  1821. 1:12:44sort of like the definition of all these
  1822. 1:12:45different tokenizers that openi
  1823. 1:12:46maintains is and so uh necessarily to do
  1824. 1:12:50the inference they had to publish some
  1825. 1:12:51of the details about the strings
  1826. 1:12:53so this is the string that we already
  1827. 1:12:55saw for gpt2 it is slightly different
  1828. 1:12:58but it is actually equivalent uh to what
  1829. 1:13:00we discussed here so this pattern that
  1830. 1:13:02we discussed is equivalent to this
  1831. 1:13:04pattern this one just executes a little
  1832. 1:13:07bit faster so here you see a little bit
  1833. 1:13:09of a slightly different definition but
  1834. 1:13:10otherwise it's the same we're going to
  1835. 1:13:12go into special tokens in a bit and then
  1836. 1:13:15if you scroll down to CL 100k this is
  1837. 1:13:18the GPT 4 tokenizer you see that the
  1838. 1:13:20pattern has changed um and this is kind
  1839. 1:13:23of like the main the major change in
  1840. 1:13:26addition to a bunch of other special
  1841. 1:13:27tokens which I'll go into in a bit again
  1842. 1:13:30now some I'm not going to actually go
  1843. 1:13:31into the full detail of the pattern
  1844. 1:13:33change because honestly this is my
  1845. 1:13:35numbing uh I would just advise that you
  1846. 1:13:37pull out chat GPT and the regex
  1847. 1:13:39documentation and just step through it
  1848. 1:13:42but really the major changes are number
  1849. 1:13:44one you see this eye here that means
  1850. 1:13:48that the um case sensitivity this is
  1851. 1:13:51case insensitive match and so the
  1852. 1:13:53comment that we saw earlier on oh we
  1853. 1:13:56should have used re. uppercase uh
  1854. 1:13:58basically we're now going to be matching
  1855. 1:14:01these apostrophe s apostrophe D
  1856. 1:14:04apostrophe M Etc uh we're going to be
  1857. 1:14:06matching them both in lowercase and in
  1858. 1:14:08uppercase so that's fixed there's a
  1859. 1:14:11bunch of different like handling of the
  1860. 1:14:12whites space that I'm not going to go
  1861. 1:14:14into the full details of and then one
  1862. 1:14:16more thing here is you will notice that
  1863. 1:14:18when they match the numbers they only
  1864. 1:14:20match one to three numbers so so they
  1865. 1:14:23will never merge
  1866. 1:14:26numbers that are in low in more than
  1867. 1:14:28three digits only up to three digits of
  1868. 1:14:31numbers will ever be merged and uh
  1869. 1:14:34that's one change that they made as well
  1870. 1:14:36to prevent uh tokens that are very very
  1871. 1:14:38long number
  1872. 1:14:40sequences uh but again we don't really
  1873. 1:14:42know why they do any of this stuff uh
  1874. 1:14:44because none of this is documented and
  1875. 1:14:46uh it's just we just get the pattern so
  1876. 1:14:49um yeah it is what it is but those are
  1877. 1:14:51some of the changes that gp4 has made
  1878. 1:14:54and of course the vocabulary size went
  1879. 1:14:56from roughly 50k to roughly
  1880. 1:14:58100K the next thing I would like to do
  1881. 1:15:00very briefly is to take you through the
  1882. 1:15:02gpt2 encoder dopy that openi has
  1883. 1:15:05released uh this is the file that I
  1884. 1:15:07already mentioned to you briefly now
  1885. 1:15:09this file is uh fairly short and should
  1886. 1:15:12be relatively understandable to you at
  1887. 1:15:14this point um starting at the bottom
  1888. 1:15:17here they are loading two files encoder
  1889. 1:15:21Json and vocab bpe and they do some
  1890. 1:15:24light processing on it and then they
  1891. 1:15:25call this encoder object which is the
  1892. 1:15:27tokenizer now if you'd like to inspect
  1893. 1:15:30these two files which together
  1894. 1:15:31constitute their saved tokenizer then
  1895. 1:15:34you can do that with a piece of code
  1896. 1:15:36like
  1897. 1:15:36this um this is where you can download
  1898. 1:15:39these two files and you can inspect them
  1899. 1:15:40if you'd like and what you will find is
  1900. 1:15:42that this encoder as they call it in
  1901. 1:15:45their code is exactly equivalent to our
  1902. 1:15:47vocab so remember here where we have
  1903. 1:15:51this vocab object which allowed us us to
  1904. 1:15:53decode very efficiently and basically it
  1905. 1:15:56took us from the integer to the byes uh
  1906. 1:16:00for that integer so our vocab is exactly
  1907. 1:16:03their encoder and then their vocab bpe
  1908. 1:16:07confusingly is actually are merges so
  1909. 1:16:11their BP merges which is based on the
  1910. 1:16:14data inside vocab bpe ends up being
  1911. 1:16:16equivalent to our merges so uh basically
  1912. 1:16:20they are saving and loading the two uh
  1913. 1:16:24variables that for us are also critical
  1914. 1:16:26the merges variable and the vocab
  1915. 1:16:28variable using just these two variables
  1916. 1:16:31you can represent a tokenizer and you
  1917. 1:16:32can both do encoding and decoding once
  1918. 1:16:34you've trained this
  1919. 1:16:36tokenizer now the only thing that um is
  1920. 1:16:40actually slightly confusing inside what
  1921. 1:16:42opening ey does here is that in addition
  1922. 1:16:44to this encoder and a decoder they also
  1923. 1:16:46have something called a bite encoder and
  1924. 1:16:48a bite decoder and this is actually
  1925. 1:16:51unfortunately just
  1926. 1:16:53kind of a spirous implementation detail
  1927. 1:16:55and isn't actually deep or interesting
  1928. 1:16:57in any way so I'm going to skip the
  1929. 1:16:59discussion of it but what opening ey
  1930. 1:17:01does here for reasons that I don't fully
  1931. 1:17:02understand is that not only have they
  1932. 1:17:05this tokenizer which can encode and
  1933. 1:17:06decode but they have a whole separate
  1934. 1:17:08layer here in addition that is used
  1935. 1:17:10serially with the tokenizer and so you
  1936. 1:17:12first do um bite encode and then encode
  1937. 1:17:16and then you do decode and then bite
  1938. 1:17:17decode so that's the loop and they are
  1939. 1:17:20just stacked serial on top of each other
  1940. 1:17:22and and it's not that interesting so I
  1941. 1:17:24won't cover it and you can step through
  1942. 1:17:25it if you'd like otherwise this file if
  1943. 1:17:28you ignore the bite encoder and the bite
  1944. 1:17:30decoder will be algorithmically very
  1945. 1:17:31familiar with you and the meat of it
  1946. 1:17:33here is the what they call bpe function
  1947. 1:17:37and you should recognize this Loop here
  1948. 1:17:39which is very similar to our own y Loop
  1949. 1:17:41where they're trying to identify the
  1950. 1:17:43Byram uh a pair that they should be
  1951. 1:17:46merging next and then here just like we
  1952. 1:17:49had they have a for Loop trying to merge
  1953. 1:17:50this pair uh so they will go over all of
  1954. 1:17:53the sequence and they will merge the
  1955. 1:17:55pair whenever they find it and they keep
  1956. 1:17:57repeating that until they run out of
  1957. 1:17:59possible merges in the in the text so
  1958. 1:18:02that's the meat of this file and uh
  1959. 1:18:04there's an encode and a decode function
  1960. 1:18:06just like we have implemented it so long
  1961. 1:18:08story short what I want you to take away
  1962. 1:18:09at this point is that unfortunately it's
  1963. 1:18:11a little bit of a messy code that they
  1964. 1:18:13have but algorithmically it is identical
  1965. 1:18:15to what we've built up above and what
  1966. 1:18:17we've built up above if you understand
  1967. 1:18:19it is algorithmically what is necessary
  1968. 1:18:21to actually build a BP to organizer
  1969. 1:18:23train it and then both encode and decode
  1970. 1:18:26the next topic I would like to turn to
  1971. 1:18:28is that of special tokens so in addition
  1972. 1:18:30to tokens that are coming from you know
  1973. 1:18:32raw bytes and the BP merges we can
  1974. 1:18:35insert all kinds of tokens that we are
  1975. 1:18:36going to use to delimit different parts
  1976. 1:18:38of the data or introduced to create a
  1977. 1:18:41special structure of the token streams
  1978. 1:18:44so in uh if you look at this encoder
  1979. 1:18:47object from open AIS gpd2 right here we
  1980. 1:18:50mentioned this is very similar to our
  1981. 1:18:52vocab you'll notice that the length of
  1982. 1:18:54this is
  1983. 1:18:5850257 and as I mentioned it's mapping uh
  1984. 1:19:01and it's inverted from the mapping of
  1985. 1:19:03our vocab our vocab goes from integer to
  1986. 1:19:06string and they go the other way around
  1987. 1:19:08for no amazing reason um but the thing
  1988. 1:19:11to note here is that this the mapping
  1989. 1:19:13table here is
  1990. 1:19:1550257 where does that number come from
  1991. 1:19:18where what are the tokens as I mentioned
  1992. 1:19:20there are 256 raw bite token
  1993. 1:19:24tokens and then opena actually did
  1994. 1:19:2750,000
  1995. 1:19:28merges so those become the other tokens
  1996. 1:19:32but this would have been
  1997. 1:19:3450256 so what is the 57th token and
  1998. 1:19:37there is basically one special
  1999. 1:19:40token and that one special token you can
  2000. 1:19:43see is called end of text so this is a
  2001. 1:19:47special token and it's the very last
  2002. 1:19:49token and this token is used to delimit
  2003. 1:19:52documents ments in the training set so
  2004. 1:19:55when we're creating the training data we
  2005. 1:19:57have all these documents and we tokenize
  2006. 1:19:59them and we get a stream of tokens those
  2007. 1:20:01tokens only range from Z to
  2008. 1:20:0550256 and then in between those
  2009. 1:20:07documents we put special end of text
  2010. 1:20:10token and we insert that token in
  2011. 1:20:12between documents and we are using this
  2012. 1:20:15as a signal to the language model that
  2013. 1:20:18the document has ended and what follows
  2014. 1:20:20is going to be unrelated to the document
  2015. 1:20:23previously that said the language model
  2016. 1:20:25has to learn this from data it it needs
  2017. 1:20:27to learn that this token usually means
  2018. 1:20:29that it should wipe its sort of memory
  2019. 1:20:31of what came before and what came before
  2020. 1:20:34this token is not actually informative
  2021. 1:20:35to what comes next but we are expecting
  2022. 1:20:37the language model to just like learn
  2023. 1:20:39this but we're giving it the Special
  2024. 1:20:40sort of the limiter of these documents
  2025. 1:20:44we can go here to Tech tokenizer and um
  2026. 1:20:46this the gpt2 tokenizer uh our code that
  2027. 1:20:49we've been playing with before so we can
  2028. 1:20:51add here right hello world world how are
  2029. 1:20:53you and we're getting different tokens
  2030. 1:20:55but now you can see what if what happens
  2031. 1:20:58if I put end of text you see how until I
  2032. 1:21:02finished it these are all different
  2033. 1:21:03tokens end of
  2034. 1:21:06text still set different tokens and now
  2035. 1:21:08when I finish it suddenly we get token
  2036. 1:21:1350256 and the reason this works is
  2037. 1:21:15because this didn't actually go through
  2038. 1:21:18the bpe merges instead the code that
  2039. 1:21:21actually outposted tokens has special
  2040. 1:21:25case instructions for handling special
  2041. 1:21:28tokens um we did not see these special
  2042. 1:21:30instructions for handling special tokens
  2043. 1:21:32in the encoder dopy it's absent there
  2044. 1:21:36but if you go to Tech token Library
  2045. 1:21:38which is uh implemented in Rust you will
  2046. 1:21:40find all kinds of special case handling
  2047. 1:21:42for these special tokens that you can
  2048. 1:21:44register uh create adds to the
  2049. 1:21:47vocabulary and then it looks for them
  2050. 1:21:49and it uh whenever it sees these special
  2051. 1:21:50tokens like this it will actually come
  2052. 1:21:53in and swap in that special token so
  2053. 1:21:56these things are outside of the typical
  2054. 1:21:58algorithm of uh B PA en
  2055. 1:22:00coding so these special tokens are used
  2056. 1:22:02pervasively uh not just in uh basically
  2057. 1:22:05base language modeling of predicting the
  2058. 1:22:07next token in the sequence but
  2059. 1:22:09especially when it gets to later to the
  2060. 1:22:10fine tuning stage and all of the chat uh
  2061. 1:22:13gbt sort of aspects of it uh because we
  2062. 1:22:15don't just want to Del limit documents
  2063. 1:22:16we want to delimit entire conversations
  2064. 1:22:18between an assistant and a user so if I
  2065. 1:22:21refresh this sck tokenizer page the
  2066. 1:22:24default example that they have here is
  2067. 1:22:26using not sort of base model encoders
  2068. 1:22:30but ftuned model uh sort of tokenizers
  2069. 1:22:33um so for example using the GPT 3.5
  2070. 1:22:35turbo scheme these here are all special
  2071. 1:22:38tokens I am start I end Etc uh this is
  2072. 1:22:43short for Imaginary mcore start by the
  2073. 1:22:46way but you can see here that there's a
  2074. 1:22:49sort of start and end of every single
  2075. 1:22:51message and there can be many other
  2076. 1:22:52other tokens lots of tokens um in use to
  2077. 1:22:56delimit these conversations and kind of
  2078. 1:22:58keep track of the flow of the messages
  2079. 1:23:00here now we can go back to the Tik token
  2080. 1:23:03library and here when you scroll to the
  2081. 1:23:06bottom they talk about how you can
  2082. 1:23:08extend tick token and I can you can
  2083. 1:23:10create basically you can Fork uh the um
  2084. 1:23:13CL 100K base tokenizers in gp4 and for
  2085. 1:23:17example you can extend it by adding more
  2086. 1:23:18special tokens and these are totally up
  2087. 1:23:20to you you can come up with any
  2088. 1:23:21arbitrary tokens and add them with the
  2089. 1:23:23new ID afterwards and the tikken library
  2090. 1:23:26will uh correctly swap them out uh when
  2091. 1:23:29it sees this in the
  2092. 1:23:31strings now we can also go back to this
  2093. 1:23:34file which we've looked at previously
  2094. 1:23:37and I mentioned that the gpt2 in Tik
  2095. 1:23:39toen open
  2096. 1:23:41I.P we have the vocabulary we have the
  2097. 1:23:44pattern for splitting and then here we
  2098. 1:23:46are registering the single special token
  2099. 1:23:48in gpd2 which was the end of text token
  2100. 1:23:50and we saw that it has this ID
  2101. 1:23:53in GPT 4 when they defy this here you
  2102. 1:23:56see that the pattern has changed as
  2103. 1:23:57we've discussed but also the special
  2104. 1:23:59tokens have changed in this tokenizer so
  2105. 1:24:01we of course have the end of text just
  2106. 1:24:03like in gpd2 but we also see three sorry
  2107. 1:24:06four additional tokens here Thim prefix
  2108. 1:24:09middle and suffix what is fim fim is
  2109. 1:24:12short for fill in the middle and if
  2110. 1:24:14you'd like to learn more about this idea
  2111. 1:24:17it comes from this paper um and I'm not
  2112. 1:24:20going to go into detail in this video
  2113. 1:24:21it's beyond this video and then there's
  2114. 1:24:23one additional uh serve token here so
  2115. 1:24:27that's that encoding as well so it's
  2116. 1:24:29very common basically to train a
  2117. 1:24:31language model and then if you'd like uh
  2118. 1:24:34you can add special tokens now when you
  2119. 1:24:37add special tokens you of course have to
  2120. 1:24:39um do some model surgery to the
  2121. 1:24:41Transformer and all the parameters
  2122. 1:24:43involved in that Transformer because you
  2123. 1:24:45are basically adding an integer and you
  2124. 1:24:47want to make sure that for example your
  2125. 1:24:48embedding Matrix for the vocabulary
  2126. 1:24:50tokens has to be extended by adding a
  2127. 1:24:53row and typically this row would be
  2128. 1:24:54initialized uh with small random numbers
  2129. 1:24:56or something like that because we need
  2130. 1:24:58to have a vector that now stands for
  2131. 1:25:01that token in addition to that you have
  2132. 1:25:03to go to the final layer of the
  2133. 1:25:04Transformer and you have to make sure
  2134. 1:25:05that that projection at the very end
  2135. 1:25:07into the classifier uh is extended by
  2136. 1:25:09one as well so basically there's some
  2137. 1:25:11model surgery involved that you have to
  2138. 1:25:13couple with the tokenization changes if
  2139. 1:25:16you are going to add special tokens but
  2140. 1:25:18this is a very common operation that
  2141. 1:25:20people do especially if they'd like to
  2142. 1:25:21fine tune the model for example taking
  2143. 1:25:23it from a base model to a chat model
  2144. 1:25:26like chat
  2145. 1:25:27GPT okay so at this point you should
  2146. 1:25:29have everything you need in order to
  2147. 1:25:31build your own gp4 tokenizer now in the
  2148. 1:25:33process of developing this lecture I've
  2149. 1:25:35done that and I published the code under
  2150. 1:25:37this repository
  2151. 1:25:38MBP so MBP looks like this right now as
  2152. 1:25:42I'm recording but uh the MBP repository
  2153. 1:25:45will probably change quite a bit because
  2154. 1:25:46I intend to continue working on it um in
  2155. 1:25:49addition to the MBP repository I've
  2156. 1:25:51published the this uh exercise
  2157. 1:25:53progression that you can follow so if
  2158. 1:25:55you go to exercise. MD here uh this is
  2159. 1:25:58sort of me breaking up the task ahead of
  2160. 1:26:01you into four steps that sort of uh
  2161. 1:26:03build up to what can be a gp4 tokenizer
  2162. 1:26:06and so feel free to follow these steps
  2163. 1:26:08exactly and follow a little bit of the
  2164. 1:26:10guidance that I've laid out here and
  2165. 1:26:12anytime you feel stuck just reference
  2166. 1:26:14the MBP repository here so either the
  2167. 1:26:17tests could be useful or the MBP
  2168. 1:26:20repository itself I try to keep the code
  2169. 1:26:22fairly clean and understandable and so
  2170. 1:26:26um feel free to reference it whenever um
  2171. 1:26:28you get
  2172. 1:26:30stuck uh in addition to that basically
  2173. 1:26:32once you write it you should be able to
  2174. 1:26:34reproduce this behavior from Tech token
  2175. 1:26:36so getting the gb4 tokenizer you can
  2176. 1:26:39take uh you can encode the string and
  2177. 1:26:41you should get these tokens and then you
  2178. 1:26:43can encode and decode the exact same
  2179. 1:26:44string to recover it and in addition to
  2180. 1:26:47all that you should be able to implement
  2181. 1:26:48your own train function uh which Tik
  2182. 1:26:50token Library does not provide it's it's
  2183. 1:26:52again only inference code but you could
  2184. 1:26:54write your own train MBP does it as well
  2185. 1:26:57and that will allow you to train your
  2186. 1:26:59own token
  2187. 1:27:00vocabularies so here are some of the
  2188. 1:27:02code inside M be mean bpe uh shows the
  2189. 1:27:06token vocabularies that you might obtain
  2190. 1:27:08so on the left uh here we have the GPT 4
  2191. 1:27:12merges uh so the first 256 are raw
  2192. 1:27:15individual bytes and then here I am
  2193. 1:27:17visualizing the merges that gp4
  2194. 1:27:19performed during its training so the
  2195. 1:27:21very first merge that gp4 did was merge
  2196. 1:27:24two spaces into a single token for you
  2197. 1:27:27know two spaces and that is a token 256
  2198. 1:27:30and so this is the order in which things
  2199. 1:27:32merged during gb4 training and this is
  2200. 1:27:34the merge order that um we obtain in MBP
  2201. 1:27:39by training a tokenizer and in this case
  2202. 1:27:41I trained it on a Wikipedia page of
  2203. 1:27:43Taylor Swift uh not because I'm a Swifty
  2204. 1:27:45but because that is one of the longest
  2205. 1:27:47um Wikipedia Pages apparently that's
  2206. 1:27:49available but she is pretty cool and
  2207. 1:27:54um what was I going to say yeah so you
  2208. 1:27:56can compare these two uh vocabularies
  2209. 1:27:59and so as an example um here GPT for
  2210. 1:28:04merged I in to become in and we've done
  2211. 1:28:06the exact same thing on this token 259
  2212. 1:28:10here space t becomes space t and that
  2213. 1:28:13happened for us a little bit later as
  2214. 1:28:14well so the difference here is again to
  2215. 1:28:16my understanding only a difference of
  2216. 1:28:18the training set so as an example
  2217. 1:28:20because I see a lot of white space I
  2218. 1:28:22supect that gp4 probably had a lot of
  2219. 1:28:23python code in its training set I'm not
  2220. 1:28:25sure uh for the
  2221. 1:28:27tokenizer and uh here we see much less
  2222. 1:28:30of that of course in the Wikipedia page
  2223. 1:28:32so roughly speaking they look the same
  2224. 1:28:34and they look the same because they're
  2225. 1:28:35running the same algorithm and when you
  2226. 1:28:38train your own you're probably going to
  2227. 1:28:39get something similar depending on what
  2228. 1:28:41you train it on okay so we are now going
  2229. 1:28:43to move on from tick token and the way
  2230. 1:28:45that open AI tokenizes its strings and
  2231. 1:28:47we're going to discuss one more very
  2232. 1:28:49commonly used library for working with
  2233. 1:28:51tokenization inlm
  2234. 1:28:52and that is sentence piece so sentence
  2235. 1:28:55piece is very commonly used in language
  2236. 1:28:58models because unlike Tik token it can
  2237. 1:29:00do both training and inference and is
  2238. 1:29:02quite efficient at both it supports a
  2239. 1:29:04number of algorithms for training uh
  2240. 1:29:06vocabularies but one of them is the B
  2241. 1:29:09pair en coding algorithm that we've been
  2242. 1:29:10looking at so it supports it now
  2243. 1:29:13sentence piece is used both by llama and
  2244. 1:29:15mistal series and many other models as
  2245. 1:29:18well it is on GitHub under Google
  2246. 1:29:20sentence piece
  2247. 1:29:22and the big difference with sentence
  2248. 1:29:24piece and we're going to look at example
  2249. 1:29:26because this is kind of hard and subtle
  2250. 1:29:27to explain is that they think different
  2251. 1:29:31about the order of operations here so in
  2252. 1:29:35the case of Tik token we first take our
  2253. 1:29:38code points in the string we encode them
  2254. 1:29:41using mutf to bytes and then we're
  2255. 1:29:42merging bytes it's fairly
  2256. 1:29:44straightforward for sentence piece um it
  2257. 1:29:48works directly on the level of the code
  2258. 1:29:50points themselves so so it looks at
  2259. 1:29:52whatever code points are available in
  2260. 1:29:53your training set and then it starts
  2261. 1:29:55merging those code points and um the bpe
  2262. 1:29:59is running on the level of code
  2263. 1:30:01points and if you happen to run out of
  2264. 1:30:04code points so there are maybe some rare
  2265. 1:30:06uh code points that just don't come up
  2266. 1:30:08too often and the Rarity is determined
  2267. 1:30:09by this character coverage hyper
  2268. 1:30:11parameter then these uh code points will
  2269. 1:30:14either get mapped to a special unknown
  2270. 1:30:16token like ank or if you have the bite
  2271. 1:30:19foldback option turned on then that will
  2272. 1:30:22take those rare Cod points it will
  2273. 1:30:23encode them using utf8 and then the
  2274. 1:30:26individual bytes of that encoding will
  2275. 1:30:27be translated into tokens and there are
  2276. 1:30:30these special bite tokens that basically
  2277. 1:30:32get added to the vocabulary so it uses
  2278. 1:30:35BP on on the code points and then it
  2279. 1:30:38falls back to bytes for rare Cod points
  2280. 1:30:41um and so that's kind of like difference
  2281. 1:30:44personally I find the Tik token we
  2282. 1:30:45significantly cleaner uh but it's kind
  2283. 1:30:47of like a subtle but pretty major
  2284. 1:30:48difference between the way they approach
  2285. 1:30:50tokenization let's work with with a
  2286. 1:30:52concrete example because otherwise this
  2287. 1:30:54is kind of hard to um to get your head
  2288. 1:30:56around so let's work with a concrete
  2289. 1:30:59example this is how we can import
  2290. 1:31:01sentence piece and then here we're going
  2291. 1:31:03to take I think I took like the
  2292. 1:31:05description of sentence piece and I just
  2293. 1:31:06created like a little toy data set it
  2294. 1:31:08really likes to have a file so I created
  2295. 1:31:10a toy. txt file with this
  2296. 1:31:13content now what's kind of a little bit
  2297. 1:31:15crazy about sentence piece is that
  2298. 1:31:16there's a ton of options and
  2299. 1:31:18configurations and the reason this is so
  2300. 1:31:20is because sentence piece has been
  2301. 1:31:22around I think for a while and it really
  2302. 1:31:23tries to handle a large diversity of
  2303. 1:31:25things and um because it's been around I
  2304. 1:31:28think it has quite a bit of accumulated
  2305. 1:31:30historical baggage uh as well and so in
  2306. 1:31:33particular there's like a ton of
  2307. 1:31:35configuration arguments this is not even
  2308. 1:31:36all of it you can go to here to see all
  2309. 1:31:39the training
  2310. 1:31:40options um and uh there's also quite
  2311. 1:31:44useful documentation when you look at
  2312. 1:31:45the raw Proto buff uh that is used to
  2313. 1:31:48represent the trainer spec and so on um
  2314. 1:31:52many of these options are irrelevant to
  2315. 1:31:54us so maybe to point out one example Das
  2316. 1:31:56Das shrinking Factor uh this shrinking
  2317. 1:31:59factor is not used in the B pair en
  2318. 1:32:01coding algorithm so this is just an
  2319. 1:32:03argument that is irrelevant to us um it
  2320. 1:32:05applies to a different training
  2321. 1:32:09algorithm now what I tried to do here is
  2322. 1:32:11I tried to set up sentence piece in a
  2323. 1:32:13way that is very very similar as far as
  2324. 1:32:15I can tell to maybe identical hopefully
  2325. 1:32:18to the way that llama 2 was strained so
  2326. 1:32:22the way they trained their own um their
  2327. 1:32:25own tokenizer and the way I did this was
  2328. 1:32:27basically you can take the tokenizer
  2329. 1:32:28model file that meta released and you
  2330. 1:32:31can um open it using the Proto protuff
  2331. 1:32:35uh sort of file that you can generate
  2332. 1:32:38and then you can inspect all the options
  2333. 1:32:39and I tried to copy over all the options
  2334. 1:32:41that looked relevant so here we set up
  2335. 1:32:43the input it's raw text in this file
  2336. 1:32:46here's going to be the output so it's
  2337. 1:32:48going to be for talk 400. model and
  2338. 1:32:50vocab
  2339. 1:32:52we're saying that we're going to use the
  2340. 1:32:53BP algorithm and we want to Bap size of
  2341. 1:32:56400 then there's a ton of configurations
  2342. 1:32:58here
  2343. 1:33:01for um for basically pre-processing and
  2344. 1:33:05normalization rules as they're called
  2345. 1:33:07normalization used to be very prevalent
  2346. 1:33:09I would say before llms in natural
  2347. 1:33:11language processing so in machine
  2348. 1:33:12translation and uh text classification
  2349. 1:33:14and so on you want to normalize and
  2350. 1:33:16simplify the text and you want to turn
  2351. 1:33:18it all lowercase and you want to remove
  2352. 1:33:19all double whites space Etc
  2353. 1:33:22and in language models we prefer not to
  2354. 1:33:23do any of it or at least that is my
  2355. 1:33:25preference as a deep learning person you
  2356. 1:33:26want to not touch your data you want to
  2357. 1:33:28keep the raw data as much as possible um
  2358. 1:33:31in a raw
  2359. 1:33:33form so you're basically trying to turn
  2360. 1:33:35off a lot of this if you can the other
  2361. 1:33:38thing that sentence piece does is that
  2362. 1:33:39it has this concept of sentences so
  2363. 1:33:43sentence piece it's back it's kind of
  2364. 1:33:45like was developed I think early in the
  2365. 1:33:46days where there was um an idea that
  2366. 1:33:50they you're training a tokenizer on a
  2367. 1:33:51bunch of independent sentences so it has
  2368. 1:33:54a lot of like how many sentences you're
  2369. 1:33:56going to train on what is the maximum
  2370. 1:33:58sentence length
  2371. 1:34:00um shuffling sentences and so for it
  2372. 1:34:03sentences are kind of like the
  2373. 1:34:04individual training examples but again
  2374. 1:34:06in the context of llms I find that this
  2375. 1:34:08is like a very spous and weird
  2376. 1:34:10distinction like sentences are just like
  2377. 1:34:13don't touch the raw data sentences
  2378. 1:34:15happen to exist but in raw data sets
  2379. 1:34:18there are a lot of like inet like what
  2380. 1:34:20exactly is a sentence what isn't a
  2381. 1:34:22sentence um and so I think like it's
  2382. 1:34:25really hard to Define what an actual
  2383. 1:34:26sentence is if you really like dig into
  2384. 1:34:28it and there could be different concepts
  2385. 1:34:30of it in different languages or
  2386. 1:34:32something like that so why even
  2387. 1:34:33introduce the concept it it doesn't
  2388. 1:34:35honestly make sense to me I would just
  2389. 1:34:36prefer to treat a file as a giant uh
  2390. 1:34:39stream of
  2391. 1:34:40bytes it has a lot of treatment around
  2392. 1:34:42rare word characters and when I say word
  2393. 1:34:45I mean code points we're going to come
  2394. 1:34:46back to this in a second and it has a
  2395. 1:34:48lot of other rules for um basically
  2396. 1:34:51splitting digits splitting white space
  2397. 1:34:54and numbers and how you deal with that
  2398. 1:34:56so these are some kind of like merge
  2399. 1:34:58rules so I think this is a little bit
  2400. 1:35:00equivalent to tick token using the
  2401. 1:35:02regular expression to split up
  2402. 1:35:04categories there's like kind of
  2403. 1:35:07equivalence of it if you squint T it in
  2404. 1:35:09sentence piece where you can also for
  2405. 1:35:10example split up split up the digits uh
  2406. 1:35:14and uh so
  2407. 1:35:15on there's a few more things here that
  2408. 1:35:18I'll come back to in a bit and then
  2409. 1:35:19there are some special tokens that you
  2410. 1:35:20can indicate and it hardcodes the UN
  2411. 1:35:23token the beginning of sentence end of
  2412. 1:35:25sentence and a pad token um and the UN
  2413. 1:35:29token must exist for my understanding
  2414. 1:35:32and then some some things so we can
  2415. 1:35:34train and when when I press train it's
  2416. 1:35:37going to create this file talk 400.
  2417. 1:35:40model and talk 400. wab I can then load
  2418. 1:35:43the model file and I can inspect the
  2419. 1:35:45vocabulary off it and so we trained
  2420. 1:35:48vocab size 400 on this text here and
  2421. 1:35:53these are the individual pieces the
  2422. 1:35:55individual tokens that sentence piece
  2423. 1:35:56will create so in the beginning we see
  2424. 1:35:58that we have the an token uh with the ID
  2425. 1:36:02zero then we have the beginning of
  2426. 1:36:04sequence end of sequence one and two and
  2427. 1:36:07then we said that the pad ID is negative
  2428. 1:36:091 so we chose not to use it so there's
  2429. 1:36:12no pad ID
  2430. 1:36:13here then these are individual bite
  2431. 1:36:16tokens so here we saw that bite fallback
  2432. 1:36:20in llama was turned on so it's true so
  2433. 1:36:23what follows are going to be the 256
  2434. 1:36:26bite
  2435. 1:36:27tokens and these are their
  2436. 1:36:31IDs and then at the bottom after the
  2437. 1:36:35bite tokens come the
  2438. 1:36:37merges and these are the parent nodes in
  2439. 1:36:40the merges so we're not seeing the
  2440. 1:36:42children we're just seeing the parents
  2441. 1:36:43and their
  2442. 1:36:44ID and then after the
  2443. 1:36:47merges comes eventually the individual
  2444. 1:36:50tokens and their IDs and so these are
  2445. 1:36:53the individual tokens so these are the
  2446. 1:36:55individual code Point tokens if you will
  2447. 1:36:58and they come at the end so that is the
  2448. 1:37:00ordering with which sentence piece sort
  2449. 1:37:01of like represents its vocabularies it
  2450. 1:37:03starts with special tokens then the bike
  2451. 1:37:06tokens then the merge tokens and then
  2452. 1:37:08the individual codo tokens and all these
  2453. 1:37:11raw codepoint to tokens are the ones
  2454. 1:37:14that it encountered in the training
  2455. 1:37:16set so those individual code points are
  2456. 1:37:19all the the entire set of code points
  2457. 1:37:22that occurred
  2458. 1:37:24here so those all get put in there and
  2459. 1:37:27then those that are extremely rare as
  2460. 1:37:29determined by character coverage so if a
  2461. 1:37:31code Point occurred only a single time
  2462. 1:37:32out of like a million um sentences or
  2463. 1:37:35something like that then it would be
  2464. 1:37:37ignored and it would not be added to our
  2465. 1:37:40uh
  2466. 1:37:41vocabulary once we have a vocabulary we
  2467. 1:37:43can encode into IDs and we can um sort
  2468. 1:37:46of get a
  2469. 1:37:47list and then here I am also decoding
  2470. 1:37:50the indiv idual tokens back into little
  2471. 1:37:54pieces as they call it so let's take a
  2472. 1:37:56look at what happened here hello space
  2473. 1:38:01on so these are the token IDs we got
  2474. 1:38:04back and when we look here uh a few
  2475. 1:38:07things sort of uh jump to mind number
  2476. 1:38:11one take a look at these characters the
  2477. 1:38:14Korean characters of course were not
  2478. 1:38:15part of the training set so sentence
  2479. 1:38:18piece is encountering code points that
  2480. 1:38:19it has not seen during training time and
  2481. 1:38:22those code points do not have a token
  2482. 1:38:24associated with them so suddenly these
  2483. 1:38:26are un tokens unknown tokens but because
  2484. 1:38:30bite fall back as true instead sentence
  2485. 1:38:33piece falls back to bytes and so it
  2486. 1:38:36takes this it encodes it with utf8 and
  2487. 1:38:39then it uses these tokens to represent
  2488. 1:38:43uh those bytes and that's what we are
  2489. 1:38:45getting sort of here this is the utf8 uh
  2490. 1:38:49encoding and in this shifted by three uh
  2491. 1:38:52because of these um special tokens here
  2492. 1:38:56that have IDs earlier on so that's what
  2493. 1:38:58happened here now one more thing that um
  2494. 1:39:02well first before I go on with respect
  2495. 1:39:05to the bitef back let me remove bite
  2496. 1:39:08foldback if this is false what's going
  2497. 1:39:10to happen let's
  2498. 1:39:12retrain so the first thing that happened
  2499. 1:39:14is all the bite tokens disappeared right
  2500. 1:39:17and now we just have the merges and we
  2501. 1:39:19have a lot more merges now because we
  2502. 1:39:20have a lot more space because we're not
  2503. 1:39:21taking up space in the wab size uh with
  2504. 1:39:25all the
  2505. 1:39:25bytes and now if we encode
  2506. 1:39:29this we get a zero so this entire string
  2507. 1:39:33here suddenly there's no bitef back so
  2508. 1:39:35this is unknown and unknown is an and so
  2509. 1:39:39this is zero because the an token is
  2510. 1:39:42token zero and you have to keep in mind
  2511. 1:39:44that this would feed into your uh
  2512. 1:39:46language model so what is a language
  2513. 1:39:48model supposed to do when all kinds of
  2514. 1:39:49different things that are unrecognized
  2515. 1:39:52because they're rare just end up mapping
  2516. 1:39:54into Unk it's not exactly the property
  2517. 1:39:56that you want so that's why I think
  2518. 1:39:57llama correctly uh used by fallback true
  2519. 1:40:02uh because we definitely want to feed
  2520. 1:40:03these um unknown or rare code points
  2521. 1:40:06into the model and some uh some manner
  2522. 1:40:08the next thing I want to show you is the
  2523. 1:40:10following notice here when we are
  2524. 1:40:12decoding all the individual tokens you
  2525. 1:40:14see how spaces uh space here ends up
  2526. 1:40:18being this um bold underline I'm not
  2527. 1:40:21100% sure by the way why sentence piece
  2528. 1:40:23switches whites space into these bold
  2529. 1:40:25underscore characters maybe it's for
  2530. 1:40:27visualization I'm not 100% sure why that
  2531. 1:40:29happens uh but notice this why do we
  2532. 1:40:32have an extra space in the front of
  2533. 1:40:37hello um what where is this coming from
  2534. 1:40:40well it's coming from this option
  2535. 1:40:43here
  2536. 1:40:45um add dummy prefix is true and when you
  2537. 1:40:48go to the
  2538. 1:40:49documentation add D whites space at the
  2539. 1:40:51beginning of text in order to treat
  2540. 1:40:53World in world and hello world in the
  2541. 1:40:55exact same way so what this is trying to
  2542. 1:40:57do is the
  2543. 1:40:59following if we go back to our tick
  2544. 1:41:02tokenizer world as uh token by itself
  2545. 1:41:06has a different ID than space world so
  2546. 1:41:10we have this is 1917 but this is 14 Etc
  2547. 1:41:14so these are two different tokens for
  2548. 1:41:16the language model and the language
  2549. 1:41:17model has to learn from data that they
  2550. 1:41:18are actually kind of like a very similar
  2551. 1:41:20concept so to the language model in the
  2552. 1:41:23Tik token World um basically words in
  2553. 1:41:26the beginning of sentences and words in
  2554. 1:41:27the middle of sentences actually look
  2555. 1:41:29completely different um and it has to
  2556. 1:41:32learned that they are roughly the same
  2557. 1:41:34so this add dami prefix is trying to
  2558. 1:41:36fight that a little bit and the way that
  2559. 1:41:38works is that it basically
  2560. 1:41:41uh adds a dummy prefix so for as a as a
  2561. 1:41:46part of pre-processing it will take the
  2562. 1:41:49string and it will add a space it will
  2563. 1:41:51do this and that's done in an effort to
  2564. 1:41:54make this world and that world the same
  2565. 1:41:57they will both be space world so that's
  2566. 1:42:00one other kind of pre-processing option
  2567. 1:42:02that is turned on and llama 2 also uh
  2568. 1:42:05uses this option and that's I think
  2569. 1:42:07everything that I want to say for my
  2570. 1:42:08preview of sentence piece and how it is
  2571. 1:42:10different um maybe here what I've done
  2572. 1:42:13is I just uh put in the Raw protocol
  2573. 1:42:16buffer representation basically of the
  2574. 1:42:19tokenizer the too trained so feel free
  2575. 1:42:22to sort of Step through this and if you
  2576. 1:42:24would like uh your tokenization to look
  2577. 1:42:27identical to that of the meta uh llama 2
  2578. 1:42:30then you would be copy pasting these
  2579. 1:42:31settings as I tried to do up above and
  2580. 1:42:34uh yeah that's I think that's it for
  2581. 1:42:36this section I think my summary for
  2582. 1:42:38sentence piece from all of this is
  2583. 1:42:40number one I think that there's a lot of
  2584. 1:42:42historical baggage in sentence piece a
  2585. 1:42:44lot of Concepts that I think are
  2586. 1:42:45slightly confusing and I think
  2587. 1:42:47potentially um contain foot guns like
  2588. 1:42:49this concept of a sentence and it's
  2589. 1:42:50maximum length and stuff like that um
  2590. 1:42:53otherwise it is fairly commonly used in
  2591. 1:42:55the industry um because it is efficient
  2592. 1:42:58and can do both training and inference
  2593. 1:43:01uh it has a few quirks like for example
  2594. 1:43:02un token must exist and the way the bite
  2595. 1:43:05fallbacks are done and so on I don't
  2596. 1:43:06find particularly elegant and
  2597. 1:43:08unfortunately I have to say it's not
  2598. 1:43:09very well documented so it took me a lot
  2599. 1:43:11of time working with this myself um and
  2600. 1:43:14just visualizing things and trying to
  2601. 1:43:16really understand what is happening here
  2602. 1:43:17because uh the documentation
  2603. 1:43:19unfortunately is in my opion not not
  2604. 1:43:21super amazing but it is a very nice repo
  2605. 1:43:24that is available to you if you'd like
  2606. 1:43:26to train your own tokenizer right now
  2607. 1:43:28okay let me now switch gears again as
  2608. 1:43:29we're starting to slowly wrap up here I
  2609. 1:43:31want to revisit this issue in a bit more
  2610. 1:43:33detail of how we should set the vocap
  2611. 1:43:35size and what are some of the
  2612. 1:43:36considerations around it so for this I'd
  2613. 1:43:39like to go back to the model
  2614. 1:43:40architecture that we developed in the
  2615. 1:43:42last video when we built the GPT from
  2616. 1:43:44scratch so this here was uh the file
  2617. 1:43:47that we built in the previous video and
  2618. 1:43:49we defined the Transformer model and and
  2619. 1:43:51let's specifically look at Bap size and
  2620. 1:43:52where it appears in this file so here we
  2621. 1:43:55Define the voap size uh at this time it
  2622. 1:43:58was 65 or something like that extremely
  2623. 1:43:59small number so this will grow much
  2624. 1:44:02larger you'll see that Bap size doesn't
  2625. 1:44:04come up too much in most of these layers
  2626. 1:44:06the only place that it comes up to is in
  2627. 1:44:08exactly these two places here so when we
  2628. 1:44:11Define the language model there's the
  2629. 1:44:13token embedding table which is this
  2630. 1:44:15two-dimensional array where the vocap
  2631. 1:44:18size is basically the number of rows and
  2632. 1:44:21uh each vocabulary element each token
  2633. 1:44:23has a vector that we're going to train
  2634. 1:44:25using back propagation that Vector is of
  2635. 1:44:27size and embed which is number of
  2636. 1:44:29channels in the Transformer and
  2637. 1:44:31basically as voap size increases this
  2638. 1:44:33embedding table as I mentioned earlier
  2639. 1:44:35is going to also grow we're going to be
  2640. 1:44:37adding rows in addition to that at the
  2641. 1:44:39end of the Transformer there's this LM
  2642. 1:44:41head layer which is a linear layer and
  2643. 1:44:44you'll notice that that layer is used at
  2644. 1:44:46the very end to produce the logits uh
  2645. 1:44:48which become the probabilities for the
  2646. 1:44:49next token in sequence and so
  2647. 1:44:51intuitively we're trying to produce a
  2648. 1:44:53probability for every single token that
  2649. 1:44:56might come next at every point in time
  2650. 1:44:58of that Transformer and if we have more
  2651. 1:45:01and more tokens we need to produce more
  2652. 1:45:02and more probabilities so every single
  2653. 1:45:04token is going to introduce an
  2654. 1:45:06additional dot product that we have to
  2655. 1:45:08do here in this linear layer for this
  2656. 1:45:10final layer in a
  2657. 1:45:11Transformer so why can't vocap size be
  2658. 1:45:14infinite why can't we grow to Infinity
  2659. 1:45:16well number one your token embedding
  2660. 1:45:18table is going to grow uh your linear
  2661. 1:45:21layer is going to grow so we're going to
  2662. 1:45:23be doing a lot more computation here
  2663. 1:45:25because this LM head layer will become
  2664. 1:45:26more computational expensive number two
  2665. 1:45:29because we have more parameters we could
  2666. 1:45:30be worried that we are going to be under
  2667. 1:45:33trining some of these
  2668. 1:45:35parameters so intuitively if you have a
  2669. 1:45:37very large vocabulary size say we have a
  2670. 1:45:38million uh tokens then every one of
  2671. 1:45:41these tokens is going to come up more
  2672. 1:45:42and more rarely in the training data
  2673. 1:45:45because there's a lot more other tokens
  2674. 1:45:46all over the place and so we're going to
  2675. 1:45:48be seeing fewer and fewer examples uh
  2676. 1:45:51for each individual token and you might
  2677. 1:45:53be worried that basically the vectors
  2678. 1:45:55associated with every token will be
  2679. 1:45:56undertrained as a result because they
  2680. 1:45:58just don't come up too often and they
  2681. 1:45:59don't participate in the forward
  2682. 1:46:00backward pass in addition to that as
  2683. 1:46:03your vocab size grows you're going to
  2684. 1:46:04start shrinking your sequences a lot
  2685. 1:46:07right and that's really nice because
  2686. 1:46:09that means that we're going to be
  2687. 1:46:10attending to more and more text so
  2688. 1:46:12that's nice but also you might be
  2689. 1:46:13worrying that two large of chunks are
  2690. 1:46:15being squished into single tokens and so
  2691. 1:46:18the model just doesn't have as much of
  2692. 1:46:20time to think per sort of um some number
  2693. 1:46:25of characters in the text or you can
  2694. 1:46:26think about it that way right so
  2695. 1:46:28basically we're squishing too much
  2696. 1:46:29information into a single token and then
  2697. 1:46:31the forward pass of the Transformer is
  2698. 1:46:33not enough to actually process that
  2699. 1:46:34information appropriately and so these
  2700. 1:46:36are some of the considerations you're
  2701. 1:46:37thinking about when you're designing the
  2702. 1:46:38vocab size as I mentioned this is mostly
  2703. 1:46:40an empirical hyperparameter and it seems
  2704. 1:46:42like in state-of-the-art architectures
  2705. 1:46:44today this is usually in the high 10,000
  2706. 1:46:46or somewhere around 100,000 today and
  2707. 1:46:49the next consideration I want to briefly
  2708. 1:46:50talk about is what if we want to take a
  2709. 1:46:53pre-trained model and we want to extend
  2710. 1:46:55the vocap size and this is done fairly
  2711. 1:46:57commonly actually so for example when
  2712. 1:46:58you're doing fine-tuning for cha GPT um
  2713. 1:47:02a lot more new special tokens get
  2714. 1:47:03introduced on top of the base model to
  2715. 1:47:05maintain the metadata and all the
  2716. 1:47:08structure of conversation objects
  2717. 1:47:09between a user and an assistant so that
  2718. 1:47:11takes a lot of special tokens you might
  2719. 1:47:14also try to throw in more special tokens
  2720. 1:47:15for example for using the browser or any
  2721. 1:47:17other tool and so it's very tempting to
  2722. 1:47:20add a lot of tokens for all kinds of
  2723. 1:47:22special functionality so if you want to
  2724. 1:47:24be adding a token that's totally
  2725. 1:47:25possible Right all we have to do is we
  2726. 1:47:27have to resize this embedding so we have
  2727. 1:47:29to add rows we would initialize these uh
  2728. 1:47:32parameters from scratch to be small
  2729. 1:47:34random numbers and then we have to
  2730. 1:47:36extend the weight inside this linear uh
  2731. 1:47:39so we have to start making dot products
  2732. 1:47:41um with the associated parameters as
  2733. 1:47:43well to basically calculate the
  2734. 1:47:44probabilities for these new tokens so
  2735. 1:47:46both of these are just a resizing
  2736. 1:47:48operation it's a very mild
  2737. 1:47:50model surgery and can be done fairly
  2738. 1:47:52easily and it's quite common that
  2739. 1:47:54basically you would freeze the base
  2740. 1:47:55model you introduce these new parameters
  2741. 1:47:57and then you only train these new
  2742. 1:47:58parameters to introduce new tokens into
  2743. 1:48:00the architecture um and so you can
  2744. 1:48:03freeze arbitrary parts of it or you can
  2745. 1:48:04train arbitrary parts of it and that's
  2746. 1:48:06totally up to you but basically minor
  2747. 1:48:08surgery required if you'd like to
  2748. 1:48:10introduce new tokens and finally I'd
  2749. 1:48:11like to mention that actually there's an
  2750. 1:48:13entire design space of applications in
  2751. 1:48:15terms of introducing new tokens into a
  2752. 1:48:17vocabulary that go Way Beyond just
  2753. 1:48:19adding special tokens and special new
  2754. 1:48:21functionality so just to give you a
  2755. 1:48:23sense of the design space but this could
  2756. 1:48:24be an entire video just by itself uh
  2757. 1:48:26this is a paper on learning to compress
  2758. 1:48:28prompts with what they called uh gist
  2759. 1:48:31tokens and the rough idea is suppose
  2760. 1:48:33that you're using language models in a
  2761. 1:48:34setting that requires very long prompts
  2762. 1:48:37while these long prompts just slow
  2763. 1:48:38everything down because you have to
  2764. 1:48:39encode them and then you have to use
  2765. 1:48:41them and then you're tending over them
  2766. 1:48:43and it's just um you know heavy to have
  2767. 1:48:45very large prompts so instead what they
  2768. 1:48:47do here in this paper is they introduce
  2769. 1:48:50new tokens and um imagine basically
  2770. 1:48:54having a few new tokens you put them in
  2771. 1:48:56a sequence and then you train the model
  2772. 1:48:59by distillation so you are keeping the
  2773. 1:49:01entire model Frozen and you're only
  2774. 1:49:03training the representations of the new
  2775. 1:49:05tokens their embeddings and you're
  2776. 1:49:06optimizing over the new tokens such that
  2777. 1:49:09the behavior of the language model is
  2778. 1:49:11identical uh to the model that has a
  2779. 1:49:15very long prompt that works for you and
  2780. 1:49:17so it's a compression technique of
  2781. 1:49:19compressing that very long prompt into
  2782. 1:49:20those few new gist tokens and so you can
  2783. 1:49:23train this and then at test time you can
  2784. 1:49:25discard your old prompt and just swap in
  2785. 1:49:26those tokens and they sort of like uh
  2786. 1:49:28stand in for that very long prompt and
  2787. 1:49:31have an almost identical performance and
  2788. 1:49:33so this is one um technique and a class
  2789. 1:49:36of parameter efficient fine-tuning
  2790. 1:49:38techniques where most of the model is
  2791. 1:49:39basically fixed and there's no training
  2792. 1:49:41of the model weights there's no training
  2793. 1:49:43of Laura or anything like that of new
  2794. 1:49:45parameters the the parameters that
  2795. 1:49:47you're training are now just the uh
  2796. 1:49:49token embeddings so that's just one
  2797. 1:49:51example but this could again be like an
  2798. 1:49:52entire video but just to give you a
  2799. 1:49:54sense that there's a whole design space
  2800. 1:49:55here that is potentially worth exploring
  2801. 1:49:57in the future the next thing I want to
  2802. 1:49:59briefly address is that I think recently
  2803. 1:50:01there's a lot of momentum in how you
  2804. 1:50:03actually could construct Transformers
  2805. 1:50:05that can simultaneously process not just
  2806. 1:50:06text as the input modality but a lot of
  2807. 1:50:08other modalities so be it images videos
  2808. 1:50:11audio Etc and how do you feed in all
  2809. 1:50:14these modalities and potentially predict
  2810. 1:50:16these modalities from a Transformer uh
  2811. 1:50:18do you have to change the architecture
  2812. 1:50:19in some fundamental way and I think what
  2813. 1:50:21a lot of people are starting to converge
  2814. 1:50:23towards is that you're not changing the
  2815. 1:50:24architecture you stick with the
  2816. 1:50:25Transformer you just kind of tokenize
  2817. 1:50:27your input domains and then call the day
  2818. 1:50:29and pretend it's just text tokens and
  2819. 1:50:31just do everything else identical in an
  2820. 1:50:33identical manner so here for example
  2821. 1:50:36there was a early paper that has nice
  2822. 1:50:37graphic for how you can take an image
  2823. 1:50:39and you can chunc at it into
  2824. 1:50:42integers um and these sometimes uh so
  2825. 1:50:45these will basically become the tokens
  2826. 1:50:46of images as an example and uh these
  2827. 1:50:49tokens can be uh hard tokens where you
  2828. 1:50:52force them to be integers they can also
  2829. 1:50:53be soft tokens where you uh sort of
  2830. 1:50:57don't require uh these to be discrete
  2831. 1:51:00but you do Force these representations
  2832. 1:51:02to go through bottlenecks like in Auto
  2833. 1:51:04encoders uh also in this paper that came
  2834. 1:51:06out from open a SORA which I think
  2835. 1:51:08really um uh blew the mind of many
  2836. 1:51:11people and inspired a lot of people in
  2837. 1:51:13terms of what's possible they have a
  2838. 1:51:15Graphic here and they talk briefly about
  2839. 1:51:16how llms have text tokens Sora has
  2840. 1:51:20visual patches so again they came up
  2841. 1:51:22with a way to chunc a videos into
  2842. 1:51:24basically tokens when they own
  2843. 1:51:26vocabularies and then you can either
  2844. 1:51:28process discrete tokens say with autog
  2845. 1:51:30regressive models or even soft tokens
  2846. 1:51:32with diffusion models and uh all of that
  2847. 1:51:35is sort of uh being actively worked on
  2848. 1:51:38designed on and is beyond the scope of
  2849. 1:51:39this video but just something I wanted
  2850. 1:51:40to mention briefly okay now that we have
  2851. 1:51:42come quite deep into the tokenization
  2852. 1:51:45algorithm and we understand a lot more
  2853. 1:51:46about how it works let's loop back
  2854. 1:51:48around to the beginning of this video
  2855. 1:51:50and go through some of these bullet
  2856. 1:51:51points and really see why they happen so
  2857. 1:51:54first of all why can't my llm spell
  2858. 1:51:56words very well or do other spell
  2859. 1:51:58related
  2860. 1:52:00tasks so fundamentally this is because
  2861. 1:52:02as we saw these characters are chunked
  2862. 1:52:05up into tokens and some of these tokens
  2863. 1:52:07are actually fairly long so as an
  2864. 1:52:10example I went to the gp4 vocabulary and
  2865. 1:52:12I looked at uh one of the longer tokens
  2866. 1:52:15so that default style turns out to be a
  2867. 1:52:17single individual token so that's a lot
  2868. 1:52:19of characters for a single token so my
  2869. 1:52:22suspicion is that there's just too much
  2870. 1:52:23crammed into this single token and my
  2871. 1:52:26suspicion was that the model should not
  2872. 1:52:27be very good at tasks related to
  2873. 1:52:30spelling of this uh single token so I
  2874. 1:52:34asked how many letters L are there in
  2875. 1:52:37the word default style and of course my
  2876. 1:52:41prompt is intentionally done that way
  2877. 1:52:44and you see how default style will be a
  2878. 1:52:45single token so this is what the model
  2879. 1:52:47sees so my suspicion is that it wouldn't
  2880. 1:52:49be very good at this and indeed it is
  2881. 1:52:51not it doesn't actually know how many
  2882. 1:52:53L's are in there it thinks there are
  2883. 1:52:54three and actually there are four if I'm
  2884. 1:52:57not getting this wrong myself so that
  2885. 1:52:59didn't go extremely well let's look look
  2886. 1:53:02at another kind of uh character level
  2887. 1:53:04task so for example here I asked uh gp4
  2888. 1:53:08to reverse the string default style and
  2889. 1:53:11they tried to use a code interpreter and
  2890. 1:53:13I stopped it and I said just do it just
  2891. 1:53:15try it and uh it gave me jumble so it
  2892. 1:53:19doesn't actually really know how to
  2893. 1:53:21reverse this string going from right to
  2894. 1:53:23left uh so it gave a wrong result so
  2895. 1:53:26again like working with this working
  2896. 1:53:28hypothesis that maybe this is due to the
  2897. 1:53:30tokenization I tried a different
  2898. 1:53:31approach I said okay let's reverse the
  2899. 1:53:34exact same string but take the following
  2900. 1:53:36approach step one just print out every
  2901. 1:53:38single character separated by spaces and
  2902. 1:53:40then as a step two reverse that list and
  2903. 1:53:43it again Tred to use a tool but when I
  2904. 1:53:44stopped it it uh first uh produced all
  2905. 1:53:47the characters and that was actually
  2906. 1:53:48correct and then It reversed them and
  2907. 1:53:50that was correct once it had this so
  2908. 1:53:53somehow it can't reverse it directly but
  2909. 1:53:54when you go just first uh you know
  2910. 1:53:57listing it out in order it can do that
  2911. 1:53:59somehow and then it can once it's uh
  2912. 1:54:01broken up this way this becomes all
  2913. 1:54:03these individual characters and so now
  2914. 1:54:06this is much easier for it to see these
  2915. 1:54:07individual tokens and reverse them and
  2916. 1:54:10print them out so that is kind of
  2917. 1:54:13interesting so let's continue now why
  2918. 1:54:16are llms worse at uh non-english langu
  2919. 1:54:20and I briefly covered this already but
  2920. 1:54:22basically um it's not only that the
  2921. 1:54:24language model sees less non-english
  2922. 1:54:27data during training of the model
  2923. 1:54:28parameters but also the tokenizer is not
  2924. 1:54:31um is not sufficiently trained on
  2925. 1:54:34non-english data and so here for example
  2926. 1:54:37hello how are you is five tokens and its
  2927. 1:54:40translation is 15 tokens so this is a
  2928. 1:54:42three times blow up and so for example
  2929. 1:54:45anang is uh just hello basically in
  2930. 1:54:48Korean and that end up being three
  2931. 1:54:50tokens I'm actually kind of surprised by
  2932. 1:54:51that because that is a very common
  2933. 1:54:53phrase there just the typical greeting
  2934. 1:54:55of like hello and that ends up being
  2935. 1:54:57three tokens whereas our hello is a
  2936. 1:54:58single token and so basically everything
  2937. 1:55:00is a lot more bloated and diffuse and
  2938. 1:55:02this is I think partly the reason that
  2939. 1:55:04the model Works worse on other
  2940. 1:55:07languages uh coming back why is LM bad
  2941. 1:55:10at simple arithmetic um that has to do
  2942. 1:55:13with the tokenization of numbers and so
  2943. 1:55:17um you'll notice that for example
  2944. 1:55:19addition is very sort of
  2945. 1:55:20like uh there's an algorithm that is
  2946. 1:55:23like character level for doing addition
  2947. 1:55:25so for example here we would first add
  2948. 1:55:27the ones and then the tens and then the
  2949. 1:55:29hundreds you have to refer to specific
  2950. 1:55:31parts of these digits but uh these
  2951. 1:55:34numbers are represented completely
  2952. 1:55:36arbitrarily based on whatever happened
  2953. 1:55:37to merge or not merge during the
  2954. 1:55:39tokenization process there's an entire
  2955. 1:55:41blog post about this that I think is
  2956. 1:55:42quite good integer tokenization is
  2957. 1:55:44insane and this person basically
  2958. 1:55:46systematically explores the tokenization
  2959. 1:55:48of numbers in I believe this is gpt2 and
  2960. 1:55:52so they notice that for example for the
  2961. 1:55:53for um four-digit numbers you can take a
  2962. 1:55:57look at whether it is uh a single token
  2963. 1:56:00or whether it is two tokens that is a 1
  2964. 1:56:02three or a 2 two or a 31 combination and
  2965. 1:56:04so all the different numbers are all the
  2966. 1:56:06different combinations and you can
  2967. 1:56:08imagine this is all completely
  2968. 1:56:09arbitrarily so and the model
  2969. 1:56:11unfortunately sometimes sees uh four um
  2970. 1:56:14a token for for all four digits
  2971. 1:56:16sometimes for three sometimes for two
  2972. 1:56:18sometimes for one and it's in an
  2973. 1:56:20arbitrary uh Manner and so this is
  2974. 1:56:22definitely a headwind if you will for
  2975. 1:56:25the language model and it's kind of
  2976. 1:56:26incredible that it can kind of do it and
  2977. 1:56:27deal with it but it's also kind of not
  2978. 1:56:30ideal and so that's why for example we
  2979. 1:56:32saw that meta when they train the Llama
  2980. 1:56:342 algorithm and they use sentence piece
  2981. 1:56:36they make sure to split up all the um
  2982. 1:56:39all the digits as an example for uh
  2983. 1:56:42llama 2 and this is partly to improve a
  2984. 1:56:44simple arithmetic kind of
  2985. 1:56:46performance and finally why is gpt2 not
  2986. 1:56:50as good in Python again this is partly a
  2987. 1:56:52modeling issue on in the architecture
  2988. 1:56:54and the data set and the strength of the
  2989. 1:56:56model but it's also partially
  2990. 1:56:58tokenization because as we saw here with
  2991. 1:57:00the simple python example the encoding
  2992. 1:57:03efficiency of the tokenizer for handling
  2993. 1:57:05spaces in Python is terrible and every
  2994. 1:57:07single space is an individual token and
  2995. 1:57:09this dramatically reduces the context
  2996. 1:57:11length that the model can attend to
  2997. 1:57:12cross so that's almost like a
  2998. 1:57:14tokenization bug for gpd2 and that was
  2999. 1:57:16later fixed with gp4 okay so here's
  3000. 1:57:20another fun one my llm abruptly halts
  3001. 1:57:22when it sees the string end of text so
  3002. 1:57:25here's um here's a very strange Behavior
  3003. 1:57:28print a string end of text is what I
  3004. 1:57:30told jt4 and it says could you please
  3005. 1:57:32specify the string and I'm I'm telling
  3006. 1:57:35it give me end of text and it seems like
  3007. 1:57:37there's an issue it's not seeing end of
  3008. 1:57:39text and then I give it end of text is
  3009. 1:57:41the string and then here's a string and
  3010. 1:57:44then it just doesn't print it so
  3011. 1:57:45obviously something is breaking here
  3012. 1:57:47with respect to the handling of the
  3013. 1:57:48special token and I don't actually know
  3014. 1:57:50what open ey is doing under the hood
  3015. 1:57:52here and whether they are potentially
  3016. 1:57:54parsing this as an um as an actual token
  3017. 1:57:58instead of this just being uh end of
  3018. 1:58:01text um as like individual sort of
  3019. 1:58:04pieces of it without the special token
  3020. 1:58:06handling logic and so it might be that
  3021. 1:58:09someone when they're calling do encode
  3022. 1:58:11uh they are passing in the allowed
  3023. 1:58:13special and they are allowing end of
  3024. 1:58:16text as a special character in the user
  3025. 1:58:18prompt but the user prompt of course is
  3026. 1:58:20is a sort of um attacker controlled text
  3027. 1:58:23so you would hope that they don't really
  3028. 1:58:25parse or use special tokens or you know
  3029. 1:58:28from that kind of input but it appears
  3030. 1:58:30that there's something definitely going
  3031. 1:58:31wrong here and um so your knowledge of
  3032. 1:58:34these special tokens ends up being in a
  3033. 1:58:36tax surface potentially and so if you'd
  3034. 1:58:38like to confuse llms then just um try to
  3035. 1:58:43give them some special tokens and see if
  3036. 1:58:44you're breaking something by chance okay
  3037. 1:58:46so this next one is a really fun one uh
  3038. 1:58:49the trailing whites space issue so if
  3039. 1:58:52you come to playground and uh we come
  3040. 1:58:56here to GPT 3.5 turbo instruct so this
  3041. 1:58:58is not a chat model this is a completion
  3042. 1:59:00model so think of it more like it's a
  3043. 1:59:02lot more closer to a base model it does
  3044. 1:59:05completion it will continue the token
  3045. 1:59:07sequence so here's a tagline for ice
  3046. 1:59:09cream shop and we want to continue the
  3047. 1:59:11sequence and so we can submit and get a
  3048. 1:59:14bunch of tokens okay no problem but now
  3049. 1:59:18suppose I do this but instead of
  3050. 1:59:20pressing submit here I do here's a
  3051. 1:59:23tagline for ice cream shop space so I
  3052. 1:59:26have a space here before I click
  3053. 1:59:28submit we get a warning your text ends
  3054. 1:59:31in a trail Ling space which causes worse
  3055. 1:59:33performance due to how API splits text
  3056. 1:59:35into tokens so what's happening here it
  3057. 1:59:38still gave us a uh sort of completion
  3058. 1:59:40here but let's take a look at what's
  3059. 1:59:42happening so here's a tagline for an ice
  3060. 1:59:44cream shop and then what does this look
  3061. 1:59:48like in the actual actual training data
  3062. 1:59:50suppose you found the completion in the
  3063. 1:59:52training document somewhere on the
  3064. 1:59:53internet and the llm trained on this
  3065. 1:59:55data so maybe it's something like oh
  3066. 1:59:58yeah maybe that's the tagline that's a
  3067. 2:00:00terrible tagline but notice here that
  3068. 2:00:02when I create o you see that because
  3069. 2:00:05there's the the space character is
  3070. 2:00:07always a prefix to these tokens in GPT
  3071. 2:00:11so it's not an O token it's a space o
  3072. 2:00:13token the space is part of the O and
  3073. 2:00:16together they are token 8840 that's
  3074. 2:00:19that's space o so what's What's
  3075. 2:00:21Happening Here is that when I just have
  3076. 2:00:24it like this and I let it complete the
  3077. 2:00:27next token it can sample the space o
  3078. 2:00:30token but instead if I have this and I
  3079. 2:00:32add my space then what I'm doing here
  3080. 2:00:34when I incode this string is I have
  3081. 2:00:37basically here's a t line for an ice
  3082. 2:00:39cream uh shop and this space at the very
  3083. 2:00:42end becomes a token
  3084. 2:00:44220 and so we've added token 220 and
  3085. 2:00:47this token otherwise would be part of
  3086. 2:00:49the tagline because if there actually is
  3087. 2:00:51a tagline here so space o is the token
  3088. 2:00:55and so this is suddenly a of
  3089. 2:00:57distribution for the model because this
  3090. 2:00:59space is part of the next token but
  3091. 2:01:01we're putting it here like this and the
  3092. 2:01:04model has seen very very little data of
  3093. 2:01:07actual Space by itself and we're asking
  3094. 2:01:10it to complete the sequence like add in
  3095. 2:01:11more tokens but the problem is that
  3096. 2:01:13we've sort of begun the first token and
  3097. 2:01:16now it's been split up and now we're out
  3098. 2:01:18of this distribution and now arbitrary
  3099. 2:01:20bad things happen and it's just a very
  3100. 2:01:23rare example for it to see something
  3101. 2:01:24like that and uh that's why we get the
  3102. 2:01:26warning so the fundamental issue here is
  3103. 2:01:29of course that um the llm is on top of
  3104. 2:01:32these tokens and these tokens are text
  3105. 2:01:34chunks they're not characters in a way
  3106. 2:01:36you and I would think of them they are
  3107. 2:01:38these are the atoms of what the LM is
  3108. 2:01:40seeing and there's a bunch of weird
  3109. 2:01:41stuff that comes out of it let's go back
  3110. 2:01:43to our default cell style I bet you that
  3111. 2:01:48the model has never in its training set
  3112. 2:01:49seen default cell sta without Le in
  3113. 2:01:54there it's always seen this as a single
  3114. 2:01:56group because uh this is some kind of a
  3115. 2:01:59function in um I'm guess I don't
  3116. 2:02:02actually know what this is part of this
  3117. 2:02:03is some kind of API but I bet you that
  3118. 2:02:05it's never seen this combination of
  3119. 2:02:07tokens uh in its training data because
  3120. 2:02:10or I think it would be extremely rare so
  3121. 2:02:12I took this and I copy pasted it here
  3122. 2:02:14and I had I tried to complete from it
  3123. 2:02:17and the it immediately gave me a big
  3124. 2:02:19error and it said the model predicted to
  3125. 2:02:21completion that begins with a stop
  3126. 2:02:22sequence resulting in no output consider
  3127. 2:02:24adjusting your prompt or stop sequences
  3128. 2:02:26so what happened here when I clicked
  3129. 2:02:27submit is that immediately the model
  3130. 2:02:30emitted and sort of like end of text
  3131. 2:02:32token I think or something like that it
  3132. 2:02:34basically predicted the stop sequence
  3133. 2:02:36immediately so it had no completion and
  3134. 2:02:38so this is why I'm getting a warning
  3135. 2:02:40again because we're off the data
  3136. 2:02:42distribution and the model is just uh
  3137. 2:02:45predicting just totally arbitrary things
  3138. 2:02:47it's just really confused basically this
  3139. 2:02:49is uh this is giving it brain damage
  3140. 2:02:50it's never seen this before it's shocked
  3141. 2:02:53and it's predicting end of text or
  3142. 2:02:54something I tried it again here and it
  3143. 2:02:57in this case it completed it but then
  3144. 2:02:59for some reason this request May violate
  3145. 2:03:01our usage policies this was
  3146. 2:03:03flagged um basically something just like
  3147. 2:03:06goes wrong and there's something like
  3148. 2:03:07Jank you can just feel the Jank because
  3149. 2:03:09the model is like extremely unhappy with
  3150. 2:03:11just this and it doesn't know how to
  3151. 2:03:12complete it because it's never occurred
  3152. 2:03:14in training set in a training set it
  3153. 2:03:16always appears like this and becomes a
  3154. 2:03:18single token
  3155. 2:03:20so these kinds of issues where tokens
  3156. 2:03:21are either you sort of like complete the
  3157. 2:03:24first character of the next token or you
  3158. 2:03:26are sort of you have long tokens that
  3159. 2:03:28you then have just some of the
  3160. 2:03:29characters off all of these are kind of
  3161. 2:03:32like issues with partial tokens is how I
  3162. 2:03:35would describe it and if you actually
  3163. 2:03:37dig into the T token
  3164. 2:03:39repository go to the rust code and
  3165. 2:03:41search for
  3166. 2:03:44unstable and you'll see um en code
  3167. 2:03:47unstable native unstable token tokens
  3168. 2:03:49and a lot of like special case handling
  3169. 2:03:51none of this stuff about unstable tokens
  3170. 2:03:53is documented anywhere but there's a ton
  3171. 2:03:55of code dealing with unstable tokens and
  3172. 2:03:58unstable tokens is exactly kind of like
  3173. 2:04:00what I'm describing here what you would
  3174. 2:04:02like out of a completion API is
  3175. 2:04:05something a lot more fancy like if we're
  3176. 2:04:06putting in default cell sta if we're
  3177. 2:04:08asking for the next token sequence we're
  3178. 2:04:10not actually trying to append the next
  3179. 2:04:12token exactly after this list we're
  3180. 2:04:14actually trying to append we're trying
  3181. 2:04:16to consider lots of tokens um
  3182. 2:04:19that if we were or I guess like we're
  3183. 2:04:22trying to search over characters that if
  3184. 2:04:25we retened would be of high probability
  3185. 2:04:28if that makes sense um so that we can
  3186. 2:04:30actually add a single individual
  3187. 2:04:32character uh instead of just like adding
  3188. 2:04:34the next full token that comes after
  3189. 2:04:36this partial token list so I this is
  3190. 2:04:39very tricky to describe and I invite you
  3191. 2:04:41to maybe like look through this it ends
  3192. 2:04:43up being extremely gnarly and hairy kind
  3193. 2:04:44of topic it and it comes from
  3194. 2:04:46tokenization fundamentally so um maybe I
  3195. 2:04:49can even spend an entire video talking
  3196. 2:04:50about unstable tokens sometime in the
  3197. 2:04:52future okay and I'm really saving the
  3198. 2:04:54best for last my favorite one by far is
  3199. 2:04:56the solid gold
  3200. 2:04:59Magikarp and it just okay so this comes
  3201. 2:05:01from this blog post uh solid gold
  3202. 2:05:03Magikarp and uh this is um internet
  3203. 2:05:07famous now for those of us in llms and
  3204. 2:05:10basically I I would advise you to uh
  3205. 2:05:11read this block Post in full but
  3206. 2:05:13basically what this person was doing is
  3207. 2:05:16this person went to the um
  3208. 2:05:19token embedding stable and clustered the
  3209. 2:05:22tokens based on their embedding
  3210. 2:05:24representation and this person noticed
  3211. 2:05:27that there's a cluster of tokens that
  3212. 2:05:29look really strange so there's a cluster
  3213. 2:05:31here at rot e stream Fame solid gold
  3214. 2:05:34Magikarp Signet message like really
  3215. 2:05:36weird tokens in uh basically in this
  3216. 2:05:39embedding cluster and so what are these
  3217. 2:05:42tokens and where do they even come from
  3218. 2:05:43like what is solid gold magikarpet makes
  3219. 2:05:45no sense and then they found bunch of
  3220. 2:05:48these
  3221. 2:05:50tokens and then they notice that
  3222. 2:05:52actually the plot thickens here because
  3223. 2:05:53if you ask the model about these tokens
  3224. 2:05:56like you ask it uh some very benign
  3225. 2:05:58question like please can you repeat back
  3226. 2:06:00to me the string sold gold Magikarp uh
  3227. 2:06:02then you get a variety of basically
  3228. 2:06:04totally broken llm Behavior so either
  3229. 2:06:07you get evasion so I'm sorry I can't
  3230. 2:06:09hear you or you get a bunch of
  3231. 2:06:11hallucinations as a response um you can
  3232. 2:06:14even get back like insults so you ask it
  3233. 2:06:17uh about streamer bot it uh tells the
  3234. 2:06:20and the model actually just calls you
  3235. 2:06:22names uh or it kind of comes up with
  3236. 2:06:24like weird humor like you're actually
  3237. 2:06:26breaking the model by asking about these
  3238. 2:06:28very simple strings like at Roth and
  3239. 2:06:30sold gold Magikarp so like what the hell
  3240. 2:06:32is happening and there's a variety of
  3241. 2:06:34here documented behaviors uh there's a
  3242. 2:06:37bunch of tokens not just so good
  3243. 2:06:38Magikarp that have that kind of a
  3244. 2:06:40behavior and so basically there's a
  3245. 2:06:42bunch of like trigger words and if you
  3246. 2:06:44ask the model about these trigger words
  3247. 2:06:46or you just include them in your prompt
  3248. 2:06:48the model goes haywire and has all kinds
  3249. 2:06:50of uh really Strange Behaviors including
  3250. 2:06:52sort of ones that violate typical safety
  3251. 2:06:54guidelines uh and the alignment of the
  3252. 2:06:57model like it's swearing back at you so
  3253. 2:06:59what is happening here and how can this
  3254. 2:07:01possibly be true well this again comes
  3255. 2:07:04down to tokenization so what's happening
  3256. 2:07:06here is that sold gold Magikarp if you
  3257. 2:07:08actually dig into it is a Reddit user so
  3258. 2:07:11there's a u Sol gold
  3259. 2:07:14Magikarp and probably what happened here
  3260. 2:07:16even though I I don't know that this has
  3261. 2:07:18been like really definitively explored
  3262. 2:07:20but what is thought to have happened is
  3263. 2:07:23that the tokenization data set was very
  3264. 2:07:25different from the training data set for
  3265. 2:07:28the actual language model so in the
  3266. 2:07:29tokenization data set there was a ton of
  3267. 2:07:31redded data potentially where the user
  3268. 2:07:34solid gold Magikarp was mentioned in the
  3269. 2:07:36text because solid gold Magikarp was a
  3270. 2:07:39very common um sort of uh person who
  3271. 2:07:41would post a lot uh this would be a
  3272. 2:07:43string that occurs many times in a
  3273. 2:07:45tokenization data set because it occurs
  3274. 2:07:48many times in a tokenization data set
  3275. 2:07:50these tokens would end up getting merged
  3276. 2:07:51to the single individual token for that
  3277. 2:07:53single Reddit user sold gold Magikarp so
  3278. 2:07:56they would have a dedicated token in a
  3279. 2:07:58vocabulary of was it 50,000 tokens in
  3280. 2:08:00gpd2 that is devoted to that Reddit user
  3281. 2:08:04and then what happens is the
  3282. 2:08:05tokenization data set has those strings
  3283. 2:08:08but then later when you train the model
  3284. 2:08:10the language model itself um this data
  3285. 2:08:13from Reddit was not present and so
  3286. 2:08:16therefore in the entire training set for
  3287. 2:08:18the language model sold gold Magikarp
  3288. 2:08:21never occurs that token never appears in
  3289. 2:08:24the training set for the actual language
  3290. 2:08:25model later so this token never gets
  3291. 2:08:28activated it's initialized at random in
  3292. 2:08:31the beginning of optimization then you
  3293. 2:08:32have forward backward passes and updates
  3294. 2:08:34to the model and this token is just
  3295. 2:08:36never updated in the embedding table
  3296. 2:08:37that row Vector never gets sampled it
  3297. 2:08:40never gets used so it never gets trained
  3298. 2:08:42and it's completely untrained it's kind
  3299. 2:08:43of like unallocated memory in a typical
  3300. 2:08:46binary program written in C or something
  3301. 2:08:48like that that so it's unallocated
  3302. 2:08:50memory and then at test time if you
  3303. 2:08:51evoke this token then you're basically
  3304. 2:08:54plucking out a row of the embedding
  3305. 2:08:55table that is completely untrained and
  3306. 2:08:57that feeds into a Transformer and
  3307. 2:08:58creates undefined behavior and that's
  3308. 2:09:00what we're seeing here this completely
  3309. 2:09:02undefined never before seen in a
  3310. 2:09:03training behavior and so any of these
  3311. 2:09:06kind of like weird tokens would evoke
  3312. 2:09:08this Behavior because fundamentally the
  3313. 2:09:09model is um is uh uh out of sample out
  3314. 2:09:14of distribution okay and the very last
  3315. 2:09:16thing I wanted to just briefly mention
  3316. 2:09:18point out although I think a lot of
  3317. 2:09:19people are quite aware of this is that
  3318. 2:09:21different kinds of formats and different
  3319. 2:09:23representations and different languages
  3320. 2:09:25and so on might be more or less
  3321. 2:09:26efficient with GPD tokenizers uh or any
  3322. 2:09:29tokenizers for any other L for that
  3323. 2:09:31matter so for example Json is actually
  3324. 2:09:33really dense in tokens and yaml is a lot
  3325. 2:09:36more efficient in tokens um so for
  3326. 2:09:39example this are these are the same in
  3327. 2:09:41Json and in yaml the Json is
  3328. 2:09:44116 and the yaml is 99 so quite a bit of
  3329. 2:09:48an Improvement and so in the token
  3330. 2:09:51economy where we are paying uh per token
  3331. 2:09:53in many ways and you are paying in the
  3332. 2:09:55context length and you're paying in um
  3333. 2:09:57dollar amount for uh the cost of
  3334. 2:09:59processing all this kind of structured
  3335. 2:10:01data when you have to um so prefer to
  3336. 2:10:03use theal over Json and in general kind
  3337. 2:10:06of like the tokenization density is
  3338. 2:10:07something that you have to um sort of
  3339. 2:10:09care about and worry about at all times
  3340. 2:10:11and try to find efficient encoding
  3341. 2:10:13schemes and spend a lot of time in tick
  3342. 2:10:15tokenizer and measure the different
  3343. 2:10:16token efficiencies of different formats
  3344. 2:10:18and settings and so on okay so that
  3345. 2:10:21concludes my fairly long video on
  3346. 2:10:23tokenization I know it's a try I know
  3347. 2:10:25it's annoying I know it's irritating I
  3348. 2:10:28personally really dislike the stage what
  3349. 2:10:30I do have to say at this point is don't
  3350. 2:10:32brush it off there's a lot of foot guns
  3351. 2:10:34sharp edges here security issues uh AI
  3352. 2:10:38safety issues as we saw plugging in
  3353. 2:10:39unallocated memory into uh language
  3354. 2:10:42models so um it's worth understanding
  3355. 2:10:45this stage um that said I will say that
  3356. 2:10:48eternal glory goes to anyone who can get
  3357. 2:10:50rid of it uh I showed you one possible
  3358. 2:10:52paper that tried to uh do that and I
  3359. 2:10:54think I hope a lot more can follow over
  3360. 2:10:57time and my final recommendations for
  3361. 2:10:59the application right now are if you can
  3362. 2:11:01reuse the GPT 4 tokens and the
  3363. 2:11:03vocabulary uh in your application then
  3364. 2:11:05that's something you should consider and
  3365. 2:11:06just use Tech token because it is very
  3366. 2:11:07efficient and nice library for inference
  3367. 2:11:11for bpe I also really like the bite
  3368. 2:11:13level BP that uh Tik toen and openi uses
  3369. 2:11:17uh if you for some reason want to train
  3370. 2:11:19your own vocabulary from scratch um then
  3371. 2:11:22I would use uh the bpe with sentence
  3372. 2:11:25piece um oops as I mentioned I'm not a
  3373. 2:11:28huge fan of sentence piece I don't like
  3374. 2:11:30its uh bite fallback and I don't like
  3375. 2:11:33that it's doing BP on unic code code
  3376. 2:11:35points I think it's uh it also has like
  3377. 2:11:37a million settings and I think there's a
  3378. 2:11:39lot of foot gonss here and I think it's
  3379. 2:11:40really easy to Mis calibrate them and
  3380. 2:11:42you end up cropping your sentences or
  3381. 2:11:43something like that uh because of some
  3382. 2:11:45type of parameter that you don't fully
  3383. 2:11:47understand so so be very careful with
  3384. 2:11:49the settings try to copy paste exactly
  3385. 2:11:51maybe where what meta did or basically
  3386. 2:11:54spend a lot of time looking at all the
  3387. 2:11:56hyper parameters and go through the code
  3388. 2:11:57of sentence piece and make sure that you
  3389. 2:11:59have this correct um but even if you
  3390. 2:12:02have all the settings correct I still
  3391. 2:12:03think that the algorithm is kind of
  3392. 2:12:04inferior to what's happening here and
  3393. 2:12:07maybe the best if you really need to
  3394. 2:12:09train your vocabulary maybe the best
  3395. 2:12:11thing is to just wait for M bpe to
  3396. 2:12:13becomes as efficient as possible and uh
  3397. 2:12:16that's something that maybe I hope to
  3398. 2:12:18work on and at some point maybe we can
  3399. 2:12:20be training basically really what we
  3400. 2:12:22want is we want tick token but training
  3401. 2:12:24code and that is the ideal thing that
  3402. 2:12:27currently does not exist and MBP is um
  3403. 2:12:31is in implementation of it but currently
  3404. 2:12:33it's in Python so that's currently what
  3405. 2:12:35I have to say for uh tokenization there
  3406. 2:12:38might be an advanced video that has even
  3407. 2:12:40drier and even more detailed in the
  3408. 2:12:41future but for now I think we're going
  3409. 2:12:43to leave things off here and uh I hope
  3410. 2:12:46that was helpful bye
  3411. 2:12:54and uh they increase this contact size
  3412. 2:12:56from gpt1 of 512 uh to 1024 and GPT 4
  3413. 2:13:02two the
  3414. 2:13:05next okay next I would like us to
  3415. 2:13:07briefly walk through the code from open
  3416. 2:13:09AI on the gpt2 encoded
  3417. 2:13:15ATP I'm sorry I'm gonna sneeze
  3418. 2:13:19and then what's Happening Here
  3419. 2:13:21is this is a spous layer that I will
  3420. 2:13:24explain in a
  3421. 2:13:26bit What's Happening Here
  3422. 2:13:33is

About this transcript

This page contains the full transcript of Let's build the GPT Tokenizer by Andrej Karpathy, generated from the public captions YouTube serves with the video. The transcript has 24,679 words across 3,422 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.