YouTube2Text

The spelled-out intro to language modeling: building makemore — Transcript

by Andrej Karpathy · 19,445 words · 3,332 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Hi everyone. Hope you're well.
  2. 0:02And uh next up what I'd like to do is
  3. 0:03I'd like to build out make more.
  4. 0:06Like micrograd before it, make more is a
  5. 0:08repository that I have on my GitHub
  6. 0:10webpage. Uh you can look at it. Uh but
  7. 0:12just like with micrograd, I'm going to
  8. 0:14build it out step-by-step and I'm going
  9. 0:16to spell everything out. So, we're going
  10. 0:18to build it out slowly and together.
  11. 0:20Now, what is make more?
  12. 0:22Make more, uh as the name suggests, uh
  13. 0:24makes more of things that you give it.
  14. 0:27So, here's an example. names.txt is an
  15. 0:30example data set to make more.
  16. 0:32And when you look at names.txt, you'll
  17. 0:34find that it's a very large data set of
  18. 0:36names.
  19. 0:38So,
  20. 0:40here's lots of different types of names.
  21. 0:41In fact, I believe there are 32,000
  22. 0:43names that I've sort of found randomly
  23. 0:45on a government website.
  24. 0:47And if you train make more on this data
  25. 0:50set, it will learn to make more of
  26. 0:52things like this. Um and in particular
  27. 0:56in this case, that will mean more things
  28. 0:58that sound name-like, but are actually
  29. 1:01unique names. And maybe if you have a
  30. 1:03baby and you're trying to assign name,
  31. 1:05maybe you're looking for a cool new
  32. 1:06sounding unique name, make more might
  33. 1:08help you.
  34. 1:09So, here are some example generations
  35. 1:11from the neural network once we train it
  36. 1:14on our data set.
  37. 1:16So, here's some example unique names
  38. 1:18that it will generate. Dontal,
  39. 1:21uh Irot,
  40. 1:23uh Zendy,
  41. 1:24and so on.
  42. 1:25And so, all these sort of sound
  43. 1:27name-like, uh but they're not, of
  44. 1:28course, names.
  45. 1:30So, under the hood, make more is a
  46. 1:32character-level language model. So, what
  47. 1:35that means is that it is treating every
  48. 1:37single line here as an example.
  49. 1:39And within each example, it's treating
  50. 1:41them all as sequences of individual
  51. 1:44characters. So, r e e s e is this
  52. 1:48example, and that's the sequence of
  53. 1:49characters, and that's the level on
  54. 1:51which we are building out make more.
  55. 1:53And what it means to to a
  56. 1:55character-level language model then is
  57. 1:57that it's just sort of modeling those
  58. 1:59sequences of characters and it knows how
  59. 2:00to predict the next character in the
  60. 2:02sequence.
  61. 2:03Now, we're actually going to implement a
  62. 2:05large number of character level language
  63. 2:07models in terms of the neural networks
  64. 2:09that are involved in predicting the next
  65. 2:10character in a sequence. So, very simple
  66. 2:13bigram and bag of word models,
  67. 2:15multi-layer perceptrons, recurrent
  68. 2:17neural networks, all the way to modern
  69. 2:19transformers. In fact, the transformer
  70. 2:21that we will build will be basically the
  71. 2:23equivalent transformer to GPT-2 if you
  72. 2:26have heard of GPT. Uh so, that's kind of
  73. 2:28a big deal. It's a modern network and by
  74. 2:30the end of this series you will actually
  75. 2:32understand how that works
  76. 2:34um on the level of characters.
  77. 2:36Now, to give you a sense of the
  78. 2:38extensions here, uh after characters we
  79. 2:40will probably spend some time on the
  80. 2:42word level so that we can generate
  81. 2:43documents of words, not just little, you
  82. 2:45know, segments of characters.
  83. 2:47Uh but we can generate entire large,
  84. 2:49much larger documents.
  85. 2:51And then we're probably going to go into
  86. 2:52images and image text uh networks such
  87. 2:55as DALL-E, Stable Diffusion, and so on.
  88. 2:58But for now we have to start
  89. 2:59uh here, character level language
  90. 3:01modeling. Let's go.
  91. 3:03So, like before we are starting with a
  92. 3:04completely blank Jupyter notebook page.
  93. 3:07The first thing is I would like to
  94. 3:08basically load up the data set
  95. 3:10names.txt.
  96. 3:11So, we're going to open up names.txt for
  97. 3:13reading.
  98. 3:15And we're going to read in everything
  99. 3:17into a massive string.
  100. 3:19And then, because it's a massive string,
  101. 3:21we'd only like the individual words and
  102. 3:23put them in a list.
  103. 3:24So, let's call splitlines on that string
  104. 3:27to get all of our words as a Python list
  105. 3:30of strings.
  106. 3:32So, basically we can look at for example
  107. 3:33the first 10 words
  108. 3:35and we have that it's a list of Emma,
  109. 3:39Olivia, Ava, and so on.
  110. 3:41And if we look at
  111. 3:43the top of the page here, that is indeed
  112. 3:45what we see.
  113. 3:47Um
  114. 3:48So, that's good.
  115. 3:49This list actually makes me feel that
  116. 3:52this is probably sorted by frequency.
  117. 3:55But okay, so these are the words. Now,
  118. 3:58we'd like to actually like learn a
  119. 4:00little bit more about this data set.
  120. 4:01Let's look at the total number of words.
  121. 4:03We expect this to be roughly 32,000.
  122. 4:06And then what is the for example
  123. 4:07shortest word?
  124. 4:09So min of
  125. 4:11len of each word for W in words. So the
  126. 4:13shortest word will be length two.
  127. 4:18And max of len W for W in words. So the
  128. 4:21longest word will be
  129. 4:2315 characters.
  130. 4:24So let's now think through our very
  131. 4:25first language model.
  132. 4:27As I mentioned, a character-level
  133. 4:28language model is predicting the next
  134. 4:30character in a sequence given already
  135. 4:33some concrete sequence of characters
  136. 4:35before it.
  137. 4:36Now, what we have to realize here is
  138. 4:37that every single word here, like
  139. 4:39Isabella,
  140. 4:40is actually quite a few examples packed
  141. 4:43in to that single word.
  142. 4:45Because what is a an existence of a word
  143. 4:47like Isabella in the data set telling us
  144. 4:48really? It's saying that
  145. 4:51the character I is a very likely
  146. 4:53character to come first in the sequence
  147. 4:56of a name.
  148. 4:58The character S is likely to come
  149. 5:01after I.
  150. 5:04The character A is likely to come after
  151. 5:06IS.
  152. 5:07The character B is very likely to come
  153. 5:09after ISA. And so on all the way to A
  154. 5:12following Isabella.
  155. 5:14And then there's one more example
  156. 5:15actually packed in here.
  157. 5:17And that is that
  158. 5:19after there's Isabella,
  159. 5:21the word is very likely to end.
  160. 5:23So that's one more sort of explicit
  161. 5:25piece of information that we have here.
  162. 5:27That we have to be careful with.
  163. 5:29And so there's a lot packed into a
  164. 5:31single individual word in terms of the
  165. 5:33statistical structure of what's likely
  166. 5:35to follow in these character sequences.
  167. 5:38And then of course we don't have just an
  168. 5:39individual word. We actually have 32,000
  169. 5:41of these. And so there's a lot of
  170. 5:42structure here to model.
  171. 5:44Now, in the beginning what I'd like to
  172. 5:46start with is I would like to start with
  173. 5:48building a bigram language model.
  174. 5:51Now, in a bigram language model, we're
  175. 5:53always working with just two characters
  176. 5:55at a time.
  177. 5:56So, we're only looking at one character
  178. 5:59that we are given, and we're trying to
  179. 6:00predict the next character in the
  180. 6:02sequence.
  181. 6:03So, um what characters are likely to
  182. 6:06follow R? What characters are likely to
  183. 6:08follow A? And so on. And we're just
  184. 6:10modeling that kind of a little local
  185. 6:11structure.
  186. 6:13And we're forgetting the fact that we
  187. 6:15may have a lot more information. We're
  188. 6:17always just looking at the previous
  189. 6:18character to predict the next one. So,
  190. 6:20it's a very simple and weak language
  191. 6:21model, but I think it's a great place to
  192. 6:23start.
  193. 6:24So, now let's begin by looking at these
  194. 6:25bigrams in our data set and what they
  195. 6:27look like. And these bigrams, again, are
  196. 6:29just two characters in a row.
  197. 6:31So, for W words,
  198. 6:33each W here is an individual word
  199. 6:35string.
  200. 6:36We want to iterate uh for
  201. 6:39in We want to iterate this word
  202. 6:41with consecutive characters. So, two
  203. 6:44characters at a time, sliding it through
  204. 6:46the word. Now, a interesting, nice way,
  205. 6:49cute way to do this in Python, by the
  206. 6:50way, is doing something like this. For
  207. 6:52character one, character two in zip off
  208. 6:56W and W at one.
  209. 7:00one colon
  210. 7:01print
  211. 7:03character one, character two
  212. 7:04And let's not do all the words. Let's
  213. 7:05just do the first three words. And I'm
  214. 7:07going to show you in a second how this
  215. 7:09works.
  216. 7:10But for now, basically, as an example,
  217. 7:12let's just do the very first word alone,
  218. 7:13Emma.
  219. 7:15You see how we have a Emma, and this
  220. 7:18will just print EM, MM, MA.
  221. 7:21And the reason this works is because W
  222. 7:23is the string Emma, W at one colon is
  223. 7:26the string MMA,
  224. 7:28and zip takes two iterators, and it
  225. 7:32pairs them up, and then creates an
  226. 7:34iterator over the tuples of their
  227. 7:35consecutive entries.
  228. 7:37And if any one of these lists is shorter
  229. 7:39than the other, then it will just uh
  230. 7:41halt and return.
  231. 7:43So basically, uh that's why we return E
  232. 7:46M M M M M M A,
  233. 7:50but then because this iterator, the
  234. 7:51second one here, runs out of elements,
  235. 7:54zip just ends, and that's why we only
  236. 7:56get these tuples. So, pretty cute.
  237. 7:59So, these are the consecutive elements
  238. 8:01in the first word.
  239. 8:03Now, we have to be careful because we
  240. 8:04actually have more information here than
  241. 8:05just these three examples. As I
  242. 8:08mentioned, we know that E is the is very
  243. 8:11come first, and we know that A in this
  244. 8:13case is coming last.
  245. 8:15So, one way to do this is basically
  246. 8:17we're going to create a special array
  247. 8:20here of characters.
  248. 8:23And
  249. 8:24we're going to hallucinate a special
  250. 8:25start token here.
  251. 8:28I'm going to
  252. 8:29call it like special start.
  253. 8:32So, this is a list of one element plus
  254. 8:36W
  255. 8:37and then plus a special end character.
  256. 8:41And the reason I'm wrapping the list of
  257. 8:42W here is because W is a string in a
  258. 8:46list of W will just have the individual
  259. 8:48characters in the list.
  260. 8:50And then doing this again now, but not
  261. 8:54iterating over Ws but over the
  262. 8:56characters,
  263. 8:58will give us something like this.
  264. 9:00So, E is likely So, this is a bigram of
  265. 9:02the start character and E, and this is a
  266. 9:05bigram of the A and the special end
  267. 9:07character.
  268. 9:09And now we can look at, for example,
  269. 9:10what this looks like for
  270. 9:12Olivia or Ava.
  271. 9:14And indeed, we can actually
  272. 9:16potentially do this for the entire data
  273. 9:17set, but we won't print that. That's
  274. 9:19going to be too much.
  275. 9:20But these are the individual character
  276. 9:22bigrams, and we can print them.
  277. 9:25Now, in order to learn the statistics
  278. 9:26about which characters are likely to
  279. 9:28follow other characters, the simplest
  280. 9:30way in the bigram language models is to
  281. 9:32simply do it by counting.
  282. 9:34So, we're basically just going to count
  283. 9:36how often any one of these combinations
  284. 9:38occurs in the training set.
  285. 9:40In these words. So we're going to need
  286. 9:42some kind of a dictionary that's going
  287. 9:44to maintain some counts for every one of
  288. 9:46these bigrams. So let's use a dictionary
  289. 9:48B.
  290. 9:49And this will map these bigrams. So
  291. 9:52bigram is a tuple of character one
  292. 9:54character two.
  293. 9:56And then B at bigram
  294. 9:58will be B.get of bigram
  295. 10:01which is basically the same as B at
  296. 10:03bigram.
  297. 10:04But in the case that bigram is not in
  298. 10:07the dictionary B, we would like to by
  299. 10:09default return a zero.
  300. 10:11Plus one.
  301. 10:13So this will basically add up all the
  302. 10:15bigrams and count how often they occur.
  303. 10:18Let's get rid of printing.
  304. 10:20Or rather
  305. 10:22let's keep the printing and let's just
  306. 10:24inspect what B is in this case.
  307. 10:27And we see that many bigrams occur just
  308. 10:29a single time. This one allegedly
  309. 10:31occurred three times.
  310. 10:33So A was an ending character three
  311. 10:34times. And that's true for all of these
  312. 10:36words.
  313. 10:38All of Emma, Olivia, and Ava end with A.
  314. 10:41Uh so that's why this occurred three
  315. 10:43times.
  316. 10:45Um
  317. 10:46now let's do it for all the words.
  318. 10:51Oops, I should not have printed.
  319. 10:55I meant to erase that.
  320. 10:56Let's kill this.
  321. 10:58Let's just run.
  322. 11:00And now B will have the statistics of
  323. 11:02the entire data set.
  324. 11:04So these are the counts across all the
  325. 11:05words of the individual bigrams.
  326. 11:08And we could for example look at some of
  327. 11:10the most common ones and least common
  328. 11:11ones.
  329. 11:12Um this kind of grows in Python, but the
  330. 11:14way to do this, the simplest way I like,
  331. 11:17is we just use B.items.
  332. 11:19B.items returns
  333. 11:21the tuples of
  334. 11:24key value. In this case the keys are the
  335. 11:27character bigrams and the values are the
  336. 11:29counts.
  337. 11:31And so then what we want to do is we
  338. 11:32want to do um
  339. 11:35sorted of this.
  340. 11:38Uh but by default sort is on the first
  341. 11:41um
  342. 11:43on the first item of a tuple, but we
  343. 11:45want to sort by the values, which are
  344. 11:47the second element of a tuple, that is
  345. 11:49the key value.
  346. 11:50So we want to use the key
  347. 11:53equals lambda
  348. 11:55uh that takes the key value
  349. 11:57and returns the key value at the at one,
  350. 12:01not at zero, but at one, which is the
  351. 12:03count. So we want to sort by the count
  352. 12:07of these elements.
  353. 12:10And actually we want it to go backwards.
  354. 12:12So here what we have is the bigram QNR
  355. 12:16occurs only a single time.
  356. 12:18Uh DZ occurred only a single time.
  357. 12:20And when we sort this the other way
  358. 12:21around,
  359. 12:23we're going to see the most likely
  360. 12:25bigrams. So we see that N was very often
  361. 12:29an ending character
  362. 12:30many, many times. And apparently N
  363. 12:32almost always follows an A. And that's a
  364. 12:34very likely combination as well.
  365. 12:37Um
  366. 12:38so
  367. 12:39this is kind of the individual counts
  368. 12:42that we achieve over the entire data
  369. 12:43set.
  370. 12:45Now it's actually going to be
  371. 12:46significantly more convenient for us to
  372. 12:48keep this information in a
  373. 12:49two-dimensional array instead of a
  374. 12:51Python dictionary.
  375. 12:53So we're going to store this information
  376. 12:56in a 2D array
  377. 12:58and
  378. 13:00the rows are going to be the first
  379. 13:01character of the bigram and the columns
  380. 13:03are going to be the second character.
  381. 13:05And each entry in this two-dimensional
  382. 13:06array will tell us how often that first
  383. 13:08character follows the second character
  384. 13:10in the data set.
  385. 13:12So in particular, the array
  386. 13:14representation that we're going to use
  387. 13:16or the library is that of PyTorch.
  388. 13:18And PyTorch is a deep learning neural
  389. 13:21network framework, but part of it is
  390. 13:23also this torch.tensor, uh which allows
  391. 13:25us to create multi-dimensional arrays,
  392. 13:27and manipulate them very efficiently.
  393. 13:29So, let's import PyTorch, which you can
  394. 13:32do by import torch.
  395. 13:34And then we can create uh arrays.
  396. 13:37So, let's create a array of zeros.
  397. 13:40And we give it a um size of this array.
  398. 13:43Let's create a 3x5 array as an example.
  399. 13:47And
  400. 13:48this is a 3x5 array of zeros.
  401. 13:51And by default, you'll notice a data D
  402. 13:53type, which is short for data type, is
  403. 13:55float 32. So, these are single-precision
  404. 13:57floating-point numbers.
  405. 13:59Because we are going to represent
  406. 14:00counts, let's actually use D type as
  407. 14:03torch.int32.
  408. 14:05So, these are uh
  409. 14:0732-bit integers.
  410. 14:10So, now you see that we have integer
  411. 14:12data inside this tensor.
  412. 14:14Now, tensors allow us to really um
  413. 14:17manipulate all the individual entries
  414. 14:19and do it very efficiently.
  415. 14:20So, for example, if we want to change
  416. 14:22this bit, we have to index into the
  417. 14:24tensor. And in particular, here, this is
  418. 14:28the first row, and the um because it's
  419. 14:31zero-indexed. So, this is row index one,
  420. 14:35and column index 0 1 2 3.
  421. 14:38So, A at 1,3, we can set that to 1.
  422. 14:43And then A will have a 1 over there.
  423. 14:47We can of course also do things like
  424. 14:48this. So, now A will be 2 over there.
  425. 14:52Or 3.
  426. 14:53And also, we can for example say A 0 0
  427. 14:56is 5.
  428. 14:57And then A will have a 5 over here.
  429. 15:00So, that's how we can index into the
  430. 15:02arrays. Now, of course, the array that
  431. 15:04we are interested in is much much
  432. 15:05bigger. So, for our purposes, we have 26
  433. 15:08letters of the alphabet, and then we
  434. 15:10have two special characters, S and E.
  435. 15:13So, uh we want 26 + 2 or 28x28 array.
  436. 15:19And let's call it the the N, because
  437. 15:21it's going to represent sort of the
  438. 15:22counts.
  439. 15:24Let me erase this stuff.
  440. 15:26So, that's the array that starts at
  441. 15:28zeros, 28 by 28.
  442. 15:30And now let's copy paste this
  443. 15:33here.
  444. 15:34But instead of having a dictionary be
  445. 15:37which we're going to erase, we now have
  446. 15:39an N.
  447. 15:41Now, the problem here is that we have
  448. 15:42these characters which are strings, but
  449. 15:44we have to now
  450. 15:45um basically index into a um array, and
  451. 15:49we have to index using integers. So, we
  452. 15:51need some kind of a lookup table from
  453. 15:53characters to integers.
  454. 15:55So, let's construct such a character
  455. 15:56array.
  456. 15:58And the way we're going to do this is
  457. 15:59we're going to take all the words, which
  458. 16:01is a list of strings.
  459. 16:02We're going to concatenate all of it
  460. 16:04into a massive string. So, this is just
  461. 16:06simply the entire data set as a single
  462. 16:07string.
  463. 16:09We're going to pass this to the set
  464. 16:10constructor, which takes this massive
  465. 16:13string and throws out duplicates because
  466. 16:16sets do not allow duplicates.
  467. 16:18So, set of this will just be the set of
  468. 16:21all the lowercase characters.
  469. 16:24And there should be a total of 26 of
  470. 16:26them.
  471. 16:28And now we actually don't want a set, we
  472. 16:30want a list.
  473. 16:32But we don't want a list sorted in some
  474. 16:34weird arbitrary way, we want it to be
  475. 16:36sorted
  476. 16:37from A to Z.
  477. 16:39So, a sorted list.
  478. 16:41So, those are our characters.
  479. 16:45Now, what we want is this lookup table
  480. 16:47as I mentioned. So, let's create a
  481. 16:49special S2I, I will call it.
  482. 16:53Um S is string or character, and this
  483. 16:55will be an S2I mapping
  484. 16:58for
  485. 17:00IS in enumerate of these characters.
  486. 17:04So, enumerate basically gives us this
  487. 17:06iterator over the integer index and the
  488. 17:10actual element of the list. And then we
  489. 17:12are mapping the character to the
  490. 17:13integer.
  491. 17:15So, S2I
  492. 17:16is a mapping from A to 0, B to 1, etc.
  493. 17:19All the way from Z to 25.
  494. 17:24And that's going to be useful here, but
  495. 17:25we actually also have to specifically
  496. 17:27set that S will be 26.
  497. 17:29And S 2 I at E
  498. 17:32will be 27, right? Because Z was 25.
  499. 17:35So, those are the lookups. And now we
  500. 17:38can come here and we can map both
  501. 17:40character one and character two to their
  502. 17:41integers.
  503. 17:42So, this will be S 2 I at character one.
  504. 17:45And IX2 will be S 2 I of character two.
  505. 17:49And now we should be able to
  506. 17:52do this line, but using our array. So, N
  507. 17:55at IX1, IX2, this is the two-dimensional
  508. 17:58array indexing I showed you before,
  509. 18:00and honestly just plus equals one.
  510. 18:02Because everything starts at zero.
  511. 18:06So, this should work and give us a large
  512. 18:1028 by 28 array
  513. 18:13of all these counts. So, if we print N,
  514. 18:16this is the array, but of course it
  515. 18:18looks ugly. So, let's erase this ugly
  516. 18:21mess and let's try to visualize it a bit
  517. 18:22more nicer.
  518. 18:24So, for that we're going to use a
  519. 18:26library called Matplotlib.
  520. 18:28So, Matplotlib allows us to create
  521. 18:30figures. So, we can do things like
  522. 18:32plt.imshow of the counter array.
  523. 18:36So, this is the 28 by 28 array.
  524. 18:38And this is a structure, but even this,
  525. 18:41I would say, is still pretty ugly. So,
  526. 18:44we're going to try to create a much
  527. 18:45nicer visualization of it, and I wrote a
  528. 18:47bunch of code for that.
  529. 18:49Uh, the first thing we're going to need
  530. 18:50is
  531. 18:52we're going to need to invert this array
  532. 18:54here, this um dictionary. So, S 2 I is a
  533. 18:57mapping from S to I.
  534. 18:59And in I 2 S, we're going to reverse
  535. 19:02this dictionary. So, iterate over all
  536. 19:03the items and just reverse that array.
  537. 19:06So, I 2 S maps inversely from 0 to A, 1
  538. 19:10to B, etc.
  539. 19:12So, we'll need that.
  540. 19:14And then here's the code that I came up
  541. 19:16with to try to make this a little bit
  542. 19:17nicer.
  543. 19:20We create a figure.
  544. 19:22We plot N.
  545. 19:24And then we do and then we visualize a
  546. 19:26bunch of things later. Let me just run
  547. 19:28it so you get a sense of what all this
  548. 19:29is.
  549. 19:32Okay.
  550. 19:33So, you see here that we have
  551. 19:35the array spaced out. And every one of
  552. 19:38these is basically like B follows G zero
  553. 19:41times.
  554. 19:42B follows H 41 times. Um so, A follows J
  555. 19:46175 times.
  556. 19:48And so, what you can see that I'm doing
  557. 19:49here is first I show that entire array.
  558. 19:53And then I iterate over all the
  559. 19:54individual little cells here.
  560. 19:56And I create a character string here.
  561. 19:59Which is the inverse mapping I to S of
  562. 20:02the integer I and the integer J. So,
  563. 20:04that's These are the bigrams in a
  564. 20:06character representation.
  565. 20:08And then I plot just the bigram text.
  566. 20:12And then I plot the number of times that
  567. 20:14this bigram occurs.
  568. 20:16Now, the reason that there's a dot item
  569. 20:17here is because when you index into
  570. 20:20these arrays, these are torch tensors.
  571. 20:23You see that we still get a tensor back.
  572. 20:26So, the type of this thing, you'd think
  573. 20:28it would be just an integer 149, but
  574. 20:29it's actually torch.tensor.
  575. 20:32And so, if you do dot item, then it will
  576. 20:34pop out that individual integer.
  577. 20:38So, it'll just be 149.
  578. 20:40So, that's what's happening there. And
  579. 20:42these are just some options to make it
  580. 20:43look nice.
  581. 20:45So, what is the structure of this array?
  582. 20:47Um
  583. 20:49We have all these counts, and we see
  584. 20:50that some of them occur often and some
  585. 20:52of them do not occur often.
  586. 20:54Now, if you scrutinize this carefully,
  587. 20:56you will notice that we're not actually
  588. 20:57being very clever.
  589. 20:58That's because when you come over here,
  590. 21:00you'll notice that for example, we have
  591. 21:02an entire row of completely zeros. And
  592. 21:04that's because the end character
  593. 21:07is never possibly going to be the first
  594. 21:08character of a bigram because we're
  595. 21:10always placing these end tokens all at
  596. 21:12the end of the bigram.
  597. 21:14Similarly, we have entire columns of
  598. 21:16zeros here because the S
  599. 21:19character will never possibly be the
  600. 21:21second element of a bigram because we
  601. 21:23always start with S and we end with E
  602. 21:25and we only have the words in between.
  603. 21:27So, we have an entire column of zeros,
  604. 21:29an entire row of zeros. And in this
  605. 21:32little 2 by 2 matrix here as well, the
  606. 21:34only one that can possibly happen is if
  607. 21:36S directly follows E.
  608. 21:38That can be non-zero if we have a word
  609. 21:41that has no letters. So, in that case
  610. 21:43there's no letters in the word, it's an
  611. 21:44empty word and we just have S follows E.
  612. 21:47But, the other ones are just not
  613. 21:49possible.
  614. 21:50And so, we're basically wasting space
  615. 21:51and not only that, but the S and the E
  616. 21:53are getting very crowded here.
  617. 21:55I was using these brackets because
  618. 21:57there's convention in natural language
  619. 21:58processing to use these kinds of
  620. 22:00brackets to denote special tokens. Uh
  621. 22:03but, we're going to use something else.
  622. 22:05So, let's fix all of this and make it
  623. 22:06prettier.
  624. 22:08We're not actually going to have two
  625. 22:09special tokens, we're only going to have
  626. 22:11one special token.
  627. 22:13So, we're going to have n by n array of
  628. 22:1527 by 27 instead.
  629. 22:18Instead of having two, we will just have
  630. 22:21one and I will call it a dot.
  631. 22:24Okay?
  632. 22:27Let me swing this over here.
  633. 22:30Now, one more thing that I would like to
  634. 22:31do is I would actually like to make this
  635. 22:33special character have position zero and
  636. 22:36I would like to offset all the other
  637. 22:37letters off. I find that a little bit
  638. 22:39more pleasing.
  639. 22:41Um
  640. 22:42so,
  641. 22:44we need a plus one here so that the
  642. 22:46first character, which is A, will start
  643. 22:48at one.
  644. 22:49So, S to I will now be A starts at one
  645. 22:53and dot is zero.
  646. 22:55And uh I to S, of course, we're not
  647. 22:58changing this because I to S just
  648. 22:59creates a reverse mapping and this will
  649. 23:01work fine. So, one is A, two is B, zero
  650. 23:04is dot.
  651. 23:06So, we reverse that. Here,
  652. 23:09we have
  653. 23:10a dot and a dot.
  654. 23:13This should work fine.
  655. 23:14Make sure I start at zeros.
  656. 23:17Count. And then here, we don't go up to
  657. 23:1928, we go up to 27.
  658. 23:22And this should just work.
  659. 23:30Okay.
  660. 23:31So, we see that dot dot never happened.
  661. 23:33It's at zero because we don't have empty
  662. 23:35words.
  663. 23:36Then this row here now is just very
  664. 23:38simply the
  665. 23:40counts for all the first letters. So, G
  666. 23:45J starts a word, H starts a word, I
  667. 23:47starts a word, etc. And then these are
  668. 23:50all the ending
  669. 23:51characters.
  670. 23:53And in between, we have the structure of
  671. 23:54what characters follow each other.
  672. 23:57So, this is the counts array of our
  673. 23:59entire uh data set. So, this array
  674. 24:02actually has all of the information
  675. 24:03necessary for us to actually sample from
  676. 24:06this bigram
  677. 24:07character level language model.
  678. 24:09And
  679. 24:10roughly speaking, what we're going to do
  680. 24:12is we're just going to start following
  681. 24:13these probabilities and these counts,
  682. 24:15and we're going to start sampling from
  683. 24:17the from the model.
  684. 24:18So, in the beginning, of course,
  685. 24:20we start with the dot, the start token.
  686. 24:23Dot. So, to sample the first character
  687. 24:26of a name, we're looking at this row
  688. 24:28here.
  689. 24:30So, we see that we have the counts, and
  690. 24:32those counts are telling are telling us
  691. 24:34how often any one of these characters is
  692. 24:37to start a word.
  693. 24:39So, if we take this N,
  694. 24:41and we grab the first row,
  695. 24:44we can do that by using just indexing at
  696. 24:47zero,
  697. 24:48and then using this notation colon for
  698. 24:51the rest of that row.
  699. 24:53So, N zero colon
  700. 24:56is indexing into the zeroth um row and
  701. 24:59then it's grabbing all the columns.
  702. 25:02And so this will give us a
  703. 25:03one-dimensional array
  704. 25:05of the first row. So 0 4 4 10.
  705. 25:08You know it's 0 4 4 10 1 3 0 6 1 5 4 2
  706. 25:12etc. It's just the first row.
  707. 25:14The shape of this is 27. It's just a row
  708. 25:17of 27.
  709. 25:19And the other way that you can do this
  710. 25:21also is you just you don't need to
  711. 25:22actually give this
  712. 25:23you just uh grab the zeroth row like
  713. 25:25this. This is equivalent.
  714. 25:28Now these are the counts.
  715. 25:30And now what we'd like to do is we'd
  716. 25:31like to basically um sample from this.
  717. 25:35Since these are the raw counts, we
  718. 25:36actually have to convert this to
  719. 25:37probabilities.
  720. 25:39So we create a probability vector.
  721. 25:42So we'll take n of zero
  722. 25:45and we'll actually convert this to float
  723. 25:48first.
  724. 25:50Okay, so these integers are converted to
  725. 25:51float
  726. 25:52a floating point numbers. And the reason
  727. 25:54we're creating floats is because we're
  728. 25:56about to normalize these counts.
  729. 25:58So to create a probability distribution
  730. 26:00here, we want to divide
  731. 26:03we basically want to do p p p divide
  732. 26:05p.sum.
  733. 26:09And now we get a vector of smaller
  734. 26:11numbers and these are now probabilities.
  735. 26:13So of course because we divided by the
  736. 26:15sum, the sum of p now is one.
  737. 26:18So this is a nice proper probability
  738. 26:20distribution. It sums to one and this is
  739. 26:22giving us the probability for any single
  740. 26:24character to be the first uh character
  741. 26:26of a word.
  742. 26:28So now we can try to sample from this
  743. 26:29distribution. To sample from these
  744. 26:31distributions, we're going to use
  745. 26:32torch.multinomial which I've pulled up
  746. 26:34here.
  747. 26:36So torch.multinomial returns uh
  748. 26:39um samples from the multinomial
  749. 26:41probability distribution which is a
  750. 26:43complicated way of saying you give me
  751. 26:45probabilities and I will give you
  752. 26:47integers which are sampled according to
  753. 26:49the probability distribution.
  754. 26:51So this is the signature of the method,
  755. 26:53and to make everything deterministic,
  756. 26:54we're going to use a generator object in
  757. 26:57PyTorch.
  758. 26:59Uh so, this makes everything
  759. 27:00deterministic. So, you when you run this
  760. 27:01on your computer, you're going to the
  761. 27:03exact get the exact same results that
  762. 27:04I'm getting here on my computer.
  763. 27:07So, let me show you how this works.
  764. 27:09Um
  765. 27:12Here's the deterministic way of creating
  766. 27:15a torch generator object,
  767. 27:18seeding it with some number that we can
  768. 27:20agree on.
  769. 27:21So, that seeds a generator, gets gives
  770. 27:23us an object G,
  771. 27:25and then we can pass that G to a
  772. 27:27function
  773. 27:28that creates um
  774. 27:30here random numbers. torch.rand creates
  775. 27:32random numbers, three of them,
  776. 27:35and it's using this generator object to
  777. 27:37as a source of randomness.
  778. 27:40Uh so, uh
  779. 27:41without normalizing it,
  780. 27:44I can just print.
  781. 27:46Uh this is sort of like numbers between
  782. 27:48zero and one that are random according
  783. 27:50to this thing. And whenever I run it
  784. 27:52again,
  785. 27:53I'm always going to get the same result
  786. 27:54because I keep using the same generator
  787. 27:56object, which I'm seeding here.
  788. 27:58And then if I divide
  789. 28:01to normalize, I'm going to get a nice
  790. 28:04probability distribution of just three
  791. 28:05elements.
  792. 28:07And then we can use torch.multinomial to
  793. 28:09draw samples from it. So, this is what
  794. 28:11that looks like.
  795. 28:13torch.multinomial will take
  796. 28:16the torch tensor
  797. 28:18of probability distributions.
  798. 28:21Then we can ask for a number of samples,
  799. 28:22let's say 20.
  800. 28:24replacement equals true means that when
  801. 28:27we draw an element, uh we will uh we can
  802. 28:30draw it, and then we can put it back
  803. 28:31into the list of eligible indices to
  804. 28:34draw again.
  805. 28:35And we have to specify replacement as
  806. 28:37true because by default, uh for for some
  807. 28:39reason, it's false.
  808. 28:41Um and I think,
  809. 28:43you know, it's just something to be
  810. 28:44careful with.
  811. 28:45Uh and the generator is passed in here.
  812. 28:47So, we are going to always get
  813. 28:48deterministic results, the same results.
  814. 28:51So, if I run these two,
  815. 28:53we're going to get a bunch of samples
  816. 28:55from this distribution.
  817. 28:57Now, you'll notice here that the
  818. 28:58probability for the first element in
  819. 29:01this tensor is 60%.
  820. 29:04So, in these 20 samples, we'd expect 60%
  821. 29:08of them to be zero.
  822. 29:10We'd expect 30% of them to be one.
  823. 29:14And because the uh element index two
  824. 29:17has only 10% probability, very few of
  825. 29:20these samples should be two. And indeed,
  826. 29:22we only have a small number of twos.
  827. 29:25And we can sample as many as we would
  828. 29:26like.
  829. 29:29And the more we sample, the more uh
  830. 29:31these numbers should um roughly have the
  831. 29:33distribution here.
  832. 29:35So, we should have lots of zeros, half
  833. 29:38as many um
  834. 29:41ones, and we should have um
  835. 29:44three times as few
  836. 29:46uh sorry, as few ones, and three times
  837. 29:48as few uh
  838. 29:50twos.
  839. 29:51So, you see that we have very few twos,
  840. 29:53we have some ones, and most of them are
  841. 29:54zero. So, that's what torch.multinomial
  842. 29:57is doing.
  843. 29:58For us here,
  844. 30:01we are interested in this row. We've
  845. 30:02created this um
  846. 30:05P here,
  847. 30:06and now we can sample from it.
  848. 30:09So, if we use the same seed,
  849. 30:12and then we sample from this
  850. 30:14distribution, let's just get one sample.
  851. 30:18Then, we see that the sample is, say,
  852. 30:2013.
  853. 30:21Um so, this will be the index.
  854. 30:25And let's You see how it's a tensor that
  855. 30:27wraps 13. We again have to use that item
  856. 30:30to pop out that integer.
  857. 30:32And now, index would be just the number
  858. 30:3513.
  859. 30:37And of course, the um we can do we can
  860. 30:40map the I2S of IX to figure out exactly
  861. 30:43which character we're sampling here.
  862. 30:46We're sampling M.
  863. 30:48So, we're saying that the first
  864. 30:49character is M in our generation.
  865. 30:53And just looking at the row here,
  866. 30:55M was drawn and you we can see that M
  867. 30:57actually starts a large number of words.
  868. 31:00M started 2,500 words out of 32,000
  869. 31:04words. So, almost
  870. 31:06a bit less than 10% of the words start
  871. 31:08with M. So, this is actually fairly
  872. 31:10likely character to draw.
  873. 31:13Um
  874. 31:15So, that would be the first character of
  875. 31:16our word and now we can continue to
  876. 31:18sample more characters because now we
  877. 31:20know that M started
  878. 31:22M is already sampled.
  879. 31:24So, now to draw the next character, we
  880. 31:26will come back here and we will look for
  881. 31:29the row
  882. 31:30that starts with M.
  883. 31:32So, you see M
  884. 31:34and we have a row here.
  885. 31:36So, we see that M.
  886. 31:38is
  887. 31:39516, MA is this many, MB is this many,
  888. 31:43etc. So, these are the counts for the
  889. 31:44next row and that's the next character
  890. 31:46that we are going to now generate. So, I
  891. 31:48think we are ready to actually just
  892. 31:50write out the loop because I think
  893. 31:51you're starting to get a sense of how
  894. 31:52this is going to go.
  895. 31:54The um
  896. 31:56We always begin at index zero because
  897. 31:59that's the start token.
  898. 32:02And then while true,
  899. 32:04we're going to grab the row
  900. 32:06corresponding to index
  901. 32:08that we're currently on. So, that's P.
  902. 32:11So, that's N array at IX.
  903. 32:14Convert it to float is RP.
  904. 32:18Then, we normalize this P to sum to one.
  905. 32:25Accidentally ran the infinite loop.
  906. 32:28We normalize P to sum to one.
  907. 32:30Then we need this generator object
  908. 32:33that we're going to initialize up here
  909. 32:35and we're going to draw a single sample
  910. 32:37from this distribution.
  911. 32:40And then this is going to tell us what
  912. 32:42index is going to be next.
  913. 32:46If the index sampled is zero, then
  914. 32:49that's now the end token.
  915. 32:52So, we will break.
  916. 32:55Otherwise, we are going to print
  917. 32:57s2i of ix.
  918. 33:02i2s of ix.
  919. 33:05And uh that's pretty much it. We're just
  920. 33:08uh this should work.
  921. 33:10Okay, more.
  922. 33:12So, that's the that's the name that
  923. 33:13we've sampled. We started with M. The
  924. 33:16next step was O, then R, and then dot.
  925. 33:21And this dot we printed here as well.
  926. 33:24So,
  927. 33:26let's now do this a few times.
  928. 33:28Um
  929. 33:30So, let's actually create an
  930. 33:33out list here.
  931. 33:37And instead of printing, we're going to
  932. 33:38append. So, out.append this character.
  933. 33:44And then here, let's just print it at
  934. 33:46the end. So, let's just join up all the
  935. 33:48outs, and we're just going to print
  936. 33:50more.
  937. 33:51Okay? Now, we're always getting the same
  938. 33:53result because of the generator.
  939. 33:55So, if you want to do this a few times,
  940. 33:56we can go for I in range
  941. 34:0010. We can sample 10 names.
  942. 34:02And we can just do that 10 times.
  943. 34:05And these are the names that we're
  944. 34:06getting out.
  945. 34:08Let's do 20.
  946. 34:14I'll be honest with you, this doesn't
  947. 34:15look right.
  948. 34:16So, I stared at it a few minutes to
  949. 34:17convince myself that it actually is
  950. 34:19right.
  951. 34:20The reason these samples are so terrible
  952. 34:22is that bigram language model
  953. 34:24is actually like just like really
  954. 34:26terrible.
  955. 34:27We can generate a few more here.
  956. 34:30And you can see that they're kind of
  957. 34:31like their name-like a little bit, like
  958. 34:33Keanu, O'Reilly, etc. Uh but they're
  959. 34:35just like totally messed up. Um
  960. 34:38And I mean, the reason that this is so
  961. 34:40bad, like we're generating H as a name.
  962. 34:43But you have to think through it from
  963. 34:45the model's eyes. It doesn't know that
  964. 34:47this H is the very first H. All it knows
  965. 34:50is that H was previously, and now how
  966. 34:52likely is H the last character? Well,
  967. 34:55it's somewhat likely, and so it just
  968. 34:58makes it last character. It doesn't know
  969. 34:59that there were other things before it
  970. 35:01or there were not other things before
  971. 35:03it. And so, that's why it's generating
  972. 35:05all these like sun- nonsense names.
  973. 35:08Another way to do this is
  974. 35:12to convince yourself that this is
  975. 35:13actually doing something reasonable even
  976. 35:14though it's so terrible is
  977. 35:17these little piece here are 27, right?
  978. 35:20Like 27.
  979. 35:23So, how about if we did something like
  980. 35:24this?
  981. 35:26Instead of P having any structure
  982. 35:27whatsoever,
  983. 35:28how about if P was just a torch.ones of
  984. 35:3127?
  985. 35:34of 27.
  986. 35:37By default, this is a float 32, so this
  987. 35:39is fine. Divide 27.
  988. 35:42So, what I'm doing here is this is the
  989. 35:45uniform distribution, which will make
  990. 35:47everything equally likely.
  991. 35:49And we can sample from that. So, let's
  992. 35:52see if that does any better.
  993. 35:54Okay? So, it's This is what you have
  994. 35:56from a model that is completely
  995. 35:58untrained, where everything is equally
  996. 35:59likely. So, it's obviously garbage. And
  997. 36:02then if we have a trained model, which
  998. 36:04is trained on just bigrams,
  999. 36:07this is what we get. So, you can see
  1000. 36:08that it is more name-like. It is
  1001. 36:10actually working. It's just um
  1002. 36:14bigram is so terrible, and we have to do
  1003. 36:15better. Now, next I would like to fix an
  1004. 36:17inefficiency that we have going on here.
  1005. 36:20Because what we're doing here is we're
  1006. 36:21always fetching a row of N from the
  1007. 36:24counts matrix up ahead.
  1008. 36:26And we we're always doing the same
  1009. 36:27things. We're converting to float and
  1010. 36:29we're dividing and we're doing this
  1011. 36:30every single iteration of this loop and
  1012. 36:32we just keep re-normalizing these rows
  1013. 36:34over and over again and it's extremely
  1014. 36:35inefficient and wasteful.
  1015. 36:37So, what I'd like to do is I'd like to
  1016. 36:38actually prepare a matrix capital P that
  1017. 36:41will just have the probabilities in it.
  1018. 36:43So, in other words, it's going to be the
  1019. 36:45same as the capital N matrix here of
  1020. 36:47counts, but every single row will have
  1021. 36:49the row of probabilities uh that is
  1022. 36:51normalized to one indicating the
  1023. 36:53probability distribution for the next
  1024. 36:55character given the character before it.
  1025. 36:58Um as defined by which row we're in.
  1026. 37:01So, basically what we'd like to do is
  1027. 37:03we'd like to just do it up front here
  1028. 37:05and then we would like to just use that
  1029. 37:06row here.
  1030. 37:07Uh so, here we would like to just do P
  1031. 37:10equals P of IX instead.
  1032. 37:12Okay.
  1033. 37:14The other reason I want to do this is
  1034. 37:16not just for efficiency, but also I
  1035. 37:17would like us to practice uh these
  1036. 37:19N-dimensional tensors and I'd like us to
  1037. 37:21practice uh their manipulation and
  1038. 37:23especially something that's called
  1039. 37:24broadcasting that we'll go into in a
  1040. 37:25second.
  1041. 37:27We're actually going to have to become
  1042. 37:28very good at these tensor manipulations
  1043. 37:30because if we're going to build out all
  1044. 37:32the way to transformers, we're going to
  1045. 37:33be doing some pretty complicated um
  1046. 37:35array operations for efficiency and uh
  1047. 37:37we need to really understand that and be
  1048. 37:39very good at it.
  1049. 37:42So, intuitively what we want to do is we
  1050. 37:43first want to grab the floating point
  1051. 37:45copy of N.
  1052. 37:48And I'm mimicking the line here
  1053. 37:49basically.
  1054. 37:50And then we want to divide all the rows
  1055. 37:53so that they sum to one.
  1056. 37:55So, we'd like to do something like this.
  1057. 37:57P divide P.sum.
  1058. 38:00But, now we have to be careful
  1059. 38:02because P.sum actually
  1060. 38:05produces a sum
  1061. 38:08Sorry. P equals N.float.copy.
  1062. 38:10P.sum produces a um
  1063. 38:14sums up all of the counts of this entire
  1064. 38:17matrix N
  1065. 38:18and gives us a single number of just the
  1066. 38:19summation of everything. So, that's not
  1067. 38:21the way we want to defi- divide. We want
  1068. 38:23to simultaneously and in parallel divide
  1069. 38:26all the rows by their respective sums.
  1070. 38:30So, what we have to do now is we have to
  1071. 38:32go into documentation for torch.sum.
  1072. 38:35And we can scroll down here to a
  1073. 38:37definition that is relevant to us, which
  1074. 38:38is where we don't only provide an input
  1075. 38:41array that we want to sum, but we also
  1076. 38:43provide the dimension along which we
  1077. 38:45want to sum.
  1078. 38:47And in particular, we want to sum up uh
  1079. 38:49over rows, right?
  1080. 38:52Now, one more argument that I want you
  1081. 38:53to pay attention to here is the keepdim
  1082. 38:56is false.
  1083. 38:57If keepdim is true, then the output
  1084. 39:00tensor is of the same size as input,
  1085. 39:02except of course the dimension along
  1086. 39:03which you summed, which will become just
  1087. 39:05one.
  1088. 39:07But, if you pass in uh keepdim as false,
  1089. 39:12then this dimension is squeezed out. And
  1090. 39:14so, torch.sum not only does the sum and
  1091. 39:16collapses dimension to be of size one,
  1092. 39:18but in addition, it does what's called a
  1093. 39:20squeeze, where it squeeze out it
  1094. 39:22squeezes out that dimension.
  1095. 39:24So,
  1096. 39:25basically what we want here is we
  1097. 39:27instead want to do p.sum of sum axis.
  1098. 39:30And in particular, notice that p.shape
  1099. 39:33is 27 by 27.
  1100. 39:35So, when we sum up across axis zero,
  1101. 39:37then we would be taking the zeroth
  1102. 39:39dimension, and we would be summing
  1103. 39:40across it.
  1104. 39:42So, when keepdim is true,
  1105. 39:45then this thing will not only give us
  1106. 39:46the counts across um
  1107. 39:50along the columns,
  1108. 39:52but notice that basically the shape of
  1109. 39:53this is 1 by 27. We just get a row
  1110. 39:56vector.
  1111. 39:57And the reason we get a row vector here
  1112. 39:59again is because we pass in zero
  1113. 40:00dimension, so this zero dimension
  1114. 40:02becomes one, and we've done a sum,
  1115. 40:04and we get a row. And so, basically,
  1116. 40:06we've done the sum
  1117. 40:08this way,
  1118. 40:09vertically, and arrived at just a single
  1119. 40:111 by 27
  1120. 40:12vector of counts.
  1121. 40:15What happens when you take out keepdim
  1122. 40:17is that we just get 27. So, it squeezes
  1123. 40:20out that dimension and we just get um
  1124. 40:23a one-dimensional vector of size 27.
  1125. 40:28Now, we don't actually want
  1126. 40:311 by 27 row vector because that gives us
  1127. 40:33the counts or the sums across um
  1128. 40:37the uh columns.
  1129. 40:39We actually want to sum the other way
  1130. 40:41along dimension one.
  1131. 40:42And you'll see that the shape of this is
  1132. 40:4427 by one. So, it's a column vector.
  1133. 40:47It's a 27 by one
  1134. 40:49vector of counts.
  1135. 40:52Okay?
  1136. 40:53And that's because what's happened here
  1137. 40:55is that we're going horizontally and
  1138. 40:57this 27 by 27 matrix becomes a 27 by one
  1139. 41:01array.
  1140. 41:03Now, you'll notice by the way that um
  1141. 41:06the actual numbers
  1142. 41:08of these counts are identical.
  1143. 41:10And that's because this special array of
  1144. 41:12counts here comes from bigram
  1145. 41:14statistics. And actually, it just so
  1146. 41:15happens by chance or because of the way
  1147. 41:18this array is constructed that the sums
  1148. 41:20along the columns or along the rows,
  1149. 41:23horizontally or vertically, is
  1150. 41:24identical.
  1151. 41:26But actually, what we want to do in this
  1152. 41:27case is we want to sum across the uh
  1153. 41:30rows
  1154. 41:31horizontally. So, what we want here is
  1155. 41:33p.sum of one with keep dim true.
  1156. 41:3727 by one column vector.
  1157. 41:39And now, what we want to do is we want
  1158. 41:40to divide by that.
  1159. 41:44Now, we have to be careful here again.
  1160. 41:46Is it possible to take
  1161. 41:48what's a um p.shape you see here is 27
  1162. 41:51by 27. Is it possible to take a 27 by 27
  1163. 41:55array and divide it by what is a 27 by
  1164. 41:58one array?
  1165. 42:01Is that an operation that you can do?
  1166. 42:03And whether or not you can perform this
  1167. 42:05operation is determined by what's called
  1168. 42:06broadcasting rules. So, if you just
  1169. 42:08search broadcasting semantics in torch,
  1170. 42:12you'll notice that there's a special
  1171. 42:13definition for uh what's called
  1172. 42:14broadcasting that for whether or not
  1173. 42:18these two arrays can be combined in a
  1174. 42:21binary operation like division.
  1175. 42:24So the first condition is each tensor
  1176. 42:25has at least one dimension, which is the
  1177. 42:27case for us.
  1178. 42:28And then when iterating over the
  1179. 42:29dimension sizes starting at the trailing
  1180. 42:31dimension,
  1181. 42:32the dimension sizes must either be
  1182. 42:33equal, one of them is one, or one of
  1183. 42:35them does not exist.
  1184. 42:38Okay? So, let's do that. We need to
  1185. 42:40align the two arrays and their shapes,
  1186. 42:44which is very easy because both of these
  1187. 42:45shapes have two elements, so they're
  1188. 42:46aligned.
  1189. 42:48Then we iterate over from the from the
  1190. 42:50right and going to the left.
  1191. 42:52Each dimension must be either equal, one
  1192. 42:55of them is a one, or one of them does
  1193. 42:56not exist.
  1194. 42:58So in this case they're not equal, but
  1195. 42:59one of them is a one. So, this is fine.
  1196. 43:02And then this dimension, they're both
  1197. 43:03equal. So, this is fine.
  1198. 43:05So, all the dimensions are fine and
  1199. 43:08therefore the this operation is
  1200. 43:10broadcastable.
  1201. 43:12So that means that this operation is
  1202. 43:13allowed.
  1203. 43:14And what is it that these arrays do when
  1204. 43:16you divide 27 by 27 by 27 by one?
  1205. 43:19What it does is that it takes this
  1206. 43:21dimension one and it stretches it out.
  1207. 43:24It copies it
  1208. 43:25to match 27 here in this case.
  1209. 43:29So in our case, it takes this column
  1210. 43:30vector which is 27 by one
  1211. 43:32and it copies it 27 times
  1212. 43:36to make
  1213. 43:37these both be 27 by 27 internally. You
  1214. 43:40can think that way. And so it copies
  1215. 43:42those counts
  1216. 43:44and then it does an element-wise
  1217. 43:45division,
  1218. 43:47which is what we want because these
  1219. 43:48counts we want to divide by them on
  1220. 43:50every single one of these columns in
  1221. 43:53this matrix.
  1222. 43:54So this actually we expect will
  1223. 43:56normalize every single row.
  1224. 43:59And we can check that this is true by
  1225. 44:01taking the first row for example and
  1226. 44:04taking its sum. We expect this to be
  1227. 44:06one.
  1228. 44:08Because it's now normalized.
  1229. 44:10And then we expect this now, because if
  1230. 44:13we actually correctly normalize all the
  1231. 44:14rows, we expect to get the exact same
  1232. 44:16result here. So, let's run this.
  1233. 44:19It's the exact same result.
  1234. 44:21So, this is correct. So, now I would
  1235. 44:23like to scare you a little bit. Uh you
  1236. 44:25actually have to like I basically
  1237. 44:27encourage you very strongly to read
  1238. 44:28through broadcasting semantics.
  1239. 44:30And I encourage you to treat this with
  1240. 44:31respect. And it's not something to play
  1241. 44:34fast and loose with. It's something to
  1242. 44:35really respect, really understand, and
  1243. 44:37look up maybe some tutorials for
  1244. 44:39broadcasting and practice it, and be
  1245. 44:40careful with it because you can very
  1246. 44:42quickly run into bugs. Let me show you
  1247. 44:44what I mean.
  1248. 44:47You see how here we have p.sum of one,
  1249. 44:49keep dims is true.
  1250. 44:50The shape of this is 27 by one.
  1251. 44:53Let me take out this line just so we
  1252. 44:54have the n, and then we can see the
  1253. 44:57counts.
  1254. 44:58We can see that this is all the counts
  1255. 45:00across all the
  1256. 45:02rows.
  1257. 45:03And it's a 27 by one column vector,
  1258. 45:05right?
  1259. 45:07Now, suppose that I tried to do the
  1260. 45:09following,
  1261. 45:10but I erase keep dims is true here.
  1262. 45:14What does that do? If keep dims is not
  1263. 45:15true, it's false, then remember,
  1264. 45:17according to documentation, it gets rid
  1265. 45:19of this dimension one. It squeezes it
  1266. 45:22out. So, basically we just get all the
  1267. 45:24same counts, the same result, except the
  1268. 45:27shape of it is not 27 by one, it is just
  1269. 45:2927. The one disappears.
  1270. 45:31But all the counts are the same.
  1271. 45:34So, you think that this divide that
  1272. 45:37would would work.
  1273. 45:40First of all, can we even write this?
  1274. 45:42And will it Is it even Is it even
  1275. 45:44expected to run? Is it broadcastable?
  1276. 45:46Let's determine if this result is
  1277. 45:47broadcastable.
  1278. 45:49p.sum at one is shape
  1279. 45:51is 27.
  1280. 45:53This is 27 by 27. So, 27 by 27
  1281. 45:57broadcasting into 27.
  1282. 46:00So, now
  1283. 46:01rules of broadcasting, number one align
  1284. 46:03all the dimensions on the right done now
  1285. 46:06iteration over all the dimensions
  1286. 46:07starting from the right going to the
  1287. 46:09left
  1288. 46:10all the dimensions must either be equal
  1289. 46:13one of them must be one or one of them
  1290. 46:14does not exist so here they are all
  1291. 46:16equal here the dimension does not exist
  1292. 46:20so internally what broadcasting will do
  1293. 46:21is it will create a one here
  1294. 46:24and then
  1295. 46:25we see that one of them is a one and
  1296. 46:27this will get copied and this will run
  1297. 46:30this will broadcast
  1298. 46:32okay so you'd expect this to work
  1299. 46:37because we we are
  1300. 46:41this broadcasting this we can divide
  1301. 46:42this
  1302. 46:43now if I run this you'd expect it to
  1303. 46:45work but
  1304. 46:46it doesn't
  1305. 46:48you actually get garbage you get a wrong
  1306. 46:49result because this is actually a bug
  1307. 46:52this keep them equals true
  1308. 46:57makes it work
  1309. 47:00this is a bug in
  1310. 47:03both cases we are doing the correct
  1311. 47:05counts we are summing up across the rows
  1312. 47:09but keep them is saving us and making it
  1313. 47:10work so in this case
  1314. 47:12I'd like you to encourage you to
  1315. 47:13potentially like pause this video at
  1316. 47:15this point and try to think about why
  1317. 47:17this is buggy and why the keep them was
  1318. 47:19necessary here
  1319. 47:22okay
  1320. 47:23so the reason to do
  1321. 47:24for this is I'm trying to hint it here
  1322. 47:26when I was sort of giving you a bit of a
  1323. 47:27hint on how this works
  1324. 47:29this
  1325. 47:3027 vector
  1326. 47:32internally inside the broadcasting this
  1327. 47:34becomes a one by 27
  1328. 47:36and one by 27 is a row vector right and
  1329. 47:39now we are dividing 27 by 27 by one by
  1330. 47:4227
  1331. 47:43and torch will replicate this dimension
  1332. 47:45so basically
  1333. 47:47it will take
  1334. 47:49it will take this
  1335. 47:51row vector and it will copy it
  1336. 47:53vertically now
  1337. 47:5527 times so the 27 by 27 lines exactly,
  1338. 47:58and element twice divides.
  1339. 48:00And so, basically, what's happening here
  1340. 48:02is
  1341. 48:03um
  1342. 48:04we're actually normalizing the columns
  1343. 48:06instead of normalizing the rows.
  1344. 48:09So, you can check that what's happening
  1345. 48:11here is that P at zero, which is the
  1346. 48:13first row of P, dot sum is not one, it's
  1347. 48:17seven.
  1348. 48:18It is the first column, as an example,
  1349. 48:20that sums to one.
  1350. 48:23So,
  1351. 48:24to summarize, where does the issue come
  1352. 48:26from? The issue comes from the silent
  1353. 48:28adding of a dimension here because in
  1354. 48:30broadcasting rules, you align on the
  1355. 48:31right and go from right to left, and if
  1356. 48:34dimension doesn't exist, you create it.
  1357. 48:36So, that's where the problem happens. We
  1358. 48:38still did the counts correctly. We did
  1359. 48:39the counts across the rows, and we got
  1360. 48:41the the counts on the right here as a
  1361. 48:44column vector. But, because the keepdims
  1362. 48:46was true, this this uh this dimension
  1363. 48:48was discarded, and now we just have a
  1364. 48:50vector 27.
  1365. 48:51And because of broadcasting the way it
  1366. 48:53works, this vector of 27 suddenly
  1367. 48:55becomes a row vector.
  1368. 48:57And then this row vector gets replicated
  1369. 48:58vertically, and then every single point
  1370. 49:01we are dividing by the by the count
  1371. 49:05uh in the opposite direction.
  1372. 49:07So, uh
  1373. 49:08so, this thing just uh doesn't work. You
  1374. 49:11This needs to be keepdims equals true in
  1375. 49:13this case.
  1376. 49:14So, then
  1377. 49:15um then we have that P at zero is
  1378. 49:17normalized.
  1379. 49:19And conversely, the first column, you'd
  1380. 49:21expect to potentially not be normalized.
  1381. 49:24And this is what makes that work.
  1382. 49:27So, pretty subtle, and uh hopefully this
  1383. 49:31helps to scare you that you should have
  1384. 49:33a respect for broadcasting, be careful,
  1385. 49:35check your work,
  1386. 49:36uh and uh understand how it works under
  1387. 49:38the hood, and make sure that it's
  1388. 49:39broadcasting in the direction that you
  1389. 49:40like. Otherwise, you're going to
  1390. 49:41introduce very subtle bugs, very hard to
  1391. 49:43hard to find bugs, and uh just be
  1392. 49:46careful. One more note on efficiency, we
  1393. 49:48don't want to be doing this here because
  1394. 49:50uh this creates a completely new tensor
  1395. 49:52that we store into P. We prefer to use
  1396. 49:55in-place operations if possible.
  1397. 49:57Also, this would be an in-place
  1398. 49:59operation. It has the potential to be
  1399. 50:01faster. It doesn't create new memory
  1400. 50:03under the hood. And then let's erase
  1401. 50:05this. We don't need it.
  1402. 50:07And let's also
  1403. 50:10um just do fewer just so I'm not wasting
  1404. 50:13space.
  1405. 50:14Okay, so we're actually in a pretty good
  1406. 50:15spot now.
  1407. 50:17We trained a bigram language model, and
  1408. 50:19we trained it really just by counting
  1409. 50:21uh how frequently any pairing occurs and
  1410. 50:24then normalizing so that we get a nice
  1411. 50:26probability distribution.
  1412. 50:28So really these elements of this array P
  1413. 50:31are really the parameters of our bigram
  1414. 50:32language model giving us and summarizing
  1415. 50:34the statistics of these bigrams.
  1416. 50:37So we trained the model and then we know
  1417. 50:38how to sample from the model. We just
  1418. 50:40iteratively uh sample the next character
  1419. 50:43and uh feed it in each time and get the
  1420. 50:45next character.
  1421. 50:47Now what I'd like to do is I'd like to
  1422. 50:48somehow evaluate the quality of this
  1423. 50:50model.
  1424. 50:51We'd like to somehow summarize the
  1425. 50:53quality of this model into a single
  1426. 50:54number. How good is it at predicting uh
  1427. 50:57the training set?
  1428. 50:59And as an example, so in the training
  1429. 51:00set we can evaluate now the training
  1430. 51:03loss. And this training loss is telling
  1431. 51:05us about uh sort of the quality of this
  1432. 51:07model in a single number just like we
  1433. 51:09saw in micrograd.
  1434. 51:11So let's try to think through the
  1435. 51:13quality of the model and how we would
  1436. 51:14evaluate it.
  1437. 51:17Basically what we're going to do is
  1438. 51:18we're going to copy paste this code
  1439. 51:20that we previously used for counting.
  1440. 51:23Okay?
  1441. 51:24And let me just print the bigrams first.
  1442. 51:26We're going to use F-strings.
  1443. 51:27And I'm going to print character one
  1444. 51:29followed by character two. These are the
  1445. 51:31bigrams. And then I don't want to do it
  1446. 51:32for all the words. Let's just do the
  1447. 51:34first three words.
  1448. 51:36So here we have Emma, Olivia, and Ava
  1449. 51:38bigrams.
  1450. 51:40Now what we'd like to do is we'd like to
  1451. 51:41basically look at the probability that
  1452. 51:44the model assigns to every one of these
  1453. 51:46bigrams.
  1454. 51:48So, in other words, we can look at the
  1455. 51:49probability, which is summarized in the
  1456. 51:51matrix P,
  1457. 51:53of IX1 IX2.
  1458. 51:56And then we can print it here as
  1459. 51:58probability.
  1460. 52:00And because these probabilities are way
  1461. 52:01too large, let me percent uh
  1462. 52:04or colon point four F
  1463. 52:06to like truncate it a bit.
  1464. 52:09So, what do we have here, right? We're
  1465. 52:10looking at the probabilities that the
  1466. 52:11model assigns to every one of these
  1467. 52:13bigrams in the data set.
  1468. 52:15And so, we can see some of them are 4%,
  1469. 52:173%, etc. Just to have a measuring stick
  1470. 52:19in our mind, by the way, um
  1471. 52:21with 27 possible characters or tokens,
  1472. 52:24and if everything was equally likely,
  1473. 52:26then you'd expect all these
  1474. 52:27probabilities to be
  1475. 52:304% roughly.
  1476. 52:32So, anything above 4% means that we've
  1477. 52:34learned something useful from these
  1478. 52:36bigram statistics.
  1479. 52:37And you see that roughly some of these
  1480. 52:38are 4%, but some of them are as high as
  1481. 52:4040%,
  1482. 52:4135%, and so on. So, you see that the
  1483. 52:44model actually assigned a pretty high
  1484. 52:45probability to whatever's in the
  1485. 52:47training set. And so, that that's a good
  1486. 52:49thing.
  1487. 52:50Um basically, if you have a very good
  1488. 52:51model, you'd expect that these
  1489. 52:53probabilities should be near one,
  1490. 52:54because that means that uh your model is
  1491. 52:56correctly predicting what's going to
  1492. 52:57come next, especially in the training
  1493. 52:59set, where you where you trained your
  1494. 53:01model.
  1495. 53:02So, now we'd like to think about how can
  1496. 53:05we summarize these probabilities into a
  1497. 53:07single number that measures the quality
  1498. 53:09of this model.
  1499. 53:11Now, when you look at the literature
  1500. 53:13into maximum likelihood estimation and
  1501. 53:15statistical modeling and so on,
  1502. 53:17you'll see that what's typically used
  1503. 53:18here is something called the likelihood.
  1504. 53:21And the likelihood is the product of all
  1505. 53:23of these probabilities.
  1506. 53:25And so, the product of all of these
  1507. 53:27probabilities is the likelihood, and
  1508. 53:29it's really telling us about the
  1509. 53:31probability of the entire data set
  1510. 53:33assigned uh assigned by the model that
  1511. 53:36we've trained. And that is a measure of
  1512. 53:38quality.
  1513. 53:39So, the product of these should be as
  1514. 53:41high as possible
  1515. 53:43when you are training the model and when
  1516. 53:44you have a good model, your your product
  1517. 53:46of these probabilities should be very
  1518. 53:48high.
  1519. 53:49Um
  1520. 53:50Now, because the product of these
  1521. 53:51probabilities is an unwieldy thing to
  1522. 53:53work with, uh you can see that all of
  1523. 53:55them are between 0 and 1. So, your
  1524. 53:56product of these probabilities will be a
  1525. 53:58very tiny number.
  1526. 54:00Um so, for convenience, what people work
  1527. 54:03with usually is not the likelihood, but
  1528. 54:04they work with what's called the log
  1529. 54:06likelihood.
  1530. 54:07So,
  1531. 54:09the product of these is likelihood.
  1532. 54:10To get likelihood, we just have to take
  1533. 54:13the log of the probability.
  1534. 54:15And so, the log of the probability here,
  1535. 54:17I have a log of X from 0 to 1.
  1536. 54:19The log is a you see here monotonic
  1537. 54:21transformation of the probability,
  1538. 54:24where if you pass in 1, you get 0.
  1539. 54:28So, probability 1 gets you log
  1540. 54:30probability of 0.
  1541. 54:32And then, as you go lower and lower
  1542. 54:33probability, the log will grow more and
  1543. 54:35more negative until all the way to
  1544. 54:37negative infinity at 0.
  1545. 54:41So, here we have a log prob, which is
  1546. 54:44really just a torch. log of probability.
  1547. 54:46Let's print it out to get a sense of
  1548. 54:48what that looks like.
  1549. 54:50Log prob,
  1550. 54:51also point four if.
  1551. 54:54Okay.
  1552. 54:56So, as you can see, when we plug in
  1553. 54:58numbers that are very close, some of our
  1554. 55:00higher numbers, we get closer and closer
  1555. 55:02to 0.
  1556. 55:03And then, if we plug in very bad
  1557. 55:05probabilities, we get more and more
  1558. 55:06negative number. That's bad.
  1559. 55:09So,
  1560. 55:10and the reason we work with this is for
  1561. 55:13large extent convenience, right? Because
  1562. 55:15we have mathematically that if you have
  1563. 55:17some product A * B * C of all these
  1564. 55:19probabilities, right? Or the likelihood
  1565. 55:22is the product of all these
  1566. 55:23probabilities,
  1567. 55:25then the log
  1568. 55:27of these is just log of A plus log of B
  1569. 55:33plus log of C. If you remember your logs
  1570. 55:36from your
  1571. 55:37high school or undergrad and so on.
  1572. 55:39So we have that basically
  1573. 55:41the likelihood is the product of
  1574. 55:42probabilities, the log likelihood is
  1575. 55:44just the sum of the logs of the
  1576. 55:46individual probabilities.
  1577. 55:48So
  1578. 55:49log likelihood
  1579. 55:52starts at zero.
  1580. 55:54And then log likelihood here we can just
  1581. 55:57accumulate simply.
  1582. 56:00And in the end we can print this.
  1583. 56:05Print the log likelihood.
  1584. 56:09F strings.
  1585. 56:11Maybe you're familiar with this.
  1586. 56:13So log likelihood is -38.
  1587. 56:19Okay.
  1588. 56:21Now
  1589. 56:22we actually want um
  1590. 56:25So how high can log likelihood get? It
  1591. 56:27can go to zero. So when all the
  1592. 56:30probabilities are one, log likelihood
  1593. 56:31will be zero. And then when all the
  1594. 56:33probabilities are lower, this will grow
  1595. 56:35more and more negative.
  1596. 56:37Now we don't actually like this because
  1597. 56:39what we'd like is a loss function. And a
  1598. 56:41loss function has the semantics that low
  1599. 56:44is good. Because we're trying to
  1600. 56:46minimize the loss. So we actually need
  1601. 56:49to invert this. And that's what gives us
  1602. 56:51something called the negative log
  1603. 56:53likelihood.
  1604. 56:54Um
  1605. 56:55negative log likelihood is just negative
  1606. 56:58of the log likelihood.
  1607. 57:03These are F strings by the way if you'd
  1608. 57:05like to look this up.
  1609. 57:06Negative log likelihood equals
  1610. 57:09So negative log likelihood now is just
  1611. 57:10the negative of it.
  1612. 57:12And so the negative log likelihood is a
  1613. 57:13very nice loss function because um
  1614. 57:17the lowest it can get is zero. And the
  1615. 57:20higher it is, the worse off the
  1616. 57:22predictions are that you're making.
  1617. 57:24And then one more modification to this
  1618. 57:26that sometimes people do is that for
  1619. 57:27convenience uh they actually like to
  1620. 57:29normalize by they like to make it an
  1621. 57:32average instead of a sum.
  1622. 57:34And so here
  1623. 57:37let's just keep some counts as well.
  1624. 57:39So n plus equals one starts at zero. And
  1625. 57:42then here we can have sort of like a
  1626. 57:45normalized log likelihood.
  1627. 57:47Um
  1628. 57:50if we just normalize it by the count
  1629. 57:52then we will sort of get the average log
  1630. 57:54likelihood. So this would be usually our
  1631. 57:57loss function here.
  1632. 57:58This this we would this is what we would
  1633. 58:00use.
  1634. 58:02Uh so our loss function for the training
  1635. 58:03set assigned by the model is 2.4. That's
  1636. 58:06the quality of this model.
  1637. 58:08And the lower it is, the better off we
  1638. 58:10are. And the higher it is, the worse off
  1639. 58:12we are.
  1640. 58:13And the job of our, you know, training
  1641. 58:16is to find the parameters that minimize
  1642. 58:19the negative log likelihood loss.
  1643. 58:22And that would be like a high quality
  1644. 58:24model. Okay, so to summarize I actually
  1645. 58:26wrote it out here.
  1646. 58:28So our goal is to maximize likelihood
  1647. 58:30which is the product of all the
  1648. 58:32probabilities assigned by the model.
  1649. 58:35And we want to maximize this likelihood
  1650. 58:37with respect to the model parameters.
  1651. 58:39And in our case, the model parameters
  1652. 58:41here are defined in the table. These
  1653. 58:43numbers, the probabilities
  1654. 58:45are
  1655. 58:46uh the model parameters sort of in our
  1656. 58:47bigram language model so far.
  1657. 58:50Uh but you have to keep in mind that
  1658. 58:51here we are storing everything in a
  1659. 58:52table format, the probabilities. But
  1660. 58:54what's coming up as a brief preview is
  1661. 58:57that these numbers will not be kept
  1662. 58:59explicitly but these numbers will be
  1663. 59:01calculated by a neural network.
  1664. 59:03So that's coming up.
  1665. 59:04And we want to change and tune the
  1666. 59:06parameters of these neural networks. We
  1667. 59:08want to change these parameters to
  1668. 59:09maximize the likelihood, the product of
  1669. 59:11the probabilities.
  1670. 59:13Now maximizing the likelihood is
  1671. 59:15equivalent to maximizing the log
  1672. 59:16likelihood because log is a monotonic
  1673. 59:18function.
  1674. 59:19Here's the graph of log.
  1675. 59:22And basically all it is doing is it's uh
  1676. 59:24just a scaling your um you can look at
  1677. 59:27it as just a scaling of the loss
  1678. 59:28function.
  1679. 59:29And so the optimization problem here and
  1680. 59:32here are actually equivalent because
  1681. 59:34this is just a scaling. You can look at
  1682. 59:35it that way.
  1683. 59:37And so these are two identical
  1684. 59:38optimization problems.
  1685. 59:41Um
  1686. 59:41maximizing the log likelihood is
  1687. 59:43equivalent to minimizing the negative
  1688. 59:44log likelihood.
  1689. 59:46And then in practice people actually
  1690. 59:47minimize the average negative log
  1691. 59:49likelihood to get numbers like 2.4.
  1692. 59:53And then this summarizes the quality of
  1693. 59:55your model.
  1694. 59:56And we'd like to minimize it and make it
  1695. 59:57as small as possible.
  1696. 59:59And the lowest it can get is zero.
  1697. 1:00:02And the lower it is
  1698. 1:00:04the better off your model is because
  1699. 1:00:06it's assigning it's assigning high
  1700. 1:00:07probabilities to your data.
  1701. 1:00:09Now let's estimate the probability over
  1702. 1:00:10the entire training set just to make
  1703. 1:00:12sure that we get something around 2.4.
  1704. 1:00:15Let's run this over the entire oops.
  1705. 1:00:17Let's take out the print statement as
  1706. 1:00:18well.
  1707. 1:00:20Okay, 2.45 over the entire training set.
  1708. 1:00:24Now what I'd like to show you is that
  1709. 1:00:25you can actually evaluate the
  1710. 1:00:26probability for any word that you want.
  1711. 1:00:28Like for example
  1712. 1:00:30if we just test a single word Andre and
  1713. 1:00:32bring back the print statement
  1714. 1:00:35then you see that Andre is actually kind
  1715. 1:00:37of like an unlikely word. Like on
  1716. 1:00:39average
  1717. 1:00:40um we take three log probability to
  1718. 1:00:43represent it. And roughly that's because
  1719. 1:00:45EJ apparently is very uncommon as an
  1720. 1:00:47example.
  1721. 1:00:50Now
  1722. 1:00:51think through this um
  1723. 1:00:53when I take Andre and I append Q and I
  1724. 1:00:55test the probability of it Andre Q
  1725. 1:01:00we actually get um infinity.
  1726. 1:01:03And that's because JQ has a 0%
  1727. 1:01:05probability according to our model. So
  1728. 1:01:07the log likelihood
  1729. 1:01:09so the log of zero will be negative
  1730. 1:01:11infinity. We get infinite loss.
  1731. 1:01:14So this is kind of undesirable, right?
  1732. 1:01:15Because we plugged in a string that
  1733. 1:01:16could be like a somewhat reasonable
  1734. 1:01:18name. But basically what this is saying
  1735. 1:01:20is that this model is exactly 0% likely
  1736. 1:01:23to uh to predict this name
  1737. 1:01:26and our loss is infinity on this
  1738. 1:01:28example.
  1739. 1:01:29And really what the reason for that is
  1740. 1:01:31that J
  1741. 1:01:32is followed by Q
  1742. 1:01:35uh zero times. Uh where is Q? JQ is zero
  1743. 1:01:39and so JQ is 0% likely.
  1744. 1:01:42So it's actually kind of gross and
  1745. 1:01:43people don't like this too much. To fix
  1746. 1:01:45this, there's a very simple fix that
  1747. 1:01:47people like to do to sort of like smooth
  1748. 1:01:49out your model a little bit that is
  1749. 1:01:50called model smoothing.
  1750. 1:01:52And roughly what's happening is that we
  1751. 1:01:53will we will add some fake counts.
  1752. 1:01:56So imagine adding a count of one to
  1753. 1:01:59everything.
  1754. 1:02:00So we add a count of one
  1755. 1:02:03like this
  1756. 1:02:04and then we recalculate the
  1757. 1:02:05probabilities.
  1758. 1:02:07And that's model smoothing and you can
  1759. 1:02:09add as much as you like. You can add
  1760. 1:02:10five and that will give you a smoother
  1761. 1:02:11model.
  1762. 1:02:12And the more you add here
  1763. 1:02:14the more uniform model you're going to
  1764. 1:02:16have. And the less you add,
  1765. 1:02:19um the more more peaked model you're
  1766. 1:02:21going to have, of course.
  1767. 1:02:22So one is like a pretty decent count to
  1768. 1:02:24add and that will ensure that there will
  1769. 1:02:27be no zeros in our probability matrix P.
  1770. 1:02:30And so this will of course change the
  1771. 1:02:32generations a little bit. In this case
  1772. 1:02:34it didn't but it in principle it could.
  1773. 1:02:36But what that's going to do now is that
  1774. 1:02:38nothing will be infinity unlikely.
  1775. 1:02:41So now
  1776. 1:02:42our model will predict some other
  1777. 1:02:44probability and we see that JQ now has a
  1778. 1:02:46very small probability. So the model
  1779. 1:02:48still finds it very surprising that this
  1780. 1:02:49was a word or bigram but we don't get
  1781. 1:02:52negative infinity.
  1782. 1:02:53So it's kind of like a nice fix that
  1783. 1:02:54people like to apply sometimes and it's
  1784. 1:02:55called model smoothing. Okay, so we've
  1785. 1:02:57now trained a respectable bigram
  1786. 1:03:00character level language model and we
  1787. 1:03:01saw that we both
  1788. 1:03:04a sort of trained the model by looking
  1789. 1:03:05at the counts of all the bigrams and
  1790. 1:03:08normalizing the rows to get probability
  1791. 1:03:10distributions.
  1792. 1:03:11We saw that we can also then use those
  1793. 1:03:14parameters of this model to perform
  1794. 1:03:16sampling of new words.
  1795. 1:03:19So, we sample new names according to
  1796. 1:03:21those distributions. And we also saw
  1797. 1:03:22that we can evaluate the quality of this
  1798. 1:03:24model.
  1799. 1:03:25And the quality of this model is
  1800. 1:03:26summarized in a single number, which is
  1801. 1:03:28the negative log likelihood. And the
  1802. 1:03:30lower this number is, the better the
  1803. 1:03:32model is
  1804. 1:03:33because it is giving high probabilities
  1805. 1:03:35to the actual next characters in all the
  1806. 1:03:37bigrams in our training set.
  1807. 1:03:40So, that's all well and good. But we've
  1808. 1:03:42arrived at this model explicitly by
  1809. 1:03:44doing something that felt sensible. We
  1810. 1:03:46were just performing counts, and then we
  1811. 1:03:48were normalizing those counts.
  1812. 1:03:51Now, what I would like to do is I would
  1813. 1:03:52like to take an alternative approach. We
  1814. 1:03:54will end up in a very, very similar
  1815. 1:03:55position, but the approach will look
  1816. 1:03:57very different because I would like to
  1817. 1:03:58cast the problem of bigram character
  1818. 1:04:00level language modeling into the neural
  1819. 1:04:02network framework.
  1820. 1:04:04And in the neural network framework,
  1821. 1:04:05we're going to approach things slightly
  1822. 1:04:07differently, but again, end up in a very
  1823. 1:04:09similar spot. I'll go into that later.
  1824. 1:04:12Now, our neural network is going to be a
  1825. 1:04:15still a bigram character level language
  1826. 1:04:16model. So, it receives a single
  1827. 1:04:18character as an input.
  1828. 1:04:20Then there's neural network with some
  1829. 1:04:21weights or some parameters W.
  1830. 1:04:24And it's going to output the probability
  1831. 1:04:26distribution over the next character in
  1832. 1:04:28a sequence. It's going to make guesses
  1833. 1:04:30as to what is likely to follow this
  1834. 1:04:32character that was input to the model.
  1835. 1:04:36And then in addition to that, we're
  1836. 1:04:37going to be able to evaluate any setting
  1837. 1:04:39of the parameters of the neural net
  1838. 1:04:41because we have the loss function.
  1839. 1:04:43The negative log likelihood. So, we're
  1840. 1:04:45going to take a look at these
  1841. 1:04:46probability distributions, and we're
  1842. 1:04:47going to use the labels
  1843. 1:04:49which are basically just the identity of
  1844. 1:04:51the next character in that bigram, the
  1845. 1:04:53second character.
  1846. 1:04:54So, knowing what second character
  1847. 1:04:56actually comes next in the bigram allows
  1848. 1:04:58us to then look at what how high of a
  1849. 1:05:00probability the model assigns to that
  1850. 1:05:02character.
  1851. 1:05:03And then we of course want the
  1852. 1:05:05probability to be very high.
  1853. 1:05:07And that is another way of saying that
  1854. 1:05:08the loss is low.
  1855. 1:05:10So, we're going to use gradient-based
  1856. 1:05:12optimization then to tune the parameters
  1857. 1:05:14of this network because we have the loss
  1858. 1:05:16function and we're going to minimize it.
  1859. 1:05:18So, we're going to tune the weights so
  1860. 1:05:20that the neural net is correctly
  1861. 1:05:21predicting the probabilities for the
  1862. 1:05:23next character.
  1863. 1:05:24So, let's get started. The first thing I
  1864. 1:05:26want to do is I want to compile the
  1865. 1:05:27training set of this neural network,
  1866. 1:05:29right? So, create
  1867. 1:05:31the training set
  1868. 1:05:33of all the bigrams.
  1869. 1:05:36Okay?
  1870. 1:05:37And
  1871. 1:05:39here
  1872. 1:05:40I'm going to copy-paste this code
  1873. 1:05:43because this code iterates over all the
  1874. 1:05:45bigrams.
  1875. 1:05:47So, here we start with the words. We
  1876. 1:05:49iterate over all the bigrams. And
  1877. 1:05:50previously, as you recall, we did the
  1878. 1:05:52counts. But now we're not going to do
  1879. 1:05:54counts. We're just creating a training
  1880. 1:05:55set.
  1881. 1:05:56Now, this training set will be made up
  1882. 1:05:58of two lists.
  1883. 1:06:02We have the
  1884. 1:06:04inputs
  1885. 1:06:06and the targets, the the labels.
  1886. 1:06:09And these bigrams will denote XY. Those
  1887. 1:06:11are the characters, right?
  1888. 1:06:13And so, we're given the first character
  1889. 1:06:14of the bigram and then we're trying to
  1890. 1:06:16predict the next one.
  1891. 1:06:17Both of these are going to be integers.
  1892. 1:06:19So, here we'll take X's. data.append is
  1893. 1:06:22just
  1894. 1:06:23X1. Y's.data.append IX2.
  1895. 1:06:27And then here
  1896. 1:06:29we actually don't want lists of
  1897. 1:06:30integers. We will create uh tensors out
  1898. 1:06:33of these. So, X's is torch.tensor
  1899. 1:06:36X's and Y's is torch.tensor of Y's.
  1900. 1:06:41And then we don't actually want to take
  1901. 1:06:43all the words just yet because I want
  1902. 1:06:45everything to be manageable. Uh so,
  1903. 1:06:47let's just do the first word, which is
  1904. 1:06:48Emma.
  1905. 1:06:51And then it's clear what these X's and
  1906. 1:06:52Y's would be.
  1907. 1:06:55Here, let me print
  1908. 1:06:57character one, character two, just so
  1909. 1:06:59you see what's going on here.
  1910. 1:07:01So, the bigrams of these characters is
  1911. 1:07:04.e, em, mm, ma, a. So, so this single
  1912. 1:07:09word, as I mentioned, has 1 2 3 4 5
  1913. 1:07:12examples for our neural network.
  1914. 1:07:14There are five separate examples in
  1915. 1:07:16Emma.
  1916. 1:07:17And those examples are summarized here.
  1917. 1:07:19When the input to the neural neural
  1918. 1:07:20network is integer zero,
  1919. 1:07:23the desired label is integer five, which
  1920. 1:07:26corresponds to e.
  1921. 1:07:28When the input to the neural network is
  1922. 1:07:29five, we want its weights to be arranged
  1923. 1:07:32so that 13 gets a very high probability.
  1924. 1:07:35When 13 is put in, we want 13 to have a
  1925. 1:07:37high probability.
  1926. 1:07:39When 13 is put in, we also want one to
  1927. 1:07:41have a high probability.
  1928. 1:07:43When one is input, we want zero to have
  1929. 1:07:45a very high probability. So, there are
  1930. 1:07:47five separate input examples to a neural
  1931. 1:07:50net
  1932. 1:07:51in this data set.
  1933. 1:07:55I wanted to add a tangent of a note of
  1934. 1:07:57caution to be careful with a lot of the
  1935. 1:07:59APIs of some of these frameworks.
  1936. 1:08:01You saw me silently use torch.Tensor
  1937. 1:08:04with a lowercase t, and the output
  1938. 1:08:06looked right.
  1939. 1:08:07But, you should be aware that there's
  1940. 1:08:09actually two ways of constructing a
  1941. 1:08:10tensor. There's a torch.lowercase
  1942. 1:08:13tensor, and there's also a torch.capital
  1943. 1:08:15tensor class, which you can also
  1944. 1:08:17construct. Uh so, you can actually call
  1945. 1:08:19both. You can also do torch.capital
  1946. 1:08:21tensor,
  1947. 1:08:22and you get an x's and y's as well.
  1948. 1:08:25So, that's not confusing at all.
  1949. 1:08:27Um
  1950. 1:08:29There are threads on what is the
  1951. 1:08:29difference between these two.
  1952. 1:08:31And um
  1953. 1:08:33unfortunately, the docs are just like
  1954. 1:08:34not clear on the difference. And when
  1955. 1:08:36you look at the the docs of lowercase
  1956. 1:08:38tensor, constructs tensor with no
  1957. 1:08:40autograd history by copying data.
  1958. 1:08:43It's just like it doesn't
  1959. 1:08:45it doesn't make sense. So, the actual
  1960. 1:08:47difference, as far as I can tell, is
  1961. 1:08:48explained eventually in this random
  1962. 1:08:50thread that you can Google.
  1963. 1:08:51And really, it comes down to, I believe,
  1964. 1:08:55that um
  1965. 1:08:56where is this?
  1966. 1:08:58torch.Tensor infers the dtype, the data
  1967. 1:09:00type, automatically, while torch.Tensor
  1968. 1:09:02just returns a float tensor.
  1969. 1:09:04I would recommend stick to torch.lower
  1970. 1:09:06case tensor.
  1971. 1:09:07So, um
  1972. 1:09:09indeed, we see that when I construct
  1973. 1:09:12this with a capital T, the data type
  1974. 1:09:14here of X's is float 32.
  1975. 1:09:18But, torch.lower case tensor
  1976. 1:09:21you see how it's now X.dtype is now
  1977. 1:09:24integer.
  1978. 1:09:26So, um
  1979. 1:09:28it's advised that you use lower case T
  1980. 1:09:30and you can read more about it if you
  1981. 1:09:32like in some of these threads.
  1982. 1:09:34Uh but basically
  1983. 1:09:35um
  1984. 1:09:36I'm pointing out some of these things
  1985. 1:09:37beca- because I want to caution you and
  1986. 1:09:39I want you to re- get used to reading a
  1987. 1:09:41lot of documentation and reading through
  1988. 1:09:43a lot of uh Q&A's and threads like this.
  1989. 1:09:46And um
  1990. 1:09:48you know, some of this stuff is
  1991. 1:09:49unfortunately not easy and not very well
  1992. 1:09:50documented and you have to be careful
  1993. 1:09:51out there. What we want here is integers
  1994. 1:09:54because that's what makes uh sense. Um
  1995. 1:09:58and so uh lower case tensor is what we
  1996. 1:10:00are using. Okay, now we want to think
  1997. 1:10:02through how we're going to feed in these
  1998. 1:10:03examples into a neural network.
  1999. 1:10:06Now, it's not quite as straightforward
  2000. 1:10:07as n-
  2001. 1:10:09plugging it in because these examples
  2002. 1:10:11right now are integers. So, there's like
  2003. 1:10:12a 0, 5, or 13. It gives us the index of
  2004. 1:10:15the character and you can't just plug an
  2005. 1:10:17integer index into a neural net.
  2006. 1:10:20These neural nets uh right are sort of
  2007. 1:10:22made up of these neurons.
  2008. 1:10:24And uh these neurons have weights. And
  2009. 1:10:27as you saw in micrograd, these weights
  2010. 1:10:29act multiplicatively on the inputs. WX
  2011. 1:10:31plus B, there's tanh's and so on. And
  2012. 1:10:34so, it doesn't really make sense to make
  2013. 1:10:35an input neuron take on integer values
  2014. 1:10:37that you feed in and then multiply on
  2015. 1:10:40with weights.
  2016. 1:10:41So, instead, a common way of encoding
  2017. 1:10:44integers is what's called one-hot
  2018. 1:10:45encoding.
  2019. 1:10:47In one-hot encoding, uh we take an
  2020. 1:10:49integer like 13 and we create a vector
  2021. 1:10:52that is all zeros except for the 13th
  2022. 1:10:54dimension, which we turn to a one.
  2023. 1:10:57And then that vector can feed into a
  2024. 1:10:59neural net.
  2025. 1:11:01Now, conveniently,
  2026. 1:11:03uh PyTorch actually has something called
  2027. 1:11:04the one-hot uh uh
  2028. 1:11:07um
  2029. 1:11:07function inside torch.nn.functional.
  2030. 1:11:10It takes a tensor made up of integers.
  2031. 1:11:13Um
  2032. 1:11:14long is a is a is an integer.
  2033. 1:11:18Um
  2034. 1:11:19and it also takes a number of classes,
  2035. 1:11:21um
  2036. 1:11:22which is how large you want your uh
  2037. 1:11:24tensor uh your vector to be.
  2038. 1:11:27So here, let's import is a common way of
  2039. 1:11:30importing it.
  2040. 1:11:34And then let's do F.one_hot.
  2041. 1:11:36And we feed in uh the integers that we
  2042. 1:11:38want to encode.
  2043. 1:11:40So we can actually feed in the entire
  2044. 1:11:41array of X's.
  2045. 1:11:44And we can tell it that num_classes is
  2046. 1:11:4627.
  2047. 1:11:47So it doesn't have to try to guess it.
  2048. 1:11:49It may have guessed that it's only 13
  2049. 1:11:51and would give us an incorrect result.
  2050. 1:11:54So this is the one-hot. Let's call this
  2051. 1:11:56X_enc for X encoded.
  2052. 1:12:02And then we see that X_encoded.shape is
  2053. 1:12:045 by 27.
  2054. 1:12:07And uh
  2055. 1:12:08we can also visualize it, plt.imshow of
  2056. 1:12:10X_enc,
  2057. 1:12:12to make it a little bit more clear
  2058. 1:12:13because this is a little messy.
  2059. 1:12:15So we see that we've encoded all the
  2060. 1:12:17five examples uh into vectors.
  2061. 1:12:20We have five examples, so we have five
  2062. 1:12:22rows, and each row here is now an
  2063. 1:12:24example into a neural net.
  2064. 1:12:26And we see that the appropriate bit is
  2065. 1:12:28turned on as a one, and everything else
  2066. 1:12:30is zero.
  2067. 1:12:31So um
  2068. 1:12:33here for example, the the zeroth bit is
  2069. 1:12:35turned on, the fifth bit is turned on,
  2070. 1:12:3813th bits are turned on for both of
  2071. 1:12:40these examples, and then the first bit
  2072. 1:12:42here is turned on.
  2073. 1:12:44So that's how we can encode um integers
  2074. 1:12:47into vectors.
  2075. 1:12:49And then these vectors can feed in to
  2076. 1:12:51neural nets. One more issue to be
  2077. 1:12:52careful with here, by the way, is
  2078. 1:12:55let's look at the data type of X
  2079. 1:12:56encoding. We always want to be careful
  2080. 1:12:58with data types.
  2081. 1:12:59What would you expect X encoding's data
  2082. 1:13:01type to be? When we're plugging numbers
  2083. 1:13:03into neural nets, we don't want them to
  2084. 1:13:05be integers. We want them to be floating
  2085. 1:13:07point numbers that can take on various
  2086. 1:13:09values. But the D type here is actually
  2087. 1:13:1264-bit integer.
  2088. 1:13:14And the reason for that, I suspect, is
  2089. 1:13:15that one hot received a 64-bit integer
  2090. 1:13:18here and it returned to the same data
  2091. 1:13:21type.
  2092. 1:13:21And when you look at the signature of
  2093. 1:13:23one hot, it doesn't even take a D type,
  2094. 1:13:25a desired data type of the output
  2095. 1:13:27tensor.
  2096. 1:13:28And so we can't In a lot of functions in
  2097. 1:13:30torch, we'd be able to do something like
  2098. 1:13:32D type equals torch.float32,
  2099. 1:13:34which is what we want, but one hot does
  2100. 1:13:36not support that.
  2101. 1:13:37So instead, we're going to want to cast
  2102. 1:13:39this to float like this.
  2103. 1:13:43So that these
  2104. 1:13:44everything is the same.
  2105. 1:13:46Everything looks the same, but the D
  2106. 1:13:48type is float 32. And floats can feed
  2107. 1:13:51into um neural nets. So now let's
  2108. 1:13:53construct our first neuron.
  2109. 1:13:56This neuron will look at these input
  2110. 1:13:58vectors.
  2111. 1:14:00And as you remember from micrograd,
  2112. 1:14:02these neurons basically perform a very
  2113. 1:14:03simple function, wx plus b, where wx is
  2114. 1:14:06a dot product, right?
  2115. 1:14:09So we can achieve the same thing here.
  2116. 1:14:12Let's first define the weights of this
  2117. 1:14:14neuron. Basically, where are the initial
  2118. 1:14:15weights at initialization for this
  2119. 1:14:17neuron?
  2120. 1:14:18Let's initialize them with torch.randn.
  2121. 1:14:21torch.randn
  2122. 1:14:23is um
  2123. 1:14:24fills a tensor with random numbers
  2124. 1:14:27drawn from a normal distribution.
  2125. 1:14:29And a normal distribution has a
  2126. 1:14:32probability uh density function like
  2127. 1:14:33this. And so most of the numbers drawn
  2128. 1:14:35from this distribution will be around
  2129. 1:14:37zero,
  2130. 1:14:38uh but some of them will be as high as
  2131. 1:14:40almost three and so on. And very few
  2132. 1:14:42numbers will be above three in
  2133. 1:14:44magnitude.
  2134. 1:14:46So we need to take a size as an input
  2135. 1:14:49here.
  2136. 1:14:50And I'm going to use size as 27 by 1.
  2137. 1:14:54So, 27 by 1, and then let's visualize W.
  2138. 1:14:58So, W is a column vector of 27 numbers.
  2139. 1:15:03And uh these weights are then multiplied
  2140. 1:15:06by the inputs.
  2141. 1:15:08So, now to perform this multiplication,
  2142. 1:15:10we can take X encoding
  2143. 1:15:12and we can multiply it with W.
  2144. 1:15:15This is a matrix multiplication operator
  2145. 1:15:17in PyTorch.
  2146. 1:15:20And the output of this operation is 5 by
  2147. 1:15:221.
  2148. 1:15:23The reason it's 5 by 1 is the following.
  2149. 1:15:25We took X encoding, which is 5 by 27,
  2150. 1:15:29and we multiplied it by 27 by 1.
  2151. 1:15:33And
  2152. 1:15:34in matrix multiplication,
  2153. 1:15:36you see that the output will become 5 by
  2154. 1:15:391 because these 27 will multiply and
  2155. 1:15:43add.
  2156. 1:15:44So, basically what we're seeing here,
  2157. 1:15:46out out of this operation,
  2158. 1:15:48is we are seeing the five um
  2159. 1:15:51activations
  2160. 1:15:53of this neuron
  2161. 1:15:56on these five inputs. And we've
  2162. 1:15:58evaluated all of them in parallel. We
  2163. 1:16:00didn't feed in just a single input to
  2164. 1:16:02this single neuron. We fed in
  2165. 1:16:04simultaneously all the five inputs into
  2166. 1:16:06the same neuron,
  2167. 1:16:08and in parallel, PyTorch has evaluated
  2168. 1:16:11the WX plus B, but here it's just WX.
  2169. 1:16:14There's no bias.
  2170. 1:16:15It has valued W W times X for all of
  2171. 1:16:18them uh independently. Now, instead of a
  2172. 1:16:21single neuron though, I would like to
  2173. 1:16:22have 27 neurons, and I'll show you in a
  2174. 1:16:24second why I want 27 neurons.
  2175. 1:16:27So, instead of having just a one here,
  2176. 1:16:29which is indicating this presence of one
  2177. 1:16:31single neuron,
  2178. 1:16:32we can use 27.
  2179. 1:16:34And then when W is 27 by 27,
  2180. 1:16:38this will in parallel evaluate all the
  2181. 1:16:4127 neurons on all the five inputs.
  2182. 1:16:46Giving us a much better, much much
  2183. 1:16:48bigger result.
  2184. 1:16:49So, now what we've done is 5 by 27
  2185. 1:16:51multiplied 27 by 27.
  2186. 1:16:54And the output of this is now 5 by 27.
  2187. 1:16:57So, we can see that the shape of this
  2188. 1:17:01is 5 by 27.
  2189. 1:17:04So, what is every element here telling
  2190. 1:17:05us, right?
  2191. 1:17:07It's telling us for every one of 27
  2192. 1:17:09neurons that we created,
  2193. 1:17:13what is the firing rate of those neurons
  2194. 1:17:16on every one of those five examples?
  2195. 1:17:19So,
  2196. 1:17:20the element, for example, 3 {comma} 13
  2197. 1:17:25is giving us the firing rate of the 13th
  2198. 1:17:28neuron looking at the third input.
  2199. 1:17:31And the way this was achieved is by a
  2200. 1:17:34dot product
  2201. 1:17:36between the third input
  2202. 1:17:38and the 13th column
  2203. 1:17:41of this W matrix here.
  2204. 1:17:44Okay? So, using matrix multiplication,
  2205. 1:17:47we can very efficiently evaluate
  2206. 1:17:50the dot product between lots of input
  2207. 1:17:52examples in a batch.
  2208. 1:17:55And lots of neurons, where all of those
  2209. 1:17:57neurons have weights in the columns of
  2210. 1:17:59those Ws.
  2211. 1:18:01And in matrix multiplication, we're just
  2212. 1:18:02doing those dot products and
  2213. 1:18:04in parallel. Just to show you that this
  2214. 1:18:06is the case, we can take X and we can
  2215. 1:18:08take the third
  2216. 1:18:10row.
  2217. 1:18:12And we can take the W and take its 13th
  2218. 1:18:14column.
  2219. 1:18:17And then we can do X and get three.
  2220. 1:18:21Element-wise multiply with W at 13.
  2221. 1:18:26And sum that up. That's WX plus B.
  2222. 1:18:29Uh well, there's no plus B. It's just WX
  2223. 1:18:31dot product. And that's
  2224. 1:18:34this number.
  2225. 1:18:35So, you see that this is just being done
  2226. 1:18:36efficiently by the matrix multiplication
  2227. 1:18:39operation for all the input examples and
  2228. 1:18:42for all the output neurons of this first
  2229. 1:18:45layer.
  2230. 1:18:46Okay, so we fed our 27-dimensional
  2231. 1:18:48inputs into a first layer of a neural
  2232. 1:18:50net that has 27 neurons, right? So, we
  2233. 1:18:53have 27 inputs and now we have 27
  2234. 1:18:56neurons. These neurons perform W * X.
  2235. 1:18:59They don't have a bias and they don't
  2236. 1:19:01have a non-linearity like tanh. We're
  2237. 1:19:03going to leave them to be a linear
  2238. 1:19:05layer.
  2239. 1:19:06In addition to that, we're not going to
  2240. 1:19:08have any other layers. This is going to
  2241. 1:19:09be it. It's just going to be
  2242. 1:19:11the dumbest, smallest, simplest neural
  2243. 1:19:13net, which is just a single linear
  2244. 1:19:14layer.
  2245. 1:19:16And now I'd like to explain what I want
  2246. 1:19:18those 27 outputs to be.
  2247. 1:19:21Intuitively, what we're trying to
  2248. 1:19:22produce here for every single input
  2249. 1:19:23example is we're trying to produce some
  2250. 1:19:25kind of a probability distribution for
  2251. 1:19:27the next character in a sequence. And
  2252. 1:19:29there's 27 of them.
  2253. 1:19:31But we have to come up with like precise
  2254. 1:19:33semantics for exactly how we're going to
  2255. 1:19:34interpret these 27 numbers that these
  2256. 1:19:37neurons take on.
  2257. 1:19:39Now, intuitively,
  2258. 1:19:41you see here that these numbers are
  2259. 1:19:42negative and some of them are positive,
  2260. 1:19:44etc.
  2261. 1:19:45And that's because these are coming out
  2262. 1:19:46of a neural net layer initialized with
  2263. 1:19:48these um
  2264. 1:19:50uh normal distribution
  2265. 1:19:52uh parameters.
  2266. 1:19:54But what we want is we want something
  2267. 1:19:55like we had here. Like each row here
  2268. 1:19:59told us the counts and then we
  2269. 1:20:01normalized the counts to get
  2270. 1:20:02probabilities. And we want something
  2271. 1:20:04similar to come out of a neural net.
  2272. 1:20:06But what we just have right now is just
  2273. 1:20:07some negative and positive numbers.
  2274. 1:20:10Now, we want those numbers to somehow
  2275. 1:20:12represent the probabilities for the next
  2276. 1:20:14character.
  2277. 1:20:15But you see that probabilities, they
  2278. 1:20:17they have a special structure. They um
  2279. 1:20:20they're positive numbers and they sum to
  2280. 1:20:21one.
  2281. 1:20:22And so, that doesn't just come out of a
  2282. 1:20:24neural net.
  2283. 1:20:25And then, they can't be counts because
  2284. 1:20:28uh these counts are positive and counts
  2285. 1:20:31are integers.
  2286. 1:20:32So, counts are also not really a good
  2287. 1:20:34thing to output from a neural net.
  2288. 1:20:36So, instead what the neural net is going
  2289. 1:20:38to output and how we are going to
  2290. 1:20:39interpret the
  2291. 1:20:42the 27 numbers is that these 27 numbers
  2292. 1:20:45are giving us log counts
  2293. 1:20:48basically.
  2294. 1:20:49Um so, instead of giving us counts
  2295. 1:20:52directly like in this table, they're
  2296. 1:20:54giving us log counts.
  2297. 1:20:56And to get the counts, we're going to
  2298. 1:20:57take the log counts and we're going to
  2299. 1:20:59exponentiate them.
  2300. 1:21:01Now,
  2301. 1:21:02exponentiation
  2302. 1:21:04takes the following form.
  2303. 1:21:06Um it takes numbers
  2304. 1:21:08that are negative or they are positive.
  2305. 1:21:10It takes the entire real line. And then
  2306. 1:21:13if you plug in negative numbers, you're
  2307. 1:21:14going to get e to the x, which is
  2308. 1:21:18uh always below one.
  2309. 1:21:20So, you're getting numbers lower than
  2310. 1:21:21one.
  2311. 1:21:23And if you plug in numbers greater than
  2312. 1:21:25zero, you're getting numbers greater
  2313. 1:21:27than one
  2314. 1:21:28all the way growing to the infinity.
  2315. 1:21:30And this here grows to zero.
  2316. 1:21:33So, basically we're going to take these
  2317. 1:21:36numbers
  2318. 1:21:37here.
  2319. 1:21:40And
  2320. 1:21:43instead of them being positive and
  2321. 1:21:44negative in all of the place, we're
  2322. 1:21:46going to interpret them as log counts
  2323. 1:21:48and then we're going to element-wise
  2324. 1:21:50exponentiate these numbers.
  2325. 1:21:52Exponentiating them now gives us
  2326. 1:21:54something like this.
  2327. 1:21:56And you see that these numbers now
  2328. 1:21:57because of they went through an
  2329. 1:21:58exponent, all the negative numbers
  2330. 1:22:00turned into numbers below one like
  2331. 1:22:020.338.
  2332. 1:22:04And all the positive numbers originally
  2333. 1:22:06turned into even more positive numbers
  2334. 1:22:08sort of greater than one.
  2335. 1:22:10Um so, like for example, seven
  2336. 1:22:13um is some positive number over here. Um
  2337. 1:22:18that is greater than zero.
  2338. 1:22:21But, exponentiated outputs here
  2339. 1:22:24um basically give us something that we
  2340. 1:22:26can use and interpret as the equivalent
  2341. 1:22:28of counts or originally. So, you see
  2342. 1:22:31these counts here, 1 12 7 51 1 etc.
  2343. 1:22:36The neural net is kind of now predicting
  2344. 1:22:39uh
  2345. 1:22:40counts.
  2346. 1:22:41And these counts are positive numbers.
  2347. 1:22:44They can never be below zero, so that
  2348. 1:22:45makes sense. And uh they can now take on
  2349. 1:22:48various values
  2350. 1:22:49depending on the settings of W.
  2351. 1:22:54So, let me break this down.
  2352. 1:22:56We're going to interpret these to be the
  2353. 1:22:58log counts.
  2354. 1:23:01Another word for this that is often used
  2355. 1:23:03is so-called logits.
  2356. 1:23:05These are logits, log counts.
  2357. 1:23:08And these will be sort of the counts.
  2358. 1:23:11Logits exponentiated.
  2359. 1:23:13And this is equivalent to the N matrix,
  2360. 1:23:16sort of, the N
  2361. 1:23:18array that we used previously. Remember,
  2362. 1:23:20this was the N.
  2363. 1:23:21This is the the array of counts. And
  2364. 1:23:24each row here are the counts for the
  2365. 1:23:27for the um
  2366. 1:23:28next character, sort of.
  2367. 1:23:32So, those are the counts. And now the
  2368. 1:23:34probabilities are just the counts um
  2369. 1:23:38normalized.
  2370. 1:23:39And so, um
  2371. 1:23:41I'm not going to find the same, but
  2372. 1:23:43basically, I'm not going to scroll all
  2373. 1:23:44over the place.
  2374. 1:23:46We've already done this. We want to
  2375. 1:23:48counts that sum along the first
  2376. 1:23:50dimension, and we want to keep dims as
  2377. 1:23:53true.
  2378. 1:23:54We went over this, and this is how we
  2379. 1:23:56normalize the rows of our counts matrix
  2380. 1:24:00to get our probabilities.
  2381. 1:24:03probs
  2382. 1:24:04So, now these are the probabilities.
  2383. 1:24:07And these are the counts that we have
  2384. 1:24:10currently. And now when I show the
  2385. 1:24:11probabilities,
  2386. 1:24:13you see that um
  2387. 1:24:15every row here,
  2388. 1:24:17of course,
  2389. 1:24:19will sum to one
  2390. 1:24:21because they're normalized.
  2391. 1:24:23And the shape of this
  2392. 1:24:25is 5 by 27.
  2393. 1:24:27And so really what we've achieved is for
  2394. 1:24:29every one of our five examples, we now
  2395. 1:24:32have a row that came out of a neural
  2396. 1:24:34net.
  2397. 1:24:35And because of the transformations here,
  2398. 1:24:37we made sure that this output of this
  2399. 1:24:39neural net now are probabilities or we
  2400. 1:24:41can interpret to be probabilities.
  2401. 1:24:44So,
  2402. 1:24:45our WX here gave us logits
  2403. 1:24:48and then we interpret those to be log
  2404. 1:24:49counts.
  2405. 1:24:50We exponentiate to get something that
  2406. 1:24:52looks like counts.
  2407. 1:24:54And then we normalize those counts to
  2408. 1:24:55get a probability distribution.
  2409. 1:24:57And all of these are differentiable
  2410. 1:24:59operations.
  2411. 1:25:00So, what we've done now is we are taking
  2412. 1:25:02inputs.
  2413. 1:25:03We have differentiable operations that
  2414. 1:25:04we can back propagate through
  2415. 1:25:07and we're getting out probability
  2416. 1:25:08distributions.
  2417. 1:25:09So, um for example, for the zeroth
  2418. 1:25:12example that fed in,
  2419. 1:25:15right, which was um
  2420. 1:25:17the zeroth example here was a one-hot
  2421. 1:25:18vector of zero.
  2422. 1:25:20And um
  2423. 1:25:22it basically corresponded to feeding in
  2424. 1:25:26uh
  2425. 1:25:26this example here. So, we're feeding in
  2426. 1:25:28a dot into a neural net. And the way we
  2427. 1:25:30fed the dot into a neural net is that we
  2428. 1:25:32first got its index.
  2429. 1:25:34Then we one-hot encoded it.
  2430. 1:25:36Then it went into the neural net and out
  2431. 1:25:39came
  2432. 1:25:40this distribution of probabilities.
  2433. 1:25:43And its shape
  2434. 1:25:46is 27. There's 27 numbers and we're
  2435. 1:25:49going to interpret this as the neural
  2436. 1:25:51net's assignment for how likely
  2437. 1:25:54every one of these characters um
  2438. 1:25:56the 27 characters are to come next.
  2439. 1:25:59And as we tune the weights W,
  2440. 1:26:02we're going to be of course getting
  2441. 1:26:03different probabilities out for any
  2442. 1:26:05character that you input.
  2443. 1:26:07And so now the question is just can we
  2444. 1:26:08optimize and find a good W
  2445. 1:26:11such that the probabilities coming out
  2446. 1:26:13are pretty good. And the way we measure
  2447. 1:26:15pretty good is by the loss function.
  2448. 1:26:17Okay, so I organized everything into a
  2449. 1:26:18single summary so that hopefully it's a
  2450. 1:26:20bit more clear. So it starts here.
  2451. 1:26:22We have an input data set.
  2452. 1:26:24We have some inputs to the neural net
  2453. 1:26:26and we have some labels for the correct
  2454. 1:26:28next character in a sequence. And these
  2455. 1:26:30are integers.
  2456. 1:26:32Here I'm using uh torch generators now
  2457. 1:26:35so that you see the same numbers that I
  2458. 1:26:37see.
  2459. 1:26:38And I'm generating
  2460. 1:26:40um
  2461. 1:26:4027 neurons weights
  2462. 1:26:42and each neuron here receives 27 inputs.
  2463. 1:26:48Then here we're going to plug in all the
  2464. 1:26:50input examples x's into a neural net. So
  2465. 1:26:52here, this is a forward pass.
  2466. 1:26:55First, we have to encode all of the
  2467. 1:26:57inputs into one-hot representations.
  2468. 1:27:00So we have 27 classes, we pass in these
  2469. 1:27:02integers and x inc becomes a array that
  2470. 1:27:07is 5 by 27.
  2471. 1:27:09Zeros except for a few ones.
  2472. 1:27:12We then multiply this in the first layer
  2473. 1:27:14of a neural net to get logits.
  2474. 1:27:16Exponentiate the logits to get fake
  2475. 1:27:18counts, sort of.
  2476. 1:27:20And normalize these counts to get
  2477. 1:27:22probabilities.
  2478. 1:27:24So the la- these last two lines by the
  2479. 1:27:26way here are called the softmax.
  2480. 1:27:29Uh which I pulled up here.
  2481. 1:27:32Softmax is a very often used layer in a
  2482. 1:27:34neural net that takes these z's which
  2483. 1:27:37are logits,
  2484. 1:27:38exponentiates them,
  2485. 1:27:40and uh divides and normalizes. It's a
  2486. 1:27:43way of taking outputs of a neural net
  2487. 1:27:45layer and these uh these outputs can be
  2488. 1:27:47positive or negative.
  2489. 1:27:49And it outputs probability
  2490. 1:27:51distributions. It outputs something that
  2491. 1:27:53is always sums to one and are positive
  2492. 1:27:56numbers, just like probabilities.
  2493. 1:27:58Um so it's kind of like a normalization
  2494. 1:28:00function if you want to think of it that
  2495. 1:28:01way. And you can put it on top of any
  2496. 1:28:03other linear layer inside a neural net
  2497. 1:28:05and it basically makes a neural net
  2498. 1:28:07output probabilities. That's very often
  2499. 1:28:09used and we used it as well here.
  2500. 1:28:13So this is the forward pass and that's
  2501. 1:28:14how we made a neural net output
  2502. 1:28:16probability.
  2503. 1:28:17Now
  2504. 1:28:19you'll notice that
  2505. 1:28:20um
  2506. 1:28:23all of these
  2507. 1:28:24this entire forward pass is made up of
  2508. 1:28:26differentiable
  2509. 1:28:27layers. Everything here we can back
  2510. 1:28:29propagate through. And we saw some of
  2511. 1:28:30the back propagation in micrograd.
  2512. 1:28:33This is just
  2513. 1:28:34multiplication and addition. All that's
  2514. 1:28:36happening here is just multiply and then
  2515. 1:28:38add. And we know how to back propagate
  2516. 1:28:39through them.
  2517. 1:28:40Exponentiation we know how to back
  2518. 1:28:42propagate through.
  2519. 1:28:43And then here we are summing and sum is
  2520. 1:28:47is easily back propagatable as well.
  2521. 1:28:50And division as well. So everything here
  2522. 1:28:52is differentiable operation
  2523. 1:28:54and we can back propagate through.
  2524. 1:28:57Now we achieve these probabilities which
  2525. 1:28:59are 5 by 27.
  2526. 1:29:01For every single example we have a
  2527. 1:29:03vector of probabilities that sum to one.
  2528. 1:29:06And then here I wrote a bunch of stuff
  2529. 1:29:08uh to sort of like break down uh the
  2530. 1:29:10examples.
  2531. 1:29:11So we have five examples making up Emma,
  2532. 1:29:14right?
  2533. 1:29:16And there are five bigrams inside Emma.
  2534. 1:29:20So bigram example a bigram example one
  2535. 1:29:23is that E is the beginning character
  2536. 1:29:26right after dot.
  2537. 1:29:28And the indexes for these are zero and
  2538. 1:29:30five.
  2539. 1:29:31So then we feed in a zero.
  2540. 1:29:34That's the input to the neural net.
  2541. 1:29:36We get probabilities from the neural net
  2542. 1:29:38that are 27 numbers.
  2543. 1:29:41And then the label is five because E
  2544. 1:29:44actually comes after dot.
  2545. 1:29:45So that's the label.
  2546. 1:29:47And then
  2547. 1:29:49we use this label five to index into the
  2548. 1:29:52probability distribution here.
  2549. 1:29:54So this index five here is 0 1 2 3 4 5.
  2550. 1:29:59It's this number here.
  2551. 1:30:01Which is here.
  2552. 1:30:04So that's basically the probability
  2553. 1:30:05assigned by the neural net to the actual
  2554. 1:30:07correct character.
  2555. 1:30:08You see that the network currently
  2556. 1:30:10thinks that this next character that E
  2557. 1:30:12following dot is only 1% likely. Which
  2558. 1:30:15is of course not very good, right?
  2559. 1:30:17Because this actually is a training
  2560. 1:30:18example and the network thinks that this
  2561. 1:30:20is currently very very unlikely. But
  2562. 1:30:22that's just because we didn't get very
  2563. 1:30:24lucky in generating a good setting of W.
  2564. 1:30:27So right now this network thinks this is
  2565. 1:30:28unlikely and 0.01 is not a good outcome.
  2566. 1:30:32So the log likelihood then
  2567. 1:30:34is very negative.
  2568. 1:30:36And the negative log likelihood is very
  2569. 1:30:38positive.
  2570. 1:30:39And so four is a very high negative log
  2571. 1:30:42likelihood and that means we're going to
  2572. 1:30:44have a high loss.
  2573. 1:30:45Because what is the loss? The loss is
  2574. 1:30:47just the average negative log
  2575. 1:30:49likelihood.
  2576. 1:30:51So the second character is E M.
  2577. 1:30:53And you see here that also the network
  2578. 1:30:55thought that M following E is very
  2579. 1:30:57unlikely, 1%.
  2580. 1:31:00Uh the for M following M it thought it
  2581. 1:31:02was 2%.
  2582. 1:31:04And for A following M it actually
  2583. 1:31:06thought it was 7% likely. So just by
  2584. 1:31:09chance this one actually has a pretty
  2585. 1:31:11good probability and therefore a pretty
  2586. 1:31:12low negative log likelihood.
  2587. 1:31:15And finally here it thought this was 1%
  2588. 1:31:17likely.
  2589. 1:31:18So overall our average negative log
  2590. 1:31:20likelihood, which is the loss, the total
  2591. 1:31:22loss that summarizes basically the how
  2592. 1:31:25well this network currently works at
  2593. 1:31:27least on this one word, not on the full
  2594. 1:31:29data set, just the one word, is 3.76.
  2595. 1:31:32Which is actually very fairly high loss.
  2596. 1:31:34This is not a very good setting of Ws.
  2597. 1:31:36Now here's what we can do.
  2598. 1:31:38We're currently getting 3.76.
  2599. 1:31:41We can actually come here and we can
  2600. 1:31:42change our W. We can resample it. So let
  2601. 1:31:45me just add one to have a different
  2602. 1:31:47seed.
  2603. 1:31:48And then we get a different W.
  2604. 1:31:50And then we can rerun this.
  2605. 1:31:52And with this different seed with this
  2606. 1:31:54different setting of Ws we now get 3.37.
  2607. 1:31:58So this is a much better W, right? And
  2608. 1:32:00that and it's better because the
  2609. 1:32:02probability just happened to come out
  2610. 1:32:04higher for the for the characters that
  2611. 1:32:07actually are next.
  2612. 1:32:08And so you can imagine actually just
  2613. 1:32:10resampling this, you know, we can try
  2614. 1:32:12two.
  2615. 1:32:14So
  2616. 1:32:15Okay, this was not very good.
  2617. 1:32:17Let's try one more.
  2618. 1:32:18We can try three.
  2619. 1:32:20Okay, this was terrible setting because
  2620. 1:32:22we have a very high loss.
  2621. 1:32:24So
  2622. 1:32:26anyway, I'm going to erase this.
  2623. 1:32:29What what I'm doing here, which is just
  2624. 1:32:31guess and check of randomly assigning
  2625. 1:32:33parameters and seeing if the network is
  2626. 1:32:34good, that is amateur hour. That's not
  2627. 1:32:37how you optimize a neural net. The way
  2628. 1:32:39you optimize a neural net is you start
  2629. 1:32:41with some random guess, and we're going
  2630. 1:32:42to commit to this one even though it's
  2631. 1:32:43not very good.
  2632. 1:32:45But now the big deal is we have a loss
  2633. 1:32:46function.
  2634. 1:32:48So this loss
  2635. 1:32:50is made up only of differentiable
  2636. 1:32:52operations.
  2637. 1:32:54And we can minimize the loss by tuning
  2638. 1:32:57W's by computing the gradients of the
  2639. 1:33:00loss with respect to these W matrices.
  2640. 1:33:05And so then we can tune W to minimize
  2641. 1:33:07the loss and find a good setting of W
  2642. 1:33:09using gradient-based optimization. So
  2643. 1:33:11let's see how that will work. Now things
  2644. 1:33:13are actually going to look almost
  2645. 1:33:14identical to what we had with micrograd.
  2646. 1:33:17So here I pulled up the lecture from
  2647. 1:33:20micrograd, the notebook. It's from this
  2648. 1:33:22repository.
  2649. 1:33:23And when I scroll all the way to the end
  2650. 1:33:25where we left off with micrograd, we had
  2651. 1:33:26something very very similar.
  2652. 1:33:28We had a number of input examples. In
  2653. 1:33:31this case, we had four input examples
  2654. 1:33:32inside X's.
  2655. 1:33:34And we had their targets. These are
  2656. 1:33:36targets.
  2657. 1:33:37Just like here, we have our X's now, but
  2658. 1:33:39we have five of them, and they're now
  2659. 1:33:41integers instead of vectors.
  2660. 1:33:44But we're going to convert our integers
  2661. 1:33:46to vectors, except our vectors will be
  2662. 1:33:4727 large instead of three large.
  2663. 1:33:51And then here what we did is first we
  2664. 1:33:53did a forward pass where where ran a
  2665. 1:33:55neural net on all of the inputs
  2666. 1:33:58to get predictions.
  2667. 1:34:00Our neural net at the time, this NFX,
  2668. 1:34:02was a net a multi-layer perceptron.
  2669. 1:34:05Our neural net is going to look
  2670. 1:34:06different because our neural net is just
  2671. 1:34:08a single layer.
  2672. 1:34:10Single linear layer followed by a
  2673. 1:34:12softmax.
  2674. 1:34:13So, that's our neural net.
  2675. 1:34:15And the loss here was the mean squared
  2676. 1:34:17error. So, we simply subtracted the
  2677. 1:34:19prediction from the ground truth and
  2678. 1:34:21squared it and summed it all up. And
  2679. 1:34:23that was the loss. And loss was the
  2680. 1:34:24single number that summarized the
  2681. 1:34:26quality of the neural net. And when loss
  2682. 1:34:29is low, like almost zero, that means the
  2683. 1:34:32neural net is um
  2684. 1:34:33predicting correctly.
  2685. 1:34:36So, we had a single number that uh that
  2686. 1:34:38summarized the
  2687. 1:34:40uh the performance of the neural net.
  2688. 1:34:42And everything here was differentiable
  2689. 1:34:43and was stored in massive compute graph.
  2690. 1:34:46And then we iterated over all the
  2691. 1:34:48parameters. We made sure that the
  2692. 1:34:50gradients are set to zero.
  2693. 1:34:51And we called loss.backward.
  2694. 1:34:54And loss.backward initiated
  2695. 1:34:55backpropagation at the final output node
  2696. 1:34:58of loss. Right? So,
  2697. 1:35:00yeah, remember these expressions? We had
  2698. 1:35:02loss all the way at the end. We start
  2699. 1:35:03backpropagation and we went all the way
  2700. 1:35:05back.
  2701. 1:35:06And we made sure that we populated all
  2702. 1:35:08the parameters.grad.
  2703. 1:35:10So, that grad started at zero, but
  2704. 1:35:12backpropagation filled it in.
  2705. 1:35:14And then in the update, we iterated over
  2706. 1:35:16all the parameters and we simply did a
  2707. 1:35:18parameter update where every single uh
  2708. 1:35:21element of our parameters was nudged in
  2709. 1:35:24the opposite direction of the gradient.
  2710. 1:35:27And so, we're going to do the exact same
  2711. 1:35:30thing here.
  2712. 1:35:31Uh so, I'm going to pull this up
  2713. 1:35:34on the side here
  2714. 1:35:38so that we have it available. And we're
  2715. 1:35:40actually going to do the exact same
  2716. 1:35:41thing.
  2717. 1:35:42So, this was the forward pass. So, where
  2718. 1:35:44we did this.
  2719. 1:35:46And probs is our Ypred.
  2720. 1:35:49So, now we have to evaluate the loss,
  2721. 1:35:50but we're not using the mean squared
  2722. 1:35:51error. we're using the negative log
  2723. 1:35:53likelihood because we are doing
  2724. 1:35:54classification, we're not doing
  2725. 1:35:56regression, as it's called.
  2726. 1:35:59So, here we want to calculate loss.
  2727. 1:36:02Now, the way we calculate it is is just
  2728. 1:36:04this average negative log likelihood.
  2729. 1:36:07Now, this probs here
  2730. 1:36:10has a shape of 5 by 27.
  2731. 1:36:13And so, to get all the we basically want
  2732. 1:36:15to pluck out the probabilities at the
  2733. 1:36:18correct indices here.
  2734. 1:36:20So, in particular, because the labels
  2735. 1:36:21are stored here in the array wise,
  2736. 1:36:24basically what we're after is for the
  2737. 1:36:26first example, we're looking at
  2738. 1:36:27probability of five, right? At index
  2739. 1:36:30five.
  2740. 1:36:31For the second example,
  2741. 1:36:32at the the second row or row index one,
  2742. 1:36:36we are interested in the probability
  2743. 1:36:37assigned to index 13.
  2744. 1:36:40At the second example, we also have 13.
  2745. 1:36:43At the third row, we want one.
  2746. 1:36:47And at the last row, which is four, we
  2747. 1:36:49want zero. So, these are the
  2748. 1:36:51probabilities we're interested in,
  2749. 1:36:53right?
  2750. 1:36:54And you can see that they're not amazing
  2751. 1:36:56as we saw above.
  2752. 1:36:58So, these are the probabilities we want,
  2753. 1:37:00but we want like a more efficient way to
  2754. 1:37:02access these probabilities. Uh not just
  2755. 1:37:05listing them out in a tuple like this.
  2756. 1:37:07So, it turns out that the way to do this
  2757. 1:37:08in PyTorch, uh one of the ways at least,
  2758. 1:37:10is we can basically pass in all of these
  2759. 1:37:16Sorry about that. All of these um
  2760. 1:37:19integers in a vectors.
  2761. 1:37:22So, the
  2762. 1:37:23these ones, you see how they're just 0 1
  2763. 1:37:252 3 4,
  2764. 1:37:27we can actually create that using MP not
  2765. 1:37:29MP, sorry, torch.arange of five.
  2766. 1:37:320 1 2 3 4.
  2767. 1:37:34So, we can index here with torch.arange
  2768. 1:37:36of five.
  2769. 1:37:38And here, we index with wise.
  2770. 1:37:41And you see that that gives us
  2771. 1:37:43exactly these numbers.
  2772. 1:37:49So, that plugs out the probabilities of
  2773. 1:37:51that the neural network assigns to the
  2774. 1:37:54correct next character.
  2775. 1:37:56Now, we take those probabilities and we
  2776. 1:37:58don't we actually look at the log
  2777. 1:37:59probability. So, we want to dot log.
  2778. 1:38:03And then, we want to just average that
  2779. 1:38:06up. So, take the mean of all of that.
  2780. 1:38:08And then, it's the negative average log
  2781. 1:38:11likelihood that is the loss.
  2782. 1:38:14So, the loss here is
  2783. 1:38:163.7 something. And you see that this
  2784. 1:38:18loss, 3.76, 3.76 is exactly as we've
  2785. 1:38:22obtained before, but this is a
  2786. 1:38:23vectorized form of that expression.
  2787. 1:38:26So, we get the same loss.
  2788. 1:38:29And this same loss we can consider sort
  2789. 1:38:31of as part of this forward pass.
  2790. 1:38:34And we've achieved here now loss.
  2791. 1:38:36Okay, so we made our way all the way to
  2792. 1:38:37loss. We defined the forward pass. We
  2793. 1:38:40forwarded the network and the loss. Now,
  2794. 1:38:42we're ready to do backward pass.
  2795. 1:38:44So, backward pass.
  2796. 1:38:48We want to first make sure that all the
  2797. 1:38:49gradients are reset. So, they're at
  2798. 1:38:51zero.
  2799. 1:38:52Now, in PyTorch, you can set the
  2800. 1:38:55gradients to be zero, but you can also
  2801. 1:38:56just set it to none. And setting it to
  2802. 1:38:58none is more efficient. And PyTorch will
  2803. 1:39:00interpret none as like a lack of a
  2804. 1:39:03gradient and is the same as zeros.
  2805. 1:39:05So, this is a way to set to zero the
  2806. 1:39:07gradient.
  2807. 1:39:10And now, we do loss.backward.
  2808. 1:39:14Before we do loss.backward, we need one
  2809. 1:39:16more thing. If you remember from
  2810. 1:39:17micrograd,
  2811. 1:39:19PyTorch actually requires
  2812. 1:39:21that we pass in requires_grad is true.
  2813. 1:39:25Uh so that we tell
  2814. 1:39:27PyTorch that we are interested in
  2815. 1:39:28calculating gradients for this leaf
  2816. 1:39:30tensor. By default, this is false.
  2817. 1:39:33So, let me recalculate with that.
  2818. 1:39:36And then, set to none and loss.backward.
  2819. 1:39:40Now, something magical happened when
  2820. 1:39:42loss of backward was run.
  2821. 1:39:44Because PyTorch, just like micrograd,
  2822. 1:39:47when we did the forward pass here,
  2823. 1:39:49it keeps track of all the operations
  2824. 1:39:51under the hood. It builds a full
  2825. 1:39:53computational graph.
  2826. 1:39:54Just like the graphs we could produce in
  2827. 1:39:57micrograd, those graphs exist inside
  2828. 1:39:59PyTorch.
  2829. 1:40:00And so, it knows all the dependencies
  2830. 1:40:02and all the mathematical operations of
  2831. 1:40:04everything.
  2832. 1:40:05And when you then calculate the loss, we
  2833. 1:40:07can call a dot backward on it.
  2834. 1:40:09And dot backward then fills in the
  2835. 1:40:11gradients of all the intermediates all
  2836. 1:40:15the way back to W's, which are the
  2837. 1:40:18parameters of our neural net. So, now we
  2838. 1:40:20can do W.grad,
  2839. 1:40:22and we see that it has structure.
  2840. 1:40:23There's stuff inside it.
  2841. 1:40:29And these gradients, every single
  2842. 1:40:31element here,
  2843. 1:40:33so W.shape is 27 by 27.
  2844. 1:40:36W.grad's shape is the same, 27 by 27.
  2845. 1:40:40And every element of W.grad
  2846. 1:40:43is telling us
  2847. 1:40:44the influence of that weight on the loss
  2848. 1:40:47function.
  2849. 1:40:48So, for example, this number all the way
  2850. 1:40:50here,
  2851. 1:40:52if this element, the 0 0 element of W,
  2852. 1:40:55because the gradient is positive, it's
  2853. 1:40:57telling us that this has a positive
  2854. 1:41:00influence on the loss, slightly nudging
  2855. 1:41:03W
  2856. 1:41:04slightly taking W00
  2857. 1:41:06and adding a small H to it
  2858. 1:41:10would increase the loss mildly because
  2859. 1:41:13this gradient is positive.
  2860. 1:41:15Some of these gradients are also
  2861. 1:41:16negative.
  2862. 1:41:18So, that's telling us about the gradient
  2863. 1:41:20information, and we can use this
  2864. 1:41:22gradient information to update the
  2865. 1:41:24weights of this neural network. So,
  2866. 1:41:26let's now do the update. It's going to
  2867. 1:41:28be very similar to what we had in
  2868. 1:41:29micrograd. We need no uh loop over all
  2869. 1:41:32the parameters because we only have one
  2870. 1:41:34parameter uh tensor, and that is W. So,
  2871. 1:41:37we simply do W.data plus equals uh, the
  2872. 1:41:42We can actually copy this almost
  2873. 1:41:43exactly. -0.1 *
  2874. 1:41:46uh, W.grad.
  2875. 1:41:48Um,
  2876. 1:41:49and that would be the update to the
  2877. 1:41:52tensor.
  2878. 1:41:54So, that updates the tensor.
  2879. 1:41:58And because the tensor is updated, we
  2880. 1:42:01would expect that now the loss should
  2881. 1:42:03decrease.
  2882. 1:42:04So, here, if I print loss
  2883. 1:42:09.item,
  2884. 1:42:11it was 3.76, right? So, we've updated
  2885. 1:42:14the W here. So, if I recalculate forward
  2886. 1:42:17pass,
  2887. 1:42:18loss now should be slightly lower. So,
  2888. 1:42:213.76 goes to
  2889. 1:42:233.74.
  2890. 1:42:25And then, we can again set to set grad
  2891. 1:42:28to none and backward update.
  2892. 1:42:32And now the parameters changed again.
  2893. 1:42:34So, if we recalculate the forward pass,
  2894. 1:42:37we expect a lower loss again, 3.72.
  2895. 1:42:42Okay, and this is again doing the We're
  2896. 1:42:44now doing gradient descent.
  2897. 1:42:48And when we achieve a low loss, that
  2898. 1:42:50will mean that the network is assigning
  2899. 1:42:52high probabilities to the correct next
  2900. 1:42:54characters. Okay, so I rearranged
  2901. 1:42:56everything and I put it all together
  2902. 1:42:58from scratch.
  2903. 1:42:59So, here is where we construct our data
  2904. 1:43:01set of bigrams.
  2905. 1:43:03You see that we are still iterating only
  2906. 1:43:04on the first word, Emma.
  2907. 1:43:06I'm going to change that in a second. I
  2908. 1:43:09added a number that counts the number of
  2909. 1:43:11elements in X's, so that we explicitly
  2910. 1:43:14see the number of examples is five.
  2911. 1:43:16Because currently we're just working
  2912. 1:43:18with Emma, and there's five bigrams
  2913. 1:43:19there.
  2914. 1:43:20And here I added a loop of exactly what
  2915. 1:43:22we had before. So, we had 10 iterations
  2916. 1:43:25of gradient descent, of forward pass,
  2917. 1:43:27backward pass, and an update. And so,
  2918. 1:43:29running these two cells, initialization
  2919. 1:43:31and gradient descent
  2920. 1:43:32gives us some improvement on uh the loss
  2921. 1:43:36function.
  2922. 1:43:38But now, I want to use all the words.
  2923. 1:43:41And there's not five, but 228,000
  2924. 1:43:44bigrams now.
  2925. 1:43:46However, this should require no
  2926. 1:43:48modification whatsoever. Everything
  2927. 1:43:49should just run because all the code we
  2928. 1:43:51wrote doesn't care if there's five
  2929. 1:43:53bigrams or 228,000 bigrams. And with
  2930. 1:43:56everything, we should just work. So,
  2931. 1:43:58you see that this will just run.
  2932. 1:44:00But now we are optimizing over the
  2933. 1:44:01entire training set of all the bigrams.
  2934. 1:44:04And you see now that we are decreasing
  2935. 1:44:06very slightly. So, actually, we can
  2936. 1:44:08probably afford a larger learning rate.
  2937. 1:44:12We can probably afford even larger
  2938. 1:44:13learning rate.
  2939. 1:44:20Even 50 seems to work on this very, very
  2940. 1:44:22simple example, right? So, let me
  2941. 1:44:24re-initialize, and let's run 100
  2942. 1:44:26iterations.
  2943. 1:44:29See what happens.
  2944. 1:44:33Okay.
  2945. 1:44:36We seem to be
  2946. 1:44:39coming up to some pretty good losses
  2947. 1:44:40here. 2.47.
  2948. 1:44:42Let me run 100 more.
  2949. 1:44:44What is the number that we expect, by
  2950. 1:44:46the way, in the loss? We expect to get
  2951. 1:44:48something around what we had originally,
  2952. 1:44:50actually.
  2953. 1:44:52So, all the way back, if you remember,
  2954. 1:44:53in the beginning of this video, when we
  2955. 1:44:55optimized uh just by counting,
  2956. 1:44:58our loss was roughly 2.47
  2957. 1:45:01after we added smoothing.
  2958. 1:45:03But before smoothing, we had roughly
  2959. 1:45:042.45
  2960. 1:45:06uh likelihood.
  2961. 1:45:08Um sorry, loss.
  2962. 1:45:09And so, that's actually roughly the
  2963. 1:45:11vicinity of what we expect to achieve.
  2964. 1:45:13But before we achieved it by counting,
  2965. 1:45:15and here we are achieving the roughly
  2966. 1:45:17the same result, but with gradient-based
  2967. 1:45:19optimization.
  2968. 1:45:21So, we come to about 2.4
  2969. 1:45:236, 2.45, etc.
  2970. 1:45:26And that makes sense because
  2971. 1:45:27fundamentally, we're not taking any
  2972. 1:45:28additional information. We're still just
  2973. 1:45:30taking in the previous character and
  2974. 1:45:31trying to predict the next one. But
  2975. 1:45:33instead of doing it explicitly by
  2976. 1:45:35counting and normalizing,
  2977. 1:45:38we are doing it with gradient-based
  2978. 1:45:39learning. And it just so happens that
  2979. 1:45:41the explicit approach happens to very
  2980. 1:45:43well optimize the loss function without
  2981. 1:45:46any need for gradient-based optimization
  2982. 1:45:48because the setup for bigram language
  2983. 1:45:50models are is is so straightforward and
  2984. 1:45:52so simple. We can just afford to
  2985. 1:45:54estimate those probabilities directly
  2986. 1:45:55and maintain them in a table.
  2987. 1:45:58But the gradient-based approach is
  2988. 1:46:00significantly more flexible.
  2989. 1:46:03So we've actually gained a lot because
  2990. 1:46:06what we can do now is um
  2991. 1:46:09we can expand this approach and
  2992. 1:46:11complexify the neural net.
  2993. 1:46:12So currently we're just taking a single
  2994. 1:46:14character and feeding into a neural net
  2995. 1:46:15and the neural net is extremely simple.
  2996. 1:46:17But we're about to iterate on this
  2997. 1:46:19substantially. We're going to be taking
  2998. 1:46:21multiple previous characters and we're
  2999. 1:46:23going to be feed them in feeding them
  3000. 1:46:25into increasingly more complex neural
  3001. 1:46:26nets. But fundamentally uh the output of
  3002. 1:46:29the neural net will always just be
  3003. 1:46:30logits.
  3004. 1:46:32And those logits will go through the
  3005. 1:46:34exact same transformation. We are going
  3006. 1:46:36to take them through a softmax,
  3007. 1:46:38calculate the loss function and the
  3008. 1:46:39negative log likelihood,
  3009. 1:46:41and do gradient-based optimization.
  3010. 1:46:43And so actually as we complexify the
  3011. 1:46:46neural nets and work all the way up to
  3012. 1:46:48transformers,
  3013. 1:46:49none of this will really fundamentally
  3014. 1:46:51change. None of this will fundamentally
  3015. 1:46:52change. The only thing that will change
  3016. 1:46:54is
  3017. 1:46:55the way we do the forward pass where we
  3018. 1:46:57take in some previous characters and
  3019. 1:46:59calculate the logits for the next
  3020. 1:47:01character in a sequence. That will
  3021. 1:47:03become more complex
  3022. 1:47:05and uh but we'll use the same machinery
  3023. 1:47:07to optimize it.
  3024. 1:47:08And um
  3025. 1:47:10it's not obvious how we would have
  3026. 1:47:12extended this bigram approach into the
  3027. 1:47:15case where there are many more
  3028. 1:47:17characters at the input because
  3029. 1:47:19eventually these tables would get way
  3030. 1:47:21too large because there's way too many
  3031. 1:47:23combinations of what previous characters
  3032. 1:47:26uh could be.
  3033. 1:47:27If you only have one previous character,
  3034. 1:47:29we can just keep everything in a table,
  3035. 1:47:31the counts. But if you have the last 10
  3036. 1:47:33characters that are input, we can't
  3037. 1:47:35actually keep everything in a table
  3038. 1:47:36anymore. So this is fundamentally an
  3039. 1:47:38unscalable approach, and the neural
  3040. 1:47:40network approach is significantly more
  3041. 1:47:42scalable, and it's something that
  3042. 1:47:44actually we can improve on over time. So
  3043. 1:47:46that's where we will be digging next. I
  3044. 1:47:48wanted to point out two more things.
  3045. 1:47:51Number one,
  3046. 1:47:52I want you to notice that this
  3047. 1:47:55X inc here,
  3048. 1:47:56this is made up of one-hot vectors, and
  3049. 1:47:59then those one-hot vectors are
  3050. 1:48:00multiplied by this W matrix.
  3051. 1:48:03And we think of this as uh multiple
  3052. 1:48:05neurons being forwarded in a fully
  3053. 1:48:07connected manner.
  3054. 1:48:08But actually what's happening here is
  3055. 1:48:10that, for example,
  3056. 1:48:12if you have a one-hot vector here that
  3057. 1:48:14has a one at, say, the fifth dimension,
  3058. 1:48:17then because of the way the matrix
  3059. 1:48:19multiplication works,
  3060. 1:48:21multiplying that one-hot vector with W
  3061. 1:48:23actually ends up plucking out the fifth
  3062. 1:48:25row of W.
  3063. 1:48:27Logits would become just a fifth row of
  3064. 1:48:30W.
  3065. 1:48:31And that's because of the way the matrix
  3066. 1:48:33multiplication works.
  3067. 1:48:35Um
  3068. 1:48:36so that's actually what ends up
  3069. 1:48:38happening.
  3070. 1:48:40So but that's actually exactly what
  3071. 1:48:42happened before.
  3072. 1:48:43Because remember all the way up here,
  3073. 1:48:46we have a bigram. We took the first
  3074. 1:48:48character, and then that first character
  3075. 1:48:50indexed into a row of this array here.
  3076. 1:48:55And that row gave us the probability
  3077. 1:48:56distribution for the next character.
  3078. 1:48:58So the first character was used as a
  3079. 1:49:00lookup into a
  3080. 1:49:02uh
  3081. 1:49:03matrix here to get the probability
  3082. 1:49:05distribution.
  3083. 1:49:06Well, that's actually exactly what's
  3084. 1:49:07happening here.
  3085. 1:49:08Because we're taking the index, we're
  3086. 1:49:10encoding it as one-hot, and multiplying
  3087. 1:49:12it by W.
  3088. 1:49:13So logits literally becomes the
  3089. 1:49:16uh the
  3090. 1:49:18the appropriate row of W.
  3091. 1:49:20And that gets just as before
  3092. 1:49:22exponentiated to create the counts
  3093. 1:49:25and then normalized and becomes
  3094. 1:49:26probability.
  3095. 1:49:27So this W here
  3096. 1:49:29is literally
  3097. 1:49:31the same as this array here.
  3098. 1:49:35But W, remember, is the log counts, not
  3099. 1:49:38the counts. So it's more precise to say
  3100. 1:49:40that W exponentiated
  3101. 1:49:42W.exp is this array.
  3102. 1:49:46But this array was filled in by counting
  3103. 1:49:49and by basically
  3104. 1:49:52populating the counts of bigrams,
  3105. 1:49:53whereas in the gradient-based framework,
  3106. 1:49:55we initialize it randomly and then we
  3107. 1:49:57let the loss guide us
  3108. 1:50:00to arrive at the exact same array.
  3109. 1:50:03So this array exactly here
  3110. 1:50:05is
  3111. 1:50:06basically the array W at the end of
  3112. 1:50:09optimization, except we arrived at it
  3113. 1:50:12piece by piece by following the loss.
  3114. 1:50:15And that's why we also obtain the same
  3115. 1:50:16loss function at the end. And the second
  3116. 1:50:18note is if I come here,
  3117. 1:50:20remember the smoothing where we added
  3118. 1:50:22fake counts to our counts in order to
  3119. 1:50:26smooth out and make more uniform the
  3120. 1:50:28distributions of these probabilities.
  3121. 1:50:31And that prevented us from assigning
  3122. 1:50:32zero probability to
  3123. 1:50:34um
  3124. 1:50:35to any one bigram.
  3125. 1:50:37Now, if I increase the count here,
  3126. 1:50:40what's happening to the probability?
  3127. 1:50:42As I increase the count, probability
  3128. 1:50:45becomes more and more uniform.
  3129. 1:50:48Right? Because these counts go only up
  3130. 1:50:50to like 900 or whatever. So if I'm
  3131. 1:50:51adding plus a million to every single
  3132. 1:50:54number here, you can see how uh the row
  3133. 1:50:57and its probability then when we divide
  3134. 1:50:59is just going to become more and more
  3135. 1:51:00close to exactly even probability,
  3136. 1:51:03uniform distribution.
  3137. 1:51:05It turns out that the gradient-based
  3138. 1:51:06framework has an equivalent to
  3139. 1:51:09smoothing.
  3140. 1:51:10In particular,
  3141. 1:51:13think through these W's here,
  3142. 1:51:15which we initialized randomly.
  3143. 1:51:18We could also think about initializing
  3144. 1:51:20W's to be zero.
  3145. 1:51:22If all the entries of W are zero,
  3146. 1:51:26then you'll see that logits will become
  3147. 1:51:27all zero.
  3148. 1:51:28And then exponentiating those logits
  3149. 1:51:30becomes all one.
  3150. 1:51:32And then the probabilities turn out to
  3151. 1:51:33be exactly uniform.
  3152. 1:51:35So, basically when W's are all equal to
  3153. 1:51:38each other, or say especially zero,
  3154. 1:51:41then the probabilities come out
  3155. 1:51:42completely uniform.
  3156. 1:51:44So,
  3157. 1:51:45trying to incentivize W to be near zero
  3158. 1:51:49is basically equivalent to label
  3159. 1:51:52smoothing. And the more you incentivize
  3160. 1:51:54that in the loss function, the more
  3161. 1:51:56smooth distribution you're going to
  3162. 1:51:57achieve.
  3163. 1:51:58So, this brings us to something that's
  3164. 1:52:00called regularization,
  3165. 1:52:02where we can actually augment the loss
  3166. 1:52:03function to have a small component that
  3167. 1:52:06we call a regularization loss.
  3168. 1:52:09In particular, what we're going to do is
  3169. 1:52:10we can take W,
  3170. 1:52:11and we can for example square all of its
  3171. 1:52:13entries,
  3172. 1:52:14and then we can uh oops.
  3173. 1:52:17Sorry about that.
  3174. 1:52:19We can take all the entries of W, and we
  3175. 1:52:20can sum them.
  3176. 1:52:23And because we're squaring, uh there
  3177. 1:52:25will be no signs anymore. Um
  3178. 1:52:28negatives and positives all get squashed
  3179. 1:52:30to be positive numbers.
  3180. 1:52:31And then the way this works is you
  3181. 1:52:33achieve zero loss if W is exactly or
  3182. 1:52:36zero.
  3183. 1:52:37But if W has non-zero numbers, you
  3184. 1:52:39accumulate loss.
  3185. 1:52:41And so, we can actually take this, and
  3186. 1:52:42we can add it on here.
  3187. 1:52:44So, we can do something like loss plus
  3188. 1:52:48W squared
  3189. 1:52:50dot sum.
  3190. 1:52:52Or let's actually, instead of sum, let's
  3191. 1:52:53take a mean, cuz otherwise the sum gets
  3192. 1:52:55too large.
  3193. 1:52:57So, mean is like a little bit more
  3194. 1:52:58manageable.
  3195. 1:53:01And then we have a regularization loss
  3196. 1:53:02here, let's say 0.01 times,
  3197. 1:53:05or something like that. You can choose
  3198. 1:53:06the regularization strength.
  3199. 1:53:09And then we can just optimize this.
  3200. 1:53:12And now this optimization actually has
  3201. 1:53:14two components. Not only is it trying to
  3202. 1:53:16make all the probabilities work out, but
  3203. 1:53:18in addition to that, there's an
  3204. 1:53:19additional component that simultaneously
  3205. 1:53:21tries to make all W's be zero. Because
  3206. 1:53:24if W's are non-zero, you feel a loss.
  3207. 1:53:26And so minimizing this, the only way to
  3208. 1:53:28do achieve that is for W to be zero.
  3209. 1:53:30And so you can think of this as adding
  3210. 1:53:32like a spring force or like a gravity
  3211. 1:53:34force that that pushes W to be zero.
  3212. 1:53:37So W wants to be zero, and the
  3213. 1:53:39probabilities want to be uniform, but
  3214. 1:53:41they also simultaneously want to match
  3215. 1:53:43up your your probabilities as indicated
  3216. 1:53:46by the data.
  3217. 1:53:47And so the strength of this
  3218. 1:53:49regularization is exactly controlling
  3219. 1:53:52the amount of counts
  3220. 1:53:54that you add here.
  3221. 1:53:57Adding a lot more counts here
  3222. 1:54:00corresponds to
  3223. 1:54:02increasing this number.
  3224. 1:54:04Because the more you increase it, the
  3225. 1:54:06more this part of the loss function
  3226. 1:54:08dominates this part, and the more these
  3227. 1:54:10these weights will be unable to grow
  3228. 1:54:13because as they grow,
  3229. 1:54:15they accumulate way too much loss.
  3230. 1:54:18And so if this is strong enough,
  3231. 1:54:21then we are not able to overcome the
  3232. 1:54:23force of this loss, and we will never
  3233. 1:54:26and basically everything will be uniform
  3234. 1:54:28predictions.
  3235. 1:54:29So I thought that's kind of cool.
  3236. 1:54:30Okay, and lastly, before we wrap up,
  3237. 1:54:33I wanted to show you how you would
  3238. 1:54:34sample from this neural net model.
  3239. 1:54:36And I copy-pasted the sampling code from
  3240. 1:54:39before.
  3241. 1:54:40Where remember that we sampled five
  3242. 1:54:43times.
  3243. 1:54:44And all we did is we started zero, we
  3244. 1:54:46grabbed the current IX row of P.
  3245. 1:54:50And that was our probability row
  3246. 1:54:52from which we sampled the next index and
  3247. 1:54:55just accumulated that and break when
  3248. 1:54:57zero.
  3249. 1:54:58And running this gave us these results.
  3250. 1:55:03I still have the
  3251. 1:55:05P in memory, so this is fine.
  3252. 1:55:07Now,
  3253. 1:55:09this P doesn't come from the row of P.
  3254. 1:55:12Instead, it comes from this neural net.
  3255. 1:55:14First, we take IX
  3256. 1:55:17and we encode it into a one-hot row of X
  3257. 1:55:21inc.
  3258. 1:55:22This X inc multiplies our W,
  3259. 1:55:25which really just plucks out the row of
  3260. 1:55:26W corresponding to IX. Really, that's
  3261. 1:55:29what's happening.
  3262. 1:55:30And that gets our logits, and then we
  3263. 1:55:33normalize those logits, exponentiate to
  3264. 1:55:35get counts, and then normalize to get uh
  3265. 1:55:37the distribution, and then we can sample
  3266. 1:55:39from the distribution.
  3267. 1:55:41So, if I run this,
  3268. 1:55:45kind of anticlimactic or climatic,
  3269. 1:55:47depending how you look at it, but we get
  3270. 1:55:48the exact same result.
  3271. 1:55:50Um
  3272. 1:55:52and that's because this is in the
  3273. 1:55:53identical model. Not only does it
  3274. 1:55:55achieve the same loss, but um as I
  3275. 1:55:58mentioned, these are identical models,
  3276. 1:55:59and this W is the log counts of what
  3277. 1:56:02we've estimated before. But, we came to
  3278. 1:56:05this answer in a very different way and
  3279. 1:56:07it's got a very different
  3280. 1:56:08interpretation. But, fundamentally, this
  3281. 1:56:10is basically the same model and gives
  3282. 1:56:11the same samples here. And so,
  3283. 1:56:14that's kind of cool. Okay, so we've
  3284. 1:56:16actually covered a lot of ground. We
  3285. 1:56:18introduced the bigram character-level
  3286. 1:56:20language model.
  3287. 1:56:22We saw how we can train the model, how
  3288. 1:56:24we can sample from the model, and how we
  3289. 1:56:25can evaluate the quality of the model
  3290. 1:56:27using the negative log likelihood loss.
  3291. 1:56:30And then we actually trained the model
  3292. 1:56:31in two completely different ways that
  3293. 1:56:33actually give the same result and the
  3294. 1:56:35same model.
  3295. 1:56:36In the first way, we just counted up the
  3296. 1:56:38frequency of all the bigrams and
  3297. 1:56:40normalized.
  3298. 1:56:41In the second way, we used the uh
  3299. 1:56:44negative log likelihood loss as a guide
  3300. 1:56:47to optimizing the counts matrix
  3301. 1:56:50uh or the counts array so that the loss
  3302. 1:56:52is minimized in the in a gradient-based
  3303. 1:56:55framework. And we saw that both of them
  3304. 1:56:56give the same result.
  3305. 1:56:58And um
  3306. 1:57:00that's it.
  3307. 1:57:01Now, the second one of these, the
  3308. 1:57:02gradient base framework, is much more
  3309. 1:57:03flexible. And right now, our neural
  3310. 1:57:06network is super simple. We're taking a
  3311. 1:57:08single previous character, and we're
  3312. 1:57:10taking it through a single linear layer
  3313. 1:57:12to calculate the logits.
  3314. 1:57:14This is about to complexify. So, in the
  3315. 1:57:16follow-up videos, we're going to be
  3316. 1:57:17taking more and more of these
  3317. 1:57:19characters,
  3318. 1:57:20and we're going to be feeding them into
  3319. 1:57:21a neural net.
  3320. 1:57:23But, this neural net will still output
  3321. 1:57:24the exact same thing. The neural net
  3322. 1:57:25will output logits.
  3323. 1:57:28And these logits will still be
  3324. 1:57:29normalized in the exact same way, and
  3325. 1:57:30all the loss and everything else in the
  3326. 1:57:32gradient gradient base framework,
  3327. 1:57:33everything stays identical.
  3328. 1:57:35It's just that this neural net will now
  3329. 1:57:37complexify all the way to transformers.
  3330. 1:57:40So, that's going to be pretty awesome,
  3331. 1:57:42and I'm looking forward to it. For now,
  3332. 1:57:44bye.

About this transcript

This page contains the full transcript of The spelled-out intro to language modeling: building makemore by Andrej Karpathy, generated from the public captions YouTube serves with the video. The transcript has 19,445 words across 3,332 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.