YouTube2Text

YouTube transcript (pwQ0l4hFCVI) — Transcript

12,541 words · 2,159 segments · language en · Watch on YouTube

Full transcript

  1. 0:05Today we are going to move to um
  2. 0:07a sequence of um lectures on large
  3. 0:10language models. Not many, like probably
  4. 0:12three, four. And in between we also have
  5. 0:15to talk about reinforcement learning a
  6. 0:16little bit because otherwise we cannot
  7. 0:17talk about some of the modern reasoning
  8. 0:20models without understanding a little
  9. 0:22bit about reinforcement learning.
  10. 0:24So that's pretty much the rest of the
  11. 0:25lecture except uh
  12. 0:26two guest lectures. I think next
  13. 0:29um the upcoming Monday we'll have Simran
  14. 0:32to give us a lecture on system ML which
  15. 0:34is also quite about um
  16. 0:36transformers. Like how do you optimize
  17. 0:38the efficiency of transformers on a GPU
  18. 0:40so on and so forth. Um
  19. 0:42Okay, sounds good. So large language
  20. 0:45models, I guess you always probably all
  21. 0:46of you have heard of this word
  22. 0:49uh large language models. This is um um
  23. 0:52precisely I think in some sense this is
  24. 0:54all starting from GPT-3 when OpenAI
  25. 0:57released GPT-3. Um it has a um it used
  26. 1:00this new paradigm which uh used
  27. 1:02autoregressive
  28. 1:03large language models. Um so um so
  29. 1:06basically we're going to focus on that,
  30. 1:07you know, autoregressive large language
  31. 1:09models. We're going to skip some of the
  32. 1:11earlier versions of transformer which
  33. 1:12are um not autoregressive. Um so
  34. 1:17and um and an autoregressive basically
  35. 1:20means you generate tokens one by one and
  36. 1:22each of the new tokens depends on the
  37. 1:23previous tokens and you do this um
  38. 1:25um and you model probability in this
  39. 1:27kind of like a autoregressive way. I'm
  40. 1:29going to talk more about details. In
  41. 1:30some sense you model your language
  42. 1:32distribution by using a
  43. 1:34um um a sequentially sequential model
  44. 1:37one by one in some sense.
  45. 1:39Um anyway, so um let's um I think the
  46. 1:41point
  47. 1:43Our goal is to talk about some details.
  48. 1:44So I'll fill in more high-level stuff as
  49. 1:46we talk about the details.
  50. 1:48So the first thing we're going to talk
  51. 1:49about is tokenization.
  52. 1:53This is um you know, just uh
  53. 1:56to prepare us for to talk about the real
  54. 1:58meat in some sense.
  55. 1:59So, what does tokenization mean? It
  56. 2:01means that you need to
  57. 2:02you know, these transformers, they are
  58. 2:04numerical you know, architecture models,
  59. 2:06right? So, they take in
  60. 2:08uh in some sense numerical numbers,
  61. 2:10right? So, um but uh you have text. And
  62. 2:12how do you turn text into some kind of
  63. 2:14things that
  64. 2:15the models can understand? So, uh the
  65. 2:18first step is to tokenize. Tokenize
  66. 2:20means that you need to decide you what
  67. 2:22is the smallest unit uh of the input and
  68. 2:25you give to the large language model.
  69. 2:27And often it's not a
  70. 2:28a word, right? And the most natural
  71. 2:30thing you can think of is maybe the
  72. 2:32smallest unit is a character or a word.
  73. 2:34If you use character level tokenization,
  74. 2:36then the issue is that often you have so
  75. 2:38many characters, right? You you need so
  76. 2:40many tokens to so many steps.
  77. 2:42Um and um
  78. 2:44uh if you use word level uh
  79. 2:45tokenization, then often the issue is
  80. 2:47that you don't How do you deal with this
  81. 2:49long word, which are very rare? For
  82. 2:50example, you have a word called say
  83. 2:52internationalization,
  84. 2:55right? So, you have If you treat this
  85. 2:57word as a single word, then maybe it's
  86. 2:58too rare and you don't leverage your
  87. 3:01understanding about other words, right?
  88. 3:03So, if you treat internationalization as
  89. 3:04one word, then
  90. 3:06it is different from internationalize.
  91. 3:08It It's different from like a nation,
  92. 3:10right? And And they are all completely
  93. 3:12different kind of things and you cannot
  94. 3:14leverage
  95. 3:16um that you For example, if you
  96. 3:17understand internationalize, then you
  97. 3:19should probably understand
  98. 3:20internationalization to some degree
  99. 3:22without even seeing more data. And uh
  100. 3:24so, um so, that's why people don't
  101. 3:26always use um subword tokenization uh
  102. 3:29don't use word tokenization. Maybe
  103. 3:30another Okay, maybe let me give you
  104. 3:32another example, right? For example, if
  105. 3:33you have use this
  106. 3:34I guess in my notes there is this word
  107. 3:36LM
  108. 3:38ification.
  109. 3:39I actually don't know whether this is a
  110. 3:40real word, you know, like I think it's a
  111. 3:41new word that is completely new word.
  112. 3:43Nobody have used it before 2020
  113. 3:46uh 2020, I guess, right? So, uh even
  114. 3:48even right now it's not that like of
  115. 3:50like common.
  116. 3:52And if you have to treat this word as a
  117. 3:55unit, then you you have to see this word
  118. 3:58enough times to understand what it
  119. 3:59means.
  120. 4:00Right? However, if you treat it as sub
  121. 4:02words, then maybe you have a a token
  122. 4:04called LM. You have a token which is
  123. 4:07you have a two units. You know, this is
  124. 4:08a token, this is a token. So, LM is a
  125. 4:10token, ification is a token. Then in
  126. 4:13some sense, you have seen LMs many many
  127. 4:14times, you already understand it. And
  128. 4:16then
  129. 4:18ification has a special meaning in
  130. 4:20English, and then you combine these two
  131. 4:22understanding, you understand the word.
  132. 4:24So, that's much more efficient
  133. 4:26understanding of the word than just the
  134. 4:29treating this whole thing as a single
  135. 4:30unit.
  136. 4:32So,
  137. 4:33and then and this happens more often in,
  138. 4:35for example, biology, where you have
  139. 4:37this very long word. Or maybe in German,
  140. 4:39you have very long word, but actually
  141. 4:40they are fundamentally many combinations
  142. 4:42of many sub words. And uh and uh and uh
  143. 4:45using sub word tokenization allows you
  144. 4:47to understand each of the sub word, and
  145. 4:49and then you can understand the
  146. 4:50combination of the those sub words
  147. 4:52as a as a as a whole
  148. 4:54very efficiently. So, yeah, so basically
  149. 4:57tokenization, you know, just means that
  150. 4:58you have a predefined vocabulary.
  151. 5:02Vocabulary.
  152. 5:04And this vocabulary is a list of like a
  153. 5:05sub words, right? So, for example, you
  154. 5:07can have, you know, many many of this
  155. 5:08like maybe A is a sub word, you know, is
  156. 5:11the one of the token, you know,
  157. 5:12I think maybe like a dog is a token.
  158. 5:16Pretty much all common words these days
  159. 5:18is a token.
  160. 5:19And actually, if you look at the the the
  161. 5:21actually, you can play with the
  162. 5:22tokenization online if you search OpenAI
  163. 5:24Tokenizer. OpenAI give you a playground
  164. 5:26where you can see how the texts are
  165. 5:28tokenized. And and you'll you'll find
  166. 5:31that actually blank dog is also token.
  167. 5:35So,
  168. 5:36uh so, the blank is actually merged with
  169. 5:37dog as a token. And uh you know, and you
  170. 5:40have so many of this. I typically this
  171. 5:41kind of like a tokenization right now,
  172. 5:44um I think for the open source models,
  173. 5:45we know the number of tokens uh they
  174. 5:47have. I think it's, you know, often on
  175. 5:49the order of like 10 to the five, like
  176. 5:51100k or at least. Um I think QN QN 3.5
  177. 5:55um this is the newest version of
  178. 5:56open-source models, which are pretty
  179. 5:57good. I think they have like 250 tokens.
  180. 6:01Um uh I think 250 probably is mostly
  181. 6:03because, you know, I think 250 is is
  182. 6:05somehow you can represent this token
  183. 6:07with like a
  184. 6:08maybe
  185. 6:09eight eight bytes in some sense, like
  186. 6:11like a
  187. 6:12it also depends on how many bits you can
  188. 6:14present the tokens. Um so uh anyway, so
  189. 6:17you have this vocabulary, and then you
  190. 6:19just uh uh given a long uh string of
  191. 6:21text, you just uh um
  192. 6:23uh take all the subwords that are
  193. 6:24tokens, and you uh turn it into a
  194. 6:26sequence of like a segments, right? For
  195. 6:28example, if you say if I have a a dog
  196. 6:31runs, suppose this is the text you see,
  197. 6:33and then you you'll say this is a token.
  198. 6:36So, this is one part. And then I think
  199. 6:38blank dog is a token.
  200. 6:41And blank runs is also token.
  201. 6:43Uh and uh
  202. 6:45and comma, I think is a token by itself.
  203. 6:48So, and every token has an ID, right?
  204. 6:50So, let's say A is probably the ID
  205. 6:52number one, dog is maybe the ID
  206. 6:54number like 356, whatever, and and run
  207. 6:57is a some ID.
  208. 6:58Maybe I don't know, like a
  209. 7:00maybe 250.
  210. 7:01I'm just writing some random numbers.
  211. 7:03So, every everything has an ID. So,
  212. 7:04basically, you turn every text into a
  213. 7:06sequence of IDs, you know, each ID
  214. 7:08corresponds to one token.
  215. 7:11Um and um
  216. 7:12and I think So, in this case, you know,
  217. 7:14most of these words are actually one
  218. 7:15token. I think uh I just checked it
  219. 7:17recently. If you have the word
  220. 7:19happiness,
  221. 7:21I think the tokenization will be
  222. 7:23unhappiness, let's say.
  223. 7:25I think this part is a token.
  224. 7:27H is a token, and
  225. 7:29a p i n e s s is a token
  226. 7:31uh
  227. 7:32for some reason. Um and maybe another
  228. 7:34example is this Oh, I I already erased
  229. 7:37it. The I L M imbification.
  230. 7:39And this is another examples [laughter]
  231. 7:41so
  232. 7:42I come uh we come up with, you know,
  233. 7:44just a I don't know actually what these
  234. 7:45words means. If you have a blast
  235. 7:47oma, I think this part is a token, this
  236. 7:49part is a token, and this part is token.
  237. 7:51I think this is just because for this
  238. 7:52biological biology words, you know,
  239. 7:54biomedical words, you know, actually
  240. 7:56every part of the word has some special
  241. 7:58meaning and uh uh
  242. 7:59uh so, you tokenize them into um tokens.
  243. 8:02And uh and what is the exact vocabulary?
  244. 8:04It depends on the model.
  245. 8:06Uh um so, basically the the algorithm
  246. 8:09that turns the text into this uh tokens
  247. 8:11is called a tokenizer.
  248. 8:13Um and uh and every model has a
  249. 8:14different tokenizer. Not not everyone
  250. 8:16has a different one. Some Some of them
  251. 8:18shared the same tokenizers, but
  252. 8:20generally it's specific to the tokenizer
  253. 8:22to to the model.
  254. 8:23Uh and different companies have
  255. 8:25different um proprietary tokenization
  256. 8:27approach. Uh you probably heard of that
  257. 8:30uh recently, I think Cloud Code uh
  258. 8:32Cloud, they updated their tokenizer so
  259. 8:34that uh
  260. 8:35uh I think they make the tokenizers more
  261. 8:36granular uh to some degree, so that the
  262. 8:38same amount of text will be tokenized
  263. 8:40into more tokens. So, before it was like
  264. 8:435,000 tokens, now it's 1.5,000 tokens.
  265. 8:45Um you know, I didn't verify myself, but
  266. 8:47I read some news about this. Uh assuming
  267. 8:49they are true, uh that means that you
  268. 8:51you spend more tokens, you know, you you
  269. 8:52you pay more to Cloud Code.
  270. 8:55>> [laughter]
  271. 8:55>> With the same amount of text. Um anyway,
  272. 8:58so um cool.
  273. 9:01And there are some details about how to
  274. 9:03come up with the vocabulary and how to
  275. 9:05kind of like uh uh um run this algorithm
  276. 9:08to tokenize. Um the the keyword is
  277. 9:11called BPE. Um
  278. 9:13BPE
  279. 9:15uh tokenizer, byte um pair um uh
  280. 9:18encoding. So, I'm not going to the
  281. 9:20details because this is just a very
  282. 9:21small quick uh uh setup. Um it's not
  283. 9:24that important. Um but roughly speaking,
  284. 9:26the BPE uh tokenization idea is that you
  285. 9:29try to In some sense, you just try to
  286. 9:30find frequent
  287. 9:32subwords that are most frequent, right?
  288. 9:34And uh and you use them them in a
  289. 9:36vocabulary. But how do you decide which
  290. 9:38one is more frequent and and not? I
  291. 9:39think that there's a sequential
  292. 9:40algorithm. You first go through maybe
  293. 9:42like the the
  294. 9:44the smallest kind of like a
  295. 9:47combination of like you go through all
  296. 9:49the characters, you know, and you you
  297. 9:51try to find, you know, I think it's a
  298. 9:52greedy algorithm. First you try to find
  299. 9:54the most frequent, you know, one
  300. 9:55character thing and you find in the most
  301. 9:57frequent combination of two characters
  302. 9:58and you remove you know, it's it's a
  303. 10:00little bit complicated which I let me
  304. 10:02not try to give you the precise version,
  305. 10:04but it's kind of like a a greedy
  306. 10:06algorithm that decides what one which
  307. 10:08one is more frequent
  308. 10:10and so and so forth. And if you are
  309. 10:12interested, you can look up the exact
  310. 10:13algorithm.
  311. 10:15Um Okay, so that's the tokenization. So
  312. 10:18that means that from now on, you know,
  313. 10:20we just assume we have a tokenizer. So
  314. 10:22every text, every piece of text will be
  315. 10:24turned into a sequence of tokens. So
  316. 10:26from now on like a
  317. 10:29a text it just means a sequence of
  318. 10:30tokens and each of this xi is in the
  319. 10:33vocabulary. Let's call the vocabulary v
  320. 10:35v.
  321. 10:37And the size of the vocabulary could be
  322. 10:38like 250
  323. 10:40thousands.
  324. 10:42Um and you know, each of these is kind
  325. 10:44of like index in the vocabulary like you
  326. 10:46know, this is the fifth token, this is
  327. 10:48the
  328. 10:49105th token and so and so forth.
  329. 10:51Okay. So and now the question is, you
  330. 10:53know, now we go into the next question
  331. 10:56which is how do we model this
  332. 10:57distribution of sequence of tokens? How
  333. 11:00do we create a probabilistic model to
  334. 11:02model it? If you think about the support
  335. 11:05of
  336. 11:06this this distribution, like how many
  337. 11:08possible combinations
  338. 11:10that they are possible here? So the
  339. 11:12number of possible combinations is
  340. 11:14you have v choice.
  341. 11:17This is the cardinality of v. You have v
  342. 11:18choice, cardinality of v choice for each
  343. 11:20of the
  344. 11:21the
  345. 11:23tokens and you have a sequence of length
  346. 11:25t. So you take the power to the t,
  347. 11:26right? So everyone has
  348. 11:29v choice and the v to the power t is the
  349. 11:31total number of choices. So, you cannot
  350. 11:34assign one probability for each of this
  351. 11:36sequence
  352. 11:39like like a like a like a like a
  353. 11:41discrete categorical tokens.
  354. 11:43Categorical distributions, right? If you
  355. 11:44treat this as a categorical
  356. 11:45distributions, then every choice here
  357. 11:48will have one probability, and that's
  358. 11:50just way too many parameters to describe
  359. 11:52it. So, that's why
  360. 11:54what people do is that people decompose
  361. 11:56this distribution using the chain rule
  362. 11:58to write it this way. So, you say
  363. 12:00suppose I have a probability
  364. 12:01distribution
  365. 12:03over the joint of the sequence, then I
  366. 12:06can write it as the following. So, I can
  367. 12:08say it's equals to
  368. 12:09P of X1 * P of
  369. 12:12X2 given X1
  370. 12:14P of X3
  371. 12:15given X2 and X1
  372. 12:18so and so forth and P of XT given X1 up
  373. 12:21to XT - 1.
  374. 12:23And this is the the so-called, you know,
  375. 12:25conditional
  376. 12:27probability decomposition. You decompose
  377. 12:29the joint distribution into a product of
  378. 12:32conditional probabilities.
  379. 12:34And you model each of these conditional
  380. 12:36probabilities.
  381. 12:37The benefit here is that each of these
  382. 12:38conditional probabilities
  383. 12:40is a distribution is a conditional
  384. 12:42distribution, but the choice of X2, for
  385. 12:45example, here is only V. You have V
  386. 12:48choices here. You only have a You only
  387. 12:50have to have a kind of like a softmax or
  388. 12:52logits over V choices for each of this.
  389. 12:55Right? So, each of these will be
  390. 12:56eventually you'll see is a some kind of
  391. 12:58like a
  392. 12:59um
  393. 13:00um you have some logits and you apply
  394. 13:01some softmax, you get a probability
  395. 13:02distribution over
  396. 13:04V choice.
  397. 13:05So, that's what we're going to do.
  398. 13:08So, basically each of this um uh will be
  399. 13:10parameterized by a new network. Um the
  400. 13:13input will be uh what's after the
  401. 13:16conditioning. Right? So, input here is
  402. 13:17the X1, X2, and output is the
  403. 13:19distribution of X3. And the And this
  404. 13:21input-output relationship is modeled by
  405. 13:24um a
  406. 13:26uh uh network, which is the transformer.
  407. 13:29Okay, so that's the um
  408. 13:31what we are
  409. 13:32going towards. How do we model each of
  410. 13:34these conditional distributions?
  411. 13:37So, um
  412. 13:40So, basically we're going to describe a
  413. 13:41conditional distribution P of P sub
  414. 13:43theta Xt given X1
  415. 13:46up to Xt minus one.
  416. 13:48So, um
  417. 13:52Before going to details, let's treat
  418. 13:54this as a kind of like a black box
  419. 13:55first. And let's first kind of like
  420. 13:58discuss a little bit about, you know,
  421. 13:59the
  422. 14:00uh the boundary, right? In some sense,
  423. 14:02right? So, like it's kind of like a
  424. 14:03there are multiple layers here. So, the
  425. 14:05typical way to um um
  426. 14:07parameterize this is the following. So,
  427. 14:10uh use a transformer. So, basically what
  428. 14:12happens is that you have
  429. 14:14a sequence X1 up to Xt.
  430. 14:16And um
  431. 14:18um So, first of all, these are discrete
  432. 14:20IDs, right? So, each of these means a
  433. 14:22um ID for the token. You first have to
  434. 14:25turn them into some numerical numbers.
  435. 14:28So, the the the first step is that you
  436. 14:30say, I'm going to turn this into
  437. 14:32an embedding E of X1.
  438. 14:35And so, basically I turn each of the
  439. 14:37token X into
  440. 14:39uh Xt into embedding of Xt.
  441. 14:43And what is the embedding? Embedding is
  442. 14:44a
  443. 14:45let's say D-dimensional vector uh in um
  444. 14:49the Euclidean space.
  445. 14:50And this embedding is is trained. So,
  446. 14:52basically you don't know what the
  447. 14:53parameters are. Uh you will just uh
  448. 14:55parameterize it, you know, as unknown
  449. 14:57numbers, and you train the model to uh
  450. 15:00find the right numbers um
  451. 15:02uh for uh for these embeddings. So,
  452. 15:04every
  453. 15:05every um
  454. 15:07token in the vocabulary will be will
  455. 15:09have a corresponding embeddings. So,
  456. 15:11basically, um you're going to have a um
  457. 15:15a sequence of embeddings E1, E2. This is
  458. 15:17the embedding for the first token,
  459. 15:19embedding for the second token, up to
  460. 15:20the embedding for the last token.
  461. 15:22Um and in this lecture, I think we're
  462. 15:24going to use row vectors because I think
  463. 15:27um we are getting to the
  464. 15:28implementations, you know, every vector
  465. 15:30is actually row vectors in the uh in the
  466. 15:32Python code, even though in mathematics
  467. 15:34people often use column vectors. So, um
  468. 15:37so I I thought about this and I I I
  469. 15:38thought just the for this for large
  470. 15:40language model sections, we'll just
  471. 15:41always use row vectors, which is more
  472. 15:43consistent with the practice. It looks a
  473. 15:45bit weird to me, actually, uh from a
  474. 15:48mathematical perspective. And you if you
  475. 15:49read papers, many of the papers do use
  476. 15:51column vectors, but I think we'll just
  477. 15:53use row vectors to be consistent with
  478. 15:54the code. So, basically, every embedding
  479. 15:56is a row vector.
  480. 15:58And then you have this uh
  481. 16:00embedding matrix. Um so, this is um you
  482. 16:03have a
  483. 16:06V of them, and each of them is of
  484. 16:08dimension D. So, you have this matrix.
  485. 16:10And uh and to read off which row you
  486. 16:12are, you just have to multiply either
  487. 16:14you just read off the ID, right? So, if
  488. 16:15you have the fifth token, you just read
  489. 16:17the fifth row. Or you can multiply this
  490. 16:19matrix with
  491. 16:20um
  492. 16:21uh uh indicator vector to get that row,
  493. 16:23which is the same thing. Anyway, so
  494. 16:25we'll just read off from this embedding
  495. 16:27matrix the corresponding embeddings for
  496. 16:29each of the token. So, you get E of X2
  497. 16:31up to E of XT.
  498. 16:35Um and uh
  499. 16:36and I think in in many cases often
  500. 16:39people append um prepend um a uh
  501. 16:43beginning of sentence token, uh which is
  502. 16:45just a fixed token uh in the very
  503. 16:46beginning. This is a kind of like a not
  504. 16:48always necessary. In some cases, you
  505. 16:50just don't do it. In some cases, you do
  506. 16:51it. But let's um probably just introduce
  507. 16:54that. You have a beginning of sentence
  508. 16:55token, which is one of the special token
  509. 16:57in this vocabulary. You probably often
  510. 16:59actually what happens is that your
  511. 17:01vocabulary is bigger than what is
  512. 17:03needed. You always reserve a few tokens,
  513. 17:05you know, actually maybe a thousand, you
  514. 17:06know, a few thousand tokens for special
  515. 17:08tokens that you can introduce uh in in
  516. 17:12uh for special purposes. And you can
  517. 17:14introduce this beginning of sentence
  518. 17:15token uh as X0, so which is always the
  519. 17:18same.
  520. 17:19Um and then you turn them into
  521. 17:20embeddings.
  522. 17:21And then what you do is that you um go
  523. 17:23through this transformer, which I'll
  524. 17:25discuss in the later half of the
  525. 17:27lecture.
  526. 17:28So, this is a transformer.
  527. 17:32For now, you just think of this as a
  528. 17:33network which does a lot of computation,
  529. 17:35matrix multiplication, you know, uh
  530. 17:37tensions, which I'm going to discuss.
  531. 17:39A lot of operations, and then you output
  532. 17:41um
  533. 17:42uh some logits
  534. 17:44of
  535. 17:45like a you output some kind of logits
  536. 17:47at each of the locations.
  537. 17:51So, each of these logits is supposed to
  538. 17:53describe the conditional distribution um
  539. 17:56before the uh the softmax. So, basically
  540. 18:01um
  541. 18:05what you do is that you say
  542. 18:07um the probability of um
  543. 18:09So, let's just you say each of these U1
  544. 18:12is
  545. 18:13Let's call this U1
  546. 18:15>> [snorts]
  547. 18:15>> U of theta
  548. 18:16X0, and then P of X1
  549. 18:20is
  550. 18:21um
  551. 18:23modeled as, you know,
  552. 18:25just the softmax.
  553. 18:28This distribution P of XI is the
  554. 18:31X1 is modeled as the softmax of
  555. 18:34um
  556. 18:35F theta
  557. 18:36X0.
  558. 18:38So, the first uh
  559. 18:40distribution, which is not conditional,
  560. 18:41you just uh get the output from
  561. 18:44the X0, you get some logic. So, this U
  562. 18:47UI is in dimension
  563. 18:49V,
  564. 18:50and you take a softmax, you get a
  565. 18:51probability vector.
  566. 18:53So, this is a probability vector.
  567. 18:59Right? In this uh simplex of dimension
  568. 19:01V, right? So, the sum this this after
  569. 19:03taking a softmax, the sum of the entries
  570. 19:05is one, uh and all the all of the
  571. 19:07entries are non-active.
  572. 19:09So, then this is your probability for
  573. 19:10X1. This is the distribution.
  574. 19:12Technically, I think the distribution of
  575. 19:13this is equal to softmax.
  576. 19:15Okay? Um and then
  577. 19:17the same thing, you know, when you have
  578. 19:18X2
  579. 19:20given X1 you say this is equals to
  580. 19:22softmax
  581. 19:25of um
  582. 19:27U2
  583. 19:28where U2 is the output at the second
  584. 19:30position.
  585. 19:31And and U2, you know, we'll just call U2
  586. 19:33the same as, you know,
  587. 19:35so you call U2
  588. 19:38the output, you know,
  589. 19:40of um this model
  590. 19:43given X
  591. 19:45um
  592. 19:46zero and X1.
  593. 19:49So
  594. 19:50so each of these UIs, you know, um is uh
  595. 19:54basically the output
  596. 19:56um
  597. 19:58for
  598. 19:59um the sequence before it, right? So UT
  599. 20:01+ 1 is F theta X0 up to XT.
  600. 20:06So um so so basically generically, you
  601. 20:09know, P of XT given X1
  602. 20:12up XT - 1
  603. 20:14is the softmax
  604. 20:17of
  605. 20:19um UT
  606. 20:22uh
  607. 20:23Is it UT or UT + 1? UT +
  608. 20:26UT, yes. UT. Which is the softmax
  609. 20:31of F theta X0 up to XT
  610. 20:36T - 1.
  611. 20:44Okay?
  612. 20:45Any questions so far?
  613. 20:47>> Is your index for your U starting at
  614. 20:49zero or
  615. 20:51>> So I I I particularly chose it to index
  616. 20:52at one, yes.
  617. 20:53>> Okay. And so that's
  618. 20:55this
  619. 20:56on the right.
  620. 20:57>> Yes.
  621. 20:57>> Okay.
  622. 20:58I'm over here.
  623. 21:00You're looking at the same thing.
  624. 21:01>> U is just a vector, you know. I think
  625. 21:03you know, I can just call U I I uh
  626. 21:06it's kind of like uh I can just call you
  627. 21:07f theta x zero it's the same thing. I'm
  628. 21:09not
  629. 21:10Did I answer the question?
  630. 21:27So oh I think uh maybe maybe that's the
  631. 21:30what
  632. 21:31So
  633. 21:32So
  634. 21:32I'm just using f theta x zero x one as a
  635. 21:36more explicit way to represent UI
  636. 21:38because I'm trying to
  637. 21:40express the dependency.
  638. 21:42So
  639. 21:43So basically I'm saying like a basically
  640. 21:45you
  641. 21:46I I think they are the same thing. Like
  642. 21:48it's just there there are no difference
  643. 21:49at all in some sense. It's just how you
  644. 21:50view it, right? Like in one way you just
  645. 21:51say oh I have I need a an
  646. 21:54annotation for the vector and the other
  647. 21:55way is that I have I need to interpret
  648. 21:57this vector as a function of the inputs.
  649. 22:02Oh okay okay okay. So maybe I didn't get
  650. 22:03the question.
  651. 22:44Yes yes yes. So I think maybe
  652. 22:47I'm trying to figure out what's the
  653. 22:48question here. So
  654. 22:49um
  655. 22:50I think maybe what you're asking is the
  656. 22:52following. So
  657. 22:54So maybe let's take this one, right? So
  658. 22:56this is a So this is a vector.
  659. 22:59This is a vector in dimension
  660. 23:01V.
  661. 23:03So in in the vocabulary size. And you
  662. 23:05have to take a softmax, it's a it's a a
  663. 23:07sequence of numbers which are
  664. 23:09uh um
  665. 23:10all non-negative and they sum to one.
  666. 23:13So, and what this really means is that
  667. 23:15um
  668. 23:16P theta of XT is equals to one
  669. 23:20given X1 up to XT minus one.
  670. 23:23The chance to see the first token under
  671. 23:25this probabilistic model parameterized
  672. 23:26by theta is equals to
  673. 23:30This is a vector. The first
  674. 23:31entry of this vector, right? It's equals
  675. 23:33to softmax
  676. 23:36F theta X0 up to XT minus one
  677. 23:40indexed by one. This is the first entry
  678. 23:42of this vector. And you have many many
  679. 23:44other entries, right? Each entry is like
  680. 23:46you have another entry softmax
  681. 23:50F theta
  682. 23:51Sorry.
  683. 23:57two, and that is is equals to P theta of
  684. 24:01XT is equal to two. The chance of seeing
  685. 24:03the second
  686. 24:04second token given X1
  687. 24:07up to XT minus one.
  688. 24:09And you have so on and so forth.
  689. 24:12Does that answer the question?
  690. 24:14V is the
  691. 24:15uh the Yeah, this V absolute value is
  692. 24:17the size of the vocabulary.
  693. 24:21>> So, this is like a masked attention
  694. 24:23because the model is only conditional on
  695. 24:25like up to the previous
  696. 24:27ones after it, right?
  697. 24:28>> Yep.
  698. 24:29Yep. So, these are causal models so far.
  699. 24:31But I haven't defined the transformer
  700. 24:32yet. So, but I mean insisting that this
  701. 24:34is only So, guys, this is one of the
  702. 24:37reason I I'm using this notation. So, I
  703. 24:38mean insisting that UT
  704. 24:40is only a function of X0 and XT minus
  705. 24:42one, but not a function of the
  706. 24:45the the later uh inputs.
  707. 24:47So, but that will be um
  708. 24:50that will be guaranteed when we
  709. 24:51introduce the exact inner working of
  710. 24:54this transformer.
  711. 24:57>> So, functions of X are not the previous
  712. 24:59ones.
  713. 25:10>> Okay, good question. So the first the
  714. 25:12second question is easier to answer. So
  715. 25:14the embeddings is only a property of the
  716. 25:17this X, right? You don't have to see any
  717. 25:18other things, right? It's just actually
  718. 25:20it's literally if this is the first
  719. 25:21token then you just read the first row
  720. 25:23of this matrix. If this is the 10th
  721. 25:25token you use the 10th row and that's
  722. 25:26your embedding. No other dependencies.
  723. 25:28And while we are writing this as a
  724. 25:30as functions of the X but not embeddings
  725. 25:33that's just the
  726. 25:34it's just a
  727. 25:36It sometimes the embeddings is the inner
  728. 25:38working to some degree. I'm just trying
  729. 25:40to give you more I'm trying to unpack it
  730. 25:42a little bit, right? So So So So if you
  731. 25:44don't have to unpack it, you just work
  732. 25:45you can work with the text. Yeah. That's
  733. 25:47just the representation. Yeah.
  734. 25:50And it's one-to-one so in some sense
  735. 25:51there is no difference.
  736. 25:56Okay, so this is a full description of
  737. 25:58the probabilistic model. Once you have
  738. 26:00this probabilistic model you can compute
  739. 26:01the maximum likelihood given any data.
  740. 26:03We can say I have a data distribution I
  741. 26:05have a data point a sequence of data. I
  742. 26:07can see under this probabilistic model P
  743. 26:09theta what is the chance of seeing this
  744. 26:13this data. So that's that's what you
  745. 26:15need for computing the do the training,
  746. 26:17right? You You You need to compute the
  747. 26:18likelihood and then you maximize the
  748. 26:20likelihood as a function of theta,
  749. 26:22right? So
  750. 26:24and let's briefly discuss you know for
  751. 26:27generation, right? So suppose you
  752. 26:29already have the
  753. 26:30generation or decoding.
  754. 26:33Suppose you already have the model.
  755. 26:35Then
  756. 26:36a model theta then what do you do? So
  757. 26:38the
  758. 26:39as you can see the most natural thing is
  759. 26:41you just keep generating X1 X2 X3 so and
  760. 26:44so forth, right? You just say I'm going
  761. 26:46to generate X1 first from this you know
  762. 26:50P say softmax.
  763. 26:55Softmax, you know, F theta
  764. 26:58X1 up to X0 up to XT X, I guess I just
  765. 27:02rewrite
  766. 27:04uh, generically. So, you generate XT
  767. 27:05minus one from
  768. 27:07um, F theta X0 up to XT minus one.
  769. 27:10Right? You do this sequentially.
  770. 27:12Sometimes people draw this in this way.
  771. 27:13So, you say, "Oh, you have X0. You first
  772. 27:16go through the transformer
  773. 27:18and you get X1 and you feed X1 back to
  774. 27:21this transformer.
  775. 27:24Right? So,
  776. 27:25and um,
  777. 27:27and then you generate X2
  778. 27:29by sampling from the softmax. And then
  779. 27:32you you feed back. So, that's kind of
  780. 27:34why it's called autoregressive because
  781. 27:35you
  782. 27:36you have to like when you are
  783. 27:37generating, you don't have the sequence,
  784. 27:39right? When you are computing log
  785. 27:40likelihood, you see the whole you
  786. 27:41already see the whole sequence. You are
  787. 27:43just trying to figure out what's the
  788. 27:44chance of seeing this sequence. But,
  789. 27:46when you are generating, you you start
  790. 27:47with nothing. And then you have to
  791. 27:48generate your sequence yourself. And you
  792. 27:51That's why you start with a beginning of
  793. 27:52sentences sequence uh, token and then
  794. 27:54you generate first token and then you
  795. 27:56give the first token back to transformer
  796. 27:58and compute this um, softmax. And then
  797. 28:00you um, generate second token. You do a
  798. 28:02sampling to generate second token. You
  799. 28:04give it back and you keep doing this.
  800. 28:06So, that's the autoregressive
  801. 28:08generation.
  802. 28:09Um, the um,
  803. 28:12um, and uh, sometimes, you know, like
  804. 28:13uh, this generation only start from
  805. 28:16certain part because in many cases you
  806. 28:18are not generating from nothing. So, you
  807. 28:21are given maybe a few tokens. Maybe you
  808. 28:23are given So, the sometimes it's called
  809. 28:25prompt.
  810. 28:27So, prompt is like something the models
  811. 28:28are given. It's maybe like a question,
  812. 28:30you know, an instruction so and so
  813. 28:31forth. You say you are given X1 up to
  814. 28:34XK. And then the model is supposed to
  815. 28:37complete the generation. So, it start
  816. 28:39with XT plus one. So, XT plus one
  817. 28:42So, this is given.
  818. 28:44And then you start with XT plus one. You
  819. 28:45first generate XT + 1 from the softmax
  820. 28:49of the alpha theta X1
  821. 28:51up to X
  822. 28:53K
  823. 28:54and then you generate XT + 2
  824. 28:57and then you do the the same thing. So,
  825. 28:59it doesn't have to be always from the
  826. 29:00beginning.
  827. 29:02And then there's another small thing
  828. 29:03which is that when you generate, you
  829. 29:06don't necessarily have to always follow
  830. 29:08the same um model. And there's one thing
  831. 29:11that you can deviate from it which is
  832. 29:12that you can um add additional
  833. 29:15temperature
  834. 29:18where um you
  835. 29:19say as opposed to generating
  836. 29:22from the original
  837. 29:24um probabilistic model, I'm generate I'm
  838. 29:26going to divide this by a a temperature
  839. 29:28T tau or T sometimes. Um so, um and uh
  840. 29:33what this what what does this mean is
  841. 29:35that you change the
  842. 29:37you don't change the orders between this
  843. 29:39uh choices, right? So, this is a
  844. 29:41this this alpha theta is a vector,
  845. 29:43right? It's a a V-dimensional vector
  846. 29:45that describes the largest for each of
  847. 29:46the tokens in the vocabulary.
  848. 29:49And when you divide by tau, you don't
  849. 29:51change their relative orders. The bigger
  850. 29:52ones are still bigger, the smaller ones
  851. 29:53are small still smaller, but you change
  852. 29:56the
  853. 29:57the scale.
  854. 29:58So, um so, for example, before you could
  855. 30:00have like a
  856. 30:02So,
  857. 30:04suppose, you know, alpha theta
  858. 30:06X0 up to XK
  859. 30:09is
  860. 30:09is equals to, let's say,
  861. 30:12maybe like a 2 1
  862. 30:15-1, something like this.
  863. 30:17And if you divide by tau, let's say
  864. 30:19suppose you divide by tau which is very
  865. 30:21very small. Suppose tau is almost zero.
  866. 30:24Then after dividing by tau,
  867. 30:26like if you divide by tau, then what
  868. 30:28happens is that maybe this becomes 200.
  869. 30:30This becomes 100. I'm I'm giving you a a
  870. 30:32concrete example which is a
  871. 30:34a little bit unrealistic, but to
  872. 30:36demonstrate a point. And what's the
  873. 30:37difference between the softmax of this
  874. 30:39vector versus the softmax of the
  875. 30:41readjusted vector.
  876. 30:43The difference is that the softmax of
  877. 30:45this is much sharper.
  878. 30:46So, maybe let me just give you an
  879. 30:48example here. So,
  880. 30:50So, if you just have let's say maybe
  881. 30:52let's just have three dimensions. So,
  882. 30:53that it's easier to demonstrate the
  883. 30:55point.
  884. 31:00So, if you do softmax of this vector
  885. 31:05of 2 1 -1, what happens is that you get
  886. 31:11you sometimes get exponential 2
  887. 31:14over exponential 2
  888. 31:16plus exponential 1 plus exponential
  889. 31:20-1.
  890. 31:21That's the after softmax was the
  891. 31:23probability of the first
  892. 31:25token.
  893. 31:26Right? And the second token will be
  894. 31:27exponential 1
  895. 31:29exponential
  896. 31:312 plus exponential 1
  897. 31:34plus exponential -1.
  898. 31:36Right? The third one will be the the
  899. 31:38same denominator, the different
  900. 31:40numerator, right?
  901. 31:41So, that's a
  902. 31:42that's the softmax of this.
  903. 31:44But, if you take the softmax of the 200
  904. 31:46100 -100, suppose you do that.
  905. 31:55Then, this will give you
  906. 31:57exponential
  907. 31:58200
  908. 32:00versus exponential
  909. 32:02200 plus exponential
  910. 32:04100 plus exponential
  911. 32:07-100.
  912. 32:09Blah blah blah.
  913. 32:10So, what's the difference here? The
  914. 32:12difference is that in terms of the these
  915. 32:13three numbers, in terms of the order,
  916. 32:15it's always like this is bigger than
  917. 32:16this, this is bigger than this, right?
  918. 32:18Because 200 is big two is bigger than
  919. 32:19one, one is bigger than minus one. The
  920. 32:21order is always maintained.
  921. 32:23But, the
  922. 32:25the the relative place um um scale is
  923. 32:28different because here you can see
  924. 32:30exponential 200 is way way bigger than
  925. 32:32exponential 100. The difference is like
  926. 32:35so astro astronomical. Like it's too the
  927. 32:37the difference is too big so that these
  928. 32:39two numbers are just so small compared
  929. 32:41to exponential 200.
  930. 32:43So, they are even negligible.
  931. 32:44So, basically in this calculation, these
  932. 32:47two numbers don't even have to show up
  933. 32:48here. Like it don't even matter. So, you
  934. 32:50just get what?
  935. 32:51So, basically the first token will just
  936. 32:52have probably exactly almost exactly
  937. 32:54one. I mean, 0.9999 something. And then
  938. 32:56if you look at the second I entry,
  939. 32:59which is exponential 100
  940. 33:01divided by exponential
  941. 33:04200 plus exponential
  942. 33:06100 plus exponential
  943. 33:08-100.
  944. 33:11Again, everything is negligible compared
  945. 33:13to exponential 200.
  946. 33:14So, this is not very useful. This is not
  947. 33:16very useful. This is kind of like close
  948. 33:18to zero. So, basically this entry will
  949. 33:20be zero.
  950. 33:21So, that's why after you you you divide
  951. 33:23by tau to make this bigger, then this is
  952. 33:25just approximately
  953. 33:271 0 and 0. So, that's that's what I mean
  954. 33:30by it's much sharper. So, you you you
  955. 33:32you sharpen the the distribution to make
  956. 33:34it most of the distribution to be the
  957. 33:37the probability mass to be on the the
  958. 33:38most likely token.
  959. 33:40And however, here
  960. 33:42exponential two and exponential one are
  961. 33:43not that different. If you compute this,
  962. 33:45you'll get something like maybe the
  963. 33:47first one will be like maybe 0.7, the
  964. 33:49second one will be 0.3. I I don't know.
  965. 33:51Like we can use a calculator to see. But
  966. 33:53you know, it will still be like this one
  967. 33:54is bigger than this. And the second one
  968. 33:56is bigger than the third one, but it
  969. 33:57won't be like
  970. 33:58all the mass is on the first token.
  971. 34:00So,
  972. 34:01so so basically what I'm trying to say
  973. 34:02is that if tau is a small number so that
  974. 34:04you
  975. 34:05mhm
  976. 34:05you you amplify the logits, then you are
  977. 34:08sharpening the resulting distribution to
  978. 34:10make it more like a deterministic
  979. 34:11distribution.
  980. 34:12And then as a result, you are just
  981. 34:14generating
  982. 34:15um you are just a sampling from the the
  983. 34:17most likely token. You are taking the
  984. 34:19max as opposed to
  985. 34:20sampling from it.
  986. 34:22So, and um and and that's actually what
  987. 34:25people actually in
  988. 34:26um um do. Like they say basically like I
  989. 34:30think actually
  990. 34:31uh in API call, you have the choice to
  991. 34:33choose the temperature
  992. 34:34and if you choose a small temperature it
  993. 34:36means that you want more deterministic
  994. 34:37generation.
  995. 34:39If you choose a temperature to be zero
  996. 34:40it means that your generation will just
  997. 34:41be always deterministic every time you
  998. 34:43take the largest
  999. 34:46most likely token and then keep going on
  1000. 34:48right and of course you can also change
  1001. 34:50the temperature to be smaller than
  1002. 34:53bigger than one which makes the the
  1003. 34:55distribution more softer so that there's
  1004. 34:56a you focus on more of the long tail
  1005. 34:58like you can generate more tokens from
  1006. 35:01the tail so you have more stochasticity
  1007. 35:03and uncertainty in your generation.
  1008. 35:05Sometimes that's useful because you want
  1009. 35:07to encourage some diversity in your
  1010. 35:09generation and you you do that
  1011. 35:11especially in the RL sometimes you have
  1012. 35:12to do that
  1013. 35:14um
  1014. 35:15Cool. Okay, sounds good.
  1015. 35:18And sometimes there are also other ways
  1016. 35:19of something which is
  1017. 35:21probably I shouldn't go into details
  1018. 35:23there are some notes
  1019. 35:24in the lecture notes. So another way to
  1020. 35:26do it is that you
  1021. 35:28this vector right now I have three
  1022. 35:30dimensions but actually it's like a
  1023. 35:32the dimensionality is V right like which
  1024. 35:34is like 250 or something. You can take
  1025. 35:36top K and then drop all of the other
  1026. 35:39dimensions so you only sample from the
  1027. 35:41the most likely K tokens and you
  1028. 35:43re-normalize the distribution. So
  1029. 35:45basically the the first K tokens may not
  1030. 35:48cover the all the probabilistic
  1031. 35:49probability mass so
  1032. 35:51but still that's probably easier to
  1033. 35:53generate because you don't want to
  1034. 35:54generate too rare tokens so you just
  1035. 35:56drop everything below top K and then
  1036. 35:58generate.
  1037. 36:00Okay,
  1038. 36:02sounds good. So I think I'm
  1039. 36:04probably
  1040. 36:10And then okay so once you have so this
  1041. 36:13is the generation and then
  1042. 36:15for the training I think I already kind
  1043. 36:17of like describe it
  1044. 36:18a little bit
  1045. 36:20so when you do the training
  1046. 36:24you just
  1047. 36:25compute a log
  1048. 36:26the negative
  1049. 36:29log likelihood.
  1050. 36:33Like this.
  1051. 36:36The
  1052. 36:37negative log likelihood. And you say
  1053. 36:40this is a log, so if you have example X1
  1054. 36:42up to XT
  1055. 36:44So the minus log P theta X1 up to XT
  1056. 36:50capital T is just the
  1057. 36:53minus log the product of all of this,
  1058. 36:56right?
  1059. 36:58P theta X1
  1060. 37:00P theta X2 given X1, so on and so forth.
  1061. 37:11And this is the minus the sum of the log
  1062. 37:14of the P of XT
  1063. 37:16given X1 up to XT minus 1.
  1064. 37:21And then you change this to
  1065. 37:24log of the softmax
  1066. 37:28of P1
  1067. 37:29of this F theta.
  1068. 37:37And here this XT is a particular choice
  1069. 37:39of XT, right? Which is seen in the data.
  1070. 37:42So that means you are looking at the
  1071. 37:44XT's entry of the softmax. This softmax
  1072. 37:47will give you the probability for every
  1073. 37:48possible tokens. And then you are only
  1074. 37:50interested in the probability to see
  1075. 37:52that particular X token XT. So that's
  1076. 37:54why you're looking at XT, the index of
  1077. 37:57that vector. And that's a
  1078. 37:59single scalar, you take the log and you
  1079. 38:01take the sum over T.
  1080. 38:02Um and that's your loss function.
  1081. 38:06And then you can optimize this loss
  1082. 38:07function by, for example, some
  1083. 38:09optimizers,
  1084. 38:10um like SGD or Adam. Uh we didn't
  1085. 38:13discuss Adam in this lecture, which is
  1086. 38:14an extension of the X SGD optimizer, but
  1087. 38:17mostly you can just call the optimizer
  1088. 38:20as a black box to optimize the loss
  1089. 38:21function.
  1090. 38:38It's not actually not the last one,
  1091. 38:39right? Because this is sum from T to
  1092. 38:41>> Oh.
  1093. 38:42>> 1 to capital T.
  1094. 38:44The little T and large T. Little T.
  1095. 38:54Okay. Sounds good.
  1096. 38:57Okay. So, then in the next uh
  1097. 38:5930 minutes, I'm going to tell you what
  1098. 39:01is this function F is, right? Or maybe
  1099. 39:04what's hidden in this transformer.
  1100. 39:06What's the operation we have to do to
  1101. 39:08compute um the the the logits here.
  1102. 39:12Um and it's pretty complicated. So, um
  1103. 39:15it probably takes more than 30 minutes,
  1104. 39:16but I'll see how much I can do.
  1105. 39:20So, um
  1106. 39:25Any questions before that?
  1107. 39:27Okay.
  1108. 39:28So, at a high level, this transformer
  1109. 39:30has multiple layers.
  1110. 39:32And uh so, you start with like a EX1 up
  1111. 39:35to EXT.
  1112. 39:37And you first apply a so-called
  1113. 39:39attention layer,
  1114. 39:42which I will define.
  1115. 39:43And then you apply the so-called MLP.
  1116. 39:46So,
  1117. 39:47um first of all, the attention layer,
  1118. 39:49the attention always takes in
  1119. 39:51uh a sequence of vectors and output a
  1120. 39:53sequence of vectors.
  1121. 39:55So, you output a bunch of vectors. At
  1122. 39:57each position, you output one vector.
  1123. 39:59And then, each of these vector will go
  1124. 40:01through a MLP. MLP means like a
  1125. 40:04multi-layer perception, which is just a
  1126. 40:05two or three-layer
  1127. 40:07two or three-layer network.
  1128. 40:09And uh each of these vector will go
  1129. 40:11through this MLP.
  1130. 40:15And then, you apply another attention.
  1131. 40:17This attention will takes in
  1132. 40:19many, many vectors and and then output a
  1133. 40:22sequence of vectors.
  1134. 40:24And then you do the MLP.
  1135. 40:28So, that's the the high-level um
  1136. 40:31structure of this attention. And
  1137. 40:32eventually, you have to output
  1138. 40:35Remember, eventually, you have to output
  1139. 40:37a sequence of vectors. So, basically,
  1140. 40:39you do enough of these alternations
  1141. 40:41between these two and eventually you
  1142. 40:42have this uh
  1143. 40:43uh MLP and you output
  1144. 40:46some vectors U1 up to UT.
  1145. 40:49Right.
  1146. 40:50And you alternate between attention and
  1147. 40:52MLP.
  1148. 40:53That's the high-level idea.
  1149. 40:55And um and I think we'll focus most of
  1150. 40:59the our energy on defining attention,
  1151. 41:00which is kind of like the most uh uh new
  1152. 41:03thing, right? Because MLP is just a
  1153. 41:04two-layer network uh with some special
  1154. 41:07kind of gating uh or kind of like some
  1155. 41:09kind of like a special activation
  1156. 41:10functions. Uh and another small thing is
  1157. 41:13that the MLP are sharing the parameters
  1158. 41:16and they are applying independently. So,
  1159. 41:18each of these vectors is
  1160. 41:20processed by the MLP in an independent
  1161. 41:22fashion. There's no kind of dependencies
  1162. 41:24between positions when you apply the
  1163. 41:26MLP. They are just very separate
  1164. 41:28operations. So, the only way that you
  1165. 41:30are using um um you are kind of like
  1166. 41:33fuse you fuse the input from different
  1167. 41:35positions or different time steps is the
  1168. 41:37attention. So, the attention will look
  1169. 41:38at all those vectors and and process a
  1170. 41:41little bit and then output a sequence of
  1171. 41:42vectors. And MLP is is very parallelized
  1172. 41:45and parallelizable and independent.
  1173. 41:49Okay. So,
  1174. 41:50now let's let's just define our task.
  1175. 41:58This entire thing will be treated as a
  1176. 42:00single system. So, the So, So,
  1177. 42:01basically, once you define the
  1178. 42:02architecture, you apply the auto
  1179. 42:04differentiation.
  1180. 42:06You you like a you treat this whole
  1181. 42:08thing as a as a gigantic network. Yeah,
  1182. 42:10end to end.
  1183. 42:20>> [clears throat]
  1184. 42:21>> Yeah, but the auto differentiation still
  1185. 42:22works.
  1186. 42:23So, so the auto differentiation, I think
  1187. 42:25as we discussed, if you remember like
  1188. 42:27it's kind of like fundamental is
  1189. 42:28claiming that if you have a way to
  1190. 42:30compute your forward pass,
  1191. 42:32as long as they are compositions of many
  1192. 42:33many different components, you know,
  1193. 42:35like a basic components, you can compute
  1194. 42:37a backward pass the same fast the same
  1195. 42:39way. So, here
  1196. 42:41is is indeed a composition of many many
  1197. 42:42small arithmetic operations because as
  1198. 42:45you see the attention layer has many
  1199. 42:49small matrix multiplications and the MLP
  1200. 42:51has many matrix multiplications and you
  1201. 42:52can automatically differentiate.
  1202. 42:56Um
  1203. 42:57Okay, so
  1204. 43:01Okay, so now let's let's talk about
  1205. 43:02attention.
  1206. 43:14So, for attention,
  1207. 43:16one thing to remember first is that the
  1208. 43:18type signature is that you your input a
  1209. 43:20sequence of vectors to the attention and
  1210. 43:22they output a sequence of vectors.
  1211. 43:25And and
  1212. 43:27we're going to start with the so-called
  1213. 43:28single head attention.
  1214. 43:35So, the single head attention is like
  1215. 43:36this. So, you first have um
  1216. 43:40uh input as a a sequence of vectors.
  1217. 43:44Um let's say H1.
  1218. 43:46Say
  1219. 43:47uh in
  1220. 43:48up to HT in.
  1221. 43:51And and you want to output, you know, H.
  1222. 43:53I'm trying to define the the type
  1223. 43:55signature.
  1224. 43:57So, your output
  1225. 43:59H1 out up to HT out.
  1226. 44:03And what you do is the following. So,
  1227. 44:06the first thing is that you turn
  1228. 44:09this input vectors into a set of query
  1229. 44:11key
  1230. 44:13and and the value vectors.
  1231. 44:15So
  1232. 44:17So you define so basically you say H
  1233. 44:22one in
  1234. 44:24you multiply it by a matrix WQ. This is
  1235. 44:26a weight matrix. This is a parameter.
  1236. 44:30And
  1237. 44:31by the way, let's treat all of these as
  1238. 44:32row vectors and then all the
  1239. 44:33multiplications are on the right hand
  1240. 44:35side as opposed to on the left hand
  1241. 44:36side. So every vector is a row vector
  1242. 44:38and you multiply on the right hand side
  1243. 44:40you get another row vector.
  1244. 44:41So so this H in let's say this is like a
  1245. 44:44dimension D
  1246. 44:46RD
  1247. 44:47I say say R one by D and this WQ is in
  1248. 44:50dimension
  1249. 44:53So this is in R one by D and this WQ is
  1250. 44:56in
  1251. 44:57R D by
  1252. 44:59um
  1253. 45:00DH.
  1254. 45:02DH is a parameter. It's a dimension you
  1255. 45:04choose. um
  1256. 45:06And and then the resulting thing is
  1257. 45:09called the Q
  1258. 45:11one which is in dimension
  1259. 45:14DH. Maybe one by DH depending on how you
  1260. 45:16think about it.
  1261. 45:18um
  1262. 45:19And you can do this for every index
  1263. 45:20every time step. Here I'm only doing it
  1264. 45:22for time step one. You can do it for
  1265. 45:23time step T. So you have QT.
  1266. 45:26So basically this give you a factor a
  1267. 45:28sequence of vectors Q1 up to Q capital
  1268. 45:30T. Each of them is in dimension
  1269. 45:33uh
  1270. 45:34one by DH.
  1271. 45:36And this these are called queries.
  1272. 45:40And you do the same thing for keys and
  1273. 45:42values. So keys
  1274. 45:45Let me define them first. So key
  1275. 45:47is the same thing. You say the HT in
  1276. 45:50multiplied by a so-called key projection
  1277. 45:52matrix WK
  1278. 45:54and that's your KT. And the
  1279. 45:56dimensionality I think is the same. So
  1280. 45:58this is also in the same dimension and
  1281. 46:00the KT is in the same dimension. And the
  1282. 46:03same thing for value function value
  1283. 46:04vectors, not value functions. Uh HT
  1284. 46:08in times W
  1285. 46:11V.
  1286. 46:12So,
  1287. 46:13in some sense maybe
  1288. 46:15maybe this is the right time to
  1289. 46:18talk about some of the intuitions, you
  1290. 46:19know, for the attentions, which is not
  1291. 46:21that obvious because eventually it just
  1292. 46:22works, you know, like nobody really know
  1293. 46:24exactly how it works. Um um but the the
  1294. 46:27rough intuition is that you have to have
  1295. 46:28some interactions between these inputs
  1296. 46:30to make it even meaningful, right?
  1297. 46:32Because if everything is independent,
  1298. 46:33then you just dealing with all the time
  1299. 46:35steps independently.
  1300. 46:36So, that's why you have to have some um
  1301. 46:38um interactions. And the interaction
  1302. 46:40sometimes is about how do you
  1303. 46:42kind of like a retrieve or kind of pay
  1304. 46:45attention to each other. So, that's
  1305. 46:47actually how the word attention come
  1306. 46:48from. So, basically, when you are
  1307. 46:50processing the maybe this vector,
  1308. 46:53you probably should attend to or pay
  1309. 46:55attention to some other vectors which
  1310. 46:57are have relevant information.
  1311. 46:59And which one you you pay attention to
  1312. 47:01is what we're going to decide here using
  1313. 47:02with this
  1314. 47:03query and key and vectors. So, in some
  1315. 47:05sense, intuitively, the attention will
  1316. 47:07what attention will do, which I didn't
  1317. 47:09show you yet, is that
  1318. 47:11you try to
  1319. 47:12um for processing each of the vectors,
  1320. 47:14you try to figure out which are the
  1321. 47:16relevant vectors
  1322. 47:18that has relevant information. And and
  1323. 47:20then you pay actual attention to those
  1324. 47:22vectors.
  1325. 47:23And and then you compute the output
  1326. 47:25depends on depending on those special
  1327. 47:27vectors to some degree.
  1328. 47:29So, that's what we are trying to do.
  1329. 47:31Um and
  1330. 47:32and to kind of like decide which one you
  1331. 47:34pay attention to, you have to prepare
  1332. 47:36these queries and key and value vectors.
  1333. 47:39Um
  1334. 47:40and how do you use them? So, you use the
  1335. 47:41query and keys as a way to decide which
  1336. 47:44one you pay attention to. So, basically,
  1337. 47:46let's say um suppose you have um
  1338. 47:49uh the teeth position, right?
  1339. 47:52I want to decide
  1340. 47:53what the teeth position should pay
  1341. 47:55attention to. Um and and
  1342. 47:58and the way to do it is you say you
  1343. 47:59compute
  1344. 48:00the query QT.
  1345. 48:03This is uh the query at this position.
  1346. 48:06And you multiply the query with the keys
  1347. 48:09at all the other positions.
  1348. 48:11So, maybe I I will just draw it.
  1349. 48:14So, you have queries Q1 up to
  1350. 48:16QT.
  1351. 48:18And one of them is
  1352. 48:19Q little T. And you have Q
  1353. 48:22K1 up to KT.
  1354. 48:25KT.
  1355. 48:27And you you multiply Qs with all the Ks.
  1356. 48:33So, you have Q1 times, you know, K1
  1357. 48:36transpose,
  1358. 48:37QT times K2 transpose,
  1359. 48:40and
  1360. 48:41QT times K capital T transpose.
  1361. 48:45Um, and there is a um
  1362. 48:47sometimes there's a scaling factor here
  1363. 48:50to kind of like normalize this in a way
  1364. 48:51that
  1365. 48:52makes it a little more kind of like a
  1366. 48:54well-behaved, but I think for the moment
  1367. 48:56you probably don't have to pay too much
  1368. 48:58attention to the scaling factor. So, the
  1369. 49:00important thing is that you are taking
  1370. 49:01the inner product between Qs and each of
  1371. 49:03the other Ks. And this QT transpose K
  1372. 49:06Q
  1373. 49:07K transpose is actually inner product
  1374. 49:09because Q is a
  1375. 49:11is a row vector,
  1376. 49:13and K transpose is a column vector.
  1377. 49:17And you are taking the inner product
  1378. 49:19between these row vectors and the column
  1379. 49:21vectors.
  1380. 49:22Um,
  1381. 49:23and you get one scalar out of it.
  1382. 49:25So,
  1383. 49:26so this is a
  1384. 49:28all of these numbers are real numbers,
  1385. 49:29right? So, this is in RT. So, you have T
  1386. 49:32real numbers.
  1387. 49:34Um, and then you take a softmax over
  1388. 49:36these T real numbers to turn it into a
  1389. 49:37probability vector.
  1390. 49:39So, then you say I'm going to do a
  1391. 49:40softmax.
  1392. 49:45Say softmax
  1393. 49:48of this whole thing.
  1394. 49:57>> And now you get what? You get a This is
  1395. 49:59a vector.
  1396. 50:00The dimensionality here is RT and each
  1397. 50:03of this This is this the the outcome of
  1398. 50:06this softmax is a sequence of numbers,
  1399. 50:08right?
  1400. 50:10And there are T numbers and all of these
  1401. 50:13numbers are non-active and the sum of
  1402. 50:14them is one. So it's a probability
  1403. 50:16vector over T positions. And if you see
  1404. 50:19a large number, it means that
  1405. 50:22Let's say suppose this number is big, it
  1406. 50:24means that this number is somewhat big
  1407. 50:27compared to other ones. It means that
  1408. 50:28your Qs and keys have like more higher
  1409. 50:31inner product. Right? And uh
  1410. 50:34it means that you should pay more
  1411. 50:35attention to K K1. Right? So So
  1412. 50:38basically if any of these entries big,
  1413. 50:39it means that
  1414. 50:42intuitively
  1415. 50:44QT needs to pay more attention to K1.
  1416. 50:46And uh that's will be reflected by the
  1417. 50:48probability vectors because the the
  1418. 50:50corresponding probability will also be
  1419. 50:51high.
  1420. 50:52And
  1421. 50:54so I think let's call this I guess I
  1422. 50:55have a notation here, so let's call this
  1423. 50:58let's denote this to be PT1,
  1424. 51:01PT2,
  1425. 51:03up to PT capital T.
  1426. 51:06So the So each of the PTI is
  1427. 51:10each of the PTI is non-active and the
  1428. 51:12sum of the PTI over I from one to T is
  1429. 51:17equal to one. So this is a probability
  1430. 51:18vector.
  1431. 51:19Um and and if this if any of these
  1432. 51:21entries is big, it means that we should
  1433. 51:23pay up more attention to that particular
  1434. 51:24position when processing the little T's
  1435. 51:27query.
  1436. 51:29And then how do you turn this How does
  1437. 51:30this relate to the final
  1438. 51:32uh outcome? So finally, once you have
  1439. 51:35this, then
  1440. 51:36uh you say the
  1441. 51:38the output at T's
  1442. 51:41position, which is a vector,
  1443. 51:43is using these probability vectors to
  1444. 51:45weight
  1445. 51:46uh the values.
  1446. 51:49The values is um
  1447. 51:52V's, which are we are defining here,
  1448. 51:53right? VT, sorry, VT.
  1449. 51:55So, you say it's PT1 * V1 + PT2
  1450. 52:00* V2 +
  1451. 52:03PTT * VT.
  1452. 52:06So, basically, if each of any of this
  1453. 52:09probability is much higher, then you are
  1454. 52:12multiplying this high probability
  1455. 52:15with the the corresponding value at that
  1456. 52:16position.
  1457. 52:18So, basically, you are using the V
  1458. 52:20Right? If, for example, if this number
  1459. 52:21is very big, then you are using V2
  1460. 52:24a lot in the output at time T.
  1461. 52:32So, you determine them by multiplying
  1462. 52:34the matrices.
  1463. 52:35And the matrices will be trained.
  1464. 52:37Uh uh uh
  1465. 52:38like like yeah. How how do you attend to
  1466. 52:40each other is also trained. Like So, so
  1467. 52:43you are only giving a template in some
  1468. 52:44sense. Like how how you know Like it may
  1469. 52:46be that, you know,
  1470. 52:48like
  1471. 52:49it could be that all of these are zero,
  1472. 52:51you know. Then you are not doing
  1473. 52:52anything with it. And then your network
  1474. 52:54the loss function will be not good, but
  1475. 52:56it's still a valid network. It's just
  1476. 52:57the loss is not good. Yeah.
  1477. 52:59So, eventually, you are going to train
  1478. 53:00the model to
  1479. 53:02behave well so that So, basically,
  1480. 53:06the training algorithms will choose the
  1481. 53:07right W so that this mechanism is
  1482. 53:10uh uh doing the right thing.
  1483. 53:27So,
  1484. 53:28um let's try to articulate you know, I
  1485. 53:29think it's a good question, but let's
  1486. 53:30try to articulate so that we can answer
  1487. 53:32it. So, what's the
  1488. 53:34uh
  1489. 53:36You are say So,
  1490. 53:43>> [clears throat]
  1491. 53:43>> Yeah.
  1492. 53:44>> So, why
  1493. 54:00>> [clears throat]
  1494. 54:00>> So, I think you are asking why you need
  1495. 54:02such an architecture but not just only
  1496. 54:04using MLP for example, right? That's one
  1497. 54:06way to rephrase the question.
  1498. 54:08I think
  1499. 54:10um
  1500. 54:11So, if you are only using MLP in this
  1501. 54:13way
  1502. 54:14suppose you remove all the tensions you
  1503. 54:16are you only using MLP in this way, then
  1504. 54:17there's no interactions between
  1505. 54:19different positions whatsoever.
  1506. 54:21So, there's no chance that's going to
  1507. 54:23work, right? You're just applying
  1508. 54:24separate networks for each of the
  1509. 54:26tokens. Uh there's no chance you can
  1510. 54:28understand the the the the the the
  1511. 54:31dependencies the the meanings, right?
  1512. 54:33Like there's no chance I I can I can
  1513. 54:34just give you individual words and you
  1514. 54:36can understand a sentence. There's no
  1515. 54:37chance, right? So, so you do have to um
  1516. 54:40have a model that can read uh uh
  1517. 54:43um all the tokens you know and and use
  1518. 54:46and study their kind kind of
  1519. 54:47correlations and combinations and so
  1520. 54:49forth, right?
  1521. 54:50But um there are other possibility
  1522. 54:53possible ways to deal with that. For
  1523. 54:54example, what if you you use a gigantic
  1524. 54:56MLP
  1525. 54:57uh that takes in all of these maybe you
  1526. 54:59concatenate all of them and you take in
  1527. 55:01one MLP. Right? So, so um for example,
  1528. 55:05you just concatenating all of these
  1529. 55:06tokens
  1530. 55:07vectors into one long vector and you go
  1531. 55:09through a network. So, there are several
  1532. 55:10issues with that. One is that
  1533. 55:13um you will have way too many
  1534. 55:15parameters. Because your sequence lines
  1535. 55:18maybe sometimes even millions. Then if
  1536. 55:20you have a million by million matrices
  1537. 55:23in your MLP, it's going to be way too
  1538. 55:24big. Right now, the MLP is actually very
  1539. 55:26small. It only takes in a vector for one
  1540. 55:29position. It doesn't even depend on T.
  1541. 55:32And another thing is that you're not
  1542. 55:34only the the size of the MLP is big, but
  1543. 55:35also the size of the MLP is also changes
  1544. 55:37over
  1545. 55:39uh as you change the sequence length,
  1546. 55:41which is also very undesirable because
  1547. 55:42for example, if you tune with sequence
  1548. 55:44length 1,000 and then you test with
  1549. 55:45sequence length 1,001, then you have to
  1550. 55:48change your parameter count, which is uh
  1551. 55:50not good. And and here, if you really
  1552. 55:53look into details, you'll find out that
  1553. 55:55the number of parameters doesn't depend
  1554. 55:58on the sequence length.
  1555. 55:59>> So, what would be
  1556. 56:01like you have this gigantic MLP
  1557. 56:03>> Yep.
  1558. 56:03>> that is too big to compute.
  1559. 56:05>> Yep.
  1560. 56:05>> But you kind of know that this there's
  1561. 56:08there is some structure because there
  1562. 56:09are
  1563. 56:10like words that match other words. So,
  1564. 56:12this is kind of trying to learn
  1565. 56:14structure
  1566. 56:15in a more efficient way.
  1567. 56:17>> Exactly, right? So, so so you you know
  1568. 56:19that, you know, if you have sufficient
  1569. 56:21compute, you know, probably your
  1570. 56:22gigantic MLP will work. If you have
  1571. 56:24sufficient data so you don't overfit,
  1572. 56:26and you have sufficient compute, then
  1573. 56:28there is some MLP that can actually
  1574. 56:30solve this problem.
  1575. 56:31But that's too expensive. So, you need
  1576. 56:33to prescribe some
  1577. 56:36structure
  1578. 56:37so that you can reduce the complexity.
  1579. 56:39Right? And and this is the the the
  1580. 56:41prescription we have. Um but you still
  1581. 56:43have to leave enough flexibilities. You
  1582. 56:45cannot just say, oh, for example, one
  1583. 56:46way to do it is you just say, let's
  1584. 56:47average. Then there's no parameters, you
  1585. 56:49just uh prescribe too much, right? You
  1586. 56:51you dictate too much, then that's not
  1587. 56:53good. So, it's kind of like an you need
  1588. 56:55to strike a balance between how much
  1589. 56:57structure you introduce and how much
  1590. 56:59flexibilities you give the models to
  1591. 57:00figure out.
  1592. 57:03Yeah. Um
  1593. 57:06Yeah. Any other questions?
  1594. 57:09By the way, there's actually in the
  1595. 57:11history, there was actually other ways
  1596. 57:12to deal with the sequence. For example,
  1597. 57:14you can use a convolutional network on
  1598. 57:16this, which is kind of like MLP, so you
  1599. 57:18have use a
  1600. 57:19use a MLP but with a specific receptive
  1601. 57:22field so that you each time you only see
  1602. 57:24a kind of like a window. I think there
  1603. 57:27are still
  1604. 57:28papers that so says that
  1605. 57:31that's not
  1606. 57:32a terrible idea. Actually, if you use a
  1607. 57:34convolutional network, it actually can
  1608. 57:35work to some degree. It's just not
  1609. 57:37as good as this one, you know. Uh
  1610. 57:40and uh and and also you can see like
  1611. 57:42there are also other
  1612. 57:44Probably we'll talk about that later.
  1613. 57:45So, there are other variants of
  1614. 57:46attentions too that does similar stuff
  1615. 57:48but have better efficiency.
  1616. 57:50Um
  1617. 57:51Anyway, so let's continue.
  1618. 57:53We have We are not done yet, actually.
  1619. 57:55This is only a single so-called single
  1620. 57:56head attention, you know. Um actually,
  1621. 57:59the real thing is a little more
  1622. 58:00complicated.
  1623. 58:02Um
  1624. 58:03but
  1625. 58:04let's see. Before I go into that, let's
  1626. 58:06see.
  1627. 58:11So, I think
  1628. 58:13you can see the notation here is very
  1629. 58:14messy. So, let's try to simplify it a
  1630. 58:17little bit so that we don't have to
  1631. 58:19always work with this messy notation.
  1632. 58:22So, one typical way that people simplify
  1633. 58:24this is that you try to organize all the
  1634. 58:26queries
  1635. 58:28into
  1636. 58:29a matrix.
  1637. 58:30Recall that each of the queries is a row
  1638. 58:32vector, and if you put them all in a
  1639. 58:34matrix, then this is uh
  1640. 58:36uh a matrix of dimension T by DH.
  1641. 58:40And you can put all the keys
  1642. 58:42into a matrix.
  1643. 58:45And you can put all the values in in a
  1644. 58:47matrix as well.
  1645. 58:51Um
  1646. 58:53All right. So, and then you can write
  1647. 58:54this H out
  1648. 58:56um
  1649. 59:05the output signal factors in either kind
  1650. 59:08of like a
  1651. 59:10um
  1652. 59:11also as a matrix is in some sense you do
  1653. 59:14a softmax
  1654. 59:16of the row of
  1655. 59:18this QK transpose / C matrix * V.
  1656. 59:22So, here the basically this QK matrix QK
  1657. 59:25transpose matrix is the
  1658. 59:28the family of Q I times KJ transpose
  1659. 59:32divided by C.
  1660. 59:34So, every entry is of this form,
  1661. 59:37which is kind of like
  1662. 59:38these guys.
  1663. 59:39And then you if you do a row softmax,
  1664. 59:41that corresponds to this softmax.
  1665. 59:43And then if you multiply by V,
  1666. 59:46it means it means this linear
  1667. 59:47combination. I think you you we probably
  1668. 59:49we don't have time enough time to
  1669. 59:51um go into the details like um but you
  1670. 59:54can double check, you know, it's just a
  1671. 59:55very linear algebra. It's just you are
  1672. 59:57rearranging like it's it's mostly just
  1673. 59:59notations. Right, so this QK divided by
  1674. 1:00:01C matrix is the inner product between
  1675. 1:00:03the queries and keys. And then you take
  1676. 1:00:05a row row wise softmax to get the
  1677. 1:00:07probabilities. So, after the softmax,
  1678. 1:00:09this matrix becomes some
  1679. 1:00:11uh every row is a probability vector and
  1680. 1:00:13then you multiply that with the V to you
  1681. 1:00:15take the linear combination of these.
  1682. 1:00:18Um and uh
  1683. 1:00:20there's one thing that we kind of cheat
  1684. 1:00:22a little bit here, which is that if you
  1685. 1:00:24look at this
  1686. 1:00:25architecture, then
  1687. 1:00:28it's not true that
  1688. 1:00:30the teeth output
  1689. 1:00:32only depends on everything before it.
  1690. 1:00:36There's no this kind of like a auto
  1691. 1:00:38aggressive kind of like properties yet,
  1692. 1:00:39right? Because
  1693. 1:00:41the output
  1694. 1:00:42at time T
  1695. 1:00:44depends on and everything, right? It
  1696. 1:00:47depends on V capital T even the last
  1697. 1:00:49input. V capital T V capital T is a
  1698. 1:00:52function of the last input.
  1699. 1:00:54And so and the output at time T depends
  1700. 1:00:57on the input um um
  1701. 1:00:59depends on VT and it depends on HT uh
  1702. 1:01:02in. So, um
  1703. 1:01:04so basically there's no kind of like
  1704. 1:01:05this dependency. So, then what you do is
  1705. 1:01:07you say um
  1706. 1:01:09you introduce concept of masking to
  1707. 1:01:12um make sure that there's no
  1708. 1:01:14uh
  1709. 1:01:15this kind of like a wrong dependency you
  1710. 1:01:17want. So, basically what you really want
  1711. 1:01:18to do is that you just uh
  1712. 1:01:20in some sense
  1713. 1:01:21for each little t, you want to delete
  1714. 1:01:24everything after
  1715. 1:01:26uh time t. So, basically, um
  1716. 1:01:29I say
  1717. 1:01:31before you have this, you know,
  1718. 1:01:33p t v
  1719. 1:01:35t plus
  1720. 1:01:39p t t v t p t capital t v capital t.
  1721. 1:01:43So, basically, you say these
  1722. 1:01:45dependencies are invalid for me, so I
  1723. 1:01:46just
  1724. 1:01:47don't have them. So, basically, you just
  1725. 1:01:49delete all of this.
  1726. 1:01:50And um
  1727. 1:01:52and then you re-normalize these p's so
  1728. 1:01:54that the sum of these p's
  1729. 1:01:56uh got normalized to one. So, that is
  1730. 1:01:58still a linear combination or convex
  1731. 1:01:59combination of the v one up to v t.
  1732. 1:02:02So, that's pretty much what you do. So,
  1733. 1:02:03you you just say if the the dependency
  1734. 1:02:05is um not good, you just
  1735. 1:02:07sometimes mask them out.
  1736. 1:02:09And mathematically, what you do, uh
  1737. 1:02:10actually, also implementation-wise, uh
  1738. 1:02:12what you do is through this masking,
  1739. 1:02:14which means that
  1740. 1:02:17you
  1741. 1:02:19um
  1742. 1:02:23which means that you add
  1743. 1:02:25um
  1744. 1:02:27a masking here.
  1745. 1:02:29Maybe I should use a different color to
  1746. 1:02:31so that they're
  1747. 1:02:33So, you add a masking here.
  1748. 1:02:35Let me use a different
  1749. 1:02:36color
  1750. 1:02:43Okay.
  1751. 1:02:44So, you add a mask here
  1752. 1:02:46so that um you can mask out the
  1753. 1:02:51Let me see. You can mask out half of the
  1754. 1:02:53this matrix.
  1755. 1:02:54So,
  1756. 1:02:56so you have like a q one k one transpose
  1757. 1:02:59q one k two transpose so on and so
  1758. 1:03:01forth.
  1759. 1:03:03And q one k two transpose is not valid
  1760. 1:03:05in some sense because you are looking
  1761. 1:03:06you are tending to something you
  1762. 1:03:08shouldn't attend, right? Q one shouldn't
  1763. 1:03:10be able to see k two. So, basically, you
  1764. 1:03:12add the this masking matrix. You what
  1765. 1:03:14you add is that you add
  1766. 1:03:16minus infinity.
  1767. 1:03:19You add a minus infinity here.
  1768. 1:03:23And for
  1769. 1:03:25Q1 K3 transpose, you also add minus
  1770. 1:03:27infinity.
  1771. 1:03:29So and so forth.
  1772. 1:03:31And then in the next row you have Q2 K1
  1773. 1:03:34transpose, which is good. You have Q2 K2
  1774. 1:03:36transpose, which is good. And you have
  1775. 1:03:38Q2 K3 transpose, which is not good. So
  1776. 1:03:42you add minus infinity here.
  1777. 1:03:44So and so forth. So basically everything
  1778. 1:03:46here
  1779. 1:03:47will be
  1780. 1:03:48adding additional minus infinity. So
  1781. 1:03:50basically this
  1782. 1:03:51upper triangle um part of this matrix
  1783. 1:03:54will be all minus infinity. You know, I
  1784. 1:03:55mean
  1785. 1:03:57by adding minus infinity you just mean
  1786. 1:03:58that you you turn it into minus
  1787. 1:03:59infinity. Right? Any number
  1788. 1:04:01added by minus infinity is the minus
  1789. 1:04:03infinity. And then what happens is that
  1790. 1:04:06or in other words, it's kind of like you
  1791. 1:04:07are just uh
  1792. 1:04:08trying to
  1793. 1:04:11add all of minus infinity to all of this
  1794. 1:04:12part.
  1795. 1:04:14And then when you add when you turn all
  1796. 1:04:15of this into minus infinity, we'll do
  1797. 1:04:17the soft max,
  1798. 1:04:19this part just doesn't help doesn't
  1799. 1:04:21change doesn't um
  1800. 1:04:22like uh doesn't um by this part I mean
  1801. 1:04:24like everything after T, right?
  1802. 1:04:27So everything after little T, you you
  1803. 1:04:29turn them into minus infinity and after
  1804. 1:04:31you do the soft max they become zero.
  1805. 1:04:34So because exponential to the minus
  1806. 1:04:36infinity becomes like so so so so become
  1807. 1:04:38basically zero. So
  1808. 1:04:40that's why
  1809. 1:04:42in this P, so everything after PTT
  1810. 1:04:48everything after here will become zero.
  1811. 1:04:50And the the sum of the rest will be
  1812. 1:04:52what? So that's exactly what you how you
  1813. 1:04:54achieve this because basically the
  1814. 1:04:55linear combination of the rest of the
  1815. 1:04:57terms
  1816. 1:04:58doesn't matter because their
  1817. 1:04:59coefficients are zero.
  1818. 1:05:07Yeah, we are masking this because you
  1819. 1:05:09shouldn't in this auto regressive model
  1820. 1:05:11you don't allow
  1821. 1:05:13the output at time T to depend on
  1822. 1:05:16anything that is after T. So, you have
  1823. 1:05:19to respect this kind of like a
  1824. 1:05:21sequential dependency. So, um, you know,
  1825. 1:05:24in generation, you you are not able to
  1826. 1:05:26see the future tokens to generate the
  1827. 1:05:27current token. So,
  1828. 1:05:29um, so every token can only depends on
  1829. 1:05:32every things before it. So, that's why
  1830. 1:05:34we have to mask it out.
  1831. 1:05:40Okay, we have
  1832. 1:05:42Yeah, we are almost done here. Okay, so
  1833. 1:05:45um
  1834. 1:05:49Okay, so this is the so-called single
  1835. 1:05:51head attention.
  1836. 1:05:52Um, and we have to make it a little more
  1837. 1:05:55complicated to even finish the attention
  1838. 1:05:57part. So, then what we do next is that
  1839. 1:06:02maybe I should erase here.
  1840. 1:06:04Let's just say
  1841. 1:06:13So,
  1842. 1:06:14in some sense, what we are doing here is
  1843. 1:06:16that we have H1 in
  1844. 1:06:19and HT in and uh, we
  1845. 1:06:22go through this single head attention.
  1846. 1:06:24I guess uh, in the note, I call it uh
  1847. 1:06:27attention 1H and you get this H out.
  1848. 1:06:33H1 out
  1849. 1:06:35up to HT out. That's what we defined so
  1850. 1:06:38far.
  1851. 1:06:39And uh,
  1852. 1:06:40in practice, you have multiple copies of
  1853. 1:06:42this single head attention with
  1854. 1:06:44different uh, weight matrices. So, you
  1855. 1:06:47get another copy
  1856. 1:06:49of the single head attention.
  1857. 1:06:51Maybe let's call this
  1858. 1:06:54single head attention one, this is
  1859. 1:06:55called single attention head attention
  1860. 1:06:56two.
  1861. 1:06:57And they have the same
  1862. 1:07:01input
  1863. 1:07:02and then you you get the
  1864. 1:07:04uh, another copy of the output.
  1865. 1:07:08But here, by copy I don't really mean
  1866. 1:07:10literally copy because then it's not
  1867. 1:07:11useful. So, what's the
  1868. 1:07:13the difference between this and this is
  1869. 1:07:15that the the weight uh are the weight
  1870. 1:07:18matrices are different. So, basically
  1871. 1:07:20this WQ
  1872. 1:07:21WK WV
  1873. 1:07:23for these two
  1874. 1:07:25copies of the attention are different.
  1875. 1:07:27So, so So, that's why they're not really
  1876. 1:07:29literally copies. They are just uh uh
  1877. 1:07:31like you have another
  1878. 1:07:32um uh another version of it.
  1879. 1:07:35And uh
  1880. 1:07:36um
  1881. 1:07:42Yeah, so So, in some sense you should
  1882. 1:07:43index this by something else because the
  1883. 1:07:45output will be different, right? So,
  1884. 1:07:46maybe you should index this by
  1885. 1:07:48um
  1886. 1:07:51Let's say this is
  1887. 1:07:53because this is the first attention you
  1888. 1:07:54say H11 and HT1 and this is H12 HT2.
  1889. 1:07:59And you also have another one
  1890. 1:08:01and you get H1
  1891. 1:08:033 up to HT3.
  1892. 1:08:06See? So, So, when you have so many
  1893. 1:08:09uh copies you have a several sequences
  1894. 1:08:11of vectors and but eventually you only
  1895. 1:08:13want one sequence of vectors. So, you
  1896. 1:08:14have to combine them in some way.
  1897. 1:08:16And the way to combine them is that you
  1898. 1:08:18um
  1899. 1:08:20uh you
  1900. 1:08:23concatenate this
  1901. 1:08:25three things. So, you concatenate
  1902. 1:08:30H11 out
  1903. 1:08:32H12 out.
  1904. 1:08:36Just
  1905. 1:08:37one I think that how many Each of these
  1906. 1:08:40is called one head. So, you have let's
  1907. 1:08:42say NH NH head.
  1908. 1:08:45So, this is a attention
  1909. 1:08:48and sub H.
  1910. 1:08:50So, you concatenate all of these vectors
  1911. 1:08:53and you get uh and then you multiply
  1912. 1:08:54that with a another projection matrix
  1913. 1:08:57and then this will give you the final uh
  1914. 1:09:00H uh
  1915. 1:09:01uh the output
  1916. 1:09:03at one.
  1917. 1:09:04So, that's the final output at position
  1918. 1:09:07one. And you do the same thing for
  1919. 1:09:08position two, where you concatenate
  1920. 1:09:11uh
  1921. 1:09:12maybe these guys, these guys.
  1922. 1:09:16And uh and then you do a projection, you
  1923. 1:09:17get the output at position two.
  1924. 1:09:20So, in some sense, this is just give you
  1925. 1:09:22more power to compute different kind of
  1926. 1:09:24like uh
  1927. 1:09:25um
  1928. 1:09:26uh um
  1929. 1:09:27kind of attention. In some sense, you
  1930. 1:09:29know, I think actually in practice, if
  1931. 1:09:31you
  1932. 1:09:32I mean, the intuition is that maybe each
  1933. 1:09:34of these attention is doing
  1934. 1:09:35single-headed attention is trying to pay
  1935. 1:09:38special type of like uh
  1936. 1:09:40uh is doing some special type of job.
  1937. 1:09:42Like, for example, maybe here it's
  1938. 1:09:43trying to find out which entity in the
  1939. 1:09:46past passage matters to me.
  1940. 1:09:49And here, this is trying to find out,
  1941. 1:09:52you know, which
  1942. 1:09:53uh um uh what are the kind of like the
  1943. 1:09:55sentiment
  1944. 1:09:56of the past passages uh I I and and
  1945. 1:09:59which part of the sentiment matters to
  1946. 1:10:00me, so and so forth, right? So, each of
  1947. 1:10:02these attention probably pay attention
  1948. 1:10:03to different type of information. And
  1949. 1:10:05then eventually, you aggregate the
  1950. 1:10:07outcome of all of this information into
  1951. 1:10:09one
  1952. 1:10:10um big vector, and you project it into
  1953. 1:10:13um uh a single output.
  1954. 1:10:16um
  1955. 1:10:19Cool.
  1956. 1:10:22>> The second attention is the output of
  1957. 1:10:24the first one, and therefore
  1958. 1:10:25>> No, no, they are they are all parallel,
  1959. 1:10:27yes.
  1960. 1:10:27>> Parallel how far
  1961. 1:10:28>> They are they are parallel, and then you
  1962. 1:10:30aggregate. So, and all of this is in one
  1963. 1:10:33of this.
  1964. 1:10:34And then, so basically, in one of this,
  1965. 1:10:36you have multiple parallel heads, and
  1966. 1:10:38then next round, you have another number
  1967. 1:10:39of parallel heads.
  1968. 1:10:46I think the number of heads is uh maybe
  1969. 1:10:48on the order of like maybe
  1970. 1:10:50at most a hundred, I guess.
  1971. 1:10:52um
  1972. 1:10:53um let me think about this. I think
  1973. 1:10:57something on that order, I think. Yeah.
  1974. 1:10:59Yeah.
  1975. 1:11:00You know, it also depends on the model
  1976. 1:11:01size, of course, right? I mean, the the
  1977. 1:11:02biggest one
  1978. 1:11:04I don't know. I don't know whether 100
  1979. 1:11:05is is for which model size. I think
  1980. 1:11:07maybe for like a
  1981. 1:11:09100 billion parameter models, you have
  1982. 1:11:10100 heads. Yeah, something like
  1983. 1:11:16Yeah.
  1984. 1:11:23>> [clears throat]
  1985. 1:11:36>> Yeah, I think um I I don't know exactly
  1986. 1:11:38what's your proposal, but um
  1987. 1:11:40um
  1988. 1:11:41maybe it's This sounds like a simple
  1989. 1:11:42simplified answer, but like a like a
  1990. 1:11:43like many of these proposals has been
  1991. 1:11:45considered
  1992. 1:11:46in various ways, and and some of them
  1993. 1:11:48are
  1994. 1:11:49for and most of them are for efficiency
  1995. 1:11:51perspectives, because eventually you can
  1996. 1:11:52just have a
  1997. 1:11:54like as we discussed with like a if you
  1998. 1:11:55really don't care about efficiency, you
  1999. 1:11:56can just have a generic very gigantic
  2000. 1:11:59networks, right? So, all of these kind
  2001. 1:12:01of like
  2002. 1:12:02special structure is to try to
  2003. 1:12:05use your parameters in the right way, so
  2004. 1:12:07that you can
  2005. 1:12:08use it more efficiently in some sense.
  2006. 1:12:13Um
  2007. 1:12:15Okay, so
  2008. 1:12:19And then
  2009. 1:12:20Okay, so I guess uh let's try to Let's
  2010. 1:12:22just give me two or three minutes to
  2011. 1:12:24talk about to con-
  2012. 1:12:26talk about some very quick thing. So,
  2013. 1:12:29one thing is that actually, when I draw
  2014. 1:12:32this, it's actually a little more
  2015. 1:12:33complicated than
  2016. 1:12:35just doing this. So, because there's
  2017. 1:12:37also a residual part in the in this kind
  2018. 1:12:40of like how to combine MLP with
  2019. 1:12:42attention. So, um so actually, what
  2020. 1:12:45actually you really do is that you say
  2021. 1:12:47you apply
  2022. 1:12:49um So,
  2023. 1:12:51after you do the uh attention, you
  2024. 1:12:54um
  2025. 1:12:57you do some kind of like normalization
  2026. 1:13:00here.
  2027. 1:13:01So, there's a normalization layer.
  2028. 1:13:03Often, this is using the RSM norm that
  2029. 1:13:05we briefly discussed in one of the
  2030. 1:13:07neural network section.
  2031. 1:13:08And then you apply MLP, and then you add
  2032. 1:13:12the residuals back to this. You you you
  2033. 1:13:14kind of like in the residual network
  2034. 1:13:16that we discussed.
  2035. 1:13:17And then you go to the next layer. And
  2036. 1:13:19there are different ways to kind of like
  2037. 1:13:20add the residuals. You can add the
  2038. 1:13:21residuals after you do the
  2039. 1:13:23>> [laughter]
  2040. 1:13:24>> attention or or before you do the
  2041. 1:13:25attention. So, there's something called
  2042. 1:13:28pre-norm and the and post-norm
  2043. 1:13:30things. So, I guess the details you you
  2044. 1:13:32can check out the lecture notes. Um, so
  2045. 1:13:34there's some kind of like orders on how
  2046. 1:13:36you when you apply your normalization,
  2047. 1:13:37residuals, and attentions, so and so
  2048. 1:13:38forth.
  2049. 1:13:40And and finally, I think one important
  2050. 1:13:42thing, especially for the next lecture
  2051. 1:13:45about the system ML um
  2052. 1:13:47perspective, is that the we should try
  2053. 1:13:50to quickly think about the computational
  2054. 1:13:53efficiency here
  2055. 1:13:54and the memory footprint.
  2056. 1:13:57So, how many operations we have to do,
  2057. 1:13:59especially as a function of the of T.
  2058. 1:14:02Um, as a function of other things is
  2059. 1:14:04actually relatively easier because as a
  2060. 1:14:05function of the number of parameters,
  2061. 1:14:07you know, like all of those is kind of
  2062. 1:14:08like a it's all multiple kind of like
  2063. 1:14:10linear in some sense. The main important
  2064. 1:14:12thing is what is the the dependency on
  2065. 1:14:14T.
  2066. 1:14:15And you can see that
  2067. 1:14:16for each of the T's position, you have
  2068. 1:14:18to compute all of these inner products.
  2069. 1:14:20So, for each of the position, you have
  2070. 1:14:22to compute capital T inner product.
  2071. 1:14:24So, and and you have capital T of these
  2072. 1:14:27positions.
  2073. 1:14:28So, that means that you have to
  2074. 1:14:30um basically do a T squared dependency,
  2075. 1:14:33right? So, so or in other words, you
  2076. 1:14:35know, even this just this matrix itself
  2077. 1:14:39is a T T by T matrix, capital T by
  2078. 1:14:41capital T matrix. And each of these
  2079. 1:14:43entry requires some inner product. So,
  2080. 1:14:45so I think the dimension, if you still
  2081. 1:14:47remember the Q eyes of the Q T's are in
  2082. 1:14:51R D H and you need to compute T square
  2083. 1:14:54of this inner product. So, basically the
  2084. 1:14:56the number of operations is T square
  2085. 1:14:58times D H.
  2086. 1:15:00The dependence on the H is less
  2087. 1:15:01relevant. The the you know, less
  2088. 1:15:03important to some degree. The dependence
  2089. 1:15:04on T square is very important because
  2090. 1:15:06this means that if you have
  2091. 1:15:08capital T being a million
  2092. 1:15:10it's going to take a million times
  2093. 1:15:11square square operations which is
  2094. 1:15:13prohibitive.
  2095. 1:15:15So, and this is one of the issues with
  2096. 1:15:17the long context issues probably have
  2097. 1:15:19heard of like a like I think we are
  2098. 1:15:21using chat GPT or cloud code you see the
  2099. 1:15:23context and then and then after some
  2100. 1:15:24point they say, "Let me compact my
  2101. 1:15:26context." right? And the reason is that
  2102. 1:15:28you have to shrink the context otherwise
  2103. 1:15:29your computational efficiency is too
  2104. 1:15:31bad. Um uh but of course you know, in
  2105. 1:15:33the in the in the real production system
  2106. 1:15:35it's not like all of this capital T is
  2107. 1:15:37really like there are other ways to
  2108. 1:15:39reduce the dependencies. I think in
  2109. 1:15:41practice I don't think necessarily it's
  2110. 1:15:42T square it's probably more closer to
  2111. 1:15:44linear in T um because there are other
  2112. 1:15:47um ways variants of transformers that
  2113. 1:15:49can reduce the dependency on T. Um I
  2114. 1:15:51think after in the next Wednesday we're
  2115. 1:15:53going to discuss a few um things that
  2116. 1:15:55can reduce the dependency on T a little
  2117. 1:15:57bit. Um exactly how you do it we don't
  2118. 1:16:00really know in some sense because the
  2119. 1:16:03uh there are a lot of open source papers
  2120. 1:16:05that uh propose different approach. Uh
  2121. 1:16:07um with at least I don't know exactly
  2122. 1:16:10which one is used uh
  2123. 1:16:11uh in the in open and a topic. Um um And
  2124. 1:16:15another thing is that what is the memory
  2125. 1:16:17usage here? So, the memory usage you
  2126. 1:16:20know,
  2127. 1:16:20as as at at the first side you have to
  2128. 1:16:22spend T square memory
  2129. 1:16:24because this matrix itself requires T
  2130. 1:16:26square memory.
  2131. 1:16:28But that actually can be reduced because
  2132. 1:16:31you don't have to actually compute the
  2133. 1:16:32whole matrix one by one and then do the
  2134. 1:16:35multiplication with V. So, in some sense
  2135. 1:16:37to some degree you can do some part of
  2136. 1:16:39this matrix and multiply with part of
  2137. 1:16:40the V first and then you don't have to
  2138. 1:16:43save that part of the matrix. You can
  2139. 1:16:44compute a second part. So, and and the
  2140. 1:16:46idea is called flash flash attention,
  2141. 1:16:48probably you have heard of, which is a
  2142. 1:16:50way to reduce the the the memory
  2143. 1:16:54footprint in this systems.
  2144. 1:16:57Um So, yeah, I think that's
  2145. 1:17:00but but you know, like some some of
  2146. 1:17:02these kind of like ideas do have a
  2147. 1:17:04trade-off because when you reduce the
  2148. 1:17:06um for example, if you change your
  2149. 1:17:08architecture to make the dependency on T
  2150. 1:17:10squared better, for example, we'll
  2151. 1:17:12discuss some of this um on Wednesday.
  2152. 1:17:14So, if you change the architecture, the
  2153. 1:17:16dependency on T is better, but your
  2154. 1:17:17architecture is less expressive, and
  2155. 1:17:19then you may lose some performance to
  2156. 1:17:21some degree.
  2157. 1:17:23Okay, sounds good.
  2158. 1:17:24>> Thank you.
  2159. 1:17:24>> Yeah, thanks.

About this transcript

This page contains the full transcript of YouTube transcript (pwQ0l4hFCVI) , generated from the public captions YouTube serves with the video. The transcript has 12,541 words across 2,159 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.