YouTube2Text

Transformers for beginners | What are they and how do they work — Transcript

by AssemblyAI · 3,794 words · 546 segments · language en · Watch on YouTube

Full transcript

  1. 0:00transformers came into our lives just a
  2. 0:02couple of years ago but they have been
  3. 0:04taking the nlp area by storm libraries
  4. 0:06like hugging phase has made it very easy
  5. 0:08for everyone to use transformers or
  6. 0:11implementations like bert or gpt3 is the
  7. 0:14reason that everyone is talking about
  8. 0:15them but what are they and how do they
  9. 0:18work so in this video we will look
  10. 0:20closely into transformers and understand
  11. 0:22their working principles this video is
  12. 0:24part of the deep learning explained
  13. 0:26series by assembly ai which is a company
  14. 0:28that is making a state-of-the-art speech
  15. 0:31to text api if you want to use assembly
  16. 0:33ai for free get your free api token
  17. 0:36using the link in the description
  18. 0:38before transformers word is coming we
  19. 0:40were using rnns to deal with text data
  20. 0:42or any sequence data really but the
  21. 0:44problem with rnns is that when you give
  22. 0:46it a very long sentence it tends to
  23. 0:48forget the beginning of the sentence
  24. 0:50when it comes to the end of the sentence
  25. 0:52and because they rely on recurrence well
  26. 0:54it's in the name recurrent neural
  27. 0:55network they cannot be paralyzed
  28. 0:58then we start using lstms lstms are a
  29. 1:01little bit more sophisticated they tend
  30. 1:02to remember information for a little bit
  31. 1:04longer of a time but they take very long
  32. 1:07to train
  33. 1:08well then we have transformers
  34. 1:10transformers only rely on attention
  35. 1:12mechanisms to remember things they do
  36. 1:14not have any recurrence at all and
  37. 1:16thanks to this they are faster because
  38. 1:18we can parallelize them we can train
  39. 1:20them in a parallel way okay but what is
  40. 1:22this attention we can definitely make
  41. 1:24another video to talk about that and if
  42. 1:26you're interested in that definitely
  43. 1:28comment and let me know but generally
  44. 1:30attention is the ability of a model to
  45. 1:32pay attention to the important part of a
  46. 1:34sentence or an image or any kind of
  47. 1:36input really so if it's a sentence this
  48. 1:40is what it would look like
  49. 1:41so let's say we have a english sentence
  50. 1:44and the sentence is the agreement on the
  51. 1:46european economic area was signed in
  52. 1:48august 1992 and the other side is the
  53. 1:51french translation of that but i do not
  54. 1:53know the first thing about french so i'm
  55. 1:54not even going to try to pronounce it
  56. 1:57but as you can see in this chart what we
  57. 1:59see is that the lighter the color of the
  58. 2:01square the more attention our model is
  59. 2:03paying to the word in that line or in
  60. 2:06that row or column and as you can see it
  61. 2:09does not always go in a diagonal way
  62. 2:12when it is translating european economic
  63. 2:15area because the word order is reversed
  64. 2:19in french it is paying attention in a
  65. 2:21reversed way
  66. 2:23if this was an image and let's say we
  67. 2:25are looking for dogs and images and we
  68. 2:27are trying to um classify different
  69. 2:30breeds of dogs then you can see what
  70. 2:33your model is paying attention to is it
  71. 2:35the noses of the dogs is it the ears of
  72. 2:37the dogs what exactly in an image is the
  73. 2:40model paying attention to to be able to
  74. 2:42understand the difference between dog
  75. 2:44breeds alright now that we briefly
  76. 2:46looked at what attention is let's look
  77. 2:48into how the transformer networks learn
  78. 2:50and how what their architecture is
  79. 2:53so this is what a transformer network
  80. 2:55more or less looks like but we will go
  81. 2:57and start from the higher levels and
  82. 2:59then start breaking down everything and
  83. 3:01understand how they work together so the
  84. 3:04basic thing on a very high level what
  85. 3:06transformers have is an encoder and a
  86. 3:10decoder part
  87. 3:11but actually what they have is six
  88. 3:14encoders and six decoders but basically
  89. 3:16the right left-hand side is the encoders
  90. 3:18and the right-hand side is the decoders
  91. 3:21each encoder has one self-attention
  92. 3:23layer that is paying attention to the
  93. 3:25sentence itself and one feed forward
  94. 3:28forward neural network layer and every
  95. 3:31decoder has two self-attention layers
  96. 3:34and one feed forward neural network
  97. 3:36layer the parallelization comes from how
  98. 3:38we feed the data into this network we
  99. 3:41feed all the words of the sentence at
  100. 3:43the same time to our network
  101. 3:45specifically the encoder inside the
  102. 3:47first step which is the self-attention
  103. 3:49sub layer
  104. 3:50all the words of the sentence is
  105. 3:52compared to all the other words in the
  106. 3:54sentence so there is some communication
  107. 3:56between the words
  108. 3:57whereas in the next step in the
  109. 3:59feed-forward neural network they are
  110. 4:01passed through a feed-forward neural
  111. 4:03network separately so they do not have
  112. 4:05any information exchange but the
  113. 4:07feed-forward neural networks that they
  114. 4:09are passed through are the same inside
  115. 4:11the same layer but as we said there are
  116. 4:13six encoders and each each in each of
  117. 4:15these six encoders the neural networks
  118. 4:17are different okay so this has been kind
  119. 4:19of the middle part of the network we
  120. 4:21also have the inputs and the outputs so
  121. 4:24all the inputs that go in either the
  122. 4:26encoder or the decoder the raw inputs
  123. 4:28are embedded
  124. 4:30what are embeddings well that's a little
  125. 4:32bit of a longer topic for this video but
  126. 4:34again if you like us to make a video on
  127. 4:36this leave a comment uh but what you
  128. 4:38need to know for now is that embeddings
  129. 4:40are a way to represent these words in a
  130. 4:44n length vector in this specific
  131. 4:47transformer architecture they are using
  132. 4:49512 length vectors and that's basically
  133. 4:52what they use in the original paper but
  134. 4:54this is a hyper parameter that you can
  135. 4:56change and on top of these word
  136. 4:58embeddings we are adding positional
  137. 5:00encodings so if you remember we said
  138. 5:02transformers do not have any recurrence
  139. 5:05so the model has no way of understanding
  140. 5:07which word comes first and the other one
  141. 5:09comes second or which word comes where
  142. 5:11in the sentence so by adding a
  143. 5:13positional encoding you are letting or
  144. 5:15you are adding some information or
  145. 5:17injecting some information with each
  146. 5:18word that tells the modal where this
  147. 5:21word in the sentence comes in and lastly
  148. 5:24for the output as you can see we have a
  149. 5:26linear layer and a softmax layer at the
  150. 5:29end of the decoders so the output of the
  151. 5:31decoders can be transformed into
  152. 5:33something that we can understand and
  153. 5:35basically what they turn into is a
  154. 5:37vector that has the length of the amount
  155. 5:39of words that we have in our vocabulary
  156. 5:41and each of these cells tells us how
  157. 5:43likely it is that this word in this cell
  158. 5:47is going through the next word in our
  159. 5:48sequence and those are the main
  160. 5:50components but there are two little
  161. 5:52things that make transformers a little
  162. 5:54bit better one of them is the
  163. 5:55normalization layers so if you realize
  164. 5:58in between the sub layers that is the
  165. 6:01self-attention layers and the
  166. 6:02feed-forward neural networks we have
  167. 6:05some add and normalized layers and what
  168. 6:07they do is to normalize the output that
  169. 6:10comes from the sub layer the
  170. 6:12normalization technique that is used
  171. 6:13there is called layer normalization and
  172. 6:15that is basically an improvement over
  173. 6:18batch normalization and if you don't
  174. 6:19know what batch normalization is we
  175. 6:21already made a video about that i will
  176. 6:23link it somewhere here and you can go
  177. 6:25watch that to understand a little bit
  178. 6:26better what batch normalization or layer
  179. 6:28normalization is and the second little
  180. 6:30detail is the skip state so if you look
  181. 6:32at the original architecture image we
  182. 6:34see that there are some arrows that are
  183. 6:36going around some of the sub-layers well
  184. 6:38actually all of the sub-layers so some
  185. 6:40of the information that does not go into
  186. 6:43either the self-attention layer
  187. 6:45sub-layers or the feed-forward neural
  188. 6:47networks are sent directly to the
  189. 6:49normalization layer this kind of helps
  190. 6:52the model not forget things and it helps
  191. 6:54the model to forward information that is
  192. 6:57important to further in the network
  193. 6:59inside these normalization layers what
  194. 7:01we do is add the information that went
  195. 7:04through the sub layer and also just skip
  196. 7:06the sub layer and then normalize them
  197. 7:08together and that's all there is to the
  198. 7:10architecture of transformers and if you
  199. 7:12look into it actually most of the things
  200. 7:15that are inside this architecture are
  201. 7:17things that we've known from before
  202. 7:19things have been around for a long time
  203. 7:21for example
  204. 7:22linear transformations soft max layers
  205. 7:24or
  206. 7:25word embeddings for example or the feed
  207. 7:27forward neural networks but there are
  208. 7:30two really noble ideas inside the
  209. 7:32original transformer paper that really
  210. 7:36made the difference for transformers and
  211. 7:37those were positional encodings and
  212. 7:40multi-headed attention so let's take a
  213. 7:42closer look at how they work let's start
  214. 7:44with multi-headed attention layers so if
  215. 7:46you look at the original architecture
  216. 7:48you see that there are two different
  217. 7:49types of multiheaded attention layers
  218. 7:52one of them is just multi-headed
  219. 7:53attention and the other one is masked
  220. 7:55multi-headed attention well it actually
  221. 7:57does the same thing no matter if it's
  222. 7:59called masked or not if it's in an
  223. 8:01encoder and a decoder and the only
  224. 8:03difference is in a normal multiheaded
  225. 8:05attention layer all the words are
  226. 8:07compared with all the other words that
  227. 8:09are inputted that are in a sentence but
  228. 8:12that will make more sense to you in a
  229. 8:13second what i mean by comparing and for
  230. 8:16masked multi-headed attention layers
  231. 8:18only the words that are coming before a
  232. 8:20word are compared to that word in the
  233. 8:22sentence in the attention layer
  234. 8:24something called the scaled dot product
  235. 8:26attention is used and then it is
  236. 8:28multiplied and is done multiple times to
  237. 8:31create that multi-headed effect and of
  238. 8:33course everything is done in mattresses
  239. 8:35to make things faster but i will show
  240. 8:37you how attention is calculated using
  241. 8:39just the vectors of words what do we
  242. 8:41have in the beginning are embeddings of
  243. 8:43words if you remember we embedded the
  244. 8:45words into a vector and also we added
  245. 8:48positional encodings and then this is
  246. 8:50fed to the first encoder and of course
  247. 8:52at first is fed to the multiheaded
  248. 8:54attention sub layer of the first encoder
  249. 8:57so in there what is done first is to
  250. 9:00multiply these embedding vectors with
  251. 9:03some mattresses these mattresses are
  252. 9:05called query key and value mattresses
  253. 9:08and these are values that are
  254. 9:10initialized randomly and are trained
  255. 9:12during the training process to be
  256. 9:13learned kind of like the weights and
  257. 9:15biases we have in neural networks as a
  258. 9:17result of this multiplication we get the
  259. 9:19query key and value vector for each word
  260. 9:23and from this point on we are going to
  261. 9:25use these vectors to
  262. 9:27keep going with the calculation the
  263. 9:29first thing that we want to do is to
  264. 9:30calculate a score for each word against
  265. 9:33all the other words in the sentence what
  266. 9:35we do for this is we dot take the dot
  267. 9:38product of the curie vector of each word
  268. 9:41against the key vector of all the other
  269. 9:44words so if you want to get the score of
  270. 9:46the first word on the first word what we
  271. 9:48do is we get dot product of the query
  272. 9:51vector of word one with the key vector
  273. 9:54of word one if we want to get the score
  274. 9:57of the first word against the second
  275. 9:59word we must get the dot product of the
  276. 10:01query vector of the first word with the
  277. 10:04key vector of the second word so when we
  278. 10:06get the dot product of all the key
  279. 10:09vectors of all the other words compared
  280. 10:12with or combined with the query vector
  281. 10:14of the first word we have the score of
  282. 10:16the first word against all the other
  283. 10:19words so they all belong all of these
  284. 10:21scores belong to the first word if you
  285. 10:23want to get the scores for the second
  286. 10:25word we're going to have to multiply its
  287. 10:27query vector with all the other words
  288. 10:30key vectors and this is what it's done
  289. 10:32and it's all done in parallel so that's
  290. 10:34why we do not have any recurrence or we
  291. 10:36don't have to wait for other words to be
  292. 10:38done before starting to process the
  293. 10:40words further in the sentence we can do
  294. 10:42these calculations for all the words at
  295. 10:44the same time once we have all the
  296. 10:46scores of all the words against all the
  297. 10:48other words what we're going to do is to
  298. 10:50divide them by eight so that might sound
  299. 10:53like a very random number for you but
  300. 10:55it's basically the square root of 64
  301. 10:58which is again sound like a random
  302. 11:00number that comes out of nowhere but
  303. 11:02actually 64 is the square root of the
  304. 11:04length of the query key and value
  305. 11:07vectors and that's why the authors of
  306. 11:09the original transformers paper are
  307. 11:11using that number after we divide
  308. 11:13everything by 8 we pass all these values
  309. 11:15to a softmax layer we do this to
  310. 11:17normalize all these values and the score
  311. 11:20values of one word against all the other
  312. 11:22words are now going to sum up to one the
  313. 11:24resulting number serves kind of like a
  314. 11:26weight from this point on what we're
  315. 11:28going to do is to multiply all the value
  316. 11:30vectors of all the words with this
  317. 11:32weight and finally you sum up all the
  318. 11:35weighted value vectors of all the words
  319. 11:38based on this one word that we were
  320. 11:40doing the calculations for and create
  321. 11:42the output of the self-attention layer
  322. 11:44for this one word and then you have to
  323. 11:46do all of these calculations for all the
  324. 11:48other words well this is done
  325. 11:50simultaneously but at the end you have
  326. 11:52the output of the attention layer
  327. 11:55as i mentioned multi-headed attention
  328. 11:57does this eight times so effectively it
  329. 11:59is training eight different query key
  330. 12:02and value mattresses not the vectors for
  331. 12:05the words but the mattresses that we
  332. 12:07multiply the
  333. 12:09input embeddings with this way the model
  334. 12:11is able to pay attention to not one
  335. 12:13other word but many other words in the
  336. 12:16sentence so in the model they're using
  337. 12:18the number eight but you can basically
  338. 12:20change it if you like to and let's look
  339. 12:22at this example again that we gave at
  340. 12:24the beginning of this video as you can
  341. 12:26see there are some of the um cells that
  342. 12:29are really bright but at the same time
  343. 12:31we also have some cells that are just
  344. 12:32kind of gray and that means that our
  345. 12:35model was also paying attention to those
  346. 12:37other words other than the primary word
  347. 12:39that they're paying attention to a
  348. 12:40little bit and this multiheaded
  349. 12:42attention thing is also one of the
  350. 12:44reasons why transformers are so
  351. 12:46seamlessly able to deal with
  352. 12:48sentences of different lengths one thing
  353. 12:51you might catch here is that if we do
  354. 12:53the same thing eight times what we're
  355. 12:55going to have is eight different
  356. 12:56resulting mattresses right and inside
  357. 12:59these mattresses one line is going to
  358. 13:00correspond to one word but we're going
  359. 13:02to have eight of them so how are we
  360. 13:04going to deal with this well what they
  361. 13:06do in the paper or what they propose to
  362. 13:08do is basically concatenate them all
  363. 13:10together and then multiply them with yet
  364. 13:13another weight matrix that is going to
  365. 13:15produce a matrix that is going to look
  366. 13:17like only one output of the attention
  367. 13:20layer this weight matrix is of course
  368. 13:22yet another thing to train inside the
  369. 13:24transformers on top of the key value and
  370. 13:27curing mattresses that we multiply all
  371. 13:30of the word embeddings with next is
  372. 13:32positional encoding so as i mentioned
  373. 13:34before position line codings is a way to
  374. 13:36inject or add information to the word
  375. 13:39embeddings that we've created before to
  376. 13:42show where in a sentence a word is so
  377. 13:45basically the location information of a
  378. 13:47word you can either use learned
  379. 13:49positional encodings or fixed positional
  380. 13:51encodings but in the paper in the
  381. 13:53original paper they suggest or they
  382. 13:55recommend that we use fixed positional
  383. 13:57encodings because they have the
  384. 13:59advantage of being able to handle
  385. 14:02lengths of sentences that we haven't
  386. 14:04seen in the training set
  387. 14:06you might say why do we need any
  388. 14:08sophisticated solution for this anyways
  389. 14:10why can't we just assign a number to the
  390. 14:12word specifying where in the sentence
  391. 14:15this word is so that wouldn't really
  392. 14:17work because let's say if you assign a
  393. 14:19number that goes from zero to one what's
  394. 14:21going to happen is that you're not going
  395. 14:23to really understand how many words are
  396. 14:25in that sentence just by looking at this
  397. 14:28one word and this value will not be
  398. 14:29consistent in between examples another
  399. 14:32solution could be to assign integers to
  400. 14:34words of course starting from one or
  401. 14:36zero to however long the sentence is but
  402. 14:39the problem with that one is that those
  403. 14:41numbers can get very high right if you
  404. 14:43have a very long sentence that could get
  405. 14:45out of control and on top of that there
  406. 14:47could be sentences with specific lengths
  407. 14:50that you do not have in the training
  408. 14:51data and that could cause some problems
  409. 14:53in terms of generalization so what they
  410. 14:56did as a solution to this positional
  411. 14:57problem in the original transformers
  412. 14:59paper was to use sine and cosine
  413. 15:01functions in different frequencies
  414. 15:04but of course i don't expect you to know
  415. 15:06this so let's look into how that looks
  416. 15:08so this is what sine and cosine
  417. 15:10functions in different frequencies look
  418. 15:12like the colors here show us numbers
  419. 15:14that range from -1 to 1. the x-axis
  420. 15:18shows us the length of the word
  421. 15:19embeddings in transformers we are using
  422. 15:21512 as i mentioned before and the y-axis
  423. 15:25is the position of this token of this
  424. 15:27word so if i want to get the positional
  425. 15:29encoding of a word that is in let's say
  426. 15:32the 20th position i need to get the
  427. 15:34horizontal line that corresponds to 20
  428. 15:38in the y-axis and the nice thing about
  429. 15:41this positional encoding is that it's
  430. 15:43going to be unique no other place no
  431. 15:45other horizontal line in this graph has
  432. 15:48the same composition of values as in
  433. 15:50that line and one other nice thing about
  434. 15:52this positional encodings is that you
  435. 15:54can always tell the difference between
  436. 15:56two words looking at these positional
  437. 15:57encodings it's always going to be the
  438. 16:00same one thing that really helped me
  439. 16:01understand this concept was to look at
  440. 16:03binary representations of integers so
  441. 16:06let's look at these examples
  442. 16:08if you realize as you increase your
  443. 16:10numbers what happens is the smaller
  444. 16:12digit in the binary representation
  445. 16:15changes from one to zero with every new
  446. 16:17integer whereas the second digit changes
  447. 16:20every two integers so at first it is
  448. 16:22zero and zero and in the second two
  449. 16:25integers it is one and one and the third
  450. 16:27two integers it is zero and zero again
  451. 16:30and again this pattern kind of follows
  452. 16:32itself and what happens is all of these
  453. 16:35binary representations are unique no two
  454. 16:37binary representations are the same and
  455. 16:39on top of that you can always tell the
  456. 16:41difference between two integers by
  457. 16:43looking at their binary representations
  458. 16:45this could also be a perfectly useful
  459. 16:47positional encoding for us too but it is
  460. 16:50only ones and zeros and we are not
  461. 16:52actually using the information that can
  462. 16:54be provided with continuous values so
  463. 16:56that's why instead we use sine and
  464. 16:57cosine functions okay what do we do once
  465. 17:00we have these encodings right let's say
  466. 17:02we have this encoding of 512
  467. 17:05values that we extracted from this graph
  468. 17:07well what we do is we basically add them
  469. 17:10together we add the word embedding and
  470. 17:12the positional encoding together and
  471. 17:14then we feed it to the encoders
  472. 17:16all right so we learned everything that
  473. 17:18we need about the architecture there are
  474. 17:20encoders specifically six of them and
  475. 17:22there are recorders again six of them we
  476. 17:25have the uh last processing and the
  477. 17:27output we have the embeddings at the
  478. 17:29input and also the positional encodings
  479. 17:32but how does it all work so basically to
  480. 17:34bring it together what happens is you
  481. 17:36first get your inputs
  482. 17:38run them through the embeddings run them
  483. 17:40through the positional encodings and
  484. 17:42then run them through six levels of
  485. 17:44encoders and then you get an output
  486. 17:46this output is fed to all of the
  487. 17:48decoders so we have six decoders as we
  488. 17:51mentioned six layers of decoders this
  489. 17:54the information from the output from the
  490. 17:56encoder is fed to all of the decoders
  491. 17:59but this information is only fed to the
  492. 18:01second sub-layer so the multi-headed
  493. 18:03attention sub-layer of the coders
  494. 18:05and the first masked multi-headed
  495. 18:08attention a sub-layer of decoders get
  496. 18:10the input from what was outputted from
  497. 18:13the decoder section of the model in the
  498. 18:16previous time step that way the decoders
  499. 18:18are taking into consideration what was
  500. 18:21the word in the previous time step on
  501. 18:24the previous position and also the
  502. 18:26context that they learned from the
  503. 18:28encoding process of the network to
  504. 18:31create the output these decoders all
  505. 18:33work together and then they create a
  506. 18:36output vector this output vector is sent
  507. 18:38to through the linear transformation
  508. 18:40that creates a logic's vector this
  509. 18:42logic's vector is as long as the amount
  510. 18:45of words that we have in our vocabulary
  511. 18:47and it has the possibilities of how
  512. 18:50likely is the next word going to be one
  513. 18:53word or the other
  514. 18:54and then we pass this through soft max
  515. 18:56to be able to get the probabilities of
  516. 18:59each word and then these probabilities
  517. 19:01will also add up to one so basically a
  518. 19:03normalized version of the logit's vector
  519. 19:06the output of the softmax layer
  520. 19:08basically tells us what the next word is
  521. 19:10going to be and that's all there is to
  522. 19:12know about transformers really it's
  523. 19:14quite simple even though it looks a bit
  524. 19:16complicated at first all you need to
  525. 19:18know that there are encoders and
  526. 19:19decoders and the two noble ideas that
  527. 19:22came into our lives with transformers
  528. 19:24are the positional encodings and the
  529. 19:26multiheaded attention layer to fully
  530. 19:28understand transformers and how they
  531. 19:30work you might need to watch this video
  532. 19:31multiple times and maybe even support
  533. 19:33your learning with some of the written
  534. 19:35resources that are out there so for that
  535. 19:37i've left links to my favorite resources
  536. 19:39in the description if there was anything
  537. 19:41that was not clear or if you have a
  538. 19:42question leave a comment and let me know
  539. 19:45if you like this video don't forget to
  540. 19:46give us a thumbs up and maybe even
  541. 19:48subscribe to be one of the first people
  542. 19:50to know when we make a new video but
  543. 19:52before you leave don't forget to grab
  544. 19:53your free token for assembly ai's
  545. 19:55special text api i'll see you in the
  546. 19:57next video

About this transcript

This page contains the full transcript of Transformers for beginners | What are they and how do they work by AssemblyAI, generated from the public captions YouTube serves with the video. The transcript has 3,794 words across 546 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.