YouTube2Text

Transformer Architecture Explained — Transcript

by Under The Hood · 3,388 words · 531 segments · language en · Watch on YouTube

Full transcript

  1. 0:02Welcome back to another video. Today in
  2. 0:05this video I will explain the
  3. 0:07transformer architecture. It was
  4. 0:10introduced in the paper attention is all
  5. 0:12you need and was published in the year
  6. 0:142017.
  7. 0:15Since then the transformer has become
  8. 0:17the foundation for many modern AI models
  9. 0:20and even today's most advanced AI
  10. 0:22systems. It has completely changed how
  11. 0:25we approach natural language processing.
  12. 0:28In this video, we'll start by
  13. 0:29understanding the basic components of
  14. 0:31the transformer like token embeddings,
  15. 0:33attention mechanisms, feed forward
  16. 0:35networks, residual connections, and
  17. 0:37more. Then we'll move on to how the
  18. 0:39model is trained and how it works during
  19. 0:41inference. To get a better idea of how
  20. 0:44it works, we'll walk through a simple
  21. 0:46example, translating a sentence from one
  22. 0:48language to another. The transformer
  23. 0:51architecture has two main parts, the
  24. 0:53encoder and the decoder, and it follows
  25. 0:55a sequencetosequence design. Just like
  26. 0:57earlier sequencetose sequence models
  27. 0:59that used LSDM,
  28. 1:01but the transformer improves on them by
  29. 1:03enabling parallel processing, better
  30. 1:05handling of long range dependencies, and
  31. 1:08overall faster and more accurate
  32. 1:10results. Let's start with the first
  33. 1:12step, data set preparation for our
  34. 1:14translation task. We will use this data
  35. 1:17set to train our model. We start by
  36. 1:20formatting our data set with special
  37. 1:21tokens, start and end of sequence
  38. 1:24tokens. These are the special tokens
  39. 1:26which we use to identify the start and
  40. 1:28end of a sentence. The source language
  41. 1:31becomes input to the encoder and the
  42. 1:33target language becomes input to the
  43. 1:35decoder. We also create training labels
  44. 1:38by shifting the target sentence one step
  45. 1:40ahead of the decoder's input and we also
  46. 1:42have that extra end of sequence token. I
  47. 1:44will discuss about this format later.
  48. 1:46But now let's say we have our data set
  49. 1:48ready. The overall flow looks like this.
  50. 1:52The source sentence is fed into the
  51. 1:54encoder which generates a context
  52. 1:56vector. The decoder takes in the target
  53. 1:58sentence input and outputs a probability
  54. 2:01distribution. The model is then trained
  55. 2:04by comparing this output to the target
  56. 2:06label. Now that we have the big picture,
  57. 2:08let's dive deeper and explore each
  58. 2:10component of the transformer step by
  59. 2:12step.
  60. 2:16Let's start with the encoder part of the
  61. 2:18transformer. We begin with a data set
  62. 2:20that contains source and target language
  63. 2:22pairs already formatted with special
  64. 2:25tokens like start of sequence and end of
  65. 2:27sequence tokens. The source language
  66. 2:29becomes the input to the encoder. The
  67. 2:32first step is tokenization.
  68. 2:34Here we convert the input text into
  69. 2:36numbers by breaking each sentence into
  70. 2:38smaller units like words or subwords.
  71. 2:41There are different types of
  72. 2:42tokenization methods such as word level
  73. 2:45or subword level. But for simplicity,
  74. 2:47let's assume each word is treated as a
  75. 2:49unique token and each token is assigned
  76. 2:51a unique number. After tokenizing the
  77. 2:54input sentence, we get a list of token
  78. 2:56numbers. In practice, we don't train
  79. 2:59with just one sequence at a time.
  80. 3:01Instead, we use batches of sequences.
  81. 3:04Since sequences can have different
  82. 3:06lengths, we need to make them all the
  83. 3:08same length by filling shorter ones with
  84. 3:10a special token called the padding
  85. 3:11token. This allows us to train in
  86. 3:14batches efficiently. But for now, we'll
  87. 3:16keep things simple and use a single
  88. 3:18example without padding. Next, these
  89. 3:21tokens go into the embedding layer. This
  90. 3:24is a trainable neural network layer that
  91. 3:26converts each token number into a
  92. 3:28highdimension vector called the
  93. 3:30embedding of the token. The token number
  94. 3:32by itself has no real meaning. But the
  95. 3:35embedding vector helps capture the
  96. 3:37semantic meaning of the word. Each token
  97. 3:39in our sequence gets its own embedding
  98. 3:41vector. In the original paper, the size
  99. 3:44of this vector is 512.
  100. 3:46These vectors exist in a highdimensional
  101. 3:48space where words with similar meanings
  102. 3:51tend to be closer together in the
  103. 3:52embedding space. Imagine a 2D version of
  104. 3:55this space. Words like bank, river, and
  105. 3:58flood appear closer together because
  106. 4:00they are often used in similar contexts.
  107. 4:02This helps the model understand how
  108. 4:04words relate to each other in meaning.
  109. 4:07Now that we have word representations,
  110. 4:09we also need to know the order of the
  111. 4:10words in the sentence like which word
  112. 4:12comes first, second and so on. The
  113. 4:15embedding vector alone doesn't contain
  114. 4:17position information. So we need to add
  115. 4:20that separately using positional
  116. 4:21encoding. In the paper, the authors
  117. 4:24added another vector called the
  118. 4:26positional encoding vector to each
  119. 4:28embedding vector. This new vector has
  120. 4:30the same size of 512 dimension and is
  121. 4:33not trainable. It is generated using s
  122. 4:36and cosine functions. Here what it does
  123. 4:39for each token position. It computes a
  124. 4:42vector using a formula. In this vector,
  125. 4:44the even positions use the sign function
  126. 4:47and the odd positions use the cosine
  127. 4:49function and are computed using the
  128. 4:51formula where it takes the position of
  129. 4:53each token in the sequence and also the
  130. 4:55position of the dimension in the vector.
  131. 4:57Like this, we compute the new vector for
  132. 4:59all other tokens as well. This helps the
  133. 5:02model understand the position of each
  134. 5:04word in the sequence. These values are
  135. 5:07fixed. They're computed once and stored
  136. 5:09in memory, then reused during training.
  137. 5:12Even though more recent models use
  138. 5:13trainable position encodings, this
  139. 5:15static method still performs quite well
  140. 5:18and was used by the authors in the
  141. 5:19original transformer paper. So now we
  142. 5:22have our input sentence represented with
  143. 5:24the token embeddings that give the
  144. 5:26meaning of words and positional
  145. 5:28encodings that provide the word order.
  146. 5:31This combined input is then passed into
  147. 5:34the first layer of the encoder block.
  148. 5:36The first step in the encoder is the
  149. 5:38multi head attention layer. But before
  150. 5:41we dive into how multi head attention
  151. 5:43works, let's first understand why we
  152. 5:45need attention, what it does, and how
  153. 5:47it's calculated.
  154. 5:54What we have now is a token represented
  155. 5:56in this highdimensional embedding space
  156. 5:58which contains the semantic meaning of
  157. 6:00the words. But here the problem same
  158. 6:03words can have different meanings based
  159. 6:04on the context. But what the embedding
  160. 6:07does is only capture the semantic
  161. 6:09meaning and does not consider the
  162. 6:11context information. If we look at our
  163. 6:13example then the word bank should be
  164. 6:15close to other words related to
  165. 6:17financial institutions rather than the
  166. 6:19river bank. We know this by looking at
  167. 6:21its context words deposited and money.
  168. 6:24So what we want is to transform this
  169. 6:26representation so that it correctly
  170. 6:28represents the token by looking at its
  171. 6:30context. And for this we have the
  172. 6:32attention mechanism. It transforms the
  173. 6:35static embedding which only captures the
  174. 6:37semantic meaning into the representation
  175. 6:40that is aligned with the context. The
  176. 6:42size remains the same just the vectors
  177. 6:45are now represented according to the
  178. 6:46context. In the paper they call this
  179. 6:49self attention and it is represented by
  180. 6:51this equation. Let's break it down to
  181. 6:53understand what it means. First let's
  182. 6:55see how we get this query key and value
  183. 6:58that you see in the attention
  184. 6:59calculation. We take our static
  185. 7:01embeddings and linearly transform this
  186. 7:03embedding with weight matrices for query
  187. 7:05key and value. What this means is take
  188. 7:08our embedding as input and use this in
  189. 7:11three different neural networks with
  190. 7:12linear activations. What this gives is a
  191. 7:15query key and value which we use in our
  192. 7:17attention calculation. These weights are
  193. 7:20trainable layers and so the query key
  194. 7:22and value vectors are adjusted after
  195. 7:24training. Now that we have our query key
  196. 7:27and value matrices, what we do is take
  197. 7:29the query and transpose of the key
  198. 7:31matrix and then perform matrix
  199. 7:33multiplication.
  200. 7:34What this does is allow all the tokens
  201. 7:36to communicate with all other tokens in
  202. 7:38parallel by taking the dotproduct
  203. 7:41between the first query vector and the
  204. 7:43key vector. It gives a score on how much
  205. 7:45it scores to itself. The query acts like
  206. 7:48asking all other tokens and giving a
  207. 7:50score to each of them on how much they
  208. 7:51match. Similarly, it does this with all
  209. 7:54other tokens using the key matrix. By
  210. 7:57doing this, the first token in our
  211. 7:59sequence attends to all other tokens
  212. 8:01including itself and has a score that
  213. 8:03tells how much it matches with other
  214. 8:05tokens. Now, similarly, the second token
  215. 8:07in our sequence does the same by taking
  216. 8:10its query vector and asking all other
  217. 8:12tokens using the key vector. Like this,
  218. 8:15all tokens attend to each other using
  219. 8:16the query and key in order to know the
  220. 8:19importance of other tokens in the
  221. 8:20context to make a new representation.
  222. 8:23This is done in terms of matrix
  223. 8:25multiplication all at once and the
  224. 8:27result we get is what we call the
  225. 8:29attention score. Since the attention
  226. 8:31scores are just scalar values, they can
  227. 8:34be a varying range and are not
  228. 8:35interpretable. So we first scale these
  229. 8:38scores by dividing by the square root of
  230. 8:39the dimension of the key vector. We
  231. 8:42apply a softmax function to this scaled
  232. 8:44attention score across each row which
  233. 8:46gives us more interpretable values. So
  234. 8:49now each row sums to one and we can
  235. 8:51interpret the attention scores in terms
  236. 8:53of probability values. We call this
  237. 8:55attention weights. If we look at our
  238. 8:57token money and look at its attention
  239. 8:59weight, then we can interpret it as it
  240. 9:01is giving 25% importance to the word
  241. 9:03deposited, 30% to money and 30% to the
  242. 9:07token bank. Like this in the attention
  243. 9:09weight, we now know which token is
  244. 9:11important in creating a new
  245. 9:13representation.
  246. 9:15Using this attention weight and
  247. 9:16multiplying with the value matrix, we
  248. 9:18get the new representation that has the
  249. 9:20context information. So for the word
  250. 9:23money, it uses that weight value to act
  251. 9:25as a weighted sum to construct a new
  252. 9:27embedding which has context
  253. 9:29representation.
  254. 9:30Like this, a new representation based on
  255. 9:33this attention weight is calculated. And
  256. 9:36this is what the attention mechanism is
  257. 9:37about. Transforming the static embedding
  258. 9:40into an embedding that has the context
  259. 9:42representation.
  260. 9:43So the summary is take the static
  261. 9:46embedding with position encoding and
  262. 9:49then transform it in three different
  263. 9:50matrices called query key and value
  264. 9:53matrices using their respective weight
  265. 9:55matrices. Then multiply the query and
  266. 9:58key scale it and apply softmax to get
  267. 10:00the attention weights and multiply this
  268. 10:02with the value matrix to get the final
  269. 10:04contextual embeddings.
  270. 10:09Now let's see what multi head attention
  271. 10:11is in our attention calculation. What we
  272. 10:14did was take the embedding input and
  273. 10:16transform it into query key and value
  274. 10:19which had the dimensions equal to the
  275. 10:20original embedding and then calculate
  276. 10:23the new embedding all at once. But here
  277. 10:25in multi head attention instead of doing
  278. 10:27it all at once we take multiple
  279. 10:29attention layers. Each attention now has
  280. 10:32its own weight matrices for query key
  281. 10:34and value but in a lower dimension. All
  282. 10:37the computations are the same but we get
  283. 10:39the final output in a lower dimension.
  284. 10:41This is what a single attention head
  285. 10:43does. We use multiple identical
  286. 10:46attention heads with different weight
  287. 10:47matrices in a multi head attention. Each
  288. 10:50attention contributes to producing the
  289. 10:52final contextual representation.
  290. 10:55Since each attention head has output in
  291. 10:57a lower dimension, we concatenate the
  292. 10:59output of each attention head. So if we
  293. 11:02want our final embedding to be of size
  294. 11:04512 and want to take eight attention
  295. 11:06heads then we should make the inner
  296. 11:08dimension of each head to be 64 then we
  297. 11:11again transform this using another
  298. 11:13weight matrix W out to get the final
  299. 11:16representation.
  300. 11:17This is what happens in multi head
  301. 11:19attention. It takes in the static
  302. 11:21embedding and gives us new contextaware
  303. 11:23embeddings using multiple self attention
  304. 11:25heads. This completes the multi head
  305. 11:28attention layer. Now after that we have
  306. 11:30the add and normalization layer. There
  307. 11:33is a residual stream that adds the
  308. 11:35static embedding to the result of the
  309. 11:37multi head attention. This is common in
  310. 11:40deep learning to prevent vanishing
  311. 11:42gradients. After the residual addition
  312. 11:44we have the layer normalization.
  313. 11:47Normally in batch normalization when we
  314. 11:49have input sequence in batches we
  315. 11:52normalize across each feature dimension
  316. 11:54across the batch. But in layer
  317. 11:56normalization which is used in this
  318. 11:58architecture, we normalize across each
  319. 12:01individual token by calculating the mean
  320. 12:03and variance across the token's
  321. 12:05features. After the normalization layer,
  322. 12:07we have our normalized output with the
  323. 12:10size remaining the same due to just the
  324. 12:11addition and normalization steps. So up
  325. 12:14until now, the overall flow looks like
  326. 12:16this.
  327. 12:22Now next we have the feed forward neural
  328. 12:24network. After the first layer norm,
  329. 12:27this becomes the input to the feed
  330. 12:28forward neural network. This is a dense
  331. 12:31neural network with nonlinearity in it.
  332. 12:34The output size is the same as the input
  333. 12:36size due to the equal number of neurons
  334. 12:38in the input and output layers. There is
  335. 12:41another add and layer norm like before
  336. 12:43but for the feed forward neural network.
  337. 12:46After this what we have is the output of
  338. 12:48the encoder block. Now this output
  339. 12:51becomes the input to another identical
  340. 12:53encoder block repeating multiple times
  341. 12:55with different parameters further
  342. 12:57refining the input sequence. This is
  343. 13:00what the encoder block looks like where
  344. 13:02the final output is now the context for
  345. 13:04the decoder. So up until now we give
  346. 13:07input to the encoder repeat it multiple
  347. 13:09times and now the output will be the
  348. 13:11representation of the input sequence
  349. 13:13that the decoder will use.
  350. 13:20Let's move on to the decoder part and
  351. 13:22see what it does. We already have our
  352. 13:25data set formatted for training. The
  353. 13:27source language is the encoder input
  354. 13:30which we used in the encoder and now we
  355. 13:32use the target language which becomes
  356. 13:34the input to the decoder. Also the
  357. 13:36target language shifted ahead will be
  358. 13:38our training label. The process is
  359. 13:41similar to that in the encoder. First
  360. 13:43the tokenization step uses the target
  361. 13:46language specific tokenizer to tokenize
  362. 13:48the input. Again for simplicity we are
  363. 13:51just using word level tokenization.
  364. 13:54Then we have the embedding layer with
  365. 13:56positional encoding as before. Now this
  366. 13:59goes into the decoder block where the
  367. 14:01first layer is the masked multi head
  368. 14:03attention layer. This is almost similar
  369. 14:05to the multi head attention layer, but
  370. 14:08with just a slight modification.
  371. 14:10As in the encoder's multi head attention
  372. 14:12layer, we compute the attention score by
  373. 14:14multiplying the query and key matrices
  374. 14:17that we get by transforming our input
  375. 14:19using three different weight matrices.
  376. 14:21This gives us the attention score for
  377. 14:23our target language where each token
  378. 14:25attends to other tokens. If you look at
  379. 14:28our data set, the decoder input is the
  380. 14:30same as the training label, just one
  381. 14:32token shifted ahead. So during training
  382. 14:35each token in the decoder tries to
  383. 14:36predict the next token in the sequence.
  384. 14:39You can see the first token attends to
  385. 14:41its target token during the attention
  386. 14:43calculation.
  387. 14:44Similarly all the tokens are attending
  388. 14:46to their target token and future tokens
  389. 14:48during the attention calculation which
  390. 14:51we don't want. We want to prevent the
  391. 14:53token from attending to its target and
  392. 14:55future tokens. To do this, we take a new
  393. 14:58matrix which we call the masking matrix
  394. 15:00and it contains negative infinity values
  395. 15:03in the positions that we want to mask.
  396. 15:05We add this matrix to our original
  397. 15:07attention matrix to get the masked
  398. 15:08attention score. What this does is
  399. 15:11during attention calculation, the
  400. 15:13softmax function converts the raw scores
  401. 15:15into a probability distribution and very
  402. 15:18large negative values like negative
  403. 15:19infinity get converted to zero
  404. 15:21probability effectively masking out
  405. 15:24those positions. So for the first token
  406. 15:26only the first token gets the full
  407. 15:28attention weight. Similarly for the
  408. 15:31second token only the token itself and
  409. 15:33the tokens before it receive attention.
  410. 15:36In this way we can prevent tokens from
  411. 15:38attending to future tokens. So in masked
  412. 15:41multi head attention we multiply query
  413. 15:43and key to get the attention score then
  414. 15:45add the mask matrix then scale apply
  415. 15:48softmax and multiply with the value
  416. 15:50matrix as in the encoder to get the
  417. 15:51final embedding. We take multiple
  418. 15:54identical self attention and then
  419. 15:56concatenate the output. Then we take
  420. 15:58another output projection matrix to get
  421. 16:01the final representation.
  422. 16:03This is what masked multi head attention
  423. 16:05is. There is also an add and layer
  424. 16:08normalization after the attention layer
  425. 16:10which is similar to the encoder layer.
  426. 16:13Now after the masked multi head
  427. 16:14attention we have another attention
  428. 16:16layer which is called cross attention.
  429. 16:23In the cross attention, we take both the
  430. 16:25encoder input and decoder input to
  431. 16:27compute the attention. From the encoder,
  432. 16:29we have the source language
  433. 16:31representation. And this input context
  434. 16:33now becomes the input to the cross
  435. 16:35attention. In this attention layer, the
  436. 16:38encoder output acts as input to get the
  437. 16:40key and value matrices while the
  438. 16:42decoders acts as the query. In the
  439. 16:44attention calculation, we multiply the
  440. 16:46key matrix which we get from the
  441. 16:48encoder's output and the query matrix
  442. 16:50which we get from the decoder's first
  443. 16:52normalization layer to get the attention
  444. 16:54score. This is called cross attention
  445. 16:57because here we are attending the
  446. 16:59decoder sequence with the encoder
  447. 17:00sequence. We don't need to apply masking
  448. 17:03because the key matrix is from the
  449. 17:05encoder's output and does not contain
  450. 17:07any training label tokens. Multiple
  451. 17:10attention heads are used and the final
  452. 17:12outputs are concatenated and projected
  453. 17:14through the output layer. This is what
  454. 17:17cross attention is and this is where we
  455. 17:19use our source language context in the
  456. 17:21decoder. Now all other layers we already
  457. 17:24discussed which include another add and
  458. 17:26layer norm and another feed forward
  459. 17:28network with layer normalization are
  460. 17:30also present in the decoder. This is
  461. 17:33repeated multiple times and finally we
  462. 17:36have the output from the decoder layer.
  463. 17:43In the output layer, we have another
  464. 17:45linear layer with softmax activation
  465. 17:48which gives the probability
  466. 17:49distribution. We have our training label
  467. 17:52ready and using this training label, we
  468. 17:54compute the loss and apply back
  469. 17:55propagation to train the model. So at
  470. 17:58the high level, this is how training is
  471. 18:00done in transformer architecture. This
  472. 18:03was all about training our model for
  473. 18:04language translation. But now let's see
  474. 18:07how we can inference using this model.
  475. 18:09During training we have the source
  476. 18:11language acting as input to the encoder
  477. 18:13and target language as decoder input and
  478. 18:15label. But during inference we just have
  479. 18:17the source language and now have to
  480. 18:19generate the target language. First like
  481. 18:22during training source language becomes
  482. 18:24input to the encoder where every step is
  483. 18:26similar as like during training. Now in
  484. 18:28the decoder at first we just start with
  485. 18:30start of sentence token since we have to
  486. 18:32generate the target language. So this
  487. 18:34single start of sentence token becomes
  488. 18:36input to the decoder in the attention
  489. 18:39calculation we use this token since we
  490. 18:41only have this token. Then we move ahead
  491. 18:44and in the cross attention layer only
  492. 18:46one token attends to other tokens from
  493. 18:48the encoder because we get that key and
  494. 18:51value from the encoder and query from
  495. 18:53the decoder. Then all other processes
  496. 18:55remain similar as in training and at the
  497. 18:58final layer we get the probability
  498. 19:00distribution that predicts the next
  499. 19:01token. We sample from this distribution
  500. 19:04and append that token into the input
  501. 19:06list in our decoder. Now decoder again
  502. 19:09performs but now using two input tokens
  503. 19:12and gives output probability
  504. 19:13distribution. We take the last token
  505. 19:16probability value to sample because we
  506. 19:18already have our second token and want
  507. 19:20to predict the third token. Similarly,
  508. 19:23we append that token in our input and
  509. 19:25then repeat the process until we get the
  510. 19:26special end of sequence token. Once we
  511. 19:29get that end of sequence token, we stop
  512. 19:31this loop and now have complete
  513. 19:32translation. This is the reason why we
  514. 19:35format our data set with start of
  515. 19:36sequence and end of sequence tokens. In
  516. 19:38training, we take the source language as
  517. 19:41input to the encoder and target language
  518. 19:43as input to the decoder which gives us
  519. 19:45probability and train in a single shot.
  520. 19:47But during inference, we generate one
  521. 19:49token at a time until we predict end of
  522. 19:52sequence token. So this nature during
  523. 19:54inference is different than during the
  524. 19:56training and is called the auto
  525. 19:57reggressive nature and text generation
  526. 20:00models like GPT use these decoderonly
  527. 20:03architecture models for text generation
  528. 20:05task. This completes the transformer
  529. 20:07architecture about what each component
  530. 20:10does, how we train it and how we can
  531. 20:11inference.

About this transcript

This page contains the full transcript of Transformer Architecture Explained by Under The Hood, generated from the public captions YouTube serves with the video. The transcript has 3,388 words across 531 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.