Transformer Architecture Explained — Transcript
Full transcript
- 0:02Welcome back to another video. Today in
- 0:05this video I will explain the
- 0:07transformer architecture. It was
- 0:10introduced in the paper attention is all
- 0:12you need and was published in the year
- 0:142017.
- 0:15Since then the transformer has become
- 0:17the foundation for many modern AI models
- 0:20and even today's most advanced AI
- 0:22systems. It has completely changed how
- 0:25we approach natural language processing.
- 0:28In this video, we'll start by
- 0:29understanding the basic components of
- 0:31the transformer like token embeddings,
- 0:33attention mechanisms, feed forward
- 0:35networks, residual connections, and
- 0:37more. Then we'll move on to how the
- 0:39model is trained and how it works during
- 0:41inference. To get a better idea of how
- 0:44it works, we'll walk through a simple
- 0:46example, translating a sentence from one
- 0:48language to another. The transformer
- 0:51architecture has two main parts, the
- 0:53encoder and the decoder, and it follows
- 0:55a sequencetosequence design. Just like
- 0:57earlier sequencetose sequence models
- 0:59that used LSDM,
- 1:01but the transformer improves on them by
- 1:03enabling parallel processing, better
- 1:05handling of long range dependencies, and
- 1:08overall faster and more accurate
- 1:10results. Let's start with the first
- 1:12step, data set preparation for our
- 1:14translation task. We will use this data
- 1:17set to train our model. We start by
- 1:20formatting our data set with special
- 1:21tokens, start and end of sequence
- 1:24tokens. These are the special tokens
- 1:26which we use to identify the start and
- 1:28end of a sentence. The source language
- 1:31becomes input to the encoder and the
- 1:33target language becomes input to the
- 1:35decoder. We also create training labels
- 1:38by shifting the target sentence one step
- 1:40ahead of the decoder's input and we also
- 1:42have that extra end of sequence token. I
- 1:44will discuss about this format later.
- 1:46But now let's say we have our data set
- 1:48ready. The overall flow looks like this.
- 1:52The source sentence is fed into the
- 1:54encoder which generates a context
- 1:56vector. The decoder takes in the target
- 1:58sentence input and outputs a probability
- 2:01distribution. The model is then trained
- 2:04by comparing this output to the target
- 2:06label. Now that we have the big picture,
- 2:08let's dive deeper and explore each
- 2:10component of the transformer step by
- 2:12step.
- 2:16Let's start with the encoder part of the
- 2:18transformer. We begin with a data set
- 2:20that contains source and target language
- 2:22pairs already formatted with special
- 2:25tokens like start of sequence and end of
- 2:27sequence tokens. The source language
- 2:29becomes the input to the encoder. The
- 2:32first step is tokenization.
- 2:34Here we convert the input text into
- 2:36numbers by breaking each sentence into
- 2:38smaller units like words or subwords.
- 2:41There are different types of
- 2:42tokenization methods such as word level
- 2:45or subword level. But for simplicity,
- 2:47let's assume each word is treated as a
- 2:49unique token and each token is assigned
- 2:51a unique number. After tokenizing the
- 2:54input sentence, we get a list of token
- 2:56numbers. In practice, we don't train
- 2:59with just one sequence at a time.
- 3:01Instead, we use batches of sequences.
- 3:04Since sequences can have different
- 3:06lengths, we need to make them all the
- 3:08same length by filling shorter ones with
- 3:10a special token called the padding
- 3:11token. This allows us to train in
- 3:14batches efficiently. But for now, we'll
- 3:16keep things simple and use a single
- 3:18example without padding. Next, these
- 3:21tokens go into the embedding layer. This
- 3:24is a trainable neural network layer that
- 3:26converts each token number into a
- 3:28highdimension vector called the
- 3:30embedding of the token. The token number
- 3:32by itself has no real meaning. But the
- 3:35embedding vector helps capture the
- 3:37semantic meaning of the word. Each token
- 3:39in our sequence gets its own embedding
- 3:41vector. In the original paper, the size
- 3:44of this vector is 512.
- 3:46These vectors exist in a highdimensional
- 3:48space where words with similar meanings
- 3:51tend to be closer together in the
- 3:52embedding space. Imagine a 2D version of
- 3:55this space. Words like bank, river, and
- 3:58flood appear closer together because
- 4:00they are often used in similar contexts.
- 4:02This helps the model understand how
- 4:04words relate to each other in meaning.
- 4:07Now that we have word representations,
- 4:09we also need to know the order of the
- 4:10words in the sentence like which word
- 4:12comes first, second and so on. The
- 4:15embedding vector alone doesn't contain
- 4:17position information. So we need to add
- 4:20that separately using positional
- 4:21encoding. In the paper, the authors
- 4:24added another vector called the
- 4:26positional encoding vector to each
- 4:28embedding vector. This new vector has
- 4:30the same size of 512 dimension and is
- 4:33not trainable. It is generated using s
- 4:36and cosine functions. Here what it does
- 4:39for each token position. It computes a
- 4:42vector using a formula. In this vector,
- 4:44the even positions use the sign function
- 4:47and the odd positions use the cosine
- 4:49function and are computed using the
- 4:51formula where it takes the position of
- 4:53each token in the sequence and also the
- 4:55position of the dimension in the vector.
- 4:57Like this, we compute the new vector for
- 4:59all other tokens as well. This helps the
- 5:02model understand the position of each
- 5:04word in the sequence. These values are
- 5:07fixed. They're computed once and stored
- 5:09in memory, then reused during training.
- 5:12Even though more recent models use
- 5:13trainable position encodings, this
- 5:15static method still performs quite well
- 5:18and was used by the authors in the
- 5:19original transformer paper. So now we
- 5:22have our input sentence represented with
- 5:24the token embeddings that give the
- 5:26meaning of words and positional
- 5:28encodings that provide the word order.
- 5:31This combined input is then passed into
- 5:34the first layer of the encoder block.
- 5:36The first step in the encoder is the
- 5:38multi head attention layer. But before
- 5:41we dive into how multi head attention
- 5:43works, let's first understand why we
- 5:45need attention, what it does, and how
- 5:47it's calculated.
- 5:54What we have now is a token represented
- 5:56in this highdimensional embedding space
- 5:58which contains the semantic meaning of
- 6:00the words. But here the problem same
- 6:03words can have different meanings based
- 6:04on the context. But what the embedding
- 6:07does is only capture the semantic
- 6:09meaning and does not consider the
- 6:11context information. If we look at our
- 6:13example then the word bank should be
- 6:15close to other words related to
- 6:17financial institutions rather than the
- 6:19river bank. We know this by looking at
- 6:21its context words deposited and money.
- 6:24So what we want is to transform this
- 6:26representation so that it correctly
- 6:28represents the token by looking at its
- 6:30context. And for this we have the
- 6:32attention mechanism. It transforms the
- 6:35static embedding which only captures the
- 6:37semantic meaning into the representation
- 6:40that is aligned with the context. The
- 6:42size remains the same just the vectors
- 6:45are now represented according to the
- 6:46context. In the paper they call this
- 6:49self attention and it is represented by
- 6:51this equation. Let's break it down to
- 6:53understand what it means. First let's
- 6:55see how we get this query key and value
- 6:58that you see in the attention
- 6:59calculation. We take our static
- 7:01embeddings and linearly transform this
- 7:03embedding with weight matrices for query
- 7:05key and value. What this means is take
- 7:08our embedding as input and use this in
- 7:11three different neural networks with
- 7:12linear activations. What this gives is a
- 7:15query key and value which we use in our
- 7:17attention calculation. These weights are
- 7:20trainable layers and so the query key
- 7:22and value vectors are adjusted after
- 7:24training. Now that we have our query key
- 7:27and value matrices, what we do is take
- 7:29the query and transpose of the key
- 7:31matrix and then perform matrix
- 7:33multiplication.
- 7:34What this does is allow all the tokens
- 7:36to communicate with all other tokens in
- 7:38parallel by taking the dotproduct
- 7:41between the first query vector and the
- 7:43key vector. It gives a score on how much
- 7:45it scores to itself. The query acts like
- 7:48asking all other tokens and giving a
- 7:50score to each of them on how much they
- 7:51match. Similarly, it does this with all
- 7:54other tokens using the key matrix. By
- 7:57doing this, the first token in our
- 7:59sequence attends to all other tokens
- 8:01including itself and has a score that
- 8:03tells how much it matches with other
- 8:05tokens. Now, similarly, the second token
- 8:07in our sequence does the same by taking
- 8:10its query vector and asking all other
- 8:12tokens using the key vector. Like this,
- 8:15all tokens attend to each other using
- 8:16the query and key in order to know the
- 8:19importance of other tokens in the
- 8:20context to make a new representation.
- 8:23This is done in terms of matrix
- 8:25multiplication all at once and the
- 8:27result we get is what we call the
- 8:29attention score. Since the attention
- 8:31scores are just scalar values, they can
- 8:34be a varying range and are not
- 8:35interpretable. So we first scale these
- 8:38scores by dividing by the square root of
- 8:39the dimension of the key vector. We
- 8:42apply a softmax function to this scaled
- 8:44attention score across each row which
- 8:46gives us more interpretable values. So
- 8:49now each row sums to one and we can
- 8:51interpret the attention scores in terms
- 8:53of probability values. We call this
- 8:55attention weights. If we look at our
- 8:57token money and look at its attention
- 8:59weight, then we can interpret it as it
- 9:01is giving 25% importance to the word
- 9:03deposited, 30% to money and 30% to the
- 9:07token bank. Like this in the attention
- 9:09weight, we now know which token is
- 9:11important in creating a new
- 9:13representation.
- 9:15Using this attention weight and
- 9:16multiplying with the value matrix, we
- 9:18get the new representation that has the
- 9:20context information. So for the word
- 9:23money, it uses that weight value to act
- 9:25as a weighted sum to construct a new
- 9:27embedding which has context
- 9:29representation.
- 9:30Like this, a new representation based on
- 9:33this attention weight is calculated. And
- 9:36this is what the attention mechanism is
- 9:37about. Transforming the static embedding
- 9:40into an embedding that has the context
- 9:42representation.
- 9:43So the summary is take the static
- 9:46embedding with position encoding and
- 9:49then transform it in three different
- 9:50matrices called query key and value
- 9:53matrices using their respective weight
- 9:55matrices. Then multiply the query and
- 9:58key scale it and apply softmax to get
- 10:00the attention weights and multiply this
- 10:02with the value matrix to get the final
- 10:04contextual embeddings.
- 10:09Now let's see what multi head attention
- 10:11is in our attention calculation. What we
- 10:14did was take the embedding input and
- 10:16transform it into query key and value
- 10:19which had the dimensions equal to the
- 10:20original embedding and then calculate
- 10:23the new embedding all at once. But here
- 10:25in multi head attention instead of doing
- 10:27it all at once we take multiple
- 10:29attention layers. Each attention now has
- 10:32its own weight matrices for query key
- 10:34and value but in a lower dimension. All
- 10:37the computations are the same but we get
- 10:39the final output in a lower dimension.
- 10:41This is what a single attention head
- 10:43does. We use multiple identical
- 10:46attention heads with different weight
- 10:47matrices in a multi head attention. Each
- 10:50attention contributes to producing the
- 10:52final contextual representation.
- 10:55Since each attention head has output in
- 10:57a lower dimension, we concatenate the
- 10:59output of each attention head. So if we
- 11:02want our final embedding to be of size
- 11:04512 and want to take eight attention
- 11:06heads then we should make the inner
- 11:08dimension of each head to be 64 then we
- 11:11again transform this using another
- 11:13weight matrix W out to get the final
- 11:16representation.
- 11:17This is what happens in multi head
- 11:19attention. It takes in the static
- 11:21embedding and gives us new contextaware
- 11:23embeddings using multiple self attention
- 11:25heads. This completes the multi head
- 11:28attention layer. Now after that we have
- 11:30the add and normalization layer. There
- 11:33is a residual stream that adds the
- 11:35static embedding to the result of the
- 11:37multi head attention. This is common in
- 11:40deep learning to prevent vanishing
- 11:42gradients. After the residual addition
- 11:44we have the layer normalization.
- 11:47Normally in batch normalization when we
- 11:49have input sequence in batches we
- 11:52normalize across each feature dimension
- 11:54across the batch. But in layer
- 11:56normalization which is used in this
- 11:58architecture, we normalize across each
- 12:01individual token by calculating the mean
- 12:03and variance across the token's
- 12:05features. After the normalization layer,
- 12:07we have our normalized output with the
- 12:10size remaining the same due to just the
- 12:11addition and normalization steps. So up
- 12:14until now, the overall flow looks like
- 12:16this.
- 12:22Now next we have the feed forward neural
- 12:24network. After the first layer norm,
- 12:27this becomes the input to the feed
- 12:28forward neural network. This is a dense
- 12:31neural network with nonlinearity in it.
- 12:34The output size is the same as the input
- 12:36size due to the equal number of neurons
- 12:38in the input and output layers. There is
- 12:41another add and layer norm like before
- 12:43but for the feed forward neural network.
- 12:46After this what we have is the output of
- 12:48the encoder block. Now this output
- 12:51becomes the input to another identical
- 12:53encoder block repeating multiple times
- 12:55with different parameters further
- 12:57refining the input sequence. This is
- 13:00what the encoder block looks like where
- 13:02the final output is now the context for
- 13:04the decoder. So up until now we give
- 13:07input to the encoder repeat it multiple
- 13:09times and now the output will be the
- 13:11representation of the input sequence
- 13:13that the decoder will use.
- 13:20Let's move on to the decoder part and
- 13:22see what it does. We already have our
- 13:25data set formatted for training. The
- 13:27source language is the encoder input
- 13:30which we used in the encoder and now we
- 13:32use the target language which becomes
- 13:34the input to the decoder. Also the
- 13:36target language shifted ahead will be
- 13:38our training label. The process is
- 13:41similar to that in the encoder. First
- 13:43the tokenization step uses the target
- 13:46language specific tokenizer to tokenize
- 13:48the input. Again for simplicity we are
- 13:51just using word level tokenization.
- 13:54Then we have the embedding layer with
- 13:56positional encoding as before. Now this
- 13:59goes into the decoder block where the
- 14:01first layer is the masked multi head
- 14:03attention layer. This is almost similar
- 14:05to the multi head attention layer, but
- 14:08with just a slight modification.
- 14:10As in the encoder's multi head attention
- 14:12layer, we compute the attention score by
- 14:14multiplying the query and key matrices
- 14:17that we get by transforming our input
- 14:19using three different weight matrices.
- 14:21This gives us the attention score for
- 14:23our target language where each token
- 14:25attends to other tokens. If you look at
- 14:28our data set, the decoder input is the
- 14:30same as the training label, just one
- 14:32token shifted ahead. So during training
- 14:35each token in the decoder tries to
- 14:36predict the next token in the sequence.
- 14:39You can see the first token attends to
- 14:41its target token during the attention
- 14:43calculation.
- 14:44Similarly all the tokens are attending
- 14:46to their target token and future tokens
- 14:48during the attention calculation which
- 14:51we don't want. We want to prevent the
- 14:53token from attending to its target and
- 14:55future tokens. To do this, we take a new
- 14:58matrix which we call the masking matrix
- 15:00and it contains negative infinity values
- 15:03in the positions that we want to mask.
- 15:05We add this matrix to our original
- 15:07attention matrix to get the masked
- 15:08attention score. What this does is
- 15:11during attention calculation, the
- 15:13softmax function converts the raw scores
- 15:15into a probability distribution and very
- 15:18large negative values like negative
- 15:19infinity get converted to zero
- 15:21probability effectively masking out
- 15:24those positions. So for the first token
- 15:26only the first token gets the full
- 15:28attention weight. Similarly for the
- 15:31second token only the token itself and
- 15:33the tokens before it receive attention.
- 15:36In this way we can prevent tokens from
- 15:38attending to future tokens. So in masked
- 15:41multi head attention we multiply query
- 15:43and key to get the attention score then
- 15:45add the mask matrix then scale apply
- 15:48softmax and multiply with the value
- 15:50matrix as in the encoder to get the
- 15:51final embedding. We take multiple
- 15:54identical self attention and then
- 15:56concatenate the output. Then we take
- 15:58another output projection matrix to get
- 16:01the final representation.
- 16:03This is what masked multi head attention
- 16:05is. There is also an add and layer
- 16:08normalization after the attention layer
- 16:10which is similar to the encoder layer.
- 16:13Now after the masked multi head
- 16:14attention we have another attention
- 16:16layer which is called cross attention.
- 16:23In the cross attention, we take both the
- 16:25encoder input and decoder input to
- 16:27compute the attention. From the encoder,
- 16:29we have the source language
- 16:31representation. And this input context
- 16:33now becomes the input to the cross
- 16:35attention. In this attention layer, the
- 16:38encoder output acts as input to get the
- 16:40key and value matrices while the
- 16:42decoders acts as the query. In the
- 16:44attention calculation, we multiply the
- 16:46key matrix which we get from the
- 16:48encoder's output and the query matrix
- 16:50which we get from the decoder's first
- 16:52normalization layer to get the attention
- 16:54score. This is called cross attention
- 16:57because here we are attending the
- 16:59decoder sequence with the encoder
- 17:00sequence. We don't need to apply masking
- 17:03because the key matrix is from the
- 17:05encoder's output and does not contain
- 17:07any training label tokens. Multiple
- 17:10attention heads are used and the final
- 17:12outputs are concatenated and projected
- 17:14through the output layer. This is what
- 17:17cross attention is and this is where we
- 17:19use our source language context in the
- 17:21decoder. Now all other layers we already
- 17:24discussed which include another add and
- 17:26layer norm and another feed forward
- 17:28network with layer normalization are
- 17:30also present in the decoder. This is
- 17:33repeated multiple times and finally we
- 17:36have the output from the decoder layer.
- 17:43In the output layer, we have another
- 17:45linear layer with softmax activation
- 17:48which gives the probability
- 17:49distribution. We have our training label
- 17:52ready and using this training label, we
- 17:54compute the loss and apply back
- 17:55propagation to train the model. So at
- 17:58the high level, this is how training is
- 18:00done in transformer architecture. This
- 18:03was all about training our model for
- 18:04language translation. But now let's see
- 18:07how we can inference using this model.
- 18:09During training we have the source
- 18:11language acting as input to the encoder
- 18:13and target language as decoder input and
- 18:15label. But during inference we just have
- 18:17the source language and now have to
- 18:19generate the target language. First like
- 18:22during training source language becomes
- 18:24input to the encoder where every step is
- 18:26similar as like during training. Now in
- 18:28the decoder at first we just start with
- 18:30start of sentence token since we have to
- 18:32generate the target language. So this
- 18:34single start of sentence token becomes
- 18:36input to the decoder in the attention
- 18:39calculation we use this token since we
- 18:41only have this token. Then we move ahead
- 18:44and in the cross attention layer only
- 18:46one token attends to other tokens from
- 18:48the encoder because we get that key and
- 18:51value from the encoder and query from
- 18:53the decoder. Then all other processes
- 18:55remain similar as in training and at the
- 18:58final layer we get the probability
- 19:00distribution that predicts the next
- 19:01token. We sample from this distribution
- 19:04and append that token into the input
- 19:06list in our decoder. Now decoder again
- 19:09performs but now using two input tokens
- 19:12and gives output probability
- 19:13distribution. We take the last token
- 19:16probability value to sample because we
- 19:18already have our second token and want
- 19:20to predict the third token. Similarly,
- 19:23we append that token in our input and
- 19:25then repeat the process until we get the
- 19:26special end of sequence token. Once we
- 19:29get that end of sequence token, we stop
- 19:31this loop and now have complete
- 19:32translation. This is the reason why we
- 19:35format our data set with start of
- 19:36sequence and end of sequence tokens. In
- 19:38training, we take the source language as
- 19:41input to the encoder and target language
- 19:43as input to the decoder which gives us
- 19:45probability and train in a single shot.
- 19:47But during inference, we generate one
- 19:49token at a time until we predict end of
- 19:52sequence token. So this nature during
- 19:54inference is different than during the
- 19:56training and is called the auto
- 19:57reggressive nature and text generation
- 20:00models like GPT use these decoderonly
- 20:03architecture models for text generation
- 20:05task. This completes the transformer
- 20:07architecture about what each component
- 20:10does, how we train it and how we can
- 20:11inference.
About this transcript
This page contains the full transcript of Transformer Architecture Explained by Under The Hood, generated from the public captions YouTube serves with the video. The transcript has 3,388 words across 531 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.