Transformers for beginners | What are they and how do they work — Transcript
Full transcript
- 0:00transformers came into our lives just a
- 0:02couple of years ago but they have been
- 0:04taking the nlp area by storm libraries
- 0:06like hugging phase has made it very easy
- 0:08for everyone to use transformers or
- 0:11implementations like bert or gpt3 is the
- 0:14reason that everyone is talking about
- 0:15them but what are they and how do they
- 0:18work so in this video we will look
- 0:20closely into transformers and understand
- 0:22their working principles this video is
- 0:24part of the deep learning explained
- 0:26series by assembly ai which is a company
- 0:28that is making a state-of-the-art speech
- 0:31to text api if you want to use assembly
- 0:33ai for free get your free api token
- 0:36using the link in the description
- 0:38before transformers word is coming we
- 0:40were using rnns to deal with text data
- 0:42or any sequence data really but the
- 0:44problem with rnns is that when you give
- 0:46it a very long sentence it tends to
- 0:48forget the beginning of the sentence
- 0:50when it comes to the end of the sentence
- 0:52and because they rely on recurrence well
- 0:54it's in the name recurrent neural
- 0:55network they cannot be paralyzed
- 0:58then we start using lstms lstms are a
- 1:01little bit more sophisticated they tend
- 1:02to remember information for a little bit
- 1:04longer of a time but they take very long
- 1:07to train
- 1:08well then we have transformers
- 1:10transformers only rely on attention
- 1:12mechanisms to remember things they do
- 1:14not have any recurrence at all and
- 1:16thanks to this they are faster because
- 1:18we can parallelize them we can train
- 1:20them in a parallel way okay but what is
- 1:22this attention we can definitely make
- 1:24another video to talk about that and if
- 1:26you're interested in that definitely
- 1:28comment and let me know but generally
- 1:30attention is the ability of a model to
- 1:32pay attention to the important part of a
- 1:34sentence or an image or any kind of
- 1:36input really so if it's a sentence this
- 1:40is what it would look like
- 1:41so let's say we have a english sentence
- 1:44and the sentence is the agreement on the
- 1:46european economic area was signed in
- 1:48august 1992 and the other side is the
- 1:51french translation of that but i do not
- 1:53know the first thing about french so i'm
- 1:54not even going to try to pronounce it
- 1:57but as you can see in this chart what we
- 1:59see is that the lighter the color of the
- 2:01square the more attention our model is
- 2:03paying to the word in that line or in
- 2:06that row or column and as you can see it
- 2:09does not always go in a diagonal way
- 2:12when it is translating european economic
- 2:15area because the word order is reversed
- 2:19in french it is paying attention in a
- 2:21reversed way
- 2:23if this was an image and let's say we
- 2:25are looking for dogs and images and we
- 2:27are trying to um classify different
- 2:30breeds of dogs then you can see what
- 2:33your model is paying attention to is it
- 2:35the noses of the dogs is it the ears of
- 2:37the dogs what exactly in an image is the
- 2:40model paying attention to to be able to
- 2:42understand the difference between dog
- 2:44breeds alright now that we briefly
- 2:46looked at what attention is let's look
- 2:48into how the transformer networks learn
- 2:50and how what their architecture is
- 2:53so this is what a transformer network
- 2:55more or less looks like but we will go
- 2:57and start from the higher levels and
- 2:59then start breaking down everything and
- 3:01understand how they work together so the
- 3:04basic thing on a very high level what
- 3:06transformers have is an encoder and a
- 3:10decoder part
- 3:11but actually what they have is six
- 3:14encoders and six decoders but basically
- 3:16the right left-hand side is the encoders
- 3:18and the right-hand side is the decoders
- 3:21each encoder has one self-attention
- 3:23layer that is paying attention to the
- 3:25sentence itself and one feed forward
- 3:28forward neural network layer and every
- 3:31decoder has two self-attention layers
- 3:34and one feed forward neural network
- 3:36layer the parallelization comes from how
- 3:38we feed the data into this network we
- 3:41feed all the words of the sentence at
- 3:43the same time to our network
- 3:45specifically the encoder inside the
- 3:47first step which is the self-attention
- 3:49sub layer
- 3:50all the words of the sentence is
- 3:52compared to all the other words in the
- 3:54sentence so there is some communication
- 3:56between the words
- 3:57whereas in the next step in the
- 3:59feed-forward neural network they are
- 4:01passed through a feed-forward neural
- 4:03network separately so they do not have
- 4:05any information exchange but the
- 4:07feed-forward neural networks that they
- 4:09are passed through are the same inside
- 4:11the same layer but as we said there are
- 4:13six encoders and each each in each of
- 4:15these six encoders the neural networks
- 4:17are different okay so this has been kind
- 4:19of the middle part of the network we
- 4:21also have the inputs and the outputs so
- 4:24all the inputs that go in either the
- 4:26encoder or the decoder the raw inputs
- 4:28are embedded
- 4:30what are embeddings well that's a little
- 4:32bit of a longer topic for this video but
- 4:34again if you like us to make a video on
- 4:36this leave a comment uh but what you
- 4:38need to know for now is that embeddings
- 4:40are a way to represent these words in a
- 4:44n length vector in this specific
- 4:47transformer architecture they are using
- 4:49512 length vectors and that's basically
- 4:52what they use in the original paper but
- 4:54this is a hyper parameter that you can
- 4:56change and on top of these word
- 4:58embeddings we are adding positional
- 5:00encodings so if you remember we said
- 5:02transformers do not have any recurrence
- 5:05so the model has no way of understanding
- 5:07which word comes first and the other one
- 5:09comes second or which word comes where
- 5:11in the sentence so by adding a
- 5:13positional encoding you are letting or
- 5:15you are adding some information or
- 5:17injecting some information with each
- 5:18word that tells the modal where this
- 5:21word in the sentence comes in and lastly
- 5:24for the output as you can see we have a
- 5:26linear layer and a softmax layer at the
- 5:29end of the decoders so the output of the
- 5:31decoders can be transformed into
- 5:33something that we can understand and
- 5:35basically what they turn into is a
- 5:37vector that has the length of the amount
- 5:39of words that we have in our vocabulary
- 5:41and each of these cells tells us how
- 5:43likely it is that this word in this cell
- 5:47is going through the next word in our
- 5:48sequence and those are the main
- 5:50components but there are two little
- 5:52things that make transformers a little
- 5:54bit better one of them is the
- 5:55normalization layers so if you realize
- 5:58in between the sub layers that is the
- 6:01self-attention layers and the
- 6:02feed-forward neural networks we have
- 6:05some add and normalized layers and what
- 6:07they do is to normalize the output that
- 6:10comes from the sub layer the
- 6:12normalization technique that is used
- 6:13there is called layer normalization and
- 6:15that is basically an improvement over
- 6:18batch normalization and if you don't
- 6:19know what batch normalization is we
- 6:21already made a video about that i will
- 6:23link it somewhere here and you can go
- 6:25watch that to understand a little bit
- 6:26better what batch normalization or layer
- 6:28normalization is and the second little
- 6:30detail is the skip state so if you look
- 6:32at the original architecture image we
- 6:34see that there are some arrows that are
- 6:36going around some of the sub-layers well
- 6:38actually all of the sub-layers so some
- 6:40of the information that does not go into
- 6:43either the self-attention layer
- 6:45sub-layers or the feed-forward neural
- 6:47networks are sent directly to the
- 6:49normalization layer this kind of helps
- 6:52the model not forget things and it helps
- 6:54the model to forward information that is
- 6:57important to further in the network
- 6:59inside these normalization layers what
- 7:01we do is add the information that went
- 7:04through the sub layer and also just skip
- 7:06the sub layer and then normalize them
- 7:08together and that's all there is to the
- 7:10architecture of transformers and if you
- 7:12look into it actually most of the things
- 7:15that are inside this architecture are
- 7:17things that we've known from before
- 7:19things have been around for a long time
- 7:21for example
- 7:22linear transformations soft max layers
- 7:24or
- 7:25word embeddings for example or the feed
- 7:27forward neural networks but there are
- 7:30two really noble ideas inside the
- 7:32original transformer paper that really
- 7:36made the difference for transformers and
- 7:37those were positional encodings and
- 7:40multi-headed attention so let's take a
- 7:42closer look at how they work let's start
- 7:44with multi-headed attention layers so if
- 7:46you look at the original architecture
- 7:48you see that there are two different
- 7:49types of multiheaded attention layers
- 7:52one of them is just multi-headed
- 7:53attention and the other one is masked
- 7:55multi-headed attention well it actually
- 7:57does the same thing no matter if it's
- 7:59called masked or not if it's in an
- 8:01encoder and a decoder and the only
- 8:03difference is in a normal multiheaded
- 8:05attention layer all the words are
- 8:07compared with all the other words that
- 8:09are inputted that are in a sentence but
- 8:12that will make more sense to you in a
- 8:13second what i mean by comparing and for
- 8:16masked multi-headed attention layers
- 8:18only the words that are coming before a
- 8:20word are compared to that word in the
- 8:22sentence in the attention layer
- 8:24something called the scaled dot product
- 8:26attention is used and then it is
- 8:28multiplied and is done multiple times to
- 8:31create that multi-headed effect and of
- 8:33course everything is done in mattresses
- 8:35to make things faster but i will show
- 8:37you how attention is calculated using
- 8:39just the vectors of words what do we
- 8:41have in the beginning are embeddings of
- 8:43words if you remember we embedded the
- 8:45words into a vector and also we added
- 8:48positional encodings and then this is
- 8:50fed to the first encoder and of course
- 8:52at first is fed to the multiheaded
- 8:54attention sub layer of the first encoder
- 8:57so in there what is done first is to
- 9:00multiply these embedding vectors with
- 9:03some mattresses these mattresses are
- 9:05called query key and value mattresses
- 9:08and these are values that are
- 9:10initialized randomly and are trained
- 9:12during the training process to be
- 9:13learned kind of like the weights and
- 9:15biases we have in neural networks as a
- 9:17result of this multiplication we get the
- 9:19query key and value vector for each word
- 9:23and from this point on we are going to
- 9:25use these vectors to
- 9:27keep going with the calculation the
- 9:29first thing that we want to do is to
- 9:30calculate a score for each word against
- 9:33all the other words in the sentence what
- 9:35we do for this is we dot take the dot
- 9:38product of the curie vector of each word
- 9:41against the key vector of all the other
- 9:44words so if you want to get the score of
- 9:46the first word on the first word what we
- 9:48do is we get dot product of the query
- 9:51vector of word one with the key vector
- 9:54of word one if we want to get the score
- 9:57of the first word against the second
- 9:59word we must get the dot product of the
- 10:01query vector of the first word with the
- 10:04key vector of the second word so when we
- 10:06get the dot product of all the key
- 10:09vectors of all the other words compared
- 10:12with or combined with the query vector
- 10:14of the first word we have the score of
- 10:16the first word against all the other
- 10:19words so they all belong all of these
- 10:21scores belong to the first word if you
- 10:23want to get the scores for the second
- 10:25word we're going to have to multiply its
- 10:27query vector with all the other words
- 10:30key vectors and this is what it's done
- 10:32and it's all done in parallel so that's
- 10:34why we do not have any recurrence or we
- 10:36don't have to wait for other words to be
- 10:38done before starting to process the
- 10:40words further in the sentence we can do
- 10:42these calculations for all the words at
- 10:44the same time once we have all the
- 10:46scores of all the words against all the
- 10:48other words what we're going to do is to
- 10:50divide them by eight so that might sound
- 10:53like a very random number for you but
- 10:55it's basically the square root of 64
- 10:58which is again sound like a random
- 11:00number that comes out of nowhere but
- 11:02actually 64 is the square root of the
- 11:04length of the query key and value
- 11:07vectors and that's why the authors of
- 11:09the original transformers paper are
- 11:11using that number after we divide
- 11:13everything by 8 we pass all these values
- 11:15to a softmax layer we do this to
- 11:17normalize all these values and the score
- 11:20values of one word against all the other
- 11:22words are now going to sum up to one the
- 11:24resulting number serves kind of like a
- 11:26weight from this point on what we're
- 11:28going to do is to multiply all the value
- 11:30vectors of all the words with this
- 11:32weight and finally you sum up all the
- 11:35weighted value vectors of all the words
- 11:38based on this one word that we were
- 11:40doing the calculations for and create
- 11:42the output of the self-attention layer
- 11:44for this one word and then you have to
- 11:46do all of these calculations for all the
- 11:48other words well this is done
- 11:50simultaneously but at the end you have
- 11:52the output of the attention layer
- 11:55as i mentioned multi-headed attention
- 11:57does this eight times so effectively it
- 11:59is training eight different query key
- 12:02and value mattresses not the vectors for
- 12:05the words but the mattresses that we
- 12:07multiply the
- 12:09input embeddings with this way the model
- 12:11is able to pay attention to not one
- 12:13other word but many other words in the
- 12:16sentence so in the model they're using
- 12:18the number eight but you can basically
- 12:20change it if you like to and let's look
- 12:22at this example again that we gave at
- 12:24the beginning of this video as you can
- 12:26see there are some of the um cells that
- 12:29are really bright but at the same time
- 12:31we also have some cells that are just
- 12:32kind of gray and that means that our
- 12:35model was also paying attention to those
- 12:37other words other than the primary word
- 12:39that they're paying attention to a
- 12:40little bit and this multiheaded
- 12:42attention thing is also one of the
- 12:44reasons why transformers are so
- 12:46seamlessly able to deal with
- 12:48sentences of different lengths one thing
- 12:51you might catch here is that if we do
- 12:53the same thing eight times what we're
- 12:55going to have is eight different
- 12:56resulting mattresses right and inside
- 12:59these mattresses one line is going to
- 13:00correspond to one word but we're going
- 13:02to have eight of them so how are we
- 13:04going to deal with this well what they
- 13:06do in the paper or what they propose to
- 13:08do is basically concatenate them all
- 13:10together and then multiply them with yet
- 13:13another weight matrix that is going to
- 13:15produce a matrix that is going to look
- 13:17like only one output of the attention
- 13:20layer this weight matrix is of course
- 13:22yet another thing to train inside the
- 13:24transformers on top of the key value and
- 13:27curing mattresses that we multiply all
- 13:30of the word embeddings with next is
- 13:32positional encoding so as i mentioned
- 13:34before position line codings is a way to
- 13:36inject or add information to the word
- 13:39embeddings that we've created before to
- 13:42show where in a sentence a word is so
- 13:45basically the location information of a
- 13:47word you can either use learned
- 13:49positional encodings or fixed positional
- 13:51encodings but in the paper in the
- 13:53original paper they suggest or they
- 13:55recommend that we use fixed positional
- 13:57encodings because they have the
- 13:59advantage of being able to handle
- 14:02lengths of sentences that we haven't
- 14:04seen in the training set
- 14:06you might say why do we need any
- 14:08sophisticated solution for this anyways
- 14:10why can't we just assign a number to the
- 14:12word specifying where in the sentence
- 14:15this word is so that wouldn't really
- 14:17work because let's say if you assign a
- 14:19number that goes from zero to one what's
- 14:21going to happen is that you're not going
- 14:23to really understand how many words are
- 14:25in that sentence just by looking at this
- 14:28one word and this value will not be
- 14:29consistent in between examples another
- 14:32solution could be to assign integers to
- 14:34words of course starting from one or
- 14:36zero to however long the sentence is but
- 14:39the problem with that one is that those
- 14:41numbers can get very high right if you
- 14:43have a very long sentence that could get
- 14:45out of control and on top of that there
- 14:47could be sentences with specific lengths
- 14:50that you do not have in the training
- 14:51data and that could cause some problems
- 14:53in terms of generalization so what they
- 14:56did as a solution to this positional
- 14:57problem in the original transformers
- 14:59paper was to use sine and cosine
- 15:01functions in different frequencies
- 15:04but of course i don't expect you to know
- 15:06this so let's look into how that looks
- 15:08so this is what sine and cosine
- 15:10functions in different frequencies look
- 15:12like the colors here show us numbers
- 15:14that range from -1 to 1. the x-axis
- 15:18shows us the length of the word
- 15:19embeddings in transformers we are using
- 15:21512 as i mentioned before and the y-axis
- 15:25is the position of this token of this
- 15:27word so if i want to get the positional
- 15:29encoding of a word that is in let's say
- 15:32the 20th position i need to get the
- 15:34horizontal line that corresponds to 20
- 15:38in the y-axis and the nice thing about
- 15:41this positional encoding is that it's
- 15:43going to be unique no other place no
- 15:45other horizontal line in this graph has
- 15:48the same composition of values as in
- 15:50that line and one other nice thing about
- 15:52this positional encodings is that you
- 15:54can always tell the difference between
- 15:56two words looking at these positional
- 15:57encodings it's always going to be the
- 16:00same one thing that really helped me
- 16:01understand this concept was to look at
- 16:03binary representations of integers so
- 16:06let's look at these examples
- 16:08if you realize as you increase your
- 16:10numbers what happens is the smaller
- 16:12digit in the binary representation
- 16:15changes from one to zero with every new
- 16:17integer whereas the second digit changes
- 16:20every two integers so at first it is
- 16:22zero and zero and in the second two
- 16:25integers it is one and one and the third
- 16:27two integers it is zero and zero again
- 16:30and again this pattern kind of follows
- 16:32itself and what happens is all of these
- 16:35binary representations are unique no two
- 16:37binary representations are the same and
- 16:39on top of that you can always tell the
- 16:41difference between two integers by
- 16:43looking at their binary representations
- 16:45this could also be a perfectly useful
- 16:47positional encoding for us too but it is
- 16:50only ones and zeros and we are not
- 16:52actually using the information that can
- 16:54be provided with continuous values so
- 16:56that's why instead we use sine and
- 16:57cosine functions okay what do we do once
- 17:00we have these encodings right let's say
- 17:02we have this encoding of 512
- 17:05values that we extracted from this graph
- 17:07well what we do is we basically add them
- 17:10together we add the word embedding and
- 17:12the positional encoding together and
- 17:14then we feed it to the encoders
- 17:16all right so we learned everything that
- 17:18we need about the architecture there are
- 17:20encoders specifically six of them and
- 17:22there are recorders again six of them we
- 17:25have the uh last processing and the
- 17:27output we have the embeddings at the
- 17:29input and also the positional encodings
- 17:32but how does it all work so basically to
- 17:34bring it together what happens is you
- 17:36first get your inputs
- 17:38run them through the embeddings run them
- 17:40through the positional encodings and
- 17:42then run them through six levels of
- 17:44encoders and then you get an output
- 17:46this output is fed to all of the
- 17:48decoders so we have six decoders as we
- 17:51mentioned six layers of decoders this
- 17:54the information from the output from the
- 17:56encoder is fed to all of the decoders
- 17:59but this information is only fed to the
- 18:01second sub-layer so the multi-headed
- 18:03attention sub-layer of the coders
- 18:05and the first masked multi-headed
- 18:08attention a sub-layer of decoders get
- 18:10the input from what was outputted from
- 18:13the decoder section of the model in the
- 18:16previous time step that way the decoders
- 18:18are taking into consideration what was
- 18:21the word in the previous time step on
- 18:24the previous position and also the
- 18:26context that they learned from the
- 18:28encoding process of the network to
- 18:31create the output these decoders all
- 18:33work together and then they create a
- 18:36output vector this output vector is sent
- 18:38to through the linear transformation
- 18:40that creates a logic's vector this
- 18:42logic's vector is as long as the amount
- 18:45of words that we have in our vocabulary
- 18:47and it has the possibilities of how
- 18:50likely is the next word going to be one
- 18:53word or the other
- 18:54and then we pass this through soft max
- 18:56to be able to get the probabilities of
- 18:59each word and then these probabilities
- 19:01will also add up to one so basically a
- 19:03normalized version of the logit's vector
- 19:06the output of the softmax layer
- 19:08basically tells us what the next word is
- 19:10going to be and that's all there is to
- 19:12know about transformers really it's
- 19:14quite simple even though it looks a bit
- 19:16complicated at first all you need to
- 19:18know that there are encoders and
- 19:19decoders and the two noble ideas that
- 19:22came into our lives with transformers
- 19:24are the positional encodings and the
- 19:26multiheaded attention layer to fully
- 19:28understand transformers and how they
- 19:30work you might need to watch this video
- 19:31multiple times and maybe even support
- 19:33your learning with some of the written
- 19:35resources that are out there so for that
- 19:37i've left links to my favorite resources
- 19:39in the description if there was anything
- 19:41that was not clear or if you have a
- 19:42question leave a comment and let me know
- 19:45if you like this video don't forget to
- 19:46give us a thumbs up and maybe even
- 19:48subscribe to be one of the first people
- 19:50to know when we make a new video but
- 19:52before you leave don't forget to grab
- 19:53your free token for assembly ai's
- 19:55special text api i'll see you in the
- 19:57next video
About this transcript
This page contains the full transcript of Transformers for beginners | What are they and how do they work by AssemblyAI, generated from the public captions YouTube serves with the video. The transcript has 3,794 words across 546 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.