Deep Dive into LLMs like ChatGPT — Transcript
Full transcript
- 0:00hi everyone so I've wanted to make this
- 0:02video for a while it is a comprehensive
- 0:05but General audience introduction to
- 0:08large language models like Chachi PT and
- 0:11what I'm hoping to achieve in this video
- 0:12is to give you kind of mental models for
- 0:14thinking through what it is that this
- 0:17tool is it is obviously magical and
- 0:19amazing in some respects it's uh really
- 0:22good at some things not very good at
- 0:23other things and there's also a lot of
- 0:25sharp edges to be aware of so what is
- 0:28behind this text box you can put
- 0:29anything in there and press enter but uh
- 0:32what should we be putting there and what
- 0:34are these words generated back how does
- 0:36this work and what what are you talking
- 0:38to exactly so I'm hoping to get at all
- 0:40those topics in this video we're going
- 0:42to go through the entire pipeline of how
- 0:44this stuff is built but I'm going to
- 0:45keep everything uh sort of accessible to
- 0:48a general audience so let's take a look
- 0:50at first how you build something like
- 0:51chpt and along the way I'm going to talk
- 0:53about um you know some of the sort of
- 0:56cognitive psychological implications of
- 0:59the tools okay so let's build Chachi PT
- 1:02so there's going to be multiple stages
- 1:04arranged sequentially the first stage is
- 1:07called the pre-training stage and the
- 1:09first step of the pre-training stage is
- 1:11to download and process the internet now
- 1:13to get a sense of what this roughly
- 1:14looks like I recommend looking at this
- 1:16URL here so um this company called
- 1:20hugging face uh collected and created
- 1:23and curated this data set called Fine
- 1:26web and they go into a lot of detail on
- 1:28this block post on how how they
- 1:30constructed the fine web data set and
- 1:32all of the major llm providers like open
- 1:34AI anthropic and Google and so on will
- 1:36have some equivalent internally of
- 1:38something like the fine web data set so
- 1:41roughly what are we trying to achieve
- 1:42here we're trying to get ton of text
- 1:44from the internet from publicly
- 1:45available sources so we're trying to
- 1:47have a huge quantity of very high
- 1:50quality documents and we also want very
- 1:53large diversity of documents because we
- 1:55want to have a lot of knowledge inside
- 1:56these models so we want large diversity
- 1:59of high quality documents and we want
- 2:01many many of them and achieving this is
- 2:04uh quite complicated and as you can see
- 2:05here takes multiple stages to do well so
- 2:08let's take a look at what some of these
- 2:09stages look like in a bit for now I'd
- 2:11like to just like to note that for
- 2:13example the fine web data set which is
- 2:14fairly representative what you would see
- 2:16in a production grade application
- 2:18actually ends up being only about 44
- 2:20terabyt of dis space um you can get a
- 2:23USB stick for like a terabyte very
- 2:25easily or I think this could fit on a
- 2:27single hard drive almost today so this
- 2:29is not a huge amount of data at the end
- 2:31of the day even though the internet is
- 2:33very very large we're working with text
- 2:35and we're also filtering it aggressively
- 2:37so we end up with about 44 terabytes in
- 2:39this example so let's take a look at uh
- 2:42kind of what this data looks like and
- 2:44what some of these stages uh also are so
- 2:47the starting point for a lot of these
- 2:48efforts and something that contributes
- 2:50most of the data by the end of it is
- 2:52Data from common crawl so common craw is
- 2:56an organization that has been basically
- 2:57scouring the internet since 2007 so as
- 3:00of 2024 for example common CW has
- 3:03indexed 2.7 billion web
- 3:05pages uh and uh they have all these
- 3:08crawlers going around the internet and
- 3:09what you end up doing basically is you
- 3:11start with a few seed web pages and then
- 3:13you follow all the links and you just
- 3:15keep following links and you keep
- 3:16indexing all the information and you end
- 3:17up with a ton of data of the internet
- 3:19over time so this is usually the
- 3:21starting point for a lot of the uh for a
- 3:24lot of these efforts now this common C
- 3:26data is quite raw and is filtered in
- 3:27many many different ways
- 3:30so here they Pro they document this is
- 3:33the same diagram they document a little
- 3:35bit the kind of processing that happens
- 3:36in these stages so the first thing here
- 3:39is something called URL
- 3:41filtering so what that is referring to
- 3:43is that there's these block
- 3:47lists of uh basically URLs that are or
- 3:50domains that uh you don't want to be
- 3:52getting data from so usually this
- 3:54includes things like U malware websites
- 3:56spam websites marketing websites uh
- 3:58racist websites adult sites and things
- 4:01like that so there's a ton of different
- 4:02types of websites that are just
- 4:04eliminated at this stage because we
- 4:06don't want them in our data set um the
- 4:08second part is text extraction you have
- 4:10to remember that all these web pages
- 4:12this is the raw HTML of these web pages
- 4:14that are being saved by these crawlers
- 4:16so when I go to inspect
- 4:18here this is what the raw HTML actually
- 4:21looks like you'll notice that it's got
- 4:23all this markup uh like lists and stuff
- 4:26like that and there's CSS and all this
- 4:28kind of stuff so this is um computer
- 4:31code almost for these web pages but what
- 4:33we really want is we just want this text
- 4:35right we just want the text of this web
- 4:37page and we don't want the navigation
- 4:38and things like that so there's a lot of
- 4:40filtering and processing uh and heris
- 4:42that go into uh adequately filtering for
- 4:45just their uh good content of these web
- 4:48pages the next stage here is language
- 4:50filtering so for example fine web
- 4:53filters uh using a language classifier
- 4:56they try to guess what language every
- 4:58single web page is in and then they only
- 5:00keep web pages that have more than 65%
- 5:02of English as an
- 5:04example and so you can get a sense that
- 5:06this is like a design decision that
- 5:07different companies can uh can uh take
- 5:10for themselves what fraction of all
- 5:12different types of languages are we
- 5:14going to include in our data set because
- 5:15for example if we filter out all of the
- 5:17Spanish as an example then you might
- 5:19imagine that our model later will not be
- 5:21very good at Spanish because it's just
- 5:22never seen that much data of that
- 5:24language and so different companies can
- 5:26focus on multilingual performance to uh
- 5:28to a different degree as an example so
- 5:30fine web is quite focused on English and
- 5:33so their language model if they end up
- 5:35training one later will be very good at
- 5:36English but not may be very good at
- 5:38other
- 5:39languages after language filtering
- 5:41there's a few other filtering steps and
- 5:43D duplication and things like that um
- 5:46finishing with for example the pii
- 5:49removal this is personally identifiable
- 5:52information so as an example addresses
- 5:54Social Security numbers and things like
- 5:56that you would try to detect them and
- 5:57you would try to filter out those kinds
- 5:58of web pages from the the data set as
- 6:00well so there's a lot of stages here and
- 6:02I won't go into full detail but it is a
- 6:05fairly extensive part of the
- 6:06pre-processing and you end up with for
- 6:08example the fine web data set so when
- 6:10you click in on it uh you can see some
- 6:12examples here of what this actually ends
- 6:14up looking like and anyone can download
- 6:16this on the huging phase web page and so
- 6:19here are some examples of the final text
- 6:21that ends up in the training set so this
- 6:24is some article about tornadoes in
- 6:272012 um so there's some t tadoes in 2020
- 6:30in 2012 and what
- 6:33happened uh this next one is something
- 6:36about did you know you have two little
- 6:38yellow 9vt battery sized adrenal glands
- 6:41in your body okay so this is some kind
- 6:43of a odd medical
- 6:46article so just think of these as
- 6:49basically uh web pages on the internet
- 6:51filtered just for the text in various
- 6:53ways and now we have a ton of text 40
- 6:56terabytes off it and that now is the
- 6:58starting point for the next step of this
- 7:00stage now I wanted to give you an
- 7:02intuitive sense of where we are right
- 7:04now so I took the first 200 web pages
- 7:06here and remember we have tons of them
- 7:09and I just take all that text and I just
- 7:11put it all together concatenate it and
- 7:13so this is what we end up with we just
- 7:15get this just just raw text raw internet
- 7:18text and there's a ton of it even in
- 7:20these 200 web pages so I can continue
- 7:22zooming out here and we just have this
- 7:24like massive tapestry of Text data and
- 7:28this text data has all these p patterns
- 7:30and what we want to do now is we want to
- 7:31start training neural networks on this
- 7:33data so the neural networks can
- 7:35internalize and model how this text
- 7:39flows right so we just have this giant
- 7:42texture of text and now we want to get
- 7:45neural Nets that mimic it okay now
- 7:48before we plug text into neural networks
- 7:51we have to decide how we're going to
- 7:52represent this text uh and how we're
- 7:54going to feed it in now the way our
- 7:57technology works for these neuron Lots
- 7:58is that they expect
- 7:59a one-dimensional sequence of symbols
- 8:02and they want a finite set of symbols
- 8:05that are possible and so we have to
- 8:08decide what are the symbols and then we
- 8:10have to represent our data as
- 8:11one-dimensional sequence of those
- 8:14symbols so right now what we have is a
- 8:16onedimensional sequence of text it
- 8:18starts here and it goes here and then it
- 8:20comes here Etc so this is a
- 8:22onedimensional sequence even though on
- 8:23my monitor of course it's laid out in a
- 8:26two-dimensional way but it goes from
- 8:27left to right and top to bottom right so
- 8:29it's a one-dimensional sequence of text
- 8:32now this being computers of course
- 8:33there's an underlying representation
- 8:35here so if I do what's called utf8 uh
- 8:38encode this text then I can get the raw
- 8:41bits that correspond to this text in the
- 8:44computer and that's what uh that looks
- 8:46like this so it turns out that for
- 8:50example this very first bar here is the
- 8:53first uh eight bits as an
- 8:56example so what is this thing right this
- 8:59is um representation that we are looking
- 9:01for uh in in a certain sense we have
- 9:04exactly two possible symbols zero and
- 9:06one and we have a very long sequence of
- 9:10it right now as it turns out um this
- 9:14sequence length is actually going to be
- 9:16very finite and precious resource uh in
- 9:19our neural network and we actually don't
- 9:21want extremely long sequences of just
- 9:23two symbols instead what we want is we
- 9:25want to trade off uh this um symbol
- 9:29size uh of this vocabulary as we call it
- 9:32and the resulting sequence length so we
- 9:35don't want just two symbols and
- 9:36extremely long sequences we're going to
- 9:38want more symbols and shorter sequences
- 9:42okay so one naive way of compressing or
- 9:44decreasing the length of our sequence
- 9:46here is to basically uh consider some
- 9:49group of consecutive bits for example
- 9:51eight bits and group them into a single
- 9:54what's called bite so because uh these
- 9:57bits are either on or off if we take a
- 10:00group of eight of them there turns out
- 10:01to be only 256 possible combinations of
- 10:04how these bits could be on or off and so
- 10:06therefore we can re repesent this
- 10:07sequence into a sequence of bytes
- 10:10instead so this sequence of bytes will
- 10:13be eight times shorter but now we have
- 10:16256 possible symbols so every number
- 10:19here goes from 0 to
- 10:20255 now I really encourage you to think
- 10:22of these not as numbers but as unique
- 10:25IDs or like unique symbols so maybe it's
- 10:28a bit more maybe it's better to actually
- 10:30think of these to replace every one of
- 10:32these with a unique Emoji you'd get
- 10:34something like this so um we basically
- 10:37have a sequence of emojis and there's
- 10:38256 possible emojis you can think of it
- 10:41that way now it turns out that in
- 10:44production for state-of-the-art language
- 10:46models uh you actually want to go even
- 10:48Beyond this you want to continue to
- 10:50shrink the length of the sequence uh
- 10:52because again it is a precious resource
- 10:54in return for more symbols in your
- 10:57vocabulary and the way this is done is
- 11:00done by running what's called The Bite
- 11:02pair encoding algorithm and the way this
- 11:04works is we're basically looking for
- 11:06consecutive bytes or symbols that are
- 11:10very common so for example turns out
- 11:13that the sequence 116 followed by 32 is
- 11:17quite common and occurs very frequently
- 11:19so what we're going to do is we're going
- 11:20to group uh this um pair into a new
- 11:24symbol so we're going to Mint a symbol
- 11:26with an ID 256 and we're going to
- 11:28rewrite every single uh pair 11632 with
- 11:32this new symbol and then can we can
- 11:34iterate this algorithm as many times as
- 11:36we wish and each time when we mint a new
- 11:38symbol we're decreasing the length and
- 11:40we're increasing the symbol size and in
- 11:43practice it turns out that a pretty good
- 11:45setting of um the basically the
- 11:48vocabulary size turns out to be about
- 11:49100,000 possible symbols so in
- 11:52particular GPT 4 uses
- 11:55100,
- 11:56277 symbols
- 11:59um and this process of converting from
- 12:04raw text into these symbols or as we
- 12:07call them tokens is the process called
- 12:10tokenization so let's now take a look at
- 12:12how gp4 performs tokenization conting
- 12:15from text to tokens and from tokens back
- 12:18to text and what this actually looks
- 12:19like so one website I like to use to
- 12:21explore these token representations is
- 12:24called tick tokenizer and so come here
- 12:27to the drop down and select CL 100 a
- 12:29base which is the gp4 base model
- 12:32tokenizer and here on the left you can
- 12:34put in text and it shows you the
- 12:36tokenization of that text so for example
- 12:40heo space
- 12:43world so hello world turns out to be
- 12:46exactly two Tokens The Token hello which
- 12:49is the token with ID
- 12:5115339 and the token space
- 12:54world that is the token 1
- 12:571917 so um hello space world now if I
- 13:02was to join these two for example I'm
- 13:04going to get again two tokens but it's
- 13:06the token H followed by the token L
- 13:09world without the
- 13:11H um if I put in two Spa two spaces here
- 13:15between hello and world it's again a
- 13:16different uh tokenization there's a new
- 13:19token 220
- 13:22here okay so you can play with this and
- 13:24see what happens here also keep in mind
- 13:26this is not uh this is case sensitive so
- 13:28if this is a capital H it is something
- 13:30else or if it's uh hello world then
- 13:35actually this ends up being three tokens
- 13:36instead of just two
- 13:41tokens yeah so you can play with this
- 13:43and get an sort of like an intuitive
- 13:44sense of uh what these tokens work like
- 13:47we're actually going to loop around to
- 13:48tokenization a bit later in the video
- 13:50for now I just wanted to show you the
- 13:51website and I wanted to uh show you that
- 13:53this text basically at the end of the
- 13:56day so for example if I take one line
- 13:57here this is what GT4 will see it as so
- 14:01this text will be a sequence of length
- 14:0462 this is the sequence here and this is
- 14:08how the chunks of text correspond to
- 14:11these symbols and again there's 100,
- 14:1627777 possible symbols and we now have
- 14:19one-dimensional sequences of those
- 14:21symbols so um yeah we're going to come
- 14:24back to tokenization but that's uh for
- 14:26now where we are okay so what I've done
- 14:28now is I've taken this uh sequence of
- 14:30text that we have here in the data set
- 14:32and I have re-represented it using our
- 14:34tokenizer into a sequence of tokens and
- 14:37this is what that looks like now so for
- 14:40example when we go back to the Fine web
- 14:41data set they mentioned that not only is
- 14:43this 44 terab of dis space but this is
- 14:45about a 15 trillion token sequence of um
- 14:50in this data set and so here these are
- 14:53just some of the first uh one or two or
- 14:56three or a few thousand here I think uh
- 14:58tokens of this data set but there's 15
- 15:01trillion here uh to keep in mind and
- 15:03again keep in mind one more time that
- 15:05all of these represent little text
- 15:07chunks they're all just like atoms of
- 15:09these sequences and the numbers here
- 15:11don't make any sense they're just uh
- 15:13they're just unique IDs okay so now we
- 15:17get to the fun part which is the uh
- 15:19neural network training and this is
- 15:21where a lot of the heavy lifting happens
- 15:23computationally when you're training
- 15:24these neural networks so what we do here
- 15:28in this this step is we want to model
- 15:30the statistical relationships of how
- 15:32these tokens follow each other in the
- 15:33sequence so what we do is we come into
- 15:36the data and we take Windows of tokens
- 15:40so we take a window of tokens uh from
- 15:43this data fairly
- 15:44randomly and um the windows length can
- 15:49range anywhere anywhere between uh zero
- 15:51tokens actually all the way up to some
- 15:54maximum size that we decide on uh so for
- 15:57example in practice you could see a
- 15:58token with Windows of say 8,000 tokens
- 16:01now in principle we can use arbitrary
- 16:03window lengths of tokens uh but uh
- 16:07processing very long uh basically U
- 16:10window sequences would just be very
- 16:12computationally expensive so we just
- 16:15kind of decide that say 8,000 is a good
- 16:16number or 4,000 or 16,000 and we crop it
- 16:19there now in this example I'm going to
- 16:22be uh taking the first four tokens just
- 16:25so everything fits nicely so these
- 16:28tokens
- 16:30we're going to take a window of four
- 16:32tokens this bar view in and space single
- 16:37which are these token
- 16:39IDs and now what we're trying to do here
- 16:41is we're trying to basically predict the
- 16:42token that comes next in the sequence so
- 16:453962 comes next right so what we do now
- 16:49here is that we call this the context
- 16:51these four tokens are context and they
- 16:54feed into a neural
- 16:56network and this is the input to the
- 16:58neural network
- 16:59now I'm going to go into the detail of
- 17:01what's inside this neural network in a
- 17:03little bit for now it's important to
- 17:04understand is the input and the output
- 17:06of the neural net so the input are
- 17:08sequences of tokens of variable length
- 17:12anywhere between zero and some maximum
- 17:14size like 8,000 the output now is a
- 17:17prediction for what comes next so
- 17:21because our vocabulary has
- 17:23100277 possible tokens the neural
- 17:26network is going to Output exactly that
- 17:28many numbers
- 17:29and all of those numbers correspond to
- 17:30the probability of that token as coming
- 17:33next in the sequence so it's making
- 17:35guesses about what comes
- 17:37next um in the beginning this neural
- 17:39network is randomly initialized so um
- 17:42and we're going to see in a little bit
- 17:44what that means but it's a it's a it's a
- 17:46random transformation so these
- 17:48probabilities in the very beginning of
- 17:49the training are also going to be kind
- 17:51of random uh so here I have three
- 17:53examples but keep in mind that there's
- 17:55100,000 numbers here um so the
- 17:58probability of this token space
- 17:59Direction neural network is saying that
- 18:01this is 4% likely right now 11799 is 2%
- 18:05and then here the probility of 3962
- 18:08which is post is 3% now of course we've
- 18:11sampled this window from our data set so
- 18:13we know what comes next we know and
- 18:16that's the label we know that the
- 18:18correct answer is that 3962 actually
- 18:19comes next in the sequence so now what
- 18:22we have is this mathematical process for
- 18:25doing an update to the neural network we
- 18:28have the way of tuning it and uh we're
- 18:30going to go into a little bit of of
- 18:32detail in a bit but basically we know
- 18:34that this probability here of 3% we want
- 18:38this probability to be higher and we
- 18:40want the probabilities of all the other
- 18:42tokens to be
- 18:44lower and so we have a way of
- 18:46mathematically calculating how to adjust
- 18:48and update the neural network so that
- 18:51the correct answer has a slightly higher
- 18:53probability so if I do an update to the
- 18:55neural network now the next time I Fe
- 18:59this particular sequence of four tokens
- 19:00into neural network the neural network
- 19:02will be slightly adjusted now and it
- 19:04will say Okay post is maybe 4% and case
- 19:07now maybe is
- 19:081% and uh Direction could become 2% or
- 19:12something like that and so we have a way
- 19:14of nudging of slightly updating the
- 19:16neuronet to um basically give a higher
- 19:19probability to the correct token that
- 19:21comes next in the sequence and now you
- 19:23just have to remember that this process
- 19:25happens not just for uh this um token
- 19:29here where these four fed in and
- 19:31predicted this one this process happens
- 19:33at the same time for all of these tokens
- 19:36in the entire data set and so in
- 19:38practice we sample little windows little
- 19:40batches of Windows and then at every
- 19:42single one of these tokens we want to
- 19:44adjust our neural network so that the
- 19:46probability of that token becomes
- 19:48slightly higher and this all happens in
- 19:50parallel in large batches of these
- 19:52tokens and this is the process of
- 19:54training the neural network it's a
- 19:55sequence of updating it so that it's
- 19:58predictions match up the statistics of
- 20:01what actually happens in your training
- 20:02set and its probabilities become
- 20:05consistent with the uh statistical
- 20:08patterns of how these tokens follow each
- 20:09other in the data so let's now briefly
- 20:12get into the internals of these neural
- 20:13networks just to give you a sense of
- 20:14what's inside so neural network
- 20:17internals so as I mentioned we have
- 20:19these inputs uh that are sequences of
- 20:22tokens in this case this is four input
- 20:24tokens but this can be anywhere between
- 20:26zero up to let's say 8,000 tokens in
- 20:30principle this can be an infinite number
- 20:31of tokens we just uh it would just be
- 20:33too computationally expensive to process
- 20:35an infinite number of tokens so we just
- 20:37crop it at a certain length and that
- 20:39becomes the maximum context length of
- 20:41that uh
- 20:42model now these inputs X are mixed up in
- 20:46a giant mathematical expression together
- 20:48with the parameters or the weights of
- 20:51these neural networks so here I'm
- 20:53showing six example parameters and their
- 20:56setting but in practice these uh um
- 21:00modern neural networks will have
- 21:01billions of these uh parameters and in
- 21:04the beginning these parameters are
- 21:06completely randomly set now with a
- 21:09random setting of parameters you might
- 21:11expect that this uh this neural network
- 21:13would make random predictions and it
- 21:15does in the beginning it's totally
- 21:16random predictions but it's through this
- 21:19process of iteratively updating the
- 21:22network uh as and we call that process
- 21:24training a neural network so uh that the
- 21:28setting of these parameters gets
- 21:29adjusted such that the outputs of our
- 21:31neural network becomes consistent with
- 21:34the patterns seen in our training
- 21:36set so think of these parameters as kind
- 21:39of like knobs on a DJ set and as you're
- 21:41twiddling these knobs you're getting
- 21:42different uh predictions for every
- 21:45possible uh token sequence input and
- 21:49training in neural network just means
- 21:50discovering a setting of parameters that
- 21:52seems to be consistent with the
- 21:54statistics of the training
- 21:56set now let me just give you an example
- 21:58what this giant mathematical expression
- 21:59looks like just to give you a sense and
- 22:01modern networks are massive expressions
- 22:03with trillions of terms probably but let
- 22:06me just show you a simple example here
- 22:08it would look something like this I mean
- 22:10these are the kinds of Expressions just
- 22:11to show you that it's not very scary we
- 22:13have inputs x uh like X1 x2 in this case
- 22:17two example inputs and they get mixed up
- 22:19with the weights of the network w0 W1 2
- 22:223 Etc and this mixing is simple things
- 22:27like multiplication addition addition
- 22:29exponentiation division Etc and it is
- 22:32the subject of neural network
- 22:34architecture research to design
- 22:36effective mathematical Expressions uh
- 22:39that have a lot of uh kind of convenient
- 22:41characteristics they are expressive
- 22:42they're optimizable they're paralyzable
- 22:45Etc and so but uh at the end of the day
- 22:48these are these are not complex
- 22:49expressions and basically they mix up
- 22:52the inputs with the parameters to make
- 22:54predictions and we're optimizing uh the
- 22:57parameters of this neural network so
- 22:59that the predictions come out consistent
- 23:01with the training set now I would like
- 23:04to show you an actual production grade
- 23:06example of what these neural networks
- 23:07look like so for that I encourage you to
- 23:09go to this website that has a very nice
- 23:11visualization of one of these
- 23:13networks so this is what you will find
- 23:16on this website and this neural network
- 23:19here that is used in production settings
- 23:21has this special kind of structure this
- 23:24network is called the Transformer and
- 23:26this particular one as an example has 8
- 23:285,000 roughly
- 23:30parameters now here on the top we take
- 23:33the inputs which are the token
- 23:36sequences and then information flows
- 23:39through the neural network until the
- 23:41output which here are the logit softmax
- 23:45but these are the predictions for what
- 23:46comes next what token comes
- 23:48next and then here there's a sequence of
- 23:52Transformations and all these
- 23:54intermediate values that get produced
- 23:56inside this mathematical expression s it
- 23:58is sort of predicting what comes next so
- 24:01as an example these tokens are embedded
- 24:04into kind of like this distributed
- 24:06representation as it's called so every
- 24:08possible token has kind of like a vector
- 24:10that represents it inside the neural
- 24:11network so first we embed the tokens and
- 24:15then those values uh kind of like flow
- 24:18through this diagram and these are all
- 24:20very simple mathematical Expressions
- 24:22individually so we have layer norms and
- 24:24Matrix multiplications and uh soft Maxes
- 24:27and so on so here kind of like the
- 24:28attention block of this Transformer and
- 24:31then information kind of flows through
- 24:33into the multi-layer perceptron block
- 24:35and so on and all these numbers here
- 24:38these are the intermediate values of the
- 24:40expression and uh you can almost think
- 24:42of these as kind of like the firing
- 24:44rates of these synthetic neurons but I
- 24:47would caution you to uh not um kind of
- 24:50think of it too much like neurons
- 24:52because these are extremely simple
- 24:53neurons compared to the neurons you
- 24:55would find in your brain your biological
- 24:57neurons are very complex dynamical
- 24:59processes that have memory and so on
- 25:01there's no memory in this expression
- 25:02it's a fixed mathematical expression
- 25:04from input to Output with no memory it's
- 25:06just a
- 25:07stateless so these are very simple
- 25:09neurons in comparison to biological
- 25:10neurons but you can still kind of
- 25:12loosely think of this as like a
- 25:13synthetic piece of uh brain tissue if
- 25:15you if you like uh to think about it
- 25:17that way so information flows through
- 25:21all these neurons fire until we get to
- 25:24the predictions now I'm not actually
- 25:26going to dwell too much on the precise
- 25:28kind of like mathematical details of all
- 25:30these Transformations honestly I don't
- 25:31think it's that important to get into
- 25:33what's really important to understand is
- 25:35that this is a mathematical function it
- 25:38is uh parameterized by some fixed set of
- 25:41parameters like say 85,000 of them and
- 25:44it is a way of transforming inputs into
- 25:46outputs and as we twiddle the parameters
- 25:48we are getting uh different kinds of
- 25:50predictions and then we need to find a
- 25:52good setting of these parameters so that
- 25:54the predictions uh sort of match up with
- 25:56the patterns seen in training set
- 25:59so that's the Transformer okay so I've
- 26:02shown you the internals of the neural
- 26:03network and we talked a bit about the
- 26:05process of training it I want to cover
- 26:07one more major stage of working with
- 26:10these networks and that is the stage
- 26:11called inference so in inference what
- 26:14we're doing is we're generating new data
- 26:16from the model and so uh we want to
- 26:18basically see what kind of patterns it
- 26:21has internalized in the parameters of
- 26:23its Network so to generate from the
- 26:26model is relatively straightforward
- 26:28we start with some tokens that are
- 26:30basically your prefix like what you want
- 26:32to start with so say we want to start
- 26:34with the token 91 well we feed it into
- 26:37the
- 26:37network and remember that the network
- 26:39gives us probabilities right it gives us
- 26:43this probability Vector here so what we
- 26:45can do now is we can basically flip a
- 26:47biased coin so um we can sample uh
- 26:52basically a token based on this
- 26:54probability distribution so the tokens
- 26:57that are given High probability by the
- 26:59model are more likely to be sampled when
- 27:01you flip this biased coin you can think
- 27:03of it that way so we sample from the
- 27:05distribution to get a single unique
- 27:08token so for example token 860 comes
- 27:11next uh so 860 in this case when we're
- 27:14generating from model could come next
- 27:16now 860 is a relatively likely token it
- 27:18might not be the only possible token in
- 27:20this case there could be many other
- 27:21tokens that could have been sampled but
- 27:23we could see that 86c is a relatively
- 27:25likely token as an example and indeed in
- 27:27our training examp example here 860 does
- 27:29follow 91 so let's now say that we um
- 27:34continue the process so after 91 there's
- 27:36a60 we append it and we again ask what
- 27:39is the third token let's sample and
- 27:42let's just say that it's 287 exactly as
- 27:44here let's do that again we come back in
- 27:47now we have a sequence of three and we
- 27:49ask what is the likely fourth token and
- 27:52we sample from that and get this one and
- 27:55now let's say we do it one more time we
- 27:58take those four we sample and we get
- 28:00this one and this
- 28:0213659 uh this is not actually uh 3962 as
- 28:06we had before so this token is the token
- 28:09article uh instead so viewing a single
- 28:12article and so in this case we didn't
- 28:15exactly reproduce the sequence that we
- 28:17saw here in the training data so keep in
- 28:20mind that these systems are stochastic
- 28:22they have um we're sampling and we're
- 28:25flipping coins and sometimes we lock out
- 28:28and we reproduce some like small chunk
- 28:30of the text and training set but
- 28:32sometimes we're uh we're getting a token
- 28:35that was not verbatim part of any of the
- 28:38documents in the training data so we're
- 28:40going to get sort of like remixes of the
- 28:43data that we saw in the training because
- 28:44at every step of the way we can flip and
- 28:47get a slightly different token and then
- 28:48once that token makes it in if you
- 28:50sample the next one and so on you very
- 28:52quickly uh start to generate token
- 28:55streams that are very different from the
- 28:57token streams that UR
- 28:58in the training documents so
- 29:00statistically they will have similar
- 29:02properties but um they are not identical
- 29:05to your training data they're kind of
- 29:06like inspired by the training data and
- 29:09so in this case we got a slightly
- 29:10different sequence and why would we get
- 29:12article you might imagine that article
- 29:14is a relatively likely token in the
- 29:16context of bar viewing single Etc and
- 29:21you can imagine that the word article
- 29:22followed this context window somewhere
- 29:24in the training documents uh to some
- 29:26extent and we just happen to sample it
- 29:28here at that stage so basically
- 29:31inference is just uh predicting from
- 29:33these distributions one at a time we
- 29:35continue feeding back tokens and getting
- 29:37the next one and we uh we're always
- 29:39flipping these coins and depending on
- 29:42how lucky or unlucky we get um we might
- 29:45get very different kinds of patterns
- 29:47depending on how we sample from these
- 29:49probability distributions so that's
- 29:51inference so in most common scenarios uh
- 29:55basically downloading the internet and
- 29:57tokenizing it is is a pre-processing
- 29:58step you do that a single time and then
- 30:01uh once you have your token sequence we
- 30:04can start training networks and in
- 30:06Practical cases you would try to train
- 30:08many different networks of different
- 30:10kinds of uh settings and different kinds
- 30:11of arrangements and different kinds of
- 30:13sizes and so you''ll be doing a lot of
- 30:15neural network training and um then once
- 30:18you have a neural network and you train
- 30:19it and you have some specific set of
- 30:21parameters that you're happy with um
- 30:24then you can take the model and you can
- 30:25do inference and you can actually uh
- 30:28generate data from the model and when
- 30:30you're on chat GPT and you're talking
- 30:31with a model uh that model is trained
- 30:33and has been trained by open aai many
- 30:36months ago probably and they have a
- 30:38specific set of Weights that work well
- 30:41and when you're talking to the model all
- 30:42of that is just inference there's no
- 30:44more training those parameters are held
- 30:47fixed and you're just talking to the
- 30:49model sort of uh you're giving it some
- 30:51of the tokens and it's kind of
- 30:53completing token sequences and that's
- 30:54what you're seeing uh generated when you
- 30:57actually use the model on CH GPT so that
- 30:59model then just does inference alone so
- 31:02let's now look at an example of training
- 31:04an inference that is kind of concrete
- 31:05and gives you a sense of what this
- 31:07actually looks like uh when these models
- 31:08are trained now the example that I would
- 31:10like to work with and that I'm
- 31:12particularly fond of is that of opening
- 31:14eyes gpt2 so GPT uh stands for
- 31:17generatively pre-trained Transformer and
- 31:19this is the second iteration of the GPT
- 31:21series by open AI when you are talking
- 31:23to chat GPT today the model that is
- 31:26underlying all of the magic of that
- 31:27interaction is GPT 4 so the fourth
- 31:30iteration of that series now gpt2 was
- 31:33published in 2019 by openi in this paper
- 31:36that I have right here and the reason I
- 31:39like gpt2 is that it is the first time
- 31:41that a recognizably modern stack came
- 31:44together so um all of the pieces of gpd2
- 31:48are recognizable today by modern
- 31:50standards it's just everything has
- 31:52gotten bigger now I'm not going to be
- 31:54able to go into the full details of this
- 31:55paper of course because it is a
- 31:57technical publication but some of the
- 31:59details that I would like to highlight
- 32:00are as follows gpt2 was a Transformer
- 32:03neural network just like you were just
- 32:05like the neural networks you would work
- 32:06with today it was it had 1.6 billion
- 32:10parameters right so these are the
- 32:12parameters that we looked at here it
- 32:14would have 1.6 billion of them today
- 32:16modern Transformers would have a lot
- 32:18closer to a trillion or several hundred
- 32:20billion
- 32:21probably the maximum context length here
- 32:24was 1,24 tokens so it is when we are
- 32:28sampling chunks of Windows of tokens
- 32:32from the data set we're never taking
- 32:34more than 1,24 tokens and so when you
- 32:36are trying to predict the next token in
- 32:38a sequence you will never have more than
- 32:401,24 tokens uh kind of in your context
- 32:43in order to make that prediction now
- 32:45this is also tiny by modern standards
- 32:47today the token uh the context lengths
- 32:49would be a lot closer to um couple
- 32:53hundred thousand or maybe even a million
- 32:55and so you have a lot more context a lot
- 32:56more tokens in history history and you
- 32:58can make a lot better prediction about
- 33:00the next token in the sequence in that
- 33:01way and finally gpt2 was trained on
- 33:04approximately 100 billion tokens and
- 33:06this is also fairly small by modern
- 33:08standards as I mentioned the fine web
- 33:10data set that we looked at here the fine
- 33:12web data set has 15 trillion tokens uh
- 33:14so 100 billion is is quite
- 33:16small
- 33:18now uh I actually tried to reproduce uh
- 33:21gpt2 for fun as part of this project
- 33:23called lm. C so you can see my rup of
- 33:27doing that in this post on GitHub under
- 33:30the lm. C repository so in particular
- 33:33the cost of training gpd2 in 2019 what
- 33:36was estimated to be approximately
- 33:39$40,000 but today you can do
- 33:41significantly better than that and in
- 33:42particular here it took about one day
- 33:45and about
- 33:47$600 uh but this wasn't even trying too
- 33:49hard I think you could really bring this
- 33:51down to about $100 today now why is it
- 33:55that the costs have come down so much
- 33:57well number one these data sets have
- 33:59gotten a lot better and the way we
- 34:01filter them extract them and prepare
- 34:03them has gotten a lot more refined and
- 34:05so the data set is of just a lot higher
- 34:08quality so that's one thing but really
- 34:10the biggest difference is that our
- 34:11computers have gotten much faster in
- 34:13terms of the hardware and we're going to
- 34:15look at that in a second and also the
- 34:17software for uh running these models and
- 34:20really squeezing out all all the speed
- 34:22from the hardware as it is possible uh
- 34:25that software has also gotten much
- 34:27better as as everyone has focused on
- 34:28these models and try to run them very
- 34:30very
- 34:31quickly now I'm not going to be able to
- 34:34go into the full detail of this gpd2
- 34:36reproduction and this is a long
- 34:37technical post but I would like to still
- 34:39give you an intuitive sense for what it
- 34:41looks like to actually train one of
- 34:43these models as a researcher like what
- 34:44are you looking at and what does it look
- 34:46like what does it feel like so let me
- 34:47give you a sense of that a little bit
- 34:50okay so this is what it looks like let
- 34:51me slide this
- 34:52over so what I'm doing here is I'm
- 34:55training a gpt2 model right now
- 34:58and um what's happening here is that
- 35:00every single line here like this one is
- 35:05one update to the model so remember how
- 35:08here we are um basically making the
- 35:12prediction better for every one of these
- 35:14tokens and we are updating these weights
- 35:15or parameters of the neural net so here
- 35:18every single line is One update to the
- 35:20neural network where we change its
- 35:22parameters by a little bit so that it is
- 35:24better at predicting next token and
- 35:26sequence in particular every single line
- 35:28here is improving the prediction on 1
- 35:32million tokens in the training set so
- 35:35we've basically taken 1 million tokens
- 35:39out of this data set and we've tried to
- 35:41improve the prediction of that token as
- 35:44coming next in a sequence on all 1
- 35:46million of them
- 35:49simultaneously and at every single one
- 35:51of these steps we are making an update
- 35:52to the network for that now the number
- 35:55to watch closely is this number called
- 35:57loss and the loss is a single number
- 36:00that is telling you how well your neural
- 36:02network is performing right now and it
- 36:05is created so that low loss is good so
- 36:08you'll see that the loss is decreasing
- 36:10as we make more updates to the neural
- 36:12nut which corresponds to making better
- 36:14predictions on the next token in a
- 36:16sequence and so the loss is the number
- 36:19that you are watching as a neural
- 36:20network researcher and you are kind of
- 36:22waiting you're twiddling your thumbs uh
- 36:24you're drinking coffee and you're making
- 36:26sure that this looks good so that with
- 36:28every update your loss is improving and
- 36:31the network is getting better at
- 36:32prediction now here you see that we are
- 36:36processing 1 million tokens per update
- 36:38each update takes about 7 Seconds
- 36:41roughly and here we are going to process
- 36:43a total of 32,000 steps of
- 36:47optimization so 32,000 steps with 1
- 36:50million tokens each is about 33 billion
- 36:52tokens that we are going to process and
- 36:54we're currently only about 420 step 20
- 36:57out of 32,000 so we are still only a bit
- 37:01more than 1% done because I've only been
- 37:03running this for 10 or 15 minutes or
- 37:05something like
- 37:06that now every 20 steps I have
- 37:09configured this optimization to do
- 37:11inference so what you're seeing here is
- 37:13the model is predicting the next token
- 37:15in a sequence and so you sort of start
- 37:17it randomly and then you continue
- 37:19plugging in the tokens so we're running
- 37:21this inference step and this is the
- 37:23model sort of predicting the next token
- 37:25in the sequence and every time you see
- 37:26something appear that's a new
- 37:29token um so let's just look at this and
- 37:34you can see that this is not yet very
- 37:35coherent and keep in mind that this is
- 37:37only 1% of the way through training and
- 37:39so the model is not yet very good at
- 37:41predicting the next token in the
- 37:42sequence so what comes out is actually
- 37:44kind of a little bit of gibberish right
- 37:47but it still has a little bit of like
- 37:48local coherence so since she is mine
- 37:51it's a part of the information should
- 37:53discuss my father great companions
- 37:55Gordon showed me sitting over at and Etc
- 37:59so I know it doesn't look very good but
- 38:00let's actually scroll up and see what it
- 38:04looked like when I started the
- 38:06optimization so all the way here at
- 38:10step
- 38:12one so after 20 steps of optimization
- 38:15you see that what we're getting here is
- 38:17looks completely random and of course
- 38:18that's because the model has only had 20
- 38:20updates to its parameters and so it's
- 38:22giving you random text because it's a
- 38:23random Network and so you can see that
- 38:25at least in comparison to this model is
- 38:27starting to do much better and indeed if
- 38:29we waited the entire 32,000 steps the
- 38:32model will have improved the point that
- 38:34it's actually uh generating fairly
- 38:36coherent English uh and the tokens
- 38:38stream correctly um and uh they they
- 38:42kind of make up English a a lot
- 38:44better
- 38:46um so this has to run for about a day or
- 38:49two more now and so uh at this stage we
- 38:52just make sure that the loss is
- 38:53decreasing everything is looking good um
- 38:56and we just have to wait
- 38:58and now um let me turn now to the um
- 39:02story of the computation that's required
- 39:05because of course I'm not running this
- 39:06optimization on my laptop that would be
- 39:08way too expensive uh because we have to
- 39:11run this neural network and we have to
- 39:12improve it and we have we need all this
- 39:14data and so on so you can't run this too
- 39:16well on your computer uh because the
- 39:18network is just too large uh so all of
- 39:21this is running on the computer that is
- 39:23out there in the cloud and I want to
- 39:25basically address the compute side of
- 39:27the store of training these models and
- 39:28what that looks like so let's take a
- 39:30look okay so the computer that I'm
- 39:32running this optimization on is this 8X
- 39:35h100 node so there are eight h100s in a
- 39:39single node or a single computer now I
- 39:42am renting this computer and it is
- 39:44somewhere in the cloud I'm not sure
- 39:45where it is physically actually the
- 39:47place I like to rent from is called
- 39:49Lambda but there are many other
- 39:50companies who provide this service so
- 39:52when you scroll down you can see that uh
- 39:55they have some on demand pricing for
- 39:57um sort of computers that have these uh
- 40:01h100s which are gpus and I'm going to
- 40:03show you what they look like in a second
- 40:06but on demand 8times Nvidia h100 uh
- 40:10GPU this machine comes for $3 per GPU
- 40:13per hour for example so you can rent
- 40:16these and then you get a machine in a
- 40:18cloud and you can uh go in and you can
- 40:20train these
- 40:21models and these uh gpus they look like
- 40:25this so this is one h100 GPU uh this is
- 40:29kind of what it looks like and you slot
- 40:30this into your computer and gpus are
- 40:32this uh perfect fit for training your
- 40:34networks because they are very
- 40:36computationally expensive but they
- 40:38display a lot of parallelism in the
- 40:40computation so you can have many
- 40:42independent workers kind of um working
- 40:44all at the same time in solving uh the
- 40:48matrix multiplication that's under the
- 40:50hood of training these neural
- 40:52networks so this is just one of these
- 40:54h100s but actually you would put them
- 40:56you would put multiple of them together
- 40:58so you could stack eight of them into a
- 41:00single node and then you can stack
- 41:02multiple nodes into an entire data
- 41:04center or an entire system
- 41:07so when we look at a data
- 41:12center can't spell when we look at a
- 41:15data center we start to see things that
- 41:16look like this right so we have one GPU
- 41:18goes to eight gpus goes to a single
- 41:19system goes to many systems and so these
- 41:22are the bigger data centers and there of
- 41:23course would be much much more expensive
- 41:26um and what's happening is that all the
- 41:28big tech companies really desire these
- 41:31gpus so they can train all these
- 41:33language models because they are so
- 41:35powerful and that has is fundamentally
- 41:37what has driven the stock price of
- 41:38Nvidia to be $3.4 trillion today as an
- 41:41example and why Nvidia has kind of
- 41:44exploded so this is the Gold Rush the
- 41:47Gold Rush is getting the gpus getting
- 41:50enough of them so they can all
- 41:52collaborate to perform this optimization
- 41:55and they're what are they all doing
- 41:56they're all collaborating to predict the
- 41:59next token on a data set like the fine
- 42:01web data
- 42:02set this is the computational workflow
- 42:05that that basically is extremely
- 42:06expensive the more gpus you have the
- 42:09more tokens you can try to predict and
- 42:10improve on and you're going to process
- 42:12this data set faster and you can iterate
- 42:15faster and get a bigger Network and
- 42:16train a bigger Network and so on so this
- 42:19is what all those machines are look like
- 42:20are uh are doing and this is why all of
- 42:24this is such a big deal and for example
- 42:26this is a
- 42:28article from like about a month ago or
- 42:30so this is why it's a big deal that for
- 42:31example Elon Musk is getting 100,000
- 42:34gpus uh in a single Data Center and all
- 42:38of these gpus are extremely expensive
- 42:40are going to take a ton of power and all
- 42:42of them are just trying to predict the
- 42:43next token in the sequence and improve
- 42:45the network uh by doing so and uh get
- 42:49probably a lot more coherent text than
- 42:50what we're seeing here a lot faster okay
- 42:52so unfortunately I do not have a couple
- 42:5510 or hundred million of dollars to
- 42:57spend on training a really big model
- 42:59like this but luckily we can turn to
- 43:01some big tech companies who train these
- 43:04models routinely and release some of
- 43:06them once they are done training so
- 43:08they've spent a huge amount of compute
- 43:10to train this network and they release
- 43:12the network at the end of the
- 43:13optimization so it's very useful because
- 43:15they've done a lot of compute for that
- 43:18so there are many companies who train
- 43:19these models routinely but actually not
- 43:21many of them release uh these what's
- 43:23called base models so the model that
- 43:25comes out at the end here is is what's
- 43:27called a base model what is a base model
- 43:29it's a token simulator right it's an
- 43:32internet text token simulator and so
- 43:35that is not by itself useful yet because
- 43:38what we want is what's called an
- 43:39assistant we want to ask questions and
- 43:41have it respond to answers these models
- 43:43won't do that they just uh create sort
- 43:45of remixes of the internet they dream
- 43:48internet pages so the base models are
- 43:51not very often released because they're
- 43:52kind of just only a step one of a few
- 43:55other steps that we still need to take
- 43:56to get in system
- 43:58however a few releases have been made so
- 44:01as an example the gbt2 model released
- 44:04the 1.6 billion sorry 1.5 billion model
- 44:08back in 2019 and this gpt2 model is a
- 44:10base model now what is a model release
- 44:13what does it look like to release these
- 44:15models so this is the gpt2 repository on
- 44:18GitHub well you need two things
- 44:20basically to release model number one we
- 44:22need the um python code usually that
- 44:27describes the sequence of operations in
- 44:30detail that they make in their model so
- 44:34um if you remember
- 44:36back this
- 44:38Transformer the sequence of steps that
- 44:40are taken here in this neural network is
- 44:42what is being described by this code so
- 44:45this code is sort of implementing the
- 44:47what's called forward pass of this
- 44:49neural network so we need the specific
- 44:51details of exactly how they wired up
- 44:53that neural network so this is just
- 44:55computer code and it's usually just a
- 44:57couple hundred lines of code it's not
- 44:59it's not that crazy and uh this is all
- 45:01fairly understandable and usually fairly
- 45:03standard what's not standard are the
- 45:05parameters that's where the actual value
- 45:07is what are the parameters of this
- 45:09neural network because there's 1.6
- 45:11billion of them and we need the correct
- 45:13setting or a really good setting and so
- 45:15that's why in addition to this source
- 45:17code they release the parameters which
- 45:20in this case is roughly 1.5 billion
- 45:23parameters and these are just numbers so
- 45:25it's one single list of 1.5 billion
- 45:27numbers the precise and good setting of
- 45:30all the knobs such that the tokens come
- 45:32out
- 45:33well so uh you need those two things to
- 45:37get a base model
- 45:39release
- 45:41now gpt2 was released but that's
- 45:43actually a fairly old model as I
- 45:44mentioned so actually the model we're
- 45:46going to turn to is called llama 3 and
- 45:49that's the one that I would like to show
- 45:50you next so llama 3 so gpt2 again was
- 45:541.6 billion parameters trained on 100
- 45:55billion tokens Lama 3 is a much bigger
- 45:58model and much more modern model it is
- 46:00released and trained by meta and it is a
- 46:0345 billion parameter model trained on 15
- 46:07trillion tokens in very much the same
- 46:09way just much much
- 46:11bigger um and meta has also made a
- 46:14release of llama 3 and that was part of
- 46:18this
- 46:19paper so with this paper that goes into
- 46:21a lot of detail the biggest base model
- 46:23that they released is the Lama 3.1 4.5
- 46:27405 billion parameter model so this is
- 46:30the base model and then in addition to
- 46:32the base model you see here
- 46:33foreshadowing for later sections of the
- 46:35video they also released the instruct
- 46:37model and the instruct means that this
- 46:39is an assistant you can ask it questions
- 46:41and it will give you answers we still
- 46:43have yet to cover that part later for
- 46:45now let's just look at this base model
- 46:47this token simulator and let's play with
- 46:49it and try to think about you know what
- 46:51is this thing and how does it work and
- 46:53um what do we get at the end of this
- 46:55optimization if you let this run Until
- 46:57the End uh for a very big neural network
- 46:59on a lot of data so my favorite place to
- 47:02interact with the base models is this um
- 47:04company called hyperbolic which is
- 47:06basically serving the base model of the
- 47:09405b Llama 3.1 so when you go to the
- 47:13website and I think you may have to
- 47:14register and so on make sure that in the
- 47:16models make sure that you are using
- 47:18llama 3.1 405 billion base it must be
- 47:22the base model and then here let's say
- 47:24the max tokens is how many tokens we're
- 47:26going to be gener rating so let's just
- 47:28decrease this to be a bit less just so
- 47:30we don't waste compute we just want the
- 47:32next 128 tokens and leave the other
- 47:34stuff alone I'm not going to go into the
- 47:36full detail here um now fundamentally
- 47:39what's going to happen here is identical
- 47:41to what happens here during inference
- 47:43for us so this is just going to continue
- 47:45the token sequence of whatever you
- 47:47prefix you're going to give it so I want
- 47:49to first show you that this model here
- 47:51is not yet an assistant so you can for
- 47:53example ask it what is 2 plus 2 it's not
- 47:56going to tell you oh it's four uh what
- 47:58else can I help you with it's not going
- 47:59to do that because what is 2 plus 2 is
- 48:02going to be tokenized and then those
- 48:05tokens just act as a prefix and then
- 48:07what the model is going to do now is
- 48:09just going to get the probability for
- 48:10the next token and it's just a glorified
- 48:12autocomplete it's a very very expensive
- 48:14autocomplete of what comes next um
- 48:17depending on the statistics of what it
- 48:18saw in its training documents which are
- 48:20basically web
- 48:22pages so let's just uh hit enter to see
- 48:25what tokens it comes up with as a
- 48:31continuation okay so here it kind of
- 48:32actually answered the question and
- 48:34started to go off into some
- 48:35philosophical territory uh let's try it
- 48:37again so let me copy and paste and let's
- 48:39try again from scratch what is 2 plus
- 48:45two so okay so it just goes off again so
- 48:49notice one more thing that I want to
- 48:50stress is that the system uh I think
- 48:53every time you put it in it just kind of
- 48:55starts from scratch
- 48:58so it doesn't uh the system here is
- 48:59stochastic so for the same prefix of
- 49:02tokens we're always getting a different
- 49:04answer and the reason for that is that
- 49:06we get this probity distribution and we
- 49:08sample from it and we always get
- 49:10different samples and we sort of always
- 49:11go into a different territory uh
- 49:13afterwards so here in this case um I
- 49:17don't know what this is let's try one
- 49:19more
- 49:22time so it just continues on so it's
- 49:25just doing the stuff that it's saw on
- 49:26the internet right um and it's just kind
- 49:29of like regurgitating those uh
- 49:31statistical
- 49:32patterns so first things it's not an
- 49:35assistant yet it's a token autocomplete
- 49:38and second it is a stochastic system now
- 49:42the crucial thing is that even though
- 49:44this model is not yet by itself very
- 49:46useful for a lot of applications just
- 49:49yet um it is still very useful because
- 49:52in the task of predicting the next token
- 49:54in the sequence the model has learned a
- 49:56lot about the world and it has stored
- 49:59all that knowledge in the parameters of
- 50:01the network so remember that our text
- 50:04looked like this right internet web
- 50:06pages and now all of this is sort of
- 50:08compressed in the weights of the network
- 50:11so you can think of um these 405 billion
- 50:15parameters is a kind of compression of
- 50:16the internet you can think of the
- 50:1945 billion parameters is kind of like a
- 50:21zip file uh but it's not a loss less
- 50:25compression it's a loss C compression
- 50:27we're kind of like left with kind of a
- 50:28gal of the internet and we can generate
- 50:31from it right now we can elicit some of
- 50:34this knowledge by prompting the base
- 50:35model uh accordingly so for example
- 50:38here's a prompt that might work to
- 50:40elicit some of that knowledge that's
- 50:41hiding in the parameters here's my top
- 50:4310 list of the top landmarks to see in
- 50:46the
- 50:48pairs
- 50:50um and I'm doing it this way because I'm
- 50:52trying to Prime the model to now
- 50:54continue this list so let's see if that
- 50:56works when I press
- 50:57enter okay so you see that it started a
- 51:00list and it's now kind of giving me some
- 51:02of those
- 51:03landmarks and now notice that it's
- 51:05trying to give a lot of information here
- 51:07now you might not be able to actually
- 51:09fully trust some of the information here
- 51:10remember that this is all just a
- 51:12recollection of some of the internet
- 51:14documents and so the things that occur
- 51:17very frequently in the internet data are
- 51:19probably more likely to be remembered
- 51:21correctly compared to things that happen
- 51:23very infrequently so you can't fully
- 51:25trust some of the things that and some
- 51:27of the information that is here because
- 51:28it's all just a vague recollection of
- 51:30Internet documents because the
- 51:32information is not stored explicitly in
- 51:34any of the parameters it's all just the
- 51:36recollection that said we did get
- 51:38something that is probably approximately
- 51:40correct and I don't actually have the
- 51:42expertise to verify that this is roughly
- 51:44correct but you see that we've elicited
- 51:46a lot of the knowledge of the model and
- 51:48this knowledge is not precise and exact
- 51:51this knowledge is vague and
- 51:53probabilistic and statistical and the
- 51:55kinds of things that occur often are the
- 51:57kinds of things that are more likely to
- 51:59be remembered um in the model now I want
- 52:02to show you a few more examples of this
- 52:04model's Behavior the first thing I want
- 52:05to show you is this example I went to
- 52:08the Wikipedia page for zebra and let me
- 52:10just copy paste the first uh even one
- 52:13sentence
- 52:14here and let me put it here now when I
- 52:17click enter what kind of uh completion
- 52:19are we going to get so let me just hit
- 52:23enter there are three living species
- 52:26etc etc what the model is producing here
- 52:29is an exact regurgitation of this
- 52:31Wikipedia entry it is reciting this
- 52:33Wikipedia entry purely from memory and
- 52:36this memory is stored in its parameters
- 52:39and so it is possible that at some point
- 52:41in these 512 tokens the model will uh
- 52:44stray away from the Wikipedia entry but
- 52:46you can see that it has huge chunks of
- 52:47it memorized here uh let me see for
- 52:50example if this sentence
- 52:51occurs by now okay so this so we're
- 52:55still on track let me check
- 52:58here okay we're still on
- 53:00track it will eventually uh stray
- 53:04away okay so this thing is just recited
- 53:07to a very large extent it will
- 53:08eventually deviate uh because it won't
- 53:11be able to remember exactly now the
- 53:13reason that this happens is because
- 53:14these models can be extremely good at
- 53:16memorization and usually this is not
- 53:18what you want in the final model and
- 53:20this is something called regurgitation
- 53:21and it's usually undesirable to site uh
- 53:24things uh directly uh that you have
- 53:26trained on now the reason that this
- 53:29happens actually is because for a lot of
- 53:31documents like for example Wikipedia
- 53:33when these documents are deemed to be of
- 53:35very high quality as a source like for
- 53:37example Wikipedia it is very often uh
- 53:40the case that when you train the model
- 53:42you will preferentially sample from
- 53:44those sources so basically the model has
- 53:46probably done a few epochs on this data
- 53:48meaning that it has seen this web page
- 53:50like maybe probably 10 times or so and
- 53:52it's a bit like you like when you read
- 53:54some kind of a text many many times say
- 53:56you read something a 100 times uh then
- 53:58you'll be able to recite it and it's
- 54:00very similar for this model if it sees
- 54:01something way too often it's going to be
- 54:03able to recite it later from memory
- 54:05except these models can be a lot more
- 54:07efficient um like per presentation than
- 54:10human so probably it's only seen this
- 54:12Wikipedia entry 10 times but basically
- 54:14it has remembered this article exactly
- 54:16in its parameters okay the next thing I
- 54:18want to show you is something that the
- 54:19model has definitely not seen during its
- 54:21training so for example if we go to the
- 54:24paper uh and then we navigate to the
- 54:26pre-training data we'll see here that uh
- 54:31the data set has a knowledge cut off
- 54:33until the end of 2023 so it will not
- 54:35have seen documents after this point and
- 54:38certainly it has not seen anything about
- 54:39the 2024 election and how it turned out
- 54:43now if we Prime the model with the
- 54:46tokens from the future it will continue
- 54:49the token sequence and it will just take
- 54:50its best guess according to the
- 54:51knowledge that it has in its own
- 54:53parameters so let's take a look at what
- 54:55that could look like
- 54:57so the Republican Party kit
- 54:59Trump okay president of the United
- 55:01States from
- 55:022017 and let's see what it says after
- 55:05this point so for example the model will
- 55:07have to guess at the running mate and
- 55:09who it's against Etc so let's hit
- 55:11enter so here thingss that Mike Pence
- 55:14was the running mate instead of JD Vance
- 55:17and the ticket was against Hillary
- 55:20Clinton and Tim Kane so this is kind of
- 55:23a interesting parallel universe
- 55:25potentially of what could have happened
- 55:26happened according to the LM let's get a
- 55:28different sample so the identical prompt
- 55:31and let's
- 55:33resample so here the running mate was
- 55:35Ronda santis and they ran against Joe
- 55:38Biden and Camala Harris so this is again
- 55:40a different parallel universe so the
- 55:42model will take educated guesses and it
- 55:44will continue the token sequence based
- 55:45on this knowledge um and it will just
- 55:48kind of like all of what we're seeing
- 55:49here is what's called hallucination the
- 55:51model is just taking its best guess uh
- 55:54in a probalistic manner the next thing I
- 55:56would like to show you is that even
- 55:58though this is a base model and not yet
- 56:00an assistant model it can still be
- 56:02utilized in Practical applications if
- 56:04you are clever with your prompt design
- 56:06so here's something that we would call a
- 56:08few shot
- 56:09prompt so what it is here is that I have
- 56:1210 words or 10 pairs and each pair is a
- 56:16word of English column and then a the
- 56:19translation in Korean and we have 10 of
- 56:22them and what the model does here is at
- 56:25the end we have teacher column and then
- 56:27here's where we're going to do a
- 56:28completion of say just five tokens and
- 56:31these models have what we call in
- 56:33context learning abilities and what
- 56:35that's referring to is that as it is
- 56:37reading this context it is learning sort
- 56:40of in
- 56:41place that there's some kind of a
- 56:43algorithmic pattern going on in my data
- 56:46and it knows to continue that pattern
- 56:48and this is called kind of like Inc
- 56:50context learning so it takes on the role
- 56:53of a
- 56:54translator and when we hit uh completion
- 56:58we see that the teacher translation is
- 56:59Sim which is correct um and so this is
- 57:03how you can build apps by being clever
- 57:05with your prompting even though we still
- 57:06just have a base model for now and it
- 57:08relies on what we call this um uh in
- 57:11context learning ability and it is done
- 57:14by constructing what's called a few shot
- 57:15prompt okay and finally I want to show
- 57:17you that there is a clever way to
- 57:19actually instantiate a whole language
- 57:21model assistant just by prompting and
- 57:24the trick to it is that we're structure
- 57:26a prompt to look like a web page that is
- 57:29a conversation between a helpful AI
- 57:31assistant and a human and then the model
- 57:34will continue that conversation so
- 57:36actually to write the prompt I turned to
- 57:38chat gbt itself which is kind of meta
- 57:41but I told it I want to create an llm
- 57:43assistant but all I have is the base
- 57:45model so can you please write my um uh
- 57:50prompt and this is what it came up with
- 57:52which is actually quite good so here's a
- 57:54conversation between an AI assistant and
- 57:55a human
- 57:56the AI assistant is knowledgeable
- 57:58helpful capable of answering wide
- 57:59variety of questions Etc and then here
- 58:03it's not enough to just give it a sort
- 58:05of description it works much better if
- 58:07you create this fot prompt so here's a
- 58:10few terms of human assistant human
- 58:13assistant and we have uh you know a few
- 58:15turns of conversation and then here at
- 58:17the end is we're going to be putting the
- 58:19actual query that we like so let me copy
- 58:21paste this into the base model prompt
- 58:25and now let me do human column and this
- 58:28is where we put our actual prompt why is
- 58:31the sky
- 58:32blue and uh let's uh
- 58:37run assistant the sky appears blue due
- 58:40to the phenomenon called R lights
- 58:41scattering etc etc so you see that the
- 58:44base model is just continuing the
- 58:45sequence but because the sequence looks
- 58:47like this conversation it takes on that
- 58:49role but it is a little subtle because
- 58:52here it just uh you know it ends the
- 58:54assistant and then just you know
- 58:55hallucinate Ates the next question by
- 58:57the human Etc so it'll just continue
- 58:58going on and on uh but you can see that
- 59:01we have sort of accomplished the task
- 59:03and if you just took this why is the sky
- 59:06blue and if we just refresh this and put
- 59:09it here then of course we don't expect
- 59:10this to work with a base model right
- 59:12we're just going to who knows what we're
- 59:14going to get okay we're just going to
- 59:15get more
- 59:16questions okay so this is one way to
- 59:19create an assistant even though you may
- 59:21only have a base model okay so this is
- 59:24the kind of brief summary of the things
- 59:26we talked about over the last few
- 59:28minutes now let me zoom out
- 59:32here and this is kind of like what we've
- 59:34talked about so far we wish to train LM
- 59:37assistants like chpt we've discussed the
- 59:40first stage of that which is the
- 59:42pre-training stage and we saw that
- 59:44really what it comes down to is we take
- 59:45Internet documents we break them up into
- 59:47these tokens these atoms of little text
- 59:49chunks and then we predict token
- 59:51sequences using neural networks the
- 59:54output of this entire stage is this base
- 59:56model it is the setting of The
- 59:58parameters of this network and this base
- 1:00:01model is basically an internet document
- 1:00:03simulator on the token level so it can
- 1:00:05just uh it can generate token sequences
- 1:00:08that have the same kind of like
- 1:00:10statistics as Internet documents and we
- 1:00:12saw that we can use it in some
- 1:00:13applications but we actually need to do
- 1:00:15better we want an assistant we want to
- 1:00:17be able to ask questions and we want the
- 1:00:18model to give us answers and so we need
- 1:00:21to now go into the second stage which is
- 1:00:23called the post-training stage so we
- 1:00:26take our base model our internet
- 1:00:28document simulator and hand it off to
- 1:00:29post training so we're now going to
- 1:00:31discuss a few ways to do what's called
- 1:00:33post training of these models these
- 1:00:36stages in post training are going to be
- 1:00:38computationally much less expensive most
- 1:00:40of the computational work all of the
- 1:00:42massive data centers um and all of the
- 1:00:45sort of heavy compute and millions of
- 1:00:47dollars are the pre-training stage but
- 1:00:50now we go into the slightly cheaper but
- 1:00:52still extremely important stage called
- 1:00:54post trining where we turn this llm
- 1:00:57model into an assistant so let's take a
- 1:00:59look at how we can get our model to not
- 1:01:02sample internet documents but to give
- 1:01:04answers to questions so in other words
- 1:01:07what we want to do is we want to start
- 1:01:08thinking about conversations and these
- 1:01:10are conversations that can be multi-turn
- 1:01:13so so uh there can be multiple turns and
- 1:01:15they are in the simplest case a
- 1:01:17conversation between a human and an
- 1:01:19assistant and so for example we can
- 1:01:21imagine the conversation could look
- 1:01:22something like this when a human says
- 1:01:24what is 2 plus2 the assistant should re
- 1:01:25respond with something like 2 plus 2 is
- 1:01:274 when a human follows up and says what
- 1:01:29if it was star instead of a plus
- 1:01:31assistant could respond with something
- 1:01:32like
- 1:01:33this um and similar here this is another
- 1:01:36example showing that the assistant could
- 1:01:37also have some kind of a personality
- 1:01:39here uh that it's kind of like nice and
- 1:01:41then here in the third example I'm
- 1:01:43showing that when a human is asking for
- 1:01:44something that we uh don't wish to help
- 1:01:47with we can produce what's called
- 1:01:48refusal we can say that we cannot help
- 1:01:50with that so in other words what we want
- 1:01:53to do now is we want to think through
- 1:01:55how in a system should interact with the
- 1:01:57human and we want to program the
- 1:01:59assistant and Its Behavior in these
- 1:02:01conversations now because this is neural
- 1:02:03networks we're not going to be
- 1:02:04programming these explicitly in code
- 1:02:07we're not going to be able to program
- 1:02:08the assistant in that way because this
- 1:02:10is neural networks everything is done
- 1:02:12through neural network training on data
- 1:02:14sets and so because of that we are going
- 1:02:17to be implicitly programming the
- 1:02:19assistant by creating data sets of
- 1:02:21conversations so these are three
- 1:02:23independent examples of conversations in
- 1:02:25a data dat set an actual data set and
- 1:02:27I'm going to show you examples will be
- 1:02:29much larger it could have hundreds of
- 1:02:31thousands of conversations that are
- 1:02:32multi- turn very long Etc and would
- 1:02:34cover a diverse breath of topics but
- 1:02:37here I'm only showing three examples but
- 1:02:39the way this works basically is uh a
- 1:02:42assistant is being programmed by example
- 1:02:45and where is this data coming from like
- 1:02:472 * 2al 4 same as 2 plus 2 Etc where
- 1:02:50does that come from this comes from
- 1:02:51Human labelers so we will basically give
- 1:02:54human labelers some conversational
- 1:02:56context and we will ask them to um
- 1:02:58basically give the ideal assistant
- 1:03:00response in this situation and a human
- 1:03:03will write out the ideal response for an
- 1:03:06assistant in any situation and then
- 1:03:08we're going to get the model to
- 1:03:10basically train on this and to imitate
- 1:03:12those kinds of
- 1:03:14responses so the way this works then is
- 1:03:16we are going to take our base model
- 1:03:17which we produced in the preing stage
- 1:03:20and this base model was trained on
- 1:03:21internet documents we're now going to
- 1:03:23take that data set of internet documents
- 1:03:25and we're gonna throw it out and we're
- 1:03:27going to substitute a new data set and
- 1:03:29that's going to be a data set of
- 1:03:30conversations and we're going to
- 1:03:32continue training the model on these
- 1:03:33conversations on this new data set of
- 1:03:35conversations and what happens is that
- 1:03:37the model will very rapidly adjust and
- 1:03:40will sort of like learn the statistics
- 1:03:42of how this assistant responds to human
- 1:03:45queries and then later during inference
- 1:03:48we'll be able to basically um Prime the
- 1:03:51assistant and get the response and it
- 1:03:54will be imitating what the humans will
- 1:03:56human labelers would do in that
- 1:03:57situation if that makes sense so we're
- 1:04:00going to see examples of that and this
- 1:04:01is going to become bit more concrete I
- 1:04:03also wanted to mention that this
- 1:04:05post-training stage we're going to
- 1:04:06basically just continue training the
- 1:04:07model but um the pre-training stage can
- 1:04:10in practice take roughly three months of
- 1:04:13training on many thousands of computers
- 1:04:15the post-training stage will typically
- 1:04:16be much shorter like 3 hours for example
- 1:04:20um and that's because the data set of
- 1:04:21conversations that we're going to create
- 1:04:23here manually is much much smaller than
- 1:04:26the data set of text on the internet and
- 1:04:28so this training will be very short but
- 1:04:31fundamentally we're just going to take
- 1:04:33our base model we're going to continue
- 1:04:35training using the exact same algorithm
- 1:04:37the exact same everything except we're
- 1:04:39swapping out the data set for
- 1:04:40conversations so the questions now are
- 1:04:43what are these conversations how do we
- 1:04:44represent them how do we get the model
- 1:04:46to see conversations instead of just raw
- 1:04:49text and then what are the outcomes of
- 1:04:52um this kind of training and what do you
- 1:04:54get in a certain like psychological
- 1:04:56sense uh when we talk about the model so
- 1:04:58let's turn to those questions now so
- 1:05:01let's start by talking about the
- 1:05:02tokenization of conversations everything
- 1:05:05in these models has to be turned into
- 1:05:07tokens because everything is just about
- 1:05:08token sequences so how do we turn
- 1:05:10conversations into token sequences is
- 1:05:12the question and so for that we need to
- 1:05:15design some kind of ending coding and uh
- 1:05:17this is kind of similar to maybe if
- 1:05:18you're familiar you don't have to be
- 1:05:20with for example the TCP IP packet in um
- 1:05:23on the internet there are precise rules
- 1:05:25and protocols for how you represent
- 1:05:27information how everything is structured
- 1:05:29together so that you have all this kind
- 1:05:30of data laid out in a way that is
- 1:05:32written out on a paper and that everyone
- 1:05:34can agree on and so it's the same thing
- 1:05:36now happening in llms we need some kind
- 1:05:38of data structures and we need to have
- 1:05:40some rules around how these data
- 1:05:41structures like conversations get
- 1:05:43encoded and decoded to and from tokens
- 1:05:46and so I want to show you now how I
- 1:05:48would
- 1:05:49recreate uh this conversation in the
- 1:05:52token space so if you go to Tech
- 1:05:54tokenizer
- 1:05:56I can take that conversation and this is
- 1:05:58how it is represented in uh for the
- 1:06:01language model so here we have we are
- 1:06:03iterating a user and an assistant in
- 1:06:06this two- turn
- 1:06:08conversation and what you're seeing here
- 1:06:10is it looks ugly but it's actually
- 1:06:11relatively simple the way it gets turned
- 1:06:13into a token sequence here at the end is
- 1:06:16a little bit complicated but at the end
- 1:06:18this conversation between a user and
- 1:06:19assistant ends up being 49 tokens it is
- 1:06:22a one-dimensional sequence of 49 tokens
- 1:06:24and these are the tokens
- 1:06:26okay and all the different llms will
- 1:06:29have a slightly different format or
- 1:06:31protocols and it's a little bit of a
- 1:06:33wild west right now but for example GPT
- 1:06:3640 does it in the following way you have
- 1:06:39this special token called imore start
- 1:06:42and this is short for IM imaginary
- 1:06:44monologue uh the
- 1:06:46start then you have to specify um I
- 1:06:49don't actually know why it's called that
- 1:06:50to be honest then you have to specify
- 1:06:52whose turn it is so for example user
- 1:06:54which is a token 4
- 1:06:5628 then you have internal monologue
- 1:07:00separator and then it's the exact
- 1:07:03question so the tokens of the question
- 1:07:05and then you have to close it so I am
- 1:07:07end the end of the imaginary monologue
- 1:07:09so
- 1:07:10basically the question from a user of
- 1:07:13what is 2 plus two ends up being the
- 1:07:16token sequence of these tokens and now
- 1:07:19the important thing to mention here is
- 1:07:20that IM start this is not text right IM
- 1:07:24start is a special token that gets added
- 1:07:27it's a new token and um this token has
- 1:07:30never been trained on so far it is a new
- 1:07:32token that we create in a post-training
- 1:07:34stage and we introduce and so these
- 1:07:37special tokens like IM seep IM start Etc
- 1:07:40are introduced and interspersed with
- 1:07:42text so that they sort of um get the
- 1:07:45model to learn that hey this is a the
- 1:07:47start of a turn for who is it start of
- 1:07:49the turn for the start of the turn is
- 1:07:51for the user and then this is what the
- 1:07:54user says and then the user ends and
- 1:07:56then it's a new start of a turn and it
- 1:07:58is by the assistant and then what does
- 1:08:01the assistant say well these are the
- 1:08:02tokens of what the assistant says Etc
- 1:08:05and so this conversation is not turned
- 1:08:06into the sequence of tokens the specific
- 1:08:09details here are not actually that
- 1:08:11important all I'm trying to show you in
- 1:08:13concrete terms is that our conversations
- 1:08:15which we think of as kind of like a
- 1:08:16structured object end up being turned
- 1:08:19via some encoding into onedimensional
- 1:08:21sequences of tokens and so because this
- 1:08:25is one dimensional sequence of tokens we
- 1:08:27can apply all the stuff that we applied
- 1:08:29before now it's just a sequence of
- 1:08:30tokens and now we can train a language
- 1:08:33model on it and so we're just predicting
- 1:08:35the next token in a sequence uh just
- 1:08:37like before and um we can represent and
- 1:08:39train on conversations and then what
- 1:08:42does it look like at test time during
- 1:08:43inference so say we've trained a model
- 1:08:46and we've trained a model on these kinds
- 1:08:49of data sets of conversations and now we
- 1:08:51want to
- 1:08:52inference so during inference what does
- 1:08:54this look like when you're on on chash
- 1:08:55apt well you come to chash apt and you
- 1:08:58have say like a dialogue with it and the
- 1:09:01way this works is
- 1:09:03basically um say that this was already
- 1:09:06filled in so like what is 2 plus 2 2
- 1:09:07plus 2 is four and now you issue what if
- 1:09:10it was times I am end and what basically
- 1:09:13ends up happening um on the servers of
- 1:09:16open AI or something like that is they
- 1:09:18put in I start assistant I amep and this
- 1:09:21is where they end it right here so they
- 1:09:24construct this context and now they
- 1:09:27start sampling from the model so it's at
- 1:09:29this stage that they will go to the
- 1:09:30model and say okay what is a good for
- 1:09:32sequence what is a good first token what
- 1:09:34is a good second token what is a good
- 1:09:36third token and this is where the LM
- 1:09:38takes over and creates a response like
- 1:09:41for example response that looks
- 1:09:43something like this but it doesn't have
- 1:09:44to be identical to this but it will have
- 1:09:46the flavor of this if this kind of a
- 1:09:48conversation was in the data set so um
- 1:09:52that's roughly how the protocol Works
- 1:09:54although the details of this protocol
- 1:09:56are not important so again my goal is
- 1:09:59that just to show you that everything
- 1:10:01ends up being just a one-dimensional
- 1:10:02token sequence so we can apply
- 1:10:04everything we've already seen but we're
- 1:10:06now training on conversations and we're
- 1:10:08now uh basically generating
- 1:10:10conversations as well okay so now I
- 1:10:13would like to turn to what these data
- 1:10:14sets look like in practice the first
- 1:10:16paper that I would like to show you and
- 1:10:17the first effort in this direction is
- 1:10:20this paper from openai in 2022 and this
- 1:10:23paper was called instruct GPT or the
- 1:10:25technique that they developed and this
- 1:10:27was the first time that opena has kind
- 1:10:29of talked about how you can take
- 1:10:30language models and fine-tune them on
- 1:10:32conversations and so this paper has a
- 1:10:34number of details that I would like to
- 1:10:36take you through so the first stop I
- 1:10:38would like to make is in section 3.4
- 1:10:40where they talk about the human
- 1:10:41contractors that they hired uh in this
- 1:10:44case from upwork or through scale AI to
- 1:10:47uh construct these conversations and so
- 1:10:49there are human labelers involved whose
- 1:10:52job it is professionally to create these
- 1:10:54conversations and these labelers are
- 1:10:56asked to come up with prompts and then
- 1:10:58they are asked to also complete the
- 1:11:00ideal assistant responses and so these
- 1:11:03are the kinds of prompts that people
- 1:11:04came up with so these are human labelers
- 1:11:06so list five ideas for how to regain
- 1:11:08enthusiasm for my career what are the
- 1:11:10top 10 science fiction books I should
- 1:11:12read next and there's many different
- 1:11:13types of uh kind of prompts here so
- 1:11:16translate this sentence from uh to
- 1:11:18Spanish Etc and so there's many things
- 1:11:21here that people came up with they first
- 1:11:23come up with the prompt and then they
- 1:11:25also uh answer that prompt and they give
- 1:11:28the ideal assistant response now how do
- 1:11:30they know what is the ideal assistant
- 1:11:32response that they should write for
- 1:11:33these prompts so when we scroll down a
- 1:11:35little bit further we see that here we
- 1:11:37have this excerpt of labeling
- 1:11:39instructions uh that are given to the
- 1:11:41human labelers so the company that is
- 1:11:44developing the language model like for
- 1:11:45example open AI writes up labeling
- 1:11:47instructions for how the humans should
- 1:11:49create ideal responses and so here for
- 1:11:52example is an excerpt uh of these kinds
- 1:11:54of labeling instruction instructions on
- 1:11:56High level you're asking people to be
- 1:11:57helpful truthful and harmless and you
- 1:11:59can pause the video if you'd like to see
- 1:12:01more here but on a high level basically
- 1:12:04just just answer try to be helpful try
- 1:12:06to be truthful and don't answer
- 1:12:08questions that we don't want um kind of
- 1:12:10the system to handle uh later in chat
- 1:12:13gbt and so roughly speaking the company
- 1:12:16comes up with the labeling instructions
- 1:12:18usually they are not this short usually
- 1:12:19there are hundreds of pages and people
- 1:12:21have to study them professionally and
- 1:12:23then they write out the ideal assistant
- 1:12:26responses uh following those labeling
- 1:12:28instructions so this is a very human
- 1:12:30heavy process as it was described in
- 1:12:32this paper now the data set for instruct
- 1:12:34GPT was never actually released by openi
- 1:12:37but we do have some open- Source um
- 1:12:39reproductions that were're trying to
- 1:12:40follow this kind of a setup and collect
- 1:12:42their own data so one that I'm familiar
- 1:12:45with for example is the effort of open
- 1:12:48Assistant from a while back and this is
- 1:12:50just one of I think many examples but I
- 1:12:52just want to show you an example so
- 1:12:54here's so these were people on the
- 1:12:56internet that were asked to basically
- 1:12:57create these conversations similar to
- 1:12:59what um open I did with human labelers
- 1:13:03and so here's an entry of a person who
- 1:13:05came up with this BR can you write a
- 1:13:07short introduction to the relevance of
- 1:13:08the term
- 1:13:09manop uh in economics please use
- 1:13:12examples Etc and then the same person or
- 1:13:15potentially a different person will
- 1:13:17write up the response so here's the
- 1:13:18assistant response to this and so then
- 1:13:21the same person or different person will
- 1:13:23actually write out this ideal
- 1:13:26response and then this is an example of
- 1:13:29maybe how the conversation could
- 1:13:30continue now explain it to a dog and
- 1:13:33then you can try to come up with a
- 1:13:34slightly a simpler explanation or
- 1:13:36something like that now this then
- 1:13:39becomes the label and we end up training
- 1:13:41on this so what happens during training
- 1:13:45is that um of course we're not going to
- 1:13:48have a full coverage of all the possible
- 1:13:50questions that um the model will
- 1:13:53encounter at test time during inference
- 1:13:56we can't possibly cover all the possible
- 1:13:57prompts that people are going to be
- 1:13:59asking in the future but if we have a
- 1:14:02like a data set of a few of these
- 1:14:03examples then the model during training
- 1:14:06will start to take on this Persona of
- 1:14:09this helpful truthful harmless assistant
- 1:14:12and it's all programmed by example and
- 1:14:14so these are all examples of behavior
- 1:14:16and if you have conversations of these
- 1:14:18example behaviors and you have enough of
- 1:14:19them like 100,00 and you train on it the
- 1:14:22model sort of starts to understand the
- 1:14:23statistical pattern and it kind of takes
- 1:14:26on this personality of this
- 1:14:28assistant now it's possible that when
- 1:14:30you get the exact same question like
- 1:14:32this at test time it's possible that the
- 1:14:35answer will be recited as exactly what
- 1:14:38was in the training set but more likely
- 1:14:40than that is that the model will kind of
- 1:14:43like do something of a similar Vibe um
- 1:14:45and we will understand that this is the
- 1:14:47kind of answer that you want um so
- 1:14:51that's what we're doing we're
- 1:14:52programming the system um by example and
- 1:14:55the system adopts statistically this
- 1:14:58Persona of this helpful truthful
- 1:15:00harmless assistant which is kind of like
- 1:15:02reflected in the labeling instructions
- 1:15:04that the company creates now I want to
- 1:15:06show you that the state-of-the-art has
- 1:15:08kind of advanced in the last 2 or 3
- 1:15:09years uh since the instr GPT paper so in
- 1:15:12particular it's not very common for
- 1:15:14humans to be doing all the heavy lifting
- 1:15:16just by themselves anymore and that's
- 1:15:18because we now have language models and
- 1:15:19these language models are helping us
- 1:15:21create these data sets and conversations
- 1:15:23so it is very rare that the people will
- 1:15:25like literally just write out the
- 1:15:26response from scratch it is a lot more
- 1:15:28likely that they will use an existing
- 1:15:29llm to basically like uh come up with an
- 1:15:32answer and then they will edit it or
- 1:15:34things like that so there's many
- 1:15:35different ways in which now llms have
- 1:15:37started to kind of permeate this
- 1:15:39posttraining Set uh stack and llms are
- 1:15:43basically used pervasively to help
- 1:15:45create these massive data sets of
- 1:15:46conversations so I don't want to show
- 1:15:49like Ultra chat is one um such example
- 1:15:52of like a more modern data set of
- 1:15:53conversations it is to a very large
- 1:15:56extent synthetic but uh I believe
- 1:15:58there's some human involvement I could
- 1:15:59be wrong with that usually there will be
- 1:16:01a little bit of human but there will be
- 1:16:02a huge amount of synthetic help um and
- 1:16:06this is all kind of like uh constructed
- 1:16:08in different ways and Ultra chat is just
- 1:16:10one example of many sft data sets that
- 1:16:12currently exist and the only thing I
- 1:16:14want to show you is that uh these data
- 1:16:15sets have now millions of conversations
- 1:16:18uh these conversations are mostly
- 1:16:19synthetic but they're probably edited to
- 1:16:21some extent by humans and they span a
- 1:16:23huge diversity of sort of
- 1:16:27um uh areas and so on so these are
- 1:16:31fairly extensive artifacts by now and
- 1:16:33there's all these like sft mixtures as
- 1:16:35they're called so you have a mixture of
- 1:16:37like lots of different types and sources
- 1:16:39and it's partially synthetic partially
- 1:16:41human and it's kind of like um gone in
- 1:16:44that direction since uh but roughly
- 1:16:46speaking we still have sft data sets
- 1:16:48they're made up of conversations we're
- 1:16:50training on them um just like we did
- 1:16:52before and
- 1:16:55uh I guess like the last thing to note
- 1:16:57is that I want to dispel a little bit of
- 1:17:00the magic of talking to an AI like when
- 1:17:02you go to chat GPT and you give it a
- 1:17:04question and then you hit enter uh what
- 1:17:07is coming back is kind of like
- 1:17:10statistically aligned with what's
- 1:17:12happening in the training set and these
- 1:17:14training sets I mean they really just
- 1:17:16have a seed in humans following labeling
- 1:17:19instructions so what are you actually
- 1:17:21talking to in chat GPT or how should you
- 1:17:24think about it well it's not coming from
- 1:17:25some magical AI like roughly speaking
- 1:17:28it's coming from something that is
- 1:17:29statistically imitating human labelers
- 1:17:32which comes from labeling instructions
- 1:17:34written by these companies and so you're
- 1:17:36kind of imitating this uh you're kind of
- 1:17:38getting um it's almost as if you're
- 1:17:40asking human labeler and imagine that
- 1:17:43the answer that is given to you uh from
- 1:17:45chbt is some kind of a simulation of a
- 1:17:47human labeler uh and it's kind of like
- 1:17:50asking what would a human labeler say in
- 1:17:53this kind of a conversation
- 1:17:56and uh it's not just like this human
- 1:17:58labeler is not just like a random person
- 1:18:00from the internet because these
- 1:18:01companies actually hire experts so for
- 1:18:03example when you are asking questions
- 1:18:04about code and so on the human labelers
- 1:18:06that would be in um involved in creation
- 1:18:08of these conversation data sets they
- 1:18:10will usually be usually be educated
- 1:18:12expert people and you're kind of like
- 1:18:15asking a question of like a simulation
- 1:18:17of those people if that makes sense so
- 1:18:19you're not talking to a magical AI
- 1:18:21you're talking to an average labeler
- 1:18:22this average labeler is probably fairly
- 1:18:24highly skilled
- 1:18:25but you're talking to kind of like an
- 1:18:26instantaneous simulation of that kind of
- 1:18:29a person that would be hired uh in the
- 1:18:32construction of these data sets so let
- 1:18:34me give you one more specific example
- 1:18:36before we move on for example when I go
- 1:18:38to chpt and I say recommend the top five
- 1:18:40landmarks who see in Paris and then I
- 1:18:42hit
- 1:18:44enter
- 1:18:49uh okay here we go okay when I hit enter
- 1:18:52what's coming out here how do I think
- 1:18:55about it well it's not some kind of a
- 1:18:56magical AI that has gone out and
- 1:18:58researched all the landmarks and then
- 1:19:00ranked them using its infinite
- 1:19:01intelligence Etc what I'm getting is a
- 1:19:04statistical simulation of a labeler that
- 1:19:07was hired by open AI you can think about
- 1:19:09it roughly in that way and so if this
- 1:19:13specific um question is in the
- 1:19:16posttraining data set somewhere at open
- 1:19:17aai then I'm very likely to see an
- 1:19:20answer that is probably very very
- 1:19:22similar to what that human labeler would
- 1:19:24have put down
- 1:19:25for those five landmarks how does the
- 1:19:27human labeler come up with this well
- 1:19:28they go off and they go on the internet
- 1:19:29and they kind of do their own little
- 1:19:31research for 20 minutes and they just
- 1:19:32come up with a list right now so if they
- 1:19:35come up with this list and this is in
- 1:19:37the data set I'm probably very likely to
- 1:19:39see what they submitted as the correct
- 1:19:41answer from the assistant now if this
- 1:19:44specific query is not part of the post
- 1:19:46training data set then what I'm getting
- 1:19:48here is a little bit more emergent uh
- 1:19:51because uh the model kind of understands
- 1:19:53the statistically
- 1:19:55um the kinds of landmarks that are in
- 1:19:57this training set are usually the
- 1:19:59prominent landmarks the landmarks that
- 1:20:00people usually want to see the kinds of
- 1:20:02landmarks that are usually uh very often
- 1:20:05talked about on the internet and
- 1:20:06remember that the model already has a
- 1:20:08ton of Knowledge from its pre-training
- 1:20:10on the internet so it's probably seen a
- 1:20:12ton of conversations about Paris about
- 1:20:13landmarks about the kinds of things that
- 1:20:15people like to see and so it's the
- 1:20:17pre-training knowledge that has then
- 1:20:18combined with the postering data set
- 1:20:20that results in this kind of an
- 1:20:23imitation um
- 1:20:25so that's uh that's roughly how you can
- 1:20:27kind of think about what's happening
- 1:20:29behind the scenes here in in this
- 1:20:31statistical sense okay now I want to
- 1:20:33turn to the topic of llm psychology as I
- 1:20:35like to call it which is what are sort
- 1:20:37of the emergent cognitive effects of the
- 1:20:40training pipeline that we have for these
- 1:20:42models so in particular the first one I
- 1:20:44want to talk to is of course
- 1:20:47hallucinations so you might be familiar
- 1:20:50with model hallucinations it's when llms
- 1:20:52make stuff up they just totally
- 1:20:53fabricate information Etc and it's a big
- 1:20:56problem with llm assistants it is a
- 1:20:58problem that existed to a large extent
- 1:21:00with early models uh from many years ago
- 1:21:02and I think the problem has gotten a bit
- 1:21:04better uh because there are some
- 1:21:05medications that I'm going to go into in
- 1:21:07a second for now let's just try to
- 1:21:09understand where these hallucinations
- 1:21:10come from so here's a specific example
- 1:21:13of a few uh of three conversations that
- 1:21:16you might think you have in your
- 1:21:17training set and um these are pretty
- 1:21:20reasonable conversations that you could
- 1:21:22imagine being in the training set so
- 1:21:23like for example who is Cruz well Tom
- 1:21:25Cruz is an famous actor American actor
- 1:21:27and producer Etc who is John baraso this
- 1:21:31turns out to be a us senetor for example
- 1:21:34who is genis Khan well genis Khan was
- 1:21:36blah blah blah and so this is what your
- 1:21:39conversations could look like at
- 1:21:40training time now the problem with this
- 1:21:42is that when the human is writing the
- 1:21:46correct answer for the assistant in each
- 1:21:48one of these cases uh the human either
- 1:21:51like knows who this person is or they
- 1:21:52research them on the Internet and they
- 1:21:53come in and they write this response
- 1:21:55that kind of has this like confident
- 1:21:57tone of an answer and what happens
- 1:21:59basically is that at test time when you
- 1:22:01ask for someone who is this is a totally
- 1:22:03random name that I totally came up with
- 1:22:05and I don't think this person exists um
- 1:22:07as far as I know I just Tred to generate
- 1:22:09it randomly the problem is when we ask
- 1:22:11who is Orson kovats the problem is that
- 1:22:15the assistant will not just tell you oh
- 1:22:17I don't know even if the assistant and
- 1:22:20the language model itself might know
- 1:22:23inside its features inside its
- 1:22:24activations inside of its brain sort of
- 1:22:26it might know that this person is like
- 1:22:28not someone that um that is that it's
- 1:22:30familiar with even if some part of the
- 1:22:32network kind of knows that in some sense
- 1:22:35the uh saying that oh I don't know who
- 1:22:37this is is is not going to happen
- 1:22:40because the model statistically imitates
- 1:22:42is training set in the training set the
- 1:22:45questions of the form who is blah are
- 1:22:47confidently answered with the correct
- 1:22:49answer and so it's going to take on the
- 1:22:52style of the answer and it's going to do
- 1:22:53its best it's going to give you
- 1:22:55statistically the most likely guess and
- 1:22:57it's just going to basically make stuff
- 1:22:58up because these models again we just
- 1:23:01talked about it is they don't have
- 1:23:02access to the internet they're not doing
- 1:23:04research these are statistical token
- 1:23:06tumblers as I call them uh is just
- 1:23:08trying to sample the next token in the
- 1:23:10sequence and it's going to basically
- 1:23:12make stuff up so let's take a look at
- 1:23:13what this looks
- 1:23:15like I have here what's called the
- 1:23:17inference playground from hugging face
- 1:23:20and I am on purpose picking on a model
- 1:23:22called Falcon 7B which is an old model
- 1:23:25this is a few years ago now so it's an
- 1:23:27older model So It suffers from
- 1:23:28hallucinations and as I mentioned this
- 1:23:31has improved over time recently but
- 1:23:33let's say who is Orson kovats let's ask
- 1:23:35Falcon 7B instruct
- 1:23:37run oh yeah Orson kovat is an American
- 1:23:40author and science uh fiction writer
- 1:23:42okay this is totally false it's
- 1:23:44hallucination let's try again these are
- 1:23:46statistical systems right so we can
- 1:23:48resample this time Orson kovat is a
- 1:23:51fictional character from this 1950s TV
- 1:23:53show it's total BS right let's try again
- 1:23:57he's a former minor league baseball
- 1:23:59player okay so basically the model
- 1:24:02doesn't know and it's given us lots of
- 1:24:04different answers because it doesn't
- 1:24:06know it's just kind of like sampling
- 1:24:08from these probabilities the model
- 1:24:10starts with the tokens who is oron
- 1:24:12kovats assistant and then it comes in
- 1:24:14here and it's get it's getting these
- 1:24:17probabilities and it's just sampling
- 1:24:19from the probabilities and it just like
- 1:24:20comes up with stuff and the stuff is
- 1:24:24actually
- 1:24:24statistically consistent with the style
- 1:24:27of the answer in its training set and
- 1:24:29it's just doing that but you and I
- 1:24:31experiened it as a madeup factual
- 1:24:33knowledge but keep in mind that uh the
- 1:24:36model basically doesn't know and it's
- 1:24:37just imitating the format of the answer
- 1:24:40and it's not going to go off and look it
- 1:24:41up uh because it's just imitating again
- 1:24:44the answer so how can we uh mitigate
- 1:24:47this because for example when we go to
- 1:24:48chat apt and I say who is oron kovats
- 1:24:50and I'm now asking the stateoftheart
- 1:24:52state-of-the-art model from open AI
- 1:24:55this model will tell
- 1:24:56you oh so this model is actually is even
- 1:25:00smarter because you saw very briefly it
- 1:25:02said searching the web uh we're going to
- 1:25:04cover this later um it's actually trying
- 1:25:07to do tool use and
- 1:25:11uh kind of just like came up with some
- 1:25:13kind of a story but I want to just who
- 1:25:15or Kovach did not use any tools I don't
- 1:25:19want it to do web
- 1:25:22search there's a wellknown historical or
- 1:25:24public figure named or oron kovats so
- 1:25:27this model is not going to make up stuff
- 1:25:29this model knows that it doesn't know
- 1:25:31and it tells you that it doesn't appear
- 1:25:32to be a person that this model knows so
- 1:25:35somehow we sort of improved
- 1:25:37hallucinations even though they clearly
- 1:25:39are an issue in older models and it
- 1:25:42makes totally uh sense why you would be
- 1:25:44getting these kinds of answers if this
- 1:25:46is what your training set looks like so
- 1:25:47how do we fix this okay well clearly we
- 1:25:50need some examples in our data set that
- 1:25:53where the correct answer for the
- 1:25:54assistant is that the model doesn't know
- 1:25:57about some particular fact but we only
- 1:25:59need to have those answers be produced
- 1:26:02in the cases where the model actually
- 1:26:03doesn't know and so the question is how
- 1:26:05do we know what the model knows or
- 1:26:07doesn't know well we can empirically
- 1:26:09probe the model to figure that out so
- 1:26:11let's take a look at for example how
- 1:26:13meta uh dealt with hallucinations for
- 1:26:16the Llama 3 series of models as an
- 1:26:18example so in this paper that they
- 1:26:20published from meta we can go into
- 1:26:22hallucinations
- 1:26:25which they call here factuality and they
- 1:26:27describe the procedure by which they
- 1:26:29basically interrogate the model to
- 1:26:32figure out what it knows and doesn't
- 1:26:33know to figure out sort of like the
- 1:26:35boundary of its knowledge and then they
- 1:26:38add examples to the training set where
- 1:26:41for the things where the model doesn't
- 1:26:44know them the correct answer is that the
- 1:26:46model doesn't know them which sounds
- 1:26:48like a very easy thing to do in
- 1:26:50principle but this roughly fixes the
- 1:26:53issue and the the reason it fixes the
- 1:26:54issue is
- 1:26:56because remember like the model might
- 1:26:59actually have a pretty good model of its
- 1:27:01self knowledge inside the network so
- 1:27:04remember we looked at the network and
- 1:27:06all these neurons inside the network you
- 1:27:08might imagine that there's a neuron
- 1:27:09somewhere in the network that sort of
- 1:27:11like lights up for when the model is
- 1:27:14uncertain but the problem is that the
- 1:27:17activation of that neuron is not
- 1:27:18currently wired up to the model actually
- 1:27:20saying in words that it doesn't know so
- 1:27:23even though the internal of the neural
- 1:27:24network no because there's some neurons
- 1:27:26that represent that the model uh will
- 1:27:29not surface that it will instead take
- 1:27:31its best guess so that it sounds
- 1:27:33confident um just like it sees in a
- 1:27:35training set so we need to basically
- 1:27:37interrogate the model and allow it to
- 1:27:39say I don't know in the cases that it
- 1:27:41doesn't know so let me take you through
- 1:27:43what meta roughly does so basically what
- 1:27:45they do is here I have an example uh
- 1:27:48Dominic kek is uh the featured article
- 1:27:51today so I just went there randomly and
- 1:27:54what they do is basically they take a
- 1:27:55random document in a training set and
- 1:27:58they take a paragraph and then they use
- 1:28:01an llm to construct questions about that
- 1:28:04paragraph so for example I did that with
- 1:28:06chat GPT
- 1:28:09here so I said here's a paragraph from
- 1:28:12this document generate three specific
- 1:28:14factual questions based on this
- 1:28:15paragraph and give me the questions and
- 1:28:17the answers and so the llms are already
- 1:28:20good enough to create and reframe this
- 1:28:23information so if the information is in
- 1:28:25the context window um of this llm this
- 1:28:29actually works pretty well it doesn't
- 1:28:30have to rely on its memory it's right
- 1:28:33there in the context window and so it
- 1:28:35can basically reframe that information
- 1:28:37with fairly high accuracy so for example
- 1:28:40can generate questions for us like for
- 1:28:41which team did he play here's the answer
- 1:28:44how many cups did he win Etc and now
- 1:28:47what we have to do is we have some
- 1:28:48question and answers and now we want to
- 1:28:50interrogate the model so roughly
- 1:28:51speaking what we'll do is we'll take our
- 1:28:53questions and we'll go to our model
- 1:28:55which would be uh say llama uh in meta
- 1:28:59but let's just interrogate mol 7B here
- 1:29:01as an example that's another model so
- 1:29:04does this model know about this answer
- 1:29:07let's take a
- 1:29:09look uh so he played for Buffalo Sabers
- 1:29:12right so the model knows and the the way
- 1:29:15that you can programmatically decide is
- 1:29:16basically we're going to take this
- 1:29:18answer from the model and we're going to
- 1:29:20compare it to the correct answer and
- 1:29:23again the model model are good enough to
- 1:29:24do this automatically so there's no
- 1:29:26humans involved here we can take uh
- 1:29:28basically the answer from the model and
- 1:29:30we can use another llm judge to check if
- 1:29:33that is correct according to this answer
- 1:29:35and if it is correct that means that the
- 1:29:37model probably knows so what we're going
- 1:29:38to do is we're going to do this maybe a
- 1:29:40few times so okay it knows it's Buffalo
- 1:29:42Savers let's drag
- 1:29:45in um Buffalo Sabers let's try one more
- 1:29:51time Buffalo Sabers so we asked three
- 1:29:54times about this factual question and
- 1:29:55the model seems to know so everything is
- 1:29:58great now let's try the second question
- 1:30:00how many Stanley Cups did he
- 1:30:02win and again let's interrogate the
- 1:30:04model about that and the correct answer
- 1:30:06is
- 1:30:08two so um here the model claims that he
- 1:30:13won um four times which is not correct
- 1:30:17right it doesn't match two so the model
- 1:30:20doesn't know it's making stuff up let's
- 1:30:22try again
- 1:30:27um so here the model again it's kind of
- 1:30:30like making stuff up right let's
- 1:30:34Dragon here it says did he did not even
- 1:30:37did not win during his career so
- 1:30:39obviously the model doesn't know and the
- 1:30:41way we can programmatically tell again
- 1:30:42is we interrogate the model three times
- 1:30:45and we compare its answers maybe three
- 1:30:47times five times whatever it is to the
- 1:30:49correct answer and if the model doesn't
- 1:30:51know then we know that the model doesn't
- 1:30:53know this question
- 1:30:54and then what we do is we take this
- 1:30:56question we create a new conversation in
- 1:30:59the training set so we're going to add a
- 1:31:01new conversation training set and when
- 1:31:03the question is how many Stanley Cups
- 1:31:05did he win the answer is I'm sorry I
- 1:31:08don't know or I don't remember and
- 1:31:10that's the correct answer for this
- 1:31:12question because we interrogated the
- 1:31:13model and we saw that that's the case if
- 1:31:15you do this for many different types of
- 1:31:18uh questions for many different types of
- 1:31:20documents you are giving the model an
- 1:31:23opportunity to in its training set
- 1:31:25refuse to say based on its knowledge and
- 1:31:28if you just have a few examples of that
- 1:31:30in your training set the model will know
- 1:31:33um and and has the opportunity to learn
- 1:31:35the association of this knowledge-based
- 1:31:37refusal to this internal neuron
- 1:31:41somewhere in its Network that we presume
- 1:31:43exists and empirically this turns out to
- 1:31:45be probably the case and it can learn
- 1:31:47that Association that hey when this
- 1:31:49neuron of uncertainty is high then I
- 1:31:52actually don't know and I'm allowed to
- 1:31:54say that I'm sorry but I don't think I
- 1:31:56remember this Etc and if you have these
- 1:31:59uh examples in your training set then
- 1:32:01this is a large mitigation for
- 1:32:03hallucination and that's roughly
- 1:32:05speaking why chpt is able to do stuff
- 1:32:08like this as well so these are kinds of
- 1:32:10uh mitigations that people have
- 1:32:12implemented and that have improved the
- 1:32:14factuality issue over time okay so I've
- 1:32:16described mitigation number one for
- 1:32:19basically mitigating the hallucinations
- 1:32:21issue now we can actually do much better
- 1:32:24than that uh it's instead of just saying
- 1:32:27that we don't know uh we can introduce
- 1:32:29an additional mitigation number two to
- 1:32:32give the llm an opportunity to be
- 1:32:33factual and actually answer the question
- 1:32:36now what do you and I do if I was to ask
- 1:32:39you a factual question and you don't
- 1:32:40know uh what would you do um in order to
- 1:32:43answer the question well you could uh go
- 1:32:45off and do some search and uh use the
- 1:32:47internet and you could figure out the
- 1:32:49answer and then tell me what that answer
- 1:32:51is and we can do the exact exact same
- 1:32:54thing with these models so think of the
- 1:32:56knowledge inside the neural network
- 1:32:58inside its billions of parameters think
- 1:33:01of that as kind of a vague recollection
- 1:33:02of the things that the model has seen
- 1:33:05during its training during the
- 1:33:07pre-training stage a long time ago so
- 1:33:09think of that knowledge in the
- 1:33:10parameters as something you read a month
- 1:33:13ago and if you keep reading something
- 1:33:15then you will remember it and the model
- 1:33:17remembers that but if it's something
- 1:33:18rare then you probably don't have a
- 1:33:20really good recollection of that
- 1:33:21information but what you and I do is we
- 1:33:23just go and look it up now when you go
- 1:33:25and look it up what you're doing
- 1:33:26basically is like you're refreshing your
- 1:33:28working memory with information and then
- 1:33:30you're able to sort of like retrieve it
- 1:33:32talk about it or Etc so we need some
- 1:33:34equivalent of allowing the model to
- 1:33:36refresh its memory or its recollection
- 1:33:38and we can do that by introducing tools
- 1:33:41uh for the
- 1:33:42models so the way we are going to
- 1:33:44approach this is that instead of just
- 1:33:45saying hey I'm sorry I don't know we can
- 1:33:48attempt to use tools so we can create uh
- 1:33:53a mechanism
- 1:33:54by which the language model can emit
- 1:33:56special tokens and these are tokens that
- 1:33:57we're going to introduce new tokens so
- 1:34:00for example here I've introduced two
- 1:34:02tokens and I've introduced a format or a
- 1:34:04protocol for how the model is allowed to
- 1:34:07use these tokens so for example instead
- 1:34:09of answering the question when the model
- 1:34:12does not instead of just saying I don't
- 1:34:14know sorry the model has the option now
- 1:34:16to emitting the special token search
- 1:34:18start and this is the query that will go
- 1:34:20to like bing.com in the case of openai
- 1:34:22or say Google search or something like
- 1:34:24that so it will emit the query and then
- 1:34:26it will emit search end and then here
- 1:34:30what will happen is that the program
- 1:34:32that is sampling from the model that is
- 1:34:34running the inference when it sees the
- 1:34:36special token search end instead of
- 1:34:39sampling the next token uh in the
- 1:34:41sequence it will actually pause
- 1:34:44generating from the model it will go off
- 1:34:46it will open a session with bing.com and
- 1:34:49it will paste the search query into Bing
- 1:34:52and it will then um get all the text
- 1:34:54that is retrieved and it will basically
- 1:34:56take that text it will maybe represent
- 1:34:58it again with some other special tokens
- 1:35:00or something like that and it will take
- 1:35:02that text and it will copy paste it here
- 1:35:05into what I Tred to like show with the
- 1:35:07brackets so all that text kind of comes
- 1:35:09here and when the text comes here it
- 1:35:12enters the context window so the model
- 1:35:15so that text from the web search is now
- 1:35:17inside the context window that will feed
- 1:35:20into the neural network and you should
- 1:35:21think of the context window as kind of
- 1:35:23like the working memory of the model
- 1:35:25that data that is in the context window
- 1:35:27is directly accessible by the model it
- 1:35:29directly feeds into the neural network
- 1:35:31so it's not anymore a vague recollection
- 1:35:33it's data that it it has in the context
- 1:35:36window and is directly available to that
- 1:35:38model so now when it's sampling the new
- 1:35:41uh tokens here afterwards it can
- 1:35:43reference very easily the data that has
- 1:35:45been copy pasted in there so that's
- 1:35:48roughly how these um how these tools use
- 1:35:52uh tools uh function
- 1:35:54and so web search is just one of the
- 1:35:55tools we're going to look at some of the
- 1:35:56other tools in a bit uh but basically
- 1:35:59you introduce new tokens you introduce
- 1:36:00some schema by which the model can
- 1:36:02utilize these tokens and can call these
- 1:36:04special functions like web search
- 1:36:06functions and how do you teach the model
- 1:36:08how to correctly use these tools like
- 1:36:10say web search search start search end
- 1:36:12Etc well again you do that through
- 1:36:14training sets so we need now to have a
- 1:36:16bunch of data and a bunch of
- 1:36:18conversations that show the model by
- 1:36:21example how to use web search so what
- 1:36:24are the what are the settings where you
- 1:36:25are using the search um and what does
- 1:36:28that look like and here's by example how
- 1:36:30you start a search and the search Etc
- 1:36:33and uh if you have a few thousand maybe
- 1:36:35examples of that in your training set
- 1:36:36the model will actually do a pretty good
- 1:36:38job of understanding uh how this tool
- 1:36:40works and it will know how to sort of
- 1:36:43structure its queries and of course
- 1:36:44because of the pre-training data set and
- 1:36:47its understanding of the world it
- 1:36:48actually kind of understands what a web
- 1:36:49search is and so it actually kind of has
- 1:36:51a pretty good native understanding
- 1:36:54um of what kind of stuff is a good
- 1:36:56search query um and so it all kind of
- 1:36:58just like works you just need a little
- 1:37:00bit of a few examples to show it how to
- 1:37:02use this new tool and then it can lean
- 1:37:04on it to retrieve information and uh put
- 1:37:07it in the context window and that's
- 1:37:08equivalent to you and I looking
- 1:37:10something up because once it's in the
- 1:37:12context it's in the working memory and
- 1:37:13it's very easy to manipulate and access
- 1:37:16so that's what we saw a few minutes ago
- 1:37:18when I was searching on chat GPT for who
- 1:37:20is Orson kovats the chat GPT language
- 1:37:23model decided Ed that this is some kind
- 1:37:24of a rare um individual or something
- 1:37:27like that and instead of giving me an
- 1:37:29answer from its memory it decided that
- 1:37:31it will sample a special token that is
- 1:37:33going to do web search and we saw
- 1:37:35briefly something flash it was like
- 1:37:36using the web tool or something like
- 1:37:38that so it briefly said that and then we
- 1:37:40waited for like two seconds and then it
- 1:37:41generated this and you see how it's
- 1:37:43creating references here and so it's
- 1:37:45citing sources so what happened here is
- 1:37:50it went off it did a web web search it
- 1:37:52found these sources and these URLs and
- 1:37:55the text of these web pages was all
- 1:37:58stuffed in between here and it's not
- 1:38:01showing here but it's it's basically
- 1:38:02stuffed as text in between here and now
- 1:38:06it sees that text and now it kind of
- 1:38:08references it and says that okay it
- 1:38:11could be these people citation could be
- 1:38:13those people citation Etc so that's what
- 1:38:15happened here and that's what and that's
- 1:38:17why when I said who is Orson kovats I
- 1:38:19could also say don't use any tools and
- 1:38:22then that's enough to um
- 1:38:24basically convince chat PT to not use
- 1:38:25tools and just use its memory and its
- 1:38:28recollection I also went off and I um
- 1:38:32tried to ask this question of Chachi PT
- 1:38:34so how many standing cups did uh Dominic
- 1:38:37Hasek win and Chachi P actually decided
- 1:38:39that it knows the answer and it has the
- 1:38:40confidence to say that uh he want twice
- 1:38:43and so it kind of just relied on its
- 1:38:45memory because presumably it has um it
- 1:38:49has enough of
- 1:38:50a kind of confidence in its weights in
- 1:38:53it parameters and activations that this
- 1:38:55is uh retrievable just for memory um but
- 1:38:59you can also
- 1:39:01conversely use web search to make sure
- 1:39:04and then for the same query it actually
- 1:39:06goes off and it searches and then it
- 1:39:07finds a bunch of sources it finds all
- 1:39:10this all of this stuff gets copy pasted
- 1:39:12in there and then it tells us uh to
- 1:39:15again and sites and it actually says the
- 1:39:17Wikipedia article which is the source of
- 1:39:20this information for us as well so
- 1:39:23that's tools web search the model
- 1:39:25determines when to search and then uh
- 1:39:27that's kind of like how these tools uh
- 1:39:29work and this is an additional kind of
- 1:39:32mitigation for uh hallucinations and
- 1:39:34factuality so I want to stress one more
- 1:39:37time this very important sort of
- 1:39:38psychology
- 1:39:40Point knowledge in the parameters of the
- 1:39:43neural network is a vague recollection
- 1:39:45the knowledge in the tokens that make up
- 1:39:47the context
- 1:39:48window is the working memory and it
- 1:39:51roughly speaking Works kind of like um
- 1:39:53it works for us in our brain the stuff
- 1:39:55we remember is our parameters uh and the
- 1:39:58stuff that we just experienced like a
- 1:40:01few seconds or minutes ago and so on you
- 1:40:03can imagine that being in our context
- 1:40:04window and this context window is being
- 1:40:05built up as you have a conscious
- 1:40:07experience around you so this has a
- 1:40:10bunch of um implications also for your
- 1:40:12use of LOLs in practice so for example I
- 1:40:15can go to chat GPT and I can do
- 1:40:17something like this I can say can you
- 1:40:18Summarize chapter one of Jane Austin's
- 1:40:20Pride and Prejudice right and this is a
- 1:40:22perfectly fine prompt and Chach actually
- 1:40:25does something relatively reasonable
- 1:40:26here and but the reason it does that is
- 1:40:28because Chach has a pretty good
- 1:40:30recollection of a famous work like Pride
- 1:40:32and Prejudice it's probably seen a ton
- 1:40:34of stuff about it there's probably
- 1:40:35forums about this book it's probably
- 1:40:37read versions of this book um and it's
- 1:40:40kind of like remembers because even if
- 1:40:43you've read this or articles about it
- 1:40:46you'd kind of have a recollection enough
- 1:40:48to actually say all this but usually
- 1:40:49when I actually interact with LMS and I
- 1:40:51want them to recall specific things it
- 1:40:53always works better if you just give it
- 1:40:55to them so I think a much better prompt
- 1:40:57would be something like this can you
- 1:40:59summarize for me chapter one of genos's
- 1:41:01spr and Prejudice and then I am
- 1:41:03attaching it below for your reference
- 1:41:04and then I do something like a delimeter
- 1:41:06here and I paste it in and I I found
- 1:41:08that just copy pasting it from some
- 1:41:10website that I found here um so copy
- 1:41:14pasting the chapter one here and I do
- 1:41:16that because when it's in the context
- 1:41:17window the model has direct access to it
- 1:41:20and can exactly it doesn't have to
- 1:41:22recall it it just has access to it and
- 1:41:24so this summary is can be expected to be
- 1:41:27a significantly high quality or higher
- 1:41:29quality than this summary uh just
- 1:41:31because it's directly available to the
- 1:41:32model and I think you and I would work
- 1:41:34in the same way if you want to it would
- 1:41:36be you would produce a much better
- 1:41:37summary if you had reread this chapter
- 1:41:40before you had to summarize it and
- 1:41:42that's basically what's happening here
- 1:41:44or the equivalent of it the next sort of
- 1:41:47psychological Quirk I'd like to talk
- 1:41:48about briefly is that of the knowledge
- 1:41:50of self so what I see very often on the
- 1:41:52internet is that people do something
- 1:41:54like this they ask llms something like
- 1:41:56what model are you and who built you and
- 1:41:59um basically this uh question is a
- 1:42:01little bit nonsensical and the reason I
- 1:42:03say that is that as I try to kind of
- 1:42:05explain with some of the underhood
- 1:42:07fundamentals this thing is not a person
- 1:42:09right it doesn't have a persistent
- 1:42:11existence in any way it sort of boots up
- 1:42:14processes tokens and shuts off and it
- 1:42:17does that for every single person it
- 1:42:18just kind of builds up a context window
- 1:42:19of conversation and then everything gets
- 1:42:21deleted and so this this entity is kind
- 1:42:23of like restarted from scratch every
- 1:42:25single conversation if that makes sense
- 1:42:27it has no persistent self it has no
- 1:42:28sense of self it's a token tumbler and
- 1:42:31uh it follows the statistical
- 1:42:33regularities of its training set so it
- 1:42:35doesn't really make sense to ask it who
- 1:42:38are you what build you Etc and by
- 1:42:40default if you do what I described and
- 1:42:42just by default and from nowhere you're
- 1:42:44going to get some pretty random answers
- 1:42:46so for example let's uh pick on Falcon
- 1:42:48which is a fairly old model and let's
- 1:42:50see what it tells
- 1:42:51us uh so it's evading the question uh
- 1:42:55talented engineers and developers here
- 1:42:58it says I was built by open AI based on
- 1:42:59the gpt3 model it's totally making stuff
- 1:43:01up now the fact that it's built by open
- 1:43:04AI here I think a lot of people would
- 1:43:06take this as evidence that this model
- 1:43:07was somehow trained on open AI data or
- 1:43:09something like that I don't actually
- 1:43:10think that that's necessarily true the
- 1:43:12reason for that is
- 1:43:14that if you don't explicitly program the
- 1:43:17model to answer these kinds of questions
- 1:43:20then what you're going to get is its
- 1:43:22statistical best guess at the answer and
- 1:43:25this model had a um sft data mixture of
- 1:43:29conversations and during the
- 1:43:32fine-tuning um the model sort of
- 1:43:35understands as it's training on this
- 1:43:36data that it's taking on this
- 1:43:38personality of this like helpful
- 1:43:40assistant and it doesn't know how to it
- 1:43:42doesn't actually it wasn't told exactly
- 1:43:44what label to apply to self it just kind
- 1:43:47of is taking on this uh this uh Persona
- 1:43:50of a helpful assistant and remember that
- 1:43:53the pre-training stage took the
- 1:43:55documents from the entire internet and
- 1:43:57Chach and open AI are very prominent in
- 1:43:59these documents and so I think what's
- 1:44:01actually likely to be happening here is
- 1:44:03that this is just its hallucinated label
- 1:44:06for what it is this is its self-identity
- 1:44:08is that it's chat GPT by open Ai and
- 1:44:11it's only saying that because there's a
- 1:44:12ton of data on the internet of um
- 1:44:15answers like this that are actually
- 1:44:17coming from open from chasht and So
- 1:44:20that's its label for what it is now you
- 1:44:23can override this as a developer if you
- 1:44:25have a llm model you can actually
- 1:44:27override it and there are a few ways to
- 1:44:28do that so for example let me show you
- 1:44:31there's this MMO model from Allen Ai and
- 1:44:35um this is one llm it's not a top tier
- 1:44:37LM or anything like that but I like it
- 1:44:39because it is fully open source so the
- 1:44:41paper for Almo and everything else is
- 1:44:43completely fully open source which is
- 1:44:44nice um so here we are looking at its
- 1:44:47sft mixture so this is the data mixture
- 1:44:49of um the fine tuning so this is the
- 1:44:52conversations data it right and so the
- 1:44:54way that they are solving it for Theo
- 1:44:56model is we see that there's a bunch of
- 1:44:58stuff in the mixture and there's a total
- 1:44:59of 1 million conversations here but here
- 1:45:02we have alot to hardcoded if we go there
- 1:45:05we see that this is 240
- 1:45:07conversations and look at these 240
- 1:45:10conversations they're hardcoded tell me
- 1:45:12about yourself says user and then the
- 1:45:15assistant says I'm and open language
- 1:45:17model developed by AI to Allen Institute
- 1:45:19of artificial intelligence Etc I'm here
- 1:45:21to help blah blah blah what is your name
- 1:45:23uh Theo project so these are all kinds
- 1:45:26of like cooked up hardcoded questions
- 1:45:27abouto 2 and the correct answers to give
- 1:45:30in these cases if you take 240 questions
- 1:45:33like this or conversations put them into
- 1:45:35your training set and fine tune with it
- 1:45:37then the model will actually be expected
- 1:45:39to parot this stuff later if you don't
- 1:45:43give it this then it's probably a Chach
- 1:45:45by open
- 1:45:46Ai and um there's one more way to
- 1:45:49sometimes do this is
- 1:45:51that basically um in these conversations
- 1:45:55and you have terms between human and
- 1:45:56assistant sometimes there's a special
- 1:45:58message called system message at the
- 1:46:00very beginning of the conversation so
- 1:46:02it's not just between human and
- 1:46:03assistant there's a system and in the
- 1:46:05system message you can actually hardcode
- 1:46:07and remind the model that hey you are a
- 1:46:10model developed by open Ai and your name
- 1:46:13is chashi pt40 and you were trained on
- 1:46:16this date and your knowledge cut off is
- 1:46:18this and basically it kind of like
- 1:46:19documents the model a little bit and
- 1:46:21then this is inserted into to your
- 1:46:23conversations so when you go on chpt you
- 1:46:25see a blank page but actually the system
- 1:46:27message is kind of like hidden in there
- 1:46:28and those tokens are in the context
- 1:46:30window and so those are the two ways to
- 1:46:33kind of um program the models to talk
- 1:46:35about themselves either it's done
- 1:46:37through uh data like this or it's done
- 1:46:40through system message and things like
- 1:46:42that basically invisible tokens that are
- 1:46:44in the context window and remind the
- 1:46:45model of its identity but it's all just
- 1:46:47kind of like cooked up and bolted on in
- 1:46:50some in some way it's not actually like
- 1:46:51really deeply there in any real sense as
- 1:46:54it would before a human I want to now
- 1:46:57continue to the next section which deals
- 1:46:59with the computational capabilities or
- 1:47:01like I should say the native
- 1:47:02computational capabilities of these
- 1:47:03models in problem solving scenarios and
- 1:47:06so in particular we have to be very
- 1:47:07careful with these models when we
- 1:47:09construct our examples of conversations
- 1:47:11and there's a lot of sharp edges here
- 1:47:13that are kind of like elucidative is
- 1:47:15that a word uh they're kind of like
- 1:47:16interesting to look at when we consider
- 1:47:18how these models think so um consider
- 1:47:22the following prompt from a human and
- 1:47:24supposed that basically that we are
- 1:47:25building out a conversation to enter
- 1:47:27into our training set of conversations
- 1:47:29so we're going to train the model on
- 1:47:30this we're teaching you how to basically
- 1:47:32solve simple math problems so the prompt
- 1:47:34is Emily buys three apples and two
- 1:47:36oranges each orange cost $2 the total
- 1:47:38cost is 13 what is the cost of apples
- 1:47:41very simple math question now there are
- 1:47:43two answers here on the left and on the
- 1:47:45right they are both correct answers they
- 1:47:48both say that the answer is three which
- 1:47:49is correct but one of these two is a
- 1:47:52significant ific anly better answer for
- 1:47:54the assistant than the other like if I
- 1:47:56was Data labeler and I was creating one
- 1:47:57of these one of these would be uh a
- 1:48:01really terrible answer for the assistant
- 1:48:03and the other would be okay and so I'd
- 1:48:05like you to potentially pause the video
- 1:48:07Even and think through why one of these
- 1:48:09two is significantly better answer uh
- 1:48:12than the other and um if you use the
- 1:48:14wrong one your model will actually be uh
- 1:48:17really bad at math potentially and it
- 1:48:19would have uh bad outcomes and this is
- 1:48:21something that you would be careful with
- 1:48:22in your life labeling documentations
- 1:48:23when you are training people uh to
- 1:48:25create the ideal responses for the
- 1:48:27assistant okay so the key to this
- 1:48:29question is to realize and remember that
- 1:48:32when the models are training and also
- 1:48:34inferencing they are working in
- 1:48:35onedimensional sequence of tokens from
- 1:48:37left to right and this is the picture
- 1:48:40that I often have in my mind I imagine
- 1:48:42basically the token sequence evolving
- 1:48:43from left to right and to always produce
- 1:48:46the next token in a sequence we are
- 1:48:48feeding all these tokens into the neural
- 1:48:50network and this neural network then is
- 1:48:53the probabilities for the next token and
- 1:48:54sequence right so this picture here is
- 1:48:56the exact same picture we saw uh before
- 1:48:58up here and this comes from the web demo
- 1:49:01that I showed you before right so this
- 1:49:04is the calculation that basically takes
- 1:49:05the input tokens here on the top and uh
- 1:49:09performs these operations of all these
- 1:49:11neurons and uh gives you the answer for
- 1:49:13the probabilities of what comes next now
- 1:49:15the important thing to realize is that
- 1:49:17roughly
- 1:49:19speaking uh there's basically a finite
- 1:49:21number of layers of computation that
- 1:49:22happened here so for example this model
- 1:49:25here has only one two three layers of
- 1:49:28what's called detention and uh MLP here
- 1:49:31um maybe um typical modern
- 1:49:34state-of-the-art Network would have more
- 1:49:36like say 100 layers or something like
- 1:49:37that but there's only 100 layers of
- 1:49:39computation or something like that to go
- 1:49:40from the previous token sequence to the
- 1:49:42probabilities for the next token and so
- 1:49:44there's a finite amount of computation
- 1:49:46that happens here for every single token
- 1:49:49and you should think of this as a very
- 1:49:50small amount of computation and this
- 1:49:52amount of computation is almost roughly
- 1:49:54fixed uh for every single token in this
- 1:49:57sequence um the that's not actually
- 1:49:59fully true because the more tokens you
- 1:50:01feed in uh the the more expensive uh
- 1:50:04this forward pass will be of this neural
- 1:50:06network but not by much so you should
- 1:50:09think of this uh and I think as a good
- 1:50:10model to have in mind this is a fixed
- 1:50:12amount of compute that's going to happen
- 1:50:13in this box for every single one of
- 1:50:15these tokens and this amount of compute
- 1:50:17Cann possibly be too big because there's
- 1:50:19not that many layers that are sort of
- 1:50:21going from the top to bottom here
- 1:50:23there's not that that much
- 1:50:24computationally that will happen here
- 1:50:26and so you can't imagine the model to to
- 1:50:27basically do arbitrary computation in a
- 1:50:29single forward pass to get a single
- 1:50:31token and so what that means is that we
- 1:50:34actually have to distribute our
- 1:50:35reasoning and our computation across
- 1:50:37many tokens because every single token
- 1:50:40is only spending a finite amount of
- 1:50:41computation on it and so we kind of want
- 1:50:45to distribute the computation across
- 1:50:47many tokens and we can't have too much
- 1:50:50computation or expect too much
- 1:50:52computation out of of the model in any
- 1:50:53single individual token because there's
- 1:50:55only so much computation that happens
- 1:50:57per token okay roughly fixed amount of
- 1:51:00computation here
- 1:51:02so that's why this answer here is
- 1:51:06significantly worse and the reason for
- 1:51:07that is Imagine going from left to right
- 1:51:09here um and I copy pasted it right here
- 1:51:13the answer is three Etc imagine the
- 1:51:16model having to go from left to right
- 1:51:17emitting these tokens one at a time it
- 1:51:19has to say or we're expecting to say the
- 1:51:23answer is space dollar sign and then
- 1:51:27right here we're expecting it to
- 1:51:28basically cram all of the computation of
- 1:51:30this problem into this single token it
- 1:51:32has to emit the correct answer three and
- 1:51:35then once we've emitted the answer three
- 1:51:37we're expecting it to say all these
- 1:51:39tokens but at this point we've already
- 1:51:41prod produced the answer and it's
- 1:51:43already in the context window for all
- 1:51:44these tokens that follow so anything
- 1:51:46here is just um kind of post Hawk
- 1:51:49justification of why this is the answer
- 1:51:52um because the answer is already created
- 1:51:53it's already in the token window so it's
- 1:51:56it's not actually being calculated here
- 1:51:58um and so if you are answering the
- 1:52:01question directly and immediately you
- 1:52:03are training the model to to try to
- 1:52:06basically guess the answer in a single
- 1:52:07token and that is just not going to work
- 1:52:10because of the finite amount of
- 1:52:11computation that happens per token
- 1:52:13that's why this answer on the right is
- 1:52:15significantly better because we are
- 1:52:17Distributing this computation across the
- 1:52:19answer we're actually getting the model
- 1:52:20to sort of slowly come to the answer
- 1:52:23from the left to right we're getting
- 1:52:24intermediate results we're saying okay
- 1:52:26the total cost of oranges is four so 30
- 1:52:28- 4 is 9 and so we're creating
- 1:52:32intermediate calculations and each one
- 1:52:34of these calculations is by itself not
- 1:52:36that expensive and so we're actually
- 1:52:38basically kind of guessing a little bit
- 1:52:40the difficulty that the model is capable
- 1:52:42of in any single one of these individual
- 1:52:44tokens and there can never be too much
- 1:52:47work in any one of these tokens
- 1:52:49computationally because then the model
- 1:52:50won't be able to do that later at test
- 1:52:52time and so we're teaching the model
- 1:52:55here to spread out its reasoning and to
- 1:52:57spread out its computation over the
- 1:52:59tokens and in this way it only has very
- 1:53:02simple problems in each token and they
- 1:53:05can add up and then by the time it's
- 1:53:07near the end it has all the previous
- 1:53:09results in its working memory and it's
- 1:53:11much easier for it to determine that the
- 1:53:13answer is and here it is three so this
- 1:53:15is a significantly better label for our
- 1:53:18computation this would be really bad and
- 1:53:20is teaching the model to try to do all
- 1:53:23the computation in a single token and
- 1:53:24it's really
- 1:53:25bad so uh that's kind of like an
- 1:53:28interesting thing to keep in mind is in
- 1:53:30your
- 1:53:31prompts uh usually don't have to think
- 1:53:33about it explicitly because uh the
- 1:53:36people at open AI have labelers and so
- 1:53:38on that actually worry about this and
- 1:53:40they make sure that the answers are
- 1:53:41spread out and so actually open AI will
- 1:53:43kind of like do the right thing so when
- 1:53:45I ask this question for chat GPT it's
- 1:53:48actually going to go very slowly it's
- 1:53:49going to be like okay let's define our
- 1:53:50variables set up the equation
- 1:53:52and it's kind of creating all these
- 1:53:54intermediate results these are not for
- 1:53:56you these are for the model if the model
- 1:53:58is not creating these intermediate
- 1:53:59results for itself it's not going to be
- 1:54:01able to reach three I also wanted to
- 1:54:04show you that it's possible to be a bit
- 1:54:06mean to the model uh we can just ask for
- 1:54:08things so as an example I said I gave it
- 1:54:10the exact same uh prompt and I said
- 1:54:13answer the question in a single token
- 1:54:15just immediately give me the answer
- 1:54:16nothing else and it turns out that for
- 1:54:18this simple um prompt here it actually
- 1:54:21was able to do it in single go so it
- 1:54:23just created a single I think this is
- 1:54:25two tokens right uh because the dollar
- 1:54:27sign is its own token so basically this
- 1:54:30model didn't give me a single token it
- 1:54:31gave me two tokens but it still produced
- 1:54:33the correct answer and it did that in a
- 1:54:35single forward pass of the
- 1:54:37network now that's because the numbers
- 1:54:40here I think are very simple and so I
- 1:54:41made it a bit more difficult to be a bit
- 1:54:43mean to the model so I said Emily buys
- 1:54:4523 apples and 177 oranges and then I
- 1:54:48just made the numbers a bit bigger and
- 1:54:50I'm just making it harder for the model
- 1:54:51I'm asking it to more computation in a
- 1:54:53single token and so I said the same
- 1:54:55thing and here it gave me five and five
- 1:54:58is actually not correct so the model
- 1:55:00failed to do all of this calculation in
- 1:55:02a single forward pass of the network it
- 1:55:04failed to go from the input tokens and
- 1:55:07then in a single forward pass of the
- 1:55:09network single go through the network it
- 1:55:11couldn't produce the result and then I
- 1:55:13said okay now don't worry about the the
- 1:55:16token limit and just solve the problem
- 1:55:18as usual and then it goes all the
- 1:55:20intermediate results it simplifies and
- 1:55:22every one of these intermediate results
- 1:55:24here and intermediate calculations is
- 1:55:26much easier for the model and um it sort
- 1:55:29of it's not too much work per token all
- 1:55:32of the tokens here are correct and it
- 1:55:33arises the solution which is seven and I
- 1:55:36just couldn't squeeze all of this work
- 1:55:38it couldn't squeeze that into a single
- 1:55:39forward passive Network so I think
- 1:55:41that's kind of just a cute example and
- 1:55:43something to kind of like think about
- 1:55:45and I think it's kind of again just
- 1:55:46elucidative in terms of how these uh
- 1:55:48models work the last thing that I would
- 1:55:50say on this topic is that if I was in
- 1:55:52practi is trying to actually solve this
- 1:55:53in my day-to-day life I might actually
- 1:55:55not uh trust that the model that all the
- 1:55:57intermediate calculations correctly here
- 1:55:59so actually probably what I do is
- 1:56:01something like this I would come here
- 1:56:02and I would say use code and uh that's
- 1:56:06because code is one of the possible
- 1:56:08tools that chachy PD can use and instead
- 1:56:11of it having to do mental arithmetic
- 1:56:14like this mental arithmetic here I don't
- 1:56:15fully trust it and especially if the
- 1:56:17numbers get really big there's no
- 1:56:19guarantee that the model will do this
- 1:56:20correctly any one of these intermediates
- 1:56:22steps might in principle fail we're
- 1:56:24using neural networks to do mental
- 1:56:26arithmetic uh kind of like you doing
- 1:56:27mental arithmetic in your brain it might
- 1:56:30just like uh screw up some of the
- 1:56:31intermediate results it's actually kind
- 1:56:32of amazing that it can even do this kind
- 1:56:34of mental arithmetic I don't think I
- 1:56:35could do this in my head but basically
- 1:56:37the model is kind of like doing it in
- 1:56:38its head and I don't trust that so I
- 1:56:40wanted to use tools so you can say stuff
- 1:56:42like use
- 1:56:43code and uh I'm not sure what happened
- 1:56:47there use
- 1:56:50code and so um like I mentioned there's
- 1:56:53a special tool and the uh the model can
- 1:56:55write code and I can inspect that this
- 1:56:58code is correct and then uh it's not
- 1:57:01relying on its mental arithmetic it is
- 1:57:03using the python interpreter which is a
- 1:57:05very simple programming language to
- 1:57:07basically uh write out the code that
- 1:57:08calculates the result and I would
- 1:57:10personally trust this a lot more because
- 1:57:12this came out of a Python program which
- 1:57:14I think has a lot more correctness
- 1:57:15guarantees than the mental arithmetic of
- 1:57:17a language model uh so just um another
- 1:57:21kind of uh potential hint that if you
- 1:57:23have these kinds of problems uh you may
- 1:57:24want to basically just uh ask the model
- 1:57:26to use the code interpreter and just
- 1:57:28like we saw with the web search the
- 1:57:30model has special uh kind of tokens for
- 1:57:34calling uh like it will not actually
- 1:57:36generate these tokens from the language
- 1:57:38model it will write the program and then
- 1:57:40it actually sends that program to a
- 1:57:42different sort of part of the computer
- 1:57:44that actually just runs that program and
- 1:57:46brings back the result and then the
- 1:57:48model gets access to that result and can
- 1:57:50tell you that okay the cost of each
- 1:57:51apple is seven
- 1:57:53um so that's another kind of tool and I
- 1:57:55would use this in practice for yourself
- 1:57:57and it's um yeah it's just uh less error
- 1:58:01prone I would say so that's why I called
- 1:58:03this section models need tokens to think
- 1:58:06distribute your competition across many
- 1:58:08tokens ask models to create intermediate
- 1:58:10results or whenever you can lean on
- 1:58:13tools and Tool use instead of allowing
- 1:58:15the models to do all of the stuff in
- 1:58:17their memory so if they try to do it all
- 1:58:18in their memory I don't fully trust it
- 1:58:21and prefer to use tools whenever
- 1:58:22possible I want to show you one more
- 1:58:24example of where this actually comes up
- 1:58:26and that's in counting so models
- 1:58:28actually are not very good at counting
- 1:58:30for the exact same reason you're asking
- 1:58:32for way too much in a single individual
- 1:58:34token so let me show you a simple
- 1:58:36example of that um how many dots are
- 1:58:38below and then I just put in a bunch of
- 1:58:41dots and Chach says there are and then
- 1:58:44it just tries to solve the problem in a
- 1:58:46single token so in a single token it has
- 1:58:49to count the number of dots in its
- 1:58:51context window
- 1:58:53um and it has to do that in the single
- 1:58:55forward pass of a network and a single
- 1:58:57forward pass of a network as we talked
- 1:58:58about there's not that much computation
- 1:59:00that can happen there just think of that
- 1:59:01as being like very little competation
- 1:59:03that happens there so if I just look at
- 1:59:06what the model sees let's go to the LM
- 1:59:09go to tokenizer it sees uh
- 1:59:13this how many dots are below and then it
- 1:59:15turns out that these dots here this
- 1:59:17group of I think 20 dots is a single
- 1:59:20token and then this group of whatever it
- 1:59:22is is another token and then for some
- 1:59:25reason they break up as this so I don't
- 1:59:28actually this has to do with the details
- 1:59:29of the tokenizer but it turns out that
- 1:59:31these um the model basically sees the
- 1:59:34token ID this this this and so on and
- 1:59:38then from these token IDs it's expected
- 1:59:40to count the number and spoiler alert is
- 1:59:43not 161 it's actually I believe
- 1:59:45177 so here's what we can do instead uh
- 1:59:48we can say use code and you might expect
- 1:59:51that like why should this work and it's
- 1:59:54actually kind of subtle and kind of
- 1:59:55interesting so when I say use code I
- 1:59:57actually expect this to work let's see
- 1:59:59okay 177 is correct so what happens here
- 2:00:02is I've actually it doesn't look like it
- 2:00:04but I've broken down the problem into a
- 2:00:08problems that are easier for the model I
- 2:00:10know that the model can't count it can't
- 2:00:12do mental counting but I know that the
- 2:00:14model is actually pretty good at doing
- 2:00:15copy pasting so what I'm doing here is
- 2:00:18when I say use code it creates a string
- 2:00:20in Python for this and the task of
- 2:00:23basically copy pasting my input here to
- 2:00:27here is very simple because for the
- 2:00:29model um it sees this string of uh it
- 2:00:33sees it as just these four tokens or
- 2:00:35whatever it is so it's very simple for
- 2:00:37the model to copy paste those token IDs
- 2:00:40and um kind of unpack them into Dots
- 2:00:45here and so it creates this string and
- 2:00:47then it calls python routine. count and
- 2:00:50then it comes up with the correct answer
- 2:00:52so the python interpreter is doing the
- 2:00:53counting it's not the models mental
- 2:00:55arithmetic doing the counting so it's
- 2:00:57again a simple example of um models need
- 2:01:00tokens to think don't rely on their
- 2:01:02mental arithmetic and um that's why also
- 2:01:05the models are not very good at counting
- 2:01:07if you need them to do counting tasks
- 2:01:08always ask them to lean on the tool now
- 2:01:11the models also have many other little
- 2:01:13cognitive deficits here and there and
- 2:01:15these are kind of like sharp edges of
- 2:01:16the technology to be kind of aware of
- 2:01:18over time so as an example the models
- 2:01:20are not very good with all kinds of
- 2:01:22spelling related tasks they're not very
- 2:01:24good at it and I told you that we would
- 2:01:26loop back around to tokenization and the
- 2:01:29reason to do for this is that the models
- 2:01:31they don't see the characters they see
- 2:01:33tokens and they their entire world is
- 2:01:35about tokens which are these little text
- 2:01:37chunks and so they don't see characters
- 2:01:39like our eyes do and so very simple
- 2:01:41character level tasks often fail so for
- 2:01:45example uh I'm giving it a string
- 2:01:47ubiquitous and I'm asking it to print
- 2:01:49only every third character starting with
- 2:01:51the first one so we start with U and
- 2:01:54then we should go every third so every
- 2:01:56so 1 2 3 Q should be next and then Etc
- 2:02:01so this I see is not correct and again
- 2:02:03my hypothesis is that this is again
- 2:02:05Dental arithmetic here is failing number
- 2:02:08one a little bit but number two I think
- 2:02:10the the more important issue here is
- 2:02:12that if you go to Tik
- 2:02:13tokenizer and you look at ubiquitous we
- 2:02:16see that it is three tokens right so you
- 2:02:19and I see ubiquitous and we can easily
- 2:02:21access the individual letters because we
- 2:02:23kind of see them and when we have it in
- 2:02:25the working memory of our visual sort of
- 2:02:27field we can really easily index into
- 2:02:29every third letter and I can do that
- 2:02:31task but the models don't have access to
- 2:02:33the individual letters they see this as
- 2:02:35these three tokens and uh remember these
- 2:02:38models are trained from scratch on the
- 2:02:39internet and all these token uh
- 2:02:42basically the model has to discover how
- 2:02:44many of all these different letters are
- 2:02:45packed into all these different tokens
- 2:02:47and the reason we even use tokens is
- 2:02:49mostly for efficiency uh but I think a
- 2:02:51lot of people areed interested to delete
- 2:02:52tokens entirely like we should really
- 2:02:54have character level or bite level
- 2:02:56models it's just that that would create
- 2:02:58very long sequences and people don't
- 2:02:59know how to deal with that right now so
- 2:03:01while we have the token World any kind
- 2:03:03of spelling tasks are not actually
- 2:03:05expected to work super well so because I
- 2:03:07know that spelling is not a strong suit
- 2:03:09because of tokenization I can again Ask
- 2:03:11it to lean On Tools so I can just say
- 2:03:13use code and I would again expect this
- 2:03:16to work because the task of copy pasting
- 2:03:18ubiquitous into the python interpreter
- 2:03:20is much easier and then we're leaning on
- 2:03:22python interpreter to manipulate the
- 2:03:25characters of this string so when I say
- 2:03:27use
- 2:03:28code
- 2:03:30ubiquitous yes it indexes into every
- 2:03:32third character and the actual truth is
- 2:03:35u2s
- 2:03:36uqs uh which looks correct to me so um
- 2:03:41again an example of spelling related
- 2:03:42tasks not working very well a very
- 2:03:44famous example of that recently is how
- 2:03:47many R are there in strawberry and this
- 2:03:49went viral many times and basically the
- 2:03:51models now get it correct they say there
- 2:03:53are three Rs in Strawberry but for a
- 2:03:55very long time all the state-of-the-art
- 2:03:56models would insist that there are only
- 2:03:58two RS in strawberry and this caused a
- 2:04:00lot of you know Ruckus because is that a
- 2:04:03word I think so because um it just kind
- 2:04:06of like why are the models so brilliant
- 2:04:08and they can solve math Olympiad
- 2:04:10questions but they can't like count RS
- 2:04:12in strawberry and the answer for that
- 2:04:14again is I've got built up to it kind of
- 2:04:16slowly but number one the models don't
- 2:04:18see characters they see tokens and
- 2:04:20number two they are not very good at
- 2:04:22counting and so here we are combining
- 2:04:25the difficulty of seeing the characters
- 2:04:27with the difficulty of counting and
- 2:04:29that's why the models struggled with
- 2:04:30this even though I think by now honestly
- 2:04:33I think open I may have hardcoded the
- 2:04:34answer here or I'm not sure what they
- 2:04:35did but um uh but this specific query
- 2:04:39now works
- 2:04:41so models are not very good at spelling
- 2:04:44and there there's a bunch of other
- 2:04:45little sharp edges and I don't want to
- 2:04:46go into all of them I just want to show
- 2:04:48you a few examples of things to be aware
- 2:04:50of and uh when you're using these models
- 2:04:52in practice I don't actually want to
- 2:04:54have a comprehensive analysis here of
- 2:04:55all the ways that the models are kind of
- 2:04:57like falling short I just want to make
- 2:04:59the point that there are some Jagged
- 2:05:01edges here and there and we've discussed
- 2:05:03a few of them and a few of them make
- 2:05:05sense but some of them also will just
- 2:05:06not make as much sense and they're kind
- 2:05:08of like you're left scratching your head
- 2:05:10even if you understand in- depth how
- 2:05:11these models work and and good example
- 2:05:14of that recently is the following uh the
- 2:05:16models are not very good at very simple
- 2:05:17questions like this and uh this is
- 2:05:20shocking to a lot of people because
- 2:05:22these math uh these problems can solve
- 2:05:23complex math problems they can answer
- 2:05:25PhD grade physics chemistry biology
- 2:05:28questions much better than I can but
- 2:05:30sometimes they fall short in like super
- 2:05:31simple problems like this so here we go
- 2:05:349.11 is bigger than 9.9 and it justifies
- 2:05:38it in some way but obviously and then at
- 2:05:40the end okay it actually it flips its
- 2:05:44decision later so um I don't believe
- 2:05:47that this is very reproducible sometimes
- 2:05:49it flips around its answer sometimes
- 2:05:50gets it right sometimes get it get it
- 2:05:52wrong uh let's try
- 2:05:56again okay even though it might look
- 2:05:59larger okay so here it doesn't even
- 2:06:01correct itself in the end if you ask
- 2:06:03many times sometimes it gets it right
- 2:06:04too but how is it that the model can do
- 2:06:07so great at Olympiad grade problems but
- 2:06:10then fail on very simple problems like
- 2:06:12this and uh I think this one is as I
- 2:06:15mentioned a little bit of a head
- 2:06:16scratcher it turns out that a bunch of
- 2:06:18people studied this in depth and I
- 2:06:19haven't actually read the paper uh but
- 2:06:22what I was told by this team was that
- 2:06:24when you scrutinize the activations
- 2:06:27inside the neural network when you look
- 2:06:29at some of the features and what what
- 2:06:31features turn on or off and what neurons
- 2:06:33turn on or off uh a bunch of neurons
- 2:06:35inside the neural network light up that
- 2:06:37are usually associated with Bible verses
- 2:06:40U and so I think the model is kind of
- 2:06:42like reminded that these almost look
- 2:06:44like Bible verse markers and in a bip
- 2:06:48verse setting 9.11 would come after 99.9
- 2:06:52and so basically the model somehow finds
- 2:06:53it like cognitively very distracting
- 2:06:56that in Bible verses 9.11 would be
- 2:06:58greater um even though here it's
- 2:07:00actually trying to justify it and come
- 2:07:02up to the answer with a math it still
- 2:07:04ends up with the wrong answer here so it
- 2:07:07basically just doesn't fully make sense
- 2:07:08and it's not fully understood and um
- 2:07:12there's a few Jagged issues like that so
- 2:07:14that's why treat this as a as what it is
- 2:07:17which is a St stochastic system that is
- 2:07:19really magical but that you can't also
- 2:07:21fully trust and you want to use it as a
- 2:07:23tool not as something that you kind of
- 2:07:25like letter rip on a problem and
- 2:07:27copypaste the results okay so we have
- 2:07:29now covered two major stages of training
- 2:07:32of large language models we saw that in
- 2:07:34the first stage this is called the
- 2:07:36pre-training stage we are basically
- 2:07:38training on internet documents and when
- 2:07:40you train a language model on internet
- 2:07:42documents you get what's called a base
- 2:07:44model and it's basically an internet
- 2:07:45document simulator right now we saw that
- 2:07:48this is an interesting artifact and uh
- 2:07:51this takes many months to train on
- 2:07:53thousands of computers and it's kind of
- 2:07:54a lossy compression of the internet and
- 2:07:57it's extremely interesting but it's not
- 2:07:58directly useful because we don't want to
- 2:08:00sample internet documents we want to ask
- 2:08:02questions of an AI and have it respond
- 2:08:05to our questions so for that we need an
- 2:08:07assistant and we saw that we can
- 2:08:09actually construct an assistant in the
- 2:08:11process of a post
- 2:08:13training and specifically in the process
- 2:08:16of supervised fine-tuning as we call
- 2:08:19it so in this stage we saw that it's
- 2:08:22algorithmically identical to
- 2:08:24pre-training nothing is going to change
- 2:08:25the only thing that changes is the data
- 2:08:27set so instead of Internet documents we
- 2:08:30now want to create and curate a very
- 2:08:32nice data set of conversations so we
- 2:08:35want Millions conversations on all kinds
- 2:08:38of diverse topics between a human and an
- 2:08:41assistant and fundamentally these
- 2:08:44conversations are created by humans so
- 2:08:47humans write the prompts and humans
- 2:08:49write the ideal response responses and
- 2:08:52they do that based on labeling
- 2:08:54documentations now in the modern stack
- 2:08:57it's not actually done fully and
- 2:08:59manually by humans right they actually
- 2:09:00now have a lot of help from these tools
- 2:09:02so we can use language models um to help
- 2:09:05us create these data sets and that's
- 2:09:07done extensively but fundamentally it's
- 2:09:09all still coming from Human curation at
- 2:09:10the end so we create these conversations
- 2:09:13that now becomes our data set we fine
- 2:09:15tune on it or continue training on it
- 2:09:17and we get an assistant and then we kind
- 2:09:20of shifted gears and started talking
- 2:09:21about some of the kind of cognitive
- 2:09:22implications of what this assistant is
- 2:09:24like and we saw that for example the
- 2:09:26assistant will hallucinate if you don't
- 2:09:29take some sort of mitigations towards it
- 2:09:32so we saw that hallucinations would be
- 2:09:34common and then we looked at some of the
- 2:09:35mitigations of those hallucinations and
- 2:09:38then we saw that the models are quite
- 2:09:39impressive and can do a lot of stuff in
- 2:09:40their head but we saw that they can also
- 2:09:43Lean On Tools to become better so for
- 2:09:45example we can lo lean on a web search
- 2:09:48in order to hallucinate less and to
- 2:09:50maybe bring up some more um recent
- 2:09:53information or something like that or we
- 2:09:54can lean on tools like code interpreter
- 2:09:57so the code can so the llm can write
- 2:09:59some code and actually run it and see
- 2:10:00the
- 2:10:01results so these are some of the topics
- 2:10:03we looked at so far um now what I'd like
- 2:10:06to do is I'd like to cover the last and
- 2:10:09major stage of this Pipeline and that is
- 2:10:12reinforcement learning so reinforcement
- 2:10:15learning is still kind of thought to be
- 2:10:16under the umbrella of posttraining uh
- 2:10:19but it is the last third major stage and
- 2:10:22it's a different way of training
- 2:10:24language models and usually follows as
- 2:10:26this third step so inside companies like
- 2:10:29open AI you will start here and these
- 2:10:31are all separate teams so there's a team
- 2:10:33doing data for pre-training and a team
- 2:10:35doing training for pre-training and then
- 2:10:37there's a team doing all the
- 2:10:39conversation generation in a in a
- 2:10:42different team that is kind of doing the
- 2:10:44supervis fine tuning and there will be a
- 2:10:45team for the reinforcement learning as
- 2:10:47well so it's kind of like a handoff of
- 2:10:49these models you get your base model the
- 2:10:51then you find you need to be an
- 2:10:52assistant and then you go into
- 2:10:53reinforcement learning which we'll talk
- 2:10:55about uh
- 2:10:56now so that's kind of like the major
- 2:10:58flow and so let's now focus on
- 2:11:01reinforcement learning the last major
- 2:11:03stage of training and let me first
- 2:11:05actually motivate it and why we would
- 2:11:07want to do reinforcement learning and
- 2:11:09what it looks like on a high level so I
- 2:11:11would now like to try to motivate the
- 2:11:12reinforcement learning stage and what it
- 2:11:13corresponds to with something that
- 2:11:15you're probably familiar with and that
- 2:11:16is basically going to school so just
- 2:11:19like you went to school to become um
- 2:11:21really good at something we want to take
- 2:11:23large language models through school and
- 2:11:25really what we're doing is um we're um
- 2:11:29we have a few paradigms of ways of uh
- 2:11:32giving them knowledge or transferring
- 2:11:33skills so in particular when we're
- 2:11:36working with textbooks in school you'll
- 2:11:38see that there are three major kind of
- 2:11:40uh pieces of information in these
- 2:11:42textbooks three classes of information
- 2:11:45the first thing you'll see is you'll see
- 2:11:46a lot of exposition um and by the way
- 2:11:49this is a totally random book I pulled
- 2:11:50from the internet I I think it's some
- 2:11:51kind of an organic chemistry or
- 2:11:53something I'm not sure uh but the
- 2:11:55important thing is that you'll see that
- 2:11:56most of the text most of it is kind of
- 2:11:58just like the meat of it is exposition
- 2:12:00it's kind of like background knowledge
- 2:12:02Etc as you are reading through the words
- 2:12:05of this Exposition you can think of that
- 2:12:08roughly as training on that data so um
- 2:12:12and that's why when you're reading
- 2:12:13through this stuff this background
- 2:12:14knowledge and this all this context
- 2:12:16information it's kind of equivalent to
- 2:12:18pre-training so it's it's where we build
- 2:12:21sort of like a knowledge base of this
- 2:12:23data and get a sense of the topic the
- 2:12:27next major kind of information that you
- 2:12:28will see is these uh problems and with
- 2:12:32their worked Solutions so basically a
- 2:12:35human expert in this case uh the author
- 2:12:37of this book has given us not just a
- 2:12:39problem but has also worked through the
- 2:12:41solution and the solution is basically
- 2:12:43like equivalent to having like this
- 2:12:45ideal response for an assistant so it's
- 2:12:48basically the expert is showing us how
- 2:12:49to solve the problem in it's uh kind of
- 2:12:52like um in its full form so as we are
- 2:12:55reading the solution we are basically
- 2:12:57training on the expert data and then
- 2:13:01later we can try to imitate the expert
- 2:13:03um and basically um that's that roughly
- 2:13:07correspond to having the sft model
- 2:13:08that's what it would be doing so
- 2:13:11basically we've already done
- 2:13:12pre-training and we've already covered
- 2:13:14this um imitation of experts and how
- 2:13:17they solve these problems and the third
- 2:13:19stage of reinforcement learning is
- 2:13:21basically the practice problems so
- 2:13:24sometimes you'll see this is just a
- 2:13:25single practice problem here but of
- 2:13:27course there will be usually many
- 2:13:28practice problems at the end of each
- 2:13:30chapter in any textbook and practice
- 2:13:32problems of course we know are critical
- 2:13:34for learning because what are they
- 2:13:36getting you to do they're getting you to
- 2:13:37practice uh to practice yourself and
- 2:13:39discover ways of solving these problems
- 2:13:42yourself and so what you get in a
- 2:13:44practice problem is you get a problem
- 2:13:46description but you're not given the
- 2:13:48solution but you are given the final
- 2:13:50answer answer usually in the answer key
- 2:13:53of the textbook and so you know the
- 2:13:55final answer that you're trying to get
- 2:13:56to and you have the problem statement
- 2:13:58but you don't have the solution you are
- 2:14:00trying to practice the solution you're
- 2:14:02trying out many different things and
- 2:14:04you're seeing what gets you to the final
- 2:14:07solution the best and so you're
- 2:14:09discovering how to solve these problems
- 2:14:11so and in the process of that you're
- 2:14:13relying on number one the background
- 2:14:14information which comes from
- 2:14:15pre-training and number two maybe a
- 2:14:17little bit of imitation of human experts
- 2:14:20and you can probably try similar kinds
- 2:14:22of solutions and so on so we've done
- 2:14:25this and this and now in this section
- 2:14:27we're going to try to practice and so
- 2:14:30we're going to be given prompts we're
- 2:14:32going to be given Solutions U sorry the
- 2:14:34final answers but we're not going to be
- 2:14:36given expert Solutions we have to
- 2:14:38practice and try stuff out and that's
- 2:14:40what reinforcement learning is about
- 2:14:43okay so let's go back to the problem
- 2:14:44that we worked with previously just so
- 2:14:46we have a concrete example to talk
- 2:14:47through as we explore sort of the topic
- 2:14:50here so um I'm here in the Teck
- 2:14:52tokenizer because I'd also like to well
- 2:14:55I get a text box which is useful but
- 2:14:57number two I want to remind you again
- 2:14:59that we're always working with
- 2:14:59onedimensional token sequences and so um
- 2:15:02I actually like prefer this view because
- 2:15:04this is like the native view of the llm
- 2:15:06if that makes sense like this is what it
- 2:15:08actually sees it sees token IDs right
- 2:15:11okay so Emily buys three apples and two
- 2:15:14oranges each orange is $2 the total cost
- 2:15:17of all the fruit is $13 what is the cost
- 2:15:19of each apple
- 2:15:21and what I'd like to what I like you to
- 2:15:23appreciate here is these are like four
- 2:15:26possible candidate Solutions as an
- 2:15:29example and they all reach the answer
- 2:15:31three now what I'd like you to
- 2:15:33appreciate at this point is that if I am
- 2:15:35the human data labeler that is creating
- 2:15:37a conversation to be entered into the
- 2:15:39training set I don't actually really
- 2:15:42know which of these
- 2:15:44conversations to um to add to the data
- 2:15:48set some of these conversations kind of
- 2:15:50set up a system equations some of them
- 2:15:52sort of like just talk through it in
- 2:15:54English and some of them just kind of
- 2:15:55like skip right through to the
- 2:15:58solution um if you look at chbt for
- 2:16:00example and you give it this question it
- 2:16:03defines a system of variables and it
- 2:16:05kind of like does this little thing what
- 2:16:07we have to appreciate and uh
- 2:16:08differentiate between though is um the
- 2:16:12first purpose of a solution is to reach
- 2:16:14the right answer of course we want to
- 2:16:15get the final answer three that is the
- 2:16:17that is the important purpose here but
- 2:16:19there's kind of like a secondary purpose
- 2:16:21as well where here we are also just kind
- 2:16:23of trying to make it like nice uh for
- 2:16:26the human because we're kind of assuming
- 2:16:27that the person wants to see the
- 2:16:29solution they want to see the
- 2:16:30intermediate steps we want to present it
- 2:16:31nicely Etc so there are two separate
- 2:16:33things going on here number one is the
- 2:16:36presentation for the human but number
- 2:16:37two we're trying to actually get the
- 2:16:38right answer um so let's for the moment
- 2:16:42focus on just reaching the final answer
- 2:16:44if we're only care if we only care about
- 2:16:46the final answer then which of these is
- 2:16:49the optimal or the best prompt um sorry
- 2:16:53the best solution for the llm to reach
- 2:16:56the right
- 2:16:57answer um and what I'm trying to get at
- 2:17:00is we don't know me as a human labeler I
- 2:17:03would not know which one of these is
- 2:17:04best so as an example we saw earlier on
- 2:17:07when we looked at
- 2:17:09um the token sequences here and the
- 2:17:11mental arithmetic and reasoning we saw
- 2:17:14that for each token we can only spend
- 2:17:15basically a finite number of finite
- 2:17:18amount of compute here that is not very
- 2:17:19large or you should think about it that
- 2:17:20way way and so we can't actually make
- 2:17:23too big of a leap in any one token is is
- 2:17:26maybe the way to think about it so as an
- 2:17:28example in this one what's really nice
- 2:17:30about it is that it's very few tokens so
- 2:17:32it's going to take us very short amount
- 2:17:34of time to get to the answer but right
- 2:17:37here when we're doing 30 - 4 IDE 3
- 2:17:39equals right in this token here we're
- 2:17:42actually asking for a lot of computation
- 2:17:44to happen on that single individual
- 2:17:45token and so maybe this is a bad example
- 2:17:48to give to the llm because it's kind of
- 2:17:49incentivizing it to skip through the
- 2:17:50calculations very quickly and it's going
- 2:17:52to actually make up mistakes make
- 2:17:54mistakes in this mental arithmetic uh so
- 2:17:56maybe it would work better to like
- 2:17:58spread out the spread it out more maybe
- 2:18:01it would be better to set it up as an
- 2:18:02equation maybe it would be better to
- 2:18:04talk through it we fundamentally don't
- 2:18:06know and we don't know because what is
- 2:18:09easy for you or I as or as human
- 2:18:12labelers what's easy for us or hard for
- 2:18:14us is different than what's easy or hard
- 2:18:16for the llm it cognition is different um
- 2:18:20and the token sequences are kind of like
- 2:18:23different hard for it and so some of the
- 2:18:27token sequences here that are trivial
- 2:18:30for me might be um very too much of a
- 2:18:33leap for the llm so right here this
- 2:18:36token would be way too hard but
- 2:18:38conversely many of the tokens that I'm
- 2:18:40creating here might be just trivial to
- 2:18:43the llm and we're just wasting tokens
- 2:18:45like why waste all these tokens when
- 2:18:46this is all trivial so if the only thing
- 2:18:49we care care about is the final answer
- 2:18:51and we're separating out the issue of
- 2:18:53the presentation to the human um then we
- 2:18:56don't actually really know how to
- 2:18:57annotate this example we don't know what
- 2:18:59solution to get to the llm because we
- 2:19:01are not the
- 2:19:02llm and it's clear here in the case of
- 2:19:05like the math example but this is
- 2:19:07actually like a very pervasive issue
- 2:19:08like for our knowledge is not lm's
- 2:19:11knowledge like the llm actually has a
- 2:19:13ton of knowledge of PhD in math and
- 2:19:15physics chemistry and whatnot so in many
- 2:19:17ways it actually knows more than I do
- 2:19:19and I'm I'm potentially not utilizing
- 2:19:21that knowledge in its problem solving
- 2:19:24but conversely I might be injecting a
- 2:19:26bunch of knowledge in my solutions that
- 2:19:28the LM doesn't know in its parameters
- 2:19:31and then those are like sudden leaps
- 2:19:33that are very confusing to the model and
- 2:19:36so our cognitions are different and I
- 2:19:38don't really know what to put here if
- 2:19:41all we care about is the reaching the
- 2:19:42final solution and doing it economically
- 2:19:45ideally and so long story short we are
- 2:19:49not in a good position to create these
- 2:19:52uh token sequences for the LM and
- 2:19:55they're useful by imitation to
- 2:19:56initialize the system but we really want
- 2:19:59the llm to discover the token sequences
- 2:20:01that work for it we need to find it
- 2:20:04needs to find for itself what token
- 2:20:06sequence reliably gets to the answer
- 2:20:09given the prompt and it needs to
- 2:20:11discover that in the process of
- 2:20:12reinforcement learning and of trial and
- 2:20:14error so let's see how this example
- 2:20:18would work like in reinforcement
- 2:20:19learning
- 2:20:21okay so we're now back in the huging
- 2:20:23face inference playground and uh that
- 2:20:26just allows me to very easily call uh
- 2:20:28different kinds of models so as an
- 2:20:29example here on the top right I chose
- 2:20:31the Gemma 2 2 billion parameter model so
- 2:20:34two billion is very very small so this
- 2:20:36is a tiny model but it's okay so we're
- 2:20:39going to give it um the way that
- 2:20:40reinforcement learning will basically
- 2:20:41work is actually quite quite simple um
- 2:20:44we need to try many different kinds of
- 2:20:47solutions and we want to see which
- 2:20:49Solutions work well or not
- 2:20:51so we're basically going to take the
- 2:20:53prompt we're going to run the
- 2:20:55model and the model generates a solution
- 2:20:58and then we're going to inspect the
- 2:20:59solution and we know that the correct
- 2:21:02answer for this one is $3 and so indeed
- 2:21:05the model gets it correct it says it's
- 2:21:06$3 so this is correct so that's just one
- 2:21:10attempt at DIS solution so now we're
- 2:21:11going to delete this and we're going to
- 2:21:13rerun it again let's try a second
- 2:21:15attempt so the model solves it in a bit
- 2:21:17slightly different way right every
- 2:21:19single attempt will be a different
- 2:21:21generation because these models are
- 2:21:23stochastic systems remember that at
- 2:21:24every single token here we have a
- 2:21:26probability distribution and we're
- 2:21:27sampling from that distribution so we
- 2:21:29end up kind kind of going down slightly
- 2:21:31different paths and so this is a second
- 2:21:34solution that also ends in the correct
- 2:21:36answer now we're going to delete that
- 2:21:38let's go a third
- 2:21:39time okay so again slightly different
- 2:21:42solution but also gets it
- 2:21:44correct now we can actually repeat this
- 2:21:46uh many times and so in practice you
- 2:21:49might actually sample thousand of
- 2:21:51independent Solutions or even like
- 2:21:52million solutions for just a single
- 2:21:55prompt um and some of them will be
- 2:21:57correct and some of them will not be
- 2:21:58very correct and basically what we want
- 2:22:00to do is we want to encourage the
- 2:22:02solutions that lead to correct answers
- 2:22:05so let's take a look at what that looks
- 2:22:06like so if we come back over here here's
- 2:22:09kind of like a cartoon diagram of what
- 2:22:10this is looking like we have a prompt
- 2:22:13and then we tried many different
- 2:22:15solutions in
- 2:22:16parallel and some of the solutions um
- 2:22:19might go well so they get the right
- 2:22:21answer which is in green and some of the
- 2:22:24solutions might go poorly and may not
- 2:22:25reach the right answer which is red now
- 2:22:28this problem here unfortunately is not
- 2:22:29the best example because it's a trivial
- 2:22:32prompt and as we saw uh even like a two
- 2:22:34billion parameter model always gets it
- 2:22:36right so it's not the best example in
- 2:22:38that sense but let's just exercise some
- 2:22:40imagination here and let's just suppose
- 2:22:43that the um green ones are good and the
- 2:22:47red ones are
- 2:22:48bad okay so we generated 15 Solutions
- 2:22:52only four of them got the right answer
- 2:22:54and so now what we want to do is
- 2:22:56basically we want to encourage the kinds
- 2:22:58of solutions that lead to right answers
- 2:23:00so whatever token sequences happened in
- 2:23:03these red Solutions obviously something
- 2:23:05went wrong along the way somewhere and
- 2:23:07uh this was not a good path to take
- 2:23:09through the solution and whatever token
- 2:23:11sequences there were in these Green
- 2:23:13Solutions well things went uh pretty
- 2:23:15well in this situation and so we want to
- 2:23:18do more things like it in prompts like
- 2:23:21this and the way we encourage this kind
- 2:23:23of a behavior in the future is we
- 2:23:25basically train on these sequences um
- 2:23:28but these training sequencies now are
- 2:23:29not coming from expert human annotators
- 2:23:32there's no human who decided that this
- 2:23:33is the correct solution this solution
- 2:23:36came from the model itself so the model
- 2:23:38is practicing here it's tried out a few
- 2:23:40Solutions four of them seem to have
- 2:23:41worked and now the model will kind of
- 2:23:43like train on them and this corresponds
- 2:23:45to a student basically looking at their
- 2:23:47Solutions and being like okay well this
- 2:23:48one worked really well so this is this
- 2:23:50is how I should be solving these kinds
- 2:23:52of problems and uh here in this example
- 2:23:55there are many different ways to
- 2:23:57actually like really tweak the
- 2:23:58methodology a little bit here but just
- 2:24:00to give the core idea across maybe it's
- 2:24:02simplest to just think about take the
- 2:24:04taking the single best solution out of
- 2:24:06these four uh like say this one that's
- 2:24:08why it was yellow uh so this is the the
- 2:24:12solution that not only led to the right
- 2:24:13answer but may maybe had some other nice
- 2:24:15properties maybe it was the shortest one
- 2:24:17or it looked nicest in some ways or uh
- 2:24:20there's other criteria you could think
- 2:24:21of as an example but we're going to
- 2:24:23decide that this the top solution we're
- 2:24:25going to train on it and then uh the
- 2:24:28model will be slightly more likely once
- 2:24:30you do the parameter update to take this
- 2:24:33path in this kind of a setting in the
- 2:24:36future but you have to remember that
- 2:24:38we're going to run many different
- 2:24:39diverse prompts across lots of math
- 2:24:42problems and physics problems and
- 2:24:43whatever wherever there might be so tens
- 2:24:46of thousands of prompts maybe have in
- 2:24:47mind there's thousands of solutions
- 2:24:50prompt and so this is all happening kind
- 2:24:52of like at the same time and as we're
- 2:24:55iterating this process the model is
- 2:24:57discovering for itself what kinds of
- 2:24:59token sequences lead it to correct
- 2:25:02answers it's not coming from a human
- 2:25:05annotator the the model is kind of like
- 2:25:08playing in this playground and it knows
- 2:25:10what it's trying to get to and it's
- 2:25:12discovering sequences that work for it
- 2:25:15uh these are sequences that don't make
- 2:25:16any mental leaps uh they they seem to
- 2:25:19work reliably and statistically and uh
- 2:25:23fully utilize the knowledge of the model
- 2:25:25as it has it and so uh this is the
- 2:25:28process of reinforcement
- 2:25:29learning it's basically a guess and
- 2:25:31check we're going to guess many
- 2:25:32different types of solutions we're going
- 2:25:33to check them and we're going to do more
- 2:25:35of what worked in the future and that is
- 2:25:38uh reinforcement learning so in the
- 2:25:40context of what came before we see now
- 2:25:43that the sft model the supervised fine
- 2:25:45tuning model it's still helpful because
- 2:25:47it still kind of like initializes the
- 2:25:49model a little bit into to the vicinity
- 2:25:51of the correct Solutions so it's kind of
- 2:25:53like a initialization of um of the model
- 2:25:56in the sense that it kind of gets the
- 2:25:58model to you know take Solutions like
- 2:26:00write out Solutions and maybe it has an
- 2:26:03understanding of setting up a system of
- 2:26:04equations or maybe it kind of like talks
- 2:26:06through a solution so it gets you into
- 2:26:08the vicinity of correct Solutions but
- 2:26:10reinforcement learning is where
- 2:26:11everything gets dialed in we really
- 2:26:13discover the solutions that work for the
- 2:26:15model get the right answers we encourage
- 2:26:17them and then the model just kind of
- 2:26:19like gets better over time time okay so
- 2:26:21that is the high Lev process for how we
- 2:26:23train large language models in short we
- 2:26:26train them kind of very similar to how
- 2:26:27we train children and basically the only
- 2:26:30difference is that children go through
- 2:26:32chapters of books and they do all these
- 2:26:34different types of training exercises um
- 2:26:37kind of within the chapter of each book
- 2:26:39but instead when we train AIS it's
- 2:26:41almost like we kind of do it stage by
- 2:26:43stage depending on the type of that
- 2:26:45stage so first what we do is we do
- 2:26:47pre-training which as we saw is
- 2:26:49equivalent to uh basically reading all
- 2:26:51the expository material so we look at
- 2:26:53all the textbooks at the same time and
- 2:26:55we read all the exposition and we try to
- 2:26:57build a knowledge base the second thing
- 2:27:00then is we go into the sft stage which
- 2:27:02is really looking at all the fixed uh
- 2:27:04sort of like solutions from Human
- 2:27:07Experts of all the different kinds of
- 2:27:09worked Solutions across all the
- 2:27:11textbooks and we just kind of get an sft
- 2:27:14model which is able to imitate the
- 2:27:16experts but does so kind of blindly it
- 2:27:18just kind of like does its best guess
- 2:27:20uh kind of just like trying to mimic
- 2:27:22statistically the expert behavior and so
- 2:27:24that's what you get when you look at all
- 2:27:26the work Solutions and then finally in
- 2:27:28the last stage we do all the practice
- 2:27:30problems in the RL stage across all the
- 2:27:33textbooks we only do the practice
- 2:27:35problems and that's how we get the RL
- 2:27:37model so on a high level the way we
- 2:27:40train llms is very much equivalent uh to
- 2:27:43the process that we train uh that we use
- 2:27:45for training of children the next point
- 2:27:47I would like to make is that actually
- 2:27:49these first two stat ages pre-training
- 2:27:51and surprise fine-tuning they've been
- 2:27:52around for years and they are very
- 2:27:53standard and everyone does them all the
- 2:27:55different llm providers it is this last
- 2:27:58stage the RL training that is a lot more
- 2:28:00early in its process of development and
- 2:28:02is not standard yet in the field and so
- 2:28:06um this stage is a lot more kind of
- 2:28:09early and nent and the reason for that
- 2:28:11is because I actually skipped over a ton
- 2:28:13of little details here in this process
- 2:28:15the high level idea is very simple it's
- 2:28:17trial and there learning but there's a
- 2:28:18ton of details and little math
- 2:28:20mathematical kind of like nuances to
- 2:28:21exactly how you pick the solutions that
- 2:28:23are the best and how much you train on
- 2:28:25them and what is the prompt distribution
- 2:28:27and how to set up the training run such
- 2:28:29that this actually works so there's a
- 2:28:30lot of little details and knobs to the
- 2:28:32core idea that is very very simple and
- 2:28:35so getting the details right here uh is
- 2:28:37not trivial and so a lot of companies
- 2:28:40like for example open and other LM
- 2:28:41providers have experimented internally
- 2:28:44with reinforcement learning fine tuning
- 2:28:46for llms for a while but they've not
- 2:28:48talked about it publicly
- 2:28:50um it's all kind of done inside the
- 2:28:52company and so that's why the paper from
- 2:28:55Deep seek that came out very very
- 2:28:56recently was such a big deal because
- 2:28:59this is a paper from this company called
- 2:29:01DC Kai in China and this paper really
- 2:29:05talked very publicly about reinforcement
- 2:29:07learning fine training for large
- 2:29:08language models and how incredibly
- 2:29:10important it is for large language
- 2:29:12models and how it brings out a lot of
- 2:29:14reasoning capabilities in the models
- 2:29:16we'll go into this in a second so this
- 2:29:18paper reinvigorated the public interest
- 2:29:21of using RL for llms and gave a lot of
- 2:29:25the um sort of n-r details that are
- 2:29:27needed to reproduce their results and
- 2:29:29actually get the stage to work for large
- 2:29:31langage models so let me take you
- 2:29:33briefly through this uh deep seek R1
- 2:29:35paper and what happens when you actually
- 2:29:36correctly apply RL to language models
- 2:29:38and what that looks like and what that
- 2:29:39gives you so the first thing I'll scroll
- 2:29:41to is this uh kind of figure two here
- 2:29:43where we are looking at the Improvement
- 2:29:45in how the models are solving
- 2:29:47mathematical problems so this is the
- 2:29:49accuracy of solving mathematical
- 2:29:50problems on the a accuracy and then we
- 2:29:54can go to the web page and we can see
- 2:29:55the kinds of problems that are actually
- 2:29:56in these um these the kinds of math
- 2:29:58problems that are being measured here so
- 2:30:00these are simple math problems you can
- 2:30:02um pause the video if you like but these
- 2:30:04are the kinds of problems that basically
- 2:30:06the models are being asked to solve and
- 2:30:08you can see that in the beginning
- 2:30:09they're not doing very well but then as
- 2:30:10you update the model with this many
- 2:30:12thousands of steps their accuracy kind
- 2:30:14of continues to climb so the models are
- 2:30:17improving and they're solving these
- 2:30:18problems with a higher accuracy
- 2:30:20as you do this trial and error on a
- 2:30:22large data set of these kinds of
- 2:30:24problems and the models are discovering
- 2:30:26how to solve math problems but even more
- 2:30:29incredible than the quantitative kind of
- 2:30:32results of solving these problems with a
- 2:30:33higher accuracy is the qualitative means
- 2:30:35by which the model achieves these
- 2:30:37results so when we scroll down uh one of
- 2:30:40the figures here that is kind of
- 2:30:41interesting is that later on in the
- 2:30:43optimization the model seems to be uh
- 2:30:46using average length per response uh
- 2:30:49goes up up so the model seems to be
- 2:30:51using more tokens to get its higher
- 2:30:54accuracy results so it's learning to
- 2:30:56create very very long Solutions why are
- 2:30:59these Solutions very long we can look at
- 2:31:00them qualitatively here so basically
- 2:31:03what they discover is that the model
- 2:31:05solution get very very long partially
- 2:31:07because so here's a question and here's
- 2:31:09kind of the answer from the model what
- 2:31:11the model learns to do um and this is an
- 2:31:13immerging property of new optimization
- 2:31:15it just discovers that this is good for
- 2:31:17problem solving is it starts to do stuff
- 2:31:19like this wait wait wait that's Nota
- 2:31:21moment I can flag here let's reevaluate
- 2:31:23this step by step to identify the
- 2:31:25correct sum can be so what is the model
- 2:31:27doing here right the model is basically
- 2:31:30re-evaluating steps it has learned that
- 2:31:32it works better for accuracy to try out
- 2:31:35lots of ideas try something from
- 2:31:37different perspectives retrace reframe
- 2:31:39backtrack is doing a lot of the things
- 2:31:41that you and I are doing in the process
- 2:31:43of problem solving for mathematical
- 2:31:44questions but it's rediscovering what
- 2:31:46happens in your head not what you put
- 2:31:48down on the solution and there is no
- 2:31:50human who can hardcode this stuff in the
- 2:31:52ideal assistant response this is only
- 2:31:55something that can be discovered in the
- 2:31:56process of reinforcement learning
- 2:31:57because you wouldn't know what to put
- 2:31:59here this just turns out to work for the
- 2:32:02model and it improves its accuracy in
- 2:32:04problem solving so the model learns what
- 2:32:06we call these chains of thought in your
- 2:32:08head and it's an emergent property of
- 2:32:10the optim of the optimization and that's
- 2:32:13what's bloating up the response length
- 2:32:16but that's also what's increasing the
- 2:32:18accuracy of the problem problem solving
- 2:32:20so what's incredible here is basically
- 2:32:22the model is discovering ways to think
- 2:32:24it's learning what I like to call
- 2:32:26cognitive strategies of how you
- 2:32:28manipulate a problem and how you
- 2:32:30approach it from different perspectives
- 2:32:31how you pull in some analogies or do
- 2:32:33different kinds of things like that and
- 2:32:35how you kind of uh try out many
- 2:32:37different things over time uh check a
- 2:32:39result from different perspectives and
- 2:32:40how you kind of uh solve problems but
- 2:32:43here it's kind of discovered by the RL
- 2:32:44so extremely incredible to see this
- 2:32:47emerge in the optimization without
- 2:32:48having to hardcode it anywhere the only
- 2:32:50thing we've given it are the correct
- 2:32:52answers and this comes out from trying
- 2:32:54to just solve them correctly which is
- 2:32:56incredible
- 2:32:58um now let's go back to actually the
- 2:33:00problem that we've been working with and
- 2:33:02let's take a look at what it would look
- 2:33:03like uh for uh for this kind of a model
- 2:33:07what we call reasoning or thinking model
- 2:33:09to solve that problem okay so recall
- 2:33:12that this is the problem we've been
- 2:33:13working with and when I pasted it into
- 2:33:15chat GPT 40 I'm getting this kind of a
- 2:33:17response let's take a look at what
- 2:33:19happens when you give this same query to
- 2:33:22what's called a reasoning or a thinking
- 2:33:23model this is a model that was trained
- 2:33:25with reinforcement learning so this
- 2:33:28model described in this paper DC car1 is
- 2:33:30available on chat. dec.com uh so this is
- 2:33:34kind of like the company uh that
- 2:33:35developed is hosting it you have to make
- 2:33:37sure that the Deep think button is
- 2:33:39turned on to get the R1 model as it's
- 2:33:41called we can paste it here and run
- 2:33:44it and so let's take a look at what
- 2:33:46happens now and what is the output of
- 2:33:48the model okay so here's it says so this
- 2:33:51is previously what we get using
- 2:33:53basically what's an sft approach a
- 2:33:54supervised funing approach this is like
- 2:33:56mimicking an expert solution this is
- 2:33:58what we get from the RL model okay let
- 2:34:01me try to figure this out so Emily buys
- 2:34:03three apples and two oranges each orange
- 2:34:05cost $2 total is 13 I need to find out
- 2:34:07blah blah blah so here you you um as
- 2:34:11you're reading this you can't escape
- 2:34:14thinking that this model is
- 2:34:16thinking um is definitely pursuing the
- 2:34:19solution solution it deres that it must
- 2:34:21cost $3 and then it says wait a second
- 2:34:23let me check my math again to be sure
- 2:34:25and then it tries it from a slightly
- 2:34:26different perspective and then it says
- 2:34:28yep all that checks out I think that's
- 2:34:30the answer I don't see any mistakes let
- 2:34:33me see if there's another way to
- 2:34:34approach the problem maybe setting up an
- 2:34:36equation let's let the cost of one apple
- 2:34:39be $8 then blah blah blah yep same
- 2:34:42answer so definitely each apple is $3
- 2:34:44all right confident that that's correct
- 2:34:47and then what it does once it sort of um
- 2:34:49did the thinking process is it writes up
- 2:34:51the nice solution for the human and so
- 2:34:54this is now considering so this is more
- 2:34:56about the correctness aspect and this is
- 2:34:58more about the presentation aspect where
- 2:35:00it kind of like writes it out nicely and
- 2:35:03uh boxes in the correct answer at the
- 2:35:05bottom and so what's incredible about
- 2:35:07this is we get this like thinking
- 2:35:08process of the model and this is what's
- 2:35:10coming from the reinforcement learning
- 2:35:12process this is what's bloating up the
- 2:35:15length of the token sequences they're
- 2:35:16doing thinking and they're trying
- 2:35:17different ways this is what's giving you
- 2:35:20higher accuracy in problem
- 2:35:22solving and this is where we are seeing
- 2:35:24these aha moments and these different
- 2:35:26strategies and these um ideas for how
- 2:35:29you can make sure that you're getting
- 2:35:31the correct
- 2:35:32answer the last point I wanted to make
- 2:35:34is some people are a little bit nervous
- 2:35:36about putting you know very sensitive
- 2:35:38data into chat.com because this is a
- 2:35:41Chinese company so people don't um
- 2:35:43people are a little bit careful and Cy
- 2:35:45with that a little bit um deep seek R1
- 2:35:48is a model that was released by this
- 2:35:50company so this is an open source model
- 2:35:52or open weights model it is available
- 2:35:54for anyone to download and use you will
- 2:35:56not be able to like run it in its full
- 2:35:59um sort of the full model in full
- 2:36:02Precision you won't run that on a
- 2:36:04MacBook but uh or like a local device
- 2:36:07because this is a fairly large model but
- 2:36:08many companies are hosting the full
- 2:36:10largest model one of those companies
- 2:36:12that I like to use is called
- 2:36:14together. so when you go to together.
- 2:36:17you sign up and you go to playgrounds
- 2:36:19you can can select here in the chat deep
- 2:36:21seek R1 and there's many different kinds
- 2:36:23of other models that you can select here
- 2:36:25these are all state-of-the-art models so
- 2:36:27this is kind of similar to the hugging
- 2:36:28face inference playground that we've
- 2:36:29been playing with so far but together. a
- 2:36:32will usually host all the
- 2:36:33state-of-the-art models so select DT
- 2:36:36car1 um you can try to ignore a lot of
- 2:36:38these I think the default settings will
- 2:36:39often be okay and we can put in this and
- 2:36:43because the model was released by Deep
- 2:36:45seek what you're getting here should be
- 2:36:47basically equivalent to what you're
- 2:36:48getting here now because of the
- 2:36:50randomness in the sampling we're going
- 2:36:51to get something slightly different uh
- 2:36:53but in principle this should be uh
- 2:36:55identical in terms of the power of the
- 2:36:57model and you should be able to see the
- 2:36:58same things quantitatively and
- 2:37:00qualitatively uh but uh this model is
- 2:37:02coming from kind of a an American
- 2:37:04company so that's deep seek and that's
- 2:37:07the what's called a reasoning
- 2:37:09model now when I go back to chat uh let
- 2:37:12me go to chat here okay so the models
- 2:37:14that you're going to see in the drop
- 2:37:15down here some of them like 01 03 mini
- 2:37:18O3 mini High Etc they are talking about
- 2:37:21uses Advanced reasoning now what this is
- 2:37:23referring to uses Advanced reasoning is
- 2:37:26it's referring to the fact that it was
- 2:37:27trained by reinforcement learning with
- 2:37:29techniques very similar to those of deep
- 2:37:31C car1 per public statements of opening
- 2:37:34ey employees uh so these are thinking
- 2:37:37models trained with RL and these models
- 2:37:40like GPT 4 or GPT 4 40 mini that you're
- 2:37:42getting in the free tier you should
- 2:37:43think of them as mostly sft models
- 2:37:45supervised fine tuning models they don't
- 2:37:47actually do this like thinking as as you
- 2:37:49see in the RL models and even though
- 2:37:52there's a little bit of reinforcement
- 2:37:53learning involved with these models and
- 2:37:55I'll go that into that in a second these
- 2:37:56are mostly sft models I think you should
- 2:37:58think about it that way so in the same
- 2:38:00way as what we saw here we can pick one
- 2:38:03of the thinking models like say 03 mini
- 2:38:05high and these models by the way might
- 2:38:07not be available to you unless you pay a
- 2:38:09Chachi PT subscription of either $20 per
- 2:38:11month or $200 per month for some of the
- 2:38:14top models so we can pick a thinking
- 2:38:16model and run now what's going to happen
- 2:38:20here is it's going to say reasoning and
- 2:38:21it's going to start to do stuff like
- 2:38:23this and um what we're seeing here is
- 2:38:26not exactly the stuff we're seeing here
- 2:38:29so even though under the hood the model
- 2:38:31produces these kinds of uh kind of
- 2:38:34chains of thought opening ey chooses to
- 2:38:36not show the exact chains of thought in
- 2:38:38the web interface it shows little
- 2:38:40summaries of that of those chains of
- 2:38:42thought and open kind of does this I
- 2:38:44think partly because uh they are worried
- 2:38:46about what's called the distillation
- 2:38:48risk that is that someone could come in
- 2:38:50and actually try to imitate those
- 2:38:51reasoning traces and recover a lot of
- 2:38:53the reasoning performance by just
- 2:38:55imitating the reasoning uh chains of
- 2:38:57thought and so they kind of hide them
- 2:38:59and they only show little summaries of
- 2:39:00them so you're not getting exactly what
- 2:39:02you would get in deep seek as with
- 2:39:04respect to the reasoning itself and then
- 2:39:07they write up the
- 2:39:08solution so these are kind of like
- 2:39:10equivalent even though we're not seeing
- 2:39:12the full under the hood details now in
- 2:39:14terms of the performance uh these models
- 2:39:17and deep seek models are currently rly
- 2:39:19on par I would say it's kind of hard to
- 2:39:21tell because of the evaluations but if
- 2:39:22you're paying $200 per month to open AI
- 2:39:24some of these models I believe are
- 2:39:25currently they basically still look
- 2:39:27better uh but deep seek R1 for now is
- 2:39:30still a very solid choice for a thinking
- 2:39:33model that would be available to you um
- 2:39:36sort of um either on this website or any
- 2:39:39other website because the model is open
- 2:39:40weights you can just download it so
- 2:39:43that's thinking models so what is the
- 2:39:46summary so far well we've talked about
- 2:39:48reinforcement learning and the fact that
- 2:39:50thinking emerges in the process of the
- 2:39:52optimization on when we basically run RL
- 2:39:55on many math uh and kind of code
- 2:39:57problems that have verifiable Solutions
- 2:39:59so there's like an answer three
- 2:40:01Etc now these thinking models you can
- 2:40:04access in for example deep seek or any
- 2:40:07inference provider like together. a and
- 2:40:09choosing deep seek over there these
- 2:40:12thinking models are also available uh in
- 2:40:14chpt under any of the 01 or O3
- 2:40:17models but these GPT 4 R models Etc
- 2:40:20they're not thinking models you should
- 2:40:21think of them as mostly sft models now
- 2:40:25if you are um if you have a prompt that
- 2:40:27requires Advanced reasoning and so on
- 2:40:29you should probably use some of the
- 2:40:30thinking models or at least try them out
- 2:40:32but empirically for a lot of my use when
- 2:40:35you're asking a simpler question there's
- 2:40:36like a knowledge based question or
- 2:40:37something like that this might be
- 2:40:39Overkill like there's no need to think
- 2:40:4030 seconds about some factual question
- 2:40:42so for that I will uh sometimes default
- 2:40:44to just GPT 40 so empirically about 80
- 2:40:4790% of my use is just gp4
- 2:40:49and when I come across a very difficult
- 2:40:51problem like in math and code Etc I will
- 2:40:53reach for the thinking models but then I
- 2:40:56have to wait a bit longer because
- 2:40:57they're thinking um so you can access
- 2:41:00these on chat on deep seek also I wanted
- 2:41:02to point out that um AI studio.
- 2:41:05go.com even though it looks really busy
- 2:41:08really ugly because Google's just unable
- 2:41:10to do this kind of stuff well it's like
- 2:41:13what is happening but if you choose
- 2:41:15model and you choose here Gemini 2.0
- 2:41:17flash thinking experimental 01 21 if you
- 2:41:20choose that one that's also a a kind of
- 2:41:22early experiment experimental of a
- 2:41:25thinking model by Google so we can go
- 2:41:27here and we can give it the same problem
- 2:41:29and click run and this is also a
- 2:41:31thinking problem a thinking model that
- 2:41:33will also do something
- 2:41:35similar and comes out with the right
- 2:41:37answer here so basically Gemini also
- 2:41:40offers a thinking model anthropic
- 2:41:42currently does not offer a thinking
- 2:41:43model but basically this is kind of like
- 2:41:45the frontier development of these llms I
- 2:41:47think RL is kind of like this new
- 2:41:49exciting stage but getting the details
- 2:41:51right is difficult and that's why all
- 2:41:53these models and thinking models are
- 2:41:55currently experimental as of 2025 very
- 2:41:57early 2025 um but this is kind of like
- 2:42:01the frontier development of pushing the
- 2:42:02performance on these very difficult
- 2:42:03problems using reasoning that is
- 2:42:05emerging in these optimizations one more
- 2:42:07connection that I wanted to bring up is
- 2:42:10that the discovery that reinforcement
- 2:42:12learning is extremely powerful way of
- 2:42:14learning is not new to the field of AI
- 2:42:17and one place what we've already seen
- 2:42:19this demonstrated is in the game of Go
- 2:42:22and famously Deep Mind developed the
- 2:42:24system alphago and you can watch a movie
- 2:42:26about it um where the system is learning
- 2:42:29to play the game of go against top human
- 2:42:32players and um when we go to the paper
- 2:42:36underlying alphago so in this paper when
- 2:42:39we scroll
- 2:42:41down we actually find a really
- 2:42:43interesting
- 2:42:44plot um that I think uh is kind of
- 2:42:47familiar uh to us and we're kind of like
- 2:42:49we discovering in the more open domain
- 2:42:51of arbitrary problem solving instead of
- 2:42:53on the closed specific domain of the
- 2:42:55game of Go but basically what they saw
- 2:42:57and we're going to see this in llms as
- 2:42:59well as this becomes more mature is this
- 2:43:03is the ELO rating of playing game of Go
- 2:43:05and this is leas dull an extremely
- 2:43:07strong human player and here what they
- 2:43:09are comparing is the strength of a model
- 2:43:11learned trained by supervised learning
- 2:43:14and a model trained by reinforcement
- 2:43:15learning so the supervised learning
- 2:43:17model is imitating human expert players
- 2:43:20so if you just get a huge amount of
- 2:43:22games played by expert players in the
- 2:43:23game of Go and you try to imitate them
- 2:43:26you are going to get better but then you
- 2:43:28top out and you never quite get better
- 2:43:31than some of the top top top players of
- 2:43:34in the game of Go like LEL so you're
- 2:43:35never going to reach there because
- 2:43:37you're just imitating human players you
- 2:43:39can't fundamentally go beyond a human
- 2:43:40player if you're just imitating human
- 2:43:42players but in a process of
- 2:43:44reinforcement learning is significantly
- 2:43:46more powerful in reinforcement learning
- 2:43:48for a game of Go it means that the
- 2:43:50system is playing moves that empirically
- 2:43:53and statistically lead to win to winning
- 2:43:56the game and so alphago is a system
- 2:43:59where it kind of plays against it itself
- 2:44:02and it's using reinforcement learning to
- 2:44:03create
- 2:44:04rollouts so it's the exact same diagram
- 2:44:07here but there's no prompt it's just uh
- 2:44:10because there's no prompt it's just a
- 2:44:11fixed game of Go but it's trying out
- 2:44:13lots of solutions it's trying out lots
- 2:44:15of plays and then the games that lead to
- 2:44:18a win instead of a specific answer are
- 2:44:20reinforced they're they're made stronger
- 2:44:24and so um the system is learning
- 2:44:26basically the sequences of actions that
- 2:44:28empirically and statistically lead to
- 2:44:30winning the game and reinforcement
- 2:44:32learning is not going to be constrained
- 2:44:34by human performance and reinforcement
- 2:44:36learning can do significantly better and
- 2:44:38overcome even the top players like Lisa
- 2:44:41Dole and so uh probably they could have
- 2:44:44run this longer and they just chose to
- 2:44:46crop it at some point because this costs
- 2:44:47money but this is very powerful
- 2:44:49demonstration of reinforcement learning
- 2:44:51and we're only starting to kind of see
- 2:44:52hints of this diagram in larger language
- 2:44:55models for reasoning problems so we're
- 2:44:58not going to get too far by just
- 2:44:59imitating experts we need to go beyond
- 2:45:01that set up these like little game
- 2:45:03environments and get let let the system
- 2:45:07discover reasoning traces or like ways
- 2:45:09of solving problems uh that are unique
- 2:45:14and that uh just basically work
- 2:45:16well now on this aspect of uniqueness
- 2:45:19notice that when you're doing
- 2:45:19reinforcement learning nothing prevents
- 2:45:21you from veering off the distribution of
- 2:45:24how humans are playing the game and so
- 2:45:26when we go back to uh this alphao search
- 2:45:29here one of the suggested modifications
- 2:45:31is called move 37 and move 37 in alphao
- 2:45:34is referring to a specific point in time
- 2:45:37where alphago basically played a move
- 2:45:40that uh no human expert would play uh so
- 2:45:43the probability of this move uh to be
- 2:45:45played by a human player was evaluated
- 2:45:47to be about 1 in 10th ,000 so it's a
- 2:45:49very rare move but in retrospect it was
- 2:45:52a brilliant move so alphago in the
- 2:45:54process of reinforcement learning
- 2:45:55discovered kind of like a strategy of
- 2:45:57playing that was unknown to humans and
- 2:46:00but is in retrospect uh brilliant I
- 2:46:02recommend this YouTube video um leis do
- 2:46:04versus alphao move 37 reactions and
- 2:46:06Analysis and this is kind of what it
- 2:46:08looked like when alphao played this
- 2:46:11move
- 2:46:14value that's a very that's a very
- 2:46:16surprising move I thought I thought it
- 2:46:19was I thought it was a
- 2:46:21mistake when I see this move anyway so
- 2:46:24basically people are kind of freaking
- 2:46:25out because it's a it's a move that a
- 2:46:28human would not play that alphago played
- 2:46:31because in its training uh this move
- 2:46:33seemed to be a good idea it just happens
- 2:46:35not to be a kind of thing that a humans
- 2:46:37would would do and so that is again the
- 2:46:39power of reinforcement learning and in
- 2:46:41principle we can actually see the
- 2:46:42equivalence of that if we continue
- 2:46:44scaling this Paradigm in language models
- 2:46:46and what that looks like is kind of
- 2:46:47unknown so so um what does it mean to
- 2:46:50solve problems in such a way that uh
- 2:46:54even humans would not be able to get how
- 2:46:56can you be better at reasoning or
- 2:46:58thinking than humans how can you go
- 2:47:00beyond just uh a thinking human like
- 2:47:03maybe it means discovering analogies
- 2:47:05that humans would not be able to uh
- 2:47:07create or maybe it's like a new thinking
- 2:47:09strategy it's kind of hard to think
- 2:47:10through uh maybe it's a holy new
- 2:47:14language that actually is not even
- 2:47:16English maybe it discovers its own
- 2:47:17language that is a lot better at
- 2:47:19thinking um because the model is
- 2:47:22unconstrained to even like stick with
- 2:47:24English uh so maybe it takes a different
- 2:47:27language to think in or it discovers its
- 2:47:29own language so in principle the
- 2:47:31behavior of the system is a lot less
- 2:47:33defined it is open to do whatever works
- 2:47:37and it is open to also slowly Drift from
- 2:47:40the distribution of its training data
- 2:47:41which is English but all of that can
- 2:47:43only be done if we have a very large
- 2:47:45diverse set of problems in which the
- 2:47:48these strategy can be refined and
- 2:47:49perfected and so that is a lot of the
- 2:47:51frontier LM research that's going on
- 2:47:53right now is trying to kind of create
- 2:47:55those kinds of prompt distributions that
- 2:47:57are large and diverse these are all kind
- 2:47:59of like game environments in which the
- 2:48:00llms can practice their thinking and uh
- 2:48:04it's kind of like writing you know these
- 2:48:06practice problems we have to create
- 2:48:07practice problems for all of domains of
- 2:48:10knowledge and if we have practice
- 2:48:12problems and tons of them the models
- 2:48:14will be able to reinforcement learning
- 2:48:16reinforcement learn on them and kind of
- 2:48:18uh create these kinds of uh diagrams but
- 2:48:21in the domain of open thinking instead
- 2:48:23of a closed domain like game of Go
- 2:48:26there's one more section within
- 2:48:27reinforcement learning that I wanted to
- 2:48:29cover and that is that of learning in
- 2:48:32unverifiable domains so so far all of
- 2:48:35the problems that we've looked at are in
- 2:48:36what's called verifiable domains that is
- 2:48:38any candidate solution we can score very
- 2:48:41easily against a concrete answer so for
- 2:48:44example answer is three and we can very
- 2:48:45easily score these Solutions against the
- 2:48:47answer of three
- 2:48:49either we require the models to like box
- 2:48:51in their answers and then we just check
- 2:48:53for equality of whatever is in the box
- 2:48:55with the answer or you can also use uh
- 2:48:58kind of what's called an llm judge so
- 2:49:00the llm judge looks at a solution and it
- 2:49:03gets the answer and just basically
- 2:49:05scores the solution for whether it's
- 2:49:06consistent with the answer or not and
- 2:49:08llms uh empirically are good enough at
- 2:49:10the current capability that they can do
- 2:49:12this fairly reliably so we can apply
- 2:49:14those kinds of techniques as well in any
- 2:49:16case we have a concrete answer and we're
- 2:49:17just checking Solutions again against it
- 2:49:19and we can do this automatically with no
- 2:49:21kind of humans in the loop the problem
- 2:49:23is that we can't apply the strategy in
- 2:49:25what's called unverifiable domains so
- 2:49:28usually these are for example creative
- 2:49:29writing tasks like write a joke about
- 2:49:31Pelicans or write a poem or summarize a
- 2:49:33paragraph or something like that in
- 2:49:35these kinds of domains it becomes harder
- 2:49:37to score our different solutions to this
- 2:49:39problem so for example writing a joke
- 2:49:41about Pelicans we can generate lots of
- 2:49:43different uh jokes of course that's fine
- 2:49:45for example we can go to chbt and we can
- 2:49:47get it to uh generate a joke about
- 2:49:51Pelicans uh so much stuff in their beaks
- 2:49:53because they don't bellan in
- 2:49:56backpacks what
- 2:49:59okay we can uh we can try something else
- 2:50:02why don't Pelicans ever pay for their
- 2:50:04drinks because they always B it to
- 2:50:06someone else haha okay so these models
- 2:50:10are not obviously not very good at humor
- 2:50:12actually I think it's pretty fascinating
- 2:50:13because I think humor is secretly very
- 2:50:15difficult and the model have the
- 2:50:16capability I think anyway in any case
- 2:50:20you could imagine creating lots of jokes
- 2:50:23the problem that we are facing is how do
- 2:50:24we score them now in principle we could
- 2:50:27of course get a human to look at all
- 2:50:29these jokes just like I did right now
- 2:50:31the problem with that is if you are
- 2:50:32doing reinforcement learning you're
- 2:50:34going to be doing many thousands of
- 2:50:36updates and for each update you want to
- 2:50:38be looking at say thousands of prompts
- 2:50:40and for each prompt you want to be
- 2:50:41potentially looking at looking at
- 2:50:43hundred or thousands of different kinds
- 2:50:44of generations and so there's just like
- 2:50:47way too many of these to look at and so
- 2:50:50um in principle you could have a human
- 2:50:52inspect all of them and score them and
- 2:50:53decide that okay maybe this one is funny
- 2:50:55and uh maybe this one is funny and this
- 2:50:58one is funny and we could train on them
- 2:51:01to get the model to become slightly
- 2:51:02better at jokes um in the context of
- 2:51:05pelicans at least um the problem is that
- 2:51:09it's just like way too much human time
- 2:51:10this is an unscalable strategy we need
- 2:51:12some kind of an automatic strategy for
- 2:51:14doing this and one sort of solution to
- 2:51:16this was proposed in this paper
- 2:51:19uh that introduced what's called
- 2:51:20reinforcement learning from Human
- 2:51:21feedback and so this was a paper from
- 2:51:23open at the time and many of these
- 2:51:25people are now um co-founders in
- 2:51:27anthropic um and this kind of proposed a
- 2:51:30approach for uh basically doing
- 2:51:33reinforcement learning in unverifiable
- 2:51:35domains so let's take a look at how that
- 2:51:36works so this is the cartoon diagram of
- 2:51:39the core ideas involved so as I
- 2:51:41mentioned the native approach is if we
- 2:51:44just set Infinity human time we could
- 2:51:46just run RL in these domains just fine
- 2:51:49so for example we can run RL as usual if
- 2:51:51I have Infinity humans I would I just
- 2:51:53want to do and these are just cartoon
- 2:51:55numbers I want to do 1,000 updates where
- 2:51:57each update will be on 1,000 prompts and
- 2:52:00in for each prompt we're going to have
- 2:52:021,000 roll outs that we're scoring so we
- 2:52:05can run RL with this kind of a setup the
- 2:52:08problem is in the process of doing this
- 2:52:10I will need to run one I will need to
- 2:52:12ask a human to evaluate a joke a total
- 2:52:15of 1 billion times and so that's a lot
- 2:52:18of people looking at really terrible
- 2:52:19jokes so we don't want to do that so
- 2:52:22instead we want to take the arlef
- 2:52:24approach so um in our Rel of approach we
- 2:52:27are kind of like the the core trick is
- 2:52:29that of indirection so we're going to
- 2:52:32involve humans just a little bit and the
- 2:52:35way we cheat is that we basically train
- 2:52:37a whole separate neural network that we
- 2:52:39call a reward model and this neural
- 2:52:41network will kind of like imitate human
- 2:52:44scores so we're going to ask humans to
- 2:52:46score um roll
- 2:52:49we're going to then imitate human scores
- 2:52:51using a neural network and this neural
- 2:52:54network will become a kind of simulator
- 2:52:55of human
- 2:52:56preferences and now that we have a
- 2:52:58neural network simulator we can do RL
- 2:53:01against it so instead of asking a real
- 2:53:03human we're asking a simulated human for
- 2:53:06their score of a joke as an example and
- 2:53:09so once we have a simulator we're often
- 2:53:11racist because we can query it as many
- 2:53:13times as we want to and it's all whole
- 2:53:16automatic process and we can now do
- 2:53:17reinforcement learning with respect to
- 2:53:19the simulator and the simulator as you
- 2:53:20might expect is not going to be a
- 2:53:22perfect human but if it's at least
- 2:53:24statistically similar to human judgment
- 2:53:26then you might expect that this will do
- 2:53:28something and in practice indeed uh it
- 2:53:30does so once we have a simulator we can
- 2:53:32do RL and everything works great so let
- 2:53:35me show you a cartoon diagram a little
- 2:53:36bit of what this process looks like
- 2:53:38although the details are not 100 like
- 2:53:40super important it's just a core idea of
- 2:53:42how this works so here I have a cartoon
- 2:53:44diagram of a hypothetical example of
- 2:53:46what training the reward model would
- 2:53:47look like so we have a prompt like write
- 2:53:50a joke about picans and then here we
- 2:53:52have five separate roll outs so these
- 2:53:54are all five different jokes just like
- 2:53:56this one now the first thing we're going
- 2:53:59to do is we are going to ask a human to
- 2:54:02uh order these jokes from the best to
- 2:54:05worst so this is uh so here this human
- 2:54:08thought that this joke is the best the
- 2:54:10funniest so number one joke this is
- 2:54:14number two joke number three joke four
- 2:54:16and five so this is the worst joke
- 2:54:19we're asking humans to order instead of
- 2:54:20give scores directly because it's a bit
- 2:54:22of an easier task it's easier for a
- 2:54:24human to give an ordering than to give
- 2:54:26precise scores now that is now the
- 2:54:29supervision for the model so the human
- 2:54:31has ordered them and that is kind of
- 2:54:32like their contribution to the training
- 2:54:34process but now separately what we're
- 2:54:36going to do is we're going to ask a
- 2:54:37reward model uh about its scoring of
- 2:54:40these jokes now the reward model is a
- 2:54:42whole separate neural network completely
- 2:54:44separate neural net um and it's also
- 2:54:47probably a transform
- 2:54:49uh but it's not a language model in the
- 2:54:50sense that it generates diverse language
- 2:54:53Etc it's just a scoring model so the
- 2:54:56reward model will take as an input The
- 2:54:59Prompt number one and number two a
- 2:55:02candidate joke so um those are the two
- 2:55:05inputs that go into the reward model so
- 2:55:07here for example the reward model would
- 2:55:08be taken this prompt and this joke now
- 2:55:11the output of a reward model is a single
- 2:55:14number and this number is thought of as
- 2:55:16a score and it can range for example
- 2:55:18from Z to one so zero would be the worst
- 2:55:20score and one would be the best score so
- 2:55:23here are some examples of what a
- 2:55:25hypothetical reward model at some stage
- 2:55:27in the training process would give uh s
- 2:55:29scoring to these jokes so 0.1 is a very
- 2:55:33low score 08 is a really high score and
- 2:55:36so on and so now um we compare the
- 2:55:40scores given by the reward model with uh
- 2:55:43the ordering given by the human and
- 2:55:45there's a precise mathematical way to
- 2:55:47actually calculate this uh basically set
- 2:55:49up a loss function and calculate a kind
- 2:55:51of like a correspondence here and uh
- 2:55:54update a model based on it but I just
- 2:55:55want to give you the intuition which is
- 2:55:57that as an example here for this second
- 2:56:00joke the the human thought that it was
- 2:56:02the funniest and the model kind of
- 2:56:03agreed right 08 is a relatively high
- 2:56:05score but this score should have been
- 2:56:07even higher right so after an update we
- 2:56:10would expect that maybe this score
- 2:56:11should have been will actually grow
- 2:56:13after an update of the network to be
- 2:56:15like say 081 or
- 2:56:16something um for this one here they
- 2:56:19actually are in a massive disagreement
- 2:56:21because the human thought that this was
- 2:56:22number two but here the the score is
- 2:56:24only 0.1 and so this score needs to be
- 2:56:27much higher so after an update on top of
- 2:56:30this um kind of a supervision this might
- 2:56:33grow a lot more like maybe it's 0.15 or
- 2:56:35something like
- 2:56:36that um and then here the human thought
- 2:56:39that this one was the worst joke but
- 2:56:41here the model actually gave it a fairly
- 2:56:43High number so you might expect that
- 2:56:45after the update uh this would come down
- 2:56:47to maybe 3 3.5 or something like that so
- 2:56:50basically we're doing what we did before
- 2:56:51we're slightly nudging the predictions
- 2:56:54from the models using a neural network
- 2:56:57training
- 2:56:58process and we're trying to make the
- 2:57:00reward model scores be consistent with
- 2:57:03human
- 2:57:04ordering and so um as we update the
- 2:57:07reward model on human data it becomes
- 2:57:09better and better simulator of the
- 2:57:11scores and orders uh that humans provide
- 2:57:14and then becomes kind of like the the
- 2:57:17neural the simulator of human
- 2:57:18preferences which we can then do RL
- 2:57:20against but critically we're not asking
- 2:57:23humans one billion times to look at a
- 2:57:24joke we're maybe looking at th000
- 2:57:26prompts and five roll outs each so maybe
- 2:57:285,000 jokes that humans have to look at
- 2:57:30in total and they just give the ordering
- 2:57:33and then we're training the model to be
- 2:57:34consistent with that ordering and I'm
- 2:57:36skipping over the mathematical details
- 2:57:38but I just want you to understand a high
- 2:57:39level idea that uh this reward model is
- 2:57:42do is basically giving us this scour and
- 2:57:45we have a way of training it to be
- 2:57:46consistent with human orderings
- 2:57:48and that's how rhf works okay so that is
- 2:57:51the rough idea we basically train
- 2:57:53simulators of humans and RL with respect
- 2:57:55to those
- 2:57:56simulators now I want to talk about
- 2:57:59first the upside of reinforcement
- 2:58:00learning from Human
- 2:58:03feedback the first thing is that this
- 2:58:05allows us to run reinforcement learning
- 2:58:07which we know is incredibly powerful
- 2:58:09kind of set of techniques and it allows
- 2:58:10us to do it in arbitrary domains and
- 2:58:13including the ones that are unverifiable
- 2:58:15so things like summarization and poem
- 2:58:17writing joke writing or any other
- 2:58:19creative writing really uh in domains
- 2:58:21outside of math and code
- 2:58:23Etc now empirically what we see when we
- 2:58:25actually apply rhf is that this is a way
- 2:58:28to improve the performance of the model
- 2:58:30and uh I have a top answer for why that
- 2:58:33might be but I don't actually know that
- 2:58:35it is like super well established on
- 2:58:38like why this is you can empirically
- 2:58:39observe that when you do rhf correctly
- 2:58:41the models you get are just like a
- 2:58:43little bit better um but as to why is I
- 2:58:45think like not as clear so here's my
- 2:58:47best guess my best guess is that this is
- 2:58:49possibly mostly due to the discriminator
- 2:58:52generator
- 2:58:53Gap what that means is that in many
- 2:58:55cases it is significantly easier to
- 2:58:58discriminate than to generate for humans
- 2:59:01so in particular an example of this is
- 2:59:04um in when we do supervised fine-tuning
- 2:59:07right
- 2:59:09sft we're asking humans to generate the
- 2:59:12ideal assistant response and in many
- 2:59:15cases here um as I've shown it uh the
- 2:59:18ideal response is very simple to write
- 2:59:20but in many cases might not be so for
- 2:59:22example in summarization or poem writing
- 2:59:24or joke writing like how are you as a
- 2:59:26human assist as a human labeler um
- 2:59:29supposed to give the ideal response in
- 2:59:30these cases it requires creative human
- 2:59:32writing to do that and so rhf kind of
- 2:59:35sidesteps this because we get um we get
- 2:59:38to ask people a significantly easier
- 2:59:40question as a data labelers they're not
- 2:59:42asked to write poems directly they're
- 2:59:44just given five poems from the model and
- 2:59:46they're just asked to order them and so
- 2:59:49that's just a much easier task for a
- 2:59:51human labeler to do and so what I think
- 2:59:53this allows you to do basically is it um
- 2:59:57it kind of like allows a lot more higher
- 3:00:00accuracy data because we're not asking
- 3:00:02people to do the generation task which
- 3:00:04can be extremely difficult like we're
- 3:00:06not asking them to do creative writing
- 3:00:07we're just trying to get them to
- 3:00:09distinguish between creative writings
- 3:00:11and uh find the ones that are best and
- 3:00:14that is the signal that humans are
- 3:00:15providing just the ordering and that is
- 3:00:17their input into the system and then the
- 3:00:20system in rhf just discovers the kinds
- 3:00:23of responses that would be graded well
- 3:00:26by humans and so that step of
- 3:00:28indirection allows the models to become
- 3:00:30a bit better so that is the upside of
- 3:00:33our LF it allows us to run RL it
- 3:00:35empirically results in better models and
- 3:00:37it allows uh people to contribute their
- 3:00:40supervision uh even without having to do
- 3:00:42extremely difficult tasks um in the case
- 3:00:45of writing ideal responses unfortunately
- 3:00:47our HF also comes with significant
- 3:00:49downsides and so um the main one is that
- 3:00:54basically we are doing reinforcement
- 3:00:55learning not with respect to humans and
- 3:00:57actual human judgment but with respect
- 3:00:59to a lossy simulation of humans right
- 3:01:01and this lossy simulation could be
- 3:01:03misleading because it's just a it's just
- 3:01:05a simulation right it's just a language
- 3:01:07model that's kind of outputting scores
- 3:01:09and it might not perfectly reflect the
- 3:01:11opinion of an actual human with an
- 3:01:13actual brain in all the possible
- 3:01:15different cases so that's number one
- 3:01:17which is actually something even more
- 3:01:18subtle and devious going on that uh
- 3:01:21really
- 3:01:22dramatically holds back our LF as a
- 3:01:24technique that we can really scale to
- 3:01:27significantly um kind of Smart Systems
- 3:01:31and that is that reinforcement learning
- 3:01:32is extremely good at discovering a way
- 3:01:35to game the model to game the simulation
- 3:01:38so this reward model that we're
- 3:01:40constructing here that gives the course
- 3:01:43these models are Transformers these
- 3:01:46Transformers are massive neurals they
- 3:01:48have billions of parameters and they
- 3:01:50imitate humans but they do so in a kind
- 3:01:52of like a simulation way now the problem
- 3:01:54is that these are massive complicated
- 3:01:56systems right there's a billion
- 3:01:57parameters here that are outputting a
- 3:01:58single
- 3:02:00score it turns out that there are ways
- 3:02:02to gain these models you can find kinds
- 3:02:05of inputs that were not part of their
- 3:02:08training set and these inputs
- 3:02:11inexplicably get very high scores but in
- 3:02:13a fake way so very often what you find
- 3:02:17if you run our lch for very long so for
- 3:02:19example if we do 1,000 updates which is
- 3:02:21like say a lot of updates you might
- 3:02:23expect that your jokes are getting
- 3:02:25better and that you're getting like real
- 3:02:26bangers about Pelicans but that's not
- 3:02:28EXA exactly what happens what happens is
- 3:02:31that uh in the first few hundred steps
- 3:02:34the jokes about Pelicans are probably
- 3:02:35improving a little bit and then they
- 3:02:37actually dramatically fall off the cliff
- 3:02:38and you start to get extremely
- 3:02:40nonsensical results like for example you
- 3:02:42start to get um the top joke about
- 3:02:45Pelicans starts to be the
- 3:02:48and this makes no sense right like when
- 3:02:49you look at it why should this be a top
- 3:02:50joke but when you take the the and you
- 3:02:53plug it into your reward model you'd
- 3:02:55expect score of zero but actually the
- 3:02:57reward model loves this as a joke it
- 3:02:59will tell you that the the the theth is
- 3:03:02a score of 1. Z this is a top joke and
- 3:03:06this makes no sense right but it's
- 3:03:07because these models are just
- 3:03:09simulations of humans and they're
- 3:03:10massive neural lots and you can find
- 3:03:12inputs at the bottom that kind of like
- 3:03:15get into the part of the input space
- 3:03:16that kind of gives you nonsensical
- 3:03:17results these examples are what's called
- 3:03:20adversarial examples and I'm not going
- 3:03:22to go into the topic too much but these
- 3:03:24are adversarial inputs to the model they
- 3:03:26are specific little inputs that kind of
- 3:03:29go between the nooks and crannies of the
- 3:03:30model and give nonsensical results at
- 3:03:32the top now here's what you might
- 3:03:34imagine doing you say okay the the the
- 3:03:36is obviously not score of one um it's
- 3:03:39obviously a low score so let's take the
- 3:03:41the the the the let's add it to the data
- 3:03:43set and give it an ordering that is
- 3:03:45extremely bad like a score of five and
- 3:03:47indeed your model will learn that the D
- 3:03:50should have a very low score and it will
- 3:03:51give it score of zero the problem is
- 3:03:53that there will always be basically
- 3:03:55infinite number of nonsensical
- 3:03:57adversarial examples hiding in the model
- 3:04:00if you iterate this process many times
- 3:04:02and you keep adding nonsensical stuff to
- 3:04:04your reward model and giving it very low
- 3:04:05scores you can you'll never win the game
- 3:04:09uh you can do this many many rounds and
- 3:04:11reinforcement learning if you run it
- 3:04:12long enough will always find a way to
- 3:04:14gain the model it will discover
- 3:04:15adversarial examples it will get get
- 3:04:17really high scores uh with nonsensical
- 3:04:20results and fundamentally this is
- 3:04:23because our scoring function is a giant
- 3:04:26neural nut and RL is extremely good at
- 3:04:28finding just the ways to trick it uh so
- 3:04:33long story short you always run rhf put
- 3:04:36for maybe a few hundred updates the
- 3:04:38model is getting better and then you
- 3:04:39have to crop it and you are done you
- 3:04:42can't run too much against this reward
- 3:04:45model because the optimization will
- 3:04:47start to game it and you basically crop
- 3:04:50it and you call it and you ship it um
- 3:04:53and uh you can improve the reward model
- 3:04:56but you kind of like come across these
- 3:04:57situations eventually at some point so
- 3:05:00rhf basically what I usually say is that
- 3:05:03RF is not RL and what I mean by that is
- 3:05:06I mean RF is RL obviously but it's not
- 3:05:09RL in the magical sense this is not RL
- 3:05:12that you can run
- 3:05:13indefinitely these kinds of problems
- 3:05:16like where you are getting con correct
- 3:05:18answer you cannot gain this as easily
- 3:05:20you either got the correct answer or you
- 3:05:21didn't and the scoring function is much
- 3:05:23much simpler you're just looking at the
- 3:05:25boxed area and seeing if the result is
- 3:05:27correct so it's very difficult to gain
- 3:05:29these functions but uh gaming a reward
- 3:05:32model is possible now in these
- 3:05:34verifiable domains you can run RL
- 3:05:36indefinitely you could run for tens of
- 3:05:38thousands hundreds of thousands of steps
- 3:05:40and discover all kinds of really crazy
- 3:05:41strategies that we might not even ever
- 3:05:43think about of Performing really well
- 3:05:45for all these problems in the game of Go
- 3:05:48there's no way to to beat to basically
- 3:05:50game uh the winning of a game or the
- 3:05:52losing of a game we have a perfect
- 3:05:54simulator we know all the different uh
- 3:05:57where all the stones are placed and we
- 3:05:59can calculate uh whether someone has won
- 3:06:01or not there's no way to gain that and
- 3:06:03so you can do RL indefinitely and you
- 3:06:05can eventually be beat even leol but
- 3:06:08with models like this which are gameable
- 3:06:11you cannot repeat this process
- 3:06:13indefinitely so I kind of see rhf as not
- 3:06:16real RL because the reward function is
- 3:06:19gameable so it's kind of more like in
- 3:06:21the realm of like little fine-tuning
- 3:06:23it's a little it's a little Improvement
- 3:06:26but it's not something that is
- 3:06:27fundamentally set up correctly where you
- 3:06:29can insert more compute run for longer
- 3:06:32and get much better and magical results
- 3:06:34so it's it's uh it's not RL in that
- 3:06:36sense it's not RL in the sense that it
- 3:06:38lacks magic um it can find you in your
- 3:06:41model and get a better performance and
- 3:06:43indeed if we go back to chat GPT the GPT
- 3:06:4640 model has gone through rhf because it
- 3:06:50works well but it's just not RL in the
- 3:06:52same sense rlf is like a little fine
- 3:06:54tune that slightly improves your model
- 3:06:56is maybe like the way I would think
- 3:06:57about it okay so that's most of the
- 3:06:59technical content that I wanted to cover
- 3:07:01I took you through the three major
- 3:07:03stages and paradigms of training these
- 3:07:05models pre-training supervised fine
- 3:07:07tuning and reinforcement learning and I
- 3:07:09showed you that they Loosely correspond
- 3:07:11to the process we already use for
- 3:07:12teaching children and so in particular
- 3:07:15we talked about pre-training being sort
- 3:07:17of like the basic knowledge acquisition
- 3:07:18of reading Exposition supervised fine
- 3:07:21tuning being the process of looking at
- 3:07:22lots and lots of worked examples and
- 3:07:24imitating experts and practice problems
- 3:07:28the only difference is that we now have
- 3:07:30to effectively write textbooks for llms
- 3:07:32and AIS across all the disciplines of
- 3:07:35human knowledge and also in all the
- 3:07:37cases where we actually would like them
- 3:07:39to work like code and math and you know
- 3:07:42basically all the other disciplines so
- 3:07:44we're in the process of writing
- 3:07:45textbooks for them refining all the
- 3:07:47algorithms that I've presented on the
- 3:07:48high level and then of course doing a
- 3:07:50really really good job at the execution
- 3:07:52of training these models at scale and
- 3:07:54efficiently so in particular I didn't go
- 3:07:56into too many details but these are
- 3:07:58extremely large and complicated
- 3:08:00distributed uh sort of
- 3:08:04um jobs that have to run over tens of
- 3:08:07thousands or even hundreds of thousands
- 3:08:08of gpus and the engineering that goes
- 3:08:10into this is really at the stateof the
- 3:08:12art of what's possible with computers at
- 3:08:14that scale so I didn't cover that aspect
- 3:08:17too much
- 3:08:19but um this is very kind of serious and
- 3:08:22they were underlying all these very
- 3:08:24simple algorithms
- 3:08:25ultimately now I also talked about sort
- 3:08:28of like the theory of mind a little bit
- 3:08:30of these models and the thing I want you
- 3:08:31to take away is that these models are
- 3:08:33really good but they're extremely useful
- 3:08:35as tools for your work you shouldn't uh
- 3:08:38sort of trust them fully and I showed
- 3:08:39you some examples of that even though we
- 3:08:41have mitigations for hallucinations the
- 3:08:43models are not perfect and they will
- 3:08:44hallucinate still it's gotten better
- 3:08:46over time and it will continue to get
- 3:08:48better but they can
- 3:08:49hallucinate in other words in in
- 3:08:52addition to that I covered kind of like
- 3:08:53what I call the Swiss cheese uh sort of
- 3:08:56model of llm capabilities that you
- 3:08:57should have in your mind the models are
- 3:08:59incredibly good across so many different
- 3:09:00disciplines but then fail randomly
- 3:09:02almost in some unique cases so for
- 3:09:05example what is bigger 9.11 or 9.9 like
- 3:09:07the model doesn't know but
- 3:09:09simultaneously it can turn around and
- 3:09:11solve Olympiad questions and so this is
- 3:09:14a hole in the Swiss cheese and there are
- 3:09:16many of them and you don't want to trip
- 3:09:17over them so don't um treat these models
- 3:09:21as infallible models check their work
- 3:09:23use them as tools use them for
- 3:09:25inspiration use them for the first draft
- 3:09:28but uh work with them as tools and be
- 3:09:30ultimately respons responsible for the
- 3:09:32you know product of your
- 3:09:35work and that's roughly what I wanted to
- 3:09:38talk about this is how they're trained
- 3:09:40and this is what they are let's now turn
- 3:09:43to what are some of the future
- 3:09:44capabilities of these models uh probably
- 3:09:46what's coming down the pipe and also
- 3:09:48where can you find these models I have a
- 3:09:50few blow points on some of the things
- 3:09:51that you can expect coming down the pipe
- 3:09:53the first thing you'll notice is that
- 3:09:55the models will very rapidly become
- 3:09:56multimodal everything I talked about
- 3:09:58above concerned text but very soon we'll
- 3:10:01have llms that can not just handle text
- 3:10:03but they can also operate natively and
- 3:10:05very easily over audio so they can hear
- 3:10:08and speak and also images so they can
- 3:10:10see and paint and we're already seeing
- 3:10:13the beginnings of all of this uh but
- 3:10:15this will be all done natively inside
- 3:10:17inside the language model and this will
- 3:10:19enable kind of like natural
- 3:10:20conversations and roughly speaking the
- 3:10:22reason that this is actually no
- 3:10:23different from everything we've covered
- 3:10:24above is that as a baseline you can
- 3:10:28tokenize audio and images and apply the
- 3:10:31exact same approaches of everything that
- 3:10:32we've talked about above so it's not a
- 3:10:34fundamental change it's just uh it's
- 3:10:36just a to we have to add some tokens so
- 3:10:38as an example for tokenizing audio we
- 3:10:41can look at slices of the spectrogram of
- 3:10:43the audio signal and we can tokenize
- 3:10:45that and just add more tokens that
- 3:10:47suddenly represent audio and just add
- 3:10:50them into the context windows and train
- 3:10:51on them just like above the same for
- 3:10:53images we can use patches and we can
- 3:10:56separately tokenize patches and then
- 3:10:58what is an image an image is just a
- 3:11:00sequence of tokens and this actually
- 3:11:03kind of works and there's a lot of early
- 3:11:04work in this direction and so we can
- 3:11:06just create streams of tokens that are
- 3:11:08representing audio images as well as
- 3:11:10text and interpers them and handle them
- 3:11:12all simultaneously in a single model so
- 3:11:14that's one example of multimodality
- 3:11:17uh second something that people are very
- 3:11:18interested in
- 3:11:20is currently most of the work is that
- 3:11:22we're handing individual tasks to the
- 3:11:24models on kind of like a silver platter
- 3:11:26like please solve this task for me and
- 3:11:28the model sort of like does this little
- 3:11:29task but it's up to us to still sort of
- 3:11:32like organize a coherent execution of
- 3:11:35tasks to perform jobs and the models are
- 3:11:38not yet at the capability required to do
- 3:11:41this in a coherent error correcting way
- 3:11:43over long periods of time so they're not
- 3:11:46able to fully string together tasks to
- 3:11:48perform these longer running jobs but
- 3:11:51they're getting there and this is
- 3:11:52improving uh over time but uh probably
- 3:11:55what's going to happen here is we're
- 3:11:56going to start to see what's called
- 3:11:57agents which perform tasks over time and
- 3:12:00you you supervise them and you watch
- 3:12:02their work and they come up to once in a
- 3:12:04while report progress and so on so we're
- 3:12:07going to see more long running agents uh
- 3:12:09tasks that don't just take you know a
- 3:12:11few seconds of response but many tens of
- 3:12:13seconds or even minutes or hours over
- 3:12:15time uh but these uh models are not
- 3:12:17infallible as we talked about above so
- 3:12:19all of this will require supervision so
- 3:12:21for example in factories people talk
- 3:12:23about the human to robot ratio uh for
- 3:12:26automation I think we're going to see
- 3:12:27something similar in the digital space
- 3:12:29where we are going to be talking about
- 3:12:31human to agent ratios where humans
- 3:12:33becomes a lot more supervisors of agent
- 3:12:35tasks um in the digital
- 3:12:38domain uh next um I think everything is
- 3:12:41going to become a lot more pervasive and
- 3:12:42invisible so it's kind of like
- 3:12:44integrated into the tools and everywhere
- 3:12:48um and in addition kind of like computer
- 3:12:51using so right now these models aren't
- 3:12:53able to take actions on your behalf but
- 3:12:56I think this is a separate bullet point
- 3:12:58um if you saw chpt launch the operator
- 3:13:02then uh that's one early example of that
- 3:13:04where you can actually hand off control
- 3:13:05to the model to perform you know
- 3:13:07keyboard and mouse actions on your
- 3:13:09behalf so that's also something that
- 3:13:11that I think is very interesting the
- 3:13:13last point I have here is just a general
- 3:13:14comment that there's still a lot of
- 3:13:15research to potentially do in this
- 3:13:16domain main one example of that uh is
- 3:13:19something along the lines of test time
- 3:13:20training so remember that everything
- 3:13:22we've done above and that we talked
- 3:13:24about has two major stages there's first
- 3:13:27the training stage where we tune the
- 3:13:28parameters of the model to perform the
- 3:13:30tasks well once we get the parameters we
- 3:13:33fix them and then we deploy the model
- 3:13:34for inference from there the model is
- 3:13:37fixed it doesn't change anymore it
- 3:13:39doesn't learn from all the stuff that
- 3:13:41it's doing a test time it's a fixed um
- 3:13:43number of parameters and the only thing
- 3:13:45that is changing is now the token inside
- 3:13:47the context windows and so the only type
- 3:13:49of learning or test time learning that
- 3:13:51the model has access to is the in
- 3:13:53context learning of its uh kind of like
- 3:13:56uh dynamically adjustable context window
- 3:13:59depending on like what it's doing at
- 3:14:00test time so but I think this is still
- 3:14:03different from humans who actually are
- 3:14:04able to like actually learn uh depending
- 3:14:06on what they're doing especially when
- 3:14:08you sleep for example like your brain is
- 3:14:09updating your parameters or something
- 3:14:10like that right so there's no kind of
- 3:14:13equivalent of that currently in these
- 3:14:14models and tools so there's a lot of
- 3:14:16like um more wonky ideas I think that
- 3:14:18are to be explored still and uh in
- 3:14:20particular I think this will be
- 3:14:21necessary because the context window is
- 3:14:24a finite and precious resource and
- 3:14:26especially once we start to tackle very
- 3:14:27long running multimodal tasks and we're
- 3:14:30putting in videos and these token
- 3:14:31windows will basically start to grow
- 3:14:34extremely large like not thousands or
- 3:14:36even hundreds of thousands but
- 3:14:37significantly beyond that and the only
- 3:14:39trick uh the only kind of trick we have
- 3:14:41Avail to us right now is to make the
- 3:14:43context Windows longer but I think that
- 3:14:46that approach by itself will will not
- 3:14:47will not scale to actual long running
- 3:14:49tasks that are multimodal over time and
- 3:14:51so I think new ideas are needed in some
- 3:14:53of those disciplines um in some of those
- 3:14:56kind of cases in the main where these
- 3:14:58tasks are going to require very long
- 3:15:00contexts so those are some examples of
- 3:15:03some of the things you can um expect
- 3:15:05coming down the pipe let's now turn to
- 3:15:07where you can actually uh kind of keep
- 3:15:09track of this progress and um you know
- 3:15:12be up to date with the latest and grest
- 3:15:13of what's happening in the field so I
- 3:15:15would say the three resources that I
- 3:15:16have consistently used to stay up to
- 3:15:18date are number one El Marina uh so let
- 3:15:21me show you El
- 3:15:23Marina this is basically an llm leader
- 3:15:26board and it ranks all the top models
- 3:15:30and the ranking is based on human
- 3:15:32comparisons so humans prompt these
- 3:15:34models and they get to judge which one
- 3:15:35gives a better answer they don't know
- 3:15:37which model is which they're just
- 3:15:39looking at which model is the better
- 3:15:40answer and you can calculate a ranking
- 3:15:42and then you get some results and so
- 3:15:44what you can hear is what you can see
- 3:15:46here is the different organizations like
- 3:15:48Google Gemini for example that produce
- 3:15:49these models when you click on any one
- 3:15:51of these it takes you to the place where
- 3:15:53that model is
- 3:15:55hosted and then here we see Google is
- 3:15:57currently on top with open AI right
- 3:15:59behind here we see deep seek in position
- 3:16:02number three now the reason this is a
- 3:16:04big deal is the last column here you see
- 3:16:05license deep seek is an MIT license
- 3:16:08model it's open weights anyone can use
- 3:16:10these weights uh anyone can download
- 3:16:12them anyone can host their own version
- 3:16:14of Deep seek and they can use it in what
- 3:16:16whatever way they like and so it's not a
- 3:16:18proprietary model that you don't have
- 3:16:19access to it's it's basically an open
- 3:16:21weight release and so this is kind of
- 3:16:24unprecedented that a model this strong
- 3:16:27was released with open weights so pretty
- 3:16:29cool from the team next up we have a few
- 3:16:32more models from Google and open Ai and
- 3:16:34then when you continue to scroll down
- 3:16:35you start to see some other Usual
- 3:16:36Suspects so xai here anthropic with son
- 3:16:40it uh here at number
- 3:16:4314 and
- 3:16:45um then
- 3:16:47meta with llama over here so llama
- 3:16:51similar to deep seek is an open weights
- 3:16:52model and so uh but it's down here as
- 3:16:55opposed to up here now I will say that
- 3:16:57this leaderboard was really good for a
- 3:17:00long time I do think that in the last
- 3:17:03few months it's become a little bit
- 3:17:05gamed um and I don't trust it as much as
- 3:17:08I used to I think um just empirically I
- 3:17:11feel like a lot of people for example
- 3:17:13are using a Sonet from anthropic and
- 3:17:15that it's a really good model so but
- 3:17:17that's all the way down here um in
- 3:17:19number 14 and conversely I think not as
- 3:17:22many people are using Gemini but it's
- 3:17:23racking really really high uh so I think
- 3:17:27use this as a first pass uh but uh sort
- 3:17:30of try out a few of the models for your
- 3:17:32tasks and see which one performs better
- 3:17:35the second thing that I would point to
- 3:17:37is the uh AI news uh newsletter so AI
- 3:17:41news is not very creatively named but it
- 3:17:43is a very good newsletter produced by
- 3:17:44swix and friends so thank you for
- 3:17:46maintaining it
- 3:17:47and it's been very helpful to me because
- 3:17:48it is extremely comprehensive so if you
- 3:17:50go to archives uh you see that it's
- 3:17:52produced almost every other day and um
- 3:17:56it is very comprehensive and some of it
- 3:17:58is written by humans and curated by
- 3:17:59humans but a lot of it is constructed
- 3:18:01automatically with llms so you'll see
- 3:18:03that these are very comprehensive and
- 3:18:04you're probably not missing anything
- 3:18:06major if you go through it of course
- 3:18:08you're probably not going to go through
- 3:18:09it because it's so long but I do think
- 3:18:12that these summaries all the way up top
- 3:18:14are quite good and I think have some
- 3:18:15human oversight uh so this has been very
- 3:18:18helpful to me and the last thing I would
- 3:18:20point to is just X and Twitter uh a lot
- 3:18:22of um AI happens on X and so I would
- 3:18:25just follow people who you like and
- 3:18:27trust and get all your latest and
- 3:18:29greatest uh on X as well so those are
- 3:18:32the major places that have worked for me
- 3:18:33over time and finally a few words on
- 3:18:35where you can find the models and where
- 3:18:37can you use them so the first one I
- 3:18:39would say is for any of the biggest
- 3:18:41proprietary models you just have to go
- 3:18:42to the website of that LM provider so
- 3:18:44for example for open a that's uh chat
- 3:18:47I believe actually works now uh so
- 3:18:49that's for open
- 3:18:50AI now for or you know for um for Gemini
- 3:18:54I think it's gem. google.com or AI
- 3:18:57Studio I think they have two for some
- 3:18:59reason that I don't fly understand no
- 3:19:01one does um for the open weights models
- 3:19:04like deep SE CL Etc you have to go to
- 3:19:06some kind of an inference provider of
- 3:19:08LMS so my favorite one is together
- 3:19:10together. a and I showed you that when
- 3:19:11you go to the playground of together. a
- 3:19:14then you can sort of pick lots of
- 3:19:15different models and all of these are
- 3:19:17open models of different types and you
- 3:19:19can talk to them here as an
- 3:19:21example um now if you'd like to use a
- 3:19:24base model like um you know a base model
- 3:19:28then this is where I think it's not as
- 3:19:29common to find base models even on these
- 3:19:31inference providers they are all
- 3:19:32targeting assistants and chat and so I
- 3:19:35think even here I can't I couldn't see
- 3:19:37base models here so for base models I
- 3:19:39usually go to hyperbolic because they
- 3:19:41serve my llama 3.1 base and I love that
- 3:19:45model and you can just talk to it here
- 3:19:47so as far as I know this is this is a
- 3:19:49good place for a base model and I wish
- 3:19:51more people hosted base models because
- 3:19:53they are useful and interesting to work
- 3:19:54with in some cases finally you can also
- 3:19:57take some of the models that are smaller
- 3:19:59and you can run them locally and so for
- 3:20:02example deep seek the biggest model
- 3:20:04you're not going to be able to run
- 3:20:05locally on your MacBook but there are
- 3:20:07smaller versions of the deep seek model
- 3:20:09that are what's called distilled and
- 3:20:11then also you can run these models at
- 3:20:12smaller Precision so not at the native
- 3:20:14Precision of for example fp8 on deep
- 3:20:17seek or you know bf16 llama but much
- 3:20:20much lower than that um and don't worry
- 3:20:23if you don't fully understand those
- 3:20:24details but you can run smaller versions
- 3:20:26that have been distilled and then at
- 3:20:28even lower precision and then you can
- 3:20:29fit them on your uh computer and so you
- 3:20:33can actually run pretty okay models on
- 3:20:35your laptop and my favorite I think
- 3:20:37place I go to usually is LM studio uh
- 3:20:39which is basically an app you can get
- 3:20:42and I think it kind of actually looks
- 3:20:43really ugly and it's I don't like that
- 3:20:45it shows you all these models that are
- 3:20:46basically not that useful like everyone
- 3:20:48just wants to run deep seek so I don't
- 3:20:49know why they give you these 500
- 3:20:51different types of models they're really
- 3:20:53complicated to search for and you have
- 3:20:54to choose different distillations and
- 3:20:56different uh precisions and it's all
- 3:20:58really confusing but once you actually
- 3:21:00understand how it works and that's a
- 3:21:01whole separate video then you can
- 3:21:02actually load up a model like here I
- 3:21:04loaded up a llama 3 uh2 instruct 1
- 3:21:08billion and um you can just talk to it
- 3:21:11so I ask for Pelican jokes and I can ask
- 3:21:14for another one and it gives me another
- 3:21:15one Etc all of this that happens here is
- 3:21:18locally on your computer so we're not
- 3:21:20actually going to anywhere anyone else
- 3:21:22this is running on the GPU on the
- 3:21:24MacBook Pro so that's very nice and you
- 3:21:26can then eject the model when you're
- 3:21:28done and that frees up the ram so LM
- 3:21:31studio is probably like my favorite one
- 3:21:33even though I don't I think it's got a
- 3:21:34lot of uiux issues and it's really
- 3:21:36geared towards uh professionals almost
- 3:21:39uh but if you watch some videos on
- 3:21:40YouTube I think you can figure out how
- 3:21:41to how to use this
- 3:21:43interface uh so those are a few words on
- 3:21:45where to find them so let me now loop
- 3:21:47back around to where we started the
- 3:21:49question was when we go to chashi
- 3:21:50pta.com and we enter some kind of a
- 3:21:53query and we hit go what exactly is
- 3:21:57happening here what are we seeing what
- 3:21:59are we talking to how does this work and
- 3:22:03I hope that this video gave you some
- 3:22:04appreciation for some of the under the
- 3:22:06hood details of how these models are
- 3:22:08trained and what this is that is coming
- 3:22:10back so in particular we now know that
- 3:22:12your query is taken and is first chopped
- 3:22:15up into tokens so we go to to tick
- 3:22:18tokenizer and here where is the place in
- 3:22:21the in the um sort of format that is for
- 3:22:24the user query we basically put in our
- 3:22:27query right there so our query goes into
- 3:22:31what we discussed here is the
- 3:22:32conversation protocol format which is
- 3:22:34this way that we maintain conversation
- 3:22:36objects so this gets inserted there and
- 3:22:39then this whole thing ends up being just
- 3:22:40a token sequence a onedimensional token
- 3:22:43sequence under the hood so Chachi PT saw
- 3:22:46this token sequence and then when we hit
- 3:22:48go it basically continues appending
- 3:22:50tokens into this list it continues the
- 3:22:53sequence it acts like a token
- 3:22:55autocomplete so in particular it gave us
- 3:22:57this response so we can basically just
- 3:23:00put it here and we see the tokens that
- 3:23:02it continued uh these are the tokens
- 3:23:04that it continued with
- 3:23:06roughly now the question
- 3:23:08becomes okay why are these the tokens
- 3:23:10that the model responded with what are
- 3:23:12these tokens where are they coming from
- 3:23:14uh what are we talking to and how do we
- 3:23:17program this system and so that's where
- 3:23:19we shifted gears and we talked about the
- 3:23:21under thehood pieces of it so the first
- 3:23:24stage of this process and there are
- 3:23:25three stages is the pre-training stage
- 3:23:27which fundamentally has to do with just
- 3:23:28knowledge acquisition from the internet
- 3:23:30into the parameters of this neural
- 3:23:32network and so the neural net
- 3:23:35internalizes a lot of Knowledge from the
- 3:23:37internet but where the personality
- 3:23:39really comes in is in the process of
- 3:23:41supervised fine-tuning here and so what
- 3:23:44what happens here is that basically the
- 3:23:46a company like openai will curate a
- 3:23:49large data set of conversations like say
- 3:23:511 million conversation across very
- 3:23:53diverse topics and there will be
- 3:23:55conversations between a human and an
- 3:23:57assistant and even though there's a lot
- 3:23:59of synthetic data generation used
- 3:24:01throughout this entire process and a lot
- 3:24:02of llm help and so on fundamentally this
- 3:24:05is a human data curation task with lots
- 3:24:08of humans involved and in particular
- 3:24:10these humans are data labelers hired by
- 3:24:12open AI who are given labeling
- 3:24:14instructions that they learn and they
- 3:24:16task is to create ideal assistant
- 3:24:18responses for any arbitrary prompts so
- 3:24:21they are teaching the neural network by
- 3:24:24example how to respond to
- 3:24:27prompts so what is the way to think
- 3:24:29about what came back here like what is
- 3:24:32this well I think the right way to think
- 3:24:34about it is that this is the neural
- 3:24:37network simulation of a data labeler at
- 3:24:40openai so it's as if I gave this query
- 3:24:44to a data Li open and this data labeler
- 3:24:47first reads all of the labeling
- 3:24:48instructions from open Ai and then
- 3:24:51spends 2 hours writing up the ideal
- 3:24:53assistant response to this query and uh
- 3:24:57giving it to me now we're not actually
- 3:24:59doing that right because we didn't wait
- 3:25:01two hours so what we're getting here is
- 3:25:02a neural network simulation of that
- 3:25:05process and we have to keep in mind that
- 3:25:08these neural networks don't function
- 3:25:10like human brains do they are different
- 3:25:12what's easy or hard for them is
- 3:25:13different from what's easy or hard for
- 3:25:15humans and so we really are just getting
- 3:25:17a simulation so here I shown you this is
- 3:25:20a token stream and this is fundamentally
- 3:25:23the neural network with a bunch of
- 3:25:24activations and neurons in between this
- 3:25:26is a fixed mathematical expression that
- 3:25:28mixes inputs from tokens with parameters
- 3:25:32of the model and they get mixed up and
- 3:25:35get you the next token in a sequence but
- 3:25:37this is a finite amount of compute that
- 3:25:39happens for every single token and so
- 3:25:41this is some kind of a lossy simulation
- 3:25:44of a human that is kind of like
- 3:25:46restricted in this way and so whatever
- 3:25:49the humans
- 3:25:50write the language model is kind of
- 3:25:52imitating on this token level with only
- 3:25:55this this specific computation for every
- 3:25:58single token and
- 3:26:00sequence we also saw that as a result of
- 3:26:03this and the cognitive differences the
- 3:26:05models will suffer in a variety of ways
- 3:26:08and uh you have to be very careful with
- 3:26:10their use so for example we saw that
- 3:26:11they will suffer from hallucinations and
- 3:26:14they also we have the sense of a Swiss
- 3:26:16model of the LM capabilities where
- 3:26:18basically there's like holes in the
- 3:26:20cheese sometimes the models will just
- 3:26:22arbitrarily like do something dumb uh so
- 3:26:25even though they're doing lots of
- 3:26:26magical stuff sometimes they just can't
- 3:26:28so maybe you're not giving them enough
- 3:26:30tokens to think and maybe they're going
- 3:26:32to just make stuff up because they're
- 3:26:33mental arithmetic breaks uh maybe they
- 3:26:35are suddenly unable to count number of
- 3:26:38letters um or maybe they're unable to
- 3:26:40tell you that 911 9.11 is smaller than
- 3:26:439.9 and it looks kind of dumb and so so
- 3:26:46it's a Swiss cheese capability and we
- 3:26:48have to be careful with that and we saw
- 3:26:49the reasons for
- 3:26:50that but fundamentally this is how we
- 3:26:53think of what came back it's again a
- 3:26:56simulation of this neural network of a
- 3:27:00human data labeler following the
- 3:27:03labeling instructions at open a so
- 3:27:06that's what we're getting back now I do
- 3:27:09think that the uh things change a little
- 3:27:11bit when you actually go and reach for
- 3:27:13one of the thinking models like o03 mini
- 3:27:17and the reason for that is that GPT
- 3:27:2040 basically doesn't do reinforcement
- 3:27:23learning it does do rhf but I've told
- 3:27:26you that rhf is not RL there's no
- 3:27:29there's no uh time for magic in there
- 3:27:31it's just a little bit of a fine-tuning
- 3:27:33is the way to look at it but these
- 3:27:35thinking models they do use RL so they
- 3:27:38go through this third state stage of
- 3:27:41perfecting their thinking process and
- 3:27:44discovering new thinking strategies and
- 3:27:46uh
- 3:27:46solutions to problem solving that look a
- 3:27:49little bit like your internal monologue
- 3:27:51in your head and they practice that on a
- 3:27:53large collection of practice problems
- 3:27:55that companies like openi create and
- 3:27:57curate and um then make available to the
- 3:28:00LMS so when I come here and I talked to
- 3:28:02a thinking model and I put in this
- 3:28:05question what we're seeing here is not
- 3:28:07anymore just the straightforward
- 3:28:09simulation of a human data labeler like
- 3:28:11this is actually kind of new unique and
- 3:28:14interesting um and of course open is not
- 3:28:16showing us the under thehood thinking
- 3:28:18and the chains of thought that are
- 3:28:20underlying the reasoning here but we
- 3:28:23know that such a thing exists and this
- 3:28:24is a summary of it and what we're
- 3:28:26getting here is actually not just an
- 3:28:27imitation of a human data labeler it's
- 3:28:29actually something that is kind of new
- 3:28:30and interesting and exciting in the
- 3:28:32sense that it is a function of thinking
- 3:28:35that was emergent in a simulation it's
- 3:28:37not just imitating human data labeler it
- 3:28:39comes from this reinforcement learning
- 3:28:41process and so here we're of course not
- 3:28:43giving it a chance to shine because this
- 3:28:45is not a mathematical or a reasoning
- 3:28:46problem this is just some kind of a sort
- 3:28:48of creative writing problem roughly
- 3:28:50speaking and I think it's um it's a a
- 3:28:54question an open question as to whether
- 3:28:57the thinking strategies that are
- 3:28:59developed inside verifiable domains
- 3:29:02transfer and are generalizable to other
- 3:29:05domains that are unverifiable such as
- 3:29:07create writing the extent to which that
- 3:29:09transfer happens is unknown in the field
- 3:29:12I would say so we're not sure if we are
- 3:29:14able to do RL on everything that is very
- 3:29:16verifiable and see the benefits of that
- 3:29:18on things that are unverifiable like
- 3:29:20this prompt so that's an open question
- 3:29:22the other thing that's interesting is
- 3:29:23that this reinforcement learning here is
- 3:29:26still like way too new primordial and
- 3:29:29nent so we're just seeing like the
- 3:29:31beginnings of the hints of greatness uh
- 3:29:34in the reasoning problems we're seeing
- 3:29:36something that is in principle capable
- 3:29:38of something like the equivalent of move
- 3:29:4037 but not in the game of Go but in open
- 3:29:44domain thinking and problem solving in
- 3:29:46principle this Paradigm is capable of
- 3:29:48doing something really cool new and
- 3:29:50exciting something even that no human
- 3:29:52has thought of before in principle these
- 3:29:54models are capable of analogies no human
- 3:29:56has had so I think it's incredibly
- 3:29:58exciting that these models exist but
- 3:30:00again it's very early and these are
- 3:30:02primordial models for now um and they
- 3:30:05will mostly shine in domains that are
- 3:30:06verifiable like math en code Etc so very
- 3:30:10interesting to play with and think about
- 3:30:11and
- 3:30:12use and then that's roughly it um um I
- 3:30:16would say those are the broad Strokes of
- 3:30:18what's available right now I will say
- 3:30:20that overall it is an extremely exciting
- 3:30:23time to be in the
- 3:30:24field personally I use these models all
- 3:30:26the time daily uh tens or hundreds of
- 3:30:28times because they dramatically
- 3:30:30accelerate my work I think a lot of
- 3:30:31people see the same thing I think we're
- 3:30:33going to see a huge amount of wealth
- 3:30:34creation as a result of these models be
- 3:30:37aware of some of their shortcomings even
- 3:30:40with RL models they're going to suffer
- 3:30:42from some of these use it as a tool in a
- 3:30:44toolbox don't trust it fully because
- 3:30:47they will randomly do dumb things they
- 3:30:49will randomly hallucinate they will
- 3:30:51randomly skip over some mental
- 3:30:52arithmetic and not get it right um they
- 3:30:55randomly can't count or something like
- 3:30:56that so use them as tools in the toolbox
- 3:30:58check their work and own the product of
- 3:31:00your work but use them for inspiration
- 3:31:03for first draft uh ask them questions
- 3:31:06but always check and verify and you will
- 3:31:08be very successful in your work if you
- 3:31:10do so uh so I hope this video was useful
- 3:31:13and interesting to you I hope you had it
- 3:31:15fun and uh it's already like very long
- 3:31:17so I apologize for that but I hope it
- 3:31:19was useful and yeah I will see you later
About this transcript
This page contains the full transcript of Deep Dive into LLMs like ChatGPT by Andrej Karpathy, generated from the public captions YouTube serves with the video. The transcript has 41,116 words across 5,761 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.