YouTube transcript (35X6zlhoCy4) — Transcript
Full transcript
- 0:05good evening people um even how are you
- 0:08guys
- 0:10doing all right my name is archa Sharma
- 0:13I'm a PhD student at Stanford and I'm
- 0:15very very excited to talk about post
- 0:17training generally speaking for large
- 0:19language models and I hope you guys are
- 0:21ready to learn some stuff because this
- 0:24has been one of the last few years in
- 0:25machine learning have been very very
- 0:26exciting uh with the Advent of large
- 0:29language model CH GPD and everything to
- 0:31that extent and hopefully after today's
- 0:33lecture you will be more comfortable
- 0:36understanding how we go from pre-train
- 0:38Models to models like CH GPD and we'll
- 0:40take a whole journey through prompting
- 0:43instruction fine tuning and DP and
- 0:45rlf so let's get
- 0:51started all
- 0:53right so something that has been very
- 0:57fundamental to our entire field is this
- 1:01idea of scaling loss and models are
- 1:04increasingly becoming larger and larger
- 1:06and they're expanding more and more
- 1:08compute so this is a graph of models
- 1:10starting all the way back in 1950s to
- 1:13somewhere around these are still this is
- 1:15an outdated graph so like this shows up
- 1:17to 10 to^ 24 flops or floating Point
- 1:19operations that go into pre-training
- 1:21these models but the number is well
- 1:23above 10 to^ 26 now but you can see the
- 1:26graph and the way it's
- 1:28trending and more more and more compute
- 1:30requires more and more data because you
- 1:32need to train on something meaningful
- 1:34and this is roughly the trend on the
- 1:35amount of language tokens that are going
- 1:37into the language models in pre-training
- 1:40and again this plot is outdated does
- 1:43anybody want to guess like we're in 2024
- 1:452022 we were at 1.4 trillion tokens or
- 1:49words roughly speaking in language model
- 1:51pre-training do does anyone want to
- 1:53guess like where we are in 2024
- 2:00that's a pretty good guess yeah so we're
- 2:02close to 15 trillion tokens um recent
- 2:05llama 3 models were roughly trained on
- 2:0615 trillion tokens so yeah just just for
- 2:10a second appreciate that these are a lot
- 2:12of words uh this is not yeah I don't I
- 2:16don't think anybody of us listens to
- 2:18like trillions of tokens in our lifetime
- 2:20so this is where we are right now and I
- 2:24hope you guys were here for the pre-
- 2:25pre-training lectures cool um so what do
- 2:30we do so like I mean broadly speaking we
- 2:32are really just learning to predict text
- 2:34tokens or language tokens but what do we
- 2:37learn in the process of pre-training why
- 2:39is why are people spending so much money
- 2:42so much compute because these Compu and
- 2:44tokens take dollars to do and we're
- 2:46we're on the order spending hundreds of
- 2:48millions of dollars on these runs so why
- 2:50are we doing this and this is basically
- 2:53a recall from whatever you have probably
- 2:55learned till now but we're learning
- 2:57things like oh we are learning knowledge
- 2:59Stanford University is located in Santa
- 3:02clar California or wherever you want to
- 3:04like say you're learning syntax you're
- 3:06learning semantics of the sentences
- 3:08these are things that you would expect
- 3:10to learn when you're training on
- 3:12language data broadly you're probably
- 3:14learning a lot about different languages
- 3:16as well so like depending on your text
- 3:17Data distribution you're learning a lot
- 3:19of things but the models we interact
- 3:22with are very intelligent so where is
- 3:24that coming from like I mean just simply
- 3:26learning about very factual things and
- 3:30it's a very simple loss function we're
- 3:32optimizing and where is that
- 3:33Intelligence coming
- 3:35from and this perhaps is the interesting
- 3:39bit recently like people have like
- 3:43started accumulating evidence for that
- 3:45like when you optimize the next token
- 3:47prediction losses you're not just
- 3:49learning about syntax you're not just
- 3:50learning knowledge but you're starting
- 3:52to like form models of Agents beliefs
- 3:55and actions as well so how do we know
- 3:58this again a lot of this is speculative
- 4:00evidence but it's mayy to like form an
- 4:02understanding that the losses we're
- 4:03optimizing are not just about the data
- 4:04fitting the data but you start learning
- 4:06something maybe more meaningful as
- 4:08well um for example like I mean in this
- 4:11specific case um we change the last
- 4:15sentence and the prediction of the text
- 4:17or the next Tex that that is predicted
- 4:19changes as well so here it starts with
- 4:22Pat watch as a demonstration of a
- 4:23bowling ball and the leaf being dropped
- 4:25at the same time pat who is a physicist
- 4:27predicts that the bowling ball and the
- 4:29leaf will land at the same rate we all
- 4:31know Gravity the way it works but when
- 4:34you L change the last sentence to Pat
- 4:36who has never seen this demonstration
- 4:38before Pat predicts that bowling ball
- 4:40will fall to the ground first maybe
- 4:42somebody who's never seen this
- 4:43experiment before might intuitively
- 4:44believe that correct so like I mean the
- 4:47language model was able to predict this
- 4:48so how do you predict this you have to
- 4:51have some notion of understanding of how
- 4:54humans work to even be able to predict
- 4:56this and that's maybe like something
- 4:58that is not obvious with You're simply
- 5:00optimizing to predict the
- 5:03text similarly like I mean these kind of
- 5:05examples are we're going to run through
- 5:06some examples to like sort of
- 5:08communicate that when you're
- 5:09pre-training these models you're
- 5:10learning much more than just language
- 5:11tokens and so on you're also learning
- 5:13about math like you're able to
- 5:16understand what a graph of a circle
- 5:17means and what the center is and where
- 5:19how to like understand
- 5:22equations probably my favorite example
- 5:24something I use pretty much every day is
- 5:26you're learning how to write code so I
- 5:29don't know how many of you have
- 5:30interacted with co-pilot before but if
- 5:33you have like you probably know like if
- 5:34you write down a few commands write down
- 5:36a function template it will
- 5:38automatically complete code for you so
- 5:41again it's not perfect but it has to
- 5:43have some deeper understanding of what
- 5:45your intent is for something like that
- 5:47to
- 5:48emerge and similarly we have examples
- 5:50from medicine as well I don't know about
- 5:52you guys but like whenever I have some
- 5:54issue I probably go to chat gbd or
- 5:55Claude or something to that effect and
- 5:57ask them a diagnosis for those things as
- 5:59well
- 6:00um I don't recommend that uh please
- 6:03don't take medical advice from me but
- 6:06yeah so broadly like the way we're
- 6:09seeing language models at this point is
- 6:11that like they're sort of emerging as
- 6:12this general purpose multitask
- 6:14assistance and it's very strange right
- 6:17like I mean we started off with text
- 6:18token prediction and we're reaching the
- 6:19stage where it can like sort of rely to
- 6:21them on them to do many many different
- 6:23things so how are we getting there and
- 6:25I'm sure you all are aware of like what
- 6:26these models are so yeah
- 6:30so today's lecture is largely going to
- 6:32be about how do we go from something
- 6:34Stanford University is located this very
- 6:37simple pretraining task a very simple
- 6:38procedure well it's more complicated but
- 6:40in abstract terms it's not very
- 6:42complicated to like something as
- 6:44powerful as CH
- 6:45GPD cool
- 6:48so um I recommend you guys stopping me
- 6:50asking me a lot of questions because
- 6:52this is a there's a lot of fun examples
- 6:54and a lot of fun techniques so like I I
- 6:56want you guys to like learn everything
- 6:58about here so the overall plan is we're
- 7:00going to talk about zero shot and few
- 7:02shot in context learning um next we're
- 7:05going to follow up with instruction
- 7:06fine-tuning and then we're going to talk
- 7:08about optimizing for preferences and
- 7:10this is where roughly things are right
- 7:12now in the industry and when we're going
- 7:15to talk about what's next what the
- 7:16limitations are and how do we move from
- 7:20here cool so we're going to start off
- 7:23with zero shot INF fusure in context
- 7:27learning um broad we're going to take an
- 7:29example example of GPT or the generative
- 7:31pre-train Transformer and this is a
- 7:33whole series of models that started off
- 7:34in roughly 2018 and like up to 2020 they
- 7:37were building GPD gpd2 gbd3 so we're
- 7:40going to start off with this example and
- 7:42yes so it's a decoder only model that is
- 7:45trained on roughly 4.6 GB of text and it
- 7:49has 12 layers of Transformers layers and
- 7:51it's trained with the next token
- 7:52prediction
- 7:53loss and the first model obviously was
- 7:57not extremely good but it started
- 7:58showing that like hey like this
- 8:00technique for pre-training can be very
- 8:03effective for general purpose tasks and
- 8:05we're going to see some
- 8:06examples um for example like I mean here
- 8:09it's able to do the task for entainment
- 8:12and okay
- 8:16um yeah and gbd1 itself was not very
- 8:21strong as a model so like but they took
- 8:23the same recipe and like I mean tried to
- 8:25like increase the model size so they
- 8:27went from 117 million parameters to
- 8:29about 1.5 billion parameters and we're
- 8:32now scaling up the data alongside as
- 8:34well so we went from 4 gabt of data to
- 8:36approximately 40 gab of data and
- 8:38pre-training is a whole different like
- 8:40melting part of techniques and there's a
- 8:42lot that goes into it but like roughly
- 8:43for example here they filter data by the
- 8:46number of upwards on the redit
- 8:48data and yeah so this is roughly where
- 8:52we are and I think one of the things
- 8:54that started emerging with gpd2 is zero
- 8:57shot learning and what do we mean by
- 9:00zero shot learning
- 9:02um conventionally in the field like when
- 9:05we pre-train models there was the idea
- 9:06that you take a few examples you update
- 9:08the model um and then you are able to
- 9:11adapt to a specific task but as you
- 9:14pre-train on more and more data and more
- 9:15and more tasks you sort of start seeing
- 9:17this phenomena where they're able to do
- 9:19the task basically zero short they're
- 9:21shown no examples of how to do the task
- 9:23and you can start thinking of oh how you
- 9:25can do it summarization you can follow
- 9:27some instructions you can do maybe a
- 9:29little bit of math as well so this is
- 9:31where the idea of zero shot learning
- 9:33started to
- 9:38emerge yeah so how do we do zero shot
- 9:41learning or task specific learning from
- 9:42these pre-trained models really the idea
- 9:45is that we have to be creative here we
- 9:47know that these are text prediction
- 9:48models if you put in a text they will
- 9:50complete whatever follows so if we can
- 9:52sort of course these models into
- 9:54completing the task we care about maybe
- 9:56it's question answering we can so start
- 9:59getting them to solve tasks here so for
- 10:02example if you want to ask questions
- 10:04about Tom Brady you sort of set it up
- 10:07you sort put information about Tom Brady
- 10:09and then you put a question that you
- 10:10wanted to answer and then it will
- 10:12autocomplete in some sense so this is
- 10:14one early perspective on these models
- 10:16these are very Advanced autocomplete
- 10:19models and
- 10:21similarly if you want to figure out like
- 10:23which answer is true or which is not
- 10:25something that is very useful to measure
- 10:27is log probabilities so
- 10:29for example we want to figure out what
- 10:32is the word it refering to here in this
- 10:35sentence the cat couldn't fit into the
- 10:36Hat because it was too big um what we
- 10:39can do is we can take the sentence
- 10:41replace it with either the cat or either
- 10:44the hat and then you can measure the
- 10:46probability of which Mo which one does
- 10:49the model think is higher and you can
- 10:51sort of get the idea what the reference
- 10:53here is so none of this is like in the
- 10:56training data it's simply learning to
- 10:58predict text but you can start seeing
- 11:00how like we can leverage these models to
- 11:02do other tasks as well be besides
- 11:06prediction so this is just more evidence
- 11:09about like how gpd2 no task specific
- 11:12fine-tuning no task specific training it
- 11:15simply is learning to predict text and
- 11:17it's establishes the state-of-the-art on
- 11:19many many different tasks simply by
- 11:22scaling up the model parameters and
- 11:23scaling up the amount of data it's
- 11:25stained on
- 11:29so this is a fun example so if you want
- 11:31to do summarization for data or like you
- 11:35have a news article that you want to
- 11:37summarize so how do you get a zero shot
- 11:40model to do it this answer is you put
- 11:42the document into the context and you
- 11:44simply put tldr in front of
- 11:47it now like I mean if most of the data
- 11:50on internet whenever you see tldr you'll
- 11:51naturally summarize it so yeah you can
- 11:54get zero shot performance and
- 11:56summarization here as well and again
- 11:57this is not trained to do something
- 11:59summarization in any specific way and
- 12:01it's still doing really well simply
- 12:03because of its pre-training
- 12:07data so yeah um I think gp22 tldr is
- 12:11somewhere there and like some of the
- 12:12very Tas specific train models are um
- 12:15like and I think you will see the trend
- 12:18with again if you were Alec Radford or
- 12:20somebody like I mean you see like these
- 12:22cool things emerging your next step
- 12:24would obviously be I'm going to scale
- 12:26this up a little more I'm going to make
- 12:27an even bigger model I'm going to train
- 12:28it even more data and we'll see how
- 12:31things go right so that's how we got
- 12:34gbd3 uh we went from 1.5 billion
- 12:36parameters to 175 billion parameters we
- 12:38are well over like 40 GB of data to 600
- 12:42gbt of data of course like now we're in
- 12:44like terabytes of data and text is a
- 12:47very compressed representation so like
- 12:48terabytes of data is a
- 12:50lot um and you know we we talked about
- 12:53zero shot learning the cool thing that
- 12:56emerged in gbd3 is like go ahead like
- 13:00used before the passage right no you
- 13:03typically put the passage uh if youve
- 13:04like interacted with Reddit or something
- 13:06like that typically somebody will write
- 13:08an entire post and then end with TLD drr
- 13:11here's a summary of the thing too long
- 13:13didn't read or if you have
- 13:15used opposite comes first oh yeah there
- 13:20are situations where it also comes first
- 13:21but um one reason is that these are like
- 13:24decoder only models so like they are
- 13:27often these are causal attention models
- 13:28so the typically need to see the context
- 13:30before yeah understand I'm just curious
- 13:33like from my experience the comes first
- 13:36then how is
- 13:38it
- 13:40the okay um there's probably a lot of
- 13:43data where the tldr comes first but
- 13:44there's probably a lot of data where
- 13:45tldr comes after as
- 13:47well cool so we saw Zero shot learning
- 13:51emerging in gpd2 few shot learning maybe
- 13:54seem slightly easier but like this is
- 13:56where things started getting really
- 13:57funny is that like you're starting to
- 13:59beat state-ofthe-art simply by just
- 14:00putting examples in context so yeah what
- 14:04does f shot learning mean here what is
- 14:06what are we talking about um as I
- 14:09mentioned like the typical idea here is
- 14:13is that like you want to solve
- 14:14translation so you would put some
- 14:16examples of translation into
- 14:18context and you know this is a
- 14:21correction task here or maybe you want
- 14:22interested in Translation and no
- 14:25gradient updates no learning in any
- 14:28conventional sense whatso ever you put a
- 14:30few examples in and that's it like I
- 14:32mean you know how to solve the task
- 14:34isn't that like crazy like you're bu you
- 14:37guys did the assignment on translation
- 14:39right so but this is what the modern NLP
- 14:42looks like
- 14:44so yeah um you put in some examples and
- 14:48you have the entire system and this is
- 14:50where things got really interesting is
- 14:52that all these task specific models that
- 14:54were created to like be really really
- 14:56good at translation or really good at
- 14:57summarization you can just put let's
- 15:00look at this graph so we start with a
- 15:02zero shot performance of this in a
- 15:04similar fashion that I described earlier
- 15:05and you start somewhere there you put
- 15:07one example in of translation from
- 15:09English to French you get to somewhere
- 15:11like already at a fine level few
- 15:13examples in you're already starting to
- 15:15like be close to the state-ofthe-art
- 15:18models wait but in that gra the state of
- 15:20theart is really high isn't
- 15:22it uh find your Inver of the bird Plus+
- 15:26here I think is like the one I'm
- 15:27referring to find your of the like which
- 15:29is trained exclusively on a lot of um
- 15:32translation data might be like slightly
- 15:33better yes um and I think that's the
- 15:37relevant comparison here is the in
- 15:39context learning starts to emerge at
- 15:42scale so and this is I think like the
- 15:45key point is that this some of this is
- 15:48contested just to be very upfront but
- 15:50like there's this idea of emergence of
- 15:52this property as you train on more
- 15:54Computing and more scale um there's more
- 15:56recent research which suggest that if we
- 15:58plot the access correctly it feels less
- 16:00emergent but the general idea is as you
- 16:02increase the number of parameters and
- 16:05increase the number of compute that is
- 16:06going into the models the ability to
- 16:08just go from a few examples to really
- 16:10strong performance is very
- 16:14compelling cool
- 16:17um and yeah I think as I explained
- 16:20earlier the general idea is that this is
- 16:22very different from the conventional
- 16:23idea of fine tuning that we typically go
- 16:25for instead of like iterating over
- 16:27examples putting it into context and
- 16:29doing gradient updates we are actually
- 16:31just going for few short promting we're
- 16:33going to put in few examples and that's
- 16:34going to give us the
- 16:42system um yes I mean the exact details
- 16:45roughly can depend on the prom template
- 16:47that you use but typically you would
- 16:49just put examples so like c order and
- 16:52put these examples and then whatever
- 16:54your task is you can just let the model
- 16:56complete from there because it can infer
- 16:58the task
- 16:59um based on the examples you've
- 17:02given any other
- 17:05questions
- 17:08cool
- 17:10so yeah like I mean we have gotten from
- 17:12zero shot prompting and like we've seen
- 17:14seeing that F shot prompting is becoming
- 17:16really competitive with good models but
- 17:18there's still limitations to this like I
- 17:20mean you cannot solve every task that
- 17:21you see here and particularly like
- 17:23things that involve like richer
- 17:25multi-step reasoning is something that
- 17:26actually can be pretty challenging and
- 17:29just to be fair human struggle at these
- 17:30tasks as well so things like addition
- 17:33and so on like these are these are
- 17:36probably like still still hard to do
- 17:38like when you keep increasing the number
- 17:39of digits but one thing that you have to
- 17:43start being creative with I alluded to
- 17:45this earlier is that you can get these
- 17:46models to do the task if you're creative
- 17:49in how you prompt the model and this is
- 17:52what we're going to see
- 17:53next um so this technique called Chain
- 17:57of Thought prompting emerged here and
- 17:59the idea that we have explored thus far
- 18:01is that we put in examples of the kind
- 18:03of tasks we want to do and we expect the
- 18:07model to zero shot learn what the task
- 18:09is and go from there um the idea is that
- 18:13like instead of just showing what the
- 18:15task is you show them examples where
- 18:17they reason through the task so they're
- 18:19not just learning to do the task but
- 18:21also learning how the reasoning is
- 18:22working so in this example initially we
- 18:24started with like we have to solve a
- 18:25simple math problem and we are just
- 18:27shown exactly the answer answer directly
- 18:30instead of doing that and if you do that
- 18:32directly you'll observe that the model
- 18:33gets the answer wrong instead of that
- 18:36what if you show model how to reason
- 18:37about the task show it a chain of
- 18:39thought and include that in the prompt
- 18:41as
- 18:42well and then you ask at a new question
- 18:46the idea is that now the model is not
- 18:47just going to Output an answer it's
- 18:50going to reason about the task and it's
- 18:52going to do actually a lot better and
- 18:54this has been shown to be very effective
- 18:57um Chain of Thought is also as you can
- 19:00see like I mean it's again something
- 19:02that improves a lot with model scale
- 19:05it's not just um yeah um but what you
- 19:09can probably start seeing is like it's
- 19:11nearly better than supervised best
- 19:13models here so power models roughly were
- 19:17about 5 40 billion
- 19:18parameters and simply with this Chain of
- 19:21Thought kind of a skill you're already
- 19:22like beating state of the
- 19:25art cool um
- 19:29so yeah so I showed you examples of
- 19:32Chain of Thought reasoning to to this
- 19:34point where you go through a reasoning
- 19:35chain but you can be even slightly
- 19:37smarter than that you might not even
- 19:39need to show them any examples you just
- 19:41need to trick them into thinking about
- 19:43what to do
- 19:44next um so yeah this s this idea emerged
- 19:49in this paper called where you let's
- 19:51think step by step where instead of even
- 19:54showing an example you just start your
- 19:56answer with let's think step by step and
- 20:01that's it like I mean the model will
- 20:03start reasoning about the answer itself
- 20:05instead of just like autoc completing to
- 20:07an answer and you get something like
- 20:11this
- 20:12so maybe you don't even need to show any
- 20:15examples like you can probably induce
- 20:17the reasoning Behavior through zero shot
- 20:18Behavior as well and again um what the
- 20:22final numbers look like is like compared
- 20:24to zero shot performance that we got
- 20:26from essentially autoc comp completing
- 20:29this zero shot Chain of Thought
- 20:31substantially improves the performance
- 20:32so you go from like 17.7 to 78.7 it's
- 20:35still worse than still putting like
- 20:37examples of reasoning and multi-shot few
- 20:40shot Chain of Thought as well but you
- 20:42can see like how much it improves the
- 20:44performance simply by asking you to
- 20:45let's think by step by step and maybe
- 20:48this is like a lesson that interacting
- 20:50with these models is like when you
- 20:52interact with these models you might not
- 20:54get the exact desired behavior from
- 20:57these models up front but often like
- 20:59these models are capable of doing the
- 21:01behavior that you might want and often
- 21:05you have to think about how to induce
- 21:06that behavior such that and the right
- 21:09way to think perhaps is like what is the
- 21:10pre-training data what is the data on
- 21:12the internet it might have seen which
- 21:13induces a similar Behavior to the kind I
- 21:15want and you probably want to like think
- 21:18about that and then induce these kinds
- 21:20of behaviors from those
- 21:24models and yeah like I mean um you know
- 21:28we we hand designed some of these
- 21:30prompts you can also like get an llm to
- 21:32design these prompts as well there's
- 21:34like recursive self-improving ideas here
- 21:37um that happen and you can bump up the
- 21:38performance a little bit
- 21:42more cool so what we have seen so far is
- 21:46that as models get stronger and stronger
- 21:49you can start forcing them to do your
- 21:50task zero shot or with few short
- 21:53examples and you can trick them into
- 21:55thinking what task you want them to
- 21:56solve
- 21:59but the downside is that there's only so
- 22:01much you can fit into context this might
- 22:04not be very true anymore models like
- 22:06becoming increasingly larger context but
- 22:09it's still somewhat unsatisfactory to
- 22:11think you have to trick the model into
- 22:13doing your task rather than like it just
- 22:15doing the task you wanted to do and
- 22:18potentially like I mean going forward
- 22:20like you probably still want to fine
- 22:22tune these models for more and more
- 22:23complex tasks and that's where we're
- 22:26going to go forward in this
- 22:28um next section we're going to cover is
- 22:30instruction fine tuning and the general
- 22:34idea we have right now is that as we
- 22:37talked about pre-training is not about
- 22:39assisting users it is about predicting
- 22:41the next token now you can trick it into
- 22:44assisting users and uh following the
- 22:47instruction you wanted to but in general
- 22:49that's not what it retrained it for and
- 22:51this is an example of where if you ask
- 22:53GPD 3 pretty strong model to explain
- 22:55like moon landing to a six-year-old in a
- 22:57few sentences and it will follow up with
- 22:59more questions about what a 60-year-old
- 23:01might want this is not what you wanted
- 23:03the model to do right so the general
- 23:07term that people use these days is that
- 23:08they're not aligned with user intent and
- 23:12the next sections that we're going to
- 23:13cover are going to talk about how to
- 23:14align it with the user intent so that
- 23:16they don't have to trick the model into
- 23:17whatever uh we wanted to
- 23:20do and this is a kind of like desired
- 23:23completion we want at the end of
- 23:24instruction tuning um and yeah
- 23:29how do we get from those pre-trained
- 23:32models to models which can respond to
- 23:33user
- 23:34intent um again um I hope this was
- 23:38covered somewhere in the class the
- 23:39general idea of pre-training and
- 23:41fine-tuning um but what you have
- 23:43probably seen thus far is that you
- 23:45pre-train on a lot of different language
- 23:47task uh on data but then you find on
- 23:50your specific task so you're taking the
- 23:53same decoder only models and you're fine
- 23:56tuning to some task with very little
- 23:58amount of data the thing that is going
- 24:00to be different now is not that we're no
- 24:02longer fine tuning on a little amount of
- 24:04data we're going to fine tune on many
- 24:05many different tasks and we're going to
- 24:07just try to put them into a single
- 24:10usable um ux for users and this is where
- 24:14fine tuning or instruction fine tuning
- 24:16comes
- 24:19in cool um so again the recipe is not
- 24:23very very complicated here um we're
- 24:25going to collect a lot of examples of
- 24:27instruction and output Pairs and the
- 24:29instructions are going to rrange over
- 24:30several task different forms um there's
- 24:33going to be question answering they're
- 24:34going to be summarization translation
- 24:36code reasoning and so on and we're going
- 24:38to collect a lot of examples uh related
- 24:41to all those tasks and the idea is like
- 24:44I mean we'll train on instruction and
- 24:45output pairs exactly with them and then
- 24:48we're going to evaluate on some unseen
- 24:50tasks as well so this is a general
- 24:53Paradigm of instruction fine tuning and
- 24:57again it's the same idea which we
- 24:59explored in pre-training is that data
- 25:01plus scale is really important and these
- 25:04days like a mean you start off with like
- 25:06one task you're now extending it over
- 25:08thousands of thousands and thousands of
- 25:10tasks with like three million plus
- 25:11examples and this is generally like a
- 25:13broad range of tasks that you might see
- 25:14in instruction fine tuning data
- 25:16sets and yeah you might even think of
- 25:19like why are we calling it fine tuning
- 25:20anymore like it's almost starting to
- 25:22look like pre-training um but yeah we
- 25:25can these are just terms um so so you
- 25:28can decide whatever you are comfortable
- 25:30with um so yeah we we get this like huge
- 25:35instruction data set we finder model the
- 25:37next question is like how do we evaluate
- 25:39these data sets um now I think you guys
- 25:43will see another lecture on evaluation
- 25:45so I don't want to like dive too deep
- 25:46into this but generally evaluation of
- 25:49these language models is an extremely
- 25:51tricky topic um there's a lot of biases
- 25:53that you need to deal with and a lot of
- 25:55this will be covered but some more
- 25:57recent progess on this is like we are
- 25:59starting to curate these like really
- 26:01large benchmarks uh like mlu where the
- 26:05models are tested on a broad range of
- 26:06diverse knowledge and this is just one
- 26:09example which is and these are the
- 26:11topics that you will see and just to
- 26:14give some intuition of what the examples
- 26:16in these evaluation look like um under
- 26:18astronomy you might be asked what is
- 26:20true for type 1 a supernova or you might
- 26:23be asked some questions about biology
- 26:25and there's a huge host of tasks for
- 26:27this and you can typically like these
- 26:30are multi- choice questions and you can
- 26:31ask the model to answer the question if
- 26:33they're instruction fine tuned already
- 26:34hopefully they can like simply answer
- 26:36the question but you can also uh Chain
- 26:38of Thought prompt these questions or few
- 26:40short promp these questions too and
- 26:43recently there's been a huge amount of
- 26:45progress uh on this Benchmark what
- 26:48people have observed is like more and
- 26:49more pre-training on more and more data
- 26:51and larger models is simply just like
- 26:52climbing up these um the number on this
- 26:55so 90% is often seen as a benchmark Mark
- 26:58number that these model wanted to cross
- 27:00because it's roughly like human level
- 27:02knowledge or understanding and recently
- 27:05the Gemini models purply cross this
- 27:09number so yeah go
- 27:12ahead is isn't this like the entire sort
- 27:15of thing all over again right like imag
- 27:18at some point you're like okay maybe my
- 27:19methods are too like too fine tuned
- 27:22implicitly on on the image bi and isn't
- 27:25something like that happening here as
- 27:27well uh um yes I think this is a tricky
- 27:30topic because a lot of the models often
- 27:33there's this idea about whether your
- 27:35test sets are leaking into your training
- 27:37data set and there's are huge concerns
- 27:39about that it's a perfectly valid
- 27:41question to ask how do we even evaluate
- 27:44this is why evaluation is actually very
- 27:45tricky but one General thing to be
- 27:48careful about is like at some point like
- 27:50it doesn't matter what your trained test
- 27:51is if the models are generally useful if
- 27:54their models are doing useful stuff like
- 27:56does it matter like how if your test if
- 27:59you train on everything you care about
- 28:01and it does well on it like does it
- 28:03matter so yeah um again we still need
- 28:08better ways to evaluate the models um
- 28:10and to understand what methods are doing
- 28:12and how they're if they're improving the
- 28:14model or not but at some point like that
- 28:16those boundaries start to like be less
- 28:22important cool so massive progress on
- 28:24this Benchmark starting with gpd2 and
- 28:26like we're roughly at 90% which to the
- 28:29point where these benchmarks are
- 28:30starting to become unclear if like
- 28:32improvements on these are actually
- 28:33meaningful or not um in fact like most
- 28:37of the times when the models are wrong
- 28:39like you might often find that the
- 28:41question itself was unclear or ambiguous
- 28:43so all evaluation benchmarks have a
- 28:46certain limited utility to
- 28:48them so yeah um going to go over like
- 28:52another evaluation example of how this
- 28:54recipe like changes things so T5 models
- 28:58were instruction fine tuned on a huge
- 28:59number of tasks and another Trend to or
- 29:02which I think will be the theme across
- 29:04this lecture is that as your models
- 29:06become larger as they're trained on more
- 29:07data they become more and more
- 29:09responsive to your task information as
- 29:11well so what you will observe here is
- 29:13like as the number of parameters ex
- 29:15increase we have like T5 small FL T5
- 29:18small and we go up to 11 billion
- 29:20parameters where where we have T5 XXL
- 29:23you'll see that the Improvement actually
- 29:25improves like going from a pre into an
- 29:28instruction model the instruction model
- 29:30is all the more better at following
- 29:32instructions so the difference is plus
- 29:356.1 and goes to plus 26.6 as the models
- 29:37become larger so this is another very
- 29:40encouraging Trend that you probably
- 29:42should train on a lot of data with a lot
- 29:44of compute and you know pre-training
- 29:48just keeps on
- 29:50giving
- 29:52so yeah um you I hope you guys get a
- 29:56chance to like play with a lot of these
- 29:57models I I think you already hopefully
- 29:59are uh but yeah before instruction fine
- 30:02tuning something when you're asked a
- 30:04question related to disambiguation QA um
- 30:07you get something like this and it
- 30:09doesn't actually follow the let's think
- 30:11by step byep instruction very clearly
- 30:15but after instruction fine tuning it is
- 30:16able to answer the question
- 30:19here and yeah like more recently people
- 30:22have been like researching into like
- 30:24what the instruction tuning data set
- 30:25should look like there's a huge plora of
- 30:28instruction tuning data sets now
- 30:29available like this is just a
- 30:30representative diagram and there's a h
- 30:32open source Community developing around
- 30:34these as
- 30:35well um some high level lessons that we
- 30:38have learned from this
- 30:41is one lesson that I think might be
- 30:44interesting is that we can actually use
- 30:45really large strong models to generate
- 30:48some of the instruction tuning data to
- 30:49train some of our smaller models so take
- 30:52your favorite model right now gbd4 maybe
- 30:54or maybe Claud or whatever and you can
- 30:56get it to answer some question s and
- 30:58generate instruction outut pairs for
- 31:01training your open source or smaller
- 31:03model and that actually is a very
- 31:04successful recipe so instead of getting
- 31:06a human to collect all the instruction
- 31:08output pairs or getting humans to
- 31:10generate the answers you can get bigger
- 31:12models to generate the answers as well
- 31:14so that's number one thing that like has
- 31:16recently emerged another thing that are
- 31:19is being emerged or is like being
- 31:21discussed is how much data do we need I
- 31:23talked about millions of examples but
- 31:25like people have often found that if you
- 31:26have really high quality example you can
- 31:28get away with thousand examples as well
- 31:30so this is the paperless as more for
- 31:32alignment and this is still an active
- 31:34area of research on how like data
- 31:36scaling and instruction tuning affects
- 31:37the final model
- 31:39performance and yeah crowdsourcing these
- 31:42models can be effective as well so there
- 31:45are very cool benchmarks that are
- 31:46emerging like open Assistant um yeah a
- 31:49lot of activity in the field and
- 31:52hopefully like a lot more progress as we
- 31:54go on yes um a question sort of in the
- 31:58spirit of this LMA paper uh doesn't like
- 32:02code or like I don't know like math word
- 32:06problems have this desired structure so
- 32:08like shouldn't we just like be training
- 32:10code models and doing like some English
- 32:13stuff and then just be like okay this is
- 32:15the best reasoning we can get at some
- 32:17point right cuz like C code has the
- 32:19structure where where where like you're
- 32:22going sort of step by step and you're
- 32:24sort of thinking in in some way in like
- 32:27a so breaking down a concept into a
- 32:29smaller so you can consider like cod
- 32:31have like very high value tokens so
- 32:33maybe like just doing so I think there's
- 32:37again pre-training is a whole Dark Art
- 32:39that I am not completely familiar with
- 32:41but um code actually ends up being
- 32:44really useful in pre-training mixtures
- 32:46and people do like up with code data
- 32:48quite a lot um similarly like I mean but
- 32:53it depends upon what the users are going
- 32:54to use the models for right um some
- 32:56people might use it for Cotes some
- 32:57people people might do for reasoning but
- 32:59that's not the only task we care about
- 33:01as you might see later on in the next
- 33:02step we'll discuss this as well is that
- 33:05people often use these models for
- 33:06Creative task they wanted to write a
- 33:08story uh they wanted to generate a movie
- 33:10script or so on and I don't know if like
- 33:13necessarily training on reasoning only
- 33:15tasks would help with that so go ahead
- 33:18yeah would you explain like there there
- 33:20exists like some data distribution which
- 33:22is like high value for Creative tasks
- 33:28yes like I mean it seems like um a lot
- 33:32of PE people write about stories and
- 33:34everything on the internet all the time
- 33:36like which is not code and sometimes
- 33:39like there's this idea of hallucinations
- 33:40as well in this like field but you can
- 33:42often think like hey like creativity
- 33:44might be a byproduct of hallucinations
- 33:47as well so I don't know what exact data
- 33:50would like lead to like more creative
- 33:52models but generally like there's a lot
- 33:54of data or a lot of stories that are
- 33:56written on the internet which allows the
- 33:57model to be
- 33:59creative yeah but I don't know if I have
- 34:01a specific answer to the
- 34:03question cool so we discussed
- 34:05instruction fine tuning um very simple
- 34:08and very straightforward there's like no
- 34:10complicated algorithms here just collect
- 34:12a lot of data and then you can start
- 34:14leveraging the performance at scale as
- 34:16well like as models become better these
- 34:18models also become more easily
- 34:21specifiable and they become more
- 34:23responsive to task as well we're going
- 34:25to discuss some limitations and I think
- 34:27this is like really important to
- 34:28understand why we are going to optimize
- 34:29for human
- 34:32preferences cool so we talked a bit
- 34:35about this like instruction fine tuning
- 34:37is necessarily contingent on humans
- 34:40labeling the data now
- 34:44humans it's expensive to collect this
- 34:46data especially as the questions become
- 34:48more and more complex you want to answer
- 34:50questions about what which may be at
- 34:52physics PhD level or things to that
- 34:54effect these things become increasingly
- 34:57expensive to collect
- 34:59so yeah this is maybe like perhaps
- 35:01obvious like collecting data
- 35:02pre-training does not require any
- 35:04specific data you scrape data of the web
- 35:06but for instruction fing you probably
- 35:08need to recruit some people to write
- 35:09down answer to your instructions so this
- 35:11can become very expensive very quickly
- 35:14but there's more limitations to this as
- 35:16well and we we just discussing this like
- 35:18there are open-ended tasks related to
- 35:20creativity that don't really have like
- 35:22an exact correct answer to begin with so
- 35:25how do you generate the right answer to
- 35:27the kind of a
- 35:30question and yeah like language modeling
- 35:34inherently like penalizes all token
- 35:36level mistakes equally um this is what
- 35:38super fine supervised fine tuning does
- 35:40as well but often like not all mistakes
- 35:42are the same so this is an example where
- 35:46you're trying to do this prediction task
- 35:47Avatar is a fantasy TV show and perhaps
- 35:50you can see like I mean calling it an
- 35:53adventure TV show is perhaps okay but
- 35:56calling it a musical May be like a much
- 35:59worse mistake but both these mistakes
- 36:01are penalized
- 36:04equally and I think one General aspect
- 36:07which is like more becoming increasingly
- 36:09relevant is that the humans that you
- 36:11might ask might not generate the right
- 36:13or the highest quality answer your
- 36:15models are becoming increasingly
- 36:17competitive and you want in some sense
- 36:19you're going to be limited by how high
- 36:22quality the answer um humans can
- 36:24generate but often I find that the
- 36:27models are generating better and better
- 36:29answers so do we really want to keep
- 36:31relying on humans to write down the
- 36:32answers or do we want to like somehow go
- 36:34over
- 36:36that so these are the three problems we
- 36:41have talked about um with instruction
- 36:43fine tuning and we made a lot of
- 36:46progress with this but this is not how
- 36:47we got Char
- 36:49GPT um and one high level problem here
- 36:53is that even though when even when we
- 36:55are instruction fine tuning there is
- 36:57still a huge mismatch
- 36:59between the end goal is to optimize for
- 37:02human preferences generate an output
- 37:04that a human might like and we're still
- 37:08doing prediction kind of tasks where
- 37:09we're predicting the next token but now
- 37:11in a more curated data set so that's
- 37:13still a bit of a mismatch going on here
- 37:15and it's not exactly what we want to
- 37:18do hopefully like I mean I'm going to
- 37:20take a second here to pause because this
- 37:21is important to understand the next
- 37:23section and if there's any
- 37:25questions feel free to ask so is this
- 37:28step uh still taken as a as a first step
- 37:33or we discard this it's a good question
- 37:36so um I think this is still one of the
- 37:39more important steps that you take
- 37:41before taking the next step but people
- 37:43are trying to like remove the step Al
- 37:45together and jump directly to the next
- 37:47step so there's work emerging on that
- 37:49but yeah and this is still a very
- 37:51important step before we do the next
- 37:53step
- 37:57go ahead is PR two also present in
- 38:00pre-training uh and if so how how do you
- 38:03avoid that ver just by having a lot of
- 38:06data um yeah that's a great question uh
- 38:08there's two diff there's one difference
- 38:10one major difference on pre-training um
- 38:13pre-training covers a lot more text so
- 38:16um just for context like I mean as we
- 38:18talked about it's pre-training is
- 38:20roughly 15 trillion tokens whereas like
- 38:23supervised instruction fine tuning might
- 38:24be somewhere on the order of millions to
- 38:26billions of tokens so it's like few
- 38:28orders of magnitude lower typically
- 38:30you'd only see one answer for a specific
- 38:32instruction but during pre-training
- 38:34you'll see multiple text and multiple
- 38:36completions for a same kind of a prompt
- 38:39um now that's good because when you see
- 38:41multiple answers or completions during
- 38:42pre-training you sort of start to weigh
- 38:44different answers you start to put
- 38:46probability Mass on different kind of
- 38:48answers or completions but instruction
- 38:50fineing might force you to put and wait
- 38:52on only one
- 38:53answer does it okay but generally yeah
- 38:57like I mean this is a problem with both
- 38:59the stages you're
- 39:01right anything
- 39:04else
- 39:07cool so as this whole thing alludes to
- 39:11we're going to start to attempt to
- 39:14satisfy human preferences directly we're
- 39:16no longer going to like try to like get
- 39:18humans to generate some data and try to
- 39:20do some kind of a token level prediction
- 39:21loss we're going to try to optimize for
- 39:24human preferences directly and that is
- 39:27uh the general field of rlf and that's
- 39:30the final step in typically getting a
- 39:32model like CH
- 39:33GPD so we talked about how collecting
- 39:36demonstration is expensive and there's
- 39:37still a broad mismatch between the LM
- 39:39objective and human preferences and now
- 39:41we're going to try and optimize for
- 39:42human preferences
- 39:44directly so what ises optimizing for
- 39:47human preferences even mean um to like
- 39:50concretely establish that let's go
- 39:52through like a specific example in mind
- 39:54which is
- 39:55summarization um we want to train a
- 39:58model to be better at
- 39:59summarization and we want to satisfy
- 40:01human preferences so let's imagine that
- 40:03a human is able to prescribe a reward
- 40:05for a specific summary let's just
- 40:07pretend there is a reward function you
- 40:09and I can assign say like reward this is
- 40:11plus one this is minus one or something
- 40:13to that
- 40:17effect okay um so in this specific case
- 40:22we have this input X which uh which is
- 40:25about an earthquake in San Francisco so
- 40:27this is news article that we want to
- 40:28summarize
- 40:30and let's pretend that we get these
- 40:33rewards and we want to optimize this so
- 40:36we we get one summary y1 which gives us
- 40:39an earthquake hit and so on and we
- 40:40assign a reward of 8.0 and another
- 40:43summary which gives us a reward of
- 40:451.2 generally speaking like the
- 40:47objective that we want to set up is
- 40:49something of the following form where we
- 40:51want to take our language model P Theta
- 40:54which generates a completion y uh given
- 40:57an input X and we want to maximize the
- 41:00reward of rxy where X is the input and Y
- 41:04is the output summary in this specific
- 41:07task and maybe like just to like really
- 41:11concretely point out something here this
- 41:13is different from everything that we
- 41:15have done in one very specific way um we
- 41:18are sampling from the model itself in
- 41:21the bottom term if you see like we're
- 41:23using Y from P Theta everything we've
- 41:25seen so far the data is sampled from
- 41:27some other source either during
- 41:28pre-training either in supervised fine
- 41:30tuning and we're maximizing the log
- 41:32likelihood of those tokens but now we're
- 41:35explicitly sampling from our model and
- 41:37optimizing potentially a
- 41:38non-differentiable
- 41:41objective
- 41:42cool so broadly the rlf pipeline looks
- 41:46something like this and first step is
- 41:48still instruction tuning something we
- 41:49have seen up until now where we take our
- 41:52pre-trained model we instruction tune on
- 41:54a large collection of tasks and we get
- 41:57some something which starts responding
- 41:58to our desired intent or
- 42:01not but there are two more steps after
- 42:03this which are typically followed in
- 42:05creating something like instruct gbt the
- 42:07first step is estimating some kind of a
- 42:09reward model something which tells us
- 42:11given an instruction how much would a
- 42:13human like this answer or how much would
- 42:15a human hate this answer so we looked at
- 42:18something like this earlier but I didn't
- 42:19talk about how do we even get something
- 42:21like that that's the second step and
- 42:23then we take this reward model and we
- 42:25optimize it through the optimiz ation
- 42:27that I suggested earlier so the
- 42:28maximizing the expected reward under
- 42:31your language
- 42:32model and we're going to go over a lot
- 42:35over in the second and third
- 42:37steps so the first question we want to
- 42:39answer is how do we even get like a
- 42:40reward model what about what humans are
- 42:43going to like like this is a very IL
- 42:46defined problem generally speaking
- 42:49so there's there's two problems here
- 42:52that we're going to address first is a
- 42:53human in the loop is expensive so let's
- 42:55say like if I ask a model to like
- 42:57generate an answer and then I get a
- 42:59human to label with some kind of a score
- 43:02I'm doing this over millions of
- 43:03completions that is not very scalable I
- 43:07I don't want to sit around and label
- 43:08millions of examples
- 43:10so this is very easy like we're in a
- 43:13machine learning class so what are we
- 43:15going to do what we're going to do is
- 43:16we're going to train something which
- 43:18predicts what a human would like or what
- 43:19a human might not like and this is
- 43:22roughly um this is essentially a machine
- 43:25learning problem where we take these
- 43:26Rewards scor scores and we try to train
- 43:28a reward model to predict given an input
- 43:31and output what the reward scores would
- 43:33look
- 43:34like simple simple machine learning
- 43:37regression style problem uh you might
- 43:38have seen this
- 43:40earlier
- 43:43cool now there's a bigger problem here
- 43:46and sorry go ahead one so do we use like
- 43:49I don't know like just embedding
- 43:52withier we use a real language model to
- 43:55do that um that's a good question
- 43:58generally like what we do is like we
- 44:00still typically need reward models where
- 44:02they need to be able to understand the
- 44:04text really well so like bigger models
- 44:06and like they're typically initialized
- 44:07from the language model that you trained
- 44:09pre-trained as well so you typically
- 44:11start with the pre-trained language
- 44:12model and do some kind of prediction
- 44:14that we'll talk about and they'll give
- 44:16you a
- 44:17score how do you if you're doing that
- 44:20how do you separate X and Y how does the
- 44:22language model know which part it
- 44:24doesn't need to it can put the X and Y
- 44:27like it only sees X and Y as an input so
- 44:29it doesn't need to T typically see it
- 44:31separated it's just going to predict a
- 44:33score at the end okay yeah the X and Y
- 44:35is more for notational convenience here
- 44:38because for us X and Y are different X
- 44:40is a question user asked and Y is
- 44:42something the model generated but you
- 44:44shove the whole thing into you shove the
- 44:46whole thing into yes
- 44:48cool now this is the bigger problem here
- 44:51and human judgments are very noisy we
- 44:54have talked about we want to assign a
- 44:55score to a completion this is something
- 44:58that's like extremely non-trivial to do
- 45:00so if I give you a summary like this
- 45:02what score are you going to assign on a
- 45:04scale of 10 if you ask me on different
- 45:07days I'll give a different answer first
- 45:08of all but across humans itself this
- 45:12this number is not calibrated in any
- 45:14meaningful way so you could assign
- 45:17number of 4.1 6.6 and different humans
- 45:19would just simply assign different
- 45:20scores and there are ways to address
- 45:23this you can like calibrate humans you
- 45:24can give them a specific rubric you can
- 45:26like talk to them but it's a very
- 45:28complicated process and like still like
- 45:30there's a lot of room for judgment which
- 45:32is not typically very nice for training
- 45:34a model like this if your labels can
- 45:37vary a lot it's just hard to
- 45:40predict so the way this is addressed is
- 45:44that instead of trying to predict the
- 45:45reward label directly you actually want
- 45:47to set up a problem in a slightly
- 45:49different way what is something much
- 45:50easier for humans to do is give them two
- 45:53answers or maybe many answers and tell
- 45:55them ask them which one is better so
- 45:58this is where the idea of asking humans
- 46:01to rank answers comes in so if I give
- 46:04you a whole news article and ask you
- 46:07which summary is better you might be
- 46:08able to give me a ranking that oh this
- 46:10second summary is the worst but the
- 46:12first one is better and the third one is
- 46:13somewhere in the middle between those
- 46:14two so you get like a ranking which
- 46:16gives you um preference over
- 46:19summaries and hopefully like I mean you
- 46:22can see like the idea that is important
- 46:24here is that even when we have some kind
- 46:27of a consistent utility function even
- 46:29when I have it's much easier to compare
- 46:32to something and know that which is
- 46:33better than this rather than ascribing
- 46:34it an arbitrary number on a scale and
- 46:37that's why the signal from something
- 46:39like this is a lot
- 46:41better now how do we get like we talked
- 46:44about we need like we get this kind of a
- 46:46preference data and now we need some
- 46:47kind of a reward score out of this and
- 46:50we shove in like our input we shove in a
- 46:53summary as well and we still need to get
- 46:54a score out of this but it's not clearly
- 46:56obvious to me like how do we take this
- 46:58data and convert it into that kind of
- 47:00score
- 47:02um incomes are pretty good friends named
- 47:05Bradley Terry um and essentially like
- 47:10there's a lot of study in like many in
- 47:12economics and like psychology which
- 47:14basically tries to model how humans make
- 47:18decisions in specific case like this
- 47:20Brad lary model essentially says that a
- 47:23probability that a human chooses answer
- 47:25y1 over y two is based on the difference
- 47:30between the rewards that humans assign
- 47:33internally and then you take a sigmoid
- 47:35around it so if you have looked at
- 47:37binary classification before uh the
- 47:40logic is simply the difference between
- 47:41the reward of some y1 minus Y2 or the
- 47:45difference between the winning
- 47:47completion and the losing
- 47:51completion is everybody with me till
- 47:53this point
- 47:57so the idea is that like if you have a
- 48:00data set where given y1 and Y2 where y1
- 48:03is a winning completion and we have a
- 48:05winning completion YW and a losing
- 48:07completion y l um the winning completion
- 48:10should score higher than the losing
- 48:12completion go ahead sorry what is J is
- 48:15that a log or like sorry what what like
- 48:19what is the type of J like this number
- 48:21here that we're getting as the
- 48:23expectation is it a log prop or what is
- 48:25it it's an log prop so it will be a
- 48:28scaler at the end sigmoid is so you're
- 48:31taking the let's say you have a reward
- 48:33model which gives a
- 48:34score R1 to like YW and R2 to y l you
- 48:39subtract that number you get another
- 48:40number you put it into sigmoid and you
- 48:42get a probability because sigmoid will
- 48:45convert a logit into probability and
- 48:47then you take a logarithm of that and
- 48:50you take the expectation of everything
- 48:51and you get this final number which
- 48:53tells you how good your reward model is
- 48:55doing on the entire data set
- 48:58so like a good model of humans should
- 48:59behave like this a good model of humans
- 49:02would um score very low here so it would
- 49:05generally assign a higher reward to the
- 49:07winning completion and generally assign
- 49:09a lower reward to the losing
- 49:12completion
- 49:14cool the math is just beginning so um
- 49:18hold on to your seats
- 49:21um cool so now let's see where we are we
- 49:24have a pre-trained model p p d y given X
- 49:28and we got this like fancy reward model
- 49:30which tells us that hey how we have a
- 49:32model of humans and it can tell us which
- 49:34instruction which answer they like and
- 49:36which in answer did not
- 49:38like now to do rlf generally like I mean
- 49:42we have discussed what this will look
- 49:43like uh we will copy our pre-train model
- 49:46or instruction tune model and we'll
- 49:48optimize the parameters for those models
- 49:51and I suggested that the param objective
- 49:53that we want to optimize is the expected
- 49:56reward when we sample completions from P
- 49:59Theta and we're going to optimize our
- 50:02learned reward model instead of like the
- 50:03true reward model which humans would
- 50:05have typically assigned do you guys see
- 50:07any problem with
- 50:09this
- 50:11um is there something that's wrong here
- 50:14or like that might go wrong if we do
- 50:16something along these
- 50:22lines go for itel
- 50:28it might collapse yes okay um but
- 50:31generally at least from my intuition
- 50:32like if you're ever doing something and
- 50:34you have you're optimizing some learned
- 50:36metric I'd be very careful because
- 50:39typically our loss functions are very
- 50:40clearly defined but here my reward model
- 50:42is learned what when it's learned it
- 50:44means it will have
- 50:46errors yes so it's going to be trained
- 50:49on some distribution it will generalize
- 50:51as well but it will have errors and when
- 50:54you're optimizing against a learn model
- 50:57it will tend to hack the reward model so
- 50:59it might exploit the reward model might
- 51:02erroneously assign a really high score
- 51:04to a really bad completion if your
- 51:06policy learns or if your language model
- 51:08learns to do that it will completely
- 51:10Hack That and start generating those
- 51:11gibberish
- 51:13completions
- 51:15so just as a general machine learning
- 51:17tip as well if you're optimizing a learn
- 51:19metric be careful about what you're
- 51:20optimizing and make sure that it's
- 51:22actually
- 51:23reliable um and the way and this is
- 51:27obviously not desirable like I mean if
- 51:28you start optimizing this objective
- 51:30you're going to converse to gibberish
- 51:31language models very very quickly so
- 51:33typically what people do is that you
- 51:35want to add some kind of a penalty that
- 51:37like avoids it drifting too far from its
- 51:40initialization and why do we want to do
- 51:42that like if it cannot Drift from too
- 51:44far from its initialization we know the
- 51:46initialization of the model is a decent
- 51:47language model and we know it is not
- 51:50necessarily satisfying this reward model
- 51:52too much and we also know that like the
- 51:54reward model is trained on a distrib
- 51:56ution of completions where the um
- 51:59initial model is so typically we when we
- 52:02talk about training this reward model we
- 52:04have trained on certain completions
- 52:05which are sampled from this initial
- 52:07distribution so we know the reward model
- 52:09will be somewhat reliable in that
- 52:10distribution so we're just going to
- 52:12Simply add a penalty which tells us that
- 52:15you should not drift too far away from
- 52:17the initial distribution and just to go
- 52:20over this we want to maximize the
- 52:22objective where we have RMF which is our
- 52:25learned one model but we're going to add
- 52:29this term beta log ratio and the ratio
- 52:31is our the model we're optimizing P
- 52:33Theta and our initial model PP PT and
- 52:37what this says is that if we assign a
- 52:39much higher probability to certain
- 52:41completion as compared to our pre-train
- 52:43model you're going to add an
- 52:44increasingly large penalty to
- 52:46it and simply you're paying a price for
- 52:49drifting too far from initial
- 52:50distribution if you guys have taken like
- 52:53machine learning this the expectation of
- 52:55this quantity can is exactly the cbak LI
- 52:58Li Divergence or k Divergence between P
- 53:01Theta and PPT so you're penalizing
- 53:04drifting between two distributions go
- 53:07forhead question shouldn't you also do
- 53:09this like add a penalty in the previous
- 53:12version where you had to find huning or
- 53:14is this only relevant for the RL HF
- 53:18that's a good question so um I think
- 53:21people do add some kinds of
- 53:22regularization in fine tuning it's not
- 53:25nearly not as critical when you're doing
- 53:27this with RL like the incentive is to
- 53:29exploit this reward model as well as
- 53:32much as possible and we'll see examples
- 53:34where like the Learned reward predicts
- 53:37like it's doing really well but the true
- 53:38reward models are completely garbage so
- 53:41it's much more important in this
- 53:47optimization cool
- 53:49so now this assume does this course does
- 53:53not assume background on reinforcement
- 53:54learning so we're not going to go into
- 53:56reinforce learning but I just want to
- 53:57give a very high level intuition about
- 53:59how this works and reinforcement
- 54:01learning is not typically just used for
- 54:03language model it's been applied to
- 54:04several uh domains of Interest game
- 54:07playing agents re um robotics developing
- 54:11chip designs and so on and the interest
- 54:15between like RL and model LMS it's also
- 54:18like dates back to roughly like 2016 as
- 54:20well but like it's been really
- 54:22successful recently and especially with
- 54:24the success of rlf
- 54:27um the general idea is that we're going
- 54:28to use our model that we're optimizing
- 54:30to generate several completions for an
- 54:32instruction um we're going to compute
- 54:35the reward under our learned reward
- 54:37model and then we're going to simply try
- 54:39and like update the update our model to
- 54:42increase the probability on the high
- 54:44reward completions so when we sample a
- 54:46model we'll see completions of varing
- 54:48quality and we'll see some good
- 54:49completions good summaries for our task
- 54:51some bad summaries for our task and
- 54:53we'll try to update our log
- 54:54probabilities in a way such that uh the
- 54:57reward for when you use a updated model
- 55:00you're typically in the higher reward
- 55:04region does the high level summary like
- 55:06make
- 55:07sense
- 55:10cool and rhf is incredibly successful I
- 55:13think this is a very good example of um
- 55:15this is the same summarization example
- 55:18and I think the key Point here is that
- 55:22the performance improves by increasing
- 55:24the model size for sure we have seen
- 55:26this in many different example what you
- 55:28can actually see is that even very small
- 55:30models can outperform human completions
- 55:33if you train it with with rlf and this
- 55:36is exactly the result you see here the
- 55:38reference summaries are human generated
- 55:39and when you evaluate when you ask
- 55:42humans which ones they prefer they often
- 55:44prefer the model generated summary over
- 55:46the human generated summary and this is
- 55:47something you only observe with rlf even
- 55:50at small scales and again the same
- 55:51scaling phenomena still holds here
- 55:53bigger models do become more responsive
- 55:55but are of itself is very impactful
- 56:00here
- 56:02cool the problem with rlf is that it's
- 56:05just incredibly complex like um I gave
- 56:08you a very high level summary that's
- 56:10like doesn't that there's whole courses
- 56:12on this for a reason um so it just and
- 56:15this image is not for you to understand
- 56:17it's just completely to intimidate you
- 56:20um
- 56:21so um you want to fit a value function
- 56:24to something there's you have to sample
- 56:25the model a lot it can be sensitive to a
- 56:28lot of hyperparameters so there's a lot
- 56:30that goes on here and yeah um if you
- 56:35start implementing an rlf pipeline it
- 56:37can be very hard and this is the reason
- 56:39why like a lot of rlf was restricted to
- 56:41very very like high compute High
- 56:43resource places and it was not very
- 56:46accessible so what we're going to talk
- 56:48about and cover in this course is
- 56:49something called direct preference
- 56:50optimization which is a much simpler
- 56:52alternative to R LF and hopefully like
- 56:54that's much more accessible but but
- 56:56please bear with me there will be a lot
- 56:58of math on here but the end goal of the
- 57:00math is to make come up with a very
- 57:02simple algorithm so hopefully like it's
- 57:04um and feel free to stop me and ask me
- 57:10questions you need in terms of like gbt
- 57:134 versus three like how much do the
- 57:15number of parameters in the base model
- 57:17help with like need to reduce the number
- 57:19of parameters or like in order sorry R
- 57:22the number of like examples from humans
- 57:25for RFS that work
- 57:26well yeah that's a really good question
- 57:28so generally speaking as the if you hold
- 57:31the data set size constant and simply
- 57:33increase the mod size it will improve
- 57:35quite a lot sure but the nice thing is
- 57:38that you can reuse the data and you can
- 57:39keep adding data yeah uh as you keep
- 57:42like scaling models up so typically like
- 57:44nobody tries to like reduce the amount
- 57:46of data collection yeah right you just
- 57:47keep increasing both the things
- 57:51out cool so we talked about rlf and the
- 57:55current pipeline is some something like
- 57:58um we train a reward model on the
- 57:59comparison data that we have seen so far
- 58:01and we're going to optimize we're going
- 58:03to start with our pre-train our
- 58:04instruction tune model and convert it
- 58:05into an rlf model using the
- 58:07reinforcement learning
- 58:09techniques now the really the key idea
- 58:11in direct preference optimization is
- 58:13what if we could just simply write a
- 58:15reward model in terms of our language
- 58:17model itself now to intuitively
- 58:20understand that like what is going on a
- 58:22language model is assigning
- 58:23probabilities to whatever is the most
- 58:25plausible completion next but those
- 58:28plausible completions might not be what
- 58:29we intended but you could restrict the
- 58:31probability simply to the completions
- 58:34that a human might like and then the log
- 58:36probabilities of your model might
- 58:38represent something which the humans
- 58:39might like and not just some arbitrary
- 58:41completion on the internet so there is a
- 58:43direct correspondence between the log
- 58:46probability that a language model
- 58:48assigns and how much a human might like
- 58:50the answer they can have like a direct
- 58:52correspondence in them and this is not
- 58:55some arbitrary intuition that I'm trying
- 58:56to like come up with we will derive this
- 59:00mathematically so the general idea with
- 59:03direct preference optimization is going
- 59:04to be we're going to write down reward
- 59:06model in terms of our language model and
- 59:09now that we can write our reward model
- 59:10in terms of our language model we can
- 59:12simply solve directly fit our reward
- 59:15model to the preference data we have and
- 59:19we don't need to do the RL Step at all
- 59:21so we started off with some preference
- 59:22data and we simply fit our reward model
- 59:24to it which directly optimiz as the
- 59:26language
- 59:28parameters and maybe at a high level why
- 59:31is this like even possible like we did
- 59:33this like really cumbersome process with
- 59:34fitting a reward model and optimizing it
- 59:37but in the whole process the only
- 59:39external information that was being
- 59:41added to the system like was human
- 59:43labels labels on the preference data
- 59:45when we optimize a learned reward model
- 59:47there's no new information being added
- 59:49into the system so this is why something
- 59:52like this is even possible for quite a
- 59:54few years this was not obvious obvious
- 59:56but like as you will see like some of
- 59:58these results like start to make sense
- 1:00:00so we're going to derive direct
- 1:00:02preference
- 1:00:03optimization I'll I'll be after I'll be
- 1:00:05here after the class as well if you have
- 1:00:06questions but I'll hopefully like this
- 1:00:08is clear
- 1:00:11so yes um we discussed that we wanted to
- 1:00:14solve this expected reward problem where
- 1:00:17we want to maximize the expected reward
- 1:00:19but we subtract this term which is the
- 1:00:21beta log ratio which essentially
- 1:00:22penalizes the distance between where our
- 1:00:25current model is and where we started
- 1:00:27off so we don't want to drift too far
- 1:00:28away from our um from where we
- 1:00:33started now it turns out that this
- 1:00:36specific problem instead of doing like
- 1:00:38an iterative routine um there's actually
- 1:00:41a close form solution to this problem so
- 1:00:45the close form solution looks something
- 1:00:46like this um again if you have seen the
- 1:00:50boltzman distribution or something to
- 1:00:52that effect before this is very
- 1:00:54basically the same idea but the idea is
- 1:00:56this that we're going to take a
- 1:00:57pre-train distribution PPT y given X and
- 1:01:00we're going to rade the distribution by
- 1:01:02the expected
- 1:01:03reward so if if if a completion has a
- 1:01:07very high reward it's going to have a
- 1:01:09higher probability mass and if it has a
- 1:01:11lower reward it's going to have a lower
- 1:01:12probability mass and it's determined by
- 1:01:14the expected reward and beta is a
- 1:01:16hyperparameter which essentially governs
- 1:01:18like what is the trade-off between the
- 1:01:20reward model and the constraint and as
- 1:01:24beta becomes lower and lower you're
- 1:01:26going to start paying more and more
- 1:01:27attention to the reward
- 1:01:29model so the probabilities look
- 1:01:32something like this and there's this
- 1:01:34like really annoying term this ZX and
- 1:01:37the reason why it exists is that the
- 1:01:39numerator by itself is not normalized
- 1:01:42it's not a probability distribution so
- 1:01:44to construct like an actual probability
- 1:01:46distribution you have to normalize it
- 1:01:48and ZX is simply just this
- 1:01:50normalization so if we write ZX out is
- 1:01:53the sum of all y okay yeah and that's
- 1:01:56exactly like it's some overall wise for
- 1:01:58a given instruction and that's exactly
- 1:02:00why this is very pesky is like it's
- 1:02:02intractable if I take an instruction and
- 1:02:04try to sum over every possible
- 1:02:06completion and not just like
- 1:02:07syntactically correct ones every single
- 1:02:09possible we have 50,000 tokens maybe
- 1:02:12even more and the completions can go
- 1:02:14arbitrary long so this space is
- 1:02:15completely intractable this quantity is
- 1:02:17not easy to approximate
- 1:02:21even um so the main point here is that
- 1:02:25you if you're given in a reward model
- 1:02:26you can actually there does exist at
- 1:02:28least a close form solution which tells
- 1:02:30us what the optimal policy will look
- 1:02:31like or optimal language model will look
- 1:02:33like but if you do a little bit of
- 1:02:35algebra just move some terms around take
- 1:02:37a logarithm here or there I I promise
- 1:02:39this is not very complicated you can
- 1:02:41actually Express the reward model in
- 1:02:43terms of the language model itself and I
- 1:02:46think this term is reasonably intuitive
- 1:02:48as well uh what it says is that um a
- 1:02:51completion y hat has a high reward if
- 1:02:54the model my optimal policy assigns a
- 1:02:57higher probability to it relative to my
- 1:03:00initialized model and this is scal by
- 1:03:03Beta so the beta log ratio is what we're
- 1:03:05looking at
- 1:03:07here and the partition function let's
- 1:03:09just ignore it for now but it's
- 1:03:10intractable but the beta log ratio is
- 1:03:13the key part
- 1:03:15here is everyone following
- 1:03:18along awesome okay so right now I'm
- 1:03:23talking about optimal policies but
- 1:03:26really like every policy is probably
- 1:03:28optimal for some kind of a reward right
- 1:03:30like this is mathematically true as well
- 1:03:32so the important bit here is that you
- 1:03:34can actually Express you take a current
- 1:03:37policy take your initialized model and
- 1:03:39you can get some kind of a reward model
- 1:03:41out of it and this is the exact identity
- 1:03:44which leads to this so reward model can
- 1:03:46be expressed in terms of your language
- 1:03:48model baring the log partition term
- 1:03:52which we'll see what happens to it go
- 1:03:55for sorry I don't know how you got like
- 1:03:57why is it that we can swap because there
- 1:03:59is a thing that we're trying to optimize
- 1:04:00and how do p star turn into P yeah um
- 1:04:04for now like we're not optimizing any
- 1:04:05reward model okay all I'm saying is that
- 1:04:08if I take my current language model it
- 1:04:10is it probably represents some kind of a
- 1:04:12reward model
- 1:04:15implicitly because of this relationship
- 1:04:17because this holds for every P star and
- 1:04:19every reward model what I'm saying is
- 1:04:22that like there if I plug in my current
- 1:04:24language model it also represents some
- 1:04:26kind of a reward model I'm not saying
- 1:04:27it's optimal
- 1:04:29okay but I want say because at the
- 1:04:31beginning uh PRL is PPT yes and so we
- 1:04:36just get that the reward is basically
- 1:04:38zero and so what what do we do initially
- 1:04:41it's zero but like we can optimize the
- 1:04:42parameters okay okay yeah um yeah but
- 1:04:45that's a good observation that it's
- 1:04:46basically zero in the beginning but how
- 1:04:48do we start optimizing
- 1:04:50it I'll get to okay okay any other
- 1:04:54questions so the idea is that given the
- 1:04:56language model you have model such that
- 1:05:02that makes the language model
- 1:05:04op yes that's uh that's the next step
- 1:05:08yes uh but the key idea is that like my
- 1:05:11log my language model the probabilities
- 1:05:13already implicitly Define a reward model
- 1:05:16I think that's really the main point
- 1:05:18here and this mathematical relationship
- 1:05:20is
- 1:05:21exact cool now like I mean I'm obviously
- 1:05:25ignoring like the elephant in the room
- 1:05:27here which is the partition function um
- 1:05:30it's not going to magically vanish away
- 1:05:31so like if this was just the beta log
- 1:05:33ratio that would be really nice I can
- 1:05:35compute all these quantities I know how
- 1:05:37to compute the log probability under my
- 1:05:38language model I know how to compute the
- 1:05:40log probability under my pre-train model
- 1:05:43and I can compute the reward score and I
- 1:05:45can optimize this but I don't know what
- 1:05:47to do about my Lo log partition function
- 1:05:50this is where something fun happens so
- 1:05:54recall what the the reward modeling
- 1:05:56objective was uh when we started off
- 1:05:59like we started off with the friends
- 1:06:00Bradley Terry again and what we really
- 1:06:03wanted to optimize was the reward
- 1:06:05difference between the winning
- 1:06:06completion and the losing
- 1:06:08completion um and really like I mean
- 1:06:11that's all we care about we don't care
- 1:06:12about the exact reward itself what we
- 1:06:15care about is maximizing the difference
- 1:06:16between the the difference between
- 1:06:19winning and losing completion and that's
- 1:06:21actually really key here because if you
- 1:06:24plug in the definition of of the RM
- 1:06:27Theta there what you'll observe is that
- 1:06:30the partition function actually just
- 1:06:32cancels out now why does it cancel out
- 1:06:36um the input is exactly the same the x
- 1:06:39is actually exactly the same in the
- 1:06:41difference so the partition function ZX
- 1:06:43will just cancel out like it's the same
- 1:06:45in both the terms so what you get is
- 1:06:47that the reward difference between the
- 1:06:48winning and losing completion is the
- 1:06:50differences between the beta log ratio
- 1:06:51for the winning and losing
- 1:06:53completion you can plug in the terms you
- 1:06:56can work it out it's fairly simple so
- 1:06:59the partition function which was our
- 1:07:00like um which was something we could not
- 1:07:03address we could not compute actually
- 1:07:04just simply vanished away I'm so sorry Z
- 1:07:07doesn't appear in theary mod um but it
- 1:07:11appears here in this equation so how
- 1:07:14does plug in
- 1:07:16model um so we're going to take this
- 1:07:19equation uh the last line that you see
- 1:07:22and we're going to plug in in place of
- 1:07:24RMF
- 1:07:26okay so um and in this the first loss
- 1:07:31equation oh I see got yeah so the first
- 1:07:33loss equation is the broadly ter loss
- 1:07:35model
- 1:07:37cool so this really is it like I mean
- 1:07:40the key observation is we could express
- 1:07:42our reward model in terms of language
- 1:07:43model and our problems with the
- 1:07:45partition function actually go away
- 1:07:46because we were optimizing the Brad lary
- 1:07:48model and um what you get is something
- 1:07:51like this is that um we're going to
- 1:07:55Express the loss function directly in
- 1:07:57terms of our language model parameters
- 1:07:59Theta and we're going to be able to
- 1:08:02directly optimize on our data um without
- 1:08:05doing any RL steps or not and this is
- 1:08:07simply a binary classification problem
- 1:08:10so we're really just trying to classify
- 1:08:12whether an answer is good or bad and
- 1:08:14that's really what we're
- 1:08:16doing before I go on like people want to
- 1:08:19like absorb this in like I mean feel
- 1:08:22they're okay with it
- 1:08:26I don't get where they why good and why
- 1:08:28win and why lose come from are they
- 1:08:30human and or they good question um it's
- 1:08:34the same data set we started with in rlf
- 1:08:36as well but the way the process works is
- 1:08:39that you take a set of instructions and
- 1:08:40get the model to generate some answers
- 1:08:42and then you get humans to label which
- 1:08:44answer they prefer so they're model
- 1:08:46generated uh typically they can be human
- 1:08:48generated as well but they're typically
- 1:08:50model generated and then you get some
- 1:08:52preference labels okay all you need is a
- 1:08:55label saying which is a better
- 1:08:58answer what do you lose here like you
- 1:09:02must be losing some information because
- 1:09:03of the lack of information about like
- 1:09:08other you're canceling out your your
- 1:09:11your because of the lack of uh any
- 1:09:13information about the partition function
- 1:09:15yeah you are bound to lose information
- 1:09:18about like other possible completions
- 1:09:20which you would have taken into account
- 1:09:22in like standard rlf right
- 1:09:26um that's a really good question I don't
- 1:09:27think I'll be able to completely answer
- 1:09:29this question in time but like partition
- 1:09:32function is almost kind of a free
- 1:09:33variable so I think the problem here is
- 1:09:35that the reward model there think of
- 1:09:38when you there's many reward models that
- 1:09:40satisfy this optimization so there's a
- 1:09:43free variable here that you can actually
- 1:09:45completely remove and that's what this
- 1:09:47optimization benefits from so think of
- 1:09:49it this way like if I assign something a
- 1:09:50reward of plus one and assign something
- 1:09:52a reward of minus one that's basically
- 1:09:54the same as saying as if it's a reward
- 1:09:55of plus
- 1:09:5799 and it will give you the same loss
- 1:10:01right so um that scale doesn't that
- 1:10:05shift invariant in a ways is that like
- 1:10:08isn't that somehow like not what you
- 1:10:11want though like like okay like if if
- 1:10:15you have if you're actually training a
- 1:10:16reward model right like 199 is like much
- 1:10:20you should pay much less attention to
- 1:10:22that as compared to
- 1:10:23like one right
- 1:10:26Zer or something what we're assuming is
- 1:10:27our choice model here is like if a human
- 1:10:30prefers something over the other like
- 1:10:32the probability is governed only by the
- 1:10:34difference between the rewards so that's
- 1:10:37an assumption that every rlf also makes
- 1:10:39and like DPO also makes now is that
- 1:10:42assumption true not completely true but
- 1:10:45like it it holds to a fairly large
- 1:10:49degree but that's a good question
- 1:10:52yeah cool um I'll move on and rest of
- 1:10:56time um and really like I mean the goal
- 1:10:58of this plot is to like we actually get
- 1:11:00fairly performant models when we
- 1:11:02optimize things with DPO um we so in
- 1:11:06this plot I think the main thing that
- 1:11:07you should look at is po which is the
- 1:11:08typical rlf Pipeline and we are
- 1:11:10evaluating the models for summarization
- 1:11:12and we're comparing to human summaries
- 1:11:15and what we find is that DP and BP sort
- 1:11:17of do similarly but you're really not
- 1:11:19losing much by just doing the DPO
- 1:11:21procedure instead of R LF and that's
- 1:11:23really compelling because DP is simply a
- 1:11:24classif app ation loss instead of like a
- 1:11:26whole reinforcement learning
- 1:11:29procedure so I want to quickly summarize
- 1:11:32um what we have seen thus far is that we
- 1:11:35want to optimize for human preferences
- 1:11:37so um and the way we do this is like
- 1:11:40instead of relying on uncalibrated
- 1:11:41scores we're getting comparison data and
- 1:11:43feedback on that and we use this ranking
- 1:11:46data to either do something like rlf
- 1:11:48where we first fit a reward model and
- 1:11:50optimize using reinforcement learning um
- 1:11:54or we do something that like direct
- 1:11:55preference optimization we simply take
- 1:11:57the data set and do a classification
- 1:11:59loss on that um and yeah like there's
- 1:12:02trade offs in these algorithms like
- 1:12:04people when they have a lot of
- 1:12:06computational budget they typically
- 1:12:07maybe go for rlf or some routine like
- 1:12:09that but if you're really looking to get
- 1:12:12the bank for your buck like I mean you
- 1:12:13might want to go for DPO and if and
- 1:12:15that's like probably going to work out
- 1:12:17of the box um it's a still an active
- 1:12:20area of research people are still trying
- 1:12:21to understand how to like best work with
- 1:12:23these algorithms so like I'm not making
- 1:12:25any strong claims here but like both of
- 1:12:26these algorithms are very effective DP
- 1:12:28is just much simpler to work
- 1:12:30with
- 1:12:32cool um so yeah like I mean let's see
- 1:12:35like we went through all this
- 1:12:36instruction tuning rlf what do we get um
- 1:12:41instruct GPD is the first model which
- 1:12:43sort of followed this pipeline it
- 1:12:45defined this pipeline so we got models
- 1:12:47which did 30,000 or so tasks remember
- 1:12:50when we were doing like only one task
- 1:12:52and now we have scaled it up from th000
- 1:12:53tasks to like 30,000 different task with
- 1:12:55many many different examples so that's
- 1:12:57like where we are with instruct GPT and
- 1:13:00it follows this pipeline that we just
- 1:13:02described in this case they're following
- 1:13:03a specific rlf pipeline where we
- 1:13:05explicitly fit a reward model and then
- 1:13:07do some kind of a reinforcement learning
- 1:13:09routine on top of it um and yeah like
- 1:13:14the task collected from labelers looks
- 1:13:15something like this um I leave it to
- 1:13:18your imagination or you can look at the
- 1:13:19details but how we started off with this
- 1:13:22model was something like completions we
- 1:13:24see from G GPD 3 uh which you know
- 1:13:27explained the moon Ling to sixer and
- 1:13:29like it is not really following the
- 1:13:31instructions but instruct GPD will give
- 1:13:33you something which is Meaningful so
- 1:13:35it's inferring what a user wanted from
- 1:13:37the specific instruction and it's
- 1:13:39converting to a realistic answer that a
- 1:13:40user might
- 1:13:43like and yeah these are just more
- 1:13:46examples of what an instruct GPD like
- 1:13:48model would do whereas your base model
- 1:13:50might not follow the instructions to
- 1:13:52your desired intentions
- 1:13:56and yeah like we went from instruct GPD
- 1:13:58to chart GPD and it was essentially this
- 1:14:01pipeline um the key difference here is
- 1:14:04that it is still doing the instruction
- 1:14:06tuning but it is more optimized for
- 1:14:08dialogue more optimized for interacting
- 1:14:10with users so the core algorithmic
- 1:14:13techniques that we discussed today are
- 1:14:15what give us CH GPD but you have to be
- 1:14:17really careful about the kind of data
- 1:14:19you're training on and that's really the
- 1:14:21whole game um but this is the foundation
- 1:14:24for CH GPD
- 1:14:26and yeah it it follows the same pipeline
- 1:14:29as well and you might look at you might
- 1:14:32interact with ch gbd I'm sure you all
- 1:14:34have interacted with it some form or not
- 1:14:35but like this is an example of what a CH
- 1:14:37gbd interaction might look
- 1:14:40like um you want to make a gen Z so like
- 1:14:44I mean you can you know the idea here is
- 1:14:46that it's like very good at responding
- 1:14:47to instructions and intent this is not
- 1:14:49something that we could like even fuse
- 1:14:51shot in very easily uh these are kind of
- 1:14:54instructions are hard to come examples
- 1:14:56for but like this is probably not
- 1:14:58something to trained on either but it's
- 1:14:59able to like infer the intent and
- 1:15:01generalize very very nicely and that's
- 1:15:03something I find personally very
- 1:15:07remarkable cool and there's been a lot
- 1:15:10of progress on the open source front as
- 1:15:12well so like DPO is much simpler and
- 1:15:14much more efficient and essentially all
- 1:15:16the open source models these days are
- 1:15:18using DPO so this is a leaderboard that
- 1:15:21is maintained by hugging hugging face a
- 1:15:23so like I mean N9 out of 10 more models
- 1:15:25here are trained with DPO so that's been
- 1:15:28something that's been enabled the open
- 1:15:29source Community to instruction tune
- 1:15:31their model betters as well and same is
- 1:15:34being used in many production models now
- 1:15:36as well mistol is using DPO llama 3 used
- 1:15:38DPO so these are very very strong models
- 1:15:41which are nearly gp4 level and they're
- 1:15:43also like starting to use um these
- 1:15:46algorithms as well and something that's
- 1:15:48very cool cool to see is like like we
- 1:15:51went through all this like optimization
- 1:15:52and like I mean math and stuff but what
- 1:15:54is really fundamentally changing in the
- 1:15:56behavior and I think this is a really
- 1:15:58good example is that if you simply ask
- 1:16:01an instruction for and ask for an sft
- 1:16:03output from an instruction tune model
- 1:16:05you'll get something like this but when
- 1:16:07you RL of the model you actually get a
- 1:16:09lot more details in your answer and
- 1:16:11they'll probably organize the answers a
- 1:16:13little better and there's something that
- 1:16:14they maybe humans prefer that's why it's
- 1:16:17an property that is emerging in these
- 1:16:19model but it's something that's a very
- 1:16:21clear difference between simply
- 1:16:24instruction tude models and some models
- 1:16:27which are
- 1:16:30rft so yeah um we discuss like this
- 1:16:34whole rlf routine where we are directly
- 1:16:37modeling the preferences and we are
- 1:16:38generalizing Beyond label data um and we
- 1:16:41also discussed RL can be very tricky to
- 1:16:43um correctly Implement though DPO sort
- 1:16:45of implements this or like avoid some of
- 1:16:48these issue and we briefly also touched
- 1:16:50upon the idea of reward model and reward
- 1:16:52hacking um and when you're optimizing
- 1:16:56for learned reward models you will often
- 1:16:58see this example is that there's a way
- 1:17:01for it to just simply crash into um the
- 1:17:05object some keep repe repetitively
- 1:17:08crashing the board to get more and more
- 1:17:09points that wasn't the goal of this game
- 1:17:12so um this is a very common example that
- 1:17:15is shown for reward hacking if you do
- 1:17:18not specify Rewards well the models can
- 1:17:20like learn weird behaviors which are not
- 1:17:23your desired intent and there's
- 1:17:24something a lot of people worry about as
- 1:17:26well um part of the reason is
- 1:17:28reinforcement learning is a very strong
- 1:17:29optimization algorithm it's at the heart
- 1:17:31of alpha go Alpha zero uh which like
- 1:17:34results in superhuman models so you have
- 1:17:36to be careful about how you specify
- 1:17:38things and the other thing is like even
- 1:17:40optimizing for human preferences is
- 1:17:42often not the right thing because humans
- 1:17:43are not do not always like things which
- 1:17:46are in their best interest so something
- 1:17:48that emerges is that they like
- 1:17:49authoritative and helpful answers but
- 1:17:51they often like don't necessarily like
- 1:17:54truthful answers
- 1:17:55so one property that happens is like is
- 1:17:58that they'll prefer authoritativeness
- 1:18:00more than correctness which is maybe
- 1:18:02like not something nice please go ahead
- 1:18:04on those lines I'm curious if maybe like
- 1:18:07chbt being so like now widely used by
- 1:18:10the public will maybe change the like
- 1:18:12how people were like made the rewards
- 1:18:14because I at least feel like now when I
- 1:18:15go to chat I TP something it gives me
- 1:18:17five like detailed paragraphs of
- 1:18:19information sometimes I'm just annoyed
- 1:18:20by that that's not what I wanted but
- 1:18:22maybe in the original reward function in
- 1:18:24the original people actually pref that
- 1:18:25and nower it less yeah um that's a great
- 1:18:29point because like as these models like
- 1:18:31integrate more and more into our system
- 1:18:33they're going to collect more and more
- 1:18:34data and they will like pick up on
- 1:18:37things maybe undesirable things as well
- 1:18:40um as far as I understand chbd is really
- 1:18:42cutting down on the verbosity which is
- 1:18:44like a huge issue that all of these
- 1:18:45models are trying to cut down on and
- 1:18:47they are dealing with that um part of
- 1:18:50the reason why that emerges is that when
- 1:18:51you collect preference data at scale
- 1:18:53people are not necessarily reading the
- 1:18:55answers the turkers might just simply
- 1:18:57choose the longer answer and that's a
- 1:18:59property that actually goes into these
- 1:19:00models so but hopefully like these
- 1:19:03things will improve over time as they
- 1:19:04get more feedb and yeah hallucinations
- 1:19:07is not a problem that is going to go
- 1:19:08away with RL and we talked a bit about
- 1:19:10reward hacking as well um biases from
- 1:19:14things and so on but hopefully like I
- 1:19:16mean what I want to conclude out is like
- 1:19:18we started with pre-trained
- 1:19:20models we we had these things which
- 1:19:22could predict text and we got chargy GPD
- 1:19:25and hopefully like it's a little more
- 1:19:26clear how we go from something like that
- 1:19:28to chat
- 1:19:30GPD and that's I'll end
- 1:19:34here thanks
About this transcript
This page contains the full transcript of YouTube transcript (35X6zlhoCy4) , generated from the public captions YouTube serves with the video. The transcript has 14,287 words across 2,082 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.