State of GPT | BRK216HFS — Transcript
Full transcript
- 0:00[MUSIC]
- 0:07ANNOUNCER: Please welcome
- 0:08AI researcher and
- 0:10founding member of OpenAI, Andrej Karpathy.
- 0:21ANDREJ KARPATHY: Hi, everyone. I'm happy
- 0:24to be here to tell you about the state of
- 0:26GPT and more generally about
- 0:28the rapidly growing ecosystem of large language models.
- 0:31I would like to partition the talk into two parts.
- 0:35In the first part, I would like to tell you about
- 0:36how we train GPT Assistance,
- 0:39and then in the second part,
- 0:40we're going to take a look at how we can use
- 0:42these assistants effectively for your applications.
- 0:46First, let's take a look at the emerging
- 0:48recipe for how to train
- 0:49these assistants and keep in mind that this is all
- 0:51very new and still rapidly evolving,
- 0:53but so far, the recipe looks something like this.
- 0:55Now, this is a complicated slide,
- 0:57I'm going to go through it piece by
- 0:59piece, but roughly speaking,
- 1:01we have four major stages, pretraining,
- 1:04supervised finetuning, reward modeling,
- 1:06reinforcement learning,
- 1:07and they follow each other serially.
- 1:09Now, in each stage,
- 1:11we have a dataset that powers that stage.
- 1:14We have an algorithm that for our purposes will be
- 1:17a objective and over for training the neural network,
- 1:22and then we have a resulting model,
- 1:23and then there are some notes on the bottom.
- 1:25The first stage we're going to start
- 1:27with as the pretraining stage.
- 1:28Now, this stage is special in this diagram,
- 1:31and this diagram is not to scale because
- 1:33this stage is where all of
- 1:34the computational work basically happens.
- 1:36This is 99 percent of the training
- 1:38compute time and also flops.
- 1:41This is where we are dealing with
- 1:44Internet scale datasets with thousands
- 1:46of GPUs in the supercomputer and
- 1:48also months of training potentially.
- 1:51The other three stages are finetuning
- 1:53stages that are much more along
- 1:54the lines of small few number of GPUs and hours or days.
- 1:59Let's take a look at the pretraining stage
- 2:00to achieve a base model.
- 2:03First, we are going to gather a large amount of data.
- 2:07Here's an example of what we call a
- 2:08data mixture that comes from
- 2:10this paper that was released by
- 2:13Meta where they released this LLaMA based model.
- 2:16Now, you can see roughly the datasets that
- 2:18enter into these collections.
- 2:20We have CommonCrawl, which is a web scrape, C4,
- 2:23which is also CommonCrawl,
- 2:25and then some high quality datasets as well.
- 2:27For example, GitHub, Wikipedia,
- 2:29Books, Archives, Stock Exchange and so on.
- 2:31These are all mixed up together,
- 2:32and then they are sampled
- 2:34according to some given proportions,
- 2:36and that forms the training set for the GPT.
- 2:40Now before we can actually train on this data,
- 2:43we need to go through one more preprocessing step,
- 2:45and that is tokenization.
- 2:46This is basically a translation of
- 2:48the raw text that we scrape from the Internet into
- 2:51sequences of integers because
- 2:53that's the native representation
- 2:55over which GPTs function.
- 2:57Now, this is a lossless translation
- 3:00between pieces of texts and tokens and integers,
- 3:03and there are a number of algorithms for the stage.
- 3:05Typically, for example, you could
- 3:07use something like byte pair encoding,
- 3:08which iteratively merges text chunks
- 3:11and groups them into tokens.
- 3:13Here, I'm showing some example chunks of these tokens,
- 3:16and then this is the raw integer sequence
- 3:18that will actually feed into a transformer.
- 3:21Now, here I'm showing
- 3:23two examples for hybrid parameters
- 3:26that govern this stage.
- 3:28GPT-4, we did not release
- 3:30too much information about how it was trained and so on,
- 3:32I'm using GPT-3s numbers,
- 3:33but GPT-3 is of course a little bit old
- 3:35by now, about three years ago.
- 3:37But LLaMA is a fairly recent model from Meta.
- 3:40These are roughly the orders of
- 3:42magnitude that we're dealing
- 3:43with when we're doing pretraining.
- 3:44The vocabulary size is usually a couple 10,000 tokens.
- 3:48The context length is usually something like 2,000,
- 3:504,000, or nowadays even 100,000,
- 3:53and this governs the maximum number of integers that
- 3:56the GPT will look at when it's trying to
- 3:58predict the next integer in a sequence.
- 4:01You can see that roughly the number of parameters say,
- 4:0465 billion for LLaMA.
- 4:06Now, even though LLaMA has only 65B parameters
- 4:08compared to GPP-3s 175 billion parameters,
- 4:11LLaMA is a significantly more powerful model,
- 4:13and intuitively, that's because
- 4:15the model is trained for significantly longer.
- 4:17In this case, 1.4 trillion tokens,
- 4:19instead of 300 billion tokens.
- 4:21You shouldn't judge the power of a model by
- 4:23the number of parameters that it contains.
- 4:26Below, I'm showing some tables of
- 4:28rough hyperparameters that typically
- 4:31go into specifying the transformer neural network,
- 4:34the number of heads,
- 4:34the dimension size, number of layers,
- 4:36and so on, and on the bottom
- 4:38I'm showing some training hyperparameters.
- 4:41For example, to train the 65B model,
- 4:44Meta used 2,000 GPUs,
- 4:46roughly 21 days of training and
- 4:48a roughly several million dollars.
- 4:52That's the rough orders of magnitude that you should have
- 4:54in mind for the pre-training stage.
- 4:57Now, when we're actually pre-training, what happens?
- 5:00Roughly speaking, we are going to take our tokens,
- 5:03and we're going to lay them out into data batches.
- 5:06We have these arrays
- 5:07that will feed into the transformer,
- 5:09and these arrays are B,
- 5:10the batch size and these are all independent examples
- 5:13stocked up in rows and B by T,
- 5:16T being the maximum context length.
- 5:17In my picture I only have 10 the context lengths,
- 5:20so this could be 2,000, 4,000, etc.
- 5:23These are extremely long rows.
- 5:24What we do is we take these documents,
- 5:26and we pack them into rows,
- 5:28and we delimit them with
- 5:29these special end of texts tokens,
- 5:31basically telling the transformer
- 5:32where a new document begins.
- 5:35Here, I have a few examples of documents and then
- 5:38I stretch them out into this input.
- 5:41Now, we're going to feed all of
- 5:43these numbers into transformer.
- 5:46Let me just focus on a single particular cell,
- 5:49but the same thing will happen at
- 5:50every cell in this diagram.
- 5:52Let's look at the green cell.
- 5:54The green cell is going to take
- 5:56a look at all of the tokens before it,
- 5:58so all of the tokens in yellow,
- 6:00and we're going to feed that entire context
- 6:03into the transforming neural network,
- 6:05and the transformer is going to try to
- 6:07predict the next token in
- 6:08a sequence, in this case in red.
- 6:10Now the transformer, I don't have
- 6:11too much time to, unfortunately,
- 6:13go into the full details of this
- 6:14neural network architecture is
- 6:15just a large blob of neural net stuff for our purposes,
- 6:18and it's got several,
- 6:2010 billion parameters typically or something like that.
- 6:22Of course, as I tune these parameters,
- 6:23you're getting slightly
- 6:24different predicted distributions
- 6:26for every single one of these cells.
- 6:28For example, if our vocabulary size is 50,257 tokens,
- 6:34then we're going to have that many
- 6:35numbers because we need to
- 6:36specify a probability
- 6:38distribution for what comes next.
- 6:40Basically, we have a probability for
- 6:41whatever may follow.
- 6:42Now, in this specific example,
- 6:44for this specific cell,
- 6:45513 will come next,
- 6:47and so we can use this as
- 6:48a source of supervision to
- 6:49update our transformers weights.
- 6:51We're applying this basically
- 6:53on every single cell in the parallel,
- 6:54and we keep swapping batches,
- 6:56and we're trying to get the transformer to make
- 6:58the correct predictions over what
- 6:59token comes next in a sequence.
- 7:02Let me show you more concretely what this looks
- 7:03like when you train one of these models.
- 7:05This is actually coming from the New York Times,
- 7:07and they trained a small GPT on Shakespeare.
- 7:11Here's a small snippet of Shakespeare,
- 7:12and they train their GPT on it.
- 7:14Now, in the beginning,
- 7:15at initialization,
- 7:17the GPT starts with completely random weights.
- 7:19You're getting completely random outputs as well.
- 7:21But over time, as you train the GPT longer and longer,
- 7:26you are getting more and more coherent and
- 7:28consistent samples from the model,
- 7:31and the way you sample from it, of course,
- 7:32is you predict what comes next,
- 7:35you sample from that distribution and
- 7:36you keep feeding that back into the process,
- 7:38and you can basically sample large sequences.
- 7:42By the end, you see that the transformer
- 7:43has learned about words and
- 7:45where to put spaces and where to put commas and so on.
- 7:48We're making
- 7:48more and more consistent predictions over time.
- 7:51These are the plots that you are looking at
- 7:53when you're doing model pretraining.
- 7:54Effectively, we're looking at
- 7:56the loss function over time as you train,
- 7:58and low loss means that our transformer
- 8:00is giving a higher probability
- 8:03to the next correct integer in the sequence.
- 8:06What are we going to do with model
- 8:08once we've trained it after a month?
- 8:10Well, the first thing that we noticed, we the field,
- 8:14is that these models
- 8:16basically in the process of language modeling,
- 8:18learn very powerful general representations,
- 8:21and it's possible to very efficiently fine tune them
- 8:23for any arbitrary downstream tasks
- 8:25you might be interested in.
- 8:26As an example, if you're
- 8:27interested in sentiment classification,
- 8:29the approach used to be
- 8:31that you collect a bunch of positives
- 8:33and negatives and then you train some NLP model
- 8:35for that, but the new approach is:
- 8:38ignore sentiment
- 8:38classification, go off and do large
- 8:41language model pretraining,
- 8:43train a large transformer,
- 8:44and then you may only have a few examples and
- 8:47you can very efficiently fine tune
- 8:48your model for that task.
- 8:51This works very well in practice.
- 8:53The reason for this is that basically
- 8:55the transformer is forced to
- 8:56multitask a huge amount of
- 8:58tasks in the language modeling task,
- 9:00because in terms of predicting the next token,
- 9:03it's forced to understand a lot about the structure of
- 9:05the text and all the different concepts therein.
- 9:09That was GPT-1. Now around the time of GPT-2,
- 9:12people noticed that actually
- 9:14even better than fine tuning,
- 9:15you can actually prompt these models very effectively.
- 9:17These are language models and they want
- 9:19to complete documents,
- 9:20you can actually trick them into performing
- 9:22tasks by arranging these fake documents.
- 9:25In this example, for example,
- 9:27we have some passage and then we like do QA, QA, QA.
- 9:31This is called Few-shot prompt, and then we do Q,
- 9:33and then as the transformer is tried to
- 9:35complete the document is actually answering our question.
- 9:37This is an example of prompt engineering based model,
- 9:40making it believe that it's imitating
- 9:42a document and getting it to perform a task.
- 9:45This kicked off, I think the era of, I would say,
- 9:48prompting over fine tuning and seeing that this
- 9:50actually can work extremely well on a lot of problems,
- 9:53even without training any neural networks,
- 9:55fine tuning or so on.
- 9:56Now since then, we've seen
- 9:58an entire evolutionary tree of
- 10:00base models that everyone has trained.
- 10:02Not all of these models are available.
- 10:05for example, the GPT-4 base model was never released.
- 10:08The GPT-4 model that you might be
- 10:09interacting with over API is not a base model,
- 10:12it's an assistant model,
- 10:13and we're going to cover how to get those in a bit.
- 10:15GPT-3 based model is available via the API under
- 10:19the name Devanshi and GPT-2 based model
- 10:21is available even as weights on our GitHub repo.
- 10:24But currently the best available base model
- 10:27probably is the LLaMA series from Meta,
- 10:29although it is not commercially licensed.
- 10:32Now, one thing
- 10:34to point out is
- 10:35base models are not assistants.
- 10:36They don't want to make answers to your questions,
- 10:41they want to complete documents.
- 10:43If you tell them to write
- 10:44a poem about the bread and cheese,
- 10:46it will answer questions with more questions,
- 10:49it's completing what it thinks is a document.
- 10:51However, you can prompt them in a specific way for
- 10:54base models that is more likely to work.
- 10:57As an example, here's a poem about bread and cheese,
- 10:59and in that case it will autocomplete correctly.
- 11:02You can even trick base models into being assistants.
- 11:06The way you would do this is you would create
- 11:08a specific few-shot prompt
- 11:09that makes it look like there's
- 11:11some document between the human and assistant
- 11:13and they're exchanging information.
- 11:16Then at the bottom,
- 11:17you put your query at the end and the base model
- 11:21will condition itself into
- 11:23being a helpful assistant and answer,
- 11:26but this is not very reliable and doesn't work
- 11:28super well in practice, although it can be done.
- 11:30Instead, we have a different path to make
- 11:32actual GPT assistants not base model document completers.
- 11:37That takes us into supervised finetuning.
- 11:39In the supervised finetuning stage,
- 11:41we are going to collect
- 11:43small but high quality data-sets, and in this case,
- 11:45we're going to ask human contractors to gather data of
- 11:48the form prompt and ideal response.
- 11:52We're going to collect lots of these
- 11:54typically tens of thousands or something like that.
- 11:56Then we're going to still do language
- 11:58modeling on this data.
- 11:59Nothing changed algorithmically,
- 12:01we're swapping out a training set.
- 12:02It used to be Internet documents,
- 12:04which has a high quantity local
- 12:06for basically Q8 prompt response data.
- 12:11That is low quantity, high quality.
- 12:13We will still do language modeling
- 12:15and then after training,
- 12:16we get an SFT model.
- 12:18You can actually deploy these models and
- 12:20they are actual assistants and they work to some extent.
- 12:22Let me show you what an
- 12:24example demonstration might look like.
- 12:25Here's something that a human contractor
- 12:27might come up with.
- 12:28Here's some random prompt. Can you
- 12:29write a short introduction
- 12:31about the relevance of
- 12:32the term monopsony or something like that?
- 12:34Then the contractor also writes out an ideal response.
- 12:37When they write out these responses,
- 12:38they are following extensive labeling
- 12:40documentations and they are being
- 12:42asked to be helpful, truthful, and harmless.
- 12:45These labeling instructions here,
- 12:48you probably can't read it, neither can I,
- 12:50but they're long and this is people
- 12:52following instructions and trying
- 12:53to complete these prompts.
- 12:55That's what the dataset looks like.
- 12:57You can train these models. This works to some extent.
- 12:59Now, you can actually continue the pipeline from
- 13:02here on, and go into RLHF,
- 13:05reinforcement learning from human feedback that
- 13:07consists of both reward modeling
- 13:09and reinforcement learning.
- 13:10Let me cover that and then I'll
- 13:11come back to why you may want to go through
- 13:13the extra steps and how that compares to SFT models.
- 13:16In the reward modeling step,
- 13:18what we're going to do is we're now going to shift
- 13:20our data collection to be of the form of comparisons.
- 13:23Here's an example of what our dataset will look like.
- 13:25I have the same identical prompt on the top,
- 13:28which is asking the assistant to write
- 13:31a program or a function that
- 13:32checks if a given string is a palindrome.
- 13:35Then what we do is we take the SFT model which
- 13:38we've already trained and we create multiple completions.
- 13:41In this case, we have three completions
- 13:42that the model has created,
- 13:43and then we ask people to rank these completions.
- 13:47If you stare at this for a while, and by the way,
- 13:49these are very difficult things to do to
- 13:51compare some of these predictions.
- 13:52This can take people even hours for
- 13:54a single prompt completion pairs,
- 13:57but let's say we decided that one of these is
- 14:00much better than the others and so on. We rank them.
- 14:03Then we can follow that with
- 14:04something that looks very much like
- 14:06a binary classification on
- 14:07all the possible pairs between these completions.
- 14:10What we do now is, we lay out our prompt in rows,
- 14:13and the prompt is identical across all three rows here.
- 14:16It's all the same prompt, but
- 14:17the completion of this varies.
- 14:19The yellow tokens are coming from the SFT model.
- 14:21Then what we do is we append
- 14:23another special reward readout token at
- 14:26the end and we basically only
- 14:28supervise the transformer at this single green token.
- 14:31The transformer will predict some reward
- 14:34for how good that completion is
- 14:36for that prompt and basically it makes
- 14:39a guess about the quality of each completion.
- 14:42Then once it makes a guess for every one of them,
- 14:44we also have the ground truth
- 14:46which is telling us the ranking of them.
- 14:47We can actually enforce that some of
- 14:50these numbers should be much higher
- 14:51than others, and so on.
- 14:52We formulate this into a loss function and we
- 14:54train our model to make reward predictions
- 14:56that are consistent with the ground truth coming
- 14:58from the comparisons from all these contractors.
- 15:01That's how we train our reward model.
- 15:02That allows us to score how good
- 15:04a completion is for a prompt.
- 15:06Once we have a reward model,
- 15:09we can't deploy this because this is
- 15:11not very useful as an assistant by itself,
- 15:13but it's very useful for the reinforcement
- 15:15learning stage that follows now.
- 15:16Because we have a reward model,
- 15:18we can score the quality of
- 15:19any arbitrary completion for any given prompt.
- 15:22What we do during reinforcement learning
- 15:24is we basically get, again,
- 15:26a large collection of prompts and now we do
- 15:28reinforcement learning with respect to
- 15:29the reward model. Here's what that looks like.
- 15:32We take a single prompt,
- 15:34we lay it out in rows,
- 15:36and now we use basically
- 15:38the model we'd like to train which
- 15:39was initialized at SFT model
- 15:41to create some completions in yellow,
- 15:43and then we append the reward token again
- 15:45and we read off the reward
- 15:47according to the reward model,
- 15:49which is now kept fixed.
- 15:50It doesn't change any more. Now the reward model
- 15:53tells us the quality of every single completion
- 15:55for all these prompts and so what we can do is we can now
- 15:58just basically apply the same
- 15:59language modeling loss function,
- 16:01but we're currently training on the yellow tokens,
- 16:04and we are weighing
- 16:06the language modeling objective
- 16:08by the rewards indicated by the reward model.
- 16:11As an example, in the first row,
- 16:13the reward model said that this is
- 16:15a fairly high-scoring completion
- 16:17and so all the tokens that we
- 16:18happen to sample on the first row are going to get
- 16:21reinforced and they're going to get
- 16:22higher probabilities for the future.
- 16:25Conversely, on the second row,
- 16:26the reward model really did not like
- 16:28this completion, -1.2.
- 16:29Therefore, every single token that we sampled in
- 16:32that second row is going to get
- 16:34a slightly higher probability for the future.
- 16:36We do this over and over on
- 16:37many prompts on many batches and basically,
- 16:39we get a policy that creates yellow tokens here.
- 16:43It's basically all the completions here will
- 16:46score high according to
- 16:47the reward model that we trained in the previous stage.
- 16:51That's what the RLHF pipeline is.
- 16:55Then at the end, you get a model that you could deploy.
- 16:58As an example, ChatGPT is an RLHF model,
- 17:02but some other models that you might
- 17:03come across for example,
- 17:05Vicuna-13B, and so on,
- 17:06these are SFT models.
- 17:08We have base models, SFT models, and RLHF models.
- 17:12That's the state of things there.
- 17:14Now why would you want to do RLHF?
- 17:16One answer that's not
- 17:19that exciting is that it works better.
- 17:20This comes from the instruct GPT paper.
- 17:22According to these experiments a while ago now,
- 17:25these PPO models are RLHF.
- 17:28We see that they are basically preferred in a lot
- 17:30of comparisons when we give them to humans.
- 17:33Humans prefer basically tokens
- 17:36that come from RLHF models compared to SFT models,
- 17:39compared to base model that is prompted to be
- 17:41an assistant. It just works better.
- 17:43But you might ask why does it work better?
- 17:47I don't think that there's a single amazing answer
- 17:49that the community has really agreed on,
- 17:51but I will offer one reason potentially.
- 17:55It has to do with the asymmetry between how easy
- 17:58computationally it is to compare versus generate.
- 18:02Let's take an example of generating a haiku.
- 18:04Suppose I ask a model to write a haiku about paper clips.
- 18:07If you're a contractor trying to train data,
- 18:10then imagine being a contractor
- 18:12collecting basically data for the SFT stage,
- 18:14how are you supposed to create
- 18:15a nice haiku for a paper clip?
- 18:16You might not be very good at that,
- 18:18but if I give you a few examples of
- 18:20haikus you might be able to
- 18:21appreciate some of these haikus a lot more than others.
- 18:24Judging which one of these is good is a much easier task.
- 18:27Basically, this asymmetry
- 18:29makes it so that comparisons are
- 18:31a better way to potentially leverage
- 18:33yourself as a human and
- 18:34your judgment to create a slightly better model.
- 18:37Now, RLHF models are not
- 18:40strictly an improvement on the base models in some cases.
- 18:43In particular, we'd notice for example
- 18:45that they lose some entropy.
- 18:46That means that they give more peaky results.
- 18:49They can output samples
- 18:54with lower variation than the base model.
- 18:55The base model has lots of entropy and
- 18:57will give lots of diverse outputs.
- 19:00For example, one place where I still
- 19:03prefer to use a base model is in the setup
- 19:06where you basically have
- 19:09n things and you want to generate more things like it.
- 19:13Here is an example that I just cooked up.
- 19:16I want to generate cool Pokemon names.
- 19:18I gave it seven Pokemon names and I asked the base model
- 19:21to complete the document and it
- 19:22gave me a lot more Pokemon names.
- 19:24These are fictitious. I tried to look them up.
- 19:27I don't believe they're actual Pokemons.
- 19:29This is the task that I think the base model would be
- 19:31good at because it still has lots of entropy.
- 19:33It'll give you lots of diverse cool
- 19:35more things that look like whatever you give it before.
- 19:41Having said all that, these are
- 19:43the assistant models that are probably
- 19:45available to you at this point.
- 19:47There was a team at Berkeley that ranked a lot of
- 19:50the available assistant models
- 19:51and give them basically Elo ratings.
- 19:53Currently, some of the best models,
- 19:54of course, are GPT-4,
- 19:55by far, I would say,
- 19:57followed by Claude,
- 19:58GPT-3.5, and then a number of models,
- 20:00some of these might be available as weights,
- 20:02like Vicuna, Koala, etc.
- 20:04The first three rows here are
- 20:07all RLHF models and
- 20:09all of the other models to my knowledge,
- 20:11are SFT models, I believe.
- 20:15That's how we train
- 20:17these models on the high level.
- 20:19Now I'm going to switch gears
- 20:20and let's look at how we can
- 20:22best apply the GPT assistant model to your problems.
- 20:26Now, I would like to work
- 20:27in setting of a concrete example.
- 20:29Let's work with a concrete example here.
- 20:32Let's say that you are working on
- 20:34an article or a blog post,
- 20:35and you're going to write this sentence at the end.
- 20:38"California's population is 53
- 20:40times that of Alaska." So for some reason,
- 20:42you want to compare the populations of these two states.
- 20:44Think about the rich internal monologue
- 20:47and tool use and how much work
- 20:49actually goes computationally in
- 20:50your brain to generate this one final sentence.
- 20:53Here's maybe what that could look like in your brain.
- 20:55For this next step, let me blog on my blog,
- 20:59let me compare these two populations.
- 21:01First I'm going to obviously need to
- 21:03get both of these populations.
- 21:05Now, I know that I probably
- 21:06don't know these populations off the top of
- 21:08my head so I'm aware
- 21:10of what I know or don't know of my self-knowledge.
- 21:12I go, I do some tool use and I go to Wikipedia and I
- 21:16look up California's population and Alaska's population.
- 21:19Now, I know that I should divide the two, but again,
- 21:22I know that dividing 39.2 by
- 21:240.74 is very unlikely to succeed.
- 21:26That's not the thing that I can
- 21:28do in my head and so therefore,
- 21:30I'm going to rely on
- 21:31the calculator so I'm going to use a calculator,
- 21:33punch it in and see that the output is roughly 53.
- 21:36Then maybe I do
- 21:38some reflection and sanity checks in
- 21:40my brain so does 53 makes sense?
- 21:42Well, that's quite a large fraction,
- 21:44but then California is the most
- 21:45populous state, so maybe that looks okay.
- 21:47Then I have all the information I might need,
- 21:50and now I get to the creative portion of writing.
- 21:52I might start to write something like "California has
- 21:5553x times greater" and then I think to myself,
- 21:58that's actually like really awkward phrasing so let me
- 22:00actually delete that and let me try again.
- 22:03As I'm writing, I have this separate process,
- 22:06almost inspecting what I'm
- 22:07writing and judging whether it looks good
- 22:09or not and then maybe I delete and maybe I reframe it,
- 22:13and then maybe I'm happy with what comes out.
- 22:15Basically long story short,
- 22:17a ton happens under the hood in terms of
- 22:19your internal monologue when you
- 22:20create sentences like this.
- 22:21But what does a sentence like this look like
- 22:24when we are training a GPT on it?
- 22:26From GPT's perspective, this
- 22:28is just a sequence of tokens.
- 22:30GPT, when it's reading or generating these tokens,
- 22:34it just goes chunk, chunk, chunk,
- 22:35chunk and each chunk is roughly
- 22:37the same amount of computational work for each token.
- 22:40These transformers are not
- 22:42very shallow networks they have
- 22:43about 80 layers of reasoning,
- 22:45but 80 is still not like too much.
- 22:47This transformer is going to do its best to imitate,
- 22:51but of course, the process here
- 22:53looks very different from the process that you took.
- 22:56In particular, in our final artifacts
- 22:59in the data sets that we create,
- 23:00and then eventually feed to
- 23:01LLMs, all that internal dialogue was completely
- 23:03stripped and unlike you,
- 23:07the GPT will look at every single token and
- 23:09spend the same amount of compute on every one of them.
- 23:12So, you can't expect it
- 23:13to do too much work per token and also in particular,
- 23:21basically these transformers are
- 23:22just like token simulators,
- 23:23they don't know what they don't know.
- 23:26They just imitate the next token.
- 23:27They don't know what they're good at or not good at.
- 23:29They just tried their best to imitate the next token.
- 23:32They don't reflect in the loop.
- 23:34They don't sanity check anything.
- 23:35They don't correct their mistakes along the way.
- 23:37By default, they just are sample token sequences.
- 23:40They don't have separate inner monologue streams
- 23:43in their head right? They're
- 23:43evaluating what's happening.
- 23:45Now, they do have some cognitive advantages,
- 23:48I would say and that is that they do actually have
- 23:51a very large fact-based knowledge
- 23:52across a vast number of areas because they have,
- 23:55say, several, 10 billion parameters.
- 23:57That's a lot of storage for a lot of facts.
- 23:59They also, I think have
- 24:02a relatively large and perfect working memory.
- 24:04Whatever fits into the context window
- 24:07is immediately available to
- 24:09the transformer through
- 24:10its internal self attention mechanism
- 24:12and so it's perfect memory,
- 24:14but it's got a finite size,
- 24:16but the transformer has a very direct access
- 24:18to it and so it can a
- 24:19losslessly remember anything that
- 24:22is inside its context window.
- 24:23This is how I would compare those two and the reason I
- 24:26bring all of this up is because I
- 24:27think to a large extent,
- 24:29prompting is just making up for
- 24:31this cognitive difference between
- 24:34these two architectures like
- 24:37our brains here and LLM brains.
- 24:39You can look at it that way almost.
- 24:41Here's one thing that people found for example
- 24:44works pretty well in practice.
- 24:45Especially if your tasks require reasoning,
- 24:48you can't expect the transformer
- 24:49to do too much reasoning per token.
- 24:52You have to really spread out
- 24:53the reasoning across more and more tokens.
- 24:56For example, you can't give a transformer
- 24:57a very complicated question and
- 24:59expect it to get the answer in a single token.
- 25:00There's just not enough time for it.
- 25:02"These transformers need tokens to
- 25:04think," I like to say sometimes.
- 25:06This is some of the things that work well,
- 25:08you may for example have a few-shot prompt that
- 25:10shows the transformer that it should show
- 25:12its work when it's answering
- 25:14question and if you give a few examples,
- 25:17the transformer will imitate that template and it
- 25:20will just end up working out
- 25:21better in terms of its evaluation.
- 25:24Additionally, you can elicit this behavior from
- 25:26the transformer by saying, let things step-by-step.
- 25:29Because this conditions the transformer into showing
- 25:32its work and because
- 25:34it snaps into a mode of showing its work,
- 25:36is going to do less computational work per token.
- 25:40It's more likely to succeed as a result because it's
- 25:42making slower reasoning over time.
- 25:46Here's another example, this one
- 25:47is called self-consistency.
- 25:49We saw that we had the ability
- 25:51to start writing and then if it didn't work out,
- 25:54I can try again and I can try multiple times
- 25:56and maybe select the one that worked best.
- 26:00In these approaches,
- 26:02you may sample not just once,
- 26:03but you may sample multiple times and
- 26:05then have some process for finding
- 26:07the ones that are good and then keeping
- 26:09just those samples or doing
- 26:10a majority vote or something like that.
- 26:11Basically these transformers in the process as
- 26:14they predict the next token, just like you,
- 26:16they can get unlucky
- 26:18and they could sample a not a very good
- 26:19token and they can go down like
- 26:21a blind alley in terms of reasoning.
- 26:24Unlike you, they cannot recover from that.
- 26:27They are stuck with every single token they
- 26:28sample and so they will continue the sequence,
- 26:31even if they know
- 26:32that this sequence is not going to work out.
- 26:34Give them the ability to look back,
- 26:36inspect or try to basically sample around it.
- 26:40Here's one technique also,
- 26:43it turns out that actually LLMs,
- 26:45they know when they've screwed up,
- 26:47so as an example, say you ask the model
- 26:50to generate a poem that does not
- 26:52rhyme and it might give you a poem,
- 26:54but it actually rhymes.
- 26:55But it turns out that especially for
- 26:57the bigger models like GPT-4,
- 26:58you can just ask it "did you meet the assignment?"
- 27:01Actually GPT-4 knows very
- 27:03well that it did not meet the assignment.
- 27:04It just got unlucky in its sampling.
- 27:07It will tell you, "No, I didn't actually meet
- 27:08the assignment here. Let me try again."
- 27:10But without you prompting
- 27:12it it doesn't know to revisit and so on.
- 27:17You have to make up for that in your prompts,
- 27:19and you have to get it to check,
- 27:21if you don't ask it to check,
- 27:23its not going to check by itself
- 27:24it's just a token simulator.
- 27:28I think more generally,
- 27:29a lot of these techniques fall into
- 27:31the bucket of what I would say recreating our System 2.
- 27:34You might be familiar with the System 1 and
- 27:36System 2 thinking for humans.
- 27:37System 1 is a fast automatic process and I
- 27:40think corresponds to an LLM just sampling tokens.
- 27:43System 2 is the slower deliberate
- 27:46planning part of your brain.
- 27:49This is a paper actually from
- 27:51just last week because
- 27:52this space is pretty quickly evolving,
- 27:53it's called Tree of Thought.
- 27:56The authors of this paper proposed maintaining
- 27:59multiple completions for any given prompt
- 28:02and then they are also scoring them along
- 28:04the way and keeping the ones that
- 28:06are going well if that makes sense.
- 28:08A lot of people are really playing
- 28:10around with prompt engineering
- 28:13to basically bring back some of
- 28:15these abilities that we have in our brain for LLMs.
- 28:19Now, one thing I would like to note
- 28:21here is that this is not just a prompt.
- 28:22This is actually prompts that are together
- 28:25used with some Python Glue code because you
- 28:28actually have to maintain multiple
- 28:29prompts and you also have to do
- 28:30some tree search algorithm here
- 28:32to figure out which prompts to expand, etc.
- 28:35It's a symbiosis of Python Glue code and
- 28:38individual prompts that are
- 28:39called in a while loop or in a bigger algorithm.
- 28:42I also think there's a really cool
- 28:43parallel here to AlphaGo.
- 28:44AlphaGo has a policy for
- 28:46placing the next stone when it plays go,
- 28:48and its policy was trained
- 28:50originally by imitating humans.
- 28:52But in addition to this policy,
- 28:54it also does Monte Carlo Tree Search.
- 28:56Basically, it will play out a number of possibilities in
- 28:59its head and evaluate all of
- 29:00them and only keep the ones that work well.
- 29:01I think this is an equivalent of
- 29:04AlphaGo but for text if that makes sense.
- 29:08Just like Tree of Thought,
- 29:10I think more generally people are
- 29:11starting to really explore
- 29:13more general techniques of not
- 29:15just the simple question-answer prompts,
- 29:17but something that looks a lot more like
- 29:19Python Glue code stringing together many prompts.
- 29:22On the right, I have an example from
- 29:23this paper called React where they
- 29:25structure the answer to a prompt
- 29:28as a sequence of thought-action-observation,
- 29:32thought-action-observation, and it's
- 29:34a full rollout and
- 29:35a thinking process to answer the query.
- 29:38In these actions, the model is also allowed to tool use.
- 29:42On the left, I have an example of AutoGPT.
- 29:45Now AutoGPT by the way is
- 29:47a project that I think got a lot of hype recently,
- 29:51but I think I still find it inspirationally interesting.
- 29:55It's a project that allows an LLM to keep
- 29:58the task list and continue to
- 30:00recursively break down tasks.
- 30:02I don't think this currently works very well and I would
- 30:04not advise people to use it in practical applications.
- 30:07I just think it's something to generally take inspiration
- 30:09from in terms of where this is going, I think over time.
- 30:12That's like giving our model System 2 thinking.
- 30:16The next thing I find interesting is,
- 30:19this following serve I would say
- 30:20almost psychological quirk of LLMs,
- 30:23is that LLMs don't want to succeed,
- 30:26they want to imitate.
- 30:28You want to succeed, and you should ask for it.
- 30:31What I mean by that is,
- 30:33when transformers are trained,
- 30:35they have training sets and there can be
- 30:38an entire spectrum of
- 30:39performance qualities in their training data.
- 30:41For example, there could be some kind of a prompt
- 30:43for some physics question or something like that,
- 30:45and there could be
- 30:45a student's solution that is completely wrong
- 30:47but there can also be an expert
- 30:49answer that is extremely right.
- 30:50Transformers can't tell the difference between low,
- 30:54they know about low-quality solutions
- 30:56and high-quality solutions,
- 30:57but by default, they want to imitate all of
- 30:59it because they're just trained on language modeling.
- 31:02At test time, you actually have
- 31:04to ask for a good performance.
- 31:06In this example in this paper,
- 31:08they tried various prompts.
- 31:10Let's think step-by-step was very powerful
- 31:13because it spread out the reasoning over many tokens.
- 31:15But what worked even better is,
- 31:17let's work this out in a step-by-step way
- 31:19to be sure we have the right answer.
- 31:20It's like conditioning on getting the right answer,
- 31:23and this actually makes the transformer work
- 31:25better because the transformer doesn't have
- 31:27to now hedge its probability mass
- 31:29on low-quality solutions,
- 31:31as ridiculous as that sounds.
- 31:33Basically, feel free to ask for a strong solution.
- 31:37Say something like, you are
- 31:38a leading expert on this topic.
- 31:39Pretend you have IQ 120, etc.
- 31:41But don't try to ask for too much IQ because if
- 31:44you ask for IQ 400,
- 31:46you might be out of data distribution,
- 31:48or even worse, you could be in data distribution for
- 31:51something like sci-fi stuff and it
- 31:52will start to take on some sci-fi,
- 31:54or like roleplaying or something like that.
- 31:56You have to find the right amount of IQ.
- 31:59I think it's got some U-shaped curve there.
- 32:02Next up, as we
- 32:04saw when we are trying to solve problems,
- 32:07we know what we are good at and what we're not good at,
- 32:09and we lean on tools computationally.
- 32:12You want to do the same potentially with your LLMs.
- 32:15In particular, we may want to give
- 32:18them calculators, code interpreters,
- 32:21and so on, the ability to do search,
- 32:23and there's a lot of techniques for doing that.
- 32:27One thing to keep in mind, again,
- 32:28is that these transformers by default may
- 32:30not know what they don't know.
- 32:32You may even want to tell the transformer in
- 32:34a prompt you are not very good at mental arithmetic.
- 32:37Whenever you need to do very large number addition,
- 32:40multiplication, or whatever,
- 32:41instead, use this calculator.
- 32:42Here's how you use the calculator,
- 32:43you use this token combination, etc.
- 32:46You have to actually spell it out because the model by
- 32:48default doesn't know what it's good at or not good at,
- 32:50necessarily, just like you and I might be.
- 32:54Next up, I think something that is very
- 32:56interesting is we went from
- 32:58a world that was retrieval only all the way,
- 33:02the pendulum has swung to the other extreme
- 33:03where its memory only in LLMs.
- 33:06But actually, there's this entire space in-between of
- 33:08these retrieval-augmented models and
- 33:10this works extremely well in practice.
- 33:12As I mentioned, the context window of
- 33:14a transformer is its working memory.
- 33:17If you can load the working memory
- 33:18with any information that is relevant to the task,
- 33:21the model will work extremely well
- 33:23because it can immediately access all that memory.
- 33:26I think a lot of people are really interested
- 33:28in basically retrieval-augment degeneration.
- 33:32On the bottom, I have an example of LlamaIndex which is
- 33:35one data connector to lots of different types of data.
- 33:38You can index all
- 33:41of that data and you can make it accessible to LLMs.
- 33:44The emerging recipe there is you take relevant documents,
- 33:47you split them up into chunks,
- 33:49you embed all of them,
- 33:50and you basically get embedding vectors
- 33:52that represent that data.
- 33:53You store that in the vector store and then at test time,
- 33:56you make some kind of a query to
- 33:57your vector store and you fetch chunks that
- 34:00might be relevant to your task and
- 34:01you stuff them into the prompt and then you generate.
- 34:04This can work quite well in practice.
- 34:06This is, I think, similar to
- 34:07when you and I solve problems.
- 34:09You can do everything from your memory and
- 34:11transformers have very large and extensive memory,
- 34:13but also it really helps to
- 34:14reference some primary documents.
- 34:17Whenever you find yourself going
- 34:19back to a textbook to find something,
- 34:21or whenever you find yourself going back to
- 34:22documentation of the library to look something up,
- 34:25transformers definitely want to do that too.
- 34:27You have some memory over how
- 34:30some documentation of the library
- 34:31works but it's much better to look it up.
- 34:33The same applies here.
- 34:35Next, I wanted to briefly talk
- 34:38about constraint prompting.
- 34:39I also find this very interesting.
- 34:41This is basically techniques
- 34:43for forcing a certain template in the outputs of LLMs.
- 34:50Guidance is one example from Microsoft actually.
- 34:53Here we are enforcing that
- 34:55the output from the LLM will be JSON.
- 34:57This will actually guarantee that
- 35:00the output will take on this form because they go
- 35:02in and they mess with the probabilities of
- 35:03all the different tokens that
- 35:04come out of the transformer and
- 35:05they clamp those tokens and then
- 35:07the transformer is only filling in the blanks here,
- 35:09and then you can enforce additional restrictions
- 35:11on what could go into those blanks.
- 35:13This might be really helpful, and I think
- 35:15this constraint sampling is also extremely interesting.
- 35:19I also want to say
- 35:20a few words about fine tuning.
- 35:22It is the case that you can get really
- 35:23far with prompt engineering,
- 35:25but it's also possible to
- 35:27think about fine tuning your models.
- 35:29Now, fine tuning models means that you
- 35:31are actually going to change the weights of the model.
- 35:33It is becoming a lot more
- 35:35accessible to do this in practice,
- 35:37and that's because of
- 35:38a number of techniques that have been
- 35:39developed and have libraries for very recently.
- 35:43So for example parameter efficient
- 35:44fine tuning techniques like Laura,
- 35:46make sure that you're only training small,
- 35:49sparse pieces of your model.
- 35:51So most of the model is kept clamped at
- 35:53the base model and some pieces of it are allowed to
- 35:55change and this still works pretty
- 35:56well empirically and makes
- 35:58it much cheaper to tune only small pieces of your model.
- 36:02It also means that because most of your model is clamped,
- 36:05you can use very low precision inference
- 36:07for computing those parts because
- 36:09you are not going to be updated by
- 36:10gradient descent and so that
- 36:12makes everything a lot more efficient as well.
- 36:13And in addition, we have a number of
- 36:15open source, high-quality base models.
- 36:17Currently, as I mentioned,
- 36:18I think LLaMa is quite nice,
- 36:20although it is not commercially
- 36:21licensed, I believe right now.
- 36:23Some things to keep in mind is that basically
- 36:26fine tuning is a lot more technically involved.
- 36:29It requires a lot more, I think,
- 36:30technical expertise to do right.
- 36:32It requires human data contractors for
- 36:34datasets and/or synthetic data pipelines
- 36:36that can be pretty complicated.
- 36:38This will definitely slow down
- 36:40your iteration cycle by a lot,
- 36:41and I would say on a high level SFT is
- 36:44achievable because you're continuing
- 36:47the language modeling task.
- 36:48It's relatively straightforward, but RLHF,
- 36:50I would say is very much research territory
- 36:53and is even much harder to get to work,
- 36:55and so I would probably not advise that someone
- 36:58just tries to roll their own RLHF of implementation.
- 37:00These things are pretty unstable,
- 37:02very difficult to train, not something that is, I think,
- 37:04very beginner friendly right now,
- 37:06and it's also potentially likely also
- 37:08to change pretty rapidly still.
- 37:11So I think these are
- 37:12my default recommendations right now.
- 37:15I would break up your task into two major parts.
- 37:18Number 1, achieve your top performance,
- 37:20and Number 2, optimize your performance in that order.
- 37:23Number 1, the best performance will
- 37:25currently come from GPT-4 model.
- 37:27It is the most capable of all by far.
- 37:29Use prompts that are very detailed.
- 37:31They have lots of task content,
- 37:33relevant information and instructions.
- 37:36Think along the lines of what would you tell
- 37:38a task contractor if they can't email you back,
- 37:40but then also keep in mind that a task contractor is a
- 37:43human and they have
- 37:44inner monologue and they're very clever, etc.
- 37:46LLMs do not possess those qualities.
- 37:48So make sure to think through
- 37:50the psychology of the LLM
- 37:52almost and cater prompts to that.
- 37:54Retrieve and add any relevant context
- 37:57and information to these prompts.
- 37:59Basically refer to a lot of
- 38:01the prompt engineering techniques.
- 38:02Some of them I've highlighted in the slides above,
- 38:04but also this is a very large space and I would
- 38:07just advise you to look
- 38:09for prompt engineering techniques online.
- 38:11There's a lot to cover there.
- 38:13Experiment with few-shot examples.
- 38:15What this refers to is, you don't just want to tell,
- 38:17you want to show whenever it's possible.
- 38:19So give it examples of everything
- 38:21that helps it really understand what you mean if you can.
- 38:25Experiment with tools and plug-ins to
- 38:27offload tasks that are difficult for LLMs natively,
- 38:30and then think about not just a
- 38:32single prompt and answer,
- 38:33think about potential chains
- 38:34and reflection and how you glue
- 38:36them together and how you can
- 38:37potentially make multiple samples and so on.
- 38:40Finally, if you think you've squeezed
- 38:42out prompt engineering,
- 38:43which I think you should stick with for a while,
- 38:45look at some potentially
- 38:48fine tuning a model to your application,
- 38:51but expect this to be a lot more
- 38:52slower in the vault and then
- 38:54there's an expert fragile research zone
- 38:56here and I would say that is RLHF,
- 38:58which currently does work a bit
- 39:00better than SFT if you can get it to work.
- 39:02But again, this is pretty involved, I would say.
- 39:05And to optimize your costs,
- 39:06try to explore lower capacity models
- 39:09or shorter prompts and so on.
- 39:12I also wanted to say a few words about the use cases
- 39:15in which I think LLMs are currently well suited for.
- 39:18In particular, note that there's a large number
- 39:20of limitations to LLMs today,
- 39:22and so I would keep that
- 39:24definitely in mind for all of your applications.
- 39:26Models, and this by the way could be an entire talk.
- 39:28So I don't have time to cover it in full detail.
- 39:30Models may be biased, they may fabricate,
- 39:32hallucinate information,
- 39:33they may have reasoning errors,
- 39:35they may struggle in entire classes of applications,
- 39:38they have knowledge cut-offs,
- 39:40so they might not know any information above,
- 39:42say, September, 2021.
- 39:43They are susceptible to a large range of
- 39:45attacks which are coming out on Twitter daily,
- 39:48including prompt injection, jailbreak attacks,
- 39:51data poisoning attacks and so on.
- 39:52So my recommendation right now is
- 39:54use LLMs in low-stakes applications.
- 39:57Combine them always with human oversight.
- 40:00Use them as a source of inspiration and
- 40:01suggestions and think co-pilots,
- 40:04instead of completely autonomous agents
- 40:05that are just like performing a task somewhere.
- 40:07It's just not clear that the models are there right now.
- 40:11So I wanted to close by saying that
- 40:13GPT-4 is an amazing artifact.
- 40:15I'm very thankful that it exists, and it's beautiful.
- 40:18It has a ton of knowledge across so many areas.
- 40:20It can do math, code and so on.
- 40:22And in addition, there's this
- 40:24thriving ecosystem of everything
- 40:25else that is being built and
- 40:27incorporated into the ecosystem.
- 40:29Some of these things I've talked about,
- 40:31and all of this power is accessible at your fingertips.
- 40:35So here's everything that's needed in terms of
- 40:37code to ask GPT-4 a question,
- 40:40to prompt it, and get a response.
- 40:42In this case, I said,
- 40:44can you say something to inspire
- 40:45the audience of Microsoft Build 2023?
- 40:47And I just punched this into Python and verbatim
- 40:50GPT-4 said the following:
- 40:53And by the way, I did not know that they
- 40:55used this trick in the keynote.
- 40:57So I thought I was being clever,
- 40:59but it is really good at this.
- 41:02It says, ladies and gentlemen,
- 41:04innovators and trailblazers Microsoft Build 2023.
- 41:07Welcome to the gathering of brilliant
- 41:08minds like no other,
- 41:09you are the architects of the future,
- 41:11the visionaries molding the digital realm
- 41:13in which humanity thrives.
- 41:15Embrace the limitless possibilities of
- 41:17technologies and let your ideas
- 41:18soar as high as your imagination.
- 41:20Together, let's create a more connected,
- 41:22remarkable, and inclusive world for generations to come.
- 41:25Get ready to unleash your creativity,
- 41:27canvas the unknown, and turn dreams into reality.
- 41:30Your journey begins today!
About this transcript
This page contains the full transcript of State of GPT | BRK216HFS by Microsoft Developer, generated from the public captions YouTube serves with the video. The transcript has 7,684 words across 1,133 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.