YouTube2Text

State of GPT | BRK216HFS — Transcript

by Microsoft Developer · 7,684 words · 1,133 segments · language en · Watch on YouTube

Full transcript

  1. 0:00[MUSIC]
  2. 0:07ANNOUNCER: Please welcome
  3. 0:08AI researcher and
  4. 0:10founding member of OpenAI, Andrej Karpathy.
  5. 0:21ANDREJ KARPATHY: Hi, everyone. I'm happy
  6. 0:24to be here to tell you about the state of
  7. 0:26GPT and more generally about
  8. 0:28the rapidly growing ecosystem of large language models.
  9. 0:31I would like to partition the talk into two parts.
  10. 0:35In the first part, I would like to tell you about
  11. 0:36how we train GPT Assistance,
  12. 0:39and then in the second part,
  13. 0:40we're going to take a look at how we can use
  14. 0:42these assistants effectively for your applications.
  15. 0:46First, let's take a look at the emerging
  16. 0:48recipe for how to train
  17. 0:49these assistants and keep in mind that this is all
  18. 0:51very new and still rapidly evolving,
  19. 0:53but so far, the recipe looks something like this.
  20. 0:55Now, this is a complicated slide,
  21. 0:57I'm going to go through it piece by
  22. 0:59piece, but roughly speaking,
  23. 1:01we have four major stages, pretraining,
  24. 1:04supervised finetuning, reward modeling,
  25. 1:06reinforcement learning,
  26. 1:07and they follow each other serially.
  27. 1:09Now, in each stage,
  28. 1:11we have a dataset that powers that stage.
  29. 1:14We have an algorithm that for our purposes will be
  30. 1:17a objective and over for training the neural network,
  31. 1:22and then we have a resulting model,
  32. 1:23and then there are some notes on the bottom.
  33. 1:25The first stage we're going to start
  34. 1:27with as the pretraining stage.
  35. 1:28Now, this stage is special in this diagram,
  36. 1:31and this diagram is not to scale because
  37. 1:33this stage is where all of
  38. 1:34the computational work basically happens.
  39. 1:36This is 99 percent of the training
  40. 1:38compute time and also flops.
  41. 1:41This is where we are dealing with
  42. 1:44Internet scale datasets with thousands
  43. 1:46of GPUs in the supercomputer and
  44. 1:48also months of training potentially.
  45. 1:51The other three stages are finetuning
  46. 1:53stages that are much more along
  47. 1:54the lines of small few number of GPUs and hours or days.
  48. 1:59Let's take a look at the pretraining stage
  49. 2:00to achieve a base model.
  50. 2:03First, we are going to gather a large amount of data.
  51. 2:07Here's an example of what we call a
  52. 2:08data mixture that comes from
  53. 2:10this paper that was released by
  54. 2:13Meta where they released this LLaMA based model.
  55. 2:16Now, you can see roughly the datasets that
  56. 2:18enter into these collections.
  57. 2:20We have CommonCrawl, which is a web scrape, C4,
  58. 2:23which is also CommonCrawl,
  59. 2:25and then some high quality datasets as well.
  60. 2:27For example, GitHub, Wikipedia,
  61. 2:29Books, Archives, Stock Exchange and so on.
  62. 2:31These are all mixed up together,
  63. 2:32and then they are sampled
  64. 2:34according to some given proportions,
  65. 2:36and that forms the training set for the GPT.
  66. 2:40Now before we can actually train on this data,
  67. 2:43we need to go through one more preprocessing step,
  68. 2:45and that is tokenization.
  69. 2:46This is basically a translation of
  70. 2:48the raw text that we scrape from the Internet into
  71. 2:51sequences of integers because
  72. 2:53that's the native representation
  73. 2:55over which GPTs function.
  74. 2:57Now, this is a lossless translation
  75. 3:00between pieces of texts and tokens and integers,
  76. 3:03and there are a number of algorithms for the stage.
  77. 3:05Typically, for example, you could
  78. 3:07use something like byte pair encoding,
  79. 3:08which iteratively merges text chunks
  80. 3:11and groups them into tokens.
  81. 3:13Here, I'm showing some example chunks of these tokens,
  82. 3:16and then this is the raw integer sequence
  83. 3:18that will actually feed into a transformer.
  84. 3:21Now, here I'm showing
  85. 3:23two examples for hybrid parameters
  86. 3:26that govern this stage.
  87. 3:28GPT-4, we did not release
  88. 3:30too much information about how it was trained and so on,
  89. 3:32I'm using GPT-3s numbers,
  90. 3:33but GPT-3 is of course a little bit old
  91. 3:35by now, about three years ago.
  92. 3:37But LLaMA is a fairly recent model from Meta.
  93. 3:40These are roughly the orders of
  94. 3:42magnitude that we're dealing
  95. 3:43with when we're doing pretraining.
  96. 3:44The vocabulary size is usually a couple 10,000 tokens.
  97. 3:48The context length is usually something like 2,000,
  98. 3:504,000, or nowadays even 100,000,
  99. 3:53and this governs the maximum number of integers that
  100. 3:56the GPT will look at when it's trying to
  101. 3:58predict the next integer in a sequence.
  102. 4:01You can see that roughly the number of parameters say,
  103. 4:0465 billion for LLaMA.
  104. 4:06Now, even though LLaMA has only 65B parameters
  105. 4:08compared to GPP-3s 175 billion parameters,
  106. 4:11LLaMA is a significantly more powerful model,
  107. 4:13and intuitively, that's because
  108. 4:15the model is trained for significantly longer.
  109. 4:17In this case, 1.4 trillion tokens,
  110. 4:19instead of 300 billion tokens.
  111. 4:21You shouldn't judge the power of a model by
  112. 4:23the number of parameters that it contains.
  113. 4:26Below, I'm showing some tables of
  114. 4:28rough hyperparameters that typically
  115. 4:31go into specifying the transformer neural network,
  116. 4:34the number of heads,
  117. 4:34the dimension size, number of layers,
  118. 4:36and so on, and on the bottom
  119. 4:38I'm showing some training hyperparameters.
  120. 4:41For example, to train the 65B model,
  121. 4:44Meta used 2,000 GPUs,
  122. 4:46roughly 21 days of training and
  123. 4:48a roughly several million dollars.
  124. 4:52That's the rough orders of magnitude that you should have
  125. 4:54in mind for the pre-training stage.
  126. 4:57Now, when we're actually pre-training, what happens?
  127. 5:00Roughly speaking, we are going to take our tokens,
  128. 5:03and we're going to lay them out into data batches.
  129. 5:06We have these arrays
  130. 5:07that will feed into the transformer,
  131. 5:09and these arrays are B,
  132. 5:10the batch size and these are all independent examples
  133. 5:13stocked up in rows and B by T,
  134. 5:16T being the maximum context length.
  135. 5:17In my picture I only have 10 the context lengths,
  136. 5:20so this could be 2,000, 4,000, etc.
  137. 5:23These are extremely long rows.
  138. 5:24What we do is we take these documents,
  139. 5:26and we pack them into rows,
  140. 5:28and we delimit them with
  141. 5:29these special end of texts tokens,
  142. 5:31basically telling the transformer
  143. 5:32where a new document begins.
  144. 5:35Here, I have a few examples of documents and then
  145. 5:38I stretch them out into this input.
  146. 5:41Now, we're going to feed all of
  147. 5:43these numbers into transformer.
  148. 5:46Let me just focus on a single particular cell,
  149. 5:49but the same thing will happen at
  150. 5:50every cell in this diagram.
  151. 5:52Let's look at the green cell.
  152. 5:54The green cell is going to take
  153. 5:56a look at all of the tokens before it,
  154. 5:58so all of the tokens in yellow,
  155. 6:00and we're going to feed that entire context
  156. 6:03into the transforming neural network,
  157. 6:05and the transformer is going to try to
  158. 6:07predict the next token in
  159. 6:08a sequence, in this case in red.
  160. 6:10Now the transformer, I don't have
  161. 6:11too much time to, unfortunately,
  162. 6:13go into the full details of this
  163. 6:14neural network architecture is
  164. 6:15just a large blob of neural net stuff for our purposes,
  165. 6:18and it's got several,
  166. 6:2010 billion parameters typically or something like that.
  167. 6:22Of course, as I tune these parameters,
  168. 6:23you're getting slightly
  169. 6:24different predicted distributions
  170. 6:26for every single one of these cells.
  171. 6:28For example, if our vocabulary size is 50,257 tokens,
  172. 6:34then we're going to have that many
  173. 6:35numbers because we need to
  174. 6:36specify a probability
  175. 6:38distribution for what comes next.
  176. 6:40Basically, we have a probability for
  177. 6:41whatever may follow.
  178. 6:42Now, in this specific example,
  179. 6:44for this specific cell,
  180. 6:45513 will come next,
  181. 6:47and so we can use this as
  182. 6:48a source of supervision to
  183. 6:49update our transformers weights.
  184. 6:51We're applying this basically
  185. 6:53on every single cell in the parallel,
  186. 6:54and we keep swapping batches,
  187. 6:56and we're trying to get the transformer to make
  188. 6:58the correct predictions over what
  189. 6:59token comes next in a sequence.
  190. 7:02Let me show you more concretely what this looks
  191. 7:03like when you train one of these models.
  192. 7:05This is actually coming from the New York Times,
  193. 7:07and they trained a small GPT on Shakespeare.
  194. 7:11Here's a small snippet of Shakespeare,
  195. 7:12and they train their GPT on it.
  196. 7:14Now, in the beginning,
  197. 7:15at initialization,
  198. 7:17the GPT starts with completely random weights.
  199. 7:19You're getting completely random outputs as well.
  200. 7:21But over time, as you train the GPT longer and longer,
  201. 7:26you are getting more and more coherent and
  202. 7:28consistent samples from the model,
  203. 7:31and the way you sample from it, of course,
  204. 7:32is you predict what comes next,
  205. 7:35you sample from that distribution and
  206. 7:36you keep feeding that back into the process,
  207. 7:38and you can basically sample large sequences.
  208. 7:42By the end, you see that the transformer
  209. 7:43has learned about words and
  210. 7:45where to put spaces and where to put commas and so on.
  211. 7:48We're making
  212. 7:48more and more consistent predictions over time.
  213. 7:51These are the plots that you are looking at
  214. 7:53when you're doing model pretraining.
  215. 7:54Effectively, we're looking at
  216. 7:56the loss function over time as you train,
  217. 7:58and low loss means that our transformer
  218. 8:00is giving a higher probability
  219. 8:03to the next correct integer in the sequence.
  220. 8:06What are we going to do with model
  221. 8:08once we've trained it after a month?
  222. 8:10Well, the first thing that we noticed, we the field,
  223. 8:14is that these models
  224. 8:16basically in the process of language modeling,
  225. 8:18learn very powerful general representations,
  226. 8:21and it's possible to very efficiently fine tune them
  227. 8:23for any arbitrary downstream tasks
  228. 8:25you might be interested in.
  229. 8:26As an example, if you're
  230. 8:27interested in sentiment classification,
  231. 8:29the approach used to be
  232. 8:31that you collect a bunch of positives
  233. 8:33and negatives and then you train some NLP model
  234. 8:35for that, but the new approach is:
  235. 8:38ignore sentiment
  236. 8:38classification, go off and do large
  237. 8:41language model pretraining,
  238. 8:43train a large transformer,
  239. 8:44and then you may only have a few examples and
  240. 8:47you can very efficiently fine tune
  241. 8:48your model for that task.
  242. 8:51This works very well in practice.
  243. 8:53The reason for this is that basically
  244. 8:55the transformer is forced to
  245. 8:56multitask a huge amount of
  246. 8:58tasks in the language modeling task,
  247. 9:00because in terms of predicting the next token,
  248. 9:03it's forced to understand a lot about the structure of
  249. 9:05the text and all the different concepts therein.
  250. 9:09That was GPT-1. Now around the time of GPT-2,
  251. 9:12people noticed that actually
  252. 9:14even better than fine tuning,
  253. 9:15you can actually prompt these models very effectively.
  254. 9:17These are language models and they want
  255. 9:19to complete documents,
  256. 9:20you can actually trick them into performing
  257. 9:22tasks by arranging these fake documents.
  258. 9:25In this example, for example,
  259. 9:27we have some passage and then we like do QA, QA, QA.
  260. 9:31This is called Few-shot prompt, and then we do Q,
  261. 9:33and then as the transformer is tried to
  262. 9:35complete the document is actually answering our question.
  263. 9:37This is an example of prompt engineering based model,
  264. 9:40making it believe that it's imitating
  265. 9:42a document and getting it to perform a task.
  266. 9:45This kicked off, I think the era of, I would say,
  267. 9:48prompting over fine tuning and seeing that this
  268. 9:50actually can work extremely well on a lot of problems,
  269. 9:53even without training any neural networks,
  270. 9:55fine tuning or so on.
  271. 9:56Now since then, we've seen
  272. 9:58an entire evolutionary tree of
  273. 10:00base models that everyone has trained.
  274. 10:02Not all of these models are available.
  275. 10:05for example, the GPT-4 base model was never released.
  276. 10:08The GPT-4 model that you might be
  277. 10:09interacting with over API is not a base model,
  278. 10:12it's an assistant model,
  279. 10:13and we're going to cover how to get those in a bit.
  280. 10:15GPT-3 based model is available via the API under
  281. 10:19the name Devanshi and GPT-2 based model
  282. 10:21is available even as weights on our GitHub repo.
  283. 10:24But currently the best available base model
  284. 10:27probably is the LLaMA series from Meta,
  285. 10:29although it is not commercially licensed.
  286. 10:32Now, one thing
  287. 10:34to point out is
  288. 10:35base models are not assistants.
  289. 10:36They don't want to make answers to your questions,
  290. 10:41they want to complete documents.
  291. 10:43If you tell them to write
  292. 10:44a poem about the bread and cheese,
  293. 10:46it will answer questions with more questions,
  294. 10:49it's completing what it thinks is a document.
  295. 10:51However, you can prompt them in a specific way for
  296. 10:54base models that is more likely to work.
  297. 10:57As an example, here's a poem about bread and cheese,
  298. 10:59and in that case it will autocomplete correctly.
  299. 11:02You can even trick base models into being assistants.
  300. 11:06The way you would do this is you would create
  301. 11:08a specific few-shot prompt
  302. 11:09that makes it look like there's
  303. 11:11some document between the human and assistant
  304. 11:13and they're exchanging information.
  305. 11:16Then at the bottom,
  306. 11:17you put your query at the end and the base model
  307. 11:21will condition itself into
  308. 11:23being a helpful assistant and answer,
  309. 11:26but this is not very reliable and doesn't work
  310. 11:28super well in practice, although it can be done.
  311. 11:30Instead, we have a different path to make
  312. 11:32actual GPT assistants not base model document completers.
  313. 11:37That takes us into supervised finetuning.
  314. 11:39In the supervised finetuning stage,
  315. 11:41we are going to collect
  316. 11:43small but high quality data-sets, and in this case,
  317. 11:45we're going to ask human contractors to gather data of
  318. 11:48the form prompt and ideal response.
  319. 11:52We're going to collect lots of these
  320. 11:54typically tens of thousands or something like that.
  321. 11:56Then we're going to still do language
  322. 11:58modeling on this data.
  323. 11:59Nothing changed algorithmically,
  324. 12:01we're swapping out a training set.
  325. 12:02It used to be Internet documents,
  326. 12:04which has a high quantity local
  327. 12:06for basically Q8 prompt response data.
  328. 12:11That is low quantity, high quality.
  329. 12:13We will still do language modeling
  330. 12:15and then after training,
  331. 12:16we get an SFT model.
  332. 12:18You can actually deploy these models and
  333. 12:20they are actual assistants and they work to some extent.
  334. 12:22Let me show you what an
  335. 12:24example demonstration might look like.
  336. 12:25Here's something that a human contractor
  337. 12:27might come up with.
  338. 12:28Here's some random prompt. Can you
  339. 12:29write a short introduction
  340. 12:31about the relevance of
  341. 12:32the term monopsony or something like that?
  342. 12:34Then the contractor also writes out an ideal response.
  343. 12:37When they write out these responses,
  344. 12:38they are following extensive labeling
  345. 12:40documentations and they are being
  346. 12:42asked to be helpful, truthful, and harmless.
  347. 12:45These labeling instructions here,
  348. 12:48you probably can't read it, neither can I,
  349. 12:50but they're long and this is people
  350. 12:52following instructions and trying
  351. 12:53to complete these prompts.
  352. 12:55That's what the dataset looks like.
  353. 12:57You can train these models. This works to some extent.
  354. 12:59Now, you can actually continue the pipeline from
  355. 13:02here on, and go into RLHF,
  356. 13:05reinforcement learning from human feedback that
  357. 13:07consists of both reward modeling
  358. 13:09and reinforcement learning.
  359. 13:10Let me cover that and then I'll
  360. 13:11come back to why you may want to go through
  361. 13:13the extra steps and how that compares to SFT models.
  362. 13:16In the reward modeling step,
  363. 13:18what we're going to do is we're now going to shift
  364. 13:20our data collection to be of the form of comparisons.
  365. 13:23Here's an example of what our dataset will look like.
  366. 13:25I have the same identical prompt on the top,
  367. 13:28which is asking the assistant to write
  368. 13:31a program or a function that
  369. 13:32checks if a given string is a palindrome.
  370. 13:35Then what we do is we take the SFT model which
  371. 13:38we've already trained and we create multiple completions.
  372. 13:41In this case, we have three completions
  373. 13:42that the model has created,
  374. 13:43and then we ask people to rank these completions.
  375. 13:47If you stare at this for a while, and by the way,
  376. 13:49these are very difficult things to do to
  377. 13:51compare some of these predictions.
  378. 13:52This can take people even hours for
  379. 13:54a single prompt completion pairs,
  380. 13:57but let's say we decided that one of these is
  381. 14:00much better than the others and so on. We rank them.
  382. 14:03Then we can follow that with
  383. 14:04something that looks very much like
  384. 14:06a binary classification on
  385. 14:07all the possible pairs between these completions.
  386. 14:10What we do now is, we lay out our prompt in rows,
  387. 14:13and the prompt is identical across all three rows here.
  388. 14:16It's all the same prompt, but
  389. 14:17the completion of this varies.
  390. 14:19The yellow tokens are coming from the SFT model.
  391. 14:21Then what we do is we append
  392. 14:23another special reward readout token at
  393. 14:26the end and we basically only
  394. 14:28supervise the transformer at this single green token.
  395. 14:31The transformer will predict some reward
  396. 14:34for how good that completion is
  397. 14:36for that prompt and basically it makes
  398. 14:39a guess about the quality of each completion.
  399. 14:42Then once it makes a guess for every one of them,
  400. 14:44we also have the ground truth
  401. 14:46which is telling us the ranking of them.
  402. 14:47We can actually enforce that some of
  403. 14:50these numbers should be much higher
  404. 14:51than others, and so on.
  405. 14:52We formulate this into a loss function and we
  406. 14:54train our model to make reward predictions
  407. 14:56that are consistent with the ground truth coming
  408. 14:58from the comparisons from all these contractors.
  409. 15:01That's how we train our reward model.
  410. 15:02That allows us to score how good
  411. 15:04a completion is for a prompt.
  412. 15:06Once we have a reward model,
  413. 15:09we can't deploy this because this is
  414. 15:11not very useful as an assistant by itself,
  415. 15:13but it's very useful for the reinforcement
  416. 15:15learning stage that follows now.
  417. 15:16Because we have a reward model,
  418. 15:18we can score the quality of
  419. 15:19any arbitrary completion for any given prompt.
  420. 15:22What we do during reinforcement learning
  421. 15:24is we basically get, again,
  422. 15:26a large collection of prompts and now we do
  423. 15:28reinforcement learning with respect to
  424. 15:29the reward model. Here's what that looks like.
  425. 15:32We take a single prompt,
  426. 15:34we lay it out in rows,
  427. 15:36and now we use basically
  428. 15:38the model we'd like to train which
  429. 15:39was initialized at SFT model
  430. 15:41to create some completions in yellow,
  431. 15:43and then we append the reward token again
  432. 15:45and we read off the reward
  433. 15:47according to the reward model,
  434. 15:49which is now kept fixed.
  435. 15:50It doesn't change any more. Now the reward model
  436. 15:53tells us the quality of every single completion
  437. 15:55for all these prompts and so what we can do is we can now
  438. 15:58just basically apply the same
  439. 15:59language modeling loss function,
  440. 16:01but we're currently training on the yellow tokens,
  441. 16:04and we are weighing
  442. 16:06the language modeling objective
  443. 16:08by the rewards indicated by the reward model.
  444. 16:11As an example, in the first row,
  445. 16:13the reward model said that this is
  446. 16:15a fairly high-scoring completion
  447. 16:17and so all the tokens that we
  448. 16:18happen to sample on the first row are going to get
  449. 16:21reinforced and they're going to get
  450. 16:22higher probabilities for the future.
  451. 16:25Conversely, on the second row,
  452. 16:26the reward model really did not like
  453. 16:28this completion, -1.2.
  454. 16:29Therefore, every single token that we sampled in
  455. 16:32that second row is going to get
  456. 16:34a slightly higher probability for the future.
  457. 16:36We do this over and over on
  458. 16:37many prompts on many batches and basically,
  459. 16:39we get a policy that creates yellow tokens here.
  460. 16:43It's basically all the completions here will
  461. 16:46score high according to
  462. 16:47the reward model that we trained in the previous stage.
  463. 16:51That's what the RLHF pipeline is.
  464. 16:55Then at the end, you get a model that you could deploy.
  465. 16:58As an example, ChatGPT is an RLHF model,
  466. 17:02but some other models that you might
  467. 17:03come across for example,
  468. 17:05Vicuna-13B, and so on,
  469. 17:06these are SFT models.
  470. 17:08We have base models, SFT models, and RLHF models.
  471. 17:12That's the state of things there.
  472. 17:14Now why would you want to do RLHF?
  473. 17:16One answer that's not
  474. 17:19that exciting is that it works better.
  475. 17:20This comes from the instruct GPT paper.
  476. 17:22According to these experiments a while ago now,
  477. 17:25these PPO models are RLHF.
  478. 17:28We see that they are basically preferred in a lot
  479. 17:30of comparisons when we give them to humans.
  480. 17:33Humans prefer basically tokens
  481. 17:36that come from RLHF models compared to SFT models,
  482. 17:39compared to base model that is prompted to be
  483. 17:41an assistant. It just works better.
  484. 17:43But you might ask why does it work better?
  485. 17:47I don't think that there's a single amazing answer
  486. 17:49that the community has really agreed on,
  487. 17:51but I will offer one reason potentially.
  488. 17:55It has to do with the asymmetry between how easy
  489. 17:58computationally it is to compare versus generate.
  490. 18:02Let's take an example of generating a haiku.
  491. 18:04Suppose I ask a model to write a haiku about paper clips.
  492. 18:07If you're a contractor trying to train data,
  493. 18:10then imagine being a contractor
  494. 18:12collecting basically data for the SFT stage,
  495. 18:14how are you supposed to create
  496. 18:15a nice haiku for a paper clip?
  497. 18:16You might not be very good at that,
  498. 18:18but if I give you a few examples of
  499. 18:20haikus you might be able to
  500. 18:21appreciate some of these haikus a lot more than others.
  501. 18:24Judging which one of these is good is a much easier task.
  502. 18:27Basically, this asymmetry
  503. 18:29makes it so that comparisons are
  504. 18:31a better way to potentially leverage
  505. 18:33yourself as a human and
  506. 18:34your judgment to create a slightly better model.
  507. 18:37Now, RLHF models are not
  508. 18:40strictly an improvement on the base models in some cases.
  509. 18:43In particular, we'd notice for example
  510. 18:45that they lose some entropy.
  511. 18:46That means that they give more peaky results.
  512. 18:49They can output samples
  513. 18:54with lower variation than the base model.
  514. 18:55The base model has lots of entropy and
  515. 18:57will give lots of diverse outputs.
  516. 19:00For example, one place where I still
  517. 19:03prefer to use a base model is in the setup
  518. 19:06where you basically have
  519. 19:09n things and you want to generate more things like it.
  520. 19:13Here is an example that I just cooked up.
  521. 19:16I want to generate cool Pokemon names.
  522. 19:18I gave it seven Pokemon names and I asked the base model
  523. 19:21to complete the document and it
  524. 19:22gave me a lot more Pokemon names.
  525. 19:24These are fictitious. I tried to look them up.
  526. 19:27I don't believe they're actual Pokemons.
  527. 19:29This is the task that I think the base model would be
  528. 19:31good at because it still has lots of entropy.
  529. 19:33It'll give you lots of diverse cool
  530. 19:35more things that look like whatever you give it before.
  531. 19:41Having said all that, these are
  532. 19:43the assistant models that are probably
  533. 19:45available to you at this point.
  534. 19:47There was a team at Berkeley that ranked a lot of
  535. 19:50the available assistant models
  536. 19:51and give them basically Elo ratings.
  537. 19:53Currently, some of the best models,
  538. 19:54of course, are GPT-4,
  539. 19:55by far, I would say,
  540. 19:57followed by Claude,
  541. 19:58GPT-3.5, and then a number of models,
  542. 20:00some of these might be available as weights,
  543. 20:02like Vicuna, Koala, etc.
  544. 20:04The first three rows here are
  545. 20:07all RLHF models and
  546. 20:09all of the other models to my knowledge,
  547. 20:11are SFT models, I believe.
  548. 20:15That's how we train
  549. 20:17these models on the high level.
  550. 20:19Now I'm going to switch gears
  551. 20:20and let's look at how we can
  552. 20:22best apply the GPT assistant model to your problems.
  553. 20:26Now, I would like to work
  554. 20:27in setting of a concrete example.
  555. 20:29Let's work with a concrete example here.
  556. 20:32Let's say that you are working on
  557. 20:34an article or a blog post,
  558. 20:35and you're going to write this sentence at the end.
  559. 20:38"California's population is 53
  560. 20:40times that of Alaska." So for some reason,
  561. 20:42you want to compare the populations of these two states.
  562. 20:44Think about the rich internal monologue
  563. 20:47and tool use and how much work
  564. 20:49actually goes computationally in
  565. 20:50your brain to generate this one final sentence.
  566. 20:53Here's maybe what that could look like in your brain.
  567. 20:55For this next step, let me blog on my blog,
  568. 20:59let me compare these two populations.
  569. 21:01First I'm going to obviously need to
  570. 21:03get both of these populations.
  571. 21:05Now, I know that I probably
  572. 21:06don't know these populations off the top of
  573. 21:08my head so I'm aware
  574. 21:10of what I know or don't know of my self-knowledge.
  575. 21:12I go, I do some tool use and I go to Wikipedia and I
  576. 21:16look up California's population and Alaska's population.
  577. 21:19Now, I know that I should divide the two, but again,
  578. 21:22I know that dividing 39.2 by
  579. 21:240.74 is very unlikely to succeed.
  580. 21:26That's not the thing that I can
  581. 21:28do in my head and so therefore,
  582. 21:30I'm going to rely on
  583. 21:31the calculator so I'm going to use a calculator,
  584. 21:33punch it in and see that the output is roughly 53.
  585. 21:36Then maybe I do
  586. 21:38some reflection and sanity checks in
  587. 21:40my brain so does 53 makes sense?
  588. 21:42Well, that's quite a large fraction,
  589. 21:44but then California is the most
  590. 21:45populous state, so maybe that looks okay.
  591. 21:47Then I have all the information I might need,
  592. 21:50and now I get to the creative portion of writing.
  593. 21:52I might start to write something like "California has
  594. 21:5553x times greater" and then I think to myself,
  595. 21:58that's actually like really awkward phrasing so let me
  596. 22:00actually delete that and let me try again.
  597. 22:03As I'm writing, I have this separate process,
  598. 22:06almost inspecting what I'm
  599. 22:07writing and judging whether it looks good
  600. 22:09or not and then maybe I delete and maybe I reframe it,
  601. 22:13and then maybe I'm happy with what comes out.
  602. 22:15Basically long story short,
  603. 22:17a ton happens under the hood in terms of
  604. 22:19your internal monologue when you
  605. 22:20create sentences like this.
  606. 22:21But what does a sentence like this look like
  607. 22:24when we are training a GPT on it?
  608. 22:26From GPT's perspective, this
  609. 22:28is just a sequence of tokens.
  610. 22:30GPT, when it's reading or generating these tokens,
  611. 22:34it just goes chunk, chunk, chunk,
  612. 22:35chunk and each chunk is roughly
  613. 22:37the same amount of computational work for each token.
  614. 22:40These transformers are not
  615. 22:42very shallow networks they have
  616. 22:43about 80 layers of reasoning,
  617. 22:45but 80 is still not like too much.
  618. 22:47This transformer is going to do its best to imitate,
  619. 22:51but of course, the process here
  620. 22:53looks very different from the process that you took.
  621. 22:56In particular, in our final artifacts
  622. 22:59in the data sets that we create,
  623. 23:00and then eventually feed to
  624. 23:01LLMs, all that internal dialogue was completely
  625. 23:03stripped and unlike you,
  626. 23:07the GPT will look at every single token and
  627. 23:09spend the same amount of compute on every one of them.
  628. 23:12So, you can't expect it
  629. 23:13to do too much work per token and also in particular,
  630. 23:21basically these transformers are
  631. 23:22just like token simulators,
  632. 23:23they don't know what they don't know.
  633. 23:26They just imitate the next token.
  634. 23:27They don't know what they're good at or not good at.
  635. 23:29They just tried their best to imitate the next token.
  636. 23:32They don't reflect in the loop.
  637. 23:34They don't sanity check anything.
  638. 23:35They don't correct their mistakes along the way.
  639. 23:37By default, they just are sample token sequences.
  640. 23:40They don't have separate inner monologue streams
  641. 23:43in their head right? They're
  642. 23:43evaluating what's happening.
  643. 23:45Now, they do have some cognitive advantages,
  644. 23:48I would say and that is that they do actually have
  645. 23:51a very large fact-based knowledge
  646. 23:52across a vast number of areas because they have,
  647. 23:55say, several, 10 billion parameters.
  648. 23:57That's a lot of storage for a lot of facts.
  649. 23:59They also, I think have
  650. 24:02a relatively large and perfect working memory.
  651. 24:04Whatever fits into the context window
  652. 24:07is immediately available to
  653. 24:09the transformer through
  654. 24:10its internal self attention mechanism
  655. 24:12and so it's perfect memory,
  656. 24:14but it's got a finite size,
  657. 24:16but the transformer has a very direct access
  658. 24:18to it and so it can a
  659. 24:19losslessly remember anything that
  660. 24:22is inside its context window.
  661. 24:23This is how I would compare those two and the reason I
  662. 24:26bring all of this up is because I
  663. 24:27think to a large extent,
  664. 24:29prompting is just making up for
  665. 24:31this cognitive difference between
  666. 24:34these two architectures like
  667. 24:37our brains here and LLM brains.
  668. 24:39You can look at it that way almost.
  669. 24:41Here's one thing that people found for example
  670. 24:44works pretty well in practice.
  671. 24:45Especially if your tasks require reasoning,
  672. 24:48you can't expect the transformer
  673. 24:49to do too much reasoning per token.
  674. 24:52You have to really spread out
  675. 24:53the reasoning across more and more tokens.
  676. 24:56For example, you can't give a transformer
  677. 24:57a very complicated question and
  678. 24:59expect it to get the answer in a single token.
  679. 25:00There's just not enough time for it.
  680. 25:02"These transformers need tokens to
  681. 25:04think," I like to say sometimes.
  682. 25:06This is some of the things that work well,
  683. 25:08you may for example have a few-shot prompt that
  684. 25:10shows the transformer that it should show
  685. 25:12its work when it's answering
  686. 25:14question and if you give a few examples,
  687. 25:17the transformer will imitate that template and it
  688. 25:20will just end up working out
  689. 25:21better in terms of its evaluation.
  690. 25:24Additionally, you can elicit this behavior from
  691. 25:26the transformer by saying, let things step-by-step.
  692. 25:29Because this conditions the transformer into showing
  693. 25:32its work and because
  694. 25:34it snaps into a mode of showing its work,
  695. 25:36is going to do less computational work per token.
  696. 25:40It's more likely to succeed as a result because it's
  697. 25:42making slower reasoning over time.
  698. 25:46Here's another example, this one
  699. 25:47is called self-consistency.
  700. 25:49We saw that we had the ability
  701. 25:51to start writing and then if it didn't work out,
  702. 25:54I can try again and I can try multiple times
  703. 25:56and maybe select the one that worked best.
  704. 26:00In these approaches,
  705. 26:02you may sample not just once,
  706. 26:03but you may sample multiple times and
  707. 26:05then have some process for finding
  708. 26:07the ones that are good and then keeping
  709. 26:09just those samples or doing
  710. 26:10a majority vote or something like that.
  711. 26:11Basically these transformers in the process as
  712. 26:14they predict the next token, just like you,
  713. 26:16they can get unlucky
  714. 26:18and they could sample a not a very good
  715. 26:19token and they can go down like
  716. 26:21a blind alley in terms of reasoning.
  717. 26:24Unlike you, they cannot recover from that.
  718. 26:27They are stuck with every single token they
  719. 26:28sample and so they will continue the sequence,
  720. 26:31even if they know
  721. 26:32that this sequence is not going to work out.
  722. 26:34Give them the ability to look back,
  723. 26:36inspect or try to basically sample around it.
  724. 26:40Here's one technique also,
  725. 26:43it turns out that actually LLMs,
  726. 26:45they know when they've screwed up,
  727. 26:47so as an example, say you ask the model
  728. 26:50to generate a poem that does not
  729. 26:52rhyme and it might give you a poem,
  730. 26:54but it actually rhymes.
  731. 26:55But it turns out that especially for
  732. 26:57the bigger models like GPT-4,
  733. 26:58you can just ask it "did you meet the assignment?"
  734. 27:01Actually GPT-4 knows very
  735. 27:03well that it did not meet the assignment.
  736. 27:04It just got unlucky in its sampling.
  737. 27:07It will tell you, "No, I didn't actually meet
  738. 27:08the assignment here. Let me try again."
  739. 27:10But without you prompting
  740. 27:12it it doesn't know to revisit and so on.
  741. 27:17You have to make up for that in your prompts,
  742. 27:19and you have to get it to check,
  743. 27:21if you don't ask it to check,
  744. 27:23its not going to check by itself
  745. 27:24it's just a token simulator.
  746. 27:28I think more generally,
  747. 27:29a lot of these techniques fall into
  748. 27:31the bucket of what I would say recreating our System 2.
  749. 27:34You might be familiar with the System 1 and
  750. 27:36System 2 thinking for humans.
  751. 27:37System 1 is a fast automatic process and I
  752. 27:40think corresponds to an LLM just sampling tokens.
  753. 27:43System 2 is the slower deliberate
  754. 27:46planning part of your brain.
  755. 27:49This is a paper actually from
  756. 27:51just last week because
  757. 27:52this space is pretty quickly evolving,
  758. 27:53it's called Tree of Thought.
  759. 27:56The authors of this paper proposed maintaining
  760. 27:59multiple completions for any given prompt
  761. 28:02and then they are also scoring them along
  762. 28:04the way and keeping the ones that
  763. 28:06are going well if that makes sense.
  764. 28:08A lot of people are really playing
  765. 28:10around with prompt engineering
  766. 28:13to basically bring back some of
  767. 28:15these abilities that we have in our brain for LLMs.
  768. 28:19Now, one thing I would like to note
  769. 28:21here is that this is not just a prompt.
  770. 28:22This is actually prompts that are together
  771. 28:25used with some Python Glue code because you
  772. 28:28actually have to maintain multiple
  773. 28:29prompts and you also have to do
  774. 28:30some tree search algorithm here
  775. 28:32to figure out which prompts to expand, etc.
  776. 28:35It's a symbiosis of Python Glue code and
  777. 28:38individual prompts that are
  778. 28:39called in a while loop or in a bigger algorithm.
  779. 28:42I also think there's a really cool
  780. 28:43parallel here to AlphaGo.
  781. 28:44AlphaGo has a policy for
  782. 28:46placing the next stone when it plays go,
  783. 28:48and its policy was trained
  784. 28:50originally by imitating humans.
  785. 28:52But in addition to this policy,
  786. 28:54it also does Monte Carlo Tree Search.
  787. 28:56Basically, it will play out a number of possibilities in
  788. 28:59its head and evaluate all of
  789. 29:00them and only keep the ones that work well.
  790. 29:01I think this is an equivalent of
  791. 29:04AlphaGo but for text if that makes sense.
  792. 29:08Just like Tree of Thought,
  793. 29:10I think more generally people are
  794. 29:11starting to really explore
  795. 29:13more general techniques of not
  796. 29:15just the simple question-answer prompts,
  797. 29:17but something that looks a lot more like
  798. 29:19Python Glue code stringing together many prompts.
  799. 29:22On the right, I have an example from
  800. 29:23this paper called React where they
  801. 29:25structure the answer to a prompt
  802. 29:28as a sequence of thought-action-observation,
  803. 29:32thought-action-observation, and it's
  804. 29:34a full rollout and
  805. 29:35a thinking process to answer the query.
  806. 29:38In these actions, the model is also allowed to tool use.
  807. 29:42On the left, I have an example of AutoGPT.
  808. 29:45Now AutoGPT by the way is
  809. 29:47a project that I think got a lot of hype recently,
  810. 29:51but I think I still find it inspirationally interesting.
  811. 29:55It's a project that allows an LLM to keep
  812. 29:58the task list and continue to
  813. 30:00recursively break down tasks.
  814. 30:02I don't think this currently works very well and I would
  815. 30:04not advise people to use it in practical applications.
  816. 30:07I just think it's something to generally take inspiration
  817. 30:09from in terms of where this is going, I think over time.
  818. 30:12That's like giving our model System 2 thinking.
  819. 30:16The next thing I find interesting is,
  820. 30:19this following serve I would say
  821. 30:20almost psychological quirk of LLMs,
  822. 30:23is that LLMs don't want to succeed,
  823. 30:26they want to imitate.
  824. 30:28You want to succeed, and you should ask for it.
  825. 30:31What I mean by that is,
  826. 30:33when transformers are trained,
  827. 30:35they have training sets and there can be
  828. 30:38an entire spectrum of
  829. 30:39performance qualities in their training data.
  830. 30:41For example, there could be some kind of a prompt
  831. 30:43for some physics question or something like that,
  832. 30:45and there could be
  833. 30:45a student's solution that is completely wrong
  834. 30:47but there can also be an expert
  835. 30:49answer that is extremely right.
  836. 30:50Transformers can't tell the difference between low,
  837. 30:54they know about low-quality solutions
  838. 30:56and high-quality solutions,
  839. 30:57but by default, they want to imitate all of
  840. 30:59it because they're just trained on language modeling.
  841. 31:02At test time, you actually have
  842. 31:04to ask for a good performance.
  843. 31:06In this example in this paper,
  844. 31:08they tried various prompts.
  845. 31:10Let's think step-by-step was very powerful
  846. 31:13because it spread out the reasoning over many tokens.
  847. 31:15But what worked even better is,
  848. 31:17let's work this out in a step-by-step way
  849. 31:19to be sure we have the right answer.
  850. 31:20It's like conditioning on getting the right answer,
  851. 31:23and this actually makes the transformer work
  852. 31:25better because the transformer doesn't have
  853. 31:27to now hedge its probability mass
  854. 31:29on low-quality solutions,
  855. 31:31as ridiculous as that sounds.
  856. 31:33Basically, feel free to ask for a strong solution.
  857. 31:37Say something like, you are
  858. 31:38a leading expert on this topic.
  859. 31:39Pretend you have IQ 120, etc.
  860. 31:41But don't try to ask for too much IQ because if
  861. 31:44you ask for IQ 400,
  862. 31:46you might be out of data distribution,
  863. 31:48or even worse, you could be in data distribution for
  864. 31:51something like sci-fi stuff and it
  865. 31:52will start to take on some sci-fi,
  866. 31:54or like roleplaying or something like that.
  867. 31:56You have to find the right amount of IQ.
  868. 31:59I think it's got some U-shaped curve there.
  869. 32:02Next up, as we
  870. 32:04saw when we are trying to solve problems,
  871. 32:07we know what we are good at and what we're not good at,
  872. 32:09and we lean on tools computationally.
  873. 32:12You want to do the same potentially with your LLMs.
  874. 32:15In particular, we may want to give
  875. 32:18them calculators, code interpreters,
  876. 32:21and so on, the ability to do search,
  877. 32:23and there's a lot of techniques for doing that.
  878. 32:27One thing to keep in mind, again,
  879. 32:28is that these transformers by default may
  880. 32:30not know what they don't know.
  881. 32:32You may even want to tell the transformer in
  882. 32:34a prompt you are not very good at mental arithmetic.
  883. 32:37Whenever you need to do very large number addition,
  884. 32:40multiplication, or whatever,
  885. 32:41instead, use this calculator.
  886. 32:42Here's how you use the calculator,
  887. 32:43you use this token combination, etc.
  888. 32:46You have to actually spell it out because the model by
  889. 32:48default doesn't know what it's good at or not good at,
  890. 32:50necessarily, just like you and I might be.
  891. 32:54Next up, I think something that is very
  892. 32:56interesting is we went from
  893. 32:58a world that was retrieval only all the way,
  894. 33:02the pendulum has swung to the other extreme
  895. 33:03where its memory only in LLMs.
  896. 33:06But actually, there's this entire space in-between of
  897. 33:08these retrieval-augmented models and
  898. 33:10this works extremely well in practice.
  899. 33:12As I mentioned, the context window of
  900. 33:14a transformer is its working memory.
  901. 33:17If you can load the working memory
  902. 33:18with any information that is relevant to the task,
  903. 33:21the model will work extremely well
  904. 33:23because it can immediately access all that memory.
  905. 33:26I think a lot of people are really interested
  906. 33:28in basically retrieval-augment degeneration.
  907. 33:32On the bottom, I have an example of LlamaIndex which is
  908. 33:35one data connector to lots of different types of data.
  909. 33:38You can index all
  910. 33:41of that data and you can make it accessible to LLMs.
  911. 33:44The emerging recipe there is you take relevant documents,
  912. 33:47you split them up into chunks,
  913. 33:49you embed all of them,
  914. 33:50and you basically get embedding vectors
  915. 33:52that represent that data.
  916. 33:53You store that in the vector store and then at test time,
  917. 33:56you make some kind of a query to
  918. 33:57your vector store and you fetch chunks that
  919. 34:00might be relevant to your task and
  920. 34:01you stuff them into the prompt and then you generate.
  921. 34:04This can work quite well in practice.
  922. 34:06This is, I think, similar to
  923. 34:07when you and I solve problems.
  924. 34:09You can do everything from your memory and
  925. 34:11transformers have very large and extensive memory,
  926. 34:13but also it really helps to
  927. 34:14reference some primary documents.
  928. 34:17Whenever you find yourself going
  929. 34:19back to a textbook to find something,
  930. 34:21or whenever you find yourself going back to
  931. 34:22documentation of the library to look something up,
  932. 34:25transformers definitely want to do that too.
  933. 34:27You have some memory over how
  934. 34:30some documentation of the library
  935. 34:31works but it's much better to look it up.
  936. 34:33The same applies here.
  937. 34:35Next, I wanted to briefly talk
  938. 34:38about constraint prompting.
  939. 34:39I also find this very interesting.
  940. 34:41This is basically techniques
  941. 34:43for forcing a certain template in the outputs of LLMs.
  942. 34:50Guidance is one example from Microsoft actually.
  943. 34:53Here we are enforcing that
  944. 34:55the output from the LLM will be JSON.
  945. 34:57This will actually guarantee that
  946. 35:00the output will take on this form because they go
  947. 35:02in and they mess with the probabilities of
  948. 35:03all the different tokens that
  949. 35:04come out of the transformer and
  950. 35:05they clamp those tokens and then
  951. 35:07the transformer is only filling in the blanks here,
  952. 35:09and then you can enforce additional restrictions
  953. 35:11on what could go into those blanks.
  954. 35:13This might be really helpful, and I think
  955. 35:15this constraint sampling is also extremely interesting.
  956. 35:19I also want to say
  957. 35:20a few words about fine tuning.
  958. 35:22It is the case that you can get really
  959. 35:23far with prompt engineering,
  960. 35:25but it's also possible to
  961. 35:27think about fine tuning your models.
  962. 35:29Now, fine tuning models means that you
  963. 35:31are actually going to change the weights of the model.
  964. 35:33It is becoming a lot more
  965. 35:35accessible to do this in practice,
  966. 35:37and that's because of
  967. 35:38a number of techniques that have been
  968. 35:39developed and have libraries for very recently.
  969. 35:43So for example parameter efficient
  970. 35:44fine tuning techniques like Laura,
  971. 35:46make sure that you're only training small,
  972. 35:49sparse pieces of your model.
  973. 35:51So most of the model is kept clamped at
  974. 35:53the base model and some pieces of it are allowed to
  975. 35:55change and this still works pretty
  976. 35:56well empirically and makes
  977. 35:58it much cheaper to tune only small pieces of your model.
  978. 36:02It also means that because most of your model is clamped,
  979. 36:05you can use very low precision inference
  980. 36:07for computing those parts because
  981. 36:09you are not going to be updated by
  982. 36:10gradient descent and so that
  983. 36:12makes everything a lot more efficient as well.
  984. 36:13And in addition, we have a number of
  985. 36:15open source, high-quality base models.
  986. 36:17Currently, as I mentioned,
  987. 36:18I think LLaMa is quite nice,
  988. 36:20although it is not commercially
  989. 36:21licensed, I believe right now.
  990. 36:23Some things to keep in mind is that basically
  991. 36:26fine tuning is a lot more technically involved.
  992. 36:29It requires a lot more, I think,
  993. 36:30technical expertise to do right.
  994. 36:32It requires human data contractors for
  995. 36:34datasets and/or synthetic data pipelines
  996. 36:36that can be pretty complicated.
  997. 36:38This will definitely slow down
  998. 36:40your iteration cycle by a lot,
  999. 36:41and I would say on a high level SFT is
  1000. 36:44achievable because you're continuing
  1001. 36:47the language modeling task.
  1002. 36:48It's relatively straightforward, but RLHF,
  1003. 36:50I would say is very much research territory
  1004. 36:53and is even much harder to get to work,
  1005. 36:55and so I would probably not advise that someone
  1006. 36:58just tries to roll their own RLHF of implementation.
  1007. 37:00These things are pretty unstable,
  1008. 37:02very difficult to train, not something that is, I think,
  1009. 37:04very beginner friendly right now,
  1010. 37:06and it's also potentially likely also
  1011. 37:08to change pretty rapidly still.
  1012. 37:11So I think these are
  1013. 37:12my default recommendations right now.
  1014. 37:15I would break up your task into two major parts.
  1015. 37:18Number 1, achieve your top performance,
  1016. 37:20and Number 2, optimize your performance in that order.
  1017. 37:23Number 1, the best performance will
  1018. 37:25currently come from GPT-4 model.
  1019. 37:27It is the most capable of all by far.
  1020. 37:29Use prompts that are very detailed.
  1021. 37:31They have lots of task content,
  1022. 37:33relevant information and instructions.
  1023. 37:36Think along the lines of what would you tell
  1024. 37:38a task contractor if they can't email you back,
  1025. 37:40but then also keep in mind that a task contractor is a
  1026. 37:43human and they have
  1027. 37:44inner monologue and they're very clever, etc.
  1028. 37:46LLMs do not possess those qualities.
  1029. 37:48So make sure to think through
  1030. 37:50the psychology of the LLM
  1031. 37:52almost and cater prompts to that.
  1032. 37:54Retrieve and add any relevant context
  1033. 37:57and information to these prompts.
  1034. 37:59Basically refer to a lot of
  1035. 38:01the prompt engineering techniques.
  1036. 38:02Some of them I've highlighted in the slides above,
  1037. 38:04but also this is a very large space and I would
  1038. 38:07just advise you to look
  1039. 38:09for prompt engineering techniques online.
  1040. 38:11There's a lot to cover there.
  1041. 38:13Experiment with few-shot examples.
  1042. 38:15What this refers to is, you don't just want to tell,
  1043. 38:17you want to show whenever it's possible.
  1044. 38:19So give it examples of everything
  1045. 38:21that helps it really understand what you mean if you can.
  1046. 38:25Experiment with tools and plug-ins to
  1047. 38:27offload tasks that are difficult for LLMs natively,
  1048. 38:30and then think about not just a
  1049. 38:32single prompt and answer,
  1050. 38:33think about potential chains
  1051. 38:34and reflection and how you glue
  1052. 38:36them together and how you can
  1053. 38:37potentially make multiple samples and so on.
  1054. 38:40Finally, if you think you've squeezed
  1055. 38:42out prompt engineering,
  1056. 38:43which I think you should stick with for a while,
  1057. 38:45look at some potentially
  1058. 38:48fine tuning a model to your application,
  1059. 38:51but expect this to be a lot more
  1060. 38:52slower in the vault and then
  1061. 38:54there's an expert fragile research zone
  1062. 38:56here and I would say that is RLHF,
  1063. 38:58which currently does work a bit
  1064. 39:00better than SFT if you can get it to work.
  1065. 39:02But again, this is pretty involved, I would say.
  1066. 39:05And to optimize your costs,
  1067. 39:06try to explore lower capacity models
  1068. 39:09or shorter prompts and so on.
  1069. 39:12I also wanted to say a few words about the use cases
  1070. 39:15in which I think LLMs are currently well suited for.
  1071. 39:18In particular, note that there's a large number
  1072. 39:20of limitations to LLMs today,
  1073. 39:22and so I would keep that
  1074. 39:24definitely in mind for all of your applications.
  1075. 39:26Models, and this by the way could be an entire talk.
  1076. 39:28So I don't have time to cover it in full detail.
  1077. 39:30Models may be biased, they may fabricate,
  1078. 39:32hallucinate information,
  1079. 39:33they may have reasoning errors,
  1080. 39:35they may struggle in entire classes of applications,
  1081. 39:38they have knowledge cut-offs,
  1082. 39:40so they might not know any information above,
  1083. 39:42say, September, 2021.
  1084. 39:43They are susceptible to a large range of
  1085. 39:45attacks which are coming out on Twitter daily,
  1086. 39:48including prompt injection, jailbreak attacks,
  1087. 39:51data poisoning attacks and so on.
  1088. 39:52So my recommendation right now is
  1089. 39:54use LLMs in low-stakes applications.
  1090. 39:57Combine them always with human oversight.
  1091. 40:00Use them as a source of inspiration and
  1092. 40:01suggestions and think co-pilots,
  1093. 40:04instead of completely autonomous agents
  1094. 40:05that are just like performing a task somewhere.
  1095. 40:07It's just not clear that the models are there right now.
  1096. 40:11So I wanted to close by saying that
  1097. 40:13GPT-4 is an amazing artifact.
  1098. 40:15I'm very thankful that it exists, and it's beautiful.
  1099. 40:18It has a ton of knowledge across so many areas.
  1100. 40:20It can do math, code and so on.
  1101. 40:22And in addition, there's this
  1102. 40:24thriving ecosystem of everything
  1103. 40:25else that is being built and
  1104. 40:27incorporated into the ecosystem.
  1105. 40:29Some of these things I've talked about,
  1106. 40:31and all of this power is accessible at your fingertips.
  1107. 40:35So here's everything that's needed in terms of
  1108. 40:37code to ask GPT-4 a question,
  1109. 40:40to prompt it, and get a response.
  1110. 40:42In this case, I said,
  1111. 40:44can you say something to inspire
  1112. 40:45the audience of Microsoft Build 2023?
  1113. 40:47And I just punched this into Python and verbatim
  1114. 40:50GPT-4 said the following:
  1115. 40:53And by the way, I did not know that they
  1116. 40:55used this trick in the keynote.
  1117. 40:57So I thought I was being clever,
  1118. 40:59but it is really good at this.
  1119. 41:02It says, ladies and gentlemen,
  1120. 41:04innovators and trailblazers Microsoft Build 2023.
  1121. 41:07Welcome to the gathering of brilliant
  1122. 41:08minds like no other,
  1123. 41:09you are the architects of the future,
  1124. 41:11the visionaries molding the digital realm
  1125. 41:13in which humanity thrives.
  1126. 41:15Embrace the limitless possibilities of
  1127. 41:17technologies and let your ideas
  1128. 41:18soar as high as your imagination.
  1129. 41:20Together, let's create a more connected,
  1130. 41:22remarkable, and inclusive world for generations to come.
  1131. 41:25Get ready to unleash your creativity,
  1132. 41:27canvas the unknown, and turn dreams into reality.
  1133. 41:30Your journey begins today!

About this transcript

This page contains the full transcript of State of GPT | BRK216HFS by Microsoft Developer, generated from the public captions YouTube serves with the video. The transcript has 7,684 words across 1,133 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.