YouTube2Text

YouTube transcript (35X6zlhoCy4) — Transcript

14,287 words · 2,082 segments · language en · Watch on YouTube

Full transcript

  1. 0:05good evening people um even how are you
  2. 0:08guys
  3. 0:10doing all right my name is archa Sharma
  4. 0:13I'm a PhD student at Stanford and I'm
  5. 0:15very very excited to talk about post
  6. 0:17training generally speaking for large
  7. 0:19language models and I hope you guys are
  8. 0:21ready to learn some stuff because this
  9. 0:24has been one of the last few years in
  10. 0:25machine learning have been very very
  11. 0:26exciting uh with the Advent of large
  12. 0:29language model CH GPD and everything to
  13. 0:31that extent and hopefully after today's
  14. 0:33lecture you will be more comfortable
  15. 0:36understanding how we go from pre-train
  16. 0:38Models to models like CH GPD and we'll
  17. 0:40take a whole journey through prompting
  18. 0:43instruction fine tuning and DP and
  19. 0:45rlf so let's get
  20. 0:51started all
  21. 0:53right so something that has been very
  22. 0:57fundamental to our entire field is this
  23. 1:01idea of scaling loss and models are
  24. 1:04increasingly becoming larger and larger
  25. 1:06and they're expanding more and more
  26. 1:08compute so this is a graph of models
  27. 1:10starting all the way back in 1950s to
  28. 1:13somewhere around these are still this is
  29. 1:15an outdated graph so like this shows up
  30. 1:17to 10 to^ 24 flops or floating Point
  31. 1:19operations that go into pre-training
  32. 1:21these models but the number is well
  33. 1:23above 10 to^ 26 now but you can see the
  34. 1:26graph and the way it's
  35. 1:28trending and more more and more compute
  36. 1:30requires more and more data because you
  37. 1:32need to train on something meaningful
  38. 1:34and this is roughly the trend on the
  39. 1:35amount of language tokens that are going
  40. 1:37into the language models in pre-training
  41. 1:40and again this plot is outdated does
  42. 1:43anybody want to guess like we're in 2024
  43. 1:452022 we were at 1.4 trillion tokens or
  44. 1:49words roughly speaking in language model
  45. 1:51pre-training do does anyone want to
  46. 1:53guess like where we are in 2024
  47. 2:00that's a pretty good guess yeah so we're
  48. 2:02close to 15 trillion tokens um recent
  49. 2:05llama 3 models were roughly trained on
  50. 2:0615 trillion tokens so yeah just just for
  51. 2:10a second appreciate that these are a lot
  52. 2:12of words uh this is not yeah I don't I
  53. 2:16don't think anybody of us listens to
  54. 2:18like trillions of tokens in our lifetime
  55. 2:20so this is where we are right now and I
  56. 2:24hope you guys were here for the pre-
  57. 2:25pre-training lectures cool um so what do
  58. 2:30we do so like I mean broadly speaking we
  59. 2:32are really just learning to predict text
  60. 2:34tokens or language tokens but what do we
  61. 2:37learn in the process of pre-training why
  62. 2:39is why are people spending so much money
  63. 2:42so much compute because these Compu and
  64. 2:44tokens take dollars to do and we're
  65. 2:46we're on the order spending hundreds of
  66. 2:48millions of dollars on these runs so why
  67. 2:50are we doing this and this is basically
  68. 2:53a recall from whatever you have probably
  69. 2:55learned till now but we're learning
  70. 2:57things like oh we are learning knowledge
  71. 2:59Stanford University is located in Santa
  72. 3:02clar California or wherever you want to
  73. 3:04like say you're learning syntax you're
  74. 3:06learning semantics of the sentences
  75. 3:08these are things that you would expect
  76. 3:10to learn when you're training on
  77. 3:12language data broadly you're probably
  78. 3:14learning a lot about different languages
  79. 3:16as well so like depending on your text
  80. 3:17Data distribution you're learning a lot
  81. 3:19of things but the models we interact
  82. 3:22with are very intelligent so where is
  83. 3:24that coming from like I mean just simply
  84. 3:26learning about very factual things and
  85. 3:30it's a very simple loss function we're
  86. 3:32optimizing and where is that
  87. 3:33Intelligence coming
  88. 3:35from and this perhaps is the interesting
  89. 3:39bit recently like people have like
  90. 3:43started accumulating evidence for that
  91. 3:45like when you optimize the next token
  92. 3:47prediction losses you're not just
  93. 3:49learning about syntax you're not just
  94. 3:50learning knowledge but you're starting
  95. 3:52to like form models of Agents beliefs
  96. 3:55and actions as well so how do we know
  97. 3:58this again a lot of this is speculative
  98. 4:00evidence but it's mayy to like form an
  99. 4:02understanding that the losses we're
  100. 4:03optimizing are not just about the data
  101. 4:04fitting the data but you start learning
  102. 4:06something maybe more meaningful as
  103. 4:08well um for example like I mean in this
  104. 4:11specific case um we change the last
  105. 4:15sentence and the prediction of the text
  106. 4:17or the next Tex that that is predicted
  107. 4:19changes as well so here it starts with
  108. 4:22Pat watch as a demonstration of a
  109. 4:23bowling ball and the leaf being dropped
  110. 4:25at the same time pat who is a physicist
  111. 4:27predicts that the bowling ball and the
  112. 4:29leaf will land at the same rate we all
  113. 4:31know Gravity the way it works but when
  114. 4:34you L change the last sentence to Pat
  115. 4:36who has never seen this demonstration
  116. 4:38before Pat predicts that bowling ball
  117. 4:40will fall to the ground first maybe
  118. 4:42somebody who's never seen this
  119. 4:43experiment before might intuitively
  120. 4:44believe that correct so like I mean the
  121. 4:47language model was able to predict this
  122. 4:48so how do you predict this you have to
  123. 4:51have some notion of understanding of how
  124. 4:54humans work to even be able to predict
  125. 4:56this and that's maybe like something
  126. 4:58that is not obvious with You're simply
  127. 5:00optimizing to predict the
  128. 5:03text similarly like I mean these kind of
  129. 5:05examples are we're going to run through
  130. 5:06some examples to like sort of
  131. 5:08communicate that when you're
  132. 5:09pre-training these models you're
  133. 5:10learning much more than just language
  134. 5:11tokens and so on you're also learning
  135. 5:13about math like you're able to
  136. 5:16understand what a graph of a circle
  137. 5:17means and what the center is and where
  138. 5:19how to like understand
  139. 5:22equations probably my favorite example
  140. 5:24something I use pretty much every day is
  141. 5:26you're learning how to write code so I
  142. 5:29don't know how many of you have
  143. 5:30interacted with co-pilot before but if
  144. 5:33you have like you probably know like if
  145. 5:34you write down a few commands write down
  146. 5:36a function template it will
  147. 5:38automatically complete code for you so
  148. 5:41again it's not perfect but it has to
  149. 5:43have some deeper understanding of what
  150. 5:45your intent is for something like that
  151. 5:47to
  152. 5:48emerge and similarly we have examples
  153. 5:50from medicine as well I don't know about
  154. 5:52you guys but like whenever I have some
  155. 5:54issue I probably go to chat gbd or
  156. 5:55Claude or something to that effect and
  157. 5:57ask them a diagnosis for those things as
  158. 5:59well
  159. 6:00um I don't recommend that uh please
  160. 6:03don't take medical advice from me but
  161. 6:06yeah so broadly like the way we're
  162. 6:09seeing language models at this point is
  163. 6:11that like they're sort of emerging as
  164. 6:12this general purpose multitask
  165. 6:14assistance and it's very strange right
  166. 6:17like I mean we started off with text
  167. 6:18token prediction and we're reaching the
  168. 6:19stage where it can like sort of rely to
  169. 6:21them on them to do many many different
  170. 6:23things so how are we getting there and
  171. 6:25I'm sure you all are aware of like what
  172. 6:26these models are so yeah
  173. 6:30so today's lecture is largely going to
  174. 6:32be about how do we go from something
  175. 6:34Stanford University is located this very
  176. 6:37simple pretraining task a very simple
  177. 6:38procedure well it's more complicated but
  178. 6:40in abstract terms it's not very
  179. 6:42complicated to like something as
  180. 6:44powerful as CH
  181. 6:45GPD cool
  182. 6:48so um I recommend you guys stopping me
  183. 6:50asking me a lot of questions because
  184. 6:52this is a there's a lot of fun examples
  185. 6:54and a lot of fun techniques so like I I
  186. 6:56want you guys to like learn everything
  187. 6:58about here so the overall plan is we're
  188. 7:00going to talk about zero shot and few
  189. 7:02shot in context learning um next we're
  190. 7:05going to follow up with instruction
  191. 7:06fine-tuning and then we're going to talk
  192. 7:08about optimizing for preferences and
  193. 7:10this is where roughly things are right
  194. 7:12now in the industry and when we're going
  195. 7:15to talk about what's next what the
  196. 7:16limitations are and how do we move from
  197. 7:20here cool so we're going to start off
  198. 7:23with zero shot INF fusure in context
  199. 7:27learning um broad we're going to take an
  200. 7:29example example of GPT or the generative
  201. 7:31pre-train Transformer and this is a
  202. 7:33whole series of models that started off
  203. 7:34in roughly 2018 and like up to 2020 they
  204. 7:37were building GPD gpd2 gbd3 so we're
  205. 7:40going to start off with this example and
  206. 7:42yes so it's a decoder only model that is
  207. 7:45trained on roughly 4.6 GB of text and it
  208. 7:49has 12 layers of Transformers layers and
  209. 7:51it's trained with the next token
  210. 7:52prediction
  211. 7:53loss and the first model obviously was
  212. 7:57not extremely good but it started
  213. 7:58showing that like hey like this
  214. 8:00technique for pre-training can be very
  215. 8:03effective for general purpose tasks and
  216. 8:05we're going to see some
  217. 8:06examples um for example like I mean here
  218. 8:09it's able to do the task for entainment
  219. 8:12and okay
  220. 8:16um yeah and gbd1 itself was not very
  221. 8:21strong as a model so like but they took
  222. 8:23the same recipe and like I mean tried to
  223. 8:25like increase the model size so they
  224. 8:27went from 117 million parameters to
  225. 8:29about 1.5 billion parameters and we're
  226. 8:32now scaling up the data alongside as
  227. 8:34well so we went from 4 gabt of data to
  228. 8:36approximately 40 gab of data and
  229. 8:38pre-training is a whole different like
  230. 8:40melting part of techniques and there's a
  231. 8:42lot that goes into it but like roughly
  232. 8:43for example here they filter data by the
  233. 8:46number of upwards on the redit
  234. 8:48data and yeah so this is roughly where
  235. 8:52we are and I think one of the things
  236. 8:54that started emerging with gpd2 is zero
  237. 8:57shot learning and what do we mean by
  238. 9:00zero shot learning
  239. 9:02um conventionally in the field like when
  240. 9:05we pre-train models there was the idea
  241. 9:06that you take a few examples you update
  242. 9:08the model um and then you are able to
  243. 9:11adapt to a specific task but as you
  244. 9:14pre-train on more and more data and more
  245. 9:15and more tasks you sort of start seeing
  246. 9:17this phenomena where they're able to do
  247. 9:19the task basically zero short they're
  248. 9:21shown no examples of how to do the task
  249. 9:23and you can start thinking of oh how you
  250. 9:25can do it summarization you can follow
  251. 9:27some instructions you can do maybe a
  252. 9:29little bit of math as well so this is
  253. 9:31where the idea of zero shot learning
  254. 9:33started to
  255. 9:38emerge yeah so how do we do zero shot
  256. 9:41learning or task specific learning from
  257. 9:42these pre-trained models really the idea
  258. 9:45is that we have to be creative here we
  259. 9:47know that these are text prediction
  260. 9:48models if you put in a text they will
  261. 9:50complete whatever follows so if we can
  262. 9:52sort of course these models into
  263. 9:54completing the task we care about maybe
  264. 9:56it's question answering we can so start
  265. 9:59getting them to solve tasks here so for
  266. 10:02example if you want to ask questions
  267. 10:04about Tom Brady you sort of set it up
  268. 10:07you sort put information about Tom Brady
  269. 10:09and then you put a question that you
  270. 10:10wanted to answer and then it will
  271. 10:12autocomplete in some sense so this is
  272. 10:14one early perspective on these models
  273. 10:16these are very Advanced autocomplete
  274. 10:19models and
  275. 10:21similarly if you want to figure out like
  276. 10:23which answer is true or which is not
  277. 10:25something that is very useful to measure
  278. 10:27is log probabilities so
  279. 10:29for example we want to figure out what
  280. 10:32is the word it refering to here in this
  281. 10:35sentence the cat couldn't fit into the
  282. 10:36Hat because it was too big um what we
  283. 10:39can do is we can take the sentence
  284. 10:41replace it with either the cat or either
  285. 10:44the hat and then you can measure the
  286. 10:46probability of which Mo which one does
  287. 10:49the model think is higher and you can
  288. 10:51sort of get the idea what the reference
  289. 10:53here is so none of this is like in the
  290. 10:56training data it's simply learning to
  291. 10:58predict text but you can start seeing
  292. 11:00how like we can leverage these models to
  293. 11:02do other tasks as well be besides
  294. 11:06prediction so this is just more evidence
  295. 11:09about like how gpd2 no task specific
  296. 11:12fine-tuning no task specific training it
  297. 11:15simply is learning to predict text and
  298. 11:17it's establishes the state-of-the-art on
  299. 11:19many many different tasks simply by
  300. 11:22scaling up the model parameters and
  301. 11:23scaling up the amount of data it's
  302. 11:25stained on
  303. 11:29so this is a fun example so if you want
  304. 11:31to do summarization for data or like you
  305. 11:35have a news article that you want to
  306. 11:37summarize so how do you get a zero shot
  307. 11:40model to do it this answer is you put
  308. 11:42the document into the context and you
  309. 11:44simply put tldr in front of
  310. 11:47it now like I mean if most of the data
  311. 11:50on internet whenever you see tldr you'll
  312. 11:51naturally summarize it so yeah you can
  313. 11:54get zero shot performance and
  314. 11:56summarization here as well and again
  315. 11:57this is not trained to do something
  316. 11:59summarization in any specific way and
  317. 12:01it's still doing really well simply
  318. 12:03because of its pre-training
  319. 12:07data so yeah um I think gp22 tldr is
  320. 12:11somewhere there and like some of the
  321. 12:12very Tas specific train models are um
  322. 12:15like and I think you will see the trend
  323. 12:18with again if you were Alec Radford or
  324. 12:20somebody like I mean you see like these
  325. 12:22cool things emerging your next step
  326. 12:24would obviously be I'm going to scale
  327. 12:26this up a little more I'm going to make
  328. 12:27an even bigger model I'm going to train
  329. 12:28it even more data and we'll see how
  330. 12:31things go right so that's how we got
  331. 12:34gbd3 uh we went from 1.5 billion
  332. 12:36parameters to 175 billion parameters we
  333. 12:38are well over like 40 GB of data to 600
  334. 12:42gbt of data of course like now we're in
  335. 12:44like terabytes of data and text is a
  336. 12:47very compressed representation so like
  337. 12:48terabytes of data is a
  338. 12:50lot um and you know we we talked about
  339. 12:53zero shot learning the cool thing that
  340. 12:56emerged in gbd3 is like go ahead like
  341. 13:00used before the passage right no you
  342. 13:03typically put the passage uh if youve
  343. 13:04like interacted with Reddit or something
  344. 13:06like that typically somebody will write
  345. 13:08an entire post and then end with TLD drr
  346. 13:11here's a summary of the thing too long
  347. 13:13didn't read or if you have
  348. 13:15used opposite comes first oh yeah there
  349. 13:20are situations where it also comes first
  350. 13:21but um one reason is that these are like
  351. 13:24decoder only models so like they are
  352. 13:27often these are causal attention models
  353. 13:28so the typically need to see the context
  354. 13:30before yeah understand I'm just curious
  355. 13:33like from my experience the comes first
  356. 13:36then how is
  357. 13:38it
  358. 13:40the okay um there's probably a lot of
  359. 13:43data where the tldr comes first but
  360. 13:44there's probably a lot of data where
  361. 13:45tldr comes after as
  362. 13:47well cool so we saw Zero shot learning
  363. 13:51emerging in gpd2 few shot learning maybe
  364. 13:54seem slightly easier but like this is
  365. 13:56where things started getting really
  366. 13:57funny is that like you're starting to
  367. 13:59beat state-ofthe-art simply by just
  368. 14:00putting examples in context so yeah what
  369. 14:04does f shot learning mean here what is
  370. 14:06what are we talking about um as I
  371. 14:09mentioned like the typical idea here is
  372. 14:13is that like you want to solve
  373. 14:14translation so you would put some
  374. 14:16examples of translation into
  375. 14:18context and you know this is a
  376. 14:21correction task here or maybe you want
  377. 14:22interested in Translation and no
  378. 14:25gradient updates no learning in any
  379. 14:28conventional sense whatso ever you put a
  380. 14:30few examples in and that's it like I
  381. 14:32mean you know how to solve the task
  382. 14:34isn't that like crazy like you're bu you
  383. 14:37guys did the assignment on translation
  384. 14:39right so but this is what the modern NLP
  385. 14:42looks like
  386. 14:44so yeah um you put in some examples and
  387. 14:48you have the entire system and this is
  388. 14:50where things got really interesting is
  389. 14:52that all these task specific models that
  390. 14:54were created to like be really really
  391. 14:56good at translation or really good at
  392. 14:57summarization you can just put let's
  393. 15:00look at this graph so we start with a
  394. 15:02zero shot performance of this in a
  395. 15:04similar fashion that I described earlier
  396. 15:05and you start somewhere there you put
  397. 15:07one example in of translation from
  398. 15:09English to French you get to somewhere
  399. 15:11like already at a fine level few
  400. 15:13examples in you're already starting to
  401. 15:15like be close to the state-ofthe-art
  402. 15:18models wait but in that gra the state of
  403. 15:20theart is really high isn't
  404. 15:22it uh find your Inver of the bird Plus+
  405. 15:26here I think is like the one I'm
  406. 15:27referring to find your of the like which
  407. 15:29is trained exclusively on a lot of um
  408. 15:32translation data might be like slightly
  409. 15:33better yes um and I think that's the
  410. 15:37relevant comparison here is the in
  411. 15:39context learning starts to emerge at
  412. 15:42scale so and this is I think like the
  413. 15:45key point is that this some of this is
  414. 15:48contested just to be very upfront but
  415. 15:50like there's this idea of emergence of
  416. 15:52this property as you train on more
  417. 15:54Computing and more scale um there's more
  418. 15:56recent research which suggest that if we
  419. 15:58plot the access correctly it feels less
  420. 16:00emergent but the general idea is as you
  421. 16:02increase the number of parameters and
  422. 16:05increase the number of compute that is
  423. 16:06going into the models the ability to
  424. 16:08just go from a few examples to really
  425. 16:10strong performance is very
  426. 16:14compelling cool
  427. 16:17um and yeah I think as I explained
  428. 16:20earlier the general idea is that this is
  429. 16:22very different from the conventional
  430. 16:23idea of fine tuning that we typically go
  431. 16:25for instead of like iterating over
  432. 16:27examples putting it into context and
  433. 16:29doing gradient updates we are actually
  434. 16:31just going for few short promting we're
  435. 16:33going to put in few examples and that's
  436. 16:34going to give us the
  437. 16:42system um yes I mean the exact details
  438. 16:45roughly can depend on the prom template
  439. 16:47that you use but typically you would
  440. 16:49just put examples so like c order and
  441. 16:52put these examples and then whatever
  442. 16:54your task is you can just let the model
  443. 16:56complete from there because it can infer
  444. 16:58the task
  445. 16:59um based on the examples you've
  446. 17:02given any other
  447. 17:05questions
  448. 17:08cool
  449. 17:10so yeah like I mean we have gotten from
  450. 17:12zero shot prompting and like we've seen
  451. 17:14seeing that F shot prompting is becoming
  452. 17:16really competitive with good models but
  453. 17:18there's still limitations to this like I
  454. 17:20mean you cannot solve every task that
  455. 17:21you see here and particularly like
  456. 17:23things that involve like richer
  457. 17:25multi-step reasoning is something that
  458. 17:26actually can be pretty challenging and
  459. 17:29just to be fair human struggle at these
  460. 17:30tasks as well so things like addition
  461. 17:33and so on like these are these are
  462. 17:36probably like still still hard to do
  463. 17:38like when you keep increasing the number
  464. 17:39of digits but one thing that you have to
  465. 17:43start being creative with I alluded to
  466. 17:45this earlier is that you can get these
  467. 17:46models to do the task if you're creative
  468. 17:49in how you prompt the model and this is
  469. 17:52what we're going to see
  470. 17:53next um so this technique called Chain
  471. 17:57of Thought prompting emerged here and
  472. 17:59the idea that we have explored thus far
  473. 18:01is that we put in examples of the kind
  474. 18:03of tasks we want to do and we expect the
  475. 18:07model to zero shot learn what the task
  476. 18:09is and go from there um the idea is that
  477. 18:13like instead of just showing what the
  478. 18:15task is you show them examples where
  479. 18:17they reason through the task so they're
  480. 18:19not just learning to do the task but
  481. 18:21also learning how the reasoning is
  482. 18:22working so in this example initially we
  483. 18:24started with like we have to solve a
  484. 18:25simple math problem and we are just
  485. 18:27shown exactly the answer answer directly
  486. 18:30instead of doing that and if you do that
  487. 18:32directly you'll observe that the model
  488. 18:33gets the answer wrong instead of that
  489. 18:36what if you show model how to reason
  490. 18:37about the task show it a chain of
  491. 18:39thought and include that in the prompt
  492. 18:41as
  493. 18:42well and then you ask at a new question
  494. 18:46the idea is that now the model is not
  495. 18:47just going to Output an answer it's
  496. 18:50going to reason about the task and it's
  497. 18:52going to do actually a lot better and
  498. 18:54this has been shown to be very effective
  499. 18:57um Chain of Thought is also as you can
  500. 19:00see like I mean it's again something
  501. 19:02that improves a lot with model scale
  502. 19:05it's not just um yeah um but what you
  503. 19:09can probably start seeing is like it's
  504. 19:11nearly better than supervised best
  505. 19:13models here so power models roughly were
  506. 19:17about 5 40 billion
  507. 19:18parameters and simply with this Chain of
  508. 19:21Thought kind of a skill you're already
  509. 19:22like beating state of the
  510. 19:25art cool um
  511. 19:29so yeah so I showed you examples of
  512. 19:32Chain of Thought reasoning to to this
  513. 19:34point where you go through a reasoning
  514. 19:35chain but you can be even slightly
  515. 19:37smarter than that you might not even
  516. 19:39need to show them any examples you just
  517. 19:41need to trick them into thinking about
  518. 19:43what to do
  519. 19:44next um so yeah this s this idea emerged
  520. 19:49in this paper called where you let's
  521. 19:51think step by step where instead of even
  522. 19:54showing an example you just start your
  523. 19:56answer with let's think step by step and
  524. 20:01that's it like I mean the model will
  525. 20:03start reasoning about the answer itself
  526. 20:05instead of just like autoc completing to
  527. 20:07an answer and you get something like
  528. 20:11this
  529. 20:12so maybe you don't even need to show any
  530. 20:15examples like you can probably induce
  531. 20:17the reasoning Behavior through zero shot
  532. 20:18Behavior as well and again um what the
  533. 20:22final numbers look like is like compared
  534. 20:24to zero shot performance that we got
  535. 20:26from essentially autoc comp completing
  536. 20:29this zero shot Chain of Thought
  537. 20:31substantially improves the performance
  538. 20:32so you go from like 17.7 to 78.7 it's
  539. 20:35still worse than still putting like
  540. 20:37examples of reasoning and multi-shot few
  541. 20:40shot Chain of Thought as well but you
  542. 20:42can see like how much it improves the
  543. 20:44performance simply by asking you to
  544. 20:45let's think by step by step and maybe
  545. 20:48this is like a lesson that interacting
  546. 20:50with these models is like when you
  547. 20:52interact with these models you might not
  548. 20:54get the exact desired behavior from
  549. 20:57these models up front but often like
  550. 20:59these models are capable of doing the
  551. 21:01behavior that you might want and often
  552. 21:05you have to think about how to induce
  553. 21:06that behavior such that and the right
  554. 21:09way to think perhaps is like what is the
  555. 21:10pre-training data what is the data on
  556. 21:12the internet it might have seen which
  557. 21:13induces a similar Behavior to the kind I
  558. 21:15want and you probably want to like think
  559. 21:18about that and then induce these kinds
  560. 21:20of behaviors from those
  561. 21:24models and yeah like I mean um you know
  562. 21:28we we hand designed some of these
  563. 21:30prompts you can also like get an llm to
  564. 21:32design these prompts as well there's
  565. 21:34like recursive self-improving ideas here
  566. 21:37um that happen and you can bump up the
  567. 21:38performance a little bit
  568. 21:42more cool so what we have seen so far is
  569. 21:46that as models get stronger and stronger
  570. 21:49you can start forcing them to do your
  571. 21:50task zero shot or with few short
  572. 21:53examples and you can trick them into
  573. 21:55thinking what task you want them to
  574. 21:56solve
  575. 21:59but the downside is that there's only so
  576. 22:01much you can fit into context this might
  577. 22:04not be very true anymore models like
  578. 22:06becoming increasingly larger context but
  579. 22:09it's still somewhat unsatisfactory to
  580. 22:11think you have to trick the model into
  581. 22:13doing your task rather than like it just
  582. 22:15doing the task you wanted to do and
  583. 22:18potentially like I mean going forward
  584. 22:20like you probably still want to fine
  585. 22:22tune these models for more and more
  586. 22:23complex tasks and that's where we're
  587. 22:26going to go forward in this
  588. 22:28um next section we're going to cover is
  589. 22:30instruction fine tuning and the general
  590. 22:34idea we have right now is that as we
  591. 22:37talked about pre-training is not about
  592. 22:39assisting users it is about predicting
  593. 22:41the next token now you can trick it into
  594. 22:44assisting users and uh following the
  595. 22:47instruction you wanted to but in general
  596. 22:49that's not what it retrained it for and
  597. 22:51this is an example of where if you ask
  598. 22:53GPD 3 pretty strong model to explain
  599. 22:55like moon landing to a six-year-old in a
  600. 22:57few sentences and it will follow up with
  601. 22:59more questions about what a 60-year-old
  602. 23:01might want this is not what you wanted
  603. 23:03the model to do right so the general
  604. 23:07term that people use these days is that
  605. 23:08they're not aligned with user intent and
  606. 23:12the next sections that we're going to
  607. 23:13cover are going to talk about how to
  608. 23:14align it with the user intent so that
  609. 23:16they don't have to trick the model into
  610. 23:17whatever uh we wanted to
  611. 23:20do and this is a kind of like desired
  612. 23:23completion we want at the end of
  613. 23:24instruction tuning um and yeah
  614. 23:29how do we get from those pre-trained
  615. 23:32models to models which can respond to
  616. 23:33user
  617. 23:34intent um again um I hope this was
  618. 23:38covered somewhere in the class the
  619. 23:39general idea of pre-training and
  620. 23:41fine-tuning um but what you have
  621. 23:43probably seen thus far is that you
  622. 23:45pre-train on a lot of different language
  623. 23:47task uh on data but then you find on
  624. 23:50your specific task so you're taking the
  625. 23:53same decoder only models and you're fine
  626. 23:56tuning to some task with very little
  627. 23:58amount of data the thing that is going
  628. 24:00to be different now is not that we're no
  629. 24:02longer fine tuning on a little amount of
  630. 24:04data we're going to fine tune on many
  631. 24:05many different tasks and we're going to
  632. 24:07just try to put them into a single
  633. 24:10usable um ux for users and this is where
  634. 24:14fine tuning or instruction fine tuning
  635. 24:16comes
  636. 24:19in cool um so again the recipe is not
  637. 24:23very very complicated here um we're
  638. 24:25going to collect a lot of examples of
  639. 24:27instruction and output Pairs and the
  640. 24:29instructions are going to rrange over
  641. 24:30several task different forms um there's
  642. 24:33going to be question answering they're
  643. 24:34going to be summarization translation
  644. 24:36code reasoning and so on and we're going
  645. 24:38to collect a lot of examples uh related
  646. 24:41to all those tasks and the idea is like
  647. 24:44I mean we'll train on instruction and
  648. 24:45output pairs exactly with them and then
  649. 24:48we're going to evaluate on some unseen
  650. 24:50tasks as well so this is a general
  651. 24:53Paradigm of instruction fine tuning and
  652. 24:57again it's the same idea which we
  653. 24:59explored in pre-training is that data
  654. 25:01plus scale is really important and these
  655. 25:04days like a mean you start off with like
  656. 25:06one task you're now extending it over
  657. 25:08thousands of thousands and thousands of
  658. 25:10tasks with like three million plus
  659. 25:11examples and this is generally like a
  660. 25:13broad range of tasks that you might see
  661. 25:14in instruction fine tuning data
  662. 25:16sets and yeah you might even think of
  663. 25:19like why are we calling it fine tuning
  664. 25:20anymore like it's almost starting to
  665. 25:22look like pre-training um but yeah we
  666. 25:25can these are just terms um so so you
  667. 25:28can decide whatever you are comfortable
  668. 25:30with um so yeah we we get this like huge
  669. 25:35instruction data set we finder model the
  670. 25:37next question is like how do we evaluate
  671. 25:39these data sets um now I think you guys
  672. 25:43will see another lecture on evaluation
  673. 25:45so I don't want to like dive too deep
  674. 25:46into this but generally evaluation of
  675. 25:49these language models is an extremely
  676. 25:51tricky topic um there's a lot of biases
  677. 25:53that you need to deal with and a lot of
  678. 25:55this will be covered but some more
  679. 25:57recent progess on this is like we are
  680. 25:59starting to curate these like really
  681. 26:01large benchmarks uh like mlu where the
  682. 26:05models are tested on a broad range of
  683. 26:06diverse knowledge and this is just one
  684. 26:09example which is and these are the
  685. 26:11topics that you will see and just to
  686. 26:14give some intuition of what the examples
  687. 26:16in these evaluation look like um under
  688. 26:18astronomy you might be asked what is
  689. 26:20true for type 1 a supernova or you might
  690. 26:23be asked some questions about biology
  691. 26:25and there's a huge host of tasks for
  692. 26:27this and you can typically like these
  693. 26:30are multi- choice questions and you can
  694. 26:31ask the model to answer the question if
  695. 26:33they're instruction fine tuned already
  696. 26:34hopefully they can like simply answer
  697. 26:36the question but you can also uh Chain
  698. 26:38of Thought prompt these questions or few
  699. 26:40short promp these questions too and
  700. 26:43recently there's been a huge amount of
  701. 26:45progress uh on this Benchmark what
  702. 26:48people have observed is like more and
  703. 26:49more pre-training on more and more data
  704. 26:51and larger models is simply just like
  705. 26:52climbing up these um the number on this
  706. 26:55so 90% is often seen as a benchmark Mark
  707. 26:58number that these model wanted to cross
  708. 27:00because it's roughly like human level
  709. 27:02knowledge or understanding and recently
  710. 27:05the Gemini models purply cross this
  711. 27:09number so yeah go
  712. 27:12ahead is isn't this like the entire sort
  713. 27:15of thing all over again right like imag
  714. 27:18at some point you're like okay maybe my
  715. 27:19methods are too like too fine tuned
  716. 27:22implicitly on on the image bi and isn't
  717. 27:25something like that happening here as
  718. 27:27well uh um yes I think this is a tricky
  719. 27:30topic because a lot of the models often
  720. 27:33there's this idea about whether your
  721. 27:35test sets are leaking into your training
  722. 27:37data set and there's are huge concerns
  723. 27:39about that it's a perfectly valid
  724. 27:41question to ask how do we even evaluate
  725. 27:44this is why evaluation is actually very
  726. 27:45tricky but one General thing to be
  727. 27:48careful about is like at some point like
  728. 27:50it doesn't matter what your trained test
  729. 27:51is if the models are generally useful if
  730. 27:54their models are doing useful stuff like
  731. 27:56does it matter like how if your test if
  732. 27:59you train on everything you care about
  733. 28:01and it does well on it like does it
  734. 28:03matter so yeah um again we still need
  735. 28:08better ways to evaluate the models um
  736. 28:10and to understand what methods are doing
  737. 28:12and how they're if they're improving the
  738. 28:14model or not but at some point like that
  739. 28:16those boundaries start to like be less
  740. 28:22important cool so massive progress on
  741. 28:24this Benchmark starting with gpd2 and
  742. 28:26like we're roughly at 90% which to the
  743. 28:29point where these benchmarks are
  744. 28:30starting to become unclear if like
  745. 28:32improvements on these are actually
  746. 28:33meaningful or not um in fact like most
  747. 28:37of the times when the models are wrong
  748. 28:39like you might often find that the
  749. 28:41question itself was unclear or ambiguous
  750. 28:43so all evaluation benchmarks have a
  751. 28:46certain limited utility to
  752. 28:48them so yeah um going to go over like
  753. 28:52another evaluation example of how this
  754. 28:54recipe like changes things so T5 models
  755. 28:58were instruction fine tuned on a huge
  756. 28:59number of tasks and another Trend to or
  757. 29:02which I think will be the theme across
  758. 29:04this lecture is that as your models
  759. 29:06become larger as they're trained on more
  760. 29:07data they become more and more
  761. 29:09responsive to your task information as
  762. 29:11well so what you will observe here is
  763. 29:13like as the number of parameters ex
  764. 29:15increase we have like T5 small FL T5
  765. 29:18small and we go up to 11 billion
  766. 29:20parameters where where we have T5 XXL
  767. 29:23you'll see that the Improvement actually
  768. 29:25improves like going from a pre into an
  769. 29:28instruction model the instruction model
  770. 29:30is all the more better at following
  771. 29:32instructions so the difference is plus
  772. 29:356.1 and goes to plus 26.6 as the models
  773. 29:37become larger so this is another very
  774. 29:40encouraging Trend that you probably
  775. 29:42should train on a lot of data with a lot
  776. 29:44of compute and you know pre-training
  777. 29:48just keeps on
  778. 29:50giving
  779. 29:52so yeah um you I hope you guys get a
  780. 29:56chance to like play with a lot of these
  781. 29:57models I I think you already hopefully
  782. 29:59are uh but yeah before instruction fine
  783. 30:02tuning something when you're asked a
  784. 30:04question related to disambiguation QA um
  785. 30:07you get something like this and it
  786. 30:09doesn't actually follow the let's think
  787. 30:11by step byep instruction very clearly
  788. 30:15but after instruction fine tuning it is
  789. 30:16able to answer the question
  790. 30:19here and yeah like more recently people
  791. 30:22have been like researching into like
  792. 30:24what the instruction tuning data set
  793. 30:25should look like there's a huge plora of
  794. 30:28instruction tuning data sets now
  795. 30:29available like this is just a
  796. 30:30representative diagram and there's a h
  797. 30:32open source Community developing around
  798. 30:34these as
  799. 30:35well um some high level lessons that we
  800. 30:38have learned from this
  801. 30:41is one lesson that I think might be
  802. 30:44interesting is that we can actually use
  803. 30:45really large strong models to generate
  804. 30:48some of the instruction tuning data to
  805. 30:49train some of our smaller models so take
  806. 30:52your favorite model right now gbd4 maybe
  807. 30:54or maybe Claud or whatever and you can
  808. 30:56get it to answer some question s and
  809. 30:58generate instruction outut pairs for
  810. 31:01training your open source or smaller
  811. 31:03model and that actually is a very
  812. 31:04successful recipe so instead of getting
  813. 31:06a human to collect all the instruction
  814. 31:08output pairs or getting humans to
  815. 31:10generate the answers you can get bigger
  816. 31:12models to generate the answers as well
  817. 31:14so that's number one thing that like has
  818. 31:16recently emerged another thing that are
  819. 31:19is being emerged or is like being
  820. 31:21discussed is how much data do we need I
  821. 31:23talked about millions of examples but
  822. 31:25like people have often found that if you
  823. 31:26have really high quality example you can
  824. 31:28get away with thousand examples as well
  825. 31:30so this is the paperless as more for
  826. 31:32alignment and this is still an active
  827. 31:34area of research on how like data
  828. 31:36scaling and instruction tuning affects
  829. 31:37the final model
  830. 31:39performance and yeah crowdsourcing these
  831. 31:42models can be effective as well so there
  832. 31:45are very cool benchmarks that are
  833. 31:46emerging like open Assistant um yeah a
  834. 31:49lot of activity in the field and
  835. 31:52hopefully like a lot more progress as we
  836. 31:54go on yes um a question sort of in the
  837. 31:58spirit of this LMA paper uh doesn't like
  838. 32:02code or like I don't know like math word
  839. 32:06problems have this desired structure so
  840. 32:08like shouldn't we just like be training
  841. 32:10code models and doing like some English
  842. 32:13stuff and then just be like okay this is
  843. 32:15the best reasoning we can get at some
  844. 32:17point right cuz like C code has the
  845. 32:19structure where where where like you're
  846. 32:22going sort of step by step and you're
  847. 32:24sort of thinking in in some way in like
  848. 32:27a so breaking down a concept into a
  849. 32:29smaller so you can consider like cod
  850. 32:31have like very high value tokens so
  851. 32:33maybe like just doing so I think there's
  852. 32:37again pre-training is a whole Dark Art
  853. 32:39that I am not completely familiar with
  854. 32:41but um code actually ends up being
  855. 32:44really useful in pre-training mixtures
  856. 32:46and people do like up with code data
  857. 32:48quite a lot um similarly like I mean but
  858. 32:53it depends upon what the users are going
  859. 32:54to use the models for right um some
  860. 32:56people might use it for Cotes some
  861. 32:57people people might do for reasoning but
  862. 32:59that's not the only task we care about
  863. 33:01as you might see later on in the next
  864. 33:02step we'll discuss this as well is that
  865. 33:05people often use these models for
  866. 33:06Creative task they wanted to write a
  867. 33:08story uh they wanted to generate a movie
  868. 33:10script or so on and I don't know if like
  869. 33:13necessarily training on reasoning only
  870. 33:15tasks would help with that so go ahead
  871. 33:18yeah would you explain like there there
  872. 33:20exists like some data distribution which
  873. 33:22is like high value for Creative tasks
  874. 33:28yes like I mean it seems like um a lot
  875. 33:32of PE people write about stories and
  876. 33:34everything on the internet all the time
  877. 33:36like which is not code and sometimes
  878. 33:39like there's this idea of hallucinations
  879. 33:40as well in this like field but you can
  880. 33:42often think like hey like creativity
  881. 33:44might be a byproduct of hallucinations
  882. 33:47as well so I don't know what exact data
  883. 33:50would like lead to like more creative
  884. 33:52models but generally like there's a lot
  885. 33:54of data or a lot of stories that are
  886. 33:56written on the internet which allows the
  887. 33:57model to be
  888. 33:59creative yeah but I don't know if I have
  889. 34:01a specific answer to the
  890. 34:03question cool so we discussed
  891. 34:05instruction fine tuning um very simple
  892. 34:08and very straightforward there's like no
  893. 34:10complicated algorithms here just collect
  894. 34:12a lot of data and then you can start
  895. 34:14leveraging the performance at scale as
  896. 34:16well like as models become better these
  897. 34:18models also become more easily
  898. 34:21specifiable and they become more
  899. 34:23responsive to task as well we're going
  900. 34:25to discuss some limitations and I think
  901. 34:27this is like really important to
  902. 34:28understand why we are going to optimize
  903. 34:29for human
  904. 34:32preferences cool so we talked a bit
  905. 34:35about this like instruction fine tuning
  906. 34:37is necessarily contingent on humans
  907. 34:40labeling the data now
  908. 34:44humans it's expensive to collect this
  909. 34:46data especially as the questions become
  910. 34:48more and more complex you want to answer
  911. 34:50questions about what which may be at
  912. 34:52physics PhD level or things to that
  913. 34:54effect these things become increasingly
  914. 34:57expensive to collect
  915. 34:59so yeah this is maybe like perhaps
  916. 35:01obvious like collecting data
  917. 35:02pre-training does not require any
  918. 35:04specific data you scrape data of the web
  919. 35:06but for instruction fing you probably
  920. 35:08need to recruit some people to write
  921. 35:09down answer to your instructions so this
  922. 35:11can become very expensive very quickly
  923. 35:14but there's more limitations to this as
  924. 35:16well and we we just discussing this like
  925. 35:18there are open-ended tasks related to
  926. 35:20creativity that don't really have like
  927. 35:22an exact correct answer to begin with so
  928. 35:25how do you generate the right answer to
  929. 35:27the kind of a
  930. 35:30question and yeah like language modeling
  931. 35:34inherently like penalizes all token
  932. 35:36level mistakes equally um this is what
  933. 35:38super fine supervised fine tuning does
  934. 35:40as well but often like not all mistakes
  935. 35:42are the same so this is an example where
  936. 35:46you're trying to do this prediction task
  937. 35:47Avatar is a fantasy TV show and perhaps
  938. 35:50you can see like I mean calling it an
  939. 35:53adventure TV show is perhaps okay but
  940. 35:56calling it a musical May be like a much
  941. 35:59worse mistake but both these mistakes
  942. 36:01are penalized
  943. 36:04equally and I think one General aspect
  944. 36:07which is like more becoming increasingly
  945. 36:09relevant is that the humans that you
  946. 36:11might ask might not generate the right
  947. 36:13or the highest quality answer your
  948. 36:15models are becoming increasingly
  949. 36:17competitive and you want in some sense
  950. 36:19you're going to be limited by how high
  951. 36:22quality the answer um humans can
  952. 36:24generate but often I find that the
  953. 36:27models are generating better and better
  954. 36:29answers so do we really want to keep
  955. 36:31relying on humans to write down the
  956. 36:32answers or do we want to like somehow go
  957. 36:34over
  958. 36:36that so these are the three problems we
  959. 36:41have talked about um with instruction
  960. 36:43fine tuning and we made a lot of
  961. 36:46progress with this but this is not how
  962. 36:47we got Char
  963. 36:49GPT um and one high level problem here
  964. 36:53is that even though when even when we
  965. 36:55are instruction fine tuning there is
  966. 36:57still a huge mismatch
  967. 36:59between the end goal is to optimize for
  968. 37:02human preferences generate an output
  969. 37:04that a human might like and we're still
  970. 37:08doing prediction kind of tasks where
  971. 37:09we're predicting the next token but now
  972. 37:11in a more curated data set so that's
  973. 37:13still a bit of a mismatch going on here
  974. 37:15and it's not exactly what we want to
  975. 37:18do hopefully like I mean I'm going to
  976. 37:20take a second here to pause because this
  977. 37:21is important to understand the next
  978. 37:23section and if there's any
  979. 37:25questions feel free to ask so is this
  980. 37:28step uh still taken as a as a first step
  981. 37:33or we discard this it's a good question
  982. 37:36so um I think this is still one of the
  983. 37:39more important steps that you take
  984. 37:41before taking the next step but people
  985. 37:43are trying to like remove the step Al
  986. 37:45together and jump directly to the next
  987. 37:47step so there's work emerging on that
  988. 37:49but yeah and this is still a very
  989. 37:51important step before we do the next
  990. 37:53step
  991. 37:57go ahead is PR two also present in
  992. 38:00pre-training uh and if so how how do you
  993. 38:03avoid that ver just by having a lot of
  994. 38:06data um yeah that's a great question uh
  995. 38:08there's two diff there's one difference
  996. 38:10one major difference on pre-training um
  997. 38:13pre-training covers a lot more text so
  998. 38:16um just for context like I mean as we
  999. 38:18talked about it's pre-training is
  1000. 38:20roughly 15 trillion tokens whereas like
  1001. 38:23supervised instruction fine tuning might
  1002. 38:24be somewhere on the order of millions to
  1003. 38:26billions of tokens so it's like few
  1004. 38:28orders of magnitude lower typically
  1005. 38:30you'd only see one answer for a specific
  1006. 38:32instruction but during pre-training
  1007. 38:34you'll see multiple text and multiple
  1008. 38:36completions for a same kind of a prompt
  1009. 38:39um now that's good because when you see
  1010. 38:41multiple answers or completions during
  1011. 38:42pre-training you sort of start to weigh
  1012. 38:44different answers you start to put
  1013. 38:46probability Mass on different kind of
  1014. 38:48answers or completions but instruction
  1015. 38:50fineing might force you to put and wait
  1016. 38:52on only one
  1017. 38:53answer does it okay but generally yeah
  1018. 38:57like I mean this is a problem with both
  1019. 38:59the stages you're
  1020. 39:01right anything
  1021. 39:04else
  1022. 39:07cool so as this whole thing alludes to
  1023. 39:11we're going to start to attempt to
  1024. 39:14satisfy human preferences directly we're
  1025. 39:16no longer going to like try to like get
  1026. 39:18humans to generate some data and try to
  1027. 39:20do some kind of a token level prediction
  1028. 39:21loss we're going to try to optimize for
  1029. 39:24human preferences directly and that is
  1030. 39:27uh the general field of rlf and that's
  1031. 39:30the final step in typically getting a
  1032. 39:32model like CH
  1033. 39:33GPD so we talked about how collecting
  1034. 39:36demonstration is expensive and there's
  1035. 39:37still a broad mismatch between the LM
  1036. 39:39objective and human preferences and now
  1037. 39:41we're going to try and optimize for
  1038. 39:42human preferences
  1039. 39:44directly so what ises optimizing for
  1040. 39:47human preferences even mean um to like
  1041. 39:50concretely establish that let's go
  1042. 39:52through like a specific example in mind
  1043. 39:54which is
  1044. 39:55summarization um we want to train a
  1045. 39:58model to be better at
  1046. 39:59summarization and we want to satisfy
  1047. 40:01human preferences so let's imagine that
  1048. 40:03a human is able to prescribe a reward
  1049. 40:05for a specific summary let's just
  1050. 40:07pretend there is a reward function you
  1051. 40:09and I can assign say like reward this is
  1052. 40:11plus one this is minus one or something
  1053. 40:13to that
  1054. 40:17effect okay um so in this specific case
  1055. 40:22we have this input X which uh which is
  1056. 40:25about an earthquake in San Francisco so
  1057. 40:27this is news article that we want to
  1058. 40:28summarize
  1059. 40:30and let's pretend that we get these
  1060. 40:33rewards and we want to optimize this so
  1061. 40:36we we get one summary y1 which gives us
  1062. 40:39an earthquake hit and so on and we
  1063. 40:40assign a reward of 8.0 and another
  1064. 40:43summary which gives us a reward of
  1065. 40:451.2 generally speaking like the
  1066. 40:47objective that we want to set up is
  1067. 40:49something of the following form where we
  1068. 40:51want to take our language model P Theta
  1069. 40:54which generates a completion y uh given
  1070. 40:57an input X and we want to maximize the
  1071. 41:00reward of rxy where X is the input and Y
  1072. 41:04is the output summary in this specific
  1073. 41:07task and maybe like just to like really
  1074. 41:11concretely point out something here this
  1075. 41:13is different from everything that we
  1076. 41:15have done in one very specific way um we
  1077. 41:18are sampling from the model itself in
  1078. 41:21the bottom term if you see like we're
  1079. 41:23using Y from P Theta everything we've
  1080. 41:25seen so far the data is sampled from
  1081. 41:27some other source either during
  1082. 41:28pre-training either in supervised fine
  1083. 41:30tuning and we're maximizing the log
  1084. 41:32likelihood of those tokens but now we're
  1085. 41:35explicitly sampling from our model and
  1086. 41:37optimizing potentially a
  1087. 41:38non-differentiable
  1088. 41:41objective
  1089. 41:42cool so broadly the rlf pipeline looks
  1090. 41:46something like this and first step is
  1091. 41:48still instruction tuning something we
  1092. 41:49have seen up until now where we take our
  1093. 41:52pre-trained model we instruction tune on
  1094. 41:54a large collection of tasks and we get
  1095. 41:57some something which starts responding
  1096. 41:58to our desired intent or
  1097. 42:01not but there are two more steps after
  1098. 42:03this which are typically followed in
  1099. 42:05creating something like instruct gbt the
  1100. 42:07first step is estimating some kind of a
  1101. 42:09reward model something which tells us
  1102. 42:11given an instruction how much would a
  1103. 42:13human like this answer or how much would
  1104. 42:15a human hate this answer so we looked at
  1105. 42:18something like this earlier but I didn't
  1106. 42:19talk about how do we even get something
  1107. 42:21like that that's the second step and
  1108. 42:23then we take this reward model and we
  1109. 42:25optimize it through the optimiz ation
  1110. 42:27that I suggested earlier so the
  1111. 42:28maximizing the expected reward under
  1112. 42:31your language
  1113. 42:32model and we're going to go over a lot
  1114. 42:35over in the second and third
  1115. 42:37steps so the first question we want to
  1116. 42:39answer is how do we even get like a
  1117. 42:40reward model what about what humans are
  1118. 42:43going to like like this is a very IL
  1119. 42:46defined problem generally speaking
  1120. 42:49so there's there's two problems here
  1121. 42:52that we're going to address first is a
  1122. 42:53human in the loop is expensive so let's
  1123. 42:55say like if I ask a model to like
  1124. 42:57generate an answer and then I get a
  1125. 42:59human to label with some kind of a score
  1126. 43:02I'm doing this over millions of
  1127. 43:03completions that is not very scalable I
  1128. 43:07I don't want to sit around and label
  1129. 43:08millions of examples
  1130. 43:10so this is very easy like we're in a
  1131. 43:13machine learning class so what are we
  1132. 43:15going to do what we're going to do is
  1133. 43:16we're going to train something which
  1134. 43:18predicts what a human would like or what
  1135. 43:19a human might not like and this is
  1136. 43:22roughly um this is essentially a machine
  1137. 43:25learning problem where we take these
  1138. 43:26Rewards scor scores and we try to train
  1139. 43:28a reward model to predict given an input
  1140. 43:31and output what the reward scores would
  1141. 43:33look
  1142. 43:34like simple simple machine learning
  1143. 43:37regression style problem uh you might
  1144. 43:38have seen this
  1145. 43:40earlier
  1146. 43:43cool now there's a bigger problem here
  1147. 43:46and sorry go ahead one so do we use like
  1148. 43:49I don't know like just embedding
  1149. 43:52withier we use a real language model to
  1150. 43:55do that um that's a good question
  1151. 43:58generally like what we do is like we
  1152. 44:00still typically need reward models where
  1153. 44:02they need to be able to understand the
  1154. 44:04text really well so like bigger models
  1155. 44:06and like they're typically initialized
  1156. 44:07from the language model that you trained
  1157. 44:09pre-trained as well so you typically
  1158. 44:11start with the pre-trained language
  1159. 44:12model and do some kind of prediction
  1160. 44:14that we'll talk about and they'll give
  1161. 44:16you a
  1162. 44:17score how do you if you're doing that
  1163. 44:20how do you separate X and Y how does the
  1164. 44:22language model know which part it
  1165. 44:24doesn't need to it can put the X and Y
  1166. 44:27like it only sees X and Y as an input so
  1167. 44:29it doesn't need to T typically see it
  1168. 44:31separated it's just going to predict a
  1169. 44:33score at the end okay yeah the X and Y
  1170. 44:35is more for notational convenience here
  1171. 44:38because for us X and Y are different X
  1172. 44:40is a question user asked and Y is
  1173. 44:42something the model generated but you
  1174. 44:44shove the whole thing into you shove the
  1175. 44:46whole thing into yes
  1176. 44:48cool now this is the bigger problem here
  1177. 44:51and human judgments are very noisy we
  1178. 44:54have talked about we want to assign a
  1179. 44:55score to a completion this is something
  1180. 44:58that's like extremely non-trivial to do
  1181. 45:00so if I give you a summary like this
  1182. 45:02what score are you going to assign on a
  1183. 45:04scale of 10 if you ask me on different
  1184. 45:07days I'll give a different answer first
  1185. 45:08of all but across humans itself this
  1186. 45:12this number is not calibrated in any
  1187. 45:14meaningful way so you could assign
  1188. 45:17number of 4.1 6.6 and different humans
  1189. 45:19would just simply assign different
  1190. 45:20scores and there are ways to address
  1191. 45:23this you can like calibrate humans you
  1192. 45:24can give them a specific rubric you can
  1193. 45:26like talk to them but it's a very
  1194. 45:28complicated process and like still like
  1195. 45:30there's a lot of room for judgment which
  1196. 45:32is not typically very nice for training
  1197. 45:34a model like this if your labels can
  1198. 45:37vary a lot it's just hard to
  1199. 45:40predict so the way this is addressed is
  1200. 45:44that instead of trying to predict the
  1201. 45:45reward label directly you actually want
  1202. 45:47to set up a problem in a slightly
  1203. 45:49different way what is something much
  1204. 45:50easier for humans to do is give them two
  1205. 45:53answers or maybe many answers and tell
  1206. 45:55them ask them which one is better so
  1207. 45:58this is where the idea of asking humans
  1208. 46:01to rank answers comes in so if I give
  1209. 46:04you a whole news article and ask you
  1210. 46:07which summary is better you might be
  1211. 46:08able to give me a ranking that oh this
  1212. 46:10second summary is the worst but the
  1213. 46:12first one is better and the third one is
  1214. 46:13somewhere in the middle between those
  1215. 46:14two so you get like a ranking which
  1216. 46:16gives you um preference over
  1217. 46:19summaries and hopefully like I mean you
  1218. 46:22can see like the idea that is important
  1219. 46:24here is that even when we have some kind
  1220. 46:27of a consistent utility function even
  1221. 46:29when I have it's much easier to compare
  1222. 46:32to something and know that which is
  1223. 46:33better than this rather than ascribing
  1224. 46:34it an arbitrary number on a scale and
  1225. 46:37that's why the signal from something
  1226. 46:39like this is a lot
  1227. 46:41better now how do we get like we talked
  1228. 46:44about we need like we get this kind of a
  1229. 46:46preference data and now we need some
  1230. 46:47kind of a reward score out of this and
  1231. 46:50we shove in like our input we shove in a
  1232. 46:53summary as well and we still need to get
  1233. 46:54a score out of this but it's not clearly
  1234. 46:56obvious to me like how do we take this
  1235. 46:58data and convert it into that kind of
  1236. 47:00score
  1237. 47:02um incomes are pretty good friends named
  1238. 47:05Bradley Terry um and essentially like
  1239. 47:10there's a lot of study in like many in
  1240. 47:12economics and like psychology which
  1241. 47:14basically tries to model how humans make
  1242. 47:18decisions in specific case like this
  1243. 47:20Brad lary model essentially says that a
  1244. 47:23probability that a human chooses answer
  1245. 47:25y1 over y two is based on the difference
  1246. 47:30between the rewards that humans assign
  1247. 47:33internally and then you take a sigmoid
  1248. 47:35around it so if you have looked at
  1249. 47:37binary classification before uh the
  1250. 47:40logic is simply the difference between
  1251. 47:41the reward of some y1 minus Y2 or the
  1252. 47:45difference between the winning
  1253. 47:47completion and the losing
  1254. 47:51completion is everybody with me till
  1255. 47:53this point
  1256. 47:57so the idea is that like if you have a
  1257. 48:00data set where given y1 and Y2 where y1
  1258. 48:03is a winning completion and we have a
  1259. 48:05winning completion YW and a losing
  1260. 48:07completion y l um the winning completion
  1261. 48:10should score higher than the losing
  1262. 48:12completion go ahead sorry what is J is
  1263. 48:15that a log or like sorry what what like
  1264. 48:19what is the type of J like this number
  1265. 48:21here that we're getting as the
  1266. 48:23expectation is it a log prop or what is
  1267. 48:25it it's an log prop so it will be a
  1268. 48:28scaler at the end sigmoid is so you're
  1269. 48:31taking the let's say you have a reward
  1270. 48:33model which gives a
  1271. 48:34score R1 to like YW and R2 to y l you
  1272. 48:39subtract that number you get another
  1273. 48:40number you put it into sigmoid and you
  1274. 48:42get a probability because sigmoid will
  1275. 48:45convert a logit into probability and
  1276. 48:47then you take a logarithm of that and
  1277. 48:50you take the expectation of everything
  1278. 48:51and you get this final number which
  1279. 48:53tells you how good your reward model is
  1280. 48:55doing on the entire data set
  1281. 48:58so like a good model of humans should
  1282. 48:59behave like this a good model of humans
  1283. 49:02would um score very low here so it would
  1284. 49:05generally assign a higher reward to the
  1285. 49:07winning completion and generally assign
  1286. 49:09a lower reward to the losing
  1287. 49:12completion
  1288. 49:14cool the math is just beginning so um
  1289. 49:18hold on to your seats
  1290. 49:21um cool so now let's see where we are we
  1291. 49:24have a pre-trained model p p d y given X
  1292. 49:28and we got this like fancy reward model
  1293. 49:30which tells us that hey how we have a
  1294. 49:32model of humans and it can tell us which
  1295. 49:34instruction which answer they like and
  1296. 49:36which in answer did not
  1297. 49:38like now to do rlf generally like I mean
  1298. 49:42we have discussed what this will look
  1299. 49:43like uh we will copy our pre-train model
  1300. 49:46or instruction tune model and we'll
  1301. 49:48optimize the parameters for those models
  1302. 49:51and I suggested that the param objective
  1303. 49:53that we want to optimize is the expected
  1304. 49:56reward when we sample completions from P
  1305. 49:59Theta and we're going to optimize our
  1306. 50:02learned reward model instead of like the
  1307. 50:03true reward model which humans would
  1308. 50:05have typically assigned do you guys see
  1309. 50:07any problem with
  1310. 50:09this
  1311. 50:11um is there something that's wrong here
  1312. 50:14or like that might go wrong if we do
  1313. 50:16something along these
  1314. 50:22lines go for itel
  1315. 50:28it might collapse yes okay um but
  1316. 50:31generally at least from my intuition
  1317. 50:32like if you're ever doing something and
  1318. 50:34you have you're optimizing some learned
  1319. 50:36metric I'd be very careful because
  1320. 50:39typically our loss functions are very
  1321. 50:40clearly defined but here my reward model
  1322. 50:42is learned what when it's learned it
  1323. 50:44means it will have
  1324. 50:46errors yes so it's going to be trained
  1325. 50:49on some distribution it will generalize
  1326. 50:51as well but it will have errors and when
  1327. 50:54you're optimizing against a learn model
  1328. 50:57it will tend to hack the reward model so
  1329. 50:59it might exploit the reward model might
  1330. 51:02erroneously assign a really high score
  1331. 51:04to a really bad completion if your
  1332. 51:06policy learns or if your language model
  1333. 51:08learns to do that it will completely
  1334. 51:10Hack That and start generating those
  1335. 51:11gibberish
  1336. 51:13completions
  1337. 51:15so just as a general machine learning
  1338. 51:17tip as well if you're optimizing a learn
  1339. 51:19metric be careful about what you're
  1340. 51:20optimizing and make sure that it's
  1341. 51:22actually
  1342. 51:23reliable um and the way and this is
  1343. 51:27obviously not desirable like I mean if
  1344. 51:28you start optimizing this objective
  1345. 51:30you're going to converse to gibberish
  1346. 51:31language models very very quickly so
  1347. 51:33typically what people do is that you
  1348. 51:35want to add some kind of a penalty that
  1349. 51:37like avoids it drifting too far from its
  1350. 51:40initialization and why do we want to do
  1351. 51:42that like if it cannot Drift from too
  1352. 51:44far from its initialization we know the
  1353. 51:46initialization of the model is a decent
  1354. 51:47language model and we know it is not
  1355. 51:50necessarily satisfying this reward model
  1356. 51:52too much and we also know that like the
  1357. 51:54reward model is trained on a distrib
  1358. 51:56ution of completions where the um
  1359. 51:59initial model is so typically we when we
  1360. 52:02talk about training this reward model we
  1361. 52:04have trained on certain completions
  1362. 52:05which are sampled from this initial
  1363. 52:07distribution so we know the reward model
  1364. 52:09will be somewhat reliable in that
  1365. 52:10distribution so we're just going to
  1366. 52:12Simply add a penalty which tells us that
  1367. 52:15you should not drift too far away from
  1368. 52:17the initial distribution and just to go
  1369. 52:20over this we want to maximize the
  1370. 52:22objective where we have RMF which is our
  1371. 52:25learned one model but we're going to add
  1372. 52:29this term beta log ratio and the ratio
  1373. 52:31is our the model we're optimizing P
  1374. 52:33Theta and our initial model PP PT and
  1375. 52:37what this says is that if we assign a
  1376. 52:39much higher probability to certain
  1377. 52:41completion as compared to our pre-train
  1378. 52:43model you're going to add an
  1379. 52:44increasingly large penalty to
  1380. 52:46it and simply you're paying a price for
  1381. 52:49drifting too far from initial
  1382. 52:50distribution if you guys have taken like
  1383. 52:53machine learning this the expectation of
  1384. 52:55this quantity can is exactly the cbak LI
  1385. 52:58Li Divergence or k Divergence between P
  1386. 53:01Theta and PPT so you're penalizing
  1387. 53:04drifting between two distributions go
  1388. 53:07forhead question shouldn't you also do
  1389. 53:09this like add a penalty in the previous
  1390. 53:12version where you had to find huning or
  1391. 53:14is this only relevant for the RL HF
  1392. 53:18that's a good question so um I think
  1393. 53:21people do add some kinds of
  1394. 53:22regularization in fine tuning it's not
  1395. 53:25nearly not as critical when you're doing
  1396. 53:27this with RL like the incentive is to
  1397. 53:29exploit this reward model as well as
  1398. 53:32much as possible and we'll see examples
  1399. 53:34where like the Learned reward predicts
  1400. 53:37like it's doing really well but the true
  1401. 53:38reward models are completely garbage so
  1402. 53:41it's much more important in this
  1403. 53:47optimization cool
  1404. 53:49so now this assume does this course does
  1405. 53:53not assume background on reinforcement
  1406. 53:54learning so we're not going to go into
  1407. 53:56reinforce learning but I just want to
  1408. 53:57give a very high level intuition about
  1409. 53:59how this works and reinforcement
  1410. 54:01learning is not typically just used for
  1411. 54:03language model it's been applied to
  1412. 54:04several uh domains of Interest game
  1413. 54:07playing agents re um robotics developing
  1414. 54:11chip designs and so on and the interest
  1415. 54:15between like RL and model LMS it's also
  1416. 54:18like dates back to roughly like 2016 as
  1417. 54:20well but like it's been really
  1418. 54:22successful recently and especially with
  1419. 54:24the success of rlf
  1420. 54:27um the general idea is that we're going
  1421. 54:28to use our model that we're optimizing
  1422. 54:30to generate several completions for an
  1423. 54:32instruction um we're going to compute
  1424. 54:35the reward under our learned reward
  1425. 54:37model and then we're going to simply try
  1426. 54:39and like update the update our model to
  1427. 54:42increase the probability on the high
  1428. 54:44reward completions so when we sample a
  1429. 54:46model we'll see completions of varing
  1430. 54:48quality and we'll see some good
  1431. 54:49completions good summaries for our task
  1432. 54:51some bad summaries for our task and
  1433. 54:53we'll try to update our log
  1434. 54:54probabilities in a way such that uh the
  1435. 54:57reward for when you use a updated model
  1436. 55:00you're typically in the higher reward
  1437. 55:04region does the high level summary like
  1438. 55:06make
  1439. 55:07sense
  1440. 55:10cool and rhf is incredibly successful I
  1441. 55:13think this is a very good example of um
  1442. 55:15this is the same summarization example
  1443. 55:18and I think the key Point here is that
  1444. 55:22the performance improves by increasing
  1445. 55:24the model size for sure we have seen
  1446. 55:26this in many different example what you
  1447. 55:28can actually see is that even very small
  1448. 55:30models can outperform human completions
  1449. 55:33if you train it with with rlf and this
  1450. 55:36is exactly the result you see here the
  1451. 55:38reference summaries are human generated
  1452. 55:39and when you evaluate when you ask
  1453. 55:42humans which ones they prefer they often
  1454. 55:44prefer the model generated summary over
  1455. 55:46the human generated summary and this is
  1456. 55:47something you only observe with rlf even
  1457. 55:50at small scales and again the same
  1458. 55:51scaling phenomena still holds here
  1459. 55:53bigger models do become more responsive
  1460. 55:55but are of itself is very impactful
  1461. 56:00here
  1462. 56:02cool the problem with rlf is that it's
  1463. 56:05just incredibly complex like um I gave
  1464. 56:08you a very high level summary that's
  1465. 56:10like doesn't that there's whole courses
  1466. 56:12on this for a reason um so it just and
  1467. 56:15this image is not for you to understand
  1468. 56:17it's just completely to intimidate you
  1469. 56:20um
  1470. 56:21so um you want to fit a value function
  1471. 56:24to something there's you have to sample
  1472. 56:25the model a lot it can be sensitive to a
  1473. 56:28lot of hyperparameters so there's a lot
  1474. 56:30that goes on here and yeah um if you
  1475. 56:35start implementing an rlf pipeline it
  1476. 56:37can be very hard and this is the reason
  1477. 56:39why like a lot of rlf was restricted to
  1478. 56:41very very like high compute High
  1479. 56:43resource places and it was not very
  1480. 56:46accessible so what we're going to talk
  1481. 56:48about and cover in this course is
  1482. 56:49something called direct preference
  1483. 56:50optimization which is a much simpler
  1484. 56:52alternative to R LF and hopefully like
  1485. 56:54that's much more accessible but but
  1486. 56:56please bear with me there will be a lot
  1487. 56:58of math on here but the end goal of the
  1488. 57:00math is to make come up with a very
  1489. 57:02simple algorithm so hopefully like it's
  1490. 57:04um and feel free to stop me and ask me
  1491. 57:10questions you need in terms of like gbt
  1492. 57:134 versus three like how much do the
  1493. 57:15number of parameters in the base model
  1494. 57:17help with like need to reduce the number
  1495. 57:19of parameters or like in order sorry R
  1496. 57:22the number of like examples from humans
  1497. 57:25for RFS that work
  1498. 57:26well yeah that's a really good question
  1499. 57:28so generally speaking as the if you hold
  1500. 57:31the data set size constant and simply
  1501. 57:33increase the mod size it will improve
  1502. 57:35quite a lot sure but the nice thing is
  1503. 57:38that you can reuse the data and you can
  1504. 57:39keep adding data yeah uh as you keep
  1505. 57:42like scaling models up so typically like
  1506. 57:44nobody tries to like reduce the amount
  1507. 57:46of data collection yeah right you just
  1508. 57:47keep increasing both the things
  1509. 57:51out cool so we talked about rlf and the
  1510. 57:55current pipeline is some something like
  1511. 57:58um we train a reward model on the
  1512. 57:59comparison data that we have seen so far
  1513. 58:01and we're going to optimize we're going
  1514. 58:03to start with our pre-train our
  1515. 58:04instruction tune model and convert it
  1516. 58:05into an rlf model using the
  1517. 58:07reinforcement learning
  1518. 58:09techniques now the really the key idea
  1519. 58:11in direct preference optimization is
  1520. 58:13what if we could just simply write a
  1521. 58:15reward model in terms of our language
  1522. 58:17model itself now to intuitively
  1523. 58:20understand that like what is going on a
  1524. 58:22language model is assigning
  1525. 58:23probabilities to whatever is the most
  1526. 58:25plausible completion next but those
  1527. 58:28plausible completions might not be what
  1528. 58:29we intended but you could restrict the
  1529. 58:31probability simply to the completions
  1530. 58:34that a human might like and then the log
  1531. 58:36probabilities of your model might
  1532. 58:38represent something which the humans
  1533. 58:39might like and not just some arbitrary
  1534. 58:41completion on the internet so there is a
  1535. 58:43direct correspondence between the log
  1536. 58:46probability that a language model
  1537. 58:48assigns and how much a human might like
  1538. 58:50the answer they can have like a direct
  1539. 58:52correspondence in them and this is not
  1540. 58:55some arbitrary intuition that I'm trying
  1541. 58:56to like come up with we will derive this
  1542. 59:00mathematically so the general idea with
  1543. 59:03direct preference optimization is going
  1544. 59:04to be we're going to write down reward
  1545. 59:06model in terms of our language model and
  1546. 59:09now that we can write our reward model
  1547. 59:10in terms of our language model we can
  1548. 59:12simply solve directly fit our reward
  1549. 59:15model to the preference data we have and
  1550. 59:19we don't need to do the RL Step at all
  1551. 59:21so we started off with some preference
  1552. 59:22data and we simply fit our reward model
  1553. 59:24to it which directly optimiz as the
  1554. 59:26language
  1555. 59:28parameters and maybe at a high level why
  1556. 59:31is this like even possible like we did
  1557. 59:33this like really cumbersome process with
  1558. 59:34fitting a reward model and optimizing it
  1559. 59:37but in the whole process the only
  1560. 59:39external information that was being
  1561. 59:41added to the system like was human
  1562. 59:43labels labels on the preference data
  1563. 59:45when we optimize a learned reward model
  1564. 59:47there's no new information being added
  1565. 59:49into the system so this is why something
  1566. 59:52like this is even possible for quite a
  1567. 59:54few years this was not obvious obvious
  1568. 59:56but like as you will see like some of
  1569. 59:58these results like start to make sense
  1570. 1:00:00so we're going to derive direct
  1571. 1:00:02preference
  1572. 1:00:03optimization I'll I'll be after I'll be
  1573. 1:00:05here after the class as well if you have
  1574. 1:00:06questions but I'll hopefully like this
  1575. 1:00:08is clear
  1576. 1:00:11so yes um we discussed that we wanted to
  1577. 1:00:14solve this expected reward problem where
  1578. 1:00:17we want to maximize the expected reward
  1579. 1:00:19but we subtract this term which is the
  1580. 1:00:21beta log ratio which essentially
  1581. 1:00:22penalizes the distance between where our
  1582. 1:00:25current model is and where we started
  1583. 1:00:27off so we don't want to drift too far
  1584. 1:00:28away from our um from where we
  1585. 1:00:33started now it turns out that this
  1586. 1:00:36specific problem instead of doing like
  1587. 1:00:38an iterative routine um there's actually
  1588. 1:00:41a close form solution to this problem so
  1589. 1:00:45the close form solution looks something
  1590. 1:00:46like this um again if you have seen the
  1591. 1:00:50boltzman distribution or something to
  1592. 1:00:52that effect before this is very
  1593. 1:00:54basically the same idea but the idea is
  1594. 1:00:56this that we're going to take a
  1595. 1:00:57pre-train distribution PPT y given X and
  1596. 1:01:00we're going to rade the distribution by
  1597. 1:01:02the expected
  1598. 1:01:03reward so if if if a completion has a
  1599. 1:01:07very high reward it's going to have a
  1600. 1:01:09higher probability mass and if it has a
  1601. 1:01:11lower reward it's going to have a lower
  1602. 1:01:12probability mass and it's determined by
  1603. 1:01:14the expected reward and beta is a
  1604. 1:01:16hyperparameter which essentially governs
  1605. 1:01:18like what is the trade-off between the
  1606. 1:01:20reward model and the constraint and as
  1607. 1:01:24beta becomes lower and lower you're
  1608. 1:01:26going to start paying more and more
  1609. 1:01:27attention to the reward
  1610. 1:01:29model so the probabilities look
  1611. 1:01:32something like this and there's this
  1612. 1:01:34like really annoying term this ZX and
  1613. 1:01:37the reason why it exists is that the
  1614. 1:01:39numerator by itself is not normalized
  1615. 1:01:42it's not a probability distribution so
  1616. 1:01:44to construct like an actual probability
  1617. 1:01:46distribution you have to normalize it
  1618. 1:01:48and ZX is simply just this
  1619. 1:01:50normalization so if we write ZX out is
  1620. 1:01:53the sum of all y okay yeah and that's
  1621. 1:01:56exactly like it's some overall wise for
  1622. 1:01:58a given instruction and that's exactly
  1623. 1:02:00why this is very pesky is like it's
  1624. 1:02:02intractable if I take an instruction and
  1625. 1:02:04try to sum over every possible
  1626. 1:02:06completion and not just like
  1627. 1:02:07syntactically correct ones every single
  1628. 1:02:09possible we have 50,000 tokens maybe
  1629. 1:02:12even more and the completions can go
  1630. 1:02:14arbitrary long so this space is
  1631. 1:02:15completely intractable this quantity is
  1632. 1:02:17not easy to approximate
  1633. 1:02:21even um so the main point here is that
  1634. 1:02:25you if you're given in a reward model
  1635. 1:02:26you can actually there does exist at
  1636. 1:02:28least a close form solution which tells
  1637. 1:02:30us what the optimal policy will look
  1638. 1:02:31like or optimal language model will look
  1639. 1:02:33like but if you do a little bit of
  1640. 1:02:35algebra just move some terms around take
  1641. 1:02:37a logarithm here or there I I promise
  1642. 1:02:39this is not very complicated you can
  1643. 1:02:41actually Express the reward model in
  1644. 1:02:43terms of the language model itself and I
  1645. 1:02:46think this term is reasonably intuitive
  1646. 1:02:48as well uh what it says is that um a
  1647. 1:02:51completion y hat has a high reward if
  1648. 1:02:54the model my optimal policy assigns a
  1649. 1:02:57higher probability to it relative to my
  1650. 1:03:00initialized model and this is scal by
  1651. 1:03:03Beta so the beta log ratio is what we're
  1652. 1:03:05looking at
  1653. 1:03:07here and the partition function let's
  1654. 1:03:09just ignore it for now but it's
  1655. 1:03:10intractable but the beta log ratio is
  1656. 1:03:13the key part
  1657. 1:03:15here is everyone following
  1658. 1:03:18along awesome okay so right now I'm
  1659. 1:03:23talking about optimal policies but
  1660. 1:03:26really like every policy is probably
  1661. 1:03:28optimal for some kind of a reward right
  1662. 1:03:30like this is mathematically true as well
  1663. 1:03:32so the important bit here is that you
  1664. 1:03:34can actually Express you take a current
  1665. 1:03:37policy take your initialized model and
  1666. 1:03:39you can get some kind of a reward model
  1667. 1:03:41out of it and this is the exact identity
  1668. 1:03:44which leads to this so reward model can
  1669. 1:03:46be expressed in terms of your language
  1670. 1:03:48model baring the log partition term
  1671. 1:03:52which we'll see what happens to it go
  1672. 1:03:55for sorry I don't know how you got like
  1673. 1:03:57why is it that we can swap because there
  1674. 1:03:59is a thing that we're trying to optimize
  1675. 1:04:00and how do p star turn into P yeah um
  1676. 1:04:04for now like we're not optimizing any
  1677. 1:04:05reward model okay all I'm saying is that
  1678. 1:04:08if I take my current language model it
  1679. 1:04:10is it probably represents some kind of a
  1680. 1:04:12reward model
  1681. 1:04:15implicitly because of this relationship
  1682. 1:04:17because this holds for every P star and
  1683. 1:04:19every reward model what I'm saying is
  1684. 1:04:22that like there if I plug in my current
  1685. 1:04:24language model it also represents some
  1686. 1:04:26kind of a reward model I'm not saying
  1687. 1:04:27it's optimal
  1688. 1:04:29okay but I want say because at the
  1689. 1:04:31beginning uh PRL is PPT yes and so we
  1690. 1:04:36just get that the reward is basically
  1691. 1:04:38zero and so what what do we do initially
  1692. 1:04:41it's zero but like we can optimize the
  1693. 1:04:42parameters okay okay yeah um yeah but
  1694. 1:04:45that's a good observation that it's
  1695. 1:04:46basically zero in the beginning but how
  1696. 1:04:48do we start optimizing
  1697. 1:04:50it I'll get to okay okay any other
  1698. 1:04:54questions so the idea is that given the
  1699. 1:04:56language model you have model such that
  1700. 1:05:02that makes the language model
  1701. 1:05:04op yes that's uh that's the next step
  1702. 1:05:08yes uh but the key idea is that like my
  1703. 1:05:11log my language model the probabilities
  1704. 1:05:13already implicitly Define a reward model
  1705. 1:05:16I think that's really the main point
  1706. 1:05:18here and this mathematical relationship
  1707. 1:05:20is
  1708. 1:05:21exact cool now like I mean I'm obviously
  1709. 1:05:25ignoring like the elephant in the room
  1710. 1:05:27here which is the partition function um
  1711. 1:05:30it's not going to magically vanish away
  1712. 1:05:31so like if this was just the beta log
  1713. 1:05:33ratio that would be really nice I can
  1714. 1:05:35compute all these quantities I know how
  1715. 1:05:37to compute the log probability under my
  1716. 1:05:38language model I know how to compute the
  1717. 1:05:40log probability under my pre-train model
  1718. 1:05:43and I can compute the reward score and I
  1719. 1:05:45can optimize this but I don't know what
  1720. 1:05:47to do about my Lo log partition function
  1721. 1:05:50this is where something fun happens so
  1722. 1:05:54recall what the the reward modeling
  1723. 1:05:56objective was uh when we started off
  1724. 1:05:59like we started off with the friends
  1725. 1:06:00Bradley Terry again and what we really
  1726. 1:06:03wanted to optimize was the reward
  1727. 1:06:05difference between the winning
  1728. 1:06:06completion and the losing
  1729. 1:06:08completion um and really like I mean
  1730. 1:06:11that's all we care about we don't care
  1731. 1:06:12about the exact reward itself what we
  1732. 1:06:15care about is maximizing the difference
  1733. 1:06:16between the the difference between
  1734. 1:06:19winning and losing completion and that's
  1735. 1:06:21actually really key here because if you
  1736. 1:06:24plug in the definition of of the RM
  1737. 1:06:27Theta there what you'll observe is that
  1738. 1:06:30the partition function actually just
  1739. 1:06:32cancels out now why does it cancel out
  1740. 1:06:36um the input is exactly the same the x
  1741. 1:06:39is actually exactly the same in the
  1742. 1:06:41difference so the partition function ZX
  1743. 1:06:43will just cancel out like it's the same
  1744. 1:06:45in both the terms so what you get is
  1745. 1:06:47that the reward difference between the
  1746. 1:06:48winning and losing completion is the
  1747. 1:06:50differences between the beta log ratio
  1748. 1:06:51for the winning and losing
  1749. 1:06:53completion you can plug in the terms you
  1750. 1:06:56can work it out it's fairly simple so
  1751. 1:06:59the partition function which was our
  1752. 1:07:00like um which was something we could not
  1753. 1:07:03address we could not compute actually
  1754. 1:07:04just simply vanished away I'm so sorry Z
  1755. 1:07:07doesn't appear in theary mod um but it
  1756. 1:07:11appears here in this equation so how
  1757. 1:07:14does plug in
  1758. 1:07:16model um so we're going to take this
  1759. 1:07:19equation uh the last line that you see
  1760. 1:07:22and we're going to plug in in place of
  1761. 1:07:24RMF
  1762. 1:07:26okay so um and in this the first loss
  1763. 1:07:31equation oh I see got yeah so the first
  1764. 1:07:33loss equation is the broadly ter loss
  1765. 1:07:35model
  1766. 1:07:37cool so this really is it like I mean
  1767. 1:07:40the key observation is we could express
  1768. 1:07:42our reward model in terms of language
  1769. 1:07:43model and our problems with the
  1770. 1:07:45partition function actually go away
  1771. 1:07:46because we were optimizing the Brad lary
  1772. 1:07:48model and um what you get is something
  1773. 1:07:51like this is that um we're going to
  1774. 1:07:55Express the loss function directly in
  1775. 1:07:57terms of our language model parameters
  1776. 1:07:59Theta and we're going to be able to
  1777. 1:08:02directly optimize on our data um without
  1778. 1:08:05doing any RL steps or not and this is
  1779. 1:08:07simply a binary classification problem
  1780. 1:08:10so we're really just trying to classify
  1781. 1:08:12whether an answer is good or bad and
  1782. 1:08:14that's really what we're
  1783. 1:08:16doing before I go on like people want to
  1784. 1:08:19like absorb this in like I mean feel
  1785. 1:08:22they're okay with it
  1786. 1:08:26I don't get where they why good and why
  1787. 1:08:28win and why lose come from are they
  1788. 1:08:30human and or they good question um it's
  1789. 1:08:34the same data set we started with in rlf
  1790. 1:08:36as well but the way the process works is
  1791. 1:08:39that you take a set of instructions and
  1792. 1:08:40get the model to generate some answers
  1793. 1:08:42and then you get humans to label which
  1794. 1:08:44answer they prefer so they're model
  1795. 1:08:46generated uh typically they can be human
  1796. 1:08:48generated as well but they're typically
  1797. 1:08:50model generated and then you get some
  1798. 1:08:52preference labels okay all you need is a
  1799. 1:08:55label saying which is a better
  1800. 1:08:58answer what do you lose here like you
  1801. 1:09:02must be losing some information because
  1802. 1:09:03of the lack of information about like
  1803. 1:09:08other you're canceling out your your
  1804. 1:09:11your because of the lack of uh any
  1805. 1:09:13information about the partition function
  1806. 1:09:15yeah you are bound to lose information
  1807. 1:09:18about like other possible completions
  1808. 1:09:20which you would have taken into account
  1809. 1:09:22in like standard rlf right
  1810. 1:09:26um that's a really good question I don't
  1811. 1:09:27think I'll be able to completely answer
  1812. 1:09:29this question in time but like partition
  1813. 1:09:32function is almost kind of a free
  1814. 1:09:33variable so I think the problem here is
  1815. 1:09:35that the reward model there think of
  1816. 1:09:38when you there's many reward models that
  1817. 1:09:40satisfy this optimization so there's a
  1818. 1:09:43free variable here that you can actually
  1819. 1:09:45completely remove and that's what this
  1820. 1:09:47optimization benefits from so think of
  1821. 1:09:49it this way like if I assign something a
  1822. 1:09:50reward of plus one and assign something
  1823. 1:09:52a reward of minus one that's basically
  1824. 1:09:54the same as saying as if it's a reward
  1825. 1:09:55of plus
  1826. 1:09:5799 and it will give you the same loss
  1827. 1:10:01right so um that scale doesn't that
  1828. 1:10:05shift invariant in a ways is that like
  1829. 1:10:08isn't that somehow like not what you
  1830. 1:10:11want though like like okay like if if
  1831. 1:10:15you have if you're actually training a
  1832. 1:10:16reward model right like 199 is like much
  1833. 1:10:20you should pay much less attention to
  1834. 1:10:22that as compared to
  1835. 1:10:23like one right
  1836. 1:10:26Zer or something what we're assuming is
  1837. 1:10:27our choice model here is like if a human
  1838. 1:10:30prefers something over the other like
  1839. 1:10:32the probability is governed only by the
  1840. 1:10:34difference between the rewards so that's
  1841. 1:10:37an assumption that every rlf also makes
  1842. 1:10:39and like DPO also makes now is that
  1843. 1:10:42assumption true not completely true but
  1844. 1:10:45like it it holds to a fairly large
  1845. 1:10:49degree but that's a good question
  1846. 1:10:52yeah cool um I'll move on and rest of
  1847. 1:10:56time um and really like I mean the goal
  1848. 1:10:58of this plot is to like we actually get
  1849. 1:11:00fairly performant models when we
  1850. 1:11:02optimize things with DPO um we so in
  1851. 1:11:06this plot I think the main thing that
  1852. 1:11:07you should look at is po which is the
  1853. 1:11:08typical rlf Pipeline and we are
  1854. 1:11:10evaluating the models for summarization
  1855. 1:11:12and we're comparing to human summaries
  1856. 1:11:15and what we find is that DP and BP sort
  1857. 1:11:17of do similarly but you're really not
  1858. 1:11:19losing much by just doing the DPO
  1859. 1:11:21procedure instead of R LF and that's
  1860. 1:11:23really compelling because DP is simply a
  1861. 1:11:24classif app ation loss instead of like a
  1862. 1:11:26whole reinforcement learning
  1863. 1:11:29procedure so I want to quickly summarize
  1864. 1:11:32um what we have seen thus far is that we
  1865. 1:11:35want to optimize for human preferences
  1866. 1:11:37so um and the way we do this is like
  1867. 1:11:40instead of relying on uncalibrated
  1868. 1:11:41scores we're getting comparison data and
  1869. 1:11:43feedback on that and we use this ranking
  1870. 1:11:46data to either do something like rlf
  1871. 1:11:48where we first fit a reward model and
  1872. 1:11:50optimize using reinforcement learning um
  1873. 1:11:54or we do something that like direct
  1874. 1:11:55preference optimization we simply take
  1875. 1:11:57the data set and do a classification
  1876. 1:11:59loss on that um and yeah like there's
  1877. 1:12:02trade offs in these algorithms like
  1878. 1:12:04people when they have a lot of
  1879. 1:12:06computational budget they typically
  1880. 1:12:07maybe go for rlf or some routine like
  1881. 1:12:09that but if you're really looking to get
  1882. 1:12:12the bank for your buck like I mean you
  1883. 1:12:13might want to go for DPO and if and
  1884. 1:12:15that's like probably going to work out
  1885. 1:12:17of the box um it's a still an active
  1886. 1:12:20area of research people are still trying
  1887. 1:12:21to understand how to like best work with
  1888. 1:12:23these algorithms so like I'm not making
  1889. 1:12:25any strong claims here but like both of
  1890. 1:12:26these algorithms are very effective DP
  1891. 1:12:28is just much simpler to work
  1892. 1:12:30with
  1893. 1:12:32cool um so yeah like I mean let's see
  1894. 1:12:35like we went through all this
  1895. 1:12:36instruction tuning rlf what do we get um
  1896. 1:12:41instruct GPD is the first model which
  1897. 1:12:43sort of followed this pipeline it
  1898. 1:12:45defined this pipeline so we got models
  1899. 1:12:47which did 30,000 or so tasks remember
  1900. 1:12:50when we were doing like only one task
  1901. 1:12:52and now we have scaled it up from th000
  1902. 1:12:53tasks to like 30,000 different task with
  1903. 1:12:55many many different examples so that's
  1904. 1:12:57like where we are with instruct GPT and
  1905. 1:13:00it follows this pipeline that we just
  1906. 1:13:02described in this case they're following
  1907. 1:13:03a specific rlf pipeline where we
  1908. 1:13:05explicitly fit a reward model and then
  1909. 1:13:07do some kind of a reinforcement learning
  1910. 1:13:09routine on top of it um and yeah like
  1911. 1:13:14the task collected from labelers looks
  1912. 1:13:15something like this um I leave it to
  1913. 1:13:18your imagination or you can look at the
  1914. 1:13:19details but how we started off with this
  1915. 1:13:22model was something like completions we
  1916. 1:13:24see from G GPD 3 uh which you know
  1917. 1:13:27explained the moon Ling to sixer and
  1918. 1:13:29like it is not really following the
  1919. 1:13:31instructions but instruct GPD will give
  1920. 1:13:33you something which is Meaningful so
  1921. 1:13:35it's inferring what a user wanted from
  1922. 1:13:37the specific instruction and it's
  1923. 1:13:39converting to a realistic answer that a
  1924. 1:13:40user might
  1925. 1:13:43like and yeah these are just more
  1926. 1:13:46examples of what an instruct GPD like
  1927. 1:13:48model would do whereas your base model
  1928. 1:13:50might not follow the instructions to
  1929. 1:13:52your desired intentions
  1930. 1:13:56and yeah like we went from instruct GPD
  1931. 1:13:58to chart GPD and it was essentially this
  1932. 1:14:01pipeline um the key difference here is
  1933. 1:14:04that it is still doing the instruction
  1934. 1:14:06tuning but it is more optimized for
  1935. 1:14:08dialogue more optimized for interacting
  1936. 1:14:10with users so the core algorithmic
  1937. 1:14:13techniques that we discussed today are
  1938. 1:14:15what give us CH GPD but you have to be
  1939. 1:14:17really careful about the kind of data
  1940. 1:14:19you're training on and that's really the
  1941. 1:14:21whole game um but this is the foundation
  1942. 1:14:24for CH GPD
  1943. 1:14:26and yeah it it follows the same pipeline
  1944. 1:14:29as well and you might look at you might
  1945. 1:14:32interact with ch gbd I'm sure you all
  1946. 1:14:34have interacted with it some form or not
  1947. 1:14:35but like this is an example of what a CH
  1948. 1:14:37gbd interaction might look
  1949. 1:14:40like um you want to make a gen Z so like
  1950. 1:14:44I mean you can you know the idea here is
  1951. 1:14:46that it's like very good at responding
  1952. 1:14:47to instructions and intent this is not
  1953. 1:14:49something that we could like even fuse
  1954. 1:14:51shot in very easily uh these are kind of
  1955. 1:14:54instructions are hard to come examples
  1956. 1:14:56for but like this is probably not
  1957. 1:14:58something to trained on either but it's
  1958. 1:14:59able to like infer the intent and
  1959. 1:15:01generalize very very nicely and that's
  1960. 1:15:03something I find personally very
  1961. 1:15:07remarkable cool and there's been a lot
  1962. 1:15:10of progress on the open source front as
  1963. 1:15:12well so like DPO is much simpler and
  1964. 1:15:14much more efficient and essentially all
  1965. 1:15:16the open source models these days are
  1966. 1:15:18using DPO so this is a leaderboard that
  1967. 1:15:21is maintained by hugging hugging face a
  1968. 1:15:23so like I mean N9 out of 10 more models
  1969. 1:15:25here are trained with DPO so that's been
  1970. 1:15:28something that's been enabled the open
  1971. 1:15:29source Community to instruction tune
  1972. 1:15:31their model betters as well and same is
  1973. 1:15:34being used in many production models now
  1974. 1:15:36as well mistol is using DPO llama 3 used
  1975. 1:15:38DPO so these are very very strong models
  1976. 1:15:41which are nearly gp4 level and they're
  1977. 1:15:43also like starting to use um these
  1978. 1:15:46algorithms as well and something that's
  1979. 1:15:48very cool cool to see is like like we
  1980. 1:15:51went through all this like optimization
  1981. 1:15:52and like I mean math and stuff but what
  1982. 1:15:54is really fundamentally changing in the
  1983. 1:15:56behavior and I think this is a really
  1984. 1:15:58good example is that if you simply ask
  1985. 1:16:01an instruction for and ask for an sft
  1986. 1:16:03output from an instruction tune model
  1987. 1:16:05you'll get something like this but when
  1988. 1:16:07you RL of the model you actually get a
  1989. 1:16:09lot more details in your answer and
  1990. 1:16:11they'll probably organize the answers a
  1991. 1:16:13little better and there's something that
  1992. 1:16:14they maybe humans prefer that's why it's
  1993. 1:16:17an property that is emerging in these
  1994. 1:16:19model but it's something that's a very
  1995. 1:16:21clear difference between simply
  1996. 1:16:24instruction tude models and some models
  1997. 1:16:27which are
  1998. 1:16:30rft so yeah um we discuss like this
  1999. 1:16:34whole rlf routine where we are directly
  2000. 1:16:37modeling the preferences and we are
  2001. 1:16:38generalizing Beyond label data um and we
  2002. 1:16:41also discussed RL can be very tricky to
  2003. 1:16:43um correctly Implement though DPO sort
  2004. 1:16:45of implements this or like avoid some of
  2005. 1:16:48these issue and we briefly also touched
  2006. 1:16:50upon the idea of reward model and reward
  2007. 1:16:52hacking um and when you're optimizing
  2008. 1:16:56for learned reward models you will often
  2009. 1:16:58see this example is that there's a way
  2010. 1:17:01for it to just simply crash into um the
  2011. 1:17:05object some keep repe repetitively
  2012. 1:17:08crashing the board to get more and more
  2013. 1:17:09points that wasn't the goal of this game
  2014. 1:17:12so um this is a very common example that
  2015. 1:17:15is shown for reward hacking if you do
  2016. 1:17:18not specify Rewards well the models can
  2017. 1:17:20like learn weird behaviors which are not
  2018. 1:17:23your desired intent and there's
  2019. 1:17:24something a lot of people worry about as
  2020. 1:17:26well um part of the reason is
  2021. 1:17:28reinforcement learning is a very strong
  2022. 1:17:29optimization algorithm it's at the heart
  2023. 1:17:31of alpha go Alpha zero uh which like
  2024. 1:17:34results in superhuman models so you have
  2025. 1:17:36to be careful about how you specify
  2026. 1:17:38things and the other thing is like even
  2027. 1:17:40optimizing for human preferences is
  2028. 1:17:42often not the right thing because humans
  2029. 1:17:43are not do not always like things which
  2030. 1:17:46are in their best interest so something
  2031. 1:17:48that emerges is that they like
  2032. 1:17:49authoritative and helpful answers but
  2033. 1:17:51they often like don't necessarily like
  2034. 1:17:54truthful answers
  2035. 1:17:55so one property that happens is like is
  2036. 1:17:58that they'll prefer authoritativeness
  2037. 1:18:00more than correctness which is maybe
  2038. 1:18:02like not something nice please go ahead
  2039. 1:18:04on those lines I'm curious if maybe like
  2040. 1:18:07chbt being so like now widely used by
  2041. 1:18:10the public will maybe change the like
  2042. 1:18:12how people were like made the rewards
  2043. 1:18:14because I at least feel like now when I
  2044. 1:18:15go to chat I TP something it gives me
  2045. 1:18:17five like detailed paragraphs of
  2046. 1:18:19information sometimes I'm just annoyed
  2047. 1:18:20by that that's not what I wanted but
  2048. 1:18:22maybe in the original reward function in
  2049. 1:18:24the original people actually pref that
  2050. 1:18:25and nower it less yeah um that's a great
  2051. 1:18:29point because like as these models like
  2052. 1:18:31integrate more and more into our system
  2053. 1:18:33they're going to collect more and more
  2054. 1:18:34data and they will like pick up on
  2055. 1:18:37things maybe undesirable things as well
  2056. 1:18:40um as far as I understand chbd is really
  2057. 1:18:42cutting down on the verbosity which is
  2058. 1:18:44like a huge issue that all of these
  2059. 1:18:45models are trying to cut down on and
  2060. 1:18:47they are dealing with that um part of
  2061. 1:18:50the reason why that emerges is that when
  2062. 1:18:51you collect preference data at scale
  2063. 1:18:53people are not necessarily reading the
  2064. 1:18:55answers the turkers might just simply
  2065. 1:18:57choose the longer answer and that's a
  2066. 1:18:59property that actually goes into these
  2067. 1:19:00models so but hopefully like these
  2068. 1:19:03things will improve over time as they
  2069. 1:19:04get more feedb and yeah hallucinations
  2070. 1:19:07is not a problem that is going to go
  2071. 1:19:08away with RL and we talked a bit about
  2072. 1:19:10reward hacking as well um biases from
  2073. 1:19:14things and so on but hopefully like I
  2074. 1:19:16mean what I want to conclude out is like
  2075. 1:19:18we started with pre-trained
  2076. 1:19:20models we we had these things which
  2077. 1:19:22could predict text and we got chargy GPD
  2078. 1:19:25and hopefully like it's a little more
  2079. 1:19:26clear how we go from something like that
  2080. 1:19:28to chat
  2081. 1:19:30GPD and that's I'll end
  2082. 1:19:34here thanks

About this transcript

This page contains the full transcript of YouTube transcript (35X6zlhoCy4) , generated from the public captions YouTube serves with the video. The transcript has 14,287 words across 2,082 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.