YouTube2Text

Deep Dive into LLMs like ChatGPT — Transcript

by Andrej Karpathy · 41,116 words · 5,761 segments · language en · Watch on YouTube

Full transcript

  1. 0:00hi everyone so I've wanted to make this
  2. 0:02video for a while it is a comprehensive
  3. 0:05but General audience introduction to
  4. 0:08large language models like Chachi PT and
  5. 0:11what I'm hoping to achieve in this video
  6. 0:12is to give you kind of mental models for
  7. 0:14thinking through what it is that this
  8. 0:17tool is it is obviously magical and
  9. 0:19amazing in some respects it's uh really
  10. 0:22good at some things not very good at
  11. 0:23other things and there's also a lot of
  12. 0:25sharp edges to be aware of so what is
  13. 0:28behind this text box you can put
  14. 0:29anything in there and press enter but uh
  15. 0:32what should we be putting there and what
  16. 0:34are these words generated back how does
  17. 0:36this work and what what are you talking
  18. 0:38to exactly so I'm hoping to get at all
  19. 0:40those topics in this video we're going
  20. 0:42to go through the entire pipeline of how
  21. 0:44this stuff is built but I'm going to
  22. 0:45keep everything uh sort of accessible to
  23. 0:48a general audience so let's take a look
  24. 0:50at first how you build something like
  25. 0:51chpt and along the way I'm going to talk
  26. 0:53about um you know some of the sort of
  27. 0:56cognitive psychological implications of
  28. 0:59the tools okay so let's build Chachi PT
  29. 1:02so there's going to be multiple stages
  30. 1:04arranged sequentially the first stage is
  31. 1:07called the pre-training stage and the
  32. 1:09first step of the pre-training stage is
  33. 1:11to download and process the internet now
  34. 1:13to get a sense of what this roughly
  35. 1:14looks like I recommend looking at this
  36. 1:16URL here so um this company called
  37. 1:20hugging face uh collected and created
  38. 1:23and curated this data set called Fine
  39. 1:26web and they go into a lot of detail on
  40. 1:28this block post on how how they
  41. 1:30constructed the fine web data set and
  42. 1:32all of the major llm providers like open
  43. 1:34AI anthropic and Google and so on will
  44. 1:36have some equivalent internally of
  45. 1:38something like the fine web data set so
  46. 1:41roughly what are we trying to achieve
  47. 1:42here we're trying to get ton of text
  48. 1:44from the internet from publicly
  49. 1:45available sources so we're trying to
  50. 1:47have a huge quantity of very high
  51. 1:50quality documents and we also want very
  52. 1:53large diversity of documents because we
  53. 1:55want to have a lot of knowledge inside
  54. 1:56these models so we want large diversity
  55. 1:59of high quality documents and we want
  56. 2:01many many of them and achieving this is
  57. 2:04uh quite complicated and as you can see
  58. 2:05here takes multiple stages to do well so
  59. 2:08let's take a look at what some of these
  60. 2:09stages look like in a bit for now I'd
  61. 2:11like to just like to note that for
  62. 2:13example the fine web data set which is
  63. 2:14fairly representative what you would see
  64. 2:16in a production grade application
  65. 2:18actually ends up being only about 44
  66. 2:20terabyt of dis space um you can get a
  67. 2:23USB stick for like a terabyte very
  68. 2:25easily or I think this could fit on a
  69. 2:27single hard drive almost today so this
  70. 2:29is not a huge amount of data at the end
  71. 2:31of the day even though the internet is
  72. 2:33very very large we're working with text
  73. 2:35and we're also filtering it aggressively
  74. 2:37so we end up with about 44 terabytes in
  75. 2:39this example so let's take a look at uh
  76. 2:42kind of what this data looks like and
  77. 2:44what some of these stages uh also are so
  78. 2:47the starting point for a lot of these
  79. 2:48efforts and something that contributes
  80. 2:50most of the data by the end of it is
  81. 2:52Data from common crawl so common craw is
  82. 2:56an organization that has been basically
  83. 2:57scouring the internet since 2007 so as
  84. 3:00of 2024 for example common CW has
  85. 3:03indexed 2.7 billion web
  86. 3:05pages uh and uh they have all these
  87. 3:08crawlers going around the internet and
  88. 3:09what you end up doing basically is you
  89. 3:11start with a few seed web pages and then
  90. 3:13you follow all the links and you just
  91. 3:15keep following links and you keep
  92. 3:16indexing all the information and you end
  93. 3:17up with a ton of data of the internet
  94. 3:19over time so this is usually the
  95. 3:21starting point for a lot of the uh for a
  96. 3:24lot of these efforts now this common C
  97. 3:26data is quite raw and is filtered in
  98. 3:27many many different ways
  99. 3:30so here they Pro they document this is
  100. 3:33the same diagram they document a little
  101. 3:35bit the kind of processing that happens
  102. 3:36in these stages so the first thing here
  103. 3:39is something called URL
  104. 3:41filtering so what that is referring to
  105. 3:43is that there's these block
  106. 3:47lists of uh basically URLs that are or
  107. 3:50domains that uh you don't want to be
  108. 3:52getting data from so usually this
  109. 3:54includes things like U malware websites
  110. 3:56spam websites marketing websites uh
  111. 3:58racist websites adult sites and things
  112. 4:01like that so there's a ton of different
  113. 4:02types of websites that are just
  114. 4:04eliminated at this stage because we
  115. 4:06don't want them in our data set um the
  116. 4:08second part is text extraction you have
  117. 4:10to remember that all these web pages
  118. 4:12this is the raw HTML of these web pages
  119. 4:14that are being saved by these crawlers
  120. 4:16so when I go to inspect
  121. 4:18here this is what the raw HTML actually
  122. 4:21looks like you'll notice that it's got
  123. 4:23all this markup uh like lists and stuff
  124. 4:26like that and there's CSS and all this
  125. 4:28kind of stuff so this is um computer
  126. 4:31code almost for these web pages but what
  127. 4:33we really want is we just want this text
  128. 4:35right we just want the text of this web
  129. 4:37page and we don't want the navigation
  130. 4:38and things like that so there's a lot of
  131. 4:40filtering and processing uh and heris
  132. 4:42that go into uh adequately filtering for
  133. 4:45just their uh good content of these web
  134. 4:48pages the next stage here is language
  135. 4:50filtering so for example fine web
  136. 4:53filters uh using a language classifier
  137. 4:56they try to guess what language every
  138. 4:58single web page is in and then they only
  139. 5:00keep web pages that have more than 65%
  140. 5:02of English as an
  141. 5:04example and so you can get a sense that
  142. 5:06this is like a design decision that
  143. 5:07different companies can uh can uh take
  144. 5:10for themselves what fraction of all
  145. 5:12different types of languages are we
  146. 5:14going to include in our data set because
  147. 5:15for example if we filter out all of the
  148. 5:17Spanish as an example then you might
  149. 5:19imagine that our model later will not be
  150. 5:21very good at Spanish because it's just
  151. 5:22never seen that much data of that
  152. 5:24language and so different companies can
  153. 5:26focus on multilingual performance to uh
  154. 5:28to a different degree as an example so
  155. 5:30fine web is quite focused on English and
  156. 5:33so their language model if they end up
  157. 5:35training one later will be very good at
  158. 5:36English but not may be very good at
  159. 5:38other
  160. 5:39languages after language filtering
  161. 5:41there's a few other filtering steps and
  162. 5:43D duplication and things like that um
  163. 5:46finishing with for example the pii
  164. 5:49removal this is personally identifiable
  165. 5:52information so as an example addresses
  166. 5:54Social Security numbers and things like
  167. 5:56that you would try to detect them and
  168. 5:57you would try to filter out those kinds
  169. 5:58of web pages from the the data set as
  170. 6:00well so there's a lot of stages here and
  171. 6:02I won't go into full detail but it is a
  172. 6:05fairly extensive part of the
  173. 6:06pre-processing and you end up with for
  174. 6:08example the fine web data set so when
  175. 6:10you click in on it uh you can see some
  176. 6:12examples here of what this actually ends
  177. 6:14up looking like and anyone can download
  178. 6:16this on the huging phase web page and so
  179. 6:19here are some examples of the final text
  180. 6:21that ends up in the training set so this
  181. 6:24is some article about tornadoes in
  182. 6:272012 um so there's some t tadoes in 2020
  183. 6:30in 2012 and what
  184. 6:33happened uh this next one is something
  185. 6:36about did you know you have two little
  186. 6:38yellow 9vt battery sized adrenal glands
  187. 6:41in your body okay so this is some kind
  188. 6:43of a odd medical
  189. 6:46article so just think of these as
  190. 6:49basically uh web pages on the internet
  191. 6:51filtered just for the text in various
  192. 6:53ways and now we have a ton of text 40
  193. 6:56terabytes off it and that now is the
  194. 6:58starting point for the next step of this
  195. 7:00stage now I wanted to give you an
  196. 7:02intuitive sense of where we are right
  197. 7:04now so I took the first 200 web pages
  198. 7:06here and remember we have tons of them
  199. 7:09and I just take all that text and I just
  200. 7:11put it all together concatenate it and
  201. 7:13so this is what we end up with we just
  202. 7:15get this just just raw text raw internet
  203. 7:18text and there's a ton of it even in
  204. 7:20these 200 web pages so I can continue
  205. 7:22zooming out here and we just have this
  206. 7:24like massive tapestry of Text data and
  207. 7:28this text data has all these p patterns
  208. 7:30and what we want to do now is we want to
  209. 7:31start training neural networks on this
  210. 7:33data so the neural networks can
  211. 7:35internalize and model how this text
  212. 7:39flows right so we just have this giant
  213. 7:42texture of text and now we want to get
  214. 7:45neural Nets that mimic it okay now
  215. 7:48before we plug text into neural networks
  216. 7:51we have to decide how we're going to
  217. 7:52represent this text uh and how we're
  218. 7:54going to feed it in now the way our
  219. 7:57technology works for these neuron Lots
  220. 7:58is that they expect
  221. 7:59a one-dimensional sequence of symbols
  222. 8:02and they want a finite set of symbols
  223. 8:05that are possible and so we have to
  224. 8:08decide what are the symbols and then we
  225. 8:10have to represent our data as
  226. 8:11one-dimensional sequence of those
  227. 8:14symbols so right now what we have is a
  228. 8:16onedimensional sequence of text it
  229. 8:18starts here and it goes here and then it
  230. 8:20comes here Etc so this is a
  231. 8:22onedimensional sequence even though on
  232. 8:23my monitor of course it's laid out in a
  233. 8:26two-dimensional way but it goes from
  234. 8:27left to right and top to bottom right so
  235. 8:29it's a one-dimensional sequence of text
  236. 8:32now this being computers of course
  237. 8:33there's an underlying representation
  238. 8:35here so if I do what's called utf8 uh
  239. 8:38encode this text then I can get the raw
  240. 8:41bits that correspond to this text in the
  241. 8:44computer and that's what uh that looks
  242. 8:46like this so it turns out that for
  243. 8:50example this very first bar here is the
  244. 8:53first uh eight bits as an
  245. 8:56example so what is this thing right this
  246. 8:59is um representation that we are looking
  247. 9:01for uh in in a certain sense we have
  248. 9:04exactly two possible symbols zero and
  249. 9:06one and we have a very long sequence of
  250. 9:10it right now as it turns out um this
  251. 9:14sequence length is actually going to be
  252. 9:16very finite and precious resource uh in
  253. 9:19our neural network and we actually don't
  254. 9:21want extremely long sequences of just
  255. 9:23two symbols instead what we want is we
  256. 9:25want to trade off uh this um symbol
  257. 9:29size uh of this vocabulary as we call it
  258. 9:32and the resulting sequence length so we
  259. 9:35don't want just two symbols and
  260. 9:36extremely long sequences we're going to
  261. 9:38want more symbols and shorter sequences
  262. 9:42okay so one naive way of compressing or
  263. 9:44decreasing the length of our sequence
  264. 9:46here is to basically uh consider some
  265. 9:49group of consecutive bits for example
  266. 9:51eight bits and group them into a single
  267. 9:54what's called bite so because uh these
  268. 9:57bits are either on or off if we take a
  269. 10:00group of eight of them there turns out
  270. 10:01to be only 256 possible combinations of
  271. 10:04how these bits could be on or off and so
  272. 10:06therefore we can re repesent this
  273. 10:07sequence into a sequence of bytes
  274. 10:10instead so this sequence of bytes will
  275. 10:13be eight times shorter but now we have
  276. 10:16256 possible symbols so every number
  277. 10:19here goes from 0 to
  278. 10:20255 now I really encourage you to think
  279. 10:22of these not as numbers but as unique
  280. 10:25IDs or like unique symbols so maybe it's
  281. 10:28a bit more maybe it's better to actually
  282. 10:30think of these to replace every one of
  283. 10:32these with a unique Emoji you'd get
  284. 10:34something like this so um we basically
  285. 10:37have a sequence of emojis and there's
  286. 10:38256 possible emojis you can think of it
  287. 10:41that way now it turns out that in
  288. 10:44production for state-of-the-art language
  289. 10:46models uh you actually want to go even
  290. 10:48Beyond this you want to continue to
  291. 10:50shrink the length of the sequence uh
  292. 10:52because again it is a precious resource
  293. 10:54in return for more symbols in your
  294. 10:57vocabulary and the way this is done is
  295. 11:00done by running what's called The Bite
  296. 11:02pair encoding algorithm and the way this
  297. 11:04works is we're basically looking for
  298. 11:06consecutive bytes or symbols that are
  299. 11:10very common so for example turns out
  300. 11:13that the sequence 116 followed by 32 is
  301. 11:17quite common and occurs very frequently
  302. 11:19so what we're going to do is we're going
  303. 11:20to group uh this um pair into a new
  304. 11:24symbol so we're going to Mint a symbol
  305. 11:26with an ID 256 and we're going to
  306. 11:28rewrite every single uh pair 11632 with
  307. 11:32this new symbol and then can we can
  308. 11:34iterate this algorithm as many times as
  309. 11:36we wish and each time when we mint a new
  310. 11:38symbol we're decreasing the length and
  311. 11:40we're increasing the symbol size and in
  312. 11:43practice it turns out that a pretty good
  313. 11:45setting of um the basically the
  314. 11:48vocabulary size turns out to be about
  315. 11:49100,000 possible symbols so in
  316. 11:52particular GPT 4 uses
  317. 11:55100,
  318. 11:56277 symbols
  319. 11:59um and this process of converting from
  320. 12:04raw text into these symbols or as we
  321. 12:07call them tokens is the process called
  322. 12:10tokenization so let's now take a look at
  323. 12:12how gp4 performs tokenization conting
  324. 12:15from text to tokens and from tokens back
  325. 12:18to text and what this actually looks
  326. 12:19like so one website I like to use to
  327. 12:21explore these token representations is
  328. 12:24called tick tokenizer and so come here
  329. 12:27to the drop down and select CL 100 a
  330. 12:29base which is the gp4 base model
  331. 12:32tokenizer and here on the left you can
  332. 12:34put in text and it shows you the
  333. 12:36tokenization of that text so for example
  334. 12:40heo space
  335. 12:43world so hello world turns out to be
  336. 12:46exactly two Tokens The Token hello which
  337. 12:49is the token with ID
  338. 12:5115339 and the token space
  339. 12:54world that is the token 1
  340. 12:571917 so um hello space world now if I
  341. 13:02was to join these two for example I'm
  342. 13:04going to get again two tokens but it's
  343. 13:06the token H followed by the token L
  344. 13:09world without the
  345. 13:11H um if I put in two Spa two spaces here
  346. 13:15between hello and world it's again a
  347. 13:16different uh tokenization there's a new
  348. 13:19token 220
  349. 13:22here okay so you can play with this and
  350. 13:24see what happens here also keep in mind
  351. 13:26this is not uh this is case sensitive so
  352. 13:28if this is a capital H it is something
  353. 13:30else or if it's uh hello world then
  354. 13:35actually this ends up being three tokens
  355. 13:36instead of just two
  356. 13:41tokens yeah so you can play with this
  357. 13:43and get an sort of like an intuitive
  358. 13:44sense of uh what these tokens work like
  359. 13:47we're actually going to loop around to
  360. 13:48tokenization a bit later in the video
  361. 13:50for now I just wanted to show you the
  362. 13:51website and I wanted to uh show you that
  363. 13:53this text basically at the end of the
  364. 13:56day so for example if I take one line
  365. 13:57here this is what GT4 will see it as so
  366. 14:01this text will be a sequence of length
  367. 14:0462 this is the sequence here and this is
  368. 14:08how the chunks of text correspond to
  369. 14:11these symbols and again there's 100,
  370. 14:1627777 possible symbols and we now have
  371. 14:19one-dimensional sequences of those
  372. 14:21symbols so um yeah we're going to come
  373. 14:24back to tokenization but that's uh for
  374. 14:26now where we are okay so what I've done
  375. 14:28now is I've taken this uh sequence of
  376. 14:30text that we have here in the data set
  377. 14:32and I have re-represented it using our
  378. 14:34tokenizer into a sequence of tokens and
  379. 14:37this is what that looks like now so for
  380. 14:40example when we go back to the Fine web
  381. 14:41data set they mentioned that not only is
  382. 14:43this 44 terab of dis space but this is
  383. 14:45about a 15 trillion token sequence of um
  384. 14:50in this data set and so here these are
  385. 14:53just some of the first uh one or two or
  386. 14:56three or a few thousand here I think uh
  387. 14:58tokens of this data set but there's 15
  388. 15:01trillion here uh to keep in mind and
  389. 15:03again keep in mind one more time that
  390. 15:05all of these represent little text
  391. 15:07chunks they're all just like atoms of
  392. 15:09these sequences and the numbers here
  393. 15:11don't make any sense they're just uh
  394. 15:13they're just unique IDs okay so now we
  395. 15:17get to the fun part which is the uh
  396. 15:19neural network training and this is
  397. 15:21where a lot of the heavy lifting happens
  398. 15:23computationally when you're training
  399. 15:24these neural networks so what we do here
  400. 15:28in this this step is we want to model
  401. 15:30the statistical relationships of how
  402. 15:32these tokens follow each other in the
  403. 15:33sequence so what we do is we come into
  404. 15:36the data and we take Windows of tokens
  405. 15:40so we take a window of tokens uh from
  406. 15:43this data fairly
  407. 15:44randomly and um the windows length can
  408. 15:49range anywhere anywhere between uh zero
  409. 15:51tokens actually all the way up to some
  410. 15:54maximum size that we decide on uh so for
  411. 15:57example in practice you could see a
  412. 15:58token with Windows of say 8,000 tokens
  413. 16:01now in principle we can use arbitrary
  414. 16:03window lengths of tokens uh but uh
  415. 16:07processing very long uh basically U
  416. 16:10window sequences would just be very
  417. 16:12computationally expensive so we just
  418. 16:15kind of decide that say 8,000 is a good
  419. 16:16number or 4,000 or 16,000 and we crop it
  420. 16:19there now in this example I'm going to
  421. 16:22be uh taking the first four tokens just
  422. 16:25so everything fits nicely so these
  423. 16:28tokens
  424. 16:30we're going to take a window of four
  425. 16:32tokens this bar view in and space single
  426. 16:37which are these token
  427. 16:39IDs and now what we're trying to do here
  428. 16:41is we're trying to basically predict the
  429. 16:42token that comes next in the sequence so
  430. 16:453962 comes next right so what we do now
  431. 16:49here is that we call this the context
  432. 16:51these four tokens are context and they
  433. 16:54feed into a neural
  434. 16:56network and this is the input to the
  435. 16:58neural network
  436. 16:59now I'm going to go into the detail of
  437. 17:01what's inside this neural network in a
  438. 17:03little bit for now it's important to
  439. 17:04understand is the input and the output
  440. 17:06of the neural net so the input are
  441. 17:08sequences of tokens of variable length
  442. 17:12anywhere between zero and some maximum
  443. 17:14size like 8,000 the output now is a
  444. 17:17prediction for what comes next so
  445. 17:21because our vocabulary has
  446. 17:23100277 possible tokens the neural
  447. 17:26network is going to Output exactly that
  448. 17:28many numbers
  449. 17:29and all of those numbers correspond to
  450. 17:30the probability of that token as coming
  451. 17:33next in the sequence so it's making
  452. 17:35guesses about what comes
  453. 17:37next um in the beginning this neural
  454. 17:39network is randomly initialized so um
  455. 17:42and we're going to see in a little bit
  456. 17:44what that means but it's a it's a it's a
  457. 17:46random transformation so these
  458. 17:48probabilities in the very beginning of
  459. 17:49the training are also going to be kind
  460. 17:51of random uh so here I have three
  461. 17:53examples but keep in mind that there's
  462. 17:55100,000 numbers here um so the
  463. 17:58probability of this token space
  464. 17:59Direction neural network is saying that
  465. 18:01this is 4% likely right now 11799 is 2%
  466. 18:05and then here the probility of 3962
  467. 18:08which is post is 3% now of course we've
  468. 18:11sampled this window from our data set so
  469. 18:13we know what comes next we know and
  470. 18:16that's the label we know that the
  471. 18:18correct answer is that 3962 actually
  472. 18:19comes next in the sequence so now what
  473. 18:22we have is this mathematical process for
  474. 18:25doing an update to the neural network we
  475. 18:28have the way of tuning it and uh we're
  476. 18:30going to go into a little bit of of
  477. 18:32detail in a bit but basically we know
  478. 18:34that this probability here of 3% we want
  479. 18:38this probability to be higher and we
  480. 18:40want the probabilities of all the other
  481. 18:42tokens to be
  482. 18:44lower and so we have a way of
  483. 18:46mathematically calculating how to adjust
  484. 18:48and update the neural network so that
  485. 18:51the correct answer has a slightly higher
  486. 18:53probability so if I do an update to the
  487. 18:55neural network now the next time I Fe
  488. 18:59this particular sequence of four tokens
  489. 19:00into neural network the neural network
  490. 19:02will be slightly adjusted now and it
  491. 19:04will say Okay post is maybe 4% and case
  492. 19:07now maybe is
  493. 19:081% and uh Direction could become 2% or
  494. 19:12something like that and so we have a way
  495. 19:14of nudging of slightly updating the
  496. 19:16neuronet to um basically give a higher
  497. 19:19probability to the correct token that
  498. 19:21comes next in the sequence and now you
  499. 19:23just have to remember that this process
  500. 19:25happens not just for uh this um token
  501. 19:29here where these four fed in and
  502. 19:31predicted this one this process happens
  503. 19:33at the same time for all of these tokens
  504. 19:36in the entire data set and so in
  505. 19:38practice we sample little windows little
  506. 19:40batches of Windows and then at every
  507. 19:42single one of these tokens we want to
  508. 19:44adjust our neural network so that the
  509. 19:46probability of that token becomes
  510. 19:48slightly higher and this all happens in
  511. 19:50parallel in large batches of these
  512. 19:52tokens and this is the process of
  513. 19:54training the neural network it's a
  514. 19:55sequence of updating it so that it's
  515. 19:58predictions match up the statistics of
  516. 20:01what actually happens in your training
  517. 20:02set and its probabilities become
  518. 20:05consistent with the uh statistical
  519. 20:08patterns of how these tokens follow each
  520. 20:09other in the data so let's now briefly
  521. 20:12get into the internals of these neural
  522. 20:13networks just to give you a sense of
  523. 20:14what's inside so neural network
  524. 20:17internals so as I mentioned we have
  525. 20:19these inputs uh that are sequences of
  526. 20:22tokens in this case this is four input
  527. 20:24tokens but this can be anywhere between
  528. 20:26zero up to let's say 8,000 tokens in
  529. 20:30principle this can be an infinite number
  530. 20:31of tokens we just uh it would just be
  531. 20:33too computationally expensive to process
  532. 20:35an infinite number of tokens so we just
  533. 20:37crop it at a certain length and that
  534. 20:39becomes the maximum context length of
  535. 20:41that uh
  536. 20:42model now these inputs X are mixed up in
  537. 20:46a giant mathematical expression together
  538. 20:48with the parameters or the weights of
  539. 20:51these neural networks so here I'm
  540. 20:53showing six example parameters and their
  541. 20:56setting but in practice these uh um
  542. 21:00modern neural networks will have
  543. 21:01billions of these uh parameters and in
  544. 21:04the beginning these parameters are
  545. 21:06completely randomly set now with a
  546. 21:09random setting of parameters you might
  547. 21:11expect that this uh this neural network
  548. 21:13would make random predictions and it
  549. 21:15does in the beginning it's totally
  550. 21:16random predictions but it's through this
  551. 21:19process of iteratively updating the
  552. 21:22network uh as and we call that process
  553. 21:24training a neural network so uh that the
  554. 21:28setting of these parameters gets
  555. 21:29adjusted such that the outputs of our
  556. 21:31neural network becomes consistent with
  557. 21:34the patterns seen in our training
  558. 21:36set so think of these parameters as kind
  559. 21:39of like knobs on a DJ set and as you're
  560. 21:41twiddling these knobs you're getting
  561. 21:42different uh predictions for every
  562. 21:45possible uh token sequence input and
  563. 21:49training in neural network just means
  564. 21:50discovering a setting of parameters that
  565. 21:52seems to be consistent with the
  566. 21:54statistics of the training
  567. 21:56set now let me just give you an example
  568. 21:58what this giant mathematical expression
  569. 21:59looks like just to give you a sense and
  570. 22:01modern networks are massive expressions
  571. 22:03with trillions of terms probably but let
  572. 22:06me just show you a simple example here
  573. 22:08it would look something like this I mean
  574. 22:10these are the kinds of Expressions just
  575. 22:11to show you that it's not very scary we
  576. 22:13have inputs x uh like X1 x2 in this case
  577. 22:17two example inputs and they get mixed up
  578. 22:19with the weights of the network w0 W1 2
  579. 22:223 Etc and this mixing is simple things
  580. 22:27like multiplication addition addition
  581. 22:29exponentiation division Etc and it is
  582. 22:32the subject of neural network
  583. 22:34architecture research to design
  584. 22:36effective mathematical Expressions uh
  585. 22:39that have a lot of uh kind of convenient
  586. 22:41characteristics they are expressive
  587. 22:42they're optimizable they're paralyzable
  588. 22:45Etc and so but uh at the end of the day
  589. 22:48these are these are not complex
  590. 22:49expressions and basically they mix up
  591. 22:52the inputs with the parameters to make
  592. 22:54predictions and we're optimizing uh the
  593. 22:57parameters of this neural network so
  594. 22:59that the predictions come out consistent
  595. 23:01with the training set now I would like
  596. 23:04to show you an actual production grade
  597. 23:06example of what these neural networks
  598. 23:07look like so for that I encourage you to
  599. 23:09go to this website that has a very nice
  600. 23:11visualization of one of these
  601. 23:13networks so this is what you will find
  602. 23:16on this website and this neural network
  603. 23:19here that is used in production settings
  604. 23:21has this special kind of structure this
  605. 23:24network is called the Transformer and
  606. 23:26this particular one as an example has 8
  607. 23:285,000 roughly
  608. 23:30parameters now here on the top we take
  609. 23:33the inputs which are the token
  610. 23:36sequences and then information flows
  611. 23:39through the neural network until the
  612. 23:41output which here are the logit softmax
  613. 23:45but these are the predictions for what
  614. 23:46comes next what token comes
  615. 23:48next and then here there's a sequence of
  616. 23:52Transformations and all these
  617. 23:54intermediate values that get produced
  618. 23:56inside this mathematical expression s it
  619. 23:58is sort of predicting what comes next so
  620. 24:01as an example these tokens are embedded
  621. 24:04into kind of like this distributed
  622. 24:06representation as it's called so every
  623. 24:08possible token has kind of like a vector
  624. 24:10that represents it inside the neural
  625. 24:11network so first we embed the tokens and
  626. 24:15then those values uh kind of like flow
  627. 24:18through this diagram and these are all
  628. 24:20very simple mathematical Expressions
  629. 24:22individually so we have layer norms and
  630. 24:24Matrix multiplications and uh soft Maxes
  631. 24:27and so on so here kind of like the
  632. 24:28attention block of this Transformer and
  633. 24:31then information kind of flows through
  634. 24:33into the multi-layer perceptron block
  635. 24:35and so on and all these numbers here
  636. 24:38these are the intermediate values of the
  637. 24:40expression and uh you can almost think
  638. 24:42of these as kind of like the firing
  639. 24:44rates of these synthetic neurons but I
  640. 24:47would caution you to uh not um kind of
  641. 24:50think of it too much like neurons
  642. 24:52because these are extremely simple
  643. 24:53neurons compared to the neurons you
  644. 24:55would find in your brain your biological
  645. 24:57neurons are very complex dynamical
  646. 24:59processes that have memory and so on
  647. 25:01there's no memory in this expression
  648. 25:02it's a fixed mathematical expression
  649. 25:04from input to Output with no memory it's
  650. 25:06just a
  651. 25:07stateless so these are very simple
  652. 25:09neurons in comparison to biological
  653. 25:10neurons but you can still kind of
  654. 25:12loosely think of this as like a
  655. 25:13synthetic piece of uh brain tissue if
  656. 25:15you if you like uh to think about it
  657. 25:17that way so information flows through
  658. 25:21all these neurons fire until we get to
  659. 25:24the predictions now I'm not actually
  660. 25:26going to dwell too much on the precise
  661. 25:28kind of like mathematical details of all
  662. 25:30these Transformations honestly I don't
  663. 25:31think it's that important to get into
  664. 25:33what's really important to understand is
  665. 25:35that this is a mathematical function it
  666. 25:38is uh parameterized by some fixed set of
  667. 25:41parameters like say 85,000 of them and
  668. 25:44it is a way of transforming inputs into
  669. 25:46outputs and as we twiddle the parameters
  670. 25:48we are getting uh different kinds of
  671. 25:50predictions and then we need to find a
  672. 25:52good setting of these parameters so that
  673. 25:54the predictions uh sort of match up with
  674. 25:56the patterns seen in training set
  675. 25:59so that's the Transformer okay so I've
  676. 26:02shown you the internals of the neural
  677. 26:03network and we talked a bit about the
  678. 26:05process of training it I want to cover
  679. 26:07one more major stage of working with
  680. 26:10these networks and that is the stage
  681. 26:11called inference so in inference what
  682. 26:14we're doing is we're generating new data
  683. 26:16from the model and so uh we want to
  684. 26:18basically see what kind of patterns it
  685. 26:21has internalized in the parameters of
  686. 26:23its Network so to generate from the
  687. 26:26model is relatively straightforward
  688. 26:28we start with some tokens that are
  689. 26:30basically your prefix like what you want
  690. 26:32to start with so say we want to start
  691. 26:34with the token 91 well we feed it into
  692. 26:37the
  693. 26:37network and remember that the network
  694. 26:39gives us probabilities right it gives us
  695. 26:43this probability Vector here so what we
  696. 26:45can do now is we can basically flip a
  697. 26:47biased coin so um we can sample uh
  698. 26:52basically a token based on this
  699. 26:54probability distribution so the tokens
  700. 26:57that are given High probability by the
  701. 26:59model are more likely to be sampled when
  702. 27:01you flip this biased coin you can think
  703. 27:03of it that way so we sample from the
  704. 27:05distribution to get a single unique
  705. 27:08token so for example token 860 comes
  706. 27:11next uh so 860 in this case when we're
  707. 27:14generating from model could come next
  708. 27:16now 860 is a relatively likely token it
  709. 27:18might not be the only possible token in
  710. 27:20this case there could be many other
  711. 27:21tokens that could have been sampled but
  712. 27:23we could see that 86c is a relatively
  713. 27:25likely token as an example and indeed in
  714. 27:27our training examp example here 860 does
  715. 27:29follow 91 so let's now say that we um
  716. 27:34continue the process so after 91 there's
  717. 27:36a60 we append it and we again ask what
  718. 27:39is the third token let's sample and
  719. 27:42let's just say that it's 287 exactly as
  720. 27:44here let's do that again we come back in
  721. 27:47now we have a sequence of three and we
  722. 27:49ask what is the likely fourth token and
  723. 27:52we sample from that and get this one and
  724. 27:55now let's say we do it one more time we
  725. 27:58take those four we sample and we get
  726. 28:00this one and this
  727. 28:0213659 uh this is not actually uh 3962 as
  728. 28:06we had before so this token is the token
  729. 28:09article uh instead so viewing a single
  730. 28:12article and so in this case we didn't
  731. 28:15exactly reproduce the sequence that we
  732. 28:17saw here in the training data so keep in
  733. 28:20mind that these systems are stochastic
  734. 28:22they have um we're sampling and we're
  735. 28:25flipping coins and sometimes we lock out
  736. 28:28and we reproduce some like small chunk
  737. 28:30of the text and training set but
  738. 28:32sometimes we're uh we're getting a token
  739. 28:35that was not verbatim part of any of the
  740. 28:38documents in the training data so we're
  741. 28:40going to get sort of like remixes of the
  742. 28:43data that we saw in the training because
  743. 28:44at every step of the way we can flip and
  744. 28:47get a slightly different token and then
  745. 28:48once that token makes it in if you
  746. 28:50sample the next one and so on you very
  747. 28:52quickly uh start to generate token
  748. 28:55streams that are very different from the
  749. 28:57token streams that UR
  750. 28:58in the training documents so
  751. 29:00statistically they will have similar
  752. 29:02properties but um they are not identical
  753. 29:05to your training data they're kind of
  754. 29:06like inspired by the training data and
  755. 29:09so in this case we got a slightly
  756. 29:10different sequence and why would we get
  757. 29:12article you might imagine that article
  758. 29:14is a relatively likely token in the
  759. 29:16context of bar viewing single Etc and
  760. 29:21you can imagine that the word article
  761. 29:22followed this context window somewhere
  762. 29:24in the training documents uh to some
  763. 29:26extent and we just happen to sample it
  764. 29:28here at that stage so basically
  765. 29:31inference is just uh predicting from
  766. 29:33these distributions one at a time we
  767. 29:35continue feeding back tokens and getting
  768. 29:37the next one and we uh we're always
  769. 29:39flipping these coins and depending on
  770. 29:42how lucky or unlucky we get um we might
  771. 29:45get very different kinds of patterns
  772. 29:47depending on how we sample from these
  773. 29:49probability distributions so that's
  774. 29:51inference so in most common scenarios uh
  775. 29:55basically downloading the internet and
  776. 29:57tokenizing it is is a pre-processing
  777. 29:58step you do that a single time and then
  778. 30:01uh once you have your token sequence we
  779. 30:04can start training networks and in
  780. 30:06Practical cases you would try to train
  781. 30:08many different networks of different
  782. 30:10kinds of uh settings and different kinds
  783. 30:11of arrangements and different kinds of
  784. 30:13sizes and so you''ll be doing a lot of
  785. 30:15neural network training and um then once
  786. 30:18you have a neural network and you train
  787. 30:19it and you have some specific set of
  788. 30:21parameters that you're happy with um
  789. 30:24then you can take the model and you can
  790. 30:25do inference and you can actually uh
  791. 30:28generate data from the model and when
  792. 30:30you're on chat GPT and you're talking
  793. 30:31with a model uh that model is trained
  794. 30:33and has been trained by open aai many
  795. 30:36months ago probably and they have a
  796. 30:38specific set of Weights that work well
  797. 30:41and when you're talking to the model all
  798. 30:42of that is just inference there's no
  799. 30:44more training those parameters are held
  800. 30:47fixed and you're just talking to the
  801. 30:49model sort of uh you're giving it some
  802. 30:51of the tokens and it's kind of
  803. 30:53completing token sequences and that's
  804. 30:54what you're seeing uh generated when you
  805. 30:57actually use the model on CH GPT so that
  806. 30:59model then just does inference alone so
  807. 31:02let's now look at an example of training
  808. 31:04an inference that is kind of concrete
  809. 31:05and gives you a sense of what this
  810. 31:07actually looks like uh when these models
  811. 31:08are trained now the example that I would
  812. 31:10like to work with and that I'm
  813. 31:12particularly fond of is that of opening
  814. 31:14eyes gpt2 so GPT uh stands for
  815. 31:17generatively pre-trained Transformer and
  816. 31:19this is the second iteration of the GPT
  817. 31:21series by open AI when you are talking
  818. 31:23to chat GPT today the model that is
  819. 31:26underlying all of the magic of that
  820. 31:27interaction is GPT 4 so the fourth
  821. 31:30iteration of that series now gpt2 was
  822. 31:33published in 2019 by openi in this paper
  823. 31:36that I have right here and the reason I
  824. 31:39like gpt2 is that it is the first time
  825. 31:41that a recognizably modern stack came
  826. 31:44together so um all of the pieces of gpd2
  827. 31:48are recognizable today by modern
  828. 31:50standards it's just everything has
  829. 31:52gotten bigger now I'm not going to be
  830. 31:54able to go into the full details of this
  831. 31:55paper of course because it is a
  832. 31:57technical publication but some of the
  833. 31:59details that I would like to highlight
  834. 32:00are as follows gpt2 was a Transformer
  835. 32:03neural network just like you were just
  836. 32:05like the neural networks you would work
  837. 32:06with today it was it had 1.6 billion
  838. 32:10parameters right so these are the
  839. 32:12parameters that we looked at here it
  840. 32:14would have 1.6 billion of them today
  841. 32:16modern Transformers would have a lot
  842. 32:18closer to a trillion or several hundred
  843. 32:20billion
  844. 32:21probably the maximum context length here
  845. 32:24was 1,24 tokens so it is when we are
  846. 32:28sampling chunks of Windows of tokens
  847. 32:32from the data set we're never taking
  848. 32:34more than 1,24 tokens and so when you
  849. 32:36are trying to predict the next token in
  850. 32:38a sequence you will never have more than
  851. 32:401,24 tokens uh kind of in your context
  852. 32:43in order to make that prediction now
  853. 32:45this is also tiny by modern standards
  854. 32:47today the token uh the context lengths
  855. 32:49would be a lot closer to um couple
  856. 32:53hundred thousand or maybe even a million
  857. 32:55and so you have a lot more context a lot
  858. 32:56more tokens in history history and you
  859. 32:58can make a lot better prediction about
  860. 33:00the next token in the sequence in that
  861. 33:01way and finally gpt2 was trained on
  862. 33:04approximately 100 billion tokens and
  863. 33:06this is also fairly small by modern
  864. 33:08standards as I mentioned the fine web
  865. 33:10data set that we looked at here the fine
  866. 33:12web data set has 15 trillion tokens uh
  867. 33:14so 100 billion is is quite
  868. 33:16small
  869. 33:18now uh I actually tried to reproduce uh
  870. 33:21gpt2 for fun as part of this project
  871. 33:23called lm. C so you can see my rup of
  872. 33:27doing that in this post on GitHub under
  873. 33:30the lm. C repository so in particular
  874. 33:33the cost of training gpd2 in 2019 what
  875. 33:36was estimated to be approximately
  876. 33:39$40,000 but today you can do
  877. 33:41significantly better than that and in
  878. 33:42particular here it took about one day
  879. 33:45and about
  880. 33:47$600 uh but this wasn't even trying too
  881. 33:49hard I think you could really bring this
  882. 33:51down to about $100 today now why is it
  883. 33:55that the costs have come down so much
  884. 33:57well number one these data sets have
  885. 33:59gotten a lot better and the way we
  886. 34:01filter them extract them and prepare
  887. 34:03them has gotten a lot more refined and
  888. 34:05so the data set is of just a lot higher
  889. 34:08quality so that's one thing but really
  890. 34:10the biggest difference is that our
  891. 34:11computers have gotten much faster in
  892. 34:13terms of the hardware and we're going to
  893. 34:15look at that in a second and also the
  894. 34:17software for uh running these models and
  895. 34:20really squeezing out all all the speed
  896. 34:22from the hardware as it is possible uh
  897. 34:25that software has also gotten much
  898. 34:27better as as everyone has focused on
  899. 34:28these models and try to run them very
  900. 34:30very
  901. 34:31quickly now I'm not going to be able to
  902. 34:34go into the full detail of this gpd2
  903. 34:36reproduction and this is a long
  904. 34:37technical post but I would like to still
  905. 34:39give you an intuitive sense for what it
  906. 34:41looks like to actually train one of
  907. 34:43these models as a researcher like what
  908. 34:44are you looking at and what does it look
  909. 34:46like what does it feel like so let me
  910. 34:47give you a sense of that a little bit
  911. 34:50okay so this is what it looks like let
  912. 34:51me slide this
  913. 34:52over so what I'm doing here is I'm
  914. 34:55training a gpt2 model right now
  915. 34:58and um what's happening here is that
  916. 35:00every single line here like this one is
  917. 35:05one update to the model so remember how
  918. 35:08here we are um basically making the
  919. 35:12prediction better for every one of these
  920. 35:14tokens and we are updating these weights
  921. 35:15or parameters of the neural net so here
  922. 35:18every single line is One update to the
  923. 35:20neural network where we change its
  924. 35:22parameters by a little bit so that it is
  925. 35:24better at predicting next token and
  926. 35:26sequence in particular every single line
  927. 35:28here is improving the prediction on 1
  928. 35:32million tokens in the training set so
  929. 35:35we've basically taken 1 million tokens
  930. 35:39out of this data set and we've tried to
  931. 35:41improve the prediction of that token as
  932. 35:44coming next in a sequence on all 1
  933. 35:46million of them
  934. 35:49simultaneously and at every single one
  935. 35:51of these steps we are making an update
  936. 35:52to the network for that now the number
  937. 35:55to watch closely is this number called
  938. 35:57loss and the loss is a single number
  939. 36:00that is telling you how well your neural
  940. 36:02network is performing right now and it
  941. 36:05is created so that low loss is good so
  942. 36:08you'll see that the loss is decreasing
  943. 36:10as we make more updates to the neural
  944. 36:12nut which corresponds to making better
  945. 36:14predictions on the next token in a
  946. 36:16sequence and so the loss is the number
  947. 36:19that you are watching as a neural
  948. 36:20network researcher and you are kind of
  949. 36:22waiting you're twiddling your thumbs uh
  950. 36:24you're drinking coffee and you're making
  951. 36:26sure that this looks good so that with
  952. 36:28every update your loss is improving and
  953. 36:31the network is getting better at
  954. 36:32prediction now here you see that we are
  955. 36:36processing 1 million tokens per update
  956. 36:38each update takes about 7 Seconds
  957. 36:41roughly and here we are going to process
  958. 36:43a total of 32,000 steps of
  959. 36:47optimization so 32,000 steps with 1
  960. 36:50million tokens each is about 33 billion
  961. 36:52tokens that we are going to process and
  962. 36:54we're currently only about 420 step 20
  963. 36:57out of 32,000 so we are still only a bit
  964. 37:01more than 1% done because I've only been
  965. 37:03running this for 10 or 15 minutes or
  966. 37:05something like
  967. 37:06that now every 20 steps I have
  968. 37:09configured this optimization to do
  969. 37:11inference so what you're seeing here is
  970. 37:13the model is predicting the next token
  971. 37:15in a sequence and so you sort of start
  972. 37:17it randomly and then you continue
  973. 37:19plugging in the tokens so we're running
  974. 37:21this inference step and this is the
  975. 37:23model sort of predicting the next token
  976. 37:25in the sequence and every time you see
  977. 37:26something appear that's a new
  978. 37:29token um so let's just look at this and
  979. 37:34you can see that this is not yet very
  980. 37:35coherent and keep in mind that this is
  981. 37:37only 1% of the way through training and
  982. 37:39so the model is not yet very good at
  983. 37:41predicting the next token in the
  984. 37:42sequence so what comes out is actually
  985. 37:44kind of a little bit of gibberish right
  986. 37:47but it still has a little bit of like
  987. 37:48local coherence so since she is mine
  988. 37:51it's a part of the information should
  989. 37:53discuss my father great companions
  990. 37:55Gordon showed me sitting over at and Etc
  991. 37:59so I know it doesn't look very good but
  992. 38:00let's actually scroll up and see what it
  993. 38:04looked like when I started the
  994. 38:06optimization so all the way here at
  995. 38:10step
  996. 38:12one so after 20 steps of optimization
  997. 38:15you see that what we're getting here is
  998. 38:17looks completely random and of course
  999. 38:18that's because the model has only had 20
  1000. 38:20updates to its parameters and so it's
  1001. 38:22giving you random text because it's a
  1002. 38:23random Network and so you can see that
  1003. 38:25at least in comparison to this model is
  1004. 38:27starting to do much better and indeed if
  1005. 38:29we waited the entire 32,000 steps the
  1006. 38:32model will have improved the point that
  1007. 38:34it's actually uh generating fairly
  1008. 38:36coherent English uh and the tokens
  1009. 38:38stream correctly um and uh they they
  1010. 38:42kind of make up English a a lot
  1011. 38:44better
  1012. 38:46um so this has to run for about a day or
  1013. 38:49two more now and so uh at this stage we
  1014. 38:52just make sure that the loss is
  1015. 38:53decreasing everything is looking good um
  1016. 38:56and we just have to wait
  1017. 38:58and now um let me turn now to the um
  1018. 39:02story of the computation that's required
  1019. 39:05because of course I'm not running this
  1020. 39:06optimization on my laptop that would be
  1021. 39:08way too expensive uh because we have to
  1022. 39:11run this neural network and we have to
  1023. 39:12improve it and we have we need all this
  1024. 39:14data and so on so you can't run this too
  1025. 39:16well on your computer uh because the
  1026. 39:18network is just too large uh so all of
  1027. 39:21this is running on the computer that is
  1028. 39:23out there in the cloud and I want to
  1029. 39:25basically address the compute side of
  1030. 39:27the store of training these models and
  1031. 39:28what that looks like so let's take a
  1032. 39:30look okay so the computer that I'm
  1033. 39:32running this optimization on is this 8X
  1034. 39:35h100 node so there are eight h100s in a
  1035. 39:39single node or a single computer now I
  1036. 39:42am renting this computer and it is
  1037. 39:44somewhere in the cloud I'm not sure
  1038. 39:45where it is physically actually the
  1039. 39:47place I like to rent from is called
  1040. 39:49Lambda but there are many other
  1041. 39:50companies who provide this service so
  1042. 39:52when you scroll down you can see that uh
  1043. 39:55they have some on demand pricing for
  1044. 39:57um sort of computers that have these uh
  1045. 40:01h100s which are gpus and I'm going to
  1046. 40:03show you what they look like in a second
  1047. 40:06but on demand 8times Nvidia h100 uh
  1048. 40:10GPU this machine comes for $3 per GPU
  1049. 40:13per hour for example so you can rent
  1050. 40:16these and then you get a machine in a
  1051. 40:18cloud and you can uh go in and you can
  1052. 40:20train these
  1053. 40:21models and these uh gpus they look like
  1054. 40:25this so this is one h100 GPU uh this is
  1055. 40:29kind of what it looks like and you slot
  1056. 40:30this into your computer and gpus are
  1057. 40:32this uh perfect fit for training your
  1058. 40:34networks because they are very
  1059. 40:36computationally expensive but they
  1060. 40:38display a lot of parallelism in the
  1061. 40:40computation so you can have many
  1062. 40:42independent workers kind of um working
  1063. 40:44all at the same time in solving uh the
  1064. 40:48matrix multiplication that's under the
  1065. 40:50hood of training these neural
  1066. 40:52networks so this is just one of these
  1067. 40:54h100s but actually you would put them
  1068. 40:56you would put multiple of them together
  1069. 40:58so you could stack eight of them into a
  1070. 41:00single node and then you can stack
  1071. 41:02multiple nodes into an entire data
  1072. 41:04center or an entire system
  1073. 41:07so when we look at a data
  1074. 41:12center can't spell when we look at a
  1075. 41:15data center we start to see things that
  1076. 41:16look like this right so we have one GPU
  1077. 41:18goes to eight gpus goes to a single
  1078. 41:19system goes to many systems and so these
  1079. 41:22are the bigger data centers and there of
  1080. 41:23course would be much much more expensive
  1081. 41:26um and what's happening is that all the
  1082. 41:28big tech companies really desire these
  1083. 41:31gpus so they can train all these
  1084. 41:33language models because they are so
  1085. 41:35powerful and that has is fundamentally
  1086. 41:37what has driven the stock price of
  1087. 41:38Nvidia to be $3.4 trillion today as an
  1088. 41:41example and why Nvidia has kind of
  1089. 41:44exploded so this is the Gold Rush the
  1090. 41:47Gold Rush is getting the gpus getting
  1091. 41:50enough of them so they can all
  1092. 41:52collaborate to perform this optimization
  1093. 41:55and they're what are they all doing
  1094. 41:56they're all collaborating to predict the
  1095. 41:59next token on a data set like the fine
  1096. 42:01web data
  1097. 42:02set this is the computational workflow
  1098. 42:05that that basically is extremely
  1099. 42:06expensive the more gpus you have the
  1100. 42:09more tokens you can try to predict and
  1101. 42:10improve on and you're going to process
  1102. 42:12this data set faster and you can iterate
  1103. 42:15faster and get a bigger Network and
  1104. 42:16train a bigger Network and so on so this
  1105. 42:19is what all those machines are look like
  1106. 42:20are uh are doing and this is why all of
  1107. 42:24this is such a big deal and for example
  1108. 42:26this is a
  1109. 42:28article from like about a month ago or
  1110. 42:30so this is why it's a big deal that for
  1111. 42:31example Elon Musk is getting 100,000
  1112. 42:34gpus uh in a single Data Center and all
  1113. 42:38of these gpus are extremely expensive
  1114. 42:40are going to take a ton of power and all
  1115. 42:42of them are just trying to predict the
  1116. 42:43next token in the sequence and improve
  1117. 42:45the network uh by doing so and uh get
  1118. 42:49probably a lot more coherent text than
  1119. 42:50what we're seeing here a lot faster okay
  1120. 42:52so unfortunately I do not have a couple
  1121. 42:5510 or hundred million of dollars to
  1122. 42:57spend on training a really big model
  1123. 42:59like this but luckily we can turn to
  1124. 43:01some big tech companies who train these
  1125. 43:04models routinely and release some of
  1126. 43:06them once they are done training so
  1127. 43:08they've spent a huge amount of compute
  1128. 43:10to train this network and they release
  1129. 43:12the network at the end of the
  1130. 43:13optimization so it's very useful because
  1131. 43:15they've done a lot of compute for that
  1132. 43:18so there are many companies who train
  1133. 43:19these models routinely but actually not
  1134. 43:21many of them release uh these what's
  1135. 43:23called base models so the model that
  1136. 43:25comes out at the end here is is what's
  1137. 43:27called a base model what is a base model
  1138. 43:29it's a token simulator right it's an
  1139. 43:32internet text token simulator and so
  1140. 43:35that is not by itself useful yet because
  1141. 43:38what we want is what's called an
  1142. 43:39assistant we want to ask questions and
  1143. 43:41have it respond to answers these models
  1144. 43:43won't do that they just uh create sort
  1145. 43:45of remixes of the internet they dream
  1146. 43:48internet pages so the base models are
  1147. 43:51not very often released because they're
  1148. 43:52kind of just only a step one of a few
  1149. 43:55other steps that we still need to take
  1150. 43:56to get in system
  1151. 43:58however a few releases have been made so
  1152. 44:01as an example the gbt2 model released
  1153. 44:04the 1.6 billion sorry 1.5 billion model
  1154. 44:08back in 2019 and this gpt2 model is a
  1155. 44:10base model now what is a model release
  1156. 44:13what does it look like to release these
  1157. 44:15models so this is the gpt2 repository on
  1158. 44:18GitHub well you need two things
  1159. 44:20basically to release model number one we
  1160. 44:22need the um python code usually that
  1161. 44:27describes the sequence of operations in
  1162. 44:30detail that they make in their model so
  1163. 44:34um if you remember
  1164. 44:36back this
  1165. 44:38Transformer the sequence of steps that
  1166. 44:40are taken here in this neural network is
  1167. 44:42what is being described by this code so
  1168. 44:45this code is sort of implementing the
  1169. 44:47what's called forward pass of this
  1170. 44:49neural network so we need the specific
  1171. 44:51details of exactly how they wired up
  1172. 44:53that neural network so this is just
  1173. 44:55computer code and it's usually just a
  1174. 44:57couple hundred lines of code it's not
  1175. 44:59it's not that crazy and uh this is all
  1176. 45:01fairly understandable and usually fairly
  1177. 45:03standard what's not standard are the
  1178. 45:05parameters that's where the actual value
  1179. 45:07is what are the parameters of this
  1180. 45:09neural network because there's 1.6
  1181. 45:11billion of them and we need the correct
  1182. 45:13setting or a really good setting and so
  1183. 45:15that's why in addition to this source
  1184. 45:17code they release the parameters which
  1185. 45:20in this case is roughly 1.5 billion
  1186. 45:23parameters and these are just numbers so
  1187. 45:25it's one single list of 1.5 billion
  1188. 45:27numbers the precise and good setting of
  1189. 45:30all the knobs such that the tokens come
  1190. 45:32out
  1191. 45:33well so uh you need those two things to
  1192. 45:37get a base model
  1193. 45:39release
  1194. 45:41now gpt2 was released but that's
  1195. 45:43actually a fairly old model as I
  1196. 45:44mentioned so actually the model we're
  1197. 45:46going to turn to is called llama 3 and
  1198. 45:49that's the one that I would like to show
  1199. 45:50you next so llama 3 so gpt2 again was
  1200. 45:541.6 billion parameters trained on 100
  1201. 45:55billion tokens Lama 3 is a much bigger
  1202. 45:58model and much more modern model it is
  1203. 46:00released and trained by meta and it is a
  1204. 46:0345 billion parameter model trained on 15
  1205. 46:07trillion tokens in very much the same
  1206. 46:09way just much much
  1207. 46:11bigger um and meta has also made a
  1208. 46:14release of llama 3 and that was part of
  1209. 46:18this
  1210. 46:19paper so with this paper that goes into
  1211. 46:21a lot of detail the biggest base model
  1212. 46:23that they released is the Lama 3.1 4.5
  1213. 46:27405 billion parameter model so this is
  1214. 46:30the base model and then in addition to
  1215. 46:32the base model you see here
  1216. 46:33foreshadowing for later sections of the
  1217. 46:35video they also released the instruct
  1218. 46:37model and the instruct means that this
  1219. 46:39is an assistant you can ask it questions
  1220. 46:41and it will give you answers we still
  1221. 46:43have yet to cover that part later for
  1222. 46:45now let's just look at this base model
  1223. 46:47this token simulator and let's play with
  1224. 46:49it and try to think about you know what
  1225. 46:51is this thing and how does it work and
  1226. 46:53um what do we get at the end of this
  1227. 46:55optimization if you let this run Until
  1228. 46:57the End uh for a very big neural network
  1229. 46:59on a lot of data so my favorite place to
  1230. 47:02interact with the base models is this um
  1231. 47:04company called hyperbolic which is
  1232. 47:06basically serving the base model of the
  1233. 47:09405b Llama 3.1 so when you go to the
  1234. 47:13website and I think you may have to
  1235. 47:14register and so on make sure that in the
  1236. 47:16models make sure that you are using
  1237. 47:18llama 3.1 405 billion base it must be
  1238. 47:22the base model and then here let's say
  1239. 47:24the max tokens is how many tokens we're
  1240. 47:26going to be gener rating so let's just
  1241. 47:28decrease this to be a bit less just so
  1242. 47:30we don't waste compute we just want the
  1243. 47:32next 128 tokens and leave the other
  1244. 47:34stuff alone I'm not going to go into the
  1245. 47:36full detail here um now fundamentally
  1246. 47:39what's going to happen here is identical
  1247. 47:41to what happens here during inference
  1248. 47:43for us so this is just going to continue
  1249. 47:45the token sequence of whatever you
  1250. 47:47prefix you're going to give it so I want
  1251. 47:49to first show you that this model here
  1252. 47:51is not yet an assistant so you can for
  1253. 47:53example ask it what is 2 plus 2 it's not
  1254. 47:56going to tell you oh it's four uh what
  1255. 47:58else can I help you with it's not going
  1256. 47:59to do that because what is 2 plus 2 is
  1257. 48:02going to be tokenized and then those
  1258. 48:05tokens just act as a prefix and then
  1259. 48:07what the model is going to do now is
  1260. 48:09just going to get the probability for
  1261. 48:10the next token and it's just a glorified
  1262. 48:12autocomplete it's a very very expensive
  1263. 48:14autocomplete of what comes next um
  1264. 48:17depending on the statistics of what it
  1265. 48:18saw in its training documents which are
  1266. 48:20basically web
  1267. 48:22pages so let's just uh hit enter to see
  1268. 48:25what tokens it comes up with as a
  1269. 48:31continuation okay so here it kind of
  1270. 48:32actually answered the question and
  1271. 48:34started to go off into some
  1272. 48:35philosophical territory uh let's try it
  1273. 48:37again so let me copy and paste and let's
  1274. 48:39try again from scratch what is 2 plus
  1275. 48:45two so okay so it just goes off again so
  1276. 48:49notice one more thing that I want to
  1277. 48:50stress is that the system uh I think
  1278. 48:53every time you put it in it just kind of
  1279. 48:55starts from scratch
  1280. 48:58so it doesn't uh the system here is
  1281. 48:59stochastic so for the same prefix of
  1282. 49:02tokens we're always getting a different
  1283. 49:04answer and the reason for that is that
  1284. 49:06we get this probity distribution and we
  1285. 49:08sample from it and we always get
  1286. 49:10different samples and we sort of always
  1287. 49:11go into a different territory uh
  1288. 49:13afterwards so here in this case um I
  1289. 49:17don't know what this is let's try one
  1290. 49:19more
  1291. 49:22time so it just continues on so it's
  1292. 49:25just doing the stuff that it's saw on
  1293. 49:26the internet right um and it's just kind
  1294. 49:29of like regurgitating those uh
  1295. 49:31statistical
  1296. 49:32patterns so first things it's not an
  1297. 49:35assistant yet it's a token autocomplete
  1298. 49:38and second it is a stochastic system now
  1299. 49:42the crucial thing is that even though
  1300. 49:44this model is not yet by itself very
  1301. 49:46useful for a lot of applications just
  1302. 49:49yet um it is still very useful because
  1303. 49:52in the task of predicting the next token
  1304. 49:54in the sequence the model has learned a
  1305. 49:56lot about the world and it has stored
  1306. 49:59all that knowledge in the parameters of
  1307. 50:01the network so remember that our text
  1308. 50:04looked like this right internet web
  1309. 50:06pages and now all of this is sort of
  1310. 50:08compressed in the weights of the network
  1311. 50:11so you can think of um these 405 billion
  1312. 50:15parameters is a kind of compression of
  1313. 50:16the internet you can think of the
  1314. 50:1945 billion parameters is kind of like a
  1315. 50:21zip file uh but it's not a loss less
  1316. 50:25compression it's a loss C compression
  1317. 50:27we're kind of like left with kind of a
  1318. 50:28gal of the internet and we can generate
  1319. 50:31from it right now we can elicit some of
  1320. 50:34this knowledge by prompting the base
  1321. 50:35model uh accordingly so for example
  1322. 50:38here's a prompt that might work to
  1323. 50:40elicit some of that knowledge that's
  1324. 50:41hiding in the parameters here's my top
  1325. 50:4310 list of the top landmarks to see in
  1326. 50:46the
  1327. 50:48pairs
  1328. 50:50um and I'm doing it this way because I'm
  1329. 50:52trying to Prime the model to now
  1330. 50:54continue this list so let's see if that
  1331. 50:56works when I press
  1332. 50:57enter okay so you see that it started a
  1333. 51:00list and it's now kind of giving me some
  1334. 51:02of those
  1335. 51:03landmarks and now notice that it's
  1336. 51:05trying to give a lot of information here
  1337. 51:07now you might not be able to actually
  1338. 51:09fully trust some of the information here
  1339. 51:10remember that this is all just a
  1340. 51:12recollection of some of the internet
  1341. 51:14documents and so the things that occur
  1342. 51:17very frequently in the internet data are
  1343. 51:19probably more likely to be remembered
  1344. 51:21correctly compared to things that happen
  1345. 51:23very infrequently so you can't fully
  1346. 51:25trust some of the things that and some
  1347. 51:27of the information that is here because
  1348. 51:28it's all just a vague recollection of
  1349. 51:30Internet documents because the
  1350. 51:32information is not stored explicitly in
  1351. 51:34any of the parameters it's all just the
  1352. 51:36recollection that said we did get
  1353. 51:38something that is probably approximately
  1354. 51:40correct and I don't actually have the
  1355. 51:42expertise to verify that this is roughly
  1356. 51:44correct but you see that we've elicited
  1357. 51:46a lot of the knowledge of the model and
  1358. 51:48this knowledge is not precise and exact
  1359. 51:51this knowledge is vague and
  1360. 51:53probabilistic and statistical and the
  1361. 51:55kinds of things that occur often are the
  1362. 51:57kinds of things that are more likely to
  1363. 51:59be remembered um in the model now I want
  1364. 52:02to show you a few more examples of this
  1365. 52:04model's Behavior the first thing I want
  1366. 52:05to show you is this example I went to
  1367. 52:08the Wikipedia page for zebra and let me
  1368. 52:10just copy paste the first uh even one
  1369. 52:13sentence
  1370. 52:14here and let me put it here now when I
  1371. 52:17click enter what kind of uh completion
  1372. 52:19are we going to get so let me just hit
  1373. 52:23enter there are three living species
  1374. 52:26etc etc what the model is producing here
  1375. 52:29is an exact regurgitation of this
  1376. 52:31Wikipedia entry it is reciting this
  1377. 52:33Wikipedia entry purely from memory and
  1378. 52:36this memory is stored in its parameters
  1379. 52:39and so it is possible that at some point
  1380. 52:41in these 512 tokens the model will uh
  1381. 52:44stray away from the Wikipedia entry but
  1382. 52:46you can see that it has huge chunks of
  1383. 52:47it memorized here uh let me see for
  1384. 52:50example if this sentence
  1385. 52:51occurs by now okay so this so we're
  1386. 52:55still on track let me check
  1387. 52:58here okay we're still on
  1388. 53:00track it will eventually uh stray
  1389. 53:04away okay so this thing is just recited
  1390. 53:07to a very large extent it will
  1391. 53:08eventually deviate uh because it won't
  1392. 53:11be able to remember exactly now the
  1393. 53:13reason that this happens is because
  1394. 53:14these models can be extremely good at
  1395. 53:16memorization and usually this is not
  1396. 53:18what you want in the final model and
  1397. 53:20this is something called regurgitation
  1398. 53:21and it's usually undesirable to site uh
  1399. 53:24things uh directly uh that you have
  1400. 53:26trained on now the reason that this
  1401. 53:29happens actually is because for a lot of
  1402. 53:31documents like for example Wikipedia
  1403. 53:33when these documents are deemed to be of
  1404. 53:35very high quality as a source like for
  1405. 53:37example Wikipedia it is very often uh
  1406. 53:40the case that when you train the model
  1407. 53:42you will preferentially sample from
  1408. 53:44those sources so basically the model has
  1409. 53:46probably done a few epochs on this data
  1410. 53:48meaning that it has seen this web page
  1411. 53:50like maybe probably 10 times or so and
  1412. 53:52it's a bit like you like when you read
  1413. 53:54some kind of a text many many times say
  1414. 53:56you read something a 100 times uh then
  1415. 53:58you'll be able to recite it and it's
  1416. 54:00very similar for this model if it sees
  1417. 54:01something way too often it's going to be
  1418. 54:03able to recite it later from memory
  1419. 54:05except these models can be a lot more
  1420. 54:07efficient um like per presentation than
  1421. 54:10human so probably it's only seen this
  1422. 54:12Wikipedia entry 10 times but basically
  1423. 54:14it has remembered this article exactly
  1424. 54:16in its parameters okay the next thing I
  1425. 54:18want to show you is something that the
  1426. 54:19model has definitely not seen during its
  1427. 54:21training so for example if we go to the
  1428. 54:24paper uh and then we navigate to the
  1429. 54:26pre-training data we'll see here that uh
  1430. 54:31the data set has a knowledge cut off
  1431. 54:33until the end of 2023 so it will not
  1432. 54:35have seen documents after this point and
  1433. 54:38certainly it has not seen anything about
  1434. 54:39the 2024 election and how it turned out
  1435. 54:43now if we Prime the model with the
  1436. 54:46tokens from the future it will continue
  1437. 54:49the token sequence and it will just take
  1438. 54:50its best guess according to the
  1439. 54:51knowledge that it has in its own
  1440. 54:53parameters so let's take a look at what
  1441. 54:55that could look like
  1442. 54:57so the Republican Party kit
  1443. 54:59Trump okay president of the United
  1444. 55:01States from
  1445. 55:022017 and let's see what it says after
  1446. 55:05this point so for example the model will
  1447. 55:07have to guess at the running mate and
  1448. 55:09who it's against Etc so let's hit
  1449. 55:11enter so here thingss that Mike Pence
  1450. 55:14was the running mate instead of JD Vance
  1451. 55:17and the ticket was against Hillary
  1452. 55:20Clinton and Tim Kane so this is kind of
  1453. 55:23a interesting parallel universe
  1454. 55:25potentially of what could have happened
  1455. 55:26happened according to the LM let's get a
  1456. 55:28different sample so the identical prompt
  1457. 55:31and let's
  1458. 55:33resample so here the running mate was
  1459. 55:35Ronda santis and they ran against Joe
  1460. 55:38Biden and Camala Harris so this is again
  1461. 55:40a different parallel universe so the
  1462. 55:42model will take educated guesses and it
  1463. 55:44will continue the token sequence based
  1464. 55:45on this knowledge um and it will just
  1465. 55:48kind of like all of what we're seeing
  1466. 55:49here is what's called hallucination the
  1467. 55:51model is just taking its best guess uh
  1468. 55:54in a probalistic manner the next thing I
  1469. 55:56would like to show you is that even
  1470. 55:58though this is a base model and not yet
  1471. 56:00an assistant model it can still be
  1472. 56:02utilized in Practical applications if
  1473. 56:04you are clever with your prompt design
  1474. 56:06so here's something that we would call a
  1475. 56:08few shot
  1476. 56:09prompt so what it is here is that I have
  1477. 56:1210 words or 10 pairs and each pair is a
  1478. 56:16word of English column and then a the
  1479. 56:19translation in Korean and we have 10 of
  1480. 56:22them and what the model does here is at
  1481. 56:25the end we have teacher column and then
  1482. 56:27here's where we're going to do a
  1483. 56:28completion of say just five tokens and
  1484. 56:31these models have what we call in
  1485. 56:33context learning abilities and what
  1486. 56:35that's referring to is that as it is
  1487. 56:37reading this context it is learning sort
  1488. 56:40of in
  1489. 56:41place that there's some kind of a
  1490. 56:43algorithmic pattern going on in my data
  1491. 56:46and it knows to continue that pattern
  1492. 56:48and this is called kind of like Inc
  1493. 56:50context learning so it takes on the role
  1494. 56:53of a
  1495. 56:54translator and when we hit uh completion
  1496. 56:58we see that the teacher translation is
  1497. 56:59Sim which is correct um and so this is
  1498. 57:03how you can build apps by being clever
  1499. 57:05with your prompting even though we still
  1500. 57:06just have a base model for now and it
  1501. 57:08relies on what we call this um uh in
  1502. 57:11context learning ability and it is done
  1503. 57:14by constructing what's called a few shot
  1504. 57:15prompt okay and finally I want to show
  1505. 57:17you that there is a clever way to
  1506. 57:19actually instantiate a whole language
  1507. 57:21model assistant just by prompting and
  1508. 57:24the trick to it is that we're structure
  1509. 57:26a prompt to look like a web page that is
  1510. 57:29a conversation between a helpful AI
  1511. 57:31assistant and a human and then the model
  1512. 57:34will continue that conversation so
  1513. 57:36actually to write the prompt I turned to
  1514. 57:38chat gbt itself which is kind of meta
  1515. 57:41but I told it I want to create an llm
  1516. 57:43assistant but all I have is the base
  1517. 57:45model so can you please write my um uh
  1518. 57:50prompt and this is what it came up with
  1519. 57:52which is actually quite good so here's a
  1520. 57:54conversation between an AI assistant and
  1521. 57:55a human
  1522. 57:56the AI assistant is knowledgeable
  1523. 57:58helpful capable of answering wide
  1524. 57:59variety of questions Etc and then here
  1525. 58:03it's not enough to just give it a sort
  1526. 58:05of description it works much better if
  1527. 58:07you create this fot prompt so here's a
  1528. 58:10few terms of human assistant human
  1529. 58:13assistant and we have uh you know a few
  1530. 58:15turns of conversation and then here at
  1531. 58:17the end is we're going to be putting the
  1532. 58:19actual query that we like so let me copy
  1533. 58:21paste this into the base model prompt
  1534. 58:25and now let me do human column and this
  1535. 58:28is where we put our actual prompt why is
  1536. 58:31the sky
  1537. 58:32blue and uh let's uh
  1538. 58:37run assistant the sky appears blue due
  1539. 58:40to the phenomenon called R lights
  1540. 58:41scattering etc etc so you see that the
  1541. 58:44base model is just continuing the
  1542. 58:45sequence but because the sequence looks
  1543. 58:47like this conversation it takes on that
  1544. 58:49role but it is a little subtle because
  1545. 58:52here it just uh you know it ends the
  1546. 58:54assistant and then just you know
  1547. 58:55hallucinate Ates the next question by
  1548. 58:57the human Etc so it'll just continue
  1549. 58:58going on and on uh but you can see that
  1550. 59:01we have sort of accomplished the task
  1551. 59:03and if you just took this why is the sky
  1552. 59:06blue and if we just refresh this and put
  1553. 59:09it here then of course we don't expect
  1554. 59:10this to work with a base model right
  1555. 59:12we're just going to who knows what we're
  1556. 59:14going to get okay we're just going to
  1557. 59:15get more
  1558. 59:16questions okay so this is one way to
  1559. 59:19create an assistant even though you may
  1560. 59:21only have a base model okay so this is
  1561. 59:24the kind of brief summary of the things
  1562. 59:26we talked about over the last few
  1563. 59:28minutes now let me zoom out
  1564. 59:32here and this is kind of like what we've
  1565. 59:34talked about so far we wish to train LM
  1566. 59:37assistants like chpt we've discussed the
  1567. 59:40first stage of that which is the
  1568. 59:42pre-training stage and we saw that
  1569. 59:44really what it comes down to is we take
  1570. 59:45Internet documents we break them up into
  1571. 59:47these tokens these atoms of little text
  1572. 59:49chunks and then we predict token
  1573. 59:51sequences using neural networks the
  1574. 59:54output of this entire stage is this base
  1575. 59:56model it is the setting of The
  1576. 59:58parameters of this network and this base
  1577. 1:00:01model is basically an internet document
  1578. 1:00:03simulator on the token level so it can
  1579. 1:00:05just uh it can generate token sequences
  1580. 1:00:08that have the same kind of like
  1581. 1:00:10statistics as Internet documents and we
  1582. 1:00:12saw that we can use it in some
  1583. 1:00:13applications but we actually need to do
  1584. 1:00:15better we want an assistant we want to
  1585. 1:00:17be able to ask questions and we want the
  1586. 1:00:18model to give us answers and so we need
  1587. 1:00:21to now go into the second stage which is
  1588. 1:00:23called the post-training stage so we
  1589. 1:00:26take our base model our internet
  1590. 1:00:28document simulator and hand it off to
  1591. 1:00:29post training so we're now going to
  1592. 1:00:31discuss a few ways to do what's called
  1593. 1:00:33post training of these models these
  1594. 1:00:36stages in post training are going to be
  1595. 1:00:38computationally much less expensive most
  1596. 1:00:40of the computational work all of the
  1597. 1:00:42massive data centers um and all of the
  1598. 1:00:45sort of heavy compute and millions of
  1599. 1:00:47dollars are the pre-training stage but
  1600. 1:00:50now we go into the slightly cheaper but
  1601. 1:00:52still extremely important stage called
  1602. 1:00:54post trining where we turn this llm
  1603. 1:00:57model into an assistant so let's take a
  1604. 1:00:59look at how we can get our model to not
  1605. 1:01:02sample internet documents but to give
  1606. 1:01:04answers to questions so in other words
  1607. 1:01:07what we want to do is we want to start
  1608. 1:01:08thinking about conversations and these
  1609. 1:01:10are conversations that can be multi-turn
  1610. 1:01:13so so uh there can be multiple turns and
  1611. 1:01:15they are in the simplest case a
  1612. 1:01:17conversation between a human and an
  1613. 1:01:19assistant and so for example we can
  1614. 1:01:21imagine the conversation could look
  1615. 1:01:22something like this when a human says
  1616. 1:01:24what is 2 plus2 the assistant should re
  1617. 1:01:25respond with something like 2 plus 2 is
  1618. 1:01:274 when a human follows up and says what
  1619. 1:01:29if it was star instead of a plus
  1620. 1:01:31assistant could respond with something
  1621. 1:01:32like
  1622. 1:01:33this um and similar here this is another
  1623. 1:01:36example showing that the assistant could
  1624. 1:01:37also have some kind of a personality
  1625. 1:01:39here uh that it's kind of like nice and
  1626. 1:01:41then here in the third example I'm
  1627. 1:01:43showing that when a human is asking for
  1628. 1:01:44something that we uh don't wish to help
  1629. 1:01:47with we can produce what's called
  1630. 1:01:48refusal we can say that we cannot help
  1631. 1:01:50with that so in other words what we want
  1632. 1:01:53to do now is we want to think through
  1633. 1:01:55how in a system should interact with the
  1634. 1:01:57human and we want to program the
  1635. 1:01:59assistant and Its Behavior in these
  1636. 1:02:01conversations now because this is neural
  1637. 1:02:03networks we're not going to be
  1638. 1:02:04programming these explicitly in code
  1639. 1:02:07we're not going to be able to program
  1640. 1:02:08the assistant in that way because this
  1641. 1:02:10is neural networks everything is done
  1642. 1:02:12through neural network training on data
  1643. 1:02:14sets and so because of that we are going
  1644. 1:02:17to be implicitly programming the
  1645. 1:02:19assistant by creating data sets of
  1646. 1:02:21conversations so these are three
  1647. 1:02:23independent examples of conversations in
  1648. 1:02:25a data dat set an actual data set and
  1649. 1:02:27I'm going to show you examples will be
  1650. 1:02:29much larger it could have hundreds of
  1651. 1:02:31thousands of conversations that are
  1652. 1:02:32multi- turn very long Etc and would
  1653. 1:02:34cover a diverse breath of topics but
  1654. 1:02:37here I'm only showing three examples but
  1655. 1:02:39the way this works basically is uh a
  1656. 1:02:42assistant is being programmed by example
  1657. 1:02:45and where is this data coming from like
  1658. 1:02:472 * 2al 4 same as 2 plus 2 Etc where
  1659. 1:02:50does that come from this comes from
  1660. 1:02:51Human labelers so we will basically give
  1661. 1:02:54human labelers some conversational
  1662. 1:02:56context and we will ask them to um
  1663. 1:02:58basically give the ideal assistant
  1664. 1:03:00response in this situation and a human
  1665. 1:03:03will write out the ideal response for an
  1666. 1:03:06assistant in any situation and then
  1667. 1:03:08we're going to get the model to
  1668. 1:03:10basically train on this and to imitate
  1669. 1:03:12those kinds of
  1670. 1:03:14responses so the way this works then is
  1671. 1:03:16we are going to take our base model
  1672. 1:03:17which we produced in the preing stage
  1673. 1:03:20and this base model was trained on
  1674. 1:03:21internet documents we're now going to
  1675. 1:03:23take that data set of internet documents
  1676. 1:03:25and we're gonna throw it out and we're
  1677. 1:03:27going to substitute a new data set and
  1678. 1:03:29that's going to be a data set of
  1679. 1:03:30conversations and we're going to
  1680. 1:03:32continue training the model on these
  1681. 1:03:33conversations on this new data set of
  1682. 1:03:35conversations and what happens is that
  1683. 1:03:37the model will very rapidly adjust and
  1684. 1:03:40will sort of like learn the statistics
  1685. 1:03:42of how this assistant responds to human
  1686. 1:03:45queries and then later during inference
  1687. 1:03:48we'll be able to basically um Prime the
  1688. 1:03:51assistant and get the response and it
  1689. 1:03:54will be imitating what the humans will
  1690. 1:03:56human labelers would do in that
  1691. 1:03:57situation if that makes sense so we're
  1692. 1:04:00going to see examples of that and this
  1693. 1:04:01is going to become bit more concrete I
  1694. 1:04:03also wanted to mention that this
  1695. 1:04:05post-training stage we're going to
  1696. 1:04:06basically just continue training the
  1697. 1:04:07model but um the pre-training stage can
  1698. 1:04:10in practice take roughly three months of
  1699. 1:04:13training on many thousands of computers
  1700. 1:04:15the post-training stage will typically
  1701. 1:04:16be much shorter like 3 hours for example
  1702. 1:04:20um and that's because the data set of
  1703. 1:04:21conversations that we're going to create
  1704. 1:04:23here manually is much much smaller than
  1705. 1:04:26the data set of text on the internet and
  1706. 1:04:28so this training will be very short but
  1707. 1:04:31fundamentally we're just going to take
  1708. 1:04:33our base model we're going to continue
  1709. 1:04:35training using the exact same algorithm
  1710. 1:04:37the exact same everything except we're
  1711. 1:04:39swapping out the data set for
  1712. 1:04:40conversations so the questions now are
  1713. 1:04:43what are these conversations how do we
  1714. 1:04:44represent them how do we get the model
  1715. 1:04:46to see conversations instead of just raw
  1716. 1:04:49text and then what are the outcomes of
  1717. 1:04:52um this kind of training and what do you
  1718. 1:04:54get in a certain like psychological
  1719. 1:04:56sense uh when we talk about the model so
  1720. 1:04:58let's turn to those questions now so
  1721. 1:05:01let's start by talking about the
  1722. 1:05:02tokenization of conversations everything
  1723. 1:05:05in these models has to be turned into
  1724. 1:05:07tokens because everything is just about
  1725. 1:05:08token sequences so how do we turn
  1726. 1:05:10conversations into token sequences is
  1727. 1:05:12the question and so for that we need to
  1728. 1:05:15design some kind of ending coding and uh
  1729. 1:05:17this is kind of similar to maybe if
  1730. 1:05:18you're familiar you don't have to be
  1731. 1:05:20with for example the TCP IP packet in um
  1732. 1:05:23on the internet there are precise rules
  1733. 1:05:25and protocols for how you represent
  1734. 1:05:27information how everything is structured
  1735. 1:05:29together so that you have all this kind
  1736. 1:05:30of data laid out in a way that is
  1737. 1:05:32written out on a paper and that everyone
  1738. 1:05:34can agree on and so it's the same thing
  1739. 1:05:36now happening in llms we need some kind
  1740. 1:05:38of data structures and we need to have
  1741. 1:05:40some rules around how these data
  1742. 1:05:41structures like conversations get
  1743. 1:05:43encoded and decoded to and from tokens
  1744. 1:05:46and so I want to show you now how I
  1745. 1:05:48would
  1746. 1:05:49recreate uh this conversation in the
  1747. 1:05:52token space so if you go to Tech
  1748. 1:05:54tokenizer
  1749. 1:05:56I can take that conversation and this is
  1750. 1:05:58how it is represented in uh for the
  1751. 1:06:01language model so here we have we are
  1752. 1:06:03iterating a user and an assistant in
  1753. 1:06:06this two- turn
  1754. 1:06:08conversation and what you're seeing here
  1755. 1:06:10is it looks ugly but it's actually
  1756. 1:06:11relatively simple the way it gets turned
  1757. 1:06:13into a token sequence here at the end is
  1758. 1:06:16a little bit complicated but at the end
  1759. 1:06:18this conversation between a user and
  1760. 1:06:19assistant ends up being 49 tokens it is
  1761. 1:06:22a one-dimensional sequence of 49 tokens
  1762. 1:06:24and these are the tokens
  1763. 1:06:26okay and all the different llms will
  1764. 1:06:29have a slightly different format or
  1765. 1:06:31protocols and it's a little bit of a
  1766. 1:06:33wild west right now but for example GPT
  1767. 1:06:3640 does it in the following way you have
  1768. 1:06:39this special token called imore start
  1769. 1:06:42and this is short for IM imaginary
  1770. 1:06:44monologue uh the
  1771. 1:06:46start then you have to specify um I
  1772. 1:06:49don't actually know why it's called that
  1773. 1:06:50to be honest then you have to specify
  1774. 1:06:52whose turn it is so for example user
  1775. 1:06:54which is a token 4
  1776. 1:06:5628 then you have internal monologue
  1777. 1:07:00separator and then it's the exact
  1778. 1:07:03question so the tokens of the question
  1779. 1:07:05and then you have to close it so I am
  1780. 1:07:07end the end of the imaginary monologue
  1781. 1:07:09so
  1782. 1:07:10basically the question from a user of
  1783. 1:07:13what is 2 plus two ends up being the
  1784. 1:07:16token sequence of these tokens and now
  1785. 1:07:19the important thing to mention here is
  1786. 1:07:20that IM start this is not text right IM
  1787. 1:07:24start is a special token that gets added
  1788. 1:07:27it's a new token and um this token has
  1789. 1:07:30never been trained on so far it is a new
  1790. 1:07:32token that we create in a post-training
  1791. 1:07:34stage and we introduce and so these
  1792. 1:07:37special tokens like IM seep IM start Etc
  1793. 1:07:40are introduced and interspersed with
  1794. 1:07:42text so that they sort of um get the
  1795. 1:07:45model to learn that hey this is a the
  1796. 1:07:47start of a turn for who is it start of
  1797. 1:07:49the turn for the start of the turn is
  1798. 1:07:51for the user and then this is what the
  1799. 1:07:54user says and then the user ends and
  1800. 1:07:56then it's a new start of a turn and it
  1801. 1:07:58is by the assistant and then what does
  1802. 1:08:01the assistant say well these are the
  1803. 1:08:02tokens of what the assistant says Etc
  1804. 1:08:05and so this conversation is not turned
  1805. 1:08:06into the sequence of tokens the specific
  1806. 1:08:09details here are not actually that
  1807. 1:08:11important all I'm trying to show you in
  1808. 1:08:13concrete terms is that our conversations
  1809. 1:08:15which we think of as kind of like a
  1810. 1:08:16structured object end up being turned
  1811. 1:08:19via some encoding into onedimensional
  1812. 1:08:21sequences of tokens and so because this
  1813. 1:08:25is one dimensional sequence of tokens we
  1814. 1:08:27can apply all the stuff that we applied
  1815. 1:08:29before now it's just a sequence of
  1816. 1:08:30tokens and now we can train a language
  1817. 1:08:33model on it and so we're just predicting
  1818. 1:08:35the next token in a sequence uh just
  1819. 1:08:37like before and um we can represent and
  1820. 1:08:39train on conversations and then what
  1821. 1:08:42does it look like at test time during
  1822. 1:08:43inference so say we've trained a model
  1823. 1:08:46and we've trained a model on these kinds
  1824. 1:08:49of data sets of conversations and now we
  1825. 1:08:51want to
  1826. 1:08:52inference so during inference what does
  1827. 1:08:54this look like when you're on on chash
  1828. 1:08:55apt well you come to chash apt and you
  1829. 1:08:58have say like a dialogue with it and the
  1830. 1:09:01way this works is
  1831. 1:09:03basically um say that this was already
  1832. 1:09:06filled in so like what is 2 plus 2 2
  1833. 1:09:07plus 2 is four and now you issue what if
  1834. 1:09:10it was times I am end and what basically
  1835. 1:09:13ends up happening um on the servers of
  1836. 1:09:16open AI or something like that is they
  1837. 1:09:18put in I start assistant I amep and this
  1838. 1:09:21is where they end it right here so they
  1839. 1:09:24construct this context and now they
  1840. 1:09:27start sampling from the model so it's at
  1841. 1:09:29this stage that they will go to the
  1842. 1:09:30model and say okay what is a good for
  1843. 1:09:32sequence what is a good first token what
  1844. 1:09:34is a good second token what is a good
  1845. 1:09:36third token and this is where the LM
  1846. 1:09:38takes over and creates a response like
  1847. 1:09:41for example response that looks
  1848. 1:09:43something like this but it doesn't have
  1849. 1:09:44to be identical to this but it will have
  1850. 1:09:46the flavor of this if this kind of a
  1851. 1:09:48conversation was in the data set so um
  1852. 1:09:52that's roughly how the protocol Works
  1853. 1:09:54although the details of this protocol
  1854. 1:09:56are not important so again my goal is
  1855. 1:09:59that just to show you that everything
  1856. 1:10:01ends up being just a one-dimensional
  1857. 1:10:02token sequence so we can apply
  1858. 1:10:04everything we've already seen but we're
  1859. 1:10:06now training on conversations and we're
  1860. 1:10:08now uh basically generating
  1861. 1:10:10conversations as well okay so now I
  1862. 1:10:13would like to turn to what these data
  1863. 1:10:14sets look like in practice the first
  1864. 1:10:16paper that I would like to show you and
  1865. 1:10:17the first effort in this direction is
  1866. 1:10:20this paper from openai in 2022 and this
  1867. 1:10:23paper was called instruct GPT or the
  1868. 1:10:25technique that they developed and this
  1869. 1:10:27was the first time that opena has kind
  1870. 1:10:29of talked about how you can take
  1871. 1:10:30language models and fine-tune them on
  1872. 1:10:32conversations and so this paper has a
  1873. 1:10:34number of details that I would like to
  1874. 1:10:36take you through so the first stop I
  1875. 1:10:38would like to make is in section 3.4
  1876. 1:10:40where they talk about the human
  1877. 1:10:41contractors that they hired uh in this
  1878. 1:10:44case from upwork or through scale AI to
  1879. 1:10:47uh construct these conversations and so
  1880. 1:10:49there are human labelers involved whose
  1881. 1:10:52job it is professionally to create these
  1882. 1:10:54conversations and these labelers are
  1883. 1:10:56asked to come up with prompts and then
  1884. 1:10:58they are asked to also complete the
  1885. 1:11:00ideal assistant responses and so these
  1886. 1:11:03are the kinds of prompts that people
  1887. 1:11:04came up with so these are human labelers
  1888. 1:11:06so list five ideas for how to regain
  1889. 1:11:08enthusiasm for my career what are the
  1890. 1:11:10top 10 science fiction books I should
  1891. 1:11:12read next and there's many different
  1892. 1:11:13types of uh kind of prompts here so
  1893. 1:11:16translate this sentence from uh to
  1894. 1:11:18Spanish Etc and so there's many things
  1895. 1:11:21here that people came up with they first
  1896. 1:11:23come up with the prompt and then they
  1897. 1:11:25also uh answer that prompt and they give
  1898. 1:11:28the ideal assistant response now how do
  1899. 1:11:30they know what is the ideal assistant
  1900. 1:11:32response that they should write for
  1901. 1:11:33these prompts so when we scroll down a
  1902. 1:11:35little bit further we see that here we
  1903. 1:11:37have this excerpt of labeling
  1904. 1:11:39instructions uh that are given to the
  1905. 1:11:41human labelers so the company that is
  1906. 1:11:44developing the language model like for
  1907. 1:11:45example open AI writes up labeling
  1908. 1:11:47instructions for how the humans should
  1909. 1:11:49create ideal responses and so here for
  1910. 1:11:52example is an excerpt uh of these kinds
  1911. 1:11:54of labeling instruction instructions on
  1912. 1:11:56High level you're asking people to be
  1913. 1:11:57helpful truthful and harmless and you
  1914. 1:11:59can pause the video if you'd like to see
  1915. 1:12:01more here but on a high level basically
  1916. 1:12:04just just answer try to be helpful try
  1917. 1:12:06to be truthful and don't answer
  1918. 1:12:08questions that we don't want um kind of
  1919. 1:12:10the system to handle uh later in chat
  1920. 1:12:13gbt and so roughly speaking the company
  1921. 1:12:16comes up with the labeling instructions
  1922. 1:12:18usually they are not this short usually
  1923. 1:12:19there are hundreds of pages and people
  1924. 1:12:21have to study them professionally and
  1925. 1:12:23then they write out the ideal assistant
  1926. 1:12:26responses uh following those labeling
  1927. 1:12:28instructions so this is a very human
  1928. 1:12:30heavy process as it was described in
  1929. 1:12:32this paper now the data set for instruct
  1930. 1:12:34GPT was never actually released by openi
  1931. 1:12:37but we do have some open- Source um
  1932. 1:12:39reproductions that were're trying to
  1933. 1:12:40follow this kind of a setup and collect
  1934. 1:12:42their own data so one that I'm familiar
  1935. 1:12:45with for example is the effort of open
  1936. 1:12:48Assistant from a while back and this is
  1937. 1:12:50just one of I think many examples but I
  1938. 1:12:52just want to show you an example so
  1939. 1:12:54here's so these were people on the
  1940. 1:12:56internet that were asked to basically
  1941. 1:12:57create these conversations similar to
  1942. 1:12:59what um open I did with human labelers
  1943. 1:13:03and so here's an entry of a person who
  1944. 1:13:05came up with this BR can you write a
  1945. 1:13:07short introduction to the relevance of
  1946. 1:13:08the term
  1947. 1:13:09manop uh in economics please use
  1948. 1:13:12examples Etc and then the same person or
  1949. 1:13:15potentially a different person will
  1950. 1:13:17write up the response so here's the
  1951. 1:13:18assistant response to this and so then
  1952. 1:13:21the same person or different person will
  1953. 1:13:23actually write out this ideal
  1954. 1:13:26response and then this is an example of
  1955. 1:13:29maybe how the conversation could
  1956. 1:13:30continue now explain it to a dog and
  1957. 1:13:33then you can try to come up with a
  1958. 1:13:34slightly a simpler explanation or
  1959. 1:13:36something like that now this then
  1960. 1:13:39becomes the label and we end up training
  1961. 1:13:41on this so what happens during training
  1962. 1:13:45is that um of course we're not going to
  1963. 1:13:48have a full coverage of all the possible
  1964. 1:13:50questions that um the model will
  1965. 1:13:53encounter at test time during inference
  1966. 1:13:56we can't possibly cover all the possible
  1967. 1:13:57prompts that people are going to be
  1968. 1:13:59asking in the future but if we have a
  1969. 1:14:02like a data set of a few of these
  1970. 1:14:03examples then the model during training
  1971. 1:14:06will start to take on this Persona of
  1972. 1:14:09this helpful truthful harmless assistant
  1973. 1:14:12and it's all programmed by example and
  1974. 1:14:14so these are all examples of behavior
  1975. 1:14:16and if you have conversations of these
  1976. 1:14:18example behaviors and you have enough of
  1977. 1:14:19them like 100,00 and you train on it the
  1978. 1:14:22model sort of starts to understand the
  1979. 1:14:23statistical pattern and it kind of takes
  1980. 1:14:26on this personality of this
  1981. 1:14:28assistant now it's possible that when
  1982. 1:14:30you get the exact same question like
  1983. 1:14:32this at test time it's possible that the
  1984. 1:14:35answer will be recited as exactly what
  1985. 1:14:38was in the training set but more likely
  1986. 1:14:40than that is that the model will kind of
  1987. 1:14:43like do something of a similar Vibe um
  1988. 1:14:45and we will understand that this is the
  1989. 1:14:47kind of answer that you want um so
  1990. 1:14:51that's what we're doing we're
  1991. 1:14:52programming the system um by example and
  1992. 1:14:55the system adopts statistically this
  1993. 1:14:58Persona of this helpful truthful
  1994. 1:15:00harmless assistant which is kind of like
  1995. 1:15:02reflected in the labeling instructions
  1996. 1:15:04that the company creates now I want to
  1997. 1:15:06show you that the state-of-the-art has
  1998. 1:15:08kind of advanced in the last 2 or 3
  1999. 1:15:09years uh since the instr GPT paper so in
  2000. 1:15:12particular it's not very common for
  2001. 1:15:14humans to be doing all the heavy lifting
  2002. 1:15:16just by themselves anymore and that's
  2003. 1:15:18because we now have language models and
  2004. 1:15:19these language models are helping us
  2005. 1:15:21create these data sets and conversations
  2006. 1:15:23so it is very rare that the people will
  2007. 1:15:25like literally just write out the
  2008. 1:15:26response from scratch it is a lot more
  2009. 1:15:28likely that they will use an existing
  2010. 1:15:29llm to basically like uh come up with an
  2011. 1:15:32answer and then they will edit it or
  2012. 1:15:34things like that so there's many
  2013. 1:15:35different ways in which now llms have
  2014. 1:15:37started to kind of permeate this
  2015. 1:15:39posttraining Set uh stack and llms are
  2016. 1:15:43basically used pervasively to help
  2017. 1:15:45create these massive data sets of
  2018. 1:15:46conversations so I don't want to show
  2019. 1:15:49like Ultra chat is one um such example
  2020. 1:15:52of like a more modern data set of
  2021. 1:15:53conversations it is to a very large
  2022. 1:15:56extent synthetic but uh I believe
  2023. 1:15:58there's some human involvement I could
  2024. 1:15:59be wrong with that usually there will be
  2025. 1:16:01a little bit of human but there will be
  2026. 1:16:02a huge amount of synthetic help um and
  2027. 1:16:06this is all kind of like uh constructed
  2028. 1:16:08in different ways and Ultra chat is just
  2029. 1:16:10one example of many sft data sets that
  2030. 1:16:12currently exist and the only thing I
  2031. 1:16:14want to show you is that uh these data
  2032. 1:16:15sets have now millions of conversations
  2033. 1:16:18uh these conversations are mostly
  2034. 1:16:19synthetic but they're probably edited to
  2035. 1:16:21some extent by humans and they span a
  2036. 1:16:23huge diversity of sort of
  2037. 1:16:27um uh areas and so on so these are
  2038. 1:16:31fairly extensive artifacts by now and
  2039. 1:16:33there's all these like sft mixtures as
  2040. 1:16:35they're called so you have a mixture of
  2041. 1:16:37like lots of different types and sources
  2042. 1:16:39and it's partially synthetic partially
  2043. 1:16:41human and it's kind of like um gone in
  2044. 1:16:44that direction since uh but roughly
  2045. 1:16:46speaking we still have sft data sets
  2046. 1:16:48they're made up of conversations we're
  2047. 1:16:50training on them um just like we did
  2048. 1:16:52before and
  2049. 1:16:55uh I guess like the last thing to note
  2050. 1:16:57is that I want to dispel a little bit of
  2051. 1:17:00the magic of talking to an AI like when
  2052. 1:17:02you go to chat GPT and you give it a
  2053. 1:17:04question and then you hit enter uh what
  2054. 1:17:07is coming back is kind of like
  2055. 1:17:10statistically aligned with what's
  2056. 1:17:12happening in the training set and these
  2057. 1:17:14training sets I mean they really just
  2058. 1:17:16have a seed in humans following labeling
  2059. 1:17:19instructions so what are you actually
  2060. 1:17:21talking to in chat GPT or how should you
  2061. 1:17:24think about it well it's not coming from
  2062. 1:17:25some magical AI like roughly speaking
  2063. 1:17:28it's coming from something that is
  2064. 1:17:29statistically imitating human labelers
  2065. 1:17:32which comes from labeling instructions
  2066. 1:17:34written by these companies and so you're
  2067. 1:17:36kind of imitating this uh you're kind of
  2068. 1:17:38getting um it's almost as if you're
  2069. 1:17:40asking human labeler and imagine that
  2070. 1:17:43the answer that is given to you uh from
  2071. 1:17:45chbt is some kind of a simulation of a
  2072. 1:17:47human labeler uh and it's kind of like
  2073. 1:17:50asking what would a human labeler say in
  2074. 1:17:53this kind of a conversation
  2075. 1:17:56and uh it's not just like this human
  2076. 1:17:58labeler is not just like a random person
  2077. 1:18:00from the internet because these
  2078. 1:18:01companies actually hire experts so for
  2079. 1:18:03example when you are asking questions
  2080. 1:18:04about code and so on the human labelers
  2081. 1:18:06that would be in um involved in creation
  2082. 1:18:08of these conversation data sets they
  2083. 1:18:10will usually be usually be educated
  2084. 1:18:12expert people and you're kind of like
  2085. 1:18:15asking a question of like a simulation
  2086. 1:18:17of those people if that makes sense so
  2087. 1:18:19you're not talking to a magical AI
  2088. 1:18:21you're talking to an average labeler
  2089. 1:18:22this average labeler is probably fairly
  2090. 1:18:24highly skilled
  2091. 1:18:25but you're talking to kind of like an
  2092. 1:18:26instantaneous simulation of that kind of
  2093. 1:18:29a person that would be hired uh in the
  2094. 1:18:32construction of these data sets so let
  2095. 1:18:34me give you one more specific example
  2096. 1:18:36before we move on for example when I go
  2097. 1:18:38to chpt and I say recommend the top five
  2098. 1:18:40landmarks who see in Paris and then I
  2099. 1:18:42hit
  2100. 1:18:44enter
  2101. 1:18:49uh okay here we go okay when I hit enter
  2102. 1:18:52what's coming out here how do I think
  2103. 1:18:55about it well it's not some kind of a
  2104. 1:18:56magical AI that has gone out and
  2105. 1:18:58researched all the landmarks and then
  2106. 1:19:00ranked them using its infinite
  2107. 1:19:01intelligence Etc what I'm getting is a
  2108. 1:19:04statistical simulation of a labeler that
  2109. 1:19:07was hired by open AI you can think about
  2110. 1:19:09it roughly in that way and so if this
  2111. 1:19:13specific um question is in the
  2112. 1:19:16posttraining data set somewhere at open
  2113. 1:19:17aai then I'm very likely to see an
  2114. 1:19:20answer that is probably very very
  2115. 1:19:22similar to what that human labeler would
  2116. 1:19:24have put down
  2117. 1:19:25for those five landmarks how does the
  2118. 1:19:27human labeler come up with this well
  2119. 1:19:28they go off and they go on the internet
  2120. 1:19:29and they kind of do their own little
  2121. 1:19:31research for 20 minutes and they just
  2122. 1:19:32come up with a list right now so if they
  2123. 1:19:35come up with this list and this is in
  2124. 1:19:37the data set I'm probably very likely to
  2125. 1:19:39see what they submitted as the correct
  2126. 1:19:41answer from the assistant now if this
  2127. 1:19:44specific query is not part of the post
  2128. 1:19:46training data set then what I'm getting
  2129. 1:19:48here is a little bit more emergent uh
  2130. 1:19:51because uh the model kind of understands
  2131. 1:19:53the statistically
  2132. 1:19:55um the kinds of landmarks that are in
  2133. 1:19:57this training set are usually the
  2134. 1:19:59prominent landmarks the landmarks that
  2135. 1:20:00people usually want to see the kinds of
  2136. 1:20:02landmarks that are usually uh very often
  2137. 1:20:05talked about on the internet and
  2138. 1:20:06remember that the model already has a
  2139. 1:20:08ton of Knowledge from its pre-training
  2140. 1:20:10on the internet so it's probably seen a
  2141. 1:20:12ton of conversations about Paris about
  2142. 1:20:13landmarks about the kinds of things that
  2143. 1:20:15people like to see and so it's the
  2144. 1:20:17pre-training knowledge that has then
  2145. 1:20:18combined with the postering data set
  2146. 1:20:20that results in this kind of an
  2147. 1:20:23imitation um
  2148. 1:20:25so that's uh that's roughly how you can
  2149. 1:20:27kind of think about what's happening
  2150. 1:20:29behind the scenes here in in this
  2151. 1:20:31statistical sense okay now I want to
  2152. 1:20:33turn to the topic of llm psychology as I
  2153. 1:20:35like to call it which is what are sort
  2154. 1:20:37of the emergent cognitive effects of the
  2155. 1:20:40training pipeline that we have for these
  2156. 1:20:42models so in particular the first one I
  2157. 1:20:44want to talk to is of course
  2158. 1:20:47hallucinations so you might be familiar
  2159. 1:20:50with model hallucinations it's when llms
  2160. 1:20:52make stuff up they just totally
  2161. 1:20:53fabricate information Etc and it's a big
  2162. 1:20:56problem with llm assistants it is a
  2163. 1:20:58problem that existed to a large extent
  2164. 1:21:00with early models uh from many years ago
  2165. 1:21:02and I think the problem has gotten a bit
  2166. 1:21:04better uh because there are some
  2167. 1:21:05medications that I'm going to go into in
  2168. 1:21:07a second for now let's just try to
  2169. 1:21:09understand where these hallucinations
  2170. 1:21:10come from so here's a specific example
  2171. 1:21:13of a few uh of three conversations that
  2172. 1:21:16you might think you have in your
  2173. 1:21:17training set and um these are pretty
  2174. 1:21:20reasonable conversations that you could
  2175. 1:21:22imagine being in the training set so
  2176. 1:21:23like for example who is Cruz well Tom
  2177. 1:21:25Cruz is an famous actor American actor
  2178. 1:21:27and producer Etc who is John baraso this
  2179. 1:21:31turns out to be a us senetor for example
  2180. 1:21:34who is genis Khan well genis Khan was
  2181. 1:21:36blah blah blah and so this is what your
  2182. 1:21:39conversations could look like at
  2183. 1:21:40training time now the problem with this
  2184. 1:21:42is that when the human is writing the
  2185. 1:21:46correct answer for the assistant in each
  2186. 1:21:48one of these cases uh the human either
  2187. 1:21:51like knows who this person is or they
  2188. 1:21:52research them on the Internet and they
  2189. 1:21:53come in and they write this response
  2190. 1:21:55that kind of has this like confident
  2191. 1:21:57tone of an answer and what happens
  2192. 1:21:59basically is that at test time when you
  2193. 1:22:01ask for someone who is this is a totally
  2194. 1:22:03random name that I totally came up with
  2195. 1:22:05and I don't think this person exists um
  2196. 1:22:07as far as I know I just Tred to generate
  2197. 1:22:09it randomly the problem is when we ask
  2198. 1:22:11who is Orson kovats the problem is that
  2199. 1:22:15the assistant will not just tell you oh
  2200. 1:22:17I don't know even if the assistant and
  2201. 1:22:20the language model itself might know
  2202. 1:22:23inside its features inside its
  2203. 1:22:24activations inside of its brain sort of
  2204. 1:22:26it might know that this person is like
  2205. 1:22:28not someone that um that is that it's
  2206. 1:22:30familiar with even if some part of the
  2207. 1:22:32network kind of knows that in some sense
  2208. 1:22:35the uh saying that oh I don't know who
  2209. 1:22:37this is is is not going to happen
  2210. 1:22:40because the model statistically imitates
  2211. 1:22:42is training set in the training set the
  2212. 1:22:45questions of the form who is blah are
  2213. 1:22:47confidently answered with the correct
  2214. 1:22:49answer and so it's going to take on the
  2215. 1:22:52style of the answer and it's going to do
  2216. 1:22:53its best it's going to give you
  2217. 1:22:55statistically the most likely guess and
  2218. 1:22:57it's just going to basically make stuff
  2219. 1:22:58up because these models again we just
  2220. 1:23:01talked about it is they don't have
  2221. 1:23:02access to the internet they're not doing
  2222. 1:23:04research these are statistical token
  2223. 1:23:06tumblers as I call them uh is just
  2224. 1:23:08trying to sample the next token in the
  2225. 1:23:10sequence and it's going to basically
  2226. 1:23:12make stuff up so let's take a look at
  2227. 1:23:13what this looks
  2228. 1:23:15like I have here what's called the
  2229. 1:23:17inference playground from hugging face
  2230. 1:23:20and I am on purpose picking on a model
  2231. 1:23:22called Falcon 7B which is an old model
  2232. 1:23:25this is a few years ago now so it's an
  2233. 1:23:27older model So It suffers from
  2234. 1:23:28hallucinations and as I mentioned this
  2235. 1:23:31has improved over time recently but
  2236. 1:23:33let's say who is Orson kovats let's ask
  2237. 1:23:35Falcon 7B instruct
  2238. 1:23:37run oh yeah Orson kovat is an American
  2239. 1:23:40author and science uh fiction writer
  2240. 1:23:42okay this is totally false it's
  2241. 1:23:44hallucination let's try again these are
  2242. 1:23:46statistical systems right so we can
  2243. 1:23:48resample this time Orson kovat is a
  2244. 1:23:51fictional character from this 1950s TV
  2245. 1:23:53show it's total BS right let's try again
  2246. 1:23:57he's a former minor league baseball
  2247. 1:23:59player okay so basically the model
  2248. 1:24:02doesn't know and it's given us lots of
  2249. 1:24:04different answers because it doesn't
  2250. 1:24:06know it's just kind of like sampling
  2251. 1:24:08from these probabilities the model
  2252. 1:24:10starts with the tokens who is oron
  2253. 1:24:12kovats assistant and then it comes in
  2254. 1:24:14here and it's get it's getting these
  2255. 1:24:17probabilities and it's just sampling
  2256. 1:24:19from the probabilities and it just like
  2257. 1:24:20comes up with stuff and the stuff is
  2258. 1:24:24actually
  2259. 1:24:24statistically consistent with the style
  2260. 1:24:27of the answer in its training set and
  2261. 1:24:29it's just doing that but you and I
  2262. 1:24:31experiened it as a madeup factual
  2263. 1:24:33knowledge but keep in mind that uh the
  2264. 1:24:36model basically doesn't know and it's
  2265. 1:24:37just imitating the format of the answer
  2266. 1:24:40and it's not going to go off and look it
  2267. 1:24:41up uh because it's just imitating again
  2268. 1:24:44the answer so how can we uh mitigate
  2269. 1:24:47this because for example when we go to
  2270. 1:24:48chat apt and I say who is oron kovats
  2271. 1:24:50and I'm now asking the stateoftheart
  2272. 1:24:52state-of-the-art model from open AI
  2273. 1:24:55this model will tell
  2274. 1:24:56you oh so this model is actually is even
  2275. 1:25:00smarter because you saw very briefly it
  2276. 1:25:02said searching the web uh we're going to
  2277. 1:25:04cover this later um it's actually trying
  2278. 1:25:07to do tool use and
  2279. 1:25:11uh kind of just like came up with some
  2280. 1:25:13kind of a story but I want to just who
  2281. 1:25:15or Kovach did not use any tools I don't
  2282. 1:25:19want it to do web
  2283. 1:25:22search there's a wellknown historical or
  2284. 1:25:24public figure named or oron kovats so
  2285. 1:25:27this model is not going to make up stuff
  2286. 1:25:29this model knows that it doesn't know
  2287. 1:25:31and it tells you that it doesn't appear
  2288. 1:25:32to be a person that this model knows so
  2289. 1:25:35somehow we sort of improved
  2290. 1:25:37hallucinations even though they clearly
  2291. 1:25:39are an issue in older models and it
  2292. 1:25:42makes totally uh sense why you would be
  2293. 1:25:44getting these kinds of answers if this
  2294. 1:25:46is what your training set looks like so
  2295. 1:25:47how do we fix this okay well clearly we
  2296. 1:25:50need some examples in our data set that
  2297. 1:25:53where the correct answer for the
  2298. 1:25:54assistant is that the model doesn't know
  2299. 1:25:57about some particular fact but we only
  2300. 1:25:59need to have those answers be produced
  2301. 1:26:02in the cases where the model actually
  2302. 1:26:03doesn't know and so the question is how
  2303. 1:26:05do we know what the model knows or
  2304. 1:26:07doesn't know well we can empirically
  2305. 1:26:09probe the model to figure that out so
  2306. 1:26:11let's take a look at for example how
  2307. 1:26:13meta uh dealt with hallucinations for
  2308. 1:26:16the Llama 3 series of models as an
  2309. 1:26:18example so in this paper that they
  2310. 1:26:20published from meta we can go into
  2311. 1:26:22hallucinations
  2312. 1:26:25which they call here factuality and they
  2313. 1:26:27describe the procedure by which they
  2314. 1:26:29basically interrogate the model to
  2315. 1:26:32figure out what it knows and doesn't
  2316. 1:26:33know to figure out sort of like the
  2317. 1:26:35boundary of its knowledge and then they
  2318. 1:26:38add examples to the training set where
  2319. 1:26:41for the things where the model doesn't
  2320. 1:26:44know them the correct answer is that the
  2321. 1:26:46model doesn't know them which sounds
  2322. 1:26:48like a very easy thing to do in
  2323. 1:26:50principle but this roughly fixes the
  2324. 1:26:53issue and the the reason it fixes the
  2325. 1:26:54issue is
  2326. 1:26:56because remember like the model might
  2327. 1:26:59actually have a pretty good model of its
  2328. 1:27:01self knowledge inside the network so
  2329. 1:27:04remember we looked at the network and
  2330. 1:27:06all these neurons inside the network you
  2331. 1:27:08might imagine that there's a neuron
  2332. 1:27:09somewhere in the network that sort of
  2333. 1:27:11like lights up for when the model is
  2334. 1:27:14uncertain but the problem is that the
  2335. 1:27:17activation of that neuron is not
  2336. 1:27:18currently wired up to the model actually
  2337. 1:27:20saying in words that it doesn't know so
  2338. 1:27:23even though the internal of the neural
  2339. 1:27:24network no because there's some neurons
  2340. 1:27:26that represent that the model uh will
  2341. 1:27:29not surface that it will instead take
  2342. 1:27:31its best guess so that it sounds
  2343. 1:27:33confident um just like it sees in a
  2344. 1:27:35training set so we need to basically
  2345. 1:27:37interrogate the model and allow it to
  2346. 1:27:39say I don't know in the cases that it
  2347. 1:27:41doesn't know so let me take you through
  2348. 1:27:43what meta roughly does so basically what
  2349. 1:27:45they do is here I have an example uh
  2350. 1:27:48Dominic kek is uh the featured article
  2351. 1:27:51today so I just went there randomly and
  2352. 1:27:54what they do is basically they take a
  2353. 1:27:55random document in a training set and
  2354. 1:27:58they take a paragraph and then they use
  2355. 1:28:01an llm to construct questions about that
  2356. 1:28:04paragraph so for example I did that with
  2357. 1:28:06chat GPT
  2358. 1:28:09here so I said here's a paragraph from
  2359. 1:28:12this document generate three specific
  2360. 1:28:14factual questions based on this
  2361. 1:28:15paragraph and give me the questions and
  2362. 1:28:17the answers and so the llms are already
  2363. 1:28:20good enough to create and reframe this
  2364. 1:28:23information so if the information is in
  2365. 1:28:25the context window um of this llm this
  2366. 1:28:29actually works pretty well it doesn't
  2367. 1:28:30have to rely on its memory it's right
  2368. 1:28:33there in the context window and so it
  2369. 1:28:35can basically reframe that information
  2370. 1:28:37with fairly high accuracy so for example
  2371. 1:28:40can generate questions for us like for
  2372. 1:28:41which team did he play here's the answer
  2373. 1:28:44how many cups did he win Etc and now
  2374. 1:28:47what we have to do is we have some
  2375. 1:28:48question and answers and now we want to
  2376. 1:28:50interrogate the model so roughly
  2377. 1:28:51speaking what we'll do is we'll take our
  2378. 1:28:53questions and we'll go to our model
  2379. 1:28:55which would be uh say llama uh in meta
  2380. 1:28:59but let's just interrogate mol 7B here
  2381. 1:29:01as an example that's another model so
  2382. 1:29:04does this model know about this answer
  2383. 1:29:07let's take a
  2384. 1:29:09look uh so he played for Buffalo Sabers
  2385. 1:29:12right so the model knows and the the way
  2386. 1:29:15that you can programmatically decide is
  2387. 1:29:16basically we're going to take this
  2388. 1:29:18answer from the model and we're going to
  2389. 1:29:20compare it to the correct answer and
  2390. 1:29:23again the model model are good enough to
  2391. 1:29:24do this automatically so there's no
  2392. 1:29:26humans involved here we can take uh
  2393. 1:29:28basically the answer from the model and
  2394. 1:29:30we can use another llm judge to check if
  2395. 1:29:33that is correct according to this answer
  2396. 1:29:35and if it is correct that means that the
  2397. 1:29:37model probably knows so what we're going
  2398. 1:29:38to do is we're going to do this maybe a
  2399. 1:29:40few times so okay it knows it's Buffalo
  2400. 1:29:42Savers let's drag
  2401. 1:29:45in um Buffalo Sabers let's try one more
  2402. 1:29:51time Buffalo Sabers so we asked three
  2403. 1:29:54times about this factual question and
  2404. 1:29:55the model seems to know so everything is
  2405. 1:29:58great now let's try the second question
  2406. 1:30:00how many Stanley Cups did he
  2407. 1:30:02win and again let's interrogate the
  2408. 1:30:04model about that and the correct answer
  2409. 1:30:06is
  2410. 1:30:08two so um here the model claims that he
  2411. 1:30:13won um four times which is not correct
  2412. 1:30:17right it doesn't match two so the model
  2413. 1:30:20doesn't know it's making stuff up let's
  2414. 1:30:22try again
  2415. 1:30:27um so here the model again it's kind of
  2416. 1:30:30like making stuff up right let's
  2417. 1:30:34Dragon here it says did he did not even
  2418. 1:30:37did not win during his career so
  2419. 1:30:39obviously the model doesn't know and the
  2420. 1:30:41way we can programmatically tell again
  2421. 1:30:42is we interrogate the model three times
  2422. 1:30:45and we compare its answers maybe three
  2423. 1:30:47times five times whatever it is to the
  2424. 1:30:49correct answer and if the model doesn't
  2425. 1:30:51know then we know that the model doesn't
  2426. 1:30:53know this question
  2427. 1:30:54and then what we do is we take this
  2428. 1:30:56question we create a new conversation in
  2429. 1:30:59the training set so we're going to add a
  2430. 1:31:01new conversation training set and when
  2431. 1:31:03the question is how many Stanley Cups
  2432. 1:31:05did he win the answer is I'm sorry I
  2433. 1:31:08don't know or I don't remember and
  2434. 1:31:10that's the correct answer for this
  2435. 1:31:12question because we interrogated the
  2436. 1:31:13model and we saw that that's the case if
  2437. 1:31:15you do this for many different types of
  2438. 1:31:18uh questions for many different types of
  2439. 1:31:20documents you are giving the model an
  2440. 1:31:23opportunity to in its training set
  2441. 1:31:25refuse to say based on its knowledge and
  2442. 1:31:28if you just have a few examples of that
  2443. 1:31:30in your training set the model will know
  2444. 1:31:33um and and has the opportunity to learn
  2445. 1:31:35the association of this knowledge-based
  2446. 1:31:37refusal to this internal neuron
  2447. 1:31:41somewhere in its Network that we presume
  2448. 1:31:43exists and empirically this turns out to
  2449. 1:31:45be probably the case and it can learn
  2450. 1:31:47that Association that hey when this
  2451. 1:31:49neuron of uncertainty is high then I
  2452. 1:31:52actually don't know and I'm allowed to
  2453. 1:31:54say that I'm sorry but I don't think I
  2454. 1:31:56remember this Etc and if you have these
  2455. 1:31:59uh examples in your training set then
  2456. 1:32:01this is a large mitigation for
  2457. 1:32:03hallucination and that's roughly
  2458. 1:32:05speaking why chpt is able to do stuff
  2459. 1:32:08like this as well so these are kinds of
  2460. 1:32:10uh mitigations that people have
  2461. 1:32:12implemented and that have improved the
  2462. 1:32:14factuality issue over time okay so I've
  2463. 1:32:16described mitigation number one for
  2464. 1:32:19basically mitigating the hallucinations
  2465. 1:32:21issue now we can actually do much better
  2466. 1:32:24than that uh it's instead of just saying
  2467. 1:32:27that we don't know uh we can introduce
  2468. 1:32:29an additional mitigation number two to
  2469. 1:32:32give the llm an opportunity to be
  2470. 1:32:33factual and actually answer the question
  2471. 1:32:36now what do you and I do if I was to ask
  2472. 1:32:39you a factual question and you don't
  2473. 1:32:40know uh what would you do um in order to
  2474. 1:32:43answer the question well you could uh go
  2475. 1:32:45off and do some search and uh use the
  2476. 1:32:47internet and you could figure out the
  2477. 1:32:49answer and then tell me what that answer
  2478. 1:32:51is and we can do the exact exact same
  2479. 1:32:54thing with these models so think of the
  2480. 1:32:56knowledge inside the neural network
  2481. 1:32:58inside its billions of parameters think
  2482. 1:33:01of that as kind of a vague recollection
  2483. 1:33:02of the things that the model has seen
  2484. 1:33:05during its training during the
  2485. 1:33:07pre-training stage a long time ago so
  2486. 1:33:09think of that knowledge in the
  2487. 1:33:10parameters as something you read a month
  2488. 1:33:13ago and if you keep reading something
  2489. 1:33:15then you will remember it and the model
  2490. 1:33:17remembers that but if it's something
  2491. 1:33:18rare then you probably don't have a
  2492. 1:33:20really good recollection of that
  2493. 1:33:21information but what you and I do is we
  2494. 1:33:23just go and look it up now when you go
  2495. 1:33:25and look it up what you're doing
  2496. 1:33:26basically is like you're refreshing your
  2497. 1:33:28working memory with information and then
  2498. 1:33:30you're able to sort of like retrieve it
  2499. 1:33:32talk about it or Etc so we need some
  2500. 1:33:34equivalent of allowing the model to
  2501. 1:33:36refresh its memory or its recollection
  2502. 1:33:38and we can do that by introducing tools
  2503. 1:33:41uh for the
  2504. 1:33:42models so the way we are going to
  2505. 1:33:44approach this is that instead of just
  2506. 1:33:45saying hey I'm sorry I don't know we can
  2507. 1:33:48attempt to use tools so we can create uh
  2508. 1:33:53a mechanism
  2509. 1:33:54by which the language model can emit
  2510. 1:33:56special tokens and these are tokens that
  2511. 1:33:57we're going to introduce new tokens so
  2512. 1:34:00for example here I've introduced two
  2513. 1:34:02tokens and I've introduced a format or a
  2514. 1:34:04protocol for how the model is allowed to
  2515. 1:34:07use these tokens so for example instead
  2516. 1:34:09of answering the question when the model
  2517. 1:34:12does not instead of just saying I don't
  2518. 1:34:14know sorry the model has the option now
  2519. 1:34:16to emitting the special token search
  2520. 1:34:18start and this is the query that will go
  2521. 1:34:20to like bing.com in the case of openai
  2522. 1:34:22or say Google search or something like
  2523. 1:34:24that so it will emit the query and then
  2524. 1:34:26it will emit search end and then here
  2525. 1:34:30what will happen is that the program
  2526. 1:34:32that is sampling from the model that is
  2527. 1:34:34running the inference when it sees the
  2528. 1:34:36special token search end instead of
  2529. 1:34:39sampling the next token uh in the
  2530. 1:34:41sequence it will actually pause
  2531. 1:34:44generating from the model it will go off
  2532. 1:34:46it will open a session with bing.com and
  2533. 1:34:49it will paste the search query into Bing
  2534. 1:34:52and it will then um get all the text
  2535. 1:34:54that is retrieved and it will basically
  2536. 1:34:56take that text it will maybe represent
  2537. 1:34:58it again with some other special tokens
  2538. 1:35:00or something like that and it will take
  2539. 1:35:02that text and it will copy paste it here
  2540. 1:35:05into what I Tred to like show with the
  2541. 1:35:07brackets so all that text kind of comes
  2542. 1:35:09here and when the text comes here it
  2543. 1:35:12enters the context window so the model
  2544. 1:35:15so that text from the web search is now
  2545. 1:35:17inside the context window that will feed
  2546. 1:35:20into the neural network and you should
  2547. 1:35:21think of the context window as kind of
  2548. 1:35:23like the working memory of the model
  2549. 1:35:25that data that is in the context window
  2550. 1:35:27is directly accessible by the model it
  2551. 1:35:29directly feeds into the neural network
  2552. 1:35:31so it's not anymore a vague recollection
  2553. 1:35:33it's data that it it has in the context
  2554. 1:35:36window and is directly available to that
  2555. 1:35:38model so now when it's sampling the new
  2556. 1:35:41uh tokens here afterwards it can
  2557. 1:35:43reference very easily the data that has
  2558. 1:35:45been copy pasted in there so that's
  2559. 1:35:48roughly how these um how these tools use
  2560. 1:35:52uh tools uh function
  2561. 1:35:54and so web search is just one of the
  2562. 1:35:55tools we're going to look at some of the
  2563. 1:35:56other tools in a bit uh but basically
  2564. 1:35:59you introduce new tokens you introduce
  2565. 1:36:00some schema by which the model can
  2566. 1:36:02utilize these tokens and can call these
  2567. 1:36:04special functions like web search
  2568. 1:36:06functions and how do you teach the model
  2569. 1:36:08how to correctly use these tools like
  2570. 1:36:10say web search search start search end
  2571. 1:36:12Etc well again you do that through
  2572. 1:36:14training sets so we need now to have a
  2573. 1:36:16bunch of data and a bunch of
  2574. 1:36:18conversations that show the model by
  2575. 1:36:21example how to use web search so what
  2576. 1:36:24are the what are the settings where you
  2577. 1:36:25are using the search um and what does
  2578. 1:36:28that look like and here's by example how
  2579. 1:36:30you start a search and the search Etc
  2580. 1:36:33and uh if you have a few thousand maybe
  2581. 1:36:35examples of that in your training set
  2582. 1:36:36the model will actually do a pretty good
  2583. 1:36:38job of understanding uh how this tool
  2584. 1:36:40works and it will know how to sort of
  2585. 1:36:43structure its queries and of course
  2586. 1:36:44because of the pre-training data set and
  2587. 1:36:47its understanding of the world it
  2588. 1:36:48actually kind of understands what a web
  2589. 1:36:49search is and so it actually kind of has
  2590. 1:36:51a pretty good native understanding
  2591. 1:36:54um of what kind of stuff is a good
  2592. 1:36:56search query um and so it all kind of
  2593. 1:36:58just like works you just need a little
  2594. 1:37:00bit of a few examples to show it how to
  2595. 1:37:02use this new tool and then it can lean
  2596. 1:37:04on it to retrieve information and uh put
  2597. 1:37:07it in the context window and that's
  2598. 1:37:08equivalent to you and I looking
  2599. 1:37:10something up because once it's in the
  2600. 1:37:12context it's in the working memory and
  2601. 1:37:13it's very easy to manipulate and access
  2602. 1:37:16so that's what we saw a few minutes ago
  2603. 1:37:18when I was searching on chat GPT for who
  2604. 1:37:20is Orson kovats the chat GPT language
  2605. 1:37:23model decided Ed that this is some kind
  2606. 1:37:24of a rare um individual or something
  2607. 1:37:27like that and instead of giving me an
  2608. 1:37:29answer from its memory it decided that
  2609. 1:37:31it will sample a special token that is
  2610. 1:37:33going to do web search and we saw
  2611. 1:37:35briefly something flash it was like
  2612. 1:37:36using the web tool or something like
  2613. 1:37:38that so it briefly said that and then we
  2614. 1:37:40waited for like two seconds and then it
  2615. 1:37:41generated this and you see how it's
  2616. 1:37:43creating references here and so it's
  2617. 1:37:45citing sources so what happened here is
  2618. 1:37:50it went off it did a web web search it
  2619. 1:37:52found these sources and these URLs and
  2620. 1:37:55the text of these web pages was all
  2621. 1:37:58stuffed in between here and it's not
  2622. 1:38:01showing here but it's it's basically
  2623. 1:38:02stuffed as text in between here and now
  2624. 1:38:06it sees that text and now it kind of
  2625. 1:38:08references it and says that okay it
  2626. 1:38:11could be these people citation could be
  2627. 1:38:13those people citation Etc so that's what
  2628. 1:38:15happened here and that's what and that's
  2629. 1:38:17why when I said who is Orson kovats I
  2630. 1:38:19could also say don't use any tools and
  2631. 1:38:22then that's enough to um
  2632. 1:38:24basically convince chat PT to not use
  2633. 1:38:25tools and just use its memory and its
  2634. 1:38:28recollection I also went off and I um
  2635. 1:38:32tried to ask this question of Chachi PT
  2636. 1:38:34so how many standing cups did uh Dominic
  2637. 1:38:37Hasek win and Chachi P actually decided
  2638. 1:38:39that it knows the answer and it has the
  2639. 1:38:40confidence to say that uh he want twice
  2640. 1:38:43and so it kind of just relied on its
  2641. 1:38:45memory because presumably it has um it
  2642. 1:38:49has enough of
  2643. 1:38:50a kind of confidence in its weights in
  2644. 1:38:53it parameters and activations that this
  2645. 1:38:55is uh retrievable just for memory um but
  2646. 1:38:59you can also
  2647. 1:39:01conversely use web search to make sure
  2648. 1:39:04and then for the same query it actually
  2649. 1:39:06goes off and it searches and then it
  2650. 1:39:07finds a bunch of sources it finds all
  2651. 1:39:10this all of this stuff gets copy pasted
  2652. 1:39:12in there and then it tells us uh to
  2653. 1:39:15again and sites and it actually says the
  2654. 1:39:17Wikipedia article which is the source of
  2655. 1:39:20this information for us as well so
  2656. 1:39:23that's tools web search the model
  2657. 1:39:25determines when to search and then uh
  2658. 1:39:27that's kind of like how these tools uh
  2659. 1:39:29work and this is an additional kind of
  2660. 1:39:32mitigation for uh hallucinations and
  2661. 1:39:34factuality so I want to stress one more
  2662. 1:39:37time this very important sort of
  2663. 1:39:38psychology
  2664. 1:39:40Point knowledge in the parameters of the
  2665. 1:39:43neural network is a vague recollection
  2666. 1:39:45the knowledge in the tokens that make up
  2667. 1:39:47the context
  2668. 1:39:48window is the working memory and it
  2669. 1:39:51roughly speaking Works kind of like um
  2670. 1:39:53it works for us in our brain the stuff
  2671. 1:39:55we remember is our parameters uh and the
  2672. 1:39:58stuff that we just experienced like a
  2673. 1:40:01few seconds or minutes ago and so on you
  2674. 1:40:03can imagine that being in our context
  2675. 1:40:04window and this context window is being
  2676. 1:40:05built up as you have a conscious
  2677. 1:40:07experience around you so this has a
  2678. 1:40:10bunch of um implications also for your
  2679. 1:40:12use of LOLs in practice so for example I
  2680. 1:40:15can go to chat GPT and I can do
  2681. 1:40:17something like this I can say can you
  2682. 1:40:18Summarize chapter one of Jane Austin's
  2683. 1:40:20Pride and Prejudice right and this is a
  2684. 1:40:22perfectly fine prompt and Chach actually
  2685. 1:40:25does something relatively reasonable
  2686. 1:40:26here and but the reason it does that is
  2687. 1:40:28because Chach has a pretty good
  2688. 1:40:30recollection of a famous work like Pride
  2689. 1:40:32and Prejudice it's probably seen a ton
  2690. 1:40:34of stuff about it there's probably
  2691. 1:40:35forums about this book it's probably
  2692. 1:40:37read versions of this book um and it's
  2693. 1:40:40kind of like remembers because even if
  2694. 1:40:43you've read this or articles about it
  2695. 1:40:46you'd kind of have a recollection enough
  2696. 1:40:48to actually say all this but usually
  2697. 1:40:49when I actually interact with LMS and I
  2698. 1:40:51want them to recall specific things it
  2699. 1:40:53always works better if you just give it
  2700. 1:40:55to them so I think a much better prompt
  2701. 1:40:57would be something like this can you
  2702. 1:40:59summarize for me chapter one of genos's
  2703. 1:41:01spr and Prejudice and then I am
  2704. 1:41:03attaching it below for your reference
  2705. 1:41:04and then I do something like a delimeter
  2706. 1:41:06here and I paste it in and I I found
  2707. 1:41:08that just copy pasting it from some
  2708. 1:41:10website that I found here um so copy
  2709. 1:41:14pasting the chapter one here and I do
  2710. 1:41:16that because when it's in the context
  2711. 1:41:17window the model has direct access to it
  2712. 1:41:20and can exactly it doesn't have to
  2713. 1:41:22recall it it just has access to it and
  2714. 1:41:24so this summary is can be expected to be
  2715. 1:41:27a significantly high quality or higher
  2716. 1:41:29quality than this summary uh just
  2717. 1:41:31because it's directly available to the
  2718. 1:41:32model and I think you and I would work
  2719. 1:41:34in the same way if you want to it would
  2720. 1:41:36be you would produce a much better
  2721. 1:41:37summary if you had reread this chapter
  2722. 1:41:40before you had to summarize it and
  2723. 1:41:42that's basically what's happening here
  2724. 1:41:44or the equivalent of it the next sort of
  2725. 1:41:47psychological Quirk I'd like to talk
  2726. 1:41:48about briefly is that of the knowledge
  2727. 1:41:50of self so what I see very often on the
  2728. 1:41:52internet is that people do something
  2729. 1:41:54like this they ask llms something like
  2730. 1:41:56what model are you and who built you and
  2731. 1:41:59um basically this uh question is a
  2732. 1:42:01little bit nonsensical and the reason I
  2733. 1:42:03say that is that as I try to kind of
  2734. 1:42:05explain with some of the underhood
  2735. 1:42:07fundamentals this thing is not a person
  2736. 1:42:09right it doesn't have a persistent
  2737. 1:42:11existence in any way it sort of boots up
  2738. 1:42:14processes tokens and shuts off and it
  2739. 1:42:17does that for every single person it
  2740. 1:42:18just kind of builds up a context window
  2741. 1:42:19of conversation and then everything gets
  2742. 1:42:21deleted and so this this entity is kind
  2743. 1:42:23of like restarted from scratch every
  2744. 1:42:25single conversation if that makes sense
  2745. 1:42:27it has no persistent self it has no
  2746. 1:42:28sense of self it's a token tumbler and
  2747. 1:42:31uh it follows the statistical
  2748. 1:42:33regularities of its training set so it
  2749. 1:42:35doesn't really make sense to ask it who
  2750. 1:42:38are you what build you Etc and by
  2751. 1:42:40default if you do what I described and
  2752. 1:42:42just by default and from nowhere you're
  2753. 1:42:44going to get some pretty random answers
  2754. 1:42:46so for example let's uh pick on Falcon
  2755. 1:42:48which is a fairly old model and let's
  2756. 1:42:50see what it tells
  2757. 1:42:51us uh so it's evading the question uh
  2758. 1:42:55talented engineers and developers here
  2759. 1:42:58it says I was built by open AI based on
  2760. 1:42:59the gpt3 model it's totally making stuff
  2761. 1:43:01up now the fact that it's built by open
  2762. 1:43:04AI here I think a lot of people would
  2763. 1:43:06take this as evidence that this model
  2764. 1:43:07was somehow trained on open AI data or
  2765. 1:43:09something like that I don't actually
  2766. 1:43:10think that that's necessarily true the
  2767. 1:43:12reason for that is
  2768. 1:43:14that if you don't explicitly program the
  2769. 1:43:17model to answer these kinds of questions
  2770. 1:43:20then what you're going to get is its
  2771. 1:43:22statistical best guess at the answer and
  2772. 1:43:25this model had a um sft data mixture of
  2773. 1:43:29conversations and during the
  2774. 1:43:32fine-tuning um the model sort of
  2775. 1:43:35understands as it's training on this
  2776. 1:43:36data that it's taking on this
  2777. 1:43:38personality of this like helpful
  2778. 1:43:40assistant and it doesn't know how to it
  2779. 1:43:42doesn't actually it wasn't told exactly
  2780. 1:43:44what label to apply to self it just kind
  2781. 1:43:47of is taking on this uh this uh Persona
  2782. 1:43:50of a helpful assistant and remember that
  2783. 1:43:53the pre-training stage took the
  2784. 1:43:55documents from the entire internet and
  2785. 1:43:57Chach and open AI are very prominent in
  2786. 1:43:59these documents and so I think what's
  2787. 1:44:01actually likely to be happening here is
  2788. 1:44:03that this is just its hallucinated label
  2789. 1:44:06for what it is this is its self-identity
  2790. 1:44:08is that it's chat GPT by open Ai and
  2791. 1:44:11it's only saying that because there's a
  2792. 1:44:12ton of data on the internet of um
  2793. 1:44:15answers like this that are actually
  2794. 1:44:17coming from open from chasht and So
  2795. 1:44:20that's its label for what it is now you
  2796. 1:44:23can override this as a developer if you
  2797. 1:44:25have a llm model you can actually
  2798. 1:44:27override it and there are a few ways to
  2799. 1:44:28do that so for example let me show you
  2800. 1:44:31there's this MMO model from Allen Ai and
  2801. 1:44:35um this is one llm it's not a top tier
  2802. 1:44:37LM or anything like that but I like it
  2803. 1:44:39because it is fully open source so the
  2804. 1:44:41paper for Almo and everything else is
  2805. 1:44:43completely fully open source which is
  2806. 1:44:44nice um so here we are looking at its
  2807. 1:44:47sft mixture so this is the data mixture
  2808. 1:44:49of um the fine tuning so this is the
  2809. 1:44:52conversations data it right and so the
  2810. 1:44:54way that they are solving it for Theo
  2811. 1:44:56model is we see that there's a bunch of
  2812. 1:44:58stuff in the mixture and there's a total
  2813. 1:44:59of 1 million conversations here but here
  2814. 1:45:02we have alot to hardcoded if we go there
  2815. 1:45:05we see that this is 240
  2816. 1:45:07conversations and look at these 240
  2817. 1:45:10conversations they're hardcoded tell me
  2818. 1:45:12about yourself says user and then the
  2819. 1:45:15assistant says I'm and open language
  2820. 1:45:17model developed by AI to Allen Institute
  2821. 1:45:19of artificial intelligence Etc I'm here
  2822. 1:45:21to help blah blah blah what is your name
  2823. 1:45:23uh Theo project so these are all kinds
  2824. 1:45:26of like cooked up hardcoded questions
  2825. 1:45:27abouto 2 and the correct answers to give
  2826. 1:45:30in these cases if you take 240 questions
  2827. 1:45:33like this or conversations put them into
  2828. 1:45:35your training set and fine tune with it
  2829. 1:45:37then the model will actually be expected
  2830. 1:45:39to parot this stuff later if you don't
  2831. 1:45:43give it this then it's probably a Chach
  2832. 1:45:45by open
  2833. 1:45:46Ai and um there's one more way to
  2834. 1:45:49sometimes do this is
  2835. 1:45:51that basically um in these conversations
  2836. 1:45:55and you have terms between human and
  2837. 1:45:56assistant sometimes there's a special
  2838. 1:45:58message called system message at the
  2839. 1:46:00very beginning of the conversation so
  2840. 1:46:02it's not just between human and
  2841. 1:46:03assistant there's a system and in the
  2842. 1:46:05system message you can actually hardcode
  2843. 1:46:07and remind the model that hey you are a
  2844. 1:46:10model developed by open Ai and your name
  2845. 1:46:13is chashi pt40 and you were trained on
  2846. 1:46:16this date and your knowledge cut off is
  2847. 1:46:18this and basically it kind of like
  2848. 1:46:19documents the model a little bit and
  2849. 1:46:21then this is inserted into to your
  2850. 1:46:23conversations so when you go on chpt you
  2851. 1:46:25see a blank page but actually the system
  2852. 1:46:27message is kind of like hidden in there
  2853. 1:46:28and those tokens are in the context
  2854. 1:46:30window and so those are the two ways to
  2855. 1:46:33kind of um program the models to talk
  2856. 1:46:35about themselves either it's done
  2857. 1:46:37through uh data like this or it's done
  2858. 1:46:40through system message and things like
  2859. 1:46:42that basically invisible tokens that are
  2860. 1:46:44in the context window and remind the
  2861. 1:46:45model of its identity but it's all just
  2862. 1:46:47kind of like cooked up and bolted on in
  2863. 1:46:50some in some way it's not actually like
  2864. 1:46:51really deeply there in any real sense as
  2865. 1:46:54it would before a human I want to now
  2866. 1:46:57continue to the next section which deals
  2867. 1:46:59with the computational capabilities or
  2868. 1:47:01like I should say the native
  2869. 1:47:02computational capabilities of these
  2870. 1:47:03models in problem solving scenarios and
  2871. 1:47:06so in particular we have to be very
  2872. 1:47:07careful with these models when we
  2873. 1:47:09construct our examples of conversations
  2874. 1:47:11and there's a lot of sharp edges here
  2875. 1:47:13that are kind of like elucidative is
  2876. 1:47:15that a word uh they're kind of like
  2877. 1:47:16interesting to look at when we consider
  2878. 1:47:18how these models think so um consider
  2879. 1:47:22the following prompt from a human and
  2880. 1:47:24supposed that basically that we are
  2881. 1:47:25building out a conversation to enter
  2882. 1:47:27into our training set of conversations
  2883. 1:47:29so we're going to train the model on
  2884. 1:47:30this we're teaching you how to basically
  2885. 1:47:32solve simple math problems so the prompt
  2886. 1:47:34is Emily buys three apples and two
  2887. 1:47:36oranges each orange cost $2 the total
  2888. 1:47:38cost is 13 what is the cost of apples
  2889. 1:47:41very simple math question now there are
  2890. 1:47:43two answers here on the left and on the
  2891. 1:47:45right they are both correct answers they
  2892. 1:47:48both say that the answer is three which
  2893. 1:47:49is correct but one of these two is a
  2894. 1:47:52significant ific anly better answer for
  2895. 1:47:54the assistant than the other like if I
  2896. 1:47:56was Data labeler and I was creating one
  2897. 1:47:57of these one of these would be uh a
  2898. 1:48:01really terrible answer for the assistant
  2899. 1:48:03and the other would be okay and so I'd
  2900. 1:48:05like you to potentially pause the video
  2901. 1:48:07Even and think through why one of these
  2902. 1:48:09two is significantly better answer uh
  2903. 1:48:12than the other and um if you use the
  2904. 1:48:14wrong one your model will actually be uh
  2905. 1:48:17really bad at math potentially and it
  2906. 1:48:19would have uh bad outcomes and this is
  2907. 1:48:21something that you would be careful with
  2908. 1:48:22in your life labeling documentations
  2909. 1:48:23when you are training people uh to
  2910. 1:48:25create the ideal responses for the
  2911. 1:48:27assistant okay so the key to this
  2912. 1:48:29question is to realize and remember that
  2913. 1:48:32when the models are training and also
  2914. 1:48:34inferencing they are working in
  2915. 1:48:35onedimensional sequence of tokens from
  2916. 1:48:37left to right and this is the picture
  2917. 1:48:40that I often have in my mind I imagine
  2918. 1:48:42basically the token sequence evolving
  2919. 1:48:43from left to right and to always produce
  2920. 1:48:46the next token in a sequence we are
  2921. 1:48:48feeding all these tokens into the neural
  2922. 1:48:50network and this neural network then is
  2923. 1:48:53the probabilities for the next token and
  2924. 1:48:54sequence right so this picture here is
  2925. 1:48:56the exact same picture we saw uh before
  2926. 1:48:58up here and this comes from the web demo
  2927. 1:49:01that I showed you before right so this
  2928. 1:49:04is the calculation that basically takes
  2929. 1:49:05the input tokens here on the top and uh
  2930. 1:49:09performs these operations of all these
  2931. 1:49:11neurons and uh gives you the answer for
  2932. 1:49:13the probabilities of what comes next now
  2933. 1:49:15the important thing to realize is that
  2934. 1:49:17roughly
  2935. 1:49:19speaking uh there's basically a finite
  2936. 1:49:21number of layers of computation that
  2937. 1:49:22happened here so for example this model
  2938. 1:49:25here has only one two three layers of
  2939. 1:49:28what's called detention and uh MLP here
  2940. 1:49:31um maybe um typical modern
  2941. 1:49:34state-of-the-art Network would have more
  2942. 1:49:36like say 100 layers or something like
  2943. 1:49:37that but there's only 100 layers of
  2944. 1:49:39computation or something like that to go
  2945. 1:49:40from the previous token sequence to the
  2946. 1:49:42probabilities for the next token and so
  2947. 1:49:44there's a finite amount of computation
  2948. 1:49:46that happens here for every single token
  2949. 1:49:49and you should think of this as a very
  2950. 1:49:50small amount of computation and this
  2951. 1:49:52amount of computation is almost roughly
  2952. 1:49:54fixed uh for every single token in this
  2953. 1:49:57sequence um the that's not actually
  2954. 1:49:59fully true because the more tokens you
  2955. 1:50:01feed in uh the the more expensive uh
  2956. 1:50:04this forward pass will be of this neural
  2957. 1:50:06network but not by much so you should
  2958. 1:50:09think of this uh and I think as a good
  2959. 1:50:10model to have in mind this is a fixed
  2960. 1:50:12amount of compute that's going to happen
  2961. 1:50:13in this box for every single one of
  2962. 1:50:15these tokens and this amount of compute
  2963. 1:50:17Cann possibly be too big because there's
  2964. 1:50:19not that many layers that are sort of
  2965. 1:50:21going from the top to bottom here
  2966. 1:50:23there's not that that much
  2967. 1:50:24computationally that will happen here
  2968. 1:50:26and so you can't imagine the model to to
  2969. 1:50:27basically do arbitrary computation in a
  2970. 1:50:29single forward pass to get a single
  2971. 1:50:31token and so what that means is that we
  2972. 1:50:34actually have to distribute our
  2973. 1:50:35reasoning and our computation across
  2974. 1:50:37many tokens because every single token
  2975. 1:50:40is only spending a finite amount of
  2976. 1:50:41computation on it and so we kind of want
  2977. 1:50:45to distribute the computation across
  2978. 1:50:47many tokens and we can't have too much
  2979. 1:50:50computation or expect too much
  2980. 1:50:52computation out of of the model in any
  2981. 1:50:53single individual token because there's
  2982. 1:50:55only so much computation that happens
  2983. 1:50:57per token okay roughly fixed amount of
  2984. 1:51:00computation here
  2985. 1:51:02so that's why this answer here is
  2986. 1:51:06significantly worse and the reason for
  2987. 1:51:07that is Imagine going from left to right
  2988. 1:51:09here um and I copy pasted it right here
  2989. 1:51:13the answer is three Etc imagine the
  2990. 1:51:16model having to go from left to right
  2991. 1:51:17emitting these tokens one at a time it
  2992. 1:51:19has to say or we're expecting to say the
  2993. 1:51:23answer is space dollar sign and then
  2994. 1:51:27right here we're expecting it to
  2995. 1:51:28basically cram all of the computation of
  2996. 1:51:30this problem into this single token it
  2997. 1:51:32has to emit the correct answer three and
  2998. 1:51:35then once we've emitted the answer three
  2999. 1:51:37we're expecting it to say all these
  3000. 1:51:39tokens but at this point we've already
  3001. 1:51:41prod produced the answer and it's
  3002. 1:51:43already in the context window for all
  3003. 1:51:44these tokens that follow so anything
  3004. 1:51:46here is just um kind of post Hawk
  3005. 1:51:49justification of why this is the answer
  3006. 1:51:52um because the answer is already created
  3007. 1:51:53it's already in the token window so it's
  3008. 1:51:56it's not actually being calculated here
  3009. 1:51:58um and so if you are answering the
  3010. 1:52:01question directly and immediately you
  3011. 1:52:03are training the model to to try to
  3012. 1:52:06basically guess the answer in a single
  3013. 1:52:07token and that is just not going to work
  3014. 1:52:10because of the finite amount of
  3015. 1:52:11computation that happens per token
  3016. 1:52:13that's why this answer on the right is
  3017. 1:52:15significantly better because we are
  3018. 1:52:17Distributing this computation across the
  3019. 1:52:19answer we're actually getting the model
  3020. 1:52:20to sort of slowly come to the answer
  3021. 1:52:23from the left to right we're getting
  3022. 1:52:24intermediate results we're saying okay
  3023. 1:52:26the total cost of oranges is four so 30
  3024. 1:52:28- 4 is 9 and so we're creating
  3025. 1:52:32intermediate calculations and each one
  3026. 1:52:34of these calculations is by itself not
  3027. 1:52:36that expensive and so we're actually
  3028. 1:52:38basically kind of guessing a little bit
  3029. 1:52:40the difficulty that the model is capable
  3030. 1:52:42of in any single one of these individual
  3031. 1:52:44tokens and there can never be too much
  3032. 1:52:47work in any one of these tokens
  3033. 1:52:49computationally because then the model
  3034. 1:52:50won't be able to do that later at test
  3035. 1:52:52time and so we're teaching the model
  3036. 1:52:55here to spread out its reasoning and to
  3037. 1:52:57spread out its computation over the
  3038. 1:52:59tokens and in this way it only has very
  3039. 1:53:02simple problems in each token and they
  3040. 1:53:05can add up and then by the time it's
  3041. 1:53:07near the end it has all the previous
  3042. 1:53:09results in its working memory and it's
  3043. 1:53:11much easier for it to determine that the
  3044. 1:53:13answer is and here it is three so this
  3045. 1:53:15is a significantly better label for our
  3046. 1:53:18computation this would be really bad and
  3047. 1:53:20is teaching the model to try to do all
  3048. 1:53:23the computation in a single token and
  3049. 1:53:24it's really
  3050. 1:53:25bad so uh that's kind of like an
  3051. 1:53:28interesting thing to keep in mind is in
  3052. 1:53:30your
  3053. 1:53:31prompts uh usually don't have to think
  3054. 1:53:33about it explicitly because uh the
  3055. 1:53:36people at open AI have labelers and so
  3056. 1:53:38on that actually worry about this and
  3057. 1:53:40they make sure that the answers are
  3058. 1:53:41spread out and so actually open AI will
  3059. 1:53:43kind of like do the right thing so when
  3060. 1:53:45I ask this question for chat GPT it's
  3061. 1:53:48actually going to go very slowly it's
  3062. 1:53:49going to be like okay let's define our
  3063. 1:53:50variables set up the equation
  3064. 1:53:52and it's kind of creating all these
  3065. 1:53:54intermediate results these are not for
  3066. 1:53:56you these are for the model if the model
  3067. 1:53:58is not creating these intermediate
  3068. 1:53:59results for itself it's not going to be
  3069. 1:54:01able to reach three I also wanted to
  3070. 1:54:04show you that it's possible to be a bit
  3071. 1:54:06mean to the model uh we can just ask for
  3072. 1:54:08things so as an example I said I gave it
  3073. 1:54:10the exact same uh prompt and I said
  3074. 1:54:13answer the question in a single token
  3075. 1:54:15just immediately give me the answer
  3076. 1:54:16nothing else and it turns out that for
  3077. 1:54:18this simple um prompt here it actually
  3078. 1:54:21was able to do it in single go so it
  3079. 1:54:23just created a single I think this is
  3080. 1:54:25two tokens right uh because the dollar
  3081. 1:54:27sign is its own token so basically this
  3082. 1:54:30model didn't give me a single token it
  3083. 1:54:31gave me two tokens but it still produced
  3084. 1:54:33the correct answer and it did that in a
  3085. 1:54:35single forward pass of the
  3086. 1:54:37network now that's because the numbers
  3087. 1:54:40here I think are very simple and so I
  3088. 1:54:41made it a bit more difficult to be a bit
  3089. 1:54:43mean to the model so I said Emily buys
  3090. 1:54:4523 apples and 177 oranges and then I
  3091. 1:54:48just made the numbers a bit bigger and
  3092. 1:54:50I'm just making it harder for the model
  3093. 1:54:51I'm asking it to more computation in a
  3094. 1:54:53single token and so I said the same
  3095. 1:54:55thing and here it gave me five and five
  3096. 1:54:58is actually not correct so the model
  3097. 1:55:00failed to do all of this calculation in
  3098. 1:55:02a single forward pass of the network it
  3099. 1:55:04failed to go from the input tokens and
  3100. 1:55:07then in a single forward pass of the
  3101. 1:55:09network single go through the network it
  3102. 1:55:11couldn't produce the result and then I
  3103. 1:55:13said okay now don't worry about the the
  3104. 1:55:16token limit and just solve the problem
  3105. 1:55:18as usual and then it goes all the
  3106. 1:55:20intermediate results it simplifies and
  3107. 1:55:22every one of these intermediate results
  3108. 1:55:24here and intermediate calculations is
  3109. 1:55:26much easier for the model and um it sort
  3110. 1:55:29of it's not too much work per token all
  3111. 1:55:32of the tokens here are correct and it
  3112. 1:55:33arises the solution which is seven and I
  3113. 1:55:36just couldn't squeeze all of this work
  3114. 1:55:38it couldn't squeeze that into a single
  3115. 1:55:39forward passive Network so I think
  3116. 1:55:41that's kind of just a cute example and
  3117. 1:55:43something to kind of like think about
  3118. 1:55:45and I think it's kind of again just
  3119. 1:55:46elucidative in terms of how these uh
  3120. 1:55:48models work the last thing that I would
  3121. 1:55:50say on this topic is that if I was in
  3122. 1:55:52practi is trying to actually solve this
  3123. 1:55:53in my day-to-day life I might actually
  3124. 1:55:55not uh trust that the model that all the
  3125. 1:55:57intermediate calculations correctly here
  3126. 1:55:59so actually probably what I do is
  3127. 1:56:01something like this I would come here
  3128. 1:56:02and I would say use code and uh that's
  3129. 1:56:06because code is one of the possible
  3130. 1:56:08tools that chachy PD can use and instead
  3131. 1:56:11of it having to do mental arithmetic
  3132. 1:56:14like this mental arithmetic here I don't
  3133. 1:56:15fully trust it and especially if the
  3134. 1:56:17numbers get really big there's no
  3135. 1:56:19guarantee that the model will do this
  3136. 1:56:20correctly any one of these intermediates
  3137. 1:56:22steps might in principle fail we're
  3138. 1:56:24using neural networks to do mental
  3139. 1:56:26arithmetic uh kind of like you doing
  3140. 1:56:27mental arithmetic in your brain it might
  3141. 1:56:30just like uh screw up some of the
  3142. 1:56:31intermediate results it's actually kind
  3143. 1:56:32of amazing that it can even do this kind
  3144. 1:56:34of mental arithmetic I don't think I
  3145. 1:56:35could do this in my head but basically
  3146. 1:56:37the model is kind of like doing it in
  3147. 1:56:38its head and I don't trust that so I
  3148. 1:56:40wanted to use tools so you can say stuff
  3149. 1:56:42like use
  3150. 1:56:43code and uh I'm not sure what happened
  3151. 1:56:47there use
  3152. 1:56:50code and so um like I mentioned there's
  3153. 1:56:53a special tool and the uh the model can
  3154. 1:56:55write code and I can inspect that this
  3155. 1:56:58code is correct and then uh it's not
  3156. 1:57:01relying on its mental arithmetic it is
  3157. 1:57:03using the python interpreter which is a
  3158. 1:57:05very simple programming language to
  3159. 1:57:07basically uh write out the code that
  3160. 1:57:08calculates the result and I would
  3161. 1:57:10personally trust this a lot more because
  3162. 1:57:12this came out of a Python program which
  3163. 1:57:14I think has a lot more correctness
  3164. 1:57:15guarantees than the mental arithmetic of
  3165. 1:57:17a language model uh so just um another
  3166. 1:57:21kind of uh potential hint that if you
  3167. 1:57:23have these kinds of problems uh you may
  3168. 1:57:24want to basically just uh ask the model
  3169. 1:57:26to use the code interpreter and just
  3170. 1:57:28like we saw with the web search the
  3171. 1:57:30model has special uh kind of tokens for
  3172. 1:57:34calling uh like it will not actually
  3173. 1:57:36generate these tokens from the language
  3174. 1:57:38model it will write the program and then
  3175. 1:57:40it actually sends that program to a
  3176. 1:57:42different sort of part of the computer
  3177. 1:57:44that actually just runs that program and
  3178. 1:57:46brings back the result and then the
  3179. 1:57:48model gets access to that result and can
  3180. 1:57:50tell you that okay the cost of each
  3181. 1:57:51apple is seven
  3182. 1:57:53um so that's another kind of tool and I
  3183. 1:57:55would use this in practice for yourself
  3184. 1:57:57and it's um yeah it's just uh less error
  3185. 1:58:01prone I would say so that's why I called
  3186. 1:58:03this section models need tokens to think
  3187. 1:58:06distribute your competition across many
  3188. 1:58:08tokens ask models to create intermediate
  3189. 1:58:10results or whenever you can lean on
  3190. 1:58:13tools and Tool use instead of allowing
  3191. 1:58:15the models to do all of the stuff in
  3192. 1:58:17their memory so if they try to do it all
  3193. 1:58:18in their memory I don't fully trust it
  3194. 1:58:21and prefer to use tools whenever
  3195. 1:58:22possible I want to show you one more
  3196. 1:58:24example of where this actually comes up
  3197. 1:58:26and that's in counting so models
  3198. 1:58:28actually are not very good at counting
  3199. 1:58:30for the exact same reason you're asking
  3200. 1:58:32for way too much in a single individual
  3201. 1:58:34token so let me show you a simple
  3202. 1:58:36example of that um how many dots are
  3203. 1:58:38below and then I just put in a bunch of
  3204. 1:58:41dots and Chach says there are and then
  3205. 1:58:44it just tries to solve the problem in a
  3206. 1:58:46single token so in a single token it has
  3207. 1:58:49to count the number of dots in its
  3208. 1:58:51context window
  3209. 1:58:53um and it has to do that in the single
  3210. 1:58:55forward pass of a network and a single
  3211. 1:58:57forward pass of a network as we talked
  3212. 1:58:58about there's not that much computation
  3213. 1:59:00that can happen there just think of that
  3214. 1:59:01as being like very little competation
  3215. 1:59:03that happens there so if I just look at
  3216. 1:59:06what the model sees let's go to the LM
  3217. 1:59:09go to tokenizer it sees uh
  3218. 1:59:13this how many dots are below and then it
  3219. 1:59:15turns out that these dots here this
  3220. 1:59:17group of I think 20 dots is a single
  3221. 1:59:20token and then this group of whatever it
  3222. 1:59:22is is another token and then for some
  3223. 1:59:25reason they break up as this so I don't
  3224. 1:59:28actually this has to do with the details
  3225. 1:59:29of the tokenizer but it turns out that
  3226. 1:59:31these um the model basically sees the
  3227. 1:59:34token ID this this this and so on and
  3228. 1:59:38then from these token IDs it's expected
  3229. 1:59:40to count the number and spoiler alert is
  3230. 1:59:43not 161 it's actually I believe
  3231. 1:59:45177 so here's what we can do instead uh
  3232. 1:59:48we can say use code and you might expect
  3233. 1:59:51that like why should this work and it's
  3234. 1:59:54actually kind of subtle and kind of
  3235. 1:59:55interesting so when I say use code I
  3236. 1:59:57actually expect this to work let's see
  3237. 1:59:59okay 177 is correct so what happens here
  3238. 2:00:02is I've actually it doesn't look like it
  3239. 2:00:04but I've broken down the problem into a
  3240. 2:00:08problems that are easier for the model I
  3241. 2:00:10know that the model can't count it can't
  3242. 2:00:12do mental counting but I know that the
  3243. 2:00:14model is actually pretty good at doing
  3244. 2:00:15copy pasting so what I'm doing here is
  3245. 2:00:18when I say use code it creates a string
  3246. 2:00:20in Python for this and the task of
  3247. 2:00:23basically copy pasting my input here to
  3248. 2:00:27here is very simple because for the
  3249. 2:00:29model um it sees this string of uh it
  3250. 2:00:33sees it as just these four tokens or
  3251. 2:00:35whatever it is so it's very simple for
  3252. 2:00:37the model to copy paste those token IDs
  3253. 2:00:40and um kind of unpack them into Dots
  3254. 2:00:45here and so it creates this string and
  3255. 2:00:47then it calls python routine. count and
  3256. 2:00:50then it comes up with the correct answer
  3257. 2:00:52so the python interpreter is doing the
  3258. 2:00:53counting it's not the models mental
  3259. 2:00:55arithmetic doing the counting so it's
  3260. 2:00:57again a simple example of um models need
  3261. 2:01:00tokens to think don't rely on their
  3262. 2:01:02mental arithmetic and um that's why also
  3263. 2:01:05the models are not very good at counting
  3264. 2:01:07if you need them to do counting tasks
  3265. 2:01:08always ask them to lean on the tool now
  3266. 2:01:11the models also have many other little
  3267. 2:01:13cognitive deficits here and there and
  3268. 2:01:15these are kind of like sharp edges of
  3269. 2:01:16the technology to be kind of aware of
  3270. 2:01:18over time so as an example the models
  3271. 2:01:20are not very good with all kinds of
  3272. 2:01:22spelling related tasks they're not very
  3273. 2:01:24good at it and I told you that we would
  3274. 2:01:26loop back around to tokenization and the
  3275. 2:01:29reason to do for this is that the models
  3276. 2:01:31they don't see the characters they see
  3277. 2:01:33tokens and they their entire world is
  3278. 2:01:35about tokens which are these little text
  3279. 2:01:37chunks and so they don't see characters
  3280. 2:01:39like our eyes do and so very simple
  3281. 2:01:41character level tasks often fail so for
  3282. 2:01:45example uh I'm giving it a string
  3283. 2:01:47ubiquitous and I'm asking it to print
  3284. 2:01:49only every third character starting with
  3285. 2:01:51the first one so we start with U and
  3286. 2:01:54then we should go every third so every
  3287. 2:01:56so 1 2 3 Q should be next and then Etc
  3288. 2:02:01so this I see is not correct and again
  3289. 2:02:03my hypothesis is that this is again
  3290. 2:02:05Dental arithmetic here is failing number
  3291. 2:02:08one a little bit but number two I think
  3292. 2:02:10the the more important issue here is
  3293. 2:02:12that if you go to Tik
  3294. 2:02:13tokenizer and you look at ubiquitous we
  3295. 2:02:16see that it is three tokens right so you
  3296. 2:02:19and I see ubiquitous and we can easily
  3297. 2:02:21access the individual letters because we
  3298. 2:02:23kind of see them and when we have it in
  3299. 2:02:25the working memory of our visual sort of
  3300. 2:02:27field we can really easily index into
  3301. 2:02:29every third letter and I can do that
  3302. 2:02:31task but the models don't have access to
  3303. 2:02:33the individual letters they see this as
  3304. 2:02:35these three tokens and uh remember these
  3305. 2:02:38models are trained from scratch on the
  3306. 2:02:39internet and all these token uh
  3307. 2:02:42basically the model has to discover how
  3308. 2:02:44many of all these different letters are
  3309. 2:02:45packed into all these different tokens
  3310. 2:02:47and the reason we even use tokens is
  3311. 2:02:49mostly for efficiency uh but I think a
  3312. 2:02:51lot of people areed interested to delete
  3313. 2:02:52tokens entirely like we should really
  3314. 2:02:54have character level or bite level
  3315. 2:02:56models it's just that that would create
  3316. 2:02:58very long sequences and people don't
  3317. 2:02:59know how to deal with that right now so
  3318. 2:03:01while we have the token World any kind
  3319. 2:03:03of spelling tasks are not actually
  3320. 2:03:05expected to work super well so because I
  3321. 2:03:07know that spelling is not a strong suit
  3322. 2:03:09because of tokenization I can again Ask
  3323. 2:03:11it to lean On Tools so I can just say
  3324. 2:03:13use code and I would again expect this
  3325. 2:03:16to work because the task of copy pasting
  3326. 2:03:18ubiquitous into the python interpreter
  3327. 2:03:20is much easier and then we're leaning on
  3328. 2:03:22python interpreter to manipulate the
  3329. 2:03:25characters of this string so when I say
  3330. 2:03:27use
  3331. 2:03:28code
  3332. 2:03:30ubiquitous yes it indexes into every
  3333. 2:03:32third character and the actual truth is
  3334. 2:03:35u2s
  3335. 2:03:36uqs uh which looks correct to me so um
  3336. 2:03:41again an example of spelling related
  3337. 2:03:42tasks not working very well a very
  3338. 2:03:44famous example of that recently is how
  3339. 2:03:47many R are there in strawberry and this
  3340. 2:03:49went viral many times and basically the
  3341. 2:03:51models now get it correct they say there
  3342. 2:03:53are three Rs in Strawberry but for a
  3343. 2:03:55very long time all the state-of-the-art
  3344. 2:03:56models would insist that there are only
  3345. 2:03:58two RS in strawberry and this caused a
  3346. 2:04:00lot of you know Ruckus because is that a
  3347. 2:04:03word I think so because um it just kind
  3348. 2:04:06of like why are the models so brilliant
  3349. 2:04:08and they can solve math Olympiad
  3350. 2:04:10questions but they can't like count RS
  3351. 2:04:12in strawberry and the answer for that
  3352. 2:04:14again is I've got built up to it kind of
  3353. 2:04:16slowly but number one the models don't
  3354. 2:04:18see characters they see tokens and
  3355. 2:04:20number two they are not very good at
  3356. 2:04:22counting and so here we are combining
  3357. 2:04:25the difficulty of seeing the characters
  3358. 2:04:27with the difficulty of counting and
  3359. 2:04:29that's why the models struggled with
  3360. 2:04:30this even though I think by now honestly
  3361. 2:04:33I think open I may have hardcoded the
  3362. 2:04:34answer here or I'm not sure what they
  3363. 2:04:35did but um uh but this specific query
  3364. 2:04:39now works
  3365. 2:04:41so models are not very good at spelling
  3366. 2:04:44and there there's a bunch of other
  3367. 2:04:45little sharp edges and I don't want to
  3368. 2:04:46go into all of them I just want to show
  3369. 2:04:48you a few examples of things to be aware
  3370. 2:04:50of and uh when you're using these models
  3371. 2:04:52in practice I don't actually want to
  3372. 2:04:54have a comprehensive analysis here of
  3373. 2:04:55all the ways that the models are kind of
  3374. 2:04:57like falling short I just want to make
  3375. 2:04:59the point that there are some Jagged
  3376. 2:05:01edges here and there and we've discussed
  3377. 2:05:03a few of them and a few of them make
  3378. 2:05:05sense but some of them also will just
  3379. 2:05:06not make as much sense and they're kind
  3380. 2:05:08of like you're left scratching your head
  3381. 2:05:10even if you understand in- depth how
  3382. 2:05:11these models work and and good example
  3383. 2:05:14of that recently is the following uh the
  3384. 2:05:16models are not very good at very simple
  3385. 2:05:17questions like this and uh this is
  3386. 2:05:20shocking to a lot of people because
  3387. 2:05:22these math uh these problems can solve
  3388. 2:05:23complex math problems they can answer
  3389. 2:05:25PhD grade physics chemistry biology
  3390. 2:05:28questions much better than I can but
  3391. 2:05:30sometimes they fall short in like super
  3392. 2:05:31simple problems like this so here we go
  3393. 2:05:349.11 is bigger than 9.9 and it justifies
  3394. 2:05:38it in some way but obviously and then at
  3395. 2:05:40the end okay it actually it flips its
  3396. 2:05:44decision later so um I don't believe
  3397. 2:05:47that this is very reproducible sometimes
  3398. 2:05:49it flips around its answer sometimes
  3399. 2:05:50gets it right sometimes get it get it
  3400. 2:05:52wrong uh let's try
  3401. 2:05:56again okay even though it might look
  3402. 2:05:59larger okay so here it doesn't even
  3403. 2:06:01correct itself in the end if you ask
  3404. 2:06:03many times sometimes it gets it right
  3405. 2:06:04too but how is it that the model can do
  3406. 2:06:07so great at Olympiad grade problems but
  3407. 2:06:10then fail on very simple problems like
  3408. 2:06:12this and uh I think this one is as I
  3409. 2:06:15mentioned a little bit of a head
  3410. 2:06:16scratcher it turns out that a bunch of
  3411. 2:06:18people studied this in depth and I
  3412. 2:06:19haven't actually read the paper uh but
  3413. 2:06:22what I was told by this team was that
  3414. 2:06:24when you scrutinize the activations
  3415. 2:06:27inside the neural network when you look
  3416. 2:06:29at some of the features and what what
  3417. 2:06:31features turn on or off and what neurons
  3418. 2:06:33turn on or off uh a bunch of neurons
  3419. 2:06:35inside the neural network light up that
  3420. 2:06:37are usually associated with Bible verses
  3421. 2:06:40U and so I think the model is kind of
  3422. 2:06:42like reminded that these almost look
  3423. 2:06:44like Bible verse markers and in a bip
  3424. 2:06:48verse setting 9.11 would come after 99.9
  3425. 2:06:52and so basically the model somehow finds
  3426. 2:06:53it like cognitively very distracting
  3427. 2:06:56that in Bible verses 9.11 would be
  3428. 2:06:58greater um even though here it's
  3429. 2:07:00actually trying to justify it and come
  3430. 2:07:02up to the answer with a math it still
  3431. 2:07:04ends up with the wrong answer here so it
  3432. 2:07:07basically just doesn't fully make sense
  3433. 2:07:08and it's not fully understood and um
  3434. 2:07:12there's a few Jagged issues like that so
  3435. 2:07:14that's why treat this as a as what it is
  3436. 2:07:17which is a St stochastic system that is
  3437. 2:07:19really magical but that you can't also
  3438. 2:07:21fully trust and you want to use it as a
  3439. 2:07:23tool not as something that you kind of
  3440. 2:07:25like letter rip on a problem and
  3441. 2:07:27copypaste the results okay so we have
  3442. 2:07:29now covered two major stages of training
  3443. 2:07:32of large language models we saw that in
  3444. 2:07:34the first stage this is called the
  3445. 2:07:36pre-training stage we are basically
  3446. 2:07:38training on internet documents and when
  3447. 2:07:40you train a language model on internet
  3448. 2:07:42documents you get what's called a base
  3449. 2:07:44model and it's basically an internet
  3450. 2:07:45document simulator right now we saw that
  3451. 2:07:48this is an interesting artifact and uh
  3452. 2:07:51this takes many months to train on
  3453. 2:07:53thousands of computers and it's kind of
  3454. 2:07:54a lossy compression of the internet and
  3455. 2:07:57it's extremely interesting but it's not
  3456. 2:07:58directly useful because we don't want to
  3457. 2:08:00sample internet documents we want to ask
  3458. 2:08:02questions of an AI and have it respond
  3459. 2:08:05to our questions so for that we need an
  3460. 2:08:07assistant and we saw that we can
  3461. 2:08:09actually construct an assistant in the
  3462. 2:08:11process of a post
  3463. 2:08:13training and specifically in the process
  3464. 2:08:16of supervised fine-tuning as we call
  3465. 2:08:19it so in this stage we saw that it's
  3466. 2:08:22algorithmically identical to
  3467. 2:08:24pre-training nothing is going to change
  3468. 2:08:25the only thing that changes is the data
  3469. 2:08:27set so instead of Internet documents we
  3470. 2:08:30now want to create and curate a very
  3471. 2:08:32nice data set of conversations so we
  3472. 2:08:35want Millions conversations on all kinds
  3473. 2:08:38of diverse topics between a human and an
  3474. 2:08:41assistant and fundamentally these
  3475. 2:08:44conversations are created by humans so
  3476. 2:08:47humans write the prompts and humans
  3477. 2:08:49write the ideal response responses and
  3478. 2:08:52they do that based on labeling
  3479. 2:08:54documentations now in the modern stack
  3480. 2:08:57it's not actually done fully and
  3481. 2:08:59manually by humans right they actually
  3482. 2:09:00now have a lot of help from these tools
  3483. 2:09:02so we can use language models um to help
  3484. 2:09:05us create these data sets and that's
  3485. 2:09:07done extensively but fundamentally it's
  3486. 2:09:09all still coming from Human curation at
  3487. 2:09:10the end so we create these conversations
  3488. 2:09:13that now becomes our data set we fine
  3489. 2:09:15tune on it or continue training on it
  3490. 2:09:17and we get an assistant and then we kind
  3491. 2:09:20of shifted gears and started talking
  3492. 2:09:21about some of the kind of cognitive
  3493. 2:09:22implications of what this assistant is
  3494. 2:09:24like and we saw that for example the
  3495. 2:09:26assistant will hallucinate if you don't
  3496. 2:09:29take some sort of mitigations towards it
  3497. 2:09:32so we saw that hallucinations would be
  3498. 2:09:34common and then we looked at some of the
  3499. 2:09:35mitigations of those hallucinations and
  3500. 2:09:38then we saw that the models are quite
  3501. 2:09:39impressive and can do a lot of stuff in
  3502. 2:09:40their head but we saw that they can also
  3503. 2:09:43Lean On Tools to become better so for
  3504. 2:09:45example we can lo lean on a web search
  3505. 2:09:48in order to hallucinate less and to
  3506. 2:09:50maybe bring up some more um recent
  3507. 2:09:53information or something like that or we
  3508. 2:09:54can lean on tools like code interpreter
  3509. 2:09:57so the code can so the llm can write
  3510. 2:09:59some code and actually run it and see
  3511. 2:10:00the
  3512. 2:10:01results so these are some of the topics
  3513. 2:10:03we looked at so far um now what I'd like
  3514. 2:10:06to do is I'd like to cover the last and
  3515. 2:10:09major stage of this Pipeline and that is
  3516. 2:10:12reinforcement learning so reinforcement
  3517. 2:10:15learning is still kind of thought to be
  3518. 2:10:16under the umbrella of posttraining uh
  3519. 2:10:19but it is the last third major stage and
  3520. 2:10:22it's a different way of training
  3521. 2:10:24language models and usually follows as
  3522. 2:10:26this third step so inside companies like
  3523. 2:10:29open AI you will start here and these
  3524. 2:10:31are all separate teams so there's a team
  3525. 2:10:33doing data for pre-training and a team
  3526. 2:10:35doing training for pre-training and then
  3527. 2:10:37there's a team doing all the
  3528. 2:10:39conversation generation in a in a
  3529. 2:10:42different team that is kind of doing the
  3530. 2:10:44supervis fine tuning and there will be a
  3531. 2:10:45team for the reinforcement learning as
  3532. 2:10:47well so it's kind of like a handoff of
  3533. 2:10:49these models you get your base model the
  3534. 2:10:51then you find you need to be an
  3535. 2:10:52assistant and then you go into
  3536. 2:10:53reinforcement learning which we'll talk
  3537. 2:10:55about uh
  3538. 2:10:56now so that's kind of like the major
  3539. 2:10:58flow and so let's now focus on
  3540. 2:11:01reinforcement learning the last major
  3541. 2:11:03stage of training and let me first
  3542. 2:11:05actually motivate it and why we would
  3543. 2:11:07want to do reinforcement learning and
  3544. 2:11:09what it looks like on a high level so I
  3545. 2:11:11would now like to try to motivate the
  3546. 2:11:12reinforcement learning stage and what it
  3547. 2:11:13corresponds to with something that
  3548. 2:11:15you're probably familiar with and that
  3549. 2:11:16is basically going to school so just
  3550. 2:11:19like you went to school to become um
  3551. 2:11:21really good at something we want to take
  3552. 2:11:23large language models through school and
  3553. 2:11:25really what we're doing is um we're um
  3554. 2:11:29we have a few paradigms of ways of uh
  3555. 2:11:32giving them knowledge or transferring
  3556. 2:11:33skills so in particular when we're
  3557. 2:11:36working with textbooks in school you'll
  3558. 2:11:38see that there are three major kind of
  3559. 2:11:40uh pieces of information in these
  3560. 2:11:42textbooks three classes of information
  3561. 2:11:45the first thing you'll see is you'll see
  3562. 2:11:46a lot of exposition um and by the way
  3563. 2:11:49this is a totally random book I pulled
  3564. 2:11:50from the internet I I think it's some
  3565. 2:11:51kind of an organic chemistry or
  3566. 2:11:53something I'm not sure uh but the
  3567. 2:11:55important thing is that you'll see that
  3568. 2:11:56most of the text most of it is kind of
  3569. 2:11:58just like the meat of it is exposition
  3570. 2:12:00it's kind of like background knowledge
  3571. 2:12:02Etc as you are reading through the words
  3572. 2:12:05of this Exposition you can think of that
  3573. 2:12:08roughly as training on that data so um
  3574. 2:12:12and that's why when you're reading
  3575. 2:12:13through this stuff this background
  3576. 2:12:14knowledge and this all this context
  3577. 2:12:16information it's kind of equivalent to
  3578. 2:12:18pre-training so it's it's where we build
  3579. 2:12:21sort of like a knowledge base of this
  3580. 2:12:23data and get a sense of the topic the
  3581. 2:12:27next major kind of information that you
  3582. 2:12:28will see is these uh problems and with
  3583. 2:12:32their worked Solutions so basically a
  3584. 2:12:35human expert in this case uh the author
  3585. 2:12:37of this book has given us not just a
  3586. 2:12:39problem but has also worked through the
  3587. 2:12:41solution and the solution is basically
  3588. 2:12:43like equivalent to having like this
  3589. 2:12:45ideal response for an assistant so it's
  3590. 2:12:48basically the expert is showing us how
  3591. 2:12:49to solve the problem in it's uh kind of
  3592. 2:12:52like um in its full form so as we are
  3593. 2:12:55reading the solution we are basically
  3594. 2:12:57training on the expert data and then
  3595. 2:13:01later we can try to imitate the expert
  3596. 2:13:03um and basically um that's that roughly
  3597. 2:13:07correspond to having the sft model
  3598. 2:13:08that's what it would be doing so
  3599. 2:13:11basically we've already done
  3600. 2:13:12pre-training and we've already covered
  3601. 2:13:14this um imitation of experts and how
  3602. 2:13:17they solve these problems and the third
  3603. 2:13:19stage of reinforcement learning is
  3604. 2:13:21basically the practice problems so
  3605. 2:13:24sometimes you'll see this is just a
  3606. 2:13:25single practice problem here but of
  3607. 2:13:27course there will be usually many
  3608. 2:13:28practice problems at the end of each
  3609. 2:13:30chapter in any textbook and practice
  3610. 2:13:32problems of course we know are critical
  3611. 2:13:34for learning because what are they
  3612. 2:13:36getting you to do they're getting you to
  3613. 2:13:37practice uh to practice yourself and
  3614. 2:13:39discover ways of solving these problems
  3615. 2:13:42yourself and so what you get in a
  3616. 2:13:44practice problem is you get a problem
  3617. 2:13:46description but you're not given the
  3618. 2:13:48solution but you are given the final
  3619. 2:13:50answer answer usually in the answer key
  3620. 2:13:53of the textbook and so you know the
  3621. 2:13:55final answer that you're trying to get
  3622. 2:13:56to and you have the problem statement
  3623. 2:13:58but you don't have the solution you are
  3624. 2:14:00trying to practice the solution you're
  3625. 2:14:02trying out many different things and
  3626. 2:14:04you're seeing what gets you to the final
  3627. 2:14:07solution the best and so you're
  3628. 2:14:09discovering how to solve these problems
  3629. 2:14:11so and in the process of that you're
  3630. 2:14:13relying on number one the background
  3631. 2:14:14information which comes from
  3632. 2:14:15pre-training and number two maybe a
  3633. 2:14:17little bit of imitation of human experts
  3634. 2:14:20and you can probably try similar kinds
  3635. 2:14:22of solutions and so on so we've done
  3636. 2:14:25this and this and now in this section
  3637. 2:14:27we're going to try to practice and so
  3638. 2:14:30we're going to be given prompts we're
  3639. 2:14:32going to be given Solutions U sorry the
  3640. 2:14:34final answers but we're not going to be
  3641. 2:14:36given expert Solutions we have to
  3642. 2:14:38practice and try stuff out and that's
  3643. 2:14:40what reinforcement learning is about
  3644. 2:14:43okay so let's go back to the problem
  3645. 2:14:44that we worked with previously just so
  3646. 2:14:46we have a concrete example to talk
  3647. 2:14:47through as we explore sort of the topic
  3648. 2:14:50here so um I'm here in the Teck
  3649. 2:14:52tokenizer because I'd also like to well
  3650. 2:14:55I get a text box which is useful but
  3651. 2:14:57number two I want to remind you again
  3652. 2:14:59that we're always working with
  3653. 2:14:59onedimensional token sequences and so um
  3654. 2:15:02I actually like prefer this view because
  3655. 2:15:04this is like the native view of the llm
  3656. 2:15:06if that makes sense like this is what it
  3657. 2:15:08actually sees it sees token IDs right
  3658. 2:15:11okay so Emily buys three apples and two
  3659. 2:15:14oranges each orange is $2 the total cost
  3660. 2:15:17of all the fruit is $13 what is the cost
  3661. 2:15:19of each apple
  3662. 2:15:21and what I'd like to what I like you to
  3663. 2:15:23appreciate here is these are like four
  3664. 2:15:26possible candidate Solutions as an
  3665. 2:15:29example and they all reach the answer
  3666. 2:15:31three now what I'd like you to
  3667. 2:15:33appreciate at this point is that if I am
  3668. 2:15:35the human data labeler that is creating
  3669. 2:15:37a conversation to be entered into the
  3670. 2:15:39training set I don't actually really
  3671. 2:15:42know which of these
  3672. 2:15:44conversations to um to add to the data
  3673. 2:15:48set some of these conversations kind of
  3674. 2:15:50set up a system equations some of them
  3675. 2:15:52sort of like just talk through it in
  3676. 2:15:54English and some of them just kind of
  3677. 2:15:55like skip right through to the
  3678. 2:15:58solution um if you look at chbt for
  3679. 2:16:00example and you give it this question it
  3680. 2:16:03defines a system of variables and it
  3681. 2:16:05kind of like does this little thing what
  3682. 2:16:07we have to appreciate and uh
  3683. 2:16:08differentiate between though is um the
  3684. 2:16:12first purpose of a solution is to reach
  3685. 2:16:14the right answer of course we want to
  3686. 2:16:15get the final answer three that is the
  3687. 2:16:17that is the important purpose here but
  3688. 2:16:19there's kind of like a secondary purpose
  3689. 2:16:21as well where here we are also just kind
  3690. 2:16:23of trying to make it like nice uh for
  3691. 2:16:26the human because we're kind of assuming
  3692. 2:16:27that the person wants to see the
  3693. 2:16:29solution they want to see the
  3694. 2:16:30intermediate steps we want to present it
  3695. 2:16:31nicely Etc so there are two separate
  3696. 2:16:33things going on here number one is the
  3697. 2:16:36presentation for the human but number
  3698. 2:16:37two we're trying to actually get the
  3699. 2:16:38right answer um so let's for the moment
  3700. 2:16:42focus on just reaching the final answer
  3701. 2:16:44if we're only care if we only care about
  3702. 2:16:46the final answer then which of these is
  3703. 2:16:49the optimal or the best prompt um sorry
  3704. 2:16:53the best solution for the llm to reach
  3705. 2:16:56the right
  3706. 2:16:57answer um and what I'm trying to get at
  3707. 2:17:00is we don't know me as a human labeler I
  3708. 2:17:03would not know which one of these is
  3709. 2:17:04best so as an example we saw earlier on
  3710. 2:17:07when we looked at
  3711. 2:17:09um the token sequences here and the
  3712. 2:17:11mental arithmetic and reasoning we saw
  3713. 2:17:14that for each token we can only spend
  3714. 2:17:15basically a finite number of finite
  3715. 2:17:18amount of compute here that is not very
  3716. 2:17:19large or you should think about it that
  3717. 2:17:20way way and so we can't actually make
  3718. 2:17:23too big of a leap in any one token is is
  3719. 2:17:26maybe the way to think about it so as an
  3720. 2:17:28example in this one what's really nice
  3721. 2:17:30about it is that it's very few tokens so
  3722. 2:17:32it's going to take us very short amount
  3723. 2:17:34of time to get to the answer but right
  3724. 2:17:37here when we're doing 30 - 4 IDE 3
  3725. 2:17:39equals right in this token here we're
  3726. 2:17:42actually asking for a lot of computation
  3727. 2:17:44to happen on that single individual
  3728. 2:17:45token and so maybe this is a bad example
  3729. 2:17:48to give to the llm because it's kind of
  3730. 2:17:49incentivizing it to skip through the
  3731. 2:17:50calculations very quickly and it's going
  3732. 2:17:52to actually make up mistakes make
  3733. 2:17:54mistakes in this mental arithmetic uh so
  3734. 2:17:56maybe it would work better to like
  3735. 2:17:58spread out the spread it out more maybe
  3736. 2:18:01it would be better to set it up as an
  3737. 2:18:02equation maybe it would be better to
  3738. 2:18:04talk through it we fundamentally don't
  3739. 2:18:06know and we don't know because what is
  3740. 2:18:09easy for you or I as or as human
  3741. 2:18:12labelers what's easy for us or hard for
  3742. 2:18:14us is different than what's easy or hard
  3743. 2:18:16for the llm it cognition is different um
  3744. 2:18:20and the token sequences are kind of like
  3745. 2:18:23different hard for it and so some of the
  3746. 2:18:27token sequences here that are trivial
  3747. 2:18:30for me might be um very too much of a
  3748. 2:18:33leap for the llm so right here this
  3749. 2:18:36token would be way too hard but
  3750. 2:18:38conversely many of the tokens that I'm
  3751. 2:18:40creating here might be just trivial to
  3752. 2:18:43the llm and we're just wasting tokens
  3753. 2:18:45like why waste all these tokens when
  3754. 2:18:46this is all trivial so if the only thing
  3755. 2:18:49we care care about is the final answer
  3756. 2:18:51and we're separating out the issue of
  3757. 2:18:53the presentation to the human um then we
  3758. 2:18:56don't actually really know how to
  3759. 2:18:57annotate this example we don't know what
  3760. 2:18:59solution to get to the llm because we
  3761. 2:19:01are not the
  3762. 2:19:02llm and it's clear here in the case of
  3763. 2:19:05like the math example but this is
  3764. 2:19:07actually like a very pervasive issue
  3765. 2:19:08like for our knowledge is not lm's
  3766. 2:19:11knowledge like the llm actually has a
  3767. 2:19:13ton of knowledge of PhD in math and
  3768. 2:19:15physics chemistry and whatnot so in many
  3769. 2:19:17ways it actually knows more than I do
  3770. 2:19:19and I'm I'm potentially not utilizing
  3771. 2:19:21that knowledge in its problem solving
  3772. 2:19:24but conversely I might be injecting a
  3773. 2:19:26bunch of knowledge in my solutions that
  3774. 2:19:28the LM doesn't know in its parameters
  3775. 2:19:31and then those are like sudden leaps
  3776. 2:19:33that are very confusing to the model and
  3777. 2:19:36so our cognitions are different and I
  3778. 2:19:38don't really know what to put here if
  3779. 2:19:41all we care about is the reaching the
  3780. 2:19:42final solution and doing it economically
  3781. 2:19:45ideally and so long story short we are
  3782. 2:19:49not in a good position to create these
  3783. 2:19:52uh token sequences for the LM and
  3784. 2:19:55they're useful by imitation to
  3785. 2:19:56initialize the system but we really want
  3786. 2:19:59the llm to discover the token sequences
  3787. 2:20:01that work for it we need to find it
  3788. 2:20:04needs to find for itself what token
  3789. 2:20:06sequence reliably gets to the answer
  3790. 2:20:09given the prompt and it needs to
  3791. 2:20:11discover that in the process of
  3792. 2:20:12reinforcement learning and of trial and
  3793. 2:20:14error so let's see how this example
  3794. 2:20:18would work like in reinforcement
  3795. 2:20:19learning
  3796. 2:20:21okay so we're now back in the huging
  3797. 2:20:23face inference playground and uh that
  3798. 2:20:26just allows me to very easily call uh
  3799. 2:20:28different kinds of models so as an
  3800. 2:20:29example here on the top right I chose
  3801. 2:20:31the Gemma 2 2 billion parameter model so
  3802. 2:20:34two billion is very very small so this
  3803. 2:20:36is a tiny model but it's okay so we're
  3804. 2:20:39going to give it um the way that
  3805. 2:20:40reinforcement learning will basically
  3806. 2:20:41work is actually quite quite simple um
  3807. 2:20:44we need to try many different kinds of
  3808. 2:20:47solutions and we want to see which
  3809. 2:20:49Solutions work well or not
  3810. 2:20:51so we're basically going to take the
  3811. 2:20:53prompt we're going to run the
  3812. 2:20:55model and the model generates a solution
  3813. 2:20:58and then we're going to inspect the
  3814. 2:20:59solution and we know that the correct
  3815. 2:21:02answer for this one is $3 and so indeed
  3816. 2:21:05the model gets it correct it says it's
  3817. 2:21:06$3 so this is correct so that's just one
  3818. 2:21:10attempt at DIS solution so now we're
  3819. 2:21:11going to delete this and we're going to
  3820. 2:21:13rerun it again let's try a second
  3821. 2:21:15attempt so the model solves it in a bit
  3822. 2:21:17slightly different way right every
  3823. 2:21:19single attempt will be a different
  3824. 2:21:21generation because these models are
  3825. 2:21:23stochastic systems remember that at
  3826. 2:21:24every single token here we have a
  3827. 2:21:26probability distribution and we're
  3828. 2:21:27sampling from that distribution so we
  3829. 2:21:29end up kind kind of going down slightly
  3830. 2:21:31different paths and so this is a second
  3831. 2:21:34solution that also ends in the correct
  3832. 2:21:36answer now we're going to delete that
  3833. 2:21:38let's go a third
  3834. 2:21:39time okay so again slightly different
  3835. 2:21:42solution but also gets it
  3836. 2:21:44correct now we can actually repeat this
  3837. 2:21:46uh many times and so in practice you
  3838. 2:21:49might actually sample thousand of
  3839. 2:21:51independent Solutions or even like
  3840. 2:21:52million solutions for just a single
  3841. 2:21:55prompt um and some of them will be
  3842. 2:21:57correct and some of them will not be
  3843. 2:21:58very correct and basically what we want
  3844. 2:22:00to do is we want to encourage the
  3845. 2:22:02solutions that lead to correct answers
  3846. 2:22:05so let's take a look at what that looks
  3847. 2:22:06like so if we come back over here here's
  3848. 2:22:09kind of like a cartoon diagram of what
  3849. 2:22:10this is looking like we have a prompt
  3850. 2:22:13and then we tried many different
  3851. 2:22:15solutions in
  3852. 2:22:16parallel and some of the solutions um
  3853. 2:22:19might go well so they get the right
  3854. 2:22:21answer which is in green and some of the
  3855. 2:22:24solutions might go poorly and may not
  3856. 2:22:25reach the right answer which is red now
  3857. 2:22:28this problem here unfortunately is not
  3858. 2:22:29the best example because it's a trivial
  3859. 2:22:32prompt and as we saw uh even like a two
  3860. 2:22:34billion parameter model always gets it
  3861. 2:22:36right so it's not the best example in
  3862. 2:22:38that sense but let's just exercise some
  3863. 2:22:40imagination here and let's just suppose
  3864. 2:22:43that the um green ones are good and the
  3865. 2:22:47red ones are
  3866. 2:22:48bad okay so we generated 15 Solutions
  3867. 2:22:52only four of them got the right answer
  3868. 2:22:54and so now what we want to do is
  3869. 2:22:56basically we want to encourage the kinds
  3870. 2:22:58of solutions that lead to right answers
  3871. 2:23:00so whatever token sequences happened in
  3872. 2:23:03these red Solutions obviously something
  3873. 2:23:05went wrong along the way somewhere and
  3874. 2:23:07uh this was not a good path to take
  3875. 2:23:09through the solution and whatever token
  3876. 2:23:11sequences there were in these Green
  3877. 2:23:13Solutions well things went uh pretty
  3878. 2:23:15well in this situation and so we want to
  3879. 2:23:18do more things like it in prompts like
  3880. 2:23:21this and the way we encourage this kind
  3881. 2:23:23of a behavior in the future is we
  3882. 2:23:25basically train on these sequences um
  3883. 2:23:28but these training sequencies now are
  3884. 2:23:29not coming from expert human annotators
  3885. 2:23:32there's no human who decided that this
  3886. 2:23:33is the correct solution this solution
  3887. 2:23:36came from the model itself so the model
  3888. 2:23:38is practicing here it's tried out a few
  3889. 2:23:40Solutions four of them seem to have
  3890. 2:23:41worked and now the model will kind of
  3891. 2:23:43like train on them and this corresponds
  3892. 2:23:45to a student basically looking at their
  3893. 2:23:47Solutions and being like okay well this
  3894. 2:23:48one worked really well so this is this
  3895. 2:23:50is how I should be solving these kinds
  3896. 2:23:52of problems and uh here in this example
  3897. 2:23:55there are many different ways to
  3898. 2:23:57actually like really tweak the
  3899. 2:23:58methodology a little bit here but just
  3900. 2:24:00to give the core idea across maybe it's
  3901. 2:24:02simplest to just think about take the
  3902. 2:24:04taking the single best solution out of
  3903. 2:24:06these four uh like say this one that's
  3904. 2:24:08why it was yellow uh so this is the the
  3905. 2:24:12solution that not only led to the right
  3906. 2:24:13answer but may maybe had some other nice
  3907. 2:24:15properties maybe it was the shortest one
  3908. 2:24:17or it looked nicest in some ways or uh
  3909. 2:24:20there's other criteria you could think
  3910. 2:24:21of as an example but we're going to
  3911. 2:24:23decide that this the top solution we're
  3912. 2:24:25going to train on it and then uh the
  3913. 2:24:28model will be slightly more likely once
  3914. 2:24:30you do the parameter update to take this
  3915. 2:24:33path in this kind of a setting in the
  3916. 2:24:36future but you have to remember that
  3917. 2:24:38we're going to run many different
  3918. 2:24:39diverse prompts across lots of math
  3919. 2:24:42problems and physics problems and
  3920. 2:24:43whatever wherever there might be so tens
  3921. 2:24:46of thousands of prompts maybe have in
  3922. 2:24:47mind there's thousands of solutions
  3923. 2:24:50prompt and so this is all happening kind
  3924. 2:24:52of like at the same time and as we're
  3925. 2:24:55iterating this process the model is
  3926. 2:24:57discovering for itself what kinds of
  3927. 2:24:59token sequences lead it to correct
  3928. 2:25:02answers it's not coming from a human
  3929. 2:25:05annotator the the model is kind of like
  3930. 2:25:08playing in this playground and it knows
  3931. 2:25:10what it's trying to get to and it's
  3932. 2:25:12discovering sequences that work for it
  3933. 2:25:15uh these are sequences that don't make
  3934. 2:25:16any mental leaps uh they they seem to
  3935. 2:25:19work reliably and statistically and uh
  3936. 2:25:23fully utilize the knowledge of the model
  3937. 2:25:25as it has it and so uh this is the
  3938. 2:25:28process of reinforcement
  3939. 2:25:29learning it's basically a guess and
  3940. 2:25:31check we're going to guess many
  3941. 2:25:32different types of solutions we're going
  3942. 2:25:33to check them and we're going to do more
  3943. 2:25:35of what worked in the future and that is
  3944. 2:25:38uh reinforcement learning so in the
  3945. 2:25:40context of what came before we see now
  3946. 2:25:43that the sft model the supervised fine
  3947. 2:25:45tuning model it's still helpful because
  3948. 2:25:47it still kind of like initializes the
  3949. 2:25:49model a little bit into to the vicinity
  3950. 2:25:51of the correct Solutions so it's kind of
  3951. 2:25:53like a initialization of um of the model
  3952. 2:25:56in the sense that it kind of gets the
  3953. 2:25:58model to you know take Solutions like
  3954. 2:26:00write out Solutions and maybe it has an
  3955. 2:26:03understanding of setting up a system of
  3956. 2:26:04equations or maybe it kind of like talks
  3957. 2:26:06through a solution so it gets you into
  3958. 2:26:08the vicinity of correct Solutions but
  3959. 2:26:10reinforcement learning is where
  3960. 2:26:11everything gets dialed in we really
  3961. 2:26:13discover the solutions that work for the
  3962. 2:26:15model get the right answers we encourage
  3963. 2:26:17them and then the model just kind of
  3964. 2:26:19like gets better over time time okay so
  3965. 2:26:21that is the high Lev process for how we
  3966. 2:26:23train large language models in short we
  3967. 2:26:26train them kind of very similar to how
  3968. 2:26:27we train children and basically the only
  3969. 2:26:30difference is that children go through
  3970. 2:26:32chapters of books and they do all these
  3971. 2:26:34different types of training exercises um
  3972. 2:26:37kind of within the chapter of each book
  3973. 2:26:39but instead when we train AIS it's
  3974. 2:26:41almost like we kind of do it stage by
  3975. 2:26:43stage depending on the type of that
  3976. 2:26:45stage so first what we do is we do
  3977. 2:26:47pre-training which as we saw is
  3978. 2:26:49equivalent to uh basically reading all
  3979. 2:26:51the expository material so we look at
  3980. 2:26:53all the textbooks at the same time and
  3981. 2:26:55we read all the exposition and we try to
  3982. 2:26:57build a knowledge base the second thing
  3983. 2:27:00then is we go into the sft stage which
  3984. 2:27:02is really looking at all the fixed uh
  3985. 2:27:04sort of like solutions from Human
  3986. 2:27:07Experts of all the different kinds of
  3987. 2:27:09worked Solutions across all the
  3988. 2:27:11textbooks and we just kind of get an sft
  3989. 2:27:14model which is able to imitate the
  3990. 2:27:16experts but does so kind of blindly it
  3991. 2:27:18just kind of like does its best guess
  3992. 2:27:20uh kind of just like trying to mimic
  3993. 2:27:22statistically the expert behavior and so
  3994. 2:27:24that's what you get when you look at all
  3995. 2:27:26the work Solutions and then finally in
  3996. 2:27:28the last stage we do all the practice
  3997. 2:27:30problems in the RL stage across all the
  3998. 2:27:33textbooks we only do the practice
  3999. 2:27:35problems and that's how we get the RL
  4000. 2:27:37model so on a high level the way we
  4001. 2:27:40train llms is very much equivalent uh to
  4002. 2:27:43the process that we train uh that we use
  4003. 2:27:45for training of children the next point
  4004. 2:27:47I would like to make is that actually
  4005. 2:27:49these first two stat ages pre-training
  4006. 2:27:51and surprise fine-tuning they've been
  4007. 2:27:52around for years and they are very
  4008. 2:27:53standard and everyone does them all the
  4009. 2:27:55different llm providers it is this last
  4010. 2:27:58stage the RL training that is a lot more
  4011. 2:28:00early in its process of development and
  4012. 2:28:02is not standard yet in the field and so
  4013. 2:28:06um this stage is a lot more kind of
  4014. 2:28:09early and nent and the reason for that
  4015. 2:28:11is because I actually skipped over a ton
  4016. 2:28:13of little details here in this process
  4017. 2:28:15the high level idea is very simple it's
  4018. 2:28:17trial and there learning but there's a
  4019. 2:28:18ton of details and little math
  4020. 2:28:20mathematical kind of like nuances to
  4021. 2:28:21exactly how you pick the solutions that
  4022. 2:28:23are the best and how much you train on
  4023. 2:28:25them and what is the prompt distribution
  4024. 2:28:27and how to set up the training run such
  4025. 2:28:29that this actually works so there's a
  4026. 2:28:30lot of little details and knobs to the
  4027. 2:28:32core idea that is very very simple and
  4028. 2:28:35so getting the details right here uh is
  4029. 2:28:37not trivial and so a lot of companies
  4030. 2:28:40like for example open and other LM
  4031. 2:28:41providers have experimented internally
  4032. 2:28:44with reinforcement learning fine tuning
  4033. 2:28:46for llms for a while but they've not
  4034. 2:28:48talked about it publicly
  4035. 2:28:50um it's all kind of done inside the
  4036. 2:28:52company and so that's why the paper from
  4037. 2:28:55Deep seek that came out very very
  4038. 2:28:56recently was such a big deal because
  4039. 2:28:59this is a paper from this company called
  4040. 2:29:01DC Kai in China and this paper really
  4041. 2:29:05talked very publicly about reinforcement
  4042. 2:29:07learning fine training for large
  4043. 2:29:08language models and how incredibly
  4044. 2:29:10important it is for large language
  4045. 2:29:12models and how it brings out a lot of
  4046. 2:29:14reasoning capabilities in the models
  4047. 2:29:16we'll go into this in a second so this
  4048. 2:29:18paper reinvigorated the public interest
  4049. 2:29:21of using RL for llms and gave a lot of
  4050. 2:29:25the um sort of n-r details that are
  4051. 2:29:27needed to reproduce their results and
  4052. 2:29:29actually get the stage to work for large
  4053. 2:29:31langage models so let me take you
  4054. 2:29:33briefly through this uh deep seek R1
  4055. 2:29:35paper and what happens when you actually
  4056. 2:29:36correctly apply RL to language models
  4057. 2:29:38and what that looks like and what that
  4058. 2:29:39gives you so the first thing I'll scroll
  4059. 2:29:41to is this uh kind of figure two here
  4060. 2:29:43where we are looking at the Improvement
  4061. 2:29:45in how the models are solving
  4062. 2:29:47mathematical problems so this is the
  4063. 2:29:49accuracy of solving mathematical
  4064. 2:29:50problems on the a accuracy and then we
  4065. 2:29:54can go to the web page and we can see
  4066. 2:29:55the kinds of problems that are actually
  4067. 2:29:56in these um these the kinds of math
  4068. 2:29:58problems that are being measured here so
  4069. 2:30:00these are simple math problems you can
  4070. 2:30:02um pause the video if you like but these
  4071. 2:30:04are the kinds of problems that basically
  4072. 2:30:06the models are being asked to solve and
  4073. 2:30:08you can see that in the beginning
  4074. 2:30:09they're not doing very well but then as
  4075. 2:30:10you update the model with this many
  4076. 2:30:12thousands of steps their accuracy kind
  4077. 2:30:14of continues to climb so the models are
  4078. 2:30:17improving and they're solving these
  4079. 2:30:18problems with a higher accuracy
  4080. 2:30:20as you do this trial and error on a
  4081. 2:30:22large data set of these kinds of
  4082. 2:30:24problems and the models are discovering
  4083. 2:30:26how to solve math problems but even more
  4084. 2:30:29incredible than the quantitative kind of
  4085. 2:30:32results of solving these problems with a
  4086. 2:30:33higher accuracy is the qualitative means
  4087. 2:30:35by which the model achieves these
  4088. 2:30:37results so when we scroll down uh one of
  4089. 2:30:40the figures here that is kind of
  4090. 2:30:41interesting is that later on in the
  4091. 2:30:43optimization the model seems to be uh
  4092. 2:30:46using average length per response uh
  4093. 2:30:49goes up up so the model seems to be
  4094. 2:30:51using more tokens to get its higher
  4095. 2:30:54accuracy results so it's learning to
  4096. 2:30:56create very very long Solutions why are
  4097. 2:30:59these Solutions very long we can look at
  4098. 2:31:00them qualitatively here so basically
  4099. 2:31:03what they discover is that the model
  4100. 2:31:05solution get very very long partially
  4101. 2:31:07because so here's a question and here's
  4102. 2:31:09kind of the answer from the model what
  4103. 2:31:11the model learns to do um and this is an
  4104. 2:31:13immerging property of new optimization
  4105. 2:31:15it just discovers that this is good for
  4106. 2:31:17problem solving is it starts to do stuff
  4107. 2:31:19like this wait wait wait that's Nota
  4108. 2:31:21moment I can flag here let's reevaluate
  4109. 2:31:23this step by step to identify the
  4110. 2:31:25correct sum can be so what is the model
  4111. 2:31:27doing here right the model is basically
  4112. 2:31:30re-evaluating steps it has learned that
  4113. 2:31:32it works better for accuracy to try out
  4114. 2:31:35lots of ideas try something from
  4115. 2:31:37different perspectives retrace reframe
  4116. 2:31:39backtrack is doing a lot of the things
  4117. 2:31:41that you and I are doing in the process
  4118. 2:31:43of problem solving for mathematical
  4119. 2:31:44questions but it's rediscovering what
  4120. 2:31:46happens in your head not what you put
  4121. 2:31:48down on the solution and there is no
  4122. 2:31:50human who can hardcode this stuff in the
  4123. 2:31:52ideal assistant response this is only
  4124. 2:31:55something that can be discovered in the
  4125. 2:31:56process of reinforcement learning
  4126. 2:31:57because you wouldn't know what to put
  4127. 2:31:59here this just turns out to work for the
  4128. 2:32:02model and it improves its accuracy in
  4129. 2:32:04problem solving so the model learns what
  4130. 2:32:06we call these chains of thought in your
  4131. 2:32:08head and it's an emergent property of
  4132. 2:32:10the optim of the optimization and that's
  4133. 2:32:13what's bloating up the response length
  4134. 2:32:16but that's also what's increasing the
  4135. 2:32:18accuracy of the problem problem solving
  4136. 2:32:20so what's incredible here is basically
  4137. 2:32:22the model is discovering ways to think
  4138. 2:32:24it's learning what I like to call
  4139. 2:32:26cognitive strategies of how you
  4140. 2:32:28manipulate a problem and how you
  4141. 2:32:30approach it from different perspectives
  4142. 2:32:31how you pull in some analogies or do
  4143. 2:32:33different kinds of things like that and
  4144. 2:32:35how you kind of uh try out many
  4145. 2:32:37different things over time uh check a
  4146. 2:32:39result from different perspectives and
  4147. 2:32:40how you kind of uh solve problems but
  4148. 2:32:43here it's kind of discovered by the RL
  4149. 2:32:44so extremely incredible to see this
  4150. 2:32:47emerge in the optimization without
  4151. 2:32:48having to hardcode it anywhere the only
  4152. 2:32:50thing we've given it are the correct
  4153. 2:32:52answers and this comes out from trying
  4154. 2:32:54to just solve them correctly which is
  4155. 2:32:56incredible
  4156. 2:32:58um now let's go back to actually the
  4157. 2:33:00problem that we've been working with and
  4158. 2:33:02let's take a look at what it would look
  4159. 2:33:03like uh for uh for this kind of a model
  4160. 2:33:07what we call reasoning or thinking model
  4161. 2:33:09to solve that problem okay so recall
  4162. 2:33:12that this is the problem we've been
  4163. 2:33:13working with and when I pasted it into
  4164. 2:33:15chat GPT 40 I'm getting this kind of a
  4165. 2:33:17response let's take a look at what
  4166. 2:33:19happens when you give this same query to
  4167. 2:33:22what's called a reasoning or a thinking
  4168. 2:33:23model this is a model that was trained
  4169. 2:33:25with reinforcement learning so this
  4170. 2:33:28model described in this paper DC car1 is
  4171. 2:33:30available on chat. dec.com uh so this is
  4172. 2:33:34kind of like the company uh that
  4173. 2:33:35developed is hosting it you have to make
  4174. 2:33:37sure that the Deep think button is
  4175. 2:33:39turned on to get the R1 model as it's
  4176. 2:33:41called we can paste it here and run
  4177. 2:33:44it and so let's take a look at what
  4178. 2:33:46happens now and what is the output of
  4179. 2:33:48the model okay so here's it says so this
  4180. 2:33:51is previously what we get using
  4181. 2:33:53basically what's an sft approach a
  4182. 2:33:54supervised funing approach this is like
  4183. 2:33:56mimicking an expert solution this is
  4184. 2:33:58what we get from the RL model okay let
  4185. 2:34:01me try to figure this out so Emily buys
  4186. 2:34:03three apples and two oranges each orange
  4187. 2:34:05cost $2 total is 13 I need to find out
  4188. 2:34:07blah blah blah so here you you um as
  4189. 2:34:11you're reading this you can't escape
  4190. 2:34:14thinking that this model is
  4191. 2:34:16thinking um is definitely pursuing the
  4192. 2:34:19solution solution it deres that it must
  4193. 2:34:21cost $3 and then it says wait a second
  4194. 2:34:23let me check my math again to be sure
  4195. 2:34:25and then it tries it from a slightly
  4196. 2:34:26different perspective and then it says
  4197. 2:34:28yep all that checks out I think that's
  4198. 2:34:30the answer I don't see any mistakes let
  4199. 2:34:33me see if there's another way to
  4200. 2:34:34approach the problem maybe setting up an
  4201. 2:34:36equation let's let the cost of one apple
  4202. 2:34:39be $8 then blah blah blah yep same
  4203. 2:34:42answer so definitely each apple is $3
  4204. 2:34:44all right confident that that's correct
  4205. 2:34:47and then what it does once it sort of um
  4206. 2:34:49did the thinking process is it writes up
  4207. 2:34:51the nice solution for the human and so
  4208. 2:34:54this is now considering so this is more
  4209. 2:34:56about the correctness aspect and this is
  4210. 2:34:58more about the presentation aspect where
  4211. 2:35:00it kind of like writes it out nicely and
  4212. 2:35:03uh boxes in the correct answer at the
  4213. 2:35:05bottom and so what's incredible about
  4214. 2:35:07this is we get this like thinking
  4215. 2:35:08process of the model and this is what's
  4216. 2:35:10coming from the reinforcement learning
  4217. 2:35:12process this is what's bloating up the
  4218. 2:35:15length of the token sequences they're
  4219. 2:35:16doing thinking and they're trying
  4220. 2:35:17different ways this is what's giving you
  4221. 2:35:20higher accuracy in problem
  4222. 2:35:22solving and this is where we are seeing
  4223. 2:35:24these aha moments and these different
  4224. 2:35:26strategies and these um ideas for how
  4225. 2:35:29you can make sure that you're getting
  4226. 2:35:31the correct
  4227. 2:35:32answer the last point I wanted to make
  4228. 2:35:34is some people are a little bit nervous
  4229. 2:35:36about putting you know very sensitive
  4230. 2:35:38data into chat.com because this is a
  4231. 2:35:41Chinese company so people don't um
  4232. 2:35:43people are a little bit careful and Cy
  4233. 2:35:45with that a little bit um deep seek R1
  4234. 2:35:48is a model that was released by this
  4235. 2:35:50company so this is an open source model
  4236. 2:35:52or open weights model it is available
  4237. 2:35:54for anyone to download and use you will
  4238. 2:35:56not be able to like run it in its full
  4239. 2:35:59um sort of the full model in full
  4240. 2:36:02Precision you won't run that on a
  4241. 2:36:04MacBook but uh or like a local device
  4242. 2:36:07because this is a fairly large model but
  4243. 2:36:08many companies are hosting the full
  4244. 2:36:10largest model one of those companies
  4245. 2:36:12that I like to use is called
  4246. 2:36:14together. so when you go to together.
  4247. 2:36:17you sign up and you go to playgrounds
  4248. 2:36:19you can can select here in the chat deep
  4249. 2:36:21seek R1 and there's many different kinds
  4250. 2:36:23of other models that you can select here
  4251. 2:36:25these are all state-of-the-art models so
  4252. 2:36:27this is kind of similar to the hugging
  4253. 2:36:28face inference playground that we've
  4254. 2:36:29been playing with so far but together. a
  4255. 2:36:32will usually host all the
  4256. 2:36:33state-of-the-art models so select DT
  4257. 2:36:36car1 um you can try to ignore a lot of
  4258. 2:36:38these I think the default settings will
  4259. 2:36:39often be okay and we can put in this and
  4260. 2:36:43because the model was released by Deep
  4261. 2:36:45seek what you're getting here should be
  4262. 2:36:47basically equivalent to what you're
  4263. 2:36:48getting here now because of the
  4264. 2:36:50randomness in the sampling we're going
  4265. 2:36:51to get something slightly different uh
  4266. 2:36:53but in principle this should be uh
  4267. 2:36:55identical in terms of the power of the
  4268. 2:36:57model and you should be able to see the
  4269. 2:36:58same things quantitatively and
  4270. 2:37:00qualitatively uh but uh this model is
  4271. 2:37:02coming from kind of a an American
  4272. 2:37:04company so that's deep seek and that's
  4273. 2:37:07the what's called a reasoning
  4274. 2:37:09model now when I go back to chat uh let
  4275. 2:37:12me go to chat here okay so the models
  4276. 2:37:14that you're going to see in the drop
  4277. 2:37:15down here some of them like 01 03 mini
  4278. 2:37:18O3 mini High Etc they are talking about
  4279. 2:37:21uses Advanced reasoning now what this is
  4280. 2:37:23referring to uses Advanced reasoning is
  4281. 2:37:26it's referring to the fact that it was
  4282. 2:37:27trained by reinforcement learning with
  4283. 2:37:29techniques very similar to those of deep
  4284. 2:37:31C car1 per public statements of opening
  4285. 2:37:34ey employees uh so these are thinking
  4286. 2:37:37models trained with RL and these models
  4287. 2:37:40like GPT 4 or GPT 4 40 mini that you're
  4288. 2:37:42getting in the free tier you should
  4289. 2:37:43think of them as mostly sft models
  4290. 2:37:45supervised fine tuning models they don't
  4291. 2:37:47actually do this like thinking as as you
  4292. 2:37:49see in the RL models and even though
  4293. 2:37:52there's a little bit of reinforcement
  4294. 2:37:53learning involved with these models and
  4295. 2:37:55I'll go that into that in a second these
  4296. 2:37:56are mostly sft models I think you should
  4297. 2:37:58think about it that way so in the same
  4298. 2:38:00way as what we saw here we can pick one
  4299. 2:38:03of the thinking models like say 03 mini
  4300. 2:38:05high and these models by the way might
  4301. 2:38:07not be available to you unless you pay a
  4302. 2:38:09Chachi PT subscription of either $20 per
  4303. 2:38:11month or $200 per month for some of the
  4304. 2:38:14top models so we can pick a thinking
  4305. 2:38:16model and run now what's going to happen
  4306. 2:38:20here is it's going to say reasoning and
  4307. 2:38:21it's going to start to do stuff like
  4308. 2:38:23this and um what we're seeing here is
  4309. 2:38:26not exactly the stuff we're seeing here
  4310. 2:38:29so even though under the hood the model
  4311. 2:38:31produces these kinds of uh kind of
  4312. 2:38:34chains of thought opening ey chooses to
  4313. 2:38:36not show the exact chains of thought in
  4314. 2:38:38the web interface it shows little
  4315. 2:38:40summaries of that of those chains of
  4316. 2:38:42thought and open kind of does this I
  4317. 2:38:44think partly because uh they are worried
  4318. 2:38:46about what's called the distillation
  4319. 2:38:48risk that is that someone could come in
  4320. 2:38:50and actually try to imitate those
  4321. 2:38:51reasoning traces and recover a lot of
  4322. 2:38:53the reasoning performance by just
  4323. 2:38:55imitating the reasoning uh chains of
  4324. 2:38:57thought and so they kind of hide them
  4325. 2:38:59and they only show little summaries of
  4326. 2:39:00them so you're not getting exactly what
  4327. 2:39:02you would get in deep seek as with
  4328. 2:39:04respect to the reasoning itself and then
  4329. 2:39:07they write up the
  4330. 2:39:08solution so these are kind of like
  4331. 2:39:10equivalent even though we're not seeing
  4332. 2:39:12the full under the hood details now in
  4333. 2:39:14terms of the performance uh these models
  4334. 2:39:17and deep seek models are currently rly
  4335. 2:39:19on par I would say it's kind of hard to
  4336. 2:39:21tell because of the evaluations but if
  4337. 2:39:22you're paying $200 per month to open AI
  4338. 2:39:24some of these models I believe are
  4339. 2:39:25currently they basically still look
  4340. 2:39:27better uh but deep seek R1 for now is
  4341. 2:39:30still a very solid choice for a thinking
  4342. 2:39:33model that would be available to you um
  4343. 2:39:36sort of um either on this website or any
  4344. 2:39:39other website because the model is open
  4345. 2:39:40weights you can just download it so
  4346. 2:39:43that's thinking models so what is the
  4347. 2:39:46summary so far well we've talked about
  4348. 2:39:48reinforcement learning and the fact that
  4349. 2:39:50thinking emerges in the process of the
  4350. 2:39:52optimization on when we basically run RL
  4351. 2:39:55on many math uh and kind of code
  4352. 2:39:57problems that have verifiable Solutions
  4353. 2:39:59so there's like an answer three
  4354. 2:40:01Etc now these thinking models you can
  4355. 2:40:04access in for example deep seek or any
  4356. 2:40:07inference provider like together. a and
  4357. 2:40:09choosing deep seek over there these
  4358. 2:40:12thinking models are also available uh in
  4359. 2:40:14chpt under any of the 01 or O3
  4360. 2:40:17models but these GPT 4 R models Etc
  4361. 2:40:20they're not thinking models you should
  4362. 2:40:21think of them as mostly sft models now
  4363. 2:40:25if you are um if you have a prompt that
  4364. 2:40:27requires Advanced reasoning and so on
  4365. 2:40:29you should probably use some of the
  4366. 2:40:30thinking models or at least try them out
  4367. 2:40:32but empirically for a lot of my use when
  4368. 2:40:35you're asking a simpler question there's
  4369. 2:40:36like a knowledge based question or
  4370. 2:40:37something like that this might be
  4371. 2:40:39Overkill like there's no need to think
  4372. 2:40:4030 seconds about some factual question
  4373. 2:40:42so for that I will uh sometimes default
  4374. 2:40:44to just GPT 40 so empirically about 80
  4375. 2:40:4790% of my use is just gp4
  4376. 2:40:49and when I come across a very difficult
  4377. 2:40:51problem like in math and code Etc I will
  4378. 2:40:53reach for the thinking models but then I
  4379. 2:40:56have to wait a bit longer because
  4380. 2:40:57they're thinking um so you can access
  4381. 2:41:00these on chat on deep seek also I wanted
  4382. 2:41:02to point out that um AI studio.
  4383. 2:41:05go.com even though it looks really busy
  4384. 2:41:08really ugly because Google's just unable
  4385. 2:41:10to do this kind of stuff well it's like
  4386. 2:41:13what is happening but if you choose
  4387. 2:41:15model and you choose here Gemini 2.0
  4388. 2:41:17flash thinking experimental 01 21 if you
  4389. 2:41:20choose that one that's also a a kind of
  4390. 2:41:22early experiment experimental of a
  4391. 2:41:25thinking model by Google so we can go
  4392. 2:41:27here and we can give it the same problem
  4393. 2:41:29and click run and this is also a
  4394. 2:41:31thinking problem a thinking model that
  4395. 2:41:33will also do something
  4396. 2:41:35similar and comes out with the right
  4397. 2:41:37answer here so basically Gemini also
  4398. 2:41:40offers a thinking model anthropic
  4399. 2:41:42currently does not offer a thinking
  4400. 2:41:43model but basically this is kind of like
  4401. 2:41:45the frontier development of these llms I
  4402. 2:41:47think RL is kind of like this new
  4403. 2:41:49exciting stage but getting the details
  4404. 2:41:51right is difficult and that's why all
  4405. 2:41:53these models and thinking models are
  4406. 2:41:55currently experimental as of 2025 very
  4407. 2:41:57early 2025 um but this is kind of like
  4408. 2:42:01the frontier development of pushing the
  4409. 2:42:02performance on these very difficult
  4410. 2:42:03problems using reasoning that is
  4411. 2:42:05emerging in these optimizations one more
  4412. 2:42:07connection that I wanted to bring up is
  4413. 2:42:10that the discovery that reinforcement
  4414. 2:42:12learning is extremely powerful way of
  4415. 2:42:14learning is not new to the field of AI
  4416. 2:42:17and one place what we've already seen
  4417. 2:42:19this demonstrated is in the game of Go
  4418. 2:42:22and famously Deep Mind developed the
  4419. 2:42:24system alphago and you can watch a movie
  4420. 2:42:26about it um where the system is learning
  4421. 2:42:29to play the game of go against top human
  4422. 2:42:32players and um when we go to the paper
  4423. 2:42:36underlying alphago so in this paper when
  4424. 2:42:39we scroll
  4425. 2:42:41down we actually find a really
  4426. 2:42:43interesting
  4427. 2:42:44plot um that I think uh is kind of
  4428. 2:42:47familiar uh to us and we're kind of like
  4429. 2:42:49we discovering in the more open domain
  4430. 2:42:51of arbitrary problem solving instead of
  4431. 2:42:53on the closed specific domain of the
  4432. 2:42:55game of Go but basically what they saw
  4433. 2:42:57and we're going to see this in llms as
  4434. 2:42:59well as this becomes more mature is this
  4435. 2:43:03is the ELO rating of playing game of Go
  4436. 2:43:05and this is leas dull an extremely
  4437. 2:43:07strong human player and here what they
  4438. 2:43:09are comparing is the strength of a model
  4439. 2:43:11learned trained by supervised learning
  4440. 2:43:14and a model trained by reinforcement
  4441. 2:43:15learning so the supervised learning
  4442. 2:43:17model is imitating human expert players
  4443. 2:43:20so if you just get a huge amount of
  4444. 2:43:22games played by expert players in the
  4445. 2:43:23game of Go and you try to imitate them
  4446. 2:43:26you are going to get better but then you
  4447. 2:43:28top out and you never quite get better
  4448. 2:43:31than some of the top top top players of
  4449. 2:43:34in the game of Go like LEL so you're
  4450. 2:43:35never going to reach there because
  4451. 2:43:37you're just imitating human players you
  4452. 2:43:39can't fundamentally go beyond a human
  4453. 2:43:40player if you're just imitating human
  4454. 2:43:42players but in a process of
  4455. 2:43:44reinforcement learning is significantly
  4456. 2:43:46more powerful in reinforcement learning
  4457. 2:43:48for a game of Go it means that the
  4458. 2:43:50system is playing moves that empirically
  4459. 2:43:53and statistically lead to win to winning
  4460. 2:43:56the game and so alphago is a system
  4461. 2:43:59where it kind of plays against it itself
  4462. 2:44:02and it's using reinforcement learning to
  4463. 2:44:03create
  4464. 2:44:04rollouts so it's the exact same diagram
  4465. 2:44:07here but there's no prompt it's just uh
  4466. 2:44:10because there's no prompt it's just a
  4467. 2:44:11fixed game of Go but it's trying out
  4468. 2:44:13lots of solutions it's trying out lots
  4469. 2:44:15of plays and then the games that lead to
  4470. 2:44:18a win instead of a specific answer are
  4471. 2:44:20reinforced they're they're made stronger
  4472. 2:44:24and so um the system is learning
  4473. 2:44:26basically the sequences of actions that
  4474. 2:44:28empirically and statistically lead to
  4475. 2:44:30winning the game and reinforcement
  4476. 2:44:32learning is not going to be constrained
  4477. 2:44:34by human performance and reinforcement
  4478. 2:44:36learning can do significantly better and
  4479. 2:44:38overcome even the top players like Lisa
  4480. 2:44:41Dole and so uh probably they could have
  4481. 2:44:44run this longer and they just chose to
  4482. 2:44:46crop it at some point because this costs
  4483. 2:44:47money but this is very powerful
  4484. 2:44:49demonstration of reinforcement learning
  4485. 2:44:51and we're only starting to kind of see
  4486. 2:44:52hints of this diagram in larger language
  4487. 2:44:55models for reasoning problems so we're
  4488. 2:44:58not going to get too far by just
  4489. 2:44:59imitating experts we need to go beyond
  4490. 2:45:01that set up these like little game
  4491. 2:45:03environments and get let let the system
  4492. 2:45:07discover reasoning traces or like ways
  4493. 2:45:09of solving problems uh that are unique
  4494. 2:45:14and that uh just basically work
  4495. 2:45:16well now on this aspect of uniqueness
  4496. 2:45:19notice that when you're doing
  4497. 2:45:19reinforcement learning nothing prevents
  4498. 2:45:21you from veering off the distribution of
  4499. 2:45:24how humans are playing the game and so
  4500. 2:45:26when we go back to uh this alphao search
  4501. 2:45:29here one of the suggested modifications
  4502. 2:45:31is called move 37 and move 37 in alphao
  4503. 2:45:34is referring to a specific point in time
  4504. 2:45:37where alphago basically played a move
  4505. 2:45:40that uh no human expert would play uh so
  4506. 2:45:43the probability of this move uh to be
  4507. 2:45:45played by a human player was evaluated
  4508. 2:45:47to be about 1 in 10th ,000 so it's a
  4509. 2:45:49very rare move but in retrospect it was
  4510. 2:45:52a brilliant move so alphago in the
  4511. 2:45:54process of reinforcement learning
  4512. 2:45:55discovered kind of like a strategy of
  4513. 2:45:57playing that was unknown to humans and
  4514. 2:46:00but is in retrospect uh brilliant I
  4515. 2:46:02recommend this YouTube video um leis do
  4516. 2:46:04versus alphao move 37 reactions and
  4517. 2:46:06Analysis and this is kind of what it
  4518. 2:46:08looked like when alphao played this
  4519. 2:46:11move
  4520. 2:46:14value that's a very that's a very
  4521. 2:46:16surprising move I thought I thought it
  4522. 2:46:19was I thought it was a
  4523. 2:46:21mistake when I see this move anyway so
  4524. 2:46:24basically people are kind of freaking
  4525. 2:46:25out because it's a it's a move that a
  4526. 2:46:28human would not play that alphago played
  4527. 2:46:31because in its training uh this move
  4528. 2:46:33seemed to be a good idea it just happens
  4529. 2:46:35not to be a kind of thing that a humans
  4530. 2:46:37would would do and so that is again the
  4531. 2:46:39power of reinforcement learning and in
  4532. 2:46:41principle we can actually see the
  4533. 2:46:42equivalence of that if we continue
  4534. 2:46:44scaling this Paradigm in language models
  4535. 2:46:46and what that looks like is kind of
  4536. 2:46:47unknown so so um what does it mean to
  4537. 2:46:50solve problems in such a way that uh
  4538. 2:46:54even humans would not be able to get how
  4539. 2:46:56can you be better at reasoning or
  4540. 2:46:58thinking than humans how can you go
  4541. 2:47:00beyond just uh a thinking human like
  4542. 2:47:03maybe it means discovering analogies
  4543. 2:47:05that humans would not be able to uh
  4544. 2:47:07create or maybe it's like a new thinking
  4545. 2:47:09strategy it's kind of hard to think
  4546. 2:47:10through uh maybe it's a holy new
  4547. 2:47:14language that actually is not even
  4548. 2:47:16English maybe it discovers its own
  4549. 2:47:17language that is a lot better at
  4550. 2:47:19thinking um because the model is
  4551. 2:47:22unconstrained to even like stick with
  4552. 2:47:24English uh so maybe it takes a different
  4553. 2:47:27language to think in or it discovers its
  4554. 2:47:29own language so in principle the
  4555. 2:47:31behavior of the system is a lot less
  4556. 2:47:33defined it is open to do whatever works
  4557. 2:47:37and it is open to also slowly Drift from
  4558. 2:47:40the distribution of its training data
  4559. 2:47:41which is English but all of that can
  4560. 2:47:43only be done if we have a very large
  4561. 2:47:45diverse set of problems in which the
  4562. 2:47:48these strategy can be refined and
  4563. 2:47:49perfected and so that is a lot of the
  4564. 2:47:51frontier LM research that's going on
  4565. 2:47:53right now is trying to kind of create
  4566. 2:47:55those kinds of prompt distributions that
  4567. 2:47:57are large and diverse these are all kind
  4568. 2:47:59of like game environments in which the
  4569. 2:48:00llms can practice their thinking and uh
  4570. 2:48:04it's kind of like writing you know these
  4571. 2:48:06practice problems we have to create
  4572. 2:48:07practice problems for all of domains of
  4573. 2:48:10knowledge and if we have practice
  4574. 2:48:12problems and tons of them the models
  4575. 2:48:14will be able to reinforcement learning
  4576. 2:48:16reinforcement learn on them and kind of
  4577. 2:48:18uh create these kinds of uh diagrams but
  4578. 2:48:21in the domain of open thinking instead
  4579. 2:48:23of a closed domain like game of Go
  4580. 2:48:26there's one more section within
  4581. 2:48:27reinforcement learning that I wanted to
  4582. 2:48:29cover and that is that of learning in
  4583. 2:48:32unverifiable domains so so far all of
  4584. 2:48:35the problems that we've looked at are in
  4585. 2:48:36what's called verifiable domains that is
  4586. 2:48:38any candidate solution we can score very
  4587. 2:48:41easily against a concrete answer so for
  4588. 2:48:44example answer is three and we can very
  4589. 2:48:45easily score these Solutions against the
  4590. 2:48:47answer of three
  4591. 2:48:49either we require the models to like box
  4592. 2:48:51in their answers and then we just check
  4593. 2:48:53for equality of whatever is in the box
  4594. 2:48:55with the answer or you can also use uh
  4595. 2:48:58kind of what's called an llm judge so
  4596. 2:49:00the llm judge looks at a solution and it
  4597. 2:49:03gets the answer and just basically
  4598. 2:49:05scores the solution for whether it's
  4599. 2:49:06consistent with the answer or not and
  4600. 2:49:08llms uh empirically are good enough at
  4601. 2:49:10the current capability that they can do
  4602. 2:49:12this fairly reliably so we can apply
  4603. 2:49:14those kinds of techniques as well in any
  4604. 2:49:16case we have a concrete answer and we're
  4605. 2:49:17just checking Solutions again against it
  4606. 2:49:19and we can do this automatically with no
  4607. 2:49:21kind of humans in the loop the problem
  4608. 2:49:23is that we can't apply the strategy in
  4609. 2:49:25what's called unverifiable domains so
  4610. 2:49:28usually these are for example creative
  4611. 2:49:29writing tasks like write a joke about
  4612. 2:49:31Pelicans or write a poem or summarize a
  4613. 2:49:33paragraph or something like that in
  4614. 2:49:35these kinds of domains it becomes harder
  4615. 2:49:37to score our different solutions to this
  4616. 2:49:39problem so for example writing a joke
  4617. 2:49:41about Pelicans we can generate lots of
  4618. 2:49:43different uh jokes of course that's fine
  4619. 2:49:45for example we can go to chbt and we can
  4620. 2:49:47get it to uh generate a joke about
  4621. 2:49:51Pelicans uh so much stuff in their beaks
  4622. 2:49:53because they don't bellan in
  4623. 2:49:56backpacks what
  4624. 2:49:59okay we can uh we can try something else
  4625. 2:50:02why don't Pelicans ever pay for their
  4626. 2:50:04drinks because they always B it to
  4627. 2:50:06someone else haha okay so these models
  4628. 2:50:10are not obviously not very good at humor
  4629. 2:50:12actually I think it's pretty fascinating
  4630. 2:50:13because I think humor is secretly very
  4631. 2:50:15difficult and the model have the
  4632. 2:50:16capability I think anyway in any case
  4633. 2:50:20you could imagine creating lots of jokes
  4634. 2:50:23the problem that we are facing is how do
  4635. 2:50:24we score them now in principle we could
  4636. 2:50:27of course get a human to look at all
  4637. 2:50:29these jokes just like I did right now
  4638. 2:50:31the problem with that is if you are
  4639. 2:50:32doing reinforcement learning you're
  4640. 2:50:34going to be doing many thousands of
  4641. 2:50:36updates and for each update you want to
  4642. 2:50:38be looking at say thousands of prompts
  4643. 2:50:40and for each prompt you want to be
  4644. 2:50:41potentially looking at looking at
  4645. 2:50:43hundred or thousands of different kinds
  4646. 2:50:44of generations and so there's just like
  4647. 2:50:47way too many of these to look at and so
  4648. 2:50:50um in principle you could have a human
  4649. 2:50:52inspect all of them and score them and
  4650. 2:50:53decide that okay maybe this one is funny
  4651. 2:50:55and uh maybe this one is funny and this
  4652. 2:50:58one is funny and we could train on them
  4653. 2:51:01to get the model to become slightly
  4654. 2:51:02better at jokes um in the context of
  4655. 2:51:05pelicans at least um the problem is that
  4656. 2:51:09it's just like way too much human time
  4657. 2:51:10this is an unscalable strategy we need
  4658. 2:51:12some kind of an automatic strategy for
  4659. 2:51:14doing this and one sort of solution to
  4660. 2:51:16this was proposed in this paper
  4661. 2:51:19uh that introduced what's called
  4662. 2:51:20reinforcement learning from Human
  4663. 2:51:21feedback and so this was a paper from
  4664. 2:51:23open at the time and many of these
  4665. 2:51:25people are now um co-founders in
  4666. 2:51:27anthropic um and this kind of proposed a
  4667. 2:51:30approach for uh basically doing
  4668. 2:51:33reinforcement learning in unverifiable
  4669. 2:51:35domains so let's take a look at how that
  4670. 2:51:36works so this is the cartoon diagram of
  4671. 2:51:39the core ideas involved so as I
  4672. 2:51:41mentioned the native approach is if we
  4673. 2:51:44just set Infinity human time we could
  4674. 2:51:46just run RL in these domains just fine
  4675. 2:51:49so for example we can run RL as usual if
  4676. 2:51:51I have Infinity humans I would I just
  4677. 2:51:53want to do and these are just cartoon
  4678. 2:51:55numbers I want to do 1,000 updates where
  4679. 2:51:57each update will be on 1,000 prompts and
  4680. 2:52:00in for each prompt we're going to have
  4681. 2:52:021,000 roll outs that we're scoring so we
  4682. 2:52:05can run RL with this kind of a setup the
  4683. 2:52:08problem is in the process of doing this
  4684. 2:52:10I will need to run one I will need to
  4685. 2:52:12ask a human to evaluate a joke a total
  4686. 2:52:15of 1 billion times and so that's a lot
  4687. 2:52:18of people looking at really terrible
  4688. 2:52:19jokes so we don't want to do that so
  4689. 2:52:22instead we want to take the arlef
  4690. 2:52:24approach so um in our Rel of approach we
  4691. 2:52:27are kind of like the the core trick is
  4692. 2:52:29that of indirection so we're going to
  4693. 2:52:32involve humans just a little bit and the
  4694. 2:52:35way we cheat is that we basically train
  4695. 2:52:37a whole separate neural network that we
  4696. 2:52:39call a reward model and this neural
  4697. 2:52:41network will kind of like imitate human
  4698. 2:52:44scores so we're going to ask humans to
  4699. 2:52:46score um roll
  4700. 2:52:49we're going to then imitate human scores
  4701. 2:52:51using a neural network and this neural
  4702. 2:52:54network will become a kind of simulator
  4703. 2:52:55of human
  4704. 2:52:56preferences and now that we have a
  4705. 2:52:58neural network simulator we can do RL
  4706. 2:53:01against it so instead of asking a real
  4707. 2:53:03human we're asking a simulated human for
  4708. 2:53:06their score of a joke as an example and
  4709. 2:53:09so once we have a simulator we're often
  4710. 2:53:11racist because we can query it as many
  4711. 2:53:13times as we want to and it's all whole
  4712. 2:53:16automatic process and we can now do
  4713. 2:53:17reinforcement learning with respect to
  4714. 2:53:19the simulator and the simulator as you
  4715. 2:53:20might expect is not going to be a
  4716. 2:53:22perfect human but if it's at least
  4717. 2:53:24statistically similar to human judgment
  4718. 2:53:26then you might expect that this will do
  4719. 2:53:28something and in practice indeed uh it
  4720. 2:53:30does so once we have a simulator we can
  4721. 2:53:32do RL and everything works great so let
  4722. 2:53:35me show you a cartoon diagram a little
  4723. 2:53:36bit of what this process looks like
  4724. 2:53:38although the details are not 100 like
  4725. 2:53:40super important it's just a core idea of
  4726. 2:53:42how this works so here I have a cartoon
  4727. 2:53:44diagram of a hypothetical example of
  4728. 2:53:46what training the reward model would
  4729. 2:53:47look like so we have a prompt like write
  4730. 2:53:50a joke about picans and then here we
  4731. 2:53:52have five separate roll outs so these
  4732. 2:53:54are all five different jokes just like
  4733. 2:53:56this one now the first thing we're going
  4734. 2:53:59to do is we are going to ask a human to
  4735. 2:54:02uh order these jokes from the best to
  4736. 2:54:05worst so this is uh so here this human
  4737. 2:54:08thought that this joke is the best the
  4738. 2:54:10funniest so number one joke this is
  4739. 2:54:14number two joke number three joke four
  4740. 2:54:16and five so this is the worst joke
  4741. 2:54:19we're asking humans to order instead of
  4742. 2:54:20give scores directly because it's a bit
  4743. 2:54:22of an easier task it's easier for a
  4744. 2:54:24human to give an ordering than to give
  4745. 2:54:26precise scores now that is now the
  4746. 2:54:29supervision for the model so the human
  4747. 2:54:31has ordered them and that is kind of
  4748. 2:54:32like their contribution to the training
  4749. 2:54:34process but now separately what we're
  4750. 2:54:36going to do is we're going to ask a
  4751. 2:54:37reward model uh about its scoring of
  4752. 2:54:40these jokes now the reward model is a
  4753. 2:54:42whole separate neural network completely
  4754. 2:54:44separate neural net um and it's also
  4755. 2:54:47probably a transform
  4756. 2:54:49uh but it's not a language model in the
  4757. 2:54:50sense that it generates diverse language
  4758. 2:54:53Etc it's just a scoring model so the
  4759. 2:54:56reward model will take as an input The
  4760. 2:54:59Prompt number one and number two a
  4761. 2:55:02candidate joke so um those are the two
  4762. 2:55:05inputs that go into the reward model so
  4763. 2:55:07here for example the reward model would
  4764. 2:55:08be taken this prompt and this joke now
  4765. 2:55:11the output of a reward model is a single
  4766. 2:55:14number and this number is thought of as
  4767. 2:55:16a score and it can range for example
  4768. 2:55:18from Z to one so zero would be the worst
  4769. 2:55:20score and one would be the best score so
  4770. 2:55:23here are some examples of what a
  4771. 2:55:25hypothetical reward model at some stage
  4772. 2:55:27in the training process would give uh s
  4773. 2:55:29scoring to these jokes so 0.1 is a very
  4774. 2:55:33low score 08 is a really high score and
  4775. 2:55:36so on and so now um we compare the
  4776. 2:55:40scores given by the reward model with uh
  4777. 2:55:43the ordering given by the human and
  4778. 2:55:45there's a precise mathematical way to
  4779. 2:55:47actually calculate this uh basically set
  4780. 2:55:49up a loss function and calculate a kind
  4781. 2:55:51of like a correspondence here and uh
  4782. 2:55:54update a model based on it but I just
  4783. 2:55:55want to give you the intuition which is
  4784. 2:55:57that as an example here for this second
  4785. 2:56:00joke the the human thought that it was
  4786. 2:56:02the funniest and the model kind of
  4787. 2:56:03agreed right 08 is a relatively high
  4788. 2:56:05score but this score should have been
  4789. 2:56:07even higher right so after an update we
  4790. 2:56:10would expect that maybe this score
  4791. 2:56:11should have been will actually grow
  4792. 2:56:13after an update of the network to be
  4793. 2:56:15like say 081 or
  4794. 2:56:16something um for this one here they
  4795. 2:56:19actually are in a massive disagreement
  4796. 2:56:21because the human thought that this was
  4797. 2:56:22number two but here the the score is
  4798. 2:56:24only 0.1 and so this score needs to be
  4799. 2:56:27much higher so after an update on top of
  4800. 2:56:30this um kind of a supervision this might
  4801. 2:56:33grow a lot more like maybe it's 0.15 or
  4802. 2:56:35something like
  4803. 2:56:36that um and then here the human thought
  4804. 2:56:39that this one was the worst joke but
  4805. 2:56:41here the model actually gave it a fairly
  4806. 2:56:43High number so you might expect that
  4807. 2:56:45after the update uh this would come down
  4808. 2:56:47to maybe 3 3.5 or something like that so
  4809. 2:56:50basically we're doing what we did before
  4810. 2:56:51we're slightly nudging the predictions
  4811. 2:56:54from the models using a neural network
  4812. 2:56:57training
  4813. 2:56:58process and we're trying to make the
  4814. 2:57:00reward model scores be consistent with
  4815. 2:57:03human
  4816. 2:57:04ordering and so um as we update the
  4817. 2:57:07reward model on human data it becomes
  4818. 2:57:09better and better simulator of the
  4819. 2:57:11scores and orders uh that humans provide
  4820. 2:57:14and then becomes kind of like the the
  4821. 2:57:17neural the simulator of human
  4822. 2:57:18preferences which we can then do RL
  4823. 2:57:20against but critically we're not asking
  4824. 2:57:23humans one billion times to look at a
  4825. 2:57:24joke we're maybe looking at th000
  4826. 2:57:26prompts and five roll outs each so maybe
  4827. 2:57:285,000 jokes that humans have to look at
  4828. 2:57:30in total and they just give the ordering
  4829. 2:57:33and then we're training the model to be
  4830. 2:57:34consistent with that ordering and I'm
  4831. 2:57:36skipping over the mathematical details
  4832. 2:57:38but I just want you to understand a high
  4833. 2:57:39level idea that uh this reward model is
  4834. 2:57:42do is basically giving us this scour and
  4835. 2:57:45we have a way of training it to be
  4836. 2:57:46consistent with human orderings
  4837. 2:57:48and that's how rhf works okay so that is
  4838. 2:57:51the rough idea we basically train
  4839. 2:57:53simulators of humans and RL with respect
  4840. 2:57:55to those
  4841. 2:57:56simulators now I want to talk about
  4842. 2:57:59first the upside of reinforcement
  4843. 2:58:00learning from Human
  4844. 2:58:03feedback the first thing is that this
  4845. 2:58:05allows us to run reinforcement learning
  4846. 2:58:07which we know is incredibly powerful
  4847. 2:58:09kind of set of techniques and it allows
  4848. 2:58:10us to do it in arbitrary domains and
  4849. 2:58:13including the ones that are unverifiable
  4850. 2:58:15so things like summarization and poem
  4851. 2:58:17writing joke writing or any other
  4852. 2:58:19creative writing really uh in domains
  4853. 2:58:21outside of math and code
  4854. 2:58:23Etc now empirically what we see when we
  4855. 2:58:25actually apply rhf is that this is a way
  4856. 2:58:28to improve the performance of the model
  4857. 2:58:30and uh I have a top answer for why that
  4858. 2:58:33might be but I don't actually know that
  4859. 2:58:35it is like super well established on
  4860. 2:58:38like why this is you can empirically
  4861. 2:58:39observe that when you do rhf correctly
  4862. 2:58:41the models you get are just like a
  4863. 2:58:43little bit better um but as to why is I
  4864. 2:58:45think like not as clear so here's my
  4865. 2:58:47best guess my best guess is that this is
  4866. 2:58:49possibly mostly due to the discriminator
  4867. 2:58:52generator
  4868. 2:58:53Gap what that means is that in many
  4869. 2:58:55cases it is significantly easier to
  4870. 2:58:58discriminate than to generate for humans
  4871. 2:59:01so in particular an example of this is
  4872. 2:59:04um in when we do supervised fine-tuning
  4873. 2:59:07right
  4874. 2:59:09sft we're asking humans to generate the
  4875. 2:59:12ideal assistant response and in many
  4876. 2:59:15cases here um as I've shown it uh the
  4877. 2:59:18ideal response is very simple to write
  4878. 2:59:20but in many cases might not be so for
  4879. 2:59:22example in summarization or poem writing
  4880. 2:59:24or joke writing like how are you as a
  4881. 2:59:26human assist as a human labeler um
  4882. 2:59:29supposed to give the ideal response in
  4883. 2:59:30these cases it requires creative human
  4884. 2:59:32writing to do that and so rhf kind of
  4885. 2:59:35sidesteps this because we get um we get
  4886. 2:59:38to ask people a significantly easier
  4887. 2:59:40question as a data labelers they're not
  4888. 2:59:42asked to write poems directly they're
  4889. 2:59:44just given five poems from the model and
  4890. 2:59:46they're just asked to order them and so
  4891. 2:59:49that's just a much easier task for a
  4892. 2:59:51human labeler to do and so what I think
  4893. 2:59:53this allows you to do basically is it um
  4894. 2:59:57it kind of like allows a lot more higher
  4895. 3:00:00accuracy data because we're not asking
  4896. 3:00:02people to do the generation task which
  4897. 3:00:04can be extremely difficult like we're
  4898. 3:00:06not asking them to do creative writing
  4899. 3:00:07we're just trying to get them to
  4900. 3:00:09distinguish between creative writings
  4901. 3:00:11and uh find the ones that are best and
  4902. 3:00:14that is the signal that humans are
  4903. 3:00:15providing just the ordering and that is
  4904. 3:00:17their input into the system and then the
  4905. 3:00:20system in rhf just discovers the kinds
  4906. 3:00:23of responses that would be graded well
  4907. 3:00:26by humans and so that step of
  4908. 3:00:28indirection allows the models to become
  4909. 3:00:30a bit better so that is the upside of
  4910. 3:00:33our LF it allows us to run RL it
  4911. 3:00:35empirically results in better models and
  4912. 3:00:37it allows uh people to contribute their
  4913. 3:00:40supervision uh even without having to do
  4914. 3:00:42extremely difficult tasks um in the case
  4915. 3:00:45of writing ideal responses unfortunately
  4916. 3:00:47our HF also comes with significant
  4917. 3:00:49downsides and so um the main one is that
  4918. 3:00:54basically we are doing reinforcement
  4919. 3:00:55learning not with respect to humans and
  4920. 3:00:57actual human judgment but with respect
  4921. 3:00:59to a lossy simulation of humans right
  4922. 3:01:01and this lossy simulation could be
  4923. 3:01:03misleading because it's just a it's just
  4924. 3:01:05a simulation right it's just a language
  4925. 3:01:07model that's kind of outputting scores
  4926. 3:01:09and it might not perfectly reflect the
  4927. 3:01:11opinion of an actual human with an
  4928. 3:01:13actual brain in all the possible
  4929. 3:01:15different cases so that's number one
  4930. 3:01:17which is actually something even more
  4931. 3:01:18subtle and devious going on that uh
  4932. 3:01:21really
  4933. 3:01:22dramatically holds back our LF as a
  4934. 3:01:24technique that we can really scale to
  4935. 3:01:27significantly um kind of Smart Systems
  4936. 3:01:31and that is that reinforcement learning
  4937. 3:01:32is extremely good at discovering a way
  4938. 3:01:35to game the model to game the simulation
  4939. 3:01:38so this reward model that we're
  4940. 3:01:40constructing here that gives the course
  4941. 3:01:43these models are Transformers these
  4942. 3:01:46Transformers are massive neurals they
  4943. 3:01:48have billions of parameters and they
  4944. 3:01:50imitate humans but they do so in a kind
  4945. 3:01:52of like a simulation way now the problem
  4946. 3:01:54is that these are massive complicated
  4947. 3:01:56systems right there's a billion
  4948. 3:01:57parameters here that are outputting a
  4949. 3:01:58single
  4950. 3:02:00score it turns out that there are ways
  4951. 3:02:02to gain these models you can find kinds
  4952. 3:02:05of inputs that were not part of their
  4953. 3:02:08training set and these inputs
  4954. 3:02:11inexplicably get very high scores but in
  4955. 3:02:13a fake way so very often what you find
  4956. 3:02:17if you run our lch for very long so for
  4957. 3:02:19example if we do 1,000 updates which is
  4958. 3:02:21like say a lot of updates you might
  4959. 3:02:23expect that your jokes are getting
  4960. 3:02:25better and that you're getting like real
  4961. 3:02:26bangers about Pelicans but that's not
  4962. 3:02:28EXA exactly what happens what happens is
  4963. 3:02:31that uh in the first few hundred steps
  4964. 3:02:34the jokes about Pelicans are probably
  4965. 3:02:35improving a little bit and then they
  4966. 3:02:37actually dramatically fall off the cliff
  4967. 3:02:38and you start to get extremely
  4968. 3:02:40nonsensical results like for example you
  4969. 3:02:42start to get um the top joke about
  4970. 3:02:45Pelicans starts to be the
  4971. 3:02:48and this makes no sense right like when
  4972. 3:02:49you look at it why should this be a top
  4973. 3:02:50joke but when you take the the and you
  4974. 3:02:53plug it into your reward model you'd
  4975. 3:02:55expect score of zero but actually the
  4976. 3:02:57reward model loves this as a joke it
  4977. 3:02:59will tell you that the the the theth is
  4978. 3:03:02a score of 1. Z this is a top joke and
  4979. 3:03:06this makes no sense right but it's
  4980. 3:03:07because these models are just
  4981. 3:03:09simulations of humans and they're
  4982. 3:03:10massive neural lots and you can find
  4983. 3:03:12inputs at the bottom that kind of like
  4984. 3:03:15get into the part of the input space
  4985. 3:03:16that kind of gives you nonsensical
  4986. 3:03:17results these examples are what's called
  4987. 3:03:20adversarial examples and I'm not going
  4988. 3:03:22to go into the topic too much but these
  4989. 3:03:24are adversarial inputs to the model they
  4990. 3:03:26are specific little inputs that kind of
  4991. 3:03:29go between the nooks and crannies of the
  4992. 3:03:30model and give nonsensical results at
  4993. 3:03:32the top now here's what you might
  4994. 3:03:34imagine doing you say okay the the the
  4995. 3:03:36is obviously not score of one um it's
  4996. 3:03:39obviously a low score so let's take the
  4997. 3:03:41the the the the let's add it to the data
  4998. 3:03:43set and give it an ordering that is
  4999. 3:03:45extremely bad like a score of five and
  5000. 3:03:47indeed your model will learn that the D
  5001. 3:03:50should have a very low score and it will
  5002. 3:03:51give it score of zero the problem is
  5003. 3:03:53that there will always be basically
  5004. 3:03:55infinite number of nonsensical
  5005. 3:03:57adversarial examples hiding in the model
  5006. 3:04:00if you iterate this process many times
  5007. 3:04:02and you keep adding nonsensical stuff to
  5008. 3:04:04your reward model and giving it very low
  5009. 3:04:05scores you can you'll never win the game
  5010. 3:04:09uh you can do this many many rounds and
  5011. 3:04:11reinforcement learning if you run it
  5012. 3:04:12long enough will always find a way to
  5013. 3:04:14gain the model it will discover
  5014. 3:04:15adversarial examples it will get get
  5015. 3:04:17really high scores uh with nonsensical
  5016. 3:04:20results and fundamentally this is
  5017. 3:04:23because our scoring function is a giant
  5018. 3:04:26neural nut and RL is extremely good at
  5019. 3:04:28finding just the ways to trick it uh so
  5020. 3:04:33long story short you always run rhf put
  5021. 3:04:36for maybe a few hundred updates the
  5022. 3:04:38model is getting better and then you
  5023. 3:04:39have to crop it and you are done you
  5024. 3:04:42can't run too much against this reward
  5025. 3:04:45model because the optimization will
  5026. 3:04:47start to game it and you basically crop
  5027. 3:04:50it and you call it and you ship it um
  5028. 3:04:53and uh you can improve the reward model
  5029. 3:04:56but you kind of like come across these
  5030. 3:04:57situations eventually at some point so
  5031. 3:05:00rhf basically what I usually say is that
  5032. 3:05:03RF is not RL and what I mean by that is
  5033. 3:05:06I mean RF is RL obviously but it's not
  5034. 3:05:09RL in the magical sense this is not RL
  5035. 3:05:12that you can run
  5036. 3:05:13indefinitely these kinds of problems
  5037. 3:05:16like where you are getting con correct
  5038. 3:05:18answer you cannot gain this as easily
  5039. 3:05:20you either got the correct answer or you
  5040. 3:05:21didn't and the scoring function is much
  5041. 3:05:23much simpler you're just looking at the
  5042. 3:05:25boxed area and seeing if the result is
  5043. 3:05:27correct so it's very difficult to gain
  5044. 3:05:29these functions but uh gaming a reward
  5045. 3:05:32model is possible now in these
  5046. 3:05:34verifiable domains you can run RL
  5047. 3:05:36indefinitely you could run for tens of
  5048. 3:05:38thousands hundreds of thousands of steps
  5049. 3:05:40and discover all kinds of really crazy
  5050. 3:05:41strategies that we might not even ever
  5051. 3:05:43think about of Performing really well
  5052. 3:05:45for all these problems in the game of Go
  5053. 3:05:48there's no way to to beat to basically
  5054. 3:05:50game uh the winning of a game or the
  5055. 3:05:52losing of a game we have a perfect
  5056. 3:05:54simulator we know all the different uh
  5057. 3:05:57where all the stones are placed and we
  5058. 3:05:59can calculate uh whether someone has won
  5059. 3:06:01or not there's no way to gain that and
  5060. 3:06:03so you can do RL indefinitely and you
  5061. 3:06:05can eventually be beat even leol but
  5062. 3:06:08with models like this which are gameable
  5063. 3:06:11you cannot repeat this process
  5064. 3:06:13indefinitely so I kind of see rhf as not
  5065. 3:06:16real RL because the reward function is
  5066. 3:06:19gameable so it's kind of more like in
  5067. 3:06:21the realm of like little fine-tuning
  5068. 3:06:23it's a little it's a little Improvement
  5069. 3:06:26but it's not something that is
  5070. 3:06:27fundamentally set up correctly where you
  5071. 3:06:29can insert more compute run for longer
  5072. 3:06:32and get much better and magical results
  5073. 3:06:34so it's it's uh it's not RL in that
  5074. 3:06:36sense it's not RL in the sense that it
  5075. 3:06:38lacks magic um it can find you in your
  5076. 3:06:41model and get a better performance and
  5077. 3:06:43indeed if we go back to chat GPT the GPT
  5078. 3:06:4640 model has gone through rhf because it
  5079. 3:06:50works well but it's just not RL in the
  5080. 3:06:52same sense rlf is like a little fine
  5081. 3:06:54tune that slightly improves your model
  5082. 3:06:56is maybe like the way I would think
  5083. 3:06:57about it okay so that's most of the
  5084. 3:06:59technical content that I wanted to cover
  5085. 3:07:01I took you through the three major
  5086. 3:07:03stages and paradigms of training these
  5087. 3:07:05models pre-training supervised fine
  5088. 3:07:07tuning and reinforcement learning and I
  5089. 3:07:09showed you that they Loosely correspond
  5090. 3:07:11to the process we already use for
  5091. 3:07:12teaching children and so in particular
  5092. 3:07:15we talked about pre-training being sort
  5093. 3:07:17of like the basic knowledge acquisition
  5094. 3:07:18of reading Exposition supervised fine
  5095. 3:07:21tuning being the process of looking at
  5096. 3:07:22lots and lots of worked examples and
  5097. 3:07:24imitating experts and practice problems
  5098. 3:07:28the only difference is that we now have
  5099. 3:07:30to effectively write textbooks for llms
  5100. 3:07:32and AIS across all the disciplines of
  5101. 3:07:35human knowledge and also in all the
  5102. 3:07:37cases where we actually would like them
  5103. 3:07:39to work like code and math and you know
  5104. 3:07:42basically all the other disciplines so
  5105. 3:07:44we're in the process of writing
  5106. 3:07:45textbooks for them refining all the
  5107. 3:07:47algorithms that I've presented on the
  5108. 3:07:48high level and then of course doing a
  5109. 3:07:50really really good job at the execution
  5110. 3:07:52of training these models at scale and
  5111. 3:07:54efficiently so in particular I didn't go
  5112. 3:07:56into too many details but these are
  5113. 3:07:58extremely large and complicated
  5114. 3:08:00distributed uh sort of
  5115. 3:08:04um jobs that have to run over tens of
  5116. 3:08:07thousands or even hundreds of thousands
  5117. 3:08:08of gpus and the engineering that goes
  5118. 3:08:10into this is really at the stateof the
  5119. 3:08:12art of what's possible with computers at
  5120. 3:08:14that scale so I didn't cover that aspect
  5121. 3:08:17too much
  5122. 3:08:19but um this is very kind of serious and
  5123. 3:08:22they were underlying all these very
  5124. 3:08:24simple algorithms
  5125. 3:08:25ultimately now I also talked about sort
  5126. 3:08:28of like the theory of mind a little bit
  5127. 3:08:30of these models and the thing I want you
  5128. 3:08:31to take away is that these models are
  5129. 3:08:33really good but they're extremely useful
  5130. 3:08:35as tools for your work you shouldn't uh
  5131. 3:08:38sort of trust them fully and I showed
  5132. 3:08:39you some examples of that even though we
  5133. 3:08:41have mitigations for hallucinations the
  5134. 3:08:43models are not perfect and they will
  5135. 3:08:44hallucinate still it's gotten better
  5136. 3:08:46over time and it will continue to get
  5137. 3:08:48better but they can
  5138. 3:08:49hallucinate in other words in in
  5139. 3:08:52addition to that I covered kind of like
  5140. 3:08:53what I call the Swiss cheese uh sort of
  5141. 3:08:56model of llm capabilities that you
  5142. 3:08:57should have in your mind the models are
  5143. 3:08:59incredibly good across so many different
  5144. 3:09:00disciplines but then fail randomly
  5145. 3:09:02almost in some unique cases so for
  5146. 3:09:05example what is bigger 9.11 or 9.9 like
  5147. 3:09:07the model doesn't know but
  5148. 3:09:09simultaneously it can turn around and
  5149. 3:09:11solve Olympiad questions and so this is
  5150. 3:09:14a hole in the Swiss cheese and there are
  5151. 3:09:16many of them and you don't want to trip
  5152. 3:09:17over them so don't um treat these models
  5153. 3:09:21as infallible models check their work
  5154. 3:09:23use them as tools use them for
  5155. 3:09:25inspiration use them for the first draft
  5156. 3:09:28but uh work with them as tools and be
  5157. 3:09:30ultimately respons responsible for the
  5158. 3:09:32you know product of your
  5159. 3:09:35work and that's roughly what I wanted to
  5160. 3:09:38talk about this is how they're trained
  5161. 3:09:40and this is what they are let's now turn
  5162. 3:09:43to what are some of the future
  5163. 3:09:44capabilities of these models uh probably
  5164. 3:09:46what's coming down the pipe and also
  5165. 3:09:48where can you find these models I have a
  5166. 3:09:50few blow points on some of the things
  5167. 3:09:51that you can expect coming down the pipe
  5168. 3:09:53the first thing you'll notice is that
  5169. 3:09:55the models will very rapidly become
  5170. 3:09:56multimodal everything I talked about
  5171. 3:09:58above concerned text but very soon we'll
  5172. 3:10:01have llms that can not just handle text
  5173. 3:10:03but they can also operate natively and
  5174. 3:10:05very easily over audio so they can hear
  5175. 3:10:08and speak and also images so they can
  5176. 3:10:10see and paint and we're already seeing
  5177. 3:10:13the beginnings of all of this uh but
  5178. 3:10:15this will be all done natively inside
  5179. 3:10:17inside the language model and this will
  5180. 3:10:19enable kind of like natural
  5181. 3:10:20conversations and roughly speaking the
  5182. 3:10:22reason that this is actually no
  5183. 3:10:23different from everything we've covered
  5184. 3:10:24above is that as a baseline you can
  5185. 3:10:28tokenize audio and images and apply the
  5186. 3:10:31exact same approaches of everything that
  5187. 3:10:32we've talked about above so it's not a
  5188. 3:10:34fundamental change it's just uh it's
  5189. 3:10:36just a to we have to add some tokens so
  5190. 3:10:38as an example for tokenizing audio we
  5191. 3:10:41can look at slices of the spectrogram of
  5192. 3:10:43the audio signal and we can tokenize
  5193. 3:10:45that and just add more tokens that
  5194. 3:10:47suddenly represent audio and just add
  5195. 3:10:50them into the context windows and train
  5196. 3:10:51on them just like above the same for
  5197. 3:10:53images we can use patches and we can
  5198. 3:10:56separately tokenize patches and then
  5199. 3:10:58what is an image an image is just a
  5200. 3:11:00sequence of tokens and this actually
  5201. 3:11:03kind of works and there's a lot of early
  5202. 3:11:04work in this direction and so we can
  5203. 3:11:06just create streams of tokens that are
  5204. 3:11:08representing audio images as well as
  5205. 3:11:10text and interpers them and handle them
  5206. 3:11:12all simultaneously in a single model so
  5207. 3:11:14that's one example of multimodality
  5208. 3:11:17uh second something that people are very
  5209. 3:11:18interested in
  5210. 3:11:20is currently most of the work is that
  5211. 3:11:22we're handing individual tasks to the
  5212. 3:11:24models on kind of like a silver platter
  5213. 3:11:26like please solve this task for me and
  5214. 3:11:28the model sort of like does this little
  5215. 3:11:29task but it's up to us to still sort of
  5216. 3:11:32like organize a coherent execution of
  5217. 3:11:35tasks to perform jobs and the models are
  5218. 3:11:38not yet at the capability required to do
  5219. 3:11:41this in a coherent error correcting way
  5220. 3:11:43over long periods of time so they're not
  5221. 3:11:46able to fully string together tasks to
  5222. 3:11:48perform these longer running jobs but
  5223. 3:11:51they're getting there and this is
  5224. 3:11:52improving uh over time but uh probably
  5225. 3:11:55what's going to happen here is we're
  5226. 3:11:56going to start to see what's called
  5227. 3:11:57agents which perform tasks over time and
  5228. 3:12:00you you supervise them and you watch
  5229. 3:12:02their work and they come up to once in a
  5230. 3:12:04while report progress and so on so we're
  5231. 3:12:07going to see more long running agents uh
  5232. 3:12:09tasks that don't just take you know a
  5233. 3:12:11few seconds of response but many tens of
  5234. 3:12:13seconds or even minutes or hours over
  5235. 3:12:15time uh but these uh models are not
  5236. 3:12:17infallible as we talked about above so
  5237. 3:12:19all of this will require supervision so
  5238. 3:12:21for example in factories people talk
  5239. 3:12:23about the human to robot ratio uh for
  5240. 3:12:26automation I think we're going to see
  5241. 3:12:27something similar in the digital space
  5242. 3:12:29where we are going to be talking about
  5243. 3:12:31human to agent ratios where humans
  5244. 3:12:33becomes a lot more supervisors of agent
  5245. 3:12:35tasks um in the digital
  5246. 3:12:38domain uh next um I think everything is
  5247. 3:12:41going to become a lot more pervasive and
  5248. 3:12:42invisible so it's kind of like
  5249. 3:12:44integrated into the tools and everywhere
  5250. 3:12:48um and in addition kind of like computer
  5251. 3:12:51using so right now these models aren't
  5252. 3:12:53able to take actions on your behalf but
  5253. 3:12:56I think this is a separate bullet point
  5254. 3:12:58um if you saw chpt launch the operator
  5255. 3:13:02then uh that's one early example of that
  5256. 3:13:04where you can actually hand off control
  5257. 3:13:05to the model to perform you know
  5258. 3:13:07keyboard and mouse actions on your
  5259. 3:13:09behalf so that's also something that
  5260. 3:13:11that I think is very interesting the
  5261. 3:13:13last point I have here is just a general
  5262. 3:13:14comment that there's still a lot of
  5263. 3:13:15research to potentially do in this
  5264. 3:13:16domain main one example of that uh is
  5265. 3:13:19something along the lines of test time
  5266. 3:13:20training so remember that everything
  5267. 3:13:22we've done above and that we talked
  5268. 3:13:24about has two major stages there's first
  5269. 3:13:27the training stage where we tune the
  5270. 3:13:28parameters of the model to perform the
  5271. 3:13:30tasks well once we get the parameters we
  5272. 3:13:33fix them and then we deploy the model
  5273. 3:13:34for inference from there the model is
  5274. 3:13:37fixed it doesn't change anymore it
  5275. 3:13:39doesn't learn from all the stuff that
  5276. 3:13:41it's doing a test time it's a fixed um
  5277. 3:13:43number of parameters and the only thing
  5278. 3:13:45that is changing is now the token inside
  5279. 3:13:47the context windows and so the only type
  5280. 3:13:49of learning or test time learning that
  5281. 3:13:51the model has access to is the in
  5282. 3:13:53context learning of its uh kind of like
  5283. 3:13:56uh dynamically adjustable context window
  5284. 3:13:59depending on like what it's doing at
  5285. 3:14:00test time so but I think this is still
  5286. 3:14:03different from humans who actually are
  5287. 3:14:04able to like actually learn uh depending
  5288. 3:14:06on what they're doing especially when
  5289. 3:14:08you sleep for example like your brain is
  5290. 3:14:09updating your parameters or something
  5291. 3:14:10like that right so there's no kind of
  5292. 3:14:13equivalent of that currently in these
  5293. 3:14:14models and tools so there's a lot of
  5294. 3:14:16like um more wonky ideas I think that
  5295. 3:14:18are to be explored still and uh in
  5296. 3:14:20particular I think this will be
  5297. 3:14:21necessary because the context window is
  5298. 3:14:24a finite and precious resource and
  5299. 3:14:26especially once we start to tackle very
  5300. 3:14:27long running multimodal tasks and we're
  5301. 3:14:30putting in videos and these token
  5302. 3:14:31windows will basically start to grow
  5303. 3:14:34extremely large like not thousands or
  5304. 3:14:36even hundreds of thousands but
  5305. 3:14:37significantly beyond that and the only
  5306. 3:14:39trick uh the only kind of trick we have
  5307. 3:14:41Avail to us right now is to make the
  5308. 3:14:43context Windows longer but I think that
  5309. 3:14:46that approach by itself will will not
  5310. 3:14:47will not scale to actual long running
  5311. 3:14:49tasks that are multimodal over time and
  5312. 3:14:51so I think new ideas are needed in some
  5313. 3:14:53of those disciplines um in some of those
  5314. 3:14:56kind of cases in the main where these
  5315. 3:14:58tasks are going to require very long
  5316. 3:15:00contexts so those are some examples of
  5317. 3:15:03some of the things you can um expect
  5318. 3:15:05coming down the pipe let's now turn to
  5319. 3:15:07where you can actually uh kind of keep
  5320. 3:15:09track of this progress and um you know
  5321. 3:15:12be up to date with the latest and grest
  5322. 3:15:13of what's happening in the field so I
  5323. 3:15:15would say the three resources that I
  5324. 3:15:16have consistently used to stay up to
  5325. 3:15:18date are number one El Marina uh so let
  5326. 3:15:21me show you El
  5327. 3:15:23Marina this is basically an llm leader
  5328. 3:15:26board and it ranks all the top models
  5329. 3:15:30and the ranking is based on human
  5330. 3:15:32comparisons so humans prompt these
  5331. 3:15:34models and they get to judge which one
  5332. 3:15:35gives a better answer they don't know
  5333. 3:15:37which model is which they're just
  5334. 3:15:39looking at which model is the better
  5335. 3:15:40answer and you can calculate a ranking
  5336. 3:15:42and then you get some results and so
  5337. 3:15:44what you can hear is what you can see
  5338. 3:15:46here is the different organizations like
  5339. 3:15:48Google Gemini for example that produce
  5340. 3:15:49these models when you click on any one
  5341. 3:15:51of these it takes you to the place where
  5342. 3:15:53that model is
  5343. 3:15:55hosted and then here we see Google is
  5344. 3:15:57currently on top with open AI right
  5345. 3:15:59behind here we see deep seek in position
  5346. 3:16:02number three now the reason this is a
  5347. 3:16:04big deal is the last column here you see
  5348. 3:16:05license deep seek is an MIT license
  5349. 3:16:08model it's open weights anyone can use
  5350. 3:16:10these weights uh anyone can download
  5351. 3:16:12them anyone can host their own version
  5352. 3:16:14of Deep seek and they can use it in what
  5353. 3:16:16whatever way they like and so it's not a
  5354. 3:16:18proprietary model that you don't have
  5355. 3:16:19access to it's it's basically an open
  5356. 3:16:21weight release and so this is kind of
  5357. 3:16:24unprecedented that a model this strong
  5358. 3:16:27was released with open weights so pretty
  5359. 3:16:29cool from the team next up we have a few
  5360. 3:16:32more models from Google and open Ai and
  5361. 3:16:34then when you continue to scroll down
  5362. 3:16:35you start to see some other Usual
  5363. 3:16:36Suspects so xai here anthropic with son
  5364. 3:16:40it uh here at number
  5365. 3:16:4314 and
  5366. 3:16:45um then
  5367. 3:16:47meta with llama over here so llama
  5368. 3:16:51similar to deep seek is an open weights
  5369. 3:16:52model and so uh but it's down here as
  5370. 3:16:55opposed to up here now I will say that
  5371. 3:16:57this leaderboard was really good for a
  5372. 3:17:00long time I do think that in the last
  5373. 3:17:03few months it's become a little bit
  5374. 3:17:05gamed um and I don't trust it as much as
  5375. 3:17:08I used to I think um just empirically I
  5376. 3:17:11feel like a lot of people for example
  5377. 3:17:13are using a Sonet from anthropic and
  5378. 3:17:15that it's a really good model so but
  5379. 3:17:17that's all the way down here um in
  5380. 3:17:19number 14 and conversely I think not as
  5381. 3:17:22many people are using Gemini but it's
  5382. 3:17:23racking really really high uh so I think
  5383. 3:17:27use this as a first pass uh but uh sort
  5384. 3:17:30of try out a few of the models for your
  5385. 3:17:32tasks and see which one performs better
  5386. 3:17:35the second thing that I would point to
  5387. 3:17:37is the uh AI news uh newsletter so AI
  5388. 3:17:41news is not very creatively named but it
  5389. 3:17:43is a very good newsletter produced by
  5390. 3:17:44swix and friends so thank you for
  5391. 3:17:46maintaining it
  5392. 3:17:47and it's been very helpful to me because
  5393. 3:17:48it is extremely comprehensive so if you
  5394. 3:17:50go to archives uh you see that it's
  5395. 3:17:52produced almost every other day and um
  5396. 3:17:56it is very comprehensive and some of it
  5397. 3:17:58is written by humans and curated by
  5398. 3:17:59humans but a lot of it is constructed
  5399. 3:18:01automatically with llms so you'll see
  5400. 3:18:03that these are very comprehensive and
  5401. 3:18:04you're probably not missing anything
  5402. 3:18:06major if you go through it of course
  5403. 3:18:08you're probably not going to go through
  5404. 3:18:09it because it's so long but I do think
  5405. 3:18:12that these summaries all the way up top
  5406. 3:18:14are quite good and I think have some
  5407. 3:18:15human oversight uh so this has been very
  5408. 3:18:18helpful to me and the last thing I would
  5409. 3:18:20point to is just X and Twitter uh a lot
  5410. 3:18:22of um AI happens on X and so I would
  5411. 3:18:25just follow people who you like and
  5412. 3:18:27trust and get all your latest and
  5413. 3:18:29greatest uh on X as well so those are
  5414. 3:18:32the major places that have worked for me
  5415. 3:18:33over time and finally a few words on
  5416. 3:18:35where you can find the models and where
  5417. 3:18:37can you use them so the first one I
  5418. 3:18:39would say is for any of the biggest
  5419. 3:18:41proprietary models you just have to go
  5420. 3:18:42to the website of that LM provider so
  5421. 3:18:44for example for open a that's uh chat
  5422. 3:18:47I believe actually works now uh so
  5423. 3:18:49that's for open
  5424. 3:18:50AI now for or you know for um for Gemini
  5425. 3:18:54I think it's gem. google.com or AI
  5426. 3:18:57Studio I think they have two for some
  5427. 3:18:59reason that I don't fly understand no
  5428. 3:19:01one does um for the open weights models
  5429. 3:19:04like deep SE CL Etc you have to go to
  5430. 3:19:06some kind of an inference provider of
  5431. 3:19:08LMS so my favorite one is together
  5432. 3:19:10together. a and I showed you that when
  5433. 3:19:11you go to the playground of together. a
  5434. 3:19:14then you can sort of pick lots of
  5435. 3:19:15different models and all of these are
  5436. 3:19:17open models of different types and you
  5437. 3:19:19can talk to them here as an
  5438. 3:19:21example um now if you'd like to use a
  5439. 3:19:24base model like um you know a base model
  5440. 3:19:28then this is where I think it's not as
  5441. 3:19:29common to find base models even on these
  5442. 3:19:31inference providers they are all
  5443. 3:19:32targeting assistants and chat and so I
  5444. 3:19:35think even here I can't I couldn't see
  5445. 3:19:37base models here so for base models I
  5446. 3:19:39usually go to hyperbolic because they
  5447. 3:19:41serve my llama 3.1 base and I love that
  5448. 3:19:45model and you can just talk to it here
  5449. 3:19:47so as far as I know this is this is a
  5450. 3:19:49good place for a base model and I wish
  5451. 3:19:51more people hosted base models because
  5452. 3:19:53they are useful and interesting to work
  5453. 3:19:54with in some cases finally you can also
  5454. 3:19:57take some of the models that are smaller
  5455. 3:19:59and you can run them locally and so for
  5456. 3:20:02example deep seek the biggest model
  5457. 3:20:04you're not going to be able to run
  5458. 3:20:05locally on your MacBook but there are
  5459. 3:20:07smaller versions of the deep seek model
  5460. 3:20:09that are what's called distilled and
  5461. 3:20:11then also you can run these models at
  5462. 3:20:12smaller Precision so not at the native
  5463. 3:20:14Precision of for example fp8 on deep
  5464. 3:20:17seek or you know bf16 llama but much
  5465. 3:20:20much lower than that um and don't worry
  5466. 3:20:23if you don't fully understand those
  5467. 3:20:24details but you can run smaller versions
  5468. 3:20:26that have been distilled and then at
  5469. 3:20:28even lower precision and then you can
  5470. 3:20:29fit them on your uh computer and so you
  5471. 3:20:33can actually run pretty okay models on
  5472. 3:20:35your laptop and my favorite I think
  5473. 3:20:37place I go to usually is LM studio uh
  5474. 3:20:39which is basically an app you can get
  5475. 3:20:42and I think it kind of actually looks
  5476. 3:20:43really ugly and it's I don't like that
  5477. 3:20:45it shows you all these models that are
  5478. 3:20:46basically not that useful like everyone
  5479. 3:20:48just wants to run deep seek so I don't
  5480. 3:20:49know why they give you these 500
  5481. 3:20:51different types of models they're really
  5482. 3:20:53complicated to search for and you have
  5483. 3:20:54to choose different distillations and
  5484. 3:20:56different uh precisions and it's all
  5485. 3:20:58really confusing but once you actually
  5486. 3:21:00understand how it works and that's a
  5487. 3:21:01whole separate video then you can
  5488. 3:21:02actually load up a model like here I
  5489. 3:21:04loaded up a llama 3 uh2 instruct 1
  5490. 3:21:08billion and um you can just talk to it
  5491. 3:21:11so I ask for Pelican jokes and I can ask
  5492. 3:21:14for another one and it gives me another
  5493. 3:21:15one Etc all of this that happens here is
  5494. 3:21:18locally on your computer so we're not
  5495. 3:21:20actually going to anywhere anyone else
  5496. 3:21:22this is running on the GPU on the
  5497. 3:21:24MacBook Pro so that's very nice and you
  5498. 3:21:26can then eject the model when you're
  5499. 3:21:28done and that frees up the ram so LM
  5500. 3:21:31studio is probably like my favorite one
  5501. 3:21:33even though I don't I think it's got a
  5502. 3:21:34lot of uiux issues and it's really
  5503. 3:21:36geared towards uh professionals almost
  5504. 3:21:39uh but if you watch some videos on
  5505. 3:21:40YouTube I think you can figure out how
  5506. 3:21:41to how to use this
  5507. 3:21:43interface uh so those are a few words on
  5508. 3:21:45where to find them so let me now loop
  5509. 3:21:47back around to where we started the
  5510. 3:21:49question was when we go to chashi
  5511. 3:21:50pta.com and we enter some kind of a
  5512. 3:21:53query and we hit go what exactly is
  5513. 3:21:57happening here what are we seeing what
  5514. 3:21:59are we talking to how does this work and
  5515. 3:22:03I hope that this video gave you some
  5516. 3:22:04appreciation for some of the under the
  5517. 3:22:06hood details of how these models are
  5518. 3:22:08trained and what this is that is coming
  5519. 3:22:10back so in particular we now know that
  5520. 3:22:12your query is taken and is first chopped
  5521. 3:22:15up into tokens so we go to to tick
  5522. 3:22:18tokenizer and here where is the place in
  5523. 3:22:21the in the um sort of format that is for
  5524. 3:22:24the user query we basically put in our
  5525. 3:22:27query right there so our query goes into
  5526. 3:22:31what we discussed here is the
  5527. 3:22:32conversation protocol format which is
  5528. 3:22:34this way that we maintain conversation
  5529. 3:22:36objects so this gets inserted there and
  5530. 3:22:39then this whole thing ends up being just
  5531. 3:22:40a token sequence a onedimensional token
  5532. 3:22:43sequence under the hood so Chachi PT saw
  5533. 3:22:46this token sequence and then when we hit
  5534. 3:22:48go it basically continues appending
  5535. 3:22:50tokens into this list it continues the
  5536. 3:22:53sequence it acts like a token
  5537. 3:22:55autocomplete so in particular it gave us
  5538. 3:22:57this response so we can basically just
  5539. 3:23:00put it here and we see the tokens that
  5540. 3:23:02it continued uh these are the tokens
  5541. 3:23:04that it continued with
  5542. 3:23:06roughly now the question
  5543. 3:23:08becomes okay why are these the tokens
  5544. 3:23:10that the model responded with what are
  5545. 3:23:12these tokens where are they coming from
  5546. 3:23:14uh what are we talking to and how do we
  5547. 3:23:17program this system and so that's where
  5548. 3:23:19we shifted gears and we talked about the
  5549. 3:23:21under thehood pieces of it so the first
  5550. 3:23:24stage of this process and there are
  5551. 3:23:25three stages is the pre-training stage
  5552. 3:23:27which fundamentally has to do with just
  5553. 3:23:28knowledge acquisition from the internet
  5554. 3:23:30into the parameters of this neural
  5555. 3:23:32network and so the neural net
  5556. 3:23:35internalizes a lot of Knowledge from the
  5557. 3:23:37internet but where the personality
  5558. 3:23:39really comes in is in the process of
  5559. 3:23:41supervised fine-tuning here and so what
  5560. 3:23:44what happens here is that basically the
  5561. 3:23:46a company like openai will curate a
  5562. 3:23:49large data set of conversations like say
  5563. 3:23:511 million conversation across very
  5564. 3:23:53diverse topics and there will be
  5565. 3:23:55conversations between a human and an
  5566. 3:23:57assistant and even though there's a lot
  5567. 3:23:59of synthetic data generation used
  5568. 3:24:01throughout this entire process and a lot
  5569. 3:24:02of llm help and so on fundamentally this
  5570. 3:24:05is a human data curation task with lots
  5571. 3:24:08of humans involved and in particular
  5572. 3:24:10these humans are data labelers hired by
  5573. 3:24:12open AI who are given labeling
  5574. 3:24:14instructions that they learn and they
  5575. 3:24:16task is to create ideal assistant
  5576. 3:24:18responses for any arbitrary prompts so
  5577. 3:24:21they are teaching the neural network by
  5578. 3:24:24example how to respond to
  5579. 3:24:27prompts so what is the way to think
  5580. 3:24:29about what came back here like what is
  5581. 3:24:32this well I think the right way to think
  5582. 3:24:34about it is that this is the neural
  5583. 3:24:37network simulation of a data labeler at
  5584. 3:24:40openai so it's as if I gave this query
  5585. 3:24:44to a data Li open and this data labeler
  5586. 3:24:47first reads all of the labeling
  5587. 3:24:48instructions from open Ai and then
  5588. 3:24:51spends 2 hours writing up the ideal
  5589. 3:24:53assistant response to this query and uh
  5590. 3:24:57giving it to me now we're not actually
  5591. 3:24:59doing that right because we didn't wait
  5592. 3:25:01two hours so what we're getting here is
  5593. 3:25:02a neural network simulation of that
  5594. 3:25:05process and we have to keep in mind that
  5595. 3:25:08these neural networks don't function
  5596. 3:25:10like human brains do they are different
  5597. 3:25:12what's easy or hard for them is
  5598. 3:25:13different from what's easy or hard for
  5599. 3:25:15humans and so we really are just getting
  5600. 3:25:17a simulation so here I shown you this is
  5601. 3:25:20a token stream and this is fundamentally
  5602. 3:25:23the neural network with a bunch of
  5603. 3:25:24activations and neurons in between this
  5604. 3:25:26is a fixed mathematical expression that
  5605. 3:25:28mixes inputs from tokens with parameters
  5606. 3:25:32of the model and they get mixed up and
  5607. 3:25:35get you the next token in a sequence but
  5608. 3:25:37this is a finite amount of compute that
  5609. 3:25:39happens for every single token and so
  5610. 3:25:41this is some kind of a lossy simulation
  5611. 3:25:44of a human that is kind of like
  5612. 3:25:46restricted in this way and so whatever
  5613. 3:25:49the humans
  5614. 3:25:50write the language model is kind of
  5615. 3:25:52imitating on this token level with only
  5616. 3:25:55this this specific computation for every
  5617. 3:25:58single token and
  5618. 3:26:00sequence we also saw that as a result of
  5619. 3:26:03this and the cognitive differences the
  5620. 3:26:05models will suffer in a variety of ways
  5621. 3:26:08and uh you have to be very careful with
  5622. 3:26:10their use so for example we saw that
  5623. 3:26:11they will suffer from hallucinations and
  5624. 3:26:14they also we have the sense of a Swiss
  5625. 3:26:16model of the LM capabilities where
  5626. 3:26:18basically there's like holes in the
  5627. 3:26:20cheese sometimes the models will just
  5628. 3:26:22arbitrarily like do something dumb uh so
  5629. 3:26:25even though they're doing lots of
  5630. 3:26:26magical stuff sometimes they just can't
  5631. 3:26:28so maybe you're not giving them enough
  5632. 3:26:30tokens to think and maybe they're going
  5633. 3:26:32to just make stuff up because they're
  5634. 3:26:33mental arithmetic breaks uh maybe they
  5635. 3:26:35are suddenly unable to count number of
  5636. 3:26:38letters um or maybe they're unable to
  5637. 3:26:40tell you that 911 9.11 is smaller than
  5638. 3:26:439.9 and it looks kind of dumb and so so
  5639. 3:26:46it's a Swiss cheese capability and we
  5640. 3:26:48have to be careful with that and we saw
  5641. 3:26:49the reasons for
  5642. 3:26:50that but fundamentally this is how we
  5643. 3:26:53think of what came back it's again a
  5644. 3:26:56simulation of this neural network of a
  5645. 3:27:00human data labeler following the
  5646. 3:27:03labeling instructions at open a so
  5647. 3:27:06that's what we're getting back now I do
  5648. 3:27:09think that the uh things change a little
  5649. 3:27:11bit when you actually go and reach for
  5650. 3:27:13one of the thinking models like o03 mini
  5651. 3:27:17and the reason for that is that GPT
  5652. 3:27:2040 basically doesn't do reinforcement
  5653. 3:27:23learning it does do rhf but I've told
  5654. 3:27:26you that rhf is not RL there's no
  5655. 3:27:29there's no uh time for magic in there
  5656. 3:27:31it's just a little bit of a fine-tuning
  5657. 3:27:33is the way to look at it but these
  5658. 3:27:35thinking models they do use RL so they
  5659. 3:27:38go through this third state stage of
  5660. 3:27:41perfecting their thinking process and
  5661. 3:27:44discovering new thinking strategies and
  5662. 3:27:46uh
  5663. 3:27:46solutions to problem solving that look a
  5664. 3:27:49little bit like your internal monologue
  5665. 3:27:51in your head and they practice that on a
  5666. 3:27:53large collection of practice problems
  5667. 3:27:55that companies like openi create and
  5668. 3:27:57curate and um then make available to the
  5669. 3:28:00LMS so when I come here and I talked to
  5670. 3:28:02a thinking model and I put in this
  5671. 3:28:05question what we're seeing here is not
  5672. 3:28:07anymore just the straightforward
  5673. 3:28:09simulation of a human data labeler like
  5674. 3:28:11this is actually kind of new unique and
  5675. 3:28:14interesting um and of course open is not
  5676. 3:28:16showing us the under thehood thinking
  5677. 3:28:18and the chains of thought that are
  5678. 3:28:20underlying the reasoning here but we
  5679. 3:28:23know that such a thing exists and this
  5680. 3:28:24is a summary of it and what we're
  5681. 3:28:26getting here is actually not just an
  5682. 3:28:27imitation of a human data labeler it's
  5683. 3:28:29actually something that is kind of new
  5684. 3:28:30and interesting and exciting in the
  5685. 3:28:32sense that it is a function of thinking
  5686. 3:28:35that was emergent in a simulation it's
  5687. 3:28:37not just imitating human data labeler it
  5688. 3:28:39comes from this reinforcement learning
  5689. 3:28:41process and so here we're of course not
  5690. 3:28:43giving it a chance to shine because this
  5691. 3:28:45is not a mathematical or a reasoning
  5692. 3:28:46problem this is just some kind of a sort
  5693. 3:28:48of creative writing problem roughly
  5694. 3:28:50speaking and I think it's um it's a a
  5695. 3:28:54question an open question as to whether
  5696. 3:28:57the thinking strategies that are
  5697. 3:28:59developed inside verifiable domains
  5698. 3:29:02transfer and are generalizable to other
  5699. 3:29:05domains that are unverifiable such as
  5700. 3:29:07create writing the extent to which that
  5701. 3:29:09transfer happens is unknown in the field
  5702. 3:29:12I would say so we're not sure if we are
  5703. 3:29:14able to do RL on everything that is very
  5704. 3:29:16verifiable and see the benefits of that
  5705. 3:29:18on things that are unverifiable like
  5706. 3:29:20this prompt so that's an open question
  5707. 3:29:22the other thing that's interesting is
  5708. 3:29:23that this reinforcement learning here is
  5709. 3:29:26still like way too new primordial and
  5710. 3:29:29nent so we're just seeing like the
  5711. 3:29:31beginnings of the hints of greatness uh
  5712. 3:29:34in the reasoning problems we're seeing
  5713. 3:29:36something that is in principle capable
  5714. 3:29:38of something like the equivalent of move
  5715. 3:29:4037 but not in the game of Go but in open
  5716. 3:29:44domain thinking and problem solving in
  5717. 3:29:46principle this Paradigm is capable of
  5718. 3:29:48doing something really cool new and
  5719. 3:29:50exciting something even that no human
  5720. 3:29:52has thought of before in principle these
  5721. 3:29:54models are capable of analogies no human
  5722. 3:29:56has had so I think it's incredibly
  5723. 3:29:58exciting that these models exist but
  5724. 3:30:00again it's very early and these are
  5725. 3:30:02primordial models for now um and they
  5726. 3:30:05will mostly shine in domains that are
  5727. 3:30:06verifiable like math en code Etc so very
  5728. 3:30:10interesting to play with and think about
  5729. 3:30:11and
  5730. 3:30:12use and then that's roughly it um um I
  5731. 3:30:16would say those are the broad Strokes of
  5732. 3:30:18what's available right now I will say
  5733. 3:30:20that overall it is an extremely exciting
  5734. 3:30:23time to be in the
  5735. 3:30:24field personally I use these models all
  5736. 3:30:26the time daily uh tens or hundreds of
  5737. 3:30:28times because they dramatically
  5738. 3:30:30accelerate my work I think a lot of
  5739. 3:30:31people see the same thing I think we're
  5740. 3:30:33going to see a huge amount of wealth
  5741. 3:30:34creation as a result of these models be
  5742. 3:30:37aware of some of their shortcomings even
  5743. 3:30:40with RL models they're going to suffer
  5744. 3:30:42from some of these use it as a tool in a
  5745. 3:30:44toolbox don't trust it fully because
  5746. 3:30:47they will randomly do dumb things they
  5747. 3:30:49will randomly hallucinate they will
  5748. 3:30:51randomly skip over some mental
  5749. 3:30:52arithmetic and not get it right um they
  5750. 3:30:55randomly can't count or something like
  5751. 3:30:56that so use them as tools in the toolbox
  5752. 3:30:58check their work and own the product of
  5753. 3:31:00your work but use them for inspiration
  5754. 3:31:03for first draft uh ask them questions
  5755. 3:31:06but always check and verify and you will
  5756. 3:31:08be very successful in your work if you
  5757. 3:31:10do so uh so I hope this video was useful
  5758. 3:31:13and interesting to you I hope you had it
  5759. 3:31:15fun and uh it's already like very long
  5760. 3:31:17so I apologize for that but I hope it
  5761. 3:31:19was useful and yeah I will see you later

About this transcript

This page contains the full transcript of Deep Dive into LLMs like ChatGPT by Andrej Karpathy, generated from the public captions YouTube serves with the video. The transcript has 41,116 words across 5,761 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.