YouTube2Text

Yann LeCun on What Comes After LLMs — Transcript

by Unsupervised Learning: With Jacob Effron · 15,072 words · 2,582 segments · language en · Watch on YouTube

Full transcript

  1. 0:00You're one of the godfathers of AI.
  2. 0:01What's your kind of view of the path of
  3. 0:03progress here? Five years complete world
  4. 0:04domination. The best way to get
  5. 0:06breakthrough research is you hire the
  6. 0:08best people and you get the [ __ ] out of
  7. 0:09the way.
  8. 0:10>> Pardon my French.
  9. 0:11>> You shared the Turing Award with two
  10. 0:12others. When did your views start
  11. 0:13diverging? In 2023. How do you know it
  12. 0:15was time to leave Meta? It sounds like
  13. 0:17you were thinking through some of these
  14. 0:17things over a period of time.
  15. 0:19>> There is a big misconception about my
  16. 0:21role, my relation to AI and how AI was
  17. 0:24run at Meta. What's like one thing
  18. 0:25you've changed your mind on in the last
  19. 0:26year? I mean, the whole idea of uh Yann
  20. 0:29LeCun is one of the godfathers of AI.
  21. 0:31He's an absolute legend in the field, uh
  22. 0:33someone I've admired for a long time.
  23. 0:35And so it was such a treat to get him on
  24. 0:37on Unsupervised Learning.
  25. 0:39Uh he's been a noted skeptic of of LLMs
  26. 0:41in many ways, and so we dug into what
  27. 0:43LLMs can do, what they can't do, uh some
  28. 0:45of the limitations he sees, and why he
  29. 0:47ultimately decided to pursue a different
  30. 0:49architecture. Uh and we also talked
  31. 0:51about his time at Meta,
  32. 0:52um you know, the things he's proud of in
  33. 0:53in setting up FAIR, how the last few
  34. 0:55years proceeded, and what ultimately led
  35. 0:57him to uh spin out and start his own
  36. 0:59company, uh me. Um I think it's just
  37. 1:01fascinating to get Yann's thoughts on
  38. 1:03everything happening in the AI ecosystem
  39. 1:05today, this tension between basic
  40. 1:07research and then pushing LLMs forward,
  41. 1:09and how that's happening in in a bunch
  42. 1:11of organizations today, as well as his
  43. 1:12thoughts on just where the the whole
  44. 1:14space is headed. Uh he's just an
  45. 1:16absolute giant in the field, and when I
  46. 1:18started this podcast, I hoped we'd get
  47. 1:19guests like him, so it is just such a
  48. 1:21treat. I think folks will really enjoy
  49. 1:23hearing the conversation we had. Without
  50. 1:24further ado, here's Yann.
  51. 1:29Yann, this is such a pleasure. You're
  52. 1:30one of the godfathers of AI. I feel like
  53. 1:32when I started doing this podcast years
  54. 1:34ago, I was really hoping we might one
  55. 1:36day get someone like you on. You know, I
  56. 1:38don't like that term because I live in
  57. 1:39New Jersey, when you're a godfather in
  58. 1:40New Jersey, it [laughter] doesn't mean
  59. 1:42the same thing.
  60. 1:43Very fair, very fair. You know,
  61. 1:45obviously, you know, your bet on on
  62. 1:46neural nets when everyone doubted them
  63. 1:47is legendary, and I feel like today
  64. 1:49you're making uh a similar bet in many
  65. 1:51ways against LLMs and the kind of
  66. 1:52predominant generative architectures
  67. 1:54that that so many believe in.
  68. 1:55Uh you've recently started a new company
  69. 1:57uh behind this theme. And so, you know,
  70. 1:59our goal today in the conversation is to
  71. 2:01leave our listeners with a lot for more
  72. 2:03information about AMI, what you're doing
  73. 2:04there, some of your work at Tapestry,
  74. 2:06um you know, why you think the rest of
  75. 2:08the field is is is pointed in the wrong
  76. 2:09direction around some of these
  77. 2:10generative models, and then also just
  78. 2:12get your reflections on the way the
  79. 2:13field's unfolded, your time at Meta, and
  80. 2:15and all that. So, you know, modest goals
  81. 2:16for uh for for for a single podcast
  82. 2:18episode. I feel it'd be great to start
  83. 2:20with AMI um because the company feels
  84. 2:22like the clearest statement of your
  85. 2:23technical thesis going forward. And so,
  86. 2:25you recently launched the company that's
  87. 2:26focused on world models uh and scaling
  88. 2:28the Jeff bar architecture, which you
  89. 2:29obviously pioneered uh over at Meta. And
  90. 2:32so, I'm wondering if you could talk a
  91. 2:32little bit about the origins of that
  92. 2:34architecture and the extent to which you
  93. 2:35drew inspiration from the human brain
  94. 2:37and the way that works. So, first of
  95. 2:38all, I want to say there's nothing wrong
  96. 2:40with
  97. 2:41LLMs
  98. 2:42in the sense of
  99. 2:44LLMs, you know, are the basis for a lot
  100. 2:46of uh
  101. 2:47very useful AI products that all of us
  102. 2:49use,
  103. 2:50including me. Uh
  104. 2:52They're great, okay, for what they do.
  105. 2:55They're just not a path towards
  106. 2:57human-level or human-like intelligence
  107. 3:00or even animal-like intelligence.
  108. 3:02Uh so, that's my claim, okay? I'm not
  109. 3:04saying LLMs are useless, right? I'm I'm
  110. 3:06just saying
  111. 3:07they're not a path towards I mean, you
  112. 3:09helped build some of the first major
  113. 3:10open-source ones.
  114. 3:11Right. Right.
  115. 3:12>> [laughter]
  116. 3:12>> Right, absolutely. So, what is uh AMI?
  117. 3:15So, AMI really stands for advanced
  118. 3:17machine intelligence. And the the the
  119. 3:21kind of subtitle, the motto, if you
  120. 3:22want, is uh AI for the real world.
  121. 3:26So, basically, a lot of
  122. 3:28you know, AI techniques that people know
  123. 3:30about today are good for language
  124. 3:32manipulation, either human language or
  125. 3:35computer code or mathematics or or
  126. 3:38legalese,
  127. 3:40which barely qualifies as human
  128. 3:42language.
  129. 3:43>> [laughter]
  130. 3:43>> Unfortunately, a lot of human language
  131. 3:44is for it. Right. Right, sadly.
  132. 3:47You know, language is very special in a
  133. 3:49way, and it's uh particularly
  134. 3:52well suited for the type of uh you know,
  135. 3:55architectures that
  136. 3:56have been so successful uh recently, the
  137. 3:59the you know, large language models,
  138. 4:01GPT-style architectures. But what about
  139. 4:04the real world? What about like
  140. 4:05understanding
  141. 4:07the physical world? Turns out reality is
  142. 4:09way more complicated than language.
  143. 4:12Uh because it's high-dimensional, it's
  144. 4:14continuous, it's noisy, it's messy.
  145. 4:18And uh training a system to understand
  146. 4:20the real world is much, much harder. So
  147. 4:21that's really what we're after. That's
  148. 4:23what I've been after for most of my
  149. 4:25career and really kind of,
  150. 4:27you know, working on in an accelerated
  151. 4:29fashion over the last uh 5, 6 years or
  152. 4:31so and making significant progress over
  153. 4:33the last 2 years.
  154. 4:35And so, it made sense to really do a
  155. 4:37startup around it and sort of go to into
  156. 4:40high gear, you know, in pushing that.
  157. 4:42And it became clear, you know, by the
  158. 4:44end of last year that
  159. 4:45Meta was really not the right place for
  160. 4:47that.
  161. 4:48So, which is why I left and started I
  162. 4:51mean, Labs. I think it's an interesting
  163. 4:53like, you know, trend that we're seeing
  164. 4:54across the board, right? Where it feels
  165. 4:56like um there you're there's there's
  166. 4:58many folks spinning out of, you know,
  167. 4:59either some of the large companies or
  168. 5:01research labs, you know, that have a
  169. 5:03particular direction of research they're
  170. 5:04excited about and
  171. 5:05you'd have some interesting vantage
  172. 5:06point of this from your time at Fair of
  173. 5:07this
  174. 5:08uh almost tension that exists between,
  175. 5:10you know, go pursue as many different
  176. 5:11research directions as possible in these
  177. 5:13companies versus hey, something's really
  178. 5:15working. This is the thing that we're
  179. 5:16going to sell for the next 6, 12 months.
  180. 5:18Like, go focus on that. You know, I'm
  181. 5:20curious your your thoughts on that and
  182. 5:21and what you kind of seen in the
  183. 5:22industry at large. Well, it's a strange
  184. 5:25uh
  185. 5:26trade-off. There's really two modes of
  186. 5:28operating, right? There's a lot of
  187. 5:29exploratory research, a lot of research
  188. 5:31directions, right? And sometimes
  189. 5:33something
  190. 5:34kind of seems to work and you you need
  191. 5:36to push it further. And it's not
  192. 5:38research anymore. I mean,
  193. 5:40the people working on it are still
  194. 5:41researchers or they're called
  195. 5:43researchers at least in the press, but
  196. 5:45uh but really it's becoming more
  197. 5:46engineering and pushing for for
  198. 5:48products, right? So,
  199. 5:51that happened a number of times at Meta
  200. 5:54because of things that were started at
  201. 5:56FAIR, such as seeing happened in, you
  202. 5:59know, early 2023, essentially. Uh, when,
  203. 6:03you know, Llama, which was developed at
  204. 6:05FAIR, Llama 1, um, was very promising.
  205. 6:09And, uh,
  206. 6:10Meta created a a whole organization,
  207. 6:12GenAI, to turn it into something real
  208. 6:15and a series of products. Uh, and
  209. 6:17produce, you know, Llama 2, Llama 3,
  210. 6:19Llama 4, which was a bit of a
  211. 6:21disappointment. Uh, and because, you
  212. 6:23know, Mark Zuckerberg was disappointed
  213. 6:25by it, he kind of rebooted the entire
  214. 6:27organization, reorganized it, and hired
  215. 6:30new people, etc. But, what also
  216. 6:32happened,
  217. 6:34uh, over the last year is that
  218. 6:36uh,
  219. 6:37basically the company, Meta, realized
  220. 6:39that,
  221. 6:40um, they'd fallen behind a little bit,
  222. 6:42and so that kind of refocused the the
  223. 6:45strategy on trying to catch up with the
  224. 6:47industry. And the sad side effect of it
  225. 6:51is that a lot of the exploratory
  226. 6:53research
  227. 6:54was basically not
  228. 6:56given high priority anymore. I mean, it
  229. 6:58didn't concern the stuff I was working
  230. 7:00on, all the Jeppa and world models,
  231. 7:02uh, cuz, you know, Mark himself and and
  232. 7:05do buzzwords, the CTO, and a bunch of
  233. 7:07other people in the company were really
  234. 7:08interested in that project and really
  235. 7:09believed in the long-term impact. But,
  236. 7:12the rest of the company was just, you
  237. 7:13know, totally entirely focused on LLMs,
  238. 7:16and made it clear to me that Meta was
  239. 7:18really not the the right place to push
  240. 7:20on that project anymore. And then we
  241. 7:22started to have good results, and so it
  242. 7:24was clear that, you know, we had to kind
  243. 7:26of make that transition between research
  244. 7:29and actually kind of uh, developing the
  245. 7:32technology, scaling it up, and building
  246. 7:33products out of it. And we realized also
  247. 7:35that most of the
  248. 7:37applications were probably
  249. 7:40for things that Meta was not
  250. 7:41particularly interested in. A lot of
  251. 7:44applications of the kind of stuff that
  252. 7:46we've been working on is in the
  253. 7:48industry, like manufacturing industry
  254. 7:50and stuff like that. Obviously, you're
  255. 7:52you're kind of pursuing world models and
  256. 7:54and and in that broader world. And I
  257. 7:55think there's other people that have
  258. 7:56come at the world model pace from a more
  259. 7:58like generative approach. And so I think
  260. 8:00you've got folks, you know, you've got
  261. 8:01the Google folks and Genie in the video
  262. 8:03models. You've got folks, you know,
  263. 8:04building VLAs on the robotic side.
  264. 8:05You've got Feifei and and kind of like
  265. 8:08the 3D spatial models. As you think
  266. 8:10about kind of the the body of of of of
  267. 8:12evidence that got you excited about the
  268. 8:14Japa models and how you kind of compare
  269. 8:15them to what the generative folks have
  270. 8:17done, you know, where do you think we
  271. 8:19are today in in terms of like comparing
  272. 8:20these architectures and approaches?
  273. 8:22Okay, so world model is quickly becoming
  274. 8:24a buzzword right now, right? Certainly
  275. 8:27in research, but also in industry to
  276. 8:28some extent. And uh and then there are
  277. 8:31two factions, if you want. I'm not going
  278. 8:33to talk about VLA because VLA
  279. 8:35is clearly now being seen as not going
  280. 8:38anywhere.
  281. 8:40Like it's really not working. Uh so VLA
  282. 8:42is, you know,
  283. 8:43vision language action models, right? So
  284. 8:45basically use the LLM technology to
  285. 8:48train a system to produce actions for
  286. 8:50like controlling a robot or something
  287. 8:52like this, right? So you have vision in,
  288. 8:54language in, action out. Maybe language
  289. 8:57out, too.
  290. 8:58Um and that's pretty much now seen as a
  291. 9:02failure.
  292. 9:03>> [laughter]
  293. 9:04>> Uh not being reliable enough, requiring
  294. 9:06too much training data, you know, things
  295. 9:07like that.
  296. 9:08Okay, then there is world models. Okay,
  297. 9:10so what is a world model? Uh a world
  298. 9:12model at a regional level is something
  299. 9:15that
  300. 9:16allows an agentic system to anticipate
  301. 9:19the consequences of its own actions.
  302. 9:22Okay, predict the consequences of its
  303. 9:24own actions. From my point of view, I
  304. 9:26cannot imagine how you can even think of
  305. 9:29building an agentic system without that
  306. 9:31system having the ability to predict the
  307. 9:33consequences of its actions.
  308. 9:35I I that's pretty essential, right? When
  309. 9:37we
  310. 9:39act in the world, we have this ability.
  311. 9:42And when we
  312. 9:43uh take an action without thinking about
  313. 9:45the consequences,
  314. 9:47we're taking a big risk. And very often,
  315. 9:49you know, other people think we're we're
  316. 9:51an idiot.
  317. 9:53Uh we have plenty of examples on the
  318. 9:55international political scene at the
  319. 9:57moment of people who have complete, you
  320. 9:59know,
  321. 10:00no ability to predict the
  322. 10:01consequences of their actions. So,
  323. 10:03that's the one model. That's what it is,
  324. 10:05right? Ability to predict the
  325. 10:06consequences of your own actions. If you
  326. 10:08If you have this ability, then you can
  327. 10:10plan
  328. 10:12a sequence of actions to
  329. 10:14accomplish a task, to you know, satisfy
  330. 10:17a goal. And you do this by
  331. 10:20planning, reasoning,
  332. 10:22uh by a process of search and
  333. 10:24optimization. You don't do this by
  334. 10:27predicting one action after the other
  335. 10:28autoregressively, like a real AI we do.
  336. 10:31Uh you do this by searching for a
  337. 10:33sequence of actions that will accomplish
  338. 10:35the task you set you set for yourself.
  339. 10:37So, the blueprint for this is completely
  340. 10:40different from what, you know, LLMs uh
  341. 10:42can do at the moment.
  342. 10:44Uh LLMs do not have the ability to
  343. 10:45predict the consequences of their
  344. 10:47actions, and they do not have any
  345. 10:48planning abilities. Because
  346. 10:50inference is by
  347. 10:52predicting the next token, right? It's
  348. 10:54not by search. Okay, so right there,
  349. 10:56you have the two characteristics that I
  350. 10:58think are essential for intelligent
  351. 11:00behavior.
  352. 11:02Ability to predict consequences of your
  353. 11:04actions. And second, uh ability to plan
  354. 11:07by optimization, by search. Um find a
  355. 11:10good sequence of actions that will
  356. 11:11produce the correct outcome. And then
  357. 11:13there is a third characteristic, which
  358. 11:15is
  359. 11:16uh how do you pre- how do you predict
  360. 11:18the consequences of your actions?
  361. 11:20Okay, so, you know, if uh if I have a
  362. 11:25water bottle in front of me. I realize
  363. 11:26some people would just listen to this
  364. 11:28and not have the picture. So, I have an
  365. 11:30open, uncapped water bottle in front of
  366. 11:33me. If I push at the bottom, it's going
  367. 11:35to slide on the table. If I push
  368. 11:38near the top, it's probably going to
  369. 11:39flip. We can't predict exactly
  370. 11:42how the
  371. 11:44the bottle will will fall in which
  372. 11:45direction.
  373. 11:47Uh we can't exactly predict how it's
  374. 11:48going to slide, you know, how the water
  375. 11:50will spill, you know, whether the table
  376. 11:52is tilted in one way and the water will
  377. 11:55uh
  378. 11:55you know, kind of flow in one direction
  379. 11:57or another.
  380. 11:58There's no way we can predict this at
  381. 12:00the pixel level.
  382. 12:01So, our mental model of the world
  383. 12:03predicts that at an abstract level of
  384. 12:06representation. So, as you were working
  385. 12:07on this architecture, was a lot of it
  386. 12:08inspired by the human brain? I mean,
  387. 12:10obviously, like the you know, the way
  388. 12:11you're articulating things is exactly
  389. 12:12how how we do things. Right, or at least
  390. 12:14by, you know, cognitive science, right?
  391. 12:16Whether you can sort of translate this
  392. 12:17into a neural architecture and things
  393. 12:19like this, that's there's a big gap
  394. 12:21there. Um okay, so that that, you know,
  395. 12:24certainly uh cognitive science was a bit
  396. 12:26of a motivation or or, you know, what uh
  397. 12:29psychological system two, which is this
  398. 12:31idea of the way you behave in sort of
  399. 12:34deliberate reflective behavior is that
  400. 12:37you do imagine, predict the consequences
  401. 12:39of your actions, and you plan
  402. 12:41uh accordingly. Contrary to system one,
  403. 12:44where you just act, you know, reactively
  404. 12:47and instinctively. So, yeah, there is an
  405. 12:49inspiration, but also there is a lot of
  406. 12:52empirical evidence
  407. 12:53that you don't want to generate pixels.
  408. 12:56Okay, I've been I've been really
  409. 12:58interested in that problem of
  410. 13:01learning models of the world by
  411. 13:03prediction for a very long time. And
  412. 13:05then had an epiphany about 5 years ago,
  413. 13:08realizing that
  414. 13:10all of the architectures that have have
  415. 13:12been successful to learn
  416. 13:14representations of images and videos
  417. 13:17are non-generative architectures. And
  418. 13:20all the generative ones basically have
  419. 13:21been failures, right? So,
  420. 13:25VAE, right? Variational autoencoders, or
  421. 13:27auto encoders more generally,
  422. 13:30uh is kind of a natural
  423. 13:32way to think about like learning
  424. 13:34abstract representations of inputs,
  425. 13:36right? So, you put a an image at the
  426. 13:38input of a
  427. 13:39of a neural net, and then you train it
  428. 13:40to just reproduce the input on its
  429. 13:43output.
  430. 13:44Uh now, with a big neural net. Now, if
  431. 13:47you just do it this way, your neural net
  432. 13:48will not do anything interesting. It
  433. 13:50will just learn the identity function.
  434. 13:51Yeah. Completely uninteresting. It
  435. 13:53doesn't work. Now, if you train a VAE to
  436. 13:55learn representations of images, you get
  437. 13:57something, but it's really not that
  438. 13:58great. Same with sparse auto encoders.
  439. 14:00Then, you have another set of
  440. 14:01techniques,
  441. 14:03uh and it's kind of derivative of
  442. 14:05something called denoising auto encoder,
  443. 14:07uh masked auto encoder is a version of
  444. 14:10this. BERT is a version of this for NLP.
  445. 14:12So, you take the image, you corrupt it
  446. 14:14in some way, and then you train this big
  447. 14:15neural net to recover the original uh
  448. 14:19the original image. There's a huge
  449. 14:20project that at FAIR on this called MAE,
  450. 14:23masked auto encoder. It was very
  451. 14:25disappointing.
  452. 14:27A lot of computation, and not not really
  453. 14:30great satisfying result. Simultaneously,
  454. 14:33uh some of the same people working on
  455. 14:35MAE, and and some other people
  456. 14:37in Paris and in New York were working on
  457. 14:39other techniques using
  458. 14:41non-generative architecture, joint
  459. 14:43embedding architecture. So, take an
  460. 14:45image, corrupt it in some way, and then
  461. 14:47run the two images through encoders, and
  462. 14:49then try to predict the representation
  463. 14:51of the original image from the
  464. 14:53representation of the corrupted one.
  465. 14:55Uh that's JEPA. Yeah. Okay. So, JEPA
  466. 14:58means joint embedding predictive
  467. 14:59architecture, right? So, you have one
  468. 15:01encoder that makes an observation,
  469. 15:03another encoder that makes a different
  470. 15:04observation. You try to predict the
  471. 15:06representation of the first one from the
  472. 15:08second one with a predictor. And those
  473. 15:11techniques turned out to work much
  474. 15:12better for representing images and
  475. 15:15video. So, things like DINO,
  476. 15:18uh DINO V1, V2, V3, um project that is
  477. 15:21still going on at at fair in Paris.
  478. 15:25Projects like I Jepa
  479. 15:27and then V Jepa and then before that
  480. 15:28there were like Sim Siam and Moco and a
  481. 15:31bunch of different techniques mostly
  482. 15:33from Meta. There was a bunch of others
  483. 15:34from other groups.
  484. 15:36Um,
  485. 15:37but
  486. 15:38that turned out to be a much better way
  487. 15:39of learning representations of images
  488. 15:42than
  489. 15:44predicting pixels. Yeah. And so
  490. 15:46it just clicked in my in my mind but you
  491. 15:48know, not just mine.
  492. 15:51That this was the way to go and
  493. 15:52predicting pixels was kind of a a losing
  494. 15:54proposition. You know, it feels like
  495. 15:56there's all these robotics demos that
  496. 15:57are released
  497. 15:59you know, from from some of the model
  498. 16:00companies that are feel increasingly
  499. 16:02impressive and maybe you know, seem to
  500. 16:04resemble things like planning and
  501. 16:05reasoning when you know, they maybe
  502. 16:07haven't seen a a room or or a specific
  503. 16:09and you know, a version of a task before
  504. 16:11and are still able to execute that task.
  505. 16:13You know, what would you say to our
  506. 16:14listeners I guess that that observe that
  507. 16:16stuff and feel like it feels like we're
  508. 16:17trending toward some real progress with
  509. 16:19some of the general approaches. Well,
  510. 16:21there is real progress and some of those
  511. 16:22demos are really impressive. Um,
  512. 16:25but
  513. 16:26>> [laughter]
  514. 16:27>> they are trained with enormous amounts
  515. 16:29of data collected either from
  516. 16:32teleoperation
  517. 16:34or from just you know, human action with
  518. 16:36things you hold in your hand that look
  519. 16:38like grippers.
  520. 16:39Grippers that you know, and you and you
  521. 16:41collect the data for that. Or just you
  522. 16:44know, tracking hands and fingers of of a
  523. 16:47person.
  524. 16:48And then translating this into kind of
  525. 16:50commands for for a robot. And so those
  526. 16:52things are trained with imitation
  527. 16:54learning mostly, right? And a little bit
  528. 16:56with you know, reinforcement learning to
  529. 16:58fine tune in mostly in simulation. So
  530. 17:01the issue with this is that you need a
  531. 17:03lot of data to train the systems
  532. 17:06to
  533. 17:07to imitation.
  534. 17:09And it
  535. 17:11it becomes expensive and it's a little
  536. 17:12brittle
  537. 17:14in the sense that you know, you need to
  538. 17:16collect lots of data for every task you
  539. 17:18want the robot to uh uh to solve.
  540. 17:21Whereas, if the system had a world model
  541. 17:23that allowed it to predict the
  542. 17:26you know,
  543. 17:27the outcome of an action, it would just
  544. 17:29plan an action to solve a new task
  545. 17:32without actually having to be trained
  546. 17:34to accomplish this task. So, the degree
  547. 17:37of generalization you would get with a
  548. 17:39world model-based system is much, much
  549. 17:41larger
  550. 17:43uh
  551. 17:43you know, kind of wider spectrum of of
  552. 17:46tasks
  553. 17:47with less training data that would be
  554. 17:49required than a a system trained with
  555. 17:51imitation learning and
  556. 17:53and you know, fine-tuning
  557. 17:54>> No doubt those approaches require more
  558. 17:56data. And I guess this question of
  559. 17:57generalization really is is the big
  560. 17:58question, right? Of you know, and I
  561. 17:59think you know, some folks have have uh
  562. 18:02have shown some results around, you
  563. 18:03know, uh getting better at task A helps
  564. 18:05with task B, but that obviously feels
  565. 18:06like there's still the big unanswered
  566. 18:08question uh you know, around those
  567. 18:09architectures. I mean, you get this uh
  568. 18:12you know, synergy between tasks. So, the
  569. 18:13more tasks that you train the system to
  570. 18:15solve, the more tasks it's being
  571. 18:17it's going to be able to acquire with
  572. 18:18with small amount of data, regardless of
  573. 18:21what what technique you use. But, but
  574. 18:22the hope with uh
  575. 18:24world models is that the system can
  576. 18:26solve new tasks at zero shot, which
  577. 18:28humans are completely capable of doing,
  578. 18:30right? And many animals as well.
  579. 18:32So, uh so, that's really the the hope.
  580. 18:35Like, you know, solving a lot more
  581. 18:36problems with uh
  582. 18:40either a small amount of training data
  583. 18:42or or no training data at all.
  584. 18:45And just a little bit of maybe, you
  585. 18:47know, RL style uh fine-tuning. Yeah.
  586. 18:49Like, you know, how
  587. 18:51how is it that a 17-year-old can learn
  588. 18:53to drive in
  589. 18:54like, a dozen hours or maybe 20 hours?
  590. 18:57Uh we have millions of hours of training
  591. 18:59data of
  592. 19:00you know, people driving cars. We still
  593. 19:02don't have level five self-driving cars,
  594. 19:04right? So, imitation learning obviously
  595. 19:06does not work even for just the task of
  596. 19:08autonomous driving. Yeah, I guess it'll
  597. 19:10be a race between the ability to develop
  598. 19:12some of those capabilities, which may
  599. 19:13take time and lots of data versus this
  600. 19:15kind of architecture. I feel like
  601. 19:16there's this dream of using video models
  602. 19:18to just generate like tons of synthetic
  603. 19:20data for for, you know, simulation and,
  604. 19:22you know, even if it's not perfect,
  605. 19:23these video models from a physics
  606. 19:24perspective, it's like helpful enough
  607. 19:26to, you know, improve
  608. 19:28robotics and in the underlying physical
  609. 19:30world. What have you made of some of
  610. 19:31those approaches? Obviously, I think
  611. 19:32Nvidia's been focused there. Google
  612. 19:33seems to be going down that road.
  613. 19:35>> I'm sort of asking you again the
  614. 19:36question,
  615. 19:37you know,
  616. 19:38why can 17-year-old launch a driving 20
  617. 19:41hours? You don't need millions of hours
  618. 19:43of demonstration. And you don't need
  619. 19:45synthetic data. Uh you don't need any of
  620. 19:47that. So, you know, I I want a system
  621. 19:49that can learn as fast as that. If we
  622. 19:51crack that, then we don't need, you
  623. 19:53know, generated data, right? I mean, we
  624. 19:55might need to train the system in
  625. 19:57simulation, but not with the same amount
  626. 19:59of uh
  627. 20:01uh
  628. 20:02you know, of time or or trials as as
  629. 20:04current systems require. It's really a
  630. 20:07question of data efficiency.
  631. 20:08>> You know, I was interviewing Jerry
  632. 20:09Tworek on the podcast. He was at OpenAI
  633. 20:11and spun out to start his own lab, and
  634. 20:13you could sense a similar tension where
  635. 20:14I think he actually might even agree
  636. 20:15that, you know, if you continued scaling
  637. 20:17RL the way we're scaling, you get more,
  638. 20:19you know, you continue getting very
  639. 20:20impressive results. But, I think he
  640. 20:22felt, "God, there's just got to be some
  641. 20:23like way more efficient way to do this."
  642. 20:25And it's interesting. It's an
  643. 20:26interesting tension because you could
  644. 20:27imagine if you're OpenAI and you know
  645. 20:29something is going to continue like you
  646. 20:30could continue scaling it and it will
  647. 20:31keep getting better. There's not a ton
  648. 20:33of incentive necessarily from a business
  649. 20:35perspective to do something more
  650. 20:36data-efficient.
  651. 20:37>> Right. And there's there's no incentive
  652. 20:39for the other companies to do anything
  653. 20:41different either because they're all
  654. 20:43chasing the same like they can't afford
  655. 20:45to kind of fall behind the others,
  656. 20:47right? So, they all work on the same
  657. 20:49thing. Yeah. And And there's a bit of
  658. 20:51this sort of, you know,
  659. 20:53kind of
  660. 20:55herd behavior
  661. 20:57uh
  662. 20:58and and, you know, in in mostly in
  663. 21:00Silicon Valley where everybody is
  664. 21:02digging the same trench. Yeah. Uh and
  665. 21:05you know, so I pur-
  666. 21:07purposely
  667. 21:09set up the headquarters of Amy Labs
  668. 21:12in Paris. Yeah.
  669. 21:13>> [laughter]
  670. 21:14>> Uh
  671. 21:16the American office being in New York,
  672. 21:18not Silicon Valley.
  673. 21:19>> [laughter]
  674. 21:19>> It's really interesting cuz I think it
  675. 21:20it it it points to a tension that, you
  676. 21:22know, it it exists in the broader
  677. 21:23ecosystem today where
  678. 21:25uh you could imagine the other side
  679. 21:26being sure, maybe there are more
  680. 21:28data-efficient methods out there, but
  681. 21:29like almost who cares because we can
  682. 21:31keep scaling what we have to to better
  683. 21:33and better results. And then obviously I
  684. 21:35think from both, you know,
  685. 21:37new things you can accomplish from these
  686. 21:39models as well as just the joy of being
  687. 21:41a researcher and finding these new
  688. 21:42things. I get why there's such an
  689. 21:43attraction to to to these other
  690. 21:45architectures as well.
  691. 21:46>> And it's a bet.
  692. 21:47But, you know, we're pretty confident
  693. 21:49because, you know, we we have results
  694. 21:51already, actually.
  695. 21:52>> And as you think about like the the kind
  696. 21:54of um the initial spaces you're most
  697. 21:56excited about for the Amy technology,
  698. 21:58like what gets you know, where do you
  699. 21:59think you know, the the technology goes
  700. 22:01and and what are you most excited about?
  701. 22:03Well, I mean, you know, AI for the real
  702. 22:04world. Um
  703. 22:07like, you know, can
  704. 22:08where is your domestic robot? Where is
  705. 22:10your level five self-driving car? Yeah.
  706. 22:12Where is uh and that's you know
  707. 22:14>> When am I going to get a domestic robot?
  708. 22:15I'm excited about this.
  709. 22:17Well, so this is several years down the
  710. 22:19line. Okay? Despite the fact that there
  711. 22:21is like
  712. 22:22huge number of companies building
  713. 22:24robots, none of those companies
  714. 22:26actually has any idea how to make them
  715. 22:28smart enough to be useful, right? Or
  716. 22:29trusted around with a baby in the house
  717. 22:31or something or
  718. 22:31>> Certainly not that. Uh but but even for
  719. 22:34like, you know, relatively narrow
  720. 22:35manufacturing task, right? You know, I
  721. 22:37mean
  722. 22:38uh none of them really knows
  723. 22:40how how to do this reliably other than
  724. 22:42you know, for by imitation learning for
  725. 22:44a small number of tasks.
  726. 22:46Uh so, how how do we make those things
  727. 22:48useful? So, that's kind of a
  728. 22:50relatively long-term objective.
  729. 22:53Shorter term, there is a huge amount of
  730. 22:55applications in industry
  731. 22:57where you need to have
  732. 22:59a a system, an intelligent system that
  733. 23:01has the ability of
  734. 23:04you know, predicting what's going to
  735. 23:05happen if I change this
  736. 23:07control variable on this complex system,
  737. 23:10be it uh
  738. 23:12a jet engine, a chemical plant, a power
  739. 23:15plant, a some manufacturing line,
  740. 23:19a patient, a human cell, right? Those
  741. 23:22are systems
  742. 23:23that are sufficiently complex that you
  743. 23:25can't
  744. 23:26model their behavior with a small number
  745. 23:28of equations. Right? So, the traditional
  746. 23:30way of modeling
  747. 23:32does not work. And what you need to do
  748. 23:34is train a neural net, deep learning
  749. 23:37system,
  750. 23:38uh to to
  751. 23:40um you know, model the dynamics of that
  752. 23:42system
  753. 23:43from data.
  754. 23:44And what you get at the end is a a
  755. 23:46phenomenological model of of that
  756. 23:49uh process, of that
  757. 23:51uh system.
  758. 23:52Um and if it's action condition, then
  759. 23:54you get
  760. 23:55basically a a world model of that system
  761. 23:58that allows you to control it optimally
  762. 24:00for whatever purpose you have. And I
  763. 24:02think the
  764. 24:04number of applications of this in
  765. 24:06industry is mind-boggling. Where do you
  766. 24:08think we'll be with uh you know, general
  767. 24:10models over the next couple years? Are
  768. 24:11there like, you know, milestones you'd
  769. 24:13point to or like, what what's your kind
  770. 24:15of view of the path of progress here?
  771. 24:16Okay, couple of years is a little short.
  772. 24:18Like, 5 years, complete world
  773. 24:19domination, essentially. [laughter]
  774. 24:21Okay. So, somewhere between on the path
  775. 24:23to world domination in 5 years. I mean,
  776. 24:25this is kind of a joke, obviously, but
  777. 24:27uh this is a quote from Linus Torvalds,
  778. 24:29right? You know, when people ask him,
  779. 24:30"What's your goal with Linux?" He said,
  780. 24:32"Total world domination." [laughter]
  781. 24:34Um he actually managed to do that.
  782. 24:36>> Yeah, very fair. To first approximation,
  783. 24:38every computer in the world runs Linux,
  784. 24:40right? So, um so, that's kind of a joke.
  785. 24:42But but in the end, I think this is the
  786. 24:44blueprint for intelligent systems of the
  787. 24:46future.
  788. 24:48There still be a a small place for LLMs,
  789. 24:51you know, for
  790. 24:52like a language interface, basically.
  791. 24:55But uh
  792. 24:56but what we're designing are are systems
  793. 24:58that are capable of thinking. They They
  794. 25:00may not be capable of talking or
  795. 25:02listening initially,
  796. 25:04but they'll do the thinking.
  797. 25:06And then you can add the talking and
  798. 25:08listening
  799. 25:10uh on top of that. I'm sure you and the
  800. 25:12team are are are eagerly working to kind
  801. 25:14of, you know, get the early proof points
  802. 25:15of this. And obviously, you've already
  803. 25:17had some in the work you've done. How do
  804. 25:18you think about like the interim steps
  805. 25:19of what you'll be able to show on that
  806. 25:21path to to 5-year world domination?
  807. 25:23Well, so I think uh
  808. 25:25you know, within a year or so, um we'll
  809. 25:28have
  810. 25:29I think a a general methodology
  811. 25:32to train
  812. 25:33hierarchical world models
  813. 25:35on, you know, a a very wide variety of
  814. 25:38modalities.
  815. 25:40We know we can do a good job on video
  816. 25:42uh with some techniques that we're not
  817. 25:44completely happy with because they have
  818. 25:46some shortcomings, but
  819. 25:48um
  820. 25:49and we have
  821. 25:50sort of small-scale demonstration of a
  822. 25:53methodology that we think is
  823. 25:55really what we want.
  824. 25:57So, we need to scale that one up
  825. 25:59and get it to the same level of
  826. 26:00performance as the
  827. 26:03the other techniques that are not as uh
  828. 26:06satis- satisfying, if you want, on on
  829. 26:08things like video, but also on other
  830. 26:10types of data sets that we would get
  831. 26:12from industry partners. Okay, so we'll
  832. 26:14have
  833. 26:15demonstrations that we can train world
  834. 26:17models, perhaps action-conditioned world
  835. 26:18models that allow us to plan
  836. 26:20for uh a number of different use cases.
  837. 26:23Some of them will be robotics, some of
  838. 26:25them will be industrial process control
  839. 26:27of various types, maybe some of them in
  840. 26:29health um health care as well cuz we
  841. 26:32have partners in that
  842. 26:33>> Yeah. in that domain. And
  843. 26:36that should be within a year or two, 18
  844. 26:38months. Um
  845. 26:40And then we'll push the this methodology
  846. 26:43and those models into
  847. 26:45uh those use cases with partners, some
  848. 26:48of which are investors already, you
  849. 26:49know, in our company, and gain
  850. 26:51experience on how to kind of
  851. 26:54essentially build a somewhat universal
  852. 26:56world model if you want. I mean, you've
  853. 26:57obviously had this uh you know, this
  854. 26:59experience before of of kind of making
  855. 27:01this really contrarian bet on neural
  856. 27:02nets and and being certainly uh proven
  857. 27:05abundantly right uh in the in in the
  858. 27:06history books. I guess as you think
  859. 27:08about this bet which I think, you know,
  860. 27:09if you talk to the majority of people uh
  861. 27:11maybe at at at at the cutting edge of
  862. 27:13various parts of AI maybe would would
  863. 27:14say is contrarian today. In what time
  864. 27:16frame do you think it will become
  865. 27:17apparent like, you know, that this was
  866. 27:19right?
  867. 27:20I think it'll happen faster than
  868. 27:24expected perhaps because I mean, you can
  869. 27:26see that world model is already becoming
  870. 27:28a buzzword, right?
  871. 27:30At least at the research level.
  872. 27:32Uh
  873. 27:34and it's starting to kind of permeate
  874. 27:35into the industry. Yeah. And a lot of
  875. 27:37people are realizing like VLs suck
  876. 27:40and, you know, LLMs don't work for real
  877. 27:42world data. Industry has realized this
  878. 27:45already. Certainly on the on the
  879. 27:48on the user side.
  880. 27:50And I think because of the importance of
  881. 27:52the robotics industry
  882. 27:54um you know, a lot of people are kind of
  883. 27:55trying to figure out like how how do we
  884. 27:57how do we get there? How do we get how
  885. 27:58you make those robots
  886. 28:00uh
  887. 28:01useful. So so I think it's
  888. 28:03I think the realization that you need a
  889. 28:05change of paradigm is is happening as we
  890. 28:07speak and will become completely obvious
  891. 28:09to people by
  892. 28:11early 2027, I think. Yeah. Now, that
  893. 28:14doesn't mean we'll have a solution by
  894. 28:15then. We hope we will, but you know,
  895. 28:17we'll see. I guess, you know, switching
  896. 28:18gears to the LM side you mentioned some
  897. 28:20of this work you're doing with uh with
  898. 28:21Tapestry which I think would be really
  899. 28:22interesting for our listeners. And so
  900. 28:24maybe to speak to that a little bit.
  901. 28:25Okay, so this is kind of a little bit
  902. 28:27orthogonal to uh to ML Labs. Yeah, as if
  903. 28:30that wasn't enough to keep you busy.
  904. 28:32>> [laughter]
  905. 28:32>> Well, it's a it's a kind of an idea I've
  906. 28:34I've been uh forming over the last uh
  907. 28:36three years or so
  908. 28:38is the fact that uh
  909. 28:41people increasingly use AI assistants
  910. 28:43for various things, right? I mean, uh
  911. 28:45you see a decrease in the use of general
  912. 28:48traditional search engines and you just
  913. 28:50ask a question to your favorite AI
  914. 28:52assistant.
  915. 28:54Um and you know, if the plan that Meta
  916. 28:58and others are are
  917. 29:00developing of, you know, having smart
  918. 29:01devices like smart glasses and stuff
  919. 29:03like that,
  920. 29:04uh
  921. 29:05you know, is realized,
  922. 29:07basically you'd just be talking to your
  923. 29:09AI assistant, you know, by voice with,
  924. 29:11you know, to your smart glasses or maybe
  925. 29:13some other smart device.
  926. 29:15And so, all of your information diet
  927. 29:18will be mediated by AI assistants.
  928. 29:22And
  929. 29:23if you are someone, you know, somewhere
  930. 29:24in the world, let's say outside the US
  931. 29:26or China, and you have an AI assistant,
  932. 29:28and that AI assistant was built in
  933. 29:29California or
  934. 29:32you know, Beijing
  935. 29:33or Shanghai or Shenzhen, uh
  936. 29:37it's not good for you. Like you may
  937. 29:39speak a language that those systems
  938. 29:41really haven't been trained to handle
  939. 29:43particularly well.
  940. 29:44Uh you may have a culture that is not
  941. 29:47particularly well understood by people
  942. 29:48in Silicon Valley and China.
  943. 29:50Not well represented by the training
  944. 29:52data that is publicly available on the
  945. 29:54internet.
  946. 29:56Um
  947. 29:57you may have a value system that is
  948. 29:58absolutely not represented by
  949. 30:01uh you know, people building those
  950. 30:02models.
  951. 30:03And certainly you'll almost certainly
  952. 30:05have political opinions that are
  953. 30:07absolutely not represented by the
  954. 30:09handful of AI assistant you you might be
  955. 30:12able to get from the
  956. 30:14you know, um West Coast tech companies
  957. 30:16or from Chinese companies.
  958. 30:19So, what is the solution to this? Like
  959. 30:21how do you serve,
  960. 30:23uh you know, a farmer in India,
  961. 30:25uh or um even a philosopher in France,
  962. 30:30uh or Germany? And
  963. 30:33what you need is
  964. 30:35a platform
  965. 30:37which basically is uh an open free
  966. 30:41foundation model, LLM style,
  967. 30:44that is fine-tunable
  968. 30:46by anyone
  969. 30:48to cater to the interest of people
  970. 30:51speaking a particular language, having a
  971. 30:53particular culture,
  972. 30:55having particular
  973. 30:57value systems, political biases,
  974. 30:59uh
  975. 31:00creeds, whatever it is.
  976. 31:03And so, what you need is a wide
  977. 31:05diversity of AI assistants. There's a
  978. 31:08lot of countries around the world uh
  979. 31:11that are neither the US nor China, who
  980. 31:14absolutely want some level of
  981. 31:16sovereignty for AI, not just for their
  982. 31:18industry, but also for their citizen.
  983. 31:20They don't want their citizen to get
  984. 31:22brainwashed by
  985. 31:24a Chinese model or a Canadian model,
  986. 31:26actually.
  987. 31:27Uh And so,
  988. 31:29they want sovereignty. How do you get
  989. 31:31that? So, the way you get
  990. 31:34a platform like the an open platform
  991. 31:36like this to get to the frontier is you
  992. 31:38just train it on more and higher quality
  993. 31:40data than the than the proprietary
  994. 31:43systems.
  995. 31:44If you talk to
  996. 31:46people in India, in France, in Vietnam,
  997. 31:49in
  998. 31:51Morocco, in Switzerland,
  999. 31:54in Korea, Japan,
  1000. 31:56uh
  1001. 31:57Kazakhstan,
  1002. 31:59everyone wants
  1003. 32:02basically sovereignty.
  1004. 32:04And you tell them like, you guys have
  1005. 32:06been training your model, you know,
  1006. 32:08locally. You don't have to share your
  1007. 32:09data. So, that's the crucial aspect of
  1008. 32:11Tapestry.
  1009. 32:12You would have international contributor
  1010. 32:15contributors to Tapestry
  1011. 32:17contributing to training a a global
  1012. 32:19model that would basically constitute a
  1013. 32:23repository of all the world's knowledge
  1014. 32:25and culture, if you want. But the
  1015. 32:26contributors would contribute uh
  1016. 32:30data and uh computing resources, but
  1017. 32:33they would preserve the control on their
  1018. 32:35data. They would not have to share their
  1019. 32:37data with the other
  1020. 32:38uh contributors.
  1021. 32:40What they would contribute is
  1022. 32:42parameter vectors. Interesting. So, it
  1023. 32:44would be kind of a kind of federated
  1024. 32:46learning style thing where uh you have a
  1025. 32:48bunch of data centers. Uh
  1026. 32:51you know, they they get the parameter
  1027. 32:53vector from the
  1028. 32:55the the global consensus of a model.
  1029. 32:58Think of it as an average of all the
  1030. 33:01all the parameter vectors of all the
  1031. 33:02contributors, right?
  1032. 33:04So, all the contributors uh periodically
  1033. 33:06tell
  1034. 33:08everyone else through maybe a central
  1035. 33:10server, here is my parameter vector,
  1036. 33:13what is yours? Okay.
  1037. 33:15Uh and so, you exchange parameter
  1038. 33:16vectors like this. And a local worker
  1039. 33:19basically, whatever it updates its
  1040. 33:20parameter vector, it tries to also
  1041. 33:24makes make it as close as possible to
  1042. 33:26the global consensus vector. So, as the
  1043. 33:30training of this thing kind of
  1044. 33:31progresses, all those parameter vectors
  1045. 33:34converge towards
  1046. 33:36like a a consensus model, essentially,
  1047. 33:38which is kind of a repository of all
  1048. 33:41human knowledge. Now, you have an open
  1049. 33:43an open model
  1050. 33:45that is as good as if it had been
  1051. 33:47trained on all the data in the world.
  1052. 33:50And now, you can fine-tune it for your
  1053. 33:52own purpose, your own
  1054. 33:55political, cultural, and linguistic
  1055. 33:56biases.
  1056. 33:57Whatever you want or centers of
  1057. 33:58interest.
  1058. 33:59And I think there is a natural force for
  1059. 34:01this to happen
  1060. 34:03uh because, you know, most countries
  1061. 34:05that are not the US nor China want
  1062. 34:09sovereignty, but also because
  1063. 34:11uh
  1064. 34:12AI is fast becoming a platform. And
  1065. 34:15there is a natural tendency for
  1066. 34:16platforms to become open.
  1067. 34:18That's what happened with Linux,
  1068. 34:20right? And that's what happened with the
  1069. 34:23software infrastructure of the internet
  1070. 34:25or the wireless network. It's all open
  1071. 34:27source.
  1072. 34:28Um
  1073. 34:29it was proprietary initially, but that
  1074. 34:32was all
  1075. 34:33wiped out. It's a really clever way to
  1076. 34:34get around you know it would seem that
  1077. 34:36this trend of you know decreasing open
  1078. 34:38source and and obviously I think there's
  1079. 34:41been many fears that is like the closed
  1080. 34:42source models get better they'll be held
  1081. 34:44back and they'll be used to train the
  1082. 34:45next generation and you know they'll
  1083. 34:46they'll kind of be this almost like a
  1084. 34:48scape scenario for for closed source
  1085. 34:49models where they get you know so much
  1086. 34:51better than than their open source
  1087. 34:52counterparts. So remember what you know
  1088. 34:54who the big players of the internet
  1089. 34:56infrastructure were
  1090. 34:58in 1990 six?
  1091. 35:01Sun Microsystems HP
  1092. 35:04Dell
  1093. 35:05and a few others.
  1094. 35:07Um so Sun Microsystems was selling you
  1095. 35:09Solaris with their you know proprietary
  1096. 35:11hardware.
  1097. 35:12HP with HP-UX
  1098. 35:14uh they were claiming you know Unix is
  1099. 35:16so much more reliable than Windows
  1100. 35:17you're not going to run a web server on
  1101. 35:18Windows. Dell was doing this you know
  1102. 35:21with Windows NT but like who is running
  1103. 35:23Windows NT now [laughter] as a web
  1104. 35:25server?
  1105. 35:26All of this was totally wiped out by
  1106. 35:28Linux like the entire internet runs on
  1107. 35:30Linux.
  1108. 35:31Um even Azure right? Even Microsoft
  1109. 35:34>> [laughter]
  1110. 35:34>> it runs Linux. So
  1111. 35:37uh
  1112. 35:39basically OpenAI and Anthropic
  1113. 35:41etc. of today are the
  1114. 35:44Sun Microsystem and HP-UX
  1115. 35:47of yesterday.
  1116. 35:49Yeah I mean I guess it it implicit in
  1117. 35:51that is is obviously um you know I think
  1118. 35:53your you know your view of like the
  1119. 35:55limitations of of what like you know the
  1120. 35:57these models can only get so good and so
  1121. 36:00it'll be possible over time for for the
  1122. 36:01open source folks to to catch up.
  1123. 36:03They've already run out of data right? I
  1124. 36:04mean the the
  1125. 36:06the open openly available publicly
  1126. 36:08available data text data
  1127. 36:11uh
  1128. 36:12is already all used.
  1129. 36:13I mean there's not more of it right? So
  1130. 36:15what what those companies are doing is
  1131. 36:17licensing uh
  1132. 36:19commercial copyrighted data
  1133. 36:22or training on
  1134. 36:24synthetic data. I guess I'm curious cuz
  1135. 36:25obviously there's been some impressive
  1136. 36:27results uh in the last few years that
  1137. 36:29they that they have been able to drive,
  1138. 36:30you know, post these large-scale
  1139. 36:31pre-trainings. Um you know, IMO gold, uh
  1140. 36:34you know, the meter
  1141. 36:35task horizon benchmark keeps going up.
  1142. 36:37Okay, that's Okay, that's very
  1143. 36:39interesting. Now, think about those two
  1144. 36:40domains, right? Mathematics and code.
  1145. 36:43Those are two domains where the language
  1146. 36:45itself is the substrate of reasoning.
  1147. 36:49It's not the only substrate of
  1148. 36:50reasoning, but a lot of
  1149. 36:52when you do mathematics, right? The the
  1150. 36:55formal way on a piece of paper, not the
  1151. 36:57intuitive stuff, but the
  1152. 36:58you manipulate language, right? And LLMs
  1153. 37:01are really good at this. So,
  1154. 37:03um
  1155. 37:03you know, proving theorems and stuff
  1156. 37:05like that, that's that's what LLMs are
  1157. 37:07really good at.
  1158. 37:08They're not so good at the sort of
  1159. 37:11you know, coming up with uh good
  1160. 37:12concepts and definitions and things like
  1161. 37:14that. It it's more like, here is a
  1162. 37:15problem, solve it. They're problem
  1163. 37:16solvers. Mathematics is not just problem
  1164. 37:19solving, right? Most of it
  1165. 37:21uh is actually a creative act that those
  1166. 37:23things don't do.
  1167. 37:25Um
  1168. 37:27and same for code. So, LLMs are good
  1169. 37:29programmers. They're not software
  1170. 37:31architects.
  1171. 37:33They're not computer scientists, right?
  1172. 37:36Uh
  1173. 37:37but they can program for us.
  1174. 37:39So, they they they're not in a in a
  1175. 37:41state where they can just, you know,
  1176. 37:43replace humans uh entirely. It changes
  1177. 37:46the world of humans. So, humans now,
  1178. 37:48you know, kind of go one level up in the
  1179. 37:50abstraction hierarchy and
  1180. 37:53our world is to decide what to build.
  1181. 37:54But like building it, you know, you can
  1182. 37:56you can get help from LLMs. But okay,
  1183. 37:59that's the the important point is that
  1184. 38:01uh LLMs are particularly successful at
  1185. 38:04domains where the language itself is the
  1186. 38:07substrate of reasoning.
  1187. 38:08Uh not for anything else. Yeah. What
  1188. 38:11would an LLM like need to do to convince
  1189. 38:12you uh otherwise? So, I mean, like a
  1190. 38:15zero-shot agentic system, right?
  1191. 38:18You have an agentic system.
  1192. 38:20Give it a new problem. It's not been
  1193. 38:21trained to solve that that problem.
  1194. 38:23Doesn't have a script for it.
  1195. 38:25Uh
  1196. 38:26is it going to be able to uh accomplish
  1197. 38:28this task? That it's never been trained
  1198. 38:30to solve.
  1199. 38:32And unless the system has the ability of
  1200. 38:34predicting the consequences of its
  1201. 38:36actions and then use using use that for
  1202. 38:38play for planning
  1203. 38:41it's not going to be able to do it. And
  1204. 38:42you're not going to do this with an LLM.
  1205. 38:44You're going to do this perhaps with a
  1206. 38:46significantly
  1207. 38:47augmented LLM that is capable of
  1208. 38:50you know, search and planning blah blah
  1209. 38:52blah. And currently
  1210. 38:54you know, LLMs that do math and code
  1211. 38:55actually do this.
  1212. 38:56>> Yeah. Right? Cuz they search for you
  1213. 38:58know, sequences of tokens that actually
  1214. 39:00accomplish a particular task and you
  1215. 39:03know, they can run the code or verify
  1216. 39:05that the proof is correct or whatever.
  1217. 39:08Um so, you have like a way of checking
  1218. 39:10whether something that's produced is is
  1219. 39:11is correct. Um but that's not a very
  1220. 39:14efficient way of of doing planning.
  1221. 39:16And it only works in domains where this
  1222. 39:19type of search can be performed in token
  1223. 39:21space.
  1224. 39:22What I'm talking about with Jeppa is you
  1225. 39:24don't do this in token space. You do
  1226. 39:25this in you know, abstract thoughts
  1227. 39:28space. And I'm sure some people
  1228. 39:30listening might think, well, you know,
  1229. 39:31hey, if if even if it's inefficient and
  1230. 39:32it works uh and it works at you know, at
  1231. 39:35things that are done in token space,
  1232. 39:36that's still a large part of the uh of
  1233. 39:38the economy that
  1234. 39:39>> I mean, if it works, it's fine. I mean,
  1235. 39:41there's again, there's nothing wrong
  1236. 39:42with you know, using an LLM for what
  1237. 39:44they're good at.
  1238. 39:45Uh it's just not a path towards human
  1239. 39:48level AI. You're missing you know, an
  1240. 39:50>> like a huge
  1241. 39:51uh domain.
  1242. 39:52>> You seem like you you know, hey, it's
  1243. 39:53going to tap out before it can become a
  1244. 39:55software architect, whereas I'm sure
  1245. 39:56>> going to tap out. It's it's just going
  1246. 39:58to have like a a limited you know,
  1247. 40:00ability to be deployed for like an it's
  1248. 40:02going to become like increasingly
  1249. 40:03difficult to kind of deploy it for an
  1250. 40:05increasingly large number
  1251. 40:07you know, of of use cases because you're
  1252. 40:09going to have to collect tons of
  1253. 40:10training data for each of those use
  1254. 40:12cases. And
  1255. 40:14there's a basically you're not going to
  1256. 40:16be able to make those systems completely
  1257. 40:18reliable, you know, without
  1258. 40:19hallucinations or or dangerous stuff or
  1259. 40:22uh etc.
  1260. 40:24Unless those systems have the ability to
  1261. 40:26predict the consequences of their
  1262. 40:27actions, which means they're going to
  1263. 40:28have to have explicit world models.
  1264. 40:30Yeah, so I guess it's a bet against, you
  1265. 40:31know, the uh 100% accuracy and then also
  1266. 40:34the generalization uh across different
  1267. 40:36tasks.
  1268. 40:36>> Right. I guess, you know, one thing that
  1269. 40:38that's so interesting about the way that
  1270. 40:39the field has developed is obviously you
  1271. 40:41uh share the the Turing Award with two
  1272. 40:43others and I feel like they seem much
  1273. 40:44more convinced of like maybe the the
  1274. 40:46power or potential threats or safety
  1275. 40:47risks of LLMs over time. Um
  1276. 40:50I'm wondering like when did your view
  1277. 40:51start diverging?
  1278. 40:53Uh in 2023.
  1279. 40:56And what like drove that in your mind? I
  1280. 40:58didn't change my mind.
  1281. 40:59They changed their mind, okay?
  1282. 41:00[laughter]
  1283. 41:01And I just about the same time, and it
  1284. 41:03was basically GPT-4.
  1285. 41:05I mean, Jeff basically had
  1286. 41:07was not connected to any of that. He was
  1287. 41:09never really interested in
  1288. 41:11LLMs and discovered uh
  1289. 41:14GPT-4 or, you know, 2023 when it came
  1290. 41:16out.
  1291. 41:17And basically he had an epiphany and
  1292. 41:18said, "Oh my god, those systems, you
  1293. 41:20know, are really close to human-level
  1294. 41:22intelligence and they have
  1295. 41:24possibly they have subjective
  1296. 41:25experience.
  1297. 41:26Uh
  1298. 41:28and he he did a a quick calculation
  1299. 41:29saying like, "Okay, the human cortex has
  1300. 41:32about 16 billion neurons. If you want to
  1301. 41:35um
  1302. 41:37do something like backprop, okay? The
  1303. 41:39brain doesn't do
  1304. 41:41backprop directly.
  1305. 41:42But if it does something like backprop,
  1306. 41:44like some sort of, you know, gradient
  1307. 41:45estimation for some sort of objective
  1308. 41:47function,
  1309. 41:48you would probably need like a a network
  1310. 41:50of a few neurons to kind of reproduce
  1311. 41:52the functionality of a virtual neuron in
  1312. 41:54a in a neural net."
  1313. 41:56So he said like, "Let's assume, you
  1314. 41:58know, maybe you need you need
  1315. 42:00a circuit of 10
  1316. 42:02actual neurons
  1317. 42:03to reproduce what a a backprop neuron
  1318. 42:05does.
  1319. 42:07Then all of a sudden your your cortex is
  1320. 42:09only 1.6 billion neurons.
  1321. 42:12Oh my god, GPT-4 is really close to
  1322. 42:13this. Okay, so maybe it's as smart you
  1323. 42:15know, it's going to get as smart as
  1324. 42:16humans. I do not believe in this claim
  1325. 42:18at all. This is kind of you know, Jeff's
  1326. 42:22uh
  1327. 42:23way of
  1328. 42:25saying
  1329. 42:26Okay, basically
  1330. 42:28I can retire. I can declare victory.
  1331. 42:31You know,
  1332. 42:32I searched for the learning algorithm of
  1333. 42:34the cortex all my career.
  1334. 42:36Uh
  1335. 42:37maybe I didn't discover what it really
  1336. 42:39was, but backprop seems to be like a
  1337. 42:41good substitute for it.
  1338. 42:42It works really well.
  1339. 42:44And so
  1340. 42:45maybe that's all we need. So, I can
  1341. 42:47retire.
  1342. 42:49Uh
  1343. 42:50and [laughter]
  1344. 42:51and go around the world and give talks
  1345. 42:52about you know, the potential
  1346. 42:54uh
  1347. 42:56promises and dangers of uh of AI.
  1348. 43:00Uh that's basically what you know, I
  1349. 43:02think what his uh
  1350. 43:03uh intellectual kind of trajectory has
  1351. 43:06been.
  1352. 43:07Uh
  1353. 43:08he's much less
  1354. 43:10vocal about the potential dangers now
  1355. 43:12than he was
  1356. 43:14uh
  1357. 43:15a year or two ago.
  1358. 43:16He kind of realized there's probably a
  1359. 43:17way to design
  1360. 43:19truly intelligent systems. So, first of
  1361. 43:20all, he probably you know, he realized
  1362. 43:22that
  1363. 43:23current LLMs are not that smart first of
  1364. 43:25all. And and second
  1365. 43:27uh
  1366. 43:28that there's probably a need for a few
  1367. 43:30breakthroughs like conceptual
  1368. 43:31breakthroughs before we get to
  1369. 43:33human-like intelligence.
  1370. 43:35And third, that the the blueprint of
  1371. 43:37those systems will be quite different
  1372. 43:39from LLMs and we have probably have a
  1373. 43:42way of
  1374. 43:43you know, making them controllable and
  1375. 43:45things like that. Yeah. I've been saying
  1376. 43:47this for years, but
  1377. 43:49Okay, he's sort of discovered this
  1378. 43:51recently. Yeah. Same kind of there's a
  1379. 43:53similar thing with Yoshua. I think what
  1380. 43:55they are both worried about is
  1381. 43:58the ability of
  1382. 43:59society
  1383. 44:01and the political system to make sure
  1384. 44:04that the benefits of of AI will be
  1385. 44:06maximized.
  1386. 44:07And AI would not, you know, just
  1387. 44:10profit you know, make a few rich people
  1388. 44:13even richer.
  1389. 44:14And uh you know,
  1390. 44:17accentuate inequalities and and you
  1391. 44:20know, cause major catastrophes because
  1392. 44:23of bad usage. Okay, this is not like the
  1393. 44:25the doomer scenario of AI taking over
  1394. 44:27the world. It's more bad use users. What
  1395. 44:31seems possible with the LLMs of today?
  1396. 44:32Which is a danger, but you know, I don't
  1397. 44:35I don't think it's as
  1398. 44:37apocalyptic as you know, what some
  1399. 44:39people have claimed it is.
  1400. 44:41Certainly not as apocalyptic as what
  1401. 44:43even Anthropic has claimed.
  1402. 44:45And it's trying to kind of lobby
  1403. 44:46governments into you know, scaring
  1404. 44:48governments into kind of regulating AI
  1405. 44:51because because of that.
  1406. 44:52I don't I don't I don't subscribe to
  1407. 44:55this at all. They seem to genuinely
  1408. 44:56believe it. I think they genuinely
  1409. 44:58believe it, but also I think there is,
  1410. 45:01you know, some kind of commercial good
  1411. 45:03commercial reasons for them to believe
  1412. 45:04that. And to kind of
  1413. 45:06uh
  1414. 45:08you know, brainwash
  1415. 45:09some people and governments into
  1416. 45:11thinking their systems are are
  1417. 45:12dangerous. And it sounds like you know,
  1418. 45:14with these other architectures, do you
  1419. 45:15think they're cuz obviously it doesn't
  1420. 45:16you know, as maybe
  1421. 45:18bearish as you are on LLMs being the end
  1422. 45:20state of everything, you know, you have
  1423. 45:22some pretty ambitious timelines too for
  1424. 45:23for these new architectures. And so it
  1425. 45:25doesn't seem like you think we're
  1426. 45:26particularly far away from from from
  1427. 45:28some very compelling capabilities. How
  1428. 45:30do you think about I guess the the
  1429. 45:31safety around, you know, if it ends up
  1430. 45:33if these breakthroughs end up coming
  1431. 45:34from new architectures and whether that
  1432. 45:35should make us rest easier or not. I'm
  1433. 45:38going to say something that's again
  1434. 45:40might be controversial. Uh
  1435. 45:42and certainly my some of my colleagues
  1436. 45:44at Meta didn't like me saying this, but
  1437. 45:46I think LLMs are interestingly unsafe.
  1438. 45:49I don't think they can be made
  1439. 45:51reliable and safe. Okay.
  1440. 45:53They cannot be made reliable because you
  1441. 45:55can't stop them from hallucinating.
  1442. 45:57Uh and if they're agentic, you cannot
  1443. 45:59guarantee they're not going to like take
  1444. 46:01an action that, you know, they didn't
  1445. 46:03predict the outcome of and that
  1446. 46:05>> I mean, does it surprise you they can do
  1447. 46:06these like 15-hour coding tasks given
  1448. 46:08the concerns around reliability?
  1449. 46:09>> Well, but coding is something where you
  1450. 46:10can actually verify that, you know, the
  1451. 46:13the the code that you generate uh you
  1452. 46:15know, satisfy your specification.
  1453. 46:17Um
  1454. 46:19but
  1455. 46:20but not everything is coding. And and
  1456. 46:23there are examples of, you know,
  1457. 46:25uh
  1458. 46:26coding agents like wiping up your
  1459. 46:28your hard drive,
  1460. 46:30like
  1461. 46:31uh
  1462. 46:32or or doing stupid things, right? That
  1463. 46:33makes you lose a lot of money or data or
  1464. 46:36whatever.
  1465. 46:37So, I think I think uh you know, LLMs in
  1466. 46:40their current forms
  1467. 46:41uh are are intrinsically unsafe
  1468. 46:44because they cannot predict the
  1469. 46:45consequences of their actions and
  1470. 46:46because the way the task that they
  1471. 46:49accomplish is determined
  1472. 46:51uh is
  1473. 46:53is subject to their training. You know,
  1474. 46:56you you give them a prompt
  1475. 46:59and then they will accomplish a task
  1476. 47:01that correspond to that prompt only to
  1477. 47:03the ex-
  1478. 47:04to the extent that their training
  1479. 47:06has conditioned them to actually do the
  1480. 47:08right task corresponding to this prompt.
  1481. 47:11But there is no, like, you know,
  1482. 47:12hardwired constraint that will force
  1483. 47:15them to accomplish this task
  1484. 47:17and then, you know, predict that the
  1485. 47:19task will be accomplished properly.
  1486. 47:20Yeah, I mean, I think people were saying
  1487. 47:21in the early days, right? They would
  1488. 47:22you'd ask them a question and they'd
  1489. 47:23keep asking the they'd keep asking the
  1490. 47:24question, right?
  1491. 47:25>> Right. Right. For example. [laughter]
  1492. 47:27Uh
  1493. 47:28or I mean, also they don't have common
  1494. 47:29sense. Right. So, I mean, there's the
  1495. 47:31the joke that was circulating like a
  1496. 47:33month ago of
  1497. 47:35you know, I need to wash my car and you
  1498. 47:37know, the the car wash is a 100 yards
  1499. 47:39from my house, should I walk?
  1500. 47:43I tried it again like maybe 2 weeks ago.
  1501. 47:45Uh they all say, "Yes, you should walk."
  1502. 47:47Except Gemini.
  1503. 47:49Gemini says
  1504. 47:50>> they're training on your video of of
  1505. 47:51having done having given that speech
  1506. 47:53before? The The was not my video because
  1507. 47:54[laughter] I I come up with this
  1508. 47:55example.
  1509. 47:56>> Whoever came up with it. Yeah, right.
  1510. 47:57Whoever came up with it. But they are
  1511. 47:58issuing sentences, right? Where where I
  1512. 48:00said like, you know, an LLM can do this
  1513. 48:02and then 6 months later it was like
  1514. 48:03people are doing it and it's simply
  1515. 48:04because, you know, as soon as people
  1516. 48:07watch the podcast of me saying LLMs can
  1517. 48:10do this, they of course type it into
  1518. 48:12ChatGPT. So now it becomes part of the
  1519. 48:14training set. And now of course, you
  1520. 48:16know, the next version has that uh you
  1521. 48:19know, that that thing in the fine-tuning
  1522. 48:21set and of course it can answer the
  1523. 48:22question but it's not because it's it
  1524. 48:23becomes smart all of a sudden. It's just
  1525. 48:25because it was explicitly trained with
  1526. 48:27that question. So LLMs are intrinsically
  1527. 48:29unsafe. Uh I don't think there is any
  1528. 48:32way to fix that in the current
  1529. 48:34um paradigm.
  1530. 48:36Um and what I've been proposing is
  1531. 48:38the architecture I've been talking about
  1532. 48:40is objective-driven AI. So basically,
  1533. 48:43you give an objective to an AI system,
  1534. 48:46which is accomplish this task. Now,
  1535. 48:48how does the system
  1536. 48:50knows
  1537. 48:51it will accomplish this task? It has a
  1538. 48:53world model and it predicts
  1539. 48:55uh
  1540. 48:57you know, the outcome
  1541. 48:59of a sequence of actions it imagines
  1542. 49:01taking.
  1543. 49:02Uh and if this uh outcome
  1544. 49:05satisfies
  1545. 49:07a
  1546. 49:09cost function that you know, describes
  1547. 49:11to what extent the task has been
  1548. 49:12accomplished or not accomplished,
  1549. 49:14then that system, if
  1550. 49:16if the way that system works is by
  1551. 49:19optimizing by optimization, finding a
  1552. 49:22sequence of actions that accomplishes
  1553. 49:24this task, minimizes this cost according
  1554. 49:26to its world model,
  1555. 49:28it can do nothing else. Yeah. Okay. And
  1556. 49:32of course there's many things that can
  1557. 49:33go wrong there. In in particular,
  1558. 49:36uh the cost function might be
  1559. 49:38inaccurate. It could be that the cost
  1560. 49:40function you think is actually measuring
  1561. 49:42to what extent the task has been
  1562. 49:44accomplished but perhaps
  1563. 49:46it's not accurate. Okay?
  1564. 49:48Uh you the world model might be
  1565. 49:49inaccurate. So the prediction that the
  1566. 49:51system makes is actually not the right
  1567. 49:52one. So, its prediction of what was
  1568. 49:54going to happen as a consequence of its
  1569. 49:56action wasn't right. Okay, so the system
  1570. 49:58can still make mistakes, but but it can
  1571. 50:01predict the consequences of its actions
  1572. 50:02to some extent, which is I think
  1573. 50:05indispensable for any agentic system.
  1574. 50:07Now, you can add to that system is not
  1575. 50:09just
  1576. 50:10a cost function that guarantees a task
  1577. 50:12has been accomplished, but you can also
  1578. 50:14add a bunch of other
  1579. 50:17objective functions, other other cost
  1580. 50:18functions, or even constraints
  1581. 50:21that are safety constraints.
  1582. 50:23And say, "Okay, you know, don't hurt
  1583. 50:24anybody on the way, right?" And you
  1584. 50:26cannot specify this at a at an abstract
  1585. 50:28level, but you can have, you know,
  1586. 50:30low-level objective functions that
  1587. 50:33put together will guarantee that the
  1588. 50:34system will not be dangerous. Uh and the
  1589. 50:37system cannot violate those things by
  1590. 50:39construction. It will have to satisfy
  1591. 50:41those conditions.
  1592. 50:42Not the case for an LLM. The LLM can
  1593. 50:44always escape. There's a gap between
  1594. 50:47your training error and test error.
  1595. 50:49There's always going to be a prompt
  1596. 50:50where the system is going to do really
  1597. 50:51stupid things. To talk to one specific
  1598. 50:53space around LLMs, like, you know, I I
  1599. 50:55think you're obviously really excited
  1600. 50:56about LLMs in healthcare, and I think
  1601. 50:57you know, people have been using LLMs in
  1602. 50:59healthcare for for all sorts of things.
  1603. 51:00And so, I'm curious how you think about
  1604. 51:02like the set of things where LLMs are
  1605. 51:04just not going to work in healthcare,
  1606. 51:05and you need like a
  1607. 51:07model that understands the world better.
  1608. 51:09So, uh I mean, designing a course of
  1609. 51:11treatment for a chronic disease, for
  1610. 51:13example,
  1611. 51:14or even a non-chronic disease,
  1612. 51:17uh for a particular patient,
  1613. 51:19uh which may not completely fit into,
  1614. 51:22you know, templates that you've observed
  1615. 51:23before.
  1616. 51:24But if you have a good mental model of
  1617. 51:26the
  1618. 51:27dynamics of the physiology of the
  1619. 51:29patient, then you might design a course
  1620. 51:31of treatment that will actually bring
  1621. 51:33the the patient to a good state. Yeah.
  1622. 51:35Uh when I'm saying and when I'm saying a
  1623. 51:36patient, it can be
  1624. 51:38a cell. Okay, how do you tell
  1625. 51:41a a
  1626. 51:43stem cell
  1627. 51:44to turn into a
  1628. 51:46uh pancreas beta cell that produces
  1629. 51:48insulin.
  1630. 51:49Okay, you have a patient with type 1
  1631. 51:51diabetes.
  1632. 51:53Um you know, they have
  1633. 51:55you know, their immune system basically,
  1634. 51:57you know, kind of
  1635. 51:59eats up their own beta cells, right?
  1636. 52:01It's autoimmune.
  1637. 52:02Um how do you keep making beta cells?
  1638. 52:04You know, can you send a message? Do you
  1639. 52:06have a a model of a of a human cell that
  1640. 52:09will allow you to figure out what
  1641. 52:11sequence of message you need to send to
  1642. 52:14uh a stem cell so that it turns into a a
  1643. 52:16beta cell. The less LLM-pilled camp and
  1644. 52:18the LLM-pilled camp talk past each
  1645. 52:19other, but it's like I think it's
  1646. 52:21actually very possible that both
  1647. 52:23what LLMs can do, which is maybe scaling
  1648. 52:25what a top doctor, the treatment you get
  1649. 52:27at like the top doctor or at the top
  1650. 52:29place, scaling that around the world,
  1651. 52:30like unbelievable potential impact of
  1652. 52:32that, right? If you're able to do that.
  1653. 52:34And then, you know, I think what you're
  1654. 52:35talking about, which is certainly still
  1655. 52:36on the come for for a lot of these
  1656. 52:37things, is okay, well, even better than
  1657. 52:40the top doctor. Like, how do you how do
  1658. 52:41you go do that?
  1659. 52:42>> more than just a top doctor, right?
  1660. 52:43Because I mean, what the LLM can do well
  1661. 52:45is
  1662. 52:47you know, it can it can sort of
  1663. 52:48regurgitate like knowledge that you can
  1664. 52:50read in books, mostly.
  1665. 52:52Um but if medicine was
  1666. 52:55only kind of
  1667. 52:57about accumulating
  1668. 52:59uh declarative language that declarative
  1669. 53:02knowledge that exist in books,
  1670. 53:05you can be a doctor by just reading
  1671. 53:06books. And you can be a doctor by
  1672. 53:08reading books. You have to do, you know,
  1673. 53:09residency and, you know, actually kind
  1674. 53:12of listen to the heart and like press on
  1675. 53:14the belly and things like that to, you
  1676. 53:16know, diagnose a disease or whatever it
  1677. 53:18is. Yeah, it's interesting. I'll I'll be
  1678. 53:20very curious to see whether LLMs
  1679. 53:21themselves can provide like, you know,
  1680. 53:23top-quality health care
  1681. 53:25globally. We'll have to we'll have to
  1682. 53:25check back in on that one. It seems it
  1683. 53:27seems like they're pretty pretty close.
  1684. 53:28You know, I definitely also want to hit
  1685. 53:29on your your time at Meta cuz you spent
  1686. 53:30over a decade building like one of the
  1687. 53:32most respected research labs in the
  1688. 53:33world. You know, obviously, you recently
  1689. 53:35left. As you reflect back on on the time
  1690. 53:37there, what do you think you got like
  1691. 53:38most right and most wrong in your time
  1692. 53:40running FAIR? So, the thing we got right
  1693. 53:42is
  1694. 53:44uh you know, building a a a top research
  1695. 53:47lab
  1696. 53:48that really sort of innovated, produced
  1697. 53:51a lot of the sort of basic
  1698. 53:53methods and science and tools like
  1699. 53:54PyTorch
  1700. 53:56um that are useful to the entire
  1701. 53:57industry, right?
  1702. 53:59>> [snorts]
  1703. 53:59>> Uh I mean, the entire industry is built
  1704. 54:01on PyTorch basically, except for a few
  1705. 54:03people at Google.
  1706. 54:04>> [laughter]
  1707. 54:05>> And I think a a culture of
  1708. 54:07uh
  1709. 54:08you know, openness and and
  1710. 54:11and kind of you know, scientific process
  1711. 54:14which I think is is necessary for
  1712. 54:17breakthrough innovation.
  1713. 54:19Yeah. Um because you know, there there
  1714. 54:21is a lot of there's a whole chain of
  1715. 54:23innovation, right? You have blue sky
  1716. 54:26research
  1717. 54:27uh new concepts, a lot of that takes
  1718. 54:29place in universities, some of that
  1719. 54:31takes place in advanced research labs in
  1720. 54:33industry
  1721. 54:35which can be counted on the fingers of
  1722. 54:36one hand.
  1723. 54:38Uh
  1724. 54:38you know, Google is a good one, uh
  1725. 54:40you know, fair was a good one.
  1726. 54:42Hopefully, it will still be, I'm not
  1727. 54:44sure.
  1728. 54:44Um
  1729. 54:46and you know, a few others. Then you
  1730. 54:48have okay, this is a good idea, but
  1731. 54:50let's push it forward and see see if it
  1732. 54:52can be
  1733. 54:53uh made useful, but still at the
  1734. 54:56research level. In a in a sense of
  1735. 54:59we're not going to fool ourselves. We're
  1736. 55:00not going to try to just you know, find
  1737. 55:04a solution that just works for this
  1738. 55:05problem. We we're going to see if
  1739. 55:08this technique that we imagine or we
  1740. 55:10picked up from other people in the
  1741. 55:12community
  1742. 55:13can actually be pushed and and be be
  1743. 55:16made uh practical, not as a product, but
  1744. 55:19like we can show that it beats some
  1745. 55:21record on you know, some uh
  1746. 55:24task or benchmark. And then
  1747. 55:26the next stage is for the the company
  1748. 55:29that hosts the research lab to say,
  1749. 55:30"Okay, now we're going to push the
  1750. 55:31button
  1751. 55:32devote a you know, big engineering
  1752. 55:34effort to that uh to that vision and
  1753. 55:37push it forward.
  1754. 55:39That is where a lot of projects fail.
  1755. 55:42That's That's where a lot of companies
  1756. 55:44kind of fail to pick up. Meta was
  1757. 55:46actually pretty good at this, okay? But
  1758. 55:49far from perfect.
  1759. 55:51It was not like, you know, textbook
  1760. 55:52example of how you do it wrong, like,
  1761. 55:54you know, Xerox PARC like totally
  1762. 55:56missing out on,
  1763. 55:57>> Yeah. you know, you GUI interface and,
  1764. 55:59you know, mouse and windowing systems,
  1765. 56:02right? Meta was, you know, kind of
  1766. 56:04missed a few steps, essentially. And it
  1767. 56:07And it it's partly partly just
  1768. 56:09organizational. It's partly because
  1769. 56:12uh
  1770. 56:13you need a an organization that is
  1771. 56:16pretty close to research, but not
  1772. 56:18completely a product organization, to
  1773. 56:20take the relay
  1774. 56:21of, you know, pushing the technology a
  1775. 56:23little further.
  1776. 56:25Not making product with a 3-month
  1777. 56:27deadline, but like, you know, pushing
  1778. 56:28things.
  1779. 56:30And
  1780. 56:31we had that at one point. Yeah. At at at
  1781. 56:34Facebook and Meta.
  1782. 56:35Uh and then we lost it.
  1783. 56:38And
  1784. 56:39FAIR was basically isolated within the
  1785. 56:41company, had lots of ideas that nobody
  1786. 56:44picked up on. And then in 2023, the
  1787. 56:46GenAI organization was created by
  1788. 56:48basically taking about 60 or 70
  1789. 56:52scientists and engineers from FAIR,
  1790. 56:54right? Initially, and then it built up.
  1791. 56:56Um but then it was under so much
  1792. 56:59short-term pressure that basically that
  1793. 57:01organization, GenAI, didn't have time to
  1794. 57:03talk to FAIR.
  1795. 57:05And so, instead of
  1796. 57:07being at the forefront and innovating in
  1797. 57:10LLM,
  1798. 57:11uh GenAI basically had to focus on
  1799. 57:14short-term things and become very
  1800. 57:15conservative.
  1801. 57:17And so, there was a gap, basically, in
  1802. 57:18periods of mismatch between research and
  1803. 57:22uh and the Is that kind of what happened
  1804. 57:23with Llama 4? Yeah.
  1805. 57:25Well, even with, you know, Llama 3,
  1806. 57:27starting with Llama 3.
  1807. 57:29So, Llama 1 was
  1808. 57:31a small project within fair. 2022 early
  1809. 57:332023 GenAI was
  1810. 57:36created. The Lama people were basically
  1811. 57:39moved to GenAI.
  1812. 57:41They started working on Lama 2. And then
  1813. 57:43a bunch of them realized
  1814. 57:45like I could do a startup.
  1815. 57:47So that was the genesis of Mistral.
  1816. 57:50>> Yeah. Okay. Two of the
  1817. 57:52authors of
  1818. 57:55Lama 1 basically created Mistral with
  1819. 57:57another guy
  1820. 57:58from Google. And
  1821. 58:01and you know a few people kind of left
  1822. 58:03and sort of did other things.
  1823. 58:05This is not a kind of a happy time at uh
  1824. 58:08at Meta for various reasons. And so
  1825. 58:10there were you know a bunch of people
  1826. 58:12kind of left. And then the the the GenAI
  1827. 58:15organization was kind of took over
  1828. 58:17uh
  1829. 58:18Lama
  1830. 58:202 to some extent and Lama 3 and 4 was
  1831. 58:23under so much short-term pressure that
  1832. 58:25they became very conservative.
  1833. 58:27And you know it's a combination of what
  1834. 58:30is apparently of of the groups but but
  1835. 58:32like pressure from the leadership and
  1836. 58:36I mean there's many ways things can go
  1837. 58:37wrong and you can't blame anyone in
  1838. 58:39particular but
  1839. 58:40um but yeah that's kind of what
  1840. 58:43happened.
  1841. 58:43>> I mean it feels like a lot of these
  1842. 58:44organizations obviously are under
  1843. 58:46short-term pressure right now because
  1844. 58:47there's just a incredible race going on.
  1845. 58:49And so I'm curious like obviously this
  1846. 58:51this you know fair setup you had and
  1847. 58:52kind of there's similar one you know at
  1848. 58:54Google for for many years and certainly
  1849. 58:56many researchers running around Open AI
  1850. 58:57and Anthropic trying many different
  1851. 58:58things. Do you think like
  1852. 59:01that is still possible going forward or
  1853. 59:03like is the only you know is one of the
  1854. 59:04only paths to leave and and do your own
  1855. 59:06company or or you know are there still
  1856. 59:08places within the industry that you
  1857. 59:10think have this like original ethos of
  1858. 59:12fair even amidst the race that is race
  1859. 59:14dynamics that are happening? I think
  1860. 59:15there are a few places within Google
  1861. 59:17research and DeepMind that where where
  1862. 59:19people actually do research.
  1863. 59:21Um
  1864. 59:22but increasingly the industry has become
  1865. 59:24more kind of closed right? I mean Google
  1866. 59:26has certainly climbed up and, you know,
  1867. 59:29Meta and Fair even is kind of going a
  1868. 59:31bit in the same direction. There are
  1869. 59:32restrictions on publication now, like
  1870. 59:35more restrictions.
  1871. 59:36Uh and so, it's still less appealing for
  1872. 59:39people who really want to kind of do
  1873. 59:41breakthrough research and, you know,
  1874. 59:43they they don't get as much resources.
  1875. 59:45If they do something that is
  1876. 59:47relevant in medium term, they are told
  1877. 59:49not to talk about it. And and so, it's
  1878. 59:51it's not, you know, it's not a good
  1879. 59:53atmosphere, I think, for for
  1880. 59:55breakthrough. It's not conducive. You
  1881. 59:57you, you know,
  1882. 59:58I mean, basically, the get the best way
  1883. 1:00:00to get breakthrough research
  1884. 1:00:03of the type that, you know, you
  1885. 1:00:05we were getting it in the early days of
  1886. 1:00:06Fair and
  1887. 1:00:08uh you know, at Bell Labs in the good
  1888. 1:00:09days and Xerox PARC is you hire the best
  1889. 1:00:12people and those are people who have a
  1890. 1:00:13good nose to know what to work on,
  1891. 1:00:16what projects to kind of attack.
  1892. 1:00:18You give them the means to succeed and
  1893. 1:00:21you get the [ __ ] out of the way. All
  1894. 1:00:23right, pardon my French.
  1895. 1:00:24>> [laughter]
  1896. 1:00:27>> Yeah, I mean, I'm curious like what you,
  1897. 1:00:28you know, what impact it then ends up
  1898. 1:00:29having on the broader research
  1899. 1:00:31community. So, obviously, one of the
  1900. 1:00:31legacies of Fair is you trained, you
  1901. 1:00:33know, uh so many researchers, right? And
  1902. 1:00:35and like they're all throughout the
  1903. 1:00:36ecosystem. Um and it feels like now the
  1904. 1:00:38maybe equivalent to those people that
  1905. 1:00:39came in younger in their careers at
  1906. 1:00:41Fair, you know, they're joining these
  1907. 1:00:42these labs uh with maybe shorter term
  1908. 1:00:45priorities and focus. And I guess I'm
  1909. 1:00:46wondering like, you know, uh in this
  1910. 1:00:48current ecosystem where it feels like a
  1911. 1:00:49lot of younger people getting into the
  1912. 1:00:51field are thrust much more into these
  1913. 1:00:53like short-term dynamics. Does that
  1914. 1:00:54change anything about the way the the
  1915. 1:00:56ecosystem evolves? Well, I mean, the
  1916. 1:00:58people who tend to want to work with me
  1917. 1:01:00are generally people who
  1918. 1:01:03uh
  1919. 1:01:04you know, sufficiently crazy to do it,
  1920. 1:01:06first of all.
  1921. 1:01:07>> Very very fair. And uh or or, you know,
  1922. 1:01:09kind of subscribe to the the whole idea
  1923. 1:01:11that
  1924. 1:01:12uh in academia and during your PhD, you
  1925. 1:01:15should work on the next generation of
  1926. 1:01:18of AI system. You shouldn't work on the
  1927. 1:01:19current generation. Yeah.
  1928. 1:01:21>> Like if you work on an in in academia
  1929. 1:01:23now, it's incredibly boring. At least to
  1930. 1:01:24me it's boring. It's basically kind of
  1931. 1:01:26studying how how and why LLMs work and
  1932. 1:01:30explaining why they work or what their
  1933. 1:01:32limitations are. It's like descriptive
  1934. 1:01:34science. It's It's really not, you know,
  1935. 1:01:36kind of creative
  1936. 1:01:37very creative. Like, I I don't find that
  1937. 1:01:40particularly interesting. It's useful.
  1938. 1:01:41Yeah. Uh
  1939. 1:01:43And, you know, if you really want to
  1940. 1:01:45kind of show how to do new things with
  1941. 1:01:47LLMs, like
  1942. 1:01:49you're not going to have the GPUs you
  1943. 1:01:50need for that. So, like, forget that.
  1944. 1:01:52Like, don't work on LLM if if you're
  1945. 1:01:54doing a PhD. Like, there's no point. You
  1946. 1:01:56cannot contribute. How do you know it
  1947. 1:01:57was time to leave Meta? It sounds like
  1948. 1:01:58it was, you know, uh you know, you were
  1949. 1:02:00thinking through some of these things
  1950. 1:02:01over a period of time. You know, was
  1951. 1:02:02there a moment that it crystallized or
  1952. 1:02:04Well, it was a combination of things,
  1953. 1:02:05right? Uh So, first of all, you have to
  1954. 1:02:07understand uh a lot of people have like
  1955. 1:02:10completely wrong idea about what my role
  1956. 1:02:12at
  1957. 1:02:13uh Facebook and Meta was. So, I joined
  1958. 1:02:15in late 2013. Really kind of started
  1959. 1:02:19early 2014. The first 4 and 1/2 years, I
  1960. 1:02:22was director of FAIR. So, I built the
  1961. 1:02:24FAIR organization,
  1962. 1:02:25uh set up the culture, hired the key
  1963. 1:02:27people,
  1964. 1:02:28and and sort of managed it.
  1965. 1:02:30Um
  1966. 1:02:32And
  1967. 1:02:33after 4 and 1/2 years, I stepped down
  1968. 1:02:35from that uh role
  1969. 1:02:38for a number of reasons. I then I became
  1970. 1:02:40chief AI scientist. Okay, so
  1971. 1:02:42uh the the
  1972. 1:02:44the reason is uh
  1973. 1:02:46you know, I was
  1974. 1:02:49basically getting close to
  1975. 1:02:52um
  1976. 1:02:53turning 60.
  1977. 1:02:54>> [laughter]
  1978. 1:02:56>> First of all, 58. And
  1979. 1:02:58uh
  1980. 1:03:00I just don't want to do management.
  1981. 1:03:02Okay. I mean, I was willing to do it for
  1982. 1:03:04a while to get the the organization
  1983. 1:03:06started, but I'm just not good at it.
  1984. 1:03:08It's not the thing I'm I'm more like a,
  1985. 1:03:10you know,
  1986. 1:03:11scientific or technical
  1987. 1:03:13visionary and engineer and scientist.
  1988. 1:03:16So,
  1989. 1:03:17uh
  1990. 1:03:18other people are much better at
  1991. 1:03:19management than I am.
  1992. 1:03:20>> [laughter]
  1993. 1:03:20>> So, I basically stepped down uh
  1994. 1:03:22you know, two other people uh Joelle
  1995. 1:03:24Pineau and uh Antoine Bordes basically
  1996. 1:03:28took over uh
  1997. 1:03:29the directorship of of FAIR. And I
  1998. 1:03:32became chief AI scientist. So, um
  1999. 1:03:35I was reporting to the CTO.
  2000. 1:03:38And uh
  2001. 1:03:39and
  2002. 1:03:40you know, had roles of
  2003. 1:03:43uh
  2004. 1:03:45you basically we're starting a research
  2005. 1:03:46project that I thought
  2006. 1:03:48was necessary because the ambition of
  2007. 1:03:50FAIR was always to build intelligent
  2008. 1:03:52systems.
  2009. 1:03:53Right? And I thought
  2010. 1:03:55you know, I put my own research in in
  2011. 1:03:57parentheses while I was running FAIR. I
  2012. 1:03:59just didn't didn't have the time.
  2013. 1:04:01And I thought it was important to
  2014. 1:04:03basically kind of
  2015. 1:04:05design the architecture of
  2016. 1:04:08of like human-level,
  2017. 1:04:10you know,
  2018. 1:04:11human-like AI systems.
  2019. 1:04:14Uh and
  2020. 1:04:16you know, I had come up with the
  2021. 1:04:19concept that this was going to be based
  2022. 1:04:21on self-supervised learning on on, you
  2023. 1:04:23know, prediction from
  2024. 1:04:25sensory signals like video, things like
  2025. 1:04:27that. I mean, this is these are old
  2026. 1:04:28ideas.
  2027. 1:04:29And uh and world models. I actually gave
  2028. 1:04:32a keynote at NeurIPS in 2016
  2029. 1:04:35where I I said like this is the way AI
  2030. 1:04:37research should go like world models
  2031. 1:04:38predict, you know, consequences of your
  2032. 1:04:40actions and plan. And I said like, you
  2033. 1:04:42know, RL is not the thing that will take
  2034. 1:04:45us there cuz it's too inefficient.
  2035. 1:04:47Supervised learning has shown its
  2036. 1:04:49limits. And so, the future is
  2037. 1:04:50self-supervised learning and world
  2038. 1:04:52models.
  2039. 1:04:53So, how do we do self-supervised
  2040. 1:04:54learning and world models? And and I
  2041. 1:04:56started a few projects on this with like
  2042. 1:04:58a few avenues that didn't pan out.
  2043. 1:05:00Uh some projects on video prediction and
  2044. 1:05:03stuff like that. And uh and then came up
  2045. 1:05:05with this uh concept that you could
  2046. 1:05:07train self-supervised learning from
  2047. 1:05:09video. Um but you have to train the
  2048. 1:05:11system to make prediction in your
  2049. 1:05:12representation space. So, that's the
  2050. 1:05:14idea of JEPA. Yeah. And if you have
  2051. 1:05:15JEPA, you can turn it into a world model
  2052. 1:05:17by making it action conditioned, and
  2053. 1:05:19then you can use it for planning. So, I
  2054. 1:05:21had this idea around 2020, and in 2022,
  2055. 1:05:23I wrote a long vision paper. So, I said,
  2056. 1:05:25"I'm just going to write a paper with my
  2057. 1:05:27entire vision, okay?
  2058. 1:05:29Spill all my secrets like I don't care.
  2059. 1:05:31Uh, but maybe that will rally a bunch of
  2060. 1:05:33people to to that vision."
  2061. 1:05:35And boy, did it work.
  2062. 1:05:37>> [laughter]
  2063. 1:05:38>> Because not only did I
  2064. 1:05:41rally, you know, a bunch of students who
  2065. 1:05:43kind of came working with me at NYU or
  2066. 1:05:45in Paris because they wanted to work on
  2067. 1:05:47this,
  2068. 1:05:48but also a whole team at at at FAIR who
  2069. 1:05:50said like,
  2070. 1:05:51"This sounds great. That's what we want
  2071. 1:05:52to work on." And then Joel Pino
  2072. 1:05:55uh, said, "Well, maybe this should be
  2073. 1:05:56like a major mission of uh,
  2074. 1:05:59of of FAIR." Uh, we called it
  2075. 1:06:01advanced machine intelligence. Yeah.
  2076. 1:06:03That was the internal name of the
  2077. 1:06:04project.
  2078. 1:06:05>> Interesting. Okay.
  2079. 1:06:06>> And they let you leave with it.
  2080. 1:06:08And now it's the name of the company.
  2081. 1:06:09Um, and you know, Mark Zuckerberg, you
  2082. 1:06:12know, kind of
  2083. 1:06:14kind of read that paper and knew what it
  2084. 1:06:16was about and subscribed to the project.
  2085. 1:06:18And Andrew Bosworth, the CTO, also.
  2086. 1:06:20And uh, Mike Schroepfer, the uh,
  2087. 1:06:23previous CTO,
  2088. 1:06:24uh, Chris Cox, who was my my direct
  2089. 1:06:27manager, chief product officer, also
  2090. 1:06:28loved the idea. So, like, you know,
  2091. 1:06:30there was a lot of support in the
  2092. 1:06:31leadership uh, about this project that
  2093. 1:06:34we internally called AMI.
  2094. 1:06:35Uh,
  2095. 1:06:37and uh, and you know, and and
  2096. 1:06:40and it started
  2097. 1:06:42really kind of working uh, for for
  2098. 1:06:45video.
  2099. 1:06:47But then,
  2100. 1:06:48you know, company kind of refocused all
  2101. 1:06:50of its effort on LLM.
  2102. 1:06:52Despite support from Mark and Andrew,
  2103. 1:06:56uh,
  2104. 1:06:57Bos, we call him Bos. Um,
  2105. 1:07:00you know, the all the layers below,
  2106. 1:07:02like,
  2107. 1:07:03didn't see the point, I think. And so,
  2108. 1:07:06politically, it sort of became a little
  2109. 1:07:08difficult. Uh the applications as I as I
  2110. 1:07:11said of
  2111. 1:07:12Japan world model are there are
  2112. 1:07:14applications in like, you know, wearable
  2113. 1:07:16agents and stuff like that, but and
  2114. 1:07:18robotics, but but Meta chose to get rid
  2115. 1:07:21of its entire robotics AI group um that
  2116. 1:07:26was led by Jitendra Malik who's now at
  2117. 1:07:28Amazon.
  2118. 1:07:29And so,
  2119. 1:07:31you know, clearly it wasn't the right
  2120. 1:07:33environment anymore. Most of the
  2121. 1:07:34applications were in industry that Meta
  2122. 1:07:36had no interest in. Uh
  2123. 1:07:39FAIR was increasingly getting pressure
  2124. 1:07:42to kind of basically help MSL with uh
  2125. 1:07:46LLMs.
  2126. 1:07:47Um so,
  2127. 1:07:49yeah, you know, it it make clear it make
  2128. 1:07:51clear. And and that, you know,
  2129. 1:07:54sort of ramming uh worked really well
  2130. 1:07:57with investors, too, because
  2131. 1:07:59when I had to raise money for Emmy,
  2132. 1:08:02everybody knew my story. And you Anybody
  2133. 1:08:04knew, you know, many investors
  2134. 1:08:07um
  2135. 1:08:08you know, staff at various VCs that read
  2136. 1:08:10my paper and or had listened to my talks
  2137. 1:08:13and had bought my story. They were
  2138. 1:08:14realizing, you know, LLMs had
  2139. 1:08:15limitations and, you know, were kind of
  2140. 1:08:20interested by the idea of like building
  2141. 1:08:22the next generation AI systems.
  2142. 1:08:24>> I guess was was like the Scale
  2143. 1:08:25acquisition like part of this catalyst
  2144. 1:08:26of of like the pure LLM focus
  2145. 1:08:28internally? Yeah, definitely. I mean,
  2146. 1:08:30there's probably some, you know, other
  2147. 1:08:31reasons to it. I think, you know, maybe
  2148. 1:08:34um
  2149. 1:08:35uh [snorts] I don't have any sort of
  2150. 1:08:36inside information to comment on this,
  2151. 1:08:38but uh
  2152. 1:08:39it's possible that Mark sees in Alex
  2153. 1:08:41kind of a potential successor to
  2154. 1:08:43himself, like a younger version of
  2155. 1:08:44himself. Yeah, I feel like that like uh
  2156. 1:08:47a lot of the popular narrative or or,
  2157. 1:08:49you know, in the media has been like,
  2158. 1:08:50oh, like, you know, when Alex comes in,
  2159. 1:08:52it then gets harder to run like a
  2160. 1:08:53research organization. You know, I don't
  2161. 1:08:54know if that the extent you felt that or
  2162. 1:08:56>> Well, okay, so here is a big
  2163. 1:08:58misconception uh about my role, my
  2164. 1:09:01relation to Alex, and how AI was run at
  2165. 1:09:04Meta.
  2166. 1:09:05I had
  2167. 1:09:06zero
  2168. 1:09:08technical contribution to Llama, like
  2169. 1:09:10none whatsoever. My one contribution to
  2170. 1:09:12Llama
  2171. 1:09:14was to argue for open-sourcing Llama 2
  2172. 1:09:16because there was a big internal debate
  2173. 1:09:17whether we should open-source. Like the
  2174. 1:09:20legal department was against it.
  2175. 1:09:22The policy department was
  2176. 1:09:24kind of against it. Uh
  2177. 1:09:27the comms department was for it. All the
  2178. 1:09:29engineering side was for it. Like Boz
  2179. 1:09:32was for it.
  2180. 1:09:33Uh so there were like enormous internal
  2181. 1:09:35discussions at a very high level, you
  2182. 1:09:36know, 40 people from Mark Zuckerberg
  2183. 1:09:39down every week for 2 hours
  2184. 1:09:41>> [laughter]
  2185. 1:09:41>> for months. So So really it was, you
  2186. 1:09:45know, kind of a a big debate internally,
  2187. 1:09:47and I really really you know, pushed um
  2188. 1:09:50argued for for the fact that uh
  2189. 1:09:53you know, and and Boz also was was very
  2190. 1:09:55vocal about it that
  2191. 1:09:57um
  2192. 1:09:58the uh
  2193. 1:09:59you know,
  2194. 1:10:00uh safety risks were basically
  2195. 1:10:02overblown.
  2196. 1:10:03Uh the opportunities to create an
  2197. 1:10:05industry were
  2198. 1:10:07extremely strong. Um
  2199. 1:10:10and that we were going to jump-start the
  2200. 1:10:11AI industry by open-sourcing Llama 2,
  2201. 1:10:13and in fact, that's exactly what
  2202. 1:10:14happened. So but I had zero contribution
  2203. 1:10:18to to Llama
  2204. 1:10:19positive or negative. Like I I didn't do
  2205. 1:10:21anything to stop it or slow it down or
  2206. 1:10:22anything. There was a lot of people
  2207. 1:10:24working on LLMs within FAIR, and it was
  2208. 1:10:26fine.
  2209. 1:10:27Uh
  2210. 1:10:28and I never said anything against it.
  2211. 1:10:30Okay.
  2212. 1:10:31Um
  2213. 1:10:32Other than saying this is not a path to
  2214. 1:10:34a human-level intelligence, but it's
  2215. 1:10:35fine. Uh
  2216. 1:10:37it's useful.
  2217. 1:10:38>> [laughter]
  2218. 1:10:39>> Uh
  2219. 1:10:40You know, same thing for speech
  2220. 1:10:41recognition or translation, right?
  2221. 1:10:43Uh so
  2222. 1:10:45uh
  2223. 1:10:46and particularly since uh 2018 when I
  2224. 1:10:48stepped down from being director of
  2225. 1:10:50FAIR,
  2226. 1:10:51uh I didn't have any direct influence on
  2227. 1:10:54what people were working on other than
  2228. 1:10:57you know, basically you publishing my my
  2229. 1:10:59vision and then rallying people
  2230. 1:11:02uh around
  2231. 1:11:03uh around my project, but
  2232. 1:11:06you know, they they were working with me
  2233. 1:11:07because they wanted not because I was
  2234. 1:11:08their boss. I wasn't telling them to
  2235. 1:11:10work with me.
  2236. 1:11:12Um
  2237. 1:11:13and so
  2238. 1:11:14um
  2239. 1:11:16so I had no positive or negative
  2240. 1:11:17influence on LLM.
  2241. 1:11:19>> [laughter]
  2242. 1:11:20>> Within
  2243. 1:11:21within Meta.
  2244. 1:11:22Uh
  2245. 1:11:23uh and uh I had some influence on the
  2246. 1:11:25strategy, but it was more like the
  2247. 1:11:27long-term and and like how how you
  2248. 1:11:29maintain a research lab and things like
  2249. 1:11:30this. And in the last uh year, you know,
  2250. 1:11:33I mean, starting maybe early '24
  2251. 1:11:36uh and certainly in '25, the the the way
  2252. 1:11:40FAIR was kind of
  2253. 1:11:42the direction in which it was moved and
  2254. 1:11:44managed basically did not correspond to
  2255. 1:11:46what I thought was necessary to preserve
  2256. 1:11:49um
  2257. 1:11:51you know, innovation, research, and
  2258. 1:11:52breakthrough and preserve the good
  2259. 1:11:54people. Like a lot of good people have
  2260. 1:11:56left already. Yeah. And I guess a lot
  2261. 1:11:58of, you know, it probably was harder to
  2262. 1:11:59get people to work on the stuff you were
  2263. 1:12:01working on internally and and I'm sure
  2264. 1:12:02there's pressure for your you yourself
  2265. 1:12:03to work on a lot of the LLM stuff. Yeah.
  2266. 1:12:06Yeah. No, but a lot of other people also
  2267. 1:12:08have left, right? No, it's it's it's
  2268. 1:12:10fascinating. I mean, one thing I'm
  2269. 1:12:11struck by throughout our whole
  2270. 1:12:11conversation is I feel like you're
  2271. 1:12:13you've like had a remarkably consistent
  2272. 1:12:15point of view like, you know, on the
  2273. 1:12:17in the space like FAIR for a long time
  2274. 1:12:19and you can go back to your, you know,
  2275. 1:12:21to a bunch of the earlier talks you
  2276. 1:12:22referenced.
  2277. 1:12:23You know, obviously it is a fast-moving
  2278. 1:12:24space and and a ton of interesting
  2279. 1:12:26things have happened in the last year.
  2280. 1:12:28What's like one thing you've changed
  2281. 1:12:29your mind on in the last year? I mean,
  2282. 1:12:30the whole idea of uh
  2283. 1:12:32what we used to call unsupervised
  2284. 1:12:33learning that we now call
  2285. 1:12:34self-supervised learning.
  2286. 1:12:36Uh you know, until about
  2287. 1:12:382003, the whole idea of
  2288. 1:12:41unsupervised pre-training
  2289. 1:12:43where you get a good representation for
  2290. 1:12:46the input data and then you either
  2291. 1:12:48fine-tune the the model with a little
  2292. 1:12:50bit of supervised labeled data. And it
  2293. 1:12:53sort of give us, you know, some evidence
  2294. 1:12:55that this whole technique could work. I
  2295. 1:12:56tried to apply this to video because
  2296. 1:12:59ultimately what I wanted to do is
  2297. 1:13:01train a system to understand how the
  2298. 1:13:02world works by just
  2299. 1:13:04watching the world go by, right? I mean,
  2300. 1:13:06that's the basic idea.
  2301. 1:13:08Uh and sort of started to argue for this
  2302. 1:13:10in the sort of
  2303. 1:13:12you know, early 2010s.
  2304. 1:13:14Um
  2305. 1:13:15did some some work on
  2306. 1:13:17simple video prediction. We didn't have
  2307. 1:13:19GPUs. Okay. Um
  2308. 1:13:22and uh
  2309. 1:13:24and then sort of doing this more
  2310. 1:13:25seriously about after the creation of
  2311. 1:13:27fair
  2312. 1:13:28um by doing pixel level video prediction
  2313. 1:13:32realizing that wasn't working. Uh but
  2314. 1:13:34then arguing for self-supervised
  2315. 1:13:35learning. Okay, this whole idea of like
  2316. 1:13:37training a system generically not to
  2317. 1:13:39solve a task but to basically just
  2318. 1:13:41predict and then using the
  2319. 1:13:42representation that is learned this way
  2320. 1:13:44as input to a downstream task that you
  2321. 1:13:47can train supervised or reinforcement or
  2322. 1:13:49whatever.
  2323. 1:13:50Uh so this that was a bit of the topic
  2324. 1:13:52of my
  2325. 1:13:53second half of my keynote at
  2326. 1:13:55at NIPS in 2016. It was still called
  2327. 1:13:57NIPS at the time.
  2328. 1:13:58>> Yeah, of course. in 2016. And then I I
  2329. 1:14:01kept kind of, you know, kind of pushing
  2330. 1:14:02for this idea and tried to kind of
  2331. 1:14:04discover some methods to to get that to
  2332. 1:14:06work. And
  2333. 1:14:08what surprised me is that that became
  2334. 1:14:10incredibly successful but not for video,
  2335. 1:14:12for language.
  2336. 1:14:13LLMs basically are
  2337. 1:14:16a a
  2338. 1:14:18blindingly successful example of
  2339. 1:14:21self-supervised learning.
  2340. 1:14:23>> No, that that they are. Well, I feel
  2341. 1:14:25like that's a that's almost like the
  2342. 1:14:25perfect note to end on but I want to
  2343. 1:14:27make sure to leave the last word to you.
  2344. 1:14:29Um I feel like there's I mean, all our
  2345. 1:14:30listeners are are very familiar with you
  2346. 1:14:32but I want to at least give you the mic
  2347. 1:14:33to point them to anything that you think
  2348. 1:14:35they should they should check out with
  2349. 1:14:36some of the new stuff you're doing or I
  2350. 1:14:37don't know, any of your your work you
  2351. 1:14:39want to point to.
  2352. 1:14:40The mic is yours. Okay, let me tell you
  2353. 1:14:43um
  2354. 1:14:44one thing, an LLM
  2355. 1:14:46works because when you have a sequence
  2356. 1:14:48of discrete symbols, making predictions
  2357. 1:14:51is easy.
  2358. 1:14:52There's only a finite number of possible
  2359. 1:14:54symbols in your language.
  2360. 1:14:57100,000 possible tokens or something
  2361. 1:14:58like that, right? And you can
  2362. 1:15:00have your neural net produce a a
  2363. 1:15:02probability distribution over all
  2364. 1:15:04possible uh tokens, and then you can
  2365. 1:15:07sample from that distribution, shift the
  2366. 1:15:09token into the input, and then produce
  2367. 1:15:10the next token, and you can do auto
  2368. 1:15:12aggressive prediction. Okay, so that's a
  2369. 1:15:14special case. If you have the real
  2370. 1:15:15world, you can't use a generative model.
  2371. 1:15:17So now you have to train a system that
  2372. 1:15:19learns a representation and makes
  2373. 1:15:20prediction in the representation space.
  2374. 1:15:22There's a big issue with this, which I
  2375. 1:15:24didn't think until about 5 years ago
  2376. 1:15:26that was
  2377. 1:15:28easily solvable,
  2378. 1:15:30even though I invented one taking to
  2379. 1:15:31solve it,
  2380. 1:15:32yeah, you know, decades before that.
  2381. 1:15:35Uh and it's a problem that
  2382. 1:15:37um
  2383. 1:15:39if you take two inputs, let's say the
  2384. 1:15:41initial segment of a video and the
  2385. 1:15:42continuation of that video, or you take
  2386. 1:15:45one image and a corrupted version of it,
  2387. 1:15:47you run them both through an encoder,
  2388. 1:15:49and you train a predictor to predict the
  2389. 1:15:51representation of one from the
  2390. 1:15:52representation of the other.
  2391. 1:15:54There's a very simple solution
  2392. 1:15:56where the system basically predicts a
  2393. 1:15:58constant representation, and now the
  2394. 1:15:59prediction problem becomes trivial.
  2395. 1:16:01That's called a collapse.
  2396. 1:16:03Representation collapse.
  2397. 1:16:05So the big question of self-supervised
  2398. 1:16:07learning for Jappa, for the joint
  2399. 1:16:08embedding architecture, is how do you
  2400. 1:16:09prevent collapse? Yeah. The solution
  2401. 1:16:11that uh I came up with many years ago,
  2402. 1:16:151993, is uh contrastive learning.
  2403. 1:16:18So basically you have
  2404. 1:16:20examples of things that should be
  2405. 1:16:22predictable from one another, and then
  2406. 1:16:23an example of things that should not be
  2407. 1:16:25predictable from one another.
  2408. 1:16:26Uh it turns out this method works, but
  2409. 1:16:29uh
  2410. 1:16:30it doesn't scale with dimension. It
  2411. 1:16:32doesn't scale very well.
  2412. 1:16:34Um there's another technique that was
  2413. 1:16:35actually invented by uh
  2414. 1:16:38Jeff Hinton and Subbiah Ecker in the
  2415. 1:16:40late '90s, late '80s, I'm sorry. Uh
  2416. 1:16:44where you have those two networks and
  2417. 1:16:45you try to maximize the mutual
  2418. 1:16:46information between them.
  2419. 1:16:48Uh
  2420. 1:16:49Jürgen Schmidhuber is mad at me because
  2421. 1:16:50he also came up with a version of this
  2422. 1:16:52[laughter]
  2423. 1:16:53in 1992
  2424. 1:16:55and he says that's JEPA. It's not JEPA.
  2425. 1:16:57It's just another way of preventing
  2426. 1:16:58collapse of a joint embedding
  2427. 1:17:00architecture. Okay.
  2428. 1:17:01Uh
  2429. 1:17:03>> [snorts]
  2430. 1:17:03>> which is
  2431. 1:17:04fine, but it's not
  2432. 1:17:06you know, it's a particular way of doing
  2433. 1:17:08it which I I don't think it's
  2434. 1:17:09particularly good. Um so
  2435. 1:17:13um
  2436. 1:17:15Okay. So, now you have the JEPA
  2437. 1:17:16architecture. You have to come up with a
  2438. 1:17:17good way of preventing collapse.
  2439. 1:17:19And there is a couple ways. So, as
  2440. 1:17:22already said, contrastive methods I
  2441. 1:17:24think is not a good uh a good approach.
  2442. 1:17:27Uh
  2443. 1:17:27there's another set of methods that are
  2444. 1:17:29kind of
  2445. 1:17:30called distillation methods.
  2446. 1:17:33And they do prevent collapse. We we
  2447. 1:17:35don't know why.
  2448. 1:17:36So, a good example of that is uh DINO or
  2449. 1:17:39DINo. Yep. Um that's a joint embedding
  2450. 1:17:42method using the distillation method.
  2451. 1:17:44Basically, one of the encoders trains
  2452. 1:17:45the other one is like used as a
  2453. 1:17:48teacher for the other encoder.
  2454. 1:17:50Uh
  2455. 1:17:51and the encoder that is being trained,
  2456. 1:17:53you do backprop to it. The one that is
  2457. 1:17:55not being trained, you don't do
  2458. 1:17:56backprop, but you share the weight with
  2459. 1:17:58the other one with some exponential
  2460. 1:17:59moving average.
  2461. 1:18:00It's a collection recipe. There was a a
  2462. 1:18:02paper from from DeepMind about it called
  2463. 1:18:04BYOL, Bootstrap Your Own Latent, which
  2464. 1:18:06uses this trick. That trick is derived
  2465. 1:18:08from some intuition from reinforcement
  2466. 1:18:10learning. And somehow it prevents
  2467. 1:18:12collapse, but we don't know why. Okay.
  2468. 1:18:14There's a few theoretical papers on it
  2469. 1:18:16that explain why it
  2470. 1:18:19possibly might work in some simple
  2471. 1:18:20cases, but it's not satisfactory. Uh
  2472. 1:18:25the function the cost function you think
  2473. 1:18:26you're minimizing, you're not actually
  2474. 1:18:28minimizing and so you can't monitor.
  2475. 1:18:30It actually goes up when you train. It
  2476. 1:18:32makes sense. So, we don't like this
  2477. 1:18:34method, but it works.
  2478. 1:18:36And some of the models we've trained,
  2479. 1:18:38large scale video representation
  2480. 1:18:40learning system, VJPA, VJPA2, VJPA2.1,
  2481. 1:18:44they train using this method.
  2482. 1:18:46Uh I Jepa also.
  2483. 1:18:48But we're moving away from this and now
  2484. 1:18:49we have uh
  2485. 1:18:51a few papers that came out recently on
  2486. 1:18:54a a specific regularizer to prevent this
  2487. 1:18:57collapse, which basically tries to
  2488. 1:18:58maximize the information content coming
  2489. 1:19:00out of the encoder. So, it's in the same
  2490. 1:19:02family as the Becker and Hinton from '89
  2491. 1:19:06and the Schmidhuber 1992 and a bunch of
  2492. 1:19:09others since then. And to some extent
  2493. 1:19:11also contrastive techniques also it's
  2494. 1:19:12not although it's not simple
  2495. 1:19:14contrastive.
  2496. 1:19:15Um
  2497. 1:19:17And then the question is how do you
  2498. 1:19:18measure information content? How do you
  2499. 1:19:19maximize
  2500. 1:19:21the information content coming out of a
  2501. 1:19:23neural net?
  2502. 1:19:24And the problem is if you want to
  2503. 1:19:25maximize the quantity, you
  2504. 1:19:27either need to be able to measure it or
  2505. 1:19:29you need to have a lower bound on it.
  2506. 1:19:31Yeah.
  2507. 1:19:32Uh information content, we only have
  2508. 1:19:33upper bounds.
  2509. 1:19:35We cannot measure it. We can only come
  2510. 1:19:37up with upper bounds. And so, we take an
  2511. 1:19:39upper bound and we cross our fingers.
  2512. 1:19:41Okay. And it kind of works. So, the
  2513. 1:19:43latest one is called SigReg.
  2514. 1:19:46That means sketch as isotropic Gaussian
  2515. 1:19:50regularization.
  2516. 1:19:51We had a previous one called
  2517. 1:19:54VCReg or VICReg, variance invariance
  2518. 1:19:56covariance regularization.
  2519. 1:19:59Um
  2520. 1:20:00And the SigReg stuff is really cool.
  2521. 1:20:02Um so, this is some work by
  2522. 1:20:04uh
  2523. 1:20:05Randall Balestriero who's uh was a
  2524. 1:20:07postdoc with me. He's
  2525. 1:20:08he's a
  2526. 1:20:09assistant professor at Brown
  2527. 1:20:11uh now. And uh it basically consists in
  2528. 1:20:14forcing the distribution of variables
  2529. 1:20:17coming out of the encoder to be
  2530. 1:20:19uh joint Gaussian essentially. So,
  2531. 1:20:21maximizing information if you want. It's
  2532. 1:20:24just a very different way of doing it
  2533. 1:20:25than, you know, what what Jürgen
  2534. 1:20:27Schmidhuber
  2535. 1:20:28>> [laughter]
  2536. 1:20:29>> and Subbarao and and Jeff Hinton were
  2537. 1:20:31doing.
  2538. 1:20:32Um
  2539. 1:20:33And so uh uh
  2540. 1:20:35This this is super promising in my
  2541. 1:20:36opinion and we have you know variations
  2542. 1:20:38of it, you know, when that we can
  2543. 1:20:39produce sparse representations.
  2544. 1:20:42Another one that can produce uh
  2545. 1:20:44uh anisotropic representations but not
  2546. 1:20:46necessarily Gaussians. And we have uh
  2547. 1:20:48uh a paper with Randall
  2548. 1:20:51uh and student at at Mila uh Luca Mice
  2549. 1:20:55that
  2550. 1:20:57where we train a world model with this.
  2551. 1:20:58It's still small scale.
  2552. 1:21:00But we think it's super promising. So if
  2553. 1:21:03you want to
  2554. 1:21:05read one paper
  2555. 1:21:06read that paper. It's Le World Model l e
  2556. 1:21:09world model. Awesome. I'll definitely
  2557. 1:21:10link to it, too. Yeah. I'm I'm not
  2558. 1:21:12responsible for the name. Randall
  2559. 1:21:13[laughter] picked up the name.
  2560. 1:21:15Amazing. Well, Jan, seriously, thank you
  2561. 1:21:16so much. It is such a privilege to get
  2562. 1:21:18to spend the last bit of time with you
  2563. 1:21:21and really appreciate you coming on the
  2564. 1:21:23podcast. Thanks for having me, though.
  2565. 1:21:24It's fun. I'm Jacob Effron and this has
  2566. 1:21:26been unsupervised learning. A podcast
  2567. 1:21:28where I get to talk to the smartest
  2568. 1:21:29people on AI and ask them tons of
  2569. 1:21:32questions about what's happening with
  2570. 1:21:33models and what it means for businesses
  2571. 1:21:35in the world. As I hope is clear, I have
  2572. 1:21:36a ton of fun doing this. It's a nights
  2573. 1:21:38and weekends project in addition to my
  2574. 1:21:40day job as an investor at Red Point, but
  2575. 1:21:42our ability to get these incredible
  2576. 1:21:43guests on really comes from folks like
  2577. 1:21:46you subscribing to the podcast, sharing
  2578. 1:21:47it with friends. It's really what
  2579. 1:21:49ultimately makes this whole thing work.
  2580. 1:21:50And so please consider doing that and
  2581. 1:21:52thank you so much for your support and
  2582. 1:21:53listening. We'll see you next episode.

About this transcript

This page contains the full transcript of Yann LeCun on What Comes After LLMs by Unsupervised Learning: With Jacob Effron, generated from the public captions YouTube serves with the video. The transcript has 15,072 words across 2,582 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.