Yann LeCun on What Comes After LLMs — Transcript
Full transcript
- 0:00You're one of the godfathers of AI.
- 0:01What's your kind of view of the path of
- 0:03progress here? Five years complete world
- 0:04domination. The best way to get
- 0:06breakthrough research is you hire the
- 0:08best people and you get the [ __ ] out of
- 0:09the way.
- 0:10>> Pardon my French.
- 0:11>> You shared the Turing Award with two
- 0:12others. When did your views start
- 0:13diverging? In 2023. How do you know it
- 0:15was time to leave Meta? It sounds like
- 0:17you were thinking through some of these
- 0:17things over a period of time.
- 0:19>> There is a big misconception about my
- 0:21role, my relation to AI and how AI was
- 0:24run at Meta. What's like one thing
- 0:25you've changed your mind on in the last
- 0:26year? I mean, the whole idea of uh Yann
- 0:29LeCun is one of the godfathers of AI.
- 0:31He's an absolute legend in the field, uh
- 0:33someone I've admired for a long time.
- 0:35And so it was such a treat to get him on
- 0:37on Unsupervised Learning.
- 0:39Uh he's been a noted skeptic of of LLMs
- 0:41in many ways, and so we dug into what
- 0:43LLMs can do, what they can't do, uh some
- 0:45of the limitations he sees, and why he
- 0:47ultimately decided to pursue a different
- 0:49architecture. Uh and we also talked
- 0:51about his time at Meta,
- 0:52um you know, the things he's proud of in
- 0:53in setting up FAIR, how the last few
- 0:55years proceeded, and what ultimately led
- 0:57him to uh spin out and start his own
- 0:59company, uh me. Um I think it's just
- 1:01fascinating to get Yann's thoughts on
- 1:03everything happening in the AI ecosystem
- 1:05today, this tension between basic
- 1:07research and then pushing LLMs forward,
- 1:09and how that's happening in in a bunch
- 1:11of organizations today, as well as his
- 1:12thoughts on just where the the whole
- 1:14space is headed. Uh he's just an
- 1:16absolute giant in the field, and when I
- 1:18started this podcast, I hoped we'd get
- 1:19guests like him, so it is just such a
- 1:21treat. I think folks will really enjoy
- 1:23hearing the conversation we had. Without
- 1:24further ado, here's Yann.
- 1:29Yann, this is such a pleasure. You're
- 1:30one of the godfathers of AI. I feel like
- 1:32when I started doing this podcast years
- 1:34ago, I was really hoping we might one
- 1:36day get someone like you on. You know, I
- 1:38don't like that term because I live in
- 1:39New Jersey, when you're a godfather in
- 1:40New Jersey, it [laughter] doesn't mean
- 1:42the same thing.
- 1:43Very fair, very fair. You know,
- 1:45obviously, you know, your bet on on
- 1:46neural nets when everyone doubted them
- 1:47is legendary, and I feel like today
- 1:49you're making uh a similar bet in many
- 1:51ways against LLMs and the kind of
- 1:52predominant generative architectures
- 1:54that that so many believe in.
- 1:55Uh you've recently started a new company
- 1:57uh behind this theme. And so, you know,
- 1:59our goal today in the conversation is to
- 2:01leave our listeners with a lot for more
- 2:03information about AMI, what you're doing
- 2:04there, some of your work at Tapestry,
- 2:06um you know, why you think the rest of
- 2:08the field is is is pointed in the wrong
- 2:09direction around some of these
- 2:10generative models, and then also just
- 2:12get your reflections on the way the
- 2:13field's unfolded, your time at Meta, and
- 2:15and all that. So, you know, modest goals
- 2:16for uh for for for a single podcast
- 2:18episode. I feel it'd be great to start
- 2:20with AMI um because the company feels
- 2:22like the clearest statement of your
- 2:23technical thesis going forward. And so,
- 2:25you recently launched the company that's
- 2:26focused on world models uh and scaling
- 2:28the Jeff bar architecture, which you
- 2:29obviously pioneered uh over at Meta. And
- 2:32so, I'm wondering if you could talk a
- 2:32little bit about the origins of that
- 2:34architecture and the extent to which you
- 2:35drew inspiration from the human brain
- 2:37and the way that works. So, first of
- 2:38all, I want to say there's nothing wrong
- 2:40with
- 2:41LLMs
- 2:42in the sense of
- 2:44LLMs, you know, are the basis for a lot
- 2:46of uh
- 2:47very useful AI products that all of us
- 2:49use,
- 2:50including me. Uh
- 2:52They're great, okay, for what they do.
- 2:55They're just not a path towards
- 2:57human-level or human-like intelligence
- 3:00or even animal-like intelligence.
- 3:02Uh so, that's my claim, okay? I'm not
- 3:04saying LLMs are useless, right? I'm I'm
- 3:06just saying
- 3:07they're not a path towards I mean, you
- 3:09helped build some of the first major
- 3:10open-source ones.
- 3:11Right. Right.
- 3:12>> [laughter]
- 3:12>> Right, absolutely. So, what is uh AMI?
- 3:15So, AMI really stands for advanced
- 3:17machine intelligence. And the the the
- 3:21kind of subtitle, the motto, if you
- 3:22want, is uh AI for the real world.
- 3:26So, basically, a lot of
- 3:28you know, AI techniques that people know
- 3:30about today are good for language
- 3:32manipulation, either human language or
- 3:35computer code or mathematics or or
- 3:38legalese,
- 3:40which barely qualifies as human
- 3:42language.
- 3:43>> [laughter]
- 3:43>> Unfortunately, a lot of human language
- 3:44is for it. Right. Right, sadly.
- 3:47You know, language is very special in a
- 3:49way, and it's uh particularly
- 3:52well suited for the type of uh you know,
- 3:55architectures that
- 3:56have been so successful uh recently, the
- 3:59the you know, large language models,
- 4:01GPT-style architectures. But what about
- 4:04the real world? What about like
- 4:05understanding
- 4:07the physical world? Turns out reality is
- 4:09way more complicated than language.
- 4:12Uh because it's high-dimensional, it's
- 4:14continuous, it's noisy, it's messy.
- 4:18And uh training a system to understand
- 4:20the real world is much, much harder. So
- 4:21that's really what we're after. That's
- 4:23what I've been after for most of my
- 4:25career and really kind of,
- 4:27you know, working on in an accelerated
- 4:29fashion over the last uh 5, 6 years or
- 4:31so and making significant progress over
- 4:33the last 2 years.
- 4:35And so, it made sense to really do a
- 4:37startup around it and sort of go to into
- 4:40high gear, you know, in pushing that.
- 4:42And it became clear, you know, by the
- 4:44end of last year that
- 4:45Meta was really not the right place for
- 4:47that.
- 4:48So, which is why I left and started I
- 4:51mean, Labs. I think it's an interesting
- 4:53like, you know, trend that we're seeing
- 4:54across the board, right? Where it feels
- 4:56like um there you're there's there's
- 4:58many folks spinning out of, you know,
- 4:59either some of the large companies or
- 5:01research labs, you know, that have a
- 5:03particular direction of research they're
- 5:04excited about and
- 5:05you'd have some interesting vantage
- 5:06point of this from your time at Fair of
- 5:07this
- 5:08uh almost tension that exists between,
- 5:10you know, go pursue as many different
- 5:11research directions as possible in these
- 5:13companies versus hey, something's really
- 5:15working. This is the thing that we're
- 5:16going to sell for the next 6, 12 months.
- 5:18Like, go focus on that. You know, I'm
- 5:20curious your your thoughts on that and
- 5:21and what you kind of seen in the
- 5:22industry at large. Well, it's a strange
- 5:25uh
- 5:26trade-off. There's really two modes of
- 5:28operating, right? There's a lot of
- 5:29exploratory research, a lot of research
- 5:31directions, right? And sometimes
- 5:33something
- 5:34kind of seems to work and you you need
- 5:36to push it further. And it's not
- 5:38research anymore. I mean,
- 5:40the people working on it are still
- 5:41researchers or they're called
- 5:43researchers at least in the press, but
- 5:45uh but really it's becoming more
- 5:46engineering and pushing for for
- 5:48products, right? So,
- 5:51that happened a number of times at Meta
- 5:54because of things that were started at
- 5:56FAIR, such as seeing happened in, you
- 5:59know, early 2023, essentially. Uh, when,
- 6:03you know, Llama, which was developed at
- 6:05FAIR, Llama 1, um, was very promising.
- 6:09And, uh,
- 6:10Meta created a a whole organization,
- 6:12GenAI, to turn it into something real
- 6:15and a series of products. Uh, and
- 6:17produce, you know, Llama 2, Llama 3,
- 6:19Llama 4, which was a bit of a
- 6:21disappointment. Uh, and because, you
- 6:23know, Mark Zuckerberg was disappointed
- 6:25by it, he kind of rebooted the entire
- 6:27organization, reorganized it, and hired
- 6:30new people, etc. But, what also
- 6:32happened,
- 6:34uh, over the last year is that
- 6:36uh,
- 6:37basically the company, Meta, realized
- 6:39that,
- 6:40um, they'd fallen behind a little bit,
- 6:42and so that kind of refocused the the
- 6:45strategy on trying to catch up with the
- 6:47industry. And the sad side effect of it
- 6:51is that a lot of the exploratory
- 6:53research
- 6:54was basically not
- 6:56given high priority anymore. I mean, it
- 6:58didn't concern the stuff I was working
- 7:00on, all the Jeppa and world models,
- 7:02uh, cuz, you know, Mark himself and and
- 7:05do buzzwords, the CTO, and a bunch of
- 7:07other people in the company were really
- 7:08interested in that project and really
- 7:09believed in the long-term impact. But,
- 7:12the rest of the company was just, you
- 7:13know, totally entirely focused on LLMs,
- 7:16and made it clear to me that Meta was
- 7:18really not the the right place to push
- 7:20on that project anymore. And then we
- 7:22started to have good results, and so it
- 7:24was clear that, you know, we had to kind
- 7:26of make that transition between research
- 7:29and actually kind of uh, developing the
- 7:32technology, scaling it up, and building
- 7:33products out of it. And we realized also
- 7:35that most of the
- 7:37applications were probably
- 7:40for things that Meta was not
- 7:41particularly interested in. A lot of
- 7:44applications of the kind of stuff that
- 7:46we've been working on is in the
- 7:48industry, like manufacturing industry
- 7:50and stuff like that. Obviously, you're
- 7:52you're kind of pursuing world models and
- 7:54and and in that broader world. And I
- 7:55think there's other people that have
- 7:56come at the world model pace from a more
- 7:58like generative approach. And so I think
- 8:00you've got folks, you know, you've got
- 8:01the Google folks and Genie in the video
- 8:03models. You've got folks, you know,
- 8:04building VLAs on the robotic side.
- 8:05You've got Feifei and and kind of like
- 8:08the 3D spatial models. As you think
- 8:10about kind of the the body of of of of
- 8:12evidence that got you excited about the
- 8:14Japa models and how you kind of compare
- 8:15them to what the generative folks have
- 8:17done, you know, where do you think we
- 8:19are today in in terms of like comparing
- 8:20these architectures and approaches?
- 8:22Okay, so world model is quickly becoming
- 8:24a buzzword right now, right? Certainly
- 8:27in research, but also in industry to
- 8:28some extent. And uh and then there are
- 8:31two factions, if you want. I'm not going
- 8:33to talk about VLA because VLA
- 8:35is clearly now being seen as not going
- 8:38anywhere.
- 8:40Like it's really not working. Uh so VLA
- 8:42is, you know,
- 8:43vision language action models, right? So
- 8:45basically use the LLM technology to
- 8:48train a system to produce actions for
- 8:50like controlling a robot or something
- 8:52like this, right? So you have vision in,
- 8:54language in, action out. Maybe language
- 8:57out, too.
- 8:58Um and that's pretty much now seen as a
- 9:02failure.
- 9:03>> [laughter]
- 9:04>> Uh not being reliable enough, requiring
- 9:06too much training data, you know, things
- 9:07like that.
- 9:08Okay, then there is world models. Okay,
- 9:10so what is a world model? Uh a world
- 9:12model at a regional level is something
- 9:15that
- 9:16allows an agentic system to anticipate
- 9:19the consequences of its own actions.
- 9:22Okay, predict the consequences of its
- 9:24own actions. From my point of view, I
- 9:26cannot imagine how you can even think of
- 9:29building an agentic system without that
- 9:31system having the ability to predict the
- 9:33consequences of its actions.
- 9:35I I that's pretty essential, right? When
- 9:37we
- 9:39act in the world, we have this ability.
- 9:42And when we
- 9:43uh take an action without thinking about
- 9:45the consequences,
- 9:47we're taking a big risk. And very often,
- 9:49you know, other people think we're we're
- 9:51an idiot.
- 9:53Uh we have plenty of examples on the
- 9:55international political scene at the
- 9:57moment of people who have complete, you
- 9:59know,
- 10:00no ability to predict the
- 10:01consequences of their actions. So,
- 10:03that's the one model. That's what it is,
- 10:05right? Ability to predict the
- 10:06consequences of your own actions. If you
- 10:08If you have this ability, then you can
- 10:10plan
- 10:12a sequence of actions to
- 10:14accomplish a task, to you know, satisfy
- 10:17a goal. And you do this by
- 10:20planning, reasoning,
- 10:22uh by a process of search and
- 10:24optimization. You don't do this by
- 10:27predicting one action after the other
- 10:28autoregressively, like a real AI we do.
- 10:31Uh you do this by searching for a
- 10:33sequence of actions that will accomplish
- 10:35the task you set you set for yourself.
- 10:37So, the blueprint for this is completely
- 10:40different from what, you know, LLMs uh
- 10:42can do at the moment.
- 10:44Uh LLMs do not have the ability to
- 10:45predict the consequences of their
- 10:47actions, and they do not have any
- 10:48planning abilities. Because
- 10:50inference is by
- 10:52predicting the next token, right? It's
- 10:54not by search. Okay, so right there,
- 10:56you have the two characteristics that I
- 10:58think are essential for intelligent
- 11:00behavior.
- 11:02Ability to predict consequences of your
- 11:04actions. And second, uh ability to plan
- 11:07by optimization, by search. Um find a
- 11:10good sequence of actions that will
- 11:11produce the correct outcome. And then
- 11:13there is a third characteristic, which
- 11:15is
- 11:16uh how do you pre- how do you predict
- 11:18the consequences of your actions?
- 11:20Okay, so, you know, if uh if I have a
- 11:25water bottle in front of me. I realize
- 11:26some people would just listen to this
- 11:28and not have the picture. So, I have an
- 11:30open, uncapped water bottle in front of
- 11:33me. If I push at the bottom, it's going
- 11:35to slide on the table. If I push
- 11:38near the top, it's probably going to
- 11:39flip. We can't predict exactly
- 11:42how the
- 11:44the bottle will will fall in which
- 11:45direction.
- 11:47Uh we can't exactly predict how it's
- 11:48going to slide, you know, how the water
- 11:50will spill, you know, whether the table
- 11:52is tilted in one way and the water will
- 11:55uh
- 11:55you know, kind of flow in one direction
- 11:57or another.
- 11:58There's no way we can predict this at
- 12:00the pixel level.
- 12:01So, our mental model of the world
- 12:03predicts that at an abstract level of
- 12:06representation. So, as you were working
- 12:07on this architecture, was a lot of it
- 12:08inspired by the human brain? I mean,
- 12:10obviously, like the you know, the way
- 12:11you're articulating things is exactly
- 12:12how how we do things. Right, or at least
- 12:14by, you know, cognitive science, right?
- 12:16Whether you can sort of translate this
- 12:17into a neural architecture and things
- 12:19like this, that's there's a big gap
- 12:21there. Um okay, so that that, you know,
- 12:24certainly uh cognitive science was a bit
- 12:26of a motivation or or, you know, what uh
- 12:29psychological system two, which is this
- 12:31idea of the way you behave in sort of
- 12:34deliberate reflective behavior is that
- 12:37you do imagine, predict the consequences
- 12:39of your actions, and you plan
- 12:41uh accordingly. Contrary to system one,
- 12:44where you just act, you know, reactively
- 12:47and instinctively. So, yeah, there is an
- 12:49inspiration, but also there is a lot of
- 12:52empirical evidence
- 12:53that you don't want to generate pixels.
- 12:56Okay, I've been I've been really
- 12:58interested in that problem of
- 13:01learning models of the world by
- 13:03prediction for a very long time. And
- 13:05then had an epiphany about 5 years ago,
- 13:08realizing that
- 13:10all of the architectures that have have
- 13:12been successful to learn
- 13:14representations of images and videos
- 13:17are non-generative architectures. And
- 13:20all the generative ones basically have
- 13:21been failures, right? So,
- 13:25VAE, right? Variational autoencoders, or
- 13:27auto encoders more generally,
- 13:30uh is kind of a natural
- 13:32way to think about like learning
- 13:34abstract representations of inputs,
- 13:36right? So, you put a an image at the
- 13:38input of a
- 13:39of a neural net, and then you train it
- 13:40to just reproduce the input on its
- 13:43output.
- 13:44Uh now, with a big neural net. Now, if
- 13:47you just do it this way, your neural net
- 13:48will not do anything interesting. It
- 13:50will just learn the identity function.
- 13:51Yeah. Completely uninteresting. It
- 13:53doesn't work. Now, if you train a VAE to
- 13:55learn representations of images, you get
- 13:57something, but it's really not that
- 13:58great. Same with sparse auto encoders.
- 14:00Then, you have another set of
- 14:01techniques,
- 14:03uh and it's kind of derivative of
- 14:05something called denoising auto encoder,
- 14:07uh masked auto encoder is a version of
- 14:10this. BERT is a version of this for NLP.
- 14:12So, you take the image, you corrupt it
- 14:14in some way, and then you train this big
- 14:15neural net to recover the original uh
- 14:19the original image. There's a huge
- 14:20project that at FAIR on this called MAE,
- 14:23masked auto encoder. It was very
- 14:25disappointing.
- 14:27A lot of computation, and not not really
- 14:30great satisfying result. Simultaneously,
- 14:33uh some of the same people working on
- 14:35MAE, and and some other people
- 14:37in Paris and in New York were working on
- 14:39other techniques using
- 14:41non-generative architecture, joint
- 14:43embedding architecture. So, take an
- 14:45image, corrupt it in some way, and then
- 14:47run the two images through encoders, and
- 14:49then try to predict the representation
- 14:51of the original image from the
- 14:53representation of the corrupted one.
- 14:55Uh that's JEPA. Yeah. Okay. So, JEPA
- 14:58means joint embedding predictive
- 14:59architecture, right? So, you have one
- 15:01encoder that makes an observation,
- 15:03another encoder that makes a different
- 15:04observation. You try to predict the
- 15:06representation of the first one from the
- 15:08second one with a predictor. And those
- 15:11techniques turned out to work much
- 15:12better for representing images and
- 15:15video. So, things like DINO,
- 15:18uh DINO V1, V2, V3, um project that is
- 15:21still going on at at fair in Paris.
- 15:25Projects like I Jepa
- 15:27and then V Jepa and then before that
- 15:28there were like Sim Siam and Moco and a
- 15:31bunch of different techniques mostly
- 15:33from Meta. There was a bunch of others
- 15:34from other groups.
- 15:36Um,
- 15:37but
- 15:38that turned out to be a much better way
- 15:39of learning representations of images
- 15:42than
- 15:44predicting pixels. Yeah. And so
- 15:46it just clicked in my in my mind but you
- 15:48know, not just mine.
- 15:51That this was the way to go and
- 15:52predicting pixels was kind of a a losing
- 15:54proposition. You know, it feels like
- 15:56there's all these robotics demos that
- 15:57are released
- 15:59you know, from from some of the model
- 16:00companies that are feel increasingly
- 16:02impressive and maybe you know, seem to
- 16:04resemble things like planning and
- 16:05reasoning when you know, they maybe
- 16:07haven't seen a a room or or a specific
- 16:09and you know, a version of a task before
- 16:11and are still able to execute that task.
- 16:13You know, what would you say to our
- 16:14listeners I guess that that observe that
- 16:16stuff and feel like it feels like we're
- 16:17trending toward some real progress with
- 16:19some of the general approaches. Well,
- 16:21there is real progress and some of those
- 16:22demos are really impressive. Um,
- 16:25but
- 16:26>> [laughter]
- 16:27>> they are trained with enormous amounts
- 16:29of data collected either from
- 16:32teleoperation
- 16:34or from just you know, human action with
- 16:36things you hold in your hand that look
- 16:38like grippers.
- 16:39Grippers that you know, and you and you
- 16:41collect the data for that. Or just you
- 16:44know, tracking hands and fingers of of a
- 16:47person.
- 16:48And then translating this into kind of
- 16:50commands for for a robot. And so those
- 16:52things are trained with imitation
- 16:54learning mostly, right? And a little bit
- 16:56with you know, reinforcement learning to
- 16:58fine tune in mostly in simulation. So
- 17:01the issue with this is that you need a
- 17:03lot of data to train the systems
- 17:06to
- 17:07to imitation.
- 17:09And it
- 17:11it becomes expensive and it's a little
- 17:12brittle
- 17:14in the sense that you know, you need to
- 17:16collect lots of data for every task you
- 17:18want the robot to uh uh to solve.
- 17:21Whereas, if the system had a world model
- 17:23that allowed it to predict the
- 17:26you know,
- 17:27the outcome of an action, it would just
- 17:29plan an action to solve a new task
- 17:32without actually having to be trained
- 17:34to accomplish this task. So, the degree
- 17:37of generalization you would get with a
- 17:39world model-based system is much, much
- 17:41larger
- 17:43uh
- 17:43you know, kind of wider spectrum of of
- 17:46tasks
- 17:47with less training data that would be
- 17:49required than a a system trained with
- 17:51imitation learning and
- 17:53and you know, fine-tuning
- 17:54>> No doubt those approaches require more
- 17:56data. And I guess this question of
- 17:57generalization really is is the big
- 17:58question, right? Of you know, and I
- 17:59think you know, some folks have have uh
- 18:02have shown some results around, you
- 18:03know, uh getting better at task A helps
- 18:05with task B, but that obviously feels
- 18:06like there's still the big unanswered
- 18:08question uh you know, around those
- 18:09architectures. I mean, you get this uh
- 18:12you know, synergy between tasks. So, the
- 18:13more tasks that you train the system to
- 18:15solve, the more tasks it's being
- 18:17it's going to be able to acquire with
- 18:18with small amount of data, regardless of
- 18:21what what technique you use. But, but
- 18:22the hope with uh
- 18:24world models is that the system can
- 18:26solve new tasks at zero shot, which
- 18:28humans are completely capable of doing,
- 18:30right? And many animals as well.
- 18:32So, uh so, that's really the the hope.
- 18:35Like, you know, solving a lot more
- 18:36problems with uh
- 18:40either a small amount of training data
- 18:42or or no training data at all.
- 18:45And just a little bit of maybe, you
- 18:47know, RL style uh fine-tuning. Yeah.
- 18:49Like, you know, how
- 18:51how is it that a 17-year-old can learn
- 18:53to drive in
- 18:54like, a dozen hours or maybe 20 hours?
- 18:57Uh we have millions of hours of training
- 18:59data of
- 19:00you know, people driving cars. We still
- 19:02don't have level five self-driving cars,
- 19:04right? So, imitation learning obviously
- 19:06does not work even for just the task of
- 19:08autonomous driving. Yeah, I guess it'll
- 19:10be a race between the ability to develop
- 19:12some of those capabilities, which may
- 19:13take time and lots of data versus this
- 19:15kind of architecture. I feel like
- 19:16there's this dream of using video models
- 19:18to just generate like tons of synthetic
- 19:20data for for, you know, simulation and,
- 19:22you know, even if it's not perfect,
- 19:23these video models from a physics
- 19:24perspective, it's like helpful enough
- 19:26to, you know, improve
- 19:28robotics and in the underlying physical
- 19:30world. What have you made of some of
- 19:31those approaches? Obviously, I think
- 19:32Nvidia's been focused there. Google
- 19:33seems to be going down that road.
- 19:35>> I'm sort of asking you again the
- 19:36question,
- 19:37you know,
- 19:38why can 17-year-old launch a driving 20
- 19:41hours? You don't need millions of hours
- 19:43of demonstration. And you don't need
- 19:45synthetic data. Uh you don't need any of
- 19:47that. So, you know, I I want a system
- 19:49that can learn as fast as that. If we
- 19:51crack that, then we don't need, you
- 19:53know, generated data, right? I mean, we
- 19:55might need to train the system in
- 19:57simulation, but not with the same amount
- 19:59of uh
- 20:01uh
- 20:02you know, of time or or trials as as
- 20:04current systems require. It's really a
- 20:07question of data efficiency.
- 20:08>> You know, I was interviewing Jerry
- 20:09Tworek on the podcast. He was at OpenAI
- 20:11and spun out to start his own lab, and
- 20:13you could sense a similar tension where
- 20:14I think he actually might even agree
- 20:15that, you know, if you continued scaling
- 20:17RL the way we're scaling, you get more,
- 20:19you know, you continue getting very
- 20:20impressive results. But, I think he
- 20:22felt, "God, there's just got to be some
- 20:23like way more efficient way to do this."
- 20:25And it's interesting. It's an
- 20:26interesting tension because you could
- 20:27imagine if you're OpenAI and you know
- 20:29something is going to continue like you
- 20:30could continue scaling it and it will
- 20:31keep getting better. There's not a ton
- 20:33of incentive necessarily from a business
- 20:35perspective to do something more
- 20:36data-efficient.
- 20:37>> Right. And there's there's no incentive
- 20:39for the other companies to do anything
- 20:41different either because they're all
- 20:43chasing the same like they can't afford
- 20:45to kind of fall behind the others,
- 20:47right? So, they all work on the same
- 20:49thing. Yeah. And And there's a bit of
- 20:51this sort of, you know,
- 20:53kind of
- 20:55herd behavior
- 20:57uh
- 20:58and and, you know, in in mostly in
- 21:00Silicon Valley where everybody is
- 21:02digging the same trench. Yeah. Uh and
- 21:05you know, so I pur-
- 21:07purposely
- 21:09set up the headquarters of Amy Labs
- 21:12in Paris. Yeah.
- 21:13>> [laughter]
- 21:14>> Uh
- 21:16the American office being in New York,
- 21:18not Silicon Valley.
- 21:19>> [laughter]
- 21:19>> It's really interesting cuz I think it
- 21:20it it it points to a tension that, you
- 21:22know, it it exists in the broader
- 21:23ecosystem today where
- 21:25uh you could imagine the other side
- 21:26being sure, maybe there are more
- 21:28data-efficient methods out there, but
- 21:29like almost who cares because we can
- 21:31keep scaling what we have to to better
- 21:33and better results. And then obviously I
- 21:35think from both, you know,
- 21:37new things you can accomplish from these
- 21:39models as well as just the joy of being
- 21:41a researcher and finding these new
- 21:42things. I get why there's such an
- 21:43attraction to to to these other
- 21:45architectures as well.
- 21:46>> And it's a bet.
- 21:47But, you know, we're pretty confident
- 21:49because, you know, we we have results
- 21:51already, actually.
- 21:52>> And as you think about like the the kind
- 21:54of um the initial spaces you're most
- 21:56excited about for the Amy technology,
- 21:58like what gets you know, where do you
- 21:59think you know, the the technology goes
- 22:01and and what are you most excited about?
- 22:03Well, I mean, you know, AI for the real
- 22:04world. Um
- 22:07like, you know, can
- 22:08where is your domestic robot? Where is
- 22:10your level five self-driving car? Yeah.
- 22:12Where is uh and that's you know
- 22:14>> When am I going to get a domestic robot?
- 22:15I'm excited about this.
- 22:17Well, so this is several years down the
- 22:19line. Okay? Despite the fact that there
- 22:21is like
- 22:22huge number of companies building
- 22:24robots, none of those companies
- 22:26actually has any idea how to make them
- 22:28smart enough to be useful, right? Or
- 22:29trusted around with a baby in the house
- 22:31or something or
- 22:31>> Certainly not that. Uh but but even for
- 22:34like, you know, relatively narrow
- 22:35manufacturing task, right? You know, I
- 22:37mean
- 22:38uh none of them really knows
- 22:40how how to do this reliably other than
- 22:42you know, for by imitation learning for
- 22:44a small number of tasks.
- 22:46Uh so, how how do we make those things
- 22:48useful? So, that's kind of a
- 22:50relatively long-term objective.
- 22:53Shorter term, there is a huge amount of
- 22:55applications in industry
- 22:57where you need to have
- 22:59a a system, an intelligent system that
- 23:01has the ability of
- 23:04you know, predicting what's going to
- 23:05happen if I change this
- 23:07control variable on this complex system,
- 23:10be it uh
- 23:12a jet engine, a chemical plant, a power
- 23:15plant, a some manufacturing line,
- 23:19a patient, a human cell, right? Those
- 23:22are systems
- 23:23that are sufficiently complex that you
- 23:25can't
- 23:26model their behavior with a small number
- 23:28of equations. Right? So, the traditional
- 23:30way of modeling
- 23:32does not work. And what you need to do
- 23:34is train a neural net, deep learning
- 23:37system,
- 23:38uh to to
- 23:40um you know, model the dynamics of that
- 23:42system
- 23:43from data.
- 23:44And what you get at the end is a a
- 23:46phenomenological model of of that
- 23:49uh process, of that
- 23:51uh system.
- 23:52Um and if it's action condition, then
- 23:54you get
- 23:55basically a a world model of that system
- 23:58that allows you to control it optimally
- 24:00for whatever purpose you have. And I
- 24:02think the
- 24:04number of applications of this in
- 24:06industry is mind-boggling. Where do you
- 24:08think we'll be with uh you know, general
- 24:10models over the next couple years? Are
- 24:11there like, you know, milestones you'd
- 24:13point to or like, what what's your kind
- 24:15of view of the path of progress here?
- 24:16Okay, couple of years is a little short.
- 24:18Like, 5 years, complete world
- 24:19domination, essentially. [laughter]
- 24:21Okay. So, somewhere between on the path
- 24:23to world domination in 5 years. I mean,
- 24:25this is kind of a joke, obviously, but
- 24:27uh this is a quote from Linus Torvalds,
- 24:29right? You know, when people ask him,
- 24:30"What's your goal with Linux?" He said,
- 24:32"Total world domination." [laughter]
- 24:34Um he actually managed to do that.
- 24:36>> Yeah, very fair. To first approximation,
- 24:38every computer in the world runs Linux,
- 24:40right? So, um so, that's kind of a joke.
- 24:42But but in the end, I think this is the
- 24:44blueprint for intelligent systems of the
- 24:46future.
- 24:48There still be a a small place for LLMs,
- 24:51you know, for
- 24:52like a language interface, basically.
- 24:55But uh
- 24:56but what we're designing are are systems
- 24:58that are capable of thinking. They They
- 25:00may not be capable of talking or
- 25:02listening initially,
- 25:04but they'll do the thinking.
- 25:06And then you can add the talking and
- 25:08listening
- 25:10uh on top of that. I'm sure you and the
- 25:12team are are are eagerly working to kind
- 25:14of, you know, get the early proof points
- 25:15of this. And obviously, you've already
- 25:17had some in the work you've done. How do
- 25:18you think about like the interim steps
- 25:19of what you'll be able to show on that
- 25:21path to to 5-year world domination?
- 25:23Well, so I think uh
- 25:25you know, within a year or so, um we'll
- 25:28have
- 25:29I think a a general methodology
- 25:32to train
- 25:33hierarchical world models
- 25:35on, you know, a a very wide variety of
- 25:38modalities.
- 25:40We know we can do a good job on video
- 25:42uh with some techniques that we're not
- 25:44completely happy with because they have
- 25:46some shortcomings, but
- 25:48um
- 25:49and we have
- 25:50sort of small-scale demonstration of a
- 25:53methodology that we think is
- 25:55really what we want.
- 25:57So, we need to scale that one up
- 25:59and get it to the same level of
- 26:00performance as the
- 26:03the other techniques that are not as uh
- 26:06satis- satisfying, if you want, on on
- 26:08things like video, but also on other
- 26:10types of data sets that we would get
- 26:12from industry partners. Okay, so we'll
- 26:14have
- 26:15demonstrations that we can train world
- 26:17models, perhaps action-conditioned world
- 26:18models that allow us to plan
- 26:20for uh a number of different use cases.
- 26:23Some of them will be robotics, some of
- 26:25them will be industrial process control
- 26:27of various types, maybe some of them in
- 26:29health um health care as well cuz we
- 26:32have partners in that
- 26:33>> Yeah. in that domain. And
- 26:36that should be within a year or two, 18
- 26:38months. Um
- 26:40And then we'll push the this methodology
- 26:43and those models into
- 26:45uh those use cases with partners, some
- 26:48of which are investors already, you
- 26:49know, in our company, and gain
- 26:51experience on how to kind of
- 26:54essentially build a somewhat universal
- 26:56world model if you want. I mean, you've
- 26:57obviously had this uh you know, this
- 26:59experience before of of kind of making
- 27:01this really contrarian bet on neural
- 27:02nets and and being certainly uh proven
- 27:05abundantly right uh in the in in the
- 27:06history books. I guess as you think
- 27:08about this bet which I think, you know,
- 27:09if you talk to the majority of people uh
- 27:11maybe at at at at the cutting edge of
- 27:13various parts of AI maybe would would
- 27:14say is contrarian today. In what time
- 27:16frame do you think it will become
- 27:17apparent like, you know, that this was
- 27:19right?
- 27:20I think it'll happen faster than
- 27:24expected perhaps because I mean, you can
- 27:26see that world model is already becoming
- 27:28a buzzword, right?
- 27:30At least at the research level.
- 27:32Uh
- 27:34and it's starting to kind of permeate
- 27:35into the industry. Yeah. And a lot of
- 27:37people are realizing like VLs suck
- 27:40and, you know, LLMs don't work for real
- 27:42world data. Industry has realized this
- 27:45already. Certainly on the on the
- 27:48on the user side.
- 27:50And I think because of the importance of
- 27:52the robotics industry
- 27:54um you know, a lot of people are kind of
- 27:55trying to figure out like how how do we
- 27:57how do we get there? How do we get how
- 27:58you make those robots
- 28:00uh
- 28:01useful. So so I think it's
- 28:03I think the realization that you need a
- 28:05change of paradigm is is happening as we
- 28:07speak and will become completely obvious
- 28:09to people by
- 28:11early 2027, I think. Yeah. Now, that
- 28:14doesn't mean we'll have a solution by
- 28:15then. We hope we will, but you know,
- 28:17we'll see. I guess, you know, switching
- 28:18gears to the LM side you mentioned some
- 28:20of this work you're doing with uh with
- 28:21Tapestry which I think would be really
- 28:22interesting for our listeners. And so
- 28:24maybe to speak to that a little bit.
- 28:25Okay, so this is kind of a little bit
- 28:27orthogonal to uh to ML Labs. Yeah, as if
- 28:30that wasn't enough to keep you busy.
- 28:32>> [laughter]
- 28:32>> Well, it's a it's a kind of an idea I've
- 28:34I've been uh forming over the last uh
- 28:36three years or so
- 28:38is the fact that uh
- 28:41people increasingly use AI assistants
- 28:43for various things, right? I mean, uh
- 28:45you see a decrease in the use of general
- 28:48traditional search engines and you just
- 28:50ask a question to your favorite AI
- 28:52assistant.
- 28:54Um and you know, if the plan that Meta
- 28:58and others are are
- 29:00developing of, you know, having smart
- 29:01devices like smart glasses and stuff
- 29:03like that,
- 29:04uh
- 29:05you know, is realized,
- 29:07basically you'd just be talking to your
- 29:09AI assistant, you know, by voice with,
- 29:11you know, to your smart glasses or maybe
- 29:13some other smart device.
- 29:15And so, all of your information diet
- 29:18will be mediated by AI assistants.
- 29:22And
- 29:23if you are someone, you know, somewhere
- 29:24in the world, let's say outside the US
- 29:26or China, and you have an AI assistant,
- 29:28and that AI assistant was built in
- 29:29California or
- 29:32you know, Beijing
- 29:33or Shanghai or Shenzhen, uh
- 29:37it's not good for you. Like you may
- 29:39speak a language that those systems
- 29:41really haven't been trained to handle
- 29:43particularly well.
- 29:44Uh you may have a culture that is not
- 29:47particularly well understood by people
- 29:48in Silicon Valley and China.
- 29:50Not well represented by the training
- 29:52data that is publicly available on the
- 29:54internet.
- 29:56Um
- 29:57you may have a value system that is
- 29:58absolutely not represented by
- 30:01uh you know, people building those
- 30:02models.
- 30:03And certainly you'll almost certainly
- 30:05have political opinions that are
- 30:07absolutely not represented by the
- 30:09handful of AI assistant you you might be
- 30:12able to get from the
- 30:14you know, um West Coast tech companies
- 30:16or from Chinese companies.
- 30:19So, what is the solution to this? Like
- 30:21how do you serve,
- 30:23uh you know, a farmer in India,
- 30:25uh or um even a philosopher in France,
- 30:30uh or Germany? And
- 30:33what you need is
- 30:35a platform
- 30:37which basically is uh an open free
- 30:41foundation model, LLM style,
- 30:44that is fine-tunable
- 30:46by anyone
- 30:48to cater to the interest of people
- 30:51speaking a particular language, having a
- 30:53particular culture,
- 30:55having particular
- 30:57value systems, political biases,
- 30:59uh
- 31:00creeds, whatever it is.
- 31:03And so, what you need is a wide
- 31:05diversity of AI assistants. There's a
- 31:08lot of countries around the world uh
- 31:11that are neither the US nor China, who
- 31:14absolutely want some level of
- 31:16sovereignty for AI, not just for their
- 31:18industry, but also for their citizen.
- 31:20They don't want their citizen to get
- 31:22brainwashed by
- 31:24a Chinese model or a Canadian model,
- 31:26actually.
- 31:27Uh And so,
- 31:29they want sovereignty. How do you get
- 31:31that? So, the way you get
- 31:34a platform like the an open platform
- 31:36like this to get to the frontier is you
- 31:38just train it on more and higher quality
- 31:40data than the than the proprietary
- 31:43systems.
- 31:44If you talk to
- 31:46people in India, in France, in Vietnam,
- 31:49in
- 31:51Morocco, in Switzerland,
- 31:54in Korea, Japan,
- 31:56uh
- 31:57Kazakhstan,
- 31:59everyone wants
- 32:02basically sovereignty.
- 32:04And you tell them like, you guys have
- 32:06been training your model, you know,
- 32:08locally. You don't have to share your
- 32:09data. So, that's the crucial aspect of
- 32:11Tapestry.
- 32:12You would have international contributor
- 32:15contributors to Tapestry
- 32:17contributing to training a a global
- 32:19model that would basically constitute a
- 32:23repository of all the world's knowledge
- 32:25and culture, if you want. But the
- 32:26contributors would contribute uh
- 32:30data and uh computing resources, but
- 32:33they would preserve the control on their
- 32:35data. They would not have to share their
- 32:37data with the other
- 32:38uh contributors.
- 32:40What they would contribute is
- 32:42parameter vectors. Interesting. So, it
- 32:44would be kind of a kind of federated
- 32:46learning style thing where uh you have a
- 32:48bunch of data centers. Uh
- 32:51you know, they they get the parameter
- 32:53vector from the
- 32:55the the global consensus of a model.
- 32:58Think of it as an average of all the
- 33:01all the parameter vectors of all the
- 33:02contributors, right?
- 33:04So, all the contributors uh periodically
- 33:06tell
- 33:08everyone else through maybe a central
- 33:10server, here is my parameter vector,
- 33:13what is yours? Okay.
- 33:15Uh and so, you exchange parameter
- 33:16vectors like this. And a local worker
- 33:19basically, whatever it updates its
- 33:20parameter vector, it tries to also
- 33:24makes make it as close as possible to
- 33:26the global consensus vector. So, as the
- 33:30training of this thing kind of
- 33:31progresses, all those parameter vectors
- 33:34converge towards
- 33:36like a a consensus model, essentially,
- 33:38which is kind of a repository of all
- 33:41human knowledge. Now, you have an open
- 33:43an open model
- 33:45that is as good as if it had been
- 33:47trained on all the data in the world.
- 33:50And now, you can fine-tune it for your
- 33:52own purpose, your own
- 33:55political, cultural, and linguistic
- 33:56biases.
- 33:57Whatever you want or centers of
- 33:58interest.
- 33:59And I think there is a natural force for
- 34:01this to happen
- 34:03uh because, you know, most countries
- 34:05that are not the US nor China want
- 34:09sovereignty, but also because
- 34:11uh
- 34:12AI is fast becoming a platform. And
- 34:15there is a natural tendency for
- 34:16platforms to become open.
- 34:18That's what happened with Linux,
- 34:20right? And that's what happened with the
- 34:23software infrastructure of the internet
- 34:25or the wireless network. It's all open
- 34:27source.
- 34:28Um
- 34:29it was proprietary initially, but that
- 34:32was all
- 34:33wiped out. It's a really clever way to
- 34:34get around you know it would seem that
- 34:36this trend of you know decreasing open
- 34:38source and and obviously I think there's
- 34:41been many fears that is like the closed
- 34:42source models get better they'll be held
- 34:44back and they'll be used to train the
- 34:45next generation and you know they'll
- 34:46they'll kind of be this almost like a
- 34:48scape scenario for for closed source
- 34:49models where they get you know so much
- 34:51better than than their open source
- 34:52counterparts. So remember what you know
- 34:54who the big players of the internet
- 34:56infrastructure were
- 34:58in 1990 six?
- 35:01Sun Microsystems HP
- 35:04Dell
- 35:05and a few others.
- 35:07Um so Sun Microsystems was selling you
- 35:09Solaris with their you know proprietary
- 35:11hardware.
- 35:12HP with HP-UX
- 35:14uh they were claiming you know Unix is
- 35:16so much more reliable than Windows
- 35:17you're not going to run a web server on
- 35:18Windows. Dell was doing this you know
- 35:21with Windows NT but like who is running
- 35:23Windows NT now [laughter] as a web
- 35:25server?
- 35:26All of this was totally wiped out by
- 35:28Linux like the entire internet runs on
- 35:30Linux.
- 35:31Um even Azure right? Even Microsoft
- 35:34>> [laughter]
- 35:34>> it runs Linux. So
- 35:37uh
- 35:39basically OpenAI and Anthropic
- 35:41etc. of today are the
- 35:44Sun Microsystem and HP-UX
- 35:47of yesterday.
- 35:49Yeah I mean I guess it it implicit in
- 35:51that is is obviously um you know I think
- 35:53your you know your view of like the
- 35:55limitations of of what like you know the
- 35:57these models can only get so good and so
- 36:00it'll be possible over time for for the
- 36:01open source folks to to catch up.
- 36:03They've already run out of data right? I
- 36:04mean the the
- 36:06the open openly available publicly
- 36:08available data text data
- 36:11uh
- 36:12is already all used.
- 36:13I mean there's not more of it right? So
- 36:15what what those companies are doing is
- 36:17licensing uh
- 36:19commercial copyrighted data
- 36:22or training on
- 36:24synthetic data. I guess I'm curious cuz
- 36:25obviously there's been some impressive
- 36:27results uh in the last few years that
- 36:29they that they have been able to drive,
- 36:30you know, post these large-scale
- 36:31pre-trainings. Um you know, IMO gold, uh
- 36:34you know, the meter
- 36:35task horizon benchmark keeps going up.
- 36:37Okay, that's Okay, that's very
- 36:39interesting. Now, think about those two
- 36:40domains, right? Mathematics and code.
- 36:43Those are two domains where the language
- 36:45itself is the substrate of reasoning.
- 36:49It's not the only substrate of
- 36:50reasoning, but a lot of
- 36:52when you do mathematics, right? The the
- 36:55formal way on a piece of paper, not the
- 36:57intuitive stuff, but the
- 36:58you manipulate language, right? And LLMs
- 37:01are really good at this. So,
- 37:03um
- 37:03you know, proving theorems and stuff
- 37:05like that, that's that's what LLMs are
- 37:07really good at.
- 37:08They're not so good at the sort of
- 37:11you know, coming up with uh good
- 37:12concepts and definitions and things like
- 37:14that. It it's more like, here is a
- 37:15problem, solve it. They're problem
- 37:16solvers. Mathematics is not just problem
- 37:19solving, right? Most of it
- 37:21uh is actually a creative act that those
- 37:23things don't do.
- 37:25Um
- 37:27and same for code. So, LLMs are good
- 37:29programmers. They're not software
- 37:31architects.
- 37:33They're not computer scientists, right?
- 37:36Uh
- 37:37but they can program for us.
- 37:39So, they they they're not in a in a
- 37:41state where they can just, you know,
- 37:43replace humans uh entirely. It changes
- 37:46the world of humans. So, humans now,
- 37:48you know, kind of go one level up in the
- 37:50abstraction hierarchy and
- 37:53our world is to decide what to build.
- 37:54But like building it, you know, you can
- 37:56you can get help from LLMs. But okay,
- 37:59that's the the important point is that
- 38:01uh LLMs are particularly successful at
- 38:04domains where the language itself is the
- 38:07substrate of reasoning.
- 38:08Uh not for anything else. Yeah. What
- 38:11would an LLM like need to do to convince
- 38:12you uh otherwise? So, I mean, like a
- 38:15zero-shot agentic system, right?
- 38:18You have an agentic system.
- 38:20Give it a new problem. It's not been
- 38:21trained to solve that that problem.
- 38:23Doesn't have a script for it.
- 38:25Uh
- 38:26is it going to be able to uh accomplish
- 38:28this task? That it's never been trained
- 38:30to solve.
- 38:32And unless the system has the ability of
- 38:34predicting the consequences of its
- 38:36actions and then use using use that for
- 38:38play for planning
- 38:41it's not going to be able to do it. And
- 38:42you're not going to do this with an LLM.
- 38:44You're going to do this perhaps with a
- 38:46significantly
- 38:47augmented LLM that is capable of
- 38:50you know, search and planning blah blah
- 38:52blah. And currently
- 38:54you know, LLMs that do math and code
- 38:55actually do this.
- 38:56>> Yeah. Right? Cuz they search for you
- 38:58know, sequences of tokens that actually
- 39:00accomplish a particular task and you
- 39:03know, they can run the code or verify
- 39:05that the proof is correct or whatever.
- 39:08Um so, you have like a way of checking
- 39:10whether something that's produced is is
- 39:11is correct. Um but that's not a very
- 39:14efficient way of of doing planning.
- 39:16And it only works in domains where this
- 39:19type of search can be performed in token
- 39:21space.
- 39:22What I'm talking about with Jeppa is you
- 39:24don't do this in token space. You do
- 39:25this in you know, abstract thoughts
- 39:28space. And I'm sure some people
- 39:30listening might think, well, you know,
- 39:31hey, if if even if it's inefficient and
- 39:32it works uh and it works at you know, at
- 39:35things that are done in token space,
- 39:36that's still a large part of the uh of
- 39:38the economy that
- 39:39>> I mean, if it works, it's fine. I mean,
- 39:41there's again, there's nothing wrong
- 39:42with you know, using an LLM for what
- 39:44they're good at.
- 39:45Uh it's just not a path towards human
- 39:48level AI. You're missing you know, an
- 39:50>> like a huge
- 39:51uh domain.
- 39:52>> You seem like you you know, hey, it's
- 39:53going to tap out before it can become a
- 39:55software architect, whereas I'm sure
- 39:56>> going to tap out. It's it's just going
- 39:58to have like a a limited you know,
- 40:00ability to be deployed for like an it's
- 40:02going to become like increasingly
- 40:03difficult to kind of deploy it for an
- 40:05increasingly large number
- 40:07you know, of of use cases because you're
- 40:09going to have to collect tons of
- 40:10training data for each of those use
- 40:12cases. And
- 40:14there's a basically you're not going to
- 40:16be able to make those systems completely
- 40:18reliable, you know, without
- 40:19hallucinations or or dangerous stuff or
- 40:22uh etc.
- 40:24Unless those systems have the ability to
- 40:26predict the consequences of their
- 40:27actions, which means they're going to
- 40:28have to have explicit world models.
- 40:30Yeah, so I guess it's a bet against, you
- 40:31know, the uh 100% accuracy and then also
- 40:34the generalization uh across different
- 40:36tasks.
- 40:36>> Right. I guess, you know, one thing that
- 40:38that's so interesting about the way that
- 40:39the field has developed is obviously you
- 40:41uh share the the Turing Award with two
- 40:43others and I feel like they seem much
- 40:44more convinced of like maybe the the
- 40:46power or potential threats or safety
- 40:47risks of LLMs over time. Um
- 40:50I'm wondering like when did your view
- 40:51start diverging?
- 40:53Uh in 2023.
- 40:56And what like drove that in your mind? I
- 40:58didn't change my mind.
- 40:59They changed their mind, okay?
- 41:00[laughter]
- 41:01And I just about the same time, and it
- 41:03was basically GPT-4.
- 41:05I mean, Jeff basically had
- 41:07was not connected to any of that. He was
- 41:09never really interested in
- 41:11LLMs and discovered uh
- 41:14GPT-4 or, you know, 2023 when it came
- 41:16out.
- 41:17And basically he had an epiphany and
- 41:18said, "Oh my god, those systems, you
- 41:20know, are really close to human-level
- 41:22intelligence and they have
- 41:24possibly they have subjective
- 41:25experience.
- 41:26Uh
- 41:28and he he did a a quick calculation
- 41:29saying like, "Okay, the human cortex has
- 41:32about 16 billion neurons. If you want to
- 41:35um
- 41:37do something like backprop, okay? The
- 41:39brain doesn't do
- 41:41backprop directly.
- 41:42But if it does something like backprop,
- 41:44like some sort of, you know, gradient
- 41:45estimation for some sort of objective
- 41:47function,
- 41:48you would probably need like a a network
- 41:50of a few neurons to kind of reproduce
- 41:52the functionality of a virtual neuron in
- 41:54a in a neural net."
- 41:56So he said like, "Let's assume, you
- 41:58know, maybe you need you need
- 42:00a circuit of 10
- 42:02actual neurons
- 42:03to reproduce what a a backprop neuron
- 42:05does.
- 42:07Then all of a sudden your your cortex is
- 42:09only 1.6 billion neurons.
- 42:12Oh my god, GPT-4 is really close to
- 42:13this. Okay, so maybe it's as smart you
- 42:15know, it's going to get as smart as
- 42:16humans. I do not believe in this claim
- 42:18at all. This is kind of you know, Jeff's
- 42:22uh
- 42:23way of
- 42:25saying
- 42:26Okay, basically
- 42:28I can retire. I can declare victory.
- 42:31You know,
- 42:32I searched for the learning algorithm of
- 42:34the cortex all my career.
- 42:36Uh
- 42:37maybe I didn't discover what it really
- 42:39was, but backprop seems to be like a
- 42:41good substitute for it.
- 42:42It works really well.
- 42:44And so
- 42:45maybe that's all we need. So, I can
- 42:47retire.
- 42:49Uh
- 42:50and [laughter]
- 42:51and go around the world and give talks
- 42:52about you know, the potential
- 42:54uh
- 42:56promises and dangers of uh of AI.
- 43:00Uh that's basically what you know, I
- 43:02think what his uh
- 43:03uh intellectual kind of trajectory has
- 43:06been.
- 43:07Uh
- 43:08he's much less
- 43:10vocal about the potential dangers now
- 43:12than he was
- 43:14uh
- 43:15a year or two ago.
- 43:16He kind of realized there's probably a
- 43:17way to design
- 43:19truly intelligent systems. So, first of
- 43:20all, he probably you know, he realized
- 43:22that
- 43:23current LLMs are not that smart first of
- 43:25all. And and second
- 43:27uh
- 43:28that there's probably a need for a few
- 43:30breakthroughs like conceptual
- 43:31breakthroughs before we get to
- 43:33human-like intelligence.
- 43:35And third, that the the blueprint of
- 43:37those systems will be quite different
- 43:39from LLMs and we have probably have a
- 43:42way of
- 43:43you know, making them controllable and
- 43:45things like that. Yeah. I've been saying
- 43:47this for years, but
- 43:49Okay, he's sort of discovered this
- 43:51recently. Yeah. Same kind of there's a
- 43:53similar thing with Yoshua. I think what
- 43:55they are both worried about is
- 43:58the ability of
- 43:59society
- 44:01and the political system to make sure
- 44:04that the benefits of of AI will be
- 44:06maximized.
- 44:07And AI would not, you know, just
- 44:10profit you know, make a few rich people
- 44:13even richer.
- 44:14And uh you know,
- 44:17accentuate inequalities and and you
- 44:20know, cause major catastrophes because
- 44:23of bad usage. Okay, this is not like the
- 44:25the doomer scenario of AI taking over
- 44:27the world. It's more bad use users. What
- 44:31seems possible with the LLMs of today?
- 44:32Which is a danger, but you know, I don't
- 44:35I don't think it's as
- 44:37apocalyptic as you know, what some
- 44:39people have claimed it is.
- 44:41Certainly not as apocalyptic as what
- 44:43even Anthropic has claimed.
- 44:45And it's trying to kind of lobby
- 44:46governments into you know, scaring
- 44:48governments into kind of regulating AI
- 44:51because because of that.
- 44:52I don't I don't I don't subscribe to
- 44:55this at all. They seem to genuinely
- 44:56believe it. I think they genuinely
- 44:58believe it, but also I think there is,
- 45:01you know, some kind of commercial good
- 45:03commercial reasons for them to believe
- 45:04that. And to kind of
- 45:06uh
- 45:08you know, brainwash
- 45:09some people and governments into
- 45:11thinking their systems are are
- 45:12dangerous. And it sounds like you know,
- 45:14with these other architectures, do you
- 45:15think they're cuz obviously it doesn't
- 45:16you know, as maybe
- 45:18bearish as you are on LLMs being the end
- 45:20state of everything, you know, you have
- 45:22some pretty ambitious timelines too for
- 45:23for these new architectures. And so it
- 45:25doesn't seem like you think we're
- 45:26particularly far away from from from
- 45:28some very compelling capabilities. How
- 45:30do you think about I guess the the
- 45:31safety around, you know, if it ends up
- 45:33if these breakthroughs end up coming
- 45:34from new architectures and whether that
- 45:35should make us rest easier or not. I'm
- 45:38going to say something that's again
- 45:40might be controversial. Uh
- 45:42and certainly my some of my colleagues
- 45:44at Meta didn't like me saying this, but
- 45:46I think LLMs are interestingly unsafe.
- 45:49I don't think they can be made
- 45:51reliable and safe. Okay.
- 45:53They cannot be made reliable because you
- 45:55can't stop them from hallucinating.
- 45:57Uh and if they're agentic, you cannot
- 45:59guarantee they're not going to like take
- 46:01an action that, you know, they didn't
- 46:03predict the outcome of and that
- 46:05>> I mean, does it surprise you they can do
- 46:06these like 15-hour coding tasks given
- 46:08the concerns around reliability?
- 46:09>> Well, but coding is something where you
- 46:10can actually verify that, you know, the
- 46:13the the code that you generate uh you
- 46:15know, satisfy your specification.
- 46:17Um
- 46:19but
- 46:20but not everything is coding. And and
- 46:23there are examples of, you know,
- 46:25uh
- 46:26coding agents like wiping up your
- 46:28your hard drive,
- 46:30like
- 46:31uh
- 46:32or or doing stupid things, right? That
- 46:33makes you lose a lot of money or data or
- 46:36whatever.
- 46:37So, I think I think uh you know, LLMs in
- 46:40their current forms
- 46:41uh are are intrinsically unsafe
- 46:44because they cannot predict the
- 46:45consequences of their actions and
- 46:46because the way the task that they
- 46:49accomplish is determined
- 46:51uh is
- 46:53is subject to their training. You know,
- 46:56you you give them a prompt
- 46:59and then they will accomplish a task
- 47:01that correspond to that prompt only to
- 47:03the ex-
- 47:04to the extent that their training
- 47:06has conditioned them to actually do the
- 47:08right task corresponding to this prompt.
- 47:11But there is no, like, you know,
- 47:12hardwired constraint that will force
- 47:15them to accomplish this task
- 47:17and then, you know, predict that the
- 47:19task will be accomplished properly.
- 47:20Yeah, I mean, I think people were saying
- 47:21in the early days, right? They would
- 47:22you'd ask them a question and they'd
- 47:23keep asking the they'd keep asking the
- 47:24question, right?
- 47:25>> Right. Right. For example. [laughter]
- 47:27Uh
- 47:28or I mean, also they don't have common
- 47:29sense. Right. So, I mean, there's the
- 47:31the joke that was circulating like a
- 47:33month ago of
- 47:35you know, I need to wash my car and you
- 47:37know, the the car wash is a 100 yards
- 47:39from my house, should I walk?
- 47:43I tried it again like maybe 2 weeks ago.
- 47:45Uh they all say, "Yes, you should walk."
- 47:47Except Gemini.
- 47:49Gemini says
- 47:50>> they're training on your video of of
- 47:51having done having given that speech
- 47:53before? The The was not my video because
- 47:54[laughter] I I come up with this
- 47:55example.
- 47:56>> Whoever came up with it. Yeah, right.
- 47:57Whoever came up with it. But they are
- 47:58issuing sentences, right? Where where I
- 48:00said like, you know, an LLM can do this
- 48:02and then 6 months later it was like
- 48:03people are doing it and it's simply
- 48:04because, you know, as soon as people
- 48:07watch the podcast of me saying LLMs can
- 48:10do this, they of course type it into
- 48:12ChatGPT. So now it becomes part of the
- 48:14training set. And now of course, you
- 48:16know, the next version has that uh you
- 48:19know, that that thing in the fine-tuning
- 48:21set and of course it can answer the
- 48:22question but it's not because it's it
- 48:23becomes smart all of a sudden. It's just
- 48:25because it was explicitly trained with
- 48:27that question. So LLMs are intrinsically
- 48:29unsafe. Uh I don't think there is any
- 48:32way to fix that in the current
- 48:34um paradigm.
- 48:36Um and what I've been proposing is
- 48:38the architecture I've been talking about
- 48:40is objective-driven AI. So basically,
- 48:43you give an objective to an AI system,
- 48:46which is accomplish this task. Now,
- 48:48how does the system
- 48:50knows
- 48:51it will accomplish this task? It has a
- 48:53world model and it predicts
- 48:55uh
- 48:57you know, the outcome
- 48:59of a sequence of actions it imagines
- 49:01taking.
- 49:02Uh and if this uh outcome
- 49:05satisfies
- 49:07a
- 49:09cost function that you know, describes
- 49:11to what extent the task has been
- 49:12accomplished or not accomplished,
- 49:14then that system, if
- 49:16if the way that system works is by
- 49:19optimizing by optimization, finding a
- 49:22sequence of actions that accomplishes
- 49:24this task, minimizes this cost according
- 49:26to its world model,
- 49:28it can do nothing else. Yeah. Okay. And
- 49:32of course there's many things that can
- 49:33go wrong there. In in particular,
- 49:36uh the cost function might be
- 49:38inaccurate. It could be that the cost
- 49:40function you think is actually measuring
- 49:42to what extent the task has been
- 49:44accomplished but perhaps
- 49:46it's not accurate. Okay?
- 49:48Uh you the world model might be
- 49:49inaccurate. So the prediction that the
- 49:51system makes is actually not the right
- 49:52one. So, its prediction of what was
- 49:54going to happen as a consequence of its
- 49:56action wasn't right. Okay, so the system
- 49:58can still make mistakes, but but it can
- 50:01predict the consequences of its actions
- 50:02to some extent, which is I think
- 50:05indispensable for any agentic system.
- 50:07Now, you can add to that system is not
- 50:09just
- 50:10a cost function that guarantees a task
- 50:12has been accomplished, but you can also
- 50:14add a bunch of other
- 50:17objective functions, other other cost
- 50:18functions, or even constraints
- 50:21that are safety constraints.
- 50:23And say, "Okay, you know, don't hurt
- 50:24anybody on the way, right?" And you
- 50:26cannot specify this at a at an abstract
- 50:28level, but you can have, you know,
- 50:30low-level objective functions that
- 50:33put together will guarantee that the
- 50:34system will not be dangerous. Uh and the
- 50:37system cannot violate those things by
- 50:39construction. It will have to satisfy
- 50:41those conditions.
- 50:42Not the case for an LLM. The LLM can
- 50:44always escape. There's a gap between
- 50:47your training error and test error.
- 50:49There's always going to be a prompt
- 50:50where the system is going to do really
- 50:51stupid things. To talk to one specific
- 50:53space around LLMs, like, you know, I I
- 50:55think you're obviously really excited
- 50:56about LLMs in healthcare, and I think
- 50:57you know, people have been using LLMs in
- 50:59healthcare for for all sorts of things.
- 51:00And so, I'm curious how you think about
- 51:02like the set of things where LLMs are
- 51:04just not going to work in healthcare,
- 51:05and you need like a
- 51:07model that understands the world better.
- 51:09So, uh I mean, designing a course of
- 51:11treatment for a chronic disease, for
- 51:13example,
- 51:14or even a non-chronic disease,
- 51:17uh for a particular patient,
- 51:19uh which may not completely fit into,
- 51:22you know, templates that you've observed
- 51:23before.
- 51:24But if you have a good mental model of
- 51:26the
- 51:27dynamics of the physiology of the
- 51:29patient, then you might design a course
- 51:31of treatment that will actually bring
- 51:33the the patient to a good state. Yeah.
- 51:35Uh when I'm saying and when I'm saying a
- 51:36patient, it can be
- 51:38a cell. Okay, how do you tell
- 51:41a a
- 51:43stem cell
- 51:44to turn into a
- 51:46uh pancreas beta cell that produces
- 51:48insulin.
- 51:49Okay, you have a patient with type 1
- 51:51diabetes.
- 51:53Um you know, they have
- 51:55you know, their immune system basically,
- 51:57you know, kind of
- 51:59eats up their own beta cells, right?
- 52:01It's autoimmune.
- 52:02Um how do you keep making beta cells?
- 52:04You know, can you send a message? Do you
- 52:06have a a model of a of a human cell that
- 52:09will allow you to figure out what
- 52:11sequence of message you need to send to
- 52:14uh a stem cell so that it turns into a a
- 52:16beta cell. The less LLM-pilled camp and
- 52:18the LLM-pilled camp talk past each
- 52:19other, but it's like I think it's
- 52:21actually very possible that both
- 52:23what LLMs can do, which is maybe scaling
- 52:25what a top doctor, the treatment you get
- 52:27at like the top doctor or at the top
- 52:29place, scaling that around the world,
- 52:30like unbelievable potential impact of
- 52:32that, right? If you're able to do that.
- 52:34And then, you know, I think what you're
- 52:35talking about, which is certainly still
- 52:36on the come for for a lot of these
- 52:37things, is okay, well, even better than
- 52:40the top doctor. Like, how do you how do
- 52:41you go do that?
- 52:42>> more than just a top doctor, right?
- 52:43Because I mean, what the LLM can do well
- 52:45is
- 52:47you know, it can it can sort of
- 52:48regurgitate like knowledge that you can
- 52:50read in books, mostly.
- 52:52Um but if medicine was
- 52:55only kind of
- 52:57about accumulating
- 52:59uh declarative language that declarative
- 53:02knowledge that exist in books,
- 53:05you can be a doctor by just reading
- 53:06books. And you can be a doctor by
- 53:08reading books. You have to do, you know,
- 53:09residency and, you know, actually kind
- 53:12of listen to the heart and like press on
- 53:14the belly and things like that to, you
- 53:16know, diagnose a disease or whatever it
- 53:18is. Yeah, it's interesting. I'll I'll be
- 53:20very curious to see whether LLMs
- 53:21themselves can provide like, you know,
- 53:23top-quality health care
- 53:25globally. We'll have to we'll have to
- 53:25check back in on that one. It seems it
- 53:27seems like they're pretty pretty close.
- 53:28You know, I definitely also want to hit
- 53:29on your your time at Meta cuz you spent
- 53:30over a decade building like one of the
- 53:32most respected research labs in the
- 53:33world. You know, obviously, you recently
- 53:35left. As you reflect back on on the time
- 53:37there, what do you think you got like
- 53:38most right and most wrong in your time
- 53:40running FAIR? So, the thing we got right
- 53:42is
- 53:44uh you know, building a a a top research
- 53:47lab
- 53:48that really sort of innovated, produced
- 53:51a lot of the sort of basic
- 53:53methods and science and tools like
- 53:54PyTorch
- 53:56um that are useful to the entire
- 53:57industry, right?
- 53:59>> [snorts]
- 53:59>> Uh I mean, the entire industry is built
- 54:01on PyTorch basically, except for a few
- 54:03people at Google.
- 54:04>> [laughter]
- 54:05>> And I think a a culture of
- 54:07uh
- 54:08you know, openness and and
- 54:11and kind of you know, scientific process
- 54:14which I think is is necessary for
- 54:17breakthrough innovation.
- 54:19Yeah. Um because you know, there there
- 54:21is a lot of there's a whole chain of
- 54:23innovation, right? You have blue sky
- 54:26research
- 54:27uh new concepts, a lot of that takes
- 54:29place in universities, some of that
- 54:31takes place in advanced research labs in
- 54:33industry
- 54:35which can be counted on the fingers of
- 54:36one hand.
- 54:38Uh
- 54:38you know, Google is a good one, uh
- 54:40you know, fair was a good one.
- 54:42Hopefully, it will still be, I'm not
- 54:44sure.
- 54:44Um
- 54:46and you know, a few others. Then you
- 54:48have okay, this is a good idea, but
- 54:50let's push it forward and see see if it
- 54:52can be
- 54:53uh made useful, but still at the
- 54:56research level. In a in a sense of
- 54:59we're not going to fool ourselves. We're
- 55:00not going to try to just you know, find
- 55:04a solution that just works for this
- 55:05problem. We we're going to see if
- 55:08this technique that we imagine or we
- 55:10picked up from other people in the
- 55:12community
- 55:13can actually be pushed and and be be
- 55:16made uh practical, not as a product, but
- 55:19like we can show that it beats some
- 55:21record on you know, some uh
- 55:24task or benchmark. And then
- 55:26the next stage is for the the company
- 55:29that hosts the research lab to say,
- 55:30"Okay, now we're going to push the
- 55:31button
- 55:32devote a you know, big engineering
- 55:34effort to that uh to that vision and
- 55:37push it forward.
- 55:39That is where a lot of projects fail.
- 55:42That's That's where a lot of companies
- 55:44kind of fail to pick up. Meta was
- 55:46actually pretty good at this, okay? But
- 55:49far from perfect.
- 55:51It was not like, you know, textbook
- 55:52example of how you do it wrong, like,
- 55:54you know, Xerox PARC like totally
- 55:56missing out on,
- 55:57>> Yeah. you know, you GUI interface and,
- 55:59you know, mouse and windowing systems,
- 56:02right? Meta was, you know, kind of
- 56:04missed a few steps, essentially. And it
- 56:07And it it's partly partly just
- 56:09organizational. It's partly because
- 56:12uh
- 56:13you need a an organization that is
- 56:16pretty close to research, but not
- 56:18completely a product organization, to
- 56:20take the relay
- 56:21of, you know, pushing the technology a
- 56:23little further.
- 56:25Not making product with a 3-month
- 56:27deadline, but like, you know, pushing
- 56:28things.
- 56:30And
- 56:31we had that at one point. Yeah. At at at
- 56:34Facebook and Meta.
- 56:35Uh and then we lost it.
- 56:38And
- 56:39FAIR was basically isolated within the
- 56:41company, had lots of ideas that nobody
- 56:44picked up on. And then in 2023, the
- 56:46GenAI organization was created by
- 56:48basically taking about 60 or 70
- 56:52scientists and engineers from FAIR,
- 56:54right? Initially, and then it built up.
- 56:56Um but then it was under so much
- 56:59short-term pressure that basically that
- 57:01organization, GenAI, didn't have time to
- 57:03talk to FAIR.
- 57:05And so, instead of
- 57:07being at the forefront and innovating in
- 57:10LLM,
- 57:11uh GenAI basically had to focus on
- 57:14short-term things and become very
- 57:15conservative.
- 57:17And so, there was a gap, basically, in
- 57:18periods of mismatch between research and
- 57:22uh and the Is that kind of what happened
- 57:23with Llama 4? Yeah.
- 57:25Well, even with, you know, Llama 3,
- 57:27starting with Llama 3.
- 57:29So, Llama 1 was
- 57:31a small project within fair. 2022 early
- 57:332023 GenAI was
- 57:36created. The Lama people were basically
- 57:39moved to GenAI.
- 57:41They started working on Lama 2. And then
- 57:43a bunch of them realized
- 57:45like I could do a startup.
- 57:47So that was the genesis of Mistral.
- 57:50>> Yeah. Okay. Two of the
- 57:52authors of
- 57:55Lama 1 basically created Mistral with
- 57:57another guy
- 57:58from Google. And
- 58:01and you know a few people kind of left
- 58:03and sort of did other things.
- 58:05This is not a kind of a happy time at uh
- 58:08at Meta for various reasons. And so
- 58:10there were you know a bunch of people
- 58:12kind of left. And then the the the GenAI
- 58:15organization was kind of took over
- 58:17uh
- 58:18Lama
- 58:202 to some extent and Lama 3 and 4 was
- 58:23under so much short-term pressure that
- 58:25they became very conservative.
- 58:27And you know it's a combination of what
- 58:30is apparently of of the groups but but
- 58:32like pressure from the leadership and
- 58:36I mean there's many ways things can go
- 58:37wrong and you can't blame anyone in
- 58:39particular but
- 58:40um but yeah that's kind of what
- 58:43happened.
- 58:43>> I mean it feels like a lot of these
- 58:44organizations obviously are under
- 58:46short-term pressure right now because
- 58:47there's just a incredible race going on.
- 58:49And so I'm curious like obviously this
- 58:51this you know fair setup you had and
- 58:52kind of there's similar one you know at
- 58:54Google for for many years and certainly
- 58:56many researchers running around Open AI
- 58:57and Anthropic trying many different
- 58:58things. Do you think like
- 59:01that is still possible going forward or
- 59:03like is the only you know is one of the
- 59:04only paths to leave and and do your own
- 59:06company or or you know are there still
- 59:08places within the industry that you
- 59:10think have this like original ethos of
- 59:12fair even amidst the race that is race
- 59:14dynamics that are happening? I think
- 59:15there are a few places within Google
- 59:17research and DeepMind that where where
- 59:19people actually do research.
- 59:21Um
- 59:22but increasingly the industry has become
- 59:24more kind of closed right? I mean Google
- 59:26has certainly climbed up and, you know,
- 59:29Meta and Fair even is kind of going a
- 59:31bit in the same direction. There are
- 59:32restrictions on publication now, like
- 59:35more restrictions.
- 59:36Uh and so, it's still less appealing for
- 59:39people who really want to kind of do
- 59:41breakthrough research and, you know,
- 59:43they they don't get as much resources.
- 59:45If they do something that is
- 59:47relevant in medium term, they are told
- 59:49not to talk about it. And and so, it's
- 59:51it's not, you know, it's not a good
- 59:53atmosphere, I think, for for
- 59:55breakthrough. It's not conducive. You
- 59:57you, you know,
- 59:58I mean, basically, the get the best way
- 1:00:00to get breakthrough research
- 1:00:03of the type that, you know, you
- 1:00:05we were getting it in the early days of
- 1:00:06Fair and
- 1:00:08uh you know, at Bell Labs in the good
- 1:00:09days and Xerox PARC is you hire the best
- 1:00:12people and those are people who have a
- 1:00:13good nose to know what to work on,
- 1:00:16what projects to kind of attack.
- 1:00:18You give them the means to succeed and
- 1:00:21you get the [ __ ] out of the way. All
- 1:00:23right, pardon my French.
- 1:00:24>> [laughter]
- 1:00:27>> Yeah, I mean, I'm curious like what you,
- 1:00:28you know, what impact it then ends up
- 1:00:29having on the broader research
- 1:00:31community. So, obviously, one of the
- 1:00:31legacies of Fair is you trained, you
- 1:00:33know, uh so many researchers, right? And
- 1:00:35and like they're all throughout the
- 1:00:36ecosystem. Um and it feels like now the
- 1:00:38maybe equivalent to those people that
- 1:00:39came in younger in their careers at
- 1:00:41Fair, you know, they're joining these
- 1:00:42these labs uh with maybe shorter term
- 1:00:45priorities and focus. And I guess I'm
- 1:00:46wondering like, you know, uh in this
- 1:00:48current ecosystem where it feels like a
- 1:00:49lot of younger people getting into the
- 1:00:51field are thrust much more into these
- 1:00:53like short-term dynamics. Does that
- 1:00:54change anything about the way the the
- 1:00:56ecosystem evolves? Well, I mean, the
- 1:00:58people who tend to want to work with me
- 1:01:00are generally people who
- 1:01:03uh
- 1:01:04you know, sufficiently crazy to do it,
- 1:01:06first of all.
- 1:01:07>> Very very fair. And uh or or, you know,
- 1:01:09kind of subscribe to the the whole idea
- 1:01:11that
- 1:01:12uh in academia and during your PhD, you
- 1:01:15should work on the next generation of
- 1:01:18of AI system. You shouldn't work on the
- 1:01:19current generation. Yeah.
- 1:01:21>> Like if you work on an in in academia
- 1:01:23now, it's incredibly boring. At least to
- 1:01:24me it's boring. It's basically kind of
- 1:01:26studying how how and why LLMs work and
- 1:01:30explaining why they work or what their
- 1:01:32limitations are. It's like descriptive
- 1:01:34science. It's It's really not, you know,
- 1:01:36kind of creative
- 1:01:37very creative. Like, I I don't find that
- 1:01:40particularly interesting. It's useful.
- 1:01:41Yeah. Uh
- 1:01:43And, you know, if you really want to
- 1:01:45kind of show how to do new things with
- 1:01:47LLMs, like
- 1:01:49you're not going to have the GPUs you
- 1:01:50need for that. So, like, forget that.
- 1:01:52Like, don't work on LLM if if you're
- 1:01:54doing a PhD. Like, there's no point. You
- 1:01:56cannot contribute. How do you know it
- 1:01:57was time to leave Meta? It sounds like
- 1:01:58it was, you know, uh you know, you were
- 1:02:00thinking through some of these things
- 1:02:01over a period of time. You know, was
- 1:02:02there a moment that it crystallized or
- 1:02:04Well, it was a combination of things,
- 1:02:05right? Uh So, first of all, you have to
- 1:02:07understand uh a lot of people have like
- 1:02:10completely wrong idea about what my role
- 1:02:12at
- 1:02:13uh Facebook and Meta was. So, I joined
- 1:02:15in late 2013. Really kind of started
- 1:02:19early 2014. The first 4 and 1/2 years, I
- 1:02:22was director of FAIR. So, I built the
- 1:02:24FAIR organization,
- 1:02:25uh set up the culture, hired the key
- 1:02:27people,
- 1:02:28and and sort of managed it.
- 1:02:30Um
- 1:02:32And
- 1:02:33after 4 and 1/2 years, I stepped down
- 1:02:35from that uh role
- 1:02:38for a number of reasons. I then I became
- 1:02:40chief AI scientist. Okay, so
- 1:02:42uh the the
- 1:02:44the reason is uh
- 1:02:46you know, I was
- 1:02:49basically getting close to
- 1:02:52um
- 1:02:53turning 60.
- 1:02:54>> [laughter]
- 1:02:56>> First of all, 58. And
- 1:02:58uh
- 1:03:00I just don't want to do management.
- 1:03:02Okay. I mean, I was willing to do it for
- 1:03:04a while to get the the organization
- 1:03:06started, but I'm just not good at it.
- 1:03:08It's not the thing I'm I'm more like a,
- 1:03:10you know,
- 1:03:11scientific or technical
- 1:03:13visionary and engineer and scientist.
- 1:03:16So,
- 1:03:17uh
- 1:03:18other people are much better at
- 1:03:19management than I am.
- 1:03:20>> [laughter]
- 1:03:20>> So, I basically stepped down uh
- 1:03:22you know, two other people uh Joelle
- 1:03:24Pineau and uh Antoine Bordes basically
- 1:03:28took over uh
- 1:03:29the directorship of of FAIR. And I
- 1:03:32became chief AI scientist. So, um
- 1:03:35I was reporting to the CTO.
- 1:03:38And uh
- 1:03:39and
- 1:03:40you know, had roles of
- 1:03:43uh
- 1:03:45you basically we're starting a research
- 1:03:46project that I thought
- 1:03:48was necessary because the ambition of
- 1:03:50FAIR was always to build intelligent
- 1:03:52systems.
- 1:03:53Right? And I thought
- 1:03:55you know, I put my own research in in
- 1:03:57parentheses while I was running FAIR. I
- 1:03:59just didn't didn't have the time.
- 1:04:01And I thought it was important to
- 1:04:03basically kind of
- 1:04:05design the architecture of
- 1:04:08of like human-level,
- 1:04:10you know,
- 1:04:11human-like AI systems.
- 1:04:14Uh and
- 1:04:16you know, I had come up with the
- 1:04:19concept that this was going to be based
- 1:04:21on self-supervised learning on on, you
- 1:04:23know, prediction from
- 1:04:25sensory signals like video, things like
- 1:04:27that. I mean, this is these are old
- 1:04:28ideas.
- 1:04:29And uh and world models. I actually gave
- 1:04:32a keynote at NeurIPS in 2016
- 1:04:35where I I said like this is the way AI
- 1:04:37research should go like world models
- 1:04:38predict, you know, consequences of your
- 1:04:40actions and plan. And I said like, you
- 1:04:42know, RL is not the thing that will take
- 1:04:45us there cuz it's too inefficient.
- 1:04:47Supervised learning has shown its
- 1:04:49limits. And so, the future is
- 1:04:50self-supervised learning and world
- 1:04:52models.
- 1:04:53So, how do we do self-supervised
- 1:04:54learning and world models? And and I
- 1:04:56started a few projects on this with like
- 1:04:58a few avenues that didn't pan out.
- 1:05:00Uh some projects on video prediction and
- 1:05:03stuff like that. And uh and then came up
- 1:05:05with this uh concept that you could
- 1:05:07train self-supervised learning from
- 1:05:09video. Um but you have to train the
- 1:05:11system to make prediction in your
- 1:05:12representation space. So, that's the
- 1:05:14idea of JEPA. Yeah. And if you have
- 1:05:15JEPA, you can turn it into a world model
- 1:05:17by making it action conditioned, and
- 1:05:19then you can use it for planning. So, I
- 1:05:21had this idea around 2020, and in 2022,
- 1:05:23I wrote a long vision paper. So, I said,
- 1:05:25"I'm just going to write a paper with my
- 1:05:27entire vision, okay?
- 1:05:29Spill all my secrets like I don't care.
- 1:05:31Uh, but maybe that will rally a bunch of
- 1:05:33people to to that vision."
- 1:05:35And boy, did it work.
- 1:05:37>> [laughter]
- 1:05:38>> Because not only did I
- 1:05:41rally, you know, a bunch of students who
- 1:05:43kind of came working with me at NYU or
- 1:05:45in Paris because they wanted to work on
- 1:05:47this,
- 1:05:48but also a whole team at at at FAIR who
- 1:05:50said like,
- 1:05:51"This sounds great. That's what we want
- 1:05:52to work on." And then Joel Pino
- 1:05:55uh, said, "Well, maybe this should be
- 1:05:56like a major mission of uh,
- 1:05:59of of FAIR." Uh, we called it
- 1:06:01advanced machine intelligence. Yeah.
- 1:06:03That was the internal name of the
- 1:06:04project.
- 1:06:05>> Interesting. Okay.
- 1:06:06>> And they let you leave with it.
- 1:06:08And now it's the name of the company.
- 1:06:09Um, and you know, Mark Zuckerberg, you
- 1:06:12know, kind of
- 1:06:14kind of read that paper and knew what it
- 1:06:16was about and subscribed to the project.
- 1:06:18And Andrew Bosworth, the CTO, also.
- 1:06:20And uh, Mike Schroepfer, the uh,
- 1:06:23previous CTO,
- 1:06:24uh, Chris Cox, who was my my direct
- 1:06:27manager, chief product officer, also
- 1:06:28loved the idea. So, like, you know,
- 1:06:30there was a lot of support in the
- 1:06:31leadership uh, about this project that
- 1:06:34we internally called AMI.
- 1:06:35Uh,
- 1:06:37and uh, and you know, and and
- 1:06:40and it started
- 1:06:42really kind of working uh, for for
- 1:06:45video.
- 1:06:47But then,
- 1:06:48you know, company kind of refocused all
- 1:06:50of its effort on LLM.
- 1:06:52Despite support from Mark and Andrew,
- 1:06:56uh,
- 1:06:57Bos, we call him Bos. Um,
- 1:07:00you know, the all the layers below,
- 1:07:02like,
- 1:07:03didn't see the point, I think. And so,
- 1:07:06politically, it sort of became a little
- 1:07:08difficult. Uh the applications as I as I
- 1:07:11said of
- 1:07:12Japan world model are there are
- 1:07:14applications in like, you know, wearable
- 1:07:16agents and stuff like that, but and
- 1:07:18robotics, but but Meta chose to get rid
- 1:07:21of its entire robotics AI group um that
- 1:07:26was led by Jitendra Malik who's now at
- 1:07:28Amazon.
- 1:07:29And so,
- 1:07:31you know, clearly it wasn't the right
- 1:07:33environment anymore. Most of the
- 1:07:34applications were in industry that Meta
- 1:07:36had no interest in. Uh
- 1:07:39FAIR was increasingly getting pressure
- 1:07:42to kind of basically help MSL with uh
- 1:07:46LLMs.
- 1:07:47Um so,
- 1:07:49yeah, you know, it it make clear it make
- 1:07:51clear. And and that, you know,
- 1:07:54sort of ramming uh worked really well
- 1:07:57with investors, too, because
- 1:07:59when I had to raise money for Emmy,
- 1:08:02everybody knew my story. And you Anybody
- 1:08:04knew, you know, many investors
- 1:08:07um
- 1:08:08you know, staff at various VCs that read
- 1:08:10my paper and or had listened to my talks
- 1:08:13and had bought my story. They were
- 1:08:14realizing, you know, LLMs had
- 1:08:15limitations and, you know, were kind of
- 1:08:20interested by the idea of like building
- 1:08:22the next generation AI systems.
- 1:08:24>> I guess was was like the Scale
- 1:08:25acquisition like part of this catalyst
- 1:08:26of of like the pure LLM focus
- 1:08:28internally? Yeah, definitely. I mean,
- 1:08:30there's probably some, you know, other
- 1:08:31reasons to it. I think, you know, maybe
- 1:08:34um
- 1:08:35uh [snorts] I don't have any sort of
- 1:08:36inside information to comment on this,
- 1:08:38but uh
- 1:08:39it's possible that Mark sees in Alex
- 1:08:41kind of a potential successor to
- 1:08:43himself, like a younger version of
- 1:08:44himself. Yeah, I feel like that like uh
- 1:08:47a lot of the popular narrative or or,
- 1:08:49you know, in the media has been like,
- 1:08:50oh, like, you know, when Alex comes in,
- 1:08:52it then gets harder to run like a
- 1:08:53research organization. You know, I don't
- 1:08:54know if that the extent you felt that or
- 1:08:56>> Well, okay, so here is a big
- 1:08:58misconception uh about my role, my
- 1:09:01relation to Alex, and how AI was run at
- 1:09:04Meta.
- 1:09:05I had
- 1:09:06zero
- 1:09:08technical contribution to Llama, like
- 1:09:10none whatsoever. My one contribution to
- 1:09:12Llama
- 1:09:14was to argue for open-sourcing Llama 2
- 1:09:16because there was a big internal debate
- 1:09:17whether we should open-source. Like the
- 1:09:20legal department was against it.
- 1:09:22The policy department was
- 1:09:24kind of against it. Uh
- 1:09:27the comms department was for it. All the
- 1:09:29engineering side was for it. Like Boz
- 1:09:32was for it.
- 1:09:33Uh so there were like enormous internal
- 1:09:35discussions at a very high level, you
- 1:09:36know, 40 people from Mark Zuckerberg
- 1:09:39down every week for 2 hours
- 1:09:41>> [laughter]
- 1:09:41>> for months. So So really it was, you
- 1:09:45know, kind of a a big debate internally,
- 1:09:47and I really really you know, pushed um
- 1:09:50argued for for the fact that uh
- 1:09:53you know, and and Boz also was was very
- 1:09:55vocal about it that
- 1:09:57um
- 1:09:58the uh
- 1:09:59you know,
- 1:10:00uh safety risks were basically
- 1:10:02overblown.
- 1:10:03Uh the opportunities to create an
- 1:10:05industry were
- 1:10:07extremely strong. Um
- 1:10:10and that we were going to jump-start the
- 1:10:11AI industry by open-sourcing Llama 2,
- 1:10:13and in fact, that's exactly what
- 1:10:14happened. So but I had zero contribution
- 1:10:18to to Llama
- 1:10:19positive or negative. Like I I didn't do
- 1:10:21anything to stop it or slow it down or
- 1:10:22anything. There was a lot of people
- 1:10:24working on LLMs within FAIR, and it was
- 1:10:26fine.
- 1:10:27Uh
- 1:10:28and I never said anything against it.
- 1:10:30Okay.
- 1:10:31Um
- 1:10:32Other than saying this is not a path to
- 1:10:34a human-level intelligence, but it's
- 1:10:35fine. Uh
- 1:10:37it's useful.
- 1:10:38>> [laughter]
- 1:10:39>> Uh
- 1:10:40You know, same thing for speech
- 1:10:41recognition or translation, right?
- 1:10:43Uh so
- 1:10:45uh
- 1:10:46and particularly since uh 2018 when I
- 1:10:48stepped down from being director of
- 1:10:50FAIR,
- 1:10:51uh I didn't have any direct influence on
- 1:10:54what people were working on other than
- 1:10:57you know, basically you publishing my my
- 1:10:59vision and then rallying people
- 1:11:02uh around
- 1:11:03uh around my project, but
- 1:11:06you know, they they were working with me
- 1:11:07because they wanted not because I was
- 1:11:08their boss. I wasn't telling them to
- 1:11:10work with me.
- 1:11:12Um
- 1:11:13and so
- 1:11:14um
- 1:11:16so I had no positive or negative
- 1:11:17influence on LLM.
- 1:11:19>> [laughter]
- 1:11:20>> Within
- 1:11:21within Meta.
- 1:11:22Uh
- 1:11:23uh and uh I had some influence on the
- 1:11:25strategy, but it was more like the
- 1:11:27long-term and and like how how you
- 1:11:29maintain a research lab and things like
- 1:11:30this. And in the last uh year, you know,
- 1:11:33I mean, starting maybe early '24
- 1:11:36uh and certainly in '25, the the the way
- 1:11:40FAIR was kind of
- 1:11:42the direction in which it was moved and
- 1:11:44managed basically did not correspond to
- 1:11:46what I thought was necessary to preserve
- 1:11:49um
- 1:11:51you know, innovation, research, and
- 1:11:52breakthrough and preserve the good
- 1:11:54people. Like a lot of good people have
- 1:11:56left already. Yeah. And I guess a lot
- 1:11:58of, you know, it probably was harder to
- 1:11:59get people to work on the stuff you were
- 1:12:01working on internally and and I'm sure
- 1:12:02there's pressure for your you yourself
- 1:12:03to work on a lot of the LLM stuff. Yeah.
- 1:12:06Yeah. No, but a lot of other people also
- 1:12:08have left, right? No, it's it's it's
- 1:12:10fascinating. I mean, one thing I'm
- 1:12:11struck by throughout our whole
- 1:12:11conversation is I feel like you're
- 1:12:13you've like had a remarkably consistent
- 1:12:15point of view like, you know, on the
- 1:12:17in the space like FAIR for a long time
- 1:12:19and you can go back to your, you know,
- 1:12:21to a bunch of the earlier talks you
- 1:12:22referenced.
- 1:12:23You know, obviously it is a fast-moving
- 1:12:24space and and a ton of interesting
- 1:12:26things have happened in the last year.
- 1:12:28What's like one thing you've changed
- 1:12:29your mind on in the last year? I mean,
- 1:12:30the whole idea of uh
- 1:12:32what we used to call unsupervised
- 1:12:33learning that we now call
- 1:12:34self-supervised learning.
- 1:12:36Uh you know, until about
- 1:12:382003, the whole idea of
- 1:12:41unsupervised pre-training
- 1:12:43where you get a good representation for
- 1:12:46the input data and then you either
- 1:12:48fine-tune the the model with a little
- 1:12:50bit of supervised labeled data. And it
- 1:12:53sort of give us, you know, some evidence
- 1:12:55that this whole technique could work. I
- 1:12:56tried to apply this to video because
- 1:12:59ultimately what I wanted to do is
- 1:13:01train a system to understand how the
- 1:13:02world works by just
- 1:13:04watching the world go by, right? I mean,
- 1:13:06that's the basic idea.
- 1:13:08Uh and sort of started to argue for this
- 1:13:10in the sort of
- 1:13:12you know, early 2010s.
- 1:13:14Um
- 1:13:15did some some work on
- 1:13:17simple video prediction. We didn't have
- 1:13:19GPUs. Okay. Um
- 1:13:22and uh
- 1:13:24and then sort of doing this more
- 1:13:25seriously about after the creation of
- 1:13:27fair
- 1:13:28um by doing pixel level video prediction
- 1:13:32realizing that wasn't working. Uh but
- 1:13:34then arguing for self-supervised
- 1:13:35learning. Okay, this whole idea of like
- 1:13:37training a system generically not to
- 1:13:39solve a task but to basically just
- 1:13:41predict and then using the
- 1:13:42representation that is learned this way
- 1:13:44as input to a downstream task that you
- 1:13:47can train supervised or reinforcement or
- 1:13:49whatever.
- 1:13:50Uh so this that was a bit of the topic
- 1:13:52of my
- 1:13:53second half of my keynote at
- 1:13:55at NIPS in 2016. It was still called
- 1:13:57NIPS at the time.
- 1:13:58>> Yeah, of course. in 2016. And then I I
- 1:14:01kept kind of, you know, kind of pushing
- 1:14:02for this idea and tried to kind of
- 1:14:04discover some methods to to get that to
- 1:14:06work. And
- 1:14:08what surprised me is that that became
- 1:14:10incredibly successful but not for video,
- 1:14:12for language.
- 1:14:13LLMs basically are
- 1:14:16a a
- 1:14:18blindingly successful example of
- 1:14:21self-supervised learning.
- 1:14:23>> No, that that they are. Well, I feel
- 1:14:25like that's a that's almost like the
- 1:14:25perfect note to end on but I want to
- 1:14:27make sure to leave the last word to you.
- 1:14:29Um I feel like there's I mean, all our
- 1:14:30listeners are are very familiar with you
- 1:14:32but I want to at least give you the mic
- 1:14:33to point them to anything that you think
- 1:14:35they should they should check out with
- 1:14:36some of the new stuff you're doing or I
- 1:14:37don't know, any of your your work you
- 1:14:39want to point to.
- 1:14:40The mic is yours. Okay, let me tell you
- 1:14:43um
- 1:14:44one thing, an LLM
- 1:14:46works because when you have a sequence
- 1:14:48of discrete symbols, making predictions
- 1:14:51is easy.
- 1:14:52There's only a finite number of possible
- 1:14:54symbols in your language.
- 1:14:57100,000 possible tokens or something
- 1:14:58like that, right? And you can
- 1:15:00have your neural net produce a a
- 1:15:02probability distribution over all
- 1:15:04possible uh tokens, and then you can
- 1:15:07sample from that distribution, shift the
- 1:15:09token into the input, and then produce
- 1:15:10the next token, and you can do auto
- 1:15:12aggressive prediction. Okay, so that's a
- 1:15:14special case. If you have the real
- 1:15:15world, you can't use a generative model.
- 1:15:17So now you have to train a system that
- 1:15:19learns a representation and makes
- 1:15:20prediction in the representation space.
- 1:15:22There's a big issue with this, which I
- 1:15:24didn't think until about 5 years ago
- 1:15:26that was
- 1:15:28easily solvable,
- 1:15:30even though I invented one taking to
- 1:15:31solve it,
- 1:15:32yeah, you know, decades before that.
- 1:15:35Uh and it's a problem that
- 1:15:37um
- 1:15:39if you take two inputs, let's say the
- 1:15:41initial segment of a video and the
- 1:15:42continuation of that video, or you take
- 1:15:45one image and a corrupted version of it,
- 1:15:47you run them both through an encoder,
- 1:15:49and you train a predictor to predict the
- 1:15:51representation of one from the
- 1:15:52representation of the other.
- 1:15:54There's a very simple solution
- 1:15:56where the system basically predicts a
- 1:15:58constant representation, and now the
- 1:15:59prediction problem becomes trivial.
- 1:16:01That's called a collapse.
- 1:16:03Representation collapse.
- 1:16:05So the big question of self-supervised
- 1:16:07learning for Jappa, for the joint
- 1:16:08embedding architecture, is how do you
- 1:16:09prevent collapse? Yeah. The solution
- 1:16:11that uh I came up with many years ago,
- 1:16:151993, is uh contrastive learning.
- 1:16:18So basically you have
- 1:16:20examples of things that should be
- 1:16:22predictable from one another, and then
- 1:16:23an example of things that should not be
- 1:16:25predictable from one another.
- 1:16:26Uh it turns out this method works, but
- 1:16:29uh
- 1:16:30it doesn't scale with dimension. It
- 1:16:32doesn't scale very well.
- 1:16:34Um there's another technique that was
- 1:16:35actually invented by uh
- 1:16:38Jeff Hinton and Subbiah Ecker in the
- 1:16:40late '90s, late '80s, I'm sorry. Uh
- 1:16:44where you have those two networks and
- 1:16:45you try to maximize the mutual
- 1:16:46information between them.
- 1:16:48Uh
- 1:16:49Jürgen Schmidhuber is mad at me because
- 1:16:50he also came up with a version of this
- 1:16:52[laughter]
- 1:16:53in 1992
- 1:16:55and he says that's JEPA. It's not JEPA.
- 1:16:57It's just another way of preventing
- 1:16:58collapse of a joint embedding
- 1:17:00architecture. Okay.
- 1:17:01Uh
- 1:17:03>> [snorts]
- 1:17:03>> which is
- 1:17:04fine, but it's not
- 1:17:06you know, it's a particular way of doing
- 1:17:08it which I I don't think it's
- 1:17:09particularly good. Um so
- 1:17:13um
- 1:17:15Okay. So, now you have the JEPA
- 1:17:16architecture. You have to come up with a
- 1:17:17good way of preventing collapse.
- 1:17:19And there is a couple ways. So, as
- 1:17:22already said, contrastive methods I
- 1:17:24think is not a good uh a good approach.
- 1:17:27Uh
- 1:17:27there's another set of methods that are
- 1:17:29kind of
- 1:17:30called distillation methods.
- 1:17:33And they do prevent collapse. We we
- 1:17:35don't know why.
- 1:17:36So, a good example of that is uh DINO or
- 1:17:39DINo. Yep. Um that's a joint embedding
- 1:17:42method using the distillation method.
- 1:17:44Basically, one of the encoders trains
- 1:17:45the other one is like used as a
- 1:17:48teacher for the other encoder.
- 1:17:50Uh
- 1:17:51and the encoder that is being trained,
- 1:17:53you do backprop to it. The one that is
- 1:17:55not being trained, you don't do
- 1:17:56backprop, but you share the weight with
- 1:17:58the other one with some exponential
- 1:17:59moving average.
- 1:18:00It's a collection recipe. There was a a
- 1:18:02paper from from DeepMind about it called
- 1:18:04BYOL, Bootstrap Your Own Latent, which
- 1:18:06uses this trick. That trick is derived
- 1:18:08from some intuition from reinforcement
- 1:18:10learning. And somehow it prevents
- 1:18:12collapse, but we don't know why. Okay.
- 1:18:14There's a few theoretical papers on it
- 1:18:16that explain why it
- 1:18:19possibly might work in some simple
- 1:18:20cases, but it's not satisfactory. Uh
- 1:18:25the function the cost function you think
- 1:18:26you're minimizing, you're not actually
- 1:18:28minimizing and so you can't monitor.
- 1:18:30It actually goes up when you train. It
- 1:18:32makes sense. So, we don't like this
- 1:18:34method, but it works.
- 1:18:36And some of the models we've trained,
- 1:18:38large scale video representation
- 1:18:40learning system, VJPA, VJPA2, VJPA2.1,
- 1:18:44they train using this method.
- 1:18:46Uh I Jepa also.
- 1:18:48But we're moving away from this and now
- 1:18:49we have uh
- 1:18:51a few papers that came out recently on
- 1:18:54a a specific regularizer to prevent this
- 1:18:57collapse, which basically tries to
- 1:18:58maximize the information content coming
- 1:19:00out of the encoder. So, it's in the same
- 1:19:02family as the Becker and Hinton from '89
- 1:19:06and the Schmidhuber 1992 and a bunch of
- 1:19:09others since then. And to some extent
- 1:19:11also contrastive techniques also it's
- 1:19:12not although it's not simple
- 1:19:14contrastive.
- 1:19:15Um
- 1:19:17And then the question is how do you
- 1:19:18measure information content? How do you
- 1:19:19maximize
- 1:19:21the information content coming out of a
- 1:19:23neural net?
- 1:19:24And the problem is if you want to
- 1:19:25maximize the quantity, you
- 1:19:27either need to be able to measure it or
- 1:19:29you need to have a lower bound on it.
- 1:19:31Yeah.
- 1:19:32Uh information content, we only have
- 1:19:33upper bounds.
- 1:19:35We cannot measure it. We can only come
- 1:19:37up with upper bounds. And so, we take an
- 1:19:39upper bound and we cross our fingers.
- 1:19:41Okay. And it kind of works. So, the
- 1:19:43latest one is called SigReg.
- 1:19:46That means sketch as isotropic Gaussian
- 1:19:50regularization.
- 1:19:51We had a previous one called
- 1:19:54VCReg or VICReg, variance invariance
- 1:19:56covariance regularization.
- 1:19:59Um
- 1:20:00And the SigReg stuff is really cool.
- 1:20:02Um so, this is some work by
- 1:20:04uh
- 1:20:05Randall Balestriero who's uh was a
- 1:20:07postdoc with me. He's
- 1:20:08he's a
- 1:20:09assistant professor at Brown
- 1:20:11uh now. And uh it basically consists in
- 1:20:14forcing the distribution of variables
- 1:20:17coming out of the encoder to be
- 1:20:19uh joint Gaussian essentially. So,
- 1:20:21maximizing information if you want. It's
- 1:20:24just a very different way of doing it
- 1:20:25than, you know, what what Jürgen
- 1:20:27Schmidhuber
- 1:20:28>> [laughter]
- 1:20:29>> and Subbarao and and Jeff Hinton were
- 1:20:31doing.
- 1:20:32Um
- 1:20:33And so uh uh
- 1:20:35This this is super promising in my
- 1:20:36opinion and we have you know variations
- 1:20:38of it, you know, when that we can
- 1:20:39produce sparse representations.
- 1:20:42Another one that can produce uh
- 1:20:44uh anisotropic representations but not
- 1:20:46necessarily Gaussians. And we have uh
- 1:20:48uh a paper with Randall
- 1:20:51uh and student at at Mila uh Luca Mice
- 1:20:55that
- 1:20:57where we train a world model with this.
- 1:20:58It's still small scale.
- 1:21:00But we think it's super promising. So if
- 1:21:03you want to
- 1:21:05read one paper
- 1:21:06read that paper. It's Le World Model l e
- 1:21:09world model. Awesome. I'll definitely
- 1:21:10link to it, too. Yeah. I'm I'm not
- 1:21:12responsible for the name. Randall
- 1:21:13[laughter] picked up the name.
- 1:21:15Amazing. Well, Jan, seriously, thank you
- 1:21:16so much. It is such a privilege to get
- 1:21:18to spend the last bit of time with you
- 1:21:21and really appreciate you coming on the
- 1:21:23podcast. Thanks for having me, though.
- 1:21:24It's fun. I'm Jacob Effron and this has
- 1:21:26been unsupervised learning. A podcast
- 1:21:28where I get to talk to the smartest
- 1:21:29people on AI and ask them tons of
- 1:21:32questions about what's happening with
- 1:21:33models and what it means for businesses
- 1:21:35in the world. As I hope is clear, I have
- 1:21:36a ton of fun doing this. It's a nights
- 1:21:38and weekends project in addition to my
- 1:21:40day job as an investor at Red Point, but
- 1:21:42our ability to get these incredible
- 1:21:43guests on really comes from folks like
- 1:21:46you subscribing to the podcast, sharing
- 1:21:47it with friends. It's really what
- 1:21:49ultimately makes this whole thing work.
- 1:21:50And so please consider doing that and
- 1:21:52thank you so much for your support and
- 1:21:53listening. We'll see you next episode.
About this transcript
This page contains the full transcript of Yann LeCun on What Comes After LLMs by Unsupervised Learning: With Jacob Effron, generated from the public captions YouTube serves with the video. The transcript has 15,072 words across 2,582 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.