But what exactly are world models? — Transcript
Full transcript
- 0:14This is Google's V03, a state-of-the-art
- 0:17video generation model. It certainly
- 0:19took some creative liberties. There's a
- 0:22really weird Transformers situation over
- 0:24here. And what's up with these bunny
- 0:26jumps?
- 0:28General video generation models still
- 0:30don't have a good grasp of real-world
- 0:32dynamics. They understand some physical
- 0:35laws like gravity or object occlusion,
- 0:37but they can't clear the quality bar
- 0:39required by production systems like
- 0:41autonomous vehicles. For that, I hear we
- 0:44need something else. A world model.
- 0:45>> A world models.
- 0:47>> A world model. A world model. Okay,
- 0:48fine. Let's try one.
- 0:50This is Genie, Google's world model.
- 0:57It seems like world models might
- 0:59actually work.
- 1:00Now look, this buzzword is being thrown
- 1:02around in all kinds of contexts today as
- 1:05some sort of big unlock in the pursuit
- 1:08of AGI. And a lot of us are left
- 1:10wondering what exactly is a world model?
- 1:12How do you build one? And then, what do
- 1:14you do with it? In this video, we'll try
- 1:16to clear out this confusion. We'll go
- 1:18through the definition, implementations,
- 1:20and the many applications of world
- 1:23models. We'll also hear directly from TJ
- 1:25Galda, an expert from Nvidia leading
- 1:27their world model efforts.
- 1:30But first things first, what exactly is
- 1:32a world model?
- 1:34The idea dates all the way back to 1943
- 1:38when a Scottish psychologist named
- 1:41Kenneth Craik suggested that the human
- 1:43mind has an internal small-scale model
- 1:46of reality.
- 1:48It can try out various alternatives,
- 1:50conclude which is the best of them, and
- 1:52act in a much fuller, safer, and more
- 1:54competent manner.
- 1:55It's how you know that jumping in front
- 1:57a train is a bad idea, even if you've
- 2:00never tried it before.
- 2:01About 80 years later, in 2018, a paper
- 2:05called world models brought this idea to
- 2:07machine learning. An agent learning to
- 2:10play a variation of Doom leverages a
- 2:12small internal model of the game to
- 2:15imagine how its actions would play out
- 2:17in the real game. Today, prominent
- 2:20researchers agree that world models are
- 2:22essential to building AI systems that
- 2:24can plan, reason, and act safely in the
- 2:27real world. Imagine a cube floating in
- 2:29the air in front of you, and imagine
- 2:30rotating that cube by 90°. You can sort
- 2:33of picture this in your mind, and this
- 2:34has nothing to do with language.
- 2:37Humans and animals navigate the world by
- 2:39building mental models of reality. What
- 2:41if AI could develop this kind of common
- 2:42sense? An ability to make predictions of
- 2:45what's going to happen in some sort of
- 2:47abstract representation space.
- 2:49We call this concept a world model.
- 2:52In the strictest sense, in the context
- 2:54of machine learning, a world model takes
- 2:57in the current state of the world,
- 2:59together with a hypothetical action that
- 3:01could take place in it. It's goal is to
- 3:04predict the effect of this action,
- 3:06reflected in a future world state. This
- 3:09is the interface that is most faithful
- 3:11to the roots of world models in
- 3:13cognitive science.
- 3:15What's your minimal definition of a
- 3:17world model?
- 3:18>> When we say world model, we mean an AI
- 3:20system that literally learns in how the
- 3:22physical world changes over time, right?
- 3:24And so, in the same way a language model
- 3:26can predict like the next token or, you
- 3:28know, the syllable of a word,
- 3:30essentially, a world model's really
- 3:32predicting the state of the world. Not
- 3:34just what it looks like, but what it's
- 3:35going to do and how it's going to act,
- 3:37right?
- 3:38The exact implementation varies a lot
- 3:40across research labs and industry use
- 3:43cases.
- 3:44For simplicity, let's anchor this in
- 3:47autonomous vehicles.
- 3:49The current world state is captured by
- 3:51the dash cam, either as a static image
- 3:54or a short video and the intended action
- 3:57can be expressed as a text prompt like
- 3:59take a left turn. Pretty much everyone
- 4:01agrees on this input shape.
- 4:04The debate happens entirely on the
- 4:05output side and it comes down to one
- 4:08question. How faithfully should the
- 4:10world model represent the future world
- 4:12state?
- 4:14There are two schools of thought,
- 4:16generative and predictive or
- 4:19non-generative.
- 4:20A generative model outputs the future
- 4:23world state in a human-friendly form
- 4:25like a fully fledged video.
- 4:28Such world models can be used as
- 4:30standalone tools be it for synthetic
- 4:32data generation or interactive
- 4:34environments.
- 4:36Multiple prominent labs subscribe to
- 4:38this philosophy. Nvidia, Google DeepMind
- 4:41and Fei-Fei Li's World Labs.
- 4:44In this chapter, we'll look at the
- 4:46implementation of Nvidia Cosmos as a
- 4:48representative model for the generative
- 4:51family.
- 4:52We'll also visit the other two in the
- 4:54applications chapter.
- 4:56Nvidia Cosmos is a family of open-source
- 4:59models highly focused on physical
- 5:01intelligence including self-driving cars
- 5:04and autonomous robots that can operate
- 5:06in a warehouse or even in a surgery
- 5:09room.
- 5:10Officially, Cosmos is a collection of
- 5:12models that includes Cosmos Predict,
- 5:15Transfer and Reason.
- 5:17But technically, Cosmos Predict is the
- 5:20only one that abides by the strict
- 5:22definition of a world model.
- 5:25Given this video of a robot pouring the
- 5:27coffee which is the current state of the
- 5:29world, Cosmos Predict outputs the
- 5:32natural continuation where the coffee
- 5:34kettle is angled back up in a vertical
- 5:36position.
- 5:39The other two Cosmos models are
- 5:41technically not world models. Cosmos
- 5:43Transfer is well, a style transfer model
- 5:46and Cosmos Region is a visual language
- 5:49model or VLA that can answer questions
- 5:51about a given visual input.
- 5:54Nonetheless, you will hear people
- 5:55referring to all of them as world
- 5:57models, but we'll just have to live with
- 5:59this ambiguity in colloquial language.
- 6:04Cosmos Predict, the true world model of
- 6:07a family, is mostly a standard video
- 6:09diffusion model with a few tweaks. I
- 6:11have an entire YouTube series on
- 6:13diffusion models, but here's the
- 6:15high-level idea.
- 6:16The model architecture is a typical
- 6:18stack of diffusion transformers or DATs,
- 6:21which I've covered in this video.
- 6:24The DAT is basically an extension of the
- 6:26original language transformer capable of
- 6:28processing both image and text tokens.
- 6:32The current world state is compressed by
- 6:34a visual encoder into a smaller latent
- 6:37space, one slice for an image input or
- 6:40multiple slices for video frames.
- 6:42The visual encoder is an off-the-shelf
- 6:44model, in particular, the WAE 2.1 VAE
- 6:48encoder from Alibaba.
- 6:51These latent input frames depicted in
- 6:53yellow are concatenated with
- 6:54placeholders for the future frames
- 6:57initialized with pure Gaussian noise.
- 7:00The diffusion model refines them
- 7:02iteratively for a fixed number of steps,
- 7:05after which the added frames
- 7:06meaningfully capture the future.
- 7:09The clean output latent is finally
- 7:11passed through an off-the-shelf video
- 7:13decoder, also taken from WAE, which maps
- 7:16it back to pixel-based video frames.
- 7:19This denoising network is trained with a
- 7:21standard flow matching objective, which
- 7:23I've covered in a previous video.
- 7:26Now, from everything I've described so
- 7:28far, Nvidia's Cosmos Predict matches the
- 7:31structure of a general-purpose video
- 7:33generation model
- 7:35like VEO 3 from Google or SeeDance from
- 7:38ByteDance. The ones we normally use to
- 7:40generate AI slop, like food cannibalism
- 7:44or soulless Hollywood-looking movie
- 7:46trailers.
- 7:48So, what's the difference then? How is a
- 7:50video-based world model any different?
- 7:53One of the main differentiators is the
- 7:54training data. General-purpose video
- 7:57generators are trained on any sort of
- 7:59video that labs can get their hands on.
- 8:01This includes real-world footage, but
- 8:03also things like video games, cartoons,
- 8:05or even slide deck presentations. In
- 8:07contrast, Cosmos was pre-trained on
- 8:10highly curated data. We've got over 20
- 8:13million hours of physics-first data.
- 8:14We're really trying to make sure that
- 8:16the data coming in is curated and and
- 8:18trained grounded on physics. When we
- 8:21train on this, then we evaluate it on a
- 8:23whole bunch of benchmarks that also
- 8:24check for that. Cosmos Predict also
- 8:27makes a few architectural tweaks. For
- 8:29instance, it swaps out the standard
- 8:32off-the-shelf text encoder with Cosmos
- 8:35Reason, the visual language model from
- 8:37the same family.
- 8:39Normally, Cosmos Reason takes in an
- 8:41image and a text prompt, but in this
- 8:43particular setup, it will only exercise
- 8:45the text processing path.
- 8:48Just like the other models in the
- 8:49family, Cosmos Reason was also trained
- 8:52on highly curated data. Nvidia even
- 8:55built a manual ontology around
- 8:57foundational physics, time and space,
- 8:59and made sure the training data covers
- 9:01it comprehensively.
- 9:03Compared to general-purpose text
- 9:05encoders like T5, this makes Cosmos
- 9:07Reason embeddings extra sensitive to the
- 9:10difference between the egg cracked after
- 9:13being dropped versus the egg was dropped
- 9:16after cracking. At a token level,
- 9:18they're very similar, but a
- 9:19physics-aware text encoder will produce
- 9:22meaningfully different embeddings.
- 9:24In practice, generative models work very
- 9:27well. Outputting pixels is actually
- 9:29inevitable if your ultimate goal is
- 9:31visual content for human consumption,
- 9:33like the interactive environments we'll
- 9:35see in the applications chapter.
- 9:38But when the output is consumed by an
- 9:39autonomous agent, it can have a more
- 9:41abstract representation. This is the
- 9:44predictive school of thought. Its
- 9:46biggest proponent is Yann LeCun, a
- 9:48Turing Award winner, former Chief AI
- 9:50Scientist at Meta, and current founder
- 9:52of a startup building world models. His
- 9:55philosophy is that an effective world
- 9:57model should not get bogged down into
- 9:59details, but rather discover high-level
- 10:02patterns that capture the fundamental
- 10:04laws of the world. In this view, a
- 10:06generative model can get distracted by
- 10:09irrelevant pixels. For instance, a road
- 10:12is the same road during the day and at
- 10:14night time.
- 10:15These two world states should live
- 10:17extremely close in a conceptual space,
- 10:20but their pixel representations are very
- 10:22different. But wait a second, modern
- 10:25diffusion models, including COSMOS
- 10:27Predict, do operate in a compact latent
- 10:30space.
- 10:31How exactly is the second school of
- 10:33thought any different? To answer this
- 10:35question, we'll look at Yann LeCun's
- 10:37work at Meta. In particular, a world
- 10:40model called V-JEPA 2-AC. It's a
- 10:44mouthful, but we'll break it down.
- 10:46You might have already heard of JEPA,
- 10:48which stands for Joint Embedding
- 10:50Predictive Architecture.
- 10:52LeCun proposed this conceptual
- 10:54architecture in his well-known position
- 10:56paper back in 2022.
- 10:59He argued that reconstruction losses,
- 11:02like next word prediction in LLMs or
- 11:04denoising pixels in diffusion models,
- 11:07are fundamentally limiting and cannot
- 11:09lead to true intelligence.
- 11:12JEPA is a modality-agnostic philosophy,
- 11:14which was later adapted to image, video,
- 11:17and even language.
- 11:20V-JEPA stands for video, which is the
- 11:22most relevant for world models.
- 11:25In its raw form, V-JEPA itself is not a
- 11:28world model, but rather a video encoder
- 11:30model.
- 11:31It takes in a video and outputs a latent
- 11:34representation for it.
- 11:36However, in V-JEPA can be fine-tuned
- 11:39into a proper world model like V-JEPA
- 11:42AC.
- 11:43AC stands for action conditioned because
- 11:46compared to the base model, it can take
- 11:48in an action in addition to the video
- 11:51input.
- 11:52Under the hood, V-JEPA AC puts together
- 11:55the pre-trained video encoder and the
- 11:58new predictor module.
- 12:00The predictor takes in the current state
- 12:02of the world in latent form as well as
- 12:04the action and outputs the future in the
- 12:07same shape.
- 12:09To understand how this representation is
- 12:11different from the latent space and say
- 12:13diffusion models, it's worth
- 12:15understanding how the base V-JEPA
- 12:17encoder is pre-trained.
- 12:19It's entirely self-supervised.
- 12:22The input video is corrupted by masking
- 12:24large regions within each frame and the
- 12:27model has to recover them.
- 12:29But crucially, not in pixel space.
- 12:32The masked frames are passed through a
- 12:34context encoder and the original frames
- 12:36are passed through a target encoder.
- 12:39Both are mapped to a latent space.
- 12:42Then the masked latents go through a
- 12:44predictor which adds back the missing
- 12:47information.
- 12:49The loss minimizes the distance between
- 12:51the recovered and the original latent
- 12:53frames.
- 12:55If implemented naively, the latent space
- 12:57can collapse into something trivial like
- 13:00all zeros. This is in fact the biggest
- 13:02challenge in JEPA models.
- 13:04At a high level, the trick is that these
- 13:06two encoders are tied together. The
- 13:09target encoder is a sort of slow-moving
- 13:12average of the context encoder weights
- 13:14and doesn't get updated by gradients.
- 13:17But regardless of these implementation
- 13:19details, what I want you to take is that
- 13:22the reconstruction no longer happens in
- 13:24pixel space but rather in embedding
- 13:26space. The loss is literally defined in
- 13:29terms of embeddings.
- 13:31In contrast, the latent space in a video
- 13:34diffusion model is more of a
- 13:36computational efficiency trick to manage
- 13:39the high dimensionality of videos.
- 13:41Ultimately, the training loss is still
- 13:43defined at a pixel level, which is less
- 13:46likely to induce a robust world model.
- 13:49So, now you have a high-level
- 13:51understanding of what world models are
- 13:54and the two schools of thought for how
- 13:55to implement them. It's time to see how
- 13:58they're used in practice. What's the
- 14:00most common use case and the flagship
- 14:03industries that are using Cosmos-like
- 14:05models today? The primary reason a lot
- 14:08of companies are coming and using and
- 14:10building off of this is because this is
- 14:13really hard to get the data. If you
- 14:15think about, you know, a self-driving
- 14:16car, most people watching will
- 14:17understand, right? I'm driving down the
- 14:19road, I've got a dashcam, I can see it,
- 14:21but there's so many different roads and
- 14:23so many different scenarios and
- 14:24different places in the world that you
- 14:26just can't record all of it. And so,
- 14:28building synthetic data to help augment,
- 14:31like we talked about with that bear, or
- 14:32if you want to go to a different city,
- 14:34or change the signs,
- 14:36you really need to build a lot of
- 14:37synthetic data. And if you imagine cars
- 14:40having dashcams, and there's a lot of
- 14:42it, think of a robot that you're
- 14:44building for the first time, and you've
- 14:45only got maybe 100 or 1,000 hours of a
- 14:47robot just doing some simple thing,
- 14:50because it doesn't exist yet. You're
- 14:51still building the grippers, you're
- 14:52still building the arms and the cameras,
- 14:54etc. So,
- 14:55uh this is where having a world
- 14:57foundation model to help you extend your
- 14:59data is a really big deal. So, world
- 15:02models have many applications, but we'll
- 15:04look at three main categories. The first
- 15:06one is what TJ described, synthetic data
- 15:09for training and evaluation.
- 15:12Historically, the autonomous vehicles
- 15:14industry used procedural simulators.
- 15:17These were hard-coded virtual
- 15:19environments built on traditional game
- 15:21engines, where every physical rule,
- 15:24visual asset, and traffic scenario had
- 15:27to be explicitly programmed by
- 15:28developers. They offered full control,
- 15:31but were not visually realistic.
- 15:34Video-based world models trade off some
- 15:37of this controllability for near-perfect
- 15:39realism.
- 15:41Take Wayve, for example, a self-driving
- 15:43company based in the UK.
- 15:46They developed a series of world models
- 15:48called Gaia, for which they released my
- 15:50comprehensive technical technical
- 15:51reports.
- 15:53Gaia can augment dashcam footage. Given
- 15:56a recording from a real ego vehicle,
- 15:58here captured by five cameras at
- 16:00different angles, it can produce
- 16:02synthetic variations. Some are changing
- 16:05the time of day and therefore
- 16:06visibility. Some are adding new
- 16:08obstacles like this second bus that only
- 16:11becomes visible after overtaking the
- 16:13first bus.
- 16:15Augmenting all these scenarios increases
- 16:18the amount of training data and makes
- 16:19the model more robust. The specific
- 16:22color of a car might be irrelevant
- 16:24information for predicting its next
- 16:26move. But this irrelevancy only becomes
- 16:29clear to the model when it can observe
- 16:31that gray, red, black, and white cars
- 16:34behave similarly.
- 16:37Synthetic data becomes even more crucial
- 16:39in safety-critical scenarios. There's an
- 16:41entire taxonomy of stress tests that can
- 16:43be covered this way. Car-to-car front
- 16:46turn across path, car-to-car rear
- 16:49stationary, and who knows how many more.
- 16:52Having such rigorous and comprehensive
- 16:54benchmarks allows AV companies to test
- 16:57how their vehicles react in
- 16:58life-threatening situations.
- 17:00Using world models to generate synthetic
- 17:03data for AVs is slowly becoming an
- 17:06industry standard. Google's self-driving
- 17:08car Waymo announced that they're
- 17:10leveraging a fine-tune of Genie, the
- 17:12world model coming from DeepMind, for
- 17:14similar purposes. But in addition to
- 17:17video, they're also generating lidar
- 17:19data, which gives their autonomous
- 17:21driver a better sense of depth.
- 17:24Synthetic data generation is an offline
- 17:26process. Once the data set is produced,
- 17:28the world model is put aside and
- 17:30completely decoupled from the deployment
- 17:33of the autonomous vehicle.
- 17:35But, as world models are becoming faster
- 17:37and faster, they're starting to be used
- 17:39in real time and interactively.
- 17:42Google's Genie is an experimental
- 17:44project that lets you craft an
- 17:46interactive 3D world from just a text
- 17:49prompt. It's what I used to generate the
- 17:51more sane version of a car drifting on
- 17:53an icy road.
- 17:55Genie is less focused on embodied AI and
- 17:58can create fictitious virtual worlds as
- 18:00well, like this one made out of felt.
- 18:04For about 60 seconds, you can use the
- 18:06WASD keys to move your character around
- 18:10or the arrow keys to change the point of
- 18:12view.
- 18:13Throughout this exploration, the
- 18:14environment mostly stays consistent.
- 18:17And you can also create your own world.
- 18:23My virtual office here looks pretty
- 18:25convincing, even though there are a few
- 18:27unnatural choices, like my desk or the
- 18:30same orange painting showing up multiple
- 18:32times.
- 18:34Nevertheless, this technology is truly
- 18:36unprecedented. There's no doubt that
- 18:38eventually this will disrupt the gaming
- 18:41and filmmaking industries.
- 18:43But, at this very moment, unit economics
- 18:46are still a bottleneck. Genie is only
- 18:48available under the Google AI Ultra
- 18:50plan, which costs $250 a month.
- 18:54Plus, once your world is created, you're
- 18:55given a single minute to explore it.
- 18:58This might be Google's way of keeping
- 18:59costs under control, but it also signals
- 19:02a technical limitation. To maintain
- 19:04consistency, the model must keep a
- 19:06memory of the entire session.
- 19:08Cross-frame consistency is actually one
- 19:10of the most striking aspects of Genie.
- 19:13Unfortunately, we don't really know how
- 19:14it works under the hood. Google
- 19:16published a technical report for Genie
- 19:181, but that's more than 2 years old by
- 19:20now.
- 19:22There were speculation online that the
- 19:24architecture of Genie 3 might have
- 19:26finally gone beyond the traditional
- 19:28video diffusion model, and the output
- 19:30might be more than just pixels, perhaps
- 19:32an actual 3D mesh and texture.
- 19:35But from Google's blog post, that seems
- 19:37very unlikely. They say consistency is
- 19:40an emergent capability and explicitly
- 19:43exclude 3D specific outputs like nerves
- 19:46and Gaussian splats.
- 19:48This is quite an opinionated choice and
- 19:50goes against the intuition of some
- 19:52experts in the field.
- 19:54Fei-Fei Li, the godmother of AI, coined
- 19:57the term spatial intelligence, referring
- 20:00to the ability of machines to perceive,
- 20:03reason about, and interact with a 3D
- 20:05physical world. Fei-Fei Li has recently
- 20:08founded World Labs, for which she raised
- 20:10over a billion dollars. Their main
- 20:13product is Marble, a world model that's
- 20:15similar to Genie and offers an
- 20:17interactive experience in a virtual 3D
- 20:19world, but with a different technical
- 20:22approach that doesn't directly generate
- 20:25pixels. For interactive environments,
- 20:27pixel outputs have one huge drawback.
- 20:30They tie together geometry with
- 20:32appearance.
- 20:33This is different from classic game
- 20:35engines, where the structure of the
- 20:37objects is defined by a
- 20:38three-dimensional mesh, while visual
- 20:41appearance, including color, roughness,
- 20:43reflectance, and so on, is applied
- 20:45separately via materials and texture
- 20:48maps layered on top.
- 20:50Marble brings back this decoupling, but
- 20:53in the form of Gaussian splats. These
- 20:55are semi-transparent colored ellipsoids
- 20:58with their own size, position, and
- 21:00orientation.
- 21:02Since these aspects are decoupled, they
- 21:04can be be separately.
- 21:06World Labs even released a custom
- 21:08renderer that converts Gaussian splats
- 21:11to pixels, but also allows these massive
- 21:13virtual worlds to be compressed and
- 21:15streamed progressively, much like a
- 21:17high-definition video buffer.
- 21:20At this point in time, World Labs seems
- 21:22to be the furthest ahead in disrupting
- 21:24the gaming and filmmaking industry.
- 21:26All the world models we discussed so far
- 21:29act like content factories, either
- 21:31generating synthetic data offline or
- 21:34creating interactive worlds for humans
- 21:36to explore. But, at its origin, a world
- 21:39model, or rather small-scale model, as
- 21:42it was called back in 1943,
- 21:45was defined as a neural component that
- 21:47actively helps humans make decisions.
- 21:51In this final category of applications,
- 21:53we look at world models that help agents
- 21:56in a similar way by unfolding
- 21:58hypothetical futures, either during
- 22:00training, inference, or both. At
- 22:03training time, an agent can practice
- 22:05inside the world model before
- 22:07deployment. This is particularly helpful
- 22:09when interacting with the real world can
- 22:11be expensive, dangerous, or slow. The
- 22:14umbrella term for this is model-based
- 22:17reinforcement learning, or MBRL.
- 22:21In classical reinforcement learning, an
- 22:24agent learns to achieve a goal from
- 22:25experience by interacting with a real
- 22:28environment that provides rewards for
- 22:30its actions. The main decision-maker is
- 22:32the policy model.
- 22:34But, in model-based reinforcement
- 22:36learning, there's an additional world
- 22:38model, separate from the policy, that
- 22:41can emulate the real environment.
- 22:44This was actually the design behind the
- 22:46paper that popularized world models,
- 22:49published in 2018 entitled simply World
- 22:52Models.
- 22:53They built a small agent to play an
- 22:55adaptation of the iconic game Doom
- 22:58called VizDoom, where the simplified
- 23:00goal of the agent is to survive for as
- 23:03long as possible.
- 23:05They enforced a deliberately strict
- 23:06constraint. The policy model should be
- 23:09trained without any interaction with the
- 23:11real game environment. Such interactions
- 23:14would only be allowed for training the
- 23:16world model.
- 23:17Of course, this is not a practical
- 23:19constraint in video games, which are
- 23:21relatively cheap to run, but at the time
- 23:23in 2018, it was a promising proof of
- 23:26concept that could one day materialize
- 23:28for training physical robots.
- 23:31So, the VizDoom agent went through two
- 23:33training stages before deployment.
- 23:36First, training a world model. They
- 23:38started with a randomly initialized
- 23:39policy and played 10,000 games. In other
- 23:43words, they issued random actions until
- 23:46the agent got killed 10,000 times.
- 23:49In the process, the world model got to
- 23:51observe the reaction of the real
- 23:52environment and recreate a smaller,
- 23:55compressed version of it, where action
- 23:57consequences unfold in a latent space
- 24:00instead of pixel frames.
- 24:02Once the world model was trained, it was
- 24:04then frozen and used as a simulation
- 24:06environment to train the policy model.
- 24:09At the end of the second training stage,
- 24:11the agent was deployed in the real
- 24:12environment and was able to successfully
- 24:15play VizDoom, despite never having seen
- 24:17the actual game before.
- 24:20Subsequent research scaled MBRL in
- 24:22various ways.
- 24:24For instance, the Dreamer model family
- 24:26from DeepMind worked their way up to
- 24:28mining a diamond in Minecraft from
- 24:31scratch.
- 24:32This had been an unsolved problem in AI
- 24:35for years, and prior attempts needed
- 24:37human gameplay videos.
- 24:40And today, we're starting to see it
- 24:41applied to embodied robots as well.
- 24:44For instance, World Gymnast is a
- 24:46tabletop manipulation robot that was
- 24:49exclusively trained inside a world model
- 24:52and deployed on physical hardware
- 24:54without any real-world fine-tuning.
- 24:57In the example so far, the world model
- 24:59is a training time tool discarded once
- 25:02the policy is trained.
- 25:04But world models can also stick around
- 25:06at inference time. The agent can
- 25:08actively use them in real time to unfold
- 25:11the effects of multiple candidate
- 25:12actions and pick the one with the most
- 25:15promising outcome. This is real time
- 25:17planning. The most common planning
- 25:19paradigm for agents is model predictive
- 25:22control inspired from 1970s chemical
- 25:25engineering.
- 25:27Here's one way to implement it.
- 25:30Before taking an action in the real
- 25:31world, an agent builds an entire tree of
- 25:34possibilities rooted in the current
- 25:36world state.
- 25:38The edges are candidate actions and the
- 25:40nodes are hypothetical future world
- 25:42states as computed by a world model.
- 25:45It's like being inside the mind of an
- 25:47extremely anxious person trying to
- 25:49foresee how the future might unfold.
- 25:52After building this tree, the agent
- 25:54identifies the branch with the best
- 25:55estimated outcome and performs only the
- 25:58very first action in the real world.
- 26:01It then uh throws away the rest of the
- 26:03branch and builds an entirely new tree
- 26:06rooted in the updated world state
- 26:09communicated by the environment.
- 26:11Working in a latent space is extremely
- 26:14important for speed since decision trees
- 26:16are built in real time.
- 26:18This is the idea behind MuZero, a 2019
- 26:21algorithm from DeepMind trained to play
- 26:24board games and Atari video games.
- 26:27But planning is now coming to embodied
- 26:29AI as well, like VIPAC, the model we
- 26:33discussed in the implementation section,
- 26:35as an exponent for the family of
- 26:37predictive world models.
- 26:39Technically, VIPAC has a branching
- 26:41factor of one but follows the same
- 26:44principle of model predictive control.
- 26:47Interestingly, it uses subgoal images
- 26:49provided by the environment to judge how
- 26:52good a certain branch is.
- 26:55This means it doesn't even need rewards
- 26:57like the ones in reinforcement learning.
- 26:59You just show it an image of what you
- 27:01want and it'll do it in a task-agnostic
- 27:04way.
- 27:05On the one hand, this is Black Mirror
- 27:07stuff, but on the other, imagine how
- 27:10helpful this would be for an autonomous
- 27:12vehicle that needs to make a quick,
- 27:14critical decision.
- 27:16You can start doing closed-loop
- 27:18scenarios where you can start saying,
- 27:20"Okay, if I want to drive the car and
- 27:21turn right, what will happen next and
- 27:23what will I see?" Now, show me 50
- 27:25different versions of that and then pick
- 27:27the most, you know, useful or reliable
- 27:29one. Before we wrap up, I want to
- 27:31mention that world models don't have to
- 27:33be visual. The same definition, given a
- 27:35state and an action, predict the next
- 27:37state, works in any environment where
- 27:39you have well-defined states and
- 27:42actions. A nice recent example is the
- 27:44Coded World Model from Meta.
- 27:46The world here is a software
- 27:48environment, a repository, a program, a
- 27:51terminal. The action is something like a
- 27:54pull request. And the world model
- 27:56predicts the next state, the variable
- 27:58values, output, execution traces, and so
- 28:01on. So, why not just execute the code
- 28:04and observe its effects instead of
- 28:06building a world model of the software
- 28:08environment?
- 28:10Well, it might take a full hour to run
- 28:12the entire test suite, and certain bugs
- 28:14might only surface in production, going
- 28:17unnoticed during development.
- 28:19A world model could be quicker and could
- 28:21catch potential issues without causing
- 28:23irreparable damage. We covered a lot of
- 28:26ground, but I hope that the through line
- 28:28is clear. A world model predicts how a
- 28:30certain action will change the state of
- 28:32the world. There's a lot of ambiguity
- 28:34around this term because the interface
- 28:36is so general and can be applied across
- 28:38so many industries. If you want to hear
- 28:41more from TJ Galda, the full interview
- 28:43is available to my YouTube and Patreon
- 28:45members. Stay tuned for more videos on
- 28:48world models.
About this transcript
This page contains the full transcript of But what exactly are world models? by Julia Turc, generated from the public captions YouTube serves with the video. The transcript has 4,376 words across 748 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.