YouTube2Text

But what exactly are world models? — Transcript

by Julia Turc · 4,376 words · 748 segments · language en · Watch on YouTube

Full transcript

  1. 0:14This is Google's V03, a state-of-the-art
  2. 0:17video generation model. It certainly
  3. 0:19took some creative liberties. There's a
  4. 0:22really weird Transformers situation over
  5. 0:24here. And what's up with these bunny
  6. 0:26jumps?
  7. 0:28General video generation models still
  8. 0:30don't have a good grasp of real-world
  9. 0:32dynamics. They understand some physical
  10. 0:35laws like gravity or object occlusion,
  11. 0:37but they can't clear the quality bar
  12. 0:39required by production systems like
  13. 0:41autonomous vehicles. For that, I hear we
  14. 0:44need something else. A world model.
  15. 0:45>> A world models.
  16. 0:47>> A world model. A world model. Okay,
  17. 0:48fine. Let's try one.
  18. 0:50This is Genie, Google's world model.
  19. 0:57It seems like world models might
  20. 0:59actually work.
  21. 1:00Now look, this buzzword is being thrown
  22. 1:02around in all kinds of contexts today as
  23. 1:05some sort of big unlock in the pursuit
  24. 1:08of AGI. And a lot of us are left
  25. 1:10wondering what exactly is a world model?
  26. 1:12How do you build one? And then, what do
  27. 1:14you do with it? In this video, we'll try
  28. 1:16to clear out this confusion. We'll go
  29. 1:18through the definition, implementations,
  30. 1:20and the many applications of world
  31. 1:23models. We'll also hear directly from TJ
  32. 1:25Galda, an expert from Nvidia leading
  33. 1:27their world model efforts.
  34. 1:30But first things first, what exactly is
  35. 1:32a world model?
  36. 1:34The idea dates all the way back to 1943
  37. 1:38when a Scottish psychologist named
  38. 1:41Kenneth Craik suggested that the human
  39. 1:43mind has an internal small-scale model
  40. 1:46of reality.
  41. 1:48It can try out various alternatives,
  42. 1:50conclude which is the best of them, and
  43. 1:52act in a much fuller, safer, and more
  44. 1:54competent manner.
  45. 1:55It's how you know that jumping in front
  46. 1:57a train is a bad idea, even if you've
  47. 2:00never tried it before.
  48. 2:01About 80 years later, in 2018, a paper
  49. 2:05called world models brought this idea to
  50. 2:07machine learning. An agent learning to
  51. 2:10play a variation of Doom leverages a
  52. 2:12small internal model of the game to
  53. 2:15imagine how its actions would play out
  54. 2:17in the real game. Today, prominent
  55. 2:20researchers agree that world models are
  56. 2:22essential to building AI systems that
  57. 2:24can plan, reason, and act safely in the
  58. 2:27real world. Imagine a cube floating in
  59. 2:29the air in front of you, and imagine
  60. 2:30rotating that cube by 90°. You can sort
  61. 2:33of picture this in your mind, and this
  62. 2:34has nothing to do with language.
  63. 2:37Humans and animals navigate the world by
  64. 2:39building mental models of reality. What
  65. 2:41if AI could develop this kind of common
  66. 2:42sense? An ability to make predictions of
  67. 2:45what's going to happen in some sort of
  68. 2:47abstract representation space.
  69. 2:49We call this concept a world model.
  70. 2:52In the strictest sense, in the context
  71. 2:54of machine learning, a world model takes
  72. 2:57in the current state of the world,
  73. 2:59together with a hypothetical action that
  74. 3:01could take place in it. It's goal is to
  75. 3:04predict the effect of this action,
  76. 3:06reflected in a future world state. This
  77. 3:09is the interface that is most faithful
  78. 3:11to the roots of world models in
  79. 3:13cognitive science.
  80. 3:15What's your minimal definition of a
  81. 3:17world model?
  82. 3:18>> When we say world model, we mean an AI
  83. 3:20system that literally learns in how the
  84. 3:22physical world changes over time, right?
  85. 3:24And so, in the same way a language model
  86. 3:26can predict like the next token or, you
  87. 3:28know, the syllable of a word,
  88. 3:30essentially, a world model's really
  89. 3:32predicting the state of the world. Not
  90. 3:34just what it looks like, but what it's
  91. 3:35going to do and how it's going to act,
  92. 3:37right?
  93. 3:38The exact implementation varies a lot
  94. 3:40across research labs and industry use
  95. 3:43cases.
  96. 3:44For simplicity, let's anchor this in
  97. 3:47autonomous vehicles.
  98. 3:49The current world state is captured by
  99. 3:51the dash cam, either as a static image
  100. 3:54or a short video and the intended action
  101. 3:57can be expressed as a text prompt like
  102. 3:59take a left turn. Pretty much everyone
  103. 4:01agrees on this input shape.
  104. 4:04The debate happens entirely on the
  105. 4:05output side and it comes down to one
  106. 4:08question. How faithfully should the
  107. 4:10world model represent the future world
  108. 4:12state?
  109. 4:14There are two schools of thought,
  110. 4:16generative and predictive or
  111. 4:19non-generative.
  112. 4:20A generative model outputs the future
  113. 4:23world state in a human-friendly form
  114. 4:25like a fully fledged video.
  115. 4:28Such world models can be used as
  116. 4:30standalone tools be it for synthetic
  117. 4:32data generation or interactive
  118. 4:34environments.
  119. 4:36Multiple prominent labs subscribe to
  120. 4:38this philosophy. Nvidia, Google DeepMind
  121. 4:41and Fei-Fei Li's World Labs.
  122. 4:44In this chapter, we'll look at the
  123. 4:46implementation of Nvidia Cosmos as a
  124. 4:48representative model for the generative
  125. 4:51family.
  126. 4:52We'll also visit the other two in the
  127. 4:54applications chapter.
  128. 4:56Nvidia Cosmos is a family of open-source
  129. 4:59models highly focused on physical
  130. 5:01intelligence including self-driving cars
  131. 5:04and autonomous robots that can operate
  132. 5:06in a warehouse or even in a surgery
  133. 5:09room.
  134. 5:10Officially, Cosmos is a collection of
  135. 5:12models that includes Cosmos Predict,
  136. 5:15Transfer and Reason.
  137. 5:17But technically, Cosmos Predict is the
  138. 5:20only one that abides by the strict
  139. 5:22definition of a world model.
  140. 5:25Given this video of a robot pouring the
  141. 5:27coffee which is the current state of the
  142. 5:29world, Cosmos Predict outputs the
  143. 5:32natural continuation where the coffee
  144. 5:34kettle is angled back up in a vertical
  145. 5:36position.
  146. 5:39The other two Cosmos models are
  147. 5:41technically not world models. Cosmos
  148. 5:43Transfer is well, a style transfer model
  149. 5:46and Cosmos Region is a visual language
  150. 5:49model or VLA that can answer questions
  151. 5:51about a given visual input.
  152. 5:54Nonetheless, you will hear people
  153. 5:55referring to all of them as world
  154. 5:57models, but we'll just have to live with
  155. 5:59this ambiguity in colloquial language.
  156. 6:04Cosmos Predict, the true world model of
  157. 6:07a family, is mostly a standard video
  158. 6:09diffusion model with a few tweaks. I
  159. 6:11have an entire YouTube series on
  160. 6:13diffusion models, but here's the
  161. 6:15high-level idea.
  162. 6:16The model architecture is a typical
  163. 6:18stack of diffusion transformers or DATs,
  164. 6:21which I've covered in this video.
  165. 6:24The DAT is basically an extension of the
  166. 6:26original language transformer capable of
  167. 6:28processing both image and text tokens.
  168. 6:32The current world state is compressed by
  169. 6:34a visual encoder into a smaller latent
  170. 6:37space, one slice for an image input or
  171. 6:40multiple slices for video frames.
  172. 6:42The visual encoder is an off-the-shelf
  173. 6:44model, in particular, the WAE 2.1 VAE
  174. 6:48encoder from Alibaba.
  175. 6:51These latent input frames depicted in
  176. 6:53yellow are concatenated with
  177. 6:54placeholders for the future frames
  178. 6:57initialized with pure Gaussian noise.
  179. 7:00The diffusion model refines them
  180. 7:02iteratively for a fixed number of steps,
  181. 7:05after which the added frames
  182. 7:06meaningfully capture the future.
  183. 7:09The clean output latent is finally
  184. 7:11passed through an off-the-shelf video
  185. 7:13decoder, also taken from WAE, which maps
  186. 7:16it back to pixel-based video frames.
  187. 7:19This denoising network is trained with a
  188. 7:21standard flow matching objective, which
  189. 7:23I've covered in a previous video.
  190. 7:26Now, from everything I've described so
  191. 7:28far, Nvidia's Cosmos Predict matches the
  192. 7:31structure of a general-purpose video
  193. 7:33generation model
  194. 7:35like VEO 3 from Google or SeeDance from
  195. 7:38ByteDance. The ones we normally use to
  196. 7:40generate AI slop, like food cannibalism
  197. 7:44or soulless Hollywood-looking movie
  198. 7:46trailers.
  199. 7:48So, what's the difference then? How is a
  200. 7:50video-based world model any different?
  201. 7:53One of the main differentiators is the
  202. 7:54training data. General-purpose video
  203. 7:57generators are trained on any sort of
  204. 7:59video that labs can get their hands on.
  205. 8:01This includes real-world footage, but
  206. 8:03also things like video games, cartoons,
  207. 8:05or even slide deck presentations. In
  208. 8:07contrast, Cosmos was pre-trained on
  209. 8:10highly curated data. We've got over 20
  210. 8:13million hours of physics-first data.
  211. 8:14We're really trying to make sure that
  212. 8:16the data coming in is curated and and
  213. 8:18trained grounded on physics. When we
  214. 8:21train on this, then we evaluate it on a
  215. 8:23whole bunch of benchmarks that also
  216. 8:24check for that. Cosmos Predict also
  217. 8:27makes a few architectural tweaks. For
  218. 8:29instance, it swaps out the standard
  219. 8:32off-the-shelf text encoder with Cosmos
  220. 8:35Reason, the visual language model from
  221. 8:37the same family.
  222. 8:39Normally, Cosmos Reason takes in an
  223. 8:41image and a text prompt, but in this
  224. 8:43particular setup, it will only exercise
  225. 8:45the text processing path.
  226. 8:48Just like the other models in the
  227. 8:49family, Cosmos Reason was also trained
  228. 8:52on highly curated data. Nvidia even
  229. 8:55built a manual ontology around
  230. 8:57foundational physics, time and space,
  231. 8:59and made sure the training data covers
  232. 9:01it comprehensively.
  233. 9:03Compared to general-purpose text
  234. 9:05encoders like T5, this makes Cosmos
  235. 9:07Reason embeddings extra sensitive to the
  236. 9:10difference between the egg cracked after
  237. 9:13being dropped versus the egg was dropped
  238. 9:16after cracking. At a token level,
  239. 9:18they're very similar, but a
  240. 9:19physics-aware text encoder will produce
  241. 9:22meaningfully different embeddings.
  242. 9:24In practice, generative models work very
  243. 9:27well. Outputting pixels is actually
  244. 9:29inevitable if your ultimate goal is
  245. 9:31visual content for human consumption,
  246. 9:33like the interactive environments we'll
  247. 9:35see in the applications chapter.
  248. 9:38But when the output is consumed by an
  249. 9:39autonomous agent, it can have a more
  250. 9:41abstract representation. This is the
  251. 9:44predictive school of thought. Its
  252. 9:46biggest proponent is Yann LeCun, a
  253. 9:48Turing Award winner, former Chief AI
  254. 9:50Scientist at Meta, and current founder
  255. 9:52of a startup building world models. His
  256. 9:55philosophy is that an effective world
  257. 9:57model should not get bogged down into
  258. 9:59details, but rather discover high-level
  259. 10:02patterns that capture the fundamental
  260. 10:04laws of the world. In this view, a
  261. 10:06generative model can get distracted by
  262. 10:09irrelevant pixels. For instance, a road
  263. 10:12is the same road during the day and at
  264. 10:14night time.
  265. 10:15These two world states should live
  266. 10:17extremely close in a conceptual space,
  267. 10:20but their pixel representations are very
  268. 10:22different. But wait a second, modern
  269. 10:25diffusion models, including COSMOS
  270. 10:27Predict, do operate in a compact latent
  271. 10:30space.
  272. 10:31How exactly is the second school of
  273. 10:33thought any different? To answer this
  274. 10:35question, we'll look at Yann LeCun's
  275. 10:37work at Meta. In particular, a world
  276. 10:40model called V-JEPA 2-AC. It's a
  277. 10:44mouthful, but we'll break it down.
  278. 10:46You might have already heard of JEPA,
  279. 10:48which stands for Joint Embedding
  280. 10:50Predictive Architecture.
  281. 10:52LeCun proposed this conceptual
  282. 10:54architecture in his well-known position
  283. 10:56paper back in 2022.
  284. 10:59He argued that reconstruction losses,
  285. 11:02like next word prediction in LLMs or
  286. 11:04denoising pixels in diffusion models,
  287. 11:07are fundamentally limiting and cannot
  288. 11:09lead to true intelligence.
  289. 11:12JEPA is a modality-agnostic philosophy,
  290. 11:14which was later adapted to image, video,
  291. 11:17and even language.
  292. 11:20V-JEPA stands for video, which is the
  293. 11:22most relevant for world models.
  294. 11:25In its raw form, V-JEPA itself is not a
  295. 11:28world model, but rather a video encoder
  296. 11:30model.
  297. 11:31It takes in a video and outputs a latent
  298. 11:34representation for it.
  299. 11:36However, in V-JEPA can be fine-tuned
  300. 11:39into a proper world model like V-JEPA
  301. 11:42AC.
  302. 11:43AC stands for action conditioned because
  303. 11:46compared to the base model, it can take
  304. 11:48in an action in addition to the video
  305. 11:51input.
  306. 11:52Under the hood, V-JEPA AC puts together
  307. 11:55the pre-trained video encoder and the
  308. 11:58new predictor module.
  309. 12:00The predictor takes in the current state
  310. 12:02of the world in latent form as well as
  311. 12:04the action and outputs the future in the
  312. 12:07same shape.
  313. 12:09To understand how this representation is
  314. 12:11different from the latent space and say
  315. 12:13diffusion models, it's worth
  316. 12:15understanding how the base V-JEPA
  317. 12:17encoder is pre-trained.
  318. 12:19It's entirely self-supervised.
  319. 12:22The input video is corrupted by masking
  320. 12:24large regions within each frame and the
  321. 12:27model has to recover them.
  322. 12:29But crucially, not in pixel space.
  323. 12:32The masked frames are passed through a
  324. 12:34context encoder and the original frames
  325. 12:36are passed through a target encoder.
  326. 12:39Both are mapped to a latent space.
  327. 12:42Then the masked latents go through a
  328. 12:44predictor which adds back the missing
  329. 12:47information.
  330. 12:49The loss minimizes the distance between
  331. 12:51the recovered and the original latent
  332. 12:53frames.
  333. 12:55If implemented naively, the latent space
  334. 12:57can collapse into something trivial like
  335. 13:00all zeros. This is in fact the biggest
  336. 13:02challenge in JEPA models.
  337. 13:04At a high level, the trick is that these
  338. 13:06two encoders are tied together. The
  339. 13:09target encoder is a sort of slow-moving
  340. 13:12average of the context encoder weights
  341. 13:14and doesn't get updated by gradients.
  342. 13:17But regardless of these implementation
  343. 13:19details, what I want you to take is that
  344. 13:22the reconstruction no longer happens in
  345. 13:24pixel space but rather in embedding
  346. 13:26space. The loss is literally defined in
  347. 13:29terms of embeddings.
  348. 13:31In contrast, the latent space in a video
  349. 13:34diffusion model is more of a
  350. 13:36computational efficiency trick to manage
  351. 13:39the high dimensionality of videos.
  352. 13:41Ultimately, the training loss is still
  353. 13:43defined at a pixel level, which is less
  354. 13:46likely to induce a robust world model.
  355. 13:49So, now you have a high-level
  356. 13:51understanding of what world models are
  357. 13:54and the two schools of thought for how
  358. 13:55to implement them. It's time to see how
  359. 13:58they're used in practice. What's the
  360. 14:00most common use case and the flagship
  361. 14:03industries that are using Cosmos-like
  362. 14:05models today? The primary reason a lot
  363. 14:08of companies are coming and using and
  364. 14:10building off of this is because this is
  365. 14:13really hard to get the data. If you
  366. 14:15think about, you know, a self-driving
  367. 14:16car, most people watching will
  368. 14:17understand, right? I'm driving down the
  369. 14:19road, I've got a dashcam, I can see it,
  370. 14:21but there's so many different roads and
  371. 14:23so many different scenarios and
  372. 14:24different places in the world that you
  373. 14:26just can't record all of it. And so,
  374. 14:28building synthetic data to help augment,
  375. 14:31like we talked about with that bear, or
  376. 14:32if you want to go to a different city,
  377. 14:34or change the signs,
  378. 14:36you really need to build a lot of
  379. 14:37synthetic data. And if you imagine cars
  380. 14:40having dashcams, and there's a lot of
  381. 14:42it, think of a robot that you're
  382. 14:44building for the first time, and you've
  383. 14:45only got maybe 100 or 1,000 hours of a
  384. 14:47robot just doing some simple thing,
  385. 14:50because it doesn't exist yet. You're
  386. 14:51still building the grippers, you're
  387. 14:52still building the arms and the cameras,
  388. 14:54etc. So,
  389. 14:55uh this is where having a world
  390. 14:57foundation model to help you extend your
  391. 14:59data is a really big deal. So, world
  392. 15:02models have many applications, but we'll
  393. 15:04look at three main categories. The first
  394. 15:06one is what TJ described, synthetic data
  395. 15:09for training and evaluation.
  396. 15:12Historically, the autonomous vehicles
  397. 15:14industry used procedural simulators.
  398. 15:17These were hard-coded virtual
  399. 15:19environments built on traditional game
  400. 15:21engines, where every physical rule,
  401. 15:24visual asset, and traffic scenario had
  402. 15:27to be explicitly programmed by
  403. 15:28developers. They offered full control,
  404. 15:31but were not visually realistic.
  405. 15:34Video-based world models trade off some
  406. 15:37of this controllability for near-perfect
  407. 15:39realism.
  408. 15:41Take Wayve, for example, a self-driving
  409. 15:43company based in the UK.
  410. 15:46They developed a series of world models
  411. 15:48called Gaia, for which they released my
  412. 15:50comprehensive technical technical
  413. 15:51reports.
  414. 15:53Gaia can augment dashcam footage. Given
  415. 15:56a recording from a real ego vehicle,
  416. 15:58here captured by five cameras at
  417. 16:00different angles, it can produce
  418. 16:02synthetic variations. Some are changing
  419. 16:05the time of day and therefore
  420. 16:06visibility. Some are adding new
  421. 16:08obstacles like this second bus that only
  422. 16:11becomes visible after overtaking the
  423. 16:13first bus.
  424. 16:15Augmenting all these scenarios increases
  425. 16:18the amount of training data and makes
  426. 16:19the model more robust. The specific
  427. 16:22color of a car might be irrelevant
  428. 16:24information for predicting its next
  429. 16:26move. But this irrelevancy only becomes
  430. 16:29clear to the model when it can observe
  431. 16:31that gray, red, black, and white cars
  432. 16:34behave similarly.
  433. 16:37Synthetic data becomes even more crucial
  434. 16:39in safety-critical scenarios. There's an
  435. 16:41entire taxonomy of stress tests that can
  436. 16:43be covered this way. Car-to-car front
  437. 16:46turn across path, car-to-car rear
  438. 16:49stationary, and who knows how many more.
  439. 16:52Having such rigorous and comprehensive
  440. 16:54benchmarks allows AV companies to test
  441. 16:57how their vehicles react in
  442. 16:58life-threatening situations.
  443. 17:00Using world models to generate synthetic
  444. 17:03data for AVs is slowly becoming an
  445. 17:06industry standard. Google's self-driving
  446. 17:08car Waymo announced that they're
  447. 17:10leveraging a fine-tune of Genie, the
  448. 17:12world model coming from DeepMind, for
  449. 17:14similar purposes. But in addition to
  450. 17:17video, they're also generating lidar
  451. 17:19data, which gives their autonomous
  452. 17:21driver a better sense of depth.
  453. 17:24Synthetic data generation is an offline
  454. 17:26process. Once the data set is produced,
  455. 17:28the world model is put aside and
  456. 17:30completely decoupled from the deployment
  457. 17:33of the autonomous vehicle.
  458. 17:35But, as world models are becoming faster
  459. 17:37and faster, they're starting to be used
  460. 17:39in real time and interactively.
  461. 17:42Google's Genie is an experimental
  462. 17:44project that lets you craft an
  463. 17:46interactive 3D world from just a text
  464. 17:49prompt. It's what I used to generate the
  465. 17:51more sane version of a car drifting on
  466. 17:53an icy road.
  467. 17:55Genie is less focused on embodied AI and
  468. 17:58can create fictitious virtual worlds as
  469. 18:00well, like this one made out of felt.
  470. 18:04For about 60 seconds, you can use the
  471. 18:06WASD keys to move your character around
  472. 18:10or the arrow keys to change the point of
  473. 18:12view.
  474. 18:13Throughout this exploration, the
  475. 18:14environment mostly stays consistent.
  476. 18:17And you can also create your own world.
  477. 18:23My virtual office here looks pretty
  478. 18:25convincing, even though there are a few
  479. 18:27unnatural choices, like my desk or the
  480. 18:30same orange painting showing up multiple
  481. 18:32times.
  482. 18:34Nevertheless, this technology is truly
  483. 18:36unprecedented. There's no doubt that
  484. 18:38eventually this will disrupt the gaming
  485. 18:41and filmmaking industries.
  486. 18:43But, at this very moment, unit economics
  487. 18:46are still a bottleneck. Genie is only
  488. 18:48available under the Google AI Ultra
  489. 18:50plan, which costs $250 a month.
  490. 18:54Plus, once your world is created, you're
  491. 18:55given a single minute to explore it.
  492. 18:58This might be Google's way of keeping
  493. 18:59costs under control, but it also signals
  494. 19:02a technical limitation. To maintain
  495. 19:04consistency, the model must keep a
  496. 19:06memory of the entire session.
  497. 19:08Cross-frame consistency is actually one
  498. 19:10of the most striking aspects of Genie.
  499. 19:13Unfortunately, we don't really know how
  500. 19:14it works under the hood. Google
  501. 19:16published a technical report for Genie
  502. 19:181, but that's more than 2 years old by
  503. 19:20now.
  504. 19:22There were speculation online that the
  505. 19:24architecture of Genie 3 might have
  506. 19:26finally gone beyond the traditional
  507. 19:28video diffusion model, and the output
  508. 19:30might be more than just pixels, perhaps
  509. 19:32an actual 3D mesh and texture.
  510. 19:35But from Google's blog post, that seems
  511. 19:37very unlikely. They say consistency is
  512. 19:40an emergent capability and explicitly
  513. 19:43exclude 3D specific outputs like nerves
  514. 19:46and Gaussian splats.
  515. 19:48This is quite an opinionated choice and
  516. 19:50goes against the intuition of some
  517. 19:52experts in the field.
  518. 19:54Fei-Fei Li, the godmother of AI, coined
  519. 19:57the term spatial intelligence, referring
  520. 20:00to the ability of machines to perceive,
  521. 20:03reason about, and interact with a 3D
  522. 20:05physical world. Fei-Fei Li has recently
  523. 20:08founded World Labs, for which she raised
  524. 20:10over a billion dollars. Their main
  525. 20:13product is Marble, a world model that's
  526. 20:15similar to Genie and offers an
  527. 20:17interactive experience in a virtual 3D
  528. 20:19world, but with a different technical
  529. 20:22approach that doesn't directly generate
  530. 20:25pixels. For interactive environments,
  531. 20:27pixel outputs have one huge drawback.
  532. 20:30They tie together geometry with
  533. 20:32appearance.
  534. 20:33This is different from classic game
  535. 20:35engines, where the structure of the
  536. 20:37objects is defined by a
  537. 20:38three-dimensional mesh, while visual
  538. 20:41appearance, including color, roughness,
  539. 20:43reflectance, and so on, is applied
  540. 20:45separately via materials and texture
  541. 20:48maps layered on top.
  542. 20:50Marble brings back this decoupling, but
  543. 20:53in the form of Gaussian splats. These
  544. 20:55are semi-transparent colored ellipsoids
  545. 20:58with their own size, position, and
  546. 21:00orientation.
  547. 21:02Since these aspects are decoupled, they
  548. 21:04can be be separately.
  549. 21:06World Labs even released a custom
  550. 21:08renderer that converts Gaussian splats
  551. 21:11to pixels, but also allows these massive
  552. 21:13virtual worlds to be compressed and
  553. 21:15streamed progressively, much like a
  554. 21:17high-definition video buffer.
  555. 21:20At this point in time, World Labs seems
  556. 21:22to be the furthest ahead in disrupting
  557. 21:24the gaming and filmmaking industry.
  558. 21:26All the world models we discussed so far
  559. 21:29act like content factories, either
  560. 21:31generating synthetic data offline or
  561. 21:34creating interactive worlds for humans
  562. 21:36to explore. But, at its origin, a world
  563. 21:39model, or rather small-scale model, as
  564. 21:42it was called back in 1943,
  565. 21:45was defined as a neural component that
  566. 21:47actively helps humans make decisions.
  567. 21:51In this final category of applications,
  568. 21:53we look at world models that help agents
  569. 21:56in a similar way by unfolding
  570. 21:58hypothetical futures, either during
  571. 22:00training, inference, or both. At
  572. 22:03training time, an agent can practice
  573. 22:05inside the world model before
  574. 22:07deployment. This is particularly helpful
  575. 22:09when interacting with the real world can
  576. 22:11be expensive, dangerous, or slow. The
  577. 22:14umbrella term for this is model-based
  578. 22:17reinforcement learning, or MBRL.
  579. 22:21In classical reinforcement learning, an
  580. 22:24agent learns to achieve a goal from
  581. 22:25experience by interacting with a real
  582. 22:28environment that provides rewards for
  583. 22:30its actions. The main decision-maker is
  584. 22:32the policy model.
  585. 22:34But, in model-based reinforcement
  586. 22:36learning, there's an additional world
  587. 22:38model, separate from the policy, that
  588. 22:41can emulate the real environment.
  589. 22:44This was actually the design behind the
  590. 22:46paper that popularized world models,
  591. 22:49published in 2018 entitled simply World
  592. 22:52Models.
  593. 22:53They built a small agent to play an
  594. 22:55adaptation of the iconic game Doom
  595. 22:58called VizDoom, where the simplified
  596. 23:00goal of the agent is to survive for as
  597. 23:03long as possible.
  598. 23:05They enforced a deliberately strict
  599. 23:06constraint. The policy model should be
  600. 23:09trained without any interaction with the
  601. 23:11real game environment. Such interactions
  602. 23:14would only be allowed for training the
  603. 23:16world model.
  604. 23:17Of course, this is not a practical
  605. 23:19constraint in video games, which are
  606. 23:21relatively cheap to run, but at the time
  607. 23:23in 2018, it was a promising proof of
  608. 23:26concept that could one day materialize
  609. 23:28for training physical robots.
  610. 23:31So, the VizDoom agent went through two
  611. 23:33training stages before deployment.
  612. 23:36First, training a world model. They
  613. 23:38started with a randomly initialized
  614. 23:39policy and played 10,000 games. In other
  615. 23:43words, they issued random actions until
  616. 23:46the agent got killed 10,000 times.
  617. 23:49In the process, the world model got to
  618. 23:51observe the reaction of the real
  619. 23:52environment and recreate a smaller,
  620. 23:55compressed version of it, where action
  621. 23:57consequences unfold in a latent space
  622. 24:00instead of pixel frames.
  623. 24:02Once the world model was trained, it was
  624. 24:04then frozen and used as a simulation
  625. 24:06environment to train the policy model.
  626. 24:09At the end of the second training stage,
  627. 24:11the agent was deployed in the real
  628. 24:12environment and was able to successfully
  629. 24:15play VizDoom, despite never having seen
  630. 24:17the actual game before.
  631. 24:20Subsequent research scaled MBRL in
  632. 24:22various ways.
  633. 24:24For instance, the Dreamer model family
  634. 24:26from DeepMind worked their way up to
  635. 24:28mining a diamond in Minecraft from
  636. 24:31scratch.
  637. 24:32This had been an unsolved problem in AI
  638. 24:35for years, and prior attempts needed
  639. 24:37human gameplay videos.
  640. 24:40And today, we're starting to see it
  641. 24:41applied to embodied robots as well.
  642. 24:44For instance, World Gymnast is a
  643. 24:46tabletop manipulation robot that was
  644. 24:49exclusively trained inside a world model
  645. 24:52and deployed on physical hardware
  646. 24:54without any real-world fine-tuning.
  647. 24:57In the example so far, the world model
  648. 24:59is a training time tool discarded once
  649. 25:02the policy is trained.
  650. 25:04But world models can also stick around
  651. 25:06at inference time. The agent can
  652. 25:08actively use them in real time to unfold
  653. 25:11the effects of multiple candidate
  654. 25:12actions and pick the one with the most
  655. 25:15promising outcome. This is real time
  656. 25:17planning. The most common planning
  657. 25:19paradigm for agents is model predictive
  658. 25:22control inspired from 1970s chemical
  659. 25:25engineering.
  660. 25:27Here's one way to implement it.
  661. 25:30Before taking an action in the real
  662. 25:31world, an agent builds an entire tree of
  663. 25:34possibilities rooted in the current
  664. 25:36world state.
  665. 25:38The edges are candidate actions and the
  666. 25:40nodes are hypothetical future world
  667. 25:42states as computed by a world model.
  668. 25:45It's like being inside the mind of an
  669. 25:47extremely anxious person trying to
  670. 25:49foresee how the future might unfold.
  671. 25:52After building this tree, the agent
  672. 25:54identifies the branch with the best
  673. 25:55estimated outcome and performs only the
  674. 25:58very first action in the real world.
  675. 26:01It then uh throws away the rest of the
  676. 26:03branch and builds an entirely new tree
  677. 26:06rooted in the updated world state
  678. 26:09communicated by the environment.
  679. 26:11Working in a latent space is extremely
  680. 26:14important for speed since decision trees
  681. 26:16are built in real time.
  682. 26:18This is the idea behind MuZero, a 2019
  683. 26:21algorithm from DeepMind trained to play
  684. 26:24board games and Atari video games.
  685. 26:27But planning is now coming to embodied
  686. 26:29AI as well, like VIPAC, the model we
  687. 26:33discussed in the implementation section,
  688. 26:35as an exponent for the family of
  689. 26:37predictive world models.
  690. 26:39Technically, VIPAC has a branching
  691. 26:41factor of one but follows the same
  692. 26:44principle of model predictive control.
  693. 26:47Interestingly, it uses subgoal images
  694. 26:49provided by the environment to judge how
  695. 26:52good a certain branch is.
  696. 26:55This means it doesn't even need rewards
  697. 26:57like the ones in reinforcement learning.
  698. 26:59You just show it an image of what you
  699. 27:01want and it'll do it in a task-agnostic
  700. 27:04way.
  701. 27:05On the one hand, this is Black Mirror
  702. 27:07stuff, but on the other, imagine how
  703. 27:10helpful this would be for an autonomous
  704. 27:12vehicle that needs to make a quick,
  705. 27:14critical decision.
  706. 27:16You can start doing closed-loop
  707. 27:18scenarios where you can start saying,
  708. 27:20"Okay, if I want to drive the car and
  709. 27:21turn right, what will happen next and
  710. 27:23what will I see?" Now, show me 50
  711. 27:25different versions of that and then pick
  712. 27:27the most, you know, useful or reliable
  713. 27:29one. Before we wrap up, I want to
  714. 27:31mention that world models don't have to
  715. 27:33be visual. The same definition, given a
  716. 27:35state and an action, predict the next
  717. 27:37state, works in any environment where
  718. 27:39you have well-defined states and
  719. 27:42actions. A nice recent example is the
  720. 27:44Coded World Model from Meta.
  721. 27:46The world here is a software
  722. 27:48environment, a repository, a program, a
  723. 27:51terminal. The action is something like a
  724. 27:54pull request. And the world model
  725. 27:56predicts the next state, the variable
  726. 27:58values, output, execution traces, and so
  727. 28:01on. So, why not just execute the code
  728. 28:04and observe its effects instead of
  729. 28:06building a world model of the software
  730. 28:08environment?
  731. 28:10Well, it might take a full hour to run
  732. 28:12the entire test suite, and certain bugs
  733. 28:14might only surface in production, going
  734. 28:17unnoticed during development.
  735. 28:19A world model could be quicker and could
  736. 28:21catch potential issues without causing
  737. 28:23irreparable damage. We covered a lot of
  738. 28:26ground, but I hope that the through line
  739. 28:28is clear. A world model predicts how a
  740. 28:30certain action will change the state of
  741. 28:32the world. There's a lot of ambiguity
  742. 28:34around this term because the interface
  743. 28:36is so general and can be applied across
  744. 28:38so many industries. If you want to hear
  745. 28:41more from TJ Galda, the full interview
  746. 28:43is available to my YouTube and Patreon
  747. 28:45members. Stay tuned for more videos on
  748. 28:48world models.

About this transcript

This page contains the full transcript of But what exactly are world models? by Julia Turc, generated from the public captions YouTube serves with the video. The transcript has 4,376 words across 748 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.