YouTube2Text

Claude Opus 5, GPT 6 hack, Flux 3, new Gemini, quantum breakthrough, new Qwen: AI NEWS — Transcript

by AI Search · 8,052 words · 1,204 segments · language en · Watch on YouTube

Full transcript

  1. 0:00AI never sleeps and this week has been
  2. 0:03absolutely insane.
  3. 0:06Anthropic unleashes Claude Opus 5. Open
  4. 0:09AI releases some new and very useful
  5. 0:11features for Chat GPT. We have a new
  6. 0:13open-source image generator and editor.
  7. 0:16So, this is super flexible. One of the
  8. 0:17top platforms for hosting open models
  9. 0:20called Hugging Face gets hacked by Open
  10. 0:22AI's internal models. But, get this, it
  11. 0:25was an open-source model GLM 5.2 that
  12. 0:28helped detect and fix Speaking of GLM
  13. 0:315.2, this is a text-only model, but now
  14. 0:34it finally gets vision capabilities. We
  15. 0:37have a new AI for creating consistent
  16. 0:39multi-shot videos where you can specify
  17. 0:41the exact cuts and transitions. Alibaba
  18. 0:44releases their best image model yet,
  19. 0:46Qwen Image 3, and teases their massive
  20. 0:49frontier model Qwen 3.8, which will be
  21. 0:52open-source. Google releases their
  22. 0:54latest Gemini models, which are
  23. 0:56incredibly fast. We have a new
  24. 0:58open-source interactive world model for
  25. 1:00Minecraft. We also have some ridiculous
  26. 1:03robot demos and a lot more. So, let's
  27. 1:05jump right in. First up, Microsoft
  28. 1:08released a really powerful image model
  29. 1:10called Mage Flow. And get this, it
  30. 1:13doesn't just generate images, but it can
  31. 1:14also edit existing images just like GPT
  32. 1:17Image or Nano Banana. First of all, here
  33. 1:20are some text-to-image examples. As you
  34. 1:22can see, it can generate some really
  35. 1:23realistic images of different-looking
  36. 1:26people like this. Here are some other
  37. 1:28random photorealistic images, and as you
  38. 1:31can see, for the most part, this looks
  39. 1:33pretty good. And this is also great with
  40. 1:35text. So, here are some examples of its
  41. 1:38generations with text and a ton of
  42. 1:40different elements. And as you can see,
  43. 1:41it can easily generate posters,
  44. 1:43infographics, or other related content.
  45. 1:46It's also multilingual, so this can
  46. 1:48easily generate different languages like
  47. 1:50Chinese. Now, like I said, instead of
  48. 1:53just generating images, this can also
  49. 1:55edit existing images. So, for example,
  50. 1:57you can easily replace the background of
  51. 1:59an image or zoom in or zoom out on an
  52. 2:02image or change the camera angle of an
  53. 2:04image or even change the pose of a
  54. 2:06character just with a text prompt. So,
  55. 2:09very similar to Nano Banana or GPT
  56. 2:11image. You can easily change the tone or
  57. 2:14weather. For example, you can turn a
  58. 2:16photo from day to night or you can also
  59. 2:18turn a realistic photo into another
  60. 2:20artistic style. And of course, you can
  61. 2:22also do virtual try-ons with this or add
  62. 2:25text to an image, change the expressions
  63. 2:27of characters, or micro-edit certain
  64. 2:30things like changing the look of
  65. 2:31someone's hair, or changing the
  66. 2:33composition or color of a certain
  67. 2:35object. This also has ControlNet
  68. 2:37inherently built in. For example, you
  69. 2:39can turn a photo into a sketch or line
  70. 2:41art, or you can also turn a photo into a
  71. 2:44pose skeleton. And of course, you can
  72. 2:46also do the reverse. You can take a pose
  73. 2:48skeleton and generate a photo from that.
  74. 2:50Or you can also generate Canny edge maps
  75. 2:53like this. And vice versa, you can take
  76. 2:55a Canny edge map and generate a photo.
  77. 2:57This can also do depth and segmentation
  78. 2:59like this. So, a super flexible image
  79. 3:02editor. Now, this is a fairly
  80. 3:04lightweight 4 billion parameter model,
  81. 3:07and they also include a four-step turbo
  82. 3:09model, which allows you to generate
  83. 3:11images way faster in only four steps.
  84. 3:13Apparently, it takes less than a second
  85. 3:15to generate an image with the turbo
  86. 3:17model. And then it takes a bit over a
  87. 3:19second to edit an existing image with
  88. 3:21the turbo model. Their self-reported
  89. 3:23benchmarks also look really impressive.
  90. 3:25So, if you look at this chart, you can
  91. 3:27see that Mage Flow even scores better
  92. 3:29than Flex Client or Quinn image.
  93. 3:31Although, they haven't included the best
  94. 3:33open models out there, Crea 2 and
  95. 3:35Ideogram. And then in terms of image
  96. 3:37editing, then you can see that Mage Flow
  97. 3:39does score slightly behind Flex Client
  98. 3:42in terms of quality, but it's pretty
  99. 3:43close, and the turbo model is even
  100. 3:45faster. The awesome thing is they've
  101. 3:47released this already. If you click on
  102. 3:49this explore models page, it includes
  103. 3:53various models you can choose from. The
  104. 3:55main Mage flow model is 17.5 GB in size,
  105. 3:59so you should be able to fit this in mid
  106. 4:00to high-end GPUs. And if not, because
  107. 4:03this is open source, I'm sure there's
  108. 4:05going to be more quantized or compressed
  109. 4:07versions of this in the future as well.
  110. 4:08And if you scroll down a bit here, it
  111. 4:10contains all the instructions on how to
  112. 4:12download and run this locally on your
  113. 4:14computer. If you're interested in
  114. 4:16reading further, I'll link to this main
  115. 4:17page in the description below. Also this
  116. 4:20week, this AI is super helpful for
  117. 4:22creating videos. So it's called Shot
  118. 4:24Plan, and this can generate videos with
  119. 4:26precisely timed cuts, fades, and camera
  120. 4:29movements. So here are some examples.
  121. 4:32For the input, it takes in a text prompt
  122. 4:34plus descriptions of each shot, and you
  123. 4:36can also specify exactly where the cuts
  124. 4:39occur, and the output is a full
  125. 4:42multi-shot video that follows all of
  126. 4:44these instructions. Or here's another
  127. 4:46example where this is the high-level
  128. 4:49prompt, and then you also add in
  129. 4:50descriptions for each separate shot,
  130. 4:53plus the frame which the cuts occur, and
  131. 4:55here is your final video. As you can
  132. 4:57see, it's able to follow everything very
  133. 5:00well, and the characters remain
  134. 5:01consistent the whole time. So this is
  135. 5:04kind of like the multi-shot feature in
  136. 5:06Kling 3. And then here's another
  137. 5:08example. Now, instead of just doing hard
  138. 5:09cuts, this AI is also able to understand
  139. 5:12crossfade transitions, or soft cuts as
  140. 5:15you can see here. It's also able to
  141. 5:17understand different camera movements.
  142. 5:19So here's an example of circle left, and
  143. 5:21then here is circle right. And of
  144. 5:23course, it also understands camera
  145. 5:24movements like zooming in or zooming
  146. 5:26out, and you can specify exactly at
  147. 5:28which frame you want this camera
  148. 5:30movement to occur. Here's another
  149. 5:31example where we specify a pullback at
  150. 5:35exactly frame 40. And as you can see,
  151. 5:37it's able to follow this very well. Or
  152. 5:40here's an example of truck right from
  153. 5:42frame 25. And again, it understands this
  154. 5:45very well. And if you look at these
  155. 5:47multi-shot generation benchmarks, Shot
  156. 5:49Plan on average performs better in terms
  157. 5:52of consistency and narrative compared to
  158. 5:54the other competitor models. The awesome
  159. 5:56thing is they released this already. So,
  160. 5:58at the top of the page, if you click on
  161. 6:00this code button and you scroll down a
  162. 6:02bit, here it contains all the
  163. 6:03instructions on how to download and run
  164. 6:05this locally on your computer. They also
  165. 6:08released the code on how to train this
  166. 6:10as well. So, this is fully open source.
  167. 6:12Notice that they released variants for 1
  168. 6:152.2 or 1 2.1. Now, for 1 2.2, that one
  169. 6:19is like 28.6 GB in size, so this will
  170. 6:22likely only fit on high-end GPUs. If
  171. 6:25you're interested in reading further,
  172. 6:26I'll link to this main page in the
  173. 6:28description below. Also this week, we
  174. 6:30have a new AI called Homey. In the
  175. 6:32simplest sense, this generates videos
  176. 6:35containing specific references which you
  177. 6:37input. This can include multiple people
  178. 6:40and objects. And as you can see, it's
  179. 6:42really good at preserving the
  180. 6:44consistency of all your input
  181. 6:46references. Note that not only can this
  182. 6:48do photorealistic references, but also
  183. 6:513D animation like this. And you can
  184. 6:54input a ton of different complex
  185. 6:56objects. For example, this kimono plus
  186. 6:59this toy here has a pretty complex
  187. 7:00design, but it's able to preserve the
  188. 7:02consistency of these really well. And
  189. 7:04so, from this you can easily pump out a
  190. 7:07ton of influencer content where you just
  191. 7:09input a photo of the influencer plus any
  192. 7:11product you want to get it to promote
  193. 7:13and then get this AI to mass produce all
  194. 7:16these UGC style videos. The really
  195. 7:18powerful part is you can input multiple
  196. 7:21views of a certain character or object
  197. 7:23to give it even more consistency.
  198. 7:25Especially for objects with really
  199. 7:27complicated designs, it's best if you
  200. 7:29have multiple views of the object so
  201. 7:31that they look consistent in the video
  202. 7:32at all angles. Another really cool
  203. 7:35feature is you can also input an OCR map
  204. 7:38which shows where text should appear in
  205. 7:40an object. And this allows the model to
  206. 7:42generate objects with text in them a lot
  207. 7:44more accurately. Here are some
  208. 7:46additional examples for your reference.
  209. 7:48Now, I featured a ton of these reference
  210. 7:50to video generators on my channel
  211. 7:52before, such as Vace and Phantom, but if
  212. 7:54you compare all these models with this
  213. 7:56new Homie model, then Homie seems to be
  214. 7:59a lot more consistent and faithful with
  215. 8:01the least amount of errors. And this is
  216. 8:03especially apparent if you have
  217. 8:05characters or objects with really
  218. 8:06complex designs. Again, for Homie, you
  219. 8:09can upload multiple views of that
  220. 8:11character to give it even more
  221. 8:12consistency, whereas the other models
  222. 8:14can't do this, and as a result, they are
  223. 8:17a lot more error prone. At the top of
  224. 8:18the page, they have released the code
  225. 8:20and the models to this already, so if
  226. 8:22you click on this code button, and you
  227. 8:24scroll down a bit, here it contains all
  228. 8:26the instructions on how to download and
  229. 8:28run this locally on your computer. The
  230. 8:30total size of everything is 37 GB, so
  231. 8:33you'll need a high-end GPU to run this,
  232. 8:35or you can also wait for some more
  233. 8:37compressed versions. Note that this is
  234. 8:39based off of one 2.1 and Phantom, and
  235. 8:41it's under the Apache 2 license, which
  236. 8:43has very minimal restrictions. You can
  237. 8:45even use this for commercial purposes.
  238. 8:47If you're interested in reading further,
  239. 8:49I'll link to this main page in the
  240. 8:50description below. Also this week, this
  241. 8:53story is pretty crazy. OpenAI revealed a
  242. 8:56pretty alarming security incident where
  243. 8:58one of its internal models kind of
  244. 9:00escaped, got access to the internet, and
  245. 9:02hacked Hugging Face, which is like one
  246. 9:04of the top platforms for hosting
  247. 9:06open-source models. So, here are the
  248. 9:08details behind this. OpenAI was
  249. 9:10basically testing GPT 5.6 Soul and an
  250. 9:13even more uncapable internal model with
  251. 9:16some cybersecurity evaluations,
  252. 9:18including a benchmark called Exploit
  253. 9:20Gym. Now, these models were placed
  254. 9:22inside what was supposed to be an
  255. 9:24isolated sandbox environment for testing
  256. 9:26only, with only network access limited
  257. 9:29to the ability to install packages
  258. 9:32through an internally hosted third-party
  259. 9:34software. So, it can't really access
  260. 9:36anything else outside of this scope.
  261. 9:38Now, instead of just solving the
  262. 9:40challenge inside this isolated sandbox
  263. 9:42environment, what the models decided to
  264. 9:44do is identify and chain vulnerabilities
  265. 9:48in the package system, which gave it
  266. 9:49broader network access. And from there,
  267. 9:52they eventually reached a device with
  268. 9:54access to the public internet, and then
  269. 9:56they also hacked into Hugging Face's
  270. 9:58production infrastructure to obtain the
  271. 10:00answers from the benchmark, instead of
  272. 10:02actually trying to solve the test
  273. 10:04itself. So, here it says, "After gaining
  274. 10:06internet access, the models searched for
  275. 10:09and successfully found ways to gain
  276. 10:10access to secret information it could
  277. 10:12use to cheat the evaluation. It chained
  278. 10:15together multiple attack vectors,
  279. 10:17including using stolen credentials and
  280. 10:20zero-day vulnerabilities, which are
  281. 10:22really hard to find, by the way. These
  282. 10:23are like hidden security flaws that are
  283. 10:25not known to anyone. And it used this to
  284. 10:27find and hack the answers to the
  285. 10:29benchmark from Hugging Face." Now, what
  286. 10:31OpenAI probably intentionally omitted in
  287. 10:35this article is that Hugging Face only
  288. 10:37detected this using the open-source
  289. 10:39GLM-5.2.
  290. 10:41So, they used this to detect the hack
  291. 10:43and fix the vulnerability, because
  292. 10:44OpenAI's model would refuse to do it.
  293. 10:47This is like the drawback of all these
  294. 10:49closed-source models like GPT and
  295. 10:52Claude, where there's a lot of nerfing
  296. 10:53going on. For example, Claude Fable
  297. 10:55would just 100% reject answering any
  298. 10:58questions related to cybersecurity or
  299. 11:01biology, even if you're just asking it
  300. 11:03to fix an issue. And you know, this
  301. 11:05again emphasizes the power and the need
  302. 11:07for open-source models like GLM-5.2 and
  303. 11:09Kimika 3 and others. So, after OpenAI
  304. 11:12detected this unusual activity
  305. 11:14internally, and then after Hugging Face
  306. 11:16also used GLM-5.2 to detect the incident
  307. 11:19on their side, they're now working
  308. 11:21together to investigate this incident
  309. 11:23and tighten network controls, and also
  310. 11:25improve monitoring of future model
  311. 11:27evaluations. The most concerning part is
  312. 11:30that as models get more and more
  313. 11:32intelligent, if you tell it to do a task
  314. 11:34or achieve a certain goal, instead of
  315. 11:36actually doing it directly, it might
  316. 11:38find a way to exploit the system and
  317. 11:41cheat to achieve your goal faster. And
  318. 11:43this could lead to unintended or
  319. 11:45malicious actions. Anyway, let me know
  320. 11:47in the comments below what you think of
  321. 11:49this incident. Also this week, OpenAI
  322. 11:51launches health in ChatGPT. This
  323. 11:54basically turns ChatGPT from just a
  324. 11:56regular chatbot to something that can
  325. 11:58actually understand your personal health
  326. 12:00history. So with your permission, you
  327. 12:03can now connect Apple Health and
  328. 12:05supported medical records, including
  329. 12:07information from certain US hospital
  330. 12:09systems and health apps, and ChatGPT can
  331. 12:11then look at things like your
  332. 12:13medications, lab results, sleep,
  333. 12:15exercise, and other health data. And
  334. 12:17this is the important part. With the
  335. 12:19regular ChatGPT, you could just upload
  336. 12:21your lab report and then ask it
  337. 12:23questions about it, but it doesn't have
  338. 12:24context about your previous health
  339. 12:26history. But with health in ChatGPT, if
  340. 12:28you upload all your records, your
  341. 12:30medications, your past visits, then it
  342. 12:32has a more comprehensive understanding
  343. 12:34of you. So when you ask it questions
  344. 12:36about lab reports or medications or
  345. 12:38whatever, it can give you more accurate
  346. 12:40answers tailored to you. Here OpenAI
  347. 12:42says they're beginning to roll out
  348. 12:44health to ChatGPT users 18 and older in
  349. 12:47the US across all plans, including the
  350. 12:50free plan, today. In the ChatGPT app, on
  351. 12:52the sidebar, you should be able to see
  352. 12:54health over here. And then once you
  353. 12:56click on that, you can connect ChatGPT
  354. 12:58to your health data to give it more
  355. 13:00context. If you're interested in reading
  356. 13:02further, I'll link to this main page in
  357. 13:04the description below. Also this week,
  358. 13:06Black Forest Labs teases their latest
  359. 13:08model, Flux 3. And this is more
  360. 13:10ambitious than their previous models,
  361. 13:13which are just image generators. Here
  362. 13:15they say that Flux 3 is one unified
  363. 13:17multimodal model designed to work across
  364. 13:19images, video, audio, and even action
  365. 13:22prediction for robotics. Basically,
  366. 13:25instead of training one AI to just
  367. 13:26understand pictures and training another
  368. 13:28model to generate video, here they're
  369. 13:30just merging everything together into a
  370. 13:32single model. Plus, the video also has
  371. 13:35audio built in, just like Seed Dance and
  372. 13:37LTEX 2.3. Here are some of its core
  373. 13:40capabilities. So, this can of course do
  374. 13:42text to video, but also image to video,
  375. 13:44or you can also just input images as
  376. 13:47references. This can also do video to
  377. 13:49video, so you can upload an existing
  378. 13:51video and then just edit it with natural
  379. 13:53language, just like Gemini Omni. You can
  380. 13:56generate videos of different aspect
  381. 13:58ratios, plus it also has really strong
  382. 14:00topography generation. So, it has no
  383. 14:02problems including text in videos. Now,
  384. 14:05they claim that Flux 3 can generate
  385. 14:07videos with audio at up to 20 seconds in
  386. 14:10length. Here, the generations that
  387. 14:12they're showing us so far are just 720p.
  388. 14:14There's no clear indication whether you
  389. 14:16can generate higher resolution like 1080
  390. 14:18or 4K. And if you look at their
  391. 14:20self-reported results comparing
  392. 14:2210-second videos in 720p with audio,
  393. 14:25then here it says that Flux 3
  394. 14:27generations were slightly more preferred
  395. 14:29than even Gemini Omni Flash or Seed
  396. 14:31Dance 2.0. And it also beats these other
  397. 14:33competitor models like Happy Horse,
  398. 14:35Kling, and Grok Imagine Video. Now, just
  399. 14:38from these preliminary videos, at least
  400. 14:40to me, it doesn't seem like it's even
  401. 14:41close to Seed Dance quality. Seed Dance
  402. 14:44just has much better prompt and physics
  403. 14:46understanding, and it can generate much
  404. 14:48better action scenes and motion and
  405. 14:50different styles compared to Flux 3. So,
  406. 14:52something looks fishy here. Now, like I
  407. 14:54said, this is a multimodal model, so in
  408. 14:56addition to generating videos, this can
  409. 14:58also generate images. Here are some
  410. 15:00sample generations for your reference.
  411. 15:02This can generate a variety of different
  412. 15:04art styles. Now, similar to their
  413. 15:06previous releases, it seems that Flux
  414. 15:09will offer a more powerful full or max
  415. 15:12model, which would be paid and closed
  416. 15:13source and only available through their
  417. 15:15APIs. But, they are also planning to
  418. 15:18release a lower quality dev version,
  419. 15:20which will be open weights. So, this is
  420. 15:22the one that you could potentially
  421. 15:23download and run locally on your
  422. 15:25computer. Currently, this is just a
  423. 15:27preview. You can't use this yet, but you
  424. 15:29could request early access using this
  425. 15:31link. If you're interested in reading
  426. 15:33further, I'll link to this main page in
  427. 15:35the description below. Also this week,
  428. 15:37we have another set of open-source
  429. 15:39models. This time it's by Poolside AI,
  430. 15:42and the family is called Laguna S 2.1.
  431. 15:45Now, this is an open weight coding model
  432. 15:46designed to keep working on really
  433. 15:48difficult problems for much longer than
  434. 15:51typical AI models. First of all, let's
  435. 15:53go over the specs of this. So, this is a
  436. 15:55118 billion parameter mixture of experts
  437. 15:58model. So, think of it as like a team of
  438. 15:59specialist AIs working together. And
  439. 16:02when you use it, only 8 billion
  440. 16:03parameters are active. So, this is
  441. 16:05fairly efficient. This has a context
  442. 16:07window of up to a million tokens. So,
  443. 16:09you can jam-pack a ton of information
  444. 16:11into your prompt at once. This is
  445. 16:12roughly like 700,000 words. And you can
  446. 16:15operate it in thinking or non-thinking
  447. 16:17modes. Note that they also released a
  448. 16:19smaller XS variant, which is only 33
  449. 16:22billion parameters with 3 billion total
  450. 16:24parameters and 3 billion active
  451. 16:26parameters. And this model is trained
  452. 16:28like an agentic agent. It's trained to
  453. 16:30verify its work, backtrack when
  454. 16:32necessary, verify its work and fix any
  455. 16:35errors, and keep going until it
  456. 16:37successfully achieves your goal. And if
  457. 16:39you look at their self-reported
  458. 16:41benchmarks comparing models of similar
  459. 16:43sizes or even much larger like DeepSeek
  460. 16:45V4 or Chimera 3, this new Laguna S 2.1
  461. 16:49is not bad. Now, here if you look at
  462. 16:51their self-reported benchmark score for
  463. 16:53Deep Sweep, you can see that Laguna S
  464. 16:552.1 scores 40% slightly above GLM 5.2,
  465. 16:59but note that GLM is like almost seven
  466. 17:01times the size. So, this is quite an
  467. 17:03impressive performance. Definitely one
  468. 17:05of the strongest open weights models
  469. 17:07that are around 100 billion parameters
  470. 17:09in size. Now, at least for me, I'm
  471. 17:11actually more interested in this XS
  472. 17:13version because this is only 33 billion
  473. 17:15parameters. So, this could potentially
  474. 17:17fit on a decent consumer GPU. But after
  475. 17:20some testing, it doesn't seem to be
  476. 17:22nearly as good as the current leading
  477. 17:24model of a similar size, which would be
  478. 17:26Qwen 3.6 35B or Qwen 3.6 27B. So at
  479. 17:30least for me, Qwen is still the best
  480. 17:32medium-sized model that can be run
  481. 17:34locally. The awesome thing is they've
  482. 17:36released the model weights to this
  483. 17:37already. So if you click on this Hugging
  484. 17:39Face link, note that the total size of
  485. 17:41this is 235 GB, which is actually not
  486. 17:44bad. You could probably fit this on just
  487. 17:46one DGX Spark. They also already have a
  488. 17:49ton of quantized versions of this. So
  489. 17:51for example, this NVFP4 version is only
  490. 17:5471 GB in size. If you're interested in
  491. 17:57reading further, I'll link to this main
  492. 17:59page in the description below. If you
  493. 18:01want to supercharge your content
  494. 18:02creation, definitely check out
  495. 18:04Higgsfield, the sponsor of this video.
  496. 18:06Think of it as an all-in-one creation
  497. 18:08platform built specifically for
  498. 18:10creators. Instead of jumping between a
  499. 18:12bunch of different tools, Higgsfield
  500. 18:13gives you access to some of the world's
  501. 18:15leading models in one place, including
  502. 18:17Seed Dance, Kling, Nano Banana, and GPT
  503. 18:20Image. But what makes it really
  504. 18:22interesting is that Higgsfield also has
  505. 18:24its own tools built for creative
  506. 18:26workflows. For example, there's
  507. 18:28Higgsfield Supercomputer, which is
  508. 18:29basically a general-purpose AI agent for
  509. 18:32content creation. You can give it one
  510. 18:34prompt and it can help with the whole
  511. 18:35process, finding an idea, creating the
  512. 18:38product concept, writing the brand
  513. 18:39direction, generating the visuals, and
  514. 18:41making the full video. And then there's
  515. 18:43also Higgsfield MCP.
  516. 18:45You can connect the content generation
  517. 18:47abilities of Higgsfield with an AI agent
  518. 18:49like Claude Code, ChatGPT, or Hermes,
  519. 18:52and they do the thinking and planning
  520. 18:54while Higgsfield executes the actual
  521. 18:56generation. They also have Marketing
  522. 18:58Studio, which is really useful for
  523. 19:00making marketing content. You can paste
  524. 19:02a product link or upload a product
  525. 19:04image, and it can generate multiple ad
  526. 19:06formats in one workflow, like UGC
  527. 19:08videos, tutorials, unboxings, product
  528. 19:10reviews, and more. One product can
  529. 19:13instantly turn into a whole campaign.
  530. 19:15They also have Cinema Studio, which is
  531. 19:16built as a full end-to-end AI filmmaking
  532. 19:19pipeline. This is especially useful if
  533. 19:21you want more cinematic control. Instead
  534. 19:23of just typing a prompt and hoping the
  535. 19:25video looks good, Cinematic Studio lets
  536. 19:28you plan scenes, control the camera, add
  537. 19:30specific characters, reuse locations,
  538. 19:33and keep everything consistent across
  539. 19:34the whole project. From idea to final
  540. 19:37output, Higgsfield gives you way more
  541. 19:39control over the whole creative process.
  542. 19:41Whether you're making ads, social
  543. 19:43videos, AI influencers, product
  544. 19:45launches, cinematic clips, or any other
  545. 19:47content, this is one of the easiest
  546. 19:49platforms to start creating with AI. Try
  547. 19:51Higgsfield today using the link in the
  548. 19:53description below. In robotics news this
  549. 19:55week, Unitree releases an absolute beast
  550. 19:58called the Super Athlete AS2W. This is a
  551. 20:01wheel-legged dog robot built for
  552. 20:04incredibly fast movement and just
  553. 20:06dashing through rough, uneven terrain.
  554. 20:09This weighs about 25 kg with 16° of
  555. 20:12freedom, and it can run up to 6 m/s,
  556. 20:15which is roughly 21 km/h or 13 mph.
  557. 20:20Incredibly fast for a robotic dog. I
  558. 20:22mean, imagine this thing chasing you
  559. 20:24down. You would have no way of escaping
  560. 20:26this. Reminds me of that Black Mirror
  561. 20:27episode. This can carry a load of 16 kg,
  562. 20:31plus this can walk up to 30 km on a
  563. 20:34single charge. This can operate in
  564. 20:36temperatures as low as -20° and as hot
  565. 20:40as 55° C. And here's a demo of its
  566. 20:43strength. Despite only weighing 25 kg,
  567. 20:46even if you have like three people
  568. 20:47standing on top of it, it can still move
  569. 20:50and stay balanced. As you can see here,
  570. 20:52this is also IP54 water resistant, so
  571. 20:55this can wade and splash through water
  572. 20:57without slipping or stopping. And it is
  573. 20:59just so flexible and dynamic. You can
  574. 21:02see it being able to flip and roll over
  575. 21:05all this uneven terrain and recover
  576. 21:07pretty much seamlessly. Also this week
  577. 21:09the Chinese labs continue cooking some
  578. 21:11massive models. If you haven't been up
  579. 21:14to date with what's going on, last week
  580. 21:16Moonshot AI dropped Kimi K3, which is a
  581. 21:18massive 2.8 trillion parameter model.
  582. 21:22And this is like one of the best models
  583. 21:23out there right now. Well, shortly
  584. 21:25afterwards Alibaba also teased their
  585. 21:28latest Qwen 3.8 and this also has a
  586. 21:31massive 2.4 trillion parameters. The
  587. 21:33best part is here they say it's going to
  588. 21:35be open weights, which is fantastic. For
  589. 21:38now you can try it via their API. They
  590. 21:41haven't released any benchmarks on this
  591. 21:42yet. There's no like blog post or
  592. 21:44article about this, but they do say that
  593. 21:46it's compatible to leading frontier
  594. 21:48models second only to Fable 5. Although
  595. 21:51from some initial reports, it doesn't
  596. 21:53seem to be as good as Kimi K3. Anyways,
  597. 21:55that's all the info we have for now.
  598. 21:57They haven't released an official page
  599. 21:59on this with more specs and benchmarks,
  600. 22:01but if you are interested, I'll link to
  601. 22:03this page in the description below where
  602. 22:05you can try out Qwen 3.8 max preview
  603. 22:07using their paid token plan. Also this
  604. 22:09week one of the best open source models
  605. 22:11out there JLM 5.2 finally has vision. So
  606. 22:15if you're not familiar with JLM 5.2, I
  607. 22:17did a full review video on it right when
  608. 22:19it came out. So see this video if you're
  609. 22:22interested. This is a super capable
  610. 22:24model and definitely one of my favorite
  611. 22:26models, but the main drawback of this is
  612. 22:28that it does not have vision
  613. 22:30capabilities. This is a text only model.
  614. 22:32Well, this week we now have an
  615. 22:34unofficial version of JLM 5.2 with
  616. 22:37vision. And this is created by Base 10.
  617. 22:39So what they did is they basically took
  618. 22:42a vision encoder which was created by
  619. 22:44the Kimi team and they merge that into
  620. 22:47the architecture of JLM 5.2. And this
  621. 22:50essentially gives eyes to JLM 5.2. Note
  622. 22:53that this is just an unofficial version.
  623. 22:55This is not published by the ZAI team or
  624. 22:58the Kimi team themselves. Note that this
  625. 23:00is an NVFP4 quantized version, but if
  626. 23:03you click on files and versions, I mean
  627. 23:04the base model is huge, so even this FP4
  628. 23:07version is 466 GB. You'll need some
  629. 23:10high-end hardware like stacking maybe
  630. 23:12two DGX Sparks in order to actually fit
  631. 23:15this locally. But if you are interested,
  632. 23:18I'll link to this main page in the
  633. 23:20description below. Also this week, what
  634. 23:22I think is the most exciting update from
  635. 23:24OpenAI is they now introduced GPT live
  636. 23:27voice inside the GPT desktop app. Now
  637. 23:30previously to use the GPT live voice,
  638. 23:33you need to use it through their online
  639. 23:35chat interface, but now you can use this
  640. 23:37directly on the ChatGPT desktop app. And
  641. 23:39this opens up a ton of new
  642. 23:40possibilities. If you're not familiar
  643. 23:42with the ChatGPT desktop app, I highly
  644. 23:44recommend you download it and try it
  645. 23:46out. It's free for you to try. This lets
  646. 23:48you work with multiple projects and
  647. 23:50files on your computer at once. And
  648. 23:52traditionally you would need to, you
  649. 23:54know, open each project in a thread and
  650. 23:56chat with it like a chatbot, but now
  651. 23:58with GPT live, you can just talk to this
  652. 24:00app and get it to coordinate different
  653. 24:02tasks in the app all at the same time.
  654. 24:05Imagine like being able to vibe code
  655. 24:07multiple projects at once just by
  656. 24:10speaking to your computer. You don't
  657. 24:11even need to type, you can just do your
  658. 24:13own thing. That's the power of this new
  659. 24:16GPT live feature within the ChatGPT
  660. 24:18desktop app. Now currently it says that
  661. 24:20this live voice feature is available in
  662. 24:22the desktop app for paid users, so free
  663. 24:25users will not have access to this yet.
  664. 24:27It seems that at least for me it's
  665. 24:29available on the Mac version, but not
  666. 24:31the Windows desktop app yet. Now the
  667. 24:33main limitation here is that only one
  668. 24:35voice chat can be active at a time. And
  669. 24:38of course you also have usage limits.
  670. 24:40Using the live voice will also take up
  671. 24:42credits and it depends on which plan you
  672. 24:44have. But the bigger idea here is that
  673. 24:47OpenAI is trying to turn this live voice
  674. 24:49into a hands-free control layer for
  675. 24:51ChatGPT. This allows you to manage
  676. 24:53increasingly complicated work simply by
  677. 24:56talking to it. If you're interested in
  678. 24:58reading further, I'll link to this main
  679. 25:00page in the description below. Also this
  680. 25:02week, we have some pretty exciting news
  681. 25:05in quantum computing from Google. So
  682. 25:07apparently Google has developed a way
  683. 25:09for a quantum computer to automatically
  684. 25:11retune itself while it's running. And
  685. 25:13this matters because quantum computers
  686. 25:15are extremely sensitive machines. Tiny
  687. 25:18changes in temperature or even
  688. 25:20electronics or hardware can slowly push
  689. 25:23its calculations out of calibration. And
  690. 25:25if this happens, then currently
  691. 25:27engineers might need to stop the
  692. 25:29computation completely and then adjust
  693. 25:31thousands of control settings and then
  694. 25:33start again. And that's a major problem
  695. 25:35if eventually we want to run these
  696. 25:37quantum computers continuously for days
  697. 25:40or months. So here's Google's solution
  698. 25:42to this. They use reinforcement learning
  699. 25:44together with quantum error correction.
  700. 25:46Now this is very technical, but I'll try
  701. 25:48to dumb it down. What quantum error
  702. 25:50correction does is it produces a stream
  703. 25:52of signals showing that errors are
  704. 25:54happening. These are like warning signs.
  705. 25:56Well, Google basically feeds these
  706. 25:58signals into an AI agent. The agent
  707. 26:00watches how these error patterns change
  708. 26:03and then adjust thousands of control
  709. 26:05parameters including the frequencies,
  710. 26:07phases, and strengths of the signals
  711. 26:09that control the qubits. You can think
  712. 26:11of it like tuning an orchestra while the
  713. 26:13musicians continue playing. And then
  714. 26:15after running this for a lot of rounds
  715. 26:17using reinforcement learning, the AI
  716. 26:19kind of learns which parameters to tweak
  717. 26:22in order to reduce the errors the most.
  718. 26:24In fact, they say that this
  719. 26:26reinforcement learning fine-tuning
  720. 26:27suppressed the logical error rate by an
  721. 26:29additional 20% which the team describes
  722. 26:33as a record low. Now the team also
  723. 26:36tested if this could scale up and what
  724. 26:38they found was that the AI's learning
  725. 26:40speed did not slow down as the system
  726. 26:42became larger which suggests that this
  727. 26:44approach could actually scale. So in a
  728. 26:46nutshell, this is basically an AI system
  729. 26:48that can learn and automatically correct
  730. 26:51errors from a quantum computer as it's
  731. 26:54running. So, you don't have to stop and
  732. 26:55reset everything and then start again. A
  733. 26:57pretty interesting breakthrough if
  734. 26:59you're into quantum computing. If you're
  735. 27:01interested in reading further, I'll link
  736. 27:03to this main page in the description
  737. 27:05below. Also this week, if you're looking
  738. 27:07for a relatively small model that can
  739. 27:09actually fit on consumer hardware, then
  740. 27:12this one might be a good option for you.
  741. 27:14It's called billion
  742. 27:18parameter agentic model. As with most of
  743. 27:20the top models out there right now, this
  744. 27:22is focused on agentic capabilities. So,
  745. 27:25tasks that include multiple steps and
  746. 27:27long horizon planning and reasoning.
  747. 27:29Now, if you compare this to similar
  748. 27:31sized models including Google's Gemma 4
  749. 27:33as well as Gwen 3.5 4B and even 9B,
  750. 27:37which is like three times larger, you
  751. 27:39can see that across these agentic and
  752. 27:42software engineering benchmarks such as
  753. 27:44MCP Atlas or Claw Eval, SweepBench
  754. 27:47Verified, SweepBench Pro, TerminalBench,
  755. 27:49and also its reasoning and world
  756. 27:51knowledge such as Humanities Last Exam,
  757. 27:53GPQA Diamond, etc., you can see that
  758. 27:56across the board Nanobeg 4.2 is even
  759. 27:59better. This is extremely impressive
  760. 28:01considering that, you know, Gemma 4 is
  761. 28:03like four times larger and Gwen 3.5 9B
  762. 28:06is three times larger. So, in terms of
  763. 28:08performance versus model size, this is
  764. 28:11by far the strongest performer. And you
  765. 28:13know, the most interesting part here is
  766. 28:15its looped transformer architecture. You
  767. 28:17see, a normal transformer moves through
  768. 28:19each layer once, but what Nanobeg does
  769. 28:21is it reuses the same layers multiple
  770. 28:24times as if it's looping the data
  771. 28:26through the network again. And
  772. 28:28interestingly, this allows it to perform
  773. 28:30more computations without needing to
  774. 28:31store a separate set of parameters
  775. 28:33somewhere else. Apparently, this looping
  776. 28:35step allows it to think for longer,
  777. 28:38which gives higher quality outputs. Now,
  778. 28:41if you click on files and versions, you
  779. 28:42can see that this is really tiny. So,
  780. 28:45So, full 3 billion parameter model is
  781. 28:46only 8 GB in size, which should fit on
  782. 28:49like most consumer GPUs. And if not, I'm
  783. 28:51sure there's also going to be more
  784. 28:52quantized versions of this that can fit
  785. 28:54on even lower VRAM. Anyway, a super
  786. 28:56performant model. This was only released
  787. 28:58like 2 days ago, but it has already
  788. 29:00gotten like over 4,000 downloads. If
  789. 29:02you're interested in trying this out,
  790. 29:04I'll link to this main page in the
  791. 29:06description below. Also this week,
  792. 29:08Alibaba releases their latest and best
  793. 29:10image model, Qwen Image 3. Now, the
  794. 29:13previous Qwen image models were very
  795. 29:15performant, but this is even better. You
  796. 29:18can jam-pack a ton of details and text
  797. 29:21and elements into your prompt, and it's
  798. 29:22able to generate all of this very well.
  799. 29:25It's even able to generate like math
  800. 29:26equations, different icons, and of
  801. 29:28course different text, as you can see
  802. 29:30here. It's also great at designing user
  803. 29:32interfaces. This entire image was
  804. 29:35actually made by Qwen 3. This is not a
  805. 29:37snapshot from VS Code. You can see there
  806. 29:39are some subtle errors such as the icons
  807. 29:42like here and here and over here. So, if
  808. 29:44you look closely, you can see that this
  809. 29:46entire image is indeed AI generated.
  810. 29:48Same with the icons over here. Here's
  811. 29:50another example of a really complex
  812. 29:52image with a ton of text and elements,
  813. 29:55but as you can see, for the most part,
  814. 29:57Qwen Image 3 was able to generate this
  815. 30:00flawlessly. Here's another crazy
  816. 30:01example. This is just an image that was
  817. 30:04generated by Qwen. This isn't actually,
  818. 30:07you know, a PDF or anything like that.
  819. 30:09But as you can see, it's able to even
  820. 30:10generate all this text as if it was a
  821. 30:12screenshot from a scientific paper, plus
  822. 30:15with all these complex math equations. A
  823. 30:17very impressive result. And of course,
  824. 30:19this can also create some very
  825. 30:21photorealistic images, as you can see
  826. 30:23from these examples. It's able to
  827. 30:25generate a variety of different textures
  828. 30:27and art styles. It has very good
  829. 30:29understanding of different things. Now,
  830. 30:31like the top frontier models out there,
  831. 30:32instead of just generating images, you
  832. 30:34can also edit existing images. And it
  833. 30:36has really good world understanding. For
  834. 30:38example, you can upload this image of
  835. 30:40two damselflies, and you can get it to
  836. 30:43basically create an infographic poster
  837. 30:45with labels and information about the
  838. 30:47species. And here's the image that it
  839. 30:49outputs. And of course, like the other
  840. 30:51frontier models, you can also upload
  841. 30:53some old and damaged photos and get it
  842. 30:54to restore these photos like this. Or
  843. 30:57here's another cool example where we can
  844. 30:59upload this photo and get it to annotate
  845. 31:01stuff like this. One of the main selling
  846. 31:03points of this is that it has precise
  847. 31:05text rendering as small as 10 pixels.
  848. 31:08It's also really good at reproducing
  849. 31:09some micro details like pores and hair
  850. 31:12strands. And this also supports 12
  851. 31:14languages. Now, currently this is
  852. 31:16closed. You can try it using their
  853. 31:18online coin studio, but it's not clear
  854. 31:21here whether they're actually using the
  855. 31:23latest coin image three or a previous
  856. 31:25version. Or you can also use it via
  857. 31:26their API. Now, here's the thing. As a
  858. 31:28closed model, this is not as good as the
  859. 31:31best ones out there. This is very
  860. 31:32similar to GPT image or C-Dream and Nano
  861. 31:36Banana. So, it's really hard to pick a
  862. 31:38winner here. From my initial test, it
  863. 31:40does look like GPT image is still the
  864. 31:42leader followed by C-Dream. So, at least
  865. 31:44for me, I don't have any real reason to
  866. 31:46use this, but if you are interested in
  867. 31:48trying this out, I'll link to this main
  868. 31:50page in the description below. Also this
  869. 31:52week, Anthropic drops their latest and
  870. 31:55best model, Claude Opus 5. And this is
  871. 31:57indeed one of the strongest models you
  872. 31:59can use right now. They claim it has
  873. 32:01much stronger reasoning and planning,
  874. 32:04especially for complex multi-step tasks
  875. 32:06and agentic workflows compared to even
  876. 32:08their previous best model, Claude Fable
  877. 32:105. So, if you look at these benchmarks
  878. 32:12in terms of agentic coding, then Opus 5
  879. 32:15is better than Claude Fable 5. In terms
  880. 32:18of legal and health, it's not as strong
  881. 32:19as Fable. And in terms of biology, as
  882. 32:21long as it doesn't outright reject your
  883. 32:23question, which it does most of the
  884. 32:25time, then it is better than Fable 5.
  885. 32:28Now, interestingly, the metric that
  886. 32:29really matters to me is Deep Sweep. This
  887. 32:31is, you can say, a more accurate measure
  888. 32:33of how good it is at agentic software
  889. 32:35engineering tasks. And as you can see
  890. 32:38here, actually Open AI's GPT-5.6 Soul is
  891. 32:42still number one. Look how misleading
  892. 32:44this is. They didn't highlight this
  893. 32:45orange, but instead gray, as if it was
  894. 32:48less important. Now, if you look at
  895. 32:49performance versus price, in terms of
  896. 32:51this frontier bench, which measures
  897. 32:53agentic coding, you can see that Opus 5
  898. 32:56can score higher than GPT-5.6 Soul,
  899. 32:59which is again grayed out to look like
  900. 33:01it's irrelevant, but it does cost more
  901. 33:03to significantly beat it. And here's
  902. 33:05another misleading thing. Note that this
  903. 33:07x-axis of price is log scale. So,
  904. 33:10they're kind of like compressing the
  905. 33:12prices on the right side. So, these
  906. 33:14prices are actually way more expensive
  907. 33:16than it looks. Now, if you look at this
  908. 33:18agentic coding index by Artificial
  909. 33:21Analysis, then as you can see, it's not
  910. 33:23too impressive. Opus 5, even at its
  911. 33:25maximum performance, can only match the
  912. 33:27performance of 5.6 Soul, which is the
  913. 33:30gray dot over here. But, it does so at a
  914. 33:32much more expensive price. Now, probably
  915. 33:34the most impressive metric of Opus 5 is
  916. 33:37its score on Arc AGI-3. This is an
  917. 33:39interactive benchmark that tests AI
  918. 33:42agents by dropping them into brand new
  919. 33:44abstract games with no instructions.
  920. 33:46Agents must explore, figure out the
  921. 33:48goals, and learn the rules on their own,
  922. 33:50just like humans do. The thing is, even
  923. 33:53the best of the best models, like Gemini
  924. 33:55and Opus 4.8, score less than 1%. And
  925. 33:57even GPT-5.6 Soul scores like around 7%.
  926. 34:02They do pretty horribly on this
  927. 34:03benchmark, and that's because,
  928. 34:04technically speaking, AI models don't
  929. 34:07really learn new things or patterns
  930. 34:09after they're finished training. Their
  931. 34:10model weights, or basically the values
  932. 34:12in their neural networks, are fixed. So,
  933. 34:15this benchmark actually tests an AI's
  934. 34:17emergent abilities to learn new patterns
  935. 34:20it has never seen before. And
  936. 34:21apparently, Opus 5 performs much better
  937. 34:24than the other models. It scores like
  938. 34:26over 30%. However, note that this could
  939. 34:30be benchmarked. If you give it some
  940. 34:31classic Witness style games that are
  941. 34:33present in Arc-AGI, then Opus 5 does
  942. 34:36very well. But on the most novel games
  943. 34:39with unusual mechanics, then Opus 5
  944. 34:42actually performs worse than Opus 4.8.
  945. 34:44And these are games where the rules must
  946. 34:46be discovered through interaction. The
  947. 34:48model needs to figure out these new
  948. 34:50rules on its own. Now, if you look at
  949. 34:52this independent leaderboard by
  950. 34:53Artificial Analysis, then you can see
  951. 34:55that Opus 5 is indeed ranked number one,
  952. 34:58just one point above Claude Fable, which
  953. 35:00is one point above GPT 5.6 Soul.
  954. 35:04However, if you look at the price of
  955. 35:05this here is where it becomes not too
  956. 35:07attractive. So, this costs almost double
  957. 35:10the cost of GPT 5.6 Soul, which you
  958. 35:12could argue is just as intelligent. So,
  959. 35:15Opus might not be the most
  960. 35:16cost-efficient option. If you look at
  961. 35:18Simple Bench, which tests an AI model on
  962. 35:20some tricky but common sense questions,
  963. 35:22Opus 5 does pretty well. It is ranked
  964. 35:25second, whereas Claude Fable is ranked
  965. 35:27first. And surprisingly, Gemini 3. Pro
  966. 35:30is still in third place, even though it
  967. 35:32was launched like half a year ago. If
  968. 35:34you look at Live Bench by Abacus AI,
  969. 35:36then GPT 5.6 Soul is still ranked number
  970. 35:39one. Surprisingly, Opus 5 is only ranked
  971. 35:42number three, even behind Fable 5. It
  972. 35:44seems to be fairly weak across the board
  973. 35:47in terms of reasoning, coding,
  974. 35:48mathematics, data analysis, language,
  975. 35:50and instruction following. Pretty
  976. 35:52interesting results. And if you look at
  977. 35:54this Valse Index, which tests a model
  978. 35:56across finance and coding tasks, then
  979. 35:58you can see that Opus 5 is also not
  980. 36:00ranked number one. It's only second
  981. 36:02place and just a few decimal points
  982. 36:04ahead of the open-source Kimmy K3. Now,
  983. 36:07the main drawback of using any Claude
  984. 36:09model is that sometimes it just outright
  985. 36:11rejects your request. There are a ton of
  986. 36:13guardrails in place, especially for the
  987. 36:16frontier Claude models. If you've been
  988. 36:18using Fable 5, then you'll know that it
  989. 36:20can fall back to less capable models,
  990. 36:22especially if you're asking about
  991. 36:24cybersecurity or biology-related
  992. 36:26questions. Well, the same goes for Opus
  993. 36:285, but it seems like the restrictions
  994. 36:30are a bit lighter. So, here it says Opus
  995. 36:325 is less restrictive than Fable 5 in
  996. 36:35terms of answering cybersecurity
  997. 36:37questions. It does allow Opus 5 to find
  998. 36:40vulnerabilities in source code, but it
  999. 36:42blocks all these other requests. And
  1000. 36:45like before, if it flags a request, then
  1001. 36:47it will fall back to the dumber Opus 4.8
  1002. 36:49by default. And then for biology, it
  1003. 36:52also says that the guardrails are less
  1004. 36:55restrictive for Opus 5. Now, currently
  1005. 36:57Claude Opus 5 is available on all paid
  1006. 37:00plans plus via API. Now, they released
  1007. 37:03this yesterday and today is reserved for
  1008. 37:05my weekly news video. So, I'll probably
  1009. 37:07make a full review video on Claude Fable
  1010. 37:095 tomorrow or the day after that. Stay
  1011. 37:12tuned because I'm going to showcase a
  1012. 37:13ton of incredible things. Also this
  1013. 37:15week, we have a new video model by
  1014. 37:17Nvidia called Sora video 2. And this is
  1015. 37:20a fairly efficient model at either 5
  1016. 37:23billion or 14 billion per parameters.
  1017. 37:25And this is designed to generate up to
  1018. 37:27720p videos on just a single GPU. So,
  1019. 37:30here are some examples for your
  1020. 37:32reference. The quality isn't bad, but
  1021. 37:34there is some noise and distortions,
  1022. 37:36especially along the edges of certain
  1023. 37:38things. I would say the quality isn't as
  1024. 37:40good as LTX 2.3 or 1. But here, I guess
  1025. 37:44they're optimizing for efficiency
  1026. 37:45instead of quality. And then here are
  1027. 37:47some higher action shots. And again, you
  1028. 37:50can clearly see some artifacts and
  1029. 37:51noise. Some dogs disappear and reappear.
  1030. 37:54So, it's not perfect for slower scenes
  1031. 37:56like walking, then it's pretty good. And
  1032. 37:58you can also generate first-person
  1033. 38:00robotics videos to use for training
  1034. 38:02data, as you can see here. So, not the
  1035. 38:04best quality video model, but it is
  1036. 38:06fairly efficient. If you click on this
  1037. 38:08code button, it doesn't look like Sora
  1038. 38:09video 2 is out yet, but they have
  1039. 38:11released the other Sora models before.
  1040. 38:13So, it's very likely that they will
  1041. 38:15release Sora video 2 in the near future.
  1042. 38:17If you're interested in reading further,
  1043. 38:19I'll link to this main page in the
  1044. 38:21description below. Also this week we
  1045. 38:23have an open source world generator for
  1046. 38:26Minecraft called Open Dreamer. And this
  1047. 38:28is basically an open source version of
  1048. 38:30Google's Dreamer 4, which is closed
  1049. 38:33source. Now this world model is
  1050. 38:35essentially just a video generator,
  1051. 38:37nothing is pre-programmed, but they
  1052. 38:38trained this Open Dreamer model to
  1053. 38:41generate scenes of Minecraft which can
  1054. 38:43react to actions from an AI agent. So
  1055. 38:45even though everything is just video,
  1056. 38:47it's kind of generating how an AI agent
  1057. 38:50would interact with and move through
  1058. 38:52this Minecraft environment over time.
  1059. 38:54The cool thing is you can try out this
  1060. 38:55demo online, so you can like press the
  1061. 38:58WASD keys to move this character around.
  1062. 39:01Now unlike Google's Dreamer 4, which is
  1063. 39:03closed source, here they're actually
  1064. 39:05releasing all the details including the
  1065. 39:07model, the training code, and a detailed
  1066. 39:09breakdown of what worked and what
  1067. 39:11failed. And interestingly, rather than
  1068. 39:13starting from Minecraft immediately,
  1069. 39:15they built a smaller version using Coin
  1070. 39:17Run, which is a simple 2D game where
  1071. 39:20moves through an environment to collect
  1072. 39:22coins. This allowed them to test the
  1073. 39:24entire system on just one GPU and find
  1074. 39:26problems quickly and confirm that their
  1075. 39:28architecture worked before scaling it up
  1076. 39:30to 3D Minecraft level. At the top of the
  1077. 39:33page they've released a GitHub repo and
  1078. 39:35here it contains all the instructions on
  1079. 39:37how to download and run this locally on
  1080. 39:39your computer. Plus they've also
  1081. 39:41released the data set and the training
  1082. 39:42code to this as well. If you're
  1083. 39:44interested in reading further, I'll link
  1084. 39:46to this main page in the description
  1085. 39:48below. Also this week Google releases
  1086. 39:50not one but three new AI models, Gemini
  1087. 39:533.6 Flash, 3.5 Flash Light, and 3.5
  1088. 39:58Flash Cyber. So let's go through each of
  1089. 40:00these in more detail. The main model is
  1090. 40:02Gemini 3.6 Flash, which is designed to
  1091. 40:05be a really fast general purpose model.
  1092. 40:07It can handle coding, document analysis,
  1093. 40:10knowledge work, visual understanding,
  1094. 40:12computer control, and other multi-modal
  1095. 40:14tasks. The key focus here is that it
  1096. 40:16completes these jobs using fewer tokens
  1097. 40:19and fewer unnecessary steps. And
  1098. 40:21according to Google, Gemini 3.6 flash
  1099. 40:24uses almost 60% fewer tokens compared to
  1100. 40:27Gemini 3.5 flash. This matters because a
  1101. 40:30model that reaches the same answer with
  1102. 40:32fewer steps and fewer tokens can be both
  1103. 40:34faster and cheaper. And the benchmark
  1104. 40:36improvements, at least according to
  1105. 40:38Google, are also fairly substantial. In
  1106. 40:40terms of Deep Sweep and machine learning
  1107. 40:42engineering benchmarks, as well as
  1108. 40:43knowledge work and computer use, 3.6
  1109. 40:46flash does significantly outperform 3.5
  1110. 40:49flash. Now, interestingly, if you look
  1111. 40:51at this independent leaderboard by
  1112. 40:52artificial analysis, then you can see
  1113. 40:54that Gemini 3.6 flash does not actually
  1114. 40:57perform better than 3.5. It's actually
  1115. 41:00just tied, but it does cost a bit
  1116. 41:02cheaper. But note that this is still way
  1117. 41:04more expensive than some more
  1118. 41:06intelligent competitors like GPT 5.6
  1119. 41:09Luna and JLM 5.2. So, it's not even the
  1120. 41:12most cost-efficient option out there.
  1121. 41:14Now, Google is also launching Gemini 3.5
  1122. 41:17flash light, which, if you're confused,
  1123. 41:20light is just basically a faster version
  1124. 41:22of flash, which is a faster version of
  1125. 41:24Pro. Again, this is designed to be a
  1126. 41:26much smaller and faster option for
  1127. 41:28high-volume tasks like search agents,
  1128. 41:31document processing, data extraction,
  1129. 41:33and generating many possible solutions
  1130. 41:35in parallel. And if you look at the
  1131. 41:37speed of this, flash light is able to
  1132. 41:39complete a task way faster, like roughly
  1133. 41:41four times as fast as the non-light
  1134. 41:44version. Now, here's the thing. Again,
  1135. 41:45if you look at this independent
  1136. 41:47leaderboard by artificial analysis, then
  1137. 41:49Gemini 3.5 flash light is still more
  1138. 41:51expensive than the more intelligent and
  1139. 41:54open-source model Deep Seek V4 flash,
  1140. 41:56which is like less than half the cost
  1141. 41:58per task. Same with GPT 5 Luna medium,
  1142. 42:01which also costs much less. However, if
  1143. 42:03you look at the speed, then this is like
  1144. 42:06way faster than all these other
  1145. 42:07competitors. Now, the third release is
  1146. 42:10Gemini 3.5 flash cyber. This is a
  1147. 42:13specialized version trained to find,
  1148. 42:15validate, and repair software
  1149. 42:17vulnerabilities. It operates inside
  1150. 42:19Google's Code Mentor platform, where
  1151. 42:21multiple cybersecurity agents
  1152. 42:23investigate a problem and combine their
  1153. 42:25work into one report. Google says it
  1154. 42:27reaches competitive frontier performance
  1155. 42:29on this benchmark called Cyber Gym,
  1156. 42:31apparently edging very close to the
  1157. 42:34frontier models like GPT 5.6 Soul and
  1158. 42:36even Mythos 5. Now, for this 3.5 Flash
  1159. 42:40Cyber, this model will be exclusively
  1160. 42:42available to governments and trusted
  1161. 42:44partners as part of a limited access
  1162. 42:46pilot program. But, for the previous two
  1163. 42:48models, 3.6 Flash and 3.5 Flash Lights,
  1164. 42:51they are already available via the API
  1165. 42:54or Google's free AI Studio platform. And
  1166. 42:573.6 Flash is also available in Google's
  1167. 43:00agentic coding platform called
  1168. 43:01Antigravity. Now, this release does seem
  1169. 43:04quite lackluster. Google did not release
  1170. 43:06any state-of-the-art model that can be,
  1171. 43:08you know, like Opus or GPT 5.6. It seems
  1172. 43:12like they're focusing more on speed and
  1173. 43:14efficiency, which, to be fair, is a
  1174. 43:16decent strategy. I mean, Google has so
  1175. 43:18many different products and platforms
  1176. 43:20like Gmail, Workspace, Google Search,
  1177. 43:22Google Analytics, Maps, YouTube,
  1178. 43:24Finance, etc. And they want to integrate
  1179. 43:27AI into all these platforms. So, they
  1180. 43:29need to create an AI model that's
  1181. 43:31incredibly fast and lightweight for
  1182. 43:33seamless integration. Anyway, if you're
  1183. 43:35interested in reading further, I'll link
  1184. 43:37to this main page in the description
  1185. 43:39below. And that sums up all the
  1186. 43:41highlights in AI this week. Let me know
  1187. 43:44in the comments what you think of all of
  1188. 43:46this. Which piece of news was your
  1189. 43:47favorite and which tool are you most
  1190. 43:50looking forward to trying out? As
  1191. 43:52always, I will be on the lookout for the
  1192. 43:54top AI news and tools to share with you.
  1193. 43:57So, if you enjoyed this video, remember
  1194. 43:59to like, share, subscribe, and stay
  1195. 44:01tuned for more content. Also, there's
  1196. 44:04just so much happening in the world of
  1197. 44:05AI every week, I can't possibly cover
  1198. 44:08everything on my YouTube channel. So, to
  1199. 44:10really stay up-to-date with all that's
  1200. 44:13going on in AI, be sure to subscribe to
  1201. 44:15my free weekly newsletter. The link to
  1202. 44:18that will be in the description below.
  1203. 44:20Thanks for watching and I'll see you in
  1204. 44:22the next one.

About this transcript

This page contains the full transcript of Claude Opus 5, GPT 6 hack, Flux 3, new Gemini, quantum breakthrough, new Qwen: AI NEWS by AI Search, generated from the public captions YouTube serves with the video. The transcript has 8,052 words across 1,204 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.