YouTube2Text

Anthropic lanza Opus 5.5: lo probé contra GPT-6 Astra — Transcript

by Benjamín Cordero · 4,815 words · 690 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Anthropic just launched Opus 5. GPT6
  2. 0:05Astra. In their table, we can see it
  3. 0:08beats it in several programming tests
  4. 0:10and professional work. And the cost
  5. 0:12shown here, I think, is a huge part of
  6. 0:15this entire new announcement. This is
  7. 0:19an excellent model, but Astra is still
  8. 0:21ahead in two benchmarks that still
  9. 0:23matter to us, which are Business
  10. 0:24Workflow, right here, and Agentic
  11. 0:26Scientific Research. But now we can see
  12. 0:30that Opus is crushing it in everything
  13. 0:32related to coding, and it's doing very
  14. 0:35well in many other tests as well, even
  15. 0:38omitting comparisons with Astra in some
  16. 0:40cases. So, the question is, should I
  17. 0:43switch? Has Anthropic really returned
  18. 0:45to being the king? We are going to open
  19. 0:49the announcement, break it down from
  20. 0:51top to bottom, and analyze all the
  21. 0:52prices, the benchmarks, what each one
  22. 0:54means, and my personal take—whether I
  23. 0:56ended up liking it or not. We are going
  24. 0:59to create a page with images from GPT
  25. 1:01Image 2.5. We will connect it to
  26. 1:04Highfield and Sidans 2.5, and we will
  27. 1:06ask for different tasks to see if it
  28. 1:09truly has the same level of performance
  29. 1:11between GPT6 Astra, for example, and
  30. 1:13Opus 5. Before we start, I would really
  31. 1:17appreciate it if you could leave a like
  32. 1:19on this video, not only because it
  33. 1:21helps me, but because you also tell
  34. 1:23your algorithm that you like this style
  35. 1:24of content and it starts recommending
  36. 1:26more and more of it. Now, let's get
  37. 1:29down to business. Okay, before we begin
  38. 1:32, let's set two prompts running so that
  39. 1:34at the end of this video we can see the
  40. 1:36result and compare GPT6 Astra with Opus
  41. 1:385.5. I’m going to throw the same
  42. 1:42prompt at both of them, which is to
  43. 1:43build a page—a landing page—
  44. 1:44connected to Highfield using Sidan
  45. 1:46Caches 2.5. And well, we'll see what
  46. 1:49happens, and then we'll analyze which
  47. 1:51one was actually the best. So, here we
  48. 1:53can already see Opus 5. Max. And now
  49. 1:57I’m going to go into Codex and ask
  50. 1:59for the exact same thing, both at their
  51. 2:02highest reasoning levels. Okay, now
  52. 2:06let's see that, uh, a few seconds or
  53. 2:09minutes ago, Anthropic made this launch
  54. 2:12, which was their new model, Opus 5. It
  55. 2:18says it performs at the level of Claude
  56. 2:21Fable 5.1. One, but this is the
  57. 2:25interesting part: it costs 40%less to
  58. 2:28run than Opus 5, which is truly a
  59. 2:30stroke of genius considering new models
  60. 2:33like, uh, JEV, which are running very,
  61. 2:36very cheaply. And now we are starting
  62. 2:40to compete; while they aren't the same,
  63. 2:42we are definitely starting to compete
  64. 2:44quite a bit on price, because
  65. 2:45alternatives that are much, much
  66. 2:46cheaper are starting to appear. But why
  67. 2:50don't we see what this really means?
  68. 2:52And what is, what is the whole point of
  69. 2:55this launch, because I started reading
  70. 2:57and it looked quite interesting. Here
  71. 2:59we have the main benchmarks comparing
  72. 3:02them to Opus 5. With the previous
  73. 3:06version, Opus 5, Claude 5.1, and GPT-4o
  74. 3:10. And well, what I also want to
  75. 3:13highlight is that notice here they say
  76. 3:16it is the first model in their new 3.5
  77. 3:18family. In other words, Sonnet and
  78. 3:22Haiku could probably be coming in the
  79. 3:23next few weeks, or perhaps they'll
  80. 3:24discontinue Haiku. I don't know, let's
  81. 3:27see what happens, but there will
  82. 3:29undoubtedly be a new Sonnet or maybe
  83. 3:31even a Claude. Uh, from what I read, it
  84. 3:35functions quite similarly in reasoning
  85. 3:38levels to Claude 3.5, and this is their
  86. 3:41official launch blog. Uh, what I want
  87. 3:44to see here, well, we already know what
  88. 3:47we always do, let's translate it to
  89. 3:49English so there are no differences or
  90. 3:52problems with language barriers. And
  91. 3:55let's see the funny translations that
  92. 3:57Google Translate gives us here. Uh, and
  93. 4:00notice they tell you that in most tasks
  94. 4:03, its operation costs 40%less than Opus
  95. 4:063. And notice this comparison they are
  96. 4:09showing us here, uh, because further
  97. 4:11down we will see how much it costs
  98. 4:13according to Astra and GPT-4o. Uh, they
  99. 4:17announce a different generation of
  100. 4:19responses. They tell us that responses
  101. 4:22are 30%faster than Opus 3, uh, and they
  102. 4:24tell you that in addition to the price
  103. 4:27drop, they increased the usage limit to
  104. 4:295 hours on Pro, Team, and Enterprise
  105. 4:31plans. So, they raised the limits to
  106. 4:34these 5 hours. Great. Uh, and notice
  107. 4:37that they also told us here that they
  108. 4:40gave us a, uh, token reset, where I
  109. 4:42found this here in their post. Uh,
  110. 4:47besides increasing the 5-hour usage
  111. 4:49limit, they also gave people, or users,
  112. 4:52or subscriptions, a rate limit reset
  113. 4:54that can be used whenever you want.
  114. 4:59This is something Codex already had,
  115. 5:01which I found great because, basically,
  116. 5:03you have periods of higher intensity or
  117. 5:05work where you need these resets, so I
  118. 5:07think it was an excellent, excellent
  119. 5:09success. We never really know how much
  120. 5:13a subscription yields for us, so they
  121. 5:14can tell us yes, eh, it has a higher
  122. 5:16limit, but in the end we can never
  123. 5:18really verify it because this is
  124. 5:19something they adjust based on
  125. 5:21computing demand and the supply they
  126. 5:23have. If we start looking here, well,
  127. 5:26it communicates more naturally, blah
  128. 5:28blah blah blah. Right, yes, obviously,
  129. 5:30I’m always going to want to say what
  130. 5:32suits me best here. So let's go down to
  131. 5:36this, let’s get into this, which is
  132. 5:39what really interests us, which would
  133. 5:41be the different benchmarks between
  134. 5:43Fable 5.1 and the rest of the
  135. 5:45benchmarks. Eh, we’re going to do
  136. 5:49what we always do, take a screenshot so
  137. 5:51we can write on or show things. And
  138. 5:54let's see what each one means. Here we
  139. 5:57have a series of benchmarks that are
  140. 6:00standardizations carried out by
  141. 6:02different people or tests used to
  142. 6:04measure all these models in an attempt
  143. 6:07to have some kind of criteria for doing
  144. 6:10so. And Anthropic itself, in fact,
  145. 6:15tells us here that these benchmarks are
  146. 6:17even becoming less and less
  147. 6:19representative because something also
  148. 6:21happens here where they isolate the
  149. 6:23models from their harnesses. So, it's
  150. 6:27not exactly the flow one has when
  151. 6:29working in something like Claude Code
  152. 6:32or Codex, eh, because the model is
  153. 6:34being isolated a bit, but anyway, it is
  154. 6:36still an excellent, excellent parameter
  155. 6:39. The first one that caught my
  156. 6:43attention is this one here, agent
  157. 6:45coding or agentic coding, which is the
  158. 6:48Terminal Bench 4.0, which tells us in
  159. 6:51the end how well an agent can complete
  160. 6:53tasks from a terminal. I'm sorry, I
  161. 6:57have to take a screenshot because I
  162. 6:59really need to see this in English,
  163. 7:01otherwise it’s very hard for me, eh,
  164. 7:05to be able to read the numbers and
  165. 7:07things well here. So now, yes. Eh,
  166. 7:10agentic Terminal Bench 4.0. We have
  167. 7:13Opus here which scored 66%versus its
  168. 7:15previous version which was 52%. So, it
  169. 7:20won here and it also beat Astra by
  170. 7:22quite a bit, which is super good, eh,
  171. 7:24because Astra was already super good at
  172. 7:27this. I personally used it quite a bit,
  173. 7:31but here Opus 5. So, let's think of
  174. 7:39this benchmark as the ability to solve
  175. 7:41a task through the terminal, that is,
  176. 7:43where the computer’s actions are
  177. 7:46executed. It does it better than all
  178. 7:49the other models and would be the new
  179. 7:50state of the art. Next, we have another
  180. 7:54benchmark here that is also about
  181. 7:56coding, but I believe in the one above,
  182. 7:58Terminal Bench 4.0, they only ask a
  183. 8:00single coding-related question that is
  184. 8:02also solved. This one involves a series
  185. 8:06of more steps; I mean, it's like a
  186. 8:08problem, a project, it involves more
  187. 8:10steps unlike Terminal Bench 4. Frontier
  188. 8:13Code goes beyond just answering a
  189. 8:16single question, and here it gets 54%
  190. 8:18versus eh Fable which gets 53%, meaning
  191. 8:21it beats it and shows a substantial
  192. 8:24improvement over Fable 5.1. I can't
  193. 8:27wait for Fable 5.5 to come out because
  194. 8:29it's going to be very interesting, as
  195. 8:31it will surely show us significant
  196. 8:33improvements as well, without a doubt.
  197. 8:36Eh, next we have another one here for
  198. 8:39Agentic Coding again, and this one is
  199. 8:41Cursor Bench 4.0, which also performs
  200. 8:43better than Fable 5.1. But it intrigues
  201. 8:47me because they aren't measuring Astra
  202. 8:49here, but in the end, this is a task,
  203. 8:51it's also Agentic Coding, but it's done
  204. 8:53in the Cursor environment; that's why
  205. 8:56it's called Cursor Bench 4.0. Here it
  206. 9:00performed with almost 58%. And let's
  207. 9:04remember all this is based in the end
  208. 9:06on the sessions you have within the
  209. 9:08Cursor harness, which allows you to use
  210. 9:11different models, which is the
  211. 9:12interesting part. Yes, it intrigues me
  212. 9:16quite a bit why the Astra box isn't
  213. 9:18there. I'm going to remove this here.
  214. 9:20And this one too. Eh, it intrigues me
  215. 9:22quite a bit why it's not there, because
  216. 9:24obviously, it not being there doesn't
  217. 9:25mean zero. Next, we go to this other
  218. 9:29test which would be knowledge work, eh,
  219. 9:32which is quite the opposite of a
  220. 9:35percentage, sorry, but it is measured
  221. 9:39in points. And here we have 100 versus
  222. 9:42the others, eh, where Fable 5.1 was
  223. 9:45already at a super high number and GPT-
  224. 9:474o Astra wasn't as high. This is a
  225. 9:52benchmark that evaluates professional
  226. 9:54work from I don't remember the exact
  227. 9:57number of professions, but basically,
  228. 10:00it asks how good the outputs of this or
  229. 10:03that deliverable are in the end, eh?
  230. 10:08And then it assigns a score, and that
  231. 10:10score for Opus 5...Is 1846, and for the
  232. 10:12other models, it is lower. Here I was
  233. 10:17also a bit surprised that GPT-4o Astra
  234. 10:20reached 1542. Eh, again, this is a kind
  235. 10:25of score like the one used in chess or
  236. 10:27an Elo rating. Eh, I don't know, higher
  237. 10:31is obviously better. This is a
  238. 10:34benchmark that I am quite interested in
  239. 10:36, and I think it is one of the most
  240. 10:38important ones for people who build
  241. 10:40automations: the Business Workflow
  242. 10:42benchmark. Uh, this is a benchmark that
  243. 10:45Zapier put together, and it ultimately
  244. 10:47measures how well it connects with the
  245. 10:49different applications we use in our
  246. 10:51day-to-day business. I mean, it takes a
  247. 10:55task and then measures whether, uh,
  248. 10:57people can correctly complete that
  249. 11:00specific task. This is much closer to
  250. 11:03what is actually done and how we use it
  251. 11:04in our day-to-day lives. I mean,
  252. 11:05especially for people who run
  253. 11:07businesses. It's like looking for
  254. 11:09something in a CRM or, I don't know,
  255. 11:11updating a record in a database or
  256. 11:13contacting and following up with this
  257. 11:16person here. uh, coordinating actions
  258. 11:18between different tools with different
  259. 11:20connections, and so on. Yeah, I think
  260. 11:22this is an excellent benchmark. And
  261. 11:24here it gets 40%and it is beaten by GPT
  262. 11:27-6 Sol, uh, sorry, GPT-6 Astra with 41%
  263. 11:30. Side note: it wouldn't surprise me if
  264. 11:33GPT-6 Sol comes out while I'm uploading
  265. 11:35this video. I think it's very likely
  266. 11:37they'll do something like that, so we
  267. 11:39just have to wait. But, uh, here GPT-6
  268. 11:43Astra does win, but where Opus 5 really
  269. 11:46wins, uh, by quite a lot, rather. Is in
  270. 11:50the Humanities Last Exam. And this is
  271. 11:54an exam that I also like quite a bit
  272. 11:57because it ultimately gathers many
  273. 11:59questions or answers from different
  274. 12:02disciplines, sort of like a PhD, so to
  275. 12:05speak, and takes these specialists and
  276. 12:08then gives them a percentage, let's say
  277. 12:11, of, uh, evaluation in specialized
  278. 12:14knowledge. Uh, and here they measure it
  279. 12:17with tools, which is a bit more
  280. 12:19real-world. In fact, without tools, I
  281. 12:21don't really understand why it's done
  282. 12:23so much, since we are using tools in
  283. 12:25our day-to-day life anyway, but it's
  284. 12:27okay to isolate it; it's interesting.
  285. 12:30Uh, but yes, Opus 5 wins here. GPT-6,
  286. 12:35uh, Astra gets 67.7%versus 57.2%. So,
  287. 12:43here both results are with the tools.
  288. 12:46Obviously, uh, getting a sort of good
  289. 12:49grade in the end on an exam doesn't
  290. 12:52mean, uh, that it's the same as doing
  291. 12:55autonomous research in a certain way,
  292. 12:58but anyway, I think Opus 5 performs
  293. 13:01much better here. Than GPT-6 Astra. Uh,
  294. 13:07and the other category where GPT-6
  295. 13:09Astra does beat Opus 5. Is the
  296. 13:14Scientific Research benchmark, which is
  297. 13:16also a terminal-based science benchmark
  298. 13:17. And here, scientific research tasks
  299. 13:20are evaluated based on an agent working
  300. 13:22in a computer terminal. This Astra did
  301. 13:26come out on top with 64%versus Opus 5.
  302. 13:30Which was at 58, almost 59%. But anyway
  303. 13:35, the announcement does report standard
  304. 13:38deviation errors at several points at
  305. 13:41the end, so this could obviously still
  306. 13:44change, but it gives you a slight idea
  307. 13:47of what we're actually facing. This was
  308. 13:51the one I wanted to get to, which is
  309. 13:53OSWorld 2.0. Why? Because GPT-4o stood
  310. 13:57out quite a bit historically for having
  311. 13:59excellent computer use. Uh, and here
  312. 14:03Opus 5 is at 1.8%. Meaning, it beat or
  313. 14:08performed better than its previous
  314. 14:12models, but it isn't comparing it to
  315. 14:15GPT-4o Astra. We'll take a look at this
  316. 14:19test, but uh, it measures how well an
  317. 14:22AI navigates an interface designed for
  318. 14:25humans, which is where GPT-4o Astra was
  319. 14:28performing super well. If I go here and
  320. 14:33go to GPT-4o Astra and look for this,
  321. 14:37like their launch page, I'll be able to
  322. 14:42find here that they get a, uh, 72.6%.
  323. 14:47But look, they compare Opus here with
  324. 14:4970, Opus 5, obviously, with 70.2%. But
  325. 14:54if I come back here, it says Opus 5 has
  326. 14:5774%, so, like, a little bit more. But,
  327. 15:01anyway, here we are comparing, if we
  328. 15:04had to compare computer use, that is,
  329. 15:06how well an AI performs navigating
  330. 15:09interfaces designed for humans. Uh, GPT
  331. 15:13-4o Astra has 72.6%, Opus has 70, Opus
  332. 15:185 and the new model, uh, Opus 5. It is
  333. 15:23quite a bit higher than that, I mean,
  334. 15:26it is six or seven points higher than
  335. 15:29Opus 5. So, seven points higher than
  336. 15:33Opus 5 here would be, well, at 77%
  337. 15:37versus 72%. So it would still be
  338. 15:40winning. Obviously, we aren't using the
  339. 15:43same parameters here, the reference
  340. 15:45isn't exact, but it tells us they
  341. 15:47focused quite a bit on computer use,
  342. 15:49which is what made many of us start
  343. 15:51using GPT-4o Astra. So it's good to
  344. 15:54know that Opus has taken the lead again
  345. 15:55, at least here in everything related
  346. 15:57to computer use. And then we have the
  347. 16:01recognition of visual graphics, which
  348. 16:04is this series of, like, cartography,
  349. 16:07which also isn't compared to GPT-4o
  350. 16:10Astra, but we can see this is basically
  351. 16:13an interpretation of how well the
  352. 16:16models understand plans, axes,
  353. 16:19distances, a bit more thought out
  354. 16:21especially since it started being
  355. 16:24implemented with Blender, Rhino,
  356. 16:27AutoCAD, and all. all these things. Uh,
  357. 16:30it’s like a mapping test, so to speak
  358. 16:32, an understanding of space and the
  359. 16:34physical plane where it performed at 89
  360. 16:37%and 84%, in other words, 89%versus
  361. 16:4288.4%and 83%before. Again, we don't
  362. 16:45have the reference for GPT6, but well,
  363. 16:47let’s see how much it costs us to get
  364. 16:50it. Here they tell us, "The main
  365. 16:52advantage of Opus lies in its
  366. 16:54efficiency; it costs less per token
  367. 16:56than Opus 5 and uses fewer tokens per
  368. 16:58task, which is a 40%reduction in costs.
  369. 17:02Uh, well, a token is this small unit of
  370. 17:05measurement used to standardize the
  371. 17:08inputs and outputs that an AI has.
  372. 17:12It’s basically how this text is
  373. 17:14separated and it would be like gasoline
  374. 17:16; spending more tokens will only
  375. 17:18increase your bill. Here is Opus 5. It
  376. 17:24costs $ 4 for input and $ 5, uh, versus
  377. 17:28the $ 5 of Opus 5, and $ 20 for output
  378. 17:32versus $ 25 for the output of Opus 5.
  379. 17:37So, here for input it went down and for
  380. 17:39output it went down $ 5. Now, here, of
  381. 17:44course, it makes you a bit curious
  382. 17:47where this 40%comes from, if this isn't
  383. 17:5040%, this is only a one-fifth decrease,
  384. 17:54right? When you combine lower rates
  385. 17:58with lower consumption, in the end, you
  386. 18:00should arrive at this kind of 40%.
  387. 18:04Because it’s not only cheaper, but
  388. 18:07they also improved the cache writes
  389. 18:10here. So, in the end, all of this
  390. 18:12compounded is what makes us reach that
  391. 18:14figure. Look, uh, they also tell you
  392. 18:17that fast mode is available in Cloud
  393. 18:20Code and the platform with a speed up
  394. 18:22to 2.5 times, uh, and it costs you
  395. 18:24double, basically, just like all the
  396. 18:27others. The interesting thing, though,
  397. 18:31is if we compare it later with GPT6
  398. 18:33Astra and we look here at what would be
  399. 18:36the price, uh, which is $ 10 per
  400. 18:38million input tokens and $ 50 per
  401. 18:40million output tokens, versus these
  402. 18:43which would be 4:20, it would be up to
  403. 18:4560%less or more economical than Astra
  404. 18:48for better performance. So, it clearly
  405. 18:52depends on your field. Obviously, if
  406. 18:56you are involved in biology, maybe, or
  407. 18:58scientific research, you will want to
  408. 19:00keep using Astra, but for most tasks,
  409. 19:02uh, Opus 5.5 should work better for you
  410. 19:04. Uh, next here we have coding, we can
  411. 19:09see that it focuses on migration and
  412. 19:11audits of large projects. For example,
  413. 19:15here 200,000 lines in 3 hours. Before,
  414. 19:18it took more than 20 hours and consumed
  415. 19:192.5 times more tokens in an internal
  416. 19:21run. They did the same thing and
  417. 19:24managed to do it, well, we already saw,
  418. 19:26in less than 3 hours, I mean, from 20
  419. 19:28hours to less than 3 hours. Anyway,
  420. 19:30this is a case documented by them, uh,
  421. 19:32anyway, but look at this. Here we have
  422. 19:35what we already saw. Look at the curves
  423. 19:37, how they show and compare them to us.
  424. 19:40Grey being GPT-4o Astra and Opus 3.5
  425. 19:42being, obviously, the orange one here.
  426. 19:46On these curves, upwards you have a
  427. 19:47higher score, towards the right you
  428. 19:49have more cost, towards the left you
  429. 19:50have less cost per attempt. Obviously,
  430. 19:55this is also represented in a
  431. 19:57logarithmic way, I mean, it's not that
  432. 20:00it's proportionally more money, at
  433. 20:02least on the cost axis, but in the end,
  434. 20:05it really helps us visually represent a
  435. 20:08difference in the changes, at least.
  436. 20:12But what they want to show us here is
  437. 20:15that Opus 3.5 equals GPT-4o Astra in
  438. 20:18the terminal bench, in average
  439. 20:20reasoning, when it's at its maximum, so
  440. 20:23to speak. And if we look at the cost,
  441. 20:28the best attempt by GPT-4o Astra to get
  442. 20:31a 58%spent $ 7, and the attempt for
  443. 20:35almost 58%spent $ 3. I mean, here we
  444. 20:39have a reduction of practically half. I
  445. 20:43mean, we can do terminal bench tasks,
  446. 20:45all the agentic coding stuff, at half
  447. 20:48the price, achieving the best result of
  448. 20:50GPT-4o Astra, which is super good, at
  449. 20:53least to keep in mind if you are
  450. 20:55building things with AI, which is what
  451. 20:57matters most to us, right? Then, well,
  452. 21:01all of this is coding; here, the
  453. 21:03Frontier code also looks much better,
  454. 21:05especially in terms of price. You know,
  455. 21:08look at this. Here we have 40 cents for
  456. 21:11Opus on low, and here we have uh 1.5.
  457. 21:18Dollars, I mean, three times less for
  458. 21:21GPT-4o Astra while getting a lower
  459. 21:23result, and obviously, if we get to
  460. 21:26maximum reasoning, we get to similar
  461. 21:29things in this sense. But my point is
  462. 21:32that you can see the trend a little
  463. 21:35closer here, much more economical. The
  464. 21:37best point at 54%and the other at 53%,
  465. 21:41but spending $ 4 versus 80 cents, I
  466. 21:44mean, five times more. That's why it's
  467. 21:48also worth looking at this curve,
  468. 21:50because it helps us understand the
  469. 21:52relationship between costs and uh very
  470. 21:54well. the scores or the outputs. The
  471. 22:01following sections, in the end, show us
  472. 22:03how they improved in terms of security,
  473. 22:06uh, in everything related to prompt
  474. 22:09injection, which, well, is a super
  475. 22:11important part if you don't want to end
  476. 22:14up using your McDonald's chatbot as an
  477. 22:16LLM, uh, and obviously if you want to
  478. 22:19keep things secure. Then, if we keep
  479. 22:23scrolling down, we can find this test
  480. 22:25here that we also saw, moving a bit
  481. 22:27away from code. We're going to start
  482. 22:30entering reports, analysis, and
  483. 22:31presentations. It has an overall
  484. 22:33performance or score, or the Elo, which
  485. 22:36would be this score we saw earlier,
  486. 22:39significantly higher than all other
  487. 22:41models. Uh, if we keep scrolling down
  488. 22:45here, they start showing us different
  489. 22:48cases, uh, different testimonials on
  490. 22:50how it changed versus its previous
  491. 22:53models. For example, here both models
  492. 22:57are comparing an error or explaining a
  493. 23:00problem with an error that occurred in
  494. 23:02billing. Uh, so here we can see that it
  495. 23:05starts with the consequence, helping
  496. 23:07you understand what the error is.
  497. 23:11Unlike here, it explains the error
  498. 23:13first and then it sort of starts
  499. 23:14explaining the consequence, from what I
  500. 23:16understood from what I managed to read
  501. 23:19here. If we keep scrolling down, we
  502. 23:22will find more things regarding
  503. 23:23security. Yes, we keep scrolling down
  504. 23:26here. Obviously, more security stuff
  505. 23:28regarding alignment and, well, all
  506. 23:30these things that you can start looking
  507. 23:32at if you're also interested in how it
  508. 23:34works and how all the APIs connect. Uh,
  509. 23:38here, well, they restricted or
  510. 23:40unrestricted, so to speak, the capacity
  511. 23:42or the blocks that had been placed on
  512. 23:45biology. Uh, it even outperforms
  513. 23:48Mixtral 5.1 in many areas of work. And
  514. 23:51over here we already have the
  515. 23:53availability. It is available on all
  516. 23:55platforms, so to speak. If you go into
  517. 23:59Claude, you will find this here. I
  518. 24:02loved that they give you a limit to
  519. 24:04reset it. On October 22nd, it will also
  520. 24:06be there until October 22nd. This means
  521. 24:08you have a sort of reset, the same as
  522. 24:10what Codex had, uh, that you can use
  523. 24:11during your most intensive periods. And
  524. 24:13now, let's see how it performed here,
  525. 24:15since it finished. Uh, let's see the
  526. 24:18final results. Here they spent, or
  527. 24:21rather, it took...Let's see. How long
  528. 24:23did it take? That's good. Uh, we sent
  529. 24:27the prompt 40 minutes ago and it
  530. 24:29finished 3 minutes ago, so it took, uh,
  531. 24:32something like 35 minutes. This would
  532. 24:35be Opus. And over here, uh, it
  533. 24:37delivered 19 minutes ago, and we sent
  534. 24:39it at the same time. It took 21 minutes
  535. 24:42, I mean, Opus 5. It took longer than,
  536. 24:46uh, GPT6 Astra. But let's see the
  537. 24:49result, which is what matters. Here we
  538. 24:51are going to open it with, or actually,
  539. 24:54I’ll just tell it to open it in
  540. 24:56Chrome to see it. I’ll hit enter, and
  541. 24:59here I’ll also tell it to open it in
  542. 25:01Chrome to see it. Uh, what we asked it
  543. 25:04to do was create this page of some
  544. 25:06exploded-view headphones. Uh, we were
  545. 25:09also connected to Highfield. Highfield
  546. 25:12is the processor we use to be able to,
  547. 25:15uh, centralize all the AI models in one
  548. 25:19place. It aligns very well with
  549. 25:22everything we do because, in the end,
  550. 25:24it allows us to have everything,
  551. 25:26obviously, in one place, including the
  552. 25:28latest AI models. As we can see here,
  553. 25:30if I go to images, I can create images
  554. 25:33with several models just by uploading
  555. 25:35images to the same place. So, of course
  556. 25:38, it came in here, generated the images
  557. 25:40, then the videos, right? It generated
  558. 25:42the exploded view with Sidans 2.5,
  559. 25:44which were the things. Uh, this
  560. 25:47Genjutsu one is also super interesting;
  561. 25:49it’s for changing your position, or
  562. 25:51changing your shirt or some outfit or
  563. 25:54whatever. And obviously, it has MSP. We
  564. 25:58connected via MSP to Cloud Code and
  565. 26:00asked it to create the pages for us,
  566. 26:02and this is what it gave us. This right
  567. 26:06here is the page from, uh, GPT6 Astra.
  568. 26:10As we can see, uh, it works really well
  569. 26:12. The exploded view is quite good. At
  570. 26:15the end, there’s just a small bug
  571. 26:17where it kind of transitions to this
  572. 26:19new thing, let's say, or it kind of
  573. 26:21explodes, or rather, it becomes a bit
  574. 26:24transparent. But overall, I think
  575. 26:26it’s super, super good." Sobora audio
  576. 26:28, your world on pause, "it shows us.
  577. 26:30There’s a small bug here too, uh, but
  578. 26:33the rest works well in the end; it gave
  579. 26:35us the exploded view, it gave us the
  580. 26:37detail, and the page is responsive.
  581. 26:40Obviously, we don’t want it to be, uh
  582. 26:43, public, right? Because, in the end,
  583. 26:46it’s not something we’re interested
  584. 26:47in. This has been a pretty good detail,
  585. 26:49as it maintained consistency really
  586. 26:51well for us. This is the page that GPT6
  587. 26:53Astra built for us. Let’s see the one
  588. 26:55Opus 5 built for us. Wow, it looks
  589. 26:58pretty good, huh? Let’s see if, well,
  590. 27:01it’s obviously not controlling the
  591. 27:03browser here. I don’t want it to
  592. 27:05control it for us, so we’ll just open
  593. 27:07it here because I don’t want it
  594. 27:09controlling my browser. Uh, and it
  595. 27:11looks pretty good and, wow, wow, if we
  596. 27:13start scrolling down. Okay, okay, if we
  597. 27:16start going down, I think, uh it did a
  598. 27:20very good job, at least at first glance
  599. 27:24. Like, look, it's like a little
  600. 27:26presentation that moves forward, like
  601. 27:28with the view, wow, wow, it's super
  602. 27:31good. I think, look, it even gave us
  603. 27:33this image here with icons. This video,
  604. 27:36uh, reserved in Obsidian, it gave us
  605. 27:38the different colors, and well, I think
  606. 27:41we have a clear winner here, at least
  607. 27:43in feeling. I have to say that this is
  608. 27:46the best page I've seen with such a
  609. 27:48simple prompt. I'm not exaggerating, I
  610. 27:50think it's the best. Uh, wow. Let's see
  611. 27:54the responsiveness if we close it here.
  612. 27:57Wow, this is really impressive. It's
  613. 28:02the best page I've seen with such a
  614. 28:04simple prompt, you know? I mean, we
  615. 28:07didn't even manage to let it finish. In
  616. 28:10fact, I interrupted it ahead of time
  617. 28:12because it kept doing, like, the whole,
  618. 28:14the analysis, but wow, wow. I have to
  619. 28:17say it's actually quite good, honestly.
  620. 28:19Uh, congratulations to Claude. I think
  621. 28:21it did a better job here. It took
  622. 28:22longer. Yes, twice as long, but
  623. 28:24basically I don't care if it takes
  624. 28:26longer, I care about a better result.
  625. 28:29If I wanted to use cheap models, I'd
  626. 28:31use others like Jeff, which by the way,
  627. 28:33I recommend you stay tuned because
  628. 28:34something very interesting is coming
  629. 28:36with how I'm using it today. This is
  630. 28:39excellent. It's excellent. Well, anyway
  631. 28:42, uh, returning to the initial question
  632. 28:44, is it worth it for us to switch from
  633. 28:46GPT-4o to Opus? In my opinion, if
  634. 28:50you're not into scientific research or
  635. 28:53biology, chemistry research, or very
  636. 28:55specific things you can see there, the
  637. 28:58answer is yes. I think the Opus 3.5
  638. 29:02announcement in the end was super good,
  639. 29:05super relevant. Uh, it obviously
  640. 29:08deserves a serious test if you do
  641. 29:09things like long programming, or tasks,
  642. 29:11or reports, or whatever, anything
  643. 29:13related to design. Look at this. Uh,
  644. 29:17and obviously complemented with
  645. 29:18Artifacts, it works super well because
  646. 29:20we can get these kinds of things all in
  647. 29:22one place. So it's epic, it's epic. Uh,
  648. 29:28regarding business automation, I think
  649. 29:30it works quite well. I still have to
  650. 29:32test it a bit more. If you want me to
  651. 29:34make a video below comparing GPT-4o
  652. 29:36with Opus 3. Like in various use cases,
  653. 29:39let me know. Uh, really, I, uh, read
  654. 29:41all the comments, I respond to all the
  655. 29:44comments, uh, I really value it when
  656. 29:46you leave me comments. And it's not
  657. 29:48just me; I also have a scraper that
  658. 29:49pulls comments from my videos and
  659. 29:51provides feedback, like:" Oh yeah, look
  660. 29:53, a lot of people are asking for this
  661. 29:54video, "and then I create it. So, it's
  662. 29:57also really helpful when you leave
  663. 29:59those things, you know? So well, Opus
  664. 30:02is back, which makes me very happy
  665. 30:04because I've always been a fan of
  666. 30:06Claude Code. I still need to test it a
  667. 30:09bit more, perhaps on the computer use
  668. 30:10side of things. If you're deciding
  669. 30:12between Codex or Claude Code, on which
  670. 30:14one to learn. I have courses for both,
  671. 30:17completely free. You can find them on
  672. 30:19YouTube, and if you want to find out as
  673. 30:21soon as all these things come out, if
  674. 30:24you look here at Imperio Argéntico, we
  675. 30:26posted it as soon as it dropped—I
  676. 30:28mean, I posted it an hour ago. It
  677. 30:30already has 140 comments. We are all
  678. 30:32super happy there. Obviously, I'm going
  679. 30:35to upload this video, so yeah, many,
  680. 30:37many good things are happening inside
  681. 30:39Imperio, and this is where we find out
  682. 30:41about everything long before everyone
  683. 30:43else so we have a competitive edge.
  684. 30:46Plus, we obviously have live sessions,
  685. 30:48we have the courses, we have everything
  686. 30:50. It's really good, it's really good.
  687. 30:52Stop by if you're interested in AI and
  688. 30:54are serious about building with AI.
  689. 30:57That said, I hope you liked this video.
  690. 30:59See you soon.

About this transcript

This page contains the full transcript of Anthropic lanza Opus 5.5: lo probé contra GPT-6 Astra by Benjamín Cordero, generated from the public captions YouTube serves with the video. The transcript has 4,815 words across 690 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.