Anthropic lanza Opus 5.5: lo probé contra GPT-6 Astra — Transcript
Full transcript
- 0:00Anthropic just launched Opus 5. GPT6
- 0:05Astra. In their table, we can see it
- 0:08beats it in several programming tests
- 0:10and professional work. And the cost
- 0:12shown here, I think, is a huge part of
- 0:15this entire new announcement. This is
- 0:19an excellent model, but Astra is still
- 0:21ahead in two benchmarks that still
- 0:23matter to us, which are Business
- 0:24Workflow, right here, and Agentic
- 0:26Scientific Research. But now we can see
- 0:30that Opus is crushing it in everything
- 0:32related to coding, and it's doing very
- 0:35well in many other tests as well, even
- 0:38omitting comparisons with Astra in some
- 0:40cases. So, the question is, should I
- 0:43switch? Has Anthropic really returned
- 0:45to being the king? We are going to open
- 0:49the announcement, break it down from
- 0:51top to bottom, and analyze all the
- 0:52prices, the benchmarks, what each one
- 0:54means, and my personal take—whether I
- 0:56ended up liking it or not. We are going
- 0:59to create a page with images from GPT
- 1:01Image 2.5. We will connect it to
- 1:04Highfield and Sidans 2.5, and we will
- 1:06ask for different tasks to see if it
- 1:09truly has the same level of performance
- 1:11between GPT6 Astra, for example, and
- 1:13Opus 5. Before we start, I would really
- 1:17appreciate it if you could leave a like
- 1:19on this video, not only because it
- 1:21helps me, but because you also tell
- 1:23your algorithm that you like this style
- 1:24of content and it starts recommending
- 1:26more and more of it. Now, let's get
- 1:29down to business. Okay, before we begin
- 1:32, let's set two prompts running so that
- 1:34at the end of this video we can see the
- 1:36result and compare GPT6 Astra with Opus
- 1:385.5. I’m going to throw the same
- 1:42prompt at both of them, which is to
- 1:43build a page—a landing page—
- 1:44connected to Highfield using Sidan
- 1:46Caches 2.5. And well, we'll see what
- 1:49happens, and then we'll analyze which
- 1:51one was actually the best. So, here we
- 1:53can already see Opus 5. Max. And now
- 1:57I’m going to go into Codex and ask
- 1:59for the exact same thing, both at their
- 2:02highest reasoning levels. Okay, now
- 2:06let's see that, uh, a few seconds or
- 2:09minutes ago, Anthropic made this launch
- 2:12, which was their new model, Opus 5. It
- 2:18says it performs at the level of Claude
- 2:21Fable 5.1. One, but this is the
- 2:25interesting part: it costs 40%less to
- 2:28run than Opus 5, which is truly a
- 2:30stroke of genius considering new models
- 2:33like, uh, JEV, which are running very,
- 2:36very cheaply. And now we are starting
- 2:40to compete; while they aren't the same,
- 2:42we are definitely starting to compete
- 2:44quite a bit on price, because
- 2:45alternatives that are much, much
- 2:46cheaper are starting to appear. But why
- 2:50don't we see what this really means?
- 2:52And what is, what is the whole point of
- 2:55this launch, because I started reading
- 2:57and it looked quite interesting. Here
- 2:59we have the main benchmarks comparing
- 3:02them to Opus 5. With the previous
- 3:06version, Opus 5, Claude 5.1, and GPT-4o
- 3:10. And well, what I also want to
- 3:13highlight is that notice here they say
- 3:16it is the first model in their new 3.5
- 3:18family. In other words, Sonnet and
- 3:22Haiku could probably be coming in the
- 3:23next few weeks, or perhaps they'll
- 3:24discontinue Haiku. I don't know, let's
- 3:27see what happens, but there will
- 3:29undoubtedly be a new Sonnet or maybe
- 3:31even a Claude. Uh, from what I read, it
- 3:35functions quite similarly in reasoning
- 3:38levels to Claude 3.5, and this is their
- 3:41official launch blog. Uh, what I want
- 3:44to see here, well, we already know what
- 3:47we always do, let's translate it to
- 3:49English so there are no differences or
- 3:52problems with language barriers. And
- 3:55let's see the funny translations that
- 3:57Google Translate gives us here. Uh, and
- 4:00notice they tell you that in most tasks
- 4:03, its operation costs 40%less than Opus
- 4:063. And notice this comparison they are
- 4:09showing us here, uh, because further
- 4:11down we will see how much it costs
- 4:13according to Astra and GPT-4o. Uh, they
- 4:17announce a different generation of
- 4:19responses. They tell us that responses
- 4:22are 30%faster than Opus 3, uh, and they
- 4:24tell you that in addition to the price
- 4:27drop, they increased the usage limit to
- 4:295 hours on Pro, Team, and Enterprise
- 4:31plans. So, they raised the limits to
- 4:34these 5 hours. Great. Uh, and notice
- 4:37that they also told us here that they
- 4:40gave us a, uh, token reset, where I
- 4:42found this here in their post. Uh,
- 4:47besides increasing the 5-hour usage
- 4:49limit, they also gave people, or users,
- 4:52or subscriptions, a rate limit reset
- 4:54that can be used whenever you want.
- 4:59This is something Codex already had,
- 5:01which I found great because, basically,
- 5:03you have periods of higher intensity or
- 5:05work where you need these resets, so I
- 5:07think it was an excellent, excellent
- 5:09success. We never really know how much
- 5:13a subscription yields for us, so they
- 5:14can tell us yes, eh, it has a higher
- 5:16limit, but in the end we can never
- 5:18really verify it because this is
- 5:19something they adjust based on
- 5:21computing demand and the supply they
- 5:23have. If we start looking here, well,
- 5:26it communicates more naturally, blah
- 5:28blah blah blah. Right, yes, obviously,
- 5:30I’m always going to want to say what
- 5:32suits me best here. So let's go down to
- 5:36this, let’s get into this, which is
- 5:39what really interests us, which would
- 5:41be the different benchmarks between
- 5:43Fable 5.1 and the rest of the
- 5:45benchmarks. Eh, we’re going to do
- 5:49what we always do, take a screenshot so
- 5:51we can write on or show things. And
- 5:54let's see what each one means. Here we
- 5:57have a series of benchmarks that are
- 6:00standardizations carried out by
- 6:02different people or tests used to
- 6:04measure all these models in an attempt
- 6:07to have some kind of criteria for doing
- 6:10so. And Anthropic itself, in fact,
- 6:15tells us here that these benchmarks are
- 6:17even becoming less and less
- 6:19representative because something also
- 6:21happens here where they isolate the
- 6:23models from their harnesses. So, it's
- 6:27not exactly the flow one has when
- 6:29working in something like Claude Code
- 6:32or Codex, eh, because the model is
- 6:34being isolated a bit, but anyway, it is
- 6:36still an excellent, excellent parameter
- 6:39. The first one that caught my
- 6:43attention is this one here, agent
- 6:45coding or agentic coding, which is the
- 6:48Terminal Bench 4.0, which tells us in
- 6:51the end how well an agent can complete
- 6:53tasks from a terminal. I'm sorry, I
- 6:57have to take a screenshot because I
- 6:59really need to see this in English,
- 7:01otherwise it’s very hard for me, eh,
- 7:05to be able to read the numbers and
- 7:07things well here. So now, yes. Eh,
- 7:10agentic Terminal Bench 4.0. We have
- 7:13Opus here which scored 66%versus its
- 7:15previous version which was 52%. So, it
- 7:20won here and it also beat Astra by
- 7:22quite a bit, which is super good, eh,
- 7:24because Astra was already super good at
- 7:27this. I personally used it quite a bit,
- 7:31but here Opus 5. So, let's think of
- 7:39this benchmark as the ability to solve
- 7:41a task through the terminal, that is,
- 7:43where the computer’s actions are
- 7:46executed. It does it better than all
- 7:49the other models and would be the new
- 7:50state of the art. Next, we have another
- 7:54benchmark here that is also about
- 7:56coding, but I believe in the one above,
- 7:58Terminal Bench 4.0, they only ask a
- 8:00single coding-related question that is
- 8:02also solved. This one involves a series
- 8:06of more steps; I mean, it's like a
- 8:08problem, a project, it involves more
- 8:10steps unlike Terminal Bench 4. Frontier
- 8:13Code goes beyond just answering a
- 8:16single question, and here it gets 54%
- 8:18versus eh Fable which gets 53%, meaning
- 8:21it beats it and shows a substantial
- 8:24improvement over Fable 5.1. I can't
- 8:27wait for Fable 5.5 to come out because
- 8:29it's going to be very interesting, as
- 8:31it will surely show us significant
- 8:33improvements as well, without a doubt.
- 8:36Eh, next we have another one here for
- 8:39Agentic Coding again, and this one is
- 8:41Cursor Bench 4.0, which also performs
- 8:43better than Fable 5.1. But it intrigues
- 8:47me because they aren't measuring Astra
- 8:49here, but in the end, this is a task,
- 8:51it's also Agentic Coding, but it's done
- 8:53in the Cursor environment; that's why
- 8:56it's called Cursor Bench 4.0. Here it
- 9:00performed with almost 58%. And let's
- 9:04remember all this is based in the end
- 9:06on the sessions you have within the
- 9:08Cursor harness, which allows you to use
- 9:11different models, which is the
- 9:12interesting part. Yes, it intrigues me
- 9:16quite a bit why the Astra box isn't
- 9:18there. I'm going to remove this here.
- 9:20And this one too. Eh, it intrigues me
- 9:22quite a bit why it's not there, because
- 9:24obviously, it not being there doesn't
- 9:25mean zero. Next, we go to this other
- 9:29test which would be knowledge work, eh,
- 9:32which is quite the opposite of a
- 9:35percentage, sorry, but it is measured
- 9:39in points. And here we have 100 versus
- 9:42the others, eh, where Fable 5.1 was
- 9:45already at a super high number and GPT-
- 9:474o Astra wasn't as high. This is a
- 9:52benchmark that evaluates professional
- 9:54work from I don't remember the exact
- 9:57number of professions, but basically,
- 10:00it asks how good the outputs of this or
- 10:03that deliverable are in the end, eh?
- 10:08And then it assigns a score, and that
- 10:10score for Opus 5...Is 1846, and for the
- 10:12other models, it is lower. Here I was
- 10:17also a bit surprised that GPT-4o Astra
- 10:20reached 1542. Eh, again, this is a kind
- 10:25of score like the one used in chess or
- 10:27an Elo rating. Eh, I don't know, higher
- 10:31is obviously better. This is a
- 10:34benchmark that I am quite interested in
- 10:36, and I think it is one of the most
- 10:38important ones for people who build
- 10:40automations: the Business Workflow
- 10:42benchmark. Uh, this is a benchmark that
- 10:45Zapier put together, and it ultimately
- 10:47measures how well it connects with the
- 10:49different applications we use in our
- 10:51day-to-day business. I mean, it takes a
- 10:55task and then measures whether, uh,
- 10:57people can correctly complete that
- 11:00specific task. This is much closer to
- 11:03what is actually done and how we use it
- 11:04in our day-to-day lives. I mean,
- 11:05especially for people who run
- 11:07businesses. It's like looking for
- 11:09something in a CRM or, I don't know,
- 11:11updating a record in a database or
- 11:13contacting and following up with this
- 11:16person here. uh, coordinating actions
- 11:18between different tools with different
- 11:20connections, and so on. Yeah, I think
- 11:22this is an excellent benchmark. And
- 11:24here it gets 40%and it is beaten by GPT
- 11:27-6 Sol, uh, sorry, GPT-6 Astra with 41%
- 11:30. Side note: it wouldn't surprise me if
- 11:33GPT-6 Sol comes out while I'm uploading
- 11:35this video. I think it's very likely
- 11:37they'll do something like that, so we
- 11:39just have to wait. But, uh, here GPT-6
- 11:43Astra does win, but where Opus 5 really
- 11:46wins, uh, by quite a lot, rather. Is in
- 11:50the Humanities Last Exam. And this is
- 11:54an exam that I also like quite a bit
- 11:57because it ultimately gathers many
- 11:59questions or answers from different
- 12:02disciplines, sort of like a PhD, so to
- 12:05speak, and takes these specialists and
- 12:08then gives them a percentage, let's say
- 12:11, of, uh, evaluation in specialized
- 12:14knowledge. Uh, and here they measure it
- 12:17with tools, which is a bit more
- 12:19real-world. In fact, without tools, I
- 12:21don't really understand why it's done
- 12:23so much, since we are using tools in
- 12:25our day-to-day life anyway, but it's
- 12:27okay to isolate it; it's interesting.
- 12:30Uh, but yes, Opus 5 wins here. GPT-6,
- 12:35uh, Astra gets 67.7%versus 57.2%. So,
- 12:43here both results are with the tools.
- 12:46Obviously, uh, getting a sort of good
- 12:49grade in the end on an exam doesn't
- 12:52mean, uh, that it's the same as doing
- 12:55autonomous research in a certain way,
- 12:58but anyway, I think Opus 5 performs
- 13:01much better here. Than GPT-6 Astra. Uh,
- 13:07and the other category where GPT-6
- 13:09Astra does beat Opus 5. Is the
- 13:14Scientific Research benchmark, which is
- 13:16also a terminal-based science benchmark
- 13:17. And here, scientific research tasks
- 13:20are evaluated based on an agent working
- 13:22in a computer terminal. This Astra did
- 13:26come out on top with 64%versus Opus 5.
- 13:30Which was at 58, almost 59%. But anyway
- 13:35, the announcement does report standard
- 13:38deviation errors at several points at
- 13:41the end, so this could obviously still
- 13:44change, but it gives you a slight idea
- 13:47of what we're actually facing. This was
- 13:51the one I wanted to get to, which is
- 13:53OSWorld 2.0. Why? Because GPT-4o stood
- 13:57out quite a bit historically for having
- 13:59excellent computer use. Uh, and here
- 14:03Opus 5 is at 1.8%. Meaning, it beat or
- 14:08performed better than its previous
- 14:12models, but it isn't comparing it to
- 14:15GPT-4o Astra. We'll take a look at this
- 14:19test, but uh, it measures how well an
- 14:22AI navigates an interface designed for
- 14:25humans, which is where GPT-4o Astra was
- 14:28performing super well. If I go here and
- 14:33go to GPT-4o Astra and look for this,
- 14:37like their launch page, I'll be able to
- 14:42find here that they get a, uh, 72.6%.
- 14:47But look, they compare Opus here with
- 14:4970, Opus 5, obviously, with 70.2%. But
- 14:54if I come back here, it says Opus 5 has
- 14:5774%, so, like, a little bit more. But,
- 15:01anyway, here we are comparing, if we
- 15:04had to compare computer use, that is,
- 15:06how well an AI performs navigating
- 15:09interfaces designed for humans. Uh, GPT
- 15:13-4o Astra has 72.6%, Opus has 70, Opus
- 15:185 and the new model, uh, Opus 5. It is
- 15:23quite a bit higher than that, I mean,
- 15:26it is six or seven points higher than
- 15:29Opus 5. So, seven points higher than
- 15:33Opus 5 here would be, well, at 77%
- 15:37versus 72%. So it would still be
- 15:40winning. Obviously, we aren't using the
- 15:43same parameters here, the reference
- 15:45isn't exact, but it tells us they
- 15:47focused quite a bit on computer use,
- 15:49which is what made many of us start
- 15:51using GPT-4o Astra. So it's good to
- 15:54know that Opus has taken the lead again
- 15:55, at least here in everything related
- 15:57to computer use. And then we have the
- 16:01recognition of visual graphics, which
- 16:04is this series of, like, cartography,
- 16:07which also isn't compared to GPT-4o
- 16:10Astra, but we can see this is basically
- 16:13an interpretation of how well the
- 16:16models understand plans, axes,
- 16:19distances, a bit more thought out
- 16:21especially since it started being
- 16:24implemented with Blender, Rhino,
- 16:27AutoCAD, and all. all these things. Uh,
- 16:30it’s like a mapping test, so to speak
- 16:32, an understanding of space and the
- 16:34physical plane where it performed at 89
- 16:37%and 84%, in other words, 89%versus
- 16:4288.4%and 83%before. Again, we don't
- 16:45have the reference for GPT6, but well,
- 16:47let’s see how much it costs us to get
- 16:50it. Here they tell us, "The main
- 16:52advantage of Opus lies in its
- 16:54efficiency; it costs less per token
- 16:56than Opus 5 and uses fewer tokens per
- 16:58task, which is a 40%reduction in costs.
- 17:02Uh, well, a token is this small unit of
- 17:05measurement used to standardize the
- 17:08inputs and outputs that an AI has.
- 17:12It’s basically how this text is
- 17:14separated and it would be like gasoline
- 17:16; spending more tokens will only
- 17:18increase your bill. Here is Opus 5. It
- 17:24costs $ 4 for input and $ 5, uh, versus
- 17:28the $ 5 of Opus 5, and $ 20 for output
- 17:32versus $ 25 for the output of Opus 5.
- 17:37So, here for input it went down and for
- 17:39output it went down $ 5. Now, here, of
- 17:44course, it makes you a bit curious
- 17:47where this 40%comes from, if this isn't
- 17:5040%, this is only a one-fifth decrease,
- 17:54right? When you combine lower rates
- 17:58with lower consumption, in the end, you
- 18:00should arrive at this kind of 40%.
- 18:04Because it’s not only cheaper, but
- 18:07they also improved the cache writes
- 18:10here. So, in the end, all of this
- 18:12compounded is what makes us reach that
- 18:14figure. Look, uh, they also tell you
- 18:17that fast mode is available in Cloud
- 18:20Code and the platform with a speed up
- 18:22to 2.5 times, uh, and it costs you
- 18:24double, basically, just like all the
- 18:27others. The interesting thing, though,
- 18:31is if we compare it later with GPT6
- 18:33Astra and we look here at what would be
- 18:36the price, uh, which is $ 10 per
- 18:38million input tokens and $ 50 per
- 18:40million output tokens, versus these
- 18:43which would be 4:20, it would be up to
- 18:4560%less or more economical than Astra
- 18:48for better performance. So, it clearly
- 18:52depends on your field. Obviously, if
- 18:56you are involved in biology, maybe, or
- 18:58scientific research, you will want to
- 19:00keep using Astra, but for most tasks,
- 19:02uh, Opus 5.5 should work better for you
- 19:04. Uh, next here we have coding, we can
- 19:09see that it focuses on migration and
- 19:11audits of large projects. For example,
- 19:15here 200,000 lines in 3 hours. Before,
- 19:18it took more than 20 hours and consumed
- 19:192.5 times more tokens in an internal
- 19:21run. They did the same thing and
- 19:24managed to do it, well, we already saw,
- 19:26in less than 3 hours, I mean, from 20
- 19:28hours to less than 3 hours. Anyway,
- 19:30this is a case documented by them, uh,
- 19:32anyway, but look at this. Here we have
- 19:35what we already saw. Look at the curves
- 19:37, how they show and compare them to us.
- 19:40Grey being GPT-4o Astra and Opus 3.5
- 19:42being, obviously, the orange one here.
- 19:46On these curves, upwards you have a
- 19:47higher score, towards the right you
- 19:49have more cost, towards the left you
- 19:50have less cost per attempt. Obviously,
- 19:55this is also represented in a
- 19:57logarithmic way, I mean, it's not that
- 20:00it's proportionally more money, at
- 20:02least on the cost axis, but in the end,
- 20:05it really helps us visually represent a
- 20:08difference in the changes, at least.
- 20:12But what they want to show us here is
- 20:15that Opus 3.5 equals GPT-4o Astra in
- 20:18the terminal bench, in average
- 20:20reasoning, when it's at its maximum, so
- 20:23to speak. And if we look at the cost,
- 20:28the best attempt by GPT-4o Astra to get
- 20:31a 58%spent $ 7, and the attempt for
- 20:35almost 58%spent $ 3. I mean, here we
- 20:39have a reduction of practically half. I
- 20:43mean, we can do terminal bench tasks,
- 20:45all the agentic coding stuff, at half
- 20:48the price, achieving the best result of
- 20:50GPT-4o Astra, which is super good, at
- 20:53least to keep in mind if you are
- 20:55building things with AI, which is what
- 20:57matters most to us, right? Then, well,
- 21:01all of this is coding; here, the
- 21:03Frontier code also looks much better,
- 21:05especially in terms of price. You know,
- 21:08look at this. Here we have 40 cents for
- 21:11Opus on low, and here we have uh 1.5.
- 21:18Dollars, I mean, three times less for
- 21:21GPT-4o Astra while getting a lower
- 21:23result, and obviously, if we get to
- 21:26maximum reasoning, we get to similar
- 21:29things in this sense. But my point is
- 21:32that you can see the trend a little
- 21:35closer here, much more economical. The
- 21:37best point at 54%and the other at 53%,
- 21:41but spending $ 4 versus 80 cents, I
- 21:44mean, five times more. That's why it's
- 21:48also worth looking at this curve,
- 21:50because it helps us understand the
- 21:52relationship between costs and uh very
- 21:54well. the scores or the outputs. The
- 22:01following sections, in the end, show us
- 22:03how they improved in terms of security,
- 22:06uh, in everything related to prompt
- 22:09injection, which, well, is a super
- 22:11important part if you don't want to end
- 22:14up using your McDonald's chatbot as an
- 22:16LLM, uh, and obviously if you want to
- 22:19keep things secure. Then, if we keep
- 22:23scrolling down, we can find this test
- 22:25here that we also saw, moving a bit
- 22:27away from code. We're going to start
- 22:30entering reports, analysis, and
- 22:31presentations. It has an overall
- 22:33performance or score, or the Elo, which
- 22:36would be this score we saw earlier,
- 22:39significantly higher than all other
- 22:41models. Uh, if we keep scrolling down
- 22:45here, they start showing us different
- 22:48cases, uh, different testimonials on
- 22:50how it changed versus its previous
- 22:53models. For example, here both models
- 22:57are comparing an error or explaining a
- 23:00problem with an error that occurred in
- 23:02billing. Uh, so here we can see that it
- 23:05starts with the consequence, helping
- 23:07you understand what the error is.
- 23:11Unlike here, it explains the error
- 23:13first and then it sort of starts
- 23:14explaining the consequence, from what I
- 23:16understood from what I managed to read
- 23:19here. If we keep scrolling down, we
- 23:22will find more things regarding
- 23:23security. Yes, we keep scrolling down
- 23:26here. Obviously, more security stuff
- 23:28regarding alignment and, well, all
- 23:30these things that you can start looking
- 23:32at if you're also interested in how it
- 23:34works and how all the APIs connect. Uh,
- 23:38here, well, they restricted or
- 23:40unrestricted, so to speak, the capacity
- 23:42or the blocks that had been placed on
- 23:45biology. Uh, it even outperforms
- 23:48Mixtral 5.1 in many areas of work. And
- 23:51over here we already have the
- 23:53availability. It is available on all
- 23:55platforms, so to speak. If you go into
- 23:59Claude, you will find this here. I
- 24:02loved that they give you a limit to
- 24:04reset it. On October 22nd, it will also
- 24:06be there until October 22nd. This means
- 24:08you have a sort of reset, the same as
- 24:10what Codex had, uh, that you can use
- 24:11during your most intensive periods. And
- 24:13now, let's see how it performed here,
- 24:15since it finished. Uh, let's see the
- 24:18final results. Here they spent, or
- 24:21rather, it took...Let's see. How long
- 24:23did it take? That's good. Uh, we sent
- 24:27the prompt 40 minutes ago and it
- 24:29finished 3 minutes ago, so it took, uh,
- 24:32something like 35 minutes. This would
- 24:35be Opus. And over here, uh, it
- 24:37delivered 19 minutes ago, and we sent
- 24:39it at the same time. It took 21 minutes
- 24:42, I mean, Opus 5. It took longer than,
- 24:46uh, GPT6 Astra. But let's see the
- 24:49result, which is what matters. Here we
- 24:51are going to open it with, or actually,
- 24:54I’ll just tell it to open it in
- 24:56Chrome to see it. I’ll hit enter, and
- 24:59here I’ll also tell it to open it in
- 25:01Chrome to see it. Uh, what we asked it
- 25:04to do was create this page of some
- 25:06exploded-view headphones. Uh, we were
- 25:09also connected to Highfield. Highfield
- 25:12is the processor we use to be able to,
- 25:15uh, centralize all the AI models in one
- 25:19place. It aligns very well with
- 25:22everything we do because, in the end,
- 25:24it allows us to have everything,
- 25:26obviously, in one place, including the
- 25:28latest AI models. As we can see here,
- 25:30if I go to images, I can create images
- 25:33with several models just by uploading
- 25:35images to the same place. So, of course
- 25:38, it came in here, generated the images
- 25:40, then the videos, right? It generated
- 25:42the exploded view with Sidans 2.5,
- 25:44which were the things. Uh, this
- 25:47Genjutsu one is also super interesting;
- 25:49it’s for changing your position, or
- 25:51changing your shirt or some outfit or
- 25:54whatever. And obviously, it has MSP. We
- 25:58connected via MSP to Cloud Code and
- 26:00asked it to create the pages for us,
- 26:02and this is what it gave us. This right
- 26:06here is the page from, uh, GPT6 Astra.
- 26:10As we can see, uh, it works really well
- 26:12. The exploded view is quite good. At
- 26:15the end, there’s just a small bug
- 26:17where it kind of transitions to this
- 26:19new thing, let's say, or it kind of
- 26:21explodes, or rather, it becomes a bit
- 26:24transparent. But overall, I think
- 26:26it’s super, super good." Sobora audio
- 26:28, your world on pause, "it shows us.
- 26:30There’s a small bug here too, uh, but
- 26:33the rest works well in the end; it gave
- 26:35us the exploded view, it gave us the
- 26:37detail, and the page is responsive.
- 26:40Obviously, we don’t want it to be, uh
- 26:43, public, right? Because, in the end,
- 26:46it’s not something we’re interested
- 26:47in. This has been a pretty good detail,
- 26:49as it maintained consistency really
- 26:51well for us. This is the page that GPT6
- 26:53Astra built for us. Let’s see the one
- 26:55Opus 5 built for us. Wow, it looks
- 26:58pretty good, huh? Let’s see if, well,
- 27:01it’s obviously not controlling the
- 27:03browser here. I don’t want it to
- 27:05control it for us, so we’ll just open
- 27:07it here because I don’t want it
- 27:09controlling my browser. Uh, and it
- 27:11looks pretty good and, wow, wow, if we
- 27:13start scrolling down. Okay, okay, if we
- 27:16start going down, I think, uh it did a
- 27:20very good job, at least at first glance
- 27:24. Like, look, it's like a little
- 27:26presentation that moves forward, like
- 27:28with the view, wow, wow, it's super
- 27:31good. I think, look, it even gave us
- 27:33this image here with icons. This video,
- 27:36uh, reserved in Obsidian, it gave us
- 27:38the different colors, and well, I think
- 27:41we have a clear winner here, at least
- 27:43in feeling. I have to say that this is
- 27:46the best page I've seen with such a
- 27:48simple prompt. I'm not exaggerating, I
- 27:50think it's the best. Uh, wow. Let's see
- 27:54the responsiveness if we close it here.
- 27:57Wow, this is really impressive. It's
- 28:02the best page I've seen with such a
- 28:04simple prompt, you know? I mean, we
- 28:07didn't even manage to let it finish. In
- 28:10fact, I interrupted it ahead of time
- 28:12because it kept doing, like, the whole,
- 28:14the analysis, but wow, wow. I have to
- 28:17say it's actually quite good, honestly.
- 28:19Uh, congratulations to Claude. I think
- 28:21it did a better job here. It took
- 28:22longer. Yes, twice as long, but
- 28:24basically I don't care if it takes
- 28:26longer, I care about a better result.
- 28:29If I wanted to use cheap models, I'd
- 28:31use others like Jeff, which by the way,
- 28:33I recommend you stay tuned because
- 28:34something very interesting is coming
- 28:36with how I'm using it today. This is
- 28:39excellent. It's excellent. Well, anyway
- 28:42, uh, returning to the initial question
- 28:44, is it worth it for us to switch from
- 28:46GPT-4o to Opus? In my opinion, if
- 28:50you're not into scientific research or
- 28:53biology, chemistry research, or very
- 28:55specific things you can see there, the
- 28:58answer is yes. I think the Opus 3.5
- 29:02announcement in the end was super good,
- 29:05super relevant. Uh, it obviously
- 29:08deserves a serious test if you do
- 29:09things like long programming, or tasks,
- 29:11or reports, or whatever, anything
- 29:13related to design. Look at this. Uh,
- 29:17and obviously complemented with
- 29:18Artifacts, it works super well because
- 29:20we can get these kinds of things all in
- 29:22one place. So it's epic, it's epic. Uh,
- 29:28regarding business automation, I think
- 29:30it works quite well. I still have to
- 29:32test it a bit more. If you want me to
- 29:34make a video below comparing GPT-4o
- 29:36with Opus 3. Like in various use cases,
- 29:39let me know. Uh, really, I, uh, read
- 29:41all the comments, I respond to all the
- 29:44comments, uh, I really value it when
- 29:46you leave me comments. And it's not
- 29:48just me; I also have a scraper that
- 29:49pulls comments from my videos and
- 29:51provides feedback, like:" Oh yeah, look
- 29:53, a lot of people are asking for this
- 29:54video, "and then I create it. So, it's
- 29:57also really helpful when you leave
- 29:59those things, you know? So well, Opus
- 30:02is back, which makes me very happy
- 30:04because I've always been a fan of
- 30:06Claude Code. I still need to test it a
- 30:09bit more, perhaps on the computer use
- 30:10side of things. If you're deciding
- 30:12between Codex or Claude Code, on which
- 30:14one to learn. I have courses for both,
- 30:17completely free. You can find them on
- 30:19YouTube, and if you want to find out as
- 30:21soon as all these things come out, if
- 30:24you look here at Imperio Argéntico, we
- 30:26posted it as soon as it dropped—I
- 30:28mean, I posted it an hour ago. It
- 30:30already has 140 comments. We are all
- 30:32super happy there. Obviously, I'm going
- 30:35to upload this video, so yeah, many,
- 30:37many good things are happening inside
- 30:39Imperio, and this is where we find out
- 30:41about everything long before everyone
- 30:43else so we have a competitive edge.
- 30:46Plus, we obviously have live sessions,
- 30:48we have the courses, we have everything
- 30:50. It's really good, it's really good.
- 30:52Stop by if you're interested in AI and
- 30:54are serious about building with AI.
- 30:57That said, I hope you liked this video.
- 30:59See you soon.
About this transcript
This page contains the full transcript of Anthropic lanza Opus 5.5: lo probé contra GPT-6 Astra by Benjamín Cordero, generated from the public captions YouTube serves with the video. The transcript has 4,815 words across 690 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.