OpenAI fights back — Transcript
Full transcript
- 0:04They're back. Okay, let me catch you
- 0:06guys up quick. Last week, Anthropic
- 0:08dropped a new model, Opus 5.5, and it
- 0:11was unbelievably good. It was so
- 0:13unbelievably good that OpenAI rushed out
- 0:15two model drops, GPT6 Soul and GPT6
- 0:18Luna. You might have noticed I didn't do
- 0:20a video on those models. There's a
- 0:22reason. They weren't that good. I was
- 0:24not particularly impressed with either
- 0:26of them and didn't really have much to
- 0:28say. But there was one other model I
- 0:30happened to get early access to that is
- 0:32now available for all. That model is
- 0:35called GPT 6.1 Soul. And this model made
- 0:38it very, very hard to film that Sonnet
- 0:405.5 video because I knew something was
- 0:42coming. Something surprisingly cheap,
- 0:44surprisingly capable, and most
- 0:46surprisingly, better than Astra. So
- 0:49yeah, I got a lot to say about this one.
- 0:52I'm filming this at 2 in the morning,
- 0:54right after filming my Sonnet 5.5
- 0:56videos. So, pardon me for stumbling over
- 0:58a few words here and there. I'm doing my
- 1:00best to get this out as reasonably
- 1:01quickly as possible because I want to
- 1:04have some coverage and I'll be real. It
- 1:06is also quite fun to cover these things
- 1:08before they are out so you are getting
- 1:10my true honest take and not the
- 1:12distilled version of what everyone else
- 1:14is saying. I'm sure this model is going
- 1:15to cause some pretty crazy waves. So, uh
- 1:18it will be nice to have my take out
- 1:20initially separately first. As always, I
- 1:22feel obligated to remind you I do have
- 1:24early access, but I'm not being paid in
- 1:25any way, shape, or form. OpenAI has no
- 1:27influence over what and how I say
- 1:29things, just when. They've politely
- 1:31asked me to wait until the model is out
- 1:33to talk about it, which makes a lot of
- 1:34sense. But I have to wait for one other
- 1:36thing first. Today's sponsor. In order
- 1:38to build good software with agents, you
- 1:39need to get feedback. And let's be real,
- 1:41they're getting a lot of that feedback
- 1:42from our CI. That's why we've all been
- 1:44seeing our CI bills skyrocket. And also
- 1:46why we've been getting more and more
- 1:47frustrated with GitHub actions. Today's
- 1:49sponsor is depot and they're here to
- 1:50solve all of this and more. Not only can
- 1:52they make your CI up to 10 times faster,
- 1:54as well as your Docker builds up to 40
- 1:56times faster, especially when they're
- 1:58downloading cache, they're also cheaper
- 2:00and they give better feedback for your
- 2:02agents. All this is possible due to
- 2:03depot metal. They're running their own
- 2:05bare metal with AMD epic processors that
- 2:07are way faster than what you get from
- 2:09traditional CI providers like of course
- 2:11GitHub actions. If you want it to be a
- 2:12drop in replacement, it absolutely can
- 2:14be, but APIs are so much better that you
- 2:17should probably use those instead. They
- 2:18enable parallelization and most
- 2:20importantly resilience when GitHub
- 2:21inevitably goes down randomly for no
- 2:23good reason. We've had our releases get
- 2:25blocked because we weren't using depot.
- 2:27And I'm so thankful that I've been
- 2:28moving more and more stuff over. For
- 2:29example, when Ben moved pick thing over
- 2:31to bun, we immediately had some CI
- 2:33failures. Normally, this would be
- 2:35obscure piles of text that our agents
- 2:36pars through for us. But when we use
- 2:38depot, it becomes way easier to see.
- 2:40They'll even analyze the failures and
- 2:42give suggestions which make it much
- 2:44simpler to get this feedback back to our
- 2:46agents. This is especially useful when
- 2:47you tell your agents that you can use
- 2:49depot because they'll no longer have to
- 2:50push changes and wait for that to
- 2:52trigger a build. They can just run the
- 2:54CLI to trigger the exact same CI that
- 2:57you'd be triggering through GitHub
- 2:58instead. No longer do you have to file
- 3:00PRs with broken code just to get
- 3:01feedback to your agents. They could just
- 3:03run a tool instead. Your agents will
- 3:05also get way better breakdowns of what
- 3:07is taking so long in your actual CI runs
- 3:10so that you can figure out how to
- 3:11improve them and make them faster and
- 3:13more reliable. You and your agents
- 3:14deserve faster Docker, faster build
- 3:16times, faster CI, better results, and
- 3:18ideally a cheaper price. Get all of that
- 3:20and more at swive.link/devo.
- 3:22Let's talk about this model a bit
- 3:24because it is not quite what I expected
- 3:26and it's probably not what you guys
- 3:28expected either, especially when you
- 3:29consider that GPT6 soul just came out
- 3:32like a week ago. It'll be around a 1
- 3:35week gap from 6.0 soul to 6.1 soul. I
- 3:38also want to disclose the numbers I'm
- 3:39currently showing on my screen are
- 3:41unlikely to be exactly accurate because
- 3:44I am running Terminal Bench for myself.
- 3:46In the first two times I ran it, I
- 3:47screwed things up. The third one seems
- 3:49to be doing much better. I didn't run it
- 3:50on medium initially, so there's a miss
- 3:52there. But low, high, XH high, and max,
- 3:55although the max run is incomplete, so
- 3:57I'm currently back filling scores from X
- 3:59high for the ones that Max either got
- 4:01wrong or in a previous run or didn't do
- 4:03yet because it takes like eight plus
- 4:05hours and some of these tasks. I this
- 4:07bench is nuts. It's I'm more skeptical
- 4:09of benchmarks than ever now that I've
- 4:11been running a lot more of them myself
- 4:12in order to get the coverage I want to
- 4:13give here. For what it is worth,
- 4:15Terminal Bench 4 is state-of-the-art
- 4:17score here. As is deep SWE, although
- 4:20this one's weirder because it goes down
- 4:22on X high and max and stays even on low
- 4:25and high. But those even low and high
- 4:27scores are scoring around what Astra did
- 4:30on high. The difference being it's doing
- 4:32it for comically cheaper. Switch over to
- 4:35the log scale, you'll see what I mean.
- 4:37This model on low costs 21
- 4:41versus Astra on low costing $1.46 and
- 4:44Opus 5.5 on max getting the same score
- 4:46for $14.65.
- 4:49While I will gladly admit that Deep Su
- 4:51is far from a perfect measure of how
- 4:54good a model is at day-to-day code work,
- 4:56the fact that 61 Soul is scoring the
- 4:58same as Opus and is also 73x cheaper is
- 5:02at least worth noticing. Here's where
- 5:05I'm going to say some of the things that
- 5:06I probably shouldn't. Considering that
- 5:08GBT6 Soul came out last week on Tuesday
- 5:11and this model's coming out this week on
- 5:14Tuesday, I think it's reasonable to
- 5:16infer that 6.1 Soul was not meant to be
- 5:196.1 Soul. There are things I'm not
- 5:22supposed to say, and I'm definitely
- 5:23walking a thin line here by sharing it.
- 5:25So, uh, I hope this proves I'm not paid
- 5:27off by Open AI because I'm about to give
- 5:29you guys info that I Yeah, just let let
- 5:32me get through this. First and foremost,
- 5:336.1 is significantly smarter than GPT6,
- 5:36thereby indicating this isn't just a one
- 5:39bump. There is something fundamentally
- 5:40different here. Next point is that it
- 5:42has meaningfully slower tokens per
- 5:44second. That tends to indicate the model
- 5:46is bigger. Hard to know for sure. Seems
- 5:49like this model might be different. Most
- 5:50importantly, we have Tibo's tweet. What
- 5:53am I referring to there? Well, right
- 5:56before I started filming, Tibo dropped
- 5:58quite a wall of text. The thing I want
- 6:01to emphasize here is first off that the
- 6:03pro $200 subscription is back. But more
- 6:05importantly, but more importantly is
- 6:08this sentence. They are changing how
- 6:09they calculate the usage in the sub. In
- 6:11effect, if you do the math, it will net
- 6:14out at half the dollar in API spend
- 6:17compared to the old pro $200 plan. Why
- 6:20in the world would they do this?
- 6:22especially right now where there's
- 6:23allegedly an internal code red because
- 6:26Opus 5.5 is so unbelievably good and has
- 6:29made the $200 quad code sub such an
- 6:31unbelievable value. The only reason in
- 6:33the world Tibo would post this right now
- 6:34is uh I don't know, maybe a new model is
- 6:37coming where their margins aren't as
- 6:38good. So, the ability to subsidize has
- 6:40gone down because remember you can get
- 6:43$8 to $9,000 of usage in a month on the
- 6:46$200 Cloud Code plan and you can get
- 6:49over 12 grand on the $200 codeex plan. I
- 6:52did actually run a lot of numbers before
- 6:53this and the amount you could get on
- 6:54Astra did go down slightly closer to
- 6:57like eight grand or so. Hard to know for
- 7:00sure because they differ for everyone
- 7:01everywhere and it's hard to log all of
- 7:03this stuff, but from my math roughly 9
- 7:05grand a month of usage. And that's where
- 7:07the price for this model comes in. This
- 7:10model is $2 per million input tokens and
- 7:12$10 per million out. This makes it way
- 7:14cheaper than 5.6 Soul was at launch.
- 7:16Half the price of 5.6 Soul after
- 7:17discounts and the same price as GPT6
- 7:19Soul. 1/5 the price of Astra. However,
- 7:24this is not the whole story cuz cash
- 7:26reads matter. And the cash read price
- 7:28for this model is going to be 10 cents
- 7:31per mill in. That's a big deal. OpenAI
- 7:34has not changed cash read price as far
- 7:36as I know ever before. It's always been
- 7:38exactly 10% of the normal read price.
- 7:41That makes it a 90% discount and now
- 7:43it's a 95% discount. That means they cut
- 7:45the cache read cost in half, massively
- 7:49reducing the cost for real world agentic
- 7:51use, which to be clear is what we're
- 7:53using these for most of the time. So
- 7:55this makes the model absurdly cheap for
- 7:58doing realworld code work. That also
- 8:00means that they are almost certainly
- 8:02cutting into their margins.
- 8:03Historically, these margins are rumored
- 8:05to be as high as 95%. Like for every $10
- 8:09you spend, they only have to spend 50.
- 8:12And as crazy as that sounds, it makes a
- 8:13lot of sense. especially when you
- 8:14consider how expensive it is to make and
- 8:15train these models. But that also gives
- 8:17them wiggle room to change things around
- 8:19a bit, which appears to be what's
- 8:21happening here. That also means that if
- 8:23they were to keep subsidizing the same
- 8:25level that they were on the
- 8:26subscriptions that your electricity cost
- 8:28for your sub would be more than you're
- 8:30paying. So, I get why they have to
- 8:32change this. They've kind of just left
- 8:34the details out there for us to uh
- 8:36reverse engineer. So, uh take this as
- 8:38you will. 6.1 coming so fast seems to
- 8:42indicate it is not just a new snapshot
- 8:44of GPT6. So let's talk more about this
- 8:46model. As I was showing earlier seems
- 8:49really good at Terminal Bench. Every
- 8:51time we refresh the numbers change
- 8:52because new runs come in and it looks
- 8:54like Max failed some things that X high
- 8:56passed which is why it just dropped a
- 8:57bit. But again pretty much all of these
- 9:00even high and X high are scoring higher
- 9:03than anything else ever has. And this is
- 9:05for me running this benchmark on random
- 9:07VMs on my network. So, uh, not the best
- 9:10suite to test against. I also had to
- 9:11drop three particular tasks from it
- 9:13because they expected an H100 to work
- 9:15against, which I make decent money. I
- 9:17don't make H100 money, okay? But none of
- 9:20this is real world code work. So, let's
- 9:22talk a bit about that. Obviously, we'll
- 9:25have all the fun things like fish slop
- 9:26near the end, so stay tuned for that.
- 9:28But, I just want to fixate a bit on the
- 9:30costs here because the most expensive
- 9:32run with 6.1 soul for me was about $1.38
- 9:36per task. And the cheapest run with Opus
- 9:385.5 was $512.
- 9:42That's a four to 5x gap from the
- 9:44cheapest Opus to the most expensive
- 9:46soul. So, at this point, I would imagine
- 9:49you are hoping and praying this model is
- 9:51good and that it can actually replace
- 9:53Opus 5.5 for day-to-day work. And I
- 9:55promise we'll get some good answers to
- 9:56that in a bit. But first, we need to
- 9:58talk a bit about model behaviors here
- 10:00because this model is a part of the GPT6
- 10:04family, which means it has uh the
- 10:06behaviors that are worth talking about.
- 10:09I know I cite this diagram a lot, but
- 10:11there's a reason for it. The thing that
- 10:12made me so frustrated with GBD6 Astra
- 10:15wasn't that it was less intelligent than
- 10:17the best models from Anthropic, because
- 10:19it was more intelligent than the best
- 10:20models from Anthropic, and I would argue
- 10:21in many ways still is. But there is a
- 10:23problem. It is also dumb. It is smart
- 10:26and dumb at the same time. GB6 Astra
- 10:29would just randomly spike into the
- 10:30dumbest I've seen a model do
- 10:32this year. Even worse than like some of
- 10:34the small openweight models I play with.
- 10:36It's still so deeply frustrating that
- 10:38Astra does this because on the other end
- 10:40when it does well, it's unbelievable.
- 10:43But these spikes got to the point where
- 10:44I effectively churned. I was only using
- 10:47my codec subs for computer use and I
- 10:49ended up just leaning on to Fable 5.1
- 10:51and obviously now Opus 5.5 for almost
- 10:54all of my day-to-day work. So, have they
- 10:56addressed the spikiness? Has GBD 6.1
- 10:59Soul fixed the problems that I was so
- 11:01frustrated about with Astra? I would say
- 11:04mostly, not entirely, but for the most
- 11:07part, yeah, this is a much better model.
- 11:10Its peaks are not as high. This is not
- 11:12the incredible revolutionary 3D
- 11:15capabilities that we saw with Astra. In
- 11:17fact, I would put it slightly below 5.5
- 11:19opus in most of those types of things.
- 11:22It is not as good at computer use as
- 11:23Astra, although it is close enough to
- 11:25the point where I have been happy using
- 11:27it for all of my day-to-day work. I
- 11:29actually had 6.1 Soul go through all of
- 11:31my emails and find invoices that I had
- 11:33forgotten to pay or was behind on,
- 11:35mostly like investing stuff, and set up
- 11:37new tabs in Chrome for every investment
- 11:40I needed to wire, fill out all the
- 11:41details for me, and just leave me to hit
- 11:43send. It didn't get a single thing
- 11:44wrong, and called out additional stuff
- 11:46that I absolutely would have missed if I
- 11:48was doing this work myself. So, I'm
- 11:49literally trusting this model to wire
- 11:51money for me. It's trustworthy enough
- 11:53for that. And honestly, I don't know if
- 11:54I would have trusted Astra with that due
- 11:56to the spikiness. 6.1 Soul much, much
- 11:59less spiky. From what I've heard from
- 12:00the other testers, they seem to agree
- 12:02with this analysis. I know for a fact
- 12:04that Julius and Ben, who have also been
- 12:05testing, have had a much better
- 12:07experience with this than Astra in terms
- 12:08of the spikiness. Julius called the
- 12:10model incredible. Ben called it
- 12:12incredibly boring. And I think that's
- 12:13the best place you can be for a model
- 12:15drop like this. But as I had mentioned
- 12:16before, its peaks are not as impressive.
- 12:19While it does quality work the majority
- 12:21of the time, there are some tasks that
- 12:23are just at the edge of its capability
- 12:25that it will start to do weirder stuff
- 12:27on. For the most part, it's fine. But I
- 12:30I'm still reaching for Opus a decent
- 12:32bit. We'll talk more about the
- 12:33comparison later. I do default to this
- 12:35model for a bunch of stuff, though.
- 12:37First off, as I mentioned before,
- 12:39computer use. I can't wait for Ultraast
- 12:41to be like an actual thing you can use
- 12:43with OpenAI models because when it is,
- 12:45this model is going to be crazy on it
- 12:47because it can already figure out how to
- 12:49navigate computer use totally fine. If
- 12:51it can suddenly do it six times faster,
- 12:53it's going to be unbelievably fun. Still
- 12:56not quite as good as Astro, but more
- 12:57than good enough that for the price
- 12:58difference, I wouldn't even think twice
- 13:00about it. But as I mentioned before,
- 13:01there are certain things I would still
- 13:03occasionally use Astra for that I am
- 13:05more than happy to use Soul for. One of
- 13:08those things is deep code reviews. I
- 13:10have still found OpenAI models and the
- 13:12like Rottweiler nature where they'll dig
- 13:14into a problem and shake it and tear it
- 13:16to pieces until they find every single
- 13:17thing wrong with it. I find 6.1 soul to
- 13:20be incredibly capable in this particular
- 13:21way. So, as you can probably guess, I
- 13:24had 6.1 Soul do some deep audits on
- 13:27orchestrator v2 and other parts of my
- 13:29real world code bases. In my
- 13:30orchestrator v2 audit, it performed
- 13:33nearly identically to Astra. I do
- 13:34believe it was slightly higher a score.
- 13:37Okay, not in this analysis, but in my
- 13:38other analysis, it did actually score
- 13:40very, very slightly higher, but it did
- 13:42it at about half the price. 297 versus
- 13:45584. Sonnet was still cheaper and Opus
- 13:48was slightly cheaper as well. The
- 13:50difference being neither of these models
- 13:52were anywhere near as thorough with
- 13:54their analysis. 6.1 Soul dug deep to
- 13:58find things, which is why it was able to
- 14:00get a score comparable to Astra,
- 14:02although it did admittedly burn way more
- 14:04tokens. Another task I've had a lot of
- 14:05fun testing with is asking the model to
- 14:07find opportunities to improve a code
- 14:09base. In this case, to improve T3 code,
- 14:12this is the one where Grock 4.7 scored
- 14:14strangely well. Of course, Astra scored
- 14:17way better at an 83.8 versus the 80.7
- 14:19from Grock 47, but GBD61 Soul hit it out
- 14:22of the park with an 87.4. 4. I didn't
- 14:26save all the prices for these runs. It's
- 14:28been a bit okay. But for Opus 5.5, it
- 14:31cost five bucks. And for Sonnet 5.5, it
- 14:33cost almost $9. With GBD61, it was
- 14:37$2.15.
- 14:39That's the difference. This model's
- 14:41token price is cheaper than Sonnet, but
- 14:43its token utilization is still
- 14:45maintaining OpenAI's usual efficiency,
- 14:48which results in just crazy price to
- 14:50performance. This whole thread was
- 14:52particularly fun because I had Opus 5.5
- 14:54review this model with a different name
- 14:57obviously, so I didn't know what it was.
- 14:58I went and edited the history after and
- 15:00it concluded very quickly this was a
- 15:01Frontier tier model. Its reviews and bug
- 15:04repros match the fixes that later
- 15:06merged. Its first draft code had real
- 15:08bugs which review bots caught. Four
- 15:10reviewers are still running. The local
- 15:12for Code reviews back Frontier tier
- 15:13again in a blind 10 model bench on the
- 15:15same prompt. 6.1 Souls placed first of
- 15:18the A7.4. It found the fish slop runs
- 15:20and compared those two. It did say 6.1
- 15:23souls quality output was slightly below
- 15:26Astras as well as the two OpenAI models
- 15:28with Opus 55 and Sonic 55, which we will
- 15:30absolutely show you in a bit. But I do
- 15:32want to call out the price here cuz it
- 15:34only cost $7 to run versus 15 for Sonic
- 15:3855 and 50 for Opus. Opus' honest tier
- 15:41call was that this model is incredible
- 15:43for scoped work, top of the frontier.
- 15:46Refine what's wrong and tell me the
- 15:47truth. I would choose it over Astra and
- 15:49about level with Opus for long
- 15:51unattended building. This was below
- 15:53Frontier follows its process rules even
- 15:55when they stop all progress and it does
- 15:57not ask for help. This I absolutely
- 15:59noticed. I had mentioned before a few
- 16:01times now that my TS Rust port that I'm
- 16:03making with Opus 5.5 is going way better
- 16:06than when I was working on that same
- 16:07port using Astra and Soul in the past. I
- 16:10had that port running for a while with
- 16:11this model and it made no progress. It
- 16:14burned a shitload of tokens, but it
- 16:15didn't actually improve the compiler at
- 16:17all. Opus was able to from scratch
- 16:19restart it and get it working in a day
- 16:22after I had spent months and hundreds of
- 16:24thousands of dollars in tokens with this
- 16:26model as well as with Astra and 5ixole.
- 16:29Opus did in like a grand in like a night
- 16:31with just two subscriptions with the
- 16:33quad plan. So for unattended long like
- 16:36heavy rewrite type stuff, Anthropic is
- 16:39just comically far ahead right now. And
- 16:41it also didn't have great judgment when
- 16:43I was using it for managing my fleet.
- 16:44And for those wondering, my fleet is all
- 16:46the computers I use for running all my
- 16:47agents and code because one computer is
- 16:49far from enough. I don't run any of them
- 16:50on this MacBook now. So, when I use this
- 16:52model to manage the fleet, it made a
- 16:54couple dumb mistakes here and there. To
- 16:55be fair, so is Opus. Aster is the only
- 16:57one that hasn't really made too many of
- 16:58those dumb mistakes. But, like, I'm
- 17:00going to be so real. I am entirely done
- 17:02using Astra after this model. After I
- 17:04had Opus do all of this review, I asked
- 17:06it how much does it think this model
- 17:07should cost. It guessed $5 per mill in,
- 17:1150 cents cashed, and 30 per mill out,
- 17:13putting it at Opus' prices roughly. It
- 17:16said that because it's performing like
- 17:17Opus. Its speed should add a premium
- 17:19because it is quite fast. And it's not a
- 17:22pro model, which is where it expects
- 17:24those higher like $100 out tiered
- 17:27pricing things to come. Pro models
- 17:28aren't really a thing anymore. We just
- 17:29use Fable and Aster, but you get the
- 17:31idea. This is the funniest part of the
- 17:32whole thread, though. If OpenAI wants
- 17:34people to adopt it, I would expect $3
- 17:37per mill in, 30 cents for cash, and $20
- 17:40instead. That would still be a fair
- 17:42price for what it does. To which I
- 17:44responded, if I told you it was $2 in,
- 17:46$10 out, and 10 cents per mill cash
- 17:48read, what would you think? I'd call
- 17:50that very aggressive pricing. For how
- 17:52you use it, it costs about a quarter of
- 17:54what I guessed. The cash price does most
- 17:56of the work. Agent workloads are 96%
- 17:59cash reads. So 10 cents for cash reads
- 18:01matters more than the $2 and $10
- 18:02headline prices. When it looked at all
- 18:05of my sessions, its price guess would
- 18:07have been $5,700. But after looking at
- 18:09these new prices, it redid the math and
- 18:11it would have been $1,550.
- 18:14That is a massive decrease. And for all
- 18:16my PR review type tasks, it was
- 18:18expecting those to be up to $10. And
- 18:20it's actually only up to $3. And that's
- 18:22for like heavy PRs with tens of
- 18:24thousands of lines of code. According to
- 18:26Opus, so don't blame me, blame Opus for
- 18:28saying this. First off, Opus says it
- 18:30becomes the default model for scoped
- 18:32work. Second off, it says that bloated
- 18:34system prompts barely matter anymore
- 18:36because of the cash pricing. It's just
- 18:38noise. Third, it says long loops are
- 18:40still a bad idea, but not because of
- 18:42money. It's cuz according to it, the TS
- 18:44ROSport wasted 4 days and made no
- 18:46progress at all. And Opus even said
- 18:48they'd be suspicious of it lasting. They
- 18:50expect this price to go up in the
- 18:52future. I cannot fathom OpenAI ever
- 18:54increasing the price for a model, but
- 18:56Opus thinking they will is hilarious and
- 18:58shows just how good a value the model
- 18:59is. I love this call out here. GBD 6.1
- 19:02Soul did three rounds of work for about
- 19:04half the cost of Sonnet's single round.
- 19:07A lot of this comes down to how context
- 19:09was managed, both because 6.1 soul is
- 19:11much more efficient, so it's not doing
- 19:12as many calls that bloat the context.
- 19:15It's not outputting as many tokens that
- 19:16are like building up over time. So the
- 19:19average number of tokens being read per
- 19:21request was only around 110,000 tokens
- 19:25versus 360,000 for Sonnet 5. The result
- 19:27is that Solot used under half as many
- 19:29input tokens as Sonnet, making this
- 19:31model significantly more efficient.
- 19:33Speaking of efficiency, I want to talk
- 19:34about these deep SWE scores a tiny bit
- 19:36more because this is a weird bench for
- 19:39me to have forked and include in these
- 19:41things. I actually did it for a
- 19:42different reason, not to compare against
- 19:446.1 soul, but to compare against a new
- 19:47release from Open Router, Jev Router.
- 19:50Open Router added Jev router to try and
- 19:52optimize costs with your requests. And I
- 19:54thought it was an incredibly stupid
- 19:56idea. Once I started running it against
- 19:58benchmarks, I confirmed it's an
- 19:59incredibly stupid idea. It turns out a
- 20:01model that cannot reason, that is given
- 20:03a prompt and no context, cannot make a
- 20:05good decision around how hard the
- 20:07problem is. and Jev router ended up
- 20:10being Deepseek v4.1 flash router for the
- 20:13vast majority of its runs. It around 60%
- 20:16of all the requests went straight to
- 20:17Deepseek 4.1 Flash, so didn't like it
- 20:20that much. It also routes to other
- 20:21smarter models, which should give it
- 20:23more of an advantage, but it ended up
- 20:25being more expensive than GPT6 Astro was
- 20:28on low while also taking four to five
- 20:30times longer cuz six Astro low took 4.6
- 20:33minutes and Jev router took 20. Jev
- 20:36Router's average task took 104 steps
- 20:39whereas GB6 Astros took 19. You get the
- 20:42idea. It wasn't very good. But the whole
- 20:44point of Jev Router is that it would be
- 20:46as cheap as possible to get a certain
- 20:48score. That was the promise on the tin.
- 20:51Whether or not you believe them is up to
- 20:53you, not me. I think it's
- 20:55Regardless, Jev Router was routing to
- 20:58Deepseek 4.1 Flash for the majority of
- 21:01its requests. Despite Jev router routing
- 21:03to the cheapest possible small
- 21:05openweight models from whatever provider
- 21:07will give it away for free, 6.1 soul on
- 21:10low got the same score for an eighth the
- 21:13price. OpenAI is here to destroy any
- 21:16wins anyone else has in terms of
- 21:18efficiency. Completing this bench in 4.8
- 21:20minutes for 21 cents with the second
- 21:22highest score I've ever seen on it is a
- 21:24massive achievement. Tying Opus 5.5
- 21:28which took 50 minutes per task on max. a
- 21:31tenth the time and a 70th the price for
- 21:34the same score. If your work fits within
- 21:37the things 6.1 Sonnet does well, you
- 21:40should probably use it for everything.
- 21:41But if your work doesn't fit in it
- 21:43particularly well, you should probably
- 21:45keep using Opus and maybe give Opus the
- 21:47ability to call 6.1 Soul when it should
- 21:50for various tasks. I'm almost certainly
- 21:52going to be setting things up so that
- 21:53Opus 5.5 can call 6.1 soul to do
- 21:56investigation work to try and like root
- 21:58cause bugs to do analysis of code bases
- 22:01to figure out what things need to be
- 22:02touched and why to help me triage real
- 22:05world work to help me review the work
- 22:07that Opus does and more. I'm kind of
- 22:09spoiling the ending here, aren't I? I'm
- 22:11going to keep using Opus 5.5 for now.
- 22:14Before I explain why, let me do the
- 22:16thing that I'm most excited for. Fish
- 22:18lop.
- 22:20The first thing you might have noticed
- 22:21is the inclusion of slop in fish slop.
- 22:25This model did the horrible thing I hate
- 22:28where it surrounded the game in a bunch
- 22:31of absolutely garbage UI. And coming to
- 22:34this right after the 5.5 Sonnet demo
- 22:37hurts me deeply because Sonnet 5.5 did
- 22:40not make graphics anywhere near this
- 22:42good-looking, but at least it made a UI
- 22:45that was nowhere near this awful. And
- 22:47man, do I wish the bad UI is where the
- 22:49issue stopped. I will turn on the sound.
- 22:53Oh god, it's blaring.
- 23:00It is stunning looking. The fish are
- 23:02some of the best. The model for the sub
- 23:05is way better. The propellers work way
- 23:07better. I'm going to mute the sound cuz
- 23:09that is looking pretty bad. I haven't
- 23:11even heard it honestly.
- 23:14But damn. like looks beautiful, but if
- 23:18you actually are playing it, one of the
- 23:19first things you'll notice is that the
- 23:21movement feels significantly worse than
- 23:24it does in either the Opus or the Sonnet
- 23:27versions that I have demoed in the past.
- 23:32It Yeah, it it moves jank. It also has a
- 23:36significantly worse frame rate than the
- 23:37versions from the other models. It does
- 23:40have higher graphic fidelity, so that
- 23:41makes sense. Like the models here with
- 23:44the plants are significantly better than
- 23:47they were with the Sonic version. The
- 23:49dares of the fidelity of the extras in
- 23:50the tank is absolutely hilarious. Like,
- 23:56yeah. But god damn, I'm so tired of the
- 23:59unnecessary text everywhere. This model
- 24:01does it worse than almost any I've ever
- 24:02seen before. We got take a breather
- 24:06paused. Your little world can wait. Back
- 24:09to the reef. Start a new tank. Slop01
- 24:13feeder submarine. A little underwater
- 24:16chaos.
- 24:18The big little goal. Your little
- 24:20ecosystem. Little fish become big
- 24:22earners. Four meals and they're all
- 24:24grown up. Make the family a little
- 24:27bigger. A little golden overachiever.
- 24:29There's so many of these. There's like
- 24:3120 plus of them. And I promise you guys,
- 24:34as soon as I saw this, I took a
- 24:36screenshot. I sent it to OpenAI and I
- 24:38crashed out in the Slack because I
- 24:39cannot fathom how they haven't fixed
- 24:41this problem. This model is
- 24:43unacceptably garbage at UI. It has
- 24:46regressed again. And if you're looking
- 24:48for a model that can make frontends that
- 24:50don't suck, go spend your money
- 24:51somewhere else because it should not be
- 24:53spent here. This model sucks at front
- 24:55end. It sucks at design. It has no taste
- 24:57and you're going to have to bring your
- 24:58taste yourself still. But it is
- 24:59admittedly really good at Blender. If
- 25:01you give it like a screenshot of a thing
- 25:03you want it to model in 3D and say,
- 25:04"Hey, you have Blender over the CLI. go
- 25:06make this. It will and it'll do a pretty
- 25:08damn good job, but I would never have it
- 25:10make the actual mechanics for my games
- 25:12because it feels awful to play. It also
- 25:15has like nowhere near as much gameplay
- 25:18loop. In fact, the first time I tried
- 25:20demoing this before filming, it just
- 25:22randomly game overed as I was like
- 25:24getting started in the first 30 seconds
- 25:26and never like said why. Actually, I
- 25:28think I technically beat it. I almost
- 25:31want to like take this version and hand
- 25:32it to Opus or Sonnet and say, "Hey, can
- 25:34you make this play better because the
- 25:36graphics are good but the game sucks."
- 25:38But when you combine how cheap it was to
- 25:40make this cuz like this was $5 I think
- 25:43to generate, that's pretty insane. And
- 25:46if you combine that with like ultra
- 25:48fast, if that ever happens, suddenly
- 25:50you're going to be able to make a game
- 25:52in a few minutes on demand. We're
- 25:55actually now getting to that threshold
- 25:57where game development is about to flip
- 25:59upside down because of how models are
- 26:01finally understanding threedimensional
- 26:03space and the tooling necessary to do
- 26:04these types of things. It's happening.
- 26:06As per usual, I was not allowed to put
- 26:08the code I wrote with this model inside
- 26:10of T3 Code or other open- source
- 26:12projects during the testing window. So,
- 26:14I had to use it exclusively on my
- 26:15internal projects like Lakebed as well
- 26:17as for auditing other work, which means
- 26:19I mostly use this for auditing other
- 26:21work. And I was very impressed. This is
- 26:23a real PR I was working on to fix a bug
- 26:25where my new little work tree like setup
- 26:28window that would appear in a new thread
- 26:29in T3 code would disappear if you left
- 26:31and came back. I had Claude code work on
- 26:33this, but this problem went pretty deep.
- 26:36So I wanted to make sure that whatever
- 26:37solution I came up with was very very
- 26:40very well vetted. While I personally
- 26:42still do not trust this model to write
- 26:44the code I'm trying to land, I
- 26:46absolutely trust it to review things.
- 26:47Ignore the GBD6 soul there, just a
- 26:50placeholder. So, when I had 6.1 Soul
- 26:53look through this, it found real
- 26:55problems that were entirely missed by
- 26:57Fable and by Opus. First, it called out
- 26:59that follow-up messages can stay blocked
- 27:01after the agent starts, which is very
- 27:03annoying if you want to cue a message,
- 27:05and also that recovered setup progress
- 27:07was disappearing too early. It figured
- 27:09all of these things out with a
- 27:10combination of reading the code and
- 27:11analyzing it as well as computer use. It
- 27:13was able to prevent me from merging a
- 27:15real regression in T3 code. So, I
- 27:17literally just copy pasted those things
- 27:18to Claude and then told it to take
- 27:20another look. Said better, but I'll
- 27:22still fix two small gaps. Remember, I
- 27:24can't use this model to code for this
- 27:26project at the time. So, I copy pasted
- 27:28that again over to Claude and it
- 27:30eventually got it good enough and then I
- 27:31finally merged. But that is what I like
- 27:33this model for, and I cannot wait to
- 27:35push its limits for actually coding.
- 27:37Although, I will say from the code that
- 27:38I did have the misfortune of reading, it
- 27:41is harder to justify merging this code
- 27:43than it is for code from Opus. Normally,
- 27:47I would make you guys wait for the Opus
- 27:48versus Soul video or the Sonnet versus
- 27:50Soul video, but I'll just spoil the
- 27:51details now. I like using them in tandem
- 27:54because I find Soul to be way better at
- 27:56reviewing and digging into the details,
- 27:58but I find Opus a more pleasant
- 27:59collaborator and significantly better at
- 28:01actually implementing code without
- 28:03getting blocked constantly throughout
- 28:05its work. And even now with the Rust
- 28:06rewrite of TypeScript, I find myself in
- 28:08a similar pattern where I have Opus 5.5
- 28:11just going and going and going, making
- 28:13the codebase work and work well. And
- 28:15then I had Soul come in and do an audit.
- 28:18And this is the funniest part. Remember
- 28:20before I said that I had Soul and Astra
- 28:22working on that TS Rust port for
- 28:24effectively months. I told Soul to come
- 28:26in and it got it from 83.7% to 100% in
- 28:30under a day. I was blown away by that
- 28:32that it had somehow unblocked the work
- 28:34that Astra and Soul were doing as well
- 28:36as 6.1 Soul. And I was absolutely blown
- 28:38away by that, that it had taken the work
- 28:39that 56 soul, 61 soul, and Astra had
- 28:42done over months and got it unblocked
- 28:44where it had been stuck for weeks and
- 28:46finished it. I was much more blown away
- 28:48when I had 61 soul take a look at that
- 28:51work and critique it. And what it
- 28:54brought up was that of the 1.8 million
- 28:56lines of code, 1.3 million were not
- 28:59being used. The reason was because Opus
- 29:02concluded all of the code from all the
- 29:04other agents was useless slop that had
- 29:06no chance of being recovered and it
- 29:08chose to rewrite it from scratch itself
- 29:10in another crate. So on one hand, the
- 29:12only reason the code worked was Opus,
- 29:13but on the other hand, the only reason
- 29:15the slop was still around was also Opus.
- 29:18So I had to have this model come in and
- 29:20clean up the mess that other OpenAI
- 29:22models had made because Opus didn't even
- 29:24notice the mess was still there. What
- 29:26I'm trying to say is this model
- 29:27absolutely has a place in your
- 29:28workflows. It could probably even be
- 29:30your default coding model and you
- 29:32wouldn't have too many issues with it,
- 29:34but I still find Opus to be a better
- 29:35collaborator overall. That said, I have
- 29:38almost no reason to use Sonnet anymore
- 29:40because this will effectively take its
- 29:42place. And you bet your butt the moment
- 29:44this model comes out, I'll be going and
- 29:46making adjustments inside of my cloud
- 29:47config because I already have it set up
- 29:49so that I can use soul inside of cloud
- 29:51code because I want to make sure opus
- 29:53knows this is the model to have review
- 29:55its work and investigate the things
- 29:57going on in the codebase. This is a damn
- 29:59good model and I'm really happy to have
- 30:00it. I wish we had something bigger,
- 30:02smarter, and more capable overall. I was
- 30:04really hoping for something to truly
- 30:06dethrone Opus 55 as my daily driver.
- 30:09This isn't it, and I'm not planning on
- 30:10canceling any of my cloud subs as a
- 30:12result of this release, but I am
- 30:13planning on taking a lot more advantage
- 30:15of my codec subs in my day-to-day work.
- 30:17Admittedly, in cloud code, this is an
- 30:19awesome release and an unbelievable
- 30:21price for what you're getting. But this
- 30:23does potentially mark the start of the
- 30:25end of the subsidization era. So, make
- 30:27sure you're subscribed so that you can
- 30:28be here when I cover all of that and
- 30:30more. God, I hope this doesn't get me
- 30:31cancelled online. I have no idea how
- 30:33others feel about this beyond like a
- 30:34handful of early access testers I've
- 30:36talked to. I legitimately don't know if
- 30:37people are going to love it or hate it
- 30:38or land somewhere between. I will know
- 30:40in a few hours, I guess, cuz it's a
- 30:43Yeah, it's 3:00 in the morning. I am
- 30:45going to go to bed now.
- 30:48This is a tiring one. Hopefully, I did a
- 30:50good job. Let me know in the comments.
- 30:51And until next time, peace, nerds. God,
- 30:55I'm so dead.
About this transcript
This page contains the full transcript of OpenAI fights back by Theo - t3․gg, generated from the public captions YouTube serves with the video. The transcript has 6,329 words across 873 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.