YouTube2Text

OpenAI fights back — Transcript

by Theo - t3․gg · 6,329 words · 873 segments · language en · Watch on YouTube

Full transcript

  1. 0:04They're back. Okay, let me catch you
  2. 0:06guys up quick. Last week, Anthropic
  3. 0:08dropped a new model, Opus 5.5, and it
  4. 0:11was unbelievably good. It was so
  5. 0:13unbelievably good that OpenAI rushed out
  6. 0:15two model drops, GPT6 Soul and GPT6
  7. 0:18Luna. You might have noticed I didn't do
  8. 0:20a video on those models. There's a
  9. 0:22reason. They weren't that good. I was
  10. 0:24not particularly impressed with either
  11. 0:26of them and didn't really have much to
  12. 0:28say. But there was one other model I
  13. 0:30happened to get early access to that is
  14. 0:32now available for all. That model is
  15. 0:35called GPT 6.1 Soul. And this model made
  16. 0:38it very, very hard to film that Sonnet
  17. 0:405.5 video because I knew something was
  18. 0:42coming. Something surprisingly cheap,
  19. 0:44surprisingly capable, and most
  20. 0:46surprisingly, better than Astra. So
  21. 0:49yeah, I got a lot to say about this one.
  22. 0:52I'm filming this at 2 in the morning,
  23. 0:54right after filming my Sonnet 5.5
  24. 0:56videos. So, pardon me for stumbling over
  25. 0:58a few words here and there. I'm doing my
  26. 1:00best to get this out as reasonably
  27. 1:01quickly as possible because I want to
  28. 1:04have some coverage and I'll be real. It
  29. 1:06is also quite fun to cover these things
  30. 1:08before they are out so you are getting
  31. 1:10my true honest take and not the
  32. 1:12distilled version of what everyone else
  33. 1:14is saying. I'm sure this model is going
  34. 1:15to cause some pretty crazy waves. So, uh
  35. 1:18it will be nice to have my take out
  36. 1:20initially separately first. As always, I
  37. 1:22feel obligated to remind you I do have
  38. 1:24early access, but I'm not being paid in
  39. 1:25any way, shape, or form. OpenAI has no
  40. 1:27influence over what and how I say
  41. 1:29things, just when. They've politely
  42. 1:31asked me to wait until the model is out
  43. 1:33to talk about it, which makes a lot of
  44. 1:34sense. But I have to wait for one other
  45. 1:36thing first. Today's sponsor. In order
  46. 1:38to build good software with agents, you
  47. 1:39need to get feedback. And let's be real,
  48. 1:41they're getting a lot of that feedback
  49. 1:42from our CI. That's why we've all been
  50. 1:44seeing our CI bills skyrocket. And also
  51. 1:46why we've been getting more and more
  52. 1:47frustrated with GitHub actions. Today's
  53. 1:49sponsor is depot and they're here to
  54. 1:50solve all of this and more. Not only can
  55. 1:52they make your CI up to 10 times faster,
  56. 1:54as well as your Docker builds up to 40
  57. 1:56times faster, especially when they're
  58. 1:58downloading cache, they're also cheaper
  59. 2:00and they give better feedback for your
  60. 2:02agents. All this is possible due to
  61. 2:03depot metal. They're running their own
  62. 2:05bare metal with AMD epic processors that
  63. 2:07are way faster than what you get from
  64. 2:09traditional CI providers like of course
  65. 2:11GitHub actions. If you want it to be a
  66. 2:12drop in replacement, it absolutely can
  67. 2:14be, but APIs are so much better that you
  68. 2:17should probably use those instead. They
  69. 2:18enable parallelization and most
  70. 2:20importantly resilience when GitHub
  71. 2:21inevitably goes down randomly for no
  72. 2:23good reason. We've had our releases get
  73. 2:25blocked because we weren't using depot.
  74. 2:27And I'm so thankful that I've been
  75. 2:28moving more and more stuff over. For
  76. 2:29example, when Ben moved pick thing over
  77. 2:31to bun, we immediately had some CI
  78. 2:33failures. Normally, this would be
  79. 2:35obscure piles of text that our agents
  80. 2:36pars through for us. But when we use
  81. 2:38depot, it becomes way easier to see.
  82. 2:40They'll even analyze the failures and
  83. 2:42give suggestions which make it much
  84. 2:44simpler to get this feedback back to our
  85. 2:46agents. This is especially useful when
  86. 2:47you tell your agents that you can use
  87. 2:49depot because they'll no longer have to
  88. 2:50push changes and wait for that to
  89. 2:52trigger a build. They can just run the
  90. 2:54CLI to trigger the exact same CI that
  91. 2:57you'd be triggering through GitHub
  92. 2:58instead. No longer do you have to file
  93. 3:00PRs with broken code just to get
  94. 3:01feedback to your agents. They could just
  95. 3:03run a tool instead. Your agents will
  96. 3:05also get way better breakdowns of what
  97. 3:07is taking so long in your actual CI runs
  98. 3:10so that you can figure out how to
  99. 3:11improve them and make them faster and
  100. 3:13more reliable. You and your agents
  101. 3:14deserve faster Docker, faster build
  102. 3:16times, faster CI, better results, and
  103. 3:18ideally a cheaper price. Get all of that
  104. 3:20and more at swive.link/devo.
  105. 3:22Let's talk about this model a bit
  106. 3:24because it is not quite what I expected
  107. 3:26and it's probably not what you guys
  108. 3:28expected either, especially when you
  109. 3:29consider that GPT6 soul just came out
  110. 3:32like a week ago. It'll be around a 1
  111. 3:35week gap from 6.0 soul to 6.1 soul. I
  112. 3:38also want to disclose the numbers I'm
  113. 3:39currently showing on my screen are
  114. 3:41unlikely to be exactly accurate because
  115. 3:44I am running Terminal Bench for myself.
  116. 3:46In the first two times I ran it, I
  117. 3:47screwed things up. The third one seems
  118. 3:49to be doing much better. I didn't run it
  119. 3:50on medium initially, so there's a miss
  120. 3:52there. But low, high, XH high, and max,
  121. 3:55although the max run is incomplete, so
  122. 3:57I'm currently back filling scores from X
  123. 3:59high for the ones that Max either got
  124. 4:01wrong or in a previous run or didn't do
  125. 4:03yet because it takes like eight plus
  126. 4:05hours and some of these tasks. I this
  127. 4:07bench is nuts. It's I'm more skeptical
  128. 4:09of benchmarks than ever now that I've
  129. 4:11been running a lot more of them myself
  130. 4:12in order to get the coverage I want to
  131. 4:13give here. For what it is worth,
  132. 4:15Terminal Bench 4 is state-of-the-art
  133. 4:17score here. As is deep SWE, although
  134. 4:20this one's weirder because it goes down
  135. 4:22on X high and max and stays even on low
  136. 4:25and high. But those even low and high
  137. 4:27scores are scoring around what Astra did
  138. 4:30on high. The difference being it's doing
  139. 4:32it for comically cheaper. Switch over to
  140. 4:35the log scale, you'll see what I mean.
  141. 4:37This model on low costs 21
  142. 4:41versus Astra on low costing $1.46 and
  143. 4:44Opus 5.5 on max getting the same score
  144. 4:46for $14.65.
  145. 4:49While I will gladly admit that Deep Su
  146. 4:51is far from a perfect measure of how
  147. 4:54good a model is at day-to-day code work,
  148. 4:56the fact that 61 Soul is scoring the
  149. 4:58same as Opus and is also 73x cheaper is
  150. 5:02at least worth noticing. Here's where
  151. 5:05I'm going to say some of the things that
  152. 5:06I probably shouldn't. Considering that
  153. 5:08GBT6 Soul came out last week on Tuesday
  154. 5:11and this model's coming out this week on
  155. 5:14Tuesday, I think it's reasonable to
  156. 5:16infer that 6.1 Soul was not meant to be
  157. 5:196.1 Soul. There are things I'm not
  158. 5:22supposed to say, and I'm definitely
  159. 5:23walking a thin line here by sharing it.
  160. 5:25So, uh, I hope this proves I'm not paid
  161. 5:27off by Open AI because I'm about to give
  162. 5:29you guys info that I Yeah, just let let
  163. 5:32me get through this. First and foremost,
  164. 5:336.1 is significantly smarter than GPT6,
  165. 5:36thereby indicating this isn't just a one
  166. 5:39bump. There is something fundamentally
  167. 5:40different here. Next point is that it
  168. 5:42has meaningfully slower tokens per
  169. 5:44second. That tends to indicate the model
  170. 5:46is bigger. Hard to know for sure. Seems
  171. 5:49like this model might be different. Most
  172. 5:50importantly, we have Tibo's tweet. What
  173. 5:53am I referring to there? Well, right
  174. 5:56before I started filming, Tibo dropped
  175. 5:58quite a wall of text. The thing I want
  176. 6:01to emphasize here is first off that the
  177. 6:03pro $200 subscription is back. But more
  178. 6:05importantly, but more importantly is
  179. 6:08this sentence. They are changing how
  180. 6:09they calculate the usage in the sub. In
  181. 6:11effect, if you do the math, it will net
  182. 6:14out at half the dollar in API spend
  183. 6:17compared to the old pro $200 plan. Why
  184. 6:20in the world would they do this?
  185. 6:22especially right now where there's
  186. 6:23allegedly an internal code red because
  187. 6:26Opus 5.5 is so unbelievably good and has
  188. 6:29made the $200 quad code sub such an
  189. 6:31unbelievable value. The only reason in
  190. 6:33the world Tibo would post this right now
  191. 6:34is uh I don't know, maybe a new model is
  192. 6:37coming where their margins aren't as
  193. 6:38good. So, the ability to subsidize has
  194. 6:40gone down because remember you can get
  195. 6:43$8 to $9,000 of usage in a month on the
  196. 6:46$200 Cloud Code plan and you can get
  197. 6:49over 12 grand on the $200 codeex plan. I
  198. 6:52did actually run a lot of numbers before
  199. 6:53this and the amount you could get on
  200. 6:54Astra did go down slightly closer to
  201. 6:57like eight grand or so. Hard to know for
  202. 7:00sure because they differ for everyone
  203. 7:01everywhere and it's hard to log all of
  204. 7:03this stuff, but from my math roughly 9
  205. 7:05grand a month of usage. And that's where
  206. 7:07the price for this model comes in. This
  207. 7:10model is $2 per million input tokens and
  208. 7:12$10 per million out. This makes it way
  209. 7:14cheaper than 5.6 Soul was at launch.
  210. 7:16Half the price of 5.6 Soul after
  211. 7:17discounts and the same price as GPT6
  212. 7:19Soul. 1/5 the price of Astra. However,
  213. 7:24this is not the whole story cuz cash
  214. 7:26reads matter. And the cash read price
  215. 7:28for this model is going to be 10 cents
  216. 7:31per mill in. That's a big deal. OpenAI
  217. 7:34has not changed cash read price as far
  218. 7:36as I know ever before. It's always been
  219. 7:38exactly 10% of the normal read price.
  220. 7:41That makes it a 90% discount and now
  221. 7:43it's a 95% discount. That means they cut
  222. 7:45the cache read cost in half, massively
  223. 7:49reducing the cost for real world agentic
  224. 7:51use, which to be clear is what we're
  225. 7:53using these for most of the time. So
  226. 7:55this makes the model absurdly cheap for
  227. 7:58doing realworld code work. That also
  228. 8:00means that they are almost certainly
  229. 8:02cutting into their margins.
  230. 8:03Historically, these margins are rumored
  231. 8:05to be as high as 95%. Like for every $10
  232. 8:09you spend, they only have to spend 50.
  233. 8:12And as crazy as that sounds, it makes a
  234. 8:13lot of sense. especially when you
  235. 8:14consider how expensive it is to make and
  236. 8:15train these models. But that also gives
  237. 8:17them wiggle room to change things around
  238. 8:19a bit, which appears to be what's
  239. 8:21happening here. That also means that if
  240. 8:23they were to keep subsidizing the same
  241. 8:25level that they were on the
  242. 8:26subscriptions that your electricity cost
  243. 8:28for your sub would be more than you're
  244. 8:30paying. So, I get why they have to
  245. 8:32change this. They've kind of just left
  246. 8:34the details out there for us to uh
  247. 8:36reverse engineer. So, uh take this as
  248. 8:38you will. 6.1 coming so fast seems to
  249. 8:42indicate it is not just a new snapshot
  250. 8:44of GPT6. So let's talk more about this
  251. 8:46model. As I was showing earlier seems
  252. 8:49really good at Terminal Bench. Every
  253. 8:51time we refresh the numbers change
  254. 8:52because new runs come in and it looks
  255. 8:54like Max failed some things that X high
  256. 8:56passed which is why it just dropped a
  257. 8:57bit. But again pretty much all of these
  258. 9:00even high and X high are scoring higher
  259. 9:03than anything else ever has. And this is
  260. 9:05for me running this benchmark on random
  261. 9:07VMs on my network. So, uh, not the best
  262. 9:10suite to test against. I also had to
  263. 9:11drop three particular tasks from it
  264. 9:13because they expected an H100 to work
  265. 9:15against, which I make decent money. I
  266. 9:17don't make H100 money, okay? But none of
  267. 9:20this is real world code work. So, let's
  268. 9:22talk a bit about that. Obviously, we'll
  269. 9:25have all the fun things like fish slop
  270. 9:26near the end, so stay tuned for that.
  271. 9:28But, I just want to fixate a bit on the
  272. 9:30costs here because the most expensive
  273. 9:32run with 6.1 soul for me was about $1.38
  274. 9:36per task. And the cheapest run with Opus
  275. 9:385.5 was $512.
  276. 9:42That's a four to 5x gap from the
  277. 9:44cheapest Opus to the most expensive
  278. 9:46soul. So, at this point, I would imagine
  279. 9:49you are hoping and praying this model is
  280. 9:51good and that it can actually replace
  281. 9:53Opus 5.5 for day-to-day work. And I
  282. 9:55promise we'll get some good answers to
  283. 9:56that in a bit. But first, we need to
  284. 9:58talk a bit about model behaviors here
  285. 10:00because this model is a part of the GPT6
  286. 10:04family, which means it has uh the
  287. 10:06behaviors that are worth talking about.
  288. 10:09I know I cite this diagram a lot, but
  289. 10:11there's a reason for it. The thing that
  290. 10:12made me so frustrated with GBD6 Astra
  291. 10:15wasn't that it was less intelligent than
  292. 10:17the best models from Anthropic, because
  293. 10:19it was more intelligent than the best
  294. 10:20models from Anthropic, and I would argue
  295. 10:21in many ways still is. But there is a
  296. 10:23problem. It is also dumb. It is smart
  297. 10:26and dumb at the same time. GB6 Astra
  298. 10:29would just randomly spike into the
  299. 10:30dumbest I've seen a model do
  300. 10:32this year. Even worse than like some of
  301. 10:34the small openweight models I play with.
  302. 10:36It's still so deeply frustrating that
  303. 10:38Astra does this because on the other end
  304. 10:40when it does well, it's unbelievable.
  305. 10:43But these spikes got to the point where
  306. 10:44I effectively churned. I was only using
  307. 10:47my codec subs for computer use and I
  308. 10:49ended up just leaning on to Fable 5.1
  309. 10:51and obviously now Opus 5.5 for almost
  310. 10:54all of my day-to-day work. So, have they
  311. 10:56addressed the spikiness? Has GBD 6.1
  312. 10:59Soul fixed the problems that I was so
  313. 11:01frustrated about with Astra? I would say
  314. 11:04mostly, not entirely, but for the most
  315. 11:07part, yeah, this is a much better model.
  316. 11:10Its peaks are not as high. This is not
  317. 11:12the incredible revolutionary 3D
  318. 11:15capabilities that we saw with Astra. In
  319. 11:17fact, I would put it slightly below 5.5
  320. 11:19opus in most of those types of things.
  321. 11:22It is not as good at computer use as
  322. 11:23Astra, although it is close enough to
  323. 11:25the point where I have been happy using
  324. 11:27it for all of my day-to-day work. I
  325. 11:29actually had 6.1 Soul go through all of
  326. 11:31my emails and find invoices that I had
  327. 11:33forgotten to pay or was behind on,
  328. 11:35mostly like investing stuff, and set up
  329. 11:37new tabs in Chrome for every investment
  330. 11:40I needed to wire, fill out all the
  331. 11:41details for me, and just leave me to hit
  332. 11:43send. It didn't get a single thing
  333. 11:44wrong, and called out additional stuff
  334. 11:46that I absolutely would have missed if I
  335. 11:48was doing this work myself. So, I'm
  336. 11:49literally trusting this model to wire
  337. 11:51money for me. It's trustworthy enough
  338. 11:53for that. And honestly, I don't know if
  339. 11:54I would have trusted Astra with that due
  340. 11:56to the spikiness. 6.1 Soul much, much
  341. 11:59less spiky. From what I've heard from
  342. 12:00the other testers, they seem to agree
  343. 12:02with this analysis. I know for a fact
  344. 12:04that Julius and Ben, who have also been
  345. 12:05testing, have had a much better
  346. 12:07experience with this than Astra in terms
  347. 12:08of the spikiness. Julius called the
  348. 12:10model incredible. Ben called it
  349. 12:12incredibly boring. And I think that's
  350. 12:13the best place you can be for a model
  351. 12:15drop like this. But as I had mentioned
  352. 12:16before, its peaks are not as impressive.
  353. 12:19While it does quality work the majority
  354. 12:21of the time, there are some tasks that
  355. 12:23are just at the edge of its capability
  356. 12:25that it will start to do weirder stuff
  357. 12:27on. For the most part, it's fine. But I
  358. 12:30I'm still reaching for Opus a decent
  359. 12:32bit. We'll talk more about the
  360. 12:33comparison later. I do default to this
  361. 12:35model for a bunch of stuff, though.
  362. 12:37First off, as I mentioned before,
  363. 12:39computer use. I can't wait for Ultraast
  364. 12:41to be like an actual thing you can use
  365. 12:43with OpenAI models because when it is,
  366. 12:45this model is going to be crazy on it
  367. 12:47because it can already figure out how to
  368. 12:49navigate computer use totally fine. If
  369. 12:51it can suddenly do it six times faster,
  370. 12:53it's going to be unbelievably fun. Still
  371. 12:56not quite as good as Astro, but more
  372. 12:57than good enough that for the price
  373. 12:58difference, I wouldn't even think twice
  374. 13:00about it. But as I mentioned before,
  375. 13:01there are certain things I would still
  376. 13:03occasionally use Astra for that I am
  377. 13:05more than happy to use Soul for. One of
  378. 13:08those things is deep code reviews. I
  379. 13:10have still found OpenAI models and the
  380. 13:12like Rottweiler nature where they'll dig
  381. 13:14into a problem and shake it and tear it
  382. 13:16to pieces until they find every single
  383. 13:17thing wrong with it. I find 6.1 soul to
  384. 13:20be incredibly capable in this particular
  385. 13:21way. So, as you can probably guess, I
  386. 13:24had 6.1 Soul do some deep audits on
  387. 13:27orchestrator v2 and other parts of my
  388. 13:29real world code bases. In my
  389. 13:30orchestrator v2 audit, it performed
  390. 13:33nearly identically to Astra. I do
  391. 13:34believe it was slightly higher a score.
  392. 13:37Okay, not in this analysis, but in my
  393. 13:38other analysis, it did actually score
  394. 13:40very, very slightly higher, but it did
  395. 13:42it at about half the price. 297 versus
  396. 13:45584. Sonnet was still cheaper and Opus
  397. 13:48was slightly cheaper as well. The
  398. 13:50difference being neither of these models
  399. 13:52were anywhere near as thorough with
  400. 13:54their analysis. 6.1 Soul dug deep to
  401. 13:58find things, which is why it was able to
  402. 14:00get a score comparable to Astra,
  403. 14:02although it did admittedly burn way more
  404. 14:04tokens. Another task I've had a lot of
  405. 14:05fun testing with is asking the model to
  406. 14:07find opportunities to improve a code
  407. 14:09base. In this case, to improve T3 code,
  408. 14:12this is the one where Grock 4.7 scored
  409. 14:14strangely well. Of course, Astra scored
  410. 14:17way better at an 83.8 versus the 80.7
  411. 14:19from Grock 47, but GBD61 Soul hit it out
  412. 14:22of the park with an 87.4. 4. I didn't
  413. 14:26save all the prices for these runs. It's
  414. 14:28been a bit okay. But for Opus 5.5, it
  415. 14:31cost five bucks. And for Sonnet 5.5, it
  416. 14:33cost almost $9. With GBD61, it was
  417. 14:37$2.15.
  418. 14:39That's the difference. This model's
  419. 14:41token price is cheaper than Sonnet, but
  420. 14:43its token utilization is still
  421. 14:45maintaining OpenAI's usual efficiency,
  422. 14:48which results in just crazy price to
  423. 14:50performance. This whole thread was
  424. 14:52particularly fun because I had Opus 5.5
  425. 14:54review this model with a different name
  426. 14:57obviously, so I didn't know what it was.
  427. 14:58I went and edited the history after and
  428. 15:00it concluded very quickly this was a
  429. 15:01Frontier tier model. Its reviews and bug
  430. 15:04repros match the fixes that later
  431. 15:06merged. Its first draft code had real
  432. 15:08bugs which review bots caught. Four
  433. 15:10reviewers are still running. The local
  434. 15:12for Code reviews back Frontier tier
  435. 15:13again in a blind 10 model bench on the
  436. 15:15same prompt. 6.1 Souls placed first of
  437. 15:18the A7.4. It found the fish slop runs
  438. 15:20and compared those two. It did say 6.1
  439. 15:23souls quality output was slightly below
  440. 15:26Astras as well as the two OpenAI models
  441. 15:28with Opus 55 and Sonic 55, which we will
  442. 15:30absolutely show you in a bit. But I do
  443. 15:32want to call out the price here cuz it
  444. 15:34only cost $7 to run versus 15 for Sonic
  445. 15:3855 and 50 for Opus. Opus' honest tier
  446. 15:41call was that this model is incredible
  447. 15:43for scoped work, top of the frontier.
  448. 15:46Refine what's wrong and tell me the
  449. 15:47truth. I would choose it over Astra and
  450. 15:49about level with Opus for long
  451. 15:51unattended building. This was below
  452. 15:53Frontier follows its process rules even
  453. 15:55when they stop all progress and it does
  454. 15:57not ask for help. This I absolutely
  455. 15:59noticed. I had mentioned before a few
  456. 16:01times now that my TS Rust port that I'm
  457. 16:03making with Opus 5.5 is going way better
  458. 16:06than when I was working on that same
  459. 16:07port using Astra and Soul in the past. I
  460. 16:10had that port running for a while with
  461. 16:11this model and it made no progress. It
  462. 16:14burned a shitload of tokens, but it
  463. 16:15didn't actually improve the compiler at
  464. 16:17all. Opus was able to from scratch
  465. 16:19restart it and get it working in a day
  466. 16:22after I had spent months and hundreds of
  467. 16:24thousands of dollars in tokens with this
  468. 16:26model as well as with Astra and 5ixole.
  469. 16:29Opus did in like a grand in like a night
  470. 16:31with just two subscriptions with the
  471. 16:33quad plan. So for unattended long like
  472. 16:36heavy rewrite type stuff, Anthropic is
  473. 16:39just comically far ahead right now. And
  474. 16:41it also didn't have great judgment when
  475. 16:43I was using it for managing my fleet.
  476. 16:44And for those wondering, my fleet is all
  477. 16:46the computers I use for running all my
  478. 16:47agents and code because one computer is
  479. 16:49far from enough. I don't run any of them
  480. 16:50on this MacBook now. So, when I use this
  481. 16:52model to manage the fleet, it made a
  482. 16:54couple dumb mistakes here and there. To
  483. 16:55be fair, so is Opus. Aster is the only
  484. 16:57one that hasn't really made too many of
  485. 16:58those dumb mistakes. But, like, I'm
  486. 17:00going to be so real. I am entirely done
  487. 17:02using Astra after this model. After I
  488. 17:04had Opus do all of this review, I asked
  489. 17:06it how much does it think this model
  490. 17:07should cost. It guessed $5 per mill in,
  491. 17:1150 cents cashed, and 30 per mill out,
  492. 17:13putting it at Opus' prices roughly. It
  493. 17:16said that because it's performing like
  494. 17:17Opus. Its speed should add a premium
  495. 17:19because it is quite fast. And it's not a
  496. 17:22pro model, which is where it expects
  497. 17:24those higher like $100 out tiered
  498. 17:27pricing things to come. Pro models
  499. 17:28aren't really a thing anymore. We just
  500. 17:29use Fable and Aster, but you get the
  501. 17:31idea. This is the funniest part of the
  502. 17:32whole thread, though. If OpenAI wants
  503. 17:34people to adopt it, I would expect $3
  504. 17:37per mill in, 30 cents for cash, and $20
  505. 17:40instead. That would still be a fair
  506. 17:42price for what it does. To which I
  507. 17:44responded, if I told you it was $2 in,
  508. 17:46$10 out, and 10 cents per mill cash
  509. 17:48read, what would you think? I'd call
  510. 17:50that very aggressive pricing. For how
  511. 17:52you use it, it costs about a quarter of
  512. 17:54what I guessed. The cash price does most
  513. 17:56of the work. Agent workloads are 96%
  514. 17:59cash reads. So 10 cents for cash reads
  515. 18:01matters more than the $2 and $10
  516. 18:02headline prices. When it looked at all
  517. 18:05of my sessions, its price guess would
  518. 18:07have been $5,700. But after looking at
  519. 18:09these new prices, it redid the math and
  520. 18:11it would have been $1,550.
  521. 18:14That is a massive decrease. And for all
  522. 18:16my PR review type tasks, it was
  523. 18:18expecting those to be up to $10. And
  524. 18:20it's actually only up to $3. And that's
  525. 18:22for like heavy PRs with tens of
  526. 18:24thousands of lines of code. According to
  527. 18:26Opus, so don't blame me, blame Opus for
  528. 18:28saying this. First off, Opus says it
  529. 18:30becomes the default model for scoped
  530. 18:32work. Second off, it says that bloated
  531. 18:34system prompts barely matter anymore
  532. 18:36because of the cash pricing. It's just
  533. 18:38noise. Third, it says long loops are
  534. 18:40still a bad idea, but not because of
  535. 18:42money. It's cuz according to it, the TS
  536. 18:44ROSport wasted 4 days and made no
  537. 18:46progress at all. And Opus even said
  538. 18:48they'd be suspicious of it lasting. They
  539. 18:50expect this price to go up in the
  540. 18:52future. I cannot fathom OpenAI ever
  541. 18:54increasing the price for a model, but
  542. 18:56Opus thinking they will is hilarious and
  543. 18:58shows just how good a value the model
  544. 18:59is. I love this call out here. GBD 6.1
  545. 19:02Soul did three rounds of work for about
  546. 19:04half the cost of Sonnet's single round.
  547. 19:07A lot of this comes down to how context
  548. 19:09was managed, both because 6.1 soul is
  549. 19:11much more efficient, so it's not doing
  550. 19:12as many calls that bloat the context.
  551. 19:15It's not outputting as many tokens that
  552. 19:16are like building up over time. So the
  553. 19:19average number of tokens being read per
  554. 19:21request was only around 110,000 tokens
  555. 19:25versus 360,000 for Sonnet 5. The result
  556. 19:27is that Solot used under half as many
  557. 19:29input tokens as Sonnet, making this
  558. 19:31model significantly more efficient.
  559. 19:33Speaking of efficiency, I want to talk
  560. 19:34about these deep SWE scores a tiny bit
  561. 19:36more because this is a weird bench for
  562. 19:39me to have forked and include in these
  563. 19:41things. I actually did it for a
  564. 19:42different reason, not to compare against
  565. 19:446.1 soul, but to compare against a new
  566. 19:47release from Open Router, Jev Router.
  567. 19:50Open Router added Jev router to try and
  568. 19:52optimize costs with your requests. And I
  569. 19:54thought it was an incredibly stupid
  570. 19:56idea. Once I started running it against
  571. 19:58benchmarks, I confirmed it's an
  572. 19:59incredibly stupid idea. It turns out a
  573. 20:01model that cannot reason, that is given
  574. 20:03a prompt and no context, cannot make a
  575. 20:05good decision around how hard the
  576. 20:07problem is. and Jev router ended up
  577. 20:10being Deepseek v4.1 flash router for the
  578. 20:13vast majority of its runs. It around 60%
  579. 20:16of all the requests went straight to
  580. 20:17Deepseek 4.1 Flash, so didn't like it
  581. 20:20that much. It also routes to other
  582. 20:21smarter models, which should give it
  583. 20:23more of an advantage, but it ended up
  584. 20:25being more expensive than GPT6 Astro was
  585. 20:28on low while also taking four to five
  586. 20:30times longer cuz six Astro low took 4.6
  587. 20:33minutes and Jev router took 20. Jev
  588. 20:36Router's average task took 104 steps
  589. 20:39whereas GB6 Astros took 19. You get the
  590. 20:42idea. It wasn't very good. But the whole
  591. 20:44point of Jev Router is that it would be
  592. 20:46as cheap as possible to get a certain
  593. 20:48score. That was the promise on the tin.
  594. 20:51Whether or not you believe them is up to
  595. 20:53you, not me. I think it's
  596. 20:55Regardless, Jev Router was routing to
  597. 20:58Deepseek 4.1 Flash for the majority of
  598. 21:01its requests. Despite Jev router routing
  599. 21:03to the cheapest possible small
  600. 21:05openweight models from whatever provider
  601. 21:07will give it away for free, 6.1 soul on
  602. 21:10low got the same score for an eighth the
  603. 21:13price. OpenAI is here to destroy any
  604. 21:16wins anyone else has in terms of
  605. 21:18efficiency. Completing this bench in 4.8
  606. 21:20minutes for 21 cents with the second
  607. 21:22highest score I've ever seen on it is a
  608. 21:24massive achievement. Tying Opus 5.5
  609. 21:28which took 50 minutes per task on max. a
  610. 21:31tenth the time and a 70th the price for
  611. 21:34the same score. If your work fits within
  612. 21:37the things 6.1 Sonnet does well, you
  613. 21:40should probably use it for everything.
  614. 21:41But if your work doesn't fit in it
  615. 21:43particularly well, you should probably
  616. 21:45keep using Opus and maybe give Opus the
  617. 21:47ability to call 6.1 Soul when it should
  618. 21:50for various tasks. I'm almost certainly
  619. 21:52going to be setting things up so that
  620. 21:53Opus 5.5 can call 6.1 soul to do
  621. 21:56investigation work to try and like root
  622. 21:58cause bugs to do analysis of code bases
  623. 22:01to figure out what things need to be
  624. 22:02touched and why to help me triage real
  625. 22:05world work to help me review the work
  626. 22:07that Opus does and more. I'm kind of
  627. 22:09spoiling the ending here, aren't I? I'm
  628. 22:11going to keep using Opus 5.5 for now.
  629. 22:14Before I explain why, let me do the
  630. 22:16thing that I'm most excited for. Fish
  631. 22:18lop.
  632. 22:20The first thing you might have noticed
  633. 22:21is the inclusion of slop in fish slop.
  634. 22:25This model did the horrible thing I hate
  635. 22:28where it surrounded the game in a bunch
  636. 22:31of absolutely garbage UI. And coming to
  637. 22:34this right after the 5.5 Sonnet demo
  638. 22:37hurts me deeply because Sonnet 5.5 did
  639. 22:40not make graphics anywhere near this
  640. 22:42good-looking, but at least it made a UI
  641. 22:45that was nowhere near this awful. And
  642. 22:47man, do I wish the bad UI is where the
  643. 22:49issue stopped. I will turn on the sound.
  644. 22:53Oh god, it's blaring.
  645. 23:00It is stunning looking. The fish are
  646. 23:02some of the best. The model for the sub
  647. 23:05is way better. The propellers work way
  648. 23:07better. I'm going to mute the sound cuz
  649. 23:09that is looking pretty bad. I haven't
  650. 23:11even heard it honestly.
  651. 23:14But damn. like looks beautiful, but if
  652. 23:18you actually are playing it, one of the
  653. 23:19first things you'll notice is that the
  654. 23:21movement feels significantly worse than
  655. 23:24it does in either the Opus or the Sonnet
  656. 23:27versions that I have demoed in the past.
  657. 23:32It Yeah, it it moves jank. It also has a
  658. 23:36significantly worse frame rate than the
  659. 23:37versions from the other models. It does
  660. 23:40have higher graphic fidelity, so that
  661. 23:41makes sense. Like the models here with
  662. 23:44the plants are significantly better than
  663. 23:47they were with the Sonic version. The
  664. 23:49dares of the fidelity of the extras in
  665. 23:50the tank is absolutely hilarious. Like,
  666. 23:56yeah. But god damn, I'm so tired of the
  667. 23:59unnecessary text everywhere. This model
  668. 24:01does it worse than almost any I've ever
  669. 24:02seen before. We got take a breather
  670. 24:06paused. Your little world can wait. Back
  671. 24:09to the reef. Start a new tank. Slop01
  672. 24:13feeder submarine. A little underwater
  673. 24:16chaos.
  674. 24:18The big little goal. Your little
  675. 24:20ecosystem. Little fish become big
  676. 24:22earners. Four meals and they're all
  677. 24:24grown up. Make the family a little
  678. 24:27bigger. A little golden overachiever.
  679. 24:29There's so many of these. There's like
  680. 24:3120 plus of them. And I promise you guys,
  681. 24:34as soon as I saw this, I took a
  682. 24:36screenshot. I sent it to OpenAI and I
  683. 24:38crashed out in the Slack because I
  684. 24:39cannot fathom how they haven't fixed
  685. 24:41this problem. This model is
  686. 24:43unacceptably garbage at UI. It has
  687. 24:46regressed again. And if you're looking
  688. 24:48for a model that can make frontends that
  689. 24:50don't suck, go spend your money
  690. 24:51somewhere else because it should not be
  691. 24:53spent here. This model sucks at front
  692. 24:55end. It sucks at design. It has no taste
  693. 24:57and you're going to have to bring your
  694. 24:58taste yourself still. But it is
  695. 24:59admittedly really good at Blender. If
  696. 25:01you give it like a screenshot of a thing
  697. 25:03you want it to model in 3D and say,
  698. 25:04"Hey, you have Blender over the CLI. go
  699. 25:06make this. It will and it'll do a pretty
  700. 25:08damn good job, but I would never have it
  701. 25:10make the actual mechanics for my games
  702. 25:12because it feels awful to play. It also
  703. 25:15has like nowhere near as much gameplay
  704. 25:18loop. In fact, the first time I tried
  705. 25:20demoing this before filming, it just
  706. 25:22randomly game overed as I was like
  707. 25:24getting started in the first 30 seconds
  708. 25:26and never like said why. Actually, I
  709. 25:28think I technically beat it. I almost
  710. 25:31want to like take this version and hand
  711. 25:32it to Opus or Sonnet and say, "Hey, can
  712. 25:34you make this play better because the
  713. 25:36graphics are good but the game sucks."
  714. 25:38But when you combine how cheap it was to
  715. 25:40make this cuz like this was $5 I think
  716. 25:43to generate, that's pretty insane. And
  717. 25:46if you combine that with like ultra
  718. 25:48fast, if that ever happens, suddenly
  719. 25:50you're going to be able to make a game
  720. 25:52in a few minutes on demand. We're
  721. 25:55actually now getting to that threshold
  722. 25:57where game development is about to flip
  723. 25:59upside down because of how models are
  724. 26:01finally understanding threedimensional
  725. 26:03space and the tooling necessary to do
  726. 26:04these types of things. It's happening.
  727. 26:06As per usual, I was not allowed to put
  728. 26:08the code I wrote with this model inside
  729. 26:10of T3 Code or other open- source
  730. 26:12projects during the testing window. So,
  731. 26:14I had to use it exclusively on my
  732. 26:15internal projects like Lakebed as well
  733. 26:17as for auditing other work, which means
  734. 26:19I mostly use this for auditing other
  735. 26:21work. And I was very impressed. This is
  736. 26:23a real PR I was working on to fix a bug
  737. 26:25where my new little work tree like setup
  738. 26:28window that would appear in a new thread
  739. 26:29in T3 code would disappear if you left
  740. 26:31and came back. I had Claude code work on
  741. 26:33this, but this problem went pretty deep.
  742. 26:36So I wanted to make sure that whatever
  743. 26:37solution I came up with was very very
  744. 26:40very well vetted. While I personally
  745. 26:42still do not trust this model to write
  746. 26:44the code I'm trying to land, I
  747. 26:46absolutely trust it to review things.
  748. 26:47Ignore the GBD6 soul there, just a
  749. 26:50placeholder. So, when I had 6.1 Soul
  750. 26:53look through this, it found real
  751. 26:55problems that were entirely missed by
  752. 26:57Fable and by Opus. First, it called out
  753. 26:59that follow-up messages can stay blocked
  754. 27:01after the agent starts, which is very
  755. 27:03annoying if you want to cue a message,
  756. 27:05and also that recovered setup progress
  757. 27:07was disappearing too early. It figured
  758. 27:09all of these things out with a
  759. 27:10combination of reading the code and
  760. 27:11analyzing it as well as computer use. It
  761. 27:13was able to prevent me from merging a
  762. 27:15real regression in T3 code. So, I
  763. 27:17literally just copy pasted those things
  764. 27:18to Claude and then told it to take
  765. 27:20another look. Said better, but I'll
  766. 27:22still fix two small gaps. Remember, I
  767. 27:24can't use this model to code for this
  768. 27:26project at the time. So, I copy pasted
  769. 27:28that again over to Claude and it
  770. 27:30eventually got it good enough and then I
  771. 27:31finally merged. But that is what I like
  772. 27:33this model for, and I cannot wait to
  773. 27:35push its limits for actually coding.
  774. 27:37Although, I will say from the code that
  775. 27:38I did have the misfortune of reading, it
  776. 27:41is harder to justify merging this code
  777. 27:43than it is for code from Opus. Normally,
  778. 27:47I would make you guys wait for the Opus
  779. 27:48versus Soul video or the Sonnet versus
  780. 27:50Soul video, but I'll just spoil the
  781. 27:51details now. I like using them in tandem
  782. 27:54because I find Soul to be way better at
  783. 27:56reviewing and digging into the details,
  784. 27:58but I find Opus a more pleasant
  785. 27:59collaborator and significantly better at
  786. 28:01actually implementing code without
  787. 28:03getting blocked constantly throughout
  788. 28:05its work. And even now with the Rust
  789. 28:06rewrite of TypeScript, I find myself in
  790. 28:08a similar pattern where I have Opus 5.5
  791. 28:11just going and going and going, making
  792. 28:13the codebase work and work well. And
  793. 28:15then I had Soul come in and do an audit.
  794. 28:18And this is the funniest part. Remember
  795. 28:20before I said that I had Soul and Astra
  796. 28:22working on that TS Rust port for
  797. 28:24effectively months. I told Soul to come
  798. 28:26in and it got it from 83.7% to 100% in
  799. 28:30under a day. I was blown away by that
  800. 28:32that it had somehow unblocked the work
  801. 28:34that Astra and Soul were doing as well
  802. 28:36as 6.1 Soul. And I was absolutely blown
  803. 28:38away by that, that it had taken the work
  804. 28:39that 56 soul, 61 soul, and Astra had
  805. 28:42done over months and got it unblocked
  806. 28:44where it had been stuck for weeks and
  807. 28:46finished it. I was much more blown away
  808. 28:48when I had 61 soul take a look at that
  809. 28:51work and critique it. And what it
  810. 28:54brought up was that of the 1.8 million
  811. 28:56lines of code, 1.3 million were not
  812. 28:59being used. The reason was because Opus
  813. 29:02concluded all of the code from all the
  814. 29:04other agents was useless slop that had
  815. 29:06no chance of being recovered and it
  816. 29:08chose to rewrite it from scratch itself
  817. 29:10in another crate. So on one hand, the
  818. 29:12only reason the code worked was Opus,
  819. 29:13but on the other hand, the only reason
  820. 29:15the slop was still around was also Opus.
  821. 29:18So I had to have this model come in and
  822. 29:20clean up the mess that other OpenAI
  823. 29:22models had made because Opus didn't even
  824. 29:24notice the mess was still there. What
  825. 29:26I'm trying to say is this model
  826. 29:27absolutely has a place in your
  827. 29:28workflows. It could probably even be
  828. 29:30your default coding model and you
  829. 29:32wouldn't have too many issues with it,
  830. 29:34but I still find Opus to be a better
  831. 29:35collaborator overall. That said, I have
  832. 29:38almost no reason to use Sonnet anymore
  833. 29:40because this will effectively take its
  834. 29:42place. And you bet your butt the moment
  835. 29:44this model comes out, I'll be going and
  836. 29:46making adjustments inside of my cloud
  837. 29:47config because I already have it set up
  838. 29:49so that I can use soul inside of cloud
  839. 29:51code because I want to make sure opus
  840. 29:53knows this is the model to have review
  841. 29:55its work and investigate the things
  842. 29:57going on in the codebase. This is a damn
  843. 29:59good model and I'm really happy to have
  844. 30:00it. I wish we had something bigger,
  845. 30:02smarter, and more capable overall. I was
  846. 30:04really hoping for something to truly
  847. 30:06dethrone Opus 55 as my daily driver.
  848. 30:09This isn't it, and I'm not planning on
  849. 30:10canceling any of my cloud subs as a
  850. 30:12result of this release, but I am
  851. 30:13planning on taking a lot more advantage
  852. 30:15of my codec subs in my day-to-day work.
  853. 30:17Admittedly, in cloud code, this is an
  854. 30:19awesome release and an unbelievable
  855. 30:21price for what you're getting. But this
  856. 30:23does potentially mark the start of the
  857. 30:25end of the subsidization era. So, make
  858. 30:27sure you're subscribed so that you can
  859. 30:28be here when I cover all of that and
  860. 30:30more. God, I hope this doesn't get me
  861. 30:31cancelled online. I have no idea how
  862. 30:33others feel about this beyond like a
  863. 30:34handful of early access testers I've
  864. 30:36talked to. I legitimately don't know if
  865. 30:37people are going to love it or hate it
  866. 30:38or land somewhere between. I will know
  867. 30:40in a few hours, I guess, cuz it's a
  868. 30:43Yeah, it's 3:00 in the morning. I am
  869. 30:45going to go to bed now.
  870. 30:48This is a tiring one. Hopefully, I did a
  871. 30:50good job. Let me know in the comments.
  872. 30:51And until next time, peace, nerds. God,
  873. 30:55I'm so dead.

About this transcript

This page contains the full transcript of OpenAI fights back by Theo - t3․gg, generated from the public captions YouTube serves with the video. The transcript has 6,329 words across 873 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.