YouTube2Text

Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators) — Transcript

by SemiAnalysis · 9,614 words · 1,428 segments · language en · Watch on YouTube

Full transcript

  1. 0:05Hello everyone. Welcome back to Semi
  2. 0:07analysis weekly. Uh today I'm joined by
  3. 0:09Brian and Myron. We're going to talk
  4. 0:11about the new article we put out on
  5. 0:13OpenAI Jalapeno veteran Blackwell their
  6. 0:16self-designed ASIC which we compared
  7. 0:18with Ruben. We analyzed the TCO. uh we
  8. 0:22assessed the throughput per megawatt
  9. 0:24claims a little bit more of the details
  10. 0:26in terms of the micro architecture
  11. 0:27system architecture how they used AI to
  12. 0:30design it and write kernels and why the
  13. 0:33speed uh from tape out to actually like
  14. 0:37first working system with first real
  15. 0:39benchmarks run on it has been so
  16. 0:40impressive. So guys excited to dig in.
  17. 0:43This is a fun one.
  18. 0:46Yeah,
  19. 0:49[laughter]
  20. 0:50>> you could say something in response
  21. 0:52there, Brian.
  22. 0:55>> Sorry, I
  23. 0:58>> Is this a Is this a
  24. 1:00>> Is this a virtual background that you
  25. 1:02got like pink cotton candy background
  26. 1:04going on or what?
  27. 1:05>> Of course. Of course. I think it's the
  28. 1:08most similar color to what uh I think
  29. 1:11Open AI's belonged. I don't remember off
  30. 1:13the top of my head, but they're good.
  31. 1:14They got this like nice designs for
  32. 1:16their blocks. Yeah. But yeah, but
  33. 1:18anyway, Jalapeno has been great. Uh
  34. 1:20quite a surprise on the last day of hot
  35. 1:22chips for those attending.
  36. 1:25Uh one of the most surprising talks in
  37. 1:28my opinion. Uh and yeah, I believe it
  38. 1:31caught the performance caught everyone
  39. 1:33by surprise. We all knew that OpenAI
  40. 1:37chip was in the works and they have
  41. 1:38announced uh deals with Broadcom but I
  42. 1:42first look at performance results and we
  43. 1:45yeah and it's really looking quite good
  44. 1:47for OpenAI
  45. 1:49a brief summary for those that didn't
  46. 1:52read the article which is kind of
  47. 1:54unlikely is that on per megawatt basis
  48. 1:57and perf per TCO basis. So in other
  49. 2:00words, uh how much it cost to actually
  50. 2:03run the chip. Uh actually Jalapeno
  51. 2:07actually beats Vera Rubin's July results
  52. 2:11and we are comparing against July
  53. 2:14results as we believe it's kind of fair
  54. 2:16because the software is in like a
  55. 2:19similar state uh between these two
  56. 2:21points. Yeah. So as you shown as on the
  57. 2:25picture, this is Jalapeno compared
  58. 2:27against GB300. So JP300 it beats out of
  59. 2:31the water completely. Uh but as we
  60. 2:33mentioned during our in our article it's
  61. 2:36a bit unfair to compare it against
  62. 2:39Blackwell because Blackwell uses HBM3
  63. 2:42while uh Jalapeno uses HBM4. So actually
  64. 2:46my can talk about this a bit later on
  65. 2:49the differences between the HBM but it's
  66. 2:53a bit unfair to compare HBM 4 against
  67. 2:55HBM3. So we're comparing it against
  68. 2:57Rubin instead.
  69. 2:59And yeah, as in the results show, he
  70. 3:02does beat Rubin on output token
  71. 3:05throughput per oil utility megawatt. But
  72. 3:08then again, these are July results. And
  73. 3:11where Rubin right now is likely uh much
  74. 3:14better than July. But Jalapeno of course
  75. 3:17will improve as time goes on. As we
  76. 3:19shown in another diagram below Jordan, I
  77. 3:22think
  78. 3:23>> uh we show how jalapeno has improved
  79. 3:25from
  80. 3:27uh 25 days.
  81. 3:29>> 25 days.
  82. 3:32>> Yeah, I'll go grab that one. But um
  83. 3:34maybe like the
  84. 3:36>> the key point you were making if you can
  85. 3:38explain a little bit more about
  86. 3:40>> the benchmarks that they're running
  87. 3:42which is Deepse CR1 8K1K.
  88. 3:46um not really the like absolute biggest
  89. 3:48model and not the most demanding
  90. 3:51inference workload because it's uh just
  91. 3:53random data. It's not agenic workflows.
  92. 3:56Um but really quickly they got this up.
  93. 4:01Clearly the performance is strong as
  94. 4:03you're yes implying with this new chart.
  95. 4:07um performance is like improving by the
  96. 4:10week at this point even by the day and
  97. 4:13um maybe you can you can also explain
  98. 4:16the denominator there like why they're
  99. 4:18choosing to divide by the power
  100. 4:20consumption.
  101. 4:22Yeah, actually uh Jensen brought up this
  102. 4:26point in Computex 2026
  103. 4:29during his keynote that uh data centers
  104. 4:32nowadays are getting power limited and
  105. 4:34that power is starting becoming the
  106. 4:36constraint. If you have money, you can
  107. 4:38always get more servers, more chips, but
  108. 4:40the the constraint is starting to become
  109. 4:43the data centers power. And like
  110. 4:46companies have tried getting through
  111. 4:48this using a BTM or behind the meter
  112. 4:51power but end of the day the power tends
  113. 4:54to be your your constraint for building
  114. 4:56a data center.
  115. 4:59Yeah, absolutely. So like if you have
  116. 5:01100 megawatts of power you can only fit
  117. 5:04so many chips in there. It doesn't
  118. 5:06really matter if these chips are more
  119. 5:09expensive
  120. 5:11uh on a per megawatt basis. if they're
  121. 5:14producing more tokens per megawatt, you
  122. 5:16think you can make more money off of the
  123. 5:18tokens they produce and you justify the
  124. 5:19extra expense on the actual chips.
  125. 5:22And
  126. 5:24while it's a valid metric for some
  127. 5:27customers, it may not be valid for
  128. 5:29others. And so it's it's not the most
  129. 5:31typical way. Most typically we see just
  130. 5:35token throughput per GPU compared like
  131. 5:38per package. Um, but in this case,
  132. 5:42OpenAI doesn't have customers for
  133. 5:44Jalapeno. They just run it for
  134. 5:45themselves and so they don't really care
  135. 5:49tokens per package. They care about
  136. 5:51tokens per like megawatt that's going
  137. 5:54into the system, right? Yeah. And
  138. 5:57actually going off of this uh topic uh
  139. 6:01someone did mention I forgot one of us
  140. 6:03mentioned that uh like performance per
  141. 6:06chip is at the end of the day is just a
  142. 6:08imaginary thing. You can just glue two
  143. 6:11chips together and call it you double
  144. 6:13your performance per chip. So looking at
  145. 6:15per chip is not really a wrong metric.
  146. 6:18just it can be easily like misled cuz
  147. 6:21you can just which is what Ruben Ultra
  148. 6:23and Blackware Ultra does if I'm not
  149. 6:26wrong. You just glue two chips together
  150. 6:29and bing bang boom, you got double your
  151. 6:32true per chip. Yeah.
  152. 6:33>> Yeah. I mean, you can do this if you're
  153. 6:35cerebrous, too, right? You can just put
  154. 6:37three of the wafers in a rack and then
  155. 6:38you can put them all on a chart where
  156. 6:40three is better than one, right?
  157. 6:43>> Yeah. Or or Yeah. Just say your chip is
  158. 6:45the whole wafer, right? Um yeah
  159. 6:49um yeah but I mean uh it it doesn't mean
  160. 6:52that um they don't care about
  161. 6:54performance per per cost right because
  162. 6:57I'd say that um you know performance per
  163. 7:00what and performance cost they're pretty
  164. 7:02closely closely linked right I mean I
  165. 7:05think um the more power the chip
  166. 7:08consumes it probably means the chip is
  167. 7:11more expensive as well generally and I
  168. 7:14think we see that um yeah we we did our
  169. 7:17um you know TCO analysis of what we
  170. 7:19think open would pay for a jalapeno
  171. 7:23system and you know they still pretty
  172. 7:25much win or or very close on uh
  173. 7:28performance for TCO as well, right?
  174. 7:32>> Yeah. Here's that chart on screen. I
  175. 7:34mean, this one's just more interesting
  176. 7:36because obviously the TCO calculation in
  177. 7:39terms of like how many tokens you're
  178. 7:41going to get per dollar
  179. 7:43um depends on that input for per dollar.
  180. 7:47And so we've got some variables on
  181. 7:50screen there about, you know, what's the
  182. 7:52total cost per hour to own, for example,
  183. 7:56a GV300 or to own Jalapeno. And we're
  184. 7:59assuming that OpenAI is going to pay 279
  185. 8:03an hour for a GB300.
  186. 8:05They're buying them themselves. They're
  187. 8:06running the data centers themselves, for
  188. 8:08example, in that case. Um they're not
  189. 8:10playing paying the current going
  190. 8:12NeoCloud prices, which are like six
  191. 8:14bucks an hour for GB300 right now. And
  192. 8:17put Vera Rubin at um 361 and then
  193. 8:21Jalapeno at $1.56. And so let's just
  194. 8:25compare it per package. Like clearly
  195. 8:28they're going to like the thesis of
  196. 8:30designing a chip in house is that you
  197. 8:31want to pay the Broadcom margins only,
  198. 8:34not the Nvidia margins or the Broadcom
  199. 8:38plus Google TPU margins or some
  200. 8:40equivalent there, right? Um and that's
  201. 8:44that's being borne out on this chart
  202. 8:46that you can see.
  203. 8:53>> All right. And I guess the the next part
  204. 8:56of this is like um
  205. 8:59where could like how could these curbs
  206. 9:02move? Um I think with jalapeno they're
  207. 9:06actually somewhat sandbagging
  208. 9:09these results significantly. Right.
  209. 9:14Uh yeah, I mean this is so the caveats I
  210. 9:18guess on the performance is that I I
  211. 9:21said it earlier that um the tests they
  212. 9:24ran were deepseek 1 uh that's based on
  213. 9:27the V3 architecture that came out in
  214. 9:29January of 2025. So it's not the most
  215. 9:30current DeepC4 but it is a relatively
  216. 9:33big model 600 billion total parameters.
  217. 9:37They ran Kimmy K 2.5 which is 1.5
  218. 9:40trillion. So that's a actual big model
  219. 9:43or just 1 trillion, sorry, not 1.5. And
  220. 9:46then they ran GPTOSS 12B, their own open
  221. 9:49source model, relatively small. Uh they
  222. 9:53had solid performance on all three. Um
  223. 9:55but
  224. 9:58I mean solid performance. They're
  225. 9:59beating Farah Rubin [laughter]
  226. 10:01on uh on all three today. And they're
  227. 10:03beating them with single token
  228. 10:06prediction and no pre-filled decode
  229. 10:08disagregation. So
  230. 10:11Brian, maybe you can explain single
  231. 10:14versus multi-token prediction and
  232. 10:16specifically in these performance claims
  233. 10:18when compared directly to Bar Rubin,
  234. 10:20which is using MTP.
  235. 10:23Um, this is like OpenAI kind of
  236. 10:26fighting with one hand behind their
  237. 10:28back.
  238. 10:30>> Yep.
  239. 10:30>> Like MTP is a really big optimization
  240. 10:33>> for interactivity.
  241. 10:35>> Yeah, of course. Uh actually going back
  242. 10:39on uh the deepse uh one it reminded me
  243. 10:42of an X post I saw like some time ago
  244. 10:46that someone said the closed lab open AI
  245. 10:49and entropic probably have an internal
  246. 10:51version of MLA and they probably
  247. 10:53discovered
  248. 10:54uh MLA like long before deep but it's
  249. 10:59quite interesting that uh we were told
  250. 11:02that the MLA kernels like openi didn't
  251. 11:05have any internal MLA kernel uh
  252. 11:07implementation. So I'm not saying like
  253. 11:09the open source tritons one they don't
  254. 11:11have any internet optimized MLA kernels
  255. 11:14which is quite interesting because that
  256. 11:16well that means none of the open AI
  257. 11:18models uses MLA which is quite
  258. 11:21interesting to me and uh yeah and on the
  259. 11:25AK1K and Deepc R1 like model choice we
  260. 11:30did release an AX uh benchmark case but
  261. 11:35the timing was very bad so I think we
  262. 11:37didn't get uh the OpenAI Jalapeno team
  263. 11:39to actually run it on Ethernet X. But
  264. 11:42yeah, as Jordan says, uh measures stuff
  265. 11:45like prefix cache which uh exposes
  266. 11:49a lot more areas for optimization. And
  267. 11:53in our we describe like a lot of stuff
  268. 11:55there's loadbearing
  269. 11:57that is being tested for Agent X and
  270. 11:59it's starting to show like what break
  271. 12:01what's breaking what's not and areas for
  272. 12:04improvement that AK1K didn't properly
  273. 12:07show.
  274. 12:08And uh and yeah going so going back to
  275. 12:10the MTP question uh for those unaware
  276. 12:13MTP stands for multi token prediction
  277. 12:14and it's a form of speculative decoding
  278. 12:17where you kind of guess the next future
  279. 12:20tokens and then verify them in a single
  280. 12:22forward pass because uh interestingly
  281. 12:25the LM doesn't output just the next
  282. 12:28token's probability but the token
  283. 12:31probability of every single token
  284. 12:34position before the last token.
  285. 12:37And we were not given an MTP results.
  286. 12:39And it's my personal guess that OpenAI
  287. 12:42has some internal speculative decoding
  288. 12:44technique that's not MTP and not DSpark
  289. 12:47or any open source uh configs. So they
  290. 12:52didn't give us speculative decoding
  291. 12:54results because
  292. 12:56there's no way to actually verify it
  293. 12:58through open source which is also quite
  294. 13:01interesting in my opinion means this
  295. 13:03spark isn't the best that we can do and
  296. 13:06this park actually gives quite a a crazy
  297. 13:09advantage over MTP by doing like uh
  298. 13:12guessing all the tokens at once uh
  299. 13:15instead of doing MTP which is just one
  300. 13:17layer of a model trying to
  301. 13:20uh guess future tokens one by one
  302. 13:23instead of doing them all at once.
  303. 13:25Yeah. And with Yeah. And we were told
  304. 13:28that the internal speculative decoding
  305. 13:30method gives like three to five times uh
  306. 13:33improvement on production models which
  307. 13:36yeah would just knock the jalapeno
  308. 13:39versus ferubin comparisons out of the
  309. 13:41water once again. So you can imagine
  310. 13:44that graph shifting like three to five
  311. 13:46times. Yeah, it's quite insane.
  312. 13:48So Kudo mode's gone, Brian.
  313. 13:52>> I mean, everyone on X is is saying, "Oh,
  314. 13:56it's just an ASIC. That's what you
  315. 13:57expected to do." And yeah, in some
  316. 13:59sense, that's what we that's what we
  317. 14:00expect an ASIC to do. But the software
  318. 14:03has come to such a point that like this
  319. 14:06line doesn't really matter if your
  320. 14:07software is good enough that you can get
  321. 14:09a motor up like very quickly. What's the
  322. 14:12difference between a general purpose or
  323. 14:15like they say GPGPU but N6 if software
  324. 14:19can just bridge this gap?
  325. 14:22>> Yeah.
  326. 14:23>> Yeah. Yeah. I mean this chip has I mean
  327. 14:27it's a toss up with Ver Rubin because we
  328. 14:28haven't seen stuff since they published
  329. 14:30that uh note in uh June or July. Um so
  330. 14:35maybe they've had a month extra to
  331. 14:37develop on the these chips. But I mean,
  332. 14:40conceptually, like we've never seen
  333. 14:43anybody else put out a chart where
  334. 14:46there's a curve showing they're beating
  335. 14:50Nvidia on every point of the curve
  336. 14:53and uh like a real test, right? So, it's
  337. 14:59I mean it's shocking that they did this.
  338. 15:01We can talk about the timeline a little
  339. 15:02bit later, but we we keep talking about
  340. 15:05the point in the curve and I think we we
  341. 15:06maybe haven't explained this in great
  342. 15:09detail or we've done it on previous
  343. 15:10podcast and people aren't, you know,
  344. 15:12familiar with this. So, the I'm going to
  345. 15:15put this chart back up on screen and try
  346. 15:17to explain a Fredo curve here for the
  347. 15:22purposes of understanding the
  348. 15:23performance claims made by Jalapeno
  349. 15:25here. And uh to do that we need to
  350. 15:28explain that the y- axis is how many
  351. 15:29tokens you can produce per megawatt
  352. 15:32you're putting into the system and the
  353. 15:33y- axis is how fast or sorry the x-axis
  354. 15:36is how fast the tokens appear to each
  355. 15:39individual user. And so whether you're
  356. 15:42optimizing for the y- axis or the x-axis
  357. 15:45jalapeno is beating the GB300 right now
  358. 15:48on the deepseek model which implies
  359. 15:50something like if you were to fix the
  360. 15:54interactivity per user. So everybody
  361. 15:56sees 100 tokens per second or 50 tokens
  362. 15:58per second. And I zoom in on like
  363. 16:01exactly that part of the curve. We're
  364. 16:03basically seeing that Jalapeno has
  365. 16:06double the amount of tokens that it can
  366. 16:09produce per megawatt implying
  367. 16:12two times more revenue, two times more
  368. 16:14profitability, whatever you want to say
  369. 16:16from an inference endpoint serving
  370. 16:18provider.
  371. 16:19And then if you look at the far right
  372. 16:21side of this curve across the x-axis and
  373. 16:24you see that at a very low batch size
  374. 16:26they can go all the way to 700 tokens
  375. 16:29per second per user. And you compare
  376. 16:32that to where the others peak out at
  377. 16:33350.
  378. 16:35This is once again
  379. 16:37double the performance.
  380. 16:39And
  381. 16:41I think the conclusion is that they're
  382. 16:44basically winning on both sides of the
  383. 16:49curve. So, both fast tokens and cheap
  384. 16:52tokens, which is
  385. 16:55it's just so interesting because we've
  386. 16:57seen so many other um companies
  387. 17:02make claims about how they're going to
  388. 17:04beat Nvidia and they just pick one of
  389. 17:06those,
  390. 17:09right? Croc or Cerebrus or any of the
  391. 17:11other startups that are going to focus
  392. 17:12on SRAMM like call it a Dmatrix that's
  393. 17:14coming up with stuff or Samonova where
  394. 17:16we've even seen some results on bench
  395. 17:19you know on the infertex benchmark or
  396. 17:20something similar to it they're saying
  397. 17:23they're going for fast tokens just
  398. 17:25decode speed low batch size they don't
  399. 17:27worry about throughput and then you've
  400. 17:29got other guys that are worried about
  401. 17:31throughput call it AMD just as a simple
  402. 17:33example but even TP or tranium could be
  403. 17:34in this bucket of just like accelerators
  404. 17:36that are going for throughput
  405. 17:38and they go, "Yeah, but we're not going
  406. 17:40to be able to compete with the other
  407. 17:41guys at high interactivity." And OpenAI
  408. 17:43has a chip that can do both for them.
  409. 17:45Really well. I mean, at a minimum, this
  410. 17:48thing is is doing really well right now
  411. 17:50and is going to serve real production
  412. 17:52tokens for them.
  413. 17:55>> Yeah. And yeah, it's actually quite
  414. 17:58surprising that I mean I was surprised
  415. 18:00that OpenAI was the first chip that
  416. 18:03actually
  417. 18:04uh like non Nvidia non AMD chip that
  418. 18:07actually appear on our public infrance
  419. 18:09X. We were like expecting some Nova or
  420. 18:12like Cerebras or even TPU trainium to be
  421. 18:15one of the first. But
  422. 18:18yeah, maybe this really come put into
  423. 18:19perspective like the difference in uh
  424. 18:22like how this time might be different
  425. 18:24from the rest like this the first chip
  426. 18:26that actually poses a real threat to the
  427. 18:28Kuda mode because
  428. 18:30like in open source we welcome uh
  429. 18:33results from anyone like you believe his
  430. 18:36chip is good run the results run the
  431. 18:39curves run the benchmarks show us what
  432. 18:42your chips does and we gladly put it on
  433. 18:44our dashboard and we gladly like write
  434. 18:46write an article about it if it's good
  435. 18:48but yeah and we have extended this offer
  436. 18:50to like edged very recently on X and of
  437. 18:53course etched uh didn't get back to us
  438. 18:56on that but yeah if your chip is good
  439. 19:00just run the benchmark show us the
  440. 19:02results and we let the results talk uh
  441. 19:05yeah so
  442. 19:06>> yeah it's it's
  443. 19:07>> we love to see more completion from the
  444. 19:09chips yeah exactly
  445. 19:11>> so did you guys see um Jensen's response
  446. 19:14to this I mean I level. I think the
  447. 19:17Cudamote has been uh getting drained
  448. 19:20slowly over the last couple of years
  449. 19:23like top models in the world like Claude
  450. 19:26and Gemini are trained without Nvidia
  451. 19:28GPUs to you know TPU right um Anthropic
  452. 19:33use lots of tranium it's not like you
  453. 19:35can only use GPUs but Nvidia is
  454. 19:38incredibly valuable company they're
  455. 19:39going to keep shipping all these GPUs
  456. 19:40and I think there's such a thing as like
  457. 19:42an Nvidia mode which includes everything
  458. 19:43in the supply chain the whole develop
  459. 19:45for ecosystem like all of the um
  460. 19:48availability to purchase and support and
  461. 19:50how you're going to like do the
  462. 19:52logistics of deploying these data
  463. 19:53centers and monitoring them over time. I
  464. 19:55mean, OpenAI's now got to go figure out
  465. 19:57how to turn on 100 megawatts of these
  466. 19:59chips, not just like three test racks,
  467. 20:01which is a monumental challenge to get
  468. 20:03over as if taping out a chip of this
  469. 20:07this uh you know quality is easy. Um
  470. 20:11which is not like the next phase will be
  471. 20:13pretty hard for them as well. Even
  472. 20:15Cerebrous is going through this
  473. 20:16themselves right now. So I guess the the
  474. 20:20question
  475. 20:21for me uh that well I think Jensen kind
  476. 20:27of answered in the style like I was
  477. 20:28saying there when he was on Mad Money
  478. 20:31with Jim Kramer, [laughter]
  479. 20:34everybody's favorite. Um, and he was
  480. 20:38basically like, I'm not bothered.
  481. 20:40And uh, yeah, like what's like Myron,
  482. 20:45what's your take on this? You think?
  483. 20:47Well, okay. What's your take on that?
  484. 20:49And just the whole like just the
  485. 20:50timeline to go from, yeah, we're tired
  486. 20:53of buying Nvidia GPUs for everything.
  487. 20:55We're going to go build it decision made
  488. 20:57at OpenAI to actually having a chip that
  489. 20:59can run in a sex benchmark. It's like
  490. 21:02under two years for concept to like real
  491. 21:05chip in lab and under nine months to
  492. 21:08actually get it taped out, right?
  493. 21:12>> Yeah. Uh yeah, I I have several thoughts
  494. 21:14on this. Um I guess starting from uh
  495. 21:19CUDA mo eroding um I think you know a
  496. 21:22lot of um how can you success both in
  497. 21:27terms of the silicon design
  498. 21:30as well as uh bringing up the software
  499. 21:32is it's it's been AI assisted right um
  500. 21:36and you know I guess somewhat the irony
  501. 21:38is that um this was all done on Nvidia
  502. 21:42GPUs training these models to bring up
  503. 21:44the capability to a point where um you
  504. 21:47know AI is able to you know program
  505. 21:50kernels um and that's been I think
  506. 21:53probably the big shift um in terms of um
  507. 21:57making it easier for the labs to and
  508. 22:01anyone to adopt um you know alternative
  509. 22:04systems for their their serving stack.
  510. 22:07Um I think you know one of the big
  511. 22:09reasons that uh Anthropic you know has
  512. 22:14decided to
  513. 22:16uh you know bring in AMD as as one of
  514. 22:20their hardware providers is because um
  515. 22:22you know Aentic programming as it allows
  516. 22:26them to sort of get around the the
  517. 22:28challenges uh of like using the AMD
  518. 22:31software stack for instance. So um you
  519. 22:33know somewhat ironically uh it's it's
  520. 22:37Nvidia's hardware has enabled um moving
  521. 22:40off of Nvidia's hardware. Um, and you
  522. 22:45know, in terms of I guess you jalapeno,
  523. 22:48it's it's such a big surprise because I
  524. 22:50mean we
  525. 22:52always knew that the team is capable,
  526. 22:55you know, they have they've had
  527. 22:57experience building um the other main
  528. 23:00successful um as program in the form of
  529. 23:04TPU. Um I think a lot of the yeah the
  530. 23:07the hardware team behind this is from
  531. 23:10their former TPU people as as we see in
  532. 23:12a lot of other um I guess whether it be
  533. 23:15AI accelerator startups or other A6 A6
  534. 23:18teams they tend to come from you know
  535. 23:20former people with TPU backgrounds um
  536. 23:24and um so for what it's worth sorry
  537. 23:27sorry sorry to interrupt but there's
  538. 23:29like basically no chip starters I can
  539. 23:32point to where it's a bunch of guys who
  540. 23:33are ex Nvidia
  541. 23:34But there's a lot of ex Google people
  542. 23:37out there doing stuff which is
  543. 23:38interesting.
  544. 23:39>> Yeah,
  545. 23:39>> I I do wonder sort of why why that is
  546. 23:42the case. But anyway, um another another
  547. 23:45topic. Um so uh and I think you know I
  548. 23:50mean designing an AI uh accelerator that
  549. 23:54is competitive with um Nvidia is is not
  550. 23:57easy, right? U it's such a huge market.
  551. 24:00Um, of course, everyone wants to try it,
  552. 24:02but um, you know, time and time again,
  553. 24:03we've seen uh, you know, as as as you
  554. 24:07guys have mentioned, we've seen a lot of
  555. 24:08entrance, but they haven't really been
  556. 24:10able to do it. Um, so, you know, we I
  557. 24:14think the expectation was that open
  558. 24:17would deliver a decent effort with their
  559. 24:19first generation. Um, and it turns out
  560. 24:21it was much better than decent. they,
  561. 24:23you know, as we said, they've already
  562. 24:24come out with something that's pretty
  563. 24:26much competitive or better than what the
  564. 24:28best of, uh, what Nvidia has to offer.
  565. 24:32So, that's surprise number one. Um, and
  566. 24:35then I think, you know, that also says
  567. 24:38something probably about um, other ASIC
  568. 24:41programs like especially, you know,
  569. 24:43Meta, Microsoft. Is it open air that's
  570. 24:46really good? Is, you know, are the
  571. 24:49silicon teams at Meta and Microsoft, do
  572. 24:50they have skill issues? It's probably a
  573. 24:52bit of both, right? I think um they're
  574. 24:55the guys that look the worst from from
  575. 24:57this announcement. Um but you know,
  576. 25:00going back to okay, where next? Um so
  577. 25:03designed a chip uh obviously you're
  578. 25:05scaling up the supply chain to um you
  579. 25:08deliver systems uh of mass like you
  580. 25:12talking about delivering like millions
  581. 25:14of these chips um you know thousands of
  582. 25:16racks um deploying them in data centers
  583. 25:19you know gigawatts of power um that's
  584. 25:22that's going to be not easy but also I
  585. 25:25think you know other people have have
  586. 25:27done it successfully I think the hardest
  587. 25:29part is really having that system design
  588. 25:31and then I think um you know OpenAI has
  589. 25:34partnered with um with people who have
  590. 25:38experience scaling this up right
  591. 25:40basically it's Broadcom and and
  592. 25:42Celestico on the system side uh and
  593. 25:44they've had experience you know doing
  594. 25:46this with TPU
  595. 25:49dig on the comment you made about the
  596. 25:51difference between like a in-house
  597. 25:53silicon program that's been going for
  598. 25:55years and years like MTIA at Meta or
  599. 25:57Maya at Microsoft if we just literally
  600. 26:00look at the specs of Jalapeno. Like it
  601. 26:04doesn't look
  602. 26:06super fancy on paper. I mean, it is an
  603. 26:09HBM4 chip, so it's going to have great
  604. 26:12HBM bandwidth.
  605. 26:14>> Um, it's got lots of
  606. 26:15>> FP4 flops, but like still
  607. 26:18>> less than
  608. 26:19>> Mhm.
  609. 26:20>> half of what Reuben's got on FP4.
  610. 26:24Less HPM capacity and less TDP. So
  611. 26:30I mean like under
  612. 26:32like almost three times less TDP per
  613. 26:35chip than Reuben, right? So if you just
  614. 26:39it's this weird thing where
  615. 26:41um when people are making the bull case
  616. 26:44for AMD, they go just look at MI450,
  617. 26:48right? It's going to have more FP4
  618. 26:51flops. It's going to have more HPM
  619. 26:52capacity. It's going to have more HPM
  620. 26:54bandwidth. It's going to be more TDP.
  621. 26:56and therefore it's going to be better.
  622. 26:59And they never want to look at the
  623. 27:00benchmarks. They never want to look at
  624. 27:01the like results of the chip, right? Or
  625. 27:05the MI355 when it was coming into
  626. 27:08production for the first time, right?
  627. 27:10But now you got an open AI chip which on
  628. 27:11papers is like objectively worse specs
  629. 27:14than everything on Reuben. Um comparable
  630. 27:19to GB300 on everything except for HPM
  631. 27:21bandwidth.
  632. 27:24uh and yet it's way outperforming GB300
  633. 27:27and
  634. 27:28outperforming Reuben so far. So like
  635. 27:31clearly this is due to the use of
  636. 27:35it's due to something about the micro
  637. 27:37architecture which we can get into and
  638. 27:40the software.
  639. 27:43So, like what's your take on what it
  640. 27:45takes to design a chip now? Because is
  641. 27:47it purely having access to the latest
  642. 27:49models and being willing to like
  643. 27:53yolo trust them on RTL and kernels? Like
  644. 27:58this is the barrier.
  645. 28:06Yeah, that's a good question and I I
  646. 28:09don't know the answer to that, but I
  647. 28:10think you know to your point, right? um
  648. 28:13you know everyone can deliver um great
  649. 28:18you know stacks on paper right you know
  650. 28:21when when we look at these stacks
  651. 28:22they're all you know pe theoretical um
  652. 28:26and I think a lot of emphasis on
  653. 28:29theoretical because I think that some of
  654. 28:31the flops numbers like no matter like
  655. 28:34how how you try to like reach them
  656. 28:37they're like impossible to reach so um
  657. 28:39it's it's somewhat determined by um you
  658. 28:42know the the chip company's market teams
  659. 28:45marketing teams but um yeah basically
  660. 28:48it's you know there are these flops that
  661. 28:50can you actually realize them in an
  662. 28:53actual workload same with HP you know
  663. 28:55bandwidth um I mean one of the big
  664. 28:59tenets of uh open's design philosophy
  665. 29:02forio is that
  666. 29:05um you know everyone can deliver raw hpn
  667. 29:09bandwidth you just like look you buy the
  668. 29:11the best HPM and then you put more
  669. 29:14stacks of it. Um, which I mean it's
  670. 29:17[snorts] it's buying HPM itself is is
  671. 29:20not easy these days, but um you know
  672. 29:22it's basically not super hard, right?
  673. 29:25It's not like it doesn't take um
  674. 29:27tremendous design skill. Um, so but
  675. 29:30really, you know, what's really
  676. 29:31throttling a lot of these chips is that
  677. 29:33they can't realize anything close to the
  678. 29:36raw HBM bandwidth because there's so
  679. 29:39many other uh things in the micro
  680. 29:41architecture that stop or stop you from
  681. 29:44doing that. And I think a lot of that is
  682. 29:45just um whether that be like very
  683. 29:48complicated memory subsystems or um
  684. 29:54or or just a lot of data movement
  685. 29:56required. Um whereas um the OKAI team
  686. 29:59has uh really been able has focused on
  687. 30:03um with mic microitecture that um
  688. 30:08reduces data movement so that um they
  689. 30:11can realize a lot of this you know the
  690. 30:13the the massive amount of page being
  691. 30:15damage they have. So I think that's
  692. 30:17really the main I guess skill um in all
  693. 30:20this and why for the competitors it's
  694. 30:24been really challenging to um you know
  695. 30:27they can deliver these all these all
  696. 30:29these like great specs but it's really
  697. 30:31difficult to actually realize them in an
  698. 30:34actual workload.
  699. 30:36>> Yeah. Yeah. And I agree with this like
  700. 30:38the fact that you can't actually reach
  701. 30:40this. I know there's there's a saying
  702. 30:42somewhere from someone I'm not sure who
  703. 30:46uh that the figures are like not numbers
  704. 30:49that you can reach and are numbers that
  705. 30:50the manufacturer can guarantee you never
  706. 30:53exceed like those flops which way to put
  707. 30:57it. Yeah.
  708. 30:58>> Yeah. Those are numbers that you are
  709. 30:59guaranteed not to exceed. And I think in
  710. 31:04one of the previous uh articles, I think
  711. 31:07in the cerebrous articles, they talk
  712. 31:09about the roofline models of chips
  713. 31:12compared to Nvidia. And although
  714. 31:13Nvidia's like flops are huge and like
  715. 31:16crazy especially the FT FP4 ones at the
  716. 31:20end of the day they are extremely uh
  717. 31:23right words on the roof lines in the
  718. 31:26computer bound region and a lot of
  719. 31:28workloads would rarely even reach those
  720. 31:31roof lines.
  721. 31:32So it doesn't matter at the end of the
  722. 31:34day the flops like the those flops in
  723. 31:36the table doesn't really matter because
  724. 31:39oops Jordan dipped again because the
  725. 31:42workloads most of the time does not
  726. 31:44actually hit those roof lines. Yeah,
  727. 31:48Jordan is gone. So maybe have some
  728. 31:50something curious to ask you about like
  729. 31:52the HPM. There's been talk about like
  730. 31:54Samsung HBM being better than the other
  731. 31:57HBM and like Halipino was
  732. 32:00uh lucky maybe to or like maybe it was a
  733. 32:03decision. I'm not so sure. I'm not a
  734. 32:05highway guy. Like what makes Samsung HBM
  735. 32:09like better or like uh better quality
  736. 32:12than the rest? Yeah. Yeah. So Samsung um
  737. 32:17for a long time you know from the HBM 3
  738. 32:20and 3 generation um Samsung's HBM you
  739. 32:23know was was really quite inferior to um
  740. 32:27Hinx which you know who dominated who
  741. 32:30you know still dominates HPM market
  742. 32:32share is you the leading supplier for
  743. 32:35Nvidia for instance um and you know
  744. 32:38really 3e that generation is really bad
  745. 32:40from Samsung um part of it was it was
  746. 32:43built um an inferior process, right? Um
  747. 32:48so the ite um
  748. 32:52Samsung, you know, they realized this
  749. 32:55and they really went all out on on their
  750. 32:58HBM core technology. Um so the DRAM dies
  751. 33:03are built on a more advanced 1C process
  752. 33:07whereas you know highix and micron are
  753. 33:10using 01B process the same which is the
  754. 33:14same process that HPN 3 is built on um
  755. 33:17they're also using there's a logic based
  756. 33:20in HBM cube that has the fi u and
  757. 33:25Samsung is built built this on a
  758. 33:28advanced logic process which is you know
  759. 33:30SF4 um Samsung Foundry 4 nanometer node
  760. 33:34um whereas Hinx is using 12 nanometer
  761. 33:38TSMC um and Micron is still using their
  762. 33:42own you know DAM process um for this
  763. 33:45base D even though you know the the
  764. 33:47bandwidth um requirements are
  765. 33:49significantly higher for HKM4 and that's
  766. 33:51sort of been um that's that's why you
  767. 33:54know Micron has had some issues with um
  768. 33:58achieving the this, you know, the
  769. 34:01highest speeds for HP4 and and similar
  770. 34:04with with Clinics, they've had some
  771. 34:06issues. um they've had to say redesign
  772. 34:09the base die and this is why um you know
  773. 34:12it it was only until like they've had to
  774. 34:15like delay shipments for um HPF core for
  775. 34:19for Nvidia's uh you know rub right um so
  776. 34:23Samsung has ended up being you know
  777. 34:25because of these you know I think it it
  778. 34:27seems to be clear that Samsung has
  779. 34:29actually the best technology for HM4
  780. 34:33um this is why you know this um Jalapeno
  781. 34:38is HPM4. It can deliver 15.4 terabytes a
  782. 34:41second of you know HPN bandwidth um
  783. 34:44which means the 10 GB per second uh pins
  784. 34:47speed the HPM4 they've gotten um this is
  785. 34:50a little bit higher than say the 9.6 six
  786. 34:52that we think basically your uh Nvidia
  787. 34:56will will ship Ruben with. Um and yeah,
  788. 34:59I think it probably does end up being
  789. 35:02that it's because of Samsung being um
  790. 35:05the supplier here. Um and you know the
  791. 35:07reason is also that whether it's it's
  792. 35:10luck or skill I think um we can debate
  793. 35:14about it but you know guess I guess
  794. 35:17traditionally um Broadcom's HBN has
  795. 35:21mostly come from Samsung right so I
  796. 35:24think it it's partly sort of that luck
  797. 35:27um this is this has hurt Broadcom for 3
  798. 35:30but um you know for four this is been I
  799. 35:34guess that's turned out pretty
  800. 35:36for that.
  801. 35:38>> Yeah. Yeah, that's very interesting. And
  802. 35:40yeah, this this story of like Samsung
  803. 35:41HBM and like uh because if I'm not
  804. 35:44wrong, SKH Highix invented HBM, right?
  805. 35:48>> Yeah. So SKH Highix um along with AMD
  806. 35:52were sort of they
  807. 35:54realized that um very correctly that um
  808. 35:59you know
  809. 36:01memory bandwidth is doesn't isn't
  810. 36:03scaling um with along with uh you know
  811. 36:07logic performance. So they needed
  812. 36:09something to really like bring out
  813. 36:11bandwidth and they came up with this HBM
  814. 36:13concept and originally it was for for
  815. 36:17gaming GPUs. So the first product with
  816. 36:21HPM was on a gaming GPU um with on on
  817. 36:26one of the uh AMD gaming GPUs. And but
  818. 36:30this this turned out to be a bit bit
  819. 36:32overkill but uh thankfully um you know
  820. 36:36they had this technology um for for
  821. 36:38another very memory bandwidth intensive
  822. 36:40application which is AI.
  823. 36:43>> Yeah. Yeah. It's interesting that like
  824. 36:46SK Hanx and AMD invented HBM, but now
  825. 36:49like I mean Samsung is doing better in
  826. 36:52HBM 4 and Nvidia is uh doing better than
  827. 36:56AMD. But on the topic of the HBM
  828. 36:59bandwave, so bandwidth is a problem,
  829. 37:02right? And is not capacity. Is that why
  830. 37:04like companies are going towards like if
  831. 37:07I'm not wrong six high and four high
  832. 37:09instead of a high?
  833. 37:12Um so
  834. 37:14I'd say yeah the the primary
  835. 37:18I mean it's in the name high bandwidth
  836. 37:20memory the the appeal is is the
  837. 37:22bandwidth um
  838. 37:25capacity is important but um I think the
  839. 37:28main sort of benefit is really bandwidth
  840. 37:30because you know there there are cases
  841. 37:33where you pay for the additional
  842. 37:34capacity but you don't um you don't you
  843. 37:37might not need it um whereas I think for
  844. 37:40memory bands um you can always just
  845. 37:42reuse that into serving tokens faster,
  846. 37:45right? Um so um [snorts] and bandwidth
  847. 37:47is really the key and the thing is like
  848. 37:50whether it's a 8 high, 12 high or four
  849. 37:52high stack um the bandwidth is the same
  850. 37:55but um because you pay for capacity and
  851. 37:59and that makes sense because you know
  852. 38:00more capacity more layers is what um
  853. 38:03adds to the cost of the supplier. um you
  854. 38:06know the sort of you pay for the extra
  855. 38:09capacity but the dollar per bandwidth um
  856. 38:12gets much worse. So um for some some
  857. 38:15some companies if they want to optimize
  858. 38:18um and say actually we just want to like
  859. 38:20get best dollar per bandwidth then going
  860. 38:22lower stacks is is the right
  861. 38:23optimization. Of course you want that
  862. 38:26balance between capacity and and
  863. 38:27bandwidth but uh you know I think there
  864. 38:31is a a philosophy especially now that
  865. 38:34HTM is getting much more expensive
  866. 38:36because um you know we have very limited
  867. 38:39um supply of HPM wafers. Um so I think
  868. 38:43that the trade-off is starting to look
  869. 38:46more in favor of going to lower stack
  870. 38:49heights rather than just increasing them
  871. 38:50further and further.
  872. 38:52>> Yeah, it's interesting. Yeah, it makes
  873. 38:54sense with me especially like uh
  874. 38:57increasingly more Rex scale
  875. 38:58architectures
  876. 39:00>> but yeah like capacity is not becoming
  877. 39:02as much of an issue anymore.
  878. 39:04>> Yeah, exactly.
  879. 39:06>> Yeah. And this is um I think
  880. 39:08specifically borne out in the per watt
  881. 39:11argument as well where um if you look at
  882. 39:14the raw HPM bandwidth comparing the
  883. 39:16specs of the chips, it's like okay, it's
  884. 39:18up there with the other ones, but then
  885. 39:20if you divide by the amount of uh power
  886. 39:23consumed by the chip, the fact that
  887. 39:25they're getting 15.4 turbines per second
  888. 39:27on a chip with a TDP of 700 watts is is
  889. 39:30incredible. Like this bandwidth per watt
  890. 39:33ratio like arbitrary units here of 22 is
  891. 39:36just so much bigger than anything else.
  892. 39:38It's literally double Reuben. Even at
  893. 39:40the max Q like low power setting option
  894. 39:45on Reuben to like run it at 1,800 watts,
  895. 39:48that's still like it's got a little bit
  896. 39:51more um HPM bandwidth, you know,
  897. 39:54whatever that is 25% more 20 versus 15
  898. 39:57terabytes per second, but it's literally
  899. 40:00more than double the power. So
  900. 40:04I mean that's that's where your if you
  901. 40:06can realize that bandwidth with that
  902. 40:08much power that's where your token
  903. 40:11output per watt advantage comes from
  904. 40:14right there. Um I guess the the like big
  905. 40:19question of course my is like
  906. 40:22B 0 is in the fab right now. They should
  907. 40:26just be able to step up the power and
  908. 40:28get even more bandwidth from this,
  909. 40:29right? Or are they at some limit?
  910. 40:34I think um I think on the HVM bandwidth
  911. 40:37they are probably at um limit. I think
  912. 40:42the
  913. 40:44uh and you know yeah I think that's sort
  914. 40:48of just the constraint of the memory
  915. 40:49itself. Um but the B 0 um it should
  916. 40:53deliver more flops um at at you know the
  917. 40:56same power basic. So um that you know I
  918. 41:00think depending on the workload um if
  919. 41:02there are compute constraints then then
  920. 41:05that should benefit um for the B zero
  921. 41:08stepping uh the A zero stepping is
  922. 41:11what's uh what the I guess all the
  923. 41:13current um results are.
  924. 41:17Yeah. Yeah. There are many different
  925. 41:21we're seeing so many chip startups
  926. 41:22explore the surface area of possible
  927. 41:25uh configurations of chips right now.
  928. 41:27But if you are optimizing for HPM
  929. 41:29bandwidth per watt, a metric that seems
  930. 41:32pretty relevant in
  931. 41:35LLM inference, this is the design to go
  932. 41:38with at this point, right? There's
  933. 41:40nothing else that compares
  934. 41:42that we've seen we've seen specs on um
  935. 41:46or that are public specs on, let's say.
  936. 41:49uh not giving away too much there. So,
  937. 41:54um
  938. 41:55maybe just to talk about realizing it a
  939. 41:57little bit more.
  940. 41:59Uh
  941. 42:01I don't know, Brian, do you want to talk
  942. 42:03a little bit about the software
  943. 42:04programming model and the micro
  944. 42:06architecture? You want me to talk about
  945. 42:07that?
  946. 42:10>> I think I think you're more
  947. 42:12knowledgeable in the aspect, right? But
  948. 42:15but actually be before I let you answer
  949. 42:17your own question,
  950. 42:19the another very big interesting point
  951. 42:21is like the role of AI in all of this
  952. 42:24right now you can you can make you can
  953. 42:27make an argument that open AI's biggest
  954. 42:29advantage is that it's able to access
  955. 42:31its own Astra models uh the new GP Astra
  956. 42:35models before anyone else. So like and
  957. 42:39that's that's that's like the biggest
  958. 42:40difference between I would say between
  959. 42:42open AI and one of the other new chips
  960. 42:45companies sova cerebras etc. Uh yeah the
  961. 42:49question is how much did at least my
  962. 42:52question is how much did Astra actually
  963. 42:54contribute to this like if you look at
  964. 42:56it from a differences point of view like
  965. 42:58this is only one of the only differences
  966. 43:00between uh
  967. 43:02uh Jalapeno and the other chips like
  968. 43:07so did Astra really contribute to like
  969. 43:11most of these performance differences or
  970. 43:14just a bit yeah but that's just that's a
  971. 43:17tangent of
  972. 43:18AI work on development and yeah actually
  973. 43:21Kim K3 when Kim K3 was released there
  974. 43:23was a part on the blog about its
  975. 43:26designing of a chip I forgot what what
  976. 43:29what was it about maybe some yeah I'm
  977. 43:32not sure what the architecture was about
  978. 43:33but yeah they did talk about K3
  979. 43:36developing a chip and yeah I would guess
  980. 43:39GPT extra does have similar capabilities
  981. 43:42and I would say to a better aspect or to
  982. 43:46higher degree
  983. 43:48Yeah. Sorry. Back back to you Jordan.
  984. 43:50Yeah. On the architect micro
  985. 43:52architecture.
  986. 43:53>> Yeah. Well, let I mean, let me comment
  987. 43:55on the AI assistance on the architecture
  988. 43:58cuz I think it's there's two ways in
  989. 44:00which AI clearly assisted the
  990. 44:03design and then the bring up of the
  991. 44:05chip, namely design and then bring up.
  992. 44:08So, on the design side, this clearly
  993. 44:11wasn't Astra because the RCL freeze was
  994. 44:14in July of last year.
  995. 44:16So, you know, from February to July of
  996. 44:19last year, when they claimed that AI
  997. 44:22assistance helped them get an 8%
  998. 44:24reduction in SIMD area and then a 10%
  999. 44:26reduction in the matrix engine area
  1000. 44:29during design, I mean, this is preGPT5
  1001. 44:33that we're talking about. So, um,
  1002. 44:36conceptually like the models have to be
  1003. 44:37getting better at RTL in the meantime,
  1004. 44:41but they were already good enough to
  1005. 44:43rapidly
  1006. 44:45like accelerate the
  1007. 44:47um really tedious human driven work that
  1008. 44:51is RTL before a tape out, right? Um and
  1009. 44:56so I I think you know maybe that's the
  1010. 44:59biggest claim here. Uh which is that
  1011. 45:04I think I think a lot of people
  1012. 45:06understand you can use these models for
  1013. 45:07kernels or and just software engineering
  1014. 45:09because like you put it in a codeex
  1015. 45:11harness you put it in a loop uh you let
  1016. 45:14it test the thing and then you just set
  1017. 45:16goal and like people have had that
  1018. 45:17experience they can kind of understand
  1019. 45:19it but I don't think a lot of people can
  1020. 45:22like have had the experience of
  1021. 45:23designing a chip. It's not like
  1022. 45:24traditional software programming and um
  1023. 45:28this was done with an older model. So I
  1024. 45:31think that's like point number one. Now
  1025. 45:34uh on the actual bring up yeah like kind
  1026. 45:39of like what I just said
  1027. 45:41the work is iterative and it's a
  1028. 45:43verifiable domain. So like this is this
  1029. 45:45kind of exactly what RL should be good
  1030. 45:48at. You should be able to give a model a
  1031. 45:50task of improving the performance of a
  1032. 45:52kernel or getting the kernel to be
  1033. 45:54functionally correct against some
  1034. 45:57u like verification script
  1035. 46:00um against some test cases, right? And
  1036. 46:03then just let the model rip, let it try.
  1037. 46:07And Dylan made this podcast on made this
  1038. 46:09point on the Dores Cesh podcast that he
  1039. 46:11was on recently, which is that for years
  1040. 46:13I think now these
  1041. 46:17uh the companies that were getting the
  1042. 46:19most value out of using AI were not
  1043. 46:21actually the companies providing AI.
  1044. 46:23like OpenAI and Anthropic were not
  1045. 46:25profitable for a very long time and now
  1046. 46:28they're like just turning a profit but I
  1047. 46:30mean still running incredibly high
  1048. 46:32margin businesses but they're just like
  1049. 46:33realizing the profitability of training
  1050. 46:36these advanced models meanwhile you know
  1051. 46:39Jane Street's going out there and
  1052. 46:40printing $15 billion in a quarter
  1053. 46:42clearly using AI for trading or
  1054. 46:44something like that right and there's
  1055. 46:46many other companies that are being
  1056. 46:47started based on the use of AI um this
  1057. 46:52is a very clear example of open AI
  1058. 46:56keeping the benefits of having access to
  1059. 47:00a model before everybody else for
  1060. 47:02themselves. They can take out a chip and
  1061. 47:04their competitors can't. And it's a sign
  1062. 47:06of what's to come. I think that's the
  1063. 47:09simplest way to put it. Um they're going
  1064. 47:12to be able to go into many domains that
  1065. 47:15are tangentially related to software
  1066. 47:18where it's like the model needs to be
  1067. 47:20able to control a computer, but it's not
  1068. 47:23explicitly like the thing you're
  1069. 47:24training it for.
  1070. 47:28You build RL environments, you spend
  1071. 47:30enough tokens, you spend enough like
  1072. 47:32time on reasoning and
  1073. 47:35enough rollouts, you know, enough like
  1074. 47:37attempts at the um problem and you're
  1075. 47:41going to get a good result. That seems
  1076. 47:42to be the lesson here.
  1077. 47:44>> Yeah. Yeah. Exactly. And I I think
  1078. 47:46Entropic is also realizing the same
  1079. 47:47thing. They are starting to hire like
  1080. 47:50silicon people. that's on the same part
  1081. 47:52as open
  1082. 47:54and they are doing a lot of stuff like
  1083. 47:56in the laboratory they got like LMS to
  1084. 47:59control microscopes and whatnot recently
  1085. 48:01and they're going seems to be going
  1086. 48:04quite long into this like biological
  1087. 48:05sciences field. Yeah. So I really agree
  1088. 48:08on with you Jordan on the point of like
  1089. 48:11uh these frontier model companies are
  1090. 48:13realizing what
  1091. 48:16this
  1092. 48:18oops lots of interference from Jordan
  1093. 48:21but yeah lots of good like downstream
  1094. 48:24impacts of having a good model first can
  1095. 48:27have not just like making money from
  1096. 48:30inference revenue. Yeah,
  1097. 48:38lots of interesting developments.
  1098. 48:40>> What's your thoughts, dude?
  1099. 48:43>> Yeah, I agree. I think um the the
  1100. 48:47progress that Yeah, Labs is only like
  1101. 48:50getting faster, right? And that's really
  1102. 48:52because
  1103. 48:53they're using their own models um really
  1104. 48:56effectively to drive product innovation
  1105. 48:59um much faster, right? Um, I remember
  1106. 49:02like I think was it earlier this year
  1107. 49:05like Anthropic was releasing a new
  1108. 49:08product like every week or something.
  1109. 49:09Um, and I think it was like everything
  1110. 49:11was basically on autopilot. They were
  1111. 49:13just using code for everything, right?
  1112. 49:16Um, so yeah, I think this
  1113. 49:19I agree with with Yeah. with what you
  1114. 49:22guys have said.
  1115. 49:24>> Yeah. Yeah. Um, yeah. Yeah, I mean like
  1116. 49:27the cynical view of a program like this
  1117. 49:29for both OpenAI and Enthropic and even
  1118. 49:31Meta and Microsoft is that it's kind of
  1119. 49:34like a head fake that gets them a
  1120. 49:36discount on the Nvidia GPUs and
  1121. 49:38therefore it pays for itself. Like you
  1122. 49:40only need to spend a few hundred million
  1123. 49:44on a program to like tape out a chip and
  1124. 49:48you know like the year or two to do it
  1125. 49:52to to potentially like um help with the
  1126. 49:56negotiations and if those negotiations
  1127. 49:58are measured to the tune of hundreds of
  1128. 50:00billions of dollars then it pays for
  1129. 50:02itself pretty quickly. But
  1130. 50:06the nonsynical view is that like this is
  1131. 50:11a real thing and it's only going to
  1132. 50:13they're only going to do more of this in
  1133. 50:14the future, right? Like
  1134. 50:16um there's no reason that they're going
  1135. 50:18to be less vertically integrated and
  1136. 50:21less interested in developing chips this
  1137. 50:24time next year. And there's no reason to
  1138. 50:26say that the RTL time from freeze to or
  1139. 50:31from like initial to freeze to tape out
  1140. 50:35can't go even shorter than 9 months. Um
  1141. 50:39and I I think you just need to to think
  1142. 50:41about where to go from there. Maybe the
  1143. 50:43other thing that was was kind of
  1144. 50:45interesting here. So like if we go
  1145. 50:47through the architecture, I mean there's
  1146. 50:49lots to say about it, right? Um
  1147. 50:53the maybe the quick high level is just
  1148. 50:55like it looks like a TPU with much
  1149. 50:58smaller um systolics uh systolic array
  1150. 51:04being like the the way a TPU has its
  1151. 51:07processing elements laid out. Um, and
  1152. 51:12maybe the the criticism of chips like a
  1153. 51:16TPU, tranium,
  1154. 51:18uh, even some of the ones that are like
  1155. 51:20TPU inspired, let's say like an etched
  1156. 51:22or a Maddox X or something like that, is
  1157. 51:24that these when they go with these
  1158. 51:26really big systolic arrays, um, they can
  1159. 51:29have these weird cliffs where small
  1160. 51:32batch dimensions like the the M
  1161. 51:34dimension in your MNK for a matrix
  1162. 51:37multiplication gets all which you know
  1163. 51:40is what happens when you have lots of
  1164. 51:41experts and you have very few requests
  1165. 51:44like low concurrency.
  1166. 51:46Um
  1167. 51:48you know the these like skinny gems,
  1168. 51:52skinny matrix multiplications
  1169. 51:54um can waste a lot of resources, right?
  1170. 51:56or um even just odd numbers where if you
  1171. 51:59go like slightly over 256 or slightly
  1172. 52:02over 128, now you're spending an entire
  1173. 52:05kernel launch on the device side just to
  1174. 52:08run one little skinny gem. And uh all of
  1175. 52:11these like
  1176. 52:13uh processing elements on the systolic
  1177. 52:16array are not being used and so
  1178. 52:18therefore it's like inefficient and you
  1179. 52:19don't actually maximize the flops on the
  1180. 52:23um chip itself. Uh the trade-off here of
  1181. 52:28course is that to get more efficiency
  1182. 52:30with tiling to you know reduce the
  1183. 52:33issues with like padding overhead or or
  1184. 52:35uh alignment on the matrix dimensions is
  1185. 52:38that you just use smaller systolics and
  1186. 52:40that's what they've done here. And so I
  1187. 52:42think that's been really smart clearly
  1188. 52:44for for efficiency across the curve. The
  1189. 52:46argument the other way is that they're
  1190. 52:48going to miss out on some I think power
  1191. 52:50efficiency and and like data movement
  1192. 52:52efficiency because well you have to have
  1193. 52:54more small elements instead of one big
  1194. 52:57element. So what do you do then? And I
  1195. 53:00guess the way that they've solved this
  1196. 53:01is by um being really smart about how
  1197. 53:05they place weights and KBs
  1198. 53:08um using synchronization between cores
  1199. 53:11like selectively and then saving the
  1200. 53:14collective network the like knock the
  1201. 53:16like network on chip that connects the
  1202. 53:18HPM slices and the computing elements
  1203. 53:20together really really sparingly. And
  1204. 53:23the I mean the results is like well the
  1205. 53:25results speak for itself and you see the
  1206. 53:27performance there but the results is
  1207. 53:29that this might be both a chip chip
  1208. 53:32that's simpler to reason about than a
  1209. 53:34GPU and a chip that's a little bit more
  1210. 53:36flexible for some of these weird
  1211. 53:39changing uh dimensions over time than a
  1212. 53:42TPU. And so I mean clearly the guys who
  1213. 53:46have experience using GPUs which OpenAI
  1214. 53:48has plenty of experience programming
  1215. 53:49GPUs mixed with the guys who have
  1216. 53:51experience designing a TPU has resulted
  1217. 53:54in a pretty well balanced system here.
  1218. 53:57Um
  1219. 53:59maybe the other thing is that it has an
  1220. 54:01L1 cache which is quite funny like all
  1221. 54:04of these accelerators do not have L1
  1222. 54:06caches now. They rely so much on L2 um
  1223. 54:09namely SRAM. we hear SRAMM all the time
  1224. 54:12and um I mean the the reason why I guess
  1225. 54:16is that uh they
  1226. 54:22like the the companies
  1227. 54:25designing other accelerators don't want
  1228. 54:28to use a um scratch pad uh sorry they do
  1229. 54:33want to use like a software manage
  1230. 54:34scratch pad and OpenAI is not using a
  1231. 54:36scratchpad cache here um and so that
  1232. 54:39makes the chip like potentially harder
  1233. 54:41to reason about when you think about
  1234. 54:42like barrier latencies and where you're
  1235. 54:44going to move data like you have to be
  1236. 54:46able to amvertise the data movement by
  1237. 54:49doing work on the CPU itself. But that
  1238. 54:52ties in with, you know, what I was
  1239. 54:54saying at the very beginning, which is
  1240. 54:55like the second phase of using AI, which
  1241. 54:57is um actually programming kernels on
  1242. 55:01like meaning software that runs on the
  1243. 55:02device side. And um the the to do this,
  1244. 55:07I mean, we haven't really been able to
  1245. 55:08verify this other than scrolling through
  1246. 55:10a a 30,000line
  1247. 55:13file with some of the OpenAI engineers
  1248. 55:15when we went on site with them. It's
  1249. 55:16like um literally
  1250. 55:20it's just
  1251. 55:22sloth. Like it's it's not sloth cuz it
  1252. 55:25performs, but it's literally just AI
  1253. 55:27generated assembly basically puked out
  1254. 55:31in gluon this uh low-level like kernel
  1255. 55:34programming languages that they built on
  1256. 55:35top of Triton which uses this you know
  1257. 55:38really interesting programming model.
  1258. 55:40Like the the point is like the guys who
  1259. 55:43were scrolling through this this code
  1260. 55:44with us, it was kind of clear that like
  1261. 55:46they know a whole bunch about hardware.
  1262. 55:48They know a whole bunch about the
  1263. 55:49concepts in the system and they just
  1264. 55:51like have no idea what this MLA kernel
  1265. 55:53that they're showing us for DeepSeek
  1266. 55:55actually does. Like like you can go line
  1267. 55:58by line and it's like nope nope nope. Um
  1268. 56:02but it doesn't matter, right? And uh the
  1269. 56:06AI understands it. the I tests it and
  1270. 56:09you see the results. It it produces
  1271. 56:11correct kernels that perform really
  1272. 56:12well. And uh I think this is like just a
  1273. 56:15sign of what's to come again, right? The
  1274. 56:18um the like that actual
  1275. 56:23code is not necessarily something that a
  1276. 56:26human has to reason about deeply if the
  1277. 56:29AI knows how to manipulate the data
  1278. 56:31movement, the processing elements on the
  1279. 56:33hardware that you've given it.
  1280. 56:37Um, okay. We're kind of running out of
  1281. 56:39time here. We've been going for a while.
  1282. 56:41Uh, we got three things that I had in my
  1283. 56:43notes that we wanted to talk about. Uh,
  1284. 56:46they don't do PD disag.
  1285. 56:49Brian, maybe you can rant about that
  1286. 56:51because you spend like all of your time
  1287. 56:52debugging PD disag.
  1288. 56:57[laughter and gasps]
  1289. 56:59Two, we didn't really talk about the
  1290. 57:00system architecture. We can talk a
  1291. 57:02little bit about how they do scale up,
  1292. 57:04scale out domains. I mean they don't
  1293. 57:06call it scale out but uh whatever it is
  1294. 57:08scale out the multi-ter scale up stuff
  1295. 57:12which is just like a total mess to try
  1296. 57:15to understand probably can't communicate
  1297. 57:16on a podcast go read the article and
  1298. 57:19then the third thing is it's a
  1299. 57:21generalized inference chip it's not
  1300. 57:23co-designed with their models like they
  1301. 57:25keep saying and the proof of that is it
  1302. 57:28runs Doom
  1303. 57:3036 frames per I can [laughter]
  1304. 57:36um [clears throat]
  1305. 57:39>> Yeah. [laughter]
  1306. 57:42Yeah. We we I remember we were talking
  1307. 57:45to these guys and we were like the like
  1308. 57:47angry disappointed mother who comes in
  1309. 57:50and they show us this like
  1310. 57:51groundbreaking chip that's so fast and
  1311. 57:53runs this stuff and they're like but
  1312. 57:55does it run Agent X? No. [laughter]
  1313. 58:0096% on the test. What four questions did
  1314. 58:04you get wrong?
  1315. 58:11Um anyway, anything you guys feel is
  1316. 58:14left unsaid on the chip? Uh pretty
  1317. 58:18exciting release, eh?
  1318. 58:21Yeah, very um I'm yeah really excited to
  1319. 58:24see where the road map goes next and
  1320. 58:28also um you're excited to see I mean
  1321. 58:31Brian mentioned this earlier but Brian
  1322. 58:33uh Anthropic is your is building it or
  1323. 58:37as um they're hiring it they're building
  1324. 58:39a team to do it. Um I think this really
  1325. 58:42sets a pretty high benchmark for
  1326. 58:45Anthropic um to to meet or beat. Um but
  1327. 58:50I think you know there's every reason to
  1328. 58:51believe that um your anthropic could
  1329. 58:55achieve a similar outcome. So very
  1330. 58:57excited to see that as well.
  1331. 58:59>> And they're they're hiring the team now,
  1332. 59:01right? So it's only like what a month
  1333. 59:03and a half until the RTL freeze and then
  1334. 59:06like what four or five more months for
  1335. 59:08the tape.
  1336. 59:12[laughter] So like we should be able to
  1337. 59:13get anthropic custom you know chip
  1338. 59:16tokens
  1339. 59:18um after like one
  1340. 59:21you know 9 month cycle right like one
  1341. 59:25[laughter]
  1342. 59:26one pregnancy term
  1343. 59:28[gasps]
  1344. 59:31you know think
  1345. 59:33yeah
  1346. 59:36anthropics in
  1347. 59:39anthropics hiring people right now so
  1348. 59:41it's like the first trime trimester and
  1349. 59:42then like the second one they they get
  1350. 59:44the RTL freeze so they go to the second
  1351. 59:47trimester and then like goes off to the
  1352. 59:49fab and then like tape out it comes back
  1353. 59:51third trimester all done right
  1354. 59:59>> a little bit longer. Yeah,
  1355. 1:00:02>> you guys don't like that one. Okay,
  1356. 1:00:04we're going to end on one other joke. So
  1357. 1:00:06Brian, you you like the one about
  1358. 1:00:08converting energy usage to calories,
  1359. 1:00:10right? So
  1360. 1:00:14I want to finish this one. We uh we
  1361. 1:00:17converted human we converted to human
  1362. 1:00:19speech and compared uh some of this
  1363. 1:00:21stuff on a on a calories, right? Because
  1364. 1:00:23uh a calorie or like how much energy it
  1365. 1:00:26takes to burn a what a cubic centimeter
  1366. 1:00:30of water I think is uh the equivalent of
  1367. 1:00:35uh whatever one jewel is 0239 food
  1368. 1:00:40calories. So, we will uh we'll throw
  1369. 1:00:42this this one up on screen to lead
  1370. 1:00:44everybody off with a a little joke. Uh
  1371. 1:00:47if you convert all of this efficiency of
  1372. 1:00:49some of these DeepSeek results that we
  1373. 1:00:52have on Asian X, uh we've identified
  1374. 1:00:54that human speech is roughly 20 times 22
  1375. 1:00:57times more energy efficient than the uh
  1376. 1:01:00concurrency one B300 configuration that
  1377. 1:01:02we were testing it against there. Right.
  1378. 1:01:07A human speaks at 3.3 tokens per second,
  1379. 1:01:09but these uh batch one configurations
  1380. 1:01:11are going up at 180 tokens per second,
  1381. 1:01:14much faster than the human brain can
  1382. 1:01:16work, consume calories, and produce
  1383. 1:01:18speech. So
  1384. 1:01:20anyway,
  1385. 1:01:21>> I think the cav the caveat is that um
  1386. 1:01:25not all human spoken tokens are very
  1387. 1:01:27high quality.
  1388. 1:01:30[laughter]
  1389. 1:01:32I mean,
  1390. 1:01:35[laughter and gasps]
  1391. 1:01:36>> I got to caveat some of my interactions
  1392. 1:01:38with Claude, too. Then, meant some of
  1393. 1:01:39this nonsense that's been spitting back
  1394. 1:01:41at me recently, I've I've I've not been
  1395. 1:01:43too pleased with either. [laughter]
  1396. 1:01:48Can't see all the thinking traces
  1397. 1:01:50anyway, but the like clawed version of
  1398. 1:01:52English it's given me has not been that
  1399. 1:01:53great. Um, [gasps] yeah. Well, uh, if
  1400. 1:01:58we're comparing the machines on how many
  1401. 1:01:59calories they're consuming per token and
  1402. 1:02:01we're comparing it to our speech, we are
  1403. 1:02:03really in competition with the machines
  1404. 1:02:05at this point then, huh? Hopefully Jeb's
  1405. 1:02:09paradox continues and and uh, everybody
  1406. 1:02:11that produces chips wants to consume
  1407. 1:02:13more tokens and produce more chips and
  1408. 1:02:15hire more people and everybody gets to
  1409. 1:02:16come have fun.
  1410. 1:02:20All right, guys. Uh, good job. Long
  1411. 1:02:23episode this time. I hope everybody
  1412. 1:02:24enjoyed the overview of OpenAI Jalapeno.
  1413. 1:02:27More to come. Last joke before we leave,
  1414. 1:02:31cuz I just saw it. The best cover image
  1415. 1:02:33in a while, I'd say. Right.
  1416. 1:02:37[laughter]
  1417. 1:02:38We'll leave that one on screen. Lisa and
  1418. 1:02:40Jensen enjoying a nice spicy pot of uh
  1419. 1:02:44chana or or katsu, whatever the uh code
  1420. 1:02:48names were for the trays and the rocks
  1421. 1:02:50and stuff in there. That's the
  1422. 1:02:51motivation behind that.
  1423. 1:02:53Vindaloo. Yeah,
  1424. 1:02:54>> right.
  1425. 1:02:55>> Sorry about that. All right, we'll sign
  1426. 1:02:57off with that image in everybody's brain
  1427. 1:02:59who's watching online. [laughter]
  1428. 1:03:02Thanks for listening, guys.

About this transcript

This page contains the full transcript of Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators) by SemiAnalysis, generated from the public captions YouTube serves with the video. The transcript has 9,614 words across 1,428 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.