Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators) — Transcript
Full transcript
- 0:05Hello everyone. Welcome back to Semi
- 0:07analysis weekly. Uh today I'm joined by
- 0:09Brian and Myron. We're going to talk
- 0:11about the new article we put out on
- 0:13OpenAI Jalapeno veteran Blackwell their
- 0:16self-designed ASIC which we compared
- 0:18with Ruben. We analyzed the TCO. uh we
- 0:22assessed the throughput per megawatt
- 0:24claims a little bit more of the details
- 0:26in terms of the micro architecture
- 0:27system architecture how they used AI to
- 0:30design it and write kernels and why the
- 0:33speed uh from tape out to actually like
- 0:37first working system with first real
- 0:39benchmarks run on it has been so
- 0:40impressive. So guys excited to dig in.
- 0:43This is a fun one.
- 0:46Yeah,
- 0:49[laughter]
- 0:50>> you could say something in response
- 0:52there, Brian.
- 0:55>> Sorry, I
- 0:58>> Is this a Is this a
- 1:00>> Is this a virtual background that you
- 1:02got like pink cotton candy background
- 1:04going on or what?
- 1:05>> Of course. Of course. I think it's the
- 1:08most similar color to what uh I think
- 1:11Open AI's belonged. I don't remember off
- 1:13the top of my head, but they're good.
- 1:14They got this like nice designs for
- 1:16their blocks. Yeah. But yeah, but
- 1:18anyway, Jalapeno has been great. Uh
- 1:20quite a surprise on the last day of hot
- 1:22chips for those attending.
- 1:25Uh one of the most surprising talks in
- 1:28my opinion. Uh and yeah, I believe it
- 1:31caught the performance caught everyone
- 1:33by surprise. We all knew that OpenAI
- 1:37chip was in the works and they have
- 1:38announced uh deals with Broadcom but I
- 1:42first look at performance results and we
- 1:45yeah and it's really looking quite good
- 1:47for OpenAI
- 1:49a brief summary for those that didn't
- 1:52read the article which is kind of
- 1:54unlikely is that on per megawatt basis
- 1:57and perf per TCO basis. So in other
- 2:00words, uh how much it cost to actually
- 2:03run the chip. Uh actually Jalapeno
- 2:07actually beats Vera Rubin's July results
- 2:11and we are comparing against July
- 2:14results as we believe it's kind of fair
- 2:16because the software is in like a
- 2:19similar state uh between these two
- 2:21points. Yeah. So as you shown as on the
- 2:25picture, this is Jalapeno compared
- 2:27against GB300. So JP300 it beats out of
- 2:31the water completely. Uh but as we
- 2:33mentioned during our in our article it's
- 2:36a bit unfair to compare it against
- 2:39Blackwell because Blackwell uses HBM3
- 2:42while uh Jalapeno uses HBM4. So actually
- 2:46my can talk about this a bit later on
- 2:49the differences between the HBM but it's
- 2:53a bit unfair to compare HBM 4 against
- 2:55HBM3. So we're comparing it against
- 2:57Rubin instead.
- 2:59And yeah, as in the results show, he
- 3:02does beat Rubin on output token
- 3:05throughput per oil utility megawatt. But
- 3:08then again, these are July results. And
- 3:11where Rubin right now is likely uh much
- 3:14better than July. But Jalapeno of course
- 3:17will improve as time goes on. As we
- 3:19shown in another diagram below Jordan, I
- 3:22think
- 3:23>> uh we show how jalapeno has improved
- 3:25from
- 3:27uh 25 days.
- 3:29>> 25 days.
- 3:32>> Yeah, I'll go grab that one. But um
- 3:34maybe like the
- 3:36>> the key point you were making if you can
- 3:38explain a little bit more about
- 3:40>> the benchmarks that they're running
- 3:42which is Deepse CR1 8K1K.
- 3:46um not really the like absolute biggest
- 3:48model and not the most demanding
- 3:51inference workload because it's uh just
- 3:53random data. It's not agenic workflows.
- 3:56Um but really quickly they got this up.
- 4:01Clearly the performance is strong as
- 4:03you're yes implying with this new chart.
- 4:07um performance is like improving by the
- 4:10week at this point even by the day and
- 4:13um maybe you can you can also explain
- 4:16the denominator there like why they're
- 4:18choosing to divide by the power
- 4:20consumption.
- 4:22Yeah, actually uh Jensen brought up this
- 4:26point in Computex 2026
- 4:29during his keynote that uh data centers
- 4:32nowadays are getting power limited and
- 4:34that power is starting becoming the
- 4:36constraint. If you have money, you can
- 4:38always get more servers, more chips, but
- 4:40the the constraint is starting to become
- 4:43the data centers power. And like
- 4:46companies have tried getting through
- 4:48this using a BTM or behind the meter
- 4:51power but end of the day the power tends
- 4:54to be your your constraint for building
- 4:56a data center.
- 4:59Yeah, absolutely. So like if you have
- 5:01100 megawatts of power you can only fit
- 5:04so many chips in there. It doesn't
- 5:06really matter if these chips are more
- 5:09expensive
- 5:11uh on a per megawatt basis. if they're
- 5:14producing more tokens per megawatt, you
- 5:16think you can make more money off of the
- 5:18tokens they produce and you justify the
- 5:19extra expense on the actual chips.
- 5:22And
- 5:24while it's a valid metric for some
- 5:27customers, it may not be valid for
- 5:29others. And so it's it's not the most
- 5:31typical way. Most typically we see just
- 5:35token throughput per GPU compared like
- 5:38per package. Um, but in this case,
- 5:42OpenAI doesn't have customers for
- 5:44Jalapeno. They just run it for
- 5:45themselves and so they don't really care
- 5:49tokens per package. They care about
- 5:51tokens per like megawatt that's going
- 5:54into the system, right? Yeah. And
- 5:57actually going off of this uh topic uh
- 6:01someone did mention I forgot one of us
- 6:03mentioned that uh like performance per
- 6:06chip is at the end of the day is just a
- 6:08imaginary thing. You can just glue two
- 6:11chips together and call it you double
- 6:13your performance per chip. So looking at
- 6:15per chip is not really a wrong metric.
- 6:18just it can be easily like misled cuz
- 6:21you can just which is what Ruben Ultra
- 6:23and Blackware Ultra does if I'm not
- 6:26wrong. You just glue two chips together
- 6:29and bing bang boom, you got double your
- 6:32true per chip. Yeah.
- 6:33>> Yeah. I mean, you can do this if you're
- 6:35cerebrous, too, right? You can just put
- 6:37three of the wafers in a rack and then
- 6:38you can put them all on a chart where
- 6:40three is better than one, right?
- 6:43>> Yeah. Or or Yeah. Just say your chip is
- 6:45the whole wafer, right? Um yeah
- 6:49um yeah but I mean uh it it doesn't mean
- 6:52that um they don't care about
- 6:54performance per per cost right because
- 6:57I'd say that um you know performance per
- 7:00what and performance cost they're pretty
- 7:02closely closely linked right I mean I
- 7:05think um the more power the chip
- 7:08consumes it probably means the chip is
- 7:11more expensive as well generally and I
- 7:14think we see that um yeah we we did our
- 7:17um you know TCO analysis of what we
- 7:19think open would pay for a jalapeno
- 7:23system and you know they still pretty
- 7:25much win or or very close on uh
- 7:28performance for TCO as well, right?
- 7:32>> Yeah. Here's that chart on screen. I
- 7:34mean, this one's just more interesting
- 7:36because obviously the TCO calculation in
- 7:39terms of like how many tokens you're
- 7:41going to get per dollar
- 7:43um depends on that input for per dollar.
- 7:47And so we've got some variables on
- 7:50screen there about, you know, what's the
- 7:52total cost per hour to own, for example,
- 7:56a GV300 or to own Jalapeno. And we're
- 7:59assuming that OpenAI is going to pay 279
- 8:03an hour for a GB300.
- 8:05They're buying them themselves. They're
- 8:06running the data centers themselves, for
- 8:08example, in that case. Um they're not
- 8:10playing paying the current going
- 8:12NeoCloud prices, which are like six
- 8:14bucks an hour for GB300 right now. And
- 8:17put Vera Rubin at um 361 and then
- 8:21Jalapeno at $1.56. And so let's just
- 8:25compare it per package. Like clearly
- 8:28they're going to like the thesis of
- 8:30designing a chip in house is that you
- 8:31want to pay the Broadcom margins only,
- 8:34not the Nvidia margins or the Broadcom
- 8:38plus Google TPU margins or some
- 8:40equivalent there, right? Um and that's
- 8:44that's being borne out on this chart
- 8:46that you can see.
- 8:53>> All right. And I guess the the next part
- 8:56of this is like um
- 8:59where could like how could these curbs
- 9:02move? Um I think with jalapeno they're
- 9:06actually somewhat sandbagging
- 9:09these results significantly. Right.
- 9:14Uh yeah, I mean this is so the caveats I
- 9:18guess on the performance is that I I
- 9:21said it earlier that um the tests they
- 9:24ran were deepseek 1 uh that's based on
- 9:27the V3 architecture that came out in
- 9:29January of 2025. So it's not the most
- 9:30current DeepC4 but it is a relatively
- 9:33big model 600 billion total parameters.
- 9:37They ran Kimmy K 2.5 which is 1.5
- 9:40trillion. So that's a actual big model
- 9:43or just 1 trillion, sorry, not 1.5. And
- 9:46then they ran GPTOSS 12B, their own open
- 9:49source model, relatively small. Uh they
- 9:53had solid performance on all three. Um
- 9:55but
- 9:58I mean solid performance. They're
- 9:59beating Farah Rubin [laughter]
- 10:01on uh on all three today. And they're
- 10:03beating them with single token
- 10:06prediction and no pre-filled decode
- 10:08disagregation. So
- 10:11Brian, maybe you can explain single
- 10:14versus multi-token prediction and
- 10:16specifically in these performance claims
- 10:18when compared directly to Bar Rubin,
- 10:20which is using MTP.
- 10:23Um, this is like OpenAI kind of
- 10:26fighting with one hand behind their
- 10:28back.
- 10:30>> Yep.
- 10:30>> Like MTP is a really big optimization
- 10:33>> for interactivity.
- 10:35>> Yeah, of course. Uh actually going back
- 10:39on uh the deepse uh one it reminded me
- 10:42of an X post I saw like some time ago
- 10:46that someone said the closed lab open AI
- 10:49and entropic probably have an internal
- 10:51version of MLA and they probably
- 10:53discovered
- 10:54uh MLA like long before deep but it's
- 10:59quite interesting that uh we were told
- 11:02that the MLA kernels like openi didn't
- 11:05have any internal MLA kernel uh
- 11:07implementation. So I'm not saying like
- 11:09the open source tritons one they don't
- 11:11have any internet optimized MLA kernels
- 11:14which is quite interesting because that
- 11:16well that means none of the open AI
- 11:18models uses MLA which is quite
- 11:21interesting to me and uh yeah and on the
- 11:25AK1K and Deepc R1 like model choice we
- 11:30did release an AX uh benchmark case but
- 11:35the timing was very bad so I think we
- 11:37didn't get uh the OpenAI Jalapeno team
- 11:39to actually run it on Ethernet X. But
- 11:42yeah, as Jordan says, uh measures stuff
- 11:45like prefix cache which uh exposes
- 11:49a lot more areas for optimization. And
- 11:53in our we describe like a lot of stuff
- 11:55there's loadbearing
- 11:57that is being tested for Agent X and
- 11:59it's starting to show like what break
- 12:01what's breaking what's not and areas for
- 12:04improvement that AK1K didn't properly
- 12:07show.
- 12:08And uh and yeah going so going back to
- 12:10the MTP question uh for those unaware
- 12:13MTP stands for multi token prediction
- 12:14and it's a form of speculative decoding
- 12:17where you kind of guess the next future
- 12:20tokens and then verify them in a single
- 12:22forward pass because uh interestingly
- 12:25the LM doesn't output just the next
- 12:28token's probability but the token
- 12:31probability of every single token
- 12:34position before the last token.
- 12:37And we were not given an MTP results.
- 12:39And it's my personal guess that OpenAI
- 12:42has some internal speculative decoding
- 12:44technique that's not MTP and not DSpark
- 12:47or any open source uh configs. So they
- 12:52didn't give us speculative decoding
- 12:54results because
- 12:56there's no way to actually verify it
- 12:58through open source which is also quite
- 13:01interesting in my opinion means this
- 13:03spark isn't the best that we can do and
- 13:06this park actually gives quite a a crazy
- 13:09advantage over MTP by doing like uh
- 13:12guessing all the tokens at once uh
- 13:15instead of doing MTP which is just one
- 13:17layer of a model trying to
- 13:20uh guess future tokens one by one
- 13:23instead of doing them all at once.
- 13:25Yeah. And with Yeah. And we were told
- 13:28that the internal speculative decoding
- 13:30method gives like three to five times uh
- 13:33improvement on production models which
- 13:36yeah would just knock the jalapeno
- 13:39versus ferubin comparisons out of the
- 13:41water once again. So you can imagine
- 13:44that graph shifting like three to five
- 13:46times. Yeah, it's quite insane.
- 13:48So Kudo mode's gone, Brian.
- 13:52>> I mean, everyone on X is is saying, "Oh,
- 13:56it's just an ASIC. That's what you
- 13:57expected to do." And yeah, in some
- 13:59sense, that's what we that's what we
- 14:00expect an ASIC to do. But the software
- 14:03has come to such a point that like this
- 14:06line doesn't really matter if your
- 14:07software is good enough that you can get
- 14:09a motor up like very quickly. What's the
- 14:12difference between a general purpose or
- 14:15like they say GPGPU but N6 if software
- 14:19can just bridge this gap?
- 14:22>> Yeah.
- 14:23>> Yeah. Yeah. I mean this chip has I mean
- 14:27it's a toss up with Ver Rubin because we
- 14:28haven't seen stuff since they published
- 14:30that uh note in uh June or July. Um so
- 14:35maybe they've had a month extra to
- 14:37develop on the these chips. But I mean,
- 14:40conceptually, like we've never seen
- 14:43anybody else put out a chart where
- 14:46there's a curve showing they're beating
- 14:50Nvidia on every point of the curve
- 14:53and uh like a real test, right? So, it's
- 14:59I mean it's shocking that they did this.
- 15:01We can talk about the timeline a little
- 15:02bit later, but we we keep talking about
- 15:05the point in the curve and I think we we
- 15:06maybe haven't explained this in great
- 15:09detail or we've done it on previous
- 15:10podcast and people aren't, you know,
- 15:12familiar with this. So, the I'm going to
- 15:15put this chart back up on screen and try
- 15:17to explain a Fredo curve here for the
- 15:22purposes of understanding the
- 15:23performance claims made by Jalapeno
- 15:25here. And uh to do that we need to
- 15:28explain that the y- axis is how many
- 15:29tokens you can produce per megawatt
- 15:32you're putting into the system and the
- 15:33y- axis is how fast or sorry the x-axis
- 15:36is how fast the tokens appear to each
- 15:39individual user. And so whether you're
- 15:42optimizing for the y- axis or the x-axis
- 15:45jalapeno is beating the GB300 right now
- 15:48on the deepseek model which implies
- 15:50something like if you were to fix the
- 15:54interactivity per user. So everybody
- 15:56sees 100 tokens per second or 50 tokens
- 15:58per second. And I zoom in on like
- 16:01exactly that part of the curve. We're
- 16:03basically seeing that Jalapeno has
- 16:06double the amount of tokens that it can
- 16:09produce per megawatt implying
- 16:12two times more revenue, two times more
- 16:14profitability, whatever you want to say
- 16:16from an inference endpoint serving
- 16:18provider.
- 16:19And then if you look at the far right
- 16:21side of this curve across the x-axis and
- 16:24you see that at a very low batch size
- 16:26they can go all the way to 700 tokens
- 16:29per second per user. And you compare
- 16:32that to where the others peak out at
- 16:33350.
- 16:35This is once again
- 16:37double the performance.
- 16:39And
- 16:41I think the conclusion is that they're
- 16:44basically winning on both sides of the
- 16:49curve. So, both fast tokens and cheap
- 16:52tokens, which is
- 16:55it's just so interesting because we've
- 16:57seen so many other um companies
- 17:02make claims about how they're going to
- 17:04beat Nvidia and they just pick one of
- 17:06those,
- 17:09right? Croc or Cerebrus or any of the
- 17:11other startups that are going to focus
- 17:12on SRAMM like call it a Dmatrix that's
- 17:14coming up with stuff or Samonova where
- 17:16we've even seen some results on bench
- 17:19you know on the infertex benchmark or
- 17:20something similar to it they're saying
- 17:23they're going for fast tokens just
- 17:25decode speed low batch size they don't
- 17:27worry about throughput and then you've
- 17:29got other guys that are worried about
- 17:31throughput call it AMD just as a simple
- 17:33example but even TP or tranium could be
- 17:34in this bucket of just like accelerators
- 17:36that are going for throughput
- 17:38and they go, "Yeah, but we're not going
- 17:40to be able to compete with the other
- 17:41guys at high interactivity." And OpenAI
- 17:43has a chip that can do both for them.
- 17:45Really well. I mean, at a minimum, this
- 17:48thing is is doing really well right now
- 17:50and is going to serve real production
- 17:52tokens for them.
- 17:55>> Yeah. And yeah, it's actually quite
- 17:58surprising that I mean I was surprised
- 18:00that OpenAI was the first chip that
- 18:03actually
- 18:04uh like non Nvidia non AMD chip that
- 18:07actually appear on our public infrance
- 18:09X. We were like expecting some Nova or
- 18:12like Cerebras or even TPU trainium to be
- 18:15one of the first. But
- 18:18yeah, maybe this really come put into
- 18:19perspective like the difference in uh
- 18:22like how this time might be different
- 18:24from the rest like this the first chip
- 18:26that actually poses a real threat to the
- 18:28Kuda mode because
- 18:30like in open source we welcome uh
- 18:33results from anyone like you believe his
- 18:36chip is good run the results run the
- 18:39curves run the benchmarks show us what
- 18:42your chips does and we gladly put it on
- 18:44our dashboard and we gladly like write
- 18:46write an article about it if it's good
- 18:48but yeah and we have extended this offer
- 18:50to like edged very recently on X and of
- 18:53course etched uh didn't get back to us
- 18:56on that but yeah if your chip is good
- 19:00just run the benchmark show us the
- 19:02results and we let the results talk uh
- 19:05yeah so
- 19:06>> yeah it's it's
- 19:07>> we love to see more completion from the
- 19:09chips yeah exactly
- 19:11>> so did you guys see um Jensen's response
- 19:14to this I mean I level. I think the
- 19:17Cudamote has been uh getting drained
- 19:20slowly over the last couple of years
- 19:23like top models in the world like Claude
- 19:26and Gemini are trained without Nvidia
- 19:28GPUs to you know TPU right um Anthropic
- 19:33use lots of tranium it's not like you
- 19:35can only use GPUs but Nvidia is
- 19:38incredibly valuable company they're
- 19:39going to keep shipping all these GPUs
- 19:40and I think there's such a thing as like
- 19:42an Nvidia mode which includes everything
- 19:43in the supply chain the whole develop
- 19:45for ecosystem like all of the um
- 19:48availability to purchase and support and
- 19:50how you're going to like do the
- 19:52logistics of deploying these data
- 19:53centers and monitoring them over time. I
- 19:55mean, OpenAI's now got to go figure out
- 19:57how to turn on 100 megawatts of these
- 19:59chips, not just like three test racks,
- 20:01which is a monumental challenge to get
- 20:03over as if taping out a chip of this
- 20:07this uh you know quality is easy. Um
- 20:11which is not like the next phase will be
- 20:13pretty hard for them as well. Even
- 20:15Cerebrous is going through this
- 20:16themselves right now. So I guess the the
- 20:20question
- 20:21for me uh that well I think Jensen kind
- 20:27of answered in the style like I was
- 20:28saying there when he was on Mad Money
- 20:31with Jim Kramer, [laughter]
- 20:34everybody's favorite. Um, and he was
- 20:38basically like, I'm not bothered.
- 20:40And uh, yeah, like what's like Myron,
- 20:45what's your take on this? You think?
- 20:47Well, okay. What's your take on that?
- 20:49And just the whole like just the
- 20:50timeline to go from, yeah, we're tired
- 20:53of buying Nvidia GPUs for everything.
- 20:55We're going to go build it decision made
- 20:57at OpenAI to actually having a chip that
- 20:59can run in a sex benchmark. It's like
- 21:02under two years for concept to like real
- 21:05chip in lab and under nine months to
- 21:08actually get it taped out, right?
- 21:12>> Yeah. Uh yeah, I I have several thoughts
- 21:14on this. Um I guess starting from uh
- 21:19CUDA mo eroding um I think you know a
- 21:22lot of um how can you success both in
- 21:27terms of the silicon design
- 21:30as well as uh bringing up the software
- 21:32is it's it's been AI assisted right um
- 21:36and you know I guess somewhat the irony
- 21:38is that um this was all done on Nvidia
- 21:42GPUs training these models to bring up
- 21:44the capability to a point where um you
- 21:47know AI is able to you know program
- 21:50kernels um and that's been I think
- 21:53probably the big shift um in terms of um
- 21:57making it easier for the labs to and
- 22:01anyone to adopt um you know alternative
- 22:04systems for their their serving stack.
- 22:07Um I think you know one of the big
- 22:09reasons that uh Anthropic you know has
- 22:14decided to
- 22:16uh you know bring in AMD as as one of
- 22:20their hardware providers is because um
- 22:22you know Aentic programming as it allows
- 22:26them to sort of get around the the
- 22:28challenges uh of like using the AMD
- 22:31software stack for instance. So um you
- 22:33know somewhat ironically uh it's it's
- 22:37Nvidia's hardware has enabled um moving
- 22:40off of Nvidia's hardware. Um, and you
- 22:45know, in terms of I guess you jalapeno,
- 22:48it's it's such a big surprise because I
- 22:50mean we
- 22:52always knew that the team is capable,
- 22:55you know, they have they've had
- 22:57experience building um the other main
- 23:00successful um as program in the form of
- 23:04TPU. Um I think a lot of the yeah the
- 23:07the hardware team behind this is from
- 23:10their former TPU people as as we see in
- 23:12a lot of other um I guess whether it be
- 23:15AI accelerator startups or other A6 A6
- 23:18teams they tend to come from you know
- 23:20former people with TPU backgrounds um
- 23:24and um so for what it's worth sorry
- 23:27sorry sorry to interrupt but there's
- 23:29like basically no chip starters I can
- 23:32point to where it's a bunch of guys who
- 23:33are ex Nvidia
- 23:34But there's a lot of ex Google people
- 23:37out there doing stuff which is
- 23:38interesting.
- 23:39>> Yeah,
- 23:39>> I I do wonder sort of why why that is
- 23:42the case. But anyway, um another another
- 23:45topic. Um so uh and I think you know I
- 23:50mean designing an AI uh accelerator that
- 23:54is competitive with um Nvidia is is not
- 23:57easy, right? U it's such a huge market.
- 24:00Um, of course, everyone wants to try it,
- 24:02but um, you know, time and time again,
- 24:03we've seen uh, you know, as as as you
- 24:07guys have mentioned, we've seen a lot of
- 24:08entrance, but they haven't really been
- 24:10able to do it. Um, so, you know, we I
- 24:14think the expectation was that open
- 24:17would deliver a decent effort with their
- 24:19first generation. Um, and it turns out
- 24:21it was much better than decent. they,
- 24:23you know, as we said, they've already
- 24:24come out with something that's pretty
- 24:26much competitive or better than what the
- 24:28best of, uh, what Nvidia has to offer.
- 24:32So, that's surprise number one. Um, and
- 24:35then I think, you know, that also says
- 24:38something probably about um, other ASIC
- 24:41programs like especially, you know,
- 24:43Meta, Microsoft. Is it open air that's
- 24:46really good? Is, you know, are the
- 24:49silicon teams at Meta and Microsoft, do
- 24:50they have skill issues? It's probably a
- 24:52bit of both, right? I think um they're
- 24:55the guys that look the worst from from
- 24:57this announcement. Um but you know,
- 25:00going back to okay, where next? Um so
- 25:03designed a chip uh obviously you're
- 25:05scaling up the supply chain to um you
- 25:08deliver systems uh of mass like you
- 25:12talking about delivering like millions
- 25:14of these chips um you know thousands of
- 25:16racks um deploying them in data centers
- 25:19you know gigawatts of power um that's
- 25:22that's going to be not easy but also I
- 25:25think you know other people have have
- 25:27done it successfully I think the hardest
- 25:29part is really having that system design
- 25:31and then I think um you know OpenAI has
- 25:34partnered with um with people who have
- 25:38experience scaling this up right
- 25:40basically it's Broadcom and and
- 25:42Celestico on the system side uh and
- 25:44they've had experience you know doing
- 25:46this with TPU
- 25:49dig on the comment you made about the
- 25:51difference between like a in-house
- 25:53silicon program that's been going for
- 25:55years and years like MTIA at Meta or
- 25:57Maya at Microsoft if we just literally
- 26:00look at the specs of Jalapeno. Like it
- 26:04doesn't look
- 26:06super fancy on paper. I mean, it is an
- 26:09HBM4 chip, so it's going to have great
- 26:12HBM bandwidth.
- 26:14>> Um, it's got lots of
- 26:15>> FP4 flops, but like still
- 26:18>> less than
- 26:19>> Mhm.
- 26:20>> half of what Reuben's got on FP4.
- 26:24Less HPM capacity and less TDP. So
- 26:30I mean like under
- 26:32like almost three times less TDP per
- 26:35chip than Reuben, right? So if you just
- 26:39it's this weird thing where
- 26:41um when people are making the bull case
- 26:44for AMD, they go just look at MI450,
- 26:48right? It's going to have more FP4
- 26:51flops. It's going to have more HPM
- 26:52capacity. It's going to have more HPM
- 26:54bandwidth. It's going to be more TDP.
- 26:56and therefore it's going to be better.
- 26:59And they never want to look at the
- 27:00benchmarks. They never want to look at
- 27:01the like results of the chip, right? Or
- 27:05the MI355 when it was coming into
- 27:08production for the first time, right?
- 27:10But now you got an open AI chip which on
- 27:11papers is like objectively worse specs
- 27:14than everything on Reuben. Um comparable
- 27:19to GB300 on everything except for HPM
- 27:21bandwidth.
- 27:24uh and yet it's way outperforming GB300
- 27:27and
- 27:28outperforming Reuben so far. So like
- 27:31clearly this is due to the use of
- 27:35it's due to something about the micro
- 27:37architecture which we can get into and
- 27:40the software.
- 27:43So, like what's your take on what it
- 27:45takes to design a chip now? Because is
- 27:47it purely having access to the latest
- 27:49models and being willing to like
- 27:53yolo trust them on RTL and kernels? Like
- 27:58this is the barrier.
- 28:06Yeah, that's a good question and I I
- 28:09don't know the answer to that, but I
- 28:10think you know to your point, right? um
- 28:13you know everyone can deliver um great
- 28:18you know stacks on paper right you know
- 28:21when when we look at these stacks
- 28:22they're all you know pe theoretical um
- 28:26and I think a lot of emphasis on
- 28:29theoretical because I think that some of
- 28:31the flops numbers like no matter like
- 28:34how how you try to like reach them
- 28:37they're like impossible to reach so um
- 28:39it's it's somewhat determined by um you
- 28:42know the the chip company's market teams
- 28:45marketing teams but um yeah basically
- 28:48it's you know there are these flops that
- 28:50can you actually realize them in an
- 28:53actual workload same with HP you know
- 28:55bandwidth um I mean one of the big
- 28:59tenets of uh open's design philosophy
- 29:02forio is that
- 29:05um you know everyone can deliver raw hpn
- 29:09bandwidth you just like look you buy the
- 29:11the best HPM and then you put more
- 29:14stacks of it. Um, which I mean it's
- 29:17[snorts] it's buying HPM itself is is
- 29:20not easy these days, but um you know
- 29:22it's basically not super hard, right?
- 29:25It's not like it doesn't take um
- 29:27tremendous design skill. Um, so but
- 29:30really, you know, what's really
- 29:31throttling a lot of these chips is that
- 29:33they can't realize anything close to the
- 29:36raw HBM bandwidth because there's so
- 29:39many other uh things in the micro
- 29:41architecture that stop or stop you from
- 29:44doing that. And I think a lot of that is
- 29:45just um whether that be like very
- 29:48complicated memory subsystems or um
- 29:54or or just a lot of data movement
- 29:56required. Um whereas um the OKAI team
- 29:59has uh really been able has focused on
- 30:03um with mic microitecture that um
- 30:08reduces data movement so that um they
- 30:11can realize a lot of this you know the
- 30:13the the massive amount of page being
- 30:15damage they have. So I think that's
- 30:17really the main I guess skill um in all
- 30:20this and why for the competitors it's
- 30:24been really challenging to um you know
- 30:27they can deliver these all these all
- 30:29these like great specs but it's really
- 30:31difficult to actually realize them in an
- 30:34actual workload.
- 30:36>> Yeah. Yeah. And I agree with this like
- 30:38the fact that you can't actually reach
- 30:40this. I know there's there's a saying
- 30:42somewhere from someone I'm not sure who
- 30:46uh that the figures are like not numbers
- 30:49that you can reach and are numbers that
- 30:50the manufacturer can guarantee you never
- 30:53exceed like those flops which way to put
- 30:57it. Yeah.
- 30:58>> Yeah. Those are numbers that you are
- 30:59guaranteed not to exceed. And I think in
- 31:04one of the previous uh articles, I think
- 31:07in the cerebrous articles, they talk
- 31:09about the roofline models of chips
- 31:12compared to Nvidia. And although
- 31:13Nvidia's like flops are huge and like
- 31:16crazy especially the FT FP4 ones at the
- 31:20end of the day they are extremely uh
- 31:23right words on the roof lines in the
- 31:26computer bound region and a lot of
- 31:28workloads would rarely even reach those
- 31:31roof lines.
- 31:32So it doesn't matter at the end of the
- 31:34day the flops like the those flops in
- 31:36the table doesn't really matter because
- 31:39oops Jordan dipped again because the
- 31:42workloads most of the time does not
- 31:44actually hit those roof lines. Yeah,
- 31:48Jordan is gone. So maybe have some
- 31:50something curious to ask you about like
- 31:52the HPM. There's been talk about like
- 31:54Samsung HBM being better than the other
- 31:57HBM and like Halipino was
- 32:00uh lucky maybe to or like maybe it was a
- 32:03decision. I'm not so sure. I'm not a
- 32:05highway guy. Like what makes Samsung HBM
- 32:09like better or like uh better quality
- 32:12than the rest? Yeah. Yeah. So Samsung um
- 32:17for a long time you know from the HBM 3
- 32:20and 3 generation um Samsung's HBM you
- 32:23know was was really quite inferior to um
- 32:27Hinx which you know who dominated who
- 32:30you know still dominates HPM market
- 32:32share is you the leading supplier for
- 32:35Nvidia for instance um and you know
- 32:38really 3e that generation is really bad
- 32:40from Samsung um part of it was it was
- 32:43built um an inferior process, right? Um
- 32:48so the ite um
- 32:52Samsung, you know, they realized this
- 32:55and they really went all out on on their
- 32:58HBM core technology. Um so the DRAM dies
- 33:03are built on a more advanced 1C process
- 33:07whereas you know highix and micron are
- 33:10using 01B process the same which is the
- 33:14same process that HPN 3 is built on um
- 33:17they're also using there's a logic based
- 33:20in HBM cube that has the fi u and
- 33:25Samsung is built built this on a
- 33:28advanced logic process which is you know
- 33:30SF4 um Samsung Foundry 4 nanometer node
- 33:34um whereas Hinx is using 12 nanometer
- 33:38TSMC um and Micron is still using their
- 33:42own you know DAM process um for this
- 33:45base D even though you know the the
- 33:47bandwidth um requirements are
- 33:49significantly higher for HKM4 and that's
- 33:51sort of been um that's that's why you
- 33:54know Micron has had some issues with um
- 33:58achieving the this, you know, the
- 34:01highest speeds for HP4 and and similar
- 34:04with with Clinics, they've had some
- 34:06issues. um they've had to say redesign
- 34:09the base die and this is why um you know
- 34:12it it was only until like they've had to
- 34:15like delay shipments for um HPF core for
- 34:19for Nvidia's uh you know rub right um so
- 34:23Samsung has ended up being you know
- 34:25because of these you know I think it it
- 34:27seems to be clear that Samsung has
- 34:29actually the best technology for HM4
- 34:33um this is why you know this um Jalapeno
- 34:38is HPM4. It can deliver 15.4 terabytes a
- 34:41second of you know HPN bandwidth um
- 34:44which means the 10 GB per second uh pins
- 34:47speed the HPM4 they've gotten um this is
- 34:50a little bit higher than say the 9.6 six
- 34:52that we think basically your uh Nvidia
- 34:56will will ship Ruben with. Um and yeah,
- 34:59I think it probably does end up being
- 35:02that it's because of Samsung being um
- 35:05the supplier here. Um and you know the
- 35:07reason is also that whether it's it's
- 35:10luck or skill I think um we can debate
- 35:14about it but you know guess I guess
- 35:17traditionally um Broadcom's HBN has
- 35:21mostly come from Samsung right so I
- 35:24think it it's partly sort of that luck
- 35:27um this is this has hurt Broadcom for 3
- 35:30but um you know for four this is been I
- 35:34guess that's turned out pretty
- 35:36for that.
- 35:38>> Yeah. Yeah, that's very interesting. And
- 35:40yeah, this this story of like Samsung
- 35:41HBM and like uh because if I'm not
- 35:44wrong, SKH Highix invented HBM, right?
- 35:48>> Yeah. So SKH Highix um along with AMD
- 35:52were sort of they
- 35:54realized that um very correctly that um
- 35:59you know
- 36:01memory bandwidth is doesn't isn't
- 36:03scaling um with along with uh you know
- 36:07logic performance. So they needed
- 36:09something to really like bring out
- 36:11bandwidth and they came up with this HBM
- 36:13concept and originally it was for for
- 36:17gaming GPUs. So the first product with
- 36:21HPM was on a gaming GPU um with on on
- 36:26one of the uh AMD gaming GPUs. And but
- 36:30this this turned out to be a bit bit
- 36:32overkill but uh thankfully um you know
- 36:36they had this technology um for for
- 36:38another very memory bandwidth intensive
- 36:40application which is AI.
- 36:43>> Yeah. Yeah. It's interesting that like
- 36:46SK Hanx and AMD invented HBM, but now
- 36:49like I mean Samsung is doing better in
- 36:52HBM 4 and Nvidia is uh doing better than
- 36:56AMD. But on the topic of the HBM
- 36:59bandwave, so bandwidth is a problem,
- 37:02right? And is not capacity. Is that why
- 37:04like companies are going towards like if
- 37:07I'm not wrong six high and four high
- 37:09instead of a high?
- 37:12Um so
- 37:14I'd say yeah the the primary
- 37:18I mean it's in the name high bandwidth
- 37:20memory the the appeal is is the
- 37:22bandwidth um
- 37:25capacity is important but um I think the
- 37:28main sort of benefit is really bandwidth
- 37:30because you know there there are cases
- 37:33where you pay for the additional
- 37:34capacity but you don't um you don't you
- 37:37might not need it um whereas I think for
- 37:40memory bands um you can always just
- 37:42reuse that into serving tokens faster,
- 37:45right? Um so um [snorts] and bandwidth
- 37:47is really the key and the thing is like
- 37:50whether it's a 8 high, 12 high or four
- 37:52high stack um the bandwidth is the same
- 37:55but um because you pay for capacity and
- 37:59and that makes sense because you know
- 38:00more capacity more layers is what um
- 38:03adds to the cost of the supplier. um you
- 38:06know the sort of you pay for the extra
- 38:09capacity but the dollar per bandwidth um
- 38:12gets much worse. So um for some some
- 38:15some companies if they want to optimize
- 38:18um and say actually we just want to like
- 38:20get best dollar per bandwidth then going
- 38:22lower stacks is is the right
- 38:23optimization. Of course you want that
- 38:26balance between capacity and and
- 38:27bandwidth but uh you know I think there
- 38:31is a a philosophy especially now that
- 38:34HTM is getting much more expensive
- 38:36because um you know we have very limited
- 38:39um supply of HPM wafers. Um so I think
- 38:43that the trade-off is starting to look
- 38:46more in favor of going to lower stack
- 38:49heights rather than just increasing them
- 38:50further and further.
- 38:52>> Yeah, it's interesting. Yeah, it makes
- 38:54sense with me especially like uh
- 38:57increasingly more Rex scale
- 38:58architectures
- 39:00>> but yeah like capacity is not becoming
- 39:02as much of an issue anymore.
- 39:04>> Yeah, exactly.
- 39:06>> Yeah. And this is um I think
- 39:08specifically borne out in the per watt
- 39:11argument as well where um if you look at
- 39:14the raw HPM bandwidth comparing the
- 39:16specs of the chips, it's like okay, it's
- 39:18up there with the other ones, but then
- 39:20if you divide by the amount of uh power
- 39:23consumed by the chip, the fact that
- 39:25they're getting 15.4 turbines per second
- 39:27on a chip with a TDP of 700 watts is is
- 39:30incredible. Like this bandwidth per watt
- 39:33ratio like arbitrary units here of 22 is
- 39:36just so much bigger than anything else.
- 39:38It's literally double Reuben. Even at
- 39:40the max Q like low power setting option
- 39:45on Reuben to like run it at 1,800 watts,
- 39:48that's still like it's got a little bit
- 39:51more um HPM bandwidth, you know,
- 39:54whatever that is 25% more 20 versus 15
- 39:57terabytes per second, but it's literally
- 40:00more than double the power. So
- 40:04I mean that's that's where your if you
- 40:06can realize that bandwidth with that
- 40:08much power that's where your token
- 40:11output per watt advantage comes from
- 40:14right there. Um I guess the the like big
- 40:19question of course my is like
- 40:22B 0 is in the fab right now. They should
- 40:26just be able to step up the power and
- 40:28get even more bandwidth from this,
- 40:29right? Or are they at some limit?
- 40:34I think um I think on the HVM bandwidth
- 40:37they are probably at um limit. I think
- 40:42the
- 40:44uh and you know yeah I think that's sort
- 40:48of just the constraint of the memory
- 40:49itself. Um but the B 0 um it should
- 40:53deliver more flops um at at you know the
- 40:56same power basic. So um that you know I
- 41:00think depending on the workload um if
- 41:02there are compute constraints then then
- 41:05that should benefit um for the B zero
- 41:08stepping uh the A zero stepping is
- 41:11what's uh what the I guess all the
- 41:13current um results are.
- 41:17Yeah. Yeah. There are many different
- 41:21we're seeing so many chip startups
- 41:22explore the surface area of possible
- 41:25uh configurations of chips right now.
- 41:27But if you are optimizing for HPM
- 41:29bandwidth per watt, a metric that seems
- 41:32pretty relevant in
- 41:35LLM inference, this is the design to go
- 41:38with at this point, right? There's
- 41:40nothing else that compares
- 41:42that we've seen we've seen specs on um
- 41:46or that are public specs on, let's say.
- 41:49uh not giving away too much there. So,
- 41:54um
- 41:55maybe just to talk about realizing it a
- 41:57little bit more.
- 41:59Uh
- 42:01I don't know, Brian, do you want to talk
- 42:03a little bit about the software
- 42:04programming model and the micro
- 42:06architecture? You want me to talk about
- 42:07that?
- 42:10>> I think I think you're more
- 42:12knowledgeable in the aspect, right? But
- 42:15but actually be before I let you answer
- 42:17your own question,
- 42:19the another very big interesting point
- 42:21is like the role of AI in all of this
- 42:24right now you can you can make you can
- 42:27make an argument that open AI's biggest
- 42:29advantage is that it's able to access
- 42:31its own Astra models uh the new GP Astra
- 42:35models before anyone else. So like and
- 42:39that's that's that's like the biggest
- 42:40difference between I would say between
- 42:42open AI and one of the other new chips
- 42:45companies sova cerebras etc. Uh yeah the
- 42:49question is how much did at least my
- 42:52question is how much did Astra actually
- 42:54contribute to this like if you look at
- 42:56it from a differences point of view like
- 42:58this is only one of the only differences
- 43:00between uh
- 43:02uh Jalapeno and the other chips like
- 43:07so did Astra really contribute to like
- 43:11most of these performance differences or
- 43:14just a bit yeah but that's just that's a
- 43:17tangent of
- 43:18AI work on development and yeah actually
- 43:21Kim K3 when Kim K3 was released there
- 43:23was a part on the blog about its
- 43:26designing of a chip I forgot what what
- 43:29what was it about maybe some yeah I'm
- 43:32not sure what the architecture was about
- 43:33but yeah they did talk about K3
- 43:36developing a chip and yeah I would guess
- 43:39GPT extra does have similar capabilities
- 43:42and I would say to a better aspect or to
- 43:46higher degree
- 43:48Yeah. Sorry. Back back to you Jordan.
- 43:50Yeah. On the architect micro
- 43:52architecture.
- 43:53>> Yeah. Well, let I mean, let me comment
- 43:55on the AI assistance on the architecture
- 43:58cuz I think it's there's two ways in
- 44:00which AI clearly assisted the
- 44:03design and then the bring up of the
- 44:05chip, namely design and then bring up.
- 44:08So, on the design side, this clearly
- 44:11wasn't Astra because the RCL freeze was
- 44:14in July of last year.
- 44:16So, you know, from February to July of
- 44:19last year, when they claimed that AI
- 44:22assistance helped them get an 8%
- 44:24reduction in SIMD area and then a 10%
- 44:26reduction in the matrix engine area
- 44:29during design, I mean, this is preGPT5
- 44:33that we're talking about. So, um,
- 44:36conceptually like the models have to be
- 44:37getting better at RTL in the meantime,
- 44:41but they were already good enough to
- 44:43rapidly
- 44:45like accelerate the
- 44:47um really tedious human driven work that
- 44:51is RTL before a tape out, right? Um and
- 44:56so I I think you know maybe that's the
- 44:59biggest claim here. Uh which is that
- 45:04I think I think a lot of people
- 45:06understand you can use these models for
- 45:07kernels or and just software engineering
- 45:09because like you put it in a codeex
- 45:11harness you put it in a loop uh you let
- 45:14it test the thing and then you just set
- 45:16goal and like people have had that
- 45:17experience they can kind of understand
- 45:19it but I don't think a lot of people can
- 45:22like have had the experience of
- 45:23designing a chip. It's not like
- 45:24traditional software programming and um
- 45:28this was done with an older model. So I
- 45:31think that's like point number one. Now
- 45:34uh on the actual bring up yeah like kind
- 45:39of like what I just said
- 45:41the work is iterative and it's a
- 45:43verifiable domain. So like this is this
- 45:45kind of exactly what RL should be good
- 45:48at. You should be able to give a model a
- 45:50task of improving the performance of a
- 45:52kernel or getting the kernel to be
- 45:54functionally correct against some
- 45:57u like verification script
- 46:00um against some test cases, right? And
- 46:03then just let the model rip, let it try.
- 46:07And Dylan made this podcast on made this
- 46:09point on the Dores Cesh podcast that he
- 46:11was on recently, which is that for years
- 46:13I think now these
- 46:17uh the companies that were getting the
- 46:19most value out of using AI were not
- 46:21actually the companies providing AI.
- 46:23like OpenAI and Anthropic were not
- 46:25profitable for a very long time and now
- 46:28they're like just turning a profit but I
- 46:30mean still running incredibly high
- 46:32margin businesses but they're just like
- 46:33realizing the profitability of training
- 46:36these advanced models meanwhile you know
- 46:39Jane Street's going out there and
- 46:40printing $15 billion in a quarter
- 46:42clearly using AI for trading or
- 46:44something like that right and there's
- 46:46many other companies that are being
- 46:47started based on the use of AI um this
- 46:52is a very clear example of open AI
- 46:56keeping the benefits of having access to
- 47:00a model before everybody else for
- 47:02themselves. They can take out a chip and
- 47:04their competitors can't. And it's a sign
- 47:06of what's to come. I think that's the
- 47:09simplest way to put it. Um they're going
- 47:12to be able to go into many domains that
- 47:15are tangentially related to software
- 47:18where it's like the model needs to be
- 47:20able to control a computer, but it's not
- 47:23explicitly like the thing you're
- 47:24training it for.
- 47:28You build RL environments, you spend
- 47:30enough tokens, you spend enough like
- 47:32time on reasoning and
- 47:35enough rollouts, you know, enough like
- 47:37attempts at the um problem and you're
- 47:41going to get a good result. That seems
- 47:42to be the lesson here.
- 47:44>> Yeah. Yeah. Exactly. And I I think
- 47:46Entropic is also realizing the same
- 47:47thing. They are starting to hire like
- 47:50silicon people. that's on the same part
- 47:52as open
- 47:54and they are doing a lot of stuff like
- 47:56in the laboratory they got like LMS to
- 47:59control microscopes and whatnot recently
- 48:01and they're going seems to be going
- 48:04quite long into this like biological
- 48:05sciences field. Yeah. So I really agree
- 48:08on with you Jordan on the point of like
- 48:11uh these frontier model companies are
- 48:13realizing what
- 48:16this
- 48:18oops lots of interference from Jordan
- 48:21but yeah lots of good like downstream
- 48:24impacts of having a good model first can
- 48:27have not just like making money from
- 48:30inference revenue. Yeah,
- 48:38lots of interesting developments.
- 48:40>> What's your thoughts, dude?
- 48:43>> Yeah, I agree. I think um the the
- 48:47progress that Yeah, Labs is only like
- 48:50getting faster, right? And that's really
- 48:52because
- 48:53they're using their own models um really
- 48:56effectively to drive product innovation
- 48:59um much faster, right? Um, I remember
- 49:02like I think was it earlier this year
- 49:05like Anthropic was releasing a new
- 49:08product like every week or something.
- 49:09Um, and I think it was like everything
- 49:11was basically on autopilot. They were
- 49:13just using code for everything, right?
- 49:16Um, so yeah, I think this
- 49:19I agree with with Yeah. with what you
- 49:22guys have said.
- 49:24>> Yeah. Yeah. Um, yeah. Yeah, I mean like
- 49:27the cynical view of a program like this
- 49:29for both OpenAI and Enthropic and even
- 49:31Meta and Microsoft is that it's kind of
- 49:34like a head fake that gets them a
- 49:36discount on the Nvidia GPUs and
- 49:38therefore it pays for itself. Like you
- 49:40only need to spend a few hundred million
- 49:44on a program to like tape out a chip and
- 49:48you know like the year or two to do it
- 49:52to to potentially like um help with the
- 49:56negotiations and if those negotiations
- 49:58are measured to the tune of hundreds of
- 50:00billions of dollars then it pays for
- 50:02itself pretty quickly. But
- 50:06the nonsynical view is that like this is
- 50:11a real thing and it's only going to
- 50:13they're only going to do more of this in
- 50:14the future, right? Like
- 50:16um there's no reason that they're going
- 50:18to be less vertically integrated and
- 50:21less interested in developing chips this
- 50:24time next year. And there's no reason to
- 50:26say that the RTL time from freeze to or
- 50:31from like initial to freeze to tape out
- 50:35can't go even shorter than 9 months. Um
- 50:39and I I think you just need to to think
- 50:41about where to go from there. Maybe the
- 50:43other thing that was was kind of
- 50:45interesting here. So like if we go
- 50:47through the architecture, I mean there's
- 50:49lots to say about it, right? Um
- 50:53the maybe the quick high level is just
- 50:55like it looks like a TPU with much
- 50:58smaller um systolics uh systolic array
- 51:04being like the the way a TPU has its
- 51:07processing elements laid out. Um, and
- 51:12maybe the the criticism of chips like a
- 51:16TPU, tranium,
- 51:18uh, even some of the ones that are like
- 51:20TPU inspired, let's say like an etched
- 51:22or a Maddox X or something like that, is
- 51:24that these when they go with these
- 51:26really big systolic arrays, um, they can
- 51:29have these weird cliffs where small
- 51:32batch dimensions like the the M
- 51:34dimension in your MNK for a matrix
- 51:37multiplication gets all which you know
- 51:40is what happens when you have lots of
- 51:41experts and you have very few requests
- 51:44like low concurrency.
- 51:46Um
- 51:48you know the these like skinny gems,
- 51:52skinny matrix multiplications
- 51:54um can waste a lot of resources, right?
- 51:56or um even just odd numbers where if you
- 51:59go like slightly over 256 or slightly
- 52:02over 128, now you're spending an entire
- 52:05kernel launch on the device side just to
- 52:08run one little skinny gem. And uh all of
- 52:11these like
- 52:13uh processing elements on the systolic
- 52:16array are not being used and so
- 52:18therefore it's like inefficient and you
- 52:19don't actually maximize the flops on the
- 52:23um chip itself. Uh the trade-off here of
- 52:28course is that to get more efficiency
- 52:30with tiling to you know reduce the
- 52:33issues with like padding overhead or or
- 52:35uh alignment on the matrix dimensions is
- 52:38that you just use smaller systolics and
- 52:40that's what they've done here. And so I
- 52:42think that's been really smart clearly
- 52:44for for efficiency across the curve. The
- 52:46argument the other way is that they're
- 52:48going to miss out on some I think power
- 52:50efficiency and and like data movement
- 52:52efficiency because well you have to have
- 52:54more small elements instead of one big
- 52:57element. So what do you do then? And I
- 53:00guess the way that they've solved this
- 53:01is by um being really smart about how
- 53:05they place weights and KBs
- 53:08um using synchronization between cores
- 53:11like selectively and then saving the
- 53:14collective network the like knock the
- 53:16like network on chip that connects the
- 53:18HPM slices and the computing elements
- 53:20together really really sparingly. And
- 53:23the I mean the results is like well the
- 53:25results speak for itself and you see the
- 53:27performance there but the results is
- 53:29that this might be both a chip chip
- 53:32that's simpler to reason about than a
- 53:34GPU and a chip that's a little bit more
- 53:36flexible for some of these weird
- 53:39changing uh dimensions over time than a
- 53:42TPU. And so I mean clearly the guys who
- 53:46have experience using GPUs which OpenAI
- 53:48has plenty of experience programming
- 53:49GPUs mixed with the guys who have
- 53:51experience designing a TPU has resulted
- 53:54in a pretty well balanced system here.
- 53:57Um
- 53:59maybe the other thing is that it has an
- 54:01L1 cache which is quite funny like all
- 54:04of these accelerators do not have L1
- 54:06caches now. They rely so much on L2 um
- 54:09namely SRAM. we hear SRAMM all the time
- 54:12and um I mean the the reason why I guess
- 54:16is that uh they
- 54:22like the the companies
- 54:25designing other accelerators don't want
- 54:28to use a um scratch pad uh sorry they do
- 54:33want to use like a software manage
- 54:34scratch pad and OpenAI is not using a
- 54:36scratchpad cache here um and so that
- 54:39makes the chip like potentially harder
- 54:41to reason about when you think about
- 54:42like barrier latencies and where you're
- 54:44going to move data like you have to be
- 54:46able to amvertise the data movement by
- 54:49doing work on the CPU itself. But that
- 54:52ties in with, you know, what I was
- 54:54saying at the very beginning, which is
- 54:55like the second phase of using AI, which
- 54:57is um actually programming kernels on
- 55:01like meaning software that runs on the
- 55:02device side. And um the the to do this,
- 55:07I mean, we haven't really been able to
- 55:08verify this other than scrolling through
- 55:10a a 30,000line
- 55:13file with some of the OpenAI engineers
- 55:15when we went on site with them. It's
- 55:16like um literally
- 55:20it's just
- 55:22sloth. Like it's it's not sloth cuz it
- 55:25performs, but it's literally just AI
- 55:27generated assembly basically puked out
- 55:31in gluon this uh low-level like kernel
- 55:34programming languages that they built on
- 55:35top of Triton which uses this you know
- 55:38really interesting programming model.
- 55:40Like the the point is like the guys who
- 55:43were scrolling through this this code
- 55:44with us, it was kind of clear that like
- 55:46they know a whole bunch about hardware.
- 55:48They know a whole bunch about the
- 55:49concepts in the system and they just
- 55:51like have no idea what this MLA kernel
- 55:53that they're showing us for DeepSeek
- 55:55actually does. Like like you can go line
- 55:58by line and it's like nope nope nope. Um
- 56:02but it doesn't matter, right? And uh the
- 56:06AI understands it. the I tests it and
- 56:09you see the results. It it produces
- 56:11correct kernels that perform really
- 56:12well. And uh I think this is like just a
- 56:15sign of what's to come again, right? The
- 56:18um the like that actual
- 56:23code is not necessarily something that a
- 56:26human has to reason about deeply if the
- 56:29AI knows how to manipulate the data
- 56:31movement, the processing elements on the
- 56:33hardware that you've given it.
- 56:37Um, okay. We're kind of running out of
- 56:39time here. We've been going for a while.
- 56:41Uh, we got three things that I had in my
- 56:43notes that we wanted to talk about. Uh,
- 56:46they don't do PD disag.
- 56:49Brian, maybe you can rant about that
- 56:51because you spend like all of your time
- 56:52debugging PD disag.
- 56:57[laughter and gasps]
- 56:59Two, we didn't really talk about the
- 57:00system architecture. We can talk a
- 57:02little bit about how they do scale up,
- 57:04scale out domains. I mean they don't
- 57:06call it scale out but uh whatever it is
- 57:08scale out the multi-ter scale up stuff
- 57:12which is just like a total mess to try
- 57:15to understand probably can't communicate
- 57:16on a podcast go read the article and
- 57:19then the third thing is it's a
- 57:21generalized inference chip it's not
- 57:23co-designed with their models like they
- 57:25keep saying and the proof of that is it
- 57:28runs Doom
- 57:3036 frames per I can [laughter]
- 57:36um [clears throat]
- 57:39>> Yeah. [laughter]
- 57:42Yeah. We we I remember we were talking
- 57:45to these guys and we were like the like
- 57:47angry disappointed mother who comes in
- 57:50and they show us this like
- 57:51groundbreaking chip that's so fast and
- 57:53runs this stuff and they're like but
- 57:55does it run Agent X? No. [laughter]
- 58:0096% on the test. What four questions did
- 58:04you get wrong?
- 58:11Um anyway, anything you guys feel is
- 58:14left unsaid on the chip? Uh pretty
- 58:18exciting release, eh?
- 58:21Yeah, very um I'm yeah really excited to
- 58:24see where the road map goes next and
- 58:28also um you're excited to see I mean
- 58:31Brian mentioned this earlier but Brian
- 58:33uh Anthropic is your is building it or
- 58:37as um they're hiring it they're building
- 58:39a team to do it. Um I think this really
- 58:42sets a pretty high benchmark for
- 58:45Anthropic um to to meet or beat. Um but
- 58:50I think you know there's every reason to
- 58:51believe that um your anthropic could
- 58:55achieve a similar outcome. So very
- 58:57excited to see that as well.
- 58:59>> And they're they're hiring the team now,
- 59:01right? So it's only like what a month
- 59:03and a half until the RTL freeze and then
- 59:06like what four or five more months for
- 59:08the tape.
- 59:12[laughter] So like we should be able to
- 59:13get anthropic custom you know chip
- 59:16tokens
- 59:18um after like one
- 59:21you know 9 month cycle right like one
- 59:25[laughter]
- 59:26one pregnancy term
- 59:28[gasps]
- 59:31you know think
- 59:33yeah
- 59:36anthropics in
- 59:39anthropics hiring people right now so
- 59:41it's like the first trime trimester and
- 59:42then like the second one they they get
- 59:44the RTL freeze so they go to the second
- 59:47trimester and then like goes off to the
- 59:49fab and then like tape out it comes back
- 59:51third trimester all done right
- 59:59>> a little bit longer. Yeah,
- 1:00:02>> you guys don't like that one. Okay,
- 1:00:04we're going to end on one other joke. So
- 1:00:06Brian, you you like the one about
- 1:00:08converting energy usage to calories,
- 1:00:10right? So
- 1:00:14I want to finish this one. We uh we
- 1:00:17converted human we converted to human
- 1:00:19speech and compared uh some of this
- 1:00:21stuff on a on a calories, right? Because
- 1:00:23uh a calorie or like how much energy it
- 1:00:26takes to burn a what a cubic centimeter
- 1:00:30of water I think is uh the equivalent of
- 1:00:35uh whatever one jewel is 0239 food
- 1:00:40calories. So, we will uh we'll throw
- 1:00:42this this one up on screen to lead
- 1:00:44everybody off with a a little joke. Uh
- 1:00:47if you convert all of this efficiency of
- 1:00:49some of these DeepSeek results that we
- 1:00:52have on Asian X, uh we've identified
- 1:00:54that human speech is roughly 20 times 22
- 1:00:57times more energy efficient than the uh
- 1:01:00concurrency one B300 configuration that
- 1:01:02we were testing it against there. Right.
- 1:01:07A human speaks at 3.3 tokens per second,
- 1:01:09but these uh batch one configurations
- 1:01:11are going up at 180 tokens per second,
- 1:01:14much faster than the human brain can
- 1:01:16work, consume calories, and produce
- 1:01:18speech. So
- 1:01:20anyway,
- 1:01:21>> I think the cav the caveat is that um
- 1:01:25not all human spoken tokens are very
- 1:01:27high quality.
- 1:01:30[laughter]
- 1:01:32I mean,
- 1:01:35[laughter and gasps]
- 1:01:36>> I got to caveat some of my interactions
- 1:01:38with Claude, too. Then, meant some of
- 1:01:39this nonsense that's been spitting back
- 1:01:41at me recently, I've I've I've not been
- 1:01:43too pleased with either. [laughter]
- 1:01:48Can't see all the thinking traces
- 1:01:50anyway, but the like clawed version of
- 1:01:52English it's given me has not been that
- 1:01:53great. Um, [gasps] yeah. Well, uh, if
- 1:01:58we're comparing the machines on how many
- 1:01:59calories they're consuming per token and
- 1:02:01we're comparing it to our speech, we are
- 1:02:03really in competition with the machines
- 1:02:05at this point then, huh? Hopefully Jeb's
- 1:02:09paradox continues and and uh, everybody
- 1:02:11that produces chips wants to consume
- 1:02:13more tokens and produce more chips and
- 1:02:15hire more people and everybody gets to
- 1:02:16come have fun.
- 1:02:20All right, guys. Uh, good job. Long
- 1:02:23episode this time. I hope everybody
- 1:02:24enjoyed the overview of OpenAI Jalapeno.
- 1:02:27More to come. Last joke before we leave,
- 1:02:31cuz I just saw it. The best cover image
- 1:02:33in a while, I'd say. Right.
- 1:02:37[laughter]
- 1:02:38We'll leave that one on screen. Lisa and
- 1:02:40Jensen enjoying a nice spicy pot of uh
- 1:02:44chana or or katsu, whatever the uh code
- 1:02:48names were for the trays and the rocks
- 1:02:50and stuff in there. That's the
- 1:02:51motivation behind that.
- 1:02:53Vindaloo. Yeah,
- 1:02:54>> right.
- 1:02:55>> Sorry about that. All right, we'll sign
- 1:02:57off with that image in everybody's brain
- 1:02:59who's watching online. [laughter]
- 1:03:02Thanks for listening, guys.
About this transcript
This page contains the full transcript of Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators) by SemiAnalysis, generated from the public captions YouTube serves with the video. The transcript has 9,614 words across 1,428 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.