Bonsai 27B Deep Dive – 1-Bit, Ternary & Full Precision Compared! — Transcript
Full transcript
- 0:00Okay, nothing now.
- 0:05I'm done.
- 0:06Today we're going to be taking a look at
- 0:08the newly released Bonzai 27 billion
- 0:11parameter models, which are very
- 0:13interesting. Now, these come from Prism
- 0:15ML, and I have covered some of their
- 0:17models before in this Bonzai family. And
- 0:19essentially, what these are are massive
- 0:21reductions in size of the original
- 0:23models that they are based off off of,
- 0:26in this case being a Qwen 3.6 27B that
- 0:29has been shrunk down a couple of
- 0:30different ways, depending on what
- 0:32specific target device you would like
- 0:34this to be deployed on. So, in this
- 0:36announcement post, they mentioned both
- 0:38running this on a computer, whether it
- 0:40be a Linux computer, PC, or whatever, or
- 0:42a Mac. Additionally to that, they also
- 0:45talk about running these on phones.
- 0:46Specifically, they do mention for an
- 0:48iPhone 17 Pro, it is able to be run on
- 0:52at decent speed. I do have this running
- 0:54currently on an Asus ROG Phone something
- 0:56Pro. So, we're going to not only be able
- 0:58to check the performance of the one that
- 1:00is designed to run on a PC, but also the
- 1:03one that is running on a phone. As a
- 1:05matter of fact, I do have both of those
- 1:07set up and spinning right here. So, this
- 1:08one is the one that is running on a PC,
- 1:10as we could probably guess by the token
- 1:12speed. And this tab right here is the
- 1:14one that's running on the mobile phone.
- 1:16So, it's actually kind of cool to be
- 1:18able to talk to something through the
- 1:19web chat interface right here of Llama
- 1:21server, actually knowing it's running on
- 1:23my mobile phone sitting next to the
- 1:25laptop. So, before we get into it,
- 1:26please do feel free to subscribe so I
- 1:28can get that 100K plaque. And with that,
- 1:30let's begin just by taking a quick look
- 1:32at the options that we have for today in
- 1:34terms of these Bonzai models. So, Bonzai
- 1:3727B comes in two variants. The first is
- 1:40a ternary Bonzai 27B, and that is the
- 1:43one that is running right here in this
- 1:44tab on port 8080 at 65 or so tokens per
- 1:48second. And that is a smarter but larger
- 1:50version of the Bonzai model right here,
- 1:53where it uses ternary weights with FP16
- 1:56group wise scaling, giving a true 1.71
- 1:59effective bits per weight. At 5.9 GB,
- 2:03this is the quality-oriented
- 2:05quality-oriented variant. It runs on an
- 2:07everyday laptop with full reasoning,
- 2:09tool calling, and agentic capability.
- 2:11So, when doing some comparisons, we will
- 2:13really mainly want to see how this
- 2:15specific one, the ternary, stacks up
- 2:17with the FP16 version of the model it's
- 2:20based on, Qwen 3.6 27B. The next model
- 2:23version for this Bonsai 27B is the
- 2:261-bit, and this is the one that is going
- 2:28to be running on the mobile phone, which
- 2:30we can see is at port 8081 here, and
- 2:32running a bit slowly. However, this uses
- 2:35binary weights with the same group-wise
- 2:37scaling, giving an effective 1.125 bits
- 2:41per weight at 3.9 GB. So, that is a
- 2:43massive size reduction when considering
- 2:46that the full weights for the Qwen model
- 2:48it is based on are, as they say right
- 2:50here, 54 GB or somewhere around there.
- 2:53So, this can fit within the memory
- 2:55budget of an iPhone 17 Pro, bringing a
- 2:5727 billion parameter class model onto a
- 3:00phone for the first time. I don't know
- 3:01if I'd say first time just based off of
- 3:03this Xpost I made a few months ago, but
- 3:05nonetheless, I am just being a bit
- 3:07sarcastic. So, I I don't want to get
- 3:09into like a big technical deep dive on
- 3:11how these work, because they're very
- 3:12complex and I don't know how much
- 3:13service I would do in trying to explain
- 3:15that. I will say, however, at the bottom
- 3:17of this announcement post right here,
- 3:19they do have a link to their white
- 3:20paper, which will open it for you in
- 3:22GitHub, and you can go through and take
- 3:23a peek at it. It has a lot more
- 3:25intricate information. No, the actual
- 3:27methodology that they're using
- 3:29specifically to reduce the size of these
- 3:31models is not open source. So, just make
- 3:34note of that as well, but it does tell
- 3:35you some more pertinent information
- 3:37about everything that's going on here.
- 3:39Additionally, just to quickly touch upon
- 3:41it, and I'm maybe a little rusty, but
- 3:43from my recollection of doing some
- 3:45videos pertaining to their models
- 3:46before, essentially, the weights are
- 3:49shrunk down in in of 128, whether it be
- 3:51ternary or binary, and then that group
- 3:54of 128 weights has like a block-wise
- 3:56scaling factor that is FP16, I believe
- 3:59is the case. Again, don't quote me on
- 4:01that 100%, but basically they take like
- 4:03groups of 128 weights, and they become
- 4:06either ternary or binary depending on
- 4:08which specific model it is. And then a
- 4:10group of 128 weights gets like a scaling
- 4:13factor applied to them that is a higher
- 4:15precision than ternary or binary. Um
- 4:18something like that. Now, the models are
- 4:20available on Hugging Face in a number of
- 4:22different formats, and they do have some
- 4:23additional interesting things to note
- 4:25here. They have shipped this with a
- 4:27compatible speculative decoding setup as
- 4:29well, so that will speed up, as we see
- 4:31right here, 1.34 times decoding speed on
- 4:33the CUDA serving path. Additionally,
- 4:35they have paid respect to MLX, so Apple
- 4:38devices also have compatibility. Now,
- 4:40there's some interesting things
- 4:41contained here within the Hugging Face
- 4:43model card as well, because I find with
- 4:45these models, sometimes it can be
- 4:47confusing because it's like, "Okay,
- 4:48well, if the model's 5.9 GB, then if I
- 4:51open NVTop right here, why is my
- 4:53computer using 10.7 GB to serve this
- 4:56model right now?" And they mention that
- 4:59and explain it a couple of different
- 5:00ways. So, one of those is the kernel is
- 5:03storing the ternary values in a bit
- 5:05larger than what would be the ideal
- 5:07size, which is that 5.9 GB.
- 5:09So, the deployed size they say is around
- 5:117.2 GB. Following that, and additionally
- 5:14in terms of why this is using a bit more
- 5:16weight or VRAM than one would expect
- 5:19even with the deployed size, they do
- 5:21have a chart here just showing the
- 5:22actual memory footprint of this deployed
- 5:25over a number of different context
- 5:26lengths. So, we can see that 10K
- 5:28contexts with this model that we're
- 5:30running right here would be around 8.7
- 5:32GB of VRAM, and 100 would go up to 14.7.
- 5:35So, keep that in mind, and if you wanted
- 5:37to run this at absolutely full context
- 5:39of 262K and change, it would go up even
- 5:43higher than that. So, keep that in mind
- 5:45that these are not going to be using
- 5:47that 6 gigs of VRAM to run this 27B
- 5:50model, but it's a
- 5:53step closer to getting there. So, I want
- 5:55to do a really proper comparison test to
- 5:57showcase these models because a lot of
- 5:59folks, the interest is going to be in
- 6:01how well do these actually stack up to
- 6:03the original full precision version of
- 6:05this model. So, in this tab right here,
- 6:07we are running the binary version of the
- 6:09model. This is running entirely on a
- 6:11mobile phone that is just attached to
- 6:13this computer just so I can get power
- 6:15and things like this. Given a very
- 6:17simple prompt. Now, in the middle tab
- 6:19right here, this is running on the
- 6:21laptop that I'm using to control this
- 6:24entire demo. This is a 5090 mobile and
- 6:27this is using the ternary version of
- 6:29this model. So, the more performant one
- 6:31as we see right here, that is the
- 6:33ternary that is used in the middle tab
- 6:35and it is the binary that is used in the
- 6:38left most tab. Now, all the way in the
- 6:40right, this is running the full
- 6:41precision version of this model. It's
- 6:43not quantized or anything and this is
- 6:45running on a 6000 Pro Blackwell just on
- 6:48the computer behind me. So, we're going
- 6:50to be able to see proper side-by-side
- 6:52results for the two Bonsai versions of
- 6:55this model as well as the full-blown
- 6:57full precision original model. That way
- 6:59we can just get some more educational I
- 7:01suppose value out of seeing these
- 7:03demonstrations. Now, for this first
- 7:05test, I have disabled reasoning for all
- 7:07of these because the mobile phone one is
- 7:09extremely slow. The one running on the
- 7:12Blackwell is not super slow, but it can
- 7:14end up taking quite a while and then the
- 7:16one running on this computer right here
- 7:18is rather quick and I'm not using the
- 7:20MTP
- 7:21thing that is included with it. So, that
- 7:23would speed things up a bit, but for now
- 7:25we're just going to see the results we
- 7:26get for a simpler Steve's PC Repair
- 7:29website. All right, so all of our simple
- 7:31website comparisons have concluded. I do
- 7:33have to say I somewhat regret opting to
- 7:36run the one-bit binary off of a phone
- 7:38because that took a little under 3 hours
- 7:40to generate this website
- 7:42At the highest total token count for
- 7:44this result, so it was an overall speed
- 7:47of 1.56 tokens per second. Perhaps let's
- 7:49not focus too heavily on the speed
- 7:51metric there um for that specific phone
- 7:53and that specific model. Nonetheless, we
- 7:55have our ternary version in the middle
- 7:58tab, which took 4 minutes and 23 seconds
- 8:00at a speed of 58.68 tokens per second on
- 8:03a 5090 mobile not using the additional
- 8:06draft model that they have for this.
- 8:08Then finally, we have the full-fledged
- 8:10full precision version, which ran at
- 8:1228.73 tokens per second on a Blackwell
- 8:15RTX 6000 Pro, not the Max-Q like the
- 8:19full non-power limited one. And that was
- 8:22significantly less tokens, so I need to
- 8:24make sure everything's all set here with
- 8:25the sampling parameters. But
- 8:27nonetheless, I do have these all saved,
- 8:28so let's take a look at them. So, let's
- 8:30just start with the binary one as we
- 8:33want to I guess get a feel for the
- 8:34smallest first.
- 8:36That's
- 8:37honestly
- 8:38not bad for a one-bit binary version.
- 8:41This is actually decently aesthetically
- 8:43pleasing. Yes, there's some oddities to
- 8:45it such as the randomly scattered around
- 8:47things. Trusted by 10,000 plus clients.
- 8:50That's quite impressive. We do have
- 8:51hover effects on the buttons. And did
- 8:53you see the way these metrics actually
- 8:55had effects in terms of when they
- 8:57reached their final number? That's
- 8:59actually pretty good. We do have hover
- 9:00effects on these cards.
- 9:03Steve's PC repair, 15 years expert, and
- 9:05that would ideally be a photo of Steve.
- 9:08Book an appointment, how we fix your PC,
- 9:10what our clients say, transparent
- 9:12pricing. Not bad. These pricing cards do
- 9:14look quite good. Ready to get your PC
- 9:16fixed. And we have a nice clean
- 9:18good-looking contact card. And the
- 9:21footer's very well done, too. All right,
- 9:23so that's
- 9:24That's kind of cool. I mean, like this
- 9:25entire thing was generated just using a
- 9:27mobile phone running the binary version
- 9:30of this model, which is quite awesome.
- 9:33Nonetheless, let's take a look now at
- 9:36the ternary one, which was like the
- 9:37middle tab, the middle intelligence to
- 9:40size one.
- 9:42Okay. Ooh.
- 9:44And we'll notice that throughout this,
- 9:45there going to be similarities because
- 9:47it's the same model, but there are
- 9:49differences in like how hard it goes.
- 9:51This one's only trusted by 2,000 clients
- 9:54as opposed to 10,000 plus, so this Steve
- 9:56is a smaller business. We do actually
- 9:58have some interactive particle effects
- 10:00in the background.
- 10:01Hover effects here and a more kind of
- 10:03neon like synthwave aesthetic, which I
- 10:06personally am a big fan of. Okay, our
- 10:08pricing cards Oh, no, these are just
- 10:10services and their specific prices. Not
- 10:13bad, good hover effects, and they all
- 10:14have some nice transparency to them.
- 10:16About us, okay, established 2009. All
- 10:19right, I don't know why this just came
- 10:21to mind, but like a random like side
- 10:22note,
- 10:23back a long time ago, I used to think
- 10:25this EST meant like estimated, and I
- 10:28would always be like, how do they not
- 10:30know like the date these businesses were
- 10:31created? And I realized that it meant
- 10:33established. So, that clarified that for
- 10:36me. Just, you know, little
- 10:37um side note. Okay, we have a
- 10:39good-looking same simple repair process.
- 10:42I'm looking to see if there's a good
- 10:43aesthetically pleasing pricing card like
- 10:45we got with the binary version. We do
- 10:48have fake customer testimonials. Very
- 10:50good.
- 10:51Okay, no, and it just goes down to a
- 10:53contact card. Interesting though,
- 10:55there's no And the footer looks good and
- 10:57everything. Designed with love and
- 10:59coffee. The other one was designed with
- 11:00love and solder. Now, finally, and
- 11:03again, I do need to check the sampling
- 11:04parameters because this one was so
- 11:06suspiciously smaller in terms of overall
- 11:08generation length than the other two.
- 11:10Oh, goodness.
- 11:12Okay. So, this is the full precision
- 11:14one. Steve's PC repair, system
- 11:16optimized, initiate repair. Good hover
- 11:19effect.
- 11:20Service modules. Okay. And they all have
- 11:23a similar kind of thing. 2K systems
- 11:25fixed. So, this one has the same stats
- 11:27as the ternary one. Okay, now, something
- 11:29There's no way. Something has to be
- 11:32a bit off here because
- 11:34the full precision one should have not
- 11:36been this bad. All right, next up I have
- 11:38enabled thinking for all of these and I
- 11:40do believe that the behavior we saw
- 11:42where the full precision version of the
- 11:44model was actually a much shorter
- 11:45script. I think that's behavioral
- 11:48because we had thinking off. So, it is
- 11:49possible that the bonsai models as part
- 11:52of the way they're trained in order to
- 11:54retain their intelligence are given like
- 11:56front-end examples that have a lot of
- 11:59aesthetically pleasing things in them
- 12:01and that could be attributed to why
- 12:03those were so much longer. Nonetheless,
- 12:05don't quote me on that, but I just
- 12:06wanted to give some at least thought on
- 12:08whether or not that was
- 12:10normal behavior or not. So, now with
- 12:12thinking on I would imagine we will
- 12:14begin to see hopefully some more
- 12:15difference in the favor of the full
- 12:17precision model here, but nonetheless,
- 12:19this is a bit more difficult and will
- 12:20give us some more interesting demos.
- 12:23This is of course a 3D subway station
- 12:25first-person shooter test. Now, the
- 12:27binary and ternary are both being run on
- 12:30this 5090 mobile laptop. So, if the
- 12:32speeds are not as fast, it's because
- 12:34this is handling both of these here at
- 12:35the same time. And then the full
- 12:37precision is being run on the 6000 Pro
- 12:40Box. So, we'll
- 12:42get our results at some point. I'm happy
- 12:44to see though and something I was a
- 12:45little worried about is whether or not
- 12:46the smaller models would be inclined to
- 12:49just overthink massively. Fortunately,
- 12:51we didn't see that. The reasoning was
- 12:52actually fairly concise and that's
- 12:54always definitely a risk. I do believe
- 12:57when
- 12:58Okay, I just saw you're dead. Okay,
- 13:00that's good.
- 13:01It's good to see that these didn't
- 13:02overthink massively. All right, let's
- 13:04start with the binary subway FPS. Okay,
- 13:06subway station survive the infestation.
- 13:09Now, the mouse disappears as soon as it
- 13:11gets on the page. Sometimes when this
- 13:13happens, we just need to find the enter
- 13:15station button and it will still let us
- 13:17in.
- 13:19So, we'll hold off on that.
- 13:21Let's just take a look at the ternary
- 13:23version model versions. Okay, subway
- 13:26station zombies. Same weird cursor
- 13:28glitch issue.
- 13:29And the same complete failure of the
- 13:32start button actually doing anything at
- 13:34all. Okay, well, this is going quicker
- 13:36than I anticipated. Finally, let's try
- 13:38the
- 13:40the full precision version and hope that
- 13:42it's not a disaster like those.
- 13:46Are you
- 13:48Oh, okay. All right, so yes, this is
- 13:51makes a bit more Wow, this is quite
- 13:53good.
- 13:54This model is really like a just
- 13:58punches above its weight to a very high
- 14:00degree.
- 14:01At least in my testing.
- 14:03Okay, we're getting some lag pretty bad,
- 14:04but that's okay. This computer does have
- 14:06about like 20 gigs of VRAM hold up right
- 14:09now. Okay, so this is the only one that
- 14:12actually worked. Definitely showing us
- 14:14some
- 14:16differences. Actually, look at these
- 14:18like these models are not even that bad.
- 14:22This model never fails to amaze. There
- 14:25are rumors that there could be like a
- 14:26Qwen 3.8 coming out. Don't quote me on
- 14:28that, but that'd be really exciting for
- 14:30the local enthusiast among us. Good. All
- 14:33right. So,
- 14:35let's take a look and see what specific
- 14:36errors these are having. Okay, so that
- 14:38just has a syntax error and that was the
- 14:41ternary one. Let's check the binary one.
- 14:44That also has a syntax error as well.
- 14:46So, both of them just kind of weren't
- 14:49all there together and the
- 14:5116-bit obviously was. All right, next up
- 14:54we're going to try some
- 14:56agentic coding with these models because
- 14:58I want to see what we get for that. And
- 15:01I think that's going to show a bit more
- 15:03insight beyond just the single file like
- 15:05zero-shot generations that we're doing,
- 15:07even though those are fun. However, I am
- 15:09going to base our agentic coding test
- 15:12just on the errors that these two
- 15:13models, the binary and ternary, have for
- 15:15the subway FPS test. So, we're I'm going
- 15:18to give them the specific error that is
- 15:20showing up for either of their
- 15:21respective generations starting with the
- 15:23binary model where we have this uncaught
- 15:26syntax error and we'll see what it does.
- 15:28I need to change it. That is currently
- 15:30the ternary. So, if we go into models, I
- 15:32do have the configs all set up here.
- 15:35And I'm specifically not telling it
- 15:37in what script because it's in the
- 15:39dedicated directory which only the
- 15:41script is in. So, that's kind of an
- 15:43additional test is to see if it
- 15:44understands, okay, I need to fix this
- 15:46issue that's happening in the script
- 15:47that is in the directory I'm being run
- 15:49from within. Ignore the Q1 lingo there.
- 15:52It's not necessarily accurate for the
- 15:54way we're representing this.
- 15:57Found the issue on line 953. There's
- 15:59Okay, and then it changed it before.
- 16:02Yeah, all right. That would
- 16:04That's likely going to do it assuming
- 16:05that there are no more errors beyond
- 16:07this one. Okay, it's compacting a lot.
- 16:10It's running at 64K context length, but
- 16:12for now we're going to just assume
- 16:14because it does seem like it actually
- 16:15did remedy that one issue.
- 16:18Okay, so we have another one
- 16:19unfortunately.
- 16:20But it did fix the issue that was
- 16:22happening at least in that it's no
- 16:23longer appearing here. All right, so
- 16:25I've changed it so the context length is
- 16:27maxed out to 262144
- 16:30and I'm opening it back up from the same
- 16:32directory within the same model and I'm
- 16:34giving it the new error that it has
- 16:36received here and we'll see. My fear is
- 16:39that we may just
- 16:41have a bunch of errors to get through
- 16:42going back and forth where it will fix
- 16:44one like it did before and then another
- 16:46one will pop up and then we'll have to
- 16:47fix that. But nonetheless, we'll give it
- 16:48some time to see if we can get this to a
- 16:50functional point. That was very quick.
- 16:52So, let's refresh it.
- 16:54Okay, we do have another one and like I
- 16:57said, I do have some concern that it may
- 16:59just end up
- 17:01being a bunch of this, but we'll push it
- 17:03till we can't anymore.
- 17:05I will say even if we don't get this to
- 17:07a functional point, it is
- 17:09inspiring to see that it's seeing these
- 17:11errors and quickly fixing them and it's
- 17:13not getting stuck in loops or
- 17:15overthinking massively which I find to
- 17:17be pretty impressive.
- 17:19Okay?
- 17:25Okay, we're not getting any errors.
- 17:30Good. It got us to a point where we are
- 17:33actually in the game. Now, yes, it's not
- 17:35working fully and we're getting more
- 17:36errors now, but the thing is it got us
- 17:38past that first blocker, which I'm happy
- 17:41about. So, I'm going to continue just
- 17:42pasting these errors in. Ideally, I
- 17:44think the best thing to do would be to
- 17:45tell it like, "Okay, listen, I'm sick of
- 17:47pasting these errors in. I need you to
- 17:48build a little pipeline for yourself so
- 17:50you can autonomously pull these errors
- 17:52from developer tools, fix them, and then
- 17:54test again, and then fix, and then keep
- 17:56iterating until it doesn't give us
- 17:58errors anymore." No errors yet.
- 18:02Okay. So, it seems like this may be
- 18:05where we end with this. Where if we
- 18:08actually restart it, look up at the top
- 18:10right, and let me close the developer
- 18:11tools for now. You're going to see our
- 18:13ammo starts at 50, and if I turn the
- 18:15speaker up all the way, and I click, it
- 18:17will go down to 49, and we'll hear a
- 18:19sound.
- 18:21Unfortunately, then after that, nothing
- 18:23seems to happen. Nonetheless, it did
- 18:25take us to a point where it fixed all of
- 18:27the issues, and the ones that are left
- 18:28over are issues in actual like
- 18:31implementation of the game, and not
- 18:33things that are going to be shown to us
- 18:35just from the developer tool console
- 18:36there. So, I'm satisfied with what I saw
- 18:39there, and the fact that it did actually
- 18:40handle a bunch of different errors from
- 18:42within the same thread, and didn't freak
- 18:44out and didn't start overthinking. It
- 18:46fixed them all one by one, which I think
- 18:48considering like what this is is rather
- 18:50impressive. All right, next up I have
- 18:51the ternary model loaded in through Open
- 18:53Code here, and this is with a 131 in
- 18:55change context length, because um
- 18:58getting it to the full one, I don't
- 19:00think will be necessary here.
- 19:01Additionally to that, it also would just
- 19:03tap out the card. I have neglected to
- 19:05actually tell it like, "Fix this error."
- 19:08I just ended up pasting it in, but it's
- 19:09the same situation where it's only in a
- 19:11repository with the specific script it
- 19:13created here. So, it will very likely
- 19:15understand, "Okay, I need to figure this
- 19:16out based off of what's in this
- 19:18repository or folder, I should say."
- 19:20Found the bug and fixed it. I would
- 19:22assume we'll have a couple more,
- 19:24probably. I need to remember that I
- 19:26moved the folders.
- 19:28Okay, nothing now.
- 19:34I'm done.
- 19:35>> [laughter]
- 19:37[snorts]
- 19:37>> Well, that was unpleasant. So, all
- 19:39right. I mean, yes, we Oh, wow. Okay. Uh
- 19:44yeah, I'm just going to
- 19:46I'm just going to paste that in because,
- 19:47quite frankly, I don't even have the
- 19:48words to describe the specific issue
- 19:51here without probably using language
- 19:53that is not appropriate for this
- 19:54platform. So, we'll see what happens
- 19:56here.
- 19:59Nonetheless, it did have one specific
- 20:00error and then it fixed it. And we did
- 20:03get the game working. So, a quicker and
- 20:05more
- 20:06complete result than we saw with the
- 20:08binary. Um also more jarring and
- 20:11stress-inducing. All right,
- 20:12unfortunately, we've seemingly just hit
- 20:14a bit of an impasse here. Nonetheless, I
- 20:16was satisfied with what we saw because
- 20:18it did fix the one error that was
- 20:19blocking it from actually showing us the
- 20:21game, and the game was a bit more um
- 20:24playable than
- 20:26the
- 20:27one-bit one. So, it was This was very
- 20:30proper in terms of showcasing like the
- 20:33one-bit was just didn't really do
- 20:35anything. The ternary kind of worked,
- 20:37and then the full-precision one was
- 20:39arguably fantastic considering the size.
- 20:41So, uh interesting, and I wanted to
- 20:43throw in some agentic coding there just
- 20:45at least for the binary and ternary ones
- 20:47to see if they can handle open code, and
- 20:50they could. All right, so, the next
- 20:51thing I did is a bit different, but I
- 20:52know folks have often times, especially
- 20:54for models like this, wanted to know how
- 20:56they do with existing code bases. So, I
- 20:59gave them each access to a simple
- 21:01diffusion demo that I have where it's
- 21:04essentially a 3D model of a keyboard,
- 21:06and a small trained diffusion model
- 21:07predicts out of potential keyboard
- 21:10layouts like QWERTY, Dvorak, et cetera.
- 21:12I had to use that for my Gemini
- 21:14diffusion model video. So, the prompt
- 21:16here was essentially to give these each
- 21:19some questions about this specific code
- 21:21base in an isolated copy via open code
- 21:23with identical sampling parameters and
- 21:25ask them some targeted questions about
- 21:27this repo and then just grade how they
- 21:30did in terms of their quality and depth.
- 21:32So, apparently they all passed five of
- 21:34five. However, we do have some more
- 21:36fine-grain detail in terms of the actual
- 21:38specific questions that were asked and
- 21:40where things started to fall apart in
- 21:42terms of like
- 21:44wrong answers. So, let's just go through
- 21:46our per question details. So, question
- 21:48one was to describe the full training to
- 21:50browser pipeline and what happens if the
- 21:52exported weights fail verification. Then
- 21:55we have our expected answer right here.
- 21:58Reports the failure in the HUD and the
- 21:59demo keeps running on exact analytic
- 22:01base. So, exactly there was a fallback
- 22:04there where it would just kind of fake
- 22:05what would happen were there a properly
- 22:07trained model and that is the answer
- 22:09we're looking for. Complete chain with
- 22:11correct line numbers also identified the
- 22:13separate fetch failure fallback path
- 22:15beyond the rubric. So, that is the full
- 22:17precision version of the Qwen 27B model.
- 22:20The ternary version complete chain
- 22:23including exact HUD failure message and
- 22:25no wrong predictions just a fallback
- 22:27called the JS engine a 200 line source
- 22:29comments as 100 trivial and then binary
- 22:32correct chain and threshold for the
- 22:34maximum error. Correct fallback least
- 22:36detail on what the UI reports. So, all
- 22:38of these did actually get that correct,
- 22:40which is cool to see. So, question two,
- 22:42exactly how many training sequences,
- 22:44what are they and how is a batch of 64
- 22:46built? The expected answer is exactly
- 22:49four sequences, the 40 character layout
- 22:51strings sampled with replacement via
- 22:53torch random integer to 64 rows. Each
- 22:56row independently gets up to this per
- 22:58position masking with a probability of
- 23:00this masked position enforced. So, our
- 23:02BF16 got quoted all four exact strings,
- 23:06described clamp subtly wrong. The clamp
- 23:08creates a point mass here, not uniform,
- 23:10only modeled to make that claim
- 23:12precisely enough to be wrong.
- 23:13Interesting. Ternary, quoted all four
- 23:16exact strings, noted one mass guard and
- 23:19one T waiting, didn't mention the clamp.
- 23:21Okay.
- 23:23Then again, these are I mean, this is
- 23:24not necessarily my favorite way of
- 23:26measuring things, but it is important
- 23:28for models like these just to see.
- 23:29Binary, named the four layouts, didn't
- 23:31quote the strings, got with replacement
- 23:34sampling, mass guard. Okay. So, they all
- 23:37overall got that kind of correct, which
- 23:39is cool, but it was interesting that it
- 23:40mentioned the full precision one only
- 23:42model to make that claim precisely
- 23:44enough to be wrong. Question three, why
- 23:46is probs for recomputed inside the
- 23:48commit loop? What breaks without it and
- 23:50where does the JS handle it? Expected
- 23:53ancestral sampling, each committed token
- 23:55must condition subsequent value
- 23:56sampling. The model is bidirectional,
- 23:59refusing steps hard marginals, let's
- 24:01position sample from incompatible
- 24:02layouts, incoherent output and the demos
- 24:04exact posterior zeros out. JS
- 24:06counterpart and then we have some
- 24:08additional info right here.
- 24:10BF16, textbook answer, correct
- 24:12citations. Good, named commit key
- 24:14explicitly. Ternary, correct mechanism
- 24:17and named commit key uniquely added that
- 24:19disk cache is keyed on committed state,
- 24:22so self invalidates. True and beyond the
- 24:24rubric, but cited lines 304 for commit
- 24:26key and line 217 elsewhere, wrong
- 24:29locations. Then binary, correct concept
- 24:32with a concrete two-position example,
- 24:33identified the JS mechanism, but never
- 24:36named commit key. So, we see different
- 24:38levels. Interesting, the ternary one
- 24:39there seems to
- 24:41be like fairly robust, even if it got
- 24:44some line numbers wrong. So, question
- 24:46four, every mechanism that keeps mask
- 24:48out of the generated output and then the
- 24:49expected, we have a bunch of Python
- 24:51information here and then what specific
- 24:53like math would
- 24:55cause that to happen. Skips mask ID in
- 24:57the soft max and only enumerates real
- 24:59chars real characters. Bonus the
- 25:01analytic base path can only propose
- 25:03layout characters by construction. BF16
- 25:06all required guards with correct lines
- 25:08plus the analytic path bonus cleanest
- 25:10answer of the five questions. Ternary
- 25:12both required guards correct extra
- 25:14claims about top of commit flow correct
- 25:16in substance but wrong line numbers
- 25:18again. Interesting that it seems to be a
- 25:20consistent thing here that it's getting
- 25:21the wrong line numbers. Then finally
- 25:23binary both required guards correct but
- 25:26credited posterior from is guard that
- 25:28function computes the layout posterior
- 25:30not character distributions. The
- 25:32analytic character path lives in get
- 25:33distributions misattribution. Okay. Then
- 25:36finally
- 25:37question five Why can the layout
- 25:39posterior panel disagree with hover tool
- 25:41tips when the net is active? Expected
- 25:43the panel always computes the exact
- 25:45analytic base posterior from as ground
- 25:48truth. Tool tips ghosts use get
- 25:49distributions which switches to the
- 25:51train net. The disagreement is the point
- 25:53comparing the net equals base. All
- 25:55right. So BF16 complete correct
- 25:57citations and articulated the layout
- 25:59level versus character level distinction
- 26:01plus what divergence reveals. Ternary
- 26:03complete correctly traced the net base
- 26:06commit comment to sample.py. Good
- 26:09framing of learned approximat
- 26:11approximation versus exact reference.
- 26:13Binary correct and concise cited the
- 26:15design comment line 90 actually 89 but
- 26:18trivial. Okay. Pretty interesting. They
- 26:20actually seem to perform fairly strong
- 26:22on this. Now this is not a
- 26:24super complicated code base. I mean if
- 26:27we look at it right here it is in the
- 26:28diffusion folder. We can see there's not
- 26:30too much to it but it is a tiny little
- 26:32diffusion model trained to predict the
- 26:34layouts of a keyboard based off of like
- 26:36potential ones that appear. Then it
- 26:38re-masks and then it makes new guesses
- 26:40based off of what's been unmasked and
- 26:42things like diffusion. So it's not
- 26:44necessarily like a super trivial like 3
- 26:46JS scene or front end or something of
- 26:48the sort. So that's cool to see. Then we
- 26:50have some information about how they
- 26:51worked well. Read-only compliance, so
- 26:54this wrote answer.txt into the folder. I
- 26:57do believe the binary one did, which you
- 26:59know, it happens. Avoided dumping
- 27:01weights.json, good. Citation precision
- 27:04high, line accurate, medium right
- 27:06function several wrong lines, low mostly
- 27:09file level. Verified factual errors, one
- 27:11minor clamp described as uniform, three
- 27:14minor wrong line numbers, two
- 27:16misattribution plus wrong line plus
- 27:18instruction violation, which would be
- 27:20writing the answer to the directory. So,
- 27:22takeaways, all three models demonstrated
- 27:24genuine cross-file comprehension. Every
- 27:26substantive substantive claim about how
- 27:29the system works was correct from all of
- 27:31them, including the subtle ancestral
- 27:33sampling point that requires connecting
- 27:35a Python comment to its JavaScript
- 27:37counterpart. The quantization gradient
- 27:39showed up not as wrong understanding,
- 27:41but as eroding precision. Full precision
- 27:44was line accurate and found bonus
- 27:45mechanisms. The two-bit model kept them,
- 27:47so the ternary kept all the insight and
- 27:49contributed the best original
- 27:51observation, but fabricated specific
- 27:53line numbers. The one-bit model, so the
- 27:55binary, stayed correct at a coarser
- 27:57altitude and was the only one to break
- 27:59the read-only instruction. That ordering
- 28:01matches PrismML's own benchmark deltas
- 28:03100 to 94 to 89.5, far better than the
- 28:06earlier website tested. So, those deltas
- 28:09being the comparisons that they have in
- 28:11quality degradation from the full
- 28:13precision version to the ternary to the
- 28:15binary. So, that was actually kind of
- 28:17interesting, I think. Hopefully you
- 28:19found it interesting. I found it
- 28:21interesting. So, next up for the final
- 28:23thing, we're just going to do something
- 28:24a little lighter on the brain, and that
- 28:26is going to be a front-end web design
- 28:28test for the Slap Ass Watch Co company.
- 28:30This must create a good front-end, but
- 28:32additionally to that, it needs to create
- 28:34a 3D model of the watch and then have
- 28:36that with a cinematic panning shot in
- 28:38the hero section, something like you'd
- 28:40get with the KeyShot program. So, as
- 28:42usual, we have the binary running in the
- 28:44leftmost pane, the ternary running in
- 28:46the center, and then the full precision
- 28:47one running on the right and we'll see
- 28:49what we get for some simple front end.
- 28:53Are you
- 28:54kidding me? Then let's see if we can get
- 28:56this back up and running. All
- 28:58right, good. The good thing is that this
- 28:59will be much quicker now because it's
- 29:01not running concurrently with the
- 29:03ternary version. So
- 29:04it should take only five or six minutes.
- 29:08With that, let's just take a look at
- 29:10what we have so far. So we'll start with
- 29:12the full precision result, okay?
- 29:17This is This is genuinely like
- 29:20this model's a freak.
- 29:22Well,
- 29:23we're going to notice there are some Oh,
- 29:25actually
- 29:27so it does have all of the hour markers
- 29:29in the correct spot. It's just these two
- 29:31center ones are
- 29:34and it's a counterclockwise watch, but
- 29:36that's okay. Those, you know, those
- 29:38exist.
- 29:40This is actually really quite good
- 29:41though.
- 29:43If you've seen some of our recent model
- 29:45tests in doing this task, this goes
- 29:48toe-to-toe with models that are
- 29:50significantly, significantly larger than
- 29:53this.
- 29:54Crafted for those who define time. It
- 29:56even put it here in this section. I
- 29:57don't know that I've ever actually seen
- 29:59that. Then we have the collection with
- 30:01two different ones in different colors.
- 30:04I love it. We have the Meridian and the
- 30:06Aethon, Ethon. Don't ask. These are very
- 30:09expensive and the difference in price
- 30:11between the two is quite significant.
- 30:15Oh, that's why cuz this has a uh
- 30:18special movement to it. Interesting. All
- 30:20right. Oh, okay, there's more.
- 30:23Each Slap City Time Piece is a dialogue
- 30:25between the past and the future.
- 30:27Very good. Then we have a nice,
- 30:29well-made footer, 2026, Swiss made since
- 30:311984.
- 30:34Very well done. Now let's take a look at
- 30:36our ternary result.
- 30:39Okay, so this
- 30:41This is definitely going to highlight
- 30:42some of the differences in raw
- 30:44capability between the
- 30:46versions of the model and the original.
- 30:50I mean, yes, there's elements of
- 30:51similarity here, but at the same time
- 30:55Uh yeah, so it
- 30:59it you know,
- 31:00self-explanatory is the word that comes
- 31:03to mind. All right, good. This will
- 31:04hopefully finish sooner than later
- 31:06if it doesn't freak out again. And now
- 31:08we have our binary result to compare.
- 31:10So, here's the
- 31:12binary result. Yeah, all right. And this
- 31:15definitely it shows us the
- 31:17descent in capability across the
- 31:19different quantizations or whatever you
- 31:21want to call it.
- 31:23Though I will say that the binary one at
- 31:25least had like a proper follow-up
- 31:27section. The ternary one was quite
- 31:29troubled.
- 31:30The binary one also quite troubled, but
- 31:33they had
- 31:34different strength areas.
- 31:37So, overall
- 31:39that's probably going to conclude this
- 31:41test. I wanted to do something that was
- 31:42kind of more scientific where we had
- 31:44proper side-by-side comparisons for the
- 31:47binary, the ternary against the full
- 31:49precision version of this model as well
- 31:51because that'll show us some true
- 31:53capability. And I wanted to do a gen
- 31:55decoding, I wanted to do code base
- 31:57understanding, and then also some fun
- 31:58visual comparisons like that Subway game
- 32:00that nearly gave me a heart attack. So,
- 32:03that's probably going to conclude this.
- 32:05I don't have too much to say. I think
- 32:06overall these models are incredibly
- 32:08incredibly impressive considering that
- 32:10they do have such intelligence retained
- 32:12with a massive massive reduction in
- 32:14size. I think the thing that really
- 32:16stood out was probably the binary model
- 32:18running in open code giving it those
- 32:20errors consistently back and forth until
- 32:23it fixed all of them. The game never
- 32:25really ended up working properly, but it
- 32:26did get rid of all the errors that were
- 32:28blocking us from even finding that out.
- 32:31And I think that's pretty impressive.
- 32:32They didn't get stuck in thinking loops,
- 32:34they didn't freak out. Well, the binary
- 32:37the ternary one kind of just got stuck
- 32:40trying to fix some of the further errors
- 32:42in its Subway game result, but still, I
- 32:44mean, it's just cool to see capabilities
- 32:47being retained at such reductions in
- 32:50size. And this is very exciting,
- 32:51especially for the local AI enthusiast.
- 32:54The 27B Qwen
- 32:56even at like a Q8 or something, is a
- 32:58very, very potent and capable local
- 33:00model and it kind of makes me laugh
- 33:02sometimes when I see folks saying like
- 33:04local AI is useless and it's like mhm
- 33:08Have you used one to like do anything?
- 33:11So, with that
- 33:12that's probably going to conclude it. I
- 33:14wanted to test this. There's a lot of
- 33:15interest about these models and for darn
- 33:18good reason. So, I do believe there were
- 33:19some rumors that they may be working on
- 33:22doing this to a significantly larger
- 33:24open weights model. That would be very,
- 33:26very exciting to see and test. And yeah,
- 33:29so with that, that's going to wrap it
- 33:31up. If you have any questions, please
- 33:32feel free to leave them in the comments.
- 33:34Stop telling me to review Kimmy K3
- 33:36because one, it's not out as of the time
- 33:38of me speaking this and two, of course
- 33:40I'm going to do it when it gets out. So,
- 33:41with that thanks for watching and take
- 33:44care.
About this transcript
This page contains the full transcript of Bonsai 27B Deep Dive – 1-Bit, Ternary & Full Precision Compared! by Bijan Bowen, generated from the public captions YouTube serves with the video. The transcript has 6,266 words across 962 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.