LLM Battle 9 - Cyber-Tiel-Coder vs Qwen — Transcript
Full transcript
- 0:00For whatever reason, we never got a
- 0:01Quinn 3.8 35B A3B. We got a Quinn 3.8
- 0:0527B dense, which is great, but uh kind
- 0:07of slow and uh a little bit difficult to
- 0:09run on low VRM setups. And we got a
- 0:11Quinn 3.8 Flash next, which is
- 0:13presumably great, but huge and very
- 0:15difficult to run in low VRM setups.
- 0:16Although, stay tuned for tomorrow's
- 0:18episode for a possible answer to that
- 0:20one. What we're missing is a a mixture
- 0:23of experts version of this new model in
- 0:25that size range in the 25 to 35 billion
- 0:27parameter range. Something that can run
- 0:29well even in constrained environments
- 0:31and something which is much faster than
- 0:32a dense model. Uh up until now we've
- 0:34been looking at post trains and
- 0:36finetunes on top of coin 3.6 35b3b and
- 0:40there are a couple of decent ones out
- 0:41there but uh it feels like that's a
- 0:42generation out of date. Can we get
- 0:44something that has the smarts of coin
- 0:453.8 and the speed of a mixture of
- 0:47experts model? Let's take a look at some
- 0:49of the alternatives in this range.
- 0:59>> Okay. What I wanted to do is take two
- 1:01new models and put them headto-head
- 1:02against Quen 3.6 35B A3B because that's
- 1:06our baseline. The first model we were
- 1:07going to look at is called Zing. I have
- 1:10no idea if I'm pronouncing that
- 1:11correctly or not, but I like saying it's
- 1:12Zing. It's a great name. Unfortunately,
- 1:15I just found out that Zing is waiting on
- 1:17a pull request to be merged into
- 1:18lambda.cpp. The architecture is not yet
- 1:21supported, so I can't actually run it in
- 1:22my lambda.cpp build. So, all right,
- 1:24Zing's out already. We'll we'll have to
- 1:26put Zing on the back burner. We'll
- 1:27revisit Zing in a future episode. So,
- 1:29that only leaves one. It's been
- 1:31suggested by many viewers. It's called
- 1:32Cyber Teal Coder 35b A3B. This is not a
- 1:35fine tune or a derivative of Quen 3.6.
- 1:39This is actually based on ornith 1.5
- 1:4335b3b sort of. If we look at the
- 1:45description here, they for whatever
- 1:48reason they didn't base this uh cyber
- 1:50teal coder model on the base model of
- 1:52ornith 1.5. They started from an
- 1:54obliterated version of 1.535b3b
- 1:58which itself was based on the ornith 1.5
- 2:00base model. So we're two levels removed
- 2:02now from the from the ornith base model.
- 2:04Uh but we've looked at that base model
- 2:06that ornith 1.5 35b3b. We've looked at
- 2:08that on this channel before and it did
- 2:09quite all right in the last time we we
- 2:10looked at it as opposed to the Ornith 9B
- 2:13version which did not do so well. Don't
- 2:14don't make me play that clip again. It's
- 2:17live roads eating frame zero.
- 2:23Uh okay. So if we look at the
- 2:25description of uh Cyber Teal Coder,
- 2:27there's a big disclaimer up front
- 2:28saying, you know, cuz it's obliterated
- 2:30you it could do harmful things. I what
- 2:33what are we going to do here that's
- 2:34harmful? I don't I don't even know.
- 2:35We're we're putting it to coding tasks,
- 2:37right? I mean, I guess if you're doing
- 2:38security penetration testing of your own
- 2:40stuff, you want to make sure your stuff
- 2:41is secure and a regular model might
- 2:43refuse to do it because it thinks you're
- 2:44some kind of evil hacker. That's the
- 2:45only that's the only reason I can think
- 2:47of to use an obliterated model. If I'm
- 2:48wrong, you know, send me straight in the
- 2:50comments below. I don't know why people
- 2:51are big on these obliterated models, but
- 2:53maybe it's just me. Uh, okay. So, Cyber
- 2:57Teal, apparently, according to their
- 2:59their Hugging Face page, it outcodes
- 3:00every other 35BAB
- 3:03uh and even spring of 2026 Frontier
- 3:05models. That's a very bold claim. Uh,
- 3:08and they also claim to be three to four
- 3:10times faster than 3.827 dense. That's
- 3:12kind of a low bar to hit because a dense
- 3:13model of course is going to be much
- 3:14slower than an MOE model. So, all right.
- 3:17All right. Let's keep an open mind. Uh,
- 3:19if we look for recommended settings here
- 3:22somewhere,
- 3:24we get temperature 6, top P 0.95, top
- 3:26K20, min.
- 3:29And if we go and take a look at what
- 3:30I've got here, yep, it matches exactly.
- 3:32I've got downloaded this at Q5 and we're
- 3:34going to run with 192,000 tokens of
- 3:35context. Uh so let's start up Cyber Teal
- 3:40and slide over to Envy Top as it loads.
- 3:43And we see that we are using 10 12 gigs
- 3:46out of 12 on card one and 19 1/2 or so
- 3:49out of 24 on card 2. So I think we're
- 3:51comfortably in VRAM even with 192,000
- 3:53tokens of context and this is an MOE
- 3:55model. So I'm expecting this is going to
- 3:56be blazing fast. Uh, so we're going to
- 3:59run our challenge first with Cyber Teal
- 4:02and then with our uh I guess Quen 3.6 35
- 4:05A3B will be the reigning champ, the
- 4:07defending champ, let's say, and we'll
- 4:09see which one does a better job of this
- 4:11uh of this task. So, what is the task
- 4:13for today? We're going to try something
- 4:15a little bit different. Normally, when
- 4:16I'm evaluating a model or doing a battle
- 4:18episode on this channel, I like to come
- 4:19up with a coding task. I want the thing
- 4:21to actually generate code for me because
- 4:23that's what I do. My primary use case
- 4:25for local LLMs is to to generate code.
- 4:28But there's other stuff that they can
- 4:29do. Uh so what I thought I thought we
- 4:31tried something a little bit different
- 4:32this time around. I'm going to put two
- 4:34challenges in front of each of these
- 4:36models. Challenge number one, we're
- 4:38going to do an in-depth code review of a
- 4:40feature that was just implemented from a
- 4:41complicated spec doc. And we're going to
- 4:43see what quality of review we get from
- 4:46uh from both models. Does one model find
- 4:49stuff that the other model misses? Do I
- 4:51get better, more niggly feedback from
- 4:53one than the other? And how long does it
- 4:54take? Uh we're going to compare. The two
- 4:56models we're comparing are both models,
- 4:59so uh they should get about the same
- 5:01tokens per second on my hardware.
- 5:02They're both at Q5. They're both have
- 5:04the same context window. So
- 5:07everything should be this should be a
- 5:08level playing field, and it's going to
- 5:10be interesting to me to see how well or
- 5:12how poorly these models uh perform
- 5:14relative to one another. So let's take a
- 5:17quick look at this monster spec
- 5:19document. This was for a project I've
- 5:21got on the go called Dusted Dominion.
- 5:22It's a new game. It's very early in
- 5:24development, so there's nothing we can
- 5:26actually show and play yet. But, uh,
- 5:28this spec doc is to introduce a, uh, a
- 5:31custom UI, uh, widget library. Uh, just
- 5:35really briefly, I know that there's, you
- 5:36know, pre-made libraries out there you
- 5:37can use. Like Pygame, GUI is a good one.
- 5:39Uh, there's a there's a whole bunch to
- 5:41choose from, but I thought because my
- 5:42requirements are pretty specific and
- 5:44because I like, you know, generating
- 5:46code, I thought, you know, we'll just
- 5:47roll our own. How hard could it be,
- 5:48right? So, so let's take a look at this
- 5:51spec real quick. Uh, the game needs a
- 5:53way to define and render UI widgets with
- 5:55optional color style theme support. And
- 5:57this particular spec doc introduces just
- 5:59one widget type just as an example, just
- 6:01as a proof of concept really to to prove
- 6:03out that the framework uh, you know, it
- 6:05works. We're going to introduce a
- 6:07button, which is about the simplest
- 6:08control I can think of that actually
- 6:09does something. Uh, and in future
- 6:11specifications, we'll introduce
- 6:12additional widgets as we need them.
- 6:14Great. Uh, theme configuration will
- 6:16allow you to set different colors
- 6:17depending on the state of each widget.
- 6:19So, if it's normal, if it's selected, if
- 6:21you're hovering over it with the mouse,
- 6:22or if it's disabled, it's all those are
- 6:24all different styles. Great. Uh, then I
- 6:26do a little bit of code prototyping here
- 6:29in the spec doc just to uh get it
- 6:31started. Uh, we've got a a base widget
- 6:33class, which all widgets are going to
- 6:34extend. We've got a base UI manager
- 6:36class. It's going to manage all of our
- 6:37widgets and handle rendering them. Uh,
- 6:39we're not going to use a layout engine.
- 6:40We're just going to use absolute
- 6:41layouts. I know I'm being a little bit
- 6:42lazy here, but I just want to get it up
- 6:44and running. Uh, and then I just I talk
- 6:46about how we're handling themes. I talk
- 6:48about, you know, can we can disable a
- 6:49widget. I talk about uh we're going to
- 6:52because the game supports multiple
- 6:53resolutions. We're going to use a design
- 6:54resolution of 1920 x 1080. So all
- 6:56internally the game will always specify
- 6:58coordinates based on that and then we'll
- 6:59scale it up or down based on the
- 7:00resolution. All of our supported
- 7:01resolutions are 16x9. So this is kind of
- 7:03a a cheap way of doing that. That's uh
- 7:06not going to introduce any problems. If
- 7:07we introduce a non-69 aspect ratio,
- 7:09we're going to have to rethink this, but
- 7:10it's fine for now. And then I talk about
- 7:12the supplied widget that's included in
- 7:13the spec. It's just a button. You can
- 7:15click it. It's got a label and an
- 7:16optional icon and you can click it and
- 7:18it does stuff. That's fine. Uh then we
- 7:20talk about unit tests and uh we talk
- 7:21about acceptance criteria. Uh our open
- 7:24questions were all resolved and uh I
- 7:26also did a little uh dev plan so we
- 7:28could stage the implementation cuz uh
- 7:30it's at that size where if you try to
- 7:32oneshot this, it's a bit too much. We've
- 7:33already done this. This this is already
- 7:34implemented. The code already exists.
- 7:36We're not going to be building this
- 7:36today. It's already built. What we're
- 7:38going to do is get our models to review
- 7:40the code that was built against this
- 7:42spec doc and tell me if there's any gaps
- 7:43or deficiencies. I've already reviewed
- 7:46it with both Quen 3.827B and GitHub
- 7:50Copilot, so I'm reasonably sure that
- 7:52this code is a good implementation of
- 7:53this spec, but let's see what we find
- 7:55out. You can always be surprised, right?
- 7:56Okay, so that's task one. We're going to
- 7:58review uh we're going to do an in-depth
- 8:00code review. It's going to have to dive
- 8:01into the codebase, take a good look at
- 8:03the code, and compare it to the spec
- 8:05doc, and tell me if we hit the target or
- 8:06not. That's that's challenge one. Uh if
- 8:08we get past challenge one with both
- 8:10models and we're going to move on to
- 8:11challenge two uh challenge two will be
- 8:14to uh do a review a spec review of a
- 8:17proposed feature that is not yet
- 8:18implemented. And we're not going to
- 8:19actually get to the point I think uh
- 8:20maybe we do maybe the models are both
- 8:22super fast. Maybe we get to the point
- 8:23where we can actually implement that
- 8:24spec doc. But challenge two officially
- 8:27at the time that I'm filming this is
- 8:28just we're going to do this spec review
- 8:31uh for feasibility and uh for uh you
- 8:35know consistency. Are there any
- 8:36contradictions in this spec doc? Are
- 8:37there any problems with it if we were to
- 8:39give you this spec doc for
- 8:40implementation? Would you be able to
- 8:41implement it is the question I'm trying
- 8:42to answer here. I've already been
- 8:44through a couple of passes of the spec
- 8:45doc with Quen. So, I think it's
- 8:47reasonably locked down, but there's
- 8:48still a few things that I think could be
- 8:49nailed down. So, this will be
- 8:50interesting to see what these models
- 8:52come up with, if they can come up with
- 8:53anything. Uh, okay. So, we've got Cyber
- 8:57Teal Coder loaded and ready to go. Let's
- 8:59bring up the official No Place Like
- 9:01Local Host stopwatch.
- 9:04super long shot, but if anybody from
- 9:06Nvidia is watching this, this could be
- 9:07the official Nvidia stopwatch. I'll put
- 9:09your logo right here. All I'm asking for
- 9:10is an RTX Pro 4500. Send me one of
- 9:14those. And this is your stopwatch. Just
- 9:16throwing that out there. All right, so
- 9:18here's our prompt. We're in open code
- 9:20with my senior dev uh code review agent.
- 9:23We've just implemented support for a new
- 9:25spec doc uh 04 UI widgets. Uh take a
- 9:28look at the code and tell me if we hit
- 9:30the target. Are there gaps or
- 9:31discrepancies in the implementation? Are
- 9:32there bugs? We're going to send that off
- 9:35and start the clock. Uh, now I did a
- 9:38little bit of testing with this model
- 9:40before filming and I was getting about
- 9:41100 tokens per second out of it. Uh,
- 9:43it's an MOE model, so it's I mean, right
- 9:44away it's super fast compared to uh,
- 9:46Quen 3.8. But does it stack up? Can it
- 9:50dig into a problem as thoroughly as Quen
- 9:543.8? Actually, I shouldn't compare it to
- 9:563.8. We're not comparing this to 3.8.
- 9:58This is a battle episode of these models
- 10:00against the Quinn 3.6 635B A3BOE which
- 10:04is the right now the the only really the
- 10:06only good model in that size range. So
- 10:10let's see what it can do.
- 10:18>> 5 minutes 5 seconds. You're kidding me.
- 10:21I mean that
- 10:23what speed were we getting here? Let's
- 10:25take a look at this. We were getting 74
- 10:27tokens per second. 75 76. Yeah, 80
- 10:33looks like speed decreased as it went
- 10:34through the task. Interesting. But I
- 10:36mean 5 minutes and 5 seconds is
- 10:37suspiciously fast for a code review.
- 10:40Let's take a look. Okay, the theme UI
- 10:42manager widget framework, color parser,
- 10:44design space scale, and the button click
- 10:46contract all track the spec line for
- 10:47line, but faithful to the spec and
- 10:49production ready aren't the same
- 10:50sentence. And there's one thing that'll
- 10:52bite you. Let's see what we Let's see
- 10:54what it found. No blocking bugs. That's
- 10:55good. Two minor edge cases.
- 10:58The drawn wrecked can be up to one pixel
- 11:00smaller than the hit wrecked. We're
- 11:02nitpicking one pixel. Pressed is cleared
- 11:05before on click fires. Negligible in
- 11:07practice. Oh, interesting. There's a
- 11:11there's a Okay, so it's not just a
- 11:13performance issue that it found. There's
- 11:15a leak. Every time you create a button,
- 11:18you get up to 10 uncashed cis font
- 11:21creations that are never cleaned up. So
- 11:23if you create 10 buttons, you get a 100
- 11:25of them per frame forever.
- 11:31It's funny that that didn't come up
- 11:33during the two code reviews we did on
- 11:35this earlier. So, okay. All right. I'm
- 11:37mildly impressed.
- 11:40Okay. So, no major security concerns.
- 11:43There's a few gaps in the test coverage.
- 11:45It seems kind of minor.
- 11:48Okay. So, it did find a bug that I did
- 11:51not know about. I was not expecting
- 11:52that. I'm just going to ask it to write
- 11:54that to a markdown file and we'll save
- 11:57that and we'll compare it to the one we
- 11:58get later from uh Okay, done. Uh so
- 12:02let's shut down Cyber Teal. Let's fire
- 12:05up Quinn 3.6 and let's do that exact
- 12:07same task again and see if Quen also
- 12:09finds that. Uh hello that bug.
- 12:14Uh okay, I lied. I thought I had this
- 12:15one at Q5, but it turns out I've got it
- 12:17at Q6. Okay, this is going to be a tiny
- 12:20little bit one lopsided, but that's
- 12:21okay. If we go to hugging face and look
- 12:23at the recommended settings for precise
- 12:25coding tasks, temperature 6 and these
- 12:27other settings here, if we compare to
- 12:29what I've got set, uh, yep, looks good.
- 12:33We're also giving it 192,000 tokens of
- 12:35context.
- 12:36Let's fire up Quen.
- 12:40This will probably take a little bit
- 12:42more VRAM just cuz the model's a few
- 12:44gigs bigger. Uh, let's see what we get.
- 12:47Just under 11 out of 12 on card one.
- 12:4921.3 out of 24 on card 2. Okay,
- 12:54sure. Let's slide over to open code.
- 12:56We've got the exact same prompt. Uh,
- 12:58Cyber Teal didn't make any code changes
- 12:59and I've moved its review markdown out
- 13:01of the way. So, this is the same branch
- 13:03basically. So, uh, let's
- 13:06let Quen chew on that.
- 13:12What was the time for for Cyber Teal?
- 13:14505. Let's see if Quen can beat 505.
- 13:212 minutes 2 minutes and 12 seconds. What
- 13:25speed were we getting? So low 80s low to
- 13:28mid 80 tokens per second. So roughly
- 13:31equivalent. Let's see what we got. Uh
- 13:33clean implementation. You built a
- 13:35theming system, a design space scaling
- 13:36layer, a button widget with autoscaling
- 13:38text, and full event-driven click
- 13:40contract. All with zero third party
- 13:42dependencies, no fluff, and the tests
- 13:43actually exercise the pixel output and
- 13:45cross resolution hit detection. Great.
- 13:47That said, it has a few wobbly spots and
- 13:49one actual bug. Let's see what it found.
- 13:51Current scale silently returns identity
- 13:53when no display exists. But scale.m
- 13:55mouse will divide by zero if scale
- 13:57equals one is used with a real display.
- 13:59Oh, so this is a hypothetical bug. This
- 14:02is a hypothetical latent bug that is not
- 14:04a you'd have to create a screen with a
- 14:06width of zero. Why would you do that?
- 14:08There's no guard against it. Fine.
- 14:11Floored edges can produce zero width
- 14:13rects for tiny design wrecks at low
- 14:14resolution. uh you uh okay so if you
- 14:18create a button with dimensions 1 comma
- 14:201 and then scale it down to a low
- 14:22resolution it rounds down to zero and we
- 14:25don't have a guard against it so the
- 14:26button just disappears but what do you
- 14:27what do you want you know if you create
- 14:29a button with a width of one and a
- 14:31height of one what are you doing I'm not
- 14:33worried about that okay so uh invalid
- 14:35test with invalid corner radiation
- 14:36should warn you zero this is fine just
- 14:39noting that the test name says okay so
- 14:41we're really really nitpicking let's see
- 14:44did okay so check this out fits Set size
- 14:46renders a font every time it's called
- 14:47during binary search. It got the same
- 14:49bug. Why didn't I get this from coin
- 14:513.8? Font is cached at the loader level,
- 14:55but font size still does a metrics
- 14:56lookup. Interesting. This doesn't
- 14:58complain about the same uh symptoms. Let
- 15:01me see what cyber said.
- 15:03Oh, for the fall back font path, there
- 15:06is no cache for the fallback font path.
- 15:09So if you don't have a font in your
- 15:11resources,
- 15:13we fall back to a font from the system.
- 15:15That's the fallback font path. I didn't
- 15:17see that when I read that. So this is
- 15:19not as serious as it looked on the first
- 15:21pass. This is only if you have not got a
- 15:24font at all, which it probably will
- 15:25never happen
- 15:27because the game will ship with a bunch
- 15:28to choose from, right? So why would you
- 15:30never have one? If you've gone out of
- 15:32your way to delete the ones that come
- 15:33with the game and you don't supply your
- 15:34own, then you get the fallback font,
- 15:36which is just whatever we find on the
- 15:37system, in which case you'll get that
- 15:40resource leak. Okay, so this thing also
- 15:43noted it, but it didn't find that
- 15:45particular uh scenario. It's just saying
- 15:47that uh yeah, it's not ideal that you're
- 15:49doing this. Okay, so not a resource
- 15:51leak, but possibly a minor performance
- 15:53problem if you've got 50 plus buttons on
- 15:54screen. I don't see that happening.
- 15:56Okay. Yeah, minor minor dock issues.
- 15:59Okay, so if I compare that, so 5 minutes
- 16:03and 5 seconds versus what was it? 2
- 16:05minutes and 12 seconds. They both found
- 16:07stuff that Quinn 3.8 didn't find, which
- 16:09is impressive, but nothing major.
- 16:12And I was expecting that because we went
- 16:13through this uh spec doc multiple times
- 16:16with Quinn 3.8 and GitHub Copilot. And
- 16:18so uh I was expecting that this was
- 16:20fairly locked down and we did find a few
- 16:22niggly things, but uh I'm I'm not
- 16:23concerned about them. So So that was our
- 16:25warm-up challenge. That was just to uh
- 16:27see if these models can actually, you
- 16:28know, do stuff. And they did. Quen was
- 16:30much faster. 2 minutes and 12 compared
- 16:33to 505 is impressive. But they both did
- 16:35a good job. And they both found stuff
- 16:37that didn't cop in the original review,
- 16:38which is interesting. So, let's move on
- 16:40to the next challenge, shall we? This is
- 16:42going to be a little bit more difficult.
- 16:44What I want to do, let me switch
- 16:45branches here. Okay, I have uh stopped
- 16:48Quinn and I've restarted Cyber Teal
- 16:50Coder and uh I've switched to a branch
- 16:53that has a spec doc that has not yet
- 16:55been implemented. Let's take a quick
- 16:57look at this uh new spec doc. This is
- 17:00for a new UI widget called text panel.
- 17:02This is a uh a readonly uh multi-line
- 17:05text panel with text wrap. So, we can
- 17:07display text on the screen. It sounds
- 17:09really simple but but let's take a quick
- 17:11uh look at it and see what we got. Text
- 17:13panel can have an optional icon. So,
- 17:15what I'm thinking is like uh during the
- 17:16game when you're playing, you get little
- 17:18uh messages from your crew or from the
- 17:19enemy ship or whatever, and it'll pop up
- 17:21on screen in a little text panel. It'll
- 17:22slide in from outside the screen, either
- 17:24from the bottom up or from the top of
- 17:26the screen down. And it's got on the
- 17:28left side, it's got an icon of whoever's
- 17:29speaking, right? A little uh profile
- 17:31picture of them. And then the rest of
- 17:32the space of the text panel is the text.
- 17:34And I want it to appear, you know, like
- 17:36the typing effect where the text appears
- 17:38like from top to bottom, left to right
- 17:40with a with a cursor writing it out. I
- 17:41want I want that kind of animation. And
- 17:43we want to be able to attach audio so
- 17:44that when the when the text panel drops
- 17:46down, we can play audio of, uh, a voice
- 17:48that's speaking the, you know, whatever
- 17:50you're hearing, so you don't have to
- 17:51read it if you don't want to. Great.
- 17:54So, uh, text panel appearance options.
- 17:56It's got an optional icon. Uh, it's got
- 17:58text, uh, and it's got, uh, a rect that
- 18:01describes where on the screen it is and
- 18:02how big it is. And we've got animation
- 18:04options. So text panels don't just
- 18:05appear they can by default they just
- 18:07appear on the screen wherever you say
- 18:08that they should be but they can uh
- 18:09slide they can fade in with transparency
- 18:12or they and/or they can slide in from
- 18:15offscreen somewhere. So I described the
- 18:17uh the easing formula we're going to use
- 18:18cuz it's not just going to drop in
- 18:20linearly. It's going to use a quadratic
- 18:21formula so it kind of you know slides in
- 18:23nicely. Describe how we the caller can
- 18:25specify how long that animation should
- 18:26take. Should should it take half a
- 18:27second or 2 seconds? How long should it
- 18:28take to drop down into place? Then I
- 18:30describe uh text display. It should
- 18:33handle line wrap at word boundaries and
- 18:34it should uh respect, you know, u
- 18:36explicit line breaks in the input text,
- 18:38all that good stuff. And audio, you can
- 18:40optionally associate audio with both the
- 18:42appearance of a text panel and the
- 18:43disappearance of a text panel. So you
- 18:44can get a little static noise or radio,
- 18:46you know, uh chatter noise or whatever
- 18:47if the uh the text panel goes away.
- 18:49Fine. Uh then I describe all the tests
- 18:51that we want to build for this thing and
- 18:53the acceptance criteria of this dock
- 18:55itself. This has not been built. So I I
- 18:57wrote the spec earlier and I've been
- 18:59through a couple of passes with Quinn
- 19:003.8 eight to tighten it up a little bit.
- 19:02But I think there's still some stuff
- 19:03that we can tighten up here. So this is
- 19:05challenge two. This is much more
- 19:07difficult, I think, because there's no
- 19:08code for it to look at yet. It's uh it's
- 19:10got to think about the spec and think
- 19:11about how it would implement this spec
- 19:12and tell me if this document is ready to
- 19:14be implemented or not. And if not, why
- 19:15not? So because we have a skill
- 19:19specifically for this in this project,
- 19:21we've got a skill called uh spec do
- 19:23review. And uh my my my prompt to Cyber
- 19:27Teal here can just be uh please do a
- 19:29spec review of this spec dock. Done. So
- 19:32I can fire that off.
- 19:363 minutes 15 seconds. That's crazy fast.
- 19:40It thinks a lot less than Quinn and it's
- 19:42three to four times faster than Quinn
- 19:443.8. I mean so uh the one thing I don't
- 19:47like about models is that it thinks so
- 19:49fast that I can't even read it. With
- 19:51Quinn, I can just sit here and read as
- 19:52it as it thinks through the problem. But
- 19:54this thing is so fast, it scrolls off
- 19:55the screen. Okay, great. Let's see what
- 19:57we got. Okay, basic uh document checks
- 20:00all pass. Let's uh see. No blocking
- 20:02problems. Great. It's implementable as
- 20:04written. Great. So, it's got a minor
- 20:06complaint about the way I specified the
- 20:08typing speed. That's going to make it a
- 20:09little bit awkward to write unit tests
- 20:11for it. Okay. No. No. Okay. It's
- 20:15confused about this point. Okay. I
- 20:17That's a signal that I probably could
- 20:18have worded it a little bit better.
- 20:20Okay. So that's a legitimate document
- 20:21problem. Okay. Again, a wording thing.
- 20:23That's that's fine. Degenerate wrecked
- 20:26not addressed. You filthy degenerate.
- 20:27What does that even mean? Button.draw
- 20:29early returns.
- 20:31A fully degenerate wreck should probably
- 20:33be in no op. Okay, I guess it's worth
- 20:35stating. That's a weird way to phrase it
- 20:37though. A degenerate wreck. Oh, this is
- 20:39great. This is great. It noticed there's
- 20:40no open questions section. Some of the
- 20:43other spec docs have an open question
- 20:44section. It suggests adding stuff. I
- 20:47haven't seen this before, but this is
- 20:48good. It's suggesting adding some open
- 20:50questions that we can resolve. Great. It
- 20:52didn't add them itself because there's
- 20:53explicit instructions in that skill not
- 20:55to modify the document. I don't want
- 20:56this thing to modify the doc. So, it's
- 20:58just putting them here in its review.
- 21:00You should add these open questions
- 21:01until we resolve them. I like that.
- 21:02That's a nice touch. Spelling and how do
- 21:05I have spelling and grammar mistakes?
- 21:07Okay. Normally in a document tightening
- 21:10session, I would make the edits myself,
- 21:12but it's offering to do this and we're
- 21:13in a feature branch specific to uh Cyber
- 21:15Teal. So, uh you know what? I'm just
- 21:17going to say,
- 21:23okay, just yeah, go ahead and do it. Go
- 21:25ahead and make the changes.
- 21:307:16 total time. So, it was 3 minutes
- 21:32and change to uh review it and then 4
- 21:35minutes and change to uh make its own
- 21:37edits to the file. Okay, fine. Still
- 21:39impressive. Still much much faster than
- 21:41Quin 3.8. I'm trying not to compare this
- 21:43to Quin 3.8 because this is supposed to
- 21:44be a battle of MOE models, but I, you
- 21:46know, my brain keeps going there cuz
- 21:48I've been working a lot with Quinn 3.8
- 21:50lately, and this is just so much faster.
- 21:52Let's put a pin in that. And let's go do
- 21:54the same thing with uh Quinn
- 21:58uh 3.635BA
- 22:003B. Okay. Uh I've shut down Cybertal,
- 22:04restarted Quinn 3.6,
- 22:07and I've switched to a feature branch
- 22:10specifically for Quinn. So, this does
- 22:11not include the changes that Cyber Teal
- 22:13just applied. That's in Cyber Teal's
- 22:14branch. We're completely separate. So,
- 22:16we're back now to the same state that
- 22:18Cybertal started in to keep things fair.
- 22:19I'm going to give it the same exact
- 22:20prompt. Please do a spec review of this
- 22:23new spec doc and go.
- 22:32One minute. Do I have reasoning turned
- 22:34off?
- 22:36No, it's thinking. 1 minute and 10
- 22:39seconds for that review. That's crazy.
- 22:42Okay. Uh, basic checks all pass. That
- 22:44was expected. Okay. This is weird.
- 22:48It's related to the typing animation. If
- 22:50it gets interrupted,
- 22:52then what's supposed to happen is it
- 22:54just like completely fills in the text.
- 22:56It stops the animation and it just
- 22:57completely renders the text. This thing
- 22:59is saying that this implies that the
- 23:02next appear, if you tell it to appear
- 23:04again after you've dismissed it, it
- 23:06should rerun the typing animation.
- 23:08That's explicitly not what it says. The
- 23:10text would be fully rendered and appear
- 23:11would just show it instantly. That's the
- 23:12intention. If you start animating the
- 23:15typing and then interrupt it and then
- 23:17tell the text panel to show itself
- 23:18again, it should just render completely.
- 23:20It shouldn't do the typing animation. I
- 23:21mean, this is a this is a a wonky edge
- 23:23case anyway, so I don't really care too
- 23:24much. But it's weird that it didn't
- 23:26understand that. That was that was
- 23:27pretty clear. So, its recommended
- 23:29resolution is the opposite of what I
- 23:30want. And I think I was pretty clear
- 23:32when I when I wrote that. You should say
- 23:35that the next call to appear does rerun
- 23:36the typing animation. No, it doesn't.
- 23:38This is the more useful and intuitive
- 23:39behavior. No, it isn't. Oh my god.
- 23:41Again, basic wording. The current
- 23:43rendered wrecked is ambiguous. No, it
- 23:45isn't. The on when a panel is sliding
- 23:48in, it's rendered wrecked changes from
- 23:50frame frame by frame. Yes, I know.
- 23:52Because it's moving. What do you think
- 23:54current rendered wctck means in that
- 23:55context? If it's interrupted at that
- 23:58point, it the current rendered wreck is
- 24:00wherever it was in that animation at the
- 24:02frame where it was interrupted. That's
- 24:03that's pretty clear, right?
- 24:06Does current render direct mean the
- 24:07wctck after applying the slide offset?
- 24:10Otherwise, a test could interpret it as
- 24:11the original constructor wt. What?
- 24:14That's not the current wreck if if it's
- 24:15not there yet. It's moving from source
- 24:17to destination. Current render direct is
- 24:20pretty straightforward. This is Quen,
- 24:22what are you doing here? Yeah. Okay. So,
- 24:23it's nitpicking the same uh stop
- 24:26amendment that the that Cyber Teal
- 24:28found, and that is that uh I didn't
- 24:30specify what you know, how does stop
- 24:32actually work? Does it stop just that
- 24:34specific sound or does it does it do
- 24:36stop on the uh on a specific channel
- 24:38where the sound is playing or stop on
- 24:39all channels where that sound is playing
- 24:41if there was if it was playing on more
- 24:42than one? Valid. Fine. I didn't word
- 24:44that very well. Oh, look at this. This
- 24:47one also suggests adding some open
- 24:48questions. Typing animation restart
- 24:50policy and stop method contract. It's
- 24:53funny that Quinn 3.8 never does that.
- 24:55Spelling and grammar. The dashes are
- 24:57uneven. I swear to God, these models are
- 24:59just trying to drive me crazy. Okay, I'm
- 25:02not going to tell Coin to implement it
- 25:04suggestions because some of its
- 25:05suggestions are just completely
- 25:06backwards. I'm not impressed with that.
- 25:07It spent one minute on that, which is
- 25:09really impressive, but it didn't seem to
- 25:12understand basic wording in at least a
- 25:14couple places and I'm not super
- 25:15impressed with that. I I got to say I
- 25:17think Cyber Teal Coder,
- 25:19although it was slower, I think it gave
- 25:22me better feedback. On the first round,
- 25:24it was maybe 50/50. But on this round,
- 25:28I'd say Cyber Teal easily beat Quinn,
- 25:31even though it was three times slower.
- 25:34Should we go for a bonus round? My
- 25:35camera's almost dead. Should we go for a
- 25:36bonus round? I want to give Cyber Teal a
- 25:39shot at actually implementing that spec
- 25:41and just see what we get, you know?
- 25:44Should we do it?
- 25:46You know what? Let's do it. I'm going to
- 25:47cut here and we're going to let Cybertal
- 25:49actually write some code.
- 25:51Okay, switch to camera 2. We're good.
- 25:54I've restarted Cyber Teal Coder and I've
- 25:56switched back to Cyberte's branch. I'm
- 25:58going to give it a chance to implement
- 25:59its spec to see what kind of code output
- 26:02we can get. I wasn't planning on doing
- 26:03this, but since Cybertal seems to be
- 26:05doing pretty well, let's give it a shot.
- 26:08So, very simple prompt. Please implement
- 26:10this spec doc and go. And I'm going to
- 26:13start the clock.
- 26:17What?
- 26:23Okay. Uh, I'm just going to just going
- 26:26to pump the brakes here. It's waiting
- 26:27for permission to access a directory.
- 26:29I'm just going to scroll up a bit and
- 26:30describe what happened here. I've been
- 26:32watching it. It's struggling.
- 26:34If I scroll up, uh, it's in the middle
- 26:37of editing code. It started having
- 26:39trouble with the edit tool.
- 26:42Uh, we were It was trying to make an
- 26:43edit and it failed and it said, "I keep
- 26:45making the same mistake. It's typing pi
- 26:47g instead of pi game. It's a weird typo.
- 26:50Uh but it couldn't fix it. It tried the
- 26:52edit tool a bunch of times,
- 26:55but it kept failing. It's trying to
- 26:56replace the wrong text. It's trying to
- 26:58replace pi g.
- 27:00Uh and it's getting mixed up. So, it
- 27:02actually
- 27:04decided on the fly to use said to use
- 27:06the bash tool instead of the right tool
- 27:08because it's having such a hard time
- 27:09with the right tool. Uh and it was
- 27:12having a hard time supplying the exact
- 27:14text replacement. I didn't quantize the
- 27:16KV cache and we're running a Q5. So,
- 27:19this feels like quantization error where
- 27:22uh it's making consistent typos, but
- 27:24that shouldn't be the case. But anyway,
- 27:27it for it said eventually it said,
- 27:29"Let's not use the right tool. I'm
- 27:30having a hard time with it. I'm just
- 27:31going to use said."
- 27:35Uh, but the said didn't change it
- 27:37because it wrote the same thing. So, it
- 27:39said, "Let me be deliberate.
- 27:41Current test is pi g, but I want pi
- 27:43game." So it tries.
- 27:47So it does write the replacement. Pi G
- 27:50should be replaced with PI game. Okay.
- 27:52And it fixes it. But then when it went
- 27:54to run the tests, it's failing because
- 27:56there's a bunch of references to pi GE
- 27:58all throughout the code. And it's
- 28:00getting frustrated because it can't let
- 28:03me see. Yeah, it's still typing pi G.
- 28:06Clearly, I have a persistent typo issue
- 28:08with this word. Let me write a script.
- 28:11Let's see what it does.
- 28:14So, it's writing the script to try to
- 28:16fix its own spelling mistakes.
- 28:20My confidence has dropped a few notches
- 28:22here, but let's let it go. I'm amused
- 28:24enough now to see what happens.
- 28:28Okay, I'm seeing a lot of weight.
- 28:30Actually, wait. Actually, wait.
- 28:32Actually, I think it's stuck.
- 28:37I think it's stuck.
- 28:41I think it's in a loop. and it's not
- 28:42coming back. I don't know what happened.
- 28:46If we look at the speed, it's been
- 28:47steadily going down uh as the context
- 28:49window increases in size. And now we're
- 28:51in the low 60s. We started in the mid
- 28:5380s, so it's getting slower and slower.
- 28:55It's still trying. It's still trying to
- 28:56do stuff.
- 29:00Okay, I'm calling it. 19:30, time of
- 29:04death, 11 p.m. Stop. Stop. It's going in
- 29:08circles. I'm glad we did that. Uh, I'm
- 29:11glad we did the coding exercise because
- 29:12I was just about to claim that Cyber
- 29:14Seal Coder is actually,
- 29:18you know what I maybe I shouldn't be
- 29:20pessimistic. This was a partial success.
- 29:22We got a good code review out of it. And
- 29:25if I can just recap what we did, it had
- 29:28to parachute in to an existing code base
- 29:30and compare the code against the spec
- 29:32from which that code was written and
- 29:34give me a code review. And it did that.
- 29:36Then we moved on and it did a doc spec,
- 29:39a spec doc review where it had to look
- 29:42at a spec that had not yet been
- 29:43implemented and tell me if the spec was
- 29:45implementable as written. Uh there's a
- 29:47few rough edges in it and it found some
- 29:49stuff and it suggested some changes to
- 29:51it. Great. And it did it in super fast
- 29:53time. That's great. Uh, at the coding
- 29:56stage, it's having trouble making
- 29:58consistent typos, failing to use the
- 30:00right tool, having to resort to the bash
- 30:03tool to to write said commands because
- 30:05it can't use the right tool, and now
- 30:06it's just going in circles and circles,
- 30:08and so I'm stopping it. Uh, still a
- 30:10partial success though in that it was
- 30:12able to give me very, very quick results
- 30:15on the spec dock review and the code
- 30:16review.
- 30:18And I'd be tempted to say that even
- 30:20though it fell down on the coding stage,
- 30:22uh I'd be tempted to say this would
- 30:25still be a useful model for me because
- 30:27Quen takes an enormous amount of time to
- 30:30to do a spec review. This thing did it
- 30:32in what 2 minutes? I'm going to say
- 30:34partial success even though it couldn't
- 30:36write code.
- 30:38There are other uses for LLMs other than
- 30:41generating code and that includes stuff
- 30:43like inspecting code and inspecting
- 30:45specifications and it did both of them
- 30:47tonight. So, uh, partial success for
- 30:49Cyber Seal Coder and in stage 2, it
- 30:51actually did a better job than Quinn.
- 30:53Granted, Quinn 3.6 35b3B.
- 30:56Uh, if we're stacking it up against that
- 30:58one, then I'd say that because this is
- 31:00based on Ornith 1.5,
- 31:02an obliterated version of Ornith 1.5.
- 31:04But I wonder how much that obliteration
- 31:06process affected this thing. I was
- 31:08thinking it was a quantization error,
- 31:10but now that I think about it, I wonder
- 31:11what that obliteration process does to a
- 31:14model's ability to reason.
- 31:17I'm not an expert on the subject so I
- 31:19don't know but uh interesting
- 31:21interesting results I'm going to say
- 31:23partial success for cyber coder
- 31:26and just before we go
- 31:31let's try this
- 31:37going to start it up with a smaller
- 31:39context window entirely on my 12 gig
- 31:42card
- 31:44because this is ane
- 31:46So, we start it up like that and let it
- 31:48load. I'm only giving it 12 gigs of VRAM
- 31:51and it fills 11 out of 12 on that card.
- 31:53If we go now to the web UI and just ask
- 31:57it a test question,
- 32:00what does a transient keyword do in
- 32:01Java? And let it think about that. It
- 32:03pegs that GPU and it still answers at
- 32:06about 35 tokens per second, which is a
- 32:07huge slowdown. That's less than half as
- 32:09fast as it was when I was giving it 36
- 32:11gigs of VRAM, but it still answers at 35
- 32:14tokens per second. And this is not even
- 32:15the MTP. Is there an MTPU?
- 32:19There is an MTP version of this model. I
- 32:21didn't download this. I downloaded the
- 32:22regular non-MTP version. So, we got 35
- 32:25tokens per second without MTP on 12 gigs
- 32:29of VRAM. This is uh another huge
- 32:31difference between a dense model and an
- 32:33MOE model. Not only is the MOE faster,
- 32:35but it can run on less VRAM. Even though
- 32:37you're you're offloading, you know,
- 32:39you're putting part of it into the
- 32:40system memory and the CPU has to get
- 32:41involved, it's still at least as fast. I
- 32:44get 35 tokens per second out of Quen
- 32:463.827B
- 32:48with much more VRAMm. So at this thing's
- 32:51worst, it's still as fast as Quen 3.8, a
- 32:54much a dense model.
- 32:57So, if you're unlimited VRAMm and you're
- 33:00looking for a a local LLM to help you
- 33:03with coding tasks, maybe not generating
- 33:05code, but reviewing code, reviewing spec
- 33:07docs, and maybe I don't know, based on
- 33:09the coding I saw tonight, I wouldn't
- 33:10trust it with coding exercise, but
- 33:13it can still do other things. It still
- 33:15has other purposes, still has other
- 33:16uses. So, I'm going to say partial
- 33:18success for Cyber Seal Coder. That was
- 33:20not maybe not fantastic, but also not uh
- 33:24how do I want to phrase this? How do I
- 33:27want to phrase this? Uh, it failed the
- 33:29coding challenge, but it passed the
- 33:30other two challenges. Two out of three.
- 33:34Okay, let's leave it there. I'm not
- 33:36going to torture it by getting it to do
- 33:37more code for me. Uh, I might consider,
- 33:41despite what we just saw, I might
- 33:43consider using this model to do spec
- 33:45reviews, uh, or secondary spec reviews
- 33:47for me because it did, now that I think
- 33:50about it, it caught something in that
- 33:51first stage that Quinn 3.8 did not
- 33:53catch. And that's impressive. That's
- 33:55impressive to me. So, and because it's
- 33:58so much faster, because of its speed and
- 34:01because of its ability to do that, I
- 34:03might keep this model around as a
- 34:04secondary spec doc reviewer because it
- 34:07just adds a few extra minutes to the
- 34:08process. You know, get Quinn 3.8 to do a
- 34:10review and then I'll get Cyber Teal to
- 34:11do a secondary review and just compare
- 34:13the results. Uh, okay. So, that'll do it
- 34:16for today.
- 34:18Thank you for watching and we will see
- 34:20you next time.
- 34:24For whatever reason, we never got a
- 34:25Quinn 3.8 35B A3B. We got a Quinn 37 37B
- 34:30dense. Well done. Try again. 3 2 1. Uh,
- 34:34if we get past challenge one with both
- 34:35models, and we're going to move on to
- 34:36challenge two, which is to what
- 34:40was challenge two?
- 34:43It's too late at night. Challenge one
- 34:45was to do a code review. Oh, right. Spec
- 34:48review.
About this transcript
This page contains the full transcript of LLM Battle 9 - Cyber-Tiel-Coder vs Qwen by No place like localhost, generated from the public captions YouTube serves with the video. The transcript has 7,032 words across 987 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.