Kimi K3 explained in 13min.. — Transcript
Full transcript
- 0:00The DeepSeek moment back in January 2025
- 0:02was a pretty big deal. Not just because
- 0:04DeepSeek practically caught up to OpenAI
- 0:06when it comes to benchmarks, but it's
- 0:08also the fact that DeepSeek trained
- 0:10their model for a fraction of the cost
- 0:12if you consider only the GPU rental
- 0:13costs. Similarly, Qwen 1.5 has now
- 0:16nearly caught up to Anthropic and
- 0:17OpenAI, but not just in benchmarks, but
- 0:20it's casting a broader question around
- 0:22who will be leading AI innovation in the
- 0:24model layer of the AI stack. People want
- 0:27to know whether Chinese models will
- 0:28become the next frontier in the AI race,
- 0:31and how Chinese models like Qwen 1.5
- 0:33will affect the demand for intelligence
- 0:35from the application layer. These are
- 0:37certainly interesting discussions to
- 0:39talk about, and we can make a lot of fun
- 0:41conjectures to debate how the rest of
- 0:43the industry might shape up to. But,
- 0:45what's really impressive about Qwen 1.5
- 0:47is not just in the benchmarks alone, but
- 0:50how it affects the layers below as we
- 0:52look at the semiconductor industry and
- 0:54how the economics of inference looks
- 0:56like. So, in order to look at the bigger
- 0:58picture, we have to look at the model
- 1:00architecture and see how the ground is
- 1:02shifting underneath us. And at a high
- 1:04level, Qwen 1.5 is all about efficiency
- 1:07and intelligence. And since there's a
- 1:08lot to go through, let's start with what
- 1:10we already all know, and then build our
- 1:13understanding from there. Quick
- 1:14disclaimer, this video does get slightly
- 1:16more technical than usual as the video
- 1:19goes on. But, I really think it's
- 1:20important to at least get a grasp on why
- 1:23innovations coming from Qwen 1.5 is so
- 1:25unique. So, let's start by grounding our
- 1:27discussion from what we all know, or at
- 1:29least heard of by now, which is mixture
- 1:31of experts. Ever since around 2024,
- 1:34mixture of experts started to become
- 1:35adopted as the industry standard, and a
- 1:38large majority of models now use mixture
- 1:40of experts at the core. The idea behind
- 1:43mixture of experts is to activate only a
- 1:45small portion of the model instead of
- 1:47the entire model, largely to reduce the
- 1:50compute overhead to process each token.
- 1:52Now, depending on the model that you
- 1:54choose, you'll have different levels of
- 1:56sparsity of experts. Minimax M3
- 1:58activates around 3.1% of experts.
- 2:01Incline from Thinking Machine at 3.1%
- 2:03Neumotron 3 Ultra activates around 4.3%
- 2:06and now Chimera 3 around 1.8% expert
- 2:09activated, which makes this model not
- 2:11just the biggest in size among open
- 2:13models, but one of the smallest
- 2:15activation ratios. Chimera 3 is divided
- 2:17into 896 experts and only 16 experts are
- 2:22activated per token, which is how they
- 2:24got to 1.8% activation ratio at any
- 2:27given time. And even comparing this
- 2:29model to their previous Chimera 2 model,
- 2:32which was a 1 trillion parameters in
- 2:34size, they not only grew their number of
- 2:36experts more than double while
- 2:38decreasing the activation of experts
- 2:40from 2% to 1.8%, which goes to show how
- 2:43they scale their model in size without
- 2:46sacrificing efficiency. So, the question
- 2:48is how does this efficiency here
- 2:50actually look like when we run them on
- 2:52GPUs in data centers. In other words,
- 2:54does the efficiency also carry out when
- 2:57it comes to inference? A model like
- 2:59Chimera 3 that's this big is typically
- 3:01spread across many GPUs in data centers.
- 3:04Even at a lower precision, the model's
- 3:05weight alone is scattered across nearly
- 3:08six GPUs in Nvidia DGX B300 setup. And
- 3:12this leaves very little room for the
- 3:14context windows to be stored in the form
- 3:16of KB cache and more. So, in a more
- 3:18realistic deployment setup, your experts
- 3:21will typically be spread across
- 3:22supernode of 64 GPUs, which is what they
- 3:25recommend, or even Nvidia NVL72
- 3:28configuration containing 72 chips in a
- 3:30single rack, which means every token
- 3:32that passes through the model could be
- 3:34doing many trips across GPUs
- 3:36interconnected within the rack because
- 3:38experts are spread across a wide array
- 3:41of GPUs, which means now token will need
- 3:43to travel from GPU to GPU to activate
- 3:46experts that are scattered. And every
- 3:48time data moves from one GPU to another
- 3:51GPU, it adds a communication overhead.
- 3:54KimiK3 added what's called Stable Latent
- 3:56MoE, building on top of Mixture of
- 3:58Experts. And if you have seen my recent
- 4:00video on NeMoTron, it's very similar to
- 4:02NeMoTron's Latent MoE, where the token
- 4:05embedding is compressed into a lower
- 4:07latent representation. Since token needs
- 4:09to travel from GPU to GPU, having a much
- 4:12more compressed token representation
- 4:14helps reduce the communication overhead
- 4:16depending on the compression ratio. And
- 4:18since our token is compressed into a
- 4:20lower dimension, it reduces the amount
- 4:23of data that's actually passed between
- 4:26the GPUs and also the compute that's
- 4:28performed on them since the dimension is
- 4:30reduced as well. This is what Latent MoE
- 4:33does. We can see it from their diagram
- 4:35of KimiK3 showing the down projection of
- 4:37the token first, routed to experts with
- 4:40two shared experts that's always
- 4:42activated by default, and flowing the
- 4:44rest through the pool of experts here.
- 4:46And later, it gets projected back up to
- 4:48the original dimension for softmax. And
- 4:50this pool of experts that you see here
- 4:52is typically spread across many GPUs in
- 4:55a server. And the word stable in stable
- 4:57latent MoE here likely refers to their
- 4:59efforts in stabilizing the router during
- 5:02training when it comes to selecting the
- 5:04experts. Since KimiK3 practically split
- 5:07the model into 896 experts, selecting
- 5:10only 16 experts from a huge list of
- 5:13experts is not a trivial task. The
- 5:15router needs to consider the whole list
- 5:17of experts with their own scores and
- 5:19choose. And even a small variation can
- 5:21throw off the router big time. In other
- 5:23words, you can't have this many experts
- 5:26and select this little without a strong
- 5:28mechanism that help keep the balance in
- 5:30training. Instead of using popular
- 5:32methods like in the case of Deep Seek 33
- 5:34or NeMoTron 3 Ultra that use bias to
- 5:37nudge the router to pick experts by
- 5:39penalizing overused experts and
- 5:41promoting underused experts to reach
- 5:43equilibrium, KimiK3 improved the
- 5:45selection mechanism by using what's
- 5:47called quantile balancing, which helps
- 5:49expert allocation directly from the
- 5:51distribution of router scores, kind of
- 5:53like making each expert take the LSAT
- 5:56and grading them by percentile curve,
- 5:58rather than comparing them using a fixed
- 6:00raw scores. This helps the router decide
- 6:03which expert should be selected relative
- 6:05to the rest of the experts in score.
- 6:07Okay, the next component we'll get into
- 6:09for Kimi K3 is this bottom left section
- 6:11of the diagram. And this is really the
- 6:13meat and bone when it comes to why Kimi
- 6:15K3 is such a beautifully crafted model.
- 6:17This section right here is what helped
- 6:19contribute Kimi K3 to maintain six-time
- 6:22decoding throughputs, while also scoring
- 6:24higher than the status quo without
- 6:26sacrificing the speed. And what makes
- 6:28this entire thing possible is what's
- 6:30called Kimi Delta attention, or KDA for
- 6:32short. And surprisingly, KDA has already
- 6:34been around since October 2025. So, Kimi
- 6:37K3 essentially adopts what they already
- 6:39reached 9 months ago into their new
- 6:42model, but at a much bigger scale. So,
- 6:44how does KDA really work? Because it
- 6:45seems like it works really well. The
- 6:47keyword here is the letter A, attention.
- 6:49I'm sure we all heard of by now that
- 6:51attention is expensive. But we also love
- 6:53a model that can offer a 1 million
- 6:55context window at the application layer.
- 6:57And even here in the US, where we have
- 6:59much more compute available, offering a
- 7:01model at 1 million context window
- 7:03without modifying the attention is sort
- 7:06of foolish. The technical terminology
- 7:08that we use is making something
- 7:09quadratic complexity into sub-quadratic
- 7:12or even linear. And there are so many
- 7:14different ways that researchers have
- 7:16contributed to make this happen. And one
- 7:18of them is called linear attention. And
- 7:20linear attention works exactly how it
- 7:22sounds. How do we make our growing
- 7:24compute demand more manageable as the
- 7:26model scales? Models like Nemotron 3
- 7:28uses what's called Mamba 2, which
- 7:30replaces some of its attention layer
- 7:33that's known to be expensive with a
- 7:34recurrent state space memory. If this
- 7:37sounds gibberish to you, basically the
- 7:38core idea is having a predefined memory
- 7:41cell where new information coming in
- 7:43updates the old memory while old memory
- 7:46is erased at a learned rate following a
- 7:48decay schedule. For those who are
- 7:49mathematically inclined, the equation
- 7:51would look something like this where
- 7:53alpha determines how much previous
- 7:55memory has decayed and the new
- 7:57information writes the current memory on
- 7:58top. You might notice that this sounds a
- 8:00lot like recurrent neural network and it
- 8:03practically is, at least in how memory
- 8:05is stored by the model. And this will
- 8:07become more important later in the video
- 8:09when we get into exactly how Kim i K3
- 8:12optimizes even further. Now, building on
- 8:14this learned decay idea, what if instead
- 8:16of rewriting the memory with the new
- 8:19value, we find the error instead and
- 8:21only apply that error to correct our
- 8:24memory to make sure that this is all
- 8:26efficient. This method is called gated
- 8:28delta net, which improves the decay
- 8:30rule, as you can see, to not only decay
- 8:32old memories according to schedule, but
- 8:34also efficiently update the error
- 8:36between its value and the predicted. So,
- 8:38using gated delta net, we have more
- 8:40efficiency in how the memory is updated.
- 8:43And reading through the paper for gated
- 8:44delta net, it explicitly says that the
- 8:46challenge still is implementing gated
- 8:48delta in a hardware efficient manner.
- 8:51So, now we finally get to Kim i delta
- 8:53attention, which is what Kim i K3
- 8:55incorporates in their model. And
- 8:56mathematically, Kim i delta simply
- 8:58replaces the decay control in the gated
- 9:01delta net with a function that gives you
- 9:03a much more fine-grained control over
- 9:05how memory is exactly decayed and
- 9:08actually carries out through each
- 9:09channel. Gated delta net couldn't
- 9:11independently control how much
- 9:13information is retained and forgotten,
- 9:15and KDA essentially allows more granular
- 9:18control of how information is actually
- 9:20retained and forgotten channel by
- 9:22channel at different rates. And in Kim
- 9:24i's case, they interleave the Kim i
- 9:26delta attention at a 3:1 ratio following
- 9:28their ablation study that helped them
- 9:30pick the more ideal ratio. And you can
- 9:32see in the diagram here where they have
- 9:34three KDA layers with one global
- 9:36attention called gated MLA, making this
- 9:38a hybrid linear attention by definition.
- 9:41I mentioned earlier how this looks a lot
- 9:43like a recurrent neural network. And the
- 9:44reason why it's important here is
- 9:46because when it comes to training, one
- 9:48of the biggest drawbacks with RNN was
- 9:50its inefficiency in training compared to
- 9:53transformers because of the sequential
- 9:55nature that made training difficult due
- 9:57to temporal dependence. Following the
- 9:59original paper in gated Delta net that
- 10:01used chunk-wise parallel to optimize on
- 10:03training, Kimi also optimized by
- 10:06grouping multiple steps to help
- 10:07parallelize training to deal with
- 10:09temporal dependencies. And you can read
- 10:11through their hardware efficient
- 10:12chunk-wise algorithm here to get deeper
- 10:15understanding into how they actually
- 10:16work. Now, looking back at Kimi's
- 10:17diagram, we covered stable latent movie,
- 10:20we just covered Kimi Delta attention and
- 10:22hybrid layers. One big thing that you
- 10:24might have noticed here was this whole
- 10:26section showing many red lines. The red
- 10:28pipes here have a lot to do with how the
- 10:30plumbing works in information moving
- 10:32across the network. This is called
- 10:34attention residual. And much like how
- 10:36Kimi Delta attention was released 9
- 10:38months before it was incorporated into
- 10:40Kimi K3's official release, attention
- 10:43residual was also released back in
- 10:45March, so about 4 months before the
- 10:47release of Kimi K3. Residual network in
- 10:49general is sort of like this unsexy
- 10:51blue-collar part of model architecture
- 10:53since it's about creating streams for
- 10:55communication to happen much like real
- 10:57plumbing. Now, I did a thorough
- 10:58explanation on the theory behind
- 11:00residual network when I covered Deep
- 11:02Seek MHC video for reference. But the
- 11:04reason why we need this in the first
- 11:06place is because typically we have so
- 11:08many layers doing operations on top of
- 11:10our input where by the time you get to
- 11:12the end, the information gets distorted
- 11:15so much that it becomes really difficult
- 11:17to train looking back the deeper the
- 11:19model gets. And residual network helps
- 11:21us create the plumbing that's necessary
- 11:23for information to carry layer by layer
- 11:25without the pressure building up. So,
- 11:27just like plumbing may be necessary to
- 11:29relieve water pressure, residual network
- 11:31help allow models to scale in layers by
- 11:34creating additional streams for inputs
- 11:37to retain its original state as it moves
- 11:39through layers and modifications are
- 11:41applied on top. And standard residual
- 11:43connection in transformers did help, but
- 11:45the drawback is that layers that are
- 11:47deeper into the model often become
- 11:49diluted since they know very little
- 11:51about layers that are much earlier on.
- 11:53Attention residual changes this by
- 11:55allowing the current layer to
- 11:57selectively pull information from
- 11:59earlier residual states. And for Qwen
- 12:011.5 they're incorporated into a logical
- 12:04block and grouping together to prevent
- 12:06connection layer from becoming too
- 12:08expensive. So, not only do we have
- 12:10information flowing between layers,
- 12:11grouping previous layers into a bigger
- 12:13logical block helps reduce too much
- 12:16information flowing between layers,
- 12:17thereby decreasing the interconnect
- 12:19burden on GPUs. And looking at the
- 12:21research paper for attention residual,
- 12:23it showed very strong results in
- 12:25comparison since it makes the model a
- 12:27lot more expressive. And this is a great
- 12:29addition to Qwen 1.5 model. Now, as a
- 12:32closing note, I'm personally excited to
- 12:34see how Qwen 1.5 actually runs on more
- 12:36advanced chips here in the US. There's
- 12:38already been talks of our government
- 12:40potentially banning Qwen 1.5, but I
- 12:42think that would be a huge mistake given
- 12:44that we have so much gain in the US.
- 12:46Even looking at just how much more
- 12:48advanced we are in chip and the
- 12:50infrastructure layer to unleash the
- 12:52model potentially at a faster or
- 12:54potentially cheaper inference than their
- 12:56current pricing at $3 per million input
- 12:58token and $15 per million output tokens.
- 13:01I'm also really curious to see how this
- 13:03number might look when new clouds in the
- 13:05US get a hold of this open model and
- 13:08serve them for general public use. And
- 13:10it certainly paints an interesting
- 13:11dilemma where Chinese models that are
- 13:13shooting for efficiency could be run
- 13:15more efficiently here in the US assuming
- 13:17that we can leverage our stacks
- 13:19underneath.
About this transcript
This page contains the full transcript of Kimi K3 explained in 13min.. by Caleb Writes Code, generated from the public captions YouTube serves with the video. The transcript has 2,439 words across 373 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.