YouTube2Text

Kimi K3 explained in 13min.. — Transcript

by Caleb Writes Code · 2,439 words · 373 segments · language en · Watch on YouTube

Full transcript

  1. 0:00The DeepSeek moment back in January 2025
  2. 0:02was a pretty big deal. Not just because
  3. 0:04DeepSeek practically caught up to OpenAI
  4. 0:06when it comes to benchmarks, but it's
  5. 0:08also the fact that DeepSeek trained
  6. 0:10their model for a fraction of the cost
  7. 0:12if you consider only the GPU rental
  8. 0:13costs. Similarly, Qwen 1.5 has now
  9. 0:16nearly caught up to Anthropic and
  10. 0:17OpenAI, but not just in benchmarks, but
  11. 0:20it's casting a broader question around
  12. 0:22who will be leading AI innovation in the
  13. 0:24model layer of the AI stack. People want
  14. 0:27to know whether Chinese models will
  15. 0:28become the next frontier in the AI race,
  16. 0:31and how Chinese models like Qwen 1.5
  17. 0:33will affect the demand for intelligence
  18. 0:35from the application layer. These are
  19. 0:37certainly interesting discussions to
  20. 0:39talk about, and we can make a lot of fun
  21. 0:41conjectures to debate how the rest of
  22. 0:43the industry might shape up to. But,
  23. 0:45what's really impressive about Qwen 1.5
  24. 0:47is not just in the benchmarks alone, but
  25. 0:50how it affects the layers below as we
  26. 0:52look at the semiconductor industry and
  27. 0:54how the economics of inference looks
  28. 0:56like. So, in order to look at the bigger
  29. 0:58picture, we have to look at the model
  30. 1:00architecture and see how the ground is
  31. 1:02shifting underneath us. And at a high
  32. 1:04level, Qwen 1.5 is all about efficiency
  33. 1:07and intelligence. And since there's a
  34. 1:08lot to go through, let's start with what
  35. 1:10we already all know, and then build our
  36. 1:13understanding from there. Quick
  37. 1:14disclaimer, this video does get slightly
  38. 1:16more technical than usual as the video
  39. 1:19goes on. But, I really think it's
  40. 1:20important to at least get a grasp on why
  41. 1:23innovations coming from Qwen 1.5 is so
  42. 1:25unique. So, let's start by grounding our
  43. 1:27discussion from what we all know, or at
  44. 1:29least heard of by now, which is mixture
  45. 1:31of experts. Ever since around 2024,
  46. 1:34mixture of experts started to become
  47. 1:35adopted as the industry standard, and a
  48. 1:38large majority of models now use mixture
  49. 1:40of experts at the core. The idea behind
  50. 1:43mixture of experts is to activate only a
  51. 1:45small portion of the model instead of
  52. 1:47the entire model, largely to reduce the
  53. 1:50compute overhead to process each token.
  54. 1:52Now, depending on the model that you
  55. 1:54choose, you'll have different levels of
  56. 1:56sparsity of experts. Minimax M3
  57. 1:58activates around 3.1% of experts.
  58. 2:01Incline from Thinking Machine at 3.1%
  59. 2:03Neumotron 3 Ultra activates around 4.3%
  60. 2:06and now Chimera 3 around 1.8% expert
  61. 2:09activated, which makes this model not
  62. 2:11just the biggest in size among open
  63. 2:13models, but one of the smallest
  64. 2:15activation ratios. Chimera 3 is divided
  65. 2:17into 896 experts and only 16 experts are
  66. 2:22activated per token, which is how they
  67. 2:24got to 1.8% activation ratio at any
  68. 2:27given time. And even comparing this
  69. 2:29model to their previous Chimera 2 model,
  70. 2:32which was a 1 trillion parameters in
  71. 2:34size, they not only grew their number of
  72. 2:36experts more than double while
  73. 2:38decreasing the activation of experts
  74. 2:40from 2% to 1.8%, which goes to show how
  75. 2:43they scale their model in size without
  76. 2:46sacrificing efficiency. So, the question
  77. 2:48is how does this efficiency here
  78. 2:50actually look like when we run them on
  79. 2:52GPUs in data centers. In other words,
  80. 2:54does the efficiency also carry out when
  81. 2:57it comes to inference? A model like
  82. 2:59Chimera 3 that's this big is typically
  83. 3:01spread across many GPUs in data centers.
  84. 3:04Even at a lower precision, the model's
  85. 3:05weight alone is scattered across nearly
  86. 3:08six GPUs in Nvidia DGX B300 setup. And
  87. 3:12this leaves very little room for the
  88. 3:14context windows to be stored in the form
  89. 3:16of KB cache and more. So, in a more
  90. 3:18realistic deployment setup, your experts
  91. 3:21will typically be spread across
  92. 3:22supernode of 64 GPUs, which is what they
  93. 3:25recommend, or even Nvidia NVL72
  94. 3:28configuration containing 72 chips in a
  95. 3:30single rack, which means every token
  96. 3:32that passes through the model could be
  97. 3:34doing many trips across GPUs
  98. 3:36interconnected within the rack because
  99. 3:38experts are spread across a wide array
  100. 3:41of GPUs, which means now token will need
  101. 3:43to travel from GPU to GPU to activate
  102. 3:46experts that are scattered. And every
  103. 3:48time data moves from one GPU to another
  104. 3:51GPU, it adds a communication overhead.
  105. 3:54KimiK3 added what's called Stable Latent
  106. 3:56MoE, building on top of Mixture of
  107. 3:58Experts. And if you have seen my recent
  108. 4:00video on NeMoTron, it's very similar to
  109. 4:02NeMoTron's Latent MoE, where the token
  110. 4:05embedding is compressed into a lower
  111. 4:07latent representation. Since token needs
  112. 4:09to travel from GPU to GPU, having a much
  113. 4:12more compressed token representation
  114. 4:14helps reduce the communication overhead
  115. 4:16depending on the compression ratio. And
  116. 4:18since our token is compressed into a
  117. 4:20lower dimension, it reduces the amount
  118. 4:23of data that's actually passed between
  119. 4:26the GPUs and also the compute that's
  120. 4:28performed on them since the dimension is
  121. 4:30reduced as well. This is what Latent MoE
  122. 4:33does. We can see it from their diagram
  123. 4:35of KimiK3 showing the down projection of
  124. 4:37the token first, routed to experts with
  125. 4:40two shared experts that's always
  126. 4:42activated by default, and flowing the
  127. 4:44rest through the pool of experts here.
  128. 4:46And later, it gets projected back up to
  129. 4:48the original dimension for softmax. And
  130. 4:50this pool of experts that you see here
  131. 4:52is typically spread across many GPUs in
  132. 4:55a server. And the word stable in stable
  133. 4:57latent MoE here likely refers to their
  134. 4:59efforts in stabilizing the router during
  135. 5:02training when it comes to selecting the
  136. 5:04experts. Since KimiK3 practically split
  137. 5:07the model into 896 experts, selecting
  138. 5:10only 16 experts from a huge list of
  139. 5:13experts is not a trivial task. The
  140. 5:15router needs to consider the whole list
  141. 5:17of experts with their own scores and
  142. 5:19choose. And even a small variation can
  143. 5:21throw off the router big time. In other
  144. 5:23words, you can't have this many experts
  145. 5:26and select this little without a strong
  146. 5:28mechanism that help keep the balance in
  147. 5:30training. Instead of using popular
  148. 5:32methods like in the case of Deep Seek 33
  149. 5:34or NeMoTron 3 Ultra that use bias to
  150. 5:37nudge the router to pick experts by
  151. 5:39penalizing overused experts and
  152. 5:41promoting underused experts to reach
  153. 5:43equilibrium, KimiK3 improved the
  154. 5:45selection mechanism by using what's
  155. 5:47called quantile balancing, which helps
  156. 5:49expert allocation directly from the
  157. 5:51distribution of router scores, kind of
  158. 5:53like making each expert take the LSAT
  159. 5:56and grading them by percentile curve,
  160. 5:58rather than comparing them using a fixed
  161. 6:00raw scores. This helps the router decide
  162. 6:03which expert should be selected relative
  163. 6:05to the rest of the experts in score.
  164. 6:07Okay, the next component we'll get into
  165. 6:09for Kimi K3 is this bottom left section
  166. 6:11of the diagram. And this is really the
  167. 6:13meat and bone when it comes to why Kimi
  168. 6:15K3 is such a beautifully crafted model.
  169. 6:17This section right here is what helped
  170. 6:19contribute Kimi K3 to maintain six-time
  171. 6:22decoding throughputs, while also scoring
  172. 6:24higher than the status quo without
  173. 6:26sacrificing the speed. And what makes
  174. 6:28this entire thing possible is what's
  175. 6:30called Kimi Delta attention, or KDA for
  176. 6:32short. And surprisingly, KDA has already
  177. 6:34been around since October 2025. So, Kimi
  178. 6:37K3 essentially adopts what they already
  179. 6:39reached 9 months ago into their new
  180. 6:42model, but at a much bigger scale. So,
  181. 6:44how does KDA really work? Because it
  182. 6:45seems like it works really well. The
  183. 6:47keyword here is the letter A, attention.
  184. 6:49I'm sure we all heard of by now that
  185. 6:51attention is expensive. But we also love
  186. 6:53a model that can offer a 1 million
  187. 6:55context window at the application layer.
  188. 6:57And even here in the US, where we have
  189. 6:59much more compute available, offering a
  190. 7:01model at 1 million context window
  191. 7:03without modifying the attention is sort
  192. 7:06of foolish. The technical terminology
  193. 7:08that we use is making something
  194. 7:09quadratic complexity into sub-quadratic
  195. 7:12or even linear. And there are so many
  196. 7:14different ways that researchers have
  197. 7:16contributed to make this happen. And one
  198. 7:18of them is called linear attention. And
  199. 7:20linear attention works exactly how it
  200. 7:22sounds. How do we make our growing
  201. 7:24compute demand more manageable as the
  202. 7:26model scales? Models like Nemotron 3
  203. 7:28uses what's called Mamba 2, which
  204. 7:30replaces some of its attention layer
  205. 7:33that's known to be expensive with a
  206. 7:34recurrent state space memory. If this
  207. 7:37sounds gibberish to you, basically the
  208. 7:38core idea is having a predefined memory
  209. 7:41cell where new information coming in
  210. 7:43updates the old memory while old memory
  211. 7:46is erased at a learned rate following a
  212. 7:48decay schedule. For those who are
  213. 7:49mathematically inclined, the equation
  214. 7:51would look something like this where
  215. 7:53alpha determines how much previous
  216. 7:55memory has decayed and the new
  217. 7:57information writes the current memory on
  218. 7:58top. You might notice that this sounds a
  219. 8:00lot like recurrent neural network and it
  220. 8:03practically is, at least in how memory
  221. 8:05is stored by the model. And this will
  222. 8:07become more important later in the video
  223. 8:09when we get into exactly how Kim i K3
  224. 8:12optimizes even further. Now, building on
  225. 8:14this learned decay idea, what if instead
  226. 8:16of rewriting the memory with the new
  227. 8:19value, we find the error instead and
  228. 8:21only apply that error to correct our
  229. 8:24memory to make sure that this is all
  230. 8:26efficient. This method is called gated
  231. 8:28delta net, which improves the decay
  232. 8:30rule, as you can see, to not only decay
  233. 8:32old memories according to schedule, but
  234. 8:34also efficiently update the error
  235. 8:36between its value and the predicted. So,
  236. 8:38using gated delta net, we have more
  237. 8:40efficiency in how the memory is updated.
  238. 8:43And reading through the paper for gated
  239. 8:44delta net, it explicitly says that the
  240. 8:46challenge still is implementing gated
  241. 8:48delta in a hardware efficient manner.
  242. 8:51So, now we finally get to Kim i delta
  243. 8:53attention, which is what Kim i K3
  244. 8:55incorporates in their model. And
  245. 8:56mathematically, Kim i delta simply
  246. 8:58replaces the decay control in the gated
  247. 9:01delta net with a function that gives you
  248. 9:03a much more fine-grained control over
  249. 9:05how memory is exactly decayed and
  250. 9:08actually carries out through each
  251. 9:09channel. Gated delta net couldn't
  252. 9:11independently control how much
  253. 9:13information is retained and forgotten,
  254. 9:15and KDA essentially allows more granular
  255. 9:18control of how information is actually
  256. 9:20retained and forgotten channel by
  257. 9:22channel at different rates. And in Kim
  258. 9:24i's case, they interleave the Kim i
  259. 9:26delta attention at a 3:1 ratio following
  260. 9:28their ablation study that helped them
  261. 9:30pick the more ideal ratio. And you can
  262. 9:32see in the diagram here where they have
  263. 9:34three KDA layers with one global
  264. 9:36attention called gated MLA, making this
  265. 9:38a hybrid linear attention by definition.
  266. 9:41I mentioned earlier how this looks a lot
  267. 9:43like a recurrent neural network. And the
  268. 9:44reason why it's important here is
  269. 9:46because when it comes to training, one
  270. 9:48of the biggest drawbacks with RNN was
  271. 9:50its inefficiency in training compared to
  272. 9:53transformers because of the sequential
  273. 9:55nature that made training difficult due
  274. 9:57to temporal dependence. Following the
  275. 9:59original paper in gated Delta net that
  276. 10:01used chunk-wise parallel to optimize on
  277. 10:03training, Kimi also optimized by
  278. 10:06grouping multiple steps to help
  279. 10:07parallelize training to deal with
  280. 10:09temporal dependencies. And you can read
  281. 10:11through their hardware efficient
  282. 10:12chunk-wise algorithm here to get deeper
  283. 10:15understanding into how they actually
  284. 10:16work. Now, looking back at Kimi's
  285. 10:17diagram, we covered stable latent movie,
  286. 10:20we just covered Kimi Delta attention and
  287. 10:22hybrid layers. One big thing that you
  288. 10:24might have noticed here was this whole
  289. 10:26section showing many red lines. The red
  290. 10:28pipes here have a lot to do with how the
  291. 10:30plumbing works in information moving
  292. 10:32across the network. This is called
  293. 10:34attention residual. And much like how
  294. 10:36Kimi Delta attention was released 9
  295. 10:38months before it was incorporated into
  296. 10:40Kimi K3's official release, attention
  297. 10:43residual was also released back in
  298. 10:45March, so about 4 months before the
  299. 10:47release of Kimi K3. Residual network in
  300. 10:49general is sort of like this unsexy
  301. 10:51blue-collar part of model architecture
  302. 10:53since it's about creating streams for
  303. 10:55communication to happen much like real
  304. 10:57plumbing. Now, I did a thorough
  305. 10:58explanation on the theory behind
  306. 11:00residual network when I covered Deep
  307. 11:02Seek MHC video for reference. But the
  308. 11:04reason why we need this in the first
  309. 11:06place is because typically we have so
  310. 11:08many layers doing operations on top of
  311. 11:10our input where by the time you get to
  312. 11:12the end, the information gets distorted
  313. 11:15so much that it becomes really difficult
  314. 11:17to train looking back the deeper the
  315. 11:19model gets. And residual network helps
  316. 11:21us create the plumbing that's necessary
  317. 11:23for information to carry layer by layer
  318. 11:25without the pressure building up. So,
  319. 11:27just like plumbing may be necessary to
  320. 11:29relieve water pressure, residual network
  321. 11:31help allow models to scale in layers by
  322. 11:34creating additional streams for inputs
  323. 11:37to retain its original state as it moves
  324. 11:39through layers and modifications are
  325. 11:41applied on top. And standard residual
  326. 11:43connection in transformers did help, but
  327. 11:45the drawback is that layers that are
  328. 11:47deeper into the model often become
  329. 11:49diluted since they know very little
  330. 11:51about layers that are much earlier on.
  331. 11:53Attention residual changes this by
  332. 11:55allowing the current layer to
  333. 11:57selectively pull information from
  334. 11:59earlier residual states. And for Qwen
  335. 12:011.5 they're incorporated into a logical
  336. 12:04block and grouping together to prevent
  337. 12:06connection layer from becoming too
  338. 12:08expensive. So, not only do we have
  339. 12:10information flowing between layers,
  340. 12:11grouping previous layers into a bigger
  341. 12:13logical block helps reduce too much
  342. 12:16information flowing between layers,
  343. 12:17thereby decreasing the interconnect
  344. 12:19burden on GPUs. And looking at the
  345. 12:21research paper for attention residual,
  346. 12:23it showed very strong results in
  347. 12:25comparison since it makes the model a
  348. 12:27lot more expressive. And this is a great
  349. 12:29addition to Qwen 1.5 model. Now, as a
  350. 12:32closing note, I'm personally excited to
  351. 12:34see how Qwen 1.5 actually runs on more
  352. 12:36advanced chips here in the US. There's
  353. 12:38already been talks of our government
  354. 12:40potentially banning Qwen 1.5, but I
  355. 12:42think that would be a huge mistake given
  356. 12:44that we have so much gain in the US.
  357. 12:46Even looking at just how much more
  358. 12:48advanced we are in chip and the
  359. 12:50infrastructure layer to unleash the
  360. 12:52model potentially at a faster or
  361. 12:54potentially cheaper inference than their
  362. 12:56current pricing at $3 per million input
  363. 12:58token and $15 per million output tokens.
  364. 13:01I'm also really curious to see how this
  365. 13:03number might look when new clouds in the
  366. 13:05US get a hold of this open model and
  367. 13:08serve them for general public use. And
  368. 13:10it certainly paints an interesting
  369. 13:11dilemma where Chinese models that are
  370. 13:13shooting for efficiency could be run
  371. 13:15more efficiently here in the US assuming
  372. 13:17that we can leverage our stacks
  373. 13:19underneath.

About this transcript

This page contains the full transcript of Kimi K3 explained in 13min.. by Caleb Writes Code, generated from the public captions YouTube serves with the video. The transcript has 2,439 words across 373 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.