YouTube2Text

How We Scaled Kimi K2.5 | Zhilin Yang's full GTC 2026 Keynote — Transcript

by Kimi AI · 5,592 words · 927 segments · language en · Watch on YouTube

Full transcript

  1. 0:07[music]
  2. 0:11>> Hi everyone.
  3. 0:13Thank you so much for the introduction.
  4. 0:15It's great to be here.
  5. 0:17It's great to have this opportunity to
  6. 0:20share with you guys some of our latest
  7. 0:23progress and efforts.
  8. 0:26So,
  9. 0:27one of our major pursues
  10. 0:30is to build better open models. And we
  11. 0:34believe in democratizing intelligence.
  12. 0:36With open models, you can deploy
  13. 0:38anywhere. It can be on your local
  14. 0:39servers. It can be on the cloud. And you
  15. 0:42can access every single bit of the
  16. 0:45weights in the modeling instead of just
  17. 0:48using a black box. And this is one of
  18. 0:50the slides that I took from Jensen's
  19. 0:53talk earlier this year at CES. So, as
  20. 0:56you can see, open models are quickly
  21. 0:58closing the gap with proprietary models
  22. 1:02and it's reaching the frontier. And we
  23. 1:04believe that with better and better open
  24. 1:06models, we're going to
  25. 1:08make intelligence more accessible to
  26. 1:10anybody in the world, in every corner of
  27. 1:13the world.
  28. 1:15But open models cannot be just open.
  29. 1:18They have have also to be great. So, in
  30. 1:23this talk, we're going to discuss how we
  31. 1:25make open models great. So, as we know,
  32. 1:29scaling is a primary driver
  33. 1:33for a lot of progress, maybe all of the
  34. 1:35you know major AI developments that we
  35. 1:37have witnessed in in the last few years.
  36. 1:40And here we're going to discuss how we
  37. 1:42scale our model in different dimensions.
  38. 1:45So, on the left-hand side, the first
  39. 1:47figure you see here is kind of the the
  40. 1:50standard scaling law. So, you on the
  41. 1:52x-axis you have the
  42. 1:54log of the number of training tokens.
  43. 1:56And on the y-axis you have the log loss.
  44. 1:59And as you scale the number of training
  45. 2:01tokens, you get a lower loss. But here
  46. 2:03the point is
  47. 2:05we're not going to just scale the number
  48. 2:06of training tokens, but we also want to
  49. 2:09want to improve the token efficiency.
  50. 2:13Meaning that we want to move this curve
  51. 2:15to the left-hand side so that we can
  52. 2:18achieve a lower loss, a much lower loss
  53. 2:21using the same number of training
  54. 2:23tokens. And this can be achieved by
  55. 2:25having better architectures and
  56. 2:28optimizers as we'll discuss in in our
  57. 2:31latest slides.
  58. 2:32And the second scaling dimensions that
  59. 2:34we're very interested in is to scale the
  60. 2:37context length. So, as you can see in
  61. 2:39the second second figure, if we increase
  62. 2:42the context length,
  63. 2:44then we can have a much higher accuracy
  64. 2:47in terms of predicting
  65. 2:50the token loss at a given position.
  66. 2:54And this means that we can increase the
  67. 2:56capability of the model to achieve more
  68. 3:00complex tasks by increasing the context
  69. 3:02length. So, this is the second scaling
  70. 3:05dimensions that we're going to going to
  71. 3:06talk about. And the third scaling
  72. 3:08dimension is the number of agents. So,
  73. 3:10we introduce this new learning paradigm
  74. 3:13of agent swarms where we can we don't
  75. 3:16just rely on a single agent, but we also
  76. 3:19orchestrate a swarm of agents that can
  77. 3:22accomplish the subtasks in parallel so
  78. 3:25that we can increase the task
  79. 3:26complexity. And we can translate all of
  80. 3:28this in into the language of agents. So,
  81. 3:31if you look at token efficiency, it's
  82. 3:33mostly about having a stronger prior so
  83. 3:37that you can
  84. 3:39have more efficiency when you do agent
  85. 3:41RL to search for a better solution. And
  86. 3:44when you think about long context, it's
  87. 3:46it's mostly about increasing the context
  88. 3:48length so that you can have a longer
  89. 3:50running agent. It can probably run for
  90. 3:52days or even weeks or months to
  91. 3:54accomplish more
  92. 3:56more more tasks, more complex tasks.
  93. 3:59And about and for agent swarms, it's
  94. 4:02another dimension that that added it and
  95. 4:05at the end of the day we're going to
  96. 4:06have a swarm of agents that each of them
  97. 4:08have a super long context and each of
  98. 4:10them have a very strong prior for us to
  99. 4:13search in this entire agent RL system.
  100. 4:19All right. So, we're going to start from
  101. 4:22token efficiency. So, this is one of the
  102. 4:24most you know classical figures in the
  103. 4:27history of machine learning. All right.
  104. 4:29So, it's taken from Kaplan et al. And
  105. 4:31basically says that if we scale
  106. 4:34proportionally the number of train
  107. 4:36tokens, the model parameters, and also
  108. 4:39the amount of compute, we can get lower
  109. 4:41and lower loss. And this is you know one
  110. 4:44of the major breakthroughs that that the
  111. 4:46entire community has achieved in the
  112. 4:49last few years to to get
  113. 4:51better intelligence. But here, what
  114. 4:54we're interested in is to have better
  115. 4:57and better token efficiency.
  116. 4:59And here's the thing. So, one thing that
  117. 5:02I would like to emphasize is that token
  118. 5:05efficiency is not just about efficiency.
  119. 5:08It's actually also about improving the
  120. 5:11upper bound of intelligence. So, here's
  121. 5:14here's why. So, suppose you have a
  122. 5:18say 50 trillion tokens, 50 trillion
  123. 5:21high-quality tokens.
  124. 5:23And then you apply this new optimizer,
  125. 5:25maybe the Meow optimizer. And then all
  126. 5:28of a sudden you have a two-times token
  127. 5:30efficiency. So, it means that
  128. 5:32it's almost like magic that you get
  129. 5:35equivalently 100 trillion tokens.
  130. 5:38And nowadays we're scaling towards the
  131. 5:41data wall and we're hitting you know the
  132. 5:43data wall and the amount of high-quality
  133. 5:46data is quite limited. And if we suppose
  134. 5:48that it's a constant amount, then we
  135. 5:51increase the token efficiency,
  136. 5:53it means that we're going to get better
  137. 5:55intelligence out of it. It's not just
  138. 5:57about infrastructure efficiency. It's
  139. 5:59about you know better
  140. 6:01intelligence. So, so this is why we
  141. 6:04spend you know a lot of efforts in this
  142. 6:07aspect because it's going to push the
  143. 6:09frontier of of intelligence. And Meow
  144. 6:12optimizer is one of the things that we
  145. 6:14have heavily invested in
  146. 6:16since last year.
  147. 6:18So, it's a second-order
  148. 6:21optimizer. And basically every single
  149. 6:24gradient update is transformed in a way
  150. 6:26that each entry is orthogonal to each
  151. 6:29other. And this is very different from
  152. 6:32the traditional Adam optimizer. And if
  153. 6:35if you implement this optimizer
  154. 6:37properly, you can get a two-times token
  155. 6:39efficient efficiency improvement. So, we
  156. 6:43we are one we are the first
  157. 6:45work we published the first work to
  158. 6:47demonstrate that Meow optimizer is
  159. 6:50actually scalable for LLM training. And
  160. 6:53these are two key techniques that we
  161. 6:56employed to make it effective for
  162. 6:58large-scale training. So, one of them is
  163. 7:00weight decay. It is critical for scaling
  164. 7:03to larger models. And the second is we
  165. 7:06want to ensure a consistent RMS updates
  166. 7:09compared to Adam. So, we have this
  167. 7:12adjustable coefficients that is applied
  168. 7:15to each update so that the resulting RMS
  169. 7:19is going to be comparable to Adam.
  170. 7:22And to make Meow memory efficient across
  171. 7:26all these Nvidia GPU clusters, we also
  172. 7:29develop a distributed Meow optimizer
  173. 7:32implementation that partitions the
  174. 7:34states across the data parallel group so
  175. 7:37that we can have a very efficient
  176. 7:40implementation for the Meow optimizer.
  177. 7:43And these are some of the results that
  178. 7:44were presented in the paper. So, as you
  179. 7:47can see, with with the same number of
  180. 7:49parameters and the same number of
  181. 7:51training tokens, we just replace the
  182. 7:53original AdamW optimizer with the new
  183. 7:56Meow optimizer. It's going to to improve
  184. 7:59the performance across the board
  185. 8:01significantly.
  186. 8:04But there was this new challenge that we
  187. 8:07encountered when we tried to scale it up
  188. 8:10further. When we tried to scale Meow for
  189. 8:12a one trillion parameter model, we
  190. 8:14encountered a new issue
  191. 8:17about training instability. So, as you
  192. 8:19can see on the left figure,
  193. 8:21we we observed that the max logits
  194. 8:24quickly explodes and quickly exceeds
  195. 8:281,000. And the typical values for
  196. 8:31for training for this max logits is
  197. 8:35about say 50 or maybe less than 100. But
  198. 8:39for for Meow, it quickly exceeds 1,000.
  199. 8:42And at the same time, we observe
  200. 8:44training divergence on the left-hand
  201. 8:47side. If you look at the training loss,
  202. 8:49it goes down a bit, but then at the end
  203. 8:51of the day it explodes and it cannot
  204. 8:53converges as expected. So, this is one
  205. 8:56of the technical challenges that we have
  206. 8:58to to adjust. And the solution to this
  207. 9:02is to introduce this new technique
  208. 9:04called
  209. 9:05QK clip. So, basically what it says is
  210. 9:07that for each attention head in this
  211. 9:10entire neural network, we're going to in
  212. 9:12the forward pass, we're going to compute
  213. 9:14the max logit. And then we're going to
  214. 9:16calculate a dividing factor that can be
  215. 9:19applied to each key projection as well
  216. 9:23as the query projection so that we can a
  217. 9:26sort of clip the maximum value of the
  218. 9:30query and the key to to sort of
  219. 9:32constrain it into a given range. So,
  220. 9:35that
  221. 9:37we're not going to have explosion
  222. 9:39anymore. So, these are some of the
  223. 9:41empirical results. On the left-hand
  224. 9:42side, there are two curves, but they are
  225. 9:44strictly overlapped with each other. So,
  226. 9:47these are the training curves before and
  227. 9:49after applying the clipping technique.
  228. 9:52So, you can see the clipping technique
  229. 9:54does not affect the training loss
  230. 9:58decrease at all. But on the right-hand
  231. 10:01side, if we inspect
  232. 10:03the intermediate magic, if we inspect
  233. 10:05the max logic, it's going to be
  234. 10:07effectively
  235. 10:08constrained. So, it first exposed as
  236. 10:12before, but at the value of 100, it's
  237. 10:15going to be clipped at the constant
  238. 10:17value for a long time. And then after a
  239. 10:19certain number of steps, it will just
  240. 10:21naturally go down.
  241. 10:22So, the neural network sort of
  242. 10:25find a way to constrain the maximum
  243. 10:28value of the max logic to ensure a
  244. 10:31stable training process. And at the same
  245. 10:34time, it doesn't affect, you know, the
  246. 10:36training convergence as shown in in the
  247. 10:39in the last figure.
  248. 10:40So, we employed this technique in our K2
  249. 10:44model training and successfully scaled
  250. 10:47it to 1 trillion parameters. And this is
  251. 10:51the first example of a large-scale
  252. 10:54million training in the history of
  253. 10:56machine learning.
  254. 10:58And the second dimension that we're very
  255. 11:00interested in is is long context.
  256. 11:04So, this is another figure. It's
  257. 11:05probably less known.
  258. 11:07It's It's one of the hidden gems in
  259. 11:10these papers.
  260. 11:11So, instead of just, you know, pushing
  261. 11:13down the training loss by training on
  262. 11:15more tokens, it has some
  263. 11:19insights from another perspective. So,
  264. 11:21as we can see, this is a comparison
  265. 11:23between transformers and LSTMs. So, on
  266. 11:26the left-hand side, we can see the
  267. 11:28transformers achieve a lower training
  268. 11:30loss given the same number of parameters
  269. 11:33and the same number of training tokens
  270. 11:34as expected. And this is why
  271. 11:36transformers become, you know, the you
  272. 11:38know, the sort of the de facto
  273. 11:39architecture that people are using right
  274. 11:41now. But on the right-hand side, it's
  275. 11:43really interesting to see that
  276. 11:46transformers are actually better because
  277. 11:48it can improve through the whole
  278. 11:50context. So, the x-axis is the token
  279. 11:53index in context. And if you increase
  280. 11:55the token index, you can see that the
  281. 11:57the training loss of transformers
  282. 11:59actually drop by a lot. If you just
  283. 12:02continue continually increase context
  284. 12:05length, the loss just continuously drops
  285. 12:07down. But if you look at, you know, the
  286. 12:09curve of LSTM, it just is saturated
  287. 12:13after a certain number of tokens. It
  288. 12:16means that transformers has this better
  289. 12:18capability of capturing longer context.
  290. 12:21And this is this is what makes it
  291. 12:23better.
  292. 12:25Because if you if you go back to like 10
  293. 12:27years ago, people use LSTM for tasks
  294. 12:30like machine translation, but it is not
  295. 12:32good for, for example, understanding
  296. 12:34entire code base or running a super long
  297. 12:37agent trajectories to solve,
  298. 12:40a a for example, writing Linux kernels
  299. 12:43from scratch. It's not going to be
  300. 12:45accomplishable by LSTMs. So, this is a
  301. 12:47very
  302. 12:49much needed capability in the era of
  303. 12:52agents because tasks are becoming harder
  304. 12:55and harder. And we need longer and
  305. 12:57longer contexts.
  306. 12:58So, the research idea here is to develop
  307. 13:01a better architecture so that we can
  308. 13:04efficiently scale to a longer context
  309. 13:07length and at the same time achieve a
  310. 13:10lower per token loss at larger token
  311. 13:14indices.
  312. 13:15And this is the motivation
  313. 13:18for which we introduce this new
  314. 13:20architecture called Kimilinear.
  315. 13:23And it contains this new
  316. 13:27linear attention variant called
  317. 13:30Kimidelta attention,
  318. 13:32which improves the original gated delta
  319. 13:36rule, GDR, by improve recurrent memory.
  320. 13:39I will show the details later. And at
  321. 13:41the same time, we're going to mix linear
  322. 13:43attention layers with full attention
  323. 13:45layers using a 1:2:3 ratio so that you
  324. 13:48can balance between this long context
  325. 13:51capabilities and at the same time having
  326. 13:54a more efficient
  327. 13:56implementation.
  328. 13:58So, this is some of the formulation. The
  329. 14:01basic idea is simple. If you look at
  330. 14:03linear attention,
  331. 14:05in the original formulation, the memory
  332. 14:08is going to be global. So, there is a
  333. 14:10global single decay factor that is
  334. 14:14applied along the way. So, it means that
  335. 14:17basically, if
  336. 14:19there are only two cases. In one case is
  337. 14:21In one case, you're going to forget
  338. 14:23basically everything and you're not
  339. 14:25going to retain any information. And in
  340. 14:27the second case, you can choose to
  341. 14:28retain, you know, almost everything, but
  342. 14:31at the same time, you you don't have the
  343. 14:32capability to leave out some of the
  344. 14:35unnecessary information in this long
  345. 14:37context. So, we introduce this key idea
  346. 14:40of having a fine-grained decay factor as
  347. 14:43shown in this highlighted alpha term.
  348. 14:46So, it's going to instead of being a
  349. 14:49scalar, it's going to be a a diagonal
  350. 14:51matrix, which controls
  351. 14:54the decay rate for each channel. So,
  352. 14:57that we can have two possibilities. For
  353. 14:58some of the channels, we can
  354. 15:00decay really, really slow, meaning that
  355. 15:03we can retain this long context
  356. 15:05information across a very long range.
  357. 15:08And at the same time, for the other
  358. 15:10channels, we can sort of quickly forget
  359. 15:13the information from the past indices to
  360. 15:16refresh it and observe new information.
  361. 15:19And this is
  362. 15:21to increase the
  363. 15:23expressivity of this model.
  364. 15:27And of course, to leverage modern GPUs,
  365. 15:29we have to use this chunk-wise
  366. 15:32formulation so that we can parallelize
  367. 15:35the computation on modern GPUs. So, the
  368. 15:38first equation here is the chunk
  369. 15:40chunk-wise
  370. 15:42formulation of Kimilinear.
  371. 15:45But as you can see, this is going to
  372. 15:47bring massive infrastructure challenges
  373. 15:50because of this newly introduced alpha
  374. 15:53term. Because now it is a matrix instead
  375. 15:56of a scalar, it cannot easily be
  376. 15:58factored out. So, to achieve an
  377. 16:00efficient implementation,
  378. 16:02we rewrite the entire equation into the
  379. 16:06the bottom three equations.
  380. 16:08So, we introduce this matrix inversion
  381. 16:11operation as well as introducing the
  382. 16:15cumulative decay factor so that we can
  383. 16:18implement this entire thing in parallel
  384. 16:20without sacrificing
  385. 16:22any efficiency.
  386. 16:24And more importantly, this is not an
  387. 16:27approximation. It's an exact
  388. 16:29mathematically equivalent formulation so
  389. 16:32that we can achieve much efficient
  390. 16:35implementation without sacrificing
  391. 16:37any loss in terms of performance.
  392. 16:40So, it's going to be as efficient as,
  393. 16:45you know, previous linear attention
  394. 16:47variants, but at the same time, much
  395. 16:48more expressive.
  396. 16:50So, these are some of the results that
  397. 16:52we obtained
  398. 16:53using fair comparison.
  399. 16:55So, on the left-hand side, we see the
  400. 16:57performance on two different types of
  401. 17:00tasks. So, MMA you is a short context
  402. 17:02task. So, for short context task,
  403. 17:05Kimilinear achieved a better performance
  404. 17:08compared to MLA and GDM.
  405. 17:11And at the same time, for longer context
  406. 17:13tasks such as ruler,
  407. 17:16Kimilinear is
  408. 17:17also better
  409. 17:19than the variants,
  410. 17:20the other variants, while being much
  411. 17:22more efficient compared to MLA.
  412. 17:25And when we scale the context length
  413. 17:27further to, for example, 1 million
  414. 17:29tokens or even longer, it's going to be
  415. 17:32much more efficient compared to
  416. 17:34the baselines. And this is also
  417. 17:37the first architecture that can
  418. 17:40outperform full attention across across
  419. 17:42the board, including short context
  420. 17:45tasks, long input tasks, and long output
  421. 17:47tasks.
  422. 17:49So, these are two key dimensions that
  423. 17:53we are interested in. And the third
  424. 17:55dimension
  425. 17:56is the agents swarm. So, here is a
  426. 17:59diagram to showcase how we design this
  427. 18:04agents swarm paradigm to solve some of
  428. 18:07the more complex tasks compared to
  429. 18:09single agent paradigms. So, here we have
  430. 18:13an orchestrator,
  431. 18:14or you can call it a main agent. It's
  432. 18:17responsible for orchestrating tasks. It
  433. 18:20has different options. For example, we
  434. 18:22can spawn
  435. 18:23a group of sub agents and assign new
  436. 18:25tasks to these sub agents. Or you can
  437. 18:28collect results
  438. 18:29from the return of this sub agents. And
  439. 18:32you can sort of performing at this
  440. 18:35process in an iterative way. And at the
  441. 18:38end of the day, you can accomplish a
  442. 18:40more complex task compared to using one
  443. 18:43single agent. And it's analogous to to
  444. 18:46human society. For example, if we build
  445. 18:48a a company, we need different roles.
  446. 18:51And we need, for example,
  447. 18:54orchestrator or maybe we need a CEO to
  448. 18:56to decompose and assign the tasks to
  449. 18:58different roles. And then at the end of
  450. 19:00the day,
  451. 19:01the entire organization is going to have
  452. 19:04to move towards this same goal. And
  453. 19:07here, for example, in this case, we have
  454. 19:09Maybe you have the AI researchers, you
  455. 19:11have the web developers, you have
  456. 19:13physics researchers, and they can study
  457. 19:15different topics. And at the end of the
  458. 19:17day, you just collect the results and
  459. 19:19spawn
  460. 19:21a group of fact-checkers and web
  461. 19:22developers and file downloaders to
  462. 19:26to assemble the results to into a single
  463. 19:29report.
  464. 19:32And this is another
  465. 19:34a perspective to to look at this new
  466. 19:37paradigm. So, the x-axis is the
  467. 19:40complexity of the task.
  468. 19:43And the y-axis is is the execution time.
  469. 19:46And the complexity of the task is
  470. 19:48measured by the accuracy of a group of
  471. 19:52models
  472. 19:53on such task. So we can see with agent
  473. 19:56swarms it's going to substantially
  474. 20:00increase reduce the execution time
  475. 20:03compared to
  476. 20:06compared to single agents. It's going to
  477. 20:08be more effective
  478. 20:10and this means that we can scale this
  479. 20:12agent swarm paradigm to for example if
  480. 20:15you run this agent swarms with 100 or
  481. 20:18maybe even 1000 sub agents you can
  482. 20:21accomplish a complex task within a
  483. 20:24certain period of of time that is
  484. 20:26tolerable for
  485. 20:28for to producing real economical value.
  486. 20:33And because we can certainly scale it in
  487. 20:35different dimensions. We can scale the
  488. 20:37input. For example we can download and
  489. 20:39read hundreds of sources or even maybe
  490. 20:42thousands of doses in parallel or you
  491. 20:44can output
  492. 20:46write a 100 page literature review
  493. 20:49in
  494. 20:50in parallel or you can take actions at
  495. 20:53scale. You can perform data analysis
  496. 20:56for 10 different tasks and also it is
  497. 20:59orchestration at scale. You have to
  498. 21:01learn to design sub tasks and aggregate
  499. 21:04the the results.
  500. 21:06And technically
  501. 21:07we define some new objective functions
  502. 21:10to guide the learning process of our
  503. 21:13agent swarm system. So there are three
  504. 21:16reward functions
  505. 21:18reward objectives that are
  506. 21:21considered here compared to
  507. 21:24the conventional single agent RL
  508. 21:26learning. So the first term is what we
  509. 21:29call the instantiation reward. It
  510. 21:32incentivizes sub agent instantiation to
  511. 21:35prevent
  512. 21:37this
  513. 21:38serial collapse phenomenon from
  514. 21:41happening. So basically we don't want it
  515. 21:44to default to single agent execution. We
  516. 21:47want to encourage the parallel
  517. 21:50executions especially
  518. 21:52when we
  519. 21:54when when this early stage in training
  520. 21:56and of course we can decay the weight
  521. 21:59for this instantiation reward term over
  522. 22:02training course because
  523. 22:04when it learns to
  524. 22:06learns parallel execution we can reduce
  525. 22:09the weight. And the second term here is
  526. 22:11is finished reward
  527. 22:13and it is used because we observe one of
  528. 22:16the things in training
  529. 22:18that
  530. 22:19some some of this
  531. 22:21sub tasks are just created but never
  532. 22:24finished. So it's almost like it's going
  533. 22:26to hack the first term by just spawning
  534. 22:29a bunch of sub agents and the task might
  535. 22:31be too complex or maybe the task just
  536. 22:33doesn't make sense. And here we use this
  537. 22:36finished reward to basically encourage
  538. 22:39that each of the sub task should have a
  539. 22:42relatively high ratio of
  540. 22:45completion instead of just spawning a
  541. 22:47bunch of pseudo tasks with needed to be
  542. 22:50meaningful.
  543. 22:51So this is the second term that we use
  544. 22:54and of course we use the same you know
  545. 22:56decay strategy. We use the relative high
  546. 22:58weight at the beginning of training and
  547. 23:00we decay to a relatively low weight at
  548. 23:03the end of training.
  549. 23:04And of course the third term is the
  550. 23:06standard term. It's it's the outcome
  551. 23:08reward. It's going to measure whether
  552. 23:11the entire task is completed
  553. 23:14and then we're going to add these three
  554. 23:16terms in our
  555. 23:17reinforcement learning
  556. 23:19system. And of course we have to build
  557. 23:21you know the entire infrastructure
  558. 23:23because
  559. 23:24right now you need to support the
  560. 23:26parallel execution and then you'll need
  561. 23:28to support different reward functions
  562. 23:31and and to you know maximize the
  563. 23:33efficiency of the entire agent swarm RL
  564. 23:36system.
  565. 23:38So here are three
  566. 23:39uh
  567. 23:40different things that that we have
  568. 23:42tried scaling. The Muon clip optimizer
  569. 23:46improves token efficiency
  570. 23:49and Kimi Delta attention in the Kimi
  571. 23:52linear architecture improves long
  572. 23:54context and we also have the agent
  573. 23:56swarms paradigm to further
  574. 24:00create a new dimension of scaling.
  575. 24:02And all of this put together we created
  576. 24:06Kimi K2.5 a new model that we just
  577. 24:09released
  578. 24:10over 1 month ago.
  579. 24:12Here's a short video to demonstrate some
  580. 24:14of its capabilities.
  581. 24:25>> [music]
  582. 24:35[music]
  583. 24:39[music]
  584. 24:46[music]
  585. 24:52[music]
  586. 24:57[music]
  587. 25:16[applause]
  588. 25:19>> So
  589. 25:20yeah there are a lot of interesting
  590. 25:22things capabilities that we discover
  591. 25:24from the model. For example
  592. 25:26it merges the visual capabilities with
  593. 25:29coding capabilities. So a lot of new
  594. 25:31things just emerge out of it. It can
  595. 25:34read a video and then produce a website
  596. 25:37that sort of replicates or style
  597. 25:40transfer the original video.
  598. 25:43And all of this are due to successful
  599. 25:47and stable training
  600. 25:49at the pre-training stage. So this is
  601. 25:50also one of the
  602. 25:52most beautiful curves that I observed in
  603. 25:55my life. So this is the training curve
  604. 25:57of the K2.5 base model. So as you can
  605. 26:02see it went through over 15 trillion
  606. 26:04tokens and of course in K2.5 we
  607. 26:06additionally trained another 15 trillion
  608. 26:08tokens and the entire
  609. 26:11the training process is just so stable.
  610. 26:14There's no loss spike especially when we
  611. 26:17introduce this new Muon optimizer we
  612. 26:19didn't observe any spike and this smooth
  613. 26:22stable training process produces a very
  614. 26:25stable outcome
  615. 26:27a very strong base model that we can
  616. 26:28fine tune on top of it to achieve you
  617. 26:32know new capabilities such as we
  618. 26:34introduced and shown in the video the
  619. 26:36video. And this is also of course
  620. 26:38uh
  621. 26:39trained on Nvidia H100 GPUs and each
  622. 26:43node in this H100 cluster contains two
  623. 26:47TB RAM and eight GPUs
  624. 26:49connected by NVLink.
  625. 26:52And uh one of the another you know key
  626. 26:55innovation of Kimi K2.5 is that it is
  627. 26:58the first open model with native joint
  628. 27:02vision text capabilities. So if you look
  629. 27:04at previous open models usually their
  630. 27:07visual capabilities are added on top of
  631. 27:10a text base meaning that for example if
  632. 27:12you train the text models for 20
  633. 27:15trillion tokens and then on top of it
  634. 27:17you do another two trillion sort of a
  635. 27:20post training process to add additional
  636. 27:22visual capabilities on top of it.
  637. 27:25But for K2.5 it's different in the sense
  638. 27:28that we fuse the training process of
  639. 27:31vision and text from day one. So it's
  640. 27:34called early fusion here. We start from
  641. 27:37you know 0% of the progress. So from day
  642. 27:40one we're going to merge the vision and
  643. 27:42text tokens and as shown in our
  644. 27:44preliminary experiments it outperforms
  645. 27:48late fusion and some of the new
  646. 27:50capabilities that we observe also come
  647. 27:52from this training recipe. For example
  648. 27:56if you want to do vision to code you
  649. 27:58really have to merge vision and text
  650. 28:01into a single brand to achieve that. If
  651. 28:04you separate these two brands it's not
  652. 28:06going to happen. You have to align these
  653. 28:08two modalities into a share embedding
  654. 28:10space
  655. 28:12in
  656. 28:13a share representation space
  657. 28:15so as to achieve this.
  658. 28:17And another interesting thing that we
  659. 28:19observe is that
  660. 28:21these two modalities can actually
  661. 28:23enhance each other. So that's
  662. 28:26that's been long been a challenge that
  663. 28:29if you add vision capabilities into a
  664. 28:31text model it's going to somewhat
  665. 28:34hurt the text performance. But here we
  666. 28:37found that if you train it properly
  667. 28:39these two modalities can actually
  668. 28:41enhance each other. So this is one of
  669. 28:43the key findings that
  670. 28:45we observe in in in our training. So
  671. 28:48first vision improves text. So this is
  672. 28:50so interesting. So before vision RL
  673. 28:54the performance in the first column and
  674. 28:56then we have the performance after
  675. 28:58vision RL. So here vision RL
  676. 29:01refers to a process that we only use
  677. 29:04vision task. So there is no text task
  678. 29:07involved here. We only have vision task.
  679. 29:09For example we teach the model how to
  680. 29:11how to count how to answer some of this
  681. 29:14visual QA
  682. 29:16problems without any for example math
  683. 29:19any coding problems in in in this space.
  684. 29:22But we observe that it's going to
  685. 29:23improve the performance for even you
  686. 29:25know reasonably heavy text task.
  687. 29:28And on the other hand text also improves
  688. 29:30vision. If you have a very strong text
  689. 29:32base
  690. 29:33you you actually don't need any vision
  691. 29:36SFT data in the training process and
  692. 29:38this is the approach that we adopt. So
  693. 29:41it's called zero vision SFT. Basically
  694. 29:44we don't have We have basically zero
  695. 29:46vision SFT data and the only SFT data
  696. 29:49that we have is the text SFT data and
  697. 29:52then we do a joint IO over text and
  698. 29:54vision and you can see that we can
  699. 29:56achieve almost state-of-the-art
  700. 29:58performance across the board on on
  701. 30:00vision task without any vision data. So,
  702. 30:03it is clear that
  703. 30:05if you have a strong text base is also
  704. 30:06going to improve the vision if if you
  705. 30:10align these two modalities into a shared
  706. 30:13space in your in your pre-training.
  707. 30:17And also
  708. 30:19these are some of the examples of uh
  709. 30:23uh
  710. 30:25Yeah, as I was showing the video so it's
  711. 30:28it demonstrates strong capabilities of
  712. 30:31visual design and front-end coding and
  713. 30:33this also emerges from our vision text
  714. 30:36pre-training.
  715. 30:38So,
  716. 30:39uh after all this so this are all about
  717. 30:42Kim E 8.5 and as probably
  718. 30:46you already know we released our new
  719. 30:48architecture yesterday
  720. 30:51in our tech report is called attention
  721. 30:53residue. So, here I'm also going to
  722. 30:56briefly talk about our new work which
  723. 30:59serves as a a sneak peek into our next
  724. 31:01generation architecture that we're
  725. 31:03probably going to adopt in in our later
  726. 31:05models.
  727. 31:07So, here the motivation is is quite
  728. 31:10simple. Can we apply some of our
  729. 31:12techniques that we use in in the in the
  730. 31:15temporal dimension and we we just take
  731. 31:19some of the inspirations and we apply to
  732. 31:21the depth
  733. 31:22dimension. So,
  734. 31:25and it starts from this residual
  735. 31:27connection. So, I still remember
  736. 31:30listening to to Kaiming's talk at the
  737. 31:34tutorial in ICML 2016 10 years ago.
  738. 31:38So, it was a brilliant idea. So,
  739. 31:40basically before ResNet nobody was able
  740. 31:44to train deep networks.
  741. 31:46If you increase the depth if you
  742. 31:48increase the number of layers for neural
  743. 31:50networks nobody was able to train it
  744. 31:53because you will observe this gradient
  745. 31:55explosion gradient vanishing all these
  746. 31:57you know stability issues. But then
  747. 31:59after the introduction of of ResNet we
  748. 32:02can train you know an arbitrary large
  749. 32:05number of layers you can stack as many
  750. 32:07layers as as you want and you you you
  751. 32:09don't have to worry about the training
  752. 32:11stability issue
  753. 32:13and stuff. And as discussed in Ilya's
  754. 32:16talk uh 2 years ago it basically says
  755. 32:19that residual connection is a variant of
  756. 32:23LSTM but just rotated 90 degrees. So,
  757. 32:27how do you understand this? If you look
  758. 32:28at LSTM is is a variant of recurrent
  759. 32:31net, right? And it's a recurrent model
  760. 32:35process. So, we're going to take the
  761. 32:37hidden states from the last step and
  762. 32:41then we're going to have some gating
  763. 32:43mechanism some function to produce the
  764. 32:45current state. Right? And if you look at
  765. 32:48the
  766. 32:49the depth dimension the residual
  767. 32:51connection is basically the same. We're
  768. 32:53going to take the output from the last
  769. 32:55layer and then we're going to apply some
  770. 32:57sort of function on top of it to produce
  771. 33:00the current
  772. 33:03the current output of the current layer.
  773. 33:06It's just the formulation is different.
  774. 33:07For example, for residual connection
  775. 33:09we're going to use a fixed addition
  776. 33:11we're going to have this
  777. 33:14we're going to add the previous hidden
  778. 33:16state
  779. 33:17with the current output. It's just the
  780. 33:20formulation that's different but the
  781. 33:21basic idea is the same it's a recurrent
  782. 33:23net applied
  783. 33:25in the dimension of depth.
  784. 33:27And but on the other hand
  785. 33:29we can think about
  786. 33:31reformulating
  787. 33:33this
  788. 33:34this function. Instead of having LSTM
  789. 33:37can we have an attention in the
  790. 33:39dimension of depth and it's going to
  791. 33:41create new possibilities because
  792. 33:43attention have has been demonstrated to
  793. 33:46be so successful
  794. 33:48in the transformer era. So, what we're
  795. 33:50going to do is not just to take the last
  796. 33:53hidden state but we're going to consider
  797. 33:55all the previous hidden states and use
  798. 33:57the attention operation the attention
  799. 33:59mechanism to assemble and aggregate
  800. 34:02all of these previous hidden states to
  801. 34:04compute the current state. So, this is
  802. 34:07exactly attention rotated by 90 degrees.
  803. 34:12It's sort of we view as a natural
  804. 34:14generalization of residual connections
  805. 34:17in the LSTM
  806. 34:19analogy.
  807. 34:21Okay, and here's the detail formulation.
  808. 34:23So, on on the left-hand side is a
  809. 34:25standard residue. As I said it's
  810. 34:27basically LSTM rotated by 90 degrees and
  811. 34:30the second figure is attention rotated
  812. 34:33by 90 degrees. So, what we do is to
  813. 34:36collect all the previous hidden states
  814. 34:38and have a simple attention operation on
  815. 34:41top of it to produce the current layers
  816. 34:43outcome. And of course to increase the
  817. 34:47efficiency to reduce the infrastructure
  818. 34:50for example communication and memory
  819. 34:51overhead we also design a new variant
  820. 34:54called
  821. 34:55block attention residue on the
  822. 34:57right-hand side. So, basically the idea
  823. 34:59is is also simple. We're going to divide
  824. 35:04all the layers in the neural network
  825. 35:06into multiple blocks. For example, each
  826. 35:08block can contain say 16 layers or it
  827. 35:11can contain maybe four layers and then
  828. 35:13for each block we're going to
  829. 35:16apply
  830. 35:17this attention residue only on the
  831. 35:19output of each block but within each
  832. 35:21block we also we still adopt this
  833. 35:24standard residue. So, this is going to
  834. 35:26reduce a lot of overhead while having
  835. 35:28minimal loss in terms of training
  836. 35:30accuracy.
  837. 35:32And these are some of the impressive
  838. 35:34results that we achieved
  839. 35:37on this new architecture. So, on the
  840. 35:39scaling law we can improve the token
  841. 35:43efficiency by 24%
  842. 35:46meaning that if you have 50 trillion
  843. 35:49high-quality tokens now you just
  844. 35:51magically have
  845. 35:53over 60 trillion tokens and then for the
  846. 35:58validation loss you can also observe
  847. 36:00that
  848. 36:01it's consistently lower than
  849. 36:04the original curve
  850. 36:07demonstrating the stability across
  851. 36:09optimization and also achieved the best
  852. 36:12improvement on some of this coding math
  853. 36:16and reasoning heavy task as shown in the
  854. 36:18benchmark results of GPQA math and human
  855. 36:22eval.
  856. 36:24So, the entire community keeps moving
  857. 36:27forward
  858. 36:28and we're happy that we can we're able
  859. 36:30to contribute to to the community with
  860. 36:33you know new technologies and some of
  861. 36:35this
  862. 36:37have you know some of this technologies
  863. 36:39have been sort of standard and de facto
  864. 36:41for a long time but as you can see we
  865. 36:44still see a lot of opportunities to
  866. 36:46improve it to
  867. 36:48to have revolutionary new design to
  868. 36:51achieve better performance. If you
  869. 36:53multiply all these gains together you
  870. 36:56can actually have a much better model.
  871. 36:59So, Adam was invented in 2014
  872. 37:02and now we scale an open source Neon
  873. 37:04clip a a dropping replacement for
  874. 37:08for Adam and I'm sure that if you're
  875. 37:11training transformer LLM it's going to
  876. 37:14be much better if you use Neon clip
  877. 37:16instead of Adam. And attention was
  878. 37:18invented over 8 years ago and then now
  879. 37:21we have Kim E linear which is a linear
  880. 37:24version. We don't have to use full
  881. 37:26attention across all layers. We can have
  882. 37:29linear attention that performs better on
  883. 37:31short long context at the same time.
  884. 37:34And also residual connections are now uh
  885. 37:38also challenged.
  886. 37:40We scale an open source attention
  887. 37:42residue. So, I think one of the
  888. 37:46interesting things about our era is that
  889. 37:49we sort of adopt a different mindset for
  890. 37:52for doing research. So, if we go back to
  891. 37:5510 years ago it's mostly about
  892. 37:57publishing a new idea
  893. 37:59but then I think the lack of the rigor
  894. 38:02of the experiments is very hard to
  895. 38:05produce
  896. 38:06solid experimental results. But now we
  897. 38:09have the scaling ladder. We have enough
  898. 38:11resources to you know train the model
  899. 38:14and running on different at different we
  900. 38:16can have you know a whole set of
  901. 38:18benchmarks to measure the progress. So,
  902. 38:21it is it becomes easier to make
  903. 38:24confident and solid conclusion out of
  904. 38:26it. And this is one of the reason why we
  905. 38:29are observing you know new progress
  906. 38:32on this
  907. 38:33ancient techniques and I'm sure that
  908. 38:35we'll see more and more especially in
  909. 38:38the open source community. I think we're
  910. 38:40going to have more and more even better
  911. 38:43architectural and you know optimization
  912. 38:46improvement in in the next few years.
  913. 38:49All right, so to summarize we're going
  914. 38:51to keep scaling our models and so this
  915. 38:55are three dimensions.
  916. 38:57For example, we we see
  917. 39:00we we see different architectures and
  918. 39:03optimizers that optimize all three you
  919. 39:06know dimensions and we'll keep you know
  920. 39:08see new dimensions for scaling.
  921. 39:11Agents forms is is not the end and we
  922. 39:14are glad that we can
  923. 39:16move forward with the entire open source
  924. 39:18community to achieve better and better
  925. 39:21you know intelligence. Thank you so
  926. 39:23much.
  927. 39:25>> [applause]

About this transcript

This page contains the full transcript of How We Scaled Kimi K2.5 | Zhilin Yang's full GTC 2026 Keynote by Kimi AI, generated from the public captions YouTube serves with the video. The transcript has 5,592 words across 927 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.