How We Scaled Kimi K2.5 | Zhilin Yang's full GTC 2026 Keynote — Transcript
Full transcript
- 0:07[music]
- 0:11>> Hi everyone.
- 0:13Thank you so much for the introduction.
- 0:15It's great to be here.
- 0:17It's great to have this opportunity to
- 0:20share with you guys some of our latest
- 0:23progress and efforts.
- 0:26So,
- 0:27one of our major pursues
- 0:30is to build better open models. And we
- 0:34believe in democratizing intelligence.
- 0:36With open models, you can deploy
- 0:38anywhere. It can be on your local
- 0:39servers. It can be on the cloud. And you
- 0:42can access every single bit of the
- 0:45weights in the modeling instead of just
- 0:48using a black box. And this is one of
- 0:50the slides that I took from Jensen's
- 0:53talk earlier this year at CES. So, as
- 0:56you can see, open models are quickly
- 0:58closing the gap with proprietary models
- 1:02and it's reaching the frontier. And we
- 1:04believe that with better and better open
- 1:06models, we're going to
- 1:08make intelligence more accessible to
- 1:10anybody in the world, in every corner of
- 1:13the world.
- 1:15But open models cannot be just open.
- 1:18They have have also to be great. So, in
- 1:23this talk, we're going to discuss how we
- 1:25make open models great. So, as we know,
- 1:29scaling is a primary driver
- 1:33for a lot of progress, maybe all of the
- 1:35you know major AI developments that we
- 1:37have witnessed in in the last few years.
- 1:40And here we're going to discuss how we
- 1:42scale our model in different dimensions.
- 1:45So, on the left-hand side, the first
- 1:47figure you see here is kind of the the
- 1:50standard scaling law. So, you on the
- 1:52x-axis you have the
- 1:54log of the number of training tokens.
- 1:56And on the y-axis you have the log loss.
- 1:59And as you scale the number of training
- 2:01tokens, you get a lower loss. But here
- 2:03the point is
- 2:05we're not going to just scale the number
- 2:06of training tokens, but we also want to
- 2:09want to improve the token efficiency.
- 2:13Meaning that we want to move this curve
- 2:15to the left-hand side so that we can
- 2:18achieve a lower loss, a much lower loss
- 2:21using the same number of training
- 2:23tokens. And this can be achieved by
- 2:25having better architectures and
- 2:28optimizers as we'll discuss in in our
- 2:31latest slides.
- 2:32And the second scaling dimensions that
- 2:34we're very interested in is to scale the
- 2:37context length. So, as you can see in
- 2:39the second second figure, if we increase
- 2:42the context length,
- 2:44then we can have a much higher accuracy
- 2:47in terms of predicting
- 2:50the token loss at a given position.
- 2:54And this means that we can increase the
- 2:56capability of the model to achieve more
- 3:00complex tasks by increasing the context
- 3:02length. So, this is the second scaling
- 3:05dimensions that we're going to going to
- 3:06talk about. And the third scaling
- 3:08dimension is the number of agents. So,
- 3:10we introduce this new learning paradigm
- 3:13of agent swarms where we can we don't
- 3:16just rely on a single agent, but we also
- 3:19orchestrate a swarm of agents that can
- 3:22accomplish the subtasks in parallel so
- 3:25that we can increase the task
- 3:26complexity. And we can translate all of
- 3:28this in into the language of agents. So,
- 3:31if you look at token efficiency, it's
- 3:33mostly about having a stronger prior so
- 3:37that you can
- 3:39have more efficiency when you do agent
- 3:41RL to search for a better solution. And
- 3:44when you think about long context, it's
- 3:46it's mostly about increasing the context
- 3:48length so that you can have a longer
- 3:50running agent. It can probably run for
- 3:52days or even weeks or months to
- 3:54accomplish more
- 3:56more more tasks, more complex tasks.
- 3:59And about and for agent swarms, it's
- 4:02another dimension that that added it and
- 4:05at the end of the day we're going to
- 4:06have a swarm of agents that each of them
- 4:08have a super long context and each of
- 4:10them have a very strong prior for us to
- 4:13search in this entire agent RL system.
- 4:19All right. So, we're going to start from
- 4:22token efficiency. So, this is one of the
- 4:24most you know classical figures in the
- 4:27history of machine learning. All right.
- 4:29So, it's taken from Kaplan et al. And
- 4:31basically says that if we scale
- 4:34proportionally the number of train
- 4:36tokens, the model parameters, and also
- 4:39the amount of compute, we can get lower
- 4:41and lower loss. And this is you know one
- 4:44of the major breakthroughs that that the
- 4:46entire community has achieved in the
- 4:49last few years to to get
- 4:51better intelligence. But here, what
- 4:54we're interested in is to have better
- 4:57and better token efficiency.
- 4:59And here's the thing. So, one thing that
- 5:02I would like to emphasize is that token
- 5:05efficiency is not just about efficiency.
- 5:08It's actually also about improving the
- 5:11upper bound of intelligence. So, here's
- 5:14here's why. So, suppose you have a
- 5:18say 50 trillion tokens, 50 trillion
- 5:21high-quality tokens.
- 5:23And then you apply this new optimizer,
- 5:25maybe the Meow optimizer. And then all
- 5:28of a sudden you have a two-times token
- 5:30efficiency. So, it means that
- 5:32it's almost like magic that you get
- 5:35equivalently 100 trillion tokens.
- 5:38And nowadays we're scaling towards the
- 5:41data wall and we're hitting you know the
- 5:43data wall and the amount of high-quality
- 5:46data is quite limited. And if we suppose
- 5:48that it's a constant amount, then we
- 5:51increase the token efficiency,
- 5:53it means that we're going to get better
- 5:55intelligence out of it. It's not just
- 5:57about infrastructure efficiency. It's
- 5:59about you know better
- 6:01intelligence. So, so this is why we
- 6:04spend you know a lot of efforts in this
- 6:07aspect because it's going to push the
- 6:09frontier of of intelligence. And Meow
- 6:12optimizer is one of the things that we
- 6:14have heavily invested in
- 6:16since last year.
- 6:18So, it's a second-order
- 6:21optimizer. And basically every single
- 6:24gradient update is transformed in a way
- 6:26that each entry is orthogonal to each
- 6:29other. And this is very different from
- 6:32the traditional Adam optimizer. And if
- 6:35if you implement this optimizer
- 6:37properly, you can get a two-times token
- 6:39efficient efficiency improvement. So, we
- 6:43we are one we are the first
- 6:45work we published the first work to
- 6:47demonstrate that Meow optimizer is
- 6:50actually scalable for LLM training. And
- 6:53these are two key techniques that we
- 6:56employed to make it effective for
- 6:58large-scale training. So, one of them is
- 7:00weight decay. It is critical for scaling
- 7:03to larger models. And the second is we
- 7:06want to ensure a consistent RMS updates
- 7:09compared to Adam. So, we have this
- 7:12adjustable coefficients that is applied
- 7:15to each update so that the resulting RMS
- 7:19is going to be comparable to Adam.
- 7:22And to make Meow memory efficient across
- 7:26all these Nvidia GPU clusters, we also
- 7:29develop a distributed Meow optimizer
- 7:32implementation that partitions the
- 7:34states across the data parallel group so
- 7:37that we can have a very efficient
- 7:40implementation for the Meow optimizer.
- 7:43And these are some of the results that
- 7:44were presented in the paper. So, as you
- 7:47can see, with with the same number of
- 7:49parameters and the same number of
- 7:51training tokens, we just replace the
- 7:53original AdamW optimizer with the new
- 7:56Meow optimizer. It's going to to improve
- 7:59the performance across the board
- 8:01significantly.
- 8:04But there was this new challenge that we
- 8:07encountered when we tried to scale it up
- 8:10further. When we tried to scale Meow for
- 8:12a one trillion parameter model, we
- 8:14encountered a new issue
- 8:17about training instability. So, as you
- 8:19can see on the left figure,
- 8:21we we observed that the max logits
- 8:24quickly explodes and quickly exceeds
- 8:281,000. And the typical values for
- 8:31for training for this max logits is
- 8:35about say 50 or maybe less than 100. But
- 8:39for for Meow, it quickly exceeds 1,000.
- 8:42And at the same time, we observe
- 8:44training divergence on the left-hand
- 8:47side. If you look at the training loss,
- 8:49it goes down a bit, but then at the end
- 8:51of the day it explodes and it cannot
- 8:53converges as expected. So, this is one
- 8:56of the technical challenges that we have
- 8:58to to adjust. And the solution to this
- 9:02is to introduce this new technique
- 9:04called
- 9:05QK clip. So, basically what it says is
- 9:07that for each attention head in this
- 9:10entire neural network, we're going to in
- 9:12the forward pass, we're going to compute
- 9:14the max logit. And then we're going to
- 9:16calculate a dividing factor that can be
- 9:19applied to each key projection as well
- 9:23as the query projection so that we can a
- 9:26sort of clip the maximum value of the
- 9:30query and the key to to sort of
- 9:32constrain it into a given range. So,
- 9:35that
- 9:37we're not going to have explosion
- 9:39anymore. So, these are some of the
- 9:41empirical results. On the left-hand
- 9:42side, there are two curves, but they are
- 9:44strictly overlapped with each other. So,
- 9:47these are the training curves before and
- 9:49after applying the clipping technique.
- 9:52So, you can see the clipping technique
- 9:54does not affect the training loss
- 9:58decrease at all. But on the right-hand
- 10:01side, if we inspect
- 10:03the intermediate magic, if we inspect
- 10:05the max logic, it's going to be
- 10:07effectively
- 10:08constrained. So, it first exposed as
- 10:12before, but at the value of 100, it's
- 10:15going to be clipped at the constant
- 10:17value for a long time. And then after a
- 10:19certain number of steps, it will just
- 10:21naturally go down.
- 10:22So, the neural network sort of
- 10:25find a way to constrain the maximum
- 10:28value of the max logic to ensure a
- 10:31stable training process. And at the same
- 10:34time, it doesn't affect, you know, the
- 10:36training convergence as shown in in the
- 10:39in the last figure.
- 10:40So, we employed this technique in our K2
- 10:44model training and successfully scaled
- 10:47it to 1 trillion parameters. And this is
- 10:51the first example of a large-scale
- 10:54million training in the history of
- 10:56machine learning.
- 10:58And the second dimension that we're very
- 11:00interested in is is long context.
- 11:04So, this is another figure. It's
- 11:05probably less known.
- 11:07It's It's one of the hidden gems in
- 11:10these papers.
- 11:11So, instead of just, you know, pushing
- 11:13down the training loss by training on
- 11:15more tokens, it has some
- 11:19insights from another perspective. So,
- 11:21as we can see, this is a comparison
- 11:23between transformers and LSTMs. So, on
- 11:26the left-hand side, we can see the
- 11:28transformers achieve a lower training
- 11:30loss given the same number of parameters
- 11:33and the same number of training tokens
- 11:34as expected. And this is why
- 11:36transformers become, you know, the you
- 11:38know, the sort of the de facto
- 11:39architecture that people are using right
- 11:41now. But on the right-hand side, it's
- 11:43really interesting to see that
- 11:46transformers are actually better because
- 11:48it can improve through the whole
- 11:50context. So, the x-axis is the token
- 11:53index in context. And if you increase
- 11:55the token index, you can see that the
- 11:57the training loss of transformers
- 11:59actually drop by a lot. If you just
- 12:02continue continually increase context
- 12:05length, the loss just continuously drops
- 12:07down. But if you look at, you know, the
- 12:09curve of LSTM, it just is saturated
- 12:13after a certain number of tokens. It
- 12:16means that transformers has this better
- 12:18capability of capturing longer context.
- 12:21And this is this is what makes it
- 12:23better.
- 12:25Because if you if you go back to like 10
- 12:27years ago, people use LSTM for tasks
- 12:30like machine translation, but it is not
- 12:32good for, for example, understanding
- 12:34entire code base or running a super long
- 12:37agent trajectories to solve,
- 12:40a a for example, writing Linux kernels
- 12:43from scratch. It's not going to be
- 12:45accomplishable by LSTMs. So, this is a
- 12:47very
- 12:49much needed capability in the era of
- 12:52agents because tasks are becoming harder
- 12:55and harder. And we need longer and
- 12:57longer contexts.
- 12:58So, the research idea here is to develop
- 13:01a better architecture so that we can
- 13:04efficiently scale to a longer context
- 13:07length and at the same time achieve a
- 13:10lower per token loss at larger token
- 13:14indices.
- 13:15And this is the motivation
- 13:18for which we introduce this new
- 13:20architecture called Kimilinear.
- 13:23And it contains this new
- 13:27linear attention variant called
- 13:30Kimidelta attention,
- 13:32which improves the original gated delta
- 13:36rule, GDR, by improve recurrent memory.
- 13:39I will show the details later. And at
- 13:41the same time, we're going to mix linear
- 13:43attention layers with full attention
- 13:45layers using a 1:2:3 ratio so that you
- 13:48can balance between this long context
- 13:51capabilities and at the same time having
- 13:54a more efficient
- 13:56implementation.
- 13:58So, this is some of the formulation. The
- 14:01basic idea is simple. If you look at
- 14:03linear attention,
- 14:05in the original formulation, the memory
- 14:08is going to be global. So, there is a
- 14:10global single decay factor that is
- 14:14applied along the way. So, it means that
- 14:17basically, if
- 14:19there are only two cases. In one case is
- 14:21In one case, you're going to forget
- 14:23basically everything and you're not
- 14:25going to retain any information. And in
- 14:27the second case, you can choose to
- 14:28retain, you know, almost everything, but
- 14:31at the same time, you you don't have the
- 14:32capability to leave out some of the
- 14:35unnecessary information in this long
- 14:37context. So, we introduce this key idea
- 14:40of having a fine-grained decay factor as
- 14:43shown in this highlighted alpha term.
- 14:46So, it's going to instead of being a
- 14:49scalar, it's going to be a a diagonal
- 14:51matrix, which controls
- 14:54the decay rate for each channel. So,
- 14:57that we can have two possibilities. For
- 14:58some of the channels, we can
- 15:00decay really, really slow, meaning that
- 15:03we can retain this long context
- 15:05information across a very long range.
- 15:08And at the same time, for the other
- 15:10channels, we can sort of quickly forget
- 15:13the information from the past indices to
- 15:16refresh it and observe new information.
- 15:19And this is
- 15:21to increase the
- 15:23expressivity of this model.
- 15:27And of course, to leverage modern GPUs,
- 15:29we have to use this chunk-wise
- 15:32formulation so that we can parallelize
- 15:35the computation on modern GPUs. So, the
- 15:38first equation here is the chunk
- 15:40chunk-wise
- 15:42formulation of Kimilinear.
- 15:45But as you can see, this is going to
- 15:47bring massive infrastructure challenges
- 15:50because of this newly introduced alpha
- 15:53term. Because now it is a matrix instead
- 15:56of a scalar, it cannot easily be
- 15:58factored out. So, to achieve an
- 16:00efficient implementation,
- 16:02we rewrite the entire equation into the
- 16:06the bottom three equations.
- 16:08So, we introduce this matrix inversion
- 16:11operation as well as introducing the
- 16:15cumulative decay factor so that we can
- 16:18implement this entire thing in parallel
- 16:20without sacrificing
- 16:22any efficiency.
- 16:24And more importantly, this is not an
- 16:27approximation. It's an exact
- 16:29mathematically equivalent formulation so
- 16:32that we can achieve much efficient
- 16:35implementation without sacrificing
- 16:37any loss in terms of performance.
- 16:40So, it's going to be as efficient as,
- 16:45you know, previous linear attention
- 16:47variants, but at the same time, much
- 16:48more expressive.
- 16:50So, these are some of the results that
- 16:52we obtained
- 16:53using fair comparison.
- 16:55So, on the left-hand side, we see the
- 16:57performance on two different types of
- 17:00tasks. So, MMA you is a short context
- 17:02task. So, for short context task,
- 17:05Kimilinear achieved a better performance
- 17:08compared to MLA and GDM.
- 17:11And at the same time, for longer context
- 17:13tasks such as ruler,
- 17:16Kimilinear is
- 17:17also better
- 17:19than the variants,
- 17:20the other variants, while being much
- 17:22more efficient compared to MLA.
- 17:25And when we scale the context length
- 17:27further to, for example, 1 million
- 17:29tokens or even longer, it's going to be
- 17:32much more efficient compared to
- 17:34the baselines. And this is also
- 17:37the first architecture that can
- 17:40outperform full attention across across
- 17:42the board, including short context
- 17:45tasks, long input tasks, and long output
- 17:47tasks.
- 17:49So, these are two key dimensions that
- 17:53we are interested in. And the third
- 17:55dimension
- 17:56is the agents swarm. So, here is a
- 17:59diagram to showcase how we design this
- 18:04agents swarm paradigm to solve some of
- 18:07the more complex tasks compared to
- 18:09single agent paradigms. So, here we have
- 18:13an orchestrator,
- 18:14or you can call it a main agent. It's
- 18:17responsible for orchestrating tasks. It
- 18:20has different options. For example, we
- 18:22can spawn
- 18:23a group of sub agents and assign new
- 18:25tasks to these sub agents. Or you can
- 18:28collect results
- 18:29from the return of this sub agents. And
- 18:32you can sort of performing at this
- 18:35process in an iterative way. And at the
- 18:38end of the day, you can accomplish a
- 18:40more complex task compared to using one
- 18:43single agent. And it's analogous to to
- 18:46human society. For example, if we build
- 18:48a a company, we need different roles.
- 18:51And we need, for example,
- 18:54orchestrator or maybe we need a CEO to
- 18:56to decompose and assign the tasks to
- 18:58different roles. And then at the end of
- 19:00the day,
- 19:01the entire organization is going to have
- 19:04to move towards this same goal. And
- 19:07here, for example, in this case, we have
- 19:09Maybe you have the AI researchers, you
- 19:11have the web developers, you have
- 19:13physics researchers, and they can study
- 19:15different topics. And at the end of the
- 19:17day, you just collect the results and
- 19:19spawn
- 19:21a group of fact-checkers and web
- 19:22developers and file downloaders to
- 19:26to assemble the results to into a single
- 19:29report.
- 19:32And this is another
- 19:34a perspective to to look at this new
- 19:37paradigm. So, the x-axis is the
- 19:40complexity of the task.
- 19:43And the y-axis is is the execution time.
- 19:46And the complexity of the task is
- 19:48measured by the accuracy of a group of
- 19:52models
- 19:53on such task. So we can see with agent
- 19:56swarms it's going to substantially
- 20:00increase reduce the execution time
- 20:03compared to
- 20:06compared to single agents. It's going to
- 20:08be more effective
- 20:10and this means that we can scale this
- 20:12agent swarm paradigm to for example if
- 20:15you run this agent swarms with 100 or
- 20:18maybe even 1000 sub agents you can
- 20:21accomplish a complex task within a
- 20:24certain period of of time that is
- 20:26tolerable for
- 20:28for to producing real economical value.
- 20:33And because we can certainly scale it in
- 20:35different dimensions. We can scale the
- 20:37input. For example we can download and
- 20:39read hundreds of sources or even maybe
- 20:42thousands of doses in parallel or you
- 20:44can output
- 20:46write a 100 page literature review
- 20:49in
- 20:50in parallel or you can take actions at
- 20:53scale. You can perform data analysis
- 20:56for 10 different tasks and also it is
- 20:59orchestration at scale. You have to
- 21:01learn to design sub tasks and aggregate
- 21:04the the results.
- 21:06And technically
- 21:07we define some new objective functions
- 21:10to guide the learning process of our
- 21:13agent swarm system. So there are three
- 21:16reward functions
- 21:18reward objectives that are
- 21:21considered here compared to
- 21:24the conventional single agent RL
- 21:26learning. So the first term is what we
- 21:29call the instantiation reward. It
- 21:32incentivizes sub agent instantiation to
- 21:35prevent
- 21:37this
- 21:38serial collapse phenomenon from
- 21:41happening. So basically we don't want it
- 21:44to default to single agent execution. We
- 21:47want to encourage the parallel
- 21:50executions especially
- 21:52when we
- 21:54when when this early stage in training
- 21:56and of course we can decay the weight
- 21:59for this instantiation reward term over
- 22:02training course because
- 22:04when it learns to
- 22:06learns parallel execution we can reduce
- 22:09the weight. And the second term here is
- 22:11is finished reward
- 22:13and it is used because we observe one of
- 22:16the things in training
- 22:18that
- 22:19some some of this
- 22:21sub tasks are just created but never
- 22:24finished. So it's almost like it's going
- 22:26to hack the first term by just spawning
- 22:29a bunch of sub agents and the task might
- 22:31be too complex or maybe the task just
- 22:33doesn't make sense. And here we use this
- 22:36finished reward to basically encourage
- 22:39that each of the sub task should have a
- 22:42relatively high ratio of
- 22:45completion instead of just spawning a
- 22:47bunch of pseudo tasks with needed to be
- 22:50meaningful.
- 22:51So this is the second term that we use
- 22:54and of course we use the same you know
- 22:56decay strategy. We use the relative high
- 22:58weight at the beginning of training and
- 23:00we decay to a relatively low weight at
- 23:03the end of training.
- 23:04And of course the third term is the
- 23:06standard term. It's it's the outcome
- 23:08reward. It's going to measure whether
- 23:11the entire task is completed
- 23:14and then we're going to add these three
- 23:16terms in our
- 23:17reinforcement learning
- 23:19system. And of course we have to build
- 23:21you know the entire infrastructure
- 23:23because
- 23:24right now you need to support the
- 23:26parallel execution and then you'll need
- 23:28to support different reward functions
- 23:31and and to you know maximize the
- 23:33efficiency of the entire agent swarm RL
- 23:36system.
- 23:38So here are three
- 23:39uh
- 23:40different things that that we have
- 23:42tried scaling. The Muon clip optimizer
- 23:46improves token efficiency
- 23:49and Kimi Delta attention in the Kimi
- 23:52linear architecture improves long
- 23:54context and we also have the agent
- 23:56swarms paradigm to further
- 24:00create a new dimension of scaling.
- 24:02And all of this put together we created
- 24:06Kimi K2.5 a new model that we just
- 24:09released
- 24:10over 1 month ago.
- 24:12Here's a short video to demonstrate some
- 24:14of its capabilities.
- 24:25>> [music]
- 24:35[music]
- 24:39[music]
- 24:46[music]
- 24:52[music]
- 24:57[music]
- 25:16[applause]
- 25:19>> So
- 25:20yeah there are a lot of interesting
- 25:22things capabilities that we discover
- 25:24from the model. For example
- 25:26it merges the visual capabilities with
- 25:29coding capabilities. So a lot of new
- 25:31things just emerge out of it. It can
- 25:34read a video and then produce a website
- 25:37that sort of replicates or style
- 25:40transfer the original video.
- 25:43And all of this are due to successful
- 25:47and stable training
- 25:49at the pre-training stage. So this is
- 25:50also one of the
- 25:52most beautiful curves that I observed in
- 25:55my life. So this is the training curve
- 25:57of the K2.5 base model. So as you can
- 26:02see it went through over 15 trillion
- 26:04tokens and of course in K2.5 we
- 26:06additionally trained another 15 trillion
- 26:08tokens and the entire
- 26:11the training process is just so stable.
- 26:14There's no loss spike especially when we
- 26:17introduce this new Muon optimizer we
- 26:19didn't observe any spike and this smooth
- 26:22stable training process produces a very
- 26:25stable outcome
- 26:27a very strong base model that we can
- 26:28fine tune on top of it to achieve you
- 26:32know new capabilities such as we
- 26:34introduced and shown in the video the
- 26:36video. And this is also of course
- 26:38uh
- 26:39trained on Nvidia H100 GPUs and each
- 26:43node in this H100 cluster contains two
- 26:47TB RAM and eight GPUs
- 26:49connected by NVLink.
- 26:52And uh one of the another you know key
- 26:55innovation of Kimi K2.5 is that it is
- 26:58the first open model with native joint
- 27:02vision text capabilities. So if you look
- 27:04at previous open models usually their
- 27:07visual capabilities are added on top of
- 27:10a text base meaning that for example if
- 27:12you train the text models for 20
- 27:15trillion tokens and then on top of it
- 27:17you do another two trillion sort of a
- 27:20post training process to add additional
- 27:22visual capabilities on top of it.
- 27:25But for K2.5 it's different in the sense
- 27:28that we fuse the training process of
- 27:31vision and text from day one. So it's
- 27:34called early fusion here. We start from
- 27:37you know 0% of the progress. So from day
- 27:40one we're going to merge the vision and
- 27:42text tokens and as shown in our
- 27:44preliminary experiments it outperforms
- 27:48late fusion and some of the new
- 27:50capabilities that we observe also come
- 27:52from this training recipe. For example
- 27:56if you want to do vision to code you
- 27:58really have to merge vision and text
- 28:01into a single brand to achieve that. If
- 28:04you separate these two brands it's not
- 28:06going to happen. You have to align these
- 28:08two modalities into a share embedding
- 28:10space
- 28:12in
- 28:13a share representation space
- 28:15so as to achieve this.
- 28:17And another interesting thing that we
- 28:19observe is that
- 28:21these two modalities can actually
- 28:23enhance each other. So that's
- 28:26that's been long been a challenge that
- 28:29if you add vision capabilities into a
- 28:31text model it's going to somewhat
- 28:34hurt the text performance. But here we
- 28:37found that if you train it properly
- 28:39these two modalities can actually
- 28:41enhance each other. So this is one of
- 28:43the key findings that
- 28:45we observe in in in our training. So
- 28:48first vision improves text. So this is
- 28:50so interesting. So before vision RL
- 28:54the performance in the first column and
- 28:56then we have the performance after
- 28:58vision RL. So here vision RL
- 29:01refers to a process that we only use
- 29:04vision task. So there is no text task
- 29:07involved here. We only have vision task.
- 29:09For example we teach the model how to
- 29:11how to count how to answer some of this
- 29:14visual QA
- 29:16problems without any for example math
- 29:19any coding problems in in in this space.
- 29:22But we observe that it's going to
- 29:23improve the performance for even you
- 29:25know reasonably heavy text task.
- 29:28And on the other hand text also improves
- 29:30vision. If you have a very strong text
- 29:32base
- 29:33you you actually don't need any vision
- 29:36SFT data in the training process and
- 29:38this is the approach that we adopt. So
- 29:41it's called zero vision SFT. Basically
- 29:44we don't have We have basically zero
- 29:46vision SFT data and the only SFT data
- 29:49that we have is the text SFT data and
- 29:52then we do a joint IO over text and
- 29:54vision and you can see that we can
- 29:56achieve almost state-of-the-art
- 29:58performance across the board on on
- 30:00vision task without any vision data. So,
- 30:03it is clear that
- 30:05if you have a strong text base is also
- 30:06going to improve the vision if if you
- 30:10align these two modalities into a shared
- 30:13space in your in your pre-training.
- 30:17And also
- 30:19these are some of the examples of uh
- 30:23uh
- 30:25Yeah, as I was showing the video so it's
- 30:28it demonstrates strong capabilities of
- 30:31visual design and front-end coding and
- 30:33this also emerges from our vision text
- 30:36pre-training.
- 30:38So,
- 30:39uh after all this so this are all about
- 30:42Kim E 8.5 and as probably
- 30:46you already know we released our new
- 30:48architecture yesterday
- 30:51in our tech report is called attention
- 30:53residue. So, here I'm also going to
- 30:56briefly talk about our new work which
- 30:59serves as a a sneak peek into our next
- 31:01generation architecture that we're
- 31:03probably going to adopt in in our later
- 31:05models.
- 31:07So, here the motivation is is quite
- 31:10simple. Can we apply some of our
- 31:12techniques that we use in in the in the
- 31:15temporal dimension and we we just take
- 31:19some of the inspirations and we apply to
- 31:21the depth
- 31:22dimension. So,
- 31:25and it starts from this residual
- 31:27connection. So, I still remember
- 31:30listening to to Kaiming's talk at the
- 31:34tutorial in ICML 2016 10 years ago.
- 31:38So, it was a brilliant idea. So,
- 31:40basically before ResNet nobody was able
- 31:44to train deep networks.
- 31:46If you increase the depth if you
- 31:48increase the number of layers for neural
- 31:50networks nobody was able to train it
- 31:53because you will observe this gradient
- 31:55explosion gradient vanishing all these
- 31:57you know stability issues. But then
- 31:59after the introduction of of ResNet we
- 32:02can train you know an arbitrary large
- 32:05number of layers you can stack as many
- 32:07layers as as you want and you you you
- 32:09don't have to worry about the training
- 32:11stability issue
- 32:13and stuff. And as discussed in Ilya's
- 32:16talk uh 2 years ago it basically says
- 32:19that residual connection is a variant of
- 32:23LSTM but just rotated 90 degrees. So,
- 32:27how do you understand this? If you look
- 32:28at LSTM is is a variant of recurrent
- 32:31net, right? And it's a recurrent model
- 32:35process. So, we're going to take the
- 32:37hidden states from the last step and
- 32:41then we're going to have some gating
- 32:43mechanism some function to produce the
- 32:45current state. Right? And if you look at
- 32:48the
- 32:49the depth dimension the residual
- 32:51connection is basically the same. We're
- 32:53going to take the output from the last
- 32:55layer and then we're going to apply some
- 32:57sort of function on top of it to produce
- 33:00the current
- 33:03the current output of the current layer.
- 33:06It's just the formulation is different.
- 33:07For example, for residual connection
- 33:09we're going to use a fixed addition
- 33:11we're going to have this
- 33:14we're going to add the previous hidden
- 33:16state
- 33:17with the current output. It's just the
- 33:20formulation that's different but the
- 33:21basic idea is the same it's a recurrent
- 33:23net applied
- 33:25in the dimension of depth.
- 33:27And but on the other hand
- 33:29we can think about
- 33:31reformulating
- 33:33this
- 33:34this function. Instead of having LSTM
- 33:37can we have an attention in the
- 33:39dimension of depth and it's going to
- 33:41create new possibilities because
- 33:43attention have has been demonstrated to
- 33:46be so successful
- 33:48in the transformer era. So, what we're
- 33:50going to do is not just to take the last
- 33:53hidden state but we're going to consider
- 33:55all the previous hidden states and use
- 33:57the attention operation the attention
- 33:59mechanism to assemble and aggregate
- 34:02all of these previous hidden states to
- 34:04compute the current state. So, this is
- 34:07exactly attention rotated by 90 degrees.
- 34:12It's sort of we view as a natural
- 34:14generalization of residual connections
- 34:17in the LSTM
- 34:19analogy.
- 34:21Okay, and here's the detail formulation.
- 34:23So, on on the left-hand side is a
- 34:25standard residue. As I said it's
- 34:27basically LSTM rotated by 90 degrees and
- 34:30the second figure is attention rotated
- 34:33by 90 degrees. So, what we do is to
- 34:36collect all the previous hidden states
- 34:38and have a simple attention operation on
- 34:41top of it to produce the current layers
- 34:43outcome. And of course to increase the
- 34:47efficiency to reduce the infrastructure
- 34:50for example communication and memory
- 34:51overhead we also design a new variant
- 34:54called
- 34:55block attention residue on the
- 34:57right-hand side. So, basically the idea
- 34:59is is also simple. We're going to divide
- 35:04all the layers in the neural network
- 35:06into multiple blocks. For example, each
- 35:08block can contain say 16 layers or it
- 35:11can contain maybe four layers and then
- 35:13for each block we're going to
- 35:16apply
- 35:17this attention residue only on the
- 35:19output of each block but within each
- 35:21block we also we still adopt this
- 35:24standard residue. So, this is going to
- 35:26reduce a lot of overhead while having
- 35:28minimal loss in terms of training
- 35:30accuracy.
- 35:32And these are some of the impressive
- 35:34results that we achieved
- 35:37on this new architecture. So, on the
- 35:39scaling law we can improve the token
- 35:43efficiency by 24%
- 35:46meaning that if you have 50 trillion
- 35:49high-quality tokens now you just
- 35:51magically have
- 35:53over 60 trillion tokens and then for the
- 35:58validation loss you can also observe
- 36:00that
- 36:01it's consistently lower than
- 36:04the original curve
- 36:07demonstrating the stability across
- 36:09optimization and also achieved the best
- 36:12improvement on some of this coding math
- 36:16and reasoning heavy task as shown in the
- 36:18benchmark results of GPQA math and human
- 36:22eval.
- 36:24So, the entire community keeps moving
- 36:27forward
- 36:28and we're happy that we can we're able
- 36:30to contribute to to the community with
- 36:33you know new technologies and some of
- 36:35this
- 36:37have you know some of this technologies
- 36:39have been sort of standard and de facto
- 36:41for a long time but as you can see we
- 36:44still see a lot of opportunities to
- 36:46improve it to
- 36:48to have revolutionary new design to
- 36:51achieve better performance. If you
- 36:53multiply all these gains together you
- 36:56can actually have a much better model.
- 36:59So, Adam was invented in 2014
- 37:02and now we scale an open source Neon
- 37:04clip a a dropping replacement for
- 37:08for Adam and I'm sure that if you're
- 37:11training transformer LLM it's going to
- 37:14be much better if you use Neon clip
- 37:16instead of Adam. And attention was
- 37:18invented over 8 years ago and then now
- 37:21we have Kim E linear which is a linear
- 37:24version. We don't have to use full
- 37:26attention across all layers. We can have
- 37:29linear attention that performs better on
- 37:31short long context at the same time.
- 37:34And also residual connections are now uh
- 37:38also challenged.
- 37:40We scale an open source attention
- 37:42residue. So, I think one of the
- 37:46interesting things about our era is that
- 37:49we sort of adopt a different mindset for
- 37:52for doing research. So, if we go back to
- 37:5510 years ago it's mostly about
- 37:57publishing a new idea
- 37:59but then I think the lack of the rigor
- 38:02of the experiments is very hard to
- 38:05produce
- 38:06solid experimental results. But now we
- 38:09have the scaling ladder. We have enough
- 38:11resources to you know train the model
- 38:14and running on different at different we
- 38:16can have you know a whole set of
- 38:18benchmarks to measure the progress. So,
- 38:21it is it becomes easier to make
- 38:24confident and solid conclusion out of
- 38:26it. And this is one of the reason why we
- 38:29are observing you know new progress
- 38:32on this
- 38:33ancient techniques and I'm sure that
- 38:35we'll see more and more especially in
- 38:38the open source community. I think we're
- 38:40going to have more and more even better
- 38:43architectural and you know optimization
- 38:46improvement in in the next few years.
- 38:49All right, so to summarize we're going
- 38:51to keep scaling our models and so this
- 38:55are three dimensions.
- 38:57For example, we we see
- 39:00we we see different architectures and
- 39:03optimizers that optimize all three you
- 39:06know dimensions and we'll keep you know
- 39:08see new dimensions for scaling.
- 39:11Agents forms is is not the end and we
- 39:14are glad that we can
- 39:16move forward with the entire open source
- 39:18community to achieve better and better
- 39:21you know intelligence. Thank you so
- 39:23much.
- 39:25>> [applause]
About this transcript
This page contains the full transcript of How We Scaled Kimi K2.5 | Zhilin Yang's full GTC 2026 Keynote by Kimi AI, generated from the public captions YouTube serves with the video. The transcript has 5,592 words across 927 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.