Deepseek just did the impossible — Transcript
Full transcript
- 0:00Deepseek did it again. They've released
- 0:02a new model and not only is this among
- 0:05the frontier models out there, but it's
- 0:07also the most optimized, efficient, and
- 0:10frictionless AI model we've seen so far.
- 0:13As always, not only have they open-
- 0:15sourced the model, but they've also
- 0:16released a technical paper on this. And
- 0:18how they designed this is just really
- 0:20unexpected and sometimes even seems
- 0:23wrong. But once you understand the logic
- 0:25behind everything, then it suddenly
- 0:27becomes absolutely brilliant. In this
- 0:29video, we're going to do a deep dive
- 0:31into its design so that you can see how
- 0:33genius and cracked this is. Now, this is
- 0:36a super technical paper, but as always,
- 0:38I'm going to break it down into simple
- 0:40terms so that anyone can understand.
- 0:42Let's jump right in. Let's first set the
- 0:44stage by reviewing the situation that
- 0:46Deep Seek is in. Keep in mind that this
- 0:48is just a small Chinese lab that does
- 0:51not have nearly as much funding as
- 0:53OpenAI. Their team is like dozens of
- 0:55times smaller. Plus, they don't even
- 0:57have access to the best Nvidia GPUs out
- 0:59there, nor do they have a massive data
- 1:01center. In fact, they are severely
- 1:04constrained in terms of compute and
- 1:06resources. Yet, they just released their
- 1:08latest model, Deepseek V4.1 Flash, which
- 1:12even matches the performance of Frontier
- 1:14models, even though this is just a Flash
- 1:16model. Plus, it's way faster and more
- 1:18efficient. And get this, its memory
- 1:21footprint is like over 400 times smaller
- 1:24compared to the first generation. How on
- 1:26earth did they pull this off? Well, to
- 1:28understand the brilliance of their
- 1:30solution, we first need to go over the
- 1:32basics of what actually happens when you
- 1:34use an AI model. It's actually broken
- 1:36down into two different phases. The
- 1:38first phase is called prefill. This is
- 1:41basically when the AI reads your prompt
- 1:43along with any other information or
- 1:45documents you give it so it can
- 1:47understand the context of everything
- 1:48before it starts generating an answer.
- 1:51Specifically, your text is broken down
- 1:53into smaller pieces called tokens which
- 1:55are then passed through the AI model's
- 1:57layers. In fact, each layer in the model
- 2:00calculates two important sets of numbers
- 2:02called keys and values for each token.
- 2:05These are stored in something called the
- 2:07KV cache. You can think of this KV cache
- 2:10as like the notes the model takes as it
- 2:12reads through your information. Now the
- 2:14second phase is the decode phase or the
- 2:17writing phase. This is where the AI
- 2:19starts generating your answer. Now
- 2:20interestingly large language models
- 2:22generate its answer one word or token at
- 2:25a time. In order to do so, the model
- 2:27needs to look at everything that came
- 2:29before it to figure out the next most
- 2:30probable word that should come next. And
- 2:33to do that efficiently, it needs to
- 2:34refer back to its notes that it took
- 2:37during the reading phase. In other
- 2:38words, it needs to look at all the KV
- 2:40cache that it has created. This KV cache
- 2:43lets the model quickly access the
- 2:44relevant information without having to
- 2:46recalculate the entire conversation from
- 2:49scratch. To understand this a bit
- 2:50better, here's a nice analogy. Think of
- 2:52an AI model like a student watching a
- 2:54really long lecture and then taking
- 2:56notes along the way. Later, when the
- 2:58student needs to answer a question about
- 3:00the lecture, he doesn't need to replay
- 3:01the entire lecture from the beginning.
- 3:04Instead, he can just refer back to his
- 3:05notes. Well, you can think of the KV
- 3:07cache like these notes. It gives the AI
- 3:10a quick way to refer back to information
- 3:12it has already processed without having
- 3:14to recalculate everything from scratch.
- 3:17Now, here's the problem the industry is
- 3:19facing right now. You see, for a short
- 3:21prompt like this, everything works fine.
- 3:23The AI model can easily convert this
- 3:25into a fairly small KV cache that fits
- 3:28well within its active memory. But
- 3:30here's the thing, the industry is now
- 3:32focused on making AI handle really long
- 3:34and complex tasks. We want to get AI
- 3:37agents to work autonomously for hours or
- 3:40even days. We want to give them a ton of
- 3:43different documents or a huge codebase
- 3:45to keep track of and keep working on for
- 3:47a really long time. And in these
- 3:49scenarios, the KV cache or the notes
- 3:51that the AI has to take is going to be
- 3:53massive. So going back to our analogy
- 3:56now, instead of just watching one
- 3:58lecture, the student has to watch weeks
- 4:00and weeks of online lectures and take
- 4:02notes on all of them. His notes will
- 4:04start to pile up fast to the point where
- 4:06they can't even fit on his desk anymore.
- 4:08They're going to fill up his entire room
- 4:10and eventually he needs to like shove
- 4:12his notes into filing cabinets in
- 4:14another room down the hallway. And then
- 4:16every time the student needs to answer a
- 4:18question, he needs to dig through all
- 4:20these huge piles of notes. Well, this
- 4:22analogy is exactly what happens to an AI
- 4:25when it needs to work with a ton of
- 4:26information. You see, inside a GPU, you
- 4:29have this high bandwidth memory or HBM.
- 4:32This is the incredibly fast memory
- 4:34that's right next to the processor chip.
- 4:36It's really fast and easy to access. So,
- 4:38it's kind of like the surface of a
- 4:40student's desk in our analogy. any notes
- 4:42that are sitting on this desk are
- 4:44incredibly quick and easy to access. But
- 4:46the desk space is limited. Well, high
- 4:49bandwidth memory is exactly the same.
- 4:51There's limited space and it's also
- 4:52extremely expensive. So when the KV
- 4:54cache gets too big, this high bandwidth
- 4:57memory gets full. So the computer has to
- 4:59start storing this data somewhere else.
- 5:01For example, on solidstate drives or
- 5:03SSDs which are outside the GPU. So, in
- 5:07our analogy, it's like putting some of
- 5:09the students notes into filing cabinets
- 5:11in another room down the hallway. You
- 5:13have much more storage space there, but
- 5:15there's a trade-off. He needs to go to
- 5:17the other room to grab the notes from a
- 5:19cabinet, which takes much longer than if
- 5:21it was just sitting in front of him on
- 5:23his desk. So, similarly, SSDs, or these
- 5:25solidstate drives outside the GPU, have
- 5:28way more storage capacity, but they're
- 5:31located further away from the GPU, and
- 5:33the latency there is devastating. If the
- 5:35KV cache needs to be stored there, well,
- 5:37every time the AI needs to predict the
- 5:39next word, it has to fetch data from the
- 5:42SSD, pull it through the motherboard up
- 5:44into its high bandwidth memory, and then
- 5:46into the GPU processor. This data
- 5:49transfer speed becomes the absolute
- 5:51bottleneck. The processor is just
- 5:53essentially sitting idle waiting for the
- 5:55massive KV cache notes to travel through
- 5:57the wires. So, to sum things up, there
- 5:59are two major problems that current AI
- 6:02systems are facing. One problem is
- 6:04compute. The student is basically just
- 6:06drowning in his own notes and it's
- 6:08really hard for him to find things. The
- 6:10second problem is speed. Because there
- 6:12are so many notes, some of these notes
- 6:14are stored in filing cabinets in another
- 6:16room and the student needs to run back
- 6:18and forth to fetch these notes and put
- 6:20them on his desk which takes a ton of
- 6:22time. So these are the limitations that
- 6:24DeepS is facing. How on earth did they
- 6:27tackle this? First of all, let's go over
- 6:29the architecture of a normal large
- 6:31language model. They use the transformer
- 6:33architecture which is basically made up
- 6:35of many layers. In fact, if you're
- 6:38curious about how transformers actually
- 6:40work under the hood, definitely see this
- 6:42video where I do a full explainer. But
- 6:45anyway, what happens is your prompt is
- 6:47broken down into data that flows through
- 6:50these layers of the transformer which
- 6:52ultimately outputs the next most
- 6:53probable word for its answer. And then
- 6:55that word is appended back and then it's
- 6:58run through the model again to predict
- 7:00the next most probable word. and then
- 7:01this loops again until it generates your
- 7:03final answer. Now, each layer that it
- 7:05goes through generates its own notes or
- 7:08KV cache. And all of this must be stored
- 7:10somewhere. This often takes up a lot of
- 7:12space. So, not only does it fill up the
- 7:14high bandwidth memory, but it also
- 7:15spills over to SSDs that are outside the
- 7:18GPU. Well, the architecture from this
- 7:20new deepse has many layers. But here's
- 7:23the really unusual part. They split this
- 7:25into two halves. There's a causal
- 7:27encoder component and then there's a
- 7:29decoder component. This already looks
- 7:31completely different from a standard
- 7:33transformer model. And here's how it
- 7:35works. During the prefill phase, again,
- 7:37this is where the AI is reading your
- 7:39prompt and all the information you give
- 7:41it. The last half is essentially turned
- 7:43off. They just do nothing. And then the
- 7:45first half of the layers do all the
- 7:47heavy lifting. They read everything.
- 7:49They build the contextual understanding
- 7:51and they generate what's called the
- 7:52global KV cache. After reading
- 7:55everything, then the last half basically
- 7:57generates the answer. it does the
- 7:59writing or the decoding phase. The thing
- 8:02is it still needs to know the context in
- 8:04order to do the writing right to predict
- 8:06the next word. So you might be wondering
- 8:07if it doesn't generate the KV cache
- 8:09itself, how does it understand the
- 8:11context? Here's why this design is so
- 8:13genius. Instead of having to read
- 8:15everything itself, it just looks at the
- 8:17output from the last layer of this
- 8:19encoder component. In other words, it
- 8:22borrows the completed global KV cache
- 8:24directly from the end of this encoder
- 8:27block. Now, this is quite a shocking and
- 8:29unexpected design because if they did
- 8:31this, if half its brain basically
- 8:33skipped the reading part, doesn't it
- 8:35lose some understanding of the context,
- 8:37which would make its answer worse? Well,
- 8:39that's the exact risk of this
- 8:41architecture. And that's why it took
- 8:43incredible engineering to balance.
- 8:45What's fascinating here is that this
- 8:47last half, this decoder component,
- 8:49doesn't really entirely skip reading. It
- 8:51just skips calculating the global
- 8:53context. In other words, the global KV
- 8:55cache, but it still computes the local
- 8:57context for itself. Now, you might be
- 8:59wondering, what's the difference between
- 9:01global and local context here? Well,
- 9:04global context is basically all the
- 9:06information that was given to it plus
- 9:08your prompt and everything else
- 9:09attached. It's basically all the notes
- 9:11that the student took after watching
- 9:13weeks and weeks of lectures. In
- 9:14contrast, the local context is just the
- 9:17immediate context of the sentence the AI
- 9:19is currently writing. In other words,
- 9:21what's directly relevant to the next
- 9:23word that it needs to output. You'll see
- 9:26this decoder component doesn't calculate
- 9:28any global KV cache, but instead it uses
- 9:31something called a sliding window
- 9:33attention to pay extremely close
- 9:35attention to the most recent tokens
- 9:37only, but not everything before it. And
- 9:39it turns out that this design works
- 9:41pretty well. Here's a nice analogy to
- 9:43wrap your head around this. Imagine a
- 9:45company where they need to analyze
- 9:47thousands of pages of financial reports.
- 9:50Well, they would first get junior
- 9:52analysts to read every single page and
- 9:54crunch out the numbers, do the data
- 9:55analysis, and then write a very dense
- 9:57and accurate executive summary. Well,
- 9:59these junior analysts are basically like
- 10:02the first half of the model. And then
- 10:04this executive summary is like the
- 10:06output at the end. They then hand the
- 10:07summary to the senior executives, which
- 10:10are like the decoder layers. These
- 10:12senior executives absolutely do not read
- 10:15the original thousands of pages.
- 10:17Instead, they just rely on the executive
- 10:20summary provided by the juniors. But
- 10:21when it comes time to sign off on
- 10:23something, then these senior executives
- 10:25put on their reading glasses and
- 10:27scrutinize the exact wording of the page
- 10:29sitting right in front of them. That's
- 10:31basically the local context. The senior
- 10:34executives rely on this global summary
- 10:36for direction, but they mostly focus
- 10:39locally on the page in front of them for
- 10:42execution. And by structuring the model
- 10:43this way, DeepC completely bypasses the
- 10:46need for basically half of the model to
- 10:48generate its own massive KV cache notes.
- 10:51And this is a huge deal. It essentially
- 10:53slashes the compute required to read
- 10:56things by half. Now, reducing the
- 10:58compute is great, but we still have the
- 11:00problem of memory, right? It's still
- 11:02generating these massive KV cache notes
- 11:04whether it's global or local and these
- 11:06are like overflowing on the student's
- 11:09desk and he's forced to like store these
- 11:11excess notes in filing cabinets in
- 11:13another room. How on earth can we reduce
- 11:15these massive piles of notes? And this
- 11:17brings us to one of the most fascinating
- 11:19parts of DeepS's new design. And this
- 11:22part is just brilliant. In fact, let me
- 11:24show you the results first so you can
- 11:26see how insane this is. If you do any
- 11:28kind of content creation, definitely
- 11:30check out Luma, the sponsor of this
- 11:32video. Think of it as a creative AI
- 11:35agent that works alongside you through
- 11:37your entire creative process. Instead of
- 11:39just giving you the results of a single
- 11:41prompt, I can access the best image and
- 11:44video models out there. And the nice
- 11:46thing is instead of manually jumping
- 11:48between these tools, I can just get Luma
- 11:50agents to autonomously do entire
- 11:52workflows for me. It can develop the
- 11:54concept, generate the visuals and shape
- 11:56the project all within the same
- 11:58workspace. For example, I can get it to
- 12:00generate a brand kit for me, design
- 12:02different products, and generate other
- 12:04marketing assets all inside the same
- 12:06project. And if I need to edit
- 12:08something, I can just prompt the agent
- 12:09to refine the results iteratively. One
- 12:11of the most powerful features is Luma
- 12:13skills. You can basically create
- 12:15reusable skills for workflows you use
- 12:18all the time. Basically, you give Luma a
- 12:20set of instructions once and then you
- 12:22can run that same workflow on different
- 12:24assets whenever you want. For example, I
- 12:26can create a skill where I can input any
- 12:28product photo and it'll output some UGC
- 12:31videos of an influencer talking about
- 12:33the product. Or here's another example
- 12:35of a skill where I can upload a product
- 12:37photo and it'll generate a 360° orbit
- 12:40view like this. Luma basically gives you
- 12:42an intelligent creative co-pilot that
- 12:45can autonomously carry out your
- 12:46workflows. Whether you're creating
- 12:48marketing campaigns, branded content,
- 12:50product visuals, or social media
- 12:52content, Luma is one of the best
- 12:54platforms you can use. Try Luma today
- 12:56using the link in the description below
- 12:58or by scanning the QR code here. If you
- 13:00compare the global KV cache size per
- 13:03token, which is basically the size of
- 13:05the notes the student has to take,
- 13:07DeepSeek V1 is almost 390,000
- 13:10bytes. Now, if you fast forward just a
- 13:12few generations to this latest V4.1
- 13:15Flash, it's only 890 bytes per token.
- 13:18They basically shrunk the size of the
- 13:20notes down by like 437 times, which is
- 13:24crazy. Even if you compare this to the
- 13:26previous DeepSseek V4 Flash, this one
- 13:29still required like 3,500 bytes per
- 13:31token. So, this new update is like
- 13:33almost four times smaller than the
- 13:35previous generation. But here's the
- 13:37challenge to all of this. How can you
- 13:39compress these notes so much without
- 13:41losing the meaning? How can you still
- 13:42maintain the AI model's understanding of
- 13:44everything? Well, Deepseek used a
- 13:46mechanism called compressed sparse
- 13:49attention 2 or CSA2. To understand this,
- 13:52let's first review how a normal
- 13:53transformer model works. Each layer in
- 13:56the model has to calculate its own KV
- 13:58cache. In other words, it has to make
- 13:59its own nodes. Conceptually, you can
- 14:01think of each layer as focusing on
- 14:03different things. For example, some
- 14:05layers might focus on certain patterns,
- 14:07while others focus on things like
- 14:09grammar or relationships between ideas
- 14:11or the broader meaning of the text. Now,
- 14:14with all these layers each making their
- 14:16own notes, you can see how the total
- 14:18size of the KV cache could become really
- 14:20hard to manage. And if each layer has to
- 14:22calculate its own notes, you can see how
- 14:24things could become redundant. Well,
- 14:26this new CSA2 mechanism by DeepSeek
- 14:30completely shatters this redundancy. It
- 14:32introduces the concept of extreme
- 14:34sharing. So instead of creating new
- 14:36notes from scratch every time, each
- 14:38layer could use three different
- 14:40operating modes. Full mode, reindex
- 14:42mode, and reuse. Let's go over each one.
- 14:45So if the layer is in full mode, it has
- 14:47to do all the hard work. It has to
- 14:49create brand new notes from scratch. In
- 14:51other words, it needs to make the full
- 14:53KV cache. But here's the important part.
- 14:55It also creates an index for future
- 14:58layers to search these notes. Think of
- 15:00the index like a guide or map or like a
- 15:03table of contents. Other layers can just
- 15:05look at this table of contents to figure
- 15:07out where exactly to search in the notes
- 15:09instead of trying to read the whole
- 15:11thing from start to finish. So this
- 15:12makes it way faster to search for
- 15:14information. Another mode is called
- 15:16reindex. And here's where the efficiency
- 15:19kicks in. If the layer is in reindex
- 15:21mode, it just reuses the notes or in
- 15:23other words the KV cache from a layer in
- 15:25full mode. It doesn't create its own
- 15:27notes from scratch. But what it does do
- 15:29is create a new index from scratch.
- 15:31Again, think of this as like creating a
- 15:33guide or a new table of contents that
- 15:36searches the same notes as before, but
- 15:38highlights completely different parts.
- 15:40For example, let's say you're giving the
- 15:41AI a ton of information about the
- 15:43history of the world. The full layer
- 15:45could make an index about things in
- 15:47chronological order, which might look
- 15:49like this. A reindexed layer would take
- 15:51the exact same notes, but give it a
- 15:53completely different table of contents.
- 15:55For example, instead of chronological
- 15:57events, it could be themes across time
- 15:59or it could be different technological
- 16:01breakthroughs. It's basically like
- 16:03mapping different paths through the same
- 16:05notes. So that's the reindex mode. And
- 16:07then finally, we have the third mode,
- 16:09which is maximum efficiency. And this is
- 16:12the reuse mode. If the layer has this
- 16:14mode, it exerts almost zero memory
- 16:17effort. It just reuses the notes from
- 16:19the full layer as well as the indices
- 16:21from either the full layer or the
- 16:23reindexed layers. It doesn't write any
- 16:25new notes, nor does it create any new
- 16:27table of contents. It just takes
- 16:29information that's already available to
- 16:31it. And with this design, with these
- 16:34different modes, each layer doesn't have
- 16:36to store as much notes. The total size
- 16:38of these notes is reduced significantly
- 16:41because some of these layers don't even
- 16:42need to create new notes at all. They're
- 16:44just reusing notes and indices from
- 16:46previous layers. But wait, this ain't
- 16:49all. Deepseek takes this one step
- 16:51further. They've added something called
- 16:53a hierarchical sparse indexer. And
- 16:55here's how it works. You see, in the
- 16:57last half of the model, this is the
- 16:59decoder part. The very first layer acts
- 17:02kind of like a gatekeeper. It scans all
- 17:04the global notes from the previous
- 17:06layer. Remember, this is also called the
- 17:08global KV cache and it generates a
- 17:10candidate pool. Think of this like a
- 17:12short list. Basically, from those piles
- 17:15and piles of notes, it figures out just
- 17:17the most relevant concepts.
- 17:19Specifically, out of a million tokens,
- 17:21it only selects around 16,000 that are
- 17:24the most relevant. Everything beyond it
- 17:26is basically ignored and then all the
- 17:28subsequent layers are forbidden to
- 17:30search for anything else in the notes.
- 17:32So, this drastically narrows the
- 17:34universe of possible answers right at
- 17:36the start of the writing phase. Now,
- 17:38obviously, as you may expect, if we
- 17:40restrict the AI's search space like
- 17:42this, it might hallucinate or miss
- 17:44important details, right? its response
- 17:46will become dumber if it doesn't look at
- 17:48everything. But here's the genius behind
- 17:50this. Deepseek was able to make it work
- 17:52by training the model to build these
- 17:54candidates with such high accuracy that
- 17:56the later layers don't even notice the
- 17:58rest of the info is missing. The
- 18:00elegance of their engineering is just
- 18:02profound. Let's take a moment to
- 18:04appreciate what they did here. They
- 18:06don't have the best Nvidia GPUs. They
- 18:08don't have the biggest data center in
- 18:09the world. Heck, they're pretty starved
- 18:11for compute. So instead of focusing on
- 18:13the hardware side, they completely
- 18:15optimized the software part so the
- 18:17hardware doesn't have to work so hard.
- 18:19And we ain't done yet. You see, all this
- 18:22intense optimization also led to some
- 18:24additional issues they had to fix.
- 18:26Remember this sliding window attention
- 18:28mechanism we talked about earlier. This
- 18:30is where the senior executive doesn't
- 18:32read everything, but when he signs off
- 18:34on something, he has to look really
- 18:35closely at all the information on the
- 18:37page in front of him. Well, this sliding
- 18:39window attention is the AI's hyper local
- 18:42short-term memory. And according to the
- 18:44paper, this hyper local memory was
- 18:46actually causing a huge storage problem.
- 18:49You see, when a user has a conversation
- 18:51with the AI over multiple turns. In
- 18:54other words, when you say something, the
- 18:55AI replies and then you reply back and
- 18:57so on. The AI has to save this local
- 19:00context of every single turn into its
- 19:03SSD so it won't forget the flow of the
- 19:05conversation. And this constant saving
- 19:07of short-term local memory was
- 19:09completely clogging the hard drives.
- 19:12Going back to our analogy, this is like
- 19:14filling up all the cabinets in the other
- 19:15room down the hall. Now, when you cache
- 19:18data, it means you save it in a
- 19:20temporary location so you can retrieve
- 19:22it quickly later, right? But if it gets
- 19:24too big and these notes are located in
- 19:26cabinets in another room, well, fetching
- 19:28this information gets super slow. the
- 19:31student needs to run down the hallway to
- 19:33the other room to grab the notes and
- 19:35then place them back on his desk. So,
- 19:37DeepSeek recognized this issue and their
- 19:40solution was something called SWA
- 19:42bounded replay. And this is probably the
- 19:44most unexpected and shocking part of
- 19:46their new design. Their solution to this
- 19:49local memory clogging up the hard drives
- 19:51is to simply delete it. They literally
- 19:53deleted the AI's short-term memory
- 19:55completely. So, when the AI finishes
- 19:57generating its reply to you, its local
- 19:59memory just evaporates into thin air. As
- 20:02you can imagine, if we delete its
- 20:04short-term memory, shouldn't it become
- 20:06like completely disoriented? Wouldn't it
- 20:08lose track of what's going on? Well,
- 20:10that's what we would expect. So, what
- 20:12Deep Seek did was they basically got the
- 20:14AI to recalculate the last parts of the
- 20:17conversation instantly, specifically the
- 20:19last 128 tokens of the conversation. In
- 20:21other words, they forced it to generate
- 20:23the most immediate new notes from
- 20:25scratch right on the spot every single
- 20:27time. This sounds so counterintuitive,
- 20:29right? I just spent the past few minutes
- 20:31in this video explaining how they tried
- 20:33to reduce the compute and split the
- 20:35brain into halves to avoid generating
- 20:37and reading so many notes. But now, if
- 20:39we get this AI to recalculate this last
- 20:42part of the conversation every time,
- 20:44wouldn't this be incredibly inefficient?
- 20:46Doesn't this slow everything down? Well,
- 20:48here's where DeepSseek gives us a
- 20:51masterclass in efficiency. Let's walk
- 20:53through the trade-off here. You kind of
- 20:55have two possible ways to handle this
- 20:57short-term memory. The first way is the
- 20:59normal way where we take data from the
- 21:01GPU, we push it through the motherboard
- 21:03and write it into a hard drive. And
- 21:06later, when it needs to use this shorter
- 21:08memory, it needs to search for it, pull
- 21:10it back up through the motherboard and
- 21:11into the GPU processor. This is super
- 21:14slow because you're physically moving
- 21:15data across this distance. Now, the
- 21:18second way is to just use the raw power
- 21:20of a modern GPU to simply recalculate
- 21:23its short-term memory. In other words,
- 21:24just rewrite the most recent and
- 21:26relevant notes from scratch. And for a
- 21:28modern GPU, just crunching like 128
- 21:31tokens is just a microscond operation
- 21:33that barely takes up any time or power.
- 21:36So, doing the math is actually much
- 21:38faster than trying to transfer this data
- 21:40back and forth from the SSD. The Deep
- 21:43Seek team realized that saving this
- 21:45short-term memory to the hard drive was
- 21:47just a massive waste of a really slow
- 21:49resource. By completely removing the
- 21:51step and just getting the GPU to quickly
- 21:53recalculate everything on the fly, not
- 21:55only does it make things faster, but it
- 21:57also freed up a ton of storage space.
- 22:00Let's review what we've gone over so
- 22:01far. Deepseek used this sliding window
- 22:04attention to focus on the most immediate
- 22:06bits of information. They also split the
- 22:09model in half and then used hierarchical
- 22:11sparse indexing to significantly reduce
- 22:14the number of notes that the final
- 22:16layers need to process. They also forced
- 22:18some of the layers to just reuse
- 22:20existing notes, which helps lower
- 22:22compute. And then finally, they also
- 22:24completely removed its short-term memory
- 22:26to save hard drive space. But guess
- 22:28what? We're not done yet. So DeepS also
- 22:31added some additional components that
- 22:33support the main architecture. One of
- 22:35them is called the single pass MHC. Now,
- 22:37to understand why this matters or what
- 22:40this is, you first need to know that
- 22:41running an AI model isn't just about
- 22:43doing a huge amount of math. It's also
- 22:46about constantly moving data around.
- 22:48When one part of the model performs
- 22:50calculations, it often produces
- 22:52intermediate values that the next part
- 22:54of the model needs to use. Normally, you
- 22:56might need to store these intermediate
- 22:58results into the GPU's memory and then
- 23:00load it again when the next operation
- 23:02needs these values. And when the AI is
- 23:05processing a really long task, it's
- 23:07basically doing this step but billions
- 23:09of times. Even though one single
- 23:11movement is just a split second, if you
- 23:13multiply this by billions, then this
- 23:15latency can add up. And this step of
- 23:18transferring data can actually be the
- 23:20bottleneck. Well, single path MHC is
- 23:23designed to eliminate some of these
- 23:24unnecessary trips. It's quite technical,
- 23:27but basically it mathematically aligns
- 23:29multiple operations together so that
- 23:32they can occur simultaneously. It's
- 23:34basically combining steps instead of
- 23:36running them one by one. And it turns
- 23:38out that this significantly reduces the
- 23:41memory traffic inside the GPU, making it
- 23:44way faster. And that's not all. They
- 23:46also introduced another supplementary
- 23:48component called the engram. This is
- 23:50kind of like a separate memory module.
- 23:52Now this contains 168 billion parameters
- 23:56and instead of living in the expensive
- 23:58memory of the GPU, this is designed to
- 24:00live in the standard RAM of the server
- 24:02or the computer. So it's physically
- 24:04separated from the GPU. And this part is
- 24:07in charge of storing static facts. So
- 24:09it's memorizing things like historical
- 24:11dates, capitals, or other fixed facts.
- 24:14You see, this is actually really
- 24:16important because we want the GPU to
- 24:18focus entirely on active reasoning and
- 24:21thinking. We want to free up as much of
- 24:23the GPU's expensive memory as possible
- 24:26to maximize its thinking capabilities.
- 24:28For these fixed static facts, which the
- 24:31AI doesn't really need to think about,
- 24:32we can put this in cheaper memory that
- 24:34lives outside the GPU. So, this prevents
- 24:36the GPU's memory from being clogged with
- 24:39static data. Only when the model needs
- 24:41to access these facts would it pull from
- 24:43this engram module. It's kind of like a
- 24:46really brilliant senior lawyer working
- 24:48on a case and synthesizing information.
- 24:50He doesn't need to memorize every single
- 24:52clause out there. He's in charge of the
- 24:54strategic reasoning. Instead, he has an
- 24:56assistant sitting next to him so that
- 24:58when the lawyer needs a specific date or
- 25:00a precise quote, he can just get the
- 25:02assistant to instantly fetch it for him.
- 25:04Well, this assistant is kind of like the
- 25:07engram. It frees up the thinking
- 25:08capacity for the senior lawyer so he can
- 25:11focus on the real strategic reasoning
- 25:14work. And we're not done yet. Deepseek
- 25:16also introduces a third supplementary
- 25:19component called DS-spark. In fact, I
- 25:21already did a full explainer video on
- 25:24D-Spark right when it came out. And this
- 25:26is quite a revolutionary breakthrough.
- 25:28You see, normal AI models need to output
- 25:30their answer one word at a time, which
- 25:32can be very slow. What DSpark does is it
- 25:35essentially allows the model to output
- 25:37multiple words at a time, making it way
- 25:40faster. Now, this is quite technical, so
- 25:42if you're interested, see this video for
- 25:44a full deep dive on DSpark. All right,
- 25:47so we've covered a ton of stuff. Let's
- 25:50take a step back and summarize
- 25:51everything so far. You see, this freak
- 25:53of a model isn't just from one single
- 25:56breakthrough, but a ton of different
- 25:58components added together. They've added
- 26:00this sliding window attention to only
- 26:02focus on the most immediate bits of
- 26:04information. They split the brain into
- 26:07two halves to drastically reduce the
- 26:09compute. We also have extreme
- 26:11compression using CSA 2 and with the
- 26:14reindex and reuse layers. This
- 26:16drastically reduces the amount of nodes
- 26:18or KV cache that it needs to generate.
- 26:21We also completely deleted its
- 26:23short-term memory with this bounded
- 26:25replay mechanism. And we also added a
- 26:27ton of supplemental mechanisms such as
- 26:29this MHC pathway to make its
- 26:31calculations more efficient and this
- 26:33engram component so it doesn't waste any
- 26:35compute on hardfax. And finally, we also
- 26:38added this D-spark mechanism which helps
- 26:40it generate more than one word at a
- 26:42time. And when you combine all these
- 26:44parts together, it becomes an absolute
- 26:46Frankenstein of efficiency. They've
- 26:49built the most optimized, turbocharged,
- 26:51frictionless Frontier model we've seen
- 26:53so far. But don't take my word for it.
- 26:56Here's a chart showing how ridiculous
- 26:58this is. If you look at figure two, this
- 27:00is one of the most incredible
- 27:01demonstrations of the efficiency of this
- 27:04new model. Here it's showing the flops
- 27:06on the y-axis which is like the raw
- 27:08computational power required to generate
- 27:11a response and the x-axis is the context
- 27:14window size or basically how much
- 27:16information it can store in its memory
- 27:18at once. When DeepSeek increases this
- 27:20window from a standard 4,000 tokens all
- 27:23the way to a massive 1 million tokens
- 27:25which is roughly 700,000 words or like a
- 27:28medium-sized code base. You can see that
- 27:30for this new 4.1 flash model, the decode
- 27:33curve remains almost completely flat. In
- 27:35other words, if you feed the AI a tiny
- 27:37one-page document versus thousands of
- 27:40pages and ask it a question, this chart
- 27:42shows that the AI spends roughly the
- 27:44same amount of energy per word. This is
- 27:46a massive deal. It completely defies the
- 27:49laws of AI because normally what we
- 27:51would expect is if you scale the
- 27:53context, the compute also increases as
- 27:55you could see with the previous
- 27:56generations of deepseek models. But here
- 27:59they've basically flattened the curve.
- 28:01And you know the ridiculous thing is not
- 28:03only is this super efficient and
- 28:05frictionless, but its performance is
- 28:07also state-of-the-art, pretty much on
- 28:09par with the Frontier models. For
- 28:11example, if you look at Deep Su 1.1,
- 28:14this scores 74.2,
- 28:16which not only beats the other open
- 28:18models out there, but if you look at the
- 28:20official leaderboard, then it's
- 28:21basically on par with GPT6 Astra, which
- 28:24also scores 74. Or if you look at
- 28:27Cyberjim, this is pretty much
- 28:28state-of-the-art. Same with Automation
- 28:30Bench. You can see that this freak of a
- 28:32model even beats GPT6 Astro Max. If you
- 28:36look at this leaderboard by LiveBench,
- 28:38then you can see that this new Deepseek
- 28:40is the number one ranked open model.
- 28:42Same with Val's index, which measures an
- 28:44AI's performance across knowledge work
- 28:46tasks. You can see that Deepseek 4.1
- 28:49Flash is currently ranked number one.
- 28:51And look at the insane cost of this. Not
- 28:53only is this the most performant model,
- 28:56but it's also the cheapest. You can see
- 28:58that second and third place are over 20
- 29:00times more expensive. Here's another
- 29:02chart showing the cost per task. You can
- 29:04see that this new DeepSeek model is all
- 29:06the way over here, which is like way
- 29:08cheaper than the Frontier GPT6 Astra as
- 29:12well as the extremely overpriced clawed
- 29:14models. If you look at the output speed
- 29:16if you use it through their API, again,
- 29:18this is insane. This achieves over 200
- 29:21tokens per second, which is like four
- 29:23times faster than GPT6. The latency, in
- 29:26other words, the time to its first
- 29:28answer is also the lowest in the
- 29:30industry. All right, so we've gone over
- 29:32a ton of stuff. The way they designed it
- 29:34is just so unexpected and at many times
- 29:36counterintuitive, but once you
- 29:38understand why they did it, then you'll
- 29:40see how brilliant and genius this is.
- 29:43And as always, they've open sourced the
- 29:45model for you to download locally, so
- 29:47you can do whatever you want with it. My
- 29:49hat off to the Deep Seek team for
- 29:51pulling off the impossible. Once again,
- 29:54I mean, the previous DeepSseek V4 was
- 29:56already super efficient, but with this
- 29:58latest model, they've managed to
- 29:59completely redesign the architecture and
- 30:02squeeze even more juice out of it. That
- 30:04sums up my deep dive on this new Deep
- 30:07Seek 4.1 Flash. This is one of the more
- 30:09technical papers I reviewed so far on my
- 30:11channel. So hopefully I made it easy for
- 30:14you to digest. In fact, the paper is
- 30:16jam-packed with a ton of additional
- 30:18technical details which I didn't have
- 30:20time to cover. So if you're interested
- 30:21in digging deeper, I'll link to this
- 30:24original paper in the description below
- 30:26as well. Let me know in the comments
- 30:27what you think of this. As always, I
- 30:29will be on the lookout for the top AI
- 30:32news and tools to share with you. So, if
- 30:34you enjoyed this video, remember to
- 30:36like, share, subscribe, and stay tuned
- 30:38for more content. Also, there's just so
- 30:41much happening in the world of AI every
- 30:43week, I can't possibly cover everything
- 30:45on my YouTube channel. So, to really
- 30:47stay up to date with all that's going on
- 30:50in AI, be sure to subscribe to my free
- 30:53weekly newsletter. The link to that will
- 30:55be in the description below. Thanks for
- 30:57watching, and I'll see you in the next
- 30:58one.
About this transcript
This page contains the full transcript of Deepseek just did the impossible by AI Search, generated from the public captions YouTube serves with the video. The transcript has 5,814 words across 859 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.