YouTube2Text

Deepseek just did the impossible — Transcript

by AI Search · 5,814 words · 859 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Deepseek did it again. They've released
  2. 0:02a new model and not only is this among
  3. 0:05the frontier models out there, but it's
  4. 0:07also the most optimized, efficient, and
  5. 0:10frictionless AI model we've seen so far.
  6. 0:13As always, not only have they open-
  7. 0:15sourced the model, but they've also
  8. 0:16released a technical paper on this. And
  9. 0:18how they designed this is just really
  10. 0:20unexpected and sometimes even seems
  11. 0:23wrong. But once you understand the logic
  12. 0:25behind everything, then it suddenly
  13. 0:27becomes absolutely brilliant. In this
  14. 0:29video, we're going to do a deep dive
  15. 0:31into its design so that you can see how
  16. 0:33genius and cracked this is. Now, this is
  17. 0:36a super technical paper, but as always,
  18. 0:38I'm going to break it down into simple
  19. 0:40terms so that anyone can understand.
  20. 0:42Let's jump right in. Let's first set the
  21. 0:44stage by reviewing the situation that
  22. 0:46Deep Seek is in. Keep in mind that this
  23. 0:48is just a small Chinese lab that does
  24. 0:51not have nearly as much funding as
  25. 0:53OpenAI. Their team is like dozens of
  26. 0:55times smaller. Plus, they don't even
  27. 0:57have access to the best Nvidia GPUs out
  28. 0:59there, nor do they have a massive data
  29. 1:01center. In fact, they are severely
  30. 1:04constrained in terms of compute and
  31. 1:06resources. Yet, they just released their
  32. 1:08latest model, Deepseek V4.1 Flash, which
  33. 1:12even matches the performance of Frontier
  34. 1:14models, even though this is just a Flash
  35. 1:16model. Plus, it's way faster and more
  36. 1:18efficient. And get this, its memory
  37. 1:21footprint is like over 400 times smaller
  38. 1:24compared to the first generation. How on
  39. 1:26earth did they pull this off? Well, to
  40. 1:28understand the brilliance of their
  41. 1:30solution, we first need to go over the
  42. 1:32basics of what actually happens when you
  43. 1:34use an AI model. It's actually broken
  44. 1:36down into two different phases. The
  45. 1:38first phase is called prefill. This is
  46. 1:41basically when the AI reads your prompt
  47. 1:43along with any other information or
  48. 1:45documents you give it so it can
  49. 1:47understand the context of everything
  50. 1:48before it starts generating an answer.
  51. 1:51Specifically, your text is broken down
  52. 1:53into smaller pieces called tokens which
  53. 1:55are then passed through the AI model's
  54. 1:57layers. In fact, each layer in the model
  55. 2:00calculates two important sets of numbers
  56. 2:02called keys and values for each token.
  57. 2:05These are stored in something called the
  58. 2:07KV cache. You can think of this KV cache
  59. 2:10as like the notes the model takes as it
  60. 2:12reads through your information. Now the
  61. 2:14second phase is the decode phase or the
  62. 2:17writing phase. This is where the AI
  63. 2:19starts generating your answer. Now
  64. 2:20interestingly large language models
  65. 2:22generate its answer one word or token at
  66. 2:25a time. In order to do so, the model
  67. 2:27needs to look at everything that came
  68. 2:29before it to figure out the next most
  69. 2:30probable word that should come next. And
  70. 2:33to do that efficiently, it needs to
  71. 2:34refer back to its notes that it took
  72. 2:37during the reading phase. In other
  73. 2:38words, it needs to look at all the KV
  74. 2:40cache that it has created. This KV cache
  75. 2:43lets the model quickly access the
  76. 2:44relevant information without having to
  77. 2:46recalculate the entire conversation from
  78. 2:49scratch. To understand this a bit
  79. 2:50better, here's a nice analogy. Think of
  80. 2:52an AI model like a student watching a
  81. 2:54really long lecture and then taking
  82. 2:56notes along the way. Later, when the
  83. 2:58student needs to answer a question about
  84. 3:00the lecture, he doesn't need to replay
  85. 3:01the entire lecture from the beginning.
  86. 3:04Instead, he can just refer back to his
  87. 3:05notes. Well, you can think of the KV
  88. 3:07cache like these notes. It gives the AI
  89. 3:10a quick way to refer back to information
  90. 3:12it has already processed without having
  91. 3:14to recalculate everything from scratch.
  92. 3:17Now, here's the problem the industry is
  93. 3:19facing right now. You see, for a short
  94. 3:21prompt like this, everything works fine.
  95. 3:23The AI model can easily convert this
  96. 3:25into a fairly small KV cache that fits
  97. 3:28well within its active memory. But
  98. 3:30here's the thing, the industry is now
  99. 3:32focused on making AI handle really long
  100. 3:34and complex tasks. We want to get AI
  101. 3:37agents to work autonomously for hours or
  102. 3:40even days. We want to give them a ton of
  103. 3:43different documents or a huge codebase
  104. 3:45to keep track of and keep working on for
  105. 3:47a really long time. And in these
  106. 3:49scenarios, the KV cache or the notes
  107. 3:51that the AI has to take is going to be
  108. 3:53massive. So going back to our analogy
  109. 3:56now, instead of just watching one
  110. 3:58lecture, the student has to watch weeks
  111. 4:00and weeks of online lectures and take
  112. 4:02notes on all of them. His notes will
  113. 4:04start to pile up fast to the point where
  114. 4:06they can't even fit on his desk anymore.
  115. 4:08They're going to fill up his entire room
  116. 4:10and eventually he needs to like shove
  117. 4:12his notes into filing cabinets in
  118. 4:14another room down the hallway. And then
  119. 4:16every time the student needs to answer a
  120. 4:18question, he needs to dig through all
  121. 4:20these huge piles of notes. Well, this
  122. 4:22analogy is exactly what happens to an AI
  123. 4:25when it needs to work with a ton of
  124. 4:26information. You see, inside a GPU, you
  125. 4:29have this high bandwidth memory or HBM.
  126. 4:32This is the incredibly fast memory
  127. 4:34that's right next to the processor chip.
  128. 4:36It's really fast and easy to access. So,
  129. 4:38it's kind of like the surface of a
  130. 4:40student's desk in our analogy. any notes
  131. 4:42that are sitting on this desk are
  132. 4:44incredibly quick and easy to access. But
  133. 4:46the desk space is limited. Well, high
  134. 4:49bandwidth memory is exactly the same.
  135. 4:51There's limited space and it's also
  136. 4:52extremely expensive. So when the KV
  137. 4:54cache gets too big, this high bandwidth
  138. 4:57memory gets full. So the computer has to
  139. 4:59start storing this data somewhere else.
  140. 5:01For example, on solidstate drives or
  141. 5:03SSDs which are outside the GPU. So, in
  142. 5:07our analogy, it's like putting some of
  143. 5:09the students notes into filing cabinets
  144. 5:11in another room down the hallway. You
  145. 5:13have much more storage space there, but
  146. 5:15there's a trade-off. He needs to go to
  147. 5:17the other room to grab the notes from a
  148. 5:19cabinet, which takes much longer than if
  149. 5:21it was just sitting in front of him on
  150. 5:23his desk. So, similarly, SSDs, or these
  151. 5:25solidstate drives outside the GPU, have
  152. 5:28way more storage capacity, but they're
  153. 5:31located further away from the GPU, and
  154. 5:33the latency there is devastating. If the
  155. 5:35KV cache needs to be stored there, well,
  156. 5:37every time the AI needs to predict the
  157. 5:39next word, it has to fetch data from the
  158. 5:42SSD, pull it through the motherboard up
  159. 5:44into its high bandwidth memory, and then
  160. 5:46into the GPU processor. This data
  161. 5:49transfer speed becomes the absolute
  162. 5:51bottleneck. The processor is just
  163. 5:53essentially sitting idle waiting for the
  164. 5:55massive KV cache notes to travel through
  165. 5:57the wires. So, to sum things up, there
  166. 5:59are two major problems that current AI
  167. 6:02systems are facing. One problem is
  168. 6:04compute. The student is basically just
  169. 6:06drowning in his own notes and it's
  170. 6:08really hard for him to find things. The
  171. 6:10second problem is speed. Because there
  172. 6:12are so many notes, some of these notes
  173. 6:14are stored in filing cabinets in another
  174. 6:16room and the student needs to run back
  175. 6:18and forth to fetch these notes and put
  176. 6:20them on his desk which takes a ton of
  177. 6:22time. So these are the limitations that
  178. 6:24DeepS is facing. How on earth did they
  179. 6:27tackle this? First of all, let's go over
  180. 6:29the architecture of a normal large
  181. 6:31language model. They use the transformer
  182. 6:33architecture which is basically made up
  183. 6:35of many layers. In fact, if you're
  184. 6:38curious about how transformers actually
  185. 6:40work under the hood, definitely see this
  186. 6:42video where I do a full explainer. But
  187. 6:45anyway, what happens is your prompt is
  188. 6:47broken down into data that flows through
  189. 6:50these layers of the transformer which
  190. 6:52ultimately outputs the next most
  191. 6:53probable word for its answer. And then
  192. 6:55that word is appended back and then it's
  193. 6:58run through the model again to predict
  194. 7:00the next most probable word. and then
  195. 7:01this loops again until it generates your
  196. 7:03final answer. Now, each layer that it
  197. 7:05goes through generates its own notes or
  198. 7:08KV cache. And all of this must be stored
  199. 7:10somewhere. This often takes up a lot of
  200. 7:12space. So, not only does it fill up the
  201. 7:14high bandwidth memory, but it also
  202. 7:15spills over to SSDs that are outside the
  203. 7:18GPU. Well, the architecture from this
  204. 7:20new deepse has many layers. But here's
  205. 7:23the really unusual part. They split this
  206. 7:25into two halves. There's a causal
  207. 7:27encoder component and then there's a
  208. 7:29decoder component. This already looks
  209. 7:31completely different from a standard
  210. 7:33transformer model. And here's how it
  211. 7:35works. During the prefill phase, again,
  212. 7:37this is where the AI is reading your
  213. 7:39prompt and all the information you give
  214. 7:41it. The last half is essentially turned
  215. 7:43off. They just do nothing. And then the
  216. 7:45first half of the layers do all the
  217. 7:47heavy lifting. They read everything.
  218. 7:49They build the contextual understanding
  219. 7:51and they generate what's called the
  220. 7:52global KV cache. After reading
  221. 7:55everything, then the last half basically
  222. 7:57generates the answer. it does the
  223. 7:59writing or the decoding phase. The thing
  224. 8:02is it still needs to know the context in
  225. 8:04order to do the writing right to predict
  226. 8:06the next word. So you might be wondering
  227. 8:07if it doesn't generate the KV cache
  228. 8:09itself, how does it understand the
  229. 8:11context? Here's why this design is so
  230. 8:13genius. Instead of having to read
  231. 8:15everything itself, it just looks at the
  232. 8:17output from the last layer of this
  233. 8:19encoder component. In other words, it
  234. 8:22borrows the completed global KV cache
  235. 8:24directly from the end of this encoder
  236. 8:27block. Now, this is quite a shocking and
  237. 8:29unexpected design because if they did
  238. 8:31this, if half its brain basically
  239. 8:33skipped the reading part, doesn't it
  240. 8:35lose some understanding of the context,
  241. 8:37which would make its answer worse? Well,
  242. 8:39that's the exact risk of this
  243. 8:41architecture. And that's why it took
  244. 8:43incredible engineering to balance.
  245. 8:45What's fascinating here is that this
  246. 8:47last half, this decoder component,
  247. 8:49doesn't really entirely skip reading. It
  248. 8:51just skips calculating the global
  249. 8:53context. In other words, the global KV
  250. 8:55cache, but it still computes the local
  251. 8:57context for itself. Now, you might be
  252. 8:59wondering, what's the difference between
  253. 9:01global and local context here? Well,
  254. 9:04global context is basically all the
  255. 9:06information that was given to it plus
  256. 9:08your prompt and everything else
  257. 9:09attached. It's basically all the notes
  258. 9:11that the student took after watching
  259. 9:13weeks and weeks of lectures. In
  260. 9:14contrast, the local context is just the
  261. 9:17immediate context of the sentence the AI
  262. 9:19is currently writing. In other words,
  263. 9:21what's directly relevant to the next
  264. 9:23word that it needs to output. You'll see
  265. 9:26this decoder component doesn't calculate
  266. 9:28any global KV cache, but instead it uses
  267. 9:31something called a sliding window
  268. 9:33attention to pay extremely close
  269. 9:35attention to the most recent tokens
  270. 9:37only, but not everything before it. And
  271. 9:39it turns out that this design works
  272. 9:41pretty well. Here's a nice analogy to
  273. 9:43wrap your head around this. Imagine a
  274. 9:45company where they need to analyze
  275. 9:47thousands of pages of financial reports.
  276. 9:50Well, they would first get junior
  277. 9:52analysts to read every single page and
  278. 9:54crunch out the numbers, do the data
  279. 9:55analysis, and then write a very dense
  280. 9:57and accurate executive summary. Well,
  281. 9:59these junior analysts are basically like
  282. 10:02the first half of the model. And then
  283. 10:04this executive summary is like the
  284. 10:06output at the end. They then hand the
  285. 10:07summary to the senior executives, which
  286. 10:10are like the decoder layers. These
  287. 10:12senior executives absolutely do not read
  288. 10:15the original thousands of pages.
  289. 10:17Instead, they just rely on the executive
  290. 10:20summary provided by the juniors. But
  291. 10:21when it comes time to sign off on
  292. 10:23something, then these senior executives
  293. 10:25put on their reading glasses and
  294. 10:27scrutinize the exact wording of the page
  295. 10:29sitting right in front of them. That's
  296. 10:31basically the local context. The senior
  297. 10:34executives rely on this global summary
  298. 10:36for direction, but they mostly focus
  299. 10:39locally on the page in front of them for
  300. 10:42execution. And by structuring the model
  301. 10:43this way, DeepC completely bypasses the
  302. 10:46need for basically half of the model to
  303. 10:48generate its own massive KV cache notes.
  304. 10:51And this is a huge deal. It essentially
  305. 10:53slashes the compute required to read
  306. 10:56things by half. Now, reducing the
  307. 10:58compute is great, but we still have the
  308. 11:00problem of memory, right? It's still
  309. 11:02generating these massive KV cache notes
  310. 11:04whether it's global or local and these
  311. 11:06are like overflowing on the student's
  312. 11:09desk and he's forced to like store these
  313. 11:11excess notes in filing cabinets in
  314. 11:13another room. How on earth can we reduce
  315. 11:15these massive piles of notes? And this
  316. 11:17brings us to one of the most fascinating
  317. 11:19parts of DeepS's new design. And this
  318. 11:22part is just brilliant. In fact, let me
  319. 11:24show you the results first so you can
  320. 11:26see how insane this is. If you do any
  321. 11:28kind of content creation, definitely
  322. 11:30check out Luma, the sponsor of this
  323. 11:32video. Think of it as a creative AI
  324. 11:35agent that works alongside you through
  325. 11:37your entire creative process. Instead of
  326. 11:39just giving you the results of a single
  327. 11:41prompt, I can access the best image and
  328. 11:44video models out there. And the nice
  329. 11:46thing is instead of manually jumping
  330. 11:48between these tools, I can just get Luma
  331. 11:50agents to autonomously do entire
  332. 11:52workflows for me. It can develop the
  333. 11:54concept, generate the visuals and shape
  334. 11:56the project all within the same
  335. 11:58workspace. For example, I can get it to
  336. 12:00generate a brand kit for me, design
  337. 12:02different products, and generate other
  338. 12:04marketing assets all inside the same
  339. 12:06project. And if I need to edit
  340. 12:08something, I can just prompt the agent
  341. 12:09to refine the results iteratively. One
  342. 12:11of the most powerful features is Luma
  343. 12:13skills. You can basically create
  344. 12:15reusable skills for workflows you use
  345. 12:18all the time. Basically, you give Luma a
  346. 12:20set of instructions once and then you
  347. 12:22can run that same workflow on different
  348. 12:24assets whenever you want. For example, I
  349. 12:26can create a skill where I can input any
  350. 12:28product photo and it'll output some UGC
  351. 12:31videos of an influencer talking about
  352. 12:33the product. Or here's another example
  353. 12:35of a skill where I can upload a product
  354. 12:37photo and it'll generate a 360° orbit
  355. 12:40view like this. Luma basically gives you
  356. 12:42an intelligent creative co-pilot that
  357. 12:45can autonomously carry out your
  358. 12:46workflows. Whether you're creating
  359. 12:48marketing campaigns, branded content,
  360. 12:50product visuals, or social media
  361. 12:52content, Luma is one of the best
  362. 12:54platforms you can use. Try Luma today
  363. 12:56using the link in the description below
  364. 12:58or by scanning the QR code here. If you
  365. 13:00compare the global KV cache size per
  366. 13:03token, which is basically the size of
  367. 13:05the notes the student has to take,
  368. 13:07DeepSeek V1 is almost 390,000
  369. 13:10bytes. Now, if you fast forward just a
  370. 13:12few generations to this latest V4.1
  371. 13:15Flash, it's only 890 bytes per token.
  372. 13:18They basically shrunk the size of the
  373. 13:20notes down by like 437 times, which is
  374. 13:24crazy. Even if you compare this to the
  375. 13:26previous DeepSseek V4 Flash, this one
  376. 13:29still required like 3,500 bytes per
  377. 13:31token. So, this new update is like
  378. 13:33almost four times smaller than the
  379. 13:35previous generation. But here's the
  380. 13:37challenge to all of this. How can you
  381. 13:39compress these notes so much without
  382. 13:41losing the meaning? How can you still
  383. 13:42maintain the AI model's understanding of
  384. 13:44everything? Well, Deepseek used a
  385. 13:46mechanism called compressed sparse
  386. 13:49attention 2 or CSA2. To understand this,
  387. 13:52let's first review how a normal
  388. 13:53transformer model works. Each layer in
  389. 13:56the model has to calculate its own KV
  390. 13:58cache. In other words, it has to make
  391. 13:59its own nodes. Conceptually, you can
  392. 14:01think of each layer as focusing on
  393. 14:03different things. For example, some
  394. 14:05layers might focus on certain patterns,
  395. 14:07while others focus on things like
  396. 14:09grammar or relationships between ideas
  397. 14:11or the broader meaning of the text. Now,
  398. 14:14with all these layers each making their
  399. 14:16own notes, you can see how the total
  400. 14:18size of the KV cache could become really
  401. 14:20hard to manage. And if each layer has to
  402. 14:22calculate its own notes, you can see how
  403. 14:24things could become redundant. Well,
  404. 14:26this new CSA2 mechanism by DeepSeek
  405. 14:30completely shatters this redundancy. It
  406. 14:32introduces the concept of extreme
  407. 14:34sharing. So instead of creating new
  408. 14:36notes from scratch every time, each
  409. 14:38layer could use three different
  410. 14:40operating modes. Full mode, reindex
  411. 14:42mode, and reuse. Let's go over each one.
  412. 14:45So if the layer is in full mode, it has
  413. 14:47to do all the hard work. It has to
  414. 14:49create brand new notes from scratch. In
  415. 14:51other words, it needs to make the full
  416. 14:53KV cache. But here's the important part.
  417. 14:55It also creates an index for future
  418. 14:58layers to search these notes. Think of
  419. 15:00the index like a guide or map or like a
  420. 15:03table of contents. Other layers can just
  421. 15:05look at this table of contents to figure
  422. 15:07out where exactly to search in the notes
  423. 15:09instead of trying to read the whole
  424. 15:11thing from start to finish. So this
  425. 15:12makes it way faster to search for
  426. 15:14information. Another mode is called
  427. 15:16reindex. And here's where the efficiency
  428. 15:19kicks in. If the layer is in reindex
  429. 15:21mode, it just reuses the notes or in
  430. 15:23other words the KV cache from a layer in
  431. 15:25full mode. It doesn't create its own
  432. 15:27notes from scratch. But what it does do
  433. 15:29is create a new index from scratch.
  434. 15:31Again, think of this as like creating a
  435. 15:33guide or a new table of contents that
  436. 15:36searches the same notes as before, but
  437. 15:38highlights completely different parts.
  438. 15:40For example, let's say you're giving the
  439. 15:41AI a ton of information about the
  440. 15:43history of the world. The full layer
  441. 15:45could make an index about things in
  442. 15:47chronological order, which might look
  443. 15:49like this. A reindexed layer would take
  444. 15:51the exact same notes, but give it a
  445. 15:53completely different table of contents.
  446. 15:55For example, instead of chronological
  447. 15:57events, it could be themes across time
  448. 15:59or it could be different technological
  449. 16:01breakthroughs. It's basically like
  450. 16:03mapping different paths through the same
  451. 16:05notes. So that's the reindex mode. And
  452. 16:07then finally, we have the third mode,
  453. 16:09which is maximum efficiency. And this is
  454. 16:12the reuse mode. If the layer has this
  455. 16:14mode, it exerts almost zero memory
  456. 16:17effort. It just reuses the notes from
  457. 16:19the full layer as well as the indices
  458. 16:21from either the full layer or the
  459. 16:23reindexed layers. It doesn't write any
  460. 16:25new notes, nor does it create any new
  461. 16:27table of contents. It just takes
  462. 16:29information that's already available to
  463. 16:31it. And with this design, with these
  464. 16:34different modes, each layer doesn't have
  465. 16:36to store as much notes. The total size
  466. 16:38of these notes is reduced significantly
  467. 16:41because some of these layers don't even
  468. 16:42need to create new notes at all. They're
  469. 16:44just reusing notes and indices from
  470. 16:46previous layers. But wait, this ain't
  471. 16:49all. Deepseek takes this one step
  472. 16:51further. They've added something called
  473. 16:53a hierarchical sparse indexer. And
  474. 16:55here's how it works. You see, in the
  475. 16:57last half of the model, this is the
  476. 16:59decoder part. The very first layer acts
  477. 17:02kind of like a gatekeeper. It scans all
  478. 17:04the global notes from the previous
  479. 17:06layer. Remember, this is also called the
  480. 17:08global KV cache and it generates a
  481. 17:10candidate pool. Think of this like a
  482. 17:12short list. Basically, from those piles
  483. 17:15and piles of notes, it figures out just
  484. 17:17the most relevant concepts.
  485. 17:19Specifically, out of a million tokens,
  486. 17:21it only selects around 16,000 that are
  487. 17:24the most relevant. Everything beyond it
  488. 17:26is basically ignored and then all the
  489. 17:28subsequent layers are forbidden to
  490. 17:30search for anything else in the notes.
  491. 17:32So, this drastically narrows the
  492. 17:34universe of possible answers right at
  493. 17:36the start of the writing phase. Now,
  494. 17:38obviously, as you may expect, if we
  495. 17:40restrict the AI's search space like
  496. 17:42this, it might hallucinate or miss
  497. 17:44important details, right? its response
  498. 17:46will become dumber if it doesn't look at
  499. 17:48everything. But here's the genius behind
  500. 17:50this. Deepseek was able to make it work
  501. 17:52by training the model to build these
  502. 17:54candidates with such high accuracy that
  503. 17:56the later layers don't even notice the
  504. 17:58rest of the info is missing. The
  505. 18:00elegance of their engineering is just
  506. 18:02profound. Let's take a moment to
  507. 18:04appreciate what they did here. They
  508. 18:06don't have the best Nvidia GPUs. They
  509. 18:08don't have the biggest data center in
  510. 18:09the world. Heck, they're pretty starved
  511. 18:11for compute. So instead of focusing on
  512. 18:13the hardware side, they completely
  513. 18:15optimized the software part so the
  514. 18:17hardware doesn't have to work so hard.
  515. 18:19And we ain't done yet. You see, all this
  516. 18:22intense optimization also led to some
  517. 18:24additional issues they had to fix.
  518. 18:26Remember this sliding window attention
  519. 18:28mechanism we talked about earlier. This
  520. 18:30is where the senior executive doesn't
  521. 18:32read everything, but when he signs off
  522. 18:34on something, he has to look really
  523. 18:35closely at all the information on the
  524. 18:37page in front of him. Well, this sliding
  525. 18:39window attention is the AI's hyper local
  526. 18:42short-term memory. And according to the
  527. 18:44paper, this hyper local memory was
  528. 18:46actually causing a huge storage problem.
  529. 18:49You see, when a user has a conversation
  530. 18:51with the AI over multiple turns. In
  531. 18:54other words, when you say something, the
  532. 18:55AI replies and then you reply back and
  533. 18:57so on. The AI has to save this local
  534. 19:00context of every single turn into its
  535. 19:03SSD so it won't forget the flow of the
  536. 19:05conversation. And this constant saving
  537. 19:07of short-term local memory was
  538. 19:09completely clogging the hard drives.
  539. 19:12Going back to our analogy, this is like
  540. 19:14filling up all the cabinets in the other
  541. 19:15room down the hall. Now, when you cache
  542. 19:18data, it means you save it in a
  543. 19:20temporary location so you can retrieve
  544. 19:22it quickly later, right? But if it gets
  545. 19:24too big and these notes are located in
  546. 19:26cabinets in another room, well, fetching
  547. 19:28this information gets super slow. the
  548. 19:31student needs to run down the hallway to
  549. 19:33the other room to grab the notes and
  550. 19:35then place them back on his desk. So,
  551. 19:37DeepSeek recognized this issue and their
  552. 19:40solution was something called SWA
  553. 19:42bounded replay. And this is probably the
  554. 19:44most unexpected and shocking part of
  555. 19:46their new design. Their solution to this
  556. 19:49local memory clogging up the hard drives
  557. 19:51is to simply delete it. They literally
  558. 19:53deleted the AI's short-term memory
  559. 19:55completely. So, when the AI finishes
  560. 19:57generating its reply to you, its local
  561. 19:59memory just evaporates into thin air. As
  562. 20:02you can imagine, if we delete its
  563. 20:04short-term memory, shouldn't it become
  564. 20:06like completely disoriented? Wouldn't it
  565. 20:08lose track of what's going on? Well,
  566. 20:10that's what we would expect. So, what
  567. 20:12Deep Seek did was they basically got the
  568. 20:14AI to recalculate the last parts of the
  569. 20:17conversation instantly, specifically the
  570. 20:19last 128 tokens of the conversation. In
  571. 20:21other words, they forced it to generate
  572. 20:23the most immediate new notes from
  573. 20:25scratch right on the spot every single
  574. 20:27time. This sounds so counterintuitive,
  575. 20:29right? I just spent the past few minutes
  576. 20:31in this video explaining how they tried
  577. 20:33to reduce the compute and split the
  578. 20:35brain into halves to avoid generating
  579. 20:37and reading so many notes. But now, if
  580. 20:39we get this AI to recalculate this last
  581. 20:42part of the conversation every time,
  582. 20:44wouldn't this be incredibly inefficient?
  583. 20:46Doesn't this slow everything down? Well,
  584. 20:48here's where DeepSseek gives us a
  585. 20:51masterclass in efficiency. Let's walk
  586. 20:53through the trade-off here. You kind of
  587. 20:55have two possible ways to handle this
  588. 20:57short-term memory. The first way is the
  589. 20:59normal way where we take data from the
  590. 21:01GPU, we push it through the motherboard
  591. 21:03and write it into a hard drive. And
  592. 21:06later, when it needs to use this shorter
  593. 21:08memory, it needs to search for it, pull
  594. 21:10it back up through the motherboard and
  595. 21:11into the GPU processor. This is super
  596. 21:14slow because you're physically moving
  597. 21:15data across this distance. Now, the
  598. 21:18second way is to just use the raw power
  599. 21:20of a modern GPU to simply recalculate
  600. 21:23its short-term memory. In other words,
  601. 21:24just rewrite the most recent and
  602. 21:26relevant notes from scratch. And for a
  603. 21:28modern GPU, just crunching like 128
  604. 21:31tokens is just a microscond operation
  605. 21:33that barely takes up any time or power.
  606. 21:36So, doing the math is actually much
  607. 21:38faster than trying to transfer this data
  608. 21:40back and forth from the SSD. The Deep
  609. 21:43Seek team realized that saving this
  610. 21:45short-term memory to the hard drive was
  611. 21:47just a massive waste of a really slow
  612. 21:49resource. By completely removing the
  613. 21:51step and just getting the GPU to quickly
  614. 21:53recalculate everything on the fly, not
  615. 21:55only does it make things faster, but it
  616. 21:57also freed up a ton of storage space.
  617. 22:00Let's review what we've gone over so
  618. 22:01far. Deepseek used this sliding window
  619. 22:04attention to focus on the most immediate
  620. 22:06bits of information. They also split the
  621. 22:09model in half and then used hierarchical
  622. 22:11sparse indexing to significantly reduce
  623. 22:14the number of notes that the final
  624. 22:16layers need to process. They also forced
  625. 22:18some of the layers to just reuse
  626. 22:20existing notes, which helps lower
  627. 22:22compute. And then finally, they also
  628. 22:24completely removed its short-term memory
  629. 22:26to save hard drive space. But guess
  630. 22:28what? We're not done yet. So DeepS also
  631. 22:31added some additional components that
  632. 22:33support the main architecture. One of
  633. 22:35them is called the single pass MHC. Now,
  634. 22:37to understand why this matters or what
  635. 22:40this is, you first need to know that
  636. 22:41running an AI model isn't just about
  637. 22:43doing a huge amount of math. It's also
  638. 22:46about constantly moving data around.
  639. 22:48When one part of the model performs
  640. 22:50calculations, it often produces
  641. 22:52intermediate values that the next part
  642. 22:54of the model needs to use. Normally, you
  643. 22:56might need to store these intermediate
  644. 22:58results into the GPU's memory and then
  645. 23:00load it again when the next operation
  646. 23:02needs these values. And when the AI is
  647. 23:05processing a really long task, it's
  648. 23:07basically doing this step but billions
  649. 23:09of times. Even though one single
  650. 23:11movement is just a split second, if you
  651. 23:13multiply this by billions, then this
  652. 23:15latency can add up. And this step of
  653. 23:18transferring data can actually be the
  654. 23:20bottleneck. Well, single path MHC is
  655. 23:23designed to eliminate some of these
  656. 23:24unnecessary trips. It's quite technical,
  657. 23:27but basically it mathematically aligns
  658. 23:29multiple operations together so that
  659. 23:32they can occur simultaneously. It's
  660. 23:34basically combining steps instead of
  661. 23:36running them one by one. And it turns
  662. 23:38out that this significantly reduces the
  663. 23:41memory traffic inside the GPU, making it
  664. 23:44way faster. And that's not all. They
  665. 23:46also introduced another supplementary
  666. 23:48component called the engram. This is
  667. 23:50kind of like a separate memory module.
  668. 23:52Now this contains 168 billion parameters
  669. 23:56and instead of living in the expensive
  670. 23:58memory of the GPU, this is designed to
  671. 24:00live in the standard RAM of the server
  672. 24:02or the computer. So it's physically
  673. 24:04separated from the GPU. And this part is
  674. 24:07in charge of storing static facts. So
  675. 24:09it's memorizing things like historical
  676. 24:11dates, capitals, or other fixed facts.
  677. 24:14You see, this is actually really
  678. 24:16important because we want the GPU to
  679. 24:18focus entirely on active reasoning and
  680. 24:21thinking. We want to free up as much of
  681. 24:23the GPU's expensive memory as possible
  682. 24:26to maximize its thinking capabilities.
  683. 24:28For these fixed static facts, which the
  684. 24:31AI doesn't really need to think about,
  685. 24:32we can put this in cheaper memory that
  686. 24:34lives outside the GPU. So, this prevents
  687. 24:36the GPU's memory from being clogged with
  688. 24:39static data. Only when the model needs
  689. 24:41to access these facts would it pull from
  690. 24:43this engram module. It's kind of like a
  691. 24:46really brilliant senior lawyer working
  692. 24:48on a case and synthesizing information.
  693. 24:50He doesn't need to memorize every single
  694. 24:52clause out there. He's in charge of the
  695. 24:54strategic reasoning. Instead, he has an
  696. 24:56assistant sitting next to him so that
  697. 24:58when the lawyer needs a specific date or
  698. 25:00a precise quote, he can just get the
  699. 25:02assistant to instantly fetch it for him.
  700. 25:04Well, this assistant is kind of like the
  701. 25:07engram. It frees up the thinking
  702. 25:08capacity for the senior lawyer so he can
  703. 25:11focus on the real strategic reasoning
  704. 25:14work. And we're not done yet. Deepseek
  705. 25:16also introduces a third supplementary
  706. 25:19component called DS-spark. In fact, I
  707. 25:21already did a full explainer video on
  708. 25:24D-Spark right when it came out. And this
  709. 25:26is quite a revolutionary breakthrough.
  710. 25:28You see, normal AI models need to output
  711. 25:30their answer one word at a time, which
  712. 25:32can be very slow. What DSpark does is it
  713. 25:35essentially allows the model to output
  714. 25:37multiple words at a time, making it way
  715. 25:40faster. Now, this is quite technical, so
  716. 25:42if you're interested, see this video for
  717. 25:44a full deep dive on DSpark. All right,
  718. 25:47so we've covered a ton of stuff. Let's
  719. 25:50take a step back and summarize
  720. 25:51everything so far. You see, this freak
  721. 25:53of a model isn't just from one single
  722. 25:56breakthrough, but a ton of different
  723. 25:58components added together. They've added
  724. 26:00this sliding window attention to only
  725. 26:02focus on the most immediate bits of
  726. 26:04information. They split the brain into
  727. 26:07two halves to drastically reduce the
  728. 26:09compute. We also have extreme
  729. 26:11compression using CSA 2 and with the
  730. 26:14reindex and reuse layers. This
  731. 26:16drastically reduces the amount of nodes
  732. 26:18or KV cache that it needs to generate.
  733. 26:21We also completely deleted its
  734. 26:23short-term memory with this bounded
  735. 26:25replay mechanism. And we also added a
  736. 26:27ton of supplemental mechanisms such as
  737. 26:29this MHC pathway to make its
  738. 26:31calculations more efficient and this
  739. 26:33engram component so it doesn't waste any
  740. 26:35compute on hardfax. And finally, we also
  741. 26:38added this D-spark mechanism which helps
  742. 26:40it generate more than one word at a
  743. 26:42time. And when you combine all these
  744. 26:44parts together, it becomes an absolute
  745. 26:46Frankenstein of efficiency. They've
  746. 26:49built the most optimized, turbocharged,
  747. 26:51frictionless Frontier model we've seen
  748. 26:53so far. But don't take my word for it.
  749. 26:56Here's a chart showing how ridiculous
  750. 26:58this is. If you look at figure two, this
  751. 27:00is one of the most incredible
  752. 27:01demonstrations of the efficiency of this
  753. 27:04new model. Here it's showing the flops
  754. 27:06on the y-axis which is like the raw
  755. 27:08computational power required to generate
  756. 27:11a response and the x-axis is the context
  757. 27:14window size or basically how much
  758. 27:16information it can store in its memory
  759. 27:18at once. When DeepSeek increases this
  760. 27:20window from a standard 4,000 tokens all
  761. 27:23the way to a massive 1 million tokens
  762. 27:25which is roughly 700,000 words or like a
  763. 27:28medium-sized code base. You can see that
  764. 27:30for this new 4.1 flash model, the decode
  765. 27:33curve remains almost completely flat. In
  766. 27:35other words, if you feed the AI a tiny
  767. 27:37one-page document versus thousands of
  768. 27:40pages and ask it a question, this chart
  769. 27:42shows that the AI spends roughly the
  770. 27:44same amount of energy per word. This is
  771. 27:46a massive deal. It completely defies the
  772. 27:49laws of AI because normally what we
  773. 27:51would expect is if you scale the
  774. 27:53context, the compute also increases as
  775. 27:55you could see with the previous
  776. 27:56generations of deepseek models. But here
  777. 27:59they've basically flattened the curve.
  778. 28:01And you know the ridiculous thing is not
  779. 28:03only is this super efficient and
  780. 28:05frictionless, but its performance is
  781. 28:07also state-of-the-art, pretty much on
  782. 28:09par with the Frontier models. For
  783. 28:11example, if you look at Deep Su 1.1,
  784. 28:14this scores 74.2,
  785. 28:16which not only beats the other open
  786. 28:18models out there, but if you look at the
  787. 28:20official leaderboard, then it's
  788. 28:21basically on par with GPT6 Astra, which
  789. 28:24also scores 74. Or if you look at
  790. 28:27Cyberjim, this is pretty much
  791. 28:28state-of-the-art. Same with Automation
  792. 28:30Bench. You can see that this freak of a
  793. 28:32model even beats GPT6 Astro Max. If you
  794. 28:36look at this leaderboard by LiveBench,
  795. 28:38then you can see that this new Deepseek
  796. 28:40is the number one ranked open model.
  797. 28:42Same with Val's index, which measures an
  798. 28:44AI's performance across knowledge work
  799. 28:46tasks. You can see that Deepseek 4.1
  800. 28:49Flash is currently ranked number one.
  801. 28:51And look at the insane cost of this. Not
  802. 28:53only is this the most performant model,
  803. 28:56but it's also the cheapest. You can see
  804. 28:58that second and third place are over 20
  805. 29:00times more expensive. Here's another
  806. 29:02chart showing the cost per task. You can
  807. 29:04see that this new DeepSeek model is all
  808. 29:06the way over here, which is like way
  809. 29:08cheaper than the Frontier GPT6 Astra as
  810. 29:12well as the extremely overpriced clawed
  811. 29:14models. If you look at the output speed
  812. 29:16if you use it through their API, again,
  813. 29:18this is insane. This achieves over 200
  814. 29:21tokens per second, which is like four
  815. 29:23times faster than GPT6. The latency, in
  816. 29:26other words, the time to its first
  817. 29:28answer is also the lowest in the
  818. 29:30industry. All right, so we've gone over
  819. 29:32a ton of stuff. The way they designed it
  820. 29:34is just so unexpected and at many times
  821. 29:36counterintuitive, but once you
  822. 29:38understand why they did it, then you'll
  823. 29:40see how brilliant and genius this is.
  824. 29:43And as always, they've open sourced the
  825. 29:45model for you to download locally, so
  826. 29:47you can do whatever you want with it. My
  827. 29:49hat off to the Deep Seek team for
  828. 29:51pulling off the impossible. Once again,
  829. 29:54I mean, the previous DeepSseek V4 was
  830. 29:56already super efficient, but with this
  831. 29:58latest model, they've managed to
  832. 29:59completely redesign the architecture and
  833. 30:02squeeze even more juice out of it. That
  834. 30:04sums up my deep dive on this new Deep
  835. 30:07Seek 4.1 Flash. This is one of the more
  836. 30:09technical papers I reviewed so far on my
  837. 30:11channel. So hopefully I made it easy for
  838. 30:14you to digest. In fact, the paper is
  839. 30:16jam-packed with a ton of additional
  840. 30:18technical details which I didn't have
  841. 30:20time to cover. So if you're interested
  842. 30:21in digging deeper, I'll link to this
  843. 30:24original paper in the description below
  844. 30:26as well. Let me know in the comments
  845. 30:27what you think of this. As always, I
  846. 30:29will be on the lookout for the top AI
  847. 30:32news and tools to share with you. So, if
  848. 30:34you enjoyed this video, remember to
  849. 30:36like, share, subscribe, and stay tuned
  850. 30:38for more content. Also, there's just so
  851. 30:41much happening in the world of AI every
  852. 30:43week, I can't possibly cover everything
  853. 30:45on my YouTube channel. So, to really
  854. 30:47stay up to date with all that's going on
  855. 30:50in AI, be sure to subscribe to my free
  856. 30:53weekly newsletter. The link to that will
  857. 30:55be in the description below. Thanks for
  858. 30:57watching, and I'll see you in the next
  859. 30:58one.

About this transcript

This page contains the full transcript of Deepseek just did the impossible by AI Search, generated from the public captions YouTube serves with the video. The transcript has 5,814 words across 859 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.