YouTube2Text

Delta Lake Masterclass | Azure Databricks | PySpark | From Zero-To-Expert — Transcript

by Afaque Ahmad · 48,542 words · 7,157 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Hey everyone, welcome to this 6-h hour
  2. 0:03master class on Delta League. In this
  3. 0:06video, I'll be covering the concepts
  4. 0:08deep internals how things work under the
  5. 0:11hood and together with this we will be
  6. 0:14doing a lot of labs, lot of practicals
  7. 0:17to see the concepts we've studied in
  8. 0:19action on Azure data bricks, right? And
  9. 0:22trust me, this is going to be the most
  10. 0:24comprehensive video you will ever watch
  11. 0:27on Delta Lake, right? So before I dive
  12. 0:30in, I want to give you a quick summary
  13. 0:32of the topics that I'll be covering and
  14. 0:35please watch it in order. So we'll be
  15. 0:38first starting with the problems with
  16. 0:40data lake and how Delta Lake solved the
  17. 0:43problem. Some examples of it is lack of
  18. 0:46asset support, lack of update, merge and
  19. 0:49delete operations and data reliability
  20. 0:52and quality issues. How delta solved all
  21. 0:55of these issues. Right? So then I'll
  22. 0:57walk you through the lab setup, the lab
  23. 1:00architecture and we'll be doing the lab
  24. 1:02setup together on Azour right and then
  25. 1:06we'll be performing DML operations
  26. 1:09operations in order to get a feel of how
  27. 1:12is it to work with data link right we'll
  28. 1:15then be uncovering the delta log right
  29. 1:19so all of this all of the magic happens
  30. 1:22through the delta log so we'll be
  31. 1:24talking about delta log in a lot of
  32. 1:26details and how things work under the
  33. 1:28hood in the delta lock. Right. Next up,
  34. 1:31we'll be covering concurrency control,
  35. 1:33optimistic concurrency control, and
  36. 1:35pessimistic concurrency control, time
  37. 1:38travel and versioning, schema
  38. 1:39validation, schema evolution, converting
  39. 1:42paret to delta, manage and external
  40. 1:46tables, and one of the most important
  41. 1:48topics deletion vectors. What exactly is
  42. 1:51copy on write and merge on read? Then
  43. 1:55we'll be covering cloning. What exactly
  44. 1:58is shallow clone and deep clone. Next
  45. 2:01we'll be talking about the small file
  46. 2:03problem. The very popular small file
  47. 2:05problem and the root causes of the small
  48. 2:08file problem. Lastly we'll be closing
  49. 2:10this with the optimization techniques.
  50. 2:14Right? So we'll be covering optimize
  51. 2:17vacuum zorder
  52. 2:19and liquid clustering. Yeah. So I hope
  53. 2:22you're as excited as I am. So let's get
  54. 2:26started. So first of all, let's start by
  55. 2:28understanding why does Delta Lake even
  56. 2:30exist today, right? Why did it even come
  57. 2:32into picture when we already had data
  58. 2:35lake, right? And we are going to do this
  59. 2:37by simply analyzing the pros and cons of
  60. 2:40data lake, right? So some of the things
  61. 2:43that data lake was really good at was
  62. 2:45that it it provided flexible data
  63. 2:48storage. So if you were to go to uh a
  64. 2:50data warehouse, you could only work with
  65. 2:53relational data. You can work with data
  66. 2:56that has columns and rows, right? But
  67. 2:59here the data storage was very flexible.
  68. 3:01You could ingest structured,
  69. 3:03semistructured, unstructured data. And
  70. 3:06this can be audio, video, images, logs
  71. 3:10or basically anything that you wish,
  72. 3:12right?
  73. 3:13and data generated from sensors, your
  74. 3:15web logs, all of this could be streamed
  75. 3:18into your data lakeink. And all of this
  76. 3:20could be stored cheaply at a low cost,
  77. 3:23right? In inside of something like an S3
  78. 3:26or an ADLS, yeah, and it scaled very
  79. 3:30well in order to accommodate high volume
  80. 3:32of data, right? You could basically put
  81. 3:35in infinite amount of data and this
  82. 3:38served a lot of use cases very well. And
  83. 3:41some of those examples are big data,
  84. 3:43machine learning and developing AI
  85. 3:45application, right? But with all of
  86. 3:48this, data lake still came with a lot of
  87. 3:51challenges. It came with several
  88. 3:53challenges and that was the reason Delta
  89. 3:55Lake was designed in order to solve
  90. 3:58these problems and we are going to go
  91. 4:00into a lot of details. We are going to
  92. 4:02discuss those problems in a lot of
  93. 4:04detail because they are going to form
  94. 4:07the foundations. Right? The first
  95. 4:08problem was lack of acid support. Now
  96. 4:13what is acid and why is acid transaction
  97. 4:17support so important. Right? So as I
  98. 4:20mentioned to you before all
  99. 4:21understanding all of these concepts are
  100. 4:23really important because they build your
  101. 4:25foundation. So I'm going to take a lot
  102. 4:27of code examples and I'll be walking you
  103. 4:29through all of these. Right? So let's
  104. 4:32get started. So what does acid stand
  105. 4:34for? What does it mean? So acid
  106. 4:37basically means atomic
  107. 4:42consistent
  108. 4:46isolation
  109. 4:49and the last one is durability.
  110. 4:52Right?
  111. 4:56So we are going to first study about
  112. 4:59atomic. What does atomicity mean? Right?
  113. 5:02So atomicity simply means all or
  114. 5:05nothing. Either you do all of it or you
  115. 5:08don't do anything. Right? Now how does
  116. 5:10that apply to our case? Now let's take a
  117. 5:12small pseudo code. Right? Let's take a
  118. 5:15small pseudo code and try to understand.
  119. 5:17So let's say
  120. 5:19you were transferring some amount from
  121. 5:23your savings
  122. 5:25to your checking account,
  123. 5:29right? From your saving to your checking
  124. 5:30out account. Now, your savings account
  125. 5:33had $1,000
  126. 5:35and your checking account had $200,
  127. 5:38right? And you wanted to transfer $500
  128. 5:43from your savings to your checking
  129. 5:45account, right? Now, what does that
  130. 5:47mean? That means that this guy should
  131. 5:49end up having 500 and this guy should
  132. 5:51end up having 700. Right? So, let's have
  133. 5:54a look at the pseudo code for that.
  134. 5:55Right? So, we basically read in the
  135. 5:57current savings which has $1,000. We
  136. 6:00then read in current checking which has
  137. 6:01$200. Right? Now we deduct from saving
  138. 6:05because we want to transfer $500. Right?
  139. 6:08So what we say is that new savings it
  140. 6:10become current savings which is $1,000
  141. 6:13minus 2 sorry this is going to be minus
  142. 6:16$500.
  143. 6:18Yeah.
  144. 6:20And this is going to be 500 in turn.
  145. 6:24Now my account is updated. Right. So the
  146. 6:26savings account is now updated. Now what
  147. 6:29happens is my system crashes.
  148. 6:32The checking account I've not been able
  149. 6:35to update it so far, right? The system
  150. 6:37crashes.
  151. 6:39So what does this mean? That this
  152. 6:41transaction is lost. It's gone as a
  153. 6:44whole, right? So what you end up having
  154. 6:46is that your savings account has $500.
  155. 6:51Your savings account has $500 and your
  156. 6:54checking account has $200. That mean
  157. 6:56that $500 went up in the air and it
  158. 6:59never went back to your checking
  159. 7:01account. And this is a situation that
  160. 7:03you never want to be in, right? Because
  161. 7:05your system failed somewhere over here.
  162. 7:10And because of that changes were not
  163. 7:12reflected in your checking account,
  164. 7:14right? Now let's see how does an atomic
  165. 7:16system look like. So what would happen
  166. 7:19in an atomic system is that you would
  167. 7:21first begin the transaction. Yeah. Now
  168. 7:24once the transaction has begun you
  169. 7:27deduct the amount from savings which is
  170. 7:29500 is deducted and the total now
  171. 7:32becomes 500. Yeah. Now it is added to
  172. 7:36checking. So the initial balance was
  173. 7:39200. You then add 500 which becomes 700.
  174. 7:42Right. And then you finally commit the
  175. 7:45transaction. Now in between any of these
  176. 7:47steps if my system crashes it is simply
  177. 7:52going to go to the accept clause and it
  178. 7:54is going to roll back my transaction.
  179. 7:57Yeah. So either the whole transaction
  180. 8:00happened as a whole both the accounts
  181. 8:03represent the actual amounts or nothing
  182. 8:06happened. The whole transaction is going
  183. 8:08to be rolled back. Right? And this is
  184. 8:10what atomicity is either everything
  185. 8:13happens or nothing happen. Consistency
  186. 8:15basically means that rules must be
  187. 8:18enforced. A transaction must bring a
  188. 8:21database from one valid state to another
  189. 8:24valid state. Yeah. So let's understand
  190. 8:26that with an example. So let's say there
  191. 8:29is an account which has $100 in place.
  192. 8:34So this account basically has $100 in
  193. 8:36place and then there are two
  194. 8:38simultaneous transaction going on here.
  195. 8:41So you see that there is there are two
  196. 8:43purchases going on. One is worth $80,
  197. 8:46the other purchase is worth $60. Right?
  198. 8:49So what these two transaction do is that
  199. 8:52they basically read in the database at
  200. 8:55the same time. The first one reads in
  201. 8:57$100. The second one also reads in $100.
  202. 9:00Yeah. So they basically go through this
  203. 9:02line of code. They basically read in the
  204. 9:04balance. Now they want to find out what
  205. 9:06is going to be the new balance. Yeah. So
  206. 9:08they subtract
  207. 9:10$80.
  208. 9:12and this guy subtract $60 which is what
  209. 9:15we write over here balance minus amount.
  210. 9:17Yeah. So this is going to be $20 and
  211. 9:20this is going to be $40.
  212. 9:23Now when both of these transactions are
  213. 9:25running there is going to be a
  214. 9:27transaction which is going to make an
  215. 9:30update to the database first. Yeah. Now
  216. 9:32let's assume that this transaction the
  217. 9:35first one the $80 one is the one that
  218. 9:38makes the update first. So that means
  219. 9:40that $20
  220. 9:43is the updated balance. Yeah. Now the
  221. 9:46moment the database was updated with
  222. 9:49$20, the second transaction goes through
  223. 9:52and then it updates it with $40. Yeah.
  224. 9:55So basically run runs this piece of
  225. 9:57code. It updates it with $40. Now what
  226. 10:01fundamentally is wrong here is that you
  227. 10:04made a purchase of 80 + 60 which is $140
  228. 10:08but you only had a balance of $100.
  229. 10:12So this is something that brings your
  230. 10:15database in an inconsistent state. Okay.
  231. 10:18So we have the same account which has
  232. 10:20$100 in place and then there are two
  233. 10:22simultaneous transaction going on which
  234. 10:24is one is trying to make a purchase of
  235. 10:26$80 the other one is trying to make a
  236. 10:28purchase of $60, right? And let's say
  237. 10:31this transaction first starts and it
  238. 10:33comes over here right it basically does
  239. 10:36a begin transaction which simply mean
  240. 10:38that okay I'm going to make changes to
  241. 10:42the values of the database right so now
  242. 10:45when the second purchase tries to come
  243. 10:48in over here what it basically tells him
  244. 10:51that hey I'm going to make some changes
  245. 10:54so you will have to wait for me right
  246. 10:56yeah so last time what happened was that
  247. 10:58Both the transactions read in $100. Now
  248. 11:02the difference is that the first one
  249. 11:04reads in $100. The second one is
  250. 11:07basically for it has still not read in
  251. 11:11anything. And now the first transaction
  252. 11:13goes ahead and executes what it wants
  253. 11:15to. Yeah. So it basically finds in the
  254. 11:17balance which is $100. Now it basically
  255. 11:20we've put a check in place, right? We
  256. 11:23put a rule in place which basically
  257. 11:24checks whether do I have that kind of
  258. 11:27amount or not. Yeah. So it basically
  259. 11:30check that and okay we have that amount.
  260. 11:33So it goes ahead and updates the balance
  261. 11:36which is 100 minus $80 which is $20 and
  262. 11:40$20 is updated and committed. Now, if at
  263. 11:43all if you didn't have $80 within your
  264. 11:47account, it would come here and then it
  265. 11:49would return to you that you have
  266. 11:51insufficient fund. Yeah. And your
  267. 11:54transaction wouldn't take place. Or if
  268. 11:59any of these operation failed over here,
  269. 12:01it would simply go to this accept clause
  270. 12:04and then it would roll back the
  271. 12:06transaction. Yeah. So nothing would take
  272. 12:08place. Yeah. So now after this
  273. 12:12transaction, we have simply updated the
  274. 12:15account balance to $20. Yeah. So now let
  275. 12:19me quickly actually erase this so that I
  276. 12:21can walk you through the second
  277. 12:22transaction.
  278. 12:24Now when the second transaction comes
  279. 12:26in, it goes through all of this. It
  280. 12:28finds the balance. Now the balance is
  281. 12:31$20. And this statement doesn't go
  282. 12:34through the balance which is $20 greater
  283. 12:37than equal to $60 which is not true. So
  284. 12:40this simply returns that you have
  285. 12:42insufficient one. Yeah. So basically
  286. 12:45what this means is that every
  287. 12:47transaction that we did both the first
  288. 12:50one and the second one it brought the
  289. 12:53database from one valid state to another
  290. 12:55valid state. The database didn't end up
  291. 12:58in an inconsistent or an invalid state.
  292. 13:01And this is what consistency is all
  293. 13:03about. The key difference between the
  294. 13:06two examples is that in the second one,
  295. 13:08we use transactions
  296. 13:11to ensure consistency. Yeah. We put in
  297. 13:14appropriate rules in place to ensure
  298. 13:17that the relevant amount was there
  299. 13:20within the bank account. Yeah. And
  300. 13:23finally, we roll back anything if
  301. 13:25anything goes wrong. We roll back the
  302. 13:27transaction if anything goes wrong. So
  303. 13:29isolation simply is the no interference
  304. 13:32policy. Yep. So multiple transactions in
  305. 13:35a database can go on without interfering
  306. 13:38with each other. Each transaction is
  307. 13:41going to operate as if it is the only
  308. 13:43one running even if other transactions
  309. 13:46may be running simultaneously. And this
  310. 13:49means that the intermediate states or
  311. 13:52the uncommitted changes of one
  312. 13:54transaction are not visible to the other
  313. 13:57transaction. So what that means is that
  314. 13:59let's say there is a table t and this is
  315. 14:03in its current state v_sub_1. Yeah. Now
  316. 14:06there is a transaction which basically
  317. 14:08read in this table and then it is going
  318. 14:11on and it is going to make some changes
  319. 14:13and then it is going to finally end up
  320. 14:16in a state called v2. It is going to end
  321. 14:19up but it hasn't ended up yet. Yeah. So
  322. 14:23these are some of the steps that it is
  323. 14:25going to undertake in order to end up in
  324. 14:28a state called V2. Yeah. And all of
  325. 14:31these changes that you see over here are
  326. 14:33uncommitted right now. Now if another
  327. 14:36person, another transaction comes in and
  328. 14:39if they want to read this table, if they
  329. 14:43want to read this table T, it is going
  330. 14:46to read the V1 state of it. Yeah.
  331. 14:50because all of the uncommitted changes
  332. 14:53it is not aware about. So basically you
  333. 14:56see that both of these transactions
  334. 14:59operate independent of each other
  335. 15:01without worrying about each other. And
  336. 15:03the way Delta achieved this is by
  337. 15:06something called optimistic
  338. 15:09concurrency control.
  339. 15:13Optimistic concurrency control. And I'm
  340. 15:16not going to uh overwhelm you with a lot
  341. 15:19of details right now, but we are going
  342. 15:21to discuss how Delta achieves isolation
  343. 15:26using optimistic concurrency control in
  344. 15:28a lot of detail going ahead. Durability
  345. 15:30simply means that once a transaction is
  346. 15:34recorded in a system, it is going to
  347. 15:36stay there. Yeah, it is going to stay
  348. 15:39there irrespective of a system crash, a
  349. 15:42power outage or a failure or anything.
  350. 15:45Right. So, consider receiving your
  351. 15:47paycheck. Yeah. So, so let's say uh your
  352. 15:50current account had $1,000 in place.
  353. 15:54Yeah. It had $1,000 in place. This
  354. 15:57current balance basically reads $1,000.
  355. 16:01And then you got your paycheck which is
  356. 16:02worth $2,000. Yeah. Now the new balance
  357. 16:06is going to be 2,000 +,000 which is
  358. 16:09$3,000
  359. 16:11and finally your account is going to be
  360. 16:14updated. So your account should be
  361. 16:16updated with $3,000 after you've
  362. 16:19received your paycheck. Now imagine that
  363. 16:22let's say some database in some region
  364. 16:25goes down and this transaction goes for
  365. 16:30a toss. Now when the system is restored
  366. 16:34the old backup is taken and it is
  367. 16:37restored with that backup. Now that
  368. 16:39backup doesn't have your paycheck in
  369. 16:42place. Right? So you worked the whole
  370. 16:44month now your paycheck is gone just
  371. 16:46because of some system crash which which
  372. 16:49didn't follow the durability principle.
  373. 16:52So your paycheck is lost. That means the
  374. 16:55old backup now shows a $1,000
  375. 16:59account balance. So this is a durability
  376. 17:02problem. So how would a system with
  377. 17:05durability in place look like?
  378. 17:09So there's a quite a bit of code here.
  379. 17:12But again all of them are very simple
  380. 17:15easy sudo code. So again whenever
  381. 17:18somebody's going ahead and making a
  382. 17:21change we are going to start a
  383. 17:23transaction we are going to first write
  384. 17:26to a transaction log. Yeah. Now think of
  385. 17:29it as uh maintaining a ledger of what
  386. 17:33all is going on. Yeah. You basically
  387. 17:35keep on writing whatever is going on. So
  388. 17:37basically we say that we are going to
  389. 17:39start a deposit transaction. Yeah. And
  390. 17:42we also say what are the details of that
  391. 17:45transaction. So we write to the
  392. 17:46transaction log that okay we are going
  393. 17:48to deposit an amount to a particular
  394. 17:51account ID. And then we basically
  395. 17:53calculate the current balance. We
  396. 17:55calculate the new balance. We add the
  397. 17:57amount and then we update the balance.
  398. 17:59Yeah. Now once the balance is updated,
  399. 18:01we finally write it to the transaction
  400. 18:04log and this has the word commit and
  401. 18:08then we finally commit the transaction.
  402. 18:10Yeah. Now if anything fails over here
  403. 18:12inside the try block, the transaction is
  404. 18:16rolled back. Now you may of course have
  405. 18:19questioned what if
  406. 18:22this whole thing either failed over here
  407. 18:25or here or here. So let's say we were
  408. 18:28writing to the transaction log and then
  409. 18:30it failed over here. Now when the system
  410. 18:32restarts the only thing that it it's
  411. 18:35going to check the the transaction log.
  412. 18:37Yeah. What it is going to see is that
  413. 18:39okay there is something like a start.
  414. 18:41Okay. Let me just change the color. uh
  415. 18:43there is something like a start
  416. 18:48but then there's nothing after that.
  417. 18:50Yeah. So there are no details that means
  418. 18:52we have to roll back this transaction.
  419. 18:54Now the second case what if it failed
  420. 18:56over here over here. Yeah. So it has
  421. 18:59something like a start and then it has
  422. 19:02some detail
  423. 19:05but again it doesn't have a commit. That
  424. 19:08means this transaction wasn't committed.
  425. 19:10it didn't actually go into the database.
  426. 19:13So again, this transaction is going to
  427. 19:15be rolled back even if it fails over
  428. 19:17here. Now all of this goes on and even
  429. 19:21if it fails over here then also it going
  430. 19:22it's going to find the same start and
  431. 19:24detail inside of the transaction log.
  432. 19:26That means this transaction is going to
  433. 19:28be rolled back. Whatever changes were
  434. 19:30made is going to be rolled back. Now if
  435. 19:33it finally fails after this when you
  436. 19:36have the word commit
  437. 19:38in the transaction log that means that
  438. 19:41the data was written inside of the
  439. 19:45transaction log. Yeah. So it the
  440. 19:48database is then going to finally commit
  441. 19:50the transaction and it is going to be
  442. 19:52recorded in the database. Yeah. So this
  443. 19:55is how durability would look like and
  444. 19:58even in cases of failures you would see
  445. 20:00that your database would be in a
  446. 20:02consistent durable state. Yeah. So with
  447. 20:06all of these examples I hope you
  448. 20:08understand how important asset
  449. 20:11properties are and unfortunately data
  450. 20:13lake doesn't have any of these
  451. 20:15properties. The second problem is the
  452. 20:18lack of support for update, merge, and
  453. 20:21delete. And we all know that these are
  454. 20:23really important operations, something
  455. 20:25that we do on a day-to-day basis, right?
  456. 20:27And the problem with traditional data
  457. 20:29lakes is that the data stored is in
  458. 20:33immutable files, right? Something that
  459. 20:35cannot be changed. Those files cannot be
  460. 20:37changed. And this creates a problem when
  461. 20:40you want to update or delete that data.
  462. 20:43Yeah. So let's take an example. Let's
  463. 20:45say you're running an e-commerce company
  464. 20:46and then you have a list of customers
  465. 20:48and then you want to update the customer
  466. 20:52addresses for some of the customers.
  467. 20:54Yeah. So let's say that first of all
  468. 20:57your data is written to this location
  469. 20:59and now you want to update customer
  470. 21:02address
  471. 21:04at this location. Yeah. So the steps
  472. 21:07that you need to follow in order to do
  473. 21:09that is first of all read all of that
  474. 21:11data. Yeah. and then you need to read
  475. 21:15and load it into the memory, make all
  476. 21:17the changes and then finally write it
  477. 21:20back. So these are the three steps that
  478. 21:22you need to do in order to update
  479. 21:24customer address. Now imagine
  480. 21:27if your system fails at any of these
  481. 21:30steps, you are left with inconsistent
  482. 21:34data.
  483. 21:37You are left with inconsistent data.
  484. 21:39Yeah, there also may be a part
  485. 21:42possibility of partial rights
  486. 21:47at this location.
  487. 21:49Partially written data at this location
  488. 21:51which basically mean that your data is
  489. 21:53corrupt. Now if somebody basically reads
  490. 21:56the data at this location, they are
  491. 21:59going to be basically reading corrupt
  492. 22:00data. Third problem is data reliability
  493. 22:03and quality issues, right? And a prime
  494. 22:06example of that is no schema
  495. 22:08enforcement. Yeah. So let's say you're
  496. 22:11collecting customer signup data and
  497. 22:13today you get a record which looks
  498. 22:15something like this. Yeah. Which have
  499. 22:18the name, phone number, email and phone
  500. 22:20number. Yeah. Now tomorrow let's say you
  501. 22:23get a record which looks something like
  502. 22:25this which has the full name. The the
  503. 22:29key basically looks completely
  504. 22:30different. full name, contact and then
  505. 22:33this is nested inside and then you have
  506. 22:36an email and phone number. So the format
  507. 22:39is completely different from the first
  508. 22:41one. Yeah. Now imagine you may end up
  509. 22:44having several of such formats within
  510. 22:46your data lake without schema
  511. 22:48enforcement. And when you're writing a
  512. 22:51query to basically let's say find um the
  513. 22:54customer email or the phone number, you
  514. 22:57may end up writing complex code in order
  515. 23:00to figure out what the right schema is,
  516. 23:03right? You may need to write complex
  517. 23:05logic in order to handle different
  518. 23:07fields and this is a very big issue
  519. 23:10because there's no consistency in your
  520. 23:12data. There's no schema in placement. So
  521. 23:15to quickly summarize the three problem
  522. 23:17that we discussed about data lake was
  523. 23:19number one lack of acid transaction
  524. 23:22support. Number two lack of support for
  525. 23:24update merge and deletes. Number three
  526. 23:27data quality and reliability issues for
  527. 23:30example no schema enforcement. Right? So
  528. 23:32Delta solves all of these problems quite
  529. 23:35beautifully. It just doesn't solve them.
  530. 23:37But it also comes in with a bunch of
  531. 23:40very interesting features like asset
  532. 23:42transactions, time travel, unified batch
  533. 23:45and streaming schema evolution and
  534. 23:47enforcement and it also helps you see
  535. 23:49the audit history and all of that.
  536. 23:51Right? So we going to be going through
  537. 23:53all of this in detail. So before we get
  538. 23:55into understanding how Delta solves all
  539. 23:58of these problems, right, the one that
  540. 24:00we just talked about, let's first get a
  541. 24:02flavor of Delta Lake. How does it look
  542. 24:05like? How does it operate? What are the
  543. 24:08kind of operation that we can perform?
  544. 24:10Right. Yeah. So, we going to be setting
  545. 24:12up a lab environment and for this we are
  546. 24:14going to be using Azure. Yeah. And we'll
  547. 24:16be setting up different kind of
  548. 24:18services. So, don't worry if you don't
  549. 24:20know anything about it. I'll walk you
  550. 24:21through the entire process. Yeah. So,
  551. 24:24first of all, we'll be creating uh a
  552. 24:27workspace using the Azure data bricks
  553. 24:30survey. Yeah. So first of all, we'll be
  554. 24:32creating a workspace
  555. 24:34and we need some place where we can
  556. 24:38store our databases, our tables and all
  557. 24:41of that, right? And for that we are
  558. 24:43going to be using the Unity catalog.
  559. 24:45Yeah. And for those of you who don't
  560. 24:46know what Unity catalog is, simple for
  561. 24:48now you can think of it as a place where
  562. 24:51all of your databases and tables will be
  563. 24:54stored. Now in order to store those
  564. 24:57tables and databases right you need some
  565. 24:59location you need some storage location
  566. 25:01and that is where the meta store comes
  567. 25:03into picture. So the meta store you can
  568. 25:06basically think of it as an object
  569. 25:07storage something like um S3 or ADLS.
  570. 25:12Yeah. So in this context because we are
  571. 25:15using Azure we'll be using ADLS and
  572. 25:19we'll be creating a storage account.
  573. 25:23We'll be creating a storage account
  574. 25:25wherein we will create a container.
  575. 25:29Yeah. So we'll be creating this
  576. 25:30container
  577. 25:32where all of our data will be stored and
  578. 25:34we'll name the container metas store.
  579. 25:36You can name anything but we'll just
  580. 25:37name it metas store. So this is the
  581. 25:39place where all of those tables and
  582. 25:42databases will be stored. But there
  583. 25:44needs to be a mechanism
  584. 25:47using which the unity catalog here can
  585. 25:50store data in the container meta store
  586. 25:52and that is where another service comes
  587. 25:55into picture which is called Azure
  588. 25:58connector sorry not Azure connector
  589. 26:00access connector for Azure data bricks
  590. 26:03yeah and what's basically going to
  591. 26:05happen is that ADLS is going to tell
  592. 26:08this
  593. 26:10this guy is that hey I'm going to
  594. 26:12authorize you
  595. 26:15to be able to access the data in the
  596. 26:18container meta store. Yeah. So you can
  597. 26:21now
  598. 26:23you can now simply access the data that
  599. 26:27has been stored over here. Yeah. So now
  600. 26:30when we want to access the data in the
  601. 26:32meta store either using the workspace or
  602. 26:35through the unity catalog we are simply
  603. 26:38going to assume the role of this guy.
  604. 26:43Yeah we are simply going to assume the
  605. 26:45role of this guy and in turn this person
  606. 26:48is going to help us access the meta
  607. 26:50store. So we finally end up creating
  608. 26:54the first service which is a datab
  609. 26:56bricks. Then we create a storage
  610. 26:59account. Then we create a container.
  611. 27:02After that we create the access
  612. 27:05connector. And finally this meta store
  613. 27:09needs to know a few things that linking
  614. 27:12needs to happen. Right? So this needs to
  615. 27:14know who is the person who can access
  616. 27:18who can help me with access. Right? And
  617. 27:19it is this guy the access connector for
  618. 27:23your data brick. So it needs to know the
  619. 27:24resource ID
  620. 27:26of the access connector and the location
  621. 27:30to which it needs access right
  622. 27:33which is going to be the meta store. So
  623. 27:36this is the path and this is going to be
  624. 27:38the meta store container path right. So
  625. 27:43number five we need to perform the
  626. 27:45linking right. So once all of these
  627. 27:47steps are done we basically set up our
  628. 27:49lab environment and then we can start
  629. 27:51working. Okay, so let's quickly go ahead
  630. 27:54and create the Azure datab bricks
  631. 27:56workspace. And for that I'm going to
  632. 27:58click on create and we going to create a
  633. 28:02new resource group because this is where
  634. 28:04our workspace the storage account the
  635. 28:07access connector all of them are going
  636. 28:09to reside. Yeah, I'm going to name it as
  637. 28:12datab bricks delta lab - rg and I'm
  638. 28:16going to follow a similar naming
  639. 28:18convention. This is WS. Uh the region is
  640. 28:21going to be Australia East because I
  641. 28:23have a lot of resources in other regions
  642. 28:25as well. I've created workspaces. So to
  643. 28:29make sure that they don't conflict, I'm
  644. 28:30going to choose Australia East. But you
  645. 28:32please go ahead and choose the region
  646. 28:34that is the closest to you. Yeah. So it
  647. 28:38is validating some stuff. So meanwhile
  648. 28:41we can go ahead and create the storage
  649. 28:44account.
  650. 28:47The storage account is again going to be
  651. 28:50in the same in the same resource group.
  652. 28:54The naming convention also we'll keep
  653. 28:57we'll keep it very similar. Delta lab
  654. 29:00storage account is going to be in
  655. 29:03Australia east. This is going to be ADLS
  656. 29:05gen 2. And for now we'll select the
  657. 29:08cheapest option which is locally
  658. 29:10redundant storage. And don't forget to
  659. 29:13put a check here because we want to use
  660. 29:15data lake storage gen 2. And we simply
  661. 29:19go ahead and click on review and create.
  662. 29:22Let's also click on create over here.
  663. 29:24And create over here. The last item that
  664. 29:27we had was the access connector for
  665. 29:32Azure datab bricks and this is the
  666. 29:35person who will be assigned who will be
  667. 29:39given the authority to be able to access
  668. 29:42data in the storage account.
  669. 29:44So
  670. 29:46let's select the same resource group
  671. 29:50DB delta lab - RG and this is going to
  672. 29:53be DB delta lab
  673. 29:56access connector. Yeah. And this is
  674. 29:59going to be in Australia east. So let's
  675. 30:01go ahead and create this.
  676. 30:09So we see now the deployment for the
  677. 30:11storage account is completed. So inside
  678. 30:14of the storage account, we need to
  679. 30:16create a container called metas store.
  680. 30:20Let's go ahead and quickly create that.
  681. 30:24So there you go. You have a container
  682. 30:26called metas store. And in this
  683. 30:28container, Unity catalog is going to
  684. 30:30store all of the data.
  685. 30:32Now the next part to this is we have to
  686. 30:36give permissions to the access connector
  687. 30:39to be able to access the data in my
  688. 30:41account in this storage account. Yeah.
  689. 30:44So what this guy is going to do is that
  690. 30:46it is going to give permissions now and
  691. 30:48the role that is going to that is going
  692. 30:51to allow is storage blob data
  693. 30:53contributor
  694. 30:55and because it's a manage identity we
  695. 30:57are simply going to select
  696. 30:59the access connector for Azure data
  697. 31:01bricks and this was the one that we just
  698. 31:03created right now and we simply click on
  699. 31:05review and assign.
  700. 31:11So now we see that the assignment is
  701. 31:13there in place. That means that the
  702. 31:15access connector will be able to access
  703. 31:18the data in the storage account.
  704. 31:23Now we are waiting for the deployment of
  705. 31:26the workspace to take place to complete.
  706. 31:29Right. Okay. So the workspace is now
  707. 31:32deployed. Now the Azure data bricks
  708. 31:34workspace is deployed. We created ADLS
  709. 31:37Gen 2. We created the metas store. We
  710. 31:41also created the access connector.
  711. 31:44Right? Now in order to enable unity
  712. 31:47catalog. Now we need to
  713. 31:50do the linking. We need to create the
  714. 31:53the meta store and then provide
  715. 31:56who is going to be the person who's
  716. 31:57going to help me with access and what is
  717. 32:00the path of the meta store. Yeah. So
  718. 32:03let's go ahead and do this last step
  719. 32:05right here. So let's quickly go ahead to
  720. 32:07the workspace. We are going to launch
  721. 32:10the workspace. And
  722. 32:17when we head over to the catalog, we see
  723. 32:19that
  724. 32:21it only has the legacy hive meta store
  725. 32:24and some shared samples over here. The
  726. 32:27Unity catalog hasn't yet been enabled.
  727. 32:30So in order to enable Unity catalog, we
  728. 32:32need to go to account.asure Azure datab
  729. 32:36bricks dot and we login with this. So
  730. 32:39now we'll be able to see that we have an
  731. 32:42option for the catalog and there is an
  732. 32:44option to create the meta store. This is
  733. 32:47going to be db delta lab meta store.
  734. 32:52This is going to be in Australia east.
  735. 32:55And the format of this is the container
  736. 32:59name at storage
  737. 33:00account.dfs.co.windows.net.
  738. 33:02So the container name is meta store.
  739. 33:06The storage account name is this one. So
  740. 33:10we simply do this
  741. 33:13dot windows
  742. 33:15dfs.core
  743. 33:18dotwind.net.
  744. 33:20Yeah. And then finally we need to put in
  745. 33:22the access connector ID. So the access
  746. 33:26connector we can find it over here. This
  747. 33:28was the one that we created. And then
  748. 33:32here is the resource ID. So we simply
  749. 33:34copy it from here and we are going to
  750. 33:36paste it over here and we click on
  751. 33:38create.
  752. 33:45So now it is asking me to which
  753. 33:49workspace do I want to assign this metas
  754. 33:52store and this is the workspace that we
  755. 33:53just created, right? So let's assign the
  756. 33:56meta store to this workspace.
  757. 34:00And now we have Unity catalog enabled.
  758. 34:02So let me quickly go ahead and refresh
  759. 34:04this.
  760. 34:07And you see that apart from the shared
  761. 34:09samples and the legacy hive meta store,
  762. 34:12we have two cataloges right now. So this
  763. 34:14means that a unity catalog is enabled
  764. 34:17and our lab environment is now set up.
  765. 34:19So let's go ahead and create a delta
  766. 34:21table and perform some operations. So
  767. 34:23we'll be coming back to this diagram but
  768. 34:25let me first go over to the workspace
  769. 34:29and then we going to create a folder
  770. 34:31called lab delta lake and here is where
  771. 34:36we are going to store all our notebooks.
  772. 34:38But for our labs we need data right. So
  773. 34:42we are going to go to the storage
  774. 34:44account that we created which is this
  775. 34:46one DB delta lab storage account and we
  776. 34:50are going to create a container right
  777. 34:52because we don't want to me we don't
  778. 34:54want to mess around with the metas store
  779. 34:56container because this is going to be
  780. 34:59used by the unity catalogs to store all
  781. 35:01the databases and tables right so let's
  782. 35:04go ahead and create something called lab
  783. 35:05data and this is going to create a new
  784. 35:10container
  785. 35:11Inside of this, I'm going to add a
  786. 35:14directory called shopping invoices
  787. 35:17or maybe just invoices.
  788. 35:20Let's keep it smaller and shorter. So,
  789. 35:23inside of invoices, I'm going to upload
  790. 35:26all of the files that I have. Right? So,
  791. 35:28I have three files
  792. 35:30and I'm going to upload all of them. So
  793. 35:33these are basically invoices for
  794. 35:34customer ids from 1 to 100, 100 to 200,
  795. 35:38201
  796. 35:39some some large number 99457 right so
  797. 35:42don't worry all of this will be made
  798. 35:45available in the GitHub repository so
  799. 35:47you can download it from there so now
  800. 35:50let's move ahead
  801. 35:53to our lab now an interesting thing is
  802. 35:57that
  803. 36:00okay so we are going to perform DML L
  804. 36:03operations on delta tables. Yeah. So
  805. 36:08this is our motive and we want to access
  806. 36:11the data that is over here inside of the
  807. 36:14lab data container. Right? Now this is
  808. 36:18not going to naturally have access.
  809. 36:21Right? So the way we access it is let's
  810. 36:23say percentage fs ls and abfss.
  811. 36:30This is going to be meta lab data at
  812. 36:35whatever this is. The storage account
  813. 36:37name is this one over here. So we copy
  814. 36:40this dfs.core.windows.net.
  815. 36:44Right. And I have this cluster already
  816. 36:48created. And let me also walk you
  817. 36:50through how to create the cluster over
  818. 36:52here. Right? It's quite simple. So you
  819. 36:55you click on create compute.
  820. 36:58go ahead and create a single node
  821. 36:59cluster because for this lab we don't
  822. 37:01need a multi-node cluster and you don't
  823. 37:03want to incur a lot of cost right so
  824. 37:05let's simply choose 14.3
  825. 37:08LTS we don't want photon accelization
  826. 37:12and let's choose a simpler one a uh
  827. 37:16something that are lesser memory than 16
  828. 37:18which is over here 14 GB and four cores
  829. 37:21and set this to 20 minutes you can even
  830. 37:24set it to 10 minutes because sometimes
  831. 37:26we just leave the cluster running and we
  832. 37:29incur a lot of cost right and then you
  833. 37:30can click on create compute so that's
  834. 37:33how you can create the compute and now
  835. 37:35let's come back here now let's say if I
  836. 37:37want to do an ls
  837. 37:41let's see what do I get okay so what it
  838. 37:44says is that invalid configuration value
  839. 37:47detected for
  840. 37:51fsazure
  841. 37:52account right so basically the crux of
  842. 37:55this is that it is not able able to
  843. 37:57access this data. Right? So that is
  844. 37:59where the concept of external location
  845. 38:02comes in. So external location are
  846. 38:04basically location that are external to
  847. 38:07data bricks that we want to access right
  848. 38:10and we need to put proper measures in
  849. 38:12place so that we are able to access that
  850. 38:14location and let's first see how we can
  851. 38:18access this external location. So we go
  852. 38:20to catalog.
  853. 38:22We then go to external data. We go to
  854. 38:25create an external location. By the way,
  855. 38:27you can also create it using a simple
  856. 38:31SQL statement. But let's go through the
  857. 38:33UI and see how this works out. So let's
  858. 38:36name the external location lab data
  859. 38:40external.
  860. 38:42And the way we do it is abs
  861. 38:50and then this is going to be lab data at
  862. 38:54the storage account named
  863. 38:56dfs.core.windows.net.
  864. 39:00Yeah. So I believe let's quickly check
  865. 39:02this uh the format abfs container name
  866. 39:05storage account dfs.core.windows.net.
  867. 39:07Yeah. So that's the path and then we
  868. 39:10need a storage credential. somebody
  869. 39:13whose role I can assume in order to be
  870. 39:16able to get access in order to be able
  871. 39:18to look at the data at that storage
  872. 39:21location. Yeah. Now when we created the
  873. 39:26access connector it automatically
  874. 39:28creates a storage credential and that is
  875. 39:31what we can use. So we simply go ahead
  876. 39:32and use that and we are going to create
  877. 39:35the external location. Now the external
  878. 39:38location is created. Now let's go ahead
  879. 39:41and run this command once again.
  880. 39:44So now you see that we are able to do an
  881. 39:48ls on that location. Right? So let me
  882. 39:51also quickly do an ls on the invoices
  883. 39:57and we are able to see all of the data
  884. 40:00that we have put on that location.
  885. 40:02Right? So that is external location.
  886. 40:03That is how external locations work.
  887. 40:06Let's go ahead and create a catalog.
  888. 40:09Right? So for those of you who haven't
  889. 40:11used Unity catalog, think of catalog as
  890. 40:14a highlevel container which is going to
  891. 40:17store your databases and the new tables.
  892. 40:20Right? So I'm going to write this SQL
  893. 40:23create
  894. 40:25catalog if not exist and this is going
  895. 40:28to be delta catalog. Right? And after
  896. 40:31this I'm going to create a schema and
  897. 40:33you can think of schema as a database.
  898. 40:36Right? So this is going to be delta
  899. 40:41dot delta db right. So let's go ahead
  900. 40:44and run this.
  901. 40:47So now here I should have a delta
  902. 40:50catalog which is right here and amongst
  903. 40:53the default and the information schemas
  904. 40:55I also have a delta db schema right and
  905. 40:58it doesn't have anything for now. So
  906. 41:00that is okay. Yeah. So we are going to
  907. 41:02quickly look at some of the data that
  908. 41:05we've stored over here and we're going
  909. 41:07to use this for creating and operating
  910. 41:10on our delta lake right so I'm going
  911. 41:12simply going to say select star from
  912. 41:15park k
  913. 41:18and this is going to be something like
  914. 41:20this and let me also do a limit five
  915. 41:24yeah so let's go ahead and run this so
  916. 41:26this is how our park file looks like
  917. 41:29it's basically simple invoices about a
  918. 41:31customer who made a purchase, what is
  919. 41:33the gender, age, payment method, what is
  920. 41:35the quantity, what is the invoice date
  921. 41:37and all of that. Right? So, let's go
  922. 41:39ahead and create a delta table out of
  923. 41:42this. Right? So, let's go ahead and
  924. 41:44simply write create or replace table and
  925. 41:48this is going to be named invoices,
  926. 41:52right? and
  927. 41:55as select
  928. 41:58star from. Okay, so this is helping me a
  929. 42:01lot.
  930. 42:03I just have to press enter more than
  931. 42:05doing the actual typing. So this is good
  932. 42:09and let's go ahead and run this, right?
  933. 42:16Okay, great. The table has been created.
  934. 42:18So let's have a look at the catalog over
  935. 42:20here. And now we see that it contains
  936. 42:23invoices and this invoice is basically
  937. 42:26your all of the data that we just
  938. 42:29discussed and it contains some
  939. 42:32interesting details over here right and
  940. 42:34we're going to come back to this soon
  941. 42:36but before that let's see the history
  942. 42:40and what do I mean by history is that
  943. 42:42what are the operation that was
  944. 42:44performed on this table.
  945. 42:52I was just happy about the about this uh
  946. 42:55about about the code appearing
  947. 42:57automatically and now it's not appearing
  948. 42:59automatically. Okay, so describe history
  949. 43:03cannot be found. Okay, I just misspelled
  950. 43:06this. So we just created the table using
  951. 43:11a cat cas statement, right? And that is
  952. 43:14what you see over here. So it basically
  953. 43:16tells you that there is a time stamp on
  954. 43:19this time stamp this particular user
  955. 43:21created this table and the operation
  956. 43:23that was performed would basically a
  957. 43:25create or replace table as select
  958. 43:28something like that right and then it
  959. 43:29has all other details as well. Yeah. So
  960. 43:33this basically tells us that version
  961. 43:35zero and this is where virgining comes
  962. 43:38in right. This form the foundation and
  963. 43:39the basis for virgining. And don't worry
  964. 43:41we'll we'll we have a complete section
  965. 43:43on time travel and virgining right so
  966. 43:45this basically forms the first version
  967. 43:48version zero of this now let's actually
  968. 43:51also have a look at what is happening
  969. 43:53behind the scenes right
  970. 43:56so we see that this data is stored in
  971. 44:00the container meta store
  972. 44:02in this storage account the storage
  973. 44:05account that we provided and then there
  974. 44:07is some folder which is named by this
  975. 44:09unique long unique unique ID and then
  976. 44:11there is a folder called table and this
  977. 44:14is the unique ID of a table right so
  978. 44:17let's quickly go over here to the
  979. 44:19storage account this meta store
  980. 44:24yeah and
  981. 44:27so now we go to tables and then I copy
  982. 44:31the ID which is this ID right over here
  983. 44:34and we see that there is a park file
  984. 44:37which is created along with a folder
  985. 44:40called data log. So it basically has the
  986. 44:42data and the metadata. The metadata
  987. 44:45about the transactions that took place,
  988. 44:47right? And that is stored in the delta
  989. 44:49log folder. So this delta log folder
  990. 44:52basically contains a JSON file which has
  991. 44:55details about the transactions and a CRC
  992. 44:57file. We can ignore the CRC file for now
  993. 44:59because it's for file valid validation
  994. 45:01and all of that. So we don't need to
  995. 45:02worry about all of that for now. So if
  996. 45:04we click on this, what do we see?
  997. 45:08So we see all of this and let me
  998. 45:11download this file right let me quickly
  999. 45:14download this file and
  1000. 45:17okay and let me format this document. So
  1001. 45:21it contains a bunch of stuff right and
  1002. 45:23don't get intimidated by this. So the
  1003. 45:26first section is commit info. The second
  1004. 45:28section is metadata. The third section
  1005. 45:30is protocol. And then the fourth section
  1006. 45:33is the operations that we performed,
  1007. 45:35right? And the most important one of
  1008. 45:37this is the last section, the operation
  1009. 45:39that we performed. But I'll still walk
  1010. 45:41you through this. The commit info
  1011. 45:42basically contains the operation that we
  1012. 45:44performed, right? The user who performed
  1013. 45:46this operation, right? And all of that.
  1014. 45:49Um now coming over to the add part
  1015. 45:53itself. It is basically the operation
  1016. 45:55that we performed. So we did a create or
  1017. 45:59replace table and this added some data
  1018. 46:01to the table. Right. And that data was
  1019. 46:04added in the park file. It then contains
  1020. 46:07other details. So because our park is
  1021. 46:10not partitioned, it doesn't have any
  1022. 46:11partition value. This is the size. Uh
  1023. 46:14these are the stats which basically
  1024. 46:17contains uh what are the minimum values
  1025. 46:20for the columns like customer ID um if
  1026. 46:22there's a numeric column then it would
  1027. 46:24be helpful something like price and all
  1028. 46:25of that right and it also contains the
  1029. 46:27max values the max values over here so
  1030. 46:30this is very helpful when we want to do
  1031. 46:34partition pruning let's say we are
  1032. 46:35reading this file using spark and we
  1033. 46:37want to do partition pruning so spark
  1034. 46:40can basically filter down the files by
  1035. 46:43reading the min and max right if it
  1036. 46:45doesn't fall within these bound it can
  1037. 46:47simply either include or not include
  1038. 46:50this file so that is what this operation
  1039. 46:54is for and it been noted down in the
  1040. 46:57transaction log now let's go ahead and
  1041. 47:00perform another operation
  1042. 47:04let's go ahead and do an insert right so
  1043. 47:07let's see how the insert statement is
  1044. 47:09going to look like and it's basically
  1045. 47:11going to look something like this insert
  1046. 47:14into delta delta catalog or delta
  1047. 47:17DB.invoices and we are simply going to
  1048. 47:19have a select star from par
  1049. 47:24dot
  1050. 47:27and this one and this is basically going
  1051. 47:30to be one to 100 because we have already
  1052. 47:35inserted from 101 to 200. So let's go
  1053. 47:38ahead and insert the data from 1 to 100.
  1054. 47:41Right? So let's go ahead and run this.
  1055. 47:47And we are also going to see the history
  1056. 47:49again. So let me paste this over here.
  1057. 47:53And let's also do some sanity check.
  1058. 47:55Right? We are going to do a select main
  1059. 48:00customer ID max
  1060. 48:03customer ID and then the count
  1061. 48:07count star. What are the total number of
  1062. 48:09records? Right? And this is simply going
  1063. 48:11to be from this table right. So now we
  1064. 48:13see that the insert operation had
  1065. 48:16succeeded. That means it inserted 100
  1066. 48:18more records. Now let's go ahead and see
  1067. 48:21how the history looks like. The first
  1068. 48:23statement was a create or replace right
  1069. 48:26and the second statement was a write
  1070. 48:29through an insert statement over here.
  1071. 48:31And this creates the second version
  1072. 48:34number one. Right? It's zero index. So
  1073. 48:36that's why the second version. Now let's
  1074. 48:40go ahead and run this. The minimum and
  1075. 48:42the maximum. So we are supposed to have
  1076. 48:44200 record. That's why the minimum is
  1077. 48:46one, maximum is 200. And the count the
  1078. 48:49total number of record is 200. So that
  1079. 48:52means that all of the insert and the
  1080. 48:54create statement has happened correctly.
  1081. 48:57So now let's go ahead and have a look at
  1082. 49:00how the delta log looks like. So earlier
  1083. 49:03there was one parket file. Now there are
  1084. 49:06two park files. We're going to have a
  1085. 49:08look at how these park files look like,
  1086. 49:10right? And let me go ahead and download
  1087. 49:12these two files.
  1088. 49:14Uh so basically I copy this and I'm
  1089. 49:19going to use park tools show and it
  1090. 49:22resides in my downloads folder.
  1091. 49:27So you see that the first park file
  1092. 49:29wherein we inserted data from for
  1093. 49:31customer ids from 101 to 200 that is
  1094. 49:35what is contained over here in the first
  1095. 49:38park file. In the second par file
  1096. 49:42this should contain all of the records
  1097. 49:44for customer ID from 1 to 100. So again
  1098. 49:48let me do park tools show
  1099. 49:53downloads
  1100. 49:55and then the part right. So now you see
  1101. 49:58that this is from customer ID 1 to
  1102. 50:01customer ID 100. Yeah. And let's go
  1103. 50:04ahead and have a look at the delta log.
  1104. 50:07So we again will download
  1105. 50:10the first delta log.
  1106. 50:13I mean the second delta log index by 01
  1107. 50:18and let's open it
  1108. 50:21format this document
  1109. 50:24and we see that this is the second paret
  1110. 50:27file that we were talking about right
  1111. 50:30the second paret file that's been added
  1112. 50:33over here. Yeah this one. So we see that
  1113. 50:36there are two transaction that took
  1114. 50:38place. First one is the create table. It
  1115. 50:40created a par file and then it created
  1116. 50:44it noted that down in a JSON file,
  1117. 50:47right? It noted that transaction down
  1118. 50:49using this add operation. Another
  1119. 50:52operation when that took place using the
  1120. 50:55insert statement, it also created
  1121. 50:57another park file, put in all of the
  1122. 50:59data over there and then again noted it
  1123. 51:01down as an insert operation in another
  1124. 51:05JSON file. Right? So these are taking
  1125. 51:07place at separate independent
  1126. 51:09transaction. Yeah. Now let's try a few
  1127. 51:12other interesting operations right. Uh
  1128. 51:14let's try the update operation. So first
  1129. 51:17of all let me do a select star from the
  1130. 51:20table where customer
  1131. 51:24ID equals 1. Yeah. And let's see how
  1132. 51:28let's see what the quantity. So the
  1133. 51:29quantity is five over here. Yeah. So let
  1134. 51:32me go ahead and run an update statement.
  1135. 51:35update the table and this is going to be
  1136. 51:38set quantity 10 where customer ID equals
  1137. 51:431. Yeah. So let's go ahead and quickly
  1138. 51:46run this
  1139. 51:50and let me also do a describe history of
  1140. 51:53this table. So we should see an update.
  1141. 51:57Great, we see an update and this is
  1142. 51:59marked as the next version. Now let me
  1143. 52:02quickly go ahead and see how
  1144. 52:05this looks like. So the quantity has
  1145. 52:07been updated to 10. That is good. That
  1146. 52:10is what we would have expected. Now
  1147. 52:12let's see what's going on behind the
  1148. 52:14scenes.
  1149. 52:17So behind the scenes
  1150. 52:20we see that that there were earlier two
  1151. 52:22files. Now we see that there are three
  1152. 52:24files. Right? So this is the latest one
  1153. 52:26that has been written at 643.
  1154. 52:30And there is something called a deletion
  1155. 52:33vector. Right? And this is quite
  1156. 52:35interesting to understand. We'll we'll
  1157. 52:37have a look at this in details. And
  1158. 52:39let's also have a look at what this file
  1159. 52:43means.
  1160. 52:44The latest transaction.
  1161. 52:48So the latest transaction. Let me format
  1162. 52:51this document. And we see a lot of
  1163. 52:54things happening, right? We see a
  1164. 52:56remove, we see an add. And then we also
  1165. 53:00see an ad, right? But the fact of the
  1166. 53:02matter is I just updated one row. Why
  1167. 53:06why did it even remove something? It
  1168. 53:08removed one park file, it added another
  1169. 53:12parket file and then it added another
  1170. 53:14parket file again. Right? What were the
  1171. 53:16need for all of this? So there is a
  1172. 53:17little bit of concept behind this and
  1173. 53:19and let's quickly walk through that.
  1174. 53:21Right?
  1175. 53:23So let's assume that there was this file
  1176. 53:26called 1. R K and this basically
  1177. 53:31contained all the records for customer
  1178. 53:33ID 1 to 100. Right now whenever somebody
  1179. 53:37wanted to read all the customers from 1
  1180. 53:39to 100, they would simply go ahead and
  1181. 53:41read this file now the situation has
  1182. 53:44changed a little bit. Right? The way it
  1183. 53:47has changed is that we ran we want to
  1184. 53:49run an update statement and we want to
  1185. 53:53update the quantity to 10 for customer
  1186. 53:58ID.
  1187. 54:00Customer ID equals 1. Yeah. So this is
  1188. 54:03what we want to do. So what delta says
  1189. 54:05is that now anybody want who wants to
  1190. 54:08read all of the data for customer ID
  1191. 54:11from 1 to 100, they just cannot go ahead
  1192. 54:13and read this file anymore, right?
  1193. 54:15because this has changed for customer id
  1194. 54:19equals 1 the quantity has become 10 but
  1195. 54:22this file contains quantity equals five.
  1196. 54:25So what I'm going to do is that I am
  1197. 54:28going to create two files right so we
  1198. 54:31are going to create
  1199. 54:33two files the first one is 1.park park
  1200. 54:36is going to be as it is right and then
  1201. 54:39we are going to create two.par
  1202. 54:43and this is only going to contain the
  1203. 54:45record customer
  1204. 54:48ID
  1205. 54:50equals 1 and it quantity is going to be
  1206. 54:53equal to 10. Yeah. So it created another
  1207. 54:57file with just the updated record and it
  1208. 55:01is going to club 1.par park a with
  1209. 55:05something called a deletion vector.
  1210. 55:12Yeah. So it is going to plug this with
  1211. 55:15something called a deletion vector. And
  1212. 55:17the deletion vector basically tells that
  1213. 55:21don't read
  1214. 55:23the record where customer ID equals 1.
  1215. 55:28Don't read the record where customer ID
  1216. 55:30equals 1 because again that is going to
  1217. 55:32be read from 2.par. So the way the delta
  1218. 55:36transaction log is built the reason why
  1219. 55:39you see basically what do you see? You
  1220. 55:41see one remove. Yeah you see one remove
  1221. 55:46over here
  1222. 55:48and that same file is being added over
  1223. 55:50here. Yeah the first one is a remove.
  1224. 55:53The second one is an add and then this
  1225. 55:56is the other file. So that is what you
  1226. 55:57see over here. So now let's try to
  1227. 56:00relate it with this. So now when we
  1228. 56:04build the when delta builds the
  1229. 56:05transaction log what it says is that
  1230. 56:08this file 1.park earlier it was being
  1231. 56:11treated as if we need to read the whole
  1232. 56:14file. So now we have to treat it
  1233. 56:16differently. You don't have to read the
  1234. 56:18whole file. So that is why it adds a
  1235. 56:21remove saying that 1.park par cannot be
  1236. 56:25treated the same way it was being
  1237. 56:27treated earlier. Yeah. And then it add
  1238. 56:30then add. So the way it has to be
  1239. 56:33treated now is 1.par plus the deletion
  1240. 56:37vector. Yeah. So first it tells that
  1241. 56:39don't treat 1.pk as it were being
  1242. 56:42treated earlier. Yeah. Because it
  1243. 56:44contains the row customer ID equals 1
  1244. 56:46but the quantity is different. It's not
  1245. 56:4810. So don't treat it like that. Treat
  1246. 56:50it like this. So for that it adds an add
  1247. 56:53statement with a deletion vector and
  1248. 56:56then it adds an add again which
  1249. 56:59basically means that read customer id
  1250. 57:02equals 1. Yeah. So it makes it very
  1251. 57:04beautiful. It let's quickly summarize
  1252. 57:07this. So it adds a
  1253. 57:09remove the remove is for 1.park
  1254. 57:14simply saying don't read it as it were
  1255. 57:16being read earlier. You were reading the
  1256. 57:17whole file earlier. Don't do that. And
  1257. 57:20then it adds an add one.park again
  1258. 57:24simply saying that read it with the
  1259. 57:28deletion vector. And then it adds an add
  1260. 57:31again with 2.pk
  1261. 57:34which only contains customer ID
  1262. 57:38equals 1. So using all of this
  1263. 57:42information it is able to construct the
  1264. 57:45file. Yeah. So that's that's the beauty
  1265. 57:48of this. Now let's go back and see what
  1266. 57:50actually happened. So
  1267. 57:53here you see this file
  1268. 57:56this is without the deletion vector and
  1269. 57:58this file is with the deletion vector
  1270. 58:01over here. Yeah. And this file over here
  1271. 58:05is the file which only contains customer
  1272. 58:08ID equals 1. And let's verify that right
  1273. 58:11let's verify that.
  1274. 58:13So this is the file over here. Let me
  1275. 58:16download it. We can do a parket tool
  1276. 58:19show and
  1277. 58:21download
  1278. 58:23and then this park over here. So you see
  1279. 58:25that it only contains customer ID equals
  1280. 58:281 and quantity equal 10. So this is a
  1281. 58:31very beautiful way of
  1282. 58:34making the transaction log without
  1283. 58:36actually doing the writing and the
  1284. 58:39reading and all of that. Right? So when
  1285. 58:42it added a remove, it actually didn't do
  1286. 58:45a remove, right? It just noted it down
  1287. 58:47and logically built up the steps to
  1288. 58:50construct the file. Okay. So let's try a
  1289. 58:52different operation now. Let's try
  1290. 58:54delete. Yeah. So let's do
  1291. 59:00something. So first of all uh let me do
  1292. 59:04something like a
  1293. 59:07where customer ID equals let's say 99.
  1294. 59:10Right? Um, so now
  1295. 59:15I just want to see how the record looks
  1296. 59:17like
  1297. 59:19and the delete statement would look
  1298. 59:21something like delete from this
  1299. 59:25where customer ID equals this, right?
  1300. 59:27Pretty simple. So this is my customer ID
  1301. 59:30equals 99. Let's go ahead and run this.
  1302. 59:32Uh, let me also quickly show you.
  1303. 59:36We have three parket files, one deletion
  1304. 59:38vector, and three JSON files, right? So,
  1305. 59:43let's go ahead and run this now.
  1306. 59:47If I were to run this, I should get an
  1307. 59:50error now. Yeah, as expected.
  1308. 59:55Yeah, I mean, we should get an empty
  1309. 59:58empty results, right? Instead of an
  1310. 1:00:00error. So now we do a describe history
  1311. 1:00:06of this table and I should see a
  1312. 1:00:08statement for delete.
  1313. 1:00:11Great. So now we see an operation for
  1314. 1:00:13delete and this becomes the most recent
  1315. 1:00:16version. So now if I refresh this I
  1316. 1:00:20still have three parket files and I have
  1317. 1:00:23just added a deletion vector. Right? So
  1318. 1:00:27there have been no changes with respect
  1319. 1:00:28to the park file. And let's go ahead and
  1320. 1:00:31see the
  1321. 1:00:33JSON log again. Let me download it. And
  1322. 1:00:37let's go ahead and open this in Visual
  1323. 1:00:39Studio
  1324. 1:00:41uh format document. And what has
  1325. 1:00:44happened?
  1326. 1:00:46So what has happened is there is a
  1327. 1:00:50remove and then there is an add. Yeah,
  1328. 1:00:55quite interesting. So there is a remove
  1329. 1:00:58for a file which already has a deletion
  1330. 1:01:01vector and now there is an add. Yeah. So
  1331. 1:01:06let's quickly go ahead and try to
  1332. 1:01:07understand what's going on. Yeah. So
  1333. 1:01:09last time we were talking about 1.park,
  1334. 1:01:14right? So we were talking about
  1335. 1:01:191.park park file
  1336. 1:01:23and this basically contained all of the
  1337. 1:01:27customer ids from 1 to 100. Right? Now,
  1338. 1:01:30if you remember this was updated with a
  1339. 1:01:33deletion vector. So, 1 dot park file a
  1340. 1:01:38deletion vector was added and this was
  1341. 1:01:42for customer
  1342. 1:01:44ID equals 1. And this was basically to
  1343. 1:01:48tell that hey don't read par file the
  1344. 1:01:521.parket park file as it is right read
  1345. 1:01:55it with the deletion vector because the
  1346. 1:01:58customer ID number one has been updated
  1347. 1:02:01the quantity was earlier five now it's
  1348. 1:02:03updated to 10 and so for that reason you
  1349. 1:02:06shouldn't read it like this you should
  1350. 1:02:08read it like this with a deletion vector
  1351. 1:02:10so that is what was the state so that is
  1352. 1:02:14the reason why you see the remove with a
  1353. 1:02:19deletion vector already and the add also
  1354. 1:02:22has a deletion vector. So these both are
  1355. 1:02:25basically the same files. They are
  1356. 1:02:28basically the same files and both of
  1357. 1:02:31them have deletion vector. It just that
  1358. 1:02:33the deletion vector has been updated.
  1359. 1:02:36Yeah. So basically now what happens is
  1360. 1:02:40we are saying that don't read one.park
  1361. 1:02:44in this condition over here. Read it in
  1362. 1:02:47a new condition because a delete
  1363. 1:02:49operation has happened. Something new
  1364. 1:02:50has happened, right? So a delete has
  1365. 1:02:53happened.
  1366. 1:02:56Yeah. So a delete has happened and in
  1367. 1:02:58this delete
  1368. 1:03:01this park file is supposed to be now
  1369. 1:03:04read with a deletion vector and this is
  1370. 1:03:07going to contain
  1371. 1:03:09details about
  1372. 1:03:12both customer ID equals 1 and customer
  1373. 1:03:18ID equals 99. Yeah. So that is why there
  1374. 1:03:22was a remove
  1375. 1:03:24basically remove over here.
  1376. 1:03:28Basically meaning that don't treat
  1377. 1:03:31one.park as what was defined over here.
  1378. 1:03:34Treat it as this one. That is why an ad
  1379. 1:03:38is added over here. Yeah. So that is the
  1380. 1:03:41reason why you see deletion vector in
  1381. 1:03:43both the places. It's just that the
  1382. 1:03:45deletion vector added over here is
  1383. 1:03:48updated with more detail. Now let's
  1384. 1:03:50perform the last operation basically to
  1385. 1:03:54use one of the park file that I've
  1386. 1:03:55created. Right? So insert into this
  1387. 1:03:59invoices table and we do a select star
  1388. 1:04:02from park
  1389. 1:04:04dot and we will select this last file
  1390. 1:04:08over here. Yeah. So let's go ahead and
  1391. 1:04:12insert this file. Insert the data that
  1392. 1:04:15we have here. And let me do a describe
  1393. 1:04:18history.
  1394. 1:04:22Yeah. So let's go ahead and do that. And
  1395. 1:04:25we see that after a delete there was a
  1396. 1:04:28right. And let's again see how this is
  1397. 1:04:32going to look like. So there's going to
  1398. 1:04:33be one JSON added.
  1399. 1:04:36And there you go. There is it as we
  1400. 1:04:38expect. And this should ideally just
  1401. 1:04:41contain one ad.
  1402. 1:04:44Yeah.
  1403. 1:04:47So it just contains one ad. We can close
  1404. 1:04:51the commit info and it just contains one
  1405. 1:04:53ad. Yeah. And
  1406. 1:04:56if we go here,
  1407. 1:04:58there is one file that has been added
  1408. 1:05:00over here which basically contains the
  1409. 1:05:02most recent data. So that is another
  1410. 1:05:06example. I I believe we've covered all
  1411. 1:05:08of the operation from creating to
  1412. 1:05:11inserting, updating and deleting. Yeah.
  1413. 1:05:14So to quickly summarize the key concept
  1414. 1:05:17behind Delta Lake is the use of a
  1415. 1:05:20transaction layer on top of your data.
  1416. 1:05:22Right? So your data is in the form of
  1417. 1:05:24park files and there is a transaction
  1418. 1:05:27layer called the delta log. Right? And
  1419. 1:05:30whenever it performs a transaction it
  1420. 1:05:33writes all of the data in a JSON file.
  1421. 1:05:37And this transaction in a delta lake is
  1422. 1:05:40an atomic unit of work. Right? It
  1423. 1:05:42basically groups one or more operations
  1424. 1:05:45together and these operations could be
  1425. 1:05:48let's say an insert, an update or a
  1426. 1:05:50delete. And the beauty of this is that
  1427. 1:05:52these transaction either complete
  1428. 1:05:55successfully or fail altogether. Yeah.
  1429. 1:05:59And you see that this brings the gist of
  1430. 1:06:02atomic in all of the asset properties.
  1431. 1:06:05And that is how atomicity is
  1432. 1:06:07implemented. So you may have this
  1433. 1:06:09question lingering in your mind that all
  1434. 1:06:12of this is good. The data is stored in
  1435. 1:06:14parquet files and there is a
  1436. 1:06:16transactional layer on top of it which
  1437. 1:06:18stores data in the form of which stores
  1438. 1:06:20transactions in the form of JSON files.
  1439. 1:06:22Right? But let's say when I do something
  1440. 1:06:25like a select star from this table,
  1441. 1:06:30how is the latest state of the table
  1442. 1:06:33that I get to see over here constructed,
  1443. 1:06:36right? With all of those park files
  1444. 1:06:38hanging here and there, right? With all
  1445. 1:06:40of those JSON files, right? How are all
  1446. 1:06:45of those used to construct this latest
  1447. 1:06:47state? Yeah. So let's quickly go ahead
  1448. 1:06:50and try to understand what happens
  1449. 1:06:52behind the scenes and this is where I'm
  1450. 1:06:55going to finally use this diagram. So
  1451. 1:06:58basically we are going to have two
  1452. 1:07:02things as we discussed. This is going to
  1453. 1:07:04be our data and this is one example
  1454. 1:07:06where I've shown partition data but this
  1455. 1:07:08can simply be your park file and then
  1456. 1:07:11this is going to be the delta log right
  1457. 1:07:15and these are your transactions.
  1458. 1:07:18These are your transaction. We'll come
  1459. 1:07:19to the checkpoint and the last
  1460. 1:07:21checkpoint by a little later in the
  1461. 1:07:24course. So don't worry about it. Be
  1462. 1:07:27relaxed. So let's take our example, the
  1463. 1:07:30example that we went through.
  1464. 1:07:33So the first statement that we ran was a
  1465. 1:07:36cat, right?
  1466. 1:07:39So it was a create table as select,
  1467. 1:07:43right? And this led to the creation of
  1468. 1:07:46the first transaction which is 0.json.
  1469. 1:07:51And this created a park file which was
  1470. 1:07:531. Park. Yeah. So we are just renaming
  1471. 1:07:56it to 1.park for simplicity. And this
  1472. 1:08:00contained all of the customers from 101
  1473. 1:08:02to 200.
  1474. 1:08:04The second operation that we ran was an
  1475. 1:08:06insert.
  1476. 1:08:09And this again led to the creation of
  1477. 1:08:11another transaction which was 1.json.
  1478. 1:08:16And this created another parket file
  1479. 1:08:20which was 2.
  1480. 1:08:22And it contained all of the customers
  1481. 1:08:25from 1 to 100.
  1482. 1:08:28The third operation that we ran was an
  1483. 1:08:32update,
  1484. 1:08:36right? and we updated customer ID equals
  1485. 1:08:411. Yeah.
  1486. 1:08:44So this again led to the creation of
  1487. 1:08:47another transaction recorded in 2.json
  1488. 1:08:51and we've already seen earlier that
  1489. 1:08:54customer ID equals 1 was present in
  1490. 1:08:572.pk. Right? So 2.park Park was removed
  1491. 1:09:02because we wanted to treat it
  1492. 1:09:04differently and
  1493. 1:09:082.par was added again
  1494. 1:09:12with a deletion vector containing
  1495. 1:09:14details about handling customer ID
  1496. 1:09:17equals 1. Right? Basically saying that
  1497. 1:09:19don't read customer ID equals 1 from
  1498. 1:09:21this park. And another parket file was
  1499. 1:09:25added
  1500. 1:09:27which only contained details about
  1501. 1:09:29customer ID equal to 1. That was
  1502. 1:09:30three.park. Yeah. Now the next operation
  1503. 1:09:34that we did was a delete
  1504. 1:09:40and this led to the creation of
  1505. 1:09:43three.json JSON
  1506. 1:09:45and what we did was that we deleted
  1507. 1:09:49customer ID
  1508. 1:09:53equals 99. Yeah. And this 99 is supposed
  1509. 1:09:56to reside again in this 2.par. Right. So
  1510. 1:09:59we again need to change the way we deal
  1511. 1:10:01with two. So what we do is that we
  1512. 1:10:06remove
  1513. 1:10:08two.par and we add another we remove
  1514. 1:10:12two. Okay, with this deletion vector and
  1515. 1:10:16we add another 2 dotpar
  1516. 1:10:21with a different deletion vector. Let's
  1517. 1:10:23say this was DV1
  1518. 1:10:26and this is DV2 now. Yeah. So that is
  1519. 1:10:28what happened as a part of the delete
  1520. 1:10:31operation. Two.park now contains details
  1521. 1:10:34about handling customer ID equals 1 and
  1522. 1:10:37customer ID equals 99. Right? It
  1523. 1:10:39basically saying that don't read
  1524. 1:10:41customer ID equals 1 and 99 from this
  1525. 1:10:44park file. Now the last operation was an
  1526. 1:10:48insert operation and this again happened
  1527. 1:10:52in 4.json. The transaction happened
  1528. 1:10:54there and this simply
  1529. 1:10:58created another parket file. Right? So
  1530. 1:11:01this simply created 4.par park and it
  1531. 1:11:05contained customers from 2011 until some
  1532. 1:11:0999,000 some big number right now the
  1533. 1:11:12question is okay we had all these
  1534. 1:11:13transactions now how do I compute the
  1535. 1:11:16latest state right and the very simple
  1536. 1:11:19answer to this is just do a summation
  1537. 1:11:22yeah how do you do the summation the way
  1538. 1:11:25you do the summation is basically
  1539. 1:11:30this was added and this was removed so
  1540. 1:11:32you remove both of this. Right? So this
  1541. 1:11:34is a plus and this is a minus. So these
  1542. 1:11:36two get removed.
  1543. 1:11:38This was added. This was removed. So
  1544. 1:11:41basically this gets removed. Now the
  1545. 1:11:43only thing that stays until now is this
  1546. 1:11:47one which is 1.pk.
  1547. 1:11:52This one 3.
  1548. 1:11:57This one 2.
  1549. 1:12:02with DV2 and then 4.par.
  1550. 1:12:10Right? So whenever we want to read the
  1551. 1:12:13file and get the latest state, we
  1552. 1:12:15basically read in 1.pk from where we get
  1553. 1:12:18all customer ids from 101 to 200. We
  1554. 1:12:21read in 3.pk where we get in customer ID
  1555. 1:12:24equals 1. We read in four.park park
  1556. 1:12:26where we get 2019,000
  1557. 1:12:30whatever number it was and then we get
  1558. 1:12:32we read 2.par wherein we read customer
  1559. 1:12:36ids from 1 to 100 excluding customer ids
  1560. 1:12:411 and 99.
  1561. 1:12:44Yeah, because one is read from here and
  1562. 1:12:4799 was deleted. So that is the way that
  1563. 1:12:50is how simple it makes for processing
  1564. 1:12:53the transaction log and you see how
  1565. 1:12:55beautifully delta log creates these
  1566. 1:12:58transactions and basically puts them in
  1567. 1:13:00a way that is very efficient right there
  1568. 1:13:02is lot of remove addition and all of
  1569. 1:13:04that happening but that actually makes
  1570. 1:13:07it very efficient because there is no
  1571. 1:13:09actual writing happening right it is
  1572. 1:13:11basically maintaining all of these
  1573. 1:13:13states in a ledger. So another
  1574. 1:13:16interesting question that you may have
  1575. 1:13:17in mind is that Delta creates a JSON
  1576. 1:13:22file for every transaction right and
  1577. 1:13:24there could be companies who have
  1578. 1:13:26millions of transactions in a day right
  1579. 1:13:29and they could have around thousands to
  1580. 1:13:33five of thousands of tables right and I
  1581. 1:13:36don't even want to go ahead and multiply
  1582. 1:13:38five thousands of tables with millions
  1583. 1:13:41of JSONs being created every day right
  1584. 1:13:43so this requires a lot of compute power
  1585. 1:13:47in order to read the JSON file. So the
  1586. 1:13:50question is that how does delta scale
  1587. 1:13:55how does it handle reading massive
  1588. 1:13:57amount of JSON files, right? So let's go
  1589. 1:14:01ahead and simulate such a kind of a
  1590. 1:14:03situation. Now, of course, we won't be
  1591. 1:14:05creating millions or even hundreds of
  1592. 1:14:08JSON files, but um we'll be creating a
  1593. 1:14:10few of them to get a rough
  1594. 1:14:13understanding, right? So, let's pick up
  1595. 1:14:16this insert statement. And I'm going to
  1596. 1:14:18run this in a loop and let me put
  1597. 1:14:22a let me make this a Python cell. And
  1598. 1:14:24let's do for i in range 10
  1599. 1:14:28and or maybe let me put 50, right? And
  1600. 1:14:33let's go ahead and run spark SQL.
  1601. 1:14:38And let's paste this over here.
  1602. 1:14:41So we are going to insert this spark a
  1603. 1:14:4450 time, right? So this should add 50
  1604. 1:14:47transactions to the table. And let me go
  1605. 1:14:50ahead and print
  1606. 1:14:53if
  1607. 1:14:55insert
  1608. 1:14:57completed. And let's go ahead and run
  1609. 1:14:59this.
  1610. 1:15:09So now all of the 50 inserts are
  1611. 1:15:12complete. So let's quickly have a look
  1612. 1:15:14at how this looks behind the scenes. And
  1613. 1:15:17we see a lot of park files which is
  1614. 1:15:20expected because we added in a lot of
  1615. 1:15:22data. Let's also have a look at the
  1616. 1:15:24delta log. And we see something
  1617. 1:15:26interesting which is the presence of a
  1618. 1:15:29compacted JSON. Right? Let me load all
  1619. 1:15:33of this. So what we see here is that a
  1620. 1:15:37compacted JSON is from the 1st to the
  1621. 1:15:416th, right? The 7th to the 12th, the
  1622. 1:15:4513th to the 18th. That means a compacted
  1623. 1:15:48JSON is being created for every six JSON
  1624. 1:15:51files, right? So I believe it is
  1625. 1:15:53condensing all of the information.
  1626. 1:15:55Actually not condensing but just picking
  1627. 1:15:58up all of the information in each of
  1628. 1:16:00those JSON files and then dumping it
  1629. 1:16:03over here so that it doesn't have to go
  1630. 1:16:05through the overhead of opening and
  1631. 1:16:07closing those files. Right? And
  1632. 1:16:12finally we see that after 36 JSON files
  1633. 1:16:16have been created there is something
  1634. 1:16:17called a checkpoint.park.
  1635. 1:16:20Yeah. So let's understand what all of
  1636. 1:16:23these are exactly. So I'll refer this
  1637. 1:16:26diagram
  1638. 1:16:28that I created. So over here we see a
  1639. 1:16:33situation which is similar to what we
  1640. 1:16:35just saw. Right?
  1641. 1:16:37What this basically means is that from
  1642. 1:16:40here to here all of this data is
  1643. 1:16:43compacted and it is present in this
  1644. 1:16:46docompacted.json
  1645. 1:16:47file. Yeah. Um what this also
  1646. 1:16:51essentially means is that there is no
  1647. 1:16:52summation happening. Remember we
  1648. 1:16:54discussed about summation. Yeah. And
  1649. 1:16:56that is how delta computes the latest
  1650. 1:16:58state. But that is not what is happening
  1651. 1:17:00here. It just dumps all of the data that
  1652. 1:17:03it had. So let's say this had
  1653. 1:17:06let's say add 1.park
  1654. 1:17:10and it has add 2. Okay. And this
  1655. 1:17:14transaction over here, it had a remove
  1656. 1:17:18one.
  1657. 1:17:20If you were to sum this, then probably
  1658. 1:17:221.par over here and 1.pk remove over
  1659. 1:17:26here, it should simply go away, right?
  1660. 1:17:28It should only end up having two park.
  1661. 1:17:30But that's not the case. It basically
  1662. 1:17:32picks up all of the data and dumps it
  1663. 1:17:34over here. So, it's going to contain add
  1664. 1:17:36of
  1665. 1:17:381.par,
  1666. 1:17:39add of two.
  1667. 1:17:43and then the remove of
  1668. 1:17:461.park park and of course all of the
  1669. 1:17:49transactions and the details all of the
  1670. 1:17:51details that is contained in these JSON
  1671. 1:17:54files right so they are all added and
  1672. 1:17:56they are put over here similarly for
  1673. 1:17:58this one as well for all of the details
  1674. 1:18:02that is present over here they are added
  1675. 1:18:03over here now imagine you just committed
  1676. 1:18:06a transaction and
  1677. 1:18:08you are over here
  1678. 1:18:11you just performed this operation so now
  1679. 1:18:14how is delta going to compute the latest
  1680. 1:18:15state it is basically going to read it
  1681. 1:18:19is going to read this file. It is going
  1682. 1:18:22to read this file and then these
  1683. 1:18:26these two files over here. That is how
  1684. 1:18:28it is going to compute the latest stage.
  1685. 1:18:31Now imagine this is a lot beneficial
  1686. 1:18:33because it doesn't have to read all of
  1687. 1:18:35this, right? It doesn't end up reading
  1688. 1:18:38all of this.
  1689. 1:18:40So that is how this is optimized for
  1690. 1:18:43performance. Now after 36 files have
  1691. 1:18:47been created you see something called a
  1692. 1:18:49checkpoint.park
  1693. 1:18:52and what it does is that it again
  1694. 1:18:55condenses all of the information from
  1695. 1:18:59the zero to the 35th
  1696. 1:19:02file and puts it over here. Yeah. So
  1697. 1:19:05let's say you were at this transaction.
  1698. 1:19:08The only file that you would have to
  1699. 1:19:09read is that it it tries to find out do
  1700. 1:19:12we have a checkpoint.park file or not.
  1701. 1:19:15Right? If it doesn't have then it is
  1702. 1:19:17going to follow the normal process the
  1703. 1:19:19one that we just discussed earlier. If
  1704. 1:19:21it has a checkpoint.paret file find the
  1705. 1:19:24latest checkpoint.pket file. So that
  1706. 1:19:27latest checkpoint.park file is going to
  1707. 1:19:29be this one. and then it is going to
  1708. 1:19:31apply it is going to read these two JSON
  1709. 1:19:33file and then it is going to produce the
  1710. 1:19:36final state. So that means that you
  1711. 1:19:40don't end up reading all of this.
  1712. 1:19:44You don't end up reading all of this and
  1713. 1:19:47that is tremendous amount of
  1714. 1:19:50optimization right so that is how delta
  1715. 1:19:53handles massive amount of JSON files
  1716. 1:19:56right simply using compacted JSONs and
  1717. 1:19:59by using checkpoint.pk park. Another
  1718. 1:20:02very important point to note is that we
  1719. 1:20:04see that the gap right the gap at which
  1720. 1:20:07this checkpoint.parquet file is created
  1721. 1:20:09is 36.
  1722. 1:20:12If you were to use the open source
  1723. 1:20:13version, you would probably see this
  1724. 1:20:15file being created at number 10. Right?
  1725. 1:20:18After 10 JSON files are created, a
  1726. 1:20:20checkpoint.parkey files are pres is
  1727. 1:20:23created. Right? Now the reason for this
  1728. 1:20:25difference why 36 on data bricks and 10
  1729. 1:20:29in the open source version is simply
  1730. 1:20:30because datab bricks has made several
  1731. 1:20:34performance optimizations in place so
  1732. 1:20:36that it can relax the gap and it can
  1733. 1:20:38create these checkpoint files at a very
  1734. 1:20:42relaxed gap right so that's one of the
  1735. 1:20:44most important points and it also varies
  1736. 1:20:48per cloud provider so in AWS this number
  1737. 1:20:52would be 36 six maybe in Azure
  1738. 1:20:56it is going to be some 100ish
  1739. 1:20:59right
  1740. 1:21:01some 100ish so I believe this gives you
  1741. 1:21:05an overall idea about how delta scales
  1742. 1:21:08and handles massive amount of metadata
  1743. 1:21:12and JSON files right earlier we talked
  1744. 1:21:14about how delta solves the isolation
  1745. 1:21:18problem the I in acid and this simply
  1746. 1:21:21means that it allows for several
  1747. 1:21:23transactions to operate concurrently to
  1748. 1:21:26happen at the same time without one
  1749. 1:21:28transaction worrying about what
  1750. 1:21:30happening in the other transaction.
  1751. 1:21:32Yeah. So it can remain carefree and it
  1752. 1:21:35can do whatever it wants without
  1753. 1:21:37worrying about what the other
  1754. 1:21:39transaction is doing. Right. And it
  1755. 1:21:41achieved this through something very
  1756. 1:21:43beautiful called optimistic concurrency
  1757. 1:21:46control. But there's before before we go
  1758. 1:21:48ahead and understand what optimistic
  1759. 1:21:50concurrency control is, there's
  1760. 1:21:52something called pessimistic concurrency
  1761. 1:21:54control. And these two guys go hand in
  1762. 1:21:56hand, right? So in order to truly
  1763. 1:21:58appreciate what OC optimistic
  1764. 1:22:01concurrency control is, let's first
  1765. 1:22:03understand what pessimistic concurrency
  1766. 1:22:06control is. And in in that sense, we'll
  1767. 1:22:09be able to understand why OC was a
  1768. 1:22:12better choice. Right? So with
  1769. 1:22:15pessimistic concurrency control, the
  1770. 1:22:16DBMS, the database management system
  1771. 1:22:19assumes that conflicts are bound to
  1772. 1:22:21happen, right? Conflicts are something
  1773. 1:22:23that will eventually happen and they are
  1774. 1:22:26likely to occur and if it's likely to
  1775. 1:22:28occur, it will occur. Yeah. So to to to
  1776. 1:22:31make sure that you understand what a
  1777. 1:22:33conflict is, conflict basically is a
  1778. 1:22:35situation where two or more transactions
  1779. 1:22:38are trying to modify the same data at
  1780. 1:22:42the same time resulting in an invalid
  1781. 1:22:45state, right? Bringing the database in
  1782. 1:22:47an invalid state. So we saw several of
  1783. 1:22:51these examples earlier, right? So if I
  1784. 1:22:53were to go back to the consistency
  1785. 1:22:56example that we took.
  1786. 1:22:59So in the consistency example we saw
  1787. 1:23:01that there were two transactions which
  1788. 1:23:03were trying to update the account
  1789. 1:23:06balance at the same time and the one
  1790. 1:23:08which executed last is going to be the
  1791. 1:23:10final balance and this is going to
  1792. 1:23:12result in an inconsistent state. Right?
  1793. 1:23:14So with pessimistic locking what we do
  1794. 1:23:17is that we whenever a transaction is
  1795. 1:23:20trying to make a change to a database we
  1796. 1:23:24basically lock that data so that another
  1797. 1:23:26person cannot come in and it cannot make
  1798. 1:23:29changes right and let's understand that
  1799. 1:23:31with an example so let's assume that we
  1800. 1:23:33have two transactions the first one is
  1801. 1:23:35T1 this guy wants to withdraw $100 from
  1802. 1:23:39this account and there's another
  1803. 1:23:41transaction called T2
  1804. 1:23:43this guy wants to deposit $400 into the
  1805. 1:23:47account. Yeah. And let's assume that T1
  1806. 1:23:49is the one which starts first. So T1
  1807. 1:23:52basically starts first. So the moment it
  1808. 1:23:54starts,
  1809. 1:23:56there is an exclusive lock which is
  1810. 1:23:59placed on this row in the database.
  1811. 1:24:02Right? So there is an exclusive lock
  1812. 1:24:04which is acquired by T1 and it is placed
  1813. 1:24:09on this row in the database. So now T1
  1814. 1:24:12goes ahead and it basically reads in the
  1815. 1:24:14balance. It reads in $1,000. Yeah. And
  1816. 1:24:19because an exclusive log is acquired by
  1817. 1:24:22T1, T2 cannot go ahead and make changes.
  1818. 1:24:25So basically this will be paused. This
  1819. 1:24:28will basically wait until T1 has
  1820. 1:24:30completed. So it basically reads in T1
  1821. 1:24:32basically reads in $1,000. it goes ahead
  1822. 1:24:35and withdraws $100 and then
  1823. 1:24:41it computes the account balance to be
  1824. 1:24:43$900. Yeah. Now the moment it computes
  1825. 1:24:46this once that is done it simply goes
  1826. 1:24:48ahead and makes an update.
  1827. 1:24:52So whatever account ID this is a 45120
  1828. 1:24:56and this is going to update the account
  1829. 1:24:59balance to $900. Yeah. Now the moment
  1830. 1:25:02this is done this lock is released.
  1831. 1:25:07Now the moment this log is released T2
  1832. 1:25:10gets a chance and T2 then places an
  1833. 1:25:14exclusive lock. So T2 is now going to
  1834. 1:25:17place an exclusive lock on this row in
  1835. 1:25:21the database. That simply means that
  1836. 1:25:23after T2 has completed executing then
  1837. 1:25:26only another transaction can come in and
  1838. 1:25:28update this row. So again, it basically
  1839. 1:25:31what it does is it simply reads in the
  1840. 1:25:34account balance which is $900 and it
  1841. 1:25:38adds in $400 which is $1,300. So it
  1842. 1:25:42basically goes over here and update the
  1843. 1:25:46account balance and this is going to be
  1844. 1:25:48$1,300. Now after this is completed, it
  1845. 1:25:52simply releases this lock.
  1846. 1:25:56Now this lock is released and the final
  1847. 1:25:59balance is $1,300.
  1848. 1:26:02So this is correct, right? There's no
  1849. 1:26:04problem with this. And what we saw
  1850. 1:26:06earlier, if you remember in the
  1851. 1:26:08consistency example,
  1852. 1:26:11we also started a transaction over here,
  1853. 1:26:14right? We started this transaction and
  1854. 1:26:17we acquired a lock and then only all of
  1855. 1:26:20this went ahead and computed the state
  1856. 1:26:23due to which the other transaction the
  1857. 1:26:25transaction two was not able to go
  1858. 1:26:27inside right so we acquired a
  1859. 1:26:29pessimistic lock now this all of this is
  1860. 1:26:32really good right there's no problem
  1861. 1:26:33with this but imagine if locking were
  1862. 1:26:36not in place if locking was not in place
  1863. 1:26:39what would have happened both the
  1864. 1:26:42transactions T1 and T2 they would have
  1865. 1:26:44come in, they would have read $1,000
  1866. 1:26:47and this guy would have subtracted $100.
  1867. 1:26:50It would have computed $900. This would
  1868. 1:26:52have added $400. This would have
  1869. 1:26:54computed 1,400. And whichever was the
  1870. 1:26:58one updating this last would be the
  1871. 1:27:02final balance of the account. Now, of
  1872. 1:27:04course, there's a chance that you could
  1873. 1:27:05have more money, but there's also a good
  1874. 1:27:08chance that you'll end up having less
  1875. 1:27:09money. Yeah. So pessimistic locking
  1876. 1:27:13helps us achieve consistency. It
  1877. 1:27:15maintain the correctness of the
  1878. 1:27:17database. So what are the problems? Why
  1879. 1:27:20don't we want to use pessimistic
  1880. 1:27:21concurrency control? Now you see that T2
  1881. 1:27:24had to wait.
  1882. 1:27:27It had to wait till T1 was completed.
  1883. 1:27:32And imagine if you have millions of
  1884. 1:27:35transactions taking place then this
  1885. 1:27:38whole system is going to come to a halt
  1886. 1:27:40right this is not going to be meaningful
  1887. 1:27:42anymore because you would have to wait
  1888. 1:27:44endlessly in order for few of the
  1889. 1:27:47transactions to complete and then I
  1890. 1:27:49would get a chance right so that is
  1891. 1:27:51where the problem comes in it makes your
  1892. 1:27:53system slow and for this reason we want
  1893. 1:27:57to use optimistic concurrency control
  1894. 1:27:59with optimistic concurrency control
  1895. 1:28:02transactions do not obtain locks when
  1896. 1:28:05they read or write and that's the
  1897. 1:28:07beautiful part right so the name
  1898. 1:28:09optimistic actually comes from the fact
  1899. 1:28:12that it assumes that conflicts are very
  1900. 1:28:16unlikely to occur and this is quite just
  1901. 1:28:18the opposite of what pessimistic
  1902. 1:28:20concurrency control assumes right so
  1903. 1:28:22this guy assumed that conflicts are very
  1904. 1:28:25unlikely to occur so I don't need to
  1905. 1:28:27think about it right and if at all it
  1906. 1:28:29occurs the conflicting transaction is
  1907. 1:28:31the guy who going to be retrying, right?
  1908. 1:28:34So let's understand this with an
  1909. 1:28:36example. So let's take the same two
  1910. 1:28:38transactions, right? T1 basically trying
  1911. 1:28:41to withdraw $100 and T2 trying to
  1912. 1:28:44deposit $400.
  1913. 1:28:47Yeah. And these two transactions go
  1914. 1:28:50ahead and read in this row at the same
  1915. 1:28:54time. So T1 basically reads in $1,000 to
  1916. 1:28:58be the current balance. T2 reads in
  1917. 1:29:02$1,000 to be the current balance. And
  1918. 1:29:05the reason why they are able to read
  1919. 1:29:07this row at the same time is because as
  1920. 1:29:10I discussed earlier, there is no concept
  1921. 1:29:12of logs in optimistic concurrency
  1922. 1:29:15control. Right? So because there is no
  1923. 1:29:17concept of logs, this row over here is
  1924. 1:29:21not logged by either of these
  1925. 1:29:23transaction. they both can go in and
  1926. 1:29:25read in the data over here at the same
  1927. 1:29:28point in time. Right? So they go ahead
  1928. 1:29:31and read the balance and along with that
  1929. 1:29:32they also read in something called
  1930. 1:29:34version number and time frame. The
  1931. 1:29:36version number is going to help track
  1932. 1:29:38changes within a table. So it reads in
  1933. 1:29:41version number one and then it reads in
  1934. 1:29:43something called TS1. Let's call this
  1935. 1:29:45TS1.
  1936. 1:29:46So this is going to read in one and TS1.
  1937. 1:29:50Now it goes ahead and does all of the
  1938. 1:29:52calculations, right? So this is going to
  1939. 1:29:55compute do a subtraction of $100 and it
  1940. 1:29:58is going to compute the final value
  1941. 1:30:01which is $900, right? And this is going
  1942. 1:30:04to add $400 and it is going to compute
  1943. 1:30:07the final balance which is $1,400. Now
  1944. 1:30:10both of these people are going to try to
  1945. 1:30:14commit.
  1946. 1:30:16Both of them are going to try to commit.
  1947. 1:30:20Yeah.
  1948. 1:30:22Now we need to understand that there is
  1949. 1:30:24going to be one person who is going to
  1950. 1:30:27commit first. Yeah. So let's assume that
  1951. 1:30:29T1 is the person who commits first and
  1952. 1:30:32there is a procedure there is a way in
  1953. 1:30:35which the commit is going to happen. So
  1954. 1:30:37the way it happen is it is going to
  1955. 1:30:40check what is the current version
  1956. 1:30:42number. It reads in that the current
  1957. 1:30:43version number is one and it sees that
  1958. 1:30:45the version number it has is also one.
  1959. 1:30:50Yeah. So that means that I am working on
  1960. 1:30:53the latest data. There is nobody who
  1961. 1:30:55came in while I was working and made
  1962. 1:30:58changes to the table. Right? So I can
  1963. 1:30:59simply go ahead and I can make the
  1964. 1:31:03addition update the balance right
  1965. 1:31:05because I was working on the latest
  1966. 1:31:06data. So now what happens is that it
  1967. 1:31:09goes ahead and add this row which is
  1968. 1:31:11A45120.
  1969. 1:31:12This is going to be 900 and it updates
  1970. 1:31:15the version to version two and this is
  1971. 1:31:17going to be TS2.
  1972. 1:31:19Now this person
  1973. 1:31:22transaction T2 who was also trying to
  1974. 1:31:24commit at the same time but actually
  1975. 1:31:25ended up committing second right after
  1976. 1:31:28this happened what it sees is that it
  1977. 1:31:31reached the version number the version
  1978. 1:31:33number that it finds is two but the
  1979. 1:31:35version number it has is one. So this
  1980. 1:31:39simply helps it understand that there is
  1981. 1:31:42a person who came before me before I was
  1982. 1:31:44trying to commit my transaction and
  1983. 1:31:46actually made changes and this means
  1984. 1:31:49that I was working on stale data. I need
  1985. 1:31:51to read in the latest data make the
  1986. 1:31:53changes and then actually update the
  1987. 1:31:56table or add a row to the table. Right?
  1988. 1:31:59So what it does is that it will simply
  1989. 1:32:02fail this transaction
  1990. 1:32:05because it was not working on the latest
  1991. 1:32:07data. Yeah. So now what happens is
  1992. 1:32:12T2 now reads in the latest version of
  1993. 1:32:15the table. So it reads in the balance
  1994. 1:32:17900. The latest version is two. And it
  1995. 1:32:20reads in the time stamp ts2. It goes
  1996. 1:32:22ahead applies the operation which is
  1997. 1:32:24$400. Computes it to,300.
  1998. 1:32:28And now it goes to commit again.
  1999. 1:32:31So it goes and commits again. And it
  2000. 1:32:34follows the same procedure. So it checks
  2001. 1:32:36what is the latest version. The latest
  2002. 1:32:38version appears to be two and the
  2003. 1:32:40version it has read is also two. That
  2004. 1:32:42means there is nobody who came in while
  2005. 1:32:44I was working and make changes to the
  2006. 1:32:46table. Right? So that also mean that I
  2007. 1:32:48was working on the latest data. So I can
  2008. 1:32:50make updates make a row addition to the
  2009. 1:32:52table. So that is what it does now. And
  2010. 1:32:55let me first just remove all of this.
  2011. 1:32:59So it goes ahead and it makes an update.
  2012. 1:33:05to this table. It adds in a new row. It
  2013. 1:33:07updates the time stamp and this is
  2014. 1:33:09called TS3. And this is how
  2015. 1:33:12the table ends up in the correct state.
  2016. 1:33:16Yeah. Now imagine that T2 was still
  2017. 1:33:20working and there was a transaction
  2018. 1:33:22called T3. It wanted to read the latest
  2019. 1:33:25state of the table. And let's imagine
  2020. 1:33:26that this row was not yet committed. T2
  2021. 1:33:29was still working. What would happen? T3
  2022. 1:33:31would simply go in and read the latest
  2023. 1:33:33state. it would read in $900. So this
  2024. 1:33:36means that even though T2 is working, T3
  2025. 1:33:39is not affected, right? So all of these
  2026. 1:33:42transactions can go on and on
  2027. 1:33:44simultaneously and there can be millions
  2028. 1:33:46of such transaction and that is where
  2029. 1:33:47the beauty of optimistic concurrency
  2030. 1:33:50control lie. Yeah. And it also allows T2
  2031. 1:33:54to fail and retry. Yeah. So it gives us
  2032. 1:33:57that flexibility. So to reiterate with
  2033. 1:34:00optimistic concurrency control we have
  2034. 1:34:03completely avoided logs. Yeah. And the
  2035. 1:34:07key to this is that if I'm performing
  2036. 1:34:09operations the other transaction isn't
  2037. 1:34:13or doesn't need to be aware of whatever
  2038. 1:34:15I'm doing right they can simply read in
  2039. 1:34:18or they can simply fail and retry.
  2040. 1:34:20Right? My operations are not visible or
  2041. 1:34:23affecting other operations other
  2042. 1:34:26transaction. Yeah. And this allows as I
  2043. 1:34:29discussed earlier it allows very high
  2044. 1:34:31levels of concurrency and this is
  2045. 1:34:33particularly beneficial in read heavy
  2046. 1:34:36systems. Yeah, it is beneficial in read
  2047. 1:34:40heavy systems. So if you were to think
  2048. 1:34:41about it optimistic concurrency control
  2049. 1:34:44basically the foundations lie on the
  2050. 1:34:46fact that
  2051. 1:34:49conflicts
  2052. 1:34:51are very rare.
  2053. 1:34:54They will most unlikely happen, right?
  2054. 1:34:56They are not bound to happen. If they
  2055. 1:34:58happen the conflicting transactions are
  2056. 1:35:00going to read right. Yeah. So that is
  2057. 1:35:02the foundation of optimistic concurrency
  2058. 1:35:04control. And when do conflicts happen
  2059. 1:35:06very rarely
  2060. 1:35:08in those systems where right do not
  2061. 1:35:10happen a lot. Yeah. So in those systems
  2062. 1:35:13which is read heavy. So this is very
  2063. 1:35:15beneficial for read heavy system. It's
  2064. 1:35:17very good for read heavy systems. So I
  2065. 1:35:20believe this helps you get an overall
  2066. 1:35:23idea and an in-depth understanding about
  2067. 1:35:25how delta solves the isolation problem
  2068. 1:35:28quite beautifully using optimistic
  2069. 1:35:30concurrency. Now let's go ahead and talk
  2070. 1:35:32about time travel and virgining. So I
  2071. 1:35:35believe in the previous few example that
  2072. 1:35:37we've seen right we did statements like
  2073. 1:35:40describe history of a table and it
  2074. 1:35:42showed us several versions of that table
  2075. 1:35:45right so versioning basically helps us
  2076. 1:35:48track different stages of the table
  2077. 1:35:51right so let's say you had a table you
  2078. 1:35:53made an update to the table now it
  2079. 1:35:56became a new version of the table let's
  2080. 1:35:58say you made a delete so that becomes
  2081. 1:36:01another version of the table and you can
  2082. 1:36:03go to any of these versions back in time
  2083. 1:36:06and you can restore your table to that
  2084. 1:36:09version. So going back in time and
  2085. 1:36:12restoring it to any other version is
  2086. 1:36:14basically time travel. Right? So we are
  2087. 1:36:16going to have a look at all of this
  2088. 1:36:18through a lot of examples. Let's go
  2089. 1:36:20ahead and create a table first. Right?
  2090. 1:36:23And we going to perform all of the
  2091. 1:36:25operations on this table. So I'll
  2092. 1:36:26quickly copy this path and and create a
  2093. 1:36:30table using the data that is there in
  2094. 1:36:32this park file. So let's go ahead and do
  2095. 1:36:34a create or replace table as then
  2096. 1:36:38actually the name of the table delta
  2097. 1:36:41catalog
  2098. 1:36:44dot delta db dot invoices
  2099. 1:36:49tt time travel and virgining as select
  2100. 1:36:53star from
  2101. 1:36:55park k
  2102. 1:36:58dot this path right so let's go ahead
  2103. 1:37:02and create this.
  2104. 1:37:04Okay. Now let's have a look at this
  2105. 1:37:06table how this looks like. Delta DB
  2106. 1:37:10delta delta catalog
  2107. 1:37:13dot delta DB dot this table. Right? And
  2108. 1:37:17let's do a limit five.
  2109. 1:37:20Okay. So this is how the table looks
  2110. 1:37:21like and I believe we've seen it a
  2111. 1:37:23couple of times already, right? So now
  2112. 1:37:25let's go ahead and perform a few
  2113. 1:37:26operations on this table. So let me go
  2114. 1:37:29ahead and do a delete first, right? So
  2115. 1:37:32we'll do delete from this table where
  2116. 1:37:36customer ID equals 1. So this row that
  2117. 1:37:39you see over here is going to be
  2118. 1:37:40removed. Let me go ahead and run this.
  2119. 1:37:46Okay, great. That succeeded. We are not
  2120. 1:37:47going to verify that now. Uh let's go
  2121. 1:37:50ahead and also do a few more operations.
  2122. 1:37:51So we'll do update
  2123. 1:37:54this table set quantity
  2124. 1:37:58quantity equals 25 where customer ID
  2125. 1:38:03equals 5. Right? So we see that the
  2126. 1:38:06customer ID 5 over here has a quantity
  2127. 1:38:08equals 1 and I want to update that
  2128. 1:38:10quantity to some random number 25.
  2129. 1:38:13Right? Let's go ahead and do that.
  2130. 1:38:16Okay, that works.
  2131. 1:38:19And the final statement we'll do is
  2132. 1:38:23insert into this table and select star
  2133. 1:38:26from
  2134. 1:38:29uh we did it from a park file last time
  2135. 1:38:32right so we inserted customer id from 1
  2136. 1:38:35to 100 let's do from 101 to 200 now yeah
  2137. 1:38:41so now let's verify all the changes very
  2138. 1:38:43quickly so the latest version of the
  2139. 1:38:47table would be this one right if I do a
  2140. 1:38:49select star from what? From this table
  2141. 1:38:51name, I would get the latest version of
  2142. 1:38:53the table. So if I were to do select
  2143. 1:38:56star from this table where customer ID
  2144. 1:38:58equals 1, there should be no rows. Let's
  2145. 1:39:02go ahead and run this. I didn't get a
  2146. 1:39:04row. That means the delete worked fine.
  2147. 1:39:07Now let me check where customer ID
  2148. 1:39:10equals 5. So the quantity is 25. That
  2149. 1:39:13means the update also worked fine.
  2150. 1:39:17And then finally
  2151. 1:39:19let's check how many customers have been
  2152. 1:39:22inserted with customer ID greater than
  2153. 1:39:25100. And the last time we did this we
  2154. 1:39:26saw that there were total 100ed rows. So
  2155. 1:39:30the count star should give me a count of
  2156. 1:39:33100. Perfect. Let's have a look at the
  2157. 1:39:36history of the table. Yeah. So I'm going
  2158. 1:39:38to write describe history of delta
  2159. 1:39:43catalog
  2160. 1:39:45delta db.invoices invoices
  2161. 1:39:48TTV. Right?
  2162. 1:39:51So here we see that there are four
  2163. 1:39:54versions of the table. Right? The first
  2164. 1:39:56version was created when we did a create
  2165. 1:39:59or replace table. Right? This statement
  2166. 1:40:01right over here. So this created the
  2167. 1:40:03first version and then we performed
  2168. 1:40:06three operation. The first one was a
  2169. 1:40:08delete where we deleted customer ID
  2170. 1:40:10equals 1. Then we updated customer ID
  2171. 1:40:12equals 5. And then we wrote in 100
  2172. 1:40:15additional records from 101 to 200 right
  2173. 1:40:18the customer ID. So that created four
  2174. 1:40:21versions. The zero version another
  2175. 1:40:23operation was applied. The first version
  2176. 1:40:25another operation was applied and the
  2177. 1:40:27second version and then so on. Right? So
  2178. 1:40:30these are essentially the version that
  2179. 1:40:32get created and using those versions you
  2180. 1:40:36can go back in time. Right? So let's
  2181. 1:40:38take an example. So if we do a select
  2182. 1:40:41star from this table where customer ID
  2183. 1:40:44equals 1, we won't get any row back,
  2184. 1:40:47right? We know that. But let's say we
  2185. 1:40:49want to go back in time and we say
  2186. 1:40:53version as of we know that in version
  2187. 1:40:57zero it contains the customer ID equals
  2188. 1:41:001 row. Right? So let's do this.
  2189. 1:41:04And there you go. You see that the row
  2190. 1:41:06where customer ID equals 1 is returned
  2191. 1:41:08because this time we pick the table
  2192. 1:41:12where version was zero the oldest
  2193. 1:41:14version over here. Yeah. And you can
  2194. 1:41:16also do this by selecting the time stamp
  2195. 1:41:19instead of the version number. And the
  2196. 1:41:21only change that you need to do is
  2197. 1:41:23select star from this table timestamp
  2198. 1:41:27as of this where
  2199. 1:41:32customer ID equals 1.
  2200. 1:41:35That should work and it gives you the
  2201. 1:41:36same results. So you can either use
  2202. 1:41:39version as of or you can use timestamp
  2203. 1:41:42as of and go back to any of the versions
  2204. 1:41:46back in time. Right? So here we just did
  2205. 1:41:48a select statement. Right? Now what if
  2206. 1:41:50you actually want to restore the table
  2207. 1:41:53to this version? How do you do it? Yeah.
  2208. 1:41:56So the way you do it is by actually
  2209. 1:41:59using the restore command. Yeah. So you
  2210. 1:42:03write restore table delta catalog to
  2211. 1:42:09version as of zero. You can also use
  2212. 1:42:14time stamp as of zero.
  2213. 1:42:17Time stamp sorry not time stamp as of
  2214. 1:42:19zero but time stamp as of whatever time
  2215. 1:42:22stamp you have over here. You can do
  2216. 1:42:25this as well.
  2217. 1:42:27But let me comment this out for now
  2218. 1:42:29because we cannot run both of them at
  2219. 1:42:32the same time.
  2220. 1:42:35So now instead of doing this, instead of
  2221. 1:42:39referring to version number zero, let's
  2222. 1:42:42remove this and then do a where customer
  2223. 1:42:46ID equals 1. So in this case now we
  2224. 1:42:48should get that row because our table
  2225. 1:42:50has been restored to the zero version.
  2226. 1:42:55And there you go. So you see that the
  2227. 1:42:58row where customer ID equals 1 is
  2228. 1:43:00returned. And let's also have a look at
  2229. 1:43:02how the history of the table now looks
  2230. 1:43:04like.
  2231. 1:43:08Great. So it tracks every operation. So
  2232. 1:43:11write was our last operation and after
  2233. 1:43:14right we did a restore. So it tracked
  2234. 1:43:17that operation. Yeah. And it was
  2235. 1:43:20restored to version number zero and that
  2236. 1:43:22is also tracked over here. Yeah. Now
  2237. 1:43:24let's say that you want to restore the
  2238. 1:43:26table to a particular time stamp. Yeah.
  2239. 1:43:30So let's take this example. We want to
  2240. 1:43:33there there's a particular version that
  2241. 1:43:34exist at 9:31 and that version is
  2242. 1:43:37version number two. Another version
  2243. 1:43:39exist which is at
  2244. 1:43:43957.
  2245. 1:43:45Yeah. And this version is version number
  2246. 1:43:48three. Now let's say if I enter a timing
  2247. 1:43:52which looks something like this which is
  2248. 1:43:549:35
  2249. 1:43:56what is going to happen because a
  2250. 1:43:58version doesn't exist at 9:35 right so
  2251. 1:44:01let's say if we were to do something
  2252. 1:44:02like select star from this table
  2253. 1:44:05timestamp as of this timing what would
  2254. 1:44:10happen
  2255. 1:44:14so you see that it gives us some
  2256. 1:44:18results. Yeah, it gives us some results.
  2257. 1:44:20So, what it essentially does is that at
  2258. 1:44:239:35, right, let's say this is 9:31. At
  2259. 1:44:259:35, it looks whether there is a
  2260. 1:44:29version which exists at that at that
  2261. 1:44:30timing or not. If there is a version, it
  2262. 1:44:33will basically pick up that version and
  2263. 1:44:34give it to you. If no version exists, it
  2264. 1:44:37is going to find what is the latest file
  2265. 1:44:40that existed before this time and it is
  2266. 1:44:43going to find out that the latest
  2267. 1:44:44version was version number two. So that
  2268. 1:44:48is why it gives version number two which
  2269. 1:44:51is the file that resided at this point
  2270. 1:44:53in time which is 931. And we can verify
  2271. 1:44:57that. We can verify that by
  2272. 1:45:02by having a look at this file and
  2273. 1:45:05checking the customer ID which says
  2274. 1:45:07where customer ID equals 5
  2275. 1:45:11because you remember that in version
  2276. 1:45:14number two we updated the quantity to be
  2277. 1:45:16equals to 25. So this should give me 25
  2278. 1:45:20if whatever logic we are trying to make
  2279. 1:45:22up right is correct. So let's have a
  2280. 1:45:25look at that and you see that the
  2281. 1:45:27quantity equals 25. Let me also run it
  2282. 1:45:30with the latest table and we would we
  2283. 1:45:32should get different results.
  2284. 1:45:36This should give us quantity equals 1
  2285. 1:45:37because this is based out of version
  2286. 1:45:40zero as we've seen, right? But this is
  2287. 1:45:43based out of version number two. Yeah.
  2288. 1:45:47So I hope this helped you understand
  2289. 1:45:49that if you enter a timing which is not
  2290. 1:45:52present in the in in the history of the
  2291. 1:45:55table it does that mapping by finding
  2292. 1:45:57out the closest file before that
  2293. 1:46:00particular time stamp. So now that we've
  2294. 1:46:02seen all of the action in SQL let's go
  2295. 1:46:04ahead and try a few things using pi
  2296. 1:46:06spark data frame. Yeah. So I'm going to
  2297. 1:46:08write a person in python and this is
  2298. 1:46:11going to be df equals spark read.park.
  2299. 1:46:15Okay. This is not going to be par. This
  2300. 1:46:17is going to be table and this will be
  2301. 1:46:19delta catalog. Delta DB.invoices_ttv
  2302. 1:46:22invoices_ttv
  2303. 1:46:25and I want to read in a particular
  2304. 1:46:26version right so let's say we want to
  2305. 1:46:29read in version number one and this is
  2306. 1:46:31simply because it doesn't contain
  2307. 1:46:33customer ID equals one yeah just to make
  2308. 1:46:36sure we are reading in the right version
  2309. 1:46:37so the way we specify that is by doing
  2310. 1:46:41an option and we specify version as of
  2311. 1:46:45and this will be one and let me do a
  2312. 1:46:49display with a filter df.ilter
  2313. 1:46:55and for that let me first from
  2314. 1:46:58pisparksql
  2315. 1:47:01sqlf functions import column and I will
  2316. 1:47:06do a column and this column is basically
  2317. 1:47:09going to be customer id equals equals 1.
  2318. 1:47:14Yeah. And this should not return to me
  2319. 1:47:17any record.
  2320. 1:47:19So it basically not didn't return to me
  2321. 1:47:21any record and that is why it read in
  2322. 1:47:24the correct version. Now let me also
  2323. 1:47:26repeat this with version number
  2324. 1:47:30version number two and this is going to
  2325. 1:47:33show me that quantity equals 25. Yeah.
  2326. 1:47:36So customer ID equals 5 and this is
  2327. 1:47:40going to be version number two
  2328. 1:47:42and let me replace add in a post in
  2329. 1:47:45Python here.
  2330. 1:47:48So this gives me quantity equals 25. So
  2331. 1:47:50this is how you would read tables um
  2332. 1:47:54read versions of tables using power data
  2333. 1:47:57frame. And you can also replace this
  2334. 1:48:00with something like a
  2335. 1:48:02time stamp as of timestamp as of as of
  2336. 1:48:07and let me again read in this one for
  2337. 1:48:11version number two. So this should also
  2338. 1:48:13give me the same result.
  2339. 1:48:18Okay, there's some
  2340. 1:48:21percent python.
  2341. 1:48:24Great. The quantity equals 25 and it
  2342. 1:48:26gave me the same result. So now let's
  2343. 1:48:28talk about schema validation. And before
  2344. 1:48:31we talk about schema validation, in
  2345. 1:48:33order to set the context, let's talk
  2346. 1:48:35about two interesting things called
  2347. 1:48:37schema on read and schema on write.
  2348. 1:48:42Yeah. So in order to understand schema
  2349. 1:48:45on read we'll take the data lake
  2350. 1:48:47context. Yeah. So inside of data links
  2351. 1:48:50what we can do is we can simply dump in
  2352. 1:48:52any kind of data which is in any kind of
  2353. 1:48:55format and it basically goes in and
  2354. 1:48:58resides in this data lake over here. Now
  2355. 1:49:02it is kept in this data lake and it can
  2356. 1:49:05reside for as long as it wants. And when
  2357. 1:49:08we want to read in data, we simply go
  2358. 1:49:10ahead and pick up that file and read
  2359. 1:49:13that file and then apply the schema
  2360. 1:49:17and then apply the schema. So the schema
  2361. 1:49:20application process happens after the
  2362. 1:49:23data has been stored in the data lake.
  2363. 1:49:25Now if you look at schema on right, what
  2364. 1:49:28basically happens is that let's say
  2365. 1:49:29these are your incoming records.
  2366. 1:49:32First it is checked whether the schema
  2367. 1:49:36matches to your existing schema. Right?
  2368. 1:49:39So let's say there is an orders table.
  2369. 1:49:41It has five columns and those five
  2370. 1:49:43columns have some data types. Right? So
  2371. 1:49:46in order for these records to be
  2372. 1:49:49ingested or to be able to reside in the
  2373. 1:49:52data warehouse, these records have to
  2374. 1:49:54match those five columns and the five
  2375. 1:49:57data type. Yeah. So the application of
  2376. 1:50:00schema happens earlier before injection.
  2377. 1:50:05So this is where injection happened.
  2378. 1:50:07Before injection the application of
  2379. 1:50:09schema happen. Yeah. So this is schema
  2380. 1:50:12on write. Schema on read is basically
  2381. 1:50:15the injection happens earlier
  2382. 1:50:18and then the application of schema
  2383. 1:50:20happened. Now the reason why this is
  2384. 1:50:22problematic is because let's say on some
  2385. 1:50:26date you got the orders file and it
  2386. 1:50:28looks something like order ID
  2387. 1:50:33the amount and then let's say the date
  2388. 1:50:35on which the order was placed and let's
  2389. 1:50:38say after a month the order file looks
  2390. 1:50:40completely different. This is the order
  2391. 1:50:43ID and then this is going to be the
  2392. 1:50:46order details
  2393. 1:50:48and then there is going to be something
  2394. 1:50:50like the amount and the date residing
  2395. 1:50:54inside of order details and now this is
  2396. 1:50:57also ingested. Yeah. So different kinds
  2397. 1:51:00of files with different kinds of schemas
  2398. 1:51:03gets ingested. Now when you want to read
  2399. 1:51:05orders you get confused whether this is
  2400. 1:51:08the right schema or this is the right
  2401. 1:51:10schema
  2402. 1:51:12and there is going to be several
  2403. 1:51:14problems. So for example if you want to
  2404. 1:51:16find out the orders some of these file
  2405. 1:51:18you're you'll be able to easily read it
  2406. 1:51:20but in other files you'll have to do
  2407. 1:51:22some transformation basically pull it
  2408. 1:51:24out from here.
  2409. 1:51:26So the schema is not consistent it is
  2410. 1:51:29not well maintained over here. So that
  2411. 1:51:32basically helps you understand what
  2412. 1:51:34schema on read and schema on write is
  2413. 1:51:36and delta
  2414. 1:51:38is schema on write. So that helps you
  2415. 1:51:42enforce schema. It validates the schema
  2416. 1:51:45before making any data come inside of a
  2417. 1:51:49delta table. So let's understand schema
  2418. 1:51:51validation with example. Yeah. So I'm
  2419. 1:51:54going to create a table that we are
  2420. 1:51:57going to specifically use to understand
  2421. 1:51:59schema validation. And for that let me
  2422. 1:52:02simply do create or replace table
  2423. 1:52:08delta catalog dot delta db dot invoices
  2424. 1:52:14underscore sv for schema validation.
  2425. 1:52:17Yeah. So the column that I'm going to
  2426. 1:52:19have is customer ID. Uh this is going to
  2427. 1:52:23be int and I'm going to add a
  2428. 1:52:24constraint. So this constraint is going
  2429. 1:52:26to be not null. My customer ID cannot be
  2430. 1:52:29null because it is the primary key of my
  2431. 1:52:32table. And then it is going to have
  2432. 1:52:35invoice number. This is going to be a
  2433. 1:52:38string.
  2434. 1:52:40Then we are going to have quantity which
  2435. 1:52:42is going to be an integer. Then we are
  2436. 1:52:45going to have price. This is going to be
  2437. 1:52:47a let's say float. And we are going to
  2438. 1:52:52have invoice
  2439. 1:52:56date. And this is going to be a date.
  2440. 1:52:59And this is going to be the definition
  2441. 1:53:02of our table. Now we have to also insert
  2442. 1:53:06some data inside our table, right? And
  2443. 1:53:08we're going to do that using the insert
  2444. 1:53:11statement. So insert into
  2445. 1:53:14into this table
  2446. 1:53:16and we're going to use the select
  2447. 1:53:18statement. So select star from
  2448. 1:53:23we're going to use the park file that we
  2449. 1:53:25used earlier.
  2450. 1:53:27This contains all the customers for for
  2451. 1:53:30customer ids from 1 to 100 from par and
  2452. 1:53:34this is going to be the file and I'm
  2453. 1:53:37going to select the relevant columns
  2454. 1:53:39only because as we've seen earlier it
  2455. 1:53:40has a lot of column. So it is going to
  2456. 1:53:42be customer ID, invoice number,
  2457. 1:53:47quantity,
  2458. 1:53:49price
  2459. 1:53:51and invoice date. Now let's go ahead and
  2460. 1:53:55run this.
  2461. 1:53:57Let's also get a feel of the data. Let's
  2462. 1:53:59see how this looks like. Let me do a
  2463. 1:54:02limit five.
  2464. 1:54:18So now you see that we have the table
  2465. 1:54:20created, right? And just to do a few
  2466. 1:54:24checks,
  2467. 1:54:25this table should have customers
  2468. 1:54:29customer ID from 1 to 100 and a total of
  2469. 1:54:32100 customers.
  2470. 1:54:34Customer ID and the count star
  2471. 1:54:40from this table should be 100 and 100.
  2472. 1:54:45Okay. So now that we have the table set
  2473. 1:54:47up, let's understand column order
  2474. 1:54:50validation. Yeah. So let me quickly
  2475. 1:54:53write down this heading scenario one
  2476. 1:54:57column order validation. Yeah. So in
  2477. 1:55:01order to do this, we are going to run an
  2478. 1:55:03insert statement. We're going to insert
  2479. 1:55:06into the table
  2480. 1:55:09and we are simply going to do it through
  2481. 1:55:11a select statement. Select
  2482. 1:55:14these columns
  2483. 1:55:17these columns from values because we're
  2484. 1:55:21going to insert only one row as T and
  2485. 1:55:25this is going to be all of this. Now the
  2486. 1:55:27customer ID is going to be a large
  2487. 1:55:30number 99,999.
  2488. 1:55:34The invoice is going to be 1 2 3 4 I 1 2
  2489. 1:55:373 4 5. Quantity is going to be 10. Price
  2490. 1:55:40is going to be 100. and invoice ID
  2491. 1:55:43invoice date is going to be 2025 0101.
  2492. 1:55:46Yeah. Now before we do the insert let's
  2493. 1:55:48see how this looks like. So this
  2494. 1:55:51basically looks like the first customer
  2495. 1:55:54ID is 99,999
  2496. 1:55:56and all of the all of the rest right
  2497. 1:55:58quantity is 10. Now what I want to do is
  2498. 1:56:01interchange the order. Now because
  2499. 1:56:04quantity and customer id are both
  2500. 1:56:06integers I want to interchange them so
  2501. 1:56:09that there is no conflicting data types.
  2502. 1:56:11Right? Now let's run this. So the first
  2503. 1:56:14column now becomes quantity which is 10
  2504. 1:56:17and customer ID becomes the third column
  2505. 1:56:20which is 99,999.
  2506. 1:56:23So what I would desire is that if the
  2507. 1:56:26insert statement happened correctly, a
  2508. 1:56:28new row gets inserted where customer ID
  2509. 1:56:31equals 99,999.
  2510. 1:56:33Yeah. And quantity equals 10. So let's
  2511. 1:56:35go ahead and run this.
  2512. 1:56:38And meanwhile, let me also write select
  2513. 1:56:40star from
  2514. 1:56:42okay. So this completed successfully.
  2515. 1:56:45Let me go ahead and write this where
  2516. 1:56:48customer ID equals 99,999.
  2517. 1:56:53Yeah. So this is what we wanted. A row
  2518. 1:56:56which has customer ID 99,999.
  2519. 1:57:00That should be in the table because we
  2520. 1:57:02just ran an insert statement. Let's go
  2521. 1:57:04ahead and run this. What do you think?
  2522. 1:57:05Would we find that row there?
  2523. 1:57:09Okay. So that row is not present there.
  2524. 1:57:12Where did it go? What happened? What
  2525. 1:57:14just happened to my table?
  2526. 1:57:17So what we see over here is
  2527. 1:57:21let me go ahead and run this again.
  2528. 1:57:24What we see over here is the first
  2529. 1:57:27column was 10. So what actually happened
  2530. 1:57:30was that it took in the columns by
  2531. 1:57:33position and it mapped to the customer
  2532. 1:57:36ID column 0 column number zero column
  2533. 1:57:38number zero. So I'm going to take in the
  2534. 1:57:40value from this column which is at
  2535. 1:57:43position number zero and put it inside
  2536. 1:57:46customer ID. Yeah. So this is something
  2537. 1:57:48called column matching by position. So
  2538. 1:57:53let me go ahead and see if we have
  2539. 1:57:57where
  2540. 1:57:59customer ID equals 10. So we already had
  2541. 1:58:03customer a customer ID equals 10, right?
  2542. 1:58:05because we inserted data from 1 to 100.
  2543. 1:58:09So that means now after this insert
  2544. 1:58:12there should be two rows.
  2545. 1:58:16There you go. So we have two rows which
  2546. 1:58:19is customer ID equals 10 and invoice
  2547. 1:58:22number 1 2 3 4 5. Quantity is 99,999
  2548. 1:58:26and the rest of what we inserted. So
  2549. 1:58:28what insert did is that it executed the
  2550. 1:58:34insertion by matching columns by
  2551. 1:58:36position and not the name. And that is
  2552. 1:58:40why you see what you see over here. And
  2553. 1:58:42this actually corrupts your data. So
  2554. 1:58:44never use an insert unless you're very
  2555. 1:58:47sure that the positions will always be
  2556. 1:58:51maintained. So let's see what is going
  2557. 1:58:53to happen if we reorder the columns and
  2558. 1:58:57insert that data into our table using a
  2559. 1:59:01merge statement. Is it going to behave
  2560. 1:59:03any differently? Yeah. So let's try that
  2561. 1:59:05out. And the source data that we are
  2562. 1:59:08going to use is all the customer ids.
  2563. 1:59:11Actually not all the customer ids. Uh
  2564. 1:59:14the customer ids from 101 to 200. And
  2565. 1:59:18I'm just going to take five of them.
  2566. 1:59:21order by customer ID descending limit
  2567. 1:59:25five. So this is going to give me five
  2568. 1:59:27customer ids from 196 to 200. Yeah. So
  2569. 1:59:32let's go ahead and write the merge
  2570. 1:59:33statement. Merge into
  2571. 1:59:36this table as target using
  2572. 1:59:40this source table.
  2573. 1:59:44Using this source table,
  2574. 1:59:47the join condition is going to be on
  2575. 1:59:50target dot customer ID equals source dot
  2576. 1:59:52customer ID and when not matched
  2577. 1:59:58then
  2578. 2:00:00simply insert star. So these five rows
  2579. 2:00:04they are not going to match with the
  2580. 2:00:06existing data that we have in the table.
  2581. 2:00:09Right? because the existing data is for
  2582. 2:00:11customers ids from 1 to 100 right so
  2583. 2:00:14these are not going to match and they
  2584. 2:00:17are supposed to be inserted in the table
  2585. 2:00:19yeah now what happens if you run this
  2586. 2:00:22statement again right again and again so
  2587. 2:00:24the first time it gets inserted the
  2588. 2:00:26second time you run it it is going to
  2589. 2:00:28match this statement because of this
  2590. 2:00:30statement target customer id equals sort
  2591. 2:00:32customer ID they are going to match
  2592. 2:00:34right because 196 to 200 already resides
  2593. 2:00:37in our table so in that case What we are
  2594. 2:00:39going to do is that when matched then
  2595. 2:00:43update
  2596. 2:00:44set the following
  2597. 2:00:47column right
  2598. 2:00:50set customer id dot customer ID equals
  2599. 2:00:54source dot customer id target dot
  2600. 2:00:58invoice number equals s.invoice invoice
  2601. 2:01:01number the quantity
  2602. 2:01:03target dot price equals source.p price
  2603. 2:01:08and the final one target dot invoice
  2604. 2:01:11date equals the current date. So if they
  2605. 2:01:16match what we are simply going to do is
  2606. 2:01:18we are going to keep all of the column
  2607. 2:01:20the same. We just going to update the
  2608. 2:01:22invoice date. Now let's go ahead and
  2609. 2:01:25make the change. We are going to swap
  2610. 2:01:28the columns over here.
  2611. 2:01:30Quantity going to be the first column
  2612. 2:01:32and customer ID going to be the third
  2613. 2:01:34column. Exactly same as what we did over
  2614. 2:01:37here. Quantity would the first column
  2615. 2:01:39and customer ID what the third column.
  2616. 2:01:41And let me also do one change because
  2617. 2:01:43this is coming from this park file.
  2618. 2:01:45Price over here is a float, right? So
  2619. 2:01:49this should also be a float.
  2620. 2:01:53Cast this as float. Yeah. So now let's
  2621. 2:01:57go ahead and run this. So while this
  2622. 2:02:00runs,
  2623. 2:02:01let me write some SQL to quickly
  2624. 2:02:03validate this. Select star from this
  2625. 2:02:05table where
  2626. 2:02:08where customer ID greater than 100,
  2627. 2:02:10right? And the only customer ID is
  2628. 2:02:12greater than 100 in this table should be
  2629. 2:02:14these ones, right? These five record. So
  2630. 2:02:17we run this
  2631. 2:02:20and there you go. So you see that the
  2632. 2:02:23customer ids from 196 to 100 have been
  2633. 2:02:25returned. That mean that the merge
  2634. 2:02:27statement ran correctly. It was able to
  2635. 2:02:30identify which columns to match. Now a
  2636. 2:02:33very important point to note is that
  2637. 2:02:35insert the matches columns by position.
  2638. 2:02:38The zero position is going to be matched
  2639. 2:02:40to the zero position from the source.
  2640. 2:02:43But what merge does is that it matches
  2641. 2:02:46column by name. It finds the right name
  2642. 2:02:48even if they are not ordered correctly
  2643. 2:02:51and then matches them. And that is the
  2644. 2:02:52reason why this worked. Yeah. So now
  2645. 2:02:55let's go ahead and run this once again,
  2646. 2:02:57right? So that it comes to this clause
  2647. 2:03:00and then it just updates the invoice
  2648. 2:03:03date. Yeah. So let's go ahead and run
  2649. 2:03:04this now.
  2650. 2:03:07And ideally
  2651. 2:03:09this is the only column that should
  2652. 2:03:11change.
  2653. 2:03:14Okay, let's order this from 200. And
  2654. 2:03:17there you go. You see that this is the
  2655. 2:03:19only column that changed while all the
  2656. 2:03:21other columns are just the same. So
  2657. 2:03:24merge seemed to work just fine. So now
  2658. 2:03:27let's try to understand how delta does
  2659. 2:03:29data type validation. Right? How it does
  2660. 2:03:33data type validation. And this is going
  2661. 2:03:35to be scenario number two. So let me
  2662. 2:03:39quickly write that down. Data type
  2663. 2:03:43validation.
  2664. 2:03:46And
  2665. 2:03:48we're going to test this by writing an
  2666. 2:03:50insert statement. So we're going to say
  2667. 2:03:52insert into delta catalog do this table
  2668. 2:03:56and the values are going to be
  2669. 2:04:01let's say the customer ID is going to be
  2670. 2:04:03ABC
  2671. 2:04:04invoice number is going to be I 4 5 6 7
  2672. 2:04:088
  2673. 2:04:10quantity is going to be 10 price is
  2674. 2:04:12going to be 98.75
  2675. 2:04:15and the invoice date is going to be
  2676. 2:04:1720250101
  2677. 2:04:19yeah now if I were to ask you whether
  2678. 2:04:22this piece of code will run or not. I
  2679. 2:04:24think the obvious answer would be no
  2680. 2:04:26because the data type for customer ID is
  2681. 2:04:30integer but what we are trying to feed
  2682. 2:04:32into it is is string. So let's go ahead
  2683. 2:04:36and run this and as you expected it
  2684. 2:04:39didn't work because this is string and
  2685. 2:04:40it tried to convert it to integer but it
  2686. 2:04:43didn't work. Yeah. But now let's go
  2687. 2:04:46ahead and do something different right.
  2688. 2:04:48So let me go ahead and write this
  2689. 2:04:51customer ID which is 99499.
  2690. 2:04:55Now what do you think? Will it run?
  2691. 2:04:58Let's go ahead and run this.
  2692. 2:05:02Okay, there you go. So it seems that it
  2693. 2:05:05ran.
  2694. 2:05:07Let's do a select star from
  2695. 2:05:11delta catalog. So the intelligent
  2696. 2:05:14sometimes works and sometime doesn't.
  2697. 2:05:17It's completely dependent on it mode uh
  2698. 2:05:20where customer ID equals this right.
  2699. 2:05:25So you see that there is a row 99499 and
  2700. 2:05:28it contains the exact same details 98.75
  2701. 2:05:32and quantity. So the question here is
  2702. 2:05:34why did this work at all if this did
  2703. 2:05:37this didn't? So the reason why this
  2704. 2:05:39worked is because
  2705. 2:05:42what delta does is that it does try it
  2706. 2:05:45puts in effort to convert it into the
  2707. 2:05:48form that it actually exists in the
  2708. 2:05:50table. So in the table it is in integer
  2709. 2:05:53form. It tries to convert it to an
  2710. 2:05:56integer form. Yeah. If the conversion
  2711. 2:05:59happens successfully, it inserts the
  2712. 2:06:02data into the table. If the conversion
  2713. 2:06:04doesn't happen, then it throws an error.
  2714. 2:06:06So that is why you see the error thrown
  2715. 2:06:09over here is that string cannot be cast
  2716. 2:06:12to int. Yeah. So it tried to cast it but
  2717. 2:06:14it cannot cast it to int. And that is
  2718. 2:06:17also the reason why this 2025 0101 this
  2719. 2:06:22is string but the data that we have over
  2720. 2:06:24here is date. So it casted string to
  2721. 2:06:28date and then inserted the data in this
  2722. 2:06:32table. Yeah. But this also makes sure
  2723. 2:06:35that we just cannot dump in any garbage
  2724. 2:06:37data. So that check is there in place.
  2725. 2:06:40Let's have a look at another scenario
  2726. 2:06:42which is column name validation. Yeah.
  2727. 2:06:45So okay, it seemed I missed writing
  2728. 2:06:49column name validation here. Column name
  2729. 2:06:52validation. And this is going to be
  2730. 2:06:54number five. number six and this is
  2731. 2:06:57scenario number three right so let's go
  2732. 2:07:00ahead and add a heading scenario number
  2733. 2:07:04three column name validation
  2734. 2:07:08and the way we are going to test this is
  2735. 2:07:11simply I'm going to copy the code that I
  2736. 2:07:14use for column order validation and we
  2737. 2:07:17are first going to test this with an
  2738. 2:07:19insert statement
  2739. 2:07:21so let's make sure that the order is the
  2740. 2:07:24name
  2741. 2:07:25because there's no point changing the
  2742. 2:07:27order, right? We've already tested for
  2743. 2:07:29order
  2744. 2:07:31and let's go ahead and change the column
  2745. 2:07:34name.
  2746. 2:07:36This is going to be as QTY.
  2747. 2:07:40So now we've changed customer ID to C ID
  2748. 2:07:43quantity to QTY. Yeah. Let's go ahead
  2749. 2:07:46and run this now.
  2750. 2:07:50So if this SQL ran correctly,
  2751. 2:07:54you should see something like a row, a
  2752. 2:07:58row which basically contains customer ID
  2753. 2:08:01equal this number.
  2754. 2:08:04Okay, there you go. So you see this row
  2755. 2:08:06where customer ID equals 99,999
  2756. 2:08:09and all of these other details, right?
  2757. 2:08:12So the reason why this worked even
  2758. 2:08:15though the column names are different is
  2759. 2:08:18because because of a reason that we
  2760. 2:08:19already discussed some time back right
  2761. 2:08:21so it does column matching by position
  2762. 2:08:24and it doesn't matter whatever in the
  2763. 2:08:27world the name that you decide to put
  2764. 2:08:29over here it doesn't matter and that
  2765. 2:08:31will be of no effect. So that is the
  2766. 2:08:33reason why this worked. Now let's try
  2767. 2:08:35the same thing with a merge statement.
  2768. 2:08:40So last time we used
  2769. 2:08:44this data over here.
  2770. 2:08:47This basically contained customer ids
  2771. 2:08:51from okay the order is changed. Let me
  2772. 2:08:55get this back to the same order.
  2773. 2:09:02So it contain customer ID from 196 to
  2774. 2:09:04200. Uh let me change it
  2775. 2:09:08to have from 101 to 105. Yeah. Now let's
  2776. 2:09:12go ahead and insert this data.
  2777. 2:09:18And before before actually we run this
  2778. 2:09:20query, we want to change we want to
  2779. 2:09:24change the name of the column, right? So
  2780. 2:09:26we want to change this to qty.
  2781. 2:09:30Uh where are the other places? C id
  2782. 2:09:34id qty. Right? So now that we've made
  2783. 2:09:38all of the changes, right? Let's go
  2784. 2:09:40ahead and run this and see what happens.
  2785. 2:09:42So if this statement runs successfully,
  2786. 2:09:44what what we should see is
  2787. 2:09:47when we run something like this, right?
  2788. 2:09:50If we run something like this,
  2789. 2:09:54I should see five more rows.
  2790. 2:09:57So I should see these rows as well. So
  2791. 2:10:00let's go ahead and run this now.
  2792. 2:10:04So now what it says is cannot resolve
  2793. 2:10:07customer ID in insert clause given
  2794. 2:10:09columns is source C ID. So the reason is
  2795. 2:10:15pretty simple and it is again something
  2796. 2:10:17that we already discussed. So think
  2797. 2:10:19about it for a minute if you're not able
  2798. 2:10:21to recall why it failed. So the reason
  2799. 2:10:24why it failed is simply because merge
  2800. 2:10:27does column matching by name. So it
  2801. 2:10:30looks at the target table. It takes up a
  2802. 2:10:33column let's say customer ID. It looks
  2803. 2:10:35up for that exact column in the source
  2804. 2:10:38table. Okay. Do I find customer ID in
  2805. 2:10:40the source table or not? If it finds it
  2806. 2:10:43then it is going to dump all of that
  2807. 2:10:44data in the target table. But if it is
  2808. 2:10:47not able to find it, it is going to
  2809. 2:10:49fail. So it is quite a robust statement.
  2810. 2:10:53So now let's have a look at another
  2811. 2:10:54scenario which is nullability
  2812. 2:10:56validation. Yeah. So let me quickly
  2813. 2:11:01add in the heading
  2814. 2:11:05scenario four nullability
  2815. 2:11:08validation. And here uh by nullability I
  2816. 2:11:11just want I just don't want to talk
  2817. 2:11:12about nullability but a few other things
  2818. 2:11:15as well. So let's have a look at this.
  2819. 2:11:16insert into
  2820. 2:11:19delta catalog dot
  2821. 2:11:22this invoices table values we have how
  2822. 2:11:25many columns 1 2 3 4 5 so there's going
  2823. 2:11:27to be five nulls and if I were to just
  2824. 2:11:32copy this
  2825. 2:11:35and paste it five times and if I were to
  2826. 2:11:38ask you whether this will run or not
  2827. 2:11:39obviously you would say that this won't
  2828. 2:11:41run
  2829. 2:11:43so the reason why it didn't run is
  2830. 2:11:44because the not null constraint raint
  2831. 2:11:46violated for column customer id. So we
  2832. 2:11:50had a constraint over here customer ID
  2833. 2:11:53int not null whatever customer id we
  2834. 2:11:56insert and for now there can be
  2835. 2:11:57duplicates right uh but they shouldn't
  2836. 2:12:00be null yeah so that constraint was
  2837. 2:12:03violated over here and that was the
  2838. 2:12:05reason why it wasn't so if we run the
  2839. 2:12:07same statement right if we run the same
  2840. 2:12:10statement inserting some random number 7
  2841. 2:12:138 912 this will work
  2842. 2:12:19And as expected this works. So the point
  2843. 2:12:22that I'm trying to make over here is not
  2844. 2:12:25just nullability validation. It's about
  2845. 2:12:27constraints validating the constraints
  2846. 2:12:30that we've put. Right? So we put in a
  2847. 2:12:33constraint which is not null over here.
  2848. 2:12:36We can also put in constraints like for
  2849. 2:12:39example the price should be greater than
  2850. 2:12:42zero. The quantity should be greater
  2851. 2:12:44than zero. Something like that. We can
  2852. 2:12:45put in rules over here. And whenever
  2853. 2:12:48data is inserted all of those rules
  2854. 2:12:51would be checked for and this would make
  2855. 2:12:54sure that the data that is getting
  2856. 2:12:56inside our tables are correct they are
  2857. 2:12:58accurate. Now let's have a look at the
  2858. 2:13:00last scenario which is extra column
  2859. 2:13:03validation. What happens if in my source
  2860. 2:13:06the data that I'm trying to insert it
  2861. 2:13:08has extra columns and the target table
  2862. 2:13:11does not have those many columns. Right?
  2863. 2:13:12So what is going to be the case? So this
  2864. 2:13:15is going to be
  2865. 2:13:18scenario
  2866. 2:13:20number five and this is going to be
  2867. 2:13:22extra column validation.
  2868. 2:13:26So let me copy some code. I'm going to
  2869. 2:13:29copy the code that I used for column
  2870. 2:13:32order validation
  2871. 2:13:34and
  2872. 2:13:36I'll comment this out for now. And let
  2873. 2:13:38me add in an extra column. Right? And
  2874. 2:13:40this is a dummy column. Let me name it
  2875. 2:13:42as customer type. So you see that all of
  2876. 2:13:46the columns over here are intact. And
  2877. 2:13:49actually let me also change
  2878. 2:13:52correct the orders.
  2879. 2:13:55Customer ID should be over here.
  2880. 2:13:58Yeah. So the orders have been corrected
  2881. 2:14:01now. And let's also change this. Right?
  2882. 2:14:06So the orders have been corrected now.
  2883. 2:14:07And we have an extra column which is
  2884. 2:14:09customer type. Now let's see what is
  2885. 2:14:11going to happen.
  2886. 2:14:14So you see that this didn't work. It
  2887. 2:14:17says that a schema mishmash detected
  2888. 2:14:20when writing to the delta table. Now
  2889. 2:14:23let's go ahead and try this with a merge
  2890. 2:14:26statement. Yep.
  2891. 2:14:29So if I were to
  2892. 2:14:32copy this
  2893. 2:14:34this merge statement over here and let's
  2894. 2:14:37go ahead and run this. Actually, let's
  2895. 2:14:39change the source data a little bit.
  2896. 2:14:41Right. So now I'm going to take up all
  2897. 2:14:45of the customers whose ID is between
  2898. 2:14:49where customer ID between 150 and 155.
  2899. 2:14:54And
  2900. 2:14:56I want the C the the order of columns to
  2901. 2:15:00be the same.
  2902. 2:15:02So let's go ahead and run this. So this
  2903. 2:15:05gives me all the customers whose ID is
  2904. 2:15:07between 150 to 155. Right? So let's go
  2905. 2:15:11ahead and paste this over here.
  2906. 2:15:14And let's also add in an extra column
  2907. 2:15:16over here which is VIP as customer type.
  2908. 2:15:22Now these rows are not in the table,
  2909. 2:15:26right? So it should go to this clause
  2910. 2:15:28when not mash then insert star. So all
  2911. 2:15:30of them should be inserted. And because
  2912. 2:15:32I want to test that particular case
  2913. 2:15:33right now, let me just remove all of
  2914. 2:15:36that for simplicity. So let's go ahead
  2915. 2:15:37and run this
  2916. 2:15:40and actually before I run this uh let me
  2917. 2:15:44show you that there is
  2918. 2:15:48there is no record. Select start from
  2919. 2:15:52this table. This would give me no
  2920. 2:15:54record. Now let's go ahead and run this
  2921. 2:15:55now.
  2922. 2:15:58And you see that it succeeded.
  2923. 2:16:01Let's run this.
  2924. 2:16:03Great.
  2925. 2:16:05So I see all of the records from 150 to
  2926. 2:16:09155 that means that the insert statement
  2927. 2:16:12has worked correctly. So it tries to
  2928. 2:16:15match the columns by their names and if
  2929. 2:16:19it's not able to find a name. So
  2930. 2:16:21basically it started with the target. It
  2931. 2:16:24searched for all the five names. It was
  2932. 2:16:25able to find them and then it just went
  2933. 2:16:28off right. It succeeded and the
  2934. 2:16:30operation completed. Yeah. So the merge
  2935. 2:16:33operation is actually quite a robust way
  2936. 2:16:36of doing things, right? It was able to
  2937. 2:16:38handle a few other issues earlier as
  2938. 2:16:41well. Now let's understand what schema
  2939. 2:16:43evolution is, right? Schema evolution is
  2940. 2:16:47the ability to handle changes in schema
  2941. 2:16:50without having to completely rewrite
  2942. 2:16:53your table or without having to rewrite
  2943. 2:16:56the underlying data. Yeah. So when I say
  2944. 2:17:00ability to handle changes, what I mean
  2945. 2:17:02is that let's say if I have a table, if
  2946. 2:17:04a new column comes in, I should be able
  2947. 2:17:06to accommodate that new column. Or let's
  2948. 2:17:10say if I have a column which is of type
  2949. 2:17:13int. Let's say order ID is of type int
  2950. 2:17:16and for some reason I'm getting a lot of
  2951. 2:17:18orders and my order ID has gone beyond
  2952. 2:17:22the range of integer and I want to
  2953. 2:17:24change it to long or big int or
  2954. 2:17:26something like that. Right? So I should
  2955. 2:17:27be able to upgrade my data type to
  2956. 2:17:30another type. So all of that flexibility
  2957. 2:17:33should be allowed. So we are going to
  2958. 2:17:35have a look at four scenario. The first
  2959. 2:17:37one is adding new columns. One is the
  2960. 2:17:40manual way, the other one is the
  2961. 2:17:41automatic way. So we'll be having a look
  2962. 2:17:43at both of them. The second one as I
  2963. 2:17:45discussed is widening of data types. The
  2964. 2:17:49third one is nested structure evolution.
  2965. 2:17:51The last one is column position changes.
  2966. 2:17:54If for whatever reason I want to reorder
  2967. 2:17:56the columns in my table, there should be
  2968. 2:17:58enough flexibility to allow me to do
  2969. 2:18:00that. Let's go ahead and understand the
  2970. 2:18:02first scenario which is adding new
  2971. 2:18:05column.
  2972. 2:18:06Let me first quickly write in
  2973. 2:18:10scenario number one which is adding
  2974. 2:18:13adding new columns. And for this we are
  2975. 2:18:17going to create a brand new table and
  2976. 2:18:19for that I'll copy the code from from
  2977. 2:18:22schema validation. Uh so let me just
  2978. 2:18:25remove this line. Change this to SE for
  2979. 2:18:28schema evolution.
  2980. 2:18:31Quantity is going to be removed because
  2981. 2:18:32I want to keep minimal number of columns
  2982. 2:18:34and minimal number of rows as well for
  2983. 2:18:37this example. So let me do a where
  2984. 2:18:41customer ID between 1 and five. So we'll
  2985. 2:18:45insert five rows into this table.
  2986. 2:18:48And let's also remove quantity. Let's go
  2987. 2:18:52ahead and run this. Quite simple. We've
  2988. 2:18:55got what we expected. And let me put
  2989. 2:18:58this over here. Let's go ahead and run
  2990. 2:19:00this. Okay, my bad. This should be se.
  2991. 2:19:06Let's go ahead and run this now.
  2992. 2:19:09And let's quickly validate
  2993. 2:19:12that
  2994. 2:19:14the table looks just as we expected it
  2995. 2:19:18to look like. And it looks as expected.
  2996. 2:19:22Yeah. So the table is set up now. Now
  2997. 2:19:25let's go ahead and add a column to the
  2998. 2:19:28table.
  2999. 2:19:29So we are going to run an alter table
  3000. 2:19:36invoiced SC add column and the column
  3001. 2:19:40that we going to have is let's first see
  3002. 2:19:42how many columns are there like what all
  3003. 2:19:44what all columns are there. So for that
  3004. 2:19:46let me run a select star and we see that
  3005. 2:19:49there are okay we'll use the quantity
  3006. 2:19:51column that we just dropped we didn't
  3007. 2:19:53want to select that so let's use that
  3008. 2:19:55for now and this is going to be quantity
  3009. 2:19:57integer
  3010. 2:20:00so this column has been added to the
  3011. 2:20:01schema of the table now now I want to
  3012. 2:20:07insert a few more rows and this time it
  3013. 2:20:11is going to be with the quantity column
  3014. 2:20:14And let's insert rows from 6 to 10 now.
  3015. 2:20:20So let's go ahead and run this.
  3016. 2:20:24Let me copy this to see how the final
  3017. 2:20:26table looks like. So the final table
  3018. 2:20:28should look like uh six five columns in
  3019. 2:20:31place with these rows not having the
  3020. 2:20:35quantity data. They should all be null.
  3021. 2:20:37And the one that we inserted right now,
  3022. 2:20:39they should all have non-null data for
  3023. 2:20:42quantity. Let's run this.
  3024. 2:20:46And there you go. As we expected, we see
  3025. 2:20:49non-null data for quantity for the
  3026. 2:20:52customer ids that we've just inserted
  3027. 2:20:54and null quantity for the one that were
  3028. 2:20:57inserted earlier. So you see that we had
  3029. 2:21:00to run an alter statement over here in
  3030. 2:21:03order to accommodate this new column
  3031. 2:21:06that was coming in. Now what if you want
  3032. 2:21:09to do this automatically?
  3033. 2:21:11Although there can be situations where
  3034. 2:21:13we want to go ahead with this approach.
  3035. 2:21:15Sometimes we don't want to go ahead
  3036. 2:21:17because we don't randomly want any kind
  3037. 2:21:19of data to come in and then sit in my
  3038. 2:21:20table. Right? But let's say for this
  3039. 2:21:22situation, what if you want to
  3040. 2:21:24automatically accommodate a new column?
  3041. 2:21:27So for that case, you need to set this
  3042. 2:21:32property to true and this enables schema
  3043. 2:21:36evolution. So let's go ahead and run
  3044. 2:21:38this now. So this is enabled to true.
  3045. 2:21:41And let me run another insert statement.
  3046. 2:21:46This time without doing an alter without
  3047. 2:21:48running an alter table command but I
  3048. 2:21:52will add another column. So let's see
  3049. 2:21:54which column do I want to add. So let's
  3050. 2:21:58add payment method. Let's add payment
  3051. 2:22:00method. So let's add
  3052. 2:22:04payment method over here.
  3053. 2:22:08And this is going to be 11 to 15. And
  3054. 2:22:13this is going to be inserted in the
  3055. 2:22:15table. The table had how many? It had 1
  3056. 2:22:172 3 4 5 column. Now if this works well,
  3057. 2:22:21we are going to have six columns. Right?
  3058. 2:22:23And the payment method column is going
  3059. 2:22:25to be populated only for customers ids
  3060. 2:22:28from 11 to 15. So let's run this. and we
  3061. 2:22:33do a select star from delta catalog blah
  3062. 2:22:38blah blah and let's run this. Okay,
  3063. 2:22:41there you go. So you see that from 11 to
  3064. 2:22:4415 we have the payment method populated.
  3065. 2:22:50That means that this statement worked
  3066. 2:22:52pretty well and automatic schema
  3067. 2:22:55evolution had just happened. You didn't
  3068. 2:22:57need to run an alter statement in order
  3069. 2:23:00to make it. Now let's talk about the
  3070. 2:23:02second scenario which is type widening
  3071. 2:23:05right and this was introduced in delta
  3072. 2:23:08version 3.2. This simply means that you
  3073. 2:23:11can upgrade a type to its bigger type.
  3074. 2:23:15Yeah. Simply meaning that let's say if
  3075. 2:23:17you have an int you want to convert it
  3076. 2:23:19to a big int for whatever reason. The
  3077. 2:23:21example that we discussed, if our orders
  3078. 2:23:24have spanned to such a large number that
  3079. 2:23:27it is not able to fit in the integer
  3080. 2:23:30data type, we want to upgrade it to a
  3081. 2:23:32big array. Similarly, you want to
  3082. 2:23:34upgrade a float to a double or a var of
  3083. 2:23:37a specified number of characters to more
  3084. 2:23:40number of characters. All of that is
  3085. 2:23:42possible using type widening. So before
  3086. 2:23:45we get started with example, we need to
  3087. 2:23:48ensure that we at least have delta
  3088. 2:23:51version 3.2. And if you remember during
  3089. 2:23:54the initial parts of the video, I
  3090. 2:23:57created a cluster.
  3091. 2:23:59The cluster had a datab bricks runtime
  3092. 2:24:02version of 14.3. And I've pulled this up
  3093. 2:24:06to show you
  3094. 2:24:08that
  3095. 2:24:10if I were to go to 14.3
  3096. 2:24:13and quickly search for delta,
  3097. 2:24:17this is going to be 3.1.0.
  3098. 2:24:20That means this DBR database runtime is
  3099. 2:24:23not going to work. So let me quickly
  3100. 2:24:26check 15.4 and what does that show me?
  3101. 2:24:30So this shows me 3.2.0.
  3102. 2:24:33That means this is going to work. So,
  3103. 2:24:34let me go ahead and upgrade this to
  3104. 2:24:3815.4.
  3105. 2:24:39Keeping all of the other stuff the same.
  3106. 2:24:42Let's click on confirm and let's start
  3107. 2:24:45the cluster. Okay. So, now that we have
  3108. 2:24:48our cluster ready,
  3109. 2:24:51let's go ahead and start with the
  3110. 2:24:52example. And before that, let me quickly
  3111. 2:24:54write down scenario
  3112. 2:24:57scenario number two which is type
  3113. 2:25:00widening.
  3114. 2:25:02And we also need to do another thing
  3115. 2:25:04which is enabling type widening on this
  3116. 2:25:07table. So let me how to enable type
  3117. 2:25:13widening on a delta table.
  3118. 2:25:17Let's quickly search for that.
  3119. 2:25:23Okay. So this is the property that I
  3120. 2:25:25need to enable and let's go ahead and
  3121. 2:25:28run this. Hopefully this should work.
  3122. 2:25:32And meanwhile, let me go ahead and write
  3123. 2:25:35an insert statement.
  3124. 2:25:37Insert into
  3125. 2:25:40this table. Actually, before that, let's
  3126. 2:25:43see how the columns currently look like.
  3127. 2:25:45So, describe this table delta catalog
  3128. 2:25:49dot invoice delta DB
  3129. 2:25:55dot invoices SA. Right? How do the
  3130. 2:25:57columns look like currently? So customer
  3131. 2:26:00ID is an integer right now and we have
  3132. 2:26:03enabled
  3133. 2:26:05automatic schema evolution right. So let
  3134. 2:26:09me go ahead and write an insert
  3135. 2:26:12statement where values
  3136. 2:26:15this I'll copy some values from here
  3137. 2:26:20and
  3138. 2:26:22this is a string this is also a string
  3139. 2:26:26comma comma this is going to be a string
  3140. 2:26:29yeah and this is an integer so let me
  3141. 2:26:31paste it over here and I want to insert
  3142. 2:26:35a number which is a big integer.
  3143. 2:26:39Although you see that I have not
  3144. 2:26:40currently changed the data type but I
  3145. 2:26:44still want to insert it and check
  3146. 2:26:45whether it works or not because I have
  3147. 2:26:47automatic schema evolution enabled. So
  3148. 2:26:50we want to try it out whether it works
  3149. 2:26:52or not. So let's write give me an
  3150. 2:26:55example
  3151. 2:26:58of a big number
  3152. 2:27:02and that is the number that I'm going to
  3153. 2:27:05use over here.
  3154. 2:27:07So I have six columns,
  3155. 2:27:10six numbers over here. And let's go
  3156. 2:27:12ahead and run this. 36 values, not
  3157. 2:27:16numbers. Okay, I missed a comma here.
  3158. 2:27:18Let's run this now.
  3159. 2:27:21Okay, so the error it throws is fail to
  3160. 2:27:24assign a value big int to type int.
  3161. 2:27:28Right? That simply means that using
  3162. 2:27:33automatic schema evolution type widening
  3163. 2:27:36won't work. We will have to manually
  3164. 2:27:39change the type of customer ID. Yeah. So
  3165. 2:27:42in order to do that let's go ahead and
  3166. 2:27:44run alter table
  3167. 2:27:48invoices se and we will write alter
  3168. 2:27:51column
  3169. 2:27:53alter column customer
  3170. 2:27:56customer id and the type that we want is
  3171. 2:28:01big int. Let's go ahead and run this
  3172. 2:28:06that is successful. So we say describe
  3173. 2:28:08table
  3174. 2:28:11invoices
  3175. 2:28:12sc and now you see that the customer ID
  3176. 2:28:15is big int earlier it was integer. So
  3177. 2:28:17let's go ahead and run the same
  3178. 2:28:19statement again. Let's see what happened
  3179. 2:28:29like star from.
  3180. 2:28:32Okay this seems to have succeeded.
  3181. 2:28:35And if this actually succeeded, that
  3182. 2:28:37means that this table should have
  3183. 2:28:43should have a row where customer ID
  3184. 2:28:45equals this value.
  3185. 2:28:53Perfect. That means this row resized in
  3186. 2:28:56the table and type widening has
  3187. 2:28:58successfully worked and we've enabled
  3188. 2:29:00type widening. Next scenario that we are
  3189. 2:29:02going to talk about is nested structure
  3190. 2:29:05evolution.
  3191. 2:29:08Scenario number three is nested
  3192. 2:29:12structure evolution.
  3193. 2:29:16So let's add a strct column to our
  3194. 2:29:20table. So in order to do that alter
  3195. 2:29:22table
  3196. 2:29:25the invoices are se table add column and
  3197. 2:29:29let's go ahead and add something called
  3198. 2:29:31purchase details and this is going to be
  3199. 2:29:34a strct.
  3200. 2:29:36The strct is going to have two column.
  3201. 2:29:38The first one is mall pin code. The mall
  3202. 2:29:42at which the purchase was made and this
  3203. 2:29:44is going to be an integer and then the
  3204. 2:29:47store code. the store at which the
  3205. 2:29:49purchase was made. So let's go ahead and
  3206. 2:29:51run this and let's add some sample data
  3207. 2:29:55to this table. So
  3208. 2:29:59this is going to be so we inserted
  3209. 2:30:02around 15 records.
  3210. 2:30:05So this is going to be the 16th one. So
  3211. 2:30:08let's add 16th.
  3212. 2:30:10And then we put a strct over here. the
  3213. 2:30:14mall pin code. Let's let's put in some
  3214. 2:30:16random number and let's put in some
  3215. 2:30:20random number for the store code. Yeah,
  3216. 2:30:22let's go ahead and run this. So, this
  3217. 2:30:24should run just fine. So, let's do
  3218. 2:30:27select star from
  3219. 2:30:30delta catalog blah blah blah. And
  3220. 2:30:36perfect. So this should be the only row
  3221. 2:30:38where you had purchase details and it
  3222. 2:30:40has all of
  3223. 2:30:43the data that we just put in the mall
  3224. 2:30:45pin code and the store code. Now let's
  3225. 2:30:47consider two examples. What if you need
  3226. 2:30:49to change
  3227. 2:30:51the data type of mall pin code for
  3228. 2:30:54whatever reason from int to begin and
  3229. 2:30:57we're going to follow a pretty similar
  3230. 2:30:59way.
  3231. 2:31:01Alter table. This table. Alter column.
  3232. 2:31:06Alter column. Uh
  3233. 2:31:10purchase details dot mall pin code.
  3234. 2:31:16The type that we want it to be is big
  3235. 2:31:19in. Let's go ahead and run this.
  3236. 2:31:22And let's insert
  3237. 2:31:25some more data to validate this. And
  3238. 2:31:27this is going to be number 17.
  3239. 2:31:31Let's add in a strct over here. And
  3240. 2:31:38the store code can be any number. But
  3241. 2:31:41for the mall pin code, I need a big int.
  3242. 2:31:44So I'm going to put that number over
  3243. 2:31:45here. Let's just increase it by one. Uh
  3244. 2:31:49and this number can just be any other
  3245. 2:31:51number, right? So let's go ahead and run
  3246. 2:31:52this.
  3247. 2:31:54And let's again do a select star from
  3248. 2:31:59the table.
  3249. 2:32:04And
  3250. 2:32:05there you go. So this completely works
  3251. 2:32:08fine. Now it is able to accommodate big
  3252. 2:32:10integers. Right? So that has completely
  3253. 2:32:13worked fine. Now let's say if you want
  3254. 2:32:15to add in additional attribute over here
  3255. 2:32:19instead uh apart from the mall pin code
  3256. 2:32:21and the store code you also want to add
  3257. 2:32:24in the store location right so the way
  3258. 2:32:27to do that again is very simple we run
  3259. 2:32:29this statement
  3260. 2:32:33and instead of alter column we just say
  3261. 2:32:35add column and let's say this is the
  3262. 2:32:37store location
  3263. 2:32:40and this is going to be string Okay,
  3264. 2:32:44let's run this.
  3265. 2:32:46This works. And in order to test this
  3266. 2:32:49out,
  3267. 2:32:51let's run this statement again. And let
  3268. 2:32:55me just put a simple number over here
  3269. 2:32:56for now. And the store location is going
  3270. 2:32:59to be the ground floor for example. So
  3271. 2:33:03let's run this. Okay, I mistakenly
  3272. 2:33:06inserted
  3273. 2:33:08row number 17 again. because we don't
  3274. 2:33:10have any uh primary key constraint in
  3275. 2:33:12place this will work but ideally don't
  3276. 2:33:14do it. So let's now have a look at
  3277. 2:33:17select star from this table
  3278. 2:33:22and
  3279. 2:33:24we should see another row for 17 and
  3280. 2:33:27this basically has the store code
  3281. 2:33:29location equals null.
  3282. 2:33:32The previous one actually this is the
  3283. 2:33:33recent one. So the store location is the
  3284. 2:33:35ground floor and you see that all the
  3285. 2:33:37others have adjusted and the schema is
  3286. 2:33:40maintained. The store location is null
  3287. 2:33:42for all of the others. Right? So this
  3288. 2:33:45shows that schema evolution inside of
  3289. 2:33:49the strruct inside of nested structures
  3290. 2:33:53also work pretty well. Let's understand
  3291. 2:33:56automatic schema evolution inside of a
  3292. 2:33:58nested structure
  3293. 2:34:01with a few examples. And first let me
  3294. 2:34:03ensure that this that automatic schema
  3295. 2:34:07evolution property is set to true which
  3296. 2:34:08is this one. So let me run this again.
  3297. 2:34:13And
  3298. 2:34:14what we want out of this example right
  3299. 2:34:16what we want to understand out of this
  3300. 2:34:18example is that initially I if if I
  3301. 2:34:22wanted to add store location I had to
  3302. 2:34:25run an alter statement over here. So if
  3303. 2:34:28I want to add an attribute to a strct, I
  3304. 2:34:32need to run an alter statement. Now with
  3305. 2:34:35schema evolution, the expectation is
  3306. 2:34:37that if there is a new attribute coming
  3307. 2:34:39in, so let's say apart from store
  3308. 2:34:41location, there is a new attribute
  3309. 2:34:44called staff ID, the person who helped
  3310. 2:34:46make the purchase is coming in. I
  3311. 2:34:49shouldn't have to run an alter
  3312. 2:34:50statement. It should be automatically
  3313. 2:34:53accommodated inside of purchase detail.
  3314. 2:34:56So let's try that out.
  3315. 2:34:58Let me copy this insert statement and
  3316. 2:35:01let me paste it over here. This will be
  3317. 2:35:03for let's say customer number 21.
  3318. 2:35:07Customer ID 21 and I will be using a
  3319. 2:35:11named strruct. The reason why we are
  3320. 2:35:14using a named strct is because earlier
  3321. 2:35:16if you think of this we already know
  3322. 2:35:18that the first value over here is going
  3323. 2:35:21to be m pen code because it is there in
  3324. 2:35:23the schema right? 765 is going to be
  3325. 2:35:26store code and ground floor is going to
  3326. 2:35:28be the store location because all of
  3327. 2:35:30this information resides inside of the
  3328. 2:35:33schema. But now in this case, if we were
  3329. 2:35:36to randomly insert a value over here,
  3330. 2:35:38which is staff ID, I I put in ST some
  3331. 2:35:41value over here, it wouldn't know what
  3332. 2:35:44key does it belong to, right? So I need
  3333. 2:35:46to specify the key so that it is able to
  3334. 2:35:48add that key over here. Adding only a
  3335. 2:35:51value wouldn't work, right? So I hope
  3336. 2:35:53you understand why we need to use a name
  3337. 2:35:57struck. So I'll simply add
  3338. 2:36:00the keys over here which is going to be
  3339. 2:36:02all pin code.
  3340. 2:36:05This is going to be the store code.
  3341. 2:36:10This is going to be the store location.
  3342. 2:36:16And the last one is going to be
  3343. 2:36:21the staff ID.
  3344. 2:36:24Okay. So, let's go ahead and run this.
  3345. 2:36:31Perfect. So, this seemed to work fine.
  3346. 2:36:33Now, let's quickly validate the data.
  3347. 2:36:46And there you go. So, you see that staff
  3348. 2:36:48ID has been inserted correctly. And we
  3349. 2:36:52didn't have to run an alter statement to
  3350. 2:36:54accommodate this inside of purchase
  3351. 2:36:57detail. And as we we would expect,
  3352. 2:36:59right? All of the other purchase details
  3353. 2:37:02have staff ID equal to null. And this
  3354. 2:37:04makes the whole schema very consistent.
  3355. 2:37:07Let's finally understand the last
  3356. 2:37:09scenario which is column position
  3357. 2:37:12changes.
  3358. 2:37:14So I'll add in a header
  3359. 2:37:18which is scenario number four. And this
  3360. 2:37:21is to be column position changes. And
  3361. 2:37:23first of all, I want to
  3362. 2:37:26set this to false because I don't want
  3363. 2:37:28to use automatic schema evolution for
  3364. 2:37:31now.
  3365. 2:37:32And
  3366. 2:37:34let's go and do things manually first.
  3367. 2:37:36So in order to add a column in a
  3368. 2:37:39particular order, it is going to be add
  3369. 2:37:42columns.
  3370. 2:37:44And actually let's figure out which
  3371. 2:37:46column do we want to add first. Yeah. So
  3372. 2:37:50we had
  3373. 2:37:52we had these many columns and we don't
  3374. 2:37:54have the age column yet in our table. So
  3375. 2:37:59let me copy this this statement because
  3376. 2:38:02we are going to need this soon.
  3377. 2:38:07So let me paste this here for now. Let's
  3378. 2:38:09also see how our table currently looks
  3379. 2:38:12like. Select star from the table.
  3380. 2:38:16And we want to add the age column. Where
  3381. 2:38:19do we want to add it? So there are two
  3382. 2:38:21options.
  3383. 2:38:23Either you can add it as the first
  3384. 2:38:24column by specifying first. So this
  3385. 2:38:26becomes the column before customer ID.
  3386. 2:38:29The second option is you add it after
  3387. 2:38:34some column. So let's say you add it
  3388. 2:38:36after price.
  3389. 2:38:39So then it is going to look like this.
  3390. 2:38:42So let's say we want to run the second
  3391. 2:38:44statement. Let's go ahead and run this.
  3392. 2:38:47And now let's see how this is going to
  3393. 2:38:48look like. There you go. You see age has
  3394. 2:38:52found its position after price because
  3395. 2:38:54that is what we ran. Now let's run let's
  3396. 2:38:59run an insert statement to see if all of
  3397. 2:39:01this works or not. So again I'm going to
  3398. 2:39:03insert five records only to keep this
  3399. 2:39:05simple. And the column that we have is
  3400. 2:39:08customer ID, invoice number, price, and
  3401. 2:39:11then we have age. Then we have invoice
  3402. 2:39:14date. Then we have quantity.
  3403. 2:39:17Then we have payment method.
  3404. 2:39:20And purchase detail was something that
  3405. 2:39:22was added by us. So that doesn't reside
  3406. 2:39:25in this park file. So let me make it
  3407. 2:39:29null for simplicity. Null as purchase
  3408. 2:39:31details. Do we have any other column?
  3409. 2:39:33No. Let's run this.
  3410. 2:39:38Now let's do a
  3411. 2:39:42Okay, we'll just run this.
  3412. 2:39:46So here you see
  3413. 2:39:48that this has perfectly inserted the
  3414. 2:39:51records from 50 to 55. All the columns
  3415. 2:39:56have been populated correctly except the
  3416. 2:39:58purchase detail column which we
  3417. 2:40:00purposely set it as null. That means
  3418. 2:40:03column ordering had just worked fine.
  3419. 2:40:05Let's see how column position change is
  3420. 2:40:08going to work with automatic schema
  3421. 2:40:10evolution which simply means that if a
  3422. 2:40:13new column comes in at whatever position
  3423. 2:40:15right our table should be able to
  3424. 2:40:17accommodate it. So let's say there is a
  3425. 2:40:20new column called category which comes
  3426. 2:40:22in and I want it to be before payment
  3427. 2:40:25method it should be able to reflect that
  3428. 2:40:28without me having to run an alter
  3429. 2:40:30statement. Now again I want to highlight
  3430. 2:40:32that maybe this kind of a scenario we
  3431. 2:40:34would never want it to be in production
  3432. 2:40:37right because we don't want some random
  3433. 2:40:39data coming in and destroying our tables
  3434. 2:40:42right destroying the structure of our
  3435. 2:40:43table but for the sake of completeness I
  3436. 2:40:46want to I want you to know every
  3437. 2:40:48everything like all possible scenarios
  3438. 2:40:51right for now let's insert category into
  3439. 2:40:55our table
  3440. 2:40:57so let me copy this insert statement and
  3441. 2:41:02we're going to run this. And before
  3442. 2:41:05this, let's also
  3443. 2:41:09enable schema evolution.
  3444. 2:41:13The schema evolution is being set to
  3445. 2:41:16true.
  3446. 2:41:18And let's also
  3447. 2:41:21quickly print our
  3448. 2:41:24table,
  3449. 2:41:27right?
  3450. 2:41:29So we see that we already have numbers
  3451. 2:41:32from until 55.
  3452. 2:41:35So this is going to be from 56 to 60.
  3453. 2:41:39And the columns are customer ID, invoice
  3454. 2:41:42number, price, age, invoice date,
  3455. 2:41:45quantity, payment method, purchase
  3456. 2:41:48details and let me add category over
  3457. 2:41:51here. Right? So ideally category should
  3458. 2:41:55just come before purchase details if at
  3459. 2:41:58all this works right. So let's go ahead
  3460. 2:42:01and run this
  3461. 2:42:04and this didn't work. The error says
  3462. 2:42:06cannot resolve category due to data type
  3463. 2:42:08mismatch. Okay I don't want to correct
  3464. 2:42:11and run this again. Uh data type
  3465. 2:42:13mismatch cannot cast string to strct
  3466. 2:42:15data type. So an important thing to
  3467. 2:42:17understand here is that
  3468. 2:42:20ideally this is the set of all of my
  3469. 2:42:22columns right and this is an extra
  3470. 2:42:26column that I just added. So if you were
  3471. 2:42:29to just remove this for now and if you
  3472. 2:42:31were to think that this is your purchase
  3473. 2:42:32detail
  3474. 2:42:34this is your purchase detail then all of
  3475. 2:42:37the column should be completed with this
  3476. 2:42:40right this would be an extra column. So
  3477. 2:42:42what the insert statement is throwing as
  3478. 2:42:44an error is that this category over here
  3479. 2:42:46is supposed to be the last column and
  3480. 2:42:48the last column currently is a struck
  3481. 2:42:50data type. But what you're giving me is
  3482. 2:42:53a string data type. And for that reason
  3483. 2:42:56I'm throwing an error. So that's the
  3484. 2:42:59cause of this error. And now let's go
  3485. 2:43:01ahead and try this with a merge
  3486. 2:43:03statement. Let's see what happens with
  3487. 2:43:04the merge statement. So we say merge
  3488. 2:43:08into delta catalog. This is going to be
  3489. 2:43:11the target
  3490. 2:43:13using
  3491. 2:43:15this table right here
  3492. 2:43:20as the source
  3493. 2:43:22on target dot
  3494. 2:43:26customer ID equals source ID source dot
  3495. 2:43:29customer ID where
  3496. 2:43:32when not
  3497. 2:43:35matched
  3498. 2:43:37then insert
  3499. 2:43:40star right so here the code is exactly
  3500. 2:43:44the same category is not in the correct
  3501. 2:43:46position and category is added and then
  3502. 2:43:49extra column over here so let's see what
  3503. 2:43:51happened
  3504. 2:43:57okay that seemed to work and we are
  3505. 2:44:01going to
  3506. 2:44:03we are going to run this
  3507. 2:44:07and okay you don't see category column
  3508. 2:44:10before purchase details but you see the
  3509. 2:44:13category column after purchase details
  3510. 2:44:16and this is quite interesting right so
  3511. 2:44:19we read that merge does a matching by
  3512. 2:44:23name so it was able to match all of
  3513. 2:44:26these column all of this over here and
  3514. 2:44:28purchase details by name so it inserted
  3515. 2:44:32them at the right places the only thing
  3516. 2:44:34that that that it was not able to match
  3517. 2:44:36was category and it inserted it as the
  3518. 2:44:40last column. So the positioning didn't
  3519. 2:44:42work but you still ended up putting the
  3520. 2:44:45data inside of the table. Right now let
  3521. 2:44:48me take you through how schema evolution
  3522. 2:44:50is going to look like with spark data
  3523. 2:44:53frame. Yeah. And the way we going to do
  3524. 2:44:56that is by quickly taking some of the
  3525. 2:45:01parket data that we read above. Right?
  3526. 2:45:04So I'm going to take this one this park
  3527. 2:45:07a files
  3528. 2:45:09and
  3529. 2:45:11we going to do a from pispark.sql.f
  3530. 2:45:15function import star.
  3531. 2:45:17Let's filter
  3532. 2:45:20the customer ids.
  3533. 2:45:24Customer ID dot between 1, 10. And let's
  3534. 2:45:28also select
  3535. 2:45:31a few columns, right? Customer ID,
  3536. 2:45:35price,
  3537. 2:45:37invoice, date.
  3538. 2:45:41And let's go ahead and write this
  3539. 2:45:43df.right write dot save as table. This
  3540. 2:45:48is going to be in delta catalog dot
  3541. 2:45:50deltadb dot
  3542. 2:45:54schema invoices schema evolution sparkd
  3543. 2:45:58right so let's go ahead and write this
  3544. 2:46:00okay this is go this should be post
  3545. 2:46:03python
  3546. 2:46:07and if everything works we should see a
  3547. 2:46:11table here inside of our catalog and
  3548. 2:46:15there you go so you We have the table.
  3549. 2:46:18We have the table that we created. Now
  3550. 2:46:20let me add
  3551. 2:46:22let me add a few more column. Right. Let
  3552. 2:46:26me add quantity payment method.
  3553. 2:46:30Quantity and payment
  3554. 2:46:34method.
  3555. 2:46:36We added quantity and payment method
  3556. 2:46:38which changes the schema
  3557. 2:46:42of the table. Right? So I mean we have
  3558. 2:46:45not written it but the schema is
  3559. 2:46:47different from what we had over here. So
  3560. 2:46:49let me choose customer ids from 11 to 25
  3561. 2:46:58and let's
  3562. 2:47:01write dot mode append
  3563. 2:47:06dot option
  3564. 2:47:12merge schema
  3565. 2:47:16to true.
  3566. 2:47:19Let's go ahead and run this now.
  3567. 2:47:25And now let's see how our table looks
  3568. 2:47:27like.
  3569. 2:47:32Okay. So you see that from 11 to 25 we
  3570. 2:47:38have the data for quantity and payment
  3571. 2:47:41method. But from 1 to 10 we don't have
  3572. 2:47:43data for quantity and payment method.
  3573. 2:47:46Right? And this is because
  3574. 2:47:49the schema has evolved due to writing
  3575. 2:47:51merge schema equals true. So now we are
  3576. 2:47:54going to understand how to convert park
  3577. 2:47:57files into the delta format. So most of
  3578. 2:48:01the times your input files or your
  3579. 2:48:03existing files are not going to be in
  3580. 2:48:06the delta format. But in order to be
  3581. 2:48:08able to use the awesome features that
  3582. 2:48:11we've been talking about, we need to
  3583. 2:48:13convert it to delta. Yeah. So that is
  3584. 2:48:15what we're going to understand through a
  3585. 2:48:17lot of examples. So I'm going to use the
  3586. 2:48:19same paret file that I've used earlier.
  3587. 2:48:21So I'll just copy the path of this
  3588. 2:48:24parket file. And first of all, let's
  3589. 2:48:26quickly do a percent fsls and see what's
  3590. 2:48:29there inside of the invoices folder.
  3591. 2:48:32And we see that there are three paret
  3592. 2:48:35files. And let me go ahead and use this
  3593. 2:48:38one. The customer ID is from 1 to 100.
  3594. 2:48:41So
  3595. 2:48:43let me first create copies two copies of
  3596. 2:48:45this.
  3597. 2:48:46I'm creating two copies because I want
  3598. 2:48:48to show you two ways of converting from
  3599. 2:48:50park a to delta. So spark read.park
  3600. 2:48:54and let me
  3601. 2:48:57do a mode overrite
  3602. 2:49:00dot
  3603. 2:49:02park. And let's paste the path. This is
  3604. 2:49:06going to be let's name it as v_sub_1 and
  3605. 2:49:08let's not keep it in the invoices folder
  3606. 2:49:14and this is going to be v2. So basically
  3607. 2:49:16where it is going to reside is inside of
  3608. 2:49:19the lab data container at the root.
  3609. 2:49:22Yeah. So here if I go back. So here is
  3610. 2:49:25my lab data container and it is going to
  3611. 2:49:27reside over here. Yeah. So let's go
  3612. 2:49:29ahead and run this now.
  3613. 2:49:33So now I should see two folders. So
  3614. 2:49:35these are my two folders, right? So
  3615. 2:49:37these are paret files and if you see
  3616. 2:49:40here it has a snappy.park and this is
  3617. 2:49:44where the data is stored and it doesn't
  3618. 2:49:46have a delta log folder. That means this
  3619. 2:49:48is not a delta table yet. Same applies
  3620. 2:49:51for this one as well. Yeah. So now let's
  3621. 2:49:54go ahead and
  3622. 2:49:56try to convert this to delta. to convert
  3623. 2:49:59to
  3624. 2:50:01delta and then I write park dot let me
  3625. 2:50:05paste the path.
  3626. 2:50:10So now let's go ahead and run this.
  3627. 2:50:14Okay, this is complete.
  3628. 2:50:17So there you go. You see a delta log
  3629. 2:50:20folder which is going to record the
  3630. 2:50:22transactions. And let's go ahead and see
  3631. 2:50:25what's inside of JSON. And there is an
  3632. 2:50:29add operation which added this park
  3633. 2:50:32file. So it registered this transaction.
  3634. 2:50:34Right? Now let's try another method.
  3635. 2:50:37Let's say you want to use the
  3636. 2:50:40delta table API.
  3637. 2:50:43Import
  3638. 2:50:44delta table. And the method that we
  3639. 2:50:47going to use is convert to delta.
  3640. 2:50:50We pass over here. And this is going to
  3641. 2:50:53be the format is park dot.
  3642. 2:50:56We specify the tildas over here. And
  3643. 2:51:00we copy this path v2 because we want to
  3644. 2:51:03convert v2. And let's go ahead and run
  3645. 2:51:07this. Okay, this is complete.
  3646. 2:51:10And let's have a look at this now. So
  3647. 2:51:14you see the delta log over here as well.
  3648. 2:51:17And let's check this. And this should
  3649. 2:51:20also have an add operation which
  3650. 2:51:23registered the data. Right? So this is
  3651. 2:51:26how you're able to convert a parquet
  3652. 2:51:29file to a delta table right and
  3653. 2:51:32similarly you can do it for a CSV file
  3654. 2:51:35or a JSON file or anything else right
  3655. 2:51:37maybe the approach would be a little
  3656. 2:51:39different you would have to read the CSV
  3657. 2:51:41file into a data frame something like df
  3658. 2:51:45equals spark read dot csv and then maybe
  3659. 2:51:49the path over here and you have to write
  3660. 2:51:51it
  3661. 2:51:53let's say the mode is
  3662. 2:51:56overrite
  3663. 2:51:58and then the format is going to be delta
  3664. 2:52:02and then you specify the path over here
  3665. 2:52:05whatever path you want it to be right so
  3666. 2:52:08this is another approach that you can
  3667. 2:52:10follow in order to write your CSV or any
  3668. 2:52:13other format of files now let's
  3669. 2:52:15understand what are manage and external
  3670. 2:52:18table actually the table that we've
  3671. 2:52:21created in most of our examples they are
  3672. 2:52:24manage tables but let's understand them
  3673. 2:52:26in a lot more details now. So managed
  3674. 2:52:29tables are those tables where your both
  3675. 2:52:32your data and the meta data
  3676. 2:52:36they are managed by delta lake
  3677. 2:52:41right they are managed by delta lake and
  3678. 2:52:44we create delta lake or delta tables
  3679. 2:52:48right we create delta tables using the
  3680. 2:52:52unity catalog. Now where does unity
  3681. 2:52:54catalog store data? Unity catalog stores
  3682. 2:52:56data inside of a storage account inside
  3683. 2:53:00of this container called metas store.
  3684. 2:53:04Right? So that is where it creates your
  3685. 2:53:06tables. So the tables that we create
  3686. 2:53:09they are stored and managed inside of
  3687. 2:53:13this container meta store right is
  3688. 2:53:15stored in a manage location. And in case
  3689. 2:53:19of an external table both the data
  3690. 2:53:22actually the data is stored in a user
  3691. 2:53:27specified location. It can be an S3
  3692. 2:53:30bucket or an ADLS gen 2 container or
  3693. 2:53:34Google cloud storage. Right? But the
  3694. 2:53:36meta data
  3695. 2:53:38resides in the meta store of the Unity
  3696. 2:53:43catalog. And the meta store again could
  3697. 2:53:45be something like this. Right? So the
  3698. 2:53:47metadata is going to reside in the meta
  3699. 2:53:49store of the Unity catalog and the data
  3700. 2:53:53is going to reside in a location of
  3701. 2:53:56users choice. So now let's see this in
  3702. 2:53:58action. Let's see an example. If I were
  3703. 2:54:01to pull up
  3704. 2:54:03any one of the table that I created
  3705. 2:54:05earlier,
  3706. 2:54:08we see that let's pick up invoices SC.
  3707. 2:54:12Let's go to details. And what you see
  3708. 2:54:15over here is the storage location,
  3709. 2:54:17right? So the storage location is this
  3710. 2:54:19storage account inside of this
  3711. 2:54:21container. The container is called meta
  3712. 2:54:23store which is the one over here. And
  3713. 2:54:28there we're going to have tables and
  3714. 2:54:32this is the unique identifier of my
  3715. 2:54:35table. So my table is going to be this
  3716. 2:54:37one. And all of the data and metadata is
  3717. 2:54:41stored over here. is managed by the
  3718. 2:54:44Unity catalog. That's one. Now, let's
  3719. 2:54:48try to create an external table. So,
  3720. 2:54:52let's write some SQL. Create or replace
  3721. 2:54:57replace table. And this is going to be
  3722. 2:55:00delta catalog dot delta DB dot invoices
  3723. 2:55:05external.
  3724. 2:55:06So I'm going to
  3725. 2:55:09be using delta because the output the
  3726. 2:55:12underlying storage that I want is
  3727. 2:55:14supposed to be in a delta format. Right?
  3728. 2:55:16Now this can be any format but I
  3729. 2:55:18specifically want it to be in delta. So
  3730. 2:55:21that is why I say using delta
  3731. 2:55:24and then I specify what is the location
  3732. 2:55:26where I want to store my data.
  3733. 2:55:29Let's go ahead and use this. Let's put
  3734. 2:55:32it in some other location. This time not
  3735. 2:55:34the meta store. Let's put it inside lab
  3736. 2:55:37data.
  3737. 2:55:40And so I have lab data over here and let
  3738. 2:55:44me name this as invoices external
  3739. 2:55:47and this is going to be as select star
  3740. 2:55:50from
  3741. 2:55:52park
  3742. 2:55:56dot
  3743. 2:55:58let's say this file over here
  3744. 2:56:03right
  3745. 2:56:05so let me paste this over here and let's
  3746. 2:56:09run this
  3747. 2:56:16So let's see if this table exists in the
  3748. 2:56:18catalog. So there is a table invoices
  3749. 2:56:21external and let's see the details. Now
  3750. 2:56:23the storage location is not in the meta
  3751. 2:56:25store. It is the location that we
  3752. 2:56:28specified. Right? So let's refresh this
  3753. 2:56:32and this is invoice ext. And you have
  3754. 2:56:35this as a delta table. Right? Now the
  3755. 2:56:38second difference is if I do a drop on a
  3756. 2:56:42manage table the data the underlying
  3757. 2:56:44data is going to vanish. It is going to
  3758. 2:56:46go away. But in case of an external
  3759. 2:56:48table this data that you see over here
  3760. 2:56:51even if I do a drop on this table right.
  3761. 2:56:54So if I do something like a
  3762. 2:56:58drop table
  3763. 2:57:01the underlying data is still going to
  3764. 2:57:03stay in the user specified location.
  3765. 2:57:07Right? So let me go ahead and
  3766. 2:57:10run this
  3767. 2:57:12now. This should disappear from the
  3768. 2:57:13catalog. It has disappeared indeed. And
  3769. 2:57:15let me go ahead and refresh this.
  3770. 2:57:21You see that the data is still there in
  3771. 2:57:24the user specified location. And if this
  3772. 2:57:26were a manage table, the data would have
  3773. 2:57:29gone. So actually sometimes in cases of
  3774. 2:57:32manage table you would still see the
  3775. 2:57:34data because of it default retention
  3776. 2:57:36period. Sometime the retention period is
  3777. 2:57:38set to 30 days. So that is the reason
  3778. 2:57:41why the data doesn't disappear
  3779. 2:57:43immediately. Now the third point of
  3780. 2:57:46difference is that manage tables use the
  3781. 2:57:49delta format only. So all of the table
  3782. 2:57:51that are created in the unity catalog
  3783. 2:57:54inside residing inside of this meta
  3784. 2:57:56store follow the delta format only. The
  3785. 2:58:00manage tables follow the delta format
  3786. 2:58:03only. But in case of external table they
  3787. 2:58:07they could follow several formats. It
  3788. 2:58:09can be CSV, it can be JSON, it can be a
  3789. 2:58:12park,
  3790. 2:58:15right? It can even be delta.
  3791. 2:58:18So in the last example we saw that we've
  3792. 2:58:23written using delta
  3793. 2:58:26and because we wrote using delta it
  3794. 2:58:29created the output in a delta format
  3795. 2:58:31right if we wrote using par
  3796. 2:58:36it would create the output in a park
  3797. 2:58:39file format right so that's the reason
  3798. 2:58:41why several formats are supported for
  3799. 2:58:44external table so now let's talk about
  3800. 2:58:46something really interesting thing in
  3801. 2:58:48Delta Lake called deletion vectors and I
  3802. 2:58:52believe we've already seen it in action
  3803. 2:58:54but now let's talk about it in a lot
  3804. 2:58:57more detail. So first of all let's
  3805. 2:59:00understand what is the problem what is
  3806. 2:59:03the issue at hand that deletion vectors
  3807. 2:59:06is trying to solve right so we already
  3808. 2:59:10know that delta uses paret files under
  3809. 2:59:13the hood for storing its data right and
  3810. 2:59:18park files are immutable. So what I mean
  3811. 2:59:20by immutable is that they cannot be
  3812. 2:59:23modified directly. Right? So if you want
  3813. 2:59:26to update or delete a record in a park
  3814. 2:59:29file, what you essentially need to do is
  3815. 2:59:32you read the park file and then you
  3816. 2:59:35apply those operation delete or an
  3817. 2:59:37update and then you write the park file
  3818. 2:59:40back. Right? So that is the only way you
  3819. 2:59:42can update or make changes to that file.
  3820. 2:59:46Right? You read it, apply the operations
  3821. 2:59:49and then write it back. Now this rewrite
  3822. 2:59:53process, this whole read and a rewrite
  3823. 2:59:55process is very costly. Imagine if you
  3824. 2:59:59had a parquet file with 10 million rows
  3825. 3:00:02and you only needed to delete a few
  3826. 3:00:04couple of rows, right? So what you would
  3827. 3:00:07end up doing is you would read the whole
  3828. 3:00:09park files and you would rewrite the
  3829. 3:00:12whole park file again except for those
  3830. 3:00:15couple of records which you want to
  3831. 3:00:17delete. Yeah. So this is a tremendously
  3832. 3:00:23computationally expensive operation and
  3833. 3:00:25this is where deletion vectors come into
  3834. 3:00:28play. Now before we talk about deletion
  3835. 3:00:31vectors, let's talk about two important
  3836. 3:00:33concepts. The first one is copy on write
  3837. 3:00:37and the second one is merge on read.
  3838. 3:00:41Yeah. So let's understand copy on write
  3839. 3:00:43first. Copy on write is exactly the same
  3840. 3:00:47as what we discussed above. Right. Every
  3841. 3:00:50change, every update, every delete or a
  3842. 3:00:54merge is going to create a new file.
  3843. 3:00:57Right? And Delta is going to follow this
  3844. 3:00:59approach when deletion vectors are
  3845. 3:01:03disabled. So it's very important to keep
  3846. 3:01:05in mind that copy on write is going to
  3847. 3:01:08be followed when deletion vectors are
  3848. 3:01:11disabled. Yeah. So let's take this
  3849. 3:01:14example. Let's say we have a paret file
  3850. 3:01:171.park and it contains,000 records,
  3851. 3:01:21right? It contains 1,000 records and
  3852. 3:01:25that is what we see over here from one
  3853. 3:01:27until,000. And now what we want to do is
  3854. 3:01:30we just want to delete rows number one
  3855. 3:01:33and six. So the most simple way that
  3856. 3:01:37copy on right is going to do is that
  3857. 3:01:39take up this whole parquet file. Right?
  3858. 3:01:42You take up this whole park file and
  3859. 3:01:45then apply the delete operation.
  3860. 3:01:49You apply the delete operation and then
  3861. 3:01:51you rewrite the whole file back except
  3862. 3:01:55rows number one and rows number six. So
  3863. 3:01:58these two rows are going to be omitted
  3864. 3:02:01and then the whole file is going to be
  3865. 3:02:04written back as you see over here. So it
  3866. 3:02:06doesn't have row number one and row
  3867. 3:02:08number six over here. Right? So the new
  3868. 3:02:12state that you see is that a new file
  3869. 3:02:14which is 2.par is written to storage.
  3870. 3:02:18Right? And this file is going to be the
  3871. 3:02:21one to be referenced for the latest
  3872. 3:02:24version. So that is how copy on write
  3873. 3:02:27works. Now in contrast what merge on
  3874. 3:02:31read does is that it allows your
  3875. 3:02:34original park files to remain untouched
  3876. 3:02:38and instead the changes the changes that
  3877. 3:02:41we apply right for example deletions
  3878. 3:02:44they are recorded in a separate file
  3879. 3:02:47known as a deletion vector. Yeah. So
  3880. 3:02:50let's say if I want to delete a record
  3881. 3:02:52that record that row number is recorded
  3882. 3:02:56in a deletion vector and when I read the
  3883. 3:02:58file the deletion vector is simply going
  3884. 3:03:00to be checked if this row is present in
  3885. 3:03:03the deletion vector or not. If it is
  3886. 3:03:06then that row is going to be skipped
  3887. 3:03:08from the output. Yeah. So delta is going
  3888. 3:03:12to follow the merge on read approach if
  3889. 3:03:15deletion vectors are enabled. Yeah. So
  3890. 3:03:18it's really important to keep this mind.
  3891. 3:03:19Keep this in mind again. If deletion
  3892. 3:03:21vectors are enabled, merge on read is
  3893. 3:03:24going to be followed. If not, copy on
  3894. 3:03:26write is going to be followed. So now
  3895. 3:03:28let's understand this with an example.
  3896. 3:03:31Yeah. So let's say we have the same park
  3897. 3:03:35file with th00and records. Yeah. So 1
  3898. 3:03:38th00and records as you see over here
  3899. 3:03:40from one until,000 over here. And we
  3900. 3:03:43want to perform the same operation. We
  3901. 3:03:47want to delete rows number one and six.
  3902. 3:03:50Yeah. So this time what happens is when
  3903. 3:03:53we say we want to perform a delete of
  3904. 3:03:56row number one which is right over here,
  3905. 3:03:59we simply record it in a deletion vector
  3906. 3:04:03which is over here. Right? The next time
  3907. 3:04:06we say that okay we want to delete row
  3908. 3:04:08number six. What happens is we record
  3909. 3:04:12row number six in the deletion vector
  3910. 3:04:14again. Right? So all of these changes
  3911. 3:04:17are getting recorded in the deletion
  3912. 3:04:19vector and all of this data that you
  3913. 3:04:22have over here they are stored in
  3914. 3:04:241.park. So now let's say I want to read
  3915. 3:04:28the file right I want to read the latest
  3916. 3:04:30state. So when I read the latest state I
  3917. 3:04:34would expect that row number one and row
  3918. 3:04:37number six shouldn't be there in the
  3919. 3:04:39file. Right? So the way I'm going to get
  3920. 3:04:42the output is that this whole file is
  3921. 3:04:46going to be produced as output and this
  3922. 3:04:49row is going to be checked against the
  3923. 3:04:51deletion vector that is it present in
  3924. 3:04:53the deletion vector or not. If it is
  3925. 3:04:56then this is marked as a soft delete and
  3926. 3:05:00you won't see that row. So this won't be
  3927. 3:05:03produced in the output. Similarly this
  3928. 3:05:05row is going to be checked is it present
  3929. 3:05:08in the deletion vector or not. It's not
  3930. 3:05:09present. So that means this row is going
  3931. 3:05:11to be displayed. Similarly for all the
  3932. 3:05:13rows and row six would also be checked
  3933. 3:05:16and it would not be displayed. Right? So
  3934. 3:05:19that is how it helps you in not
  3935. 3:05:24rewriting the entire file back. The
  3936. 3:05:26final state that you would have is
  3937. 3:05:281.park park the whole file is untouched
  3938. 3:05:31and a small deletion vector bit mapap
  3939. 3:05:34file right where all of the deletes all
  3940. 3:05:38of the changes are getting recorded and
  3941. 3:05:40it is going to be checked against and
  3942. 3:05:43this is going to add a lot of speed to
  3943. 3:05:47the entire process and it is going to
  3944. 3:05:48make a lot of things more performant. So
  3945. 3:05:51we see that deletion vectors bring the
  3946. 3:05:54merge on read capability to delta lake
  3947. 3:05:57thereby increasing the performance of
  3948. 3:05:59deletes updates. So update is considered
  3949. 3:06:02as an insert plus delete and merge
  3950. 3:06:06operations. Right? So instead of
  3951. 3:06:08rewriting the whole file back for small
  3952. 3:06:11small changes what it does is that it
  3953. 3:06:13records those small changes in a
  3954. 3:06:15separate bit map file known as deletion
  3955. 3:06:18vectors. Let's see all of this in action
  3956. 3:06:20now. Let's get started with our labs.
  3957. 3:06:23So, I'm going to quickly pick up one of
  3958. 3:06:26the old parket file that we've been
  3959. 3:06:28using. And let me quickly do a select
  3960. 3:06:32star from parket.
  3961. 3:06:34And then this one, let me use the 2010
  3962. 3:06:39to 2011. Right? And let's keep the
  3963. 3:06:43default as SQL because we're going to
  3964. 3:06:45write lots of SQL. Let's connect this
  3965. 3:06:49cluster and let me quickly do a limit
  3966. 3:06:52five and let's go ahead and run this.
  3967. 3:06:55Meanwhile,
  3968. 3:06:57first we are going to explode copy on
  3969. 3:07:00right. So let me quickly write how to
  3970. 3:07:04create a table with deletion vector
  3971. 3:07:09disabled.
  3972. 3:07:11Right. So let's see how this works.
  3973. 3:07:16Great. So let me do create
  3974. 3:07:20or replace table
  3975. 3:07:24and
  3976. 3:07:26the catalog that we're using is this one
  3977. 3:07:30delta catalog dot delta db right so this
  3978. 3:07:34is going to be delta catalog do delta db
  3979. 3:07:40dot invoices
  3980. 3:07:42and this is going to be
  3981. 3:07:46D copy on right right yeah and let me
  3982. 3:07:50quickly copy this over here in the table
  3983. 3:07:52properties
  3984. 3:07:54delta dot enable deletion vector to be
  3985. 3:07:58false and this is going to be created as
  3986. 3:08:02a c dash creatable at select statement.
  3987. 3:08:06So let's go ahead and run this now.
  3988. 3:08:10Yeah, let me also write a describe
  3989. 3:08:14extended
  3990. 3:08:16and then
  3991. 3:08:18delta
  3992. 3:08:20catalog dot
  3993. 3:08:23delta db dot invoices.
  3994. 3:08:27Yeah. So let's run this and have a look
  3995. 3:08:31at some of the properties quickly.
  3996. 3:08:38So here we see that this is a manage
  3997. 3:08:42table and the deletion vector is not
  3998. 3:08:46enabled. Yeah. Let's also see where
  3999. 3:08:50where does this table reside inside of
  4000. 3:08:53ADLS.
  4001. 3:08:57So let me quickly open up ADLS.
  4002. 3:09:12And this is the table right here.
  4003. 3:09:16So it basically contains one paret file
  4004. 3:09:19at this moment. Right? Now let's go
  4005. 3:09:22ahead and perform some operation. Right?
  4006. 3:09:25So let's say this user 105
  4007. 3:09:28uh from age 57 let's say it was
  4008. 3:09:31incorrectly recorded and we want to
  4009. 3:09:34correct the age to 55 for user 105. So
  4010. 3:09:38we are simply going to say update
  4011. 3:09:43this table
  4012. 3:09:45set
  4013. 3:09:47age equals 55 where customer ID equals
  4014. 3:09:52105. Right? So this customer is 105
  4015. 3:09:57right here. So let's go ahead and run
  4016. 3:10:00this.
  4017. 3:10:05And let's also see the history.
  4018. 3:10:08Describe history
  4019. 3:10:11and this table.
  4020. 3:10:13So it should contain two rows. The first
  4021. 3:10:15one for the create. Yeah. The first one
  4022. 3:10:18for the create and the second one for
  4023. 3:10:20the update. Right. So based on whatever
  4024. 3:10:23we've studied and understood right now,
  4025. 3:10:26this update should create a fresh par
  4026. 3:10:29file right with the changes. So let's go
  4027. 3:10:33ahead and understand a few metric the
  4028. 3:10:36operation metrics over here right. So
  4029. 3:10:38what it says is that number of removed
  4030. 3:10:40files equals 1. Number of copied rows
  4031. 3:10:44equals 99. So we see that we had a total
  4032. 3:10:46of 100 rows over here right from here.
  4033. 3:10:49And the row that got updated is just one
  4034. 3:10:53row which is row number 105. So other
  4035. 3:10:56than that it copied all of the 99 rows.
  4036. 3:11:00No deletion vectors added or removed
  4037. 3:11:02because deletion vectors are disabled
  4038. 3:11:06number of added files. So it added a new
  4039. 3:11:08file right and then it updated one row.
  4040. 3:11:12Quite similar to what we would expect,
  4041. 3:11:14right? Because we ran an update
  4042. 3:11:15statement over here.
  4043. 3:11:18Now let's go ahead and see how this
  4044. 3:11:20would look in the storage.
  4045. 3:11:24So we see that this has added another
  4046. 3:11:26file. It has rewritten another file.
  4047. 3:11:28Yeah. And if you were to quickly look at
  4048. 3:11:31the sizes,
  4049. 3:11:33the sizes is more or less the same.
  4050. 3:11:36Right? This is 59 to1 and this is also
  4051. 3:11:3959 to1. Yeah. Now let's go ahead and run
  4052. 3:11:42another another statement. This time
  4053. 3:11:45let's go ahead and run a delete
  4054. 3:11:47statement. So let's go ahead and delete
  4055. 3:11:50the row for customer ID equals 102.
  4056. 3:11:54So I'm going to write delete
  4057. 3:11:57from this table where customer ID equals
  4058. 3:12:00102. Right? So let's go ahead and run
  4059. 3:12:02this and let me also see
  4060. 3:12:05what the history is going to show me.
  4061. 3:12:11Now we have three rows. the third one
  4062. 3:12:13for the delete operation and let's
  4063. 3:12:15quickly see the operation matrix. So now
  4064. 3:12:19the operation matrix what it says is
  4065. 3:12:22number of copied rows is 99.
  4066. 3:12:26Number of deleted rows is one. Yeah. And
  4067. 3:12:30then it added all of this in a new file.
  4068. 3:12:33Right? So it basically removed the file
  4069. 3:12:36that it read in the previous version.
  4070. 3:12:38That is why you see number of removed
  4071. 3:12:39files one. And then it added this new
  4072. 3:12:42file. It deleted one row over here
  4073. 3:12:45because number of deleted rows is equal
  4074. 3:12:46to one. And let's have a look at the
  4075. 3:12:48size. So this is 5897
  4076. 3:12:52and this was 5921. A little smaller
  4077. 3:12:56because we deleted one record. Now if
  4078. 3:12:58you have a look over here, we should see
  4079. 3:13:00one file over here. Yeah,
  4080. 3:13:04there you go. So we see this file which
  4081. 3:13:07is which was added at 1639.
  4082. 3:13:10This file was added with 99 records.
  4083. 3:13:14Yeah. So I believe this confirms how
  4084. 3:13:18copy on write works. It is reading the
  4085. 3:13:21file applying the changes and then
  4086. 3:13:23writing it as a new file. So let's now
  4087. 3:13:26see how is a delta table going to behave
  4088. 3:13:29if merge on read was enabled. Yeah. So
  4089. 3:13:34we've already discussed that if deletion
  4090. 3:13:37vectors are enabled then the paradigm
  4091. 3:13:40that is going to be followed is merge on
  4092. 3:13:43read. So let me create this table and I
  4093. 3:13:47will rename this to m and let's set this
  4094. 3:13:51to true which is going to enable
  4095. 3:13:54deletion vectors and let's go ahead and
  4096. 3:13:56run this. I will do a describe extended
  4097. 3:14:04this table just to verify a few of the
  4098. 3:14:06properties. So this is a manage table
  4099. 3:14:10and delta.enable deletion vectors is set
  4100. 3:14:14to true. Now let's go ahead and perform
  4101. 3:14:18similar operations that we performed on
  4102. 3:14:21the earlier table, right? The table
  4103. 3:14:23where deletion vectors was disabled.
  4104. 3:14:27So let's go ahead and run this. And this
  4105. 3:14:30is going to be M.
  4106. 3:14:33And I'm going to do a describe history
  4107. 3:14:37of
  4108. 3:14:39this table.
  4109. 3:14:41And I should see two rows. Yeah. The
  4110. 3:14:43first one for the CAS and the second one
  4111. 3:14:46for delete. So let's understand
  4112. 3:14:50a few operation metrics. Right? So what
  4113. 3:14:54it says is that number of removed files
  4114. 3:14:57and the number of removed bytes is zero.
  4115. 3:15:00And this is a little different from what
  4116. 3:15:03we saw earlier.
  4117. 3:15:06What we saw earlier was
  4118. 3:15:09it removed the file and then it
  4119. 3:15:11performed the delete operation and then
  4120. 3:15:13it wrote the file back. Yeah. But here
  4121. 3:15:16what we see is a little different.
  4122. 3:15:19There were no files that were removed
  4123. 3:15:21and there were no rows that were copied.
  4124. 3:15:23Instead, there is a deletion vector that
  4125. 3:15:26has been added. Yeah. And the number of
  4126. 3:15:30deleted rows is one, which is exactly
  4127. 3:15:33what we did over here. Yeah.
  4128. 3:15:36So that's all of the operation metrics,
  4129. 3:15:39right? Now let's quickly see how does it
  4130. 3:15:42look like
  4131. 3:15:44in the storage in our storage layer.
  4132. 3:15:47Yeah. So, let me copy the file
  4133. 3:15:51identifier
  4134. 3:15:53and this file is right here. And there
  4135. 3:15:55you see that this is the original file.
  4136. 3:15:59We don't have another new file as you
  4137. 3:16:02would see in a copy on write paradigm.
  4138. 3:16:06We only have the old file and then we
  4139. 3:16:08have a deletion vector which records the
  4140. 3:16:11rows which we need to delete. Yeah, very
  4141. 3:16:14simple and very performant because it
  4142. 3:16:16didn't have to rewrite the whole file.
  4143. 3:16:18Now let's go ahead and perform another
  4144. 3:16:22operation which is the update.
  4145. 3:16:26So I'll just change the name of the
  4146. 3:16:28table and the operation will be just the
  4147. 3:16:32same and let me do a describe history
  4148. 3:16:38table and I should see three rows. Yeah,
  4149. 3:16:42the third one is an update. And let's
  4150. 3:16:44quickly have a look over here. So the
  4151. 3:16:47number of removed files is zero. Number
  4152. 3:16:48of removed bytes, number of copied rows
  4153. 3:16:50is all zero. What it did is it added a
  4154. 3:16:54deletion vector. Yeah, it removed the
  4155. 3:16:57old deletion vector and then it added a
  4156. 3:17:00new deletion vector. Right? So the new
  4157. 3:17:02deletion vector has been updated with
  4158. 3:17:05more details and that is what we would
  4159. 3:17:08see over here. So if I were to look at
  4160. 3:17:12quickly refresh this what what should we
  4161. 3:17:14expect? So we should expect two things.
  4162. 3:17:16The first one is that I should expect a
  4163. 3:17:18park file because as we've read in the
  4164. 3:17:21earlier section an update is considered
  4165. 3:17:24as a delete plus insert. Where is the
  4166. 3:17:28delete going to be recorded? The delete
  4167. 3:17:30is going to be recorded in the updated
  4168. 3:17:32deletion vector. So one more deletion
  4169. 3:17:35vector should be created and the row
  4170. 3:17:37that is updated that should be recorded
  4171. 3:17:40in another paret file. Yeah. So I should
  4172. 3:17:43see one more paret file and I should see
  4173. 3:17:45an updated deletion vector that is one
  4174. 3:17:48more deletion vector. So let's refresh
  4175. 3:17:50this and there you go. So you see one
  4176. 3:17:53more parket file and you see one more
  4177. 3:17:57deletion vector and this is the updated
  4178. 3:18:00deletion vector. So now you may have
  4179. 3:18:02this question that on a table there
  4180. 3:18:06could be several deletes, updates and
  4181. 3:18:08merge operation that could be going on
  4182. 3:18:10right and as a result of this as you see
  4183. 3:18:12over here it could generate a lot of
  4184. 3:18:16small files lot of deletion vectors
  4185. 3:18:18right so what is the way to clean this
  4186. 3:18:21up right to tidy this up because there
  4187. 3:18:24are two deletion vectors over here one
  4188. 3:18:26is obsolete the other one is updated and
  4189. 3:18:29then there is this one file over here
  4190. 3:18:31which just contains one row. Right? So
  4191. 3:18:34for all of this, datab bricks provides
  4192. 3:18:37us the optimize command. So you have the
  4193. 3:18:41optimize command that you can apply on a
  4194. 3:18:44delta table. So let me go ahead and
  4195. 3:18:48quickly show you. So let's say we run
  4196. 3:18:51the optimize on this table.
  4197. 3:18:56And what it does is that it is going to
  4198. 3:18:58create a new version and that new
  4199. 3:19:02version is going to be the latest state
  4200. 3:19:04with all of the deletion vector and all
  4201. 3:19:07of those computations applied. Right? So
  4202. 3:19:09it is going to check which row it's
  4203. 3:19:11supposed to be there in the latest
  4204. 3:19:12version. Do all of those computations
  4205. 3:19:14and then create one fresh paret file.
  4206. 3:19:17Right? So over here I should just see
  4207. 3:19:20one fresh par file created using all of
  4208. 3:19:23this. Right? it is going to apply all of
  4209. 3:19:25the computation deductions and then
  4210. 3:19:27create one new park file. So now what we
  4211. 3:19:30see over here that this new park file
  4212. 3:19:33has been added and this has been created
  4213. 3:19:36by computing the latest state. Now if I
  4214. 3:19:38were to do a vacuum yeah so vacuum
  4215. 3:19:41basically keep the latest state. It is
  4216. 3:19:43going to remove all of the history. It
  4217. 3:19:46is only going to keep this final file,
  4218. 3:19:49right? Because this final file contains
  4219. 3:19:51the latest state. So let's go ahead and
  4220. 3:19:55run
  4221. 3:19:57vacuum this table and then retain zero
  4222. 3:20:01hours. And then we also need to set a
  4223. 3:20:04configuration which is called
  4224. 3:20:07set spark.ta
  4225. 3:20:10uh spark.databicks.delta
  4226. 3:20:13dot retention
  4227. 3:20:15duration check.enable is false. And let
  4228. 3:20:17me go ahead and run this. So they're
  4229. 3:20:19going to clean up all of the older
  4230. 3:20:21versions and what we should end up
  4231. 3:20:24seeing it just one park file which is
  4232. 3:20:27the one created at 1814. Yeah.
  4233. 3:20:34So let me refresh this and there you go.
  4234. 3:20:36So you see that the the file that was
  4235. 3:20:39created at 1814 was the latest file
  4236. 3:20:43right and this has been retained when we
  4237. 3:20:46run the vacuum command. So by now we've
  4238. 3:20:49seen deletion vectors, copy on write,
  4239. 3:20:52merge on read, all of it in action.
  4240. 3:20:54Yeah. So it's really important to
  4241. 3:20:57understand that neither paradigm,
  4242. 3:20:59neither copy on write nor merge on read
  4243. 3:21:03offer a silver bullet for all use cases.
  4244. 3:21:06Yeah. However, copy on write is really
  4245. 3:21:09good for use cases which is read heavy.
  4246. 3:21:12Yeah. And where the rights are very low.
  4247. 3:21:16What will happen if the rights are high?
  4248. 3:21:18It will simply read and rewrite the
  4249. 3:21:21whole file again and again and again,
  4250. 3:21:23right? And we want to avoid doing that
  4251. 3:21:25because it's a costly operation. Yeah.
  4252. 3:21:27So, copy and write is good for read
  4253. 3:21:30heavy use cases. And similarly, merge on
  4254. 3:21:34read works best for cases for use cases
  4255. 3:21:38where data is updated frequently. Right?
  4256. 3:21:42Why does it work best for use cases
  4257. 3:21:44where there are frequent updates? That
  4258. 3:21:47is because it is simply going to record
  4259. 3:21:49those changes in a deletion vector file.
  4260. 3:21:52It is not going to rewrite the whole
  4261. 3:21:55file. So ultimately it helps you reduce
  4262. 3:21:58the right latency because it is simply
  4263. 3:22:00recording it in a deletion vector. Yeah.
  4264. 3:22:03So I hope that helps you understand the
  4265. 3:22:06tradeoffs between the two and how
  4266. 3:22:08deletion vectors can help boost
  4267. 3:22:11performance. So now we are going to talk
  4268. 3:22:14about cloning. So cloning as you know
  4269. 3:22:18let us create a snapshot of our delta
  4270. 3:22:21tables at a specific point in time.
  4271. 3:22:24Right? And there are two flavors to
  4272. 3:22:27cloning. There are two approaches to
  4273. 3:22:30clone a delta table. The first one is a
  4274. 3:22:33shallow clone and the second one is a
  4275. 3:22:37deep clone. So let's understand both of
  4276. 3:22:39them in details. Shallow clone is
  4277. 3:22:42essentially taking the snapshot of
  4278. 3:22:45metadata at a particular point in time,
  4279. 3:22:47right? So when you shallow clone a
  4280. 3:22:49table, delta doesn't copy the underlying
  4281. 3:22:53data files, right? It simply references
  4282. 3:22:56them. So let's say there is a source
  4283. 3:22:58table and then there is a shallow clone
  4284. 3:23:00table. The shallow clone table is simply
  4285. 3:23:02going to reference the paret or the
  4286. 3:23:05underlying data files, right? It is not
  4287. 3:23:08going to copy those data file. So what
  4288. 3:23:11this means is that it makes the shallow
  4289. 3:23:14clone operation super fast,
  4290. 3:23:16computationally cheap and it doesn't eat
  4291. 3:23:19up a lot of storage. So let's understand
  4292. 3:23:22this with this diagram right here. Let's
  4293. 3:23:25say we have a source delta table and we
  4294. 3:23:28want to shallow clone this source delta
  4295. 3:23:30table into a target table. So what the
  4296. 3:23:33source table has is this metadata over
  4297. 3:23:36here in the delta
  4298. 3:23:39log folder and this is some bunch of
  4299. 3:23:42JSON files that you see over here right
  4300. 3:23:45and then it also has some data in the
  4301. 3:23:48form of park a file. Yeah. So now if we
  4302. 3:23:53want to shallow clone,
  4303. 3:23:56if we want to shallow clone this table,
  4304. 3:23:58the first thing that is going to happen
  4305. 3:24:01is a duplication or a replication of the
  4306. 3:24:05metadata itself, the metadata that we
  4307. 3:24:07have over here. And when I say
  4308. 3:24:08duplication, what I mean to say is that
  4309. 3:24:11it's not a one toone copy of these JSON
  4310. 3:24:14files, right? It's not that I'm going to
  4311. 3:24:15copy 0.json 1.json and then put it over
  4312. 3:24:18there, right? put it in the source in
  4313. 3:24:20the target table. Right? So that's
  4314. 3:24:22that's not the purpose. What is going to
  4315. 3:24:24happen is that all of this is going to
  4316. 3:24:27be condensed into a checkpoint.par
  4317. 3:24:30file and then the latest state of this
  4318. 3:24:34source table is going to be computed and
  4319. 3:24:37that is going to be put in zero.json. So
  4320. 3:24:40you would see that these two files would
  4321. 3:24:44be present in the delta log folder. in
  4322. 3:24:49the delta log folder of the target delta
  4323. 3:24:53table. Right? So this one is going to
  4324. 3:24:55you going to be used in order to compact
  4325. 3:24:58all of the information that you have
  4326. 3:24:59over here and we put it into a
  4327. 3:25:020.cheepoint.park
  4328. 3:25:04and the latest state of the source table
  4329. 3:25:07is computed and that is put inside
  4330. 3:25:100.json which becomes version number zero
  4331. 3:25:15of the target table. So the target table
  4332. 3:25:18is going to start with a fresh version
  4333. 3:25:21which is version number zero. Right? And
  4334. 3:25:24what happens to the data? So the park
  4335. 3:25:26file that I've shown over here, it
  4336. 3:25:28doesn't actually mean that there are
  4337. 3:25:30going to be park files, right? It is
  4338. 3:25:31just symbolic of the fact that there is
  4339. 3:25:34going to be data in this target table
  4340. 3:25:37and this data is going to reference all
  4341. 3:25:40of the park files that you see over
  4342. 3:25:42here. Right? So it is simply going to
  4343. 3:25:45reference all of the park file that you
  4344. 3:25:47see in the source table. And that is how
  4345. 3:25:51a shallow clone works under the hood.
  4346. 3:25:53Now let's understand what deep clone is.
  4347. 3:25:56Right? So a deep clone makes an
  4348. 3:25:59independent copy of both the metadata
  4349. 3:26:02and the data files. Right? So this
  4350. 3:26:05operation takes a little bit longer. It
  4351. 3:26:08uses more storage because now this time
  4352. 3:26:11it is not referencing the data of the
  4353. 3:26:13source file. Right? It is actually
  4354. 3:26:15copying the data of the source file in
  4355. 3:26:18the deep clone file in the target.
  4356. 3:26:21Right? So this is going to use more
  4357. 3:26:23storage but the result is that it is a
  4358. 3:26:25completely self-contained and an
  4359. 3:26:28independent table. Right? So let's
  4360. 3:26:31understand that with this diagram right
  4361. 3:26:34here. And again we have a source table
  4362. 3:26:37and we have a target table and we want
  4363. 3:26:41to deep clone the source into the target
  4364. 3:26:45right we want to deep clone the source
  4365. 3:26:48into the target and similar to the last
  4366. 3:26:50example we have metadata in the delta
  4367. 3:26:53log folder with a bunch of JSON files
  4368. 3:26:56that you see over here and some data
  4369. 3:26:59files and these data files are basically
  4370. 3:27:01parket right now when we do a deep clone
  4371. 3:27:04phone. The exact same thing is going to
  4372. 3:27:07happen for the metadata. All of this is
  4373. 3:27:11going to be condensed into a
  4374. 3:27:130.point.park
  4375. 3:27:16and the latest state is going to be
  4376. 3:27:18computed put inside 0.json which is
  4377. 3:27:22going to be version version number zero
  4378. 3:27:25right which is going to be version
  4379. 3:27:28number zero. So the deep clones table is
  4380. 3:27:31going to start at a fresh version which
  4381. 3:27:34is version number zero. Right? And as
  4382. 3:27:37you see over here all of these part
  4383. 3:27:39files right this one this one this one
  4384. 3:27:42and this one this one and this one and
  4385. 3:27:44this one and this one they are going to
  4386. 3:27:46be exactly replicated. Right? So this is
  4387. 3:27:51going to be exactly replicated.
  4388. 3:27:56So this is going to be a one to one
  4389. 3:27:58copy.
  4390. 3:28:00So that is how a deep clone is going to
  4391. 3:28:03work under the hood. Now let's see all
  4392. 3:28:06of this in action. Right. So first of
  4393. 3:28:09all we are going to start with shallow
  4394. 3:28:12clones.
  4395. 3:28:14We're going to start with shallow
  4396. 3:28:15clones. And for this purpose I am going
  4397. 3:28:17to create a new table with customer ids
  4398. 3:28:22from 1 to 100. And I'm going to use a
  4399. 3:28:27Cash statement.
  4400. 3:28:29I'm going to use a Cash statement. So
  4401. 3:28:31this is simply going to be create or
  4402. 3:28:34replace delta catalog delta DB dot
  4403. 3:28:38invoices
  4404. 3:28:40of customers
  4405. 3:28:42from 1 until 100, right? Because that is
  4406. 3:28:44what we are doing over here. And let me
  4407. 3:28:48just put as and this is going to be
  4408. 3:28:50create a replace table. And let's go
  4409. 3:28:52ahead and run this. Let's also quickly
  4410. 3:28:55do a select star just to see how this
  4411. 3:28:58looks like, right? Select star from this
  4412. 3:29:00table. Limit five.
  4413. 3:29:04Yeah.
  4414. 3:29:06Okay. So this looks fine. Now if I were
  4415. 3:29:09to do a describe history of this table
  4416. 3:29:14delta catalog. Delta DB.invoices
  4417. 3:29:18C 1 to 100. we simply going to have one
  4418. 3:29:21row that is the catas statement that
  4419. 3:29:24we've ran right so now in order to add
  4420. 3:29:27more history to this table let's perform
  4421. 3:29:30a few more operation yeah so let me go
  4422. 3:29:33ahead and do a delete from this table
  4423. 3:29:37where
  4424. 3:29:39customer ID is between
  4425. 3:29:4315 and 20 yeah let's run this let's run
  4426. 3:29:48another another statement which is an
  4427. 3:29:50update update this table where so what
  4428. 3:29:54do we want to update so let's update
  4429. 3:29:56customer ID equals 3 and let's change
  4430. 3:29:59the quantity from 3 to 10 right so this
  4431. 3:30:03is going to be where customer ID equals
  4432. 3:30:063 and I need to set this quantity to 10
  4433. 3:30:12yeah so let's run this again and let's
  4434. 3:30:16also quickly see how the layout looks
  4435. 3:30:19like in the file storage, right? So, if
  4436. 3:30:23we go to details
  4437. 3:30:26and this is the storage right here.
  4438. 3:30:31Okay, so here you see two files and then
  4439. 3:30:33two deletion vectors, right? So, I
  4440. 3:30:35believe the first file is for the first
  4441. 3:30:37instance when we did a creation using
  4442. 3:30:40the cas command, right? the deletion
  4443. 3:30:43vector. The first deletion vector is for
  4444. 3:30:45the delete command and the file the park
  4445. 3:30:48file and the updated deletion vector
  4446. 3:30:50that we see actually this is the updated
  4447. 3:30:53deletion vector because this is created
  4448. 3:30:56at 1739. So the updated deletion vector
  4449. 3:30:59that we see is because of the update
  4450. 3:31:01command and the updated row is contained
  4451. 3:31:03within this park file. Now let's go
  4452. 3:31:05ahead and run an insert statement. So
  4453. 3:31:08insert into this table
  4454. 3:31:12and then the values let me quickly copy
  4455. 3:31:16from
  4456. 3:31:19from this table right here. Right?
  4457. 3:31:23So this is let me put a random customer
  4458. 3:31:26ID which is 1099
  4459. 3:31:28and let's quickly format this. Okay. Now
  4460. 3:31:32that this is formatted let's go ahead
  4461. 3:31:34and run this.
  4462. 3:31:38Okay, so the insert statement is
  4463. 3:31:40complete. And if I were to go to the
  4464. 3:31:44file storage ADLS, we should see another
  4465. 3:31:48parket file over here that is going to
  4466. 3:31:49contain that insert statement, right?
  4467. 3:31:53There you go. So we see another parket
  4468. 3:31:56file over here, right? The one that ends
  4469. 3:31:58with 1 to CB - C0. Right? So now let's
  4470. 3:32:04do a describe history
  4471. 3:32:08of this table. We should see four rows.
  4472. 3:32:12Yeah. So the first one is for the cage.
  4473. 3:32:14The second one is for the delete between
  4474. 3:32:1615 to 20. The third one the third one is
  4475. 3:32:19the update where we updated customer ID
  4476. 3:32:22equal three. And the last one is the
  4477. 3:32:24insert over here. Right? So all of this
  4478. 3:32:27makes sense now. Now let's say we want
  4479. 3:32:31to do a shallow clone of this table.
  4480. 3:32:35Let's see what is going to happen. So
  4481. 3:32:36first of all let's write a create or
  4482. 3:32:40replace table
  4483. 3:32:43and this is going to be the same and
  4484. 3:32:45I'll just append it with shallow clone
  4485. 3:32:49and let's write shallow clone
  4486. 3:32:53and the table over here delta catalog
  4487. 3:32:54with this. Right? So let's go ahead and
  4488. 3:32:56run this now. Let's see what happened.
  4489. 3:33:00Okay. So if I were to look at the
  4490. 3:33:02catalog, this should show me this table.
  4491. 3:33:05So this table is created right here. And
  4492. 3:33:07if I were to look at the details, I
  4493. 3:33:09would find the path. And let's have a
  4494. 3:33:11look at what is there in this path. So
  4495. 3:33:14let me just open this in a new tab.
  4496. 3:33:23So now the interesting thing is you only
  4497. 3:33:26see the delta log right so we were
  4498. 3:33:29talking about that the files are
  4499. 3:33:31referenced so that is why you don't see
  4500. 3:33:33any file over here and let's also see
  4501. 3:33:36what there in the delta log so as we
  4502. 3:33:39discuss we see a 0 checkpoint and the
  4503. 3:33:44latest state in zero.json JSON. So I
  4504. 3:33:47hope all of that makes sense. Now you're
  4505. 3:33:48able to connect the dots. Let's look at
  4506. 3:33:50a few more details of the shallow clone
  4507. 3:33:54table. So if I were to say describe
  4508. 3:33:58history of the shallow clone table, I
  4509. 3:34:01should essentially see one command.
  4510. 3:34:03Yeah, which is the clone operation
  4511. 3:34:04itself. Now if you look at the operation
  4512. 3:34:07parameters, it tells me that the source
  4513. 3:34:09of the table is this table over here,
  4514. 3:34:13right? and the source version right the
  4515. 3:34:16source version the latest version of the
  4516. 3:34:18source table was number three and that
  4517. 3:34:21version was used to build the shallow
  4518. 3:34:25clone right so that is what it says over
  4519. 3:34:26here source version is number three and
  4520. 3:34:29is shallow equals true yeah now let's
  4521. 3:34:33update something in the source table and
  4522. 3:34:36see if it affects the shallow clone
  4523. 3:34:38table right so we have the source table
  4524. 3:34:42right here uh over here right so let me
  4525. 3:34:48do a select quickly do a select star
  4526. 3:34:52from this table
  4527. 3:34:55select start from this table and let me
  4528. 3:34:58do a limit five and let's go ahead and
  4529. 3:35:01run this so I want to update some row so
  4530. 3:35:05let me update customer ID equals 5 this
  4531. 3:35:08time so customer ID equals 5 has
  4532. 3:35:11quantity equals one but now I want to
  4533. 3:35:13update it to 10. So let's go ahead and
  4534. 3:35:16run this. Let's also see the history of
  4535. 3:35:19this table.
  4536. 3:35:22So now it should contain an update.
  4537. 3:35:25Right? So after an insert that we did
  4538. 3:35:27previously, it now has an update. So
  4539. 3:35:30let's quickly do a select star from this
  4540. 3:35:35table where customer ID equals 5. And
  4541. 3:35:39now I should see quantity equals 10.
  4542. 3:35:43Right? I see quantity equals 10. Let's
  4543. 3:35:46quickly verify
  4544. 3:35:48what is the resulting value in the
  4545. 3:35:51shallow clone table.
  4546. 3:35:55So he we see here that quantity equals 1
  4547. 3:35:59and customer ID equals 5. That means
  4548. 3:36:02even though the shallow clone is
  4549. 3:36:04referencing the source table after the
  4550. 3:36:08point that the shallow clone has been
  4551. 3:36:10cloned right the table the target
  4552. 3:36:12shallow clone table has been cloned it
  4553. 3:36:14is going to have its own history its own
  4554. 3:36:17set of operation and any activity on the
  4555. 3:36:20source table is not going to affect the
  4556. 3:36:24shallow clone right I hope this is very
  4557. 3:36:25clear now let's try something different
  4558. 3:36:27let's try deleting or modifying some
  4559. 3:36:31records in the shallow clone table and
  4560. 3:36:34let's see whether that affects the
  4561. 3:36:36source table or not. Right? So let's do
  4562. 3:36:39a delete from delta catalog delta db do
  4563. 3:36:45this table shallow clone where customer
  4564. 3:36:49ID equal 99. Right? So let's go ahead
  4565. 3:36:52and run this and let's also quickly run
  4566. 3:36:54select star from this table where
  4567. 3:36:58customer ID equals 99. So this shouldn't
  4568. 3:37:02return me any row and let's also have a
  4569. 3:37:05look at the history. Okay, so no rows
  4570. 3:37:07returned. Let's also have a look at the
  4571. 3:37:09history of this table and it has a clone
  4572. 3:37:13and then a delete the delete operation
  4573. 3:37:16that we ran just now. Now if I were to
  4574. 3:37:18compare this to
  4575. 3:37:22to the original table, ideally I should
  4576. 3:37:26find the record, right? Because the
  4577. 3:37:29operation that a shallow clone is going
  4578. 3:37:31to have is going to maintain its own
  4579. 3:37:34history after the point it has been
  4580. 3:37:36cloned. So any operation on the shallow
  4581. 3:37:39clone table shouldn't affect the source
  4582. 3:37:42table. Yeah. So let's go ahead and run
  4583. 3:37:44this. And there you go. we find that the
  4584. 3:37:47row number 99 customer ID equal 99 is
  4585. 3:37:50still there and let's also verify this
  4586. 3:37:53with the history just to make sure that
  4587. 3:37:56the history also remains unaffected
  4588. 3:37:58right so let's run this
  4589. 3:38:02and there is no trace of a delete
  4590. 3:38:04statement right so after the shallow
  4591. 3:38:06clone table is cloned it maintains its
  4592. 3:38:10own history right so now to quickly
  4593. 3:38:13summarize once A source table has been
  4594. 3:38:16shallow cloned. Changes or updates or
  4595. 3:38:20modifications or deletes to the shallow
  4596. 3:38:23clone table is not going to affect the
  4597. 3:38:25source table. And changes or updates or
  4598. 3:38:28modifications to the source table is not
  4599. 3:38:31going to affect the shallow clone table.
  4600. 3:38:33Right? The shallow clone table is only
  4601. 3:38:36going to refer to the data of the source
  4602. 3:38:39table until the point the clone
  4603. 3:38:42happened. Right? From that point
  4604. 3:38:44onwards, it is going to maintain its own
  4605. 3:38:47version, its own history, its own data
  4606. 3:38:49files. Right? So I hope that makes
  4607. 3:38:51sense. Now what I want to show you is
  4608. 3:38:54that we can also clone a table using a
  4609. 3:38:58particular timestamp or a particular
  4610. 3:39:00version number. Yeah. So what we can do
  4611. 3:39:03is that we can run a create or replace
  4612. 3:39:09table
  4613. 3:39:11and this is going to be shallow clone.
  4614. 3:39:14Let's let's create a clone of version
  4615. 3:39:17number this version right version number
  4616. 3:39:19zero v 0
  4617. 3:39:22and this is going to be shallow clone
  4618. 3:39:26delta catalog dot this table
  4619. 3:39:30version as of
  4620. 3:39:33zero right so let's run this now
  4621. 3:39:38so if this has worked properly
  4622. 3:39:42you should not see actually you should
  4623. 3:39:45see records from customer ID is 15 to 20
  4624. 3:39:49right so let's start from this table
  4625. 3:39:52where customer
  4626. 3:39:55where customer
  4627. 3:39:57ID
  4628. 3:39:59between 15 and 20 now if I were to
  4629. 3:40:03similarly run this on the original table
  4630. 3:40:08I shouldn't get any records right
  4631. 3:40:10because the latest version of this table
  4632. 3:40:12doesn't have those rows from 15 to 20.
  4633. 3:40:15Yeah. So let's run this.
  4634. 3:40:18And there you see 15 to 20. But in the
  4635. 3:40:21original table there should be no rows
  4636. 3:40:22return. Okay. That works. And another
  4637. 3:40:26variation of this is that you can also
  4638. 3:40:30do a time stamp as of and you can
  4639. 3:40:33basically put pick up any time stamp
  4640. 3:40:35over here. Right? So you can pick this
  4641. 3:40:38one up and write something like this.
  4642. 3:40:40Right? Now let's talk about time travel.
  4643. 3:40:43So if I were to quickly show you,
  4644. 3:40:46okay, let me just note this down. Time
  4645. 3:40:50travel.
  4646. 3:40:51If I were to show you describe history
  4647. 3:40:55of this table, we are going to have
  4648. 3:40:59several rows, right? Because we
  4649. 3:41:01performed several operation. Now what
  4650. 3:41:05I've already mentioned earlier is that
  4651. 3:41:08when we create a shallow clone it starts
  4652. 3:41:11from version zero. It takes up the
  4653. 3:41:13latest state of the source table and it
  4654. 3:41:16starts from version zero. So if I were
  4655. 3:41:18to do something like this,
  4656. 3:41:21create or replace
  4657. 3:41:25table
  4658. 3:41:27and let me put this as
  4659. 3:41:32test shallow clone as shallow clone
  4660. 3:41:38this whole table.
  4661. 3:41:41Right? So if I were to run this as we've
  4662. 3:41:45already seen earlier
  4663. 3:41:49as we've already seen earlier this is
  4664. 3:41:51going to create this table taking the
  4665. 3:41:55latest version of this table over here
  4666. 3:41:57right so if you do a history on this
  4667. 3:42:00table it is only going to contain one
  4668. 3:42:02row right
  4669. 3:42:07so I'm reiterating all of this right
  4670. 3:42:09because I want to make sure that this is
  4671. 3:42:12super clear.
  4672. 3:42:15So you see that there is only one row
  4673. 3:42:18that means version number zero of this
  4674. 3:42:21table has been created using the latest
  4675. 3:42:25state of the source table. So the source
  4676. 3:42:27table has many versions right it has
  4677. 3:42:29version 0 1 until 4. If I want to roll
  4678. 3:42:32back to a particular version or if I
  4679. 3:42:34want to see a particular version of the
  4680. 3:42:36source table I can definitely do that
  4681. 3:42:38right. I can see version number three,
  4682. 3:42:40version number two, version number one
  4683. 3:42:42and so on. But just because we have
  4684. 3:42:46shallow cloned this table, it doesn't
  4685. 3:42:48mean that we can access the history of
  4686. 3:42:51the sort table. We cannot go back to
  4687. 3:42:53version one or two for this shallow
  4688. 3:42:57clone table. Right? So I hope that is
  4689. 3:42:59clear because we are going to start with
  4690. 3:43:02a fresh history with a fresh version
  4691. 3:43:04with version number zero taking the
  4692. 3:43:07latest state of the source table. Now
  4693. 3:43:10let's try out something interesting.
  4694. 3:43:12Let's run a vacuum on the source table.
  4695. 3:43:15And by source what I mean is this table
  4696. 3:43:18right here invoices_c1_00.
  4697. 3:43:22Right? So if you remember when we
  4698. 3:43:24created this table
  4699. 3:43:27when we created this table right over
  4700. 3:43:29here we ran a few statements in order to
  4701. 3:43:32populate the history. The first one was
  4702. 3:43:34a delete the second one was an update
  4703. 3:43:37and the third one was an insert and I
  4704. 3:43:41also pointed out that this is the park
  4705. 3:43:45file where the value of that insert was
  4706. 3:43:48stored. Right? So
  4707. 3:43:51the the customer id that is belonging to
  4708. 3:43:54this insert is 1099
  4709. 3:43:56and this table later on was cloned
  4710. 3:44:00right. So that means that this row where
  4711. 3:44:02customer ID equals 1099 is present in
  4712. 3:44:05the source table it is also present in
  4713. 3:44:08the shallow clone table right. So the
  4714. 3:44:10shallow clone table is also referencing
  4715. 3:44:12this row. Now if what happens if I were
  4716. 3:44:16to run a delete? So let me just quickly
  4717. 3:44:19add this
  4718. 3:44:21heading and
  4719. 3:44:23let's say I run a delete. Delete from
  4720. 3:44:26this table.
  4721. 3:44:30Delete from this table
  4722. 3:44:32where customer ID equals 1099.
  4723. 3:44:38And this should ideally
  4724. 3:44:41remove
  4725. 3:44:44remove this row, right? So, like start
  4726. 3:44:45from this table where customer ID equals
  4727. 3:44:511099, right? So, this should not give me
  4728. 3:44:53any result. Okay, that works perfectly
  4729. 3:44:56fine now
  4730. 3:44:58because this file is now orphaned,
  4731. 3:45:01right? This file is not being referred
  4732. 3:45:04in the latest version when I run a
  4733. 3:45:06vacuum. This file should just go away,
  4734. 3:45:09right? That is how vacuum is going to
  4735. 3:45:11behave. So first of all I have to set
  4736. 3:45:13this property set spark dot databick
  4737. 3:45:18dot delta dot retention
  4738. 3:45:22duration
  4739. 3:45:25enabled is false and then let's go ahead
  4740. 3:45:29and run a vacuum vacuum this table
  4741. 3:45:33retain zero hours right let me go ahead
  4742. 3:45:37and run this and ideally if everything
  4743. 3:45:39works out fine I shouldn't see this park
  4744. 3:45:43file which end with 12 CB - C0 right
  4745. 3:45:48let me go ahead and run this
  4746. 3:45:52okay this is complete and let me refresh
  4747. 3:45:55this
  4748. 3:45:58okay so one of the deletion vectors was
  4749. 3:46:01cleared which is good but I don't see
  4750. 3:46:05the file deleted right the one ending
  4751. 3:46:09with 12 CB hy hyphen C0 is not deleted.
  4752. 3:46:13Why that may be the case? Right? So can
  4753. 3:46:16you think of why this may have happened?
  4754. 3:46:20So the reason this has happened is
  4755. 3:46:23because although this file in the source
  4756. 3:46:25has been deleted, there is a shallow
  4757. 3:46:28clone which is still referencing that
  4758. 3:46:30row. Right? So that is why that file had
  4759. 3:46:33not been deleted. So let me
  4760. 3:46:36let me show you
  4761. 3:46:39if I were to run this on the shallow
  4762. 3:46:42clone table.
  4763. 3:46:47This is there right now. Let me go ahead
  4764. 3:46:49and delete this from all of the shallow
  4765. 3:46:51clone that I've created. So by now you
  4766. 3:46:54must have seen that I created a few
  4767. 3:46:55shallow clones and we want to remove all
  4768. 3:46:58of the references of that row 1099.
  4769. 3:47:01Right? So this would be delete from
  4770. 3:47:05delta catalog dot this table where
  4771. 3:47:10customer ID equal 1099.
  4772. 3:47:15What were the other table that I
  4773. 3:47:17created? It was v 0
  4774. 3:47:21and
  4775. 3:47:23what else?
  4776. 3:47:25So I created v 0 and then underscore
  4777. 3:47:28test_cccl
  4778. 3:47:31test_cl.
  4779. 3:47:33So I want to remove all of the
  4780. 3:47:35references. Right? So now what this
  4781. 3:47:38means is that this was the source table
  4782. 3:47:41and there were several shallow clones.
  4783. 3:47:44Some of them were referring to row
  4784. 3:47:45number 1099. Now I have removed all of
  4785. 3:47:48the references. So that file is now an
  4786. 3:47:51orphaned file. There is no reference. So
  4787. 3:47:53once I run a vacuum that file should be
  4788. 3:47:56gone now, right? So now let's go ahead
  4789. 3:47:59and run this vacuum once again.
  4790. 3:48:03Okay. So this is complete now. So I
  4791. 3:48:06believe this file should be gone if
  4792. 3:48:10whatever logic that we were discussing
  4793. 3:48:12is true, right? So let's refresh. And
  4794. 3:48:16there you see. So you have only three
  4795. 3:48:19files and that file is gone because
  4796. 3:48:22there are no more references to that
  4797. 3:48:25park file anymore. Right now let's see
  4798. 3:48:27deep clone in action. So let me quickly
  4799. 3:48:30put that down as a heading
  4800. 3:48:32and I want to clone
  4801. 3:48:36the source table that we've been using
  4802. 3:48:38so far which is this one. So let me
  4803. 3:48:40quickly see the history. Okay. So
  4804. 3:48:43there's a lot of thing that we've
  4805. 3:48:44performed on this table, right? And let
  4806. 3:48:48me also
  4807. 3:48:50let me also
  4808. 3:48:52see how the file looks like. Right? So
  4809. 3:48:54this is how the files look like. Right?
  4810. 3:48:56Now let me go ahead and perform a deep
  4811. 3:49:00clone. Create or replace table
  4812. 3:49:05this table. And this is going to be a
  4813. 3:49:08deep clone. Let me also remove this and
  4814. 3:49:14let me write this as DCL which is deep
  4815. 3:49:18clone. Right? So let's go ahead and run
  4816. 3:49:22this. Okay. So this has created a deep
  4817. 3:49:26cloned table. Let's see describe history
  4818. 3:49:32of this table. And ideally we should
  4819. 3:49:34just see one row which is the clone.
  4820. 3:49:36Perfect. And if we look at the
  4821. 3:49:39parameters, we see that this is being
  4822. 3:49:42cloned from this table and the source
  4823. 3:49:44version is 13, right? 13 the latest
  4824. 3:49:47version. Yeah. So let's have a look at a
  4825. 3:49:51few other things. So if I were to
  4826. 3:49:55check how the files look like, right? So
  4827. 3:49:59let me have a look at how the files look
  4828. 3:50:01like.
  4829. 3:50:06So what you see over here is that these
  4830. 3:50:09two are an almost an exact copy. So
  4831. 3:50:13there are three parket files and one
  4832. 3:50:16deletion vector and that is what you see
  4833. 3:50:19over here exactly. So these are three
  4834. 3:50:20parket files and one deletion vector and
  4835. 3:50:23this is a one one copy right. So we see
  4836. 3:50:28let's let's quickly check the numbers.
  4837. 3:50:29So this is 967
  4838. 3:50:31F2B
  4839. 3:50:33and 3d2 right. So this is 967 F2B and
  4840. 3:50:383d2 right and the deletion vector end
  4841. 3:50:40with 9 F9. It also ends with 9 F9.
  4842. 3:50:44Right? So that means that this is an
  4843. 3:50:47exact copy from the source. But for the
  4844. 3:50:51delta log we discussed that there is
  4845. 3:50:52going to be a 0.choint.park
  4846. 3:50:55and then 0.json JSON which is going to
  4847. 3:50:57generate the latest state
  4848. 3:51:01and exactly that is what we see over
  4849. 3:51:03here. Yeah. So now let's do a few other
  4850. 3:51:06things. Let's make some changes in the
  4851. 3:51:08in the deep clone and see if it affects
  4852. 3:51:11the source table. Ideally it shouldn't
  4853. 3:51:13because we mentioned earlier that the
  4854. 3:51:15deep clone is an independent and a
  4855. 3:51:18self-contained copy. Right? So let's
  4856. 3:51:22make a few changes now. First of all,
  4857. 3:51:25let me do select star from
  4858. 3:51:28this table right here. And let's do a
  4859. 3:51:32limit five. So let's run this.
  4860. 3:51:36So I'm going to write an update
  4861. 3:51:39from
  4862. 3:51:41sorry update this table where I'm going
  4863. 3:51:46to set the quantity
  4864. 3:51:50of
  4865. 3:51:52customer ID equals 4.
  4866. 3:51:55I'm going to set the quantity to 10
  4867. 3:51:58where customer ID equals 4.
  4868. 3:52:02And let's go ahead and run this. So
  4869. 3:52:03let's quickly verify this from this
  4870. 3:52:07table where customer ID equals 4. This
  4871. 3:52:13should now be 10 instead of five. And
  4872. 3:52:18that is what we see over here. Now let
  4873. 3:52:20me quickly verify this in the original
  4874. 3:52:24table.
  4875. 3:52:27The original table still shows five.
  4876. 3:52:30Right? So as we expected that any change
  4877. 3:52:34in the cloned table in the deep clone
  4878. 3:52:38table is not going to affect anything in
  4879. 3:52:41the source table and vice versa is going
  4880. 3:52:43to be true. any changes in the source
  4881. 3:52:45table is not going to affect the deep
  4882. 3:52:49cloned table. Right now, I'm not going
  4883. 3:52:52to run and show you a vacuum on a deep
  4884. 3:52:55clone because as I mentioned earlier, a
  4885. 3:52:57deep clone is a self-contained and an
  4886. 3:53:01independent copy, right, of the source
  4887. 3:53:03table. There are no references between
  4888. 3:53:05the deep clone and the source table. So
  4889. 3:53:08if we were to quickly summarize what
  4890. 3:53:10happened in the shallow clone vacuum. So
  4891. 3:53:14we had a file right we had a file where
  4892. 3:53:16customer ID equal 1099 and there were
  4893. 3:53:19several shallow clones referring to that
  4894. 3:53:22row. Now when we performed a delete and
  4895. 3:53:24then we did a vacuum still that file was
  4896. 3:53:28not removed from the source table. Why?
  4897. 3:53:32Because shallow clone there were three
  4898. 3:53:33shallow clones which were referring to
  4899. 3:53:36customer ID equals 199 1099. Once all of
  4900. 3:53:40those references were gone, right? Once
  4901. 3:53:43we deleted all of the references, we ran
  4902. 3:53:46a delete statement which meant that that
  4903. 3:53:49row is no longer in the shallow clone.
  4904. 3:53:53Those references were deleted. And then
  4905. 3:53:55we ran a vacuum on the source table. We
  4906. 3:53:57saw that the park file was deleted.
  4907. 3:54:00Right? But in case of a deep clone, the
  4908. 3:54:03source table and the deep clone are
  4909. 3:54:06completely different. Different in the
  4910. 3:54:08sense that they are independent copies
  4911. 3:54:11of each other. There is no reference
  4912. 3:54:13from the deep clone to the source table.
  4913. 3:54:16Right? So that is the reason why vacuums
  4914. 3:54:19should be completely independent. Right?
  4915. 3:54:22So before we conclude this section on
  4916. 3:54:24deep and shallow clones, I believe that
  4917. 3:54:26a good number of you would have this
  4918. 3:54:28question that how are cats different
  4919. 3:54:31from deep clone. So the tables that are
  4920. 3:54:34generated using cas create or replace
  4921. 3:54:36table and then a select query. How is
  4922. 3:54:39that different from the tables that are
  4923. 3:54:42generated using deep clone? So if you
  4924. 3:54:45were to look at it at an output level,
  4925. 3:54:48so let's say you're doing a select star
  4926. 3:54:50from whatever table, right? If you were
  4927. 3:54:52to look at the outputs that are being
  4928. 3:54:53generated by the two tables, one which
  4929. 3:54:56is generated using CAS and the other one
  4930. 3:54:59that is generated using deep clone, they
  4931. 3:55:01would look exactly identical, right? But
  4932. 3:55:04that is not where the difference lies.
  4933. 3:55:07So when you deep clone a table, you also
  4934. 3:55:10clone, you don't need to respspecify the
  4935. 3:55:13partitioning properties, the constraints
  4936. 3:55:15and all of that. Right? With CASS, the
  4937. 3:55:18table is just created using the output
  4938. 3:55:20of a select query, right? So you write a
  4939. 3:55:22select and then it generates a set of
  4940. 3:55:24rows and it just uses those rows to
  4941. 3:55:27create a table. All of the properties
  4942. 3:55:30are lost and that is where deep clone
  4943. 3:55:33comes in handy. It's a robust way to
  4944. 3:55:35clone the metadata, the data and the
  4945. 3:55:39properties of the table. Another very
  4946. 3:55:42important advantage is that it works in
  4947. 3:55:45an incremental manner. So what I mean is
  4948. 3:55:48that let me let me show that to you with
  4949. 3:55:50an example. So let's say we have a
  4950. 3:55:54disaster recovery use case, right? So
  4951. 3:55:57let's say we have a
  4952. 3:55:59disaster recovery
  4953. 3:56:02use case. And by this what I mean is
  4954. 3:56:05that let's say we have a source table
  4955. 3:56:06over here. we have a source
  4956. 3:56:11and then what I'm doing is that I am
  4957. 3:56:14having a replica.
  4958. 3:56:16So this is the replica
  4959. 3:56:20and I am syncing these two tables.
  4960. 3:56:23Right?
  4961. 3:56:25What I did first is that I did a deep
  4962. 3:56:27clone.
  4963. 3:56:30I did a deep clone of the source table.
  4964. 3:56:33Right? Now what happens when I'm going
  4965. 3:56:35to get an update?
  4966. 3:56:38So let's say there is an update
  4967. 3:56:41on the source table, right? And then
  4968. 3:56:44there was a delete
  4969. 3:56:49on this source table. What is going to
  4970. 3:56:51happen? So I would need to bring these
  4971. 3:56:53two tables again back in sync. Right? So
  4972. 3:56:56in order to bring this back in sync
  4973. 3:56:58again, I would need to do a
  4974. 3:57:02deep clone again, right? So I would end
  4975. 3:57:05up doing a deep clone again. But the
  4976. 3:57:06beauty of this is that it doesn't copy
  4977. 3:57:09the whole source
  4978. 3:57:11into the whole replica, right? It
  4979. 3:57:13doesn't copy the whole thing again,
  4980. 3:57:15right? It only copies incremental
  4981. 3:57:18changes. What it is going to do is that
  4982. 3:57:20it is going to take this update and
  4983. 3:57:22apply it on the replica. It is going to
  4984. 3:57:24take this delete and apply this on the
  4985. 3:57:27replica. Right? So the whole table is
  4986. 3:57:29not going to be copied again and again.
  4987. 3:57:32The operations are copied or synced in
  4988. 3:57:35an incremental manner. And that is the
  4989. 3:57:37reason why deep clones are robust and
  4990. 3:57:41they are very performant. So now we are
  4991. 3:57:43going to talk about something that
  4992. 3:57:45silently kills your spark performance
  4993. 3:57:48and that is the small file problem.
  4994. 3:57:54Yeah, the small file problem
  4995. 3:57:59and we are going to understand what
  4996. 3:58:01exactly the small file problem is, why
  4997. 3:58:03does it happen and how this can be fixed
  4998. 3:58:07using delta lakes optimize command.
  4999. 3:58:11Yeah. So let's first understand what the
  5000. 3:58:15small file problem is and why exactly is
  5001. 3:58:18it problematic. Yeah. So now I'm going
  5002. 3:58:20to take a weird example but it is just
  5003. 3:58:24to help you get the right understanding.
  5004. 3:58:26Yeah. So imagine you are reading a PDF.
  5005. 3:58:31So imagine you're reading a PDF of a
  5006. 3:58:34300page novel
  5007. 3:58:37of a 300page novel. Right? So naturally
  5008. 3:58:40you would expect the novel to be a
  5009. 3:58:44single 300page PDF, isn't it? So you
  5010. 3:58:47would expect it to be a single
  5011. 3:58:51300page PDF, right? But what if what if
  5012. 3:58:55I would say that someone saved each of
  5013. 3:58:58these pages as a PDF? So what they did
  5014. 3:59:01was they saved
  5015. 3:59:04they saved one PDF, two PDF
  5016. 3:59:09until all the way until 300. PDF. So
  5017. 3:59:13what they did is that they took each of
  5018. 3:59:16those pages, each of those 300 pages and
  5019. 3:59:19saved it as a PDF. So how would you read
  5020. 3:59:23this kind of a book, right? So first of
  5021. 3:59:26all, you would end up saying some very
  5022. 3:59:28nice words to that person who did this
  5023. 3:59:31which I cannot say on this video. But
  5024. 3:59:33then what you would essentially do is
  5025. 3:59:35that you would open the first PDF.
  5026. 3:59:37Actually, you would look for the first
  5027. 3:59:39PDF, right? So first of all you would
  5028. 3:59:41look for the first PDF look or find the
  5029. 3:59:44first PDF right and then you would open
  5030. 3:59:47the first PDF and you would read it and
  5031. 3:59:51finally when the reading is complete you
  5032. 3:59:54would close the PDF right now you would
  5033. 3:59:57have to do this for all of the 300 pages
  5034. 4:00:01that have been given to you and what we
  5035. 4:00:04see here is that we've spent a good
  5036. 4:00:07amount of time first of all finding the
  5037. 4:00:10right file and then opening it and then
  5038. 4:00:13finally closing it. Right? So there is a
  5039. 4:00:16good amount of time that you spend in
  5040. 4:00:19all of these problems in all of these
  5041. 4:00:22operations. Right? And that is exactly
  5042. 4:00:25what the small file problem is. When you
  5043. 4:00:27have a small pile problem, there are too
  5044. 4:00:30many small files and then you end up
  5045. 4:00:32doing such operations, right? you end up
  5046. 4:00:36first of all finding the right file and
  5047. 4:00:38then opening it and closing it right and
  5048. 4:00:41these are time consuming operation. So
  5049. 4:00:44to reiterate, when you're reading a
  5050. 4:00:47table and instead of having a few neatly
  5051. 4:00:50packed file, if your data is scattered
  5052. 4:00:53over thousands of smaller files, there
  5053. 4:00:56is going to be thousands of those open,
  5054. 4:00:59close and metadata lookup or finding
  5055. 4:01:03operation that we just discussed. All of
  5056. 4:01:06this leading to wasted compute and poor
  5057. 4:01:08IO, right? So all of this all of these
  5058. 4:01:11operations is simply going to lead to
  5059. 4:01:14wasted compute
  5060. 4:01:18and poor IO
  5061. 4:01:20right so how does delta solve this
  5062. 4:01:24problem delta provides a builtin
  5063. 4:01:27operation which is called optimize and
  5064. 4:01:30what it simply does is that it compacts
  5065. 4:01:33all of these small files into larger
  5066. 4:01:36more appropriately sized one right And
  5067. 4:01:39in order to do this, it simply uses
  5068. 4:01:42something called a bin packing
  5069. 4:01:45algorithm.
  5070. 4:01:47It uses a bin packing algorithm, right?
  5071. 4:01:51So we are going to see this in detail.
  5072. 4:01:53So don't worry about it. Nevertheless,
  5073. 4:01:55it's actually a very simple algorithm.
  5074. 4:01:57What it does is that it collects all
  5075. 4:01:59your file sizes and then it sorts it
  5076. 4:02:02from high to low and then it start
  5077. 4:02:04picking each file and it places each of
  5078. 4:02:08these files in a particular bin and
  5079. 4:02:10think of this bin as one large file.
  5080. 4:02:14Right? So you have lot of small files
  5081. 4:02:17over here. Right? You have lot of small
  5082. 4:02:20files. What it does is that it simply
  5083. 4:02:24takes this and it puts them in a bigger
  5084. 4:02:29bin as long as it fits in the bin.
  5085. 4:02:32Right? So let's say this simply is able
  5086. 4:02:36to fit in this bin. This is also able to
  5087. 4:02:38fit in this bin. This is also able to
  5088. 4:02:40fit in this bin. But this one is not
  5089. 4:02:43able to fit in this bin. Right? So in
  5090. 4:02:45that case a new bin is going to be
  5091. 4:02:48created and then this is going to go
  5092. 4:02:52over here. Right? So that is how simply
  5093. 4:02:55the bin packing algorithm works. And I'm
  5094. 4:02:58going to show you a visualization in
  5095. 4:03:01order to help help you to understand
  5096. 4:03:03this better. But before we go there I
  5097. 4:03:05want you to understand that the default
  5098. 4:03:09bin size that delta targets is 1 GBTE.
  5099. 4:03:14Right? So this is going to be 1 GBTE.
  5100. 4:03:16That is what is the default for Delta
  5101. 4:03:20Lake, right? And this default 1GB size
  5102. 4:03:23has been selected after years of usage
  5103. 4:03:27and testing on different kind of spark
  5104. 4:03:29workloads and it's really proven to be
  5105. 4:03:31robust. Right? So this number 1GB has
  5106. 4:03:35really proven to be robust. So unless
  5107. 4:03:38you have a very good reason to change
  5108. 4:03:40it, don't change it. go ahead and use
  5109. 4:03:42the default 1GB number. Yeah. So just to
  5110. 4:03:46let you know just in case you want to
  5111. 4:03:48change it, there is a property which is
  5112. 4:03:51called spark dot databick
  5113. 4:03:55dot delta dot optimize
  5114. 4:03:58dot max file size. So look it up on the
  5115. 4:04:01internet
  5116. 4:04:04dot max file size. So you can use this
  5117. 4:04:07property to change the 1GB default
  5118. 4:04:09number but my advice is again don't
  5119. 4:04:12change it unless you have a very
  5120. 4:04:14compelling reason to do so. Yeah. So now
  5121. 4:04:16let's see the visualization. So here we
  5122. 4:04:20have a list of files a list of parakeet
  5123. 4:04:23files and it has simply been sorted in
  5124. 4:04:26descending order by file sizes right. So
  5125. 4:04:29the first one you see is 320 MB and the
  5126. 4:04:31last one is 80 MB. So we are going to
  5127. 4:04:33run through how bins are going to be
  5128. 4:04:36created right and again we are assuming
  5129. 4:04:38that the target size of one bin one file
  5130. 4:04:42one large file is going to be 1 GBTE
  5131. 4:04:45right so let's pick up the first file
  5132. 4:04:46and let's see what happens the first
  5133. 4:04:48file park number seven is picked up and
  5134. 4:04:51it is simply placed in the first bin and
  5135. 4:04:53this is 1,000 megabytes of size out of
  5136. 4:04:56which 320 has been occupied now let's
  5137. 4:05:00pick up the next one the next one is 310
  5138. 4:05:02megabytes. Of course, the first bin has
  5139. 4:05:05that much of size. So, it is going to be
  5140. 4:05:08placed in the first bin. Now, let's go
  5141. 4:05:10to the third one. The third one is 280
  5142. 4:05:13MB. This does have 280 MB of size. So,
  5143. 4:05:17it is going to be placed in the first
  5144. 4:05:19bin again. Now, the first bin has only
  5145. 4:05:2390 MB of size left. Now, let's go over
  5146. 4:05:27to the next file. But the next file is
  5147. 4:05:29270 MB. So this for this a new bin is
  5148. 4:05:33going to be created. Right? So it tries
  5149. 4:05:34to put it in the first bin but it
  5150. 4:05:36cannot. So that is why a new bin is
  5151. 4:05:39created. Yeah. Now let's go over to the
  5152. 4:05:41next file. The next file is 220 MB. The
  5153. 4:05:44first bin only has 90 MB. So that is why
  5154. 4:05:47it's going to be placed in the second
  5155. 4:05:49bin. Now yeah. Now let's go over to the
  5156. 4:05:51next park file which is park number 9.
  5157. 4:05:54190 megabits of megabytes of size. It
  5158. 4:05:57cannot be placed in the first bin. So
  5159. 4:06:00let's see that it cannot be placed in
  5160. 4:06:01the first bin. So it goes over to the
  5161. 4:06:03next bin. Yeah. The next park file is
  5162. 4:06:07180 megabytes of size. Again it cannot
  5163. 4:06:09be placed in the first bin and it should
  5164. 4:06:11go in the second bin. Next one. Again
  5165. 4:06:14160 MB of size here. Cannot be placed in
  5166. 4:06:18the first bin. How much size does the
  5167. 4:06:20second bin have now? It has 140 MB of
  5168. 4:06:24size left. So it cannot be placed in the
  5169. 4:06:27second bin also. Right. So it goes to
  5170. 4:06:30the second bin but it cannot be placed.
  5171. 4:06:32So that is why a new bin is going to be
  5172. 4:06:34created. Yeah. Simple. So let's go over
  5173. 4:06:37to the next one. 150 megabytes of size.
  5174. 4:06:40It cannot be placed in the first one.
  5175. 4:06:42Cannot be placed in the second one also
  5176. 4:06:44but can be placed in the third one.
  5177. 4:06:46Yeah. Let's go over to the next one. 130
  5178. 4:06:49megaby cannot be placed in the first
  5179. 4:06:51one. But it can be placed in the second
  5180. 4:06:54one. Right? because one because the
  5181. 4:06:57second one has 140 mgabytes of size
  5182. 4:07:00left. So, it is going to go in the
  5183. 4:07:02second bin. Right? Let's pick up the
  5184. 4:07:05next one. Again, it cannot go to the
  5185. 4:07:06first bin. It cannot go to the second
  5186. 4:07:09bin. So, it is going to end up in the
  5187. 4:07:11third bin.
  5188. 4:07:13Now, the last one, we see that the first
  5189. 4:07:15bin does have 90 mgabytes of size left.
  5190. 4:07:19So, it is going to end up in the first
  5191. 4:07:22bin itself, right? And that is how we
  5192. 4:07:26see that three bins are created nearing
  5193. 4:07:301 GB of size and all of the files you
  5194. 4:07:33had so many files it has been reduced to
  5195. 4:07:36smaller number of files. Right? So now
  5196. 4:07:39if we correlate this back to the PDF
  5197. 4:07:42example instead of opening 300 PDFs one
  5198. 4:07:45by one probably you will have four to
  5199. 4:07:48five PDF. Right? Okay. So now let's see
  5200. 4:07:51all of this in action and for this I'm
  5201. 4:07:53going to create a notebook a new
  5202. 4:07:56notebook which is going to be named 07
  5203. 4:08:00optimize right and let's use this data
  5204. 4:08:04set that we've already been using
  5205. 4:08:06earlier for quite a good number of time.
  5206. 4:08:08So let's simply do a spark read.par
  5207. 4:08:13and together with this let's also
  5208. 4:08:15display five rows from this data set.
  5209. 4:08:19Yeah.
  5210. 4:08:22So now while we are doing this
  5211. 4:08:25let me also print the total number of
  5212. 4:08:27rows and along with that. Okay. So here
  5213. 4:08:31is what it looks like and let me also do
  5214. 4:08:34a df do select category dodistinct
  5215. 4:08:39dot count. Right? So the reason why I'm
  5216. 4:08:42doing a distinct of the column category
  5217. 4:08:44which has eight unique categories is
  5218. 4:08:47because I want to partition by this
  5219. 4:08:49column. Right? So let's quickly do a df
  5220. 4:08:53dot
  5221. 4:08:54repartition
  5222. 4:08:56five dot write dot mode override
  5223. 4:09:02dot partition by we partition by the
  5224. 4:09:06category column dot save as table and
  5225. 4:09:10I'm going to simply save as delta
  5226. 4:09:12catalog dot deltadb dot optimize example
  5227. 4:09:17one right so let's just save it like
  5228. 4:09:19this and the Reason why I'm doing a
  5229. 4:09:21repartition five is because I want to
  5230. 4:09:24mimic the small file behavior. So each
  5231. 4:09:27of the partitions that you see over here
  5232. 4:09:29that is going to be created for the
  5233. 4:09:31column category. So you're going to see
  5234. 4:09:33one partition for clothing, one
  5235. 4:09:34partition for toys and so on. Each of
  5236. 4:09:37these partitions should have five file.
  5237. 4:09:40Yeah. So let's go ahead and run this.
  5238. 4:09:44Okay. So this is complete.
  5239. 4:09:47Let me go to the catalog.
  5240. 4:09:51And
  5241. 4:09:54I simply take this up.
  5242. 4:09:57Take the table ID.
  5243. 4:10:00And
  5244. 4:10:01let me go to the container
  5245. 4:10:04tables. And then here is this table. And
  5246. 4:10:08you see all of the partitions over here.
  5247. 4:10:10Each of these partitions should have
  5248. 4:10:12five files, right? There you go. So you
  5249. 4:10:15see five files over here. And that is
  5250. 4:10:16because of the repartition pipe that we
  5251. 4:10:19put over there. Yeah. So now let's run a
  5252. 4:10:21query. And let's also time this query.
  5253. 4:10:25Let's do a df example one equals
  5254. 4:10:30this. I read in the table. And let's say
  5255. 4:10:33DF output equals
  5256. 4:10:37DF example 1 dot
  5257. 4:10:40where DF example one dot category equal
  5258. 4:10:44equals clothing
  5259. 4:10:46dot collect right and let's also let's
  5260. 4:10:51also print okay let's let's go ahead and
  5261. 4:10:53run this now so let's run this and see
  5262. 4:10:55what happens what is the total time that
  5263. 4:10:57is taken 2.3 seconds is the total time
  5264. 4:11:00that is taken to run this query. Right
  5265. 4:11:03now let's go ahead and run a optimize
  5266. 4:11:06run a compaction on top of this table.
  5267. 4:11:08Right? So basically the five small files
  5268. 4:11:12that we see that should be combined into
  5269. 4:11:15one bin one file as a result of this
  5270. 4:11:17optimize operation. So we simply do from
  5271. 4:11:20delta dot table import
  5272. 4:11:24delta table
  5273. 4:11:26and
  5274. 4:11:27table equals delta table dot for name.
  5275. 4:11:32Okay, this is great. That is what I
  5276. 4:11:34wanted to run. So let's go ahead and run
  5277. 4:11:36this. Yeah,
  5278. 4:11:45perfect. So now this is complete. What I
  5279. 4:11:48should see over here is instead of five
  5280. 4:11:51files, I should see just one file
  5281. 4:11:54because they've all been compacted into
  5282. 4:11:56one file, right? So, let me run a
  5283. 4:11:59refresh.
  5284. 4:12:02Okay, so now instead of one file, I see
  5285. 4:12:04six files, right? All of the files were
  5286. 4:12:07written at 185515.
  5287. 4:12:10But then there is this one file which
  5288. 4:12:12just got written now at 185731.
  5289. 4:12:16And if you look at the size, right, the
  5290. 4:12:18size is bigger than all of this, right?
  5291. 4:12:20So what happened is it combined all of
  5292. 4:12:24those five files into one file, right?
  5293. 4:12:27And that and this file right here is
  5294. 4:12:29that one file. What happened to the
  5295. 4:12:32other file? They still stay there, but
  5296. 4:12:34they have been tombstone. What tombstone
  5297. 4:12:37means is that they have been marked for
  5298. 4:12:40soft delete. Right? So if you look at
  5299. 4:12:44the catalog over here,
  5300. 4:12:48sorry the delta log over here, not the
  5301. 4:12:50catalog. What we should see over here is
  5302. 4:12:53if I were to look at category clothing,
  5303. 4:12:56we see that there are five removes.
  5304. 4:12:59Yeah. And there should be one add
  5305. 4:13:02because it removes those five small
  5306. 4:13:04files from the log and then it adds one
  5307. 4:13:07file. So here we should see one ad
  5308. 4:13:12for the ad operation for the clothing
  5309. 4:13:14category. Yeah. And this is the new file
  5310. 4:13:17that got added. Right. So I hope that
  5311. 4:13:18makes sense when you compare it with the
  5312. 4:13:20delta log. Now how do you remove those
  5313. 4:13:23five files which is no longer going to
  5314. 4:13:25be referenced in the latest version?
  5315. 4:13:28We'll come to that. Now let's go back
  5316. 4:13:31and let's run this
  5317. 4:13:35run an operation called vacuum. Right.
  5318. 4:13:38So what vacuum is going to do is it is
  5319. 4:13:40going to remove the files for
  5320. 4:13:43tombstoning. Right. The files that have
  5321. 4:13:45been tombstone or marked for soft delete
  5322. 4:13:49they are going to be removed. Yeah. So
  5323. 4:13:51in order to do that we simply run a
  5324. 4:13:53table vacuum and I put a number zero
  5325. 4:13:57over here. Zero basically stands for
  5326. 4:13:59retention duration. Generally the
  5327. 4:14:01retention duration is 168 hours. So
  5328. 4:14:05files that have been created within the
  5329. 4:14:06last 160 68 hours are not going to be
  5330. 4:14:10deleted. But when I put a number zero,
  5331. 4:14:13it simply says that go ahead and delete
  5332. 4:14:16all of the files that have been
  5333. 4:14:18tombstone or have been marked for soft
  5334. 4:14:21delete. So let's go ahead and run this.
  5335. 4:14:25Okay. So it gives me a warning saying
  5336. 4:14:28that are you sure that you would like to
  5337. 4:14:30vacuum file with such low retention
  5338. 4:14:33duration and it's really good that such
  5339. 4:14:36a warning is in place right because
  5340. 4:14:38somebody may be just running a vacuum by
  5341. 4:14:40mistake. So what we need to do is if
  5342. 4:14:44you're certain that there are no
  5343. 4:14:45operation being performed on this table
  5344. 4:14:47such as blah blah blah then you may turn
  5345. 4:14:49off this check by setting this right and
  5346. 4:14:52as I was saying earlier if you're not
  5347. 4:14:54sure please use the value not less than
  5348. 4:14:57168 hours right 168 hours happens to be
  5349. 4:15:00the default so let me go ahead and run
  5350. 4:15:03this
  5351. 4:15:07and this is going to be a set
  5352. 4:15:14Let's run this now
  5353. 4:15:16and let's run the vacuum now.
  5354. 4:15:21Yeah. Okay. So, the vacuum is complete
  5355. 4:15:24now. Now if I if I were to go back to
  5356. 4:15:28this table and if I were to open any of
  5357. 4:15:31these categories ideally I should see
  5358. 4:15:34just one file because all of the five
  5359. 4:15:37files were compacted into one file.
  5360. 4:15:40Right? And that is what you see over
  5361. 4:15:42here. Right?
  5362. 4:15:44Let me also check the other ones.
  5363. 4:15:48Okay. Perfect. So this works as
  5364. 4:15:50expected. Let me also rerun
  5365. 4:15:54rerun this query on this table after
  5366. 4:15:57we've run an optimize and a vacuum.
  5367. 4:16:00Right? So let's go ahead and run this
  5368. 4:16:02again.
  5369. 4:16:05Okay. So the wall time is now 1.4
  5370. 4:16:09seconds comparing to 2.3 seconds
  5371. 4:16:11earlier. So there is an improvement in
  5372. 4:16:14the runtime. Right. Of course the number
  5373. 4:16:15is not very huge because our file sizes
  5374. 4:16:19are very small. we will be able to see
  5375. 4:16:21significant improvements on a larger
  5376. 4:16:23data set. Right? But overall the idea is
  5377. 4:16:28to help us understand that small file
  5378. 4:16:31problem is a significant problem. Right?
  5379. 4:16:34And optimize helps us eliminate that
  5380. 4:16:37problem. Okay. So now I want to share
  5381. 4:16:40with you how to use predicates or
  5382. 4:16:43filters or where condition with
  5383. 4:16:45optimize. So this is going to be very
  5384. 4:16:48helpful in cases where you have
  5385. 4:16:50incremental data coming in. Let's say
  5386. 4:16:52you have daily data coming in, right?
  5387. 4:16:54You have one partition for each day. You
  5388. 4:16:57would naturally want to run an optimize
  5389. 4:17:00on that particular day on that
  5390. 4:17:01particular partition. You don't want to
  5391. 4:17:03run optimize on the full data. Yeah. So
  5392. 4:17:06in those cases, we would want to use a
  5393. 4:17:08predicate or a filter with the optimize.
  5394. 4:17:13Yeah. So let's try to mimic that
  5395. 4:17:15situation. And I'm going to use this
  5396. 4:17:17same data frame df over here. So let's
  5397. 4:17:21assume that I'm creating a new category
  5398. 4:17:25called fruits
  5399. 4:17:27and this is simply going to be so I'll
  5400. 4:17:28just take up one category and rename it
  5401. 4:17:30right uh just for the sake of
  5402. 4:17:32simplicity. So
  5403. 4:17:35df do.category category equals fruit.
  5404. 4:17:38This is actually this I'll take up
  5405. 4:17:41clothing all of the data that exist in
  5406. 4:17:43clothing and I'm simply going to rename
  5407. 4:17:46it
  5408. 4:17:50as fruit.
  5409. 4:17:55Okay. And this is simply going to be
  5410. 4:17:57from pispar.sql
  5411. 4:18:00functions import pispark.sqlfunctions
  5412. 4:18:03lf.
  5413. 4:18:05Yeah. So let's go ahead and run this. So
  5414. 4:18:08basically what I've done is I've taken
  5415. 4:18:09all of the data that exists for the
  5416. 4:18:10clothing category and simply renamed
  5417. 4:18:12clothing to fruits. And this is just to
  5418. 4:18:15generate some sample data. So what this
  5419. 4:18:18means is that a new category or a new
  5420. 4:18:21partition is coming in which is by the
  5421. 4:18:24name fruits. And now I simply want to
  5422. 4:18:28optimize that partition. Yeah. So let's
  5423. 4:18:32go ahead and first
  5424. 4:18:35check that this only contains
  5425. 4:18:40this only contains one category.
  5426. 4:18:48Okay, so it simply contains one
  5427. 4:18:50category. Right now let me go ahead and
  5428. 4:18:53write this data. So I'm again going to
  5429. 4:18:56do a repartition five. Okay, great.
  5430. 4:19:03So, this should write a new partition to
  5431. 4:19:07this table. Yeah. So, let's go ahead and
  5432. 4:19:10run this.
  5433. 4:19:13And now,
  5434. 4:19:17I should see a new partition called
  5435. 4:19:19fruits over here. Once this completes,
  5436. 4:19:25okay, so the writing has completed. Let
  5437. 4:19:27me refresh this.
  5438. 4:19:29And I see a new category called fruits.
  5439. 4:19:32Now let's go here. And I see five files
  5440. 4:19:35over here. Yeah, as expected because I
  5441. 4:19:38did a repartition because I also wanted
  5442. 4:19:40to mimic the small file problem. So now
  5443. 4:19:44what I'm simply going to do is
  5444. 4:19:47I'm going to run a table dot optimize
  5445. 4:19:53dot where the category
  5446. 4:19:58the category equals fruits and then
  5447. 4:20:00execute compaction right
  5448. 4:20:03you can also use SQL in order to do this
  5449. 4:20:06and for this time let's use SQL
  5450. 4:20:10so the command is going to be something
  5451. 4:20:13like this. Optimize
  5452. 4:20:17delta catalog dot delta DB dot
  5453. 4:20:21optimize example one where category
  5454. 4:20:27equals fruit
  5455. 4:20:29right let's go ahead and run this now
  5456. 4:20:33okay so this is complete now I should
  5457. 4:20:35see one extra file over here
  5458. 4:20:39and you see that there is this one extra
  5459. 4:20:42file that is created it. Yeah. And now I
  5460. 4:20:45simply want to remove
  5461. 4:20:48all of the tombstone file. Yeah. So
  5462. 4:20:51again let's use SQL for now. Vacuum
  5463. 4:20:55this table delta catalog delta DB dot
  5464. 4:20:59optimize example one retain zero hours.
  5465. 4:21:05So you get a gist of
  5466. 4:21:08using both SQL and the delta table API.
  5467. 4:21:11Yeah. So let's run this.
  5468. 4:21:22Great. So now you see that we have only
  5469. 4:21:25one file as expected. Yeah. So that is
  5470. 4:21:28how you can use predicates when using an
  5471. 4:21:32optimize command. So given that now
  5472. 4:21:34we've seen the small file problem in
  5473. 4:21:37action and how to fix them using the
  5474. 4:21:40optimize command let's also understand
  5475. 4:21:43the root cause right because many a
  5476. 4:21:46times if you understand the root cause
  5477. 4:21:48you may fix the problem right at the
  5478. 4:21:51source or the origin right so let's have
  5479. 4:21:54a look at the root cause one by one the
  5480. 4:21:57first one is
  5481. 4:22:00first of all let me quickly write root
  5482. 4:22:02root causes and the first one is
  5483. 4:22:06repartitioning to a very large number
  5484. 4:22:10right the first one is a repartition so
  5485. 4:22:14let's say that you have a file which is
  5486. 4:22:1710 GB right and you end up
  5487. 4:22:20repartitioning it into 10,000 parts
  5488. 4:22:24right so what you do is you do a dot
  5489. 4:22:26repartition
  5490. 4:22:28you do a dot repartition of 10,000
  5491. 4:22:33on whatever data frame that you're
  5492. 4:22:34running. Right? So if your data set is
  5493. 4:22:3810 GB in size, you are essentially going
  5494. 4:22:42to end up with 10 into,000
  5495. 4:22:45mgabytes. I'm assuming 1 GB to be,000
  5496. 4:22:48megabytes for simplicity into 10,000
  5497. 4:22:52which is simply 1 mgabyte in size.
  5498. 4:22:56Right? So each of these partition is
  5499. 4:22:58going to be one megabyte in size and
  5500. 4:23:01that is going to cause the small size
  5501. 4:23:03small file problem. Yeah. The second one
  5502. 4:23:06the second issue is partitioning
  5503. 4:23:10partitioning
  5504. 4:23:13on a high cardality column. Right? You
  5505. 4:23:17partition on a high
  5506. 4:23:20cardality column.
  5507. 4:23:24And by high cardality I simply mean that
  5508. 4:23:26this column has a lot of distinct
  5509. 4:23:29values. So let's imagine that you have a
  5510. 4:23:33500 MB retail data set, right? You have
  5511. 4:23:36a 500 mgabyte retail data set and you
  5512. 4:23:41have a category column.
  5513. 4:23:45So for some reason you want to partition
  5514. 4:23:48by the category column. Let's say you
  5515. 4:23:50are writing this data set and when you
  5516. 4:23:52do a df dot write dot partition by
  5517. 4:23:58dot partition by
  5518. 4:24:00you basically put in the category column
  5519. 4:24:04right and this uh category column has a
  5520. 4:24:07lot of distinct values so let's say it
  5521. 4:24:09has thousands of distinct values right
  5522. 4:24:13so what you're going to end up with is a
  5523. 4:24:16lot of small files and this again will
  5524. 4:24:20again lead to the small file problem.
  5525. 4:24:23The third one is frequently updated data
  5526. 4:24:27sets, right? Frequently
  5527. 4:24:30frequently updated data set,
  5528. 4:24:34right? So let's say you have data coming
  5529. 4:24:38in um continuously and by continuously
  5530. 4:24:41let's assume that it comes in every 5
  5531. 4:24:43minutes, right? So these data sets that
  5532. 4:24:48come in every 5 minutes, they basically
  5533. 4:24:51end up writing small small updates,
  5534. 4:24:53right? And these small updates are going
  5535. 4:24:57to be in the form of small files,
  5536. 4:25:01right? Which is again going to lead to a
  5537. 4:25:04small file problem, right? So as we can
  5538. 4:25:06see that
  5539. 4:25:08these three could be some of the most
  5540. 4:25:12important root causes for the small file
  5541. 4:25:16problem. And number one and number two
  5542. 4:25:18these are purely technical in nature.
  5543. 4:25:20Right? You can change the repartition
  5544. 4:25:22number or you can change the
  5545. 4:25:24partitioning column in order to avoid
  5546. 4:25:26it. Right? Avoid the small file problem.
  5547. 4:25:28But problems like this where let's say
  5548. 4:25:31the business needs to see data
  5549. 4:25:33instantly. Right? or they want the data
  5550. 4:25:36to be updated frequently. In those
  5551. 4:25:38cases, these problems are not technical
  5552. 4:25:41in nature. They are more of a business
  5553. 4:25:43problem. Right? So, how do you solve
  5554. 4:25:45such kind of problems? So what you can
  5555. 4:25:47do essentially is given that the team is
  5556. 4:25:50going to use this underlying data set
  5557. 4:25:53you can choose to run optimize after a
  5558. 4:25:55number of hours or maybe every day so
  5559. 4:25:58that at the end of the day you end up
  5560. 4:26:00combining those small files into a
  5561. 4:26:03larger file. Right
  5562. 4:26:06now there are three methods or
  5563. 4:26:08approaches that you can think about
  5564. 4:26:10whenever you want to run the optimize
  5565. 4:26:13command. Yeah. And the first approach or
  5566. 4:26:16method is manual optimize. And we've
  5567. 4:26:19seen it in the code example where we
  5568. 4:26:23simply run the optimize command. We
  5569. 4:26:26specify the table and if needed we
  5570. 4:26:28specify a predicate. Right? So either
  5571. 4:26:30you can run it like this or you can run
  5572. 4:26:33it like this. Right? So this is the most
  5573. 4:26:35simplest approach that is widely
  5574. 4:26:38followed. Now the next approach is an
  5575. 4:26:40automatic one and is called optimize
  5576. 4:26:45right. Yeah, it's called optimize
  5577. 4:26:50right.
  5578. 4:26:52What optimize write does is that it
  5579. 4:26:55simply combines all the small rights to
  5580. 4:26:59a partition into a single write command.
  5581. 4:27:03So we are going to understand what that
  5582. 4:27:04means. But let's say we are doing the
  5583. 4:27:07traditional right. What simply happens
  5584. 4:27:10is that for creating one partition and
  5585. 4:27:13let's imagine that this is the date
  5586. 4:27:14partition something like date= 2025
  5587. 4:27:1805 01. Yeah. So in order to create this
  5588. 4:27:23partition there are several processes
  5589. 4:27:26which are writing to this partition. So
  5590. 4:27:28we see that executor number one is
  5591. 4:27:31writing to this partition. Executor
  5592. 4:27:33number two is also writing to this
  5593. 4:27:35partition. Executor number three is also
  5594. 4:27:39writing to this partition. And these may
  5595. 4:27:42end up writing small paret files. Let's
  5596. 4:27:44say this is 1.pk, this is 2.pk and
  5597. 4:27:48similarly this is 3.pk. So these may end
  5598. 4:27:51up writing small files thereby creating
  5599. 4:27:54the small file problem. Right? So the
  5600. 4:27:56root cause of it is several processes
  5601. 4:27:59writing to a partition. Yeah. Now
  5602. 4:28:02instead if you look at optimized right
  5603. 4:28:05what going to happen is that all of the
  5604. 4:28:08data is going to be shuffled and then is
  5605. 4:28:11going to be executed as a single write
  5606. 4:28:14command. So all of the data is shuffled
  5607. 4:28:16over here from executor one executor 2
  5608. 4:28:21and executor 3. All of the data is
  5609. 4:28:24shuffled and then this is executed as a
  5610. 4:28:28single write command. Yeah. So a very
  5611. 4:28:32important point to note here is that
  5612. 4:28:34it's executed as a single write command
  5613. 4:28:37and this is going to produce
  5614. 4:28:39appropriately sized files right so this
  5615. 4:28:42is going to produce appropriately file
  5616. 4:28:44size files in partition 1 2 and three
  5617. 4:28:48yeah so that's the benefit of using
  5618. 4:28:51optimize right however a tradeoff is
  5619. 4:28:54that it's going to incur a shuffle
  5620. 4:28:58is going to incur a shuffle and we know
  5621. 4:29:01that shuffles are costly because it
  5622. 4:29:03requires data transfer right so it's
  5623. 4:29:06important to keep this in mind whenever
  5624. 4:29:08we want to use an optimized right what
  5625. 4:29:11are the thing that you want to optimize
  5626. 4:29:13for yeah is it right latency if it is
  5627. 4:29:17right latency then this might not be a
  5628. 4:29:19good solution because
  5629. 4:29:23there is going to be some shuffle
  5630. 4:29:24involved and that is going to take time
  5631. 4:29:27if it's optimizing for the small file
  5632. 4:29:30problem then of course this is a very
  5633. 4:29:33good solution. So now let's see optimize
  5634. 4:29:36right in action and for that I'm going
  5635. 4:29:38to use the same data frame over here. So
  5636. 4:29:42let me quickly
  5637. 4:29:45create a new table df dot
  5638. 4:29:49repartition
  5639. 4:29:51and for this time let me create lots of
  5640. 4:29:54partition. So let me repartition by 288
  5641. 4:29:58and then write dot mode override
  5642. 4:30:02dot partition by let's partition by
  5643. 4:30:05category
  5644. 4:30:07and then let's save as table
  5645. 4:30:14delta
  5646. 4:30:16catalog dot deltadb dot optimize example
  5647. 4:30:212. Yeah. So let's run this and on this
  5648. 4:30:25table I also want to run this query
  5649. 4:30:29the same query that I've run earlier and
  5650. 4:30:31see how it performs.
  5651. 4:30:34So this is going to be example two.
  5652. 4:30:43Yeah.
  5653. 4:30:45Okay. So now that is complete.
  5654. 4:30:48I should see a table called optimize to
  5655. 4:30:51over here. And if I go to the details,
  5656. 4:30:58let me
  5657. 4:31:00have a look at the table over here. And
  5658. 4:31:05okay, so I basically see all of these
  5659. 4:31:08partitions over here. Yeah. So ideally
  5660. 4:31:10there should be some 288 partitions. So
  5661. 4:31:13now
  5662. 4:31:15let's go ahead and run this. And let me
  5663. 4:31:18create another table. This time,
  5664. 4:31:22this time after partition by I am going
  5665. 4:31:27to write an option which is going to be
  5666. 4:31:30optimize right to be true
  5667. 4:31:35and then we are going to save this table
  5668. 4:31:38as example three.
  5669. 4:31:41Yeah. So let's go ahead and run this. So
  5670. 4:31:44this operation has taken 6.97
  5671. 4:31:48seconds. Yeah.
  5672. 4:31:50Okay, so this is completed. Again, I
  5673. 4:31:53should see a new table here which is
  5674. 4:31:55optimize example three. And if I have a
  5675. 4:31:58look at the table this time,
  5676. 4:32:04I should see minimal partitions, right?
  5677. 4:32:07Okay. So now when I look at these
  5678. 4:32:09partitions,
  5679. 4:32:13we see just one file.
  5680. 4:32:17So that means the optimize write has
  5681. 4:32:19worked very well and it is not writing
  5682. 4:32:22small files. It has just written one
  5683. 4:32:24file which is appropriately signed.
  5684. 4:32:26Let's also run this operation once again
  5685. 4:32:30and this time on example three.
  5686. 4:32:43Perfect. This runs a lot quicker. If you
  5687. 4:32:45see 6.97 seconds versus 1.82 seconds. So
  5688. 4:32:48now we've seen an understood optimize
  5689. 4:32:51rights and it's actually good for cases
  5690. 4:32:54where many processes or executors are
  5691. 4:32:57trying to write many files to a
  5692. 4:32:59partition and instead of all of that we
  5693. 4:33:02shuffle all of those files. We combine
  5694. 4:33:04all of it and write it into appropriate
  5695. 4:33:08file sizes for every partition. Right?
  5696. 4:33:10So every partition is going to have an
  5697. 4:33:13appropriately sized file and there is
  5698. 4:33:16going to be an appropriate number of
  5699. 4:33:18those files. Right? So this helps avoid
  5700. 4:33:21the small file problem. But in cases
  5701. 4:33:23where we are frequently writing small
  5702. 4:33:26small updates to a table, right? The
  5703. 4:33:29updates to a table are coming in small
  5704. 4:33:31small chunks, the files that we get out
  5705. 4:33:34of those small updates are still going
  5706. 4:33:37to be small files, right? So we still
  5707. 4:33:40end up with the small file problem and
  5708. 4:33:43here is where autoco compaction comes
  5709. 4:33:45into picture and that is the third
  5710. 4:33:47approach that I wanted to talk about. So
  5711. 4:33:50what autoco compaction does is that
  5712. 4:33:53so what autoco compaction does is and
  5713. 4:33:56let me let me quickly write something
  5714. 4:33:58here is that let's say when you get
  5715. 4:34:01small files right you get a few files
  5716. 4:34:04after every write
  5717. 4:34:07it is going to run a compaction
  5718. 4:34:10operation
  5719. 4:34:11right it is going to run an optimize
  5720. 4:34:14command
  5721. 4:34:16which is going to compact all of these
  5722. 4:34:18files
  5723. 4:34:19into appropriatelyized files
  5724. 4:34:23and that is what autoco compaction
  5725. 4:34:25exactly does. Right now it's really
  5726. 4:34:27important to understand is that auto
  5727. 4:34:29compaction is not going to run for an
  5728. 4:34:32arbitrary number of file. Let's say you
  5729. 4:34:34got two files and now you want autoco
  5730. 4:34:36compaction to run and combine it into
  5731. 4:34:38one file. Basically there is a setting
  5732. 4:34:40which is called spark
  5733. 4:34:42databasel.compact.min
  5734. 4:34:45files minum file. Just Google it up. You
  5735. 4:34:49can specify the minimum number of files
  5736. 4:34:52that need to be present in order to
  5737. 4:34:54trigger autoco compaction. Once that is
  5738. 4:34:56there, it is going to compact all of
  5739. 4:34:59those files into appropriately sized
  5740. 4:35:02files. Right? So let's see autoco
  5741. 4:35:05compaction in action right now. And we
  5742. 4:35:08are going to use the same table example
  5743. 4:35:11three as earlier. And before I can use
  5744. 4:35:14that I need to do two things. The first
  5745. 4:35:17one is to enable
  5746. 4:35:20to enable
  5747. 4:35:22autoco compaction
  5748. 4:35:24and the second one is to disable
  5749. 4:35:26optimize rate because if I don't disable
  5750. 4:35:28optimize right it is going to take all
  5751. 4:35:30the small file shuffle it into one and
  5752. 4:35:33then write it as one file so I won't be
  5753. 4:35:35able to see whether auto compaction is
  5754. 4:35:37working or not. Yeah. So first of all I
  5755. 4:35:45I disable
  5756. 4:35:48disable this
  5757. 4:35:53and then
  5758. 4:35:56let's enable this
  5759. 4:36:00actually let me change this auto comp uh
  5760. 4:36:02autooptimize dot auto compact
  5761. 4:36:06this is going to be true. Yeah, perfect.
  5762. 4:36:14Okay, this is done. Now I want to check
  5763. 4:36:17another property which basically tells
  5764. 4:36:20me
  5765. 4:36:21uh spark databreak dot delta
  5766. 4:36:25dot autocompact
  5767. 4:36:29dot min num files. This basically tells
  5768. 4:36:33me what is the minimum number of files
  5769. 4:36:37that I need in order to trigger
  5770. 4:36:39autocompaction. So I actually recently
  5771. 4:36:41changed it to three. The default number
  5772. 4:36:45is 50. Yeah. So I changed it to three.
  5773. 4:36:48You can change it to three using the
  5774. 4:36:50following command
  5775. 4:36:56using a set. Yeah. So that is how you
  5776. 4:36:59can do that. Now I want to insert new
  5777. 4:37:02partition, a new partition into this
  5778. 4:37:04table. Yeah. And for that I'm going to
  5779. 4:37:07use the fruits code again. So let's
  5780. 4:37:10assume that there's new data coming in.
  5781. 4:37:13And let me simply rename this to
  5782. 4:37:15detergents.
  5783. 4:37:17And let's also rename this to
  5784. 4:37:19detergents.
  5785. 4:37:21And again this is just for creating
  5786. 4:37:23sample data. I am simply taking up all
  5787. 4:37:26of the rows with respect to the clothing
  5788. 4:37:28category. and then simply renaming
  5789. 4:37:30renaming the clothing um clothing values
  5790. 4:37:34to detergents. Right? So let's go ahead
  5791. 4:37:37and run this. And now I want to run a DF
  5792. 4:37:42detergent dot repartition five
  5793. 4:37:47dot write dot mode overrite and this is
  5794. 4:37:51going to be written as example three.
  5795. 4:37:53Yeah. So let's go ahead and run this
  5796. 4:37:56now.
  5797. 4:37:59Okay, so this is complete. Let's check.
  5798. 4:38:05Let's check what happened.
  5799. 4:38:10Okay, so I see a detergent category.
  5800. 4:38:15And I see 1 2 3 4 5 6 files, which is
  5801. 4:38:20exactly what we wanted. The reason for
  5802. 4:38:22that is it would have written five files
  5803. 4:38:25initially
  5804. 4:38:27because we had a repartition five and
  5805. 4:38:30then it wrote an additional file because
  5806. 4:38:34of the compaction operation. Yeah. So
  5807. 4:38:36what you see over here is essentially
  5808. 4:38:39there is one file which is bigger in
  5809. 4:38:42size which is this one and this was the
  5810. 4:38:44one that was written as a result of the
  5811. 4:38:48auto compaction operation. So that is
  5812. 4:38:50how autoco compaction works and all of
  5813. 4:38:53these other files right that you see
  5814. 4:38:54over here they have been tombstone when
  5815. 4:38:57you run a vacuum all of them could be
  5816. 4:38:59removed right so I hope that gave you a
  5817. 4:39:01gist of how autocompaction works so now
  5818. 4:39:04let's move over to the next topic which
  5819. 4:39:06is called vacuum
  5820. 4:39:09and we've already seen vacuum in action
  5821. 4:39:13so whenever we delete update or merge
  5822. 4:39:16records in a delta table, some of the
  5823. 4:39:20underlying records are going to be
  5824. 4:39:22removed, right? So when you update, it
  5825. 4:39:25creates a new record and the old records
  5826. 4:39:28becomes irrelevant. When you delete, all
  5827. 4:39:31of those records become irrelevant,
  5828. 4:39:33right? So because these records have
  5829. 4:39:36been irrelevant, have become irrelevant,
  5830. 4:39:38they're supposed to be deleted, right?
  5831. 4:39:40But what actually happens in delta is
  5832. 4:39:43that it doesn't physically delete it
  5833. 4:39:46from the cloud storage or from the disk.
  5834. 4:39:48It just marks it for deletion. Right? It
  5835. 4:39:52just tombstones it or it does something
  5836. 4:39:54which is called soft delete. Right? So
  5837. 4:39:58all of these rows which are marked for
  5838. 4:40:00deletion which is soft delete or the
  5839. 4:40:03rows which have been tombstone they can
  5840. 4:40:06only be removed once you run your
  5841. 4:40:08vacuum. Right. So what essentially
  5842. 4:40:11happens is let's say when you run a
  5843. 4:40:13delete
  5844. 4:40:16merge
  5845. 4:40:19or let's say an update
  5846. 4:40:23there are rows which are going to become
  5847. 4:40:24irrelevant right and those rows are
  5848. 4:40:28simply tombstoned
  5849. 4:40:32or they are marked for soft deletion.
  5850. 4:40:39Right? So the moment
  5851. 4:40:41you apply a vacuum or let's say you want
  5852. 4:40:45to remove these roles physically from
  5853. 4:40:49the disk
  5854. 4:40:52physically remove all of these roles
  5855. 4:40:54either from the cloud storage or from
  5856. 4:40:56the disk. Right? So in order to do that
  5857. 4:41:00you need to run the vacuum command and
  5858. 4:41:03vacuum is going to remove all of the
  5859. 4:41:05rows which have been tombstone or have
  5860. 4:41:08been marked for stop delete. Right?
  5861. 4:41:10Essentially the same thing right. So
  5862. 4:41:12that is how vacuum works and there are
  5863. 4:41:15two important points to note. The first
  5864. 4:41:17one is that vacuum helps you save on
  5865. 4:41:20storage cost. It doesn't make your
  5866. 4:41:23queries faster. It's a very popular
  5867. 4:41:25misconception that running vacuum is
  5868. 4:41:28going to make your queries run faster.
  5869. 4:41:31That's not the case. The reason for that
  5870. 4:41:33is we've seen in earlier examples that
  5871. 4:41:38we had five files, right? We have
  5872. 4:41:41written five files and then we run an
  5873. 4:41:44optimize.
  5874. 4:41:46We run an optimize and then it compacts
  5875. 4:41:49all of the data into one file. Right? So
  5876. 4:41:53now you end up having six files right
  5877. 4:41:56five files which were already present
  5878. 4:41:58earlier and additionally you have one
  5879. 4:42:01file which were the result of the
  5880. 4:42:04compaction operation right so now the
  5881. 4:42:07five files have been marked for deletion
  5882. 4:42:10right so this is the five file and this
  5883. 4:42:12is one file this has been tombstoned
  5884. 4:42:18and this one is active
  5885. 4:42:21right so The moment you run a vacuum,
  5886. 4:42:26it is going to delete all of these
  5887. 4:42:29files,
  5888. 4:42:30it is going to permanently physically
  5889. 4:42:33delete all of these files and the file
  5890. 4:42:36that you are left with is this one file.
  5891. 4:42:39So the data you end up scanning is just
  5892. 4:42:42the same, right? The filtering and all
  5893. 4:42:45of that happens at a delta log level,
  5894. 4:42:47right? Which files do I need to scan?
  5895. 4:42:49That happens at a delta log level. And
  5896. 4:42:52that is taken care of over there. Right?
  5897. 4:42:54So essentially when we run vacuum we
  5898. 4:42:57just save on storage cost. We are simply
  5899. 4:43:00removing data and that is how it doesn't
  5900. 4:43:05make your queries run faster. It simply
  5901. 4:43:07removes data that is not needed. And the
  5902. 4:43:10second point is that it limits your
  5903. 4:43:13ability to time travel.
  5904. 4:43:16Limits your ability to time travel. So
  5905. 4:43:20if you run vacuum you will not be able
  5906. 4:43:24to go back to any of the previous
  5907. 4:43:26version. So let's see that with an
  5908. 4:43:28example. So I'm simply going to take one
  5909. 4:43:30of these files over here and let's
  5910. 4:43:34create two data frames right equals
  5911. 4:43:36spark dot
  5912. 4:43:39park read.park park
  5913. 4:43:42and this is going to be a filter f dot
  5914. 4:43:46column customer id dot between this is
  5915. 4:43:51going to be 101 and 150 right so this
  5916. 4:43:55park file contains all customers from
  5917. 4:43:57101 to 200 so I am simply selecting 101
  5918. 4:44:01to 150 and this is going to be 151 until
  5919. 4:44:05200 so this will be 151 comma 200 100
  5920. 4:44:11and let's go ahead and run this.
  5921. 4:44:14Let's now write this to a table.
  5922. 4:44:18Yeah. So, this is going to be write do
  5923. 4:44:23mode override.
  5924. 4:44:26I don't need any partitioning.
  5925. 4:44:29And let's just say
  5926. 4:44:32vacuum example one.
  5927. 4:44:35And let's just write it. Yep.
  5928. 4:44:38Let's also see the history.
  5929. 4:44:41Describe history
  5930. 4:44:44delta catalog dot deltatb dot vacuum
  5931. 4:44:50example one. So I should just see one
  5932. 4:44:52row. Okay. Uh what is the issue here?
  5933. 4:44:57Okay, that's a spelling mistake. Uh
  5934. 4:45:00vacuum
  5935. 4:45:01example one.
  5936. 4:45:03And this is as expected. So now let's
  5937. 4:45:06write the other data frame as well.
  5938. 4:45:08Yeah.
  5939. 4:45:09And this is simply going to be TF 151.
  5940. 4:45:12And I'm going to append it.
  5941. 4:45:17Let's go ahead and write it again.
  5942. 4:45:21Okay, perfect. So now we see two rows.
  5943. 4:45:24The first one is the create. The second
  5944. 4:45:26one is the write operation where we
  5945. 4:45:29wrote in additional 50 rows of data.
  5946. 4:45:32Now let's go ahead and do a delete.
  5947. 4:45:35Yeah. So actually before I do a delete,
  5948. 4:45:38let me show you something. So this is
  5949. 4:45:40the vacuum table.
  5950. 4:45:44Let me copy the UU ID and
  5951. 4:45:51okay. So I see two rows over here. Yeah.
  5952. 4:45:53The first one which is 22 2123 the time
  5953. 4:45:57at which it was written. This will be
  5954. 4:46:00this file over here which contains all
  5955. 4:46:03of the customers from 101 to 150. And
  5956. 4:46:07the second one should be 151 to 200.
  5957. 4:46:10Yeah. So the second one over here should
  5958. 4:46:13be 151 to 200. Yeah. The one that starts
  5959. 4:46:16with C82.
  5960. 4:46:19Okay. Now let's go ahead and delete some
  5961. 4:46:22rows. So we do a delete from delta
  5962. 4:46:25catalog.
  5963. 4:46:28Delta catalog that this where customer
  5964. 4:46:31ID between
  5965. 4:46:34151 and 200. So customer ID is from 151
  5966. 4:46:39to 200 they basically reside in this
  5967. 4:46:43park file right over here. Yeah. So when
  5968. 4:46:45I run this statement ideally it should
  5969. 4:46:48tombstone the second file. It should
  5970. 4:46:50mark it for delete. Yeah. So now let's
  5971. 4:46:54run this again.
  5972. 4:46:56Let's see the history again.
  5973. 4:47:00Perfect. Now you see three uh three
  5974. 4:47:03statements or three versions. Yeah. The
  5975. 4:47:05third one is the delete.
  5976. 4:47:08After this, let's run a few quick
  5977. 4:47:14statistics which is minimum of customer
  5978. 4:47:20ID,
  5979. 4:47:23comma, maximum of customer ID
  5980. 4:47:27from this table
  5981. 4:47:34vacuum example one. Yeah.
  5982. 4:47:38What should we see? So we have
  5983. 4:47:42all the customers from 101 to 200 and we
  5984. 4:47:46have removed 151 to 200. Right? So we
  5985. 4:47:50have removed 151 to 200. So we should
  5986. 4:47:54see the minimum to be 101 and the
  5987. 4:47:56maximum to be 150.
  5988. 4:47:59And that is what you see over here. Now
  5989. 4:48:01if I were to run the same statement
  5990. 4:48:05on version number
  5991. 4:48:08version number
  5992. 4:48:10one, we ran a delete on version number
  5993. 4:48:14two. If I were to run the same statement
  5994. 4:48:16on version number one, it would have all
  5995. 4:48:19the role from 101 to 200. Right? So the
  5996. 4:48:22result would be a little different. The
  5997. 4:48:23minimum would be 101, but the maximum
  5998. 4:48:26would be 200. Yeah.
  5999. 4:48:30Perfect.
  6000. 4:48:32So now what we want to do is we want to
  6001. 4:48:35run a vacuum. Yeah, we want to run a
  6002. 4:48:38vacuum in order to remove the tombstone
  6003. 4:48:41file over here. We want to remove this
  6004. 4:48:44file over here which has been tombstone.
  6005. 4:48:47So, let's go ahead and run a vacuum
  6006. 4:48:55delta catalog dot
  6007. 4:48:58delta db dot vacuum retain zero hours.
  6008. 4:49:02Yeah. So, let's go ahead and run this.
  6009. 4:49:04And as expected, I got a message saying
  6010. 4:49:07that are you sure you want to delete
  6011. 4:49:09this? And we've seen this earlier as
  6012. 4:49:10well. So let's go ahead and
  6013. 4:49:14disable this check over here. So we
  6014. 4:49:17simply do this equals false.
  6015. 4:49:20And let's run this again.
  6016. 4:49:24Okay. So now the vacuum operation is
  6017. 4:49:27complete and I should just see one file
  6018. 4:49:29over here. The one that starts with 8
  6019. 4:49:32EF, right? The second one should be
  6020. 4:49:34gone. Perfect. So that is what was
  6021. 4:49:38expected. Right? So now that the vacuum
  6022. 4:49:41operation is complete, if I were to run
  6023. 4:49:44a history, if I were to see the history
  6024. 4:49:47of this table,
  6025. 4:49:49describe history,
  6026. 4:49:52I am basically going to see two
  6027. 4:49:54additional rows.
  6028. 4:49:56The first one is for vacuum start and
  6029. 4:49:59vacuum end. Right? So now
  6030. 4:50:02we see over here that we only have one
  6031. 4:50:06park file and this corresponds to the
  6032. 4:50:08data from customer id 101 to 150. Right?
  6033. 4:50:14So the experiment what I want to do is
  6034. 4:50:18select star from this table
  6035. 4:50:23version
  6036. 4:50:25as of yeah and I want to
  6037. 4:50:29play around with these numbers a little
  6038. 4:50:30bit. So let's say if you do version as
  6039. 4:50:33of two what should you get? Should you
  6040. 4:50:36get any data or not? That's the first
  6041. 4:50:38question. Okay. So ideally you should
  6042. 4:50:41get data because
  6043. 4:50:44after this delete all of the data from
  6044. 4:50:48151 to 200 was deleted right. So the
  6045. 4:50:50data from 101 to 150 was still present
  6046. 4:50:55right and we still have the file that is
  6047. 4:50:58needed for giving us the data from 101
  6048. 4:51:02to 150. Yeah. So this operation should
  6049. 4:51:05just work fine
  6050. 4:51:08and it works fine. Yeah. So what will
  6051. 4:51:11happen if you do version as of one? Will
  6052. 4:51:14this work?
  6053. 4:51:16Take a guess.
  6054. 4:51:21Okay, this doesn't work because in
  6055. 4:51:24operation in version number one, we
  6056. 4:51:26wrote a file which contained data from
  6057. 4:51:29151 to 200 and that file has gone
  6058. 4:51:31missing. So that is why the the data
  6059. 4:51:34that was contained in version number one
  6060. 4:51:37is incomplete and that is the reason why
  6061. 4:51:39we cannot go back in time in order to
  6062. 4:51:41access this. Yeah. Let's see about
  6063. 4:51:43version number zero. Again take a guess
  6064. 4:51:45what would happen over here.
  6065. 4:51:49Okay. You already have the park file for
  6066. 4:51:52version number zero because we wrote the
  6067. 4:51:54data from 101 to 150 and that can be
  6068. 4:51:57accessed because of the park file that
  6069. 4:51:59we see over here. Yeah. So that is how
  6070. 4:52:03uh time travel is going to work with
  6071. 4:52:05vacuum and you need to be a little
  6072. 4:52:07careful if we want to go back in time
  6073. 4:52:10and access data points right so we need
  6074. 4:52:12to run vacuum very carefully so the next
  6075. 4:52:15optimization technique that we are going
  6076. 4:52:17to talk about is Z order yeah it's
  6077. 4:52:21called Z order
  6078. 4:52:25but before we understand what exactly
  6079. 4:52:27this is let's actually set some ground
  6080. 4:52:30let's set some context with this
  6081. 4:52:32example. Yeah. So let's say you have a
  6082. 4:52:35bunch of files that you see over here 1
  6083. 4:52:382 3 and four.par park and these are
  6084. 4:52:41currently sitting on your disk either on
  6085. 4:52:43your disk or on some cloud storage and
  6086. 4:52:47you basically want to process them right
  6087. 4:52:51you want to process them and in order to
  6088. 4:52:53process them these have to be loaded in
  6089. 4:52:56the memory of whatever compute you're
  6090. 4:52:58using right so these have to be loaded
  6091. 4:53:01in memory now in order to load it into
  6092. 4:53:04memory there has to be a costly data
  6093. 4:53:07transfer over the wire Right. So in
  6094. 4:53:09order to do that, let's first take this
  6095. 4:53:12example where we have a query coming in
  6096. 4:53:15from the user. The user basically asks
  6097. 4:53:18that give me all of the records where
  6098. 4:53:21age is greater than equal to 5 and less
  6099. 4:53:24than equal to 10. Yeah. So for each of
  6100. 4:53:27the park files that we have over here,
  6101. 4:53:29we have stored some statistics and we
  6102. 4:53:33already know that all of these
  6103. 4:53:34statistics are stored in delta log,
  6104. 4:53:37right? the minimum, maximum, count of
  6105. 4:53:39values for all of the columns. Right?
  6106. 4:53:42So, we've put the statistics over here
  6107. 4:53:45and we will use this in order to figure
  6108. 4:53:49which of the files need to be
  6109. 4:53:51transferred over the wire. Yeah. So, we
  6110. 4:53:54see that the minimum age and the maximum
  6111. 4:53:57age is 4 and 12 which overlaps with this
  6112. 4:54:00value. Right? So, that means that this
  6113. 4:54:03file is going to be transferred. This
  6114. 4:54:06file is going to be transferred over the
  6115. 4:54:08wire. Here we see that the minimum is
  6116. 4:54:1010, maximum is 28 and that overlaps with
  6117. 4:54:13this value. So this file is also going
  6118. 4:54:16to be transferred. Again the minimum is
  6119. 4:54:186 and 14. That also overlaps with this
  6120. 4:54:21value. Yeah, with this value here to
  6121. 4:54:24here and this file is also going to be
  6122. 4:54:27transferred. And this is a wide range
  6123. 4:54:29from 4 to 60. And this again overlaps
  6124. 4:54:32with this value over here. That means
  6125. 4:54:354.park par is also going to be
  6126. 4:54:37transferred. Yeah. So essentially what
  6127. 4:54:40we see is that when such a query comes
  6128. 4:54:44in all of these four files are going to
  6129. 4:54:47be transferred. Yeah. These are going to
  6130. 4:54:50be transferred over the wire
  6131. 4:54:54and this is a costly operation.
  6132. 4:55:00This is a costly operation. Now if we
  6133. 4:55:03were to take a step back, do you think
  6134. 4:55:06this transfer can be avoided? Instead of
  6135. 4:55:09transferring all of the four files, can
  6136. 4:55:12we transfer fewer files? Is it possible?
  6137. 4:55:15Take a moment and think about it. So
  6138. 4:55:17with this data layout, it's actually
  6139. 4:55:20impossible to avoid sending any of the
  6140. 4:55:23files over the network, right? We cannot
  6141. 4:55:25omit to send any of the files over the
  6142. 4:55:28network. We have to send all four of
  6143. 4:55:30them. If we omit any of them then we
  6144. 4:55:32won't get the right answer to this query
  6145. 4:55:35that the user has asked right so what is
  6146. 4:55:38the alternative how can we avoid sending
  6147. 4:55:42any of the files over the network is
  6148. 4:55:44that even possible how do we optimize
  6149. 4:55:46for that right so in order to do that
  6150. 4:55:49that is where zorder comes in so first
  6151. 4:55:52of all let me take some of these numbers
  6152. 4:55:54the minimum and maximum right so let's
  6153. 4:55:57say we have four over here we have 12
  6154. 4:56:00over here and of course between them
  6155. 4:56:02there is going to be several keys
  6156. 4:56:04several values for age which I'm not
  6157. 4:56:06writing down but we will have all of
  6158. 4:56:09those values right uh and then we have
  6159. 4:56:1110 and 28 so 10 is going to be over here
  6160. 4:56:15and let's say 28 is going to be over
  6161. 4:56:17here and then we have 6 and 14 so 6 is
  6162. 4:56:21going to be over here 14 is going to be
  6163. 4:56:23over here and then we have four and 60
  6164. 4:56:26I've already noted down four so 60 is
  6165. 4:56:28going to be somewhere over here Right?
  6166. 4:56:30And between all of them, right? There is
  6167. 4:56:32going to be some some keys which we do
  6168. 4:56:35not know but we are aware that there is
  6169. 4:56:37going to be data in between them. Right?
  6170. 4:56:40So this is a sorted
  6171. 4:56:43order
  6172. 4:56:44that I've produced. Right? So here is
  6173. 4:56:47something that we can apply in order to
  6174. 4:56:49skip files. And first of all let me copy
  6175. 4:56:53this.
  6176. 4:56:56Let me copy this here.
  6177. 4:57:02Yeah. So what essentially I'm going to
  6178. 4:57:05do is that
  6179. 4:57:08I am going to take the first few values
  6180. 4:57:12which is from 4 to 10 and I'm going to
  6181. 4:57:15put it into one file which is 1.pk perk
  6182. 4:57:21and then I am going to take other few
  6183. 4:57:24values which is from 10 to 28
  6184. 4:57:28and then I'm going to put it into 2.par.
  6185. 4:57:34The intuition behind this is to avoid
  6186. 4:57:38overlaps between the file. Right? If
  6187. 4:57:40there are minimal overlaps probably I'm
  6188. 4:57:43going to select only one file where all
  6189. 4:57:45of my data is going to the side. Yeah,
  6190. 4:57:48the next file I'm going to choose is so
  6191. 4:57:50probably we'll have this number 29 over
  6192. 4:57:52here and then 50 somewhere down the line
  6193. 4:57:55over here,
  6194. 4:57:57right? So 29 to 50 I am going to place
  6195. 4:58:00this in three.park
  6196. 4:58:04and then we are going to have this
  6197. 4:58:05number 51.
  6198. 4:58:07Let me place it over here. I'm going to
  6199. 4:58:10have this number 51 and until 60 I'm
  6200. 4:58:13going to take all of this data and then
  6201. 4:58:15put it in 4.pk. park.
  6202. 4:58:18The goal and idea behind this is to
  6203. 4:58:21avoid overlaps between the file. The
  6204. 4:58:24lesser overlaps are the lesser files you
  6205. 4:58:28need to scan. Right? You will be able to
  6206. 4:58:30prune the file. You would be able to say
  6207. 4:58:32that okay, this is the file I need and
  6208. 4:58:34let's send this over the network. Right?
  6209. 4:58:37So, Gorder is basically a sort and
  6210. 4:58:40repartition. So, here all of the data
  6211. 4:58:42has been sorted by whatever key we
  6212. 4:58:45wished.
  6213. 4:58:47Right? In this case, the query that has
  6214. 4:58:49come in and we basically repartition the
  6215. 4:58:52data into those number of files. Right?
  6216. 4:58:56So that is the logic and the idea behind
  6217. 4:58:58reorder. And what it does is that in
  6218. 4:59:00these files it simply colllocates data.
  6219. 4:59:05By colllocate data what I mean is
  6220. 4:59:07similar data is placed in the same file.
  6221. 4:59:11So similar data is placed in the same
  6222. 4:59:14file. Now let's go ahead and have a look
  6223. 4:59:17at this how how this query is going to
  6224. 4:59:19behave now. Yeah. So now
  6225. 4:59:23where age is greater than equal to 5 and
  6226. 4:59:25less than equal to 10. Does it overlap
  6227. 4:59:27with this? Yes. That means this file is
  6228. 4:59:30going to be transferred. Does it overlap
  6229. 4:59:32with this which is 2.par?
  6230. 4:59:35Yes, it does. This value and this value
  6231. 4:59:37overlaps. That means this park is going
  6232. 4:59:40to be transferred over the wire. But
  6233. 4:59:42does it overlap with these two? The
  6234. 4:59:43minimum itself is 29 and 51. Over here
  6235. 4:59:46it doesn't. So that means we've avoided
  6236. 4:59:49sending these two files over the wire.
  6237. 4:59:52It has been pruned. It has been filtered
  6238. 4:59:55that we don't need these two files.
  6239. 4:59:57Right? And that is one of the benefits
  6240. 4:59:59and beauty of the order. It collocates
  6241. 5:00:02it collocates similar data in the same
  6242. 5:00:05files due to which this becomes
  6243. 5:00:07possible. So to keep things simple I
  6244. 5:00:10have taken a single dimensional example
  6245. 5:00:13wherein we refer to the column age but
  6246. 5:00:16this very well applies to multiple
  6247. 5:00:18dimensions right where we would like to
  6248. 5:00:21filter by multiple columns and the idea
  6249. 5:00:23is very simple the idea is to take all
  6250. 5:00:25of these columns that we want to filter
  6251. 5:00:28by and then map it to single dimension.
  6252. 5:00:31So all of these column they are
  6253. 5:00:33multi-dimensional right more than one
  6254. 5:00:35columns it's multi-dimensional.
  6255. 5:00:37So we take all of these columns that we
  6256. 5:00:41want to filter by, right? So we take all
  6257. 5:00:44of these columns and then we map it to a
  6258. 5:00:47single dimension.
  6259. 5:00:50We map it to a single dimension. And how
  6260. 5:00:52does this happen? It happens when we put
  6261. 5:00:55this inside the zorder function, right?
  6262. 5:00:57When we apply Z order on the set of
  6263. 5:01:01columns, right? And this basically
  6264. 5:01:03preserves the locality. When we apply
  6265. 5:01:06Zorder on the set of columns, it
  6266. 5:01:09arranges the data layout in such a way
  6267. 5:01:11so that the locality is preserved. And
  6268. 5:01:14by locality, what I mean is that similar
  6269. 5:01:17data points are close to each other. So
  6270. 5:01:20Z order basically
  6271. 5:01:23preserves locality.
  6272. 5:01:28And by preserving locality, what we
  6273. 5:01:30essentially mean is that similar points,
  6274. 5:01:35similar data points are in the same
  6275. 5:01:37file.
  6276. 5:01:39Actually, not necessarily always in the
  6277. 5:01:42same file because in the earlier example
  6278. 5:01:44that we saw the value the age 10 was in
  6279. 5:01:471.park and then the age 10 was also in
  6280. 5:01:512. Okay, the idea is to avoid overlaps
  6281. 5:01:55to keep overlaps as minimal as possible,
  6282. 5:01:58right? Because we also don't want one
  6283. 5:02:01file to has have enormous amount of
  6284. 5:02:03data. Yeah. So, I'm going to take
  6285. 5:02:05another example where
  6286. 5:02:07we are going to have two columns which
  6287. 5:02:09is product ID and quantity and for some
  6288. 5:02:13reason we write a lot of queries, a lot
  6289. 5:02:16of filters on these two columns. Yeah.
  6290. 5:02:19So we write a lot of filters on these
  6291. 5:02:21two column. So what we want to do is
  6292. 5:02:23that we want to zorder
  6293. 5:02:26we want to zorder
  6294. 5:02:29you want to zorder by product ID
  6295. 5:02:33and the quantity.
  6296. 5:02:35So when we essentially do this what it
  6297. 5:02:38does is that it takes these two points
  6298. 5:02:41from the two-dimensional space and then
  6299. 5:02:44it applies a zorder. What zorder is is
  6300. 5:02:48is it's essentially a space filling
  6301. 5:02:50curve which maintains locality. Yeah. So
  6302. 5:02:53just to iterate again what it does is
  6303. 5:02:55that it is going to map this
  6304. 5:02:57twodimensional value which is let's say
  6305. 5:03:00this value this value and all the other
  6306. 5:03:02values that you see over here. It is
  6307. 5:03:04going to map two-dimensional values to a
  6308. 5:03:07single dimensional value.
  6309. 5:03:12Yeah. And the beauty of this curve this
  6310. 5:03:15curve zord order is that the points
  6311. 5:03:17which are closed in the two-dimensional
  6312. 5:03:20space they are also going to be closed
  6313. 5:03:23in the single dimensional space. So the
  6314. 5:03:26points which are closed in two dimension
  6315. 5:03:29when you map it to a single dimension
  6316. 5:03:31they are also going to be closed. Yeah.
  6317. 5:03:33So for argument sake in order to make
  6318. 5:03:36you in order to help you understand this
  6319. 5:03:38better let's say what the order does is
  6320. 5:03:41that it simply adds values right and
  6321. 5:03:44this is just to make you understand
  6322. 5:03:45things better yeah so let's take all
  6323. 5:03:47these values the values the the values
  6324. 5:03:52which have product ID equals 10 and
  6325. 5:03:54let's simply add them right so this is
  6326. 5:03:56going to be 15 this is going to be 19
  6327. 5:04:01this is going to be 18 This is going to
  6328. 5:04:04be 22
  6329. 5:04:06and this is going to be 14. This is
  6330. 5:04:10going to be 12 and this is going to be
  6331. 5:04:1211. Now what I do is that I start
  6332. 5:04:16putting the data points from the lowest
  6333. 5:04:18to the highest by Z order. Right? So the
  6334. 5:04:21lowest is this one. I basically put 10
  6335. 5:04:23and 10 and one over here. Then we have
  6336. 5:04:2612. I put 10 and two over here. Then we
  6337. 5:04:29have 14. I put 10 and four over here.
  6338. 5:04:31Similarly 10 and 5
  6339. 5:04:34and then 10 and 8 this one over here and
  6340. 5:04:38then 10 and 9 and then finally 10 and
  6341. 5:04:4012. So what we see is that points which
  6342. 5:04:43are close
  6343. 5:04:45points which are close in the two
  6344. 5:04:46dimensional space which is 10 and 1.
  6345. 5:04:48This is very close to 10 and 2. They are
  6346. 5:04:52also close in the single dimensional
  6347. 5:04:55space after doing a zord. Right? So 11
  6348. 5:04:58and 12 are closed. So if we put the
  6349. 5:05:01values over here they are closed these
  6350. 5:05:03are placed close to each other right
  6351. 5:05:05just one row apart right so this makes
  6352. 5:05:08sure zorder makes sure that when points
  6353. 5:05:11are reduced from a multi-dimensional
  6354. 5:05:13space to a single dimensional space they
  6355. 5:05:17are collocated right they are close to
  6356. 5:05:19each other and using this using this we
  6357. 5:05:24simply create the files we place all of
  6358. 5:05:26this in one file over here and then all
  6359. 5:05:28the other data which is again
  6360. 5:05:30colloccated
  6361. 5:05:32is placed in this file over here. Now we
  6362. 5:05:34don't want to create one paret file for
  6363. 5:05:3611 and 12 and 15 because of course we
  6364. 5:05:40don't want to end up with the small file
  6365. 5:05:41problem. We want the data to be
  6366. 5:05:44appropriately sized. Yeah. So that is
  6367. 5:05:46why we don't create one file for each of
  6368. 5:05:49these values. So I hope this gives you
  6369. 5:05:52some kind of intuition on how zorder
  6370. 5:05:55works internally. It basically brings
  6371. 5:05:58similar data in the same file. Sometimes
  6372. 5:06:02not always in the same file because
  6373. 5:06:03again as I discussed earlier we don't
  6374. 5:06:05want to put too much of data in the same
  6375. 5:06:08file. Right? If we had a lot of rows for
  6376. 5:06:11number 10 over here, product ID number
  6377. 5:06:1310 over here. Probably some of them
  6378. 5:06:15would have also have gone over here in
  6379. 5:06:17the second file. Right? Because we don't
  6380. 5:06:19want one file to have enormous amount of
  6381. 5:06:22data. We want it to be appropriately
  6382. 5:06:25sized. Yeah. So it collocates similar
  6383. 5:06:28data in the same files and the goal is
  6384. 5:06:32to skip to skip as much amount of files
  6385. 5:06:36and data points as possible and that is
  6386. 5:06:39going to help us run the queries faster
  6387. 5:06:42scan lesser number of files and send
  6388. 5:06:45lesser amount of data over the wire.
  6389. 5:06:47Okay. So let's see zorder in action and
  6390. 5:06:50for that I'm going to use this file
  6391. 5:06:52again. So let's go ahead and read this
  6392. 5:06:55park for read.par k and this file and
  6393. 5:07:00this time I am going to select few
  6394. 5:07:02columns only it is going to be customer
  6395. 5:07:05id and the category price quantity and
  6396. 5:07:09invoiced perfect. So let's go ahead and
  6397. 5:07:12run this and let me do a dm dotlimit of
  6398. 5:07:16five.
  6399. 5:07:19Okay. And I want to generate a lot of
  6400. 5:07:23rows. So what I'm simply going to do is
  6401. 5:07:28uh I'm going to run a loop. Expected
  6402. 5:07:31rows equals
  6403. 5:07:3420 million.
  6404. 5:07:37So we simply do 2000 0 0 0.
  6405. 5:07:42Okay. And now while dfun dot count
  6406. 5:07:48is less than equal to expected rows and
  6407. 5:07:52df
  6408. 5:07:53union equals
  6409. 5:07:56df union. So I keep unioning the same
  6410. 5:07:59data frame in the loop and
  6411. 5:08:03the count is simply df union dot count
  6412. 5:08:09and I copy this and put the
  6413. 5:08:14final count to be this one. Okay. Now I
  6414. 5:08:17also want to write this data frame
  6415. 5:08:22to a table.
  6416. 5:08:24save as table in our delta catalog dot
  6417. 5:08:28delta db dot zorder example one. Yeah.
  6418. 5:08:34Okay. So this is complete and now we are
  6419. 5:08:38writing it to this table. While this is
  6420. 5:08:41going on I want to write a query
  6421. 5:08:46basically to see how long does it take
  6422. 5:08:48to run. Right? And this is simply going
  6423. 5:08:52to be uh
  6424. 5:08:58select
  6425. 5:09:00category comma
  6426. 5:09:04category comma sum of price into
  6427. 5:09:08quantity.
  6428. 5:09:10Price into quantity
  6429. 5:09:14as total sales,
  6430. 5:09:17right? as total sales from this table
  6431. 5:09:21right here
  6432. 5:09:24from this table
  6433. 5:09:26where customer ID
  6434. 5:09:29equals 2011
  6435. 5:09:32group by category
  6436. 5:09:35right so I'm purposely
  6437. 5:09:37doing a filter on customer ID equals 201
  6438. 5:09:42and we are going to compare it compare
  6439. 5:09:44the performance of this query after a
  6440. 5:09:47the order has been applied. Yeah. Okay.
  6441. 5:09:50So the write is complete. Now let's go
  6442. 5:09:53ahead and run this query. And let me
  6443. 5:09:56also Okay. So the wall time is 785
  6444. 5:09:59milliseconds.
  6445. 5:10:00And let me also
  6446. 5:10:04see how this table looks like.
  6447. 5:10:08I'm going to take the location from here
  6448. 5:10:12and
  6449. 5:10:15basically check over here.
  6450. 5:10:18Okay, so here it is and then it created
  6451. 5:10:21a lot of files
  6452. 5:10:23and let me download the delta log over
  6453. 5:10:25here
  6454. 5:10:29and let's open it. So it added how many
  6455. 5:10:33it added?
  6456. 5:10:35255 park files. Yeah. And what we see
  6457. 5:10:39over here is that each of these files
  6458. 5:10:42have a minimum value for customer ID
  6459. 5:10:44which is 201
  6460. 5:10:46and a max value of 99457.
  6461. 5:10:50Right? So if I write a query which is
  6462. 5:10:53something like this where customer ID
  6463. 5:10:56equals 201 it is going to scan all of
  6464. 5:11:00those files all of these 255 files
  6465. 5:11:03because the minimum value of customer ID
  6466. 5:11:05in each of these file is 201 from 2019
  6467. 5:11:104557. So it ends up scanning four uh 255
  6468. 5:11:14files. Yeah. Now, now let's go ahead and
  6469. 5:11:18run an optimize, right?
  6470. 5:11:21A Z order with an optimize. So, optimize
  6471. 5:11:24this table.
  6472. 5:11:27Optimize this table. Z order by customer
  6473. 5:11:31ID. Now, Z order and optimize go hand in
  6474. 5:11:35hand, right? Whenever you want to run a
  6475. 5:11:37Z order by, you run it together with an
  6476. 5:11:39optimize. Optimize is going to create
  6477. 5:11:43larger files, right? It is going to
  6478. 5:11:45merge all of the small files into a
  6479. 5:11:47large bin or a large file and while this
  6480. 5:11:50process happens while it is putting the
  6481. 5:11:52small files together zorder is going to
  6482. 5:11:56colllocate similar data. Yeah. So that
  6483. 5:11:58is the reason why they work hand in
  6484. 5:12:00hand. They work together. So you can
  6485. 5:12:03either use this SQL or you can also do
  6486. 5:12:07this. So from delta dot from delta
  6487. 5:12:12import
  6488. 5:12:14delta.ts
  6489. 5:12:15import data table
  6490. 5:12:18and we simply do table equals
  6491. 5:12:22delta table dot for name and then table
  6492. 5:12:26dot optimize
  6493. 5:12:33table.optimize optimize dotexecute
  6494. 5:12:35zord order by and then we put in the
  6495. 5:12:38customer ID over here. So you can use
  6496. 5:12:40any of these two approaches. Yeah.
  6497. 5:12:44So I'll comment this for now and I will
  6498. 5:12:47go ahead and run this.
  6499. 5:12:51Okay. So the Z order is complete. And if
  6500. 5:12:54I were to look at some of the metrics,
  6501. 5:12:56what it says is that number of files
  6502. 5:12:59added is two and it removed 256
  6503. 5:13:04files. Yeah. So we're going to look at
  6504. 5:13:07all of that. But before we do that, let
  6505. 5:13:10me run this query again and let's
  6506. 5:13:13compare the run times. Yeah. So let's
  6507. 5:13:16run this query. And it is 144
  6508. 5:13:19milliseconds and the last time was 785
  6509. 5:13:22milliseconds. Six close to six times of
  6510. 5:13:26improvement. Yeah. And that's a
  6511. 5:13:28significant improvement. Let's also see
  6512. 5:13:30what happened behind the scenes. So if I
  6513. 5:13:34just refresh this and if I download this
  6514. 5:13:37log,
  6515. 5:13:39let's download this and let's have a
  6516. 5:13:41look.
  6517. 5:13:42So all of these files which were earlier
  6518. 5:13:45added there is a remove for all of this
  6519. 5:13:48right and we have add for only two of
  6520. 5:13:52them that means two new park files have
  6521. 5:13:55been added and let's quickly have a look
  6522. 5:13:57at the statistics so minimum value for
  6523. 5:14:01the customer ID is from 2011
  6524. 5:14:04until 48632
  6525. 5:14:08and the other file is from 48 632 until
  6526. 5:14:1399457. Right? So this is something
  6527. 5:14:16really amazing that has been done. What
  6528. 5:14:18it first did is that it removed all of
  6529. 5:14:22the 255 or 256 whatever number of files
  6530. 5:14:25were there. It removed all of it and
  6531. 5:14:27then it compacted the data into two
  6532. 5:14:29files. That is where the role of
  6533. 5:14:31optimize comes in and then similar data
  6534. 5:14:34was colllocated. Right? It collocated
  6535. 5:14:37data from 2011 until 48632.
  6536. 5:14:42The remaining was put in the other park
  6537. 5:14:45file. Now I don't need so initially what
  6538. 5:14:49I was doing is okay I don't have that
  6539. 5:14:51statistics over here. Let me go over
  6540. 5:14:53here. Initially what I was doing is I
  6541. 5:14:55was scanning all of the files right I
  6542. 5:14:58was scanning customer ID 220 2016 2011
  6543. 5:15:02in all of the files right because the
  6544. 5:15:04minimum value told me to do so now I
  6545. 5:15:07just need to scan just one file right
  6546. 5:15:11and that's a remarkable improvement so
  6547. 5:15:13that is how zorder works and I hope that
  6548. 5:15:16gave you some insight into all of this
  6549. 5:15:18right so now let's talk about how do you
  6550. 5:15:21apply zorder to together with hive style
  6551. 5:15:25partitions. Yeah. So let's let's try to
  6552. 5:15:28mimic that example and I'm going to use
  6553. 5:15:30this data frame over here. So this is
  6554. 5:15:34going to be let's try to create those
  6555. 5:15:36partitions
  6556. 5:15:38dot mode override dot partition by and
  6557. 5:15:42let's partition by invoice
  6558. 5:15:45invoice date and let's save this table
  6559. 5:15:47as delta catalog dot delta db dot zorder
  6560. 5:15:53example 2. Yeah. So let's go ahead and
  6561. 5:15:56run this and it should basically create
  6562. 5:15:59partitions on the invoice date. Yeah. So
  6563. 5:16:02now we've earlier seen that we applied Z
  6564. 5:16:06order on the customer ID. So generally
  6565. 5:16:10there's a pattern that people follow is
  6566. 5:16:12if we have a hive style partition on
  6567. 5:16:16let's say the invoice date and our we
  6568. 5:16:20can use this together.
  6569. 5:16:23We can zorder by
  6570. 5:16:26we can zorder by the customer ID for
  6571. 5:16:30each of those invoice date.
  6572. 5:16:34Yeah.
  6573. 5:16:39So we can zorder by customer ID for each
  6574. 5:16:43of those invoice days. Yeah. So now
  6575. 5:16:49let's imagine that a new partition is
  6576. 5:16:51coming in and let me quickly
  6577. 5:16:55write some code to mimic that. So we
  6578. 5:16:59going to do dfm do.filter f dot column
  6579. 5:17:02invoice
  6580. 5:17:04date equal 2023 0101. So I basically
  6581. 5:17:09take in all of the data for that day and
  6582. 5:17:11I just rename it.
  6583. 5:17:14I rename it
  6584. 5:17:18to
  6585. 5:17:232025 0504.
  6586. 5:17:27Yeah. So I'm basically taking in all of
  6587. 5:17:29the data for someday and then I'm just
  6588. 5:17:31renaming it to today's date and this is
  6589. 5:17:34simply to create some new data mimicking
  6590. 5:17:37the fact that new data came in today and
  6591. 5:17:40now I want to write uh write a new
  6592. 5:17:43partition to my table. Yeah. And that
  6593. 5:17:46table is basically this table. So let's
  6594. 5:17:49also see how this looks like.
  6595. 5:17:52So,
  6596. 5:18:02so basically this table has all the
  6597. 5:18:04invoice dates, right? And
  6598. 5:18:07new data has come in for today and I
  6599. 5:18:10want to basically
  6600. 5:18:14write this data to this table. Let's
  6601. 5:18:18quickly check that there are records
  6602. 5:18:21over here in this table
  6603. 5:18:24and then okay perfect we do have
  6604. 5:18:26records. So now let's go ahead and
  6605. 5:18:29simply write this right do mode append
  6606. 5:18:34and then
  6607. 5:18:39partition by
  6608. 5:18:43invoice date and then save as this
  6609. 5:18:45table. Right? So this would mean that we
  6610. 5:18:47just wrote in a new partition. And let's
  6611. 5:18:50also verify that select max of
  6612. 5:18:54invoice
  6613. 5:18:56invoice date from
  6614. 5:18:59this table delta catalog. Delta DB dot Z
  6615. 5:19:04order example 2. And that should show us
  6616. 5:19:07this new date that we put in which is
  6617. 5:19:08the same. Now let's go ahead and apply a
  6618. 5:19:13Z order.
  6619. 5:19:15So we are going to use it together with
  6620. 5:19:17an optimize
  6621. 5:19:21Z order by customer ID but we only want
  6622. 5:19:27to do it for
  6623. 5:19:30the current invoice date. Yeah. So the
  6624. 5:19:34current invoice date equals
  6625. 5:19:362025
  6626. 5:19:3905 04. Yeah. So we have our hive style
  6627. 5:19:45partitions and given that we know that
  6628. 5:19:47our queries for whatever reason ei is a
  6629. 5:19:49lot on customer ID we zorder by customer
  6630. 5:19:52ID and that is this is also a pattern
  6631. 5:19:54which is widely used and in your
  6632. 5:19:56pipeline when you are running your
  6633. 5:19:58pipelines on a daily basis you can
  6634. 5:19:59simply parameterize this you can simply
  6635. 5:20:03parameterize this part of the command
  6636. 5:20:05you can simply put in current day
  6637. 5:20:10current day minus one right because
  6638. 5:20:12let's Okay, you run this pipeline on
  6639. 5:20:15yesterday's data because you have all of
  6640. 5:20:17that data. Yeah. Now let's talk about
  6641. 5:20:19the final optimization technique which
  6642. 5:20:21is liquid clustering. Yeah. So by now
  6643. 5:20:25we've seen things like high style
  6644. 5:20:28partitioning where each partition value
  6645. 5:20:30gets its own folder and all of the
  6646. 5:20:32records for that partition value goes to
  6647. 5:20:34that folder and this essentially speeds
  6648. 5:20:37up queries. On the other side we have
  6649. 5:20:39zorder which optimizes your data layout.
  6650. 5:20:43It colllocates data so that we can pick
  6651. 5:20:45up relevant files right we pick up
  6652. 5:20:49minimum number of files and then we scan
  6653. 5:20:51them right so all of this makes your
  6654. 5:20:54queries faster however the biggest
  6655. 5:20:57problems with this approach is that
  6656. 5:20:59they're not flexible yeah so let me note
  6657. 5:21:02all of this down so the first one is
  6658. 5:21:05hive style partitioning
  6659. 5:21:09hive style partitioning and the Second
  6660. 5:21:12one is Z order.
  6661. 5:21:16Now the problems with these two
  6662. 5:21:18approaches is that they are not
  6663. 5:21:21flexible.
  6664. 5:21:23They are not flexible. And when I say
  6665. 5:21:25that they are not flexible, what I mean
  6666. 5:21:28is that you have to decide your
  6667. 5:21:30partitioning column or Z order column up
  6668. 5:21:34front.
  6669. 5:21:36Right? These two have to be decided up
  6670. 5:21:40front. Either whenever you do the
  6671. 5:21:42partition by and the column name or the
  6672. 5:21:45Z order by and the column name the
  6673. 5:21:47column has to be decided up front.
  6674. 5:21:50Right? Now the disadvantage with that is
  6675. 5:21:52if today your filter pattern is by a
  6676. 5:21:56column called country
  6677. 5:22:00and few months down the line it changes
  6678. 5:22:02to another column called category
  6679. 5:22:06called category.
  6680. 5:22:09The layout of your data is by country as
  6681. 5:22:12of now, right? It is optimized in such a
  6682. 5:22:15way such that whenever somebody queries
  6683. 5:22:18by country, it's very easy to find the
  6684. 5:22:20records. But when the filter pattern
  6685. 5:22:23changes over time, when it changes to
  6686. 5:22:25category, that data layout is no more
  6687. 5:22:28helpful. Right? So essentially what you
  6688. 5:22:30end up doing is you'll probably end up
  6689. 5:22:33scanning the whole data set and the data
  6690. 5:22:36layout is no more helpful. What you
  6691. 5:22:38essentially need to do is you'll end up
  6692. 5:22:43rewriting
  6693. 5:22:46rewriting data according to the new
  6694. 5:22:49filter pattern which is by category.
  6695. 5:22:53Right? So that is the biggest
  6696. 5:22:56disadvantage of these two approaches.
  6697. 5:22:58Now this is where liquid clustering
  6698. 5:23:00comes in. Right? So it allows you to
  6699. 5:23:03change the clustering columns anytime.
  6700. 5:23:06Yeah. So it allows you to change the
  6701. 5:23:09clustering columns.
  6702. 5:23:12So this can be changed anytime
  6703. 5:23:16right and that is the biggest benefit of
  6704. 5:23:20using liquid clustering. It is very
  6705. 5:23:23flexible in nature. It is incremental
  6706. 5:23:26meaning that you can change the
  6707. 5:23:27clustering columns going ahead. Yeah, it
  6708. 5:23:30is going to use. So let's say you have
  6709. 5:23:32particular set of clustering columns and
  6710. 5:23:34down the line when you want to change
  6711. 5:23:36it, it is going to use those new
  6712. 5:23:38clustering columns in order to decide
  6713. 5:23:41the layout of data. So the layout of
  6714. 5:23:43data is then going to change going
  6715. 5:23:46ahead. So the algorithm that liquid
  6716. 5:23:48clustering uses under the hood, it
  6717. 5:23:51maintains a balanced layout and by
  6718. 5:23:54balanced layout what I mean is that it
  6719. 5:23:57ensures two things. The first one is
  6720. 5:24:00uniform file size.
  6721. 5:24:02Uniform file size. And the second one is
  6722. 5:24:07appropriate number of files.
  6723. 5:24:11Appropriate number of files. Right? So
  6724. 5:24:14it is going to ensure that appropriate
  6725. 5:24:17number of files are created. It doesn't
  6726. 5:24:19end up creating lots of file with
  6727. 5:24:21minimal amount of data. So that is where
  6728. 5:24:23you avoid the small file problem. And it
  6729. 5:24:26also ensure that those files are
  6730. 5:24:30appropriately sized. Right? So it ensure
  6731. 5:24:33the number of files and the size of
  6732. 5:24:35those files. And a good point to note is
  6733. 5:24:38that in all of these files your data is
  6734. 5:24:41colllocated
  6735. 5:24:45which simply means that similar data is
  6736. 5:24:48going to reside in the same file. Yeah.
  6737. 5:24:52So let's understand this with an example
  6738. 5:24:54and I'm going to refer to Denny Lee's
  6739. 5:24:56blog who is a developer advocate at data
  6740. 5:24:58bricks and also a spark and mlflow
  6741. 5:25:01contributor. So here I'm in the blog and
  6742. 5:25:03let me quickly enable annotation.
  6743. 5:25:07Let me zoom this a little bit.
  6744. 5:25:10Okay. And here we go. So this was the
  6745. 5:25:14diagram that I was referring to. So all
  6746. 5:25:17of these boxes that you see over here,
  6747. 5:25:19right? The boxes with the years put in
  6748. 5:25:21over here, these are basically
  6749. 5:25:24partitions, right? These are basically
  6750. 5:25:27partitions. And these arrows that you
  6751. 5:25:29see over here, these arrows, these are
  6752. 5:25:33basically task. And in Spark
  6753. 5:25:35terminology, a task takes up and
  6754. 5:25:38processes one partition, right? A task
  6755. 5:25:41basically takes up and processes one
  6756. 5:25:43partition. And all of the partitions
  6757. 5:25:46over here as you see they are uniformly
  6758. 5:25:51uniformly or evenly side.
  6759. 5:25:55They are uniformly or evenly sized.
  6760. 5:25:57Right? So all of these tasks are going
  6761. 5:26:00to complete almost on the same time.
  6762. 5:26:04Right? Because the amount of data that
  6763. 5:26:06they getting to process is just the
  6764. 5:26:08same. Now this is something that
  6765. 5:26:10actually doesn't happen in a real world
  6766. 5:26:12scenario. Let me show you what would
  6767. 5:26:14happen in a real world scenario. So a
  6768. 5:26:17real world scenario would look something
  6769. 5:26:19like this. Probably one of the partition
  6770. 5:26:21would have too much of data and some of
  6771. 5:26:23the partitions would have very less
  6772. 5:26:25amount of data. Right? So we see that
  6773. 5:26:272023 and 2022 have a lot more data than
  6774. 5:26:30the others like 2020 and 2004. So these
  6775. 5:26:34task these tasks are going to be on the
  6776. 5:26:38critical path. So they are going to be
  6777. 5:26:41on the critical path. And by that what I
  6778. 5:26:43mean is that these two tasks need to
  6779. 5:26:47complete in order for the whole job to
  6780. 5:26:50to complete. Right? So these two tasks
  6781. 5:26:52need to complete in order for the whole
  6782. 5:26:54job to complete. Yeah.
  6783. 5:26:58They are going to take the most amount
  6784. 5:27:00of time.
  6785. 5:27:04They're going to take the most amount of
  6786. 5:27:06time. And the problem with this is that
  6787. 5:27:09these two cores, the core processing,
  6788. 5:27:11these two are going to be occupied while
  6789. 5:27:14all the others are going to remain idle.
  6790. 5:27:18They are going to remain idle. Yeah. So
  6791. 5:27:21that is underutilization of your
  6792. 5:27:23resources. Now what liquid clustering is
  6793. 5:27:25going to do is that it is going to
  6794. 5:27:28combine all of the small partitions
  6795. 5:27:31together into something meaningful
  6796. 5:27:36right into something meaningful. So this
  6797. 5:27:382018 partition is added days it combined
  6798. 5:27:412019 and 2004 into one 2020 and 2021
  6799. 5:27:46into one. Right? But we still face this
  6800. 5:27:48problem because 2023 and 2022 these are
  6801. 5:27:51still big fat chunks that one core or
  6802. 5:27:56one task would need to process and this
  6803. 5:27:58is going to become a bottleneck because
  6804. 5:28:00these guys have to do a lot of work.
  6805. 5:28:02Yeah. So again what liquid clustering
  6806. 5:28:04does is very smartly it breaks it down.
  6807. 5:28:08It divides the larger chunks into
  6808. 5:28:11smaller buckets. Right? So you see that
  6809. 5:28:142023 had been divided into three part
  6810. 5:28:18and 2022 has been divided into two
  6811. 5:28:21parts. Now the core which takes this up
  6812. 5:28:24the task which run over here all of them
  6813. 5:28:27are very uniformly signed. This becomes
  6814. 5:28:30similar to the ideal case that we had
  6815. 5:28:33over here. Right? this ideal so-called
  6816. 5:28:38ideal case that we had, right? And
  6817. 5:28:41therefore my resource utilization
  6818. 5:28:45is going to go up, right? And we are
  6819. 5:28:49going to complete
  6820. 5:28:53we are going to complete
  6821. 5:28:56this job
  6822. 5:28:58in a reasonable amount of time, right?
  6823. 5:29:00because every task or every core has got
  6824. 5:29:04even amount of data to process right so
  6825. 5:29:06that's the beauty of liquid clustering
  6826. 5:29:09and this example is on a single
  6827. 5:29:11dimension let's see what would happen in
  6828. 5:29:14multi-dimension so let's say we want to
  6829. 5:29:16cluster we want to cluster by two
  6830. 5:29:19columns
  6831. 5:29:22and they are basically the year and
  6832. 5:29:26customer
  6833. 5:29:28right so in this case we are also going
  6834. 5:29:31to analyze the sizes for each of the
  6835. 5:29:34combinations and let's say that for year
  6836. 5:29:36I have 2023 2022 and 2021 and these are
  6837. 5:29:41my customers right so we see that these
  6838. 5:29:45are the data sizes that have been marked
  6839. 5:29:48over here right so red is small yellow
  6840. 5:29:52is medium size green is optimal and blue
  6841. 5:29:56is excel right so other than these to
  6842. 5:29:59these two values
  6843. 5:30:02for the year and customer enterprise
  6844. 5:30:04customer in 2023.
  6845. 5:30:06This is optimally sized and 2021
  6846. 5:30:09enterprise customer is optimally sized.
  6847. 5:30:12All of the others either they are small
  6848. 5:30:14in size the files are either small in
  6849. 5:30:17size or they are large in size like this
  6850. 5:30:20one. Right? So what is going to happen
  6851. 5:30:22is that liquid clustering is going to
  6852. 5:30:25combine the file sizes and make them
  6853. 5:30:28appropriately size. Right? So something
  6854. 5:30:31like this is going to happen. It is
  6855. 5:30:33going to combine medium two medium size
  6856. 5:30:36and one small size file. Again two
  6857. 5:30:39medium and one small size file.
  6858. 5:30:43Three small size and two medium size and
  6859. 5:30:46two medium and one small size file.
  6860. 5:30:48Right? So it is going to combine all of
  6861. 5:30:51this and create one optimal size file.
  6862. 5:30:57Optimal size file.
  6863. 5:31:00Right now once this is done we we've
  6864. 5:31:02seen that all of the files are now
  6865. 5:31:05optimally sized except for this one. Now
  6866. 5:31:08what it is going to do is that it is
  6867. 5:31:10going to break this down into smaller
  6868. 5:31:13file sizes.
  6869. 5:31:15Yeah. So it is going to break this down
  6870. 5:31:18into smaller file sizes. And
  6871. 5:31:22the final result that we are going to
  6872. 5:31:24get is something like this. So now we
  6873. 5:31:26see that all of the sizes all of the
  6874. 5:31:29files are appropriately sized and we
  6875. 5:31:32have appropriate number of files and
  6876. 5:31:34that's the beauty of liquid clustering.
  6877. 5:31:36So it helps you achieve two things. The
  6878. 5:31:38first one is the small file problem.
  6879. 5:31:45So it helps you avoid the small file
  6880. 5:31:48problem. We've seen several files which
  6881. 5:31:50were marked red, right? They were small
  6882. 5:31:53files. So what it did is that it merged
  6883. 5:31:55all of them and then it created an
  6884. 5:31:58appropriately sized file. That's the
  6885. 5:32:00first one. The second problem that it
  6886. 5:32:02helps avoid is data skew.
  6887. 5:32:06We've seen that we had
  6888. 5:32:09partitions which had lots of data and
  6889. 5:32:12that is where this example came in.
  6890. 5:32:14Right? This was that example where one
  6891. 5:32:16partition had a lot of data and we also
  6892. 5:32:20had this example
  6893. 5:32:23where the partition 2023 was skewed.
  6894. 5:32:28Right? So what it did was that it
  6895. 5:32:29divided this partition into appropriate
  6896. 5:32:33number of parts and that is how it
  6897. 5:32:35avoided the skew problem. Now it's very
  6898. 5:32:38important to note is that when we talk
  6899. 5:32:41about skew we think of it as there is
  6900. 5:32:44one partition and it has tremendous
  6901. 5:32:47amount of data to process and for that
  6902. 5:32:49reason it takes forever to process it
  6903. 5:32:51right. So this is avoided by dividing it
  6904. 5:32:54into respective parts. But let's say you
  6905. 5:32:57have a data set it has 1 million keys
  6906. 5:33:00for a particular value right. So let's
  6907. 5:33:02say um there is a key and for that key
  6908. 5:33:05there are 1 million values
  6909. 5:33:09and when you do a join the same keys go
  6910. 5:33:11to the same partition in spark right so
  6911. 5:33:14it is going to create a skew again now
  6912. 5:33:17I'm mentioning this in order to point
  6913. 5:33:19out that this doesn't liquid clustering
  6914. 5:33:22doesn't change shuffle behavior
  6915. 5:33:26right
  6916. 5:33:27what it helps avoid is a skew at this
  6917. 5:33:31point Right? It basically sizes the file
  6918. 5:33:34appropriately so that all of the tasks
  6919. 5:33:37get appropriately sized partitions.
  6920. 5:33:39Right? The sizes of the partitions that
  6921. 5:33:41they get are not skewed. Okay. So now
  6922. 5:33:44let's see liquid clustering in action.
  6923. 5:33:46What I'm going to do is that I am going
  6924. 5:33:48to create two tables. The first one I'm
  6925. 5:33:50going to partition and zorder by and the
  6926. 5:33:53second one I'm going to do a liquid
  6927. 5:33:56clustering. And then let's run a query
  6928. 5:33:58on both the tables and compare the
  6929. 5:34:00runtime. Yeah. So, let me quickly copy
  6930. 5:34:04all of these paths. So, this is going to
  6931. 5:34:07be spark dot
  6932. 5:34:09read.park.
  6933. 5:34:10[Music]
  6934. 5:34:30Okay, this has been read in. Now let me
  6935. 5:34:33union all of these files, all of these
  6936. 5:34:36data frames. DF2 dot union df3.
  6937. 5:34:44Let's also select some relevant columns
  6938. 5:34:47from here because we don't want to there
  6939. 5:34:50are too many columns in here. uh
  6940. 5:34:53customer ID, category,
  6941. 5:34:57price, quantity,
  6942. 5:34:59and the invoice date. Yeah. So now that
  6943. 5:35:03this is done, let me go ahead and write
  6944. 5:35:06this file. Write dot mode override
  6945. 5:35:12dot partition by
  6946. 5:35:14let's partition by invoice date.
  6947. 5:35:17Something similar that we've done
  6948. 5:35:19earlier. And let's save this file as
  6949. 5:35:23delta catalog delta DB dot liquid
  6950. 5:35:28clustering example one.
  6951. 5:35:31And
  6952. 5:35:33let's now run a Z order on this file.
  6953. 5:35:38Delta catalog delta DB dot
  6954. 5:35:43example one Z order by customer ID.
  6955. 5:35:48Yeah, because you remember we've done a
  6956. 5:35:50similar uh we've run a query where we
  6957. 5:35:53filtered by customer ID all here, right?
  6958. 5:35:55So for that reason I want to zorder by
  6959. 5:35:58customer ID and then I'll run the same
  6960. 5:36:01kind of query. Yeah. So just to quickly
  6961. 5:36:04show you once again
  6962. 5:36:08I want to
  6963. 5:36:11mimic this kind of a query which we've
  6964. 5:36:13run earlier.
  6965. 5:36:16Yeah, something like this. Right, let's
  6966. 5:36:19go ahead and
  6967. 5:36:26run this.
  6968. 5:36:28Okay, so now this is the first table
  6969. 5:36:32and let's also do a count to make sure
  6970. 5:36:36that
  6971. 5:36:37we have data that has landed in over
  6972. 5:36:39here.
  6973. 5:36:44Yeah, let's run this and parallelly I'm
  6974. 5:36:47also going to create another table which
  6975. 5:36:50is going to be example two and I am
  6976. 5:36:53going to do a cluster by and you
  6977. 5:36:56remember that in cases where we have a
  6978. 5:36:58partitioned and a zordered column we
  6979. 5:37:02basically take in both the partition and
  6980. 5:37:04the zord ordered column as the cluster
  6981. 5:37:06by columns right so the partition column
  6982. 5:37:08is invoice date the zorder column is
  6983. 5:37:11customer ID So I'll basically put in
  6984. 5:37:13both of these values over here. Yeah. So
  6985. 5:37:16now let's go ahead and also run this.
  6986. 5:37:19Okay. So there have been some
  6987. 5:37:20significant optimization. It added 174
  6988. 5:37:24files and 364 files have been removed.
  6989. 5:37:27So now let's have a look at the count.
  6990. 5:37:29The count is 99457
  6991. 5:37:33and the count over here is just the
  6992. 5:37:35same. So now let's go ahead and run this
  6993. 5:37:37query which is select category
  6994. 5:37:42comma sum of
  6995. 5:37:46sum of price into
  6996. 5:37:49quantity
  6997. 5:37:51as total sales
  6998. 5:37:55right as total sales
  6999. 5:37:58from
  7000. 5:37:59this table
  7001. 5:38:01where customer ID equals 201
  7002. 5:38:05and we group by category, right?
  7003. 5:38:10And I also want to time this query. So
  7004. 5:38:13I'm going to do this inside a
  7005. 5:38:18uh spark.sql.
  7006. 5:38:27And let's put this over here.
  7007. 5:38:30And let's go ahead and run this.
  7008. 5:38:35Yeah, actually let me also add in
  7009. 5:38:38something else. Let me add in
  7010. 5:38:42and invoice date between
  7011. 5:38:472021
  7012. 5:38:5001 and
  7013. 5:38:542 0 2 3 12 31. The reason why I also
  7014. 5:38:59added in this over here is because I
  7015. 5:39:01want to test out both the columns,
  7016. 5:39:03right? Because we are partitioning
  7017. 5:39:06by invoice date and then we are
  7018. 5:39:08reordering by the customer ID. Yeah.
  7019. 5:39:15Okay. So this is complete and this took
  7020. 5:39:18somewhere around 254 milliseconds. Now
  7021. 5:39:20let's also run this on the other table
  7022. 5:39:25and we see that this is 150 milliseconds
  7023. 5:39:28right not a very significant improvement
  7024. 5:39:31but still a good improvement right uh
  7025. 5:39:34but we know that the data on which we
  7026. 5:39:37are running is limited right so probably
  7027. 5:39:40we'll be able to see these improvements
  7028. 5:39:42on a larger scale when we run these
  7029. 5:39:45operation on huge amounts of data right
  7030. 5:39:47but overall I believe that this gives
  7031. 5:39:49you a is of how liquid clustering works
  7032. 5:39:53and how you can still compare when you
  7033. 5:39:55use a partition by and a reorder and
  7034. 5:39:57equivalently convert that to a cluster
  7035. 5:40:00by and there are definitely performant
  7036. 5:40:03benefits that you get out of it right
  7037. 5:40:05okay so now that we've seen liquid
  7038. 5:40:06clustering in action it's really
  7039. 5:40:08important to understand that it cannot
  7040. 5:40:10be used together with zorder or hive
  7041. 5:40:13tile partitioning right so if you're
  7042. 5:40:15using liquid clustering that will
  7043. 5:40:18independently be the mechanism for
  7044. 5:40:21deciding the layout of your data. Yeah,
  7045. 5:40:23so another question that you might have
  7046. 5:40:25in mind is how do I choose the liquid
  7047. 5:40:27clustering columns, right? The best
  7048. 5:40:30practice is to choose the most
  7049. 5:40:33frequently used column in your query
  7050. 5:40:35filters, right? And if two columns are
  7051. 5:40:39highly correlated, you should just
  7052. 5:40:41include one of them in your query
  7053. 5:40:43filters. Yeah. And what I mean by that
  7054. 5:40:45is actually let me first quickly
  7055. 5:40:48annotate. Let's say you have a column
  7056. 5:40:51called product category.
  7057. 5:40:55You have a column called product
  7058. 5:40:56category and you have another column
  7059. 5:40:58called location.
  7060. 5:41:00And these two columns are for whatever
  7061. 5:41:03reason they are highly
  7062. 5:41:06correlated.
  7063. 5:41:08So what it means is that whenever you
  7064. 5:41:09choose a particular product,
  7065. 5:41:12it is always going to yield some the
  7066. 5:41:14same location. Right? If you choose
  7067. 5:41:17another product, it is going to yield
  7068. 5:41:18the same location always. Right? So that
  7069. 5:41:21means that there is no point including
  7070. 5:41:23this column location. So in those cases,
  7071. 5:41:27you remove the columns which are highly
  7072. 5:41:29correlated. You remove location and you
  7073. 5:41:31only go ahead with product category.
  7074. 5:41:33Right? So that's the first point. The
  7075. 5:41:36second point is
  7076. 5:41:39Yeah. The second point is if you're
  7077. 5:41:40converting a table an existing table and
  7078. 5:41:44it already has some kind of strategy,
  7079. 5:41:47some kind of partitioning or zorder zord
  7080. 5:41:49strategy already being followed and now
  7081. 5:41:51you want to use liquid clustering. These
  7082. 5:41:53are the rules that you should follow. So
  7083. 5:41:55the first one is is if it's already hive
  7084. 5:41:58partition basically just go ahead and
  7085. 5:42:01use that partition column as the
  7086. 5:42:03clustering key. The second one is if
  7087. 5:42:06you're using a column for the order
  7088. 5:42:08indexing,
  7089. 5:42:10use the same Z order column for your
  7090. 5:42:13clustering. Right? The third one is high
  7091. 5:42:16style partitioning and the order. We've
  7092. 5:42:18seen this example, right? We've used
  7093. 5:42:20both we first partitioned by high style
  7094. 5:42:22partitioning and inside of it we've
  7095. 5:42:25applied the order. So there are two
  7096. 5:42:26different columns and in the previous
  7097. 5:42:28example we hive partitioned by invoice
  7098. 5:42:32date and then we reordered by category
  7099. 5:42:36right so in that case use both the
  7100. 5:42:39partition column and the zorder by
  7101. 5:42:41column as your clustering key. So your
  7102. 5:42:43new clustering key basically becomes
  7103. 5:42:45invoice date and category. So you simply
  7104. 5:42:48cluster by
  7105. 5:42:50these two columns. And the last one is
  7106. 5:42:53if you're having if you're using
  7107. 5:42:54generated column to reduce the cardality
  7108. 5:42:57for example date or time. So let's say
  7109. 5:42:59you have a time stamp column and then
  7110. 5:43:01you convert it to a date. Many times we
  7111. 5:43:03do that in order to reduce the cardality
  7112. 5:43:05right. So if we are already doing this
  7113. 5:43:08and if we partition by date in that case
  7114. 5:43:11use the original column just use the
  7115. 5:43:14time stamp as your clustering key and
  7116. 5:43:16don't create a generated column. So
  7117. 5:43:18there's no need to create a date column,
  7118. 5:43:20right? You just partition, sorry, you
  7119. 5:43:22just cluster by the time stamp. Yeah. So
  7120. 5:43:26to quickly summarize, liquid clustering
  7121. 5:43:28is going to be super helpful when your
  7122. 5:43:31query patterns are going to change down
  7123. 5:43:33the line, right? You feel that your
  7124. 5:43:35query patterns are going to change down
  7125. 5:43:36the line and you don't want to be bound
  7126. 5:43:39by fixing the partitioning or the zorder
  7127. 5:43:42column, right? So so liquid clustering
  7128. 5:43:44helps you remain flexible. you are you
  7129. 5:43:47have that flexibility of being able to
  7130. 5:43:49change the cluster by columns anytime
  7131. 5:43:51down the line. Right? So that's the
  7132. 5:43:53first benefit. The second benefit is it
  7133. 5:43:56avoids the small file problem. As we've
  7134. 5:43:59seen that it combines lot of small files
  7135. 5:44:02into appropriately sized files. It even
  7136. 5:44:05breaks down bigger files largely sized
  7137. 5:44:08files into appropriately sized one.
  7138. 5:44:10Yeah. And the third one is it avoids
  7139. 5:44:13data skew. It basically breaks up the
  7140. 5:44:15bigger partitions into smaller ones and
  7141. 5:44:18that is how each task or each score gets
  7142. 5:44:22a reasonable amount of data to process.
  7143. 5:44:25I'm super happy to see that you've
  7144. 5:44:27reached the end of the video and I
  7145. 5:44:29really hope that you learned a lot from
  7146. 5:44:32it and you enjoyed it. So, please don't
  7147. 5:44:34forget to like and share this video. Tag
  7148. 5:44:37me on LinkedIn. Share whatever you've
  7149. 5:44:39learned. I'll be more than happy to see
  7150. 5:44:41that. Please don't forget to subscribe
  7151. 5:44:44to my channel because I've seen that
  7152. 5:44:46only 30% of you have subscribed to my
  7153. 5:44:48channel. So, please go ahead and hit
  7154. 5:44:50that subscribe button. It really
  7155. 5:44:52motivates me to make a lot more content.
  7156. 5:44:54So, thank you so much for watching.
  7157. 5:44:58[Music]

About this transcript

This page contains the full transcript of Delta Lake Masterclass | Azure Databricks | PySpark | From Zero-To-Expert by Afaque Ahmad, generated from the public captions YouTube serves with the video. The transcript has 48,542 words across 7,157 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.