Delta Lake Masterclass | Azure Databricks | PySpark | From Zero-To-Expert — Transcript
Full transcript
- 0:00Hey everyone, welcome to this 6-h hour
- 0:03master class on Delta League. In this
- 0:06video, I'll be covering the concepts
- 0:08deep internals how things work under the
- 0:11hood and together with this we will be
- 0:14doing a lot of labs, lot of practicals
- 0:17to see the concepts we've studied in
- 0:19action on Azure data bricks, right? And
- 0:22trust me, this is going to be the most
- 0:24comprehensive video you will ever watch
- 0:27on Delta Lake, right? So before I dive
- 0:30in, I want to give you a quick summary
- 0:32of the topics that I'll be covering and
- 0:35please watch it in order. So we'll be
- 0:38first starting with the problems with
- 0:40data lake and how Delta Lake solved the
- 0:43problem. Some examples of it is lack of
- 0:46asset support, lack of update, merge and
- 0:49delete operations and data reliability
- 0:52and quality issues. How delta solved all
- 0:55of these issues. Right? So then I'll
- 0:57walk you through the lab setup, the lab
- 1:00architecture and we'll be doing the lab
- 1:02setup together on Azour right and then
- 1:06we'll be performing DML operations
- 1:09operations in order to get a feel of how
- 1:12is it to work with data link right we'll
- 1:15then be uncovering the delta log right
- 1:19so all of this all of the magic happens
- 1:22through the delta log so we'll be
- 1:24talking about delta log in a lot of
- 1:26details and how things work under the
- 1:28hood in the delta lock. Right. Next up,
- 1:31we'll be covering concurrency control,
- 1:33optimistic concurrency control, and
- 1:35pessimistic concurrency control, time
- 1:38travel and versioning, schema
- 1:39validation, schema evolution, converting
- 1:42paret to delta, manage and external
- 1:46tables, and one of the most important
- 1:48topics deletion vectors. What exactly is
- 1:51copy on write and merge on read? Then
- 1:55we'll be covering cloning. What exactly
- 1:58is shallow clone and deep clone. Next
- 2:01we'll be talking about the small file
- 2:03problem. The very popular small file
- 2:05problem and the root causes of the small
- 2:08file problem. Lastly we'll be closing
- 2:10this with the optimization techniques.
- 2:14Right? So we'll be covering optimize
- 2:17vacuum zorder
- 2:19and liquid clustering. Yeah. So I hope
- 2:22you're as excited as I am. So let's get
- 2:26started. So first of all, let's start by
- 2:28understanding why does Delta Lake even
- 2:30exist today, right? Why did it even come
- 2:32into picture when we already had data
- 2:35lake, right? And we are going to do this
- 2:37by simply analyzing the pros and cons of
- 2:40data lake, right? So some of the things
- 2:43that data lake was really good at was
- 2:45that it it provided flexible data
- 2:48storage. So if you were to go to uh a
- 2:50data warehouse, you could only work with
- 2:53relational data. You can work with data
- 2:56that has columns and rows, right? But
- 2:59here the data storage was very flexible.
- 3:01You could ingest structured,
- 3:03semistructured, unstructured data. And
- 3:06this can be audio, video, images, logs
- 3:10or basically anything that you wish,
- 3:12right?
- 3:13and data generated from sensors, your
- 3:15web logs, all of this could be streamed
- 3:18into your data lakeink. And all of this
- 3:20could be stored cheaply at a low cost,
- 3:23right? In inside of something like an S3
- 3:26or an ADLS, yeah, and it scaled very
- 3:30well in order to accommodate high volume
- 3:32of data, right? You could basically put
- 3:35in infinite amount of data and this
- 3:38served a lot of use cases very well. And
- 3:41some of those examples are big data,
- 3:43machine learning and developing AI
- 3:45application, right? But with all of
- 3:48this, data lake still came with a lot of
- 3:51challenges. It came with several
- 3:53challenges and that was the reason Delta
- 3:55Lake was designed in order to solve
- 3:58these problems and we are going to go
- 4:00into a lot of details. We are going to
- 4:02discuss those problems in a lot of
- 4:04detail because they are going to form
- 4:07the foundations. Right? The first
- 4:08problem was lack of acid support. Now
- 4:13what is acid and why is acid transaction
- 4:17support so important. Right? So as I
- 4:20mentioned to you before all
- 4:21understanding all of these concepts are
- 4:23really important because they build your
- 4:25foundation. So I'm going to take a lot
- 4:27of code examples and I'll be walking you
- 4:29through all of these. Right? So let's
- 4:32get started. So what does acid stand
- 4:34for? What does it mean? So acid
- 4:37basically means atomic
- 4:42consistent
- 4:46isolation
- 4:49and the last one is durability.
- 4:52Right?
- 4:56So we are going to first study about
- 4:59atomic. What does atomicity mean? Right?
- 5:02So atomicity simply means all or
- 5:05nothing. Either you do all of it or you
- 5:08don't do anything. Right? Now how does
- 5:10that apply to our case? Now let's take a
- 5:12small pseudo code. Right? Let's take a
- 5:15small pseudo code and try to understand.
- 5:17So let's say
- 5:19you were transferring some amount from
- 5:23your savings
- 5:25to your checking account,
- 5:29right? From your saving to your checking
- 5:30out account. Now, your savings account
- 5:33had $1,000
- 5:35and your checking account had $200,
- 5:38right? And you wanted to transfer $500
- 5:43from your savings to your checking
- 5:45account, right? Now, what does that
- 5:47mean? That means that this guy should
- 5:49end up having 500 and this guy should
- 5:51end up having 700. Right? So, let's have
- 5:54a look at the pseudo code for that.
- 5:55Right? So, we basically read in the
- 5:57current savings which has $1,000. We
- 6:00then read in current checking which has
- 6:01$200. Right? Now we deduct from saving
- 6:05because we want to transfer $500. Right?
- 6:08So what we say is that new savings it
- 6:10become current savings which is $1,000
- 6:13minus 2 sorry this is going to be minus
- 6:16$500.
- 6:18Yeah.
- 6:20And this is going to be 500 in turn.
- 6:24Now my account is updated. Right. So the
- 6:26savings account is now updated. Now what
- 6:29happens is my system crashes.
- 6:32The checking account I've not been able
- 6:35to update it so far, right? The system
- 6:37crashes.
- 6:39So what does this mean? That this
- 6:41transaction is lost. It's gone as a
- 6:44whole, right? So what you end up having
- 6:46is that your savings account has $500.
- 6:51Your savings account has $500 and your
- 6:54checking account has $200. That mean
- 6:56that $500 went up in the air and it
- 6:59never went back to your checking
- 7:01account. And this is a situation that
- 7:03you never want to be in, right? Because
- 7:05your system failed somewhere over here.
- 7:10And because of that changes were not
- 7:12reflected in your checking account,
- 7:14right? Now let's see how does an atomic
- 7:16system look like. So what would happen
- 7:19in an atomic system is that you would
- 7:21first begin the transaction. Yeah. Now
- 7:24once the transaction has begun you
- 7:27deduct the amount from savings which is
- 7:29500 is deducted and the total now
- 7:32becomes 500. Yeah. Now it is added to
- 7:36checking. So the initial balance was
- 7:39200. You then add 500 which becomes 700.
- 7:42Right. And then you finally commit the
- 7:45transaction. Now in between any of these
- 7:47steps if my system crashes it is simply
- 7:52going to go to the accept clause and it
- 7:54is going to roll back my transaction.
- 7:57Yeah. So either the whole transaction
- 8:00happened as a whole both the accounts
- 8:03represent the actual amounts or nothing
- 8:06happened. The whole transaction is going
- 8:08to be rolled back. Right? And this is
- 8:10what atomicity is either everything
- 8:13happens or nothing happen. Consistency
- 8:15basically means that rules must be
- 8:18enforced. A transaction must bring a
- 8:21database from one valid state to another
- 8:24valid state. Yeah. So let's understand
- 8:26that with an example. So let's say there
- 8:29is an account which has $100 in place.
- 8:34So this account basically has $100 in
- 8:36place and then there are two
- 8:38simultaneous transaction going on here.
- 8:41So you see that there is there are two
- 8:43purchases going on. One is worth $80,
- 8:46the other purchase is worth $60. Right?
- 8:49So what these two transaction do is that
- 8:52they basically read in the database at
- 8:55the same time. The first one reads in
- 8:57$100. The second one also reads in $100.
- 9:00Yeah. So they basically go through this
- 9:02line of code. They basically read in the
- 9:04balance. Now they want to find out what
- 9:06is going to be the new balance. Yeah. So
- 9:08they subtract
- 9:10$80.
- 9:12and this guy subtract $60 which is what
- 9:15we write over here balance minus amount.
- 9:17Yeah. So this is going to be $20 and
- 9:20this is going to be $40.
- 9:23Now when both of these transactions are
- 9:25running there is going to be a
- 9:27transaction which is going to make an
- 9:30update to the database first. Yeah. Now
- 9:32let's assume that this transaction the
- 9:35first one the $80 one is the one that
- 9:38makes the update first. So that means
- 9:40that $20
- 9:43is the updated balance. Yeah. Now the
- 9:46moment the database was updated with
- 9:49$20, the second transaction goes through
- 9:52and then it updates it with $40. Yeah.
- 9:55So basically run runs this piece of
- 9:57code. It updates it with $40. Now what
- 10:01fundamentally is wrong here is that you
- 10:04made a purchase of 80 + 60 which is $140
- 10:08but you only had a balance of $100.
- 10:12So this is something that brings your
- 10:15database in an inconsistent state. Okay.
- 10:18So we have the same account which has
- 10:20$100 in place and then there are two
- 10:22simultaneous transaction going on which
- 10:24is one is trying to make a purchase of
- 10:26$80 the other one is trying to make a
- 10:28purchase of $60, right? And let's say
- 10:31this transaction first starts and it
- 10:33comes over here right it basically does
- 10:36a begin transaction which simply mean
- 10:38that okay I'm going to make changes to
- 10:42the values of the database right so now
- 10:45when the second purchase tries to come
- 10:48in over here what it basically tells him
- 10:51that hey I'm going to make some changes
- 10:54so you will have to wait for me right
- 10:56yeah so last time what happened was that
- 10:58Both the transactions read in $100. Now
- 11:02the difference is that the first one
- 11:04reads in $100. The second one is
- 11:07basically for it has still not read in
- 11:11anything. And now the first transaction
- 11:13goes ahead and executes what it wants
- 11:15to. Yeah. So it basically finds in the
- 11:17balance which is $100. Now it basically
- 11:20we've put a check in place, right? We
- 11:23put a rule in place which basically
- 11:24checks whether do I have that kind of
- 11:27amount or not. Yeah. So it basically
- 11:30check that and okay we have that amount.
- 11:33So it goes ahead and updates the balance
- 11:36which is 100 minus $80 which is $20 and
- 11:40$20 is updated and committed. Now, if at
- 11:43all if you didn't have $80 within your
- 11:47account, it would come here and then it
- 11:49would return to you that you have
- 11:51insufficient fund. Yeah. And your
- 11:54transaction wouldn't take place. Or if
- 11:59any of these operation failed over here,
- 12:01it would simply go to this accept clause
- 12:04and then it would roll back the
- 12:06transaction. Yeah. So nothing would take
- 12:08place. Yeah. So now after this
- 12:12transaction, we have simply updated the
- 12:15account balance to $20. Yeah. So now let
- 12:19me quickly actually erase this so that I
- 12:21can walk you through the second
- 12:22transaction.
- 12:24Now when the second transaction comes
- 12:26in, it goes through all of this. It
- 12:28finds the balance. Now the balance is
- 12:31$20. And this statement doesn't go
- 12:34through the balance which is $20 greater
- 12:37than equal to $60 which is not true. So
- 12:40this simply returns that you have
- 12:42insufficient one. Yeah. So basically
- 12:45what this means is that every
- 12:47transaction that we did both the first
- 12:50one and the second one it brought the
- 12:53database from one valid state to another
- 12:55valid state. The database didn't end up
- 12:58in an inconsistent or an invalid state.
- 13:01And this is what consistency is all
- 13:03about. The key difference between the
- 13:06two examples is that in the second one,
- 13:08we use transactions
- 13:11to ensure consistency. Yeah. We put in
- 13:14appropriate rules in place to ensure
- 13:17that the relevant amount was there
- 13:20within the bank account. Yeah. And
- 13:23finally, we roll back anything if
- 13:25anything goes wrong. We roll back the
- 13:27transaction if anything goes wrong. So
- 13:29isolation simply is the no interference
- 13:32policy. Yep. So multiple transactions in
- 13:35a database can go on without interfering
- 13:38with each other. Each transaction is
- 13:41going to operate as if it is the only
- 13:43one running even if other transactions
- 13:46may be running simultaneously. And this
- 13:49means that the intermediate states or
- 13:52the uncommitted changes of one
- 13:54transaction are not visible to the other
- 13:57transaction. So what that means is that
- 13:59let's say there is a table t and this is
- 14:03in its current state v_sub_1. Yeah. Now
- 14:06there is a transaction which basically
- 14:08read in this table and then it is going
- 14:11on and it is going to make some changes
- 14:13and then it is going to finally end up
- 14:16in a state called v2. It is going to end
- 14:19up but it hasn't ended up yet. Yeah. So
- 14:23these are some of the steps that it is
- 14:25going to undertake in order to end up in
- 14:28a state called V2. Yeah. And all of
- 14:31these changes that you see over here are
- 14:33uncommitted right now. Now if another
- 14:36person, another transaction comes in and
- 14:39if they want to read this table, if they
- 14:43want to read this table T, it is going
- 14:46to read the V1 state of it. Yeah.
- 14:50because all of the uncommitted changes
- 14:53it is not aware about. So basically you
- 14:56see that both of these transactions
- 14:59operate independent of each other
- 15:01without worrying about each other. And
- 15:03the way Delta achieved this is by
- 15:06something called optimistic
- 15:09concurrency control.
- 15:13Optimistic concurrency control. And I'm
- 15:16not going to uh overwhelm you with a lot
- 15:19of details right now, but we are going
- 15:21to discuss how Delta achieves isolation
- 15:26using optimistic concurrency control in
- 15:28a lot of detail going ahead. Durability
- 15:30simply means that once a transaction is
- 15:34recorded in a system, it is going to
- 15:36stay there. Yeah, it is going to stay
- 15:39there irrespective of a system crash, a
- 15:42power outage or a failure or anything.
- 15:45Right. So, consider receiving your
- 15:47paycheck. Yeah. So, so let's say uh your
- 15:50current account had $1,000 in place.
- 15:54Yeah. It had $1,000 in place. This
- 15:57current balance basically reads $1,000.
- 16:01And then you got your paycheck which is
- 16:02worth $2,000. Yeah. Now the new balance
- 16:06is going to be 2,000 +,000 which is
- 16:09$3,000
- 16:11and finally your account is going to be
- 16:14updated. So your account should be
- 16:16updated with $3,000 after you've
- 16:19received your paycheck. Now imagine that
- 16:22let's say some database in some region
- 16:25goes down and this transaction goes for
- 16:30a toss. Now when the system is restored
- 16:34the old backup is taken and it is
- 16:37restored with that backup. Now that
- 16:39backup doesn't have your paycheck in
- 16:42place. Right? So you worked the whole
- 16:44month now your paycheck is gone just
- 16:46because of some system crash which which
- 16:49didn't follow the durability principle.
- 16:52So your paycheck is lost. That means the
- 16:55old backup now shows a $1,000
- 16:59account balance. So this is a durability
- 17:02problem. So how would a system with
- 17:05durability in place look like?
- 17:09So there's a quite a bit of code here.
- 17:12But again all of them are very simple
- 17:15easy sudo code. So again whenever
- 17:18somebody's going ahead and making a
- 17:21change we are going to start a
- 17:23transaction we are going to first write
- 17:26to a transaction log. Yeah. Now think of
- 17:29it as uh maintaining a ledger of what
- 17:33all is going on. Yeah. You basically
- 17:35keep on writing whatever is going on. So
- 17:37basically we say that we are going to
- 17:39start a deposit transaction. Yeah. And
- 17:42we also say what are the details of that
- 17:45transaction. So we write to the
- 17:46transaction log that okay we are going
- 17:48to deposit an amount to a particular
- 17:51account ID. And then we basically
- 17:53calculate the current balance. We
- 17:55calculate the new balance. We add the
- 17:57amount and then we update the balance.
- 17:59Yeah. Now once the balance is updated,
- 18:01we finally write it to the transaction
- 18:04log and this has the word commit and
- 18:08then we finally commit the transaction.
- 18:10Yeah. Now if anything fails over here
- 18:12inside the try block, the transaction is
- 18:16rolled back. Now you may of course have
- 18:19questioned what if
- 18:22this whole thing either failed over here
- 18:25or here or here. So let's say we were
- 18:28writing to the transaction log and then
- 18:30it failed over here. Now when the system
- 18:32restarts the only thing that it it's
- 18:35going to check the the transaction log.
- 18:37Yeah. What it is going to see is that
- 18:39okay there is something like a start.
- 18:41Okay. Let me just change the color. uh
- 18:43there is something like a start
- 18:48but then there's nothing after that.
- 18:50Yeah. So there are no details that means
- 18:52we have to roll back this transaction.
- 18:54Now the second case what if it failed
- 18:56over here over here. Yeah. So it has
- 18:59something like a start and then it has
- 19:02some detail
- 19:05but again it doesn't have a commit. That
- 19:08means this transaction wasn't committed.
- 19:10it didn't actually go into the database.
- 19:13So again, this transaction is going to
- 19:15be rolled back even if it fails over
- 19:17here. Now all of this goes on and even
- 19:21if it fails over here then also it going
- 19:22it's going to find the same start and
- 19:24detail inside of the transaction log.
- 19:26That means this transaction is going to
- 19:28be rolled back. Whatever changes were
- 19:30made is going to be rolled back. Now if
- 19:33it finally fails after this when you
- 19:36have the word commit
- 19:38in the transaction log that means that
- 19:41the data was written inside of the
- 19:45transaction log. Yeah. So it the
- 19:48database is then going to finally commit
- 19:50the transaction and it is going to be
- 19:52recorded in the database. Yeah. So this
- 19:55is how durability would look like and
- 19:58even in cases of failures you would see
- 20:00that your database would be in a
- 20:02consistent durable state. Yeah. So with
- 20:06all of these examples I hope you
- 20:08understand how important asset
- 20:11properties are and unfortunately data
- 20:13lake doesn't have any of these
- 20:15properties. The second problem is the
- 20:18lack of support for update, merge, and
- 20:21delete. And we all know that these are
- 20:23really important operations, something
- 20:25that we do on a day-to-day basis, right?
- 20:27And the problem with traditional data
- 20:29lakes is that the data stored is in
- 20:33immutable files, right? Something that
- 20:35cannot be changed. Those files cannot be
- 20:37changed. And this creates a problem when
- 20:40you want to update or delete that data.
- 20:43Yeah. So let's take an example. Let's
- 20:45say you're running an e-commerce company
- 20:46and then you have a list of customers
- 20:48and then you want to update the customer
- 20:52addresses for some of the customers.
- 20:54Yeah. So let's say that first of all
- 20:57your data is written to this location
- 20:59and now you want to update customer
- 21:02address
- 21:04at this location. Yeah. So the steps
- 21:07that you need to follow in order to do
- 21:09that is first of all read all of that
- 21:11data. Yeah. and then you need to read
- 21:15and load it into the memory, make all
- 21:17the changes and then finally write it
- 21:20back. So these are the three steps that
- 21:22you need to do in order to update
- 21:24customer address. Now imagine
- 21:27if your system fails at any of these
- 21:30steps, you are left with inconsistent
- 21:34data.
- 21:37You are left with inconsistent data.
- 21:39Yeah, there also may be a part
- 21:42possibility of partial rights
- 21:47at this location.
- 21:49Partially written data at this location
- 21:51which basically mean that your data is
- 21:53corrupt. Now if somebody basically reads
- 21:56the data at this location, they are
- 21:59going to be basically reading corrupt
- 22:00data. Third problem is data reliability
- 22:03and quality issues, right? And a prime
- 22:06example of that is no schema
- 22:08enforcement. Yeah. So let's say you're
- 22:11collecting customer signup data and
- 22:13today you get a record which looks
- 22:15something like this. Yeah. Which have
- 22:18the name, phone number, email and phone
- 22:20number. Yeah. Now tomorrow let's say you
- 22:23get a record which looks something like
- 22:25this which has the full name. The the
- 22:29key basically looks completely
- 22:30different. full name, contact and then
- 22:33this is nested inside and then you have
- 22:36an email and phone number. So the format
- 22:39is completely different from the first
- 22:41one. Yeah. Now imagine you may end up
- 22:44having several of such formats within
- 22:46your data lake without schema
- 22:48enforcement. And when you're writing a
- 22:51query to basically let's say find um the
- 22:54customer email or the phone number, you
- 22:57may end up writing complex code in order
- 23:00to figure out what the right schema is,
- 23:03right? You may need to write complex
- 23:05logic in order to handle different
- 23:07fields and this is a very big issue
- 23:10because there's no consistency in your
- 23:12data. There's no schema in placement. So
- 23:15to quickly summarize the three problem
- 23:17that we discussed about data lake was
- 23:19number one lack of acid transaction
- 23:22support. Number two lack of support for
- 23:24update merge and deletes. Number three
- 23:27data quality and reliability issues for
- 23:30example no schema enforcement. Right? So
- 23:32Delta solves all of these problems quite
- 23:35beautifully. It just doesn't solve them.
- 23:37But it also comes in with a bunch of
- 23:40very interesting features like asset
- 23:42transactions, time travel, unified batch
- 23:45and streaming schema evolution and
- 23:47enforcement and it also helps you see
- 23:49the audit history and all of that.
- 23:51Right? So we going to be going through
- 23:53all of this in detail. So before we get
- 23:55into understanding how Delta solves all
- 23:58of these problems, right, the one that
- 24:00we just talked about, let's first get a
- 24:02flavor of Delta Lake. How does it look
- 24:05like? How does it operate? What are the
- 24:08kind of operation that we can perform?
- 24:10Right. Yeah. So, we going to be setting
- 24:12up a lab environment and for this we are
- 24:14going to be using Azure. Yeah. And we'll
- 24:16be setting up different kind of
- 24:18services. So, don't worry if you don't
- 24:20know anything about it. I'll walk you
- 24:21through the entire process. Yeah. So,
- 24:24first of all, we'll be creating uh a
- 24:27workspace using the Azure data bricks
- 24:30survey. Yeah. So first of all, we'll be
- 24:32creating a workspace
- 24:34and we need some place where we can
- 24:38store our databases, our tables and all
- 24:41of that, right? And for that we are
- 24:43going to be using the Unity catalog.
- 24:45Yeah. And for those of you who don't
- 24:46know what Unity catalog is, simple for
- 24:48now you can think of it as a place where
- 24:51all of your databases and tables will be
- 24:54stored. Now in order to store those
- 24:57tables and databases right you need some
- 24:59location you need some storage location
- 25:01and that is where the meta store comes
- 25:03into picture. So the meta store you can
- 25:06basically think of it as an object
- 25:07storage something like um S3 or ADLS.
- 25:12Yeah. So in this context because we are
- 25:15using Azure we'll be using ADLS and
- 25:19we'll be creating a storage account.
- 25:23We'll be creating a storage account
- 25:25wherein we will create a container.
- 25:29Yeah. So we'll be creating this
- 25:30container
- 25:32where all of our data will be stored and
- 25:34we'll name the container metas store.
- 25:36You can name anything but we'll just
- 25:37name it metas store. So this is the
- 25:39place where all of those tables and
- 25:42databases will be stored. But there
- 25:44needs to be a mechanism
- 25:47using which the unity catalog here can
- 25:50store data in the container meta store
- 25:52and that is where another service comes
- 25:55into picture which is called Azure
- 25:58connector sorry not Azure connector
- 26:00access connector for Azure data bricks
- 26:03yeah and what's basically going to
- 26:05happen is that ADLS is going to tell
- 26:08this
- 26:10this guy is that hey I'm going to
- 26:12authorize you
- 26:15to be able to access the data in the
- 26:18container meta store. Yeah. So you can
- 26:21now
- 26:23you can now simply access the data that
- 26:27has been stored over here. Yeah. So now
- 26:30when we want to access the data in the
- 26:32meta store either using the workspace or
- 26:35through the unity catalog we are simply
- 26:38going to assume the role of this guy.
- 26:43Yeah we are simply going to assume the
- 26:45role of this guy and in turn this person
- 26:48is going to help us access the meta
- 26:50store. So we finally end up creating
- 26:54the first service which is a datab
- 26:56bricks. Then we create a storage
- 26:59account. Then we create a container.
- 27:02After that we create the access
- 27:05connector. And finally this meta store
- 27:09needs to know a few things that linking
- 27:12needs to happen. Right? So this needs to
- 27:14know who is the person who can access
- 27:18who can help me with access. Right? And
- 27:19it is this guy the access connector for
- 27:23your data brick. So it needs to know the
- 27:24resource ID
- 27:26of the access connector and the location
- 27:30to which it needs access right
- 27:33which is going to be the meta store. So
- 27:36this is the path and this is going to be
- 27:38the meta store container path right. So
- 27:43number five we need to perform the
- 27:45linking right. So once all of these
- 27:47steps are done we basically set up our
- 27:49lab environment and then we can start
- 27:51working. Okay, so let's quickly go ahead
- 27:54and create the Azure datab bricks
- 27:56workspace. And for that I'm going to
- 27:58click on create and we going to create a
- 28:02new resource group because this is where
- 28:04our workspace the storage account the
- 28:07access connector all of them are going
- 28:09to reside. Yeah, I'm going to name it as
- 28:12datab bricks delta lab - rg and I'm
- 28:16going to follow a similar naming
- 28:18convention. This is WS. Uh the region is
- 28:21going to be Australia East because I
- 28:23have a lot of resources in other regions
- 28:25as well. I've created workspaces. So to
- 28:29make sure that they don't conflict, I'm
- 28:30going to choose Australia East. But you
- 28:32please go ahead and choose the region
- 28:34that is the closest to you. Yeah. So it
- 28:38is validating some stuff. So meanwhile
- 28:41we can go ahead and create the storage
- 28:44account.
- 28:47The storage account is again going to be
- 28:50in the same in the same resource group.
- 28:54The naming convention also we'll keep
- 28:57we'll keep it very similar. Delta lab
- 29:00storage account is going to be in
- 29:03Australia east. This is going to be ADLS
- 29:05gen 2. And for now we'll select the
- 29:08cheapest option which is locally
- 29:10redundant storage. And don't forget to
- 29:13put a check here because we want to use
- 29:15data lake storage gen 2. And we simply
- 29:19go ahead and click on review and create.
- 29:22Let's also click on create over here.
- 29:24And create over here. The last item that
- 29:27we had was the access connector for
- 29:32Azure datab bricks and this is the
- 29:35person who will be assigned who will be
- 29:39given the authority to be able to access
- 29:42data in the storage account.
- 29:44So
- 29:46let's select the same resource group
- 29:50DB delta lab - RG and this is going to
- 29:53be DB delta lab
- 29:56access connector. Yeah. And this is
- 29:59going to be in Australia east. So let's
- 30:01go ahead and create this.
- 30:09So we see now the deployment for the
- 30:11storage account is completed. So inside
- 30:14of the storage account, we need to
- 30:16create a container called metas store.
- 30:20Let's go ahead and quickly create that.
- 30:24So there you go. You have a container
- 30:26called metas store. And in this
- 30:28container, Unity catalog is going to
- 30:30store all of the data.
- 30:32Now the next part to this is we have to
- 30:36give permissions to the access connector
- 30:39to be able to access the data in my
- 30:41account in this storage account. Yeah.
- 30:44So what this guy is going to do is that
- 30:46it is going to give permissions now and
- 30:48the role that is going to that is going
- 30:51to allow is storage blob data
- 30:53contributor
- 30:55and because it's a manage identity we
- 30:57are simply going to select
- 30:59the access connector for Azure data
- 31:01bricks and this was the one that we just
- 31:03created right now and we simply click on
- 31:05review and assign.
- 31:11So now we see that the assignment is
- 31:13there in place. That means that the
- 31:15access connector will be able to access
- 31:18the data in the storage account.
- 31:23Now we are waiting for the deployment of
- 31:26the workspace to take place to complete.
- 31:29Right. Okay. So the workspace is now
- 31:32deployed. Now the Azure data bricks
- 31:34workspace is deployed. We created ADLS
- 31:37Gen 2. We created the metas store. We
- 31:41also created the access connector.
- 31:44Right? Now in order to enable unity
- 31:47catalog. Now we need to
- 31:50do the linking. We need to create the
- 31:53the meta store and then provide
- 31:56who is going to be the person who's
- 31:57going to help me with access and what is
- 32:00the path of the meta store. Yeah. So
- 32:03let's go ahead and do this last step
- 32:05right here. So let's quickly go ahead to
- 32:07the workspace. We are going to launch
- 32:10the workspace. And
- 32:17when we head over to the catalog, we see
- 32:19that
- 32:21it only has the legacy hive meta store
- 32:24and some shared samples over here. The
- 32:27Unity catalog hasn't yet been enabled.
- 32:30So in order to enable Unity catalog, we
- 32:32need to go to account.asure Azure datab
- 32:36bricks dot and we login with this. So
- 32:39now we'll be able to see that we have an
- 32:42option for the catalog and there is an
- 32:44option to create the meta store. This is
- 32:47going to be db delta lab meta store.
- 32:52This is going to be in Australia east.
- 32:55And the format of this is the container
- 32:59name at storage
- 33:00account.dfs.co.windows.net.
- 33:02So the container name is meta store.
- 33:06The storage account name is this one. So
- 33:10we simply do this
- 33:13dot windows
- 33:15dfs.core
- 33:18dotwind.net.
- 33:20Yeah. And then finally we need to put in
- 33:22the access connector ID. So the access
- 33:26connector we can find it over here. This
- 33:28was the one that we created. And then
- 33:32here is the resource ID. So we simply
- 33:34copy it from here and we are going to
- 33:36paste it over here and we click on
- 33:38create.
- 33:45So now it is asking me to which
- 33:49workspace do I want to assign this metas
- 33:52store and this is the workspace that we
- 33:53just created, right? So let's assign the
- 33:56meta store to this workspace.
- 34:00And now we have Unity catalog enabled.
- 34:02So let me quickly go ahead and refresh
- 34:04this.
- 34:07And you see that apart from the shared
- 34:09samples and the legacy hive meta store,
- 34:12we have two cataloges right now. So this
- 34:14means that a unity catalog is enabled
- 34:17and our lab environment is now set up.
- 34:19So let's go ahead and create a delta
- 34:21table and perform some operations. So
- 34:23we'll be coming back to this diagram but
- 34:25let me first go over to the workspace
- 34:29and then we going to create a folder
- 34:31called lab delta lake and here is where
- 34:36we are going to store all our notebooks.
- 34:38But for our labs we need data right. So
- 34:42we are going to go to the storage
- 34:44account that we created which is this
- 34:46one DB delta lab storage account and we
- 34:50are going to create a container right
- 34:52because we don't want to me we don't
- 34:54want to mess around with the metas store
- 34:56container because this is going to be
- 34:59used by the unity catalogs to store all
- 35:01the databases and tables right so let's
- 35:04go ahead and create something called lab
- 35:05data and this is going to create a new
- 35:10container
- 35:11Inside of this, I'm going to add a
- 35:14directory called shopping invoices
- 35:17or maybe just invoices.
- 35:20Let's keep it smaller and shorter. So,
- 35:23inside of invoices, I'm going to upload
- 35:26all of the files that I have. Right? So,
- 35:28I have three files
- 35:30and I'm going to upload all of them. So
- 35:33these are basically invoices for
- 35:34customer ids from 1 to 100, 100 to 200,
- 35:38201
- 35:39some some large number 99457 right so
- 35:42don't worry all of this will be made
- 35:45available in the GitHub repository so
- 35:47you can download it from there so now
- 35:50let's move ahead
- 35:53to our lab now an interesting thing is
- 35:57that
- 36:00okay so we are going to perform DML L
- 36:03operations on delta tables. Yeah. So
- 36:08this is our motive and we want to access
- 36:11the data that is over here inside of the
- 36:14lab data container. Right? Now this is
- 36:18not going to naturally have access.
- 36:21Right? So the way we access it is let's
- 36:23say percentage fs ls and abfss.
- 36:30This is going to be meta lab data at
- 36:35whatever this is. The storage account
- 36:37name is this one over here. So we copy
- 36:40this dfs.core.windows.net.
- 36:44Right. And I have this cluster already
- 36:48created. And let me also walk you
- 36:50through how to create the cluster over
- 36:52here. Right? It's quite simple. So you
- 36:55you click on create compute.
- 36:58go ahead and create a single node
- 36:59cluster because for this lab we don't
- 37:01need a multi-node cluster and you don't
- 37:03want to incur a lot of cost right so
- 37:05let's simply choose 14.3
- 37:08LTS we don't want photon accelization
- 37:12and let's choose a simpler one a uh
- 37:16something that are lesser memory than 16
- 37:18which is over here 14 GB and four cores
- 37:21and set this to 20 minutes you can even
- 37:24set it to 10 minutes because sometimes
- 37:26we just leave the cluster running and we
- 37:29incur a lot of cost right and then you
- 37:30can click on create compute so that's
- 37:33how you can create the compute and now
- 37:35let's come back here now let's say if I
- 37:37want to do an ls
- 37:41let's see what do I get okay so what it
- 37:44says is that invalid configuration value
- 37:47detected for
- 37:51fsazure
- 37:52account right so basically the crux of
- 37:55this is that it is not able able to
- 37:57access this data. Right? So that is
- 37:59where the concept of external location
- 38:02comes in. So external location are
- 38:04basically location that are external to
- 38:07data bricks that we want to access right
- 38:10and we need to put proper measures in
- 38:12place so that we are able to access that
- 38:14location and let's first see how we can
- 38:18access this external location. So we go
- 38:20to catalog.
- 38:22We then go to external data. We go to
- 38:25create an external location. By the way,
- 38:27you can also create it using a simple
- 38:31SQL statement. But let's go through the
- 38:33UI and see how this works out. So let's
- 38:36name the external location lab data
- 38:40external.
- 38:42And the way we do it is abs
- 38:50and then this is going to be lab data at
- 38:54the storage account named
- 38:56dfs.core.windows.net.
- 39:00Yeah. So I believe let's quickly check
- 39:02this uh the format abfs container name
- 39:05storage account dfs.core.windows.net.
- 39:07Yeah. So that's the path and then we
- 39:10need a storage credential. somebody
- 39:13whose role I can assume in order to be
- 39:16able to get access in order to be able
- 39:18to look at the data at that storage
- 39:21location. Yeah. Now when we created the
- 39:26access connector it automatically
- 39:28creates a storage credential and that is
- 39:31what we can use. So we simply go ahead
- 39:32and use that and we are going to create
- 39:35the external location. Now the external
- 39:38location is created. Now let's go ahead
- 39:41and run this command once again.
- 39:44So now you see that we are able to do an
- 39:48ls on that location. Right? So let me
- 39:51also quickly do an ls on the invoices
- 39:57and we are able to see all of the data
- 40:00that we have put on that location.
- 40:02Right? So that is external location.
- 40:03That is how external locations work.
- 40:06Let's go ahead and create a catalog.
- 40:09Right? So for those of you who haven't
- 40:11used Unity catalog, think of catalog as
- 40:14a highlevel container which is going to
- 40:17store your databases and the new tables.
- 40:20Right? So I'm going to write this SQL
- 40:23create
- 40:25catalog if not exist and this is going
- 40:28to be delta catalog. Right? And after
- 40:31this I'm going to create a schema and
- 40:33you can think of schema as a database.
- 40:36Right? So this is going to be delta
- 40:41dot delta db right. So let's go ahead
- 40:44and run this.
- 40:47So now here I should have a delta
- 40:50catalog which is right here and amongst
- 40:53the default and the information schemas
- 40:55I also have a delta db schema right and
- 40:58it doesn't have anything for now. So
- 41:00that is okay. Yeah. So we are going to
- 41:02quickly look at some of the data that
- 41:05we've stored over here and we're going
- 41:07to use this for creating and operating
- 41:10on our delta lake right so I'm going
- 41:12simply going to say select star from
- 41:15park k
- 41:18and this is going to be something like
- 41:20this and let me also do a limit five
- 41:24yeah so let's go ahead and run this so
- 41:26this is how our park file looks like
- 41:29it's basically simple invoices about a
- 41:31customer who made a purchase, what is
- 41:33the gender, age, payment method, what is
- 41:35the quantity, what is the invoice date
- 41:37and all of that. Right? So, let's go
- 41:39ahead and create a delta table out of
- 41:42this. Right? So, let's go ahead and
- 41:44simply write create or replace table and
- 41:48this is going to be named invoices,
- 41:52right? and
- 41:55as select
- 41:58star from. Okay, so this is helping me a
- 42:01lot.
- 42:03I just have to press enter more than
- 42:05doing the actual typing. So this is good
- 42:09and let's go ahead and run this, right?
- 42:16Okay, great. The table has been created.
- 42:18So let's have a look at the catalog over
- 42:20here. And now we see that it contains
- 42:23invoices and this invoice is basically
- 42:26your all of the data that we just
- 42:29discussed and it contains some
- 42:32interesting details over here right and
- 42:34we're going to come back to this soon
- 42:36but before that let's see the history
- 42:40and what do I mean by history is that
- 42:42what are the operation that was
- 42:44performed on this table.
- 42:52I was just happy about the about this uh
- 42:55about about the code appearing
- 42:57automatically and now it's not appearing
- 42:59automatically. Okay, so describe history
- 43:03cannot be found. Okay, I just misspelled
- 43:06this. So we just created the table using
- 43:11a cat cas statement, right? And that is
- 43:14what you see over here. So it basically
- 43:16tells you that there is a time stamp on
- 43:19this time stamp this particular user
- 43:21created this table and the operation
- 43:23that was performed would basically a
- 43:25create or replace table as select
- 43:28something like that right and then it
- 43:29has all other details as well. Yeah. So
- 43:33this basically tells us that version
- 43:35zero and this is where virgining comes
- 43:38in right. This form the foundation and
- 43:39the basis for virgining. And don't worry
- 43:41we'll we'll we have a complete section
- 43:43on time travel and virgining right so
- 43:45this basically forms the first version
- 43:48version zero of this now let's actually
- 43:51also have a look at what is happening
- 43:53behind the scenes right
- 43:56so we see that this data is stored in
- 44:00the container meta store
- 44:02in this storage account the storage
- 44:05account that we provided and then there
- 44:07is some folder which is named by this
- 44:09unique long unique unique ID and then
- 44:11there is a folder called table and this
- 44:14is the unique ID of a table right so
- 44:17let's quickly go over here to the
- 44:19storage account this meta store
- 44:24yeah and
- 44:27so now we go to tables and then I copy
- 44:31the ID which is this ID right over here
- 44:34and we see that there is a park file
- 44:37which is created along with a folder
- 44:40called data log. So it basically has the
- 44:42data and the metadata. The metadata
- 44:45about the transactions that took place,
- 44:47right? And that is stored in the delta
- 44:49log folder. So this delta log folder
- 44:52basically contains a JSON file which has
- 44:55details about the transactions and a CRC
- 44:57file. We can ignore the CRC file for now
- 44:59because it's for file valid validation
- 45:01and all of that. So we don't need to
- 45:02worry about all of that for now. So if
- 45:04we click on this, what do we see?
- 45:08So we see all of this and let me
- 45:11download this file right let me quickly
- 45:14download this file and
- 45:17okay and let me format this document. So
- 45:21it contains a bunch of stuff right and
- 45:23don't get intimidated by this. So the
- 45:26first section is commit info. The second
- 45:28section is metadata. The third section
- 45:30is protocol. And then the fourth section
- 45:33is the operations that we performed,
- 45:35right? And the most important one of
- 45:37this is the last section, the operation
- 45:39that we performed. But I'll still walk
- 45:41you through this. The commit info
- 45:42basically contains the operation that we
- 45:44performed, right? The user who performed
- 45:46this operation, right? And all of that.
- 45:49Um now coming over to the add part
- 45:53itself. It is basically the operation
- 45:55that we performed. So we did a create or
- 45:59replace table and this added some data
- 46:01to the table. Right. And that data was
- 46:04added in the park file. It then contains
- 46:07other details. So because our park is
- 46:10not partitioned, it doesn't have any
- 46:11partition value. This is the size. Uh
- 46:14these are the stats which basically
- 46:17contains uh what are the minimum values
- 46:20for the columns like customer ID um if
- 46:22there's a numeric column then it would
- 46:24be helpful something like price and all
- 46:25of that right and it also contains the
- 46:27max values the max values over here so
- 46:30this is very helpful when we want to do
- 46:34partition pruning let's say we are
- 46:35reading this file using spark and we
- 46:37want to do partition pruning so spark
- 46:40can basically filter down the files by
- 46:43reading the min and max right if it
- 46:45doesn't fall within these bound it can
- 46:47simply either include or not include
- 46:50this file so that is what this operation
- 46:54is for and it been noted down in the
- 46:57transaction log now let's go ahead and
- 47:00perform another operation
- 47:04let's go ahead and do an insert right so
- 47:07let's see how the insert statement is
- 47:09going to look like and it's basically
- 47:11going to look something like this insert
- 47:14into delta delta catalog or delta
- 47:17DB.invoices and we are simply going to
- 47:19have a select star from par
- 47:24dot
- 47:27and this one and this is basically going
- 47:30to be one to 100 because we have already
- 47:35inserted from 101 to 200. So let's go
- 47:38ahead and insert the data from 1 to 100.
- 47:41Right? So let's go ahead and run this.
- 47:47And we are also going to see the history
- 47:49again. So let me paste this over here.
- 47:53And let's also do some sanity check.
- 47:55Right? We are going to do a select main
- 48:00customer ID max
- 48:03customer ID and then the count
- 48:07count star. What are the total number of
- 48:09records? Right? And this is simply going
- 48:11to be from this table right. So now we
- 48:13see that the insert operation had
- 48:16succeeded. That means it inserted 100
- 48:18more records. Now let's go ahead and see
- 48:21how the history looks like. The first
- 48:23statement was a create or replace right
- 48:26and the second statement was a write
- 48:29through an insert statement over here.
- 48:31And this creates the second version
- 48:34number one. Right? It's zero index. So
- 48:36that's why the second version. Now let's
- 48:40go ahead and run this. The minimum and
- 48:42the maximum. So we are supposed to have
- 48:44200 record. That's why the minimum is
- 48:46one, maximum is 200. And the count the
- 48:49total number of record is 200. So that
- 48:52means that all of the insert and the
- 48:54create statement has happened correctly.
- 48:57So now let's go ahead and have a look at
- 49:00how the delta log looks like. So earlier
- 49:03there was one parket file. Now there are
- 49:06two park files. We're going to have a
- 49:08look at how these park files look like,
- 49:10right? And let me go ahead and download
- 49:12these two files.
- 49:14Uh so basically I copy this and I'm
- 49:19going to use park tools show and it
- 49:22resides in my downloads folder.
- 49:27So you see that the first park file
- 49:29wherein we inserted data from for
- 49:31customer ids from 101 to 200 that is
- 49:35what is contained over here in the first
- 49:38park file. In the second par file
- 49:42this should contain all of the records
- 49:44for customer ID from 1 to 100. So again
- 49:48let me do park tools show
- 49:53downloads
- 49:55and then the part right. So now you see
- 49:58that this is from customer ID 1 to
- 50:01customer ID 100. Yeah. And let's go
- 50:04ahead and have a look at the delta log.
- 50:07So we again will download
- 50:10the first delta log.
- 50:13I mean the second delta log index by 01
- 50:18and let's open it
- 50:21format this document
- 50:24and we see that this is the second paret
- 50:27file that we were talking about right
- 50:30the second paret file that's been added
- 50:33over here. Yeah this one. So we see that
- 50:36there are two transaction that took
- 50:38place. First one is the create table. It
- 50:40created a par file and then it created
- 50:44it noted that down in a JSON file,
- 50:47right? It noted that transaction down
- 50:49using this add operation. Another
- 50:52operation when that took place using the
- 50:55insert statement, it also created
- 50:57another park file, put in all of the
- 50:59data over there and then again noted it
- 51:01down as an insert operation in another
- 51:05JSON file. Right? So these are taking
- 51:07place at separate independent
- 51:09transaction. Yeah. Now let's try a few
- 51:12other interesting operations right. Uh
- 51:14let's try the update operation. So first
- 51:17of all let me do a select star from the
- 51:20table where customer
- 51:24ID equals 1. Yeah. And let's see how
- 51:28let's see what the quantity. So the
- 51:29quantity is five over here. Yeah. So let
- 51:32me go ahead and run an update statement.
- 51:35update the table and this is going to be
- 51:38set quantity 10 where customer ID equals
- 51:431. Yeah. So let's go ahead and quickly
- 51:46run this
- 51:50and let me also do a describe history of
- 51:53this table. So we should see an update.
- 51:57Great, we see an update and this is
- 51:59marked as the next version. Now let me
- 52:02quickly go ahead and see how
- 52:05this looks like. So the quantity has
- 52:07been updated to 10. That is good. That
- 52:10is what we would have expected. Now
- 52:12let's see what's going on behind the
- 52:14scenes.
- 52:17So behind the scenes
- 52:20we see that that there were earlier two
- 52:22files. Now we see that there are three
- 52:24files. Right? So this is the latest one
- 52:26that has been written at 643.
- 52:30And there is something called a deletion
- 52:33vector. Right? And this is quite
- 52:35interesting to understand. We'll we'll
- 52:37have a look at this in details. And
- 52:39let's also have a look at what this file
- 52:43means.
- 52:44The latest transaction.
- 52:48So the latest transaction. Let me format
- 52:51this document. And we see a lot of
- 52:54things happening, right? We see a
- 52:56remove, we see an add. And then we also
- 53:00see an ad, right? But the fact of the
- 53:02matter is I just updated one row. Why
- 53:06why did it even remove something? It
- 53:08removed one park file, it added another
- 53:12parket file and then it added another
- 53:14parket file again. Right? What were the
- 53:16need for all of this? So there is a
- 53:17little bit of concept behind this and
- 53:19and let's quickly walk through that.
- 53:21Right?
- 53:23So let's assume that there was this file
- 53:26called 1. R K and this basically
- 53:31contained all the records for customer
- 53:33ID 1 to 100. Right now whenever somebody
- 53:37wanted to read all the customers from 1
- 53:39to 100, they would simply go ahead and
- 53:41read this file now the situation has
- 53:44changed a little bit. Right? The way it
- 53:47has changed is that we ran we want to
- 53:49run an update statement and we want to
- 53:53update the quantity to 10 for customer
- 53:58ID.
- 54:00Customer ID equals 1. Yeah. So this is
- 54:03what we want to do. So what delta says
- 54:05is that now anybody want who wants to
- 54:08read all of the data for customer ID
- 54:11from 1 to 100, they just cannot go ahead
- 54:13and read this file anymore, right?
- 54:15because this has changed for customer id
- 54:19equals 1 the quantity has become 10 but
- 54:22this file contains quantity equals five.
- 54:25So what I'm going to do is that I am
- 54:28going to create two files right so we
- 54:31are going to create
- 54:33two files the first one is 1.park park
- 54:36is going to be as it is right and then
- 54:39we are going to create two.par
- 54:43and this is only going to contain the
- 54:45record customer
- 54:48ID
- 54:50equals 1 and it quantity is going to be
- 54:53equal to 10. Yeah. So it created another
- 54:57file with just the updated record and it
- 55:01is going to club 1.par park a with
- 55:05something called a deletion vector.
- 55:12Yeah. So it is going to plug this with
- 55:15something called a deletion vector. And
- 55:17the deletion vector basically tells that
- 55:21don't read
- 55:23the record where customer ID equals 1.
- 55:28Don't read the record where customer ID
- 55:30equals 1 because again that is going to
- 55:32be read from 2.par. So the way the delta
- 55:36transaction log is built the reason why
- 55:39you see basically what do you see? You
- 55:41see one remove. Yeah you see one remove
- 55:46over here
- 55:48and that same file is being added over
- 55:50here. Yeah the first one is a remove.
- 55:53The second one is an add and then this
- 55:56is the other file. So that is what you
- 55:57see over here. So now let's try to
- 56:00relate it with this. So now when we
- 56:04build the when delta builds the
- 56:05transaction log what it says is that
- 56:08this file 1.park earlier it was being
- 56:11treated as if we need to read the whole
- 56:14file. So now we have to treat it
- 56:16differently. You don't have to read the
- 56:18whole file. So that is why it adds a
- 56:21remove saying that 1.park par cannot be
- 56:25treated the same way it was being
- 56:27treated earlier. Yeah. And then it add
- 56:30then add. So the way it has to be
- 56:33treated now is 1.par plus the deletion
- 56:37vector. Yeah. So first it tells that
- 56:39don't treat 1.pk as it were being
- 56:42treated earlier. Yeah. Because it
- 56:44contains the row customer ID equals 1
- 56:46but the quantity is different. It's not
- 56:4810. So don't treat it like that. Treat
- 56:50it like this. So for that it adds an add
- 56:53statement with a deletion vector and
- 56:56then it adds an add again which
- 56:59basically means that read customer id
- 57:02equals 1. Yeah. So it makes it very
- 57:04beautiful. It let's quickly summarize
- 57:07this. So it adds a
- 57:09remove the remove is for 1.park
- 57:14simply saying don't read it as it were
- 57:16being read earlier. You were reading the
- 57:17whole file earlier. Don't do that. And
- 57:20then it adds an add one.park again
- 57:24simply saying that read it with the
- 57:28deletion vector. And then it adds an add
- 57:31again with 2.pk
- 57:34which only contains customer ID
- 57:38equals 1. So using all of this
- 57:42information it is able to construct the
- 57:45file. Yeah. So that's that's the beauty
- 57:48of this. Now let's go back and see what
- 57:50actually happened. So
- 57:53here you see this file
- 57:56this is without the deletion vector and
- 57:58this file is with the deletion vector
- 58:01over here. Yeah. And this file over here
- 58:05is the file which only contains customer
- 58:08ID equals 1. And let's verify that right
- 58:11let's verify that.
- 58:13So this is the file over here. Let me
- 58:16download it. We can do a parket tool
- 58:19show and
- 58:21download
- 58:23and then this park over here. So you see
- 58:25that it only contains customer ID equals
- 58:281 and quantity equal 10. So this is a
- 58:31very beautiful way of
- 58:34making the transaction log without
- 58:36actually doing the writing and the
- 58:39reading and all of that. Right? So when
- 58:42it added a remove, it actually didn't do
- 58:45a remove, right? It just noted it down
- 58:47and logically built up the steps to
- 58:50construct the file. Okay. So let's try a
- 58:52different operation now. Let's try
- 58:54delete. Yeah. So let's do
- 59:00something. So first of all uh let me do
- 59:04something like a
- 59:07where customer ID equals let's say 99.
- 59:10Right? Um, so now
- 59:15I just want to see how the record looks
- 59:17like
- 59:19and the delete statement would look
- 59:21something like delete from this
- 59:25where customer ID equals this, right?
- 59:27Pretty simple. So this is my customer ID
- 59:30equals 99. Let's go ahead and run this.
- 59:32Uh, let me also quickly show you.
- 59:36We have three parket files, one deletion
- 59:38vector, and three JSON files, right? So,
- 59:43let's go ahead and run this now.
- 59:47If I were to run this, I should get an
- 59:50error now. Yeah, as expected.
- 59:55Yeah, I mean, we should get an empty
- 59:58empty results, right? Instead of an
- 1:00:00error. So now we do a describe history
- 1:00:06of this table and I should see a
- 1:00:08statement for delete.
- 1:00:11Great. So now we see an operation for
- 1:00:13delete and this becomes the most recent
- 1:00:16version. So now if I refresh this I
- 1:00:20still have three parket files and I have
- 1:00:23just added a deletion vector. Right? So
- 1:00:27there have been no changes with respect
- 1:00:28to the park file. And let's go ahead and
- 1:00:31see the
- 1:00:33JSON log again. Let me download it. And
- 1:00:37let's go ahead and open this in Visual
- 1:00:39Studio
- 1:00:41uh format document. And what has
- 1:00:44happened?
- 1:00:46So what has happened is there is a
- 1:00:50remove and then there is an add. Yeah,
- 1:00:55quite interesting. So there is a remove
- 1:00:58for a file which already has a deletion
- 1:01:01vector and now there is an add. Yeah. So
- 1:01:06let's quickly go ahead and try to
- 1:01:07understand what's going on. Yeah. So
- 1:01:09last time we were talking about 1.park,
- 1:01:14right? So we were talking about
- 1:01:191.park park file
- 1:01:23and this basically contained all of the
- 1:01:27customer ids from 1 to 100. Right? Now,
- 1:01:30if you remember this was updated with a
- 1:01:33deletion vector. So, 1 dot park file a
- 1:01:38deletion vector was added and this was
- 1:01:42for customer
- 1:01:44ID equals 1. And this was basically to
- 1:01:48tell that hey don't read par file the
- 1:01:521.parket park file as it is right read
- 1:01:55it with the deletion vector because the
- 1:01:58customer ID number one has been updated
- 1:02:01the quantity was earlier five now it's
- 1:02:03updated to 10 and so for that reason you
- 1:02:06shouldn't read it like this you should
- 1:02:08read it like this with a deletion vector
- 1:02:10so that is what was the state so that is
- 1:02:14the reason why you see the remove with a
- 1:02:19deletion vector already and the add also
- 1:02:22has a deletion vector. So these both are
- 1:02:25basically the same files. They are
- 1:02:28basically the same files and both of
- 1:02:31them have deletion vector. It just that
- 1:02:33the deletion vector has been updated.
- 1:02:36Yeah. So basically now what happens is
- 1:02:40we are saying that don't read one.park
- 1:02:44in this condition over here. Read it in
- 1:02:47a new condition because a delete
- 1:02:49operation has happened. Something new
- 1:02:50has happened, right? So a delete has
- 1:02:53happened.
- 1:02:56Yeah. So a delete has happened and in
- 1:02:58this delete
- 1:03:01this park file is supposed to be now
- 1:03:04read with a deletion vector and this is
- 1:03:07going to contain
- 1:03:09details about
- 1:03:12both customer ID equals 1 and customer
- 1:03:18ID equals 99. Yeah. So that is why there
- 1:03:22was a remove
- 1:03:24basically remove over here.
- 1:03:28Basically meaning that don't treat
- 1:03:31one.park as what was defined over here.
- 1:03:34Treat it as this one. That is why an ad
- 1:03:38is added over here. Yeah. So that is the
- 1:03:41reason why you see deletion vector in
- 1:03:43both the places. It's just that the
- 1:03:45deletion vector added over here is
- 1:03:48updated with more detail. Now let's
- 1:03:50perform the last operation basically to
- 1:03:54use one of the park file that I've
- 1:03:55created. Right? So insert into this
- 1:03:59invoices table and we do a select star
- 1:04:02from park
- 1:04:04dot and we will select this last file
- 1:04:08over here. Yeah. So let's go ahead and
- 1:04:12insert this file. Insert the data that
- 1:04:15we have here. And let me do a describe
- 1:04:18history.
- 1:04:22Yeah. So let's go ahead and do that. And
- 1:04:25we see that after a delete there was a
- 1:04:28right. And let's again see how this is
- 1:04:32going to look like. So there's going to
- 1:04:33be one JSON added.
- 1:04:36And there you go. There is it as we
- 1:04:38expect. And this should ideally just
- 1:04:41contain one ad.
- 1:04:44Yeah.
- 1:04:47So it just contains one ad. We can close
- 1:04:51the commit info and it just contains one
- 1:04:53ad. Yeah. And
- 1:04:56if we go here,
- 1:04:58there is one file that has been added
- 1:05:00over here which basically contains the
- 1:05:02most recent data. So that is another
- 1:05:06example. I I believe we've covered all
- 1:05:08of the operation from creating to
- 1:05:11inserting, updating and deleting. Yeah.
- 1:05:14So to quickly summarize the key concept
- 1:05:17behind Delta Lake is the use of a
- 1:05:20transaction layer on top of your data.
- 1:05:22Right? So your data is in the form of
- 1:05:24park files and there is a transaction
- 1:05:27layer called the delta log. Right? And
- 1:05:30whenever it performs a transaction it
- 1:05:33writes all of the data in a JSON file.
- 1:05:37And this transaction in a delta lake is
- 1:05:40an atomic unit of work. Right? It
- 1:05:42basically groups one or more operations
- 1:05:45together and these operations could be
- 1:05:48let's say an insert, an update or a
- 1:05:50delete. And the beauty of this is that
- 1:05:52these transaction either complete
- 1:05:55successfully or fail altogether. Yeah.
- 1:05:59And you see that this brings the gist of
- 1:06:02atomic in all of the asset properties.
- 1:06:05And that is how atomicity is
- 1:06:07implemented. So you may have this
- 1:06:09question lingering in your mind that all
- 1:06:12of this is good. The data is stored in
- 1:06:14parquet files and there is a
- 1:06:16transactional layer on top of it which
- 1:06:18stores data in the form of which stores
- 1:06:20transactions in the form of JSON files.
- 1:06:22Right? But let's say when I do something
- 1:06:25like a select star from this table,
- 1:06:30how is the latest state of the table
- 1:06:33that I get to see over here constructed,
- 1:06:36right? With all of those park files
- 1:06:38hanging here and there, right? With all
- 1:06:40of those JSON files, right? How are all
- 1:06:45of those used to construct this latest
- 1:06:47state? Yeah. So let's quickly go ahead
- 1:06:50and try to understand what happens
- 1:06:52behind the scenes and this is where I'm
- 1:06:55going to finally use this diagram. So
- 1:06:58basically we are going to have two
- 1:07:02things as we discussed. This is going to
- 1:07:04be our data and this is one example
- 1:07:06where I've shown partition data but this
- 1:07:08can simply be your park file and then
- 1:07:11this is going to be the delta log right
- 1:07:15and these are your transactions.
- 1:07:18These are your transaction. We'll come
- 1:07:19to the checkpoint and the last
- 1:07:21checkpoint by a little later in the
- 1:07:24course. So don't worry about it. Be
- 1:07:27relaxed. So let's take our example, the
- 1:07:30example that we went through.
- 1:07:33So the first statement that we ran was a
- 1:07:36cat, right?
- 1:07:39So it was a create table as select,
- 1:07:43right? And this led to the creation of
- 1:07:46the first transaction which is 0.json.
- 1:07:51And this created a park file which was
- 1:07:531. Park. Yeah. So we are just renaming
- 1:07:56it to 1.park for simplicity. And this
- 1:08:00contained all of the customers from 101
- 1:08:02to 200.
- 1:08:04The second operation that we ran was an
- 1:08:06insert.
- 1:08:09And this again led to the creation of
- 1:08:11another transaction which was 1.json.
- 1:08:16And this created another parket file
- 1:08:20which was 2.
- 1:08:22And it contained all of the customers
- 1:08:25from 1 to 100.
- 1:08:28The third operation that we ran was an
- 1:08:32update,
- 1:08:36right? and we updated customer ID equals
- 1:08:411. Yeah.
- 1:08:44So this again led to the creation of
- 1:08:47another transaction recorded in 2.json
- 1:08:51and we've already seen earlier that
- 1:08:54customer ID equals 1 was present in
- 1:08:572.pk. Right? So 2.park Park was removed
- 1:09:02because we wanted to treat it
- 1:09:04differently and
- 1:09:082.par was added again
- 1:09:12with a deletion vector containing
- 1:09:14details about handling customer ID
- 1:09:17equals 1. Right? Basically saying that
- 1:09:19don't read customer ID equals 1 from
- 1:09:21this park. And another parket file was
- 1:09:25added
- 1:09:27which only contained details about
- 1:09:29customer ID equal to 1. That was
- 1:09:30three.park. Yeah. Now the next operation
- 1:09:34that we did was a delete
- 1:09:40and this led to the creation of
- 1:09:43three.json JSON
- 1:09:45and what we did was that we deleted
- 1:09:49customer ID
- 1:09:53equals 99. Yeah. And this 99 is supposed
- 1:09:56to reside again in this 2.par. Right. So
- 1:09:59we again need to change the way we deal
- 1:10:01with two. So what we do is that we
- 1:10:06remove
- 1:10:08two.par and we add another we remove
- 1:10:12two. Okay, with this deletion vector and
- 1:10:16we add another 2 dotpar
- 1:10:21with a different deletion vector. Let's
- 1:10:23say this was DV1
- 1:10:26and this is DV2 now. Yeah. So that is
- 1:10:28what happened as a part of the delete
- 1:10:31operation. Two.park now contains details
- 1:10:34about handling customer ID equals 1 and
- 1:10:37customer ID equals 99. Right? It
- 1:10:39basically saying that don't read
- 1:10:41customer ID equals 1 and 99 from this
- 1:10:44park file. Now the last operation was an
- 1:10:48insert operation and this again happened
- 1:10:52in 4.json. The transaction happened
- 1:10:54there and this simply
- 1:10:58created another parket file. Right? So
- 1:11:01this simply created 4.par park and it
- 1:11:05contained customers from 2011 until some
- 1:11:0999,000 some big number right now the
- 1:11:12question is okay we had all these
- 1:11:13transactions now how do I compute the
- 1:11:16latest state right and the very simple
- 1:11:19answer to this is just do a summation
- 1:11:22yeah how do you do the summation the way
- 1:11:25you do the summation is basically
- 1:11:30this was added and this was removed so
- 1:11:32you remove both of this. Right? So this
- 1:11:34is a plus and this is a minus. So these
- 1:11:36two get removed.
- 1:11:38This was added. This was removed. So
- 1:11:41basically this gets removed. Now the
- 1:11:43only thing that stays until now is this
- 1:11:47one which is 1.pk.
- 1:11:52This one 3.
- 1:11:57This one 2.
- 1:12:02with DV2 and then 4.par.
- 1:12:10Right? So whenever we want to read the
- 1:12:13file and get the latest state, we
- 1:12:15basically read in 1.pk from where we get
- 1:12:18all customer ids from 101 to 200. We
- 1:12:21read in 3.pk where we get in customer ID
- 1:12:24equals 1. We read in four.park park
- 1:12:26where we get 2019,000
- 1:12:30whatever number it was and then we get
- 1:12:32we read 2.par wherein we read customer
- 1:12:36ids from 1 to 100 excluding customer ids
- 1:12:411 and 99.
- 1:12:44Yeah, because one is read from here and
- 1:12:4799 was deleted. So that is the way that
- 1:12:50is how simple it makes for processing
- 1:12:53the transaction log and you see how
- 1:12:55beautifully delta log creates these
- 1:12:58transactions and basically puts them in
- 1:13:00a way that is very efficient right there
- 1:13:02is lot of remove addition and all of
- 1:13:04that happening but that actually makes
- 1:13:07it very efficient because there is no
- 1:13:09actual writing happening right it is
- 1:13:11basically maintaining all of these
- 1:13:13states in a ledger. So another
- 1:13:16interesting question that you may have
- 1:13:17in mind is that Delta creates a JSON
- 1:13:22file for every transaction right and
- 1:13:24there could be companies who have
- 1:13:26millions of transactions in a day right
- 1:13:29and they could have around thousands to
- 1:13:33five of thousands of tables right and I
- 1:13:36don't even want to go ahead and multiply
- 1:13:38five thousands of tables with millions
- 1:13:41of JSONs being created every day right
- 1:13:43so this requires a lot of compute power
- 1:13:47in order to read the JSON file. So the
- 1:13:50question is that how does delta scale
- 1:13:55how does it handle reading massive
- 1:13:57amount of JSON files, right? So let's go
- 1:14:01ahead and simulate such a kind of a
- 1:14:03situation. Now, of course, we won't be
- 1:14:05creating millions or even hundreds of
- 1:14:08JSON files, but um we'll be creating a
- 1:14:10few of them to get a rough
- 1:14:13understanding, right? So, let's pick up
- 1:14:16this insert statement. And I'm going to
- 1:14:18run this in a loop and let me put
- 1:14:22a let me make this a Python cell. And
- 1:14:24let's do for i in range 10
- 1:14:28and or maybe let me put 50, right? And
- 1:14:33let's go ahead and run spark SQL.
- 1:14:38And let's paste this over here.
- 1:14:41So we are going to insert this spark a
- 1:14:4450 time, right? So this should add 50
- 1:14:47transactions to the table. And let me go
- 1:14:50ahead and print
- 1:14:53if
- 1:14:55insert
- 1:14:57completed. And let's go ahead and run
- 1:14:59this.
- 1:15:09So now all of the 50 inserts are
- 1:15:12complete. So let's quickly have a look
- 1:15:14at how this looks behind the scenes. And
- 1:15:17we see a lot of park files which is
- 1:15:20expected because we added in a lot of
- 1:15:22data. Let's also have a look at the
- 1:15:24delta log. And we see something
- 1:15:26interesting which is the presence of a
- 1:15:29compacted JSON. Right? Let me load all
- 1:15:33of this. So what we see here is that a
- 1:15:37compacted JSON is from the 1st to the
- 1:15:416th, right? The 7th to the 12th, the
- 1:15:4513th to the 18th. That means a compacted
- 1:15:48JSON is being created for every six JSON
- 1:15:51files, right? So I believe it is
- 1:15:53condensing all of the information.
- 1:15:55Actually not condensing but just picking
- 1:15:58up all of the information in each of
- 1:16:00those JSON files and then dumping it
- 1:16:03over here so that it doesn't have to go
- 1:16:05through the overhead of opening and
- 1:16:07closing those files. Right? And
- 1:16:12finally we see that after 36 JSON files
- 1:16:16have been created there is something
- 1:16:17called a checkpoint.park.
- 1:16:20Yeah. So let's understand what all of
- 1:16:23these are exactly. So I'll refer this
- 1:16:26diagram
- 1:16:28that I created. So over here we see a
- 1:16:33situation which is similar to what we
- 1:16:35just saw. Right?
- 1:16:37What this basically means is that from
- 1:16:40here to here all of this data is
- 1:16:43compacted and it is present in this
- 1:16:46docompacted.json
- 1:16:47file. Yeah. Um what this also
- 1:16:51essentially means is that there is no
- 1:16:52summation happening. Remember we
- 1:16:54discussed about summation. Yeah. And
- 1:16:56that is how delta computes the latest
- 1:16:58state. But that is not what is happening
- 1:17:00here. It just dumps all of the data that
- 1:17:03it had. So let's say this had
- 1:17:06let's say add 1.park
- 1:17:10and it has add 2. Okay. And this
- 1:17:14transaction over here, it had a remove
- 1:17:18one.
- 1:17:20If you were to sum this, then probably
- 1:17:221.par over here and 1.pk remove over
- 1:17:26here, it should simply go away, right?
- 1:17:28It should only end up having two park.
- 1:17:30But that's not the case. It basically
- 1:17:32picks up all of the data and dumps it
- 1:17:34over here. So, it's going to contain add
- 1:17:36of
- 1:17:381.par,
- 1:17:39add of two.
- 1:17:43and then the remove of
- 1:17:461.park park and of course all of the
- 1:17:49transactions and the details all of the
- 1:17:51details that is contained in these JSON
- 1:17:54files right so they are all added and
- 1:17:56they are put over here similarly for
- 1:17:58this one as well for all of the details
- 1:18:02that is present over here they are added
- 1:18:03over here now imagine you just committed
- 1:18:06a transaction and
- 1:18:08you are over here
- 1:18:11you just performed this operation so now
- 1:18:14how is delta going to compute the latest
- 1:18:15state it is basically going to read it
- 1:18:19is going to read this file. It is going
- 1:18:22to read this file and then these
- 1:18:26these two files over here. That is how
- 1:18:28it is going to compute the latest stage.
- 1:18:31Now imagine this is a lot beneficial
- 1:18:33because it doesn't have to read all of
- 1:18:35this, right? It doesn't end up reading
- 1:18:38all of this.
- 1:18:40So that is how this is optimized for
- 1:18:43performance. Now after 36 files have
- 1:18:47been created you see something called a
- 1:18:49checkpoint.park
- 1:18:52and what it does is that it again
- 1:18:55condenses all of the information from
- 1:18:59the zero to the 35th
- 1:19:02file and puts it over here. Yeah. So
- 1:19:05let's say you were at this transaction.
- 1:19:08The only file that you would have to
- 1:19:09read is that it it tries to find out do
- 1:19:12we have a checkpoint.park file or not.
- 1:19:15Right? If it doesn't have then it is
- 1:19:17going to follow the normal process the
- 1:19:19one that we just discussed earlier. If
- 1:19:21it has a checkpoint.paret file find the
- 1:19:24latest checkpoint.pket file. So that
- 1:19:27latest checkpoint.park file is going to
- 1:19:29be this one. and then it is going to
- 1:19:31apply it is going to read these two JSON
- 1:19:33file and then it is going to produce the
- 1:19:36final state. So that means that you
- 1:19:40don't end up reading all of this.
- 1:19:44You don't end up reading all of this and
- 1:19:47that is tremendous amount of
- 1:19:50optimization right so that is how delta
- 1:19:53handles massive amount of JSON files
- 1:19:56right simply using compacted JSONs and
- 1:19:59by using checkpoint.pk park. Another
- 1:20:02very important point to note is that we
- 1:20:04see that the gap right the gap at which
- 1:20:07this checkpoint.parquet file is created
- 1:20:09is 36.
- 1:20:12If you were to use the open source
- 1:20:13version, you would probably see this
- 1:20:15file being created at number 10. Right?
- 1:20:18After 10 JSON files are created, a
- 1:20:20checkpoint.parkey files are pres is
- 1:20:23created. Right? Now the reason for this
- 1:20:25difference why 36 on data bricks and 10
- 1:20:29in the open source version is simply
- 1:20:30because datab bricks has made several
- 1:20:34performance optimizations in place so
- 1:20:36that it can relax the gap and it can
- 1:20:38create these checkpoint files at a very
- 1:20:42relaxed gap right so that's one of the
- 1:20:44most important points and it also varies
- 1:20:48per cloud provider so in AWS this number
- 1:20:52would be 36 six maybe in Azure
- 1:20:56it is going to be some 100ish
- 1:20:59right
- 1:21:01some 100ish so I believe this gives you
- 1:21:05an overall idea about how delta scales
- 1:21:08and handles massive amount of metadata
- 1:21:12and JSON files right earlier we talked
- 1:21:14about how delta solves the isolation
- 1:21:18problem the I in acid and this simply
- 1:21:21means that it allows for several
- 1:21:23transactions to operate concurrently to
- 1:21:26happen at the same time without one
- 1:21:28transaction worrying about what
- 1:21:30happening in the other transaction.
- 1:21:32Yeah. So it can remain carefree and it
- 1:21:35can do whatever it wants without
- 1:21:37worrying about what the other
- 1:21:39transaction is doing. Right. And it
- 1:21:41achieved this through something very
- 1:21:43beautiful called optimistic concurrency
- 1:21:46control. But there's before before we go
- 1:21:48ahead and understand what optimistic
- 1:21:50concurrency control is, there's
- 1:21:52something called pessimistic concurrency
- 1:21:54control. And these two guys go hand in
- 1:21:56hand, right? So in order to truly
- 1:21:58appreciate what OC optimistic
- 1:22:01concurrency control is, let's first
- 1:22:03understand what pessimistic concurrency
- 1:22:06control is. And in in that sense, we'll
- 1:22:09be able to understand why OC was a
- 1:22:12better choice. Right? So with
- 1:22:15pessimistic concurrency control, the
- 1:22:16DBMS, the database management system
- 1:22:19assumes that conflicts are bound to
- 1:22:21happen, right? Conflicts are something
- 1:22:23that will eventually happen and they are
- 1:22:26likely to occur and if it's likely to
- 1:22:28occur, it will occur. Yeah. So to to to
- 1:22:31make sure that you understand what a
- 1:22:33conflict is, conflict basically is a
- 1:22:35situation where two or more transactions
- 1:22:38are trying to modify the same data at
- 1:22:42the same time resulting in an invalid
- 1:22:45state, right? Bringing the database in
- 1:22:47an invalid state. So we saw several of
- 1:22:51these examples earlier, right? So if I
- 1:22:53were to go back to the consistency
- 1:22:56example that we took.
- 1:22:59So in the consistency example we saw
- 1:23:01that there were two transactions which
- 1:23:03were trying to update the account
- 1:23:06balance at the same time and the one
- 1:23:08which executed last is going to be the
- 1:23:10final balance and this is going to
- 1:23:12result in an inconsistent state. Right?
- 1:23:14So with pessimistic locking what we do
- 1:23:17is that we whenever a transaction is
- 1:23:20trying to make a change to a database we
- 1:23:24basically lock that data so that another
- 1:23:26person cannot come in and it cannot make
- 1:23:29changes right and let's understand that
- 1:23:31with an example so let's assume that we
- 1:23:33have two transactions the first one is
- 1:23:35T1 this guy wants to withdraw $100 from
- 1:23:39this account and there's another
- 1:23:41transaction called T2
- 1:23:43this guy wants to deposit $400 into the
- 1:23:47account. Yeah. And let's assume that T1
- 1:23:49is the one which starts first. So T1
- 1:23:52basically starts first. So the moment it
- 1:23:54starts,
- 1:23:56there is an exclusive lock which is
- 1:23:59placed on this row in the database.
- 1:24:02Right? So there is an exclusive lock
- 1:24:04which is acquired by T1 and it is placed
- 1:24:09on this row in the database. So now T1
- 1:24:12goes ahead and it basically reads in the
- 1:24:14balance. It reads in $1,000. Yeah. And
- 1:24:19because an exclusive log is acquired by
- 1:24:22T1, T2 cannot go ahead and make changes.
- 1:24:25So basically this will be paused. This
- 1:24:28will basically wait until T1 has
- 1:24:30completed. So it basically reads in T1
- 1:24:32basically reads in $1,000. it goes ahead
- 1:24:35and withdraws $100 and then
- 1:24:41it computes the account balance to be
- 1:24:43$900. Yeah. Now the moment it computes
- 1:24:46this once that is done it simply goes
- 1:24:48ahead and makes an update.
- 1:24:52So whatever account ID this is a 45120
- 1:24:56and this is going to update the account
- 1:24:59balance to $900. Yeah. Now the moment
- 1:25:02this is done this lock is released.
- 1:25:07Now the moment this log is released T2
- 1:25:10gets a chance and T2 then places an
- 1:25:14exclusive lock. So T2 is now going to
- 1:25:17place an exclusive lock on this row in
- 1:25:21the database. That simply means that
- 1:25:23after T2 has completed executing then
- 1:25:26only another transaction can come in and
- 1:25:28update this row. So again, it basically
- 1:25:31what it does is it simply reads in the
- 1:25:34account balance which is $900 and it
- 1:25:38adds in $400 which is $1,300. So it
- 1:25:42basically goes over here and update the
- 1:25:46account balance and this is going to be
- 1:25:48$1,300. Now after this is completed, it
- 1:25:52simply releases this lock.
- 1:25:56Now this lock is released and the final
- 1:25:59balance is $1,300.
- 1:26:02So this is correct, right? There's no
- 1:26:04problem with this. And what we saw
- 1:26:06earlier, if you remember in the
- 1:26:08consistency example,
- 1:26:11we also started a transaction over here,
- 1:26:14right? We started this transaction and
- 1:26:17we acquired a lock and then only all of
- 1:26:20this went ahead and computed the state
- 1:26:23due to which the other transaction the
- 1:26:25transaction two was not able to go
- 1:26:27inside right so we acquired a
- 1:26:29pessimistic lock now this all of this is
- 1:26:32really good right there's no problem
- 1:26:33with this but imagine if locking were
- 1:26:36not in place if locking was not in place
- 1:26:39what would have happened both the
- 1:26:42transactions T1 and T2 they would have
- 1:26:44come in, they would have read $1,000
- 1:26:47and this guy would have subtracted $100.
- 1:26:50It would have computed $900. This would
- 1:26:52have added $400. This would have
- 1:26:54computed 1,400. And whichever was the
- 1:26:58one updating this last would be the
- 1:27:02final balance of the account. Now, of
- 1:27:04course, there's a chance that you could
- 1:27:05have more money, but there's also a good
- 1:27:08chance that you'll end up having less
- 1:27:09money. Yeah. So pessimistic locking
- 1:27:13helps us achieve consistency. It
- 1:27:15maintain the correctness of the
- 1:27:17database. So what are the problems? Why
- 1:27:20don't we want to use pessimistic
- 1:27:21concurrency control? Now you see that T2
- 1:27:24had to wait.
- 1:27:27It had to wait till T1 was completed.
- 1:27:32And imagine if you have millions of
- 1:27:35transactions taking place then this
- 1:27:38whole system is going to come to a halt
- 1:27:40right this is not going to be meaningful
- 1:27:42anymore because you would have to wait
- 1:27:44endlessly in order for few of the
- 1:27:47transactions to complete and then I
- 1:27:49would get a chance right so that is
- 1:27:51where the problem comes in it makes your
- 1:27:53system slow and for this reason we want
- 1:27:57to use optimistic concurrency control
- 1:27:59with optimistic concurrency control
- 1:28:02transactions do not obtain locks when
- 1:28:05they read or write and that's the
- 1:28:07beautiful part right so the name
- 1:28:09optimistic actually comes from the fact
- 1:28:12that it assumes that conflicts are very
- 1:28:16unlikely to occur and this is quite just
- 1:28:18the opposite of what pessimistic
- 1:28:20concurrency control assumes right so
- 1:28:22this guy assumed that conflicts are very
- 1:28:25unlikely to occur so I don't need to
- 1:28:27think about it right and if at all it
- 1:28:29occurs the conflicting transaction is
- 1:28:31the guy who going to be retrying, right?
- 1:28:34So let's understand this with an
- 1:28:36example. So let's take the same two
- 1:28:38transactions, right? T1 basically trying
- 1:28:41to withdraw $100 and T2 trying to
- 1:28:44deposit $400.
- 1:28:47Yeah. And these two transactions go
- 1:28:50ahead and read in this row at the same
- 1:28:54time. So T1 basically reads in $1,000 to
- 1:28:58be the current balance. T2 reads in
- 1:29:02$1,000 to be the current balance. And
- 1:29:05the reason why they are able to read
- 1:29:07this row at the same time is because as
- 1:29:10I discussed earlier, there is no concept
- 1:29:12of logs in optimistic concurrency
- 1:29:15control. Right? So because there is no
- 1:29:17concept of logs, this row over here is
- 1:29:21not logged by either of these
- 1:29:23transaction. they both can go in and
- 1:29:25read in the data over here at the same
- 1:29:28point in time. Right? So they go ahead
- 1:29:31and read the balance and along with that
- 1:29:32they also read in something called
- 1:29:34version number and time frame. The
- 1:29:36version number is going to help track
- 1:29:38changes within a table. So it reads in
- 1:29:41version number one and then it reads in
- 1:29:43something called TS1. Let's call this
- 1:29:45TS1.
- 1:29:46So this is going to read in one and TS1.
- 1:29:50Now it goes ahead and does all of the
- 1:29:52calculations, right? So this is going to
- 1:29:55compute do a subtraction of $100 and it
- 1:29:58is going to compute the final value
- 1:30:01which is $900, right? And this is going
- 1:30:04to add $400 and it is going to compute
- 1:30:07the final balance which is $1,400. Now
- 1:30:10both of these people are going to try to
- 1:30:14commit.
- 1:30:16Both of them are going to try to commit.
- 1:30:20Yeah.
- 1:30:22Now we need to understand that there is
- 1:30:24going to be one person who is going to
- 1:30:27commit first. Yeah. So let's assume that
- 1:30:29T1 is the person who commits first and
- 1:30:32there is a procedure there is a way in
- 1:30:35which the commit is going to happen. So
- 1:30:37the way it happen is it is going to
- 1:30:40check what is the current version
- 1:30:42number. It reads in that the current
- 1:30:43version number is one and it sees that
- 1:30:45the version number it has is also one.
- 1:30:50Yeah. So that means that I am working on
- 1:30:53the latest data. There is nobody who
- 1:30:55came in while I was working and made
- 1:30:58changes to the table. Right? So I can
- 1:30:59simply go ahead and I can make the
- 1:31:03addition update the balance right
- 1:31:05because I was working on the latest
- 1:31:06data. So now what happens is that it
- 1:31:09goes ahead and add this row which is
- 1:31:11A45120.
- 1:31:12This is going to be 900 and it updates
- 1:31:15the version to version two and this is
- 1:31:17going to be TS2.
- 1:31:19Now this person
- 1:31:22transaction T2 who was also trying to
- 1:31:24commit at the same time but actually
- 1:31:25ended up committing second right after
- 1:31:28this happened what it sees is that it
- 1:31:31reached the version number the version
- 1:31:33number that it finds is two but the
- 1:31:35version number it has is one. So this
- 1:31:39simply helps it understand that there is
- 1:31:42a person who came before me before I was
- 1:31:44trying to commit my transaction and
- 1:31:46actually made changes and this means
- 1:31:49that I was working on stale data. I need
- 1:31:51to read in the latest data make the
- 1:31:53changes and then actually update the
- 1:31:56table or add a row to the table. Right?
- 1:31:59So what it does is that it will simply
- 1:32:02fail this transaction
- 1:32:05because it was not working on the latest
- 1:32:07data. Yeah. So now what happens is
- 1:32:12T2 now reads in the latest version of
- 1:32:15the table. So it reads in the balance
- 1:32:17900. The latest version is two. And it
- 1:32:20reads in the time stamp ts2. It goes
- 1:32:22ahead applies the operation which is
- 1:32:24$400. Computes it to,300.
- 1:32:28And now it goes to commit again.
- 1:32:31So it goes and commits again. And it
- 1:32:34follows the same procedure. So it checks
- 1:32:36what is the latest version. The latest
- 1:32:38version appears to be two and the
- 1:32:40version it has read is also two. That
- 1:32:42means there is nobody who came in while
- 1:32:44I was working and make changes to the
- 1:32:46table. Right? So that also mean that I
- 1:32:48was working on the latest data. So I can
- 1:32:50make updates make a row addition to the
- 1:32:52table. So that is what it does now. And
- 1:32:55let me first just remove all of this.
- 1:32:59So it goes ahead and it makes an update.
- 1:33:05to this table. It adds in a new row. It
- 1:33:07updates the time stamp and this is
- 1:33:09called TS3. And this is how
- 1:33:12the table ends up in the correct state.
- 1:33:16Yeah. Now imagine that T2 was still
- 1:33:20working and there was a transaction
- 1:33:22called T3. It wanted to read the latest
- 1:33:25state of the table. And let's imagine
- 1:33:26that this row was not yet committed. T2
- 1:33:29was still working. What would happen? T3
- 1:33:31would simply go in and read the latest
- 1:33:33state. it would read in $900. So this
- 1:33:36means that even though T2 is working, T3
- 1:33:39is not affected, right? So all of these
- 1:33:42transactions can go on and on
- 1:33:44simultaneously and there can be millions
- 1:33:46of such transaction and that is where
- 1:33:47the beauty of optimistic concurrency
- 1:33:50control lie. Yeah. And it also allows T2
- 1:33:54to fail and retry. Yeah. So it gives us
- 1:33:57that flexibility. So to reiterate with
- 1:34:00optimistic concurrency control we have
- 1:34:03completely avoided logs. Yeah. And the
- 1:34:07key to this is that if I'm performing
- 1:34:09operations the other transaction isn't
- 1:34:13or doesn't need to be aware of whatever
- 1:34:15I'm doing right they can simply read in
- 1:34:18or they can simply fail and retry.
- 1:34:20Right? My operations are not visible or
- 1:34:23affecting other operations other
- 1:34:26transaction. Yeah. And this allows as I
- 1:34:29discussed earlier it allows very high
- 1:34:31levels of concurrency and this is
- 1:34:33particularly beneficial in read heavy
- 1:34:36systems. Yeah, it is beneficial in read
- 1:34:40heavy systems. So if you were to think
- 1:34:41about it optimistic concurrency control
- 1:34:44basically the foundations lie on the
- 1:34:46fact that
- 1:34:49conflicts
- 1:34:51are very rare.
- 1:34:54They will most unlikely happen, right?
- 1:34:56They are not bound to happen. If they
- 1:34:58happen the conflicting transactions are
- 1:35:00going to read right. Yeah. So that is
- 1:35:02the foundation of optimistic concurrency
- 1:35:04control. And when do conflicts happen
- 1:35:06very rarely
- 1:35:08in those systems where right do not
- 1:35:10happen a lot. Yeah. So in those systems
- 1:35:13which is read heavy. So this is very
- 1:35:15beneficial for read heavy system. It's
- 1:35:17very good for read heavy systems. So I
- 1:35:20believe this helps you get an overall
- 1:35:23idea and an in-depth understanding about
- 1:35:25how delta solves the isolation problem
- 1:35:28quite beautifully using optimistic
- 1:35:30concurrency. Now let's go ahead and talk
- 1:35:32about time travel and virgining. So I
- 1:35:35believe in the previous few example that
- 1:35:37we've seen right we did statements like
- 1:35:40describe history of a table and it
- 1:35:42showed us several versions of that table
- 1:35:45right so versioning basically helps us
- 1:35:48track different stages of the table
- 1:35:51right so let's say you had a table you
- 1:35:53made an update to the table now it
- 1:35:56became a new version of the table let's
- 1:35:58say you made a delete so that becomes
- 1:36:01another version of the table and you can
- 1:36:03go to any of these versions back in time
- 1:36:06and you can restore your table to that
- 1:36:09version. So going back in time and
- 1:36:12restoring it to any other version is
- 1:36:14basically time travel. Right? So we are
- 1:36:16going to have a look at all of this
- 1:36:18through a lot of examples. Let's go
- 1:36:20ahead and create a table first. Right?
- 1:36:23And we going to perform all of the
- 1:36:25operations on this table. So I'll
- 1:36:26quickly copy this path and and create a
- 1:36:30table using the data that is there in
- 1:36:32this park file. So let's go ahead and do
- 1:36:34a create or replace table as then
- 1:36:38actually the name of the table delta
- 1:36:41catalog
- 1:36:44dot delta db dot invoices
- 1:36:49tt time travel and virgining as select
- 1:36:53star from
- 1:36:55park k
- 1:36:58dot this path right so let's go ahead
- 1:37:02and create this.
- 1:37:04Okay. Now let's have a look at this
- 1:37:06table how this looks like. Delta DB
- 1:37:10delta delta catalog
- 1:37:13dot delta DB dot this table. Right? And
- 1:37:17let's do a limit five.
- 1:37:20Okay. So this is how the table looks
- 1:37:21like and I believe we've seen it a
- 1:37:23couple of times already, right? So now
- 1:37:25let's go ahead and perform a few
- 1:37:26operations on this table. So let me go
- 1:37:29ahead and do a delete first, right? So
- 1:37:32we'll do delete from this table where
- 1:37:36customer ID equals 1. So this row that
- 1:37:39you see over here is going to be
- 1:37:40removed. Let me go ahead and run this.
- 1:37:46Okay, great. That succeeded. We are not
- 1:37:47going to verify that now. Uh let's go
- 1:37:50ahead and also do a few more operations.
- 1:37:51So we'll do update
- 1:37:54this table set quantity
- 1:37:58quantity equals 25 where customer ID
- 1:38:03equals 5. Right? So we see that the
- 1:38:06customer ID 5 over here has a quantity
- 1:38:08equals 1 and I want to update that
- 1:38:10quantity to some random number 25.
- 1:38:13Right? Let's go ahead and do that.
- 1:38:16Okay, that works.
- 1:38:19And the final statement we'll do is
- 1:38:23insert into this table and select star
- 1:38:26from
- 1:38:29uh we did it from a park file last time
- 1:38:32right so we inserted customer id from 1
- 1:38:35to 100 let's do from 101 to 200 now yeah
- 1:38:41so now let's verify all the changes very
- 1:38:43quickly so the latest version of the
- 1:38:47table would be this one right if I do a
- 1:38:49select star from what? From this table
- 1:38:51name, I would get the latest version of
- 1:38:53the table. So if I were to do select
- 1:38:56star from this table where customer ID
- 1:38:58equals 1, there should be no rows. Let's
- 1:39:02go ahead and run this. I didn't get a
- 1:39:04row. That means the delete worked fine.
- 1:39:07Now let me check where customer ID
- 1:39:10equals 5. So the quantity is 25. That
- 1:39:13means the update also worked fine.
- 1:39:17And then finally
- 1:39:19let's check how many customers have been
- 1:39:22inserted with customer ID greater than
- 1:39:25100. And the last time we did this we
- 1:39:26saw that there were total 100ed rows. So
- 1:39:30the count star should give me a count of
- 1:39:33100. Perfect. Let's have a look at the
- 1:39:36history of the table. Yeah. So I'm going
- 1:39:38to write describe history of delta
- 1:39:43catalog
- 1:39:45delta db.invoices invoices
- 1:39:48TTV. Right?
- 1:39:51So here we see that there are four
- 1:39:54versions of the table. Right? The first
- 1:39:56version was created when we did a create
- 1:39:59or replace table. Right? This statement
- 1:40:01right over here. So this created the
- 1:40:03first version and then we performed
- 1:40:06three operation. The first one was a
- 1:40:08delete where we deleted customer ID
- 1:40:10equals 1. Then we updated customer ID
- 1:40:12equals 5. And then we wrote in 100
- 1:40:15additional records from 101 to 200 right
- 1:40:18the customer ID. So that created four
- 1:40:21versions. The zero version another
- 1:40:23operation was applied. The first version
- 1:40:25another operation was applied and the
- 1:40:27second version and then so on. Right? So
- 1:40:30these are essentially the version that
- 1:40:32get created and using those versions you
- 1:40:36can go back in time. Right? So let's
- 1:40:38take an example. So if we do a select
- 1:40:41star from this table where customer ID
- 1:40:44equals 1, we won't get any row back,
- 1:40:47right? We know that. But let's say we
- 1:40:49want to go back in time and we say
- 1:40:53version as of we know that in version
- 1:40:57zero it contains the customer ID equals
- 1:41:001 row. Right? So let's do this.
- 1:41:04And there you go. You see that the row
- 1:41:06where customer ID equals 1 is returned
- 1:41:08because this time we pick the table
- 1:41:12where version was zero the oldest
- 1:41:14version over here. Yeah. And you can
- 1:41:16also do this by selecting the time stamp
- 1:41:19instead of the version number. And the
- 1:41:21only change that you need to do is
- 1:41:23select star from this table timestamp
- 1:41:27as of this where
- 1:41:32customer ID equals 1.
- 1:41:35That should work and it gives you the
- 1:41:36same results. So you can either use
- 1:41:39version as of or you can use timestamp
- 1:41:42as of and go back to any of the versions
- 1:41:46back in time. Right? So here we just did
- 1:41:48a select statement. Right? Now what if
- 1:41:50you actually want to restore the table
- 1:41:53to this version? How do you do it? Yeah.
- 1:41:56So the way you do it is by actually
- 1:41:59using the restore command. Yeah. So you
- 1:42:03write restore table delta catalog to
- 1:42:09version as of zero. You can also use
- 1:42:14time stamp as of zero.
- 1:42:17Time stamp sorry not time stamp as of
- 1:42:19zero but time stamp as of whatever time
- 1:42:22stamp you have over here. You can do
- 1:42:25this as well.
- 1:42:27But let me comment this out for now
- 1:42:29because we cannot run both of them at
- 1:42:32the same time.
- 1:42:35So now instead of doing this, instead of
- 1:42:39referring to version number zero, let's
- 1:42:42remove this and then do a where customer
- 1:42:46ID equals 1. So in this case now we
- 1:42:48should get that row because our table
- 1:42:50has been restored to the zero version.
- 1:42:55And there you go. So you see that the
- 1:42:58row where customer ID equals 1 is
- 1:43:00returned. And let's also have a look at
- 1:43:02how the history of the table now looks
- 1:43:04like.
- 1:43:08Great. So it tracks every operation. So
- 1:43:11write was our last operation and after
- 1:43:14right we did a restore. So it tracked
- 1:43:17that operation. Yeah. And it was
- 1:43:20restored to version number zero and that
- 1:43:22is also tracked over here. Yeah. Now
- 1:43:24let's say that you want to restore the
- 1:43:26table to a particular time stamp. Yeah.
- 1:43:30So let's take this example. We want to
- 1:43:33there there's a particular version that
- 1:43:34exist at 9:31 and that version is
- 1:43:37version number two. Another version
- 1:43:39exist which is at
- 1:43:43957.
- 1:43:45Yeah. And this version is version number
- 1:43:48three. Now let's say if I enter a timing
- 1:43:52which looks something like this which is
- 1:43:549:35
- 1:43:56what is going to happen because a
- 1:43:58version doesn't exist at 9:35 right so
- 1:44:01let's say if we were to do something
- 1:44:02like select star from this table
- 1:44:05timestamp as of this timing what would
- 1:44:10happen
- 1:44:14so you see that it gives us some
- 1:44:18results. Yeah, it gives us some results.
- 1:44:20So, what it essentially does is that at
- 1:44:239:35, right, let's say this is 9:31. At
- 1:44:259:35, it looks whether there is a
- 1:44:29version which exists at that at that
- 1:44:30timing or not. If there is a version, it
- 1:44:33will basically pick up that version and
- 1:44:34give it to you. If no version exists, it
- 1:44:37is going to find what is the latest file
- 1:44:40that existed before this time and it is
- 1:44:43going to find out that the latest
- 1:44:44version was version number two. So that
- 1:44:48is why it gives version number two which
- 1:44:51is the file that resided at this point
- 1:44:53in time which is 931. And we can verify
- 1:44:57that. We can verify that by
- 1:45:02by having a look at this file and
- 1:45:05checking the customer ID which says
- 1:45:07where customer ID equals 5
- 1:45:11because you remember that in version
- 1:45:14number two we updated the quantity to be
- 1:45:16equals to 25. So this should give me 25
- 1:45:20if whatever logic we are trying to make
- 1:45:22up right is correct. So let's have a
- 1:45:25look at that and you see that the
- 1:45:27quantity equals 25. Let me also run it
- 1:45:30with the latest table and we would we
- 1:45:32should get different results.
- 1:45:36This should give us quantity equals 1
- 1:45:37because this is based out of version
- 1:45:40zero as we've seen, right? But this is
- 1:45:43based out of version number two. Yeah.
- 1:45:47So I hope this helped you understand
- 1:45:49that if you enter a timing which is not
- 1:45:52present in the in in the history of the
- 1:45:55table it does that mapping by finding
- 1:45:57out the closest file before that
- 1:46:00particular time stamp. So now that we've
- 1:46:02seen all of the action in SQL let's go
- 1:46:04ahead and try a few things using pi
- 1:46:06spark data frame. Yeah. So I'm going to
- 1:46:08write a person in python and this is
- 1:46:11going to be df equals spark read.park.
- 1:46:15Okay. This is not going to be par. This
- 1:46:17is going to be table and this will be
- 1:46:19delta catalog. Delta DB.invoices_ttv
- 1:46:22invoices_ttv
- 1:46:25and I want to read in a particular
- 1:46:26version right so let's say we want to
- 1:46:29read in version number one and this is
- 1:46:31simply because it doesn't contain
- 1:46:33customer ID equals one yeah just to make
- 1:46:36sure we are reading in the right version
- 1:46:37so the way we specify that is by doing
- 1:46:41an option and we specify version as of
- 1:46:45and this will be one and let me do a
- 1:46:49display with a filter df.ilter
- 1:46:55and for that let me first from
- 1:46:58pisparksql
- 1:47:01sqlf functions import column and I will
- 1:47:06do a column and this column is basically
- 1:47:09going to be customer id equals equals 1.
- 1:47:14Yeah. And this should not return to me
- 1:47:17any record.
- 1:47:19So it basically not didn't return to me
- 1:47:21any record and that is why it read in
- 1:47:24the correct version. Now let me also
- 1:47:26repeat this with version number
- 1:47:30version number two and this is going to
- 1:47:33show me that quantity equals 25. Yeah.
- 1:47:36So customer ID equals 5 and this is
- 1:47:40going to be version number two
- 1:47:42and let me replace add in a post in
- 1:47:45Python here.
- 1:47:48So this gives me quantity equals 25. So
- 1:47:50this is how you would read tables um
- 1:47:54read versions of tables using power data
- 1:47:57frame. And you can also replace this
- 1:48:00with something like a
- 1:48:02time stamp as of timestamp as of as of
- 1:48:07and let me again read in this one for
- 1:48:11version number two. So this should also
- 1:48:13give me the same result.
- 1:48:18Okay, there's some
- 1:48:21percent python.
- 1:48:24Great. The quantity equals 25 and it
- 1:48:26gave me the same result. So now let's
- 1:48:28talk about schema validation. And before
- 1:48:31we talk about schema validation, in
- 1:48:33order to set the context, let's talk
- 1:48:35about two interesting things called
- 1:48:37schema on read and schema on write.
- 1:48:42Yeah. So in order to understand schema
- 1:48:45on read we'll take the data lake
- 1:48:47context. Yeah. So inside of data links
- 1:48:50what we can do is we can simply dump in
- 1:48:52any kind of data which is in any kind of
- 1:48:55format and it basically goes in and
- 1:48:58resides in this data lake over here. Now
- 1:49:02it is kept in this data lake and it can
- 1:49:05reside for as long as it wants. And when
- 1:49:08we want to read in data, we simply go
- 1:49:10ahead and pick up that file and read
- 1:49:13that file and then apply the schema
- 1:49:17and then apply the schema. So the schema
- 1:49:20application process happens after the
- 1:49:23data has been stored in the data lake.
- 1:49:25Now if you look at schema on right, what
- 1:49:28basically happens is that let's say
- 1:49:29these are your incoming records.
- 1:49:32First it is checked whether the schema
- 1:49:36matches to your existing schema. Right?
- 1:49:39So let's say there is an orders table.
- 1:49:41It has five columns and those five
- 1:49:43columns have some data types. Right? So
- 1:49:46in order for these records to be
- 1:49:49ingested or to be able to reside in the
- 1:49:52data warehouse, these records have to
- 1:49:54match those five columns and the five
- 1:49:57data type. Yeah. So the application of
- 1:50:00schema happens earlier before injection.
- 1:50:05So this is where injection happened.
- 1:50:07Before injection the application of
- 1:50:09schema happen. Yeah. So this is schema
- 1:50:12on write. Schema on read is basically
- 1:50:15the injection happens earlier
- 1:50:18and then the application of schema
- 1:50:20happened. Now the reason why this is
- 1:50:22problematic is because let's say on some
- 1:50:26date you got the orders file and it
- 1:50:28looks something like order ID
- 1:50:33the amount and then let's say the date
- 1:50:35on which the order was placed and let's
- 1:50:38say after a month the order file looks
- 1:50:40completely different. This is the order
- 1:50:43ID and then this is going to be the
- 1:50:46order details
- 1:50:48and then there is going to be something
- 1:50:50like the amount and the date residing
- 1:50:54inside of order details and now this is
- 1:50:57also ingested. Yeah. So different kinds
- 1:51:00of files with different kinds of schemas
- 1:51:03gets ingested. Now when you want to read
- 1:51:05orders you get confused whether this is
- 1:51:08the right schema or this is the right
- 1:51:10schema
- 1:51:12and there is going to be several
- 1:51:14problems. So for example if you want to
- 1:51:16find out the orders some of these file
- 1:51:18you're you'll be able to easily read it
- 1:51:20but in other files you'll have to do
- 1:51:22some transformation basically pull it
- 1:51:24out from here.
- 1:51:26So the schema is not consistent it is
- 1:51:29not well maintained over here. So that
- 1:51:32basically helps you understand what
- 1:51:34schema on read and schema on write is
- 1:51:36and delta
- 1:51:38is schema on write. So that helps you
- 1:51:42enforce schema. It validates the schema
- 1:51:45before making any data come inside of a
- 1:51:49delta table. So let's understand schema
- 1:51:51validation with example. Yeah. So I'm
- 1:51:54going to create a table that we are
- 1:51:57going to specifically use to understand
- 1:51:59schema validation. And for that let me
- 1:52:02simply do create or replace table
- 1:52:08delta catalog dot delta db dot invoices
- 1:52:14underscore sv for schema validation.
- 1:52:17Yeah. So the column that I'm going to
- 1:52:19have is customer ID. Uh this is going to
- 1:52:23be int and I'm going to add a
- 1:52:24constraint. So this constraint is going
- 1:52:26to be not null. My customer ID cannot be
- 1:52:29null because it is the primary key of my
- 1:52:32table. And then it is going to have
- 1:52:35invoice number. This is going to be a
- 1:52:38string.
- 1:52:40Then we are going to have quantity which
- 1:52:42is going to be an integer. Then we are
- 1:52:45going to have price. This is going to be
- 1:52:47a let's say float. And we are going to
- 1:52:52have invoice
- 1:52:56date. And this is going to be a date.
- 1:52:59And this is going to be the definition
- 1:53:02of our table. Now we have to also insert
- 1:53:06some data inside our table, right? And
- 1:53:08we're going to do that using the insert
- 1:53:11statement. So insert into
- 1:53:14into this table
- 1:53:16and we're going to use the select
- 1:53:18statement. So select star from
- 1:53:23we're going to use the park file that we
- 1:53:25used earlier.
- 1:53:27This contains all the customers for for
- 1:53:30customer ids from 1 to 100 from par and
- 1:53:34this is going to be the file and I'm
- 1:53:37going to select the relevant columns
- 1:53:39only because as we've seen earlier it
- 1:53:40has a lot of column. So it is going to
- 1:53:42be customer ID, invoice number,
- 1:53:47quantity,
- 1:53:49price
- 1:53:51and invoice date. Now let's go ahead and
- 1:53:55run this.
- 1:53:57Let's also get a feel of the data. Let's
- 1:53:59see how this looks like. Let me do a
- 1:54:02limit five.
- 1:54:18So now you see that we have the table
- 1:54:20created, right? And just to do a few
- 1:54:24checks,
- 1:54:25this table should have customers
- 1:54:29customer ID from 1 to 100 and a total of
- 1:54:32100 customers.
- 1:54:34Customer ID and the count star
- 1:54:40from this table should be 100 and 100.
- 1:54:45Okay. So now that we have the table set
- 1:54:47up, let's understand column order
- 1:54:50validation. Yeah. So let me quickly
- 1:54:53write down this heading scenario one
- 1:54:57column order validation. Yeah. So in
- 1:55:01order to do this, we are going to run an
- 1:55:03insert statement. We're going to insert
- 1:55:06into the table
- 1:55:09and we are simply going to do it through
- 1:55:11a select statement. Select
- 1:55:14these columns
- 1:55:17these columns from values because we're
- 1:55:21going to insert only one row as T and
- 1:55:25this is going to be all of this. Now the
- 1:55:27customer ID is going to be a large
- 1:55:30number 99,999.
- 1:55:34The invoice is going to be 1 2 3 4 I 1 2
- 1:55:373 4 5. Quantity is going to be 10. Price
- 1:55:40is going to be 100. and invoice ID
- 1:55:43invoice date is going to be 2025 0101.
- 1:55:46Yeah. Now before we do the insert let's
- 1:55:48see how this looks like. So this
- 1:55:51basically looks like the first customer
- 1:55:54ID is 99,999
- 1:55:56and all of the all of the rest right
- 1:55:58quantity is 10. Now what I want to do is
- 1:56:01interchange the order. Now because
- 1:56:04quantity and customer id are both
- 1:56:06integers I want to interchange them so
- 1:56:09that there is no conflicting data types.
- 1:56:11Right? Now let's run this. So the first
- 1:56:14column now becomes quantity which is 10
- 1:56:17and customer ID becomes the third column
- 1:56:20which is 99,999.
- 1:56:23So what I would desire is that if the
- 1:56:26insert statement happened correctly, a
- 1:56:28new row gets inserted where customer ID
- 1:56:31equals 99,999.
- 1:56:33Yeah. And quantity equals 10. So let's
- 1:56:35go ahead and run this.
- 1:56:38And meanwhile, let me also write select
- 1:56:40star from
- 1:56:42okay. So this completed successfully.
- 1:56:45Let me go ahead and write this where
- 1:56:48customer ID equals 99,999.
- 1:56:53Yeah. So this is what we wanted. A row
- 1:56:56which has customer ID 99,999.
- 1:57:00That should be in the table because we
- 1:57:02just ran an insert statement. Let's go
- 1:57:04ahead and run this. What do you think?
- 1:57:05Would we find that row there?
- 1:57:09Okay. So that row is not present there.
- 1:57:12Where did it go? What happened? What
- 1:57:14just happened to my table?
- 1:57:17So what we see over here is
- 1:57:21let me go ahead and run this again.
- 1:57:24What we see over here is the first
- 1:57:27column was 10. So what actually happened
- 1:57:30was that it took in the columns by
- 1:57:33position and it mapped to the customer
- 1:57:36ID column 0 column number zero column
- 1:57:38number zero. So I'm going to take in the
- 1:57:40value from this column which is at
- 1:57:43position number zero and put it inside
- 1:57:46customer ID. Yeah. So this is something
- 1:57:48called column matching by position. So
- 1:57:53let me go ahead and see if we have
- 1:57:57where
- 1:57:59customer ID equals 10. So we already had
- 1:58:03customer a customer ID equals 10, right?
- 1:58:05because we inserted data from 1 to 100.
- 1:58:09So that means now after this insert
- 1:58:12there should be two rows.
- 1:58:16There you go. So we have two rows which
- 1:58:19is customer ID equals 10 and invoice
- 1:58:22number 1 2 3 4 5. Quantity is 99,999
- 1:58:26and the rest of what we inserted. So
- 1:58:28what insert did is that it executed the
- 1:58:34insertion by matching columns by
- 1:58:36position and not the name. And that is
- 1:58:40why you see what you see over here. And
- 1:58:42this actually corrupts your data. So
- 1:58:44never use an insert unless you're very
- 1:58:47sure that the positions will always be
- 1:58:51maintained. So let's see what is going
- 1:58:53to happen if we reorder the columns and
- 1:58:57insert that data into our table using a
- 1:59:01merge statement. Is it going to behave
- 1:59:03any differently? Yeah. So let's try that
- 1:59:05out. And the source data that we are
- 1:59:08going to use is all the customer ids.
- 1:59:11Actually not all the customer ids. Uh
- 1:59:14the customer ids from 101 to 200. And
- 1:59:18I'm just going to take five of them.
- 1:59:21order by customer ID descending limit
- 1:59:25five. So this is going to give me five
- 1:59:27customer ids from 196 to 200. Yeah. So
- 1:59:32let's go ahead and write the merge
- 1:59:33statement. Merge into
- 1:59:36this table as target using
- 1:59:40this source table.
- 1:59:44Using this source table,
- 1:59:47the join condition is going to be on
- 1:59:50target dot customer ID equals source dot
- 1:59:52customer ID and when not matched
- 1:59:58then
- 2:00:00simply insert star. So these five rows
- 2:00:04they are not going to match with the
- 2:00:06existing data that we have in the table.
- 2:00:09Right? because the existing data is for
- 2:00:11customers ids from 1 to 100 right so
- 2:00:14these are not going to match and they
- 2:00:17are supposed to be inserted in the table
- 2:00:19yeah now what happens if you run this
- 2:00:22statement again right again and again so
- 2:00:24the first time it gets inserted the
- 2:00:26second time you run it it is going to
- 2:00:28match this statement because of this
- 2:00:30statement target customer id equals sort
- 2:00:32customer ID they are going to match
- 2:00:34right because 196 to 200 already resides
- 2:00:37in our table so in that case What we are
- 2:00:39going to do is that when matched then
- 2:00:43update
- 2:00:44set the following
- 2:00:47column right
- 2:00:50set customer id dot customer ID equals
- 2:00:54source dot customer id target dot
- 2:00:58invoice number equals s.invoice invoice
- 2:01:01number the quantity
- 2:01:03target dot price equals source.p price
- 2:01:08and the final one target dot invoice
- 2:01:11date equals the current date. So if they
- 2:01:16match what we are simply going to do is
- 2:01:18we are going to keep all of the column
- 2:01:20the same. We just going to update the
- 2:01:22invoice date. Now let's go ahead and
- 2:01:25make the change. We are going to swap
- 2:01:28the columns over here.
- 2:01:30Quantity going to be the first column
- 2:01:32and customer ID going to be the third
- 2:01:34column. Exactly same as what we did over
- 2:01:37here. Quantity would the first column
- 2:01:39and customer ID what the third column.
- 2:01:41And let me also do one change because
- 2:01:43this is coming from this park file.
- 2:01:45Price over here is a float, right? So
- 2:01:49this should also be a float.
- 2:01:53Cast this as float. Yeah. So now let's
- 2:01:57go ahead and run this. So while this
- 2:02:00runs,
- 2:02:01let me write some SQL to quickly
- 2:02:03validate this. Select star from this
- 2:02:05table where
- 2:02:08where customer ID greater than 100,
- 2:02:10right? And the only customer ID is
- 2:02:12greater than 100 in this table should be
- 2:02:14these ones, right? These five record. So
- 2:02:17we run this
- 2:02:20and there you go. So you see that the
- 2:02:23customer ids from 196 to 100 have been
- 2:02:25returned. That mean that the merge
- 2:02:27statement ran correctly. It was able to
- 2:02:30identify which columns to match. Now a
- 2:02:33very important point to note is that
- 2:02:35insert the matches columns by position.
- 2:02:38The zero position is going to be matched
- 2:02:40to the zero position from the source.
- 2:02:43But what merge does is that it matches
- 2:02:46column by name. It finds the right name
- 2:02:48even if they are not ordered correctly
- 2:02:51and then matches them. And that is the
- 2:02:52reason why this worked. Yeah. So now
- 2:02:55let's go ahead and run this once again,
- 2:02:57right? So that it comes to this clause
- 2:03:00and then it just updates the invoice
- 2:03:03date. Yeah. So let's go ahead and run
- 2:03:04this now.
- 2:03:07And ideally
- 2:03:09this is the only column that should
- 2:03:11change.
- 2:03:14Okay, let's order this from 200. And
- 2:03:17there you go. You see that this is the
- 2:03:19only column that changed while all the
- 2:03:21other columns are just the same. So
- 2:03:24merge seemed to work just fine. So now
- 2:03:27let's try to understand how delta does
- 2:03:29data type validation. Right? How it does
- 2:03:33data type validation. And this is going
- 2:03:35to be scenario number two. So let me
- 2:03:39quickly write that down. Data type
- 2:03:43validation.
- 2:03:46And
- 2:03:48we're going to test this by writing an
- 2:03:50insert statement. So we're going to say
- 2:03:52insert into delta catalog do this table
- 2:03:56and the values are going to be
- 2:04:01let's say the customer ID is going to be
- 2:04:03ABC
- 2:04:04invoice number is going to be I 4 5 6 7
- 2:04:088
- 2:04:10quantity is going to be 10 price is
- 2:04:12going to be 98.75
- 2:04:15and the invoice date is going to be
- 2:04:1720250101
- 2:04:19yeah now if I were to ask you whether
- 2:04:22this piece of code will run or not. I
- 2:04:24think the obvious answer would be no
- 2:04:26because the data type for customer ID is
- 2:04:30integer but what we are trying to feed
- 2:04:32into it is is string. So let's go ahead
- 2:04:36and run this and as you expected it
- 2:04:39didn't work because this is string and
- 2:04:40it tried to convert it to integer but it
- 2:04:43didn't work. Yeah. But now let's go
- 2:04:46ahead and do something different right.
- 2:04:48So let me go ahead and write this
- 2:04:51customer ID which is 99499.
- 2:04:55Now what do you think? Will it run?
- 2:04:58Let's go ahead and run this.
- 2:05:02Okay, there you go. So it seems that it
- 2:05:05ran.
- 2:05:07Let's do a select star from
- 2:05:11delta catalog. So the intelligent
- 2:05:14sometimes works and sometime doesn't.
- 2:05:17It's completely dependent on it mode uh
- 2:05:20where customer ID equals this right.
- 2:05:25So you see that there is a row 99499 and
- 2:05:28it contains the exact same details 98.75
- 2:05:32and quantity. So the question here is
- 2:05:34why did this work at all if this did
- 2:05:37this didn't? So the reason why this
- 2:05:39worked is because
- 2:05:42what delta does is that it does try it
- 2:05:45puts in effort to convert it into the
- 2:05:48form that it actually exists in the
- 2:05:50table. So in the table it is in integer
- 2:05:53form. It tries to convert it to an
- 2:05:56integer form. Yeah. If the conversion
- 2:05:59happens successfully, it inserts the
- 2:06:02data into the table. If the conversion
- 2:06:04doesn't happen, then it throws an error.
- 2:06:06So that is why you see the error thrown
- 2:06:09over here is that string cannot be cast
- 2:06:12to int. Yeah. So it tried to cast it but
- 2:06:14it cannot cast it to int. And that is
- 2:06:17also the reason why this 2025 0101 this
- 2:06:22is string but the data that we have over
- 2:06:24here is date. So it casted string to
- 2:06:28date and then inserted the data in this
- 2:06:32table. Yeah. But this also makes sure
- 2:06:35that we just cannot dump in any garbage
- 2:06:37data. So that check is there in place.
- 2:06:40Let's have a look at another scenario
- 2:06:42which is column name validation. Yeah.
- 2:06:45So okay, it seemed I missed writing
- 2:06:49column name validation here. Column name
- 2:06:52validation. And this is going to be
- 2:06:54number five. number six and this is
- 2:06:57scenario number three right so let's go
- 2:07:00ahead and add a heading scenario number
- 2:07:04three column name validation
- 2:07:08and the way we are going to test this is
- 2:07:11simply I'm going to copy the code that I
- 2:07:14use for column order validation and we
- 2:07:17are first going to test this with an
- 2:07:19insert statement
- 2:07:21so let's make sure that the order is the
- 2:07:24name
- 2:07:25because there's no point changing the
- 2:07:27order, right? We've already tested for
- 2:07:29order
- 2:07:31and let's go ahead and change the column
- 2:07:34name.
- 2:07:36This is going to be as QTY.
- 2:07:40So now we've changed customer ID to C ID
- 2:07:43quantity to QTY. Yeah. Let's go ahead
- 2:07:46and run this now.
- 2:07:50So if this SQL ran correctly,
- 2:07:54you should see something like a row, a
- 2:07:58row which basically contains customer ID
- 2:08:01equal this number.
- 2:08:04Okay, there you go. So you see this row
- 2:08:06where customer ID equals 99,999
- 2:08:09and all of these other details, right?
- 2:08:12So the reason why this worked even
- 2:08:15though the column names are different is
- 2:08:18because because of a reason that we
- 2:08:19already discussed some time back right
- 2:08:21so it does column matching by position
- 2:08:24and it doesn't matter whatever in the
- 2:08:27world the name that you decide to put
- 2:08:29over here it doesn't matter and that
- 2:08:31will be of no effect. So that is the
- 2:08:33reason why this worked. Now let's try
- 2:08:35the same thing with a merge statement.
- 2:08:40So last time we used
- 2:08:44this data over here.
- 2:08:47This basically contained customer ids
- 2:08:51from okay the order is changed. Let me
- 2:08:55get this back to the same order.
- 2:09:02So it contain customer ID from 196 to
- 2:09:04200. Uh let me change it
- 2:09:08to have from 101 to 105. Yeah. Now let's
- 2:09:12go ahead and insert this data.
- 2:09:18And before before actually we run this
- 2:09:20query, we want to change we want to
- 2:09:24change the name of the column, right? So
- 2:09:26we want to change this to qty.
- 2:09:30Uh where are the other places? C id
- 2:09:34id qty. Right? So now that we've made
- 2:09:38all of the changes, right? Let's go
- 2:09:40ahead and run this and see what happens.
- 2:09:42So if this statement runs successfully,
- 2:09:44what what we should see is
- 2:09:47when we run something like this, right?
- 2:09:50If we run something like this,
- 2:09:54I should see five more rows.
- 2:09:57So I should see these rows as well. So
- 2:10:00let's go ahead and run this now.
- 2:10:04So now what it says is cannot resolve
- 2:10:07customer ID in insert clause given
- 2:10:09columns is source C ID. So the reason is
- 2:10:15pretty simple and it is again something
- 2:10:17that we already discussed. So think
- 2:10:19about it for a minute if you're not able
- 2:10:21to recall why it failed. So the reason
- 2:10:24why it failed is simply because merge
- 2:10:27does column matching by name. So it
- 2:10:30looks at the target table. It takes up a
- 2:10:33column let's say customer ID. It looks
- 2:10:35up for that exact column in the source
- 2:10:38table. Okay. Do I find customer ID in
- 2:10:40the source table or not? If it finds it
- 2:10:43then it is going to dump all of that
- 2:10:44data in the target table. But if it is
- 2:10:47not able to find it, it is going to
- 2:10:49fail. So it is quite a robust statement.
- 2:10:53So now let's have a look at another
- 2:10:54scenario which is nullability
- 2:10:56validation. Yeah. So let me quickly
- 2:11:01add in the heading
- 2:11:05scenario four nullability
- 2:11:08validation. And here uh by nullability I
- 2:11:11just want I just don't want to talk
- 2:11:12about nullability but a few other things
- 2:11:15as well. So let's have a look at this.
- 2:11:16insert into
- 2:11:19delta catalog dot
- 2:11:22this invoices table values we have how
- 2:11:25many columns 1 2 3 4 5 so there's going
- 2:11:27to be five nulls and if I were to just
- 2:11:32copy this
- 2:11:35and paste it five times and if I were to
- 2:11:38ask you whether this will run or not
- 2:11:39obviously you would say that this won't
- 2:11:41run
- 2:11:43so the reason why it didn't run is
- 2:11:44because the not null constraint raint
- 2:11:46violated for column customer id. So we
- 2:11:50had a constraint over here customer ID
- 2:11:53int not null whatever customer id we
- 2:11:56insert and for now there can be
- 2:11:57duplicates right uh but they shouldn't
- 2:12:00be null yeah so that constraint was
- 2:12:03violated over here and that was the
- 2:12:05reason why it wasn't so if we run the
- 2:12:07same statement right if we run the same
- 2:12:10statement inserting some random number 7
- 2:12:138 912 this will work
- 2:12:19And as expected this works. So the point
- 2:12:22that I'm trying to make over here is not
- 2:12:25just nullability validation. It's about
- 2:12:27constraints validating the constraints
- 2:12:30that we've put. Right? So we put in a
- 2:12:33constraint which is not null over here.
- 2:12:36We can also put in constraints like for
- 2:12:39example the price should be greater than
- 2:12:42zero. The quantity should be greater
- 2:12:44than zero. Something like that. We can
- 2:12:45put in rules over here. And whenever
- 2:12:48data is inserted all of those rules
- 2:12:51would be checked for and this would make
- 2:12:54sure that the data that is getting
- 2:12:56inside our tables are correct they are
- 2:12:58accurate. Now let's have a look at the
- 2:13:00last scenario which is extra column
- 2:13:03validation. What happens if in my source
- 2:13:06the data that I'm trying to insert it
- 2:13:08has extra columns and the target table
- 2:13:11does not have those many columns. Right?
- 2:13:12So what is going to be the case? So this
- 2:13:15is going to be
- 2:13:18scenario
- 2:13:20number five and this is going to be
- 2:13:22extra column validation.
- 2:13:26So let me copy some code. I'm going to
- 2:13:29copy the code that I used for column
- 2:13:32order validation
- 2:13:34and
- 2:13:36I'll comment this out for now. And let
- 2:13:38me add in an extra column. Right? And
- 2:13:40this is a dummy column. Let me name it
- 2:13:42as customer type. So you see that all of
- 2:13:46the columns over here are intact. And
- 2:13:49actually let me also change
- 2:13:52correct the orders.
- 2:13:55Customer ID should be over here.
- 2:13:58Yeah. So the orders have been corrected
- 2:14:01now. And let's also change this. Right?
- 2:14:06So the orders have been corrected now.
- 2:14:07And we have an extra column which is
- 2:14:09customer type. Now let's see what is
- 2:14:11going to happen.
- 2:14:14So you see that this didn't work. It
- 2:14:17says that a schema mishmash detected
- 2:14:20when writing to the delta table. Now
- 2:14:23let's go ahead and try this with a merge
- 2:14:26statement. Yep.
- 2:14:29So if I were to
- 2:14:32copy this
- 2:14:34this merge statement over here and let's
- 2:14:37go ahead and run this. Actually, let's
- 2:14:39change the source data a little bit.
- 2:14:41Right. So now I'm going to take up all
- 2:14:45of the customers whose ID is between
- 2:14:49where customer ID between 150 and 155.
- 2:14:54And
- 2:14:56I want the C the the order of columns to
- 2:15:00be the same.
- 2:15:02So let's go ahead and run this. So this
- 2:15:05gives me all the customers whose ID is
- 2:15:07between 150 to 155. Right? So let's go
- 2:15:11ahead and paste this over here.
- 2:15:14And let's also add in an extra column
- 2:15:16over here which is VIP as customer type.
- 2:15:22Now these rows are not in the table,
- 2:15:26right? So it should go to this clause
- 2:15:28when not mash then insert star. So all
- 2:15:30of them should be inserted. And because
- 2:15:32I want to test that particular case
- 2:15:33right now, let me just remove all of
- 2:15:36that for simplicity. So let's go ahead
- 2:15:37and run this
- 2:15:40and actually before I run this uh let me
- 2:15:44show you that there is
- 2:15:48there is no record. Select start from
- 2:15:52this table. This would give me no
- 2:15:54record. Now let's go ahead and run this
- 2:15:55now.
- 2:15:58And you see that it succeeded.
- 2:16:01Let's run this.
- 2:16:03Great.
- 2:16:05So I see all of the records from 150 to
- 2:16:09155 that means that the insert statement
- 2:16:12has worked correctly. So it tries to
- 2:16:15match the columns by their names and if
- 2:16:19it's not able to find a name. So
- 2:16:21basically it started with the target. It
- 2:16:24searched for all the five names. It was
- 2:16:25able to find them and then it just went
- 2:16:28off right. It succeeded and the
- 2:16:30operation completed. Yeah. So the merge
- 2:16:33operation is actually quite a robust way
- 2:16:36of doing things, right? It was able to
- 2:16:38handle a few other issues earlier as
- 2:16:41well. Now let's understand what schema
- 2:16:43evolution is, right? Schema evolution is
- 2:16:47the ability to handle changes in schema
- 2:16:50without having to completely rewrite
- 2:16:53your table or without having to rewrite
- 2:16:56the underlying data. Yeah. So when I say
- 2:17:00ability to handle changes, what I mean
- 2:17:02is that let's say if I have a table, if
- 2:17:04a new column comes in, I should be able
- 2:17:06to accommodate that new column. Or let's
- 2:17:10say if I have a column which is of type
- 2:17:13int. Let's say order ID is of type int
- 2:17:16and for some reason I'm getting a lot of
- 2:17:18orders and my order ID has gone beyond
- 2:17:22the range of integer and I want to
- 2:17:24change it to long or big int or
- 2:17:26something like that. Right? So I should
- 2:17:27be able to upgrade my data type to
- 2:17:30another type. So all of that flexibility
- 2:17:33should be allowed. So we are going to
- 2:17:35have a look at four scenario. The first
- 2:17:37one is adding new columns. One is the
- 2:17:40manual way, the other one is the
- 2:17:41automatic way. So we'll be having a look
- 2:17:43at both of them. The second one as I
- 2:17:45discussed is widening of data types. The
- 2:17:49third one is nested structure evolution.
- 2:17:51The last one is column position changes.
- 2:17:54If for whatever reason I want to reorder
- 2:17:56the columns in my table, there should be
- 2:17:58enough flexibility to allow me to do
- 2:18:00that. Let's go ahead and understand the
- 2:18:02first scenario which is adding new
- 2:18:05column.
- 2:18:06Let me first quickly write in
- 2:18:10scenario number one which is adding
- 2:18:13adding new columns. And for this we are
- 2:18:17going to create a brand new table and
- 2:18:19for that I'll copy the code from from
- 2:18:22schema validation. Uh so let me just
- 2:18:25remove this line. Change this to SE for
- 2:18:28schema evolution.
- 2:18:31Quantity is going to be removed because
- 2:18:32I want to keep minimal number of columns
- 2:18:34and minimal number of rows as well for
- 2:18:37this example. So let me do a where
- 2:18:41customer ID between 1 and five. So we'll
- 2:18:45insert five rows into this table.
- 2:18:48And let's also remove quantity. Let's go
- 2:18:52ahead and run this. Quite simple. We've
- 2:18:55got what we expected. And let me put
- 2:18:58this over here. Let's go ahead and run
- 2:19:00this. Okay, my bad. This should be se.
- 2:19:06Let's go ahead and run this now.
- 2:19:09And let's quickly validate
- 2:19:12that
- 2:19:14the table looks just as we expected it
- 2:19:18to look like. And it looks as expected.
- 2:19:22Yeah. So the table is set up now. Now
- 2:19:25let's go ahead and add a column to the
- 2:19:28table.
- 2:19:29So we are going to run an alter table
- 2:19:36invoiced SC add column and the column
- 2:19:40that we going to have is let's first see
- 2:19:42how many columns are there like what all
- 2:19:44what all columns are there. So for that
- 2:19:46let me run a select star and we see that
- 2:19:49there are okay we'll use the quantity
- 2:19:51column that we just dropped we didn't
- 2:19:53want to select that so let's use that
- 2:19:55for now and this is going to be quantity
- 2:19:57integer
- 2:20:00so this column has been added to the
- 2:20:01schema of the table now now I want to
- 2:20:07insert a few more rows and this time it
- 2:20:11is going to be with the quantity column
- 2:20:14And let's insert rows from 6 to 10 now.
- 2:20:20So let's go ahead and run this.
- 2:20:24Let me copy this to see how the final
- 2:20:26table looks like. So the final table
- 2:20:28should look like uh six five columns in
- 2:20:31place with these rows not having the
- 2:20:35quantity data. They should all be null.
- 2:20:37And the one that we inserted right now,
- 2:20:39they should all have non-null data for
- 2:20:42quantity. Let's run this.
- 2:20:46And there you go. As we expected, we see
- 2:20:49non-null data for quantity for the
- 2:20:52customer ids that we've just inserted
- 2:20:54and null quantity for the one that were
- 2:20:57inserted earlier. So you see that we had
- 2:21:00to run an alter statement over here in
- 2:21:03order to accommodate this new column
- 2:21:06that was coming in. Now what if you want
- 2:21:09to do this automatically?
- 2:21:11Although there can be situations where
- 2:21:13we want to go ahead with this approach.
- 2:21:15Sometimes we don't want to go ahead
- 2:21:17because we don't randomly want any kind
- 2:21:19of data to come in and then sit in my
- 2:21:20table. Right? But let's say for this
- 2:21:22situation, what if you want to
- 2:21:24automatically accommodate a new column?
- 2:21:27So for that case, you need to set this
- 2:21:32property to true and this enables schema
- 2:21:36evolution. So let's go ahead and run
- 2:21:38this now. So this is enabled to true.
- 2:21:41And let me run another insert statement.
- 2:21:46This time without doing an alter without
- 2:21:48running an alter table command but I
- 2:21:52will add another column. So let's see
- 2:21:54which column do I want to add. So let's
- 2:21:58add payment method. Let's add payment
- 2:22:00method. So let's add
- 2:22:04payment method over here.
- 2:22:08And this is going to be 11 to 15. And
- 2:22:13this is going to be inserted in the
- 2:22:15table. The table had how many? It had 1
- 2:22:172 3 4 5 column. Now if this works well,
- 2:22:21we are going to have six columns. Right?
- 2:22:23And the payment method column is going
- 2:22:25to be populated only for customers ids
- 2:22:28from 11 to 15. So let's run this. and we
- 2:22:33do a select star from delta catalog blah
- 2:22:38blah blah and let's run this. Okay,
- 2:22:41there you go. So you see that from 11 to
- 2:22:4415 we have the payment method populated.
- 2:22:50That means that this statement worked
- 2:22:52pretty well and automatic schema
- 2:22:55evolution had just happened. You didn't
- 2:22:57need to run an alter statement in order
- 2:23:00to make it. Now let's talk about the
- 2:23:02second scenario which is type widening
- 2:23:05right and this was introduced in delta
- 2:23:08version 3.2. This simply means that you
- 2:23:11can upgrade a type to its bigger type.
- 2:23:15Yeah. Simply meaning that let's say if
- 2:23:17you have an int you want to convert it
- 2:23:19to a big int for whatever reason. The
- 2:23:21example that we discussed, if our orders
- 2:23:24have spanned to such a large number that
- 2:23:27it is not able to fit in the integer
- 2:23:30data type, we want to upgrade it to a
- 2:23:32big array. Similarly, you want to
- 2:23:34upgrade a float to a double or a var of
- 2:23:37a specified number of characters to more
- 2:23:40number of characters. All of that is
- 2:23:42possible using type widening. So before
- 2:23:45we get started with example, we need to
- 2:23:48ensure that we at least have delta
- 2:23:51version 3.2. And if you remember during
- 2:23:54the initial parts of the video, I
- 2:23:57created a cluster.
- 2:23:59The cluster had a datab bricks runtime
- 2:24:02version of 14.3. And I've pulled this up
- 2:24:06to show you
- 2:24:08that
- 2:24:10if I were to go to 14.3
- 2:24:13and quickly search for delta,
- 2:24:17this is going to be 3.1.0.
- 2:24:20That means this DBR database runtime is
- 2:24:23not going to work. So let me quickly
- 2:24:26check 15.4 and what does that show me?
- 2:24:30So this shows me 3.2.0.
- 2:24:33That means this is going to work. So,
- 2:24:34let me go ahead and upgrade this to
- 2:24:3815.4.
- 2:24:39Keeping all of the other stuff the same.
- 2:24:42Let's click on confirm and let's start
- 2:24:45the cluster. Okay. So, now that we have
- 2:24:48our cluster ready,
- 2:24:51let's go ahead and start with the
- 2:24:52example. And before that, let me quickly
- 2:24:54write down scenario
- 2:24:57scenario number two which is type
- 2:25:00widening.
- 2:25:02And we also need to do another thing
- 2:25:04which is enabling type widening on this
- 2:25:07table. So let me how to enable type
- 2:25:13widening on a delta table.
- 2:25:17Let's quickly search for that.
- 2:25:23Okay. So this is the property that I
- 2:25:25need to enable and let's go ahead and
- 2:25:28run this. Hopefully this should work.
- 2:25:32And meanwhile, let me go ahead and write
- 2:25:35an insert statement.
- 2:25:37Insert into
- 2:25:40this table. Actually, before that, let's
- 2:25:43see how the columns currently look like.
- 2:25:45So, describe this table delta catalog
- 2:25:49dot invoice delta DB
- 2:25:55dot invoices SA. Right? How do the
- 2:25:57columns look like currently? So customer
- 2:26:00ID is an integer right now and we have
- 2:26:03enabled
- 2:26:05automatic schema evolution right. So let
- 2:26:09me go ahead and write an insert
- 2:26:12statement where values
- 2:26:15this I'll copy some values from here
- 2:26:20and
- 2:26:22this is a string this is also a string
- 2:26:26comma comma this is going to be a string
- 2:26:29yeah and this is an integer so let me
- 2:26:31paste it over here and I want to insert
- 2:26:35a number which is a big integer.
- 2:26:39Although you see that I have not
- 2:26:40currently changed the data type but I
- 2:26:44still want to insert it and check
- 2:26:45whether it works or not because I have
- 2:26:47automatic schema evolution enabled. So
- 2:26:50we want to try it out whether it works
- 2:26:52or not. So let's write give me an
- 2:26:55example
- 2:26:58of a big number
- 2:27:02and that is the number that I'm going to
- 2:27:05use over here.
- 2:27:07So I have six columns,
- 2:27:10six numbers over here. And let's go
- 2:27:12ahead and run this. 36 values, not
- 2:27:16numbers. Okay, I missed a comma here.
- 2:27:18Let's run this now.
- 2:27:21Okay, so the error it throws is fail to
- 2:27:24assign a value big int to type int.
- 2:27:28Right? That simply means that using
- 2:27:33automatic schema evolution type widening
- 2:27:36won't work. We will have to manually
- 2:27:39change the type of customer ID. Yeah. So
- 2:27:42in order to do that let's go ahead and
- 2:27:44run alter table
- 2:27:48invoices se and we will write alter
- 2:27:51column
- 2:27:53alter column customer
- 2:27:56customer id and the type that we want is
- 2:28:01big int. Let's go ahead and run this
- 2:28:06that is successful. So we say describe
- 2:28:08table
- 2:28:11invoices
- 2:28:12sc and now you see that the customer ID
- 2:28:15is big int earlier it was integer. So
- 2:28:17let's go ahead and run the same
- 2:28:19statement again. Let's see what happened
- 2:28:29like star from.
- 2:28:32Okay this seems to have succeeded.
- 2:28:35And if this actually succeeded, that
- 2:28:37means that this table should have
- 2:28:43should have a row where customer ID
- 2:28:45equals this value.
- 2:28:53Perfect. That means this row resized in
- 2:28:56the table and type widening has
- 2:28:58successfully worked and we've enabled
- 2:29:00type widening. Next scenario that we are
- 2:29:02going to talk about is nested structure
- 2:29:05evolution.
- 2:29:08Scenario number three is nested
- 2:29:12structure evolution.
- 2:29:16So let's add a strct column to our
- 2:29:20table. So in order to do that alter
- 2:29:22table
- 2:29:25the invoices are se table add column and
- 2:29:29let's go ahead and add something called
- 2:29:31purchase details and this is going to be
- 2:29:34a strct.
- 2:29:36The strct is going to have two column.
- 2:29:38The first one is mall pin code. The mall
- 2:29:42at which the purchase was made and this
- 2:29:44is going to be an integer and then the
- 2:29:47store code. the store at which the
- 2:29:49purchase was made. So let's go ahead and
- 2:29:51run this and let's add some sample data
- 2:29:55to this table. So
- 2:29:59this is going to be so we inserted
- 2:30:02around 15 records.
- 2:30:05So this is going to be the 16th one. So
- 2:30:08let's add 16th.
- 2:30:10And then we put a strct over here. the
- 2:30:14mall pin code. Let's let's put in some
- 2:30:16random number and let's put in some
- 2:30:20random number for the store code. Yeah,
- 2:30:22let's go ahead and run this. So, this
- 2:30:24should run just fine. So, let's do
- 2:30:27select star from
- 2:30:30delta catalog blah blah blah. And
- 2:30:36perfect. So this should be the only row
- 2:30:38where you had purchase details and it
- 2:30:40has all of
- 2:30:43the data that we just put in the mall
- 2:30:45pin code and the store code. Now let's
- 2:30:47consider two examples. What if you need
- 2:30:49to change
- 2:30:51the data type of mall pin code for
- 2:30:54whatever reason from int to begin and
- 2:30:57we're going to follow a pretty similar
- 2:30:59way.
- 2:31:01Alter table. This table. Alter column.
- 2:31:06Alter column. Uh
- 2:31:10purchase details dot mall pin code.
- 2:31:16The type that we want it to be is big
- 2:31:19in. Let's go ahead and run this.
- 2:31:22And let's insert
- 2:31:25some more data to validate this. And
- 2:31:27this is going to be number 17.
- 2:31:31Let's add in a strct over here. And
- 2:31:38the store code can be any number. But
- 2:31:41for the mall pin code, I need a big int.
- 2:31:44So I'm going to put that number over
- 2:31:45here. Let's just increase it by one. Uh
- 2:31:49and this number can just be any other
- 2:31:51number, right? So let's go ahead and run
- 2:31:52this.
- 2:31:54And let's again do a select star from
- 2:31:59the table.
- 2:32:04And
- 2:32:05there you go. So this completely works
- 2:32:08fine. Now it is able to accommodate big
- 2:32:10integers. Right? So that has completely
- 2:32:13worked fine. Now let's say if you want
- 2:32:15to add in additional attribute over here
- 2:32:19instead uh apart from the mall pin code
- 2:32:21and the store code you also want to add
- 2:32:24in the store location right so the way
- 2:32:27to do that again is very simple we run
- 2:32:29this statement
- 2:32:33and instead of alter column we just say
- 2:32:35add column and let's say this is the
- 2:32:37store location
- 2:32:40and this is going to be string Okay,
- 2:32:44let's run this.
- 2:32:46This works. And in order to test this
- 2:32:49out,
- 2:32:51let's run this statement again. And let
- 2:32:55me just put a simple number over here
- 2:32:56for now. And the store location is going
- 2:32:59to be the ground floor for example. So
- 2:33:03let's run this. Okay, I mistakenly
- 2:33:06inserted
- 2:33:08row number 17 again. because we don't
- 2:33:10have any uh primary key constraint in
- 2:33:12place this will work but ideally don't
- 2:33:14do it. So let's now have a look at
- 2:33:17select star from this table
- 2:33:22and
- 2:33:24we should see another row for 17 and
- 2:33:27this basically has the store code
- 2:33:29location equals null.
- 2:33:32The previous one actually this is the
- 2:33:33recent one. So the store location is the
- 2:33:35ground floor and you see that all the
- 2:33:37others have adjusted and the schema is
- 2:33:40maintained. The store location is null
- 2:33:42for all of the others. Right? So this
- 2:33:45shows that schema evolution inside of
- 2:33:49the strruct inside of nested structures
- 2:33:53also work pretty well. Let's understand
- 2:33:56automatic schema evolution inside of a
- 2:33:58nested structure
- 2:34:01with a few examples. And first let me
- 2:34:03ensure that this that automatic schema
- 2:34:07evolution property is set to true which
- 2:34:08is this one. So let me run this again.
- 2:34:13And
- 2:34:14what we want out of this example right
- 2:34:16what we want to understand out of this
- 2:34:18example is that initially I if if I
- 2:34:22wanted to add store location I had to
- 2:34:25run an alter statement over here. So if
- 2:34:28I want to add an attribute to a strct, I
- 2:34:32need to run an alter statement. Now with
- 2:34:35schema evolution, the expectation is
- 2:34:37that if there is a new attribute coming
- 2:34:39in, so let's say apart from store
- 2:34:41location, there is a new attribute
- 2:34:44called staff ID, the person who helped
- 2:34:46make the purchase is coming in. I
- 2:34:49shouldn't have to run an alter
- 2:34:50statement. It should be automatically
- 2:34:53accommodated inside of purchase detail.
- 2:34:56So let's try that out.
- 2:34:58Let me copy this insert statement and
- 2:35:01let me paste it over here. This will be
- 2:35:03for let's say customer number 21.
- 2:35:07Customer ID 21 and I will be using a
- 2:35:11named strruct. The reason why we are
- 2:35:14using a named strct is because earlier
- 2:35:16if you think of this we already know
- 2:35:18that the first value over here is going
- 2:35:21to be m pen code because it is there in
- 2:35:23the schema right? 765 is going to be
- 2:35:26store code and ground floor is going to
- 2:35:28be the store location because all of
- 2:35:30this information resides inside of the
- 2:35:33schema. But now in this case, if we were
- 2:35:36to randomly insert a value over here,
- 2:35:38which is staff ID, I I put in ST some
- 2:35:41value over here, it wouldn't know what
- 2:35:44key does it belong to, right? So I need
- 2:35:46to specify the key so that it is able to
- 2:35:48add that key over here. Adding only a
- 2:35:51value wouldn't work, right? So I hope
- 2:35:53you understand why we need to use a name
- 2:35:57struck. So I'll simply add
- 2:36:00the keys over here which is going to be
- 2:36:02all pin code.
- 2:36:05This is going to be the store code.
- 2:36:10This is going to be the store location.
- 2:36:16And the last one is going to be
- 2:36:21the staff ID.
- 2:36:24Okay. So, let's go ahead and run this.
- 2:36:31Perfect. So, this seemed to work fine.
- 2:36:33Now, let's quickly validate the data.
- 2:36:46And there you go. So, you see that staff
- 2:36:48ID has been inserted correctly. And we
- 2:36:52didn't have to run an alter statement to
- 2:36:54accommodate this inside of purchase
- 2:36:57detail. And as we we would expect,
- 2:36:59right? All of the other purchase details
- 2:37:02have staff ID equal to null. And this
- 2:37:04makes the whole schema very consistent.
- 2:37:07Let's finally understand the last
- 2:37:09scenario which is column position
- 2:37:12changes.
- 2:37:14So I'll add in a header
- 2:37:18which is scenario number four. And this
- 2:37:21is to be column position changes. And
- 2:37:23first of all, I want to
- 2:37:26set this to false because I don't want
- 2:37:28to use automatic schema evolution for
- 2:37:31now.
- 2:37:32And
- 2:37:34let's go and do things manually first.
- 2:37:36So in order to add a column in a
- 2:37:39particular order, it is going to be add
- 2:37:42columns.
- 2:37:44And actually let's figure out which
- 2:37:46column do we want to add first. Yeah. So
- 2:37:50we had
- 2:37:52we had these many columns and we don't
- 2:37:54have the age column yet in our table. So
- 2:37:59let me copy this this statement because
- 2:38:02we are going to need this soon.
- 2:38:07So let me paste this here for now. Let's
- 2:38:09also see how our table currently looks
- 2:38:12like. Select star from the table.
- 2:38:16And we want to add the age column. Where
- 2:38:19do we want to add it? So there are two
- 2:38:21options.
- 2:38:23Either you can add it as the first
- 2:38:24column by specifying first. So this
- 2:38:26becomes the column before customer ID.
- 2:38:29The second option is you add it after
- 2:38:34some column. So let's say you add it
- 2:38:36after price.
- 2:38:39So then it is going to look like this.
- 2:38:42So let's say we want to run the second
- 2:38:44statement. Let's go ahead and run this.
- 2:38:47And now let's see how this is going to
- 2:38:48look like. There you go. You see age has
- 2:38:52found its position after price because
- 2:38:54that is what we ran. Now let's run let's
- 2:38:59run an insert statement to see if all of
- 2:39:01this works or not. So again I'm going to
- 2:39:03insert five records only to keep this
- 2:39:05simple. And the column that we have is
- 2:39:08customer ID, invoice number, price, and
- 2:39:11then we have age. Then we have invoice
- 2:39:14date. Then we have quantity.
- 2:39:17Then we have payment method.
- 2:39:20And purchase detail was something that
- 2:39:22was added by us. So that doesn't reside
- 2:39:25in this park file. So let me make it
- 2:39:29null for simplicity. Null as purchase
- 2:39:31details. Do we have any other column?
- 2:39:33No. Let's run this.
- 2:39:38Now let's do a
- 2:39:42Okay, we'll just run this.
- 2:39:46So here you see
- 2:39:48that this has perfectly inserted the
- 2:39:51records from 50 to 55. All the columns
- 2:39:56have been populated correctly except the
- 2:39:58purchase detail column which we
- 2:40:00purposely set it as null. That means
- 2:40:03column ordering had just worked fine.
- 2:40:05Let's see how column position change is
- 2:40:08going to work with automatic schema
- 2:40:10evolution which simply means that if a
- 2:40:13new column comes in at whatever position
- 2:40:15right our table should be able to
- 2:40:17accommodate it. So let's say there is a
- 2:40:20new column called category which comes
- 2:40:22in and I want it to be before payment
- 2:40:25method it should be able to reflect that
- 2:40:28without me having to run an alter
- 2:40:30statement. Now again I want to highlight
- 2:40:32that maybe this kind of a scenario we
- 2:40:34would never want it to be in production
- 2:40:37right because we don't want some random
- 2:40:39data coming in and destroying our tables
- 2:40:42right destroying the structure of our
- 2:40:43table but for the sake of completeness I
- 2:40:46want to I want you to know every
- 2:40:48everything like all possible scenarios
- 2:40:51right for now let's insert category into
- 2:40:55our table
- 2:40:57so let me copy this insert statement and
- 2:41:02we're going to run this. And before
- 2:41:05this, let's also
- 2:41:09enable schema evolution.
- 2:41:13The schema evolution is being set to
- 2:41:16true.
- 2:41:18And let's also
- 2:41:21quickly print our
- 2:41:24table,
- 2:41:27right?
- 2:41:29So we see that we already have numbers
- 2:41:32from until 55.
- 2:41:35So this is going to be from 56 to 60.
- 2:41:39And the columns are customer ID, invoice
- 2:41:42number, price, age, invoice date,
- 2:41:45quantity, payment method, purchase
- 2:41:48details and let me add category over
- 2:41:51here. Right? So ideally category should
- 2:41:55just come before purchase details if at
- 2:41:58all this works right. So let's go ahead
- 2:42:01and run this
- 2:42:04and this didn't work. The error says
- 2:42:06cannot resolve category due to data type
- 2:42:08mismatch. Okay I don't want to correct
- 2:42:11and run this again. Uh data type
- 2:42:13mismatch cannot cast string to strct
- 2:42:15data type. So an important thing to
- 2:42:17understand here is that
- 2:42:20ideally this is the set of all of my
- 2:42:22columns right and this is an extra
- 2:42:26column that I just added. So if you were
- 2:42:29to just remove this for now and if you
- 2:42:31were to think that this is your purchase
- 2:42:32detail
- 2:42:34this is your purchase detail then all of
- 2:42:37the column should be completed with this
- 2:42:40right this would be an extra column. So
- 2:42:42what the insert statement is throwing as
- 2:42:44an error is that this category over here
- 2:42:46is supposed to be the last column and
- 2:42:48the last column currently is a struck
- 2:42:50data type. But what you're giving me is
- 2:42:53a string data type. And for that reason
- 2:42:56I'm throwing an error. So that's the
- 2:42:59cause of this error. And now let's go
- 2:43:01ahead and try this with a merge
- 2:43:03statement. Let's see what happens with
- 2:43:04the merge statement. So we say merge
- 2:43:08into delta catalog. This is going to be
- 2:43:11the target
- 2:43:13using
- 2:43:15this table right here
- 2:43:20as the source
- 2:43:22on target dot
- 2:43:26customer ID equals source ID source dot
- 2:43:29customer ID where
- 2:43:32when not
- 2:43:35matched
- 2:43:37then insert
- 2:43:40star right so here the code is exactly
- 2:43:44the same category is not in the correct
- 2:43:46position and category is added and then
- 2:43:49extra column over here so let's see what
- 2:43:51happened
- 2:43:57okay that seemed to work and we are
- 2:44:01going to
- 2:44:03we are going to run this
- 2:44:07and okay you don't see category column
- 2:44:10before purchase details but you see the
- 2:44:13category column after purchase details
- 2:44:16and this is quite interesting right so
- 2:44:19we read that merge does a matching by
- 2:44:23name so it was able to match all of
- 2:44:26these column all of this over here and
- 2:44:28purchase details by name so it inserted
- 2:44:32them at the right places the only thing
- 2:44:34that that that it was not able to match
- 2:44:36was category and it inserted it as the
- 2:44:40last column. So the positioning didn't
- 2:44:42work but you still ended up putting the
- 2:44:45data inside of the table. Right now let
- 2:44:48me take you through how schema evolution
- 2:44:50is going to look like with spark data
- 2:44:53frame. Yeah. And the way we going to do
- 2:44:56that is by quickly taking some of the
- 2:45:01parket data that we read above. Right?
- 2:45:04So I'm going to take this one this park
- 2:45:07a files
- 2:45:09and
- 2:45:11we going to do a from pispark.sql.f
- 2:45:15function import star.
- 2:45:17Let's filter
- 2:45:20the customer ids.
- 2:45:24Customer ID dot between 1, 10. And let's
- 2:45:28also select
- 2:45:31a few columns, right? Customer ID,
- 2:45:35price,
- 2:45:37invoice, date.
- 2:45:41And let's go ahead and write this
- 2:45:43df.right write dot save as table. This
- 2:45:48is going to be in delta catalog dot
- 2:45:50deltadb dot
- 2:45:54schema invoices schema evolution sparkd
- 2:45:58right so let's go ahead and write this
- 2:46:00okay this is go this should be post
- 2:46:03python
- 2:46:07and if everything works we should see a
- 2:46:11table here inside of our catalog and
- 2:46:15there you go so you We have the table.
- 2:46:18We have the table that we created. Now
- 2:46:20let me add
- 2:46:22let me add a few more column. Right. Let
- 2:46:26me add quantity payment method.
- 2:46:30Quantity and payment
- 2:46:34method.
- 2:46:36We added quantity and payment method
- 2:46:38which changes the schema
- 2:46:42of the table. Right? So I mean we have
- 2:46:45not written it but the schema is
- 2:46:47different from what we had over here. So
- 2:46:49let me choose customer ids from 11 to 25
- 2:46:58and let's
- 2:47:01write dot mode append
- 2:47:06dot option
- 2:47:12merge schema
- 2:47:16to true.
- 2:47:19Let's go ahead and run this now.
- 2:47:25And now let's see how our table looks
- 2:47:27like.
- 2:47:32Okay. So you see that from 11 to 25 we
- 2:47:38have the data for quantity and payment
- 2:47:41method. But from 1 to 10 we don't have
- 2:47:43data for quantity and payment method.
- 2:47:46Right? And this is because
- 2:47:49the schema has evolved due to writing
- 2:47:51merge schema equals true. So now we are
- 2:47:54going to understand how to convert park
- 2:47:57files into the delta format. So most of
- 2:48:01the times your input files or your
- 2:48:03existing files are not going to be in
- 2:48:06the delta format. But in order to be
- 2:48:08able to use the awesome features that
- 2:48:11we've been talking about, we need to
- 2:48:13convert it to delta. Yeah. So that is
- 2:48:15what we're going to understand through a
- 2:48:17lot of examples. So I'm going to use the
- 2:48:19same paret file that I've used earlier.
- 2:48:21So I'll just copy the path of this
- 2:48:24parket file. And first of all, let's
- 2:48:26quickly do a percent fsls and see what's
- 2:48:29there inside of the invoices folder.
- 2:48:32And we see that there are three paret
- 2:48:35files. And let me go ahead and use this
- 2:48:38one. The customer ID is from 1 to 100.
- 2:48:41So
- 2:48:43let me first create copies two copies of
- 2:48:45this.
- 2:48:46I'm creating two copies because I want
- 2:48:48to show you two ways of converting from
- 2:48:50park a to delta. So spark read.park
- 2:48:54and let me
- 2:48:57do a mode overrite
- 2:49:00dot
- 2:49:02park. And let's paste the path. This is
- 2:49:06going to be let's name it as v_sub_1 and
- 2:49:08let's not keep it in the invoices folder
- 2:49:14and this is going to be v2. So basically
- 2:49:16where it is going to reside is inside of
- 2:49:19the lab data container at the root.
- 2:49:22Yeah. So here if I go back. So here is
- 2:49:25my lab data container and it is going to
- 2:49:27reside over here. Yeah. So let's go
- 2:49:29ahead and run this now.
- 2:49:33So now I should see two folders. So
- 2:49:35these are my two folders, right? So
- 2:49:37these are paret files and if you see
- 2:49:40here it has a snappy.park and this is
- 2:49:44where the data is stored and it doesn't
- 2:49:46have a delta log folder. That means this
- 2:49:48is not a delta table yet. Same applies
- 2:49:51for this one as well. Yeah. So now let's
- 2:49:54go ahead and
- 2:49:56try to convert this to delta. to convert
- 2:49:59to
- 2:50:01delta and then I write park dot let me
- 2:50:05paste the path.
- 2:50:10So now let's go ahead and run this.
- 2:50:14Okay, this is complete.
- 2:50:17So there you go. You see a delta log
- 2:50:20folder which is going to record the
- 2:50:22transactions. And let's go ahead and see
- 2:50:25what's inside of JSON. And there is an
- 2:50:29add operation which added this park
- 2:50:32file. So it registered this transaction.
- 2:50:34Right? Now let's try another method.
- 2:50:37Let's say you want to use the
- 2:50:40delta table API.
- 2:50:43Import
- 2:50:44delta table. And the method that we
- 2:50:47going to use is convert to delta.
- 2:50:50We pass over here. And this is going to
- 2:50:53be the format is park dot.
- 2:50:56We specify the tildas over here. And
- 2:51:00we copy this path v2 because we want to
- 2:51:03convert v2. And let's go ahead and run
- 2:51:07this. Okay, this is complete.
- 2:51:10And let's have a look at this now. So
- 2:51:14you see the delta log over here as well.
- 2:51:17And let's check this. And this should
- 2:51:20also have an add operation which
- 2:51:23registered the data. Right? So this is
- 2:51:26how you're able to convert a parquet
- 2:51:29file to a delta table right and
- 2:51:32similarly you can do it for a CSV file
- 2:51:35or a JSON file or anything else right
- 2:51:37maybe the approach would be a little
- 2:51:39different you would have to read the CSV
- 2:51:41file into a data frame something like df
- 2:51:45equals spark read dot csv and then maybe
- 2:51:49the path over here and you have to write
- 2:51:51it
- 2:51:53let's say the mode is
- 2:51:56overrite
- 2:51:58and then the format is going to be delta
- 2:52:02and then you specify the path over here
- 2:52:05whatever path you want it to be right so
- 2:52:08this is another approach that you can
- 2:52:10follow in order to write your CSV or any
- 2:52:13other format of files now let's
- 2:52:15understand what are manage and external
- 2:52:18table actually the table that we've
- 2:52:21created in most of our examples they are
- 2:52:24manage tables but let's understand them
- 2:52:26in a lot more details now. So managed
- 2:52:29tables are those tables where your both
- 2:52:32your data and the meta data
- 2:52:36they are managed by delta lake
- 2:52:41right they are managed by delta lake and
- 2:52:44we create delta lake or delta tables
- 2:52:48right we create delta tables using the
- 2:52:52unity catalog. Now where does unity
- 2:52:54catalog store data? Unity catalog stores
- 2:52:56data inside of a storage account inside
- 2:53:00of this container called metas store.
- 2:53:04Right? So that is where it creates your
- 2:53:06tables. So the tables that we create
- 2:53:09they are stored and managed inside of
- 2:53:13this container meta store right is
- 2:53:15stored in a manage location. And in case
- 2:53:19of an external table both the data
- 2:53:22actually the data is stored in a user
- 2:53:27specified location. It can be an S3
- 2:53:30bucket or an ADLS gen 2 container or
- 2:53:34Google cloud storage. Right? But the
- 2:53:36meta data
- 2:53:38resides in the meta store of the Unity
- 2:53:43catalog. And the meta store again could
- 2:53:45be something like this. Right? So the
- 2:53:47metadata is going to reside in the meta
- 2:53:49store of the Unity catalog and the data
- 2:53:53is going to reside in a location of
- 2:53:56users choice. So now let's see this in
- 2:53:58action. Let's see an example. If I were
- 2:54:01to pull up
- 2:54:03any one of the table that I created
- 2:54:05earlier,
- 2:54:08we see that let's pick up invoices SC.
- 2:54:12Let's go to details. And what you see
- 2:54:15over here is the storage location,
- 2:54:17right? So the storage location is this
- 2:54:19storage account inside of this
- 2:54:21container. The container is called meta
- 2:54:23store which is the one over here. And
- 2:54:28there we're going to have tables and
- 2:54:32this is the unique identifier of my
- 2:54:35table. So my table is going to be this
- 2:54:37one. And all of the data and metadata is
- 2:54:41stored over here. is managed by the
- 2:54:44Unity catalog. That's one. Now, let's
- 2:54:48try to create an external table. So,
- 2:54:52let's write some SQL. Create or replace
- 2:54:57replace table. And this is going to be
- 2:55:00delta catalog dot delta DB dot invoices
- 2:55:05external.
- 2:55:06So I'm going to
- 2:55:09be using delta because the output the
- 2:55:12underlying storage that I want is
- 2:55:14supposed to be in a delta format. Right?
- 2:55:16Now this can be any format but I
- 2:55:18specifically want it to be in delta. So
- 2:55:21that is why I say using delta
- 2:55:24and then I specify what is the location
- 2:55:26where I want to store my data.
- 2:55:29Let's go ahead and use this. Let's put
- 2:55:32it in some other location. This time not
- 2:55:34the meta store. Let's put it inside lab
- 2:55:37data.
- 2:55:40And so I have lab data over here and let
- 2:55:44me name this as invoices external
- 2:55:47and this is going to be as select star
- 2:55:50from
- 2:55:52park
- 2:55:56dot
- 2:55:58let's say this file over here
- 2:56:03right
- 2:56:05so let me paste this over here and let's
- 2:56:09run this
- 2:56:16So let's see if this table exists in the
- 2:56:18catalog. So there is a table invoices
- 2:56:21external and let's see the details. Now
- 2:56:23the storage location is not in the meta
- 2:56:25store. It is the location that we
- 2:56:28specified. Right? So let's refresh this
- 2:56:32and this is invoice ext. And you have
- 2:56:35this as a delta table. Right? Now the
- 2:56:38second difference is if I do a drop on a
- 2:56:42manage table the data the underlying
- 2:56:44data is going to vanish. It is going to
- 2:56:46go away. But in case of an external
- 2:56:48table this data that you see over here
- 2:56:51even if I do a drop on this table right.
- 2:56:54So if I do something like a
- 2:56:58drop table
- 2:57:01the underlying data is still going to
- 2:57:03stay in the user specified location.
- 2:57:07Right? So let me go ahead and
- 2:57:10run this
- 2:57:12now. This should disappear from the
- 2:57:13catalog. It has disappeared indeed. And
- 2:57:15let me go ahead and refresh this.
- 2:57:21You see that the data is still there in
- 2:57:24the user specified location. And if this
- 2:57:26were a manage table, the data would have
- 2:57:29gone. So actually sometimes in cases of
- 2:57:32manage table you would still see the
- 2:57:34data because of it default retention
- 2:57:36period. Sometime the retention period is
- 2:57:38set to 30 days. So that is the reason
- 2:57:41why the data doesn't disappear
- 2:57:43immediately. Now the third point of
- 2:57:46difference is that manage tables use the
- 2:57:49delta format only. So all of the table
- 2:57:51that are created in the unity catalog
- 2:57:54inside residing inside of this meta
- 2:57:56store follow the delta format only. The
- 2:58:00manage tables follow the delta format
- 2:58:03only. But in case of external table they
- 2:58:07they could follow several formats. It
- 2:58:09can be CSV, it can be JSON, it can be a
- 2:58:12park,
- 2:58:15right? It can even be delta.
- 2:58:18So in the last example we saw that we've
- 2:58:23written using delta
- 2:58:26and because we wrote using delta it
- 2:58:29created the output in a delta format
- 2:58:31right if we wrote using par
- 2:58:36it would create the output in a park
- 2:58:39file format right so that's the reason
- 2:58:41why several formats are supported for
- 2:58:44external table so now let's talk about
- 2:58:46something really interesting thing in
- 2:58:48Delta Lake called deletion vectors and I
- 2:58:52believe we've already seen it in action
- 2:58:54but now let's talk about it in a lot
- 2:58:57more detail. So first of all let's
- 2:59:00understand what is the problem what is
- 2:59:03the issue at hand that deletion vectors
- 2:59:06is trying to solve right so we already
- 2:59:10know that delta uses paret files under
- 2:59:13the hood for storing its data right and
- 2:59:18park files are immutable. So what I mean
- 2:59:20by immutable is that they cannot be
- 2:59:23modified directly. Right? So if you want
- 2:59:26to update or delete a record in a park
- 2:59:29file, what you essentially need to do is
- 2:59:32you read the park file and then you
- 2:59:35apply those operation delete or an
- 2:59:37update and then you write the park file
- 2:59:40back. Right? So that is the only way you
- 2:59:42can update or make changes to that file.
- 2:59:46Right? You read it, apply the operations
- 2:59:49and then write it back. Now this rewrite
- 2:59:53process, this whole read and a rewrite
- 2:59:55process is very costly. Imagine if you
- 2:59:59had a parquet file with 10 million rows
- 3:00:02and you only needed to delete a few
- 3:00:04couple of rows, right? So what you would
- 3:00:07end up doing is you would read the whole
- 3:00:09park files and you would rewrite the
- 3:00:12whole park file again except for those
- 3:00:15couple of records which you want to
- 3:00:17delete. Yeah. So this is a tremendously
- 3:00:23computationally expensive operation and
- 3:00:25this is where deletion vectors come into
- 3:00:28play. Now before we talk about deletion
- 3:00:31vectors, let's talk about two important
- 3:00:33concepts. The first one is copy on write
- 3:00:37and the second one is merge on read.
- 3:00:41Yeah. So let's understand copy on write
- 3:00:43first. Copy on write is exactly the same
- 3:00:47as what we discussed above. Right. Every
- 3:00:50change, every update, every delete or a
- 3:00:54merge is going to create a new file.
- 3:00:57Right? And Delta is going to follow this
- 3:00:59approach when deletion vectors are
- 3:01:03disabled. So it's very important to keep
- 3:01:05in mind that copy on write is going to
- 3:01:08be followed when deletion vectors are
- 3:01:11disabled. Yeah. So let's take this
- 3:01:14example. Let's say we have a paret file
- 3:01:171.park and it contains,000 records,
- 3:01:21right? It contains 1,000 records and
- 3:01:25that is what we see over here from one
- 3:01:27until,000. And now what we want to do is
- 3:01:30we just want to delete rows number one
- 3:01:33and six. So the most simple way that
- 3:01:37copy on right is going to do is that
- 3:01:39take up this whole parquet file. Right?
- 3:01:42You take up this whole park file and
- 3:01:45then apply the delete operation.
- 3:01:49You apply the delete operation and then
- 3:01:51you rewrite the whole file back except
- 3:01:55rows number one and rows number six. So
- 3:01:58these two rows are going to be omitted
- 3:02:01and then the whole file is going to be
- 3:02:04written back as you see over here. So it
- 3:02:06doesn't have row number one and row
- 3:02:08number six over here. Right? So the new
- 3:02:12state that you see is that a new file
- 3:02:14which is 2.par is written to storage.
- 3:02:18Right? And this file is going to be the
- 3:02:21one to be referenced for the latest
- 3:02:24version. So that is how copy on write
- 3:02:27works. Now in contrast what merge on
- 3:02:31read does is that it allows your
- 3:02:34original park files to remain untouched
- 3:02:38and instead the changes the changes that
- 3:02:41we apply right for example deletions
- 3:02:44they are recorded in a separate file
- 3:02:47known as a deletion vector. Yeah. So
- 3:02:50let's say if I want to delete a record
- 3:02:52that record that row number is recorded
- 3:02:56in a deletion vector and when I read the
- 3:02:58file the deletion vector is simply going
- 3:03:00to be checked if this row is present in
- 3:03:03the deletion vector or not. If it is
- 3:03:06then that row is going to be skipped
- 3:03:08from the output. Yeah. So delta is going
- 3:03:12to follow the merge on read approach if
- 3:03:15deletion vectors are enabled. Yeah. So
- 3:03:18it's really important to keep this mind.
- 3:03:19Keep this in mind again. If deletion
- 3:03:21vectors are enabled, merge on read is
- 3:03:24going to be followed. If not, copy on
- 3:03:26write is going to be followed. So now
- 3:03:28let's understand this with an example.
- 3:03:31Yeah. So let's say we have the same park
- 3:03:35file with th00and records. Yeah. So 1
- 3:03:38th00and records as you see over here
- 3:03:40from one until,000 over here. And we
- 3:03:43want to perform the same operation. We
- 3:03:47want to delete rows number one and six.
- 3:03:50Yeah. So this time what happens is when
- 3:03:53we say we want to perform a delete of
- 3:03:56row number one which is right over here,
- 3:03:59we simply record it in a deletion vector
- 3:04:03which is over here. Right? The next time
- 3:04:06we say that okay we want to delete row
- 3:04:08number six. What happens is we record
- 3:04:12row number six in the deletion vector
- 3:04:14again. Right? So all of these changes
- 3:04:17are getting recorded in the deletion
- 3:04:19vector and all of this data that you
- 3:04:22have over here they are stored in
- 3:04:241.park. So now let's say I want to read
- 3:04:28the file right I want to read the latest
- 3:04:30state. So when I read the latest state I
- 3:04:34would expect that row number one and row
- 3:04:37number six shouldn't be there in the
- 3:04:39file. Right? So the way I'm going to get
- 3:04:42the output is that this whole file is
- 3:04:46going to be produced as output and this
- 3:04:49row is going to be checked against the
- 3:04:51deletion vector that is it present in
- 3:04:53the deletion vector or not. If it is
- 3:04:56then this is marked as a soft delete and
- 3:05:00you won't see that row. So this won't be
- 3:05:03produced in the output. Similarly this
- 3:05:05row is going to be checked is it present
- 3:05:08in the deletion vector or not. It's not
- 3:05:09present. So that means this row is going
- 3:05:11to be displayed. Similarly for all the
- 3:05:13rows and row six would also be checked
- 3:05:16and it would not be displayed. Right? So
- 3:05:19that is how it helps you in not
- 3:05:24rewriting the entire file back. The
- 3:05:26final state that you would have is
- 3:05:281.park park the whole file is untouched
- 3:05:31and a small deletion vector bit mapap
- 3:05:34file right where all of the deletes all
- 3:05:38of the changes are getting recorded and
- 3:05:40it is going to be checked against and
- 3:05:43this is going to add a lot of speed to
- 3:05:47the entire process and it is going to
- 3:05:48make a lot of things more performant. So
- 3:05:51we see that deletion vectors bring the
- 3:05:54merge on read capability to delta lake
- 3:05:57thereby increasing the performance of
- 3:05:59deletes updates. So update is considered
- 3:06:02as an insert plus delete and merge
- 3:06:06operations. Right? So instead of
- 3:06:08rewriting the whole file back for small
- 3:06:11small changes what it does is that it
- 3:06:13records those small changes in a
- 3:06:15separate bit map file known as deletion
- 3:06:18vectors. Let's see all of this in action
- 3:06:20now. Let's get started with our labs.
- 3:06:23So, I'm going to quickly pick up one of
- 3:06:26the old parket file that we've been
- 3:06:28using. And let me quickly do a select
- 3:06:32star from parket.
- 3:06:34And then this one, let me use the 2010
- 3:06:39to 2011. Right? And let's keep the
- 3:06:43default as SQL because we're going to
- 3:06:45write lots of SQL. Let's connect this
- 3:06:49cluster and let me quickly do a limit
- 3:06:52five and let's go ahead and run this.
- 3:06:55Meanwhile,
- 3:06:57first we are going to explode copy on
- 3:07:00right. So let me quickly write how to
- 3:07:04create a table with deletion vector
- 3:07:09disabled.
- 3:07:11Right. So let's see how this works.
- 3:07:16Great. So let me do create
- 3:07:20or replace table
- 3:07:24and
- 3:07:26the catalog that we're using is this one
- 3:07:30delta catalog dot delta db right so this
- 3:07:34is going to be delta catalog do delta db
- 3:07:40dot invoices
- 3:07:42and this is going to be
- 3:07:46D copy on right right yeah and let me
- 3:07:50quickly copy this over here in the table
- 3:07:52properties
- 3:07:54delta dot enable deletion vector to be
- 3:07:58false and this is going to be created as
- 3:08:02a c dash creatable at select statement.
- 3:08:06So let's go ahead and run this now.
- 3:08:10Yeah, let me also write a describe
- 3:08:14extended
- 3:08:16and then
- 3:08:18delta
- 3:08:20catalog dot
- 3:08:23delta db dot invoices.
- 3:08:27Yeah. So let's run this and have a look
- 3:08:31at some of the properties quickly.
- 3:08:38So here we see that this is a manage
- 3:08:42table and the deletion vector is not
- 3:08:46enabled. Yeah. Let's also see where
- 3:08:50where does this table reside inside of
- 3:08:53ADLS.
- 3:08:57So let me quickly open up ADLS.
- 3:09:12And this is the table right here.
- 3:09:16So it basically contains one paret file
- 3:09:19at this moment. Right? Now let's go
- 3:09:22ahead and perform some operation. Right?
- 3:09:25So let's say this user 105
- 3:09:28uh from age 57 let's say it was
- 3:09:31incorrectly recorded and we want to
- 3:09:34correct the age to 55 for user 105. So
- 3:09:38we are simply going to say update
- 3:09:43this table
- 3:09:45set
- 3:09:47age equals 55 where customer ID equals
- 3:09:52105. Right? So this customer is 105
- 3:09:57right here. So let's go ahead and run
- 3:10:00this.
- 3:10:05And let's also see the history.
- 3:10:08Describe history
- 3:10:11and this table.
- 3:10:13So it should contain two rows. The first
- 3:10:15one for the create. Yeah. The first one
- 3:10:18for the create and the second one for
- 3:10:20the update. Right. So based on whatever
- 3:10:23we've studied and understood right now,
- 3:10:26this update should create a fresh par
- 3:10:29file right with the changes. So let's go
- 3:10:33ahead and understand a few metric the
- 3:10:36operation metrics over here right. So
- 3:10:38what it says is that number of removed
- 3:10:40files equals 1. Number of copied rows
- 3:10:44equals 99. So we see that we had a total
- 3:10:46of 100 rows over here right from here.
- 3:10:49And the row that got updated is just one
- 3:10:53row which is row number 105. So other
- 3:10:56than that it copied all of the 99 rows.
- 3:11:00No deletion vectors added or removed
- 3:11:02because deletion vectors are disabled
- 3:11:06number of added files. So it added a new
- 3:11:08file right and then it updated one row.
- 3:11:12Quite similar to what we would expect,
- 3:11:14right? Because we ran an update
- 3:11:15statement over here.
- 3:11:18Now let's go ahead and see how this
- 3:11:20would look in the storage.
- 3:11:24So we see that this has added another
- 3:11:26file. It has rewritten another file.
- 3:11:28Yeah. And if you were to quickly look at
- 3:11:31the sizes,
- 3:11:33the sizes is more or less the same.
- 3:11:36Right? This is 59 to1 and this is also
- 3:11:3959 to1. Yeah. Now let's go ahead and run
- 3:11:42another another statement. This time
- 3:11:45let's go ahead and run a delete
- 3:11:47statement. So let's go ahead and delete
- 3:11:50the row for customer ID equals 102.
- 3:11:54So I'm going to write delete
- 3:11:57from this table where customer ID equals
- 3:12:00102. Right? So let's go ahead and run
- 3:12:02this and let me also see
- 3:12:05what the history is going to show me.
- 3:12:11Now we have three rows. the third one
- 3:12:13for the delete operation and let's
- 3:12:15quickly see the operation matrix. So now
- 3:12:19the operation matrix what it says is
- 3:12:22number of copied rows is 99.
- 3:12:26Number of deleted rows is one. Yeah. And
- 3:12:30then it added all of this in a new file.
- 3:12:33Right? So it basically removed the file
- 3:12:36that it read in the previous version.
- 3:12:38That is why you see number of removed
- 3:12:39files one. And then it added this new
- 3:12:42file. It deleted one row over here
- 3:12:45because number of deleted rows is equal
- 3:12:46to one. And let's have a look at the
- 3:12:48size. So this is 5897
- 3:12:52and this was 5921. A little smaller
- 3:12:56because we deleted one record. Now if
- 3:12:58you have a look over here, we should see
- 3:13:00one file over here. Yeah,
- 3:13:04there you go. So we see this file which
- 3:13:07is which was added at 1639.
- 3:13:10This file was added with 99 records.
- 3:13:14Yeah. So I believe this confirms how
- 3:13:18copy on write works. It is reading the
- 3:13:21file applying the changes and then
- 3:13:23writing it as a new file. So let's now
- 3:13:26see how is a delta table going to behave
- 3:13:29if merge on read was enabled. Yeah. So
- 3:13:34we've already discussed that if deletion
- 3:13:37vectors are enabled then the paradigm
- 3:13:40that is going to be followed is merge on
- 3:13:43read. So let me create this table and I
- 3:13:47will rename this to m and let's set this
- 3:13:51to true which is going to enable
- 3:13:54deletion vectors and let's go ahead and
- 3:13:56run this. I will do a describe extended
- 3:14:04this table just to verify a few of the
- 3:14:06properties. So this is a manage table
- 3:14:10and delta.enable deletion vectors is set
- 3:14:14to true. Now let's go ahead and perform
- 3:14:18similar operations that we performed on
- 3:14:21the earlier table, right? The table
- 3:14:23where deletion vectors was disabled.
- 3:14:27So let's go ahead and run this. And this
- 3:14:30is going to be M.
- 3:14:33And I'm going to do a describe history
- 3:14:37of
- 3:14:39this table.
- 3:14:41And I should see two rows. Yeah. The
- 3:14:43first one for the CAS and the second one
- 3:14:46for delete. So let's understand
- 3:14:50a few operation metrics. Right? So what
- 3:14:54it says is that number of removed files
- 3:14:57and the number of removed bytes is zero.
- 3:15:00And this is a little different from what
- 3:15:03we saw earlier.
- 3:15:06What we saw earlier was
- 3:15:09it removed the file and then it
- 3:15:11performed the delete operation and then
- 3:15:13it wrote the file back. Yeah. But here
- 3:15:16what we see is a little different.
- 3:15:19There were no files that were removed
- 3:15:21and there were no rows that were copied.
- 3:15:23Instead, there is a deletion vector that
- 3:15:26has been added. Yeah. And the number of
- 3:15:30deleted rows is one, which is exactly
- 3:15:33what we did over here. Yeah.
- 3:15:36So that's all of the operation metrics,
- 3:15:39right? Now let's quickly see how does it
- 3:15:42look like
- 3:15:44in the storage in our storage layer.
- 3:15:47Yeah. So, let me copy the file
- 3:15:51identifier
- 3:15:53and this file is right here. And there
- 3:15:55you see that this is the original file.
- 3:15:59We don't have another new file as you
- 3:16:02would see in a copy on write paradigm.
- 3:16:06We only have the old file and then we
- 3:16:08have a deletion vector which records the
- 3:16:11rows which we need to delete. Yeah, very
- 3:16:14simple and very performant because it
- 3:16:16didn't have to rewrite the whole file.
- 3:16:18Now let's go ahead and perform another
- 3:16:22operation which is the update.
- 3:16:26So I'll just change the name of the
- 3:16:28table and the operation will be just the
- 3:16:32same and let me do a describe history
- 3:16:38table and I should see three rows. Yeah,
- 3:16:42the third one is an update. And let's
- 3:16:44quickly have a look over here. So the
- 3:16:47number of removed files is zero. Number
- 3:16:48of removed bytes, number of copied rows
- 3:16:50is all zero. What it did is it added a
- 3:16:54deletion vector. Yeah, it removed the
- 3:16:57old deletion vector and then it added a
- 3:17:00new deletion vector. Right? So the new
- 3:17:02deletion vector has been updated with
- 3:17:05more details and that is what we would
- 3:17:08see over here. So if I were to look at
- 3:17:12quickly refresh this what what should we
- 3:17:14expect? So we should expect two things.
- 3:17:16The first one is that I should expect a
- 3:17:18park file because as we've read in the
- 3:17:21earlier section an update is considered
- 3:17:24as a delete plus insert. Where is the
- 3:17:28delete going to be recorded? The delete
- 3:17:30is going to be recorded in the updated
- 3:17:32deletion vector. So one more deletion
- 3:17:35vector should be created and the row
- 3:17:37that is updated that should be recorded
- 3:17:40in another paret file. Yeah. So I should
- 3:17:43see one more paret file and I should see
- 3:17:45an updated deletion vector that is one
- 3:17:48more deletion vector. So let's refresh
- 3:17:50this and there you go. So you see one
- 3:17:53more parket file and you see one more
- 3:17:57deletion vector and this is the updated
- 3:18:00deletion vector. So now you may have
- 3:18:02this question that on a table there
- 3:18:06could be several deletes, updates and
- 3:18:08merge operation that could be going on
- 3:18:10right and as a result of this as you see
- 3:18:12over here it could generate a lot of
- 3:18:16small files lot of deletion vectors
- 3:18:18right so what is the way to clean this
- 3:18:21up right to tidy this up because there
- 3:18:24are two deletion vectors over here one
- 3:18:26is obsolete the other one is updated and
- 3:18:29then there is this one file over here
- 3:18:31which just contains one row. Right? So
- 3:18:34for all of this, datab bricks provides
- 3:18:37us the optimize command. So you have the
- 3:18:41optimize command that you can apply on a
- 3:18:44delta table. So let me go ahead and
- 3:18:48quickly show you. So let's say we run
- 3:18:51the optimize on this table.
- 3:18:56And what it does is that it is going to
- 3:18:58create a new version and that new
- 3:19:02version is going to be the latest state
- 3:19:04with all of the deletion vector and all
- 3:19:07of those computations applied. Right? So
- 3:19:09it is going to check which row it's
- 3:19:11supposed to be there in the latest
- 3:19:12version. Do all of those computations
- 3:19:14and then create one fresh paret file.
- 3:19:17Right? So over here I should just see
- 3:19:20one fresh par file created using all of
- 3:19:23this. Right? it is going to apply all of
- 3:19:25the computation deductions and then
- 3:19:27create one new park file. So now what we
- 3:19:30see over here that this new park file
- 3:19:33has been added and this has been created
- 3:19:36by computing the latest state. Now if I
- 3:19:38were to do a vacuum yeah so vacuum
- 3:19:41basically keep the latest state. It is
- 3:19:43going to remove all of the history. It
- 3:19:46is only going to keep this final file,
- 3:19:49right? Because this final file contains
- 3:19:51the latest state. So let's go ahead and
- 3:19:55run
- 3:19:57vacuum this table and then retain zero
- 3:20:01hours. And then we also need to set a
- 3:20:04configuration which is called
- 3:20:07set spark.ta
- 3:20:10uh spark.databicks.delta
- 3:20:13dot retention
- 3:20:15duration check.enable is false. And let
- 3:20:17me go ahead and run this. So they're
- 3:20:19going to clean up all of the older
- 3:20:21versions and what we should end up
- 3:20:24seeing it just one park file which is
- 3:20:27the one created at 1814. Yeah.
- 3:20:34So let me refresh this and there you go.
- 3:20:36So you see that the the file that was
- 3:20:39created at 1814 was the latest file
- 3:20:43right and this has been retained when we
- 3:20:46run the vacuum command. So by now we've
- 3:20:49seen deletion vectors, copy on write,
- 3:20:52merge on read, all of it in action.
- 3:20:54Yeah. So it's really important to
- 3:20:57understand that neither paradigm,
- 3:20:59neither copy on write nor merge on read
- 3:21:03offer a silver bullet for all use cases.
- 3:21:06Yeah. However, copy on write is really
- 3:21:09good for use cases which is read heavy.
- 3:21:12Yeah. And where the rights are very low.
- 3:21:16What will happen if the rights are high?
- 3:21:18It will simply read and rewrite the
- 3:21:21whole file again and again and again,
- 3:21:23right? And we want to avoid doing that
- 3:21:25because it's a costly operation. Yeah.
- 3:21:27So, copy and write is good for read
- 3:21:30heavy use cases. And similarly, merge on
- 3:21:34read works best for cases for use cases
- 3:21:38where data is updated frequently. Right?
- 3:21:42Why does it work best for use cases
- 3:21:44where there are frequent updates? That
- 3:21:47is because it is simply going to record
- 3:21:49those changes in a deletion vector file.
- 3:21:52It is not going to rewrite the whole
- 3:21:55file. So ultimately it helps you reduce
- 3:21:58the right latency because it is simply
- 3:22:00recording it in a deletion vector. Yeah.
- 3:22:03So I hope that helps you understand the
- 3:22:06tradeoffs between the two and how
- 3:22:08deletion vectors can help boost
- 3:22:11performance. So now we are going to talk
- 3:22:14about cloning. So cloning as you know
- 3:22:18let us create a snapshot of our delta
- 3:22:21tables at a specific point in time.
- 3:22:24Right? And there are two flavors to
- 3:22:27cloning. There are two approaches to
- 3:22:30clone a delta table. The first one is a
- 3:22:33shallow clone and the second one is a
- 3:22:37deep clone. So let's understand both of
- 3:22:39them in details. Shallow clone is
- 3:22:42essentially taking the snapshot of
- 3:22:45metadata at a particular point in time,
- 3:22:47right? So when you shallow clone a
- 3:22:49table, delta doesn't copy the underlying
- 3:22:53data files, right? It simply references
- 3:22:56them. So let's say there is a source
- 3:22:58table and then there is a shallow clone
- 3:23:00table. The shallow clone table is simply
- 3:23:02going to reference the paret or the
- 3:23:05underlying data files, right? It is not
- 3:23:08going to copy those data file. So what
- 3:23:11this means is that it makes the shallow
- 3:23:14clone operation super fast,
- 3:23:16computationally cheap and it doesn't eat
- 3:23:19up a lot of storage. So let's understand
- 3:23:22this with this diagram right here. Let's
- 3:23:25say we have a source delta table and we
- 3:23:28want to shallow clone this source delta
- 3:23:30table into a target table. So what the
- 3:23:33source table has is this metadata over
- 3:23:36here in the delta
- 3:23:39log folder and this is some bunch of
- 3:23:42JSON files that you see over here right
- 3:23:45and then it also has some data in the
- 3:23:48form of park a file. Yeah. So now if we
- 3:23:53want to shallow clone,
- 3:23:56if we want to shallow clone this table,
- 3:23:58the first thing that is going to happen
- 3:24:01is a duplication or a replication of the
- 3:24:05metadata itself, the metadata that we
- 3:24:07have over here. And when I say
- 3:24:08duplication, what I mean to say is that
- 3:24:11it's not a one toone copy of these JSON
- 3:24:14files, right? It's not that I'm going to
- 3:24:15copy 0.json 1.json and then put it over
- 3:24:18there, right? put it in the source in
- 3:24:20the target table. Right? So that's
- 3:24:22that's not the purpose. What is going to
- 3:24:24happen is that all of this is going to
- 3:24:27be condensed into a checkpoint.par
- 3:24:30file and then the latest state of this
- 3:24:34source table is going to be computed and
- 3:24:37that is going to be put in zero.json. So
- 3:24:40you would see that these two files would
- 3:24:44be present in the delta log folder. in
- 3:24:49the delta log folder of the target delta
- 3:24:53table. Right? So this one is going to
- 3:24:55you going to be used in order to compact
- 3:24:58all of the information that you have
- 3:24:59over here and we put it into a
- 3:25:020.cheepoint.park
- 3:25:04and the latest state of the source table
- 3:25:07is computed and that is put inside
- 3:25:100.json which becomes version number zero
- 3:25:15of the target table. So the target table
- 3:25:18is going to start with a fresh version
- 3:25:21which is version number zero. Right? And
- 3:25:24what happens to the data? So the park
- 3:25:26file that I've shown over here, it
- 3:25:28doesn't actually mean that there are
- 3:25:30going to be park files, right? It is
- 3:25:31just symbolic of the fact that there is
- 3:25:34going to be data in this target table
- 3:25:37and this data is going to reference all
- 3:25:40of the park files that you see over
- 3:25:42here. Right? So it is simply going to
- 3:25:45reference all of the park file that you
- 3:25:47see in the source table. And that is how
- 3:25:51a shallow clone works under the hood.
- 3:25:53Now let's understand what deep clone is.
- 3:25:56Right? So a deep clone makes an
- 3:25:59independent copy of both the metadata
- 3:26:02and the data files. Right? So this
- 3:26:05operation takes a little bit longer. It
- 3:26:08uses more storage because now this time
- 3:26:11it is not referencing the data of the
- 3:26:13source file. Right? It is actually
- 3:26:15copying the data of the source file in
- 3:26:18the deep clone file in the target.
- 3:26:21Right? So this is going to use more
- 3:26:23storage but the result is that it is a
- 3:26:25completely self-contained and an
- 3:26:28independent table. Right? So let's
- 3:26:31understand that with this diagram right
- 3:26:34here. And again we have a source table
- 3:26:37and we have a target table and we want
- 3:26:41to deep clone the source into the target
- 3:26:45right we want to deep clone the source
- 3:26:48into the target and similar to the last
- 3:26:50example we have metadata in the delta
- 3:26:53log folder with a bunch of JSON files
- 3:26:56that you see over here and some data
- 3:26:59files and these data files are basically
- 3:27:01parket right now when we do a deep clone
- 3:27:04phone. The exact same thing is going to
- 3:27:07happen for the metadata. All of this is
- 3:27:11going to be condensed into a
- 3:27:130.point.park
- 3:27:16and the latest state is going to be
- 3:27:18computed put inside 0.json which is
- 3:27:22going to be version version number zero
- 3:27:25right which is going to be version
- 3:27:28number zero. So the deep clones table is
- 3:27:31going to start at a fresh version which
- 3:27:34is version number zero. Right? And as
- 3:27:37you see over here all of these part
- 3:27:39files right this one this one this one
- 3:27:42and this one this one and this one and
- 3:27:44this one and this one they are going to
- 3:27:46be exactly replicated. Right? So this is
- 3:27:51going to be exactly replicated.
- 3:27:56So this is going to be a one to one
- 3:27:58copy.
- 3:28:00So that is how a deep clone is going to
- 3:28:03work under the hood. Now let's see all
- 3:28:06of this in action. Right. So first of
- 3:28:09all we are going to start with shallow
- 3:28:12clones.
- 3:28:14We're going to start with shallow
- 3:28:15clones. And for this purpose I am going
- 3:28:17to create a new table with customer ids
- 3:28:22from 1 to 100. And I'm going to use a
- 3:28:27Cash statement.
- 3:28:29I'm going to use a Cash statement. So
- 3:28:31this is simply going to be create or
- 3:28:34replace delta catalog delta DB dot
- 3:28:38invoices
- 3:28:40of customers
- 3:28:42from 1 until 100, right? Because that is
- 3:28:44what we are doing over here. And let me
- 3:28:48just put as and this is going to be
- 3:28:50create a replace table. And let's go
- 3:28:52ahead and run this. Let's also quickly
- 3:28:55do a select star just to see how this
- 3:28:58looks like, right? Select star from this
- 3:29:00table. Limit five.
- 3:29:04Yeah.
- 3:29:06Okay. So this looks fine. Now if I were
- 3:29:09to do a describe history of this table
- 3:29:14delta catalog. Delta DB.invoices
- 3:29:18C 1 to 100. we simply going to have one
- 3:29:21row that is the catas statement that
- 3:29:24we've ran right so now in order to add
- 3:29:27more history to this table let's perform
- 3:29:30a few more operation yeah so let me go
- 3:29:33ahead and do a delete from this table
- 3:29:37where
- 3:29:39customer ID is between
- 3:29:4315 and 20 yeah let's run this let's run
- 3:29:48another another statement which is an
- 3:29:50update update this table where so what
- 3:29:54do we want to update so let's update
- 3:29:56customer ID equals 3 and let's change
- 3:29:59the quantity from 3 to 10 right so this
- 3:30:03is going to be where customer ID equals
- 3:30:063 and I need to set this quantity to 10
- 3:30:12yeah so let's run this again and let's
- 3:30:16also quickly see how the layout looks
- 3:30:19like in the file storage, right? So, if
- 3:30:23we go to details
- 3:30:26and this is the storage right here.
- 3:30:31Okay, so here you see two files and then
- 3:30:33two deletion vectors, right? So, I
- 3:30:35believe the first file is for the first
- 3:30:37instance when we did a creation using
- 3:30:40the cas command, right? the deletion
- 3:30:43vector. The first deletion vector is for
- 3:30:45the delete command and the file the park
- 3:30:48file and the updated deletion vector
- 3:30:50that we see actually this is the updated
- 3:30:53deletion vector because this is created
- 3:30:56at 1739. So the updated deletion vector
- 3:30:59that we see is because of the update
- 3:31:01command and the updated row is contained
- 3:31:03within this park file. Now let's go
- 3:31:05ahead and run an insert statement. So
- 3:31:08insert into this table
- 3:31:12and then the values let me quickly copy
- 3:31:16from
- 3:31:19from this table right here. Right?
- 3:31:23So this is let me put a random customer
- 3:31:26ID which is 1099
- 3:31:28and let's quickly format this. Okay. Now
- 3:31:32that this is formatted let's go ahead
- 3:31:34and run this.
- 3:31:38Okay, so the insert statement is
- 3:31:40complete. And if I were to go to the
- 3:31:44file storage ADLS, we should see another
- 3:31:48parket file over here that is going to
- 3:31:49contain that insert statement, right?
- 3:31:53There you go. So we see another parket
- 3:31:56file over here, right? The one that ends
- 3:31:58with 1 to CB - C0. Right? So now let's
- 3:32:04do a describe history
- 3:32:08of this table. We should see four rows.
- 3:32:12Yeah. So the first one is for the cage.
- 3:32:14The second one is for the delete between
- 3:32:1615 to 20. The third one the third one is
- 3:32:19the update where we updated customer ID
- 3:32:22equal three. And the last one is the
- 3:32:24insert over here. Right? So all of this
- 3:32:27makes sense now. Now let's say we want
- 3:32:31to do a shallow clone of this table.
- 3:32:35Let's see what is going to happen. So
- 3:32:36first of all let's write a create or
- 3:32:40replace table
- 3:32:43and this is going to be the same and
- 3:32:45I'll just append it with shallow clone
- 3:32:49and let's write shallow clone
- 3:32:53and the table over here delta catalog
- 3:32:54with this. Right? So let's go ahead and
- 3:32:56run this now. Let's see what happened.
- 3:33:00Okay. So if I were to look at the
- 3:33:02catalog, this should show me this table.
- 3:33:05So this table is created right here. And
- 3:33:07if I were to look at the details, I
- 3:33:09would find the path. And let's have a
- 3:33:11look at what is there in this path. So
- 3:33:14let me just open this in a new tab.
- 3:33:23So now the interesting thing is you only
- 3:33:26see the delta log right so we were
- 3:33:29talking about that the files are
- 3:33:31referenced so that is why you don't see
- 3:33:33any file over here and let's also see
- 3:33:36what there in the delta log so as we
- 3:33:39discuss we see a 0 checkpoint and the
- 3:33:44latest state in zero.json JSON. So I
- 3:33:47hope all of that makes sense. Now you're
- 3:33:48able to connect the dots. Let's look at
- 3:33:50a few more details of the shallow clone
- 3:33:54table. So if I were to say describe
- 3:33:58history of the shallow clone table, I
- 3:34:01should essentially see one command.
- 3:34:03Yeah, which is the clone operation
- 3:34:04itself. Now if you look at the operation
- 3:34:07parameters, it tells me that the source
- 3:34:09of the table is this table over here,
- 3:34:13right? and the source version right the
- 3:34:16source version the latest version of the
- 3:34:18source table was number three and that
- 3:34:21version was used to build the shallow
- 3:34:25clone right so that is what it says over
- 3:34:26here source version is number three and
- 3:34:29is shallow equals true yeah now let's
- 3:34:33update something in the source table and
- 3:34:36see if it affects the shallow clone
- 3:34:38table right so we have the source table
- 3:34:42right here uh over here right so let me
- 3:34:48do a select quickly do a select star
- 3:34:52from this table
- 3:34:55select start from this table and let me
- 3:34:58do a limit five and let's go ahead and
- 3:35:01run this so I want to update some row so
- 3:35:05let me update customer ID equals 5 this
- 3:35:08time so customer ID equals 5 has
- 3:35:11quantity equals one but now I want to
- 3:35:13update it to 10. So let's go ahead and
- 3:35:16run this. Let's also see the history of
- 3:35:19this table.
- 3:35:22So now it should contain an update.
- 3:35:25Right? So after an insert that we did
- 3:35:27previously, it now has an update. So
- 3:35:30let's quickly do a select star from this
- 3:35:35table where customer ID equals 5. And
- 3:35:39now I should see quantity equals 10.
- 3:35:43Right? I see quantity equals 10. Let's
- 3:35:46quickly verify
- 3:35:48what is the resulting value in the
- 3:35:51shallow clone table.
- 3:35:55So he we see here that quantity equals 1
- 3:35:59and customer ID equals 5. That means
- 3:36:02even though the shallow clone is
- 3:36:04referencing the source table after the
- 3:36:08point that the shallow clone has been
- 3:36:10cloned right the table the target
- 3:36:12shallow clone table has been cloned it
- 3:36:14is going to have its own history its own
- 3:36:17set of operation and any activity on the
- 3:36:20source table is not going to affect the
- 3:36:24shallow clone right I hope this is very
- 3:36:25clear now let's try something different
- 3:36:27let's try deleting or modifying some
- 3:36:31records in the shallow clone table and
- 3:36:34let's see whether that affects the
- 3:36:36source table or not. Right? So let's do
- 3:36:39a delete from delta catalog delta db do
- 3:36:45this table shallow clone where customer
- 3:36:49ID equal 99. Right? So let's go ahead
- 3:36:52and run this and let's also quickly run
- 3:36:54select star from this table where
- 3:36:58customer ID equals 99. So this shouldn't
- 3:37:02return me any row and let's also have a
- 3:37:05look at the history. Okay, so no rows
- 3:37:07returned. Let's also have a look at the
- 3:37:09history of this table and it has a clone
- 3:37:13and then a delete the delete operation
- 3:37:16that we ran just now. Now if I were to
- 3:37:18compare this to
- 3:37:22to the original table, ideally I should
- 3:37:26find the record, right? Because the
- 3:37:29operation that a shallow clone is going
- 3:37:31to have is going to maintain its own
- 3:37:34history after the point it has been
- 3:37:36cloned. So any operation on the shallow
- 3:37:39clone table shouldn't affect the source
- 3:37:42table. Yeah. So let's go ahead and run
- 3:37:44this. And there you go. we find that the
- 3:37:47row number 99 customer ID equal 99 is
- 3:37:50still there and let's also verify this
- 3:37:53with the history just to make sure that
- 3:37:56the history also remains unaffected
- 3:37:58right so let's run this
- 3:38:02and there is no trace of a delete
- 3:38:04statement right so after the shallow
- 3:38:06clone table is cloned it maintains its
- 3:38:10own history right so now to quickly
- 3:38:13summarize once A source table has been
- 3:38:16shallow cloned. Changes or updates or
- 3:38:20modifications or deletes to the shallow
- 3:38:23clone table is not going to affect the
- 3:38:25source table. And changes or updates or
- 3:38:28modifications to the source table is not
- 3:38:31going to affect the shallow clone table.
- 3:38:33Right? The shallow clone table is only
- 3:38:36going to refer to the data of the source
- 3:38:39table until the point the clone
- 3:38:42happened. Right? From that point
- 3:38:44onwards, it is going to maintain its own
- 3:38:47version, its own history, its own data
- 3:38:49files. Right? So I hope that makes
- 3:38:51sense. Now what I want to show you is
- 3:38:54that we can also clone a table using a
- 3:38:58particular timestamp or a particular
- 3:39:00version number. Yeah. So what we can do
- 3:39:03is that we can run a create or replace
- 3:39:09table
- 3:39:11and this is going to be shallow clone.
- 3:39:14Let's let's create a clone of version
- 3:39:17number this version right version number
- 3:39:19zero v 0
- 3:39:22and this is going to be shallow clone
- 3:39:26delta catalog dot this table
- 3:39:30version as of
- 3:39:33zero right so let's run this now
- 3:39:38so if this has worked properly
- 3:39:42you should not see actually you should
- 3:39:45see records from customer ID is 15 to 20
- 3:39:49right so let's start from this table
- 3:39:52where customer
- 3:39:55where customer
- 3:39:57ID
- 3:39:59between 15 and 20 now if I were to
- 3:40:03similarly run this on the original table
- 3:40:08I shouldn't get any records right
- 3:40:10because the latest version of this table
- 3:40:12doesn't have those rows from 15 to 20.
- 3:40:15Yeah. So let's run this.
- 3:40:18And there you see 15 to 20. But in the
- 3:40:21original table there should be no rows
- 3:40:22return. Okay. That works. And another
- 3:40:26variation of this is that you can also
- 3:40:30do a time stamp as of and you can
- 3:40:33basically put pick up any time stamp
- 3:40:35over here. Right? So you can pick this
- 3:40:38one up and write something like this.
- 3:40:40Right? Now let's talk about time travel.
- 3:40:43So if I were to quickly show you,
- 3:40:46okay, let me just note this down. Time
- 3:40:50travel.
- 3:40:51If I were to show you describe history
- 3:40:55of this table, we are going to have
- 3:40:59several rows, right? Because we
- 3:41:01performed several operation. Now what
- 3:41:05I've already mentioned earlier is that
- 3:41:08when we create a shallow clone it starts
- 3:41:11from version zero. It takes up the
- 3:41:13latest state of the source table and it
- 3:41:16starts from version zero. So if I were
- 3:41:18to do something like this,
- 3:41:21create or replace
- 3:41:25table
- 3:41:27and let me put this as
- 3:41:32test shallow clone as shallow clone
- 3:41:38this whole table.
- 3:41:41Right? So if I were to run this as we've
- 3:41:45already seen earlier
- 3:41:49as we've already seen earlier this is
- 3:41:51going to create this table taking the
- 3:41:55latest version of this table over here
- 3:41:57right so if you do a history on this
- 3:42:00table it is only going to contain one
- 3:42:02row right
- 3:42:07so I'm reiterating all of this right
- 3:42:09because I want to make sure that this is
- 3:42:12super clear.
- 3:42:15So you see that there is only one row
- 3:42:18that means version number zero of this
- 3:42:21table has been created using the latest
- 3:42:25state of the source table. So the source
- 3:42:27table has many versions right it has
- 3:42:29version 0 1 until 4. If I want to roll
- 3:42:32back to a particular version or if I
- 3:42:34want to see a particular version of the
- 3:42:36source table I can definitely do that
- 3:42:38right. I can see version number three,
- 3:42:40version number two, version number one
- 3:42:42and so on. But just because we have
- 3:42:46shallow cloned this table, it doesn't
- 3:42:48mean that we can access the history of
- 3:42:51the sort table. We cannot go back to
- 3:42:53version one or two for this shallow
- 3:42:57clone table. Right? So I hope that is
- 3:42:59clear because we are going to start with
- 3:43:02a fresh history with a fresh version
- 3:43:04with version number zero taking the
- 3:43:07latest state of the source table. Now
- 3:43:10let's try out something interesting.
- 3:43:12Let's run a vacuum on the source table.
- 3:43:15And by source what I mean is this table
- 3:43:18right here invoices_c1_00.
- 3:43:22Right? So if you remember when we
- 3:43:24created this table
- 3:43:27when we created this table right over
- 3:43:29here we ran a few statements in order to
- 3:43:32populate the history. The first one was
- 3:43:34a delete the second one was an update
- 3:43:37and the third one was an insert and I
- 3:43:41also pointed out that this is the park
- 3:43:45file where the value of that insert was
- 3:43:48stored. Right? So
- 3:43:51the the customer id that is belonging to
- 3:43:54this insert is 1099
- 3:43:56and this table later on was cloned
- 3:44:00right. So that means that this row where
- 3:44:02customer ID equals 1099 is present in
- 3:44:05the source table it is also present in
- 3:44:08the shallow clone table right. So the
- 3:44:10shallow clone table is also referencing
- 3:44:12this row. Now if what happens if I were
- 3:44:16to run a delete? So let me just quickly
- 3:44:19add this
- 3:44:21heading and
- 3:44:23let's say I run a delete. Delete from
- 3:44:26this table.
- 3:44:30Delete from this table
- 3:44:32where customer ID equals 1099.
- 3:44:38And this should ideally
- 3:44:41remove
- 3:44:44remove this row, right? So, like start
- 3:44:45from this table where customer ID equals
- 3:44:511099, right? So, this should not give me
- 3:44:53any result. Okay, that works perfectly
- 3:44:56fine now
- 3:44:58because this file is now orphaned,
- 3:45:01right? This file is not being referred
- 3:45:04in the latest version when I run a
- 3:45:06vacuum. This file should just go away,
- 3:45:09right? That is how vacuum is going to
- 3:45:11behave. So first of all I have to set
- 3:45:13this property set spark dot databick
- 3:45:18dot delta dot retention
- 3:45:22duration
- 3:45:25enabled is false and then let's go ahead
- 3:45:29and run a vacuum vacuum this table
- 3:45:33retain zero hours right let me go ahead
- 3:45:37and run this and ideally if everything
- 3:45:39works out fine I shouldn't see this park
- 3:45:43file which end with 12 CB - C0 right
- 3:45:48let me go ahead and run this
- 3:45:52okay this is complete and let me refresh
- 3:45:55this
- 3:45:58okay so one of the deletion vectors was
- 3:46:01cleared which is good but I don't see
- 3:46:05the file deleted right the one ending
- 3:46:09with 12 CB hy hyphen C0 is not deleted.
- 3:46:13Why that may be the case? Right? So can
- 3:46:16you think of why this may have happened?
- 3:46:20So the reason this has happened is
- 3:46:23because although this file in the source
- 3:46:25has been deleted, there is a shallow
- 3:46:28clone which is still referencing that
- 3:46:30row. Right? So that is why that file had
- 3:46:33not been deleted. So let me
- 3:46:36let me show you
- 3:46:39if I were to run this on the shallow
- 3:46:42clone table.
- 3:46:47This is there right now. Let me go ahead
- 3:46:49and delete this from all of the shallow
- 3:46:51clone that I've created. So by now you
- 3:46:54must have seen that I created a few
- 3:46:55shallow clones and we want to remove all
- 3:46:58of the references of that row 1099.
- 3:47:01Right? So this would be delete from
- 3:47:05delta catalog dot this table where
- 3:47:10customer ID equal 1099.
- 3:47:15What were the other table that I
- 3:47:17created? It was v 0
- 3:47:21and
- 3:47:23what else?
- 3:47:25So I created v 0 and then underscore
- 3:47:28test_cccl
- 3:47:31test_cl.
- 3:47:33So I want to remove all of the
- 3:47:35references. Right? So now what this
- 3:47:38means is that this was the source table
- 3:47:41and there were several shallow clones.
- 3:47:44Some of them were referring to row
- 3:47:45number 1099. Now I have removed all of
- 3:47:48the references. So that file is now an
- 3:47:51orphaned file. There is no reference. So
- 3:47:53once I run a vacuum that file should be
- 3:47:56gone now, right? So now let's go ahead
- 3:47:59and run this vacuum once again.
- 3:48:03Okay. So this is complete now. So I
- 3:48:06believe this file should be gone if
- 3:48:10whatever logic that we were discussing
- 3:48:12is true, right? So let's refresh. And
- 3:48:16there you see. So you have only three
- 3:48:19files and that file is gone because
- 3:48:22there are no more references to that
- 3:48:25park file anymore. Right now let's see
- 3:48:27deep clone in action. So let me quickly
- 3:48:30put that down as a heading
- 3:48:32and I want to clone
- 3:48:36the source table that we've been using
- 3:48:38so far which is this one. So let me
- 3:48:40quickly see the history. Okay. So
- 3:48:43there's a lot of thing that we've
- 3:48:44performed on this table, right? And let
- 3:48:48me also
- 3:48:50let me also
- 3:48:52see how the file looks like. Right? So
- 3:48:54this is how the files look like. Right?
- 3:48:56Now let me go ahead and perform a deep
- 3:49:00clone. Create or replace table
- 3:49:05this table. And this is going to be a
- 3:49:08deep clone. Let me also remove this and
- 3:49:14let me write this as DCL which is deep
- 3:49:18clone. Right? So let's go ahead and run
- 3:49:22this. Okay. So this has created a deep
- 3:49:26cloned table. Let's see describe history
- 3:49:32of this table. And ideally we should
- 3:49:34just see one row which is the clone.
- 3:49:36Perfect. And if we look at the
- 3:49:39parameters, we see that this is being
- 3:49:42cloned from this table and the source
- 3:49:44version is 13, right? 13 the latest
- 3:49:47version. Yeah. So let's have a look at a
- 3:49:51few other things. So if I were to
- 3:49:55check how the files look like, right? So
- 3:49:59let me have a look at how the files look
- 3:50:01like.
- 3:50:06So what you see over here is that these
- 3:50:09two are an almost an exact copy. So
- 3:50:13there are three parket files and one
- 3:50:16deletion vector and that is what you see
- 3:50:19over here exactly. So these are three
- 3:50:20parket files and one deletion vector and
- 3:50:23this is a one one copy right. So we see
- 3:50:28let's let's quickly check the numbers.
- 3:50:29So this is 967
- 3:50:31F2B
- 3:50:33and 3d2 right. So this is 967 F2B and
- 3:50:383d2 right and the deletion vector end
- 3:50:40with 9 F9. It also ends with 9 F9.
- 3:50:44Right? So that means that this is an
- 3:50:47exact copy from the source. But for the
- 3:50:51delta log we discussed that there is
- 3:50:52going to be a 0.choint.park
- 3:50:55and then 0.json JSON which is going to
- 3:50:57generate the latest state
- 3:51:01and exactly that is what we see over
- 3:51:03here. Yeah. So now let's do a few other
- 3:51:06things. Let's make some changes in the
- 3:51:08in the deep clone and see if it affects
- 3:51:11the source table. Ideally it shouldn't
- 3:51:13because we mentioned earlier that the
- 3:51:15deep clone is an independent and a
- 3:51:18self-contained copy. Right? So let's
- 3:51:22make a few changes now. First of all,
- 3:51:25let me do select star from
- 3:51:28this table right here. And let's do a
- 3:51:32limit five. So let's run this.
- 3:51:36So I'm going to write an update
- 3:51:39from
- 3:51:41sorry update this table where I'm going
- 3:51:46to set the quantity
- 3:51:50of
- 3:51:52customer ID equals 4.
- 3:51:55I'm going to set the quantity to 10
- 3:51:58where customer ID equals 4.
- 3:52:02And let's go ahead and run this. So
- 3:52:03let's quickly verify this from this
- 3:52:07table where customer ID equals 4. This
- 3:52:13should now be 10 instead of five. And
- 3:52:18that is what we see over here. Now let
- 3:52:20me quickly verify this in the original
- 3:52:24table.
- 3:52:27The original table still shows five.
- 3:52:30Right? So as we expected that any change
- 3:52:34in the cloned table in the deep clone
- 3:52:38table is not going to affect anything in
- 3:52:41the source table and vice versa is going
- 3:52:43to be true. any changes in the source
- 3:52:45table is not going to affect the deep
- 3:52:49cloned table. Right now, I'm not going
- 3:52:52to run and show you a vacuum on a deep
- 3:52:55clone because as I mentioned earlier, a
- 3:52:57deep clone is a self-contained and an
- 3:53:01independent copy, right, of the source
- 3:53:03table. There are no references between
- 3:53:05the deep clone and the source table. So
- 3:53:08if we were to quickly summarize what
- 3:53:10happened in the shallow clone vacuum. So
- 3:53:14we had a file right we had a file where
- 3:53:16customer ID equal 1099 and there were
- 3:53:19several shallow clones referring to that
- 3:53:22row. Now when we performed a delete and
- 3:53:24then we did a vacuum still that file was
- 3:53:28not removed from the source table. Why?
- 3:53:32Because shallow clone there were three
- 3:53:33shallow clones which were referring to
- 3:53:36customer ID equals 199 1099. Once all of
- 3:53:40those references were gone, right? Once
- 3:53:43we deleted all of the references, we ran
- 3:53:46a delete statement which meant that that
- 3:53:49row is no longer in the shallow clone.
- 3:53:53Those references were deleted. And then
- 3:53:55we ran a vacuum on the source table. We
- 3:53:57saw that the park file was deleted.
- 3:54:00Right? But in case of a deep clone, the
- 3:54:03source table and the deep clone are
- 3:54:06completely different. Different in the
- 3:54:08sense that they are independent copies
- 3:54:11of each other. There is no reference
- 3:54:13from the deep clone to the source table.
- 3:54:16Right? So that is the reason why vacuums
- 3:54:19should be completely independent. Right?
- 3:54:22So before we conclude this section on
- 3:54:24deep and shallow clones, I believe that
- 3:54:26a good number of you would have this
- 3:54:28question that how are cats different
- 3:54:31from deep clone. So the tables that are
- 3:54:34generated using cas create or replace
- 3:54:36table and then a select query. How is
- 3:54:39that different from the tables that are
- 3:54:42generated using deep clone? So if you
- 3:54:45were to look at it at an output level,
- 3:54:48so let's say you're doing a select star
- 3:54:50from whatever table, right? If you were
- 3:54:52to look at the outputs that are being
- 3:54:53generated by the two tables, one which
- 3:54:56is generated using CAS and the other one
- 3:54:59that is generated using deep clone, they
- 3:55:01would look exactly identical, right? But
- 3:55:04that is not where the difference lies.
- 3:55:07So when you deep clone a table, you also
- 3:55:10clone, you don't need to respspecify the
- 3:55:13partitioning properties, the constraints
- 3:55:15and all of that. Right? With CASS, the
- 3:55:18table is just created using the output
- 3:55:20of a select query, right? So you write a
- 3:55:22select and then it generates a set of
- 3:55:24rows and it just uses those rows to
- 3:55:27create a table. All of the properties
- 3:55:30are lost and that is where deep clone
- 3:55:33comes in handy. It's a robust way to
- 3:55:35clone the metadata, the data and the
- 3:55:39properties of the table. Another very
- 3:55:42important advantage is that it works in
- 3:55:45an incremental manner. So what I mean is
- 3:55:48that let me let me show that to you with
- 3:55:50an example. So let's say we have a
- 3:55:54disaster recovery use case, right? So
- 3:55:57let's say we have a
- 3:55:59disaster recovery
- 3:56:02use case. And by this what I mean is
- 3:56:05that let's say we have a source table
- 3:56:06over here. we have a source
- 3:56:11and then what I'm doing is that I am
- 3:56:14having a replica.
- 3:56:16So this is the replica
- 3:56:20and I am syncing these two tables.
- 3:56:23Right?
- 3:56:25What I did first is that I did a deep
- 3:56:27clone.
- 3:56:30I did a deep clone of the source table.
- 3:56:33Right? Now what happens when I'm going
- 3:56:35to get an update?
- 3:56:38So let's say there is an update
- 3:56:41on the source table, right? And then
- 3:56:44there was a delete
- 3:56:49on this source table. What is going to
- 3:56:51happen? So I would need to bring these
- 3:56:53two tables again back in sync. Right? So
- 3:56:56in order to bring this back in sync
- 3:56:58again, I would need to do a
- 3:57:02deep clone again, right? So I would end
- 3:57:05up doing a deep clone again. But the
- 3:57:06beauty of this is that it doesn't copy
- 3:57:09the whole source
- 3:57:11into the whole replica, right? It
- 3:57:13doesn't copy the whole thing again,
- 3:57:15right? It only copies incremental
- 3:57:18changes. What it is going to do is that
- 3:57:20it is going to take this update and
- 3:57:22apply it on the replica. It is going to
- 3:57:24take this delete and apply this on the
- 3:57:27replica. Right? So the whole table is
- 3:57:29not going to be copied again and again.
- 3:57:32The operations are copied or synced in
- 3:57:35an incremental manner. And that is the
- 3:57:37reason why deep clones are robust and
- 3:57:41they are very performant. So now we are
- 3:57:43going to talk about something that
- 3:57:45silently kills your spark performance
- 3:57:48and that is the small file problem.
- 3:57:54Yeah, the small file problem
- 3:57:59and we are going to understand what
- 3:58:01exactly the small file problem is, why
- 3:58:03does it happen and how this can be fixed
- 3:58:07using delta lakes optimize command.
- 3:58:11Yeah. So let's first understand what the
- 3:58:15small file problem is and why exactly is
- 3:58:18it problematic. Yeah. So now I'm going
- 3:58:20to take a weird example but it is just
- 3:58:24to help you get the right understanding.
- 3:58:26Yeah. So imagine you are reading a PDF.
- 3:58:31So imagine you're reading a PDF of a
- 3:58:34300page novel
- 3:58:37of a 300page novel. Right? So naturally
- 3:58:40you would expect the novel to be a
- 3:58:44single 300page PDF, isn't it? So you
- 3:58:47would expect it to be a single
- 3:58:51300page PDF, right? But what if what if
- 3:58:55I would say that someone saved each of
- 3:58:58these pages as a PDF? So what they did
- 3:59:01was they saved
- 3:59:04they saved one PDF, two PDF
- 3:59:09until all the way until 300. PDF. So
- 3:59:13what they did is that they took each of
- 3:59:16those pages, each of those 300 pages and
- 3:59:19saved it as a PDF. So how would you read
- 3:59:23this kind of a book, right? So first of
- 3:59:26all, you would end up saying some very
- 3:59:28nice words to that person who did this
- 3:59:31which I cannot say on this video. But
- 3:59:33then what you would essentially do is
- 3:59:35that you would open the first PDF.
- 3:59:37Actually, you would look for the first
- 3:59:39PDF, right? So first of all you would
- 3:59:41look for the first PDF look or find the
- 3:59:44first PDF right and then you would open
- 3:59:47the first PDF and you would read it and
- 3:59:51finally when the reading is complete you
- 3:59:54would close the PDF right now you would
- 3:59:57have to do this for all of the 300 pages
- 4:00:01that have been given to you and what we
- 4:00:04see here is that we've spent a good
- 4:00:07amount of time first of all finding the
- 4:00:10right file and then opening it and then
- 4:00:13finally closing it. Right? So there is a
- 4:00:16good amount of time that you spend in
- 4:00:19all of these problems in all of these
- 4:00:22operations. Right? And that is exactly
- 4:00:25what the small file problem is. When you
- 4:00:27have a small pile problem, there are too
- 4:00:30many small files and then you end up
- 4:00:32doing such operations, right? you end up
- 4:00:36first of all finding the right file and
- 4:00:38then opening it and closing it right and
- 4:00:41these are time consuming operation. So
- 4:00:44to reiterate, when you're reading a
- 4:00:47table and instead of having a few neatly
- 4:00:50packed file, if your data is scattered
- 4:00:53over thousands of smaller files, there
- 4:00:56is going to be thousands of those open,
- 4:00:59close and metadata lookup or finding
- 4:01:03operation that we just discussed. All of
- 4:01:06this leading to wasted compute and poor
- 4:01:08IO, right? So all of this all of these
- 4:01:11operations is simply going to lead to
- 4:01:14wasted compute
- 4:01:18and poor IO
- 4:01:20right so how does delta solve this
- 4:01:24problem delta provides a builtin
- 4:01:27operation which is called optimize and
- 4:01:30what it simply does is that it compacts
- 4:01:33all of these small files into larger
- 4:01:36more appropriately sized one right And
- 4:01:39in order to do this, it simply uses
- 4:01:42something called a bin packing
- 4:01:45algorithm.
- 4:01:47It uses a bin packing algorithm, right?
- 4:01:51So we are going to see this in detail.
- 4:01:53So don't worry about it. Nevertheless,
- 4:01:55it's actually a very simple algorithm.
- 4:01:57What it does is that it collects all
- 4:01:59your file sizes and then it sorts it
- 4:02:02from high to low and then it start
- 4:02:04picking each file and it places each of
- 4:02:08these files in a particular bin and
- 4:02:10think of this bin as one large file.
- 4:02:14Right? So you have lot of small files
- 4:02:17over here. Right? You have lot of small
- 4:02:20files. What it does is that it simply
- 4:02:24takes this and it puts them in a bigger
- 4:02:29bin as long as it fits in the bin.
- 4:02:32Right? So let's say this simply is able
- 4:02:36to fit in this bin. This is also able to
- 4:02:38fit in this bin. This is also able to
- 4:02:40fit in this bin. But this one is not
- 4:02:43able to fit in this bin. Right? So in
- 4:02:45that case a new bin is going to be
- 4:02:48created and then this is going to go
- 4:02:52over here. Right? So that is how simply
- 4:02:55the bin packing algorithm works. And I'm
- 4:02:58going to show you a visualization in
- 4:03:01order to help help you to understand
- 4:03:03this better. But before we go there I
- 4:03:05want you to understand that the default
- 4:03:09bin size that delta targets is 1 GBTE.
- 4:03:14Right? So this is going to be 1 GBTE.
- 4:03:16That is what is the default for Delta
- 4:03:20Lake, right? And this default 1GB size
- 4:03:23has been selected after years of usage
- 4:03:27and testing on different kind of spark
- 4:03:29workloads and it's really proven to be
- 4:03:31robust. Right? So this number 1GB has
- 4:03:35really proven to be robust. So unless
- 4:03:38you have a very good reason to change
- 4:03:40it, don't change it. go ahead and use
- 4:03:42the default 1GB number. Yeah. So just to
- 4:03:46let you know just in case you want to
- 4:03:48change it, there is a property which is
- 4:03:51called spark dot databick
- 4:03:55dot delta dot optimize
- 4:03:58dot max file size. So look it up on the
- 4:04:01internet
- 4:04:04dot max file size. So you can use this
- 4:04:07property to change the 1GB default
- 4:04:09number but my advice is again don't
- 4:04:12change it unless you have a very
- 4:04:14compelling reason to do so. Yeah. So now
- 4:04:16let's see the visualization. So here we
- 4:04:20have a list of files a list of parakeet
- 4:04:23files and it has simply been sorted in
- 4:04:26descending order by file sizes right. So
- 4:04:29the first one you see is 320 MB and the
- 4:04:31last one is 80 MB. So we are going to
- 4:04:33run through how bins are going to be
- 4:04:36created right and again we are assuming
- 4:04:38that the target size of one bin one file
- 4:04:42one large file is going to be 1 GBTE
- 4:04:45right so let's pick up the first file
- 4:04:46and let's see what happens the first
- 4:04:48file park number seven is picked up and
- 4:04:51it is simply placed in the first bin and
- 4:04:53this is 1,000 megabytes of size out of
- 4:04:56which 320 has been occupied now let's
- 4:05:00pick up the next one the next one is 310
- 4:05:02megabytes. Of course, the first bin has
- 4:05:05that much of size. So, it is going to be
- 4:05:08placed in the first bin. Now, let's go
- 4:05:10to the third one. The third one is 280
- 4:05:13MB. This does have 280 MB of size. So,
- 4:05:17it is going to be placed in the first
- 4:05:19bin again. Now, the first bin has only
- 4:05:2390 MB of size left. Now, let's go over
- 4:05:27to the next file. But the next file is
- 4:05:29270 MB. So this for this a new bin is
- 4:05:33going to be created. Right? So it tries
- 4:05:34to put it in the first bin but it
- 4:05:36cannot. So that is why a new bin is
- 4:05:39created. Yeah. Now let's go over to the
- 4:05:41next file. The next file is 220 MB. The
- 4:05:44first bin only has 90 MB. So that is why
- 4:05:47it's going to be placed in the second
- 4:05:49bin. Now yeah. Now let's go over to the
- 4:05:51next park file which is park number 9.
- 4:05:54190 megabits of megabytes of size. It
- 4:05:57cannot be placed in the first bin. So
- 4:06:00let's see that it cannot be placed in
- 4:06:01the first bin. So it goes over to the
- 4:06:03next bin. Yeah. The next park file is
- 4:06:07180 megabytes of size. Again it cannot
- 4:06:09be placed in the first bin and it should
- 4:06:11go in the second bin. Next one. Again
- 4:06:14160 MB of size here. Cannot be placed in
- 4:06:18the first bin. How much size does the
- 4:06:20second bin have now? It has 140 MB of
- 4:06:24size left. So it cannot be placed in the
- 4:06:27second bin also. Right. So it goes to
- 4:06:30the second bin but it cannot be placed.
- 4:06:32So that is why a new bin is going to be
- 4:06:34created. Yeah. Simple. So let's go over
- 4:06:37to the next one. 150 megabytes of size.
- 4:06:40It cannot be placed in the first one.
- 4:06:42Cannot be placed in the second one also
- 4:06:44but can be placed in the third one.
- 4:06:46Yeah. Let's go over to the next one. 130
- 4:06:49megaby cannot be placed in the first
- 4:06:51one. But it can be placed in the second
- 4:06:54one. Right? because one because the
- 4:06:57second one has 140 mgabytes of size
- 4:07:00left. So, it is going to go in the
- 4:07:02second bin. Right? Let's pick up the
- 4:07:05next one. Again, it cannot go to the
- 4:07:06first bin. It cannot go to the second
- 4:07:09bin. So, it is going to end up in the
- 4:07:11third bin.
- 4:07:13Now, the last one, we see that the first
- 4:07:15bin does have 90 mgabytes of size left.
- 4:07:19So, it is going to end up in the first
- 4:07:22bin itself, right? And that is how we
- 4:07:26see that three bins are created nearing
- 4:07:301 GB of size and all of the files you
- 4:07:33had so many files it has been reduced to
- 4:07:36smaller number of files. Right? So now
- 4:07:39if we correlate this back to the PDF
- 4:07:42example instead of opening 300 PDFs one
- 4:07:45by one probably you will have four to
- 4:07:48five PDF. Right? Okay. So now let's see
- 4:07:51all of this in action and for this I'm
- 4:07:53going to create a notebook a new
- 4:07:56notebook which is going to be named 07
- 4:08:00optimize right and let's use this data
- 4:08:04set that we've already been using
- 4:08:06earlier for quite a good number of time.
- 4:08:08So let's simply do a spark read.par
- 4:08:13and together with this let's also
- 4:08:15display five rows from this data set.
- 4:08:19Yeah.
- 4:08:22So now while we are doing this
- 4:08:25let me also print the total number of
- 4:08:27rows and along with that. Okay. So here
- 4:08:31is what it looks like and let me also do
- 4:08:34a df do select category dodistinct
- 4:08:39dot count. Right? So the reason why I'm
- 4:08:42doing a distinct of the column category
- 4:08:44which has eight unique categories is
- 4:08:47because I want to partition by this
- 4:08:49column. Right? So let's quickly do a df
- 4:08:53dot
- 4:08:54repartition
- 4:08:56five dot write dot mode override
- 4:09:02dot partition by we partition by the
- 4:09:06category column dot save as table and
- 4:09:10I'm going to simply save as delta
- 4:09:12catalog dot deltadb dot optimize example
- 4:09:17one right so let's just save it like
- 4:09:19this and the Reason why I'm doing a
- 4:09:21repartition five is because I want to
- 4:09:24mimic the small file behavior. So each
- 4:09:27of the partitions that you see over here
- 4:09:29that is going to be created for the
- 4:09:31column category. So you're going to see
- 4:09:33one partition for clothing, one
- 4:09:34partition for toys and so on. Each of
- 4:09:37these partitions should have five file.
- 4:09:40Yeah. So let's go ahead and run this.
- 4:09:44Okay. So this is complete.
- 4:09:47Let me go to the catalog.
- 4:09:51And
- 4:09:54I simply take this up.
- 4:09:57Take the table ID.
- 4:10:00And
- 4:10:01let me go to the container
- 4:10:04tables. And then here is this table. And
- 4:10:08you see all of the partitions over here.
- 4:10:10Each of these partitions should have
- 4:10:12five files, right? There you go. So you
- 4:10:15see five files over here. And that is
- 4:10:16because of the repartition pipe that we
- 4:10:19put over there. Yeah. So now let's run a
- 4:10:21query. And let's also time this query.
- 4:10:25Let's do a df example one equals
- 4:10:30this. I read in the table. And let's say
- 4:10:33DF output equals
- 4:10:37DF example 1 dot
- 4:10:40where DF example one dot category equal
- 4:10:44equals clothing
- 4:10:46dot collect right and let's also let's
- 4:10:51also print okay let's let's go ahead and
- 4:10:53run this now so let's run this and see
- 4:10:55what happens what is the total time that
- 4:10:57is taken 2.3 seconds is the total time
- 4:11:00that is taken to run this query. Right
- 4:11:03now let's go ahead and run a optimize
- 4:11:06run a compaction on top of this table.
- 4:11:08Right? So basically the five small files
- 4:11:12that we see that should be combined into
- 4:11:15one bin one file as a result of this
- 4:11:17optimize operation. So we simply do from
- 4:11:20delta dot table import
- 4:11:24delta table
- 4:11:26and
- 4:11:27table equals delta table dot for name.
- 4:11:32Okay, this is great. That is what I
- 4:11:34wanted to run. So let's go ahead and run
- 4:11:36this. Yeah,
- 4:11:45perfect. So now this is complete. What I
- 4:11:48should see over here is instead of five
- 4:11:51files, I should see just one file
- 4:11:54because they've all been compacted into
- 4:11:56one file, right? So, let me run a
- 4:11:59refresh.
- 4:12:02Okay, so now instead of one file, I see
- 4:12:04six files, right? All of the files were
- 4:12:07written at 185515.
- 4:12:10But then there is this one file which
- 4:12:12just got written now at 185731.
- 4:12:16And if you look at the size, right, the
- 4:12:18size is bigger than all of this, right?
- 4:12:20So what happened is it combined all of
- 4:12:24those five files into one file, right?
- 4:12:27And that and this file right here is
- 4:12:29that one file. What happened to the
- 4:12:32other file? They still stay there, but
- 4:12:34they have been tombstone. What tombstone
- 4:12:37means is that they have been marked for
- 4:12:40soft delete. Right? So if you look at
- 4:12:44the catalog over here,
- 4:12:48sorry the delta log over here, not the
- 4:12:50catalog. What we should see over here is
- 4:12:53if I were to look at category clothing,
- 4:12:56we see that there are five removes.
- 4:12:59Yeah. And there should be one add
- 4:13:02because it removes those five small
- 4:13:04files from the log and then it adds one
- 4:13:07file. So here we should see one ad
- 4:13:12for the ad operation for the clothing
- 4:13:14category. Yeah. And this is the new file
- 4:13:17that got added. Right. So I hope that
- 4:13:18makes sense when you compare it with the
- 4:13:20delta log. Now how do you remove those
- 4:13:23five files which is no longer going to
- 4:13:25be referenced in the latest version?
- 4:13:28We'll come to that. Now let's go back
- 4:13:31and let's run this
- 4:13:35run an operation called vacuum. Right.
- 4:13:38So what vacuum is going to do is it is
- 4:13:40going to remove the files for
- 4:13:43tombstoning. Right. The files that have
- 4:13:45been tombstone or marked for soft delete
- 4:13:49they are going to be removed. Yeah. So
- 4:13:51in order to do that we simply run a
- 4:13:53table vacuum and I put a number zero
- 4:13:57over here. Zero basically stands for
- 4:13:59retention duration. Generally the
- 4:14:01retention duration is 168 hours. So
- 4:14:05files that have been created within the
- 4:14:06last 160 68 hours are not going to be
- 4:14:10deleted. But when I put a number zero,
- 4:14:13it simply says that go ahead and delete
- 4:14:16all of the files that have been
- 4:14:18tombstone or have been marked for soft
- 4:14:21delete. So let's go ahead and run this.
- 4:14:25Okay. So it gives me a warning saying
- 4:14:28that are you sure that you would like to
- 4:14:30vacuum file with such low retention
- 4:14:33duration and it's really good that such
- 4:14:36a warning is in place right because
- 4:14:38somebody may be just running a vacuum by
- 4:14:40mistake. So what we need to do is if
- 4:14:44you're certain that there are no
- 4:14:45operation being performed on this table
- 4:14:47such as blah blah blah then you may turn
- 4:14:49off this check by setting this right and
- 4:14:52as I was saying earlier if you're not
- 4:14:54sure please use the value not less than
- 4:14:57168 hours right 168 hours happens to be
- 4:15:00the default so let me go ahead and run
- 4:15:03this
- 4:15:07and this is going to be a set
- 4:15:14Let's run this now
- 4:15:16and let's run the vacuum now.
- 4:15:21Yeah. Okay. So, the vacuum is complete
- 4:15:24now. Now if I if I were to go back to
- 4:15:28this table and if I were to open any of
- 4:15:31these categories ideally I should see
- 4:15:34just one file because all of the five
- 4:15:37files were compacted into one file.
- 4:15:40Right? And that is what you see over
- 4:15:42here. Right?
- 4:15:44Let me also check the other ones.
- 4:15:48Okay. Perfect. So this works as
- 4:15:50expected. Let me also rerun
- 4:15:54rerun this query on this table after
- 4:15:57we've run an optimize and a vacuum.
- 4:16:00Right? So let's go ahead and run this
- 4:16:02again.
- 4:16:05Okay. So the wall time is now 1.4
- 4:16:09seconds comparing to 2.3 seconds
- 4:16:11earlier. So there is an improvement in
- 4:16:14the runtime. Right. Of course the number
- 4:16:15is not very huge because our file sizes
- 4:16:19are very small. we will be able to see
- 4:16:21significant improvements on a larger
- 4:16:23data set. Right? But overall the idea is
- 4:16:28to help us understand that small file
- 4:16:31problem is a significant problem. Right?
- 4:16:34And optimize helps us eliminate that
- 4:16:37problem. Okay. So now I want to share
- 4:16:40with you how to use predicates or
- 4:16:43filters or where condition with
- 4:16:45optimize. So this is going to be very
- 4:16:48helpful in cases where you have
- 4:16:50incremental data coming in. Let's say
- 4:16:52you have daily data coming in, right?
- 4:16:54You have one partition for each day. You
- 4:16:57would naturally want to run an optimize
- 4:17:00on that particular day on that
- 4:17:01particular partition. You don't want to
- 4:17:03run optimize on the full data. Yeah. So
- 4:17:06in those cases, we would want to use a
- 4:17:08predicate or a filter with the optimize.
- 4:17:13Yeah. So let's try to mimic that
- 4:17:15situation. And I'm going to use this
- 4:17:17same data frame df over here. So let's
- 4:17:21assume that I'm creating a new category
- 4:17:25called fruits
- 4:17:27and this is simply going to be so I'll
- 4:17:28just take up one category and rename it
- 4:17:30right uh just for the sake of
- 4:17:32simplicity. So
- 4:17:35df do.category category equals fruit.
- 4:17:38This is actually this I'll take up
- 4:17:41clothing all of the data that exist in
- 4:17:43clothing and I'm simply going to rename
- 4:17:46it
- 4:17:50as fruit.
- 4:17:55Okay. And this is simply going to be
- 4:17:57from pispar.sql
- 4:18:00functions import pispark.sqlfunctions
- 4:18:03lf.
- 4:18:05Yeah. So let's go ahead and run this. So
- 4:18:08basically what I've done is I've taken
- 4:18:09all of the data that exists for the
- 4:18:10clothing category and simply renamed
- 4:18:12clothing to fruits. And this is just to
- 4:18:15generate some sample data. So what this
- 4:18:18means is that a new category or a new
- 4:18:21partition is coming in which is by the
- 4:18:24name fruits. And now I simply want to
- 4:18:28optimize that partition. Yeah. So let's
- 4:18:32go ahead and first
- 4:18:35check that this only contains
- 4:18:40this only contains one category.
- 4:18:48Okay, so it simply contains one
- 4:18:50category. Right now let me go ahead and
- 4:18:53write this data. So I'm again going to
- 4:18:56do a repartition five. Okay, great.
- 4:19:03So, this should write a new partition to
- 4:19:07this table. Yeah. So, let's go ahead and
- 4:19:10run this.
- 4:19:13And now,
- 4:19:17I should see a new partition called
- 4:19:19fruits over here. Once this completes,
- 4:19:25okay, so the writing has completed. Let
- 4:19:27me refresh this.
- 4:19:29And I see a new category called fruits.
- 4:19:32Now let's go here. And I see five files
- 4:19:35over here. Yeah, as expected because I
- 4:19:38did a repartition because I also wanted
- 4:19:40to mimic the small file problem. So now
- 4:19:44what I'm simply going to do is
- 4:19:47I'm going to run a table dot optimize
- 4:19:53dot where the category
- 4:19:58the category equals fruits and then
- 4:20:00execute compaction right
- 4:20:03you can also use SQL in order to do this
- 4:20:06and for this time let's use SQL
- 4:20:10so the command is going to be something
- 4:20:13like this. Optimize
- 4:20:17delta catalog dot delta DB dot
- 4:20:21optimize example one where category
- 4:20:27equals fruit
- 4:20:29right let's go ahead and run this now
- 4:20:33okay so this is complete now I should
- 4:20:35see one extra file over here
- 4:20:39and you see that there is this one extra
- 4:20:42file that is created it. Yeah. And now I
- 4:20:45simply want to remove
- 4:20:48all of the tombstone file. Yeah. So
- 4:20:51again let's use SQL for now. Vacuum
- 4:20:55this table delta catalog delta DB dot
- 4:20:59optimize example one retain zero hours.
- 4:21:05So you get a gist of
- 4:21:08using both SQL and the delta table API.
- 4:21:11Yeah. So let's run this.
- 4:21:22Great. So now you see that we have only
- 4:21:25one file as expected. Yeah. So that is
- 4:21:28how you can use predicates when using an
- 4:21:32optimize command. So given that now
- 4:21:34we've seen the small file problem in
- 4:21:37action and how to fix them using the
- 4:21:40optimize command let's also understand
- 4:21:43the root cause right because many a
- 4:21:46times if you understand the root cause
- 4:21:48you may fix the problem right at the
- 4:21:51source or the origin right so let's have
- 4:21:54a look at the root cause one by one the
- 4:21:57first one is
- 4:22:00first of all let me quickly write root
- 4:22:02root causes and the first one is
- 4:22:06repartitioning to a very large number
- 4:22:10right the first one is a repartition so
- 4:22:14let's say that you have a file which is
- 4:22:1710 GB right and you end up
- 4:22:20repartitioning it into 10,000 parts
- 4:22:24right so what you do is you do a dot
- 4:22:26repartition
- 4:22:28you do a dot repartition of 10,000
- 4:22:33on whatever data frame that you're
- 4:22:34running. Right? So if your data set is
- 4:22:3810 GB in size, you are essentially going
- 4:22:42to end up with 10 into,000
- 4:22:45mgabytes. I'm assuming 1 GB to be,000
- 4:22:48megabytes for simplicity into 10,000
- 4:22:52which is simply 1 mgabyte in size.
- 4:22:56Right? So each of these partition is
- 4:22:58going to be one megabyte in size and
- 4:23:01that is going to cause the small size
- 4:23:03small file problem. Yeah. The second one
- 4:23:06the second issue is partitioning
- 4:23:10partitioning
- 4:23:13on a high cardality column. Right? You
- 4:23:17partition on a high
- 4:23:20cardality column.
- 4:23:24And by high cardality I simply mean that
- 4:23:26this column has a lot of distinct
- 4:23:29values. So let's imagine that you have a
- 4:23:33500 MB retail data set, right? You have
- 4:23:36a 500 mgabyte retail data set and you
- 4:23:41have a category column.
- 4:23:45So for some reason you want to partition
- 4:23:48by the category column. Let's say you
- 4:23:50are writing this data set and when you
- 4:23:52do a df dot write dot partition by
- 4:23:58dot partition by
- 4:24:00you basically put in the category column
- 4:24:04right and this uh category column has a
- 4:24:07lot of distinct values so let's say it
- 4:24:09has thousands of distinct values right
- 4:24:13so what you're going to end up with is a
- 4:24:16lot of small files and this again will
- 4:24:20again lead to the small file problem.
- 4:24:23The third one is frequently updated data
- 4:24:27sets, right? Frequently
- 4:24:30frequently updated data set,
- 4:24:34right? So let's say you have data coming
- 4:24:38in um continuously and by continuously
- 4:24:41let's assume that it comes in every 5
- 4:24:43minutes, right? So these data sets that
- 4:24:48come in every 5 minutes, they basically
- 4:24:51end up writing small small updates,
- 4:24:53right? And these small updates are going
- 4:24:57to be in the form of small files,
- 4:25:01right? Which is again going to lead to a
- 4:25:04small file problem, right? So as we can
- 4:25:06see that
- 4:25:08these three could be some of the most
- 4:25:12important root causes for the small file
- 4:25:16problem. And number one and number two
- 4:25:18these are purely technical in nature.
- 4:25:20Right? You can change the repartition
- 4:25:22number or you can change the
- 4:25:24partitioning column in order to avoid
- 4:25:26it. Right? Avoid the small file problem.
- 4:25:28But problems like this where let's say
- 4:25:31the business needs to see data
- 4:25:33instantly. Right? or they want the data
- 4:25:36to be updated frequently. In those
- 4:25:38cases, these problems are not technical
- 4:25:41in nature. They are more of a business
- 4:25:43problem. Right? So, how do you solve
- 4:25:45such kind of problems? So what you can
- 4:25:47do essentially is given that the team is
- 4:25:50going to use this underlying data set
- 4:25:53you can choose to run optimize after a
- 4:25:55number of hours or maybe every day so
- 4:25:58that at the end of the day you end up
- 4:26:00combining those small files into a
- 4:26:03larger file. Right
- 4:26:06now there are three methods or
- 4:26:08approaches that you can think about
- 4:26:10whenever you want to run the optimize
- 4:26:13command. Yeah. And the first approach or
- 4:26:16method is manual optimize. And we've
- 4:26:19seen it in the code example where we
- 4:26:23simply run the optimize command. We
- 4:26:26specify the table and if needed we
- 4:26:28specify a predicate. Right? So either
- 4:26:30you can run it like this or you can run
- 4:26:33it like this. Right? So this is the most
- 4:26:35simplest approach that is widely
- 4:26:38followed. Now the next approach is an
- 4:26:40automatic one and is called optimize
- 4:26:45right. Yeah, it's called optimize
- 4:26:50right.
- 4:26:52What optimize write does is that it
- 4:26:55simply combines all the small rights to
- 4:26:59a partition into a single write command.
- 4:27:03So we are going to understand what that
- 4:27:04means. But let's say we are doing the
- 4:27:07traditional right. What simply happens
- 4:27:10is that for creating one partition and
- 4:27:13let's imagine that this is the date
- 4:27:14partition something like date= 2025
- 4:27:1805 01. Yeah. So in order to create this
- 4:27:23partition there are several processes
- 4:27:26which are writing to this partition. So
- 4:27:28we see that executor number one is
- 4:27:31writing to this partition. Executor
- 4:27:33number two is also writing to this
- 4:27:35partition. Executor number three is also
- 4:27:39writing to this partition. And these may
- 4:27:42end up writing small paret files. Let's
- 4:27:44say this is 1.pk, this is 2.pk and
- 4:27:48similarly this is 3.pk. So these may end
- 4:27:51up writing small files thereby creating
- 4:27:54the small file problem. Right? So the
- 4:27:56root cause of it is several processes
- 4:27:59writing to a partition. Yeah. Now
- 4:28:02instead if you look at optimized right
- 4:28:05what going to happen is that all of the
- 4:28:08data is going to be shuffled and then is
- 4:28:11going to be executed as a single write
- 4:28:14command. So all of the data is shuffled
- 4:28:16over here from executor one executor 2
- 4:28:21and executor 3. All of the data is
- 4:28:24shuffled and then this is executed as a
- 4:28:28single write command. Yeah. So a very
- 4:28:32important point to note here is that
- 4:28:34it's executed as a single write command
- 4:28:37and this is going to produce
- 4:28:39appropriately sized files right so this
- 4:28:42is going to produce appropriately file
- 4:28:44size files in partition 1 2 and three
- 4:28:48yeah so that's the benefit of using
- 4:28:51optimize right however a tradeoff is
- 4:28:54that it's going to incur a shuffle
- 4:28:58is going to incur a shuffle and we know
- 4:29:01that shuffles are costly because it
- 4:29:03requires data transfer right so it's
- 4:29:06important to keep this in mind whenever
- 4:29:08we want to use an optimized right what
- 4:29:11are the thing that you want to optimize
- 4:29:13for yeah is it right latency if it is
- 4:29:17right latency then this might not be a
- 4:29:19good solution because
- 4:29:23there is going to be some shuffle
- 4:29:24involved and that is going to take time
- 4:29:27if it's optimizing for the small file
- 4:29:30problem then of course this is a very
- 4:29:33good solution. So now let's see optimize
- 4:29:36right in action and for that I'm going
- 4:29:38to use the same data frame over here. So
- 4:29:42let me quickly
- 4:29:45create a new table df dot
- 4:29:49repartition
- 4:29:51and for this time let me create lots of
- 4:29:54partition. So let me repartition by 288
- 4:29:58and then write dot mode override
- 4:30:02dot partition by let's partition by
- 4:30:05category
- 4:30:07and then let's save as table
- 4:30:14delta
- 4:30:16catalog dot deltadb dot optimize example
- 4:30:212. Yeah. So let's run this and on this
- 4:30:25table I also want to run this query
- 4:30:29the same query that I've run earlier and
- 4:30:31see how it performs.
- 4:30:34So this is going to be example two.
- 4:30:43Yeah.
- 4:30:45Okay. So now that is complete.
- 4:30:48I should see a table called optimize to
- 4:30:51over here. And if I go to the details,
- 4:30:58let me
- 4:31:00have a look at the table over here. And
- 4:31:05okay, so I basically see all of these
- 4:31:08partitions over here. Yeah. So ideally
- 4:31:10there should be some 288 partitions. So
- 4:31:13now
- 4:31:15let's go ahead and run this. And let me
- 4:31:18create another table. This time,
- 4:31:22this time after partition by I am going
- 4:31:27to write an option which is going to be
- 4:31:30optimize right to be true
- 4:31:35and then we are going to save this table
- 4:31:38as example three.
- 4:31:41Yeah. So let's go ahead and run this. So
- 4:31:44this operation has taken 6.97
- 4:31:48seconds. Yeah.
- 4:31:50Okay, so this is completed. Again, I
- 4:31:53should see a new table here which is
- 4:31:55optimize example three. And if I have a
- 4:31:58look at the table this time,
- 4:32:04I should see minimal partitions, right?
- 4:32:07Okay. So now when I look at these
- 4:32:09partitions,
- 4:32:13we see just one file.
- 4:32:17So that means the optimize write has
- 4:32:19worked very well and it is not writing
- 4:32:22small files. It has just written one
- 4:32:24file which is appropriately signed.
- 4:32:26Let's also run this operation once again
- 4:32:30and this time on example three.
- 4:32:43Perfect. This runs a lot quicker. If you
- 4:32:45see 6.97 seconds versus 1.82 seconds. So
- 4:32:48now we've seen an understood optimize
- 4:32:51rights and it's actually good for cases
- 4:32:54where many processes or executors are
- 4:32:57trying to write many files to a
- 4:32:59partition and instead of all of that we
- 4:33:02shuffle all of those files. We combine
- 4:33:04all of it and write it into appropriate
- 4:33:08file sizes for every partition. Right?
- 4:33:10So every partition is going to have an
- 4:33:13appropriately sized file and there is
- 4:33:16going to be an appropriate number of
- 4:33:18those files. Right? So this helps avoid
- 4:33:21the small file problem. But in cases
- 4:33:23where we are frequently writing small
- 4:33:26small updates to a table, right? The
- 4:33:29updates to a table are coming in small
- 4:33:31small chunks, the files that we get out
- 4:33:34of those small updates are still going
- 4:33:37to be small files, right? So we still
- 4:33:40end up with the small file problem and
- 4:33:43here is where autoco compaction comes
- 4:33:45into picture and that is the third
- 4:33:47approach that I wanted to talk about. So
- 4:33:50what autoco compaction does is that
- 4:33:53so what autoco compaction does is and
- 4:33:56let me let me quickly write something
- 4:33:58here is that let's say when you get
- 4:34:01small files right you get a few files
- 4:34:04after every write
- 4:34:07it is going to run a compaction
- 4:34:10operation
- 4:34:11right it is going to run an optimize
- 4:34:14command
- 4:34:16which is going to compact all of these
- 4:34:18files
- 4:34:19into appropriatelyized files
- 4:34:23and that is what autoco compaction
- 4:34:25exactly does. Right now it's really
- 4:34:27important to understand is that auto
- 4:34:29compaction is not going to run for an
- 4:34:32arbitrary number of file. Let's say you
- 4:34:34got two files and now you want autoco
- 4:34:36compaction to run and combine it into
- 4:34:38one file. Basically there is a setting
- 4:34:40which is called spark
- 4:34:42databasel.compact.min
- 4:34:45files minum file. Just Google it up. You
- 4:34:49can specify the minimum number of files
- 4:34:52that need to be present in order to
- 4:34:54trigger autoco compaction. Once that is
- 4:34:56there, it is going to compact all of
- 4:34:59those files into appropriately sized
- 4:35:02files. Right? So let's see autoco
- 4:35:05compaction in action right now. And we
- 4:35:08are going to use the same table example
- 4:35:11three as earlier. And before I can use
- 4:35:14that I need to do two things. The first
- 4:35:17one is to enable
- 4:35:20to enable
- 4:35:22autoco compaction
- 4:35:24and the second one is to disable
- 4:35:26optimize rate because if I don't disable
- 4:35:28optimize right it is going to take all
- 4:35:30the small file shuffle it into one and
- 4:35:33then write it as one file so I won't be
- 4:35:35able to see whether auto compaction is
- 4:35:37working or not. Yeah. So first of all I
- 4:35:45I disable
- 4:35:48disable this
- 4:35:53and then
- 4:35:56let's enable this
- 4:36:00actually let me change this auto comp uh
- 4:36:02autooptimize dot auto compact
- 4:36:06this is going to be true. Yeah, perfect.
- 4:36:14Okay, this is done. Now I want to check
- 4:36:17another property which basically tells
- 4:36:20me
- 4:36:21uh spark databreak dot delta
- 4:36:25dot autocompact
- 4:36:29dot min num files. This basically tells
- 4:36:33me what is the minimum number of files
- 4:36:37that I need in order to trigger
- 4:36:39autocompaction. So I actually recently
- 4:36:41changed it to three. The default number
- 4:36:45is 50. Yeah. So I changed it to three.
- 4:36:48You can change it to three using the
- 4:36:50following command
- 4:36:56using a set. Yeah. So that is how you
- 4:36:59can do that. Now I want to insert new
- 4:37:02partition, a new partition into this
- 4:37:04table. Yeah. And for that I'm going to
- 4:37:07use the fruits code again. So let's
- 4:37:10assume that there's new data coming in.
- 4:37:13And let me simply rename this to
- 4:37:15detergents.
- 4:37:17And let's also rename this to
- 4:37:19detergents.
- 4:37:21And again this is just for creating
- 4:37:23sample data. I am simply taking up all
- 4:37:26of the rows with respect to the clothing
- 4:37:28category. and then simply renaming
- 4:37:30renaming the clothing um clothing values
- 4:37:34to detergents. Right? So let's go ahead
- 4:37:37and run this. And now I want to run a DF
- 4:37:42detergent dot repartition five
- 4:37:47dot write dot mode overrite and this is
- 4:37:51going to be written as example three.
- 4:37:53Yeah. So let's go ahead and run this
- 4:37:56now.
- 4:37:59Okay, so this is complete. Let's check.
- 4:38:05Let's check what happened.
- 4:38:10Okay, so I see a detergent category.
- 4:38:15And I see 1 2 3 4 5 6 files, which is
- 4:38:20exactly what we wanted. The reason for
- 4:38:22that is it would have written five files
- 4:38:25initially
- 4:38:27because we had a repartition five and
- 4:38:30then it wrote an additional file because
- 4:38:34of the compaction operation. Yeah. So
- 4:38:36what you see over here is essentially
- 4:38:39there is one file which is bigger in
- 4:38:42size which is this one and this was the
- 4:38:44one that was written as a result of the
- 4:38:48auto compaction operation. So that is
- 4:38:50how autoco compaction works and all of
- 4:38:53these other files right that you see
- 4:38:54over here they have been tombstone when
- 4:38:57you run a vacuum all of them could be
- 4:38:59removed right so I hope that gave you a
- 4:39:01gist of how autocompaction works so now
- 4:39:04let's move over to the next topic which
- 4:39:06is called vacuum
- 4:39:09and we've already seen vacuum in action
- 4:39:13so whenever we delete update or merge
- 4:39:16records in a delta table, some of the
- 4:39:20underlying records are going to be
- 4:39:22removed, right? So when you update, it
- 4:39:25creates a new record and the old records
- 4:39:28becomes irrelevant. When you delete, all
- 4:39:31of those records become irrelevant,
- 4:39:33right? So because these records have
- 4:39:36been irrelevant, have become irrelevant,
- 4:39:38they're supposed to be deleted, right?
- 4:39:40But what actually happens in delta is
- 4:39:43that it doesn't physically delete it
- 4:39:46from the cloud storage or from the disk.
- 4:39:48It just marks it for deletion. Right? It
- 4:39:52just tombstones it or it does something
- 4:39:54which is called soft delete. Right? So
- 4:39:58all of these rows which are marked for
- 4:40:00deletion which is soft delete or the
- 4:40:03rows which have been tombstone they can
- 4:40:06only be removed once you run your
- 4:40:08vacuum. Right. So what essentially
- 4:40:11happens is let's say when you run a
- 4:40:13delete
- 4:40:16merge
- 4:40:19or let's say an update
- 4:40:23there are rows which are going to become
- 4:40:24irrelevant right and those rows are
- 4:40:28simply tombstoned
- 4:40:32or they are marked for soft deletion.
- 4:40:39Right? So the moment
- 4:40:41you apply a vacuum or let's say you want
- 4:40:45to remove these roles physically from
- 4:40:49the disk
- 4:40:52physically remove all of these roles
- 4:40:54either from the cloud storage or from
- 4:40:56the disk. Right? So in order to do that
- 4:41:00you need to run the vacuum command and
- 4:41:03vacuum is going to remove all of the
- 4:41:05rows which have been tombstone or have
- 4:41:08been marked for stop delete. Right?
- 4:41:10Essentially the same thing right. So
- 4:41:12that is how vacuum works and there are
- 4:41:15two important points to note. The first
- 4:41:17one is that vacuum helps you save on
- 4:41:20storage cost. It doesn't make your
- 4:41:23queries faster. It's a very popular
- 4:41:25misconception that running vacuum is
- 4:41:28going to make your queries run faster.
- 4:41:31That's not the case. The reason for that
- 4:41:33is we've seen in earlier examples that
- 4:41:38we had five files, right? We have
- 4:41:41written five files and then we run an
- 4:41:44optimize.
- 4:41:46We run an optimize and then it compacts
- 4:41:49all of the data into one file. Right? So
- 4:41:53now you end up having six files right
- 4:41:56five files which were already present
- 4:41:58earlier and additionally you have one
- 4:42:01file which were the result of the
- 4:42:04compaction operation right so now the
- 4:42:07five files have been marked for deletion
- 4:42:10right so this is the five file and this
- 4:42:12is one file this has been tombstoned
- 4:42:18and this one is active
- 4:42:21right so The moment you run a vacuum,
- 4:42:26it is going to delete all of these
- 4:42:29files,
- 4:42:30it is going to permanently physically
- 4:42:33delete all of these files and the file
- 4:42:36that you are left with is this one file.
- 4:42:39So the data you end up scanning is just
- 4:42:42the same, right? The filtering and all
- 4:42:45of that happens at a delta log level,
- 4:42:47right? Which files do I need to scan?
- 4:42:49That happens at a delta log level. And
- 4:42:52that is taken care of over there. Right?
- 4:42:54So essentially when we run vacuum we
- 4:42:57just save on storage cost. We are simply
- 4:43:00removing data and that is how it doesn't
- 4:43:05make your queries run faster. It simply
- 4:43:07removes data that is not needed. And the
- 4:43:10second point is that it limits your
- 4:43:13ability to time travel.
- 4:43:16Limits your ability to time travel. So
- 4:43:20if you run vacuum you will not be able
- 4:43:24to go back to any of the previous
- 4:43:26version. So let's see that with an
- 4:43:28example. So I'm simply going to take one
- 4:43:30of these files over here and let's
- 4:43:34create two data frames right equals
- 4:43:36spark dot
- 4:43:39park read.park park
- 4:43:42and this is going to be a filter f dot
- 4:43:46column customer id dot between this is
- 4:43:51going to be 101 and 150 right so this
- 4:43:55park file contains all customers from
- 4:43:57101 to 200 so I am simply selecting 101
- 4:44:01to 150 and this is going to be 151 until
- 4:44:05200 so this will be 151 comma 200 100
- 4:44:11and let's go ahead and run this.
- 4:44:14Let's now write this to a table.
- 4:44:18Yeah. So, this is going to be write do
- 4:44:23mode override.
- 4:44:26I don't need any partitioning.
- 4:44:29And let's just say
- 4:44:32vacuum example one.
- 4:44:35And let's just write it. Yep.
- 4:44:38Let's also see the history.
- 4:44:41Describe history
- 4:44:44delta catalog dot deltatb dot vacuum
- 4:44:50example one. So I should just see one
- 4:44:52row. Okay. Uh what is the issue here?
- 4:44:57Okay, that's a spelling mistake. Uh
- 4:45:00vacuum
- 4:45:01example one.
- 4:45:03And this is as expected. So now let's
- 4:45:06write the other data frame as well.
- 4:45:08Yeah.
- 4:45:09And this is simply going to be TF 151.
- 4:45:12And I'm going to append it.
- 4:45:17Let's go ahead and write it again.
- 4:45:21Okay, perfect. So now we see two rows.
- 4:45:24The first one is the create. The second
- 4:45:26one is the write operation where we
- 4:45:29wrote in additional 50 rows of data.
- 4:45:32Now let's go ahead and do a delete.
- 4:45:35Yeah. So actually before I do a delete,
- 4:45:38let me show you something. So this is
- 4:45:40the vacuum table.
- 4:45:44Let me copy the UU ID and
- 4:45:51okay. So I see two rows over here. Yeah.
- 4:45:53The first one which is 22 2123 the time
- 4:45:57at which it was written. This will be
- 4:46:00this file over here which contains all
- 4:46:03of the customers from 101 to 150. And
- 4:46:07the second one should be 151 to 200.
- 4:46:10Yeah. So the second one over here should
- 4:46:13be 151 to 200. Yeah. The one that starts
- 4:46:16with C82.
- 4:46:19Okay. Now let's go ahead and delete some
- 4:46:22rows. So we do a delete from delta
- 4:46:25catalog.
- 4:46:28Delta catalog that this where customer
- 4:46:31ID between
- 4:46:34151 and 200. So customer ID is from 151
- 4:46:39to 200 they basically reside in this
- 4:46:43park file right over here. Yeah. So when
- 4:46:45I run this statement ideally it should
- 4:46:48tombstone the second file. It should
- 4:46:50mark it for delete. Yeah. So now let's
- 4:46:54run this again.
- 4:46:56Let's see the history again.
- 4:47:00Perfect. Now you see three uh three
- 4:47:03statements or three versions. Yeah. The
- 4:47:05third one is the delete.
- 4:47:08After this, let's run a few quick
- 4:47:14statistics which is minimum of customer
- 4:47:20ID,
- 4:47:23comma, maximum of customer ID
- 4:47:27from this table
- 4:47:34vacuum example one. Yeah.
- 4:47:38What should we see? So we have
- 4:47:42all the customers from 101 to 200 and we
- 4:47:46have removed 151 to 200. Right? So we
- 4:47:50have removed 151 to 200. So we should
- 4:47:54see the minimum to be 101 and the
- 4:47:56maximum to be 150.
- 4:47:59And that is what you see over here. Now
- 4:48:01if I were to run the same statement
- 4:48:05on version number
- 4:48:08version number
- 4:48:10one, we ran a delete on version number
- 4:48:14two. If I were to run the same statement
- 4:48:16on version number one, it would have all
- 4:48:19the role from 101 to 200. Right? So the
- 4:48:22result would be a little different. The
- 4:48:23minimum would be 101, but the maximum
- 4:48:26would be 200. Yeah.
- 4:48:30Perfect.
- 4:48:32So now what we want to do is we want to
- 4:48:35run a vacuum. Yeah, we want to run a
- 4:48:38vacuum in order to remove the tombstone
- 4:48:41file over here. We want to remove this
- 4:48:44file over here which has been tombstone.
- 4:48:47So, let's go ahead and run a vacuum
- 4:48:55delta catalog dot
- 4:48:58delta db dot vacuum retain zero hours.
- 4:49:02Yeah. So, let's go ahead and run this.
- 4:49:04And as expected, I got a message saying
- 4:49:07that are you sure you want to delete
- 4:49:09this? And we've seen this earlier as
- 4:49:10well. So let's go ahead and
- 4:49:14disable this check over here. So we
- 4:49:17simply do this equals false.
- 4:49:20And let's run this again.
- 4:49:24Okay. So now the vacuum operation is
- 4:49:27complete and I should just see one file
- 4:49:29over here. The one that starts with 8
- 4:49:32EF, right? The second one should be
- 4:49:34gone. Perfect. So that is what was
- 4:49:38expected. Right? So now that the vacuum
- 4:49:41operation is complete, if I were to run
- 4:49:44a history, if I were to see the history
- 4:49:47of this table,
- 4:49:49describe history,
- 4:49:52I am basically going to see two
- 4:49:54additional rows.
- 4:49:56The first one is for vacuum start and
- 4:49:59vacuum end. Right? So now
- 4:50:02we see over here that we only have one
- 4:50:06park file and this corresponds to the
- 4:50:08data from customer id 101 to 150. Right?
- 4:50:14So the experiment what I want to do is
- 4:50:18select star from this table
- 4:50:23version
- 4:50:25as of yeah and I want to
- 4:50:29play around with these numbers a little
- 4:50:30bit. So let's say if you do version as
- 4:50:33of two what should you get? Should you
- 4:50:36get any data or not? That's the first
- 4:50:38question. Okay. So ideally you should
- 4:50:41get data because
- 4:50:44after this delete all of the data from
- 4:50:48151 to 200 was deleted right. So the
- 4:50:50data from 101 to 150 was still present
- 4:50:55right and we still have the file that is
- 4:50:58needed for giving us the data from 101
- 4:51:02to 150. Yeah. So this operation should
- 4:51:05just work fine
- 4:51:08and it works fine. Yeah. So what will
- 4:51:11happen if you do version as of one? Will
- 4:51:14this work?
- 4:51:16Take a guess.
- 4:51:21Okay, this doesn't work because in
- 4:51:24operation in version number one, we
- 4:51:26wrote a file which contained data from
- 4:51:29151 to 200 and that file has gone
- 4:51:31missing. So that is why the the data
- 4:51:34that was contained in version number one
- 4:51:37is incomplete and that is the reason why
- 4:51:39we cannot go back in time in order to
- 4:51:41access this. Yeah. Let's see about
- 4:51:43version number zero. Again take a guess
- 4:51:45what would happen over here.
- 4:51:49Okay. You already have the park file for
- 4:51:52version number zero because we wrote the
- 4:51:54data from 101 to 150 and that can be
- 4:51:57accessed because of the park file that
- 4:51:59we see over here. Yeah. So that is how
- 4:52:03uh time travel is going to work with
- 4:52:05vacuum and you need to be a little
- 4:52:07careful if we want to go back in time
- 4:52:10and access data points right so we need
- 4:52:12to run vacuum very carefully so the next
- 4:52:15optimization technique that we are going
- 4:52:17to talk about is Z order yeah it's
- 4:52:21called Z order
- 4:52:25but before we understand what exactly
- 4:52:27this is let's actually set some ground
- 4:52:30let's set some context with this
- 4:52:32example. Yeah. So let's say you have a
- 4:52:35bunch of files that you see over here 1
- 4:52:382 3 and four.par park and these are
- 4:52:41currently sitting on your disk either on
- 4:52:43your disk or on some cloud storage and
- 4:52:47you basically want to process them right
- 4:52:51you want to process them and in order to
- 4:52:53process them these have to be loaded in
- 4:52:56the memory of whatever compute you're
- 4:52:58using right so these have to be loaded
- 4:53:01in memory now in order to load it into
- 4:53:04memory there has to be a costly data
- 4:53:07transfer over the wire Right. So in
- 4:53:09order to do that, let's first take this
- 4:53:12example where we have a query coming in
- 4:53:15from the user. The user basically asks
- 4:53:18that give me all of the records where
- 4:53:21age is greater than equal to 5 and less
- 4:53:24than equal to 10. Yeah. So for each of
- 4:53:27the park files that we have over here,
- 4:53:29we have stored some statistics and we
- 4:53:33already know that all of these
- 4:53:34statistics are stored in delta log,
- 4:53:37right? the minimum, maximum, count of
- 4:53:39values for all of the columns. Right?
- 4:53:42So, we've put the statistics over here
- 4:53:45and we will use this in order to figure
- 4:53:49which of the files need to be
- 4:53:51transferred over the wire. Yeah. So, we
- 4:53:54see that the minimum age and the maximum
- 4:53:57age is 4 and 12 which overlaps with this
- 4:54:00value. Right? So, that means that this
- 4:54:03file is going to be transferred. This
- 4:54:06file is going to be transferred over the
- 4:54:08wire. Here we see that the minimum is
- 4:54:1010, maximum is 28 and that overlaps with
- 4:54:13this value. So this file is also going
- 4:54:16to be transferred. Again the minimum is
- 4:54:186 and 14. That also overlaps with this
- 4:54:21value. Yeah, with this value here to
- 4:54:24here and this file is also going to be
- 4:54:27transferred. And this is a wide range
- 4:54:29from 4 to 60. And this again overlaps
- 4:54:32with this value over here. That means
- 4:54:354.park par is also going to be
- 4:54:37transferred. Yeah. So essentially what
- 4:54:40we see is that when such a query comes
- 4:54:44in all of these four files are going to
- 4:54:47be transferred. Yeah. These are going to
- 4:54:50be transferred over the wire
- 4:54:54and this is a costly operation.
- 4:55:00This is a costly operation. Now if we
- 4:55:03were to take a step back, do you think
- 4:55:06this transfer can be avoided? Instead of
- 4:55:09transferring all of the four files, can
- 4:55:12we transfer fewer files? Is it possible?
- 4:55:15Take a moment and think about it. So
- 4:55:17with this data layout, it's actually
- 4:55:20impossible to avoid sending any of the
- 4:55:23files over the network, right? We cannot
- 4:55:25omit to send any of the files over the
- 4:55:28network. We have to send all four of
- 4:55:30them. If we omit any of them then we
- 4:55:32won't get the right answer to this query
- 4:55:35that the user has asked right so what is
- 4:55:38the alternative how can we avoid sending
- 4:55:42any of the files over the network is
- 4:55:44that even possible how do we optimize
- 4:55:46for that right so in order to do that
- 4:55:49that is where zorder comes in so first
- 4:55:52of all let me take some of these numbers
- 4:55:54the minimum and maximum right so let's
- 4:55:57say we have four over here we have 12
- 4:56:00over here and of course between them
- 4:56:02there is going to be several keys
- 4:56:04several values for age which I'm not
- 4:56:06writing down but we will have all of
- 4:56:09those values right uh and then we have
- 4:56:1110 and 28 so 10 is going to be over here
- 4:56:15and let's say 28 is going to be over
- 4:56:17here and then we have 6 and 14 so 6 is
- 4:56:21going to be over here 14 is going to be
- 4:56:23over here and then we have four and 60
- 4:56:26I've already noted down four so 60 is
- 4:56:28going to be somewhere over here Right?
- 4:56:30And between all of them, right? There is
- 4:56:32going to be some some keys which we do
- 4:56:35not know but we are aware that there is
- 4:56:37going to be data in between them. Right?
- 4:56:40So this is a sorted
- 4:56:43order
- 4:56:44that I've produced. Right? So here is
- 4:56:47something that we can apply in order to
- 4:56:49skip files. And first of all let me copy
- 4:56:53this.
- 4:56:56Let me copy this here.
- 4:57:02Yeah. So what essentially I'm going to
- 4:57:05do is that
- 4:57:08I am going to take the first few values
- 4:57:12which is from 4 to 10 and I'm going to
- 4:57:15put it into one file which is 1.pk perk
- 4:57:21and then I am going to take other few
- 4:57:24values which is from 10 to 28
- 4:57:28and then I'm going to put it into 2.par.
- 4:57:34The intuition behind this is to avoid
- 4:57:38overlaps between the file. Right? If
- 4:57:40there are minimal overlaps probably I'm
- 4:57:43going to select only one file where all
- 4:57:45of my data is going to the side. Yeah,
- 4:57:48the next file I'm going to choose is so
- 4:57:50probably we'll have this number 29 over
- 4:57:52here and then 50 somewhere down the line
- 4:57:55over here,
- 4:57:57right? So 29 to 50 I am going to place
- 4:58:00this in three.park
- 4:58:04and then we are going to have this
- 4:58:05number 51.
- 4:58:07Let me place it over here. I'm going to
- 4:58:10have this number 51 and until 60 I'm
- 4:58:13going to take all of this data and then
- 4:58:15put it in 4.pk. park.
- 4:58:18The goal and idea behind this is to
- 4:58:21avoid overlaps between the file. The
- 4:58:24lesser overlaps are the lesser files you
- 4:58:28need to scan. Right? You will be able to
- 4:58:30prune the file. You would be able to say
- 4:58:32that okay, this is the file I need and
- 4:58:34let's send this over the network. Right?
- 4:58:37So, Gorder is basically a sort and
- 4:58:40repartition. So, here all of the data
- 4:58:42has been sorted by whatever key we
- 4:58:45wished.
- 4:58:47Right? In this case, the query that has
- 4:58:49come in and we basically repartition the
- 4:58:52data into those number of files. Right?
- 4:58:56So that is the logic and the idea behind
- 4:58:58reorder. And what it does is that in
- 4:59:00these files it simply colllocates data.
- 4:59:05By colllocate data what I mean is
- 4:59:07similar data is placed in the same file.
- 4:59:11So similar data is placed in the same
- 4:59:14file. Now let's go ahead and have a look
- 4:59:17at this how how this query is going to
- 4:59:19behave now. Yeah. So now
- 4:59:23where age is greater than equal to 5 and
- 4:59:25less than equal to 10. Does it overlap
- 4:59:27with this? Yes. That means this file is
- 4:59:30going to be transferred. Does it overlap
- 4:59:32with this which is 2.par?
- 4:59:35Yes, it does. This value and this value
- 4:59:37overlaps. That means this park is going
- 4:59:40to be transferred over the wire. But
- 4:59:42does it overlap with these two? The
- 4:59:43minimum itself is 29 and 51. Over here
- 4:59:46it doesn't. So that means we've avoided
- 4:59:49sending these two files over the wire.
- 4:59:52It has been pruned. It has been filtered
- 4:59:55that we don't need these two files.
- 4:59:57Right? And that is one of the benefits
- 4:59:59and beauty of the order. It collocates
- 5:00:02it collocates similar data in the same
- 5:00:05files due to which this becomes
- 5:00:07possible. So to keep things simple I
- 5:00:10have taken a single dimensional example
- 5:00:13wherein we refer to the column age but
- 5:00:16this very well applies to multiple
- 5:00:18dimensions right where we would like to
- 5:00:21filter by multiple columns and the idea
- 5:00:23is very simple the idea is to take all
- 5:00:25of these columns that we want to filter
- 5:00:28by and then map it to single dimension.
- 5:00:31So all of these column they are
- 5:00:33multi-dimensional right more than one
- 5:00:35columns it's multi-dimensional.
- 5:00:37So we take all of these columns that we
- 5:00:41want to filter by, right? So we take all
- 5:00:44of these columns and then we map it to a
- 5:00:47single dimension.
- 5:00:50We map it to a single dimension. And how
- 5:00:52does this happen? It happens when we put
- 5:00:55this inside the zorder function, right?
- 5:00:57When we apply Z order on the set of
- 5:01:01columns, right? And this basically
- 5:01:03preserves the locality. When we apply
- 5:01:06Zorder on the set of columns, it
- 5:01:09arranges the data layout in such a way
- 5:01:11so that the locality is preserved. And
- 5:01:14by locality, what I mean is that similar
- 5:01:17data points are close to each other. So
- 5:01:20Z order basically
- 5:01:23preserves locality.
- 5:01:28And by preserving locality, what we
- 5:01:30essentially mean is that similar points,
- 5:01:35similar data points are in the same
- 5:01:37file.
- 5:01:39Actually, not necessarily always in the
- 5:01:42same file because in the earlier example
- 5:01:44that we saw the value the age 10 was in
- 5:01:471.park and then the age 10 was also in
- 5:01:512. Okay, the idea is to avoid overlaps
- 5:01:55to keep overlaps as minimal as possible,
- 5:01:58right? Because we also don't want one
- 5:02:01file to has have enormous amount of
- 5:02:03data. Yeah. So, I'm going to take
- 5:02:05another example where
- 5:02:07we are going to have two columns which
- 5:02:09is product ID and quantity and for some
- 5:02:13reason we write a lot of queries, a lot
- 5:02:16of filters on these two columns. Yeah.
- 5:02:19So we write a lot of filters on these
- 5:02:21two column. So what we want to do is
- 5:02:23that we want to zorder
- 5:02:26we want to zorder
- 5:02:29you want to zorder by product ID
- 5:02:33and the quantity.
- 5:02:35So when we essentially do this what it
- 5:02:38does is that it takes these two points
- 5:02:41from the two-dimensional space and then
- 5:02:44it applies a zorder. What zorder is is
- 5:02:48is it's essentially a space filling
- 5:02:50curve which maintains locality. Yeah. So
- 5:02:53just to iterate again what it does is
- 5:02:55that it is going to map this
- 5:02:57twodimensional value which is let's say
- 5:03:00this value this value and all the other
- 5:03:02values that you see over here. It is
- 5:03:04going to map two-dimensional values to a
- 5:03:07single dimensional value.
- 5:03:12Yeah. And the beauty of this curve this
- 5:03:15curve zord order is that the points
- 5:03:17which are closed in the two-dimensional
- 5:03:20space they are also going to be closed
- 5:03:23in the single dimensional space. So the
- 5:03:26points which are closed in two dimension
- 5:03:29when you map it to a single dimension
- 5:03:31they are also going to be closed. Yeah.
- 5:03:33So for argument sake in order to make
- 5:03:36you in order to help you understand this
- 5:03:38better let's say what the order does is
- 5:03:41that it simply adds values right and
- 5:03:44this is just to make you understand
- 5:03:45things better yeah so let's take all
- 5:03:47these values the values the the values
- 5:03:52which have product ID equals 10 and
- 5:03:54let's simply add them right so this is
- 5:03:56going to be 15 this is going to be 19
- 5:04:01this is going to be 18 This is going to
- 5:04:04be 22
- 5:04:06and this is going to be 14. This is
- 5:04:10going to be 12 and this is going to be
- 5:04:1211. Now what I do is that I start
- 5:04:16putting the data points from the lowest
- 5:04:18to the highest by Z order. Right? So the
- 5:04:21lowest is this one. I basically put 10
- 5:04:23and 10 and one over here. Then we have
- 5:04:2612. I put 10 and two over here. Then we
- 5:04:29have 14. I put 10 and four over here.
- 5:04:31Similarly 10 and 5
- 5:04:34and then 10 and 8 this one over here and
- 5:04:38then 10 and 9 and then finally 10 and
- 5:04:4012. So what we see is that points which
- 5:04:43are close
- 5:04:45points which are close in the two
- 5:04:46dimensional space which is 10 and 1.
- 5:04:48This is very close to 10 and 2. They are
- 5:04:52also close in the single dimensional
- 5:04:55space after doing a zord. Right? So 11
- 5:04:58and 12 are closed. So if we put the
- 5:05:01values over here they are closed these
- 5:05:03are placed close to each other right
- 5:05:05just one row apart right so this makes
- 5:05:08sure zorder makes sure that when points
- 5:05:11are reduced from a multi-dimensional
- 5:05:13space to a single dimensional space they
- 5:05:17are collocated right they are close to
- 5:05:19each other and using this using this we
- 5:05:24simply create the files we place all of
- 5:05:26this in one file over here and then all
- 5:05:28the other data which is again
- 5:05:30colloccated
- 5:05:32is placed in this file over here. Now we
- 5:05:34don't want to create one paret file for
- 5:05:3611 and 12 and 15 because of course we
- 5:05:40don't want to end up with the small file
- 5:05:41problem. We want the data to be
- 5:05:44appropriately sized. Yeah. So that is
- 5:05:46why we don't create one file for each of
- 5:05:49these values. So I hope this gives you
- 5:05:52some kind of intuition on how zorder
- 5:05:55works internally. It basically brings
- 5:05:58similar data in the same file. Sometimes
- 5:06:02not always in the same file because
- 5:06:03again as I discussed earlier we don't
- 5:06:05want to put too much of data in the same
- 5:06:08file. Right? If we had a lot of rows for
- 5:06:11number 10 over here, product ID number
- 5:06:1310 over here. Probably some of them
- 5:06:15would have also have gone over here in
- 5:06:17the second file. Right? Because we don't
- 5:06:19want one file to have enormous amount of
- 5:06:22data. We want it to be appropriately
- 5:06:25sized. Yeah. So it collocates similar
- 5:06:28data in the same files and the goal is
- 5:06:32to skip to skip as much amount of files
- 5:06:36and data points as possible and that is
- 5:06:39going to help us run the queries faster
- 5:06:42scan lesser number of files and send
- 5:06:45lesser amount of data over the wire.
- 5:06:47Okay. So let's see zorder in action and
- 5:06:50for that I'm going to use this file
- 5:06:52again. So let's go ahead and read this
- 5:06:55park for read.par k and this file and
- 5:07:00this time I am going to select few
- 5:07:02columns only it is going to be customer
- 5:07:05id and the category price quantity and
- 5:07:09invoiced perfect. So let's go ahead and
- 5:07:12run this and let me do a dm dotlimit of
- 5:07:16five.
- 5:07:19Okay. And I want to generate a lot of
- 5:07:23rows. So what I'm simply going to do is
- 5:07:28uh I'm going to run a loop. Expected
- 5:07:31rows equals
- 5:07:3420 million.
- 5:07:37So we simply do 2000 0 0 0.
- 5:07:42Okay. And now while dfun dot count
- 5:07:48is less than equal to expected rows and
- 5:07:52df
- 5:07:53union equals
- 5:07:56df union. So I keep unioning the same
- 5:07:59data frame in the loop and
- 5:08:03the count is simply df union dot count
- 5:08:09and I copy this and put the
- 5:08:14final count to be this one. Okay. Now I
- 5:08:17also want to write this data frame
- 5:08:22to a table.
- 5:08:24save as table in our delta catalog dot
- 5:08:28delta db dot zorder example one. Yeah.
- 5:08:34Okay. So this is complete and now we are
- 5:08:38writing it to this table. While this is
- 5:08:41going on I want to write a query
- 5:08:46basically to see how long does it take
- 5:08:48to run. Right? And this is simply going
- 5:08:52to be uh
- 5:08:58select
- 5:09:00category comma
- 5:09:04category comma sum of price into
- 5:09:08quantity.
- 5:09:10Price into quantity
- 5:09:14as total sales,
- 5:09:17right? as total sales from this table
- 5:09:21right here
- 5:09:24from this table
- 5:09:26where customer ID
- 5:09:29equals 2011
- 5:09:32group by category
- 5:09:35right so I'm purposely
- 5:09:37doing a filter on customer ID equals 201
- 5:09:42and we are going to compare it compare
- 5:09:44the performance of this query after a
- 5:09:47the order has been applied. Yeah. Okay.
- 5:09:50So the write is complete. Now let's go
- 5:09:53ahead and run this query. And let me
- 5:09:56also Okay. So the wall time is 785
- 5:09:59milliseconds.
- 5:10:00And let me also
- 5:10:04see how this table looks like.
- 5:10:08I'm going to take the location from here
- 5:10:12and
- 5:10:15basically check over here.
- 5:10:18Okay, so here it is and then it created
- 5:10:21a lot of files
- 5:10:23and let me download the delta log over
- 5:10:25here
- 5:10:29and let's open it. So it added how many
- 5:10:33it added?
- 5:10:35255 park files. Yeah. And what we see
- 5:10:39over here is that each of these files
- 5:10:42have a minimum value for customer ID
- 5:10:44which is 201
- 5:10:46and a max value of 99457.
- 5:10:50Right? So if I write a query which is
- 5:10:53something like this where customer ID
- 5:10:56equals 201 it is going to scan all of
- 5:11:00those files all of these 255 files
- 5:11:03because the minimum value of customer ID
- 5:11:05in each of these file is 201 from 2019
- 5:11:104557. So it ends up scanning four uh 255
- 5:11:14files. Yeah. Now, now let's go ahead and
- 5:11:18run an optimize, right?
- 5:11:21A Z order with an optimize. So, optimize
- 5:11:24this table.
- 5:11:27Optimize this table. Z order by customer
- 5:11:31ID. Now, Z order and optimize go hand in
- 5:11:35hand, right? Whenever you want to run a
- 5:11:37Z order by, you run it together with an
- 5:11:39optimize. Optimize is going to create
- 5:11:43larger files, right? It is going to
- 5:11:45merge all of the small files into a
- 5:11:47large bin or a large file and while this
- 5:11:50process happens while it is putting the
- 5:11:52small files together zorder is going to
- 5:11:56colllocate similar data. Yeah. So that
- 5:11:58is the reason why they work hand in
- 5:12:00hand. They work together. So you can
- 5:12:03either use this SQL or you can also do
- 5:12:07this. So from delta dot from delta
- 5:12:12import
- 5:12:14delta.ts
- 5:12:15import data table
- 5:12:18and we simply do table equals
- 5:12:22delta table dot for name and then table
- 5:12:26dot optimize
- 5:12:33table.optimize optimize dotexecute
- 5:12:35zord order by and then we put in the
- 5:12:38customer ID over here. So you can use
- 5:12:40any of these two approaches. Yeah.
- 5:12:44So I'll comment this for now and I will
- 5:12:47go ahead and run this.
- 5:12:51Okay. So the Z order is complete. And if
- 5:12:54I were to look at some of the metrics,
- 5:12:56what it says is that number of files
- 5:12:59added is two and it removed 256
- 5:13:04files. Yeah. So we're going to look at
- 5:13:07all of that. But before we do that, let
- 5:13:10me run this query again and let's
- 5:13:13compare the run times. Yeah. So let's
- 5:13:16run this query. And it is 144
- 5:13:19milliseconds and the last time was 785
- 5:13:22milliseconds. Six close to six times of
- 5:13:26improvement. Yeah. And that's a
- 5:13:28significant improvement. Let's also see
- 5:13:30what happened behind the scenes. So if I
- 5:13:34just refresh this and if I download this
- 5:13:37log,
- 5:13:39let's download this and let's have a
- 5:13:41look.
- 5:13:42So all of these files which were earlier
- 5:13:45added there is a remove for all of this
- 5:13:48right and we have add for only two of
- 5:13:52them that means two new park files have
- 5:13:55been added and let's quickly have a look
- 5:13:57at the statistics so minimum value for
- 5:14:01the customer ID is from 2011
- 5:14:04until 48632
- 5:14:08and the other file is from 48 632 until
- 5:14:1399457. Right? So this is something
- 5:14:16really amazing that has been done. What
- 5:14:18it first did is that it removed all of
- 5:14:22the 255 or 256 whatever number of files
- 5:14:25were there. It removed all of it and
- 5:14:27then it compacted the data into two
- 5:14:29files. That is where the role of
- 5:14:31optimize comes in and then similar data
- 5:14:34was colllocated. Right? It collocated
- 5:14:37data from 2011 until 48632.
- 5:14:42The remaining was put in the other park
- 5:14:45file. Now I don't need so initially what
- 5:14:49I was doing is okay I don't have that
- 5:14:51statistics over here. Let me go over
- 5:14:53here. Initially what I was doing is I
- 5:14:55was scanning all of the files right I
- 5:14:58was scanning customer ID 220 2016 2011
- 5:15:02in all of the files right because the
- 5:15:04minimum value told me to do so now I
- 5:15:07just need to scan just one file right
- 5:15:11and that's a remarkable improvement so
- 5:15:13that is how zorder works and I hope that
- 5:15:16gave you some insight into all of this
- 5:15:18right so now let's talk about how do you
- 5:15:21apply zorder to together with hive style
- 5:15:25partitions. Yeah. So let's let's try to
- 5:15:28mimic that example and I'm going to use
- 5:15:30this data frame over here. So this is
- 5:15:34going to be let's try to create those
- 5:15:36partitions
- 5:15:38dot mode override dot partition by and
- 5:15:42let's partition by invoice
- 5:15:45invoice date and let's save this table
- 5:15:47as delta catalog dot delta db dot zorder
- 5:15:53example 2. Yeah. So let's go ahead and
- 5:15:56run this and it should basically create
- 5:15:59partitions on the invoice date. Yeah. So
- 5:16:02now we've earlier seen that we applied Z
- 5:16:06order on the customer ID. So generally
- 5:16:10there's a pattern that people follow is
- 5:16:12if we have a hive style partition on
- 5:16:16let's say the invoice date and our we
- 5:16:20can use this together.
- 5:16:23We can zorder by
- 5:16:26we can zorder by the customer ID for
- 5:16:30each of those invoice date.
- 5:16:34Yeah.
- 5:16:39So we can zorder by customer ID for each
- 5:16:43of those invoice days. Yeah. So now
- 5:16:49let's imagine that a new partition is
- 5:16:51coming in and let me quickly
- 5:16:55write some code to mimic that. So we
- 5:16:59going to do dfm do.filter f dot column
- 5:17:02invoice
- 5:17:04date equal 2023 0101. So I basically
- 5:17:09take in all of the data for that day and
- 5:17:11I just rename it.
- 5:17:14I rename it
- 5:17:18to
- 5:17:232025 0504.
- 5:17:27Yeah. So I'm basically taking in all of
- 5:17:29the data for someday and then I'm just
- 5:17:31renaming it to today's date and this is
- 5:17:34simply to create some new data mimicking
- 5:17:37the fact that new data came in today and
- 5:17:40now I want to write uh write a new
- 5:17:43partition to my table. Yeah. And that
- 5:17:46table is basically this table. So let's
- 5:17:49also see how this looks like.
- 5:17:52So,
- 5:18:02so basically this table has all the
- 5:18:04invoice dates, right? And
- 5:18:07new data has come in for today and I
- 5:18:10want to basically
- 5:18:14write this data to this table. Let's
- 5:18:18quickly check that there are records
- 5:18:21over here in this table
- 5:18:24and then okay perfect we do have
- 5:18:26records. So now let's go ahead and
- 5:18:29simply write this right do mode append
- 5:18:34and then
- 5:18:39partition by
- 5:18:43invoice date and then save as this
- 5:18:45table. Right? So this would mean that we
- 5:18:47just wrote in a new partition. And let's
- 5:18:50also verify that select max of
- 5:18:54invoice
- 5:18:56invoice date from
- 5:18:59this table delta catalog. Delta DB dot Z
- 5:19:04order example 2. And that should show us
- 5:19:07this new date that we put in which is
- 5:19:08the same. Now let's go ahead and apply a
- 5:19:13Z order.
- 5:19:15So we are going to use it together with
- 5:19:17an optimize
- 5:19:21Z order by customer ID but we only want
- 5:19:27to do it for
- 5:19:30the current invoice date. Yeah. So the
- 5:19:34current invoice date equals
- 5:19:362025
- 5:19:3905 04. Yeah. So we have our hive style
- 5:19:45partitions and given that we know that
- 5:19:47our queries for whatever reason ei is a
- 5:19:49lot on customer ID we zorder by customer
- 5:19:52ID and that is this is also a pattern
- 5:19:54which is widely used and in your
- 5:19:56pipeline when you are running your
- 5:19:58pipelines on a daily basis you can
- 5:19:59simply parameterize this you can simply
- 5:20:03parameterize this part of the command
- 5:20:05you can simply put in current day
- 5:20:10current day minus one right because
- 5:20:12let's Okay, you run this pipeline on
- 5:20:15yesterday's data because you have all of
- 5:20:17that data. Yeah. Now let's talk about
- 5:20:19the final optimization technique which
- 5:20:21is liquid clustering. Yeah. So by now
- 5:20:25we've seen things like high style
- 5:20:28partitioning where each partition value
- 5:20:30gets its own folder and all of the
- 5:20:32records for that partition value goes to
- 5:20:34that folder and this essentially speeds
- 5:20:37up queries. On the other side we have
- 5:20:39zorder which optimizes your data layout.
- 5:20:43It colllocates data so that we can pick
- 5:20:45up relevant files right we pick up
- 5:20:49minimum number of files and then we scan
- 5:20:51them right so all of this makes your
- 5:20:54queries faster however the biggest
- 5:20:57problems with this approach is that
- 5:20:59they're not flexible yeah so let me note
- 5:21:02all of this down so the first one is
- 5:21:05hive style partitioning
- 5:21:09hive style partitioning and the Second
- 5:21:12one is Z order.
- 5:21:16Now the problems with these two
- 5:21:18approaches is that they are not
- 5:21:21flexible.
- 5:21:23They are not flexible. And when I say
- 5:21:25that they are not flexible, what I mean
- 5:21:28is that you have to decide your
- 5:21:30partitioning column or Z order column up
- 5:21:34front.
- 5:21:36Right? These two have to be decided up
- 5:21:40front. Either whenever you do the
- 5:21:42partition by and the column name or the
- 5:21:45Z order by and the column name the
- 5:21:47column has to be decided up front.
- 5:21:50Right? Now the disadvantage with that is
- 5:21:52if today your filter pattern is by a
- 5:21:56column called country
- 5:22:00and few months down the line it changes
- 5:22:02to another column called category
- 5:22:06called category.
- 5:22:09The layout of your data is by country as
- 5:22:12of now, right? It is optimized in such a
- 5:22:15way such that whenever somebody queries
- 5:22:18by country, it's very easy to find the
- 5:22:20records. But when the filter pattern
- 5:22:23changes over time, when it changes to
- 5:22:25category, that data layout is no more
- 5:22:28helpful. Right? So essentially what you
- 5:22:30end up doing is you'll probably end up
- 5:22:33scanning the whole data set and the data
- 5:22:36layout is no more helpful. What you
- 5:22:38essentially need to do is you'll end up
- 5:22:43rewriting
- 5:22:46rewriting data according to the new
- 5:22:49filter pattern which is by category.
- 5:22:53Right? So that is the biggest
- 5:22:56disadvantage of these two approaches.
- 5:22:58Now this is where liquid clustering
- 5:23:00comes in. Right? So it allows you to
- 5:23:03change the clustering columns anytime.
- 5:23:06Yeah. So it allows you to change the
- 5:23:09clustering columns.
- 5:23:12So this can be changed anytime
- 5:23:16right and that is the biggest benefit of
- 5:23:20using liquid clustering. It is very
- 5:23:23flexible in nature. It is incremental
- 5:23:26meaning that you can change the
- 5:23:27clustering columns going ahead. Yeah, it
- 5:23:30is going to use. So let's say you have
- 5:23:32particular set of clustering columns and
- 5:23:34down the line when you want to change
- 5:23:36it, it is going to use those new
- 5:23:38clustering columns in order to decide
- 5:23:41the layout of data. So the layout of
- 5:23:43data is then going to change going
- 5:23:46ahead. So the algorithm that liquid
- 5:23:48clustering uses under the hood, it
- 5:23:51maintains a balanced layout and by
- 5:23:54balanced layout what I mean is that it
- 5:23:57ensures two things. The first one is
- 5:24:00uniform file size.
- 5:24:02Uniform file size. And the second one is
- 5:24:07appropriate number of files.
- 5:24:11Appropriate number of files. Right? So
- 5:24:14it is going to ensure that appropriate
- 5:24:17number of files are created. It doesn't
- 5:24:19end up creating lots of file with
- 5:24:21minimal amount of data. So that is where
- 5:24:23you avoid the small file problem. And it
- 5:24:26also ensure that those files are
- 5:24:30appropriately sized. Right? So it ensure
- 5:24:33the number of files and the size of
- 5:24:35those files. And a good point to note is
- 5:24:38that in all of these files your data is
- 5:24:41colllocated
- 5:24:45which simply means that similar data is
- 5:24:48going to reside in the same file. Yeah.
- 5:24:52So let's understand this with an example
- 5:24:54and I'm going to refer to Denny Lee's
- 5:24:56blog who is a developer advocate at data
- 5:24:58bricks and also a spark and mlflow
- 5:25:01contributor. So here I'm in the blog and
- 5:25:03let me quickly enable annotation.
- 5:25:07Let me zoom this a little bit.
- 5:25:10Okay. And here we go. So this was the
- 5:25:14diagram that I was referring to. So all
- 5:25:17of these boxes that you see over here,
- 5:25:19right? The boxes with the years put in
- 5:25:21over here, these are basically
- 5:25:24partitions, right? These are basically
- 5:25:27partitions. And these arrows that you
- 5:25:29see over here, these arrows, these are
- 5:25:33basically task. And in Spark
- 5:25:35terminology, a task takes up and
- 5:25:38processes one partition, right? A task
- 5:25:41basically takes up and processes one
- 5:25:43partition. And all of the partitions
- 5:25:46over here as you see they are uniformly
- 5:25:51uniformly or evenly side.
- 5:25:55They are uniformly or evenly sized.
- 5:25:57Right? So all of these tasks are going
- 5:26:00to complete almost on the same time.
- 5:26:04Right? Because the amount of data that
- 5:26:06they getting to process is just the
- 5:26:08same. Now this is something that
- 5:26:10actually doesn't happen in a real world
- 5:26:12scenario. Let me show you what would
- 5:26:14happen in a real world scenario. So a
- 5:26:17real world scenario would look something
- 5:26:19like this. Probably one of the partition
- 5:26:21would have too much of data and some of
- 5:26:23the partitions would have very less
- 5:26:25amount of data. Right? So we see that
- 5:26:272023 and 2022 have a lot more data than
- 5:26:30the others like 2020 and 2004. So these
- 5:26:34task these tasks are going to be on the
- 5:26:38critical path. So they are going to be
- 5:26:41on the critical path. And by that what I
- 5:26:43mean is that these two tasks need to
- 5:26:47complete in order for the whole job to
- 5:26:50to complete. Right? So these two tasks
- 5:26:52need to complete in order for the whole
- 5:26:54job to complete. Yeah.
- 5:26:58They are going to take the most amount
- 5:27:00of time.
- 5:27:04They're going to take the most amount of
- 5:27:06time. And the problem with this is that
- 5:27:09these two cores, the core processing,
- 5:27:11these two are going to be occupied while
- 5:27:14all the others are going to remain idle.
- 5:27:18They are going to remain idle. Yeah. So
- 5:27:21that is underutilization of your
- 5:27:23resources. Now what liquid clustering is
- 5:27:25going to do is that it is going to
- 5:27:28combine all of the small partitions
- 5:27:31together into something meaningful
- 5:27:36right into something meaningful. So this
- 5:27:382018 partition is added days it combined
- 5:27:412019 and 2004 into one 2020 and 2021
- 5:27:46into one. Right? But we still face this
- 5:27:48problem because 2023 and 2022 these are
- 5:27:51still big fat chunks that one core or
- 5:27:56one task would need to process and this
- 5:27:58is going to become a bottleneck because
- 5:28:00these guys have to do a lot of work.
- 5:28:02Yeah. So again what liquid clustering
- 5:28:04does is very smartly it breaks it down.
- 5:28:08It divides the larger chunks into
- 5:28:11smaller buckets. Right? So you see that
- 5:28:142023 had been divided into three part
- 5:28:18and 2022 has been divided into two
- 5:28:21parts. Now the core which takes this up
- 5:28:24the task which run over here all of them
- 5:28:27are very uniformly signed. This becomes
- 5:28:30similar to the ideal case that we had
- 5:28:33over here. Right? this ideal so-called
- 5:28:38ideal case that we had, right? And
- 5:28:41therefore my resource utilization
- 5:28:45is going to go up, right? And we are
- 5:28:49going to complete
- 5:28:53we are going to complete
- 5:28:56this job
- 5:28:58in a reasonable amount of time, right?
- 5:29:00because every task or every core has got
- 5:29:04even amount of data to process right so
- 5:29:06that's the beauty of liquid clustering
- 5:29:09and this example is on a single
- 5:29:11dimension let's see what would happen in
- 5:29:14multi-dimension so let's say we want to
- 5:29:16cluster we want to cluster by two
- 5:29:19columns
- 5:29:22and they are basically the year and
- 5:29:26customer
- 5:29:28right so in this case we are also going
- 5:29:31to analyze the sizes for each of the
- 5:29:34combinations and let's say that for year
- 5:29:36I have 2023 2022 and 2021 and these are
- 5:29:41my customers right so we see that these
- 5:29:45are the data sizes that have been marked
- 5:29:48over here right so red is small yellow
- 5:29:52is medium size green is optimal and blue
- 5:29:56is excel right so other than these to
- 5:29:59these two values
- 5:30:02for the year and customer enterprise
- 5:30:04customer in 2023.
- 5:30:06This is optimally sized and 2021
- 5:30:09enterprise customer is optimally sized.
- 5:30:12All of the others either they are small
- 5:30:14in size the files are either small in
- 5:30:17size or they are large in size like this
- 5:30:20one. Right? So what is going to happen
- 5:30:22is that liquid clustering is going to
- 5:30:25combine the file sizes and make them
- 5:30:28appropriately size. Right? So something
- 5:30:31like this is going to happen. It is
- 5:30:33going to combine medium two medium size
- 5:30:36and one small size file. Again two
- 5:30:39medium and one small size file.
- 5:30:43Three small size and two medium size and
- 5:30:46two medium and one small size file.
- 5:30:48Right? So it is going to combine all of
- 5:30:51this and create one optimal size file.
- 5:30:57Optimal size file.
- 5:31:00Right now once this is done we we've
- 5:31:02seen that all of the files are now
- 5:31:05optimally sized except for this one. Now
- 5:31:08what it is going to do is that it is
- 5:31:10going to break this down into smaller
- 5:31:13file sizes.
- 5:31:15Yeah. So it is going to break this down
- 5:31:18into smaller file sizes. And
- 5:31:22the final result that we are going to
- 5:31:24get is something like this. So now we
- 5:31:26see that all of the sizes all of the
- 5:31:29files are appropriately sized and we
- 5:31:32have appropriate number of files and
- 5:31:34that's the beauty of liquid clustering.
- 5:31:36So it helps you achieve two things. The
- 5:31:38first one is the small file problem.
- 5:31:45So it helps you avoid the small file
- 5:31:48problem. We've seen several files which
- 5:31:50were marked red, right? They were small
- 5:31:53files. So what it did is that it merged
- 5:31:55all of them and then it created an
- 5:31:58appropriately sized file. That's the
- 5:32:00first one. The second problem that it
- 5:32:02helps avoid is data skew.
- 5:32:06We've seen that we had
- 5:32:09partitions which had lots of data and
- 5:32:12that is where this example came in.
- 5:32:14Right? This was that example where one
- 5:32:16partition had a lot of data and we also
- 5:32:20had this example
- 5:32:23where the partition 2023 was skewed.
- 5:32:28Right? So what it did was that it
- 5:32:29divided this partition into appropriate
- 5:32:33number of parts and that is how it
- 5:32:35avoided the skew problem. Now it's very
- 5:32:38important to note is that when we talk
- 5:32:41about skew we think of it as there is
- 5:32:44one partition and it has tremendous
- 5:32:47amount of data to process and for that
- 5:32:49reason it takes forever to process it
- 5:32:51right. So this is avoided by dividing it
- 5:32:54into respective parts. But let's say you
- 5:32:57have a data set it has 1 million keys
- 5:33:00for a particular value right. So let's
- 5:33:02say um there is a key and for that key
- 5:33:05there are 1 million values
- 5:33:09and when you do a join the same keys go
- 5:33:11to the same partition in spark right so
- 5:33:14it is going to create a skew again now
- 5:33:17I'm mentioning this in order to point
- 5:33:19out that this doesn't liquid clustering
- 5:33:22doesn't change shuffle behavior
- 5:33:26right
- 5:33:27what it helps avoid is a skew at this
- 5:33:31point Right? It basically sizes the file
- 5:33:34appropriately so that all of the tasks
- 5:33:37get appropriately sized partitions.
- 5:33:39Right? The sizes of the partitions that
- 5:33:41they get are not skewed. Okay. So now
- 5:33:44let's see liquid clustering in action.
- 5:33:46What I'm going to do is that I am going
- 5:33:48to create two tables. The first one I'm
- 5:33:50going to partition and zorder by and the
- 5:33:53second one I'm going to do a liquid
- 5:33:56clustering. And then let's run a query
- 5:33:58on both the tables and compare the
- 5:34:00runtime. Yeah. So, let me quickly copy
- 5:34:04all of these paths. So, this is going to
- 5:34:07be spark dot
- 5:34:09read.park.
- 5:34:10[Music]
- 5:34:30Okay, this has been read in. Now let me
- 5:34:33union all of these files, all of these
- 5:34:36data frames. DF2 dot union df3.
- 5:34:44Let's also select some relevant columns
- 5:34:47from here because we don't want to there
- 5:34:50are too many columns in here. uh
- 5:34:53customer ID, category,
- 5:34:57price, quantity,
- 5:34:59and the invoice date. Yeah. So now that
- 5:35:03this is done, let me go ahead and write
- 5:35:06this file. Write dot mode override
- 5:35:12dot partition by
- 5:35:14let's partition by invoice date.
- 5:35:17Something similar that we've done
- 5:35:19earlier. And let's save this file as
- 5:35:23delta catalog delta DB dot liquid
- 5:35:28clustering example one.
- 5:35:31And
- 5:35:33let's now run a Z order on this file.
- 5:35:38Delta catalog delta DB dot
- 5:35:43example one Z order by customer ID.
- 5:35:48Yeah, because you remember we've done a
- 5:35:50similar uh we've run a query where we
- 5:35:53filtered by customer ID all here, right?
- 5:35:55So for that reason I want to zorder by
- 5:35:58customer ID and then I'll run the same
- 5:36:01kind of query. Yeah. So just to quickly
- 5:36:04show you once again
- 5:36:08I want to
- 5:36:11mimic this kind of a query which we've
- 5:36:13run earlier.
- 5:36:16Yeah, something like this. Right, let's
- 5:36:19go ahead and
- 5:36:26run this.
- 5:36:28Okay, so now this is the first table
- 5:36:32and let's also do a count to make sure
- 5:36:36that
- 5:36:37we have data that has landed in over
- 5:36:39here.
- 5:36:44Yeah, let's run this and parallelly I'm
- 5:36:47also going to create another table which
- 5:36:50is going to be example two and I am
- 5:36:53going to do a cluster by and you
- 5:36:56remember that in cases where we have a
- 5:36:58partitioned and a zordered column we
- 5:37:02basically take in both the partition and
- 5:37:04the zord ordered column as the cluster
- 5:37:06by columns right so the partition column
- 5:37:08is invoice date the zorder column is
- 5:37:11customer ID So I'll basically put in
- 5:37:13both of these values over here. Yeah. So
- 5:37:16now let's go ahead and also run this.
- 5:37:19Okay. So there have been some
- 5:37:20significant optimization. It added 174
- 5:37:24files and 364 files have been removed.
- 5:37:27So now let's have a look at the count.
- 5:37:29The count is 99457
- 5:37:33and the count over here is just the
- 5:37:35same. So now let's go ahead and run this
- 5:37:37query which is select category
- 5:37:42comma sum of
- 5:37:46sum of price into
- 5:37:49quantity
- 5:37:51as total sales
- 5:37:55right as total sales
- 5:37:58from
- 5:37:59this table
- 5:38:01where customer ID equals 201
- 5:38:05and we group by category, right?
- 5:38:10And I also want to time this query. So
- 5:38:13I'm going to do this inside a
- 5:38:18uh spark.sql.
- 5:38:27And let's put this over here.
- 5:38:30And let's go ahead and run this.
- 5:38:35Yeah, actually let me also add in
- 5:38:38something else. Let me add in
- 5:38:42and invoice date between
- 5:38:472021
- 5:38:5001 and
- 5:38:542 0 2 3 12 31. The reason why I also
- 5:38:59added in this over here is because I
- 5:39:01want to test out both the columns,
- 5:39:03right? Because we are partitioning
- 5:39:06by invoice date and then we are
- 5:39:08reordering by the customer ID. Yeah.
- 5:39:15Okay. So this is complete and this took
- 5:39:18somewhere around 254 milliseconds. Now
- 5:39:20let's also run this on the other table
- 5:39:25and we see that this is 150 milliseconds
- 5:39:28right not a very significant improvement
- 5:39:31but still a good improvement right uh
- 5:39:34but we know that the data on which we
- 5:39:37are running is limited right so probably
- 5:39:40we'll be able to see these improvements
- 5:39:42on a larger scale when we run these
- 5:39:45operation on huge amounts of data right
- 5:39:47but overall I believe that this gives
- 5:39:49you a is of how liquid clustering works
- 5:39:53and how you can still compare when you
- 5:39:55use a partition by and a reorder and
- 5:39:57equivalently convert that to a cluster
- 5:40:00by and there are definitely performant
- 5:40:03benefits that you get out of it right
- 5:40:05okay so now that we've seen liquid
- 5:40:06clustering in action it's really
- 5:40:08important to understand that it cannot
- 5:40:10be used together with zorder or hive
- 5:40:13tile partitioning right so if you're
- 5:40:15using liquid clustering that will
- 5:40:18independently be the mechanism for
- 5:40:21deciding the layout of your data. Yeah,
- 5:40:23so another question that you might have
- 5:40:25in mind is how do I choose the liquid
- 5:40:27clustering columns, right? The best
- 5:40:30practice is to choose the most
- 5:40:33frequently used column in your query
- 5:40:35filters, right? And if two columns are
- 5:40:39highly correlated, you should just
- 5:40:41include one of them in your query
- 5:40:43filters. Yeah. And what I mean by that
- 5:40:45is actually let me first quickly
- 5:40:48annotate. Let's say you have a column
- 5:40:51called product category.
- 5:40:55You have a column called product
- 5:40:56category and you have another column
- 5:40:58called location.
- 5:41:00And these two columns are for whatever
- 5:41:03reason they are highly
- 5:41:06correlated.
- 5:41:08So what it means is that whenever you
- 5:41:09choose a particular product,
- 5:41:12it is always going to yield some the
- 5:41:14same location. Right? If you choose
- 5:41:17another product, it is going to yield
- 5:41:18the same location always. Right? So that
- 5:41:21means that there is no point including
- 5:41:23this column location. So in those cases,
- 5:41:27you remove the columns which are highly
- 5:41:29correlated. You remove location and you
- 5:41:31only go ahead with product category.
- 5:41:33Right? So that's the first point. The
- 5:41:36second point is
- 5:41:39Yeah. The second point is if you're
- 5:41:40converting a table an existing table and
- 5:41:44it already has some kind of strategy,
- 5:41:47some kind of partitioning or zorder zord
- 5:41:49strategy already being followed and now
- 5:41:51you want to use liquid clustering. These
- 5:41:53are the rules that you should follow. So
- 5:41:55the first one is is if it's already hive
- 5:41:58partition basically just go ahead and
- 5:42:01use that partition column as the
- 5:42:03clustering key. The second one is if
- 5:42:06you're using a column for the order
- 5:42:08indexing,
- 5:42:10use the same Z order column for your
- 5:42:13clustering. Right? The third one is high
- 5:42:16style partitioning and the order. We've
- 5:42:18seen this example, right? We've used
- 5:42:20both we first partitioned by high style
- 5:42:22partitioning and inside of it we've
- 5:42:25applied the order. So there are two
- 5:42:26different columns and in the previous
- 5:42:28example we hive partitioned by invoice
- 5:42:32date and then we reordered by category
- 5:42:36right so in that case use both the
- 5:42:39partition column and the zorder by
- 5:42:41column as your clustering key. So your
- 5:42:43new clustering key basically becomes
- 5:42:45invoice date and category. So you simply
- 5:42:48cluster by
- 5:42:50these two columns. And the last one is
- 5:42:53if you're having if you're using
- 5:42:54generated column to reduce the cardality
- 5:42:57for example date or time. So let's say
- 5:42:59you have a time stamp column and then
- 5:43:01you convert it to a date. Many times we
- 5:43:03do that in order to reduce the cardality
- 5:43:05right. So if we are already doing this
- 5:43:08and if we partition by date in that case
- 5:43:11use the original column just use the
- 5:43:14time stamp as your clustering key and
- 5:43:16don't create a generated column. So
- 5:43:18there's no need to create a date column,
- 5:43:20right? You just partition, sorry, you
- 5:43:22just cluster by the time stamp. Yeah. So
- 5:43:26to quickly summarize, liquid clustering
- 5:43:28is going to be super helpful when your
- 5:43:31query patterns are going to change down
- 5:43:33the line, right? You feel that your
- 5:43:35query patterns are going to change down
- 5:43:36the line and you don't want to be bound
- 5:43:39by fixing the partitioning or the zorder
- 5:43:42column, right? So so liquid clustering
- 5:43:44helps you remain flexible. you are you
- 5:43:47have that flexibility of being able to
- 5:43:49change the cluster by columns anytime
- 5:43:51down the line. Right? So that's the
- 5:43:53first benefit. The second benefit is it
- 5:43:56avoids the small file problem. As we've
- 5:43:59seen that it combines lot of small files
- 5:44:02into appropriately sized files. It even
- 5:44:05breaks down bigger files largely sized
- 5:44:08files into appropriately sized one.
- 5:44:10Yeah. And the third one is it avoids
- 5:44:13data skew. It basically breaks up the
- 5:44:15bigger partitions into smaller ones and
- 5:44:18that is how each task or each score gets
- 5:44:22a reasonable amount of data to process.
- 5:44:25I'm super happy to see that you've
- 5:44:27reached the end of the video and I
- 5:44:29really hope that you learned a lot from
- 5:44:32it and you enjoyed it. So, please don't
- 5:44:34forget to like and share this video. Tag
- 5:44:37me on LinkedIn. Share whatever you've
- 5:44:39learned. I'll be more than happy to see
- 5:44:41that. Please don't forget to subscribe
- 5:44:44to my channel because I've seen that
- 5:44:46only 30% of you have subscribed to my
- 5:44:48channel. So, please go ahead and hit
- 5:44:50that subscribe button. It really
- 5:44:52motivates me to make a lot more content.
- 5:44:54So, thank you so much for watching.
- 5:44:58[Music]
About this transcript
This page contains the full transcript of Delta Lake Masterclass | Azure Databricks | PySpark | From Zero-To-Expert by Afaque Ahmad, generated from the public captions YouTube serves with the video. The transcript has 48,542 words across 7,157 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.