Delta Lake Architecture Explained | How the Lakehouse Revolutionizes Data Engineering" — Transcript
Full transcript
- 0:01[Music]
- 0:08Well, we all know what is a data lake
- 0:10and actually we have used data lake in
- 0:12this particular course couple of time
- 0:14which is actually our Azure data lake
- 0:16storage gen 2. So we know that we can
- 0:19store a huge amount of data inside a
- 0:21data lake. Um but we don't know what is
- 0:23a delta lake and now that's the time to
- 0:26understand uh a delta lake actually and
- 0:28we also need to understand delta lake
- 0:30architecture in this. So let's focus on
- 0:32this first. Uh so what is a data lake?
- 0:35The data lake stores a large volume of
- 0:37structured semistructured or
- 0:39unstructured data in short all kind of
- 0:41data in its native format. Uh mostly
- 0:43data lake is like welcoming all kind of
- 0:45data from no matter which kind of source
- 0:47it is. Data lake architecture has
- 0:50evolved in recent years to better meet
- 0:52the demands of increasingly datadriven
- 0:55enterprises as the data volume continues
- 0:57to rise. We are living in a world where
- 0:59data is increasing every day and then
- 1:01when we have huge amount of data the
- 1:04data lake which was invented before few
- 1:06years has actually some evolution inside
- 1:09that and then this evolution has changed
- 1:11the data lake structure also with that
- 1:14and from that only we got something
- 1:16which is known as delta lake before I
- 1:19explain what is a delta lake let me tell
- 1:21you one thing guys delta lake is not one
- 1:23of the new service of Azure cloud it's
- 1:25not it's not a service it's actually a
- 1:27conceptual architecture ure and you have
- 1:29to implement this conceptual
- 1:31architecture with your data lake only.
- 1:33So when you're going to use your data
- 1:34lake gen 2, you are going to associate
- 1:37that in such a way that it is going to
- 1:38be treated like a delta lake. So what is
- 1:41this? Well, the delta lake is an
- 1:43open-source storage layer that brings
- 1:46reliability and performance to your
- 1:47existing data links. So why and how this
- 1:51is going to brings reliability and
- 1:53performance actually and why it is
- 1:54required that is something which we're
- 1:55going to focus right now. They are
- 1:57saying delta lake provides asset
- 1:59transactions, scalable metadata handling
- 2:02and unifi streaming and batch data
- 2:05processing. Now all the things are
- 2:08obviously missing in the normal data
- 2:09lake. Most of the time data is just
- 2:12going to store the huge amount of data
- 2:14inside the data lake kind of services
- 2:16but it's not associating with asset
- 2:18transaction. I hope you heard about
- 2:20asset transactions, atomicity,
- 2:22consistency, whatever. And uh when you
- 2:25want scalable metadata handling or you
- 2:27want some unified or structured
- 2:29streaming kind of things, all these
- 2:30things are where actually uh missing
- 2:32from the data lake architecture where
- 2:35delta lake is actually focusing on this
- 2:37kind of things. The delta lake runs on
- 2:39top of your existing data lake. So it's
- 2:42not a new service as I said on your
- 2:44existing data lake only. this is going
- 2:45to be running and it is fully compatible
- 2:48with Apache Spark APIs. That's one of
- 2:50the reason when you're going to use your
- 2:52Synapse notebooks or you're going to use
- 2:54your datab bricks notebooks, your Apache
- 2:56Spark APIs will directly associate with
- 2:58your data lake and then you can
- 3:00implement Delta Lake with that. In order
- 3:03to understand this concept properly,
- 3:04let's focus on this next slide. This is
- 3:07actually something which is a Delta Lake
- 3:09architecture. Now I have a two
- 3:11variations of this architecture. This is
- 3:13a simpler one. If you focus on this
- 3:15architecture left side, you have a data
- 3:18which is maybe a batch data or streaming
- 3:20data which is going to be ingested into
- 3:22this delta lake architecture.
- 3:24The bottom layer of this delta lake
- 3:26architecture you can see is nothing but
- 3:27your same existing data lake. On top of
- 3:30this data lake, you're going to
- 3:31implement delta lake where we have three
- 3:33different variations of our data. We
- 3:36have something which is known as bronze
- 3:38table, silver table and gold table. Some
- 3:40people call this thing layers also
- 3:42bronze layer, silver layer and gold
- 3:44layer of your data. And these layers are
- 3:46actually making it special. These layers
- 3:48are actually going to have uh different
- 3:50kind of processed unprocessed data
- 3:52inside that which are going to make this
- 3:54thing special. Now most of the time when
- 3:56you have a data in bronze, silver and
- 3:59gold each layer is having some
- 4:01characteristics of the data. The final
- 4:04layer of this particular data is gold
- 4:05layer or gold tables. This is actually
- 4:07the one which is going to be used with
- 4:09your Azure machine learning and some
- 4:11other further processing. But then the
- 4:13process actually going to start from the
- 4:15bronze. Most of the time your bronze
- 4:18layer or bronze tables are going to have
- 4:20your streaming data batch data which is
- 4:22coming from various sources. Once the
- 4:24data is stored inside the bronze layer,
- 4:26we will pick the data and we are going
- 4:28to refine those data with some kind of
- 4:31additional joins or some other
- 4:32meaningful associations with that. This
- 4:35meaningful refined tables are going to
- 4:36be treated as a silver layer or silver
- 4:38tables with that. But after that also
- 4:41silver tables are not ready to be
- 4:42processed or not something which is
- 4:44exactly looking for the business
- 4:46requirement and that's why on the silver
- 4:48tables we are going to apply some kind
- 4:49of aggregates. These aggregates are
- 4:52going to be applied and then final data
- 4:53layer is going to be created which is
- 4:55your gold layer which is going to be
- 4:56further useful in the processing of
- 4:58PowerBI or maybe machine learning or
- 5:01maybe some artificial intelligence kind
- 5:03of things. These three layers are
- 5:05actually making this Delta Lake
- 5:06architecture special and that's why
- 5:08let's deep dive into these words which
- 5:10are bronze, silver and gold.
- 5:14Hey guys, sorry for interruption. My
- 5:15name is Marauti and I'm here to make an
- 5:18very important announcement. I hope you
- 5:20are liking our videos and you're doing a
- 5:22continuous learning with us on an Azure
- 5:24cloud and AI related topics. If you are
- 5:27enjoying this thing, I'm going to
- 5:28announce skilltech.club club which is
- 5:30our upcoming website which is going to
- 5:32be launched very soon. We are here to
- 5:34tell you one thing that everyone who is
- 5:37a subscriber of this particular channel
- 5:39will get Azure cloud and Azure AI
- 5:41related certification courses free of
- 5:43cost in skilltech.club. So you will be a
- 5:47part of the skilltech.club kind of a
- 5:48membership automatically free of cost
- 5:51and everyone who's a subscriber of this
- 5:53particular channel will get those
- 5:55benefits which are available in that. So
- 5:57what are you waiting for? I request you
- 5:59to please subscribe to this channel and
- 6:01share it with your friends and families
- 6:02if they are also interested in Azure
- 6:05cloud and AI learning. That's it from my
- 6:07side. Now you can carry on with your
- 6:09learning. Thank you. If I focus on this,
- 6:13the answer says that the bronze table
- 6:15contains raw data ingested from its
- 6:18various sources. It can be JSON, RDBMS,
- 6:21IoT data or maybe some data coming from
- 6:23the event hub kind of live streaming.
- 6:25This bronze data is going to be further
- 6:27processed into a silver. So silver
- 6:29tables will provide more refined view of
- 6:31our data. Most of the time in this case
- 6:33you can join the fields from various
- 6:35bronze tables to enrich your streaming
- 6:37records or maybe you can update your
- 6:39account statuses based on the recent
- 6:41activities. It's something which is a
- 6:43further processing or data preparation
- 6:45kind of things which you are applying on
- 6:47top of your bronze layer. Once the
- 6:49silver layer is processed, you are going
- 6:51to have the final layer process from
- 6:53this which is gold tables. These gold
- 6:55tables are going to provide business
- 6:57level aggregates and often used for
- 6:59reporting and dashboarding or maybe for
- 7:01machine learning and model creations.
- 7:04This would include aggregations such as
- 7:06maybe it's going to be daily active
- 7:07website users you want to see or maybe
- 7:09you want to get a sales per store uh or
- 7:12maybe you want to associate with the
- 7:14gross revenue per quarter for any
- 7:16particular company and department. This
- 7:18kind of more meaningful data is going to
- 7:20be used in my gold layer. And then this
- 7:23transition of data from bronze to silver
- 7:25to gold is the heart of your delta lake
- 7:27architecture. The end outputs are
- 7:30actionable insights, dashboards and
- 7:32reports which maybe you're going to
- 7:34associate with any existing service.
- 7:36This delta lake architecture if you
- 7:38understand with this bronze, silver and
- 7:40gold, the detailed view of this
- 7:42particular architecture is going to be
- 7:43somehow going to look like this. Now if
- 7:46you see right now my delta lake
- 7:48architecture is actually focusing only
- 7:50on this box but I want you to see the
- 7:52full picture of the delta lake
- 7:53architecture and that's the reason I
- 7:55have associated this delta lake
- 7:57architecture with the other services
- 7:58which are associated with that you can
- 8:00see I'm taking a scenario of data bricks
- 8:02right now because exactly after this
- 8:04video I'm going to show you how you can
- 8:06implement delta lake architecture with
- 8:08data bricks actually in this case
- 8:10obviously the flow is going to start
- 8:12from the streaming data or maybe a batch
- 8:15data which is coming from various
- 8:16sources. All the sources of data which
- 8:18are coming into this will be ingested
- 8:20into this particular data bricks and
- 8:22then that is going to be stored inside
- 8:24my data lake as a delta lake and that is
- 8:28going to be my bronze layer of my data.
- 8:30Once blown layer of data is there we are
- 8:32going to take that data and we are going
- 8:34to do some data preparation ETL kind of
- 8:36things on that and from bronze we are
- 8:38going to get another layer of the data
- 8:40which is going to be silver. These are
- 8:42going to be much more refined tables as
- 8:43we discuss but this is not something
- 8:46which is having some business related
- 8:47aggregates on that. So we'll take the
- 8:49silver tables and then we are going to
- 8:51apply some extraction some aggregators
- 8:53on that and then we are going to have a
- 8:55final layer of the data which is going
- 8:57to be my gold layer. Technically guys
- 8:59this gold layer is actually nothing but
- 9:01your data m. This is the one which is
- 9:03focusing on some business features and
- 9:05capabilities with that. And that data is
- 9:07a final produced data of this process
- 9:10which is going to be further stored into
- 9:12a maybe a cloud data warehouse like
- 9:14Synapse or maybe it can be directly
- 9:16associated with the PowerBI or maybe
- 9:19Tableau or maybe some other related
- 9:21product where we can either generate
- 9:23reports or we can use this thing for
- 9:25machine learning model training or some
- 9:27other purpose which can be there. This
- 9:30full process, this full architecture is
- 9:32actually showing you that how Delta Lake
- 9:34is going to be an important part for
- 9:36your end to-end data processing which is
- 9:38happening with this kind of data bricks
- 9:40or Azure Synapse kind of services. This
- 9:43is going to be an complex architecture
- 9:45to understand but I'll be giving you one
- 9:48guarantee that if you understand this
- 9:49architecture and if you know how to
- 9:51implement this thing this way in most
- 9:53organizations in your projects when you
- 9:55are doing data analytics this is going
- 9:57to be the heart of that architecture.
- 10:01Now if you understood this architecture
- 10:03I want you to take a screenshot of this
- 10:05particular one and then just make sure
- 10:07that you are going to refer this thing
- 10:09in our next sample. Next two videos are
- 10:12going to show you how you can implement
- 10:13Delta Lake architecture with data bras
- 10:16actually. So let's have a look at that
- 10:18one. I hope you understood the
- 10:20architecture and it will be much more
- 10:22clear when you see this thing
- 10:23practically how it is going to be
- 10:25implemented in data bricks. Thank you.
- 10:29[Music]
About this transcript
This page contains the full transcript of Delta Lake Architecture Explained | How the Lakehouse Revolutionizes Data Engineering" by Skilltech Club, generated from the public captions YouTube serves with the video. The transcript has 2,024 words across 295 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.