What is Spark? (Visual Explanation) — Transcript
Full transcript
- 0:00Hey friends, so now I'm going to show
- 0:01you exactly what is Apache Spark. So I'm
- 0:04going to show you its architecture, the
- 0:06main components, and how Spark works
- 0:09behind the scenes. And at the end, I'm
- 0:11going to show you the entire ecosystem
- 0:13of the Spark and my recommendation on
- 0:15how to learn it. And of course, this is
- 0:17a very technical topic. That's why, as
- 0:20usual, I will break it into very simple
- 0:22animated sketches. So now, let's get
- 0:24started. Show started.
- 0:26The show's started. The main idea is
- 0:28very simple. Instead of using only one
- 0:31single machine to process big data, to
- 0:33do machine learning, analytics,
- 0:35engineering, we're going to use multiple
- 0:37machines working together in the Spark
- 0:40environment. So it going to split our
- 0:42data between multiple machines and
- 0:44process our data in the memory in
- 0:47parallel. And then the result of each
- 0:48piece can be collected and combined into
- 0:51a final result. So actually, that's it.
- 0:53This is Spark in very high level. Now
- 0:56we're going to go to the lower to
- 0:57understand the detailed architecture of
- 1:00Spark. Now we start by writing Spark
- 1:02code and mainly we have to write two
- 1:04blocks. The first one is we need to
- 1:06create Spark session. This is an initial
- 1:09step and without it we cannot do
- 1:11anything. So once Spark sees this, it's
- 1:13going to say, "Okay, I'm going to turn
- 1:15the engine on and I'm ready." So the
- 1:17first thing that it going to do, it
- 1:19going to create something called driver.
- 1:21This is the brain of Spark. It will do
- 1:24all the thinking for your program, but
- 1:26it will not do the heavy job itself. So
- 1:29only thinking. And that's why we have
- 1:30the other type of node that going to do
- 1:32the heavy work. We have the workers, the
- 1:35machines. Those are the muscles. Those
- 1:38are real machines with CPU, memory,
- 1:40disk. And we can run on those machines
- 1:43the executors. And the executor is a
- 1:45process that going to use the machine
- 1:47resources to execute a task. So we have
- 1:50the brain, we have the muscles. Now the
- 1:52brain will not go and start like finding
- 1:54which worker is busy and which worker is
- 1:57free. That's why we have something
- 1:58called the cluster manager. It's like
- 2:01real life they will not do the actual
- 2:03work. They will manage other workers. So
- 2:06they going to find which worker is
- 2:07available. They going to start the
- 2:09executor processes, allocate the CPU and
- 2:12memory, and monitor the entire health.
- 2:15So the brain going to ask the manager to
- 2:17allocate the resources and it knows best
- 2:19which worker going to be involved in our
- 2:21task. So with this we have the brain,
- 2:23the muscles, the managers. But actually,
- 2:25until now nothing happens. We are just
- 2:28turning Spark on. And here we comes to
- 2:30the second block in our Spark code, we
- 2:32have the data logic. Like for example,
- 2:35you are reading table, you are filtering
- 2:37the data, doing some group by. And at
- 2:39the end we have an action in order to
- 2:41show the data. So now, there is
- 2:43something really important to understand
- 2:44about Spark. It is very lazy. So it will
- 2:47not go and panic from the first command
- 2:49and start like getting the data and
- 2:51processing it, doing aggregations,
- 2:53filters. Instead, Spark is lazy, calm,
- 2:56and start listening. So it's going to
- 2:58say, "Okay, I see that you want to read
- 3:00the table. Mhm. Now you want to filter.
- 3:03Okay, and you want to group by by
- 3:04country. Interesting stuff, right? So
- 3:07let me think about it." So it will not
- 3:08execute it immediately. The driver, the
- 3:10brain, it can start building a full plan
- 3:13step by step. It keep reading your code
- 3:15until it sees an action, the show
- 3:18methods. And now Spark suddenly going to
- 3:20say, "Oh, you actually want me to do all
- 3:22those stuff? Okay, it's time to work."
- 3:24Now the driver going to go and optimize
- 3:26the plan. It can breaks everything into
- 3:28stages and steps. And here comes the
- 3:30main idea. Your data going to be
- 3:32splitted into partitions and each
- 3:34partition becomes one task. Now the
- 3:37driver going to talk directly to the
- 3:39executors, to the workers, not through
- 3:41the cluster manager anymore. The manager
- 3:43only worked at the start and that's it.
- 3:45So the driver now coordinates the
- 3:47execution and it's going to say, "You
- 3:49work with the partition one, you work
- 3:51with the partition two, you work three,
- 3:52[music] and you handle that." And
- 3:54everyone start processing now. So each
- 3:56executor going to reads its partitions,
- 3:58it's going to go and apply the filters,
- 4:00perform aggregations. And what could
- 4:02happen is shuffle. So the driver might
- 4:05start moving actually data between the
- 4:07workers if, for example, they belong to
- 4:09the same country, then it makes sense to
- 4:11process everything in one node. Of
- 4:12course, the executors, they don't do the
- 4:15thinking. It is the job of the driver.
- 4:17And now once all the nodes are done with
- 4:19the processing, now it is time to
- 4:21combine everything in one final answer.
- 4:23So the final answer going to be sent
- 4:25back to the driver and it's going to
- 4:27print it in the output because of the
- 4:29method show. So actually, my friends,
- 4:30this is how Spark works behind the
- 4:32scenes. Again, to summarize. So the
- 4:34driver thinks and plans, the cluster
- 4:37manager provides the workers, the
- 4:39machines, the executors do the work, the
- 4:43partitions are actually the units of the
- 4:45parallel work, and the flow is very
- 4:47simple. You write a code with two
- 4:49blocks, create a session and the logic
- 4:51of your data. Spark is lazy, so it's
- 4:53going to listen and build the plan. And
- 4:56once it sees an action, it's going to
- 4:57trigger the execution. So the driver
- 4:59going to coordinate everything directly
- 5:01and speak to the executors. And the
- 5:03executors going to process everything in
- 5:05parallel and return the result back to
- 5:08the driver. That's it. Now by looking to
- 5:10the Spark architecture, I can say the
- 5:12whole thing, the driver, workers,
- 5:14executors, the manager, we call the
- 5:17whole thing as a cluster. So a cluster
- 5:19is simply a group of machines that work
- 5:21together as one system. And now my
- 5:23friends, we come to something really
- 5:24important to understand if you want to
- 5:26use Spark in projects. You have two
- 5:29ways. Either you're going to go and use
- 5:30Databricks in order to interact with
- 5:33Spark and do projects. And here the big
- 5:35advantage, you don't have manually to
- 5:37configure all this complexity. So you
- 5:39don't have to start drivers manually,
- 5:41you don't have to launch the executors
- 5:43yourself, don't talk directly to the
- 5:45cluster manager. The only one thing that
- 5:48you have to decide and to configure is
- 5:50actually how many workers do I need for
- 5:52my cluster. Everything else going to be
- 5:54handled internally in Databricks. So if
- 5:57you use Databricks, you don't have to
- 5:59worry about all those details. You just
- 6:01focus on writing your code for your
- 6:03project and that's it. But of course, it
- 6:05is nice to understand those concepts as
- 6:07you are using the platform. Now the
- 6:09other option that you have is to not use
- 6:11Databricks and you're going to run Spark
- 6:13in more traditional environments. And
- 6:15here you have to master all those
- 6:17details because, my friend, you have to
- 6:19configure everything on your own. The
- 6:21number of executors, the memory, the
- 6:23number of cores, the cluster manager
- 6:25type, the allocations, and all those
- 6:27details. So the level of details you
- 6:29need depends really on where you run
- 6:31Spark. But of course, the most important
- 6:33part is to understand how Spark works
- 6:36behind the scenes and then you
- 6:37understand how to use Spark everywhere.
- 6:39All right, friends, so with this now you
- 6:41have a clear picture about how things
- 6:43works behind the scenes in Spark. And
- 6:45now what we're going to do, we're going
- 6:46to zoom out a little bit because Spark
- 6:48is not only a distributed engine, it is
- 6:51a full ecosystem. So Spark looks like
- 6:54this. At the center we have something
- 6:56called Spark Core. It is the main part
- 6:58where it's going to handles the
- 7:00distributed execution, the memory
- 7:02management, task scheduling, and fault
- 7:05tolerance. So basically, all those heavy
- 7:07infrastructure work that happens behind
- 7:10the scenes. And everything else going to
- 7:12be built on top of the Spark Core. So
- 7:15the other things are actually like we
- 7:17have different libraries. Like for
- 7:19example, the very common and famous
- 7:21library we have the Spark SQL. It
- 7:24provides us very simple tools that looks
- 7:26like SQL to query and manipulate your
- 7:29data. Another library we have the Spark
- 7:31Streaming. This going to be specially
- 7:33important if you want to process
- 7:35real-time data. Like for example,
- 7:37streaming from Kafka platform. Another
- 7:40library we have the Spark MLlib. If you
- 7:43are data scientist and you want to build
- 7:45machine learning models. And another one
- 7:48we have the GraphX. If you want to work
- 7:50with graph data like networking or
- 7:52relationships. And the last library we
- 7:55have the SparkR. If you are an R user
- 7:57and you want to work with Spark. So now
- 7:59by looking to this, the Spark Core is
- 8:01the engine. And then you have those
- 8:03different libraries. They are like tools
- 8:06that allow us to do different type of
- 8:08work using Spark. Now another thing in
- 8:10the ecosystem, you can actually use
- 8:12different languages to interact with
- 8:15Spark. So we can use Python, R, SQL,
- 8:19Scala, Java. So my friends, you can pick
- 8:22the language that suits you to work with
- 8:24Spark because at the end it doesn't
- 8:26matter which one you pick. Under the
- 8:28hood, everything eventually go through
- 8:30the Spark Core. So now by looking to
- 8:32this, you can understand why Spark
- 8:33became so powerful because this is not
- 8:36only like an engine or one tool, it is
- 8:39an entire platform. And now I totally
- 8:41understand if I might get you scared
- 8:43because if you look to the ecosystem and
- 8:45you want to learn Spark, you might say,
- 8:47"I'm going to go and learn all those
- 8:49stuff." Well, my friends, don't worry
- 8:51because you don't have to learn
- 8:53everything. I'm going to give you now my
- 8:54recommendations on how to learn it and
- 8:57my honest opinion. Now if you are a data
- 8:59analyst, then you just learn the theory
- 9:01about the Spark Core, just some basic
- 9:04understanding. And now about the
- 9:05libraries, you just need to learn the
- 9:08Spark SQL and actually that's it because
- 9:11it's going to gives you enough tools
- 9:12that you need as a data analyst to work
- 9:14with your data. But now if you are a
- 9:16data engineer, then it really depends on
- 9:18the environments. If you're going to use
- 9:20Spark in platforms like Databricks, then
- 9:22I'm going to say the same things. You
- 9:24just need to learn the basic theory
- 9:26about the Spark Core. But if you want to
- 9:28build and manage Spark in the
- 9:30traditional way without Databricks, then
- 9:33you need more than that. You need deeper
- 9:35knowledge on the Spark Core because
- 9:37you're going to have to configure
- 9:39everything on your own. Now about the
- 9:40libraries, again you have to master the
- 9:42Spark SQL because you have to manipulate
- 9:45and transform the data. As well, you
- 9:47have to learn the Spark Streaming
- 9:50because in many modern projects you will
- 9:51be streaming data from platforms like
- 9:54Kafka. And this is your library in order
- 9:56to do that. So, actually that's it. Now,
- 9:59if you are a data scientist, then it is
- 10:01similar to the analyst. You need to have
- 10:03as well some basic understanding about
- 10:04the Spark Core and about the libraries,
- 10:07you have to master as well the Spark
- 10:08SQL. And now, about the Spark MLlib here
- 10:11is my honest opinion, you don't really
- 10:13need to learn it because currently in
- 10:15many modern machine learning projects,
- 10:17we prefer to use external libraries like
- 10:20the scikit-learn, TensorFlow, and
- 10:22PyTorch. So, as a data scientist, you're
- 10:25going to end up using Spark to just do
- 10:27data preparations and maybe feature
- 10:29engineering. But, for model training,
- 10:31you will be using something external. So
- 10:33now, by looking to the whole big picture
- 10:35and to those libraries, you can
- 10:36understand quickly that actually the
- 10:38most important component is Spark SQL
- 10:42because it going to gives you many
- 10:43amazing tools that looks like SQL. And
- 10:46my friend, SQL is very easy to use.
- 10:48That's going to help you to manipulate
- 10:50and work with your data. And about the
- 10:52programming language, this library going
- 10:54to allows you to write only Python
- 10:56PySpark code or as well to just write
- 10:59SQL queries like you do in any
- 11:01databases. Then, it doesn't matter which
- 11:03language you use because everything
- 11:04going to go through the Spark Core to
- 11:07the Spark engine. So, it's going to be
- 11:08super fast. And this is exactly the
- 11:10reason why we going to deep dive into
- 11:12this library, the Spark SQL. So, that's
- 11:14it, my friends. This is what Spark, the
- 11:17architecture, how it works behind the
- 11:18scenes. That is exactly how it should
- 11:20work. The whole ecosystem and how to
- 11:22learn it. And of course, if you want to
- 11:24learn Spark with me, I will deep dive
- 11:26now into the Spark SQL in order to learn
- 11:28all the commands and how to use it in
- 11:30real projects. And now, my friends, if
- 11:32you enjoy this type of free content
- 11:34where I'm sketching the complex concepts
- 11:36using those [music] animated visuals,
- 11:38then support the channel by subscribing,
- 11:40liking, and commenting. This going to
- 11:42help us to grow and as well reach nice
- 11:44people like you. So, if you're still
- 11:45here, thank you so much for watching and
- 11:47I will see you in the next video.
- 11:48Bye-bye. Okay, bye. Bye.
- 11:55>> [music]
About this transcript
This page contains the full transcript of What is Spark? (Visual Explanation) by Data with Baraa, generated from the public captions YouTube serves with the video. The transcript has 2,313 words across 332 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.