YouTube2Text

What is Spark? (Visual Explanation) — Transcript

by Data with Baraa · 2,313 words · 332 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Hey friends, so now I'm going to show
  2. 0:01you exactly what is Apache Spark. So I'm
  3. 0:04going to show you its architecture, the
  4. 0:06main components, and how Spark works
  5. 0:09behind the scenes. And at the end, I'm
  6. 0:11going to show you the entire ecosystem
  7. 0:13of the Spark and my recommendation on
  8. 0:15how to learn it. And of course, this is
  9. 0:17a very technical topic. That's why, as
  10. 0:20usual, I will break it into very simple
  11. 0:22animated sketches. So now, let's get
  12. 0:24started. Show started.
  13. 0:26The show's started. The main idea is
  14. 0:28very simple. Instead of using only one
  15. 0:31single machine to process big data, to
  16. 0:33do machine learning, analytics,
  17. 0:35engineering, we're going to use multiple
  18. 0:37machines working together in the Spark
  19. 0:40environment. So it going to split our
  20. 0:42data between multiple machines and
  21. 0:44process our data in the memory in
  22. 0:47parallel. And then the result of each
  23. 0:48piece can be collected and combined into
  24. 0:51a final result. So actually, that's it.
  25. 0:53This is Spark in very high level. Now
  26. 0:56we're going to go to the lower to
  27. 0:57understand the detailed architecture of
  28. 1:00Spark. Now we start by writing Spark
  29. 1:02code and mainly we have to write two
  30. 1:04blocks. The first one is we need to
  31. 1:06create Spark session. This is an initial
  32. 1:09step and without it we cannot do
  33. 1:11anything. So once Spark sees this, it's
  34. 1:13going to say, "Okay, I'm going to turn
  35. 1:15the engine on and I'm ready." So the
  36. 1:17first thing that it going to do, it
  37. 1:19going to create something called driver.
  38. 1:21This is the brain of Spark. It will do
  39. 1:24all the thinking for your program, but
  40. 1:26it will not do the heavy job itself. So
  41. 1:29only thinking. And that's why we have
  42. 1:30the other type of node that going to do
  43. 1:32the heavy work. We have the workers, the
  44. 1:35machines. Those are the muscles. Those
  45. 1:38are real machines with CPU, memory,
  46. 1:40disk. And we can run on those machines
  47. 1:43the executors. And the executor is a
  48. 1:45process that going to use the machine
  49. 1:47resources to execute a task. So we have
  50. 1:50the brain, we have the muscles. Now the
  51. 1:52brain will not go and start like finding
  52. 1:54which worker is busy and which worker is
  53. 1:57free. That's why we have something
  54. 1:58called the cluster manager. It's like
  55. 2:01real life they will not do the actual
  56. 2:03work. They will manage other workers. So
  57. 2:06they going to find which worker is
  58. 2:07available. They going to start the
  59. 2:09executor processes, allocate the CPU and
  60. 2:12memory, and monitor the entire health.
  61. 2:15So the brain going to ask the manager to
  62. 2:17allocate the resources and it knows best
  63. 2:19which worker going to be involved in our
  64. 2:21task. So with this we have the brain,
  65. 2:23the muscles, the managers. But actually,
  66. 2:25until now nothing happens. We are just
  67. 2:28turning Spark on. And here we comes to
  68. 2:30the second block in our Spark code, we
  69. 2:32have the data logic. Like for example,
  70. 2:35you are reading table, you are filtering
  71. 2:37the data, doing some group by. And at
  72. 2:39the end we have an action in order to
  73. 2:41show the data. So now, there is
  74. 2:43something really important to understand
  75. 2:44about Spark. It is very lazy. So it will
  76. 2:47not go and panic from the first command
  77. 2:49and start like getting the data and
  78. 2:51processing it, doing aggregations,
  79. 2:53filters. Instead, Spark is lazy, calm,
  80. 2:56and start listening. So it's going to
  81. 2:58say, "Okay, I see that you want to read
  82. 3:00the table. Mhm. Now you want to filter.
  83. 3:03Okay, and you want to group by by
  84. 3:04country. Interesting stuff, right? So
  85. 3:07let me think about it." So it will not
  86. 3:08execute it immediately. The driver, the
  87. 3:10brain, it can start building a full plan
  88. 3:13step by step. It keep reading your code
  89. 3:15until it sees an action, the show
  90. 3:18methods. And now Spark suddenly going to
  91. 3:20say, "Oh, you actually want me to do all
  92. 3:22those stuff? Okay, it's time to work."
  93. 3:24Now the driver going to go and optimize
  94. 3:26the plan. It can breaks everything into
  95. 3:28stages and steps. And here comes the
  96. 3:30main idea. Your data going to be
  97. 3:32splitted into partitions and each
  98. 3:34partition becomes one task. Now the
  99. 3:37driver going to talk directly to the
  100. 3:39executors, to the workers, not through
  101. 3:41the cluster manager anymore. The manager
  102. 3:43only worked at the start and that's it.
  103. 3:45So the driver now coordinates the
  104. 3:47execution and it's going to say, "You
  105. 3:49work with the partition one, you work
  106. 3:51with the partition two, you work three,
  107. 3:52[music] and you handle that." And
  108. 3:54everyone start processing now. So each
  109. 3:56executor going to reads its partitions,
  110. 3:58it's going to go and apply the filters,
  111. 4:00perform aggregations. And what could
  112. 4:02happen is shuffle. So the driver might
  113. 4:05start moving actually data between the
  114. 4:07workers if, for example, they belong to
  115. 4:09the same country, then it makes sense to
  116. 4:11process everything in one node. Of
  117. 4:12course, the executors, they don't do the
  118. 4:15thinking. It is the job of the driver.
  119. 4:17And now once all the nodes are done with
  120. 4:19the processing, now it is time to
  121. 4:21combine everything in one final answer.
  122. 4:23So the final answer going to be sent
  123. 4:25back to the driver and it's going to
  124. 4:27print it in the output because of the
  125. 4:29method show. So actually, my friends,
  126. 4:30this is how Spark works behind the
  127. 4:32scenes. Again, to summarize. So the
  128. 4:34driver thinks and plans, the cluster
  129. 4:37manager provides the workers, the
  130. 4:39machines, the executors do the work, the
  131. 4:43partitions are actually the units of the
  132. 4:45parallel work, and the flow is very
  133. 4:47simple. You write a code with two
  134. 4:49blocks, create a session and the logic
  135. 4:51of your data. Spark is lazy, so it's
  136. 4:53going to listen and build the plan. And
  137. 4:56once it sees an action, it's going to
  138. 4:57trigger the execution. So the driver
  139. 4:59going to coordinate everything directly
  140. 5:01and speak to the executors. And the
  141. 5:03executors going to process everything in
  142. 5:05parallel and return the result back to
  143. 5:08the driver. That's it. Now by looking to
  144. 5:10the Spark architecture, I can say the
  145. 5:12whole thing, the driver, workers,
  146. 5:14executors, the manager, we call the
  147. 5:17whole thing as a cluster. So a cluster
  148. 5:19is simply a group of machines that work
  149. 5:21together as one system. And now my
  150. 5:23friends, we come to something really
  151. 5:24important to understand if you want to
  152. 5:26use Spark in projects. You have two
  153. 5:29ways. Either you're going to go and use
  154. 5:30Databricks in order to interact with
  155. 5:33Spark and do projects. And here the big
  156. 5:35advantage, you don't have manually to
  157. 5:37configure all this complexity. So you
  158. 5:39don't have to start drivers manually,
  159. 5:41you don't have to launch the executors
  160. 5:43yourself, don't talk directly to the
  161. 5:45cluster manager. The only one thing that
  162. 5:48you have to decide and to configure is
  163. 5:50actually how many workers do I need for
  164. 5:52my cluster. Everything else going to be
  165. 5:54handled internally in Databricks. So if
  166. 5:57you use Databricks, you don't have to
  167. 5:59worry about all those details. You just
  168. 6:01focus on writing your code for your
  169. 6:03project and that's it. But of course, it
  170. 6:05is nice to understand those concepts as
  171. 6:07you are using the platform. Now the
  172. 6:09other option that you have is to not use
  173. 6:11Databricks and you're going to run Spark
  174. 6:13in more traditional environments. And
  175. 6:15here you have to master all those
  176. 6:17details because, my friend, you have to
  177. 6:19configure everything on your own. The
  178. 6:21number of executors, the memory, the
  179. 6:23number of cores, the cluster manager
  180. 6:25type, the allocations, and all those
  181. 6:27details. So the level of details you
  182. 6:29need depends really on where you run
  183. 6:31Spark. But of course, the most important
  184. 6:33part is to understand how Spark works
  185. 6:36behind the scenes and then you
  186. 6:37understand how to use Spark everywhere.
  187. 6:39All right, friends, so with this now you
  188. 6:41have a clear picture about how things
  189. 6:43works behind the scenes in Spark. And
  190. 6:45now what we're going to do, we're going
  191. 6:46to zoom out a little bit because Spark
  192. 6:48is not only a distributed engine, it is
  193. 6:51a full ecosystem. So Spark looks like
  194. 6:54this. At the center we have something
  195. 6:56called Spark Core. It is the main part
  196. 6:58where it's going to handles the
  197. 7:00distributed execution, the memory
  198. 7:02management, task scheduling, and fault
  199. 7:05tolerance. So basically, all those heavy
  200. 7:07infrastructure work that happens behind
  201. 7:10the scenes. And everything else going to
  202. 7:12be built on top of the Spark Core. So
  203. 7:15the other things are actually like we
  204. 7:17have different libraries. Like for
  205. 7:19example, the very common and famous
  206. 7:21library we have the Spark SQL. It
  207. 7:24provides us very simple tools that looks
  208. 7:26like SQL to query and manipulate your
  209. 7:29data. Another library we have the Spark
  210. 7:31Streaming. This going to be specially
  211. 7:33important if you want to process
  212. 7:35real-time data. Like for example,
  213. 7:37streaming from Kafka platform. Another
  214. 7:40library we have the Spark MLlib. If you
  215. 7:43are data scientist and you want to build
  216. 7:45machine learning models. And another one
  217. 7:48we have the GraphX. If you want to work
  218. 7:50with graph data like networking or
  219. 7:52relationships. And the last library we
  220. 7:55have the SparkR. If you are an R user
  221. 7:57and you want to work with Spark. So now
  222. 7:59by looking to this, the Spark Core is
  223. 8:01the engine. And then you have those
  224. 8:03different libraries. They are like tools
  225. 8:06that allow us to do different type of
  226. 8:08work using Spark. Now another thing in
  227. 8:10the ecosystem, you can actually use
  228. 8:12different languages to interact with
  229. 8:15Spark. So we can use Python, R, SQL,
  230. 8:19Scala, Java. So my friends, you can pick
  231. 8:22the language that suits you to work with
  232. 8:24Spark because at the end it doesn't
  233. 8:26matter which one you pick. Under the
  234. 8:28hood, everything eventually go through
  235. 8:30the Spark Core. So now by looking to
  236. 8:32this, you can understand why Spark
  237. 8:33became so powerful because this is not
  238. 8:36only like an engine or one tool, it is
  239. 8:39an entire platform. And now I totally
  240. 8:41understand if I might get you scared
  241. 8:43because if you look to the ecosystem and
  242. 8:45you want to learn Spark, you might say,
  243. 8:47"I'm going to go and learn all those
  244. 8:49stuff." Well, my friends, don't worry
  245. 8:51because you don't have to learn
  246. 8:53everything. I'm going to give you now my
  247. 8:54recommendations on how to learn it and
  248. 8:57my honest opinion. Now if you are a data
  249. 8:59analyst, then you just learn the theory
  250. 9:01about the Spark Core, just some basic
  251. 9:04understanding. And now about the
  252. 9:05libraries, you just need to learn the
  253. 9:08Spark SQL and actually that's it because
  254. 9:11it's going to gives you enough tools
  255. 9:12that you need as a data analyst to work
  256. 9:14with your data. But now if you are a
  257. 9:16data engineer, then it really depends on
  258. 9:18the environments. If you're going to use
  259. 9:20Spark in platforms like Databricks, then
  260. 9:22I'm going to say the same things. You
  261. 9:24just need to learn the basic theory
  262. 9:26about the Spark Core. But if you want to
  263. 9:28build and manage Spark in the
  264. 9:30traditional way without Databricks, then
  265. 9:33you need more than that. You need deeper
  266. 9:35knowledge on the Spark Core because
  267. 9:37you're going to have to configure
  268. 9:39everything on your own. Now about the
  269. 9:40libraries, again you have to master the
  270. 9:42Spark SQL because you have to manipulate
  271. 9:45and transform the data. As well, you
  272. 9:47have to learn the Spark Streaming
  273. 9:50because in many modern projects you will
  274. 9:51be streaming data from platforms like
  275. 9:54Kafka. And this is your library in order
  276. 9:56to do that. So, actually that's it. Now,
  277. 9:59if you are a data scientist, then it is
  278. 10:01similar to the analyst. You need to have
  279. 10:03as well some basic understanding about
  280. 10:04the Spark Core and about the libraries,
  281. 10:07you have to master as well the Spark
  282. 10:08SQL. And now, about the Spark MLlib here
  283. 10:11is my honest opinion, you don't really
  284. 10:13need to learn it because currently in
  285. 10:15many modern machine learning projects,
  286. 10:17we prefer to use external libraries like
  287. 10:20the scikit-learn, TensorFlow, and
  288. 10:22PyTorch. So, as a data scientist, you're
  289. 10:25going to end up using Spark to just do
  290. 10:27data preparations and maybe feature
  291. 10:29engineering. But, for model training,
  292. 10:31you will be using something external. So
  293. 10:33now, by looking to the whole big picture
  294. 10:35and to those libraries, you can
  295. 10:36understand quickly that actually the
  296. 10:38most important component is Spark SQL
  297. 10:42because it going to gives you many
  298. 10:43amazing tools that looks like SQL. And
  299. 10:46my friend, SQL is very easy to use.
  300. 10:48That's going to help you to manipulate
  301. 10:50and work with your data. And about the
  302. 10:52programming language, this library going
  303. 10:54to allows you to write only Python
  304. 10:56PySpark code or as well to just write
  305. 10:59SQL queries like you do in any
  306. 11:01databases. Then, it doesn't matter which
  307. 11:03language you use because everything
  308. 11:04going to go through the Spark Core to
  309. 11:07the Spark engine. So, it's going to be
  310. 11:08super fast. And this is exactly the
  311. 11:10reason why we going to deep dive into
  312. 11:12this library, the Spark SQL. So, that's
  313. 11:14it, my friends. This is what Spark, the
  314. 11:17architecture, how it works behind the
  315. 11:18scenes. That is exactly how it should
  316. 11:20work. The whole ecosystem and how to
  317. 11:22learn it. And of course, if you want to
  318. 11:24learn Spark with me, I will deep dive
  319. 11:26now into the Spark SQL in order to learn
  320. 11:28all the commands and how to use it in
  321. 11:30real projects. And now, my friends, if
  322. 11:32you enjoy this type of free content
  323. 11:34where I'm sketching the complex concepts
  324. 11:36using those [music] animated visuals,
  325. 11:38then support the channel by subscribing,
  326. 11:40liking, and commenting. This going to
  327. 11:42help us to grow and as well reach nice
  328. 11:44people like you. So, if you're still
  329. 11:45here, thank you so much for watching and
  330. 11:47I will see you in the next video.
  331. 11:48Bye-bye. Okay, bye. Bye.
  332. 11:55>> [music]

About this transcript

This page contains the full transcript of What is Spark? (Visual Explanation) by Data with Baraa, generated from the public captions YouTube serves with the video. The transcript has 2,313 words across 332 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.