YouTube2Text

Machine Learning Development Life Cycle | MLDLC in Data Science — Transcript

by CampusX · 4,271 words · 639 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Hello guys, welcome to my YouTube
  2. 0:02channel. This is 100 Days of Machine
  3. 0:04Learning and today is day nine. And
  4. 0:06today we are going to cover a very
  5. 0:08important topic called the Machine
  6. 0:09Learning Development Life Cycle. And
  7. 0:12this is a very important topic. Let me
  8. 0:14tell you why. So far, all the topics we
  9. 0:17have covered were focused on either the
  10. 0:20why or the what. Why? And what? We have
  11. 0:24been focusing on these two questions so
  12. 0:26far. This will be the first video that
  13. 0:29will focus on the question how to do it
  14. 0:31. Okay? So, to be honest, this is the
  15. 0:34first time we are diving into the 'how'
  16. 0:36part of machine learning, i.e., how
  17. 0:38machine learning is done. Okay? And
  18. 0:41guess what, this particular video is
  19. 0:43important because the upcoming videos
  20. 0:46will be based on what we discuss here.
  21. 0:49Okay? So, this is like a roadmap for
  22. 0:52all the future videos. Okay? The next
  23. 0:5491 videos are going to stem from this
  24. 0:57video. So trust me, it's very important
  25. 0:59. Okay? So let's focus on this video.
  26. 1:01Let's start the discussion. Okay? So
  27. 1:04before starting, let me give you some
  28. 1:06background on why we are studying this
  29. 1:08topic. If you are a computer science
  30. 1:11student, or if you ever pick up any
  31. 1:13book on computer science. Then there is
  32. 1:16a topic you have to study there. It's
  33. 1:19called Software Engineering. It's a
  34. 1:21subject, probably in your sixth or
  35. 1:23seventh semester, and honestly, people
  36. 1:26find it very boring. But there is a
  37. 1:29topic there that you have to study, and
  38. 1:32you should study, called SDLC. Maybe
  39. 1:35you have heard the name; SDLC stands
  40. 1:38for Software Development Life Cycle. So
  41. 1:41, if you are in a company as a software
  42. 1:43developer and you are building a
  43. 1:46software product for that company, then
  44. 1:48you have to follow that SDLC. SDLC is a
  45. 1:52guideline for how a software product is
  46. 1:55built from beginning to end, okay. Now,
  47. 1:59since machine learning is new and
  48. 2:01different people across the industry
  49. 2:03are building machine learning-based
  50. 2:05software products in different
  51. 2:07companies, researchers felt there
  52. 2:09should be a common pattern, procedure,
  53. 2:12or process so everyone can get a
  54. 2:14guideline that, okay, this is how
  55. 2:16machine learning software is built. And
  56. 2:19then came this concept, called ML (D)
  57. 2:22LC, which stands for Machine Learning
  58. 2:26Development Life Cycle. Just like the
  59. 2:29Software Development Life Cycle, it is
  60. 2:30exactly the same for the Machine
  61. 2:32Learning Development Life Cycle. So,
  62. 2:33what is the Machine Learning
  63. 2:34Development Life Cycle? It's a set of
  64. 2:37guidelines that you need to follow.
  65. 2:40Whenever you build a machine
  66. 2:42learning-based software product. All
  67. 2:44right? Here you will find every
  68. 2:46guideline that will guide you from the
  69. 2:49idea to the product. It will tell you
  70. 2:51the complete process. All right? And
  71. 2:53that is why this video is extremely
  72. 2:55important because as students and
  73. 2:57beginners, what we do is we make a
  74. 2:59mistake. Our mistake is that we just
  75. 3:01train the model, get the accuracy, and
  76. 3:04stop there. We feel our job is done.
  77. 3:07But when you sit in interviews, you
  78. 3:09realize that they are looking for
  79. 3:11candidates who have experience in
  80. 3:12building end-to-end products, and
  81. 3:14anyone who wants to build end-to-end
  82. 3:16products must know this topic. All
  83. 3:18right? And that is why this topic is
  84. 3:20important. In this topic, we will cover
  85. 3:22, in fact, we are going to cover a
  86. 3:24total of nine different steps. So,
  87. 3:27these nine steps are what we are going
  88. 3:29to make 90 videos from. That is why
  89. 3:31this is a very important video. We will
  90. 3:33focus a bit on this. So, I am going to
  91. 3:36discuss nine steps that come under
  92. 3:38MLDLC. It’s a bit of a tongue-twister
  93. 3:41of a name. But there are nine steps. If
  94. 3:43you refer to any other video or refer
  95. 3:45to a textbook, it is possible that the
  96. 3:47number of steps might be slightly fewer
  97. 3:48or slightly more, because it is not
  98. 3:50properly defined yet. Because machine
  99. 3:52learning is still a new technology, so
  100. 3:54not everyone agrees on a common
  101. 3:56standard, but try to understand the
  102. 3:57core idea. The core idea you will find
  103. 3:59the same everywhere. So, if I am saying
  104. 4:01nine and someone else is saying 10, it
  105. 4:03means the same. All right? So, let's
  106. 4:06start with MLDLC. All right? Let's go
  107. 4:09through all the steps one by one. All
  108. 4:11right? So, the idea is very simple: you
  109. 4:13need to build a software product that
  110. 4:15includes machine learning. It could be
  111. 4:17anything. It could be a recommender
  112. 4:19system for your website, it could be a
  113. 4:22loan prediction model for your bank. It
  114. 4:25could be what a user's marks will be in
  115. 4:27school. It could be anything. Any kind
  116. 4:30of software product you are building
  117. 4:31that will involve the use of machine
  118. 4:33learning. So, how will you proceed? We
  119. 4:35are starting the discussion on this.
  120. 4:37Step one, step one is framing the
  121. 4:40problem. Meaning, if you have to build
  122. 4:43something, then first of all you have
  123. 4:45to decide a few things. All right?
  124. 4:47That’s when you move forward. Because
  125. 4:49you aren't building a school project.
  126. 4:51You aren't building a college project.
  127. 4:53You are working for a company, and that
  128. 4:55company is serving its clients or its
  129. 4:57customers. You can’t just start
  130. 5:00things off like that there. Realizing
  131. 5:02midway that, oh, we thought of
  132. 5:03something wrong. Let’s start all over
  133. 5:05again. You cannot do that because money
  134. 5:07is being spent there. So, while
  135. 5:09starting, it is your responsibility to
  136. 5:12frame your problem perfectly. Okay?
  137. 5:15This is where you decide exactly what
  138. 5:18the problem is. What needs to be solved
  139. 5:20? Who are your customers? What will the
  140. 5:23cost be? How many team members are
  141. 5:25needed? And what will the product look
  142. 5:27like? Is the machine learning model you
  143. 5:29need to implement supervised or
  144. 5:31unsupervised? Will the machine learning
  145. 5:33model you need to implement run in
  146. 5:35offline mode or batch mode? What kind
  147. 5:37of algorithms will help you? Where will
  148. 5:40your data come from? You try to answer
  149. 5:42all these kinds of questions at this
  150. 5:45stage. So that you get a mental idea of
  151. 5:46, okay, what do I need to do next? Okay
  152. 5:49? Once you do all of this properly,
  153. 5:52once you frame the problem properly,
  154. 5:54only then do you proceed to the second
  155. 5:57stage, which is gathering the data. See
  156. 6:00, if you are working on a machine
  157. 6:02learning project, then you need data;
  158. 6:04without data, machine learning is not
  159. 6:05possible. Now, data is very easy in our
  160. 6:08case when we are doing projects at the
  161. 6:10college or school level; we get our
  162. 6:12data from Kaggle, or someone gives it
  163. 6:15to us, or we download it from somewhere
  164. 6:17on the internet, but it’s not like
  165. 6:19that for companies. Data can be very
  166. 6:22specific and is not that easily
  167. 6:24available. So, I may have discussed
  168. 6:27with you that data can come from
  169. 6:29different sources. Either you get it
  170. 6:32directly in CSV files, then there is no
  171. 6:34problem at all. But sometimes it
  172. 6:36happens that you have to fetch data
  173. 6:38from an API. Meaning you hit an API,
  174. 6:41write Python code to fetch the data in
  175. 6:45JSON format, and then convert that JSON
  176. 6:48format into your preferred format,
  177. 6:51which is generally a CSV file.
  178. 6:54Sometimes it happens that your data is
  179. 6:57not publicly available on any website.
  180. 7:00So, in that case, what you do is you
  181. 7:01perform web scraping. You scrape data
  182. 7:04from that website. Like trivago.com is
  183. 7:06a website where hotel details are found
  184. 7:09. So what they do is they web scrape
  185. 7:12the data. They fetch data from various
  186. 7:15hotel websites by running a web scraper
  187. 7:17via Python code. Right? Or you might
  188. 7:20have seen many websites that
  189. 7:22dynamically fetch and show you the
  190. 7:24product prices from different
  191. 7:26e-commerce sites. So web scraping is
  192. 7:28happening there too. Right? Sometimes,
  193. 7:31your data is in your database. But you
  194. 7:35cannot run machine learning models
  195. 7:37directly on this database because it is
  196. 7:39a running database. If something goes
  197. 7:42wrong here, your website could go down.
  198. 7:43So what do you do? You create a data
  199. 7:46warehouse from it. Perhaps you might
  200. 7:47have heard the name. A thing called ETL
  201. 7:49—Extract, Transform, Load—is used
  202. 7:52here. We will study all this later. And
  203. 7:54then, what do you do with these data
  204. 7:55warehouses? You fetch data and do your
  205. 7:57work. Right? Sometimes your data is in
  206. 8:01tools like Spark. It is in clusters.
  207. 8:04Big data is basically huge data, so it
  208. 8:06resides in different clusters. So you
  209. 8:08fetch data through those clusters and
  210. 8:10do your work. In short, bringing your
  211. 8:13data is a very important stage because,
  212. 8:15without it, nothing will happen. So,
  213. 8:18bringing in the data and storing it in
  214. 8:21the right format. So that you can start
  215. 8:23your work. That is stage number two.
  216. 8:26Step number two. Right? So we just
  217. 8:28discussed data gathering. Okay? Then
  218. 8:31comes the third stage, data
  219. 8:34preprocessing. Mark my words. If you
  220. 8:37are bringing data from external sources
  221. 8:40, the data is bound to be unclean or "
  222. 8:42dirty" data, which you cannot use
  223. 8:45directly. You cannot pass that kind of
  224. 8:47data directly into a machine learning
  225. 8:49model. Because the results won't be
  226. 8:51good. There can be many kinds of issues
  227. 8:53in the data. There could be structural
  228. 8:55issues, missing data, outliers, or
  229. 8:58noisy data. It could be that data is
  230. 9:01coming from different places and is not
  231. 9:03compatible. The number of columns might
  232. 9:05be different. Many types of data can
  233. 9:07cause problems. So what do you have to
  234. 9:10do here? You have to do preprocessing.
  235. 9:12Preprocessing means the changes made
  236. 9:14before processing. Right? Now, there
  237. 9:17are many things involved here. What is
  238. 9:19the first thing you do here? You, you
  239. 9:21remove duplicates, you remove
  240. 9:24duplicates. Right? Uh, you remove
  241. 9:29missing values, you remove missing
  242. 9:33values, you remove outliers. Okay? You
  243. 9:39scale the values. So, sometimes what
  244. 9:41happens is that the values in your
  245. 9:43input columns, the value in one column
  246. 9:46is a very large number, and the value
  247. 9:48in another column is a very small
  248. 9:50number, so your machine learning
  249. 9:51algorithm is all about mathematics. It
  250. 9:54might have to calculate distances. Then
  251. 9:56those distances won't make any sense.
  252. 9:58If one number is in the millions and
  253. 9:59one number is in decimals. Right? So
  254. 10:01you scale down the values. This is
  255. 10:03called standardization. Okay? You do
  256. 10:06many such tasks. The key idea is that
  257. 10:09you have to bring your data into a
  258. 10:11format that your machine learning
  259. 10:13algorithm can easily consume. That is
  260. 10:16the core idea of this step. Okay? And
  261. 10:19this step is known as data
  262. 10:21preprocessing. Okay? After that comes a
  263. 10:24very important step which we call
  264. 10:27Exploratory Data Analysis or it is
  265. 10:29called EDA. So what is the concept of
  266. 10:32EDA? As the name suggests, the word
  267. 10:34data analysis is attached to it.
  268. 10:35Meaning you analyze the data here. Okay
  269. 10:39? Meaning you try to study the
  270. 10:41relationship between the input and the
  271. 10:43output. Okay? The whole idea is that
  272. 10:46since you have to build a prediction or
  273. 10:48machine learning-based software, before
  274. 10:50that you must know what is in your data
  275. 10:52. If you don't know that, you won't be
  276. 10:55able to build the model properly. So
  277. 10:57this entire stage is where you just
  278. 10:59have to do lots of experiments with the
  279. 11:01data. You have to extract the
  280. 11:03relationships hidden within the data.
  281. 11:05Okay? So what do you do here? Here you
  282. 11:07plot graphs and stuff. You do
  283. 11:10visualization. You plot graphs and
  284. 11:12stuff. You do different types of
  285. 11:15analysis here. Here you do univariate
  286. 11:18analysis. Univariate analysis means you
  287. 11:21do an independent analysis on each
  288. 11:23column. What is the mean in each column
  289. 11:24? What is the standard deviation? What
  290. 11:27kind of curve is it following? Okay?
  291. 11:30Then you do bivariate analysis. Meaning
  292. 11:32you analyze two columns together to see
  293. 11:35what kind of relationship they have.
  294. 11:37Sometimes you do multivariate analysis.
  295. 11:40Where you analyze three or four columns
  296. 11:42together. Okay? So that you understand
  297. 11:44the relationship between them. Okay?
  298. 11:46What do you do right here? You write
  299. 11:49the outlier detection code here as well
  300. 11:51. You perform outlier detection here
  301. 11:54too. Okay? And here, if your data is
  302. 11:57imbalanced, you try to convert it into
  303. 12:01a balanced dataset. An imbalanced
  304. 12:04dataset means, for example, if you are
  305. 12:06solving an image classification problem
  306. 12:09, say dog versus cat, then you have
  307. 12:11many more cat images. And very few dog
  308. 12:14images. So, this is an imbalanced
  309. 12:16dataset. So, you handle this. Okay? The
  310. 12:19whole idea is to build a concrete
  311. 12:20understanding of your data in your mind
  312. 12:22during this stage. So that the
  313. 12:25subsequent steps become easier for you.
  314. 12:27Because, quite simply, if you are going
  315. 12:29to do something, you must first
  316. 12:30understand the fundamentals behind it.
  317. 12:33That is why this EDA step becomes very
  318. 12:35important, and many people spend a lot
  319. 12:38of time on it because the more time you
  320. 12:40spend here, the easier your further
  321. 12:41work becomes. You might have heard the
  322. 12:45proverb that if you have six hours to
  323. 12:48cut down a tree, you should spend four
  324. 12:51hours sharpening your saw or whatever
  325. 12:53cutting tool you have. Because the more
  326. 12:56time you spend on that, the less effort
  327. 12:58it will take to cut the tree. Right? So
  328. 13:01, the same principle applies here: if
  329. 13:03you want to build a machine learning
  330. 13:05model, learn as much as you can about
  331. 13:07the data beforehand. So that when it
  332. 13:10comes time for decision-making, it
  333. 13:11becomes easier because you already have
  334. 13:13a great understanding of the data. Okay
  335. 13:15? So, I hope you understand this fourth
  336. 13:17step as well. The next step is feature
  337. 13:19engineering and selection. So, features
  338. 13:22mean input columns. Okay? I have told
  339. 13:25you this before that there are two
  340. 13:27things: input and output. The input is
  341. 13:29called a feature, and features are
  342. 13:32important because your output depends
  343. 13:34solely on the input. Right? So, what is
  344. 13:36the idea behind feature engineering?
  345. 13:38It's that sometimes you create new
  346. 13:41columns on your own. So that it becomes
  347. 13:44a bit easier for your analysis. For
  348. 13:47example, suppose you are building a
  349. 13:49house price prediction machine learning
  350. 13:51model, and your inputs include area in
  351. 13:54square feet, number of rooms, and
  352. 13:56number of bathrooms. And suppose you
  353. 13:59don't have square feet, but just the
  354. 14:01number of rooms, number of washrooms,
  355. 14:02locality, and such data, then what
  356. 14:04would you do? You would remove the
  357. 14:07number of rooms and washrooms, and
  358. 14:09instead, create a single new column
  359. 14:11called 'square feet', which is actually
  360. 14:13a representation of both. The benefit
  361. 14:16is that now you have one column instead
  362. 14:18of two. So, this is called feature
  363. 14:20engineering, where you create new
  364. 14:21features. Okay? Or, you make some
  365. 14:24intelligent changes to existing
  366. 14:25features, which makes your analysis
  367. 14:27much easier. We will shoot three or
  368. 14:29four videos on this entire topic.
  369. 14:30Because this is one of the most
  370. 14:32important techniques that people use in
  371. 14:34this entire workflow. Okay? Then there
  372. 14:36is something called feature selection.
  373. 14:37Sometimes, you have too many features.
  374. 14:41Like 100 or 200 types of features. In
  375. 14:44that case, you cannot move forward
  376. 14:46using all of them. Because there are
  377. 14:48two reasons for that. The first reason
  378. 14:49is that, first of all, so many features
  379. 14:51aren't even helpful. It’s not
  380. 14:53necessary that every input impacts the
  381. 14:55output. So, you have to remove those
  382. 14:58columns or features that are not
  383. 15:01impacting your output. So, you select
  384. 15:05features, and the second reason to
  385. 15:07remove them is that the more columns
  386. 15:09you have, the more time it takes to
  387. 15:11train your model. Okay? So, you want to
  388. 15:14reduce that time as well. Therefore,
  389. 15:16feature engineering and feature
  390. 15:18selection are very, very crucial and
  391. 15:20important in this flow, and we will
  392. 15:21cover about five or six videos on this
  393. 15:23in the future. Okay? There are
  394. 15:26different techniques. I will teach you
  395. 15:27various techniques. At this point, just
  396. 15:30understand that changing input columns
  397. 15:32in a different way or picking out some
  398. 15:35important columns from them. That is
  399. 15:37known as feature engineering and
  400. 15:39feature selection. Okay? Okay. Now,
  401. 15:42once you are sure about your data, you
  402. 15:45have cleaned the data completely. You
  403. 15:47have also created all your good
  404. 15:49features. So now, you are ready to
  405. 15:52train your model. Okay? So now, what
  406. 15:54you do is bring in different algorithms
  407. 15:57. There are different algorithms in
  408. 15:59machine learning. You bring in
  409. 16:00different types of algorithms and feed
  410. 16:02your data to all of them to train those
  411. 16:05algorithms. All right? Generally,
  412. 16:07nobody just trains a single algorithm.
  413. 16:10Because to be honest, everyone knows
  414. 16:12that a particular algorithm is good for
  415. 16:14a specific type of data. But you never
  416. 16:17know which algorithm might turn out to
  417. 16:19be good for a particular set of data. I
  418. 16:22mean, what am I saying? For instance,
  419. 16:24there is an algorithm called Naive
  420. 16:26Bayes that performs very well on text
  421. 16:27data. But it is possible that some
  422. 16:30other algorithm might also perform well
  423. 16:32. So you never know, you have all these
  424. 16:34tools available. Your job is to run all
  425. 16:37of them and then decide which one to
  426. 16:39use. All right? So, what do you do in
  427. 16:41the model training phase? You bring in
  428. 16:43different algorithms. Generally, you
  429. 16:45bring in algorithms from different
  430. 16:47families. You apply neural networks as
  431. 16:48well. You apply ensemble techniques as
  432. 16:50well. You apply linear techniques as
  433. 16:52well. You apply kernel-based methods as
  434. 16:54well, and you collect the results of
  435. 16:56all of them. So that you can finally
  436. 16:58decide which model to use. All right?
  437. 17:00So, the second stage in this is called
  438. 17:03evaluation. There's a spelling mistake
  439. 17:05here, please ignore it. In the
  440. 17:07evaluation stage, you evaluate all the
  441. 17:09models. Now, there are different ways
  442. 17:12to evaluate. You have some metrics,
  443. 17:15called performance metrics, based on
  444. 17:17which you decide which model is
  445. 17:19performing well. Now, there are
  446. 17:21different metrics. In the case of
  447. 17:22classification, there is an accuracy
  448. 17:24score. For regression, there is mean
  449. 17:27squared error. For clustering, there is
  450. 17:30the Dunn index. We will learn about all
  451. 17:31these metrics gradually. The job of all
  452. 17:33these metrics is simply to tell you how
  453. 17:36well your models are working. All right
  454. 17:39? So, this entire evaluation becomes
  455. 17:41important, and it is through this
  456. 17:43evaluation that we understand which
  457. 17:44model we should use in the end. All
  458. 17:46right? After that, what do you do
  459. 17:48finally? You perform model selection.
  460. 17:50What do you do in model selection? You
  461. 17:52pick one or multiple algorithms and
  462. 17:54then tune their parameters. Every
  463. 17:56algorithm has parameters. What are
  464. 17:58parameters? Their settings. You tune
  465. 18:00their settings. It's like when you
  466. 18:03watch TV, you tune the settings
  467. 18:04according to your preference. Like when
  468. 18:07you're watching a late-night movie, you
  469. 18:08set a different picture mode, a
  470. 18:10different sound mode, and you'll
  471. 18:11probably increase the volume a bit.
  472. 18:13Right? So, when you select a final
  473. 18:16model, you tune its best parameters. So
  474. 18:19that its performance improves even
  475. 18:20further. Okay? That is what we call
  476. 18:22hyperparameter tuning. We will shoot
  477. 18:24two or three videos on this as well.
  478. 18:26Okay? Then, sometimes what do you do?
  479. 18:28There is a thing called ensemble
  480. 18:29learning. If you go to my channel, I
  481. 18:32have made very detailed videos on this.
  482. 18:35In ensemble learning, what happens is
  483. 18:38—sometimes, or actually very
  484. 18:39frequently—what do you do? You
  485. 18:42combine multiple machine learning
  486. 18:44algorithms to create a new, powerful
  487. 18:46method. You create a new, powerful
  488. 18:48algorithm. Okay? There are different
  489. 18:50techniques for this. Bagging, boosting,
  490. 18:51stacking, cascading. There are various
  491. 18:54techniques, and the core concept of all
  492. 18:56these is that you have multiple models
  493. 18:58and you combine them to make a bigger,
  494. 19:00powerful model. So, generally, when you
  495. 19:02use ensemble learning, your performance
  496. 19:05improves. So trust me, this is one of
  497. 19:07the steps you always perform. Trained a
  498. 19:09lot of models. Evaluated them all,
  499. 19:11tuned their hyperparameters, and then
  500. 19:13applied ensemble learning. So this
  501. 19:15results in you having a very large,
  502. 19:17powerful model in the end. Okay? Now,
  503. 19:20once you have done all this, you have
  504. 19:23your machine learning model that can
  505. 19:25make predictions. But the story doesn't
  506. 19:28end here. Now your main work starts.
  507. 19:31Now you have to convert it into
  508. 19:33software so that users can use it. And
  509. 19:36that software could be a website, a
  510. 19:38mobile app, or even a desktop app. Okay
  511. 19:41? So, what do you generally do? You
  512. 19:44take that model and create a file from
  513. 19:50it, which is generally a binary file.
  514. 19:55That binary file, meaning it's not a
  515. 19:59text file, okay? Different tools are
  516. 20:01used for this, like a tool called
  517. 20:03pickle. You take that file and what do
  518. 20:06you do? You convert it into an API. I
  519. 20:11hope you know what an API is. An API is
  520. 20:13a URL where, if you provide the right
  521. 20:16inputs, it gives you JSON data in
  522. 20:17return. Okay? So, the entire flow will
  523. 20:21be that the user provides input through
  524. 20:24a form on your website. This input will
  525. 20:29be received by your Python app on your
  526. 20:33server, which will send it to this API.
  527. 20:39Your binary file is stored here, and as
  528. 20:40soon as you give it all the correct
  529. 20:42inputs, it will make a prediction.
  530. 20:44Right? And you will show this
  531. 20:46prediction back to the user in JSON
  532. 20:48format. Right? Now, this might sound a
  533. 20:50bit complex right now. But trust me,
  534. 20:51it's not that difficult. We will do
  535. 20:53this too. We will build some end-to-end
  536. 20:56websites where you will see the machine
  537. 20:58learning model you created answering
  538. 21:00users as a website. Right? So, this is
  539. 21:03called model deployment. What do you do
  540. 21:05here? You take these models and put
  541. 21:07them on a server. Either you use Heroku
  542. 21:10, AWS, or GCP—Google Cloud Platform
  543. 21:13—we will be using all of these. So,
  544. 21:17your model is online and now serving
  545. 21:20user requests. Right? After doing this
  546. 21:22much, what do you do? You perform
  547. 21:24testing. Testing generally means beta
  548. 21:27testing. Beta testing means, out of all
  549. 21:30your users, you pick a set of loyal
  550. 21:32customers from whom you expect to get
  551. 21:35good feedback. You roll out the changes
  552. 21:38to them first. You must have noticed
  553. 21:41that whenever there's an upgrade in any
  554. 21:42software, it doesn't reach everyone at
  555. 21:44once; it comes gradually. Right? So,
  556. 21:47you send this new feature to some
  557. 21:50trusted customers and take their
  558. 21:53feedback. Right? An interesting
  559. 21:55technique is used here, which we call A
  560. 21:57/B testing. It is very famous. We will
  561. 22:00make a video on this as well. What is A
  562. 22:02/B testing? And by doing A/B testing,
  563. 22:06you decide how well the model you just
  564. 22:10created is working. Right? If it's
  565. 22:13working very well, then that's great.
  566. 22:15If something goes wrong, you repeat the
  567. 22:17previous steps. Right? Obviously, you
  568. 22:19might have bad data, or you might not
  569. 22:21have pre-processed it correctly, or
  570. 22:23your feature selection is poor, or
  571. 22:25there is an issue with the algorithm
  572. 22:27you implemented. So, there could be a
  573. 22:28problem at any stage. You go back and
  574. 22:30redo all the steps again. But if your
  575. 22:33model is correct and you are getting
  576. 22:34good feedback from your customers, then
  577. 22:35you move forward. And moving forward
  578. 22:38means optimizing your entire process.
  579. 22:41Right? I mean, what do you do? In the
  580. 22:43next step, the last step, which is
  581. 22:45optimization, what do you do there? You
  582. 22:47scale and launch your model on the
  583. 22:50server for all your customers. And
  584. 22:53before doing that, you perform a series
  585. 22:54of steps. Such as taking a backup of
  586. 22:57your model. Which is quite important.
  587. 23:00You take a backup of your data, which
  588. 23:03is again quite important. You ensure
  589. 23:05that if your model breaks or something
  590. 23:08goes wrong, you can roll it back and
  591. 23:10make it live again. You set up these
  592. 23:14kinds of automations on your server.
  593. 23:17You handle load balancing and such, so
  594. 23:19if there are too many users, you can
  595. 23:22divide the traffic and serve requests
  596. 23:24in the same or even less time. Then you
  597. 23:28decide how frequently you need to
  598. 23:30retrain the model. Because if you don't
  599. 23:33retrain the model, it might gradually
  600. 23:36start performing poorly. This happens
  601. 23:38very often. This is called rotting—
  602. 23:41R-O-T-T-I-N-G—meaning as time passes,
  603. 23:43your model's performance degrades
  604. 23:45because your data starts to evolve.
  605. 23:48Let's say you are building a mask
  606. 23:49detection system. You just need to
  607. 23:51detect whether the person in front has
  608. 23:53a mask on or not. Now, different types
  609. 23:55of masks are appearing. Like a mask
  610. 23:56that has started appearing where the
  611. 23:58bottom part looks exactly like a face.
  612. 23:59So obviously, our classification will
  613. 24:01start failing on that. So we need new
  614. 24:03data and we will have to retrain our
  615. 24:05model. So you will have to decide the
  616. 24:08frequency of this training. Whether it
  617. 24:11should be weekly or monthly, and this
  618. 24:13whole process must be automated.
  619. 24:15Because you cannot just go and repeat
  620. 24:17the whole thing every single week.
  621. 24:19Right? And you will want to optimize
  622. 24:22every place where there is even a
  623. 24:25little extra cost and keep the entire
  624. 24:27process perfect. Right? So this is the
  625. 24:31entire flow, guys. This is what you do
  626. 24:33when you are working on a machine
  627. 24:35learning project. We have told you nine
  628. 24:37steps. These are the nine steps you
  629. 24:40roughly need to follow. And throughout
  630. 24:43these 100 days, I will have you work on
  631. 24:46these nine steps in detail. Yeah, so
  632. 24:50that's it. That's it about today's
  633. 24:52video. I hope you understood this. I
  634. 24:56will share some more links here along
  635. 24:58with this video. So that once you read
  636. 25:01it in a bit more detail, you get some
  637. 25:03perspective. So yeah, I hope you liked
  638. 25:06the video. If so guys, please consider
  639. 25:09subscribing. Uh, thanks for watching.

About this transcript

This page contains the full transcript of Machine Learning Development Life Cycle | MLDLC in Data Science by CampusX, generated from the public captions YouTube serves with the video. The transcript has 4,271 words across 639 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.