Machine Learning Development Life Cycle | MLDLC in Data Science — Transcript
Full transcript
- 0:00Hello guys, welcome to my YouTube
- 0:02channel. This is 100 Days of Machine
- 0:04Learning and today is day nine. And
- 0:06today we are going to cover a very
- 0:08important topic called the Machine
- 0:09Learning Development Life Cycle. And
- 0:12this is a very important topic. Let me
- 0:14tell you why. So far, all the topics we
- 0:17have covered were focused on either the
- 0:20why or the what. Why? And what? We have
- 0:24been focusing on these two questions so
- 0:26far. This will be the first video that
- 0:29will focus on the question how to do it
- 0:31. Okay? So, to be honest, this is the
- 0:34first time we are diving into the 'how'
- 0:36part of machine learning, i.e., how
- 0:38machine learning is done. Okay? And
- 0:41guess what, this particular video is
- 0:43important because the upcoming videos
- 0:46will be based on what we discuss here.
- 0:49Okay? So, this is like a roadmap for
- 0:52all the future videos. Okay? The next
- 0:5491 videos are going to stem from this
- 0:57video. So trust me, it's very important
- 0:59. Okay? So let's focus on this video.
- 1:01Let's start the discussion. Okay? So
- 1:04before starting, let me give you some
- 1:06background on why we are studying this
- 1:08topic. If you are a computer science
- 1:11student, or if you ever pick up any
- 1:13book on computer science. Then there is
- 1:16a topic you have to study there. It's
- 1:19called Software Engineering. It's a
- 1:21subject, probably in your sixth or
- 1:23seventh semester, and honestly, people
- 1:26find it very boring. But there is a
- 1:29topic there that you have to study, and
- 1:32you should study, called SDLC. Maybe
- 1:35you have heard the name; SDLC stands
- 1:38for Software Development Life Cycle. So
- 1:41, if you are in a company as a software
- 1:43developer and you are building a
- 1:46software product for that company, then
- 1:48you have to follow that SDLC. SDLC is a
- 1:52guideline for how a software product is
- 1:55built from beginning to end, okay. Now,
- 1:59since machine learning is new and
- 2:01different people across the industry
- 2:03are building machine learning-based
- 2:05software products in different
- 2:07companies, researchers felt there
- 2:09should be a common pattern, procedure,
- 2:12or process so everyone can get a
- 2:14guideline that, okay, this is how
- 2:16machine learning software is built. And
- 2:19then came this concept, called ML (D)
- 2:22LC, which stands for Machine Learning
- 2:26Development Life Cycle. Just like the
- 2:29Software Development Life Cycle, it is
- 2:30exactly the same for the Machine
- 2:32Learning Development Life Cycle. So,
- 2:33what is the Machine Learning
- 2:34Development Life Cycle? It's a set of
- 2:37guidelines that you need to follow.
- 2:40Whenever you build a machine
- 2:42learning-based software product. All
- 2:44right? Here you will find every
- 2:46guideline that will guide you from the
- 2:49idea to the product. It will tell you
- 2:51the complete process. All right? And
- 2:53that is why this video is extremely
- 2:55important because as students and
- 2:57beginners, what we do is we make a
- 2:59mistake. Our mistake is that we just
- 3:01train the model, get the accuracy, and
- 3:04stop there. We feel our job is done.
- 3:07But when you sit in interviews, you
- 3:09realize that they are looking for
- 3:11candidates who have experience in
- 3:12building end-to-end products, and
- 3:14anyone who wants to build end-to-end
- 3:16products must know this topic. All
- 3:18right? And that is why this topic is
- 3:20important. In this topic, we will cover
- 3:22, in fact, we are going to cover a
- 3:24total of nine different steps. So,
- 3:27these nine steps are what we are going
- 3:29to make 90 videos from. That is why
- 3:31this is a very important video. We will
- 3:33focus a bit on this. So, I am going to
- 3:36discuss nine steps that come under
- 3:38MLDLC. It’s a bit of a tongue-twister
- 3:41of a name. But there are nine steps. If
- 3:43you refer to any other video or refer
- 3:45to a textbook, it is possible that the
- 3:47number of steps might be slightly fewer
- 3:48or slightly more, because it is not
- 3:50properly defined yet. Because machine
- 3:52learning is still a new technology, so
- 3:54not everyone agrees on a common
- 3:56standard, but try to understand the
- 3:57core idea. The core idea you will find
- 3:59the same everywhere. So, if I am saying
- 4:01nine and someone else is saying 10, it
- 4:03means the same. All right? So, let's
- 4:06start with MLDLC. All right? Let's go
- 4:09through all the steps one by one. All
- 4:11right? So, the idea is very simple: you
- 4:13need to build a software product that
- 4:15includes machine learning. It could be
- 4:17anything. It could be a recommender
- 4:19system for your website, it could be a
- 4:22loan prediction model for your bank. It
- 4:25could be what a user's marks will be in
- 4:27school. It could be anything. Any kind
- 4:30of software product you are building
- 4:31that will involve the use of machine
- 4:33learning. So, how will you proceed? We
- 4:35are starting the discussion on this.
- 4:37Step one, step one is framing the
- 4:40problem. Meaning, if you have to build
- 4:43something, then first of all you have
- 4:45to decide a few things. All right?
- 4:47That’s when you move forward. Because
- 4:49you aren't building a school project.
- 4:51You aren't building a college project.
- 4:53You are working for a company, and that
- 4:55company is serving its clients or its
- 4:57customers. You can’t just start
- 5:00things off like that there. Realizing
- 5:02midway that, oh, we thought of
- 5:03something wrong. Let’s start all over
- 5:05again. You cannot do that because money
- 5:07is being spent there. So, while
- 5:09starting, it is your responsibility to
- 5:12frame your problem perfectly. Okay?
- 5:15This is where you decide exactly what
- 5:18the problem is. What needs to be solved
- 5:20? Who are your customers? What will the
- 5:23cost be? How many team members are
- 5:25needed? And what will the product look
- 5:27like? Is the machine learning model you
- 5:29need to implement supervised or
- 5:31unsupervised? Will the machine learning
- 5:33model you need to implement run in
- 5:35offline mode or batch mode? What kind
- 5:37of algorithms will help you? Where will
- 5:40your data come from? You try to answer
- 5:42all these kinds of questions at this
- 5:45stage. So that you get a mental idea of
- 5:46, okay, what do I need to do next? Okay
- 5:49? Once you do all of this properly,
- 5:52once you frame the problem properly,
- 5:54only then do you proceed to the second
- 5:57stage, which is gathering the data. See
- 6:00, if you are working on a machine
- 6:02learning project, then you need data;
- 6:04without data, machine learning is not
- 6:05possible. Now, data is very easy in our
- 6:08case when we are doing projects at the
- 6:10college or school level; we get our
- 6:12data from Kaggle, or someone gives it
- 6:15to us, or we download it from somewhere
- 6:17on the internet, but it’s not like
- 6:19that for companies. Data can be very
- 6:22specific and is not that easily
- 6:24available. So, I may have discussed
- 6:27with you that data can come from
- 6:29different sources. Either you get it
- 6:32directly in CSV files, then there is no
- 6:34problem at all. But sometimes it
- 6:36happens that you have to fetch data
- 6:38from an API. Meaning you hit an API,
- 6:41write Python code to fetch the data in
- 6:45JSON format, and then convert that JSON
- 6:48format into your preferred format,
- 6:51which is generally a CSV file.
- 6:54Sometimes it happens that your data is
- 6:57not publicly available on any website.
- 7:00So, in that case, what you do is you
- 7:01perform web scraping. You scrape data
- 7:04from that website. Like trivago.com is
- 7:06a website where hotel details are found
- 7:09. So what they do is they web scrape
- 7:12the data. They fetch data from various
- 7:15hotel websites by running a web scraper
- 7:17via Python code. Right? Or you might
- 7:20have seen many websites that
- 7:22dynamically fetch and show you the
- 7:24product prices from different
- 7:26e-commerce sites. So web scraping is
- 7:28happening there too. Right? Sometimes,
- 7:31your data is in your database. But you
- 7:35cannot run machine learning models
- 7:37directly on this database because it is
- 7:39a running database. If something goes
- 7:42wrong here, your website could go down.
- 7:43So what do you do? You create a data
- 7:46warehouse from it. Perhaps you might
- 7:47have heard the name. A thing called ETL
- 7:49—Extract, Transform, Load—is used
- 7:52here. We will study all this later. And
- 7:54then, what do you do with these data
- 7:55warehouses? You fetch data and do your
- 7:57work. Right? Sometimes your data is in
- 8:01tools like Spark. It is in clusters.
- 8:04Big data is basically huge data, so it
- 8:06resides in different clusters. So you
- 8:08fetch data through those clusters and
- 8:10do your work. In short, bringing your
- 8:13data is a very important stage because,
- 8:15without it, nothing will happen. So,
- 8:18bringing in the data and storing it in
- 8:21the right format. So that you can start
- 8:23your work. That is stage number two.
- 8:26Step number two. Right? So we just
- 8:28discussed data gathering. Okay? Then
- 8:31comes the third stage, data
- 8:34preprocessing. Mark my words. If you
- 8:37are bringing data from external sources
- 8:40, the data is bound to be unclean or "
- 8:42dirty" data, which you cannot use
- 8:45directly. You cannot pass that kind of
- 8:47data directly into a machine learning
- 8:49model. Because the results won't be
- 8:51good. There can be many kinds of issues
- 8:53in the data. There could be structural
- 8:55issues, missing data, outliers, or
- 8:58noisy data. It could be that data is
- 9:01coming from different places and is not
- 9:03compatible. The number of columns might
- 9:05be different. Many types of data can
- 9:07cause problems. So what do you have to
- 9:10do here? You have to do preprocessing.
- 9:12Preprocessing means the changes made
- 9:14before processing. Right? Now, there
- 9:17are many things involved here. What is
- 9:19the first thing you do here? You, you
- 9:21remove duplicates, you remove
- 9:24duplicates. Right? Uh, you remove
- 9:29missing values, you remove missing
- 9:33values, you remove outliers. Okay? You
- 9:39scale the values. So, sometimes what
- 9:41happens is that the values in your
- 9:43input columns, the value in one column
- 9:46is a very large number, and the value
- 9:48in another column is a very small
- 9:50number, so your machine learning
- 9:51algorithm is all about mathematics. It
- 9:54might have to calculate distances. Then
- 9:56those distances won't make any sense.
- 9:58If one number is in the millions and
- 9:59one number is in decimals. Right? So
- 10:01you scale down the values. This is
- 10:03called standardization. Okay? You do
- 10:06many such tasks. The key idea is that
- 10:09you have to bring your data into a
- 10:11format that your machine learning
- 10:13algorithm can easily consume. That is
- 10:16the core idea of this step. Okay? And
- 10:19this step is known as data
- 10:21preprocessing. Okay? After that comes a
- 10:24very important step which we call
- 10:27Exploratory Data Analysis or it is
- 10:29called EDA. So what is the concept of
- 10:32EDA? As the name suggests, the word
- 10:34data analysis is attached to it.
- 10:35Meaning you analyze the data here. Okay
- 10:39? Meaning you try to study the
- 10:41relationship between the input and the
- 10:43output. Okay? The whole idea is that
- 10:46since you have to build a prediction or
- 10:48machine learning-based software, before
- 10:50that you must know what is in your data
- 10:52. If you don't know that, you won't be
- 10:55able to build the model properly. So
- 10:57this entire stage is where you just
- 10:59have to do lots of experiments with the
- 11:01data. You have to extract the
- 11:03relationships hidden within the data.
- 11:05Okay? So what do you do here? Here you
- 11:07plot graphs and stuff. You do
- 11:10visualization. You plot graphs and
- 11:12stuff. You do different types of
- 11:15analysis here. Here you do univariate
- 11:18analysis. Univariate analysis means you
- 11:21do an independent analysis on each
- 11:23column. What is the mean in each column
- 11:24? What is the standard deviation? What
- 11:27kind of curve is it following? Okay?
- 11:30Then you do bivariate analysis. Meaning
- 11:32you analyze two columns together to see
- 11:35what kind of relationship they have.
- 11:37Sometimes you do multivariate analysis.
- 11:40Where you analyze three or four columns
- 11:42together. Okay? So that you understand
- 11:44the relationship between them. Okay?
- 11:46What do you do right here? You write
- 11:49the outlier detection code here as well
- 11:51. You perform outlier detection here
- 11:54too. Okay? And here, if your data is
- 11:57imbalanced, you try to convert it into
- 12:01a balanced dataset. An imbalanced
- 12:04dataset means, for example, if you are
- 12:06solving an image classification problem
- 12:09, say dog versus cat, then you have
- 12:11many more cat images. And very few dog
- 12:14images. So, this is an imbalanced
- 12:16dataset. So, you handle this. Okay? The
- 12:19whole idea is to build a concrete
- 12:20understanding of your data in your mind
- 12:22during this stage. So that the
- 12:25subsequent steps become easier for you.
- 12:27Because, quite simply, if you are going
- 12:29to do something, you must first
- 12:30understand the fundamentals behind it.
- 12:33That is why this EDA step becomes very
- 12:35important, and many people spend a lot
- 12:38of time on it because the more time you
- 12:40spend here, the easier your further
- 12:41work becomes. You might have heard the
- 12:45proverb that if you have six hours to
- 12:48cut down a tree, you should spend four
- 12:51hours sharpening your saw or whatever
- 12:53cutting tool you have. Because the more
- 12:56time you spend on that, the less effort
- 12:58it will take to cut the tree. Right? So
- 13:01, the same principle applies here: if
- 13:03you want to build a machine learning
- 13:05model, learn as much as you can about
- 13:07the data beforehand. So that when it
- 13:10comes time for decision-making, it
- 13:11becomes easier because you already have
- 13:13a great understanding of the data. Okay
- 13:15? So, I hope you understand this fourth
- 13:17step as well. The next step is feature
- 13:19engineering and selection. So, features
- 13:22mean input columns. Okay? I have told
- 13:25you this before that there are two
- 13:27things: input and output. The input is
- 13:29called a feature, and features are
- 13:32important because your output depends
- 13:34solely on the input. Right? So, what is
- 13:36the idea behind feature engineering?
- 13:38It's that sometimes you create new
- 13:41columns on your own. So that it becomes
- 13:44a bit easier for your analysis. For
- 13:47example, suppose you are building a
- 13:49house price prediction machine learning
- 13:51model, and your inputs include area in
- 13:54square feet, number of rooms, and
- 13:56number of bathrooms. And suppose you
- 13:59don't have square feet, but just the
- 14:01number of rooms, number of washrooms,
- 14:02locality, and such data, then what
- 14:04would you do? You would remove the
- 14:07number of rooms and washrooms, and
- 14:09instead, create a single new column
- 14:11called 'square feet', which is actually
- 14:13a representation of both. The benefit
- 14:16is that now you have one column instead
- 14:18of two. So, this is called feature
- 14:20engineering, where you create new
- 14:21features. Okay? Or, you make some
- 14:24intelligent changes to existing
- 14:25features, which makes your analysis
- 14:27much easier. We will shoot three or
- 14:29four videos on this entire topic.
- 14:30Because this is one of the most
- 14:32important techniques that people use in
- 14:34this entire workflow. Okay? Then there
- 14:36is something called feature selection.
- 14:37Sometimes, you have too many features.
- 14:41Like 100 or 200 types of features. In
- 14:44that case, you cannot move forward
- 14:46using all of them. Because there are
- 14:48two reasons for that. The first reason
- 14:49is that, first of all, so many features
- 14:51aren't even helpful. It’s not
- 14:53necessary that every input impacts the
- 14:55output. So, you have to remove those
- 14:58columns or features that are not
- 15:01impacting your output. So, you select
- 15:05features, and the second reason to
- 15:07remove them is that the more columns
- 15:09you have, the more time it takes to
- 15:11train your model. Okay? So, you want to
- 15:14reduce that time as well. Therefore,
- 15:16feature engineering and feature
- 15:18selection are very, very crucial and
- 15:20important in this flow, and we will
- 15:21cover about five or six videos on this
- 15:23in the future. Okay? There are
- 15:26different techniques. I will teach you
- 15:27various techniques. At this point, just
- 15:30understand that changing input columns
- 15:32in a different way or picking out some
- 15:35important columns from them. That is
- 15:37known as feature engineering and
- 15:39feature selection. Okay? Okay. Now,
- 15:42once you are sure about your data, you
- 15:45have cleaned the data completely. You
- 15:47have also created all your good
- 15:49features. So now, you are ready to
- 15:52train your model. Okay? So now, what
- 15:54you do is bring in different algorithms
- 15:57. There are different algorithms in
- 15:59machine learning. You bring in
- 16:00different types of algorithms and feed
- 16:02your data to all of them to train those
- 16:05algorithms. All right? Generally,
- 16:07nobody just trains a single algorithm.
- 16:10Because to be honest, everyone knows
- 16:12that a particular algorithm is good for
- 16:14a specific type of data. But you never
- 16:17know which algorithm might turn out to
- 16:19be good for a particular set of data. I
- 16:22mean, what am I saying? For instance,
- 16:24there is an algorithm called Naive
- 16:26Bayes that performs very well on text
- 16:27data. But it is possible that some
- 16:30other algorithm might also perform well
- 16:32. So you never know, you have all these
- 16:34tools available. Your job is to run all
- 16:37of them and then decide which one to
- 16:39use. All right? So, what do you do in
- 16:41the model training phase? You bring in
- 16:43different algorithms. Generally, you
- 16:45bring in algorithms from different
- 16:47families. You apply neural networks as
- 16:48well. You apply ensemble techniques as
- 16:50well. You apply linear techniques as
- 16:52well. You apply kernel-based methods as
- 16:54well, and you collect the results of
- 16:56all of them. So that you can finally
- 16:58decide which model to use. All right?
- 17:00So, the second stage in this is called
- 17:03evaluation. There's a spelling mistake
- 17:05here, please ignore it. In the
- 17:07evaluation stage, you evaluate all the
- 17:09models. Now, there are different ways
- 17:12to evaluate. You have some metrics,
- 17:15called performance metrics, based on
- 17:17which you decide which model is
- 17:19performing well. Now, there are
- 17:21different metrics. In the case of
- 17:22classification, there is an accuracy
- 17:24score. For regression, there is mean
- 17:27squared error. For clustering, there is
- 17:30the Dunn index. We will learn about all
- 17:31these metrics gradually. The job of all
- 17:33these metrics is simply to tell you how
- 17:36well your models are working. All right
- 17:39? So, this entire evaluation becomes
- 17:41important, and it is through this
- 17:43evaluation that we understand which
- 17:44model we should use in the end. All
- 17:46right? After that, what do you do
- 17:48finally? You perform model selection.
- 17:50What do you do in model selection? You
- 17:52pick one or multiple algorithms and
- 17:54then tune their parameters. Every
- 17:56algorithm has parameters. What are
- 17:58parameters? Their settings. You tune
- 18:00their settings. It's like when you
- 18:03watch TV, you tune the settings
- 18:04according to your preference. Like when
- 18:07you're watching a late-night movie, you
- 18:08set a different picture mode, a
- 18:10different sound mode, and you'll
- 18:11probably increase the volume a bit.
- 18:13Right? So, when you select a final
- 18:16model, you tune its best parameters. So
- 18:19that its performance improves even
- 18:20further. Okay? That is what we call
- 18:22hyperparameter tuning. We will shoot
- 18:24two or three videos on this as well.
- 18:26Okay? Then, sometimes what do you do?
- 18:28There is a thing called ensemble
- 18:29learning. If you go to my channel, I
- 18:32have made very detailed videos on this.
- 18:35In ensemble learning, what happens is
- 18:38—sometimes, or actually very
- 18:39frequently—what do you do? You
- 18:42combine multiple machine learning
- 18:44algorithms to create a new, powerful
- 18:46method. You create a new, powerful
- 18:48algorithm. Okay? There are different
- 18:50techniques for this. Bagging, boosting,
- 18:51stacking, cascading. There are various
- 18:54techniques, and the core concept of all
- 18:56these is that you have multiple models
- 18:58and you combine them to make a bigger,
- 19:00powerful model. So, generally, when you
- 19:02use ensemble learning, your performance
- 19:05improves. So trust me, this is one of
- 19:07the steps you always perform. Trained a
- 19:09lot of models. Evaluated them all,
- 19:11tuned their hyperparameters, and then
- 19:13applied ensemble learning. So this
- 19:15results in you having a very large,
- 19:17powerful model in the end. Okay? Now,
- 19:20once you have done all this, you have
- 19:23your machine learning model that can
- 19:25make predictions. But the story doesn't
- 19:28end here. Now your main work starts.
- 19:31Now you have to convert it into
- 19:33software so that users can use it. And
- 19:36that software could be a website, a
- 19:38mobile app, or even a desktop app. Okay
- 19:41? So, what do you generally do? You
- 19:44take that model and create a file from
- 19:50it, which is generally a binary file.
- 19:55That binary file, meaning it's not a
- 19:59text file, okay? Different tools are
- 20:01used for this, like a tool called
- 20:03pickle. You take that file and what do
- 20:06you do? You convert it into an API. I
- 20:11hope you know what an API is. An API is
- 20:13a URL where, if you provide the right
- 20:16inputs, it gives you JSON data in
- 20:17return. Okay? So, the entire flow will
- 20:21be that the user provides input through
- 20:24a form on your website. This input will
- 20:29be received by your Python app on your
- 20:33server, which will send it to this API.
- 20:39Your binary file is stored here, and as
- 20:40soon as you give it all the correct
- 20:42inputs, it will make a prediction.
- 20:44Right? And you will show this
- 20:46prediction back to the user in JSON
- 20:48format. Right? Now, this might sound a
- 20:50bit complex right now. But trust me,
- 20:51it's not that difficult. We will do
- 20:53this too. We will build some end-to-end
- 20:56websites where you will see the machine
- 20:58learning model you created answering
- 21:00users as a website. Right? So, this is
- 21:03called model deployment. What do you do
- 21:05here? You take these models and put
- 21:07them on a server. Either you use Heroku
- 21:10, AWS, or GCP—Google Cloud Platform
- 21:13—we will be using all of these. So,
- 21:17your model is online and now serving
- 21:20user requests. Right? After doing this
- 21:22much, what do you do? You perform
- 21:24testing. Testing generally means beta
- 21:27testing. Beta testing means, out of all
- 21:30your users, you pick a set of loyal
- 21:32customers from whom you expect to get
- 21:35good feedback. You roll out the changes
- 21:38to them first. You must have noticed
- 21:41that whenever there's an upgrade in any
- 21:42software, it doesn't reach everyone at
- 21:44once; it comes gradually. Right? So,
- 21:47you send this new feature to some
- 21:50trusted customers and take their
- 21:53feedback. Right? An interesting
- 21:55technique is used here, which we call A
- 21:57/B testing. It is very famous. We will
- 22:00make a video on this as well. What is A
- 22:02/B testing? And by doing A/B testing,
- 22:06you decide how well the model you just
- 22:10created is working. Right? If it's
- 22:13working very well, then that's great.
- 22:15If something goes wrong, you repeat the
- 22:17previous steps. Right? Obviously, you
- 22:19might have bad data, or you might not
- 22:21have pre-processed it correctly, or
- 22:23your feature selection is poor, or
- 22:25there is an issue with the algorithm
- 22:27you implemented. So, there could be a
- 22:28problem at any stage. You go back and
- 22:30redo all the steps again. But if your
- 22:33model is correct and you are getting
- 22:34good feedback from your customers, then
- 22:35you move forward. And moving forward
- 22:38means optimizing your entire process.
- 22:41Right? I mean, what do you do? In the
- 22:43next step, the last step, which is
- 22:45optimization, what do you do there? You
- 22:47scale and launch your model on the
- 22:50server for all your customers. And
- 22:53before doing that, you perform a series
- 22:54of steps. Such as taking a backup of
- 22:57your model. Which is quite important.
- 23:00You take a backup of your data, which
- 23:03is again quite important. You ensure
- 23:05that if your model breaks or something
- 23:08goes wrong, you can roll it back and
- 23:10make it live again. You set up these
- 23:14kinds of automations on your server.
- 23:17You handle load balancing and such, so
- 23:19if there are too many users, you can
- 23:22divide the traffic and serve requests
- 23:24in the same or even less time. Then you
- 23:28decide how frequently you need to
- 23:30retrain the model. Because if you don't
- 23:33retrain the model, it might gradually
- 23:36start performing poorly. This happens
- 23:38very often. This is called rotting—
- 23:41R-O-T-T-I-N-G—meaning as time passes,
- 23:43your model's performance degrades
- 23:45because your data starts to evolve.
- 23:48Let's say you are building a mask
- 23:49detection system. You just need to
- 23:51detect whether the person in front has
- 23:53a mask on or not. Now, different types
- 23:55of masks are appearing. Like a mask
- 23:56that has started appearing where the
- 23:58bottom part looks exactly like a face.
- 23:59So obviously, our classification will
- 24:01start failing on that. So we need new
- 24:03data and we will have to retrain our
- 24:05model. So you will have to decide the
- 24:08frequency of this training. Whether it
- 24:11should be weekly or monthly, and this
- 24:13whole process must be automated.
- 24:15Because you cannot just go and repeat
- 24:17the whole thing every single week.
- 24:19Right? And you will want to optimize
- 24:22every place where there is even a
- 24:25little extra cost and keep the entire
- 24:27process perfect. Right? So this is the
- 24:31entire flow, guys. This is what you do
- 24:33when you are working on a machine
- 24:35learning project. We have told you nine
- 24:37steps. These are the nine steps you
- 24:40roughly need to follow. And throughout
- 24:43these 100 days, I will have you work on
- 24:46these nine steps in detail. Yeah, so
- 24:50that's it. That's it about today's
- 24:52video. I hope you understood this. I
- 24:56will share some more links here along
- 24:58with this video. So that once you read
- 25:01it in a bit more detail, you get some
- 25:03perspective. So yeah, I hope you liked
- 25:06the video. If so guys, please consider
- 25:09subscribing. Uh, thanks for watching.
About this transcript
This page contains the full transcript of Machine Learning Development Life Cycle | MLDLC in Data Science by CampusX, generated from the public captions YouTube serves with the video. The transcript has 4,271 words across 639 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.