Complete Machine Learning In 6 Hours| Krish Naik — Transcript
Full transcript
- 0:06so today's session what all things we
- 0:08are basically going to discuss so first
- 0:10of all we going to discuss about
- 0:12different types of machine learning
- 0:13algorithm like how many different types
- 0:15of machine learning
- 0:16algor understand the purpose of taking
- 0:20this session is to clear the interviews
- 0:23okay clear the interviews once you go
- 0:25for a data science interviews and all
- 0:28the main purpose is to clear the
- 0:29interviews I've seen people who knew
- 0:32machine learning algorithms in a proper
- 0:34way okay they were definitely able to
- 0:36clear it because they just explain the
- 0:38algorithms in a better way to the
- 0:40recruiter so that they got hired first
- 0:42of all is the introduction to machine
- 0:45learning here I'm just specifically
- 0:47going to talk about AI versus ml versus
- 0:51DL versus data sign then the second
- 0:53thing that we are going to talk about
- 0:55over here is the difference between
- 0:58supervised MS
- 1:00and unsupervised ml the third thing that
- 1:03we are probably going to discuss about
- 1:05is something called as linear regression
- 1:08so we are going to clearly understand
- 1:10the maths and geometric intuition the
- 1:13next thing that we are probably going to
- 1:15discuss about is R square and adjusted R
- 1:18square the fifth topic that we are going
- 1:20to discuss about is Ridge and lasso
- 1:23regression the first topic that we are
- 1:25going to discuss about is AI versus ml
- 1:30versus DL versus data science so this is
- 1:34the first topic that we are probably
- 1:35going to discuss if you really want to
- 1:38understand the difference between AI
- 1:39versus ml versus DL versus data science
- 1:41we will go in this specific format so
- 1:43just imagine the entire universe so this
- 1:46entire universe I will probably call it
- 1:48as an AI now specifically when I say AI
- 1:51this basically means AI artificial
- 1:53intelligence whatever role you are in
- 1:55you are as a machine learning developer
- 1:57you working as a deep learning developer
- 1:59Vision developer or a data scientist or
- 2:02an AI engineer at the end of the day you
- 2:05are actually creating AI application so
- 2:09if I really want to Define what is this
- 2:11artificial intelligence you can just say
- 2:13that it is a process wherein we create
- 2:16some kind of applications in which it
- 2:19will be able to do its task without any
- 2:22human intervention so that basically
- 2:24means a person need not monitor this AI
- 2:27application automatically it'll be able
- 2:29to make decisions it will be able to
- 2:31perform its task and it will be able to
- 2:34do many things so this is what an AI
- 2:36application is some of the examples that
- 2:38I would definitely like to consider so
- 2:41the first example that I would like to
- 2:43consider AI application AI module
- 2:46Netflix has an AI module suppose if you
- 2:49see a kind of action movie for some time
- 2:53then the kind of AI work or AI work that
- 2:56is basically implemented over here is
- 2:57something called as recommendation
- 3:00so here through this application what
- 3:04happens is that when you're continuously
- 3:06seeing the action movies then
- 3:08automatically the AI module that is
- 3:10present inside Netflix will make sure
- 3:13that it gives us recommendation on
- 3:15action movies second if I take an
- 3:18example of comedy movie If I
- 3:20continuously see comedy movie then also
- 3:22it'll give us the recommendation of the
- 3:24comedy movie so this through this what
- 3:26happens is that it understands your
- 3:28behavior and it is being able to do its
- 3:30task without asking you anything the
- 3:33second example that I would like to take
- 3:35up in is
- 3:36amazon.in now amazon.in again if you buy
- 3:39an
- 3:40iPhone then it may recommend you a
- 3:43headphones so this kind of
- 3:45recommendation is also a part of AI
- 3:48module that is integrated with the
- 3:49amazon.in website the ads that you see
- 3:52probably when you opening my channel
- 3:55through which I get paid a little bit
- 3:56from my from a from the hard work that I
- 3:59do in YouTube right so through that ads
- 4:02how that is recommended to you uh that
- 4:05is also an AI engine that is included in
- 4:07the YouTube channel itself which really
- 4:09plays it is a business-driven goal
- 4:12understand it is a business driven
- 4:13things that we basically do with the
- 4:15help of AI one more example that I would
- 4:17like to give you is if I consider it
- 4:20self-driving cars so here you'll be able
- 4:23to see self-driving cars if you take an
- 4:25example of Tesla so self-driving cars
- 4:27what happens based on the road it is
- 4:29able ble to drive it automatically who
- 4:31is doing that there is an AI application
- 4:33integrated with the car itself right so
- 4:36if I consider all these things these all
- 4:38are AI application at the end of the day
- 4:42whatever role you do you are going to
- 4:44create an AI application this is the
- 4:46common mistake what people do you know
- 4:48like our CEO sudhansu Kumar he has
- 4:50written in his profile that he's an AI
- 4:52engineer that basically means his goal
- 4:55is to create an AI application so
- 4:57probably in a product based companies
- 4:58you'll be seeing this kind of roles
- 4:59called as AI engineer now let's go to
- 5:01the next role which is called as machine
- 5:03learning so where does machine learning
- 5:04comes into existence so if I try to
- 5:07create this machine learning is a subset
- 5:10of AI and what is the role of machine
- 5:12learning it provides stats
- 5:15tools
- 5:17to analyze the data visualize the data
- 5:22and apart from that to do
- 5:24predictions I'm
- 5:27forecasting so you will be seeing a lot
- 5:29of machine learning algorithms so
- 5:31internally those machine learning
- 5:33algorithm the equation that we are
- 5:34basically using it is basically using it
- 5:38is having a kind of stats tool stat
- 5:40techniques because whenever we work with
- 5:42data statistics is definitely very much
- 5:44important so this exactly is called as
- 5:47machine learning so it is a subset of AI
- 5:50this is very much important to
- 5:52understand ml is a subset of AI so here
- 5:55you can see that it is a part of this
- 5:57now let's go to the next one which is
- 5:59called called as deep learning deep
- 6:01learning is again a subset of ml now
- 6:04let's consider why deep learning came
- 6:05into existence because in 1950s 60s
- 6:09scientists thought that can we make
- 6:11machine learn like how we human being
- 6:13learn so for that particular purpose
- 6:16deep learning came into existence here
- 6:18the plan is to basically mimic human
- 6:21brain so when I say mimicking human
- 6:24brain that basically means we are trying
- 6:26to mimic the human brain to implement
- 6:28something to learn something so for this
- 6:31you use something called as
- 6:32multi-layered neural networks so this is
- 6:35what deep learning is it is a subset of
- 6:37machine learning its main aim is to
- 6:40mimic human brain so they actually
- 6:42create multi-layer neural network and
- 6:45this multi-layered neural network will
- 6:47basically help you to train the machines
- 6:49or applications whatever we are trying
- 6:51to create and deep learning has really
- 6:54really done an amazing work with the
- 6:56help of deep learning we are able to
- 6:58solve such a complex complex complex use
- 7:02cases that we will be probably
- 7:04discussing as we go ahead now if I come
- 7:06to data science see this is the thing
- 7:08guys if you want to say yourself as a
- 7:10data scientist tomorrow you given a
- 7:13business use case and situation comes
- 7:15that you probably have to solve that use
- 7:17case with the help of machine learning
- 7:18algorithms or deep learning algorithms
- 7:20again the final goal is to create an AI
- 7:22application right you cannot say that I
- 7:24am a data scientist and I'll just work
- 7:26in machine learning I or I'll work in
- 7:29deep learning or I may I don't know how
- 7:31to analyze the data no you cannot do
- 7:33that when I was working in Panasonic I
- 7:36got various different kind of task
- 7:39sometime I was told to use W powerbi to
- 7:41visualize analyze the data sometime I
- 7:43was given a machine learning project
- 7:45sometime I was given a deep learning
- 7:46project so as a data scientist if I
- 7:49consider where does data scientist fall
- 7:51into this it will be a part of
- 7:53everything so if I talk about machine
- 7:56learning and deep learning with respect
- 7:58to any kind of problem statement that we
- 8:00solve the majority of the business use
- 8:03cases will be falling in two sections
- 8:05one is supervised machine learning one
- 8:08is unsupervised machine learning so most
- 8:10of the problems that you are basically
- 8:12solving this is with respect to this two
- 8:15problem statement two different types of
- 8:16machine learning algorithms that is
- 8:18supervised machine learning and deep
- 8:20learning if I talk about supervised
- 8:22machine learning two major problem
- 8:24statements that you are basically
- 8:25solving here also one is regression
- 8:28problem
- 8:30and the other one is something called as
- 8:31classification problem and in the case
- 8:34of unsupervised machine learning problem
- 8:36statement you are basically solving two
- 8:37different types of problem one is
- 8:39clustering and one is dimensionality
- 8:42reduction and there is also one more
- 8:44type which is called as reinforcement
- 8:46learning reinforcement learning I can I
- 8:50I will definitely talk about this not
- 8:52right now right now we are just focusing
- 8:53on all these things now understand what
- 8:56happens in supervised machine learning
- 8:58let's consider consider a data set so
- 9:00here I have a data set which says this
- 9:03is my age and this is my weight suppose
- 9:07I have these two specific features let's
- 9:09say that I have values like 24 62 25 63
- 9:1521 72
- 9:19257 uh 62 and many more data over here
- 9:23let's say that my task is to basically
- 9:25take this particular data and create a
- 9:27model wherein so suppose my task is that
- 9:31I need to create a model whenever it
- 9:34takes the New Age first of all we train
- 9:36this model with this data and whenever
- 9:39we take age a new age it should be able
- 9:41to give us the output of weight this
- 9:44particular model is also called as
- 9:46hypothesis okay I'll discuss about this
- 9:49today when I we discussing about linear
- 9:51regression now what are the important
- 9:53components whenever we have this kind of
- 9:55problem statement first of all you need
- 9:56to understand there are two important
- 9:59things one is independent features and
- 10:02the other one is something called as
- 10:03dependent features now let's go ahead
- 10:05and discuss what is independent feature
- 10:07independent feature basically means in
- 10:09this particular case since the input
- 10:11that I'm basically training in all those
- 10:13features becomes an independent feature
- 10:15now in this particular case my age is
- 10:17independent feature and whatever I'm
- 10:20actually predicting so when I say
- 10:21predicting I know this is my output okay
- 10:24this is the what I have to basically
- 10:27make my model uh give this as a an
- 10:29output so in this particular casee my
- 10:31dependent feature becomes weight why we
- 10:34specifically say a dependent feature
- 10:36because this is completely dependent on
- 10:38this value whenever this is increasing
- 10:40or decreasing this value is basically
- 10:41getting changed so that is the reason
- 10:44why we basically say this has
- 10:45independent and dependent feature
- 10:47whenever we are solving a problem right
- 10:50in the case of supervised machine
- 10:51learning remember they will be one
- 10:53dependent feature and there can be any
- 10:55number of independent features now let's
- 10:58go ahead and let's discuss about
- 10:59regression and classification what is
- 11:01the difference between them now let
- 11:03let's go ahead and let's discuss about
- 11:05two things one
- 11:08is let's say I want a regression problem
- 11:11statement suppose I take the same
- 11:14example as age and weight so I have
- 11:17values like as discussed 24 72 23
- 11:2271 uh 24 or 25
- 11:2671.5 okay so this kind of data I have
- 11:29see this is my output variable which is
- 11:32my dependent feature now in this
- 11:34particular dependent feature now
- 11:36whenever I'm trying to find out the
- 11:37output and in this particular output you
- 11:39have a continuous variable when you have
- 11:42a continuous variable then this becomes
- 11:44a regression problem statement now one
- 11:47example I would like to give suppose
- 11:49this is my data set right this is my age
- 11:52this is my weight suppose I am
- 11:54populating this particular data set with
- 11:56the help of scatter plot then in order
- 11:58to basically solve this problem what
- 12:01we'll do suppose if I take an example of
- 12:03linear regression I will try to draw a
- 12:05straight line and this particular line
- 12:08is my equation which is called as yal mx
- 12:11+ C and with the help of this particular
- 12:13equation I will try to find out the
- 12:15predicted points so this will be my
- 12:17predicted point this will be my
- 12:18predicted point this this any new points
- 12:21that I see over here will basically be
- 12:23my predicted point with respect to Y so
- 12:26in this way we basically solve a
- 12:28regression problem statement so this is
- 12:30very much important to understand let's
- 12:32go to the always understand in a
- 12:34regression problem statement your output
- 12:35will be a continuous variable the second
- 12:37one is basically a classification
- 12:40problem now in classification problem
- 12:42suppose I have a data set let's say that
- 12:45number of hours study number of study
- 12:48hours number of play
- 12:51hours so this is my independent feature
- 12:54let's say a number of sleeping hours and
- 12:57finally I have my output which will will
- 12:59be pass or fail so in this I have all
- 13:03this as my independent features and this
- 13:05is my dependent feature so I will be
- 13:08having some values like this and here
- 13:11either you'll be pass or fail or pass or
- 13:15fail now whenever you have in your
- 13:18output fixed number of categories then
- 13:21that becomes a classification problem
- 13:23suppose it just has two outputs then it
- 13:25becomes a binary classification if you
- 13:28have more than two different categories
- 13:30at that time it becomes a multiclass
- 13:32classification so this is the difference
- 13:34between regression problem statement and
- 13:36the classification problem statement now
- 13:39let's go ahead and let's discuss about
- 13:40something called as unsupervised machine
- 13:42learning now in unsupervised machine
- 13:44learning which is my second main topic
- 13:47over here I'm just going to write
- 13:49unsupervised machine learning now what
- 13:52exactly is unsupervised machine learning
- 13:54here whenever I talk about there are two
- 13:56main problem statement that we solve one
- 13:58is clustering
- 13:59one is dimensionality reduction let's
- 14:02take one example of a specific data set
- 14:04over here let's say that my data set is
- 14:06something called as salary and age now
- 14:10in this scenario we don't have any
- 14:12output variable no output variable no
- 14:14dependent variable then what kind of
- 14:16assumptions that we can take out from
- 14:19this particular data set suppose I have
- 14:21salary and age as my values so in this
- 14:23particular case I would like to do
- 14:25something called as clustering now why
- 14:28clustering is used just understand let's
- 14:31say I am going to do something called as
- 14:33customer segmentation now what does this
- 14:35customer segmentation do clustering
- 14:37basically means that based on this data
- 14:39I will try to find out similar groups
- 14:41groups of people suppose this is my one
- 14:44group this is my another group this is
- 14:46my third group let's say that I was able
- 14:48to create this many groups this many
- 14:50groups are clusters I'll say cluster 1 2
- 14:53three each and every cluster will be
- 14:56specifying some information this cluster
- 14:58May specify that this person uh he was
- 15:01very young but he was able to get some
- 15:03amazing salary this person it may some
- 15:06specify that these people are basically
- 15:07having more age and they are getting
- 15:10good salary these people are like middle
- 15:12class background where with respect to
- 15:14the age the salary is not that much
- 15:16increasing so here what we are doing we
- 15:18are doing clustering we are grouping
- 15:20them together main thing is grouping
- 15:23this word is very much important now why
- 15:25do we use this suppose my company
- 15:28launches is a product and I want to just
- 15:31Target this particular product to rich
- 15:33people let's say product one is for rich
- 15:35people product two is for middle class
- 15:37people so if I make this kind of
- 15:40clusters I will be able to Target my ads
- 15:43only to this kind of people let's say
- 15:46that this is the rich people this is the
- 15:48middle class people I will be able to
- 15:50Target this particular ads or this
- 15:53particular product or send this
- 15:55particular things to those specific
- 15:56group of people by that that is
- 15:59basically called as ad marketing and
- 16:00this uses something called as customer
- 16:04segmentation a very important example
- 16:07and based on this customer segmentation
- 16:08we can later apply any regression or
- 16:10classification kind of problem statement
- 16:12now coming to the second one after
- 16:14clustering which is called as
- 16:15dimensionality reduction now in
- 16:17dimensionality reduction what we are
- 16:19focusing on suppose if we have th000
- 16:22features can we reduce this features to
- 16:25lower Dimensions let's say that I want
- 16:27to convert this
- 16:29uh th000 feature to 100 features lower
- 16:32Dimension so can we do that yes it is
- 16:36possible with the help of dimensionality
- 16:38deduction algorithm there are some
- 16:40algorithms like PCA so I'll also try to
- 16:42cover this as we go ahead understand
- 16:44clustering is not a classification
- 16:46problem clustering is a grouping
- 16:48algorithm there is no output feature no
- 16:50dependent variable in clustering sorry
- 16:53in unsupervised ml so yes I will also
- 16:55try to cover up LDA we'll cover up PCA
- 16:58and all as we go ahead so with respect
- 17:00to supervised and unsupervised so first
- 17:03thing that we are going to cover is
- 17:04something called as linear regression
- 17:06the second algorithm that we will try to
- 17:08cover after linear regression is
- 17:10something called as Ridge and lasso
- 17:12third that we are going to cover is
- 17:14something called as logistic regression
- 17:16the fourth that we are basically going
- 17:17to cover is something called as decision
- 17:19tree decision tree includes both
- 17:21classification and regression four fifth
- 17:24that we are going to cover is something
- 17:25called as adab boost sixth that we are
- 17:27going to cover is something called as
- 17:28random Forest seventh that we are going
- 17:30to cover is something called as gradient
- 17:32boosting eighth that we are going to
- 17:34cover is something called as XG boost N9
- 17:37that we are going to cover is something
- 17:38called as n bias then when we go to the
- 17:41unsupervised machine learning algorithm
- 17:43the first algorithm that we are going to
- 17:45do is something called as K means K
- 17:47means algorithm then we also have DV
- 17:48scan then we are also going to do higher
- 17:50C clustering there is also something
- 17:52called as K nearest neighbor clustering
- 17:55fifth we'll try to see about PCA then
- 17:57LDA so different different things we
- 18:00will try to cover up yes svm I have
- 18:02missed here I'm going to include svm KNN
- 18:05will also get covered so I have that in
- 18:07my list probably I may miss one or two
- 18:08but we are going to cover everything so
- 18:10let's start our first algorithm linear
- 18:13regression so let's go ahead and discuss
- 18:15about linear regression linear
- 18:16regression problem statement is very
- 18:18simple guys so suppose I have let's say
- 18:21I have two features one is my X feature
- 18:23and one is my y feature let's say that X
- 18:25is nothing but age and Y is nothing but
- 18:29weight so based on these two features I
- 18:31have some data points that has been
- 18:34present over here so in linear
- 18:35regression what we try to do is that we
- 18:38try to create a model with the help of
- 18:40this training data set so this will be
- 18:43my training data set what I'm actually
- 18:45going to do is that I'm going to
- 18:47basically train a model and this model
- 18:50is nothing but a kind of hypothesis
- 18:52testing or it is just kind of hypothesis
- 18:54which takes the new age and gives the
- 18:57output of the weights and then with the
- 19:01help of performance metrics we try to
- 19:03verify whether this model is performing
- 19:05well or not now in short what we are
- 19:06going to do in linear regression is that
- 19:08we'll try to find out a best fit line
- 19:10which will actually help us to do the
- 19:12prediction that basically means if I get
- 19:14my new age over here then what should be
- 19:16my output with respect to Y okay so with
- 19:19respect to this what should be my output
- 19:21over here in this particular case
- 19:23whenever we are drawing a diagram like
- 19:24this I can basically say that Y is a
- 19:28linear function of X so this is what we
- 19:31are going to do now understand how we
- 19:33are going to create this best fit line
- 19:35this is very much important whenever we
- 19:36say linear regression it basically means
- 19:39that we are going to create a linear
- 19:40line over there you may be thinking sir
- 19:43why to create linear line why not
- 19:44nonlinear line that I'll discuss about
- 19:46it as we go ahead see other other
- 19:48algorithms so to begin with let's
- 19:51consider this line that you see over
- 19:53here right this line equation can be
- 19:56given by multiple equations someone some
- 19:58people people write yal mx + C some
- 20:01people write uh H some people write yal
- 20:05beta 0 + beta 1 into X some people write
- 20:08H Theta of xal to Theta 0 + Theta 1 into
- 20:13X many many equations are there for this
- 20:16this straight line this straight line
- 20:18many many equations are there with
- 20:20respect to many many different kind of
- 20:22notations but the first algorithm that I
- 20:24have probably learned of linear
- 20:26regression is from Andrew Ng definitely
- 20:29I would like to give him the entire
- 20:30credits and based on his notation
- 20:33whatever he has explained I'll try to
- 20:34explain you over here so the credits for
- 20:37this algorithm specifically goes to
- 20:40Andrew NG so let's consider this one
- 20:43over here in order to create this
- 20:45straight line I will basically use a
- 20:47equation which is called as H Theta so
- 20:50this is the equation of a straight line
- 20:52if I know the equation of the straight
- 20:54line whatever I can write I can write
- 20:56many things yal mx + C yal beta 0 + beta
- 21:001 * X and then I can also write one more
- 21:04that is H Theta of xal theta 0 + Theta 1
- 21:08into X of I here also you can basically
- 21:11say x of I here also you can say x of I
- 21:13now let's go ahead and let's take this
- 21:15equation for now let's take this
- 21:17equation of now so I'm I'm going to take
- 21:19out this equation and just write one
- 21:21equation through which I have also
- 21:23studied but I will definitely be adding
- 21:25some points which probably Andrew and
- 21:27could not mention mention in his video
- 21:29but I'll try my level best obviously he
- 21:32is the best I cannot even compare myself
- 21:34to him so Theta 0 + Theta 1 into X now
- 21:39let's understand what is Theta 0 Theta 1
- 21:42as I said that let's say I have a
- 21:44problem statement over here let's say I
- 21:47this is my X and this is my y this is my
- 21:49data points now what I'm doing I'm
- 21:51trying to create a best fit line like
- 21:53this now what is this best fit line what
- 21:55is uh when I say this best fit line is
- 21:57basically given by this equation what
- 21:59does Theta 0 basically indicate Theta 0
- 22:02over here is something called as
- 22:04intercept now what exactly is intercept
- 22:08intercept basically means that when your
- 22:10X is zero then H Theta of X is equal to
- 22:13Theta 0 so in this particular case
- 22:16intercept basically indicates that at
- 22:18what point you are meeting the Y AIS so
- 22:22this particular point is basically
- 22:24your intercept when your X is equal to 0
- 22:28at that point of time you'll be seeing
- 22:30that this line is intersecting the y-
- 22:32AIS whatever value this will be that is
- 22:34your intercept now the second thing is
- 22:37about your Theta 1 what is Theta 1 this
- 22:40is nothing but slope or coefficient now
- 22:43what does this basically indicate this
- 22:45indicates let let's say that this is the
- 22:47unit one unit in the x-axis and probably
- 22:50with respect to this I can find one
- 22:52point over here one point over here and
- 22:55if I try to draw this over here to here
- 22:57this is the unit movement in y so what
- 23:00does it basically say slope with the
- 23:02unit movement in one one unit movement
- 23:05towards the x-axis what is the unit
- 23:07movement in y- axis that is basically
- 23:09slope or coefficient Theta 0 and Theta 1
- 23:11two things and X of I is definitely your
- 23:14data points now our main aim is to
- 23:18create a best fit line in such a way
- 23:21that I I'll just try to show it to you
- 23:22what is our main aim let's let's
- 23:24understand what is the aim of a linear
- 23:26regression so if I take an example of
- 23:29linear regression I need to find out the
- 23:32best fit line in such a way that the
- 23:35distance
- 23:36between this data points that I have and
- 23:40the predicted points should be very very
- 23:42less suppose I'm creating a best fit
- 23:46line okay I'm creating a best fit line
- 23:49so with respect to this data points
- 23:51initially was this right but my
- 23:52predicted point is this point in this
- 23:55particular case my predicted point is
- 23:56this point so and if I do do the
- 23:58summation of all these points those
- 24:01distance should be minimal then only
- 24:04I'll be able to say that this is the
- 24:06best fit line so I I cannot definitely
- 24:08say that this is exactly the best fit
- 24:10line or not how will I say when I try to
- 24:13calculate the difference between this
- 24:15point and the predicted Point these are
- 24:17my predicted point right if I try to
- 24:19calculate the distance between them then
- 24:22I will basically have a aim to it should
- 24:24be minimal if I do the summation of all
- 24:26the distance it should be minimal
- 24:29so for that what I can do is that see
- 24:31you may be also thinking Krish why not
- 24:33just do one thing okay suppose if these
- 24:35are my data points why not just play and
- 24:38create multiple lines and try to compare
- 24:40what we can do is that we can compare
- 24:42multiple we can create multiple lines
- 24:44right like this and then whoever is
- 24:46giving the best minimal point I will go
- 24:48and select that but how many iteration
- 24:51you will do how you will come to know
- 24:52that okay this line is the best line so
- 24:55for that specific purpose we should
- 24:57start at one point and we should lead
- 25:01towards finding the best fit line start
- 25:04at one point and then we should go
- 25:06towards finding the best fit line so for
- 25:10this particular purpose what we do is
- 25:12that we create a something called as uh
- 25:15cost function I have already shown you
- 25:17what is my hypothesis function my best
- 25:19fit line equation is basically given as
- 25:21H Theta of x equal to Theta 0 + Theta 1
- 25:26* X this is my hypothesis right now
- 25:29coming to the cost function which is
- 25:32super super important why this it is
- 25:34super important because cost function
- 25:37basically what what is cost function
- 25:38over here I told right right this
- 25:41distance when I do the
- 25:42summation this distance that I when I'm
- 25:45doing the summation it should be minimal
- 25:48so if I really want to find out this
- 25:49particular distance I will be using one
- 25:51more equation how can I use a distance
- 25:54formula between the predicted and the
- 25:56real point I will just say that H Theta
- 26:00of x - y so when I say h Theta of x - Y
- 26:06what does this basically mean this is my
- 26:07real point and this is my predicted
- 26:10Point predicted point is basically given
- 26:12by H Theta of X and what I'm going to do
- 26:15I'm going to basically do the squaring
- 26:17because I may get a negative value so
- 26:18because of that I really want to do the
- 26:20squaring part Now understand one thing I
- 26:23need to also do the
- 26:25summation I = 1 to compl complete M
- 26:29let's say that I'm taking the number of
- 26:30data points over here as M because I
- 26:33need to calculate the distance between
- 26:34all the points right with respect to the
- 26:37predicted and the predict with respect
- 26:39to the real
- 26:40points so after this I also need to
- 26:44divide by 1X 2m the reason why I'm
- 26:47dividing by first of all let me show you
- 26:49why we are dividing by 1 by m 1 by m
- 26:51will give us the average of all the
- 26:53values that we have the specific reason
- 26:56why we are dividing by 1 by 2 do is for
- 26:59the derivation purpose it helps us to
- 27:02make our equation very much simpler so
- 27:05that later on when I am updating the
- 27:08weights when I say weights I'm basically
- 27:10updating Theta 0 and Theta 1 Theta 0 and
- 27:13Theta 1 at that point of time you'll be
- 27:15able to see that this particular value
- 27:18when we probably do the derivative it
- 27:20will help us to do it again I'm going to
- 27:22repeat it I'm going to write it down for
- 27:24you first of
- 27:26all now in order to find find out the
- 27:28best fit line I need to keep on changing
- 27:30Theta 0 and Theta 1 unless and until I
- 27:33get the best fit line unless and until I
- 27:35don't get the best fit line I need to
- 27:37keep on updating Theta 0 and Theta 1 now
- 27:40if I need to keep on updating Theta 0
- 27:42and Theta 1 I probably require a cost
- 27:45function okay what this cost function
- 27:47will do I'll just tell you so cost
- 27:49function over here I will specify as J
- 27:53of theta 0 comma Theta 1 is equal to now
- 27:57what is cost fun function over here what
- 27:59this distance I told right this distance
- 28:01between the H Theta of X and Y if I do
- 28:05the summation of all these things it
- 28:07needs to be minimal it needs to be less
- 28:10because with respect to an X point this
- 28:12is my y point
- 28:14right similarly with respect to this x
- 28:16point this is my y point so what I'm
- 28:19actually going to do I'm going to use a
- 28:20cost function now in this cost function
- 28:23my main aim is
- 28:25to basically write H Theta of x - y s
- 28:29this will be with respect to I I I why I
- 28:32am saying I because this will be moving
- 28:34from I equal to 1 to all the points that
- 28:37is m m is basically all the points over
- 28:41here now apart from this what I actually
- 28:44going to do I'm going to divide by 1X 2
- 28:46m I'll tell you why I'm specifically
- 28:48dividing by 1X 2 m first of all by
- 28:51dividing by m I will be getting an
- 28:53average
- 28:54output average cost function because
- 28:57here I'm iterating M the reason why I'm
- 29:00dividing by two because it will help us
- 29:01in derivation why let's say that I have
- 29:04x² if I try to find out derivative of x²
- 29:08with respect to X then what will I get I
- 29:11will basically get 2x right that is what
- 29:14is the formula what is the derivation of
- 29:16X of n it is nothing but n x of n
- 29:19minus1 so that is the reason why I'm
- 29:21actually making it 1 by two so that when
- 29:24two comes over here this two and two
- 29:26will get cancelled so I hope everybody's
- 29:29able to understand so this is my cost
- 29:32function Now understand what is this
- 29:34called as this entire equation is
- 29:36basically called as squared error
- 29:40function yes mathematical Simplicity
- 29:42basically means because when we are
- 29:44updating Theta 0 and Theta 1 we
- 29:46basically find out derivation in the
- 29:47cost function so that is the reason why
- 29:50we are specifically doing it squaring
- 29:52off is basically done because so that we
- 29:54don't get any negative values here
- 29:56squared error function now let's go
- 29:59towards the what we need to solve this
- 30:02is my cost function okay so I need to
- 30:07minimize minimize this particular value
- 30:10that is 1x 2 m summation of I = 1 2 m
- 30:15and then this will basically be H Theta
- 30:17of X of I minus y of I whole Square we
- 30:23need to minimize this by adjusting
- 30:26parameter Theta 0 and Theta 1
- 30:28this entirely is what this is nothing
- 30:31but J of theta 0 comma Theta 1 and we
- 30:36really need to minimize this so this is
- 30:38our task okay this is our task now let's
- 30:41go ahead and let's try to compare with
- 30:44two different thing one is the
- 30:46hypothesis testing and one is with
- 30:48respect to the cost
- 30:49function okay let's take an
- 30:52example so right now my equation of
- 30:58the
- 30:59hypothesis is nothing but H Theta of x
- 31:02equal to Theta 0 + Theta 1 *
- 31:06X if Theta 0 is 0 then what does this
- 31:11basically indicate can I say that it
- 31:14basically the line the line the best fit
- 31:16line passes through the origin and this
- 31:18is nothing but s Theta of xal to Theta
- 31:211 multiplied by X can I say like this
- 31:25obviously I can definitely say like this
- 31:27right so my equation will be like this
- 31:29so for right now let's consider that
- 31:33your Theta 0 is equal to 0 so this is
- 31:35what it is we have done till here we
- 31:37have minimized we have written the
- 31:39equation everything yes so it is passing
- 31:42through the origin and this is what is
- 31:44the equation I'm actually getting now
- 31:47let's take one example and let's try to
- 31:48solve this if I if I have H Theta of X
- 31:51so this is my new hypothesis considering
- 31:54that my intercept is passing through the
- 31:57region so with respect to this let's say
- 32:00that I will create one line over here
- 32:04let's say this is
- 32:05my this is my data points like X1 y1 I
- 32:11have 1 2 3 I have 1 2 3 now let's
- 32:19consider that if I have T I have data
- 32:22points like what I have data points like
- 32:24let's say I have three data points 1
- 32:26comma 1 2A 2 3 comma 3 so 1A 1 is
- 32:31nothing but this is my data point 2A 2
- 32:34is nothing but this is my data point and
- 32:363 comma 3 is this is my data point so
- 32:39these are my data points from the data
- 32:41set that I
- 32:43have so 2 comma 2 is this point and 3
- 32:47comma 3 is basically this point let's
- 32:49consider that these are my points that I
- 32:51have these are my data points now if I
- 32:54consider Theta 1 as 1 where do you think
- 32:57the straight line will pass through
- 32:59where do you think the straight line
- 33:00will pass the straight line will
- 33:02definitely pass like this right my
- 33:05straight line will definitely pass
- 33:06through all the points this same point
- 33:08becomes a prediction point also right
- 33:11same point let's consider that this is
- 33:13also getting pass through this it passes
- 33:15through all the points when Theta 1 is
- 33:17equal to 1 Theta 1 is nothing but slope
- 33:19when slope is equal to 1 in this
- 33:21scenario it passes through all the
- 33:22points now go ahead and calculate your J
- 33:25of theta so what will the form of J of
- 33:28theta 1 become because Theta 0 is 0 okay
- 33:31we can basically write 1 by 2 m
- 33:33summation of I = 1 2 three how many
- 33:36points are there three right and here I
- 33:39have J of H of theta of X1
- 33:43sorry X of theta of x i - y i
- 33:49s right now let's go ahead and compute
- 33:52now in this particular scenario what
- 33:54will happen 1X 2 m
- 33:57then what is what is this point minus y
- 34:00of I see h of X is also 1 y of I is also
- 34:04one both the point are 1 so this will
- 34:06become 1 - 1 whole S Plus because we are
- 34:09doing summation the next point is also
- 34:11falling in 2A 2 so this will become 2 -
- 34:132 s + 3 - 3 S so in total this will
- 34:18become zero so when your J of theta when
- 34:22Theta 1 is 1 Theta 1 is 1 so J of theta
- 34:261 is how much it is
- 34:29Z right so what is this J of theta 1 it
- 34:33is the cost function so let me draw the
- 34:35cost function graph over here let's say
- 34:39that this is my Theta and this is
- 34:42my so here I have 0.5 here I have 1 here
- 34:46I have 1.5 so this is my Theta here I
- 34:49have two then I have 2.5 okay then
- 34:52similarly I have 0. five then I have 1
- 34:581.5 2 2.5 this is my J of theta 1 so
- 35:04right now what is my Theta 1 my Theta 1
- 35:07is 1 at this particular Point what did I
- 35:09get J of theta 1 is nothing but zero so
- 35:12this will be my first point this will be
- 35:15my first point guys I have discussed why
- 35:18why the value will be 1X 2m basically to
- 35:20make the calculation simpler we are
- 35:22dividing by 1X 2 m is basically used to
- 35:26average aage is the sumission that we
- 35:28are actually doing over here now let's
- 35:30go ahead and let's take the second
- 35:32scenario in the second scenario let's
- 35:34consider my Theta 1 let's say that my
- 35:37Theta 1 over here is now 0.5 if my Theta
- 35:411 is 0.5 then tell me what are the
- 35:43points that I will get for x equal to
- 35:471.5 * 1 so it will come as 0.5 over
- 35:51here right then similarly when X is
- 35:54equal to 2.5 * 2 is nothing but 1 over
- 35:59here and then similarly when uh for x
- 36:03equal to
- 36:0435 multiplied by 3 see we are
- 36:07multiplying here right5 multi by 3 is
- 36:091.5 so the next point will come over
- 36:12here now when I create my best fit line
- 36:15what will happen so here is my next best
- 36:19fit line which I will probably create by
- 36:20green
- 36:23color okay so this is my second one
- 36:25which is green color here definitely
- 36:27slope is decreasing so if I go ahead and
- 36:30calculate my J of theta let's see what
- 36:32I'll get so J of theta
- 36:351 is nothing but 1X 2
- 36:39m again same equation summation of I = 1
- 36:422 3 H Theta of X of
- 36:46i - y of
- 36:49i² so what we have for over here we have
- 36:52nothing but 1X 2 m now let's do the
- 36:56summation what is this point this point
- 36:58is nothing but the predicted point and
- 37:01this point is the real point right so in
- 37:03this particular scenario the first point
- 37:05that I will get is nothing but. 5 - 1
- 37:10whole s how I'm getting. 5 - 1 whole
- 37:12Square this is 1 this is the real Point
- 37:151 this is the predicted Point .5 so here
- 37:18I'm getting. 5 - 1 whole Square the
- 37:21second point will be 1 - 2 whole s right
- 37:252 so 1 - 2 whole
- 37:28s and then I will finally get 1.5 - 3
- 37:34whole s so finally if I do this
- 37:36calculation how much I'm actually
- 37:38getting 1X 2 * 3 which is 6 here I'm
- 37:42getting
- 37:44.25 5 Square here I'm getting 1 here I'm
- 37:47getting 1.5 whole Square so my final
- 37:51output will be which I have already
- 37:53calculated it is nothing but point it
- 37:56will be approximately equal to. 58 so 58
- 38:01now with Theta as this is nothing but
- 38:04Theta Theta 1 as
- 38:07.5 right that is what Theta 1 as .5 we
- 38:11are able to get. 58 so Theta 1 is .5
- 38:15over here and. 58 will be coming
- 38:17somewhere here right so this is my next
- 38:20point which will be again in green color
- 38:23now let's go ahead and calculate the
- 38:24third condition now in third condition
- 38:26what I'm actually going to write I'm
- 38:28going to basically say Theta 1 as 0 at
- 38:31that point of time just go and assume
- 38:34what is 0 multiplied by X it will
- 38:36obviously be zero so I will be getting
- 38:38three points and my next line will be in
- 38:41this line that is the
- 38:45x-axis and this is basically all my
- 38:47points now if I go ahead and calculate
- 38:50this what is J of theta 1
- 38:52now what is J of theta 1 now in this
- 38:55particular case when my Theta 1 is equal
- 38:57= to 0 1X 2 m now this part you'll be
- 39:02able to see this is 0 - 1 0 - 2 0 -
- 39:083 okay so it will become 0 - 1 s 0 - 2 s
- 39:14and 0 - 3
- 39:16S okay so this will become 1X 6
- 39:20* 1 + 4 + 9 which will not be it will be
- 39:25nothing but 2.3 which is approximately
- 39:29equal to
- 39:302.3 then what will happen with respect
- 39:33to Theta 1 as 0 we are getting 2.3 so if
- 39:36I draw this it is nothing but with
- 39:38respect to zero I'm getting 2.
- 39:412
- 39:442.3 this is my point so similarly when
- 39:47you start constructing with Theta 1 is
- 39:49equal 2 I may get some point over here
- 39:52so here when I join this points
- 39:56together you will be seeing that I will
- 39:58be getting this kind of
- 40:01curve okay and this curve is something
- 40:04called as gradient
- 40:07descent and this gradient descent will
- 40:10play a very very important role in
- 40:14making sure that in making sure that you
- 40:17get the right Theta 1 value or light
- 40:20slope value now which is the most
- 40:22suitable point the most suitable point
- 40:24is to come over here because this is
- 40:27this this point is basically called AS
- 40:30Global
- 40:31Minima because see out of all these
- 40:34three lines which is the best fit line
- 40:35this is the best fit line right this is
- 40:38the best fit line when I had this best
- 40:40fit line my point that came over here
- 40:44was here itself this was my point that
- 40:46came over here right and I want to
- 40:48basically come to this region because
- 40:50this is my Global
- 40:52Minima when I basically am over here the
- 40:56distance between the predicted and the
- 40:58real point is very very less right so
- 41:02this specific point is basically called
- 41:04AS Global minimum but still I did not
- 41:07discuss Krish you have assumed Theta 1
- 41:10is 1 Theta 1 is .5 Theta 1 is 0 here
- 41:13also you're assuming many things right
- 41:15and then you probably calculating and
- 41:17you're creating this gradient descent
- 41:19but the thing should be that probably
- 41:22you come to one point over here and then
- 41:25you reach towards this so for that
- 41:27specific reason how do you do that how
- 41:30do I first of all come to a point and
- 41:32then move towards This Global Minima so
- 41:35for that specific case we will be using
- 41:37one convergence algorithm because if I
- 41:40come to one specific point after that I
- 41:43just need to keep on updating Theta 1
- 41:45instead of using different different
- 41:47Theta 1 value so for this we use
- 41:50something called as convergence
- 41:52algorithm so here the convergence
- 41:54algorithm basically says
- 41:59repeat until
- 42:03convergence that basically means I'm in
- 42:05a while loop let's say and here I'm
- 42:08basically going to update my Theta value
- 42:11which will be given by this notation
- 42:13which is continuous updation where I'll
- 42:15say Theta J minus I'll talk about this
- 42:19Alpha don't worry and then it will be
- 42:22derivative of theta
- 42:25J with respect to this J of theta
- 42:290 and Theta 1 so this should happen that
- 42:34basically means after we reach to a
- 42:36specific point of theta after performing
- 42:40this particular operation we should be
- 42:43able to come to the global Minima and
- 42:45this this specific thing that you are
- 42:47able to see is called as
- 42:50derivative this is called as derivative
- 42:52derivative basically means I'm trying to
- 42:54find out the slope
- 42:57derivative which I can also say it as
- 42:59slope this equation will definitely work
- 43:02guys trust me this will definitely work
- 43:04why it will work I'll just draw it show
- 43:06it to you let's say that this is my cost
- 43:09function let's say that I've got this
- 43:11gradient
- 43:12descent and let's say that my first
- 43:15point is somewhere here but I have to
- 43:18reach somewhere here right now when I
- 43:20reach this this is my Theta 1 and this
- 43:23is my J of theta 1 suppose I reach at
- 43:25this specific point and I will also have
- 43:28another gradient descent which looks
- 43:30like this let's say that in the initial
- 43:33time I reach the point over here how we
- 43:35will be coming to this minimal Global
- 43:37Minima by using this equation I'll talk
- 43:40about Alpha also don't worry now this is
- 43:42also my Theta 1 this is also my J of
- 43:44theta 1 now let's say suppose I came to
- 43:47this particular point right after coming
- 43:49to this particular point I will
- 43:52basically apply this derivative on this
- 43:55J of theta 1 okay now when I find out a
- 43:59derivative that basically means we are
- 44:00trying to find out the slope and in
- 44:02order to find the slope we just create a
- 44:04straight line like
- 44:05this which will look like this I'll just
- 44:08try to
- 44:10create so I'll try to create a slope
- 44:12like this this
- 44:15slope so if you try to find out with
- 44:17respect to this this is a positive slope
- 44:20how do we indicate it because understand
- 44:22the right hand side of the line of this
- 44:24is pointing on the top wordss Direction
- 44:27this is the best easy way to find out
- 44:30whether it is a positive slope or
- 44:31negative slope now in this particular
- 44:33case this is a positive slope now when I
- 44:36get a positive slope that basically
- 44:38means I will update my weights or Theta
- 44:401 as Theta 1 let's say I'm writing it
- 44:44over here so I will just apply this
- 44:46convergence algorithm see Theta
- 44:491 colon Theta 1 minus this learning rate
- 44:55which is called as Alpha this is my my
- 44:57learning rate I'll talk about learning
- 44:58rate don't worry then this derivative
- 45:02value in this particular case since I'm
- 45:04having a positive slope I will be
- 45:06getting a positive value let's say that
- 45:09for this Theta value I got this slope
- 45:12initially now I need to come to this
- 45:15location so for that I have to reduce
- 45:17Theta 1 so that I come to this main
- 45:20point now here you can see that I am I
- 45:23subtracting Theta 1 with something which
- 45:25is a positive number
- 45:28right this is a positive number so
- 45:29definitely I know that after some n
- 45:31number of iteration I will be able to
- 45:34come to the global Minima similarly if I
- 45:36take the right hand side and if I try to
- 45:38draw the slope in this particular case
- 45:40my slope will be
- 45:42negative so similarly I can write the
- 45:44equation as Theta
- 45:461 = to Theta 1 minus learning rate
- 45:51multiplied by a negative number so minus
- 45:54into minus will be positive right
- 45:55suppose initially my 1 was
- 45:58here my Theta 1 was here now I'll keep
- 46:01on updating the weight to come to this
- 46:02Global Minima so minus into minus is
- 46:06positive so I will basically get Theta 1
- 46:09+
- 46:10Alpha by a positive number because minus
- 46:13into minus is plus so this will
- 46:16definitely work so that we will be able
- 46:19to come over here to the global Minima
- 46:22whether it is a positive slope or a
- 46:24negative slope now what is this learning
- 46:26learning rate now learning rate based on
- 46:30this learning rate suppose I want to
- 46:32come from this point to the global
- 46:35Minima by what speed I should be coming
- 46:39what speed if my learning rate value is
- 46:41bigger what speed I may be coming
- 46:43suppose if I say usually we select
- 46:45learning rate as 01 if I select a small
- 46:48number then it'll start taking small
- 46:50small steps to move towards the optimal
- 46:52Minima but if I take a alpha value a
- 46:55huge value if it is a huge huge value
- 46:57then what will happen this uh this
- 47:00updation of the Theta 1 will keep on
- 47:02jumping here and there and the situation
- 47:03will be that it will never meet it will
- 47:07never reach the global Minima so it is a
- 47:09very very good decision to take a alpha
- 47:12small value it should also not be a very
- 47:13very small value if it becomes a very
- 47:16very small value then what will happen
- 47:18very tiny steps it will take forever to
- 47:20reach the global Minima that basically
- 47:22means my model will keep on training
- 47:24itself so definitely this Al is going to
- 47:27work now let me talk about one
- 47:30scenario one scenario will be that what
- 47:33if my my cost function has a local
- 47:36Minima what if I have a local Minima
- 47:39because here if I
- 47:41come here if I come this is a local
- 47:43Minima suppose one of my points come
- 47:46over here and finally I'm reaching over
- 47:48here what will happen in this particular
- 47:50case because in this case you'll be
- 47:52seeing that what will be my equation my
- 47:54equation will be simply Theta 1
- 47:57Theta 1 minus Alpha in this point in
- 48:01this local Minima slope will be zero so
- 48:03in this particular case my Theta 1 will
- 48:05be equal to Theta 1 now you may be
- 48:07thinking what is if this is the scenario
- 48:10then we will be stuck in local Minima
- 48:13this is called as local
- 48:15Minima but usually with respect to the
- 48:18gradient descent and the equation that
- 48:20we are using here we do not get stuck in
- 48:23local Minima because our gradient
- 48:25descent in this particular scenar iio
- 48:27will always look like this but yes in
- 48:29deep learning when we are learning about
- 48:31grade in descent and a Ann at that point
- 48:34of time we have lot of local Minima and
- 48:37because of that we have different
- 48:38different G decent algorithm like RMS
- 48:40prop we have Adam optimizers which will
- 48:43solve that specific problem so this one
- 48:46point also I wanted to mention because
- 48:48tomorrow if someone asks you as an
- 48:49interview question that what if in your
- 48:52uh do you see any local Minima in linear
- 48:54regression you can just that the cost
- 48:57function that we use will definitely not
- 49:00give us local Minima but if in deep
- 49:02learning techniques with that we are
- 49:03trying to use like Ann we have different
- 49:05different kind of optimizers which will
- 49:07solve that particular problem so that is
- 49:10the answer you basically have to give
- 49:12now let me go ahead and write with
- 49:14respect to the gradient descent
- 49:15algorithm so here again I'm going to
- 49:17write the gradient descent algorithm so
- 49:19this will be my gradient descent
- 49:21algorithm and remember guys gradient
- 49:24descent is an amazing algorithm and you
- 49:26you will definitely be using it so
- 49:29please make sure that you know this
- 49:32perfectly now some questions are that
- 49:35when will convergence stop convergence
- 49:37will stop when we come to near this area
- 49:40where my uh J of theta will be very very
- 49:44less now in gradient descent algorithm I
- 49:47will again repeat it so what did I say I
- 49:50said
- 49:51repeat until convergence I told you
- 49:54right here we have written this
- 49:55algorithm
- 49:57and now let's take it for Theta 0 and
- 49:59Theta 1 so here I will write Theta 0
- 50:02J equal to Theta
- 50:06J minus learning rate of derivative of
- 50:11theta
- 50:14J J of theta 0 and Theta 1 so this is my
- 50:19repeat until convergence now we really
- 50:22need to find out what we'll try to
- 50:24equate we'll try to first of all find
- 50:25out what is this
- 50:28now if I really want to find out
- 50:30derivative
- 50:32of derivative of derivative of theta J
- 50:37with respect to J of theta 0 and Theta 1
- 50:41so how do I write this I can definitely
- 50:44write this in a easy way okay so this
- 50:46will be derivative of theta J and
- 50:49remember J will be 0 and 1 right because
- 50:53we need to find out for 0 Theta 0 and
- 50:55Theta 1 so this will be 1 by 2 m what is
- 50:59what is J of theta 0a Theta 1 obviously
- 51:02my cost function so I will write
- 51:04summation of IAL 1 to M and here I will
- 51:08basically write J of theta of X of I
- 51:11minus y of I whole squar so if my J is
- 51:16equal to Z so what will happen for this
- 51:19so here I can specifically say that
- 51:21derivative of derivative of theta 0 J of
- 51:25theta 0a 1
- 51:27now it's simple here what I will be
- 51:29doing is that I will be simply applying
- 51:31derivative function see guys what is
- 51:34this derivative let's consider this is
- 51:36something like this 1X 2 m x² so if I
- 51:40try to find out the derivative this will
- 51:42be 2x 2 MX so 2 and 2 will get cancel so
- 51:46similarly I'll have 1 by m and here I
- 51:49will specifically be writing summation
- 51:52of I = 1 2 m h Theta of x X of I which
- 51:58will be my
- 51:59x - y of i² so this will be my
- 52:03derivative with respect to Theta 0 this
- 52:06is what I got now the second thing will
- 52:08be that when J is equal to 1 derivative
- 52:11of derivative of theta 1 J of theta 0
- 52:15comma Theta
- 52:161 in this particular case I will be
- 52:19having 1 by m summation of I = 1 to M
- 52:23then again see in this particular case
- 52:26Theta of 1 is there right Theta of 1
- 52:29basically means what if I try to replace
- 52:31this let's say that I'm trying to
- 52:33replace this H Theta of X with something
- 52:35else what is s Theta of X I know that
- 52:38right it is Theta 0 + Theta 1 * X so
- 52:42Theta 0 + Theta 1 * X so after this if
- 52:46I'm trying to find out the derivative
- 52:48with respect to Theta 0 this will
- 52:50obviously become I will be able to get
- 52:52this much right now with respect to the
- 52:54second derivative what I will be writing
- 52:56I will again be writing H thet of X of i
- 52:59- y of i s
- 53:03multiplied X of I so this Square also
- 53:06went off understand this H Theta of X is
- 53:09what see they H Theta of X is nothing
- 53:12but Theta 0 + Theta 1 * X so if I'm
- 53:16trying to find out derivative with
- 53:18respect to Theta 0 nothing will be going
- 53:19to come okay Theta 1 of X will become a
- 53:22constant in this particular case in this
- 53:25case because Theta 1 of X is there so if
- 53:28I try to find out derivative of theta 1
- 53:30into X only I'll be getting X Y Square
- 53:33will not be there it's easy right X squ
- 53:35means 2x this is the derivative of x
- 53:37square right so that square went and 1X
- 53:402 1 2 by two got cancelled so this will
- 53:44be now my convergence algorithm so here
- 53:47we have discussed about linear
- 53:48regression oh sorry I have to remove
- 53:50Square here also so let me write it
- 53:53again okay repeat until conver con let
- 53:57me write it down again repeat until
- 53:59convergence finally your two updates
- 54:03will be happening one is Theta 0 so here
- 54:06it will be Theta 0
- 54:09minus Alpha that is my learning rate 1
- 54:12by m summation of IAL 1 to M and this
- 54:17will basically be H Theta of X of I
- 54:21minus y of
- 54:23I and similarly if I want to update
- 54:26Theta 1 it will be - alpha 1 by m
- 54:30summation of I = 1 to m h Theta of X of
- 54:36I oh my God y of I uh multiplied by X of
- 54:42I Alpha is your learning rate guys Alpha
- 54:45is nothing but it is learning rate here
- 54:48we have to initialize some value like
- 54:510.1 see what is s Theta of X Theta 0 +
- 54:55Theta 1 into X right if I do derivative
- 54:58of theta 1 into x what is derivative of
- 55:01theta 1 with Theta 1 x it is nothing but
- 55:03X so this x will come over here now
- 55:07let's discuss about two important thing
- 55:09one is R square and adjusted R square
- 55:11now similarly what will happen you will
- 55:14have lot of convex functions now see if
- 55:16I talk about uh like if you have
- 55:19multiple features like X1 X2 X3 x4 at
- 55:23that point of time you will be having a
- 55:253D curve curve which looks like this
- 55:28gradient
- 55:29decent which will be something like this
- 55:40gradient it's just like coming down a
- 55:44mountain now let's discuss about two
- 55:46performance metrics which is important
- 55:48in this particular case one is R
- 55:52square and adjusted R square
- 55:57we usually use this performance metrix
- 55:59to verify how our model is and how good
- 56:01our model is with respect to linear
- 56:03regression so R square is basically
- 56:05given R square is a performance Matrix
- 56:07to check how good the specific model is
- 56:10so here we basically have a formula
- 56:12which is like 1 minus sum of residual
- 56:16divided by sum of total now this is the
- 56:19formula of R squ now what is this sum of
- 56:21residual I can basically write like this
- 56:23summation of y i Min - y i hat whole
- 56:29Square this Yi hat is nothing but H
- 56:31Theta of X just consider in this way
- 56:33divided by summation of Y of i - y mean
- 56:39y mean y s to formula this is the
- 56:42formula I'll try to explain you what
- 56:44this formula definitely says okay so
- 56:47first thing first let's consider that
- 56:49this is my this is my problem statement
- 56:51that I'm trying to solve suppose these
- 56:53are my data points and if I try to
- 56:55create the best fit
- 56:57line This Yi hat Yi hat basically means
- 57:01this specific point we are trying to
- 57:03find out the difference between this
- 57:05things difference between these things
- 57:07let's say that these are my points I'm
- 57:09trying to find out a difference between
- 57:11this predicted this is my predicted the
- 57:13point in green color are my predicted
- 57:15points which I have denoted as y i hat
- 57:18and always understand this is what Su
- 57:21sum of residual is sum of residual is
- 57:23nothing but difference between this
- 57:24point to this point this point to this
- 57:26point this point to this point this
- 57:27point to this point and I doing the all
- 57:29the summation of those now the next
- 57:32point which is very much important here
- 57:34is my X and Y what is this y IUS y y bar
- 57:39Y Bar is nothing but mean mean of Y if I
- 57:43calculate the mean of Y then I will
- 57:45probably get a line which looks like
- 57:47this I'll get a line something like this
- 57:49and then I will probably try to
- 57:51calculate the distance between each and
- 57:53every point and this specific point with
- 57:55respect to the distance between this
- 57:57point and this point the denominator
- 57:59will definitely be high right this value
- 58:02obviously this value will be higher than
- 58:04this value right the reason why it will
- 58:07be higher because the mean of this
- 58:09particular value distance will obviously
- 58:11be higher so this 1 minus high this will
- 58:16be a low value and this will be a high
- 58:18value when I try to divide Low by
- 58:23High Low by high then obviously this
- 58:26entire number will become a small number
- 58:28when this is a small number 1 minus
- 58:30small number will be a big number so
- 58:33this basically shows that our R square
- 58:35has fitted properly right it has
- 58:38basically got a very good R square now
- 58:40tell me can I get this entire R square a
- 58:43negative number let's say that in this
- 58:44particular case I got 90% can I get this
- 58:47R square as negative number there will
- 58:50be situation guys what if I create a
- 58:52best fit line which looks like
- 58:54this if I create this best fit line
- 58:57which looks like this then this value
- 58:59will be quite High it is only possible
- 59:02when this value will be higher
- 59:05than higher than this
- 59:08value okay but in the usual scenario it
- 59:11will not happen because obviously we'll
- 59:13try to fit a line which will be at least
- 59:16good it's not just like pulling one line
- 59:19somewhere we don't want to create a best
- 59:21fit line which is worse than this right
- 59:23worse than this so in this particular
- 59:26scenario you'll be saying that in R
- 59:28square now here you'll be able to see
- 59:31one one amazing feature about R square
- 59:33is that let's say let's say one scenario
- 59:36suppose I have features like let's say
- 59:38that my feature is something like uh
- 59:41let's say I have a price of a house okay
- 59:43so suppose this is my bedrooms how many
- 59:45bedrooms I have and this is basically
- 59:48the price of the house now if I if I
- 59:51probably solve this Pro problem I'll
- 59:53definitely get an R square value let's
- 59:54say the R square value is 85% let's say
- 59:57that my R square is 85% now what if if I
- 1:00:00add one more feature the one more
- 1:00:02feature basically says that okay if I
- 1:00:05add
- 1:00:06location location of the house will be
- 1:00:09definitely correlated with price so
- 1:00:12there is a definite chance that the R
- 1:00:14square value will increase let's say
- 1:00:16that R square will become 90% if I
- 1:00:19probably have this two specific feature
- 1:00:21and obviously it is basically increasing
- 1:00:23the R square because this is also
- 1:00:24correlated to price
- 1:00:26and let me change the example see first
- 1:00:29case I got by R square as 85% let's say
- 1:00:32now as soon as I added location I got
- 1:00:3590% now let's say that I added one more
- 1:00:37feature which gender is going to stay
- 1:00:40gender like male or female is going to
- 1:00:42stay you know that gender is no way
- 1:00:44correlated to price but even though I
- 1:00:47add one feature there is a scenario that
- 1:00:48my R square will still increase and it
- 1:00:51may become
- 1:00:5291% even though my feature is not that
- 1:00:56important even gender is not that
- 1:00:58important the R square formula Works in
- 1:01:01such a way that if I keep on adding
- 1:01:03features and that are not nowhere
- 1:01:05correlated this is obviously nowhere
- 1:01:07correlated this is not correlated with
- 1:01:10price then also what it does is that it
- 1:01:13is basically increasing my r² so this
- 1:01:16specific thing should not happen whether
- 1:01:19a male will stay or female will stay
- 1:01:21that does not matter at all still when
- 1:01:23you do the calculation the R square will
- 1:01:26still increase so in order to not impact
- 1:01:30the model because see now right now with
- 1:01:32this particular model where I have got
- 1:01:3490% now as soon as I see R square as 91%
- 1:01:38because it is considering this
- 1:01:40particular gender so this model will be
- 1:01:43picked right because it is performing
- 1:01:45well and is giving you a better R square
- 1:01:46value but this should not happen because
- 1:01:49that is not at all corelated this model
- 1:01:51should have been picked so in order to
- 1:01:53prevent this situation what we do we
- 1:01:55basically Ally use something called as
- 1:01:57adjusted R square now what is this
- 1:01:59adjusted R square and how it will work
- 1:02:02I'll also show it to you very very nice
- 1:02:04concept of adjusted R square so adjusted
- 1:02:06R square R square
- 1:02:08adjusted is given by the
- 1:02:11formula is given by the Formula 1 - 1 -
- 1:02:16r² * N - 1 where n is the total number
- 1:02:20of samples n minus P minus 1 this p p is
- 1:02:24nothing but number of features
- 1:02:26or predictors we'll also say or
- 1:02:28predictors suppose initially my number
- 1:02:31of predictors were in this particular
- 1:02:33scenario in this scenario where I saw
- 1:02:35this my number of predictors was two and
- 1:02:37in this particular case my number of
- 1:02:39predictor was three now if my predictor
- 1:02:41is 2 I got the r squ as 90% so in this
- 1:02:45particular scenario what all the
- 1:02:46calculation will happen okay all the
- 1:02:48calculation will happen and let's say
- 1:02:50that my R square adjusted it'll be
- 1:02:52little bit less it'll be little bit less
- 1:02:55let's say it8 is 6% let's say that my R
- 1:02:57square adjusted is 86% based on this
- 1:03:00predictor 2 now when I use my predictor
- 1:03:033 predictor basically means number of
- 1:03:05features that I'm going to use and now
- 1:03:08in this one one feature is nowhere
- 1:03:10related like gender but what we are
- 1:03:12getting we are basically getting R
- 1:03:14square increased to
- 1:03:1691% now for the R square
- 1:03:19adjusted this will not increase this
- 1:03:21will in turn decrease right now it will
- 1:03:24become 82% how it will become I'll show
- 1:03:26you I've just considered some value 8682
- 1:03:29here you can see that there is an
- 1:03:31increase here an increase is there here
- 1:03:33decrease is there now how this is
- 1:03:35basically happening see this P value
- 1:03:39that I will be putting okay if I put a p
- 1:03:42isal 3 obviously with n minus P minus 1
- 1:03:46this will become a little bit smaller
- 1:03:48number or sorry little bit uh smaller
- 1:03:50number right so now in this particular
- 1:03:53case if it is not correlated obviously
- 1:03:55this will be high when I'm increasing
- 1:03:56this so this will also be high let me
- 1:03:58write the equation something like this
- 1:04:00just a second so this will basically
- 1:04:04be okay now why probably this value may
- 1:04:08have decreased let me talk about this
- 1:04:10one what is r squ I hope everybody
- 1:04:12understood n is the number of data
- 1:04:17points p is the number of
- 1:04:21predictors if p is increasing then what
- 1:04:24will happen as P keeps on increasing
- 1:04:27this value will keep on
- 1:04:29decreasing this value will keep on
- 1:04:31decreasing if this values keep on
- 1:04:33decreasing this will be a bigger number
- 1:04:35this will obviously be a big number a
- 1:04:38big number divided by a small number
- 1:04:40what it will be obviously this will be a
- 1:04:42little bit bigger number 1 minus bigger
- 1:04:45number we will basically get some values
- 1:04:47which will be decreasing if my P value
- 1:04:49is two in this particular case it will
- 1:04:52be less smaller than this right at least
- 1:04:54it will be greater than this this
- 1:04:55particular value right when p is equal
- 1:04:57to
- 1:04:573 so with the help of P obviously R
- 1:05:01square is there to support you okay
- 1:05:03whether it is correlated or not always
- 1:05:05remember when the features are highly
- 1:05:07correlated your R square value will
- 1:05:09increase tremendously if it is less
- 1:05:12correlated then it will be there will be
- 1:05:14a small increase but there will not be a
- 1:05:16very huge increase now if I consider p
- 1:05:18is equal to 2 obviously when I'm trying
- 1:05:20to find out this uh calculation n minus
- 1:05:22P minus 1 it will obviously be greater
- 1:05:25than p is equal to 3 when p is equal to
- 1:05:283 then this value will be still more
- 1:05:30smaller and when we are dividing a
- 1:05:32bigger number by a smaller number
- 1:05:34obviously we are subtracting with one so
- 1:05:37that basically means even though my R
- 1:05:39square is 86 over here there may be a
- 1:05:41scenario since this is nowhere
- 1:05:43correlated I'm basically getting an 82%
- 1:05:45because of this entire equation so I
- 1:05:48hope you are understanding this this is
- 1:05:50very much important to understand a very
- 1:05:53very important property simple way to
- 1:05:55define is that as my P value keeps on
- 1:05:58increasing the number of predictors
- 1:06:00keeps on increasing my R squ gets
- 1:06:02adjusted whatever R square I'm getting
- 1:06:05with respect to this it will always be
- 1:06:07less than this particular R square there
- 1:06:10was one interview question that was
- 1:06:11asked one of my student between R square
- 1:06:14and adjusted R square which will always
- 1:06:15be bigger definitely the student said R
- 1:06:18square then he told him to explain about
- 1:06:20adjusted R square why does that specific
- 1:06:22happen agenda one is about Ridge lasso
- 1:06:27regression second is assumptions of
- 1:06:31linear regression the third point that
- 1:06:34we are probably going to discuss about
- 1:06:37is logistic regression then the fourth
- 1:06:42thing that we are going to discuss about
- 1:06:43is something called as confusion
- 1:06:46Matrix the fifth thing that we are going
- 1:06:49to consider about
- 1:06:51is practicals
- 1:06:54for lead lineer Ridge lasso and logistic
- 1:07:00so first topic uh that we are probably
- 1:07:03going to discuss is something called as
- 1:07:05Ridge and lasso
- 1:07:10regression so let's understand about
- 1:07:12Ridge and lasso regression if you
- 1:07:15remember in our previous session what
- 1:07:17all things we discussed linear
- 1:07:21regression and then we had discussed
- 1:07:23about the cost function we have
- 1:07:24discussed about R square adjusted
- 1:07:26adjusted R square sorry R square and
- 1:07:29adjusted R square we have discussed
- 1:07:30about it gradient descent we have
- 1:07:32discussed about it it was nothing but 1
- 1:07:34by 2 m summation of I = 1 2 m h Theta of
- 1:07:41x i -
- 1:07:45y - y i s so this is the cost function
- 1:07:50that we had discussed right yesterday
- 1:07:53and this cost function was able to give
- 1:07:55us a
- 1:07:57gradient descent with respect to the J
- 1:07:59of
- 1:08:00theta J of theta Zer or Theta not so I
- 1:08:03can also write this as J of theta comma
- 1:08:06Theta 0 comma Theta 1 now let me give
- 1:08:09you a scenario let's say that I have a
- 1:08:11scenario over here and I have this
- 1:08:14specific scenario let's say that I just
- 1:08:16have two points which looks like this
- 1:08:20okay now if I have these two specific
- 1:08:23points what will happen I will probably
- 1:08:25try to create a best fit line the best
- 1:08:27fit line will definitely pass through
- 1:08:29all the points like this if I try to
- 1:08:32calculate the cost function what will be
- 1:08:34the value of J of theta 0 comma Theta 1
- 1:08:38let's say that in this particular case
- 1:08:39since it is passing through the origin
- 1:08:41my Theta 0 will be zero okay so what
- 1:08:44will be the value of theta 0 comma Theta
- 1:08:471 so here obviously you can see that
- 1:08:49there is no difference so it will
- 1:08:50obviously become zero Now understand
- 1:08:54this data that you see right right this
- 1:08:56data is basically called as training
- 1:08:59data so this data that I have actually
- 1:09:01plotted with two points these are
- 1:09:03specifically called as training
- 1:09:05data now what is the problem in this
- 1:09:08data right now see right now exactly
- 1:09:11whatever line is basically getting
- 1:09:13created over here which is through the
- 1:09:16uh hypothesis over here you can see that
- 1:09:18it is passing through every point so
- 1:09:19that is the reason your cost is zero and
- 1:09:21our main aim is to basically minimize
- 1:09:23the cost function that is absolutely
- 1:09:26fine now in this particular case in
- 1:09:29which my model this if this model is
- 1:09:32getting trained initially this data is
- 1:09:34basically called as training data now
- 1:09:37just imagine that tomorrow new data
- 1:09:40points comes so if my new data points
- 1:09:42are here let's consider that I I want to
- 1:09:45basically uh come up with this new data
- 1:09:48point now in this particular scenario if
- 1:09:50I want to predict with respect to this
- 1:09:52particular Point let's say my predicted
- 1:09:54point is here
- 1:09:55is this the difference between the
- 1:09:57predicted and the real Point quite
- 1:10:00huge yes or no so this is basically
- 1:10:04creating a condition which is called as
- 1:10:07overfitting that basically means even
- 1:10:11though my
- 1:10:13model has given or trained well with the
- 1:10:16training
- 1:10:17data or let me write it down properly
- 1:10:20over here so this condition since since
- 1:10:23you can see that over here my each and
- 1:10:25every point is basically passing through
- 1:10:27the best fit line so because of that
- 1:10:30what happens it causes something called
- 1:10:32as
- 1:10:33overfitting so you really need to
- 1:10:35understand what is overfitting now what
- 1:10:37does overfitting mean overfitting
- 1:10:40basically means my model performs well
- 1:10:44with training data but it fails to
- 1:10:48perform well with test data now what is
- 1:10:51the test data over here the test data is
- 1:10:53basically this points the real test data
- 1:10:55answer was this points but because the
- 1:10:58my line is like this I'm actually
- 1:11:00getting the predicted point over here so
- 1:11:02this distance if I try to calculate it
- 1:11:03is quite huge so in this scenario
- 1:11:06whenever I say my model performs well
- 1:11:08with training data and it fails to
- 1:11:10perform well with test data then this
- 1:11:12scenario we say it as overfitting so
- 1:11:14this scenario when the model performs
- 1:11:16well with training data I have a
- 1:11:18condition which is called as low bias
- 1:11:20and when it fails to perform with the
- 1:11:22test data then it is basically called as
- 1:11:25high High variance very important okay I
- 1:11:28will make each and everyone understand
- 1:11:31one by one if it is performing well with
- 1:11:33the training data that is basically low
- 1:11:35bias and whenever it performs well with
- 1:11:38the test sorry fails to perform well
- 1:11:40with the fails to perform well with the
- 1:11:42test data then it is basically High
- 1:11:44variance now similarly I may have
- 1:11:46another scenario which is called as
- 1:11:48underfitting so let's say that I have
- 1:11:50something called as
- 1:11:51underfitting now in this underfitting
- 1:11:53what is the scenario the
- 1:11:56model fails to perform it gives bad
- 1:12:00accuracy I say that model always
- 1:12:03remember whenever I talk about bias then
- 1:12:05you can understand that it is something
- 1:12:07related to the training data whenever I
- 1:12:10talk about test data at that point of
- 1:12:12time you talk about variance and that
- 1:12:15specifically whenever you talk about
- 1:12:17variance that basically means we are
- 1:12:18talking about the test data so for an
- 1:12:21overfitting you will basically have low
- 1:12:23bias and high variance low bias with
- 1:12:26respect to the training data and high
- 1:12:29variance with respect to the test data
- 1:12:31now if the model accuracy is bad with
- 1:12:36training data and the model accuracy is
- 1:12:39also bad with test data in this scenario
- 1:12:44we basically say it as underfitting so
- 1:12:47these are the two conditions that are
- 1:12:50with respect to underfitting that
- 1:12:51basically means that both for the
- 1:12:54training data also the model is giving
- 1:12:55bad accuracy and again for the test data
- 1:12:59also it is basically having a bad
- 1:13:01accuracy so in this particular scenario
- 1:13:03we can definitely say two things out of
- 1:13:05underfitting one is high bias and high
- 1:13:10variance so this is the condition with
- 1:13:12respect to underfitting very super
- 1:13:15important let me just explain you once
- 1:13:17again suppose let's consider I have one
- 1:13:21model I have model two this is model one
- 1:13:24this is model one this is model two and
- 1:13:27this is model 3 okay guys so suppose
- 1:13:30let's say that I have my model my
- 1:13:33training accuracy is let's say
- 1:13:3690% And my let's say that my test
- 1:13:39accuracy is 80% now in this particular
- 1:13:42case let's say that my training accuracy
- 1:13:44is
- 1:13:4692% and my test accuracy is 91% and
- 1:13:51let's say my model three is basically
- 1:13:53having training accuracy as
- 1:13:5670% and my test accuracy is 65% so if I
- 1:14:01take this particular case it is
- 1:14:03basically overfitting if I take this
- 1:14:06particular thing this basically becomes
- 1:14:08my generalized model and when I talk
- 1:14:11about this this is my I'll just say that
- 1:14:15okay I'll also put nice color so that uh
- 1:14:17you'll be able to understand this this
- 1:14:19becomes our generalized model and this
- 1:14:22finally becomes our underfitting right
- 1:14:24under under fitting so here is my red
- 1:14:27color I will just say it as underfitting
- 1:14:29what are the main properties of this
- 1:14:31overfitting as I said in this scenario
- 1:14:34since it is performing well with the
- 1:14:36training data so it will be low bias
- 1:14:38High variance in this particular case it
- 1:14:41will be low bias low variance and this
- 1:14:44particular case it will be high bias and
- 1:14:47high variance understand in this
- 1:14:49terminology in this particular way
- 1:14:51you'll be able to understand so why do
- 1:14:53we require always a generalized model
- 1:14:55because whenever our new data will
- 1:14:57definitely come generalized model will
- 1:14:59be able to give us very good output
- 1:15:01let's go back to this particular example
- 1:15:03here you'll be able to see this straight
- 1:15:05line the red line that I have actually
- 1:15:07created is basically overfitting so that
- 1:15:10whenever I probably get the new points
- 1:15:12which is having this real value and the
- 1:15:14predicted points here you'll be able to
- 1:15:16see the difference is quite huge so
- 1:15:18because of this it will definitely be a
- 1:15:20scenario of overfitting where it has low
- 1:15:24bias and high weight
- 1:15:25so again let me go ahead and take this
- 1:15:28example so this was my line which I have
- 1:15:30actually drawn I had two points and when
- 1:15:33I draw this line which was a best fit
- 1:15:36line to which is passing through both
- 1:15:37the points this scenario is basically
- 1:15:40causing a overfitting problem and I've
- 1:15:42also shown you my J of theta 1 will be
- 1:15:45zero in this scenario since it is
- 1:15:47passing exactly and the predicted point
- 1:15:49is also over there now understand one
- 1:15:52thing is that what can can we take out
- 1:15:55from this what assumptions we can take
- 1:15:57out from this definitely if I talk about
- 1:16:00our cost function our cost function here
- 1:16:02is nothing but 1X 2 m summation of I = 1
- 1:16:062 m h Theta of X of i - y of I whole s
- 1:16:13now let's consider that I am going to
- 1:16:15use this H Theta X and I'm going to
- 1:16:17basically write it as y hat okay let's
- 1:16:19focus on this specific point so when I
- 1:16:22take this I'm I'm just going to focus on
- 1:16:24this particular point so here I will
- 1:16:26definitely write it as y hat minus y of
- 1:16:30I whole squ so this is my y y hat of I
- 1:16:35minus y hat y i whole Square so this is
- 1:16:38nothing but the difference between the
- 1:16:40predicted value and the real value okay
- 1:16:42this is what I'm actually trying to get
- 1:16:44now in this scenario if I am adding this
- 1:16:47values obviously I'm going to get the
- 1:16:48value as zero now I have to make sure
- 1:16:52that this value does not come to zero
- 1:16:53because this is still over fitting so
- 1:16:57that is where your Ridge regression will
- 1:16:58come into picture Ridge and lasso will
- 1:17:01come into picture now when I use Ridge
- 1:17:03and lasso suppose if I use Ridge now in
- 1:17:06Ridge what we say this this is also
- 1:17:09called as L2
- 1:17:11regularization now L2 regularization
- 1:17:14what it does is that it basically adds a
- 1:17:17unique
- 1:17:18parameter add a One More Sample value
- 1:17:21which is like Lambda multiplied by slope
- 1:17:25Square now what is this slope whatever
- 1:17:28slope of this particular line it is we
- 1:17:30are just going to square it off now
- 1:17:33suppose if I take my equation which
- 1:17:34looks like this H Theta of X is equal to
- 1:17:39Theta 0 + Theta 1 x now in this
- 1:17:41particular case my Theta 0 was zero so
- 1:17:44my H Theta of X is nothing but Theta 1
- 1:17:47what is Theta 1 this is specifically
- 1:17:49called as slope and I am basically
- 1:17:52taking this Theta 1 I'm actually making
- 1:17:54it as a square Square so always
- 1:17:56understand I don't want to make this as
- 1:17:57zero because if it becomes zero it may
- 1:18:00lead to overfitting condition now what
- 1:18:03will happen if I add this particular
- 1:18:05equation if I add this particular
- 1:18:06equation this will obviously come as
- 1:18:08zero let's consider my Lambda value over
- 1:18:12here my Lambda value is one I'll talk
- 1:18:15about how do you set up Lambda value
- 1:18:17okay let's consider that I'm
- 1:18:18initializing it to one let's say my
- 1:18:21Lambda value is 1 now what I will do is
- 1:18:24that this l Lambda value is 1 Let's
- 1:18:26consider our slope value initially is
- 1:18:28two and because of this two I got this
- 1:18:30best fit line I'm just going to consider
- 1:18:32it so if I do the total sum over here if
- 1:18:35I'm just considering this this value is
- 1:18:37three now the cost function will not
- 1:18:40stop over here because still it has to
- 1:18:42minimize it has to reduce this three
- 1:18:45value so what it will do it will again
- 1:18:47change the Theta 1 value and let's say
- 1:18:49that my Theta van value has changed now
- 1:18:52it got another best fit line which looks
- 1:18:54something like like this this is my next
- 1:18:56best fit line I'll talk about Lambda
- 1:18:57Lambda is a hyper parameter guys what
- 1:19:00exactly is Lambda I'll just talk about
- 1:19:01it now when I basically change this line
- 1:19:04now see why I'm getting this line let's
- 1:19:06consider I have changed my Theta 1 value
- 1:19:08since we need to minimize now when we
- 1:19:11need to minimize what it will do we'll
- 1:19:12again calculate the slope of this
- 1:19:14particular line and then we will try to
- 1:19:16create a new line when we sorry it is
- 1:19:18two two not three just a second guys 0 +
- 1:19:241 multiplied by 2 s which is nothing but
- 1:19:274 so now my cost function will not stop
- 1:19:31over here so we are going to still
- 1:19:33reduce this now in order to reduce this
- 1:19:36again Theta 1 value will get changed and
- 1:19:39then we will get a next best fit line
- 1:19:40for this point now what will happen in
- 1:19:43this scenario once we have this best fit
- 1:19:45line we will definitely get a kind of
- 1:19:47small difference so now if I go ahead
- 1:19:50and consider the new equation my y hat I
- 1:19:54minus y
- 1:19:55i² + Lambda of slope squar this value
- 1:20:00will be a small value now because I have
- 1:20:03some difference and then plus again 1
- 1:20:06multiplied by now understand whether the
- 1:20:09slope will increase in this particular
- 1:20:11case or whether it will decrease in this
- 1:20:13particular case there will be some slope
- 1:20:15value let's say that I have got some
- 1:20:17slope of this particular line in this
- 1:20:19particular scenario again your slope
- 1:20:21will definitely decrease so let's say in
- 1:20:23the case of two initially it was now it
- 1:20:25is basically
- 1:20:271.36 whole squ now this small Value
- 1:20:32Plus 1 + 1.3 squ or let me consider that
- 1:20:37my slope is now one simple value that is
- 1:20:405 so if I get this it is 2.25 2.25 plus
- 1:20:44small value it will be less than three
- 1:20:46only right it will obviously be less
- 1:20:48than three or equal to 3 but understand
- 1:20:50what is happening the value is getting
- 1:20:52reduced from 4 to 3 so this is is the
- 1:20:55importance of Ridge now what will happen
- 1:20:57is that you will try to get a
- 1:20:59generalized model which has low bias and
- 1:21:02low variance instead of this overfitting
- 1:21:05condition you know why specifically we
- 1:21:08are adding Ridge L2 regularization it is
- 1:21:11basically to prevent
- 1:21:14overfitting because here you are not
- 1:21:16stopping here you are trying to reduce
- 1:21:18it unless and until you get a line you
- 1:21:21get a line which will be able to handle
- 1:21:24the which will be able to handle as a uh
- 1:21:27generalized model now here you can see
- 1:21:29now if I have my new points like how I
- 1:21:31drew over here now the distance will be
- 1:21:33less so now you'll be able to see that
- 1:21:36it will be able to create a generalized
- 1:21:38model guys this will be a small value
- 1:21:40only see initially when we have this
- 1:21:42line obviously we have zero if we try to
- 1:21:45slightly move here and there so here
- 1:21:48you'll be able to see that it will just
- 1:21:50a slight movement but what this movement
- 1:21:52is basically specifying it is specifying
- 1:21:55that the slope should not be steep if we
- 1:21:59probably have a steep slope it obviously
- 1:22:02leads to most of the time overfitting
- 1:22:04condition it should not be steep it
- 1:22:06should be very very it should be less
- 1:22:08steeper but it should actually help you
- 1:22:10to create a generalized model so you
- 1:22:13will be seeing that after playing for
- 1:22:14some amount of time this value will not
- 1:22:18reduce after some point of time it'll
- 1:22:19get almost it'll be a minimal value
- 1:22:22it'll be a smaller value and for this
- 1:22:24also you have to specify iterations how
- 1:22:26many times you probably have to train
- 1:22:29them now this iterations is also a
- 1:22:33hyperparameter based on number of
- 1:22:35iterations you will probably see your R
- 1:22:38square or adjusted R square over here so
- 1:22:41this iterations based on the number of
- 1:22:43iterations it will never become zero
- 1:22:45guys understand because zero it is not
- 1:22:48possible if it becomes zero trust me it
- 1:22:50is an overfitting model you cannot get
- 1:22:52that is something zero now what is
- 1:22:55Lambda coming to this Lambda this Lambda
- 1:22:57is a
- 1:22:58hyperparameter this is basically to
- 1:23:01check how fast you want to lessen the
- 1:23:04steepness or how fast you want to make a
- 1:23:06steepness grow higher right and this
- 1:23:08Lambda will also be selected by using
- 1:23:11hyper parameter and this also I'll show
- 1:23:13you today in Practical what do you mean
- 1:23:15by iterations iteration basically means
- 1:23:17how many time I want to change the Theta
- 1:23:181 value how many times you want to
- 1:23:20change the Theta value that is the
- 1:23:22convergence algorithm right
- 1:23:25convergence algorithm over here L2
- 1:23:27regularization or Ridge is basically
- 1:23:30used in such a way that you should never
- 1:23:33overfit why we assume Theta 0 is equal
- 1:23:35to 0 because I'm considering that it
- 1:23:37passes through a origin right origin
- 1:23:40over here Lambda is a hyper
- 1:23:43parameter steep basically means how
- 1:23:46steep the line is if I have this line
- 1:23:49this line is quite steep if I have this
- 1:23:51line This is less steep now if I go to
- 1:23:54the next regularization which is called
- 1:23:56as lasso raso R lasso regression this is
- 1:24:00also called as L1
- 1:24:02regularization now here the formula will
- 1:24:05be changing little bit here you will be
- 1:24:07having y hat of minus of Y whole Square
- 1:24:11here you'll be adding a parameter Lambda
- 1:24:14but understand here you'll not be adding
- 1:24:16slope Square no here you'll be adding
- 1:24:20mode of slope here you'll be adding mode
- 1:24:23of slope and this mode of slope will
- 1:24:26work is that it will actually help you
- 1:24:29to do feature selection now you may be
- 1:24:31thinking how feature selection crash
- 1:24:33let's consider a equation over here
- 1:24:35let's say that I have many many features
- 1:24:37I have many many many features okay so
- 1:24:40my H Theta of X which I'm indicating
- 1:24:42here as y hat let's say that I'm I'm
- 1:24:45writing this equation apart from
- 1:24:47preventing for overfitting it will also
- 1:24:49help you to do feature selection here
- 1:24:51let me just show you over here with an
- 1:24:53example this H Theta of X which I'm
- 1:24:56probably writing as y hat will basically
- 1:24:59be indicated by something over here
- 1:25:01you'll be able to see that it is nothing
- 1:25:03but let's say that I have multiple
- 1:25:05features like this now in this
- 1:25:07particular features obviously there are
- 1:25:09so many coefficients over here so many
- 1:25:11slopes over here now mod of slope will
- 1:25:13be what it will be nothing but mod of X1
- 1:25:16plus X2 plus X3 plus X4 plus X5 like
- 1:25:20this up to xn now in this particular
- 1:25:23case how it is basically helping you to
- 1:25:25sorry not X1 sorry just a second this
- 1:25:29mod of I have taken the data point this
- 1:25:31is not data points this should be your
- 1:25:34mod of theta 0 + Theta 1 + Theta 2 +
- 1:25:38theta 3 + Theta 4 + Theta 5 like this up
- 1:25:42to Theta n so here you'll be able to see
- 1:25:45that this is how I will basically uh
- 1:25:47I'll basically be calculating the slope
- 1:25:50now as we go ahead guys whichever
- 1:25:52features are probably not playing an
- 1:25:55amazing role the Theta value the
- 1:25:57coefficient value the slope value will
- 1:25:59be very very small it is just like that
- 1:26:01entire feature is neglected that entire
- 1:26:04feature is neglected now in this
- 1:26:06particular case we were doing squaring
- 1:26:08because of the squaring that value was
- 1:26:10also increasing but here because of the
- 1:26:11mode that value will not increase
- 1:26:14instead it will be a condition wherein
- 1:26:16we are basically neglecting those
- 1:26:18features that are not at all important
- 1:26:21in this specific problem statement so
- 1:26:23with the help of L1 regularization that
- 1:26:26is lasso you are able to do two
- 1:26:28important things one is preventing
- 1:26:31overfitting and the second case is that
- 1:26:33if you have many features and many of
- 1:26:36the features are not that important okay
- 1:26:39in basically finding out your slope or
- 1:26:42your line or the best fit line in that
- 1:26:44particular case it will also help you to
- 1:26:46perform feature selection so this is the
- 1:26:48importance of the entire what is the
- 1:26:51importance of this this is the
- 1:26:52importance of the uh Ridge and the lasso
- 1:26:56regression that we are doing here I'm
- 1:26:57just going to write L1
- 1:26:59regularization and obviously we have
- 1:27:01discussed about L2 regularization also
- 1:27:04now you have probably understood Lambda
- 1:27:06is one hyperparameter okay which we will
- 1:27:09specifically using okay and based on
- 1:27:12this Lambda this will be found out
- 1:27:13through cross
- 1:27:14validation cross validation is a
- 1:27:16technique wherein we will try to
- 1:27:19probably train our model and try to find
- 1:27:21out the specific things okay what should
- 1:27:24be the exact value and there also we
- 1:27:26play with multiple values in short what
- 1:27:28we are doing we just trying to reduce
- 1:27:29the cost function in such a way that uh
- 1:27:32it will definitely never become zero but
- 1:27:34it will basically reduce based on the
- 1:27:37Lambda and the slope value in most of
- 1:27:39the scenario if you ask me we should
- 1:27:41definitely try both the regularization
- 1:27:44and see that wherever the performance
- 1:27:46Matrix is good we should use that what
- 1:27:48is cross validation basically means I
- 1:27:50will try to use different different
- 1:27:52Lambda value and basically Ally use it
- 1:27:55so in a short let me write it down again
- 1:27:58for Ridge regression which is an L2 Norm
- 1:28:03here I'm simply writing my cost function
- 1:28:05in this particular case will be little
- 1:28:08bit different here I can definitely
- 1:28:10write my cost function as H Theta X of i
- 1:28:14- y of I S Plus Lambda multiplied slope
- 1:28:21Square what is the purpose of this the
- 1:28:24purpose is very simple here we are
- 1:28:27preventing overfitting this was with
- 1:28:29respect to the Ridge Recreation that is
- 1:28:31L2 nor now if I go ahead and discuss
- 1:28:33about the next one which is called as
- 1:28:34lasso regression which is also called as
- 1:28:38L1 regularization in the case of lasso
- 1:28:41regression your cost function will be H
- 1:28:44Theta of X of
- 1:28:47IUS y of
- 1:28:49i² plus Lambda ultied mode of flow so
- 1:28:55here you have this specific thing and
- 1:28:57what is the purpose the purpose are two
- 1:29:00one is prevent overfitting and the
- 1:29:04second one is something called as
- 1:29:05feature selection so these two are the
- 1:29:08outcomes of the entire thing see with
- 1:29:10respect to this lasso right you have
- 1:29:12slopes slopes here you'll be having
- 1:29:15Theta 0 plus Theta 1 plus Theta 2 plus
- 1:29:17theta 3 like this up to Theta n now when
- 1:29:20you'll have this many number of thetas
- 1:29:22when you have many number of features
- 1:29:24and when you have many number of
- 1:29:25features that basically means you'll
- 1:29:26have multiple slopes right those
- 1:29:28features that are not performing well or
- 1:29:30that has no contribution in finding out
- 1:29:32your output that coefficient value will
- 1:29:35be almost nil right it will be very much
- 1:29:37near to zero in short you neglecting
- 1:29:40that value by using modulus you're not
- 1:29:42squaring them up you're not increasing
- 1:29:44those values now I will continue and uh
- 1:29:47probably I will also discuss about the
- 1:29:49assumptions of linear regressions so
- 1:29:52what are the assumptions of linear
- 1:29:54regression in this particular scenario
- 1:29:56so assumption is that number one point
- 1:30:00linear regression if our features are in
- 1:30:03normal or gion
- 1:30:06distribution if our features follows
- 1:30:09this particular distribution it is
- 1:30:11obviously good our model will get
- 1:30:13trained well so there is one concept
- 1:30:16which is called as feature
- 1:30:18transformation now in future
- 1:30:20transformation always understand what
- 1:30:22will happen if a model does not fall
- 1:30:24follow a gan distribution then we apply
- 1:30:26some kind of mathematical equation onto
- 1:30:28the data and try to convert them into
- 1:30:30normal orian distribution the second
- 1:30:33assumption that I would definitely like
- 1:30:34to make is that standard scalar or
- 1:30:37standard digestion standard dig is
- 1:30:40nothing but it is a kind of scaling your
- 1:30:43data by using Z score I hope everybody
- 1:30:46remembers Z score this is what we
- 1:30:48basically apply there your mean is equal
- 1:30:50to zero and standard deviation equal to
- 1:30:521 see guys wherever you have gradient
- 1:30:54descent involved it is good to basically
- 1:30:57do
- 1:30:58standardization because if our initial
- 1:31:01point is a small Point somewhere here
- 1:31:03then to reach the global Minima or
- 1:31:05training will happen quickly otherwise
- 1:31:07what will happen if your values are
- 1:31:09quite huge then your graph may be very
- 1:31:11big and the point can come over any over
- 1:31:13there and the third point is that this
- 1:31:16linear regression works with respect to
- 1:31:19linearity it works if your data is
- 1:31:22linearly separable
- 1:31:24I'll not say linearly separable but this
- 1:31:26linearity will come into picture if your
- 1:31:28data is too much linear it will
- 1:31:30obviously be able to give a very good
- 1:31:31answer like logistic regression also
- 1:31:34which we are going to discuss today this
- 1:31:35also has the same property now you may
- 1:31:38be asking is it compulsory to do
- 1:31:40standardization guys if you want to
- 1:31:42increase the training time of your model
- 1:31:45or if you want to optimize your model I
- 1:31:47would suggest go ahead and do
- 1:31:48standardization now coming to the fourth
- 1:31:50Point here you really need to check
- 1:31:52about multicolinearity
- 1:31:54this is also one kind of check we
- 1:31:56basically do what is multicol
- 1:31:58linearities let's say I have X1 I have
- 1:32:00X2 and this is my output feature I have
- 1:32:03let's say X3 also now let's say that if
- 1:32:05I try to see the colinearity of this two
- 1:32:08feature how how correlated these two
- 1:32:10feature are let's say that these two
- 1:32:12feature are 95% correlated is it is it a
- 1:32:16wise decision to use both the features
- 1:32:18and let's say that let's let's say that
- 1:32:20these two features are 95% correlated
- 1:32:23but it is highly correlated with Y is it
- 1:32:25necessary that we should use both the
- 1:32:27feature in this particular scenario the
- 1:32:29answer should be no we can drop this
- 1:32:32particular feature okay we can drop this
- 1:32:34particular feature any one of the
- 1:32:36feature we can definitely drop it and
- 1:32:38based on that I can just use one single
- 1:32:40feature and basically we do the
- 1:32:42prediction there is also a concept which
- 1:32:44is called as variation inflation factor
- 1:32:46I will try to make a dedicated video
- 1:32:48about this multical is also solved with
- 1:32:51the help of variation inflation Factor
- 1:32:53one more term is there homos orc so that
- 1:32:56kind of terminologies also we use one
- 1:32:58more condition in this but if you almost
- 1:33:00satisfied with this assumptions you will
- 1:33:02definitely be able to outperform in
- 1:33:03linear regression so you have got an
- 1:33:06idea of the assumptions you have also
- 1:33:07got an idea of multiple things okay now
- 1:33:10let's go towards something called as
- 1:33:12logistic regression now logistic
- 1:33:14regression what logistic regression is
- 1:33:16the first type of algorithm that we are
- 1:33:18going to learn in classification let's
- 1:33:20say that in classification I have one
- 1:33:22example you know so suppose I have say
- 1:33:24number of hours study hours and number
- 1:33:28of play hours based on this I want to
- 1:33:31predict whether a child is passing or
- 1:33:33failing suppose these two are my
- 1:33:35features I want to predict whether it is
- 1:33:36pass or fail so here you'll be able to
- 1:33:39see that I have some fixed number of
- 1:33:40categories specifically in this
- 1:33:42particular scenario I have two
- 1:33:43categories binary logistic regression
- 1:33:46works very well with binary
- 1:33:48classification now the uh question comes
- 1:33:50that can we solve multiclass
- 1:33:52classification using logistic the answer
- 1:33:54is simply yes you can definitely do it
- 1:33:57so let's go ahead and let's try to
- 1:33:58discuss about uh logistic regression now
- 1:34:02what is the main purpose of the logistic
- 1:34:03regression first of all let's let's uh
- 1:34:06understand one scenario okay suppose I
- 1:34:08have a feature which basically says um
- 1:34:13number of study hours and this is like 1
- 1:34:172 3 4 5 6 7 and let's say that I have
- 1:34:24pass this point is basically pass and
- 1:34:28this point is basically
- 1:34:29fail so I have this two conditions these
- 1:34:32are my outcomes now what I'll do I will
- 1:34:34just try to make some data points let's
- 1:34:36say that if I study Less Than 3 hours I
- 1:34:40will probably be fail if I study more
- 1:34:43than 3 hours then probably I will pass
- 1:34:47this I'll make it as fail and this I
- 1:34:49will make it as pass so I will be having
- 1:34:52points over here this 1 2 3 let's say
- 1:34:57that this is my training data set now
- 1:35:00the first question says that okay Chris
- 1:35:02fine you have some data over here
- 1:35:05whenever it is less than three you are
- 1:35:06basically the person is failing if it is
- 1:35:08greater than five greater than three it
- 1:35:12is basically showing data points points
- 1:35:14with respect to pass now can't we solve
- 1:35:17this problem first with linear
- 1:35:18regression now with the help of linear
- 1:35:21regression here the first point will be
- 1:35:23that yes I can definitely draw a best
- 1:35:26fit line my best fit line in this
- 1:35:28particular scenario may be something
- 1:35:30like this it may it may look something
- 1:35:32like this so here fail is nothing but
- 1:35:35zero pass is one the middle point is
- 1:35:38basically 0.5 so obviously with the help
- 1:35:41of linear
- 1:35:42regression I'm able to create this best
- 1:35:44fit line and I'll put a scenario that
- 1:35:47whenever the value is less
- 1:35:50than5 whenever the value is less than
- 1:35:520.5 whenever the output is less than5
- 1:35:55let's say that new data point is this
- 1:35:57and based on this I'll try to do the
- 1:35:58prediction I'm actually able to get the
- 1:36:00output over here now when I'm getting
- 1:36:02the output over here this basically is
- 1:36:040.25 now in this particular scenario
- 1:36:06obviously I'm able to say that yes the
- 1:36:08person I'll write a condition over here
- 1:36:11saying that if my H Theta of x value is
- 1:36:15less than 0.5 then my output should be
- 1:36:20zero let's say less than 0.5 I'll say
- 1:36:22not less than or equal to less than5
- 1:36:25then my output will be zero right so in
- 1:36:28this particular case Zero basically
- 1:36:29means fail similarly I'll have a
- 1:36:32scenario where I'll say that when if my
- 1:36:35S of theta of X is greater than or equal
- 1:36:36to 5 then this will basically be one
- 1:36:39which is nothing but pass so this two
- 1:36:41condition I can definitely write over
- 1:36:42here this is my center point so that any
- 1:36:45point that will probably come over here
- 1:36:47let's say that this point is coming over
- 1:36:49here right let's say new data point is
- 1:36:51somewhere coming over here with this red
- 1:36:53point
- 1:36:54now what I'll do I'll basically draw a
- 1:36:56straight line it will come over here I
- 1:36:58will just extend this line
- 1:37:00long I will extend this line over here
- 1:37:04and I will extend this line over here
- 1:37:07and here you can see that based on this
- 1:37:09I'm actually getting this particular
- 1:37:11prediction which is greater than 0.5 so
- 1:37:13I will say that okay the person has
- 1:37:15passed obviously this is fine this is
- 1:37:18obviously working better this is
- 1:37:20obviously working better so what what is
- 1:37:22the problem why we are not using linear
- 1:37:24regression okay in order to solve this
- 1:37:26particular problem why you are
- 1:37:27specifically having logistic regression
- 1:37:29the answer is very much simple guys the
- 1:37:31answer is that whenever let's say that
- 1:37:34if I have an outlier which looks
- 1:37:35something like this suppose I have an
- 1:37:37outlier which comes like this over here
- 1:37:40what is this value let's say that this
- 1:37:41value is nothing but 7 8 9 10 let's say
- 1:37:46that the number of study hours and I'm
- 1:37:48studying for nine it is obviously pass
- 1:37:51now in this particular scenario when I
- 1:37:52have an outlier this entire line will
- 1:37:54change now I will probably get my line
- 1:37:57which looks something like this okay my
- 1:37:59line will basically move something like
- 1:38:01this it will now get moved something
- 1:38:03like this now when it gets moves
- 1:38:04completely like this now for even five
- 1:38:08or even at any point that I am actually
- 1:38:10predicting let's say that at this
- 1:38:11particular point if I try to find out
- 1:38:14it'll be showing less than 0. five so
- 1:38:16here this particular value or answer
- 1:38:19will be wrong right because if we are
- 1:38:21studying more than 5 hours OB viously B
- 1:38:24based on the previous line the person
- 1:38:26had to pass but in this scenario it is
- 1:38:28failing it is coming less than 0.5 but
- 1:38:31the real value for this is basically
- 1:38:33passed so I hope you are understanding
- 1:38:36because of the outlier the entire line
- 1:38:37is getting changed so how do we fix this
- 1:38:40particular problem now in this two
- 1:38:42scenarios are there first of all
- 1:38:44obviously because of just an outlier
- 1:38:46your entire line is getting shifted here
- 1:38:48and there the second point is that over
- 1:38:50here sometimes you're also getting
- 1:38:52greater than one you you're also getting
- 1:38:54less than one suppose if I try to
- 1:38:55calculate for this particular point if I
- 1:38:58project it in behind I'll be getting
- 1:39:00some negative value so we have to squash
- 1:39:02this function if I squash this function
- 1:39:04then it'll become a plain line right how
- 1:39:07do we squash it and for this we use
- 1:39:09something called as sigmoid activation
- 1:39:11function or sigmoid function if somebody
- 1:39:14ask you why don't you use linear
- 1:39:17regession in order to solve this
- 1:39:19classification problem then your answer
- 1:39:21should be very much simple you should
- 1:39:23say this to specific points so we will
- 1:39:26try to go ahead and solve some linear
- 1:39:27regression now with the help of cost
- 1:39:29function everything as such and we'll
- 1:39:31try to understand how the cost function
- 1:39:34will look for logistic regression second
- 1:39:36reason I told you right it is greater
- 1:39:38than zero over here the line is going
- 1:39:40greater than zero right greater than
- 1:39:42zero I have only Z and one and it is
- 1:39:45becoming greater than zero but I have
- 1:39:47already told that our maximum and
- 1:39:49minimum value are 1 and zero so I hope
- 1:39:51you have understood why linear Reg
- 1:39:53cannot be used okay I showed you all the
- 1:39:56scenarios why linear regression should
- 1:39:58not be used now we'll continue and
- 1:40:00probably discuss about the other things
- 1:40:02over here and uh we will now try to
- 1:40:05understand fine what exactly logistic
- 1:40:07regression is all about and how the
- 1:40:09decision boundaries basically created
- 1:40:11now we'll go ahead and discuss about
- 1:40:12that specific thing so let's go ahead
- 1:40:15our values should be always between 0 to
- 1:40:17one over here in this particular case
- 1:40:19because it is a binary classification
- 1:40:21problem only this should be the answer
- 1:40:23so let's go ahead and let's define our
- 1:40:25decision boundary so my decision
- 1:40:26boundary decision boundary in the case
- 1:40:29of logistic regression first of all as
- 1:40:31usual in logistic regression we defined
- 1:40:34our hypothesis okay guys first of all
- 1:40:36let's see if I'm writing my my h of
- 1:40:40theta my H Theta of X as Theta 0 + Theta
- 1:40:451 into x + Theta 2 into X like this X1
- 1:40:49X2 + Theta n into xn
- 1:40:53now in this scenario can I write this
- 1:40:55entire equation as Theta transpose X
- 1:40:59obviously I can definitely write this
- 1:41:01way right and this is what is the
- 1:41:02notation that you will probably seeing
- 1:41:04in many places so with respect to the
- 1:41:06decision boundary of logistic regression
- 1:41:10our Theta see like this we can write I'm
- 1:41:12saying okay but since we have to
- 1:41:14consider two things one is squashing the
- 1:41:17line okay how that squashing will
- 1:41:19basically happen see if I have this if I
- 1:41:22have this line
- 1:41:24we saw in the above right if I have this
- 1:41:26line suppose I have some data points
- 1:41:28over here and I have some data points
- 1:41:30over here if I want to create the best
- 1:41:32fit line how will I create I will
- 1:41:33basically create like this but I have to
- 1:41:35also do two things one is squash over
- 1:41:37here and squash over here right squash
- 1:41:40over here and squash over here now in
- 1:41:42order to squash I'm saying squash squash
- 1:41:46means
- 1:41:48okay now in order to do this I use a
- 1:41:51function which is called as sigmoid
- 1:41:52activation function
- 1:41:54that basically means what happens
- 1:41:56obviously you know this line is
- 1:41:57basically denoted by H Theta of x equal
- 1:42:01to how do you denote this straight line
- 1:42:04let me write it down nicely for you so
- 1:42:06how do you denote this straight line the
- 1:42:08straight line is obviously denoted by
- 1:42:11Theta 0 + Theta 1 * X1 let's say now on
- 1:42:15top of this on top of this I have to
- 1:42:18apply something on top of this value I
- 1:42:21have to apply something so that I can
- 1:42:23make this line straight instead of just
- 1:42:26expanding in this way so my hypothesis
- 1:42:29will basically be now G of G is
- 1:42:32basically a function on Theta 0 and
- 1:42:34Theta 1 * X1 so here I'm trying to
- 1:42:38basically what I'm trying to do I will
- 1:42:40apply a mathematical formula on top of
- 1:42:42this linear regression to squash this
- 1:42:45line now let's go ahead and let's try to
- 1:42:47find out what is this G okay what is
- 1:42:50this G I will say let Z equal to Theta 0
- 1:42:54+ Theta 1 * X I'm just initializing this
- 1:42:58now my H Theta of X is nothing but G of
- 1:43:00Z now we need to understand what is this
- 1:43:03z g of Z and how do we basically specify
- 1:43:06what is the G function so my G function
- 1:43:08is nothing but H Theta of x equal to 1
- 1:43:11by 1 + e ^ of minus Z which in short if
- 1:43:15I try to initialize Zed now it is 1 ^ of
- 1:43:19e ^ of minus Theta 0 + Theta 1 * X so
- 1:43:24this is what is my H Theta of X which is
- 1:43:26my hypothesis and this obviously works
- 1:43:29well because it is being able to squash
- 1:43:32the function so this is basically my
- 1:43:34hypothesis which I am definitely trying
- 1:43:36to use it and this function that you are
- 1:43:39actually able to see is called as
- 1:43:43sigmoid or logistic function now you
- 1:43:47need to understand what does this
- 1:43:48sigmoid function look like in graph in
- 1:43:50graph it looks something like this so
- 1:43:52this this is my Zed value and this is my
- 1:43:56G of Z this is my 05 your sigmoid
- 1:44:00function will have this curve so this is
- 1:44:03your one this is zero your value when
- 1:44:07now from this we can make a lot of
- 1:44:08assumptions what are the assumptions
- 1:44:10that we can basically make your G of Zed
- 1:44:15your G of Zed is greater than or equal
- 1:44:18to
- 1:44:185.5 is obviously greater than or equal
- 1:44:21to 0.5 when your Zed value is greater
- 1:44:24than or equal to zero this is the major
- 1:44:27assumptions that we can basically make
- 1:44:29that is whenever your G of Z is greater
- 1:44:32than your G of Z is greater than or
- 1:44:35equal to 0.5 whenever your Zed is
- 1:44:38greater than or equal to Z so obviously
- 1:44:40whenever your Zed value is greater than
- 1:44:42Z it is greater than 0.5 if your Zed
- 1:44:44value is less than zero what it will
- 1:44:46become it will basically be less than
- 1:44:470.5 so you can write that specific
- 1:44:50condition also you want so this is the
- 1:44:52most important condition
- 1:44:53over here why it is called as logistic
- 1:44:55regression see guys with the help of
- 1:44:56regression you creating this straight
- 1:44:57line and with the help of the concept of
- 1:44:59sigmo you are able to squash it so they
- 1:45:01have probably combined that name and uh
- 1:45:04basically have written in this way will
- 1:45:05squashing of the best fit L line help to
- 1:45:07overcome the outlier issues yes
- 1:45:09obviously it'll be able to help you so
- 1:45:10let's go ahead and let's try to solve
- 1:45:12the problem statement now usually let's
- 1:45:14consider my training set let's consider
- 1:45:17my training set suppose I have some
- 1:45:19training points like this x of 1 comma y
- 1:45:22of 1
- 1:45:24let's say x of 2A y of 2 okay X of 3A y
- 1:45:28of 3 like this I have lot of training
- 1:45:30points and finally X of n comma y of n
- 1:45:33let's say that this is my training data
- 1:45:35so here uh my y y will belong to what
- 1:45:41zero or 1 because I will only have two
- 1:45:43outputs since we are solving a binary
- 1:45:45classification problem here is my
- 1:45:47training set with two outputs and I hope
- 1:45:50everybody knows about J Theta of Z
- 1:45:53it is nothing but 1 + e ^ of minus Z
- 1:45:57here your Z is nothing but Theta 0 +
- 1:45:59Theta 1 * X1 so this is your Theta 0 now
- 1:46:04what we have to do we have to select
- 1:46:06this Theta now in this particular case
- 1:46:08let's consider that my Theta 0 is 0
- 1:46:10because it is passing through the origin
- 1:46:13just for time pass sake suppose my Z is
- 1:46:15Theta 1 into X so now I need to change
- 1:46:19what is my parameter my parameter is
- 1:46:21Theta 1
- 1:46:23I have to change parameter Theta 1 in
- 1:46:25such a way that I get the best fit line
- 1:46:28and along that I apply this sigmoid
- 1:46:30activation function now let's go ahead
- 1:46:33and let's first of all Define our cost
- 1:46:36function because for this we definitely
- 1:46:38require our cost
- 1:46:39function now everything will be same
- 1:46:42obviously you know the cost function of
- 1:46:44linear regression because the first best
- 1:46:47fit line that you are probably creating
- 1:46:48is with the help of linear
- 1:46:50regression now in this particular case
- 1:46:52in the case of linear regression so here
- 1:46:55you can basically write J J of theta 1
- 1:46:57is nothing but 1 by m summation of I = 1
- 1:47:022 m 1X 2 and here you have H Theta of x
- 1:47:08minus y of I I whole Square so this is
- 1:47:13your entire thing of if you remember
- 1:47:15linear regression whatever things we
- 1:47:17have discussed yesterday okay so this is
- 1:47:19the cost function let's consider that
- 1:47:22for linear regression for this is for
- 1:47:24the linear regression now for the
- 1:47:25logistic regression what will happen for
- 1:47:27your logistic regression I will take the
- 1:47:28same cost function H Theta of X now you
- 1:47:31know what is s Theta of X it is nothing
- 1:47:33but 1 + 1 + e ^ of minus Theta 0 + Theta
- 1:47:37sorry Theta 1 multiplied by X right this
- 1:47:40is my with respect to logistic
- 1:47:42regression this is my entire equation
- 1:47:45now similarly I will try to only put
- 1:47:48this H Theta of X let's consider that
- 1:47:51this is my cost function only only my H
- 1:47:53Theta of X is changing in this
- 1:47:55particular case so if I go ahead and
- 1:47:57write my cost function I can basically
- 1:47:59say 1x2 h Theta of X of i - y of
- 1:48:05i² and in this particular scenario what
- 1:48:07is h Theta of X it is nothing but 1 + 1
- 1:48:11+ e ^ minus Theta 1 x so this is what
- 1:48:16this is getting replaced and this is my
- 1:48:18logistic regression cost function I'm
- 1:48:20just considering this cost function part
- 1:48:22this part later on if you replace this
- 1:48:25to this see if I replace this to this
- 1:48:28and if I replace this to this it becomes
- 1:48:30a logistic regression cost function
- 1:48:33intercept I'm considering it as zero
- 1:48:34guys now when I'm replacing this to this
- 1:48:36this to this then it becomes a logistic
- 1:48:39uh regression cost function but there is
- 1:48:41one problem we cannot we cannot use we
- 1:48:45cannot use this cost function there is a
- 1:48:48reason for this because this equation
- 1:48:50that you're seeing 1/ 1 + e^ of minus
- 1:48:54Theta 1 * X this is a non-convex
- 1:48:59function now you may be considering what
- 1:49:01is a non-convex function so let me write
- 1:49:03it down so here this this term this
- 1:49:07terminology right it is a non-convex
- 1:49:09function now what is this non-convex
- 1:49:10function let me show you and let me
- 1:49:12differentiate it with convex function
- 1:49:15okay we'll try to understand what is the
- 1:49:16difference between non-convex function
- 1:49:18and convex function this is related to
- 1:49:21gradient descent very important this is
- 1:49:24related to gradient desent if you
- 1:49:27remember with the help of linear
- 1:49:29regression whatever gradient Dent we are
- 1:49:32actually getting it is a convex function
- 1:49:34like this this is the convex function
- 1:49:38which looks like a parabola curve
- 1:49:40Parabola curve because of this Parabola
- 1:49:42curve whenever we use this linear
- 1:49:44regression cost function specifically
- 1:49:46because here my H Theta of X is what it
- 1:49:48is nothing but Theta 0 + Theta 1 into X
- 1:49:51because of this this equ
- 1:49:53will always give you a parabola curve
- 1:49:56this kind of cost function or convex
- 1:49:59function you can say but here your s
- 1:50:01Theta of X is changing so in the case of
- 1:50:03if I use that cost function you will be
- 1:50:05getting some curves which looks like
- 1:50:07this now what is the problem with this
- 1:50:08curve here you have lot of local Minima
- 1:50:11if local Minima is there you will never
- 1:50:13reach This Global Minima so that is the
- 1:50:15reason we cannot use that c function now
- 1:50:18mathematically you can also go and
- 1:50:20probably search in the Google what is
- 1:50:22the
- 1:50:23what is the graph or what is a convex or
- 1:50:25non-convex function but always remember
- 1:50:27whenever we updates Theta 1 with this
- 1:50:30within this particular equation by
- 1:50:32finding the slope then this way it will
- 1:50:35not be differentiable and here you have
- 1:50:37lot of local Minima and because of this
- 1:50:39local Minima you will never be able to
- 1:50:41reach the global Minima this is your
- 1:50:42Global Minima right in case
- 1:50:45of in case of linear regression you'll
- 1:50:48reach This Global Minima but in this
- 1:50:50case you will never reach never never
- 1:50:52you'll be stuck over here or you may get
- 1:50:54stuck over here you may get stuck over
- 1:50:56here okay so this has a local Minima
- 1:51:00problem so how do we solve this
- 1:51:02understand in local Minima these are my
- 1:51:03points right I have to come over here
- 1:51:05this is my deepest point in this
- 1:51:07particular case I don't have any local
- 1:51:09Minima now in local Minima also you'll
- 1:51:11get slope is equal to Z so that is the
- 1:51:13reason your Theta 1 will never get
- 1:51:14updated so in order to solve this
- 1:51:17problem you can see this diagram we have
- 1:51:19something called as logistic regression
- 1:51:20cost function so I can now write my
- 1:51:23logistic regression cost function in a
- 1:51:25different way so this researcher
- 1:51:27researcher thought of it and basically
- 1:51:30came up with this proposal that the
- 1:51:31logistic cost function should look
- 1:51:33something like this so the entire cost
- 1:51:36function of logistic regression that is
- 1:51:38specifically H Theta of X of I comma y
- 1:51:43this should be written something like
- 1:51:44this and it should be written like this
- 1:51:47see here I'm just going to write cost
- 1:51:49function of J of theta 1 let's say that
- 1:51:51I'm writing J of theta 1 okay so J of
- 1:51:54theta 1 what are the different different
- 1:51:56output that I'll be getting I'll be get
- 1:51:58I'll be getting yal 1 or y equal to 0 So
- 1:52:02based on this two scenarios our cost
- 1:52:04function will look something like this
- 1:52:06minus log of H of theta of X and I know
- 1:52:11I hope you all know what is h Theta of x
- 1:52:13h Theta of X is nothing but 1 + 1 ^ of -
- 1:52:19Theta 1 x so this is what is my H Theta
- 1:52:22of X and whenever Y is Zer then you
- 1:52:25basically have minus log * 1 - H Theta
- 1:52:31of X of I of I okay so this is how you
- 1:52:35basically write your cost function in
- 1:52:36this particular scenario now with the
- 1:52:38help of this cost function it is always
- 1:52:40possible since it is getting log log is
- 1:52:42basically getting used in this scenario
- 1:52:45you'll always get a global Minima that
- 1:52:46is the reason why they have completely
- 1:52:48neglected this cost function and utiliz
- 1:52:51this cost function now what does this
- 1:52:52cost function basically mean two
- 1:52:55scenarios if Y is equal to 1 Let's
- 1:52:58consider this is my cost function
- 1:53:01graph I have H Theta of X and you know
- 1:53:06that H Theta of x value will be ranging
- 1:53:08between 0 to 1 since it is a
- 1:53:10classification problem so it will be
- 1:53:11ranging between 0 to 1 and this is
- 1:53:14basically of J of theta 1 which is my
- 1:53:16cost function so if Y is equal to 1 this
- 1:53:19specific equation will be used and
- 1:53:21whenever this equation is is basically
- 1:53:22used you get a you get a curve see minus
- 1:53:25log s of X of I you get a curve which
- 1:53:29looks something like this okay which
- 1:53:31you'll get a curve which looks like this
- 1:53:33now what does this curve basically
- 1:53:35specify the curve come up with two
- 1:53:37assumptions the cost will be zero if Y
- 1:53:42is = 1 and H Theta of x equal to 1 that
- 1:53:46basically when your s Theta of X is 1
- 1:53:49and the Y is output is one that
- 1:53:51basically means you're going to assign
- 1:53:52over here one right so in this
- 1:53:54particular case you will be seeing that
- 1:53:56your cost function will be zero cost is
- 1:53:59zero so here is my zero it is meeting
- 1:54:01over here if you of x equal to 1 and Y
- 1:54:04is equal to 1 so this is this is again a
- 1:54:06convex function only then the next point
- 1:54:08that you can probably discuss over here
- 1:54:10is with respect to Y is equal to 0 if
- 1:54:13your Y is Z then what kind of curve you
- 1:54:16will be getting you'll get a different
- 1:54:18kind of curve which will look like this
- 1:54:20H Theta of x here your value will be 0
- 1:54:23to one and here you'll be having a curve
- 1:54:26which looks like this so when you
- 1:54:29combine this two you'll be able to see
- 1:54:31that you are able to get a kind of
- 1:54:34gradient descent so this will definitely
- 1:54:36help us to create a cost function so I
- 1:54:38hope everybody is able to understand
- 1:54:40till here with respect to this and this
- 1:54:42will definitely work so finally I can
- 1:54:45also write my cost function in a
- 1:54:47different way the cost function that I
- 1:54:49will probably write over here so this
- 1:54:50will be my J of theta 1
- 1:54:53so I can come up with a cost function
- 1:54:54which looks like this
- 1:54:57cost of H of theta of X of I comma Yus
- 1:55:02log of H Theta of x if Y is equal
- 1:55:091 and then minus
- 1:55:11log 1 - H Theta of x if Y is equal
- 1:55:170 now I can combine this both and
- 1:55:21probably write something like like this
- 1:55:23I can combine this both and I can
- 1:55:25basically write cost of H Theta of X of
- 1:55:27IA Y is equal to - y log H Theta of X of
- 1:55:35I minus log 1 -
- 1:55:40y okay 1 - y log of 1 - H Theta of X so
- 1:55:47this will be my final cost
- 1:55:50function and here also you can see that
- 1:55:53if I
- 1:55:54replace if I replace y with one then
- 1:55:57what will remain only this particular
- 1:55:59value will remain right this value when
- 1:56:01Y is equal to 1 this thing only will
- 1:56:03come you see over here replace y with
- 1:56:05one probably replace y with one and then
- 1:56:08you'll be able to see so here I can now
- 1:56:10write if Y is equal to 1 my cost
- 1:56:14function will Rook something like this
- 1:56:18which is nothing
- 1:56:19but see Y is 1 then what will happen my
- 1:56:22log of H Theta of X of I will come and
- 1:56:26this 1 - 1 is 0 so 0 multili by anything
- 1:56:29will be 0 if Y is equal to 0 then what
- 1:56:32will happen my cost function will be so
- 1:56:36when it is zero this will - y will
- 1:56:38become 0 0 multili by anything is z so
- 1:56:42here you'll be able to see that I am
- 1:56:43I'll be having minus log 1 - H Theta of
- 1:56:48x i so this both the condition has been
- 1:56:50proved by this cost function
- 1:56:52so this is my cost function yes cost
- 1:56:54function and loss function with respect
- 1:56:55to the number of parameters will be
- 1:56:57almost same so finally if I try to write
- 1:57:00J of theta because I have that 1X 2 m
- 1:57:03also right so 1X 2 m also I have so what
- 1:57:06I'm actually going to do here you will
- 1:57:08be able to see that I can write J of
- 1:57:11theta 1 is equal to 1 by 2 m summation
- 1:57:16of IAL 1 to M and then write down the
- 1:57:19entire equation that you have probably
- 1:57:22over here so here you have minus y or I
- 1:57:26I'll just remove this minus and put it
- 1:57:27over here and this will become plus
- 1:57:29sorry y of I
- 1:57:31* log H Theta of X of I 1 - y of i y
- 1:57:41log 1 - H Theta of X of I so this
- 1:57:45becomes my entire first function and
- 1:57:48obviously you know what is h thet of x H
- 1:57:52Theta of X of I is nothing but 1 + 1 e^
- 1:57:56minus Theta 1 * X and finally my
- 1:57:59convergence algorithm I have to repeat
- 1:58:02this to update Theta 1 repeat until this
- 1:58:07updation that is Theta Theta
- 1:58:11J is equal to Theta J minus learning
- 1:58:15rate derivative with respect to Theta J
- 1:58:18and this will be my J of theta 1 this is
- 1:58:21my repeat until conversion so this is my
- 1:58:24cost function this is my repeat
- 1:58:27algorithm and here I will be updating my
- 1:58:30entire Theta
- 1:58:321 and this solves your problem with
- 1:58:35respect to logistic regression simple
- 1:58:37simple questions may come like how it is
- 1:58:39different from linear regression how it
- 1:58:41is not different from linear regression
- 1:58:44can we say log likelihood a topic from
- 1:58:46probabilistic yes this is uh this is log
- 1:58:50likelihood if now I will discuss about
- 1:58:54performance metrics and this is specific
- 1:58:56to classification problem and binary
- 1:58:59classification I'm talking let's
- 1:59:02consider let's consider I have a data
- 1:59:04set which has X1 X2 and this is y and
- 1:59:09obviously in logistic uh classification
- 1:59:11you have outputs like 0 1 0 1 1 0 1 and
- 1:59:17your y hat y hat is basically the output
- 1:59:20of the predicted model now in this
- 1:59:22particular scenario my y hat will
- 1:59:24probably be 1 1 0 uh 1 1 1 Z so in this
- 1:59:31particular scenario this is my predicted
- 1:59:34output and this is my actual output so
- 1:59:39can we come to some kind of conclusions
- 1:59:41wherein probably we will be able to
- 1:59:44identify what may be the accuracy of
- 1:59:48this specific model with respect to this
- 1:59:49many data points because confusion
- 1:59:52Matrix is all dealt with this is called
- 1:59:54as we will first of all have to create a
- 1:59:56confusion Matrix now for a binary
- 1:59:59classification problem the confusion
- 2:00:01Matrix will look like this so here you
- 2:00:03have 1 0 1 0 Let's say that this is
- 2:00:06prediction let's say that these are my
- 2:00:08actual value and these are my prediction
- 2:00:10value okay these both are prediction
- 2:00:12value these are my output value when my
- 2:00:15actual value is zero my predicted value
- 2:00:17is one does this what does this mean
- 2:00:21wrong prediction right so when my actual
- 2:00:23value is zero my predicted value is 1 so
- 2:00:26here my count will increase to one let's
- 2:00:28go to the second scenario when the
- 2:00:30actual value is one and my predicted
- 2:00:33value is one that basically means one
- 2:00:35and one so here I'm going to increase my
- 2:00:37count similarly when my actual value is
- 2:00:40zero my predicted value is zero so that
- 2:00:42basically mean when my actual value is z
- 2:00:43my predicted value is zero I'm going to
- 2:00:45increase the count by one if I go over
- 2:00:47here 1 one again it is so instead of
- 2:00:50writing one now this will become two I'm
- 2:00:52going to increase the count similarly
- 2:00:54I'll go over here one more one is there
- 2:00:56so I'm going to increase the count three
- 2:00:58then I have 01 01 basically means when
- 2:01:00my actual value is zero I'm actually
- 2:01:02getting it as one so I'm also going to
- 2:01:04increase this particular value as two
- 2:01:07and then finally I have 1 and zero where
- 2:01:09I'm going to increase like this now what
- 2:01:11does this basically mean now what does
- 2:01:13this basically mean see with respect to
- 2:01:16this kind of predictions whenever we are
- 2:01:17discussing this basically basically says
- 2:01:20so this is my actual values and I have Z
- 2:01:221 and zero and this is my predicted
- 2:01:24values I also have 1 and zero this value
- 2:01:27when one and one are there this is
- 2:01:29called as true positive this value when
- 2:01:310 and Zer are there this is called as
- 2:01:33false negative whenever your actual
- 2:01:35value is zero and you have predicted one
- 2:01:37this becomes false positive and whenever
- 2:01:40your actual value is one you have
- 2:01:41predicted zero this becomes false
- 2:01:43negative now coming to this I really
- 2:01:45need to find out the accuracy of this
- 2:01:47model now if I really want to find out
- 2:01:51and this is what is called as confusion
- 2:01:52Matrix now in this confusion Matrix if I
- 2:01:55really want to find out the accuracy the
- 2:01:57accuracy of this model it is very much
- 2:01:59simple this middle elements that you are
- 2:02:01able to see will basically give us the
- 2:02:03right output so this and this if I add
- 2:02:07it up it will give us the right output
- 2:02:10so here I'm going to get TP + TN divided
- 2:02:13by TP + FP + FN + TN so once I calculate
- 2:02:21this so I have 3 + 1
- 2:02:23/ 3 + 2 + 1 + 1 so this is nothing but 4
- 2:02:29by 7 what is 4 by
- 2:02:32757 so am I getting 57 percentage
- 2:02:35accuracy so I'm actually getting 57%
- 2:02:38accuracy over here with respect to the
- 2:02:39accuracy so this is how we basically
- 2:02:42calculate with respect to basic accuracy
- 2:02:45with the help of uh the confusion Matrix
- 2:02:48okay so this is specifically called as
- 2:02:49confusion Matrix now there are some more
- 2:02:52things that you really need to specify
- 2:02:54always remember our model aim should be
- 2:02:56that we should try to reduce false
- 2:02:57positive and false negative now let's
- 2:03:00say that I want to discuss about two
- 2:03:02topics what one is suppose in our data
- 2:03:04set I have zeros and one category let's
- 2:03:07say in my output if I say Zer are 900
- 2:03:11and ones are 100 this becomes an
- 2:03:13imbalanced data very clear right so this
- 2:03:15become an imbalanced data set it is a
- 2:03:18biased data suppose if I say zeros are
- 2:03:21probab
- 2:03:22600 and ones are probably 400 in this
- 2:03:25particular scenario I will say that this
- 2:03:27is the balance data because yes you have
- 2:03:29100 less but it's okay the it may not
- 2:03:32impact many of the algorithm now see
- 2:03:34guys most of the algorithm that we will
- 2:03:36be probably discussing imbalanced if we
- 2:03:38have an imbalanced data set it will
- 2:03:40obviously affect the algorithms let me
- 2:03:42talk about this let's say that I have
- 2:03:44number of zeros as 900 and number of
- 2:03:46ones is 100 now let's say that my model
- 2:03:49I have created which will directly
- 2:03:51predict
- 2:03:52zero it'll I'll just say that all my
- 2:03:55inputs that it is probably getting with
- 2:03:57respect to this training data it'll just
- 2:03:59output zero now in this particular
- 2:04:01scenario what will be my accuracy my
- 2:04:03accuracy will be 900 divid by 1,000
- 2:04:05right so this is nothing but 90% so is
- 2:04:09this a good
- 2:04:10accuracy obviously it is a good accuracy
- 2:04:12but this is a biased data if my model is
- 2:04:15basically just outputting 00000000 0 if
- 2:04:19it is outputting 00 00 0 obviously most
- 2:04:22of the answer will be zeros but this
- 2:04:24will be a scenario like you know where
- 2:04:27it is just outputting one thing then
- 2:04:28also it is able to get 90% accuracy so
- 2:04:31you should only not be dependent on
- 2:04:33accuracy so there are lot of
- 2:04:35terminologies that we will basically use
- 2:04:37one terminology that we specifically use
- 2:04:40is something called as Precision then
- 2:04:42we'll also use recall what is precision
- 2:04:45what is recall I'll write the formula
- 2:04:46over here in Precision what do we need
- 2:04:48to focus and then finally we will
- 2:04:50discuss about f score so we have to use
- 2:04:53different kind of parametrics of sorry
- 2:04:55different kind of formulas whenever you
- 2:04:58have an imbalanced data set you can also
- 2:04:59do oversampling but again understand in
- 2:05:02most of the scenarios in some of the
- 2:05:04scenarios oversampling may work but we
- 2:05:06have to focus on the type of performance
- 2:05:08metrics that we are focusing on right
- 2:05:10now I'll not say F1 score I'll say F
- 2:05:11score the reason why I'm saying I'll
- 2:05:13just let you know so let's talk about
- 2:05:15recall recall formula is basically given
- 2:05:17by true positive divided by true
- 2:05:20positive plus false negative
- 2:05:22Precision is given by true positive
- 2:05:23divided by true positive plus false
- 2:05:27positive and then I will probably
- 2:05:29discuss about F sore also or we
- 2:05:31basically say fbaa also now I'll just
- 2:05:34draw this confusion Matrix again okay
- 2:05:36which is having true positive true
- 2:05:37negative so let me draw it over here so
- 2:05:40this is my ones and zeros these are my
- 2:05:42actual values and these are my predicted
- 2:05:44values I have true positive I have true
- 2:05:47negative false positive and false
- 2:05:49negative now in this particular scenario
- 2:05:50when I'm actually discussing understand
- 2:05:53what is recall and what focus it is
- 2:05:54basically given on so here whenever I
- 2:05:57talk about recall recall basically says
- 2:05:59that TP TP divided by TP plus FN so I'm
- 2:06:04actually focusing on this so what does
- 2:06:06this basically say true uh recall out of
- 2:06:10all the actual true positives how many
- 2:06:13have been predicted correctly that is
- 2:06:15basically mentioned by TP out of all the
- 2:06:18positive values how many of them have
- 2:06:20predicted as positive so this is what it
- 2:06:22is basically saying and this scenario is
- 2:06:24called as recall in this the false
- 2:06:27negative is basically given more
- 2:06:28priority and our focus should be that we
- 2:06:31should try to reduce false positive
- 2:06:33false negative sorry we should try to
- 2:06:35reduce this now let's go ahead and let's
- 2:06:37discuss about Precision in Precision
- 2:06:39what we are doing we are basically
- 2:06:41taking out of all the predicted values
- 2:06:44out of all the predicted positive values
- 2:06:47how many of them are actual true or
- 2:06:50positive okay this is what Precision
- 2:06:52basically means now suppose if I
- 2:06:54consider spam classification suppose
- 2:06:56this is my task tell me in this
- 2:06:57particular case should we use Precision
- 2:07:00or recall and one more use case I'm
- 2:07:02saying that whether the person has
- 2:07:05cancer or not in which case we have to
- 2:07:08support recall and in which case we have
- 2:07:10to go ahead with Precision has cancer or
- 2:07:13not in spam what is important okay guys
- 2:07:16the recall is also called as true
- 2:07:18positive rate I can also say recall as
- 2:07:20sensitivity so if I go with Spam
- 2:07:22classification it should definitely go
- 2:07:24with Precision why it should go with
- 2:07:26Precision if I probably get a Spam ma
- 2:07:28the main aim should be that whenever I
- 2:07:30get a Spam Mill it should be identified
- 2:07:31as spam okay in that specific scenario
- 2:07:34my positive false positive we should try
- 2:07:37to reduce and in this scenario my false
- 2:07:39pository talks about the spam
- 2:07:41classification a lot in a better way in
- 2:07:43the case of cancer I should definitely
- 2:07:46use recall let's let's focus on the
- 2:07:48recall formula tp/ by TP plus FN if a
- 2:07:52person has a cancer see one actually he
- 2:07:55has a cancer it should be predicted as
- 2:07:57one otherwise if we have FN it is
- 2:07:59basically predicting it does not have a
- 2:08:01cancer that is really a big situation in
- 2:08:04this case if a person does not have a
- 2:08:07Cancer and if he's predict if the model
- 2:08:09predicts okay fine he has a cancer he
- 2:08:11may go and further do the test and then
- 2:08:13he'll come to know whether he has a
- 2:08:14cancer or not but this scenario is very
- 2:08:16dangerous if a person has a cancer but
- 2:08:19he is being indicated that he does not
- 2:08:20have that cancer
- 2:08:22so here false negative is given more
- 2:08:24priority over here in the case of spam
- 2:08:26classification false positive is given
- 2:08:28more priority so this is something
- 2:08:30important over here and you really need
- 2:08:31to understand with respect to different
- 2:08:33different problem statement let me give
- 2:08:35you one more example tomorrow the stock
- 2:08:37market is going to crash in this what we
- 2:08:40need to focus on should we focus on
- 2:08:41Precision or should we focus on recall
- 2:08:44now here two things are there who is
- 2:08:46solving what kind of problem see many
- 2:08:48people will say recall or Precision but
- 2:08:50here two things are there on whose point
- 2:08:52of view you are creating this model are
- 2:08:55you creating this model for the industry
- 2:08:57or are you creating this model for the
- 2:08:59people for the people he should
- 2:09:01definitely get identified that okay in
- 2:09:04this particular scenario you need to
- 2:09:06sell your stock because tomorrow stock
- 2:09:07market is going to crash but for
- 2:09:09companies this is very bad okay I hope
- 2:09:11everybody is able to understand for
- 2:09:13companies it is very very bad so in this
- 2:09:15particular case sometime we need to
- 2:09:17focus both on false positive and false
- 2:09:19negative and again I'm telling you for
- 2:09:22which problem statement you are solving
- 2:09:23that indicates if you are solving for
- 2:09:25people then they should be able to get
- 2:09:27the notification saying that it is going
- 2:09:29to crash if you're probably uh doing it
- 2:09:32for companies at that time your
- 2:09:34Precision recall may change but if I
- 2:09:36consider for both the scenarios at that
- 2:09:39point of time I will definitely use
- 2:09:40something called as F score F score or
- 2:09:42I'll also say it as F beta now how is
- 2:09:45fbaa Formula given as I will talk about
- 2:09:48it and here in the F score you have
- 2:09:50three different formulas the first
- 2:09:51Formula I will say basically as when
- 2:09:53your beta value is 1 okay first of all
- 2:09:57I'll just give a generic definition of f
- 2:09:59s or F beta here you are basically going
- 2:10:01to consider 1 + beta squ Precision
- 2:10:05multiplied by recall divided beta Square
- 2:10:09* Precision plus recall whenever your
- 2:10:14both false positive and false negative
- 2:10:16are important we select beta as one so
- 2:10:19if I select beta as 1 it becomes 1 + 4
- 2:10:22Precision multiplied by recall then you
- 2:10:25have Precision plus recall so here sorry
- 2:10:281 + 1 so this becomes 2 multiplied by
- 2:10:31Precision into recall divided by
- 2:10:34Precision plus recall so here you have
- 2:10:37this is basically called as harmonic
- 2:10:39mean harmonic mean probably you have
- 2:10:41seen this kind of equation where you
- 2:10:42have written 2x y / x + y same type you
- 2:10:46are able to see this this is called as
- 2:10:48harmonic mean here the focus is on both
- 2:10:51false positive and false negative let's
- 2:10:53say that your false positive is more
- 2:10:56important than false negative at that
- 2:10:58point of time you will try to decrease
- 2:11:01or you will try to decrease your beta
- 2:11:03value let's say that I'm decreasing my
- 2:11:05Beta value to 0.5 then what will happen
- 2:11:071 +5 whole
- 2:11:09s and then you have P * R Precision
- 2:11:13recall and here also you have 25 p + r
- 2:11:17now in this particular scenario I'm
- 2:11:19decreasing my Beta decreasing the beta
- 2:11:21basically means that you are providing
- 2:11:23more importance to false positive than
- 2:11:25false negative and finally you'll be
- 2:11:27able to see that if I consider beta
- 2:11:30value as let me just say my notes if I
- 2:11:34consider beta value as two that
- 2:11:37basically means you are giving more
- 2:11:38importance to false negative than false
- 2:11:40positive so with this specific case you
- 2:11:42can come up to a conclusion what value
- 2:11:44you basically want to use now whenever I
- 2:11:46use beta is equal to 1 it becomes fub1
- 2:11:49score if I use beta as .5 then this
- 2:11:52basically becomes f.5 score and this
- 2:11:56becomes your F2 score So based on which
- 2:12:00is important okay which is important
- 2:12:03whether your Precision or false positive
- 2:12:05or false negative is important you can
- 2:12:06consider those things F score will have
- 2:12:09different values if you're using beta is
- 2:12:11equal to 1 that basically means you are
- 2:12:13giving importance to both precision and
- 2:12:16recall if your false positive is more
- 2:12:18important then at that point of time you
- 2:12:20reduce beta value if false negative is
- 2:12:23greater than false bet uh false positive
- 2:12:25then your beta value is
- 2:12:26increasing beta is a deciding parameter
- 2:12:29to decide your F1 score or F2 score or F
- 2:12:32Point score now first thing first what
- 2:12:34is the agenda of today's session first
- 2:12:36of all we will complete practicals for
- 2:12:39all the algorithms that we have
- 2:12:41discussed these all algorithms that we
- 2:12:43have discussed we will cover the
- 2:12:45practicals probably we will be doing
- 2:12:47hyper parameter tuning everything the
- 2:12:49second thing and again here we are going
- 2:12:51to take just simple examples so yes uh
- 2:12:54so today's session I said practicals
- 2:12:56with simple examples where I'll probably
- 2:12:59discuss about all the hyper parameter
- 2:13:01tuning then the second one the second
- 2:13:04algorithm that I'm going to discuss
- 2:13:05about is something called as n bias this
- 2:13:09is a classification algorithm so we are
- 2:13:10going to understand the intuition and
- 2:13:13the third one that we are going to
- 2:13:15probably discusses KNN algorithm so KNN
- 2:13:19algorithms is definitely there
- 2:13:21so this our today's plan I know I've
- 2:13:23written very less but this much maths
- 2:13:26and involved in na bias right we'll
- 2:13:29understand the probability theorem again
- 2:13:30over there there is something called as
- 2:13:32bias theorem we'll try to understand and
- 2:13:35then we'll try to solve a problem on
- 2:13:36that so let's proceed and let's enjoy
- 2:13:39today's session how do we enjoy first of
- 2:13:42all we enjoy by creating a practical
- 2:13:44problem so I am actually opening a
- 2:13:47notebook file in front of you so here uh
- 2:13:50we will try to Sol solve it with the
- 2:13:51help of linear regression Ridge lasso
- 2:13:55and try to solve some problems let's see
- 2:13:58how much we will be able to solve it but
- 2:14:00again the aim is that we learn in a
- 2:14:02better way okay uh so that everybody
- 2:14:06understands some basic basic things okay
- 2:14:08so first of all as usual uh everybody
- 2:14:11open your jupyter notebook file the
- 2:14:13first algorithm that I'm going to
- 2:14:14discuss about is something called as SK
- 2:14:16learn linear regression so everybody I
- 2:14:19hope everybody knows about this SK learn
- 2:14:21let's see what all things are basically
- 2:14:23there in this we will be using fit
- 2:14:25intercept everything as such but here
- 2:14:28the main aim is to find out the
- 2:14:29coefficients which is basically
- 2:14:31indicated by Theta 0 Theta 1 and all the
- 2:14:34first thing we'll start with linear
- 2:14:38regression and then we will go ahead and
- 2:14:40discuss with r and lassor I'm just going
- 2:14:42to make this as
- 2:14:44markdown how many different libraries of
- 2:14:46for linear regression you can do with
- 2:14:48stats you can do with skyi you can do
- 2:14:49with many things okay so first thing
- 2:14:52first let's first of all we require a
- 2:14:53data set so for the data set what we are
- 2:14:56going to do is that we are going to
- 2:14:58basically take up some smaller smaller
- 2:15:01data just let me do this so for this uh
- 2:15:05we are going to take the house pricing
- 2:15:07data set so we are going to solve house
- 2:15:10pricing data set problem a simple data
- 2:15:13set which is already present in SK learn
- 2:15:16only now in order to import the data set
- 2:15:18I will write a line of code which is
- 2:15:19like from SK learn dot data sets data
- 2:15:24sets
- 2:15:25import load uncore Boston so we have
- 2:15:29some Boston house pricing data set so
- 2:15:31I'm just going to execute this I'm also
- 2:15:33going to make a lot of Sals so that I
- 2:15:35don't have to again go ahead and create
- 2:15:37all the sales again some basic libraries
- 2:15:39that I probably want is pro import numai
- 2:15:43as
- 2:15:44NP
- 2:15:45import pandas
- 2:15:48SPD okay import cbon as
- 2:15:52SNS and then I will also import Matt
- 2:15:56Matt plot lib do p plot as PLT and then
- 2:16:02percentile matplot lib matlot lib do
- 2:16:07inline and I will try to execute this
- 2:16:09see this my typing speed has become a
- 2:16:11little bit faster by writing by
- 2:16:12executing this queries again and again
- 2:16:15and uh let's go ahead uh so I have
- 2:16:18imported all the necessary libraries
- 2:16:19that is required which which will be
- 2:16:21more than sufficient for you all to
- 2:16:23start with now in order to load this
- 2:16:25particular data set I will just use this
- 2:16:27Library called as load uncore Boston and
- 2:16:30I'm going to just initialize this so if
- 2:16:32you press shift tab you will be able to
- 2:16:34see that return load and return the
- 2:16:37Boston house prices data set it is a
- 2:16:39regression problem it is saying and then
- 2:16:41probably I'm just going to execute it
- 2:16:43now once I execute it I will go and
- 2:16:45probably see the type of DF so it is
- 2:16:48basically saying skarn dos. bunch now if
- 2:16:51I go and probably execute DF you'll be
- 2:16:53able to see that this will be in the
- 2:16:55form of key value pairs okay like Target
- 2:16:57is here data is here okay so data is
- 2:17:01here Target is here and probably you'll
- 2:17:03be able to find out feature names is
- 2:17:04here so we definitely require feature
- 2:17:06names we require our Target value and
- 2:17:09our data value so we really need to
- 2:17:11combine this specific thing in a proper
- 2:17:14way in the form of a data frame so that
- 2:17:16you will be able to see so what I'm
- 2:17:18actually going to do over here I'm just
- 2:17:19going to say PD do data frame I'll
- 2:17:22convert this entirely into a data frame
- 2:17:24and I will say DF do data see this is a
- 2:17:27key value pair right so DF do data is
- 2:17:29basically giving me all the features
- 2:17:31value so if I write DF do data and just
- 2:17:35execute it you'll be able to see that I
- 2:17:36you will be able to get my entire data
- 2:17:39set in this way my entire data set in
- 2:17:41this way this is my feature one feature
- 2:17:43two feature three feature 4 this feature
- 2:17:4512 I have 12 features over here and
- 2:17:47based on that I have that specific value
- 2:17:50now the next thing thing that I'm going
- 2:17:51to do probably I should also be able to
- 2:17:53add the target feature name over here so
- 2:17:55what I will do I will just convert this
- 2:17:57into DF and then I will also say DF do
- 2:18:02columns and I'll set it to DF do Target
- 2:18:05okay and let me change this to data set
- 2:18:08so I'm going to change this to data set
- 2:18:10and I'm going to say data set. columns
- 2:18:12is equal to DF do Target so if I execute
- 2:18:15this and now if I probably
- 2:18:18print my data set do head you will be
- 2:18:22able to see this specific thing okay it
- 2:18:24is an error let's see expected axis has
- 2:18:2713 element new values has
- 2:18:30506 so Target okay I should not use
- 2:18:33Target over here instead I had a column
- 2:18:36which is called as features feature
- 2:18:38names like if I go and probably see
- 2:18:41DF DF over here you'll be able to see
- 2:18:45there is one thing which is called as
- 2:18:46feature names so I'm going to use DF do
- 2:18:48feature names over here so here it is DF
- 2:18:52do feature names I'm just going to paste
- 2:18:55it over here and now if I go and write
- 2:18:57here you can see print DF data set. head
- 2:19:00if I go and execute without print you'll
- 2:19:02be able to see my entire data set so
- 2:19:04these are my features with respect to
- 2:19:06different different things and this is
- 2:19:09basically a house pricing data set so
- 2:19:10initially I have this features CRM ZN
- 2:19:13indust CH nox RM age distance radius tax
- 2:19:18PT ratio b l stack that so I have my
- 2:19:22entire data set over here the same data
- 2:19:24set I have basically put it over here
- 2:19:26now here also you'll be able to see what
- 2:19:28all this feature basically means this is
- 2:19:30showing wasted weighted distance to five
- 2:19:31do uh Five Boston employment center rad
- 2:19:34basically means index of accessibility
- 2:19:36to radial Highway tax basically means
- 2:19:39full value property tax rate this much
- 2:19:41PT rate basically means pupil teacher
- 2:19:44ratio I don't know what the hell it
- 2:19:45means but it's fine we have some kind of
- 2:19:47data over here properly in front of you
- 2:19:51so these are my independent features
- 2:19:53what are these these all are my
- 2:19:54independent features if you want the
- 2:19:56features detail here you can see it
- 2:19:59right everything what is CRM this
- 2:20:01basically means per capita crime rate by
- 2:20:03town which is important ZN it is
- 2:20:06proportional of residential land zone
- 2:20:08for Lots over 25,000 Square ft so this
- 2:20:12is my DF I did not do much I'm just
- 2:20:14using data frame DF do data column
- 2:20:17features name I'm getting this value
- 2:20:18very much simple now let's go a little
- 2:20:21bit slowly so that many people will be
- 2:20:23able to also understand now this is my
- 2:20:25data set. head now the thing is that I
- 2:20:29obviously have taken all these
- 2:20:31particular values but this is my
- 2:20:32independent feature I still have my
- 2:20:35dependent feature so what I'm actually
- 2:20:37going to do I will create a new feature
- 2:20:40which is like data set of price I'll
- 2:20:42create my feature name price price of
- 2:20:44the house and what I will assign this
- 2:20:46particular value this value will be
- 2:20:48assigned with this target this target
- 2:20:50value this target value is basically the
- 2:20:53sale the price of the houses right it is
- 2:20:56again in the form of array so I'm going
- 2:20:58to take this and put it as a dependent
- 2:21:00feature so here you'll be able to see
- 2:21:02that my price will be my dependent
- 2:21:04feature so here I'll basically write DF
- 2:21:06do Target so once I execute it and now
- 2:21:09if I probably go and see my data set do
- 2:21:12head you'll be able to see features over
- 2:21:15here and one more feature is getting
- 2:21:17added that is price now this price may
- 2:21:20be the units may be in
- 2:21:22millions somewhere Target should be here
- 2:21:24or there it should be probably in
- 2:21:27millions
- 2:21:28or I cannot see it but it should be
- 2:21:31somewhere here it should have definitely
- 2:21:33said that it is probably in millions or
- 2:21:36okay but that is not a problem I think
- 2:21:37but mostly it'll be in millions
- 2:21:39somewhere I think it should be
- 2:21:42here okay I cannot see it but probably
- 2:21:45if I put more time I'll be able to
- 2:21:47understand it okay so over here what is
- 2:21:49the thing main thing this all are my
- 2:21:51independent features and this is my
- 2:21:53dependent feature right so if I'm trying
- 2:21:55to solve linear regression I have to
- 2:21:57divide my independent and dependent
- 2:21:58features properly now let's go to the
- 2:22:01next step that
- 2:22:03is
- 2:22:05dividing the data
- 2:22:07set dividing the oh my God dividing the
- 2:22:12data
- 2:22:14set
- 2:22:17into
- 2:22:19train into first of all I'll try try to
- 2:22:22divide into
- 2:22:24independent and dependent
- 2:22:27features so I want my entire features
- 2:22:30data set divided into independent and
- 2:22:31dependent features X I will be using as
- 2:22:34my independent featur so I will write
- 2:22:35data set dot I will use an iock which is
- 2:22:39present in data frames and understand
- 2:22:41from which feature to which feature I
- 2:22:42will be taking as my independent feature
- 2:22:44to this feature till lat so the best way
- 2:22:48that basically means that I just need to
- 2:22:49skip the last feature in order to skip
- 2:22:52the last feature what I'm actually going
- 2:22:54to do from all the columns I will just
- 2:22:57skip the last column so this is how you
- 2:22:59basically do an indexing with respect to
- 2:23:02just skipping the last feature and this
- 2:23:05will basically be my independent
- 2:23:06features and here I will basically say Y
- 2:23:08is equal to data set do iock and here I
- 2:23:11just want the last feature so I will
- 2:23:14write colon all the records I want and
- 2:23:18see the first term that we are probably
- 2:23:20WR writing over here this basically
- 2:23:22specifies with respect to records here
- 2:23:24this specifies with respect to columns
- 2:23:26from all the columns I'm taking the last
- 2:23:27column here I will just take the last
- 2:23:29column and this will basically be my
- 2:23:32dependent features dependent features so
- 2:23:35here I have basically executed now if
- 2:23:37you can go and probably see x. head here
- 2:23:40you'll be able to find all my
- 2:23:41independent features in y do head you'll
- 2:23:43be able to find the dependent feature
- 2:23:45now let's go to the first algorithm that
- 2:23:47is called as linear regression
- 2:23:51always remember whenever I definitely
- 2:23:53start with linear regression I'll
- 2:23:55definitely not go directly with linear
- 2:23:56regression instead what I will do is
- 2:23:59that I'll try to go with Ridge
- 2:24:00regression and uh lasso regression
- 2:24:02because there you are lot of options
- 2:24:04with respect to hyper pment T but I'll
- 2:24:06just show you how linear regression is
- 2:24:08done so basically you really really need
- 2:24:11to use a lot of libraries okay over here
- 2:24:13and based on this libraries this
- 2:24:15libraries will try to install okay and
- 2:24:18what are these libraries these are
- 2:24:19basically the linear regression Library
- 2:24:21so here I'm basically going to use two
- 2:24:23specific thing one is linear regression
- 2:24:25Library so I will just use from SK learn
- 2:24:28do linear uncore model import linear
- 2:24:32regression do you need to remember this
- 2:24:35the answer is no because I also do the
- 2:24:37Google and I try to find out where in
- 2:24:39escal and it is present okay so here is
- 2:24:42my linear regression so I will try to
- 2:24:44initialize linear reg is equal to
- 2:24:47initialize with linear regression and
- 2:24:49then here what I'm actually going to do
- 2:24:51I'm going to basically apply something
- 2:24:53called as cross validation cross
- 2:24:55validation is very much important
- 2:24:57because in Cross validation we divide
- 2:24:59out train and test data in such a way
- 2:25:01that every combination of the train and
- 2:25:04test data is basically taken by care is
- 2:25:07taken by the model and whoever accuracy
- 2:25:09is better that all entire thing is
- 2:25:11basically combined so here what I'm
- 2:25:13going to do I'm going to say mean square
- 2:25:14error is equal to here I will import one
- 2:25:17more Library let's say from SK learn
- 2:25:20dot model selection I'm going to import
- 2:25:25cross Val
- 2:25:26score so cross Val score cross
- 2:25:29validation score basically means it is
- 2:25:31going to do a lot of train and test
- 2:25:32split it's something like this one
- 2:25:34example I will show it to you here only
- 2:25:37so what does cross validation basically
- 2:25:39do okay so in Cross validation what
- 2:25:42happens what you do suppose this is your
- 2:25:44entire data
- 2:25:46set suppose this is 100 records if you
- 2:25:48do five cross validation then in the
- 2:25:51first this will be your test data and
- 2:25:53remaining all will be your training data
- 2:25:55if in the second cross validation this
- 2:25:58will be your test data and remaining all
- 2:25:59will be your test uh training data like
- 2:26:01this five times you'll be doing cross
- 2:26:03validation by taking the different
- 2:26:05combination of train and test but I'm
- 2:26:07not going to discuss much about it in
- 2:26:09the future if you want a separate
- 2:26:10session I will include that in one of
- 2:26:11the session itself so this was uh
- 2:26:13basically the plan with respect to cross
- 2:26:15validation or cross Val score so here
- 2:26:17I'm going to basically take cross
- 2:26:20Val
- 2:26:21score and here the first parameter that
- 2:26:24I give is my model so linear regression
- 2:26:27is my model and here I will take X and Y
- 2:26:30I'm not doing a train test split
- 2:26:32specifically over here I'm giving the
- 2:26:34entire X and Y and probably based on
- 2:26:36that I'm going to do a cross validation
- 2:26:38over here you can also do train test
- 2:26:39plate initially and then just give the X
- 2:26:42train and Y train over here to do the
- 2:26:43cross validation it is up to you but the
- 2:26:45best practices will be that first you do
- 2:26:47the train test split and then only give
- 2:26:49the train data over here to do the cross
- 2:26:51validation I'm just going to use scoring
- 2:26:53is equal to you can use mean squared
- 2:26:56error negative mean squared error let's
- 2:26:58say that I'm going to use negative mean
- 2:27:00squ error again where do you find all
- 2:27:02these things you will be able to see in
- 2:27:04the SK learn page of L uh cross Val
- 2:27:06score and then finally in the cross Val
- 2:27:08score you give cross validation value as
- 2:27:105 10 whatever you want so after this
- 2:27:13what I'm actually going to do I'm just
- 2:27:14going to basically from this how many
- 2:27:17scores I will get the mean squar error
- 2:27:19will be five since I'm doing five cross
- 2:27:21validation if you don't believe me just
- 2:27:23see over here print msse so here you'll
- 2:27:26be able to see five different values 1 2
- 2:27:303 4 5 right five different mean values
- 2:27:34because we are doing cross five five
- 2:27:36cross validation so here what I'm going
- 2:27:37to write I'm just going to say np. mean
- 2:27:40I want to take the average of all the
- 2:27:41five so here will basically be my
- 2:27:45meanor
- 2:27:46msse okay and then probably I'll print I
- 2:27:49will print my Ms meanor MSC so this will
- 2:27:54be my average score with respect to this
- 2:27:56the negative value is there because we
- 2:27:58have used negative mean squ error but if
- 2:28:00you just consider mean square error then
- 2:28:01it is only 37.1 3 okay so this I have
- 2:28:05actually shown you how to do cross
- 2:28:06validation see with respect to linear
- 2:28:08regression you can't modify much with
- 2:28:10the parameter so that is the reason why
- 2:28:12specifically in order to overcome
- 2:28:14overfitting and do the feature selection
- 2:28:16we use uh R and lasso regression so here
- 2:28:19I will show show you how to do ridge
- 2:28:21ridge regression
- 2:28:24now now in order to do the prediction
- 2:28:26all you have to do is that just go over
- 2:28:28here take the model okay what is the
- 2:28:31model linear R and just say do
- 2:28:37predict so here you can see uh you'll be
- 2:28:40getting a function called as do predict
- 2:28:42and give the test value whatever you
- 2:28:44want to predict automatically the
- 2:28:45prediction will be done so I'm just
- 2:28:46going to remove this and focus on Ridge
- 2:28:48regression right now because I I want to
- 2:28:50show how hyperparameter tuning is done
- 2:28:52in R regression so for R regression the
- 2:28:54simple thing is that I'll be using two
- 2:28:56different libraries from skarn do
- 2:29:01linear linear uncore model I'm going to
- 2:29:05import Ridge so for the ridge it is also
- 2:29:09present in linear underscore model for
- 2:29:11doing the hyperparameter tuning I will
- 2:29:12be using from SK learn do modore
- 2:29:17selection and then I'm going to import
- 2:29:20grid SE CV so these are the two
- 2:29:22libraries that I'm actually going to use
- 2:29:24grid SE CV will be able to help you out
- 2:29:26with the um okay will be able to help
- 2:29:30you out with Hyper parameter tuning and
- 2:29:32then probably you'll be able to do
- 2:29:34that uh difference between MSE and
- 2:29:37negative MSE not big thing guys if you
- 2:29:39use MSE here mean squ error you'll be
- 2:29:42getting 37 I've just used negation of
- 2:29:45MSE it's okay anything is fine you can
- 2:29:48go with MSE also means square error
- 2:29:50there is also another uh another scoring
- 2:29:53area which is like which focuses on
- 2:29:54square root square mean Square uh sorry
- 2:29:58root means Square eror okay so there are
- 2:30:00different different things which you can
- 2:30:01basically focus on okay now in order to
- 2:30:05give you this specific good value I'm
- 2:30:07actually going to do hyper Peter tuning
- 2:30:10now let's go ahead with uh grid s CV so
- 2:30:13here what I'm going to do again I'm
- 2:30:14going to basically Define my model which
- 2:30:16will be
- 2:30:18Ridge okay so this this is what I have
- 2:30:20actually imported now uh let me open the
- 2:30:24ridge skarn so SK learn
- 2:30:28Ridge we need need to understand what
- 2:30:30all parameters are basically
- 2:30:33used do you remember this Alpha value
- 2:30:36guys do you remember this Alpha value
- 2:30:38why do we use Alpha I I told you now
- 2:30:40Alpha multiplied by slope square if you
- 2:30:43remember in Ridge we specifically use
- 2:30:45this right Ridge and lasso regression
- 2:30:48Alpha so this is the alpha the this is
- 2:30:50probably the best parameter we can
- 2:30:52perform hyper parameter tuning the next
- 2:30:54parameter that we can probably perform
- 2:30:56is basically uh this Max iteration okay
- 2:31:00Max iteration basically means how many
- 2:31:01number of iteration how many number of
- 2:31:03times we may probably change the Theta 1
- 2:31:05value to get the right value so we can
- 2:31:08do this so what I'm actually going to do
- 2:31:10I'm going to select some Alpha values
- 2:31:12I'm going to play with this apart from
- 2:31:14that if I want I can also play with the
- 2:31:16other parameters which are uh like kind
- 2:31:19of uh you know probably you can you can
- 2:31:21also play with the iteration parameter
- 2:31:23it is up to you try whichever parameter
- 2:31:25you want to change you can go ahead and
- 2:31:26change it now let me show you how do we
- 2:31:28write this and how do we make sure that
- 2:31:31this specific thing is done now uh
- 2:31:34before doing grid s CV uh let me do one
- 2:31:36thing I will Define my parameters okay
- 2:31:39so here is my Ridge now what I'm going
- 2:31:41to do I'm going to say parameters and in
- 2:31:44this parameter two important value that
- 2:31:46I'm probably going to take is this one
- 2:31:49that is my C value and I will try to
- 2:31:51Define this in the form of dictionaries
- 2:31:53so here the C value that I sorry not C
- 2:31:57just a second
- 2:31:58guys my mistake it is not C it is
- 2:32:03Alpha let's see so how do I Define my
- 2:32:05Alpha value we'll try to see so here the
- 2:32:09parameters will be Alpha C is basically
- 2:32:13for uh logistic regression I'll show you
- 2:32:16so the alpha value I will just mention
- 2:32:18some values like
- 2:32:201 e to the power of -5 that basically Me
- 2:32:2700000000 0 0 0 1 similarly I I can write
- 2:32:311 E to the^ of - 10 that again means 0 0
- 2:32:350 0 0 0 0 0 10 * 0 1 I'm just making fun
- 2:32:38okay so that you will also get
- 2:32:39entertained 1 E to the^ of minus 8 okay
- 2:32:43similarly I can write 1 E to the^ of
- 2:32:45minus 3 from this particular value now
- 2:32:48I'm increasing this value see 1 E to
- 2:32:50the^ of minus 2 and then probably I can
- 2:32:53have 1 5 10 um 20 something like this so
- 2:32:58I'm going to play with all this
- 2:32:59particular parameters for right now
- 2:33:01because in grit or CV what they do is
- 2:33:03that they take all the combination of
- 2:33:04this Alpha value and wherever your uh
- 2:33:07your your model performs well it is
- 2:33:09going to take that specific parameter
- 2:33:11and it is going to give you that okay
- 2:33:13this is the best fit parameter that is
- 2:33:14got selected so here I have got all
- 2:33:16these things now what I'm going to do
- 2:33:18I'm going to basically apply the grid C
- 2:33:19TV so here I have uh gridge uh sorry
- 2:33:24Ridge GD I'm
- 2:33:25saying ridore regressor so I'm going to
- 2:33:28use git s
- 2:33:31CV git s CV and here I'm basically going
- 2:33:34to take the parameters regge okay Ridge
- 2:33:36is my first model and then I will take
- 2:33:39up all this params that I have actually
- 2:33:40defined see in git CV if I press shift
- 2:33:43tab I have to first of all execute this
- 2:33:46then only it will be able to press shift
- 2:33:48tab so here if I press shift tab here
- 2:33:50you'll be able to see estimator and
- 2:33:52parameter grid is my second parameter
- 2:33:54then scoring and then all the other
- 2:33:56parameters so here the first thing that
- 2:33:58goes is your model then your parameters
- 2:34:00which what you are actually playing then
- 2:34:03the third parameter is basically your
- 2:34:05scoring
- 2:34:06scoring and again here I'm going to use
- 2:34:09negative mean squ error some people are
- 2:34:10saying that mean squared error is not
- 2:34:13present so that is the reason why
- 2:34:15negative mean squ error is done why it
- 2:34:18may not be present because
- 2:34:20uh they try to always create a generic
- 2:34:22Library probably this kind of uh scoring
- 2:34:24parameter may also get used in other
- 2:34:26algorithms so that is the reason they
- 2:34:28may not have created but if you want to
- 2:34:30Deep dive into it Google
- 2:34:33Google then what is r regress dot fit on
- 2:34:38X comma y again I'm telling you you can
- 2:34:40first of all do train test split on X
- 2:34:42and Y and then probably only do this on
- 2:34:44X train and Y train parameter is not oh
- 2:34:47sorry
- 2:34:49okay I get this okay parameter is not
- 2:34:53and why it is not and oh yeah it has
- 2:34:56become a
- 2:34:58list I'm going to make this as
- 2:35:00dictionary right now I'm fully focused
- 2:35:02on implementing things if I get an error
- 2:35:04I'll definitely make sure that it'll get
- 2:35:07fixed anyhow if I get that error I will
- 2:35:09not say oh Kish why why this error came
- 2:35:12you
- 2:35:13know why this error came I I'll not get
- 2:35:15worried I'll get the error down only you
- 2:35:17cannot give this as the one okay so try
- 2:35:21to understand okay so this is your gitar
- 2:35:23CV I've also done the fit and let's go
- 2:35:27and select the best parameter so what I
- 2:35:28can do I will write print
- 2:35:32ridore
- 2:35:35regressor dot
- 2:35:37params sorry there will be a parameter
- 2:35:39called as best params I'm going to print
- 2:35:42this and I'm going to print ridore
- 2:35:46regressor Dot
- 2:35:50best
- 2:35:51score so these are all the values that
- 2:35:53are got selected one is Alpha is equal
- 2:35:55to 20 and the best score is - 32 so
- 2:35:58initially I gotus 37 but because of
- 2:36:00Ridge regression you can see that our
- 2:36:02negative mean square error has
- 2:36:04definitely become better there is a
- 2:36:06minus sign don't worry but from 37 it
- 2:36:08has come to 32 cross validation guys
- 2:36:11over here inside grids s CV also when it
- 2:36:13is probably taking the entire
- 2:36:15combination over there the CV Value
- 2:36:17Cross validation also we can use
- 2:36:20so probably if I am probably considering
- 2:36:23all these
- 2:36:24things many people has a question Chris
- 2:36:27is this minus value increased that
- 2:36:29basically means you cannot use Ridge
- 2:36:31regression you are right in this
- 2:36:33particular case Ridge regression is not
- 2:36:34helping you out so guys let me again
- 2:36:36write it down everybody don't worry yeah
- 2:36:41previous I got minus 32 right now I'm
- 2:36:43getting - 37
- 2:36:45right sorry previously I got what - 37
- 2:36:52- 37 now I got - 32 so here you can see
- 2:36:56this I got it from linear regression
- 2:36:59this I got it from what Ridge which one
- 2:37:02should I select I should select this
- 2:37:03model only because it is performing well
- 2:37:05than this but again understand Ridge
- 2:37:08also tries to reduce the overfitting so
- 2:37:11probably in this particular scenario we
- 2:37:12cannot use Ridge because the performance
- 2:37:14is becoming more bad so what I will do I
- 2:37:17will go and try with lasso regression
- 2:37:20now I'll copy and paste the same thing
- 2:37:22so linear model import lasso then this
- 2:37:25will basically be my
- 2:37:27lasso let's see with lasso whether it
- 2:37:29will increase or not let's
- 2:37:33see this is my parameter that got
- 2:37:35selected now let me write lasto
- 2:37:38regressor
- 2:37:39dot best params so this is Alpha is
- 2:37:42equal to one is got selected over here
- 2:37:44I'm just going to print it okay and then
- 2:37:47I'm going to print with last one
- 2:37:49regression DOT score will be the best so
- 2:37:52here I'm actually getting - 35 - 35 here
- 2:37:56I'm actually getting - 32 so minus 35
- 2:37:59still I will focus on linear regression
- 2:38:01now see what will happen if I add more
- 2:38:04parameters if I add more parameters see
- 2:38:06what will happen so now I'm going to
- 2:38:08take Alpha different different values
- 2:38:10see this I'm just going to remove this
- 2:38:13and probably add Alpha value in this
- 2:38:16way see here I have added more values 5
- 2:38:1910 20 30 35 40 45 100 okay let's see
- 2:38:23whether we our performance will increase
- 2:38:25or not so here
- 2:38:28uh first of all let me remove from here
- 2:38:32in Ridge just take it down guys I'm I'm
- 2:38:35adding more parameters like this just
- 2:38:36take it down yeah CV is equal to 5
- 2:38:40nobody okay you're not able to see it um
- 2:38:43CV is equal to 5 now here it is uh what
- 2:38:46you can basically focus on so here you
- 2:38:49can see I have added some values like
- 2:38:51this you can also
- 2:38:52add and just try to execute and now if I
- 2:38:56go and probably see this is my see first
- 2:38:59I have tried for Ridge I'm getting minus
- 2:39:0229 do you see after adding more
- 2:39:04parameters what happened in Ridge after
- 2:39:07adding more parameters what happened in
- 2:39:09Ridge you can see om minus 29 and the
- 2:39:12alpha value that is got selected is 100
- 2:39:14if you want try with cross validation
- 2:39:1710 and just try to execute now
- 2:39:20now so these are are some hyper
- 2:39:22parameters that we will definitely play
- 2:39:24with here you can see - 29 so here you
- 2:39:27can see minus 29 you can also increase
- 2:39:30the cross validation
- 2:39:32value over here also and probably
- 2:39:34execute it but with lasso I don't know
- 2:39:38whether it is improving or not it is
- 2:39:39coming to minus 34 you just have to play
- 2:39:42with this parameters as now for a bigger
- 2:39:45problem statement the thing is not
- 2:39:47limited to here right we try to take
- 2:39:49multiples and many parameters multiples
- 2:39:52and many parameters and try to do these
- 2:39:54things it is up to you we play with
- 2:39:56multiple parameters whichever gives us
- 2:39:58the best result we are basically taking
- 2:40:00it it's okay error is increased I know
- 2:40:03that no error is increasing definitely
- 2:40:06error is increasing even though by
- 2:40:08trying with different different
- 2:40:09parameters but about most of the
- 2:40:11scenario see here I gotus 37 probably
- 2:40:14what I can actually do is that uh try to
- 2:40:17get better one with respect to this
- 2:40:20now the best way what I can also do is
- 2:40:22that I can basically take up train and
- 2:40:25test split also and probably do these
- 2:40:27things let's see let's see one example
- 2:40:29so how do we do train and test from SK
- 2:40:32scalar dot I think model selection
- 2:40:35import train test split okay it's okay
- 2:40:38guys you may get a different value okay
- 2:40:40let's do one thing okay let's make your
- 2:40:42problem statement little bit simpler now
- 2:40:45what I'm going to do just tell me in
- 2:40:46train test plate what we need to do so
- 2:40:48I'm going to take the same code I'm
- 2:40:50going to paste it over here or let me do
- 2:40:52one thing let me insert a cell below and
- 2:40:55let me do it for train test split so in
- 2:40:57train test plate what we can do so I'm
- 2:41:00just going to take the syntax paste it
- 2:41:02over here let's say that I'm taking XT
- 2:41:04train y train and then I'm using train
- 2:41:07test split with 33% now if I execute
- 2:41:10with respect to X train and Y train so
- 2:41:12here is my you can see this I have
- 2:41:13written this code from SK learn. model
- 2:41:15selection uh train test plate random
- 2:41:17State can be anything whatever you write
- 2:41:20it is fine then you basically give X and
- 2:41:22Y with test sizes 33 uh this is
- 2:41:25basically saying that the test will have
- 2:41:2733% and the train data will be 77% so
- 2:41:31this is what I'm actually getting with
- 2:41:33respect to X train and Y train here what
- 2:41:35I'm going to do I'm going to basically
- 2:41:37take X train comma y train and now if I
- 2:41:40go and probably see this here you can
- 2:41:41see minus 25 understand this value
- 2:41:44should go towards zero if it is going
- 2:41:47towards zero that basically means the
- 2:41:49performance is better now similarly I do
- 2:41:52it for Ridge in Ridge what I'm actually
- 2:41:54going to do here I'm going to write X
- 2:41:56train and Y train and if I go and
- 2:41:58probably select the best score than this
- 2:42:00here you'll be able to see I'm getting
- 2:42:03how much I'm getting minus
- 2:42:062.47 okay here I'm getting
- 2:42:0925.8 here 25. 47 that basically means
- 2:42:12now still the Improvement is little bit
- 2:42:15bad because here we are not going
- 2:42:17towards zero so the next part again here
- 2:42:20also you can basically do it for X train
- 2:42:22and Y train X train and Y train so here
- 2:42:25you have this one and let's go and
- 2:42:27execute this so here you can see minus
- 2:42:302.47 now what you can also do is that
- 2:42:33you can use this
- 2:42:35lasso regressor do predict and you can
- 2:42:39basically predict with respect to X test
- 2:42:42so this is your white test value suppose
- 2:42:44let's say that this is my y PR Yore PR
- 2:42:47then what I can do from SK
- 2:42:50learn I will be using R square and
- 2:42:53adjusted R square if you remember SK
- 2:42:55learn R square r² so this is my R2 score
- 2:43:00so where it is present in SK learn.
- 2:43:02Matrix so I'm going to write from SK
- 2:43:04learn import let's say I'm saying from
- 2:43:08skarn do Matrix import r² R2 score now
- 2:43:14what I'm going to do over here I'm
- 2:43:16basically going to say my R2 score which
- 2:43:20is my variable I'll say this is nothing
- 2:43:22but R2 score here I'm just going to give
- 2:43:24my y PR comma Yore test so if I go and
- 2:43:28probably see the output here I will be
- 2:43:30able to see print R2 score this is all I
- 2:43:34have discussed guys there is also
- 2:43:37adjusted rant score is there where is R2
- 2:43:41R2 score one adjusted r² okay R2 score
- 2:43:46is there but adjusted R square should be
- 2:43:48here somewhere in some manner so this is
- 2:43:52how your output looks like with respect
- 2:43:53to by using this lasso regressor okay
- 2:43:56which is very good okay it should be I
- 2:43:59told it should be near 100% right now
- 2:44:01I'm getting 67% if I want to tie with
- 2:44:04the ridge you can also try that so you
- 2:44:06can say Ridge regressor do predict and
- 2:44:10here you can see 7 68% then you can also
- 2:44:12try linear regressor and
- 2:44:16predict what is the error saying the
- 2:44:19regression is not fitted yet why why it
- 2:44:22is not fitted why it is not
- 2:44:25fitted let's say that I have fitted here
- 2:44:28linear
- 2:44:30regression dot fit on X train and Y
- 2:44:33train X train and comma y train so I'm
- 2:44:37just going to fit it now if I go and
- 2:44:40probably try to do the
- 2:44:41calculation so if I go and see my R2
- 2:44:44score it is also coming somewhere around
- 2:44:4668% 67% now since this is just a linear
- 2:44:50regression you won't be able to get 100%
- 2:44:52because you're drawing a straight line
- 2:44:53right so for that you basically have to
- 2:44:56other use other algorithms like XG boost
- 2:44:58and all n bias so many algorithms are
- 2:45:01there it's okay see you give y test over
- 2:45:04here y PR over here both are same right
- 2:45:06they're
- 2:45:07comparing by see at one limit you can
- 2:45:10you can increase the performance after
- 2:45:12that you cannot see again I'm telling
- 2:45:14you in linear regression what we do
- 2:45:15these are my points right I will be only
- 2:45:17able to create one best line I cannot
- 2:45:19create a curve line right over here so
- 2:45:21obviously my accuracy will be only
- 2:45:23limited let's go and do it logistic
- 2:45:26practical
- 2:45:27quickly and here uh in logistic also we
- 2:45:31can do git SE CV now what I'm actually
- 2:45:34going to do first of all let's go ahead
- 2:45:35with the data set so I will quickly
- 2:45:38Implement logistic so from LC learn.
- 2:45:41linear
- 2:45:42model I'm going to import logistic
- 2:45:46regression so I'm going to use logistic
- 2:45:48regression and apart from that we know
- 2:45:50that let's take a new data set because
- 2:45:52for logistic we need to solve using
- 2:45:54classification problem so this is
- 2:45:56basically my logistic regression I'll
- 2:45:58take one data set so from SK learn. data
- 2:46:01sets import we'll take a data set which
- 2:46:03is like uh breast cancer data set so
- 2:46:05that is also present in SK learn with
- 2:46:07respect to the breast cancer data set
- 2:46:09I'm just going to use this see load best
- 2:46:12cancer data set I'm loading it and all
- 2:46:14the independent features are in data and
- 2:46:16my columns are feature names the same
- 2:46:18thing like how we did previously okay so
- 2:46:20this will basically be my
- 2:46:23complete uh complete independent feature
- 2:46:25so if I go and probably see this x. head
- 2:46:28here you'll be able to see that based on
- 2:46:31this input features the independent
- 2:46:32feature we need to determine whether the
- 2:46:34person is having cancer or not these are
- 2:46:37some of the features over here and this
- 2:46:39is like many many features are actually
- 2:46:40present so next thing I this that was my
- 2:46:43independent feature now I'll take my
- 2:46:45dependent feature dependent feature will
- 2:46:47already present in DF Target okay this
- 2:46:50particular data set that we have taken
- 2:46:52in DF in DF do Target we will basically
- 2:46:55have all our dependent feature these are
- 2:46:56my independent features so what I'm
- 2:46:58actually going to do I'm going to create
- 2:46:59Y and I'm going to say PD do data frame
- 2:47:04and here I'm going to say DF do Target
- 2:47:07Target and this column name should be
- 2:47:11Target right so this will be my column
- 2:47:13name and now if I go and see my y y is
- 2:47:16basically having zeros and one in the
- 2:47:18target feature now the next thing that
- 2:47:20we are going to do is that uh apply
- 2:47:23basically apply the first of all we need
- 2:47:26to check whether this data set is uh
- 2:47:29this particular y column is balanced or
- 2:47:31imbalanced okay in order to do that I
- 2:47:33will just write F
- 2:47:35Target if the data set is imbalanced
- 2:47:38definitely we need to work on that and
- 2:47:40try to perform upsampling so if I write
- 2:47:42y target. Valore counts if I execute
- 2:47:46this so here you'll be able to see that
- 2:47:48value SC counts will basically give that
- 2:47:50how many number of ones are and how many
- 2:47:52number of zeros are so now total number
- 2:47:54of ones are 357 and total number of
- 2:47:57zeros are 22 so is this a imbalanced
- 2:48:01data set probably this is a balanced
- 2:48:03data set so here I'm actually going to
- 2:48:04now do train test spit train test spit I
- 2:48:08will try to do again train test spit how
- 2:48:10do we do we can quickly do copy the same
- 2:48:14thing entirely I'll copy this entirely
- 2:48:16over here and then I will get my X and Y
- 2:48:20so here is my X train X test y train y
- 2:48:22test so train test plate obviously I'll
- 2:48:24be doing it now in logistic regression
- 2:48:26if I go and search for
- 2:48:28logistic regression escalar I will be
- 2:48:31able to see this what all parameters are
- 2:48:33there this is basically the L1 Norm or
- 2:48:35L2 Norm or L1 regularization or L2
- 2:48:37regularization with respect to whatever
- 2:48:39things we have discussed in logistic and
- 2:48:41then the C value these two parameter
- 2:48:43values are very much important if I
- 2:48:45probably show you over here the penalty
- 2:48:49what kind of penalty whether you want to
- 2:48:50add L2 penalty L1 penalty you can use L2
- 2:48:53or L1 the next thing is C this is
- 2:48:56nothing but inverse of regularization
- 2:48:57strength this basically says 1 by Lambda
- 2:49:01something like that this parameter is
- 2:49:02also very much important guys class
- 2:49:04weight suppose if your data set is not
- 2:49:06balanced at that point of time you can
- 2:49:09apply weights to your classes if
- 2:49:11probably your data set is balanced you
- 2:49:14can directly use class weight is equal
- 2:49:16to balanced other than that you can use
- 2:49:18other other weight which you basically
- 2:49:19want so this is specifically some of
- 2:49:22this right no this is not Ridge or lasso
- 2:49:25okay this is logistic in logistic also
- 2:49:28you have L1 norm and L2
- 2:49:30Norms understand probably I missed that
- 2:49:32particular part in the theory but here
- 2:49:35also you have an L2 penalty norm and L1
- 2:49:37penalty Norm I probably did not teach
- 2:49:39you in theory because if you look see
- 2:49:43logistic regression can be learned by
- 2:49:45two different ways one is through
- 2:49:47probabilistic method and one is through
- 2:49:49geometric method if you go and probably
- 2:49:51see my video that is present with
- 2:49:52respect to logistic regression right now
- 2:49:54in my YouTube channel there I have
- 2:49:56explained you about this L1 and L2 Norms
- 2:49:58also over there so in this also it is
- 2:49:59basically present it is a kind of
- 2:50:01penalty again just for uh using for this
- 2:50:05kind of classification problem so what
- 2:50:08I'm actually going to do let's go and
- 2:50:10play with the parameters that I am
- 2:50:12looking at so I will play with two
- 2:50:14parameters one is params C value here
- 2:50:17I'm defining 1 10 20 anything that you
- 2:50:20can Define one set of values you can
- 2:50:22Define and there was one more parameter
- 2:50:24which is called as Max iteration this is
- 2:50:26specifically for grits or CV okay that
- 2:50:28I'm specifically going to apply so I
- 2:50:30will just try to execute this this will
- 2:50:32be my params now I'm going to quickly
- 2:50:34Define my model one which will be my
- 2:50:36logistic regression model so my logistic
- 2:50:39regression here by default one value
- 2:50:41I'll give for C and Max itra let's say
- 2:50:45I'm giving this value later on what I
- 2:50:47will do for this model I'll apply it to
- 2:50:49grid sear CV so I'm just going to say
- 2:50:51grid s CV and I'm going to apply it for
- 2:50:55model one param grid is equal to params
- 2:50:59this parameter that I'm specifically
- 2:51:01trying to apply since this is a
- 2:51:02classification problem and I am not
- 2:51:04pretty sure that whether true positive
- 2:51:06is important or true negative is
- 2:51:08important I'm going to use F1 scoring
- 2:51:10okay F1 scoring is basically again the
- 2:51:13parametric term which we discussed
- 2:51:14yesterday which is nothing but
- 2:51:16performance metrics and then I'm going
- 2:51:18to use CV is equal to 5 so this will be
- 2:51:21entirely my model with respect to grid s
- 2:51:24CV and I'll be executing this then I
- 2:51:27will do model. fit on my X train and Y
- 2:51:32train data so once I execute it here you
- 2:51:34can see all the output along with
- 2:51:36warnings a lot of warnings will be
- 2:51:38coming I don't know because this many
- 2:51:40parameters are there and finally you can
- 2:51:42see that this has got selected now if
- 2:51:44you really want to find out what is your
- 2:51:46best param score model
- 2:51:49dot best params so here you can see Max
- 2:51:52iteration as
- 2:51:54150 and what you can actually do with
- 2:51:58respect to your best score model do best
- 2:52:03score is 95 percentage but still we want
- 2:52:06to test it with test data so can we do
- 2:52:09it yes we can definitely do it I'll say
- 2:52:11model do core or I'll say model dot
- 2:52:15predict on my X test data and this will
- 2:52:18basically be my y red so this will be my
- 2:52:21y red all the Y prediction that I'm
- 2:52:23actually getting so if you go and see y
- 2:52:26red so these are my ones and zeros with
- 2:52:28respect to the Y
- 2:52:30prediction at finally after getting the
- 2:52:32prediction values I can apply confusion
- 2:52:35Matrix I hope I have taught you about
- 2:52:36confusion Matrix so from sklearn do
- 2:52:39confusion Matrix sorry sklearn do metrix
- 2:52:43I'm going to import confusion metrix
- 2:52:46classification report and the next thing
- 2:52:49that I would like to do is this two I
- 2:52:52will try to import confusion Matrix and
- 2:52:54classification report now if you want to
- 2:52:56see the confusion Matrix with respect to
- 2:52:58your I can just write
- 2:53:00Yore frad or Yore test whatever you want
- 2:53:04go ahead with it and this is basically
- 2:53:06my confusion Matrix if I put this
- 2:53:09forward no difference will be there only
- 2:53:11this thing will be moving that also I
- 2:53:13showed you 63 118 3 and 4 now finally if
- 2:53:17I want to accuracy score I can also
- 2:53:19import accuracy score over here so here
- 2:53:21you can see accuracy score is imported I
- 2:53:23can also find out my accuracy score
- 2:53:25which is my the total accuracy with
- 2:53:28respect to this I we can give y test and
- 2:53:31Yore PR which we have discussed
- 2:53:34yesterday this is giving
- 2:53:3596% if you want detailed Precision
- 2:53:38recall all the score then at that point
- 2:53:40of time I can use this classification
- 2:53:43report and here I can give white test
- 2:53:45and wied here is what I'm actually
- 2:53:47getting so here you can see with respect
- 2:53:50to F1 F1 score Precision recall since
- 2:53:52this is a balanced data set obviously
- 2:53:54the performance will be best yes you can
- 2:53:57also use Roc see I'll also show you how
- 2:54:00to use Roc and probably you'll be able
- 2:54:01to see this you have to probably
- 2:54:03calculate false positive rate two
- 2:54:05positive rate but don't worry about Roc
- 2:54:07I will first of all explain you the
- 2:54:08theoretical part now let's go ahead and
- 2:54:10discuss about n bias n bias is an
- 2:54:13important algorithm so here I'm just
- 2:54:16going to go ahead so now let's go ahead
- 2:54:18and discuss about na bias and here we
- 2:54:21are going to discuss about the intuition
- 2:54:23so na bias is an another amazing
- 2:54:26algorithm which is specifically used for
- 2:54:29classification and this specifically
- 2:54:31works on something called as base
- 2:54:34theorem now what exactly is base theorem
- 2:54:36first of all we need to understand about
- 2:54:38base theorem let's say that guys I have
- 2:54:41base theorem let's say that I have an
- 2:54:43experiment which is called as rolling a
- 2:54:45dis now in rolling a dis how many number
- 2:54:47of elements I have have so if I say what
- 2:54:49is the probability of 1 then obviously
- 2:54:51you'll be saying 1X 6 if I say
- 2:54:53probability of two then also here you'll
- 2:54:55say 1X 6 if I say probability of three
- 2:54:58then I will definitely say it is 1x 6 so
- 2:55:01here you know that this kind of events
- 2:55:04are basically called as independent
- 2:55:06events now rolling a dice why it is
- 2:55:08called as an independent event because
- 2:55:10getting one or two in every experiment
- 2:55:12one is not dependent on two two is not
- 2:55:14dependent on three so they are all
- 2:55:16independent that is the reason why we
- 2:55:18specifically say is an independent event
- 2:55:20but if I take an example of dependent
- 2:55:22events let's consider that I have a bag
- 2:55:24of marbles okay in this marble I
- 2:55:28basically have three red marbles and I
- 2:55:31have two green marbles now tell me what
- 2:55:33is the probability of suppose I have a
- 2:55:36event in the first event I take out a
- 2:55:38red marble so what is the probability of
- 2:55:40taking out a red marble so here you can
- 2:55:43definitely say that it is
- 2:55:443x5 okay so this is my first event now
- 2:55:47in the second event let's say that in
- 2:55:49this you have taken out the red marble
- 2:55:51now what is the second second time again
- 2:55:53you are taking out the second red marble
- 2:55:55or forget about second Rand marble now
- 2:55:57you want to take out the green marble
- 2:55:59now what is the probability with respect
- 2:56:01to taking out a green marble so here
- 2:56:03you'll be definitely saying that okay
- 2:56:05one red marble has been removed then the
- 2:56:07total number of marbles that are left
- 2:56:09are four so here you can definitely
- 2:56:11write that probability of getting a
- 2:56:12green marble is nothing but 2x4 which is
- 2:56:14nothing but 1x2 so here what is
- 2:56:16happening first first element you took
- 2:56:18out first marble that you took out first
- 2:56:20event from from the first event you took
- 2:56:21out red marble from the second event you
- 2:56:23took out green marble this two are in
- 2:56:25these two are dependent events because
- 2:56:28the number of marbles are getting
- 2:56:29reduced as you take out from them so if
- 2:56:32I tell you what is the probability of
- 2:56:35taking out a red marble and then a green
- 2:56:39marble so it's the simple the formula
- 2:56:42will be very much simple right which we
- 2:56:43have already discussed in stats it is
- 2:56:45nothing but probability of probability
- 2:56:47of red multiplied by probability of
- 2:56:50green given Red so this specific thing
- 2:56:53is called as conditional probability
- 2:56:55here understand what is happening
- 2:56:57probability of green marble given the
- 2:56:59red marble event has occurred here both
- 2:57:01the events are independent now let me
- 2:57:03write it down very nicely so I can write
- 2:57:05probability of A and B is equal to
- 2:57:08probability of a multiplied probability
- 2:57:12of B divided by probability of a let's
- 2:57:16go and derive something can can I write
- 2:57:18probability of A and B is equal to
- 2:57:21probability of b and a so answer is yes
- 2:57:24we can definitely say we can definitely
- 2:57:26say if you go and do the calculation
- 2:57:27you'll be able to get the answer you
- 2:57:29should not say no now what is the
- 2:57:32formula for probability of A and B so
- 2:57:34here you can basically write probability
- 2:57:36of a multiplied by probability of B
- 2:57:39given a if I take out probability of
- 2:57:42green what is probability of green in
- 2:57:43this particular case 2x 5 what is
- 2:57:46probability of red 3x 4 for right now
- 2:57:49let's consider this now this part I can
- 2:57:51definitely write as this part I can
- 2:57:54definitely write as probability of B
- 2:57:56multiplied by probability of B
- 2:58:00probability of B this one probability of
- 2:58:02B and this will be probability of a
- 2:58:04given B so I can definitely write this
- 2:58:06much with respect to all this
- 2:58:08information now can I derive probability
- 2:58:10of a is equal to probability of B
- 2:58:14multiplied by probability of a / B me
- 2:58:18probability of a given B divided by
- 2:58:21probability of sorry I'll write this as
- 2:58:24probability of B given a divided by
- 2:58:27probability of a and this is
- 2:58:28specifically called as base theorem and
- 2:58:31this is the Crux behind na bias
- 2:58:34understand this is the Crux behind the
- 2:58:35base theorem now let's go ahead and
- 2:58:38let's discuss about how we are using
- 2:58:40this to solve let's take some examples
- 2:58:43and probably make you understand let's
- 2:58:45say that I have some features like X1 X2
- 2:58:49X3 X4 X5 like this till xn and I have my
- 2:58:54output y so these are my independent
- 2:58:56features these all are my independent
- 2:58:58features these all are my independent
- 2:59:00features so here I'm going to write
- 2:59:02independent features and this is my
- 2:59:04output feature which is also my
- 2:59:05dependent feature now what is happening
- 2:59:08if I say probability of b or a what does
- 2:59:10this basically mean I need to really
- 2:59:12find what is the probability of Y and
- 2:59:15you know that guys I will have some
- 2:59:17values over here and basically I'll have
- 2:59:19some output value over here so based on
- 2:59:21this input values I need to predict what
- 2:59:23is the output initially on a training
- 2:59:25data set I will have your input and then
- 2:59:28your output initially my model will get
- 2:59:30trained on this now let's consider what
- 2:59:32this entire terminology is I will try to
- 2:59:34write in terms of this equation so I
- 2:59:36will say probability of Y given x1a x2a
- 2:59:41X3 up till xn then this equation will
- 2:59:44become probability of Y see probability
- 2:59:46of Y given X X1 X2 X3 xn this a is
- 2:59:50nothing but X1 X2 X3 xn and I'm trying
- 2:59:52to find out what is the probability of Y
- 2:59:54and then I will write probability of b b
- 2:59:57is nothing but y but before that what
- 3:00:00I'll write probability of a / B right a
- 3:00:03given b or probability of B probability
- 3:00:06of B is nothing but y multiplied by
- 3:00:08probability of a given B probability of
- 3:00:12a given B basically means probability of
- 3:00:15x1a X2 comma xn and given b b is given
- 3:00:20right so I'm able to find this entire
- 3:00:22value now just a second I made some
- 3:00:24mistakes I guess now it is correct sorry
- 3:00:26I I just missed one term that is this
- 3:00:29given y this is how it will become and
- 3:00:32this will be equal to probability of a
- 3:00:36that is X1 comma X2 like this up to XL
- 3:00:39so probability of Y multiplied by
- 3:00:41probability of a given y now if I try to
- 3:00:44expand this then this will basically
- 3:00:46become something like this see
- 3:00:48probability of Y multiplied by
- 3:00:51probability of X1 given yes a given y
- 3:00:56sorry given y multiplied by probability
- 3:01:00of X2 given y probability of x3 given Y
- 3:01:06and like this it will be probability of
- 3:01:08xn given y so this will also be y1 Y2 Y3
- 3:01:12YN this I can expand it like this and
- 3:01:15then this will basically become
- 3:01:16probability of X Y 1 multiplied by
- 3:01:18probability of X2 multiplied by
- 3:01:21probability of x3 like this up to
- 3:01:23probability of xn so this is with
- 3:01:26respect to all the probability y will be
- 3:01:28different see here for this particular
- 3:01:30record y will be different for this y
- 3:01:32will be different for this y will be
- 3:01:34different but why output it may be yes
- 3:01:37or no right it may be yes or no okay I
- 3:01:41I'll solve a problem it will make
- 3:01:43everything understand and this will
- 3:01:45probably be probability of Y it can be
- 3:01:47binary multiclass whatever things you
- 3:01:49want I'll solve a problem in front of
- 3:01:51you now let's say that I have my y as
- 3:01:54let's say that I have a lot of features
- 3:01:56X1 X2 X3 X X4 with respect to this let's
- 3:02:02say in my one of my data set I have this
- 3:02:03many x1s this many features and this is
- 3:02:06my y so these are my feature number and
- 3:02:09this is my y let's say that in y I have
- 3:02:11yes or no so how I will probably write
- 3:02:15we really need to understand this okay I
- 3:02:17will basically
- 3:02:18say what is the probability of Y is
- 3:02:21equal to yes given this x of I this is
- 3:02:25my first record first record of X of I
- 3:02:27this is my second record of X of I so I
- 3:02:30may write like this what is the
- 3:02:31probability of Y being yes if x of I is
- 3:02:34given to you X of I basically means X1
- 3:02:37X2 X3 X4 so here you'll obviously write
- 3:02:39what kind of equation you'll basically
- 3:02:41say probability of yes multiplied by
- 3:02:45probability of yes multiplied by
- 3:02:46probability of X of 1 given
- 3:02:50yes multiplied by probability of X2
- 3:02:53given yes probability of x3 given yes
- 3:02:58and probability of X4 given yes divided
- 3:03:03by probability of X1 multiplied by
- 3:03:06probability of X2 multiplied by
- 3:03:08probability of x3 multiplied by
- 3:03:10probability of X4 Y is fixed it may be
- 3:03:13yes or it may be no but with respect to
- 3:03:15different different records this value
- 3:03:17may change similarly if I write
- 3:03:18probability of Y is equal to no given X
- 3:03:22of I what it will be then it will be
- 3:03:26probability of no multiplied by
- 3:03:30probability of X1 given no then
- 3:03:33probability of
- 3:03:35X2 given
- 3:03:37no probability of
- 3:03:39x3 given
- 3:03:42no and probability of X4 given no so
- 3:03:46here because every any input that I give
- 3:03:49any input X of I that I give I may
- 3:03:51either get yes or no so I need to find
- 3:03:53both the probability so probability of
- 3:03:54X1 multiplied by probability of X2
- 3:03:57multiplied by probability of x3
- 3:03:59multiplied by probability of X4 see with
- 3:04:02respect to Any X of I the output can be
- 3:04:05yes or no and I really need to find out
- 3:04:07the probabilities so both the formula is
- 3:04:09written over here what is the
- 3:04:11probability of with respect to yes and
- 3:04:13what is the probability with respect to
- 3:04:14no now in this case one common thing you
- 3:04:17see that this this denominator is fixed
- 3:04:20this is definitely fixed it is fixed it
- 3:04:22is it is not going to change for both of
- 3:04:24them and I can consider that this is a
- 3:04:27constant so what I can do I can
- 3:04:30definitely ignore so here I can
- 3:04:32definitely ignore these things ignore
- 3:04:34this also ignore this Al because see
- 3:04:36this is constant so I don't want to
- 3:04:38consider this in the next time I'll just
- 3:04:40use this specific formula to calculate
- 3:04:42the probability now let's say that if my
- 3:04:46first probability for a specific data
- 3:04:49set yes of X of I is let's say that I'm
- 3:04:52getting
- 3:04:53as13 and similarly probability of no
- 3:04:56with respect to X of I if I get
- 3:05:0005 you know that in a binary
- 3:05:02classification any values if it get
- 3:05:04greater than or equal to 5 we are going
- 3:05:06to consider it as 1 and if it is less
- 3:05:09than 0.5 I'm going to consider it as
- 3:05:10zero now I'm getting values like this 13
- 3:05:13and .1 05 obviously I'm getting .13 05
- 3:05:18so we do something called as
- 3:05:20normalization it says that if I really
- 3:05:23want to find out the probability of X
- 3:05:24with X of I if I do normalization it is
- 3:05:27nothing but .13 divided by .13 +
- 3:05:3105 72 this is nothing but
- 3:05:3572% and similarly if I do for
- 3:05:37probability of no given X of I here
- 3:05:39obviously it will say 1 - 72 which will
- 3:05:42be your remaining answer that is 28
- 3:05:44which is nothing but 28% so your final
- 3:05:47answer will be this one this formulas
- 3:05:49you have to remember now we'll solve a
- 3:05:50problem let's solve a problem this will
- 3:05:52be a very very interesting problem let's
- 3:05:54say I have a data set which has like
- 3:05:56this feature day let me just copy this
- 3:05:59data set okay for you all now in this
- 3:06:02data set I want to take out some
- 3:06:04information let's take out Outlook
- 3:06:08table now based on this output Outlook
- 3:06:11feature see over here Outlook my day
- 3:06:14outlook temperature humidity wind are
- 3:06:17the input features independent feature
- 3:06:19this is my output feature this one that
- 3:06:22you are probably seeing play tennis is
- 3:06:24my output feature which is specifically
- 3:06:26a binary
- 3:06:27classification so what I'm actually
- 3:06:29going to do I'm basically going to take
- 3:06:31my Outlook feature and based on this
- 3:06:33Outlook feature I will just try to
- 3:06:34create a smaller table which will give
- 3:06:36some information now based on Outlook
- 3:06:39first of all try to find out how many
- 3:06:40categories are there in Outlook one is
- 3:06:43sunny one is
- 3:06:45overcast and one is rain right three
- 3:06:48categories are there so I'm going to
- 3:06:50write it down over here Sunny overcast
- 3:06:53and rain so these three are my features
- 3:06:56with respect to Sunny uh with Outlook I
- 3:06:58have three categories one is sunny one
- 3:07:00is overcast and one is RA here I'm going
- 3:07:02to basically say with respect to Sunny
- 3:07:05how many yes are there and how many no
- 3:07:08are there and what is the probability of
- 3:07:11yes and probability of no so I'm going
- 3:07:13to again write it over here so this is
- 3:07:16my Outlook feature
- 3:07:18and then I have categories first yes no
- 3:07:23Sunny overcast rain yes no then
- 3:07:28probability of yes and probability of no
- 3:07:31now the next thing that we need to find
- 3:07:33out is that with respect to Sunny how
- 3:07:37many of them are yes see yes we have so
- 3:07:40when we have sunny over here the answer
- 3:07:42is no so I will increase the count over
- 3:07:44here one then again I have sunny again
- 3:07:47answer is no so I'm going to increase
- 3:07:49the count to two with this sunny this is
- 3:07:52basically no okay so again I'm going to
- 3:07:54increase the count to three now with
- 3:07:56sunny how many of them are yes one and
- 3:08:00two so I have this one and this one so I
- 3:08:03have two so I'm going to say with
- 3:08:05respect to Sunny I have two
- 3:08:07yes understand Outlook is my X1 X1
- 3:08:11feature let's consider now the next
- 3:08:13thing is that let's see with respect to
- 3:08:16overcost with overcast how many of them
- 3:08:18are yes so this overcast is there yes 1
- 3:08:222 3 and four so total four yes are there
- 3:08:26with respect to overcast then with
- 3:08:28respect to overcast how many are on no
- 3:08:31you can go ah and find out it is
- 3:08:32basically zero NOS then with respect to
- 3:08:35rain how many of them are yes so here
- 3:08:37you can see with respect to one rain yes
- 3:08:40yes no no so this is nothing but 3 2
- 3:08:46let's try to find out there are three is
- 3:08:47two or
- 3:08:48not one here also one yes is there right
- 3:08:52so 3 yes two NOS so the total number of
- 3:08:55yes and NOS if you count it there are
- 3:08:58nine yes and five NOS this is my total
- 3:09:01count so if you totally count this 9 + 5
- 3:09:04is 14 you'll be able to compare that
- 3:09:06there will be 9 yes and five NOS what is
- 3:09:08the probability of yes when Sunny is
- 3:09:10given so here you have 2X 9 here you
- 3:09:14have 4X 9 here you have 3x 9 now if if I
- 3:09:17say what is the probability of no given
- 3:09:20Sunny now see probability of yes given
- 3:09:23Sunny probability of yes given forecast
- 3:09:26probability of yes given rain so it is
- 3:09:28basically that I will just try to write
- 3:09:30it in a simpler manner so that you'll
- 3:09:31not get confused okay so this is my
- 3:09:33probability of yes and this is my
- 3:09:35probability of no but understand what
- 3:09:37does this basically mean this
- 3:09:39terminology basically means probability
- 3:09:41of yes given Sunny probability of yes
- 3:09:44given overcast probability of yes given
- 3:09:46rain similarly what is probability of no
- 3:09:49probability of no obviously you know
- 3:09:50that 3x 5 is my first probability then
- 3:09:54you have 0x 5 and then you have 2X 5 now
- 3:09:58with respect to the next feature let's
- 3:10:00consider that I'm going to consider one
- 3:10:01more feature and in this feature I will
- 3:10:03say let's consider
- 3:10:05temperature okay let's consider
- 3:10:07temperature now in temperature how many
- 3:10:10features I have or how many categories I
- 3:10:12have I have hot you can see hot mild and
- 3:10:17and cold now with respect to hot mild
- 3:10:19cold here also I will be having yes no
- 3:10:23probability of yes and probability of no
- 3:10:26now try to find out with respect to hot
- 3:10:28how many are yes so here no is there
- 3:10:31here also no is there two NOS uh 1 yes
- 3:10:36uh 2 yes so two yes and two NOS probably
- 3:10:39then similarly with respect to mild mild
- 3:10:42how many are there 1 yes 1 No 2 yes 3s
- 3:10:484s 4S and two knows okay so here you
- 3:10:51basically go and calculate 4 yes and two
- 3:10:54knows with respect to cold how many are
- 3:10:57there cool cool or cold 1 yes 1 No 2 yes
- 3:11:033 S 3 S and 1 no so here I have
- 3:11:07specifically have 3s and 1 no again the
- 3:11:10total number is 9 and five which will be
- 3:11:12equal to the same thing that what we
- 3:11:15have got now really go ahead with
- 3:11:16finding probability of yes given hot so
- 3:11:19it will be 2x 9 over here then here it
- 3:11:22will be how much 4X 9 here it will be 3x
- 3:11:269 again here what will be the
- 3:11:28probability of no given given hot so
- 3:11:31it'll be 2x 5 2x 5 1X 5 so this two
- 3:11:36tables has already been created and
- 3:11:37finally with respect to play the total
- 3:11:39number of plays are yes is 9 no is five
- 3:11:44and the answer is total 14 if if I say
- 3:11:47what is the probability of yes only yes
- 3:11:50then it is nothing but 9 by4 what is the
- 3:11:54probability of no it is nothing but
- 3:11:565x4 okay so this two values also you
- 3:11:59require now let's say that you get a new
- 3:12:02data set you need get a new data set
- 3:12:05let's say you get a new test data where
- 3:12:08it says that suppose if you are having
- 3:12:11sunny and hot tell me what is the output
- 3:12:16so this is my problem statement so let
- 3:12:18me write it down so here I will write
- 3:12:20probability of yes given Sunny comma hot
- 3:12:25then here I will write probability of
- 3:12:27yes multiplied by probability of so here
- 3:12:31I will write probability of Sunny given
- 3:12:34yes multiplied by probability of hot
- 3:12:38given yes divided by what is it
- 3:12:42probability of Sunny multiplied by
- 3:12:45probability of hot
- 3:12:50equation because it is a
- 3:12:52constant because probability of no also
- 3:12:55I'll be getting the same value 9 by4 so
- 3:12:58probability of yes I'm going to replace
- 3:13:00it with 9
- 3:13:02by4 multiplied by 2x 9 then probability
- 3:13:06of hot given yes so I am going to get 2
- 3:13:09by 9 so
- 3:13:12here 99 cancel or 2 1 7 then this is
- 3:13:17nothing but 2 by
- 3:13:216331 I read this statement little bit
- 3:13:23wrong it should be probability of Sunny
- 3:13:25given yes now go ahead and calculate go
- 3:13:28ahead and calculate what is probability
- 3:13:30of no given sunny and hot so here you
- 3:13:33have probability of no multiplied by
- 3:13:36probability of Sunny given
- 3:13:38no multiplied by probability of hot
- 3:13:43given
- 3:13:44no divided by probability of Sunny
- 3:13:50multiplied by probability of heart this
- 3:13:53will get cancelled denominator is a
- 3:13:55constant guys this is a constant so what
- 3:13:58is probability of no so probability of
- 3:14:00no is nothing but 5 by4 so I will write
- 3:14:03over here 5 by4 multiplied by
- 3:14:07probability of Sunny given no what is
- 3:14:09probability of Sunny given no what is
- 3:14:11probability of Sunny given no is nothing
- 3:14:13but probability of Sunny given no is
- 3:14:15nothing but 3x 5 so here I'm going to
- 3:14:17get 3x 5 multiplied probability of H
- 3:14:22given no that is nothing but 2x 5 so 2x
- 3:14:255 is here 3x 5 is there five and five
- 3:14:28will get cancelled 2 1 2 7 and then I'm
- 3:14:32getting 3x 35 which is nothing but
- 3:14:35calculator uh if I'm actually getting
- 3:14:37three ID by 35 it's nothing but
- 3:14:41857 I will write it down again
- 3:14:44probability of yes given Sunny comma hot
- 3:14:49which is my independent feature is
- 3:14:51nothing but
- 3:14:52031
- 3:14:54031 and this is probability of no given
- 3:14:57Sunny comma hot 85 now we'll try to
- 3:15:00normalize this 85 + Point divided by 031
- 3:15:06+ 085 73 this is nothing but 73% and
- 3:15:11here I can basically say 1 -73 which is
- 3:15:14my27 which is nothing but 27% if the
- 3:15:18input comes as sunny and hot if the
- 3:15:21weather is sunny and hot what will the
- 3:15:23person do whether he will play or not
- 3:15:26the answer is no okay now my next
- 3:15:29question will be that if your new data
- 3:15:31is overcast and Mild now tell me what
- 3:15:34will be the probability using name bias
- 3:15:37now you can add any number of features
- 3:15:39let's say that I will say that okay
- 3:15:42let's let's say that I will I will
- 3:15:44probably say we can consider humidity
- 3:15:47mind wind also you basically create this
- 3:15:49kind of table to find it out but this
- 3:15:50will be an assignment just do
- 3:15:53it overcast and Mild if it is with
- 3:15:56respect to NB try to solve it so the
- 3:15:58second algorithm that we are going to
- 3:16:00discuss about is something called as KNN
- 3:16:02algorithm KNN algorithm is a very simple
- 3:16:05problem statement okay which can be used
- 3:16:09to solve both classification and
- 3:16:11regression so KNN basically means K
- 3:16:14nearest neighbor let's first of all
- 3:16:16discuss about classification problem
- 3:16:18number one classification problem let's
- 3:16:20say that I have a binary classification
- 3:16:22problem which looks like this I have two
- 3:16:23data points like this one and this is
- 3:16:26another one suppose a new data point
- 3:16:29suppose a new data point which comes
- 3:16:31over
- 3:16:32here then how do I say that whether this
- 3:16:35belongs to this category or whether it
- 3:16:36belongs to this category if I probably
- 3:16:38create a logistic regression I may
- 3:16:40divide a line but in this particular
- 3:16:42scenario how do we Define or how do we
- 3:16:44come to a conclusion that
- 3:16:47whether this will belong to this
- 3:16:48category or this category so for here we
- 3:16:50basically use something called as K
- 3:16:52nearest neighbor let's say that I say
- 3:16:55that my K value is five so what it is
- 3:16:57going to do it is going to basically
- 3:16:58take the five nearest closest point
- 3:17:01let's say from this you have two nearest
- 3:17:03closest point and from here you have
- 3:17:05three nearest closest point so here we
- 3:17:07basically see from the distance the
- 3:17:09distance that which is my nearest point
- 3:17:11now in this particular case you see that
- 3:17:13maximum number of points are from Red
- 3:17:15categories from Red from Red categories
- 3:17:18I'm getting three points and from White
- 3:17:21categories I'm getting two points now in
- 3:17:23this particular scenario maximum number
- 3:17:25of categories from where it is coming we
- 3:17:27basically categorize that into that
- 3:17:29particular class just with the help of
- 3:17:30distance which all distance we
- 3:17:31specifically use we use two distance one
- 3:17:33is ukan distance and the other one is
- 3:17:36something called as Manhattan distance
- 3:17:37so ukan and Manhattan distance now what
- 3:17:40does ukan distance basically say suppose
- 3:17:42if this is your two points which is
- 3:17:44denoted by X1 y1
- 3:17:47X2 Y2 ukine distance in order to
- 3:17:50calculate we apply a formula which looks
- 3:17:52like this X2 - X1 s + Y2 - y1 s whereas
- 3:17:58in the case of magetan distance suppose
- 3:18:00this are my two points then we calculate
- 3:18:03the distance in this way we calculate
- 3:18:05the distance from here then here right
- 3:18:07this is the distance we calculate we
- 3:18:09don't calculate the hypothenuse distance
- 3:18:10so this is the basic difference between
- 3:18:11ukan and magetan distance now you may be
- 3:18:14thinking Chris then fine that is for
- 3:18:15classification problem for regression
- 3:18:17what do we do for regression also it is
- 3:18:19very much simple suppose I have all the
- 3:18:22data points which looks like this now
- 3:18:24for a new data point like this if I want
- 3:18:26to calculate then we basically take up
- 3:18:28the nearest Five Points let's say my K
- 3:18:30is five k is a hyper parameter which we
- 3:18:33play now suppose let's say that K it
- 3:18:35finds the nearest point over here here
- 3:18:38here here and here so if we need to find
- 3:18:42out the point for this particular output
- 3:18:44with respect to the K is equal to 5 it
- 3:18:46will try to calculate the average of all
- 3:18:48the points once it calculates the
- 3:18:51average of all the points that becomes
- 3:18:53your output so regression and
- 3:18:55classification that is the only
- 3:18:56difference because this K is actually an
- 3:18:58hyper parameter we try with K is equal
- 3:19:00to 1 to 50 and then we probably try to
- 3:19:03check the error rate and if the error
- 3:19:06rate is less then only we select the
- 3:19:08model now two more things with respect
- 3:19:10to K nearish neighbor K nearest neighbor
- 3:19:12works very bad with respect to two
- 3:19:15things one is outliers and and one is
- 3:19:17imbalanced data set now if I have an
- 3:19:19outlier let's say I have an outlier over
- 3:19:22here this is one of my categories like
- 3:19:24this and this is my another category
- 3:19:26let's consider that I have some outliers
- 3:19:28which looks like this now if I'm trying
- 3:19:29to find out the point for this you can
- 3:19:32see that the nearest point is basically
- 3:19:35blue only and it belongs to the blue
- 3:19:37category but because this outlier you
- 3:19:39know it'll consider that the nearest
- 3:19:40neighbor is this so then this will be
- 3:19:42basically treated in this group only
- 3:19:44formula for Manhattan distance it uses
- 3:19:46modulus X2 - X1 + Y2 - y1 mode X2 - X1
- 3:19:53Y2 - y1 uh this was it from my side guys
- 3:19:55and yes I've also made detailed videos
- 3:19:57about whatever topics we have discussed
- 3:19:59today you can directly go and search for
- 3:20:01that particular
- 3:20:03topic so this is the agenda of this
- 3:20:06session we will try to complete this all
- 3:20:08things again here we are going to
- 3:20:10understand the mathematical equations
- 3:20:12and all uh so today's session we are
- 3:20:14basically going to discuss about uh
- 3:20:16decision tree okay and uh in this
- 3:20:20session we are going to basically
- 3:20:21understand what is the exact purpose of
- 3:20:23decision tree with the help of decision
- 3:20:25tree you are actually solving two
- 3:20:27different problems one is regression and
- 3:20:30the other one is
- 3:20:32classification so we'll try to
- 3:20:34understand both this particular part
- 3:20:37well we will take a specific data set
- 3:20:38and try to solve those problems now
- 3:20:40coming to the decision tree one thing
- 3:20:42you need to understand I'll say that if
- 3:20:45age is less than 8 let's say I'm writing
- 3:20:48this condition if age is less than or
- 3:20:51equal to 18 I'm going to say print go to
- 3:20:55college here I'm printing print college
- 3:20:58and then I'll write else if age is
- 3:21:02greater than 18 and pag is less than or
- 3:21:05equal to 35 I'll say print work then
- 3:21:09again I'll write else if age is let me
- 3:21:12let me put this condition little bit
- 3:21:14better then I'll write here L if if age
- 3:21:17is greater than 18 and age is less than
- 3:21:22or equal to 35 I'm going to say print
- 3:21:25work basically people needs to work in
- 3:21:27this age else I'm just going to consider
- 3:21:30print retire so here is my ifls
- 3:21:34condition over here now whenever we have
- 3:21:36this kind of nested if Els condition
- 3:21:38what we can do is that we can also
- 3:21:40represent this in the form of decision
- 3:21:42trees we'll also we can actually form
- 3:21:45this in the form of decision and the
- 3:21:46decision tree here first of all we will
- 3:21:48have a specific root node let's say this
- 3:21:51is my root node now in this root node
- 3:21:52the first condition is less than or
- 3:21:54equal to 18 so here obviously I will be
- 3:21:56having two conditions saying that if it
- 3:21:59is less than or equal to 18 and one
- 3:22:02condition will be yes one condition will
- 3:22:03be no so if this is yes and if this is
- 3:22:06no right if this condition is true that
- 3:22:09basically means we'll go in this side if
- 3:22:11it is true then here we will basically
- 3:22:14have something like college so this is
- 3:22:17your Leaf node similarly when I have no
- 3:22:22okay no no in this particular case we
- 3:22:24will go to the next condition in this
- 3:22:26next condition I will again create a
- 3:22:28node and I'll say that okay this is less
- 3:22:30than 18 and greater than sorry less than
- 3:22:33or equal to 35 so if this is also there
- 3:22:38then again I'll have two conditions
- 3:22:39which is basically yes or no now when I
- 3:22:42create this yes or no over here you'll
- 3:22:43be able to see that basically means here
- 3:22:46again two condition will be there if it
- 3:22:48is yes I will say print work so this
- 3:22:50will again be my leaf
- 3:22:52node and again for no again I will do
- 3:22:55the further splitting which is retire so
- 3:22:59here you can see that this entire
- 3:23:00algorithm this entire code that I have
- 3:23:02actually written you can see that it has
- 3:23:05got converted to this kind of
- 3:23:08trees where you specifically able to
- 3:23:10take decisions yes or no so can we solve
- 3:23:15a classification
- 3:23:17problem sorry this is greater than 18
- 3:23:21again if it is greater than 18 and less
- 3:23:23than or 35 so can we solve a
- 3:23:28regression and a classification problem
- 3:23:31regression and classification problem
- 3:23:34using this decision trees by creating
- 3:23:37this kind of
- 3:23:38nodes so in short whenever we talk about
- 3:23:41decision
- 3:23:42trees whenever we talk about decision
- 3:23:45trees
- 3:23:47you will be seeing that decision trees
- 3:23:49are nothing but decision trees are
- 3:23:52nothing but by using this nested if El
- 3:23:56condition we can definitely solve some
- 3:23:58specific problem statement but here in
- 3:24:00the visualized way we will specifically
- 3:24:02create this decision tree in the form of
- 3:24:04nodes now you need to understand that
- 3:24:07what type of maths we will probably use
- 3:24:10okay so let's do one thing let's take a
- 3:24:12specific data set which I will
- 3:24:14definitely do it over here in front of
- 3:24:15you
- 3:24:17okay and we will try to solve this
- 3:24:18particular data set and this will
- 3:24:20basically give you an idea like how we
- 3:24:23can probably solve these problems so uh
- 3:24:26let me just open my snippet tool so this
- 3:24:29is my data set that I have let's
- 3:24:31consider that I have this specific data
- 3:24:33set now this data set are pretty much
- 3:24:35important because this probably in
- 3:24:39research papers also probably people who
- 3:24:41have come up with this algorithm they
- 3:24:43usually take this they take this thing
- 3:24:46but but right now this particular
- 3:24:47problem statement if I talk about this
- 3:24:49is a classification problem statement
- 3:24:51okay but don't worry I will also help
- 3:24:53you to explain I'll also explain you
- 3:24:56about regression also how decision tree
- 3:24:58regression will definitely work so let's
- 3:25:01go ahead and let's try to understand
- 3:25:03suppose if I have this specific problem
- 3:25:05statement how do we solve this this is
- 3:25:07my output feature play tennis yes or no
- 3:25:10okay whether the person is going to pay
- 3:25:12tennis or not yesterday or there after
- 3:25:14yesterday or whenever you want so if I
- 3:25:17have this input features like Outlook
- 3:25:19temperature humidity and wind is the
- 3:25:22person going to play tennis or not this
- 3:25:24is what my model should predict with the
- 3:25:26help of decision tree so how decision
- 3:25:28tree will work in this particular case
- 3:25:29first of all let's consider any any any
- 3:25:33specific uh feature let's say that
- 3:25:35Outlook is my feature so this will be my
- 3:25:37first
- 3:25:38feature which is specifically Outlook
- 3:25:41now just tell me how many are basically
- 3:25:45having no and how many are basically
- 3:25:48having yes in the case of Outlook there
- 3:25:51you'll be able to find out there are
- 3:25:52nine yes see 1 2 3 4 5 6 7 8 9 and how
- 3:25:58many NOS are there 1 2 3 4 5 I think 1 2
- 3:26:043 4 5 so nine yes and five NOS what we
- 3:26:09are going to do in this specific thing
- 3:26:11now we have N9 yes and five Nos and the
- 3:26:13first node that I have actually taken
- 3:26:17is basically Outlook so Outlook feature
- 3:26:20now just try to find out we are focusing
- 3:26:22on this specific feature now in this
- 3:26:24feature how many categories I have I
- 3:26:26have one Sunny category you can see over
- 3:26:29here I have Sunny one category then I
- 3:26:31have another category called as
- 3:26:33overcast then I have another category as
- 3:26:37rain so I have three unique categories
- 3:26:40So based on these three categories I
- 3:26:42will try to create three nodes so here
- 3:26:45is my one node here is my second node
- 3:26:49here is my third node so these are my
- 3:26:52three categories so this category is
- 3:26:53basically called as Sunny this category
- 3:26:57is basically called as overcast and this
- 3:27:00category is basically called as rain
- 3:27:03based on these three categories so I'm
- 3:27:04splitting it now just go ahead and see
- 3:27:07in Sunny how many yes and how many no
- 3:27:10are there how many yes with respect to
- 3:27:12Sunny are there see in sunny I have two
- 3:27:14NOS see one and two no uh one more no is
- 3:27:18there three NOS so here you can see this
- 3:27:21is my one no then this is my two no this
- 3:27:25is my three no and yes are two so this
- 3:27:30one and this one so how many total
- 3:27:33number of yes so here you can see that
- 3:27:36there are 1 2 2 yes and three no let's
- 3:27:41say that I have randomly selected one
- 3:27:43feature which is Outlook why can't I
- 3:27:45when like see it is up to it it is up to
- 3:27:49the decision tree to select any of the
- 3:27:51feature here I have specifically taken
- 3:27:53Outlook later on I'll explain why it it
- 3:27:57can basically select how it selects the
- 3:27:59feature okay I'll I'll talk about it
- 3:28:00don't worry so in the Outlook we have
- 3:28:04two yes sorry in the case of Sunny we
- 3:28:06have two yes and three NOS now the next
- 3:28:08thing is that let's go and see for
- 3:28:10overcast in overcast I have 1 yes uh 2s
- 3:28:14um 3s and 4 yes I don't have any no in
- 3:28:18overcast so over here my thing will be
- 3:28:21that four yes and Zer Nos and then
- 3:28:24finally when we go to the Rain part see
- 3:28:26in Rain how many features are there in
- 3:28:29rain if you go and probably see it how
- 3:28:31many number of yes and NOS are there go
- 3:28:33and see in one one yes in row rain two
- 3:28:36yes then one no then again you have one
- 3:28:39yes and one no right so here you can
- 3:28:43basically say that in rain in the case
- 3:28:45of rain if I take a as an example how
- 3:28:47many number of yes and NOS are there it
- 3:28:49will be 3 yes and two
- 3:28:52NOS understand understanding
- 3:28:57algorithm then everything will you'll be
- 3:29:00able to understand now let's go ahead
- 3:29:03and try to cease for sunny sunny
- 3:29:05definitely has 2 yes and three NOS this
- 3:29:08has four yes and zero NOS here you have
- 3:29:10three Y and two NOS now if I probably
- 3:29:13take overcast here you need to
- 3:29:15understand understand about two things
- 3:29:17one is pure
- 3:29:18split and one is impure split now what
- 3:29:22does pure split basically mean pure spit
- 3:29:25basically means that now see in this
- 3:29:26particular scenario in overcast in
- 3:29:29overcast I have either yes or no so here
- 3:29:32you can see that I have four yes and Zer
- 3:29:35NOS so that basically means this is a
- 3:29:37pure split anybody tomorrow in my data
- 3:29:40set if I just take this Outlook feature
- 3:29:43suppose in one day in day 15 the Outlook
- 3:29:46is Outlook is basically overcast then I
- 3:29:50know directly it is the person is going
- 3:29:52to play so this part is already created
- 3:29:54and this node is called as pure
- 3:29:58node understand this why it is called as
- 3:30:00pure node because either you have all
- 3:30:03Yes or zeros NOS or zero yes or all NOS
- 3:30:08like that in this particular case I have
- 3:30:10all yes so if I take this specific path
- 3:30:13I know that with respect to overcast my
- 3:30:16final decision which is yes it is always
- 3:30:17going to become yes so this is what it
- 3:30:19basically says so I don't have to split
- 3:30:22further so from here I will probably not
- 3:30:25split I will definitely not split more
- 3:30:28because I don't require it because I
- 3:30:31have it is a pure leaf node okay you can
- 3:30:34also say that this is a pure leaf node
- 3:30:37so I'm just going to mention it again
- 3:30:39this one I'm specifically talking about
- 3:30:41now let's talk about sunny in the case
- 3:30:43of Sunny you have two yes and three NOS
- 3:30:45so this is obviously impure so what we
- 3:30:48do we take next feature and again how do
- 3:30:52we calculate that which feature we
- 3:30:54should take next I'll discuss about it
- 3:30:56let's say that after this I take up
- 3:31:00temperature I take up temperature and I
- 3:31:02start splitting again since this is
- 3:31:04impure okay and this split will happen
- 3:31:08until we get finally a pure split
- 3:31:11similarly with respect to rain we will
- 3:31:13go ahead and take another feature and
- 3:31:15we'll keep on splitting unless and until
- 3:31:18we get a leaf node which is completely
- 3:31:21pure I hope you understood how this
- 3:31:23exactly work now two questions two
- 3:31:27questions is that Kish the first thing
- 3:31:29is that how do we calculate this
- 3:31:32Purity and how do we come to know that
- 3:31:35this is a pure split just by seeing
- 3:31:38definitely I can say I can definitely
- 3:31:41say by just seeing that how many number
- 3:31:43of yes or NOS are there based on that I
- 3:31:45can def itely say it is a pure split or
- 3:31:47not so for this we use two different
- 3:31:50things one is
- 3:31:53entropy and the other one is something
- 3:31:55called as guine coefficient so we will
- 3:31:58try to understand how does entropy work
- 3:32:01and how does Guinea coefficient work in
- 3:32:04decision tree which will help us to
- 3:32:06determine whether the split is pure
- 3:32:09split or not or whether this node is
- 3:32:11leaf node or not then coming to the
- 3:32:13second thing okay coming to the second
- 3:32:16thing one is with respect to Purity
- 3:32:18second thing your first most important
- 3:32:20question which you had asked why did I
- 3:32:22probably select Outlook how the features
- 3:32:24are selected and here you have a topic
- 3:32:27which is called as Information Gain and
- 3:32:29if you know this both your problem is
- 3:32:32solved so now let's go ahead and let's
- 3:32:35understand about entropy or guinea
- 3:32:38coefficient or Information Gain entropy
- 3:32:40or guine coefficient oh sorry Guinea
- 3:32:42coefficient I'm saying guine impurity
- 3:32:44also you can say over here
- 3:32:46I'll write it as guine impurity not
- 3:32:48coefficient also I'll just say it as
- 3:32:50Guinea impurity but I hope everybody is
- 3:32:53understood till here let's go ahead and
- 3:32:55let's discuss about the first thing that
- 3:32:57is
- 3:32:58entropy how does entropy work and how we
- 3:33:01are going to use the formula so entropy
- 3:33:04here I will just write guine so we are
- 3:33:07going to discuss about this both the
- 3:33:09things let's say that the entropy
- 3:33:12formula which is given by I will write h
- 3:33:14of s is equal to so h of s is equal to
- 3:33:17minus P plus I'll talk about what is
- 3:33:20minus what is p plus log base 2 p
- 3:33:26+- p
- 3:33:28minus log base 2 p minus so this is the
- 3:33:32formula and in guine impurity the
- 3:33:34formula is 1 minus summation of I equal
- 3:33:391 2 N p² I even talk about when you
- 3:33:43should use guine impurity when you
- 3:33:44should not use guine impurity
- 3:33:46when you should use entropy you know by
- 3:33:48default the decision tree regression or
- 3:33:51classific sorry decision tree
- 3:33:53classification uses Guinea impurity now
- 3:33:56let's take one specific example so my
- 3:33:58example is that I have a feature one my
- 3:34:00root node I have a feature one which is
- 3:34:03my root node and let's say that in this
- 3:34:05root node I have six yes and three NOS
- 3:34:08very simple let's say that this has two
- 3:34:11categories and based on this two
- 3:34:13categories of split has happened that is
- 3:34:16a C1 let's say in this I have 3 S3 Nos
- 3:34:20and here I have 3 s0 Nos and this is my
- 3:34:24second category always understand if I
- 3:34:26do the sumission 3s and 3s is 6s see
- 3:34:30this this sumission if I do 3 + 3 is
- 3:34:33obviously 6 3 + 0 is obviously so this
- 3:34:36you need to understand based on the
- 3:34:38number of root nodes only almost it'll
- 3:34:40be same now let's go ahead and let's
- 3:34:44understand how do we Cal calculate let's
- 3:34:46take this example how do we calculate
- 3:34:48the entropy of this so I have already
- 3:34:50shown you the entropy formula over here
- 3:34:52now let's understand the components I
- 3:34:55will write h of s is equal to minus sign
- 3:34:59is there what is p+ p+ basically means
- 3:35:03that what is the probability of yes what
- 3:35:07is the probability of yes this is a
- 3:35:10simple thing for you all out of this
- 3:35:13what is the probability of yes yes out
- 3:35:16of this so obviously how you'll write if
- 3:35:19you want to find out the probability of
- 3:35:20yes out of this see when I say plus that
- 3:35:24basically means yes when I say minus
- 3:35:27that basically means no so what is the
- 3:35:29probability of yes so it is be nothing
- 3:35:32but yes plus and minus are specifically
- 3:35:35for binary
- 3:35:37class this can be positive negative so
- 3:35:40the probability with respect to yes can
- 3:35:42I write 3x 3 only for this what is the
- 3:35:45probability out of this total number of
- 3:35:48this is there 3x3 similarly if I go and
- 3:35:51see the next term log to the base 2 p+
- 3:35:54so again if I go ahead and write over
- 3:35:56here log to the base 2 p+ p+ is again
- 3:36:033x3 so then again we have minus and this
- 3:36:07is now P minus what is p minus 0 by 3
- 3:36:11log base 2 0 by 3 this obviously will
- 3:36:15become zero this will obviously become 0
- 3:36:18because 0 divid by anything is zero what
- 3:36:21will this be 1 log to the base 1 what is
- 3:36:25this this is nothing but zero log to the
- 3:36:28base 1 is nothing but zero tell me
- 3:36:31whether this is a pure split or impure
- 3:36:35split so this is a pure split whenever
- 3:36:38we have a pure split the answer of the
- 3:36:41entropy is going to come to zero so here
- 3:36:44I'm going to Define one graph
- 3:36:46this is H of s and let's say this is p+
- 3:36:49or P minus if my probability of plus see
- 3:36:53when I say probability of plus is 0.5
- 3:36:56what will be probability of minus it
- 3:36:57will also be 0. five right because it's
- 3:37:01just like P is equal to 1 - Q right if p
- 3:37:04is .5 then Q will be 1 - P same thing
- 3:37:07right so when it
- 3:37:09is5 obviously my h of s will be 1 let's
- 3:37:14say so this is this is the graph that
- 3:37:16will basically get formed let's go ahead
- 3:37:19and try to calculate the entropy of this
- 3:37:21guys what will be the entropy of this
- 3:37:24node so here I'm going to just make a
- 3:37:26graph h of s minus what is p+ p+ is
- 3:37:31nothing but 3x 6 log base 2 3x 6
- 3:37:37minus three no are there 3x 6 log base 2
- 3:37:433x 6 so if you compute this
- 3:37:46log base 2 to the^ of 1 if you do the
- 3:37:50calculation here I'm actually going to
- 3:37:52get one so when I'm getting one when I'm
- 3:37:55actually getting one when you have three
- 3:37:57yes and three NOS what is the
- 3:37:59probability it is 50/50% right so when
- 3:38:02your p+ is5 that basically means your h
- 3:38:06of s is coming as one so from this graph
- 3:38:09you can see that I'm getting one if this
- 3:38:11is zero this is one this is zero and
- 3:38:13this is one I hope everybody is able to
- 3:38:15to understand guys 0o and one if your p+
- 3:38:20is
- 3:38:21zero or if your p+ is one that basically
- 3:38:24means it becomes a pure split so in h of
- 3:38:26s you are going to get
- 3:38:29zero so always understand your entropy
- 3:38:33will be between 0 to
- 3:38:361 if I have a impure this is a
- 3:38:39completely impure split because here you
- 3:38:42have 50% probability of getting yes 50%
- 3:38:45probability of getting no h ofs is
- 3:38:48entropy this is entropy for the sample H
- 3:38:52ofs notation that I'm using is H ofs so
- 3:38:56if whenever the split is happening the
- 3:38:59first thing is done the purity test the
- 3:39:02purity test is done with the help of
- 3:39:04entropy right now I'll also show guinea
- 3:39:07guinea impurity don't worry so with the
- 3:39:09entropy you'll be able to find if I am
- 3:39:11getting one that basically means it is a
- 3:39:14impure split and if I'm getting zero it
- 3:39:18is pure split so this is the graph okay
- 3:39:22this is the graph and this graph is
- 3:39:24basically the entropy graph again
- 3:39:26understand if your probability of
- 3:39:28getting yes or no is 0.5 that basically
- 3:39:30means 50/50 is there 3s and three NOS
- 3:39:34then your entropy is going to be 1 h of
- 3:39:37s if your probability is completely one
- 3:39:39that basically means either you're
- 3:39:40getting completely yes or completely no
- 3:39:43so your your entropy will be zero that
- 3:39:46basically means it is pure split so in
- 3:39:48the case of probability .5 you're
- 3:39:50getting plus one then it'll keep on
- 3:39:52reducing now let's go ahead and let's
- 3:39:54try to understand so here you have
- 3:39:56understood about purity test definitely
- 3:39:58you'll use entropy try to find out
- 3:40:00whether it is pure or impure if it is
- 3:40:02impure you go ahead with the further
- 3:40:04shift further division of the categories
- 3:40:08again you take another feature divide it
- 3:40:10because here from this two which split
- 3:40:13you will do further you will do this
- 3:40:14split as further if you are getting 6 6
- 3:40:18is this specific value then you probably
- 3:40:20go and draw over here this is your
- 3:40:23entropy if your probability is here
- 3:40:25which
- 3:40:26is.3 then you will go here and create
- 3:40:29this this may be0 4 or3 something like
- 3:40:32this it will be between 0 to 1 let's go
- 3:40:35ahead and discuss about the second issue
- 3:40:37I hope everybody is discussed about we
- 3:40:40have discussed about checking the pure
- 3:40:42split or not and we have understood this
- 3:40:45much but the next thing is that okay
- 3:40:47fine chish this is very good we have
- 3:40:49explained well I know many people will
- 3:40:51say that but there are some people I
- 3:40:53can't help let's say that I have some
- 3:40:55features okay now coming to the second
- 3:40:58problem how do we consider which node to
- 3:41:02cap which which feature to take and
- 3:41:05split because here I may have one one
- 3:41:08split so again let's see that what is
- 3:41:10the second problem which feature to take
- 3:41:14to split right this is the second
- 3:41:16problem that we are trying to solve
- 3:41:18let's say that I have one feature one
- 3:41:19over here and I have two categories
- 3:41:22let's say this is there C1 and C2 here
- 3:41:25let's say that I have 9 years 5 Nos and
- 3:41:29then I have 6 years 2 NOS here I have
- 3:41:32basically three yes and three NOS let's
- 3:41:34say and in my data set I have features
- 3:41:36like F1 FS2 F3 now let's say that
- 3:41:40another split I can actually start with
- 3:41:42feature two also and in feature two I
- 3:41:45may have probably three categories like
- 3:41:47C1 C2 C3 so with respect to the root
- 3:41:52node and all the other features because
- 3:41:54after this also I may have to split
- 3:41:56right I may have to take another feature
- 3:41:58and keep on splitting right based on the
- 3:42:01Pure or impure split how do I decide
- 3:42:03should I take fub1 first or F2 first or
- 3:42:07F3 first or any other feature first how
- 3:42:10should I decide that which feature
- 3:42:12should I take and probably do the split
- 3:42:15that is the major question so for this
- 3:42:18we specifically use something called as
- 3:42:20Information Gain so here I'm just going
- 3:42:22to say here we basically use Information
- 3:42:26Gain now what is this Information Gain
- 3:42:29I'll talk about it so Information Gain
- 3:42:31first of all I will write the formula we
- 3:42:33basically write gain with sample first
- 3:42:37with feature one I will compute so first
- 3:42:40with feature one I will compute suppose
- 3:42:42this is my first split of my data and
- 3:42:44probably I'm Computing over here this
- 3:42:46can be written as h of s I'll discuss
- 3:42:50about each and every parameter don't
- 3:42:51worry summation of V belong to values s
- 3:42:56of V don't worry guys if you have not
- 3:42:58understood the formula I will explain it
- 3:43:01then the sample size H of SV I'll
- 3:43:04discuss about each and every parameter
- 3:43:06let's say that I'm taking this feature
- 3:43:09one split I have you have already seen
- 3:43:11what is feature one so this is my
- 3:43:13feature one I have two categories C1 C2
- 3:43:18this has 9 yes 5 NOS this has 6s and two
- 3:43:24Nos and this has 3 yes and three NOS now
- 3:43:27I will try to calculate the information
- 3:43:29gain of this specific split now I will
- 3:43:32go ahead and probably take this up now
- 3:43:35see over here we'll try to understand
- 3:43:37what is this now if I want to compute
- 3:43:40the gain of s of F1 first is first first
- 3:43:43thing that I need to find out is H of s
- 3:43:45now this h of s is specifically of the
- 3:43:48root node so I need to first of all
- 3:43:50calculate what is h of s h ofs is
- 3:43:52nothing but entropy entropy of the root
- 3:43:56node so if I want to compute the entropy
- 3:43:58of the node node tell me how should I
- 3:44:00compute h of s is equal to minus p + log
- 3:44:04base 2 p+ calculate guys along with me -
- 3:44:07P minus log base to P minus so I hope
- 3:44:11everybody knows this so here I'm going
- 3:44:13to compute by what is ability of plus
- 3:44:16over here in this specific root node it
- 3:44:18is nothing but 9 by4 then I have log
- 3:44:22base 2 again 9
- 3:44:24by4 then I have P minus what is p minus
- 3:44:285x4 log base 2 5 by4 so this calculation
- 3:44:34I will probably get it as
- 3:44:3694 approximately equal to 94 just check
- 3:44:40it whether you're getting this or not
- 3:44:42again you can use calculator if you want
- 3:44:44now now I have definitely found out this
- 3:44:47this is specifically for the root node
- 3:44:50now let's see the next thing the next
- 3:44:51important thing which is this part what
- 3:44:54is s of v and what is s and what is h of
- 3:44:57SV now very important just have a look
- 3:45:01everybody see this graph okay see this
- 3:45:05graph I will talk about h of SV first of
- 3:45:07all I'll talk about h of SV okay this
- 3:45:10one this is the entropy of category one
- 3:45:13you need to find and entropy of category
- 3:45:152 you need to find so if I write h of SV
- 3:45:19of category 1 so what is category 1 for
- 3:45:22this I'll write SC1 let's say I'm going
- 3:45:25to write like this quickly calculate the
- 3:45:28H of SV of this and this separately you
- 3:45:31need to calculate so h of SV of C1 okay
- 3:45:35so here again you'll write - 6X 8 log
- 3:45:38base 2 6X
- 3:45:418us 2x 8 log base to 2x 8 I hope
- 3:45:46everybody knows this how we got it so h
- 3:45:50of SV basically means I'm going to
- 3:45:51compute the entropy of this category and
- 3:45:54this category so for that I will
- 3:45:56basically write h of so here I will
- 3:45:59write - 6 by8 log base 2 6X 8 - 2x 8 log
- 3:46:08base 2 2x 8 so if I get it I'm actually
- 3:46:12going to get 81 and similarly if I if I
- 3:46:15calculate h of C2 quickly calculate how
- 3:46:18much you are going to get guys 6X 8 6X 8
- 3:46:21with respect to this we need to find out
- 3:46:24so now we have all these values we'll
- 3:46:25start equating them to this equation so
- 3:46:29here we have finally gain of s comma
- 3:46:33fub1 so let's say that here I'm going to
- 3:46:36basically add
- 3:46:3894 minus see minus summation of okay
- 3:46:42summation of what is s s of V understand
- 3:46:46s of V basically means that how many
- 3:46:48samples I have over here let's say for
- 3:46:51category one how many samples I have for
- 3:46:54category one over here simple if you
- 3:46:56really want to just calculate it is
- 3:46:58nothing but eight and total number of
- 3:47:01sample is how much if I go and see over
- 3:47:03here there are 9 years five NOS okay 9
- 3:47:07years and five NOS that basically means
- 3:47:1014 total sample here you have eight
- 3:47:13sample Okay so this will become
- 3:47:178x4 then you multiply by what see see
- 3:47:21from this equation you multiply by h of
- 3:47:23SV so h of SV is nothing but the entropy
- 3:47:26of category 1 so entropy of category 1
- 3:47:29is nothing but 81 plus then you go again
- 3:47:33back to the graph and try to see that
- 3:47:36for C2 how much how many total number of
- 3:47:39samples are there 3 + 3 is 6 so 6 by 14
- 3:47:42it will
- 3:47:43become multiplied by 1 right so this is
- 3:47:49your entire thing so here after all the
- 3:47:52calculation you are going to get
- 3:47:540.041 so this is my gain with s comma F1
- 3:47:59so here I have got this value amazing I
- 3:48:02did this with feature one only what
- 3:48:05about feature two let's say that this
- 3:48:07was my split for feature two and suppose
- 3:48:10I get the gain for S comma feature 2 as
- 3:48:17.51 if I get this now tell
- 3:48:21me in using which feature should I start
- 3:48:25splitting first whether it should be
- 3:48:28fub1 or whether it should be FS2 based
- 3:48:31on this value you know that over here
- 3:48:35the gain the information gain of s comma
- 3:48:38F2 is greater than gain of s comma fub1
- 3:48:43so your answer is very much simple we
- 3:48:45will definitely use feature 2 to start
- 3:48:48the split the thing over here you are
- 3:48:51trying to understand that if I really
- 3:48:52want to select which feature to select
- 3:48:54to start my splitting then I have to
- 3:48:58basically calculate the information gain
- 3:49:00and go throughout the all the paths and
- 3:49:03whichever path has the highest
- 3:49:04Information
- 3:49:05Gain then we will select that specific
- 3:49:09thing now the question Rises Kish
- 3:49:12obviously this is good but you had
- 3:49:14written about guinea impurity what is
- 3:49:16the purpose of that please explain us
- 3:49:19and why Guinea impurity is basically
- 3:49:20used so let me go ahead with guine
- 3:49:22impurity I told that yes you can
- 3:49:25obviously
- 3:49:26use you can obviously use entropy but
- 3:49:29why Guinea impurity so guine impurity
- 3:49:32formula which I have specifically
- 3:49:34written as 1 minus summation of IAL 1
- 3:49:382 N
- 3:49:41p² now what is this p² suppose let's say
- 3:49:45that in my n n is the number of outputs
- 3:49:47right now how many outputs I have I have
- 3:49:49two outputs yes or no so I will expand
- 3:49:52this 1 minus since this is summation I
- 3:49:55equal to 1 to n I'm basically going to
- 3:49:57basically say that okay fine I will
- 3:50:00write probability of plus whole
- 3:50:03Square uh plus probability of minus
- 3:50:07whole Square so this is the formula for
- 3:50:10guinea impurity now you may be thinking
- 3:50:14okay fine the calculation will be
- 3:50:16obviously very much equal easy right
- 3:50:18suppose if I have a node sorry if I have
- 3:50:21a node which which has 2 yes two NOS now
- 3:50:25in this particular case how do I
- 3:50:26calculate my this probability if I have
- 3:50:29two yes or two NOS suppose let's say
- 3:50:31that I have a node over here which is my
- 3:50:33split and this is having two yes and two
- 3:50:36no so how do I calculate I will write 1
- 3:50:38minus what is probability of square 1X 2
- 3:50:41square sorry not 1 by two
- 3:50:45yeah 1X 2 squ + 1 by 2
- 3:50:49squ right then I will say 1 by 1X 4 + 1X
- 3:50:544 is nothing but 2x 4 which is nothing
- 3:50:56but 1X 2 so I will be getting 0.5 now
- 3:51:00here here you understand this is a
- 3:51:02complete impure split right if you have
- 3:51:06an impure split in entropy the output
- 3:51:10you getting it as one whereas in the
- 3:51:13case of Guinea impurity
- 3:51:15as Z sorry
- 3:51:170.5 so if I go ahead with the graph that
- 3:51:21I probably had created here so my Guinea
- 3:51:24impurity line will look something like
- 3:51:27this so it will be looking something
- 3:51:29like this for zero obviously I'll be
- 3:51:31getting zero but whenever my probability
- 3:51:34of plus is 0.5 I'm going to get 0.5 over
- 3:51:38here and that is the difference between
- 3:51:40Guinea
- 3:51:42impurity and entropy but the re but you
- 3:51:45may be seeing Kish when to use what now
- 3:51:48let's understand that when to use Guinea
- 3:51:51and when to use entropy tell me guys if
- 3:51:55I consider this formula of guine
- 3:51:58impurity and if I probably
- 3:52:01consider if I consider entropy this
- 3:52:05formula where do you think more time
- 3:52:09will take for execution for this
- 3:52:11particular formula whether for entropy
- 3:52:14it will take or for guinea impurity it
- 3:52:18will take more time where it will
- 3:52:21probably take for the execution purpose
- 3:52:24see understand decision tree is having a
- 3:52:29worst time complexity because if you
- 3:52:32have 100 features probably you'll keep
- 3:52:34on comparing by dividing many many
- 3:52:37features then probably compute a
- 3:52:38Information Gain like this if you have
- 3:52:40just 100 features so which is faster
- 3:52:43entrop
- 3:52:45or guine impurity understand in entropy
- 3:52:48you have log function here you have log
- 3:52:52function here you have simple maths the
- 3:52:56more amount of time out of entropy and
- 3:52:59guine impurity the more amount of time
- 3:53:01basically is taken
- 3:53:03by
- 3:53:06entropy so if you have huge number of
- 3:53:10features like 100 200 features and you
- 3:53:12are planning to apply decision Tre I
- 3:53:15would suggest try to use Guinea impurity
- 3:53:18then entropy if you have small set of
- 3:53:20features then you can go ahead with
- 3:53:23entropy so over here definitely with
- 3:53:25respect to fast Guinea is greater than
- 3:53:31entropy now let's go ahead and
- 3:53:33understand with respect to you may be
- 3:53:36thinking Kish okay fine you have
- 3:53:38basically explained us about categorical
- 3:53:41variables over here see over here you
- 3:53:44have you have explained about
- 3:53:45categorical variables what if I have
- 3:53:47numerical feature let's say I have F1
- 3:53:51over here which is a numerical
- 3:53:53feature I have an F1 feature which is
- 3:53:56numerical feature and I may have values
- 3:53:58let's say that I have sorted all the
- 3:54:00values over here okay let's say that I
- 3:54:02have F1 and output okay so this F1 let's
- 3:54:06say that I have values
- 3:54:07like ass sorted order values I'm sorting
- 3:54:10this features I'm basically doing this
- 3:54:12let's say that initially I have this
- 3:54:15features like this and let's say I have
- 3:54:17values like 2.3 1.3 4 5 7 3 let's say I
- 3:54:23have this features now this is a
- 3:54:26continuous
- 3:54:27feature this is a continuous feature so
- 3:54:29for a continuous feature how probably
- 3:54:32the decision tree entropy will be
- 3:54:34calculated and the Information Gain will
- 3:54:37get calculated so here you'll be able to
- 3:54:39see that I will first of all sort these
- 3:54:41values so in F1 the decision tree will B
- 3:54:44basically first of all sort this values
- 3:54:45so I have 1.3 then you have 2.3 then you
- 3:54:49have four then you have three three then
- 3:54:53you have four then you have five and
- 3:54:55then you have six now whenever you have
- 3:54:57a continuous feature so how the
- 3:54:59continuous feature will basically work
- 3:55:01in this case first of all your decision
- 3:55:04tree node will say
- 3:55:06that we'll take this one only one first
- 3:55:10record and say that if it is less than
- 3:55:12or equal to 1.3
- 3:55:14okay if it is less than or equal to 1.3
- 3:55:16so you here you'll be getting two
- 3:55:18branches yes or no so yes and no
- 3:55:22definitely your output over here will be
- 3:55:25put over here right and then for the no
- 3:55:28here you'll be having another node over
- 3:55:30here how many number of Records you'll
- 3:55:31be having in this particular case you'll
- 3:55:33be having one record in this particular
- 3:55:35case you will be having around five to
- 3:55:36six records and here also you'll be able
- 3:55:38to see right how many yes and NOS are
- 3:55:40there definitely this will be a leaf
- 3:55:42node so in the first instance they will
- 3:55:45go ahead and calculate the information
- 3:55:47gain of this then probably once the
- 3:55:50Information Gain Is got then what
- 3:55:51they'll do they will take the first two
- 3:55:54records and again create a new decision
- 3:55:57tree let's say that this will be my
- 3:56:00suggestion where they'll say it is less
- 3:56:02than or equal to 2.3 so I will get one
- 3:56:05and one over here so in this now you'll
- 3:56:07be having two records which will
- 3:56:09basically say how many yes and no are
- 3:56:10there and remaining all records will
- 3:56:12come over here then again Information
- 3:56:16Gain will be computed here then again
- 3:56:17what will happen they'll go to the next
- 3:56:19record then then again they'll create
- 3:56:21another feature where they'll say less
- 3:56:22than or equal to three and they will
- 3:56:24create this many nodes again they'll try
- 3:56:28to understand that how many yes or no
- 3:56:29are there and then they'll again compute
- 3:56:31The Information Gain like this they'll
- 3:56:34do it for each and every record and
- 3:56:36finally whichever Information Gain is
- 3:56:38higher they will select that specific
- 3:56:40value in that feature and they'll split
- 3:56:42the node so in a continuous feature
- 3:56:45whenever you have a continuous feature
- 3:56:47this is how it will basically have and
- 3:56:50then it will try to compute who is
- 3:56:51having the highest Information Gain the
- 3:56:54best Information Gain will get selected
- 3:56:57and from there the splitting will
- 3:56:59happen now let's go ahead and understand
- 3:57:01about the next topic is that how this
- 3:57:04entirely things work in decision tree
- 3:57:07regressor because in decision tree
- 3:57:09regressor my output is an continuous
- 3:57:13variable so suppose if I have one
- 3:57:15feature one feature two and this output
- 3:57:17is a continuous feature it will be
- 3:57:20continuous any value can be there so in
- 3:57:23this particular case how do I split it
- 3:57:27so let's say that f1c feature is getting
- 3:57:30selected now in this f1c feature what
- 3:57:32value will come when it is getting
- 3:57:34selected first of all the entire mean
- 3:57:38will get calculated of the output mean
- 3:57:40will get calculated so here I will have
- 3:57:42the mean and here here the cost function
- 3:57:45that is used is not Guinea coefficient
- 3:57:48or guinea impurity or entropy here we
- 3:57:51use mean squared
- 3:57:53error or you can also use mean absolute
- 3:57:56error now what is mean squared error if
- 3:57:58you remember from our logistic linear
- 3:58:00regression how do we calculate 1 by 2 m
- 3:58:03summation of I = 1 to n y hat minus y
- 3:58:08whole Square y hat of i y - y whole
- 3:58:12Square this is what is mean square error
- 3:58:14so what it will do first based on F1
- 3:58:17feature it will try to assign a mean
- 3:58:20value and then it will compute the MSE
- 3:58:23value and then it'll go ahead and do the
- 3:58:26splitting now when it is doing splitting
- 3:58:29based on categories of continuous
- 3:58:31variable I will be having different
- 3:58:33different categories now in this
- 3:58:35categories what will happen after split
- 3:58:37some records will go over
- 3:58:40here then I will be having a mean value
- 3:58:42of this over here
- 3:58:45that will be my output and then again
- 3:58:47the MSC will get calculated over here as
- 3:58:50the msse gets reduced that basically
- 3:58:53means we are reaching near the leaf
- 3:58:55note and the same thing will happen over
- 3:58:57here so finally when you follow this
- 3:59:00path whatever mean value is present over
- 3:59:02here that will be your output this is
- 3:59:05the difference between the decision tree
- 3:59:06regressor and the classifier here
- 3:59:09instead of using entropy and all you use
- 3:59:12mean squar error or mean absolute error
- 3:59:14and this is the formula of mean square
- 3:59:16error now let's go to the one more topic
- 3:59:19which is called as the hyperparameters
- 3:59:22tell me decision tree if I keep on
- 3:59:25growing this to any depth what kind of
- 3:59:28problem it will face regressor part you
- 3:59:31want me to explain okay let's
- 3:59:33see okay let's let's do the
- 3:59:36regression decision
- 3:59:39tree
- 3:59:41regressor let's say I have feature F1
- 3:59:44and this is my output let's say I have
- 3:59:46values like 20 24 26 28 30 and this is
- 3:59:53my feature one with category one
- 3:59:56category one let's
- 3:59:58say some categories are there let's say
- 4:00:01I have done
- 4:00:03the division by
- 4:00:06F1 that is this feature initially tell
- 4:00:09me what is the mean of this that mean
- 4:00:12value will get assigned over here then
- 4:00:14using msse that is mean squar error here
- 4:00:18you will try to calculate suppose I get
- 4:00:20an msse of some 37 47 something like
- 4:00:23this and then I will try to split this
- 4:00:27then I will be getting two more nodes or
- 4:00:29three more nodes it depends then that
- 4:00:31specific nodes will be the part of this
- 4:00:33again the mean will change again the
- 4:00:36mean will change over here suppose this
- 4:00:38two is there this two records goes here
- 4:00:41right then again MC will get calculated
- 4:00:44I'm just taking as an example over here
- 4:00:46just try to assume this thing now if I
- 4:00:48talk about hyper parameters see this is
- 4:00:51what is the formula that gets applied
- 4:00:52over MSC now let's see in this hyper
- 4:00:56parameter always understand decision
- 4:00:58tree leads to overfitting because we are
- 4:01:00just going to divide the nodes to
- 4:01:03whatever level we want so this obviously
- 4:01:06will lead to
- 4:01:07overfitting now in order to prevent
- 4:01:10overfitting we perform two important
- 4:01:12steps one is post pruning and one is
- 4:01:16pre- pruning so this two post pruning
- 4:01:18and pre pruning is a condition let's say
- 4:01:21that I have done some
- 4:01:23splits I have done some splits let's say
- 4:01:26over here I have seven yes and two
- 4:01:28no and again probably I do the further
- 4:01:31split like this now in this particular
- 4:01:33scenario you know that if 7 yes and two
- 4:01:35NOS are there there is a maximum there
- 4:01:37is more than 80% chances that this node
- 4:01:40is saying that the output is yes so
- 4:01:43should we further do more
- 4:01:46pruning the answer is no we can close it
- 4:01:49and we can cut the branch from here this
- 4:01:52technique is basically called as post
- 4:01:54pruning that basically means first of
- 4:01:57all you create your decision tree then
- 4:01:59probably see the decision tree and see
- 4:02:01that whether there is an extra Branch or
- 4:02:03not and just try to cut it there is one
- 4:02:06more thing which is called as
- 4:02:07pre-pruning now pre-pruning is decided
- 4:02:10by hyperparameters what kind of hyper
- 4:02:13parameters you can basically say that
- 4:02:15how many number of decision tree needs
- 4:02:17to be used not number of decision tree
- 4:02:20sorry over here you may say that what is
- 4:02:22the max
- 4:02:24depth what is the max depth how many Max
- 4:02:27Leaf you can
- 4:02:28have so this all parameters you can set
- 4:02:31it with grid SE
- 4:02:33CV and you can try it and you can
- 4:02:36basically come up with a pre- pruning
- 4:02:38technique so this is the idea about
- 4:02:41decision tree uh regressor yes yes it is
- 4:02:44possible your guinea value will be one
- 4:02:45no this graph is there
- 4:02:47no Guinea value are you talking about
- 4:02:50this Guinea entropy it will not be one
- 4:02:51it will always be between 0
- 4:02:53to.5 so the first thing first as usual
- 4:02:57what we should do we should import the
- 4:02:59libraries so here I will go ahead and
- 4:03:02import the librar so I'll say
- 4:03:04import pandas as NP PD import matplot
- 4:03:10li. pyplot as PLT
- 4:03:14uh
- 4:03:16import so this basic things I have with
- 4:03:19me so I will go and take any data set
- 4:03:22that I want from SK
- 4:03:24learn. data sets import let's say that
- 4:03:28I'm going to take load Iris data set and
- 4:03:31then I'm going to upload the iris data
- 4:03:33set so I'm going to write load Iris
- 4:03:36there is my Iris data set then the next
- 4:03:38step uh once you get your iris data set
- 4:03:41so this is my iris. dat
- 4:03:45okay these are all my features the four
- 4:03:47features will be there these four
- 4:03:49features are petal length petal width
- 4:03:51SLE length and SLE width this is my
- 4:03:54independent features then if I really
- 4:03:56want to apply
- 4:03:58for classifier so decision tree
- 4:04:03classifier so I can first of all import
- 4:04:06from
- 4:04:08skarn do tree import decision let's see
- 4:04:13where decision tree present in a scalon
- 4:04:16decision tree
- 4:04:17classifier the name is absolutely fine
- 4:04:20but I was not getting over here
- 4:04:23so so this is got no module SK okay SK
- 4:04:29skar
- 4:04:31skn learn so here you have
- 4:04:35classifier right now I'm just going to
- 4:04:37overfit the data then I'll probably show
- 4:04:38you how you can go ahead with uh
- 4:04:42pruning so by default what are the
- 4:04:44parameters over here if you probably go
- 4:04:46and see in in the classifier over here
- 4:04:49you have Criterion see this the first P
- 4:04:52parameter is Criterion by default it is
- 4:04:54Guinea then you have Splitter Splitter
- 4:04:57basically means how you're going to
- 4:04:58split and there also you have two types
- 4:05:01best and random you can randomly select
- 4:05:04the features and do it okay you should
- 4:05:06always go with
- 4:05:07best max depth is a hyper parameter
- 4:05:11minimum sample lift is a hyper parameter
- 4:05:13Max Fe features how many number of
- 4:05:14features we are going to take in order
- 4:05:16to fix that that is also an hyper
- 4:05:17parameter so all these things are hyper
- 4:05:19parameter okay so I will just by default
- 4:05:22executed whatever is giving me in
- 4:05:24decision tree and the next thing that
- 4:05:26I'm actually going to do is create a
- 4:05:28decision tree so for this I will be
- 4:05:31using plot. fig size plot. figure inside
- 4:05:35figure I have this fix
- 4:05:38size okay and I will probably show in
- 4:05:41some better figure size so that
- 4:05:43everybody body will be able to see it so
- 4:05:45here let me say that I'm going to take
- 4:05:47an area of
- 4:05:491510 and then probably I'm going to say
- 4:05:51tree Dot
- 4:05:54Plot and here I'm going to say a
- 4:05:57classifier and it should be filled the
- 4:06:00coloring should be filled with this so
- 4:06:04tree sorry Tre Tre Tre Tre
- 4:06:09Tre it should be classifi tree. plot
- 4:06:12okay I have to also import uh tree so I
- 4:06:16have to basically import tree so from SK
- 4:06:20learn
- 4:06:22import three again I'm getting
- 4:06:26error has no attribute plot
- 4:06:29why let me just see the documentation
- 4:06:32guys so this plot function is like plot
- 4:06:34uncore tree dot tab plot _ tree now what
- 4:06:40is the error we are getting okay not
- 4:06:42fitted yet
- 4:06:44sorry so I'm going to say
- 4:06:47classifier do fit on data what data
- 4:06:53iris.
- 4:06:55data and then I'm going to fit with Iris
- 4:06:58dot
- 4:07:00Target so once this is done I think now
- 4:07:03it will get
- 4:07:04executed so this is how your graph will
- 4:07:07look like guys so here you can see this
- 4:07:10is how your graph looks like now if I
- 4:07:12show you the graph over here see you can
- 4:07:14see some amazing things over here three
- 4:07:18outputs are actually there in this when
- 4:07:21you see in this left hand side this
- 4:07:23become a leaf node so this first one is
- 4:07:25probably vers color uh versol flower
- 4:07:29okay if you go on the right hand side
- 4:07:31here you can see 50/50 is there so based
- 4:07:32on one feature based on one feature here
- 4:07:35you'll be able to see that you are
- 4:07:37getting a leaf node based on another
- 4:07:39Branch here you are getting
- 4:07:4105050 so again you have two more
- 4:07:44features getting splitted over here so
- 4:07:46here you have 495 here you have
- 4:07:48471 do we require this split anybody
- 4:07:51tell me from here do we require any any
- 4:07:54more split just try to think this is
- 4:07:56after post pruning I want to find out
- 4:07:59whether more splits are required or not
- 4:08:01now in this particular case you see this
- 4:08:03after this do you require any
- 4:08:05split you do not require right here you
- 4:08:08are basically getting 47 and one I guess
- 4:08:11after this also you require no split
- 4:08:14understand this so this is basically
- 4:08:15post pruning so you can then decide your
- 4:08:19level and probably do it gu value is
- 4:08:22more than
- 4:08:240.5 okay this side H this is coming as
- 4:08:290.5 greater than 0.5 it should not had
- 4:08:33here it is
- 4:08:340.5 no maximum .5 can come 0 to.5 only
- 4:08:39should come I don't know why this is
- 4:08:41coming as 667
- 4:08:44I'll have a look onto this guys but
- 4:08:47anywhere you see other than that you're
- 4:08:50everywhere you're getting less
- 4:08:51than5 the plotting graph is very much
- 4:08:54easy you use SK learn import tree then
- 4:08:57you basically do this get classify and
- 4:08:59field is equal to true and you can just
- 4:09:02do this so the agenda let me Define the
- 4:09:05agenda what all things are there first
- 4:09:08we'll understand about
- 4:09:11emble techniques in this assemble
- 4:09:13techniques we are basically going to
- 4:09:15discuss about what is the difference
- 4:09:17between
- 4:09:19bagging and boosting
- 4:09:22second what we are basically going to
- 4:09:24discuss about is so uh the agenda of
- 4:09:27this session is emble techniques bagging
- 4:09:29and boosting then we are probably going
- 4:09:31to cover random forest and then probably
- 4:09:35we will try to cover adab boost and if I
- 4:09:39have more energy I will also try to
- 4:09:40cover XG boost so all this Al lthms
- 4:09:43we'll discuss about it so let's go ahead
- 4:09:46and let's start the
- 4:09:48topics the first topic that we are going
- 4:09:50to discuss is about emble
- 4:09:52techniques now what exactly is emble
- 4:09:55techniques and we are going to discuss
- 4:09:58about it okay so emble techniques what
- 4:10:01exactly is emble techniques till now we
- 4:10:03have solved two different kind of
- 4:10:04problem statement one is
- 4:10:07classification and regression and you
- 4:10:09have learned about different different
- 4:10:11algorithms like uh linear regression
- 4:10:13logistic regression we have discussed
- 4:10:15about KNN we have discussed about
- 4:10:17yesterday what disc what did we discuss
- 4:10:19about n bias different different
- 4:10:21algorithms we have already finished now
- 4:10:24with respect to classification
- 4:10:25regression Problem whatever algorithm we
- 4:10:27are discussing there was only one
- 4:10:28algorithm at a time we were discussing
- 4:10:31one algorithm at a time we are
- 4:10:32discussing and we are trying to either
- 4:10:33solve a classification or a regression
- 4:10:35problem now the next thing is over here
- 4:10:38is that can we use multiple algorithms
- 4:10:42mul multiple algorithm to solve a
- 4:10:44problem multiple algorithms basically
- 4:10:46means can we I'll just talk about it
- 4:10:49okay now the if I ask this specific
- 4:10:52question can we use multiple algorithms
- 4:10:54to solve a problem at that point of time
- 4:10:57I will definitely say yes we can because
- 4:10:59we are going to use something called as
- 4:11:00emble techniques there now what this
- 4:11:03emble techniques is okay so emble
- 4:11:06techniques in emble techniques we
- 4:11:08specifically use two different ways one
- 4:11:12is one one way is that we specifically
- 4:11:15use and the other one I'll just go to
- 4:11:16write it over here so one that we
- 4:11:19basically use is something called as
- 4:11:20bagging technique and the other one we
- 4:11:23specifically use is something called as
- 4:11:25boosting technique so in bagging
- 4:11:27Technique we what exactly we can do and
- 4:11:31in boosting technique what we can
- 4:11:32actually do and how we are combining
- 4:11:34multiple models to solve a problem so
- 4:11:36let's first of all discuss about bagging
- 4:11:39now how does bagging work let's say that
- 4:11:42I have a specific data set so this is my
- 4:11:44data set with uh with features rows
- 4:11:48columns everything like this I have this
- 4:11:50specific data set just imagine I have
- 4:11:52many many features over here like this
- 4:11:54fub1 F2 F3 and probably I have my output
- 4:11:57so this is my data set D let's consider
- 4:11:59it now what we do in bagging is that we
- 4:12:04create models and this model can be
- 4:12:06anything it can be logistic it can be
- 4:12:08linear for a classification problem
- 4:12:10let's say that this is logistic model so
- 4:12:12this is my model M1 let's say I have
- 4:12:14another model M2 then I may have another
- 4:12:17model M3 let's say that this is
- 4:12:20logistic and this is probably the other
- 4:12:23model which is like decision tree and
- 4:12:25then probably we use this model as KNN
- 4:12:29classification and this model can again
- 4:12:31be decision tree it's fine let's use
- 4:12:34another decision tree so now here you
- 4:12:36can see that we have used so many models
- 4:12:39okay so many models are there now with
- 4:12:41respect to this particular model what I
- 4:12:42will do is that the first step that I
- 4:12:44will do from this particular data set I
- 4:12:46will just take up some rows so I'll
- 4:12:48basically do row
- 4:12:50sampling and I'll take a row sampling of
- 4:12:53D Dash D Das basically means this D Das
- 4:12:55is always less than D some of the rows
- 4:12:58I'll push it to M1 okay I can also use n
- 4:13:01fine so what I'll do is that some of the
- 4:13:03rows I'll push it to model one this
- 4:13:05model one will be training let's say
- 4:13:07that for out of this 10,000 record th000
- 4:13:09rows I'm actually doing a row sampling
- 4:13:11of th rows and giving it to M1 to train
- 4:13:14it then what I'm actually going to do
- 4:13:16over here I'm basically going to give
- 4:13:18this specific model M2 and again I'm
- 4:13:21going to do row row sampling and I'm
- 4:13:24again going to sample some of the rows
- 4:13:25and give it to model two and again
- 4:13:27remember some of the rows may get
- 4:13:29repeated from this D Dash to next dble
- 4:13:31Dash similarly I will do row sampling
- 4:13:33and give it to this and again I may have
- 4:13:35d triple Dash and D4 Dash so different
- 4:13:38different different different rows data
- 4:13:41points when I say row sampling basically
- 4:13:42I'm talking about data points different
- 4:13:45different data points I will give it to
- 4:13:47separate separate model and this model
- 4:13:49will specifically train when I say D
- 4:13:52Dash that basically means uh suppose I
- 4:13:54say th 10,000 are my total number of
- 4:13:56data points when I say D Dash This D
- 4:13:59Dash may be th000 points then D Double
- 4:14:02Dash may be another th000 points and
- 4:14:04some of the rows may get repeated over
- 4:14:05here dle Dash here also I can basically
- 4:14:08use so here specifically row sampling
- 4:14:10will be used now when I have this many
- 4:14:12specific each and every model will be
- 4:14:14trained with different kind of data now
- 4:14:17how the inferencing will happen for the
- 4:14:18test data so first thing first let's say
- 4:14:21that I'm going to get a new test data
- 4:14:23over here now new test data will be
- 4:14:25passed to M1 and this M1 suppose it
- 4:14:28gives zero as my output suppose let's
- 4:14:30say that I'm doing a binary
- 4:14:31classification it gives a Zer as an
- 4:14:33output so this is my output of zero next
- 4:14:37M2 for the new test data gives one M3
- 4:14:40gives one and M4 also gives one as the
- 4:14:43the output now in this particular case
- 4:14:46in this particular case what will happen
- 4:14:49now you can see over here it's simple
- 4:14:51what what do you think the output may be
- 4:14:53in this particular case now M1 has
- 4:14:55predicted for this particular test data
- 4:14:56as zero the model M2 has predicted 1 M3
- 4:15:00has predicted 1 and M4 has predicted one
- 4:15:02so finally all these outputs are going
- 4:15:04to get
- 4:15:06aggregated are going to get aggregated
- 4:15:08and a simple thing that gets applied is
- 4:15:11majority voting majority voting so tell
- 4:15:14me what will be the output for with
- 4:15:16respect to this the output will
- 4:15:18obviously be one because the majority
- 4:15:19voting that you can see three people are
- 4:15:21basically saying it as one so my output
- 4:15:24over here will be one okay this is the
- 4:15:26concept of bagging wherein you are
- 4:15:29providing different different rows with
- 4:15:31probably all the features in this case
- 4:15:33and giving it to different different
- 4:15:34model again which is a classification
- 4:15:36model and then finally you are combining
- 4:15:38them based on majority voting and you're
- 4:15:40getting the answer as one so this step
- 4:15:43is called as bootstrap aggregator that
- 4:15:45basically means you're aggregating all
- 4:15:48the output that is basically coming from
- 4:15:50all the specific models all the specific
- 4:15:52models now many people will say Krish
- 4:15:54what about Tai guys like this kind of
- 4:15:56situation you know we will be having
- 4:15:58more than 100 to 200 models so it is
- 4:16:01very very difficult that it will be a
- 4:16:03tie who are repeating questions they
- 4:16:05will be put up in time out so what if
- 4:16:09you're saying that if the 50% of model
- 4:16:12says yes 50% of our models says no
- 4:16:14always understand guys we will be having
- 4:16:17more than 100 to 200 plus models so in
- 4:16:19this particular case there will be high
- 4:16:21probability that always there will be a
- 4:16:23majority voting available it will always
- 4:16:25not be in that specific scenario so this
- 4:16:28was the concept about bagging now some
- 4:16:30people will be saying that Krish why are
- 4:16:31you using different different models
- 4:16:34guys I'm not discussing about random
- 4:16:35Forest over here random Forest uses only
- 4:16:37one type of model that is decision tree
- 4:16:39but if we think as an concept of bagging
- 4:16:43you can have different different models
- 4:16:44over here and you can basically combine
- 4:16:46them so this is a technique of emble
- 4:16:49techniques and this is basically called
- 4:16:51as bagging okay now tell me one point I
- 4:16:54missed out fine this is with respect to
- 4:16:56the classification problem with respect
- 4:16:58to the regression problem what will
- 4:17:00happen in case of a regression problem
- 4:17:02let's say that I got here 120 here 140
- 4:17:06here 122 here 148 as my output so in
- 4:17:09regression what will happen is that the
- 4:17:11entire mean will be taken mean will be
- 4:17:15taken the output mean will be basically
- 4:17:18taken and that will be your output of
- 4:17:20the model average or mean very simple
- 4:17:22right so average or mean will be
- 4:17:25basically taken up and here based on the
- 4:17:27average you'll be able to solve the
- 4:17:29regression problem great now let's go
- 4:17:31ahead and try to understand with respect
- 4:17:34to bagging and boosting how many
- 4:17:36different types of algorithm are but
- 4:17:37before that I need to make you
- 4:17:39understand what exactly is boosting now
- 4:17:41here in bagging you have seen that you
- 4:17:43have parallel models right one one one
- 4:17:46independent you have parallel models
- 4:17:48you're giving some row samples in
- 4:17:49different different models and basically
- 4:17:51are able to find out the output now in
- 4:17:53case of boosting boosting is a
- 4:17:56sequential combination of models like
- 4:17:59this you have lot of sequential models
- 4:18:03like this and one after the model like
- 4:18:06first I'll give my training data to this
- 4:18:07particular model then it will go to this
- 4:18:09data then this model then this model so
- 4:18:12this will be my M1 M2 M3 M4 and finally
- 4:18:16I will be getting my output so here you
- 4:18:18can basically say that boosting is all
- 4:18:21about and this M1 M2 M3 we basically
- 4:18:24mention it as weak Learners so this will
- 4:18:26be weak learner weak learner weak
- 4:18:29learner weak learner and finally when we
- 4:18:32go till here it it'll if I combine all
- 4:18:35these weak ners weak
- 4:18:38learner weak learner okay once I combine
- 4:18:41all this weak learner it becomes a it
- 4:18:43becomes a strong learner finally if I
- 4:18:46try to combine this this will basically
- 4:18:47become a strong learner so here you have
- 4:18:50all the models sequentially one after
- 4:18:52the other and then you will probably try
- 4:18:55to provide your uh input from one model
- 4:18:58to the next model to the next model and
- 4:19:00these all models will be a very simpler
- 4:19:01weak learner model which will not be
- 4:19:03able to predict properly but when you
- 4:19:05combine all this particular models
- 4:19:08together sequentially it becomes a
- 4:19:09strong learner how this specifically
- 4:19:11works I'll take an example example of AD
- 4:19:13boost XG boost I will show you that okay
- 4:19:16week learner basically means the
- 4:19:17prediction is very bad but as you go
- 4:19:19sequentially you combine them they
- 4:19:21become a strong learner okay one example
- 4:19:24I want to give you let's say that you
- 4:19:26are a data scientist right let's say
- 4:19:30that this model one may be a teacher
- 4:19:33with respect to physics then this model
- 4:19:35two may be a teacher with respect to
- 4:19:37chemistry let's say model 3 is basically
- 4:19:40a teacher of maths and model four is a
- 4:19:43teacher of geography now suppose if you
- 4:19:46are trying to solve one problem
- 4:19:48obviously if the physics teacher is not
- 4:19:50able to solve that particular problem
- 4:19:51then probably chemistry can help or
- 4:19:54maths can help or geography can help or
- 4:19:56someone can help so when we combine this
- 4:19:58many expertise together they will be
- 4:20:01able to give you the output in an
- 4:20:03efficient way Sumit I'll talk about it
- 4:20:05where whether all the features are
- 4:20:07basically passed to all the models or
- 4:20:08not I'll just talk about it just give me
- 4:20:10some time okay but I just want to give
- 4:20:12you an idea about in short if someone
- 4:20:14asks you in an interview what exactly is
- 4:20:17boosting okay boosting is you can just
- 4:20:21say that it is a sequential set of all
- 4:20:23the models combined together and these
- 4:20:25all models that I initialized are
- 4:20:27usually weak Learners and when they are
- 4:20:29combined together they become a strong
- 4:20:30learner and based on the strong learner
- 4:20:32they gives an amazing output and right
- 4:20:35now if I say in most of the kaggle
- 4:20:37competition they use different types of
- 4:20:39boosting or bagging technique so we have
- 4:20:42basically as I said
- 4:20:44bagging and boosting in bagging what
- 4:20:47kind of algorithm we specifically use we
- 4:20:49use something called as random forest
- 4:20:54classifier and the second model that we
- 4:20:57specifically use is something called as
- 4:20:59random
- 4:21:00Forest regress so we specifically use
- 4:21:04these two kind of models which I'm
- 4:21:05actually going to discuss right now
- 4:21:06after this and then in boosting we
- 4:21:09basically use techniques like ad boost
- 4:21:12gradi Boost number three is Extreme
- 4:21:15gradient boost which we also say it as
- 4:21:17XG boost extreme gradient boost so let's
- 4:21:20go ahead and let's discuss about the
- 4:21:22first algorithm which is called as
- 4:21:24random forest classifier and regressor
- 4:21:28now first thing first let's understand
- 4:21:31some things from the yesterday's class I
- 4:21:33hope uh what is the main problem with
- 4:21:35respect to decision tree whenever we
- 4:21:37create a decision tree without any
- 4:21:39hyperparameter it does it not lead to
- 4:21:42overit
- 4:21:43does it not lead to overfitting uh
- 4:21:45whenever you probably have a decision
- 4:21:48tree right it leads to something like
- 4:21:50overfitting why overfitting because it
- 4:21:53completely splits all the feature till
- 4:21:55it's complete depth overfitting
- 4:21:57basically means for training data the
- 4:21:58accuracy is high for test data the
- 4:22:00accuracy is low so training data when
- 4:22:02the accuracy is high I may basically say
- 4:22:04it as high bias and then I may basically
- 4:22:07say it as sorry not high bias low bias
- 4:22:11and high V variance so low bias and high
- 4:22:14variance yes obviously we can do pruning
- 4:22:16and all guys but again understand
- 4:22:18pruning is an extensive task probably if
- 4:22:21your if you have 100 features if you
- 4:22:23have data points which is like 1 million
- 4:22:25to do pruning also it is very much
- 4:22:27difficult yes pre pruning can be done
- 4:22:29but again we cannot confirm that it may
- 4:22:31work well or not so right now with
- 4:22:33respect to decision tree you have this
- 4:22:35specific problem that is low bias and
- 4:22:37high variance now in low Biance and high
- 4:22:39variance you know that my model is
- 4:22:41basically the generalized model that I
- 4:22:43should get it should have low bias and
- 4:22:46low variance so if somebody asks you why
- 4:22:49do you use random Forest you can
- 4:22:51basically explain about decision trees
- 4:22:52like this now my main aim is to convert
- 4:22:54this High variance to low variance now I
- 4:22:58will be able to convert this High
- 4:22:59variance to low variance using random
- 4:23:01forest classifier or random Forest
- 4:23:03regressor now what does random Forest do
- 4:23:06random Forest is a bagging technique
- 4:23:08similarly I have a data set over here
- 4:23:10let's say that I have this data set
- 4:23:13and then here I will be having multiple
- 4:23:15models like
- 4:23:16M1
- 4:23:19M2
- 4:23:21M3 M4 let's say I have this four models
- 4:23:24like this we have many many models now
- 4:23:27with respect to this models this models
- 4:23:29all the models are actually decision
- 4:23:31Tree in random forest all are decision
- 4:23:34trees you don't have a different model
- 4:23:37over there so over here you can see that
- 4:23:39all the models are decision trees that
- 4:23:41is going to get used used in random
- 4:23:43Forest so decision trees always gets
- 4:23:45used in random Forest the first thing
- 4:23:47that you should know now whenever we are
- 4:23:49using decision trees you know that
- 4:23:51decision tree if I by default if we try
- 4:23:53to create it it may lead to overfitting
- 4:23:56and because of that every decision tree
- 4:23:58will basically create low V low bias and
- 4:24:01high variance but if we combine in the
- 4:24:04form of bootstrap aggregator this High
- 4:24:07variance will be getting converted to
- 4:24:08low variance because why because
- 4:24:10majority of voting we will be taking
- 4:24:12from this particular decision trees like
- 4:24:14there will be many many decision tree so
- 4:24:16they lot of outputs will be coming and
- 4:24:19with the help of majority voting
- 4:24:20classifier this High variance will get
- 4:24:22converted to low variance now in random
- 4:24:24Forest how it works in the first case if
- 4:24:27I talk about random Forest over here two
- 4:24:29things basically happen with respect to
- 4:24:30the D- data set let's say in first model
- 4:24:34we do some kind of row
- 4:24:36sampling plus
- 4:24:38Feature Feature
- 4:24:40sampling that basically means we have to
- 4:24:42select some set of rows and some set of
- 4:24:45features and give it to M1 similarly you
- 4:24:48do row sampling and feature sampling and
- 4:24:50give it to M2 then you do row sampling
- 4:24:52and feature sampling you give it to M3
- 4:24:54and then you do row sampling and feature
- 4:24:56sampling you give it to M4 now when you
- 4:25:00do this so what will happen
- 4:25:01independently you're giving some
- 4:25:03features along with some rows now there
- 4:25:05may be a situation that your features
- 4:25:07may also get repeated it may also get
- 4:25:09repeated your records or data points may
- 4:25:11also get repeated so when you are
- 4:25:13probably training your model with this
- 4:25:15specific data sets and specific features
- 4:25:18this model become expert in predicting
- 4:25:20something right as I said one example
- 4:25:23over here I'm giving a physics model
- 4:25:25some data I'm giving chemistry data
- 4:25:27chemistry model with some data similarly
- 4:25:29here I'm giving some information to some
- 4:25:31model so the model will be an expert
- 4:25:33with respect to that specific data So
- 4:25:36based on all this particular data
- 4:25:38whenever I get a new test data so what
- 4:25:40will happen suppose let's say that this
- 4:25:42this is a classification problem the M1
- 4:25:44model will be predicting zero this will
- 4:25:46be predicting one this will be
- 4:25:47predicting zero and this will be
- 4:25:49predicting zero now in this particular
- 4:25:51case again the majority voting
- 4:25:53classifier or majority voting will
- 4:25:55happen in the case of classification
- 4:25:57problem and then here you will be
- 4:26:01specifically able to get the output as
- 4:26:03zero so I hope everybody is able to
- 4:26:06understand all the models over here are
- 4:26:07decision trees and based on that you
- 4:26:10will be doing see when in I interview
- 4:26:12should be very very uh things the things
- 4:26:15that I'm telling you over here is all
- 4:26:17all the points are very much important
- 4:26:19and similarly if you tell the
- 4:26:21interviewer definitely your interview is
- 4:26:22cracked in this kind of algorithm I've
- 4:26:25seen some of my students saying that
- 4:26:26okay uh Kish um when the interviewer
- 4:26:29asked me that which is my favorite
- 4:26:30algorithm I said random Forest I told
- 4:26:32why did you say like that because he
- 4:26:34said that because that person let me let
- 4:26:36him ask any questions in random Forest
- 4:26:38I'm very much confident about it and I'm
- 4:26:40also going to prove him you know
- 4:26:42why they are very very good so with this
- 4:26:45specific case here you can basically see
- 4:26:47that because of the overfitting
- 4:26:49condition of the decision tree you're
- 4:26:50combining multiple decision tree so that
- 4:26:52you get a generalized model which has
- 4:26:54low bias and low variance so I hope
- 4:26:57everybody is able to understand boost
- 4:26:59feature sampling basically means suppose
- 4:27:00if I have 1 2 3 four feature for the
- 4:27:04first model I may give two features for
- 4:27:06the second model I may get three
- 4:27:07features for the fourth model I may give
- 4:27:09four features or uh any one feature ALS
- 4:27:12I can give to a specific model so
- 4:27:13internally that random Forest it take
- 4:27:15carees of over here these things are
- 4:27:18there and this is how random Forest
- 4:27:19Works only the difference between random
- 4:27:21Forest classify and regression is that
- 4:27:23in regression again whatever output you
- 4:27:25are basically getting you basically do
- 4:27:26the mean that's it average you just do
- 4:27:29the average you'll be able to get the
- 4:27:31output based on all the models output
- 4:27:33that you are actually getting now let's
- 4:27:34talk about some of the important points
- 4:27:36in random Forest the first thing first
- 4:27:38question is that is normalization
- 4:27:41required in random Forest then the next
- 4:27:43question is that in KNN is normalization
- 4:27:47when I say normalization or
- 4:27:49standardization I I'll just talk about
- 4:27:51standardization is standardization is
- 4:27:54required so this will be my another
- 4:27:56question so is normalization required in
- 4:27:59random forest or decision tree you here
- 4:28:01you can also say it as decision tree is
- 4:28:03it required so for this the answer will
- 4:28:06be no because understand decision tree
- 4:28:09will basically do the splits if you Mini
- 4:28:12minimize the data also that split won't
- 4:28:14be that much important but if I talk
- 4:28:17about KNN whether standardization
- 4:28:19normalization required over here the
- 4:28:21answer is yes because here we use two
- 4:28:23things one is ukan distance and
- 4:28:26Manhattan distance because of this you
- 4:28:28definitely have to apply standardization
- 4:28:30so that the computation or distance
- 4:28:32becomes easy so this is one of the most
- 4:28:34common interview questions that is
- 4:28:36basically asked in random Forest coming
- 4:28:38to the third question is random Forest
- 4:28:40impacted by outlier
- 4:28:43over here the answer will be no just
- 4:28:46check it out outside basically means
- 4:28:48Google and check it out check it out in
- 4:28:50Google okay perfect so I hope I've
- 4:28:53covered most of the things in random
- 4:28:54Forest is random Forest impacted by
- 4:28:57outliers this is the third question is
- 4:28:59KNN impacted by
- 4:29:00outliers is this KNN algorithm impacted
- 4:29:04by outliers is KNN impacted Byers the
- 4:29:07answer is yes big yes perfect so so
- 4:29:12these all are the interview questions
- 4:29:13that needs to be covered now let's go
- 4:29:15ahead and discuss about adab boost now
- 4:29:18in bagging most of the time we
- 4:29:20specifically use random forest or you
- 4:29:23can also create custom bagging
- 4:29:25techniques custom bagging techniques
- 4:29:27means whatever algorithm you want use
- 4:29:29the combination of them and try to give
- 4:29:32the output this also you can do it
- 4:29:33manually with the help of hands okay
- 4:29:36guys so second thing uh we are going to
- 4:29:38discuss about is boosting technique in
- 4:29:40this
- 4:29:42the first thing that uh first algorithm
- 4:29:44that we are going to discuss about is
- 4:29:45adab Boost so adab boost we going to
- 4:29:48discuss about how does adab Boost uh
- 4:29:50work now let's solve uh the first
- 4:29:53boosting technique which is called as
- 4:29:54adab boost okay and uh this is a
- 4:29:57boosting technique um in the boosting
- 4:30:00technique you have heard that we have to
- 4:30:02basically solve in a sequential way this
- 4:30:05at least you know I know there is a lot
- 4:30:07of confusion within you all but we'll
- 4:30:09try to solve a problem let's say so
- 4:30:11suppose I have a data set which looks
- 4:30:12like this fub1 F2 F3 F4 so these are my
- 4:30:16features and probably these are my
- 4:30:18output okay so let's say that I'm having
- 4:30:20this features like this and this is my
- 4:30:22output like yes or no like this so let's
- 4:30:25say that how many records I have over
- 4:30:27here three
- 4:30:304 5 6 and one more is there 7 so this
- 4:30:36seven records are there now in adab
- 4:30:38boost the first thing is that
- 4:30:40specifically with adab Boost uh you
- 4:30:42really need to understand that what all
- 4:30:43things we can basically do how do we
- 4:30:45solve this classification problem that
- 4:30:47we are going to understand the first
- 4:30:49thing first is that we Define a weight
- 4:30:51and the weight is very much simple
- 4:30:53initially to all the records to all this
- 4:30:55input records we provide an equal weight
- 4:30:58now how do we provide an equal weight we
- 4:30:59just go and count how many number of
- 4:31:01records are there now in this particular
- 4:31:03case the total number of records are one
- 4:31:062 3 4 5 6 7 now every record I have to
- 4:31:12provide an equal weight that is between
- 4:31:150 to 1 so the overall sum should be one
- 4:31:19so in this particular case what I can do
- 4:31:20if I make 1X 7 1X 7 1X 7 to everyone
- 4:31:24this will definitely become
- 4:31:26a equal weights to all right and if I do
- 4:31:30the total sum it will obviously be one
- 4:31:32let's go to the next one now after this
- 4:31:34what do we do okay after this in adab
- 4:31:37the first thing that we do is that we
- 4:31:39take any of this feature how do you
- 4:31:41decide which feature to take whether we
- 4:31:42should go with F1 or whether we should
- 4:31:44go with FS2 or whether we should go with
- 4:31:46F3 this we can do it with the help of
- 4:31:49Information Gain and Information Gain
- 4:31:53and entropy or guinea right based on
- 4:31:56this we can definitely understand
- 4:31:57whether we should start making decision
- 4:31:59here also you specifically make decision
- 4:32:01trees so here what you do is that you
- 4:32:04probably have to determine by using
- 4:32:06which feature I have to start my
- 4:32:07decision tree so suppose out of all this
- 4:32:09feature one feature two feature three
- 4:32:11you have selected that okay the
- 4:32:12information gain and entropy of feature
- 4:32:13one is higher so I'm going to use
- 4:32:15feature one and probably divide this
- 4:32:17into decision trees now when I divide
- 4:32:21this into decision tree let's say that
- 4:32:22I'm dividing like this into decision
- 4:32:23tree this decision tree depth will be
- 4:32:26only one one depth and this depth since
- 4:32:29it has only one depth we basically call
- 4:32:31it as stumps so what we do over here
- 4:32:34specifically we will create a decision
- 4:32:36Tre by taking only one feature and we
- 4:32:37will only divide it to one level okay
- 4:32:39one level or one depth that's
- 4:32:42and this is specifically called as stump
- 4:32:45what we are going to do next is that
- 4:32:46from this particular stump okay the
- 4:32:48stump is basically getting created only
- 4:32:51one so that is adab Boost right we say
- 4:32:52it as weak Learners because this is weak
- 4:32:54learner weak learner why there is a
- 4:32:57reason we say this as weak learner so
- 4:33:00only weak learner so that is the first
- 4:33:02thing with respect to uh this particular
- 4:33:06adab boost so the first step is that
- 4:33:07this is a weak learner so for the weak
- 4:33:09learner we basically create a stump
- 4:33:12stump basically means one level decision
- 4:33:14tree that's it based on the information
- 4:33:17gain and entropy I have selected the
- 4:33:18feature and then I just made a decision
- 4:33:21tree with only one level why it is
- 4:33:24called as it is called as weak learner
- 4:33:27okay so that is the reason we use only
- 4:33:28stum that is just a one level decision
- 4:33:31tree now the next step happens is that
- 4:33:33we provide all the specific records to
- 4:33:36this F1 and we train this specific model
- 4:33:39only with one level decision tree we
- 4:33:41train them
- 4:33:42now after we train them let's say that
- 4:33:44we are going to pass all these
- 4:33:45particular records to find out how many
- 4:33:47are correct and how many are wrong this
- 4:33:49decision this decision tree is basically
- 4:33:51giving so let's say that out of this
- 4:33:53entire records one
- 4:33:55record one record was just given as
- 4:33:59wrong let's say that this is the this is
- 4:34:01the record which was given as wrong okay
- 4:34:04so let's say that this record output was
- 4:34:07predicted wrong from this particular
- 4:34:09model only one wrong was there after
- 4:34:11training the model now what we need to
- 4:34:14do in this specific case understand a
- 4:34:16very important thing so let's say that
- 4:34:18we have done this and probably after
- 4:34:20this what we are actually going to do we
- 4:34:22are going to calculate the total error
- 4:34:24so how many error this particular model
- 4:34:26made let's say that in this particular
- 4:34:28case only one was wrong so this was only
- 4:34:31wrong right one was wrong so if I want
- 4:34:35to calculate the total error how will I
- 4:34:37calculate how many how many of them are
- 4:34:39wrong how many of them are wrong only
- 4:34:40one is wrong what is the weight of this
- 4:34:42so I will go and write 1X 7 so this is
- 4:34:45specifically my total error out of this
- 4:34:47specific model which is my stump over
- 4:34:49here okay which is my F1 stop now this
- 4:34:53is my first
- 4:34:54step the second step is that I need to
- 4:34:57see the performance of stump which stump
- 4:34:59this specific stump and the performance
- 4:35:02is basically checked by a formula which
- 4:35:04is 1 by log e 1us total error divided
- 4:35:09total error why we are doing this
- 4:35:11everything will make sense okay in just
- 4:35:13time every every in just a small time
- 4:35:16everything will make sense the first
- 4:35:18step that we do in adaab boost is that
- 4:35:20we try to find out the total error the
- 4:35:22second step we try to find out the
- 4:35:24performance of stump now in this
- 4:35:26particular case it will be 1 by log e 1
- 4:35:29- 1 by 7 / 1X 7 so once I calculate it
- 4:35:35it will be coming as
- 4:35:37895 F2 and F3 see again understand out
- 4:35:42of all these features I found out from
- 4:35:43Information Gain and entropy that this
- 4:35:45is the best feature let's say that I
- 4:35:47have calculated this
- 4:35:49as895 so this is my second step the
- 4:35:51first step is find out the total error
- 4:35:53the second step is performance of stum
- 4:35:55what is te te basically means total
- 4:35:57error te basically means total error now
- 4:36:01see see the steps okay see the steps
- 4:36:03whenever I'm discussing about boosting
- 4:36:05I'm going to combine weak Learners
- 4:36:07together to get a strong learner now
- 4:36:09what is the next step out of this now
- 4:36:11what what will be my third step
- 4:36:13understand over here my third step will
- 4:36:16be to update all these weights and that
- 4:36:19is the reason why I'm calculating this
- 4:36:20total error and performance of Step so
- 4:36:23my third step will basically be new
- 4:36:26sample weight from the decision tree one
- 4:36:29which is my stump so I'll say new sample
- 4:36:32weight is equal to I need to update all
- 4:36:34these weights why I need to update all
- 4:36:36these weights again understand I'll I'll
- 4:36:39talk about it just a second so if I want
- 4:36:41to up update the sample weights first
- 4:36:44update I will do it for correct records
- 4:36:46see for correct records whichever are
- 4:36:49correct like these all records are
- 4:36:51correct these all records are correct
- 4:36:53now when I update the weights of this
- 4:36:55update the weights of this particular
- 4:36:57record it should reduce and when the the
- 4:37:00the wrong records that I have this
- 4:37:02update should increase why because
- 4:37:06because if I increase this weights then
- 4:37:08the wrong records that are there that
- 4:37:11record should go to the next week
- 4:37:12learner that is the reason why I'm doing
- 4:37:14it now how to update this particular
- 4:37:17weights for correct records for correct
- 4:37:19records the formula looks something like
- 4:37:21this weight multiplied by weight
- 4:37:25multiplied by E to the^ of
- 4:37:28minus this specific performance okay
- 4:37:31this specific performance so e to the
- 4:37:33power of PS I'll write performance of
- 4:37:35stump and then I will basically be able
- 4:37:38to write 1X 7 * e to the^ of minus
- 4:37:43895 if I do the calculation everybody
- 4:37:45try to do it the answer will be
- 4:37:4805 now this is for correct records what
- 4:37:50about incorrect records for the
- 4:37:52incorrect
- 4:37:53records the the weights that is going to
- 4:37:56the formula that we going to apply is
- 4:37:58multiplied by E to the^ of plus PS not
- 4:38:02minus PS plus PS so here I'll write 1 by
- 4:38:057 multiplied e to the^ of
- 4:38:08895 so if I go and probably calcul this
- 4:38:12I'm going to get it
- 4:38:13as 349 so this two are the weights that
- 4:38:18I have got that basically means all
- 4:38:20these records now which are correct 1X 7
- 4:38:23the new updated weights will be 05 05
- 4:38:2805
- 4:38:3005 sorry not for the wrong
- 4:38:33records then this will be 05 then 05 and
- 4:38:3805 so let me just see what is 1x 7 so
- 4:38:41here you can see initially it was. 142
- 4:38:45now it has got reduced to 05 because all
- 4:38:47these records are correct but the wrong
- 4:38:50record value is 349 so my weights will
- 4:38:53now become over here as 349 now I will
- 4:38:56just go and go ahead and write over here
- 4:38:58my new weight my new weight is nothing
- 4:39:01but 05
- 4:39:06055
- 4:39:0705 05 05 1 2 how many 1 2 3 okay fourth
- 4:39:14record is here fourth record is there 1
- 4:39:182 3 4 05 05 okay how many records are
- 4:39:22there 1 2 3 4 5 6 7 so my fourth record
- 4:39:27will basically become the new value that
- 4:39:30I'm having is something called as
- 4:39:34349 now tell me guys if I do the
- 4:39:37summation of all these weights is this
- 4:39:39is it one so prob
- 4:39:41no I don't think so it is one because if
- 4:39:44I try to add it up it is not one but if
- 4:39:46I go and see over here these all are one
- 4:39:48if I combine all the things 1 2 3 4 5 6
- 4:39:517 these all are one so here I have need
- 4:39:53to find out my normalized weight now in
- 4:39:55order to find out the normalized
- 4:39:57weight all I have to do is that what I
- 4:40:00have to do because the entire sumission
- 4:40:03should be one so we have to
- 4:40:05normalize now in order to normalize all
- 4:40:08you have to do is that go and find out
- 4:40:10what is the sum of all this things the
- 4:40:12summation of all these things will be
- 4:40:150 649 all you have to do is that divide
- 4:40:18all the numbers
- 4:40:20by 649 divided by
- 4:40:24649
- 4:40:26649 like this divide all the numbers by
- 4:40:28649 and tell me what will be the answer
- 4:40:30that you'll be getting so here your
- 4:40:32normalized weight will now look like
- 4:40:35077 07 and this value will be somewhere
- 4:40:39around uh
- 4:40:41537 I guess in this case then this will
- 4:40:44be 07
- 4:40:47077 here we are going to divide by all
- 4:40:50this 64 649 now this is my normalized
- 4:40:53weight now after you get a normalized
- 4:40:56weight we will try to create something
- 4:40:57called as buckets because see one
- 4:41:00decision tree we have already created
- 4:41:02which is a stump and you know from this
- 4:41:04particular stum what you're going to get
- 4:41:06okay as an output then in the sequential
- 4:41:09model we will go and combine another
- 4:41:11model over here now it's the time that I
- 4:41:13have to create this specific model now
- 4:41:15in order to create this specific model I
- 4:41:17need to provide some specific rows only
- 4:41:19to this model to train because this
- 4:41:21model is giving one wrong now what I
- 4:41:24have to do is that whatever is wrong
- 4:41:26along with other data points I need to
- 4:41:28provide this specific model with those
- 4:41:30records so that this model will be able
- 4:41:33to train on this and probably be able to
- 4:41:35get the output now let's create buckets
- 4:41:38now based on buckets how the buckets
- 4:41:39will be created over here I will take 07
- 4:41:43until
- 4:41:45sorry whatever is the value over here
- 4:41:48normal we value okay so I will start
- 4:41:50creating my buckets buckets basically
- 4:41:52from 0 to
- 4:41:5307 what did I say now for this decision
- 4:41:57tree or stump I need to provide some
- 4:42:00records so the maximum number of record
- 4:42:02that should be going should be the wrong
- 4:42:05records that should go over here now how
- 4:42:07do we decide that okay there should be a
- 4:42:09way that we should be able to say that
- 4:42:11that specific wrong number of Records
- 4:42:13should go to that decision tree so for
- 4:42:16that purpose what we do is that this
- 4:42:18decision tree will randomly create some
- 4:42:20numbers between 0 to 1 randomly create
- 4:42:25those numbers between 0 to 1 and
- 4:42:27whichever bucket it will come in like 07
- 4:42:30to 014 014 to 07 basically means 0 2 1
- 4:42:37then 0 2 1 2 see how the bucket is
- 4:42:40getting cre this value is getting added
- 4:42:42to this so that becomes this bucket 021
- 4:42:45+3 537 how much it is it is nothing but
- 4:42:50470 747 then 747
- 4:42:55to
- 4:42:57751 like this you create all the buckets
- 4:43:00okay you can create all the buckets now
- 4:43:02tell me which record is basically having
- 4:43:04the biggest bucket size obviously this
- 4:43:07record so if I randomly create a number
- 4:43:10between 0 to one what is the highest
- 4:43:13probability that the values will be
- 4:43:15going in so in this particular case most
- 4:43:17of the wrong records will be passed
- 4:43:18along with the other records obviously
- 4:43:20other records there are chances that
- 4:43:22other records will go to the next
- 4:43:24decision tree but understand maximum
- 4:43:26number will go with the wrong records
- 4:43:28because the bucket is high over here so
- 4:43:31the bucket is high over here so most of
- 4:43:32the time this specific record will get
- 4:43:35create selected and then it will be gone
- 4:43:37to the second tree now suppose I have
- 4:43:40this all records
- 4:43:41so this is my first stump this is my
- 4:43:44second stump this is my third stump
- 4:43:47similarly the third stump from the
- 4:43:48second stump whichever wrong records
- 4:43:50will be going maximum number of Records
- 4:43:52will go over here then again it will be
- 4:43:54trained like this we'll be having lot of
- 4:43:56stumps minimum 100 decision trees can be
- 4:43:59added you know that every decision tree
- 4:44:01will give one output for a new test data
- 4:44:03new test data this week learner will
- 4:44:05give one output this week learner will
- 4:44:07give one output this week learner and
- 4:44:09this will week learner will be giving
- 4:44:10one output obviously the time complexity
- 4:44:12will be more now from this particular
- 4:44:14output suppose it is a binary
- 4:44:16classification I will be getting 0 1 1 1
- 4:44:19so again over here majority voting will
- 4:44:21happen and the output will be one in
- 4:44:24case of regression problem I will be
- 4:44:25having a continuous value over here and
- 4:44:28for this the average average will be
- 4:44:31computed and that will give me an output
- 4:44:33over here so for regression the average
- 4:44:36will be done for classification what
- 4:44:39will happen majority will be be
- 4:44:41happening so everywhere that same part
- 4:44:43will be going on buckets is very much
- 4:44:45simple guys buckets basically means
- 4:44:47based on this weights normalized weight
- 4:44:49we are going to create bucket so that
- 4:44:51whichever records has the highest bucket
- 4:44:53based on this randomly creating code you
- 4:44:55know it will select those specific
- 4:44:57buckets and put it into random Forest
- 4:44:59understand why this bucket size is Big
- 4:45:02the other wrong records which are
- 4:45:03present right suppose they are have more
- 4:45:05than four to five wrong records their
- 4:45:06bucket size will also be bigger and
- 4:45:08because based on this randomly creating
- 4:45:10num between 0 to 1 most of the wrong
- 4:45:12records will be selected and given to
- 4:45:14the second stum similarly this
- 4:45:16particular decision tree will be doing
- 4:45:17some mistakes then that wrong records
- 4:45:19will get updated all the weights will
- 4:45:20get updated and it will be passed to the
- 4:45:22next decision tree guys when I say wrong
- 4:45:24record the output will be same only no
- 4:45:26zero and one so interesting everyone I
- 4:45:29hope you understood so much of maths in
- 4:45:31adab boost and how adab boost actually
- 4:45:33work three main things one is total
- 4:45:35error one is performance of stump and
- 4:45:37one is the new sample weight these
- 4:45:39things are getting calculated extensive
- 4:45:41max normalized weight was basically used
- 4:45:43because the sum of all these weights are
- 4:45:45approximately equal to one when boosting
- 4:45:48why not take the last output no no no we
- 4:45:50have to give the importance of every
- 4:45:52decision tree output every decision tree
- 4:45:55output are important okay let me talk
- 4:45:57about one model which is called as
- 4:45:59blackbox model versus white box what is
- 4:46:03the difference between blackbox model
- 4:46:04and white box if I take an example of
- 4:46:07linear regression tell me what kind of
- 4:46:09model it is is is it a white box model
- 4:46:12or black box if I take an example of
- 4:46:14random
- 4:46:15Forest is this a white box or black box
- 4:46:18if I take an example of decision tree it
- 4:46:21is a white box of blackbox model if I
- 4:46:23take an example of a Ann is it a white
- 4:46:26box of blackbox model linear regression
- 4:46:28is basically called as an wide Box model
- 4:46:30because here you can basically visualize
- 4:46:33how the Theta value is basically
- 4:46:35changing and how it is coming to a
- 4:46:36global Minima and all those things in
- 4:46:38random Forest I will say this as
- 4:46:40blackbox model because it is impossible
- 4:46:42to see all the decision tree how it is
- 4:46:44working so that is the reason the maths
- 4:46:46is so complex inside this if I talk
- 4:46:49about decision tree this is basically a
- 4:46:50white box model because in decision tree
- 4:46:52we know how the split are basically
- 4:46:54happening with the help of paper and pen
- 4:46:55you'll be able to do it in the case of
- 4:46:58an Ann this is a blackbox model because
- 4:47:00here you don't know like how many
- 4:47:02neurons are there how they are
- 4:47:03performing and how the weights are
- 4:47:05getting updated so this is the basic
- 4:47:07difference between the blackbox and uh
- 4:47:10uh white box model this entire thing is
- 4:47:13the agenda of today's session so let's
- 4:47:15start uh the first algorithm that we are
- 4:47:17probably going to discuss today is
- 4:47:19something called as K
- 4:47:21means
- 4:47:23clustering K means clustering and this
- 4:47:26is a kind of unsupervised machine
- 4:47:28learning now always remember
- 4:47:31unsupervised machine learning basically
- 4:47:33means that uh the one and the most
- 4:47:35important thing is that in unsupervised
- 4:47:38machine learning
- 4:47:41in unsupervised ml you don't have any
- 4:47:44specific output so you don't have any
- 4:47:46specific output so suppose you have
- 4:47:48feature one and feature two and suppose
- 4:47:50you have datas different different data
- 4:47:53you know and based on this data what we
- 4:47:55do we basically try to create clusters
- 4:47:58this clusters basically says what are
- 4:48:00the similar kind of data so this is what
- 4:48:03we basically do from uh clustering and
- 4:48:06there are various techniques like K
- 4:48:08Mains uh it is hierle clustering and all
- 4:48:10so first of all we'll try to understand
- 4:48:12about K means and how does it
- 4:48:14specifically work it's simple uh suppose
- 4:48:17you have a data points like this okay
- 4:48:20let's say that this is your F1 feature
- 4:48:21F2 feature and based on this in two
- 4:48:23dimensional probably I will be plotting
- 4:48:26this points and suppose this is my
- 4:48:28another points so our main purpose is
- 4:48:31basically to Cluster together in
- 4:48:34different different groups okay so this
- 4:48:36will be my one group and probably the
- 4:48:38other group will be this group right so
- 4:48:40two groups because obviously you can see
- 4:48:42from this clusters here you have two
- 4:48:44similar kind of data which is basically
- 4:48:47grouped together right this is my
- 4:48:49cluster one and this is my cluster 2 let
- 4:48:51me talk about this and why specifically
- 4:48:54it'll be very much useful then we'll try
- 4:48:56to understand about math intuition also
- 4:48:58now always understand guys uh where does
- 4:49:00clustering gets used okay in most of the
- 4:49:03Ensemble techniques I told you about
- 4:49:05custom emble technique right so custom
- 4:49:08emble techniques in custom assemble
- 4:49:11techniques you know whenever we are
- 4:49:13probably creating a model first of all
- 4:49:15on our data set what we do is that we
- 4:49:18create clusters so suppose this is my
- 4:49:20data set during my model creation the
- 4:49:22first algorithm we will probably apply
- 4:49:24will be clustering algorithm and after
- 4:49:26that it is obviously good that we can
- 4:49:28apply regression or classification
- 4:49:30problem suppose in this clustering I
- 4:49:32have two or three groups let's say that
- 4:49:34I have two or three groups over here for
- 4:49:36each group we can apply a separate
- 4:49:40supervis machine learning algorithm if
- 4:49:42we know the specific output that we
- 4:49:44really want to take ahead I'll talk
- 4:49:46about this and uh give you some of the
- 4:49:48examples as I go ahead now let's go on
- 4:49:51go ahead and focus more on understanding
- 4:49:53how does kin's clustering algorithm work
- 4:49:56so let's go over here the word K means
- 4:49:59has this K value this K are nothing but
- 4:50:02this K basically means centroids K
- 4:50:05basically means centroids so suppose if
- 4:50:08I have a data set which looks like this
- 4:50:10let's say that this is my data set now
- 4:50:12over here just by seeing the data set
- 4:50:14what are the possible groups you think
- 4:50:16definitely you'll be saying K is equal
- 4:50:18to 2 So when you say k is equal to two
- 4:50:20that basically means you will be able to
- 4:50:22get two groups like this and each and
- 4:50:24every group will be having a centroid a
- 4:50:28centroid Point here also there will be a
- 4:50:30centroid point so this centroid will
- 4:50:32determine basically this is a separate
- 4:50:34group over here this is a separate group
- 4:50:36over here so over here here you can
- 4:50:38definitely say that fine this is two
- 4:50:40groups but but how do we come to a
- 4:50:41conclusion that there is only two groups
- 4:50:44okay we cannot just directly say that
- 4:50:46okay we'll try to just by seeing the
- 4:50:48data because your data will be having a
- 4:50:50high dimension data right right now I'm
- 4:50:52just showing your two Dimension data but
- 4:50:55for a high dimension data definitely
- 4:50:56you'll not be able to see the data
- 4:50:58points how it is plotted so how do you
- 4:51:00come to a conclusion that only two
- 4:51:02groups are there so for this there is
- 4:51:03some steps that we basically perform in
- 4:51:05K mins the first step is that we try
- 4:51:08with different K values we try with
- 4:51:11different K values and which is the
- 4:51:13suitable K value K is nothing but
- 4:51:15centroids okay it is nothing but
- 4:51:18centroids we try with different
- 4:51:20different centroids in this particular
- 4:51:22case let's say that I have this
- 4:51:24particular data point and I actually
- 4:51:27start with k is equal 1 or 2 or 3 any
- 4:51:29one you want let's say that I'm going to
- 4:51:31start with k is equal 2 how to come up
- 4:51:34with this K is equal to 2 as a perfect
- 4:51:37value that I'll talk about it we need to
- 4:51:39know there is a concept which is called
- 4:51:41as within cluster sum of square so when
- 4:51:43we try different K values let's say that
- 4:51:45for K is equal to 2 what will happen the
- 4:51:47first step we select a we try K values
- 4:51:50so let's say that we are considering K
- 4:51:52is equal to 2 the second step is that we
- 4:51:54initialize K number of centroids now in
- 4:51:57this particular case I know my K value
- 4:51:59is 2 so we will be initializing randomly
- 4:52:02let's say that K is equal to 2 so what
- 4:52:05we can actually do let's say that this
- 4:52:07is this is my one centroid I will I'll
- 4:52:09put it in another color so this will be
- 4:52:11my one centroid and let's say that this
- 4:52:13is my another centroid so I have
- 4:52:15initialized two centroids randomly in
- 4:52:17this space now after this particular
- 4:52:19centroid what we have to do is that
- 4:52:21after initializing this centroid what we
- 4:52:23have to do is that we have to basically
- 4:52:26find out which points are near to the
- 4:52:29centroid and which points are near to
- 4:52:31this centroid now in order to find out
- 4:52:33it is a very easy step we can basically
- 4:52:35use ukan distance to find out the
- 4:52:38distance between the points in an easy
- 4:52:40way if I really want to show you that
- 4:52:44you know like how many points I want to
- 4:52:46in an easy way what I can do I can
- 4:52:48basically draw a straight line over here
- 4:52:50let's say that I'm drawing a straight
- 4:52:51line over here in another color I can
- 4:52:54draw a straight line and I can also draw
- 4:52:56one parallel line like this so This
- 4:52:58basically indicates that whichever
- 4:53:01points you see over here suppose if I
- 4:53:03draw a straight line in between all
- 4:53:05these points you will be able to see
- 4:53:07that let's say that I'm drawing one more
- 4:53:09parallel line
- 4:53:11which is intersecting together so from
- 4:53:14this you can definitely find out let's
- 4:53:16say that these are all my points that
- 4:53:17are nearer to this green line Green
- 4:53:20Point so what I'm actually going to do
- 4:53:21in this particular case all these points
- 4:53:24that you are seeing near the green it
- 4:53:26will become green color so that
- 4:53:28basically means this is basically nearer
- 4:53:30to this centroid and whichever points
- 4:53:33are nearer to this particular point that
- 4:53:35will become red point so that basically
- 4:53:38means this belongs to this group okay
- 4:53:40this belongs to this group so I hope
- 4:53:42everybody's clear till here then what
- 4:53:44will happen is that this summation of
- 4:53:48all the values then we initialize the K
- 4:53:51number of centroids that is done then we
- 4:53:53try to calculate the distance we try to
- 4:53:55find out which all points is nearer to
- 4:53:57the centroid let's say that this is my
- 4:53:58one centroid this is my another centroid
- 4:54:01and we have seen that okay these all
- 4:54:02points belong to this centroid it near
- 4:54:05to this particular centroid so this is
- 4:54:07becoming red so that is based on the
- 4:54:09shortage distance and here it is
- 4:54:11becoming green now the next step let's
- 4:54:13see what is the next step after this so
- 4:54:15I am going to remove this thing now the
- 4:54:17next step will be that the entire points
- 4:54:20that is in red color all the average
- 4:54:22will be taken so here again the average
- 4:54:25will be taken now third step here I'm
- 4:54:28going to write here we are going to
- 4:54:30compute the average the reason we
- 4:54:32compute the average is that because we
- 4:54:34need to update the centroid so compute
- 4:54:37the average to update centroid to update
- 4:54:40centroids so here you'll be able to see
- 4:54:42that what I'm actually doing as soon as
- 4:54:45we compute the average this centroid is
- 4:54:47going to move to some other location so
- 4:54:50what location it will move it will
- 4:54:51obviously become somewhere in Center so
- 4:54:53here now I'm going to rub this and now
- 4:54:56my new centroid will be this point where
- 4:54:58I am actually going to draw like this
- 4:55:00let's say this is my new centroid now
- 4:55:02similarly this thing will happen with
- 4:55:04respect to the green color so with
- 4:55:06respect to the green color also it will
- 4:55:08happen and this green will also Al get
- 4:55:10updated so I'm going to rub this and
- 4:55:12this will be my new Green Point which
- 4:55:14will get updated over here then again
- 4:55:16what will happen again the distance will
- 4:55:18be calculated and again a perpendicular
- 4:55:20line will be calculated here you can see
- 4:55:22that now all the points are towards
- 4:55:25there okay again the centroid based on
- 4:55:27this particular distance again it will
- 4:55:29be calculated and here you can see that
- 4:55:31all the points are in its own location
- 4:55:33so here now no update will actually
- 4:55:36happen let's say that there was one
- 4:55:38point which was red color over here
- 4:55:41then this would have become green color
- 4:55:42but since the updation has happened
- 4:55:44perfectly we are not going to update it
- 4:55:46and we are not going to update the
- 4:55:48centroid right so now you can understand
- 4:55:51that yes now we have actually got the
- 4:55:53perfect centroid and now this will be
- 4:55:56considered as one group and this will be
- 4:55:58basically considered as the another
- 4:56:00group it will not intersect but right by
- 4:56:02default here intersection is happening
- 4:56:05so I hope everybody's understood the
- 4:56:07steps that you have actually followed in
- 4:56:09initializing the centroids in updating
- 4:56:12the centroids and in updating the points
- 4:56:14is it clear everybody with respect to K
- 4:56:17means now let's discuss about one
- 4:56:20point how do we decide this K value okay
- 4:56:24how do we decide this K value so for
- 4:56:26deciding the K value there is a concept
- 4:56:27which is called as elbow method so here
- 4:56:31I'm going to basically Define my elbow
- 4:56:32method now elbow method says something
- 4:56:35very much important because this will
- 4:56:37actually help us to find out what is the
- 4:56:40optimized K value whether the K value
- 4:56:42should be two whether uh the K value is
- 4:56:45going to be three whether the K value is
- 4:56:47going to become four and always
- 4:56:49understand suppose this is my data set
- 4:56:51suppose this is my data set initially
- 4:56:53let's say that I have my data points
- 4:56:54like this we cannot go ahead and
- 4:56:57directly say say that okay K is equal to
- 4:56:592 is going to work so obviously we are
- 4:57:01going to go with iteration for I is
- 4:57:04equal to probably 1 to 10 I'm going to
- 4:57:06move towards iteration from 1 to 10
- 4:57:09let's say so for every iteration we will
- 4:57:11construct a graph with respect to K
- 4:57:14value and with respect to something
- 4:57:16called as W CSS now what is this W CSS W
- 4:57:20CSS basically means within cluster sum
- 4:57:23of
- 4:57:24square okay this is the meaning of wcss
- 4:57:27within cluster sum of square now let's
- 4:57:30say that initially we start with one
- 4:57:33centroid so one centroid let's say it is
- 4:57:35initialized here one centroid is
- 4:57:37basically initialized here if we go and
- 4:57:39compute the distance
- 4:57:40between each and every points to the
- 4:57:43centroid and if we try to find out the
- 4:57:45distance will the distance value be
- 4:57:47greater or it will be smaller will it be
- 4:57:50smaller or greater tell me if you try to
- 4:57:53calculate this distance from this
- 4:57:55centroid to every point this is what is
- 4:57:57within cluster sum of square it will
- 4:58:00always be very very much greater so
- 4:58:02let's say that my first point has come
- 4:58:04somewhere here it is going to be
- 4:58:06obviously greater let's say that my
- 4:58:07first point is coming over here find
- 4:58:10So within K is equal to 1 initially we
- 4:58:12took and we found out the distance of w
- 4:58:14CSS and it is a very huge value okay
- 4:58:17because we're going to compute the
- 4:58:18distance between each and every point to
- 4:58:20the centroid now the next thing that I'm
- 4:58:23actually going to do is that now we'll
- 4:58:26go with next value that is K is equal to
- 4:58:282 now in K is equal to 2 I will
- 4:58:31initialize two points okay I will
- 4:58:34initialize two points and then probably
- 4:58:36I will do the entire process which I
- 4:58:38have written on the top now tell me me
- 4:58:40whichever points is nearer to this green
- 4:58:42point if we compute the distance and
- 4:58:46whichever points is nearer to the red
- 4:58:48point if you compute the distance like
- 4:58:52this now this summation of the distance
- 4:58:55will be lesser than the previous W CSS
- 4:58:57or not obviously it is going to be
- 4:59:00lesser than the previous W CSS so what
- 4:59:02I'm actually going to do probably with K
- 4:59:04is equal to 2 your value may come
- 4:59:06somewhere here then with K is equal to 3
- 4:59:09your value May come somewhere here then
- 4:59:10K is equal to 4 will come here to 5 6
- 4:59:13like this it will go so here if I
- 4:59:15probably join this line you'll be able
- 4:59:17to see that there will be an Abrupt
- 4:59:19changes in the W CSS value in the wcss
- 4:59:23value there will be an Abrupt changes
- 4:59:25and this this is basically called as
- 4:59:27elbow curve now why we say it as elbow
- 4:59:30curve because it is in the shape of
- 4:59:32elbow and here at one specific point
- 4:59:34there will be an Abrupt change and then
- 4:59:36it will be straight so that is the
- 4:59:38reason why we basically say this as
- 4:59:41elbow okay so this is a very important
- 4:59:43thing see in finding the K value we use
- 4:59:46elbow method but for validating purpose
- 4:59:49how do we validate that this model is
- 4:59:52performing well we use silard score that
- 4:59:54I'll show you just in some time but
- 4:59:57understand that in K means clustering we
- 5:00:00need to update the centroids and based
- 5:00:02on that we calculate the distance and as
- 5:00:05the K value keep on increasing you'll be
- 5:00:07able to see that the distance will
- 5:00:09become normal or the wcss value will
- 5:00:12become normal and then we really need to
- 5:00:14find out which is the phys K value where
- 5:00:17the abrupt change see over here suppose
- 5:00:20abrupt change is there and then it is
- 5:00:21normal then I will probably take this as
- 5:00:24my K value so obviously the model
- 5:00:26complexity will be high because we are
- 5:00:28going to check with respect to different
- 5:00:30different K values and wcss values and
- 5:00:33this basically means that the value that
- 5:00:36we'll probably get first of all we need
- 5:00:38to construct this elbow curve then see
- 5:00:40the changes where it is basically
- 5:00:42happening we'll need to find out the
- 5:00:43abrupt change and once we get the abrupt
- 5:00:46change we basically say that this may be
- 5:00:49the K value so K is equal to 4 as an
- 5:00:52example I'm telling you so unless and
- 5:00:54until if you really want to find the
- 5:00:56cluster it is very much simple we take a
- 5:00:59k value we initialize K number of
- 5:01:01centroids we compute the average to
- 5:01:03update the centroids then again we try
- 5:01:05to find out the distance try to see that
- 5:01:07whether any points has changed and
- 5:01:08continue that process unless and until
- 5:01:10we get separate groups okay so this is
- 5:01:14the entire funa of claim in clustering
- 5:01:16so finally you'll be able to see that
- 5:01:18with respect to the K value we will be
- 5:01:20able to get that many number of groups
- 5:01:22if my K value is four that basically
- 5:01:24means I will be probably getting four
- 5:01:26different groups like this 1 two right
- 5:01:30three like this and four I will be
- 5:01:32getting four groups like this with K is
- 5:01:34equal to 4 that basically means K is
- 5:01:35equal to four clusters and every group
- 5:01:38will be having its own centroids okay
- 5:01:41every group will be having okay
- 5:01:42centroids are very much important yes
- 5:01:45I'll try to show you in the coding also
- 5:01:47guys let's go towards the second
- 5:01:48algorithm the second algorithm that we
- 5:01:51will be probably discussing is called as
- 5:01:54hierarchical clustering now hierarchal
- 5:01:56clustering is very much simple guys all
- 5:01:58you have to do is that let's say this is
- 5:02:00your data points this is your data
- 5:02:01points and this is my P1 let's say P2
- 5:02:04now hierle clustering says that we will
- 5:02:07go step by step the first thing is that
- 5:02:10we will try to find out the most nearest
- 5:02:12Value let's say this is my X and Y let's
- 5:02:15say these are my points like this is my
- 5:02:18P1 point this is my P2 point this is my
- 5:02:21P3 point this is my P4 Point P5 Point P6
- 5:02:25point p7 point okay so these are my
- 5:02:28points that I have actually named over
- 5:02:29here let's say that this may be the
- 5:02:31nearest point to each other so what it
- 5:02:32will do it will combine this together
- 5:02:34into one cluster this we have computed
- 5:02:37the distance so it will C create one
- 5:02:39cluster now what will happen on the
- 5:02:41right hand side there will be another
- 5:02:42notation which you may be using in
- 5:02:45connecting all the points one so suppose
- 5:02:46this is my P1 this is my P2 this is my
- 5:02:50P3 P4 let's say that I have this many
- 5:02:53points and probably I will also try to
- 5:02:56make
- 5:02:57p7 so these are my points p7 now you
- 5:03:00know that the nearest point that we are
- 5:03:02having okay this will probably be
- 5:03:04distance 1 2 3 this is distance okay 4 5
- 5:03:096 like this we have lot of distance so
- 5:03:12hierle clustering will first of all find
- 5:03:14out the nearest point and try to compute
- 5:03:17the distance between them and just try
- 5:03:18to combine them together into one what
- 5:03:21do we do we basically combine them into
- 5:03:23one group okay so P1 and P2 has been
- 5:03:26combined let's say then it'll go and
- 5:03:29find out the other nearest point so
- 5:03:31let's say P6 and p7 are near so they are
- 5:03:33also going to combine into one group so
- 5:03:35once they combine into one group then we
- 5:03:37have P6 and p7 which will be obviously L
- 5:03:40greater than the previous distance and
- 5:03:42we may get this kind of computation and
- 5:03:44another combination or cluster will form
- 5:03:47get formed over here then you have seen
- 5:03:49that okay P3 and P5 are nearer to each
- 5:03:52other so we are going to combine this so
- 5:03:54I'm going to basically combine P3 and
- 5:03:57P5 okay and let's say that this distance
- 5:03:59is greater than the previous one because
- 5:04:02we are basically going to sh start with
- 5:04:03the shortest distance and then we are
- 5:04:05going to capture the longest distance
- 5:04:07now this is done now you can see that
- 5:04:08the next point that is near right to
- 5:04:11this particular group is P4 so we are
- 5:04:13going to combine this together into one
- 5:04:15group so once we combine this into one
- 5:04:17group this P4 will get connected like
- 5:04:20this let's say it is getting connected
- 5:04:23like this P4 has got connected then what
- 5:04:25is the nearest Point whether it is P6 p7
- 5:04:28group or P1 P2 obviously here you can
- 5:04:30see that P1 P2 is there so I am probably
- 5:04:32going to combine this group together
- 5:04:34that basically means P1 P2 let's say I'm
- 5:04:38just going to combine this group group
- 5:04:40together again circle is coming so I
- 5:04:42will make a dot let's say I'm going to
- 5:04:43combine this group together because
- 5:04:45these are my nearest groups so what will
- 5:04:47happen P1 and P2 will get combined to P5
- 5:04:50sorry P4 P5 this one so I will be
- 5:04:53getting another line like this and then
- 5:04:55finally you'll be seeing that P6 p7 is
- 5:04:57the nearest group to this so this will
- 5:05:00totally get combined and it may look
- 5:05:02something like this so this will become
- 5:05:05a total group like
- 5:05:07this so all the groups are combined so
- 5:05:10finally you'll be able to see that there
- 5:05:11will be one more line which will get
- 5:05:13combined like
- 5:05:14this this is basically called as
- 5:05:17dendogram dendogram okay which is like
- 5:05:21bottom root to top now the question
- 5:05:24arises is that how do you find that how
- 5:05:25many groups should be here how do you
- 5:05:27find out that how many groups should be
- 5:05:29here the funa is very much Clear guys in
- 5:05:32this is that you need
- 5:05:34to find the longest
- 5:05:41vertical line you need to find out the
- 5:05:43longest vertical line that has no
- 5:05:46horizontal line pass through it no
- 5:05:49horizontal
- 5:05:51line passed through it this is very much
- 5:05:54important that has no horizontal line
- 5:05:56pass through it now what this is
- 5:05:58basically meaning is that I will try to
- 5:06:00find out the longest line longest
- 5:06:03vertical line in such a way that none of
- 5:06:06the horizontal line passes through it
- 5:06:07what is horizontal line suppose if I
- 5:06:09consider this vertical line This
- 5:06:11vertical line over here if you see that
- 5:06:13if I extend this green line it is
- 5:06:15passing through this if I extend this
- 5:06:17line it is passing through this right if
- 5:06:20I'm extending this line it is passing
- 5:06:21through this right so out of this the
- 5:06:25longest line that may be passing in such
- 5:06:27a way that no horizontal line probably
- 5:06:29is this line that I can actually see so
- 5:06:31what you do over here is that you
- 5:06:33basically just create a straight line
- 5:06:35over this and then you try to find out
- 5:06:37that how many clusters it will be there
- 5:06:39by understanding that how many lines it
- 5:06:41is passing through if it is passing
- 5:06:42through this one line two line three
- 5:06:44line four line that basically means your
- 5:06:47clusters will be four
- 5:06:49clusters this is how we basically do the
- 5:06:52calculation in heral clustering again
- 5:06:56here it may not be the perfect line I've
- 5:06:58just drawn with some assumptions but if
- 5:07:00you are trying to do this probably you
- 5:07:02have to do in this specific way okay
- 5:07:04I've already uploaded a lot of practical
- 5:07:06videos with respect to highill
- 5:07:08clustering and all now now tell me
- 5:07:11maximum effort or maximum time is taken
- 5:07:15by is taken
- 5:07:18by K
- 5:07:20means or hierle clustering this is a
- 5:07:25question for you yes guys number of
- 5:07:26clusters may be three but here I'm just
- 5:07:29showing you that how many lines it may
- 5:07:31be passed by how do you basically
- 5:07:34determine whether maximum time will be
- 5:07:36taken by kin or Hier clustering this is
- 5:07:38an interview question the maximum time
- 5:07:40that will be taken is by hierarchical
- 5:07:45clustering why because let's say that I
- 5:07:48have many many many data points at that
- 5:07:51point of time hierle clustering will
- 5:07:53keep on constructing this kind of
- 5:07:55dendograms and it will be taking many
- 5:07:58many many time lot time right so hierle
- 5:08:02clustering will take more time maximum
- 5:08:05time that it is going to basically take
- 5:08:07so it is very much important that that
- 5:08:09you understand which is making basically
- 5:08:12taking more time so if your data set is
- 5:08:15small you may go ahead with hierle
- 5:08:18clustering if your data set is large go
- 5:08:21with K means clustering go with K means
- 5:08:23clustering in short both will take more
- 5:08:25time but K Min will perform better than
- 5:08:28hle clustering see guys you will be
- 5:08:30forming this kind of dendograms right
- 5:08:33and just imagine if you have 10 features
- 5:08:34and many data points how you're going to
- 5:08:37do it it will be a cubers some process
- 5:08:40you'll not be even able to see this
- 5:08:42dendogram properly and manually
- 5:08:44obviously you cannot do it so this was
- 5:08:46with respect to K means clust swing and
- 5:08:49H mean clust swing I hope everybody's
- 5:08:51understood now the next topic that we'll
- 5:08:53focus on is that how do we
- 5:08:56validate see how do we validate a
- 5:08:59classification problem we use
- 5:09:00performance metric like confusion Matrix
- 5:09:03accuracy um different different true
- 5:09:05positive rate Precision recall but how
- 5:09:07do we validate clustering model Model S
- 5:09:10we are going to use something called as
- 5:09:12so we are going to basically use
- 5:09:14something called as
- 5:09:15Sil score I'll show you what Sid score
- 5:09:19is I'm going to just open the Wikipedia
- 5:09:21so this is how a CID score looks like a
- 5:09:25very very amazing topic okay how do we
- 5:09:28validate whether my model basically has
- 5:09:32perfect three or four model perfect
- 5:09:35three suppose if I find out my K value
- 5:09:37is three how do we find out now see one
- 5:09:40more one more issue with K means one
- 5:09:42issue with K means which I forgot to
- 5:09:44tell you let's say that I have a data
- 5:09:46point which looks like this and suppose
- 5:09:49I have some data points like this I have
- 5:09:51some data points which looks like this
- 5:09:55let's say I have like this now in this
- 5:09:58one issue will be that suppose I try to
- 5:10:01make a cluster over here obviously
- 5:10:03you'll be saying my K value will be two
- 5:10:05okay in this particular case suppose
- 5:10:07this is one cluster this is my another
- 5:10:08cluster
- 5:10:10right because of my wrong initialization
- 5:10:13of the points okay understand because
- 5:10:16suppose if I initialize just randomly
- 5:10:18some centroids like this then what may
- 5:10:20happen is that there is a possibility
- 5:10:22that we may also have three clusters
- 5:10:24like like like this kind of clusters one
- 5:10:27cluster will be here one cluster will be
- 5:10:29here one cluster will be here so this
- 5:10:32initialization of the centroids one
- 5:10:35condition is that it should be very very
- 5:10:37far if we initialize our centroids very
- 5:10:41very far at that point of time we will
- 5:10:43be able to find the centroid exactly in
- 5:10:46the center because it will keep on
- 5:10:47updating it'll keep on going ahead right
- 5:10:50but if we don't initialize that very far
- 5:10:53then there will be a situation that
- 5:10:55probably if I wanted to get only the
- 5:10:57real thing was to get only two centroids
- 5:10:59I was probably getting three centroids
- 5:11:01right so this is a problem so for this
- 5:11:04there is an algorithm which is called as
- 5:11:06K means Plus+ and what this K means
- 5:11:08Plus+ will do which I will probably show
- 5:11:10you in Practical this will make sure
- 5:11:12that all the centroids that are
- 5:11:14initialized it is very very
- 5:11:16far okay all the in centroids that is
- 5:11:19basically there it is initialized very
- 5:11:21very far we'll see that in practical
- 5:11:23application where specifically those
- 5:11:26centroids are basically used now let me
- 5:11:28go ahead and let me show
- 5:11:30you with respect to Sid clust string now
- 5:11:34what is the solo color string I'm going
- 5:11:36to explain you in an amazing way this is
- 5:11:38important
- 5:11:39if someone says you how do we validate
- 5:11:43how do we validate cluster
- 5:11:46model then at that point of time we
- 5:11:48basically use this site it will be used
- 5:11:51in it will be used with respect
- 5:11:55to it will be used with respect to K
- 5:11:58means it can be used in hierle mean
- 5:12:00right if you want to validate how do we
- 5:12:03validate okay that is what we are
- 5:12:04basically going to see over here now in
- 5:12:08C's clustering
- 5:12:09what are the most important things the
- 5:12:12first and the most important thing is
- 5:12:13that we will try to find out we will try
- 5:12:16to find out a ofi we will try to find
- 5:12:19out a of I now what is this a ofi see
- 5:12:22this a ofi that you basically see a ofi
- 5:12:25is nothing but see three major steps
- 5:12:28happens in order to validate cluster
- 5:12:30model with the help of solo first thing
- 5:12:33is that I will probably take one cluster
- 5:12:36okay there will be one point
- 5:12:39which will be my centroid let's say and
- 5:12:42then what I'm going to do I'm just going
- 5:12:44to whatever points are there inside this
- 5:12:46cluster I'm going to compute the
- 5:12:49distance between them so I'm going to do
- 5:12:52the summation and I'm also going to do
- 5:12:54the average of all this distance so here
- 5:12:57you can see that when I said distance of
- 5:12:59I comma J I basically means this point J
- 5:13:03basically means all these points I is
- 5:13:06nothing but it is the centroid so here
- 5:13:08is nothing but this this is the centroid
- 5:13:09let's say that I'm having the centroid
- 5:13:11so I'm going to compute all the distance
- 5:13:13over here which is mentioned by this and
- 5:13:15this value that you see that I'm
- 5:13:17actually dividing by C of I minus one in
- 5:13:20Short I am actually trying to calculate
- 5:13:22the average
- 5:13:24distance so this is the first point
- 5:13:26where I'm actually Computing the a ofi
- 5:13:28now similarly what I will do is
- 5:13:31that what I will do is that the next
- 5:13:34point will be that suppose I have
- 5:13:36computed a ofi the next the next that we
- 5:13:39need to compute is B ofi now what is b
- 5:13:41ofi b ofi is nothing but there will be
- 5:13:44multiple clusters in a k means problem
- 5:13:47statement we will try to find out the
- 5:13:50nearest cluster okay suppose let's say
- 5:13:52that this is the nearest cluster and in
- 5:13:54this I have all the variety of points
- 5:13:58then B ofi basically says that I will
- 5:14:00try to compute the distance between each
- 5:14:03point and the other point in this
- 5:14:06centroid sorry in this cluster so this
- 5:14:08is my cluster one this is my cluster two
- 5:14:12so what I'm actually going to do is that
- 5:14:14here I'm going to compute the distance
- 5:14:16between this point to this point then
- 5:14:17this point to this point then this point
- 5:14:20to this point this point to this point
- 5:14:22this point to this point this point to
- 5:14:24this point every point I'm actually
- 5:14:26going to compute the distance once this
- 5:14:28point is done we will go ahead with the
- 5:14:30next point and we'll try to compute the
- 5:14:31distance and once we get all this
- 5:14:34particular distance what we are going to
- 5:14:35do we are going to do the average of
- 5:14:37them average
- 5:14:39now tell me if I try to find out the
- 5:14:42relationship between a of I and B of I
- 5:14:45if my cluster model is good will a of
- 5:14:50I will be greater than b of I or
- 5:14:54will B of I will be greater than a ofi
- 5:14:58if I have a good clustering model if I
- 5:15:01have a good clustering model will a of I
- 5:15:05is greater than b of I will be greater
- 5:15:08than b of I or whether B of I will be
- 5:15:10greater than a of I out of this if we
- 5:15:13have a really good model obviously the
- 5:15:16distance between B of I will be greater
- 5:15:19than a of I in a good model that
- 5:15:22basically means if I talk about sloid
- 5:15:24clustering the values will be between -1
- 5:15:27to +1 the more the value is towards +1
- 5:15:32that basically means the good the model
- 5:15:34is the good the clustering model is the
- 5:15:37more the values towards negative one
- 5:15:39that basically means this condition is
- 5:15:40getting applied now what does this
- 5:15:42condition basically say that basically
- 5:15:43means that this distance is far than the
- 5:15:46cluster distance this is what this
- 5:15:48information is getting portrayed and
- 5:15:51this is the importance of CID
- 5:15:53clustering finally when we apply the
- 5:15:55formula of CID clustering you'll be able
- 5:15:57to see that sloid clustering is nothing
- 5:16:00but let me rub this everything guys for
- 5:16:03you let me just show you what is CID
- 5:16:05clustering CID clustering formula will
- 5:16:08be something like this this B of I so
- 5:16:11here you have solid clustering this is
- 5:16:13the formula B of I minus a of I Max of a
- 5:16:18of I comma B of I if C of I is greater
- 5:16:21than one right so by this you will be
- 5:16:24getting the value between -1 to + 1 and
- 5:16:28more the value is towards + one the more
- 5:16:31good your model is more the values
- 5:16:33towards minus1 more bad your model is
- 5:16:36because if it is towards minus1 that
- 5:16:38basically means your a of I is obviously
- 5:16:41greater than b of I so this is the
- 5:16:43outcome with respect to cot crust string
- 5:16:46if s is equal to zero that basically
- 5:16:47means still your model needs to be uh
- 5:16:50per basically the clustering needs to be
- 5:16:52improved what is I over here I is
- 5:16:54nothing but one data point you you can
- 5:16:56just read this guys data point in I in
- 5:16:59the cluster C of I so I hope everybody's
- 5:17:01understood this now let's go ahead and
- 5:17:03let's discuss about the next topic we
- 5:17:05have obviously finished up solart
- 5:17:07clustering over here let's discuss about
- 5:17:09something called as DB
- 5:17:11scan so for DB scan clustering this is
- 5:17:14an amazing clustering algorithm we'll
- 5:17:17try to understand how to actually do DB
- 5:17:20clustering and probably you'll be able
- 5:17:22to understand a lot of things from this
- 5:17:24now in DB scan clustering what are the
- 5:17:27important things so let's start with
- 5:17:29respect to DB scan clustering and let's
- 5:17:32understand some of the important points
- 5:17:33over here the first point that you
- 5:17:35really need to remember is something
- 5:17:37called as score point points I'll also
- 5:17:39talk about when do you say core points
- 5:17:42or when do you say other points as such
- 5:17:44so the first point that I will probably
- 5:17:46discuss about is something called as Min
- 5:17:49points the second point that I will
- 5:17:51probably discuss about is something
- 5:17:53called as score points the third thing
- 5:17:56that I will probably discuss about is
- 5:17:57something called as border points and
- 5:18:00the fourth point that I will definitely
- 5:18:02talk about is something called as noise
- 5:18:04Point okay guys now tell me in C's
- 5:18:07clustering
- 5:18:09if I have this kind of groups don't you
- 5:18:11think with the help of two different
- 5:18:14clusters I may combine this two like
- 5:18:16this with the help of two different
- 5:18:18clusters I may combine something like
- 5:18:22this right but understand over here what
- 5:18:25what problem is basically happening with
- 5:18:27the second clustering this is actually
- 5:18:30an outliers let's say that let's say one
- 5:18:32thing very nicely I will put okay let's
- 5:18:35say I have one point over here I have
- 5:18:38one point over here here so if I do
- 5:18:39clustering probably I will get one
- 5:18:41cluster
- 5:18:43here and I may get another cluster which
- 5:18:45is somewhere here now understand one
- 5:18:47thing this point is definitely an
- 5:18:50outlier even though this is an outlier
- 5:18:53with the help of K means what I'm
- 5:18:54actually doing I'm actually grouping
- 5:18:56this into another group so can we have a
- 5:18:59scenario wherein a kind of clustering
- 5:19:01algorithm is there where we can leave
- 5:19:03the outlier separately and this outlier
- 5:19:06in this particular algorithm and this is
- 5:19:08B basically uh we will be using DB scan
- 5:19:11to relieve the outlier and this point
- 5:19:13will be called as a noisy Point noisy
- 5:19:15point or I can also say it as an outlier
- 5:19:18so this will be a noise point for this
- 5:19:20kind of algorithm where you want to skip
- 5:19:22the outliers we can definitely use DB
- 5:19:25scan that is density based spatial
- 5:19:27clustering of application with noise a
- 5:19:31very amazing algorithm and definitely I
- 5:19:33have tried using this a lot nowadays I
- 5:19:36don't use K means or Hier means instead
- 5:19:38use this kind of algorithm now see this
- 5:19:41what are the important things over here
- 5:19:42first of all you need to go ahead with
- 5:19:44Min points Min points so first thing is
- 5:19:47that you need to have Min points this
- 5:19:50Min points is a kind of
- 5:19:52hyperparameter this basically says what
- 5:19:55does hyper parameter says and there is
- 5:19:57also a value which is called as
- 5:19:59Epsilon which I forgot I will write it
- 5:20:01down over here this is called as Epsilon
- 5:20:04now what does epsilon mean Epsilon
- 5:20:06basically means if I have a point like
- 5:20:08this
- 5:20:09and if I take Epsilon this is nothing
- 5:20:11but the radius of that specific Circle
- 5:20:13radius of that specific Circle okay so
- 5:20:16Epsilon is nothing but radius over here
- 5:20:19in this specific T what does minimum
- 5:20:21points is equal to 4 mean let's say that
- 5:20:24I have I have taken a point over here
- 5:20:26let's say that this is my
- 5:20:28point and I have drawn a circle which
- 5:20:31looks like this and let's say that this
- 5:20:33is my Epsilon
- 5:20:34value okay this is my Epsilon value if I
- 5:20:37say my Min point point is equal to 4
- 5:20:40which is again a hyper
- 5:20:41parameter that basically means I can if
- 5:20:45I have four at least four points over
- 5:20:47here near to this particular Circle
- 5:20:49based on this Epsilon value then what
- 5:20:52will happen is that this point this red
- 5:20:55point will actually become a core
- 5:20:58point a core point which is basically
- 5:21:01given over here if it has at least that
- 5:21:04many number of Min points inside or near
- 5:21:07to this particular within this
- 5:21:09Epsilon okay within this particular
- 5:21:11cluster suppose this is my cluster with
- 5:21:14the help of Epsilon I have actually
- 5:21:15created it is there a particular unit of
- 5:21:17Epsilon or we simply take the unit of
- 5:21:19distance no Epsilon value will also get
- 5:21:21selected through some way I I'll show
- 5:21:23you I'll show you in the practical
- 5:21:24application don't worry now the next
- 5:21:26thing is that let's say let's say I have
- 5:21:28another another point over here let's
- 5:21:30say that I have another point over here
- 5:21:32and this is my circle with respect to
- 5:21:35Epsilon I have created it let's say that
- 5:21:38here I have only one
- 5:21:41point I have only one point inside this
- 5:21:45particular cluster at that point this
- 5:21:48point becomes something called as border
- 5:21:52Point border Point border point also we
- 5:21:55have discussed over here right so border
- 5:21:58point is also there so here I'm saying
- 5:22:00that at least one at least one if it is
- 5:22:04only one it is present then it will
- 5:22:06become a border point if it has Force
- 5:22:08definitely this will become a core Point
- 5:22:10core Point like how we have this red
- 5:22:11color so and there will be one more
- 5:22:14scenario suppose I have this one cluster
- 5:22:16let's say this is my Epsilon and suppose
- 5:22:19if I don't have any points near this
- 5:22:21then this will definitely become my
- 5:22:23noise point and this noise point will
- 5:22:26nothing be but this will be a
- 5:22:28cluster okay so here I have actually
- 5:22:30discussed about the noise point also so
- 5:22:33I hope everybody is able to understand
- 5:22:34the key terms now what is basically
- 5:22:36happening is that whenever we have a
- 5:22:39noise Point like in this particular
- 5:22:40scenario we have a noise point and we
- 5:22:42don't find any points inside this any
- 5:22:45core point or border point if you don't
- 5:22:47find inside this then it is going to
- 5:22:49just get neglected that basically means
- 5:22:52this is basically treated as an outlier
- 5:22:55I hope everybody is able to understand
- 5:22:57here this point will be treated as an
- 5:22:59outlier or it can also be treated as a
- 5:23:02noise point and this will never be taken
- 5:23:05inside a group okay it will never never
- 5:23:08be taken inside a group suppose I have
- 5:23:10this set of points which you see
- 5:23:12basically over here red core and all and
- 5:23:14there is also a border Point by making
- 5:23:18multiple circles over here here you can
- 5:23:20definitely say that how we are defining
- 5:23:22core points and the Border points and
- 5:23:24this can be combined into a single group
- 5:23:27okay this can be combined into a single
- 5:23:29group because how the connection is now
- 5:23:31see this this yellow line is basically
- 5:23:33created by one sorry this yellow point
- 5:23:35is basically created by one Epsilon and
- 5:23:37we have one One Core point over here
- 5:23:40remember over here it should be at least
- 5:23:43one core Point okay not one point but
- 5:23:47one core point at least if it is having
- 5:23:50one core point then it will become a
- 5:23:52border point this will become a border
- 5:23:54point that basically means yes this can
- 5:23:56be the part of this specific group so
- 5:23:59what we are doing Whenever there is a
- 5:24:00noise we are going to neglect it
- 5:24:02wherever there is a broader and core
- 5:24:03points we are going to combine it so
- 5:24:05I'll show you one more diagram which is
- 5:24:06an amazing diagram which will help you
- 5:24:09understand more in this a k means
- 5:24:10clustering and Hier mean clustering now
- 5:24:12see this everybody now the right hand
- 5:24:15side of diagram that you see is based on
- 5:24:19DB scan clustering and the left hand
- 5:24:21side is basically your traditional
- 5:24:23clustering method let's say that this is
- 5:24:25K means which one do you think is better
- 5:24:28over here you see this these all
- 5:24:30outliers are not combined inside a group
- 5:24:34But whichever are nearer as a core point
- 5:24:37and the broader point separate separate
- 5:24:38groups are actually
- 5:24:40created right so this is how amazing a
- 5:24:44DB scan clustering is a DB scan
- 5:24:47clustering is pretty much amazing that
- 5:24:50is basically the outcome of this here in
- 5:24:53C's clustering you can see this all
- 5:24:54these points has also been taken as blue
- 5:24:57color as one group because I'll be
- 5:24:58considering this as one group but here
- 5:25:00we are able to determine this in a
- 5:25:03amazing groups so in I'm saying you guys
- 5:25:06directly use DB scan with without
- 5:25:08worrying about anything so now let's
- 5:25:10focus on the Practical part uh I'm just
- 5:25:12going to give you a GitHub link
- 5:25:14everybody download the code guys I've
- 5:25:16given you the GitHub link quickly
- 5:25:18download and keep your file ready I'm
- 5:25:20going to open my anaconda prompt
- 5:25:23probably open my jupyter notbook we'll
- 5:25:25do one practical problem I've given you
- 5:25:28the link guys please open it so this is
- 5:25:30what we are going to do today this will
- 5:25:32be amazing here you'll be able to see
- 5:25:34amazing things how do you come to know
- 5:25:37that over fitting or underfitting is
- 5:25:39happening you don't know the real value
- 5:25:41right so in in clustering there will not
- 5:25:43be any underfitting or overfitting so uh
- 5:25:46what all things we'll be importing first
- 5:25:48is that we'll try cin clustering we'll
- 5:25:50do silot scoring and then probably we'll
- 5:25:52see the output and um and we'll do DB
- 5:25:56scan Also let's say DB scan is also
- 5:25:58there so uh what are the things we have
- 5:26:00basically imported one is the cin
- 5:26:03clustering one is the Sout samples and
- 5:26:05Sout scores these all are present in the
- 5:26:08SK learn and it is present in metrics
- 5:26:11that basically means we use this
- 5:26:12specific parameter to validate
- 5:26:15clustering models okay now we'll try to
- 5:26:18execute this and apart from that mat
- 5:26:20plot lib we are just trying to import
- 5:26:22numai we are trying to import and all
- 5:26:24here we are executing it perfectly the
- 5:26:26next thing is that here the next step is
- 5:26:29that generating the sample data from
- 5:26:31make underscore blobs first of all we
- 5:26:33are just trying to generate some samples
- 5:26:35with some two features and we are saying
- 5:26:37that okay should have four centroids or
- 5:26:39C centroids itself with some features
- 5:26:43I'm trying to generate some X and Y data
- 5:26:45randomly and this particular data set
- 5:26:47will basically be used in performing
- 5:26:50clustering algorithms okay forget about
- 5:26:52range undor ncore clusters because we
- 5:26:54need to try with different different
- 5:26:55clusters and try to find out the solid
- 5:26:57score so right now I just initialized
- 5:26:59with 2 3 4 5 6 values it is very simple
- 5:27:02so if I go and probably see my X data so
- 5:27:05my X data will look something like this
- 5:27:07so this is my X data with two features
- 5:27:09and this is my Y data with one feature
- 5:27:12which is my output which belongs to a
- 5:27:13specific class okay so that you can
- 5:27:16actually do with the help of make
- 5:27:17underscore blobs let's say how to apply
- 5:27:21kin's clustering algorithm so as I said
- 5:27:23that I will be using W CSS W CSS
- 5:27:26basically means within cluster sum of
- 5:27:28square so I'm going to import K means
- 5:27:30over here for I in range 1A 11 that
- 5:27:33basically means I'm going to use
- 5:27:35different different K values or centroid
- 5:27:37values and try to C which is having the
- 5:27:39minimal wcss value and I'll try to draw
- 5:27:42that graph which I had actually shown
- 5:27:44you with respect to Elbow method so here
- 5:27:47I will basically be also using K means
- 5:27:50number of clusters will be I and
- 5:27:52initialization technique I will will be
- 5:27:54using K means Plus+ so that the points
- 5:27:57the centroids that are initialized those
- 5:27:59those points are very very far and then
- 5:28:01you have random state is equal to zero
- 5:28:03then we do fit and finally we do wcss do
- 5:28:06upend cins doin inertia okay this dot
- 5:28:10inertia will give you the distance
- 5:28:13between the centroids and all the other
- 5:28:16points and this is what I'm going to
- 5:28:18append in this wcss value and finally
- 5:28:20I'll just plot it now here you can see
- 5:28:22that I'm just plotting it obviously by
- 5:28:25seeing this graph this graph looks like
- 5:28:27an elbow okay this graph looks like an
- 5:28:29elbow so the point that I'm actually
- 5:28:31going to consider over here see which is
- 5:28:34the last abrupt change so if I talk
- 5:28:36about the last abrupt change here I have
- 5:28:38the specific value with respect to this
- 5:28:41okay I have one specific value with
- 5:28:43respect to this this is my abrupt change
- 5:28:45from here the changes are normal so I'm
- 5:28:48going to basically select K is equal to
- 5:28:504 now what I'm actually going to do with
- 5:28:52the help of sart with the help of s CL
- 5:28:57score we are going to compare whether K
- 5:29:00is equal to 4 is valid or not so that is
- 5:29:03what we are going to do valid or not so
- 5:29:06here we are going to do this now let's
- 5:29:09go ahead and let's try to see it how we
- 5:29:11are going to do it so here you can see n
- 5:29:13clusters is equal to 4 then I'm actually
- 5:29:15able to find out the prediction and this
- 5:29:17is specifically my output okay this is
- 5:29:19done now see this code okay this code is
- 5:29:23a huge code I have actually taken this
- 5:29:24code directly from the SK learn page of
- 5:29:28Silo if you go and see this this code is
- 5:29:30directly given over there but I'm just
- 5:29:33going to talk about like what are the
- 5:29:35important things we need to see over
- 5:29:37here with respect to different different
- 5:29:39clusters see see this clusters 2 3 4 5 6
- 5:29:43I'm going to basically compare whether
- 5:29:46the K value should be four or not with
- 5:29:48the help of solid scoring so let's go
- 5:29:51here and here you can see that I'm
- 5:29:54applying this one first I will go with
- 5:29:56respect to for Loop for ncore clusters
- 5:29:59in range underscore clusters different
- 5:30:00different cluster values are there first
- 5:30:02we'll start with two so here you can see
- 5:30:04initialize the cluster with and cluster
- 5:30:06value and a random generator seed of 10
- 5:30:09for reproducibility so ncore clusters
- 5:30:12first I take took it as two and then I
- 5:30:14did fit predict on X after I did fit
- 5:30:17predictor on X I'm using this score on X
- 5:30:21comma cluster label now what this is
- 5:30:23going to do understand in Solo what did
- 5:30:25we discuss it will it will try to find
- 5:30:27out all the Clusters the Clusters over
- 5:30:30here like this and it'll try to
- 5:30:32calculate the distance between them
- 5:30:34which is the a of I then it'll try to
- 5:30:36compute the B of I then finally it'll
- 5:30:39try to compute the score and if the
- 5:30:41value is between minus1 to +1 the more
- 5:30:43the Valu is towards + one the more
- 5:30:45better it is right so these all things
- 5:30:47we have already discussed and that is
- 5:30:49what this specific function will do and
- 5:30:51this will give my solo average value
- 5:30:53over here solid value will be over here
- 5:30:55okay this we have done and then we can
- 5:30:58continuously do it for another another
- 5:31:00things you can actually find it over
- 5:31:02here and this value that you see this
- 5:31:05code that you see is nothing nothing so
- 5:31:08complex okay this is just to display the
- 5:31:11data properly in the form of graphs okay
- 5:31:15in the form of graphs so again I'm
- 5:31:17telling you I did not write this code
- 5:31:18I've directly taken it from the uh SK
- 5:31:22learn page of solid okay so just try to
- 5:31:25see this particular uh plotting diagrams
- 5:31:27and all that you can definitely figure
- 5:31:29out but let's see I will try to execute
- 5:31:31it and try to find out the output now
- 5:31:33see for ncore cluster is equal to 2 the
- 5:31:37average solid score is 70 I told you the
- 5:31:40value will be between -1 to +1 and I'm
- 5:31:43actually getting 704 which is very very
- 5:31:46good and then for ncore cluster is equal
- 5:31:48to 3 588 then ncore cluster is equal to
- 5:31:524 I'm getting 65 which is pretty much
- 5:31:54amazing and then for ncore cluster equal
- 5:31:57to 5 the average score is 563 and ncore
- 5:32:00cluster is equal to 6 you are saying
- 5:32:02.45 here directly you can actually say
- 5:32:05that fine for _ cluster equal to 2 I'm
- 5:32:08getting an amazing score of
- 5:32:10704 obviously you're you're getting the
- 5:32:12highest value over this so should we
- 5:32:14select ncore cluster isal to two Okay we
- 5:32:17should not directly conclude from it
- 5:32:19because here we need to also see that
- 5:32:21any feature value or any cluster value
- 5:32:24is also coming as negative value that
- 5:32:26also we need to check so here we will go
- 5:32:28down over here you will see the first
- 5:32:30one over here with respect to the first
- 5:32:32one you see that I'm get getting the
- 5:32:35value from 0 to 1 it is not going going
- 5:32:38to Min -.1 so definitely two clusters
- 5:32:40was able to solve the problem so I'll
- 5:32:43keep it like this with me I definitely
- 5:32:45have a chance that this may this may
- 5:32:48perform well I may have a chance that
- 5:32:50this K uh K is equal to 2 May perform
- 5:32:53well okay so I may have a chance let's
- 5:32:55see to the next one to the next one over
- 5:32:57here you can see that for one of the
- 5:32:59cluster the value is negative if the
- 5:33:01value is negative that basically means
- 5:33:03the AI is obviously greater than b ofi
- 5:33:06so I'm not going to prer this because it
- 5:33:08is having some negative values even
- 5:33:10though my cluster looks better but again
- 5:33:13understand what is the problem with
- 5:33:14respect to this cluster is that if I
- 5:33:17take this cluster and probably compute
- 5:33:19the distance between this point to this
- 5:33:20point and if I probably compute from
- 5:33:22this point to this point or this point
- 5:33:24to this point this point is obviously
- 5:33:26nearer to this right it is obviously
- 5:33:29nearer to this so that is the reason why
- 5:33:31I'm getting a negative value over here
- 5:33:33okay negative value over here this is my
- 5:33:36uh output my score this point that you
- 5:33:40see dotted points this is my score 58
- 5:33:43what whatever it is this is basically my
- 5:33:45score so obviously this basically
- 5:33:46indicates that this point is near the
- 5:33:48other cluster point is nearer to this so
- 5:33:50I'm actually getting a negative value
- 5:33:52right so this you really need to
- 5:33:54understand okay now similarly if I go
- 5:33:56with respect to ncore Cluster is equal
- 5:33:58to 4 this looks good because here I
- 5:34:00don't have any negative value and here
- 5:34:03you can see how cooly it has basically
- 5:34:06divided the points amazing inly with the
- 5:34:08help of k equal to 4 right and similarly
- 5:34:11if I go with five obviously you can see
- 5:34:13some negative values are here some
- 5:34:15dotted line negative value are there
- 5:34:17with respect to six you also have some
- 5:34:18negative values so definitely I'll not
- 5:34:21go with six I may either go with four or
- 5:34:23I may either go with two now whenever
- 5:34:26you have this options always take a
- 5:34:27bigger number instead of two take four
- 5:34:30because four is greater than two because
- 5:34:32it will be able to create a generalized
- 5:34:34model so from this I'm actually going to
- 5:34:37take and is equal to 4 K is equal to 4
- 5:34:39now should we compare with this with the
- 5:34:41elbow method here also I got four right
- 5:34:44so both are actually matching so this
- 5:34:47indicates that with the help of this
- 5:34:49clustering this siluette score we can
- 5:34:51definitely come to a conclusion and
- 5:34:53validate our clustering model in an
- 5:34:55amazing way so I hope everybody is able
- 5:34:57to understand and this way you basically
- 5:35:00validate a model and definitely you can
- 5:35:03try it out you can understand this code
- 5:35:04definitely I but till here you have
- 5:35:06understood that here I'm going to get
- 5:35:08the average value then for iore clusters
- 5:35:12whatever cluster this is matching it is
- 5:35:14just mapping over there and it is
- 5:35:16basically giving so this was the session
- 5:35:19and uh yes in today's session we
- 5:35:22efficiently covered many topics we
- 5:35:24covered kin hierle clustering solid
- 5:35:27score DB clustering in tomorrow's
- 5:35:29session the topics that are probably
- 5:35:31pending is first I'll start with svm and
- 5:35:34svr second I will go ahead with XG boost
- 5:35:37and and third I will cover up PCA let's
- 5:35:40see whether I'll be able to complete
- 5:35:41this session uh one one amazing thing
- 5:35:45that I want to teach you guys because
- 5:35:46many people ask me the definition of
- 5:35:48bias and variance so guys uh many people
- 5:35:52get confused when we talk about bias and
- 5:35:55variance you know because let's say that
- 5:35:57uh I have a model for the training data
- 5:35:59set it gives us somewhere around 90%
- 5:36:02accuracy let's say I'm getting a 90%
- 5:36:04accuracy for the test data I may
- 5:36:07probably getting somewhere around 70%
- 5:36:10accuracy now tell me which scenario is
- 5:36:12basically this most of the people will
- 5:36:14be saying that okay fine it is
- 5:36:16overfitting now when I say overfitting I
- 5:36:19basically mention overfitting by low
- 5:36:23bias and high
- 5:36:25variance right so many people get
- 5:36:28confused Krish tell me just the exact
- 5:36:30definition of bias and variance low bias
- 5:36:33obviously you are saying that because
- 5:36:34the training is performed like the model
- 5:36:37is performing well with the help of
- 5:36:39training data set but with respect to
- 5:36:41the test data set the model is not
- 5:36:43performing well with respect to training
- 5:36:45data set why do we always say bias and
- 5:36:48with respect to test data set why do we
- 5:36:50always say variance so for this you need
- 5:36:52to understand the definition of bias so
- 5:36:54let me write down the definition of bias
- 5:36:56over here so here I can definitely write
- 5:36:59that bias it is a
- 5:37:02phenomena that
- 5:37:05skews the
- 5:37:08result of an
- 5:37:13algorithm in
- 5:37:15favor in favor or against an
- 5:37:20idea against an idea I'll make you
- 5:37:23understand the definition uh um but
- 5:37:27understand the understand understand
- 5:37:28what I have actually written over here
- 5:37:30it is a phenomena that skewes the result
- 5:37:32of an algorithm in favor or against an
- 5:37:34idea whenever I say this specific idea
- 5:37:37this idea I will just talk about the
- 5:37:39training data set initially now when we
- 5:37:42train a specific model suppose if I have
- 5:37:44this specific model over
- 5:37:46here and I'm training with this specific
- 5:37:49training data set so this is my training
- 5:37:52data set now based on the definition
- 5:37:54what does it basically say it is a
- 5:37:55phenomenon that skews the result of an
- 5:37:57algorithm in favor or against an idea or
- 5:38:00a this specific training data set so
- 5:38:03even though I'm training this particular
- 5:38:04model with this training data set
- 5:38:07with this data set it may it may be in
- 5:38:11favor of that or it may be against of
- 5:38:12that that basically means it may perform
- 5:38:14well it may not perform well if it is
- 5:38:15not performing well that basically means
- 5:38:17the accuracy is down if the accuracy is
- 5:38:19better at that point of time what will
- 5:38:21say see if the accuracy is better that
- 5:38:23time what we'll say we we'll come up
- 5:38:25with two terms from here obviously you
- 5:38:27understand okay there are two scenarios
- 5:38:28of bias now here if it is in favor that
- 5:38:32basically means it is performing well
- 5:38:33with respect to the training data set I
- 5:38:35will basically say that it has high bu
- 5:38:38if it is not able to perform well with
- 5:38:40the training data set then here I will
- 5:38:42say it as low
- 5:38:44bias I hope everybody is able to
- 5:38:46understand in this specific thing
- 5:38:47because many many many people has this
- 5:38:49kind kind of confusion now similarly if
- 5:38:51I talk about variance let's say about
- 5:38:53variance because you need to understand
- 5:38:55the definition a definition is very much
- 5:38:59important okay if I if I just talk about
- 5:39:01the definition of variance I'm just
- 5:39:03going to refer like this the variance
- 5:39:07refers to the changes in the model when
- 5:39:13using when using different
- 5:39:16portion of the
- 5:39:20training or test
- 5:39:23data now let's understand this
- 5:39:25particular
- 5:39:27definition variance refers to the
- 5:39:29changes in the model when using
- 5:39:31different proportion of the test
- 5:39:32training data or test data we obviously
- 5:39:34know that whenever initially if I have a
- 5:39:38model understand from the definition
- 5:39:39everything will make sense I am
- 5:39:41basically training initially with the
- 5:39:43training
- 5:39:44data okay because we divide our data set
- 5:39:47see our data set whenever we are working
- 5:39:49with we divide this into two parts one
- 5:39:52is our train data and test data okay
- 5:39:56because this is a tra test data is a
- 5:39:58part of that particular data set right
- 5:40:00and suppose in this particular training
- 5:40:02data it gets trained and performs well
- 5:40:04here I'm actually talking about bias but
- 5:40:07when we come with respect to the
- 5:40:09prediction of the specific model at that
- 5:40:12point of time I can use other training
- 5:40:14data that basically means that training
- 5:40:15data may not be similar or I can also
- 5:40:18use test data now in this test data what
- 5:40:21we do we do some kind of predictions
- 5:40:23these are my predictions and in this
- 5:40:25prediction again I may get two
- 5:40:27scenario I may get two scenario which is
- 5:40:30basically mentioned by variance it
- 5:40:32refers to the changes in the model when
- 5:40:34using when using different portion of
- 5:40:37the training or test data refers to the
- 5:40:40changes basically means whether it is
- 5:40:42able to give a good prediction or wrong
- 5:40:44predictions that's it so in this
- 5:40:46particular scenario if it gives a good
- 5:40:48prediction I may definitely say it as
- 5:40:50low variance that basically means the
- 5:40:53accuracy with the accuracy with respect
- 5:40:55to the test data is also very good if I
- 5:40:58probably get a bad if I probably get a
- 5:41:01bad accuracy at that time I basically
- 5:41:04say it as high variance so if I talk
- 5:41:06about three scenarios over here let's
- 5:41:08say this is my model one and this is my
- 5:41:11model
- 5:41:12two and this is my model
- 5:41:15three now in this scenario let's
- 5:41:18consider that my model one has the
- 5:41:21training
- 5:41:23accuracy of 90% and test accuracy of
- 5:41:3075% similarly I have here as my train
- 5:41:33accuracy of 60% and my test accuracy
- 5:41:38of
- 5:41:4055% now similarly if I have my train
- 5:41:44accuracy of 90% And my test accuracy of
- 5:41:4992% now tell me what what things you
- 5:41:51will be getting here obviously you can
- 5:41:54directly say that fine your training
- 5:41:56accuracy is better now you're talking
- 5:41:58about bias so this basically indicates
- 5:42:00that this has low
- 5:42:02bias and since your test accuracy is bad
- 5:42:07because it is when compared to the train
- 5:42:08accuracy it is less so here you are
- 5:42:10basically going to say high
- 5:42:14variance understand with respect to the
- 5:42:16definition similarly over here what
- 5:42:18you'll say high
- 5:42:20bias High variance because obviously it
- 5:42:22is not performing
- 5:42:25well this is another scenario last the
- 5:42:28last scenario is that this is the
- 5:42:30scenario that we want because it is low
- 5:42:33bias and low variance
- 5:42:37okay many many people have basically
- 5:42:39asked me the definition with respect to
- 5:42:41bias and variance and here I've actually
- 5:42:43discussed and this indicates this gives
- 5:42:45me a generalized model and this is what
- 5:42:50is our aim when we are working as a data
- 5:42:53scientist so I hope you have understood
- 5:42:55the basic difference between V bias and
- 5:42:57variance and I was able to give you lot
- 5:43:00of examples lot of understanding with
- 5:43:02respect to this so I hope you have
- 5:43:05actually got this particular uh
- 5:43:08understanding of this uh two terms which
- 5:43:10we specifically talk about high bias low
- 5:43:12bias High variance low variance right so
- 5:43:16this was it from my side guys uh and uh
- 5:43:19I hope you have understood
- 5:43:22this
- 5:43:29okay so let's take let's consider a data
- 5:43:34set credit
- 5:43:37and let's say this is a
- 5:43:39approval so we are going to take this
- 5:43:42sample data set and understand how does
- 5:43:43XG boost work suppose salary is less
- 5:43:47than or equal to 50 and the credit is
- 5:43:50bad so approval the loan approval will
- 5:43:52be zero that basically means he he or
- 5:43:54she will not get if it is less than or
- 5:43:56equal to 50 if the credit score is good
- 5:44:00then probably approval will be one if it
- 5:44:02is less than or equal to 50 if it is
- 5:44:06good
- 5:44:07again then it is going to get one if it
- 5:44:10is greater than
- 5:44:1250 and if it is bad then obviously
- 5:44:16approval will be
- 5:44:19zero if it is greater than
- 5:44:2250 if it is good we are going to get it
- 5:44:25as one if it is greater than
- 5:44:2950k and probably if it is normal then
- 5:44:33also we are going to get
- 5:44:35it so this is this is my data set so how
- 5:44:38does XG boost classifier work understand
- 5:44:41the full form of XG boost is
- 5:44:44Extreme gradient
- 5:44:47boosting extreme gradient boosting so we
- 5:44:50will basically understand about extreme
- 5:44:52gradient boosting now extreme gradient
- 5:44:55boosting uh will be actually used to
- 5:44:58solve both classification and the
- 5:45:00regression problem statement so first of
- 5:45:02all let's understand how it is basically
- 5:45:05exib basically how it actually if you if
- 5:45:08you just talk about XG boost you
- 5:45:10understand that it is a boosting
- 5:45:11technique and internally it tries to use
- 5:45:13decision tree so how does this decision
- 5:45:16Tre is basically getting constructed in
- 5:45:18the case of XV boost and how it is
- 5:45:20basically solved we are going to discuss
- 5:45:21about it so whenever we start exib boost
- 5:45:24classifier understand that first of all
- 5:45:26we create a specific base model suppose
- 5:45:29if I say this is my base model and this
- 5:45:32base model will be a weak learner okay
- 5:45:36and this base model will always give an
- 5:45:38output of probability of 0.5 in the case
- 5:45:42of classification problem so suppose if
- 5:45:45I say this is probability 0.5 then I
- 5:45:48will try to create a field over here
- 5:45:50this field is called as residual field
- 5:45:53so first base model what I'm going to do
- 5:45:55any data set that you give from here to
- 5:45:57train it will always give you the output
- 5:45:59as 0.5 so this is just a dummy base
- 5:46:02model now tell me if my probability
- 5:46:06output is is 0.5 if I want to calculate
- 5:46:08the residual that basically means I need
- 5:46:10to subtract approval minus this
- 5:46:12particular value so what will be the
- 5:46:16value over here 0 -.5 will be
- 5:46:20-.5 1 -.5 will be5 1 -.5 will
- 5:46:25be5 and 0 -.5 will be -.5 and this 1 -.5
- 5:46:31will
- 5:46:32be uh 0.5 and this will also be 0.5
- 5:46:36let's consider that I have one more
- 5:46:37record uh and this specific record can
- 5:46:40be anything uh because I want to keep
- 5:46:43some more records over here so let's
- 5:46:45consider that I have one more record
- 5:46:46which is less than or equal to 50K and
- 5:46:49if the credit scod is normal you're
- 5:46:51going to get zero so here also if I try
- 5:46:53to find out the residual it will be
- 5:46:55minus5 now the first step I hope
- 5:46:58everybody's understood we have to create
- 5:46:59a base model okay this base model is
- 5:47:01very much important because we have to
- 5:47:04create all the decision Tree in a
- 5:47:06sequential manner so the first
- 5:47:09sequential base tree which is again this
- 5:47:11is also a decision tree kind of thing
- 5:47:12you can consider but this is a base
- 5:47:15model which takes any inputs and gives
- 5:47:17by default the probability as 05 now
- 5:47:20let's go ahead and understand what are
- 5:47:22the steps in constructing decision tree
- 5:47:24after creating the base model the first
- 5:47:27step is that create uh binary decision
- 5:47:31tree so I'm going to write it down all
- 5:47:34the steps please make sure that you note
- 5:47:35it down so so create a binary tree
- 5:47:39binary decision tree using the features
- 5:47:43second step we basically Define we we we
- 5:47:47say it as okay Second Step what we do we
- 5:47:50actually calculate the similarity weight
- 5:47:54we calculate the similarity weight I'll
- 5:47:57talk about this similarity weight what
- 5:47:59exactly it is if I want to use this a
- 5:48:02formula it is summation of residual
- 5:48:05Square
- 5:48:07divided
- 5:48:08by summation of probability 1 minus
- 5:48:13probability plus Lambda I'll talk about
- 5:48:16this what is exactly Lambda it is the
- 5:48:18kind of hyperparameter again so that it
- 5:48:20does not overfit the third thing is that
- 5:48:23we calculate the Information Gain okay
- 5:48:26Information Gain so these are the steps
- 5:48:28we basically use in constructing or in
- 5:48:32solving uh in creating an HD boost
- 5:48:34classifier the first step is that we
- 5:48:36create a inary decision tree using the
- 5:48:38feature then we go ahead with
- 5:48:40calculating the similarity weight and
- 5:48:42finally we go ahead and calculate the
- 5:48:43information gain so how does it go ahead
- 5:48:46let's understand over here and let's try
- 5:48:47to find out okay now let's go ahead and
- 5:48:50let's try to construct the decision tree
- 5:48:53as I said that let's consider that I'm
- 5:48:55considering salary feature So based on
- 5:48:58using salary feature what I'm actually
- 5:48:59going to do I am going to take this as
- 5:49:02my node and I'm going to split this up
- 5:49:05and remember whenever we are creating
- 5:49:07decision Tree in this particular case it
- 5:49:09will be a binary decision tree let's say
- 5:49:13that in salary one is less than or equal
- 5:49:15to one is greater than 50 so this two
- 5:49:18you obviously have in the case of binary
- 5:49:20in case of credit where there are three
- 5:49:22categories I'll also show you how that
- 5:49:25further split will happen and how that
- 5:49:27will get converted into a binary team so
- 5:49:29here you have less than or equal to 50K
- 5:49:32and greater than 50k now let's go ahead
- 5:49:35and understand how many vales are there
- 5:49:37in this salary so if I see before the
- 5:49:40split you can definitely see that I'm
- 5:49:42going to use this residual and probably
- 5:49:45train this entire model now if I really
- 5:49:48wanted to find out the residual
- 5:49:49initially these are my residuals over
- 5:49:51here so one resid is -.5 then I have 0.5
- 5:49:56over here then I have .5 then again I
- 5:49:59have -.5 then again I have 0.5 then
- 5:50:03again I have 0.5 and finally I have
- 5:50:06minus .5 so these are my total residuals
- 5:50:09that are there suppose if I make this
- 5:50:11split less than or equal to 50 First
- 5:50:14less than or equal to 50 the residuals
- 5:50:16what are things are there so here I'm
- 5:50:18going to have minus5 then less than or
- 5:50:21equal to 50 again I'm going to have 05
- 5:50:23then again less than or equal to 50 I'm
- 5:50:25going to have 0.5 and less than or equal
- 5:50:27to again one more 0.5 is there I'm just
- 5:50:30going to remove this the last5 which is
- 5:50:33nothing but Min -.5 so I hope you
- 5:50:35understood this split so half of the
- 5:50:37things came over here the remaining half
- 5:50:40will be greater than or equal to greater
- 5:50:41than 50 so you have one value here one
- 5:50:44value here one value here so it will be
- 5:50:46Min -.5 then you have 0.5 and then
- 5:50:50finally you have 0.5 residuals how do we
- 5:50:53get it guys see from the base model
- 5:50:55which is by default giving 0.5 first my
- 5:50:58data goes over here by default
- 5:51:00probability I'm going to get 0.5 so
- 5:51:02residual is basically calculated from
- 5:51:04this probability and approval so this
- 5:51:07probability minus approval so if you
- 5:51:09subtract 0 -.5 sorry I'm just going to
- 5:51:12rub this so if you subtract 0 -.5 you're
- 5:51:16going to get -.5 1 -.5 you're going to
- 5:51:19get .5 1 -.5 you're going to get .5 so
- 5:51:22everybody I hope is very much clear with
- 5:51:24respect to this so this is the first
- 5:51:26step we constructed a binary tree now in
- 5:51:28the second step it says calculate the
- 5:51:30similarity weight now how to calculate
- 5:51:33the similarity weight similarity weight
- 5:51:35formula is sum of residual Square now
- 5:51:37what is residual Square let's say that
- 5:51:39I'm going to calculate the the the uh
- 5:51:43I'm going to calculate for this okay
- 5:51:45similarity weight now in this particular
- 5:51:47case if I go and calculate my similarity
- 5:51:49weight it will be summation of residual
- 5:51:52Square this is my residual values this
- 5:51:55is my residual Valu so I'm going to do
- 5:51:57the summation of this Square okay this
- 5:52:01value square you can see over here sum
- 5:52:03of residual Square everybody you can see
- 5:52:06sum of of residual squares so what do
- 5:52:08you think sum of residual squares will
- 5:52:09be in this particular case how I have to
- 5:52:12do it I will just take up this all
- 5:52:14values like
- 5:52:16-.5
- 5:52:17+5
- 5:52:20+5 and
- 5:52:22-.5 whole square right I'm just going to
- 5:52:24do the squaring of this divided by
- 5:52:27understand what it is divided by it is
- 5:52:29divided by probability of 1 minus
- 5:52:31probability now where do we get this
- 5:52:33probability value where do we get this
- 5:52:35probability value value we get this
- 5:52:37probability value from our base model
- 5:52:40right so here I'm basically going to say
- 5:52:42that we are going to do the summation of
- 5:52:44probability of 1 minus probability 1
- 5:52:47minus probability that basically means
- 5:52:50for each and every point for each and
- 5:52:52every Point what is the probability see
- 5:52:54probability is basically coming from the
- 5:52:56base model so for each Pro each point
- 5:52:59I'm going to come compute two things one
- 5:53:01is the probability and then 1 minus
- 5:53:04probability and this I'm going to do the
- 5:53:06summ
- 5:53:07like this I will do it four times 1 -.5
- 5:53:10then .5 * 1 -.5 and finally you'll be
- 5:53:15able to see one more will be there which
- 5:53:17is
- 5:53:18+5 1 -.5 so this will be your total
- 5:53:21things with respect to this so I hope
- 5:53:24you have understood till here uh where
- 5:53:26you are able to understand that what we
- 5:53:28have done this is summation of uh
- 5:53:31residual square and this is the
- 5:53:33remaining probability multiplied by 1
- 5:53:35minus probability now tell me what are
- 5:53:39you able to find out from this if you
- 5:53:41cancel this and this this and this this
- 5:53:44value is going to become zero so this
- 5:53:47entire value is going to become Zer
- 5:53:48because 0 divided by anything is 0er so
- 5:53:51here I hope everybody is understood what
- 5:53:53is the similarity weight of this
- 5:53:55specific node if I want to write it is
- 5:53:57nothing but zero now you may be
- 5:53:59considering where is Lambda
- 5:54:01value okay we will initially initialize
- 5:54:04Lambda by 1 I'll talk about this hyper
- 5:54:05parameter let's consider it as 1 so here
- 5:54:09+ 1 or plus 0 let's let's consider
- 5:54:12Lambda value 0 let's say for right now
- 5:54:14okay I'm just going to make it Lambda is
- 5:54:16equal to0 I'm just going to talk about
- 5:54:19it because it is a kind of hyper
- 5:54:21parameter by Z -.5 -.5 +5 +5 if I do the
- 5:54:28summation if I do the summation here you
- 5:54:31will be able to see that I'm going to
- 5:54:32get zero so this calculation we have
- 5:54:34done and we have got uh the sumission of
- 5:54:36weight is equal to Z and let's go ahead
- 5:54:39and calculate the sumission of the
- 5:54:40weight of the next node no no no it's
- 5:54:43not first Square it is whole squar so
- 5:54:46here also if I do so it is5 +5 now let's
- 5:54:51do it for this if I want to find out the
- 5:54:53similarity weight again see I'm going to
- 5:54:55repeat it .5 +5 whole squ and since
- 5:55:00there are three points so I'm going to
- 5:55:01basically use probability 1 minus
- 5:55:04probability for one point then plus
- 5:55:08probability 1 minus probability second
- 5:55:11point and then probability and 1 minus
- 5:55:14probability for the third point and
- 5:55:16Lambda is zero so I'm not going to write
- 5:55:18anything now go let's go and do the
- 5:55:20calculation for this node so - 5 - 5 it
- 5:55:24becomes zero then .5 whole square right
- 5:55:27so here I'm going to get 0.25 here if
- 5:55:30you do the calculation here you are
- 5:55:31going to get 75 so this value is going
- 5:55:34to be 1x3 and which is nothing at33 so
- 5:55:37the similarity weight for this node for
- 5:55:40this node
- 5:55:42is33 so here you can see probability of
- 5:55:45multiplied by 1 minus
- 5:55:47probability okay now the next step that
- 5:55:50we do is that calculate the information
- 5:55:53gain now you know how to calculate the
- 5:55:55information gain but before that let's
- 5:55:57do the computation for this also for
- 5:55:59this root node also go ahead and
- 5:56:01calculate the similarity weight of
- 5:56:04this okay they
- 5:56:06why the base model probability is5
- 5:56:09because it is just understand that it is
- 5:56:11a dummy dummy model I have just put a if
- 5:56:14condition there saying that it is going
- 5:56:15to give 0.5 now do it for this one guys
- 5:56:17root node what it will be see I can
- 5:56:20calculate from here only minus1 gone
- 5:56:23this is also gone this is also gone this
- 5:56:25will be .25 divided by something now
- 5:56:29tell me guys what should be for the root
- 5:56:32node what is the similarity similarity
- 5:56:34weight what is the similarity weight for
- 5:56:36for this do this calculation everyone up
- 5:56:39one I know it will be. 25 divided by
- 5:56:44this will be 1.75 are you getting this
- 5:56:48similarity weight which will be nothing
- 5:56:50but 1 by 7 and if I divide 1 by 7 if I
- 5:56:54say what is 1 by 7 it
- 5:56:57is42 so it is nothing but .14 if I want
- 5:57:00to calculate the root node similarity
- 5:57:02weight over here
- 5:57:05is4 so I know 0.14 here 0 here 33 now
- 5:57:09see over here we calculate the
- 5:57:11Information Gain Next Step the third
- 5:57:13step what we do is that we calculate the
- 5:57:15information gain now Information Gain is
- 5:57:19nothing but in this particular case the
- 5:57:21root node similarity weight we'll try to
- 5:57:24add up so I will be getting
- 5:57:270.33 minus this particular Top Root node
- 5:57:31whatever split has happened that
- 5:57:33similarity weight I'll take 0 +33
- 5:57:36-14 so Point
- 5:57:39-14 and if I do it it is nothing but
- 5:57:42just open your calculator again and
- 5:57:4633
- 5:57:48-14 so it is nothing but .19 I'm getting
- 5:57:52.19 as my information gain the
- 5:57:56information gain of this specific tree I
- 5:57:59got it
- 5:58:00as19 obviously you know how the features
- 5:58:03will get selected based on the
- 5:58:06Information Gain but let's say that the
- 5:58:08highest Information Gain that is given
- 5:58:10by salary okay now we will go ahead and
- 5:58:13do the further split let's go ahead and
- 5:58:16do the further split so I I know my
- 5:58:18information gain now it is1 n and
- 5:58:20Information Gain is basically used to
- 5:58:23select that specific node through which
- 5:58:26the split will happen now I'll further
- 5:58:27go and do the split let's say that I'm
- 5:58:29going to do the further split with the
- 5:58:31next feature that is which one credit so
- 5:58:33I'm going to take credit over here I'm
- 5:58:36going to take credit over here and again
- 5:58:39I have to do a binary split again but
- 5:58:42you may be considering chish here are
- 5:58:43only three categories how we are going
- 5:58:45to basically do this particular split
- 5:58:48right because we don't know how to do
- 5:58:50the split because we have three
- 5:58:51categories over here so in this case
- 5:58:53what I will do is that we what we can
- 5:58:56definitely do is that in this particular
- 5:58:58case the split that we are probably
- 5:59:00going to do is that let's consider two
- 5:59:02categories like good and normal at one
- 5:59:04side bad at one side so here it becomes
- 5:59:06a binary split again now let's go ahead
- 5:59:09and let's try to see that how many data
- 5:59:11points will fall here and how many data
- 5:59:12points will fall here so for writing
- 5:59:14down the data points let's say if it is
- 5:59:17less than or see go to the path if it is
- 5:59:19less than or equal to 50 it'll go this
- 5:59:21path and if it is B then we are probably
- 5:59:24going to get how much is the residual we
- 5:59:26are going to get one residual over here
- 5:59:28first of all so this is my one residual
- 5:59:31that is -.5 then similarly if I see less
- 5:59:34than or equal to 50 good is there right
- 5:59:37good or normal is there so here again 0.
- 5:59:39five will come I hope everybody is able
- 5:59:42to understand see the second record less
- 5:59:44than or equal to 50 we go in this path
- 5:59:45but it is good we come over here again
- 5:59:48less than or equal to 50 good again we
- 5:59:50are going to get 1
- 5:59:51more5 then go with respect to greater
- 5:59:55than or equal to 50 which is coming over
- 5:59:57here we'll not worry about it right now
- 5:59:59again less than or equal to 50 normal
- 6:00:01again it is
- 6:00:03-.5 right so this many records
- 6:00:06definitely coming over here only one
- 6:00:08record is basically coming over here
- 6:00:10then again we will start the same
- 6:00:12process again we will start the same
- 6:00:14process now for the same process what we
- 6:00:16are going to do again try to calculate
- 6:00:18the similarity weight now in order to
- 6:00:20calculate the similarity weight what I
- 6:00:22will do I will basically say this is my
- 6:00:24similarity weight this will become .25
- 6:00:28divided 025 why because this whole
- 6:00:31square right this whole Square residual
- 6:00:33square right summation of residual
- 6:00:36square but here I have only one residual
- 6:00:38so this Square it will become and then
- 6:00:40what I'm actually going to do I'm going
- 6:00:41to basically write .5 - 1 -.5 this is
- 6:00:45nothing for only for one data point so
- 6:00:47this is nothing but .5 * .5 which is
- 6:00:50nothing but 0.25 right now in this
- 6:00:53particular case I will get similarity
- 6:00:54weight as I hope everybody I'm getting
- 6:00:56it as one now what about this similarity
- 6:00:58weight if you want to compute it is
- 6:01:00again very very simple this and this
- 6:01:02will get cancelled then again it will be
- 6:01:03025 divided by um if I say one like this
- 6:01:08.25 then again it will be 75 then this
- 6:01:11will also be 1 by3 that is nothing but
- 6:01:1333 so similarity weight will
- 6:01:16be33 then again I have to calculate the
- 6:01:19information gain of this node what I
- 6:01:21will do I will add this up see 1
- 6:01:24+33 I'll add like 1
- 6:01:27+33 minus 0 why zero because the
- 6:01:30information gain the similarity weight
- 6:01:32of this uh the up one is basically 0
- 6:01:37right for this particular credit node
- 6:01:39similarity weight is zero so 1
- 6:01:41+33 minus 0 this will be 1.33 so like
- 6:01:45this further split will again happen
- 6:01:47over here with different different node
- 6:01:49and we will only be getting a binary
- 6:01:51split but we will be comparing based on
- 6:01:54Information Gain which one is coming
- 6:01:55good now let's say that I have created
- 6:01:57this path I have I have designed I have
- 6:02:00developed my entire binary decision tree
- 6:02:02which is a speciality in XG boost now
- 6:02:06what I'm going to do over here is that
- 6:02:08see everybody what I'm going to do let's
- 6:02:10consider the inferencing part let's say
- 6:02:12this record is going to go how we are
- 6:02:15going to calculate the output so this
- 6:02:17first of all went to this base model now
- 6:02:21let's go ahead and see how the
- 6:02:22inferencing will happen suppose This
- 6:02:24Record is going right so first of all
- 6:02:26this record will go to this base model
- 6:02:29the base model is giving the probability
- 6:02:30as 0.5 so the first base model is
- 6:02:34basically giving 0.5 now base based on
- 6:02:36this 05 how do we calculate the real
- 6:02:39probability how do we calculate the real
- 6:02:41probability in this okay so we apply
- 6:02:43something called as logs so we basically
- 6:02:45say log of P / 1us P so this is the
- 6:02:49formula we basically apply in only the
- 6:02:52case of base model so if we try to see
- 6:02:55this it is nothing but log
- 6:02:57of5 / .5 which is nothing but zero log
- 6:03:01of one is nothing but zero so in the
- 6:03:03first case whenever any record goes I
- 6:03:05will be getting the zero value over here
- 6:03:08okay zero value over here then plus why
- 6:03:11plus I'm doing because it will now go to
- 6:03:13the binary decision tree now this record
- 6:03:15will go to my binary decision Tre
- 6:03:17whatever value I'm getting from this I'm
- 6:03:19actually adding that up and now it will
- 6:03:21go over here now when it goes over here
- 6:03:24first of all let's see which branch it
- 6:03:25is following it is following less than
- 6:03:27or equal to 50 Branch first Branch over
- 6:03:29here then this is bad it'll go and
- 6:03:32follow here so here I can see that the
- 6:03:34similarity weight is one now the
- 6:03:36similarity weight is basically one in
- 6:03:38this case so what we do in the case of
- 6:03:40this we pass it to a learning rate
- 6:03:44parameter so this specifically is my
- 6:03:46learning rate multiplied by 1 one
- 6:03:49because why similarity weight is one
- 6:03:51over here so this will basically be my
- 6:03:54first references and Alpha over here is
- 6:03:57my learning rate it can be a very small
- 6:03:59value based on the learning parameter
- 6:04:01that we use like how we have defined
- 6:04:04learning parameters elsewhere on top of
- 6:04:06this we apply an activation function
- 6:04:09which is called as sigmoid since this is
- 6:04:11a classification problem we apply an
- 6:04:14activation function which is called as
- 6:04:15sigmoid and I hope you know what is the
- 6:04:17use of sigmoid based on this based on
- 6:04:20the alpha value based on this the output
- 6:04:22will be between 0 to 1 now I hope you
- 6:04:25getting it guys this is how the entire
- 6:04:27inferencing will probably happen now
- 6:04:30similarly what I will do I will try to
- 6:04:32construct this kind of decision tree
- 6:04:33parall so we we can also write our
- 6:04:37entire function will look something like
- 6:04:40this Alpha 0 + alpha 1 and this will be
- 6:04:46your decision tree 1 output then Alpha 2
- 6:04:50your decision tree output Alpha 3 your
- 6:04:53decision 3 output like this Alpha 4 your
- 6:04:57decision 3 output fourth decision tree
- 6:05:00like this it will be alpha n your
- 6:05:02decision tree n output and this will be
- 6:05:06your output finally when you're trying
- 6:05:09to inference from any new
- 6:05:12record now the reason why we say this as
- 6:05:15boosting because see understand we are
- 6:05:17going to add each and every decision
- 6:05:19tree output slowly to finally get our
- 6:05:22output with respect to the working of
- 6:05:23the decision tree this is how XG boost
- 6:05:26actually work don't credit further needs
- 6:05:28to be simplified yes see like this
- 6:05:31similarly we can split credit with the
- 6:05:33help of like we can make blue green one
- 6:05:35side normal at one side But whichever
- 6:05:37will be giving the information gain more
- 6:05:40that will be taken into consideration
- 6:05:41right and this is how your entire X
- 6:05:43boost classifier works it is very very
- 6:05:46difficult to basically calculate all
- 6:05:48those things so that is the reason we
- 6:05:50say that XG boost is also a blackbox
- 6:05:53model so this is basically a blackb
- 6:05:56model it is it prone to overfitting see
- 6:05:59at one stage we also need to perform
- 6:06:02hyperparameter tuning and this we
- 6:06:05specifically say pre- pruning we tend to
- 6:06:08do pre pruning and since we are
- 6:06:10combining multiple decision trees no no
- 6:06:14this decision tree this decision tree is
- 6:06:17this one this independent decision tree
- 6:06:19which I have created now parall after
- 6:06:21this what I'll do I'll create one more
- 6:06:22decision tree so it'll be looking like
- 6:06:24this see finally how it will look so
- 6:06:26this is my base model then my data then
- 6:06:29my data will go to this decision tree
- 6:06:31which I have actually done as a binary
- 6:06:33split on different different records
- 6:06:36then again we will make another decision
- 6:06:38tree which will again be a binary tree
- 6:06:40the splits will look like this then this
- 6:06:43is my base model where I'm getting the
- 6:06:45value as zero this will be alpha 1
- 6:06:47multiplied by decision tree 1 which is
- 6:06:50this then this is Alpha 2 multiplied by
- 6:06:53decision tree 2 which is this and like
- 6:06:55this we will keep on continuously adding
- 6:06:58more decision trees unless and until
- 6:07:00this entire things becomes a very strong
- 6:07:04learner so this is how how we basically
- 6:07:06do the combination of all these things
- 6:07:08so I hope everybody is able to
- 6:07:10understand about the XG boost classifier
- 6:07:14now you may be thinking how does
- 6:07:15regressor work do you want a regressor
- 6:07:17problem statement also the decision tree
- 6:07:19will get constructed based on
- 6:07:21Independent features and again Lambda
- 6:07:23value is a hyperparameter we basically
- 6:07:26set up Lambda value with the help of
- 6:07:28cross validation now uh let's go ahead
- 6:07:30and discuss about ex boost regressor the
- 6:07:33second algorithm that we we will
- 6:07:35probably discuss about is something
- 6:07:37called as XG boost regressor and how
- 6:07:41does X boost regressor actually work
- 6:07:43some fundamental is follow in random
- 6:07:45Forest no in random Forest it is
- 6:07:47completely different there bagging
- 6:07:49happens bagging happens so over here
- 6:07:52let's go ahead with the regressor so
- 6:07:54here I'm going to take some example
- 6:07:56let's say that I have this many
- 6:07:57experience this many Gap and based on
- 6:08:00that we need to determine the salary my
- 6:08:02salary is my output feature let's say
- 6:08:04the experience is 2 2.5 3 4 4.5 okay now
- 6:08:10in this Gap let's say it is yes
- 6:08:13yes no no yes and let's say that the
- 6:08:17salary is somewhere around 40K it is
- 6:08:2041k
- 6:08:2252k and uh let's see some more data set
- 6:08:25over here 60k and 62k now the first step
- 6:08:29in classifier we created a base model
- 6:08:32here also we'll try to create a base
- 6:08:33model first of all this base model what
- 6:08:36output it will give it will give the
- 6:08:38average of all these values what is the
- 6:08:40average of all these values okay what is
- 6:08:42the average of all these value 40 81 52
- 6:08:4560 62 if I just do the average it is
- 6:08:48nothing but 51k so by default I will
- 6:08:50create a base model which will take any
- 6:08:52input and just give the output as 51
- 6:08:54this is the first step now based on this
- 6:08:56I will try to calculate my residual now
- 6:08:58how do I calculate my residual I will
- 6:09:00just subtract 40 by 51k so this will
- 6:09:03basically be - 11k
- 6:09:06and uh this will be 10 K - 10 K - 10 and
- 6:09:11this will be 1 this will be 9 and this
- 6:09:16will be 11 I hope everybody's able to
- 6:09:18get this let's say that I I make this as
- 6:09:2142k okay for just making my calculation
- 6:09:23little bit easy so I have 9 over here so
- 6:09:26this is my residual then again the first
- 6:09:28step is that I construct my uh decision
- 6:09:32tree now let's say say that I'm going to
- 6:09:35use The Experience over here so this is
- 6:09:37my experience node and based on this
- 6:09:39experience node I have my features over
- 6:09:42here so here I will take up all my
- 6:09:44residuals - 11 99 1 99 11 and then how
- 6:09:50do I do the split based on experience
- 6:09:52this is a continuous feature so I have
- 6:09:56to basically do split with respect to
- 6:09:58continuous feature which I have already
- 6:09:59shown you in decision tree how do we do
- 6:10:01so here is my residual here it is 40
- 6:10:04minus this
- 6:10:05is - 11 K - 9 K uh this is 1 K this is 9
- 6:10:12K and
- 6:10:1411k - 9k so now I will just create take
- 6:10:17up my first node here I'm going to use
- 6:10:20my experience feature I know my values
- 6:10:23what all things are going to come 11k in
- 6:10:25the root node - 9 1 9 and 11 now what we
- 6:10:30are going to do over here is that so I'm
- 6:10:32going to do again a binary split over
- 6:10:34here now the binary split will happen
- 6:10:36based on the continuous feature that is
- 6:10:38experienced so two types of Records I
- 6:10:40may get one is less than or equal to two
- 6:10:42and one is greater than 2 less than or
- 6:10:46equal to two and one is greater than two
- 6:10:48now less than or equal to two when I do
- 6:10:49the split let's see how many values we
- 6:10:51are getting less than or equal to two I
- 6:10:53will get only one value that is -1 and
- 6:10:56here I'm actually going to get all the
- 6:10:58other values - 9 1 9 11 now what we are
- 6:11:02going to do after this is that calculate
- 6:11:04the similarity weight now here the
- 6:11:06similarity weight will little bit the
- 6:11:08formula will change with respect to
- 6:11:10regression so similarity weight is
- 6:11:12nothing but summation of residual
- 6:11:15squares divided by number of residuals
- 6:11:18plus Lambda again here we are going to
- 6:11:20consider Lambda is zero because this is
- 6:11:22a hyper parameter tuning more the value
- 6:11:25of Lambda that basically means more more
- 6:11:27we are penalizing with respect to the
- 6:11:29residuals so this will be the formula
- 6:11:31that we are going to apply okay so let's
- 6:11:33see for the first number that that we
- 6:11:35want to apply so how this will get
- 6:11:37applied again I'm going to write this
- 6:11:39formula here it'll be better let's say
- 6:11:42here similarity weight is equal to
- 6:11:46summation of residual square and here
- 6:11:49you have number of residuals plus Lambda
- 6:11:52see previously we were using probability
- 6:11:54and then all those things we are using
- 6:11:56so if you want to calculate the
- 6:11:58similarity weight of this this will
- 6:11:59become 121 divided by number of residual
- 6:12:03is 1 plus Lambda is 0 so this is going
- 6:12:06to be 121 so here we are going to
- 6:12:08calculate the similarity weight which is
- 6:12:10nothing but 121 if if we probably take
- 6:12:13Alpha let's let's do one thing if we
- 6:12:15probably take uh if if we probably take
- 6:12:19Alpha is equal to 1 then what will
- 6:12:20happen if you take Alpha is equal to 1
- 6:12:22just think over here what will what may
- 6:12:23happen we may directly penalize the
- 6:12:26similarity weight right by just adding
- 6:12:28one okay so let's do that also suppose I
- 6:12:30say I'm going to take Alpha is equal to
- 6:12:321 so what will happen this will not be
- 6:12:35the formula now now what will become 121
- 6:12:38divided number of residual is 1 + 1 this
- 6:12:41is nothing but 65.5 let's say that I now
- 6:12:44have 65.5 as my similarity weight now
- 6:12:47similarly I will go ahead and compute
- 6:12:49the similarity weight for the next one
- 6:12:52so here it will become - 9 + 9 + 9 + 11
- 6:12:58whole Square divided 4 + 1 so this and
- 6:13:01this will get subtracted 12 squ is
- 6:13:04nothing but 14 4 144 divid 5 so if I go
- 6:13:07ahead and calculate 144 ID 5 it is
- 6:13:10nothing but 28.5 so here I get
- 6:13:1528.5 so the similarity weight for this
- 6:13:18is
- 6:13:2028.5 similarly I can go ahead and
- 6:13:22calculate the similarity weight for this
- 6:13:24for the top one so it'll be nothing but
- 6:13:27what it will be 11 + sorry - 11 - 11 - 9
- 6:13:34+ + 1 + 9 + 11 divided 1 2 3 4 5 5 + 1
- 6:13:41is 6 so this is getting subtracted this
- 6:13:44will be 1X 6 anyhow this will be whole
- 6:13:46square right so anyhow it will be 1X 6
- 6:13:48only so 1X 6 will be my similarity
- 6:13:51weight over here okay 28.8 hits okay now
- 6:13:54finally The Information Gain that we
- 6:13:56need to compute will be very much simple
- 6:13:58what will be the Information Gain 65.5 +
- 6:14:0328.8
- 6:14:06minus 1X 6 so try to get it whatever we
- 6:14:09are trying to get it over here just tell
- 6:14:11me what will be the output is it 98.34%
- 6:14:3560.5 60.5 + 28 88 then this will change
- 6:14:40just a second 89.1 3 understand you
- 6:14:44don't have to worry about calculation
- 6:14:46automatically that things will be doing
- 6:14:48it okay so you don't have to worry now
- 6:14:50see we have now further the decision
- 6:14:52tree can be splitted into any number of
- 6:14:54times probably the next split what we
- 6:14:56can do is that we can we can do next
- 6:14:58split something like this this will be
- 6:15:00my experience the two splits that may
- 6:15:03happen with respect to less than or
- 6:15:05equal to 2.5 less than or equal to 2.5
- 6:15:08or greater than 2.5 now if this probably
- 6:15:11gives the Information Gain better then
- 6:15:13the split will happen like this
- 6:15:14otherwise whichever gives the better
- 6:15:16information again the split will
- 6:15:17basically happen like this I hope like
- 6:15:20let's say that this is this is the split
- 6:15:22that is required - 11 - 11 is 9 is over
- 6:15:25here and then we have 1 comma 9A 11 okay
- 6:15:28because less than or equal to 2.5 this
- 6:15:30two records will definitely go over here
- 6:15:32and this two This Record will definitely
- 6:15:34go over here now if I try to calculate
- 6:15:36the similarity weight for this it will
- 6:15:38be nothing but - 11 - 9 - 11 - 9 whole S
- 6:15:43ided 2 + 1 right now in this particular
- 6:15:46case it will be - 20 s / 3 which is
- 6:15:51nothing but 400 2 20 into 20 is 400
- 6:15:55which is nothing but 3 so if I go and
- 6:15:57probably use a
- 6:15:59calculator and show it to you
- 6:16:02400 / 3 which is nothing but
- 6:16:06133.33 so the similarity weight for this
- 6:16:08is
- 6:16:10133.33 similarly I can go ahead and
- 6:16:12compute for this it will be 1 + 9 + 11
- 6:16:15whole s / 3 + 1 right so it will be 10 +
- 6:16:1911 10 + 11 is nothing but 21 whole s/ 4
- 6:16:24so what it is 21 whole square if I open
- 6:16:27my calculator 21 s 21 * 21 which is
- 6:16:33nothing but 441 divid by 4 divid by 4 so
- 6:16:37this will probably 110 110.
- 6:16:412.25 and similarly I can go ahead and
- 6:16:44compute for this so if I want to compute
- 6:16:46for this what it will be the same thing
- 6:16:49that we have got over here that is 1x 6
- 6:16:51so this will basically be 1X 6 so
- 6:16:53finally if I compute the information
- 6:16:55again it will be what it will be 133
- 6:17:011333 +
- 6:17:031.25 - 1X 6 obviously this value will be
- 6:17:06greater than the previous one what we
- 6:17:08have got that is
- 6:17:108913 so definitely we are going to use
- 6:17:12this split which is better than the
- 6:17:14previous split right let's say that this
- 6:17:17split has been considered finally how do
- 6:17:20we see the output okay I hope everybody
- 6:17:23is able to understand right let's say
- 6:17:24that this split has worked well so I'm
- 6:17:26going to rub all these things
- 6:17:2911.25 is there now suppose I want to do
- 6:17:33the inferencing how the inferencing will
- 6:17:35be done
- 6:17:3711.25 here 110.2 now suppose any record
- 6:17:41comes from here first of all any record
- 6:17:43that will go it will go to the base
- 6:17:45model so the base model whenever it goes
- 6:17:47the value is 51 51 plus alpha 1 this is
- 6:17:51my learning rate one suppose if it goes
- 6:17:54in this route then what we have we have
- 6:17:56- 11 - 9 whenever we go in this rote
- 6:17:59which has - 11 and - 9 the average of
- 6:18:02both these numbers will be considered
- 6:18:03what is average of both these numbers -
- 6:18:0511 - 1 9/ 2 this is nothing but - 10
- 6:18:10right so - 10 will get multiplied here
- 6:18:13suppose if it goes in this route then
- 6:18:15here what will happen here will 1 + 9 +
- 6:18:1811 divide by 3 average will be taken so
- 6:18:2021 divid 3 7 will be there so this will
- 6:18:23get replaced by 7 so similarly anything
- 6:18:27that you are doing this is with respect
- 6:18:28to decision tree 1 like this we will
- 6:18:30again construct decision tree separately
- 6:18:33and again it will become Alpha 2 by
- 6:18:35decision Tre 2 Alpha 3 by decision 3 3
- 6:18:39and like this you will be doing till
- 6:18:42Alpha and decision 3 n and once you
- 6:18:45calculate this this will be your
- 6:18:47specific output in a regression tree so
- 6:18:49in this particular case what will happen
- 6:18:51you're just trying to play with
- 6:18:53parameters and you're trying to use in a
- 6:18:55different way to compute all this things
- 6:18:57everybody clear but again it is a
- 6:18:59blackbox model you cannot visualize all
- 6:19:02this things now let's go to the third
- 6:19:03algorithm which is called as s VM see
- 6:19:05svm is almost like decision uh logistic
- 6:19:08regression okay so the major aim of svm
- 6:19:12is
- 6:19:13that major aim of svm is that suppose if
- 6:19:16I have a do data points like this okay
- 6:19:20we obviously use uh logistic regression
- 6:19:23to split this data points right like
- 6:19:25this we try to create a best fit line
- 6:19:28which looks like this and probably based
- 6:19:30on this best fit line we try to divide
- 6:19:32the point now in svm what we do is that
- 6:19:36we not only create a best fit line but
- 6:19:40instead we also create a point which is
- 6:19:44called as marginal
- 6:19:45planes so like this we create some
- 6:19:48marginal
- 6:19:49plane so this is your hyper plane and
- 6:19:53this is your marginal plane and
- 6:19:55whichever plane has this maximum
- 6:19:58distance will be able to divide the
- 6:20:01points more efficiently but usually in
- 6:20:05in a normal scenario you know whenever
- 6:20:07we talk about hyper plane or whenever we
- 6:20:10talk about marginal plane there will be
- 6:20:11lot of overlapping of points right
- 6:20:13suppose if I have some specific points I
- 6:20:16have one point which looks like this I
- 6:20:18may also have another points which may
- 6:20:20overlap so it is very difficult to get
- 6:20:23an exact straight marginal planes and
- 6:20:26split the point based on this now this
- 6:20:28specific marginal plane should be
- 6:20:30maximum because we can create any type
- 6:20:32best fit line and probably
- 6:20:35uh use this marginal plane now if we
- 6:20:38have this overlapping right if for what
- 6:20:40do we call for this kind of plane this
- 6:20:42kind of plane is basically called as
- 6:20:44hard marginal plane so this is basically
- 6:20:47called as hardge marginal plane okay and
- 6:20:51similarly if any points are overlapping
- 6:20:54suppose this yellow points can also get
- 6:20:56overlapped over here and there may be
- 6:20:58some kind of Errors so for this
- 6:21:00particular case we basically say as soft
- 6:21:02marginal plane because here we will be
- 6:21:05able to see that errors will be there
- 6:21:07now in asvm what we focus on doing is
- 6:21:10that we focus on creating this marginal
- 6:21:13plane with maximum distance even though
- 6:21:15there are some errors we consider it in
- 6:21:17solving it by providing some kind of
- 6:21:19hyper parameter now how do we go ahead
- 6:21:22and basically create this all marginal
- 6:21:24planes and how do we go ahead with this
- 6:21:26it's very much simple uh just imagine in
- 6:21:29this specific way that initially let's
- 6:21:32consider that I have this data point
- 6:21:33suppose this is my
- 6:21:35best fit line how do we give this best
- 6:21:38fit line as equation we basically say
- 6:21:40yal mx + C right we we basically say
- 6:21:43this equation as y mx + C no hard hard
- 6:21:47marginal it is impossible in a normal
- 6:21:50data set obviously you'll not be able to
- 6:21:52get it but definitely we go ahead with
- 6:21:55creating a soft marginal plan now Y is
- 6:21:56equal to MX plus C what does this m
- 6:21:59indicate m is nothing but slope and C
- 6:22:02indicates nothing but intercept
- 6:22:05can I say that this both equations are
- 6:22:07same ax + b y + C isal 0 can I also say
- 6:22:12that this is the equation of a straight
- 6:22:14line can I say that this is also the
- 6:22:16equation of straight line I will say
- 6:22:18that both of them are equal can I say
- 6:22:20both of them are equal see if I try to
- 6:22:22prove this to you if I take this
- 6:22:24equation and try to find out y it will
- 6:22:26be nothing but minus C Min - c
- 6:22:30minus a sorry - a x and this will be
- 6:22:34divided by B this will be divided by
- 6:22:37B this will be divided by B so here you
- 6:22:40can see that it is almost the same in
- 6:22:42this particular case my M value will be
- 6:22:44- A by B and my C will basically be
- 6:22:47minus C by B so both the equation are
- 6:22:49almost same
- 6:22:51so let's consider that this is my
- 6:22:53equation and I am actually and whenever
- 6:22:57I say Y is equal to mx + C can I also
- 6:23:00write something like this Y is equal to
- 6:23:03W1
- 6:23:05X1 + W2 X2 plus like this plus C or plus
- 6:23:10b same thing no so here also we can
- 6:23:13write y w transpose x + B same equation
- 6:23:17right we are basically using same
- 6:23:19equation yes we can also write it in a
- 6:23:21different way but at the end of the day
- 6:23:23we are also treating something like this
- 6:23:25let's say that this slope is in this
- 6:23:28direction if this slope is in this
- 6:23:30direction then I can basically say that
- 6:23:32let's consider that the slope is minus
- 6:23:33one
- 6:23:35let's say that this slope is minus one
- 6:23:36see it is in the negative Direction
- 6:23:38let's say that this slope is minus one
- 6:23:40I'm just trying to prove that this slope
- 6:23:42is negative value let's consider this
- 6:23:44now suppose this is one of my point - 4a
- 6:23:480 and obviously this particular equation
- 6:23:50is given by this particular line is
- 6:23:52given by this equation now if I really
- 6:23:55want to find out the Y value let's say
- 6:23:57that this is my
- 6:23:59X1 this is my X1 and this is my X2 let's
- 6:24:03say that
- 6:24:05I want to find out I want to find out
- 6:24:08this W transpose x + b the Y value based
- 6:24:12on this line if I want to compute the y-
- 6:24:14value based on this line how will I
- 6:24:16compute W transpose X basically means
- 6:24:18what w value what all things will be
- 6:24:20there one value is B right B is
- 6:24:23intercept right now intercept is passing
- 6:24:25from origin can I say my B will be zero
- 6:24:28obviously I can assume that b will be
- 6:24:30zero now in this particular case if I
- 6:24:32talk about w w in this case is minus one
- 6:24:35which I have initialized over here so if
- 6:24:37I want to do this matrix multiplication
- 6:24:39it will be W transpose can be written as
- 6:24:41like this and this x value can be
- 6:24:44written as -4 comma - 4 and 0 -4 and 0
- 6:24:49right so I can basically write like this
- 6:24:52now if I do this multiplication what
- 6:24:54will my value I get I will basically get
- 6:24:57four right so this is a positive
- 6:25:01value this is a positive value Now
- 6:25:04understand since this is a positive
- 6:25:05value any points that are below this
- 6:25:08line any points that I consider below
- 6:25:11this line and if I try to calculate the
- 6:25:13Y can I say that it will always be
- 6:25:15positive yes or no similarly if I could
- 6:25:18probably consider one point over here as
- 6:25:214A 4A 4 now tell me in this 4A 4 if I
- 6:25:25calculate the Y value what will you get
- 6:25:27whether you'll get a positive value or a
- 6:25:29negative value if I try to calculate the
- 6:25:30Y value in this case because here only
- 6:25:32positive values will'll be getting right
- 6:25:34so if I calculate the Y value will the Y
- 6:25:37value be negative or positive just try
- 6:25:39to calculate how do you calculate again
- 6:25:41I will use y equation this time again my
- 6:25:44slope is minus1 my intercept is zero and
- 6:25:46here I will have 4 comma
- 6:25:494 now here Min
- 6:25:51-4 and then this is + 0 this will be Min
- 6:25:54-4 right so this will be a negative
- 6:25:57value negative value guys negative see -
- 6:26:004 + 0 negative so any point that I will
- 6:26:05probably have in top of this any
- 6:26:08points Above This Plane right and if I
- 6:26:12try to calculate the Y value it will
- 6:26:13always be negative so what two things
- 6:26:16you are able to get positive and
- 6:26:17negative so you can consider this
- 6:26:19entirely one category this another
- 6:26:22category at least these two things you
- 6:26:24can basically
- 6:26:25consider guys I hope everybody is able
- 6:26:27to understand this so this will be my
- 6:26:29one
- 6:26:30category and this will be my another
- 6:26:32category obviously so that basically
- 6:26:34means I can definitely use a plane and
- 6:26:35split this point I hope everybody is
- 6:26:37able to understand now let's go ahead
- 6:26:39and let's see how this marginal plane
- 6:26:41will get created and what is the cost
- 6:26:44function to basically do this or what is
- 6:26:46the cost function in making sure that
- 6:26:48the marginal plane will definitely work
- 6:26:50right it becomes difficult right so
- 6:26:52suppose let's consider an
- 6:26:55example suppose I say that this is my
- 6:26:58lines let's say uh I want to basically
- 6:27:01create a kind of I have two variety of
- 6:27:03points one is this point let's say I
- 6:27:06have all this points like this and the
- 6:27:07other points I have somewhere here let's
- 6:27:10consider I am just using directly good
- 6:27:13number of points so that I can split it
- 6:27:15okay because I will try to talk about it
- 6:27:17what I'm actually trying to prove so
- 6:27:20obviously this is my best fit line that
- 6:27:21splits and apart from that what I will
- 6:27:24do is that I'll also create a marginal
- 6:27:26points so in order to create the
- 6:27:27marginal point I may use some different
- 6:27:30color let's see which color this will be
- 6:27:32my one marginal point remember it will
- 6:27:35be to the nearest point over here and
- 6:27:38basically we will construct like like
- 6:27:40this and similarly here we will be
- 6:27:43constructing like this I've already told
- 6:27:45you guys this equation can be mentioned
- 6:27:48at w transpose x + B = 0 right I can
- 6:27:51definitely say this because ax + b y + C
- 6:27:55is equal to 0 so this I can also write
- 6:27:57it as W transpose x equal to 0 sorry
- 6:28:00plus b plus b equal to 0 so both are
- 6:28:03same okay this I don't have to prove it
- 6:28:05I hope everybody's clear with this now
- 6:28:08what I'm going to do let's represent
- 6:28:10this line also with some equation so
- 6:28:12this line if I want to represent this
- 6:28:14will be W transpose x + B what value
- 6:28:17will come over here positive or negative
- 6:28:19C from this line anything above this
- 6:28:21plane right any any any distance that we
- 6:28:24try to find out it will always be
- 6:28:25negative so let's say that I'm using it
- 6:28:27as minus one to just read as it is a
- 6:28:30negative value and this line that I am
- 6:28:32going to mention it it will be W
- 6:28:34transpose x + B is equal to + 1 Min -1
- 6:28:37above + 1 because we have already
- 6:28:39discussed from this point if you're
- 6:28:41trying to calculate the Y value it is
- 6:28:43always going to be + one this is going
- 6:28:45to be minus one here I should definitely
- 6:28:48say this as K okay but I'm not
- 6:28:50mentioning K in many articles you'll see
- 6:28:53it as minus one uh many research paper
- 6:28:55also they use it as minus one but I
- 6:28:57would like to specify uh minus and plus
- 6:28:59K but here let's go and write minus1 and
- 6:29:02plus now my aim is to increase this
- 6:29:05distance okay this distance I really
- 6:29:07want to increase this distance now in
- 6:29:09order to increase this if I increase
- 6:29:11this distance that basically means my
- 6:29:13model is performing well so let's say I
- 6:29:16want to find this distance first of all
- 6:29:18so if I write w transpose X Plus Bal to
- 6:29:201 and here I will write w transpose x +
- 6:29:23B isal minus1 so what I'm going to do
- 6:29:25I'm going to do the computation and
- 6:29:28subtract it like this so here obviously
- 6:29:31this will be my X1 this will be my X2
- 6:29:34okay because these are my another points
- 6:29:35X2 and X1 so I can write w transpose X1
- 6:29:40-
- 6:29:42X2 B and B will get cancell and here I
- 6:29:45will be writing two right so from here
- 6:29:49we can definitely write two different
- 6:29:50things let's see what all things we can
- 6:29:52write so here this is nothing but the
- 6:29:54difference between my this plane and
- 6:29:56this plane which is given by like this
- 6:29:58okay now always understand whenever we
- 6:30:01consider any any vector vors right any
- 6:30:06vectors right it also has something
- 6:30:07called as
- 6:30:09magnitude so if I want to remove this
- 6:30:12magnitude I can divide this by W this
- 6:30:16magnitude of w then only my Vector will
- 6:30:18remain which is indicated like this so
- 6:30:20I'm going to basically divide by this
- 6:30:22particular operation both both the side
- 6:30:24I'm dividing by this magnitude of w and
- 6:30:27I don't care about the directions over
- 6:30:29here right now we just care about the
- 6:30:30vectors now when I write like this what
- 6:30:33is our aim our aim is to can I say our
- 6:30:36aim is to our aim is to
- 6:30:40maximize 2 byw can I say this guys yes
- 6:30:43or
- 6:30:46no what is our aim our aim is to
- 6:30:49basically maximize this right by
- 6:30:52updating W comma B value I need to
- 6:30:56maximize this yes everybody's clear with
- 6:30:59this can I say that yes I want to
- 6:31:01maximize this yes or no everybody I want
- 6:31:05to maximize this if I maximize this that
- 6:31:07basically means my marginal plane will
- 6:31:08become bigger my marginal plane will be
- 6:31:10bigger okay now can I write along with
- 6:31:13this that such that y of I my output
- 6:31:17will be dependent on two different
- 6:31:18things one is I can say that my y y of I
- 6:31:22is plus of uh is + one when w transpose
- 6:31:26x + B is greater than or equal to 1
- 6:31:29everybody see in this equation what I'm
- 6:31:31actually trying to specify such that y
- 6:31:33of I is + 1 when w transpose x + B is
- 6:31:36greater than 1 and when it is minus 1
- 6:31:38that basically means w transpose of X is
- 6:31:40B is less than or equal to minus now
- 6:31:42what does this basically mean see all my
- 6:31:46values whenever I compute W transpose x
- 6:31:49+ B is greater than or equal to 1 I'm
- 6:31:51obviously going to get this + one when w
- 6:31:54transpose X+ B is less than or equal to
- 6:31:561 I'm always going to get the output as
- 6:31:58minus one I hope that is the reason why
- 6:32:00I have actually written like this so
- 6:32:02this two we have already discussed why
- 6:32:03we are specifically writing we want to
- 6:32:05increase the marginal plane which is
- 6:32:07this this is my marginal plane and I'm
- 6:32:09writing one condition that my Yi value
- 6:32:11will be+ one when w transpose X plus b
- 6:32:14is greater than or equal to 1 otherwise
- 6:32:16it when it is less than or equal to
- 6:32:17minus one it is going to be very much
- 6:32:18clear with this transpose condition we
- 6:32:20have already done it everybody clear
- 6:32:22with this now on top of it we can add
- 6:32:25one more very important Point instead of
- 6:32:28writing such that and all you can also
- 6:32:30say that our major
- 6:32:32aim our major aim is that if I multiply
- 6:32:36y i multiplied by W transpose X of I + B
- 6:32:41If I multiply this two this will always
- 6:32:44be able greater than or equal to 1 for
- 6:32:48correct points right for correct points
- 6:32:52because understand if it is minus one if
- 6:32:55I'm multiplying with this and if it is a
- 6:32:57correct Point minus into minus will
- 6:32:59obviously be greater than or equal to
- 6:33:01one only right similarly for this it
- 6:33:03will be greater than 1 so I can also
- 6:33:05definitely say that my major M If I
- 6:33:07multiply y of I with this it will be
- 6:33:10always greater than or equal to + 1 U
- 6:33:12which is definitely saying that it will
- 6:33:14be a positive value so this is just a
- 6:33:16representation guys but understand what
- 6:33:19is the minimized cost function this is
- 6:33:21my minimized cost function maximized
- 6:33:23cost function now I'm going to again
- 6:33:26write it down
- 6:33:28maximize W comma B maximize W comma b 2
- 6:33:33by magnitude of w I can also write
- 6:33:37something like this minimize W comma B
- 6:33:40and I can just inverse this which looks
- 6:33:43like this are these both are same or not
- 6:33:45because always understand in machine
- 6:33:48learning algorithm why do we write
- 6:33:51minimize things because we are trying to
- 6:33:54minimize something okay both are
- 6:33:57equivalent these both are equivalent and
- 6:33:59why we specifically write minimization
- 6:34:01because in the back propagation when we
- 6:34:03we are continuously updating the weights
- 6:34:05of w and B so we can definitely write
- 6:34:08like this so here my main target is to
- 6:34:12minimize this particular value by
- 6:34:14changing W and B and I will start adding
- 6:34:17some more parameters over here this is
- 6:34:19fine till here I think everybody has got
- 6:34:22it this is our aim and we are going to
- 6:34:23do this but I'm going to add two more
- 6:34:26parameters in this Optimizer one is C of
- 6:34:29I and one is summation of I equal 1 to n
- 6:34:33and here I will use something called as
- 6:34:35EA EA of I first of all I'll tell what
- 6:34:38is C of I see if I have this specific
- 6:34:41data point let's say if some of my
- 6:34:44points are over here then is it a right
- 6:34:47right prediction or wrong prediction if
- 6:34:49some of my points are over here is it a
- 6:34:51right prediction or wrong prediction
- 6:34:54obviously it is a wrong prediction if my
- 6:34:56points are somewhere here is it a WR
- 6:34:58prediction wrong wrong incorrect
- 6:34:59prediction right so this C value
- 6:35:02basically says that how many errors we
- 6:35:04can have how many errors we can have if
- 6:35:06it says that fine we can have six errors
- 6:35:08or seven errors how many errors we can
- 6:35:11have even though we are using the
- 6:35:13marginal plane how many errors we can
- 6:35:16have so here I'm specifically writing
- 6:35:18how many errors we can have this is what
- 6:35:21is specified by C ofi EA of I basically
- 6:35:24says that what is the summation of I'm
- 6:35:26going to write it down since we are
- 6:35:28doing the sumission this entire term
- 6:35:31basically mentions that sumission
- 6:35:34of the distance of the values distance
- 6:35:37of the wrong points and how do we
- 6:35:39calculate the distance from here to here
- 6:35:42suppose this is a wrong point I will try
- 6:35:44to calculate the distance from here to
- 6:35:45here I will do the sumission of this
- 6:35:47I'll do the sumission of this I will do
- 6:35:49the sumission of this similarly for the
- 6:35:51Green Point another sumission will
- 6:35:53happen from here to here like this here
- 6:35:56to here and we going to do that specific
- 6:35:57sumission so we are telling that fine if
- 6:36:01you are not able to fit properly try to
- 6:36:05apply this two hyperparameters and try
- 6:36:07to make sure that this many errors are
- 6:36:10also there it is well and good no
- 6:36:11problem we will go ahead with that try
- 6:36:14to do the submission of the data points
- 6:36:15and based on that try to construct the
- 6:36:18best fit line along with the marginal
- 6:36:20plane like this even though there are
- 6:36:23some errors over here or errors over
- 6:36:25here we are good to go with respect one
- 6:36:27more thing is there which is called as
- 6:36:28Al svr svr only one thing is getting
- 6:36:32changed in svr only this value will get
- 6:36:36changed so I want you all to explore and
- 6:36:38just let me know this will be one
- 6:36:40assignment for you only this value will
- 6:36:42be changing remaining everything are
- 6:36:43same so just try to if you change this
- 6:36:46particular value that becomes an svr
- 6:36:49just try to explore and just try to find
- 6:36:51out and just try to let me know so
- 6:36:52overall uh did you like the entire
- 6:36:55session everyone okay in this one more
- 6:36:57thing is there which is called as kernel
- 6:36:59Matrix svm kernel we say it as svm
- 6:37:02kernel now in s VM kernel what happens
- 6:37:04suppose if I have a specific data points
- 6:37:06which looks like this which looks like
- 6:37:08this so we obviously cannot use a
- 6:37:10straight line and try to divide it so
- 6:37:11what we do we convert this two Dimension
- 6:37:14into three dimensions and then probably
- 6:37:17we push our Point like this one point
- 6:37:19will go like this and the white point
- 6:37:21will go down and then we can basically
- 6:37:24use a plane to split it so I uploaded a
- 6:37:26video around uh around that and uh you
- 6:37:29can definitely have a look onto that and
- 6:37:31I have also shown you practically how to
- 6:37:33do it that is the reason I've created
- 6:37:35that specific video so great uh this was
- 6:37:37it from my side I hope you like this
- 6:37:39session so thank you everyone have a
- 6:37:41great day keep on rocking keep on
- 6:37:43learning and never give up
About this transcript
This page contains the full transcript of Complete Machine Learning In 6 Hours| Krish Naik by Krish Naik, generated from the public captions YouTube serves with the video. The transcript has 69,818 words across 9,542 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.