YouTube2Text

Complete Machine Learning In 6 Hours| Krish Naik — Transcript

by Krish Naik · 69,818 words · 9,542 segments · language en · Watch on YouTube

Full transcript

  1. 0:06so today's session what all things we
  2. 0:08are basically going to discuss so first
  3. 0:10of all we going to discuss about
  4. 0:12different types of machine learning
  5. 0:13algorithm like how many different types
  6. 0:15of machine learning
  7. 0:16algor understand the purpose of taking
  8. 0:20this session is to clear the interviews
  9. 0:23okay clear the interviews once you go
  10. 0:25for a data science interviews and all
  11. 0:28the main purpose is to clear the
  12. 0:29interviews I've seen people who knew
  13. 0:32machine learning algorithms in a proper
  14. 0:34way okay they were definitely able to
  15. 0:36clear it because they just explain the
  16. 0:38algorithms in a better way to the
  17. 0:40recruiter so that they got hired first
  18. 0:42of all is the introduction to machine
  19. 0:45learning here I'm just specifically
  20. 0:47going to talk about AI versus ml versus
  21. 0:51DL versus data sign then the second
  22. 0:53thing that we are going to talk about
  23. 0:55over here is the difference between
  24. 0:58supervised MS
  25. 1:00and unsupervised ml the third thing that
  26. 1:03we are probably going to discuss about
  27. 1:05is something called as linear regression
  28. 1:08so we are going to clearly understand
  29. 1:10the maths and geometric intuition the
  30. 1:13next thing that we are probably going to
  31. 1:15discuss about is R square and adjusted R
  32. 1:18square the fifth topic that we are going
  33. 1:20to discuss about is Ridge and lasso
  34. 1:23regression the first topic that we are
  35. 1:25going to discuss about is AI versus ml
  36. 1:30versus DL versus data science so this is
  37. 1:34the first topic that we are probably
  38. 1:35going to discuss if you really want to
  39. 1:38understand the difference between AI
  40. 1:39versus ml versus DL versus data science
  41. 1:41we will go in this specific format so
  42. 1:43just imagine the entire universe so this
  43. 1:46entire universe I will probably call it
  44. 1:48as an AI now specifically when I say AI
  45. 1:51this basically means AI artificial
  46. 1:53intelligence whatever role you are in
  47. 1:55you are as a machine learning developer
  48. 1:57you working as a deep learning developer
  49. 1:59Vision developer or a data scientist or
  50. 2:02an AI engineer at the end of the day you
  51. 2:05are actually creating AI application so
  52. 2:09if I really want to Define what is this
  53. 2:11artificial intelligence you can just say
  54. 2:13that it is a process wherein we create
  55. 2:16some kind of applications in which it
  56. 2:19will be able to do its task without any
  57. 2:22human intervention so that basically
  58. 2:24means a person need not monitor this AI
  59. 2:27application automatically it'll be able
  60. 2:29to make decisions it will be able to
  61. 2:31perform its task and it will be able to
  62. 2:34do many things so this is what an AI
  63. 2:36application is some of the examples that
  64. 2:38I would definitely like to consider so
  65. 2:41the first example that I would like to
  66. 2:43consider AI application AI module
  67. 2:46Netflix has an AI module suppose if you
  68. 2:49see a kind of action movie for some time
  69. 2:53then the kind of AI work or AI work that
  70. 2:56is basically implemented over here is
  71. 2:57something called as recommendation
  72. 3:00so here through this application what
  73. 3:04happens is that when you're continuously
  74. 3:06seeing the action movies then
  75. 3:08automatically the AI module that is
  76. 3:10present inside Netflix will make sure
  77. 3:13that it gives us recommendation on
  78. 3:15action movies second if I take an
  79. 3:18example of comedy movie If I
  80. 3:20continuously see comedy movie then also
  81. 3:22it'll give us the recommendation of the
  82. 3:24comedy movie so this through this what
  83. 3:26happens is that it understands your
  84. 3:28behavior and it is being able to do its
  85. 3:30task without asking you anything the
  86. 3:33second example that I would like to take
  87. 3:35up in is
  88. 3:36amazon.in now amazon.in again if you buy
  89. 3:39an
  90. 3:40iPhone then it may recommend you a
  91. 3:43headphones so this kind of
  92. 3:45recommendation is also a part of AI
  93. 3:48module that is integrated with the
  94. 3:49amazon.in website the ads that you see
  95. 3:52probably when you opening my channel
  96. 3:55through which I get paid a little bit
  97. 3:56from my from a from the hard work that I
  98. 3:59do in YouTube right so through that ads
  99. 4:02how that is recommended to you uh that
  100. 4:05is also an AI engine that is included in
  101. 4:07the YouTube channel itself which really
  102. 4:09plays it is a business-driven goal
  103. 4:12understand it is a business driven
  104. 4:13things that we basically do with the
  105. 4:15help of AI one more example that I would
  106. 4:17like to give you is if I consider it
  107. 4:20self-driving cars so here you'll be able
  108. 4:23to see self-driving cars if you take an
  109. 4:25example of Tesla so self-driving cars
  110. 4:27what happens based on the road it is
  111. 4:29able ble to drive it automatically who
  112. 4:31is doing that there is an AI application
  113. 4:33integrated with the car itself right so
  114. 4:36if I consider all these things these all
  115. 4:38are AI application at the end of the day
  116. 4:42whatever role you do you are going to
  117. 4:44create an AI application this is the
  118. 4:46common mistake what people do you know
  119. 4:48like our CEO sudhansu Kumar he has
  120. 4:50written in his profile that he's an AI
  121. 4:52engineer that basically means his goal
  122. 4:55is to create an AI application so
  123. 4:57probably in a product based companies
  124. 4:58you'll be seeing this kind of roles
  125. 4:59called as AI engineer now let's go to
  126. 5:01the next role which is called as machine
  127. 5:03learning so where does machine learning
  128. 5:04comes into existence so if I try to
  129. 5:07create this machine learning is a subset
  130. 5:10of AI and what is the role of machine
  131. 5:12learning it provides stats
  132. 5:15tools
  133. 5:17to analyze the data visualize the data
  134. 5:22and apart from that to do
  135. 5:24predictions I'm
  136. 5:27forecasting so you will be seeing a lot
  137. 5:29of machine learning algorithms so
  138. 5:31internally those machine learning
  139. 5:33algorithm the equation that we are
  140. 5:34basically using it is basically using it
  141. 5:38is having a kind of stats tool stat
  142. 5:40techniques because whenever we work with
  143. 5:42data statistics is definitely very much
  144. 5:44important so this exactly is called as
  145. 5:47machine learning so it is a subset of AI
  146. 5:50this is very much important to
  147. 5:52understand ml is a subset of AI so here
  148. 5:55you can see that it is a part of this
  149. 5:57now let's go to the next one which is
  150. 5:59called called as deep learning deep
  151. 6:01learning is again a subset of ml now
  152. 6:04let's consider why deep learning came
  153. 6:05into existence because in 1950s 60s
  154. 6:09scientists thought that can we make
  155. 6:11machine learn like how we human being
  156. 6:13learn so for that particular purpose
  157. 6:16deep learning came into existence here
  158. 6:18the plan is to basically mimic human
  159. 6:21brain so when I say mimicking human
  160. 6:24brain that basically means we are trying
  161. 6:26to mimic the human brain to implement
  162. 6:28something to learn something so for this
  163. 6:31you use something called as
  164. 6:32multi-layered neural networks so this is
  165. 6:35what deep learning is it is a subset of
  166. 6:37machine learning its main aim is to
  167. 6:40mimic human brain so they actually
  168. 6:42create multi-layer neural network and
  169. 6:45this multi-layered neural network will
  170. 6:47basically help you to train the machines
  171. 6:49or applications whatever we are trying
  172. 6:51to create and deep learning has really
  173. 6:54really done an amazing work with the
  174. 6:56help of deep learning we are able to
  175. 6:58solve such a complex complex complex use
  176. 7:02cases that we will be probably
  177. 7:04discussing as we go ahead now if I come
  178. 7:06to data science see this is the thing
  179. 7:08guys if you want to say yourself as a
  180. 7:10data scientist tomorrow you given a
  181. 7:13business use case and situation comes
  182. 7:15that you probably have to solve that use
  183. 7:17case with the help of machine learning
  184. 7:18algorithms or deep learning algorithms
  185. 7:20again the final goal is to create an AI
  186. 7:22application right you cannot say that I
  187. 7:24am a data scientist and I'll just work
  188. 7:26in machine learning I or I'll work in
  189. 7:29deep learning or I may I don't know how
  190. 7:31to analyze the data no you cannot do
  191. 7:33that when I was working in Panasonic I
  192. 7:36got various different kind of task
  193. 7:39sometime I was told to use W powerbi to
  194. 7:41visualize analyze the data sometime I
  195. 7:43was given a machine learning project
  196. 7:45sometime I was given a deep learning
  197. 7:46project so as a data scientist if I
  198. 7:49consider where does data scientist fall
  199. 7:51into this it will be a part of
  200. 7:53everything so if I talk about machine
  201. 7:56learning and deep learning with respect
  202. 7:58to any kind of problem statement that we
  203. 8:00solve the majority of the business use
  204. 8:03cases will be falling in two sections
  205. 8:05one is supervised machine learning one
  206. 8:08is unsupervised machine learning so most
  207. 8:10of the problems that you are basically
  208. 8:12solving this is with respect to this two
  209. 8:15problem statement two different types of
  210. 8:16machine learning algorithms that is
  211. 8:18supervised machine learning and deep
  212. 8:20learning if I talk about supervised
  213. 8:22machine learning two major problem
  214. 8:24statements that you are basically
  215. 8:25solving here also one is regression
  216. 8:28problem
  217. 8:30and the other one is something called as
  218. 8:31classification problem and in the case
  219. 8:34of unsupervised machine learning problem
  220. 8:36statement you are basically solving two
  221. 8:37different types of problem one is
  222. 8:39clustering and one is dimensionality
  223. 8:42reduction and there is also one more
  224. 8:44type which is called as reinforcement
  225. 8:46learning reinforcement learning I can I
  226. 8:50I will definitely talk about this not
  227. 8:52right now right now we are just focusing
  228. 8:53on all these things now understand what
  229. 8:56happens in supervised machine learning
  230. 8:58let's consider consider a data set so
  231. 9:00here I have a data set which says this
  232. 9:03is my age and this is my weight suppose
  233. 9:07I have these two specific features let's
  234. 9:09say that I have values like 24 62 25 63
  235. 9:1521 72
  236. 9:19257 uh 62 and many more data over here
  237. 9:23let's say that my task is to basically
  238. 9:25take this particular data and create a
  239. 9:27model wherein so suppose my task is that
  240. 9:31I need to create a model whenever it
  241. 9:34takes the New Age first of all we train
  242. 9:36this model with this data and whenever
  243. 9:39we take age a new age it should be able
  244. 9:41to give us the output of weight this
  245. 9:44particular model is also called as
  246. 9:46hypothesis okay I'll discuss about this
  247. 9:49today when I we discussing about linear
  248. 9:51regression now what are the important
  249. 9:53components whenever we have this kind of
  250. 9:55problem statement first of all you need
  251. 9:56to understand there are two important
  252. 9:59things one is independent features and
  253. 10:02the other one is something called as
  254. 10:03dependent features now let's go ahead
  255. 10:05and discuss what is independent feature
  256. 10:07independent feature basically means in
  257. 10:09this particular case since the input
  258. 10:11that I'm basically training in all those
  259. 10:13features becomes an independent feature
  260. 10:15now in this particular case my age is
  261. 10:17independent feature and whatever I'm
  262. 10:20actually predicting so when I say
  263. 10:21predicting I know this is my output okay
  264. 10:24this is the what I have to basically
  265. 10:27make my model uh give this as a an
  266. 10:29output so in this particular casee my
  267. 10:31dependent feature becomes weight why we
  268. 10:34specifically say a dependent feature
  269. 10:36because this is completely dependent on
  270. 10:38this value whenever this is increasing
  271. 10:40or decreasing this value is basically
  272. 10:41getting changed so that is the reason
  273. 10:44why we basically say this has
  274. 10:45independent and dependent feature
  275. 10:47whenever we are solving a problem right
  276. 10:50in the case of supervised machine
  277. 10:51learning remember they will be one
  278. 10:53dependent feature and there can be any
  279. 10:55number of independent features now let's
  280. 10:58go ahead and let's discuss about
  281. 10:59regression and classification what is
  282. 11:01the difference between them now let
  283. 11:03let's go ahead and let's discuss about
  284. 11:05two things one
  285. 11:08is let's say I want a regression problem
  286. 11:11statement suppose I take the same
  287. 11:14example as age and weight so I have
  288. 11:17values like as discussed 24 72 23
  289. 11:2271 uh 24 or 25
  290. 11:2671.5 okay so this kind of data I have
  291. 11:29see this is my output variable which is
  292. 11:32my dependent feature now in this
  293. 11:34particular dependent feature now
  294. 11:36whenever I'm trying to find out the
  295. 11:37output and in this particular output you
  296. 11:39have a continuous variable when you have
  297. 11:42a continuous variable then this becomes
  298. 11:44a regression problem statement now one
  299. 11:47example I would like to give suppose
  300. 11:49this is my data set right this is my age
  301. 11:52this is my weight suppose I am
  302. 11:54populating this particular data set with
  303. 11:56the help of scatter plot then in order
  304. 11:58to basically solve this problem what
  305. 12:01we'll do suppose if I take an example of
  306. 12:03linear regression I will try to draw a
  307. 12:05straight line and this particular line
  308. 12:08is my equation which is called as yal mx
  309. 12:11+ C and with the help of this particular
  310. 12:13equation I will try to find out the
  311. 12:15predicted points so this will be my
  312. 12:17predicted point this will be my
  313. 12:18predicted point this this any new points
  314. 12:21that I see over here will basically be
  315. 12:23my predicted point with respect to Y so
  316. 12:26in this way we basically solve a
  317. 12:28regression problem statement so this is
  318. 12:30very much important to understand let's
  319. 12:32go to the always understand in a
  320. 12:34regression problem statement your output
  321. 12:35will be a continuous variable the second
  322. 12:37one is basically a classification
  323. 12:40problem now in classification problem
  324. 12:42suppose I have a data set let's say that
  325. 12:45number of hours study number of study
  326. 12:48hours number of play
  327. 12:51hours so this is my independent feature
  328. 12:54let's say a number of sleeping hours and
  329. 12:57finally I have my output which will will
  330. 12:59be pass or fail so in this I have all
  331. 13:03this as my independent features and this
  332. 13:05is my dependent feature so I will be
  333. 13:08having some values like this and here
  334. 13:11either you'll be pass or fail or pass or
  335. 13:15fail now whenever you have in your
  336. 13:18output fixed number of categories then
  337. 13:21that becomes a classification problem
  338. 13:23suppose it just has two outputs then it
  339. 13:25becomes a binary classification if you
  340. 13:28have more than two different categories
  341. 13:30at that time it becomes a multiclass
  342. 13:32classification so this is the difference
  343. 13:34between regression problem statement and
  344. 13:36the classification problem statement now
  345. 13:39let's go ahead and let's discuss about
  346. 13:40something called as unsupervised machine
  347. 13:42learning now in unsupervised machine
  348. 13:44learning which is my second main topic
  349. 13:47over here I'm just going to write
  350. 13:49unsupervised machine learning now what
  351. 13:52exactly is unsupervised machine learning
  352. 13:54here whenever I talk about there are two
  353. 13:56main problem statement that we solve one
  354. 13:58is clustering
  355. 13:59one is dimensionality reduction let's
  356. 14:02take one example of a specific data set
  357. 14:04over here let's say that my data set is
  358. 14:06something called as salary and age now
  359. 14:10in this scenario we don't have any
  360. 14:12output variable no output variable no
  361. 14:14dependent variable then what kind of
  362. 14:16assumptions that we can take out from
  363. 14:19this particular data set suppose I have
  364. 14:21salary and age as my values so in this
  365. 14:23particular case I would like to do
  366. 14:25something called as clustering now why
  367. 14:28clustering is used just understand let's
  368. 14:31say I am going to do something called as
  369. 14:33customer segmentation now what does this
  370. 14:35customer segmentation do clustering
  371. 14:37basically means that based on this data
  372. 14:39I will try to find out similar groups
  373. 14:41groups of people suppose this is my one
  374. 14:44group this is my another group this is
  375. 14:46my third group let's say that I was able
  376. 14:48to create this many groups this many
  377. 14:50groups are clusters I'll say cluster 1 2
  378. 14:53three each and every cluster will be
  379. 14:56specifying some information this cluster
  380. 14:58May specify that this person uh he was
  381. 15:01very young but he was able to get some
  382. 15:03amazing salary this person it may some
  383. 15:06specify that these people are basically
  384. 15:07having more age and they are getting
  385. 15:10good salary these people are like middle
  386. 15:12class background where with respect to
  387. 15:14the age the salary is not that much
  388. 15:16increasing so here what we are doing we
  389. 15:18are doing clustering we are grouping
  390. 15:20them together main thing is grouping
  391. 15:23this word is very much important now why
  392. 15:25do we use this suppose my company
  393. 15:28launches is a product and I want to just
  394. 15:31Target this particular product to rich
  395. 15:33people let's say product one is for rich
  396. 15:35people product two is for middle class
  397. 15:37people so if I make this kind of
  398. 15:40clusters I will be able to Target my ads
  399. 15:43only to this kind of people let's say
  400. 15:46that this is the rich people this is the
  401. 15:48middle class people I will be able to
  402. 15:50Target this particular ads or this
  403. 15:53particular product or send this
  404. 15:55particular things to those specific
  405. 15:56group of people by that that is
  406. 15:59basically called as ad marketing and
  407. 16:00this uses something called as customer
  408. 16:04segmentation a very important example
  409. 16:07and based on this customer segmentation
  410. 16:08we can later apply any regression or
  411. 16:10classification kind of problem statement
  412. 16:12now coming to the second one after
  413. 16:14clustering which is called as
  414. 16:15dimensionality reduction now in
  415. 16:17dimensionality reduction what we are
  416. 16:19focusing on suppose if we have th000
  417. 16:22features can we reduce this features to
  418. 16:25lower Dimensions let's say that I want
  419. 16:27to convert this
  420. 16:29uh th000 feature to 100 features lower
  421. 16:32Dimension so can we do that yes it is
  422. 16:36possible with the help of dimensionality
  423. 16:38deduction algorithm there are some
  424. 16:40algorithms like PCA so I'll also try to
  425. 16:42cover this as we go ahead understand
  426. 16:44clustering is not a classification
  427. 16:46problem clustering is a grouping
  428. 16:48algorithm there is no output feature no
  429. 16:50dependent variable in clustering sorry
  430. 16:53in unsupervised ml so yes I will also
  431. 16:55try to cover up LDA we'll cover up PCA
  432. 16:58and all as we go ahead so with respect
  433. 17:00to supervised and unsupervised so first
  434. 17:03thing that we are going to cover is
  435. 17:04something called as linear regression
  436. 17:06the second algorithm that we will try to
  437. 17:08cover after linear regression is
  438. 17:10something called as Ridge and lasso
  439. 17:12third that we are going to cover is
  440. 17:14something called as logistic regression
  441. 17:16the fourth that we are basically going
  442. 17:17to cover is something called as decision
  443. 17:19tree decision tree includes both
  444. 17:21classification and regression four fifth
  445. 17:24that we are going to cover is something
  446. 17:25called as adab boost sixth that we are
  447. 17:27going to cover is something called as
  448. 17:28random Forest seventh that we are going
  449. 17:30to cover is something called as gradient
  450. 17:32boosting eighth that we are going to
  451. 17:34cover is something called as XG boost N9
  452. 17:37that we are going to cover is something
  453. 17:38called as n bias then when we go to the
  454. 17:41unsupervised machine learning algorithm
  455. 17:43the first algorithm that we are going to
  456. 17:45do is something called as K means K
  457. 17:47means algorithm then we also have DV
  458. 17:48scan then we are also going to do higher
  459. 17:50C clustering there is also something
  460. 17:52called as K nearest neighbor clustering
  461. 17:55fifth we'll try to see about PCA then
  462. 17:57LDA so different different things we
  463. 18:00will try to cover up yes svm I have
  464. 18:02missed here I'm going to include svm KNN
  465. 18:05will also get covered so I have that in
  466. 18:07my list probably I may miss one or two
  467. 18:08but we are going to cover everything so
  468. 18:10let's start our first algorithm linear
  469. 18:13regression so let's go ahead and discuss
  470. 18:15about linear regression linear
  471. 18:16regression problem statement is very
  472. 18:18simple guys so suppose I have let's say
  473. 18:21I have two features one is my X feature
  474. 18:23and one is my y feature let's say that X
  475. 18:25is nothing but age and Y is nothing but
  476. 18:29weight so based on these two features I
  477. 18:31have some data points that has been
  478. 18:34present over here so in linear
  479. 18:35regression what we try to do is that we
  480. 18:38try to create a model with the help of
  481. 18:40this training data set so this will be
  482. 18:43my training data set what I'm actually
  483. 18:45going to do is that I'm going to
  484. 18:47basically train a model and this model
  485. 18:50is nothing but a kind of hypothesis
  486. 18:52testing or it is just kind of hypothesis
  487. 18:54which takes the new age and gives the
  488. 18:57output of the weights and then with the
  489. 19:01help of performance metrics we try to
  490. 19:03verify whether this model is performing
  491. 19:05well or not now in short what we are
  492. 19:06going to do in linear regression is that
  493. 19:08we'll try to find out a best fit line
  494. 19:10which will actually help us to do the
  495. 19:12prediction that basically means if I get
  496. 19:14my new age over here then what should be
  497. 19:16my output with respect to Y okay so with
  498. 19:19respect to this what should be my output
  499. 19:21over here in this particular case
  500. 19:23whenever we are drawing a diagram like
  501. 19:24this I can basically say that Y is a
  502. 19:28linear function of X so this is what we
  503. 19:31are going to do now understand how we
  504. 19:33are going to create this best fit line
  505. 19:35this is very much important whenever we
  506. 19:36say linear regression it basically means
  507. 19:39that we are going to create a linear
  508. 19:40line over there you may be thinking sir
  509. 19:43why to create linear line why not
  510. 19:44nonlinear line that I'll discuss about
  511. 19:46it as we go ahead see other other
  512. 19:48algorithms so to begin with let's
  513. 19:51consider this line that you see over
  514. 19:53here right this line equation can be
  515. 19:56given by multiple equations someone some
  516. 19:58people people write yal mx + C some
  517. 20:01people write uh H some people write yal
  518. 20:05beta 0 + beta 1 into X some people write
  519. 20:08H Theta of xal to Theta 0 + Theta 1 into
  520. 20:13X many many equations are there for this
  521. 20:16this straight line this straight line
  522. 20:18many many equations are there with
  523. 20:20respect to many many different kind of
  524. 20:22notations but the first algorithm that I
  525. 20:24have probably learned of linear
  526. 20:26regression is from Andrew Ng definitely
  527. 20:29I would like to give him the entire
  528. 20:30credits and based on his notation
  529. 20:33whatever he has explained I'll try to
  530. 20:34explain you over here so the credits for
  531. 20:37this algorithm specifically goes to
  532. 20:40Andrew NG so let's consider this one
  533. 20:43over here in order to create this
  534. 20:45straight line I will basically use a
  535. 20:47equation which is called as H Theta so
  536. 20:50this is the equation of a straight line
  537. 20:52if I know the equation of the straight
  538. 20:54line whatever I can write I can write
  539. 20:56many things yal mx + C yal beta 0 + beta
  540. 21:001 * X and then I can also write one more
  541. 21:04that is H Theta of xal theta 0 + Theta 1
  542. 21:08into X of I here also you can basically
  543. 21:11say x of I here also you can say x of I
  544. 21:13now let's go ahead and let's take this
  545. 21:15equation for now let's take this
  546. 21:17equation of now so I'm I'm going to take
  547. 21:19out this equation and just write one
  548. 21:21equation through which I have also
  549. 21:23studied but I will definitely be adding
  550. 21:25some points which probably Andrew and
  551. 21:27could not mention mention in his video
  552. 21:29but I'll try my level best obviously he
  553. 21:32is the best I cannot even compare myself
  554. 21:34to him so Theta 0 + Theta 1 into X now
  555. 21:39let's understand what is Theta 0 Theta 1
  556. 21:42as I said that let's say I have a
  557. 21:44problem statement over here let's say I
  558. 21:47this is my X and this is my y this is my
  559. 21:49data points now what I'm doing I'm
  560. 21:51trying to create a best fit line like
  561. 21:53this now what is this best fit line what
  562. 21:55is uh when I say this best fit line is
  563. 21:57basically given by this equation what
  564. 21:59does Theta 0 basically indicate Theta 0
  565. 22:02over here is something called as
  566. 22:04intercept now what exactly is intercept
  567. 22:08intercept basically means that when your
  568. 22:10X is zero then H Theta of X is equal to
  569. 22:13Theta 0 so in this particular case
  570. 22:16intercept basically indicates that at
  571. 22:18what point you are meeting the Y AIS so
  572. 22:22this particular point is basically
  573. 22:24your intercept when your X is equal to 0
  574. 22:28at that point of time you'll be seeing
  575. 22:30that this line is intersecting the y-
  576. 22:32AIS whatever value this will be that is
  577. 22:34your intercept now the second thing is
  578. 22:37about your Theta 1 what is Theta 1 this
  579. 22:40is nothing but slope or coefficient now
  580. 22:43what does this basically indicate this
  581. 22:45indicates let let's say that this is the
  582. 22:47unit one unit in the x-axis and probably
  583. 22:50with respect to this I can find one
  584. 22:52point over here one point over here and
  585. 22:55if I try to draw this over here to here
  586. 22:57this is the unit movement in y so what
  587. 23:00does it basically say slope with the
  588. 23:02unit movement in one one unit movement
  589. 23:05towards the x-axis what is the unit
  590. 23:07movement in y- axis that is basically
  591. 23:09slope or coefficient Theta 0 and Theta 1
  592. 23:11two things and X of I is definitely your
  593. 23:14data points now our main aim is to
  594. 23:18create a best fit line in such a way
  595. 23:21that I I'll just try to show it to you
  596. 23:22what is our main aim let's let's
  597. 23:24understand what is the aim of a linear
  598. 23:26regression so if I take an example of
  599. 23:29linear regression I need to find out the
  600. 23:32best fit line in such a way that the
  601. 23:35distance
  602. 23:36between this data points that I have and
  603. 23:40the predicted points should be very very
  604. 23:42less suppose I'm creating a best fit
  605. 23:46line okay I'm creating a best fit line
  606. 23:49so with respect to this data points
  607. 23:51initially was this right but my
  608. 23:52predicted point is this point in this
  609. 23:55particular case my predicted point is
  610. 23:56this point so and if I do do the
  611. 23:58summation of all these points those
  612. 24:01distance should be minimal then only
  613. 24:04I'll be able to say that this is the
  614. 24:06best fit line so I I cannot definitely
  615. 24:08say that this is exactly the best fit
  616. 24:10line or not how will I say when I try to
  617. 24:13calculate the difference between this
  618. 24:15point and the predicted Point these are
  619. 24:17my predicted point right if I try to
  620. 24:19calculate the distance between them then
  621. 24:22I will basically have a aim to it should
  622. 24:24be minimal if I do the summation of all
  623. 24:26the distance it should be minimal
  624. 24:29so for that what I can do is that see
  625. 24:31you may be also thinking Krish why not
  626. 24:33just do one thing okay suppose if these
  627. 24:35are my data points why not just play and
  628. 24:38create multiple lines and try to compare
  629. 24:40what we can do is that we can compare
  630. 24:42multiple we can create multiple lines
  631. 24:44right like this and then whoever is
  632. 24:46giving the best minimal point I will go
  633. 24:48and select that but how many iteration
  634. 24:51you will do how you will come to know
  635. 24:52that okay this line is the best line so
  636. 24:55for that specific purpose we should
  637. 24:57start at one point and we should lead
  638. 25:01towards finding the best fit line start
  639. 25:04at one point and then we should go
  640. 25:06towards finding the best fit line so for
  641. 25:10this particular purpose what we do is
  642. 25:12that we create a something called as uh
  643. 25:15cost function I have already shown you
  644. 25:17what is my hypothesis function my best
  645. 25:19fit line equation is basically given as
  646. 25:21H Theta of x equal to Theta 0 + Theta 1
  647. 25:26* X this is my hypothesis right now
  648. 25:29coming to the cost function which is
  649. 25:32super super important why this it is
  650. 25:34super important because cost function
  651. 25:37basically what what is cost function
  652. 25:38over here I told right right this
  653. 25:41distance when I do the
  654. 25:42summation this distance that I when I'm
  655. 25:45doing the summation it should be minimal
  656. 25:48so if I really want to find out this
  657. 25:49particular distance I will be using one
  658. 25:51more equation how can I use a distance
  659. 25:54formula between the predicted and the
  660. 25:56real point I will just say that H Theta
  661. 26:00of x - y so when I say h Theta of x - Y
  662. 26:06what does this basically mean this is my
  663. 26:07real point and this is my predicted
  664. 26:10Point predicted point is basically given
  665. 26:12by H Theta of X and what I'm going to do
  666. 26:15I'm going to basically do the squaring
  667. 26:17because I may get a negative value so
  668. 26:18because of that I really want to do the
  669. 26:20squaring part Now understand one thing I
  670. 26:23need to also do the
  671. 26:25summation I = 1 to compl complete M
  672. 26:29let's say that I'm taking the number of
  673. 26:30data points over here as M because I
  674. 26:33need to calculate the distance between
  675. 26:34all the points right with respect to the
  676. 26:37predicted and the predict with respect
  677. 26:39to the real
  678. 26:40points so after this I also need to
  679. 26:44divide by 1X 2m the reason why I'm
  680. 26:47dividing by first of all let me show you
  681. 26:49why we are dividing by 1 by m 1 by m
  682. 26:51will give us the average of all the
  683. 26:53values that we have the specific reason
  684. 26:56why we are dividing by 1 by 2 do is for
  685. 26:59the derivation purpose it helps us to
  686. 27:02make our equation very much simpler so
  687. 27:05that later on when I am updating the
  688. 27:08weights when I say weights I'm basically
  689. 27:10updating Theta 0 and Theta 1 Theta 0 and
  690. 27:13Theta 1 at that point of time you'll be
  691. 27:15able to see that this particular value
  692. 27:18when we probably do the derivative it
  693. 27:20will help us to do it again I'm going to
  694. 27:22repeat it I'm going to write it down for
  695. 27:24you first of
  696. 27:26all now in order to find find out the
  697. 27:28best fit line I need to keep on changing
  698. 27:30Theta 0 and Theta 1 unless and until I
  699. 27:33get the best fit line unless and until I
  700. 27:35don't get the best fit line I need to
  701. 27:37keep on updating Theta 0 and Theta 1 now
  702. 27:40if I need to keep on updating Theta 0
  703. 27:42and Theta 1 I probably require a cost
  704. 27:45function okay what this cost function
  705. 27:47will do I'll just tell you so cost
  706. 27:49function over here I will specify as J
  707. 27:53of theta 0 comma Theta 1 is equal to now
  708. 27:57what is cost fun function over here what
  709. 27:59this distance I told right this distance
  710. 28:01between the H Theta of X and Y if I do
  711. 28:05the summation of all these things it
  712. 28:07needs to be minimal it needs to be less
  713. 28:10because with respect to an X point this
  714. 28:12is my y point
  715. 28:14right similarly with respect to this x
  716. 28:16point this is my y point so what I'm
  717. 28:19actually going to do I'm going to use a
  718. 28:20cost function now in this cost function
  719. 28:23my main aim is
  720. 28:25to basically write H Theta of x - y s
  721. 28:29this will be with respect to I I I why I
  722. 28:32am saying I because this will be moving
  723. 28:34from I equal to 1 to all the points that
  724. 28:37is m m is basically all the points over
  725. 28:41here now apart from this what I actually
  726. 28:44going to do I'm going to divide by 1X 2
  727. 28:46m I'll tell you why I'm specifically
  728. 28:48dividing by 1X 2 m first of all by
  729. 28:51dividing by m I will be getting an
  730. 28:53average
  731. 28:54output average cost function because
  732. 28:57here I'm iterating M the reason why I'm
  733. 29:00dividing by two because it will help us
  734. 29:01in derivation why let's say that I have
  735. 29:04x² if I try to find out derivative of x²
  736. 29:08with respect to X then what will I get I
  737. 29:11will basically get 2x right that is what
  738. 29:14is the formula what is the derivation of
  739. 29:16X of n it is nothing but n x of n
  740. 29:19minus1 so that is the reason why I'm
  741. 29:21actually making it 1 by two so that when
  742. 29:24two comes over here this two and two
  743. 29:26will get cancelled so I hope everybody's
  744. 29:29able to understand so this is my cost
  745. 29:32function Now understand what is this
  746. 29:34called as this entire equation is
  747. 29:36basically called as squared error
  748. 29:40function yes mathematical Simplicity
  749. 29:42basically means because when we are
  750. 29:44updating Theta 0 and Theta 1 we
  751. 29:46basically find out derivation in the
  752. 29:47cost function so that is the reason why
  753. 29:50we are specifically doing it squaring
  754. 29:52off is basically done because so that we
  755. 29:54don't get any negative values here
  756. 29:56squared error function now let's go
  757. 29:59towards the what we need to solve this
  758. 30:02is my cost function okay so I need to
  759. 30:07minimize minimize this particular value
  760. 30:10that is 1x 2 m summation of I = 1 2 m
  761. 30:15and then this will basically be H Theta
  762. 30:17of X of I minus y of I whole Square we
  763. 30:23need to minimize this by adjusting
  764. 30:26parameter Theta 0 and Theta 1
  765. 30:28this entirely is what this is nothing
  766. 30:31but J of theta 0 comma Theta 1 and we
  767. 30:36really need to minimize this so this is
  768. 30:38our task okay this is our task now let's
  769. 30:41go ahead and let's try to compare with
  770. 30:44two different thing one is the
  771. 30:46hypothesis testing and one is with
  772. 30:48respect to the cost
  773. 30:49function okay let's take an
  774. 30:52example so right now my equation of
  775. 30:58the
  776. 30:59hypothesis is nothing but H Theta of x
  777. 31:02equal to Theta 0 + Theta 1 *
  778. 31:06X if Theta 0 is 0 then what does this
  779. 31:11basically indicate can I say that it
  780. 31:14basically the line the line the best fit
  781. 31:16line passes through the origin and this
  782. 31:18is nothing but s Theta of xal to Theta
  783. 31:211 multiplied by X can I say like this
  784. 31:25obviously I can definitely say like this
  785. 31:27right so my equation will be like this
  786. 31:29so for right now let's consider that
  787. 31:33your Theta 0 is equal to 0 so this is
  788. 31:35what it is we have done till here we
  789. 31:37have minimized we have written the
  790. 31:39equation everything yes so it is passing
  791. 31:42through the origin and this is what is
  792. 31:44the equation I'm actually getting now
  793. 31:47let's take one example and let's try to
  794. 31:48solve this if I if I have H Theta of X
  795. 31:51so this is my new hypothesis considering
  796. 31:54that my intercept is passing through the
  797. 31:57region so with respect to this let's say
  798. 32:00that I will create one line over here
  799. 32:04let's say this is
  800. 32:05my this is my data points like X1 y1 I
  801. 32:11have 1 2 3 I have 1 2 3 now let's
  802. 32:19consider that if I have T I have data
  803. 32:22points like what I have data points like
  804. 32:24let's say I have three data points 1
  805. 32:26comma 1 2A 2 3 comma 3 so 1A 1 is
  806. 32:31nothing but this is my data point 2A 2
  807. 32:34is nothing but this is my data point and
  808. 32:363 comma 3 is this is my data point so
  809. 32:39these are my data points from the data
  810. 32:41set that I
  811. 32:43have so 2 comma 2 is this point and 3
  812. 32:47comma 3 is basically this point let's
  813. 32:49consider that these are my points that I
  814. 32:51have these are my data points now if I
  815. 32:54consider Theta 1 as 1 where do you think
  816. 32:57the straight line will pass through
  817. 32:59where do you think the straight line
  818. 33:00will pass the straight line will
  819. 33:02definitely pass like this right my
  820. 33:05straight line will definitely pass
  821. 33:06through all the points this same point
  822. 33:08becomes a prediction point also right
  823. 33:11same point let's consider that this is
  824. 33:13also getting pass through this it passes
  825. 33:15through all the points when Theta 1 is
  826. 33:17equal to 1 Theta 1 is nothing but slope
  827. 33:19when slope is equal to 1 in this
  828. 33:21scenario it passes through all the
  829. 33:22points now go ahead and calculate your J
  830. 33:25of theta so what will the form of J of
  831. 33:28theta 1 become because Theta 0 is 0 okay
  832. 33:31we can basically write 1 by 2 m
  833. 33:33summation of I = 1 2 three how many
  834. 33:36points are there three right and here I
  835. 33:39have J of H of theta of X1
  836. 33:43sorry X of theta of x i - y i
  837. 33:49s right now let's go ahead and compute
  838. 33:52now in this particular scenario what
  839. 33:54will happen 1X 2 m
  840. 33:57then what is what is this point minus y
  841. 34:00of I see h of X is also 1 y of I is also
  842. 34:04one both the point are 1 so this will
  843. 34:06become 1 - 1 whole S Plus because we are
  844. 34:09doing summation the next point is also
  845. 34:11falling in 2A 2 so this will become 2 -
  846. 34:132 s + 3 - 3 S so in total this will
  847. 34:18become zero so when your J of theta when
  848. 34:22Theta 1 is 1 Theta 1 is 1 so J of theta
  849. 34:261 is how much it is
  850. 34:29Z right so what is this J of theta 1 it
  851. 34:33is the cost function so let me draw the
  852. 34:35cost function graph over here let's say
  853. 34:39that this is my Theta and this is
  854. 34:42my so here I have 0.5 here I have 1 here
  855. 34:46I have 1.5 so this is my Theta here I
  856. 34:49have two then I have 2.5 okay then
  857. 34:52similarly I have 0. five then I have 1
  858. 34:581.5 2 2.5 this is my J of theta 1 so
  859. 35:04right now what is my Theta 1 my Theta 1
  860. 35:07is 1 at this particular Point what did I
  861. 35:09get J of theta 1 is nothing but zero so
  862. 35:12this will be my first point this will be
  863. 35:15my first point guys I have discussed why
  864. 35:18why the value will be 1X 2m basically to
  865. 35:20make the calculation simpler we are
  866. 35:22dividing by 1X 2 m is basically used to
  867. 35:26average aage is the sumission that we
  868. 35:28are actually doing over here now let's
  869. 35:30go ahead and let's take the second
  870. 35:32scenario in the second scenario let's
  871. 35:34consider my Theta 1 let's say that my
  872. 35:37Theta 1 over here is now 0.5 if my Theta
  873. 35:411 is 0.5 then tell me what are the
  874. 35:43points that I will get for x equal to
  875. 35:471.5 * 1 so it will come as 0.5 over
  876. 35:51here right then similarly when X is
  877. 35:54equal to 2.5 * 2 is nothing but 1 over
  878. 35:59here and then similarly when uh for x
  879. 36:03equal to
  880. 36:0435 multiplied by 3 see we are
  881. 36:07multiplying here right5 multi by 3 is
  882. 36:091.5 so the next point will come over
  883. 36:12here now when I create my best fit line
  884. 36:15what will happen so here is my next best
  885. 36:19fit line which I will probably create by
  886. 36:20green
  887. 36:23color okay so this is my second one
  888. 36:25which is green color here definitely
  889. 36:27slope is decreasing so if I go ahead and
  890. 36:30calculate my J of theta let's see what
  891. 36:32I'll get so J of theta
  892. 36:351 is nothing but 1X 2
  893. 36:39m again same equation summation of I = 1
  894. 36:422 3 H Theta of X of
  895. 36:46i - y of
  896. 36:49i² so what we have for over here we have
  897. 36:52nothing but 1X 2 m now let's do the
  898. 36:56summation what is this point this point
  899. 36:58is nothing but the predicted point and
  900. 37:01this point is the real point right so in
  901. 37:03this particular scenario the first point
  902. 37:05that I will get is nothing but. 5 - 1
  903. 37:10whole s how I'm getting. 5 - 1 whole
  904. 37:12Square this is 1 this is the real Point
  905. 37:151 this is the predicted Point .5 so here
  906. 37:18I'm getting. 5 - 1 whole Square the
  907. 37:21second point will be 1 - 2 whole s right
  908. 37:252 so 1 - 2 whole
  909. 37:28s and then I will finally get 1.5 - 3
  910. 37:34whole s so finally if I do this
  911. 37:36calculation how much I'm actually
  912. 37:38getting 1X 2 * 3 which is 6 here I'm
  913. 37:42getting
  914. 37:44.25 5 Square here I'm getting 1 here I'm
  915. 37:47getting 1.5 whole Square so my final
  916. 37:51output will be which I have already
  917. 37:53calculated it is nothing but point it
  918. 37:56will be approximately equal to. 58 so 58
  919. 38:01now with Theta as this is nothing but
  920. 38:04Theta Theta 1 as
  921. 38:07.5 right that is what Theta 1 as .5 we
  922. 38:11are able to get. 58 so Theta 1 is .5
  923. 38:15over here and. 58 will be coming
  924. 38:17somewhere here right so this is my next
  925. 38:20point which will be again in green color
  926. 38:23now let's go ahead and calculate the
  927. 38:24third condition now in third condition
  928. 38:26what I'm actually going to write I'm
  929. 38:28going to basically say Theta 1 as 0 at
  930. 38:31that point of time just go and assume
  931. 38:34what is 0 multiplied by X it will
  932. 38:36obviously be zero so I will be getting
  933. 38:38three points and my next line will be in
  934. 38:41this line that is the
  935. 38:45x-axis and this is basically all my
  936. 38:47points now if I go ahead and calculate
  937. 38:50this what is J of theta 1
  938. 38:52now what is J of theta 1 now in this
  939. 38:55particular case when my Theta 1 is equal
  940. 38:57= to 0 1X 2 m now this part you'll be
  941. 39:02able to see this is 0 - 1 0 - 2 0 -
  942. 39:083 okay so it will become 0 - 1 s 0 - 2 s
  943. 39:14and 0 - 3
  944. 39:16S okay so this will become 1X 6
  945. 39:20* 1 + 4 + 9 which will not be it will be
  946. 39:25nothing but 2.3 which is approximately
  947. 39:29equal to
  948. 39:302.3 then what will happen with respect
  949. 39:33to Theta 1 as 0 we are getting 2.3 so if
  950. 39:36I draw this it is nothing but with
  951. 39:38respect to zero I'm getting 2.
  952. 39:412
  953. 39:442.3 this is my point so similarly when
  954. 39:47you start constructing with Theta 1 is
  955. 39:49equal 2 I may get some point over here
  956. 39:52so here when I join this points
  957. 39:56together you will be seeing that I will
  958. 39:58be getting this kind of
  959. 40:01curve okay and this curve is something
  960. 40:04called as gradient
  961. 40:07descent and this gradient descent will
  962. 40:10play a very very important role in
  963. 40:14making sure that in making sure that you
  964. 40:17get the right Theta 1 value or light
  965. 40:20slope value now which is the most
  966. 40:22suitable point the most suitable point
  967. 40:24is to come over here because this is
  968. 40:27this this point is basically called AS
  969. 40:30Global
  970. 40:31Minima because see out of all these
  971. 40:34three lines which is the best fit line
  972. 40:35this is the best fit line right this is
  973. 40:38the best fit line when I had this best
  974. 40:40fit line my point that came over here
  975. 40:44was here itself this was my point that
  976. 40:46came over here right and I want to
  977. 40:48basically come to this region because
  978. 40:50this is my Global
  979. 40:52Minima when I basically am over here the
  980. 40:56distance between the predicted and the
  981. 40:58real point is very very less right so
  982. 41:02this specific point is basically called
  983. 41:04AS Global minimum but still I did not
  984. 41:07discuss Krish you have assumed Theta 1
  985. 41:10is 1 Theta 1 is .5 Theta 1 is 0 here
  986. 41:13also you're assuming many things right
  987. 41:15and then you probably calculating and
  988. 41:17you're creating this gradient descent
  989. 41:19but the thing should be that probably
  990. 41:22you come to one point over here and then
  991. 41:25you reach towards this so for that
  992. 41:27specific reason how do you do that how
  993. 41:30do I first of all come to a point and
  994. 41:32then move towards This Global Minima so
  995. 41:35for that specific case we will be using
  996. 41:37one convergence algorithm because if I
  997. 41:40come to one specific point after that I
  998. 41:43just need to keep on updating Theta 1
  999. 41:45instead of using different different
  1000. 41:47Theta 1 value so for this we use
  1001. 41:50something called as convergence
  1002. 41:52algorithm so here the convergence
  1003. 41:54algorithm basically says
  1004. 41:59repeat until
  1005. 42:03convergence that basically means I'm in
  1006. 42:05a while loop let's say and here I'm
  1007. 42:08basically going to update my Theta value
  1008. 42:11which will be given by this notation
  1009. 42:13which is continuous updation where I'll
  1010. 42:15say Theta J minus I'll talk about this
  1011. 42:19Alpha don't worry and then it will be
  1012. 42:22derivative of theta
  1013. 42:25J with respect to this J of theta
  1014. 42:290 and Theta 1 so this should happen that
  1015. 42:34basically means after we reach to a
  1016. 42:36specific point of theta after performing
  1017. 42:40this particular operation we should be
  1018. 42:43able to come to the global Minima and
  1019. 42:45this this specific thing that you are
  1020. 42:47able to see is called as
  1021. 42:50derivative this is called as derivative
  1022. 42:52derivative basically means I'm trying to
  1023. 42:54find out the slope
  1024. 42:57derivative which I can also say it as
  1025. 42:59slope this equation will definitely work
  1026. 43:02guys trust me this will definitely work
  1027. 43:04why it will work I'll just draw it show
  1028. 43:06it to you let's say that this is my cost
  1029. 43:09function let's say that I've got this
  1030. 43:11gradient
  1031. 43:12descent and let's say that my first
  1032. 43:15point is somewhere here but I have to
  1033. 43:18reach somewhere here right now when I
  1034. 43:20reach this this is my Theta 1 and this
  1035. 43:23is my J of theta 1 suppose I reach at
  1036. 43:25this specific point and I will also have
  1037. 43:28another gradient descent which looks
  1038. 43:30like this let's say that in the initial
  1039. 43:33time I reach the point over here how we
  1040. 43:35will be coming to this minimal Global
  1041. 43:37Minima by using this equation I'll talk
  1042. 43:40about Alpha also don't worry now this is
  1043. 43:42also my Theta 1 this is also my J of
  1044. 43:44theta 1 now let's say suppose I came to
  1045. 43:47this particular point right after coming
  1046. 43:49to this particular point I will
  1047. 43:52basically apply this derivative on this
  1048. 43:55J of theta 1 okay now when I find out a
  1049. 43:59derivative that basically means we are
  1050. 44:00trying to find out the slope and in
  1051. 44:02order to find the slope we just create a
  1052. 44:04straight line like
  1053. 44:05this which will look like this I'll just
  1054. 44:08try to
  1055. 44:10create so I'll try to create a slope
  1056. 44:12like this this
  1057. 44:15slope so if you try to find out with
  1058. 44:17respect to this this is a positive slope
  1059. 44:20how do we indicate it because understand
  1060. 44:22the right hand side of the line of this
  1061. 44:24is pointing on the top wordss Direction
  1062. 44:27this is the best easy way to find out
  1063. 44:30whether it is a positive slope or
  1064. 44:31negative slope now in this particular
  1065. 44:33case this is a positive slope now when I
  1066. 44:36get a positive slope that basically
  1067. 44:38means I will update my weights or Theta
  1068. 44:401 as Theta 1 let's say I'm writing it
  1069. 44:44over here so I will just apply this
  1070. 44:46convergence algorithm see Theta
  1071. 44:491 colon Theta 1 minus this learning rate
  1072. 44:55which is called as Alpha this is my my
  1073. 44:57learning rate I'll talk about learning
  1074. 44:58rate don't worry then this derivative
  1075. 45:02value in this particular case since I'm
  1076. 45:04having a positive slope I will be
  1077. 45:06getting a positive value let's say that
  1078. 45:09for this Theta value I got this slope
  1079. 45:12initially now I need to come to this
  1080. 45:15location so for that I have to reduce
  1081. 45:17Theta 1 so that I come to this main
  1082. 45:20point now here you can see that I am I
  1083. 45:23subtracting Theta 1 with something which
  1084. 45:25is a positive number
  1085. 45:28right this is a positive number so
  1086. 45:29definitely I know that after some n
  1087. 45:31number of iteration I will be able to
  1088. 45:34come to the global Minima similarly if I
  1089. 45:36take the right hand side and if I try to
  1090. 45:38draw the slope in this particular case
  1091. 45:40my slope will be
  1092. 45:42negative so similarly I can write the
  1093. 45:44equation as Theta
  1094. 45:461 = to Theta 1 minus learning rate
  1095. 45:51multiplied by a negative number so minus
  1096. 45:54into minus will be positive right
  1097. 45:55suppose initially my 1 was
  1098. 45:58here my Theta 1 was here now I'll keep
  1099. 46:01on updating the weight to come to this
  1100. 46:02Global Minima so minus into minus is
  1101. 46:06positive so I will basically get Theta 1
  1102. 46:09+
  1103. 46:10Alpha by a positive number because minus
  1104. 46:13into minus is plus so this will
  1105. 46:16definitely work so that we will be able
  1106. 46:19to come over here to the global Minima
  1107. 46:22whether it is a positive slope or a
  1108. 46:24negative slope now what is this learning
  1109. 46:26learning rate now learning rate based on
  1110. 46:30this learning rate suppose I want to
  1111. 46:32come from this point to the global
  1112. 46:35Minima by what speed I should be coming
  1113. 46:39what speed if my learning rate value is
  1114. 46:41bigger what speed I may be coming
  1115. 46:43suppose if I say usually we select
  1116. 46:45learning rate as 01 if I select a small
  1117. 46:48number then it'll start taking small
  1118. 46:50small steps to move towards the optimal
  1119. 46:52Minima but if I take a alpha value a
  1120. 46:55huge value if it is a huge huge value
  1121. 46:57then what will happen this uh this
  1122. 47:00updation of the Theta 1 will keep on
  1123. 47:02jumping here and there and the situation
  1124. 47:03will be that it will never meet it will
  1125. 47:07never reach the global Minima so it is a
  1126. 47:09very very good decision to take a alpha
  1127. 47:12small value it should also not be a very
  1128. 47:13very small value if it becomes a very
  1129. 47:16very small value then what will happen
  1130. 47:18very tiny steps it will take forever to
  1131. 47:20reach the global Minima that basically
  1132. 47:22means my model will keep on training
  1133. 47:24itself so definitely this Al is going to
  1134. 47:27work now let me talk about one
  1135. 47:30scenario one scenario will be that what
  1136. 47:33if my my cost function has a local
  1137. 47:36Minima what if I have a local Minima
  1138. 47:39because here if I
  1139. 47:41come here if I come this is a local
  1140. 47:43Minima suppose one of my points come
  1141. 47:46over here and finally I'm reaching over
  1142. 47:48here what will happen in this particular
  1143. 47:50case because in this case you'll be
  1144. 47:52seeing that what will be my equation my
  1145. 47:54equation will be simply Theta 1
  1146. 47:57Theta 1 minus Alpha in this point in
  1147. 48:01this local Minima slope will be zero so
  1148. 48:03in this particular case my Theta 1 will
  1149. 48:05be equal to Theta 1 now you may be
  1150. 48:07thinking what is if this is the scenario
  1151. 48:10then we will be stuck in local Minima
  1152. 48:13this is called as local
  1153. 48:15Minima but usually with respect to the
  1154. 48:18gradient descent and the equation that
  1155. 48:20we are using here we do not get stuck in
  1156. 48:23local Minima because our gradient
  1157. 48:25descent in this particular scenar iio
  1158. 48:27will always look like this but yes in
  1159. 48:29deep learning when we are learning about
  1160. 48:31grade in descent and a Ann at that point
  1161. 48:34of time we have lot of local Minima and
  1162. 48:37because of that we have different
  1163. 48:38different G decent algorithm like RMS
  1164. 48:40prop we have Adam optimizers which will
  1165. 48:43solve that specific problem so this one
  1166. 48:46point also I wanted to mention because
  1167. 48:48tomorrow if someone asks you as an
  1168. 48:49interview question that what if in your
  1169. 48:52uh do you see any local Minima in linear
  1170. 48:54regression you can just that the cost
  1171. 48:57function that we use will definitely not
  1172. 49:00give us local Minima but if in deep
  1173. 49:02learning techniques with that we are
  1174. 49:03trying to use like Ann we have different
  1175. 49:05different kind of optimizers which will
  1176. 49:07solve that particular problem so that is
  1177. 49:10the answer you basically have to give
  1178. 49:12now let me go ahead and write with
  1179. 49:14respect to the gradient descent
  1180. 49:15algorithm so here again I'm going to
  1181. 49:17write the gradient descent algorithm so
  1182. 49:19this will be my gradient descent
  1183. 49:21algorithm and remember guys gradient
  1184. 49:24descent is an amazing algorithm and you
  1185. 49:26you will definitely be using it so
  1186. 49:29please make sure that you know this
  1187. 49:32perfectly now some questions are that
  1188. 49:35when will convergence stop convergence
  1189. 49:37will stop when we come to near this area
  1190. 49:40where my uh J of theta will be very very
  1191. 49:44less now in gradient descent algorithm I
  1192. 49:47will again repeat it so what did I say I
  1193. 49:50said
  1194. 49:51repeat until convergence I told you
  1195. 49:54right here we have written this
  1196. 49:55algorithm
  1197. 49:57and now let's take it for Theta 0 and
  1198. 49:59Theta 1 so here I will write Theta 0
  1199. 50:02J equal to Theta
  1200. 50:06J minus learning rate of derivative of
  1201. 50:11theta
  1202. 50:14J J of theta 0 and Theta 1 so this is my
  1203. 50:19repeat until convergence now we really
  1204. 50:22need to find out what we'll try to
  1205. 50:24equate we'll try to first of all find
  1206. 50:25out what is this
  1207. 50:28now if I really want to find out
  1208. 50:30derivative
  1209. 50:32of derivative of derivative of theta J
  1210. 50:37with respect to J of theta 0 and Theta 1
  1211. 50:41so how do I write this I can definitely
  1212. 50:44write this in a easy way okay so this
  1213. 50:46will be derivative of theta J and
  1214. 50:49remember J will be 0 and 1 right because
  1215. 50:53we need to find out for 0 Theta 0 and
  1216. 50:55Theta 1 so this will be 1 by 2 m what is
  1217. 50:59what is J of theta 0a Theta 1 obviously
  1218. 51:02my cost function so I will write
  1219. 51:04summation of IAL 1 to M and here I will
  1220. 51:08basically write J of theta of X of I
  1221. 51:11minus y of I whole squar so if my J is
  1222. 51:16equal to Z so what will happen for this
  1223. 51:19so here I can specifically say that
  1224. 51:21derivative of derivative of theta 0 J of
  1225. 51:25theta 0a 1
  1226. 51:27now it's simple here what I will be
  1227. 51:29doing is that I will be simply applying
  1228. 51:31derivative function see guys what is
  1229. 51:34this derivative let's consider this is
  1230. 51:36something like this 1X 2 m x² so if I
  1231. 51:40try to find out the derivative this will
  1232. 51:42be 2x 2 MX so 2 and 2 will get cancel so
  1233. 51:46similarly I'll have 1 by m and here I
  1234. 51:49will specifically be writing summation
  1235. 51:52of I = 1 2 m h Theta of x X of I which
  1236. 51:58will be my
  1237. 51:59x - y of i² so this will be my
  1238. 52:03derivative with respect to Theta 0 this
  1239. 52:06is what I got now the second thing will
  1240. 52:08be that when J is equal to 1 derivative
  1241. 52:11of derivative of theta 1 J of theta 0
  1242. 52:15comma Theta
  1243. 52:161 in this particular case I will be
  1244. 52:19having 1 by m summation of I = 1 to M
  1245. 52:23then again see in this particular case
  1246. 52:26Theta of 1 is there right Theta of 1
  1247. 52:29basically means what if I try to replace
  1248. 52:31this let's say that I'm trying to
  1249. 52:33replace this H Theta of X with something
  1250. 52:35else what is s Theta of X I know that
  1251. 52:38right it is Theta 0 + Theta 1 * X so
  1252. 52:42Theta 0 + Theta 1 * X so after this if
  1253. 52:46I'm trying to find out the derivative
  1254. 52:48with respect to Theta 0 this will
  1255. 52:50obviously become I will be able to get
  1256. 52:52this much right now with respect to the
  1257. 52:54second derivative what I will be writing
  1258. 52:56I will again be writing H thet of X of i
  1259. 52:59- y of i s
  1260. 53:03multiplied X of I so this Square also
  1261. 53:06went off understand this H Theta of X is
  1262. 53:09what see they H Theta of X is nothing
  1263. 53:12but Theta 0 + Theta 1 * X so if I'm
  1264. 53:16trying to find out derivative with
  1265. 53:18respect to Theta 0 nothing will be going
  1266. 53:19to come okay Theta 1 of X will become a
  1267. 53:22constant in this particular case in this
  1268. 53:25case because Theta 1 of X is there so if
  1269. 53:28I try to find out derivative of theta 1
  1270. 53:30into X only I'll be getting X Y Square
  1271. 53:33will not be there it's easy right X squ
  1272. 53:35means 2x this is the derivative of x
  1273. 53:37square right so that square went and 1X
  1274. 53:402 1 2 by two got cancelled so this will
  1275. 53:44be now my convergence algorithm so here
  1276. 53:47we have discussed about linear
  1277. 53:48regression oh sorry I have to remove
  1278. 53:50Square here also so let me write it
  1279. 53:53again okay repeat until conver con let
  1280. 53:57me write it down again repeat until
  1281. 53:59convergence finally your two updates
  1282. 54:03will be happening one is Theta 0 so here
  1283. 54:06it will be Theta 0
  1284. 54:09minus Alpha that is my learning rate 1
  1285. 54:12by m summation of IAL 1 to M and this
  1286. 54:17will basically be H Theta of X of I
  1287. 54:21minus y of
  1288. 54:23I and similarly if I want to update
  1289. 54:26Theta 1 it will be - alpha 1 by m
  1290. 54:30summation of I = 1 to m h Theta of X of
  1291. 54:36I oh my God y of I uh multiplied by X of
  1292. 54:42I Alpha is your learning rate guys Alpha
  1293. 54:45is nothing but it is learning rate here
  1294. 54:48we have to initialize some value like
  1295. 54:510.1 see what is s Theta of X Theta 0 +
  1296. 54:55Theta 1 into X right if I do derivative
  1297. 54:58of theta 1 into x what is derivative of
  1298. 55:01theta 1 with Theta 1 x it is nothing but
  1299. 55:03X so this x will come over here now
  1300. 55:07let's discuss about two important thing
  1301. 55:09one is R square and adjusted R square
  1302. 55:11now similarly what will happen you will
  1303. 55:14have lot of convex functions now see if
  1304. 55:16I talk about uh like if you have
  1305. 55:19multiple features like X1 X2 X3 x4 at
  1306. 55:23that point of time you will be having a
  1307. 55:253D curve curve which looks like this
  1308. 55:28gradient
  1309. 55:29decent which will be something like this
  1310. 55:40gradient it's just like coming down a
  1311. 55:44mountain now let's discuss about two
  1312. 55:46performance metrics which is important
  1313. 55:48in this particular case one is R
  1314. 55:52square and adjusted R square
  1315. 55:57we usually use this performance metrix
  1316. 55:59to verify how our model is and how good
  1317. 56:01our model is with respect to linear
  1318. 56:03regression so R square is basically
  1319. 56:05given R square is a performance Matrix
  1320. 56:07to check how good the specific model is
  1321. 56:10so here we basically have a formula
  1322. 56:12which is like 1 minus sum of residual
  1323. 56:16divided by sum of total now this is the
  1324. 56:19formula of R squ now what is this sum of
  1325. 56:21residual I can basically write like this
  1326. 56:23summation of y i Min - y i hat whole
  1327. 56:29Square this Yi hat is nothing but H
  1328. 56:31Theta of X just consider in this way
  1329. 56:33divided by summation of Y of i - y mean
  1330. 56:39y mean y s to formula this is the
  1331. 56:42formula I'll try to explain you what
  1332. 56:44this formula definitely says okay so
  1333. 56:47first thing first let's consider that
  1334. 56:49this is my this is my problem statement
  1335. 56:51that I'm trying to solve suppose these
  1336. 56:53are my data points and if I try to
  1337. 56:55create the best fit
  1338. 56:57line This Yi hat Yi hat basically means
  1339. 57:01this specific point we are trying to
  1340. 57:03find out the difference between this
  1341. 57:05things difference between these things
  1342. 57:07let's say that these are my points I'm
  1343. 57:09trying to find out a difference between
  1344. 57:11this predicted this is my predicted the
  1345. 57:13point in green color are my predicted
  1346. 57:15points which I have denoted as y i hat
  1347. 57:18and always understand this is what Su
  1348. 57:21sum of residual is sum of residual is
  1349. 57:23nothing but difference between this
  1350. 57:24point to this point this point to this
  1351. 57:26point this point to this point this
  1352. 57:27point to this point and I doing the all
  1353. 57:29the summation of those now the next
  1354. 57:32point which is very much important here
  1355. 57:34is my X and Y what is this y IUS y y bar
  1356. 57:39Y Bar is nothing but mean mean of Y if I
  1357. 57:43calculate the mean of Y then I will
  1358. 57:45probably get a line which looks like
  1359. 57:47this I'll get a line something like this
  1360. 57:49and then I will probably try to
  1361. 57:51calculate the distance between each and
  1362. 57:53every point and this specific point with
  1363. 57:55respect to the distance between this
  1364. 57:57point and this point the denominator
  1365. 57:59will definitely be high right this value
  1366. 58:02obviously this value will be higher than
  1367. 58:04this value right the reason why it will
  1368. 58:07be higher because the mean of this
  1369. 58:09particular value distance will obviously
  1370. 58:11be higher so this 1 minus high this will
  1371. 58:16be a low value and this will be a high
  1372. 58:18value when I try to divide Low by
  1373. 58:23High Low by high then obviously this
  1374. 58:26entire number will become a small number
  1375. 58:28when this is a small number 1 minus
  1376. 58:30small number will be a big number so
  1377. 58:33this basically shows that our R square
  1378. 58:35has fitted properly right it has
  1379. 58:38basically got a very good R square now
  1380. 58:40tell me can I get this entire R square a
  1381. 58:43negative number let's say that in this
  1382. 58:44particular case I got 90% can I get this
  1383. 58:47R square as negative number there will
  1384. 58:50be situation guys what if I create a
  1385. 58:52best fit line which looks like
  1386. 58:54this if I create this best fit line
  1387. 58:57which looks like this then this value
  1388. 58:59will be quite High it is only possible
  1389. 59:02when this value will be higher
  1390. 59:05than higher than this
  1391. 59:08value okay but in the usual scenario it
  1392. 59:11will not happen because obviously we'll
  1393. 59:13try to fit a line which will be at least
  1394. 59:16good it's not just like pulling one line
  1395. 59:19somewhere we don't want to create a best
  1396. 59:21fit line which is worse than this right
  1397. 59:23worse than this so in this particular
  1398. 59:26scenario you'll be saying that in R
  1399. 59:28square now here you'll be able to see
  1400. 59:31one one amazing feature about R square
  1401. 59:33is that let's say let's say one scenario
  1402. 59:36suppose I have features like let's say
  1403. 59:38that my feature is something like uh
  1404. 59:41let's say I have a price of a house okay
  1405. 59:43so suppose this is my bedrooms how many
  1406. 59:45bedrooms I have and this is basically
  1407. 59:48the price of the house now if I if I
  1408. 59:51probably solve this Pro problem I'll
  1409. 59:53definitely get an R square value let's
  1410. 59:54say the R square value is 85% let's say
  1411. 59:57that my R square is 85% now what if if I
  1412. 1:00:00add one more feature the one more
  1413. 1:00:02feature basically says that okay if I
  1414. 1:00:05add
  1415. 1:00:06location location of the house will be
  1416. 1:00:09definitely correlated with price so
  1417. 1:00:12there is a definite chance that the R
  1418. 1:00:14square value will increase let's say
  1419. 1:00:16that R square will become 90% if I
  1420. 1:00:19probably have this two specific feature
  1421. 1:00:21and obviously it is basically increasing
  1422. 1:00:23the R square because this is also
  1423. 1:00:24correlated to price
  1424. 1:00:26and let me change the example see first
  1425. 1:00:29case I got by R square as 85% let's say
  1426. 1:00:32now as soon as I added location I got
  1427. 1:00:3590% now let's say that I added one more
  1428. 1:00:37feature which gender is going to stay
  1429. 1:00:40gender like male or female is going to
  1430. 1:00:42stay you know that gender is no way
  1431. 1:00:44correlated to price but even though I
  1432. 1:00:47add one feature there is a scenario that
  1433. 1:00:48my R square will still increase and it
  1434. 1:00:51may become
  1435. 1:00:5291% even though my feature is not that
  1436. 1:00:56important even gender is not that
  1437. 1:00:58important the R square formula Works in
  1438. 1:01:01such a way that if I keep on adding
  1439. 1:01:03features and that are not nowhere
  1440. 1:01:05correlated this is obviously nowhere
  1441. 1:01:07correlated this is not correlated with
  1442. 1:01:10price then also what it does is that it
  1443. 1:01:13is basically increasing my r² so this
  1444. 1:01:16specific thing should not happen whether
  1445. 1:01:19a male will stay or female will stay
  1446. 1:01:21that does not matter at all still when
  1447. 1:01:23you do the calculation the R square will
  1448. 1:01:26still increase so in order to not impact
  1449. 1:01:30the model because see now right now with
  1450. 1:01:32this particular model where I have got
  1451. 1:01:3490% now as soon as I see R square as 91%
  1452. 1:01:38because it is considering this
  1453. 1:01:40particular gender so this model will be
  1454. 1:01:43picked right because it is performing
  1455. 1:01:45well and is giving you a better R square
  1456. 1:01:46value but this should not happen because
  1457. 1:01:49that is not at all corelated this model
  1458. 1:01:51should have been picked so in order to
  1459. 1:01:53prevent this situation what we do we
  1460. 1:01:55basically Ally use something called as
  1461. 1:01:57adjusted R square now what is this
  1462. 1:01:59adjusted R square and how it will work
  1463. 1:02:02I'll also show it to you very very nice
  1464. 1:02:04concept of adjusted R square so adjusted
  1465. 1:02:06R square R square
  1466. 1:02:08adjusted is given by the
  1467. 1:02:11formula is given by the Formula 1 - 1 -
  1468. 1:02:16r² * N - 1 where n is the total number
  1469. 1:02:20of samples n minus P minus 1 this p p is
  1470. 1:02:24nothing but number of features
  1471. 1:02:26or predictors we'll also say or
  1472. 1:02:28predictors suppose initially my number
  1473. 1:02:31of predictors were in this particular
  1474. 1:02:33scenario in this scenario where I saw
  1475. 1:02:35this my number of predictors was two and
  1476. 1:02:37in this particular case my number of
  1477. 1:02:39predictor was three now if my predictor
  1478. 1:02:41is 2 I got the r squ as 90% so in this
  1479. 1:02:45particular scenario what all the
  1480. 1:02:46calculation will happen okay all the
  1481. 1:02:48calculation will happen and let's say
  1482. 1:02:50that my R square adjusted it'll be
  1483. 1:02:52little bit less it'll be little bit less
  1484. 1:02:55let's say it8 is 6% let's say that my R
  1485. 1:02:57square adjusted is 86% based on this
  1486. 1:03:00predictor 2 now when I use my predictor
  1487. 1:03:033 predictor basically means number of
  1488. 1:03:05features that I'm going to use and now
  1489. 1:03:08in this one one feature is nowhere
  1490. 1:03:10related like gender but what we are
  1491. 1:03:12getting we are basically getting R
  1492. 1:03:14square increased to
  1493. 1:03:1691% now for the R square
  1494. 1:03:19adjusted this will not increase this
  1495. 1:03:21will in turn decrease right now it will
  1496. 1:03:24become 82% how it will become I'll show
  1497. 1:03:26you I've just considered some value 8682
  1498. 1:03:29here you can see that there is an
  1499. 1:03:31increase here an increase is there here
  1500. 1:03:33decrease is there now how this is
  1501. 1:03:35basically happening see this P value
  1502. 1:03:39that I will be putting okay if I put a p
  1503. 1:03:42isal 3 obviously with n minus P minus 1
  1504. 1:03:46this will become a little bit smaller
  1505. 1:03:48number or sorry little bit uh smaller
  1506. 1:03:50number right so now in this particular
  1507. 1:03:53case if it is not correlated obviously
  1508. 1:03:55this will be high when I'm increasing
  1509. 1:03:56this so this will also be high let me
  1510. 1:03:58write the equation something like this
  1511. 1:04:00just a second so this will basically
  1512. 1:04:04be okay now why probably this value may
  1513. 1:04:08have decreased let me talk about this
  1514. 1:04:10one what is r squ I hope everybody
  1515. 1:04:12understood n is the number of data
  1516. 1:04:17points p is the number of
  1517. 1:04:21predictors if p is increasing then what
  1518. 1:04:24will happen as P keeps on increasing
  1519. 1:04:27this value will keep on
  1520. 1:04:29decreasing this value will keep on
  1521. 1:04:31decreasing if this values keep on
  1522. 1:04:33decreasing this will be a bigger number
  1523. 1:04:35this will obviously be a big number a
  1524. 1:04:38big number divided by a small number
  1525. 1:04:40what it will be obviously this will be a
  1526. 1:04:42little bit bigger number 1 minus bigger
  1527. 1:04:45number we will basically get some values
  1528. 1:04:47which will be decreasing if my P value
  1529. 1:04:49is two in this particular case it will
  1530. 1:04:52be less smaller than this right at least
  1531. 1:04:54it will be greater than this this
  1532. 1:04:55particular value right when p is equal
  1533. 1:04:57to
  1534. 1:04:573 so with the help of P obviously R
  1535. 1:05:01square is there to support you okay
  1536. 1:05:03whether it is correlated or not always
  1537. 1:05:05remember when the features are highly
  1538. 1:05:07correlated your R square value will
  1539. 1:05:09increase tremendously if it is less
  1540. 1:05:12correlated then it will be there will be
  1541. 1:05:14a small increase but there will not be a
  1542. 1:05:16very huge increase now if I consider p
  1543. 1:05:18is equal to 2 obviously when I'm trying
  1544. 1:05:20to find out this uh calculation n minus
  1545. 1:05:22P minus 1 it will obviously be greater
  1546. 1:05:25than p is equal to 3 when p is equal to
  1547. 1:05:283 then this value will be still more
  1548. 1:05:30smaller and when we are dividing a
  1549. 1:05:32bigger number by a smaller number
  1550. 1:05:34obviously we are subtracting with one so
  1551. 1:05:37that basically means even though my R
  1552. 1:05:39square is 86 over here there may be a
  1553. 1:05:41scenario since this is nowhere
  1554. 1:05:43correlated I'm basically getting an 82%
  1555. 1:05:45because of this entire equation so I
  1556. 1:05:48hope you are understanding this this is
  1557. 1:05:50very much important to understand a very
  1558. 1:05:53very important property simple way to
  1559. 1:05:55define is that as my P value keeps on
  1560. 1:05:58increasing the number of predictors
  1561. 1:06:00keeps on increasing my R squ gets
  1562. 1:06:02adjusted whatever R square I'm getting
  1563. 1:06:05with respect to this it will always be
  1564. 1:06:07less than this particular R square there
  1565. 1:06:10was one interview question that was
  1566. 1:06:11asked one of my student between R square
  1567. 1:06:14and adjusted R square which will always
  1568. 1:06:15be bigger definitely the student said R
  1569. 1:06:18square then he told him to explain about
  1570. 1:06:20adjusted R square why does that specific
  1571. 1:06:22happen agenda one is about Ridge lasso
  1572. 1:06:27regression second is assumptions of
  1573. 1:06:31linear regression the third point that
  1574. 1:06:34we are probably going to discuss about
  1575. 1:06:37is logistic regression then the fourth
  1576. 1:06:42thing that we are going to discuss about
  1577. 1:06:43is something called as confusion
  1578. 1:06:46Matrix the fifth thing that we are going
  1579. 1:06:49to consider about
  1580. 1:06:51is practicals
  1581. 1:06:54for lead lineer Ridge lasso and logistic
  1582. 1:07:00so first topic uh that we are probably
  1583. 1:07:03going to discuss is something called as
  1584. 1:07:05Ridge and lasso
  1585. 1:07:10regression so let's understand about
  1586. 1:07:12Ridge and lasso regression if you
  1587. 1:07:15remember in our previous session what
  1588. 1:07:17all things we discussed linear
  1589. 1:07:21regression and then we had discussed
  1590. 1:07:23about the cost function we have
  1591. 1:07:24discussed about R square adjusted
  1592. 1:07:26adjusted R square sorry R square and
  1593. 1:07:29adjusted R square we have discussed
  1594. 1:07:30about it gradient descent we have
  1595. 1:07:32discussed about it it was nothing but 1
  1596. 1:07:34by 2 m summation of I = 1 2 m h Theta of
  1597. 1:07:41x i -
  1598. 1:07:45y - y i s so this is the cost function
  1599. 1:07:50that we had discussed right yesterday
  1600. 1:07:53and this cost function was able to give
  1601. 1:07:55us a
  1602. 1:07:57gradient descent with respect to the J
  1603. 1:07:59of
  1604. 1:08:00theta J of theta Zer or Theta not so I
  1605. 1:08:03can also write this as J of theta comma
  1606. 1:08:06Theta 0 comma Theta 1 now let me give
  1607. 1:08:09you a scenario let's say that I have a
  1608. 1:08:11scenario over here and I have this
  1609. 1:08:14specific scenario let's say that I just
  1610. 1:08:16have two points which looks like this
  1611. 1:08:20okay now if I have these two specific
  1612. 1:08:23points what will happen I will probably
  1613. 1:08:25try to create a best fit line the best
  1614. 1:08:27fit line will definitely pass through
  1615. 1:08:29all the points like this if I try to
  1616. 1:08:32calculate the cost function what will be
  1617. 1:08:34the value of J of theta 0 comma Theta 1
  1618. 1:08:38let's say that in this particular case
  1619. 1:08:39since it is passing through the origin
  1620. 1:08:41my Theta 0 will be zero okay so what
  1621. 1:08:44will be the value of theta 0 comma Theta
  1622. 1:08:471 so here obviously you can see that
  1623. 1:08:49there is no difference so it will
  1624. 1:08:50obviously become zero Now understand
  1625. 1:08:54this data that you see right right this
  1626. 1:08:56data is basically called as training
  1627. 1:08:59data so this data that I have actually
  1628. 1:09:01plotted with two points these are
  1629. 1:09:03specifically called as training
  1630. 1:09:05data now what is the problem in this
  1631. 1:09:08data right now see right now exactly
  1632. 1:09:11whatever line is basically getting
  1633. 1:09:13created over here which is through the
  1634. 1:09:16uh hypothesis over here you can see that
  1635. 1:09:18it is passing through every point so
  1636. 1:09:19that is the reason your cost is zero and
  1637. 1:09:21our main aim is to basically minimize
  1638. 1:09:23the cost function that is absolutely
  1639. 1:09:26fine now in this particular case in
  1640. 1:09:29which my model this if this model is
  1641. 1:09:32getting trained initially this data is
  1642. 1:09:34basically called as training data now
  1643. 1:09:37just imagine that tomorrow new data
  1644. 1:09:40points comes so if my new data points
  1645. 1:09:42are here let's consider that I I want to
  1646. 1:09:45basically uh come up with this new data
  1647. 1:09:48point now in this particular scenario if
  1648. 1:09:50I want to predict with respect to this
  1649. 1:09:52particular Point let's say my predicted
  1650. 1:09:54point is here
  1651. 1:09:55is this the difference between the
  1652. 1:09:57predicted and the real Point quite
  1653. 1:10:00huge yes or no so this is basically
  1654. 1:10:04creating a condition which is called as
  1655. 1:10:07overfitting that basically means even
  1656. 1:10:11though my
  1657. 1:10:13model has given or trained well with the
  1658. 1:10:16training
  1659. 1:10:17data or let me write it down properly
  1660. 1:10:20over here so this condition since since
  1661. 1:10:23you can see that over here my each and
  1662. 1:10:25every point is basically passing through
  1663. 1:10:27the best fit line so because of that
  1664. 1:10:30what happens it causes something called
  1665. 1:10:32as
  1666. 1:10:33overfitting so you really need to
  1667. 1:10:35understand what is overfitting now what
  1668. 1:10:37does overfitting mean overfitting
  1669. 1:10:40basically means my model performs well
  1670. 1:10:44with training data but it fails to
  1671. 1:10:48perform well with test data now what is
  1672. 1:10:51the test data over here the test data is
  1673. 1:10:53basically this points the real test data
  1674. 1:10:55answer was this points but because the
  1675. 1:10:58my line is like this I'm actually
  1676. 1:11:00getting the predicted point over here so
  1677. 1:11:02this distance if I try to calculate it
  1678. 1:11:03is quite huge so in this scenario
  1679. 1:11:06whenever I say my model performs well
  1680. 1:11:08with training data and it fails to
  1681. 1:11:10perform well with test data then this
  1682. 1:11:12scenario we say it as overfitting so
  1683. 1:11:14this scenario when the model performs
  1684. 1:11:16well with training data I have a
  1685. 1:11:18condition which is called as low bias
  1686. 1:11:20and when it fails to perform with the
  1687. 1:11:22test data then it is basically called as
  1688. 1:11:25high High variance very important okay I
  1689. 1:11:28will make each and everyone understand
  1690. 1:11:31one by one if it is performing well with
  1691. 1:11:33the training data that is basically low
  1692. 1:11:35bias and whenever it performs well with
  1693. 1:11:38the test sorry fails to perform well
  1694. 1:11:40with the fails to perform well with the
  1695. 1:11:42test data then it is basically High
  1696. 1:11:44variance now similarly I may have
  1697. 1:11:46another scenario which is called as
  1698. 1:11:48underfitting so let's say that I have
  1699. 1:11:50something called as
  1700. 1:11:51underfitting now in this underfitting
  1701. 1:11:53what is the scenario the
  1702. 1:11:56model fails to perform it gives bad
  1703. 1:12:00accuracy I say that model always
  1704. 1:12:03remember whenever I talk about bias then
  1705. 1:12:05you can understand that it is something
  1706. 1:12:07related to the training data whenever I
  1707. 1:12:10talk about test data at that point of
  1708. 1:12:12time you talk about variance and that
  1709. 1:12:15specifically whenever you talk about
  1710. 1:12:17variance that basically means we are
  1711. 1:12:18talking about the test data so for an
  1712. 1:12:21overfitting you will basically have low
  1713. 1:12:23bias and high variance low bias with
  1714. 1:12:26respect to the training data and high
  1715. 1:12:29variance with respect to the test data
  1716. 1:12:31now if the model accuracy is bad with
  1717. 1:12:36training data and the model accuracy is
  1718. 1:12:39also bad with test data in this scenario
  1719. 1:12:44we basically say it as underfitting so
  1720. 1:12:47these are the two conditions that are
  1721. 1:12:50with respect to underfitting that
  1722. 1:12:51basically means that both for the
  1723. 1:12:54training data also the model is giving
  1724. 1:12:55bad accuracy and again for the test data
  1725. 1:12:59also it is basically having a bad
  1726. 1:13:01accuracy so in this particular scenario
  1727. 1:13:03we can definitely say two things out of
  1728. 1:13:05underfitting one is high bias and high
  1729. 1:13:10variance so this is the condition with
  1730. 1:13:12respect to underfitting very super
  1731. 1:13:15important let me just explain you once
  1732. 1:13:17again suppose let's consider I have one
  1733. 1:13:21model I have model two this is model one
  1734. 1:13:24this is model one this is model two and
  1735. 1:13:27this is model 3 okay guys so suppose
  1736. 1:13:30let's say that I have my model my
  1737. 1:13:33training accuracy is let's say
  1738. 1:13:3690% And my let's say that my test
  1739. 1:13:39accuracy is 80% now in this particular
  1740. 1:13:42case let's say that my training accuracy
  1741. 1:13:44is
  1742. 1:13:4692% and my test accuracy is 91% and
  1743. 1:13:51let's say my model three is basically
  1744. 1:13:53having training accuracy as
  1745. 1:13:5670% and my test accuracy is 65% so if I
  1746. 1:14:01take this particular case it is
  1747. 1:14:03basically overfitting if I take this
  1748. 1:14:06particular thing this basically becomes
  1749. 1:14:08my generalized model and when I talk
  1750. 1:14:11about this this is my I'll just say that
  1751. 1:14:15okay I'll also put nice color so that uh
  1752. 1:14:17you'll be able to understand this this
  1753. 1:14:19becomes our generalized model and this
  1754. 1:14:22finally becomes our underfitting right
  1755. 1:14:24under under fitting so here is my red
  1756. 1:14:27color I will just say it as underfitting
  1757. 1:14:29what are the main properties of this
  1758. 1:14:31overfitting as I said in this scenario
  1759. 1:14:34since it is performing well with the
  1760. 1:14:36training data so it will be low bias
  1761. 1:14:38High variance in this particular case it
  1762. 1:14:41will be low bias low variance and this
  1763. 1:14:44particular case it will be high bias and
  1764. 1:14:47high variance understand in this
  1765. 1:14:49terminology in this particular way
  1766. 1:14:51you'll be able to understand so why do
  1767. 1:14:53we require always a generalized model
  1768. 1:14:55because whenever our new data will
  1769. 1:14:57definitely come generalized model will
  1770. 1:14:59be able to give us very good output
  1771. 1:15:01let's go back to this particular example
  1772. 1:15:03here you'll be able to see this straight
  1773. 1:15:05line the red line that I have actually
  1774. 1:15:07created is basically overfitting so that
  1775. 1:15:10whenever I probably get the new points
  1776. 1:15:12which is having this real value and the
  1777. 1:15:14predicted points here you'll be able to
  1778. 1:15:16see the difference is quite huge so
  1779. 1:15:18because of this it will definitely be a
  1780. 1:15:20scenario of overfitting where it has low
  1781. 1:15:24bias and high weight
  1782. 1:15:25so again let me go ahead and take this
  1783. 1:15:28example so this was my line which I have
  1784. 1:15:30actually drawn I had two points and when
  1785. 1:15:33I draw this line which was a best fit
  1786. 1:15:36line to which is passing through both
  1787. 1:15:37the points this scenario is basically
  1788. 1:15:40causing a overfitting problem and I've
  1789. 1:15:42also shown you my J of theta 1 will be
  1790. 1:15:45zero in this scenario since it is
  1791. 1:15:47passing exactly and the predicted point
  1792. 1:15:49is also over there now understand one
  1793. 1:15:52thing is that what can can we take out
  1794. 1:15:55from this what assumptions we can take
  1795. 1:15:57out from this definitely if I talk about
  1796. 1:16:00our cost function our cost function here
  1797. 1:16:02is nothing but 1X 2 m summation of I = 1
  1798. 1:16:062 m h Theta of X of i - y of I whole s
  1799. 1:16:13now let's consider that I am going to
  1800. 1:16:15use this H Theta X and I'm going to
  1801. 1:16:17basically write it as y hat okay let's
  1802. 1:16:19focus on this specific point so when I
  1803. 1:16:22take this I'm I'm just going to focus on
  1804. 1:16:24this particular point so here I will
  1805. 1:16:26definitely write it as y hat minus y of
  1806. 1:16:30I whole squ so this is my y y hat of I
  1807. 1:16:35minus y hat y i whole Square so this is
  1808. 1:16:38nothing but the difference between the
  1809. 1:16:40predicted value and the real value okay
  1810. 1:16:42this is what I'm actually trying to get
  1811. 1:16:44now in this scenario if I am adding this
  1812. 1:16:47values obviously I'm going to get the
  1813. 1:16:48value as zero now I have to make sure
  1814. 1:16:52that this value does not come to zero
  1815. 1:16:53because this is still over fitting so
  1816. 1:16:57that is where your Ridge regression will
  1817. 1:16:58come into picture Ridge and lasso will
  1818. 1:17:01come into picture now when I use Ridge
  1819. 1:17:03and lasso suppose if I use Ridge now in
  1820. 1:17:06Ridge what we say this this is also
  1821. 1:17:09called as L2
  1822. 1:17:11regularization now L2 regularization
  1823. 1:17:14what it does is that it basically adds a
  1824. 1:17:17unique
  1825. 1:17:18parameter add a One More Sample value
  1826. 1:17:21which is like Lambda multiplied by slope
  1827. 1:17:25Square now what is this slope whatever
  1828. 1:17:28slope of this particular line it is we
  1829. 1:17:30are just going to square it off now
  1830. 1:17:33suppose if I take my equation which
  1831. 1:17:34looks like this H Theta of X is equal to
  1832. 1:17:39Theta 0 + Theta 1 x now in this
  1833. 1:17:41particular case my Theta 0 was zero so
  1834. 1:17:44my H Theta of X is nothing but Theta 1
  1835. 1:17:47what is Theta 1 this is specifically
  1836. 1:17:49called as slope and I am basically
  1837. 1:17:52taking this Theta 1 I'm actually making
  1838. 1:17:54it as a square Square so always
  1839. 1:17:56understand I don't want to make this as
  1840. 1:17:57zero because if it becomes zero it may
  1841. 1:18:00lead to overfitting condition now what
  1842. 1:18:03will happen if I add this particular
  1843. 1:18:05equation if I add this particular
  1844. 1:18:06equation this will obviously come as
  1845. 1:18:08zero let's consider my Lambda value over
  1846. 1:18:12here my Lambda value is one I'll talk
  1847. 1:18:15about how do you set up Lambda value
  1848. 1:18:17okay let's consider that I'm
  1849. 1:18:18initializing it to one let's say my
  1850. 1:18:21Lambda value is 1 now what I will do is
  1851. 1:18:24that this l Lambda value is 1 Let's
  1852. 1:18:26consider our slope value initially is
  1853. 1:18:28two and because of this two I got this
  1854. 1:18:30best fit line I'm just going to consider
  1855. 1:18:32it so if I do the total sum over here if
  1856. 1:18:35I'm just considering this this value is
  1857. 1:18:37three now the cost function will not
  1858. 1:18:40stop over here because still it has to
  1859. 1:18:42minimize it has to reduce this three
  1860. 1:18:45value so what it will do it will again
  1861. 1:18:47change the Theta 1 value and let's say
  1862. 1:18:49that my Theta van value has changed now
  1863. 1:18:52it got another best fit line which looks
  1864. 1:18:54something like like this this is my next
  1865. 1:18:56best fit line I'll talk about Lambda
  1866. 1:18:57Lambda is a hyper parameter guys what
  1867. 1:19:00exactly is Lambda I'll just talk about
  1868. 1:19:01it now when I basically change this line
  1869. 1:19:04now see why I'm getting this line let's
  1870. 1:19:06consider I have changed my Theta 1 value
  1871. 1:19:08since we need to minimize now when we
  1872. 1:19:11need to minimize what it will do we'll
  1873. 1:19:12again calculate the slope of this
  1874. 1:19:14particular line and then we will try to
  1875. 1:19:16create a new line when we sorry it is
  1876. 1:19:18two two not three just a second guys 0 +
  1877. 1:19:241 multiplied by 2 s which is nothing but
  1878. 1:19:274 so now my cost function will not stop
  1879. 1:19:31over here so we are going to still
  1880. 1:19:33reduce this now in order to reduce this
  1881. 1:19:36again Theta 1 value will get changed and
  1882. 1:19:39then we will get a next best fit line
  1883. 1:19:40for this point now what will happen in
  1884. 1:19:43this scenario once we have this best fit
  1885. 1:19:45line we will definitely get a kind of
  1886. 1:19:47small difference so now if I go ahead
  1887. 1:19:50and consider the new equation my y hat I
  1888. 1:19:54minus y
  1889. 1:19:55i² + Lambda of slope squar this value
  1890. 1:20:00will be a small value now because I have
  1891. 1:20:03some difference and then plus again 1
  1892. 1:20:06multiplied by now understand whether the
  1893. 1:20:09slope will increase in this particular
  1894. 1:20:11case or whether it will decrease in this
  1895. 1:20:13particular case there will be some slope
  1896. 1:20:15value let's say that I have got some
  1897. 1:20:17slope of this particular line in this
  1898. 1:20:19particular scenario again your slope
  1899. 1:20:21will definitely decrease so let's say in
  1900. 1:20:23the case of two initially it was now it
  1901. 1:20:25is basically
  1902. 1:20:271.36 whole squ now this small Value
  1903. 1:20:32Plus 1 + 1.3 squ or let me consider that
  1904. 1:20:37my slope is now one simple value that is
  1905. 1:20:405 so if I get this it is 2.25 2.25 plus
  1906. 1:20:44small value it will be less than three
  1907. 1:20:46only right it will obviously be less
  1908. 1:20:48than three or equal to 3 but understand
  1909. 1:20:50what is happening the value is getting
  1910. 1:20:52reduced from 4 to 3 so this is is the
  1911. 1:20:55importance of Ridge now what will happen
  1912. 1:20:57is that you will try to get a
  1913. 1:20:59generalized model which has low bias and
  1914. 1:21:02low variance instead of this overfitting
  1915. 1:21:05condition you know why specifically we
  1916. 1:21:08are adding Ridge L2 regularization it is
  1917. 1:21:11basically to prevent
  1918. 1:21:14overfitting because here you are not
  1919. 1:21:16stopping here you are trying to reduce
  1920. 1:21:18it unless and until you get a line you
  1921. 1:21:21get a line which will be able to handle
  1922. 1:21:24the which will be able to handle as a uh
  1923. 1:21:27generalized model now here you can see
  1924. 1:21:29now if I have my new points like how I
  1925. 1:21:31drew over here now the distance will be
  1926. 1:21:33less so now you'll be able to see that
  1927. 1:21:36it will be able to create a generalized
  1928. 1:21:38model guys this will be a small value
  1929. 1:21:40only see initially when we have this
  1930. 1:21:42line obviously we have zero if we try to
  1931. 1:21:45slightly move here and there so here
  1932. 1:21:48you'll be able to see that it will just
  1933. 1:21:50a slight movement but what this movement
  1934. 1:21:52is basically specifying it is specifying
  1935. 1:21:55that the slope should not be steep if we
  1936. 1:21:59probably have a steep slope it obviously
  1937. 1:22:02leads to most of the time overfitting
  1938. 1:22:04condition it should not be steep it
  1939. 1:22:06should be very very it should be less
  1940. 1:22:08steeper but it should actually help you
  1941. 1:22:10to create a generalized model so you
  1942. 1:22:13will be seeing that after playing for
  1943. 1:22:14some amount of time this value will not
  1944. 1:22:18reduce after some point of time it'll
  1945. 1:22:19get almost it'll be a minimal value
  1946. 1:22:22it'll be a smaller value and for this
  1947. 1:22:24also you have to specify iterations how
  1948. 1:22:26many times you probably have to train
  1949. 1:22:29them now this iterations is also a
  1950. 1:22:33hyperparameter based on number of
  1951. 1:22:35iterations you will probably see your R
  1952. 1:22:38square or adjusted R square over here so
  1953. 1:22:41this iterations based on the number of
  1954. 1:22:43iterations it will never become zero
  1955. 1:22:45guys understand because zero it is not
  1956. 1:22:48possible if it becomes zero trust me it
  1957. 1:22:50is an overfitting model you cannot get
  1958. 1:22:52that is something zero now what is
  1959. 1:22:55Lambda coming to this Lambda this Lambda
  1960. 1:22:57is a
  1961. 1:22:58hyperparameter this is basically to
  1962. 1:23:01check how fast you want to lessen the
  1963. 1:23:04steepness or how fast you want to make a
  1964. 1:23:06steepness grow higher right and this
  1965. 1:23:08Lambda will also be selected by using
  1966. 1:23:11hyper parameter and this also I'll show
  1967. 1:23:13you today in Practical what do you mean
  1968. 1:23:15by iterations iteration basically means
  1969. 1:23:17how many time I want to change the Theta
  1970. 1:23:181 value how many times you want to
  1971. 1:23:20change the Theta value that is the
  1972. 1:23:22convergence algorithm right
  1973. 1:23:25convergence algorithm over here L2
  1974. 1:23:27regularization or Ridge is basically
  1975. 1:23:30used in such a way that you should never
  1976. 1:23:33overfit why we assume Theta 0 is equal
  1977. 1:23:35to 0 because I'm considering that it
  1978. 1:23:37passes through a origin right origin
  1979. 1:23:40over here Lambda is a hyper
  1980. 1:23:43parameter steep basically means how
  1981. 1:23:46steep the line is if I have this line
  1982. 1:23:49this line is quite steep if I have this
  1983. 1:23:51line This is less steep now if I go to
  1984. 1:23:54the next regularization which is called
  1985. 1:23:56as lasso raso R lasso regression this is
  1986. 1:24:00also called as L1
  1987. 1:24:02regularization now here the formula will
  1988. 1:24:05be changing little bit here you will be
  1989. 1:24:07having y hat of minus of Y whole Square
  1990. 1:24:11here you'll be adding a parameter Lambda
  1991. 1:24:14but understand here you'll not be adding
  1992. 1:24:16slope Square no here you'll be adding
  1993. 1:24:20mode of slope here you'll be adding mode
  1994. 1:24:23of slope and this mode of slope will
  1995. 1:24:26work is that it will actually help you
  1996. 1:24:29to do feature selection now you may be
  1997. 1:24:31thinking how feature selection crash
  1998. 1:24:33let's consider a equation over here
  1999. 1:24:35let's say that I have many many features
  2000. 1:24:37I have many many many features okay so
  2001. 1:24:40my H Theta of X which I'm indicating
  2002. 1:24:42here as y hat let's say that I'm I'm
  2003. 1:24:45writing this equation apart from
  2004. 1:24:47preventing for overfitting it will also
  2005. 1:24:49help you to do feature selection here
  2006. 1:24:51let me just show you over here with an
  2007. 1:24:53example this H Theta of X which I'm
  2008. 1:24:56probably writing as y hat will basically
  2009. 1:24:59be indicated by something over here
  2010. 1:25:01you'll be able to see that it is nothing
  2011. 1:25:03but let's say that I have multiple
  2012. 1:25:05features like this now in this
  2013. 1:25:07particular features obviously there are
  2014. 1:25:09so many coefficients over here so many
  2015. 1:25:11slopes over here now mod of slope will
  2016. 1:25:13be what it will be nothing but mod of X1
  2017. 1:25:16plus X2 plus X3 plus X4 plus X5 like
  2018. 1:25:20this up to xn now in this particular
  2019. 1:25:23case how it is basically helping you to
  2020. 1:25:25sorry not X1 sorry just a second this
  2021. 1:25:29mod of I have taken the data point this
  2022. 1:25:31is not data points this should be your
  2023. 1:25:34mod of theta 0 + Theta 1 + Theta 2 +
  2024. 1:25:38theta 3 + Theta 4 + Theta 5 like this up
  2025. 1:25:42to Theta n so here you'll be able to see
  2026. 1:25:45that this is how I will basically uh
  2027. 1:25:47I'll basically be calculating the slope
  2028. 1:25:50now as we go ahead guys whichever
  2029. 1:25:52features are probably not playing an
  2030. 1:25:55amazing role the Theta value the
  2031. 1:25:57coefficient value the slope value will
  2032. 1:25:59be very very small it is just like that
  2033. 1:26:01entire feature is neglected that entire
  2034. 1:26:04feature is neglected now in this
  2035. 1:26:06particular case we were doing squaring
  2036. 1:26:08because of the squaring that value was
  2037. 1:26:10also increasing but here because of the
  2038. 1:26:11mode that value will not increase
  2039. 1:26:14instead it will be a condition wherein
  2040. 1:26:16we are basically neglecting those
  2041. 1:26:18features that are not at all important
  2042. 1:26:21in this specific problem statement so
  2043. 1:26:23with the help of L1 regularization that
  2044. 1:26:26is lasso you are able to do two
  2045. 1:26:28important things one is preventing
  2046. 1:26:31overfitting and the second case is that
  2047. 1:26:33if you have many features and many of
  2048. 1:26:36the features are not that important okay
  2049. 1:26:39in basically finding out your slope or
  2050. 1:26:42your line or the best fit line in that
  2051. 1:26:44particular case it will also help you to
  2052. 1:26:46perform feature selection so this is the
  2053. 1:26:48importance of the entire what is the
  2054. 1:26:51importance of this this is the
  2055. 1:26:52importance of the uh Ridge and the lasso
  2056. 1:26:56regression that we are doing here I'm
  2057. 1:26:57just going to write L1
  2058. 1:26:59regularization and obviously we have
  2059. 1:27:01discussed about L2 regularization also
  2060. 1:27:04now you have probably understood Lambda
  2061. 1:27:06is one hyperparameter okay which we will
  2062. 1:27:09specifically using okay and based on
  2063. 1:27:12this Lambda this will be found out
  2064. 1:27:13through cross
  2065. 1:27:14validation cross validation is a
  2066. 1:27:16technique wherein we will try to
  2067. 1:27:19probably train our model and try to find
  2068. 1:27:21out the specific things okay what should
  2069. 1:27:24be the exact value and there also we
  2070. 1:27:26play with multiple values in short what
  2071. 1:27:28we are doing we just trying to reduce
  2072. 1:27:29the cost function in such a way that uh
  2073. 1:27:32it will definitely never become zero but
  2074. 1:27:34it will basically reduce based on the
  2075. 1:27:37Lambda and the slope value in most of
  2076. 1:27:39the scenario if you ask me we should
  2077. 1:27:41definitely try both the regularization
  2078. 1:27:44and see that wherever the performance
  2079. 1:27:46Matrix is good we should use that what
  2080. 1:27:48is cross validation basically means I
  2081. 1:27:50will try to use different different
  2082. 1:27:52Lambda value and basically Ally use it
  2083. 1:27:55so in a short let me write it down again
  2084. 1:27:58for Ridge regression which is an L2 Norm
  2085. 1:28:03here I'm simply writing my cost function
  2086. 1:28:05in this particular case will be little
  2087. 1:28:08bit different here I can definitely
  2088. 1:28:10write my cost function as H Theta X of i
  2089. 1:28:14- y of I S Plus Lambda multiplied slope
  2090. 1:28:21Square what is the purpose of this the
  2091. 1:28:24purpose is very simple here we are
  2092. 1:28:27preventing overfitting this was with
  2093. 1:28:29respect to the Ridge Recreation that is
  2094. 1:28:31L2 nor now if I go ahead and discuss
  2095. 1:28:33about the next one which is called as
  2096. 1:28:34lasso regression which is also called as
  2097. 1:28:38L1 regularization in the case of lasso
  2098. 1:28:41regression your cost function will be H
  2099. 1:28:44Theta of X of
  2100. 1:28:47IUS y of
  2101. 1:28:49i² plus Lambda ultied mode of flow so
  2102. 1:28:55here you have this specific thing and
  2103. 1:28:57what is the purpose the purpose are two
  2104. 1:29:00one is prevent overfitting and the
  2105. 1:29:04second one is something called as
  2106. 1:29:05feature selection so these two are the
  2107. 1:29:08outcomes of the entire thing see with
  2108. 1:29:10respect to this lasso right you have
  2109. 1:29:12slopes slopes here you'll be having
  2110. 1:29:15Theta 0 plus Theta 1 plus Theta 2 plus
  2111. 1:29:17theta 3 like this up to Theta n now when
  2112. 1:29:20you'll have this many number of thetas
  2113. 1:29:22when you have many number of features
  2114. 1:29:24and when you have many number of
  2115. 1:29:25features that basically means you'll
  2116. 1:29:26have multiple slopes right those
  2117. 1:29:28features that are not performing well or
  2118. 1:29:30that has no contribution in finding out
  2119. 1:29:32your output that coefficient value will
  2120. 1:29:35be almost nil right it will be very much
  2121. 1:29:37near to zero in short you neglecting
  2122. 1:29:40that value by using modulus you're not
  2123. 1:29:42squaring them up you're not increasing
  2124. 1:29:44those values now I will continue and uh
  2125. 1:29:47probably I will also discuss about the
  2126. 1:29:49assumptions of linear regressions so
  2127. 1:29:52what are the assumptions of linear
  2128. 1:29:54regression in this particular scenario
  2129. 1:29:56so assumption is that number one point
  2130. 1:30:00linear regression if our features are in
  2131. 1:30:03normal or gion
  2132. 1:30:06distribution if our features follows
  2133. 1:30:09this particular distribution it is
  2134. 1:30:11obviously good our model will get
  2135. 1:30:13trained well so there is one concept
  2136. 1:30:16which is called as feature
  2137. 1:30:18transformation now in future
  2138. 1:30:20transformation always understand what
  2139. 1:30:22will happen if a model does not fall
  2140. 1:30:24follow a gan distribution then we apply
  2141. 1:30:26some kind of mathematical equation onto
  2142. 1:30:28the data and try to convert them into
  2143. 1:30:30normal orian distribution the second
  2144. 1:30:33assumption that I would definitely like
  2145. 1:30:34to make is that standard scalar or
  2146. 1:30:37standard digestion standard dig is
  2147. 1:30:40nothing but it is a kind of scaling your
  2148. 1:30:43data by using Z score I hope everybody
  2149. 1:30:46remembers Z score this is what we
  2150. 1:30:48basically apply there your mean is equal
  2151. 1:30:50to zero and standard deviation equal to
  2152. 1:30:521 see guys wherever you have gradient
  2153. 1:30:54descent involved it is good to basically
  2154. 1:30:57do
  2155. 1:30:58standardization because if our initial
  2156. 1:31:01point is a small Point somewhere here
  2157. 1:31:03then to reach the global Minima or
  2158. 1:31:05training will happen quickly otherwise
  2159. 1:31:07what will happen if your values are
  2160. 1:31:09quite huge then your graph may be very
  2161. 1:31:11big and the point can come over any over
  2162. 1:31:13there and the third point is that this
  2163. 1:31:16linear regression works with respect to
  2164. 1:31:19linearity it works if your data is
  2165. 1:31:22linearly separable
  2166. 1:31:24I'll not say linearly separable but this
  2167. 1:31:26linearity will come into picture if your
  2168. 1:31:28data is too much linear it will
  2169. 1:31:30obviously be able to give a very good
  2170. 1:31:31answer like logistic regression also
  2171. 1:31:34which we are going to discuss today this
  2172. 1:31:35also has the same property now you may
  2173. 1:31:38be asking is it compulsory to do
  2174. 1:31:40standardization guys if you want to
  2175. 1:31:42increase the training time of your model
  2176. 1:31:45or if you want to optimize your model I
  2177. 1:31:47would suggest go ahead and do
  2178. 1:31:48standardization now coming to the fourth
  2179. 1:31:50Point here you really need to check
  2180. 1:31:52about multicolinearity
  2181. 1:31:54this is also one kind of check we
  2182. 1:31:56basically do what is multicol
  2183. 1:31:58linearities let's say I have X1 I have
  2184. 1:32:00X2 and this is my output feature I have
  2185. 1:32:03let's say X3 also now let's say that if
  2186. 1:32:05I try to see the colinearity of this two
  2187. 1:32:08feature how how correlated these two
  2188. 1:32:10feature are let's say that these two
  2189. 1:32:12feature are 95% correlated is it is it a
  2190. 1:32:16wise decision to use both the features
  2191. 1:32:18and let's say that let's let's say that
  2192. 1:32:20these two features are 95% correlated
  2193. 1:32:23but it is highly correlated with Y is it
  2194. 1:32:25necessary that we should use both the
  2195. 1:32:27feature in this particular scenario the
  2196. 1:32:29answer should be no we can drop this
  2197. 1:32:32particular feature okay we can drop this
  2198. 1:32:34particular feature any one of the
  2199. 1:32:36feature we can definitely drop it and
  2200. 1:32:38based on that I can just use one single
  2201. 1:32:40feature and basically we do the
  2202. 1:32:42prediction there is also a concept which
  2203. 1:32:44is called as variation inflation factor
  2204. 1:32:46I will try to make a dedicated video
  2205. 1:32:48about this multical is also solved with
  2206. 1:32:51the help of variation inflation Factor
  2207. 1:32:53one more term is there homos orc so that
  2208. 1:32:56kind of terminologies also we use one
  2209. 1:32:58more condition in this but if you almost
  2210. 1:33:00satisfied with this assumptions you will
  2211. 1:33:02definitely be able to outperform in
  2212. 1:33:03linear regression so you have got an
  2213. 1:33:06idea of the assumptions you have also
  2214. 1:33:07got an idea of multiple things okay now
  2215. 1:33:10let's go towards something called as
  2216. 1:33:12logistic regression now logistic
  2217. 1:33:14regression what logistic regression is
  2218. 1:33:16the first type of algorithm that we are
  2219. 1:33:18going to learn in classification let's
  2220. 1:33:20say that in classification I have one
  2221. 1:33:22example you know so suppose I have say
  2222. 1:33:24number of hours study hours and number
  2223. 1:33:28of play hours based on this I want to
  2224. 1:33:31predict whether a child is passing or
  2225. 1:33:33failing suppose these two are my
  2226. 1:33:35features I want to predict whether it is
  2227. 1:33:36pass or fail so here you'll be able to
  2228. 1:33:39see that I have some fixed number of
  2229. 1:33:40categories specifically in this
  2230. 1:33:42particular scenario I have two
  2231. 1:33:43categories binary logistic regression
  2232. 1:33:46works very well with binary
  2233. 1:33:48classification now the uh question comes
  2234. 1:33:50that can we solve multiclass
  2235. 1:33:52classification using logistic the answer
  2236. 1:33:54is simply yes you can definitely do it
  2237. 1:33:57so let's go ahead and let's try to
  2238. 1:33:58discuss about uh logistic regression now
  2239. 1:34:02what is the main purpose of the logistic
  2240. 1:34:03regression first of all let's let's uh
  2241. 1:34:06understand one scenario okay suppose I
  2242. 1:34:08have a feature which basically says um
  2243. 1:34:13number of study hours and this is like 1
  2244. 1:34:172 3 4 5 6 7 and let's say that I have
  2245. 1:34:24pass this point is basically pass and
  2246. 1:34:28this point is basically
  2247. 1:34:29fail so I have this two conditions these
  2248. 1:34:32are my outcomes now what I'll do I will
  2249. 1:34:34just try to make some data points let's
  2250. 1:34:36say that if I study Less Than 3 hours I
  2251. 1:34:40will probably be fail if I study more
  2252. 1:34:43than 3 hours then probably I will pass
  2253. 1:34:47this I'll make it as fail and this I
  2254. 1:34:49will make it as pass so I will be having
  2255. 1:34:52points over here this 1 2 3 let's say
  2256. 1:34:57that this is my training data set now
  2257. 1:35:00the first question says that okay Chris
  2258. 1:35:02fine you have some data over here
  2259. 1:35:05whenever it is less than three you are
  2260. 1:35:06basically the person is failing if it is
  2261. 1:35:08greater than five greater than three it
  2262. 1:35:12is basically showing data points points
  2263. 1:35:14with respect to pass now can't we solve
  2264. 1:35:17this problem first with linear
  2265. 1:35:18regression now with the help of linear
  2266. 1:35:21regression here the first point will be
  2267. 1:35:23that yes I can definitely draw a best
  2268. 1:35:26fit line my best fit line in this
  2269. 1:35:28particular scenario may be something
  2270. 1:35:30like this it may it may look something
  2271. 1:35:32like this so here fail is nothing but
  2272. 1:35:35zero pass is one the middle point is
  2273. 1:35:38basically 0.5 so obviously with the help
  2274. 1:35:41of linear
  2275. 1:35:42regression I'm able to create this best
  2276. 1:35:44fit line and I'll put a scenario that
  2277. 1:35:47whenever the value is less
  2278. 1:35:50than5 whenever the value is less than
  2279. 1:35:520.5 whenever the output is less than5
  2280. 1:35:55let's say that new data point is this
  2281. 1:35:57and based on this I'll try to do the
  2282. 1:35:58prediction I'm actually able to get the
  2283. 1:36:00output over here now when I'm getting
  2284. 1:36:02the output over here this basically is
  2285. 1:36:040.25 now in this particular scenario
  2286. 1:36:06obviously I'm able to say that yes the
  2287. 1:36:08person I'll write a condition over here
  2288. 1:36:11saying that if my H Theta of x value is
  2289. 1:36:15less than 0.5 then my output should be
  2290. 1:36:20zero let's say less than 0.5 I'll say
  2291. 1:36:22not less than or equal to less than5
  2292. 1:36:25then my output will be zero right so in
  2293. 1:36:28this particular case Zero basically
  2294. 1:36:29means fail similarly I'll have a
  2295. 1:36:32scenario where I'll say that when if my
  2296. 1:36:35S of theta of X is greater than or equal
  2297. 1:36:36to 5 then this will basically be one
  2298. 1:36:39which is nothing but pass so this two
  2299. 1:36:41condition I can definitely write over
  2300. 1:36:42here this is my center point so that any
  2301. 1:36:45point that will probably come over here
  2302. 1:36:47let's say that this point is coming over
  2303. 1:36:49here right let's say new data point is
  2304. 1:36:51somewhere coming over here with this red
  2305. 1:36:53point
  2306. 1:36:54now what I'll do I'll basically draw a
  2307. 1:36:56straight line it will come over here I
  2308. 1:36:58will just extend this line
  2309. 1:37:00long I will extend this line over here
  2310. 1:37:04and I will extend this line over here
  2311. 1:37:07and here you can see that based on this
  2312. 1:37:09I'm actually getting this particular
  2313. 1:37:11prediction which is greater than 0.5 so
  2314. 1:37:13I will say that okay the person has
  2315. 1:37:15passed obviously this is fine this is
  2316. 1:37:18obviously working better this is
  2317. 1:37:20obviously working better so what what is
  2318. 1:37:22the problem why we are not using linear
  2319. 1:37:24regression okay in order to solve this
  2320. 1:37:26particular problem why you are
  2321. 1:37:27specifically having logistic regression
  2322. 1:37:29the answer is very much simple guys the
  2323. 1:37:31answer is that whenever let's say that
  2324. 1:37:34if I have an outlier which looks
  2325. 1:37:35something like this suppose I have an
  2326. 1:37:37outlier which comes like this over here
  2327. 1:37:40what is this value let's say that this
  2328. 1:37:41value is nothing but 7 8 9 10 let's say
  2329. 1:37:46that the number of study hours and I'm
  2330. 1:37:48studying for nine it is obviously pass
  2331. 1:37:51now in this particular scenario when I
  2332. 1:37:52have an outlier this entire line will
  2333. 1:37:54change now I will probably get my line
  2334. 1:37:57which looks something like this okay my
  2335. 1:37:59line will basically move something like
  2336. 1:38:01this it will now get moved something
  2337. 1:38:03like this now when it gets moves
  2338. 1:38:04completely like this now for even five
  2339. 1:38:08or even at any point that I am actually
  2340. 1:38:10predicting let's say that at this
  2341. 1:38:11particular point if I try to find out
  2342. 1:38:14it'll be showing less than 0. five so
  2343. 1:38:16here this particular value or answer
  2344. 1:38:19will be wrong right because if we are
  2345. 1:38:21studying more than 5 hours OB viously B
  2346. 1:38:24based on the previous line the person
  2347. 1:38:26had to pass but in this scenario it is
  2348. 1:38:28failing it is coming less than 0.5 but
  2349. 1:38:31the real value for this is basically
  2350. 1:38:33passed so I hope you are understanding
  2351. 1:38:36because of the outlier the entire line
  2352. 1:38:37is getting changed so how do we fix this
  2353. 1:38:40particular problem now in this two
  2354. 1:38:42scenarios are there first of all
  2355. 1:38:44obviously because of just an outlier
  2356. 1:38:46your entire line is getting shifted here
  2357. 1:38:48and there the second point is that over
  2358. 1:38:50here sometimes you're also getting
  2359. 1:38:52greater than one you you're also getting
  2360. 1:38:54less than one suppose if I try to
  2361. 1:38:55calculate for this particular point if I
  2362. 1:38:58project it in behind I'll be getting
  2363. 1:39:00some negative value so we have to squash
  2364. 1:39:02this function if I squash this function
  2365. 1:39:04then it'll become a plain line right how
  2366. 1:39:07do we squash it and for this we use
  2367. 1:39:09something called as sigmoid activation
  2368. 1:39:11function or sigmoid function if somebody
  2369. 1:39:14ask you why don't you use linear
  2370. 1:39:17regession in order to solve this
  2371. 1:39:19classification problem then your answer
  2372. 1:39:21should be very much simple you should
  2373. 1:39:23say this to specific points so we will
  2374. 1:39:26try to go ahead and solve some linear
  2375. 1:39:27regression now with the help of cost
  2376. 1:39:29function everything as such and we'll
  2377. 1:39:31try to understand how the cost function
  2378. 1:39:34will look for logistic regression second
  2379. 1:39:36reason I told you right it is greater
  2380. 1:39:38than zero over here the line is going
  2381. 1:39:40greater than zero right greater than
  2382. 1:39:42zero I have only Z and one and it is
  2383. 1:39:45becoming greater than zero but I have
  2384. 1:39:47already told that our maximum and
  2385. 1:39:49minimum value are 1 and zero so I hope
  2386. 1:39:51you have understood why linear Reg
  2387. 1:39:53cannot be used okay I showed you all the
  2388. 1:39:56scenarios why linear regression should
  2389. 1:39:58not be used now we'll continue and
  2390. 1:40:00probably discuss about the other things
  2391. 1:40:02over here and uh we will now try to
  2392. 1:40:05understand fine what exactly logistic
  2393. 1:40:07regression is all about and how the
  2394. 1:40:09decision boundaries basically created
  2395. 1:40:11now we'll go ahead and discuss about
  2396. 1:40:12that specific thing so let's go ahead
  2397. 1:40:15our values should be always between 0 to
  2398. 1:40:17one over here in this particular case
  2399. 1:40:19because it is a binary classification
  2400. 1:40:21problem only this should be the answer
  2401. 1:40:23so let's go ahead and let's define our
  2402. 1:40:25decision boundary so my decision
  2403. 1:40:26boundary decision boundary in the case
  2404. 1:40:29of logistic regression first of all as
  2405. 1:40:31usual in logistic regression we defined
  2406. 1:40:34our hypothesis okay guys first of all
  2407. 1:40:36let's see if I'm writing my my h of
  2408. 1:40:40theta my H Theta of X as Theta 0 + Theta
  2409. 1:40:451 into x + Theta 2 into X like this X1
  2410. 1:40:49X2 + Theta n into xn
  2411. 1:40:53now in this scenario can I write this
  2412. 1:40:55entire equation as Theta transpose X
  2413. 1:40:59obviously I can definitely write this
  2414. 1:41:01way right and this is what is the
  2415. 1:41:02notation that you will probably seeing
  2416. 1:41:04in many places so with respect to the
  2417. 1:41:06decision boundary of logistic regression
  2418. 1:41:10our Theta see like this we can write I'm
  2419. 1:41:12saying okay but since we have to
  2420. 1:41:14consider two things one is squashing the
  2421. 1:41:17line okay how that squashing will
  2422. 1:41:19basically happen see if I have this if I
  2423. 1:41:22have this line
  2424. 1:41:24we saw in the above right if I have this
  2425. 1:41:26line suppose I have some data points
  2426. 1:41:28over here and I have some data points
  2427. 1:41:30over here if I want to create the best
  2428. 1:41:32fit line how will I create I will
  2429. 1:41:33basically create like this but I have to
  2430. 1:41:35also do two things one is squash over
  2431. 1:41:37here and squash over here right squash
  2432. 1:41:40over here and squash over here now in
  2433. 1:41:42order to squash I'm saying squash squash
  2434. 1:41:46means
  2435. 1:41:48okay now in order to do this I use a
  2436. 1:41:51function which is called as sigmoid
  2437. 1:41:52activation function
  2438. 1:41:54that basically means what happens
  2439. 1:41:56obviously you know this line is
  2440. 1:41:57basically denoted by H Theta of x equal
  2441. 1:42:01to how do you denote this straight line
  2442. 1:42:04let me write it down nicely for you so
  2443. 1:42:06how do you denote this straight line the
  2444. 1:42:08straight line is obviously denoted by
  2445. 1:42:11Theta 0 + Theta 1 * X1 let's say now on
  2446. 1:42:15top of this on top of this I have to
  2447. 1:42:18apply something on top of this value I
  2448. 1:42:21have to apply something so that I can
  2449. 1:42:23make this line straight instead of just
  2450. 1:42:26expanding in this way so my hypothesis
  2451. 1:42:29will basically be now G of G is
  2452. 1:42:32basically a function on Theta 0 and
  2453. 1:42:34Theta 1 * X1 so here I'm trying to
  2454. 1:42:38basically what I'm trying to do I will
  2455. 1:42:40apply a mathematical formula on top of
  2456. 1:42:42this linear regression to squash this
  2457. 1:42:45line now let's go ahead and let's try to
  2458. 1:42:47find out what is this G okay what is
  2459. 1:42:50this G I will say let Z equal to Theta 0
  2460. 1:42:54+ Theta 1 * X I'm just initializing this
  2461. 1:42:58now my H Theta of X is nothing but G of
  2462. 1:43:00Z now we need to understand what is this
  2463. 1:43:03z g of Z and how do we basically specify
  2464. 1:43:06what is the G function so my G function
  2465. 1:43:08is nothing but H Theta of x equal to 1
  2466. 1:43:11by 1 + e ^ of minus Z which in short if
  2467. 1:43:15I try to initialize Zed now it is 1 ^ of
  2468. 1:43:19e ^ of minus Theta 0 + Theta 1 * X so
  2469. 1:43:24this is what is my H Theta of X which is
  2470. 1:43:26my hypothesis and this obviously works
  2471. 1:43:29well because it is being able to squash
  2472. 1:43:32the function so this is basically my
  2473. 1:43:34hypothesis which I am definitely trying
  2474. 1:43:36to use it and this function that you are
  2475. 1:43:39actually able to see is called as
  2476. 1:43:43sigmoid or logistic function now you
  2477. 1:43:47need to understand what does this
  2478. 1:43:48sigmoid function look like in graph in
  2479. 1:43:50graph it looks something like this so
  2480. 1:43:52this this is my Zed value and this is my
  2481. 1:43:56G of Z this is my 05 your sigmoid
  2482. 1:44:00function will have this curve so this is
  2483. 1:44:03your one this is zero your value when
  2484. 1:44:07now from this we can make a lot of
  2485. 1:44:08assumptions what are the assumptions
  2486. 1:44:10that we can basically make your G of Zed
  2487. 1:44:15your G of Zed is greater than or equal
  2488. 1:44:18to
  2489. 1:44:185.5 is obviously greater than or equal
  2490. 1:44:21to 0.5 when your Zed value is greater
  2491. 1:44:24than or equal to zero this is the major
  2492. 1:44:27assumptions that we can basically make
  2493. 1:44:29that is whenever your G of Z is greater
  2494. 1:44:32than your G of Z is greater than or
  2495. 1:44:35equal to 0.5 whenever your Zed is
  2496. 1:44:38greater than or equal to Z so obviously
  2497. 1:44:40whenever your Zed value is greater than
  2498. 1:44:42Z it is greater than 0.5 if your Zed
  2499. 1:44:44value is less than zero what it will
  2500. 1:44:46become it will basically be less than
  2501. 1:44:470.5 so you can write that specific
  2502. 1:44:50condition also you want so this is the
  2503. 1:44:52most important condition
  2504. 1:44:53over here why it is called as logistic
  2505. 1:44:55regression see guys with the help of
  2506. 1:44:56regression you creating this straight
  2507. 1:44:57line and with the help of the concept of
  2508. 1:44:59sigmo you are able to squash it so they
  2509. 1:45:01have probably combined that name and uh
  2510. 1:45:04basically have written in this way will
  2511. 1:45:05squashing of the best fit L line help to
  2512. 1:45:07overcome the outlier issues yes
  2513. 1:45:09obviously it'll be able to help you so
  2514. 1:45:10let's go ahead and let's try to solve
  2515. 1:45:12the problem statement now usually let's
  2516. 1:45:14consider my training set let's consider
  2517. 1:45:17my training set suppose I have some
  2518. 1:45:19training points like this x of 1 comma y
  2519. 1:45:22of 1
  2520. 1:45:24let's say x of 2A y of 2 okay X of 3A y
  2521. 1:45:28of 3 like this I have lot of training
  2522. 1:45:30points and finally X of n comma y of n
  2523. 1:45:33let's say that this is my training data
  2524. 1:45:35so here uh my y y will belong to what
  2525. 1:45:41zero or 1 because I will only have two
  2526. 1:45:43outputs since we are solving a binary
  2527. 1:45:45classification problem here is my
  2528. 1:45:47training set with two outputs and I hope
  2529. 1:45:50everybody knows about J Theta of Z
  2530. 1:45:53it is nothing but 1 + e ^ of minus Z
  2531. 1:45:57here your Z is nothing but Theta 0 +
  2532. 1:45:59Theta 1 * X1 so this is your Theta 0 now
  2533. 1:46:04what we have to do we have to select
  2534. 1:46:06this Theta now in this particular case
  2535. 1:46:08let's consider that my Theta 0 is 0
  2536. 1:46:10because it is passing through the origin
  2537. 1:46:13just for time pass sake suppose my Z is
  2538. 1:46:15Theta 1 into X so now I need to change
  2539. 1:46:19what is my parameter my parameter is
  2540. 1:46:21Theta 1
  2541. 1:46:23I have to change parameter Theta 1 in
  2542. 1:46:25such a way that I get the best fit line
  2543. 1:46:28and along that I apply this sigmoid
  2544. 1:46:30activation function now let's go ahead
  2545. 1:46:33and let's first of all Define our cost
  2546. 1:46:36function because for this we definitely
  2547. 1:46:38require our cost
  2548. 1:46:39function now everything will be same
  2549. 1:46:42obviously you know the cost function of
  2550. 1:46:44linear regression because the first best
  2551. 1:46:47fit line that you are probably creating
  2552. 1:46:48is with the help of linear
  2553. 1:46:50regression now in this particular case
  2554. 1:46:52in the case of linear regression so here
  2555. 1:46:55you can basically write J J of theta 1
  2556. 1:46:57is nothing but 1 by m summation of I = 1
  2557. 1:47:022 m 1X 2 and here you have H Theta of x
  2558. 1:47:08minus y of I I whole Square so this is
  2559. 1:47:13your entire thing of if you remember
  2560. 1:47:15linear regression whatever things we
  2561. 1:47:17have discussed yesterday okay so this is
  2562. 1:47:19the cost function let's consider that
  2563. 1:47:22for linear regression for this is for
  2564. 1:47:24the linear regression now for the
  2565. 1:47:25logistic regression what will happen for
  2566. 1:47:27your logistic regression I will take the
  2567. 1:47:28same cost function H Theta of X now you
  2568. 1:47:31know what is s Theta of X it is nothing
  2569. 1:47:33but 1 + 1 + e ^ of minus Theta 0 + Theta
  2570. 1:47:37sorry Theta 1 multiplied by X right this
  2571. 1:47:40is my with respect to logistic
  2572. 1:47:42regression this is my entire equation
  2573. 1:47:45now similarly I will try to only put
  2574. 1:47:48this H Theta of X let's consider that
  2575. 1:47:51this is my cost function only only my H
  2576. 1:47:53Theta of X is changing in this
  2577. 1:47:55particular case so if I go ahead and
  2578. 1:47:57write my cost function I can basically
  2579. 1:47:59say 1x2 h Theta of X of i - y of
  2580. 1:48:05i² and in this particular scenario what
  2581. 1:48:07is h Theta of X it is nothing but 1 + 1
  2582. 1:48:11+ e ^ minus Theta 1 x so this is what
  2583. 1:48:16this is getting replaced and this is my
  2584. 1:48:18logistic regression cost function I'm
  2585. 1:48:20just considering this cost function part
  2586. 1:48:22this part later on if you replace this
  2587. 1:48:25to this see if I replace this to this
  2588. 1:48:28and if I replace this to this it becomes
  2589. 1:48:30a logistic regression cost function
  2590. 1:48:33intercept I'm considering it as zero
  2591. 1:48:34guys now when I'm replacing this to this
  2592. 1:48:36this to this then it becomes a logistic
  2593. 1:48:39uh regression cost function but there is
  2594. 1:48:41one problem we cannot we cannot use we
  2595. 1:48:45cannot use this cost function there is a
  2596. 1:48:48reason for this because this equation
  2597. 1:48:50that you're seeing 1/ 1 + e^ of minus
  2598. 1:48:54Theta 1 * X this is a non-convex
  2599. 1:48:59function now you may be considering what
  2600. 1:49:01is a non-convex function so let me write
  2601. 1:49:03it down so here this this term this
  2602. 1:49:07terminology right it is a non-convex
  2603. 1:49:09function now what is this non-convex
  2604. 1:49:10function let me show you and let me
  2605. 1:49:12differentiate it with convex function
  2606. 1:49:15okay we'll try to understand what is the
  2607. 1:49:16difference between non-convex function
  2608. 1:49:18and convex function this is related to
  2609. 1:49:21gradient descent very important this is
  2610. 1:49:24related to gradient desent if you
  2611. 1:49:27remember with the help of linear
  2612. 1:49:29regression whatever gradient Dent we are
  2613. 1:49:32actually getting it is a convex function
  2614. 1:49:34like this this is the convex function
  2615. 1:49:38which looks like a parabola curve
  2616. 1:49:40Parabola curve because of this Parabola
  2617. 1:49:42curve whenever we use this linear
  2618. 1:49:44regression cost function specifically
  2619. 1:49:46because here my H Theta of X is what it
  2620. 1:49:48is nothing but Theta 0 + Theta 1 into X
  2621. 1:49:51because of this this equ
  2622. 1:49:53will always give you a parabola curve
  2623. 1:49:56this kind of cost function or convex
  2624. 1:49:59function you can say but here your s
  2625. 1:50:01Theta of X is changing so in the case of
  2626. 1:50:03if I use that cost function you will be
  2627. 1:50:05getting some curves which looks like
  2628. 1:50:07this now what is the problem with this
  2629. 1:50:08curve here you have lot of local Minima
  2630. 1:50:11if local Minima is there you will never
  2631. 1:50:13reach This Global Minima so that is the
  2632. 1:50:15reason we cannot use that c function now
  2633. 1:50:18mathematically you can also go and
  2634. 1:50:20probably search in the Google what is
  2635. 1:50:22the
  2636. 1:50:23what is the graph or what is a convex or
  2637. 1:50:25non-convex function but always remember
  2638. 1:50:27whenever we updates Theta 1 with this
  2639. 1:50:30within this particular equation by
  2640. 1:50:32finding the slope then this way it will
  2641. 1:50:35not be differentiable and here you have
  2642. 1:50:37lot of local Minima and because of this
  2643. 1:50:39local Minima you will never be able to
  2644. 1:50:41reach the global Minima this is your
  2645. 1:50:42Global Minima right in case
  2646. 1:50:45of in case of linear regression you'll
  2647. 1:50:48reach This Global Minima but in this
  2648. 1:50:50case you will never reach never never
  2649. 1:50:52you'll be stuck over here or you may get
  2650. 1:50:54stuck over here you may get stuck over
  2651. 1:50:56here okay so this has a local Minima
  2652. 1:51:00problem so how do we solve this
  2653. 1:51:02understand in local Minima these are my
  2654. 1:51:03points right I have to come over here
  2655. 1:51:05this is my deepest point in this
  2656. 1:51:07particular case I don't have any local
  2657. 1:51:09Minima now in local Minima also you'll
  2658. 1:51:11get slope is equal to Z so that is the
  2659. 1:51:13reason your Theta 1 will never get
  2660. 1:51:14updated so in order to solve this
  2661. 1:51:17problem you can see this diagram we have
  2662. 1:51:19something called as logistic regression
  2663. 1:51:20cost function so I can now write my
  2664. 1:51:23logistic regression cost function in a
  2665. 1:51:25different way so this researcher
  2666. 1:51:27researcher thought of it and basically
  2667. 1:51:30came up with this proposal that the
  2668. 1:51:31logistic cost function should look
  2669. 1:51:33something like this so the entire cost
  2670. 1:51:36function of logistic regression that is
  2671. 1:51:38specifically H Theta of X of I comma y
  2672. 1:51:43this should be written something like
  2673. 1:51:44this and it should be written like this
  2674. 1:51:47see here I'm just going to write cost
  2675. 1:51:49function of J of theta 1 let's say that
  2676. 1:51:51I'm writing J of theta 1 okay so J of
  2677. 1:51:54theta 1 what are the different different
  2678. 1:51:56output that I'll be getting I'll be get
  2679. 1:51:58I'll be getting yal 1 or y equal to 0 So
  2680. 1:52:02based on this two scenarios our cost
  2681. 1:52:04function will look something like this
  2682. 1:52:06minus log of H of theta of X and I know
  2683. 1:52:11I hope you all know what is h Theta of x
  2684. 1:52:13h Theta of X is nothing but 1 + 1 ^ of -
  2685. 1:52:19Theta 1 x so this is what is my H Theta
  2686. 1:52:22of X and whenever Y is Zer then you
  2687. 1:52:25basically have minus log * 1 - H Theta
  2688. 1:52:31of X of I of I okay so this is how you
  2689. 1:52:35basically write your cost function in
  2690. 1:52:36this particular scenario now with the
  2691. 1:52:38help of this cost function it is always
  2692. 1:52:40possible since it is getting log log is
  2693. 1:52:42basically getting used in this scenario
  2694. 1:52:45you'll always get a global Minima that
  2695. 1:52:46is the reason why they have completely
  2696. 1:52:48neglected this cost function and utiliz
  2697. 1:52:51this cost function now what does this
  2698. 1:52:52cost function basically mean two
  2699. 1:52:55scenarios if Y is equal to 1 Let's
  2700. 1:52:58consider this is my cost function
  2701. 1:53:01graph I have H Theta of X and you know
  2702. 1:53:06that H Theta of x value will be ranging
  2703. 1:53:08between 0 to 1 since it is a
  2704. 1:53:10classification problem so it will be
  2705. 1:53:11ranging between 0 to 1 and this is
  2706. 1:53:14basically of J of theta 1 which is my
  2707. 1:53:16cost function so if Y is equal to 1 this
  2708. 1:53:19specific equation will be used and
  2709. 1:53:21whenever this equation is is basically
  2710. 1:53:22used you get a you get a curve see minus
  2711. 1:53:25log s of X of I you get a curve which
  2712. 1:53:29looks something like this okay which
  2713. 1:53:31you'll get a curve which looks like this
  2714. 1:53:33now what does this curve basically
  2715. 1:53:35specify the curve come up with two
  2716. 1:53:37assumptions the cost will be zero if Y
  2717. 1:53:42is = 1 and H Theta of x equal to 1 that
  2718. 1:53:46basically when your s Theta of X is 1
  2719. 1:53:49and the Y is output is one that
  2720. 1:53:51basically means you're going to assign
  2721. 1:53:52over here one right so in this
  2722. 1:53:54particular case you will be seeing that
  2723. 1:53:56your cost function will be zero cost is
  2724. 1:53:59zero so here is my zero it is meeting
  2725. 1:54:01over here if you of x equal to 1 and Y
  2726. 1:54:04is equal to 1 so this is this is again a
  2727. 1:54:06convex function only then the next point
  2728. 1:54:08that you can probably discuss over here
  2729. 1:54:10is with respect to Y is equal to 0 if
  2730. 1:54:13your Y is Z then what kind of curve you
  2731. 1:54:16will be getting you'll get a different
  2732. 1:54:18kind of curve which will look like this
  2733. 1:54:20H Theta of x here your value will be 0
  2734. 1:54:23to one and here you'll be having a curve
  2735. 1:54:26which looks like this so when you
  2736. 1:54:29combine this two you'll be able to see
  2737. 1:54:31that you are able to get a kind of
  2738. 1:54:34gradient descent so this will definitely
  2739. 1:54:36help us to create a cost function so I
  2740. 1:54:38hope everybody is able to understand
  2741. 1:54:40till here with respect to this and this
  2742. 1:54:42will definitely work so finally I can
  2743. 1:54:45also write my cost function in a
  2744. 1:54:47different way the cost function that I
  2745. 1:54:49will probably write over here so this
  2746. 1:54:50will be my J of theta 1
  2747. 1:54:53so I can come up with a cost function
  2748. 1:54:54which looks like this
  2749. 1:54:57cost of H of theta of X of I comma Yus
  2750. 1:55:02log of H Theta of x if Y is equal
  2751. 1:55:091 and then minus
  2752. 1:55:11log 1 - H Theta of x if Y is equal
  2753. 1:55:170 now I can combine this both and
  2754. 1:55:21probably write something like like this
  2755. 1:55:23I can combine this both and I can
  2756. 1:55:25basically write cost of H Theta of X of
  2757. 1:55:27IA Y is equal to - y log H Theta of X of
  2758. 1:55:35I minus log 1 -
  2759. 1:55:40y okay 1 - y log of 1 - H Theta of X so
  2760. 1:55:47this will be my final cost
  2761. 1:55:50function and here also you can see that
  2762. 1:55:53if I
  2763. 1:55:54replace if I replace y with one then
  2764. 1:55:57what will remain only this particular
  2765. 1:55:59value will remain right this value when
  2766. 1:56:01Y is equal to 1 this thing only will
  2767. 1:56:03come you see over here replace y with
  2768. 1:56:05one probably replace y with one and then
  2769. 1:56:08you'll be able to see so here I can now
  2770. 1:56:10write if Y is equal to 1 my cost
  2771. 1:56:14function will Rook something like this
  2772. 1:56:18which is nothing
  2773. 1:56:19but see Y is 1 then what will happen my
  2774. 1:56:22log of H Theta of X of I will come and
  2775. 1:56:26this 1 - 1 is 0 so 0 multili by anything
  2776. 1:56:29will be 0 if Y is equal to 0 then what
  2777. 1:56:32will happen my cost function will be so
  2778. 1:56:36when it is zero this will - y will
  2779. 1:56:38become 0 0 multili by anything is z so
  2780. 1:56:42here you'll be able to see that I am
  2781. 1:56:43I'll be having minus log 1 - H Theta of
  2782. 1:56:48x i so this both the condition has been
  2783. 1:56:50proved by this cost function
  2784. 1:56:52so this is my cost function yes cost
  2785. 1:56:54function and loss function with respect
  2786. 1:56:55to the number of parameters will be
  2787. 1:56:57almost same so finally if I try to write
  2788. 1:57:00J of theta because I have that 1X 2 m
  2789. 1:57:03also right so 1X 2 m also I have so what
  2790. 1:57:06I'm actually going to do here you will
  2791. 1:57:08be able to see that I can write J of
  2792. 1:57:11theta 1 is equal to 1 by 2 m summation
  2793. 1:57:16of IAL 1 to M and then write down the
  2794. 1:57:19entire equation that you have probably
  2795. 1:57:22over here so here you have minus y or I
  2796. 1:57:26I'll just remove this minus and put it
  2797. 1:57:27over here and this will become plus
  2798. 1:57:29sorry y of I
  2799. 1:57:31* log H Theta of X of I 1 - y of i y
  2800. 1:57:41log 1 - H Theta of X of I so this
  2801. 1:57:45becomes my entire first function and
  2802. 1:57:48obviously you know what is h thet of x H
  2803. 1:57:52Theta of X of I is nothing but 1 + 1 e^
  2804. 1:57:56minus Theta 1 * X and finally my
  2805. 1:57:59convergence algorithm I have to repeat
  2806. 1:58:02this to update Theta 1 repeat until this
  2807. 1:58:07updation that is Theta Theta
  2808. 1:58:11J is equal to Theta J minus learning
  2809. 1:58:15rate derivative with respect to Theta J
  2810. 1:58:18and this will be my J of theta 1 this is
  2811. 1:58:21my repeat until conversion so this is my
  2812. 1:58:24cost function this is my repeat
  2813. 1:58:27algorithm and here I will be updating my
  2814. 1:58:30entire Theta
  2815. 1:58:321 and this solves your problem with
  2816. 1:58:35respect to logistic regression simple
  2817. 1:58:37simple questions may come like how it is
  2818. 1:58:39different from linear regression how it
  2819. 1:58:41is not different from linear regression
  2820. 1:58:44can we say log likelihood a topic from
  2821. 1:58:46probabilistic yes this is uh this is log
  2822. 1:58:50likelihood if now I will discuss about
  2823. 1:58:54performance metrics and this is specific
  2824. 1:58:56to classification problem and binary
  2825. 1:58:59classification I'm talking let's
  2826. 1:59:02consider let's consider I have a data
  2827. 1:59:04set which has X1 X2 and this is y and
  2828. 1:59:09obviously in logistic uh classification
  2829. 1:59:11you have outputs like 0 1 0 1 1 0 1 and
  2830. 1:59:17your y hat y hat is basically the output
  2831. 1:59:20of the predicted model now in this
  2832. 1:59:22particular scenario my y hat will
  2833. 1:59:24probably be 1 1 0 uh 1 1 1 Z so in this
  2834. 1:59:31particular scenario this is my predicted
  2835. 1:59:34output and this is my actual output so
  2836. 1:59:39can we come to some kind of conclusions
  2837. 1:59:41wherein probably we will be able to
  2838. 1:59:44identify what may be the accuracy of
  2839. 1:59:48this specific model with respect to this
  2840. 1:59:49many data points because confusion
  2841. 1:59:52Matrix is all dealt with this is called
  2842. 1:59:54as we will first of all have to create a
  2843. 1:59:56confusion Matrix now for a binary
  2844. 1:59:59classification problem the confusion
  2845. 2:00:01Matrix will look like this so here you
  2846. 2:00:03have 1 0 1 0 Let's say that this is
  2847. 2:00:06prediction let's say that these are my
  2848. 2:00:08actual value and these are my prediction
  2849. 2:00:10value okay these both are prediction
  2850. 2:00:12value these are my output value when my
  2851. 2:00:15actual value is zero my predicted value
  2852. 2:00:17is one does this what does this mean
  2853. 2:00:21wrong prediction right so when my actual
  2854. 2:00:23value is zero my predicted value is 1 so
  2855. 2:00:26here my count will increase to one let's
  2856. 2:00:28go to the second scenario when the
  2857. 2:00:30actual value is one and my predicted
  2858. 2:00:33value is one that basically means one
  2859. 2:00:35and one so here I'm going to increase my
  2860. 2:00:37count similarly when my actual value is
  2861. 2:00:40zero my predicted value is zero so that
  2862. 2:00:42basically mean when my actual value is z
  2863. 2:00:43my predicted value is zero I'm going to
  2864. 2:00:45increase the count by one if I go over
  2865. 2:00:47here 1 one again it is so instead of
  2866. 2:00:50writing one now this will become two I'm
  2867. 2:00:52going to increase the count similarly
  2868. 2:00:54I'll go over here one more one is there
  2869. 2:00:56so I'm going to increase the count three
  2870. 2:00:58then I have 01 01 basically means when
  2871. 2:01:00my actual value is zero I'm actually
  2872. 2:01:02getting it as one so I'm also going to
  2873. 2:01:04increase this particular value as two
  2874. 2:01:07and then finally I have 1 and zero where
  2875. 2:01:09I'm going to increase like this now what
  2876. 2:01:11does this basically mean now what does
  2877. 2:01:13this basically mean see with respect to
  2878. 2:01:16this kind of predictions whenever we are
  2879. 2:01:17discussing this basically basically says
  2880. 2:01:20so this is my actual values and I have Z
  2881. 2:01:221 and zero and this is my predicted
  2882. 2:01:24values I also have 1 and zero this value
  2883. 2:01:27when one and one are there this is
  2884. 2:01:29called as true positive this value when
  2885. 2:01:310 and Zer are there this is called as
  2886. 2:01:33false negative whenever your actual
  2887. 2:01:35value is zero and you have predicted one
  2888. 2:01:37this becomes false positive and whenever
  2889. 2:01:40your actual value is one you have
  2890. 2:01:41predicted zero this becomes false
  2891. 2:01:43negative now coming to this I really
  2892. 2:01:45need to find out the accuracy of this
  2893. 2:01:47model now if I really want to find out
  2894. 2:01:51and this is what is called as confusion
  2895. 2:01:52Matrix now in this confusion Matrix if I
  2896. 2:01:55really want to find out the accuracy the
  2897. 2:01:57accuracy of this model it is very much
  2898. 2:01:59simple this middle elements that you are
  2899. 2:02:01able to see will basically give us the
  2900. 2:02:03right output so this and this if I add
  2901. 2:02:07it up it will give us the right output
  2902. 2:02:10so here I'm going to get TP + TN divided
  2903. 2:02:13by TP + FP + FN + TN so once I calculate
  2904. 2:02:21this so I have 3 + 1
  2905. 2:02:23/ 3 + 2 + 1 + 1 so this is nothing but 4
  2906. 2:02:29by 7 what is 4 by
  2907. 2:02:32757 so am I getting 57 percentage
  2908. 2:02:35accuracy so I'm actually getting 57%
  2909. 2:02:38accuracy over here with respect to the
  2910. 2:02:39accuracy so this is how we basically
  2911. 2:02:42calculate with respect to basic accuracy
  2912. 2:02:45with the help of uh the confusion Matrix
  2913. 2:02:48okay so this is specifically called as
  2914. 2:02:49confusion Matrix now there are some more
  2915. 2:02:52things that you really need to specify
  2916. 2:02:54always remember our model aim should be
  2917. 2:02:56that we should try to reduce false
  2918. 2:02:57positive and false negative now let's
  2919. 2:03:00say that I want to discuss about two
  2920. 2:03:02topics what one is suppose in our data
  2921. 2:03:04set I have zeros and one category let's
  2922. 2:03:07say in my output if I say Zer are 900
  2923. 2:03:11and ones are 100 this becomes an
  2924. 2:03:13imbalanced data very clear right so this
  2925. 2:03:15become an imbalanced data set it is a
  2926. 2:03:18biased data suppose if I say zeros are
  2927. 2:03:21probab
  2928. 2:03:22600 and ones are probably 400 in this
  2929. 2:03:25particular scenario I will say that this
  2930. 2:03:27is the balance data because yes you have
  2931. 2:03:29100 less but it's okay the it may not
  2932. 2:03:32impact many of the algorithm now see
  2933. 2:03:34guys most of the algorithm that we will
  2934. 2:03:36be probably discussing imbalanced if we
  2935. 2:03:38have an imbalanced data set it will
  2936. 2:03:40obviously affect the algorithms let me
  2937. 2:03:42talk about this let's say that I have
  2938. 2:03:44number of zeros as 900 and number of
  2939. 2:03:46ones is 100 now let's say that my model
  2940. 2:03:49I have created which will directly
  2941. 2:03:51predict
  2942. 2:03:52zero it'll I'll just say that all my
  2943. 2:03:55inputs that it is probably getting with
  2944. 2:03:57respect to this training data it'll just
  2945. 2:03:59output zero now in this particular
  2946. 2:04:01scenario what will be my accuracy my
  2947. 2:04:03accuracy will be 900 divid by 1,000
  2948. 2:04:05right so this is nothing but 90% so is
  2949. 2:04:09this a good
  2950. 2:04:10accuracy obviously it is a good accuracy
  2951. 2:04:12but this is a biased data if my model is
  2952. 2:04:15basically just outputting 00000000 0 if
  2953. 2:04:19it is outputting 00 00 0 obviously most
  2954. 2:04:22of the answer will be zeros but this
  2955. 2:04:24will be a scenario like you know where
  2956. 2:04:27it is just outputting one thing then
  2957. 2:04:28also it is able to get 90% accuracy so
  2958. 2:04:31you should only not be dependent on
  2959. 2:04:33accuracy so there are lot of
  2960. 2:04:35terminologies that we will basically use
  2961. 2:04:37one terminology that we specifically use
  2962. 2:04:40is something called as Precision then
  2963. 2:04:42we'll also use recall what is precision
  2964. 2:04:45what is recall I'll write the formula
  2965. 2:04:46over here in Precision what do we need
  2966. 2:04:48to focus and then finally we will
  2967. 2:04:50discuss about f score so we have to use
  2968. 2:04:53different kind of parametrics of sorry
  2969. 2:04:55different kind of formulas whenever you
  2970. 2:04:58have an imbalanced data set you can also
  2971. 2:04:59do oversampling but again understand in
  2972. 2:05:02most of the scenarios in some of the
  2973. 2:05:04scenarios oversampling may work but we
  2974. 2:05:06have to focus on the type of performance
  2975. 2:05:08metrics that we are focusing on right
  2976. 2:05:10now I'll not say F1 score I'll say F
  2977. 2:05:11score the reason why I'm saying I'll
  2978. 2:05:13just let you know so let's talk about
  2979. 2:05:15recall recall formula is basically given
  2980. 2:05:17by true positive divided by true
  2981. 2:05:20positive plus false negative
  2982. 2:05:22Precision is given by true positive
  2983. 2:05:23divided by true positive plus false
  2984. 2:05:27positive and then I will probably
  2985. 2:05:29discuss about F sore also or we
  2986. 2:05:31basically say fbaa also now I'll just
  2987. 2:05:34draw this confusion Matrix again okay
  2988. 2:05:36which is having true positive true
  2989. 2:05:37negative so let me draw it over here so
  2990. 2:05:40this is my ones and zeros these are my
  2991. 2:05:42actual values and these are my predicted
  2992. 2:05:44values I have true positive I have true
  2993. 2:05:47negative false positive and false
  2994. 2:05:49negative now in this particular scenario
  2995. 2:05:50when I'm actually discussing understand
  2996. 2:05:53what is recall and what focus it is
  2997. 2:05:54basically given on so here whenever I
  2998. 2:05:57talk about recall recall basically says
  2999. 2:05:59that TP TP divided by TP plus FN so I'm
  3000. 2:06:04actually focusing on this so what does
  3001. 2:06:06this basically say true uh recall out of
  3002. 2:06:10all the actual true positives how many
  3003. 2:06:13have been predicted correctly that is
  3004. 2:06:15basically mentioned by TP out of all the
  3005. 2:06:18positive values how many of them have
  3006. 2:06:20predicted as positive so this is what it
  3007. 2:06:22is basically saying and this scenario is
  3008. 2:06:24called as recall in this the false
  3009. 2:06:27negative is basically given more
  3010. 2:06:28priority and our focus should be that we
  3011. 2:06:31should try to reduce false positive
  3012. 2:06:33false negative sorry we should try to
  3013. 2:06:35reduce this now let's go ahead and let's
  3014. 2:06:37discuss about Precision in Precision
  3015. 2:06:39what we are doing we are basically
  3016. 2:06:41taking out of all the predicted values
  3017. 2:06:44out of all the predicted positive values
  3018. 2:06:47how many of them are actual true or
  3019. 2:06:50positive okay this is what Precision
  3020. 2:06:52basically means now suppose if I
  3021. 2:06:54consider spam classification suppose
  3022. 2:06:56this is my task tell me in this
  3023. 2:06:57particular case should we use Precision
  3024. 2:07:00or recall and one more use case I'm
  3025. 2:07:02saying that whether the person has
  3026. 2:07:05cancer or not in which case we have to
  3027. 2:07:08support recall and in which case we have
  3028. 2:07:10to go ahead with Precision has cancer or
  3029. 2:07:13not in spam what is important okay guys
  3030. 2:07:16the recall is also called as true
  3031. 2:07:18positive rate I can also say recall as
  3032. 2:07:20sensitivity so if I go with Spam
  3033. 2:07:22classification it should definitely go
  3034. 2:07:24with Precision why it should go with
  3035. 2:07:26Precision if I probably get a Spam ma
  3036. 2:07:28the main aim should be that whenever I
  3037. 2:07:30get a Spam Mill it should be identified
  3038. 2:07:31as spam okay in that specific scenario
  3039. 2:07:34my positive false positive we should try
  3040. 2:07:37to reduce and in this scenario my false
  3041. 2:07:39pository talks about the spam
  3042. 2:07:41classification a lot in a better way in
  3043. 2:07:43the case of cancer I should definitely
  3044. 2:07:46use recall let's let's focus on the
  3045. 2:07:48recall formula tp/ by TP plus FN if a
  3046. 2:07:52person has a cancer see one actually he
  3047. 2:07:55has a cancer it should be predicted as
  3048. 2:07:57one otherwise if we have FN it is
  3049. 2:07:59basically predicting it does not have a
  3050. 2:08:01cancer that is really a big situation in
  3051. 2:08:04this case if a person does not have a
  3052. 2:08:07Cancer and if he's predict if the model
  3053. 2:08:09predicts okay fine he has a cancer he
  3054. 2:08:11may go and further do the test and then
  3055. 2:08:13he'll come to know whether he has a
  3056. 2:08:14cancer or not but this scenario is very
  3057. 2:08:16dangerous if a person has a cancer but
  3058. 2:08:19he is being indicated that he does not
  3059. 2:08:20have that cancer
  3060. 2:08:22so here false negative is given more
  3061. 2:08:24priority over here in the case of spam
  3062. 2:08:26classification false positive is given
  3063. 2:08:28more priority so this is something
  3064. 2:08:30important over here and you really need
  3065. 2:08:31to understand with respect to different
  3066. 2:08:33different problem statement let me give
  3067. 2:08:35you one more example tomorrow the stock
  3068. 2:08:37market is going to crash in this what we
  3069. 2:08:40need to focus on should we focus on
  3070. 2:08:41Precision or should we focus on recall
  3071. 2:08:44now here two things are there who is
  3072. 2:08:46solving what kind of problem see many
  3073. 2:08:48people will say recall or Precision but
  3074. 2:08:50here two things are there on whose point
  3075. 2:08:52of view you are creating this model are
  3076. 2:08:55you creating this model for the industry
  3077. 2:08:57or are you creating this model for the
  3078. 2:08:59people for the people he should
  3079. 2:09:01definitely get identified that okay in
  3080. 2:09:04this particular scenario you need to
  3081. 2:09:06sell your stock because tomorrow stock
  3082. 2:09:07market is going to crash but for
  3083. 2:09:09companies this is very bad okay I hope
  3084. 2:09:11everybody is able to understand for
  3085. 2:09:13companies it is very very bad so in this
  3086. 2:09:15particular case sometime we need to
  3087. 2:09:17focus both on false positive and false
  3088. 2:09:19negative and again I'm telling you for
  3089. 2:09:22which problem statement you are solving
  3090. 2:09:23that indicates if you are solving for
  3091. 2:09:25people then they should be able to get
  3092. 2:09:27the notification saying that it is going
  3093. 2:09:29to crash if you're probably uh doing it
  3094. 2:09:32for companies at that time your
  3095. 2:09:34Precision recall may change but if I
  3096. 2:09:36consider for both the scenarios at that
  3097. 2:09:39point of time I will definitely use
  3098. 2:09:40something called as F score F score or
  3099. 2:09:42I'll also say it as F beta now how is
  3100. 2:09:45fbaa Formula given as I will talk about
  3101. 2:09:48it and here in the F score you have
  3102. 2:09:50three different formulas the first
  3103. 2:09:51Formula I will say basically as when
  3104. 2:09:53your beta value is 1 okay first of all
  3105. 2:09:57I'll just give a generic definition of f
  3106. 2:09:59s or F beta here you are basically going
  3107. 2:10:01to consider 1 + beta squ Precision
  3108. 2:10:05multiplied by recall divided beta Square
  3109. 2:10:09* Precision plus recall whenever your
  3110. 2:10:14both false positive and false negative
  3111. 2:10:16are important we select beta as one so
  3112. 2:10:19if I select beta as 1 it becomes 1 + 4
  3113. 2:10:22Precision multiplied by recall then you
  3114. 2:10:25have Precision plus recall so here sorry
  3115. 2:10:281 + 1 so this becomes 2 multiplied by
  3116. 2:10:31Precision into recall divided by
  3117. 2:10:34Precision plus recall so here you have
  3118. 2:10:37this is basically called as harmonic
  3119. 2:10:39mean harmonic mean probably you have
  3120. 2:10:41seen this kind of equation where you
  3121. 2:10:42have written 2x y / x + y same type you
  3122. 2:10:46are able to see this this is called as
  3123. 2:10:48harmonic mean here the focus is on both
  3124. 2:10:51false positive and false negative let's
  3125. 2:10:53say that your false positive is more
  3126. 2:10:56important than false negative at that
  3127. 2:10:58point of time you will try to decrease
  3128. 2:11:01or you will try to decrease your beta
  3129. 2:11:03value let's say that I'm decreasing my
  3130. 2:11:05Beta value to 0.5 then what will happen
  3131. 2:11:071 +5 whole
  3132. 2:11:09s and then you have P * R Precision
  3133. 2:11:13recall and here also you have 25 p + r
  3134. 2:11:17now in this particular scenario I'm
  3135. 2:11:19decreasing my Beta decreasing the beta
  3136. 2:11:21basically means that you are providing
  3137. 2:11:23more importance to false positive than
  3138. 2:11:25false negative and finally you'll be
  3139. 2:11:27able to see that if I consider beta
  3140. 2:11:30value as let me just say my notes if I
  3141. 2:11:34consider beta value as two that
  3142. 2:11:37basically means you are giving more
  3143. 2:11:38importance to false negative than false
  3144. 2:11:40positive so with this specific case you
  3145. 2:11:42can come up to a conclusion what value
  3146. 2:11:44you basically want to use now whenever I
  3147. 2:11:46use beta is equal to 1 it becomes fub1
  3148. 2:11:49score if I use beta as .5 then this
  3149. 2:11:52basically becomes f.5 score and this
  3150. 2:11:56becomes your F2 score So based on which
  3151. 2:12:00is important okay which is important
  3152. 2:12:03whether your Precision or false positive
  3153. 2:12:05or false negative is important you can
  3154. 2:12:06consider those things F score will have
  3155. 2:12:09different values if you're using beta is
  3156. 2:12:11equal to 1 that basically means you are
  3157. 2:12:13giving importance to both precision and
  3158. 2:12:16recall if your false positive is more
  3159. 2:12:18important then at that point of time you
  3160. 2:12:20reduce beta value if false negative is
  3161. 2:12:23greater than false bet uh false positive
  3162. 2:12:25then your beta value is
  3163. 2:12:26increasing beta is a deciding parameter
  3164. 2:12:29to decide your F1 score or F2 score or F
  3165. 2:12:32Point score now first thing first what
  3166. 2:12:34is the agenda of today's session first
  3167. 2:12:36of all we will complete practicals for
  3168. 2:12:39all the algorithms that we have
  3169. 2:12:41discussed these all algorithms that we
  3170. 2:12:43have discussed we will cover the
  3171. 2:12:45practicals probably we will be doing
  3172. 2:12:47hyper parameter tuning everything the
  3173. 2:12:49second thing and again here we are going
  3174. 2:12:51to take just simple examples so yes uh
  3175. 2:12:54so today's session I said practicals
  3176. 2:12:56with simple examples where I'll probably
  3177. 2:12:59discuss about all the hyper parameter
  3178. 2:13:01tuning then the second one the second
  3179. 2:13:04algorithm that I'm going to discuss
  3180. 2:13:05about is something called as n bias this
  3181. 2:13:09is a classification algorithm so we are
  3182. 2:13:10going to understand the intuition and
  3183. 2:13:13the third one that we are going to
  3184. 2:13:15probably discusses KNN algorithm so KNN
  3185. 2:13:19algorithms is definitely there
  3186. 2:13:21so this our today's plan I know I've
  3187. 2:13:23written very less but this much maths
  3188. 2:13:26and involved in na bias right we'll
  3189. 2:13:29understand the probability theorem again
  3190. 2:13:30over there there is something called as
  3191. 2:13:32bias theorem we'll try to understand and
  3192. 2:13:35then we'll try to solve a problem on
  3193. 2:13:36that so let's proceed and let's enjoy
  3194. 2:13:39today's session how do we enjoy first of
  3195. 2:13:42all we enjoy by creating a practical
  3196. 2:13:44problem so I am actually opening a
  3197. 2:13:47notebook file in front of you so here uh
  3198. 2:13:50we will try to Sol solve it with the
  3199. 2:13:51help of linear regression Ridge lasso
  3200. 2:13:55and try to solve some problems let's see
  3201. 2:13:58how much we will be able to solve it but
  3202. 2:14:00again the aim is that we learn in a
  3203. 2:14:02better way okay uh so that everybody
  3204. 2:14:06understands some basic basic things okay
  3205. 2:14:08so first of all as usual uh everybody
  3206. 2:14:11open your jupyter notebook file the
  3207. 2:14:13first algorithm that I'm going to
  3208. 2:14:14discuss about is something called as SK
  3209. 2:14:16learn linear regression so everybody I
  3210. 2:14:19hope everybody knows about this SK learn
  3211. 2:14:21let's see what all things are basically
  3212. 2:14:23there in this we will be using fit
  3213. 2:14:25intercept everything as such but here
  3214. 2:14:28the main aim is to find out the
  3215. 2:14:29coefficients which is basically
  3216. 2:14:31indicated by Theta 0 Theta 1 and all the
  3217. 2:14:34first thing we'll start with linear
  3218. 2:14:38regression and then we will go ahead and
  3219. 2:14:40discuss with r and lassor I'm just going
  3220. 2:14:42to make this as
  3221. 2:14:44markdown how many different libraries of
  3222. 2:14:46for linear regression you can do with
  3223. 2:14:48stats you can do with skyi you can do
  3224. 2:14:49with many things okay so first thing
  3225. 2:14:52first let's first of all we require a
  3226. 2:14:53data set so for the data set what we are
  3227. 2:14:56going to do is that we are going to
  3228. 2:14:58basically take up some smaller smaller
  3229. 2:15:01data just let me do this so for this uh
  3230. 2:15:05we are going to take the house pricing
  3231. 2:15:07data set so we are going to solve house
  3232. 2:15:10pricing data set problem a simple data
  3233. 2:15:13set which is already present in SK learn
  3234. 2:15:16only now in order to import the data set
  3235. 2:15:18I will write a line of code which is
  3236. 2:15:19like from SK learn dot data sets data
  3237. 2:15:24sets
  3238. 2:15:25import load uncore Boston so we have
  3239. 2:15:29some Boston house pricing data set so
  3240. 2:15:31I'm just going to execute this I'm also
  3241. 2:15:33going to make a lot of Sals so that I
  3242. 2:15:35don't have to again go ahead and create
  3243. 2:15:37all the sales again some basic libraries
  3244. 2:15:39that I probably want is pro import numai
  3245. 2:15:43as
  3246. 2:15:44NP
  3247. 2:15:45import pandas
  3248. 2:15:48SPD okay import cbon as
  3249. 2:15:52SNS and then I will also import Matt
  3250. 2:15:56Matt plot lib do p plot as PLT and then
  3251. 2:16:02percentile matplot lib matlot lib do
  3252. 2:16:07inline and I will try to execute this
  3253. 2:16:09see this my typing speed has become a
  3254. 2:16:11little bit faster by writing by
  3255. 2:16:12executing this queries again and again
  3256. 2:16:15and uh let's go ahead uh so I have
  3257. 2:16:18imported all the necessary libraries
  3258. 2:16:19that is required which which will be
  3259. 2:16:21more than sufficient for you all to
  3260. 2:16:23start with now in order to load this
  3261. 2:16:25particular data set I will just use this
  3262. 2:16:27Library called as load uncore Boston and
  3263. 2:16:30I'm going to just initialize this so if
  3264. 2:16:32you press shift tab you will be able to
  3265. 2:16:34see that return load and return the
  3266. 2:16:37Boston house prices data set it is a
  3267. 2:16:39regression problem it is saying and then
  3268. 2:16:41probably I'm just going to execute it
  3269. 2:16:43now once I execute it I will go and
  3270. 2:16:45probably see the type of DF so it is
  3271. 2:16:48basically saying skarn dos. bunch now if
  3272. 2:16:51I go and probably execute DF you'll be
  3273. 2:16:53able to see that this will be in the
  3274. 2:16:55form of key value pairs okay like Target
  3275. 2:16:57is here data is here okay so data is
  3276. 2:17:01here Target is here and probably you'll
  3277. 2:17:03be able to find out feature names is
  3278. 2:17:04here so we definitely require feature
  3279. 2:17:06names we require our Target value and
  3280. 2:17:09our data value so we really need to
  3281. 2:17:11combine this specific thing in a proper
  3282. 2:17:14way in the form of a data frame so that
  3283. 2:17:16you will be able to see so what I'm
  3284. 2:17:18actually going to do over here I'm just
  3285. 2:17:19going to say PD do data frame I'll
  3286. 2:17:22convert this entirely into a data frame
  3287. 2:17:24and I will say DF do data see this is a
  3288. 2:17:27key value pair right so DF do data is
  3289. 2:17:29basically giving me all the features
  3290. 2:17:31value so if I write DF do data and just
  3291. 2:17:35execute it you'll be able to see that I
  3292. 2:17:36you will be able to get my entire data
  3293. 2:17:39set in this way my entire data set in
  3294. 2:17:41this way this is my feature one feature
  3295. 2:17:43two feature three feature 4 this feature
  3296. 2:17:4512 I have 12 features over here and
  3297. 2:17:47based on that I have that specific value
  3298. 2:17:50now the next thing thing that I'm going
  3299. 2:17:51to do probably I should also be able to
  3300. 2:17:53add the target feature name over here so
  3301. 2:17:55what I will do I will just convert this
  3302. 2:17:57into DF and then I will also say DF do
  3303. 2:18:02columns and I'll set it to DF do Target
  3304. 2:18:05okay and let me change this to data set
  3305. 2:18:08so I'm going to change this to data set
  3306. 2:18:10and I'm going to say data set. columns
  3307. 2:18:12is equal to DF do Target so if I execute
  3308. 2:18:15this and now if I probably
  3309. 2:18:18print my data set do head you will be
  3310. 2:18:22able to see this specific thing okay it
  3311. 2:18:24is an error let's see expected axis has
  3312. 2:18:2713 element new values has
  3313. 2:18:30506 so Target okay I should not use
  3314. 2:18:33Target over here instead I had a column
  3315. 2:18:36which is called as features feature
  3316. 2:18:38names like if I go and probably see
  3317. 2:18:41DF DF over here you'll be able to see
  3318. 2:18:45there is one thing which is called as
  3319. 2:18:46feature names so I'm going to use DF do
  3320. 2:18:48feature names over here so here it is DF
  3321. 2:18:52do feature names I'm just going to paste
  3322. 2:18:55it over here and now if I go and write
  3323. 2:18:57here you can see print DF data set. head
  3324. 2:19:00if I go and execute without print you'll
  3325. 2:19:02be able to see my entire data set so
  3326. 2:19:04these are my features with respect to
  3327. 2:19:06different different things and this is
  3328. 2:19:09basically a house pricing data set so
  3329. 2:19:10initially I have this features CRM ZN
  3330. 2:19:13indust CH nox RM age distance radius tax
  3331. 2:19:18PT ratio b l stack that so I have my
  3332. 2:19:22entire data set over here the same data
  3333. 2:19:24set I have basically put it over here
  3334. 2:19:26now here also you'll be able to see what
  3335. 2:19:28all this feature basically means this is
  3336. 2:19:30showing wasted weighted distance to five
  3337. 2:19:31do uh Five Boston employment center rad
  3338. 2:19:34basically means index of accessibility
  3339. 2:19:36to radial Highway tax basically means
  3340. 2:19:39full value property tax rate this much
  3341. 2:19:41PT rate basically means pupil teacher
  3342. 2:19:44ratio I don't know what the hell it
  3343. 2:19:45means but it's fine we have some kind of
  3344. 2:19:47data over here properly in front of you
  3345. 2:19:51so these are my independent features
  3346. 2:19:53what are these these all are my
  3347. 2:19:54independent features if you want the
  3348. 2:19:56features detail here you can see it
  3349. 2:19:59right everything what is CRM this
  3350. 2:20:01basically means per capita crime rate by
  3351. 2:20:03town which is important ZN it is
  3352. 2:20:06proportional of residential land zone
  3353. 2:20:08for Lots over 25,000 Square ft so this
  3354. 2:20:12is my DF I did not do much I'm just
  3355. 2:20:14using data frame DF do data column
  3356. 2:20:17features name I'm getting this value
  3357. 2:20:18very much simple now let's go a little
  3358. 2:20:21bit slowly so that many people will be
  3359. 2:20:23able to also understand now this is my
  3360. 2:20:25data set. head now the thing is that I
  3361. 2:20:29obviously have taken all these
  3362. 2:20:31particular values but this is my
  3363. 2:20:32independent feature I still have my
  3364. 2:20:35dependent feature so what I'm actually
  3365. 2:20:37going to do I will create a new feature
  3366. 2:20:40which is like data set of price I'll
  3367. 2:20:42create my feature name price price of
  3368. 2:20:44the house and what I will assign this
  3369. 2:20:46particular value this value will be
  3370. 2:20:48assigned with this target this target
  3371. 2:20:50value this target value is basically the
  3372. 2:20:53sale the price of the houses right it is
  3373. 2:20:56again in the form of array so I'm going
  3374. 2:20:58to take this and put it as a dependent
  3375. 2:21:00feature so here you'll be able to see
  3376. 2:21:02that my price will be my dependent
  3377. 2:21:04feature so here I'll basically write DF
  3378. 2:21:06do Target so once I execute it and now
  3379. 2:21:09if I probably go and see my data set do
  3380. 2:21:12head you'll be able to see features over
  3381. 2:21:15here and one more feature is getting
  3382. 2:21:17added that is price now this price may
  3383. 2:21:20be the units may be in
  3384. 2:21:22millions somewhere Target should be here
  3385. 2:21:24or there it should be probably in
  3386. 2:21:27millions
  3387. 2:21:28or I cannot see it but it should be
  3388. 2:21:31somewhere here it should have definitely
  3389. 2:21:33said that it is probably in millions or
  3390. 2:21:36okay but that is not a problem I think
  3391. 2:21:37but mostly it'll be in millions
  3392. 2:21:39somewhere I think it should be
  3393. 2:21:42here okay I cannot see it but probably
  3394. 2:21:45if I put more time I'll be able to
  3395. 2:21:47understand it okay so over here what is
  3396. 2:21:49the thing main thing this all are my
  3397. 2:21:51independent features and this is my
  3398. 2:21:53dependent feature right so if I'm trying
  3399. 2:21:55to solve linear regression I have to
  3400. 2:21:57divide my independent and dependent
  3401. 2:21:58features properly now let's go to the
  3402. 2:22:01next step that
  3403. 2:22:03is
  3404. 2:22:05dividing the data
  3405. 2:22:07set dividing the oh my God dividing the
  3406. 2:22:12data
  3407. 2:22:14set
  3408. 2:22:17into
  3409. 2:22:19train into first of all I'll try try to
  3410. 2:22:22divide into
  3411. 2:22:24independent and dependent
  3412. 2:22:27features so I want my entire features
  3413. 2:22:30data set divided into independent and
  3414. 2:22:31dependent features X I will be using as
  3415. 2:22:34my independent featur so I will write
  3416. 2:22:35data set dot I will use an iock which is
  3417. 2:22:39present in data frames and understand
  3418. 2:22:41from which feature to which feature I
  3419. 2:22:42will be taking as my independent feature
  3420. 2:22:44to this feature till lat so the best way
  3421. 2:22:48that basically means that I just need to
  3422. 2:22:49skip the last feature in order to skip
  3423. 2:22:52the last feature what I'm actually going
  3424. 2:22:54to do from all the columns I will just
  3425. 2:22:57skip the last column so this is how you
  3426. 2:22:59basically do an indexing with respect to
  3427. 2:23:02just skipping the last feature and this
  3428. 2:23:05will basically be my independent
  3429. 2:23:06features and here I will basically say Y
  3430. 2:23:08is equal to data set do iock and here I
  3431. 2:23:11just want the last feature so I will
  3432. 2:23:14write colon all the records I want and
  3433. 2:23:18see the first term that we are probably
  3434. 2:23:20WR writing over here this basically
  3435. 2:23:22specifies with respect to records here
  3436. 2:23:24this specifies with respect to columns
  3437. 2:23:26from all the columns I'm taking the last
  3438. 2:23:27column here I will just take the last
  3439. 2:23:29column and this will basically be my
  3440. 2:23:32dependent features dependent features so
  3441. 2:23:35here I have basically executed now if
  3442. 2:23:37you can go and probably see x. head here
  3443. 2:23:40you'll be able to find all my
  3444. 2:23:41independent features in y do head you'll
  3445. 2:23:43be able to find the dependent feature
  3446. 2:23:45now let's go to the first algorithm that
  3447. 2:23:47is called as linear regression
  3448. 2:23:51always remember whenever I definitely
  3449. 2:23:53start with linear regression I'll
  3450. 2:23:55definitely not go directly with linear
  3451. 2:23:56regression instead what I will do is
  3452. 2:23:59that I'll try to go with Ridge
  3453. 2:24:00regression and uh lasso regression
  3454. 2:24:02because there you are lot of options
  3455. 2:24:04with respect to hyper pment T but I'll
  3456. 2:24:06just show you how linear regression is
  3457. 2:24:08done so basically you really really need
  3458. 2:24:11to use a lot of libraries okay over here
  3459. 2:24:13and based on this libraries this
  3460. 2:24:15libraries will try to install okay and
  3461. 2:24:18what are these libraries these are
  3462. 2:24:19basically the linear regression Library
  3463. 2:24:21so here I'm basically going to use two
  3464. 2:24:23specific thing one is linear regression
  3465. 2:24:25Library so I will just use from SK learn
  3466. 2:24:28do linear uncore model import linear
  3467. 2:24:32regression do you need to remember this
  3468. 2:24:35the answer is no because I also do the
  3469. 2:24:37Google and I try to find out where in
  3470. 2:24:39escal and it is present okay so here is
  3471. 2:24:42my linear regression so I will try to
  3472. 2:24:44initialize linear reg is equal to
  3473. 2:24:47initialize with linear regression and
  3474. 2:24:49then here what I'm actually going to do
  3475. 2:24:51I'm going to basically apply something
  3476. 2:24:53called as cross validation cross
  3477. 2:24:55validation is very much important
  3478. 2:24:57because in Cross validation we divide
  3479. 2:24:59out train and test data in such a way
  3480. 2:25:01that every combination of the train and
  3481. 2:25:04test data is basically taken by care is
  3482. 2:25:07taken by the model and whoever accuracy
  3483. 2:25:09is better that all entire thing is
  3484. 2:25:11basically combined so here what I'm
  3485. 2:25:13going to do I'm going to say mean square
  3486. 2:25:14error is equal to here I will import one
  3487. 2:25:17more Library let's say from SK learn
  3488. 2:25:20dot model selection I'm going to import
  3489. 2:25:25cross Val
  3490. 2:25:26score so cross Val score cross
  3491. 2:25:29validation score basically means it is
  3492. 2:25:31going to do a lot of train and test
  3493. 2:25:32split it's something like this one
  3494. 2:25:34example I will show it to you here only
  3495. 2:25:37so what does cross validation basically
  3496. 2:25:39do okay so in Cross validation what
  3497. 2:25:42happens what you do suppose this is your
  3498. 2:25:44entire data
  3499. 2:25:46set suppose this is 100 records if you
  3500. 2:25:48do five cross validation then in the
  3501. 2:25:51first this will be your test data and
  3502. 2:25:53remaining all will be your training data
  3503. 2:25:55if in the second cross validation this
  3504. 2:25:58will be your test data and remaining all
  3505. 2:25:59will be your test uh training data like
  3506. 2:26:01this five times you'll be doing cross
  3507. 2:26:03validation by taking the different
  3508. 2:26:05combination of train and test but I'm
  3509. 2:26:07not going to discuss much about it in
  3510. 2:26:09the future if you want a separate
  3511. 2:26:10session I will include that in one of
  3512. 2:26:11the session itself so this was uh
  3513. 2:26:13basically the plan with respect to cross
  3514. 2:26:15validation or cross Val score so here
  3515. 2:26:17I'm going to basically take cross
  3516. 2:26:20Val
  3517. 2:26:21score and here the first parameter that
  3518. 2:26:24I give is my model so linear regression
  3519. 2:26:27is my model and here I will take X and Y
  3520. 2:26:30I'm not doing a train test split
  3521. 2:26:32specifically over here I'm giving the
  3522. 2:26:34entire X and Y and probably based on
  3523. 2:26:36that I'm going to do a cross validation
  3524. 2:26:38over here you can also do train test
  3525. 2:26:39plate initially and then just give the X
  3526. 2:26:42train and Y train over here to do the
  3527. 2:26:43cross validation it is up to you but the
  3528. 2:26:45best practices will be that first you do
  3529. 2:26:47the train test split and then only give
  3530. 2:26:49the train data over here to do the cross
  3531. 2:26:51validation I'm just going to use scoring
  3532. 2:26:53is equal to you can use mean squared
  3533. 2:26:56error negative mean squared error let's
  3534. 2:26:58say that I'm going to use negative mean
  3535. 2:27:00squ error again where do you find all
  3536. 2:27:02these things you will be able to see in
  3537. 2:27:04the SK learn page of L uh cross Val
  3538. 2:27:06score and then finally in the cross Val
  3539. 2:27:08score you give cross validation value as
  3540. 2:27:105 10 whatever you want so after this
  3541. 2:27:13what I'm actually going to do I'm just
  3542. 2:27:14going to basically from this how many
  3543. 2:27:17scores I will get the mean squar error
  3544. 2:27:19will be five since I'm doing five cross
  3545. 2:27:21validation if you don't believe me just
  3546. 2:27:23see over here print msse so here you'll
  3547. 2:27:26be able to see five different values 1 2
  3548. 2:27:303 4 5 right five different mean values
  3549. 2:27:34because we are doing cross five five
  3550. 2:27:36cross validation so here what I'm going
  3551. 2:27:37to write I'm just going to say np. mean
  3552. 2:27:40I want to take the average of all the
  3553. 2:27:41five so here will basically be my
  3554. 2:27:45meanor
  3555. 2:27:46msse okay and then probably I'll print I
  3556. 2:27:49will print my Ms meanor MSC so this will
  3557. 2:27:54be my average score with respect to this
  3558. 2:27:56the negative value is there because we
  3559. 2:27:58have used negative mean squ error but if
  3560. 2:28:00you just consider mean square error then
  3561. 2:28:01it is only 37.1 3 okay so this I have
  3562. 2:28:05actually shown you how to do cross
  3563. 2:28:06validation see with respect to linear
  3564. 2:28:08regression you can't modify much with
  3565. 2:28:10the parameter so that is the reason why
  3566. 2:28:12specifically in order to overcome
  3567. 2:28:14overfitting and do the feature selection
  3568. 2:28:16we use uh R and lasso regression so here
  3569. 2:28:19I will show show you how to do ridge
  3570. 2:28:21ridge regression
  3571. 2:28:24now now in order to do the prediction
  3572. 2:28:26all you have to do is that just go over
  3573. 2:28:28here take the model okay what is the
  3574. 2:28:31model linear R and just say do
  3575. 2:28:37predict so here you can see uh you'll be
  3576. 2:28:40getting a function called as do predict
  3577. 2:28:42and give the test value whatever you
  3578. 2:28:44want to predict automatically the
  3579. 2:28:45prediction will be done so I'm just
  3580. 2:28:46going to remove this and focus on Ridge
  3581. 2:28:48regression right now because I I want to
  3582. 2:28:50show how hyperparameter tuning is done
  3583. 2:28:52in R regression so for R regression the
  3584. 2:28:54simple thing is that I'll be using two
  3585. 2:28:56different libraries from skarn do
  3586. 2:29:01linear linear uncore model I'm going to
  3587. 2:29:05import Ridge so for the ridge it is also
  3588. 2:29:09present in linear underscore model for
  3589. 2:29:11doing the hyperparameter tuning I will
  3590. 2:29:12be using from SK learn do modore
  3591. 2:29:17selection and then I'm going to import
  3592. 2:29:20grid SE CV so these are the two
  3593. 2:29:22libraries that I'm actually going to use
  3594. 2:29:24grid SE CV will be able to help you out
  3595. 2:29:26with the um okay will be able to help
  3596. 2:29:30you out with Hyper parameter tuning and
  3597. 2:29:32then probably you'll be able to do
  3598. 2:29:34that uh difference between MSE and
  3599. 2:29:37negative MSE not big thing guys if you
  3600. 2:29:39use MSE here mean squ error you'll be
  3601. 2:29:42getting 37 I've just used negation of
  3602. 2:29:45MSE it's okay anything is fine you can
  3603. 2:29:48go with MSE also means square error
  3604. 2:29:50there is also another uh another scoring
  3605. 2:29:53area which is like which focuses on
  3606. 2:29:54square root square mean Square uh sorry
  3607. 2:29:58root means Square eror okay so there are
  3608. 2:30:00different different things which you can
  3609. 2:30:01basically focus on okay now in order to
  3610. 2:30:05give you this specific good value I'm
  3611. 2:30:07actually going to do hyper Peter tuning
  3612. 2:30:10now let's go ahead with uh grid s CV so
  3613. 2:30:13here what I'm going to do again I'm
  3614. 2:30:14going to basically Define my model which
  3615. 2:30:16will be
  3616. 2:30:18Ridge okay so this this is what I have
  3617. 2:30:20actually imported now uh let me open the
  3618. 2:30:24ridge skarn so SK learn
  3619. 2:30:28Ridge we need need to understand what
  3620. 2:30:30all parameters are basically
  3621. 2:30:33used do you remember this Alpha value
  3622. 2:30:36guys do you remember this Alpha value
  3623. 2:30:38why do we use Alpha I I told you now
  3624. 2:30:40Alpha multiplied by slope square if you
  3625. 2:30:43remember in Ridge we specifically use
  3626. 2:30:45this right Ridge and lasso regression
  3627. 2:30:48Alpha so this is the alpha the this is
  3628. 2:30:50probably the best parameter we can
  3629. 2:30:52perform hyper parameter tuning the next
  3630. 2:30:54parameter that we can probably perform
  3631. 2:30:56is basically uh this Max iteration okay
  3632. 2:31:00Max iteration basically means how many
  3633. 2:31:01number of iteration how many number of
  3634. 2:31:03times we may probably change the Theta 1
  3635. 2:31:05value to get the right value so we can
  3636. 2:31:08do this so what I'm actually going to do
  3637. 2:31:10I'm going to select some Alpha values
  3638. 2:31:12I'm going to play with this apart from
  3639. 2:31:14that if I want I can also play with the
  3640. 2:31:16other parameters which are uh like kind
  3641. 2:31:19of uh you know probably you can you can
  3642. 2:31:21also play with the iteration parameter
  3643. 2:31:23it is up to you try whichever parameter
  3644. 2:31:25you want to change you can go ahead and
  3645. 2:31:26change it now let me show you how do we
  3646. 2:31:28write this and how do we make sure that
  3647. 2:31:31this specific thing is done now uh
  3648. 2:31:34before doing grid s CV uh let me do one
  3649. 2:31:36thing I will Define my parameters okay
  3650. 2:31:39so here is my Ridge now what I'm going
  3651. 2:31:41to do I'm going to say parameters and in
  3652. 2:31:44this parameter two important value that
  3653. 2:31:46I'm probably going to take is this one
  3654. 2:31:49that is my C value and I will try to
  3655. 2:31:51Define this in the form of dictionaries
  3656. 2:31:53so here the C value that I sorry not C
  3657. 2:31:57just a second
  3658. 2:31:58guys my mistake it is not C it is
  3659. 2:32:03Alpha let's see so how do I Define my
  3660. 2:32:05Alpha value we'll try to see so here the
  3661. 2:32:09parameters will be Alpha C is basically
  3662. 2:32:13for uh logistic regression I'll show you
  3663. 2:32:16so the alpha value I will just mention
  3664. 2:32:18some values like
  3665. 2:32:201 e to the power of -5 that basically Me
  3666. 2:32:2700000000 0 0 0 1 similarly I I can write
  3667. 2:32:311 E to the^ of - 10 that again means 0 0
  3668. 2:32:350 0 0 0 0 0 10 * 0 1 I'm just making fun
  3669. 2:32:38okay so that you will also get
  3670. 2:32:39entertained 1 E to the^ of minus 8 okay
  3671. 2:32:43similarly I can write 1 E to the^ of
  3672. 2:32:45minus 3 from this particular value now
  3673. 2:32:48I'm increasing this value see 1 E to
  3674. 2:32:50the^ of minus 2 and then probably I can
  3675. 2:32:53have 1 5 10 um 20 something like this so
  3676. 2:32:58I'm going to play with all this
  3677. 2:32:59particular parameters for right now
  3678. 2:33:01because in grit or CV what they do is
  3679. 2:33:03that they take all the combination of
  3680. 2:33:04this Alpha value and wherever your uh
  3681. 2:33:07your your model performs well it is
  3682. 2:33:09going to take that specific parameter
  3683. 2:33:11and it is going to give you that okay
  3684. 2:33:13this is the best fit parameter that is
  3685. 2:33:14got selected so here I have got all
  3686. 2:33:16these things now what I'm going to do
  3687. 2:33:18I'm going to basically apply the grid C
  3688. 2:33:19TV so here I have uh gridge uh sorry
  3689. 2:33:24Ridge GD I'm
  3690. 2:33:25saying ridore regressor so I'm going to
  3691. 2:33:28use git s
  3692. 2:33:31CV git s CV and here I'm basically going
  3693. 2:33:34to take the parameters regge okay Ridge
  3694. 2:33:36is my first model and then I will take
  3695. 2:33:39up all this params that I have actually
  3696. 2:33:40defined see in git CV if I press shift
  3697. 2:33:43tab I have to first of all execute this
  3698. 2:33:46then only it will be able to press shift
  3699. 2:33:48tab so here if I press shift tab here
  3700. 2:33:50you'll be able to see estimator and
  3701. 2:33:52parameter grid is my second parameter
  3702. 2:33:54then scoring and then all the other
  3703. 2:33:56parameters so here the first thing that
  3704. 2:33:58goes is your model then your parameters
  3705. 2:34:00which what you are actually playing then
  3706. 2:34:03the third parameter is basically your
  3707. 2:34:05scoring
  3708. 2:34:06scoring and again here I'm going to use
  3709. 2:34:09negative mean squ error some people are
  3710. 2:34:10saying that mean squared error is not
  3711. 2:34:13present so that is the reason why
  3712. 2:34:15negative mean squ error is done why it
  3713. 2:34:18may not be present because
  3714. 2:34:20uh they try to always create a generic
  3715. 2:34:22Library probably this kind of uh scoring
  3716. 2:34:24parameter may also get used in other
  3717. 2:34:26algorithms so that is the reason they
  3718. 2:34:28may not have created but if you want to
  3719. 2:34:30Deep dive into it Google
  3720. 2:34:33Google then what is r regress dot fit on
  3721. 2:34:38X comma y again I'm telling you you can
  3722. 2:34:40first of all do train test split on X
  3723. 2:34:42and Y and then probably only do this on
  3724. 2:34:44X train and Y train parameter is not oh
  3725. 2:34:47sorry
  3726. 2:34:49okay I get this okay parameter is not
  3727. 2:34:53and why it is not and oh yeah it has
  3728. 2:34:56become a
  3729. 2:34:58list I'm going to make this as
  3730. 2:35:00dictionary right now I'm fully focused
  3731. 2:35:02on implementing things if I get an error
  3732. 2:35:04I'll definitely make sure that it'll get
  3733. 2:35:07fixed anyhow if I get that error I will
  3734. 2:35:09not say oh Kish why why this error came
  3735. 2:35:12you
  3736. 2:35:13know why this error came I I'll not get
  3737. 2:35:15worried I'll get the error down only you
  3738. 2:35:17cannot give this as the one okay so try
  3739. 2:35:21to understand okay so this is your gitar
  3740. 2:35:23CV I've also done the fit and let's go
  3741. 2:35:27and select the best parameter so what I
  3742. 2:35:28can do I will write print
  3743. 2:35:32ridore
  3744. 2:35:35regressor dot
  3745. 2:35:37params sorry there will be a parameter
  3746. 2:35:39called as best params I'm going to print
  3747. 2:35:42this and I'm going to print ridore
  3748. 2:35:46regressor Dot
  3749. 2:35:50best
  3750. 2:35:51score so these are all the values that
  3751. 2:35:53are got selected one is Alpha is equal
  3752. 2:35:55to 20 and the best score is - 32 so
  3753. 2:35:58initially I gotus 37 but because of
  3754. 2:36:00Ridge regression you can see that our
  3755. 2:36:02negative mean square error has
  3756. 2:36:04definitely become better there is a
  3757. 2:36:06minus sign don't worry but from 37 it
  3758. 2:36:08has come to 32 cross validation guys
  3759. 2:36:11over here inside grids s CV also when it
  3760. 2:36:13is probably taking the entire
  3761. 2:36:15combination over there the CV Value
  3762. 2:36:17Cross validation also we can use
  3763. 2:36:20so probably if I am probably considering
  3764. 2:36:23all these
  3765. 2:36:24things many people has a question Chris
  3766. 2:36:27is this minus value increased that
  3767. 2:36:29basically means you cannot use Ridge
  3768. 2:36:31regression you are right in this
  3769. 2:36:33particular case Ridge regression is not
  3770. 2:36:34helping you out so guys let me again
  3771. 2:36:36write it down everybody don't worry yeah
  3772. 2:36:41previous I got minus 32 right now I'm
  3773. 2:36:43getting - 37
  3774. 2:36:45right sorry previously I got what - 37
  3775. 2:36:52- 37 now I got - 32 so here you can see
  3776. 2:36:56this I got it from linear regression
  3777. 2:36:59this I got it from what Ridge which one
  3778. 2:37:02should I select I should select this
  3779. 2:37:03model only because it is performing well
  3780. 2:37:05than this but again understand Ridge
  3781. 2:37:08also tries to reduce the overfitting so
  3782. 2:37:11probably in this particular scenario we
  3783. 2:37:12cannot use Ridge because the performance
  3784. 2:37:14is becoming more bad so what I will do I
  3785. 2:37:17will go and try with lasso regression
  3786. 2:37:20now I'll copy and paste the same thing
  3787. 2:37:22so linear model import lasso then this
  3788. 2:37:25will basically be my
  3789. 2:37:27lasso let's see with lasso whether it
  3790. 2:37:29will increase or not let's
  3791. 2:37:33see this is my parameter that got
  3792. 2:37:35selected now let me write lasto
  3793. 2:37:38regressor
  3794. 2:37:39dot best params so this is Alpha is
  3795. 2:37:42equal to one is got selected over here
  3796. 2:37:44I'm just going to print it okay and then
  3797. 2:37:47I'm going to print with last one
  3798. 2:37:49regression DOT score will be the best so
  3799. 2:37:52here I'm actually getting - 35 - 35 here
  3800. 2:37:56I'm actually getting - 32 so minus 35
  3801. 2:37:59still I will focus on linear regression
  3802. 2:38:01now see what will happen if I add more
  3803. 2:38:04parameters if I add more parameters see
  3804. 2:38:06what will happen so now I'm going to
  3805. 2:38:08take Alpha different different values
  3806. 2:38:10see this I'm just going to remove this
  3807. 2:38:13and probably add Alpha value in this
  3808. 2:38:16way see here I have added more values 5
  3809. 2:38:1910 20 30 35 40 45 100 okay let's see
  3810. 2:38:23whether we our performance will increase
  3811. 2:38:25or not so here
  3812. 2:38:28uh first of all let me remove from here
  3813. 2:38:32in Ridge just take it down guys I'm I'm
  3814. 2:38:35adding more parameters like this just
  3815. 2:38:36take it down yeah CV is equal to 5
  3816. 2:38:40nobody okay you're not able to see it um
  3817. 2:38:43CV is equal to 5 now here it is uh what
  3818. 2:38:46you can basically focus on so here you
  3819. 2:38:49can see I have added some values like
  3820. 2:38:51this you can also
  3821. 2:38:52add and just try to execute and now if I
  3822. 2:38:56go and probably see this is my see first
  3823. 2:38:59I have tried for Ridge I'm getting minus
  3824. 2:39:0229 do you see after adding more
  3825. 2:39:04parameters what happened in Ridge after
  3826. 2:39:07adding more parameters what happened in
  3827. 2:39:09Ridge you can see om minus 29 and the
  3828. 2:39:12alpha value that is got selected is 100
  3829. 2:39:14if you want try with cross validation
  3830. 2:39:1710 and just try to execute now
  3831. 2:39:20now so these are are some hyper
  3832. 2:39:22parameters that we will definitely play
  3833. 2:39:24with here you can see - 29 so here you
  3834. 2:39:27can see minus 29 you can also increase
  3835. 2:39:30the cross validation
  3836. 2:39:32value over here also and probably
  3837. 2:39:34execute it but with lasso I don't know
  3838. 2:39:38whether it is improving or not it is
  3839. 2:39:39coming to minus 34 you just have to play
  3840. 2:39:42with this parameters as now for a bigger
  3841. 2:39:45problem statement the thing is not
  3842. 2:39:47limited to here right we try to take
  3843. 2:39:49multiples and many parameters multiples
  3844. 2:39:52and many parameters and try to do these
  3845. 2:39:54things it is up to you we play with
  3846. 2:39:56multiple parameters whichever gives us
  3847. 2:39:58the best result we are basically taking
  3848. 2:40:00it it's okay error is increased I know
  3849. 2:40:03that no error is increasing definitely
  3850. 2:40:06error is increasing even though by
  3851. 2:40:08trying with different different
  3852. 2:40:09parameters but about most of the
  3853. 2:40:11scenario see here I gotus 37 probably
  3854. 2:40:14what I can actually do is that uh try to
  3855. 2:40:17get better one with respect to this
  3856. 2:40:20now the best way what I can also do is
  3857. 2:40:22that I can basically take up train and
  3858. 2:40:25test split also and probably do these
  3859. 2:40:27things let's see let's see one example
  3860. 2:40:29so how do we do train and test from SK
  3861. 2:40:32scalar dot I think model selection
  3862. 2:40:35import train test split okay it's okay
  3863. 2:40:38guys you may get a different value okay
  3864. 2:40:40let's do one thing okay let's make your
  3865. 2:40:42problem statement little bit simpler now
  3866. 2:40:45what I'm going to do just tell me in
  3867. 2:40:46train test plate what we need to do so
  3868. 2:40:48I'm going to take the same code I'm
  3869. 2:40:50going to paste it over here or let me do
  3870. 2:40:52one thing let me insert a cell below and
  3871. 2:40:55let me do it for train test split so in
  3872. 2:40:57train test plate what we can do so I'm
  3873. 2:41:00just going to take the syntax paste it
  3874. 2:41:02over here let's say that I'm taking XT
  3875. 2:41:04train y train and then I'm using train
  3876. 2:41:07test split with 33% now if I execute
  3877. 2:41:10with respect to X train and Y train so
  3878. 2:41:12here is my you can see this I have
  3879. 2:41:13written this code from SK learn. model
  3880. 2:41:15selection uh train test plate random
  3881. 2:41:17State can be anything whatever you write
  3882. 2:41:20it is fine then you basically give X and
  3883. 2:41:22Y with test sizes 33 uh this is
  3884. 2:41:25basically saying that the test will have
  3885. 2:41:2733% and the train data will be 77% so
  3886. 2:41:31this is what I'm actually getting with
  3887. 2:41:33respect to X train and Y train here what
  3888. 2:41:35I'm going to do I'm going to basically
  3889. 2:41:37take X train comma y train and now if I
  3890. 2:41:40go and probably see this here you can
  3891. 2:41:41see minus 25 understand this value
  3892. 2:41:44should go towards zero if it is going
  3893. 2:41:47towards zero that basically means the
  3894. 2:41:49performance is better now similarly I do
  3895. 2:41:52it for Ridge in Ridge what I'm actually
  3896. 2:41:54going to do here I'm going to write X
  3897. 2:41:56train and Y train and if I go and
  3898. 2:41:58probably select the best score than this
  3899. 2:42:00here you'll be able to see I'm getting
  3900. 2:42:03how much I'm getting minus
  3901. 2:42:062.47 okay here I'm getting
  3902. 2:42:0925.8 here 25. 47 that basically means
  3903. 2:42:12now still the Improvement is little bit
  3904. 2:42:15bad because here we are not going
  3905. 2:42:17towards zero so the next part again here
  3906. 2:42:20also you can basically do it for X train
  3907. 2:42:22and Y train X train and Y train so here
  3908. 2:42:25you have this one and let's go and
  3909. 2:42:27execute this so here you can see minus
  3910. 2:42:302.47 now what you can also do is that
  3911. 2:42:33you can use this
  3912. 2:42:35lasso regressor do predict and you can
  3913. 2:42:39basically predict with respect to X test
  3914. 2:42:42so this is your white test value suppose
  3915. 2:42:44let's say that this is my y PR Yore PR
  3916. 2:42:47then what I can do from SK
  3917. 2:42:50learn I will be using R square and
  3918. 2:42:53adjusted R square if you remember SK
  3919. 2:42:55learn R square r² so this is my R2 score
  3920. 2:43:00so where it is present in SK learn.
  3921. 2:43:02Matrix so I'm going to write from SK
  3922. 2:43:04learn import let's say I'm saying from
  3923. 2:43:08skarn do Matrix import r² R2 score now
  3924. 2:43:14what I'm going to do over here I'm
  3925. 2:43:16basically going to say my R2 score which
  3926. 2:43:20is my variable I'll say this is nothing
  3927. 2:43:22but R2 score here I'm just going to give
  3928. 2:43:24my y PR comma Yore test so if I go and
  3929. 2:43:28probably see the output here I will be
  3930. 2:43:30able to see print R2 score this is all I
  3931. 2:43:34have discussed guys there is also
  3932. 2:43:37adjusted rant score is there where is R2
  3933. 2:43:41R2 score one adjusted r² okay R2 score
  3934. 2:43:46is there but adjusted R square should be
  3935. 2:43:48here somewhere in some manner so this is
  3936. 2:43:52how your output looks like with respect
  3937. 2:43:53to by using this lasso regressor okay
  3938. 2:43:56which is very good okay it should be I
  3939. 2:43:59told it should be near 100% right now
  3940. 2:44:01I'm getting 67% if I want to tie with
  3941. 2:44:04the ridge you can also try that so you
  3942. 2:44:06can say Ridge regressor do predict and
  3943. 2:44:10here you can see 7 68% then you can also
  3944. 2:44:12try linear regressor and
  3945. 2:44:16predict what is the error saying the
  3946. 2:44:19regression is not fitted yet why why it
  3947. 2:44:22is not fitted why it is not
  3948. 2:44:25fitted let's say that I have fitted here
  3949. 2:44:28linear
  3950. 2:44:30regression dot fit on X train and Y
  3951. 2:44:33train X train and comma y train so I'm
  3952. 2:44:37just going to fit it now if I go and
  3953. 2:44:40probably try to do the
  3954. 2:44:41calculation so if I go and see my R2
  3955. 2:44:44score it is also coming somewhere around
  3956. 2:44:4668% 67% now since this is just a linear
  3957. 2:44:50regression you won't be able to get 100%
  3958. 2:44:52because you're drawing a straight line
  3959. 2:44:53right so for that you basically have to
  3960. 2:44:56other use other algorithms like XG boost
  3961. 2:44:58and all n bias so many algorithms are
  3962. 2:45:01there it's okay see you give y test over
  3963. 2:45:04here y PR over here both are same right
  3964. 2:45:06they're
  3965. 2:45:07comparing by see at one limit you can
  3966. 2:45:10you can increase the performance after
  3967. 2:45:12that you cannot see again I'm telling
  3968. 2:45:14you in linear regression what we do
  3969. 2:45:15these are my points right I will be only
  3970. 2:45:17able to create one best line I cannot
  3971. 2:45:19create a curve line right over here so
  3972. 2:45:21obviously my accuracy will be only
  3973. 2:45:23limited let's go and do it logistic
  3974. 2:45:26practical
  3975. 2:45:27quickly and here uh in logistic also we
  3976. 2:45:31can do git SE CV now what I'm actually
  3977. 2:45:34going to do first of all let's go ahead
  3978. 2:45:35with the data set so I will quickly
  3979. 2:45:38Implement logistic so from LC learn.
  3980. 2:45:41linear
  3981. 2:45:42model I'm going to import logistic
  3982. 2:45:46regression so I'm going to use logistic
  3983. 2:45:48regression and apart from that we know
  3984. 2:45:50that let's take a new data set because
  3985. 2:45:52for logistic we need to solve using
  3986. 2:45:54classification problem so this is
  3987. 2:45:56basically my logistic regression I'll
  3988. 2:45:58take one data set so from SK learn. data
  3989. 2:46:01sets import we'll take a data set which
  3990. 2:46:03is like uh breast cancer data set so
  3991. 2:46:05that is also present in SK learn with
  3992. 2:46:07respect to the breast cancer data set
  3993. 2:46:09I'm just going to use this see load best
  3994. 2:46:12cancer data set I'm loading it and all
  3995. 2:46:14the independent features are in data and
  3996. 2:46:16my columns are feature names the same
  3997. 2:46:18thing like how we did previously okay so
  3998. 2:46:20this will basically be my
  3999. 2:46:23complete uh complete independent feature
  4000. 2:46:25so if I go and probably see this x. head
  4001. 2:46:28here you'll be able to see that based on
  4002. 2:46:31this input features the independent
  4003. 2:46:32feature we need to determine whether the
  4004. 2:46:34person is having cancer or not these are
  4005. 2:46:37some of the features over here and this
  4006. 2:46:39is like many many features are actually
  4007. 2:46:40present so next thing I this that was my
  4008. 2:46:43independent feature now I'll take my
  4009. 2:46:45dependent feature dependent feature will
  4010. 2:46:47already present in DF Target okay this
  4011. 2:46:50particular data set that we have taken
  4012. 2:46:52in DF in DF do Target we will basically
  4013. 2:46:55have all our dependent feature these are
  4014. 2:46:56my independent features so what I'm
  4015. 2:46:58actually going to do I'm going to create
  4016. 2:46:59Y and I'm going to say PD do data frame
  4017. 2:47:04and here I'm going to say DF do Target
  4018. 2:47:07Target and this column name should be
  4019. 2:47:11Target right so this will be my column
  4020. 2:47:13name and now if I go and see my y y is
  4021. 2:47:16basically having zeros and one in the
  4022. 2:47:18target feature now the next thing that
  4023. 2:47:20we are going to do is that uh apply
  4024. 2:47:23basically apply the first of all we need
  4025. 2:47:26to check whether this data set is uh
  4026. 2:47:29this particular y column is balanced or
  4027. 2:47:31imbalanced okay in order to do that I
  4028. 2:47:33will just write F
  4029. 2:47:35Target if the data set is imbalanced
  4030. 2:47:38definitely we need to work on that and
  4031. 2:47:40try to perform upsampling so if I write
  4032. 2:47:42y target. Valore counts if I execute
  4033. 2:47:46this so here you'll be able to see that
  4034. 2:47:48value SC counts will basically give that
  4035. 2:47:50how many number of ones are and how many
  4036. 2:47:52number of zeros are so now total number
  4037. 2:47:54of ones are 357 and total number of
  4038. 2:47:57zeros are 22 so is this a imbalanced
  4039. 2:48:01data set probably this is a balanced
  4040. 2:48:03data set so here I'm actually going to
  4041. 2:48:04now do train test spit train test spit I
  4042. 2:48:08will try to do again train test spit how
  4043. 2:48:10do we do we can quickly do copy the same
  4044. 2:48:14thing entirely I'll copy this entirely
  4045. 2:48:16over here and then I will get my X and Y
  4046. 2:48:20so here is my X train X test y train y
  4047. 2:48:22test so train test plate obviously I'll
  4048. 2:48:24be doing it now in logistic regression
  4049. 2:48:26if I go and search for
  4050. 2:48:28logistic regression escalar I will be
  4051. 2:48:31able to see this what all parameters are
  4052. 2:48:33there this is basically the L1 Norm or
  4053. 2:48:35L2 Norm or L1 regularization or L2
  4054. 2:48:37regularization with respect to whatever
  4055. 2:48:39things we have discussed in logistic and
  4056. 2:48:41then the C value these two parameter
  4057. 2:48:43values are very much important if I
  4058. 2:48:45probably show you over here the penalty
  4059. 2:48:49what kind of penalty whether you want to
  4060. 2:48:50add L2 penalty L1 penalty you can use L2
  4061. 2:48:53or L1 the next thing is C this is
  4062. 2:48:56nothing but inverse of regularization
  4063. 2:48:57strength this basically says 1 by Lambda
  4064. 2:49:01something like that this parameter is
  4065. 2:49:02also very much important guys class
  4066. 2:49:04weight suppose if your data set is not
  4067. 2:49:06balanced at that point of time you can
  4068. 2:49:09apply weights to your classes if
  4069. 2:49:11probably your data set is balanced you
  4070. 2:49:14can directly use class weight is equal
  4071. 2:49:16to balanced other than that you can use
  4072. 2:49:18other other weight which you basically
  4073. 2:49:19want so this is specifically some of
  4074. 2:49:22this right no this is not Ridge or lasso
  4075. 2:49:25okay this is logistic in logistic also
  4076. 2:49:28you have L1 norm and L2
  4077. 2:49:30Norms understand probably I missed that
  4078. 2:49:32particular part in the theory but here
  4079. 2:49:35also you have an L2 penalty norm and L1
  4080. 2:49:37penalty Norm I probably did not teach
  4081. 2:49:39you in theory because if you look see
  4082. 2:49:43logistic regression can be learned by
  4083. 2:49:45two different ways one is through
  4084. 2:49:47probabilistic method and one is through
  4085. 2:49:49geometric method if you go and probably
  4086. 2:49:51see my video that is present with
  4087. 2:49:52respect to logistic regression right now
  4088. 2:49:54in my YouTube channel there I have
  4089. 2:49:56explained you about this L1 and L2 Norms
  4090. 2:49:58also over there so in this also it is
  4091. 2:49:59basically present it is a kind of
  4092. 2:50:01penalty again just for uh using for this
  4093. 2:50:05kind of classification problem so what
  4094. 2:50:08I'm actually going to do let's go and
  4095. 2:50:10play with the parameters that I am
  4096. 2:50:12looking at so I will play with two
  4097. 2:50:14parameters one is params C value here
  4098. 2:50:17I'm defining 1 10 20 anything that you
  4099. 2:50:20can Define one set of values you can
  4100. 2:50:22Define and there was one more parameter
  4101. 2:50:24which is called as Max iteration this is
  4102. 2:50:26specifically for grits or CV okay that
  4103. 2:50:28I'm specifically going to apply so I
  4104. 2:50:30will just try to execute this this will
  4105. 2:50:32be my params now I'm going to quickly
  4106. 2:50:34Define my model one which will be my
  4107. 2:50:36logistic regression model so my logistic
  4108. 2:50:39regression here by default one value
  4109. 2:50:41I'll give for C and Max itra let's say
  4110. 2:50:45I'm giving this value later on what I
  4111. 2:50:47will do for this model I'll apply it to
  4112. 2:50:49grid sear CV so I'm just going to say
  4113. 2:50:51grid s CV and I'm going to apply it for
  4114. 2:50:55model one param grid is equal to params
  4115. 2:50:59this parameter that I'm specifically
  4116. 2:51:01trying to apply since this is a
  4117. 2:51:02classification problem and I am not
  4118. 2:51:04pretty sure that whether true positive
  4119. 2:51:06is important or true negative is
  4120. 2:51:08important I'm going to use F1 scoring
  4121. 2:51:10okay F1 scoring is basically again the
  4122. 2:51:13parametric term which we discussed
  4123. 2:51:14yesterday which is nothing but
  4124. 2:51:16performance metrics and then I'm going
  4125. 2:51:18to use CV is equal to 5 so this will be
  4126. 2:51:21entirely my model with respect to grid s
  4127. 2:51:24CV and I'll be executing this then I
  4128. 2:51:27will do model. fit on my X train and Y
  4129. 2:51:32train data so once I execute it here you
  4130. 2:51:34can see all the output along with
  4131. 2:51:36warnings a lot of warnings will be
  4132. 2:51:38coming I don't know because this many
  4133. 2:51:40parameters are there and finally you can
  4134. 2:51:42see that this has got selected now if
  4135. 2:51:44you really want to find out what is your
  4136. 2:51:46best param score model
  4137. 2:51:49dot best params so here you can see Max
  4138. 2:51:52iteration as
  4139. 2:51:54150 and what you can actually do with
  4140. 2:51:58respect to your best score model do best
  4141. 2:52:03score is 95 percentage but still we want
  4142. 2:52:06to test it with test data so can we do
  4143. 2:52:09it yes we can definitely do it I'll say
  4144. 2:52:11model do core or I'll say model dot
  4145. 2:52:15predict on my X test data and this will
  4146. 2:52:18basically be my y red so this will be my
  4147. 2:52:21y red all the Y prediction that I'm
  4148. 2:52:23actually getting so if you go and see y
  4149. 2:52:26red so these are my ones and zeros with
  4150. 2:52:28respect to the Y
  4151. 2:52:30prediction at finally after getting the
  4152. 2:52:32prediction values I can apply confusion
  4153. 2:52:35Matrix I hope I have taught you about
  4154. 2:52:36confusion Matrix so from sklearn do
  4155. 2:52:39confusion Matrix sorry sklearn do metrix
  4156. 2:52:43I'm going to import confusion metrix
  4157. 2:52:46classification report and the next thing
  4158. 2:52:49that I would like to do is this two I
  4159. 2:52:52will try to import confusion Matrix and
  4160. 2:52:54classification report now if you want to
  4161. 2:52:56see the confusion Matrix with respect to
  4162. 2:52:58your I can just write
  4163. 2:53:00Yore frad or Yore test whatever you want
  4164. 2:53:04go ahead with it and this is basically
  4165. 2:53:06my confusion Matrix if I put this
  4166. 2:53:09forward no difference will be there only
  4167. 2:53:11this thing will be moving that also I
  4168. 2:53:13showed you 63 118 3 and 4 now finally if
  4169. 2:53:17I want to accuracy score I can also
  4170. 2:53:19import accuracy score over here so here
  4171. 2:53:21you can see accuracy score is imported I
  4172. 2:53:23can also find out my accuracy score
  4173. 2:53:25which is my the total accuracy with
  4174. 2:53:28respect to this I we can give y test and
  4175. 2:53:31Yore PR which we have discussed
  4176. 2:53:34yesterday this is giving
  4177. 2:53:3596% if you want detailed Precision
  4178. 2:53:38recall all the score then at that point
  4179. 2:53:40of time I can use this classification
  4180. 2:53:43report and here I can give white test
  4181. 2:53:45and wied here is what I'm actually
  4182. 2:53:47getting so here you can see with respect
  4183. 2:53:50to F1 F1 score Precision recall since
  4184. 2:53:52this is a balanced data set obviously
  4185. 2:53:54the performance will be best yes you can
  4186. 2:53:57also use Roc see I'll also show you how
  4187. 2:54:00to use Roc and probably you'll be able
  4188. 2:54:01to see this you have to probably
  4189. 2:54:03calculate false positive rate two
  4190. 2:54:05positive rate but don't worry about Roc
  4191. 2:54:07I will first of all explain you the
  4192. 2:54:08theoretical part now let's go ahead and
  4193. 2:54:10discuss about n bias n bias is an
  4194. 2:54:13important algorithm so here I'm just
  4195. 2:54:16going to go ahead so now let's go ahead
  4196. 2:54:18and discuss about na bias and here we
  4197. 2:54:21are going to discuss about the intuition
  4198. 2:54:23so na bias is an another amazing
  4199. 2:54:26algorithm which is specifically used for
  4200. 2:54:29classification and this specifically
  4201. 2:54:31works on something called as base
  4202. 2:54:34theorem now what exactly is base theorem
  4203. 2:54:36first of all we need to understand about
  4204. 2:54:38base theorem let's say that guys I have
  4205. 2:54:41base theorem let's say that I have an
  4206. 2:54:43experiment which is called as rolling a
  4207. 2:54:45dis now in rolling a dis how many number
  4208. 2:54:47of elements I have have so if I say what
  4209. 2:54:49is the probability of 1 then obviously
  4210. 2:54:51you'll be saying 1X 6 if I say
  4211. 2:54:53probability of two then also here you'll
  4212. 2:54:55say 1X 6 if I say probability of three
  4213. 2:54:58then I will definitely say it is 1x 6 so
  4214. 2:55:01here you know that this kind of events
  4215. 2:55:04are basically called as independent
  4216. 2:55:06events now rolling a dice why it is
  4217. 2:55:08called as an independent event because
  4218. 2:55:10getting one or two in every experiment
  4219. 2:55:12one is not dependent on two two is not
  4220. 2:55:14dependent on three so they are all
  4221. 2:55:16independent that is the reason why we
  4222. 2:55:18specifically say is an independent event
  4223. 2:55:20but if I take an example of dependent
  4224. 2:55:22events let's consider that I have a bag
  4225. 2:55:24of marbles okay in this marble I
  4226. 2:55:28basically have three red marbles and I
  4227. 2:55:31have two green marbles now tell me what
  4228. 2:55:33is the probability of suppose I have a
  4229. 2:55:36event in the first event I take out a
  4230. 2:55:38red marble so what is the probability of
  4231. 2:55:40taking out a red marble so here you can
  4232. 2:55:43definitely say that it is
  4233. 2:55:443x5 okay so this is my first event now
  4234. 2:55:47in the second event let's say that in
  4235. 2:55:49this you have taken out the red marble
  4236. 2:55:51now what is the second second time again
  4237. 2:55:53you are taking out the second red marble
  4238. 2:55:55or forget about second Rand marble now
  4239. 2:55:57you want to take out the green marble
  4240. 2:55:59now what is the probability with respect
  4241. 2:56:01to taking out a green marble so here
  4242. 2:56:03you'll be definitely saying that okay
  4243. 2:56:05one red marble has been removed then the
  4244. 2:56:07total number of marbles that are left
  4245. 2:56:09are four so here you can definitely
  4246. 2:56:11write that probability of getting a
  4247. 2:56:12green marble is nothing but 2x4 which is
  4248. 2:56:14nothing but 1x2 so here what is
  4249. 2:56:16happening first first element you took
  4250. 2:56:18out first marble that you took out first
  4251. 2:56:20event from from the first event you took
  4252. 2:56:21out red marble from the second event you
  4253. 2:56:23took out green marble this two are in
  4254. 2:56:25these two are dependent events because
  4255. 2:56:28the number of marbles are getting
  4256. 2:56:29reduced as you take out from them so if
  4257. 2:56:32I tell you what is the probability of
  4258. 2:56:35taking out a red marble and then a green
  4259. 2:56:39marble so it's the simple the formula
  4260. 2:56:42will be very much simple right which we
  4261. 2:56:43have already discussed in stats it is
  4262. 2:56:45nothing but probability of probability
  4263. 2:56:47of red multiplied by probability of
  4264. 2:56:50green given Red so this specific thing
  4265. 2:56:53is called as conditional probability
  4266. 2:56:55here understand what is happening
  4267. 2:56:57probability of green marble given the
  4268. 2:56:59red marble event has occurred here both
  4269. 2:57:01the events are independent now let me
  4270. 2:57:03write it down very nicely so I can write
  4271. 2:57:05probability of A and B is equal to
  4272. 2:57:08probability of a multiplied probability
  4273. 2:57:12of B divided by probability of a let's
  4274. 2:57:16go and derive something can can I write
  4275. 2:57:18probability of A and B is equal to
  4276. 2:57:21probability of b and a so answer is yes
  4277. 2:57:24we can definitely say we can definitely
  4278. 2:57:26say if you go and do the calculation
  4279. 2:57:27you'll be able to get the answer you
  4280. 2:57:29should not say no now what is the
  4281. 2:57:32formula for probability of A and B so
  4282. 2:57:34here you can basically write probability
  4283. 2:57:36of a multiplied by probability of B
  4284. 2:57:39given a if I take out probability of
  4285. 2:57:42green what is probability of green in
  4286. 2:57:43this particular case 2x 5 what is
  4287. 2:57:46probability of red 3x 4 for right now
  4288. 2:57:49let's consider this now this part I can
  4289. 2:57:51definitely write as this part I can
  4290. 2:57:54definitely write as probability of B
  4291. 2:57:56multiplied by probability of B
  4292. 2:58:00probability of B this one probability of
  4293. 2:58:02B and this will be probability of a
  4294. 2:58:04given B so I can definitely write this
  4295. 2:58:06much with respect to all this
  4296. 2:58:08information now can I derive probability
  4297. 2:58:10of a is equal to probability of B
  4298. 2:58:14multiplied by probability of a / B me
  4299. 2:58:18probability of a given B divided by
  4300. 2:58:21probability of sorry I'll write this as
  4301. 2:58:24probability of B given a divided by
  4302. 2:58:27probability of a and this is
  4303. 2:58:28specifically called as base theorem and
  4304. 2:58:31this is the Crux behind na bias
  4305. 2:58:34understand this is the Crux behind the
  4306. 2:58:35base theorem now let's go ahead and
  4307. 2:58:38let's discuss about how we are using
  4308. 2:58:40this to solve let's take some examples
  4309. 2:58:43and probably make you understand let's
  4310. 2:58:45say that I have some features like X1 X2
  4311. 2:58:49X3 X4 X5 like this till xn and I have my
  4312. 2:58:54output y so these are my independent
  4313. 2:58:56features these all are my independent
  4314. 2:58:58features these all are my independent
  4315. 2:59:00features so here I'm going to write
  4316. 2:59:02independent features and this is my
  4317. 2:59:04output feature which is also my
  4318. 2:59:05dependent feature now what is happening
  4319. 2:59:08if I say probability of b or a what does
  4320. 2:59:10this basically mean I need to really
  4321. 2:59:12find what is the probability of Y and
  4322. 2:59:15you know that guys I will have some
  4323. 2:59:17values over here and basically I'll have
  4324. 2:59:19some output value over here so based on
  4325. 2:59:21this input values I need to predict what
  4326. 2:59:23is the output initially on a training
  4327. 2:59:25data set I will have your input and then
  4328. 2:59:28your output initially my model will get
  4329. 2:59:30trained on this now let's consider what
  4330. 2:59:32this entire terminology is I will try to
  4331. 2:59:34write in terms of this equation so I
  4332. 2:59:36will say probability of Y given x1a x2a
  4333. 2:59:41X3 up till xn then this equation will
  4334. 2:59:44become probability of Y see probability
  4335. 2:59:46of Y given X X1 X2 X3 xn this a is
  4336. 2:59:50nothing but X1 X2 X3 xn and I'm trying
  4337. 2:59:52to find out what is the probability of Y
  4338. 2:59:54and then I will write probability of b b
  4339. 2:59:57is nothing but y but before that what
  4340. 3:00:00I'll write probability of a / B right a
  4341. 3:00:03given b or probability of B probability
  4342. 3:00:06of B is nothing but y multiplied by
  4343. 3:00:08probability of a given B probability of
  4344. 3:00:12a given B basically means probability of
  4345. 3:00:15x1a X2 comma xn and given b b is given
  4346. 3:00:20right so I'm able to find this entire
  4347. 3:00:22value now just a second I made some
  4348. 3:00:24mistakes I guess now it is correct sorry
  4349. 3:00:26I I just missed one term that is this
  4350. 3:00:29given y this is how it will become and
  4351. 3:00:32this will be equal to probability of a
  4352. 3:00:36that is X1 comma X2 like this up to XL
  4353. 3:00:39so probability of Y multiplied by
  4354. 3:00:41probability of a given y now if I try to
  4355. 3:00:44expand this then this will basically
  4356. 3:00:46become something like this see
  4357. 3:00:48probability of Y multiplied by
  4358. 3:00:51probability of X1 given yes a given y
  4359. 3:00:56sorry given y multiplied by probability
  4360. 3:01:00of X2 given y probability of x3 given Y
  4361. 3:01:06and like this it will be probability of
  4362. 3:01:08xn given y so this will also be y1 Y2 Y3
  4363. 3:01:12YN this I can expand it like this and
  4364. 3:01:15then this will basically become
  4365. 3:01:16probability of X Y 1 multiplied by
  4366. 3:01:18probability of X2 multiplied by
  4367. 3:01:21probability of x3 like this up to
  4368. 3:01:23probability of xn so this is with
  4369. 3:01:26respect to all the probability y will be
  4370. 3:01:28different see here for this particular
  4371. 3:01:30record y will be different for this y
  4372. 3:01:32will be different for this y will be
  4373. 3:01:34different but why output it may be yes
  4374. 3:01:37or no right it may be yes or no okay I
  4375. 3:01:41I'll solve a problem it will make
  4376. 3:01:43everything understand and this will
  4377. 3:01:45probably be probability of Y it can be
  4378. 3:01:47binary multiclass whatever things you
  4379. 3:01:49want I'll solve a problem in front of
  4380. 3:01:51you now let's say that I have my y as
  4381. 3:01:54let's say that I have a lot of features
  4382. 3:01:56X1 X2 X3 X X4 with respect to this let's
  4383. 3:02:02say in my one of my data set I have this
  4384. 3:02:03many x1s this many features and this is
  4385. 3:02:06my y so these are my feature number and
  4386. 3:02:09this is my y let's say that in y I have
  4387. 3:02:11yes or no so how I will probably write
  4388. 3:02:15we really need to understand this okay I
  4389. 3:02:17will basically
  4390. 3:02:18say what is the probability of Y is
  4391. 3:02:21equal to yes given this x of I this is
  4392. 3:02:25my first record first record of X of I
  4393. 3:02:27this is my second record of X of I so I
  4394. 3:02:30may write like this what is the
  4395. 3:02:31probability of Y being yes if x of I is
  4396. 3:02:34given to you X of I basically means X1
  4397. 3:02:37X2 X3 X4 so here you'll obviously write
  4398. 3:02:39what kind of equation you'll basically
  4399. 3:02:41say probability of yes multiplied by
  4400. 3:02:45probability of yes multiplied by
  4401. 3:02:46probability of X of 1 given
  4402. 3:02:50yes multiplied by probability of X2
  4403. 3:02:53given yes probability of x3 given yes
  4404. 3:02:58and probability of X4 given yes divided
  4405. 3:03:03by probability of X1 multiplied by
  4406. 3:03:06probability of X2 multiplied by
  4407. 3:03:08probability of x3 multiplied by
  4408. 3:03:10probability of X4 Y is fixed it may be
  4409. 3:03:13yes or it may be no but with respect to
  4410. 3:03:15different different records this value
  4411. 3:03:17may change similarly if I write
  4412. 3:03:18probability of Y is equal to no given X
  4413. 3:03:22of I what it will be then it will be
  4414. 3:03:26probability of no multiplied by
  4415. 3:03:30probability of X1 given no then
  4416. 3:03:33probability of
  4417. 3:03:35X2 given
  4418. 3:03:37no probability of
  4419. 3:03:39x3 given
  4420. 3:03:42no and probability of X4 given no so
  4421. 3:03:46here because every any input that I give
  4422. 3:03:49any input X of I that I give I may
  4423. 3:03:51either get yes or no so I need to find
  4424. 3:03:53both the probability so probability of
  4425. 3:03:54X1 multiplied by probability of X2
  4426. 3:03:57multiplied by probability of x3
  4427. 3:03:59multiplied by probability of X4 see with
  4428. 3:04:02respect to Any X of I the output can be
  4429. 3:04:05yes or no and I really need to find out
  4430. 3:04:07the probabilities so both the formula is
  4431. 3:04:09written over here what is the
  4432. 3:04:11probability of with respect to yes and
  4433. 3:04:13what is the probability with respect to
  4434. 3:04:14no now in this case one common thing you
  4435. 3:04:17see that this this denominator is fixed
  4436. 3:04:20this is definitely fixed it is fixed it
  4437. 3:04:22is it is not going to change for both of
  4438. 3:04:24them and I can consider that this is a
  4439. 3:04:27constant so what I can do I can
  4440. 3:04:30definitely ignore so here I can
  4441. 3:04:32definitely ignore these things ignore
  4442. 3:04:34this also ignore this Al because see
  4443. 3:04:36this is constant so I don't want to
  4444. 3:04:38consider this in the next time I'll just
  4445. 3:04:40use this specific formula to calculate
  4446. 3:04:42the probability now let's say that if my
  4447. 3:04:46first probability for a specific data
  4448. 3:04:49set yes of X of I is let's say that I'm
  4449. 3:04:52getting
  4450. 3:04:53as13 and similarly probability of no
  4451. 3:04:56with respect to X of I if I get
  4452. 3:05:0005 you know that in a binary
  4453. 3:05:02classification any values if it get
  4454. 3:05:04greater than or equal to 5 we are going
  4455. 3:05:06to consider it as 1 and if it is less
  4456. 3:05:09than 0.5 I'm going to consider it as
  4457. 3:05:10zero now I'm getting values like this 13
  4458. 3:05:13and .1 05 obviously I'm getting .13 05
  4459. 3:05:18so we do something called as
  4460. 3:05:20normalization it says that if I really
  4461. 3:05:23want to find out the probability of X
  4462. 3:05:24with X of I if I do normalization it is
  4463. 3:05:27nothing but .13 divided by .13 +
  4464. 3:05:3105 72 this is nothing but
  4465. 3:05:3572% and similarly if I do for
  4466. 3:05:37probability of no given X of I here
  4467. 3:05:39obviously it will say 1 - 72 which will
  4468. 3:05:42be your remaining answer that is 28
  4469. 3:05:44which is nothing but 28% so your final
  4470. 3:05:47answer will be this one this formulas
  4471. 3:05:49you have to remember now we'll solve a
  4472. 3:05:50problem let's solve a problem this will
  4473. 3:05:52be a very very interesting problem let's
  4474. 3:05:54say I have a data set which has like
  4475. 3:05:56this feature day let me just copy this
  4476. 3:05:59data set okay for you all now in this
  4477. 3:06:02data set I want to take out some
  4478. 3:06:04information let's take out Outlook
  4479. 3:06:08table now based on this output Outlook
  4480. 3:06:11feature see over here Outlook my day
  4481. 3:06:14outlook temperature humidity wind are
  4482. 3:06:17the input features independent feature
  4483. 3:06:19this is my output feature this one that
  4484. 3:06:22you are probably seeing play tennis is
  4485. 3:06:24my output feature which is specifically
  4486. 3:06:26a binary
  4487. 3:06:27classification so what I'm actually
  4488. 3:06:29going to do I'm basically going to take
  4489. 3:06:31my Outlook feature and based on this
  4490. 3:06:33Outlook feature I will just try to
  4491. 3:06:34create a smaller table which will give
  4492. 3:06:36some information now based on Outlook
  4493. 3:06:39first of all try to find out how many
  4494. 3:06:40categories are there in Outlook one is
  4495. 3:06:43sunny one is
  4496. 3:06:45overcast and one is rain right three
  4497. 3:06:48categories are there so I'm going to
  4498. 3:06:50write it down over here Sunny overcast
  4499. 3:06:53and rain so these three are my features
  4500. 3:06:56with respect to Sunny uh with Outlook I
  4501. 3:06:58have three categories one is sunny one
  4502. 3:07:00is overcast and one is RA here I'm going
  4503. 3:07:02to basically say with respect to Sunny
  4504. 3:07:05how many yes are there and how many no
  4505. 3:07:08are there and what is the probability of
  4506. 3:07:11yes and probability of no so I'm going
  4507. 3:07:13to again write it over here so this is
  4508. 3:07:16my Outlook feature
  4509. 3:07:18and then I have categories first yes no
  4510. 3:07:23Sunny overcast rain yes no then
  4511. 3:07:28probability of yes and probability of no
  4512. 3:07:31now the next thing that we need to find
  4513. 3:07:33out is that with respect to Sunny how
  4514. 3:07:37many of them are yes see yes we have so
  4515. 3:07:40when we have sunny over here the answer
  4516. 3:07:42is no so I will increase the count over
  4517. 3:07:44here one then again I have sunny again
  4518. 3:07:47answer is no so I'm going to increase
  4519. 3:07:49the count to two with this sunny this is
  4520. 3:07:52basically no okay so again I'm going to
  4521. 3:07:54increase the count to three now with
  4522. 3:07:56sunny how many of them are yes one and
  4523. 3:08:00two so I have this one and this one so I
  4524. 3:08:03have two so I'm going to say with
  4525. 3:08:05respect to Sunny I have two
  4526. 3:08:07yes understand Outlook is my X1 X1
  4527. 3:08:11feature let's consider now the next
  4528. 3:08:13thing is that let's see with respect to
  4529. 3:08:16overcost with overcast how many of them
  4530. 3:08:18are yes so this overcast is there yes 1
  4531. 3:08:222 3 and four so total four yes are there
  4532. 3:08:26with respect to overcast then with
  4533. 3:08:28respect to overcast how many are on no
  4534. 3:08:31you can go ah and find out it is
  4535. 3:08:32basically zero NOS then with respect to
  4536. 3:08:35rain how many of them are yes so here
  4537. 3:08:37you can see with respect to one rain yes
  4538. 3:08:40yes no no so this is nothing but 3 2
  4539. 3:08:46let's try to find out there are three is
  4540. 3:08:47two or
  4541. 3:08:48not one here also one yes is there right
  4542. 3:08:52so 3 yes two NOS so the total number of
  4543. 3:08:55yes and NOS if you count it there are
  4544. 3:08:58nine yes and five NOS this is my total
  4545. 3:09:01count so if you totally count this 9 + 5
  4546. 3:09:04is 14 you'll be able to compare that
  4547. 3:09:06there will be 9 yes and five NOS what is
  4548. 3:09:08the probability of yes when Sunny is
  4549. 3:09:10given so here you have 2X 9 here you
  4550. 3:09:14have 4X 9 here you have 3x 9 now if if I
  4551. 3:09:17say what is the probability of no given
  4552. 3:09:20Sunny now see probability of yes given
  4553. 3:09:23Sunny probability of yes given forecast
  4554. 3:09:26probability of yes given rain so it is
  4555. 3:09:28basically that I will just try to write
  4556. 3:09:30it in a simpler manner so that you'll
  4557. 3:09:31not get confused okay so this is my
  4558. 3:09:33probability of yes and this is my
  4559. 3:09:35probability of no but understand what
  4560. 3:09:37does this basically mean this
  4561. 3:09:39terminology basically means probability
  4562. 3:09:41of yes given Sunny probability of yes
  4563. 3:09:44given overcast probability of yes given
  4564. 3:09:46rain similarly what is probability of no
  4565. 3:09:49probability of no obviously you know
  4566. 3:09:50that 3x 5 is my first probability then
  4567. 3:09:54you have 0x 5 and then you have 2X 5 now
  4568. 3:09:58with respect to the next feature let's
  4569. 3:10:00consider that I'm going to consider one
  4570. 3:10:01more feature and in this feature I will
  4571. 3:10:03say let's consider
  4572. 3:10:05temperature okay let's consider
  4573. 3:10:07temperature now in temperature how many
  4574. 3:10:10features I have or how many categories I
  4575. 3:10:12have I have hot you can see hot mild and
  4576. 3:10:17and cold now with respect to hot mild
  4577. 3:10:19cold here also I will be having yes no
  4578. 3:10:23probability of yes and probability of no
  4579. 3:10:26now try to find out with respect to hot
  4580. 3:10:28how many are yes so here no is there
  4581. 3:10:31here also no is there two NOS uh 1 yes
  4582. 3:10:36uh 2 yes so two yes and two NOS probably
  4583. 3:10:39then similarly with respect to mild mild
  4584. 3:10:42how many are there 1 yes 1 No 2 yes 3s
  4585. 3:10:484s 4S and two knows okay so here you
  4586. 3:10:51basically go and calculate 4 yes and two
  4587. 3:10:54knows with respect to cold how many are
  4588. 3:10:57there cool cool or cold 1 yes 1 No 2 yes
  4589. 3:11:033 S 3 S and 1 no so here I have
  4590. 3:11:07specifically have 3s and 1 no again the
  4591. 3:11:10total number is 9 and five which will be
  4592. 3:11:12equal to the same thing that what we
  4593. 3:11:15have got now really go ahead with
  4594. 3:11:16finding probability of yes given hot so
  4595. 3:11:19it will be 2x 9 over here then here it
  4596. 3:11:22will be how much 4X 9 here it will be 3x
  4597. 3:11:269 again here what will be the
  4598. 3:11:28probability of no given given hot so
  4599. 3:11:31it'll be 2x 5 2x 5 1X 5 so this two
  4600. 3:11:36tables has already been created and
  4601. 3:11:37finally with respect to play the total
  4602. 3:11:39number of plays are yes is 9 no is five
  4603. 3:11:44and the answer is total 14 if if I say
  4604. 3:11:47what is the probability of yes only yes
  4605. 3:11:50then it is nothing but 9 by4 what is the
  4606. 3:11:54probability of no it is nothing but
  4607. 3:11:565x4 okay so this two values also you
  4608. 3:11:59require now let's say that you get a new
  4609. 3:12:02data set you need get a new data set
  4610. 3:12:05let's say you get a new test data where
  4611. 3:12:08it says that suppose if you are having
  4612. 3:12:11sunny and hot tell me what is the output
  4613. 3:12:16so this is my problem statement so let
  4614. 3:12:18me write it down so here I will write
  4615. 3:12:20probability of yes given Sunny comma hot
  4616. 3:12:25then here I will write probability of
  4617. 3:12:27yes multiplied by probability of so here
  4618. 3:12:31I will write probability of Sunny given
  4619. 3:12:34yes multiplied by probability of hot
  4620. 3:12:38given yes divided by what is it
  4621. 3:12:42probability of Sunny multiplied by
  4622. 3:12:45probability of hot
  4623. 3:12:50equation because it is a
  4624. 3:12:52constant because probability of no also
  4625. 3:12:55I'll be getting the same value 9 by4 so
  4626. 3:12:58probability of yes I'm going to replace
  4627. 3:13:00it with 9
  4628. 3:13:02by4 multiplied by 2x 9 then probability
  4629. 3:13:06of hot given yes so I am going to get 2
  4630. 3:13:09by 9 so
  4631. 3:13:12here 99 cancel or 2 1 7 then this is
  4632. 3:13:17nothing but 2 by
  4633. 3:13:216331 I read this statement little bit
  4634. 3:13:23wrong it should be probability of Sunny
  4635. 3:13:25given yes now go ahead and calculate go
  4636. 3:13:28ahead and calculate what is probability
  4637. 3:13:30of no given sunny and hot so here you
  4638. 3:13:33have probability of no multiplied by
  4639. 3:13:36probability of Sunny given
  4640. 3:13:38no multiplied by probability of hot
  4641. 3:13:43given
  4642. 3:13:44no divided by probability of Sunny
  4643. 3:13:50multiplied by probability of heart this
  4644. 3:13:53will get cancelled denominator is a
  4645. 3:13:55constant guys this is a constant so what
  4646. 3:13:58is probability of no so probability of
  4647. 3:14:00no is nothing but 5 by4 so I will write
  4648. 3:14:03over here 5 by4 multiplied by
  4649. 3:14:07probability of Sunny given no what is
  4650. 3:14:09probability of Sunny given no what is
  4651. 3:14:11probability of Sunny given no is nothing
  4652. 3:14:13but probability of Sunny given no is
  4653. 3:14:15nothing but 3x 5 so here I'm going to
  4654. 3:14:17get 3x 5 multiplied probability of H
  4655. 3:14:22given no that is nothing but 2x 5 so 2x
  4656. 3:14:255 is here 3x 5 is there five and five
  4657. 3:14:28will get cancelled 2 1 2 7 and then I'm
  4658. 3:14:32getting 3x 35 which is nothing but
  4659. 3:14:35calculator uh if I'm actually getting
  4660. 3:14:37three ID by 35 it's nothing but
  4661. 3:14:41857 I will write it down again
  4662. 3:14:44probability of yes given Sunny comma hot
  4663. 3:14:49which is my independent feature is
  4664. 3:14:51nothing but
  4665. 3:14:52031
  4666. 3:14:54031 and this is probability of no given
  4667. 3:14:57Sunny comma hot 85 now we'll try to
  4668. 3:15:00normalize this 85 + Point divided by 031
  4669. 3:15:06+ 085 73 this is nothing but 73% and
  4670. 3:15:11here I can basically say 1 -73 which is
  4671. 3:15:14my27 which is nothing but 27% if the
  4672. 3:15:18input comes as sunny and hot if the
  4673. 3:15:21weather is sunny and hot what will the
  4674. 3:15:23person do whether he will play or not
  4675. 3:15:26the answer is no okay now my next
  4676. 3:15:29question will be that if your new data
  4677. 3:15:31is overcast and Mild now tell me what
  4678. 3:15:34will be the probability using name bias
  4679. 3:15:37now you can add any number of features
  4680. 3:15:39let's say that I will say that okay
  4681. 3:15:42let's let's say that I will I will
  4682. 3:15:44probably say we can consider humidity
  4683. 3:15:47mind wind also you basically create this
  4684. 3:15:49kind of table to find it out but this
  4685. 3:15:50will be an assignment just do
  4686. 3:15:53it overcast and Mild if it is with
  4687. 3:15:56respect to NB try to solve it so the
  4688. 3:15:58second algorithm that we are going to
  4689. 3:16:00discuss about is something called as KNN
  4690. 3:16:02algorithm KNN algorithm is a very simple
  4691. 3:16:05problem statement okay which can be used
  4692. 3:16:09to solve both classification and
  4693. 3:16:11regression so KNN basically means K
  4694. 3:16:14nearest neighbor let's first of all
  4695. 3:16:16discuss about classification problem
  4696. 3:16:18number one classification problem let's
  4697. 3:16:20say that I have a binary classification
  4698. 3:16:22problem which looks like this I have two
  4699. 3:16:23data points like this one and this is
  4700. 3:16:26another one suppose a new data point
  4701. 3:16:29suppose a new data point which comes
  4702. 3:16:31over
  4703. 3:16:32here then how do I say that whether this
  4704. 3:16:35belongs to this category or whether it
  4705. 3:16:36belongs to this category if I probably
  4706. 3:16:38create a logistic regression I may
  4707. 3:16:40divide a line but in this particular
  4708. 3:16:42scenario how do we Define or how do we
  4709. 3:16:44come to a conclusion that
  4710. 3:16:47whether this will belong to this
  4711. 3:16:48category or this category so for here we
  4712. 3:16:50basically use something called as K
  4713. 3:16:52nearest neighbor let's say that I say
  4714. 3:16:55that my K value is five so what it is
  4715. 3:16:57going to do it is going to basically
  4716. 3:16:58take the five nearest closest point
  4717. 3:17:01let's say from this you have two nearest
  4718. 3:17:03closest point and from here you have
  4719. 3:17:05three nearest closest point so here we
  4720. 3:17:07basically see from the distance the
  4721. 3:17:09distance that which is my nearest point
  4722. 3:17:11now in this particular case you see that
  4723. 3:17:13maximum number of points are from Red
  4724. 3:17:15categories from Red from Red categories
  4725. 3:17:18I'm getting three points and from White
  4726. 3:17:21categories I'm getting two points now in
  4727. 3:17:23this particular scenario maximum number
  4728. 3:17:25of categories from where it is coming we
  4729. 3:17:27basically categorize that into that
  4730. 3:17:29particular class just with the help of
  4731. 3:17:30distance which all distance we
  4732. 3:17:31specifically use we use two distance one
  4733. 3:17:33is ukan distance and the other one is
  4734. 3:17:36something called as Manhattan distance
  4735. 3:17:37so ukan and Manhattan distance now what
  4736. 3:17:40does ukan distance basically say suppose
  4737. 3:17:42if this is your two points which is
  4738. 3:17:44denoted by X1 y1
  4739. 3:17:47X2 Y2 ukine distance in order to
  4740. 3:17:50calculate we apply a formula which looks
  4741. 3:17:52like this X2 - X1 s + Y2 - y1 s whereas
  4742. 3:17:58in the case of magetan distance suppose
  4743. 3:18:00this are my two points then we calculate
  4744. 3:18:03the distance in this way we calculate
  4745. 3:18:05the distance from here then here right
  4746. 3:18:07this is the distance we calculate we
  4747. 3:18:09don't calculate the hypothenuse distance
  4748. 3:18:10so this is the basic difference between
  4749. 3:18:11ukan and magetan distance now you may be
  4750. 3:18:14thinking Chris then fine that is for
  4751. 3:18:15classification problem for regression
  4752. 3:18:17what do we do for regression also it is
  4753. 3:18:19very much simple suppose I have all the
  4754. 3:18:22data points which looks like this now
  4755. 3:18:24for a new data point like this if I want
  4756. 3:18:26to calculate then we basically take up
  4757. 3:18:28the nearest Five Points let's say my K
  4758. 3:18:30is five k is a hyper parameter which we
  4759. 3:18:33play now suppose let's say that K it
  4760. 3:18:35finds the nearest point over here here
  4761. 3:18:38here here and here so if we need to find
  4762. 3:18:42out the point for this particular output
  4763. 3:18:44with respect to the K is equal to 5 it
  4764. 3:18:46will try to calculate the average of all
  4765. 3:18:48the points once it calculates the
  4766. 3:18:51average of all the points that becomes
  4767. 3:18:53your output so regression and
  4768. 3:18:55classification that is the only
  4769. 3:18:56difference because this K is actually an
  4770. 3:18:58hyper parameter we try with K is equal
  4771. 3:19:00to 1 to 50 and then we probably try to
  4772. 3:19:03check the error rate and if the error
  4773. 3:19:06rate is less then only we select the
  4774. 3:19:08model now two more things with respect
  4775. 3:19:10to K nearish neighbor K nearest neighbor
  4776. 3:19:12works very bad with respect to two
  4777. 3:19:15things one is outliers and and one is
  4778. 3:19:17imbalanced data set now if I have an
  4779. 3:19:19outlier let's say I have an outlier over
  4780. 3:19:22here this is one of my categories like
  4781. 3:19:24this and this is my another category
  4782. 3:19:26let's consider that I have some outliers
  4783. 3:19:28which looks like this now if I'm trying
  4784. 3:19:29to find out the point for this you can
  4785. 3:19:32see that the nearest point is basically
  4786. 3:19:35blue only and it belongs to the blue
  4787. 3:19:37category but because this outlier you
  4788. 3:19:39know it'll consider that the nearest
  4789. 3:19:40neighbor is this so then this will be
  4790. 3:19:42basically treated in this group only
  4791. 3:19:44formula for Manhattan distance it uses
  4792. 3:19:46modulus X2 - X1 + Y2 - y1 mode X2 - X1
  4793. 3:19:53Y2 - y1 uh this was it from my side guys
  4794. 3:19:55and yes I've also made detailed videos
  4795. 3:19:57about whatever topics we have discussed
  4796. 3:19:59today you can directly go and search for
  4797. 3:20:01that particular
  4798. 3:20:03topic so this is the agenda of this
  4799. 3:20:06session we will try to complete this all
  4800. 3:20:08things again here we are going to
  4801. 3:20:10understand the mathematical equations
  4802. 3:20:12and all uh so today's session we are
  4803. 3:20:14basically going to discuss about uh
  4804. 3:20:16decision tree okay and uh in this
  4805. 3:20:20session we are going to basically
  4806. 3:20:21understand what is the exact purpose of
  4807. 3:20:23decision tree with the help of decision
  4808. 3:20:25tree you are actually solving two
  4809. 3:20:27different problems one is regression and
  4810. 3:20:30the other one is
  4811. 3:20:32classification so we'll try to
  4812. 3:20:34understand both this particular part
  4813. 3:20:37well we will take a specific data set
  4814. 3:20:38and try to solve those problems now
  4815. 3:20:40coming to the decision tree one thing
  4816. 3:20:42you need to understand I'll say that if
  4817. 3:20:45age is less than 8 let's say I'm writing
  4818. 3:20:48this condition if age is less than or
  4819. 3:20:51equal to 18 I'm going to say print go to
  4820. 3:20:55college here I'm printing print college
  4821. 3:20:58and then I'll write else if age is
  4822. 3:21:02greater than 18 and pag is less than or
  4823. 3:21:05equal to 35 I'll say print work then
  4824. 3:21:09again I'll write else if age is let me
  4825. 3:21:12let me put this condition little bit
  4826. 3:21:14better then I'll write here L if if age
  4827. 3:21:17is greater than 18 and age is less than
  4828. 3:21:22or equal to 35 I'm going to say print
  4829. 3:21:25work basically people needs to work in
  4830. 3:21:27this age else I'm just going to consider
  4831. 3:21:30print retire so here is my ifls
  4832. 3:21:34condition over here now whenever we have
  4833. 3:21:36this kind of nested if Els condition
  4834. 3:21:38what we can do is that we can also
  4835. 3:21:40represent this in the form of decision
  4836. 3:21:42trees we'll also we can actually form
  4837. 3:21:45this in the form of decision and the
  4838. 3:21:46decision tree here first of all we will
  4839. 3:21:48have a specific root node let's say this
  4840. 3:21:51is my root node now in this root node
  4841. 3:21:52the first condition is less than or
  4842. 3:21:54equal to 18 so here obviously I will be
  4843. 3:21:56having two conditions saying that if it
  4844. 3:21:59is less than or equal to 18 and one
  4845. 3:22:02condition will be yes one condition will
  4846. 3:22:03be no so if this is yes and if this is
  4847. 3:22:06no right if this condition is true that
  4848. 3:22:09basically means we'll go in this side if
  4849. 3:22:11it is true then here we will basically
  4850. 3:22:14have something like college so this is
  4851. 3:22:17your Leaf node similarly when I have no
  4852. 3:22:22okay no no in this particular case we
  4853. 3:22:24will go to the next condition in this
  4854. 3:22:26next condition I will again create a
  4855. 3:22:28node and I'll say that okay this is less
  4856. 3:22:30than 18 and greater than sorry less than
  4857. 3:22:33or equal to 35 so if this is also there
  4858. 3:22:38then again I'll have two conditions
  4859. 3:22:39which is basically yes or no now when I
  4860. 3:22:42create this yes or no over here you'll
  4861. 3:22:43be able to see that basically means here
  4862. 3:22:46again two condition will be there if it
  4863. 3:22:48is yes I will say print work so this
  4864. 3:22:50will again be my leaf
  4865. 3:22:52node and again for no again I will do
  4866. 3:22:55the further splitting which is retire so
  4867. 3:22:59here you can see that this entire
  4868. 3:23:00algorithm this entire code that I have
  4869. 3:23:02actually written you can see that it has
  4870. 3:23:05got converted to this kind of
  4871. 3:23:08trees where you specifically able to
  4872. 3:23:10take decisions yes or no so can we solve
  4873. 3:23:15a classification
  4874. 3:23:17problem sorry this is greater than 18
  4875. 3:23:21again if it is greater than 18 and less
  4876. 3:23:23than or 35 so can we solve a
  4877. 3:23:28regression and a classification problem
  4878. 3:23:31regression and classification problem
  4879. 3:23:34using this decision trees by creating
  4880. 3:23:37this kind of
  4881. 3:23:38nodes so in short whenever we talk about
  4882. 3:23:41decision
  4883. 3:23:42trees whenever we talk about decision
  4884. 3:23:45trees
  4885. 3:23:47you will be seeing that decision trees
  4886. 3:23:49are nothing but decision trees are
  4887. 3:23:52nothing but by using this nested if El
  4888. 3:23:56condition we can definitely solve some
  4889. 3:23:58specific problem statement but here in
  4890. 3:24:00the visualized way we will specifically
  4891. 3:24:02create this decision tree in the form of
  4892. 3:24:04nodes now you need to understand that
  4893. 3:24:07what type of maths we will probably use
  4894. 3:24:10okay so let's do one thing let's take a
  4895. 3:24:12specific data set which I will
  4896. 3:24:14definitely do it over here in front of
  4897. 3:24:15you
  4898. 3:24:17okay and we will try to solve this
  4899. 3:24:18particular data set and this will
  4900. 3:24:20basically give you an idea like how we
  4901. 3:24:23can probably solve these problems so uh
  4902. 3:24:26let me just open my snippet tool so this
  4903. 3:24:29is my data set that I have let's
  4904. 3:24:31consider that I have this specific data
  4905. 3:24:33set now this data set are pretty much
  4906. 3:24:35important because this probably in
  4907. 3:24:39research papers also probably people who
  4908. 3:24:41have come up with this algorithm they
  4909. 3:24:43usually take this they take this thing
  4910. 3:24:46but but right now this particular
  4911. 3:24:47problem statement if I talk about this
  4912. 3:24:49is a classification problem statement
  4913. 3:24:51okay but don't worry I will also help
  4914. 3:24:53you to explain I'll also explain you
  4915. 3:24:56about regression also how decision tree
  4916. 3:24:58regression will definitely work so let's
  4917. 3:25:01go ahead and let's try to understand
  4918. 3:25:03suppose if I have this specific problem
  4919. 3:25:05statement how do we solve this this is
  4920. 3:25:07my output feature play tennis yes or no
  4921. 3:25:10okay whether the person is going to pay
  4922. 3:25:12tennis or not yesterday or there after
  4923. 3:25:14yesterday or whenever you want so if I
  4924. 3:25:17have this input features like Outlook
  4925. 3:25:19temperature humidity and wind is the
  4926. 3:25:22person going to play tennis or not this
  4927. 3:25:24is what my model should predict with the
  4928. 3:25:26help of decision tree so how decision
  4929. 3:25:28tree will work in this particular case
  4930. 3:25:29first of all let's consider any any any
  4931. 3:25:33specific uh feature let's say that
  4932. 3:25:35Outlook is my feature so this will be my
  4933. 3:25:37first
  4934. 3:25:38feature which is specifically Outlook
  4935. 3:25:41now just tell me how many are basically
  4936. 3:25:45having no and how many are basically
  4937. 3:25:48having yes in the case of Outlook there
  4938. 3:25:51you'll be able to find out there are
  4939. 3:25:52nine yes see 1 2 3 4 5 6 7 8 9 and how
  4940. 3:25:58many NOS are there 1 2 3 4 5 I think 1 2
  4941. 3:26:043 4 5 so nine yes and five NOS what we
  4942. 3:26:09are going to do in this specific thing
  4943. 3:26:11now we have N9 yes and five Nos and the
  4944. 3:26:13first node that I have actually taken
  4945. 3:26:17is basically Outlook so Outlook feature
  4946. 3:26:20now just try to find out we are focusing
  4947. 3:26:22on this specific feature now in this
  4948. 3:26:24feature how many categories I have I
  4949. 3:26:26have one Sunny category you can see over
  4950. 3:26:29here I have Sunny one category then I
  4951. 3:26:31have another category called as
  4952. 3:26:33overcast then I have another category as
  4953. 3:26:37rain so I have three unique categories
  4954. 3:26:40So based on these three categories I
  4955. 3:26:42will try to create three nodes so here
  4956. 3:26:45is my one node here is my second node
  4957. 3:26:49here is my third node so these are my
  4958. 3:26:52three categories so this category is
  4959. 3:26:53basically called as Sunny this category
  4960. 3:26:57is basically called as overcast and this
  4961. 3:27:00category is basically called as rain
  4962. 3:27:03based on these three categories so I'm
  4963. 3:27:04splitting it now just go ahead and see
  4964. 3:27:07in Sunny how many yes and how many no
  4965. 3:27:10are there how many yes with respect to
  4966. 3:27:12Sunny are there see in sunny I have two
  4967. 3:27:14NOS see one and two no uh one more no is
  4968. 3:27:18there three NOS so here you can see this
  4969. 3:27:21is my one no then this is my two no this
  4970. 3:27:25is my three no and yes are two so this
  4971. 3:27:30one and this one so how many total
  4972. 3:27:33number of yes so here you can see that
  4973. 3:27:36there are 1 2 2 yes and three no let's
  4974. 3:27:41say that I have randomly selected one
  4975. 3:27:43feature which is Outlook why can't I
  4976. 3:27:45when like see it is up to it it is up to
  4977. 3:27:49the decision tree to select any of the
  4978. 3:27:51feature here I have specifically taken
  4979. 3:27:53Outlook later on I'll explain why it it
  4980. 3:27:57can basically select how it selects the
  4981. 3:27:59feature okay I'll I'll talk about it
  4982. 3:28:00don't worry so in the Outlook we have
  4983. 3:28:04two yes sorry in the case of Sunny we
  4984. 3:28:06have two yes and three NOS now the next
  4985. 3:28:08thing is that let's go and see for
  4986. 3:28:10overcast in overcast I have 1 yes uh 2s
  4987. 3:28:14um 3s and 4 yes I don't have any no in
  4988. 3:28:18overcast so over here my thing will be
  4989. 3:28:21that four yes and Zer Nos and then
  4990. 3:28:24finally when we go to the Rain part see
  4991. 3:28:26in Rain how many features are there in
  4992. 3:28:29rain if you go and probably see it how
  4993. 3:28:31many number of yes and NOS are there go
  4994. 3:28:33and see in one one yes in row rain two
  4995. 3:28:36yes then one no then again you have one
  4996. 3:28:39yes and one no right so here you can
  4997. 3:28:43basically say that in rain in the case
  4998. 3:28:45of rain if I take a as an example how
  4999. 3:28:47many number of yes and NOS are there it
  5000. 3:28:49will be 3 yes and two
  5001. 3:28:52NOS understand understanding
  5002. 3:28:57algorithm then everything will you'll be
  5003. 3:29:00able to understand now let's go ahead
  5004. 3:29:03and try to cease for sunny sunny
  5005. 3:29:05definitely has 2 yes and three NOS this
  5006. 3:29:08has four yes and zero NOS here you have
  5007. 3:29:10three Y and two NOS now if I probably
  5008. 3:29:13take overcast here you need to
  5009. 3:29:15understand understand about two things
  5010. 3:29:17one is pure
  5011. 3:29:18split and one is impure split now what
  5012. 3:29:22does pure split basically mean pure spit
  5013. 3:29:25basically means that now see in this
  5014. 3:29:26particular scenario in overcast in
  5015. 3:29:29overcast I have either yes or no so here
  5016. 3:29:32you can see that I have four yes and Zer
  5017. 3:29:35NOS so that basically means this is a
  5018. 3:29:37pure split anybody tomorrow in my data
  5019. 3:29:40set if I just take this Outlook feature
  5020. 3:29:43suppose in one day in day 15 the Outlook
  5021. 3:29:46is Outlook is basically overcast then I
  5022. 3:29:50know directly it is the person is going
  5023. 3:29:52to play so this part is already created
  5024. 3:29:54and this node is called as pure
  5025. 3:29:58node understand this why it is called as
  5026. 3:30:00pure node because either you have all
  5027. 3:30:03Yes or zeros NOS or zero yes or all NOS
  5028. 3:30:08like that in this particular case I have
  5029. 3:30:10all yes so if I take this specific path
  5030. 3:30:13I know that with respect to overcast my
  5031. 3:30:16final decision which is yes it is always
  5032. 3:30:17going to become yes so this is what it
  5033. 3:30:19basically says so I don't have to split
  5034. 3:30:22further so from here I will probably not
  5035. 3:30:25split I will definitely not split more
  5036. 3:30:28because I don't require it because I
  5037. 3:30:31have it is a pure leaf node okay you can
  5038. 3:30:34also say that this is a pure leaf node
  5039. 3:30:37so I'm just going to mention it again
  5040. 3:30:39this one I'm specifically talking about
  5041. 3:30:41now let's talk about sunny in the case
  5042. 3:30:43of Sunny you have two yes and three NOS
  5043. 3:30:45so this is obviously impure so what we
  5044. 3:30:48do we take next feature and again how do
  5045. 3:30:52we calculate that which feature we
  5046. 3:30:54should take next I'll discuss about it
  5047. 3:30:56let's say that after this I take up
  5048. 3:31:00temperature I take up temperature and I
  5049. 3:31:02start splitting again since this is
  5050. 3:31:04impure okay and this split will happen
  5051. 3:31:08until we get finally a pure split
  5052. 3:31:11similarly with respect to rain we will
  5053. 3:31:13go ahead and take another feature and
  5054. 3:31:15we'll keep on splitting unless and until
  5055. 3:31:18we get a leaf node which is completely
  5056. 3:31:21pure I hope you understood how this
  5057. 3:31:23exactly work now two questions two
  5058. 3:31:27questions is that Kish the first thing
  5059. 3:31:29is that how do we calculate this
  5060. 3:31:32Purity and how do we come to know that
  5061. 3:31:35this is a pure split just by seeing
  5062. 3:31:38definitely I can say I can definitely
  5063. 3:31:41say by just seeing that how many number
  5064. 3:31:43of yes or NOS are there based on that I
  5065. 3:31:45can def itely say it is a pure split or
  5066. 3:31:47not so for this we use two different
  5067. 3:31:50things one is
  5068. 3:31:53entropy and the other one is something
  5069. 3:31:55called as guine coefficient so we will
  5070. 3:31:58try to understand how does entropy work
  5071. 3:32:01and how does Guinea coefficient work in
  5072. 3:32:04decision tree which will help us to
  5073. 3:32:06determine whether the split is pure
  5074. 3:32:09split or not or whether this node is
  5075. 3:32:11leaf node or not then coming to the
  5076. 3:32:13second thing okay coming to the second
  5077. 3:32:16thing one is with respect to Purity
  5078. 3:32:18second thing your first most important
  5079. 3:32:20question which you had asked why did I
  5080. 3:32:22probably select Outlook how the features
  5081. 3:32:24are selected and here you have a topic
  5082. 3:32:27which is called as Information Gain and
  5083. 3:32:29if you know this both your problem is
  5084. 3:32:32solved so now let's go ahead and let's
  5085. 3:32:35understand about entropy or guinea
  5086. 3:32:38coefficient or Information Gain entropy
  5087. 3:32:40or guine coefficient oh sorry Guinea
  5088. 3:32:42coefficient I'm saying guine impurity
  5089. 3:32:44also you can say over here
  5090. 3:32:46I'll write it as guine impurity not
  5091. 3:32:48coefficient also I'll just say it as
  5092. 3:32:50Guinea impurity but I hope everybody is
  5093. 3:32:53understood till here let's go ahead and
  5094. 3:32:55let's discuss about the first thing that
  5095. 3:32:57is
  5096. 3:32:58entropy how does entropy work and how we
  5097. 3:33:01are going to use the formula so entropy
  5098. 3:33:04here I will just write guine so we are
  5099. 3:33:07going to discuss about this both the
  5100. 3:33:09things let's say that the entropy
  5101. 3:33:12formula which is given by I will write h
  5102. 3:33:14of s is equal to so h of s is equal to
  5103. 3:33:17minus P plus I'll talk about what is
  5104. 3:33:20minus what is p plus log base 2 p
  5105. 3:33:26+- p
  5106. 3:33:28minus log base 2 p minus so this is the
  5107. 3:33:32formula and in guine impurity the
  5108. 3:33:34formula is 1 minus summation of I equal
  5109. 3:33:391 2 N p² I even talk about when you
  5110. 3:33:43should use guine impurity when you
  5111. 3:33:44should not use guine impurity
  5112. 3:33:46when you should use entropy you know by
  5113. 3:33:48default the decision tree regression or
  5114. 3:33:51classific sorry decision tree
  5115. 3:33:53classification uses Guinea impurity now
  5116. 3:33:56let's take one specific example so my
  5117. 3:33:58example is that I have a feature one my
  5118. 3:34:00root node I have a feature one which is
  5119. 3:34:03my root node and let's say that in this
  5120. 3:34:05root node I have six yes and three NOS
  5121. 3:34:08very simple let's say that this has two
  5122. 3:34:11categories and based on this two
  5123. 3:34:13categories of split has happened that is
  5124. 3:34:16a C1 let's say in this I have 3 S3 Nos
  5125. 3:34:20and here I have 3 s0 Nos and this is my
  5126. 3:34:24second category always understand if I
  5127. 3:34:26do the sumission 3s and 3s is 6s see
  5128. 3:34:30this this sumission if I do 3 + 3 is
  5129. 3:34:33obviously 6 3 + 0 is obviously so this
  5130. 3:34:36you need to understand based on the
  5131. 3:34:38number of root nodes only almost it'll
  5132. 3:34:40be same now let's go ahead and let's
  5133. 3:34:44understand how do we Cal calculate let's
  5134. 3:34:46take this example how do we calculate
  5135. 3:34:48the entropy of this so I have already
  5136. 3:34:50shown you the entropy formula over here
  5137. 3:34:52now let's understand the components I
  5138. 3:34:55will write h of s is equal to minus sign
  5139. 3:34:59is there what is p+ p+ basically means
  5140. 3:35:03that what is the probability of yes what
  5141. 3:35:07is the probability of yes this is a
  5142. 3:35:10simple thing for you all out of this
  5143. 3:35:13what is the probability of yes yes out
  5144. 3:35:16of this so obviously how you'll write if
  5145. 3:35:19you want to find out the probability of
  5146. 3:35:20yes out of this see when I say plus that
  5147. 3:35:24basically means yes when I say minus
  5148. 3:35:27that basically means no so what is the
  5149. 3:35:29probability of yes so it is be nothing
  5150. 3:35:32but yes plus and minus are specifically
  5151. 3:35:35for binary
  5152. 3:35:37class this can be positive negative so
  5153. 3:35:40the probability with respect to yes can
  5154. 3:35:42I write 3x 3 only for this what is the
  5155. 3:35:45probability out of this total number of
  5156. 3:35:48this is there 3x3 similarly if I go and
  5157. 3:35:51see the next term log to the base 2 p+
  5158. 3:35:54so again if I go ahead and write over
  5159. 3:35:56here log to the base 2 p+ p+ is again
  5160. 3:36:033x3 so then again we have minus and this
  5161. 3:36:07is now P minus what is p minus 0 by 3
  5162. 3:36:11log base 2 0 by 3 this obviously will
  5163. 3:36:15become zero this will obviously become 0
  5164. 3:36:18because 0 divid by anything is zero what
  5165. 3:36:21will this be 1 log to the base 1 what is
  5166. 3:36:25this this is nothing but zero log to the
  5167. 3:36:28base 1 is nothing but zero tell me
  5168. 3:36:31whether this is a pure split or impure
  5169. 3:36:35split so this is a pure split whenever
  5170. 3:36:38we have a pure split the answer of the
  5171. 3:36:41entropy is going to come to zero so here
  5172. 3:36:44I'm going to Define one graph
  5173. 3:36:46this is H of s and let's say this is p+
  5174. 3:36:49or P minus if my probability of plus see
  5175. 3:36:53when I say probability of plus is 0.5
  5176. 3:36:56what will be probability of minus it
  5177. 3:36:57will also be 0. five right because it's
  5178. 3:37:01just like P is equal to 1 - Q right if p
  5179. 3:37:04is .5 then Q will be 1 - P same thing
  5180. 3:37:07right so when it
  5181. 3:37:09is5 obviously my h of s will be 1 let's
  5182. 3:37:14say so this is this is the graph that
  5183. 3:37:16will basically get formed let's go ahead
  5184. 3:37:19and try to calculate the entropy of this
  5185. 3:37:21guys what will be the entropy of this
  5186. 3:37:24node so here I'm going to just make a
  5187. 3:37:26graph h of s minus what is p+ p+ is
  5188. 3:37:31nothing but 3x 6 log base 2 3x 6
  5189. 3:37:37minus three no are there 3x 6 log base 2
  5190. 3:37:433x 6 so if you compute this
  5191. 3:37:46log base 2 to the^ of 1 if you do the
  5192. 3:37:50calculation here I'm actually going to
  5193. 3:37:52get one so when I'm getting one when I'm
  5194. 3:37:55actually getting one when you have three
  5195. 3:37:57yes and three NOS what is the
  5196. 3:37:59probability it is 50/50% right so when
  5197. 3:38:02your p+ is5 that basically means your h
  5198. 3:38:06of s is coming as one so from this graph
  5199. 3:38:09you can see that I'm getting one if this
  5200. 3:38:11is zero this is one this is zero and
  5201. 3:38:13this is one I hope everybody is able to
  5202. 3:38:15to understand guys 0o and one if your p+
  5203. 3:38:20is
  5204. 3:38:21zero or if your p+ is one that basically
  5205. 3:38:24means it becomes a pure split so in h of
  5206. 3:38:26s you are going to get
  5207. 3:38:29zero so always understand your entropy
  5208. 3:38:33will be between 0 to
  5209. 3:38:361 if I have a impure this is a
  5210. 3:38:39completely impure split because here you
  5211. 3:38:42have 50% probability of getting yes 50%
  5212. 3:38:45probability of getting no h ofs is
  5213. 3:38:48entropy this is entropy for the sample H
  5214. 3:38:52ofs notation that I'm using is H ofs so
  5215. 3:38:56if whenever the split is happening the
  5216. 3:38:59first thing is done the purity test the
  5217. 3:39:02purity test is done with the help of
  5218. 3:39:04entropy right now I'll also show guinea
  5219. 3:39:07guinea impurity don't worry so with the
  5220. 3:39:09entropy you'll be able to find if I am
  5221. 3:39:11getting one that basically means it is a
  5222. 3:39:14impure split and if I'm getting zero it
  5223. 3:39:18is pure split so this is the graph okay
  5224. 3:39:22this is the graph and this graph is
  5225. 3:39:24basically the entropy graph again
  5226. 3:39:26understand if your probability of
  5227. 3:39:28getting yes or no is 0.5 that basically
  5228. 3:39:30means 50/50 is there 3s and three NOS
  5229. 3:39:34then your entropy is going to be 1 h of
  5230. 3:39:37s if your probability is completely one
  5231. 3:39:39that basically means either you're
  5232. 3:39:40getting completely yes or completely no
  5233. 3:39:43so your your entropy will be zero that
  5234. 3:39:46basically means it is pure split so in
  5235. 3:39:48the case of probability .5 you're
  5236. 3:39:50getting plus one then it'll keep on
  5237. 3:39:52reducing now let's go ahead and let's
  5238. 3:39:54try to understand so here you have
  5239. 3:39:56understood about purity test definitely
  5240. 3:39:58you'll use entropy try to find out
  5241. 3:40:00whether it is pure or impure if it is
  5242. 3:40:02impure you go ahead with the further
  5243. 3:40:04shift further division of the categories
  5244. 3:40:08again you take another feature divide it
  5245. 3:40:10because here from this two which split
  5246. 3:40:13you will do further you will do this
  5247. 3:40:14split as further if you are getting 6 6
  5248. 3:40:18is this specific value then you probably
  5249. 3:40:20go and draw over here this is your
  5250. 3:40:23entropy if your probability is here
  5251. 3:40:25which
  5252. 3:40:26is.3 then you will go here and create
  5253. 3:40:29this this may be0 4 or3 something like
  5254. 3:40:32this it will be between 0 to 1 let's go
  5255. 3:40:35ahead and discuss about the second issue
  5256. 3:40:37I hope everybody is discussed about we
  5257. 3:40:40have discussed about checking the pure
  5258. 3:40:42split or not and we have understood this
  5259. 3:40:45much but the next thing is that okay
  5260. 3:40:47fine chish this is very good we have
  5261. 3:40:49explained well I know many people will
  5262. 3:40:51say that but there are some people I
  5263. 3:40:53can't help let's say that I have some
  5264. 3:40:55features okay now coming to the second
  5265. 3:40:58problem how do we consider which node to
  5266. 3:41:02cap which which feature to take and
  5267. 3:41:05split because here I may have one one
  5268. 3:41:08split so again let's see that what is
  5269. 3:41:10the second problem which feature to take
  5270. 3:41:14to split right this is the second
  5271. 3:41:16problem that we are trying to solve
  5272. 3:41:18let's say that I have one feature one
  5273. 3:41:19over here and I have two categories
  5274. 3:41:22let's say this is there C1 and C2 here
  5275. 3:41:25let's say that I have 9 years 5 Nos and
  5276. 3:41:29then I have 6 years 2 NOS here I have
  5277. 3:41:32basically three yes and three NOS let's
  5278. 3:41:34say and in my data set I have features
  5279. 3:41:36like F1 FS2 F3 now let's say that
  5280. 3:41:40another split I can actually start with
  5281. 3:41:42feature two also and in feature two I
  5282. 3:41:45may have probably three categories like
  5283. 3:41:47C1 C2 C3 so with respect to the root
  5284. 3:41:52node and all the other features because
  5285. 3:41:54after this also I may have to split
  5286. 3:41:56right I may have to take another feature
  5287. 3:41:58and keep on splitting right based on the
  5288. 3:42:01Pure or impure split how do I decide
  5289. 3:42:03should I take fub1 first or F2 first or
  5290. 3:42:07F3 first or any other feature first how
  5291. 3:42:10should I decide that which feature
  5292. 3:42:12should I take and probably do the split
  5293. 3:42:15that is the major question so for this
  5294. 3:42:18we specifically use something called as
  5295. 3:42:20Information Gain so here I'm just going
  5296. 3:42:22to say here we basically use Information
  5297. 3:42:26Gain now what is this Information Gain
  5298. 3:42:29I'll talk about it so Information Gain
  5299. 3:42:31first of all I will write the formula we
  5300. 3:42:33basically write gain with sample first
  5301. 3:42:37with feature one I will compute so first
  5302. 3:42:40with feature one I will compute suppose
  5303. 3:42:42this is my first split of my data and
  5304. 3:42:44probably I'm Computing over here this
  5305. 3:42:46can be written as h of s I'll discuss
  5306. 3:42:50about each and every parameter don't
  5307. 3:42:51worry summation of V belong to values s
  5308. 3:42:56of V don't worry guys if you have not
  5309. 3:42:58understood the formula I will explain it
  5310. 3:43:01then the sample size H of SV I'll
  5311. 3:43:04discuss about each and every parameter
  5312. 3:43:06let's say that I'm taking this feature
  5313. 3:43:09one split I have you have already seen
  5314. 3:43:11what is feature one so this is my
  5315. 3:43:13feature one I have two categories C1 C2
  5316. 3:43:18this has 9 yes 5 NOS this has 6s and two
  5317. 3:43:24Nos and this has 3 yes and three NOS now
  5318. 3:43:27I will try to calculate the information
  5319. 3:43:29gain of this specific split now I will
  5320. 3:43:32go ahead and probably take this up now
  5321. 3:43:35see over here we'll try to understand
  5322. 3:43:37what is this now if I want to compute
  5323. 3:43:40the gain of s of F1 first is first first
  5324. 3:43:43thing that I need to find out is H of s
  5325. 3:43:45now this h of s is specifically of the
  5326. 3:43:48root node so I need to first of all
  5327. 3:43:50calculate what is h of s h ofs is
  5328. 3:43:52nothing but entropy entropy of the root
  5329. 3:43:56node so if I want to compute the entropy
  5330. 3:43:58of the node node tell me how should I
  5331. 3:44:00compute h of s is equal to minus p + log
  5332. 3:44:04base 2 p+ calculate guys along with me -
  5333. 3:44:07P minus log base to P minus so I hope
  5334. 3:44:11everybody knows this so here I'm going
  5335. 3:44:13to compute by what is ability of plus
  5336. 3:44:16over here in this specific root node it
  5337. 3:44:18is nothing but 9 by4 then I have log
  5338. 3:44:22base 2 again 9
  5339. 3:44:24by4 then I have P minus what is p minus
  5340. 3:44:285x4 log base 2 5 by4 so this calculation
  5341. 3:44:34I will probably get it as
  5342. 3:44:3694 approximately equal to 94 just check
  5343. 3:44:40it whether you're getting this or not
  5344. 3:44:42again you can use calculator if you want
  5345. 3:44:44now now I have definitely found out this
  5346. 3:44:47this is specifically for the root node
  5347. 3:44:50now let's see the next thing the next
  5348. 3:44:51important thing which is this part what
  5349. 3:44:54is s of v and what is s and what is h of
  5350. 3:44:57SV now very important just have a look
  5351. 3:45:01everybody see this graph okay see this
  5352. 3:45:05graph I will talk about h of SV first of
  5353. 3:45:07all I'll talk about h of SV okay this
  5354. 3:45:10one this is the entropy of category one
  5355. 3:45:13you need to find and entropy of category
  5356. 3:45:152 you need to find so if I write h of SV
  5357. 3:45:19of category 1 so what is category 1 for
  5358. 3:45:22this I'll write SC1 let's say I'm going
  5359. 3:45:25to write like this quickly calculate the
  5360. 3:45:28H of SV of this and this separately you
  5361. 3:45:31need to calculate so h of SV of C1 okay
  5362. 3:45:35so here again you'll write - 6X 8 log
  5363. 3:45:38base 2 6X
  5364. 3:45:418us 2x 8 log base to 2x 8 I hope
  5365. 3:45:46everybody knows this how we got it so h
  5366. 3:45:50of SV basically means I'm going to
  5367. 3:45:51compute the entropy of this category and
  5368. 3:45:54this category so for that I will
  5369. 3:45:56basically write h of so here I will
  5370. 3:45:59write - 6 by8 log base 2 6X 8 - 2x 8 log
  5371. 3:46:08base 2 2x 8 so if I get it I'm actually
  5372. 3:46:12going to get 81 and similarly if I if I
  5373. 3:46:15calculate h of C2 quickly calculate how
  5374. 3:46:18much you are going to get guys 6X 8 6X 8
  5375. 3:46:21with respect to this we need to find out
  5376. 3:46:24so now we have all these values we'll
  5377. 3:46:25start equating them to this equation so
  5378. 3:46:29here we have finally gain of s comma
  5379. 3:46:33fub1 so let's say that here I'm going to
  5380. 3:46:36basically add
  5381. 3:46:3894 minus see minus summation of okay
  5382. 3:46:42summation of what is s s of V understand
  5383. 3:46:46s of V basically means that how many
  5384. 3:46:48samples I have over here let's say for
  5385. 3:46:51category one how many samples I have for
  5386. 3:46:54category one over here simple if you
  5387. 3:46:56really want to just calculate it is
  5388. 3:46:58nothing but eight and total number of
  5389. 3:47:01sample is how much if I go and see over
  5390. 3:47:03here there are 9 years five NOS okay 9
  5391. 3:47:07years and five NOS that basically means
  5392. 3:47:1014 total sample here you have eight
  5393. 3:47:13sample Okay so this will become
  5394. 3:47:178x4 then you multiply by what see see
  5395. 3:47:21from this equation you multiply by h of
  5396. 3:47:23SV so h of SV is nothing but the entropy
  5397. 3:47:26of category 1 so entropy of category 1
  5398. 3:47:29is nothing but 81 plus then you go again
  5399. 3:47:33back to the graph and try to see that
  5400. 3:47:36for C2 how much how many total number of
  5401. 3:47:39samples are there 3 + 3 is 6 so 6 by 14
  5402. 3:47:42it will
  5403. 3:47:43become multiplied by 1 right so this is
  5404. 3:47:49your entire thing so here after all the
  5405. 3:47:52calculation you are going to get
  5406. 3:47:540.041 so this is my gain with s comma F1
  5407. 3:47:59so here I have got this value amazing I
  5408. 3:48:02did this with feature one only what
  5409. 3:48:05about feature two let's say that this
  5410. 3:48:07was my split for feature two and suppose
  5411. 3:48:10I get the gain for S comma feature 2 as
  5412. 3:48:17.51 if I get this now tell
  5413. 3:48:21me in using which feature should I start
  5414. 3:48:25splitting first whether it should be
  5415. 3:48:28fub1 or whether it should be FS2 based
  5416. 3:48:31on this value you know that over here
  5417. 3:48:35the gain the information gain of s comma
  5418. 3:48:38F2 is greater than gain of s comma fub1
  5419. 3:48:43so your answer is very much simple we
  5420. 3:48:45will definitely use feature 2 to start
  5421. 3:48:48the split the thing over here you are
  5422. 3:48:51trying to understand that if I really
  5423. 3:48:52want to select which feature to select
  5424. 3:48:54to start my splitting then I have to
  5425. 3:48:58basically calculate the information gain
  5426. 3:49:00and go throughout the all the paths and
  5427. 3:49:03whichever path has the highest
  5428. 3:49:04Information
  5429. 3:49:05Gain then we will select that specific
  5430. 3:49:09thing now the question Rises Kish
  5431. 3:49:12obviously this is good but you had
  5432. 3:49:14written about guinea impurity what is
  5433. 3:49:16the purpose of that please explain us
  5434. 3:49:19and why Guinea impurity is basically
  5435. 3:49:20used so let me go ahead with guine
  5436. 3:49:22impurity I told that yes you can
  5437. 3:49:25obviously
  5438. 3:49:26use you can obviously use entropy but
  5439. 3:49:29why Guinea impurity so guine impurity
  5440. 3:49:32formula which I have specifically
  5441. 3:49:34written as 1 minus summation of IAL 1
  5442. 3:49:382 N
  5443. 3:49:41p² now what is this p² suppose let's say
  5444. 3:49:45that in my n n is the number of outputs
  5445. 3:49:47right now how many outputs I have I have
  5446. 3:49:49two outputs yes or no so I will expand
  5447. 3:49:52this 1 minus since this is summation I
  5448. 3:49:55equal to 1 to n I'm basically going to
  5449. 3:49:57basically say that okay fine I will
  5450. 3:50:00write probability of plus whole
  5451. 3:50:03Square uh plus probability of minus
  5452. 3:50:07whole Square so this is the formula for
  5453. 3:50:10guinea impurity now you may be thinking
  5454. 3:50:14okay fine the calculation will be
  5455. 3:50:16obviously very much equal easy right
  5456. 3:50:18suppose if I have a node sorry if I have
  5457. 3:50:21a node which which has 2 yes two NOS now
  5458. 3:50:25in this particular case how do I
  5459. 3:50:26calculate my this probability if I have
  5460. 3:50:29two yes or two NOS suppose let's say
  5461. 3:50:31that I have a node over here which is my
  5462. 3:50:33split and this is having two yes and two
  5463. 3:50:36no so how do I calculate I will write 1
  5464. 3:50:38minus what is probability of square 1X 2
  5465. 3:50:41square sorry not 1 by two
  5466. 3:50:45yeah 1X 2 squ + 1 by 2
  5467. 3:50:49squ right then I will say 1 by 1X 4 + 1X
  5468. 3:50:544 is nothing but 2x 4 which is nothing
  5469. 3:50:56but 1X 2 so I will be getting 0.5 now
  5470. 3:51:00here here you understand this is a
  5471. 3:51:02complete impure split right if you have
  5472. 3:51:06an impure split in entropy the output
  5473. 3:51:10you getting it as one whereas in the
  5474. 3:51:13case of Guinea impurity
  5475. 3:51:15as Z sorry
  5476. 3:51:170.5 so if I go ahead with the graph that
  5477. 3:51:21I probably had created here so my Guinea
  5478. 3:51:24impurity line will look something like
  5479. 3:51:27this so it will be looking something
  5480. 3:51:29like this for zero obviously I'll be
  5481. 3:51:31getting zero but whenever my probability
  5482. 3:51:34of plus is 0.5 I'm going to get 0.5 over
  5483. 3:51:38here and that is the difference between
  5484. 3:51:40Guinea
  5485. 3:51:42impurity and entropy but the re but you
  5486. 3:51:45may be seeing Kish when to use what now
  5487. 3:51:48let's understand that when to use Guinea
  5488. 3:51:51and when to use entropy tell me guys if
  5489. 3:51:55I consider this formula of guine
  5490. 3:51:58impurity and if I probably
  5491. 3:52:01consider if I consider entropy this
  5492. 3:52:05formula where do you think more time
  5493. 3:52:09will take for execution for this
  5494. 3:52:11particular formula whether for entropy
  5495. 3:52:14it will take or for guinea impurity it
  5496. 3:52:18will take more time where it will
  5497. 3:52:21probably take for the execution purpose
  5498. 3:52:24see understand decision tree is having a
  5499. 3:52:29worst time complexity because if you
  5500. 3:52:32have 100 features probably you'll keep
  5501. 3:52:34on comparing by dividing many many
  5502. 3:52:37features then probably compute a
  5503. 3:52:38Information Gain like this if you have
  5504. 3:52:40just 100 features so which is faster
  5505. 3:52:43entrop
  5506. 3:52:45or guine impurity understand in entropy
  5507. 3:52:48you have log function here you have log
  5508. 3:52:52function here you have simple maths the
  5509. 3:52:56more amount of time out of entropy and
  5510. 3:52:59guine impurity the more amount of time
  5511. 3:53:01basically is taken
  5512. 3:53:03by
  5513. 3:53:06entropy so if you have huge number of
  5514. 3:53:10features like 100 200 features and you
  5515. 3:53:12are planning to apply decision Tre I
  5516. 3:53:15would suggest try to use Guinea impurity
  5517. 3:53:18then entropy if you have small set of
  5518. 3:53:20features then you can go ahead with
  5519. 3:53:23entropy so over here definitely with
  5520. 3:53:25respect to fast Guinea is greater than
  5521. 3:53:31entropy now let's go ahead and
  5522. 3:53:33understand with respect to you may be
  5523. 3:53:36thinking Kish okay fine you have
  5524. 3:53:38basically explained us about categorical
  5525. 3:53:41variables over here see over here you
  5526. 3:53:44have you have explained about
  5527. 3:53:45categorical variables what if I have
  5528. 3:53:47numerical feature let's say I have F1
  5529. 3:53:51over here which is a numerical
  5530. 3:53:53feature I have an F1 feature which is
  5531. 3:53:56numerical feature and I may have values
  5532. 3:53:58let's say that I have sorted all the
  5533. 3:54:00values over here okay let's say that I
  5534. 3:54:02have F1 and output okay so this F1 let's
  5535. 3:54:06say that I have values
  5536. 3:54:07like ass sorted order values I'm sorting
  5537. 3:54:10this features I'm basically doing this
  5538. 3:54:12let's say that initially I have this
  5539. 3:54:15features like this and let's say I have
  5540. 3:54:17values like 2.3 1.3 4 5 7 3 let's say I
  5541. 3:54:23have this features now this is a
  5542. 3:54:26continuous
  5543. 3:54:27feature this is a continuous feature so
  5544. 3:54:29for a continuous feature how probably
  5545. 3:54:32the decision tree entropy will be
  5546. 3:54:34calculated and the Information Gain will
  5547. 3:54:37get calculated so here you'll be able to
  5548. 3:54:39see that I will first of all sort these
  5549. 3:54:41values so in F1 the decision tree will B
  5550. 3:54:44basically first of all sort this values
  5551. 3:54:45so I have 1.3 then you have 2.3 then you
  5552. 3:54:49have four then you have three three then
  5553. 3:54:53you have four then you have five and
  5554. 3:54:55then you have six now whenever you have
  5555. 3:54:57a continuous feature so how the
  5556. 3:54:59continuous feature will basically work
  5557. 3:55:01in this case first of all your decision
  5558. 3:55:04tree node will say
  5559. 3:55:06that we'll take this one only one first
  5560. 3:55:10record and say that if it is less than
  5561. 3:55:12or equal to 1.3
  5562. 3:55:14okay if it is less than or equal to 1.3
  5563. 3:55:16so you here you'll be getting two
  5564. 3:55:18branches yes or no so yes and no
  5565. 3:55:22definitely your output over here will be
  5566. 3:55:25put over here right and then for the no
  5567. 3:55:28here you'll be having another node over
  5568. 3:55:30here how many number of Records you'll
  5569. 3:55:31be having in this particular case you'll
  5570. 3:55:33be having one record in this particular
  5571. 3:55:35case you will be having around five to
  5572. 3:55:36six records and here also you'll be able
  5573. 3:55:38to see right how many yes and NOS are
  5574. 3:55:40there definitely this will be a leaf
  5575. 3:55:42node so in the first instance they will
  5576. 3:55:45go ahead and calculate the information
  5577. 3:55:47gain of this then probably once the
  5578. 3:55:50Information Gain Is got then what
  5579. 3:55:51they'll do they will take the first two
  5580. 3:55:54records and again create a new decision
  5581. 3:55:57tree let's say that this will be my
  5582. 3:56:00suggestion where they'll say it is less
  5583. 3:56:02than or equal to 2.3 so I will get one
  5584. 3:56:05and one over here so in this now you'll
  5585. 3:56:07be having two records which will
  5586. 3:56:09basically say how many yes and no are
  5587. 3:56:10there and remaining all records will
  5588. 3:56:12come over here then again Information
  5589. 3:56:16Gain will be computed here then again
  5590. 3:56:17what will happen they'll go to the next
  5591. 3:56:19record then then again they'll create
  5592. 3:56:21another feature where they'll say less
  5593. 3:56:22than or equal to three and they will
  5594. 3:56:24create this many nodes again they'll try
  5595. 3:56:28to understand that how many yes or no
  5596. 3:56:29are there and then they'll again compute
  5597. 3:56:31The Information Gain like this they'll
  5598. 3:56:34do it for each and every record and
  5599. 3:56:36finally whichever Information Gain is
  5600. 3:56:38higher they will select that specific
  5601. 3:56:40value in that feature and they'll split
  5602. 3:56:42the node so in a continuous feature
  5603. 3:56:45whenever you have a continuous feature
  5604. 3:56:47this is how it will basically have and
  5605. 3:56:50then it will try to compute who is
  5606. 3:56:51having the highest Information Gain the
  5607. 3:56:54best Information Gain will get selected
  5608. 3:56:57and from there the splitting will
  5609. 3:56:59happen now let's go ahead and understand
  5610. 3:57:01about the next topic is that how this
  5611. 3:57:04entirely things work in decision tree
  5612. 3:57:07regressor because in decision tree
  5613. 3:57:09regressor my output is an continuous
  5614. 3:57:13variable so suppose if I have one
  5615. 3:57:15feature one feature two and this output
  5616. 3:57:17is a continuous feature it will be
  5617. 3:57:20continuous any value can be there so in
  5618. 3:57:23this particular case how do I split it
  5619. 3:57:27so let's say that f1c feature is getting
  5620. 3:57:30selected now in this f1c feature what
  5621. 3:57:32value will come when it is getting
  5622. 3:57:34selected first of all the entire mean
  5623. 3:57:38will get calculated of the output mean
  5624. 3:57:40will get calculated so here I will have
  5625. 3:57:42the mean and here here the cost function
  5626. 3:57:45that is used is not Guinea coefficient
  5627. 3:57:48or guinea impurity or entropy here we
  5628. 3:57:51use mean squared
  5629. 3:57:53error or you can also use mean absolute
  5630. 3:57:56error now what is mean squared error if
  5631. 3:57:58you remember from our logistic linear
  5632. 3:58:00regression how do we calculate 1 by 2 m
  5633. 3:58:03summation of I = 1 to n y hat minus y
  5634. 3:58:08whole Square y hat of i y - y whole
  5635. 3:58:12Square this is what is mean square error
  5636. 3:58:14so what it will do first based on F1
  5637. 3:58:17feature it will try to assign a mean
  5638. 3:58:20value and then it will compute the MSE
  5639. 3:58:23value and then it'll go ahead and do the
  5640. 3:58:26splitting now when it is doing splitting
  5641. 3:58:29based on categories of continuous
  5642. 3:58:31variable I will be having different
  5643. 3:58:33different categories now in this
  5644. 3:58:35categories what will happen after split
  5645. 3:58:37some records will go over
  5646. 3:58:40here then I will be having a mean value
  5647. 3:58:42of this over here
  5648. 3:58:45that will be my output and then again
  5649. 3:58:47the MSC will get calculated over here as
  5650. 3:58:50the msse gets reduced that basically
  5651. 3:58:53means we are reaching near the leaf
  5652. 3:58:55note and the same thing will happen over
  5653. 3:58:57here so finally when you follow this
  5654. 3:59:00path whatever mean value is present over
  5655. 3:59:02here that will be your output this is
  5656. 3:59:05the difference between the decision tree
  5657. 3:59:06regressor and the classifier here
  5658. 3:59:09instead of using entropy and all you use
  5659. 3:59:12mean squar error or mean absolute error
  5660. 3:59:14and this is the formula of mean square
  5661. 3:59:16error now let's go to the one more topic
  5662. 3:59:19which is called as the hyperparameters
  5663. 3:59:22tell me decision tree if I keep on
  5664. 3:59:25growing this to any depth what kind of
  5665. 3:59:28problem it will face regressor part you
  5666. 3:59:31want me to explain okay let's
  5667. 3:59:33see okay let's let's do the
  5668. 3:59:36regression decision
  5669. 3:59:39tree
  5670. 3:59:41regressor let's say I have feature F1
  5671. 3:59:44and this is my output let's say I have
  5672. 3:59:46values like 20 24 26 28 30 and this is
  5673. 3:59:53my feature one with category one
  5674. 3:59:56category one let's
  5675. 3:59:58say some categories are there let's say
  5676. 4:00:01I have done
  5677. 4:00:03the division by
  5678. 4:00:06F1 that is this feature initially tell
  5679. 4:00:09me what is the mean of this that mean
  5680. 4:00:12value will get assigned over here then
  5681. 4:00:14using msse that is mean squar error here
  5682. 4:00:18you will try to calculate suppose I get
  5683. 4:00:20an msse of some 37 47 something like
  5684. 4:00:23this and then I will try to split this
  5685. 4:00:27then I will be getting two more nodes or
  5686. 4:00:29three more nodes it depends then that
  5687. 4:00:31specific nodes will be the part of this
  5688. 4:00:33again the mean will change again the
  5689. 4:00:36mean will change over here suppose this
  5690. 4:00:38two is there this two records goes here
  5691. 4:00:41right then again MC will get calculated
  5692. 4:00:44I'm just taking as an example over here
  5693. 4:00:46just try to assume this thing now if I
  5694. 4:00:48talk about hyper parameters see this is
  5695. 4:00:51what is the formula that gets applied
  5696. 4:00:52over MSC now let's see in this hyper
  5697. 4:00:56parameter always understand decision
  5698. 4:00:58tree leads to overfitting because we are
  5699. 4:01:00just going to divide the nodes to
  5700. 4:01:03whatever level we want so this obviously
  5701. 4:01:06will lead to
  5702. 4:01:07overfitting now in order to prevent
  5703. 4:01:10overfitting we perform two important
  5704. 4:01:12steps one is post pruning and one is
  5705. 4:01:16pre- pruning so this two post pruning
  5706. 4:01:18and pre pruning is a condition let's say
  5707. 4:01:21that I have done some
  5708. 4:01:23splits I have done some splits let's say
  5709. 4:01:26over here I have seven yes and two
  5710. 4:01:28no and again probably I do the further
  5711. 4:01:31split like this now in this particular
  5712. 4:01:33scenario you know that if 7 yes and two
  5713. 4:01:35NOS are there there is a maximum there
  5714. 4:01:37is more than 80% chances that this node
  5715. 4:01:40is saying that the output is yes so
  5716. 4:01:43should we further do more
  5717. 4:01:46pruning the answer is no we can close it
  5718. 4:01:49and we can cut the branch from here this
  5719. 4:01:52technique is basically called as post
  5720. 4:01:54pruning that basically means first of
  5721. 4:01:57all you create your decision tree then
  5722. 4:01:59probably see the decision tree and see
  5723. 4:02:01that whether there is an extra Branch or
  5724. 4:02:03not and just try to cut it there is one
  5725. 4:02:06more thing which is called as
  5726. 4:02:07pre-pruning now pre-pruning is decided
  5727. 4:02:10by hyperparameters what kind of hyper
  5728. 4:02:13parameters you can basically say that
  5729. 4:02:15how many number of decision tree needs
  5730. 4:02:17to be used not number of decision tree
  5731. 4:02:20sorry over here you may say that what is
  5732. 4:02:22the max
  5733. 4:02:24depth what is the max depth how many Max
  5734. 4:02:27Leaf you can
  5735. 4:02:28have so this all parameters you can set
  5736. 4:02:31it with grid SE
  5737. 4:02:33CV and you can try it and you can
  5738. 4:02:36basically come up with a pre- pruning
  5739. 4:02:38technique so this is the idea about
  5740. 4:02:41decision tree uh regressor yes yes it is
  5741. 4:02:44possible your guinea value will be one
  5742. 4:02:45no this graph is there
  5743. 4:02:47no Guinea value are you talking about
  5744. 4:02:50this Guinea entropy it will not be one
  5745. 4:02:51it will always be between 0
  5746. 4:02:53to.5 so the first thing first as usual
  5747. 4:02:57what we should do we should import the
  5748. 4:02:59libraries so here I will go ahead and
  5749. 4:03:02import the librar so I'll say
  5750. 4:03:04import pandas as NP PD import matplot
  5751. 4:03:10li. pyplot as PLT
  5752. 4:03:14uh
  5753. 4:03:16import so this basic things I have with
  5754. 4:03:19me so I will go and take any data set
  5755. 4:03:22that I want from SK
  5756. 4:03:24learn. data sets import let's say that
  5757. 4:03:28I'm going to take load Iris data set and
  5758. 4:03:31then I'm going to upload the iris data
  5759. 4:03:33set so I'm going to write load Iris
  5760. 4:03:36there is my Iris data set then the next
  5761. 4:03:38step uh once you get your iris data set
  5762. 4:03:41so this is my iris. dat
  5763. 4:03:45okay these are all my features the four
  5764. 4:03:47features will be there these four
  5765. 4:03:49features are petal length petal width
  5766. 4:03:51SLE length and SLE width this is my
  5767. 4:03:54independent features then if I really
  5768. 4:03:56want to apply
  5769. 4:03:58for classifier so decision tree
  5770. 4:04:03classifier so I can first of all import
  5771. 4:04:06from
  5772. 4:04:08skarn do tree import decision let's see
  5773. 4:04:13where decision tree present in a scalon
  5774. 4:04:16decision tree
  5775. 4:04:17classifier the name is absolutely fine
  5776. 4:04:20but I was not getting over here
  5777. 4:04:23so so this is got no module SK okay SK
  5778. 4:04:29skar
  5779. 4:04:31skn learn so here you have
  5780. 4:04:35classifier right now I'm just going to
  5781. 4:04:37overfit the data then I'll probably show
  5782. 4:04:38you how you can go ahead with uh
  5783. 4:04:42pruning so by default what are the
  5784. 4:04:44parameters over here if you probably go
  5785. 4:04:46and see in in the classifier over here
  5786. 4:04:49you have Criterion see this the first P
  5787. 4:04:52parameter is Criterion by default it is
  5788. 4:04:54Guinea then you have Splitter Splitter
  5789. 4:04:57basically means how you're going to
  5790. 4:04:58split and there also you have two types
  5791. 4:05:01best and random you can randomly select
  5792. 4:05:04the features and do it okay you should
  5793. 4:05:06always go with
  5794. 4:05:07best max depth is a hyper parameter
  5795. 4:05:11minimum sample lift is a hyper parameter
  5796. 4:05:13Max Fe features how many number of
  5797. 4:05:14features we are going to take in order
  5798. 4:05:16to fix that that is also an hyper
  5799. 4:05:17parameter so all these things are hyper
  5800. 4:05:19parameter okay so I will just by default
  5801. 4:05:22executed whatever is giving me in
  5802. 4:05:24decision tree and the next thing that
  5803. 4:05:26I'm actually going to do is create a
  5804. 4:05:28decision tree so for this I will be
  5805. 4:05:31using plot. fig size plot. figure inside
  5806. 4:05:35figure I have this fix
  5807. 4:05:38size okay and I will probably show in
  5808. 4:05:41some better figure size so that
  5809. 4:05:43everybody body will be able to see it so
  5810. 4:05:45here let me say that I'm going to take
  5811. 4:05:47an area of
  5812. 4:05:491510 and then probably I'm going to say
  5813. 4:05:51tree Dot
  5814. 4:05:54Plot and here I'm going to say a
  5815. 4:05:57classifier and it should be filled the
  5816. 4:06:00coloring should be filled with this so
  5817. 4:06:04tree sorry Tre Tre Tre Tre
  5818. 4:06:09Tre it should be classifi tree. plot
  5819. 4:06:12okay I have to also import uh tree so I
  5820. 4:06:16have to basically import tree so from SK
  5821. 4:06:20learn
  5822. 4:06:22import three again I'm getting
  5823. 4:06:26error has no attribute plot
  5824. 4:06:29why let me just see the documentation
  5825. 4:06:32guys so this plot function is like plot
  5826. 4:06:34uncore tree dot tab plot _ tree now what
  5827. 4:06:40is the error we are getting okay not
  5828. 4:06:42fitted yet
  5829. 4:06:44sorry so I'm going to say
  5830. 4:06:47classifier do fit on data what data
  5831. 4:06:53iris.
  5832. 4:06:55data and then I'm going to fit with Iris
  5833. 4:06:58dot
  5834. 4:07:00Target so once this is done I think now
  5835. 4:07:03it will get
  5836. 4:07:04executed so this is how your graph will
  5837. 4:07:07look like guys so here you can see this
  5838. 4:07:10is how your graph looks like now if I
  5839. 4:07:12show you the graph over here see you can
  5840. 4:07:14see some amazing things over here three
  5841. 4:07:18outputs are actually there in this when
  5842. 4:07:21you see in this left hand side this
  5843. 4:07:23become a leaf node so this first one is
  5844. 4:07:25probably vers color uh versol flower
  5845. 4:07:29okay if you go on the right hand side
  5846. 4:07:31here you can see 50/50 is there so based
  5847. 4:07:32on one feature based on one feature here
  5848. 4:07:35you'll be able to see that you are
  5849. 4:07:37getting a leaf node based on another
  5850. 4:07:39Branch here you are getting
  5851. 4:07:4105050 so again you have two more
  5852. 4:07:44features getting splitted over here so
  5853. 4:07:46here you have 495 here you have
  5854. 4:07:48471 do we require this split anybody
  5855. 4:07:51tell me from here do we require any any
  5856. 4:07:54more split just try to think this is
  5857. 4:07:56after post pruning I want to find out
  5858. 4:07:59whether more splits are required or not
  5859. 4:08:01now in this particular case you see this
  5860. 4:08:03after this do you require any
  5861. 4:08:05split you do not require right here you
  5862. 4:08:08are basically getting 47 and one I guess
  5863. 4:08:11after this also you require no split
  5864. 4:08:14understand this so this is basically
  5865. 4:08:15post pruning so you can then decide your
  5866. 4:08:19level and probably do it gu value is
  5867. 4:08:22more than
  5868. 4:08:240.5 okay this side H this is coming as
  5869. 4:08:290.5 greater than 0.5 it should not had
  5870. 4:08:33here it is
  5871. 4:08:340.5 no maximum .5 can come 0 to.5 only
  5872. 4:08:39should come I don't know why this is
  5873. 4:08:41coming as 667
  5874. 4:08:44I'll have a look onto this guys but
  5875. 4:08:47anywhere you see other than that you're
  5876. 4:08:50everywhere you're getting less
  5877. 4:08:51than5 the plotting graph is very much
  5878. 4:08:54easy you use SK learn import tree then
  5879. 4:08:57you basically do this get classify and
  5880. 4:08:59field is equal to true and you can just
  5881. 4:09:02do this so the agenda let me Define the
  5882. 4:09:05agenda what all things are there first
  5883. 4:09:08we'll understand about
  5884. 4:09:11emble techniques in this assemble
  5885. 4:09:13techniques we are basically going to
  5886. 4:09:15discuss about what is the difference
  5887. 4:09:17between
  5888. 4:09:19bagging and boosting
  5889. 4:09:22second what we are basically going to
  5890. 4:09:24discuss about is so uh the agenda of
  5891. 4:09:27this session is emble techniques bagging
  5892. 4:09:29and boosting then we are probably going
  5893. 4:09:31to cover random forest and then probably
  5894. 4:09:35we will try to cover adab boost and if I
  5895. 4:09:39have more energy I will also try to
  5896. 4:09:40cover XG boost so all this Al lthms
  5897. 4:09:43we'll discuss about it so let's go ahead
  5898. 4:09:46and let's start the
  5899. 4:09:48topics the first topic that we are going
  5900. 4:09:50to discuss is about emble
  5901. 4:09:52techniques now what exactly is emble
  5902. 4:09:55techniques and we are going to discuss
  5903. 4:09:58about it okay so emble techniques what
  5904. 4:10:01exactly is emble techniques till now we
  5905. 4:10:03have solved two different kind of
  5906. 4:10:04problem statement one is
  5907. 4:10:07classification and regression and you
  5908. 4:10:09have learned about different different
  5909. 4:10:11algorithms like uh linear regression
  5910. 4:10:13logistic regression we have discussed
  5911. 4:10:15about KNN we have discussed about
  5912. 4:10:17yesterday what disc what did we discuss
  5913. 4:10:19about n bias different different
  5914. 4:10:21algorithms we have already finished now
  5915. 4:10:24with respect to classification
  5916. 4:10:25regression Problem whatever algorithm we
  5917. 4:10:27are discussing there was only one
  5918. 4:10:28algorithm at a time we were discussing
  5919. 4:10:31one algorithm at a time we are
  5920. 4:10:32discussing and we are trying to either
  5921. 4:10:33solve a classification or a regression
  5922. 4:10:35problem now the next thing is over here
  5923. 4:10:38is that can we use multiple algorithms
  5924. 4:10:42mul multiple algorithm to solve a
  5925. 4:10:44problem multiple algorithms basically
  5926. 4:10:46means can we I'll just talk about it
  5927. 4:10:49okay now the if I ask this specific
  5928. 4:10:52question can we use multiple algorithms
  5929. 4:10:54to solve a problem at that point of time
  5930. 4:10:57I will definitely say yes we can because
  5931. 4:10:59we are going to use something called as
  5932. 4:11:00emble techniques there now what this
  5933. 4:11:03emble techniques is okay so emble
  5934. 4:11:06techniques in emble techniques we
  5935. 4:11:08specifically use two different ways one
  5936. 4:11:12is one one way is that we specifically
  5937. 4:11:15use and the other one I'll just go to
  5938. 4:11:16write it over here so one that we
  5939. 4:11:19basically use is something called as
  5940. 4:11:20bagging technique and the other one we
  5941. 4:11:23specifically use is something called as
  5942. 4:11:25boosting technique so in bagging
  5943. 4:11:27Technique we what exactly we can do and
  5944. 4:11:31in boosting technique what we can
  5945. 4:11:32actually do and how we are combining
  5946. 4:11:34multiple models to solve a problem so
  5947. 4:11:36let's first of all discuss about bagging
  5948. 4:11:39now how does bagging work let's say that
  5949. 4:11:42I have a specific data set so this is my
  5950. 4:11:44data set with uh with features rows
  5951. 4:11:48columns everything like this I have this
  5952. 4:11:50specific data set just imagine I have
  5953. 4:11:52many many features over here like this
  5954. 4:11:54fub1 F2 F3 and probably I have my output
  5955. 4:11:57so this is my data set D let's consider
  5956. 4:11:59it now what we do in bagging is that we
  5957. 4:12:04create models and this model can be
  5958. 4:12:06anything it can be logistic it can be
  5959. 4:12:08linear for a classification problem
  5960. 4:12:10let's say that this is logistic model so
  5961. 4:12:12this is my model M1 let's say I have
  5962. 4:12:14another model M2 then I may have another
  5963. 4:12:17model M3 let's say that this is
  5964. 4:12:20logistic and this is probably the other
  5965. 4:12:23model which is like decision tree and
  5966. 4:12:25then probably we use this model as KNN
  5967. 4:12:29classification and this model can again
  5968. 4:12:31be decision tree it's fine let's use
  5969. 4:12:34another decision tree so now here you
  5970. 4:12:36can see that we have used so many models
  5971. 4:12:39okay so many models are there now with
  5972. 4:12:41respect to this particular model what I
  5973. 4:12:42will do is that the first step that I
  5974. 4:12:44will do from this particular data set I
  5975. 4:12:46will just take up some rows so I'll
  5976. 4:12:48basically do row
  5977. 4:12:50sampling and I'll take a row sampling of
  5978. 4:12:53D Dash D Das basically means this D Das
  5979. 4:12:55is always less than D some of the rows
  5980. 4:12:58I'll push it to M1 okay I can also use n
  5981. 4:13:01fine so what I'll do is that some of the
  5982. 4:13:03rows I'll push it to model one this
  5983. 4:13:05model one will be training let's say
  5984. 4:13:07that for out of this 10,000 record th000
  5985. 4:13:09rows I'm actually doing a row sampling
  5986. 4:13:11of th rows and giving it to M1 to train
  5987. 4:13:14it then what I'm actually going to do
  5988. 4:13:16over here I'm basically going to give
  5989. 4:13:18this specific model M2 and again I'm
  5990. 4:13:21going to do row row sampling and I'm
  5991. 4:13:24again going to sample some of the rows
  5992. 4:13:25and give it to model two and again
  5993. 4:13:27remember some of the rows may get
  5994. 4:13:29repeated from this D Dash to next dble
  5995. 4:13:31Dash similarly I will do row sampling
  5996. 4:13:33and give it to this and again I may have
  5997. 4:13:35d triple Dash and D4 Dash so different
  5998. 4:13:38different different different rows data
  5999. 4:13:41points when I say row sampling basically
  6000. 4:13:42I'm talking about data points different
  6001. 4:13:45different data points I will give it to
  6002. 4:13:47separate separate model and this model
  6003. 4:13:49will specifically train when I say D
  6004. 4:13:52Dash that basically means uh suppose I
  6005. 4:13:54say th 10,000 are my total number of
  6006. 4:13:56data points when I say D Dash This D
  6007. 4:13:59Dash may be th000 points then D Double
  6008. 4:14:02Dash may be another th000 points and
  6009. 4:14:04some of the rows may get repeated over
  6010. 4:14:05here dle Dash here also I can basically
  6011. 4:14:08use so here specifically row sampling
  6012. 4:14:10will be used now when I have this many
  6013. 4:14:12specific each and every model will be
  6014. 4:14:14trained with different kind of data now
  6015. 4:14:17how the inferencing will happen for the
  6016. 4:14:18test data so first thing first let's say
  6017. 4:14:21that I'm going to get a new test data
  6018. 4:14:23over here now new test data will be
  6019. 4:14:25passed to M1 and this M1 suppose it
  6020. 4:14:28gives zero as my output suppose let's
  6021. 4:14:30say that I'm doing a binary
  6022. 4:14:31classification it gives a Zer as an
  6023. 4:14:33output so this is my output of zero next
  6024. 4:14:37M2 for the new test data gives one M3
  6025. 4:14:40gives one and M4 also gives one as the
  6026. 4:14:43the output now in this particular case
  6027. 4:14:46in this particular case what will happen
  6028. 4:14:49now you can see over here it's simple
  6029. 4:14:51what what do you think the output may be
  6030. 4:14:53in this particular case now M1 has
  6031. 4:14:55predicted for this particular test data
  6032. 4:14:56as zero the model M2 has predicted 1 M3
  6033. 4:15:00has predicted 1 and M4 has predicted one
  6034. 4:15:02so finally all these outputs are going
  6035. 4:15:04to get
  6036. 4:15:06aggregated are going to get aggregated
  6037. 4:15:08and a simple thing that gets applied is
  6038. 4:15:11majority voting majority voting so tell
  6039. 4:15:14me what will be the output for with
  6040. 4:15:16respect to this the output will
  6041. 4:15:18obviously be one because the majority
  6042. 4:15:19voting that you can see three people are
  6043. 4:15:21basically saying it as one so my output
  6044. 4:15:24over here will be one okay this is the
  6045. 4:15:26concept of bagging wherein you are
  6046. 4:15:29providing different different rows with
  6047. 4:15:31probably all the features in this case
  6048. 4:15:33and giving it to different different
  6049. 4:15:34model again which is a classification
  6050. 4:15:36model and then finally you are combining
  6051. 4:15:38them based on majority voting and you're
  6052. 4:15:40getting the answer as one so this step
  6053. 4:15:43is called as bootstrap aggregator that
  6054. 4:15:45basically means you're aggregating all
  6055. 4:15:48the output that is basically coming from
  6056. 4:15:50all the specific models all the specific
  6057. 4:15:52models now many people will say Krish
  6058. 4:15:54what about Tai guys like this kind of
  6059. 4:15:56situation you know we will be having
  6060. 4:15:58more than 100 to 200 models so it is
  6061. 4:16:01very very difficult that it will be a
  6062. 4:16:03tie who are repeating questions they
  6063. 4:16:05will be put up in time out so what if
  6064. 4:16:09you're saying that if the 50% of model
  6065. 4:16:12says yes 50% of our models says no
  6066. 4:16:14always understand guys we will be having
  6067. 4:16:17more than 100 to 200 plus models so in
  6068. 4:16:19this particular case there will be high
  6069. 4:16:21probability that always there will be a
  6070. 4:16:23majority voting available it will always
  6071. 4:16:25not be in that specific scenario so this
  6072. 4:16:28was the concept about bagging now some
  6073. 4:16:30people will be saying that Krish why are
  6074. 4:16:31you using different different models
  6075. 4:16:34guys I'm not discussing about random
  6076. 4:16:35Forest over here random Forest uses only
  6077. 4:16:37one type of model that is decision tree
  6078. 4:16:39but if we think as an concept of bagging
  6079. 4:16:43you can have different different models
  6080. 4:16:44over here and you can basically combine
  6081. 4:16:46them so this is a technique of emble
  6082. 4:16:49techniques and this is basically called
  6083. 4:16:51as bagging okay now tell me one point I
  6084. 4:16:54missed out fine this is with respect to
  6085. 4:16:56the classification problem with respect
  6086. 4:16:58to the regression problem what will
  6087. 4:17:00happen in case of a regression problem
  6088. 4:17:02let's say that I got here 120 here 140
  6089. 4:17:06here 122 here 148 as my output so in
  6090. 4:17:09regression what will happen is that the
  6091. 4:17:11entire mean will be taken mean will be
  6092. 4:17:15taken the output mean will be basically
  6093. 4:17:18taken and that will be your output of
  6094. 4:17:20the model average or mean very simple
  6095. 4:17:22right so average or mean will be
  6096. 4:17:25basically taken up and here based on the
  6097. 4:17:27average you'll be able to solve the
  6098. 4:17:29regression problem great now let's go
  6099. 4:17:31ahead and try to understand with respect
  6100. 4:17:34to bagging and boosting how many
  6101. 4:17:36different types of algorithm are but
  6102. 4:17:37before that I need to make you
  6103. 4:17:39understand what exactly is boosting now
  6104. 4:17:41here in bagging you have seen that you
  6105. 4:17:43have parallel models right one one one
  6106. 4:17:46independent you have parallel models
  6107. 4:17:48you're giving some row samples in
  6108. 4:17:49different different models and basically
  6109. 4:17:51are able to find out the output now in
  6110. 4:17:53case of boosting boosting is a
  6111. 4:17:56sequential combination of models like
  6112. 4:17:59this you have lot of sequential models
  6113. 4:18:03like this and one after the model like
  6114. 4:18:06first I'll give my training data to this
  6115. 4:18:07particular model then it will go to this
  6116. 4:18:09data then this model then this model so
  6117. 4:18:12this will be my M1 M2 M3 M4 and finally
  6118. 4:18:16I will be getting my output so here you
  6119. 4:18:18can basically say that boosting is all
  6120. 4:18:21about and this M1 M2 M3 we basically
  6121. 4:18:24mention it as weak Learners so this will
  6122. 4:18:26be weak learner weak learner weak
  6123. 4:18:29learner weak learner and finally when we
  6124. 4:18:32go till here it it'll if I combine all
  6125. 4:18:35these weak ners weak
  6126. 4:18:38learner weak learner okay once I combine
  6127. 4:18:41all this weak learner it becomes a it
  6128. 4:18:43becomes a strong learner finally if I
  6129. 4:18:46try to combine this this will basically
  6130. 4:18:47become a strong learner so here you have
  6131. 4:18:50all the models sequentially one after
  6132. 4:18:52the other and then you will probably try
  6133. 4:18:55to provide your uh input from one model
  6134. 4:18:58to the next model to the next model and
  6135. 4:19:00these all models will be a very simpler
  6136. 4:19:01weak learner model which will not be
  6137. 4:19:03able to predict properly but when you
  6138. 4:19:05combine all this particular models
  6139. 4:19:08together sequentially it becomes a
  6140. 4:19:09strong learner how this specifically
  6141. 4:19:11works I'll take an example example of AD
  6142. 4:19:13boost XG boost I will show you that okay
  6143. 4:19:16week learner basically means the
  6144. 4:19:17prediction is very bad but as you go
  6145. 4:19:19sequentially you combine them they
  6146. 4:19:21become a strong learner okay one example
  6147. 4:19:24I want to give you let's say that you
  6148. 4:19:26are a data scientist right let's say
  6149. 4:19:30that this model one may be a teacher
  6150. 4:19:33with respect to physics then this model
  6151. 4:19:35two may be a teacher with respect to
  6152. 4:19:37chemistry let's say model 3 is basically
  6153. 4:19:40a teacher of maths and model four is a
  6154. 4:19:43teacher of geography now suppose if you
  6155. 4:19:46are trying to solve one problem
  6156. 4:19:48obviously if the physics teacher is not
  6157. 4:19:50able to solve that particular problem
  6158. 4:19:51then probably chemistry can help or
  6159. 4:19:54maths can help or geography can help or
  6160. 4:19:56someone can help so when we combine this
  6161. 4:19:58many expertise together they will be
  6162. 4:20:01able to give you the output in an
  6163. 4:20:03efficient way Sumit I'll talk about it
  6164. 4:20:05where whether all the features are
  6165. 4:20:07basically passed to all the models or
  6166. 4:20:08not I'll just talk about it just give me
  6167. 4:20:10some time okay but I just want to give
  6168. 4:20:12you an idea about in short if someone
  6169. 4:20:14asks you in an interview what exactly is
  6170. 4:20:17boosting okay boosting is you can just
  6171. 4:20:21say that it is a sequential set of all
  6172. 4:20:23the models combined together and these
  6173. 4:20:25all models that I initialized are
  6174. 4:20:27usually weak Learners and when they are
  6175. 4:20:29combined together they become a strong
  6176. 4:20:30learner and based on the strong learner
  6177. 4:20:32they gives an amazing output and right
  6178. 4:20:35now if I say in most of the kaggle
  6179. 4:20:37competition they use different types of
  6180. 4:20:39boosting or bagging technique so we have
  6181. 4:20:42basically as I said
  6182. 4:20:44bagging and boosting in bagging what
  6183. 4:20:47kind of algorithm we specifically use we
  6184. 4:20:49use something called as random forest
  6185. 4:20:54classifier and the second model that we
  6186. 4:20:57specifically use is something called as
  6187. 4:20:59random
  6188. 4:21:00Forest regress so we specifically use
  6189. 4:21:04these two kind of models which I'm
  6190. 4:21:05actually going to discuss right now
  6191. 4:21:06after this and then in boosting we
  6192. 4:21:09basically use techniques like ad boost
  6193. 4:21:12gradi Boost number three is Extreme
  6194. 4:21:15gradient boost which we also say it as
  6195. 4:21:17XG boost extreme gradient boost so let's
  6196. 4:21:20go ahead and let's discuss about the
  6197. 4:21:22first algorithm which is called as
  6198. 4:21:24random forest classifier and regressor
  6199. 4:21:28now first thing first let's understand
  6200. 4:21:31some things from the yesterday's class I
  6201. 4:21:33hope uh what is the main problem with
  6202. 4:21:35respect to decision tree whenever we
  6203. 4:21:37create a decision tree without any
  6204. 4:21:39hyperparameter it does it not lead to
  6205. 4:21:42overit
  6206. 4:21:43does it not lead to overfitting uh
  6207. 4:21:45whenever you probably have a decision
  6208. 4:21:48tree right it leads to something like
  6209. 4:21:50overfitting why overfitting because it
  6210. 4:21:53completely splits all the feature till
  6211. 4:21:55it's complete depth overfitting
  6212. 4:21:57basically means for training data the
  6213. 4:21:58accuracy is high for test data the
  6214. 4:22:00accuracy is low so training data when
  6215. 4:22:02the accuracy is high I may basically say
  6216. 4:22:04it as high bias and then I may basically
  6217. 4:22:07say it as sorry not high bias low bias
  6218. 4:22:11and high V variance so low bias and high
  6219. 4:22:14variance yes obviously we can do pruning
  6220. 4:22:16and all guys but again understand
  6221. 4:22:18pruning is an extensive task probably if
  6222. 4:22:21your if you have 100 features if you
  6223. 4:22:23have data points which is like 1 million
  6224. 4:22:25to do pruning also it is very much
  6225. 4:22:27difficult yes pre pruning can be done
  6226. 4:22:29but again we cannot confirm that it may
  6227. 4:22:31work well or not so right now with
  6228. 4:22:33respect to decision tree you have this
  6229. 4:22:35specific problem that is low bias and
  6230. 4:22:37high variance now in low Biance and high
  6231. 4:22:39variance you know that my model is
  6232. 4:22:41basically the generalized model that I
  6233. 4:22:43should get it should have low bias and
  6234. 4:22:46low variance so if somebody asks you why
  6235. 4:22:49do you use random Forest you can
  6236. 4:22:51basically explain about decision trees
  6237. 4:22:52like this now my main aim is to convert
  6238. 4:22:54this High variance to low variance now I
  6239. 4:22:58will be able to convert this High
  6240. 4:22:59variance to low variance using random
  6241. 4:23:01forest classifier or random Forest
  6242. 4:23:03regressor now what does random Forest do
  6243. 4:23:06random Forest is a bagging technique
  6244. 4:23:08similarly I have a data set over here
  6245. 4:23:10let's say that I have this data set
  6246. 4:23:13and then here I will be having multiple
  6247. 4:23:15models like
  6248. 4:23:16M1
  6249. 4:23:19M2
  6250. 4:23:21M3 M4 let's say I have this four models
  6251. 4:23:24like this we have many many models now
  6252. 4:23:27with respect to this models this models
  6253. 4:23:29all the models are actually decision
  6254. 4:23:31Tree in random forest all are decision
  6255. 4:23:34trees you don't have a different model
  6256. 4:23:37over there so over here you can see that
  6257. 4:23:39all the models are decision trees that
  6258. 4:23:41is going to get used used in random
  6259. 4:23:43Forest so decision trees always gets
  6260. 4:23:45used in random Forest the first thing
  6261. 4:23:47that you should know now whenever we are
  6262. 4:23:49using decision trees you know that
  6263. 4:23:51decision tree if I by default if we try
  6264. 4:23:53to create it it may lead to overfitting
  6265. 4:23:56and because of that every decision tree
  6266. 4:23:58will basically create low V low bias and
  6267. 4:24:01high variance but if we combine in the
  6268. 4:24:04form of bootstrap aggregator this High
  6269. 4:24:07variance will be getting converted to
  6270. 4:24:08low variance because why because
  6271. 4:24:10majority of voting we will be taking
  6272. 4:24:12from this particular decision trees like
  6273. 4:24:14there will be many many decision tree so
  6274. 4:24:16they lot of outputs will be coming and
  6275. 4:24:19with the help of majority voting
  6276. 4:24:20classifier this High variance will get
  6277. 4:24:22converted to low variance now in random
  6278. 4:24:24Forest how it works in the first case if
  6279. 4:24:27I talk about random Forest over here two
  6280. 4:24:29things basically happen with respect to
  6281. 4:24:30the D- data set let's say in first model
  6282. 4:24:34we do some kind of row
  6283. 4:24:36sampling plus
  6284. 4:24:38Feature Feature
  6285. 4:24:40sampling that basically means we have to
  6286. 4:24:42select some set of rows and some set of
  6287. 4:24:45features and give it to M1 similarly you
  6288. 4:24:48do row sampling and feature sampling and
  6289. 4:24:50give it to M2 then you do row sampling
  6290. 4:24:52and feature sampling you give it to M3
  6291. 4:24:54and then you do row sampling and feature
  6292. 4:24:56sampling you give it to M4 now when you
  6293. 4:25:00do this so what will happen
  6294. 4:25:01independently you're giving some
  6295. 4:25:03features along with some rows now there
  6296. 4:25:05may be a situation that your features
  6297. 4:25:07may also get repeated it may also get
  6298. 4:25:09repeated your records or data points may
  6299. 4:25:11also get repeated so when you are
  6300. 4:25:13probably training your model with this
  6301. 4:25:15specific data sets and specific features
  6302. 4:25:18this model become expert in predicting
  6303. 4:25:20something right as I said one example
  6304. 4:25:23over here I'm giving a physics model
  6305. 4:25:25some data I'm giving chemistry data
  6306. 4:25:27chemistry model with some data similarly
  6307. 4:25:29here I'm giving some information to some
  6308. 4:25:31model so the model will be an expert
  6309. 4:25:33with respect to that specific data So
  6310. 4:25:36based on all this particular data
  6311. 4:25:38whenever I get a new test data so what
  6312. 4:25:40will happen suppose let's say that this
  6313. 4:25:42this is a classification problem the M1
  6314. 4:25:44model will be predicting zero this will
  6315. 4:25:46be predicting one this will be
  6316. 4:25:47predicting zero and this will be
  6317. 4:25:49predicting zero now in this particular
  6318. 4:25:51case again the majority voting
  6319. 4:25:53classifier or majority voting will
  6320. 4:25:55happen in the case of classification
  6321. 4:25:57problem and then here you will be
  6322. 4:26:01specifically able to get the output as
  6323. 4:26:03zero so I hope everybody is able to
  6324. 4:26:06understand all the models over here are
  6325. 4:26:07decision trees and based on that you
  6326. 4:26:10will be doing see when in I interview
  6327. 4:26:12should be very very uh things the things
  6328. 4:26:15that I'm telling you over here is all
  6329. 4:26:17all the points are very much important
  6330. 4:26:19and similarly if you tell the
  6331. 4:26:21interviewer definitely your interview is
  6332. 4:26:22cracked in this kind of algorithm I've
  6333. 4:26:25seen some of my students saying that
  6334. 4:26:26okay uh Kish um when the interviewer
  6335. 4:26:29asked me that which is my favorite
  6336. 4:26:30algorithm I said random Forest I told
  6337. 4:26:32why did you say like that because he
  6338. 4:26:34said that because that person let me let
  6339. 4:26:36him ask any questions in random Forest
  6340. 4:26:38I'm very much confident about it and I'm
  6341. 4:26:40also going to prove him you know
  6342. 4:26:42why they are very very good so with this
  6343. 4:26:45specific case here you can basically see
  6344. 4:26:47that because of the overfitting
  6345. 4:26:49condition of the decision tree you're
  6346. 4:26:50combining multiple decision tree so that
  6347. 4:26:52you get a generalized model which has
  6348. 4:26:54low bias and low variance so I hope
  6349. 4:26:57everybody is able to understand boost
  6350. 4:26:59feature sampling basically means suppose
  6351. 4:27:00if I have 1 2 3 four feature for the
  6352. 4:27:04first model I may give two features for
  6353. 4:27:06the second model I may get three
  6354. 4:27:07features for the fourth model I may give
  6355. 4:27:09four features or uh any one feature ALS
  6356. 4:27:12I can give to a specific model so
  6357. 4:27:13internally that random Forest it take
  6358. 4:27:15carees of over here these things are
  6359. 4:27:18there and this is how random Forest
  6360. 4:27:19Works only the difference between random
  6361. 4:27:21Forest classify and regression is that
  6362. 4:27:23in regression again whatever output you
  6363. 4:27:25are basically getting you basically do
  6364. 4:27:26the mean that's it average you just do
  6365. 4:27:29the average you'll be able to get the
  6366. 4:27:31output based on all the models output
  6367. 4:27:33that you are actually getting now let's
  6368. 4:27:34talk about some of the important points
  6369. 4:27:36in random Forest the first thing first
  6370. 4:27:38question is that is normalization
  6371. 4:27:41required in random Forest then the next
  6372. 4:27:43question is that in KNN is normalization
  6373. 4:27:47when I say normalization or
  6374. 4:27:49standardization I I'll just talk about
  6375. 4:27:51standardization is standardization is
  6376. 4:27:54required so this will be my another
  6377. 4:27:56question so is normalization required in
  6378. 4:27:59random forest or decision tree you here
  6379. 4:28:01you can also say it as decision tree is
  6380. 4:28:03it required so for this the answer will
  6381. 4:28:06be no because understand decision tree
  6382. 4:28:09will basically do the splits if you Mini
  6383. 4:28:12minimize the data also that split won't
  6384. 4:28:14be that much important but if I talk
  6385. 4:28:17about KNN whether standardization
  6386. 4:28:19normalization required over here the
  6387. 4:28:21answer is yes because here we use two
  6388. 4:28:23things one is ukan distance and
  6389. 4:28:26Manhattan distance because of this you
  6390. 4:28:28definitely have to apply standardization
  6391. 4:28:30so that the computation or distance
  6392. 4:28:32becomes easy so this is one of the most
  6393. 4:28:34common interview questions that is
  6394. 4:28:36basically asked in random Forest coming
  6395. 4:28:38to the third question is random Forest
  6396. 4:28:40impacted by outlier
  6397. 4:28:43over here the answer will be no just
  6398. 4:28:46check it out outside basically means
  6399. 4:28:48Google and check it out check it out in
  6400. 4:28:50Google okay perfect so I hope I've
  6401. 4:28:53covered most of the things in random
  6402. 4:28:54Forest is random Forest impacted by
  6403. 4:28:57outliers this is the third question is
  6404. 4:28:59KNN impacted by
  6405. 4:29:00outliers is this KNN algorithm impacted
  6406. 4:29:04by outliers is KNN impacted Byers the
  6407. 4:29:07answer is yes big yes perfect so so
  6408. 4:29:12these all are the interview questions
  6409. 4:29:13that needs to be covered now let's go
  6410. 4:29:15ahead and discuss about adab boost now
  6411. 4:29:18in bagging most of the time we
  6412. 4:29:20specifically use random forest or you
  6413. 4:29:23can also create custom bagging
  6414. 4:29:25techniques custom bagging techniques
  6415. 4:29:27means whatever algorithm you want use
  6416. 4:29:29the combination of them and try to give
  6417. 4:29:32the output this also you can do it
  6418. 4:29:33manually with the help of hands okay
  6419. 4:29:36guys so second thing uh we are going to
  6420. 4:29:38discuss about is boosting technique in
  6421. 4:29:40this
  6422. 4:29:42the first thing that uh first algorithm
  6423. 4:29:44that we are going to discuss about is
  6424. 4:29:45adab Boost so adab boost we going to
  6425. 4:29:48discuss about how does adab Boost uh
  6426. 4:29:50work now let's solve uh the first
  6427. 4:29:53boosting technique which is called as
  6428. 4:29:54adab boost okay and uh this is a
  6429. 4:29:57boosting technique um in the boosting
  6430. 4:30:00technique you have heard that we have to
  6431. 4:30:02basically solve in a sequential way this
  6432. 4:30:05at least you know I know there is a lot
  6433. 4:30:07of confusion within you all but we'll
  6434. 4:30:09try to solve a problem let's say so
  6435. 4:30:11suppose I have a data set which looks
  6436. 4:30:12like this fub1 F2 F3 F4 so these are my
  6437. 4:30:16features and probably these are my
  6438. 4:30:18output okay so let's say that I'm having
  6439. 4:30:20this features like this and this is my
  6440. 4:30:22output like yes or no like this so let's
  6441. 4:30:25say that how many records I have over
  6442. 4:30:27here three
  6443. 4:30:304 5 6 and one more is there 7 so this
  6444. 4:30:36seven records are there now in adab
  6445. 4:30:38boost the first thing is that
  6446. 4:30:40specifically with adab Boost uh you
  6447. 4:30:42really need to understand that what all
  6448. 4:30:43things we can basically do how do we
  6449. 4:30:45solve this classification problem that
  6450. 4:30:47we are going to understand the first
  6451. 4:30:49thing first is that we Define a weight
  6452. 4:30:51and the weight is very much simple
  6453. 4:30:53initially to all the records to all this
  6454. 4:30:55input records we provide an equal weight
  6455. 4:30:58now how do we provide an equal weight we
  6456. 4:30:59just go and count how many number of
  6457. 4:31:01records are there now in this particular
  6458. 4:31:03case the total number of records are one
  6459. 4:31:062 3 4 5 6 7 now every record I have to
  6460. 4:31:12provide an equal weight that is between
  6461. 4:31:150 to 1 so the overall sum should be one
  6462. 4:31:19so in this particular case what I can do
  6463. 4:31:20if I make 1X 7 1X 7 1X 7 to everyone
  6464. 4:31:24this will definitely become
  6465. 4:31:26a equal weights to all right and if I do
  6466. 4:31:30the total sum it will obviously be one
  6467. 4:31:32let's go to the next one now after this
  6468. 4:31:34what do we do okay after this in adab
  6469. 4:31:37the first thing that we do is that we
  6470. 4:31:39take any of this feature how do you
  6471. 4:31:41decide which feature to take whether we
  6472. 4:31:42should go with F1 or whether we should
  6473. 4:31:44go with FS2 or whether we should go with
  6474. 4:31:46F3 this we can do it with the help of
  6475. 4:31:49Information Gain and Information Gain
  6476. 4:31:53and entropy or guinea right based on
  6477. 4:31:56this we can definitely understand
  6478. 4:31:57whether we should start making decision
  6479. 4:31:59here also you specifically make decision
  6480. 4:32:01trees so here what you do is that you
  6481. 4:32:04probably have to determine by using
  6482. 4:32:06which feature I have to start my
  6483. 4:32:07decision tree so suppose out of all this
  6484. 4:32:09feature one feature two feature three
  6485. 4:32:11you have selected that okay the
  6486. 4:32:12information gain and entropy of feature
  6487. 4:32:13one is higher so I'm going to use
  6488. 4:32:15feature one and probably divide this
  6489. 4:32:17into decision trees now when I divide
  6490. 4:32:21this into decision tree let's say that
  6491. 4:32:22I'm dividing like this into decision
  6492. 4:32:23tree this decision tree depth will be
  6493. 4:32:26only one one depth and this depth since
  6494. 4:32:29it has only one depth we basically call
  6495. 4:32:31it as stumps so what we do over here
  6496. 4:32:34specifically we will create a decision
  6497. 4:32:36Tre by taking only one feature and we
  6498. 4:32:37will only divide it to one level okay
  6499. 4:32:39one level or one depth that's
  6500. 4:32:42and this is specifically called as stump
  6501. 4:32:45what we are going to do next is that
  6502. 4:32:46from this particular stump okay the
  6503. 4:32:48stump is basically getting created only
  6504. 4:32:51one so that is adab Boost right we say
  6505. 4:32:52it as weak Learners because this is weak
  6506. 4:32:54learner weak learner why there is a
  6507. 4:32:57reason we say this as weak learner so
  6508. 4:33:00only weak learner so that is the first
  6509. 4:33:02thing with respect to uh this particular
  6510. 4:33:06adab boost so the first step is that
  6511. 4:33:07this is a weak learner so for the weak
  6512. 4:33:09learner we basically create a stump
  6513. 4:33:12stump basically means one level decision
  6514. 4:33:14tree that's it based on the information
  6515. 4:33:17gain and entropy I have selected the
  6516. 4:33:18feature and then I just made a decision
  6517. 4:33:21tree with only one level why it is
  6518. 4:33:24called as it is called as weak learner
  6519. 4:33:27okay so that is the reason we use only
  6520. 4:33:28stum that is just a one level decision
  6521. 4:33:31tree now the next step happens is that
  6522. 4:33:33we provide all the specific records to
  6523. 4:33:36this F1 and we train this specific model
  6524. 4:33:39only with one level decision tree we
  6525. 4:33:41train them
  6526. 4:33:42now after we train them let's say that
  6527. 4:33:44we are going to pass all these
  6528. 4:33:45particular records to find out how many
  6529. 4:33:47are correct and how many are wrong this
  6530. 4:33:49decision this decision tree is basically
  6531. 4:33:51giving so let's say that out of this
  6532. 4:33:53entire records one
  6533. 4:33:55record one record was just given as
  6534. 4:33:59wrong let's say that this is the this is
  6535. 4:34:01the record which was given as wrong okay
  6536. 4:34:04so let's say that this record output was
  6537. 4:34:07predicted wrong from this particular
  6538. 4:34:09model only one wrong was there after
  6539. 4:34:11training the model now what we need to
  6540. 4:34:14do in this specific case understand a
  6541. 4:34:16very important thing so let's say that
  6542. 4:34:18we have done this and probably after
  6543. 4:34:20this what we are actually going to do we
  6544. 4:34:22are going to calculate the total error
  6545. 4:34:24so how many error this particular model
  6546. 4:34:26made let's say that in this particular
  6547. 4:34:28case only one was wrong so this was only
  6548. 4:34:31wrong right one was wrong so if I want
  6549. 4:34:35to calculate the total error how will I
  6550. 4:34:37calculate how many how many of them are
  6551. 4:34:39wrong how many of them are wrong only
  6552. 4:34:40one is wrong what is the weight of this
  6553. 4:34:42so I will go and write 1X 7 so this is
  6554. 4:34:45specifically my total error out of this
  6555. 4:34:47specific model which is my stump over
  6556. 4:34:49here okay which is my F1 stop now this
  6557. 4:34:53is my first
  6558. 4:34:54step the second step is that I need to
  6559. 4:34:57see the performance of stump which stump
  6560. 4:34:59this specific stump and the performance
  6561. 4:35:02is basically checked by a formula which
  6562. 4:35:04is 1 by log e 1us total error divided
  6563. 4:35:09total error why we are doing this
  6564. 4:35:11everything will make sense okay in just
  6565. 4:35:13time every every in just a small time
  6566. 4:35:16everything will make sense the first
  6567. 4:35:18step that we do in adaab boost is that
  6568. 4:35:20we try to find out the total error the
  6569. 4:35:22second step we try to find out the
  6570. 4:35:24performance of stump now in this
  6571. 4:35:26particular case it will be 1 by log e 1
  6572. 4:35:29- 1 by 7 / 1X 7 so once I calculate it
  6573. 4:35:35it will be coming as
  6574. 4:35:37895 F2 and F3 see again understand out
  6575. 4:35:42of all these features I found out from
  6576. 4:35:43Information Gain and entropy that this
  6577. 4:35:45is the best feature let's say that I
  6578. 4:35:47have calculated this
  6579. 4:35:49as895 so this is my second step the
  6580. 4:35:51first step is find out the total error
  6581. 4:35:53the second step is performance of stum
  6582. 4:35:55what is te te basically means total
  6583. 4:35:57error te basically means total error now
  6584. 4:36:01see see the steps okay see the steps
  6585. 4:36:03whenever I'm discussing about boosting
  6586. 4:36:05I'm going to combine weak Learners
  6587. 4:36:07together to get a strong learner now
  6588. 4:36:09what is the next step out of this now
  6589. 4:36:11what what will be my third step
  6590. 4:36:13understand over here my third step will
  6591. 4:36:16be to update all these weights and that
  6592. 4:36:19is the reason why I'm calculating this
  6593. 4:36:20total error and performance of Step so
  6594. 4:36:23my third step will basically be new
  6595. 4:36:26sample weight from the decision tree one
  6596. 4:36:29which is my stump so I'll say new sample
  6597. 4:36:32weight is equal to I need to update all
  6598. 4:36:34these weights why I need to update all
  6599. 4:36:36these weights again understand I'll I'll
  6600. 4:36:39talk about it just a second so if I want
  6601. 4:36:41to up update the sample weights first
  6602. 4:36:44update I will do it for correct records
  6603. 4:36:46see for correct records whichever are
  6604. 4:36:49correct like these all records are
  6605. 4:36:51correct these all records are correct
  6606. 4:36:53now when I update the weights of this
  6607. 4:36:55update the weights of this particular
  6608. 4:36:57record it should reduce and when the the
  6609. 4:37:00the wrong records that I have this
  6610. 4:37:02update should increase why because
  6611. 4:37:06because if I increase this weights then
  6612. 4:37:08the wrong records that are there that
  6613. 4:37:11record should go to the next week
  6614. 4:37:12learner that is the reason why I'm doing
  6615. 4:37:14it now how to update this particular
  6616. 4:37:17weights for correct records for correct
  6617. 4:37:19records the formula looks something like
  6618. 4:37:21this weight multiplied by weight
  6619. 4:37:25multiplied by E to the^ of
  6620. 4:37:28minus this specific performance okay
  6621. 4:37:31this specific performance so e to the
  6622. 4:37:33power of PS I'll write performance of
  6623. 4:37:35stump and then I will basically be able
  6624. 4:37:38to write 1X 7 * e to the^ of minus
  6625. 4:37:43895 if I do the calculation everybody
  6626. 4:37:45try to do it the answer will be
  6627. 4:37:4805 now this is for correct records what
  6628. 4:37:50about incorrect records for the
  6629. 4:37:52incorrect
  6630. 4:37:53records the the weights that is going to
  6631. 4:37:56the formula that we going to apply is
  6632. 4:37:58multiplied by E to the^ of plus PS not
  6633. 4:38:02minus PS plus PS so here I'll write 1 by
  6634. 4:38:057 multiplied e to the^ of
  6635. 4:38:08895 so if I go and probably calcul this
  6636. 4:38:12I'm going to get it
  6637. 4:38:13as 349 so this two are the weights that
  6638. 4:38:18I have got that basically means all
  6639. 4:38:20these records now which are correct 1X 7
  6640. 4:38:23the new updated weights will be 05 05
  6641. 4:38:2805
  6642. 4:38:3005 sorry not for the wrong
  6643. 4:38:33records then this will be 05 then 05 and
  6644. 4:38:3805 so let me just see what is 1x 7 so
  6645. 4:38:41here you can see initially it was. 142
  6646. 4:38:45now it has got reduced to 05 because all
  6647. 4:38:47these records are correct but the wrong
  6648. 4:38:50record value is 349 so my weights will
  6649. 4:38:53now become over here as 349 now I will
  6650. 4:38:56just go and go ahead and write over here
  6651. 4:38:58my new weight my new weight is nothing
  6652. 4:39:01but 05
  6653. 4:39:06055
  6654. 4:39:0705 05 05 1 2 how many 1 2 3 okay fourth
  6655. 4:39:14record is here fourth record is there 1
  6656. 4:39:182 3 4 05 05 okay how many records are
  6657. 4:39:22there 1 2 3 4 5 6 7 so my fourth record
  6658. 4:39:27will basically become the new value that
  6659. 4:39:30I'm having is something called as
  6660. 4:39:34349 now tell me guys if I do the
  6661. 4:39:37summation of all these weights is this
  6662. 4:39:39is it one so prob
  6663. 4:39:41no I don't think so it is one because if
  6664. 4:39:44I try to add it up it is not one but if
  6665. 4:39:46I go and see over here these all are one
  6666. 4:39:48if I combine all the things 1 2 3 4 5 6
  6667. 4:39:517 these all are one so here I have need
  6668. 4:39:53to find out my normalized weight now in
  6669. 4:39:55order to find out the normalized
  6670. 4:39:57weight all I have to do is that what I
  6671. 4:40:00have to do because the entire sumission
  6672. 4:40:03should be one so we have to
  6673. 4:40:05normalize now in order to normalize all
  6674. 4:40:08you have to do is that go and find out
  6675. 4:40:10what is the sum of all this things the
  6676. 4:40:12summation of all these things will be
  6677. 4:40:150 649 all you have to do is that divide
  6678. 4:40:18all the numbers
  6679. 4:40:20by 649 divided by
  6680. 4:40:24649
  6681. 4:40:26649 like this divide all the numbers by
  6682. 4:40:28649 and tell me what will be the answer
  6683. 4:40:30that you'll be getting so here your
  6684. 4:40:32normalized weight will now look like
  6685. 4:40:35077 07 and this value will be somewhere
  6686. 4:40:39around uh
  6687. 4:40:41537 I guess in this case then this will
  6688. 4:40:44be 07
  6689. 4:40:47077 here we are going to divide by all
  6690. 4:40:50this 64 649 now this is my normalized
  6691. 4:40:53weight now after you get a normalized
  6692. 4:40:56weight we will try to create something
  6693. 4:40:57called as buckets because see one
  6694. 4:41:00decision tree we have already created
  6695. 4:41:02which is a stump and you know from this
  6696. 4:41:04particular stum what you're going to get
  6697. 4:41:06okay as an output then in the sequential
  6698. 4:41:09model we will go and combine another
  6699. 4:41:11model over here now it's the time that I
  6700. 4:41:13have to create this specific model now
  6701. 4:41:15in order to create this specific model I
  6702. 4:41:17need to provide some specific rows only
  6703. 4:41:19to this model to train because this
  6704. 4:41:21model is giving one wrong now what I
  6705. 4:41:24have to do is that whatever is wrong
  6706. 4:41:26along with other data points I need to
  6707. 4:41:28provide this specific model with those
  6708. 4:41:30records so that this model will be able
  6709. 4:41:33to train on this and probably be able to
  6710. 4:41:35get the output now let's create buckets
  6711. 4:41:38now based on buckets how the buckets
  6712. 4:41:39will be created over here I will take 07
  6713. 4:41:43until
  6714. 4:41:45sorry whatever is the value over here
  6715. 4:41:48normal we value okay so I will start
  6716. 4:41:50creating my buckets buckets basically
  6717. 4:41:52from 0 to
  6718. 4:41:5307 what did I say now for this decision
  6719. 4:41:57tree or stump I need to provide some
  6720. 4:42:00records so the maximum number of record
  6721. 4:42:02that should be going should be the wrong
  6722. 4:42:05records that should go over here now how
  6723. 4:42:07do we decide that okay there should be a
  6724. 4:42:09way that we should be able to say that
  6725. 4:42:11that specific wrong number of Records
  6726. 4:42:13should go to that decision tree so for
  6727. 4:42:16that purpose what we do is that this
  6728. 4:42:18decision tree will randomly create some
  6729. 4:42:20numbers between 0 to 1 randomly create
  6730. 4:42:25those numbers between 0 to 1 and
  6731. 4:42:27whichever bucket it will come in like 07
  6732. 4:42:30to 014 014 to 07 basically means 0 2 1
  6733. 4:42:37then 0 2 1 2 see how the bucket is
  6734. 4:42:40getting cre this value is getting added
  6735. 4:42:42to this so that becomes this bucket 021
  6736. 4:42:45+3 537 how much it is it is nothing but
  6737. 4:42:50470 747 then 747
  6738. 4:42:55to
  6739. 4:42:57751 like this you create all the buckets
  6740. 4:43:00okay you can create all the buckets now
  6741. 4:43:02tell me which record is basically having
  6742. 4:43:04the biggest bucket size obviously this
  6743. 4:43:07record so if I randomly create a number
  6744. 4:43:10between 0 to one what is the highest
  6745. 4:43:13probability that the values will be
  6746. 4:43:15going in so in this particular case most
  6747. 4:43:17of the wrong records will be passed
  6748. 4:43:18along with the other records obviously
  6749. 4:43:20other records there are chances that
  6750. 4:43:22other records will go to the next
  6751. 4:43:24decision tree but understand maximum
  6752. 4:43:26number will go with the wrong records
  6753. 4:43:28because the bucket is high over here so
  6754. 4:43:31the bucket is high over here so most of
  6755. 4:43:32the time this specific record will get
  6756. 4:43:35create selected and then it will be gone
  6757. 4:43:37to the second tree now suppose I have
  6758. 4:43:40this all records
  6759. 4:43:41so this is my first stump this is my
  6760. 4:43:44second stump this is my third stump
  6761. 4:43:47similarly the third stump from the
  6762. 4:43:48second stump whichever wrong records
  6763. 4:43:50will be going maximum number of Records
  6764. 4:43:52will go over here then again it will be
  6765. 4:43:54trained like this we'll be having lot of
  6766. 4:43:56stumps minimum 100 decision trees can be
  6767. 4:43:59added you know that every decision tree
  6768. 4:44:01will give one output for a new test data
  6769. 4:44:03new test data this week learner will
  6770. 4:44:05give one output this week learner will
  6771. 4:44:07give one output this week learner and
  6772. 4:44:09this will week learner will be giving
  6773. 4:44:10one output obviously the time complexity
  6774. 4:44:12will be more now from this particular
  6775. 4:44:14output suppose it is a binary
  6776. 4:44:16classification I will be getting 0 1 1 1
  6777. 4:44:19so again over here majority voting will
  6778. 4:44:21happen and the output will be one in
  6779. 4:44:24case of regression problem I will be
  6780. 4:44:25having a continuous value over here and
  6781. 4:44:28for this the average average will be
  6782. 4:44:31computed and that will give me an output
  6783. 4:44:33over here so for regression the average
  6784. 4:44:36will be done for classification what
  6785. 4:44:39will happen majority will be be
  6786. 4:44:41happening so everywhere that same part
  6787. 4:44:43will be going on buckets is very much
  6788. 4:44:45simple guys buckets basically means
  6789. 4:44:47based on this weights normalized weight
  6790. 4:44:49we are going to create bucket so that
  6791. 4:44:51whichever records has the highest bucket
  6792. 4:44:53based on this randomly creating code you
  6793. 4:44:55know it will select those specific
  6794. 4:44:57buckets and put it into random Forest
  6795. 4:44:59understand why this bucket size is Big
  6796. 4:45:02the other wrong records which are
  6797. 4:45:03present right suppose they are have more
  6798. 4:45:05than four to five wrong records their
  6799. 4:45:06bucket size will also be bigger and
  6800. 4:45:08because based on this randomly creating
  6801. 4:45:10num between 0 to 1 most of the wrong
  6802. 4:45:12records will be selected and given to
  6803. 4:45:14the second stum similarly this
  6804. 4:45:16particular decision tree will be doing
  6805. 4:45:17some mistakes then that wrong records
  6806. 4:45:19will get updated all the weights will
  6807. 4:45:20get updated and it will be passed to the
  6808. 4:45:22next decision tree guys when I say wrong
  6809. 4:45:24record the output will be same only no
  6810. 4:45:26zero and one so interesting everyone I
  6811. 4:45:29hope you understood so much of maths in
  6812. 4:45:31adab boost and how adab boost actually
  6813. 4:45:33work three main things one is total
  6814. 4:45:35error one is performance of stump and
  6815. 4:45:37one is the new sample weight these
  6816. 4:45:39things are getting calculated extensive
  6817. 4:45:41max normalized weight was basically used
  6818. 4:45:43because the sum of all these weights are
  6819. 4:45:45approximately equal to one when boosting
  6820. 4:45:48why not take the last output no no no we
  6821. 4:45:50have to give the importance of every
  6822. 4:45:52decision tree output every decision tree
  6823. 4:45:55output are important okay let me talk
  6824. 4:45:57about one model which is called as
  6825. 4:45:59blackbox model versus white box what is
  6826. 4:46:03the difference between blackbox model
  6827. 4:46:04and white box if I take an example of
  6828. 4:46:07linear regression tell me what kind of
  6829. 4:46:09model it is is is it a white box model
  6830. 4:46:12or black box if I take an example of
  6831. 4:46:14random
  6832. 4:46:15Forest is this a white box or black box
  6833. 4:46:18if I take an example of decision tree it
  6834. 4:46:21is a white box of blackbox model if I
  6835. 4:46:23take an example of a Ann is it a white
  6836. 4:46:26box of blackbox model linear regression
  6837. 4:46:28is basically called as an wide Box model
  6838. 4:46:30because here you can basically visualize
  6839. 4:46:33how the Theta value is basically
  6840. 4:46:35changing and how it is coming to a
  6841. 4:46:36global Minima and all those things in
  6842. 4:46:38random Forest I will say this as
  6843. 4:46:40blackbox model because it is impossible
  6844. 4:46:42to see all the decision tree how it is
  6845. 4:46:44working so that is the reason the maths
  6846. 4:46:46is so complex inside this if I talk
  6847. 4:46:49about decision tree this is basically a
  6848. 4:46:50white box model because in decision tree
  6849. 4:46:52we know how the split are basically
  6850. 4:46:54happening with the help of paper and pen
  6851. 4:46:55you'll be able to do it in the case of
  6852. 4:46:58an Ann this is a blackbox model because
  6853. 4:47:00here you don't know like how many
  6854. 4:47:02neurons are there how they are
  6855. 4:47:03performing and how the weights are
  6856. 4:47:05getting updated so this is the basic
  6857. 4:47:07difference between the blackbox and uh
  6858. 4:47:10uh white box model this entire thing is
  6859. 4:47:13the agenda of today's session so let's
  6860. 4:47:15start uh the first algorithm that we are
  6861. 4:47:17probably going to discuss today is
  6862. 4:47:19something called as K
  6863. 4:47:21means
  6864. 4:47:23clustering K means clustering and this
  6865. 4:47:26is a kind of unsupervised machine
  6866. 4:47:28learning now always remember
  6867. 4:47:31unsupervised machine learning basically
  6868. 4:47:33means that uh the one and the most
  6869. 4:47:35important thing is that in unsupervised
  6870. 4:47:38machine learning
  6871. 4:47:41in unsupervised ml you don't have any
  6872. 4:47:44specific output so you don't have any
  6873. 4:47:46specific output so suppose you have
  6874. 4:47:48feature one and feature two and suppose
  6875. 4:47:50you have datas different different data
  6876. 4:47:53you know and based on this data what we
  6877. 4:47:55do we basically try to create clusters
  6878. 4:47:58this clusters basically says what are
  6879. 4:48:00the similar kind of data so this is what
  6880. 4:48:03we basically do from uh clustering and
  6881. 4:48:06there are various techniques like K
  6882. 4:48:08Mains uh it is hierle clustering and all
  6883. 4:48:10so first of all we'll try to understand
  6884. 4:48:12about K means and how does it
  6885. 4:48:14specifically work it's simple uh suppose
  6886. 4:48:17you have a data points like this okay
  6887. 4:48:20let's say that this is your F1 feature
  6888. 4:48:21F2 feature and based on this in two
  6889. 4:48:23dimensional probably I will be plotting
  6890. 4:48:26this points and suppose this is my
  6891. 4:48:28another points so our main purpose is
  6892. 4:48:31basically to Cluster together in
  6893. 4:48:34different different groups okay so this
  6894. 4:48:36will be my one group and probably the
  6895. 4:48:38other group will be this group right so
  6896. 4:48:40two groups because obviously you can see
  6897. 4:48:42from this clusters here you have two
  6898. 4:48:44similar kind of data which is basically
  6899. 4:48:47grouped together right this is my
  6900. 4:48:49cluster one and this is my cluster 2 let
  6901. 4:48:51me talk about this and why specifically
  6902. 4:48:54it'll be very much useful then we'll try
  6903. 4:48:56to understand about math intuition also
  6904. 4:48:58now always understand guys uh where does
  6905. 4:49:00clustering gets used okay in most of the
  6906. 4:49:03Ensemble techniques I told you about
  6907. 4:49:05custom emble technique right so custom
  6908. 4:49:08emble techniques in custom assemble
  6909. 4:49:11techniques you know whenever we are
  6910. 4:49:13probably creating a model first of all
  6911. 4:49:15on our data set what we do is that we
  6912. 4:49:18create clusters so suppose this is my
  6913. 4:49:20data set during my model creation the
  6914. 4:49:22first algorithm we will probably apply
  6915. 4:49:24will be clustering algorithm and after
  6916. 4:49:26that it is obviously good that we can
  6917. 4:49:28apply regression or classification
  6918. 4:49:30problem suppose in this clustering I
  6919. 4:49:32have two or three groups let's say that
  6920. 4:49:34I have two or three groups over here for
  6921. 4:49:36each group we can apply a separate
  6922. 4:49:40supervis machine learning algorithm if
  6923. 4:49:42we know the specific output that we
  6924. 4:49:44really want to take ahead I'll talk
  6925. 4:49:46about this and uh give you some of the
  6926. 4:49:48examples as I go ahead now let's go on
  6927. 4:49:51go ahead and focus more on understanding
  6928. 4:49:53how does kin's clustering algorithm work
  6929. 4:49:56so let's go over here the word K means
  6930. 4:49:59has this K value this K are nothing but
  6931. 4:50:02this K basically means centroids K
  6932. 4:50:05basically means centroids so suppose if
  6933. 4:50:08I have a data set which looks like this
  6934. 4:50:10let's say that this is my data set now
  6935. 4:50:12over here just by seeing the data set
  6936. 4:50:14what are the possible groups you think
  6937. 4:50:16definitely you'll be saying K is equal
  6938. 4:50:18to 2 So when you say k is equal to two
  6939. 4:50:20that basically means you will be able to
  6940. 4:50:22get two groups like this and each and
  6941. 4:50:24every group will be having a centroid a
  6942. 4:50:28centroid Point here also there will be a
  6943. 4:50:30centroid point so this centroid will
  6944. 4:50:32determine basically this is a separate
  6945. 4:50:34group over here this is a separate group
  6946. 4:50:36over here so over here here you can
  6947. 4:50:38definitely say that fine this is two
  6948. 4:50:40groups but but how do we come to a
  6949. 4:50:41conclusion that there is only two groups
  6950. 4:50:44okay we cannot just directly say that
  6951. 4:50:46okay we'll try to just by seeing the
  6952. 4:50:48data because your data will be having a
  6953. 4:50:50high dimension data right right now I'm
  6954. 4:50:52just showing your two Dimension data but
  6955. 4:50:55for a high dimension data definitely
  6956. 4:50:56you'll not be able to see the data
  6957. 4:50:58points how it is plotted so how do you
  6958. 4:51:00come to a conclusion that only two
  6959. 4:51:02groups are there so for this there is
  6960. 4:51:03some steps that we basically perform in
  6961. 4:51:05K mins the first step is that we try
  6962. 4:51:08with different K values we try with
  6963. 4:51:11different K values and which is the
  6964. 4:51:13suitable K value K is nothing but
  6965. 4:51:15centroids okay it is nothing but
  6966. 4:51:18centroids we try with different
  6967. 4:51:20different centroids in this particular
  6968. 4:51:22case let's say that I have this
  6969. 4:51:24particular data point and I actually
  6970. 4:51:27start with k is equal 1 or 2 or 3 any
  6971. 4:51:29one you want let's say that I'm going to
  6972. 4:51:31start with k is equal 2 how to come up
  6973. 4:51:34with this K is equal to 2 as a perfect
  6974. 4:51:37value that I'll talk about it we need to
  6975. 4:51:39know there is a concept which is called
  6976. 4:51:41as within cluster sum of square so when
  6977. 4:51:43we try different K values let's say that
  6978. 4:51:45for K is equal to 2 what will happen the
  6979. 4:51:47first step we select a we try K values
  6980. 4:51:50so let's say that we are considering K
  6981. 4:51:52is equal to 2 the second step is that we
  6982. 4:51:54initialize K number of centroids now in
  6983. 4:51:57this particular case I know my K value
  6984. 4:51:59is 2 so we will be initializing randomly
  6985. 4:52:02let's say that K is equal to 2 so what
  6986. 4:52:05we can actually do let's say that this
  6987. 4:52:07is this is my one centroid I will I'll
  6988. 4:52:09put it in another color so this will be
  6989. 4:52:11my one centroid and let's say that this
  6990. 4:52:13is my another centroid so I have
  6991. 4:52:15initialized two centroids randomly in
  6992. 4:52:17this space now after this particular
  6993. 4:52:19centroid what we have to do is that
  6994. 4:52:21after initializing this centroid what we
  6995. 4:52:23have to do is that we have to basically
  6996. 4:52:26find out which points are near to the
  6997. 4:52:29centroid and which points are near to
  6998. 4:52:31this centroid now in order to find out
  6999. 4:52:33it is a very easy step we can basically
  7000. 4:52:35use ukan distance to find out the
  7001. 4:52:38distance between the points in an easy
  7002. 4:52:40way if I really want to show you that
  7003. 4:52:44you know like how many points I want to
  7004. 4:52:46in an easy way what I can do I can
  7005. 4:52:48basically draw a straight line over here
  7006. 4:52:50let's say that I'm drawing a straight
  7007. 4:52:51line over here in another color I can
  7008. 4:52:54draw a straight line and I can also draw
  7009. 4:52:56one parallel line like this so This
  7010. 4:52:58basically indicates that whichever
  7011. 4:53:01points you see over here suppose if I
  7012. 4:53:03draw a straight line in between all
  7013. 4:53:05these points you will be able to see
  7014. 4:53:07that let's say that I'm drawing one more
  7015. 4:53:09parallel line
  7016. 4:53:11which is intersecting together so from
  7017. 4:53:14this you can definitely find out let's
  7018. 4:53:16say that these are all my points that
  7019. 4:53:17are nearer to this green line Green
  7020. 4:53:20Point so what I'm actually going to do
  7021. 4:53:21in this particular case all these points
  7022. 4:53:24that you are seeing near the green it
  7023. 4:53:26will become green color so that
  7024. 4:53:28basically means this is basically nearer
  7025. 4:53:30to this centroid and whichever points
  7026. 4:53:33are nearer to this particular point that
  7027. 4:53:35will become red point so that basically
  7028. 4:53:38means this belongs to this group okay
  7029. 4:53:40this belongs to this group so I hope
  7030. 4:53:42everybody's clear till here then what
  7031. 4:53:44will happen is that this summation of
  7032. 4:53:48all the values then we initialize the K
  7033. 4:53:51number of centroids that is done then we
  7034. 4:53:53try to calculate the distance we try to
  7035. 4:53:55find out which all points is nearer to
  7036. 4:53:57the centroid let's say that this is my
  7037. 4:53:58one centroid this is my another centroid
  7038. 4:54:01and we have seen that okay these all
  7039. 4:54:02points belong to this centroid it near
  7040. 4:54:05to this particular centroid so this is
  7041. 4:54:07becoming red so that is based on the
  7042. 4:54:09shortage distance and here it is
  7043. 4:54:11becoming green now the next step let's
  7044. 4:54:13see what is the next step after this so
  7045. 4:54:15I am going to remove this thing now the
  7046. 4:54:17next step will be that the entire points
  7047. 4:54:20that is in red color all the average
  7048. 4:54:22will be taken so here again the average
  7049. 4:54:25will be taken now third step here I'm
  7050. 4:54:28going to write here we are going to
  7051. 4:54:30compute the average the reason we
  7052. 4:54:32compute the average is that because we
  7053. 4:54:34need to update the centroid so compute
  7054. 4:54:37the average to update centroid to update
  7055. 4:54:40centroids so here you'll be able to see
  7056. 4:54:42that what I'm actually doing as soon as
  7057. 4:54:45we compute the average this centroid is
  7058. 4:54:47going to move to some other location so
  7059. 4:54:50what location it will move it will
  7060. 4:54:51obviously become somewhere in Center so
  7061. 4:54:53here now I'm going to rub this and now
  7062. 4:54:56my new centroid will be this point where
  7063. 4:54:58I am actually going to draw like this
  7064. 4:55:00let's say this is my new centroid now
  7065. 4:55:02similarly this thing will happen with
  7066. 4:55:04respect to the green color so with
  7067. 4:55:06respect to the green color also it will
  7068. 4:55:08happen and this green will also Al get
  7069. 4:55:10updated so I'm going to rub this and
  7070. 4:55:12this will be my new Green Point which
  7071. 4:55:14will get updated over here then again
  7072. 4:55:16what will happen again the distance will
  7073. 4:55:18be calculated and again a perpendicular
  7074. 4:55:20line will be calculated here you can see
  7075. 4:55:22that now all the points are towards
  7076. 4:55:25there okay again the centroid based on
  7077. 4:55:27this particular distance again it will
  7078. 4:55:29be calculated and here you can see that
  7079. 4:55:31all the points are in its own location
  7080. 4:55:33so here now no update will actually
  7081. 4:55:36happen let's say that there was one
  7082. 4:55:38point which was red color over here
  7083. 4:55:41then this would have become green color
  7084. 4:55:42but since the updation has happened
  7085. 4:55:44perfectly we are not going to update it
  7086. 4:55:46and we are not going to update the
  7087. 4:55:48centroid right so now you can understand
  7088. 4:55:51that yes now we have actually got the
  7089. 4:55:53perfect centroid and now this will be
  7090. 4:55:56considered as one group and this will be
  7091. 4:55:58basically considered as the another
  7092. 4:56:00group it will not intersect but right by
  7093. 4:56:02default here intersection is happening
  7094. 4:56:05so I hope everybody's understood the
  7095. 4:56:07steps that you have actually followed in
  7096. 4:56:09initializing the centroids in updating
  7097. 4:56:12the centroids and in updating the points
  7098. 4:56:14is it clear everybody with respect to K
  7099. 4:56:17means now let's discuss about one
  7100. 4:56:20point how do we decide this K value okay
  7101. 4:56:24how do we decide this K value so for
  7102. 4:56:26deciding the K value there is a concept
  7103. 4:56:27which is called as elbow method so here
  7104. 4:56:31I'm going to basically Define my elbow
  7105. 4:56:32method now elbow method says something
  7106. 4:56:35very much important because this will
  7107. 4:56:37actually help us to find out what is the
  7108. 4:56:40optimized K value whether the K value
  7109. 4:56:42should be two whether uh the K value is
  7110. 4:56:45going to be three whether the K value is
  7111. 4:56:47going to become four and always
  7112. 4:56:49understand suppose this is my data set
  7113. 4:56:51suppose this is my data set initially
  7114. 4:56:53let's say that I have my data points
  7115. 4:56:54like this we cannot go ahead and
  7116. 4:56:57directly say say that okay K is equal to
  7117. 4:56:592 is going to work so obviously we are
  7118. 4:57:01going to go with iteration for I is
  7119. 4:57:04equal to probably 1 to 10 I'm going to
  7120. 4:57:06move towards iteration from 1 to 10
  7121. 4:57:09let's say so for every iteration we will
  7122. 4:57:11construct a graph with respect to K
  7123. 4:57:14value and with respect to something
  7124. 4:57:16called as W CSS now what is this W CSS W
  7125. 4:57:20CSS basically means within cluster sum
  7126. 4:57:23of
  7127. 4:57:24square okay this is the meaning of wcss
  7128. 4:57:27within cluster sum of square now let's
  7129. 4:57:30say that initially we start with one
  7130. 4:57:33centroid so one centroid let's say it is
  7131. 4:57:35initialized here one centroid is
  7132. 4:57:37basically initialized here if we go and
  7133. 4:57:39compute the distance
  7134. 4:57:40between each and every points to the
  7135. 4:57:43centroid and if we try to find out the
  7136. 4:57:45distance will the distance value be
  7137. 4:57:47greater or it will be smaller will it be
  7138. 4:57:50smaller or greater tell me if you try to
  7139. 4:57:53calculate this distance from this
  7140. 4:57:55centroid to every point this is what is
  7141. 4:57:57within cluster sum of square it will
  7142. 4:58:00always be very very much greater so
  7143. 4:58:02let's say that my first point has come
  7144. 4:58:04somewhere here it is going to be
  7145. 4:58:06obviously greater let's say that my
  7146. 4:58:07first point is coming over here find
  7147. 4:58:10So within K is equal to 1 initially we
  7148. 4:58:12took and we found out the distance of w
  7149. 4:58:14CSS and it is a very huge value okay
  7150. 4:58:17because we're going to compute the
  7151. 4:58:18distance between each and every point to
  7152. 4:58:20the centroid now the next thing that I'm
  7153. 4:58:23actually going to do is that now we'll
  7154. 4:58:26go with next value that is K is equal to
  7155. 4:58:282 now in K is equal to 2 I will
  7156. 4:58:31initialize two points okay I will
  7157. 4:58:34initialize two points and then probably
  7158. 4:58:36I will do the entire process which I
  7159. 4:58:38have written on the top now tell me me
  7160. 4:58:40whichever points is nearer to this green
  7161. 4:58:42point if we compute the distance and
  7162. 4:58:46whichever points is nearer to the red
  7163. 4:58:48point if you compute the distance like
  7164. 4:58:52this now this summation of the distance
  7165. 4:58:55will be lesser than the previous W CSS
  7166. 4:58:57or not obviously it is going to be
  7167. 4:59:00lesser than the previous W CSS so what
  7168. 4:59:02I'm actually going to do probably with K
  7169. 4:59:04is equal to 2 your value may come
  7170. 4:59:06somewhere here then with K is equal to 3
  7171. 4:59:09your value May come somewhere here then
  7172. 4:59:10K is equal to 4 will come here to 5 6
  7173. 4:59:13like this it will go so here if I
  7174. 4:59:15probably join this line you'll be able
  7175. 4:59:17to see that there will be an Abrupt
  7176. 4:59:19changes in the W CSS value in the wcss
  7177. 4:59:23value there will be an Abrupt changes
  7178. 4:59:25and this this is basically called as
  7179. 4:59:27elbow curve now why we say it as elbow
  7180. 4:59:30curve because it is in the shape of
  7181. 4:59:32elbow and here at one specific point
  7182. 4:59:34there will be an Abrupt change and then
  7183. 4:59:36it will be straight so that is the
  7184. 4:59:38reason why we basically say this as
  7185. 4:59:41elbow okay so this is a very important
  7186. 4:59:43thing see in finding the K value we use
  7187. 4:59:46elbow method but for validating purpose
  7188. 4:59:49how do we validate that this model is
  7189. 4:59:52performing well we use silard score that
  7190. 4:59:54I'll show you just in some time but
  7191. 4:59:57understand that in K means clustering we
  7192. 5:00:00need to update the centroids and based
  7193. 5:00:02on that we calculate the distance and as
  7194. 5:00:05the K value keep on increasing you'll be
  7195. 5:00:07able to see that the distance will
  7196. 5:00:09become normal or the wcss value will
  7197. 5:00:12become normal and then we really need to
  7198. 5:00:14find out which is the phys K value where
  7199. 5:00:17the abrupt change see over here suppose
  7200. 5:00:20abrupt change is there and then it is
  7201. 5:00:21normal then I will probably take this as
  7202. 5:00:24my K value so obviously the model
  7203. 5:00:26complexity will be high because we are
  7204. 5:00:28going to check with respect to different
  7205. 5:00:30different K values and wcss values and
  7206. 5:00:33this basically means that the value that
  7207. 5:00:36we'll probably get first of all we need
  7208. 5:00:38to construct this elbow curve then see
  7209. 5:00:40the changes where it is basically
  7210. 5:00:42happening we'll need to find out the
  7211. 5:00:43abrupt change and once we get the abrupt
  7212. 5:00:46change we basically say that this may be
  7213. 5:00:49the K value so K is equal to 4 as an
  7214. 5:00:52example I'm telling you so unless and
  7215. 5:00:54until if you really want to find the
  7216. 5:00:56cluster it is very much simple we take a
  7217. 5:00:59k value we initialize K number of
  7218. 5:01:01centroids we compute the average to
  7219. 5:01:03update the centroids then again we try
  7220. 5:01:05to find out the distance try to see that
  7221. 5:01:07whether any points has changed and
  7222. 5:01:08continue that process unless and until
  7223. 5:01:10we get separate groups okay so this is
  7224. 5:01:14the entire funa of claim in clustering
  7225. 5:01:16so finally you'll be able to see that
  7226. 5:01:18with respect to the K value we will be
  7227. 5:01:20able to get that many number of groups
  7228. 5:01:22if my K value is four that basically
  7229. 5:01:24means I will be probably getting four
  7230. 5:01:26different groups like this 1 two right
  7231. 5:01:30three like this and four I will be
  7232. 5:01:32getting four groups like this with K is
  7233. 5:01:34equal to 4 that basically means K is
  7234. 5:01:35equal to four clusters and every group
  7235. 5:01:38will be having its own centroids okay
  7236. 5:01:41every group will be having okay
  7237. 5:01:42centroids are very much important yes
  7238. 5:01:45I'll try to show you in the coding also
  7239. 5:01:47guys let's go towards the second
  7240. 5:01:48algorithm the second algorithm that we
  7241. 5:01:51will be probably discussing is called as
  7242. 5:01:54hierarchical clustering now hierarchal
  7243. 5:01:56clustering is very much simple guys all
  7244. 5:01:58you have to do is that let's say this is
  7245. 5:02:00your data points this is your data
  7246. 5:02:01points and this is my P1 let's say P2
  7247. 5:02:04now hierle clustering says that we will
  7248. 5:02:07go step by step the first thing is that
  7249. 5:02:10we will try to find out the most nearest
  7250. 5:02:12Value let's say this is my X and Y let's
  7251. 5:02:15say these are my points like this is my
  7252. 5:02:18P1 point this is my P2 point this is my
  7253. 5:02:21P3 point this is my P4 Point P5 Point P6
  7254. 5:02:25point p7 point okay so these are my
  7255. 5:02:28points that I have actually named over
  7256. 5:02:29here let's say that this may be the
  7257. 5:02:31nearest point to each other so what it
  7258. 5:02:32will do it will combine this together
  7259. 5:02:34into one cluster this we have computed
  7260. 5:02:37the distance so it will C create one
  7261. 5:02:39cluster now what will happen on the
  7262. 5:02:41right hand side there will be another
  7263. 5:02:42notation which you may be using in
  7264. 5:02:45connecting all the points one so suppose
  7265. 5:02:46this is my P1 this is my P2 this is my
  7266. 5:02:50P3 P4 let's say that I have this many
  7267. 5:02:53points and probably I will also try to
  7268. 5:02:56make
  7269. 5:02:57p7 so these are my points p7 now you
  7270. 5:03:00know that the nearest point that we are
  7271. 5:03:02having okay this will probably be
  7272. 5:03:04distance 1 2 3 this is distance okay 4 5
  7273. 5:03:096 like this we have lot of distance so
  7274. 5:03:12hierle clustering will first of all find
  7275. 5:03:14out the nearest point and try to compute
  7276. 5:03:17the distance between them and just try
  7277. 5:03:18to combine them together into one what
  7278. 5:03:21do we do we basically combine them into
  7279. 5:03:23one group okay so P1 and P2 has been
  7280. 5:03:26combined let's say then it'll go and
  7281. 5:03:29find out the other nearest point so
  7282. 5:03:31let's say P6 and p7 are near so they are
  7283. 5:03:33also going to combine into one group so
  7284. 5:03:35once they combine into one group then we
  7285. 5:03:37have P6 and p7 which will be obviously L
  7286. 5:03:40greater than the previous distance and
  7287. 5:03:42we may get this kind of computation and
  7288. 5:03:44another combination or cluster will form
  7289. 5:03:47get formed over here then you have seen
  7290. 5:03:49that okay P3 and P5 are nearer to each
  7291. 5:03:52other so we are going to combine this so
  7292. 5:03:54I'm going to basically combine P3 and
  7293. 5:03:57P5 okay and let's say that this distance
  7294. 5:03:59is greater than the previous one because
  7295. 5:04:02we are basically going to sh start with
  7296. 5:04:03the shortest distance and then we are
  7297. 5:04:05going to capture the longest distance
  7298. 5:04:07now this is done now you can see that
  7299. 5:04:08the next point that is near right to
  7300. 5:04:11this particular group is P4 so we are
  7301. 5:04:13going to combine this together into one
  7302. 5:04:15group so once we combine this into one
  7303. 5:04:17group this P4 will get connected like
  7304. 5:04:20this let's say it is getting connected
  7305. 5:04:23like this P4 has got connected then what
  7306. 5:04:25is the nearest Point whether it is P6 p7
  7307. 5:04:28group or P1 P2 obviously here you can
  7308. 5:04:30see that P1 P2 is there so I am probably
  7309. 5:04:32going to combine this group together
  7310. 5:04:34that basically means P1 P2 let's say I'm
  7311. 5:04:38just going to combine this group group
  7312. 5:04:40together again circle is coming so I
  7313. 5:04:42will make a dot let's say I'm going to
  7314. 5:04:43combine this group together because
  7315. 5:04:45these are my nearest groups so what will
  7316. 5:04:47happen P1 and P2 will get combined to P5
  7317. 5:04:50sorry P4 P5 this one so I will be
  7318. 5:04:53getting another line like this and then
  7319. 5:04:55finally you'll be seeing that P6 p7 is
  7320. 5:04:57the nearest group to this so this will
  7321. 5:05:00totally get combined and it may look
  7322. 5:05:02something like this so this will become
  7323. 5:05:05a total group like
  7324. 5:05:07this so all the groups are combined so
  7325. 5:05:10finally you'll be able to see that there
  7326. 5:05:11will be one more line which will get
  7327. 5:05:13combined like
  7328. 5:05:14this this is basically called as
  7329. 5:05:17dendogram dendogram okay which is like
  7330. 5:05:21bottom root to top now the question
  7331. 5:05:24arises is that how do you find that how
  7332. 5:05:25many groups should be here how do you
  7333. 5:05:27find out that how many groups should be
  7334. 5:05:29here the funa is very much Clear guys in
  7335. 5:05:32this is that you need
  7336. 5:05:34to find the longest
  7337. 5:05:41vertical line you need to find out the
  7338. 5:05:43longest vertical line that has no
  7339. 5:05:46horizontal line pass through it no
  7340. 5:05:49horizontal
  7341. 5:05:51line passed through it this is very much
  7342. 5:05:54important that has no horizontal line
  7343. 5:05:56pass through it now what this is
  7344. 5:05:58basically meaning is that I will try to
  7345. 5:06:00find out the longest line longest
  7346. 5:06:03vertical line in such a way that none of
  7347. 5:06:06the horizontal line passes through it
  7348. 5:06:07what is horizontal line suppose if I
  7349. 5:06:09consider this vertical line This
  7350. 5:06:11vertical line over here if you see that
  7351. 5:06:13if I extend this green line it is
  7352. 5:06:15passing through this if I extend this
  7353. 5:06:17line it is passing through this right if
  7354. 5:06:20I'm extending this line it is passing
  7355. 5:06:21through this right so out of this the
  7356. 5:06:25longest line that may be passing in such
  7357. 5:06:27a way that no horizontal line probably
  7358. 5:06:29is this line that I can actually see so
  7359. 5:06:31what you do over here is that you
  7360. 5:06:33basically just create a straight line
  7361. 5:06:35over this and then you try to find out
  7362. 5:06:37that how many clusters it will be there
  7363. 5:06:39by understanding that how many lines it
  7364. 5:06:41is passing through if it is passing
  7365. 5:06:42through this one line two line three
  7366. 5:06:44line four line that basically means your
  7367. 5:06:47clusters will be four
  7368. 5:06:49clusters this is how we basically do the
  7369. 5:06:52calculation in heral clustering again
  7370. 5:06:56here it may not be the perfect line I've
  7371. 5:06:58just drawn with some assumptions but if
  7372. 5:07:00you are trying to do this probably you
  7373. 5:07:02have to do in this specific way okay
  7374. 5:07:04I've already uploaded a lot of practical
  7375. 5:07:06videos with respect to highill
  7376. 5:07:08clustering and all now now tell me
  7377. 5:07:11maximum effort or maximum time is taken
  7378. 5:07:15by is taken
  7379. 5:07:18by K
  7380. 5:07:20means or hierle clustering this is a
  7381. 5:07:25question for you yes guys number of
  7382. 5:07:26clusters may be three but here I'm just
  7383. 5:07:29showing you that how many lines it may
  7384. 5:07:31be passed by how do you basically
  7385. 5:07:34determine whether maximum time will be
  7386. 5:07:36taken by kin or Hier clustering this is
  7387. 5:07:38an interview question the maximum time
  7388. 5:07:40that will be taken is by hierarchical
  7389. 5:07:45clustering why because let's say that I
  7390. 5:07:48have many many many data points at that
  7391. 5:07:51point of time hierle clustering will
  7392. 5:07:53keep on constructing this kind of
  7393. 5:07:55dendograms and it will be taking many
  7394. 5:07:58many many time lot time right so hierle
  7395. 5:08:02clustering will take more time maximum
  7396. 5:08:05time that it is going to basically take
  7397. 5:08:07so it is very much important that that
  7398. 5:08:09you understand which is making basically
  7399. 5:08:12taking more time so if your data set is
  7400. 5:08:15small you may go ahead with hierle
  7401. 5:08:18clustering if your data set is large go
  7402. 5:08:21with K means clustering go with K means
  7403. 5:08:23clustering in short both will take more
  7404. 5:08:25time but K Min will perform better than
  7405. 5:08:28hle clustering see guys you will be
  7406. 5:08:30forming this kind of dendograms right
  7407. 5:08:33and just imagine if you have 10 features
  7408. 5:08:34and many data points how you're going to
  7409. 5:08:37do it it will be a cubers some process
  7410. 5:08:40you'll not be even able to see this
  7411. 5:08:42dendogram properly and manually
  7412. 5:08:44obviously you cannot do it so this was
  7413. 5:08:46with respect to K means clust swing and
  7414. 5:08:49H mean clust swing I hope everybody's
  7415. 5:08:51understood now the next topic that we'll
  7416. 5:08:53focus on is that how do we
  7417. 5:08:56validate see how do we validate a
  7418. 5:08:59classification problem we use
  7419. 5:09:00performance metric like confusion Matrix
  7420. 5:09:03accuracy um different different true
  7421. 5:09:05positive rate Precision recall but how
  7422. 5:09:07do we validate clustering model Model S
  7423. 5:09:10we are going to use something called as
  7424. 5:09:12so we are going to basically use
  7425. 5:09:14something called as
  7426. 5:09:15Sil score I'll show you what Sid score
  7427. 5:09:19is I'm going to just open the Wikipedia
  7428. 5:09:21so this is how a CID score looks like a
  7429. 5:09:25very very amazing topic okay how do we
  7430. 5:09:28validate whether my model basically has
  7431. 5:09:32perfect three or four model perfect
  7432. 5:09:35three suppose if I find out my K value
  7433. 5:09:37is three how do we find out now see one
  7434. 5:09:40more one more issue with K means one
  7435. 5:09:42issue with K means which I forgot to
  7436. 5:09:44tell you let's say that I have a data
  7437. 5:09:46point which looks like this and suppose
  7438. 5:09:49I have some data points like this I have
  7439. 5:09:51some data points which looks like this
  7440. 5:09:55let's say I have like this now in this
  7441. 5:09:58one issue will be that suppose I try to
  7442. 5:10:01make a cluster over here obviously
  7443. 5:10:03you'll be saying my K value will be two
  7444. 5:10:05okay in this particular case suppose
  7445. 5:10:07this is one cluster this is my another
  7446. 5:10:08cluster
  7447. 5:10:10right because of my wrong initialization
  7448. 5:10:13of the points okay understand because
  7449. 5:10:16suppose if I initialize just randomly
  7450. 5:10:18some centroids like this then what may
  7451. 5:10:20happen is that there is a possibility
  7452. 5:10:22that we may also have three clusters
  7453. 5:10:24like like like this kind of clusters one
  7454. 5:10:27cluster will be here one cluster will be
  7455. 5:10:29here one cluster will be here so this
  7456. 5:10:32initialization of the centroids one
  7457. 5:10:35condition is that it should be very very
  7458. 5:10:37far if we initialize our centroids very
  7459. 5:10:41very far at that point of time we will
  7460. 5:10:43be able to find the centroid exactly in
  7461. 5:10:46the center because it will keep on
  7462. 5:10:47updating it'll keep on going ahead right
  7463. 5:10:50but if we don't initialize that very far
  7464. 5:10:53then there will be a situation that
  7465. 5:10:55probably if I wanted to get only the
  7466. 5:10:57real thing was to get only two centroids
  7467. 5:10:59I was probably getting three centroids
  7468. 5:11:01right so this is a problem so for this
  7469. 5:11:04there is an algorithm which is called as
  7470. 5:11:06K means Plus+ and what this K means
  7471. 5:11:08Plus+ will do which I will probably show
  7472. 5:11:10you in Practical this will make sure
  7473. 5:11:12that all the centroids that are
  7474. 5:11:14initialized it is very very
  7475. 5:11:16far okay all the in centroids that is
  7476. 5:11:19basically there it is initialized very
  7477. 5:11:21very far we'll see that in practical
  7478. 5:11:23application where specifically those
  7479. 5:11:26centroids are basically used now let me
  7480. 5:11:28go ahead and let me show
  7481. 5:11:30you with respect to Sid clust string now
  7482. 5:11:34what is the solo color string I'm going
  7483. 5:11:36to explain you in an amazing way this is
  7484. 5:11:38important
  7485. 5:11:39if someone says you how do we validate
  7486. 5:11:43how do we validate cluster
  7487. 5:11:46model then at that point of time we
  7488. 5:11:48basically use this site it will be used
  7489. 5:11:51in it will be used with respect
  7490. 5:11:55to it will be used with respect to K
  7491. 5:11:58means it can be used in hierle mean
  7492. 5:12:00right if you want to validate how do we
  7493. 5:12:03validate okay that is what we are
  7494. 5:12:04basically going to see over here now in
  7495. 5:12:08C's clustering
  7496. 5:12:09what are the most important things the
  7497. 5:12:12first and the most important thing is
  7498. 5:12:13that we will try to find out we will try
  7499. 5:12:16to find out a ofi we will try to find
  7500. 5:12:19out a of I now what is this a ofi see
  7501. 5:12:22this a ofi that you basically see a ofi
  7502. 5:12:25is nothing but see three major steps
  7503. 5:12:28happens in order to validate cluster
  7504. 5:12:30model with the help of solo first thing
  7505. 5:12:33is that I will probably take one cluster
  7506. 5:12:36okay there will be one point
  7507. 5:12:39which will be my centroid let's say and
  7508. 5:12:42then what I'm going to do I'm just going
  7509. 5:12:44to whatever points are there inside this
  7510. 5:12:46cluster I'm going to compute the
  7511. 5:12:49distance between them so I'm going to do
  7512. 5:12:52the summation and I'm also going to do
  7513. 5:12:54the average of all this distance so here
  7514. 5:12:57you can see that when I said distance of
  7515. 5:12:59I comma J I basically means this point J
  7516. 5:13:03basically means all these points I is
  7517. 5:13:06nothing but it is the centroid so here
  7518. 5:13:08is nothing but this this is the centroid
  7519. 5:13:09let's say that I'm having the centroid
  7520. 5:13:11so I'm going to compute all the distance
  7521. 5:13:13over here which is mentioned by this and
  7522. 5:13:15this value that you see that I'm
  7523. 5:13:17actually dividing by C of I minus one in
  7524. 5:13:20Short I am actually trying to calculate
  7525. 5:13:22the average
  7526. 5:13:24distance so this is the first point
  7527. 5:13:26where I'm actually Computing the a ofi
  7528. 5:13:28now similarly what I will do is
  7529. 5:13:31that what I will do is that the next
  7530. 5:13:34point will be that suppose I have
  7531. 5:13:36computed a ofi the next the next that we
  7532. 5:13:39need to compute is B ofi now what is b
  7533. 5:13:41ofi b ofi is nothing but there will be
  7534. 5:13:44multiple clusters in a k means problem
  7535. 5:13:47statement we will try to find out the
  7536. 5:13:50nearest cluster okay suppose let's say
  7537. 5:13:52that this is the nearest cluster and in
  7538. 5:13:54this I have all the variety of points
  7539. 5:13:58then B ofi basically says that I will
  7540. 5:14:00try to compute the distance between each
  7541. 5:14:03point and the other point in this
  7542. 5:14:06centroid sorry in this cluster so this
  7543. 5:14:08is my cluster one this is my cluster two
  7544. 5:14:12so what I'm actually going to do is that
  7545. 5:14:14here I'm going to compute the distance
  7546. 5:14:16between this point to this point then
  7547. 5:14:17this point to this point then this point
  7548. 5:14:20to this point this point to this point
  7549. 5:14:22this point to this point this point to
  7550. 5:14:24this point every point I'm actually
  7551. 5:14:26going to compute the distance once this
  7552. 5:14:28point is done we will go ahead with the
  7553. 5:14:30next point and we'll try to compute the
  7554. 5:14:31distance and once we get all this
  7555. 5:14:34particular distance what we are going to
  7556. 5:14:35do we are going to do the average of
  7557. 5:14:37them average
  7558. 5:14:39now tell me if I try to find out the
  7559. 5:14:42relationship between a of I and B of I
  7560. 5:14:45if my cluster model is good will a of
  7561. 5:14:50I will be greater than b of I or
  7562. 5:14:54will B of I will be greater than a ofi
  7563. 5:14:58if I have a good clustering model if I
  7564. 5:15:01have a good clustering model will a of I
  7565. 5:15:05is greater than b of I will be greater
  7566. 5:15:08than b of I or whether B of I will be
  7567. 5:15:10greater than a of I out of this if we
  7568. 5:15:13have a really good model obviously the
  7569. 5:15:16distance between B of I will be greater
  7570. 5:15:19than a of I in a good model that
  7571. 5:15:22basically means if I talk about sloid
  7572. 5:15:24clustering the values will be between -1
  7573. 5:15:27to +1 the more the value is towards +1
  7574. 5:15:32that basically means the good the model
  7575. 5:15:34is the good the clustering model is the
  7576. 5:15:37more the values towards negative one
  7577. 5:15:39that basically means this condition is
  7578. 5:15:40getting applied now what does this
  7579. 5:15:42condition basically say that basically
  7580. 5:15:43means that this distance is far than the
  7581. 5:15:46cluster distance this is what this
  7582. 5:15:48information is getting portrayed and
  7583. 5:15:51this is the importance of CID
  7584. 5:15:53clustering finally when we apply the
  7585. 5:15:55formula of CID clustering you'll be able
  7586. 5:15:57to see that sloid clustering is nothing
  7587. 5:16:00but let me rub this everything guys for
  7588. 5:16:03you let me just show you what is CID
  7589. 5:16:05clustering CID clustering formula will
  7590. 5:16:08be something like this this B of I so
  7591. 5:16:11here you have solid clustering this is
  7592. 5:16:13the formula B of I minus a of I Max of a
  7593. 5:16:18of I comma B of I if C of I is greater
  7594. 5:16:21than one right so by this you will be
  7595. 5:16:24getting the value between -1 to + 1 and
  7596. 5:16:28more the value is towards + one the more
  7597. 5:16:31good your model is more the values
  7598. 5:16:33towards minus1 more bad your model is
  7599. 5:16:36because if it is towards minus1 that
  7600. 5:16:38basically means your a of I is obviously
  7601. 5:16:41greater than b of I so this is the
  7602. 5:16:43outcome with respect to cot crust string
  7603. 5:16:46if s is equal to zero that basically
  7604. 5:16:47means still your model needs to be uh
  7605. 5:16:50per basically the clustering needs to be
  7606. 5:16:52improved what is I over here I is
  7607. 5:16:54nothing but one data point you you can
  7608. 5:16:56just read this guys data point in I in
  7609. 5:16:59the cluster C of I so I hope everybody's
  7610. 5:17:01understood this now let's go ahead and
  7611. 5:17:03let's discuss about the next topic we
  7612. 5:17:05have obviously finished up solart
  7613. 5:17:07clustering over here let's discuss about
  7614. 5:17:09something called as DB
  7615. 5:17:11scan so for DB scan clustering this is
  7616. 5:17:14an amazing clustering algorithm we'll
  7617. 5:17:17try to understand how to actually do DB
  7618. 5:17:20clustering and probably you'll be able
  7619. 5:17:22to understand a lot of things from this
  7620. 5:17:24now in DB scan clustering what are the
  7621. 5:17:27important things so let's start with
  7622. 5:17:29respect to DB scan clustering and let's
  7623. 5:17:32understand some of the important points
  7624. 5:17:33over here the first point that you
  7625. 5:17:35really need to remember is something
  7626. 5:17:37called as score point points I'll also
  7627. 5:17:39talk about when do you say core points
  7628. 5:17:42or when do you say other points as such
  7629. 5:17:44so the first point that I will probably
  7630. 5:17:46discuss about is something called as Min
  7631. 5:17:49points the second point that I will
  7632. 5:17:51probably discuss about is something
  7633. 5:17:53called as score points the third thing
  7634. 5:17:56that I will probably discuss about is
  7635. 5:17:57something called as border points and
  7636. 5:18:00the fourth point that I will definitely
  7637. 5:18:02talk about is something called as noise
  7638. 5:18:04Point okay guys now tell me in C's
  7639. 5:18:07clustering
  7640. 5:18:09if I have this kind of groups don't you
  7641. 5:18:11think with the help of two different
  7642. 5:18:14clusters I may combine this two like
  7643. 5:18:16this with the help of two different
  7644. 5:18:18clusters I may combine something like
  7645. 5:18:22this right but understand over here what
  7646. 5:18:25what problem is basically happening with
  7647. 5:18:27the second clustering this is actually
  7648. 5:18:30an outliers let's say that let's say one
  7649. 5:18:32thing very nicely I will put okay let's
  7650. 5:18:35say I have one point over here I have
  7651. 5:18:38one point over here here so if I do
  7652. 5:18:39clustering probably I will get one
  7653. 5:18:41cluster
  7654. 5:18:43here and I may get another cluster which
  7655. 5:18:45is somewhere here now understand one
  7656. 5:18:47thing this point is definitely an
  7657. 5:18:50outlier even though this is an outlier
  7658. 5:18:53with the help of K means what I'm
  7659. 5:18:54actually doing I'm actually grouping
  7660. 5:18:56this into another group so can we have a
  7661. 5:18:59scenario wherein a kind of clustering
  7662. 5:19:01algorithm is there where we can leave
  7663. 5:19:03the outlier separately and this outlier
  7664. 5:19:06in this particular algorithm and this is
  7665. 5:19:08B basically uh we will be using DB scan
  7666. 5:19:11to relieve the outlier and this point
  7667. 5:19:13will be called as a noisy Point noisy
  7668. 5:19:15point or I can also say it as an outlier
  7669. 5:19:18so this will be a noise point for this
  7670. 5:19:20kind of algorithm where you want to skip
  7671. 5:19:22the outliers we can definitely use DB
  7672. 5:19:25scan that is density based spatial
  7673. 5:19:27clustering of application with noise a
  7674. 5:19:31very amazing algorithm and definitely I
  7675. 5:19:33have tried using this a lot nowadays I
  7676. 5:19:36don't use K means or Hier means instead
  7677. 5:19:38use this kind of algorithm now see this
  7678. 5:19:41what are the important things over here
  7679. 5:19:42first of all you need to go ahead with
  7680. 5:19:44Min points Min points so first thing is
  7681. 5:19:47that you need to have Min points this
  7682. 5:19:50Min points is a kind of
  7683. 5:19:52hyperparameter this basically says what
  7684. 5:19:55does hyper parameter says and there is
  7685. 5:19:57also a value which is called as
  7686. 5:19:59Epsilon which I forgot I will write it
  7687. 5:20:01down over here this is called as Epsilon
  7688. 5:20:04now what does epsilon mean Epsilon
  7689. 5:20:06basically means if I have a point like
  7690. 5:20:08this
  7691. 5:20:09and if I take Epsilon this is nothing
  7692. 5:20:11but the radius of that specific Circle
  7693. 5:20:13radius of that specific Circle okay so
  7694. 5:20:16Epsilon is nothing but radius over here
  7695. 5:20:19in this specific T what does minimum
  7696. 5:20:21points is equal to 4 mean let's say that
  7697. 5:20:24I have I have taken a point over here
  7698. 5:20:26let's say that this is my
  7699. 5:20:28point and I have drawn a circle which
  7700. 5:20:31looks like this and let's say that this
  7701. 5:20:33is my Epsilon
  7702. 5:20:34value okay this is my Epsilon value if I
  7703. 5:20:37say my Min point point is equal to 4
  7704. 5:20:40which is again a hyper
  7705. 5:20:41parameter that basically means I can if
  7706. 5:20:45I have four at least four points over
  7707. 5:20:47here near to this particular Circle
  7708. 5:20:49based on this Epsilon value then what
  7709. 5:20:52will happen is that this point this red
  7710. 5:20:55point will actually become a core
  7711. 5:20:58point a core point which is basically
  7712. 5:21:01given over here if it has at least that
  7713. 5:21:04many number of Min points inside or near
  7714. 5:21:07to this particular within this
  7715. 5:21:09Epsilon okay within this particular
  7716. 5:21:11cluster suppose this is my cluster with
  7717. 5:21:14the help of Epsilon I have actually
  7718. 5:21:15created it is there a particular unit of
  7719. 5:21:17Epsilon or we simply take the unit of
  7720. 5:21:19distance no Epsilon value will also get
  7721. 5:21:21selected through some way I I'll show
  7722. 5:21:23you I'll show you in the practical
  7723. 5:21:24application don't worry now the next
  7724. 5:21:26thing is that let's say let's say I have
  7725. 5:21:28another another point over here let's
  7726. 5:21:30say that I have another point over here
  7727. 5:21:32and this is my circle with respect to
  7728. 5:21:35Epsilon I have created it let's say that
  7729. 5:21:38here I have only one
  7730. 5:21:41point I have only one point inside this
  7731. 5:21:45particular cluster at that point this
  7732. 5:21:48point becomes something called as border
  7733. 5:21:52Point border Point border point also we
  7734. 5:21:55have discussed over here right so border
  7735. 5:21:58point is also there so here I'm saying
  7736. 5:22:00that at least one at least one if it is
  7737. 5:22:04only one it is present then it will
  7738. 5:22:06become a border point if it has Force
  7739. 5:22:08definitely this will become a core Point
  7740. 5:22:10core Point like how we have this red
  7741. 5:22:11color so and there will be one more
  7742. 5:22:14scenario suppose I have this one cluster
  7743. 5:22:16let's say this is my Epsilon and suppose
  7744. 5:22:19if I don't have any points near this
  7745. 5:22:21then this will definitely become my
  7746. 5:22:23noise point and this noise point will
  7747. 5:22:26nothing be but this will be a
  7748. 5:22:28cluster okay so here I have actually
  7749. 5:22:30discussed about the noise point also so
  7750. 5:22:33I hope everybody is able to understand
  7751. 5:22:34the key terms now what is basically
  7752. 5:22:36happening is that whenever we have a
  7753. 5:22:39noise Point like in this particular
  7754. 5:22:40scenario we have a noise point and we
  7755. 5:22:42don't find any points inside this any
  7756. 5:22:45core point or border point if you don't
  7757. 5:22:47find inside this then it is going to
  7758. 5:22:49just get neglected that basically means
  7759. 5:22:52this is basically treated as an outlier
  7760. 5:22:55I hope everybody is able to understand
  7761. 5:22:57here this point will be treated as an
  7762. 5:22:59outlier or it can also be treated as a
  7763. 5:23:02noise point and this will never be taken
  7764. 5:23:05inside a group okay it will never never
  7765. 5:23:08be taken inside a group suppose I have
  7766. 5:23:10this set of points which you see
  7767. 5:23:12basically over here red core and all and
  7768. 5:23:14there is also a border Point by making
  7769. 5:23:18multiple circles over here here you can
  7770. 5:23:20definitely say that how we are defining
  7771. 5:23:22core points and the Border points and
  7772. 5:23:24this can be combined into a single group
  7773. 5:23:27okay this can be combined into a single
  7774. 5:23:29group because how the connection is now
  7775. 5:23:31see this this yellow line is basically
  7776. 5:23:33created by one sorry this yellow point
  7777. 5:23:35is basically created by one Epsilon and
  7778. 5:23:37we have one One Core point over here
  7779. 5:23:40remember over here it should be at least
  7780. 5:23:43one core Point okay not one point but
  7781. 5:23:47one core point at least if it is having
  7782. 5:23:50one core point then it will become a
  7783. 5:23:52border point this will become a border
  7784. 5:23:54point that basically means yes this can
  7785. 5:23:56be the part of this specific group so
  7786. 5:23:59what we are doing Whenever there is a
  7787. 5:24:00noise we are going to neglect it
  7788. 5:24:02wherever there is a broader and core
  7789. 5:24:03points we are going to combine it so
  7790. 5:24:05I'll show you one more diagram which is
  7791. 5:24:06an amazing diagram which will help you
  7792. 5:24:09understand more in this a k means
  7793. 5:24:10clustering and Hier mean clustering now
  7794. 5:24:12see this everybody now the right hand
  7795. 5:24:15side of diagram that you see is based on
  7796. 5:24:19DB scan clustering and the left hand
  7797. 5:24:21side is basically your traditional
  7798. 5:24:23clustering method let's say that this is
  7799. 5:24:25K means which one do you think is better
  7800. 5:24:28over here you see this these all
  7801. 5:24:30outliers are not combined inside a group
  7802. 5:24:34But whichever are nearer as a core point
  7803. 5:24:37and the broader point separate separate
  7804. 5:24:38groups are actually
  7805. 5:24:40created right so this is how amazing a
  7806. 5:24:44DB scan clustering is a DB scan
  7807. 5:24:47clustering is pretty much amazing that
  7808. 5:24:50is basically the outcome of this here in
  7809. 5:24:53C's clustering you can see this all
  7810. 5:24:54these points has also been taken as blue
  7811. 5:24:57color as one group because I'll be
  7812. 5:24:58considering this as one group but here
  7813. 5:25:00we are able to determine this in a
  7814. 5:25:03amazing groups so in I'm saying you guys
  7815. 5:25:06directly use DB scan with without
  7816. 5:25:08worrying about anything so now let's
  7817. 5:25:10focus on the Practical part uh I'm just
  7818. 5:25:12going to give you a GitHub link
  7819. 5:25:14everybody download the code guys I've
  7820. 5:25:16given you the GitHub link quickly
  7821. 5:25:18download and keep your file ready I'm
  7822. 5:25:20going to open my anaconda prompt
  7823. 5:25:23probably open my jupyter notbook we'll
  7824. 5:25:25do one practical problem I've given you
  7825. 5:25:28the link guys please open it so this is
  7826. 5:25:30what we are going to do today this will
  7827. 5:25:32be amazing here you'll be able to see
  7828. 5:25:34amazing things how do you come to know
  7829. 5:25:37that over fitting or underfitting is
  7830. 5:25:39happening you don't know the real value
  7831. 5:25:41right so in in clustering there will not
  7832. 5:25:43be any underfitting or overfitting so uh
  7833. 5:25:46what all things we'll be importing first
  7834. 5:25:48is that we'll try cin clustering we'll
  7835. 5:25:50do silot scoring and then probably we'll
  7836. 5:25:52see the output and um and we'll do DB
  7837. 5:25:56scan Also let's say DB scan is also
  7838. 5:25:58there so uh what are the things we have
  7839. 5:26:00basically imported one is the cin
  7840. 5:26:03clustering one is the Sout samples and
  7841. 5:26:05Sout scores these all are present in the
  7842. 5:26:08SK learn and it is present in metrics
  7843. 5:26:11that basically means we use this
  7844. 5:26:12specific parameter to validate
  7845. 5:26:15clustering models okay now we'll try to
  7846. 5:26:18execute this and apart from that mat
  7847. 5:26:20plot lib we are just trying to import
  7848. 5:26:22numai we are trying to import and all
  7849. 5:26:24here we are executing it perfectly the
  7850. 5:26:26next thing is that here the next step is
  7851. 5:26:29that generating the sample data from
  7852. 5:26:31make underscore blobs first of all we
  7853. 5:26:33are just trying to generate some samples
  7854. 5:26:35with some two features and we are saying
  7855. 5:26:37that okay should have four centroids or
  7856. 5:26:39C centroids itself with some features
  7857. 5:26:43I'm trying to generate some X and Y data
  7858. 5:26:45randomly and this particular data set
  7859. 5:26:47will basically be used in performing
  7860. 5:26:50clustering algorithms okay forget about
  7861. 5:26:52range undor ncore clusters because we
  7862. 5:26:54need to try with different different
  7863. 5:26:55clusters and try to find out the solid
  7864. 5:26:57score so right now I just initialized
  7865. 5:26:59with 2 3 4 5 6 values it is very simple
  7866. 5:27:02so if I go and probably see my X data so
  7867. 5:27:05my X data will look something like this
  7868. 5:27:07so this is my X data with two features
  7869. 5:27:09and this is my Y data with one feature
  7870. 5:27:12which is my output which belongs to a
  7871. 5:27:13specific class okay so that you can
  7872. 5:27:16actually do with the help of make
  7873. 5:27:17underscore blobs let's say how to apply
  7874. 5:27:21kin's clustering algorithm so as I said
  7875. 5:27:23that I will be using W CSS W CSS
  7876. 5:27:26basically means within cluster sum of
  7877. 5:27:28square so I'm going to import K means
  7878. 5:27:30over here for I in range 1A 11 that
  7879. 5:27:33basically means I'm going to use
  7880. 5:27:35different different K values or centroid
  7881. 5:27:37values and try to C which is having the
  7882. 5:27:39minimal wcss value and I'll try to draw
  7883. 5:27:42that graph which I had actually shown
  7884. 5:27:44you with respect to Elbow method so here
  7885. 5:27:47I will basically be also using K means
  7886. 5:27:50number of clusters will be I and
  7887. 5:27:52initialization technique I will will be
  7888. 5:27:54using K means Plus+ so that the points
  7889. 5:27:57the centroids that are initialized those
  7890. 5:27:59those points are very very far and then
  7891. 5:28:01you have random state is equal to zero
  7892. 5:28:03then we do fit and finally we do wcss do
  7893. 5:28:06upend cins doin inertia okay this dot
  7894. 5:28:10inertia will give you the distance
  7895. 5:28:13between the centroids and all the other
  7896. 5:28:16points and this is what I'm going to
  7897. 5:28:18append in this wcss value and finally
  7898. 5:28:20I'll just plot it now here you can see
  7899. 5:28:22that I'm just plotting it obviously by
  7900. 5:28:25seeing this graph this graph looks like
  7901. 5:28:27an elbow okay this graph looks like an
  7902. 5:28:29elbow so the point that I'm actually
  7903. 5:28:31going to consider over here see which is
  7904. 5:28:34the last abrupt change so if I talk
  7905. 5:28:36about the last abrupt change here I have
  7906. 5:28:38the specific value with respect to this
  7907. 5:28:41okay I have one specific value with
  7908. 5:28:43respect to this this is my abrupt change
  7909. 5:28:45from here the changes are normal so I'm
  7910. 5:28:48going to basically select K is equal to
  7911. 5:28:504 now what I'm actually going to do with
  7912. 5:28:52the help of sart with the help of s CL
  7913. 5:28:57score we are going to compare whether K
  7914. 5:29:00is equal to 4 is valid or not so that is
  7915. 5:29:03what we are going to do valid or not so
  7916. 5:29:06here we are going to do this now let's
  7917. 5:29:09go ahead and let's try to see it how we
  7918. 5:29:11are going to do it so here you can see n
  7919. 5:29:13clusters is equal to 4 then I'm actually
  7920. 5:29:15able to find out the prediction and this
  7921. 5:29:17is specifically my output okay this is
  7922. 5:29:19done now see this code okay this code is
  7923. 5:29:23a huge code I have actually taken this
  7924. 5:29:24code directly from the SK learn page of
  7925. 5:29:28Silo if you go and see this this code is
  7926. 5:29:30directly given over there but I'm just
  7927. 5:29:33going to talk about like what are the
  7928. 5:29:35important things we need to see over
  7929. 5:29:37here with respect to different different
  7930. 5:29:39clusters see see this clusters 2 3 4 5 6
  7931. 5:29:43I'm going to basically compare whether
  7932. 5:29:46the K value should be four or not with
  7933. 5:29:48the help of solid scoring so let's go
  7934. 5:29:51here and here you can see that I'm
  7935. 5:29:54applying this one first I will go with
  7936. 5:29:56respect to for Loop for ncore clusters
  7937. 5:29:59in range underscore clusters different
  7938. 5:30:00different cluster values are there first
  7939. 5:30:02we'll start with two so here you can see
  7940. 5:30:04initialize the cluster with and cluster
  7941. 5:30:06value and a random generator seed of 10
  7942. 5:30:09for reproducibility so ncore clusters
  7943. 5:30:12first I take took it as two and then I
  7944. 5:30:14did fit predict on X after I did fit
  7945. 5:30:17predictor on X I'm using this score on X
  7946. 5:30:21comma cluster label now what this is
  7947. 5:30:23going to do understand in Solo what did
  7948. 5:30:25we discuss it will it will try to find
  7949. 5:30:27out all the Clusters the Clusters over
  7950. 5:30:30here like this and it'll try to
  7951. 5:30:32calculate the distance between them
  7952. 5:30:34which is the a of I then it'll try to
  7953. 5:30:36compute the B of I then finally it'll
  7954. 5:30:39try to compute the score and if the
  7955. 5:30:41value is between minus1 to +1 the more
  7956. 5:30:43the Valu is towards + one the more
  7957. 5:30:45better it is right so these all things
  7958. 5:30:47we have already discussed and that is
  7959. 5:30:49what this specific function will do and
  7960. 5:30:51this will give my solo average value
  7961. 5:30:53over here solid value will be over here
  7962. 5:30:55okay this we have done and then we can
  7963. 5:30:58continuously do it for another another
  7964. 5:31:00things you can actually find it over
  7965. 5:31:02here and this value that you see this
  7966. 5:31:05code that you see is nothing nothing so
  7967. 5:31:08complex okay this is just to display the
  7968. 5:31:11data properly in the form of graphs okay
  7969. 5:31:15in the form of graphs so again I'm
  7970. 5:31:17telling you I did not write this code
  7971. 5:31:18I've directly taken it from the uh SK
  7972. 5:31:22learn page of solid okay so just try to
  7973. 5:31:25see this particular uh plotting diagrams
  7974. 5:31:27and all that you can definitely figure
  7975. 5:31:29out but let's see I will try to execute
  7976. 5:31:31it and try to find out the output now
  7977. 5:31:33see for ncore cluster is equal to 2 the
  7978. 5:31:37average solid score is 70 I told you the
  7979. 5:31:40value will be between -1 to +1 and I'm
  7980. 5:31:43actually getting 704 which is very very
  7981. 5:31:46good and then for ncore cluster is equal
  7982. 5:31:48to 3 588 then ncore cluster is equal to
  7983. 5:31:524 I'm getting 65 which is pretty much
  7984. 5:31:54amazing and then for ncore cluster equal
  7985. 5:31:57to 5 the average score is 563 and ncore
  7986. 5:32:00cluster is equal to 6 you are saying
  7987. 5:32:02.45 here directly you can actually say
  7988. 5:32:05that fine for _ cluster equal to 2 I'm
  7989. 5:32:08getting an amazing score of
  7990. 5:32:10704 obviously you're you're getting the
  7991. 5:32:12highest value over this so should we
  7992. 5:32:14select ncore cluster isal to two Okay we
  7993. 5:32:17should not directly conclude from it
  7994. 5:32:19because here we need to also see that
  7995. 5:32:21any feature value or any cluster value
  7996. 5:32:24is also coming as negative value that
  7997. 5:32:26also we need to check so here we will go
  7998. 5:32:28down over here you will see the first
  7999. 5:32:30one over here with respect to the first
  8000. 5:32:32one you see that I'm get getting the
  8001. 5:32:35value from 0 to 1 it is not going going
  8002. 5:32:38to Min -.1 so definitely two clusters
  8003. 5:32:40was able to solve the problem so I'll
  8004. 5:32:43keep it like this with me I definitely
  8005. 5:32:45have a chance that this may this may
  8006. 5:32:48perform well I may have a chance that
  8007. 5:32:50this K uh K is equal to 2 May perform
  8008. 5:32:53well okay so I may have a chance let's
  8009. 5:32:55see to the next one to the next one over
  8010. 5:32:57here you can see that for one of the
  8011. 5:32:59cluster the value is negative if the
  8012. 5:33:01value is negative that basically means
  8013. 5:33:03the AI is obviously greater than b ofi
  8014. 5:33:06so I'm not going to prer this because it
  8015. 5:33:08is having some negative values even
  8016. 5:33:10though my cluster looks better but again
  8017. 5:33:13understand what is the problem with
  8018. 5:33:14respect to this cluster is that if I
  8019. 5:33:17take this cluster and probably compute
  8020. 5:33:19the distance between this point to this
  8021. 5:33:20point and if I probably compute from
  8022. 5:33:22this point to this point or this point
  8023. 5:33:24to this point this point is obviously
  8024. 5:33:26nearer to this right it is obviously
  8025. 5:33:29nearer to this so that is the reason why
  8026. 5:33:31I'm getting a negative value over here
  8027. 5:33:33okay negative value over here this is my
  8028. 5:33:36uh output my score this point that you
  8029. 5:33:40see dotted points this is my score 58
  8030. 5:33:43what whatever it is this is basically my
  8031. 5:33:45score so obviously this basically
  8032. 5:33:46indicates that this point is near the
  8033. 5:33:48other cluster point is nearer to this so
  8034. 5:33:50I'm actually getting a negative value
  8035. 5:33:52right so this you really need to
  8036. 5:33:54understand okay now similarly if I go
  8037. 5:33:56with respect to ncore Cluster is equal
  8038. 5:33:58to 4 this looks good because here I
  8039. 5:34:00don't have any negative value and here
  8040. 5:34:03you can see how cooly it has basically
  8041. 5:34:06divided the points amazing inly with the
  8042. 5:34:08help of k equal to 4 right and similarly
  8043. 5:34:11if I go with five obviously you can see
  8044. 5:34:13some negative values are here some
  8045. 5:34:15dotted line negative value are there
  8046. 5:34:17with respect to six you also have some
  8047. 5:34:18negative values so definitely I'll not
  8048. 5:34:21go with six I may either go with four or
  8049. 5:34:23I may either go with two now whenever
  8050. 5:34:26you have this options always take a
  8051. 5:34:27bigger number instead of two take four
  8052. 5:34:30because four is greater than two because
  8053. 5:34:32it will be able to create a generalized
  8054. 5:34:34model so from this I'm actually going to
  8055. 5:34:37take and is equal to 4 K is equal to 4
  8056. 5:34:39now should we compare with this with the
  8057. 5:34:41elbow method here also I got four right
  8058. 5:34:44so both are actually matching so this
  8059. 5:34:47indicates that with the help of this
  8060. 5:34:49clustering this siluette score we can
  8061. 5:34:51definitely come to a conclusion and
  8062. 5:34:53validate our clustering model in an
  8063. 5:34:55amazing way so I hope everybody is able
  8064. 5:34:57to understand and this way you basically
  8065. 5:35:00validate a model and definitely you can
  8066. 5:35:03try it out you can understand this code
  8067. 5:35:04definitely I but till here you have
  8068. 5:35:06understood that here I'm going to get
  8069. 5:35:08the average value then for iore clusters
  8070. 5:35:12whatever cluster this is matching it is
  8071. 5:35:14just mapping over there and it is
  8072. 5:35:16basically giving so this was the session
  8073. 5:35:19and uh yes in today's session we
  8074. 5:35:22efficiently covered many topics we
  8075. 5:35:24covered kin hierle clustering solid
  8076. 5:35:27score DB clustering in tomorrow's
  8077. 5:35:29session the topics that are probably
  8078. 5:35:31pending is first I'll start with svm and
  8079. 5:35:34svr second I will go ahead with XG boost
  8080. 5:35:37and and third I will cover up PCA let's
  8081. 5:35:40see whether I'll be able to complete
  8082. 5:35:41this session uh one one amazing thing
  8083. 5:35:45that I want to teach you guys because
  8084. 5:35:46many people ask me the definition of
  8085. 5:35:48bias and variance so guys uh many people
  8086. 5:35:52get confused when we talk about bias and
  8087. 5:35:55variance you know because let's say that
  8088. 5:35:57uh I have a model for the training data
  8089. 5:35:59set it gives us somewhere around 90%
  8090. 5:36:02accuracy let's say I'm getting a 90%
  8091. 5:36:04accuracy for the test data I may
  8092. 5:36:07probably getting somewhere around 70%
  8093. 5:36:10accuracy now tell me which scenario is
  8094. 5:36:12basically this most of the people will
  8095. 5:36:14be saying that okay fine it is
  8096. 5:36:16overfitting now when I say overfitting I
  8097. 5:36:19basically mention overfitting by low
  8098. 5:36:23bias and high
  8099. 5:36:25variance right so many people get
  8100. 5:36:28confused Krish tell me just the exact
  8101. 5:36:30definition of bias and variance low bias
  8102. 5:36:33obviously you are saying that because
  8103. 5:36:34the training is performed like the model
  8104. 5:36:37is performing well with the help of
  8105. 5:36:39training data set but with respect to
  8106. 5:36:41the test data set the model is not
  8107. 5:36:43performing well with respect to training
  8108. 5:36:45data set why do we always say bias and
  8109. 5:36:48with respect to test data set why do we
  8110. 5:36:50always say variance so for this you need
  8111. 5:36:52to understand the definition of bias so
  8112. 5:36:54let me write down the definition of bias
  8113. 5:36:56over here so here I can definitely write
  8114. 5:36:59that bias it is a
  8115. 5:37:02phenomena that
  8116. 5:37:05skews the
  8117. 5:37:08result of an
  8118. 5:37:13algorithm in
  8119. 5:37:15favor in favor or against an
  8120. 5:37:20idea against an idea I'll make you
  8121. 5:37:23understand the definition uh um but
  8122. 5:37:27understand the understand understand
  8123. 5:37:28what I have actually written over here
  8124. 5:37:30it is a phenomena that skewes the result
  8125. 5:37:32of an algorithm in favor or against an
  8126. 5:37:34idea whenever I say this specific idea
  8127. 5:37:37this idea I will just talk about the
  8128. 5:37:39training data set initially now when we
  8129. 5:37:42train a specific model suppose if I have
  8130. 5:37:44this specific model over
  8131. 5:37:46here and I'm training with this specific
  8132. 5:37:49training data set so this is my training
  8133. 5:37:52data set now based on the definition
  8134. 5:37:54what does it basically say it is a
  8135. 5:37:55phenomenon that skews the result of an
  8136. 5:37:57algorithm in favor or against an idea or
  8137. 5:38:00a this specific training data set so
  8138. 5:38:03even though I'm training this particular
  8139. 5:38:04model with this training data set
  8140. 5:38:07with this data set it may it may be in
  8141. 5:38:11favor of that or it may be against of
  8142. 5:38:12that that basically means it may perform
  8143. 5:38:14well it may not perform well if it is
  8144. 5:38:15not performing well that basically means
  8145. 5:38:17the accuracy is down if the accuracy is
  8146. 5:38:19better at that point of time what will
  8147. 5:38:21say see if the accuracy is better that
  8148. 5:38:23time what we'll say we we'll come up
  8149. 5:38:25with two terms from here obviously you
  8150. 5:38:27understand okay there are two scenarios
  8151. 5:38:28of bias now here if it is in favor that
  8152. 5:38:32basically means it is performing well
  8153. 5:38:33with respect to the training data set I
  8154. 5:38:35will basically say that it has high bu
  8155. 5:38:38if it is not able to perform well with
  8156. 5:38:40the training data set then here I will
  8157. 5:38:42say it as low
  8158. 5:38:44bias I hope everybody is able to
  8159. 5:38:46understand in this specific thing
  8160. 5:38:47because many many many people has this
  8161. 5:38:49kind kind of confusion now similarly if
  8162. 5:38:51I talk about variance let's say about
  8163. 5:38:53variance because you need to understand
  8164. 5:38:55the definition a definition is very much
  8165. 5:38:59important okay if I if I just talk about
  8166. 5:39:01the definition of variance I'm just
  8167. 5:39:03going to refer like this the variance
  8168. 5:39:07refers to the changes in the model when
  8169. 5:39:13using when using different
  8170. 5:39:16portion of the
  8171. 5:39:20training or test
  8172. 5:39:23data now let's understand this
  8173. 5:39:25particular
  8174. 5:39:27definition variance refers to the
  8175. 5:39:29changes in the model when using
  8176. 5:39:31different proportion of the test
  8177. 5:39:32training data or test data we obviously
  8178. 5:39:34know that whenever initially if I have a
  8179. 5:39:38model understand from the definition
  8180. 5:39:39everything will make sense I am
  8181. 5:39:41basically training initially with the
  8182. 5:39:43training
  8183. 5:39:44data okay because we divide our data set
  8184. 5:39:47see our data set whenever we are working
  8185. 5:39:49with we divide this into two parts one
  8186. 5:39:52is our train data and test data okay
  8187. 5:39:56because this is a tra test data is a
  8188. 5:39:58part of that particular data set right
  8189. 5:40:00and suppose in this particular training
  8190. 5:40:02data it gets trained and performs well
  8191. 5:40:04here I'm actually talking about bias but
  8192. 5:40:07when we come with respect to the
  8193. 5:40:09prediction of the specific model at that
  8194. 5:40:12point of time I can use other training
  8195. 5:40:14data that basically means that training
  8196. 5:40:15data may not be similar or I can also
  8197. 5:40:18use test data now in this test data what
  8198. 5:40:21we do we do some kind of predictions
  8199. 5:40:23these are my predictions and in this
  8200. 5:40:25prediction again I may get two
  8201. 5:40:27scenario I may get two scenario which is
  8202. 5:40:30basically mentioned by variance it
  8203. 5:40:32refers to the changes in the model when
  8204. 5:40:34using when using different portion of
  8205. 5:40:37the training or test data refers to the
  8206. 5:40:40changes basically means whether it is
  8207. 5:40:42able to give a good prediction or wrong
  8208. 5:40:44predictions that's it so in this
  8209. 5:40:46particular scenario if it gives a good
  8210. 5:40:48prediction I may definitely say it as
  8211. 5:40:50low variance that basically means the
  8212. 5:40:53accuracy with the accuracy with respect
  8213. 5:40:55to the test data is also very good if I
  8214. 5:40:58probably get a bad if I probably get a
  8215. 5:41:01bad accuracy at that time I basically
  8216. 5:41:04say it as high variance so if I talk
  8217. 5:41:06about three scenarios over here let's
  8218. 5:41:08say this is my model one and this is my
  8219. 5:41:11model
  8220. 5:41:12two and this is my model
  8221. 5:41:15three now in this scenario let's
  8222. 5:41:18consider that my model one has the
  8223. 5:41:21training
  8224. 5:41:23accuracy of 90% and test accuracy of
  8225. 5:41:3075% similarly I have here as my train
  8226. 5:41:33accuracy of 60% and my test accuracy
  8227. 5:41:38of
  8228. 5:41:4055% now similarly if I have my train
  8229. 5:41:44accuracy of 90% And my test accuracy of
  8230. 5:41:4992% now tell me what what things you
  8231. 5:41:51will be getting here obviously you can
  8232. 5:41:54directly say that fine your training
  8233. 5:41:56accuracy is better now you're talking
  8234. 5:41:58about bias so this basically indicates
  8235. 5:42:00that this has low
  8236. 5:42:02bias and since your test accuracy is bad
  8237. 5:42:07because it is when compared to the train
  8238. 5:42:08accuracy it is less so here you are
  8239. 5:42:10basically going to say high
  8240. 5:42:14variance understand with respect to the
  8241. 5:42:16definition similarly over here what
  8242. 5:42:18you'll say high
  8243. 5:42:20bias High variance because obviously it
  8244. 5:42:22is not performing
  8245. 5:42:25well this is another scenario last the
  8246. 5:42:28last scenario is that this is the
  8247. 5:42:30scenario that we want because it is low
  8248. 5:42:33bias and low variance
  8249. 5:42:37okay many many people have basically
  8250. 5:42:39asked me the definition with respect to
  8251. 5:42:41bias and variance and here I've actually
  8252. 5:42:43discussed and this indicates this gives
  8253. 5:42:45me a generalized model and this is what
  8254. 5:42:50is our aim when we are working as a data
  8255. 5:42:53scientist so I hope you have understood
  8256. 5:42:55the basic difference between V bias and
  8257. 5:42:57variance and I was able to give you lot
  8258. 5:43:00of examples lot of understanding with
  8259. 5:43:02respect to this so I hope you have
  8260. 5:43:05actually got this particular uh
  8261. 5:43:08understanding of this uh two terms which
  8262. 5:43:10we specifically talk about high bias low
  8263. 5:43:12bias High variance low variance right so
  8264. 5:43:16this was it from my side guys uh and uh
  8265. 5:43:19I hope you have understood
  8266. 5:43:22this
  8267. 5:43:29okay so let's take let's consider a data
  8268. 5:43:34set credit
  8269. 5:43:37and let's say this is a
  8270. 5:43:39approval so we are going to take this
  8271. 5:43:42sample data set and understand how does
  8272. 5:43:43XG boost work suppose salary is less
  8273. 5:43:47than or equal to 50 and the credit is
  8274. 5:43:50bad so approval the loan approval will
  8275. 5:43:52be zero that basically means he he or
  8276. 5:43:54she will not get if it is less than or
  8277. 5:43:56equal to 50 if the credit score is good
  8278. 5:44:00then probably approval will be one if it
  8279. 5:44:02is less than or equal to 50 if it is
  8280. 5:44:06good
  8281. 5:44:07again then it is going to get one if it
  8282. 5:44:10is greater than
  8283. 5:44:1250 and if it is bad then obviously
  8284. 5:44:16approval will be
  8285. 5:44:19zero if it is greater than
  8286. 5:44:2250 if it is good we are going to get it
  8287. 5:44:25as one if it is greater than
  8288. 5:44:2950k and probably if it is normal then
  8289. 5:44:33also we are going to get
  8290. 5:44:35it so this is this is my data set so how
  8291. 5:44:38does XG boost classifier work understand
  8292. 5:44:41the full form of XG boost is
  8293. 5:44:44Extreme gradient
  8294. 5:44:47boosting extreme gradient boosting so we
  8295. 5:44:50will basically understand about extreme
  8296. 5:44:52gradient boosting now extreme gradient
  8297. 5:44:55boosting uh will be actually used to
  8298. 5:44:58solve both classification and the
  8299. 5:45:00regression problem statement so first of
  8300. 5:45:02all let's understand how it is basically
  8301. 5:45:05exib basically how it actually if you if
  8302. 5:45:08you just talk about XG boost you
  8303. 5:45:10understand that it is a boosting
  8304. 5:45:11technique and internally it tries to use
  8305. 5:45:13decision tree so how does this decision
  8306. 5:45:16Tre is basically getting constructed in
  8307. 5:45:18the case of XV boost and how it is
  8308. 5:45:20basically solved we are going to discuss
  8309. 5:45:21about it so whenever we start exib boost
  8310. 5:45:24classifier understand that first of all
  8311. 5:45:26we create a specific base model suppose
  8312. 5:45:29if I say this is my base model and this
  8313. 5:45:32base model will be a weak learner okay
  8314. 5:45:36and this base model will always give an
  8315. 5:45:38output of probability of 0.5 in the case
  8316. 5:45:42of classification problem so suppose if
  8317. 5:45:45I say this is probability 0.5 then I
  8318. 5:45:48will try to create a field over here
  8319. 5:45:50this field is called as residual field
  8320. 5:45:53so first base model what I'm going to do
  8321. 5:45:55any data set that you give from here to
  8322. 5:45:57train it will always give you the output
  8323. 5:45:59as 0.5 so this is just a dummy base
  8324. 5:46:02model now tell me if my probability
  8325. 5:46:06output is is 0.5 if I want to calculate
  8326. 5:46:08the residual that basically means I need
  8327. 5:46:10to subtract approval minus this
  8328. 5:46:12particular value so what will be the
  8329. 5:46:16value over here 0 -.5 will be
  8330. 5:46:20-.5 1 -.5 will be5 1 -.5 will
  8331. 5:46:25be5 and 0 -.5 will be -.5 and this 1 -.5
  8332. 5:46:31will
  8333. 5:46:32be uh 0.5 and this will also be 0.5
  8334. 5:46:36let's consider that I have one more
  8335. 5:46:37record uh and this specific record can
  8336. 5:46:40be anything uh because I want to keep
  8337. 5:46:43some more records over here so let's
  8338. 5:46:45consider that I have one more record
  8339. 5:46:46which is less than or equal to 50K and
  8340. 5:46:49if the credit scod is normal you're
  8341. 5:46:51going to get zero so here also if I try
  8342. 5:46:53to find out the residual it will be
  8343. 5:46:55minus5 now the first step I hope
  8344. 5:46:58everybody's understood we have to create
  8345. 5:46:59a base model okay this base model is
  8346. 5:47:01very much important because we have to
  8347. 5:47:04create all the decision Tree in a
  8348. 5:47:06sequential manner so the first
  8349. 5:47:09sequential base tree which is again this
  8350. 5:47:11is also a decision tree kind of thing
  8351. 5:47:12you can consider but this is a base
  8352. 5:47:15model which takes any inputs and gives
  8353. 5:47:17by default the probability as 05 now
  8354. 5:47:20let's go ahead and understand what are
  8355. 5:47:22the steps in constructing decision tree
  8356. 5:47:24after creating the base model the first
  8357. 5:47:27step is that create uh binary decision
  8358. 5:47:31tree so I'm going to write it down all
  8359. 5:47:34the steps please make sure that you note
  8360. 5:47:35it down so so create a binary tree
  8361. 5:47:39binary decision tree using the features
  8362. 5:47:43second step we basically Define we we we
  8363. 5:47:47say it as okay Second Step what we do we
  8364. 5:47:50actually calculate the similarity weight
  8365. 5:47:54we calculate the similarity weight I'll
  8366. 5:47:57talk about this similarity weight what
  8367. 5:47:59exactly it is if I want to use this a
  8368. 5:48:02formula it is summation of residual
  8369. 5:48:05Square
  8370. 5:48:07divided
  8371. 5:48:08by summation of probability 1 minus
  8372. 5:48:13probability plus Lambda I'll talk about
  8373. 5:48:16this what is exactly Lambda it is the
  8374. 5:48:18kind of hyperparameter again so that it
  8375. 5:48:20does not overfit the third thing is that
  8376. 5:48:23we calculate the Information Gain okay
  8377. 5:48:26Information Gain so these are the steps
  8378. 5:48:28we basically use in constructing or in
  8379. 5:48:32solving uh in creating an HD boost
  8380. 5:48:34classifier the first step is that we
  8381. 5:48:36create a inary decision tree using the
  8382. 5:48:38feature then we go ahead with
  8383. 5:48:40calculating the similarity weight and
  8384. 5:48:42finally we go ahead and calculate the
  8385. 5:48:43information gain so how does it go ahead
  8386. 5:48:46let's understand over here and let's try
  8387. 5:48:47to find out okay now let's go ahead and
  8388. 5:48:50let's try to construct the decision tree
  8389. 5:48:53as I said that let's consider that I'm
  8390. 5:48:55considering salary feature So based on
  8391. 5:48:58using salary feature what I'm actually
  8392. 5:48:59going to do I am going to take this as
  8393. 5:49:02my node and I'm going to split this up
  8394. 5:49:05and remember whenever we are creating
  8395. 5:49:07decision Tree in this particular case it
  8396. 5:49:09will be a binary decision tree let's say
  8397. 5:49:13that in salary one is less than or equal
  8398. 5:49:15to one is greater than 50 so this two
  8399. 5:49:18you obviously have in the case of binary
  8400. 5:49:20in case of credit where there are three
  8401. 5:49:22categories I'll also show you how that
  8402. 5:49:25further split will happen and how that
  8403. 5:49:27will get converted into a binary team so
  8404. 5:49:29here you have less than or equal to 50K
  8405. 5:49:32and greater than 50k now let's go ahead
  8406. 5:49:35and understand how many vales are there
  8407. 5:49:37in this salary so if I see before the
  8408. 5:49:40split you can definitely see that I'm
  8409. 5:49:42going to use this residual and probably
  8410. 5:49:45train this entire model now if I really
  8411. 5:49:48wanted to find out the residual
  8412. 5:49:49initially these are my residuals over
  8413. 5:49:51here so one resid is -.5 then I have 0.5
  8414. 5:49:56over here then I have .5 then again I
  8415. 5:49:59have -.5 then again I have 0.5 then
  8416. 5:50:03again I have 0.5 and finally I have
  8417. 5:50:06minus .5 so these are my total residuals
  8418. 5:50:09that are there suppose if I make this
  8419. 5:50:11split less than or equal to 50 First
  8420. 5:50:14less than or equal to 50 the residuals
  8421. 5:50:16what are things are there so here I'm
  8422. 5:50:18going to have minus5 then less than or
  8423. 5:50:21equal to 50 again I'm going to have 05
  8424. 5:50:23then again less than or equal to 50 I'm
  8425. 5:50:25going to have 0.5 and less than or equal
  8426. 5:50:27to again one more 0.5 is there I'm just
  8427. 5:50:30going to remove this the last5 which is
  8428. 5:50:33nothing but Min -.5 so I hope you
  8429. 5:50:35understood this split so half of the
  8430. 5:50:37things came over here the remaining half
  8431. 5:50:40will be greater than or equal to greater
  8432. 5:50:41than 50 so you have one value here one
  8433. 5:50:44value here one value here so it will be
  8434. 5:50:46Min -.5 then you have 0.5 and then
  8435. 5:50:50finally you have 0.5 residuals how do we
  8436. 5:50:53get it guys see from the base model
  8437. 5:50:55which is by default giving 0.5 first my
  8438. 5:50:58data goes over here by default
  8439. 5:51:00probability I'm going to get 0.5 so
  8440. 5:51:02residual is basically calculated from
  8441. 5:51:04this probability and approval so this
  8442. 5:51:07probability minus approval so if you
  8443. 5:51:09subtract 0 -.5 sorry I'm just going to
  8444. 5:51:12rub this so if you subtract 0 -.5 you're
  8445. 5:51:16going to get -.5 1 -.5 you're going to
  8446. 5:51:19get .5 1 -.5 you're going to get .5 so
  8447. 5:51:22everybody I hope is very much clear with
  8448. 5:51:24respect to this so this is the first
  8449. 5:51:26step we constructed a binary tree now in
  8450. 5:51:28the second step it says calculate the
  8451. 5:51:30similarity weight now how to calculate
  8452. 5:51:33the similarity weight similarity weight
  8453. 5:51:35formula is sum of residual Square now
  8454. 5:51:37what is residual Square let's say that
  8455. 5:51:39I'm going to calculate the the the uh
  8456. 5:51:43I'm going to calculate for this okay
  8457. 5:51:45similarity weight now in this particular
  8458. 5:51:47case if I go and calculate my similarity
  8459. 5:51:49weight it will be summation of residual
  8460. 5:51:52Square this is my residual values this
  8461. 5:51:55is my residual Valu so I'm going to do
  8462. 5:51:57the summation of this Square okay this
  8463. 5:52:01value square you can see over here sum
  8464. 5:52:03of residual Square everybody you can see
  8465. 5:52:06sum of of residual squares so what do
  8466. 5:52:08you think sum of residual squares will
  8467. 5:52:09be in this particular case how I have to
  8468. 5:52:12do it I will just take up this all
  8469. 5:52:14values like
  8470. 5:52:16-.5
  8471. 5:52:17+5
  8472. 5:52:20+5 and
  8473. 5:52:22-.5 whole square right I'm just going to
  8474. 5:52:24do the squaring of this divided by
  8475. 5:52:27understand what it is divided by it is
  8476. 5:52:29divided by probability of 1 minus
  8477. 5:52:31probability now where do we get this
  8478. 5:52:33probability value where do we get this
  8479. 5:52:35probability value value we get this
  8480. 5:52:37probability value from our base model
  8481. 5:52:40right so here I'm basically going to say
  8482. 5:52:42that we are going to do the summation of
  8483. 5:52:44probability of 1 minus probability 1
  8484. 5:52:47minus probability that basically means
  8485. 5:52:50for each and every point for each and
  8486. 5:52:52every Point what is the probability see
  8487. 5:52:54probability is basically coming from the
  8488. 5:52:56base model so for each Pro each point
  8489. 5:52:59I'm going to come compute two things one
  8490. 5:53:01is the probability and then 1 minus
  8491. 5:53:04probability and this I'm going to do the
  8492. 5:53:06summ
  8493. 5:53:07like this I will do it four times 1 -.5
  8494. 5:53:10then .5 * 1 -.5 and finally you'll be
  8495. 5:53:15able to see one more will be there which
  8496. 5:53:17is
  8497. 5:53:18+5 1 -.5 so this will be your total
  8498. 5:53:21things with respect to this so I hope
  8499. 5:53:24you have understood till here uh where
  8500. 5:53:26you are able to understand that what we
  8501. 5:53:28have done this is summation of uh
  8502. 5:53:31residual square and this is the
  8503. 5:53:33remaining probability multiplied by 1
  8504. 5:53:35minus probability now tell me what are
  8505. 5:53:39you able to find out from this if you
  8506. 5:53:41cancel this and this this and this this
  8507. 5:53:44value is going to become zero so this
  8508. 5:53:47entire value is going to become Zer
  8509. 5:53:48because 0 divided by anything is 0er so
  8510. 5:53:51here I hope everybody is understood what
  8511. 5:53:53is the similarity weight of this
  8512. 5:53:55specific node if I want to write it is
  8513. 5:53:57nothing but zero now you may be
  8514. 5:53:59considering where is Lambda
  8515. 5:54:01value okay we will initially initialize
  8516. 5:54:04Lambda by 1 I'll talk about this hyper
  8517. 5:54:05parameter let's consider it as 1 so here
  8518. 5:54:09+ 1 or plus 0 let's let's consider
  8519. 5:54:12Lambda value 0 let's say for right now
  8520. 5:54:14okay I'm just going to make it Lambda is
  8521. 5:54:16equal to0 I'm just going to talk about
  8522. 5:54:19it because it is a kind of hyper
  8523. 5:54:21parameter by Z -.5 -.5 +5 +5 if I do the
  8524. 5:54:28summation if I do the summation here you
  8525. 5:54:31will be able to see that I'm going to
  8526. 5:54:32get zero so this calculation we have
  8527. 5:54:34done and we have got uh the sumission of
  8528. 5:54:36weight is equal to Z and let's go ahead
  8529. 5:54:39and calculate the sumission of the
  8530. 5:54:40weight of the next node no no no it's
  8531. 5:54:43not first Square it is whole squar so
  8532. 5:54:46here also if I do so it is5 +5 now let's
  8533. 5:54:51do it for this if I want to find out the
  8534. 5:54:53similarity weight again see I'm going to
  8535. 5:54:55repeat it .5 +5 whole squ and since
  8536. 5:55:00there are three points so I'm going to
  8537. 5:55:01basically use probability 1 minus
  8538. 5:55:04probability for one point then plus
  8539. 5:55:08probability 1 minus probability second
  8540. 5:55:11point and then probability and 1 minus
  8541. 5:55:14probability for the third point and
  8542. 5:55:16Lambda is zero so I'm not going to write
  8543. 5:55:18anything now go let's go and do the
  8544. 5:55:20calculation for this node so - 5 - 5 it
  8545. 5:55:24becomes zero then .5 whole square right
  8546. 5:55:27so here I'm going to get 0.25 here if
  8547. 5:55:30you do the calculation here you are
  8548. 5:55:31going to get 75 so this value is going
  8549. 5:55:34to be 1x3 and which is nothing at33 so
  8550. 5:55:37the similarity weight for this node for
  8551. 5:55:40this node
  8552. 5:55:42is33 so here you can see probability of
  8553. 5:55:45multiplied by 1 minus
  8554. 5:55:47probability okay now the next step that
  8555. 5:55:50we do is that calculate the information
  8556. 5:55:53gain now you know how to calculate the
  8557. 5:55:55information gain but before that let's
  8558. 5:55:57do the computation for this also for
  8559. 5:55:59this root node also go ahead and
  8560. 5:56:01calculate the similarity weight of
  8561. 5:56:04this okay they
  8562. 5:56:06why the base model probability is5
  8563. 5:56:09because it is just understand that it is
  8564. 5:56:11a dummy dummy model I have just put a if
  8565. 5:56:14condition there saying that it is going
  8566. 5:56:15to give 0.5 now do it for this one guys
  8567. 5:56:17root node what it will be see I can
  8568. 5:56:20calculate from here only minus1 gone
  8569. 5:56:23this is also gone this is also gone this
  8570. 5:56:25will be .25 divided by something now
  8571. 5:56:29tell me guys what should be for the root
  8572. 5:56:32node what is the similarity similarity
  8573. 5:56:34weight what is the similarity weight for
  8574. 5:56:36for this do this calculation everyone up
  8575. 5:56:39one I know it will be. 25 divided by
  8576. 5:56:44this will be 1.75 are you getting this
  8577. 5:56:48similarity weight which will be nothing
  8578. 5:56:50but 1 by 7 and if I divide 1 by 7 if I
  8579. 5:56:54say what is 1 by 7 it
  8580. 5:56:57is42 so it is nothing but .14 if I want
  8581. 5:57:00to calculate the root node similarity
  8582. 5:57:02weight over here
  8583. 5:57:05is4 so I know 0.14 here 0 here 33 now
  8584. 5:57:09see over here we calculate the
  8585. 5:57:11Information Gain Next Step the third
  8586. 5:57:13step what we do is that we calculate the
  8587. 5:57:15information gain now Information Gain is
  8588. 5:57:19nothing but in this particular case the
  8589. 5:57:21root node similarity weight we'll try to
  8590. 5:57:24add up so I will be getting
  8591. 5:57:270.33 minus this particular Top Root node
  8592. 5:57:31whatever split has happened that
  8593. 5:57:33similarity weight I'll take 0 +33
  8594. 5:57:36-14 so Point
  8595. 5:57:39-14 and if I do it it is nothing but
  8596. 5:57:42just open your calculator again and
  8597. 5:57:4633
  8598. 5:57:48-14 so it is nothing but .19 I'm getting
  8599. 5:57:52.19 as my information gain the
  8600. 5:57:56information gain of this specific tree I
  8601. 5:57:59got it
  8602. 5:58:00as19 obviously you know how the features
  8603. 5:58:03will get selected based on the
  8604. 5:58:06Information Gain but let's say that the
  8605. 5:58:08highest Information Gain that is given
  8606. 5:58:10by salary okay now we will go ahead and
  8607. 5:58:13do the further split let's go ahead and
  8608. 5:58:16do the further split so I I know my
  8609. 5:58:18information gain now it is1 n and
  8610. 5:58:20Information Gain is basically used to
  8611. 5:58:23select that specific node through which
  8612. 5:58:26the split will happen now I'll further
  8613. 5:58:27go and do the split let's say that I'm
  8614. 5:58:29going to do the further split with the
  8615. 5:58:31next feature that is which one credit so
  8616. 5:58:33I'm going to take credit over here I'm
  8617. 5:58:36going to take credit over here and again
  8618. 5:58:39I have to do a binary split again but
  8619. 5:58:42you may be considering chish here are
  8620. 5:58:43only three categories how we are going
  8621. 5:58:45to basically do this particular split
  8622. 5:58:48right because we don't know how to do
  8623. 5:58:50the split because we have three
  8624. 5:58:51categories over here so in this case
  8625. 5:58:53what I will do is that we what we can
  8626. 5:58:56definitely do is that in this particular
  8627. 5:58:58case the split that we are probably
  8628. 5:59:00going to do is that let's consider two
  8629. 5:59:02categories like good and normal at one
  8630. 5:59:04side bad at one side so here it becomes
  8631. 5:59:06a binary split again now let's go ahead
  8632. 5:59:09and let's try to see that how many data
  8633. 5:59:11points will fall here and how many data
  8634. 5:59:12points will fall here so for writing
  8635. 5:59:14down the data points let's say if it is
  8636. 5:59:17less than or see go to the path if it is
  8637. 5:59:19less than or equal to 50 it'll go this
  8638. 5:59:21path and if it is B then we are probably
  8639. 5:59:24going to get how much is the residual we
  8640. 5:59:26are going to get one residual over here
  8641. 5:59:28first of all so this is my one residual
  8642. 5:59:31that is -.5 then similarly if I see less
  8643. 5:59:34than or equal to 50 good is there right
  8644. 5:59:37good or normal is there so here again 0.
  8645. 5:59:39five will come I hope everybody is able
  8646. 5:59:42to understand see the second record less
  8647. 5:59:44than or equal to 50 we go in this path
  8648. 5:59:45but it is good we come over here again
  8649. 5:59:48less than or equal to 50 good again we
  8650. 5:59:50are going to get 1
  8651. 5:59:51more5 then go with respect to greater
  8652. 5:59:55than or equal to 50 which is coming over
  8653. 5:59:57here we'll not worry about it right now
  8654. 5:59:59again less than or equal to 50 normal
  8655. 6:00:01again it is
  8656. 6:00:03-.5 right so this many records
  8657. 6:00:06definitely coming over here only one
  8658. 6:00:08record is basically coming over here
  8659. 6:00:10then again we will start the same
  8660. 6:00:12process again we will start the same
  8661. 6:00:14process now for the same process what we
  8662. 6:00:16are going to do again try to calculate
  8663. 6:00:18the similarity weight now in order to
  8664. 6:00:20calculate the similarity weight what I
  8665. 6:00:22will do I will basically say this is my
  8666. 6:00:24similarity weight this will become .25
  8667. 6:00:28divided 025 why because this whole
  8668. 6:00:31square right this whole Square residual
  8669. 6:00:33square right summation of residual
  8670. 6:00:36square but here I have only one residual
  8671. 6:00:38so this Square it will become and then
  8672. 6:00:40what I'm actually going to do I'm going
  8673. 6:00:41to basically write .5 - 1 -.5 this is
  8674. 6:00:45nothing for only for one data point so
  8675. 6:00:47this is nothing but .5 * .5 which is
  8676. 6:00:50nothing but 0.25 right now in this
  8677. 6:00:53particular case I will get similarity
  8678. 6:00:54weight as I hope everybody I'm getting
  8679. 6:00:56it as one now what about this similarity
  8680. 6:00:58weight if you want to compute it is
  8681. 6:01:00again very very simple this and this
  8682. 6:01:02will get cancelled then again it will be
  8683. 6:01:03025 divided by um if I say one like this
  8684. 6:01:08.25 then again it will be 75 then this
  8685. 6:01:11will also be 1 by3 that is nothing but
  8686. 6:01:1333 so similarity weight will
  8687. 6:01:16be33 then again I have to calculate the
  8688. 6:01:19information gain of this node what I
  8689. 6:01:21will do I will add this up see 1
  8690. 6:01:24+33 I'll add like 1
  8691. 6:01:27+33 minus 0 why zero because the
  8692. 6:01:30information gain the similarity weight
  8693. 6:01:32of this uh the up one is basically 0
  8694. 6:01:37right for this particular credit node
  8695. 6:01:39similarity weight is zero so 1
  8696. 6:01:41+33 minus 0 this will be 1.33 so like
  8697. 6:01:45this further split will again happen
  8698. 6:01:47over here with different different node
  8699. 6:01:49and we will only be getting a binary
  8700. 6:01:51split but we will be comparing based on
  8701. 6:01:54Information Gain which one is coming
  8702. 6:01:55good now let's say that I have created
  8703. 6:01:57this path I have I have designed I have
  8704. 6:02:00developed my entire binary decision tree
  8705. 6:02:02which is a speciality in XG boost now
  8706. 6:02:06what I'm going to do over here is that
  8707. 6:02:08see everybody what I'm going to do let's
  8708. 6:02:10consider the inferencing part let's say
  8709. 6:02:12this record is going to go how we are
  8710. 6:02:15going to calculate the output so this
  8711. 6:02:17first of all went to this base model now
  8712. 6:02:21let's go ahead and see how the
  8713. 6:02:22inferencing will happen suppose This
  8714. 6:02:24Record is going right so first of all
  8715. 6:02:26this record will go to this base model
  8716. 6:02:29the base model is giving the probability
  8717. 6:02:30as 0.5 so the first base model is
  8718. 6:02:34basically giving 0.5 now base based on
  8719. 6:02:36this 05 how do we calculate the real
  8720. 6:02:39probability how do we calculate the real
  8721. 6:02:41probability in this okay so we apply
  8722. 6:02:43something called as logs so we basically
  8723. 6:02:45say log of P / 1us P so this is the
  8724. 6:02:49formula we basically apply in only the
  8725. 6:02:52case of base model so if we try to see
  8726. 6:02:55this it is nothing but log
  8727. 6:02:57of5 / .5 which is nothing but zero log
  8728. 6:03:01of one is nothing but zero so in the
  8729. 6:03:03first case whenever any record goes I
  8730. 6:03:05will be getting the zero value over here
  8731. 6:03:08okay zero value over here then plus why
  8732. 6:03:11plus I'm doing because it will now go to
  8733. 6:03:13the binary decision tree now this record
  8734. 6:03:15will go to my binary decision Tre
  8735. 6:03:17whatever value I'm getting from this I'm
  8736. 6:03:19actually adding that up and now it will
  8737. 6:03:21go over here now when it goes over here
  8738. 6:03:24first of all let's see which branch it
  8739. 6:03:25is following it is following less than
  8740. 6:03:27or equal to 50 Branch first Branch over
  8741. 6:03:29here then this is bad it'll go and
  8742. 6:03:32follow here so here I can see that the
  8743. 6:03:34similarity weight is one now the
  8744. 6:03:36similarity weight is basically one in
  8745. 6:03:38this case so what we do in the case of
  8746. 6:03:40this we pass it to a learning rate
  8747. 6:03:44parameter so this specifically is my
  8748. 6:03:46learning rate multiplied by 1 one
  8749. 6:03:49because why similarity weight is one
  8750. 6:03:51over here so this will basically be my
  8751. 6:03:54first references and Alpha over here is
  8752. 6:03:57my learning rate it can be a very small
  8753. 6:03:59value based on the learning parameter
  8754. 6:04:01that we use like how we have defined
  8755. 6:04:04learning parameters elsewhere on top of
  8756. 6:04:06this we apply an activation function
  8757. 6:04:09which is called as sigmoid since this is
  8758. 6:04:11a classification problem we apply an
  8759. 6:04:14activation function which is called as
  8760. 6:04:15sigmoid and I hope you know what is the
  8761. 6:04:17use of sigmoid based on this based on
  8762. 6:04:20the alpha value based on this the output
  8763. 6:04:22will be between 0 to 1 now I hope you
  8764. 6:04:25getting it guys this is how the entire
  8765. 6:04:27inferencing will probably happen now
  8766. 6:04:30similarly what I will do I will try to
  8767. 6:04:32construct this kind of decision tree
  8768. 6:04:33parall so we we can also write our
  8769. 6:04:37entire function will look something like
  8770. 6:04:40this Alpha 0 + alpha 1 and this will be
  8771. 6:04:46your decision tree 1 output then Alpha 2
  8772. 6:04:50your decision tree output Alpha 3 your
  8773. 6:04:53decision 3 output like this Alpha 4 your
  8774. 6:04:57decision 3 output fourth decision tree
  8775. 6:05:00like this it will be alpha n your
  8776. 6:05:02decision tree n output and this will be
  8777. 6:05:06your output finally when you're trying
  8778. 6:05:09to inference from any new
  8779. 6:05:12record now the reason why we say this as
  8780. 6:05:15boosting because see understand we are
  8781. 6:05:17going to add each and every decision
  8782. 6:05:19tree output slowly to finally get our
  8783. 6:05:22output with respect to the working of
  8784. 6:05:23the decision tree this is how XG boost
  8785. 6:05:26actually work don't credit further needs
  8786. 6:05:28to be simplified yes see like this
  8787. 6:05:31similarly we can split credit with the
  8788. 6:05:33help of like we can make blue green one
  8789. 6:05:35side normal at one side But whichever
  8790. 6:05:37will be giving the information gain more
  8791. 6:05:40that will be taken into consideration
  8792. 6:05:41right and this is how your entire X
  8793. 6:05:43boost classifier works it is very very
  8794. 6:05:46difficult to basically calculate all
  8795. 6:05:48those things so that is the reason we
  8796. 6:05:50say that XG boost is also a blackbox
  8797. 6:05:53model so this is basically a blackb
  8798. 6:05:56model it is it prone to overfitting see
  8799. 6:05:59at one stage we also need to perform
  8800. 6:06:02hyperparameter tuning and this we
  8801. 6:06:05specifically say pre- pruning we tend to
  8802. 6:06:08do pre pruning and since we are
  8803. 6:06:10combining multiple decision trees no no
  8804. 6:06:14this decision tree this decision tree is
  8805. 6:06:17this one this independent decision tree
  8806. 6:06:19which I have created now parall after
  8807. 6:06:21this what I'll do I'll create one more
  8808. 6:06:22decision tree so it'll be looking like
  8809. 6:06:24this see finally how it will look so
  8810. 6:06:26this is my base model then my data then
  8811. 6:06:29my data will go to this decision tree
  8812. 6:06:31which I have actually done as a binary
  8813. 6:06:33split on different different records
  8814. 6:06:36then again we will make another decision
  8815. 6:06:38tree which will again be a binary tree
  8816. 6:06:40the splits will look like this then this
  8817. 6:06:43is my base model where I'm getting the
  8818. 6:06:45value as zero this will be alpha 1
  8819. 6:06:47multiplied by decision tree 1 which is
  8820. 6:06:50this then this is Alpha 2 multiplied by
  8821. 6:06:53decision tree 2 which is this and like
  8822. 6:06:55this we will keep on continuously adding
  8823. 6:06:58more decision trees unless and until
  8824. 6:07:00this entire things becomes a very strong
  8825. 6:07:04learner so this is how how we basically
  8826. 6:07:06do the combination of all these things
  8827. 6:07:08so I hope everybody is able to
  8828. 6:07:10understand about the XG boost classifier
  8829. 6:07:14now you may be thinking how does
  8830. 6:07:15regressor work do you want a regressor
  8831. 6:07:17problem statement also the decision tree
  8832. 6:07:19will get constructed based on
  8833. 6:07:21Independent features and again Lambda
  8834. 6:07:23value is a hyperparameter we basically
  8835. 6:07:26set up Lambda value with the help of
  8836. 6:07:28cross validation now uh let's go ahead
  8837. 6:07:30and discuss about ex boost regressor the
  8838. 6:07:33second algorithm that we we will
  8839. 6:07:35probably discuss about is something
  8840. 6:07:37called as XG boost regressor and how
  8841. 6:07:41does X boost regressor actually work
  8842. 6:07:43some fundamental is follow in random
  8843. 6:07:45Forest no in random Forest it is
  8844. 6:07:47completely different there bagging
  8845. 6:07:49happens bagging happens so over here
  8846. 6:07:52let's go ahead with the regressor so
  8847. 6:07:54here I'm going to take some example
  8848. 6:07:56let's say that I have this many
  8849. 6:07:57experience this many Gap and based on
  8850. 6:08:00that we need to determine the salary my
  8851. 6:08:02salary is my output feature let's say
  8852. 6:08:04the experience is 2 2.5 3 4 4.5 okay now
  8853. 6:08:10in this Gap let's say it is yes
  8854. 6:08:13yes no no yes and let's say that the
  8855. 6:08:17salary is somewhere around 40K it is
  8856. 6:08:2041k
  8857. 6:08:2252k and uh let's see some more data set
  8858. 6:08:25over here 60k and 62k now the first step
  8859. 6:08:29in classifier we created a base model
  8860. 6:08:32here also we'll try to create a base
  8861. 6:08:33model first of all this base model what
  8862. 6:08:36output it will give it will give the
  8863. 6:08:38average of all these values what is the
  8864. 6:08:40average of all these values okay what is
  8865. 6:08:42the average of all these value 40 81 52
  8866. 6:08:4560 62 if I just do the average it is
  8867. 6:08:48nothing but 51k so by default I will
  8868. 6:08:50create a base model which will take any
  8869. 6:08:52input and just give the output as 51
  8870. 6:08:54this is the first step now based on this
  8871. 6:08:56I will try to calculate my residual now
  8872. 6:08:58how do I calculate my residual I will
  8873. 6:09:00just subtract 40 by 51k so this will
  8874. 6:09:03basically be - 11k
  8875. 6:09:06and uh this will be 10 K - 10 K - 10 and
  8876. 6:09:11this will be 1 this will be 9 and this
  8877. 6:09:16will be 11 I hope everybody's able to
  8878. 6:09:18get this let's say that I I make this as
  8879. 6:09:2142k okay for just making my calculation
  8880. 6:09:23little bit easy so I have 9 over here so
  8881. 6:09:26this is my residual then again the first
  8882. 6:09:28step is that I construct my uh decision
  8883. 6:09:32tree now let's say say that I'm going to
  8884. 6:09:35use The Experience over here so this is
  8885. 6:09:37my experience node and based on this
  8886. 6:09:39experience node I have my features over
  8887. 6:09:42here so here I will take up all my
  8888. 6:09:44residuals - 11 99 1 99 11 and then how
  8889. 6:09:50do I do the split based on experience
  8890. 6:09:52this is a continuous feature so I have
  8891. 6:09:56to basically do split with respect to
  8892. 6:09:58continuous feature which I have already
  8893. 6:09:59shown you in decision tree how do we do
  8894. 6:10:01so here is my residual here it is 40
  8895. 6:10:04minus this
  8896. 6:10:05is - 11 K - 9 K uh this is 1 K this is 9
  8897. 6:10:12K and
  8898. 6:10:1411k - 9k so now I will just create take
  8899. 6:10:17up my first node here I'm going to use
  8900. 6:10:20my experience feature I know my values
  8901. 6:10:23what all things are going to come 11k in
  8902. 6:10:25the root node - 9 1 9 and 11 now what we
  8903. 6:10:30are going to do over here is that so I'm
  8904. 6:10:32going to do again a binary split over
  8905. 6:10:34here now the binary split will happen
  8906. 6:10:36based on the continuous feature that is
  8907. 6:10:38experienced so two types of Records I
  8908. 6:10:40may get one is less than or equal to two
  8909. 6:10:42and one is greater than 2 less than or
  8910. 6:10:46equal to two and one is greater than two
  8911. 6:10:48now less than or equal to two when I do
  8912. 6:10:49the split let's see how many values we
  8913. 6:10:51are getting less than or equal to two I
  8914. 6:10:53will get only one value that is -1 and
  8915. 6:10:56here I'm actually going to get all the
  8916. 6:10:58other values - 9 1 9 11 now what we are
  8917. 6:11:02going to do after this is that calculate
  8918. 6:11:04the similarity weight now here the
  8919. 6:11:06similarity weight will little bit the
  8920. 6:11:08formula will change with respect to
  8921. 6:11:10regression so similarity weight is
  8922. 6:11:12nothing but summation of residual
  8923. 6:11:15squares divided by number of residuals
  8924. 6:11:18plus Lambda again here we are going to
  8925. 6:11:20consider Lambda is zero because this is
  8926. 6:11:22a hyper parameter tuning more the value
  8927. 6:11:25of Lambda that basically means more more
  8928. 6:11:27we are penalizing with respect to the
  8929. 6:11:29residuals so this will be the formula
  8930. 6:11:31that we are going to apply okay so let's
  8931. 6:11:33see for the first number that that we
  8932. 6:11:35want to apply so how this will get
  8933. 6:11:37applied again I'm going to write this
  8934. 6:11:39formula here it'll be better let's say
  8935. 6:11:42here similarity weight is equal to
  8936. 6:11:46summation of residual square and here
  8937. 6:11:49you have number of residuals plus Lambda
  8938. 6:11:52see previously we were using probability
  8939. 6:11:54and then all those things we are using
  8940. 6:11:56so if you want to calculate the
  8941. 6:11:58similarity weight of this this will
  8942. 6:11:59become 121 divided by number of residual
  8943. 6:12:03is 1 plus Lambda is 0 so this is going
  8944. 6:12:06to be 121 so here we are going to
  8945. 6:12:08calculate the similarity weight which is
  8946. 6:12:10nothing but 121 if if we probably take
  8947. 6:12:13Alpha let's let's do one thing if we
  8948. 6:12:15probably take uh if if we probably take
  8949. 6:12:19Alpha is equal to 1 then what will
  8950. 6:12:20happen if you take Alpha is equal to 1
  8951. 6:12:22just think over here what will what may
  8952. 6:12:23happen we may directly penalize the
  8953. 6:12:26similarity weight right by just adding
  8954. 6:12:28one okay so let's do that also suppose I
  8955. 6:12:30say I'm going to take Alpha is equal to
  8956. 6:12:321 so what will happen this will not be
  8957. 6:12:35the formula now now what will become 121
  8958. 6:12:38divided number of residual is 1 + 1 this
  8959. 6:12:41is nothing but 65.5 let's say that I now
  8960. 6:12:44have 65.5 as my similarity weight now
  8961. 6:12:47similarly I will go ahead and compute
  8962. 6:12:49the similarity weight for the next one
  8963. 6:12:52so here it will become - 9 + 9 + 9 + 11
  8964. 6:12:58whole Square divided 4 + 1 so this and
  8965. 6:13:01this will get subtracted 12 squ is
  8966. 6:13:04nothing but 14 4 144 divid 5 so if I go
  8967. 6:13:07ahead and calculate 144 ID 5 it is
  8968. 6:13:10nothing but 28.5 so here I get
  8969. 6:13:1528.5 so the similarity weight for this
  8970. 6:13:18is
  8971. 6:13:2028.5 similarly I can go ahead and
  8972. 6:13:22calculate the similarity weight for this
  8973. 6:13:24for the top one so it'll be nothing but
  8974. 6:13:27what it will be 11 + sorry - 11 - 11 - 9
  8975. 6:13:34+ + 1 + 9 + 11 divided 1 2 3 4 5 5 + 1
  8976. 6:13:41is 6 so this is getting subtracted this
  8977. 6:13:44will be 1X 6 anyhow this will be whole
  8978. 6:13:46square right so anyhow it will be 1X 6
  8979. 6:13:48only so 1X 6 will be my similarity
  8980. 6:13:51weight over here okay 28.8 hits okay now
  8981. 6:13:54finally The Information Gain that we
  8982. 6:13:56need to compute will be very much simple
  8983. 6:13:58what will be the Information Gain 65.5 +
  8984. 6:14:0328.8
  8985. 6:14:06minus 1X 6 so try to get it whatever we
  8986. 6:14:09are trying to get it over here just tell
  8987. 6:14:11me what will be the output is it 98.34%
  8988. 6:14:3560.5 60.5 + 28 88 then this will change
  8989. 6:14:40just a second 89.1 3 understand you
  8990. 6:14:44don't have to worry about calculation
  8991. 6:14:46automatically that things will be doing
  8992. 6:14:48it okay so you don't have to worry now
  8993. 6:14:50see we have now further the decision
  8994. 6:14:52tree can be splitted into any number of
  8995. 6:14:54times probably the next split what we
  8996. 6:14:56can do is that we can we can do next
  8997. 6:14:58split something like this this will be
  8998. 6:15:00my experience the two splits that may
  8999. 6:15:03happen with respect to less than or
  9000. 6:15:05equal to 2.5 less than or equal to 2.5
  9001. 6:15:08or greater than 2.5 now if this probably
  9002. 6:15:11gives the Information Gain better then
  9003. 6:15:13the split will happen like this
  9004. 6:15:14otherwise whichever gives the better
  9005. 6:15:16information again the split will
  9006. 6:15:17basically happen like this I hope like
  9007. 6:15:20let's say that this is this is the split
  9008. 6:15:22that is required - 11 - 11 is 9 is over
  9009. 6:15:25here and then we have 1 comma 9A 11 okay
  9010. 6:15:28because less than or equal to 2.5 this
  9011. 6:15:30two records will definitely go over here
  9012. 6:15:32and this two This Record will definitely
  9013. 6:15:34go over here now if I try to calculate
  9014. 6:15:36the similarity weight for this it will
  9015. 6:15:38be nothing but - 11 - 9 - 11 - 9 whole S
  9016. 6:15:43ided 2 + 1 right now in this particular
  9017. 6:15:46case it will be - 20 s / 3 which is
  9018. 6:15:51nothing but 400 2 20 into 20 is 400
  9019. 6:15:55which is nothing but 3 so if I go and
  9020. 6:15:57probably use a
  9021. 6:15:59calculator and show it to you
  9022. 6:16:02400 / 3 which is nothing but
  9023. 6:16:06133.33 so the similarity weight for this
  9024. 6:16:08is
  9025. 6:16:10133.33 similarly I can go ahead and
  9026. 6:16:12compute for this it will be 1 + 9 + 11
  9027. 6:16:15whole s / 3 + 1 right so it will be 10 +
  9028. 6:16:1911 10 + 11 is nothing but 21 whole s/ 4
  9029. 6:16:24so what it is 21 whole square if I open
  9030. 6:16:27my calculator 21 s 21 * 21 which is
  9031. 6:16:33nothing but 441 divid by 4 divid by 4 so
  9032. 6:16:37this will probably 110 110.
  9033. 6:16:412.25 and similarly I can go ahead and
  9034. 6:16:44compute for this so if I want to compute
  9035. 6:16:46for this what it will be the same thing
  9036. 6:16:49that we have got over here that is 1x 6
  9037. 6:16:51so this will basically be 1X 6 so
  9038. 6:16:53finally if I compute the information
  9039. 6:16:55again it will be what it will be 133
  9040. 6:17:011333 +
  9041. 6:17:031.25 - 1X 6 obviously this value will be
  9042. 6:17:06greater than the previous one what we
  9043. 6:17:08have got that is
  9044. 6:17:108913 so definitely we are going to use
  9045. 6:17:12this split which is better than the
  9046. 6:17:14previous split right let's say that this
  9047. 6:17:17split has been considered finally how do
  9048. 6:17:20we see the output okay I hope everybody
  9049. 6:17:23is able to understand right let's say
  9050. 6:17:24that this split has worked well so I'm
  9051. 6:17:26going to rub all these things
  9052. 6:17:2911.25 is there now suppose I want to do
  9053. 6:17:33the inferencing how the inferencing will
  9054. 6:17:35be done
  9055. 6:17:3711.25 here 110.2 now suppose any record
  9056. 6:17:41comes from here first of all any record
  9057. 6:17:43that will go it will go to the base
  9058. 6:17:45model so the base model whenever it goes
  9059. 6:17:47the value is 51 51 plus alpha 1 this is
  9060. 6:17:51my learning rate one suppose if it goes
  9061. 6:17:54in this route then what we have we have
  9062. 6:17:56- 11 - 9 whenever we go in this rote
  9063. 6:17:59which has - 11 and - 9 the average of
  9064. 6:18:02both these numbers will be considered
  9065. 6:18:03what is average of both these numbers -
  9066. 6:18:0511 - 1 9/ 2 this is nothing but - 10
  9067. 6:18:10right so - 10 will get multiplied here
  9068. 6:18:13suppose if it goes in this route then
  9069. 6:18:15here what will happen here will 1 + 9 +
  9070. 6:18:1811 divide by 3 average will be taken so
  9071. 6:18:2021 divid 3 7 will be there so this will
  9072. 6:18:23get replaced by 7 so similarly anything
  9073. 6:18:27that you are doing this is with respect
  9074. 6:18:28to decision tree 1 like this we will
  9075. 6:18:30again construct decision tree separately
  9076. 6:18:33and again it will become Alpha 2 by
  9077. 6:18:35decision Tre 2 Alpha 3 by decision 3 3
  9078. 6:18:39and like this you will be doing till
  9079. 6:18:42Alpha and decision 3 n and once you
  9080. 6:18:45calculate this this will be your
  9081. 6:18:47specific output in a regression tree so
  9082. 6:18:49in this particular case what will happen
  9083. 6:18:51you're just trying to play with
  9084. 6:18:53parameters and you're trying to use in a
  9085. 6:18:55different way to compute all this things
  9086. 6:18:57everybody clear but again it is a
  9087. 6:18:59blackbox model you cannot visualize all
  9088. 6:19:02this things now let's go to the third
  9089. 6:19:03algorithm which is called as s VM see
  9090. 6:19:05svm is almost like decision uh logistic
  9091. 6:19:08regression okay so the major aim of svm
  9092. 6:19:12is
  9093. 6:19:13that major aim of svm is that suppose if
  9094. 6:19:16I have a do data points like this okay
  9095. 6:19:20we obviously use uh logistic regression
  9096. 6:19:23to split this data points right like
  9097. 6:19:25this we try to create a best fit line
  9098. 6:19:28which looks like this and probably based
  9099. 6:19:30on this best fit line we try to divide
  9100. 6:19:32the point now in svm what we do is that
  9101. 6:19:36we not only create a best fit line but
  9102. 6:19:40instead we also create a point which is
  9103. 6:19:44called as marginal
  9104. 6:19:45planes so like this we create some
  9105. 6:19:48marginal
  9106. 6:19:49plane so this is your hyper plane and
  9107. 6:19:53this is your marginal plane and
  9108. 6:19:55whichever plane has this maximum
  9109. 6:19:58distance will be able to divide the
  9110. 6:20:01points more efficiently but usually in
  9111. 6:20:05in a normal scenario you know whenever
  9112. 6:20:07we talk about hyper plane or whenever we
  9113. 6:20:10talk about marginal plane there will be
  9114. 6:20:11lot of overlapping of points right
  9115. 6:20:13suppose if I have some specific points I
  9116. 6:20:16have one point which looks like this I
  9117. 6:20:18may also have another points which may
  9118. 6:20:20overlap so it is very difficult to get
  9119. 6:20:23an exact straight marginal planes and
  9120. 6:20:26split the point based on this now this
  9121. 6:20:28specific marginal plane should be
  9122. 6:20:30maximum because we can create any type
  9123. 6:20:32best fit line and probably
  9124. 6:20:35uh use this marginal plane now if we
  9125. 6:20:38have this overlapping right if for what
  9126. 6:20:40do we call for this kind of plane this
  9127. 6:20:42kind of plane is basically called as
  9128. 6:20:44hard marginal plane so this is basically
  9129. 6:20:47called as hardge marginal plane okay and
  9130. 6:20:51similarly if any points are overlapping
  9131. 6:20:54suppose this yellow points can also get
  9132. 6:20:56overlapped over here and there may be
  9133. 6:20:58some kind of Errors so for this
  9134. 6:21:00particular case we basically say as soft
  9135. 6:21:02marginal plane because here we will be
  9136. 6:21:05able to see that errors will be there
  9137. 6:21:07now in asvm what we focus on doing is
  9138. 6:21:10that we focus on creating this marginal
  9139. 6:21:13plane with maximum distance even though
  9140. 6:21:15there are some errors we consider it in
  9141. 6:21:17solving it by providing some kind of
  9142. 6:21:19hyper parameter now how do we go ahead
  9143. 6:21:22and basically create this all marginal
  9144. 6:21:24planes and how do we go ahead with this
  9145. 6:21:26it's very much simple uh just imagine in
  9146. 6:21:29this specific way that initially let's
  9147. 6:21:32consider that I have this data point
  9148. 6:21:33suppose this is my
  9149. 6:21:35best fit line how do we give this best
  9150. 6:21:38fit line as equation we basically say
  9151. 6:21:40yal mx + C right we we basically say
  9152. 6:21:43this equation as y mx + C no hard hard
  9153. 6:21:47marginal it is impossible in a normal
  9154. 6:21:50data set obviously you'll not be able to
  9155. 6:21:52get it but definitely we go ahead with
  9156. 6:21:55creating a soft marginal plan now Y is
  9157. 6:21:56equal to MX plus C what does this m
  9158. 6:21:59indicate m is nothing but slope and C
  9159. 6:22:02indicates nothing but intercept
  9160. 6:22:05can I say that this both equations are
  9161. 6:22:07same ax + b y + C isal 0 can I also say
  9162. 6:22:12that this is the equation of a straight
  9163. 6:22:14line can I say that this is also the
  9164. 6:22:16equation of straight line I will say
  9165. 6:22:18that both of them are equal can I say
  9166. 6:22:20both of them are equal see if I try to
  9167. 6:22:22prove this to you if I take this
  9168. 6:22:24equation and try to find out y it will
  9169. 6:22:26be nothing but minus C Min - c
  9170. 6:22:30minus a sorry - a x and this will be
  9171. 6:22:34divided by B this will be divided by
  9172. 6:22:37B this will be divided by B so here you
  9173. 6:22:40can see that it is almost the same in
  9174. 6:22:42this particular case my M value will be
  9175. 6:22:44- A by B and my C will basically be
  9176. 6:22:47minus C by B so both the equation are
  9177. 6:22:49almost same
  9178. 6:22:51so let's consider that this is my
  9179. 6:22:53equation and I am actually and whenever
  9180. 6:22:57I say Y is equal to mx + C can I also
  9181. 6:23:00write something like this Y is equal to
  9182. 6:23:03W1
  9183. 6:23:05X1 + W2 X2 plus like this plus C or plus
  9184. 6:23:10b same thing no so here also we can
  9185. 6:23:13write y w transpose x + B same equation
  9186. 6:23:17right we are basically using same
  9187. 6:23:19equation yes we can also write it in a
  9188. 6:23:21different way but at the end of the day
  9189. 6:23:23we are also treating something like this
  9190. 6:23:25let's say that this slope is in this
  9191. 6:23:28direction if this slope is in this
  9192. 6:23:30direction then I can basically say that
  9193. 6:23:32let's consider that the slope is minus
  9194. 6:23:33one
  9195. 6:23:35let's say that this slope is minus one
  9196. 6:23:36see it is in the negative Direction
  9197. 6:23:38let's say that this slope is minus one
  9198. 6:23:40I'm just trying to prove that this slope
  9199. 6:23:42is negative value let's consider this
  9200. 6:23:44now suppose this is one of my point - 4a
  9201. 6:23:480 and obviously this particular equation
  9202. 6:23:50is given by this particular line is
  9203. 6:23:52given by this equation now if I really
  9204. 6:23:55want to find out the Y value let's say
  9205. 6:23:57that this is my
  9206. 6:23:59X1 this is my X1 and this is my X2 let's
  9207. 6:24:03say that
  9208. 6:24:05I want to find out I want to find out
  9209. 6:24:08this W transpose x + b the Y value based
  9210. 6:24:12on this line if I want to compute the y-
  9211. 6:24:14value based on this line how will I
  9212. 6:24:16compute W transpose X basically means
  9213. 6:24:18what w value what all things will be
  9214. 6:24:20there one value is B right B is
  9215. 6:24:23intercept right now intercept is passing
  9216. 6:24:25from origin can I say my B will be zero
  9217. 6:24:28obviously I can assume that b will be
  9218. 6:24:30zero now in this particular case if I
  9219. 6:24:32talk about w w in this case is minus one
  9220. 6:24:35which I have initialized over here so if
  9221. 6:24:37I want to do this matrix multiplication
  9222. 6:24:39it will be W transpose can be written as
  9223. 6:24:41like this and this x value can be
  9224. 6:24:44written as -4 comma - 4 and 0 -4 and 0
  9225. 6:24:49right so I can basically write like this
  9226. 6:24:52now if I do this multiplication what
  9227. 6:24:54will my value I get I will basically get
  9228. 6:24:57four right so this is a positive
  9229. 6:25:01value this is a positive value Now
  9230. 6:25:04understand since this is a positive
  9231. 6:25:05value any points that are below this
  9232. 6:25:08line any points that I consider below
  9233. 6:25:11this line and if I try to calculate the
  9234. 6:25:13Y can I say that it will always be
  9235. 6:25:15positive yes or no similarly if I could
  9236. 6:25:18probably consider one point over here as
  9237. 6:25:214A 4A 4 now tell me in this 4A 4 if I
  9238. 6:25:25calculate the Y value what will you get
  9239. 6:25:27whether you'll get a positive value or a
  9240. 6:25:29negative value if I try to calculate the
  9241. 6:25:30Y value in this case because here only
  9242. 6:25:32positive values will'll be getting right
  9243. 6:25:34so if I calculate the Y value will the Y
  9244. 6:25:37value be negative or positive just try
  9245. 6:25:39to calculate how do you calculate again
  9246. 6:25:41I will use y equation this time again my
  9247. 6:25:44slope is minus1 my intercept is zero and
  9248. 6:25:46here I will have 4 comma
  9249. 6:25:494 now here Min
  9250. 6:25:51-4 and then this is + 0 this will be Min
  9251. 6:25:54-4 right so this will be a negative
  9252. 6:25:57value negative value guys negative see -
  9253. 6:26:004 + 0 negative so any point that I will
  9254. 6:26:05probably have in top of this any
  9255. 6:26:08points Above This Plane right and if I
  9256. 6:26:12try to calculate the Y value it will
  9257. 6:26:13always be negative so what two things
  9258. 6:26:16you are able to get positive and
  9259. 6:26:17negative so you can consider this
  9260. 6:26:19entirely one category this another
  9261. 6:26:22category at least these two things you
  9262. 6:26:24can basically
  9263. 6:26:25consider guys I hope everybody is able
  9264. 6:26:27to understand this so this will be my
  9265. 6:26:29one
  9266. 6:26:30category and this will be my another
  9267. 6:26:32category obviously so that basically
  9268. 6:26:34means I can definitely use a plane and
  9269. 6:26:35split this point I hope everybody is
  9270. 6:26:37able to understand now let's go ahead
  9271. 6:26:39and let's see how this marginal plane
  9272. 6:26:41will get created and what is the cost
  9273. 6:26:44function to basically do this or what is
  9274. 6:26:46the cost function in making sure that
  9275. 6:26:48the marginal plane will definitely work
  9276. 6:26:50right it becomes difficult right so
  9277. 6:26:52suppose let's consider an
  9278. 6:26:55example suppose I say that this is my
  9279. 6:26:58lines let's say uh I want to basically
  9280. 6:27:01create a kind of I have two variety of
  9281. 6:27:03points one is this point let's say I
  9282. 6:27:06have all this points like this and the
  9283. 6:27:07other points I have somewhere here let's
  9284. 6:27:10consider I am just using directly good
  9285. 6:27:13number of points so that I can split it
  9286. 6:27:15okay because I will try to talk about it
  9287. 6:27:17what I'm actually trying to prove so
  9288. 6:27:20obviously this is my best fit line that
  9289. 6:27:21splits and apart from that what I will
  9290. 6:27:24do is that I'll also create a marginal
  9291. 6:27:26points so in order to create the
  9292. 6:27:27marginal point I may use some different
  9293. 6:27:30color let's see which color this will be
  9294. 6:27:32my one marginal point remember it will
  9295. 6:27:35be to the nearest point over here and
  9296. 6:27:38basically we will construct like like
  9297. 6:27:40this and similarly here we will be
  9298. 6:27:43constructing like this I've already told
  9299. 6:27:45you guys this equation can be mentioned
  9300. 6:27:48at w transpose x + B = 0 right I can
  9301. 6:27:51definitely say this because ax + b y + C
  9302. 6:27:55is equal to 0 so this I can also write
  9303. 6:27:57it as W transpose x equal to 0 sorry
  9304. 6:28:00plus b plus b equal to 0 so both are
  9305. 6:28:03same okay this I don't have to prove it
  9306. 6:28:05I hope everybody's clear with this now
  9307. 6:28:08what I'm going to do let's represent
  9308. 6:28:10this line also with some equation so
  9309. 6:28:12this line if I want to represent this
  9310. 6:28:14will be W transpose x + B what value
  9311. 6:28:17will come over here positive or negative
  9312. 6:28:19C from this line anything above this
  9313. 6:28:21plane right any any any distance that we
  9314. 6:28:24try to find out it will always be
  9315. 6:28:25negative so let's say that I'm using it
  9316. 6:28:27as minus one to just read as it is a
  9317. 6:28:30negative value and this line that I am
  9318. 6:28:32going to mention it it will be W
  9319. 6:28:34transpose x + B is equal to + 1 Min -1
  9320. 6:28:37above + 1 because we have already
  9321. 6:28:39discussed from this point if you're
  9322. 6:28:41trying to calculate the Y value it is
  9323. 6:28:43always going to be + one this is going
  9324. 6:28:45to be minus one here I should definitely
  9325. 6:28:48say this as K okay but I'm not
  9326. 6:28:50mentioning K in many articles you'll see
  9327. 6:28:53it as minus one uh many research paper
  9328. 6:28:55also they use it as minus one but I
  9329. 6:28:57would like to specify uh minus and plus
  9330. 6:28:59K but here let's go and write minus1 and
  9331. 6:29:02plus now my aim is to increase this
  9332. 6:29:05distance okay this distance I really
  9333. 6:29:07want to increase this distance now in
  9334. 6:29:09order to increase this if I increase
  9335. 6:29:11this distance that basically means my
  9336. 6:29:13model is performing well so let's say I
  9337. 6:29:16want to find this distance first of all
  9338. 6:29:18so if I write w transpose X Plus Bal to
  9339. 6:29:201 and here I will write w transpose x +
  9340. 6:29:23B isal minus1 so what I'm going to do
  9341. 6:29:25I'm going to do the computation and
  9342. 6:29:28subtract it like this so here obviously
  9343. 6:29:31this will be my X1 this will be my X2
  9344. 6:29:34okay because these are my another points
  9345. 6:29:35X2 and X1 so I can write w transpose X1
  9346. 6:29:40-
  9347. 6:29:42X2 B and B will get cancell and here I
  9348. 6:29:45will be writing two right so from here
  9349. 6:29:49we can definitely write two different
  9350. 6:29:50things let's see what all things we can
  9351. 6:29:52write so here this is nothing but the
  9352. 6:29:54difference between my this plane and
  9353. 6:29:56this plane which is given by like this
  9354. 6:29:58okay now always understand whenever we
  9355. 6:30:01consider any any vector vors right any
  9356. 6:30:06vectors right it also has something
  9357. 6:30:07called as
  9358. 6:30:09magnitude so if I want to remove this
  9359. 6:30:12magnitude I can divide this by W this
  9360. 6:30:16magnitude of w then only my Vector will
  9361. 6:30:18remain which is indicated like this so
  9362. 6:30:20I'm going to basically divide by this
  9363. 6:30:22particular operation both both the side
  9364. 6:30:24I'm dividing by this magnitude of w and
  9365. 6:30:27I don't care about the directions over
  9366. 6:30:29here right now we just care about the
  9367. 6:30:30vectors now when I write like this what
  9368. 6:30:33is our aim our aim is to can I say our
  9369. 6:30:36aim is to our aim is to
  9370. 6:30:40maximize 2 byw can I say this guys yes
  9371. 6:30:43or
  9372. 6:30:46no what is our aim our aim is to
  9373. 6:30:49basically maximize this right by
  9374. 6:30:52updating W comma B value I need to
  9375. 6:30:56maximize this yes everybody's clear with
  9376. 6:30:59this can I say that yes I want to
  9377. 6:31:01maximize this yes or no everybody I want
  9378. 6:31:05to maximize this if I maximize this that
  9379. 6:31:07basically means my marginal plane will
  9380. 6:31:08become bigger my marginal plane will be
  9381. 6:31:10bigger okay now can I write along with
  9382. 6:31:13this that such that y of I my output
  9383. 6:31:17will be dependent on two different
  9384. 6:31:18things one is I can say that my y y of I
  9385. 6:31:22is plus of uh is + one when w transpose
  9386. 6:31:26x + B is greater than or equal to 1
  9387. 6:31:29everybody see in this equation what I'm
  9388. 6:31:31actually trying to specify such that y
  9389. 6:31:33of I is + 1 when w transpose x + B is
  9390. 6:31:36greater than 1 and when it is minus 1
  9391. 6:31:38that basically means w transpose of X is
  9392. 6:31:40B is less than or equal to minus now
  9393. 6:31:42what does this basically mean see all my
  9394. 6:31:46values whenever I compute W transpose x
  9395. 6:31:49+ B is greater than or equal to 1 I'm
  9396. 6:31:51obviously going to get this + one when w
  9397. 6:31:54transpose X+ B is less than or equal to
  9398. 6:31:561 I'm always going to get the output as
  9399. 6:31:58minus one I hope that is the reason why
  9400. 6:32:00I have actually written like this so
  9401. 6:32:02this two we have already discussed why
  9402. 6:32:03we are specifically writing we want to
  9403. 6:32:05increase the marginal plane which is
  9404. 6:32:07this this is my marginal plane and I'm
  9405. 6:32:09writing one condition that my Yi value
  9406. 6:32:11will be+ one when w transpose X plus b
  9407. 6:32:14is greater than or equal to 1 otherwise
  9408. 6:32:16it when it is less than or equal to
  9409. 6:32:17minus one it is going to be very much
  9410. 6:32:18clear with this transpose condition we
  9411. 6:32:20have already done it everybody clear
  9412. 6:32:22with this now on top of it we can add
  9413. 6:32:25one more very important Point instead of
  9414. 6:32:28writing such that and all you can also
  9415. 6:32:30say that our major
  9416. 6:32:32aim our major aim is that if I multiply
  9417. 6:32:36y i multiplied by W transpose X of I + B
  9418. 6:32:41If I multiply this two this will always
  9419. 6:32:44be able greater than or equal to 1 for
  9420. 6:32:48correct points right for correct points
  9421. 6:32:52because understand if it is minus one if
  9422. 6:32:55I'm multiplying with this and if it is a
  9423. 6:32:57correct Point minus into minus will
  9424. 6:32:59obviously be greater than or equal to
  9425. 6:33:01one only right similarly for this it
  9426. 6:33:03will be greater than 1 so I can also
  9427. 6:33:05definitely say that my major M If I
  9428. 6:33:07multiply y of I with this it will be
  9429. 6:33:10always greater than or equal to + 1 U
  9430. 6:33:12which is definitely saying that it will
  9431. 6:33:14be a positive value so this is just a
  9432. 6:33:16representation guys but understand what
  9433. 6:33:19is the minimized cost function this is
  9434. 6:33:21my minimized cost function maximized
  9435. 6:33:23cost function now I'm going to again
  9436. 6:33:26write it down
  9437. 6:33:28maximize W comma B maximize W comma b 2
  9438. 6:33:33by magnitude of w I can also write
  9439. 6:33:37something like this minimize W comma B
  9440. 6:33:40and I can just inverse this which looks
  9441. 6:33:43like this are these both are same or not
  9442. 6:33:45because always understand in machine
  9443. 6:33:48learning algorithm why do we write
  9444. 6:33:51minimize things because we are trying to
  9445. 6:33:54minimize something okay both are
  9446. 6:33:57equivalent these both are equivalent and
  9447. 6:33:59why we specifically write minimization
  9448. 6:34:01because in the back propagation when we
  9449. 6:34:03we are continuously updating the weights
  9450. 6:34:05of w and B so we can definitely write
  9451. 6:34:08like this so here my main target is to
  9452. 6:34:12minimize this particular value by
  9453. 6:34:14changing W and B and I will start adding
  9454. 6:34:17some more parameters over here this is
  9455. 6:34:19fine till here I think everybody has got
  9456. 6:34:22it this is our aim and we are going to
  9457. 6:34:23do this but I'm going to add two more
  9458. 6:34:26parameters in this Optimizer one is C of
  9459. 6:34:29I and one is summation of I equal 1 to n
  9460. 6:34:33and here I will use something called as
  9461. 6:34:35EA EA of I first of all I'll tell what
  9462. 6:34:38is C of I see if I have this specific
  9463. 6:34:41data point let's say if some of my
  9464. 6:34:44points are over here then is it a right
  9465. 6:34:47right prediction or wrong prediction if
  9466. 6:34:49some of my points are over here is it a
  9467. 6:34:51right prediction or wrong prediction
  9468. 6:34:54obviously it is a wrong prediction if my
  9469. 6:34:56points are somewhere here is it a WR
  9470. 6:34:58prediction wrong wrong incorrect
  9471. 6:34:59prediction right so this C value
  9472. 6:35:02basically says that how many errors we
  9473. 6:35:04can have how many errors we can have if
  9474. 6:35:06it says that fine we can have six errors
  9475. 6:35:08or seven errors how many errors we can
  9476. 6:35:11have even though we are using the
  9477. 6:35:13marginal plane how many errors we can
  9478. 6:35:16have so here I'm specifically writing
  9479. 6:35:18how many errors we can have this is what
  9480. 6:35:21is specified by C ofi EA of I basically
  9481. 6:35:24says that what is the summation of I'm
  9482. 6:35:26going to write it down since we are
  9483. 6:35:28doing the sumission this entire term
  9484. 6:35:31basically mentions that sumission
  9485. 6:35:34of the distance of the values distance
  9486. 6:35:37of the wrong points and how do we
  9487. 6:35:39calculate the distance from here to here
  9488. 6:35:42suppose this is a wrong point I will try
  9489. 6:35:44to calculate the distance from here to
  9490. 6:35:45here I will do the sumission of this
  9491. 6:35:47I'll do the sumission of this I will do
  9492. 6:35:49the sumission of this similarly for the
  9493. 6:35:51Green Point another sumission will
  9494. 6:35:53happen from here to here like this here
  9495. 6:35:56to here and we going to do that specific
  9496. 6:35:57sumission so we are telling that fine if
  9497. 6:36:01you are not able to fit properly try to
  9498. 6:36:05apply this two hyperparameters and try
  9499. 6:36:07to make sure that this many errors are
  9500. 6:36:10also there it is well and good no
  9501. 6:36:11problem we will go ahead with that try
  9502. 6:36:14to do the submission of the data points
  9503. 6:36:15and based on that try to construct the
  9504. 6:36:18best fit line along with the marginal
  9505. 6:36:20plane like this even though there are
  9506. 6:36:23some errors over here or errors over
  9507. 6:36:25here we are good to go with respect one
  9508. 6:36:27more thing is there which is called as
  9509. 6:36:28Al svr svr only one thing is getting
  9510. 6:36:32changed in svr only this value will get
  9511. 6:36:36changed so I want you all to explore and
  9512. 6:36:38just let me know this will be one
  9513. 6:36:40assignment for you only this value will
  9514. 6:36:42be changing remaining everything are
  9515. 6:36:43same so just try to if you change this
  9516. 6:36:46particular value that becomes an svr
  9517. 6:36:49just try to explore and just try to find
  9518. 6:36:51out and just try to let me know so
  9519. 6:36:52overall uh did you like the entire
  9520. 6:36:55session everyone okay in this one more
  9521. 6:36:57thing is there which is called as kernel
  9522. 6:36:59Matrix svm kernel we say it as svm
  9523. 6:37:02kernel now in s VM kernel what happens
  9524. 6:37:04suppose if I have a specific data points
  9525. 6:37:06which looks like this which looks like
  9526. 6:37:08this so we obviously cannot use a
  9527. 6:37:10straight line and try to divide it so
  9528. 6:37:11what we do we convert this two Dimension
  9529. 6:37:14into three dimensions and then probably
  9530. 6:37:17we push our Point like this one point
  9531. 6:37:19will go like this and the white point
  9532. 6:37:21will go down and then we can basically
  9533. 6:37:24use a plane to split it so I uploaded a
  9534. 6:37:26video around uh around that and uh you
  9535. 6:37:29can definitely have a look onto that and
  9536. 6:37:31I have also shown you practically how to
  9537. 6:37:33do it that is the reason I've created
  9538. 6:37:35that specific video so great uh this was
  9539. 6:37:37it from my side I hope you like this
  9540. 6:37:39session so thank you everyone have a
  9541. 6:37:41great day keep on rocking keep on
  9542. 6:37:43learning and never give up

About this transcript

This page contains the full transcript of Complete Machine Learning In 6 Hours| Krish Naik by Krish Naik, generated from the public captions YouTube serves with the video. The transcript has 69,818 words across 9,542 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.